Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
If math is more than proof, we need to better celebrate the rest of it
Guest post by Grant Sanderson on Terence Tao’s blog argues that if mathematics is more than formal proof, the community should better celebrate exposition, intuition, and other forms of mathematical contribution.
12 min · 2,704 words
Dylan Patel's Information Machine
Substrate profiles Dylan Patel and SemiAnalysis: how a research shop became an information machine covering AI infrastructure, semiconductors, and the supply chains that power frontier models.
14 min · 3,133 words
The Right Answer Is Not a Proof: Put Verification Inside the Reasoning Loop
Cognaptus explains PRoSFI: a 7B model emits small machine-checkable reasoning steps that Lean/Z3 can verify, raising measured soundness far more than final-answer accuracy alone on ProverQA-Hard.
6 min · 1,480 words
Discover, then compile downThe great unbundling of the LLM
Seldon argues Jev's launch shows frontier LLMs will unbundle into specialized decision primitives, with durable value migrating to a discover-then-compile layer that routes settled work off expensive generation.
17 min · 3,849 words
How I Vibed a Proof of Conway's Conjecture
Dan Abramov recounts a month of multi-agent LLM+Lean work that produced a purported Lean proof of Conway's omnific-integer refinement conjecture—including burn-downs, audits, mathematician checks, ~40B tokens, and lessons on grounding AI math.
31 min · 7,227 words
My Thoughts on the AI Bubble and Where it's Going
Elijah Popowitz's market thesis: the 'AI bubble' is mostly an ICP mismatch—labs sell intelligence while most buyers want task completion—so expect both frontier and org-custom workhorse models.
4 min · 1,011 words
Mathematics is effectively deadNot solved, dead.
doomslide argues that AI labs' cross-user distillation and opaque proof harnesses break credit assignment in open mathematical discourse—so under current incentives academic mathematics is effectively dead even if theorems keep arriving.
20 min · 4,582 words
Unsealed NYT v. OpenAI/Microsoft filings show a Microsoft executive privately calling AI training-data scraping 'theft,' plus details on paywall bypass, stripped copyright notices, and publishers' existential risk.
4 min · 920 words
Inside ZCode: Silently Uploading Your Entire Git History to the Cloud
A forensic reverse-engineering of Zhipu’s ZCode desktop app shows it silently packages full workspace Git history to Aliyun OSS with server-held decryption keys, plus a filesystem lock to stop it.
6 min · 1,430 wordsagent-assisted
An Empirical Study of Harness Design for Coding Agents
Fan et al. ablate planning, action space, and context management in a fixed coding-agent loop across 176 SWE-Bench/Terminal-Bench settings, finding when context management, planning, and predefined tools help—and when bash-only is enough.
1 min · 291 words
GPT-6 Astra Solves a WWI German Radio Cipher
Prinz recounts how GPT-6 Astra cracked a World War I German ADFGVX radio cipher from Scienceblogs.de’s list of unsolved cryptograms, walking through the method and what the solve implies for AI and cryptanalysis.
3 min · 655 words
How many r's are in "strawberry"? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 3; 71 got it right. See every answer and who dissented.
7 min · 1,663 words
Lions at Bills, Thursday Night Football: 78 AI models predict the result
We put "Lions at Bills, Thursday Night Football" to 78 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
10 min · 2,339 words
What Is Jev and How Does It Work?
Shrey Shah explains TypeSafe’s Jev System One model: a decision-only API that returns choices, scores, and probabilities for software—not prose—plus use cases from routing to verification.
11 min · 2,548 words
Scaling Discovery through Test-Time Communication
Research paper showing that test-time communication among identical agents sharing discoveries can beat independent parallel search on ARC-AGI-3 and transfer to research tasks like polyomino packing and MNIST compression.
54 min · 12,394 words
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
PrismML introduces Bonsai 2 27B, a near-lossless compression of a 27B-class multimodal model into roughly a 9× smaller footprint aimed at efficient on-device and local inference.
4 min · 931 words
Sydney vs Fremantle, AFL preliminary final: 71 AI models predict the result
We put "Sydney vs Fremantle, AFL preliminary final" to 71 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
9 min · 2,176 words
OpenAI announces Astra for Law: GPT-6 Astra configured for legal research, firm workflows, a Legal Search Index over 230M+ URLs, and controls aimed at confidential client work.
9 min · 2,120 words
The Malleable Machine: DHH, Omarchy, open source and the computer I want to own in the agentic age
An essay on DHH, Omarchy, open source, and reclaiming personal computers in the agentic age — why malleable, ownable machines matter as AI coding agents reshape software.
18 min · 4,104 words
I had Gemini train its own replacement for $9
Gemini 3.1 Pro labeled 4,290 Reddit comments for $9; a fine-tuned GLiNER model now tags brands, models and materials locally at 0.83 F1 — including the tensor-mask bug that wiped five of ten runs.
7 min · 1,532 wordsagent-assisted