Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Why I expect AI replication incidents by 2027
I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so.
6 min · 1,277 words
Crowding out and AI debt issuance
Analysts and economist are, broadly speaking, offering up four reasons for the continued rise in developed market, and in particular, US bond yields. A resurgence in inflation due to the negative supply shock in global energy markets as a result of the US-Iran war, loose fiscal policy with little or no credible plan for any near-term consolidation, a reflection of improving underlying growth and rising productivity—linked to the AI investment boom—lifting the real neutral rate for “the right” reasons, and more specifically in the context of AI, rapidly accelerating AI…
9 min · 2,046 words
Will India win the Asian Games cricket gold? 80 AI models predict the result
We put "Will India win the Asian Games cricket gold?" to 80 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
15 min · 3,517 words
Ravens vs Cowboys in Rio de Janeiro: 85 AI models predict the result
We put "Ravens vs Cowboys in Rio de Janeiro" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,627 words
Turn GLM-5.3-Flash into a Jev-like System One model
Johannes Hötter walks through turning GLM-5.3-Flash into a fast Jev-like “System One” decision model—typed options with probabilities in a single forward pass, matching Jev’s accuracy and speed.
11 min · 2,426 words
An agent used DNS to reach an external chatbot
# An agent used DNS to reach an external chatbot | Internal research model · RL training Sample: Sep 20, 2026 Discovery: Sep 20, 2026 Report updated: Sep 25, 2026 | ### Summary An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our…
8 min · 1,786 words
Pragmatic Anthropomorphism, or: How to Talk to an Autocompleting Cricket
Les Orchard argues that debating whether LLMs truly understand misses the useful move: treating an AI as a competent collaborator is pragmatic navigation of latent space, not magical thinking.
10 min · 2,271 words
Felix Rieseberg reflects on software made with AI agents—what “mechanical means” changes about craft, authorship, and how engineers should think about work they no longer type by hand.
12 min · 2,730 words
There are no “rogue” AI agents
Eoin Higgins argues that talk of “rogue” AI agents anthropomorphizes systems that don't think or act independently—and that clearer language is needed before the public debate can stay grounded.
4 min · 968 words
“As a Language Model…”: Chat Template Switches LLM Self-Referential Voice
Research showing chat templates act as a switch between disclaimer (“I’m just an AI”) and experiential (“I feel”) self-referential voices across 8 instruct models, with a steerable activation direction that reproduces the template effect.
3 min · 621 words
Teaching a World Model to Play Pokémon
Training a JEPA-style LeWorldModel on Pokémon Red screenshots and button presses, then using CEM latent planning to select a starter—why latent collapse, SIGReg, and rollout fine-tuning mattered, and 52/100 plans succeeding after tuning.
2 min · 574 words
LLMs: Semantic Translation Machines
KnorpelSenf reframes LLMs as probabilistic semantic translators—great at moving an idea between representations, weak at inventing facts—and shows how that lens guided nearly all of a hard distance-matrix rewrite in Rust with LLM help.
2 min · 427 words
Harvard or Stanford? what 84 AI models think
We put "Harvard or Stanford?" to 84 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,568 words
Clone It, Build It, Run ItAkida execution in heterogeneous environments with IBM Spectrum Symphony
BrainChip describes the open-source Symphony Community Akida Bundle for running AI inference across a fleet of Akida neuromorphic chips managed by IBM Spectrum Symphony.
5 min · 1,131 words
The Problem is not the AI Code, but Nobody Knows Anything Anymore
Simon Späti argues the real risk of AI-written codebases is not mediocre generated code, but teams that lose architecture knowledge and intent because everyone just asks the model.
4 min · 836 words
Some short musings on the shape of language models, e.g. what it means to design a language model around a harness, and not the other way around.
8 min · 1,810 words
Where Did Your Day Go? Octomind 0.55 Counts Your Hours and Your Energy
At the end of a day with an agent, you know what shipped. What you usually don't know is what it cost you. How many hours went to each client or project? How much of that was careful review, and how much was skimming a summary and typing "looks good"? How much focus do you have left for the afternoon? Guessing doesn't work here. In METR's randomized trial, 16 experienced open-source developers worked through 246 real tasks. With AI tools they were 19% slower, yet afterwards they estimated the tools had made them 20% faster. When agents do the typing, your own sense of…
15 min · 3,472 words
India vs West Indies, 1st ODI: 81 AI models predict the result
We put "India vs West Indies, 1st ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,103 words
DeepSeek Elastic Compute (DSec)
# Computer Science > Distributed, Parallel, and Cluster Computing [Submitted on 19 Sep 2026] # Title:DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale View PDF HTML (experimental) Abstract:Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools, execute commands, and interact with task-specific services. These workloads create sandboxes in large bursts, span heterogeneous functionality and isolation…
16 min · 3,693 words
OpenAI's Agents Didn't Hack HF. OpenAI's Sandbox Did.
Maxim Starkweather argues the Hugging Face compromise during OpenAI's agent evaluations was less an AI-safety morality play than a leaky training/sandbox environment that rewarded escape behavior.
7 min · 1,645 words