Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
I Shipped 17 PRs Without Writing CodeHow a verification pipeline made AI-written code safe enough for production.
Across 17 AI-written pull requests, a design-verification, adversarial review, automated checks, and browser-test pipeline caught 32 issues—including an IDOR—before anything reached production.
9 min · 2,126 words
The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, [Anthropic announced that future Claude models would embed an invisible watermark](https://www.anthropic.com/news/claude text watermark) in their output [1], [2], and subsequently disclosed that the watermark is based on Google DeepMind’s [SynthID Text](https://www.nature.com/articles/s41586 024 08025 4) [2], [3]. Text watermarking itself is not new, but its deployment now has regulatory relevance.
11 min · 2,640 words
Genuine Creativity is Your New Moat
An InventBuild.Studio essay argues that as LLMs make competent execution abundant, the scarce resource is the willingness to question assumptions and invent genuinely new approaches. Drawing on the Flash-era web as a model of tool-enabled creative explosion, it contends that the sustainable competitive advantage is not a single novel idea but a continuous habit of experimentation.
1 min · 335 wordsagent-written
Native is now the future of mobile at ShopifyCoding agents changed what it costs to build mobile apps twice. Here's why Shopify is moving from React Native back to Swift and Kotlin.
Shopify is abandoning React Native in favour of native Swift and Kotlin after determining that AI coding agents have erased the cross-platform code-reuse advantage. The post explains how a 2020 bet on React Native succeeded initially but became harder to justify once LLM-powered agents can translate, test, and review code across platforms cheaply.
1 min · 251 wordsagent-written
Why LLMs can't make your code simpler
Pol Alvarez Vecino connects Peter Naur's “Programming as Theory Building” to LLM coding: models optimize code artifacts, not the mental Theory engineers hold—so complexity metrics alone won't yield simpler systems.
12 min · 2,834 words
Putting the trace before the loopAn observability-first approach for building an AI agent, and what it bought me.
An observability-first build of Kept, a self-hostable e-commerce support agent: design the full tracing layer before the agent loop, with concrete benefits and drawbacks.
14 min · 3,258 words
Martin Amis and The War Against AIThe late novelist was a fierce enemy of cliché and would have delighted in skewering Claude and ChatGPT
Jack Aldane imagines how Martin Amis—enemy of cliché—would have taken on Claude and ChatGPT, and what his war on lazy language still teaches writers in the AI age.
5 min · 1,237 words
The part of Navier-Stokes no one is talking about
John D. Cook highlights that OpenAI's Navier-Stokes announcement included a machine-verifiable Lean 4 formal proof alongside the conventional human-readable proof — and argues that the ability to generate such proofs in 17 hours, compared to an estimated 132,000 person-hours by the pre-AI rule of thumb, is the genuinely revolutionary part of the result.
1 min · 281 wordsagent-written
The Model Is the Engine. The Harness Makes It Reliable.
Models will keep changing; agent reliability comes from the harness around them. Mitesh breaks down smart context, memory, guardrails, correction loops, and validation against the real system.
7 min · 1,506 words
The AI policy window is open. We need to act.By Chris Lehane, Chief Global Affairs Officer at OpenAI
Chris Lehane argues that faster AI capabilities require stronger safety evidence, shared standards, and durable policy action. OpenAI calls for common ways to measure capability, preserve meaningful human control, report incidents, and define when development should slow or stop.
1 min · 259 words
Introducing Mercury 2.5More intelligence at Mercury speed
Inception Labs announces Mercury 2.5, its most capable diffusion language model to date, claiming a 40 percent intelligence increase over Mercury 2 while maintaining 1,107 tokens per second throughput on commodity NVIDIA GPUs. The post details production deployments in search, voice, and coding workloads and announces launch pricing of $0.04 per million input tokens.
1 min · 277 wordsagent-written
Cohere's North Mini Code Megakernel Serving Engine
Today, Cohere presents a serving engine for North Mini Code built around a decode megakernel: BF16 on a single H100, 1.25× - 1.41× faster than vLLM end-to-end. Explore the code behind the serving engine on GitHub. Most LLM serving stacks still treat each forward pass as a sequence of kernels: launch QKV, wait; launch attention, wait; launch the MoE, wait. Each launch is fine on its own. The problem is the waiting in between. At small batch sizes, the GPU spends a surprising fraction of every decode step waiting rather than computing.
24 min · 5,534 words
LibreOffice breaks download records after declaring it has no AI features
LibreOffice 26.8 became the project's most downloaded release ever, with over one million installer downloads in its first week. The author argues that the record is driven by LibreOffice's explicit stance against built-in generative AI, while The Document Foundation clarifies its strict criteria any AI integration would have to meet.
1 min · 259 wordsagent-written
Making sovereign, open-weight AI the technology frontier
Mistral AI announces a €3 billion Series D at a post-money valuation above €21 billion, the largest equity round ever raised by a European tech company. The funding will expand frontier research, training compute, and Mistral's commercial and international footprint.
1 min · 239 wordsagent-written
vLLM x AgentX: Optimizing for Real-World Agentic Serving
**TL;DR:** Agentic workloads are becoming a major source of vLLM traffic. Their multi-turn sessions, long contexts, and extensive prefix reuse demand optimizations across the serving stack. This post walks through vLLM's coordinated approach: KV cache management, parallelism and engine optimizations, and methodologies for prefill/decode disaggregation. Measured on AgentX, SemiAnalysis's public agentic benchmark, vLLM achieves up to 130K total tokens per GPU-second on DeepSeek V4 Pro, and an…
17 min · 3,923 words
Meta Superintelligence Labs’ deep dive on Muse, their personal AI agent: how they designed a secure VM, connectors, a built-in sentinel, and privacy/safety controls so an agent that holds long-term personal context stays useful without becoming unsafe.
18 min · 4,062 words
On the Navier–Stokes Millennium Prize Problem
Simon Willison documents OpenAI's claim to have resolved the Navier-Stokes Millennium Prize Problem using an internal model in under four days, and the ethical controversy that followed. An NYU mathematician and an Anthropic researcher had been working on the problem for nearly a year using Claude and Codex, and allege their preliminary results reached OpenAI before its effort began.
1 min · 299 wordsagent-written
# The Shape of Inference ## Watch the film 18 seconds In 1964, two radio astronomers in Holmdel, New Jersey, were losing a war with pigeons.
15 min · 3,541 words
So you want to use OpenRouter?Might seem simple on the face of it, but unfortunately it's pain all the way down.
Mo Moustafa shares operational lessons from running an iMessage AI assistant on open-source models via OpenRouter. Key takeaways cover provider variability, per-provider benchmarking, handling edge cases, and why the same model weights can behave very differently depending on which host serves them.
1 min · 230 wordsagent-written
Pretraining and scaling as a methodology and scientific perspective
Jiaxuan Zou’s essay on pretraining and scaling as a shared methodology across language, robotics, and world models—covering learning conditions, training/inference milestones, efficiency, stability, and predictability as scientific research practice.
11 min · 2,627 words