Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Gemini 3.8 text-to-speech says hello
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS—more expressive audio models for custom character voices and scene dialogue across AI Studio, the Gemini API, Enterprise, Notebook, and Vids.
6 min · 1,376 words
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Simon Willison’s first-look notes on Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna launches the same day—pricing, early impressions, and what the renewed model price war means for builders.
5 min · 1,148 words
Senior PhD student Bhavay Tyagi collects practical advice for juniors and undergrads navigating a PhD amid rapid AI change—staying abreast, choosing problems, and keeping research craft intact.
2 min · 490 words
SlopShape: Identifying AI-Generated Commercial Web Content
Research paper introducing SlopShape, a method for identifying AI-generated commercial web content from structural signals alone—motivation, method, and evaluation on web-scale data.
48 min · 11,021 words
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Better prompt caching for GPT-6
OpenAI explains GPT-6 prompt-caching improvements: higher cache hit rates, new diagnostics, explicit breakpoints, and controls aimed at cutting latency and inference cost.
3 min · 623 words
Priorities and principles for effective third party assessments
OpenAI outlines priorities and principles for rigorous, secure, independent third-party assessments of frontier models and safeguards—including access models, scope, and public reporting expectations.
9 min · 1,980 words
Unreal Labs introduces Unreal Agent, an agent harness claiming up to 40% cost savings versus Codex on production workloads and coding/science benchmarks, with details on architecture and evaluation.
5 min · 1,071 words
Carson Gross argues that Markdown has become a first-class source artifact for LLM-built software systems, and that treating it like code in /src follows the same locality and clarity principles as HTML-in-/src.
6 min · 1,398 words
Making the MiniMax H3 Video VAE 2x Faster
The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4 2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129 frame encode and decode round trip drops from 24.3 to 12.7 seconds. What changed, the technical details A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes.
2 min · 454 words
The MVUEH Break: GPT-6 Astra and a WWII Enigma message unsolved since 2005
Crypto Cellar Research documents how OpenAI’s GPT-6 Astra helped break the long-unsolved MVUEH Kriegsmarine Enigma message—methods, cribs, and what the recovered plaintext reveals.
4 min · 831 words
Tackling Robotics with (V)LM Agents
Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.
10 min · 2,304 words
Which is bigger: 9.11 or 9.9? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 9.9; 75 got it right. See every answer and who dissented.
8 min · 1,936 words
Artificial Intelligence in ResearchThe preface to a PhD thesis on AI, how research is changing, and the kind of researcher to become.
Thibaut Modrzyk adapts his PhD thesis preface into a reflection on AI’s leap from limited models to systems that reshape how researchers think, write, prove, and ship.
17 min · 3,811 words
Confused Deputy: The Old Bug That AI Agents Keep Reintroducing
Auth0 revisits Norm Hardy’s 1988 confused-deputy problem and shows how AI agents with ambient credentials recreate it—then argues for short-lived, task-scoped tokens instead of standing access.
9 min · 2,062 words
OpenAI introduces GPT-6.1 Sol, positioning it as near-Astra intelligence at a fraction of the price, with notes on capabilities, availability, and how it fits the GPT-6.1 family.
4 min · 805 words
Tom Batey of WebDepend argues that AI can run checks but cannot judge whether software satisfies expectations—so businesses still need human testers who own responsibility when AI writes and verifies code.
6 min · 1,412 words
Jev and System One Models: Calibration Beats Accuracy
A deep dive into TypeSafe’s Jev “System One” decision model: why calibrated probabilities matter more than raw accuracy for agents, games, and UIs that need millisecond choices.
9 min · 2,034 words
FLAWED’s Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization. On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
10 min · 2,234 words
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence IndexA 20% price cut, deeper cache discounts, and leading scores on agentic knowledge-work evals.
Artificial Analysis’s first look at Claude Opus 5.5: Intelligence Index score of 58 at max effort, parity with GPT-6 Astra on Terminal-Bench 4.0, stronger agentic knowledge-work results, and Anthropic’s $4/$20 pricing with cheaper cache reads.
3 min · 609 words