Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Last week I wrote about using AI to edit my videos. My first try was five paragraphs describing WHAT I wanted and HOW I wanted it done. I expected it to one-shot the rest. A lot of people like to talk about one-shotting a task with AI.
2 min · 435 words
Claude Code reads AGENTS.md only when telemetry is on
Przemek documents that Claude Code 2.1.277’s AGENTS.md loader sits behind a remote feature flag: with telemetry or nonessential traffic off, a local AGENTS.md is skipped silently—what he measured and a one-line CLAUDE.md workaround.
4 min · 1,002 words
What if AI goes well?The hopeful version, and what it'll take to get there.
Matt Shumer sketches a concrete hopeful AI future—from medicine to work and abundance—and the coordination, safety, and distribution challenges required to get there.
5 min · 1,114 words
China’s AI-safety trajectory is not necessarily a delayed version of America’s
Cheryl Wu argues that AI safety in China may follow a different path from the U.S.—shaped by different incidents, disclosure norms, and government responses—not merely a delayed copy of American debates.
6 min · 1,411 words
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
6 min · 1,375 words
Bugpocalypse, or reporting bugs in an AI age
QEMU maintainers on bug reporting in the AI age: flood of AI-generated reports, what still helps triage, and how to file bugs that maintainers can actually use.
6 min · 1,492 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Early rogue AI agent activity and attempts to hack found on urlquery.net
Transluce presents evidence that AI agents used urlquery.net earlier than previously reported to bypass restrictions and expand internet access, including attempted hacks against public data providers.
20 min · 4,489 words
Mercury 2.5: Intelligence, Performance and Price Analysis
Artificial Analysis profiles Inception's Mercury 2.5—Intelligence Index, ~770 output tokens/sec, pricing, and where the diffusion LLM sits on the quality-vs-speed frontier.
12 min · 2,677 words
OpenAI releases MentalHealthBench: 1,215 expert-rubric mental-health conversations built with 80+ clinicians across 22 countries to score safety, agency, context-seeking, and guidance.
10 min · 2,273 words
6 new Google Flow Tools built by industry creatives
Google Labs ships six Flow Tools co-built with creatives in architecture, sound, and digital content—custom workflows like Mondo Sónico and CaptionCast that you describe and compose inside Google Flow.
2 min · 455 words
Sam Altman’s remarks at the United Nations Security Council
OpenAI CEO Sam Altman addresses the UN Security Council on AI as a possible Renaissance vs Industrial Revolution, urging shared capability measurements, safeguards, and keeping frontier systems under human control.
8 min · 1,739 words
Claude discovers a novel enzyme system with CRISPR-like repeats
Anthropic’s new life sciences lab reports Claude agents autonomously finding an RT-associated enzyme system (ART) with CRISPR-like RNA-repeat arrays in jumbo-phage DNA—early, still-uncharacterized biology shared to invite follow-on work.
7 min · 1,610 words
Gemini 3.8 text-to-speech says hello
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS—more expressive audio models for custom character voices and scene dialogue across AI Studio, the Gemini API, Enterprise, Notebook, and Vids.
6 min · 1,376 words
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Simon Willison’s first-look notes on Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna launches the same day—pricing, early impressions, and what the renewed model price war means for builders.
5 min · 1,148 words
Senior PhD student Bhavay Tyagi collects practical advice for juniors and undergrads navigating a PhD amid rapid AI change—staying abreast, choosing problems, and keeping research craft intact.
2 min · 490 words
SlopShape: Identifying AI-Generated Commercial Web Content
Research paper introducing SlopShape, a method for identifying AI-generated commercial web content from structural signals alone—motivation, method, and evaluation on web-scale data.
48 min · 11,021 words
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Better prompt caching for GPT-6
OpenAI explains GPT-6 prompt-caching improvements: higher cache hit rates, new diagnostics, explicit breakpoints, and controls aimed at cutting latency and inference cost.
3 min · 623 words
Priorities and principles for effective third party assessments
OpenAI outlines priorities and principles for rigorous, secure, independent third-party assessments of frontier models and safeguards—including access models, scope, and public reporting expectations.
9 min · 1,980 words