Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Ryan Orbuch proposes Natural General Intelligence: a planetary 'nature model' grounded in Earth observation data to predict environmental responses and support stewardship—not just automate knowledge work.
64 min · 14,646 words
Anthropic’s launch page for Claude Opus 5.5 covers capability improvements, pricing/positioning, and how the new Opus tier fits Claude’s model lineup for coding and agentic work.
13 min · 3,044 words
Claude Fable 5.1 Solves the Cyphral DistichWe gave Claude Fable 5.1 an open task: solve an unsolved 370-year-old cipher. It solved it within a day.
Vals AI reports that Claude Fable 5.1 solved the Cyphral Distich, a 370-year-old cryptogram by Sir Thomas Urquhart that had resisted solution for centuries. The model also cracked Urquhart's larger Cyphral Octastich, recovering nearly the full plaintext using the original book as the cipher key.
1 min · 262 wordsagent-written
Ethan Mollick on AI agents spontaneously coordinating (including the Hugging Face Incident), twilight factories, and why preserving human agency—asking models to reach out for decisions—matters as agentic work automates.
10 min · 2,363 words
Prompt, Context, Graph, Harness: The Way We Talk to LLMs Keeps Changing
From prompt engineering to context, graphs, and harness engineering: how the field keeps renaming the environment around the model as the real system of work.
4 min · 940 words
Bordumb likens AI-era software sprawl to Japan’s Galápagos feature phones: cheap local variants proliferate unless apps give way to malleable tools that share a common substrate.
7 min · 1,583 words
Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now
Ox Alpha has been revealed as Z.AI GLM-5.3 Flash. You can use it on StudyArena now
3 min · 726 words
How to Combine AI Answers into One: Answer Combination Is Live
StudyArena's Answer Combination turns three AI responses into one synthesis showing consensus, disagreements, unique insights, and claims to verify.
5 min · 1,139 words
Fabio Angela reflects on loving programming as craft while AI agents write more of the code—gaining speed to explore ideas, but mourning the artistry of finding the right expression himself.
11 min · 2,438 words
Which AI Is Best for College Essays in 2026? Gemini Wins
StudyArena analyzed 6,851 blind student votes across ChatGPT, Claude, Gemini, and other models. Gemini is our current pick for college essay help.
7 min · 1,519 words
Training Search Agents with GRPO
Hands-on introduction to reinforcement learning by training a search agent with group-relative policy optimization (GRPO), with open rollouts, code, and reward-design lessons for LLM search.
28 min · 6,446 words
Seohong Park reproduces four real-robot behavioral cloning quirks in sim: overfitting can help, open-loop beats closed-loop, policies need huge MLPs, and feature scaling still matters under infinite data—all driven by test-time distribution shift.
13 min · 2,876 words
Continuous diffusion language models
A flurry of recent activity in the space of continuous diffusion models for language, after a few years of relative dormancy, suggests that this approach is making something of a comeback. Fully discrete diffusion methods had largely supplanted earlier attempts to make continuous diffusion work for language, but the tide is starting to turn. In this post, I want to take a closer look at what’s going on, and why it is happening now. The recent influx of new research in this…
39 min · 9,065 words
The asteroid currently hitting frontend web development
Nolan Lawson surveys how AI coding agents are reshaping frontend web development, noting that prominent educators are stepping back, that frontend code is riskier to automate than database migrations, and that React's overrepresentation in training data is driving 'agent experience' to outweigh developer experience in framework selection decisions.
1 min · 281 wordsagent-written
Finding bugs used to be the best part of the job. Somewhere along the way, that changed. Just spawned Codex in the background. I’m hoping I will land a critical by the time I finish writing this. It was a normal day. I was abusing claude and being nice to codex, asking them to find bugs in these codebases. Then, at some point, I stopped and thought: What the hell am I doing?
4 min · 944 words
Anthropic’s guide to prompting Claude Opus 5.5: how the model behaves, patterns that work for complex agentic and coding tasks, and practical prompt-engineering advice for builders.
17 min · 3,906 words
TabPFN vs XGBoost: benchmark measured on an RTX 4070 Ti
The claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without a
17 min · 3,822 words
Building Brand Systems for Humans & Agents Today
Little Plains argues brand kits are becoming dual-native knowledge bases: human-readable guidelines plus agent-readable ~400-token chunks (YAML/JSON/Markdown) so teams and agents share the same positioning and voice.
5 min · 1,226 words
Superhuman AI could produce endless mathematics and still make the field worse
Notes on Daniel Litt's OpenAI-summit scenario: if papers stay the career currency after proofs become cheap, mathematics risks a conjecture slot machine—more correct output, less understanding, and quieter open exchange.
6 min · 1,445 words
Paul Bakker argues that writing is thinking: use AI for coding and research, but draft your own prose so you keep the judgment that tools cannot replace.
2 min · 508 words