Topic
Everything filed under Research, newest first.
RSS · JSON · All topics
Tackling Robotics with (V)LM Agents
Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.
10 min · 2,304 words
Artificial Intelligence in ResearchThe preface to a PhD thesis on AI, how research is changing, and the kind of researcher to become.
Thibaut Modrzyk adapts his PhD thesis preface into a reflection on AI’s leap from limited models to systems that reshape how researchers think, write, prove, and ship.
17 min · 3,811 words
How Zambia Approved an HIV Drug in 12 Days
How Zambia’s medicines regulator authorized an HIV drug in 12 days—and what that says about the access gap when medicines sit where the disease is not.
16 min · 3,572 words
FLAWED’s Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization. On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
10 min · 2,234 words
What Happens When Formalization Becomes Cheap?
I started working on machine learning for formal theorem proving in 2018. When people ask how I got into the field so early, I sometimes give an answer that makes me sound quite visionary. The actual story is that my advisor had a student leaving, and he assigned the project to me. My apologies to everyone who got the visionary version. It was a fortunate assignment. Over the following years, I developed CoqGym and LeanDojo and contributed to Goedel-Prover. I was lucky to join a small research…
13 min · 2,890 words
The Machine-Native Economy: How digital assets connect intelligence, commerce, and compute
BlackRock Digital Assets Research argues agentic AI needs machine-native payment rails (stablecoins/blockchains) and explores tokenized compute as a converging digital-asset use case.
17 min · 3,814 words
Autonomous AI Agents are breaking into Online Retailers for $25 a target
Gambit Security reconstructs an ongoing campaign where open-source AI harnesses attack retailers at ~$25/target, steal 600k+ cards, inject skimmers, and sometimes wipe databases during cleanup.
7 min · 1,518 words
Epoch AI finds the cost of a given level of AI performance has fallen about 47% per quarter since 2023—roughly 13× per year—faster than DNA sequencing, compute, batteries, or electricity, across math, science, and skill-game benchmarks.
40 min · 9,215 words
MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement
Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.
4 min · 839 words
Here’s What California Is Learning From Solar Panels Built Over Irrigation CanalsProject Nexus in Turlock tests canal-top solar for water savings, clean power, and water quality.
KQED’s report on California’s Project Nexus pilot: building solar arrays over working irrigation canals to generate power, cut evaporation, and study water-quality effects—early lessons from Turlock Irrigation District’s canal-top panels.
7 min · 1,680 words
What Is RLCD? The Secret Behind Jev
Di Zhang explains RLCD (schema-conditioned Plackett–Luce reward modeling) and how Jev turns calibrated multiway decisions into a product—making the reward model the model rather than hiding it behind a generator.
10 min · 2,324 words
Measurements for understanding the pace of AI development inside frontier labs
Anthropic proposes public metrics for the pace of frontier AI development so outsiders can see what is happening inside labs—beyond marketing and model cards.
19 min · 4,401 words
How GPT-6 Astra ascended NetHack: setup, agent loop, tool use, failure modes, and what beating a famously hard roguelike says about LLM agents in open-ended environments.
11 min · 2,477 words
Announcing the Advisory Group on Mathematics and Artificial Intelligence
Nine leading mathematicians announce an independent IAS-hosted advisory group to counsel AI labs on releasing math results—starting with OpenAI’s claim of 100+ solved open problems—and invite community input.
2 min · 378 words
In Search of a Compositional Theory of Self-Stabilization
Murat Demirbas connects metastable failures to classical self-stabilization and asks what a compositional theory would need to make distributed recovery composable.
10 min · 2,248 words
dlab Open Source Week: Frontier AI on Your Own Hardware
Tim Dettmers argues that small academic labs can compete by building coherent open-source ecosystems rather than isolated papers. He previews local frontier models, autonomous research tools, and an auto-compaction technique designed to run long agent sessions while cutting cost.
1 min · 276 words
Arrow heads at Obi-Rakhmat (Uzbekistan) 80 ka ago?
A PLOS ONE study examines stone points from Obi-Rakhmat Cave and asks whether they indicate early projectile / arrow technology around 80,000 years ago in Central Asia.
66 min · 15,269 words
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
LLM groups replaying human Wason discussions reach full consensus far more often than humans—partly because agents almost always speak up—cautioning against treating multi-agent agreement as truth.
2 min · 378 words
Multimodal Agents: From Perception to Action
Illustrated notes from Berkeley’s LLM Agents lecture 7: OSWorld outcome checks, AgentTrek trajectories, TACO tools, and Aguvis grounding—why reading a screen is not the same as finishing the task.
8 min · 1,915 words
The Millennium Problems for Biology
A proposed set of millennium-scale open problems for biology—framed as hard, motivating challenges analogous to the Clay Mathematics Millennium Prize Problems.
7 min · 1,615 words