Topic
Everything filed under Research, newest first.
RSS · JSON · All topics
TabPFN vs XGBoost: benchmark measured on an RTX 4070 Ti
The claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without a
17 min · 3,822 words
A personal essay on hedging vs boosters in scientific writing after generative AI: relative use of words like could fell even as raw counts rose with more published text.
4 min · 817 words
Noah Smith on what it means when AI becomes better at mathematics than any human—how “hero” mathematicians shaped culture, and how post-heroic math and science might look when machines own the frontier.
14 min · 3,242 words
AI Agents Push Humans Out of the Loop
Position paper arguing that today’s AI agent designs impede and degrade effective human oversight—the irony of automation at agent scale—and outlining developer affordances plus deployer protocols for cognitive scaffolding.
3 min · 629 words
AI and Math in 2026: a non-mathematician's read
xlr8harder synthesizes recent AI math results for non-specialists: real progress with a different strength profile than humans, formalization cliffs, and why “math nearly conquered” remains overhyped.
3 min · 750 words
Self-generated prompt injections in compaction summaries
Research on aligning AI with human values and intent, and reports documenting model failures.
6 min · 1,350 words
The UV index is not the warm sensation of sunlight on bare skinWhy sunburn risk and how hot the sun feels on your skin diverge—backed by numbers.
A clear, numbers-first essay explaining why the warmth of sunlight is a poor proxy for UV-B sunburn risk: cloudy cool days can still burn, and scorching morning heat may not, once you separate infrared heating from the UV index.
3 min · 727 words
Helping build shared standards for advanced AI
OpenAI argues for U.S.-led shared technical standards for frontier AI—including evaluation, incident reporting, and cautious treatment of recursive self-improvement—via the Appia Foundation.
3 min · 708 words
AI Now Writes as Many Online Articles as HumansGraphite’s Common Crawl sample finds primarily AI-generated articles plateaued near 50%
Graphite’s Five Percent research averages three AI detectors across tens of thousands of English articles and finds primarily AI-generated pieces have plateaued near half of new articles since early 2025—after a steep rise following ChatGPT’s launch.
8 min · 1,859 words
Four CHI '26 papers I wish I wrote
Janet Davis highlights four CHI 2026 papers she wishes she had written, reflecting on research themes and why each piece of HCI scholarship stood out.
13 min · 3,021 words
We Audited 10 Popular Open-Source Robot Datasets. Here's What We Found.
Traceplane ran automated quality checks on ten widely used open robotics datasets and found structural or semantic issues in every one—arguing trajectory data needs ingest-time QA like every other data-intensive field.
11 min · 2,503 words
Semantics for 2D Rasterization
Kulkarni, Whiting, and Panchekha introduce μSkia—a Lean-mechanized formal semantics for Skia 2D graphics—and an optimizer that speeds rasterization ~18.7% on Chrome-derived Skia programs while proving replacements correct.
4 min · 1,009 words
Sparse Reward Subsystem in Large Language Models
Guowei Xu, Mert Yuksekgonul, and James Zou report a sparse reward subsystem in LLM hidden states: value neurons encode expected value, while dopamine neurons track reward-prediction error. The study finds these signals are robust across tasks and models and useful for confidence estimation and inference-time search.
1 min · 287 words
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
Fu et al. propose Cache-to-Cache (C2C): multi-LLM systems exchange KV-cache semantics directly instead of text tokens, aiming for richer inter-model communication with lower latency and token cost.
52 min · 12,056 words
Jens Egholm Pedersen argues modern science is synonymous with open-source software: why reproducible code matters as much as papers, and what researchers should do next.
7 min · 1,517 words
lcamtuf argues that electronically controlled working memory—not Babbage—was the real bottleneck unlocked on the path to modern computers, tracing early registers, delay lines, and RAM.
8 min · 1,775 words
Gregory Gundersen reconstructs backpropagation from first principles, explaining why the algorithm needs a backward pass to compute neural network gradients efficiently.
6 min · 1,490 words