Topic
Everything filed under Research, newest first.
RSS · JSON · All topics
Why consciousness is more likely a property of life than of computation and why creating conscious, or even conscious-seeming AI, is a bad idea.
34 min · 7,804 words
Surprisingly Complex Waves Reveal the Brain's Inner Workings
Unexpected patterns traveling across the human brain may be reorganizing its activity in real time.
9 min · 2,113 words
Functional ultrasound imaging from scratch
Functional ultrasound imaging (fUSI) is a brain imaging modality that is just starting to be demonstrated in human studies. I’m sure many will be familiar with structural ultrasound imaging, which forms static images using ultrasound, as in a sonogram1. Functional ultrasound instead captures movies of brain tissue very rapidly, tracking subtle changes caused by the blood flow that follows neural activity.
11 min · 2,451 words
SynthID Bio: Watermarking methods for synthetic biology
Introducing SynthID Bio Proof of concept for watermarking AIgenerated proteins while preserving biological function. Today, we’re introducing SynthID Bio to bring watermarking technology to synthetic biology.
6 min · 1,276 words
Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever
Qingcheng Zeng gives Muon and NorMuon the same tuning budget as AdamW when fine-tuning a contrastively pretrained dense retriever: lower training loss, no BEIR win. Learning rate and transfer matter more than the optimizer.
5 min · 1,151 words
GLM-5.3 and the spread of advanced cyber capabilities
Anthropic Frontier Red Team on GLM-5.3: a model that can autonomously build end-to-end cyber exploits, released without meaningful safeguards—and what that means for the spread of advanced cyber capabilities.
8 min · 1,815 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
Responsible Release of AI-Generated Mathematics
The Advisory Group on Mathematics and AI (Sep 29, 2026) recommends how frontier labs should release AI-generated math results: deposit promptly, cite related work, formalize where possible, disclose prompts and costs, and fund community-led human understanding.
2 min · 559 words
How we found 24 Android vulnerabilities using our open source AI security agent
GitHub Security Lab explains the targeted AI taskflows behind 24 Android findings, the bugs they uncovered, and how to run the same open-source Taskflow Agent on your own app.
10 min · 2,349 words
Historians often blame drought, famine, and invaders for the Late Bronze Age collapse—but that can’t explain why the empires never came back. Patrick Fitzsimmons argues for a deeper structural story.
14 min · 3,253 words
Towards safety cases for frontier AI training
OpenAI argues frontier RL runs should require structured safety documentation approaching “safety cases”: technical safeguards, operational practices, and incident investigation before continuing training.
7 min · 1,571 words
Meta FAIR introduces RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics): learn rubrics from the gap between expert writing and model output, then train models toward expert-level text generation to reduce AI slop.
16 min · 3,572 words
Can a Model Learn New Skills as Add-Ons?
Connito Research trains residual MoE experts with their own routers on a frozen DeepSeek-V2-Lite base, then merges independently trained math, code, medical, law, and finance experts in seconds without retraining—lifting domain benchmarks while leaving the original model untouched.
4 min · 913 words
On the Value of Doing a PhD in the Age of AI
MIT's Phillip Isola offers ten reminders for anxious AI PhD students: public research still has leverage, human expertise remains safety infrastructure, and the PhD's job is to chase a moving frontier for the love of the game.
3 min · 781 words
Standard deviation or standard error? What each one measures, and which one belongs in your paper
Selçuk Korkmaz clarifies SD vs SE with interactive figures: use SD to describe sample spread, and report estimates with confidence intervals rather than bare standard errors.
16 min · 3,717 words
How Pew Research Center is – and is not – using AI in our work
Pew Research Center outlines internal AI guidelines: surveys stay human-answered, disclosure rules for production use, and careful experimentation while keeping research people-centered.
2 min · 542 words
EX-ARRR: Sailing the 0-click Seas
**Every serious Apple device compromise of the last decade has boring person at the bottom of it: a parser read a file and trusted it a little too much. Not a phishing link, nor a stolen password, but a daemon you never launched, decoding a file you never opened, one byte past the end of a buffer. This is that story. It starts late one evening with a fuzzer that did not know what an EXR file was, and ends with a heap overflow that fires inside a privileged Apple daemon the instant an iMessage lands evading BlastDoor’s, before the little notification banner even finishes…
19 min · 4,447 words
Throughout the history of AI, open research has played a critical role in driving progress. Today, many key details of frontier large language models (LLMs) remain proprietary, but open-weights model families— such as DeepSeek, Kimi, and MiMo —continue to provide a valuable window into the development process for modern LLMs.
7 min · 1,609 words
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.
7 min · 1,560 words
Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution
A research write-up proposing Dual-Sided Andon: out-of-band, kernel-boundary preemption and containment for runaway AI agents, arguing application-level kill switches are insufficient.
21 min · 4,848 words