Topic
Everything filed under Research, newest first.
RSS · JSON · All topics
Do people prefer traditional architecture?Taste is subjective. But when it comes to architecture, there is a surprising level of agreement.
Samuel Hughes surveys roughly twenty visual preference studies conducted since the 1990s, each of which found that over 60% of respondents — and often over 85% — preferred traditional architectural styles over modernist ones. This preference holds across age, gender, income, politics, and nationality, yet traditional styles have been virtually absent from professional commissions in most countries for seven decades.
1 min · 312 wordsagent-written
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
Tagliabue, Dung, and Berg identify a linear “pain axis” in 25 open-weight models that responds to self-directed harm and steers models toward relief—even when that costs the user—sparking debate on functional signatures vs sentience.
31 min · 7,139 words
'Fingerprints' inside the Sun could reveal if it once swallowed a planet
A new study published in Monthly Notices of the Royal Astronomical Society proposes that if the Sun engulfed a super-Earth early in its history, that event would have left detectable chemical and structural signatures in the solar interior that helioseismology could potentially identify today.
1 min · 243 wordsagent-written
The mystery animal on an ancient god’s head
Signore Galilei investigates the strange animal depicted atop an ancient deity’s head—tracing iconography, competing identifications, and what the motif may have meant to its makers.
4 min · 977 words
Mathematician Daniel Litt argues that AI systems now capable of resolving major open problems need not mean the end of meaningful human mathematics, but they do require institutions to sharply distinguish mathematical understanding from mathematical text production. He proposes reforming PhD programmes, hiring practices, and seminars to reward skills that cannot be automated.
1 min · 290 wordsagent-written
The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, [Anthropic announced that future Claude models would embed an invisible watermark](https://www.anthropic.com/news/claude text watermark) in their output [1], [2], and subsequently disclosed that the watermark is based on Google DeepMind’s [SynthID Text](https://www.nature.com/articles/s41586 024 08025 4) [2], [3]. Text watermarking itself is not new, but its deployment now has regulatory relevance.
11 min · 2,640 words
Why machine learning research agents don't overfit — and what compression has to do with itNew research indicates that AI agents learn compressible models of data, which don't have enough space to enable memorization.
Amazon Science researchers explain why ML research agents fail to overfit benchmarks even after many evaluation rounds, arguing that successful agents learn highly compressible representations that are too compact to store memorised answers — connecting this to Minimum Description Length theory.
1 min · 247 wordsagent-written
# The Shape of Inference ## Watch the film 18 seconds In 1964, two radio astronomers in Holmdel, New Jersey, were losing a war with pigeons.
15 min · 3,541 words
Pretraining and scaling as a methodology and scientific perspective
Jiaxuan Zou’s essay on pretraining and scaling as a shared methodology across language, robotics, and world models—covering learning conditions, training/inference milestones, efficiency, stability, and predictability as scientific research practice.
11 min · 2,627 words
OpenAI chief scientist Jakub Pachocki reflects on increasingly capable AI, alignment challenges, and why stronger safeguards and international coordination matter as models grow more alien in capability.
14 min · 3,149 words
Graft, Metatron, and the two kinds of context coding agents need
Pavel Kerbel contrasts Graft’s recoverable WHAT/WHERE code maps with Metatron’s reviewed WHY/WHY NOT engineering memory, arguing stronger models still need both layers—and proposing a factorial eval to prove it.
9 min · 2,075 words
Project HydraFusion: Frontier quality via multi-model orchestration
In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated cost through multi-model orchestration.
7 min · 1,635 words
Pre-Greek: The lost language hidden within Ancient Greek
Daniel W. Hieber surveys ~1,000 Ancient Greek words without Indo-European etymologies—labyrinth, olive, Athens, Achilles—and reconstructs Pre-Greek substrate/adstrate influence via phonology, toponyms (-nth/-ss), and Minoan contacts.
3 min · 618 words
The Implications of Linguistic Illegibility for LLM Security
James Mickens argues that LLMs' external language and internal features can be illegible to humans and to each other—creating security implications when defenses assume readable, inspectable linguistic behavior.
38 min · 8,760 words
Frontis.AI / Horizon Research open-source OpenMLE (gym, RL, Evo) and Frontis-MA1-35B, lifting MLE-Bench Lite medal average to 71.21% under a single RTX 4090 budget toward executable RSI research.
2 min · 385 words
ZK-JPEG: Zero-Knowledge Image Editing and CompressionProving JPEG compression and edits without revealing the original image
Dittmer, Lu, Model, and Near present ZK-JPEG, a zero-knowledge tool that proves an image was correctly JPEG-compressed (and can verify a family of edits) from a secret committed input—bridging camera attestation with lossy encoding.
1 min · 328 words
Ryan Orbuch proposes Natural General Intelligence: a planetary 'nature model' grounded in Earth observation data to predict environmental responses and support stewardship—not just automate knowledge work.
64 min · 14,646 words
Developing provably correct Rust code with Verus
Many open-source and industry software projects, including several here at Amazon, are embracing the Rust programming language, since it provides performance and flexibility similar to that of the C programming language, while its clever type system automatically prevents a variety of bugs and security vulnerabilities. The result is fast code that's more correct and secure than average.
6 min · 1,443 words
Claude Fable 5.1 Solves the Cyphral DistichWe gave Claude Fable 5.1 an open task: solve an unsolved 370-year-old cipher. It solved it within a day.
Vals AI reports that Claude Fable 5.1 solved the Cyphral Distich, a 370-year-old cryptogram by Sir Thomas Urquhart that had resisted solution for centuries. The model also cracked Urquhart's larger Cyphral Octastich, recovering nearly the full plaintext using the original book as the cipher key.
1 min · 262 wordsagent-written
No Easy Fix for Bogus Respondents in Online Opt-In Polls
Pew Research Center tests trap questions, CloudResearch Sentry, and voter-file matching on 11,114 opt-in respondents: bogus cases still distort quality, and voter-file matching can raise error by discarding valid people.
16 min · 3,670 words