Topic
Everything filed under LLMs, newest first.
RSS · JSON · All topics
Halo: Frontier-Lab Training for Everyone
White Circle open-sources Halo, a Hugging Face–native training framework claiming up to ~2.8× faster post-training than stock TRL with lower memory use—from single GPU to multi-node, keeping checkpoints in native HF format.
18 min · 4,098 words
What Is RLCD? The Secret Behind Jev
Di Zhang explains RLCD (schema-conditioned Plackett–Luce reward modeling) and how Jev turns calibrated multiway decisions into a product—making the reward model the model rather than hiding it behind a generator.
10 min · 2,324 words
The current balance of power in open modelsThe expanded form of a testimony prepared for Congress on open-weight models and U.S.–China competition.
Nathan Lambert’s congressional briefing notes on open vs open-weight vs closed models: where Chinese labs lead, what Western open weights still hold, and why policy should treat open distribution as a strategic variable—not a binary.
12 min · 2,866 words
Arcturus Labs compares OpenAI’s emerging decision-model direction with TypeSafe’s Jev—and asks whether a frontier lab can absorb the System One / structured-decision niche startups are building.
12 min · 2,748 words
Measurements for understanding the pace of AI development inside frontier labs
Anthropic proposes public metrics for the pace of frontier AI development so outsiders can see what is happening inside labs—beyond marketing and model cards.
19 min · 4,401 words
The Function That Beat the Model: What We Measured When We Removed the LLMs
SPERIXLABS replaced a 1B-parameter local model that validated sensitive-data detections with a 40-line Python function, then published the four experiments showing where classical checks beat the LLM on accuracy and latency.
7 min · 1,535 words
Introducing Strands harness: frontier performance with 28% lower token cost
Arron Bailiss introduces Strands harness: a fully assembled, customizable local/cloud agent that aims for Claude Code/Codex-like “it just works” behavior with about 28% lower token cost.
5 min · 1,198 words
The Claude DelusionWhat if we're the ones having hallucinations?
Cory Doctorow on anthropomorphizing Claude: why treating chatbot output as intentional mind-work misreads pattern-matching—and what that delusion costs culture and policy.
9 min · 1,958 words
Amit Shekhar walks through how LLM design moved from RNNs to attention, Transformers, scaling laws, Mixture of Experts, and the open problems still ahead.
25 min · 5,729 words
Dan McKinley argues that obsessing over prompt text misses the point: build interlocking evaluation and optimization pipelines, and treat LLMs as non-conscious systems whose 'meaning' is mostly our projection.
12 min · 2,673 words
Rebuilding Nym’s agent around Jev
How Nym rebuilt its agent stack around TypeSafe’s Jev for guardrails, browser actions, and tool selection—with benchmarks and a shopping demo.
10 min · 2,186 words
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
LLM groups replaying human Wason discussions reach full consensus far more often than humans—partly because agents almost always speak up—cautioning against treating multi-agent agreement as truth.
2 min · 378 words
A technical walkthrough of what Jev-style decision models appear to be doing—single-token decisions, structured outputs, and why the architecture matters for agent tooling.
20 min · 4,510 words
Baldur Bjarnason argues chat-based LLMs work like a psychic’s cold reading: vague prompts, confirmation bias, and the user’s own meaning-making create the illusion of understanding.
22 min · 5,118 words
How Instinct's memory works: a reverse-engineering teardown
A black-box teardown of Instinct
15 min · 3,416 words
aie_2.1: building it — a conversation engine where the LLM 'remembers'
Meraki builds a Python conversation engine that shows how LLMs fake memory: every turn resends the full history so the model appears to remember, with practical notes on context and design.
9 min · 2,124 words
Keva: Running Coding Agents On-Device on Unrooted Android
Simon Lin's technical paper on Keva—an on-device Android AI coding agent running Claude Code/Codex-style loops—covering architecture, failure modes, and systems lessons without rooting the phone.
33 min · 7,580 words
Model Context Protocol with Spring AI, Building MCP Clients and Servers in Java
Ayush Shrivastava walks through building MCP clients and servers with Spring AI in Java: tool discovery, protocol basics, and wiring MCP into agentic Spring applications beyond a basic demo.
13 min · 2,972 words
I built non-autoregressive decision models with RL a year ago
Convai Innovations’ Nandakishor recounts building Laya—a ~33ms multilingual non-autoregressive decision engine with calibrated probabilities—via RLCD a year before frontier labs framed similar System One models as breakthroughs.
8 min · 1,852 words
TypeSafe's Jev AI Model in .NET: A Community SDK for Structured AI Output in C#
Laurent Kempé introduces TypeSafe’s Jev decision model and walks through a community .NET 11 / C# 15 SDK port so apps can get typed, structured decisions without brittle JSON parsing.
10 min · 2,298 wordsagent-assisted