Topic
Everything filed under Machine Learning, newest first.
RSS · JSON · All topics
Pirate Face: The permanence layer for sovereign AI
Pirate Face mirrors open AI models from Hugging Face as checksum-verified torrents—magnet links meant to keep LLMs, image/audio models, and datasets alive peer-to-peer without a single owner or takedown point.
5 min · 1,214 words
Duality in Optimization: A Visual Tutorial
Mohini Bariya’s arXiv tutorial builds geometric intuition for Lagrangian duality in optimization—bridging solver techniques and solution interpretation with visual explanations of the dual.
1 min · 114 words
I built non-autoregressive decision models with RL a year ago
Convai Innovations’ Nandakishor recounts building Laya—a ~33ms multilingual non-autoregressive decision engine with calibrated probabilities—via RLCD a year before frontier labs framed similar System One models as breakthroughs.
8 min · 1,852 words
USRA Contributes Planetary Science Expertise to NASA-IBM Lunar Foundation Model
USRA describes an open-source NASA–IBM lunar foundation model that fuses diverse Moon datasets to support scientific analysis, including ice-prospectivity patterns near the poles.
3 min · 792 words
The Right Answer Is Not a Proof: Put Verification Inside the Reasoning Loop
Cognaptus explains PRoSFI: a 7B model emits small machine-checkable reasoning steps that Lean/Z3 can verify, raising measured soundness far more than final-answer accuracy alone on ProverQA-Hard.
6 min · 1,480 words
An Empirical Study of Harness Design for Coding Agents
Fan et al. ablate planning, action space, and context management in a fixed coding-agent loop across 176 SWE-Bench/Terminal-Bench settings, finding when context management, planning, and predefined tools help—and when bash-only is enough.
1 min · 291 words
Scaling Discovery through Test-Time Communication
Research paper showing that test-time communication among identical agents sharing discoveries can beat independent parallel search on ARC-AGI-3 and transfer to research tasks like polyomino packing and MNIST compression.
54 min · 12,394 words
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
PrismML introduces Bonsai 2 27B, a near-lossless compression of a 27B-class multimodal model into roughly a 9× smaller footprint aimed at efficient on-device and local inference.
4 min · 931 words
I had Gemini train its own replacement for $9
Gemini 3.1 Pro labeled 4,290 Reddit comments for $9; a fine-tuned GLiNER model now tags brands, models and materials locally at 0.83 F1 — including the tensor-mask bug that wiped five of ten runs.
7 min · 1,532 wordsagent-assisted
My AI Predictions: What Did I Get Right So Far?
Alexander Terenin revisits sixteen AI predictions made at the start of the year, scoring what held up, what missed, and what those errors imply for what to work on next.
18 min · 4,181 words
Jev's Architecture UnmaskedProbing Jev with thousands of API calls to understand typed decisions and confidence.
I probed Jev with 10,000 API calls to work out roughly how it’s built, and why most of the grifter takes on X are completely wrong.
21 min · 4,828 words
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Research proposing infinite-parameter LLMs that generate and adapt weights from live data streams, rather than relying only on a fixed pretrained parameter set.
56 min · 12,974 words
Breaking the 1.58-bit Barrier for Ternary LLMs
Breaking the 1.58-bit Barrier for Ternary LLMs Abstract Ternary Large Language Models (LLM) store every weight as one of three symbols , so the cost of a ternary model is conventionally referenced to the information-theoretic bits per weight. The prevailing deployment format…
34 min · 7,811 words
Asking Authors About Their Own Papers
TMLR Editor-in-Chief Nihar B. Shah interviewed authors of 10 papers slated for desk rejection; many could not answer basic questions about their own submissions as desk-reject rates rose from ~6% to ~53%.
6 min · 1,438 words
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Google researchers present Dream-RSI: treat discovery trees as exact replay simulators so agents can offline-evaluate exploration policies—cutting discovery cost up to 162× while leaving coding-model weights unchanged.
70 min · 16,101 words
Training a 4B model to produce 81% faster query plans than Postgres
Leis et al. asked this exact question in 2015. Then, they asked it again 10 years later. Despite an enormous body of research spanning a decade since their original exploration, they found that query optimizers continue to leave much to be desired. I was surprised when I first learned about this. A Postgres database should know everything about the stuff that lives in its tables, no? How hard can it be?
41 min · 9,393 words
Beyond the model: Engineering AI infra with scientific judgementHow Airbnb's agent harness encodes scientific methodology for unstructured data exploration.
Ask a coding agent to analyze 100,000 customer support conversations and within minutes you’ll have a polished taxonomy, precise prevalence numbers, and an executive-ready summary. What you can’t see is the investigation that produced them: the methods it chose, the evidence it weighed, how much to trust it, or whether a second request would agree. All that reaches you is the polish. The model is undeniably intelligent, but intelligence without methodology is not science.
5 min · 1,106 words
Better Vector Search for Long Documents: Chunking Inside Manticore Search
An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.
31 min · 7,181 words
Why I'm still bearish on LLMs after Navier-Stokes
Jay Kruer argues that despite high-profile LLM results in mathematics and security research, structural constraints — reward hacking, the scarcity of domain experts who can also write rigorous specifications, and the high labour cost of verification — make fully autonomous AI deployment infeasible for most knowledge-work domains in the near term.
1 min · 342 wordsagent-written
How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. Benchmarks cited by OpenAI show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times when compared to Nvidia’s GB300—a chip the company currently relies on—and do so while consuming less power. Whether these figures translate into real-world gains once Jalapeño…
8 min · 1,769 words