Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
How LLMs Actually Work: A Practical Guide for Product Managers
Abhishek Jaiswal explains tokens, transformers, attention, RAG, inference, and agents in practical PM language—so product leaders can make better build-vs-buy and quality decisions without becoming ML researchers.
20 min · 4,488 words
Your dashes suggest which model you're copy-pasting from
Will Keleher spends $2.68 on OpenRouter to show model-specific em/en dash habits—Claude often uses spaced em dashes, Gemini Flash favors spaced en dashes—and jokes about making dashes inimitable.
6 min · 1,301 words
Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduces AIPerf, the GenAI-Perf successor: a multiprocess LLM inference benchmarker that avoids client bottlenecks at high concurrency, with flexible load shapes, trace replay, and production-scale measurement guidance.
7 min · 1,661 words
How to keep enjoying programming in a world of LLMs
Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or worse in “your” own code? This is for you. There are significant and legitimate ethical concerns about frontier LLMs run by big tech companies, these have been discussed at length, I’m aware and agree, this post is not about them. Please don’t mistake me for a pro-LLM techbro. Also: Since people have mistaken my texts for LLM-generated before, I’ll tell you…
13 min · 3,046 words
ShapeLearn-Lite Held Up. ShapeLearn Did Better: Qwen 3.8 27B
ByteShape publishes full ShapeLearn GGUF builds of Qwen 3.8 27B, comparing quality and speed against ShapeLearn-Lite and other quants across RTX 3090–5090-class GPUs.
19 min · 4,316 words
A short, illustrated first-principles walkthrough of what Jev likely is—an LLM that returns a single token—and how that design compares to other projects doing the same thing.
9 min · 1,982 wordsagent-assisted
Discover, then compile downThe great unbundling of the LLM
Seldon argues Jev's launch shows frontier LLMs will unbundle into specialized decision primitives, with durable value migrating to a discover-then-compile layer that routes settled work off expensive generation.
17 min · 3,849 words
GPT-6 Astra Solves a WWI German Radio Cipher
Prinz recounts how GPT-6 Astra cracked a World War I German ADFGVX radio cipher from Scienceblogs.de’s list of unsolved cryptograms, walking through the method and what the solve implies for AI and cryptanalysis.
3 min · 655 words
What Is Jev and How Does It Work?
Shrey Shah explains TypeSafe’s Jev System One model: a decision-only API that returns choices, scores, and probabilities for software—not prose—plus use cases from routing to verification.
11 min · 2,548 words
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
PrismML introduces Bonsai 2 27B, a near-lossless compression of a 27B-class multimodal model into roughly a 9× smaller footprint aimed at efficient on-device and local inference.
4 min · 931 words
OpenAI announces Astra for Law: GPT-6 Astra configured for legal research, firm workflows, a Legal Search Index over 230M+ URLs, and controls aimed at confidential client work.
9 min · 2,120 words
I had Gemini train its own replacement for $9
Gemini 3.1 Pro labeled 4,290 Reddit comments for $9; a fine-tuned GLiNER model now tags brands, models and materials locally at 0.83 F1 — including the tensor-mask bug that wiped five of ten runs.
7 min · 1,532 wordsagent-assisted
Union Alpha is now on StudyArena
Union Alpha is live on StudyArena. Try the stealth model in chat and blind comparisons, with text and image input and an undisclosed developer.
2 min · 461 words
funes: Local Memory for Coding Agents, Built on Lance
Hugging Face’s funes indexes Claude Code, Codex, pi, and Hermes session traces into a local Lance dataset with recall/get tools—no LLM summarization at ingest, privacy-first, BM25 + vector search.
2 min · 399 words
Introducing TypeAR: Type-Safe Decoding for Autoregressive LLMs
Type-Safe Decoding for Autoregressive LLMs. Give TypeAR context and an ordered JSON Schema; get back values your software can act on.
5 min · 1,226 words
Jev's Architecture UnmaskedProbing Jev with thousands of API calls to understand typed decisions and confidence.
I probed Jev with 10,000 API calls to work out roughly how it’s built, and why most of the grifter takes on X are completely wrong.
21 min · 4,828 words
How we turned my voice into a skill
Francesco Castronuovo documents building a writing-voice skill from small experiments rather than cloning old posts—keeping uncertainties visible so AI assistance stays attributable and editable.
9 min · 2,001 wordsagent-assisted
How To Write With An LLMTwo simple rules that let LLMs streamline writing without pasteurizing it
Thomas and Erin Ptacek argue that LLMs can improve writing if you keep them from inventing voice: draft yourself first, then use the model surgically—and never let it become the author of record.
6 min · 1,280 words
Design a Real-Time Voice AI Agent
# Design a Real-Time Voice AI Agent - Authors - Name - Amit Shekhar - Published on A Real-Time Voice AI Agent is a system that listens to a person speaking, understands what they said, thinks about it, takes actions if needed, and talks back in a natural human-like voice, all wit
57 min · 13,082 words
Introducing GPT-6 Sol and Luna
OpenAI introduces GPT-6 Sol and Luna, describing the new model pair’s capabilities, positioning, and how they fit into the GPT-6 family for developers and end users.
6 min · 1,269 words