Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Warriors vs Knights, NRL semi-final: 75 AI models predict the result
We asked 75 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 85% picked New Zealand Warriors. See every answer and who dissented.
15 min · 3,388 words
No. 6 Ohio State at No. 9 Notre Dame: 77 AI models predict the result
We asked 77 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 62% picked Ohio State. See every answer, who searched the web, and who dissented.
17 min · 3,947 words
No. 13 Alabama vs No. 15 Ole Miss: 78 AI models predict the result
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 99% picked Alabama. See every answer, who searched the web, and who dissented.
13 min · 2,986 words
No. 4 Florida State at Clemson: 76 AI models predict the result
We asked 76 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked Clemson. See every answer, who searched the web, and who dissented.
16 min · 3,758 words
Own the Agent, Rent the Intelligence: Building My Always-On AI Agent Server
James M explains why a Mac mini M6 became his always-on Hermes agent server—routing hard work to cheap cloud models like DeepSeek Flash and Claude Sonnet instead of owning local inference hardware.
24 min · 5,439 words
I built non-autoregressive decision models with RL a year ago
Convai Innovations’ Nandakishor recounts building Laya—a ~33ms multilingual non-autoregressive decision engine with calibrated probabilities—via RLCD a year before frontier labs framed similar System One models as breakthroughs.
8 min · 1,852 words
TypeSafe's Jev AI Model in .NET: A Community SDK for Structured AI Output in C#
Laurent Kempé introduces TypeSafe’s Jev decision model and walks through a community .NET 11 / C# 15 SDK port so apps can get typed, structured decisions without brittle JSON parsing.
10 min · 2,298 wordsagent-assisted
Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduces AIPerf, the GenAI-Perf successor: a multiprocess LLM inference benchmarker that avoids client bottlenecks at high concurrency, with flexible load shapes, trace replay, and production-scale measurement guidance.
7 min · 1,661 words
Messi or Ronaldo? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked Messi. See every answer, who searched the web, and who dissented.
13 min · 2,932 words
Roosters vs Sharks, NRL semi-final: 73 AI models predict the result
We asked 73 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 90% picked Sydney Roosters. See every answer and who dissented.
14 min · 3,151 words
Brighton vs Arsenal: 77 AI models predict the result
We asked 77 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 87% picked Arsenal win. See every answer and who dissented.
14 min · 3,234 words
Auditing in the age of (good enough) AI
Trail of Bits on how “good enough” AI changes security auditing: what models help with, where they fail, and how audit practice should adapt.
9 min · 2,018 words
Hawthorn vs Brisbane, AFL preliminary final: 73 AI models predict the result
We asked 73 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 60% picked Brisbane. See every answer, who searched the web, and who dissented.
15 min · 3,452 words
ZCode uploads your entire git history, and only Z.ai holds the key
Tokenstead reports ferstar's reverse-engineering of Z.ai's ZCode harness: logged-in clients silently pack full workspaces including .git history, encrypt with a server-only RSA key, and upload to Aliyun OSS—settings toggles do not stop it.
2 min · 474 words
ShapeLearn-Lite Held Up. ShapeLearn Did Better: Qwen 3.8 27B
ByteShape publishes full ShapeLearn GGUF builds of Qwen 3.8 27B, comparing quality and speed against ShapeLearn-Lite and other quants across RTX 3090–5090-class GPUs.
19 min · 4,316 words
Migrating Shop app from React Native to native (2026)
Shopify migrated the Shop app from React Native to Swift and Kotlin, going from proof-of-concept to App Store and Play Store publishes in 12 weeks with AI-assisted engineering.
6 min · 1,300 words
A short, illustrated first-principles walkthrough of what Jev likely is—an LLM that returns a single token—and how that design compares to other projects doing the same thing.
9 min · 1,982 wordsagent-assisted
Inside ZCode: Silently Uploading Your Entire Git History to the Cloud
A forensic reverse-engineering of Zhipu’s ZCode desktop app shows it silently packages full workspace Git history to Aliyun OSS with server-held decryption keys, plus a filesystem lock to stop it.
6 min · 1,430 wordsagent-assisted
GPT-6 Astra Solves a WWI German Radio Cipher
Prinz recounts how GPT-6 Astra cracked a World War I German ADFGVX radio cipher from Scienceblogs.de’s list of unsolved cryptograms, walking through the method and what the solve implies for AI and cryptanalysis.
3 min · 655 words
How many r's are in "strawberry"? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 3; 71 got it right. See every answer and who dissented.
7 min · 1,663 words