Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine

I see it as a point on the LLM pareto optimal curve in a regime that had a large revealed latent demand (no thinking, single token, low latency acceptable intelligence) that was under-invested into because of a race to higher intelligence.

  • Karpathy-san, on Jev

While everyone and their cousin is loudly building coding agents and chatbots, there’s a quieter inference revolution going on in the backend. Simple LLM transformations of data can be incredibly powerful, provided the cost-performance is good enough — just scroll social media and catch a few of the eye-popping, hack-inspiring demos of TypeSafe AI’s Jev model.