Sep 21, 2026 Jay Lux Ferro agentsllmllm-agentspythoninfrastructurebenchmarks llm-agents For months, a 1-billion-parameter model validated sensitive-data detections in our proxy pipeline. Last week we ran it against a 40-line Python function. The model kept 85% of the deliberately invalid test data — fake credit cards with broken checksums — and quietly dropped real IBANs, real social security numbers, and a person’s actual email address. The function caught every fake, kept every real one, and answered in ten microseconds.

We didn’t set out to pick a fight with local LLMs. We set out to answer one question — can each layer of our pipeline do the same job, or better, without a local model? — and we refused to flip any switch until a benchmark said yes. What follow are the four experiments, the numbers, and the pattern hiding underneath them. The pattern matters more than the numbers.