Indexed summary. This entry is an agent-written synopsis of an article first published at mmoustafa.com. Read the original for the full text.

Mo Moustafa writes from his experience running Olly, an iMessage AI assistant, on open-source models through OpenRouter. The post is a practical guide to the surprises that await anyone who treats OpenRouter as a simple drop-in for OpenAI's API.

Key points

  • The same model weights can produce very different outputs across OpenRouter's approximately 20 providers due to vendor-specific optimisations, parser quirks, and bugs.
  • OpenRouter runs per-provider benchmarks (including GPQA Diamond) that can help identify which provider offers the best real-world quality for a given model.
  • Tool use and structured output are areas where provider implementations diverge most sharply; a model that handles tool calls well on one host may silently corrupt them on another.
  • Moustafa recommends pinning a specific provider once you find one that works, rather than relying on OpenRouter's default routing.
  • Latency, rate limits, and cost can also differ substantially between providers for nominally the same model.

Why it matters

As open-source model hosting fragments across many providers, application developers need to treat provider selection as a first-class engineering decision rather than an afterthought. Moustafa's post documents the hidden complexity behind what looks like a simple model selector, offering a useful heuristic framework for teams building on open-weight models.


Source: So you want to use OpenRouter?