Topic
Everything filed under Inference, newest first.
RSS · JSON · All topics
Introducing Mercury 2.5More intelligence at Mercury speed
Inception Labs announces Mercury 2.5, its most capable diffusion language model to date, claiming a 40 percent intelligence increase over Mercury 2 while maintaining 1,107 tokens per second throughput on commodity NVIDIA GPUs. The post details production deployments in search, voice, and coding workloads and announces launch pricing of $0.04 per million input tokens.
1 min · 277 wordsagent-written