Benchmarking inference at scale: coding agents

Together AI Blog
Generative AI

Real-world inference benchmarks for coding agents: 31% TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.