Status: experiment design demo; measurements will be added after real model endpoints are connected.
Model routing cannot optimize for price per call alone. A minimum experiment records:
| Metric | Why it matters |
|---|---|
| Task success | Whether the output is actually usable |
| Time to first token | Interactive responsiveness |
| Total latency | End-to-end throughput |
| Cost per success | Avoiding models that are cheap but fail more often |
The experiment separates extraction, structured generation, and multi-step reasoning, then selects the lowest-cost model that satisfies the quality threshold for each tier.