Documentation
Lower cost, route by benchmark, compare models, and cache repeated work.
Use lower-cost Flex capacity when it meets your latency needs.
Route through weighted benchmark aliases with automatic fallback.
Stop overpaying for every request. Find the cheapest model that performs on your real production traffic.
Reuse provider-cached prompt prefixes to reduce latency and cost.