Great for customer support, content production, knowledge-base Q&A, and tool-chain agents.
Context
256K tokens
Availability
1/1 available
Reference latency
2.50s
Inferred from the model family and tags; actual calls are authoritative.
Final cost is determined at settlement.
Pricing comes from the unified pricing config; balance and budget are checked before each call.
Rate-limit fields sync from the console config; no guarantees until synced.