AI APIs106 · Module III · Lesson 13 of 14
Visuals · 8 min

Cost comparison

The per-token grid, with worked examples.

Summary

The per-token cost matrix, with worked examples.

Worked Example

Cost of one chat turn

Typical turn: 2K input tokens + 500 output tokens. On GPT-4o: 2K × $2.50/M + 500 × $10/M = $0.005 + $0.005 = $0.01. On Llama 3.3 70B via Together: 2K × $0.60/M + 500 × $0.80/M = $0.0016. Frontier is ~6× the open-model cost per turn on this shape.

Visuals
Approximate 2026 API pricing (per 1M input / output tokens, USD)

Frontier: GPT-4o ~$2.50 / ~$10; Claude Sonnet 4 ~$3 / ~$15; Gemini 2.5 Pro ~$1.25 / ~$5; Grok 4 ~$3 / ~$15. Reasoning: o3 ~$10 / ~$40; Claude thinking-mode ~$3-15 with thinking tokens billed as output. Open on Together/Fireworks: Llama 3.3 70B ~$0.60 / ~$0.80; DeepSeek V3 ~$0.30 / ~$1.10; Qwen 72B ~$0.60 / ~$0.80. Local Ollama on your GPU: $0 marginal after amortized hardware. Prices change monthly; always verify current rates before committing to a budget.

Key Ideas
  • Input tokens are cheap; output tokens are expensive.
  • Reasoning models are 5-20× per token; use where they earn it.