All posts

DeepSeek V4 vs GPT-5.5 vs Claude Opus 4.8: price and performance (2026)

· DeepSeek· Comparison· Pricing

The interesting question in mid-2026 isn't whether a Chinese open-weight model can compete with the closed frontier. DeepSeek V4-Pro already tops several coding leaderboards. The real question is what you give up, and what you save, by choosing it. Here's the data.

The contenders

  • DeepSeek V4-Pro (released April 24, 2026): 1.6T total parameters, 49B active. DeepSeek's reasoning and coding flagship.
  • DeepSeek V4-Flash: 284B / 13B active. The fast, ultra-cheap tier.
  • GPT-5.5 (OpenAI): the closed agentic-coding flagship.
  • Claude Opus 4.8 (Anthropic, released May 28, 2026): 1M-token context, strong on long agentic tasks.

Coding benchmarks

  • LiveCodeBench: DeepSeek V4-Pro 93.5, the highest score among all evaluated models, closed APIs included. GPT-5.5 scores above 85.
  • Codeforces (Elo): DeepSeek V4-Pro 3206, a grandmaster-level rating.
  • Terminal-Bench 2.0 (real terminal and agent tasks): DeepSeek V4-Pro 67.9%, ahead of Kimi K2.6 (66.7%). V4-Flash trails at 56.9%.
  • SWE-Bench Pro (real GitHub issue resolution): GPT-5.5 and Kimi K2.6 both around 58.6%; DeepSeek V4-Pro at 55.4%. (GLM-5.2, released later in June, took the open lead here at 62.1.)

The takeaway: on raw code generation and competitive programming, V4-Pro is at or above the closed frontier. On long-horizon *agentic* repair (SWE-Bench Pro), GPT-5.5 and the best open models are roughly even. The closed frontier no longer has a clear moat there.

Reasoning and knowledge

  • GPQA Diamond (hard science Q&A): DeepSeek V4-Pro 90.1%, MiniMax M3 92.68%, Gemini 3.1 Pro on top at 94.3%.
  • MMLU: GPT-5.5 around 92.4%.

V4-Pro is competitive but not the outright leader on the hardest reasoning sets. That's where the closed frontier still edges ahead.

Price: where the gap is enormous

Per 1M tokens (input / output):

  • DeepSeek V4-Pro: $0.435 / $0.87.
  • DeepSeek V4-Flash: $0.14 / $0.28, with cache hits near $0.014.
  • GPT-5.5: $5.00 / $30.00.
  • Claude Opus 4.8: $5.00 / $25.00.

On output, the dominant cost for most workloads, DeepSeek V4-Pro is roughly 34× cheaper than GPT-5.5 and about 29× cheaper than Claude Opus 4.8, while matching or beating them on coding. V4-Flash is on the order of 100× cheaper than the closed flagships for tasks that don't need the top reasoning tier.

So which should you use?

  • Coding, agents, high volume → DeepSeek V4-Pro. Frontier-level code at a fraction of the cost.
  • Cheap classification, extraction, drafting → DeepSeek V4-Flash. Near-free at scale.
  • The very hardest reasoning, or when you're already locked into one vendor → GPT-5.5 or Claude Opus 4.8 still have a small edge, at dozens of times the price.

For most teams the pragmatic answer is a mix: route the bulk to DeepSeek and reserve the closed frontier for the few prompts that truly need it.

Calling them from one place

You can reach DeepSeek V4-Pro and V4-Flash and Claude Opus 4.8 through a single OpenAI-compatible key on Turiloop, pay-as-you-go with an international card, so routing between cheap and frontier is a one-line model change. DeepSeek V4-Flash costs near-zero per token, and V4-Pro stays a fraction of closed-model rates, both discounted right now.