TL;DR — DeepSeek V4 Pro costs ~12× V4 Flash ($1.74/$3.48 vs $0.14/$0.28), with both keeping DeepSeek's signature cheap-output pricing. Every number below is from our weekly-verified pricing tracker (verified 2026-08-26).
Pricing side by side (per 1M tokens, USD)
| Model | Provider | Input | Output | |-------|----------|-------|--------| | DeepSeek V4 Pro | DeepSeek | $1.74 | $3.48 | | DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 |
Verified 2026-08-26 against official rate cards — the tracker keeps the change log.
What a real month costs
At a standard workload of 10M input + 2M output tokens/month:
- DeepSeek V4 Pro: $24.36
- DeepSeek V4 Flash: $1.96
On this mix, DeepSeek V4 Flash is the cheaper column (12.4× gap). Your ratio will differ — run your own mix through the LLM cost calculator.
Choose DeepSeek V4 Pro if
you need the reasoning tier — V4 Pro's $3.48 output is among the cheapest frontier-adjacent output rates tracked, which changes the math for generation-heavy reasoning work.
Choose DeepSeek V4 Flash if
bulk processing at minimum cost: V4 Flash's $0.28 output undercuts even Gemini 3 Flash's $0.30, making it a serious floor-tier option.
Routing decision card
- DeepSeek's family bias is cheap output at both tiers — route output-heavy work here first
- V4 Flash vs Gemini 3 Flash is the real floor fight: Gemini wins input ($0.075 vs $0.14), DeepSeek wins output ($0.28 vs $0.30)
- Two-rung DeepSeek-only routing is viable for cost-obsessed pipelines
- The router answer: most teams shouldn't hard-pick one — ClawRouters routes each request to whichever side fits it, so the comparison becomes a per-call decision instead of a bet.
Part of the LLM model comparison index. Prices move fast in 2026 — the tracker logs every change with dates.