Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

DeepSeek V4 Pro

Flagship

by DeepSeek

DeepSeek V4 Pro is the large half of the V4 family, released April 24, 2026: a mixture-of-experts with roughly 1.6 trillion total parameters and 49 billion active per token, a 1,048,576 token context window, and MIT-licensed weights on Hugging Face. The reason it is worth a page of its own in September is a retirement that did not happen. When V4.1 Flash shipped on September 10, DeepSeek said that from September 14 every deepseek-v4-pro request would route to V4.1 Flash and bill at Flash rates. On September 11, after user objections, it withdrew that plan, and its changelog now states that API service for V4 Pro continues after September 14, 2026 with the billing method unchanged. The whole cycle ran in five days, and the practical result is that there is no migration to perform: a call naming V4 Pro is still answered by V4 Pro. Pricing splits by the clock the same way the rest of the lineup does, at $1.32 per million cache-miss input tokens and $3.96 output during peak hours, halved off-peak to $0.66 and $1.98, with cache hits at $0.044 and $0.022. Peak is Monday through Friday, 01:00 to 04:00 and 06:00 to 10:00 UTC, excluding Chinese public holidays. That is a long way above the $0.30 and $1.20 V4.1 Flash charges, and DeepSeek's own table has V4.1 Flash ahead on agentic work, so the case for staying on V4 Pro is narrow: it still leads on GPQA Diamond at 92.4 against 90.9 and on Humanity's Last Exam without tools. Self-hosting is a rack proposition, with the published weights at roughly 893GB.

Input Price

$1.32

per 1M tokens

Output Price

$3.96

per 1M tokens

Context Window

1.0M

tokens

Released

2026-04

Open source

Capabilities

texttool-usecodereasoning

Key Strengths

  • ✓Retirement withdrawn on September 11, so no migration is required
  • ✓1.6T total parameters with 49B active per token
  • ✓MIT licensed open weights on Hugging Face
  • ✓1,048,576 token context window
  • ✓Still leads V4.1 Flash on GPQA Diamond, 92.4 to 90.9
  • ✓Off-peak rates halve to $0.66 and $1.98

Best For

  • ▸Workloads already pinned to V4 Pro that no longer need to move
  • ▸Knowledge and science reasoning where GPQA-style accuracy leads
  • ▸Self-hosted deployment under MIT terms at rack scale
  • ▸Batch work scheduled into the off-peak window

Benchmark Scores

BenchmarkScoreDescription
SWE-bench80.6Real-world software engineering tasks from GitHub issues (SWE-bench Verified)
MMLU-Pro91.5General knowledge and reasoning across 57 subjects
HumanEval94.8Python code generation and problem solving
GPQA Diamond92.4Graduate-level science questions verified by domain experts
MATH92.4Competition-level mathematics problems

Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.

Pricing Details

Input tokens

$1.32

per 1M tokens

Output tokens

$3.96

per 1M tokens

Estimated cost per 1K requests

$3.30

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Open Source Model

DeepSeek V4 Pro is free to download and self-host under the MIT. Hosted API pricing varies by provider (e.g., Together, Fireworks, Groq). See our open source LLM guide for deployment options.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide