Experiential
ModelsLogs
Star us on GitHubDocsSettings
Sign in
Continue with GoogleContinue with GitHub
or
Models

DeepSeek V4 Flash

deepseek-v4-flashby DeepSeekrecommendedZDR enforceable
Try in Playground

Context

1.05M

Max output

384K

Price per million tokens

$0.060in$0.018cached$0.12out

Fastest

176 tok/s

Released

Apr 2026

Waterfall

  1. 1
    Azure Foundry via Experiential Cloud$0.19/$0.5199.5% up176 tok/s1.58s
  2. 2H
    Huawei Cloud via Experiential Cloud$0.14/$0.27100% up66.4 tok/s1.24s
  3. 3
    Experiential Cloud$0.19/$0.51100% up168 tok/s1.51s

Quickstart

I want you to route my LLM calls for "deepseek-v4-flash" through the Experiential gateway instead of calling the provider directly. It speaks the OpenAI Chat Completions API, so this is a base-URL and key swap. Please:

1. Point the client at https://api.experientiallabs.ai/v1 as the base URL.
2. Authenticate with my Experiential API key from the EXPLABS_API_KEY environment variable. If it isn't set, stop and tell me to create one under Settings -> API Keys and export it.
3. Use the model id "deepseek-v4-flash" exactly.
4. Update every place my code builds an LLM client for this model to use that base URL and key, leaving streaming and tool-calls as they are.
5. Make one test call and show me the reply plus the token usage, so we confirm it runs on my Experiential credits.

Tell me which files you changed.
CapabilitiesAccepts text

Sampling

Temperature0-2Top-p0.01-1Stop sequences*

Tools & structure

ToolsParallel tool calls*Structured output*

Reasoning

Reasoningnone · low · medium · high · xhigh · maxDefaultnone

Streaming

Streaming

Limits

Max-tokens fieldmax_tokens

Sending an unsupported field? See error referenceAzure OpenAI docs

Benchmarks7
BenchmarkScoreSource
MMLU-ProHugging Face · Aug 202686.4%Hugging Face · Aug 2026
GPQA DiamondHugging Face · Aug 202688.1%Hugging Face · Aug 2026
SWE-bench VerifiedHugging Face · Aug 202679.0%Hugging Face · Aug 2026
LiveBenchpublic leaderboard · Aug 202665.5%public leaderboard · Aug 2026
AIME 2026public leaderboard · Aug 202695.8%public leaderboard · Aug 2026
Humanity's Last ExamHugging Face · Aug 202634.8%Hugging Face · Aug 2026
Terminal-BenchHugging Face · Aug 202656.9Hugging Face · Aug 2026
Providers serving this model9
Status
Experiential Clouddeepseek-v4-flash27.6 tok/s385ms99.7%1.05M384K$0.042$0.085cache$0.008-ZDRdisabledmeasured
Experiential Cloudexperientialdeepseek/deepseek-v4-flash168 tok/s1.51s100%1.05M384K$0.19Free 25% · Pro 50% · Max 75% · Ultra 100% off$0.51cache$0.028-ZDRactivemeasured
Azure Foundry via Experiential CloudexperientialDeepSeek-V4-Flash176 tok/s1.58s99.5%1.05M384K$0.19Free 25% · Pro 50% · Max 75% · Ultra 100% off$0.51cache$0.028-no dataactivemeasured
HHuawei Cloud via Experiential Cloudexperiential66.4 tok/s1.24s100%1.05M384K$0.14Free 25% · Pro 50% · Max 75% · Ultra 100% off$0.27$0.14-no dataactivemeasured
Azure FoundryFW-DeepSeek-V4-Flash176 tok/s1.58s99.5%1.05M384K$0.089≈$0.18≈cache$0.018≈-no dataactivemeasured
Azure FoundryDeepSeek-V4-Flash-2026-04-23deployed 2026-04-23176 tok/s1.58s99.5%1.05M384K$0.089≈$0.18≈cache$0.018≈-no dataactivemeasured
TTencent Cloud via Experiential Cloud89 tok/s5.64s0%1.05M384K$0.089$0.18cache$0.018-no datadisabledmeasured
Fireworksaccounts/fireworks/models/deepseek-v4-flashno datano data99.9%1.05M384K$0.089≈$0.18≈cache$0.018≈-ZDRactiveestimated
Fireworks via Experiential Cloudaccounts/fireworks/models/deepseek-v4-flash-0731no datano data99.9%1.05M384K$0.22$0.66cache$0.007-ZDRdisabled-
Hugging Face