Experiential
ModelsLogs
Star us on GitHubDocsSettings
Sign in
Continue with GoogleContinue with GitHub
or
Models

Nemotron 3 Ultra

nemotron-3-ultra-550b-a55bby NVIDIAZDR enforceableFree

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

CompareOpen in Playground

Context

512K

Max output

no data

Pricing

$0.60$0/M in$0.12$0/M cached$2.40$0/M out

Fastest

1036 tok/s

Released

Jun 2026

Promotion

FreeYour limits

Waterfall

  1. Azure Foundry via Experiential Cloudfree tierFW-Nemotron-3-Ultra-NVFP4·$0.66 /M in·$2.64 /M out$0 /M in·$0 /M out96.6% up297 tok/s1.41s
  2. 1Azure Foundry via Experiential Clouduses exp creditsFW-Nemotron-3-Ultra-NVFP4·$0.66 /M in·$2.64 /M out96.6% up297 tok/s1.41sUse
  3. 2Fireworks via Experiential Clouduses exp creditsaccounts/fireworks/models/nemotron-3-ultra-nvfp4·$0.60 /M in·$2.40 /M out98.8% up126 tok/s417msUse

Supported parameters

primary route

Sampling

Temperature0-2Top-p0-1Stop sequences

Tools & structure

ToolsParallel tool calls*Structured output*

Reasoning

Reasoningnone · low · medium · high · xhigh · maxDefaultmedium

Streaming

Streaming

Limits

Max-tokens fieldmax_tokens

* not supported on every route; the gateway adapts or drops it (with a warning) on stricter providers.

Shown for the primary route; fallback routes may differ. Sending an unsupported field? See error reference.

Azure OpenAI docs

Quickstart

I want you to route my LLM calls for "nemotron-3-ultra-550b-a55b" through the Experiential gateway instead of
calling the provider directly. It speaks the OpenAI Chat Completions API, so this is a base-URL
and key swap. Please:

1. Point the client at https://api.experientiallabs.ai/v1 as the base URL.
2. Authenticate with my Experiential API key from the EXPLABS_API_KEY environment variable. If
   it isn't set, stop and tell me to create one under Settings -> API Keys and export it.
3. Use the model id "nemotron-3-ultra-550b-a55b" exactly.
4. Update every place my code builds an LLM client for this model to use that base URL and key,
   leaving streaming and tool-calls as they are.
5. Make one test call and show me the reply plus the token usage, so we confirm it runs on my
   Experiential credits.

Tell me which files you changed.

Set EXPLABS_API_KEY to an organization API key before running your agent.

Benchmarks

Hugging Face
  • MMLU-ProHugging Face · Aug 202686.8%
  • GPQA DiamondHugging Face · Aug 202687.9%
  • SWE-bench VerifiedHugging Face · Aug 202669.5%
  • Humanity's Last ExamHugging Face · Aug 202626.1%
Status
Azure Foundry via Experiential Cloudfree tierFW-Nemotron-3-Ultra-NVFP4297 tok/s1.41s96.6%512Kno dataFreeFreeFreenvfp4no dataactivemeasured
Azure Foundry via Experiential CloudexperientialFW-Nemotron-3-Ultra-NVFP4297 tok/s1.41s96.6%512Kno data$0.66$2.64cache$0.13nvfp4no dataactivemeasured
Fireworks via Experiential Cloudexperientialaccounts/fireworks/models/nemotron-3-ultra-nvfp4126 tok/s417ms98.8%512Kno data$0.60$2.40cache$0.12nvfp4ZDRactivemeasured
Azure FoundryNemotron-3-Ultra-550B-A55B-BF16297 tok/s1.41s96.6%512Kno datano datano datano databf16no dataactivemeasured
Fireworksaccounts/fireworks/models/nemotron-3-ultra-bf16126 tok/s417ms98.8%512Kno data$0.60≈$3.60≈cache$0.20≈bf16ZDRactivemeasured
  • LMArena EloLMArena · Aug 20261445
  • Terminal-BenchHugging Face · Aug 202653.9