NVIDIA · Released Feb 15, 2026
Nemotron 3 Super
NVIDIA optimized Mixture-of-Experts model tuned with TensorRT-LLM for maximum hardware utilization on Hopper/Blackwell GPUs.
OpenRating Index
43.6
0–100 composite
Reasoning Index
60.4
0–100 composite
Cost / Task
$0.23
benchmark run
Output speed
124 tok/s
blended tok/s
Time to first token
0.5s
median latency
Input price
$0.12
per 1M tokens
Output price
$0.48
per 1M tokens
Context window
128K
tokens
Strengths
Optimized specifically for TensorRT-LLM and Triton servers
Active-expert MoE architecture with low activation footprint
High reliability for enterprise automation workflows
How we rate Nemotron 3 Super
OpenRating scores every model on the same axes. The OpenRating Index is a weighted composite of reasoning, knowledge, and agentic benchmarks, normalised to a 0–100 scale. Speed is blended output measured in tokens per second under standardised load, with time-to-first-token reported separately. Prices are list input and output prices per million tokens, and Task Cost measures end-to-end workload spending.
OpenRating Index
43.6
Reasoning Index
60.4
Token Pricing
Input per 1M tokens
$0.12
Output per 1M tokens
$0.48
Cost of 1M input + 1M output
$0.6
Task Cost Economics
Empirical cost breakdown for a standard OpenRating Index benchmark task ($0.23 total).
Related models
Compare Nemotron 3 Super with the models it is most often measured against.