OpenRating is an independent, open-source AI rating platform. Join the community

OOpenRating
All models

NVIDIA · Released Mar 5, 2026

Nemotron 3 Ultra

NVIDIA's 550B parameter flagship open architecture designed for high-concurrency enterprise inference clusters.

Open weights

OpenRating Index

38.6

0–100 composite

Reasoning Index

56.8

0–100 composite

Cost / Task

$0.38

benchmark run

Output speed

141 tok/s

blended tok/s

Time to first token

0.4s

median latency

Input price

$0.2

per 1M tokens

Output price

$0.8

per 1M tokens

Context window

256K

tokens

Strengths

  • 141 tok/s high-throughput generation on GPU clusters

  • Open weights for internal enterprise deployment

  • Extensive synthetic data curation and alignment

How we rate Nemotron 3 Ultra

OpenRating scores every model on the same axes. The OpenRating Index is a weighted composite of reasoning, knowledge, and agentic benchmarks, normalised to a 0–100 scale. Speed is blended output measured in tokens per second under standardised load, with time-to-first-token reported separately. Prices are list input and output prices per million tokens, and Task Cost measures end-to-end workload spending.

OpenRating Index

38.6

Reasoning Index

56.8

Token Pricing

Input per 1M tokens

$0.2

Output per 1M tokens

$0.8

Cost of 1M input + 1M output

$1

Task Cost Economics

Empirical cost breakdown for a standard OpenRating Index benchmark task ($0.38 total).

Reasoning (Chain of Thought)
$0.041
Answer generation
$0.024
Cache write overhead
$0.31
Cache hit queries
$0.00
Input payload
$0.004