OpenRating is an independent, open-source AI rating platform. Join the community

OOpenRating
All models

NVIDIA · Released Apr 1, 2026

Nemotron 3.5 Lightning

NVIDIA's latency-optimized open model engineered for high-concurrency microservices, instant autocomplete, and low-latency agents.

Open weights

OpenRating Index

36.6

0–100 composite

Reasoning Index

51.5

0–100 composite

Cost / Task

$0.076

benchmark run

Output speed

285 tok/s

blended tok/s

Time to first token

0.3s

median latency

Input price

$0.04

per 1M tokens

Output price

$0.16

per 1M tokens

Context window

128K

tokens

Strengths

  • 284 tok/s high-speed generation

  • Ultra-low latency (0.31s TTFT)

  • Minimal resource footprint for edge and server deployment

How we rate Nemotron 3.5 Lightning

OpenRating scores every model on the same axes. The OpenRating Index is a weighted composite of reasoning, knowledge, and agentic benchmarks, normalised to a 0–100 scale. Speed is blended output measured in tokens per second under standardised load, with time-to-first-token reported separately. Prices are list input and output prices per million tokens, and Task Cost measures end-to-end workload spending.

OpenRating Index

36.6

Reasoning Index

51.5

Token Pricing

Input per 1M tokens

$0.04

Output per 1M tokens

$0.16

Cost of 1M input + 1M output

$0.2

Task Cost Economics

Empirical cost breakdown for a standard OpenRating Index benchmark task ($0.076 total).

Reasoning (Chain of Thought)
$0.004
Answer generation
$0.003
Cache write overhead
$0.048
Cache hit queries
$0.021
Input payload
$0.001