Sovereign inference — NVIDIA Nemotron Super 3 120B

Your sovereign, private NVIDIA Nemotron Super 3 120B

A reasoning model in the extra-large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident1.56 UoI/1M input4.68 UoI/1M output164K context

How NVIDIA Nemotron Super 3 120B compares

Independent intelligence index for NVIDIA Nemotron Super 3 120B against every peer in the Extra-large tier (and the tier above where needed) — same source, same tests.

NVIDIA Nemotron Super 3 120B (this model)12.8GLM 5.344.8GLM 5.233.7GLM 527.9Kimi K2.627.0GLM 5.126.1Kimi K2.7 Code25.8

Latency — end-to-end seconds for a 500-token answer

NVIDIA Nemotron Super 3 120B16.9sGLM 5.347.3sGLM 5.240.1sGLM 554.6sKimi K2.6125.7sGLM 5.1117.2sKimi K2.7 Code44.5s

Capability indexes

Domain-weighted agentic performance for NVIDIA Nemotron Super 3 120B.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs extra-large peers

12.825.638.451.213.29.112.810.817.020.6NVIDIA Nemotron Super 3 120B44.747.642.046.747.151.2GLM 5.333.330.634.332.437.245.6GLM 5.228.322.029.430.540.5Kimi K2.626.323.328.830.237.6GLM 5.128.928.027.726.529.136.5Kimi K2.7 Code
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for NVIDIA Nemotron Super 3 120B — Elo scores as published; pass rates shown as percentages.

AA-Briefcase0.0% %AutomationBench-AA3.8% %Terminal-Bench 4.00.0% %SciCode36.2% %Humanity's Last Exam20.8% %GDP.pdf2.6% %CritPt3.1% %AA-Omniscience-4150.0% %AA-LCR (long context)65.7% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase0.0%56.9 min604419
GDPval-AA477.332.3 min24750
AutomationBench-AA3.8%1.2 min12246
Terminal-Bench 4.00.0%8.3 min87694
SciCode36.2%0.2 min1915
Humanity's Last Exam20.8%3.4 min36199
GDP.pdf2.6%0.7 min7700
CritPt3.1%5.6 min59305
AA-Omniscience-4150.0%0.2 min2274
AA-LCR (long context)65.7%0.8 min8998

Run NVIDIA Nemotron Super 3 120B privately

One key, one endpoint, three jurisdictions.

  • Size tierExtra-large
  • Input rate1.56 UoI/1M
  • Output rate4.68 UoI/1M
  • Cached input rate—
  • Context window164K tokens
  • Max output16K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call