Sovereign inference — NVIDIA Nemotron Nano 12B v2

Your sovereign, private NVIDIA Nemotron Nano 12B v2

A reasoning model in the small tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.17 UoI/1M input0.51 UoI/1M output131K context

How NVIDIA Nemotron Nano 12B v2 compares

Independent intelligence index for NVIDIA Nemotron Nano 12B v2 against every peer in the Small tier (and the tier above where needed) — same source, same tests.

NVIDIA Nemotron Nano 12B v2 (this model)5.8Qwen3.8 27B33.7Qwen 3.6 27B21.4Gemma 4 26B A4B16.7Llama 3.1 8b6.9Gemma 3 12B3.8

Latency — end-to-end seconds for a 500-token answer

NVIDIA Nemotron Nano 12B v23.6sQwen3.8 27B62.3sQwen 3.6 27B111.0sGemma 4 26B A4B4.5s

Run NVIDIA Nemotron Nano 12B v2 privately

One key, one endpoint, three jurisdictions.

  • Size tierSmall
  • Input rate0.17 UoI/1M
  • Output rate0.51 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output8K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call