Sovereign inference — NVIDIA Nemotron Nano 12B v2
Your sovereign, private NVIDIA Nemotron Nano 12B v2
A reasoning model in the small tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.
AU · EU · US resident0.17 UoI/1M input0.51 UoI/1M output131K context
How NVIDIA Nemotron Nano 12B v2 compares
Independent intelligence index for NVIDIA Nemotron Nano 12B v2 against every peer in the Small tier (and the tier above where needed) — same source, same tests.
Latency — end-to-end seconds for a 500-token answer
Run NVIDIA Nemotron Nano 12B v2 privately
One key, one endpoint, three jurisdictions.
- Size tierSmall
- Input rate0.17 UoI/1M
- Output rate0.51 UoI/1M
- Cached input rate—
- Context window131K tokens
- Max output8K tokens
- ReasoningYes — emits thinking traces
- Vision & video inputSupports image input
- Tool callingSupported
- Data residencyAU, EU and US endpoints — processing stays inside the region you call


