Sovereign inference — Llama 3.1 8b

Your sovereign, private Llama 3.1 8b

A standard model in the small tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.17 UoI/1M input0.51 UoI/1M output131K context

How Llama 3.1 8b compares

Independent intelligence index for Llama 3.1 8b against every peer in the Small tier (and the tier above where needed) — same source, same tests.

Llama 3.1 8b (this model)6.9Qwen3.8 27B33.7Qwen 3.6 27B21.4Gemma 4 26B A4B16.7NVIDIA Nemotron Nano 12B v25.8Gemma 3 12B3.8

Latency — end-to-end seconds for a 500-token answer

Llama 3.1 8b4.5sQwen3.8 27B62.3sQwen 3.6 27B111.0sGemma 4 26B A4B3.6s

Run Llama 3.1 8b privately

One key, one endpoint, three jurisdictions.

  • Size tierSmall
  • Input rate0.17 UoI/1M
  • Output rate0.51 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output131K tokens
  • ReasoningNo
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call