Sovereign inference — Llama 4 Maverick

Your sovereign, private Llama 4 Maverick

A standard model in the extra-large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident1.56 UoI/1M input4.68 UoI/1M output131K context

How Llama 4 Maverick compares

Independent intelligence index for Llama 4 Maverick against every peer in the Extra-large tier (and the tier above where needed) — same source, same tests.

Llama 4 Maverick (this model)10.0GLM 5.344.8GLM 5.233.7GLM 527.9Kimi K2.627.0GLM 5.126.1Kimi K2.7 Code25.8

Latency — end-to-end seconds for a 500-token answer

Llama 4 Maverick5.3sGLM 5.347.3sGLM 5.240.1sGLM 554.6sKimi K2.6125.7sGLM 5.1117.2sKimi K2.7 Code44.5s

Capability indexes

Domain-weighted agentic performance for Llama 4 Maverick.

Healthcare & MedicalEngineeringEconomics

Capability index vs extra-large peers

12.825.638.451.29.111.313.0Llama 4 Maverick44.747.642.046.747.151.2GLM 5.333.330.634.332.437.245.6GLM 5.228.322.029.430.540.5Kimi K2.626.323.328.830.237.6GLM 5.128.928.027.726.529.136.5Kimi K2.7 Code
Healthcare & MedicalEngineeringEconomics

Run Llama 4 Maverick privately

One key, one endpoint, three jurisdictions.

  • Size tierExtra-large
  • Input rate1.56 UoI/1M
  • Output rate4.68 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output4K tokens
  • ReasoningNo
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call