Sovereign inference — Kimi K3

Your sovereign, private Kimi K3

A reasoning model in the ultra tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident4.20 UoI/1M input21.00 UoI/1M output0.42 UoI/1M cached input800K context

How Kimi K3 compares

Independent intelligence index for Kimi K3 against every peer in the Ultra tier (and the tier above where needed) — same source, same tests.

Kimi K3 (this model)43.6Qwen3.8 Max45.4GLM 5.3 · Extra-large44.8GLM 5.2 · Extra-large33.7GLM 5 · Extra-large27.9Kimi K2.6 · Extra-large27.0GLM 5.1 · Extra-large26.1

Throughput — tokens per second, median

Kimi K337.3 t/s

Latency — end-to-end seconds for a 500-token answer

Kimi K371.2sQwen3.8 Max66.7sGLM 5.347.3sGLM 5.240.1sGLM 554.6sKimi K2.6125.7sGLM 5.1117.2s

Capability indexes

Domain-weighted agentic performance for Kimi K3.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs ultra peers

13.326.740.053.446.950.548.645.043.353.4Kimi K345.647.346.141.446.452.3Qwen3.8 Max44.747.642.046.747.151.2GLM 5.333.330.634.332.437.245.6GLM 5.228.322.029.430.540.5Kimi K2.626.323.328.830.237.6GLM 5.1
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for Kimi K3 — Elo scores as published; pass rates shown as percentages.

AutomationBench-AA58.3% %Terminal-Bench 4.012.6% %SciCode59.5% %Humanity's Last Exam46.9% %GDP.pdf22.0% %CritPt23.4% %AA-LCR (long context)88.7% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase1504.0449.5 min119873
GDPval-AA1523.9724.5 min59382
AutomationBench-AA58.3%8.2 min19920
Terminal-Bench 4.012.6%50.8 min122899
SciCode59.5%1.5 min3520
Humanity's Last Exam46.9%10.8 min26162
GDP.pdf22.0%4.0 min9597
CritPt23.4%24.5 min59238
AA-Omniscience19.703.6 min8813
AA-LCR (long context)88.7%0.6 min1529

Run Kimi K3 privately

One key, one endpoint, three jurisdictions.

  • Size tierUltra
  • Input rate4.20 UoI/1M
  • Output rate21.00 UoI/1M
  • Cached input rate0.42 UoI/1M
  • Context window800K tokens
  • Max output120K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call