Sovereign inference — Kimi K2.6

Your sovereign, private Kimi K2.6

A reasoning model in the extra-large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident1.56 UoI/1M input4.68 UoI/1M output262K context

How Kimi K2.6 compares

Independent intelligence index for Kimi K2.6 against every peer in the Extra-large tier (and the tier above where needed) — same source, same tests.

Kimi K2.6 (this model)27.0GLM 5.344.8GLM 5.233.7GLM 527.9GLM 5.126.1Kimi K2.7 Code25.8Inkling25.0

Latency — end-to-end seconds for a 500-token answer

Kimi K2.6125.7sGLM 5.347.3sGLM 5.240.1sGLM 554.6sGLM 5.1117.2sKimi K2.7 Code44.5sInkling22.5s

Capability indexes

Domain-weighted agentic performance for Kimi K2.6.

Finance & AccountingStrategy & OpsLegalEngineeringEconomics

Capability index vs extra-large peers

12.825.638.451.228.322.029.430.540.5Kimi K2.644.747.642.046.747.151.2GLM 5.333.330.634.332.437.245.6GLM 5.226.323.328.830.237.6GLM 5.128.928.027.726.529.136.5Kimi K2.7 Code29.523.231.425.427.738.5Inkling
Finance & AccountingStrategy & OpsLegalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for Kimi K2.6 — Elo scores as published; pass rates shown as percentages.

AutomationBench-AA13.0% %Terminal-Bench 4.00.5% %SciCode51.5% %Humanity's Last Exam37.5% %GDP.pdf13.0% %CritPt8.0% %AA-LCR (long context)81.0% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase814.7315.6 min42409
GDPval-AA1026.099.3 min25296
AutomationBench-AA13.0%8.5 min23221
Terminal-Bench 4.00.5%41.5 min112817
SciCode51.5%3.8 min10301
Humanity's Last Exam37.5%15.8 min42968
GDP.pdf13.0%4.7 min12836
CritPt8.0%67.3 min182820
AA-Omniscience5.302.4 min6397
AA-LCR (long context)81.0%1.1 min3028

Run Kimi K2.6 privately

One key, one endpoint, three jurisdictions.

  • Size tierExtra-large
  • Input rate1.56 UoI/1M
  • Output rate4.68 UoI/1M
  • Cached input rate—
  • Context window262K tokens
  • Max output262K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call