Sovereign inference — GLM 5.1

Your sovereign, private GLM 5.1

A reasoning model in the extra-large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident1.56 UoI/1M input4.68 UoI/1M output203K context

How GLM 5.1 compares

Independent intelligence index for GLM 5.1 against every peer in the Extra-large tier (and the tier above where needed) — same source, same tests.

GLM 5.1 (this model)26.1GLM 5.344.8GLM 5.233.7GLM 527.9Kimi K2.627.0Kimi K2.7 Code25.8Inkling25.0

Latency — end-to-end seconds for a 500-token answer

GLM 5.1117.2sGLM 5.347.3sGLM 5.240.1sGLM 554.6sKimi K2.6125.7sKimi K2.7 Code44.5sInkling22.5s

Capability indexes

Domain-weighted agentic performance for GLM 5.1.

Finance & AccountingStrategy & OpsLegalEngineeringEconomics

Capability index vs extra-large peers

12.825.638.451.226.323.328.830.237.6GLM 5.144.747.642.046.747.151.2GLM 5.333.330.634.332.437.245.6GLM 5.228.322.029.430.540.5Kimi K2.628.928.027.726.529.136.5Kimi K2.7 Code29.523.231.425.427.738.5Inkling
Finance & AccountingStrategy & OpsLegalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for GLM 5.1 — Elo scores as published; pass rates shown as percentages.

AutomationBench-AA20.3% %Terminal-Bench 4.02.0% %SciCode44.8% %Humanity's Last Exam30.1% %GDP.pdf8.4% %CritPt4.6% %AA-Omniscience85.0% %AA-LCR (long context)73.7% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase963.9217.6 min42816
GDPval-AA1103.3612.7 min30858
AutomationBench-AA20.3%10.1 min24585
Terminal-Bench 4.02.0%64.7 min156878
SciCode44.8%4.2 min10086
Humanity's Last Exam30.1%15.6 min37819
GDP.pdf8.4%2.7 min6554
CritPt4.6%36.3 min88173
AA-Omniscience85.0%0.9 min2295
AA-LCR (long context)73.7%2.2 min5243

Run GLM 5.1 privately

One key, one endpoint, three jurisdictions.

  • Size tierExtra-large
  • Input rate1.56 UoI/1M
  • Output rate4.68 UoI/1M
  • Cached input rate—
  • Context window203K tokens
  • Max output16K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call