Sovereign inference — GLM 5.2

Your sovereign, private GLM 5.2

A reasoning model in the extra-large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident1.56 UoI/1M input4.68 UoI/1M output0.33 UoI/1M cached input800K context

How GLM 5.2 compares

Independent intelligence index for GLM 5.2 against every peer in the Extra-large tier (and the tier above where needed) — same source, same tests.

GLM 5.2 (this model)33.7GLM 5.344.8GLM 527.9Kimi K2.627.0GLM 5.126.1Kimi K2.7 Code25.8Inkling25.0

Latency — end-to-end seconds for a 500-token answer

GLM 5.240.1sGLM 5.347.3sGLM 554.6sKimi K2.6125.7sGLM 5.1117.2sKimi K2.7 Code44.5sInkling22.5s

Capability indexes

Domain-weighted agentic performance for GLM 5.2.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs extra-large peers

12.825.638.451.233.330.634.332.437.245.6GLM 5.244.747.642.046.747.151.2GLM 5.328.322.029.430.540.5Kimi K2.626.323.328.830.237.6GLM 5.128.928.027.726.529.136.5Kimi K2.7 Code29.523.231.425.427.738.5Inkling
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for GLM 5.2 — Elo scores as published; pass rates shown as percentages.

AutomationBench-AA28.4% %Terminal-Bench 4.01.0% %SciCode51.2% %Humanity's Last Exam41.1% %GDP.pdf10.4% %CritPt20.9% %AA-LCR (long context)78.3% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase1232.9523.3 min115123
GDPval-AA1357.5216.0 min78757
AutomationBench-AA28.4%7.8 min38464
Terminal-Bench 4.01.0%40.0 min197218
SciCode51.2%2.4 min11809
Humanity's Last Exam41.1%8.2 min40587
GDP.pdf10.4%2.3 min11573
CritPt20.9%21.6 min106412
AA-Omniscience4.430.4 min1912
AA-LCR (long context)78.3%0.8 min4061

Run GLM 5.2 privately

One key, one endpoint, three jurisdictions.

  • Size tierExtra-large
  • Input rate1.56 UoI/1M
  • Output rate4.68 UoI/1M
  • Cached input rate0.33 UoI/1M
  • Context window800K tokens
  • Max output16K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call