Sovereign inference — GLM 5.3 Flash

Your sovereign, private GLM 5.3 Flash

A reasoning model in the medium tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.43 UoI/1M input1.51 UoI/1M output1049K context

How GLM 5.3 Flash compares

Independent intelligence index for GLM 5.3 Flash against every peer in the Medium tier (and the tier above where needed) — same source, same tests.

GLM 5.3 Flash (this model)41.8DeepSeek V4.1 Flash39.5DeepSeek V4 Flash Official34.3Minimax 329.2Gemma 4 31B19.0Qwen3 Coder 30B A3B9.6GPT OSS 20B9.0

Latency — end-to-end seconds for a 500-token answer

GLM 5.3 Flash58.6sDeepSeek V4.1 Flash11.7sDeepSeek V4 Flash Official12.5sMinimax 322.5sGemma 4 31B64.1sQwen3 Coder 30B A3B9.4sGPT OSS 20B14.7s

Capability indexes

Domain-weighted agentic performance for GLM 5.3 Flash.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs medium peers

12.925.938.851.742.045.139.944.843.848.9GLM 5.3 Flash45.051.741.340.639.344.4DeepSeek V4.1 Flash37.543.137.735.434.940.3DeepSeek V4 Flash Official30.125.031.529.731.243.7Minimax 316.720.6Gemma 4 31B7.34.87.36.611.711.3GPT OSS 20B
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for GLM 5.3 Flash — Elo scores as published; pass rates shown as percentages.

AutomationBench-AA60.4% %Terminal-Bench 4.032.8% %SciCode51.6% %Humanity's Last Exam39.9% %GDP.pdf15.4% %CritPt15.4% %AA-LCR (long context)80.0% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase1459.4138.4 min131965
GDPval-AA1640.7836.7 min125884
AutomationBench-AA60.4%5.9 min20219
Terminal-Bench 4.032.8%52.7 min181054
SciCode51.6%7.9 min27041
Humanity's Last Exam39.9%12.6 min43214
GDP.pdf15.4%4.6 min15701
CritPt15.4%23.8 min81785
AA-Omniscience7.470.3 min1159
AA-LCR (long context)80.0%1.3 min4515

Run GLM 5.3 Flash privately

One key, one endpoint, three jurisdictions.

  • Size tierMedium
  • Input rate0.43 UoI/1M
  • Output rate1.51 UoI/1M
  • Cached input rate—
  • Context window1049K tokens
  • Max output16K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call