Sovereign inference — GPT OSS 20B

Your sovereign, private GPT OSS 20B

A reasoning model in the medium tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.43 UoI/1M input1.51 UoI/1M output131K context

How GPT OSS 20B compares

Independent intelligence index for GPT OSS 20B against every peer in the Medium tier (and the tier above where needed) — same source, same tests.

GPT OSS 20B (this model)9.0GLM 5.3 Flash41.8DeepSeek V4.1 Flash39.5DeepSeek V4 Flash Official34.3Minimax 329.2Gemma 4 31B19.0Qwen3 Coder 30B A3B9.6

Latency — end-to-end seconds for a 500-token answer

GPT OSS 20B14.7sGLM 5.3 Flash58.6sDeepSeek V4.1 Flash11.7sDeepSeek V4 Flash Official12.5sMinimax 322.5sGemma 4 31B64.1sQwen3 Coder 30B A3B9.4s

Capability indexes

Domain-weighted agentic performance for GPT OSS 20B.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs medium peers

12.925.938.851.77.34.87.36.611.711.3GPT OSS 20B42.045.139.944.843.848.9GLM 5.3 Flash45.051.741.340.639.344.4DeepSeek V4.1 Flash37.543.137.735.434.940.3DeepSeek V4 Flash Official30.125.031.529.731.243.7Minimax 316.720.6Gemma 4 31B
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for GPT OSS 20B — Elo scores as published; pass rates shown as percentages.

AA-Briefcase0.0% %AutomationBench-AA0.2% %Terminal-Bench 4.00.0% %SciCode38.9% %Humanity's Last Exam11.0% %GDP.pdf2.0% %CritPt1.4% %AA-Omniscience-6305.0% %AA-LCR (long context)34.7% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase0.0%2.2 min25109
GDPval-AA326.463.3 min37012
AutomationBench-AA0.2%0.4 min4861
Terminal-Bench 4.00.0%2.2 min24838
SciCode38.9%0.9 min9747
Humanity's Last Exam11.0%1.4 min16181
GDP.pdf2.0%1.0 min10833
CritPt1.4%3.4 min37604
AA-Omniscience-6305.0%0.2 min2129
AA-LCR (long context)34.7%0.3 min3171

Run GPT OSS 20B privately

One key, one endpoint, three jurisdictions.

  • Size tierMedium
  • Input rate0.43 UoI/1M
  • Output rate1.51 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output66K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call