Sovereign inference — GPT OSS 120B

Your sovereign, private GPT OSS 120B

A reasoning model in the large tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.84 UoI/1M input2.10 UoI/1M output131K context

How GPT OSS 120B compares

Independent intelligence index for GPT OSS 120B against every peer in the Large tier (and the tier above where needed) — same source, same tests.

GPT OSS 120B (this model)11.6Kimi K2.523.5MiniMax 2.522.8MiniMax 2.722.8Qwen3 Coder Next9.2

Latency — end-to-end seconds for a 500-token answer

GPT OSS 120B14.3sKimi K2.531.0sMiniMax 2.550.5sMiniMax 2.76.8s

Capability indexes

Domain-weighted agentic performance for GPT OSS 120B.

Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

Capability index vs large peers

9.017.926.935.912.27.911.310.715.118.3GPT OSS 120B23.516.326.921.125.635.9MiniMax 2.77.75.58.39.212.0Qwen3 Coder Next
Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics

The evaluations behind the index

Independent benchmark results for GPT OSS 120B — Elo scores as published; pass rates shown as percentages.

AA-Briefcase0.0% %AutomationBench-AA0.2% %Terminal-Bench 4.00.0% %SciCode34.0% %Humanity's Last Exam19.6% %GDP.pdf4.0% %CritPt1.1% %AA-Omniscience-4925.0% %AA-LCR (long context)52.0% %

Per-evaluation economics

EvaluationScoreTime / taskOutput tokens / task
AA-Briefcase0.0%8.4 min97717
GDPval-AA595.502.7 min31411
AutomationBench-AA0.2%0.2 min2183
Terminal-Bench 4.00.0%1.8 min20583
SciCode34.0%0.5 min5707
Humanity's Last Exam19.6%2.0 min22890
GDP.pdf4.0%0.7 min7681
CritPt1.1%2.5 min29812
AA-Omniscience-4925.0%0.2 min1916
AA-LCR (long context)52.0%0.2 min1831

Run GPT OSS 120B privately

One key, one endpoint, three jurisdictions.

  • Size tierLarge
  • Input rate0.84 UoI/1M
  • Output rate2.10 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output66K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputText only
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call