Sovereign inference — Gemma 4 31B

Your sovereign, private Gemma 4 31B

A reasoning model in the medium tier, served on ToothFairyAI's AU, EU and US endpoints — private by default, billed in Units of Intelligence at the published rates below. No training on your data, no vendor lock-in.

AU · EU · US resident0.43 UoI/1M input1.51 UoI/1M output131K context

How Gemma 4 31B compares

Independent intelligence index for Gemma 4 31B against every peer in the Medium tier (and the tier above where needed) — same source, same tests.

Gemma 4 31B (this model)19.0GLM 5.3 Flash41.8DeepSeek V4.1 Flash39.5DeepSeek V4 Flash Official34.3Minimax 329.2Qwen3 Coder 30B A3B9.6GPT OSS 20B9.0

Latency — end-to-end seconds for a 500-token answer

Gemma 4 31B64.1sGLM 5.3 Flash58.6sDeepSeek V4.1 Flash11.7sDeepSeek V4 Flash Official12.5sMinimax 322.5sQwen3 Coder 30B A3B9.4sGPT OSS 20B14.7s

Capability indexes

Domain-weighted agentic performance for Gemma 4 31B.

EngineeringEconomics

Capability index vs medium peers

12.925.938.851.716.720.6Gemma 4 31B42.045.139.944.843.848.9GLM 5.3 Flash45.051.741.340.639.344.4DeepSeek V4.1 Flash37.543.137.735.434.940.3DeepSeek V4 Flash Official30.125.031.529.731.243.7Minimax 37.34.87.36.611.711.3GPT OSS 20B
EngineeringEconomics

Run Gemma 4 31B privately

One key, one endpoint, three jurisdictions.

  • Size tierMedium
  • Input rate0.43 UoI/1M
  • Output rate1.51 UoI/1M
  • Cached input rate—
  • Context window131K tokens
  • Max output66K tokens
  • ReasoningYes — emits thinking traces
  • Vision & video inputSupports image input
  • Tool callingSupported
  • Data residencyAU, EU and US endpoints — processing stays inside the region you call