ToothFairyAI is the sovereign AI and privacy-oriented platform — from bare-bone inference to full agent deployments. Sovereign LLM inference behind OpenAI-compatible endpoints — zero data retention, published per-token rates, and no model-provider lock-in.
Live now: 19 models on AU · 31 on EU · 31 on US — same rates in every region, backed by in-country data centres
ToothFairyAI is a platform-as-a-service, not a one-size-fits-all SaaS product. Engage at whatever level suits your team: call bare-bone inference directly through OpenAI-compatible API endpoints, build with the SDK, CLI and API, integrate through MCP, or start with the ready-to-use desktop, mobile and web apps.
A private model API for developers and CTOs. Open-weight and TF-own models behind OpenAI-compatible regional endpoints, sold per token with publicly listed rates.
The same sovereign compute, wrapped in the agent platform: autonomous agents that plan, reason and perform tasks 24/7 — voice, chat, documents, orchestration.
Point your existing OpenAI SDK at a ToothFairyAI regional host and you are done — choice of open-weight or TF-own models.
Self-serve sign-up, no seats or licences. Enable the intelligence you need across your models from the start.
Point your existing OpenAI-compatible SDK at an Australian, European or US endpoint — your code stays exactly the same.
Choose from the live catalog — reasoning, coding, vision and low-latency options — and scale with pay-per-use billing.
curl https://ais.toothfairyai.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "mystica_15",
"messages": [
{ "role": "user", "content": "Summarise this week's regulatory changes." }
]
}'Interactive per-token maths, straight from the public /models_list feed — pick a model, drag the volume, blend input and output, and compare the total against published flagship rates. The identical catalog and rates are served from every regional endpoint — au, eu and us — so residency never changes your price.
50M tokens/month at a 75/25 input-output blend · GLM 5.3 Flash: $0.42 in / $1.26 out per 1M tokens. Competitor rates are published standard-tier API prices (in/out per 1M tokens): Claude Opus 5 $5/$25 · Claude Fable 5.1 $10/$50 · ChatGPT (chat-latest) $5/$30 · GPT-6 Astra $10/$50 · GPT-5.6 Sol $4/$20 · Gemini Pro $1.25/$10 · Gemini Flash $0.30/$2.50 — ChatGPT Enterprise itself is quoted per seat, not per token. Estimates only; caching discounts and volume agreements are not reflected.
43 active serverless models — every one of them with capabilities, specs and live regional health on the model catalog page.
Compute is the base of the ladder; the platform above it speaks to your existing tooling through standard integrations.
Connect agents to your tools through the Model Context Protocol server.
Python and Node SDKs, CLI agents and full REST endpoints for the whole platform.
Ready-to-use apps for macOS, Windows and Linux, iOS and Android, plus the web studio.
A guided integration builder that generates production-ready code, connectors and MCP tool mappings against your stack from a single brief — powered by TF Code, our sovereign coding agent.
Build it with TF Code →A guided design studio that turns a single brief into editable prototypes, slide decks and infographics on your brand — powered by TF Design, our creative agent in the desktop app.
Design it with TF Design →Two land-and-expand workloads that go live on the same sovereign compute tier your developers already priced.
An autonomous phone agent that answers, plans, reasons and acts — booking, triaging and escalating — 24/7, on your model of choice and inside your region.
A private chat agent embedded on your site that answers from your own documents, never trains on them, and hands off to your team when needed.
Direct answers to the four questions every compliance-minded buyer asks first — scoped honestly, region by region.
Call our AU endpoint: every regional deployment runs smart routing that keeps requests inside the region you call — AU requests are processed in Australia, and the same holds for the EU and US endpoints.
Our own TF models — TF Sorcerer and TF Mystica — are deployed in every region. Availability of the broader third-party catalog per region is expanding; the models table shows which regional endpoints serve each model.
Pay-per-use: per-token rates are published openly per model, per region, in our live model list — no licences, no seats, no minimums. Platform access starts from $5.
Every rate on this page comes from the same /models_list feed our own platform reads, so there is no negotiation arbitrage: what you see quoted is what you are billed per million tokens.
No. The catalog is open-weight and TF-own models, served through standard OpenAI-compatible endpoints, so your code never changes when you swap models.
Move between models with a single string change, bring your own model under Enterprise agreements, and export your stack at any time — model-provider lock-in is not part of the design.
Never. Zero data retention is a platform-wide commitment: your prompts, outputs and customer data are never used to train AI models.
Your data stays encrypted at rest and in transit inside your chosen region, and the platform holds ISO 27001:2022 and ISO 42001 certifications with GDPR and HIPAA compliance.