The AI Inference platform

Workers AI lets you run AI inference globally with one API call. No GPUs to manage, no capacity planning. Just intelligent machine learning models running where they're needed, on Cloudflare's global network.

Serverless pricing

Pay-per-inference pricing with no idle costs. No guessing what.

Rich model catalog

50+ models running close to users in 200+ cities

Widely compatible

One API call, works with any OpenAI SDK or task type

Scale up, and down

Inference is hard to predict and spiky in nature, unlike training. GPU utilization is, on average, only 20-40% — with one-third of organizations utilizing less than 15%. Workers AI allows customers to save by only paying for usage. No guessing or committing to hardware that goes unused.