Scale up as you grow — whether you're running one virtual machine or ten thousand.

From GPU-powered inference and Kubernetes to managed databases and storage, get everything you need to build, scale, and deploy intelligent applications.

Trying to figure out the daily token budget for Tier 2 serverless inference (inference.do-ai.run) so to find out if it can be used for agentic coding work (bursts requests). I can’t find a consistent answer and would like to reach out in the community to see if anybody could help me figure this out.
Have done some requests and recorded the response (probing with “messages”: [] returns HTTP 400 and does not consume tokens, zero-cost way to pull headers):
From this I think I can conclude that the bucket behaves like capacity: 500,000 is the tier 2 TPM range and max refill rate is 5,000,000/day and the tokens stop filling up the bucket after ~2.4h, when 500,000 tokens is reached.
Reading the docs I can’t find clarity:
With a 500,000 ceiling I can’t see how Tier 2 could be used for agentic coding. DO’s OpenCode tutorial (inference-router-agentic-coding) uses ~4.1M tokens over a few hours. At 57.87 tok/s, ~20h would be needed to run through it on my tier.
So, is 500,000 the intended Tier 2 daily capacity, and where is it documented?
Ehsan Zilaei
trru
d7fd4a488ff846d6bb18c276373438