Stop a Runaway OpenAI or Anthropic Bill: Set a Daily Budget Alert in 5 Minutes
An agent loop, a retry storm or a viral feature can turn a $5 day into a $500 day before the provider dashboard updates. How to track AI spend per call and get alerted at 50%, 80% and 100% of a daily budget.
The bill that hurts is rarely the one you planned for. A retry with no backoff. An agent loop that re-sends its whole growing conversation on every step. A feature that gets shared and ten times the users show up. By the time the provider's usage page catches up, the money is spent.
What you want is not a better invoice. It is an alert while it is happening.
Why the dashboards are too slow
Provider usage pages and billing data lag behind real traffic, and some budget limits are enforced asynchronously, so spend can pass the number you set before anything stops it. Treat the platform limit as a backstop, not your only defence. Real protection has two parts:
- Know what each call costs, as it happens.
- Get told when today's total crosses a line you chose.
Step 1: record the cost of every call
FlareLog's SDK instruments AI calls with one flag. It reads the token usage that every response returns, prices it from a table, and logs the call with its cost:
import { flarelog } from "@flarelog/sdk";
const logger = flarelog({
apiKey: process.env.FLARELOG_API_KEY,
ai: true, // every fetch() to OpenAI, Anthropic and compatible APIs is captured
});
// Use your AI SDK as normal. Nothing else changes.Each call becomes a log entry with the model, tokens, latency and an estimated cost in USD. It works with OpenAI, Anthropic, Cloudflare Workers AI, the Vercel AI SDK and any OpenAI-compatible gateway.
If your negotiated price differs from the public one, override it:
const logger = flarelog({
apiKey: process.env.FLARELOG_API_KEY,
ai: { priceOverrides: { "gpt-4o": { input: 2.5, output: 10 } } },
});Step 2: set a daily budget
In your project: Settings → Cost Burn Alerts → Daily Budget (USD), for example 20.
From then on FlareLog adds up the day's spend (AI calls plus, if you use the Cloudflare Tail Worker, your Worker cost estimates) and alerts you when the total crosses:
- 50% of the budget
- 80% of the budget
- 100% of the budget
Each level fires once per UTC day, so you get three clear signals, not a flood. If a single burst jumps from 40% straight past 100%, you get one alert at 100%, not three at once. The alert goes out by email and, if you have configured a webhook, to Slack or Discord. Leave the field blank to turn the budget off.
The budget alert is part of the Pro plan's cost burn alerts. Spend is still recorded on the free plan, so upgrading later does not lose your history.
What an alert gives you
Daily budget 80% reached: $16.12 of $20.00 spent on 2026-10-06 (UTC). $15.40 of that is AI calls.
That one line answers "how bad is it" and "is it the AI". From there, the MCP server lets your assistant ask which model and which route is burning the money (get_ai_cost_trend, get_ai_calls), so you are not scrolling a dashboard at midnight.
Be honest about what this is
- It is an estimate. Costs come from the SDK's price table and the token counts in each response. A model the table does not know has no price, and its calls will not count towards the total. Check your models and add
priceOverrideswhere needed. The provider's invoice remains the source of truth. - It alerts; it does not stop the spend. A budget alert tells you within moments. To actually cap spending, also use your provider's project budgets and limits and a hard cap in your own code.
Guardrails worth adding in code
The alert is your smoke detector. These are the sprinklers:
- Cap agent steps. Set a maximum number of iterations. Agent frameworks resend the whole growing conversation each step, so a 40-step loop bills the early context dozens of times.
- Retry with backoff and a limit. Unbounded retries are the classic runaway.
- Trim context. Keep the last N turns, or enforce a maximum context size.
- Pin your model. A framework default such as "latest" can change what you pay without any change on your side.
- Use a separate key per workflow. If one goes wild you can revoke it without taking everything down.
Set it up
- Create a free project and add the SDK with
ai: true. - Make a few AI calls and check they appear in your AI dashboard.
- Upgrade to Pro and set a daily budget.
- Add a Slack or Discord webhook so the alert reaches you where you will see it.
Five minutes now, versus finding out from the invoice.
Never miss an invisible crash again
FlareLog catches the errors Cloudflare can't log. Set up the Tail Worker in 5 minutes and see every crash, timeout, and cost spike in real time.
Start free →