> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kirafin.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

Kira limits how fast you can call the API. The limit is **per account**, set when your account is provisioned, and it applies across every endpoint rather than per endpoint.

## What your account has

Two limits work together — a throughput limit, which is a rate with a burst allowance on top of it, and a quota:

| Limit          | What it controls                                                                                             |
| -------------- | ------------------------------------------------------------------------------------------------------------ |
| **Throughput** | How many requests per second you can sustain, plus the burst that may arrive at once before the rate applies |
| **A quota**    | How many you can make in total over a period                                                                 |

Most accounts are provisioned at **20 requests per second with a burst of 50**. Quotas vary more, so treat the rate as the number to design against and check your own quota with your Kira contact.

## When you hit it

You get a **`429`**.

<Warning>
  **Read the status code, not the message.** A `429` is a rate limit — the body it comes with is a generic one, so an integration that branches on the message rather than the code will look for a problem that is not there.
</Warning>

Both limits produce the same status: the per-second one when you are going too fast right now, and the quota when you have used your allowance for the period. What separates them is that the first clears in a moment and the second does not.

## Staying under it

**Back off, do not retry immediately.** Wait, then retry with the wait doubling each time. Retrying a `429` at once makes it worse and is the fastest way to stay limited.

**Do not poll.** Almost everything worth knowing announces itself — see [Webhooks](/webhooks/overview). Polling is the most common reason an integration meets a rate limit at all, and it is avoidable.

**Spread scheduled work.** A nightly job that fires everything at midnight uses your burst in a second and then queues behind your rate. Pace it.

**Ask before you need to.** If your volume is growing, your limit is a configuration change rather than a rewrite — but it is not self-serve, so ask ahead of the launch rather than during it.
