CoherenceDevelopers

Concepts

Rate limits and allowances

Two separate limits: how fast a key may call, and how much each tool may do.

Rate limit

Each key may make 120 requests a minute, counted in fixed one-minute windows. Every response says where you stand:

Headers
RateLimit-Limit: 120
RateLimit-Remaining: 117
RateLimit-Reset: 42

RateLimit-Reset is in seconds. Over the limit you get 429 rate_limited with Retry-After; wait that long, then retry.

Allowances

Each licence has daily and monthly allowances (briefs, questions, receipts), shared between the app, the API and MCP. GET /licences shows this month’s use. When one is used up, calls that need it answer 429 allowance_exceeded until it resets; reading keeps working. Long jobs also have a concurrency limit, for example two briefs at a time, answered with 429 concurrency_limit.