A usage limit is one AI model’s quota running out; a billing pause is your organization’s account refusing AI requests for one of four named reasons, one of them the monthly spending limit you set. In Insulin the two look different on screen, and each has its own fix.
A reply stops partway, or never starts, and a card sits where the answer should be. The first question is practical: press Retry, wait, or find whoever looks after billing?
The word that muddles it is limit. A model’s usage limit is a quota on that one model. Your organization’s monthly spending limit is a ceiling on what it may spend in a calendar month. Either can stop a reply, and they clear in different ways: a usage limit when the model’s limit resets, a spending limit when it is raised or removed, or when the next month starts.
Insulin keeps them apart on screen, withdraws an interrupted card when the turn finished after all, and offers Switch model only when it knows what went wrong, because a model change fixes only some causes. This post is a key to those cards, for the stops Insulin makes on its own: what each card says, what stopped the reply, and the next click.
Which stop is this?
Read the card’s words before its buttons: a usage limit names a model, a billing pause names an account reason, and an interrupted card says the reply stopped before the turn finished. Then match it here.
| What you see | What stopped the reply | What to do |
|---|---|---|
”You’ve hit the usage limit for <model>.”, plus “Your limit resets at <time>” when the provider reports one | That model’s quota is spent, and failover had nothing left to try | Follow the card to the provider’s usage page, or to Settings for a Fours-hosted model. Wait for the reset time if there is one, then Retry. The card doesn’t offer Switch model |
| A message naming the model you hit first, listing the others that are limited too, and when the first limit resets | Every connected model is limited | Check your providers’ usage pages, or try again once a limit clears |
| Monthly spending limit reached | The month’s spend hit the limit your organization set | Raise or remove the limit on Settings → Billing. Adding credit changes nothing |
| This organization is suspended, AI features are paused — your credit has run out, or All billable features are paused — your account has reached its debt limit | Another billing reason: a hold placed by Fours, a pay-as-you-go balance at zero, or a balance past the account’s negative floor | Contact support for a suspension; add credit for the other two |
| A reply cut off mid-text, with the billing explanation | The balance ran out while the reply was streaming | Add credit. Service resumes by itself about a minute after the credit lands |
| ”The response stream ended before this turn finished. Retry to run it again.”, or another interrupted card | The reply stopped arriving before the turn finished, and the turn may still have finished on the server | If the turn finished, the card is withdrawn and the reply loads in its place. If a note then says the reply could not be loaded, reload the conversation. If the card stays, Retry |
| An error card offering both Retry and Switch model | A rejected key, a model that could not be reached, or a provider that is down | Retry, or Switch model to pick a different model before trying again |
A billing reason appears as a banner at the top of Settings → Billing, and in the conversation if you were chatting when it happened. You are told only the first reason that applies.
Three other things can look like a stopped reply, and each already has its own explanation:
- An italic note at the end of a finished reply, such as “
first-modelhit its usage limit — continued withsecond-model.”, is not a stop. Failover worked, and the note names both models; it reads was unavailable instead when an error such as a provider outage or a rejected key caused the switch. Who picks the model, and what happens when it fails explains the design. - A turn with no model to run on at all stops with a message of its own: a reply from the built-in assistant, or an error card naming the cause. What people see when no model is left walks through each one.
- A failed CRM action in the Inbox chat rail leaves out Retry on purpose, because the CRM write may already have happened. The Inbox CRM action audit trail shows what to check before running it again.
What does a usage-limit card mean?
It means one model’s quota is spent and the turn had nowhere left to fail over to. When a model call fails on a usage limit, Insulin first retries the turn on your next connected provider. The card appears only when there is nothing left to try, and it carries three things:
- The model. “You’ve hit the usage limit for
<model>.” - A reset time, when there is one. It adds “Your limit resets at
<time>” when the provider reports one. - Where to look. It links to that provider’s usage page, or points you at Settings when the model was Fours-hosted.
A usage limit is not only a bring-your-own-key event. A Fours-hosted model can reach one too, which is when the card points at Settings instead of a provider’s page. Either way, the card offers Retry, not Switch model.
When every connected model is limited, one message covers them all. It names the model you hit first, lists the others that are limited too, and gives the time the first limit resets, with a suggestion to check your providers’ usage pages or try again once a limit clears.
Whose quota is it?
It depends on who owns the agent. The built-in Insulin assistant and personal agents run on the providers you connected under Settings → Integrations. Organization agents run on the providers the organization connected, which only an org ADMIN can connect or disconnect. Both prefer those connected models, with Fours-hosted models as the fallback wherever your organization allows them.
So when a card from an organization agent links to a provider’s usage page, the account behind it is the organization’s, not yours. That is the detail to pass on when you ask an admin about it.
Is a usage limit the same as your monthly spending limit?
No. A usage limit is one model’s quota; the monthly spending limit is a ceiling your organization sets on its own spend, and reaching it pauses AI requests for the rest of the month unless you raise or remove it.
| A model’s usage limit | Your monthly spending limit | |
|---|---|---|
| What it caps | One model’s quota | What the organization may spend in a calendar month |
| What it stops | That model. The turn fails over first, and stops only when nothing is left | AI requests, for the rest of the month. Automatic top-up stops too |
| What you see | The usage-limit card, or the every-model message | Monthly spending limit reached, in the conversation and at the top of Settings → Billing |
| Where to look | The provider’s usage page, or Settings for a Fours-hosted model | Settings → Billing |
| When it clears | When the model’s limit resets. The card gives the time when the provider reports it | About a minute after you raise or remove it; otherwise at the start of the next month |
Adding credit does not clear the spending limit. The limit is compared against what the organization has spent this month, not against its balance, so topping up changes nothing until the limit is raised or the month ends.
A third limit sits among the billing reasons: the debt limit, which applies when the balance falls past the negative floor set for the account. Credit clears it, and support can change the floor.
What happens to a reply when AI is paused for billing?
A request that has not started is refused before any model is called; a long reply that is already streaming is cut off where the text stops once the balance is known to have run out. Either way, the conversation shows the reason and what to do about it.
- A streaming reply is checked every ten seconds. The check reads a decision held for up to a minute, so a reply can run briefly past zero before it stops.
- If the check itself cannot be answered, the reply finishes rather than being truncated.
- You are told the first reason that applies. A suspension is checked first, then your monthly spending limit, then the balance. The four reasons Insulin pauses AI covers what each means and what keeps working while it lasts.
- Service resumes by itself, with nothing to switch back on: about a minute after credit lands, or after you raise or remove your monthly limit. With automatic top-up, allow longer: the top-up runs every five minutes, and the minute starts once it has been charged. A suspension ends only when Fours lifts it; a successful payment does not clear it on its own.
A billing cut-off and an interrupted card can both leave a reply unfinished, but they read differently. The cut-off comes with a billing reason; the interrupted card is about the stream.
What does an interrupted reply mean?
An interrupted card means the reply stopped arriving before the turn finished. It reads, for example, “The response stream ended before this turn finished. Retry to run it again.” One cause is a connection that drops while a reply is still being written.
The card is provisional, because the turn may still have finished on the server:
- If the turn finished, the card is withdrawn and the reply is loaded in its place.
- If that reply then cannot be loaded, the card is replaced by the note “This turn finished, but its reply could not be loaded here. Reload the conversation to see it.” Reload the conversation rather than retrying, because the work is already done.
- If the card stays, Retry runs the turn again. The card does not offer Switch model.
A brief drop with no reply in progress needs nothing from you: the client reconnects in the background, and no banner is shown.
What reloading or leaving the page does to a reply that is still being written is a question about stopping work rather than diagnosing it, and how to stop work in Insulin answers it.
Retry or Switch model: which button does what?
Retry re-sends the message that failed; Switch model opens the agent’s model settings so you can pick a different model before trying again. Two details decide which to press.
Retry re-sends the message immediately above that card, not whatever you typed most recently. Scroll back to an old error and Retry still re-sends the right turn.
Switch model appears for three classifications only, because a model change fixes only some causes:
| How the failure was classified | Retry | Switch model |
|---|---|---|
| Usage limit — that model’s quota is spent | Yes | No |
| Provider auth — the key was rejected | Yes | Yes |
| Model unavailable — the model itself could not be reached | Yes | Yes |
| Provider unavailable — the provider is down | Yes | Yes |
| Interrupted — the reply stream ended before the turn finished | Yes | No |
| Unclassified — the cause is not known | Yes | No |
If a custom agent’s saved Default model is unavailable, reconnect its provider or choose another model.
Frequently asked questions
Is a usage limit the same as the monthly spending limit?
No. A usage limit is one model’s quota, and a turn stops on one only when failover has nothing left to try. The monthly spending limit caps your organization’s own spend; reaching it pauses AI requests until it is raised or removed, or the month ends.
Do Fours-hosted models have usage limits?
Yes. A usage-limit card can name a Fours-hosted model as well as one from a provider you or your organization connected. For a Fours-hosted model the card points you at Settings; otherwise it links to that provider’s usage page.
What does Retry re-send?
The user message immediately above that error card, not whatever you typed most recently. Scroll back to an old error and Retry still re-sends the turn that failed.
When does Insulin offer Switch model?
Only when a failure is classified as a rejected key, a model that could not be reached, or a provider that is down. A spent quota, an interrupted reply and a failure of unknown cause offer Retry but not Switch model.
Should I retry an interrupted reply?
If the card stays, yes: Retry runs the turn again. It is provisional, though. If the turn finished on the server, the card is withdrawn and the reply loads; if a note says the reply could not be loaded, reload the conversation instead.
Can a billing pause cut off a reply that is already streaming?
Yes. A long streaming reply is checked every ten seconds and, once the balance is known to have run out, cut off where the text stops, with the billing explanation. It can run briefly past zero first; if the check cannot be answered, it finishes.
Takeaways
- Read the card’s words first: a usage-limit card names a model, a billing pause names one of four account reasons, and an interrupted card says the reply stopped before the turn finished.
- A usage limit is one model’s quota. It stops a reply only when failover has nothing left to try, and Fours-hosted models can reach one too.
- The monthly spending limit is your organization’s own ceiling. Adding credit does not clear it; raising or removing it does, within about a minute.
- An interrupted card is provisional: if a note says the turn finished, reload rather than retry.
- Retry re-sends the message above that card. Switch model appears only for a rejected key, an unreachable model or a provider that is down.
Knowing which card you are looking at decides the next click. To see how agents are scoped, grounded and shared, start with Insulin agents; the documentation on picking the turn back up and when AI pauses is the reference for every card above.
Sources
Primary sources for the platform rules cited above. Last verified September 29, 2026. Cloud providers change fees, eligibility, and program terms without notice — check the source before relying on a figure.
- Insulin Agents — Fours Doc — A turn retrying on the next connected provider after a usage limit, outage or rejected key; whose providers the built-in assistant, personal agents and organization agents use, with Fours-hosted models as the fallback; the italic failover note and its was unavailable variant; the usage-limit card text, the reset line when the provider reports one, and its link to the provider's usage page or to Settings for a Fours-hosted model; the every-model-limited message; Retry re-sending the message above that card; Switch model and the six-row classification table; the failed CRM action that leaves Retry out; the interrupted card's example text, its withdrawal and the reload note; background reconnection with no banner; a custom agent's unavailable saved Default model; a turn with no model to run on
- Insulin Billing: When AI Pauses — Fours Doc — The four billing reasons and their banner texts, in the conversation and at the top of Settings → Billing, checked in a fixed order with the first that applies reported; a request refused before the model is called; a long streaming reply checked every ten seconds and cut off once the balance is known to have run out, the decision held for up to a minute, and the reply allowed to finish when the check cannot be answered; automatic resumption about a minute after credit lands or the monthly limit is changed, the five-minute top-up run, and a suspension that a payment does not clear; credit not clearing the spending limit
- Insulin Billing: Payment, Limits, and Top-Ups — Fours Doc — The monthly spending limit as a ceiling on the organization's calendar-month spend that pauses AI requests for the rest of the month, stops automatic top-up and resets at the start of the next month; compared against spend rather than balance, so adding credit does not resume service
- Insulin Getting Started — Fours Doc — Organization integrations under Settings → Integrations: any member can view them, and only an org ADMIN can connect or disconnect them
Keep reading
Stay Updated
New posts, product updates and marketplace strategy are shared on LinkedIn as they publish.
Follow Fours on LinkedIn