Which Models Serve Insulin Chat and Knowledge Bases?

With no provider of your own, Insulin chat runs on a Fours-hosted pool and knowledge bases embed with BGE-M3. The documented roster, who serves it, and how to check.

Chengjun Yuan
Chengjun Yuan
Co-founder & CTO · Sep 29, 2026

With no provider of your own connected, Insulin chat runs on the Fours-hosted pool, served by one supplier through the Fours Hosted (DeepInfra) integration, and knowledge bases embed with the hosted BGE-M3, as documented on September 29, 2026.


A security review of an AI workspace usually opens with which vendors may see your prompts. The follow-up is narrower: if nobody connects a provider, which models answer, and who serves them?

For Fours Insulin, a general-purpose AI platform for business work built by Fours, the documentation names them. This post gathers that roster for two jobs: writing chat replies, and embedding knowledge-base documents so they can be searched. The two sit in separate tables, because an embedding model’s job is indexing and search, and the reply comes from a chat model.

Every model named here is documented as of publication, September 29, 2026. The roster can change, so re-check the Fours-hosted models in Insulin’s billing documentation before a review relies on it. Two other kinds of model call sit outside this roster: AI that a custom app runs, and the platform features Insulin runs for itself, such as summarization and reranking, which the billing documentation counts as Fours’ own AI.

Which models serve Insulin chat?

With nothing of your own connected, chat runs on the Fours-hosted pool: DeepSeek V4 Pro as its quality tier, and DeepSeek V4 Flash, GLM 5.3 Flash and Gemma 4 31B as its light tier. Connect a provider with your own key, or a personal Claude Code or Codex sign-in, and its models join ahead of the pool for the agents that can use them.

SourceChat models, documented as of publicationWho can use themWhen they answer
Fours-hosted pool, on Fours’ keyDeepSeek V4 Pro (quality tier); DeepSeek V4 Flash, GLM 5.3 Flash and Gemma 4 31B (light tier)The built-in Insulin assistant, and Personal and Organization agents alike, with the same models either wayWith nothing of yours connected, or as the fallback a turn degrades to; only while the AI Model Policy admits the pool
Anthropic, OpenAI or Gemini, on your own keyModels that the connection makes available in InsulinOrganization agents use the organization’s connections; the built-in assistant and your Personal agents use yoursAhead of the hosted pool
Cloudflare Workers AI or DigitalOcean GradientAI, on your own keyA fixed set of five chat models eachAs aboveAhead of the hosted pool
OpenRouter, Baseten, DeepInfra, Fireworks AI or Together AI, the open-source model aggregators, on your own keyA large catalog eachAs aboveAhead of the hosted pool
Claude Code or Codex, your personal Anthropic or OpenAI sign-inModels that the connection makes available in InsulinYou: the built-in Insulin assistant and your Personal agentsAhead of the hosted pool

Whatever the source, the agent’s Default model runs first. If a call fails, through a provider outage, a usage limit or a rejected key, the turn retries on your next connected provider, and the hosted pool is the fallback it degrades to. A hosted model can be the Default too: either kind of agent may pin one, and the built-in assistant’s model settings show the hosted models beside yours. The documentation charts how a turn resolves to a model, from your own providers down to the hosted pool.

The five models Cloudflare Workers AI and DigitalOcean GradientAI each offer, and what their connect checks prove, are in running Insulin agents on Cloudflare Workers AI or DigitalOcean with your own key. Inbox, which drafts email replies outside a chat, “walks the same chain of candidate models the chat agent uses,” with Fours’ hosted models as the last resort.

Why record the source beside the model?

Because the same name can sit in two rows. DeepSeek V4 Pro is the hosted pool’s quality tier and also one of DigitalOcean GradientAI’s five chat models, on your own key. DeepInfra appears twice as well: Fours Hosted (DeepInfra) is the integration the hosted pool resolves through, on Fours’ key, while DeepInfra is also an aggregator you can connect with a key of your own. A model name alone doesn’t say whose key served a request, so a review should record the source with every model.

What is the Fours-hosted pool?

The Fours-hosted pool is the set of chat models Insulin runs on Fours’ own key, and it runs on one supplier. Every hosted model resolves through a single integration, which the AI Model Policy names Fours Hosted (DeepInfra). Adding that entry to the policy “admits the pool, and nothing more”: it pins no particular model. The pool has two tiers.

Hosted chat modelTier, in the documentation’s wordsOffered
DeepSeek V4 ProQuality tier (Tier 1)Until the tier step-down in a billing period
DeepSeek V4 FlashLight tier (Tier 2)Whenever the pool is offered
GLM 5.3 FlashLight tier (Tier 2)Whenever the pool is offered
Gemma 4 31BLight tier (Tier 2)Whenever the pool is offered

Past a spend point Fours sets, the pool serves only its light tier until the next billing period. Once your organization’s spend on Fours’ own key passes that point, which the billing documentation states and your organization cannot configure, Tier 1 is withheld for the rest of the period and the pool serves Tier 2. No banner appears, no request fails and nothing is paused, and Tier 1 returns at the start of the next period. Usage on your own keys (BYOK) is not counted toward that point, and your own models are unaffected.

What a supplier’s name doesn’t tell you

A roster says which models and which supplier can receive a request, not what happens to one. Don’t read data handling, processing location, retention or training terms into a model’s name or a supplier’s; none of them follows from it. The AI Model Policy works at the same level: it decides which providers receive requests, not what they do with them.

Which embedding models serve Insulin knowledge bases?

By default, knowledge bases embed with the Fours-hosted BGE-M3; the alternatives are Qwen3 Embedding 0.6B on Fours’ key, or a model from a provider connected with your own key. An embedding model is the model that turns documents, and each search against them, into vectors so matching passages can be found. Every document in a knowledge base is embedded with its one model, searches use that same model, and the reply comes from the chat model of the agent that searched.

SourceEmbedding models, documented as of publicationWhich knowledge bases can use themWhen they’re used
Fours-hosted, on Fours’ keyBGE-M3 (the default) and Qwen3 Embedding 0.6BPersonal and organization knowledge bases alike, only while Allow Fours platform key is onWhen chosen, and BGE-M3 when nobody chooses. An organization knowledge base whose chosen provider isn’t available at its first index falls back to one, and says so
Your own provider, on your own keyOpenAI: text-embedding-3-small and text-embedding-3-large. Gemini: Gemini embedding 2. OpenRouter: BGE-M3. DeepInfra: BGE-M3 and Qwen3 Embedding 0.6B. Fireworks: Nomic Embed v1.5. Together: multilingual-e5-large-instructOrganization knowledge bases use the organization’s connections; personal knowledge bases use yours, never the organization’sWhen chosen for that knowledge base

Two things to know when you read the chooser:

  • Every option is labelled with its provider, vector size and hosting, such as openai · 1536 dims · BYOK or deepinfra · 1024 dims · Suger-hosted. The hosting label is how you tell Fours’ BGE-M3 and Qwen3 Embedding 0.6B apart from the options of the same name on your own DeepInfra key, or BGE-M3 on OpenRouter.
  • Not every chat provider embeds. Fours registers no embedding model for Cloudflare Workers AI or DigitalOcean GradientAI, so connecting either adds nothing to the chooser.

How can you tell which model answered?

Check the model settings for what runs first, the reply for any switch, and the Usage panel for what ran over a period. For a knowledge base, the answer is in its own Settings.

Where to lookWhat it tells you
Model settings: the pencil in a chat with Insulin opens Insulin model settings; a custom agent has a Default model fieldWhich model runs first. Insulin’s Models list shows the hosted models beside your own and badges the Default, and the agent picker groups models by provider. A Default whose provider was disconnected, or whose model was retired, reads ”— unavailable”, and the agent never silently switches to another
The replyA turn that switched models and then succeeded ends with a short italic note naming both: first-model hit its usage limit — continued with second-model, or was unavailable when an error caused the switch
A usage-limit cardWhen a model is out of quota and nothing is left to fail over to, the card links to that provider’s usage page, or points you at Settings when the model was Fours-hosted
Settings → Billing → UsageGroup a period by Model (“which model was used”) or by Provider (“which provider served it”), or filter by either
A knowledge base’s SettingsThe Embedding model section shows the Current model. An organization knowledge base that fell back says “Fell back to a Fours-hosted model.” A personal one never switches on its own; it marks the affected documents Failed instead

Two caveats apply to the Usage panel. It is re-aggregated every couple of hours, so the latest turns may not appear yet. And Fours doesn’t currently ask Cloudflare Workers AI, DigitalOcean GradientAI or Together AI for a token count on streamed replies, so a reply streamed from one of them on your own key may be missing from the breakdown.

Failover is surfaced on purpose, rather than left for you to infer. Who picks the model, and what happens when it fails explains why Insulin has no per-conversation model picker at all.

How do you take the hosted pool off the table?

Use Settings → Organization → AI Model Policy, which an org ADMIN changes. Leave Fours Hosted (DeepInfra) off a non-empty Allowed AI integrations list, or turn off Allow Fours platform key, which also removes the hosted embedding models from the knowledge-base chooser. Both controls fail closed, so read how the AI model allowlist decides which vendors see your prompts before changing either.

What should a model review record?

The date you read the roster, the policy that admits each source, the connections behind each kind of agent and knowledge base, and what actually ran. As a checklist:

  1. Date the roster. Record the hosted models with the day you read them. The tables here are documented as of September 29, 2026; re-check the billing documentation before each review.
  2. Read the AI Model Policy. Is the allow-list empty, or does it include Fours Hosted (DeepInfra)? Is Allow Fours platform key on? Together they decide whether the hosted pool is offered for chat, and the switch decides the hosted embedding models.
  3. List the organization’s connected providers. They serve Organization agents and organization knowledge bases. The built-in assistant and Personal agents use each person’s own connections instead, a Claude Code or Codex sign-in included, and personal knowledge bases use each person’s own embedding providers.
  4. Read each knowledge base’s Current model, and note any that fell back to a Fours-hosted model.
  5. Compare with what ran. After a period, group Usage by Model and by Provider, allowing for the couple-of-hours lag and for streamed replies the breakdown may miss.

Frequently asked questions

Which AI models does Insulin use by default?

With no provider of your own connected, chat runs on the Fours-hosted pool: DeepSeek V4 Pro as its quality tier, and DeepSeek V4 Flash, GLM 5.3 Flash and Gemma 4 31B as its light tier, as documented on September 29, 2026. Knowledge bases embed with BGE-M3.

Who serves the Fours-hosted models?

One supplier. Every hosted model resolves through one integration, which the AI Model Policy names Fours Hosted (DeepInfra). That is separate from connecting your own DeepInfra key, which makes DeepInfra’s models available on your own key.

Does a knowledge base’s embedding model write the answers?

No. An embedding model indexes a knowledge base’s documents and runs its searches. The reply comes from the chat model of the agent that searched, which is why chat and embedding models are separate rosters.

How can I tell which model answered a chat?

Check the agent’s model settings for its Default model, then read the reply: a turn that switched models ends with a note naming both. Over a period, group Settings → Billing → Usage by Model or by Provider.

Why might the hosted quality tier stop answering?

Once your organization’s spend on Fours’ own key passes a point Fours sets, the pool withholds its quality tier for the rest of the billing period and serves its light tier. No banner appears, and your own models are unaffected.

How do we keep prompts off the hosted models?

Use the AI Model Policy. Leave Fours Hosted (DeepInfra) off a non-empty allow-list, or turn off Allow Fours platform key to make the organization bring-your-own-key only. Turning the switch off also removes the hosted embedding models.

Takeaways

  • With nothing of your own connected, chat runs on the Fours-hosted pool: one supplier, a quality tier and a light tier.
  • Knowledge bases embed separately, with the hosted BGE-M3 by default. Embedding models index and search; chat models write replies.
  • Your own models are preferred and the hosted pool is the fallback; past a spend point Fours sets, it serves only its light tier until the next period.
  • The same model name can come from Fours’ key or yours, so record the provider with every model.
  • A supplier’s name says nothing about how a request is handled; removing the pool is an AI Model Policy decision.

A roster says which models can answer; an agent’s Default model decides which one answers first. That choice is made agent by agent: see Insulin agents, each scoped to one job with its own model.

Sources

Primary sources for the platform rules cited above. Last verified September 29, 2026. Cloud providers change fees, eligibility, and program terms without notice — check the source before relying on a figure.

  • Insulin Billing: What Is Metered — Fours Doc — The Fours-hosted pool running on one supplier, and the Fours-hosted models it names, read for their names and tiers only: DeepSeek V4 Pro as the pool's quality tier, and DeepSeek V4 Flash, GLM 5.3 Flash and Gemma 4 31B as its light tier; own-key usage attributed to the provider whose key was used, and a reply streamed from Cloudflare Workers AI, DigitalOcean GradientAI or Together AI possibly missing from the breakdown; model calls a Custom App runs; platform features such as summarization and reranking counted as Fours' own AI
  • Insulin Agents — Fours Doc — The Default model's sources: Anthropic, OpenAI and Gemini on your own key, a Claude Code or Codex sign-in for a Personal agent, the open-source model aggregators OpenRouter, Baseten, DeepInfra, Fireworks AI and Together AI with large catalogs, and Cloudflare Workers AI and DigitalOcean GradientAI with five chat models each; the picker grouped by provider; hosted models offered to Personal and Organization agents alike, the same models either way; whose connected providers each kind of agent adds; your own models preferred and the hosted models the fallback a turn degrades to; failover on an outage, a usage limit or a rejected key; the resolution chain through the platform-key switch, the allow-list and the Tier 1 spend threshold; either scope may pin a hosted model; Insulin model settings, its Models list and Default badge; the unavailable marker and no silent switch; the failover note naming both models; the usage-limit card pointing at Settings for a hosted model
  • Insulin Getting Started: AI Model Policy — Fours Doc — Every Fours-hosted model resolving through the one Fours Hosted (DeepInfra) integration, which admits the pool and pins no model; a non-empty allow-list without it failing closed, and an empty list restricting nothing; Allow Fours platform key on by default, and bring-your-own-key only when off; the switch deciding whether the hosted embedding models are offered; only an org ADMIN changing Organization tabs; the built-in Insulin assistant working out of the box on Fours-hosted models
  • Insulin Billing: Payment, Limits, and Top-Ups — Fours Doc — The hosted pool's two tiers, a Tier 1 quality model and Tier 2 models; Tier 1 withheld for the rest of the billing period once spend on Fours' own key passes a point Fours sets and the organization cannot configure; no banner, and nothing paused, refused or refunded; reset at the start of the next period; own-key usage not counted, and your own models unaffected
  • Insulin Knowledge Bases — Fours Doc — The embedding model used to index and search a knowledge base's documents; the Fours-hosted BGE-M3 (the default) and Qwen3 Embedding 0.6B, offered to either visibility only while the platform key is on; the own-provider embedding models for OpenAI, Gemini, OpenRouter, DeepInfra, Fireworks and Together, and whose connections each kind of knowledge base uses; the provider, vector-size and hosting labels; the Current model in Settings; an organization knowledge base's fallback note and a user knowledge base marking documents Failed
  • Insulin Billing: Reading Your Usage — Fours Doc — Grouping a period by Model (which model was used) and by Provider (which provider served it), the matching filters, and the breakdown re-aggregated every couple of hours
  • Insulin Inbox: AI Model — Fours Doc — The AI model card listing Claude Code, Codex, Anthropic, OpenAI and Gemini; Inbox's work outside a chat walking the same chain of candidate models the chat agent uses, with Fours' hosted models as the last resort
  • Cloudflare Workers AI — Fours Doc — A fixed set of five Workers AI chat models; organization connections serving Organization agents and user connections serving Personal agents; chat models only, with no Workers AI embedding model registered
  • DigitalOcean GradientAI — Fours Doc — A fixed set of five GradientAI chat models, DeepSeek V4 Pro among them; organization and user connections; chat models only, with no GradientAI embedding model registered
  • DeepInfra — Fours Doc — Connecting your own DeepInfra key makes DeepInfra's models selectable, attributed to DeepInfra, for agents and knowledge-base embeddings; organization and user connections
  • Anthropic — Fours Doc — At the user level, a personal API key or a Claude Code connection tied to your own Anthropic account
  • OpenAI — Fours Doc — Codex as OpenAI's user-level integration, for personal workspace access

Browse every post on the Insulin Blog

Stay Updated

New posts, product updates and marketplace strategy are shared on LinkedIn as they publish.

Follow Fours on LinkedIn