Servonaut AI is the hosted AI gateway included on Solo and Teams plans. Hit F2 in the TUI, log in, and start chatting with your fleet — no personal API key needed, no provider account to set up. The part we care most about: you pay from a money balance, not a per-request token quota that loses its meaning the day a vendor changes its pricing.

This post is the design rationale: why money instead of tokens, what happens when a provider has a bad day, and what we deliberately gave up.

Updated October 2026: the AI allowance is now one balance shared by every AI feature. This post describes how it works today.

What it is

The chat panel sends your request to Servonaut's gateway, which routes it to one of several upstream AI providers. If a provider errors, rate-limits or returns an empty answer, the gateway moves on to the next one. Your CLI sees one assistant — Servonaut AI — and the routing stays our problem, not yours.

Every plan with hosted AI includes a monthly AI allowance: £4.50 on Solo, and £14.50 per seat on Teams — pooled into one balance the whole team shares. Chat, server scans and memory summaries all draw on that one balance. When you need more, a one-off top-up adds to the same balance — £5 adds £5.00 and £20 adds £20.00 — and lasts 365 days from purchase. The chat panel and servonaut ai quota show what's left, and so does your account dashboard.

Why money, not tokens

Token quotas have one nice property: they're the unit the provider bills in. They have one terrible property: they have no meaning to a human.

If we said "Solo gets 200,000 tokens per month", the questions that immediately follow are:

  • Is that input tokens or output tokens?
  • For which model?
  • What about cached prompts?
  • Does a 20MB log analysis count the same as a quick question?

The answer is all of the above and none of the above. Token semantics drift across providers, across models from the same provider, and across cache states. So we sidestepped them: every request is charged for the work it actually did, priced per model and kept current as vendors change their prices, and you see one amount of money.

The trade-off: you don't get an exact "tokens left" count. You get an amount that means the same thing tomorrow as it did yesterday — plus a rough "≈ requests left", worked out from what your own requests typically cost.

What happens as the balance runs down

  • Near the end of the monthly allowance, requests are served by a faster, lower-cost model so the balance lasts longer. A top-up brings the full model back.
  • When the balance is empty, requests stop until your allowance renews or you top up. Nothing is charged beyond what you have, except that a request already in progress is allowed to finish.
  • Unused allowance doesn't roll over: each month's allowance expires when the next one arrives. Top-ups are used only after the allowance, so they last as long as they can.

What we don't do

A few things we said no to, all on purpose:

  • No per-request token cap. A big log analysis is fine; what it costs comes out of the balance like anything else.
  • No overage bills. The balance is the ceiling. You're never charged for AI you didn't pay for up front.
  • No vendor juggling for you. Which provider answers a given request is a routing decision we make and change as providers change. You see one assistant, one balance.

Free tier — by design, not by accident

Free-tier users get no hosted AI. That's intentional, not a bug:

  • The free CLI is MIT-licensed and unrestricted. You can wire up your own Anthropic / OpenAI / Gemini / Ollama key in config.json and chat to your heart's content.
  • The hosted gateway is the value-add of the paid tier. It's the one feature where "give it away free" runs us a real, monthly bill we can't claw back.

If you want AI without paying, run a local Ollama install and point the CLI at it. That works on the free tier and costs us nothing — which means we're happy to keep recommending it.

Trying it

servonaut login from the CLI, hit F2, type something. The chat panel shows your remaining balance. Solo and Teams plans on the pricing page. Free tier instructions on the quickstart page.