Blueprint mode is on — see the design decisions behind this site

Controlling AI cost: a multi-model gateway and a budget per conversation

In an AI product every answer costs money, and user growth can turn into growing losses. A multi-model gateway, the right model for each task, caching and usage caps make the cost predictable.

In ordinary software, serving the thousandth user costs about the same as serving the first. AI products are different: every message, every answer and every processed document has a direct cost.

If that cost isn't designed for from the start, the day comes when each new customer shrinks the margin.

Where the cost comes from

Language models are priced by the amount of input and output text. Three things push it up:

  • Long context. Each time you send the whole conversation history and ten documents, you pay for all of it.
  • A large model for a small task. Using the most capable model to recognise that a message says "hello".
  • Repetition. Answering the same question again and again.

The multi-model gateway

The first step is to stop the product depending directly on one provider. All requests pass through a central layer — a gateway.

This is the swappable layers pattern applied to models. The gateway does several jobs in one place:

  • Choosing the model by task type.
  • Counting usage per customer and per feature.
  • Enforcing caps and stopping abnormal usage.
  • Failing over when a provider is unavailable.
  • Logging for quality and cost review.

Without it, each of these has to be repeated in ten places.

The right model for the task

Not every task needs the strongest model. A practical split:

TaskSuitable model
Classification, intent detection, simple extractionSmall and fast
Routine answers from known contentMid-range
Multi-step reasoning, complex document analysisMost capable

In a sales assistant many messages are simple. Send those to a small model and only the hard cases to a strong one, and cost drops noticeably without quality falling where it matters.

Shortening context

  • Send only what is relevant. A few matching products, not the whole catalogue.
  • Summarise history. In long conversations keep older turns as a summary.
  • Write compact instructions. Fixed text sent with every request is repeated thousands of times.

Caching

Frequent questions have similar answers. Storing answers to identical or near-identical questions lowers cost and improves speed. Just make sure a cached answer is still valid for data that changes, such as price and stock.

Caps and budgets

In a SaaS product every customer needs a defined budget:

  • A cap that matches the plan. Usage should be proportionate to what that customer pays.
  • A warning before the cap. Not a sudden cut-off.
  • Defined behaviour after the cap. For example, dropping to a cheaper model or passing to a human operator.
  • Anomaly detection. A runaway loop or abuse must not burn the budget overnight.

This is tied to tenant isolation: each customer's usage has to be counted and limited separately.

Measure cost alongside quality

Cutting cost at the price of bad answers isn't a saving. Any change to model or context should be run against an evaluation set so you know what happened to quality.

The number to watch isn't cost per message. It is cost per outcome: how much each conversation that reached a correct answer or a sale actually cost.

What to log from day one

  • Usage by customer, feature and model.
  • Response time.
  • Errors and failovers.
  • The share of cached answers.

Without this, optimisation is guesswork.

The takeaway

AI cost isn't a finance problem to solve later. It is an architectural decision. A central gateway, the right model per task, short context, caching and usage caps keep a product out of the "the more we grow, the more we lose" trap.

Related: AI in the product: core, not decoration.

Written by

Mohammad Ali Eslamipour

Mohammad Ali Eslamipour is a digital product architect who defines, architects and — with his own engineering team — ships AI products, SaaS platforms and web apps.