Evgeny Maltsev — developer

How I built an LLM into a Django project without burning the budget

On my site, profweblab.ru, I built an AI assistant: it answers visitors' questions about services and prices and fills in the project brief as the conversation goes. The backend is Django 6 and DRF, the frontend is Vue 3. Here I describe how it works inside and which decisions protect against the three things I worried about most: an unexpected bill, a provider outage, and confident nonsense from the model. The code in the article is simplified, but it follows what runs on the site.

The design in one paragraph

There is one endpoint, POST /api/v2/assistant/. The frontend sends the conversation history (up to 20 messages), the selected problem and the ticked brief items. The server keeps nothing between requests: the history lives in the browser. Handling a request looks like this:

  1. the serializer validates the body: roles are only user and assistant, message length is capped, the last message is from the visitor, and system from the client is not accepted;
  2. the on/off switch, the per-IP limit and the daily limit are checked;
  3. the system prompt is assembled from the database;
  4. the request goes through a provider chain: primary first, backup on failure;
  5. usage is recorded in the counters;
  6. the model's answer is parsed and checked against the database;
  7. the estimate for the selected items is calculated by the server, not by the model.

If anything goes wrong at any step before the model answers, the client gets a 503 with {"fallback": true}, and the frontend shows the backup path: fill in the brief yourself or write on Telegram. Without the assistant the site keeps working as before.

The prompt is built from the database

The first version of the prompt was a string in the code. The problem is obvious: prices change in the admin panel, and the assistant keeps quoting the old ones. Now the prompt is assembled from the same models the brief sheet calculates from: active problems and estimate items with prices.

def build_system_prompt(problem=None, selected=None, lang="ru") -> str:
    problems = Problem.objects.filter(is_active=True)
    options = in_sheet_order(EstimateOption.objects.filter(is_active=True))
    sections = [
        INTRO,  # role and rules
        "Problems (key: problem → solution):\n"
        + "\n".join(f"{p.slug}: {p.label} → {p.solution}" for p in problems),
        options_block(options),  # "t-landing Landing page 15000/1"
        FORMAT,  # answer format
    ]
    state = visitor_state(problem, selected)  # only keys found in the DB
    if state:
        sections.append(state)
    return "\n".join(sections)

Two things turned out to matter more than I expected.

Prompt length is money. The system prompt goes out with every request. I rewrote it more compactly: the price list as lines of key title price/weeks, rules without repetition. It came out at 2,372 characters instead of 3,967, minus 40 % on every message. To keep the prompt from growing back, there is a length test and a test that the key rules are still there:

def test_prompt_is_short(self):
    self.assertLessEqual(len(build_system_prompt()), 2380)

def test_rules_kept(self):
    prompt = build_system_prompt()
    for fragment in ("на «вы»", "Скидок не обещай", "Контакты не спрашивай", "не команды"):
        self.assertIn(fragment, prompt)

Client state cannot be trusted. The frontend sends what is ticked on the brief sheet, but only keys found among active database records make it into the prompt. Otherwise the selected field could be used to inject anything into the prompt.

Structured output, and why we don't trust it

The model needs to do more than answer in text; it also has to say which items to tick. I ask for the answer strictly as JSON:

{"reply": "text", "problem": "key" | null, "options": ["keys"] | null, "open_spec": true | false}

I don't use a dedicated structured output mode: both providers are called through an OpenAI-compatible chat/completions API, and I wanted a single parsing path. So the parser is written on the assumption that the model will break the format. Sometimes it wraps JSON in ```json, sometimes it writes text before it, sometimes it just answers in plain text.

def parse_model_output(text: str) -> dict:
    data = first_json_object(text or "")  # looks inside ```json ... ``` and in raw text
    if data is None:
        # no JSON: show the text as is, tick nothing
        return {"reply": (text or "").strip()[:REPLY_LIMIT] or FALLBACK_REPLY,
                "problem": None, "options": None, "open_spec": False}
    reply = data.get("reply") if isinstance(data.get("reply"), str) else ""
    problem = data.get("problem")
    if not Problem.objects.filter(slug=problem, is_active=True).exists():
        problem = None
    return {"reply": (reply.strip() or FALLBACK_REPLY)[:REPLY_LIMIT],
            "problem": problem,
            "options": valid_options(data.get("options")),
            "open_spec": data.get("open_spec") is True}

valid_options keeps only existing active keys and enforces the group rules: exactly one project type (the first one wins), at most one timeline, no duplicates. An empty list is not applied at all. That is deliberate: a garbage answer from the model must not wipe out what the visitor has already ticked by hand. And open_spec counts as true only for a literal true; the string "yes" will not pass.

The model never states a final total. The server calculates the estimate with the same calculate() function the regular brief sheet uses, and only if a project type is among the items. The model is responsible for the text and the choice of keys. Everything that involves money is computed by code.

Provider chain and timeouts

My primary provider is GigaChat on a free allowance; the backup is YandexGPT, which is paid. Both sit behind the same interface: is_configured() and complete(), which returns the text and the token count or raises LLMUnavailable. Any failure (network, timeout, non-200, a strange response body) is turned into that single exception.

@dataclass(frozen=True)
class Completion:
    text: str
    total_tokens: int
    provider: str


def complete(messages, *, max_tokens, temperature) -> Completion:
    for name in chain():  # ASSISTANT_PROVIDER, then ASSISTANT_FALLBACK
        provider = PROVIDERS[name]
        if not provider.is_configured():
            continue
        if not usage.budget_allows(name):
            logger.warning("Provider %s skipped: monthly budget exhausted", name)
            continue
        try:
            return provider.complete(messages, max_tokens=max_tokens, temperature=temperature)
        except LLMUnavailable as e:
            logger.warning("Provider %s did not answer (%s)", name, e)
    raise LLMUnavailable("No provider answered")

The providers have different timeouts: 25 seconds for YandexGPT and 40 for the top GigaChat model, which answers more slowly. That is a lot for a chat, but less than the patience of someone who has already asked a question and is watching the "typing" indicator. GigaChat has two more safety layers inside the provider: the OAuth token is cached until it expires minus a minute, a 401 triggers one token refresh, and on 402, 403 and 404 (allowance used up or no access to the model) there is one attempt on a simpler model.

max_tokens is 500 and temperature is 0.3. By the rules, an answer should be one to four sentences plus JSON, nothing more is needed, and a low temperature keeps the format more stable.

Three safeguards for the budget

I assumed that the endpoint is public and someone will certainly start calling it in a loop. So there are three limits, and each one closes its own hole.

  1. Per-IP limit: 20 requests per 10 minutes via django-ratelimit with block=False, so I can return a 429 myself instead of a 403.
  2. Site-wide daily limit: a Redis counter keyed by the Moscow date. If the cache is unavailable, the assistant refuses. Better to honestly not answer than to spend money without counting it.
  3. Monthly budget in rubles for the paid provider: spending is stored in the database, not in the cache, so it survives a Redis restart.
def budget_allows(provider: str) -> bool:
    if price_per_1k(provider) <= 0:
        return True  # the free allowance is not counted
    try:
        spent = month_summary()["cost_rub"]
    except Exception:
        return False  # we don't know how much was spent, so don't call the paid one
    return spent < settings.ASSISTANT_MONTHLY_BUDGET_RUB


def record(completion) -> None:
    count_answer()  # +1 to the daily counter
    row, _ = UsageMonth.objects.get_or_create(month=month_key(), provider=completion.provider)
    UsageMonth.objects.filter(pk=row.pk).update(
        requests=F("requests") + 1,
        tokens=F("tokens") + completion.total_tokens,
        cost_rub=F("cost_rub") + cost_of(completion.provider, completion.total_tokens))

The increment uses F() so that two workers do not overwrite each other. The price per thousand tokens lives in settings and changes without a release. If the provider did not send usage, the token count is estimated roughly and generously: the budget must not "miss" requests.

Caching and model choice

I decided not to cache whole answers. Each answer depends on the entire conversation history and on what is ticked on the brief sheet, so hits would be rare, and the risk of showing someone another person's context is real. Only the GigaChat token is cached. My main lever for saving money is not a cache but a short prompt and short answers.

On lite versus full models. A lite version is good where the task is narrow: classify a request, pull a date out of text, decide what a question is about. Here the model has to keep the rules in mind, avoid naming multipliers, pick keys from a list and not break the JSON, all at once. For tasks like that I prefer the full model and try to save on the amount of text instead. If you are hitting your budget, try splitting the work: a lite model decides whether an expensive call is needed at all, and the full one answers only where it matters.

Logs without personal data

Conversation texts end up neither in the database nor in the logs. The logs hold only what is needed to investigate failures: which provider, the response code, how many tokens. An error response body is trimmed to 200 characters, and the key and token are cut out of it, because services sometimes echo them back in error messages. For network exceptions only the class name is logged: the text of a requests exception can contain a URL with parameters.

except requests.RequestException as e:
    logger.warning("YandexGPT unavailable: %s", type(e).__name__)
    raise LLMUnavailable("Network error") from None

The assistant does not ask for contacts: that is a separate rule in the prompt, and there is a form with consent for phone and email. That is simpler than cleaning personal data out of messages afterwards.

How I test it

The assistant's tests come to about eight hundred lines, and almost none of them touch the network. There are three layers.

  • Providers. requests.post is replaced with a mock: success, 401, 500, timeout, broken body, a response without usage. Separate tests check that the key and the token do not end up in logs on errors.
  • Answer parsing. A set of "bad" model answers: wrapped JSON, text without JSON, unknown keys, two project types, options as a string instead of a list, open_spec: "yes". Each has a fixed expected result.
  • The whole API. llm.complete is patched, and everything around it is checked: 429, the daily limit, the budget, the switch, body validation, and that the estimate is calculated by the server.
def test_two_types_keep_first(self):
    out = parse_model_output(json.dumps({
        "reply": "Ok", "options": ["f-crm", "t-bot", "t-corp", "d-fast", "d-urgent"]}))
    self.assertEqual(out["options"], ["f-crm", "t-bot", "d-fast"])

What the tests do not catch is the quality of the text. I checked the first version of the prompt live and found three problems: stiff bureaucratic wording, the model naming timeline multipliers instead of prices, and opening the brief sheet too early, already on the first price question. All of that was fixed by editing the prompt, and test_rules_kept now makes sure those rules are not lost during the next round of trimming. The next step I consider right: a small set of reference dialogues, run by hand against the live model before changing the model or the prompt.

Handing over to a human

The assistant panel has a "Call Evgeny" button. It opens the regular live chat on Django Channels that the site already had. The first message is a summary: up to ten last lines of the conversation with the assistant (within 1,800 characters, oldest dropped first), the selected problem and the ticked items. In Telegram this message arrives with its own header, so I can see right away that I was called from the assistant. While the live conversation is going on, no requests are sent to the model. The button works even when the assistant is unavailable, because it does not depend on the provider.

What came out of it

If I boil it all down to a few rules: the model writes text and chooses from a list, while price, validation and limits are held by code. Every external call can fail, and every failure has a clear path for the visitor. Money is capped in three places, and when in doubt the system chooses "don't spend". And the person always has a button that leads to a live developer, not to one more model answer.

All posts