Blog

AI & Machine Learning articles

Integrating Large Language Models into Business Software

An LLM integration guide for business software: architecture, structured outputs, validation, error handling, cost control, logging and evaluation.

5 min read AI & Machine Learning

Calling a large language model from code takes a few lines. Building a feature on top of one that behaves reliably for thousands of users, keeps costs predictable and does not leak data is a different job. LLM integration is mostly ordinary software engineering applied to an unusual component: one that is slow compared to a database, costs money per call, and occasionally returns something unexpected. This guide covers the design decisions that make the difference, aimed at product managers and developers adding AI features to existing business applications.

Pick tasks that suit a language model

LLMs are strong at reading and writing language: summarising, classifying, extracting fields from text, drafting replies, translating, and answering questions over supplied documents. They are weak at exact arithmetic, guaranteeing facts they were not given, and making consistent decisions where rules already exist. A good integration gives the model the language part of a task and keeps calculations, lookups and final decisions in normal code.

Architecture: keep the model behind your backend

Never call an LLM provider directly from a browser or mobile app with your API key embedded. Anyone could extract the key and run up your bill. Route every call through your own server, which:

  • Holds the API keys in a secrets store.
  • Authenticates the user and checks what they are allowed to do.
  • Builds the prompt from trusted templates plus the user's input.
  • Applies rate limits and spending caps per user or tenant.
  • Validates the response before anything acts on it.
  • Logs requests for debugging and auditing.

Wrap provider calls in a small internal service or module with one interface. Models and providers change frequently, and this isolation lets you switch or test alternatives without touching the rest of the application.

Ask for structured output, then validate it

Free-form text is hard for code to use. For extraction and classification tasks, ask for JSON matching a defined schema. Many providers offer a structured-output or JSON mode that constrains responses to a schema you supply. Even so, validate everything on your side:

TICKET_SCHEMA = {
    "type": "object",
    "properties": {
        "category": {"enum": ["billing", "delivery", "product", "account", "other"]},
        "urgency":  {"enum": ["low", "normal", "high"]},
        "summary":  {"type": "string", "maxLength": 300}
    },
    "required": ["category", "urgency", "summary"],
    "additionalProperties": False
}

def classify_ticket(text):
    raw = llm_client.generate(
        system=CLASSIFY_INSTRUCTIONS,
        user=text,
        response_schema=TICKET_SCHEMA,
        max_output_tokens=300,
        timeout_seconds=20,
    )
    data = json.loads(raw)
    jsonschema.validate(data, TICKET_SCHEMA)   # raises if invalid
    return data

(Shown as provider-neutral Python; llm_client stands for whichever SDK you use.) If validation fails, retry once, then fall back: route the ticket to the "unclassified" queue rather than crashing or guessing.

Tool calling: let the model ask, let your code decide

Most current models support tool calling (also called function calling): you describe functions such as get_order_status(order_id), and the model can respond with a request to call one. Your code then runs the function and returns the result to the model.

The crucial rule is that your code, not the model, enforces permissions. Check that the logged-in user owns the order before returning its status. Prefer read-only tools. For actions with consequences, such as refunds, cancellations or emails to customers, require explicit human confirmation. Treat every model request as untrusted input, because user text and retrieved documents can attempt prompt injection to steer the model into misusing tools.

Plan for latency and failure

  • Timeouts. Responses can take several seconds or longer. Set explicit timeouts and decide what users see when one is hit.
  • Retries with backoff for rate-limit and temporary server errors, with a cap on attempts.
  • Streaming for chat-style features, so text appears as it is generated rather than after a long pause.
  • Background jobs for bulk work, such as summarising a thousand documents, using a queue instead of a web request.
  • Graceful degradation. If the AI service is down, the core application should still work, just without the AI feature.

Control cost from day one

Providers charge per token, a chunk of text roughly the size of a short word or word fragment, counting both input and output. Costs grow quietly with long prompts, long documents and chatty outputs. Practical controls:

  • Use the smallest model that meets your quality bar; reserve larger ones for hard cases.
  • Set max_output_tokens for every call.
  • Send only relevant context, not entire documents, by using retrieval.
  • Cache results for identical inputs, and use the provider's prompt caching where it is offered for repeated instructions.
  • Track spend per feature and per customer, with alerts.

Handle data responsibly

Know exactly what data leaves your systems. Review the provider's terms on retention and on use of your data for training; business and API terms often differ from consumer apps. Remove or mask personal data that the task does not need, choose data processing regions where available, and make sure your privacy notice and contracts cover the processing. Laws on this vary by country.

Log, evaluate and monitor

Store each request's prompt template version, model name, inputs (redacted where necessary), output, latency, token counts and validation result. When a user reports a bad answer, you can then reproduce it.

Before launch, build an evaluation set: real examples with expected outputs. Run it whenever you change the prompt, model or provider, because a change that fixes one case can break others. After launch, sample live outputs for human review and track user feedback. Model versions are updated or retired by providers, so pin versions where possible and re-evaluate before upgrading.

An LLM integration pre-launch checklist

  1. API keys only on the server; per-user rate limits and spending caps in place.
  2. Structured outputs validated, with a defined fallback.
  3. Tools permission-checked in code; consequential actions confirmed by a person.
  4. Timeouts, retries and graceful degradation tested.
  5. Data handling reviewed against provider terms and your obligations.
  6. Evaluation set passing; logging and monitoring live.

Our AI and machine learning development and web application development teams build LLM features into existing products, and our cloud solutions team can host the supporting services.

Key takeaways

  • Successful LLM integration is mostly solid engineering around an unpredictable component.
  • Keep calls on your backend, request structured output and validate it.
  • Let the model suggest tool calls, but enforce permissions and confirmations in code.
  • Budget for latency and cost, log everything, and re-run evaluations on every change.

Need help with this?

Netifi helps businesses around the world with AI & Machine Learning. Tell us what you are working on.