AI Chatbot

Building a Low-Hallucination AI Chatbot With Ollama and OpenAI

Quantiq Automation (product in beta)

Building a Low-Hallucination AI Chatbot With Ollama and OpenAI — Quantiq Automation case study

A beta AI chatbot that knows your business and customers, limits what users can see, and keeps hallucination low using Ollama and OpenAI. It is a prototype in beta testing with five customers and clients, answers in seconds, and has not been deployed yet.

Prototype in beta testing, not yet deployed

5 beta clients

Client
Quantiq Automation (product in beta)
Focus
AI Chatbot · Ollama · OpenAI API · Local LLM
Published

The Challenge

Most business chatbots fail in one of two ways. Either they know too little, so they give generic answers and start inventing details the moment a question touches your own prices, orders, tickets or policies, or they know too much, so a customer who phrases a question cleverly can pull out information that was never meant for them. A support or sales chatbot sits exactly between those two failures. It needs deep context about your systems and about the person it is talking to, and it needs a hard boundary around what any one person is allowed to see.

Hallucination is the first half of the problem. A large language model produces fluent text whether or not it has the facts, and a confident wrong answer about a refund rule or a delivery date is worse than no answer, because the customer acts on it. Research on retrieval-augmented generation (RAG), where a model is given relevant source material before it answers, reports that it significantly reduces hallucinations in an enterprise application (arXiv 2404.08189). But RAG only helps when the right context actually reaches the model. Thin or missing context leaves gaps, and models fill gaps with guesses.

Data exposure is the second half. The OWASP Top 10 for LLM Applications (2025 edition) lists Prompt Injection (LLM01), Sensitive Information Disclosure (LLM02), Excessive Agency (LLM06) and Misinformation (LLM09) among the main risks of putting language models in front of users. A chatbot that has access to everything and simply trusts the model to behave is exposed to all four. Limits have to be enforced by the system around the model, not requested politely in a prompt.

Then there is the trade-off between cost, speed and capability. Sending every message to a top-tier cloud model is slower and more expensive than most questions deserve, and it means every customer message leaves your own infrastructure. Using only a small local model keeps things in-house and quick, but it hits a ceiling on complex searching and multi-step reasoning. Published research on hybrid routing (Hybrid LLM, ICLR 2024) found that sending easy queries to a small model and hard ones to a large model can cut large-model calls by up to 40 percent with no drop in response quality, although that is the researchers' benchmark and not our result.

So the brief we set ourselves was specific. Build an AI chatbot that holds the full context of the business and of each customer, keeps hallucination very low, never hands over more data than the user is allowed to see, answers in seconds, and still copes with the hard questions. And build it as a product Quantiq can offer to its own customers, not a one-off for a single account.

What We Built

We built the chatbot as two model layers working together, a local layer and a cloud reasoning layer, with a context and limits layer wrapped around both.

The local layer runs on Ollama, which serves open models on infrastructure we control instead of through a shared public chatbot service. It handles the routine questions that make up most real conversations: order and account questions, policy questions, how-do-I questions. These models are limited on purpose. They only work from the context supplied for the person asking, and they carry restrictions on what they are allowed to reveal, so the chatbot does not hand a customer all the data it can technically reach.

The reasoning layer uses the OpenAI API, and it is reserved for the very complex work: deep searching across a large amount of context, comparing many records, and multi-step reasoning that a small local model handles poorly. Keeping this layer for hard tasks is what keeps the everyday experience fast and the cost sensible. OpenAI's business data page states that it does not train its models on API or business data by default, and that qualifying organizations can use zero data retention on the API platform. We still recommend that each customer checks the current terms for their own region and contract.

The context layer is what gives the chatbot its knowledge. The chatbot is supplied with the context of your systems and of the customer in front of it, including business knowledge, policies and the customer's own history, so answers come from that material and not from the model's general memory. The behavioural rule that goes with it matters as much as the data: when the supplied context does not contain the answer, the chatbot should say so instead of guessing. This is the main lever against hallucination.

The limits layer is the principle that the system decides what can be disclosed, not the model's goodwill. Prompt injection, sensitive information disclosure and excessive agency are the risks the design is built around, which is why the local models are restricted and why a customer asking for another customer's data, or for something outside their scope, should get a refusal instead of an answer.

Here is a simplified example of how one conversation moves through the system. A customer asks where their order is. The chatbot assembles the context for that customer, the local Ollama model answers from it in seconds, and nothing needs to go to a cloud model. A few messages later the same customer asks why this month's invoice is higher than the last three. That is a comparison across several records and a contract, so the router escalates that task to the OpenAI API, which searches and reasons across the supplied context and returns a grounded answer, still under the same limits. If the context showed no reason for the difference, the correct answer is that it cannot tell, and the chatbot says exactly that.

The beta is designed to test the hardest case for hallucination: heavy context. Long, dense context can bury the one fact that matters, and that is where chatbots often start to drift. So we are deliberately feeding the prototype large amounts of business and customer context and checking whether its answers stay faithful to it. Testing is under way with five customers and clients, and the chatbot responds in seconds.

We are packaging it as a Quantiq product rather than a one-off project. The core stays the same, and each customer gets it configured with their own context, their own limits and their own tone.

The Result

Status first, because it matters: this chatbot is at prototype stage and in beta testing. It has not been deployed to production. Five customers and clients are testing it now, answers arrive in seconds, and the two-model design, local Ollama models for routine questions and the OpenAI API for complex reasoning, is running in the beta.

What the beta is checking is specific. We are loading the chatbot with heavy context on purpose to see whether it stays faithful to that context, whether it admits when the context does not contain the answer, and whether the disclosure limits hold when a user pushes for data they should not see. Before any production launch we plan to measure the grounded-answer rate (how many answers can be traced to supplied context), refusal correctness, leak and prompt-injection attempts, how often a question escalates to OpenAI, response time, and cost per conversation.

We have not attached hallucination-rate percentages, cost savings or ROI to this case study, because beta testing is still running and we only publish figures we can stand behind. The one number here that is real is the number of testers: five. The 40 percent figure quoted earlier comes from published research and is not our result.

What comes next is on the roadmap but not built yet: a handoff to a human agent with a summary of the conversation, citations that show which source each answer came from, an audit log of every question, answer and model used, a feedback loop where thumbs up and down grow a test set, and replies in the customer's own language. Buyers in regions with strict data rules, such as the UK, Europe and the UAE, tend to ask where their data is processed. Because Ollama is self-hosted, where the routine layer runs is a deployment decision, and it is one we can plan together with each customer.

Is the AI chatbot live yet? No. It is a prototype in beta testing with five customers and clients and has not been deployed to production.

Why use Ollama and OpenAI together instead of one model? Local Ollama models keep routine questions fast, in-house and restricted. The OpenAI API is used for very complex searching and reasoning that smaller models handle poorly. Using both keeps everyday answers quick while still giving hard questions the depth they need.

How does the chatbot keep hallucination low? It answers from the business and customer context it is given, and it is built to say it does not know when that context does not contain the answer. No language model can promise zero hallucination, so we measure it in beta instead of claiming it away.

Can the chatbot leak other customers' data? The design goal is that it never gives a user more than they are allowed to see, using restricted local models and disclosure limits that follow the OWASP risks for sensitive information disclosure and prompt injection. Beta testing includes pushing on those limits, and we will publish results once we can stand behind them.

Can my business get this chatbot? We are packaging it as a Quantiq offering. If you want to see how it would work with your own systems and customer data, get in touch and we will walk you through the beta.

Sources: OpenAI, Business data privacy, security, and compliance (openai.com/business-data). Ding et al., Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing, ICLR 2024 (arxiv.org/abs/2404.14618). arXiv 2404.08189, a study of retrieval-augmented generation in an enterprise workflow-generation application. OWASP Top 10 for LLM Applications 2025 (genai.owasp.org/llm-top-10).

Ollama (self-hosted open models)OpenAI APIRetrieval-Augmented Generation (RAG)Hybrid LLM routingRole-scoped data accessDisclosure limits and guardrails

Working from the UK, Europe, South Africa or the UAE?

We deliver remotely. See how an engagement works for your region:

Want something like this built for you?

Tell us about your workflow and we'll put together a tailored plan.

Book a Free Auditarrow_forward

More Case Studies