Skip to content

How to Train a Chatbot on Your Company’s Documents (RAG, Without the Jargon)

Published · Updated

Company binders and manuals next to a laptop with a chat open

No chatbot trained on your company’s documents is right 100% of the time. The best industry RAG systems answer only 63% of questions without hallucinating, according to Meta’s CRAG benchmark (NeurIPS 2024). Any vendor promising “zero hallucinations” is contradicting public evidence. With that settled, here is how it works, what it costs, and where to start.

The second uncomfortable truth: half of all “chatbot on my documents” projects are documentation projects in disguise. The system cannot cite a policy nobody ever wrote, and it cannot read a crooked scanned PDF.

What does “training” a chatbot on my documents actually mean?

Three different things get called “training,” and confusing them is the most expensive mistake in this category. About 95% of projects sold as “training” are actually the second one: RAG.

Approach What it actually does Good for Updating one fact
Fine-tuning Retrains the model’s weights Style, format, behavior A full retraining run
RAG (retrieval + generation) Searches your documents at answer time and cites the source Proprietary, changing, verifiable knowledge Replace one PDF
Prompt with context Pasting the document into the conversation Dozens of pages, one-off cases Paste the file again

The number that ends the argument comes from Microsoft. In Fine-Tuning or Retrieval? (Ovadia et al.), the authors conclude that “RAG consistently outperforms it, both for existing knowledge encountered during training and entirely new knowledge,” and that LLMs “struggle to acquire new factual information through unsupervised fine-tuning.”

The rule for an owner or CEO: fine-tuning changes how the model talks; RAG changes what it knows. If the bot doesn’t know your price list or your warranty terms, your problem is RAG. Starting with fine-tuning is spending money on the thing that doesn’t solve it.

How does RAG work if I’m not technical?

Picture a newly hired attorney: brilliant, fast, and completely unfamiliar with your business. You sit them next to a filing cabinet. Every time you ask a question, someone pulls a few index cards and sets them on the desk. The attorney reads only those cards and writes the answer. That’s RAG.

  1. Ingestion. Your PDFs, spreadsheets, and emails go into the cabinet as plain text. This is where scanned PDFs and phone photos of documents die.
  2. Chunking. Documents are cut into half-page cards. If a cut lands mid-table, the card is useless.
  3. Embeddings. Each card gets a “coordinate of meaning,” not keywords. That’s why “how much PTO do I get?” finds the card that says “paid leave accrual.”
  4. Vector search. Your question pulls the 5 to 20 cards closest in meaning.
  5. Reranking. A second reader reorders them and throws out the ones that only looked relevant.
  6. Generation. The attorney reads only those cards and drafts the answer with citations.

The honest part: the attorney doesn’t know what never reached the desk, and if the wrong cards arrive, the answer comes back just as confident. Almost every RAG failure is a filing-cabinet failure, not a model failure.

What are my options and what do they cost?

They range from “set it up this afternoon” to six months of development. Prices below are published vendor rates as of September 2026. Where a vendor does not publish amounts, we say so instead of inventing a number. Mexican peso conversions use the Banxico FIX rate of 16.9755 MXN/USD (September 1, 2026).

Option Cost (USD, Sept 2026) Time to deploy Main limit Best for
Custom GPTs / ChatGPT Business $25/user/month ($20 annual, 2–200 users) Hours Max 20 files, 512 MB each; no per-team permissions Internal pilot
Claude Projects (Team) $25/user/month ($20 annual), 2-seat minimum Hours No granular permission control Dense documents
Gemini Notebook (formerly NotebookLM) Included with Workspace or AI Pro Minutes Not a customer-facing bot, can’t be embedded Internal research
Chatbase Free $0 · Hobby $40 · Standard $150 · Pro $500/month 1–3 days 40 MB of content max (Pro) Customer support on your site
Copilot Studio $200/pack/month (25,000 credits); M365 Copilot $30/user/month 1–4 weeks Inherits SharePoint permissions, and its mess Companies already on Microsoft 365
Dify (cloud) Free sandbox (50 docs) · Professional $590 · Team $1,590/year 1–3 weeks Document ceiling per plan Control without coding
n8n + vector store Starter €20 · Pro €50/month; self-hosted free 2–4 weeks Someone has to build and maintain it Wiring RAG into processes
Vertex AI Search $1.50/1,000 queries (Standard), $4.00 (Enterprise), $5/GB/month index; 10,000 queries/month free 2–6 weeks Google Cloud lock-in High query volume
AWS Bedrock Knowledge Bases $5/GB/month index + $1/1,000 Retrieve calls ($4/1,000 agentic) 2–6 weeks AWS lock-in Companies already on AWS
Azure AI Search Does not publish amounts. Microsoft’s docs use an illustrative $100/month per search unit 3–8 weeks Real price only comes from the calculator Microsoft environments
Custom build In Mexico, $50,000–$300,000 MXN ($2,950–$17,700 USD) implementation — agency-reported market estimate, not audited — plus infrastructure 4–16 weeks Permanent maintenance cost Unique business logic

Besides Azure AI Search, neither Qdrant (Standard and Premium) nor Botpress enterprise plans publish amounts. If someone hands you those numbers as list price, ask for the formal written quote.

What a U.S. buyer should know about building this in Mexico

The custom-build range above is the reason nearshore comes up. A comparable U.S. agency engagement generally starts well above it, though that comparison is a market estimate, not a published rate. Mexico City runs on U.S. Central time, so a team there overlaps your entire workday, which matters more than hourly rate on a project whose hardest work is cleaning your documents with you.

Where the cards live: vector databases

The vector database is the filing cabinet. For most small and mid-sized companies, under one or two million vectors, the right answer is the boring one: pgvector on PostgreSQL, free and probably already in your stack.

  • pgvector / Supabase: free to 500 MB, Pro $25/month, Team $599/month USD.
  • Qdrant: free forever (1 node, 0.5 vCPU, 1 GB RAM, 4 GB disk). Standard and Premium prices are not published.
  • Chroma: free open source; Cloud Team $250/month plus usage ($2.50/GiB written, $0.33/GiB-month).
  • Weaviate: free to 100,000 objects; Flex from $45 and Premium from $400/month USD.
  • Pinecone: Starter free (2 GB), Builder $20, Standard from $50, Enterprise from $500/month USD with a 99.95% SLA.

How accurate is a chatbot trained on my documents, really?

Less accurate than you’ll be told. The public numbers:

  • 44% with naive RAG. On Meta’s CRAG benchmark, the best standalone LLMs reach 34% accuracy; adding straightforward RAG moves them to 44%. Industry RAG solutions answer 63% without hallucinating.
  • 17% to 33% hallucination in elite tools. Stanford RegLab tested 202 legal queries: Lexis+ AI and Westlaw AI hallucinate in that range despite “hallucination-free” marketing. Lexis+ AI was correct 65% of the time.
  • 25.8 points lost to digitization. OHR-Bench, across 8,500+ pages: 67.6% correct answers with standard OCR versus 91.2% with clean text.
  • Industrial documents break OCR. InduOCRBench (April 2026): GPT-4o drops from 75% to 52%, PP-StructureV3 from 86.7% to 60.3%. High OCR accuracy doesn’t translate into good RAG: 82.9% OCR produced 52.8% RAG accuracy.
  • Chunking matters as much as the model. FloTorch (2026): recursive 512-token chunks 69%, fixed-size 512 67%, semantic chunking 54%. Vectara (NAACL 2025), across 25 configurations and 48 embedding models: “chunking configuration had as much or more influence than the choice of embedding model.”

Realistic expectation: a well-built RAG system on clean documents with bounded questions lands around 70–85% correct answers. That’s the number to negotiate against. Anyone promising 95%+ or “zero hallucinations” is selling something public evidence doesn’t support.

How do I measure whether the chatbot is working?

With an evaluation set, not impressions. The most widely used framework is Ragas; TruLens, DeepEval, and Azure AI Foundry do the same job. It measures five things.

  • Faithfulness or groundedness: is the answer supported by the retrieved cards? This is the hallucination metric.
  • Response relevancy: did it answer the question asked?
  • Context precision: were the right cards retrieved?
  • Context recall: was anything missed that did exist?
  • Noise sensitivity: does it fall apart with noisy input?

Demand the same thing before you sign: a set of 50 to 100 real questions from your business with answers you already know, scored against the proposed system. If the vendor won’t run it, the conversation is over. In our AI implementations that set exists before the first line of code.

Why do document chatbot projects fail?

Gartner predicted at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. The seven causes:

1. The content doesn’t exist

Nobody wrote the policy you’re asking about, so the system fills the vacuum with a plausible invention. Write the manual first, automate second.

2. Scanned PDFs and tables

Up to 25.8 accuracy points are lost in digitization alone. Worse, tables sliced by the chunker return wrong figures in perfect formatting, the hardest error to catch.

3. Bad chunking

43-token fragments with no context, or cuts through the middle of a contract clause. The gap between best and worst strategy reaches 15 points on identical documents.

4. Contradictory or expired documents

Three versions of the price list sit in the repository and the system confidently cites the 2023 one. RAG has no sense of what’s current unless you build it.

5. Permissions that don’t carry over

The bot surfaces whatever the repository already had mispermissioned. It doesn’t create the leak; it makes the leak searchable in plain English. Audit permissions before indexing.

6. Hallucination despite RAG

17% to 33% even in top-tier legal tools, per Stanford. A bot that cites sources can cite the wrong one, or cite correctly and conclude badly.

7. Shipped with no measurement and no owner

Barnett et al. (CAIN/IEEE-ACM 2024), covering three real case studies — one with 4,017 documents and 1,000 questions — concludes that “the validation of a RAG system is only feasible during operation.” With no owner for the corpus, quality decays with every new document.

What does each level cost and how long does it take?

Level Timeline Year-one software cost In pesos
No-code pilot, 5 users 1–5 days $1,200–$1,500 USD $20,000–$25,500 MXN
Customer-facing no-code bot (Chatbase Standard) 1–3 weeks ~$1,800 USD plus content cleanup ~$30,500 MXN
Mid-tier platform (Dify or n8n) 3–8 weeks $3,000–$8,000 USD plus fees $50,000–$136,000 MXN
Enterprise cloud (Bedrock or Vertex) 6–12 weeks $6,000–$25,000 USD $100,000–$425,000 MXN
Custom build 3–6 months $2,950–$17,700 USD implementation plus $500–$2,000 USD/month infrastructure $50,000–$300,000 MXN

Converted at the 16.9755 MXN/USD FIX rate. Custom-build ranges are agency-reported market estimates, not list prices. For a firm number on your scope, request a quote with a defined scope.

What happens legally if the chatbot gets it wrong?

You’re on the hook. In Moffatt v. Air Canada (BC Civil Resolution Tribunal, February 14, 2024) the airline paid about $650 CAD plus interest over a fare its chatbot invented. The tribunal rejected the argument that the bot was a separate entity: “it should be obvious to Air Canada that it is responsible for all the information on its website.”

If your bot touches Mexican customer data, the new LFPDPPP (published March 20, 2025, effective March 21) sets fines in Article 59 of 100 to 160,000 UMA for minor violations and 200 to 320,000 UMA for serious ones. At the 2026 UMA of $117.31 MXN per day, the ceiling is $37.5 million MXN (about $2.2 million USD), doubled for sensitive data. INAI was replaced by the Secretaría Anticorrupción y Buen Gobierno.

Do the AI vendors train on my documents?

The big three say no, in writing. OpenAI: “we do not train our models on your data by default” across Business, Enterprise, and API, with up to 30-day retention, Zero Data Retention available, and SOC 2 Type 2. Anthropic: “by default, we will not use your inputs or outputs from our commercial products to train”; with explicit feedback, conversations are kept up to five years. Google Workspace: “your uploads, queries and responses will not be reviewed by humans or used to train AI models.”

Where should I start without overspending?

With a one-week test that costs almost nothing and tells you whether the problem is the technology or your filing cabinet.

  1. Week 1. Load 20 real documents into a Custom GPT or Claude Project and write 50 questions whose correct answers you already know. Score it.
  2. Under 60% correct: your documents are the problem, not the technology. Fix the corpus before spending on a platform.
  3. Over 75% correct: the use case works. Now evaluate platforms by volume, permission needs, and channel: website, WhatsApp, or intranet.
  4. Never: start with fine-tuning, or sign a custom build without running step 1 first.

What else do buyers ask before signing?

How accurate is a document-trained chatbot, really?

Roughly 70% to 85% correct on clean documents with bounded questions. The industry state of the art on open questions is 63% without hallucination, per Meta’s CRAG benchmark. No system reaches 100%, and any vendor promising it is contradicting public evidence.

Is my data safe with ChatGPT or Claude?

On business plans, yes, contractually. OpenAI and Anthropic both state they don’t train on commercial customer data by default, and Google Workspace says the same. Check retention: OpenAI keeps data up to 30 days; Anthropic keeps feedback-flagged conversations up to five years.

How long does it take to get one running?

An internal no-code pilot: hours to five days. A customer-facing site bot: one to three weeks. A platform with permissions and connectors: three to eight weeks. A custom build: three to six months. Cleaning the documents usually takes longer than the technical setup.

Do I need a developer?

Not for a pilot on Custom GPTs, Claude Projects, or Chatbase. Yes for per-team permissions, ERP integration, or WhatsApp deployment with custom logic. The practical line: if the bot has to read from or write to another system, you need development.

Which documents work and which don’t?

Digital text works: Word files, computer-generated PDFs, web pages, and knowledge bases. What breaks the system: scanned PDFs, photos of documents, engineering drawings, spreadsheets with merged cells, and slide decks with text baked into images. Standard OCR alone costs up to 25.8 accuracy points.

What’s the cheapest option that actually works?

A Custom GPT on ChatGPT Business at $25 USD per user per month ($20 USD annual), capped at 20 files. Five users runs about $1,200 to $1,500 USD in year one, roughly $20,000 to $25,500 MXN. That’s enough for a serious internal pilot.

What if the chatbot tells a customer something wrong?

You answer for it. In Moffatt v. Air Canada the tribunal held the airline liable for a fare its bot invented and refused to treat the bot as a separate entity. Mitigate with visible disclaimers, source citations on every answer, human escalation, and conversation logs.

Fine-tuning or RAG?

RAG, almost always. Microsoft’s Fine-Tuning or Retrieval? paper concludes RAG consistently outperforms fine-tuning for adding knowledge, including brand-new knowledge. Fine-tuning changes how the model talks; RAG changes what it knows. Updating a price costs a retraining run with fine-tuning; with RAG you replace a PDF.

Want to know whether your documents can support a chatbot?

Iterando scores your real corpus against a question set from your own business before proposing any platform. If the answer is that documentation is the problem and not the technology, we tell you that.

See our AI solutions for companies

Message us on WhatsApp