
RAG, short for retrieval-augmented generation, looks up the relevant passages in your documents at the moment a question is asked and hands them to the model with the question, so answers can use current company data and cite where it came from. Fine-tuning retrains a model on examples so it behaves a particular way, such as following a format, a tone or a classification scheme, but it does not keep the model current with your files. For an assistant that answers from company data, start with RAG, and fine-tune only when the model has the right facts but still answers in the wrong shape.
OpenAI's own guide to optimizing accuracy frames it the same way: RAG fixes what the model needs to know, fine-tuning fixes how the model needs to act, and prompt engineering is usually the best place to start. This guide covers how each works, what each costs as of October 2026 and what changed for fine-tuning this year. For the business case behind an internal assistant, see our AI knowledge base FAQ.
How do RAG and fine-tuning compare?
| RAG | Fine-tuning | |
|---|---|---|
| What changes | What the model sees with each question | The model's weights, through extra training |
| Best for | Answers from documents, policies, records and anything that changes | A consistent format, tone or task, such as classifying requests or filling a template |
| Keeping current | Update the documents; answers reflect the change once the index refreshes | Retrain when the task data changes |
| Permissions | Can retrieve only what the person asking may open | What the model learned can't be limited per user afterward |
| Citations | Can show the passage each answer came from | No built-in link back to a source |
| Upfront work | Clean documents, splitting them into passages, an index and permission checks | Curated example prompts and answers, training runs and evaluation |
| Running cost | A search service, plus longer prompts on every question | Per-token use of the tuned model and, on some platforms, hosting fees |
| Typical failure | Retrieving the wrong passage, so the answer is wrong | Fluent answers with outdated or invented facts |
What is retrieval-augmented generation?
Anthropic's Contextual Retrieval article describes RAG as a method that retrieves relevant information from a knowledge base and appends it to the user's prompt. In practice your documents are split into passages and indexed, a question triggers a search, and the best matches go to the model with instructions to answer from them and to say when they don't contain the answer.
You are probably using RAG already. Microsoft's Copilot Studio documentation describes tenant graph grounding as retrieval-augmented generation over your tenant's Microsoft Graph, and Microsoft 365 Copilot grounds answers in the files, mail and chats each user can open. The same pattern, built over a CRM, an ERP or a shared drive, is the usual design for a custom company assistant.
One shortcut is worth knowing. In September 2024, Anthropic noted that a knowledge base smaller than about 200,000 tokens, roughly 500 pages, can go into the prompt whole with no retrieval step, and that prompt caching cuts the cost and delay of sending it repeatedly. For a small policy library, that can be simpler than building search.
What is fine-tuning?
Microsoft's fine-tuning overview describes it as adapting a pretrained model to your task through additional training on task-specific data, adjusting the model's weights rather than adding examples to each prompt. It lists where that helps: consistent style and structure, better tool use, fewer examples needed in every prompt, and specializing a smaller, cheaper model. It is also blunt about the limits: fine-tuning adds training and hosting costs and doesn't replace retrieval for current information.
The work is in the examples. You need a set of prompts and ideal answers that represents the task, a held-out set to compare the tuned model with your current one, and a plan to retrain as the task data changes or the base model approaches retirement.
What changed for fine-tuning in 2026?
Fine-tuning became harder to buy. According to OpenAI's deprecation notice, organizations that had never run fine-tuning lost the ability to start fine-tuning jobs on May 7, 2026; since July 2, 2026, so have organizations that hadn't run inference on a fine-tuned model in the previous 60 days; and on January 6, 2027, active existing customers lose the ability to create new ones. Models already fine-tuned stay available until their base models are deprecated. Microsoft Foundry still offers managed fine-tuning for a list of models, with training and deployment choices that affect price and data residency.
For a small business, that is one more reason to keep company knowledge in a retrieval layer you control rather than inside one provider's tuned model. RAG works with whichever model you use next.
Which fits a small team?
- An assistant that answers from policies, procedures, contracts or project files: RAG, or the whole document set in the prompt if it is small.
- Answers from Microsoft 365 for staff who have Copilot: Microsoft 365 Copilot already retrieves from what each user can open; an agent or a Graph connector extends it to other sources.
- A high-volume task with a fixed output, such as sorting thousands of emails into categories: start with a well-written prompt and examples, and consider fine-tuning a smaller model only if accuracy or cost then falls short.
- Brand voice or report format: try instructions and templates first; fine-tuning is the last resort, not the first.
What does each cost?
RAG costs are mostly search and longer prompts. As of October 2026:
- Turning documents into searchable vectors: OpenAI lists text-embedding-3-small at $0.02 per million tokens, so indexing a few thousand pages costs cents.
- The search service: a Basic tier on the Azure AI Search pricing page lists at $0.101 an hour in East US, about $74 a month. OpenAI's hosted file search charges $0.10 per GB per day of storage after the first free gigabyte, plus $2.50 per 1,000 searches.
- Model use: every answer carries the retrieved passages, so prompts are longer; OpenAI lists gpt-4.1-mini at $0.40 per million input tokens.
Fine-tuning costs come in three parts. Training is priced per token or per hour: OpenAI lists $5.00 per million training tokens for gpt-4.1-mini, for customers who can still run jobs. Using the tuned model costs more than the base model: $0.80 per million input tokens and $3.20 per million output tokens for a fine-tuned gpt-4.1-mini on OpenAI, double the base rates. On Microsoft Foundry, standard deployments of fine-tuned models add hosting charges on top of per-token billing. Then add the people time to build examples, evaluate and retrain. OpenAI figures come from its API pricing page.
How NetSys helps
Our custom AI application development builds assistants on the retrieval pattern. The model reads your systems at query time through Microsoft Graph, APIs or a read replica, under the permissions of the person asking, and each answer cites the record it came from; the model never becomes a copy of your data. Work starts with a one-page specification of the questions, sources, permitted actions and users, a data readiness check and a cost-to-run estimate before any code. Builds use Anthropic Claude with Azure, Microsoft Graph and Power Platform, are tested against injected instructions before rollout, and are maintained under a month-to-month agreement. Where the data already lives in Microsoft 365, a Microsoft Graph connector may be the cleaner route.
Frequently asked questions
What is the difference between RAG and fine-tuning?
RAG retrieves relevant passages from your documents at question time and gives them to the model, so answers reflect current data and can cite sources. Fine-tuning retrains the model on examples to change how it behaves, such as its format or tone. RAG changes what the model sees for each question; fine-tuning changes how it acts.
Is RAG cheaper than fine-tuning?
For company knowledge, usually. RAG needs a search index and longer prompts, while fine-tuning adds training runs, higher per-token rates for the tuned model on some platforms, possible hosting fees and retraining whenever the data changes. Fine-tuning can save money at very high volume when it lets a smaller model replace a larger one.
Does fine-tuning teach a model our company data?
Not reliably, and not in a way you can keep current or limit by user. Microsoft states that fine-tuning doesn't replace retrieval for current information. Use retrieval for facts and documents, and fine-tuning, if at all, for behavior.
Does Microsoft 365 Copilot use RAG?
In effect, yes. Copilot retrieves the Microsoft 365 content each user can open and grounds its answers in it, and Microsoft describes Copilot Studio's tenant graph grounding explicitly as retrieval-augmented generation over your tenant's Microsoft Graph. Content outside Microsoft 365 needs a connector or a custom application before it can be retrieved.
What affects the cost, implementation and support of an AI assistant on company data?
How clean and well organized the documents are, how many systems the assistant reads, how permissions must be enforced, and how many questions it handles. Support means keeping the index current, reviewing wrong answers, watching model and search costs, and retesting when the model changes.
Discuss custom ai applications for your business.
Tell us about your current systems, the result you need and your timeline. We will discuss the work, responsibilities and pricing before you decide on an engagement.



