Book Call

RAG vs Fine-Tuning for Private AI

Choose an approach around the task and evidence, not the label on a model.

Illustrative document workflow

What is the difference?

Retrieval-augmented generation (RAG) supplies relevant source material to a model at answer time. Fine-tuning updates model parameters using training examples. Local hosting is a separate deployment choice: either approach may use local or hosted components.

Choosing an initial approach
RequirementStarting pointWhat to test
Frequently changing internal policiesPermission-aware retrievalFreshness, access boundaries, source citations
Consistent task-specific outputPrompting baseline, then fine-tuning if neededHeld-out accuracy, formatting and regression cases
No external inference connectionLocal deployment assessmentHardware capacity, latency, licensing and offline dependencies
Sources plus specialised behaviourEvaluate a combined approachWhether additional complexity improves measured results

Worked example: an internal policy assistant

This is an illustrative design exercise, not a customer result or a measured model benchmark. An employee asks: “What is the current travel approval policy for my team?”

  1. Authenticate the employee and retrieve only policies their role may access.
  2. Keep each policy's effective date and document version with its text.
  3. Ask the model to answer from the retrieved sources and identify missing or conflicting information.
  4. Show source references so the employee can verify the answer.
  5. Route approval decisions to the accountable person; a generated answer does not grant approval.

How would we evaluate it?

Use a held-out question set containing current policies, superseded policies, unsupported questions and permission-boundary cases. Record expected sources, acceptable answers and prohibited disclosures before running the test.

Do not tune against the held-out set and then present it as independent evidence. Keep test data separate from training examples and document model, prompt and retrieval versions.

What does private deployment not solve?

Running locally does not remove hallucinations, weak permissions, insecure tools or sensitive logs. Backups, administrators, model licences and update processes still require review. Confirm which components contact external services.

Sources and next steps

The original RAG paper describes retrieval-augmented generation. Evaluate application controls separately using our tool-permission demonstration. Discuss your use case through private LLM development or training and evaluation.