NEXWERK / PRACTICAL RESOURCES
RAG vs Fine-Tuning for Private AI
Choose an approach around the task and evidence, not the label on a model.

What is the difference?
Retrieval-augmented generation (RAG) supplies relevant source material to a model at answer time. Fine-tuning updates model parameters using training examples. Local hosting is a separate deployment choice: either approach may use local or hosted components.
| Requirement | Starting point | What to test |
|---|---|---|
| Frequently changing internal policies | Permission-aware retrieval | Freshness, access boundaries, source citations |
| Consistent task-specific output | Prompting baseline, then fine-tuning if needed | Held-out accuracy, formatting and regression cases |
| No external inference connection | Local deployment assessment | Hardware capacity, latency, licensing and offline dependencies |
| Sources plus specialised behaviour | Evaluate a combined approach | Whether additional complexity improves measured results |
Worked example: an internal policy assistant
This is an illustrative design exercise, not a customer result or a measured model benchmark. An employee asks: “What is the current travel approval policy for my team?”
- Authenticate the employee and retrieve only policies their role may access.
- Keep each policy's effective date and document version with its text.
- Ask the model to answer from the retrieved sources and identify missing or conflicting information.
- Show source references so the employee can verify the answer.
- Route approval decisions to the accountable person; a generated answer does not grant approval.
How would we evaluate it?
Use a held-out question set containing current policies, superseded policies, unsupported questions and permission-boundary cases. Record expected sources, acceptable answers and prohibited disclosures before running the test.
- Retrieval: did the permitted, relevant source reach the model?
- Grounding: does each material claim agree with that source?
- Abstention: does the system acknowledge missing evidence?
- Isolation: can a user obtain another team's restricted material?
- Operations: record latency and compute cost under defined conditions.
Do not tune against the held-out set and then present it as independent evidence. Keep test data separate from training examples and document model, prompt and retrieval versions.
What does private deployment not solve?
Running locally does not remove hallucinations, weak permissions, insecure tools or sensitive logs. Backups, administrators, model licences and update processes still require review. Confirm which components contact external services.
Sources and next steps
The original RAG paper describes retrieval-augmented generation. Evaluate application controls separately using our tool-permission demonstration. Discuss your use case through private LLM development or training and evaluation.