The most dangerous answer is a confident wrong one
The pipeline, with real numbers
| Stage | Setting |
|---|---|
| Source | 62-page PDF (61 pages with text) |
| Chunking | 300 words, 30% overlap (90 words), minimum 50 words |
| Embeddings | all-MiniLM-L6-v2, 384 dimensions |
| Vector store | FAISS flat index, L2-normalised for cosine similarity |
| Retrieval | Top 5 chunks |
| Threshold | 0.25 minimum similarity |
| Generation | GPT-4o-mini, temperature 0.3 |
The one line that matters
Python
What the scores actually look like
| Rank | Page | Similarity |
|---|---|---|
| 1 | 41 | 0.643 |
| 2 | 41 | 0.257 |
| 3 | 39 | 0.107 |
| 4 | 40 | −0.001 |
| 5 | 43 | −0.091 |
Why 0.25?
- Set it too high and the assistant refuses questions it could have answered.
- Set it too low and it's back to guessing.
Chunking: the boring decision that decides everything
- 300 words is big enough to hold a complete rule with its context.
- 30% overlap means a rule that crosses a chunk boundary still appears whole in at least one chunk.
- A 50-word minimum drops tiny fragments, like headers and page numbers, that match everything and mean nothing.
Lessons that carry over to any RAG project
- Retrieval always returns something. Decide what "not good enough" means and enforce it in code.
- Refusing is a feature. In a support or policy assistant, "I don't know" is far cheaper than a wrong answer.
- Use a low temperature for grounded answers. You want the model to restate the source, not get creative.
- Log scores, not just answers. You can't tune a threshold you can't see.
- Start small. A flat FAISS index on disk is enough for hundreds of pages. Add infrastructure when you need it.
Want an assistant that answers from your own documents and knows when to stop? Tell me what you're working with.