Two ways to find the right page
The step in question
Before Ask Prover answers a question about signalling standards or Prover's tools, it has to pick a handful of pages out of a large collection and put them in front of the model. That step is called retrieval.
A vector database stores each document as an embedding, a list of numbers that places the text on a map of meaning, so a question can match a page that shares no words with it. File search gives the model an engineer's terminal tools instead.
How each one works
What the vector database costs to run
A second service. Something more to run, and to keep in step with the documents.
Tokens to load. 2.3 million to write summaries, 1.2 million to embed, before anyone can ask anything.
Loading time for 437 files, leaving a 70 MB data directory.
The question was whether all that buys better answers, or only ceremony.
The test
Same model, same documents, same ten questions
Claude Opus 5 with three tools: list, search for a word, read a file
- 10 questions
- 4 min 50 s
- Cost
- $2.10
The same model and loop, no index. Semantic search over OpenViking 0.4.19 with OpenAI's text-embedding-3-small, plus read and word search inside the store
- 10 questions
- 5 min 10 s
- Cost
- $2.10
Prover's internal knowledge base: railway concepts, Swedish regulation topics, internal tool notes
- Files
- 437
- Words
- ~150,000
- Figures
- 279
- Questions
- 10
- 5 concept questions, worded deliberately away from the document headings.
- 5 exact questions, asking for a flag, an exit code, a version number, a formula or a configuration key.
- Each fixed in advance with the file that holds the answer and the exact sentence that answers it. Cost is at list prices.
Semantic search on its own
No model in the loop: is the right file in the first ten results? "Text only" filters to text documents afterwards (this version had no file-type filter); "Word" is a keyword search with one hand-picked keyword per question.
Correct file in the top ten
Word search winsThis matters most: the answers were still nine of ten in the next test, so the model made up for the misses itself.
Files skipped silently. Their formats could not be parsed by this version.
Files merged. Three sharing a source citation became one entry. Files were stored under their first heading, so a lookup by filename found nothing.
Top results were images on one Swedish regulation question: a vision model captioned the 279 figures in dense English, while the prose around them is Swedish.
Semantic search alone found the right file in its top ten for half the questions; counting any file in the right topic folder, it reached nine of ten. Speed was never the problem: 0.13 to 0.67 seconds per search, and under 0.05 seconds for word search. What got in the way was loading, shown above; worth knowing before anyone loads documents into a store like this.
With the model: the answers
Both arms end to end, model included, over the same ten questions.
Answers matching the known answer
TieThe tenth: "no" for file search, "partly" for the vector database.
Correct file opened or seen
TieCost at list prices
TieFile search re-sends its index with every question: 963,000 cache-read tokens against 33,000. Cache-read tokens are text the provider keeps ready from earlier requests, billed at a small fraction of the normal rate.
Wall time, lower is better
About evenModel turns, lower is better
File search fewerInput tokens, lower is better
File search fewerOutput tokens: 19,000 against 21,000.
Nine of ten either way, at the same cost and the same speed. The one difference was the question about the common-cause-failure formula. File search stopped at a similar formula in a neighboring file and never opened the right one. The vector database surfaced the right file at rank five and the model read it, but it still led with the wrong formula. With ten questions, a one-question difference is noise.
Why the answers came out even
Semantic search alone put the right file in the top ten only half the time, and the answers were still nine of ten. The model recovered by searching again and by falling back to word search. Whichever store you use, the model needs a plain keyword search beside it.
Vector database: good at
- Matching a question to a document that means the same thing in other words
- Speed: 0.13 to 0.67 seconds per search
- Other jobs, not tested here: per-user document spaces, memory from past sessions, parsing PDF and Word files people upload
Vector database: not good at, in this test
- Finding the exact file: top ten for 5 of 10 questions
- Looking a file up by name, or keeping files apart that share a citation
- Telling you what it skipped: 42 files dropped silently
- Mixed languages: English image captions crowded out Swedish text
- Being cheap to own: a second service, kept in step, with a loading bill
What we decided
What Ask Prover runs today
No embeddings, no vector database, no second service to run and no second token bill for loading the documents. The measurement gave no retrieval-quality reason to add one. The ten questions became Ask Prover's standing evaluation, run against the documents the site actually serves.
When we would add one
This is a decision about this version, not a verdict on the technology. A vector index is one database migration away, and it goes in the day an evaluation shows questions that are missed and that only semantic search fixes. The other reasons to run such a store are real but different, and worth re-testing when we need them:
Per-user document spaces
Each person searching their own set of documents.
Memory from past sessions
Extracting what earlier conversations established.
Uploaded PDF and Word files
Parsing the documents people bring themselves.
None of these was tested in this experiment.