← Reports

Does Ask Prover still need a vector database?

Before building Ask Prover, the question-answering assistant on Prover Labs, we had to choose how it finds things. We measured both ways in one afternoon, on the same documents and the same ten questions, and the measurement decided it.

Two ways to find the right page

The step in question

Before Ask Prover answers a question about signalling standards or Prover's tools, it has to pick a handful of pages out of a large collection and put them in front of the model. That step is called retrieval.

A vector database stores each document as an embedding, a list of numbers that places the text on a map of meaning, so a question can match a page that shares no words with it. File search gives the model an engineer's terminal tools instead.

How each one works

Vector database (OpenViking) 437 files summarize + embed once, about 35 min store: 330 entries 70 MB on disk 2.3M tokens of summaries 1.2M tokens of embeddings question nearest by meaning top 10 results Claude File search question Claude list files search for a word read a file search again its instructions carry an index of every file and its first heading, about 12,000 tokens
A vector database reads every document before the first question; file search sends the model to the files, the way an engineer would. All figures are from this experiment. A token is the unit a model is billed by, roughly three quarters of a word.

What the vector database costs to run

+1

A second service. Something more to run, and to keep in step with the documents.

3.5M

Tokens to load. 2.3 million to write summaries, 1.2 million to embed, before anyone can ask anything.

35 min

Loading time for 437 files, leaving a 70 MB data directory.

The question was whether all that buys better answers, or only ceremony.

The test

Same model, same documents, same ten questions

File search + index

Claude Opus 5 with three tools: list, search for a word, read a file

10 questions
4 min 50 s
Cost
$2.10
Vector database

The same model and loop, no index. Semantic search over OpenViking 0.4.19 with OpenAI's text-embedding-3-small, plus read and word search inside the store

10 questions
5 min 10 s
Cost
$2.10
The test

Prover's internal knowledge base: railway concepts, Swedish regulation topics, internal tool notes

Files
437
Words
~150,000
Figures
279
Questions
10
  • 5 concept questions, worded deliberately away from the document headings.
  • 5 exact questions, asking for a flag, an exit code, a version number, a formula or a configuration key.
  • Each fixed in advance with the file that holds the answer and the exact sentence that answers it. Cost is at list prices.

Semantic search on its own

No model in the loop: is the right file in the first ten results? "Text only" filters to text documents afterwards (this version had no file-type filter); "Word" is a keyword search with one hand-picked keyword per question.

Semantic searchSemantic, text onlyWord search
Most important

Correct file in the top ten

Word search wins
Semantic
5 / 10
Text only
7 / 10
Word
9 / 10

This matters most: the answers were still nine of ten in the next test, so the model made up for the misses itself.

Rank of the first correct file in the top ten results, per question; a dashed cell means it was not in the top ten. Word search records only hit or miss. Semantic search missed four of the five concept questions.
42

Files skipped silently. Their formats could not be parsed by this version.

3→1

Files merged. Three sharing a source citation became one entry. Files were stored under their first heading, so a lookup by filename found nothing.

7 / 10

Top results were images on one Swedish regulation question: a vision model captioned the 279 figures in dense English, while the prose around them is Swedish.

Semantic search alone found the right file in its top ten for half the questions; counting any file in the right topic folder, it reached nine of ten. Speed was never the problem: 0.13 to 0.67 seconds per search, and under 0.05 seconds for word search. What got in the way was loading, shown above; worth knowing before anyone loads documents into a store like this.

With the model: the answers

Both arms end to end, model included, over the same ten questions.

File search + indexVector database

Answers matching the known answer

Tie
Files
9 / 10
Vector
9 / 10

The tenth: "no" for file search, "partly" for the vector database.

Correct file opened or seen

Tie
Files
9 / 10
Vector
9 / 10

Cost at list prices

Tie
Files
$2.10
Vector
$2.10

File search re-sends its index with every question: 963,000 cache-read tokens against 33,000. Cache-read tokens are text the provider keeps ready from earlier requests, billed at a small fraction of the normal rate.

Wall time, lower is better

About even
Files
4:50
Vector
5:10

Model turns, lower is better

File search fewer
Files
37
Vector
42

Input tokens, lower is better

File search fewer
Files
228k
Vector
312k

Output tokens: 19,000 against 21,000.

Nine of ten either way, at the same cost and the same speed. The one difference was the question about the common-cause-failure formula. File search stopped at a similar formula in a neighboring file and never opened the right one. The vector database surfaced the right file at rank five and the model read it, but it still led with the wrong formula. With ten questions, a one-question difference is noise.

Why the answers came out even

Vector Load every document up front→ Find by meaning→ Model reads the top results
Files no loading step→ Search for a word→ Model opens a file, searches again

Semantic search alone put the right file in the top ten only half the time, and the answers were still nine of ten. The model recovered by searching again and by falling back to word search. Whichever store you use, the model needs a plain keyword search beside it.

Vector database: good at

  • Matching a question to a document that means the same thing in other words
  • Speed: 0.13 to 0.67 seconds per search
  • Other jobs, not tested here: per-user document spaces, memory from past sessions, parsing PDF and Word files people upload

Vector database: not good at, in this test

  • Finding the exact file: top ten for 5 of 10 questions
  • Looking a file up by name, or keeping files apart that share a citation
  • Telling you what it skipped: 42 files dropped silently
  • Mixed languages: English image captions crowded out Swedish text
  • Being cheap to own: a second service, kept in step, with a loading bill

What we decided

What Ask Prover runs today

Ask Prover PostgreSQL full-text search+ Word-pattern search→ Paged read

No embeddings, no vector database, no second service to run and no second token bill for loading the documents. The measurement gave no retrieval-quality reason to add one. The ten questions became Ask Prover's standing evaluation, run against the documents the site actually serves.

When we would add one

This is a decision about this version, not a verdict on the technology. A vector index is one database migration away, and it goes in the day an evaluation shows questions that are missed and that only semantic search fixes. The other reasons to run such a store are real but different, and worth re-testing when we need them:

Per-user document spaces

Each person searching their own set of documents.

Memory from past sessions

Extracting what earlier conversations established.

Uploaded PDF and Word files

Parsing the documents people bring themselves.

None of these was tested in this experiment.

Written for Prover Labs. Experiment run 2026-09-10; page dated 2026-09-19. One afternoon's measurement on one corpus with 10 questions: not a benchmark, and not to be read as one.