# How We Test Prisment's Local AI Assistant — and What the Numbers Can't Tell You

> How we test the local AI assistant: question sets frozen first, every answer read, results in words — and what public benchmarks can and cannot show.

- Canonical: https://prisment.io/blog/how-we-test-the-local-ai-assistant/
- Published: 2026-09-26
- Updated: 2026-10-07
- Checked against: Prisment v2.0.6
- Author: Prisment Team (Documentation & Developer Tooling)
- Publisher: Prisment, https://prisment.io/

**We test the local assistant by freezing question sets before the first run, reading every answer ourselves, and describing results in words — never as a score we gave ourselves — and we publish two public retrieval benchmarks as disclosures, with what they cannot show.** This is the evidence behind [the assistant article](https://prisment.io/blog/local-ai-assistant-for-your-documents/); read that first for what the assistant does.

## 1. Why words, and not a percentage?

**On our own folders and our own questions we are the judge, so we do not grade ourselves.** A figure we award our own product is not something a reader can check, and an early version of this work showed why: we published bar charts for the bare model, plain keyword search and the full pipeline, then found that the scorer behind them was broken, and withdrew them (section 4). What replaced them is a description of what each model does — the table in [the assistant article](https://prisment.io/blog/local-ai-assistant-for-your-documents/) headed *What should you expect from each model?* — and official-scorer results on public datasets, where someone else wrote the rules.

## 2. How does a test run?

**Write the questions, freeze them, run each several times on both models, read every answer, and only then fix — and test the fix on a fresh set.** A set that has been read cannot fairly test a change made from reading it.

![How a test runs: a question set is written and fingerprinted before its first run; it is run several times on a small and a larger model at two context windows; a person reads every answer and sorts it into right, wrong or unsafe; fixes are made from what was read; the fix is tried on a new set, never on the one that was read](https://prisment.io/images/blog/local-ai-testing-method.png)

- **The questions** are fictional workspaces of design notes, a runbook, API limits, source code in several languages, a to-do list and meeting notes: facts, questions spanning documents, questions the folder cannot answer, summaries, reviews, to-do and code questions.
- **The models** are Llama 3.2 3B and Qwen2.5-Coder 14B, 4-bit, run through Ollama on one Apple laptop, each question in a fresh conversation with the same seed per trial where an app accepts one.
- **Reading, not keyword checks.** Keyword checks pass wrong answers and fail right ones, so we read every answer and count a contradiction as wrong.

![The assistant panel after a question about a call that keeps failing: the model's name and timing, the line Context — 5 files, a Read line naming docs/observability.md, and an answer citing docs/networking.md — the lines a person reads each answer against](https://prisment.io/images/blog/local-ai-assistant-answer.png)
- **Other apps, at their defaults,** were given the same questions on the same machine — not to rank anyone, but to see what a reader gets elsewhere and where ours falls short. The clearest lesson was speed: we are slower, reviews on the larger model most of all.

## 3. What did reading the answers change?

**Reading found defects no keyword check would have, and each one is now fixed in the app rather than hidden in a prompt.**

- **A right answer was being thrown away.** The two lines it came from sat either side of a blank line, the app could not match it to them, and after one retry it replaced a correct answer with "could not produce an answer".
- **The small model kept re-reading files it had already been given whole,** so one answer took several trips to the model. The app now answers those requests itself.
- **A few replies repeated one paragraph until the model ran out of room** — the slowest answers we saw. The app now stops a reply at its first repeat.
- **Reviews of Python missed a leftover `print()`;** the same check now covers Python, Kotlin, Go, Rust and PHP.

Not fixed: the model's own count can contradict the app's correct one printed beneath it, and an answer can state two figures correctly and still conclude they match. The app says so under the answer when it sees either.

## 4. What did we get wrong, and withdraw?

**Our first scorer used phrase matching, and it produced a headline that was not accuracy.** It counted "local disk" inside "Nothing is written to local disk" as a contradiction, treated a correct answer ending "I could not find any further information" as a refusal, and gave an answer with no citations a perfect path score. We withdrew every figure it produced and now report lexical fact coverage, path validity and source recall separately, and say that a phrase matcher cannot certify that a claim is supported. The longer account is in the closed notes below.

## 5. What did leaving beta require?

**A release check we wrote down before running it, on new folders the app had never been adjusted to: no file changed by a question, no invented file names, versions and missing answers told apart every time, and answers, summaries and reviews right with only rare misses.** The first run found problems — the larger model offered an edit to a question, and a summary of a document named without its folder was thrown away — which we fixed and checked on a second, fresh set, reading every answer. **Qwen2.5-Coder 14B passed at the context window the app sets by default; Llama 3.2 3B did not.** So the recommended model is the larger one at the default window, and the small one is described rather than hidden.

## 6. What do the public benchmarks show?

**They show how well Prisment's search stage ranks documents on two public datasets, and nothing about how correct the assistant's answers are.** We ran every judged test query in SciFact and NFCorpus with the official BEIR scoring, and a 20-question pilot of ALCE for answer correctness. The tables, the settings and the caveats are in the disclosure below, closed by default.

<details>
<summary>Public benchmark results: BEIR and the ALCE pilot</summary>

[BEIR](https://github.com/beir-cellar/beir) provides public retrieval datasets. On 26 September 2026 we indexed every document and evaluated every judged test query in SciFact and NFCorpus using BEIR 2.2.0 scoring. These are two complete dataset runs, not the full BEIR suite or a leaderboard submission.

| Dataset and arm | Test queries | nDCG@10 | Recall@5 | MRR@10 |
| :--- | ---: | ---: | ---: | ---: |
| SciFact — lexical search | 300 | 0.66299 | 0.72028 | 0.62461 |
| SciFact — initial context pack | 300 | 0.66299 | 0.72028 | 0.62461 |
| NFCorpus — lexical search | 323 | 0.31565 | 0.12187 | 0.51461 |
| NFCorpus — initial context pack | 323 | 0.31268 | 0.11834 | 0.51151 |

SciFact contained 5,183 documents and NFCorpus 3,633. Titles and text became synthetic Markdown files; relevance labels never entered retrieval. Lexical search returned up to 1,000 hits. The context-pack arm retained the production candidate limit and used a 60,000-token budget to expose admission order; it was not an 8K answer-generation test. Ranking metrics are not accuracy percentages, and the two datasets have different judgments and difficulty. Without matched baseline systems, we cannot reliably call these results excellent or industry average.

Both corpora had zero reference edges, so this test establishes no graph benefit. Sixteen NFCorpus questions received empty initial packs: fifteen also had no lexical hits, and one had two lexical hits excluded by the pack.

We also tested an optional local semantic model, BGE-M3, and combined its ranking with Prisment's lexical ranking. We selected one combination on SciFact train and NFCorpus development data, froze it, and then ran the untouched test splits once.

| Complete test split | Current lexical | Experimental local hybrid | Published BM25 reference | Published SPLADE reference |
| :--- | ---: | ---: | ---: | ---: |
| SciFact | 0.66299 | **0.69892** | 0.665 | 0.699 |
| NFCorpus | 0.31565 | **0.34597** | 0.325 | 0.345 |

The hybrid beats the cited BM25 number on both datasets and the cited SPLADE number on NFCorpus. On SciFact it rounds to the same 0.699 shown in the published table but is lower by 0.00008 at full precision, so we do not claim that it beat SPLADE on both. This is a measured direction, not a shipped feature: it needs a separate 1.2 GB local embedding model, cached document vectors, memory and speed controls, and more tests before it belongs in Prisment. The comparison uses Table 2 of [Resources for Brewing BEIR](https://arxiv.org/html/2306.07471#S3.T2), a dated published reference rather than a live leaderboard rank.

[ALCE](https://github.com/princeton-nlp/ALCE) evaluates answer correctness and citation quality. Our separate ASQA pilot selected 20 of its 948 questions by a fixed hash ordering, then ran five seeds per question with Llama 3.2 3B Q4_K_M, Ollama 0.33.3, an 8,192-token window, temperature 0.2 and a 512-token output cap. Each question's supplied GTR top-100 passages became an isolated workspace. This evaluates selection and answering within that candidate pool, not full-Wikipedia retrieval.

| ALCE pilot measure | Result | What it establishes |
| :--- | ---: | :--- |
| STR-EM | 27.40% | Average reference QA-pair coverage by the official string matcher |
| STR-HIT | 17.00% | Responses matching every reference QA pair |
| Citation support | Not graded | Existing source paths do not prove that claims are supported |

All 100 responses were retained, including failures. They represent 20 unique questions, not 100 independent questions. The pinned upstream scorer's first-newline rule truncated 18 outputs for scoring; originals remain preserved. Contradictory answers and exposed lookup commands remain failures. AutoAIS citation entailment, QA-model scoring, MAUVE and ROUGE were not run; this is a partial ALCE evaluation, not a full ALCE score. Our former internal path-existence metric was not ALCE citation precision either.

The pilot also differs from the app: it uses a planned lookup allowance and fixed output cap, while the app adjusts both from the actual prompt. It omits UI, desktop, history and adaptive-state behavior. Its scores cannot establish the complete app's quality or this model's best achievable result. No public leaderboard result or registered submission is claimed.

These are retained baseline runs; they have not been rerun after every subsequent software fix. nDCG measures how highly relevant documents rank; Recall measures how many judged relevant documents were retrieved; MRR measures how early the first relevant result appears. None is a percentage of correct assistant answers.

</details>

## 7. What can a few dozen questions not show?

**They cannot show how your folder will go.** The questions are ours, the folders are fictional, the answers were graded by us, and the other apps ran at their defaults rather than tuned; it is one computer, and our repository is private, so these runs cannot be repeated outside it. A wrong answer in your own folder — the question and the passage it should have used — tells us more than any general score.

<details>
<summary>Test notes: how we observed this, and what it cannot show</summary>

- **Setup:** Llama 3.2 3B and Qwen2.5-Coder 14B, both Q4_K_M, in Ollama 0.33.3 on an Apple M4 Pro with 24 GB; an 8,192-token window, temperature 0.2, each question in a fresh conversation, the same seed per trial wherever an app accepts one. The apps ran one at a time.
- **Questions:** two sets, each a fictional workspace of design notes, a runbook, API limits, source code (Java in the first; Python and Kotlin in the second), a to-do list and meeting notes, asked facts, questions spanning documents, questions the folder cannot answer, summaries, reviews, to-do questions, code questions and comparisons. Neither set was edited after its first run.
- **Order:** the fixes came from reading the first set's answers, so they were tried on the second set instead — a set that has been read cannot fairly test a fix made from it. The release check followed the same rule with two more sets: six fictional folders in English, French and Spanish, with TypeScript, Python, Java, Kotlin, YAML and JSON, each frozen before its first run, at both 4,096- and 8,192-token windows.
- **What it cannot show:** a few dozen questions, one computer, folders we wrote and answers we graded, and other apps at their defaults rather than tuned.

</details>

<details>
<summary>Test notes: how the answer checks were tried, 26–27 September</summary>

Several seeded trials per case, through the same code the panel runs, with both models. *Development* cases were used to build the checks; one case was *held out* — a design-decision record written and fingerprinted before its first run. With both models the checks did what they are for: summaries named their sections, reviews started from the problems the app found, answers named their lines, and a figure no supplied line contained was flagged. The small model answered much sooner.

The held-out case's first run failed: the record's own *Context* heading ended the app's reading of it after one line. We fixed that general defect without changing the case's checks, so the case has now informed one fix and the next claim needs a new one. An older 4B model improved on the source questions — and on the way the checks briefly gave its wrong *"Maple signed in 2018"* a source; that is why a figure moved onto another subject is now reported and never sourced.

These are small, fixed cases on one Mac, and we wrote and graded them. They show the checks working, not how every folder will go.

</details>

<details>
<summary>Test notes: internal failures and the corrected scoring approach</summary>

The 26 September audit used `llama3.2:3b` Q4_K_M in Ollama 0.33.3, an 8,192-token window and temperature 0.2. Six questions ran five times through three arms — bare model, BM25 retrieval and the full pipeline — giving 90 retained answers. These were service-level trials; separate browser checks exercised the panel and graph.

| Observation | What it means for a reader |
| :--- | :--- |
| Most full-pipeline dead-letter answers invented `docs/dead-letter-queue.md`; the answer was in `docs/networking.md` | Correct-sounding prose can cite the wrong source |
| A secrets answer substituted environment variables for the documented Vault policy while naming a real file | A real path does not establish claim support |
| Most raw summary replies began with edit-format blocks | Answer-only requests need enforcement in the app, not just prompt instructions |
| A panel summary omitted formats and blurred Save with Export | Short answers must still preserve important distinctions |
| A panel suggestion missed duplicated sections and proposed material already present | Suggestions require evidence of an actual issue |
| The tested retry-count edit was usually, not always, recovered | A successful demonstration does not guarantee every edit will apply |

**Why the original accuracy headline was removed.** The original scorer used phrase matching. It counted “local disk” in “Nothing is written to local disk” as a contradiction, and treated a correct answer ending “I could not find any further information” as a refusal. Its “citation precision” checked whether paths existed, not whether their contents supported the answer. Answers without citations received a perfect score on that path metric.

We first published scores for the bare model, BM25 and the full pipeline from that rubric. **We withdrew them: they were outputs of a flawed rubric, not accuracy.** Re-scoring the same answers must be reported separately from improving the assistant; changing a scoring rule is not a model-quality gain. The revised reporting distinguishes lexical fact coverage, path validity, source recall and task-completion proxies. Missing citations are not applicable for path validity. Semantic claim support still requires its own evaluation; a phrase matcher cannot certify it.

</details>

<details>
<summary>Efficiency measurements and what we still need to verify</summary>

The pilot needed several model requests per answer, many of them follow-up lookups that found nothing new — missing files, invalid line ranges, partial reads — and most of each prompt was already cached. Its answer times were workload observations on one Mac, not a speed promise. The next comparisons must measure whether better evidence selection and fewer unproductive follow-ups improve supported answers within the same model, memory and time budget. Faster unsupported answers would not be a quality improvement.

The follow-up fixes retain passages found during searches when checking an answer, retry missing source references, and ask comparisons to cover each option's purpose and result without repeating the same facts. These checks can catch missing or unknown references; they do not prove that every cited passage supports every claim. Another fix ends a streamed lookup block after the four complete commands the app can use, instead of letting the model generate hundreds of commands that would be discarded. Since 27 September, a reply that has started repeating the same sentence is stopped at its first repeat, and a request to re-read a file already given in full is answered by the app without another trip through the folder. Ordinary answers and document edits keep their existing output allowance. The public baseline above predates these changes.

Graph retrieval needs separate evaluation on documents with real links: success on a generic text-retrieval dataset alone cannot establish the value of a workspace graph. Editing, undo, workspace switching and accidental-write prevention also need app-level tests beyond a question-answering benchmark. Passage-level citation support is still not graded. The internal harness is not yet publicly reproducible, and no perfect-score claim is justified by the current evidence.

</details>

## 8. What are we improving?

**Keeping useful passages, cutting unproductive searches and slow answers — reviews first — and correcting a count or a conclusion the model states against the app's own.** A faster wrong answer is not success. A local semantic-search experiment now beats the cited keyword-search scores on both public retrieval tests, but it is not part of the product yet. If the assistant gets something wrong in your folder, [tell us](https://prisment.io/help/#getting-help); that report is worth more than a general score.

## Questions this article answers

### How does Prisment test its local AI assistant?

With question sets written for the purpose and frozen before their first run, several trials per question on both recommended models and both context windows, and a person reading every answer rather than trusting keyword checks. Results are described in words per model, never as a score we gave ourselves. The method, what reading the answers changed and the public benchmark runs are in this article.

### Why does Prisment not publish an accuracy percentage for the assistant?

Because on our own folders and our own questions we are the judge, and a percentage we gave ourselves is not a measurement a reader can rely on. We published one early, found that its scorer counted a phrase match as a contradiction, and withdrew it. We describe what each model does in words, and publish only official-scorer results on public datasets.

### Did Prisment submit to a public leaderboard?

No. We ran two complete public retrieval datasets, SciFact and NFCorpus, with the official BEIR scoring, and a 20-question ALCE pilot. These are dated baselines for the search stage and a partial answer evaluation — not a leaderboard submission, and not an accuracy rating of the whole assistant.

### Can I repeat these tests myself?

Not the ones on our own folders: the repository that holds the questions, scorers and raw answers is private, so this article publishes the settings and the rules instead. The public benchmark datasets and scorers it cites are open, and you can run the same ones.
