Work

Synthetic data

Ask your documents

Answers from your own documents, offline, with the source.

A question about a synthetic company handbook returns the exact sentence with file, page and section; a question the documents do not answer is shown returning "Not found in the documents." with the reason.

Who it is for

  • Small offices with handbooks, contracts and procedures that must not leave the building.

The challenge

Handbooks, procedures and supplier terms hold the answer to everyday questions, but finding the right paragraph takes time. An answer without its source is hard to check, and many offices cannot send internal documents to a cloud service.

What we built

  1. Reads a folder of PDF (with a text layer), Markdown or text files, and cuts them into short passages that keep their file name, page and nearest heading.
  2. Searches with BM25, a standard keyword-ranking formula. No AI model, no embeddings, no network.
  3. Tries to refuse unsupported questions, under defined rules. If a name in the question appears in no document, or the best passage does not contain at least 60% of the question's key terms, the answer is "Not found in the documents.", with the missing terms listed. The rules are not perfect: see the false answer in the results below.
  4. Answers with the best-matching sentence, copied word for word, with its file, page and section. Every answer and every refusal shows a plain reason.

Optional: an Ollama model can reword the answer. It is off by default and takes no part in the search or the "not found" decision. When it is on, the question and the retrieved passage are sent to the Ollama address in the OLLAMA_HOST setting: by default this computer (localhost), but if that setting points to another machine, they leave this one. The rewording is model output and is labelled as such; wording that adds numbers or names not in the passage or the question is not shown, and the cited quote is always shown next to it.

An answer copied from a synthetic company document, shown with its file, page and section.
Answer with its source passage · synthetic data, 5 October 2026.
The reply "Not found in the documents." with the question's missing terms listed.
Not found in the documents · synthetic data, 5 October 2026.

Results

8 invented documents about an invented company (68 passages), with questions whose answers are known, including questions the documents do not answer. Per the developer's record, settings were tuned on the development set only and the hold-out questions were never used for tuning, so the hold-out is the fairer number. Both sets ask about the same 8 documents; the hold-out holds back questions, not documents. Measured 5 October 2026; figures from the project's evaluation summary.

SetSettingsAnswered correctlyCorrect "not found"False answers to unanswerable questionsOverall
hold-outdefault25 of 31 (80.6%)11 of 12136 of 43 (83.7%)
hold-outalways answer (baseline)27 of 31 (87.1%)0 of 121227 of 43 (62.8%)
developmentdefault28 of 30 (93.3%)11 of 12139 of 42 (92.9%)
developmentalways answer (baseline)28 of 30 (93.3%)2 of 121030 of 42 (71.4%)
  • The baseline uses the same search with the name and 60% refusal rules switched off; it still says "not found" when no passage shares any searchable word with the question. On the hold-out set it gave a wrong answer to all 12 unanswerable questions; with the "not found" rules, 11 of 12 were refused correctly. With only 12 such questions, all about the same 8 documents, this count is not a reliable rate.
  • The trade-off: 5 answerable hold-out questions were wrongly marked "not found", so correct answers dropped from 27 of 31 to 25 of 31.
  • One false answer remains: asked "Who is the safety officer?", the tool returned the rule that mentions the safety officer.
  • These are small synthetic sets (85 questions over 8 documents) and do not establish accuracy on your documents.

Technical background

Python, pypdf, snowballstemmer (word stems for search), pandas, Streamlit; tests with pytest and Playwright; packaged with Docker.

Limits

  • Word matching, not understanding. A question that uses other words than the document ("USB drives" against "USB storage devices") can get "not found".
  • Near-miss questions about a detail next to a covered topic can still get a wrong answer.
  • Scanned PDFs need OCR (text recognition), which is not included. Only PDFs with a text layer are read. Tables lose their layout.
  • One language at a time. An English question does not find an Italian passage, and the other way round.
  • Input limits: files over 20 MB, uploads over 100 MB or 200 files, and PDFs over 500 pages are refused. A hostile PDF within these limits can still be slow.
  • Single-user demo: document indexes are kept in the app's memory for up to an hour and are shared by everyone using the same server, so it should not be exposed to other users without authentication.
  • Not legal advice: always read the cited passage in the original document.

Source: snello-docqa README and reports/summary.json and eval_holdout.json, measured 5 October 2026 at commit fc6bcbe (measured on 5 October 2026 at commit fc6bcbe). Code is private, so it is not linked.

Related work

Synthetic data

doc2data

Invoice and receipt PDFs and scans to checked, structured data.

Test results and limits

0 field errors on 3 clean synthetic sets of 16 documents; 92.9% (131 of 141) and 91.4% (127 of 139) of fields on degraded synthetic scans (measured 2 October 2026).

Read the project: doc2data
Synthetic data

FatturaPA reader and SdI receipt explainer

Italian e-invoices (XML and signed .p7m) to one checked spreadsheet, where the implemented checks flag defined problems, with the SdI rejection code where one applies; SdI notices explained in plain words.

Test results and limits

On 3 synthetic sets of 94 invoices: 28 of 28 planted faults detected in each, 0 of 66 clean invoices flagged by mistake (measured 5 October 2026).

Read the project: FatturaPA reader and SdI receipt explainer

Bring your workflow into focus.

Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.

Start a project

Optional AI assistance

The project now includes optional Groq assistance, disabled by default. The original local checks, calculations and source passages remain the basis for results. AI output needs human review.

After administrator setup, explicit consent is required before text is sent to Groq. Explanation tools show a context preview; invoice extraction can send OCR or document text after its consent step. API credentials stay on the server. Owner access, upload limits and a shared request budget protect the local demos.

This static website does not run the assistant or accept document uploads. Earlier demo recordings and measurements cover the original local workflows; the Groq integration has not yet been evaluated with live requests.

SNELLO / TOOLS

This explains the tool; it is not a live AI session. No document is uploaded.

Our working method
Full diagram

Read the project

Discuss your project

Send up to 3 samples or describe one process, the tools you use and the result you need.

snello.contact@gmail.com

Go to the Contact page

Navigation