Evaluation results
Scores from the last published run.
These are the figures already stated on the site. The date of the last CI run is not recorded yet.
Placeholder for the run date and release id. Scores are the published numbers. CI should overwrite data/evals.json on each gated release.
Published result, not a live CI badge: 1,213 Langfuse-scored decisions, all gate.correct = 1. CI should replace this status when it overwrites the file.
How to read these numbers
Three terms. One line each.
- RAGAS
- A score out of 10 for whether the answer stays grounded in the text that was retrieved. The published average is 9.1. A RAG release below 9.0 does not ship.
- Recall@10
- How often the correct GL code is among the 10 historical lines the retriever returned. 100% means it was in that top 10 on every line in the golden set.
- gate.correct
- 1 means the gate was right: an auto-posted code was correct, or an uncertain line was sent to a person. All 1,213 scored decisions are 1.
Sample trace
One invoice line, step by step.
Anonymized and simplified from a real trace.
Placeholder. Vendor, amount, GL codes, and the threshold in this walkthrough are synthetic. Replace data/sample-trace.json with a redacted trace before treating any code as real.
01Retrieval query
- Vendor
- Sample Vendor A (synthetic)
- Description
- Monthly platform subscription (synthetic)
- Amount
- 1,250.00 (synthetic)
- Cost centre
- CC-100 (synthetic)
What to notice. The query is built from the invoice line — vendor, description, amount, cost centre — not from a blank prompt.
02Top 3 retrieved lines
- 1
- Sample Vendor A — Platform subscription — GL 6100 — confirmed (synthetic)
- 2
- Sample Vendor A — Software licence — GL 6100 — confirmed (synthetic)
- 3
- Sample Vendor B — IT services — GL 6240 — confirmed (synthetic)
What to notice. The model sees confirmed historical lines. Rank is what the retriever returned, not what the model wished were true.
03Prompt summary
- Instruction
- Propose one GL code and a confidence score from the retrieved confirmed lines. Include the reasoning. Do not propose a code none of the examples support. (Summary of the prompt, not the production template.)
What to notice. The model is asked to generalise from those confirmed lines and to return a code plus a confidence. It is not asked to memorise a rule.
04Proposed code
- Proposed GL
- 6100 (synthetic)
- Confidence
- 0.94
What to notice. 0.94 is the model's confidence on this line. It is a proposal until the gate accepts it.
05Gate decision
- Confidence
- 0.94
- Example threshold
- 0.90 (synthetic)
- Decision
- Auto-post — 0.94 is at or above the example threshold
What to notice. The gate decides the write, not the model. The 0.90 threshold is part of this simplified example, not a published production setting.
06Write-back confirmation
- Written GL
- 6100 (synthetic)
- Corpus
- Appended to the tenant's confirmed lines (simplified)
- Trace
- Retrieval, prompt, decision, and confidence recorded in Langfuse
What to notice. Only the confirmed code is written back. The corpus grows from a decision the gate accepted, not from the model's first guess.