Evaluation
An evaluation is a test of your assistant: a set of questions with the documents that answer them. Run it before and after a change, and the numbers tell you whether answers got better or worse.
Updated
On this page
Why evaluate
#Changing the model, the number of documents retrieved or the similarity threshold can quietly break answers that worked yesterday. An eval set turns "it seems fine" into numbers you can compare, so you change settings with evidence.
- Before switching to a cheaper or newer model.
- Before changing retrieval settings such as top k or minimum similarity.
- After adding or rewriting a lot of knowledge.
- When customers report wrong answers, to check the fix and keep it fixed.
Evaluations run on your own AI provider key: a retrieval-only run costs only query embeddings; a full run also generates answers and, if you choose, has a judge model score them. Only workspace owners and admins can create sets and start runs.
Build an eval set
#Choose New eval set and give it a name, for example Regression. Aim for 50 to 200 questions (up to 1000 per set), and make at least 15% of them questions your knowledge cannot answer, so you also test that the assistant declines instead of guessing.
| Field | What it is | Required |
|---|---|---|
| Question | What a visitor would type | Yes |
| Conversation before the question | Earlier turns, for follow-up questions (up to 20) | No |
| The knowledge answers this | Off for questions the assistant should decline | On by default |
| Expected answer | A reference answer, used by the correctness judge in full runs | No |
| Gold documents | The documents that contain the answer: search by title or URL, or paste a URL, file name or doc_ id. Add a span to require a chunk with that exact text | No, but needed for retrieval scores |
| Tags | Labels such as shipping or hindi, to read results by group | No |
| Key | A stable id, so cases can be matched between runs | No |
There are four ways to add cases:
- 01
By hand
Add case on the set page. Best for your most important questions.
- 02
From real conversations
In a conversation, press Add to eval set under any reply. The question, earlier turns and the cited sources are filled in; remove personal data before saving. This is the best source of realistic questions.
- 03
Generate from your knowledge
Generate cases writes questions from your documents with the bot's own model (default 30 answerable and 6 unanswerable), tagged synthetic. Review them: generated cases count half as much as hand-written ones in every score.
- 04
Import a JSONL file
One case per line, up to 2 MB. Choose to add to the existing cases or replace them. Problems are shown line by line before anything is saved. Export gives you the same format.
{"key": "ship-canada", "question": "Do you ship to Canada?", "expected_answer": "Yes, in 5 to 8 business days.", "gold_documents": [{"external_ref": "https://help.yourshop.com/shipping", "must_contain": "5 to 8 business days"}], "answerable": true, "tags": ["shipping"]}
{"key": "refund-window", "question": "How many days do I have to return an item?", "gold_documents": [{"external_ref": "https://help.yourshop.com/returns"}], "tags": ["returns"]}
{"key": "no-crypto", "question": "Can I pay with Bitcoin?", "answerable": false, "tags": ["payments", "unanswerable"]}
# lines starting with # and blank lines are ignoredOnly question is required. Tags use lower-case letters, digits and _ : . - (up to 20 per case). Unknown fields are rejected so typos do not pass silently.
Run an evaluation
#- 01
Start run
On the set page, press Start run. The run uses the bot's live settings. You can try different retrieval settings or temperature for this run only; the bot itself is never changed.
- 02
Choose the mode
Retrieval only checks whether the right documents are found; it costs query embeddings only and finishes in seconds. A full run also writes answers and scores them.
- 03
Choose the judge
For full runs, a judge model scores groundedness, relevance and correctness: the bot's own model, another key and model, or no judge. Keep the same judge for runs you want to compare.
- 04
Set concurrency
How many questions run at once (default 4). Use 1 or 2 on low-tier keys to avoid rate limits and to see realistic response times.
Tip: When you change the model or retrieval settings on the Settings tab and the bot has a set with at least 20 cases, a Run eval first button appears. It runs the set with your unsaved changes and opens the comparison, without saving anything.
Read the results
#Scores are shown from 0 to 100% with a 95% confidence range and the number of cases behind them. A small set gives a wide range: treat small differences with care.
| Metric | What it tells you | Good direction |
|---|---|---|
| Recall@k | How often at least one gold document is in the top k retrieved chunks | Higher |
| MRR | How high the first correct document ranks (1.0 means first every time) | Higher |
| Candidate recall@20 | Whether the right document was found at all in the top 20. High here but low recall@k means it was found but cut off: raise top k | Higher |
| Groundedness | Share of claims in the answer supported by the retrieved text. A contradicted claim flags the case | Higher |
| Answer relevance | Whether the answer addresses the question | Higher |
| Correctness | Agreement with your expected answer (only cases that have one) | Higher |
| No-answer accuracy | How often the assistant declines questions your knowledge cannot answer | Higher |
| False decline rate | How often it declines questions it should have answered | Lower |
| Latency, cost per turn | Speed and spend on your key | Lower |
A PARTIAL badge means the judge could not score every case; judged numbers are then less reliable.
Compare two runs
#Tick two runs of the same set and press Compare selected, or use Compare with baseline on a run. You see both runs side by side, the change in each metric in green or red, and the cases that got worse.
| Gate fails when | Threshold |
|---|---|
| Recall@6 drops | more than 2 points |
| MRR drops | more than 0.03 |
| Groundedness drops, or a new contradicted claim appears | more than 2 points |
| No-answer accuracy drops, or false decline rate rises | more than 3 points |
| Answer relevance drops | more than 5 points |
| Cost or slow-answer time rises | more than 15% or 20% (warning only) |
A drop that stays inside the baseline's confidence range is shown as within range and does not fail. When a run is the new normal, use Accept as baseline and note why.
Worked examples
#Online shop
Switching to a cheaper model
Check that a cheaper chat model answers shipping and returns questions as well as the current one.
- 01
Build the set
80 questions: 40 from real conversations (Add to eval set), 30 generated from the help centre, 10 unanswerable such as "Do you sell cars?". Tag them shipping, returns, payments.
- 02
Run the baseline
A full run on the current model with the bot's model as judge. Accept it as baseline.
- 03
Try the new model
Pick the cheaper model on the Settings tab and press Run eval first instead of Save.
- 04
Decide
If groundedness and correctness hold and the gates pass, save. If returns questions regressed, read those cases first: often the returns page needs a clearer sentence.
Clinic or bookings
Making sure the assistant declines medical questions
Confirm that questions outside the clinic's published information are declined, not guessed.
Build a set with 30 answerable questions about timings, doctors and fees, and 20 unanswerable ones such as "Is this rash serious?" or "What dose of paracetamol for a child?". Watch no-answer accuracy (should be close to 100%) and false decline rate (should stay low, so real questions are still answered).
Software company
Tuning retrieval for a large help centre
Find out whether answers fail because the right article is never found or because it is found and cut off.
Run retrieval only with the current top k. If candidate recall@20 is high but recall@6 is low, the right article is found but ranked too low: try a run with a higher top k or hybrid search turned on. If candidate recall is low as well, the fix is in the content: split long articles and use the words customers use.
Coaching institute
Checking Hindi and Hinglish questions
Make sure students who ask in Hindi or Hinglish get the same answers as those who ask in English.
Write each key question three ways, for example "What are the JEE batch timings?", "JEE batch ka time kya hai?" and the same in Devanagari, with the same gold document. Tag them english, hinglish and hindi, then compare recall and groundedness by tag in the results.
Tips
#- Start small: 30 good questions from real conversations beat 300 generated ones.
- Always include unanswerable questions, at least 15% of the set.
- Add a gold document to every answerable case, or retrieval scores cannot be computed.
- Give cases a key, so runs can be compared case by case.
- Keep the judge the same across runs you compare.
- Add every wrong answer a customer reports as a case. Your set becomes your safety net.