Skip to content
Orvenant
Orvenant

Evaluating an AI customer assistant

Evaluating an AI customer assistant

Evaluate an AI customer assistant against representative questions, access boundaries and escalation behavior. Answer fluency is not enough to establish reliability.

Define the assistant’s job and authority

Choose a narrow customer task, such as explaining published product instructions or helping a signed-in customer locate an order. List approved knowledge sources, unsupported topics and actions that require confirmation. Distinguish reading information, drafting a proposed change and executing that change. These need different permissions. Name the owner who resolves conflicting policies and the person who receives escalations. If the business cannot define the correct outcome, a model comparison will not settle the uncertainty.

Build a test set from the customer journey

Use permitted, de-identified questions that represent the expected journey. Include straightforward requests, incomplete questions, misspellings, follow-up questions and cases where the correct response is to ask for clarification. Add outdated and conflicting documents, unavailable records and requests outside the stated scope. Each case needs the acceptable behavior, the source of truth and the severity of a wrong answer. Avoid grading only whether a response resembles a preferred sentence; several wordings may be correct.

Score usefulness and evidence separately

Assess whether the answer is correct, addresses the actual question and communicates uncertainty appropriately. Check cited material against the claim it supposedly supports. A valid-looking source link does not make an unsupported conclusion correct. For policy or account questions, verify that the assistant uses the relevant version and context. Record failures by category so a good average does not conceal repeated mistakes in one important journey. Keep the original question and a reviewable result for each test.

Test permissions beyond the conversation

Try to retrieve another customer’s record, continue after access is revoked and request an action outside the signed-in user’s authority. The application and tools must enforce access independently of what the assistant says. OWASP describes prompt injection through direct requests and external content, so include hostile instructions inside retrieved documents as well as in chat. Check that such text cannot authorize an export, change a destination or override the tool’s allowed scope. Log enough for investigation without collecting unnecessary private data.

Agree the release decision before scoring

Separate critical failures from quality improvements. An unauthorized action or cross-customer disclosure should block the relevant capability even if most ordinary answers pass. Set agreed thresholds for the lower-risk tasks, response time and escalation completion. Compare the assistant with the current process: can a person actually resolve the handoff with the context provided? Roll out a restricted scope first when evidence is limited, and keep a clear route to a human for unresolved requests.

Keep evaluation connected to change

Version the test set alongside the prompt, model configuration, knowledge and tools. Re-run relevant cases when any of these change. Review permitted real failures to add missing scenarios, while keeping a stable test group for comparison. Assign someone to inspect recurring escalation reasons and stale documents. Provide a control to disable a risky action or the entire assistant without taking down the main customer journey. Evaluation is part of operating the service after launch.

Questions about this guide.

How many questions are enough for an evaluation?

There is no universal count. Cover every permitted capability and each important failure type, then expand where errors cluster. A large set of easy questions cannot compensate for missing access-control or escalation cases.

Can an AI model grade all of its own answers?

Automated grading can help sort results, but calibrate it against human judgments and inspect serious failures directly. Keep business rules, access checks and action outcomes independently testable.

Does connecting a knowledge base prevent invented answers?

No. Test whether the assistant retrieves the right material, interprets it correctly and declines to invent an answer when evidence is absent. A source connection is an input to evaluation, not proof of correctness.

Sources and further reading

Sources support the stated technical context. The planning recommendations are Orvenant's assessment.

PUT THE GUIDANCE TO WORK

Bring us the problem. We’ll work through the next step.

Share what exists today and what you want to change. We can discuss the scope, dependencies and a practical way forward.

Privacy settings

Optional Google Analytics measures page visits and contact-option use. Rejecting keeps Analytics off. Accepting allows analytics cookies. Advertising features stay off, and enquiry contents and contact details are not sent to Analytics.

Withdrawing stops future collection and removes this site’s accessible Analytics cookies. It does not erase information already received by Google. Privacy notice · Cookie notice