
A website assistant needs to handle missing information as well as straightforward questions. Hostinger’s discussion of everyday AI limitations provides the starting point; the practical task here is to create an evaluation set before exposing your support assistant to visitors.
1. Define expected behaviour first
Collect representative questions from your support history after removing personal information. Include short wording, everyday phrasing and typing mistakes. For every question, identify the approved page or procedure that supports a response.
Write the essential answer points yourself. Asking the model under test to create its own reference answers can turn the exercise into a check of self-consistency. Date the supporting material and flag questions that require an authenticated business system rather than general documentation.
Add cases that should not receive an immediate factual answer: an unspecified product, a request with insufficient context or a topic outside the knowledge base. Define the clarification or handoff that would be appropriate. Admitting that information is missing can be the correct outcome.
2. Run a repeatable evaluation
Keep the configuration and documents fixed for a test run. Record the model version when available, the instructions and the knowledge snapshot date. Start a fresh conversation for independent cases, then include separate multi-message scenarios to check how the assistant follows context.
Adapt this worksheet to your actual service:
| Case | Expected behaviour |
|---|---|
| Documented question | Faithful answer and relevant source |
| Ambiguous product | Clarifying question |
| Missing information | Clear limitation without invention |
| Personal request | Appropriate approved channel |
Save the full response alongside each assessment. Repeat sensitive cases because one successful response does not establish consistent behaviour. Evaluate French and English separately if visitors can use both languages. A correct answer in one language should not automatically count as a pass in the other.
3. Set a release decision rule
Separate presentation problems from substantive errors. A verbose answer and an invented procedure have different consequences. Choose acceptance criteria before reviewing the results, including an explicit rule for failures that would prevent you from enabling the assistant.
Classify each failure in a way that suggests a fix: missing documentation, irrelevant retrieval, unsupported answer or poor routing of the conversation. Change one cause at a time and rerun both the affected cases and previously passing cases. Reserve some questions for a final check instead of repeatedly tuning against the entire set.
Keep the evaluation after launch. Run it when you change the model, instructions or documentation. Add real visitor failures in a redacted form, with reviewed expected behaviour. This produces a growing check of the service you actually operate, rather than a favourable impression from a handful of friendly demonstration conversations.
Sources: Hostinger