Read the conversation, understand the failure, and give engineering an example they can act on. Then measure whether the change improves the task your customer came to complete.
Product managers, designers, and teams reviewing AI quality
Start with the experience
Review the customer conversation and the answer or action they actually received.
Add a clear judgment
Record why the result was useful, wrong, incomplete, or confusing rather than leaving an unexplained score.
Follow the outcome
Connect the release to task success and feedback; do not substitute model activity for product impact.
At a glance
| Product question | Useful evidence |
|---|---|
| Are people getting help? | Task outcomes plus reviewed conversations |
| What should we improve? | Repeated failure patterns and their severity |
| Did the prompt change help? | Comparable examples before release and outcomes after |
| Is usage healthy? | Adoption alongside completion, abandonment, and feedback |
| What should engineering do next? | A specific example and acceptable behavior |
Make qualitative review part of the workflow
A graph can show that usage increased without explaining whether the assistant helped. Lunary documents chats, threads, user context, and feedback tracking so a review can begin with the experience itself. Sample conversations by task type and outcome, including successful ones. Look for repeated failure patterns: unclear intent, missing context, an incorrect answer, or a technically correct answer that did not help the customer proceed.
Give engineering an actionable example
A useful report includes the conversation or run, what the user was trying to do, what happened, and what the acceptable behavior would have been. Avoid prescribing a prompt fix before the cause is understood. The failure might be missing data or a tool problem rather than wording. Agree on a small set of review categories and keep a written explanation for unusual cases.
Participate in prompt review with clear ownership
Lunary documents prompt templates and a playground. Use the workflow to inspect candidate behavior with the engineer who owns the application. Bring real examples, review both improvements and regressions, and agree on who approves changes. Check the exact prompt version and configuration used in the test so feedback refers to the same candidate. Editing a draft and changing the production behavior should be deliberate, distinguishable actions.
Measure the customer task, not just the conversation count
Choose an outcome that represents value: a correctly completed action, a resolved question, or a useful answer confirmed by review. Instrument that outcome in your product and connect it to safe task or session identifiers where appropriate. Report helpfulness, completion, and escalation rates alongside volume. A decline in support escalations is ambiguous if users are simply abandoning the assistant, so read the metrics together.
Run a review that ends with decisions
Each week, review a representative set of conversations and the highest-impact recent failures. Agree on the next changes, the examples that will test them, and the metric that could show improvement. After release, compare similar task cohorts and inspect unexpected movement. Use experiments when practical; otherwise describe the evidence as an observed change rather than claiming that the prompt caused a revenue increase.
- Bring examples of success as well as failure.
- Keep severity, frequency, and affected task separate.
- Assign an owner and a validation case to each proposed change.
- Check customer outcomes after the release, not just the test score.
Common questions
Do product reviewers need to understand every trace?
No. They should be able to explain the customer experience and expected behavior. Engineering can use the related execution context to find the cause.
Is thumbs-up feedback enough to measure quality?
It is useful but incomplete. Feedback can be sparse or biased toward unusual experiences; combine it with task outcomes and representative review.
Does Lunary automatically prove revenue impact?
No. Revenue attribution requires appropriate product and billing instrumentation, a reliable identity link, and an analysis that accounts for alternative explanations.
Sources & methodology
Lunary publishes this guide. We compare documented workflows and explain where each approach fits; this is not an independent benchmark or a hands-on product rating. Features, limits, and commercial terms can change. Check the linked vendor documentation before deciding.