Measure AI search by query type and protect what already works.
A demo with some great examples says little about daily search quality. Evaluation uses fixed query sets, human relevance assessments, and actual search behavior to determine where a change helps, where it harms, and whether the search experience remains stable.
Why one average score is not enough
A change can improve natural language while deteriorating SKU search. A general average can hide that damage. Evaluate by query type, category, language, device and possibly customer segment.
Key segments include exact identifiers, brands, categories, attributes, problem-focused questions, long-tail queries, and search queries with troubles. Each segment has its own success criteria and risks.
Build a representative query set
Use real anonymized search queries from the online store, supplemented with business-critical cases. Include both popular queries and rare difficult questions. Zero-result searches and often reformulated queries are extra valuable because existing problems are visible there.
Freeze a core set for regression tests and periodically refresh a dynamic set. This way you protect well-known quality without only optimizing for historical questions.
Capture relevance judgments
For each query, it is assessed which products are highly relevant, usable, marginal or unsuitable. Technical products include factual compatibility in that assessment. With broad inspiration questions, multiple product groups may be relevant.
Allow doubts to be assessed by people with assortment knowledge. Document why a result is relevant; this will prevent assessments from taking on a different meaning tacitly on a next round.
Offline metrics
- Precision in the first positions: How many early results are really useful?
- Recall: are important suitable products found at all?
- Ranking quality: are the best candidates higher than weaker alternatives?
- Zero-result searches: does the set remain empty while appropriate assortment exists?
- Regressions: which previously good queries have deteriorated?
No metric tells the whole story. Combine them with segment analysis and concrete error examples.
Online behavior as additional evidence
Click ratio, product detail visit, filter usage, reformulations, add to cart and conversion show what customers do. They must be interpreted with caution: a high click ratio can also arise because the customer has to open several wrong results.
Therefore, check out the entire search funnel and compare similar periods or controlled experiments. Pay attention to seasonal effects, campaigns and assortment changes.
Golden queries and release protection
Golden queries are business-critical questions that must always remain correct, such as well-known SKUs, top brands or commonly used categories. Before each release, it is automatically checked whether the expected products and positions are retained.
Add new regressions to this set. This way, every error found becomes a permanent test instead of a one-time manual correction.
From wrong to action
Classify causes: query understanding, missing entity, bad product data, wrong constraint, retrieval, ranking or merchandising. A low click ratio requires a different solution when the right products are missing than if they are only too low.
Findoviq links evaluation to search analytics and configuration management. This allows a team to adjust, retest and only release changes that demonstrably help without damaging critical search routes.
Discuss your search questions
Do you want to know how this approach fits your assortment, product data and customer behavior? Together, we look at which query types have priority and where exact, semantic and business signals need to complement each other.