Checklist

Check search changes before they are rolled out wide.

A fixed evaluation structure prevents ranking, AI signals or merchandising from being judged on a few fine examples.

Query set

Does the test set include top queries, long-tail, SKUs, brands, typos and problem queries?

Test set template →

Baseline

Is it documented how the current search performs on the same queries and KPIs?

Analytics →

Judgments

Are relevant results recorded in advance or with a consistent assessment method?

Relevance judgments

Online KPIs

Are CTR, conversion and reformulations consistently defined?

KPI Framework →

Segments

Are device, market, store and query type viewed separately where necessary?

Segmentation →

Guardrails

What metrics or critical queries should not deteriorate?

Governance →

Rate search changes on the same yardstick.

This makes experiments, regressions and rollback decisions more explainable.

See the Academy lesson