guide · Evidence checked 2026-09-02

Freshness and coverage fail in different ways

A practical, evidence-led decision guide. Product capabilities and limits are separated from anything that would require hands-on testing.

A search system can retrieve yesterday's news yet miss the authoritative manual that has existed for years. Freshness and coverage therefore need separate fixtures.

Build four query groups: time-sensitive events with known publication times; stable facts with primary sources; obscure pages known to exist; and contested topics requiring multiple viewpoints. Freeze the query wording and gold URLs before testing.

Record recall@10, time from publication to first retrieval, primary-source rate, domain diversity, duplicate results and unsupported answer claims. Repeat the fresh set at fixed intervals and the stable set weekly. Save raw results because a later rerun cannot explain what the system returned today.

Compare Brave Search, Kagi, You.com, ChatGPT Search or any API with the same locale, date and source rules. For answer engines, split prose into claims and test each citation. “Has sources” is not a metric.

Set two gates: maximum acceptable discovery delay for fresh items and minimum recall for known authoritative pages. A product must pass both for the intended job.

Primary sources