Artifact Benchmark

BrowseComp

Canonical ID https://id.searchplex.net/browsecomp/

A benchmark for evaluating browsing agents on their ability to locate hard-to-find information on the web and produce short, verifiable answers.

Scope

BrowseComp evaluates web-browsing agent behavior over live web search and browsing conditions. It is not a fixed IR test collection in the BEIR sense and should be distinguished from BrowseComp-Plus, which introduces a fixed curated corpus for more controlled evaluation.

Evaluates

Not equivalent to

  • BrowseComp is not BEIR. BrowseComp evaluates browsing agents finding hard-to-locate web information; BEIR is a fixed information-retrieval benchmark suite.
  • BrowseComp is not BrowseComp-Plus. BrowseComp evaluates browsing agents in live web conditions; BrowseComp-Plus derives from BrowseComp but uses a fixed curated corpus, supporting documents, and hard negatives for controlled evaluation.
  • BrowseComp is not Test Collection. BrowseComp is a browsing-agent benchmark; a test collection is a fixed IR evaluation resource consisting of corpus, topics or queries, and relevance judgments.
  • BrowseComp is not Web Browsing Agent. BrowseComp is a benchmark for evaluating browsing agents; a web browsing agent is the system being evaluated.

Learn more

ID