Artifact Benchmark
BrowseComp
Canonical ID https://id.searchplex.net/browsecomp/
A benchmark for evaluating browsing agents on their ability to locate hard-to-find information on the web and produce short, verifiable answers.
Scope
BrowseComp evaluates web-browsing agent behavior over live web search and browsing conditions. It is not a fixed IR test collection in the BEIR sense and should be distinguished from BrowseComp-Plus, which introduces a fixed curated corpus for more controlled evaluation.
Evaluates
Not equivalent to
- BrowseComp is not BEIR. BrowseComp evaluates browsing agents finding hard-to-locate web information; BEIR is a fixed information-retrieval benchmark suite.
- BrowseComp is not BrowseComp-Plus. BrowseComp evaluates browsing agents in live web conditions; BrowseComp-Plus derives from BrowseComp but uses a fixed curated corpus, supporting documents, and hard negatives for controlled evaluation.
- BrowseComp is not Test Collection. BrowseComp is a browsing-agent benchmark; a test collection is a fixed IR evaluation resource consisting of corpus, topics or queries, and relevance judgments.
- BrowseComp is not Web Browsing Agent. BrowseComp is a benchmark for evaluating browsing agents; a web browsing agent is the system being evaluated.
Learn more
ID
- Canonical ID
https://id.searchplex.net/browsecomp/ - Scheme
Searchplex ID - Machine
JSONJSON-LD