General Concept
Web Browsing Agent Evaluation
Canonical ID https://id.searchplex.net/web-browsing-agent-evaluation/
Evaluation of agents that search, browse, inspect web sources, and synthesize answers or results from information found on the web.
Also called
- Browsing Agent Evaluation
- Web Agent Evaluation
Scope
Web browsing agent evaluation may measure source finding, persistence, answer accuracy, evidence use, citation fidelity, and multi-step browsing behavior. It is distinct from ordinary web search evaluation and from retrieval-only benchmark suites.
Evaluated by
Not equivalent to
- Web Browsing Agent Evaluation is not BEIR. BEIR evaluates retrieval systems over fixed IR datasets; web browsing agent evaluation measures agents that search and browse the web or web-like corpora.
- Web Browsing Agent Evaluation is not MTEB. MTEB evaluates text embedding models across tasks; web browsing agent evaluation evaluates agent behavior in web information-seeking tasks.
- Web Browsing Agent Evaluation is not Web Search. Web browsing agent evaluation evaluates agent behavior across search and browsing actions; web search is a search task or system setting.
Learn more
ID
- Canonical ID
https://id.searchplex.net/web-browsing-agent-evaluation/ - Scheme
Searchplex ID - Machine
JSONJSON-LD