Artifact Benchmark
MTEB
Canonical ID https://id.searchplex.net/mteb/
The Massive Text Embedding Benchmark, a benchmark suite for evaluating text embedding models across multiple tasks, datasets, and languages.
Also called
- Massive Text Embedding Benchmark
Scope
MTEB includes retrieval and reranking tasks, but it is broader than IR evaluation alone because it also covers tasks such as classification, clustering, semantic textual similarity, summarization, pair classification, and bitext mining. It should not be treated as the same kind of benchmark suite as BEIR, which is focused on information retrieval.
Observed terms
- MTEB Leaderboard
- Embedding Benchmark
Evaluates
Not equivalent to
- MTEB is not BEIR. MTEB evaluates text embeddings across many NLP and retrieval-adjacent tasks; BEIR is an information-retrieval benchmark suite.
- MTEB is not Embedding. MTEB is a benchmark suite for evaluating embedding models; an embedding is a vector representation.
- MTEB is not Test Collection. MTEB is a benchmark suite aggregating many tasks and datasets; a test collection is an evaluation resource pattern.
- MTEB is not BrowseComp-Plus. BrowseComp-Plus evaluates deep-research agents and retrievers; MTEB evaluates text embedding models across multiple tasks.
- MTEB is not Web Browsing Agent Evaluation. MTEB evaluates text embedding models across tasks; web browsing agent evaluation evaluates agent behavior in web information-seeking tasks.
Learn more
ID
- Canonical ID
https://id.searchplex.net/mteb/ - Scheme
Searchplex ID - Machine
JSONJSON-LD