Multilingual Embeddings: Data, Training, and Understanding

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

University of Waterloo

Abstract

Embedding models have been a central component of modern information access systems, including search engines, question answering systems, retrieval-augmented generation, and nowadays agentic search pipelines. By converting text-space search into vector-space search, embedding models provide semantic matching beyond exact lexical overlap. However, their progress has been uneven across languages. While English dense retrieval has benefited from large training collections, mature benchmarks, and well-studied training recipes, many other languages still lack reliable retrieval resources, practical modeling guidance, and a clear understanding of why multilingual transfer works. This thesis studies multilingual embedding models for retrieval from three connected perspectives: data, training, and understanding. First, it introduces two multilingual retrieval resources, Mr. TYDI and MIRACL. Mr. TYDI establishes the first large-scale mono-lingual retrieval benchmark over eleven typologically diverse languages, while MIRACL expands the setting to eighteen languages with ten times richer human relevance annotations. Together, these datasets provide both the supervision needed to train multilingual dense retrievers and the benchmarks needed to evaluate them reliably across diverse languages and scripts. Second, we investigate how to train multilingual dense retrievers under realistic resource conditions. Starting from the observation that plain multilingual DPR can perform only marginally better than BM25 in unsupervised settings, this part studies cases where target-language training data, target-language pretrained models, or both may be unavailable. The analysis compares pretrained backbones, translated and in-language data, multi-stage fine-tuning, cross-lingual transfer, knowledge distillation, and monolingual versus multilingual transformers. These experiments provide practical guidance for building effective multilingual retrieval models when resources differ across languages. Third, this thesis examines how multilingual language models may understand across languages. It analyzes the roles of shared tokens across languages and their impact at the embedding finetuning stage, and then examines how language models may understand token-level semantic concepts, revealing that multilingual understanding and cross-lingual transfer largely depend on token-level semantic structures within multilingual vocabularies and embedding spaces. Overall, this thesis contributes datasets, training strategies, and model analyses that move multilingual retrieval research from infrastructure to practice to interpretation, advancing the development of retrieval systems that can support information access more reliably across languages, scripts, and resource conditions.

Description

Keywords

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By