Multilingual Embeddings: Data, Training, and Understanding

dc.contributor.authorZhang, Xinyu
dc.date.accessioned2026-08-12T15:47:32Z
dc.date.issued2026-08-12
dc.date.submitted2026-07-27
dc.description.abstractEmbedding models have been a central component of modern information access systems, including search engines, question answering systems, retrieval-augmented generation, and nowadays agentic search pipelines. By converting text-space search into vector-space search, embedding models provide semantic matching beyond exact lexical overlap. However, their progress has been uneven across languages. While English dense retrieval has benefited from large training collections, mature benchmarks, and well-studied training recipes, many other languages still lack reliable retrieval resources, practical modeling guidance, and a clear understanding of why multilingual transfer works. This thesis studies multilingual embedding models for retrieval from three connected perspectives: data, training, and understanding. First, it introduces two multilingual retrieval resources, Mr. TYDI and MIRACL. Mr. TYDI establishes the first large-scale mono-lingual retrieval benchmark over eleven typologically diverse languages, while MIRACL expands the setting to eighteen languages with ten times richer human relevance annotations. Together, these datasets provide both the supervision needed to train multilingual dense retrievers and the benchmarks needed to evaluate them reliably across diverse languages and scripts. Second, we investigate how to train multilingual dense retrievers under realistic resource conditions. Starting from the observation that plain multilingual DPR can perform only marginally better than BM25 in unsupervised settings, this part studies cases where target-language training data, target-language pretrained models, or both may be unavailable. The analysis compares pretrained backbones, translated and in-language data, multi-stage fine-tuning, cross-lingual transfer, knowledge distillation, and monolingual versus multilingual transformers. These experiments provide practical guidance for building effective multilingual retrieval models when resources differ across languages. Third, this thesis examines how multilingual language models may understand across languages. It analyzes the roles of shared tokens across languages and their impact at the embedding finetuning stage, and then examines how language models may understand token-level semantic concepts, revealing that multilingual understanding and cross-lingual transfer largely depend on token-level semantic structures within multilingual vocabularies and embedding spaces. Overall, this thesis contributes datasets, training strategies, and model analyses that move multilingual retrieval research from infrastructure to practice to interpretation, advancing the development of retrieval systems that can support information access more reliably across languages, scripts, and resource conditions.
dc.identifier.urihttps://hdl.handle.net/10012/23958
dc.language.isoen
dc.pendingfalse
dc.publisherUniversity of Waterlooen
dc.titleMultilingual Embeddings: Data, Training, and Understanding
dc.typeDoctoral Thesis
uws-etd.degreeDoctor of Philosophy
uws-etd.degree.departmentDavid R. Cheriton School of Computer Science
uws-etd.degree.disciplineComputer Science
uws-etd.degree.grantorUniversity of Waterlooen
uws-etd.embargo.terms0
uws.contributor.advisorLin, Jimmy
uws.contributor.affiliation1Faculty of Mathematics
uws.peerReviewStatusUnrevieweden
uws.published.cityWaterlooen
uws.published.countryCanadaen
uws.published.provinceOntarioen
uws.scholarLevelGraduateen
uws.typeOfResourceTexten

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhang_Xinyu.pdf
Size:
3.61 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
6.4 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections