Breaking Information Silos: Advancing Search System for Unified Information Seeking

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

University of Waterloo

Abstract

Information seeking is central to how people learn, make decisions, and advance knowledge. At the center of any search system is the retriever, the component that provides access to world knowledge: a system can surface only what its retriever can find. Retrievers, however, have traditionally been specialized, each trained for a particular task, fitted to a single language, restricted to one modality, and limited to a single retrieval pass. Each such specialization makes the retriever effective for what it was built for but walls off everything else, fragmenting the world's information into information silos: bodies of relevant information that no single search system can cohesively reach. This thesis breaks these silos by replacing traditional retrieval's specialized components, one for each task, language, and modality, with general-purpose models that learn relevance, perceive documents, and reason over information needs, lifting each underlying constraint rather than building another pipeline. The thesis then studies how to make this unified access general, efficient, verifiable, and reproducible, across three dimensions: domain and language, modality, and representation space. For domain and language silos, I show that fine-tuning a large language model (LLM) as a dense retriever yields strong generalization across retrieval tasks, even without the task-specific pretraining or data augmentation that BERT-era retrievers relied upon. I then show that this relevance-modeling capability need not remain locked inside a large model: it can be distilled into small, efficient retrievers that generalize across both tasks and languages, and it can be exploited with no relevance supervision at all by letting an LLM generate a hypothetical document that an off-the-shelf encoder grounds to the corpus. A recurring consequence is that such retrievers improve in step with the underlying language models. For modality silos, I introduce a paradigm shift from extracting text out of documents to encoding documents as they appear. Document Screenshot Embedding encodes a document's screenshot directly with a vision-language model, preserving text, images, tables, and layout in a single dense representation and outperforming parsing- and OCR-based pipelines on mixed-modality retrieval. I further argue that unifying access is only half of the goal, since a unified system must also be trustworthy: VISA extends the visual paradigm to the output side, attributing each generated answer to the specific region of the source document that supports it, and yielding a fully visual pipeline from retrieval to verifiable generation. For representation-space silos, I address the limitation that a single query embedding reaches only a local neighborhood, while the evidence for a complex, reasoning-intensive need is dispersed across the space. Crossing this silo requires an agent that can reason and search iteratively. I therefore cultivate general reasoning by scaling reinforcement learning with verifiable rewards beyond mathematics and code, bring that reasoning into the retrieval pipeline through a reranker trained to reason before ranking, and construct a benchmark over a fixed, curated corpus that, for the first time, evaluates deep-search agents fairly by disentangling the contribution of the retriever from that of the reasoning agent. This evaluation shows that, even with strong agents, retrieval quality remains pivotal. Underpinning these contributions, I develop and maintain open-source toolkits that make the unified paradigm efficient and reproducible: one for reproducible first-stage retrieval and evaluation, and one for training neural retrievers at scale and across languages and modalities. Together, these works move document retrieval from a collection of specialized pipelines toward unified, efficient, verifiable, and agentic access to world knowledge. I conclude by outlining how representations, search agents, and interaction must evolve to make such access seamless across both digital and physical worlds.

Description

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By