Towards Foundation Models for Text-Rich Multimodal Tabular Data
| dc.contributor.author | Loh, Wei Min | |
| dc.date.accessioned | 2026-08-19T16:01:24Z | |
| dc.date.issued | 2026-08-19 | |
| dc.date.submitted | 2026-08-14 | |
| dc.description.abstract | Tabular data has been a central data format in statistics for centuries. In the age of machine learning and artificial intelligence, our set of tools has increased in size yet, many areas related to tabular modalities and tasks remain underexplored. Input representation is an important consideration for improving the current state of tabular approaches. Many successful approaches in purely numeric tabular data do not place much emphasis on input representations, and that may be limiting them to the purely numeric regime. On the other hand, the topic of information transfer is also one that practitioners need to address in the real world, especially when facing limited training data and evolving attributes. Information transfer can be in the form of continuously updating a model from a stream of data, or in the form of leveraging common information from a set of related tasks. This dissertation investigates these two areas, aiming to advance the current state of practical tabular models. The issue of the continual adaptation of a tabular model is addressed by framing it as a multi-armed bandit. The clear advantage of this formulation is that the solution not only adapts to evolving needs at inference time but also has the ability to influence which data points to collect next. In this dissertation, we propose a framework based on Nadaraya-Watson kernel regression and Thompson sampling to continuously improve the existing model without updating model weight at inference time. This framework includes a provable guarantee on the upper bound on error using finite-sample analysis. Empirically, we demonstrated improvements on both the standard benchmark for multi-armed bandits and a newly proposed benchmark for realistic news recommendation tasks. On the question of input representations, we started with the investigation of the current state of tabular models and identified important desiderata when working with tabular data. To the best of our knowledge, there are very few approaches in the current literature that satisfy the desiderata. We propose a transformer-based architecture, called basis transformers from the ground up, and introduce specialized components suited to the idiosyncrasies of tabular data. This architecture was evaluated on multi-task and related task regression experiments, and demonstrated improvements over gradient boosted decision trees, finetuned large language models, and similar deep tabular models. Scaling in terms of the amount of data and the number of learnable parameters is the predominant technique for improving learned representations and performance in contemporary machine learning. In the field of tabular models, scaling is particularly challenging due to the heterogeneous formats and data types, and relatively few publicly available sources of tabular datasets because many tables contain proprietary and sensitive information. With extensive engineering and data processing efforts, we developed a multimodal tabular foundation model, capable of extracting holistic representations and taking column names into account. The experiments demonstrated two main outcomes: the ability to perform zero-shot inference, which remains limited in tabular domains, and improvements over foundation model baselines on text-rich tabular datasets. | |
| dc.identifier.uri | https://hdl.handle.net/10012/23988 | |
| dc.language.iso | en | |
| dc.pending | false | |
| dc.publisher | University of Waterloo | en |
| dc.subject | machine learning | |
| dc.subject | tabular data | |
| dc.subject | multi-armed bandit | |
| dc.subject | foundation models | |
| dc.title | Towards Foundation Models for Text-Rich Multimodal Tabular Data | |
| dc.type | Doctoral Thesis | |
| uws-etd.degree | Doctor of Philosophy | |
| uws-etd.degree.department | David R. Cheriton School of Computer Science | |
| uws-etd.degree.discipline | Computer Science | |
| uws-etd.degree.grantor | University of Waterloo | en |
| uws-etd.embargo.terms | 0 | |
| uws.contributor.advisor | Poupart, Pascal | |
| uws.contributor.affiliation1 | Faculty of Mathematics | |
| uws.peerReviewStatus | Unreviewed | en |
| uws.published.city | Waterloo | en |
| uws.published.country | Canada | en |
| uws.published.province | Ontario | en |
| uws.scholarLevel | Graduate | en |
| uws.typeOfResource | Text | en |