Evaluating Machine Learning Models for Predicting PFAS Adsorption onto Activated Carbon

dc.contributor.authorBui, Gia Thinh
dc.date.accessioned2026-09-23T19:35:18Z
dc.date.issued2026-09-23
dc.date.submitted2026-08-19
dc.description.abstractPer- and polyfluoroalkyl substances (PFAS) are a diverse class of persistent contaminants of increasing regulatory concern. Activated carbon (AC) adsorption is among the most widely implemented technologies for removing PFAS from water, with performance governed by the combined effects of PFAS physicochemical properties, AC characteristics, and water chemistry. Although isotherm experiments are commonly used to characterize PFAS adsorption on AC, they can be time-consuming and resource-intensive. Models capable of predicting PFAS adsorption could reduce the need for isotherm testing by guiding experimental design or replacing selected experiments. This thesis evaluated the applicability of machine learning (ML) regression models for predicting PFAS adsorption onto AC in two types of systems: (1) synthetic solutions containing a single PFAS and (2) real groundwater solutions containing multiple PFAS. For the single-PFAS system, isotherm data from 21 published studies were consolidated into a curated dataset comprising 658 observations across 11 PFAS and 36 AC materials. To prevent data leakage and rigorously assess model performance, an isotherm-aware data splitting approach was implemented alongside three complementary validation regimes, including 5-fold cross-validation (CV), leave-one-PFAS-out (LOPO) CV, and leave-one-adsorbent-out (LOAO) CV. Four features were selected to predict the solid-liquid distribution coefficient (log Kd): PFAS molecular weight, AC specific surface area, the difference between AC pHpzc and solution pH, and aqueous-phase PFAS concentration. Among the models evaluated, multiple linear regression demonstrated consistent predictive performance across all three validation regimes (5-fold CV: R² = 0.77, LOPO CV: R² = 0.76, LOAO R² = 0.75), comparable to more complex ML models. For the compiled dataset and selected feature space, these results indicate that the linear model offers a robust and interpretable alternative to more complex ML algorithms for predicting PFAS adsorption in relatively simple systems. The modelling approach was subsequently extended to a more complex problem of predicting multi-PFAS adsorption on colloidal activated carbon (CAC) in real groundwater. The dataset comprised 192 observations for PFOS, PFOA, PFHxS, PFHxA, and 6:2 FTS obtained from isotherm experiments conducted with seven groundwater samples collected from PFAS-impacted sites. In each experiment, the five PFAS were added to the groundwater at equimolar initial concentrations. Because the dataset was relatively small, the allocation of groundwater samples between training and test datasets was systematically varied to evaluate the sensitivity of model performance to data partitioning and differences in groundwater chemistry. Across the data splitting scenarios, extreme gradient boosting and random forest provided the most accurate and consistent predictions. However, both tree-based models showed systematic overprediction or underprediction in some data-split cases. Reserving adsorption data collected at a single CAC concentration for calibration corrected this systematic bias and substantially improved predictive performance. This calibration strategy provides a practical workflow in which a single-point experiment can be used to verify or adjust model predictions under new water chemistry conditions, thereby reducing the need for complete isotherm experiments. Overall, this thesis demonstrates the potential of data-driven models to support the characterization of PFAS adsorption in both single-PFAS synthetic solutions and multi-PFAS real groundwater solutions. The models can inform the design of bench-scale isotherm studies, prioritize experimental testing, and reduce the experimental effort required to evaluate AC-based PFAS treatment. The contrasting model choices across the two studies further demonstrate that the predictive advantage of complex ML models over linear regression is situational rather than universal and should be evaluated rather than assumed. Beyond providing practical predictive tools, this work illustrates rigorous approaches to data splitting and model evaluation for relatively small experimental datasets, providing a framework for strengthening the reliability of data-driven modelling of contaminant adsorption.
dc.identifier.urihttps://hdl.handle.net/10012/24399
dc.language.isoen
dc.pendingfalse
dc.publisherUniversity of Waterlooen
dc.relation.urihttps://doi.org/10.17632/hr2xd2b7dn.1
dc.relation.urihttps://doi.org/10.17632/t8h2pcrdw7.1
dc.titleEvaluating Machine Learning Models for Predicting PFAS Adsorption onto Activated Carbon
dc.typeMaster Thesis
uws-etd.degreeMaster of Applied Science
uws-etd.degree.departmentCivil and Environmental Engineering
uws-etd.degree.disciplineCivil Engineering
uws-etd.degree.grantorUniversity of Waterlooen
uws-etd.embargo.terms4 months
uws.contributor.advisorPham, Anh
uws.contributor.advisorYeum, Chul Min
uws.contributor.affiliation1Faculty of Engineering
uws.peerReviewStatusUnrevieweden
uws.published.cityWaterlooen
uws.published.countryCanadaen
uws.published.provinceOntarioen
uws.scholarLevelGraduateen
uws.typeOfResourceTexten

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Bui_GiaThinh.pdf
Size:
3.26 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
6.4 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections