How Well Do Large Language Models Detect Bugs in Code Changes?

dc.contributor.authorYakubu, Ayinde
dc.date.accessioned2026-09-29T19:07:19Z
dc.date.issued2026-09-29
dc.date.submitted2026-09-25
dc.description.abstractThis thesis evaluates how well general-purpose open-weight large language models detect bugs in software code changes. The evaluation uses historical development data from the Apache Kafka project obtained through the ApacheJIT dataset. From approximately 12,000 commit records, the dataset was filtered to obtain 524 one-to-one bug-inducing commit (BIC) and bug-fixing commit (BFC) relationships and 530 non-bug-inducing commits. An automated framework was developed to retrieve commit patches, submit code changes for LLM-based review, and record predictions and review comments. Three open-weight LLMs—gpt-oss-120b, gemma-4-31B-it, and Qwen3.6-35B-A3B—were evaluated under a common zero-shot prompting strategy across three repeated experimental runs. Performance was measured using precision, accuracy, recall, F1-score, balanced accuracy, Matthews correlation coefficient, and processing coverage. In addition, an LLM-as-a-Judge procedure assessed whether generated defect reports were semantically consistent with evidence from corresponding bug-fixing commits and Apache Kafka JIRA issue records. The results show that the evaluated LLMs have limited reliability as autonomous defect detectors. Although the models identified subsets of historically labelled bug-inducing changes, substantial numbers of false positives and false negatives were observed. The first gpt-oss-120b run achieved the highest reported recall of 0.5163, while the highest individual-run accuracy was 0.4872. However, comparison with a trivial always-NOBUG baseline showed that model accuracies did not exceed the corresponding baseline accuracies on successfully processed records. Across the reported runs, balanced accuracy remained below 0.5 and Matthews correlation coefficient (MCC) remained negative, indicating weak overall discrimination between BIC and non-BIC benchmark examples. The results further demonstrate that conventional classification metrics alone provide an incomplete characterisation of LLM bug-detection reliability, because useful review requires semantic correctness, actionable explanations, and sufficient project context. Classification metrics alone do not establish explanation quality. The assigned judges rated 16–23% of first-run true-positive explanations as matching the historically documented defect. The findings suggest that assistant-style use is a more appropriate direction for further evaluation than autonomous defect detection. The thesis contributes a real-world evaluation framework, a comparative empirical assessment of three open-weight LLMs, and an evidence-based methodology for assessing generated explanations against historical defect evidence. The results also highlight repository context, semantic grounding, and hallucination reduction as important directions for improving future LLM-based bug detection systems.
dc.identifier.urihttps://hdl.handle.net/10012/24442
dc.language.isoen
dc.pendingfalse
dc.publisherUniversity of Waterlooen
dc.relation.urihttps://git.uwaterloo.ca/mmath-thesis-2025-ayinde/onboarding-chatbot-platform
dc.titleHow Well Do Large Language Models Detect Bugs in Code Changes?
dc.typeMaster Thesis
uws-etd.degreeMaster of Mathematics
uws-etd.degree.departmentDavid R. Cheriton School of Computer Science
uws-etd.degree.disciplineComputer Science
uws-etd.degree.grantorUniversity of Waterlooen
uws-etd.embargo.terms0
uws.contributor.advisorNagappan, Mei
uws.contributor.affiliation1Faculty of Mathematics
uws.peerReviewStatusUnrevieweden
uws.published.cityWaterlooen
uws.published.countryCanadaen
uws.published.provinceOntarioen
uws.scholarLevelGraduateen
uws.typeOfResourceTexten

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Yakubu_Ayinde.pdf
Size:
1.37 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
6.4 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections