Advisor(s)

Chengcui Zhang

Committee Member(s)

Baocheng Geng
Brandon Oubre
Keren Li
Tianyang Wang

Document Type

Dissertation

Date of Award

6-18-2026

Degree Name

Doctor of Philosophy (PhD)

School

College of Arts and Sciences

Department

Computer and Information Sciences

Abstract

Over the past few years, multimodal foundation models have achieved remarkable progress in perception and understanding. However, two challenges limit their reliability: (1) dependence on offline training, which in most real-world settings requires large volumes of labeled data and, as a result, hinders the model’s ability to adapt to new data or domains; (2) weak cross-modal grounding, which often leads to hallucinated content generation, producing descriptions that are linguistically fluent but inconsistent with the input visual evidence. This dissertation frames hallucination mitigation as an outcome of transitioning from fixed learning (static, offline fine-tuning) to adaptive, feedback-driven lifelong learning. By incorporating test-time adaptation, it seeks to enable multimodal foundation models to evolve through interaction with real data. In this process, an evaluator assesses the overall quality of generated outputs, providing informative feedback that allows the model to adapt its internal representations and progressively reduce hallucinations during inference. This work begins with unimodal visual perception and a human-in-the-loop validation mechanism, where human evaluators assess and refine uncertain predictions to enhance overall reliability. However, this approach can only verify the generated content but is not capable of refining its model based on informative human feedback. This limitation motivates the final hallucination mitigation framework, where AI-based evaluation is transformed into a reward signal that actively updates the foundation model's parameters during inference, enabling continuous model improvement and reduced hallucination. Throughout the subsequent chapters, we progressively unfold the work, tracing its evolution from unimodal perception to multimodal understanding, and from static fine-tuning to feedback-driven, lifelong learning. The research initially explores contrastive learning to enhance intra-modal representations. Next, the study extends contrastive mechanisms to multimodality and hallucination detection. We subsequently evaluate parameter-efficient fine-tuning strategies in practical settings. Finally, the reinforcement learning-based test-time adaptation framework empowers the multimodal foundation model to refine its internal parameters and generated outputs during inference, achieving ongoing enhancement without offline retraining and effectively reducing hallucinations through iterative feedback-driven correction. Together, these studies establish a unified framework that transforms multimodal foundation models from passive generators into adaptive learners, leveraging lifelong learning to thereby achieve effective hallucination mitigation.

Keywords

Hallucination Mitigation;Image Captioning;Reinforcement Learning;Vision-Language Models

Share

COinS