1ViHallu
Visual variations and visual instruction construction improve visual-semantic alignment for LVLM hallucination mitigation.
visual variationsinstruction tuningalignment
Fine-tuning
vision-language
ACM MM 2025
First posted Jul 29, 2025
visual alignment
Authors: Ziyun Dai, Xiaoqiang Li, Shaohua Zhang·Corresponding: not specified·Affiliation: Shanghai University; Shanghai Business School
2ToR
Token reweighting jointly optimizes perception and reasoning tokens for multimodal RLVR.
RLVRtoken reweightinggrounded reasoning
Perception-grounded RL
vision-language
arXiv 2026
First posted Mar 26, 2026
RLVR / grounding
Authors: Jinda Lu, Junkang Wu, Jinghan Li·Corresponding: Jinda Lu·Affiliation: University of Science and Technology of China
3PGPO
Policy optimization with token-level visual dependency for grounded multimodal reasoning.
RLVRpolicy optimizationvisual dependency
Perception-grounded RL
vision-language
arXiv 2026
First posted Apr 2, 2026
RLVR / token credit
Authors: Zekai Ye, Qiming Li, Xiaocheng Feng·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory
4VEPO
Vision-anchored token selection combines visual sensitivity with entropy during RL.
RLVRtoken selectionvisual reasoning
Perception-grounded RL
vision-language
arXiv 2026
First posted Jun 2, 2026
RLVR / visual reasoning
Authors: Senjie Jin, Peixin Wang, Boyang Liu·Corresponding: Tao Gui·Affiliation: Fudan University
5AOD
Adversarial orthogonal disentanglement with training-free contrastive decoding for LVLM hallucination mitigation.
contrastive decodingdisentanglementtraining-free
Activation Editing
vision-language
arXiv 2026
First posted May 25, 2026
contrastive decoding
Authors: Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai·Corresponding: Xingjun Ma·Affiliation: Fudan University / Tencent; Nanjing University; Southeast University
6Search-G1
Grounded search agents via representation-based intrinsic rewards.
search agentsintrinsic rewardsgrounding
Tool-augmented
vision-language
awaiting release
arXiv pending
Author and affiliation metadata pending a public paper record.
7HIRE
Intermediate representation editing for hallucination mitigation.
representation editingactivation steeringLVLM
Activation Editing
vision-language
arXiv 2026
First posted Mar 31, 2026
representation editing
Authors: Wei Suo, Hanzu Zhang, Lijun Zhang·Corresponding: Peng Wang·Affiliation: Northwestern Polytechnical University
8NoLan
Suppress language priors dynamically during generation.
language priorsdynamic suppressionobject hallucination
Decoding-time
vision-language
arXiv 2026
First posted Feb 25, 2026
language-prior suppression
Authors: Lingfeng Ren, Weihao Yu, Runpeng Yu·Corresponding: Weihao Yu, Xinchao Wang·Affiliation: National University of Singapore; Peking University Shenzhen Graduate School
9MCoT
Mitigate reasoning-time hallucinations in multimodal CoT.
multimodal CoTreasoningmitigation
Verification-based
vision-language
CVPR 2026
First posted Mar 28, 2026
multimodal CoT
Authors: Ji Ma, Wei Suo, Peng Wang·Corresponding: Wei Suo·Affiliation: Northwestern Polytechnical University
10R-CoV
Region-aware chain-of-verification for LVLM hallucination.
region-awarechain-of-verificationobject hallucination
Verification-based
vision-language
arXiv 2026
First posted Apr 22, 2026
region verification
Authors: Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr·Corresponding: not specified·Affiliation: Max Planck Institute for Informatics / VIA Research Center; Google
11LEAD
Latent entropy-aware decoding for multimodal reasoning models.
latent entropydecodingmultimodal reasoning
Uncertainty-aware
vision-language
arXiv 2026
First posted Mar 9, 2026
entropy-aware decoding
Authors: Zhongxing Xu, Zhonghua Wang, Zhe Qian·Corresponding: not specified·Affiliation: Monash University
12Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Constructs a robust instruction tuning dataset with positive and negative samples to update the model.
instruction tuningdata curation
Fine-tuning
vision-language
ICLR 2024
First posted Jun 26, 2023
instruction tuning
Authors: Fuxiao Liu, Kevin Lin, Linjie Li·Corresponding: not specified·Affiliation: University of Maryland, College Park; Microsoft Corporation
13Aligning Large Multimodal Models with Factually Augmented RLHF
Introduces a factually augmented mechanism, using reinforcement learning from human feedback to align the model.
reinforcement learningrlhfmodel alignment
Perception-grounded RL
vision-language
ACL 2024
First posted Sep 25, 2023
rlhf
Authors: Zhiqing Sun, Sheng Shen, Shengcao Cao·Corresponding: not specified·Affiliation: UC Berkeley; CMU; UIUC; UW–Madison; UMass Amherst; Microsoft Research; MIT-IBM Watson AI Lab
14Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
Specifically trains the MLLM to generate self-feedback and revise its output based on it.
self-correctionfine-tuningmodel alignment
Fine-tuning
vision-language
NAACL 2024
First posted Nov 13, 2023
fine-tuning
Authors: Seongyun Lee, Sue Hyun Park, Yongrae Jo·Corresponding: not specified·Affiliation: Korea University; KAIST AI; LG AI Research
15HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Cleans "hallucinatory toxicity" in existing visual instruction datasets and generates counterfactual data for fine-tuning.
data curationmodel alignment
Fine-tuning
vision-language
CVPR 2024
First posted Nov 22, 2023
data curation
Authors: Qifan Yu, Juncheng Li, Longhui Wei·Corresponding: Longhui Wei·Affiliation: Zhejiang University; Huawei Cloud; Institute of Computing Technology, Chinese Academy of Sciences
16Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Constructs hallucination-aware data pairs and applies the DPO algorithm directly to update model weights.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
ICME 2024
First posted Nov 28, 2023
preference optimization
Authors: Zhiyuan Zhao, Bin Wang, Linke Ouyang·Corresponding: not specified·Affiliation: Shanghai AI Laboratory
17RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
Collects fine-grained, segment-level human correctional feedback for reward modeling and PPO optimization.
reinforcement learningreward modeling
Perception-grounded RL
vision-language
CVPR 2024
First posted Dec 1, 2023
rlhf
Authors: Tianyu Yu, Yuan Yao, Haoye Zhang·Corresponding: not specified·Affiliation: Tsinghua University; National University of Singapore
18ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
Trains a fine-grained visual reward model to guide the model toward better visual alignment.
visual groundingreward modeling
Perception-grounded RL
vision-language
CVPR 2024
First posted Feb 9, 2024
reward modeling
Authors: Siming Yan, Min Bai, Weifeng Chen·Corresponding: not specified·Affiliation: The University of Texas at Austin; AWS AI
19Visually Dehallucinative Instruction Generation: Know What You Don't Know
Teaches the model to actively output "I don't know" when evidence is insufficient via specialized instruction tuning.
instruction tuningmodel alignment
Fine-tuning
vision-language
ACL 2024
First posted Feb 15, 2024
instruction tuning
Authors: Sungguk Cha, Jusung Lee, Younghyun Lee·Corresponding: Sungguk Cha, Cheoljong Yang·Affiliation: Multimodal AI Lab., NC Research, NCSOFT Corporation
20Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Constructs a preference-based dataset to align the model with high-quality image descriptions via preference fine-tuning.
data curationpreference fine-tuningmodel alignment
Perception-grounded RL
vision-language
NeurIPS 2024
First posted Feb 18, 2024
preference fine-tuning
Authors: Yiyang Zhou, Chenhang Cui, Rafael Rafailov·Corresponding: Huaxiu Yao·Affiliation: UNC-Chapel Hill; Stanford University
21Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
Reconstructs the fine-tuning dataset to teach the model to output the End-Of-Sequence (EOS) token early to truncate generation.
instruction tuningdata curation
Fine-tuning
vision-language
ACL 2024
First posted Feb 22, 2024
instruction tuning
Authors: Zihao Yue, Liang Zhang, Qin Jin·Corresponding: not specified·Affiliation: Renmin University of China
22Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
Introduces adversarial instruction tuning to improve model robustness against misleading user queries.
instruction tuningmodel alignment
Fine-tuning
vision-language
ACL 2024
First posted Mar 15, 2024
instruction tuning
Authors: Dongmin Park, Zhaofang Qian, Guangxing Han·Corresponding: not specified·Affiliation: KRAFTON; UC San Diego; University of Central Florida
23Calibrated Self-Rewarding Vision Language Models
Proposes a self-calibrating reward mechanism enabling iterative self-generated feedback and preference alignment.
uncertaintyself-rewarding learningmodel alignment
Perception-grounded RL
vision-language
NeurIPS 2024
First posted May 23, 2024
self-rewarding learning
Authors: Yiyang Zhou, Zhiyuan Fan, Dongjie Cheng·Corresponding: not specified·Affiliation: UNC-Chapel Hill; University of Chicago; University of Maryland; Rutgers University; Independent Researcher
24Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
Constructs a reflective instruction dataset to teach the model to perform internal visual fact-checking before outputting answers.
instruction tuningdata curation
Fine-tuning
vision-language
ECCV 2024
First posted Jul 16, 2024
instruction tuning
Authors: Jinrui Zhang, Teng Wang, Haigang Zhang·Corresponding: Feng Zheng·Affiliation: Southern University of Science and Technology; The University of Hong Kong; Shenzhen Polytechnic University; The Cloud Computing and IT Institute of ZTE Corporation; Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory, Shenzhen, China
25Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
Cleans and rewrites captions in CLIP pre-training data to reduce object hallucinations at the vision-language source.
object hallucinationdata curation
Fine-tuning
vision-language
EMNLP 2024
First posted Oct 4, 2024
pre-training intervention
Authors: Yufang Liu, Tao Ji, Changzhi Sun·Corresponding: not specified·Affiliation: School of Computer Science and Technology, East China Normal University; School of Computer Science, Fudan University; Pazhou Laboratory, Huangpu
26V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
Uses a vision-guided mechanism to construct high-quality positive and negative preference pairs for DPO.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
EMNLP 2024
First posted Nov 5, 2024
preference optimization
Authors: Yuxi Xie, Guanzhen Li, Xiao Xu·Corresponding: not specified·Affiliation: National University of Singapore
27Hallucination-resistant multimodal content generation through knowledge graph-based reinforcement learning
Grounds multimodal generation in structured knowledge and optimizes graph-based rewards to suppress unsupported content.
reinforcement learningknowledge-graph reinforcement learningmodel alignment
Perception-grounded RL
vision-language
knowledge-graph reinforcement learning
Authors: Liang Zeng, Xinyi Lin, Shanping Yu·Corresponding: not specified·Affiliation:
28Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training
Selects hallucination-oriented samples, assigns severity-specific loss weights, and modulates localized visual attention during HD-DPO.
preference optimizationattention steering
Perception-grounded RL
vision-language
severity-guided dpo
Authors: Yuanyi Xu, Xiangru Zhu, Sihang Jiang·Corresponding: not specified·Affiliation: Fudan University; Renmin University of China; Alibaba Group
29EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation
Builds temporally grounded video preferences from visual and audio evidence and improves supervision with echo-layered keyframe sampling.
preference optimizationaudiovisual preference optimizationmodel alignment
Perception-grounded RL
vision-language
audiovisual preference optimization
Authors: Shuai Liu, Da Chen, Yiheng Pan·Corresponding: not specified·Affiliation: School of Software Engineering, Xi'an Jiaotong University; ByteDance; School of Cyber Science and Engineering, Xi'an Jiaotong University
30VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models
Reweights positive words by visual relevance and adapts the negative preference loss using lexical importance differences.
preference optimizationvision-guided lexical dpomodel alignment
Perception-grounded RL
vision-language
vision-guided lexical dpo
Authors: Siyuan Li, Feng Wang, Simeng Qin·Corresponding: not specified·Affiliation: School of Data Science and Intelligent Media, Communication University of China, Beijing, China; Tianjin University, Tianjin, China; Northeastern University, Shenyang, China; Alibaba Group, Hangzhou, China; Communication University of China, Beijing, China
31Multimodal Chain-of-Thought Reasoning in Language Models
Adopts a two-stage fine-tuning framework, training the model to generate a CoT reasoning process before the answer.
multimodal reasoningfine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2023
First posted Feb 2, 2023
fine-tuning
Authors: Zhuosheng Zhang, Aston Zhang, Mu Li·Corresponding: not specified·Affiliation: School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; GenAI, Meta; Amazon Web Services; Department of Computer Science and Engineering, Shanghai Jiao Tong University
32Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Introduces a hallucination-augmented contrastive learning loss function during the training phase.
contrastive learningmodel alignment
Fine-tuning
vision-language
arXiv 2023
First posted Dec 12, 2023
contrastive learning
Authors: Chaoya Jiang, Haiyang Xu, Mengfan Dong·Corresponding: not specified·Affiliation: National Engineering Research Center for Software Engineering, Peking University; Alibaba Group
33KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
Jointly trains a language model, a vision encoder, and a Graph Neural Network (GNN) to integrate knowledge graphs for reasoning.
external toolsmultimodal reasoning
Fine-tuning
vision-language
arXiv 2024
First posted Jan 23, 2024
joint training
Authors: Debjyoti Mondal, Suraj Modi, Subhadarshi Panda·Corresponding: not specified·Affiliation: Samsung R&D Institute India - Bangalore
34Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training
Proposes a novel vision-language pre-training (VLP) objective to suppress hallucinations.
pre-trainingmodel alignment
Fine-tuning
vision-language
arXiv 2023
First posted Oct 14, 2022
pre-training
Authors: Wenliang Dai, Zihan Liu, Ziwei Ji·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology; NVIDIA
35Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
Uses ChatGPT to rewrite image captions for data construction and performs two-stage fine-tuning.
fine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2023
First posted Dec 4, 2023
fine-tuning
Authors: Lei Wang, Jiabang He, Shenshen Li·Corresponding: not specified·Affiliation: Singapore Management University, Beijing Forestry University; University of Electronic Science and Technology of China
36Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
Proposes hallucination-aware DPO, utilizing fine-grained AI feedback to construct positive and negative pairs for weight updates.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Apr 22, 2024
preference optimization
Authors: Wenyi Xiao, Ziwei Huang, Leilei Gan·Corresponding: not specified·Affiliation: Zhejiang University; Alibaba Group; Fudan University
37RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
Utilizes fine-grained feedback from open-source AIs to build a preference dataset and updates weights using alignment algorithms.
reinforcement learningdata curation
Perception-grounded RL
vision-language
arXiv 2024
First posted May 27, 2024
rlaif
Authors: Tianyu Yu, Haoye Zhang, Qiming Li·Corresponding: not specified·Affiliation: Department of Computer Science and Technology, Tsinghua University; NExT++ Lab, School of Computing, National University of Singapore
38mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Proposes a conditional preference optimization algorithm to reduce hallucinations without degrading general capabilities.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Jun 17, 2024
preference optimization
Authors: Fei Wang, Wenxuan Zhou, James Y. Huang·Corresponding: not specified·Affiliation: University of Southern California; University of California, Davis; Microsoft Research
39CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
Uses a pre-trained CLIP model directly as an AI judge to generate preference pairs for DPO training.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Aug 19, 2024
preference optimization
Authors: Yassine Ouali, Adrian Bulat, Brais Martinez·Corresponding: not specified·Affiliation: Samsung AI Center Cambridge, UK; Technical University of Ia \textcommabelow si, Romania; Queen Mary University of London, UK
40EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
Proposes an efficient fine-grained unlearning framework to "forget" the tendency to hallucinate via reverse parameter updates.
unlearningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Feb 15, 2024
unlearning
Authors: Shangyu Xing, Fei Zhao, Zhen Wu·Corresponding: not specified·Affiliation: National Key Laboratory for Novel Software Technology, Nanjing University, China
41Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
Introduces a hallucination-inducing mechanism during training to construct sample pairs for parameter optimization.
contrastive tuningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted May 24, 2024
contrastive tuning
Authors: Xinyu Lyu, Beitao Chen, Lianli Gao·Corresponding: not specified·Affiliation: Center for Future Media, University of Electronic Science and Technology of China; Shenzhen Institute for Advanced Study, UESTC
42Mitigating Open-Vocabulary Caption Hallucinations
Trains a reward model based on Natural Language Inference (NLI) and uses RL to fine-tune the MLLM.
reinforcement learningreward modeling
Perception-grounded RL
vision-language
arXiv 2023
First posted Dec 6, 2023
rlhf
Authors: Assaf Ben-Kish, Moran Yanuka, Morris Alper·Corresponding: not specified·Affiliation: Tel-Aviv University
43Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
Constructs targeted repair instruction datasets for different hallucination types to perform targeted fine-tuning.
data curationfine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Apr 16, 2024
fine-tuning
Authors: Rui Hu, Yahan Tu, Shuyu Wei·Corresponding: not specified·Affiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
44See or Guess: Counterfactually Regularized Image Captioning
Introduces a counterfactual regularization term during training to force distinction between seen visuals and guessed priors.
regularization trainingmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Aug 29, 2024
regularization training
Authors: Qian Cao, Xu Chen, Ruihua Song·Corresponding: not specified·Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China; Tencent AI Lab
45Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
Generates targeted preference data specifically for hallucinations and applies DPO to penalize hallucination-prone features.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Nov 15, 2024
preference optimization
Authors: Yuhan Fu, Ruobing Xie, Xingwu Sun·Corresponding: not specified·Affiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
46Silkie: Preference Distillation for Large Visual Language Models
Uses AI to construct the VLFeedback dataset and distills preferences into the model via DPO.
preference optimizationdata curation
Perception-grounded RL
vision-language
arXiv 2023
First posted Dec 17, 2023
preference distillation
Authors: Lei Li, Zhihui Xie, Mukai Li·Corresponding: not specified·Affiliation: The University of Hong Kong; The Chinese University of Hong Kong (Shenzhen); Peking University
47Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
Constructs hard negative samples via data augmentation and introduces a contrastive loss to strengthen visual grounding.
visual groundingcontrastive tuningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted May 28, 2024
contrastive tuning
Authors: Pritam Sarkar, Sayna Ebrahimi, Ali Etemad·Corresponding: not specified·Affiliation: Queen's University and Vector Institute; Google DeepMind; Queen's University; Google Cloud AI Research
48Mitigating Hallucination in Visual Language Models with Visual Supervision
Constructs a fine-grained RAI-30k dataset and integrates the SAM model during instruction tuning.
instruction tuningdata curation
Fine-tuning
vision-language
arXiv 2023
First posted Nov 27, 2023
instruction tuning
Authors: Zhiyang Chen, Yousong Zhu, Yufei Zhan·Corresponding: not specified·Affiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Peng Cheng Laboratory; Wuhan AI Research
49Mitigating Multilingual Hallucination in Large Vision-Language Models
Builds a correctional dataset specifically for multilingual scenarios to mitigate cross-lingual hallucinations caused by language bias.
instruction tuningdata curation
Fine-tuning
vision-language
arXiv 2024
First posted Aug 1, 2024
instruction tuning
Authors: Xiaoye Qu, Mingyang Song, Wei Wei·Corresponding: Wei Wei·Affiliation: School of Computer Science & Technology, Huazhong University of Science and Technology; School of Computing Science and Technology, Fudan University; Wangxuan Institute of Computer Technology, Peking University; School of Computing Science and Technology, Zhejiang Gongshang University; Department of Computer Science and Engineering, The Chinese University of Hong Kong
50Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
Proposes a modality-fair preference optimization algorithm to balance penalties and prevent language prior dominance.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Oct 20, 2024
preference optimization
Authors: Songtao Jiang, Yan Zhang, Ruizhe Chen·Corresponding: not specified·Affiliation: Zhejiang University; National University of Singapore
51Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
Proposes a token-level preference optimization strategy based on self-calibrated visual-anchored rewards.
preference optimizationuncertainty
Perception-grounded RL
vision-language
arXiv 2024
First posted Dec 19, 2024
preference optimization
Authors: Jihao Gu, Yingyao Wang, Meng Cao·Corresponding: not specified·Affiliation: Alibaba Group; Mohamed bin Zayed University of Artificial Intelligence
52Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
Introduces Retrieval-Augmented Generation (RAG) into preference data construction to guide DPO training.
preference optimizationretrieval
Perception-grounded RL
vision-language
arXiv 2025
First posted Feb 18, 2025
preference optimization
Authors: Shuo Xing, Peiran Li, Yuping Wang·Corresponding: not specified·Affiliation: Texas A&M University; University of Michigan; UIUC; UNC Chapel Hill
53Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
Proposes symmetrical visual contrastive optimization using minimal contrastive images for alignment fine-tuning.
contrastive optimizationmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Feb 19, 2025
contrastive optimization
Authors: Shengguang Wu, Fan-Yun Sun, Kaiyue Wen·Corresponding: not specified·Affiliation: Stanford University
54FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
Uses fine-grained AI feedback to construct preference data and applies DPO for alignment training.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Apr 7, 2024
preference optimization
Authors: Liqiang Jing, Xinya Du·Corresponding: not specified·Affiliation: The University of Texas at Dallas
55TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Massively scales up text-centric visual instruction tuning data to specifically address text hallucinations in OCR tasks.
instruction tuningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Apr 19, 2024
instruction tuning
Authors: Jingqun Tang, Chunhui Lin, Zhen Zhao·Corresponding: not specified·Affiliation: ByteDance; East China Normal University; Huazhong University of Science and Technology
56Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Systematically builds datasets covering various alignment strategies to empirically verify their impact on hallucinations.
verificationdata curation
Fine-tuning
vision-language
arXiv 2024
First posted Jul 2, 2024
fine-tuning
Authors: Elmira Amirloo, Jean-Philippe Fauconnier, Christoph Roesmann·Corresponding: not specified·Affiliation: Apple Inc
57RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
Optimizes a visual reward model using auxiliary text-only preference data during RLHF to enhance hallucination discrimination.
reinforcement learningreward modeling
Perception-grounded RL
vision-language
arXiv 2024
First posted Aug 22, 2024
reward modeling
Authors: Chenglong Wang, Yang Gan, Yifu Huo·Corresponding: Chunliang Zhang·Affiliation: School of Computer Science and Engineering, Northeastern University, Shenyang, China; NiuTrans Research, Shenyang, China; CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China
58HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
Proposes a fine-tuning framework combining vision-enhanced penalty decoding with hierarchical feedback learning for behavior alignment.
feedback learningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Sep 30, 2024
feedback learning
Authors: Fan Yuan, Chi Qin, Xiaogang Xu·Corresponding: not specified·Affiliation: College of Artificial Intelligence; Nanjing University of Aeronautics and Astronautics, Nanjing, China; MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China; The Chinese University of Hong Kong, Hong Kong, China
59Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
Proposes entity-centric multimodal preference optimization, focusing penalty granularity precisely on specific visual entity tokens.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Jun 4, 2025
preference optimization
Authors: Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang·Corresponding: not specified·Affiliation: Soochow University, Suzhou, China; Baidu Inc., Beijing, China; Shandong University, Qingdao, China
60Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
Introduces faithful and concise reasoning rationales as supervision signals during training to enhance logical generation.
multimodal reasoningrationale trainingmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Apr 17, 2024
rationale training
Authors: Minghe Gao, Shuang Chen, Liang Pang·Corresponding: not specified·Affiliation: Zhejiang University; Chinese Academy of Sciences; National University of Singapore; Sun Yat-sen University
61Generating Faithful and Salient Text from Multimodal Data
Constructs a multimodal factual consistency dataset and fine-tunes the model to improve faithfulness in data-to-text generation.
data curationfine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Sep 6, 2024
fine-tuning
Authors: Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya·Corresponding: not specified·Affiliation: Department of Data Science & AI, Monash University, Australia; Department of Data Science, Monash University, Indonesia; Department of CSE, Bangladesh University of Engineering and Technology, Bangladesh
62Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
Formulates preference modeling as a next-token prediction task to update the model using fine-grained verifier data.
verificationpreference modelingmodel alignment
Perception-grounded RL
vision-language
arXiv 2024
First posted Oct 18, 2024
preference modeling
Authors: Chenhang Cui, An Zhang, Yiyang Zhou·Corresponding: An Zhang·Affiliation: National University of Singapore; UNC-Chapel Hill; Chicago University; Nanyang Technological University
63Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
Constructs a dataset targeting verb concept hallucinations and applies specific instruction tuning to correct action biases.
instruction tuningdata curation
Fine-tuning
vision-language
arXiv 2024
First posted Dec 6, 2024
instruction tuning
Authors: Zehao Wang, Xinpeng Liu, Yudonglin Zhang·Corresponding: not specified·Affiliation: Shanghai Jiao Tong University; ARC Lab, Tencent PCG
64Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution
Automatically generates and filters high-quality alignment data through an iterative self-evolution mechanism.
data curationself-evolution learningmodel alignment
Fine-tuning
vision-language
arXiv 2024
First posted Dec 20, 2024
self-evolution learning
Authors: Wentao Tan, Qiong Cao, Yibing Zhan·Corresponding: not specified·Affiliation: South China University of Technology; JD Explore Academy, Beijing; Pazhou Lab, Guangzhou
65CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
Proposes cross-modal hierarchical DPO to penalize global image-text matching and local object features at different levels.
preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Jan 28, 2025
preference optimization
Authors: Jinlan Fu, Shenzhen Huangfu, Hao Fei·Corresponding: not specified·Affiliation: National University of Singapore; Fudan University; Digital Twin Institute, Eastern Institute of Technology, Ningbo
66PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
Introduces subtle visual perturbations during training forward passes to force the learning of robust visual representations.
representation editingvisual perturbation trainingmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Mar 9, 2025
visual perturbation training
Authors: Cong Chen, Mingyu Liu, Chenchen Jing·Corresponding: not specified·Affiliation: Zhejiang University; WeChat Group; Zhejiang University of Technology
67Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy
Builds a low-level visual hallucination database and applies negative sampling to enhance awareness of low-level features.
negative sample fine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Mar 26, 2025
negative sample fine-tuning
Authors: Yinan Sun, Xiongkuo Min, Zicheng Zhang·Corresponding: Xiongkuo Min·Affiliation:
68Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
Adopts rationale-augmented instruction tuning to teach the model to generate critique logic before providing an answer.
instruction tuningmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted May 12, 2025
instruction tuning
Authors: Zexian Yang, Dian Li, Dayan Wu·Corresponding: Dayan Wu, Gang Liu·Affiliation: Institute of Information Engineering, Chinese Academy of Sciences; Foundation Technology Center, Tencent PCG
69OViP: Online Vision-Language Preference Learning for VLM Hallucination
Proposes an online preference learning framework to dynamically generate preference pairs during training.
preference learningmodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted May 21, 2025
preference learning
Authors: Shujun Liu, Siyuan Wang, Zejun Li·Corresponding: not specified·Affiliation: Fudan University; University of Southern California; ByteDance
70BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
Employs a bijective maximum likelihood learning approach to suppress hallucinations by optimizing joint probability distributions.
maximum likelihood learningmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted May 30, 2025
maximum likelihood learning
Authors: Huu-Thien Tran, Thanh-Dat Truong, Khoa Luu·Corresponding: not specified·Affiliation: CVIU Lab, University of Arkansas
71Stop learning it all to mitigate visual hallucination, Focus on the hallucination target
Stops learning redundant backgrounds and forces weight updates to focus exclusively on target areas causing hallucinations.
preference optimizationtarget-localized dpomodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Jun 13, 2025
target-localized dpo
Authors: Dokyoon Yoon, Youngsook Song, Woomyong Park·Corresponding: not specified·Affiliation: SIONIC AI
72Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
Identifies and purges hallucinated samples from fine-tuning data, using the refined high-quality data for knowledge distillation.
knowledge distillationmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Jul 7, 2025
knowledge distillation
Authors: Wenhao Li, Xiu Su, Jingyi Wu·Corresponding: not specified·Affiliation: University of Sydney; Central South University; Fudan University; Southeast University; HKUST; Sensetime Research
73ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
Proposes a closed-loop framework where the model generates predictions, performs backward verification, and uses errors as training signals.
verificationclosed-loop trainingmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Jul 7, 2025
closed-loop training
Authors: Jianjiang Yang, Yanshu li, Ziyan Huang·Corresponding: not specified·Affiliation: Department of Computer Science, University of Bristol; School of Future Technology, South China University of Technology; Brown University
74Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
Analyzes object biases introduced during pre-training and debiases the model using a reverse objective function during fine-tuning.
unlearningbias unlearningmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Aug 6, 2025
bias unlearning
Authors: Yifan Li, Kun Zhou, Wayne Xin Zhao·Corresponding: not specified·Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China; University of California, San Diego; DataCanvas Alaya NeW
75TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
Adopts an adaptive MinMax Token preference strategy to dynamically adjust reward allocation.
adaptive preference strategymodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Jul 29, 2025
adaptive preference strategy
Authors: Kejia Zhang, Keda Tao, Zhiming Luo·Corresponding: not specified·Affiliation: Xiamen University; Westlake University; DAMO Academy, Alibaba Group; AWS AI Lab, Amazon; Hupan Laboratory
76Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
Integrates the CHAIR metric, utilizing object-level matching scores as the preference reward for DPO.
preference optimizationobject-aware dpomodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Aug 27, 2025
object-aware dpo
Authors: Alberto Compagnoni, Davide Caffagni, Nicholas Moratelli·Corresponding: not specified·Affiliation: University of Modena and Reggio Emilia, Modena, Italy
77Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
Distinguishes "omission" and "fabrication" hallucination causes and constructs respective fine-tuning data for decoupled training.
decoupled fine-tuningmodel alignment
Fine-tuning
vision-language
arXiv 2025
First posted Aug 30, 2025
decoupled fine-tuning
Authors: Guangzong Si, Hao Yin, Xianfei Li·Corresponding: not specified·Affiliation: USTC; CowaRobot
78Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
Proposes semantic curriculum preference optimization, training the model in increasing difficulty to smoothly eliminate hallucinations.
preference optimizationcurriculum preference optimizationmodel alignment
Perception-grounded RL
vision-language
arXiv 2025
First posted Sep 29, 2025
curriculum preference optimization
Authors: Yuanshuai Li, Yuping Yan, Junfeng Tang·Corresponding: not specified·Affiliation: School of Engineering, Westlake University, Hangzhou, China; School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China
79Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
Proposes a universal training framework across multiple alignment formats (e.g., DPO, PPO) to eliminate vision-language biases.
preference optimizationreinforcement learning
Perception-grounded RL
vision-language
arXiv 2025
First posted Nov 21, 2025
unified alignment framework
Authors: Jiaye Qian, Ge Zheng, Yuchen Zhu·Corresponding: Sibei Yang·Affiliation: School of Computer Science and Engineering, Sun Yat-sen University; ShanghaiTech University
80OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Extracts self-attention matrices to find and penalize attention flows overly reliant on summary tokens.
attention steeringattention penaltyinternal intervention
Activation Editing
vision-language
CVPR 2024
First posted Nov 29, 2023
attention penalty
Authors: Qidong Huang, Xiaoyi Dong, Pan Zhang·Corresponding: not specified·Affiliation: University of Science and Technology of China; Shanghai AI Laboratory
81Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Fuses global and local attention feature matrices internally to enhance fine-grained perception.
attention steeringattention fusioninternal intervention
Activation Editing
vision-language
CVPR 2025
First posted Jun 18, 2024
attention fusion
Authors: Wenbin An, Feng Tian, Sicong Leng·Corresponding: not specified·Affiliation: Xi'an Jiaotong University; Nanyang Technological University; Lenovo Research; SGIT AI Lab; University of Massachusetts Boston
82Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
Adjusts cross-modal attention by forcefully amplifying the weights of visual tokens.
attention steeringattention amplificationinternal intervention
Activation Editing
vision-language
ECCV 2024
First posted Jul 31, 2024
attention amplification
Authors: Shi Liu, Kecheng Zheng, Wei Chen·Corresponding: not specified·Affiliation: State Key Lab of CAD&CG, Zhejiang University
83DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
Dives into the MLLM to mask specific attention connections overly dominated by text.
attention steeringattention maskinginternal intervention
Activation Editing
vision-language
EMNLP 2024
First posted Oct 6, 2024
attention masking
Authors: Xuan Gong, Tianshi Ming, Xinpeng Wang·Corresponding: not specified·Affiliation: Department of Computer Science and Technology, Tongji University
84Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
Uses causal analysis to block specific causal attention paths causing hallucinations during inference.
attention steeringcausal attention disconnectioninternal intervention
Activation Editing
vision-language
ICLR 2025
First posted Oct 7, 2024
causal attention disconnection
Authors: Guanyu Zhou, Yibo Yan, Xin Zou·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Tsinghua University
85Mitigating Object Hallucination via Concentric Causal Attention
Strengthens the causal association between local visual features and global semantics in attention matrices.
attention steeringconcentric causal attentioninternal intervention
Activation Editing
vision-language
NIPS 2024
First posted Oct 21, 2024
concentric causal attention
Authors: Yun Xing, Yiheng Li, Ivan Laptev·Corresponding: not specified·Affiliation: Nanyang Technological University; MBZUAI
86Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
Projects hidden states into a hallucination space and strips them via orthogonalization to clear hallucinations.
representation editinglatent space orthogonalizationinternal intervention
Activation Editing
vision-language
CVPR 2025
First posted Dec 18, 2024
latent space orthogonalization
Authors: Le Yang, Ziwei Zheng, Boxu Chen·Corresponding: not specified·Affiliation: Xi'an Jiaotong University, Xi'an 710049, China
87VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
Introduces visual-aware sparsification to directly drop irrelevant visual tokens prone to inducing hallucinations.
sparsification droppinginternal intervention
Activation Editing
vision-language
CVPR 2025
First posted Jan 11, 2025
sparsification dropping
Authors: Xianwei Zhuang, Zhihong Zhu, Yuxin Xie·Corresponding: not specified·Affiliation: SECE of Peking University
88Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
Adaptively constrains inter-layer information flow to block false causal paths causing object hallucinations.
object hallucinationinformation flow constraintinternal intervention
Activation Editing
vision-language
AAAI 2025
First posted Feb 28, 2025
information flow constraint
Authors: Jiaqi Bai, Hongcheng Guo, Zhongyuan Peng·Corresponding: Mohan Li·Affiliation: Cyberspace Institute of Advanced Technology, Guangzhou University, China; Huangpu Research School of Guangzhou University, China; CCSE, Beihang University, China; University of the Chinese Academy of Sciences, China
89ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
Amplifies low-level visual signals in multimodal layers to combat dilution by deep language features.
signal enhancementinternal intervention
Activation Editing
vision-language
CVPR 2025
First posted Mar 17, 2025
signal enhancement
Authors: Hao Yin, Guangzong Si, Zilei Wang·Corresponding: not specified·Affiliation: University of Science and Technology of China
90MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
Reconstructs multimodal attention using a causal matrix based on Manhattan distance.
attention steeringmanhattan causal replacementinternal intervention
Activation Editing
vision-language
ACMMM 2025
First posted Jul 12, 2025
manhattan causal replacement
Authors: Qiyan Zhao, Xiaofeng Zhang, Yiheng Li·Corresponding: not specified·Affiliation: FKLPRIU, Xiamen University of Technology, China; Shanghai Jiao Tong University, China; Nanyang Technological University, Singapore; Jilin University, China; Monash University, Australia; Zhejiang University, China; Huizhou University, China; Chinese Academy of Sciences, China
91Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
Intervenes from the visual processing architecture to fundamentally alter how the model acquires visual information.
pure vision reinforcementinternal intervention
Activation Editing
vision-language
EMNLP 2025
First posted Sep 17, 2025
pure vision reinforcement
Authors: Weihang Wang, Xinhao Li, Ziyue Wang·Corresponding: not specified·Affiliation: Bilibili; UESTC; University of Virginia
92Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination
Maps hallucination activations toward trusted attention regions with mixture-Gaussian bridges.
attention steeringrepresentation editing
Activation Editing
vision-language
WACV 2026
First posted Feb 10, 2026
attention manifold alignment
Authors: Ziqiang Shi, Rujie Liu, Shanshan Yu·Corresponding: not specified·Affiliation: Fujitsu Research & Development Center Co.,LTD., Beijing, China
93RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Predicts input-specific latent steering vectors with gradient correction in a single forward pass.
attention steeringrectified-flow interventioninternal intervention
Activation Editing
vision-language
rectified-flow intervention
Authors: Junyu Cheng, Zhibiao Liang, Yidong Chen·Corresponding: not specified·Affiliation: Department of Artificial Intelligence, School of Informatics, Xiamen University, China; School of Computer Science, South China Normal University, China; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China
94Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models
Filters intermediate attention maps to isolate dominant phantom text tokens and strengthen visual anchor tokens.
attention steeringdata curation
Activation Editing
vision-language
token-asymmetric filtering
Authors: Shuyi Ouyang, Hongyi Wang, Gongfan Fang·Corresponding: not specified·Affiliation: Zhejiang University; National University of Singapore
95Look Before You Speak: Visual Re-Focusing for Hallucination Mitigation in Large Vision-Language Models
Reorients the model toward image evidence before response generation to reduce visually unsupported claims.
visual re-focusinginternal intervention
Activation Editing
vision-language
visual re-focusing
Authors: Zheyuan Zhang, Yingmin Liu, Yang Shu·Corresponding: Yang Shu·Affiliation: Northwestern Polytechnical University, Xi'an, China; Zhejiang University, Hangzhou, China
96Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
Uses an attention lens to locate hallucinated features in middle layers and masks them for correction.
attention steeringmiddle layer maskinginternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Nov 23, 2024
middle layer masking
Authors: Zhangqi Jiang, Junkai Chen, Beier Zhu·Corresponding: not specified·Affiliation: National University of Defense Technology; Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and; Its Interdisciplinary Applications (Southeast University), Ministry of Education, China; Nanyang Technological University
97Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
Quantifies perception divergence and dynamically masks specific heads that lose visual alignment.
visual groundingattention steering
Activation Editing
vision-language
arXiv 2024
First posted Dec 18, 2024
attention head pruning
Authors: Jinghan He, Kuan Zhu, Haiyun Guo·Corresponding: not specified·Affiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; University of Science and Technology of China; Southeast University; National University of Singapore
98Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
Reallocates Attention Maps during inference to pull misplaced attention back to visual targets.
attention steeringattention reallocationinternal intervention
Activation Editing
vision-language
arXiv 2025
First posted Mar 11, 2025
attention reallocation
Authors: Chongjun Tu, Peng Ye, Dongzhan Zhou·Corresponding: Tao Chen·Affiliation: Fudan University; The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory; StepFun
99ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Proposes cross-level trusted intervention to block the drift of object representations toward false language priors.
representation editingcross-level blockinginternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Nov 22, 2024
cross-level blocking
Authors: Junzhe Chen, Tianshu Zhang, Shiyu Huang·Corresponding: not specified·Affiliation: Tsinghua University; The Hong Kong University of Science and Technology (Guangzhou); Zhipu AI; Chongqing University; Shanghai Jiao Tong University.
100Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
Implements causal intervention to cut off spurious feature maps hidden in attention matrices.
attention steeringcausal attention interventioninternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Dec 4, 2024
causal attention intervention
Authors: Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen·Corresponding: not specified·Affiliation: Inventec Corporation, No. 66, Hougang St., Shihlin Dist., Taipei City; University at Albany - SUNY, Albany, New York
101From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
Locates the evolution of object features and corrects visual tokens via early layer interventions.
early layer interventioninternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Oct 9, 2024
early layer intervention
Authors: Yuying Shang, Xinyi Zeng, Yutao Zhu·Corresponding: not specified·Affiliation: Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing, China; Gaoling School of Artificial Intelligence, Renmin University of China; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University, Beijing, China; Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University; Kuaishou Technology Inc., Beijing, China
102Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
Accurately locates key attention heads and specifically enhances or suppresses them during forward propagation.
attention steeringhead interventioninternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Nov 15, 2024
head intervention
Authors: Xiaofeng Zhang, Yihao Quan, Chaochen Gu·Corresponding: not specified·Affiliation: Shanghai Jiao Tong University; Alibaba Group; Beijing Jiaotong University
103Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
Calibrates the weight allocation of attention heads to weaken the implicit coverage of visual inputs by language priors.
uncertaintyattention steering
Activation Editing
vision-language
arXiv 2024
First posted May 28, 2024
attention calibration
Authors: Sangmin Woo, Donguk Kim, Jaehyuk Jang·Corresponding: not specified·Affiliation: KAIST
104Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Introduces dual-attention mechanisms to internally rebalance language priors and image feature flows.
attention steeringdual-attention balancinginternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Nov 21, 2024
dual-attention balancing
Authors: Haozhe Zhao, Shuzheng Si, Liang Chen·Corresponding: Baobao Chang·Affiliation: University of Illinois Urbana-Champaign; Peking University; Tsinghua University
105CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
Cuts off implicit semantic biases introduced during cross-lingual transitions inside the model.
attention steeringcross-lingual attention interventioninternal intervention
Activation Editing
vision-language
arXiv 2025
First posted Jun 3, 2025
cross-lingual attention intervention
Authors: Zekai Ye, Qiming Li, Xiaocheng Feng·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory; Central South University; Huawei Technologies Co., Ltd
106HalLoc: Token-level Localization of Hallucinations for Vision Language Models
Pinpoints tokens sourcing hallucinations and excises them during forward propagation.
token localization & pruninginternal intervention
Activation Editing
vision-language
arXiv 2025
First posted Jun 12, 2025
token localization & pruning
Authors: Eunkyu Park, Minyeong Kim, Gunhee Kim·Corresponding: not specified·Affiliation: Seoul National University
107Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Identifies and edits vision-language feature vectors representing hallucinations in intermediate layers.
representation editinginternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Oct 3, 2024
representation editing
Authors: Nick Jiang, Anish Kachinthaya, Suzie Petryk·Corresponding: not specified·Affiliation: University of California, Berkeley
108Reducing Hallucinations in Vision-Language Models via Latent Space Steering
Injects pre-trained steering vectors into the latent space to guide generation away from hallucinations.
attention steeringlatent space steeringinternal intervention
Activation Editing
vision-language
arXiv 2024
First posted Oct 21, 2024
latent space steering
Authors: Sheng Liu, Haotian Ye, Lei Xing·Corresponding: not specified·Affiliation: Stanford University
109Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
Penalizes decoding by contrasting the logit output distributions of original and distorted images during inference.
contrastive decodinggeneration control
Decoding-time
vision-language
CVPR 2024
First posted Nov 28, 2023
contrastive decoding
Authors: Sicong Leng, Hang Zhang, Guanzheng Chen·Corresponding: not specified·Affiliation: DAMO Academy, Alibaba Group; Nanyang Technological University; Hupan Lab, 310023, Hangzhou, China
110Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
Introduces Classifier-Free Guidance (CFG) during inference to amplify the influence of visual features in the generation process.
inference guidancegeneration control
Decoding-time
vision-language
CVPR 2024
First posted Feb 13, 2024
inference guidance
Authors: Linxi Zhao, Yihe Deng, Weitong Zhang·Corresponding: Quanquan Gu·Affiliation: Department of Computer Science, Cornell University, Ithaca, NY, USA; Department of Computer Science, University of California, Los Angeles, CA, USA; School of Data Science and Society, UNC, Chapel Hill, NC, USA
111IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
Penalizes pure language priors by contrasting the logits of inputs with and without images.
image-biased decodinggeneration control
Decoding-time
vision-language
ECCV 2024
First posted Feb 28, 2024
image-biased decoding
Authors: Lanyun Zhu, Deyi Ji, Tianrun Chen·Corresponding: not specified·Affiliation: Singapore University of Technology and Design; Alibaba Group; Zhejiang University
112HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Proposes adaptive focal-contrast decoding to fine-grainedly suppress hallucinations at the local token level.
contrastive decodingfocal contrastive decodinggeneration control
Decoding-time
vision-language
ICML 2024
First posted Mar 1, 2024
focal contrastive decoding
Authors: Zhaorun Chen, Zhuokai Zhao, Hongyin Luo·Corresponding: Zhaorun Chen, Zhuokai Zhao, Bo Li, Jiawei Zhou·Affiliation: University of Chicago, Chicago IL, USA; University of Illinois at Urbana-Champaign, Champaign IL, USA; Massachusetts Institute of Technology, Boston MA, USA; UNC-Chapel Hill, Chapel Hill NC, USA; Toyota Technological Institute at Chicago, Chicago IL, USA
113Multi-Modal Hallucination Control by Visual Information Grounding
Introduces a visual grounding mechanism during decoding to correct biases by contrasting different information streams.
visual groundingvisual information contrastgeneration control
Decoding-time
vision-language
CVPR 2024
First posted Mar 20, 2024
visual information contrast
Authors: Alessandro Favero, Luca Zancato, Matthew Trager·Corresponding: not specified·Affiliation: AWS AI Labs
114Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
Constructs positive and negative instructions, contrasting their logits during inference to guide decoding.
contrastive decodinginstruction contrastive decodinggeneration control
Decoding-time
vision-language
ACL 2024
First posted Mar 27, 2024
instruction contrastive decoding
Authors: Xintong Wang, Jingheng Pan, Liang Ding·Corresponding: not specified·Affiliation: Department of Informatics, Universität Hamburg; The University of Sydney
115Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
Introduces residual visual connections to add early visual features directly into subsequent decoding computations.
residual decodinggeneration control
Decoding-time
vision-language
ACL 2024
First posted Jun 30, 2024
residual decoding
Authors: Weihong Zhong, Xiaocheng Feng, Liang Zhao·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory
116ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
Uses attention visualization to locate hallucination regions and constructs interference images for contrastive decoding.
contrastive decodingattention steering
Decoding-time
vision-language
AAAI 2025
First posted Aug 25, 2024
visualized contrastive decoding
Authors: Yeji Park, Deokyeong Lee, Junsuk Choe·Corresponding: Junsuk Choe, Buru Chang·Affiliation: Sogang University
117Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Introduces a generative feedback loop to detect and correct the probability distribution of the current token in real time.
self-correctionself-correcting decodinggeneration control
Decoding-time
vision-language
ICLR 2025
First posted Feb 10, 2025
self-correcting decoding
Authors: Ce Zhang, Zifu Wan, Zhehan Kan·Corresponding: not specified·Affiliation: School of Computer Science, Carnegie Mellon University
118Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
Adaptively adjusts the contrastive penalty based on image complexity and local token features.
contrastive decodingdynamic contrastive decodinggeneration control
Decoding-time
vision-language
CVPR 2025
First posted Mar 1, 2025
dynamic contrastive decoding
Authors: Wei Suo, Lijun Zhang, Mengyang Sun·Corresponding: Yanning Zhang·Affiliation: School of Computer Science and Ningbo Institute, Northwestern Polytechnical University,China.; School of Cybersecurity, Northwestern Polytechnical University, China.; Computer Science, Swansea University.
119Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
Adaptively switches between different sampling algorithms based on attention strength.
attention steeringmixture of decodinggeneration control
Decoding-time
vision-language
ACL 2025
First posted May 17, 2025
mixture of decoding
Authors: Xinlong Chen, Yuanxing Zhang, Qiang Liu·Corresponding: not specified·Affiliation: New Laboratory of Pattern Recognition (NLPR); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences; Kuaishou Technology; Nanjing University
120Med-VCD: Mitigating hallucination for medical large vision language models through visual contrastive decoding
Selects visually informed tokens on the fly, removing redundant tokens while retaining critical medical image context.
contrastive decodingsparse visual contrastive decodinggeneration control
Decoding-time
vision-language
Computers in Biology and Medicine 2026
First posted Dec 1, 2025
sparse visual contrastive decoding
Authors: Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman Khaloo·Corresponding: not specified·Affiliation: Department of computer science, University of Central Florida, Orlando, USA; Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA; Department of electrical engineering, Columbia university, New York, NY, USA; Technical University of Applied Sciences Regensburg, Regensburg, Germany; Department of Surgery, University of Calgary, Calgary, Alberta, Canada; University College of Nabi Akram, Tabriz, Iran; School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
121Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
Constructs an image-free text input as a negative reference and subtracts language prior scores during logit calculation.
contrastive decodinglanguage contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Aug 6, 2024
language contrastive decoding
Authors: Avshalom Manevich, Reut Tsarfaty·Corresponding: not specified·Affiliation: Bar Ilan University
122CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
Contrasts logit distributions based on the original image versus self-generated descriptions to penalize text-reliant tokens.
contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Jun 4, 2024
contrastive decoding
Authors: Junho Kim, Hyunjun Kim, Yeonju Kim·Corresponding: Junho Kim·Affiliation: Integrated Vision and Language Lab, KAIST
123EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Suppresses event hallucinations by contrasting output probabilities between full and truncated videos.
contrastive decodingtemporal contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Sep 25, 2024
temporal contrastive decoding
Authors: Jiacheng Zhang, Yang Jiao, Shaoxiang Chen·Corresponding: not specified·Affiliation: Shanghai Key Lab of Intelligent Information Processing, School of CS, Fudan University; Shanghai Collaborative Innovation Center on Intelligent Visual Computing; Singapore University of Technology and Design; Shanghai Academy of Artificial Intelligence for Science; Meituan
124Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
Introduces an external CLIP model to score and guide candidate tokens during decoding to improve image-text consistency.
visual groundingexternal tools
Decoding-time
vision-language
arXiv 2024
First posted Feb 23, 2024
external guided decoding
Authors: Ailin Deng, Zhirui Chen, Bryan Hooi·Corresponding: not specified·Affiliation: School of Computing, National University of Singapore; Department of Industrial Systems Engineering and Management, National University of Singapore
125Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
Dynamically adjusts the logit weight distribution of text and image features during inference.
contrastive decodingre-balancing contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Sep 10, 2024
re-balancing contrastive decoding
Authors: Xiaoyu Liang, Jiayuan Yu, Lianrui Mu·Corresponding: Haoji Hu·Affiliation: College of Information Science and Electronic Engineering, Zhejiang University, China
126Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
Extracts internally generated facts as references for contrastive decoding to suppress over-extrapolation.
contrastive decodinginternal fact contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2025
First posted Feb 3, 2025
internal fact contrastive decoding
Authors: Chao Wang, Xuancheng Zhou, Weiwei Fu·Corresponding: Chao Wang, Yang Zhou·Affiliation: School of Future Technology, Shanghai University, Shanghai, 200444, China; Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China; School of Mechatronic Engineering and Automation, Shanghai, 200444, China
127HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Introduces a control parameter during decoding to force switching between pure contextual description and parameterized knowledge imagination.
inference interventiongeneration control
Decoding-time
vision-language
arXiv 2023
First posted Oct 3, 2023
inference intervention
Authors: Bohan Zhai, Shijia Yang, Chenfeng Xu·Corresponding: not specified·Affiliation: ByteDance Inc.; Stanford University; UC Berkeley; UIUC
128Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework
Introduces language contrastive decoding to weaken misleading priors from user prompts to combat sycophancy-induced hallucinations.
contrastive decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Aug 21, 2024
contrastive decoding
Authors: Yunpu Zhao, Rui Zhang, Junbin Xiao·Corresponding: Rui Zhang·Affiliation: School of Computer Science and Technology, University of Science and Technology of China; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences; Department of Computer Science, National University of Singapore; University of Illinois Urbana-Champaign; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences
129Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
Utilizes image summaries as guidance signals to intervene in output probabilities for contextual consistency.
summary-guided decodinggeneration control
Decoding-time
vision-language
arXiv 2024
First posted Oct 17, 2024
summary-guided decoding
Authors: Kyungmin Min, Minbeom Kim, Kang-il Lee·Corresponding: not specified·Affiliation: IPAI, Seoul National University; Dept. of ECE, Seoul National University
130Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
Contrasts logits from the original image and retrieved similar reference images to highlight core facts.
contrastive decodingretrieval
Decoding-time
vision-language
arXiv 2025
First posted May 26, 2025
retrieval visual contrastive decoding
Authors: Jihoon Lee, Min Song·Corresponding: not specified·Affiliation: Yonsei University; Onoma AI
131CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
Generates multi-granularity hierarchical feedback to dynamically correct outputs.
coarse-to-fine feedback decodinggeneration control
Decoding-time
vision-language
arXiv 2025
First posted Dec 29, 2025
coarse-to-fine feedback decoding
Authors: Zongsheng Cao, Yangfan He, Anran Liu·Corresponding: Anran Liu, Zepeng Wang·Affiliation: Researcher; UMN; PCIE
132Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Trains an independent Reviser model specifically to detect and reconstruct hallucinated entities in outputs.
self-correctionpost-hoc revisionexternal grounding
Verification-based
vision-language
ICLR 2023
First posted Oct 1, 2023
post-hoc revision
Authors: Yiyang Zhou, Chenhang Cui, Jaehong Yoon·Corresponding: not specified·Affiliation: UNC-Chapel Hill; Rutgers University; Columbia University; Stanford University
133Woodpecker: Hallucination Correction for Multimodal Large Language Models
Calls external visual foundation models to verify entity facts, finally using an LLM to rewrite and correct.
verificationexternal tools
Verification-based
vision-language
SCIS 2024
First posted Oct 24, 2023
agent correction
Authors: Shukang Yin, Chaoyou Fu, Sirui Zhao·Corresponding: not specified·Affiliation: School of Data Science, USTC & State Key Laboratory of Cognitive Intelligence; Tencent YouTu Lab
134Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
Constructs a pipeline system containing claim extraction, external detection verification, and corrected generation.
verificationexternal tools
Verification-based
vision-language
CVPR 2024
First posted Apr 30, 2024
pipeline agent
Authors: Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman·Corresponding: not specified·Affiliation: NVIDIA
135Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning
Modifies reference images to expose hallucinating agents and replaces consensus-only debate with evidence-based factual verification.
verificationcounterfactual multi-agent gamingexternal grounding
Verification-based
vision-language
counterfactual multi-agent gaming
Authors: Dayong Liang, Xiao-Yong Wei, Changmeng Zheng·Corresponding: not specified·Affiliation: South China University of Technology, Guangzhou, China; Peng Cheng Laboratory, Shenzhen, China; Sichuan University, Chengdu, China; The Hong Kong Polytechnic University, Hong Kong, China
136Mitigating Object and Relationship Hallucination in Large Vision Language Model with Multi-Agent Guidance
Coordinates specialized agents to identify and correct unsupported object and relationship claims.
verificationmulti-agent verificationexternal grounding
Verification-based
vision-language
multi-agent verification
Authors: Soohyun Kim, Gusang Lee, Kyuhong Shim·Corresponding: not specified·Affiliation: Seoul National University, Department of Electrical and Computer Engineering, Korea; Sungkyunkwan University, Department of Computer Science and Engineering, Korea
137CEBC: Conformal Evidence-Bounded Control for Low-Hallucination Vision-Language Generation
Calibrates an external detector on held-out data and minimally revises unsupported object mentions under explicit risk bounds.
self-correctionexternal toolsuncertainty
Verification-based
vision-language
conformal evidence-bounded editing
Authors: Ashish Mishra, Tarun Kumar, Arpit Shah·Corresponding: not specified·Affiliation: Hewlett Packard Labs, Bangalore; Hewlett Packard Labs, USA
138Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
Uses external language models to ask questions about generated claims for self-verification.
verificationexternal tools
Verification-based
vision-language
arXiv 2024
First posted Feb 18, 2024
logical loop verification
Authors: Junfei Wu, Qiang Liu, Ding Wang·Corresponding: not specified·Affiliation: Center for Research on Intelligent Perception and Computing; State Key Laboratory of Multimodal Artificial Intelligence Systems; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Nanjing University
139Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
Utilizes a bottom-up holistic reasoning framework combined with multi-perspective prompt verification.
verificationmultimodal reasoning
Verification-based
vision-language
arXiv 2024
First posted Dec 15, 2024
multi-perspective cross-checking
Authors: Shengqiong Wu, Hao Fei, Liangming Pan·Corresponding: Hao Fei·Affiliation: National University of Singapore, Singapore; University of Arizona, USA; University of California, Santa Barbara, USA; Skywork AI, Singapore; Nanyang Technological University, Singapore
140Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
Uses an external detector to locate relation errors, then applies template-based secondary correction.
external toolsuncertainty
Verification-based
vision-language
arXiv 2024
First posted Aug 18, 2024
external detection calibration
Authors: Kening Zheng, Junkai Chen, Yibo Yan·Corresponding: Xuming Hu·Affiliation: Hong Kong University of Science and Technology (Guangzhou); Guangxi Zhuang Autonomous Region Big Data Research Institute; Hong Kong University of Science and Technology
141Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
Adopts a "Retrospect-then-Compare" multi-step prompt to make the model review visual details for correction.
multi-step retrospective promptingexternal grounding
Verification-based
vision-language
arXiv 2024
First posted Mar 21, 2024
multi-step retrospective prompting
Authors: Dingchen Yang, Bowen Cao, Guang Chen·Corresponding: not specified·Affiliation: Tongji University; Peking University
142Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
Explicitly adds spatial relation logic rules into the prompt to correct spatial confusion.
spatial logic rule promptingexternal grounding
Verification-based
vision-language
arXiv 2025
First posted Feb 12, 2025
spatial logic rule prompting
Authors: Jiarui Wu, Zhuo Liu, Hangfeng He·Corresponding: not specified·Affiliation: University of Rochester
143RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
Applies random geometric transformations to construct multiple views for consistency verification and correction.
verificationmulti-view consistencyexternal grounding
Verification-based
vision-language
arXiv 2024
First posted May 28, 2024
multi-view consistency
Authors: Sangmin Woo, Jaehyuk Jang, Donguk Kim·Corresponding: not specified·Affiliation: KAIST
144Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
The model generates a draft, then independently reviews and self-scores sub-topics for correction.
topic-level self-scoringexternal grounding
Verification-based
vision-language
arXiv 2024
First posted Nov 26, 2024
topic-level self-scoring
Authors: Lehan He, Zeren Chen, Zhelun Shi·Corresponding: not specified·Affiliation: Shanghai AI Laboratory; School of Software, Beihang University; Shanghai Innovation Institute; Tsinghua University
145InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
Deploys a network of introspective and external verification agents for interactive debate.
verificationexternal tools
Verification-based
vision-language
arXiv 2025
First posted Dec 2, 2025
cross-modal introspective debate
Authors: Zhongyu Yang, Yingfang Yuan, Xuanming Jiang·Corresponding: Xuanming Jiang, Wei Pang·Affiliation: Xi'an Jiyun Technology Co., Ltd., Xi'an, China; BCML, Heriot-Watt University, Edinburgh, UK
146Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
Evaluates model uncertainty to actively trigger knowledge base retrieval for entity correction.
retrievalexternal toolsuncertainty
Tool-augmented
vision-language
TOMM 2025
First posted Aug 1, 2024
dynamic rag supplement
Authors: Xiaoye Qu, Qiyuan Chen, Wei Wei·Corresponding: not specified·Affiliation: Huazhong University of Science and Technology; Zhejiang University; Xiamen University; Zhejiang Gongshang University
147A Unified Hallucination Mitigation Framework for Large Vision-Language Models
Builds a unified cross-modal diagnosis pipeline utilizing external discriminators to intercept and modify outputs.
external toolsexternal discriminator systemexternal grounding
Tool-augmented
vision-language
TMLR 2024
First posted Sep 24, 2024
external discriminator system
Authors: Yue Chang, Liqiang Jing, Xiaopeng Zhang·Corresponding: not specified·Affiliation: The University of Texas at Dallas
148Mitigating Object Hallucinations via Sentence-Level Early Intervention
Implements factual truncation prompts in early sentence stages of generation to prevent error propagation.
early sentence truncationexternal grounding
Tool-augmented
vision-language
ICCV 2025
First posted Jul 16, 2025
early sentence truncation
Authors: Shangpin Peng, Senqiao Yang, Li Jiang·Corresponding: not specified·Affiliation: Harbin Institute of Technology, Shenzhen; The Chinese University of Hong Kong; The Chinese University of Hong Kong, Shenzhen
149Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
Overlays visual prompts (e.g., circles, bounding boxes) on the original image to force model attention.
attention steeringoriginal image overlayexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Apr 17, 2024
original image overlay
Authors: Yichi Zhang, Yinpeng Dong, Siyuan Zhang·Corresponding: not specified·Affiliation: Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center; THBI Lab, BNRist Center, Tsinghua University, Beijing 100084, China; RealAI; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong
150What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
Introduces "What if" prompting to compel the model to self-reflect on its initial judgments.
counterfactual promptingexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Mar 20, 2024
counterfactual prompting
Authors: Junho Kim, Yeon Ju Kim, Yong Man Ro·Corresponding: Yeon Ju Kim·Affiliation: Integrated Vision and Language Lab KAIST South Korea
151From Training-Free to Adaptive: Empirical Insights into MLLMs’ Understanding of Detection Information
Feeds information identified by external object detection models directly to the MLLM as context.
external toolsexternal visual guidanceexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Jan 31, 2024
external visual guidance
Authors: Qirui Jiao, Daoyuan Chen, Yilun Huang·Corresponding: not specified·Affiliation: Sun Yat-Sen University, Shenzhen, China; Alibaba Group, Hangzhou, China
152Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
Pre-supplements missing low-level visual perception details in the input prompt to bridge the gap.
perception detail supplementexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted May 24, 2024
perception detail supplement
Authors: Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar·Corresponding: not specified·Affiliation: University of Maryland, College Park, USA; Adobe, USA
153Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
Strictly controls the requested detail quantity via prompt to balance richness and hallucination rate.
prompt detail constraintexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Jun 18, 2024
prompt detail constraint
Authors: Mingqian Feng, Yunlong Tang, Zeliang Zhang·Corresponding: not specified·Affiliation: University of Rochester
154Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
Adopts multi-view prompts and multi-path reasoning, aggregating and comparing answers.
multimodal reasoningmulti-path votingexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Aug 30, 2024
multi-path voting
Authors: Xiaoye Qu, Jiashuo Sun, Wei Wei·Corresponding: not specified·Affiliation: Huazhong University of Science and Technology; Xiamen University; The Chinese University of Hong Kong
155Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
Uses prompts to guide the model to actively re-extract visual clues for confirmation during generation.
memory retracing confirmationexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Oct 4, 2024
memory retracing confirmation
Authors: Xin Zou, Yizhou Wang, Yibo Yan·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; China University of Geosciences; University of Technology Sydney
156Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Introduces an external Vision Value Model to assist tree search, guiding autoregressive generation.
external toolsexternal scoring searchexternal grounding
Tool-augmented
vision-language
arXiv 2024
First posted Dec 4, 2024
external scoring search
Authors: Xiyao Wang, Zhengyuan Yang, Linjie Li·Corresponding: not specified·Affiliation: University of Maryland, College Park; Microsoft
157Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Uses black-box optimization to automatically search for "visual prompt patches" that suppress hallucinations.
external toolsautomated patch searchexternal grounding
Tool-augmented
vision-language
arXiv 2025
First posted Apr 30, 2025
automated patch search
Authors: Sangmin Woo, Kang Zhou, Yun Zhou·Corresponding: Sangmin Woo, Haibo Ding·Affiliation: Amazon AWS AI; KAIST
158Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models
Uses uncertainty-aware contrastive calibration to suppress hallucinated actions in vision-language-action generation.
uncertaintyentropy-optimized decodinggeneration control
Uncertainty-aware
vision-language
entropy-optimized decoding
Authors: Ye Qiu, Zhaoxin Fan, Qingchen Yu·Corresponding: not specified·Affiliation:
159VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models
Uses a reinforcement-learned caption model as an auxiliary visual clue source and applies image-confidence constraints during generation.
uncertaintyvisual-clue-guided decodinggeneration control
Uncertainty-aware
vision-language
visual-clue-guided decoding
Authors: Guoqing Chen, Fu Zhang, Bingqian Liu·Corresponding: not specified·Affiliation: Northeastern University
160CTDD: Cumulative Trend Divergence Decoding for Mitigating Hallucination in Large Vision-Language Models
Tracks accumulated divergence in token-distribution trends and calibrates generation when visual and linguistic evidence separate.
uncertaintycumulative trend-divergence decodinggeneration control
Uncertainty-aware
vision-language
cumulative trend-divergence decoding
Authors: Jiani Hou, Zhixuan You, Siyu Luo·Corresponding: Bing Guo·Affiliation: College of Software Engineering, Sichuan University, Chengdu, China; International Economics and Trade, Central University of Finance and Economics, Beijing, China; College of Computer Science, Sichuan University, Chengdu, China
161Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
Filters hallucinated words by comparing the confidence of different generation paths via the model's own introspective mechanism.
uncertaintydata curation
Uncertainty-aware
vision-language
arXiv 2024
First posted Aug 4, 2024
self-introspective decoding
Authors: Fushuo Huo, Wenchao Xu, Zhong Zhang·Corresponding: Wenchao Xu, Peilin Zhao·Affiliation: Department of Computing, The Hong Kong Polytechnic University; Division of Integrative Systems and Design, Hong Kong University of Science and Technology; Tencent AI Lab; Huazhong University of Science and Technology; Tsinghua University
162ZINA: Multimodal Fine-grained Hallucination Detection and Editing
ZINA detects hallucinated spans, classifies six error types, and suggests grounded edits for MLLM outputs.
fine-grained spanserror taxonomyhallucination editing
Detection
vision-language
CVPR 2026
First posted Jun 16, 2025
span-level detection
Authors: Yuiga Wada, Kazuki Matsuda, Komei Sugiura·Corresponding: not specified·Affiliation: Keio AI Research Center; Keio University; Carnegie Mellon University
163Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
VBackChecker grounds response claims backward to image pixels with rich contextual evidence.
backward visual groundingrich-contextpixel-level grounding
Detection
vision-language
AAAI 2026
First posted Nov 15, 2025
pixel-level grounding
Authors: Pinxue Guo, Chongruo Wu, Xinyu Zhou·Corresponding: Wei Zhang, Wenqiang Zhang·Affiliation: Fudan University; Independent Researcher
164MHALO: Evaluating MLLMs as Fine-grained Hallucination Detectors
MHALO evaluates MLLMs on hallucination recognition and token-level localization with fine-grained metrics.
token-level detectionmeta-evaluation benchmarkMLLM as judge
Detection
vision-language
fine-grained F1/IoU
Authors: Yishuo Cai, Renjie Gu, Jiaxu Li·Corresponding: Xuancheng Huang·Affiliation: Central South University; Zhipu AI; Tsinghua University
165Detecting and Preventing Hallucinations in Large Vision Language Models
M-HalDetect supplies fine-grained multimodal annotations and reward-model signals for hallucination detection.
fine-grained feedbackreward modelingobject attribute relation
Detection
vision-language
AAAI 2024
First posted Aug 11, 2023
fine-grained detection
Authors: Anisha Gunjal, Jihan Yin, Erhan Bas·Corresponding: Erhan Bas·Affiliation: Scale AI
166Hal-Eval: a Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
A universal evaluator spanning object, attribute, relation, and event hallucinations with discriminative and generative modes.
event hallucinationfine-grained evaluationdiscriminative evaluation
Detection
vision-language
ACM MM 2024
First posted Feb 24, 2024
universal hallucination evaluation
Authors: Chaoya Jiang, Hongrui Jia, Mengfan Dong·Corresponding: Wei Ye·Affiliation: Peking University
167TLDR: Token-Level Detective Reward Model for Large Vision Language Models
TLDR assigns token-level rewards to expose hallucinated spans and support visual-language self-correction.
token-level rewardsself-correctionhallucination evaluation
Detection
vision-language
ICLR 2025
First posted Oct 7, 2024
token-level reward
Authors: Deqing Fu, Tong Xiao, Rui Wang·Corresponding: Deqing Fu, Lawrence Chen·Affiliation: Meta; University of Southern California
168Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
Med-HallMark and MediHallDetector provide hierarchical medical hallucination evaluation and fine-grained detection.
medical hallucinationhierarchical evaluationmultitask detection
Detection
vision-language
arXiv 2024
First posted Jun 14, 2024
medical hallucination detection
Authors: Jiawei Chen, Dingkang Yang, Tong Wu·Corresponding: Lihua Zhang·Affiliation: Fudan University; Tencent Youtu Lab; Cognition and Intelligent Technology Laboratory
169Unified Hallucination Detection for Multimodal Large Language Models
UNIHD combines auxiliary tools for claim-level hallucination detection and introduces the MHaluBench benchmark.
tool-augmented verificationclaim-level detectionmulti-category detection
Detection
vision-language
ACL 2024
First posted Feb 5, 2024
unified hallucination detection
Authors: Xiang Chen, Chenxi Wang, Yida Xue·Corresponding: Ningyu Zhang, Huajun Chen·Affiliation: Zhejiang University; Ant Group
170VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
VL-Uncertainty estimates response uncertainty under semantic-equivalent perturbations to detect hallucinations.
semantic-equivalent perturbationresponse entropyuncertainty estimation
Detection
vision-language
arXiv 2024
First posted Nov 18, 2024
uncertainty-based detection
Authors: Ruiyang Zhang, Hu Zhang, Zhedong Zheng·Corresponding: Zhedong Zheng·Affiliation: University of Macau; CSIRO Data61
171Structural Graph Probing of Vision–Language Models
Graph-based probes model neuron-correlation topology and expose structural signals associated with multimodal behavior and hallucination.
graph probingneuron topologyhallucination classification
Detection
vision-language
arXiv 2026
First posted Mar 28, 2026
structural hallucination probing
Authors: Haoyu He, Yue Zhuo, Yu Zheng·Corresponding: not specified·Affiliation: Northeastern University; Massachusetts Institute of Technology
172Lyapunov Probes for Hallucination Detection in Large Foundation Models
Stability-constrained perturbation probes distinguish factual representation regions from hallucination-prone boundaries.
stability theoryrepresentation probingperturbation analysis
Detection
vision-language
arXiv 2026
First posted Mar 6, 2026
stability-based detection
Authors: Bozhi Luan, Gen Li, Yalan Qin·Corresponding: Zhaoxin Fan·Affiliation: Beihang University; Shanghai University
173HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
HALP predicts hallucination risk before decoding by probing visual, vision-token, and query-token representations.
pre-generation detectioninternal representationslightweight probes
Detection
vision-language
EACL 2026
First posted Mar 5, 2026
pre-generation risk prediction
Authors: Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe Sun·Corresponding: not specified·Affiliation: Stony Brook University; Toyota Technological Institute at Chicago
174VADE: Visual Attention Guided Hallucination Detection and Elimination
VADE learns sequential patterns from raw visual attention maps for fine-grained hallucination detection and mitigation.
visual attentionsequence modelingfine-grained detection
Detection
vision-language
attention-map detection
Authors: Vishnu Prabhakaran, Purav Aggarwal, Vinay Kumar Verma·Corresponding: not specified·Affiliation: Amazon, India; Amazon, USA
175HALLUSHIFT++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
HALLUSHIFT++ extends internal distribution-shift analysis to hierarchical category, attribute, and relation hallucinations in MLLMs.
representation shiftshierarchical hallucinationsemantic chunking
Detection
vision-language
arXiv 2025
First posted Dec 8, 2025
hierarchical internal-state detection
Authors: Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta·Corresponding: Swagatam Das·Affiliation: Netaji Subhash Engineering College; TCG Crest; Indian Statistical Institute
176PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
PAS measures attention on preliminary output tokens as a training-free, reference-free object hallucination signal.
prelim-token attentiontraining-freeobject hallucination
Detection
vision-language
CVPR 2026
First posted Nov 14, 2025
attention-based object detection
Authors: Nhat Hoang-Xuan, Minh Vu, My T. Thai·Corresponding: not specified·Affiliation: Los Alamos National Laboratory; University of Florida
177MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
MTRE aggregates early multi-token logits with likelihood-ratio evidence to estimate VLM response reliability.
multi-token logitsreliability estimationlikelihood ratios
Detection
vision-language
arXiv 2025
First posted May 16, 2025
multi-token reliability detection
Authors: Geigh Zollicoffer, Minh Vu, Manish Bhattarai·Corresponding: not specified·Affiliation: Los Alamos National Laboratory