1See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
ViHallu uses visual variations and visual instruction construction for visual-semantic alignment.
visual variationsinstruction tuningalignment
Mitigation
vision-language
ACM MM 2025
First posted Jul 29, 2025
visual alignment
Authors: Ziyun Dai, Xiaoqiang Li, Shaohua Zhang·Corresponding: not specified·Affiliation: Shanghai University; Shanghai Business School
2Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
ToR reweights coupled perception and reasoning tokens during multimodal RLVR.
RLVRtoken reweightinggrounded reasoning
Mitigation
vision-language
arXiv 2026
First posted Mar 26, 2026
RLVR / grounding
Authors: Jinda Lu, Junkang Wu, Jinghan Li·Corresponding: Jinda Lu·Affiliation: University of Science and Technology of China
3Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
PGPO amplifies learning signals for visually dependent tokens.
RLVRpolicy optimizationvisual dependency
Mitigation
vision-language
arXiv 2026
First posted Apr 2, 2026
RLVR / token credit
Authors: Zekai Ye, Qiming Li, Xiaocheng Feng·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory
4Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection
VEPO combines visual sensitivity and token entropy for visual reasoning RL.
RLVRtoken selectionvisual reasoning
Mitigation
vision-language
arXiv 2026
First posted Jun 2, 2026
RLVR / visual reasoning
Authors: Senjie Jin, Peixin Wang, Boyang Liu·Corresponding: Tao Gui·Affiliation: Fudan University
5Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
AOD disentangles hallucination directions for training-free contrastive decoding in LVLMs.
contrastive decodingdisentanglementtraining-free
Mitigation
vision-language
arXiv 2026
First posted May 25, 2026
LVLM / decoding
Authors: Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai·Corresponding: Xingjun Ma·Affiliation: Fudan University / Tencent; Nanjing University; Southeast University
6Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Grounded search agents; public arXiv record is pending.
search agentsintrinsic rewardsgrounding
Mitigation
vision-language
agents / grounding
arXiv pending
Author and affiliation metadata pending a public paper record.
7Hallucination-aware intermediate representation edit in large vision-language models
Intermediate representation editing for LVLM hallucination mitigation.
representation editingactivation steeringLVLM
Mitigation
vision-language
arXiv 2026
First posted Mar 31, 2026
VLM / editing
Authors: Wei Suo, Hanzu Zhang, Lijun Zhang·Corresponding: Peng Wang·Affiliation: Northwestern Polytechnical University
8NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
Dynamic suppression of language priors during decoding.
language priorsdynamic suppressionobject hallucination
Mitigation
vision-language
arXiv 2026
First posted Feb 25, 2026
decoding / object
Authors: Lingfeng Ren, Weihao Yu, Runpeng Yu·Corresponding: Weihao Yu, Xinchao Wang·Affiliation: National University of Singapore; Peking University Shenzhen Graduate School
9Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
Hallucination analysis and mitigation for multimodal CoT.
multimodal CoTreasoningmitigation
Mitigation
vision-language
CVPR 2026
First posted Mar 28, 2026
reasoning / CoT
Authors: Ji Ma, Wei Suo, Peng Wang·Corresponding: Wei Suo·Affiliation: Northwestern Polytechnical University
10R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
Region-aware verification loops for object hallucinations.
region-awarechain-of-verificationobject hallucination
Verification
vision-language
arXiv 2026
First posted Apr 22, 2026
verification / region
Authors: Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr·Corresponding: not specified·Affiliation: Max Planck Institute for Informatics / VIA Research Center; Google
11Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
Entropy-aware decoding for multimodal reasoning models.
latent entropydecodingmultimodal reasoning
Uncertainty
vision-language
arXiv 2026
First posted Mar 9, 2026
uncertainty / decoding
Authors: Zhongxing Xu, Zhonghua Wang, Zhe Qian·Corresponding: not specified·Affiliation: Monash University
12Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Constructs a robust instruction tuning dataset with positive and negative samples to update the model.
instruction tuningdata curation
Mitigation
vision-language
ICLR 2024
First posted Jun 26, 2023
instruction tuning
Authors: Fuxiao Liu, Kevin Lin, Linjie Li·Corresponding: not specified·Affiliation: University of Maryland, College Park; Microsoft Corporation
13Aligning Large Multimodal Models with Factually Augmented RLHF
Introduces a factually augmented mechanism, using reinforcement learning from human feedback to align the model.
reinforcement learningrlhfmodel alignment
Mitigation
vision-language
ACL 2024
First posted Sep 25, 2023
rlhf
Authors: Zhiqing Sun, Sheng Shen, Shengcao Cao·Corresponding: not specified·Affiliation: UC Berkeley; CMU; UIUC; UW–Madison; UMass Amherst; Microsoft Research; MIT-IBM Watson AI Lab
14Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
Specifically trains the MLLM to generate self-feedback and revise its output based on it.
self-correctionfine-tuningmodel alignment
Mitigation
vision-language
NAACL 2024
First posted Nov 13, 2023
fine-tuning
Authors: Seongyun Lee, Sue Hyun Park, Yongrae Jo·Corresponding: not specified·Affiliation: Korea University; KAIST AI; LG AI Research
15HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Cleans "hallucinatory toxicity" in existing visual instruction datasets and generates counterfactual data for fine-tuning.
data curationmodel alignment
Mitigation
vision-language
CVPR 2024
First posted Nov 22, 2023
data curation
Authors: Qifan Yu, Juncheng Li, Longhui Wei·Corresponding: Longhui Wei·Affiliation: Zhejiang University; Huawei Cloud; Institute of Computing Technology, Chinese Academy of Sciences
16Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Constructs hallucination-aware data pairs and applies the DPO algorithm directly to update model weights.
preference optimizationmodel alignment
Mitigation
vision-language
ICME 2024
First posted Nov 28, 2023
preference optimization
Authors: Zhiyuan Zhao, Bin Wang, Linke Ouyang·Corresponding: not specified·Affiliation: Shanghai AI Laboratory
17RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
Collects fine-grained, segment-level human correctional feedback for reward modeling and PPO optimization.
reinforcement learningreward modeling
Mitigation
vision-language
CVPR 2024
First posted Dec 1, 2023
rlhf
Authors: Tianyu Yu, Yuan Yao, Haoye Zhang·Corresponding: not specified·Affiliation: Tsinghua University; National University of Singapore
18ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
Trains a fine-grained visual reward model to guide the model toward better visual alignment.
visual groundingreward modeling
Mitigation
vision-language
CVPR 2024
First posted Feb 9, 2024
reward modeling
Authors: Siming Yan, Min Bai, Weifeng Chen·Corresponding: not specified·Affiliation: The University of Texas at Austin; AWS AI
19Visually Dehallucinative Instruction Generation: Know What You Don't Know
Teaches the model to actively output "I don't know" when evidence is insufficient via specialized instruction tuning.
instruction tuningmodel alignment
Mitigation
vision-language
ACL 2024
First posted Feb 15, 2024
instruction tuning
Authors: Sungguk Cha, Jusung Lee, Younghyun Lee·Corresponding: Sungguk Cha, Cheoljong Yang·Affiliation: Multimodal AI Lab., NC Research, NCSOFT Corporation
20Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Constructs a preference-based dataset to align the model with high-quality image descriptions via preference fine-tuning.
data curationpreference fine-tuningmodel alignment
Mitigation
vision-language
NeurIPS 2024
First posted Feb 18, 2024
preference fine-tuning
Authors: Yiyang Zhou, Chenhang Cui, Rafael Rafailov·Corresponding: Huaxiu Yao·Affiliation: UNC-Chapel Hill; Stanford University
21Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
Reconstructs the fine-tuning dataset to teach the model to output the End-Of-Sequence (EOS) token early to truncate generation.
instruction tuningdata curation
Mitigation
vision-language
ACL 2024
First posted Feb 22, 2024
instruction tuning
Authors: Zihao Yue, Liang Zhang, Qin Jin·Corresponding: not specified·Affiliation: Renmin University of China
22Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
Introduces adversarial instruction tuning to improve model robustness against misleading user queries.
instruction tuningmodel alignment
Mitigation
vision-language
ACL 2024
First posted Mar 15, 2024
instruction tuning
Authors: Dongmin Park, Zhaofang Qian, Guangxing Han·Corresponding: not specified·Affiliation: KRAFTON; UC San Diego; University of Central Florida
23Calibrated Self-Rewarding Vision Language Models
Proposes a self-calibrating reward mechanism enabling iterative self-generated feedback and preference alignment.
uncertaintyself-rewarding learningmodel alignment
Mitigation
vision-language
NeurIPS 2024
First posted May 23, 2024
self-rewarding learning
Authors: Yiyang Zhou, Zhiyuan Fan, Dongjie Cheng·Corresponding: not specified·Affiliation: UNC-Chapel Hill; University of Chicago; University of Maryland; Rutgers University; Independent Researcher
24Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
Constructs a reflective instruction dataset to teach the model to perform internal visual fact-checking before outputting answers.
instruction tuningdata curation
Mitigation
vision-language
ECCV 2024
First posted Jul 16, 2024
instruction tuning
Authors: Jinrui Zhang, Teng Wang, Haigang Zhang·Corresponding: Feng Zheng·Affiliation: Southern University of Science and Technology; The University of Hong Kong; Shenzhen Polytechnic University; The Cloud Computing and IT Institute of ZTE Corporation; Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory, Shenzhen, China
25Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
Cleans and rewrites captions in CLIP pre-training data to reduce object hallucinations at the vision-language source.
object hallucinationdata curation
Mitigation
vision-language
EMNLP 2024
First posted Oct 4, 2024
pre-training intervention
Authors: Yufang Liu, Tao Ji, Changzhi Sun·Corresponding: not specified·Affiliation: School of Computer Science and Technology, East China Normal University; School of Computer Science, Fudan University; Pazhou Laboratory, Huangpu
26V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
Uses a vision-guided mechanism to construct high-quality positive and negative preference pairs for DPO.
preference optimizationmodel alignment
Mitigation
vision-language
EMNLP 2024
First posted Nov 5, 2024
preference optimization
Authors: Yuxi Xie, Guanzhen Li, Xiao Xu·Corresponding: not specified·Affiliation: National University of Singapore
27Hallucination-resistant multimodal content generation through knowledge graph-based reinforcement learning
Grounds multimodal generation in structured knowledge and optimizes graph-based rewards to suppress unsupported content.
reinforcement learningknowledge-graph reinforcement learningmodel alignment
Mitigation
vision-language
knowledge-graph reinforcement learning
Authors: Liang Zeng, Xinyi Lin, Shanping Yu·Corresponding: not specified·Affiliation:
28Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training
Selects hallucination-oriented samples, assigns severity-specific loss weights, and modulates localized visual attention during HD-DPO.
preference optimizationattention steering
Mitigation
vision-language
severity-guided dpo
Authors: Yuanyi Xu, Xiangru Zhu, Sihang Jiang·Corresponding: not specified·Affiliation: Fudan University; Renmin University of China; Alibaba Group
29EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation
Builds temporally grounded video preferences from visual and audio evidence and improves supervision with echo-layered keyframe sampling.
preference optimizationaudiovisual preference optimizationmodel alignment
Mitigation
vision-language
audiovisual preference optimization
Authors: Shuai Liu, Da Chen, Yiheng Pan·Corresponding: not specified·Affiliation: School of Software Engineering, Xi'an Jiaotong University; ByteDance; School of Cyber Science and Engineering, Xi'an Jiaotong University
30VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models
Reweights positive words by visual relevance and adapts the negative preference loss using lexical importance differences.
preference optimizationvision-guided lexical dpomodel alignment
Mitigation
vision-language
vision-guided lexical dpo
Authors: Siyuan Li, Feng Wang, Simeng Qin·Corresponding: not specified·Affiliation: School of Data Science and Intelligent Media, Communication University of China, Beijing, China; Tianjin University, Tianjin, China; Northeastern University, Shenyang, China; Alibaba Group, Hangzhou, China; Communication University of China, Beijing, China
31Multimodal Chain-of-Thought Reasoning in Language Models
Adopts a two-stage fine-tuning framework, training the model to generate a CoT reasoning process before the answer.
multimodal reasoningfine-tuningmodel alignment
Mitigation
vision-language
arXiv 2023
First posted Feb 2, 2023
fine-tuning
Authors: Zhuosheng Zhang, Aston Zhang, Mu Li·Corresponding: not specified·Affiliation: School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; GenAI, Meta; Amazon Web Services; Department of Computer Science and Engineering, Shanghai Jiao Tong University
32Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Introduces a hallucination-augmented contrastive learning loss function during the training phase.
contrastive learningmodel alignment
Mitigation
vision-language
arXiv 2023
First posted Dec 12, 2023
contrastive learning
Authors: Chaoya Jiang, Haiyang Xu, Mengfan Dong·Corresponding: not specified·Affiliation: National Engineering Research Center for Software Engineering, Peking University; Alibaba Group
33KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
Jointly trains a language model, a vision encoder, and a Graph Neural Network (GNN) to integrate knowledge graphs for reasoning.
external toolsmultimodal reasoning
Mitigation
vision-language
arXiv 2024
First posted Jan 23, 2024
joint training
Authors: Debjyoti Mondal, Suraj Modi, Subhadarshi Panda·Corresponding: not specified·Affiliation: Samsung R&D Institute India - Bangalore
34Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training
Proposes a novel vision-language pre-training (VLP) objective to suppress hallucinations.
pre-trainingmodel alignment
Mitigation
vision-language
arXiv 2023
First posted Oct 14, 2022
pre-training
Authors: Wenliang Dai, Zihan Liu, Ziwei Ji·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology; NVIDIA
35Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
Uses ChatGPT to rewrite image captions for data construction and performs two-stage fine-tuning.
fine-tuningmodel alignment
Mitigation
vision-language
arXiv 2023
First posted Dec 4, 2023
fine-tuning
Authors: Lei Wang, Jiabang He, Shenshen Li·Corresponding: not specified·Affiliation: Singapore Management University, Beijing Forestry University; University of Electronic Science and Technology of China
36Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
Proposes hallucination-aware DPO, utilizing fine-grained AI feedback to construct positive and negative pairs for weight updates.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Apr 22, 2024
preference optimization
Authors: Wenyi Xiao, Ziwei Huang, Leilei Gan·Corresponding: not specified·Affiliation: Zhejiang University; Alibaba Group; Fudan University
37RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
Utilizes fine-grained feedback from open-source AIs to build a preference dataset and updates weights using alignment algorithms.
reinforcement learningdata curation
Mitigation
vision-language
arXiv 2024
First posted May 27, 2024
rlaif
Authors: Tianyu Yu, Haoye Zhang, Qiming Li·Corresponding: not specified·Affiliation: Department of Computer Science and Technology, Tsinghua University; NExT++ Lab, School of Computing, National University of Singapore
38mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Proposes a conditional preference optimization algorithm to reduce hallucinations without degrading general capabilities.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Jun 17, 2024
preference optimization
Authors: Fei Wang, Wenxuan Zhou, James Y. Huang·Corresponding: not specified·Affiliation: University of Southern California; University of California, Davis; Microsoft Research
39CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
Uses a pre-trained CLIP model directly as an AI judge to generate preference pairs for DPO training.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Aug 19, 2024
preference optimization
Authors: Yassine Ouali, Adrian Bulat, Brais Martinez·Corresponding: not specified·Affiliation: Samsung AI Center Cambridge, UK; Technical University of Ia \textcommabelow si, Romania; Queen Mary University of London, UK
40EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
Proposes an efficient fine-grained unlearning framework to "forget" the tendency to hallucinate via reverse parameter updates.
unlearningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Feb 15, 2024
unlearning
Authors: Shangyu Xing, Fei Zhao, Zhen Wu·Corresponding: not specified·Affiliation: National Key Laboratory for Novel Software Technology, Nanjing University, China
41Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
Introduces a hallucination-inducing mechanism during training to construct sample pairs for parameter optimization.
contrastive tuningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted May 24, 2024
contrastive tuning
Authors: Xinyu Lyu, Beitao Chen, Lianli Gao·Corresponding: not specified·Affiliation: Center for Future Media, University of Electronic Science and Technology of China; Shenzhen Institute for Advanced Study, UESTC
42Mitigating Open-Vocabulary Caption Hallucinations
Trains a reward model based on Natural Language Inference (NLI) and uses RL to fine-tune the MLLM.
reinforcement learningreward modeling
Mitigation
vision-language
arXiv 2023
First posted Dec 6, 2023
rlhf
Authors: Assaf Ben-Kish, Moran Yanuka, Morris Alper·Corresponding: not specified·Affiliation: Tel-Aviv University
43Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
Constructs targeted repair instruction datasets for different hallucination types to perform targeted fine-tuning.
data curationfine-tuningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Apr 16, 2024
fine-tuning
Authors: Rui Hu, Yahan Tu, Shuyu Wei·Corresponding: not specified·Affiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
44See or Guess: Counterfactually Regularized Image Captioning
Introduces a counterfactual regularization term during training to force distinction between seen visuals and guessed priors.
regularization trainingmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Aug 29, 2024
regularization training
Authors: Qian Cao, Xu Chen, Ruihua Song·Corresponding: not specified·Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China; Tencent AI Lab
45Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
Generates targeted preference data specifically for hallucinations and applies DPO to penalize hallucination-prone features.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Nov 15, 2024
preference optimization
Authors: Yuhan Fu, Ruobing Xie, Xingwu Sun·Corresponding: not specified·Affiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
46Silkie: Preference Distillation for Large Visual Language Models
Uses AI to construct the VLFeedback dataset and distills preferences into the model via DPO.
preference optimizationdata curation
Mitigation
vision-language
arXiv 2023
First posted Dec 17, 2023
preference distillation
Authors: Lei Li, Zhihui Xie, Mukai Li·Corresponding: not specified·Affiliation: The University of Hong Kong; The Chinese University of Hong Kong (Shenzhen); Peking University
47Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
Constructs hard negative samples via data augmentation and introduces a contrastive loss to strengthen visual grounding.
visual groundingcontrastive tuningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted May 28, 2024
contrastive tuning
Authors: Pritam Sarkar, Sayna Ebrahimi, Ali Etemad·Corresponding: not specified·Affiliation: Queen's University and Vector Institute; Google DeepMind; Queen's University; Google Cloud AI Research
48Mitigating Hallucination in Visual Language Models with Visual Supervision
Constructs a fine-grained RAI-30k dataset and integrates the SAM model during instruction tuning.
instruction tuningdata curation
Mitigation
vision-language
arXiv 2023
First posted Nov 27, 2023
instruction tuning
Authors: Zhiyang Chen, Yousong Zhu, Yufei Zhan·Corresponding: not specified·Affiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Peng Cheng Laboratory; Wuhan AI Research
49Mitigating Multilingual Hallucination in Large Vision-Language Models
Builds a correctional dataset specifically for multilingual scenarios to mitigate cross-lingual hallucinations caused by language bias.
instruction tuningdata curation
Mitigation
vision-language
arXiv 2024
First posted Aug 1, 2024
instruction tuning
Authors: Xiaoye Qu, Mingyang Song, Wei Wei·Corresponding: Wei Wei·Affiliation: School of Computer Science & Technology, Huazhong University of Science and Technology; School of Computing Science and Technology, Fudan University; Wangxuan Institute of Computer Technology, Peking University; School of Computing Science and Technology, Zhejiang Gongshang University; Department of Computer Science and Engineering, The Chinese University of Hong Kong
50Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
Proposes a modality-fair preference optimization algorithm to balance penalties and prevent language prior dominance.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Oct 20, 2024
preference optimization
Authors: Songtao Jiang, Yan Zhang, Ruizhe Chen·Corresponding: not specified·Affiliation: Zhejiang University; National University of Singapore
51Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
Proposes a token-level preference optimization strategy based on self-calibrated visual-anchored rewards.
preference optimizationuncertainty
Mitigation
vision-language
arXiv 2024
First posted Dec 19, 2024
preference optimization
Authors: Jihao Gu, Yingyao Wang, Meng Cao·Corresponding: not specified·Affiliation: Alibaba Group; Mohamed bin Zayed University of Artificial Intelligence
52Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
Introduces Retrieval-Augmented Generation (RAG) into preference data construction to guide DPO training.
preference optimizationretrieval
Mitigation
vision-language
arXiv 2025
First posted Feb 18, 2025
preference optimization
Authors: Shuo Xing, Peiran Li, Yuping Wang·Corresponding: not specified·Affiliation: Texas A&M University; University of Michigan; UIUC; UNC Chapel Hill
53Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
Proposes symmetrical visual contrastive optimization using minimal contrastive images for alignment fine-tuning.
contrastive optimizationmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Feb 19, 2025
contrastive optimization
Authors: Shengguang Wu, Fan-Yun Sun, Kaiyue Wen·Corresponding: not specified·Affiliation: Stanford University
54FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
Uses fine-grained AI feedback to construct preference data and applies DPO for alignment training.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Apr 7, 2024
preference optimization
Authors: Liqiang Jing, Xinya Du·Corresponding: not specified·Affiliation: The University of Texas at Dallas
55TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Massively scales up text-centric visual instruction tuning data to specifically address text hallucinations in OCR tasks.
instruction tuningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Apr 19, 2024
instruction tuning
Authors: Jingqun Tang, Chunhui Lin, Zhen Zhao·Corresponding: not specified·Affiliation: ByteDance; East China Normal University; Huazhong University of Science and Technology
56Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Systematically builds datasets covering various alignment strategies to empirically verify their impact on hallucinations.
verificationdata curation
Mitigation
vision-language
arXiv 2024
First posted Jul 2, 2024
fine-tuning
Authors: Elmira Amirloo, Jean-Philippe Fauconnier, Christoph Roesmann·Corresponding: not specified·Affiliation: Apple Inc
57RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
Optimizes a visual reward model using auxiliary text-only preference data during RLHF to enhance hallucination discrimination.
reinforcement learningreward modeling
Mitigation
vision-language
arXiv 2024
First posted Aug 22, 2024
reward modeling
Authors: Chenglong Wang, Yang Gan, Yifu Huo·Corresponding: Chunliang Zhang·Affiliation: School of Computer Science and Engineering, Northeastern University, Shenyang, China; NiuTrans Research, Shenyang, China; CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China
58HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
Proposes a fine-tuning framework combining vision-enhanced penalty decoding with hierarchical feedback learning for behavior alignment.
feedback learningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Sep 30, 2024
feedback learning
Authors: Fan Yuan, Chi Qin, Xiaogang Xu·Corresponding: not specified·Affiliation: College of Artificial Intelligence; Nanjing University of Aeronautics and Astronautics, Nanjing, China; MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China; The Chinese University of Hong Kong, Hong Kong, China
59Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
Proposes entity-centric multimodal preference optimization, focusing penalty granularity precisely on specific visual entity tokens.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jun 4, 2025
preference optimization
Authors: Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang·Corresponding: not specified·Affiliation: Soochow University, Suzhou, China; Baidu Inc., Beijing, China; Shandong University, Qingdao, China
60Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
Introduces faithful and concise reasoning rationales as supervision signals during training to enhance logical generation.
multimodal reasoningrationale trainingmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Apr 17, 2024
rationale training
Authors: Minghe Gao, Shuang Chen, Liang Pang·Corresponding: not specified·Affiliation: Zhejiang University; Chinese Academy of Sciences; National University of Singapore; Sun Yat-sen University
61Generating Faithful and Salient Text from Multimodal Data
Constructs a multimodal factual consistency dataset and fine-tunes the model to improve faithfulness in data-to-text generation.
data curationfine-tuningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Sep 6, 2024
fine-tuning
Authors: Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya·Corresponding: not specified·Affiliation: Department of Data Science & AI, Monash University, Australia; Department of Data Science, Monash University, Indonesia; Department of CSE, Bangladesh University of Engineering and Technology, Bangladesh
62Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
Formulates preference modeling as a next-token prediction task to update the model using fine-grained verifier data.
verificationpreference modelingmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Oct 18, 2024
preference modeling
Authors: Chenhang Cui, An Zhang, Yiyang Zhou·Corresponding: An Zhang·Affiliation: National University of Singapore; UNC-Chapel Hill; Chicago University; Nanyang Technological University
63Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
Constructs a dataset targeting verb concept hallucinations and applies specific instruction tuning to correct action biases.
instruction tuningdata curation
Mitigation
vision-language
arXiv 2024
First posted Dec 6, 2024
instruction tuning
Authors: Zehao Wang, Xinpeng Liu, Yudonglin Zhang·Corresponding: not specified·Affiliation: Shanghai Jiao Tong University; ARC Lab, Tencent PCG
64Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution
Automatically generates and filters high-quality alignment data through an iterative self-evolution mechanism.
data curationself-evolution learningmodel alignment
Mitigation
vision-language
arXiv 2024
First posted Dec 20, 2024
self-evolution learning
Authors: Wentao Tan, Qiong Cao, Yibing Zhan·Corresponding: not specified·Affiliation: South China University of Technology; JD Explore Academy, Beijing; Pazhou Lab, Guangzhou
65CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
Proposes cross-modal hierarchical DPO to penalize global image-text matching and local object features at different levels.
preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jan 28, 2025
preference optimization
Authors: Jinlan Fu, Shenzhen Huangfu, Hao Fei·Corresponding: not specified·Affiliation: National University of Singapore; Fudan University; Digital Twin Institute, Eastern Institute of Technology, Ningbo
66PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
Introduces subtle visual perturbations during training forward passes to force the learning of robust visual representations.
representation editingvisual perturbation trainingmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Mar 9, 2025
visual perturbation training
Authors: Cong Chen, Mingyu Liu, Chenchen Jing·Corresponding: not specified·Affiliation: Zhejiang University; WeChat Group; Zhejiang University of Technology
67Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy
Builds a low-level visual hallucination database and applies negative sampling to enhance awareness of low-level features.
negative sample fine-tuningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Mar 26, 2025
negative sample fine-tuning
Authors: Yinan Sun, Xiongkuo Min, Zicheng Zhang·Corresponding: Xiongkuo Min·Affiliation:
68Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
Adopts rationale-augmented instruction tuning to teach the model to generate critique logic before providing an answer.
instruction tuningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted May 12, 2025
instruction tuning
Authors: Zexian Yang, Dian Li, Dayan Wu·Corresponding: Dayan Wu, Gang Liu·Affiliation: Institute of Information Engineering, Chinese Academy of Sciences; Foundation Technology Center, Tencent PCG
69OViP: Online Vision-Language Preference Learning for VLM Hallucination
Proposes an online preference learning framework to dynamically generate preference pairs during training.
preference learningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted May 21, 2025
preference learning
Authors: Shujun Liu, Siyuan Wang, Zejun Li·Corresponding: not specified·Affiliation: Fudan University; University of Southern California; ByteDance
70BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
Employs a bijective maximum likelihood learning approach to suppress hallucinations by optimizing joint probability distributions.
maximum likelihood learningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted May 30, 2025
maximum likelihood learning
Authors: Huu-Thien Tran, Thanh-Dat Truong, Khoa Luu·Corresponding: not specified·Affiliation: CVIU Lab, University of Arkansas
71Stop learning it all to mitigate visual hallucination, Focus on the hallucination target
Stops learning redundant backgrounds and forces weight updates to focus exclusively on target areas causing hallucinations.
preference optimizationtarget-localized dpomodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jun 13, 2025
target-localized dpo
Authors: Dokyoon Yoon, Youngsook Song, Woomyong Park·Corresponding: not specified·Affiliation: SIONIC AI
72Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
Identifies and purges hallucinated samples from fine-tuning data, using the refined high-quality data for knowledge distillation.
knowledge distillationmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jul 7, 2025
knowledge distillation
Authors: Wenhao Li, Xiu Su, Jingyi Wu·Corresponding: not specified·Affiliation: University of Sydney; Central South University; Fudan University; Southeast University; HKUST; Sensetime Research
73ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
Proposes a closed-loop framework where the model generates predictions, performs backward verification, and uses errors as training signals.
verificationclosed-loop trainingmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jul 7, 2025
closed-loop training
Authors: Jianjiang Yang, Yanshu li, Ziyan Huang·Corresponding: not specified·Affiliation: Department of Computer Science, University of Bristol; School of Future Technology, South China University of Technology; Brown University
74Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
Analyzes object biases introduced during pre-training and debiases the model using a reverse objective function during fine-tuning.
unlearningbias unlearningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Aug 6, 2025
bias unlearning
Authors: Yifan Li, Kun Zhou, Wayne Xin Zhao·Corresponding: not specified·Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China; University of California, San Diego; DataCanvas Alaya NeW
75TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
Adopts an adaptive MinMax Token preference strategy to dynamically adjust reward allocation.
adaptive preference strategymodel alignment
Mitigation
vision-language
arXiv 2025
First posted Jul 29, 2025
adaptive preference strategy
Authors: Kejia Zhang, Keda Tao, Zhiming Luo·Corresponding: not specified·Affiliation: Xiamen University; Westlake University; DAMO Academy, Alibaba Group; AWS AI Lab, Amazon; Hupan Laboratory
76Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
Integrates the CHAIR metric, utilizing object-level matching scores as the preference reward for DPO.
preference optimizationobject-aware dpomodel alignment
Mitigation
vision-language
arXiv 2025
First posted Aug 27, 2025
object-aware dpo
Authors: Alberto Compagnoni, Davide Caffagni, Nicholas Moratelli·Corresponding: not specified·Affiliation: University of Modena and Reggio Emilia, Modena, Italy
77Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
Distinguishes "omission" and "fabrication" hallucination causes and constructs respective fine-tuning data for decoupled training.
decoupled fine-tuningmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Aug 30, 2025
decoupled fine-tuning
Authors: Guangzong Si, Hao Yin, Xianfei Li·Corresponding: not specified·Affiliation: USTC; CowaRobot
78Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
Proposes semantic curriculum preference optimization, training the model in increasing difficulty to smoothly eliminate hallucinations.
preference optimizationcurriculum preference optimizationmodel alignment
Mitigation
vision-language
arXiv 2025
First posted Sep 29, 2025
curriculum preference optimization
Authors: Yuanshuai Li, Yuping Yan, Junfeng Tang·Corresponding: not specified·Affiliation: School of Engineering, Westlake University, Hangzhou, China; School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China
79Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
Proposes a universal training framework across multiple alignment formats (e.g., DPO, PPO) to eliminate vision-language biases.
preference optimizationreinforcement learning
Mitigation
vision-language
arXiv 2025
First posted Nov 21, 2025
unified alignment framework
Authors: Jiaye Qian, Ge Zheng, Yuchen Zhu·Corresponding: Sibei Yang·Affiliation: School of Computer Science and Engineering, Sun Yat-sen University; ShanghaiTech University
80OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Extracts self-attention matrices to find and penalize attention flows overly reliant on summary tokens.
attention steeringattention penaltyinternal intervention
Mitigation
vision-language
CVPR 2024
First posted Nov 29, 2023
attention penalty
Authors: Qidong Huang, Xiaoyi Dong, Pan Zhang·Corresponding: not specified·Affiliation: University of Science and Technology of China; Shanghai AI Laboratory
81Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Fuses global and local attention feature matrices internally to enhance fine-grained perception.
attention steeringattention fusioninternal intervention
Mitigation
vision-language
CVPR 2025
First posted Jun 18, 2024
attention fusion
Authors: Wenbin An, Feng Tian, Sicong Leng·Corresponding: not specified·Affiliation: Xi'an Jiaotong University; Nanyang Technological University; Lenovo Research; SGIT AI Lab; University of Massachusetts Boston
82Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
Adjusts cross-modal attention by forcefully amplifying the weights of visual tokens.
attention steeringattention amplificationinternal intervention
Mitigation
vision-language
ECCV 2024
First posted Jul 31, 2024
attention amplification
Authors: Shi Liu, Kecheng Zheng, Wei Chen·Corresponding: not specified·Affiliation: State Key Lab of CAD&CG, Zhejiang University
83DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
Dives into the MLLM to mask specific attention connections overly dominated by text.
attention steeringattention maskinginternal intervention
Mitigation
vision-language
EMNLP 2024
First posted Oct 6, 2024
attention masking
Authors: Xuan Gong, Tianshi Ming, Xinpeng Wang·Corresponding: not specified·Affiliation: Department of Computer Science and Technology, Tongji University
84Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
Uses causal analysis to block specific causal attention paths causing hallucinations during inference.
attention steeringcausal attention disconnectioninternal intervention
Mitigation
vision-language
ICLR 2025
First posted Oct 7, 2024
causal attention disconnection
Authors: Guanyu Zhou, Yibo Yan, Xin Zou·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Tsinghua University
85Mitigating Object Hallucination via Concentric Causal Attention
Strengthens the causal association between local visual features and global semantics in attention matrices.
attention steeringconcentric causal attentioninternal intervention
Mitigation
vision-language
NIPS 2024
First posted Oct 21, 2024
concentric causal attention
Authors: Yun Xing, Yiheng Li, Ivan Laptev·Corresponding: not specified·Affiliation: Nanyang Technological University; MBZUAI
86Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
Projects hidden states into a hallucination space and strips them via orthogonalization to clear hallucinations.
representation editinglatent space orthogonalizationinternal intervention
Mitigation
vision-language
CVPR 2025
First posted Dec 18, 2024
latent space orthogonalization
Authors: Le Yang, Ziwei Zheng, Boxu Chen·Corresponding: not specified·Affiliation: Xi'an Jiaotong University, Xi'an 710049, China
87VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
Introduces visual-aware sparsification to directly drop irrelevant visual tokens prone to inducing hallucinations.
sparsification droppinginternal intervention
Mitigation
vision-language
CVPR 2025
First posted Jan 11, 2025
sparsification dropping
Authors: Xianwei Zhuang, Zhihong Zhu, Yuxin Xie·Corresponding: not specified·Affiliation: SECE of Peking University
88Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
Adaptively constrains inter-layer information flow to block false causal paths causing object hallucinations.
object hallucinationinformation flow constraintinternal intervention
Mitigation
vision-language
AAAI 2025
First posted Feb 28, 2025
information flow constraint
Authors: Jiaqi Bai, Hongcheng Guo, Zhongyuan Peng·Corresponding: Mohan Li·Affiliation: Cyberspace Institute of Advanced Technology, Guangzhou University, China; Huangpu Research School of Guangzhou University, China; CCSE, Beihang University, China; University of the Chinese Academy of Sciences, China
89ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
Amplifies low-level visual signals in multimodal layers to combat dilution by deep language features.
signal enhancementinternal intervention
Mitigation
vision-language
CVPR 2025
First posted Mar 17, 2025
signal enhancement
Authors: Hao Yin, Guangzong Si, Zilei Wang·Corresponding: not specified·Affiliation: University of Science and Technology of China
90MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
Reconstructs multimodal attention using a causal matrix based on Manhattan distance.
attention steeringmanhattan causal replacementinternal intervention
Mitigation
vision-language
ACMMM 2025
First posted Jul 12, 2025
manhattan causal replacement
Authors: Qiyan Zhao, Xiaofeng Zhang, Yiheng Li·Corresponding: not specified·Affiliation: FKLPRIU, Xiamen University of Technology, China; Shanghai Jiao Tong University, China; Nanyang Technological University, Singapore; Jilin University, China; Monash University, Australia; Zhejiang University, China; Huizhou University, China; Chinese Academy of Sciences, China
91Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
Intervenes from the visual processing architecture to fundamentally alter how the model acquires visual information.
pure vision reinforcementinternal intervention
Mitigation
vision-language
EMNLP 2025
First posted Sep 17, 2025
pure vision reinforcement
Authors: Weihang Wang, Xinhao Li, Ziyue Wang·Corresponding: not specified·Affiliation: Bilibili; UESTC; University of Virginia
92Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination
Maps hallucination activations toward trusted attention regions with mixture-Gaussian bridges.
attention steeringrepresentation editing
Mitigation
vision-language
WACV 2026
First posted Feb 10, 2026
attention manifold alignment
Authors: Ziqiang Shi, Rujie Liu, Shanshan Yu·Corresponding: not specified·Affiliation: Fujitsu Research & Development Center Co.,LTD., Beijing, China
93RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Predicts input-specific latent steering vectors with gradient correction in a single forward pass.
attention steeringrectified-flow interventioninternal intervention
Mitigation
vision-language
rectified-flow intervention
Authors: Junyu Cheng, Zhibiao Liang, Yidong Chen·Corresponding: not specified·Affiliation: Department of Artificial Intelligence, School of Informatics, Xiamen University, China; School of Computer Science, South China Normal University, China; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China
94Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models
Filters intermediate attention maps to isolate dominant phantom text tokens and strengthen visual anchor tokens.
attention steeringdata curation
Mitigation
vision-language
token-asymmetric filtering
Authors: Shuyi Ouyang, Hongyi Wang, Gongfan Fang·Corresponding: not specified·Affiliation: Zhejiang University; National University of Singapore
95Look Before You Speak: Visual Re-Focusing for Hallucination Mitigation in Large Vision-Language Models
Reorients the model toward image evidence before response generation to reduce visually unsupported claims.
visual re-focusinginternal intervention
Mitigation
vision-language
visual re-focusing
Authors: Zheyuan Zhang, Yingmin Liu, Yang Shu·Corresponding: Yang Shu·Affiliation: Northwestern Polytechnical University, Xi'an, China; Zhejiang University, Hangzhou, China
96Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
Uses an attention lens to locate hallucinated features in middle layers and masks them for correction.
attention steeringmiddle layer maskinginternal intervention
Mitigation
vision-language
arXiv 2024
First posted Nov 23, 2024
middle layer masking
Authors: Zhangqi Jiang, Junkai Chen, Beier Zhu·Corresponding: not specified·Affiliation: National University of Defense Technology; Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and; Its Interdisciplinary Applications (Southeast University), Ministry of Education, China; Nanyang Technological University
97Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
Quantifies perception divergence and dynamically masks specific heads that lose visual alignment.
visual groundingattention steering
Mitigation
vision-language
arXiv 2024
First posted Dec 18, 2024
attention head pruning
Authors: Jinghan He, Kuan Zhu, Haiyun Guo·Corresponding: not specified·Affiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; University of Science and Technology of China; Southeast University; National University of Singapore
98Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
Reallocates Attention Maps during inference to pull misplaced attention back to visual targets.
attention steeringattention reallocationinternal intervention
Mitigation
vision-language
arXiv 2025
First posted Mar 11, 2025
attention reallocation
Authors: Chongjun Tu, Peng Ye, Dongzhan Zhou·Corresponding: Tao Chen·Affiliation: Fudan University; The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory; StepFun
99ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Proposes cross-level trusted intervention to block the drift of object representations toward false language priors.
representation editingcross-level blockinginternal intervention
Mitigation
vision-language
arXiv 2024
First posted Nov 22, 2024
cross-level blocking
Authors: Junzhe Chen, Tianshu Zhang, Shiyu Huang·Corresponding: not specified·Affiliation: Tsinghua University; The Hong Kong University of Science and Technology (Guangzhou); Zhipu AI; Chongqing University; Shanghai Jiao Tong University.
100Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
Implements causal intervention to cut off spurious feature maps hidden in attention matrices.
attention steeringcausal attention interventioninternal intervention
Mitigation
vision-language
arXiv 2024
First posted Dec 4, 2024
causal attention intervention
Authors: Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen·Corresponding: not specified·Affiliation: Inventec Corporation, No. 66, Hougang St., Shihlin Dist., Taipei City; University at Albany - SUNY, Albany, New York
101From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
Locates the evolution of object features and corrects visual tokens via early layer interventions.
early layer interventioninternal intervention
Mitigation
vision-language
arXiv 2024
First posted Oct 9, 2024
early layer intervention
Authors: Yuying Shang, Xinyi Zeng, Yutao Zhu·Corresponding: not specified·Affiliation: Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing, China; Gaoling School of Artificial Intelligence, Renmin University of China; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University, Beijing, China; Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University; Kuaishou Technology Inc., Beijing, China
102Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
Accurately locates key attention heads and specifically enhances or suppresses them during forward propagation.
attention steeringhead interventioninternal intervention
Mitigation
vision-language
arXiv 2024
First posted Nov 15, 2024
head intervention
Authors: Xiaofeng Zhang, Yihao Quan, Chaochen Gu·Corresponding: not specified·Affiliation: Shanghai Jiao Tong University; Alibaba Group; Beijing Jiaotong University
103Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
Calibrates the weight allocation of attention heads to weaken the implicit coverage of visual inputs by language priors.
uncertaintyattention steering
Mitigation
vision-language
arXiv 2024
First posted May 28, 2024
attention calibration
Authors: Sangmin Woo, Donguk Kim, Jaehyuk Jang·Corresponding: not specified·Affiliation: KAIST
104Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Introduces dual-attention mechanisms to internally rebalance language priors and image feature flows.
attention steeringdual-attention balancinginternal intervention
Mitigation
vision-language
arXiv 2024
First posted Nov 21, 2024
dual-attention balancing
Authors: Haozhe Zhao, Shuzheng Si, Liang Chen·Corresponding: Baobao Chang·Affiliation: University of Illinois Urbana-Champaign; Peking University; Tsinghua University
105CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
Cuts off implicit semantic biases introduced during cross-lingual transitions inside the model.
attention steeringcross-lingual attention interventioninternal intervention
Mitigation
vision-language
arXiv 2025
First posted Jun 3, 2025
cross-lingual attention intervention
Authors: Zekai Ye, Qiming Li, Xiaocheng Feng·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory; Central South University; Huawei Technologies Co., Ltd
106HalLoc: Token-level Localization of Hallucinations for Vision Language Models
Pinpoints tokens sourcing hallucinations and excises them during forward propagation.
token localization & pruninginternal intervention
Mitigation
vision-language
arXiv 2025
First posted Jun 12, 2025
token localization & pruning
Authors: Eunkyu Park, Minyeong Kim, Gunhee Kim·Corresponding: not specified·Affiliation: Seoul National University
107Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Identifies and edits vision-language feature vectors representing hallucinations in intermediate layers.
representation editinginternal intervention
Mitigation
vision-language
arXiv 2024
First posted Oct 3, 2024
representation editing
Authors: Nick Jiang, Anish Kachinthaya, Suzie Petryk·Corresponding: not specified·Affiliation: University of California, Berkeley
108Reducing Hallucinations in Vision-Language Models via Latent Space Steering
Injects pre-trained steering vectors into the latent space to guide generation away from hallucinations.
attention steeringlatent space steeringinternal intervention
Mitigation
vision-language
arXiv 2024
First posted Oct 21, 2024
latent space steering
Authors: Sheng Liu, Haotian Ye, Lei Xing·Corresponding: not specified·Affiliation: Stanford University
109Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
Penalizes decoding by contrasting the logit output distributions of original and distorted images during inference.
contrastive decodinggeneration control
Mitigation
vision-language
CVPR 2024
First posted Nov 28, 2023
contrastive decoding
Authors: Sicong Leng, Hang Zhang, Guanzheng Chen·Corresponding: not specified·Affiliation: DAMO Academy, Alibaba Group; Nanyang Technological University; Hupan Lab, 310023, Hangzhou, China
110Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
Introduces Classifier-Free Guidance (CFG) during inference to amplify the influence of visual features in the generation process.
inference guidancegeneration control
Mitigation
vision-language
CVPR 2024
First posted Feb 13, 2024
inference guidance
Authors: Linxi Zhao, Yihe Deng, Weitong Zhang·Corresponding: Quanquan Gu·Affiliation: Department of Computer Science, Cornell University, Ithaca, NY, USA; Department of Computer Science, University of California, Los Angeles, CA, USA; School of Data Science and Society, UNC, Chapel Hill, NC, USA
111IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
Penalizes pure language priors by contrasting the logits of inputs with and without images.
image-biased decodinggeneration control
Mitigation
vision-language
ECCV 2024
First posted Feb 28, 2024
image-biased decoding
Authors: Lanyun Zhu, Deyi Ji, Tianrun Chen·Corresponding: not specified·Affiliation: Singapore University of Technology and Design; Alibaba Group; Zhejiang University
112HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Proposes adaptive focal-contrast decoding to fine-grainedly suppress hallucinations at the local token level.
contrastive decodingfocal contrastive decodinggeneration control
Mitigation
vision-language
ICML 2024
First posted Mar 1, 2024
focal contrastive decoding
Authors: Zhaorun Chen, Zhuokai Zhao, Hongyin Luo·Corresponding: Zhaorun Chen, Zhuokai Zhao, Bo Li, Jiawei Zhou·Affiliation: University of Chicago, Chicago IL, USA; University of Illinois at Urbana-Champaign, Champaign IL, USA; Massachusetts Institute of Technology, Boston MA, USA; UNC-Chapel Hill, Chapel Hill NC, USA; Toyota Technological Institute at Chicago, Chicago IL, USA
113Multi-Modal Hallucination Control by Visual Information Grounding
Introduces a visual grounding mechanism during decoding to correct biases by contrasting different information streams.
visual groundingvisual information contrastgeneration control
Mitigation
vision-language
CVPR 2024
First posted Mar 20, 2024
visual information contrast
Authors: Alessandro Favero, Luca Zancato, Matthew Trager·Corresponding: not specified·Affiliation: AWS AI Labs
114Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
Constructs positive and negative instructions, contrasting their logits during inference to guide decoding.
contrastive decodinginstruction contrastive decodinggeneration control
Mitigation
vision-language
ACL 2024
First posted Mar 27, 2024
instruction contrastive decoding
Authors: Xintong Wang, Jingheng Pan, Liang Ding·Corresponding: not specified·Affiliation: Department of Informatics, Universität Hamburg; The University of Sydney
115Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
Introduces residual visual connections to add early visual features directly into subsequent decoding computations.
residual decodinggeneration control
Mitigation
vision-language
ACL 2024
First posted Jun 30, 2024
residual decoding
Authors: Weihong Zhong, Xiaocheng Feng, Liang Zhao·Corresponding: not specified·Affiliation: Harbin Institute of Technology; Peng Cheng Laboratory
116ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
Uses attention visualization to locate hallucination regions and constructs interference images for contrastive decoding.
contrastive decodingattention steering
Mitigation
vision-language
AAAI 2025
First posted Aug 25, 2024
visualized contrastive decoding
Authors: Yeji Park, Deokyeong Lee, Junsuk Choe·Corresponding: Junsuk Choe, Buru Chang·Affiliation: Sogang University
117Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Introduces a generative feedback loop to detect and correct the probability distribution of the current token in real time.
self-correctionself-correcting decodinggeneration control
Mitigation
vision-language
ICLR 2025
First posted Feb 10, 2025
self-correcting decoding
Authors: Ce Zhang, Zifu Wan, Zhehan Kan·Corresponding: not specified·Affiliation: School of Computer Science, Carnegie Mellon University
118Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
Adaptively adjusts the contrastive penalty based on image complexity and local token features.
contrastive decodingdynamic contrastive decodinggeneration control
Mitigation
vision-language
CVPR 2025
First posted Mar 1, 2025
dynamic contrastive decoding
Authors: Wei Suo, Lijun Zhang, Mengyang Sun·Corresponding: Yanning Zhang·Affiliation: School of Computer Science and Ningbo Institute, Northwestern Polytechnical University,China.; School of Cybersecurity, Northwestern Polytechnical University, China.; Computer Science, Swansea University.
119Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
Adaptively switches between different sampling algorithms based on attention strength.
attention steeringmixture of decodinggeneration control
Mitigation
vision-language
ACL 2025
First posted May 17, 2025
mixture of decoding
Authors: Xinlong Chen, Yuanxing Zhang, Qiang Liu·Corresponding: not specified·Affiliation: New Laboratory of Pattern Recognition (NLPR); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences; Kuaishou Technology; Nanjing University
120Med-VCD: Mitigating hallucination for medical large vision language models through visual contrastive decoding
Selects visually informed tokens on the fly, removing redundant tokens while retaining critical medical image context.
contrastive decodingsparse visual contrastive decodinggeneration control
Mitigation
vision-language
Computers in Biology and Medicine 2026
First posted Dec 1, 2025
sparse visual contrastive decoding
Authors: Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman Khaloo·Corresponding: not specified·Affiliation: Department of computer science, University of Central Florida, Orlando, USA; Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA; Department of electrical engineering, Columbia university, New York, NY, USA; Technical University of Applied Sciences Regensburg, Regensburg, Germany; Department of Surgery, University of Calgary, Calgary, Alberta, Canada; University College of Nabi Akram, Tabriz, Iran; School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
121Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
Constructs an image-free text input as a negative reference and subtracts language prior scores during logit calculation.
contrastive decodinglanguage contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Aug 6, 2024
language contrastive decoding
Authors: Avshalom Manevich, Reut Tsarfaty·Corresponding: not specified·Affiliation: Bar Ilan University
122CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
Contrasts logit distributions based on the original image versus self-generated descriptions to penalize text-reliant tokens.
contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Jun 4, 2024
contrastive decoding
Authors: Junho Kim, Hyunjun Kim, Yeonju Kim·Corresponding: Junho Kim·Affiliation: Integrated Vision and Language Lab, KAIST
123EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Suppresses event hallucinations by contrasting output probabilities between full and truncated videos.
contrastive decodingtemporal contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Sep 25, 2024
temporal contrastive decoding
Authors: Jiacheng Zhang, Yang Jiao, Shaoxiang Chen·Corresponding: not specified·Affiliation: Shanghai Key Lab of Intelligent Information Processing, School of CS, Fudan University; Shanghai Collaborative Innovation Center on Intelligent Visual Computing; Singapore University of Technology and Design; Shanghai Academy of Artificial Intelligence for Science; Meituan
124Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
Introduces an external CLIP model to score and guide candidate tokens during decoding to improve image-text consistency.
visual groundingexternal tools
Mitigation
vision-language
arXiv 2024
First posted Feb 23, 2024
external guided decoding
Authors: Ailin Deng, Zhirui Chen, Bryan Hooi·Corresponding: not specified·Affiliation: School of Computing, National University of Singapore; Department of Industrial Systems Engineering and Management, National University of Singapore
125Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
Dynamically adjusts the logit weight distribution of text and image features during inference.
contrastive decodingre-balancing contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Sep 10, 2024
re-balancing contrastive decoding
Authors: Xiaoyu Liang, Jiayuan Yu, Lianrui Mu·Corresponding: Haoji Hu·Affiliation: College of Information Science and Electronic Engineering, Zhejiang University, China
126Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
Extracts internally generated facts as references for contrastive decoding to suppress over-extrapolation.
contrastive decodinginternal fact contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2025
First posted Feb 3, 2025
internal fact contrastive decoding
Authors: Chao Wang, Xuancheng Zhou, Weiwei Fu·Corresponding: Chao Wang, Yang Zhou·Affiliation: School of Future Technology, Shanghai University, Shanghai, 200444, China; Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China; School of Mechatronic Engineering and Automation, Shanghai, 200444, China
127HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Introduces a control parameter during decoding to force switching between pure contextual description and parameterized knowledge imagination.
inference interventiongeneration control
Mitigation
vision-language
arXiv 2023
First posted Oct 3, 2023
inference intervention
Authors: Bohan Zhai, Shijia Yang, Chenfeng Xu·Corresponding: not specified·Affiliation: ByteDance Inc.; Stanford University; UC Berkeley; UIUC
128Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework
Introduces language contrastive decoding to weaken misleading priors from user prompts to combat sycophancy-induced hallucinations.
contrastive decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Aug 21, 2024
contrastive decoding
Authors: Yunpu Zhao, Rui Zhang, Junbin Xiao·Corresponding: Rui Zhang·Affiliation: School of Computer Science and Technology, University of Science and Technology of China; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences; Department of Computer Science, National University of Singapore; University of Illinois Urbana-Champaign; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences
129Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
Utilizes image summaries as guidance signals to intervene in output probabilities for contextual consistency.
summary-guided decodinggeneration control
Mitigation
vision-language
arXiv 2024
First posted Oct 17, 2024
summary-guided decoding
Authors: Kyungmin Min, Minbeom Kim, Kang-il Lee·Corresponding: not specified·Affiliation: IPAI, Seoul National University; Dept. of ECE, Seoul National University
130Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
Contrasts logits from the original image and retrieved similar reference images to highlight core facts.
contrastive decodingretrieval
Mitigation
vision-language
arXiv 2025
First posted May 26, 2025
retrieval visual contrastive decoding
Authors: Jihoon Lee, Min Song·Corresponding: not specified·Affiliation: Yonsei University; Onoma AI
131CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
Generates multi-granularity hierarchical feedback to dynamically correct outputs.
coarse-to-fine feedback decodinggeneration control
Mitigation
vision-language
arXiv 2025
First posted Dec 29, 2025
coarse-to-fine feedback decoding
Authors: Zongsheng Cao, Yangfan He, Anran Liu·Corresponding: Anran Liu, Zepeng Wang·Affiliation: Researcher; UMN; PCIE
132Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Trains an independent Reviser model specifically to detect and reconstruct hallucinated entities in outputs.
self-correctionpost-hoc revisionexternal grounding
Mitigation
vision-language
ICLR 2023
First posted Oct 1, 2023
post-hoc revision
Authors: Yiyang Zhou, Chenhang Cui, Jaehong Yoon·Corresponding: not specified·Affiliation: UNC-Chapel Hill; Rutgers University; Columbia University; Stanford University
133Woodpecker: Hallucination Correction for Multimodal Large Language Models
Calls external visual foundation models to verify entity facts, finally using an LLM to rewrite and correct.
verificationexternal tools
Mitigation
vision-language
SCIS 2024
First posted Oct 24, 2023
agent correction
Authors: Shukang Yin, Chaoyou Fu, Sirui Zhao·Corresponding: not specified·Affiliation: School of Data Science, USTC & State Key Laboratory of Cognitive Intelligence; Tencent YouTu Lab
134Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
Constructs a pipeline system containing claim extraction, external detection verification, and corrected generation.
verificationexternal tools
Mitigation
vision-language
CVPR 2024
First posted Apr 30, 2024
pipeline agent
Authors: Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman·Corresponding: not specified·Affiliation: NVIDIA
135Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning
Modifies reference images to expose hallucinating agents and replaces consensus-only debate with evidence-based factual verification.
verificationcounterfactual multi-agent gamingexternal grounding
Mitigation
vision-language
counterfactual multi-agent gaming
Authors: Dayong Liang, Xiao-Yong Wei, Changmeng Zheng·Corresponding: not specified·Affiliation: South China University of Technology, Guangzhou, China; Peng Cheng Laboratory, Shenzhen, China; Sichuan University, Chengdu, China; The Hong Kong Polytechnic University, Hong Kong, China
136Mitigating Object and Relationship Hallucination in Large Vision Language Model with Multi-Agent Guidance
Coordinates specialized agents to identify and correct unsupported object and relationship claims.
verificationmulti-agent verificationexternal grounding
Mitigation
vision-language
multi-agent verification
Authors: Soohyun Kim, Gusang Lee, Kyuhong Shim·Corresponding: not specified·Affiliation: Seoul National University, Department of Electrical and Computer Engineering, Korea; Sungkyunkwan University, Department of Computer Science and Engineering, Korea
137CEBC: Conformal Evidence-Bounded Control for Low-Hallucination Vision-Language Generation
Calibrates an external detector on held-out data and minimally revises unsupported object mentions under explicit risk bounds.
self-correctionexternal toolsuncertainty
Mitigation
vision-language
conformal evidence-bounded editing
Authors: Ashish Mishra, Tarun Kumar, Arpit Shah·Corresponding: not specified·Affiliation: Hewlett Packard Labs, Bangalore; Hewlett Packard Labs, USA
138Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
Uses external language models to ask questions about generated claims for self-verification.
verificationexternal tools
Mitigation
vision-language
arXiv 2024
First posted Feb 18, 2024
logical loop verification
Authors: Junfei Wu, Qiang Liu, Ding Wang·Corresponding: not specified·Affiliation: Center for Research on Intelligent Perception and Computing; State Key Laboratory of Multimodal Artificial Intelligence Systems; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Nanjing University
139Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
Utilizes a bottom-up holistic reasoning framework combined with multi-perspective prompt verification.
verificationmultimodal reasoning
Mitigation
vision-language
arXiv 2024
First posted Dec 15, 2024
multi-perspective cross-checking
Authors: Shengqiong Wu, Hao Fei, Liangming Pan·Corresponding: Hao Fei·Affiliation: National University of Singapore, Singapore; University of Arizona, USA; University of California, Santa Barbara, USA; Skywork AI, Singapore; Nanyang Technological University, Singapore
140Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
Uses an external detector to locate relation errors, then applies template-based secondary correction.
external toolsuncertainty
Mitigation
vision-language
arXiv 2024
First posted Aug 18, 2024
external detection calibration
Authors: Kening Zheng, Junkai Chen, Yibo Yan·Corresponding: Xuming Hu·Affiliation: Hong Kong University of Science and Technology (Guangzhou); Guangxi Zhuang Autonomous Region Big Data Research Institute; Hong Kong University of Science and Technology
141Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
Adopts a "Retrospect-then-Compare" multi-step prompt to make the model review visual details for correction.
multi-step retrospective promptingexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Mar 21, 2024
multi-step retrospective prompting
Authors: Dingchen Yang, Bowen Cao, Guang Chen·Corresponding: not specified·Affiliation: Tongji University; Peking University
142Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
Explicitly adds spatial relation logic rules into the prompt to correct spatial confusion.
spatial logic rule promptingexternal grounding
Mitigation
vision-language
arXiv 2025
First posted Feb 12, 2025
spatial logic rule prompting
Authors: Jiarui Wu, Zhuo Liu, Hangfeng He·Corresponding: not specified·Affiliation: University of Rochester
143RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
Applies random geometric transformations to construct multiple views for consistency verification and correction.
verificationmulti-view consistencyexternal grounding
Mitigation
vision-language
arXiv 2024
First posted May 28, 2024
multi-view consistency
Authors: Sangmin Woo, Jaehyuk Jang, Donguk Kim·Corresponding: not specified·Affiliation: KAIST
144Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
The model generates a draft, then independently reviews and self-scores sub-topics for correction.
topic-level self-scoringexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Nov 26, 2024
topic-level self-scoring
Authors: Lehan He, Zeren Chen, Zhelun Shi·Corresponding: not specified·Affiliation: Shanghai AI Laboratory; School of Software, Beihang University; Shanghai Innovation Institute; Tsinghua University
145InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
Deploys a network of introspective and external verification agents for interactive debate.
verificationexternal tools
Mitigation
vision-language
arXiv 2025
First posted Dec 2, 2025
cross-modal introspective debate
Authors: Zhongyu Yang, Yingfang Yuan, Xuanming Jiang·Corresponding: Xuanming Jiang, Wei Pang·Affiliation: Xi'an Jiyun Technology Co., Ltd., Xi'an, China; BCML, Heriot-Watt University, Edinburgh, UK
146Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
Evaluates model uncertainty to actively trigger knowledge base retrieval for entity correction.
retrievalexternal toolsuncertainty
Mitigation
vision-language
TOMM 2025
First posted Aug 1, 2024
dynamic rag supplement
Authors: Xiaoye Qu, Qiyuan Chen, Wei Wei·Corresponding: not specified·Affiliation: Huazhong University of Science and Technology; Zhejiang University; Xiamen University; Zhejiang Gongshang University
147A Unified Hallucination Mitigation Framework for Large Vision-Language Models
Builds a unified cross-modal diagnosis pipeline utilizing external discriminators to intercept and modify outputs.
external toolsexternal discriminator systemexternal grounding
Mitigation
vision-language
TMLR 2024
First posted Sep 24, 2024
external discriminator system
Authors: Yue Chang, Liqiang Jing, Xiaopeng Zhang·Corresponding: not specified·Affiliation: The University of Texas at Dallas
148Mitigating Object Hallucinations via Sentence-Level Early Intervention
Implements factual truncation prompts in early sentence stages of generation to prevent error propagation.
early sentence truncationexternal grounding
Mitigation
vision-language
ICCV 2025
First posted Jul 16, 2025
early sentence truncation
Authors: Shangpin Peng, Senqiao Yang, Li Jiang·Corresponding: not specified·Affiliation: Harbin Institute of Technology, Shenzhen; The Chinese University of Hong Kong; The Chinese University of Hong Kong, Shenzhen
149Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
Overlays visual prompts (e.g., circles, bounding boxes) on the original image to force model attention.
attention steeringoriginal image overlayexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Apr 17, 2024
original image overlay
Authors: Yichi Zhang, Yinpeng Dong, Siyuan Zhang·Corresponding: not specified·Affiliation: Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center; THBI Lab, BNRist Center, Tsinghua University, Beijing 100084, China; RealAI; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong
150What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
Introduces "What if" prompting to compel the model to self-reflect on its initial judgments.
counterfactual promptingexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Mar 20, 2024
counterfactual prompting
Authors: Junho Kim, Yeon Ju Kim, Yong Man Ro·Corresponding: Yeon Ju Kim·Affiliation: Integrated Vision and Language Lab KAIST South Korea
151From Training-Free to Adaptive: Empirical Insights into MLLMs’ Understanding of Detection Information
Feeds information identified by external object detection models directly to the MLLM as context.
external toolsexternal visual guidanceexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Jan 31, 2024
external visual guidance
Authors: Qirui Jiao, Daoyuan Chen, Yilun Huang·Corresponding: not specified·Affiliation: Sun Yat-Sen University, Shenzhen, China; Alibaba Group, Hangzhou, China
152Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
Pre-supplements missing low-level visual perception details in the input prompt to bridge the gap.
perception detail supplementexternal grounding
Mitigation
vision-language
arXiv 2024
First posted May 24, 2024
perception detail supplement
Authors: Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar·Corresponding: not specified·Affiliation: University of Maryland, College Park, USA; Adobe, USA
153Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
Strictly controls the requested detail quantity via prompt to balance richness and hallucination rate.
prompt detail constraintexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Jun 18, 2024
prompt detail constraint
Authors: Mingqian Feng, Yunlong Tang, Zeliang Zhang·Corresponding: not specified·Affiliation: University of Rochester
154Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
Adopts multi-view prompts and multi-path reasoning, aggregating and comparing answers.
multimodal reasoningmulti-path votingexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Aug 30, 2024
multi-path voting
Authors: Xiaoye Qu, Jiashuo Sun, Wei Wei·Corresponding: not specified·Affiliation: Huazhong University of Science and Technology; Xiamen University; The Chinese University of Hong Kong
155Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
Uses prompts to guide the model to actively re-extract visual clues for confirmation during generation.
memory retracing confirmationexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Oct 4, 2024
memory retracing confirmation
Authors: Xin Zou, Yizhou Wang, Yibo Yan·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; China University of Geosciences; University of Technology Sydney
156Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Introduces an external Vision Value Model to assist tree search, guiding autoregressive generation.
external toolsexternal scoring searchexternal grounding
Mitigation
vision-language
arXiv 2024
First posted Dec 4, 2024
external scoring search
Authors: Xiyao Wang, Zhengyuan Yang, Linjie Li·Corresponding: not specified·Affiliation: University of Maryland, College Park; Microsoft
157Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Uses black-box optimization to automatically search for "visual prompt patches" that suppress hallucinations.
external toolsautomated patch searchexternal grounding
Mitigation
vision-language
arXiv 2025
First posted Apr 30, 2025
automated patch search
Authors: Sangmin Woo, Kang Zhou, Yun Zhou·Corresponding: Sangmin Woo, Haibo Ding·Affiliation: Amazon AWS AI; KAIST
158Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models
Uses uncertainty-aware contrastive calibration to suppress hallucinated actions in vision-language-action generation.
uncertaintyentropy-optimized decodinggeneration control
Mitigation
vision-language
entropy-optimized decoding
Authors: Ye Qiu, Zhaoxin Fan, Qingchen Yu·Corresponding: not specified·Affiliation:
159VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models
Uses a reinforcement-learned caption model as an auxiliary visual clue source and applies image-confidence constraints during generation.
uncertaintyvisual-clue-guided decodinggeneration control
Mitigation
vision-language
visual-clue-guided decoding
Authors: Guoqing Chen, Fu Zhang, Bingqian Liu·Corresponding: not specified·Affiliation: Northeastern University
160CTDD: Cumulative Trend Divergence Decoding for Mitigating Hallucination in Large Vision-Language Models
Tracks accumulated divergence in token-distribution trends and calibrates generation when visual and linguistic evidence separate.
uncertaintycumulative trend-divergence decodinggeneration control
Mitigation
vision-language
cumulative trend-divergence decoding
Authors: Jiani Hou, Zhixuan You, Siyu Luo·Corresponding: Bing Guo·Affiliation: College of Software Engineering, Sichuan University, Chengdu, China; International Economics and Trade, Central University of Finance and Economics, Beijing, China; College of Computer Science, Sichuan University, Chengdu, China
161Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
Filters hallucinated words by comparing the confidence of different generation paths via the model's own introspective mechanism.
uncertaintydata curation
Mitigation
vision-language
arXiv 2024
First posted Aug 4, 2024
self-introspective decoding
Authors: Fushuo Huo, Wenchao Xu, Zhong Zhang·Corresponding: Wenchao Xu, Peilin Zhao·Affiliation: Department of Computing, The Hong Kong Polytechnic University; Division of Integrative Systems and Design, Hong Kong University of Science and Technology; Tencent AI Lab; Huazhong University of Science and Technology; Tsinghua University
162HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
Evaluates synthetic images / vqa with Gen tasks using Acc metrics.
object-levelgenerationsynthetic images vqa
Benchmarks
vision-language
ECCV 2024
First posted Jul 22, 2024
7,748 / Acc
Authors: Zhecan Wang, Garrett Bingham, Adams Yu·Corresponding: not specified·Affiliation: Columbia University, New York, NY 10027; Google DeepMind, Mountain View, CA 94043
163HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
Evaluates speech/sound/music with Dis & Gen tasks using Acc / Hallucination Rate metrics.
discriminationgenerationspeechsoundmusic
Benchmarks
vision-language
5,000+ / Acc / Hallucination Rate
Authors: Feiyu Zhao, Yiming Chen, Wenhuan Lu·Corresponding: Xianghu Yue·Affiliation: College of Intelligence and Computing, Tianjin University, China; ASUS Intelligent Cloud Services, Singapore
164Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
Evaluates medical context with Dis & Gen tasks using MediHall Score metrics.
object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language
arXiv 2024
First posted Jun 14, 2024
889,125 / MediHall Score
Authors: Jiawei Chen, Dingkang Yang, Tong Wu·Corresponding: not specified·Affiliation: Academy for Engineering and Technology, Fudan University, Shanghai, China; Tencent Youtu Lab, Shanghai, China; Cognition and Intelligent Technology Laboratory, Shanghai, China; Engineering Research Center of AI and Robotics, Ministry of Education, Shanghai, China; AI and Unmanned Systems Engineering Research Center of Jilin Province, Changchun, China
165VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Evaluates temporal/extrinsic with Dis & Gen tasks using Acc metrics.
object-levelrelation-leveldiscriminationgeneration
Benchmarks
vision-language
arXiv 2024
First posted Jun 24, 2024
1,800 / Acc
Authors: Yuxuan Wang, Yueqian Wang, Dongyan Zhao·Corresponding: not specified·Affiliation: Beijing Institute for General Artificial Intelligence, Beijing China; State Key Laboratory of General Artificial Intelligence, Beijing, China; Wangxuan Institute of Computer Technology, Peking University, Beijing, China; Computer Science and Engineering, University of California, Santa Cruz
166MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
Evaluates medical context with Dis & Gen tasks using Characterization Score metrics.
object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language
arXiv 2024
First posted Jul 3, 2024
2,000 / Characterization Score
Authors: Zishan Gu, Changchang Yin, Fenglin Liu·Corresponding: not specified·Affiliation: The Ohio State University; University of Oxford
167Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
Evaluates grounded medical with Gen tasks using Acc / Loc metrics.
object-levelgenerationgrounded medical
Benchmarks
vision-language
arXiv 2025
First posted Apr 30, 2025
67,000 / Acc / Loc
Authors: Dung Nguyen, Minh Khoi Ho, Huy Ta·Corresponding: not specified·Affiliation: Hanoi University of Science and Technology; University of Wollongong; Australian Institute for Machine Learning, The University of Adelaide; Griffith University; College of Medicine and Public Health, Flinders University
168MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
Evaluates multi-image reasoning with Dis & Gen tasks using Acc metrics.
object-leveldiscriminationgenerationmulti-image reasoning
Benchmarks
vision-language
arXiv 2025
First posted Aug 1, 2025
3,000 / Acc
Authors: Jiale Li, Mingrui Wu, Zixiang Jin·Corresponding: not specified·Affiliation: Xiamen University; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China; Zhongguancun Academy
169The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
Evaluates spurious images with Dis tasks using Acc metrics.
object-leveldiscriminationspurious images
Benchmarks
vision-language
arXiv 2024
First posted Feb 6, 2024
7,308 / Acc
Authors: Tianyang Han, Qing Lian, Rui Pan·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology; University of Illinois at Urbana-Champaign; The Hong Kong Polytechnic University
170Explore the Hallucination on Low-level Perception for MLLMs
Evaluates low-level perception with Gen tasks using Acc metrics.
attribute-levelgenerationlow-level perception
Benchmarks
vision-language
arXiv 2024
First posted Sep 15, 2024
4,989 / Acc
Authors: Yinan Sun, Zicheng Zhang, Haoning Wu·Corresponding: not specified·Affiliation: Shanghai Jiao Tong University; Nanyang Technological University
171JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
Evaluates imaginary vqa with Dis & Gen tasks using Acc metrics.
object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language
arXiv 2024
First posted Sep 19, 2024
13,500 / Acc
Authors: Zhecan Wang, Junzhang Liu, Chia-Wei Tang·Corresponding: not specified·Affiliation: Columbia University; UCLA; Virginia Tech
172Detecting and Preventing Hallucinations in Large Vision Language Models
Evaluates fine-grained with Dis tasks using Reward Score metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
AAAI 2024
First posted Aug 11, 2023
16,000 / Reward Score
Authors: Anisha Gunjal, Jihan Yin, Erhan Bas·Corresponding: not specified·Affiliation:
173Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
Evaluates negative pronoun with Dis tasks using Acc/METEOR metrics.
object-leveldiscriminationnegative pronoun
Benchmarks
vision-language
ACL-W 2024
First posted Oct 9, 2023
29,500 / Acc/METEOR
Authors: Holy Lovenia, Wenliang Dai, Samuel Cahyawijaya·Corresponding: not specified·Affiliation: The Hong Kong University of Science and Technology
174THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
Evaluates free-form generations with Gen tasks using P/R/F metrics.
object-levelgenerationfree-form generations
Benchmarks
vision-language
CVPR 2024
First posted May 8, 2024
5,000 / P/R/F
Authors: Prannay Kaul, Zhizhong Li, Hao Yang·Corresponding: not specified·Affiliation: VGG, University of Oxford; AWS AI Labs
175Quantity Matters: Towards Assessing and Mitigating Number Hallucination in Large Vision-Language Models
Evaluates object counting with Dis tasks using Acc / Consistency metrics.
object-levelattribute-leveldiscriminationobject counting
Benchmarks
vision-language
arXiv 2024
First posted Mar 3, 2024
20,000 / Acc / Consistency
Authors: Huixuan Zhang, Junzhe Zhang, Xiaojun Wan·Corresponding: not specified·Affiliation: School of Electronics Engineering and Computer Science, Peking University; Wangxuan Institute of Computer Technology, Peking University
176CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
Evaluates multimodal hallucination with Dis tasks using Acc/P/R/F1/Specificity metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
NeurIPS-W 2023
First posted Sep 5, 2023
40,367 / Acc/P/R/F1/Specificity
Authors: Hongyu Hu, Jiyuan Zhang, Minyi Zhao·Corresponding: not specified·Affiliation: ByteDance Inc, Shanghai
177Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
Evaluates visual grounding with Dis tasks using Accuracy metrics.
object-levelrelation-leveldiscriminationvisual grounding
Benchmarks
vision-language
CVPR 2025
First posted Dec 3, 2023
31,373 / Accuracy
Authors: Andrés Villa, Juan Carlos León Alcázar, Alvaro Soto·Corresponding: not specified·Affiliation: King Abdullah University of Science and Technology (KAUST); Pontificia Universidad Católica de Chile
178Unified Hallucination Detection for Multimodal Large Language Models
Evaluates t2i with Gen tasks using Acc/P/R/F metrics.
object-levelattribute-levelgenerationt2i
Benchmarks
vision-language
ACL 2024
First posted Feb 5, 2024
1,860 / Acc/P/R/F
Authors: Xiang Chen, Chenxi Wang, Yida Xue·Corresponding: Huajun Chen·Affiliation: College of Computer Science and Technology, Zhejiang University; School of Software Technology, Zhejiang University; Zhejiang University-Ant Group Joint Laboratory of Knowledge Graph; Ant Group
179Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
Evaluates event hallucination with Dis & Gen tasks using Acc/Score metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
ACMMM 2024
First posted Feb 24, 2024
10,000 / Acc/Score
Authors: Chaoya Jiang, Hongrui Jia, Wei Ye·Corresponding: not specified·Affiliation: National Engineering Research Center for Software Engineering, Peking University; DAMO Academy, Alibaba Group
180VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
Evaluates multimodal hallucination with Gen tasks using Faithfulness & Coverage metrics.
object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language
ACL 2024
First posted Apr 22, 2024
211 / Faithfulness & Coverage
Authors: Haoyi Qiu, Wenbo Hu, Zi-Yi Dou·Corresponding: not specified·Affiliation: University of California, Los Angeles
181AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Evaluates multimodal hallucination with Dis tasks using ASR/MASR/CASR metrics.
object-leveldiscrimination
Benchmarks
vision-language
EMNLP 2024
First posted Jun 16, 2024
5,000 / ASR/MASR/CASR
Authors: Xiyang Wu, Tianrui Guan, Dianqi Li·Corresponding: not specified·Affiliation: University of Maryland, College Park
182ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models
Evaluates open-set with Dis & Gen tasks using AMBER/Acc metrics.
object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language
CVPR 2025
First posted Sep 14, 2024
8,786 / AMBER/Acc
Authors: Yahan Tu, Rui Hu, Jitao Sang·Corresponding: not specified·Affiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
183Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
Evaluates ocr/action/counting with Dis & Gen tasks using Hallucination Rate metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
arXiv 2024
First posted Jun 24, 2024
4,000 / Hallucination Rate
Authors: Bei Yan, Jie Zhang, Zheng Yuan·Corresponding: not specified·Affiliation: Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Key Laboratory of Al Safety, Chinese Academy of Sciences
184FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
Evaluates eval method with Dis & Gen tasks using Acc/P/R/F1 metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
arXiv 2024
First posted Sep 20, 2024
N/A / Acc/P/R/F1
Authors: Bowen Yan, Zhengsong Zhang, Liqiang Jing·Corresponding: not specified·Affiliation: University of Texas at Dallas, Richardson, United States
185TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
Evaluates unanswerable qs with Dis tasks using Acc metrics.
object-levelrelation-leveldiscriminationunanswerable qs
Benchmarks
vision-language
arXiv 2024
First posted Oct 5, 2024
2,354 / Acc
Authors: Xingwei He, Qianru Zhang, A-Long Jin·Corresponding: not specified·Affiliation: The University of Hong Kong; Xi’an Jiaotong-Liverpool University; School of Computer Science and Engineering, Beihang University, Beijing, China; State Key Laboratory of Software, Development Environment; Zhongguancun Laboratory
186MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
Evaluates manipulation/ooc/veracity with Dis & Gen tasks using Acc/F1 metrics.
discriminationgenerationmanipulationoocveracity
Benchmarks
vision-language
arXiv 2024
First posted Jun 17, 2024
35,000 / Acc/F1
Authors: Shengkang Wang, Hongzhan Lin, Ziyang Luo·Corresponding: not specified·Affiliation: Beijing University of Posts and Telecommunications; Hong Kong Baptist University; Hong Kong University of Science and Technology
187Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
Evaluates perturbed inputs with Dis tasks using Acc metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
arXiv 2024
First posted Aug 2, 2024
1,260 / Acc
Authors: Peng Ding, Jingyu Wu, Jun Kuang·Corresponding: not specified·Affiliation: National Key Laboratory for Novel Software Technology, Nanjing University; College of Computer Science and Technology, Zhejiang University; Zhejiang-Singapore Innovation and AI Joint Research Lab, Zhejiang University
188CAST: Cross-modal Alignment Similarity Test for Vision Language Models
Evaluates self-consistency with Dis & Gen tasks using Similarity metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
arXiv 2024
First posted Sep 17, 2024
15,000 / Similarity
Authors: Gautier Dagan, Olga Loginova, Anil Batra·Corresponding: not specified·Affiliation: University of Edinburgh; University of Trento
189FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
Evaluates obj. counting with Gen tasks using FaithScore metrics.
object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language
EMNLP 2024
First posted Nov 2, 2023
2,000 / FaithScore
Authors: Liqiang Jing, Ruosen Li, Yunmo Chen·Corresponding: not specified·Affiliation: University of Texas at Dallas; Johns Hopkins University; University of Notre Dame
190ALOHa: A New Measure for Hallucination in Captioning Models
Evaluates captioning metric with Gen tasks using ALOHa metrics.
object-levelgenerationcaptioning metric
Benchmarks
vision-language
arXiv 2024
First posted Apr 3, 2024
N/A / ALOHa
Authors: Suzanne Petryk, David M. Chan, Anish Kachinthaya·Corresponding: not specified·Affiliation: University of California, Berkeley
191Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
Evaluates metric method with Gen tasks using CHAIR-MEN / FaithScore metrics.
object-levelgenerationmetric method
Benchmarks
vision-language
arXiv 2024
First posted Jun 20, 2024
N/A / CHAIR-MEN / FaithScore
Authors: Gregor Geigle, Radu Timofte, Goran Glavaš·Corresponding: not specified·Affiliation: WüNLP; Computer Vision Lab, CAIDAS, University of Würzburg
192Visual Hallucinations of Multi-modal Large Language Models
Evaluates visual hallucination with Dis & Gen tasks using Acc metrics.
object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language
ACL 2024
First posted Feb 22, 2024
1,200 / Acc
Authors: Wen Huang, Hongbin Liu, Minxin Guo·Corresponding: not specified·Affiliation: University of Science & Technology of China; Duke University; The University of Hong Kong
193PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
Evaluates sentiment with Dis tasks using PhD Index metrics.
object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language
CVPR 2025
First posted Mar 17, 2024
102,564 / PhD Index
Authors: Jiazhen Liu, Yuhan Fu, Ruobing Xie·Corresponding: not specified·Affiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
194Automated Multi-level Preference for MLLMs
Evaluates multi-round conv with Gen tasks using Preference Score metrics.
object-levelattribute-levelgenerationmulti-round conv
Benchmarks
vision-language
NeurIPS 2024
First posted May 18, 2024
N/A / Preference Score
Authors: Mengxi Zhang, Wenhao Wu, Yu Lu·Corresponding: not specified·Affiliation: Baidu Inc.; Tianjin University; The University of Sydney; University of Technology Sydney; Tsinghua University; Chinese Academy of Science
195Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
Evaluates multimodal hallucination with Dis tasks using Acc/P/R/F1 metrics.
object-levelrelation-leveldiscrimination
Benchmarks
vision-language
ICML 2024
First posted Jun 24, 2024
8,030 / Acc/P/R/F1
Authors: Mingrui Wu, Jiayi Ji, Oucheng Huang·Corresponding: Jiayi Ji·Affiliation: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China
196Multi-Object Hallucination in Vision-Language Models
Evaluates multi-object with Gen tasks using Accuracy metrics.
object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language
NeurIPS 2024
First posted Jul 8, 2024
5,000 / Accuracy
Authors: Xuweiyi Chen, Ziqiao Ma, Xuejun Zhang·Corresponding: not specified·Affiliation: University of Michigan; University of Virginia; New York University
197BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
Evaluates before-after changes with Dis tasks using TU/IG/SB/ID metrics.
object-leveldiscriminationbefore-after changes
Benchmarks
vision-language
ECCV 2024
First posted Jul 18, 2024
26,064 / TU/IG/SB/ID
Authors: Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi·Corresponding: not specified·Affiliation: Dept. of EE, POSTECH, Korea; Grad. School of AI, POSTECH, Korea; Institute for Convergence Research and Education in Advanced Technology, Yonsei University, Korea
198Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Evaluates cp-bench mitigation with Gen tasks using Acc metrics.
object-levelgenerationcp-bench mitigation
Benchmarks
vision-language
CVPR 2025
First posted Apr 29, 2025
N/A / Acc
Authors: Yuanchen Wu, Lu Zhang, Hang Yao·Corresponding: not specified·Affiliation: School of Computer Engineering & Science, Shanghai University; Tencent YouTu Lab
199EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflow
Evaluates ophthalmology with Dis & Gen tasks using N/A metrics.
discriminationgenerationophthalmology
Benchmarks
vision-language
Information Fusion 2026
First posted Jul 24, 2025
N/A / N/A
Authors: Xiaoyu Pan, Yang Bai, Ke Zou·Corresponding: not specified·Affiliation: Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR), 1 Fusionopolis Way, Singapore, 138632; Centre for Innovation and Precision Eye Health; and Department of Ophthalmology, NUHS Tower Block, Level 7, 1E Kent Ridge Road, Singapore, 119228; Singapore Eye Research Institute, Singapore National Eye Centre, 20 College Road, Singapore, 169856
200Evaluation and Analysis of Hallucination in Large Vision-Language Models
Evaluates multimodal hallucination with Gen tasks using LLM Rating metrics.
object-levelgeneration
Benchmarks
vision-language
arXiv 2023
First posted Aug 29, 2023
25,000 / LLM Rating
Authors: Junyang Wang, Yiyang Zhou, Guohai Xu·Corresponding: not specified·Affiliation: School of Computer and Information Technology, Beijing Jiaotong University, Beijing, China; School of Software Engineering, Xi'an Jiaotong University, Xi'an, China; School of Software, Shandong University, Jinan, China; MAIS, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China; DAMO Academy, Alibaba Group
201Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Evaluates model bias with Gen tasks using Human Assessment metrics.
generationmodel bias
Benchmarks
vision-language
arXiv 2023
First posted Nov 6, 2023
370 / Human Assessment
Authors: Chenhang Cui, Yiyang Zhou, Xinyu Yang·Corresponding: not specified·Affiliation: UNC-Chapel Hill; Carnegie Mellon University; Stanford University; Rutgers University
202MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
Evaluates instruction tuning with Gen tasks using Acc metrics.
object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language
arXiv 2024
First posted Jul 22, 2024
973,000 / Acc
Authors: Yangzhou Liu, Yue Cao, Zhangwei Gao·Corresponding: not specified·Affiliation: Shanghai AI Laboratory, Shanghai 200232, China; SenseTime Research, Shanghai 200233, China; Tsinghua University, Beijing 100084, China; Nanjing University, Nanjing 210023, China; Fudan University, Shanghai 200433, China; The Chinese University of Hong Kong, Hong Kong 999077, China; Shanghai Jiao Tong University, Shanghai 200240, China
203MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
Evaluates decoding method with Gen tasks using Acc metrics.
object-levelgenerationdecoding method
Benchmarks
vision-language
arXiv 2024
First posted Oct 15, 2024
N/A / Acc
Authors: Chenxi Wang, Xiang Chen, Ningyu Zhang·Corresponding: not specified·Affiliation: Zhejiang University; National University of Singapore, NUS-NCS Joint Lab, Singapore
204Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Evaluates unsolvable (upd) with Dis tasks using Acc metrics.
discriminationunsolvable (upd)
Benchmarks
vision-language
arXiv 2024
First posted Mar 29, 2024
2,095 / Acc
Authors: Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang·Corresponding: not specified·Affiliation: The University of Tokyo; S-Lab, Nanyang Technological University; Duke University; University of Wisconsin-Madison; LY Corporation; Tokyo University of Science
205LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
Evaluates societal bias/pref with Dis tasks using Alignment/Bias metrics.
discriminationsocietal biaspref
Benchmarks
vision-language
arXiv 2025
First posted Jul 25, 2025
N/A / Alignment/Bias
Authors: Yusuke Hirota, Boyi Li, Ryo Hachiuma·Corresponding: not specified·Affiliation: NVIDIA Research; Osaka University; Stanford University
206How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
Evaluates deceptive prompts with Gen tasks using Acc metrics.
object-levelattribute-levelgenerationdeceptive prompts
Benchmarks
vision-language
arXiv 2024
First posted Feb 20, 2024
1,000 / Acc
Authors: Yusu Qian, Haotian Zhang, Yinfei Yang·Corresponding: not specified·Affiliation: Apple
207VLind-Bench: Measuring Language Priors in Large Vision-Language Models
Evaluates language priors with Dis tasks using Acc (Y/N) metrics.
object-leveldiscriminationlanguage priors
Benchmarks
vision-language
arXiv 2024
First posted Jun 13, 2024
2,576 / Acc (Y/N)
Authors: Kang-il Lee, Minbeom Kim, Seunghyun Yoon·Corresponding: not specified·Affiliation: ECE, SNU; IPAI, SNU
208MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
Evaluates relation understanding with Dis & Gen tasks using Acc metrics.
object-levelrelation-leveldiscriminationgeneration
Benchmarks
vision-language
arXiv 2024
First posted Jun 13, 2024
22,500 / Acc
Authors: Jiahao Nie, Gongjie Zhang, Wenbin An·Corresponding: not specified·Affiliation: IGP, Nanyang Technological University; Nanyang Technological University; Alibaba DAMO Academy; Xi’an Jiaotong University
209Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
Evaluates alignment method with Dis tasks using Pfram metrics.
object-leveldiscriminationalignment method
Benchmarks
vision-language
arXiv 2024
First posted Sep 2, 2024
N/A / Pfram
Authors: Yueqian Wang, Jianxin Liang, Yuxuan Wang·Corresponding: not specified·Affiliation: Wangxuan Institute of Computer Technology, Peking University; Beijing Institute for General Artificial Intelligence; National Key Laboratory of General Artificial Intelligence
210GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
A training-free detector that combines global scene similarity with local visual grounding for object hallucination detection.
global-local similarityVisual Logit Lensobject grounding
Detection
vision-language
NeurIPS 2025
First posted Aug 27, 2025
global-local grounding
Authors: Seongheon Park, Sharon Li·Corresponding: not specified·Affiliation: University of Wisconsin–Madison
211Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models
A training-free object hallucination detector that uses calibrated local and instruction-context consistency scores.
instruction embeddingsLogit Lensobject hallucination
Detection
vision-language
ICML 2026
First posted May 12, 2026
instruction-token scoring
Authors: Runhe Lai, Xinhua Lu, Yanqi Wu·Corresponding: Weijiang Yu, Ruixuan Wang·Affiliation: Sun Yat-sen University; Peng Cheng Laboratory; Key Laboratory of Machine Intelligence and Advanced Computing, MOE
212Uncertainty Estimation in Autoregressive Structured Prediction
A general ensemble-based framework for token-level and sequence-level uncertainty estimation in autoregressive structured prediction.
ensemble uncertaintytoken-level estimatessequence-level estimates
Quantification
vision-language
ICLR 2021
First posted Feb 18, 2020
autoregressive uncertainty
Authors: Andrey Malinin, Mark Gales·Corresponding: not specified·Affiliation: Yandex; Higher School of Economics; University of Cambridge
213Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
ACT-ViT treats full layer-by-token activation tensors as image-like inputs for cross-LLM hallucination detection.
activation tensorsvision transformercross-model transfer
Detection
vision-language
NeurIPS 2025
First posted Sep 30, 2025
activation-tensor detection
Authors: Guy Bar-Shalom, Fabrizio Frasca, Yaniv Galron·Corresponding: not specified·Affiliation: Technion
214TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
A latent-state detector identifies hallucinated object tokens and guides decoding toward a shared truthful direction.
truthful directionlatent subspacepre-intervention
Detection
vision-language
ICCV 2025
First posted Mar 13, 2025
latent-state detection
Authors: Jinhao Duan, Fei Kong, Hao Cheng·Corresponding: Kaidi Xu·Affiliation: Drexel University; University of Electronic Science and Technology of China; Hong Kong University of Science and Technology (Guangzhou)
215Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
A patch-level detector combines attention dispersion and cross-modal grounding consistency to localize hallucinated tokens.
attention dispersioncross-modal groundingpatch-level detection
Detection
vision-language
CVPR 2026
First posted Apr 6, 2026
token-grounding detection
Authors: Tuan Dung Nguyen, Minh Khoi Ho, Qi Chen·Corresponding: Phi Le Nguyen, Vu Minh Hieu Phan·Affiliation: Hanoi University of Science and Technology; Australian Institute for Machine Learning, University of Adelaide; Mohamed bin Zayed University of Artificial Intelligence