Skip to content
Home›Papers›Overview
📄

Paper Index

Full titles, links, and topic tags.

215Curated entries
4Categories
299Verified links
200Dated papers

A compact reading index for quick scanning.

215 curated entries

Paper Topics

Curated Papers

1

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

ViHallu uses visual variations and visual instruction construction for visual-semantic alignment.

visual variationsinstruction tuningalignment
Mitigation
vision-language

ACM MM 2025

First posted Jul 29, 2025

visual alignment
Authors: Ziyun Dai, Xiaoqiang Li, Shaohua ZhangCorresponding: not specifiedAffiliation: Shanghai University; Shanghai Business School
2

Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs

ToR reweights coupled perception and reasoning tokens during multimodal RLVR.

RLVRtoken reweightinggrounded reasoning
Mitigation
vision-language

arXiv 2026

First posted Mar 26, 2026

RLVR / grounding
Authors: Jinda Lu, Junkang Wu, Jinghan LiCorresponding: Jinda LuAffiliation: University of Science and Technology of China
3

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

PGPO amplifies learning signals for visually dependent tokens.

RLVRpolicy optimizationvisual dependency
Mitigation
vision-language

arXiv 2026

First posted Apr 2, 2026

RLVR / token credit
Authors: Zekai Ye, Qiming Li, Xiaocheng FengCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory
4

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection

VEPO combines visual sensitivity and token entropy for visual reasoning RL.

RLVRtoken selectionvisual reasoning
Mitigation
vision-language

arXiv 2026

First posted Jun 2, 2026

RLVR / visual reasoning
Authors: Senjie Jin, Peixin Wang, Boyang LiuCorresponding: Tao GuiAffiliation: Fudan University
5

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

AOD disentangles hallucination directions for training-free contrastive decoding in LVLMs.

contrastive decodingdisentanglementtraining-free
Mitigation
vision-language

arXiv 2026

First posted May 25, 2026

LVLM / decoding
Authors: Ruoxi Cheng, Haoxuan Ma, Zhengfei HaiCorresponding: Xingjun MaAffiliation: Fudan University / Tencent; Nanjing University; Southeast University
6

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Grounded search agents; public arXiv record is pending.

search agentsintrinsic rewardsgrounding
Mitigation
vision-language

arXiv pending

agents / grounding
arXiv pending
Author and affiliation metadata pending a public paper record.
7

Hallucination-aware intermediate representation edit in large vision-language models

Intermediate representation editing for LVLM hallucination mitigation.

representation editingactivation steeringLVLM
Mitigation
vision-language

arXiv 2026

First posted Mar 31, 2026

VLM / editing
Authors: Wei Suo, Hanzu Zhang, Lijun ZhangCorresponding: Peng WangAffiliation: Northwestern Polytechnical University
8

NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors

Dynamic suppression of language priors during decoding.

language priorsdynamic suppressionobject hallucination
Mitigation
vision-language

arXiv 2026

First posted Feb 25, 2026

decoding / object
Authors: Lingfeng Ren, Weihao Yu, Runpeng YuCorresponding: Weihao Yu, Xinchao WangAffiliation: National University of Singapore; Peking University Shenzhen Graduate School
9

Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models

Hallucination analysis and mitigation for multimodal CoT.

multimodal CoTreasoningmitigation
Mitigation
vision-language

CVPR 2026

First posted Mar 28, 2026

reasoning / CoT
Authors: Ji Ma, Wei Suo, Peng WangCorresponding: Wei SuoAffiliation: Northwestern Polytechnical University
10

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

Region-aware verification loops for object hallucinations.

region-awarechain-of-verificationobject hallucination
Verification
vision-language

arXiv 2026

First posted Apr 22, 2026

verification / region
Authors: Jiahao Xie, Alessio Tonioni, Nathalie RauschmayrCorresponding: not specifiedAffiliation: Max Planck Institute for Informatics / VIA Research Center; Google
11

Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding

Entropy-aware decoding for multimodal reasoning models.

latent entropydecodingmultimodal reasoning
Uncertainty
vision-language

arXiv 2026

First posted Mar 9, 2026

uncertainty / decoding
Authors: Zhongxing Xu, Zhonghua Wang, Zhe QianCorresponding: not specifiedAffiliation: Monash University
12

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Constructs a robust instruction tuning dataset with positive and negative samples to update the model.

instruction tuningdata curation
Mitigation
vision-language

ICLR 2024

First posted Jun 26, 2023

instruction tuning
Authors: Fuxiao Liu, Kevin Lin, Linjie LiCorresponding: not specifiedAffiliation: University of Maryland, College Park; Microsoft Corporation
13

Aligning Large Multimodal Models with Factually Augmented RLHF

Introduces a factually augmented mechanism, using reinforcement learning from human feedback to align the model.

reinforcement learningrlhfmodel alignment
Mitigation
vision-language

ACL 2024

First posted Sep 25, 2023

rlhf
Authors: Zhiqing Sun, Sheng Shen, Shengcao CaoCorresponding: not specifiedAffiliation: UC Berkeley; CMU; UIUC; UW–Madison; UMass Amherst; Microsoft Research; MIT-IBM Watson AI Lab
14

Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Specifically trains the MLLM to generate self-feedback and revise its output based on it.

self-correctionfine-tuningmodel alignment
Mitigation
vision-language

NAACL 2024

First posted Nov 13, 2023

fine-tuning
Authors: Seongyun Lee, Sue Hyun Park, Yongrae JoCorresponding: not specifiedAffiliation: Korea University; KAIST AI; LG AI Research
15

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data

Cleans "hallucinatory toxicity" in existing visual instruction datasets and generates counterfactual data for fine-tuning.

data curationmodel alignment
Mitigation
vision-language

CVPR 2024

First posted Nov 22, 2023

data curation
Authors: Qifan Yu, Juncheng Li, Longhui WeiCorresponding: Longhui WeiAffiliation: Zhejiang University; Huawei Cloud; Institute of Computing Technology, Chinese Academy of Sciences
16

Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Constructs hallucination-aware data pairs and applies the DPO algorithm directly to update model weights.

preference optimizationmodel alignment
Mitigation
vision-language

ICME 2024

First posted Nov 28, 2023

preference optimization
Authors: Zhiyuan Zhao, Bin Wang, Linke OuyangCorresponding: not specifiedAffiliation: Shanghai AI Laboratory
17

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Collects fine-grained, segment-level human correctional feedback for reward modeling and PPO optimization.

reinforcement learningreward modeling
Mitigation
vision-language

CVPR 2024

First posted Dec 1, 2023

rlhf
Authors: Tianyu Yu, Yuan Yao, Haoye ZhangCorresponding: not specifiedAffiliation: Tsinghua University; National University of Singapore
18

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

Trains a fine-grained visual reward model to guide the model toward better visual alignment.

visual groundingreward modeling
Mitigation
vision-language

CVPR 2024

First posted Feb 9, 2024

reward modeling
Authors: Siming Yan, Min Bai, Weifeng ChenCorresponding: not specifiedAffiliation: The University of Texas at Austin; AWS AI
19

Visually Dehallucinative Instruction Generation: Know What You Don't Know

Teaches the model to actively output "I don't know" when evidence is insufficient via specialized instruction tuning.

instruction tuningmodel alignment
Mitigation
vision-language

ACL 2024

First posted Feb 15, 2024

instruction tuning
Authors: Sungguk Cha, Jusung Lee, Younghyun LeeCorresponding: Sungguk Cha, Cheoljong YangAffiliation: Multimodal AI Lab., NC Research, NCSOFT Corporation
20

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Constructs a preference-based dataset to align the model with high-quality image descriptions via preference fine-tuning.

data curationpreference fine-tuningmodel alignment
Mitigation
vision-language

NeurIPS 2024

First posted Feb 18, 2024

preference fine-tuning
Authors: Yiyang Zhou, Chenhang Cui, Rafael RafailovCorresponding: Huaxiu YaoAffiliation: UNC-Chapel Hill; Stanford University
21

Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reconstructs the fine-tuning dataset to teach the model to output the End-Of-Sequence (EOS) token early to truncate generation.

instruction tuningdata curation
Mitigation
vision-language

ACL 2024

First posted Feb 22, 2024

instruction tuning
Authors: Zihao Yue, Liang Zhang, Qin JinCorresponding: not specifiedAffiliation: Renmin University of China
22

Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning

Introduces adversarial instruction tuning to improve model robustness against misleading user queries.

instruction tuningmodel alignment
Mitigation
vision-language

ACL 2024

First posted Mar 15, 2024

instruction tuning
Authors: Dongmin Park, Zhaofang Qian, Guangxing HanCorresponding: not specifiedAffiliation: KRAFTON; UC San Diego; University of Central Florida
23

Calibrated Self-Rewarding Vision Language Models

Proposes a self-calibrating reward mechanism enabling iterative self-generated feedback and preference alignment.

uncertaintyself-rewarding learningmodel alignment
Mitigation
vision-language

NeurIPS 2024

First posted May 23, 2024

self-rewarding learning
Authors: Yiyang Zhou, Zhiyuan Fan, Dongjie ChengCorresponding: not specifiedAffiliation: UNC-Chapel Hill; University of Chicago; University of Maryland; Rutgers University; Independent Researcher
24

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

Constructs a reflective instruction dataset to teach the model to perform internal visual fact-checking before outputting answers.

instruction tuningdata curation
Mitigation
vision-language

ECCV 2024

First posted Jul 16, 2024

instruction tuning
Authors: Jinrui Zhang, Teng Wang, Haigang ZhangCorresponding: Feng ZhengAffiliation: Southern University of Science and Technology; The University of Hong Kong; Shenzhen Polytechnic University; The Cloud Computing and IT Institute of ZTE Corporation; Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory, Shenzhen, China
25

Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models

Cleans and rewrites captions in CLIP pre-training data to reduce object hallucinations at the vision-language source.

object hallucinationdata curation
Mitigation
vision-language

EMNLP 2024

First posted Oct 4, 2024

pre-training intervention
Authors: Yufang Liu, Tao Ji, Changzhi SunCorresponding: not specifiedAffiliation: School of Computer Science and Technology, East China Normal University; School of Computer Science, Fudan University; Pazhou Laboratory, Huangpu
26

V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

Uses a vision-guided mechanism to construct high-quality positive and negative preference pairs for DPO.

preference optimizationmodel alignment
Mitigation
vision-language

EMNLP 2024

First posted Nov 5, 2024

preference optimization
Authors: Yuxi Xie, Guanzhen Li, Xiao XuCorresponding: not specifiedAffiliation: National University of Singapore
27

Hallucination-resistant multimodal content generation through knowledge graph-based reinforcement learning

Grounds multimodal generation in structured knowledge and optimizes graph-based rewards to suppress unsupported content.

reinforcement learningknowledge-graph reinforcement learningmodel alignment
Mitigation
vision-language

Information Fusion 2026

knowledge-graph reinforcement learning
Authors: Liang Zeng, Xinyi Lin, Shanping YuCorresponding: not specifiedAffiliation:
28

Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training

Selects hallucination-oriented samples, assigns severity-specific loss weights, and modulates localized visual attention during HD-DPO.

preference optimizationattention steering
Mitigation
vision-language

AAAI 2026

severity-guided dpo
Authors: Yuanyi Xu, Xiangru Zhu, Sihang JiangCorresponding: not specifiedAffiliation: Fudan University; Renmin University of China; Alibaba Group
29

EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation

Builds temporally grounded video preferences from visual and audio evidence and improves supervision with echo-layered keyframe sampling.

preference optimizationaudiovisual preference optimizationmodel alignment
Mitigation
vision-language

AAAI 2026

audiovisual preference optimization
Authors: Shuai Liu, Da Chen, Yiheng PanCorresponding: not specifiedAffiliation: School of Software Engineering, Xi'an Jiaotong University; ByteDance; School of Cyber Science and Engineering, Xi'an Jiaotong University
30

VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models

Reweights positive words by visual relevance and adapts the negative preference loss using lexical importance differences.

preference optimizationvision-guided lexical dpomodel alignment
Mitigation
vision-language

TOMM 2026

vision-guided lexical dpo
Authors: Siyuan Li, Feng Wang, Simeng QinCorresponding: not specifiedAffiliation: School of Data Science and Intelligent Media, Communication University of China, Beijing, China; Tianjin University, Tianjin, China; Northeastern University, Shenyang, China; Alibaba Group, Hangzhou, China; Communication University of China, Beijing, China
31

Multimodal Chain-of-Thought Reasoning in Language Models

Adopts a two-stage fine-tuning framework, training the model to generate a CoT reasoning process before the answer.

multimodal reasoningfine-tuningmodel alignment
Mitigation
vision-language

arXiv 2023

First posted Feb 2, 2023

fine-tuning
Authors: Zhuosheng Zhang, Aston Zhang, Mu LiCorresponding: not specifiedAffiliation: School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; GenAI, Meta; Amazon Web Services; Department of Computer Science and Engineering, Shanghai Jiao Tong University
32

Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

Introduces a hallucination-augmented contrastive learning loss function during the training phase.

contrastive learningmodel alignment
Mitigation
vision-language

arXiv 2023

First posted Dec 12, 2023

contrastive learning
Authors: Chaoya Jiang, Haiyang Xu, Mengfan DongCorresponding: not specifiedAffiliation: National Engineering Research Center for Software Engineering, Peking University; Alibaba Group
33

KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

Jointly trains a language model, a vision encoder, and a Graph Neural Network (GNN) to integrate knowledge graphs for reasoning.

external toolsmultimodal reasoning
Mitigation
vision-language

arXiv 2024

First posted Jan 23, 2024

joint training
Authors: Debjyoti Mondal, Suraj Modi, Subhadarshi PandaCorresponding: not specifiedAffiliation: Samsung R&D Institute India - Bangalore
34

Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training

Proposes a novel vision-language pre-training (VLP) objective to suppress hallucinations.

pre-trainingmodel alignment
Mitigation
vision-language

arXiv 2023

First posted Oct 14, 2022

pre-training
Authors: Wenliang Dai, Zihan Liu, Ziwei JiCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology; NVIDIA
35

Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites

Uses ChatGPT to rewrite image captions for data construction and performs two-stage fine-tuning.

fine-tuningmodel alignment
Mitigation
vision-language

arXiv 2023

First posted Dec 4, 2023

fine-tuning
Authors: Lei Wang, Jiabang He, Shenshen LiCorresponding: not specifiedAffiliation: Singapore Management University, Beijing Forestry University; University of Electronic Science and Technology of China
36

Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

Proposes hallucination-aware DPO, utilizing fine-grained AI feedback to construct positive and negative pairs for weight updates.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Apr 22, 2024

preference optimization
Authors: Wenyi Xiao, Ziwei Huang, Leilei GanCorresponding: not specifiedAffiliation: Zhejiang University; Alibaba Group; Fudan University
37

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Utilizes fine-grained feedback from open-source AIs to build a preference dataset and updates weights using alignment algorithms.

reinforcement learningdata curation
Mitigation
vision-language

arXiv 2024

First posted May 27, 2024

rlaif
Authors: Tianyu Yu, Haoye Zhang, Qiming LiCorresponding: not specifiedAffiliation: Department of Computer Science and Technology, Tsinghua University; NExT++ Lab, School of Computing, National University of Singapore
38

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Proposes a conditional preference optimization algorithm to reduce hallucinations without degrading general capabilities.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Jun 17, 2024

preference optimization
Authors: Fei Wang, Wenxuan Zhou, James Y. HuangCorresponding: not specifiedAffiliation: University of Southern California; University of California, Davis; Microsoft Research
39

CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs

Uses a pre-trained CLIP model directly as an AI judge to generate preference pairs for DPO training.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Aug 19, 2024

preference optimization
Authors: Yassine Ouali, Adrian Bulat, Brais MartinezCorresponding: not specifiedAffiliation: Samsung AI Center Cambridge, UK; Technical University of Ia \textcommabelow si, Romania; Queen Mary University of London, UK
40

EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models

Proposes an efficient fine-grained unlearning framework to "forget" the tendency to hallucinate via reverse parameter updates.

unlearningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Feb 15, 2024

unlearning
Authors: Shangyu Xing, Fei Zhao, Zhen WuCorresponding: not specifiedAffiliation: National Key Laboratory for Novel Software Technology, Nanjing University, China
41

Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization

Introduces a hallucination-inducing mechanism during training to construct sample pairs for parameter optimization.

contrastive tuningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted May 24, 2024

contrastive tuning
Authors: Xinyu Lyu, Beitao Chen, Lianli GaoCorresponding: not specifiedAffiliation: Center for Future Media, University of Electronic Science and Technology of China; Shenzhen Institute for Advanced Study, UESTC
42

Mitigating Open-Vocabulary Caption Hallucinations

Trains a reward model based on Natural Language Inference (NLI) and uses RL to fine-tune the MLLM.

reinforcement learningreward modeling
Mitigation
vision-language

arXiv 2023

First posted Dec 6, 2023

rlhf
Authors: Assaf Ben-Kish, Moran Yanuka, Morris AlperCorresponding: not specifiedAffiliation: Tel-Aviv University
43

Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning

Constructs targeted repair instruction datasets for different hallucination types to perform targeted fine-tuning.

data curationfine-tuningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Apr 16, 2024

fine-tuning
Authors: Rui Hu, Yahan Tu, Shuyu WeiCorresponding: not specifiedAffiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
44

See or Guess: Counterfactually Regularized Image Captioning

Introduces a counterfactual regularization term during training to force distinction between seen visuals and guessed priors.

regularization trainingmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Aug 29, 2024

regularization training
Authors: Qian Cao, Xu Chen, Ruihua SongCorresponding: not specifiedAffiliation: Gaoling School of Artificial Intelligence, Renmin University of China; Tencent AI Lab
45

Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization

Generates targeted preference data specifically for hallucinations and applies DPO to penalize hallucination-prone features.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Nov 15, 2024

preference optimization
Authors: Yuhan Fu, Ruobing Xie, Xingwu SunCorresponding: not specifiedAffiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
46

Silkie: Preference Distillation for Large Visual Language Models

Uses AI to construct the VLFeedback dataset and distills preferences into the model via DPO.

preference optimizationdata curation
Mitigation
vision-language

arXiv 2023

First posted Dec 17, 2023

preference distillation
Authors: Lei Li, Zhihui Xie, Mukai LiCorresponding: not specifiedAffiliation: The University of Hong Kong; The Chinese University of Hong Kong (Shenzhen); Peking University
47

Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Constructs hard negative samples via data augmentation and introduces a contrastive loss to strengthen visual grounding.

visual groundingcontrastive tuningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted May 28, 2024

contrastive tuning
Authors: Pritam Sarkar, Sayna Ebrahimi, Ali EtemadCorresponding: not specifiedAffiliation: Queen's University and Vector Institute; Google DeepMind; Queen's University; Google Cloud AI Research
48

Mitigating Hallucination in Visual Language Models with Visual Supervision

Constructs a fine-grained RAI-30k dataset and integrates the SAM model during instruction tuning.

instruction tuningdata curation
Mitigation
vision-language

arXiv 2023

First posted Nov 27, 2023

instruction tuning
Authors: Zhiyang Chen, Yousong Zhu, Yufei ZhanCorresponding: not specifiedAffiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Peng Cheng Laboratory; Wuhan AI Research
49

Mitigating Multilingual Hallucination in Large Vision-Language Models

Builds a correctional dataset specifically for multilingual scenarios to mitigate cross-lingual hallucinations caused by language bias.

instruction tuningdata curation
Mitigation
vision-language

arXiv 2024

First posted Aug 1, 2024

instruction tuning
Authors: Xiaoye Qu, Mingyang Song, Wei WeiCorresponding: Wei WeiAffiliation: School of Computer Science & Technology, Huazhong University of Science and Technology; School of Computing Science and Technology, Fudan University; Wangxuan Institute of Computer Technology, Peking University; School of Computing Science and Technology, Zhejiang Gongshang University; Department of Computer Science and Engineering, The Chinese University of Hong Kong
50

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

Proposes a modality-fair preference optimization algorithm to balance penalties and prevent language prior dominance.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Oct 20, 2024

preference optimization
Authors: Songtao Jiang, Yan Zhang, Ruizhe ChenCorresponding: not specifiedAffiliation: Zhejiang University; National University of Singapore
51

Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation

Proposes a token-level preference optimization strategy based on self-calibrated visual-anchored rewards.

preference optimizationuncertainty
Mitigation
vision-language

arXiv 2024

First posted Dec 19, 2024

preference optimization
Authors: Jihao Gu, Yingyao Wang, Meng CaoCorresponding: not specifiedAffiliation: Alibaba Group; Mohamed bin Zayed University of Artificial Intelligence
52

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Introduces Retrieval-Augmented Generation (RAG) into preference data construction to guide DPO training.

preference optimizationretrieval
Mitigation
vision-language

arXiv 2025

First posted Feb 18, 2025

preference optimization
Authors: Shuo Xing, Peiran Li, Yuping WangCorresponding: not specifiedAffiliation: Texas A&M University; University of Michigan; UIUC; UNC Chapel Hill
53

Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images

Proposes symmetrical visual contrastive optimization using minimal contrastive images for alignment fine-tuning.

contrastive optimizationmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Feb 19, 2025

contrastive optimization
Authors: Shengguang Wu, Fan-Yun Sun, Kaiyue WenCorresponding: not specifiedAffiliation: Stanford University
54

FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback

Uses fine-grained AI feedback to construct preference data and applies DPO for alignment training.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Apr 7, 2024

preference optimization
Authors: Liqiang Jing, Xinya DuCorresponding: not specifiedAffiliation: The University of Texas at Dallas
55

TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Massively scales up text-centric visual instruction tuning data to specifically address text hallucinations in OCR tasks.

instruction tuningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Apr 19, 2024

instruction tuning
Authors: Jingqun Tang, Chunhui Lin, Zhen ZhaoCorresponding: not specifiedAffiliation: ByteDance; East China Normal University; Huazhong University of Science and Technology
56

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Systematically builds datasets covering various alignment strategies to empirically verify their impact on hallucinations.

verificationdata curation
Mitigation
vision-language

arXiv 2024

First posted Jul 2, 2024

fine-tuning
Authors: Elmira Amirloo, Jean-Philippe Fauconnier, Christoph RoesmannCorresponding: not specifiedAffiliation: Apple Inc
57

RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data

Optimizes a visual reward model using auxiliary text-only preference data during RLHF to enhance hallucination discrimination.

reinforcement learningreward modeling
Mitigation
vision-language

arXiv 2024

First posted Aug 22, 2024

reward modeling
Authors: Chenglong Wang, Yang Gan, Yifu HuoCorresponding: Chunliang ZhangAffiliation: School of Computer Science and Engineering, Northeastern University, Shenyang, China; NiuTrans Research, Shenyang, China; CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China
58

HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding

Proposes a fine-tuning framework combining vision-enhanced penalty decoding with hierarchical feedback learning for behavior alignment.

feedback learningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Sep 30, 2024

feedback learning
Authors: Fan Yuan, Chi Qin, Xiaogang XuCorresponding: not specifiedAffiliation: College of Artificial Intelligence; Nanjing University of Aeronautics and Astronautics, Nanjing, China; MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China; The Chinese University of Hong Kong, Hong Kong, China
59

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Proposes entity-centric multimodal preference optimization, focusing penalty granularity precisely on specific visual entity tokens.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jun 4, 2025

preference optimization
Authors: Jiulong Wu, Zhengliang Shi, Shuaiqiang WangCorresponding: not specifiedAffiliation: Soochow University, Suzhou, China; Baidu Inc., Beijing, China; Shandong University, Qingdao, China
60

Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales

Introduces faithful and concise reasoning rationales as supervision signals during training to enhance logical generation.

multimodal reasoningrationale trainingmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Apr 17, 2024

rationale training
Authors: Minghe Gao, Shuang Chen, Liang PangCorresponding: not specifiedAffiliation: Zhejiang University; Chinese Academy of Sciences; National University of Singapore; Sun Yat-sen University
61

Generating Faithful and Salient Text from Multimodal Data

Constructs a multimodal factual consistency dataset and fine-tunes the model to improve faithfulness in data-to-text generation.

data curationfine-tuningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Sep 6, 2024

fine-tuning
Authors: Tahsina Hashem, Weiqing Wang, Derry Tanti WijayaCorresponding: not specifiedAffiliation: Department of Data Science & AI, Monash University, Australia; Department of Data Science, Monash University, Indonesia; Department of CSE, Bangladesh University of Engineering and Technology, Bangladesh
62

Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment

Formulates preference modeling as a next-token prediction task to update the model using fine-grained verifier data.

verificationpreference modelingmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Oct 18, 2024

preference modeling
Authors: Chenhang Cui, An Zhang, Yiyang ZhouCorresponding: An ZhangAffiliation: National University of Singapore; UNC-Chapel Hill; Chicago University; Nanyang Technological University
63

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

Constructs a dataset targeting verb concept hallucinations and applies specific instruction tuning to correct action biases.

instruction tuningdata curation
Mitigation
vision-language

arXiv 2024

First posted Dec 6, 2024

instruction tuning
Authors: Zehao Wang, Xinpeng Liu, Yudonglin ZhangCorresponding: not specifiedAffiliation: Shanghai Jiao Tong University; ARC Lab, Tencent PCG
64

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution

Automatically generates and filters high-quality alignment data through an iterative self-evolution mechanism.

data curationself-evolution learningmodel alignment
Mitigation
vision-language

arXiv 2024

First posted Dec 20, 2024

self-evolution learning
Authors: Wentao Tan, Qiong Cao, Yibing ZhanCorresponding: not specifiedAffiliation: South China University of Technology; JD Explore Academy, Beijing; Pazhou Lab, Guangzhou
65

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs

Proposes cross-modal hierarchical DPO to penalize global image-text matching and local object features at different levels.

preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jan 28, 2025

preference optimization
Authors: Jinlan Fu, Shenzhen Huangfu, Hao FeiCorresponding: not specifiedAffiliation: National University of Singapore; Fudan University; Digital Twin Institute, Eastern Institute of Technology, Ningbo
66

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

Introduces subtle visual perturbations during training forward passes to force the learning of robust visual representations.

representation editingvisual perturbation trainingmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Mar 9, 2025

visual perturbation training
Authors: Cong Chen, Mingyu Liu, Chenchen JingCorresponding: not specifiedAffiliation: Zhejiang University; WeChat Group; Zhejiang University of Technology
67

Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy

Builds a low-level visual hallucination database and applies negative sampling to enhance awareness of low-level features.

negative sample fine-tuningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Mar 26, 2025

negative sample fine-tuning
Authors: Yinan Sun, Xiongkuo Min, Zicheng ZhangCorresponding: Xiongkuo MinAffiliation:
68

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning

Adopts rationale-augmented instruction tuning to teach the model to generate critique logic before providing an answer.

instruction tuningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted May 12, 2025

instruction tuning
Authors: Zexian Yang, Dian Li, Dayan WuCorresponding: Dayan Wu, Gang LiuAffiliation: Institute of Information Engineering, Chinese Academy of Sciences; Foundation Technology Center, Tencent PCG
69

OViP: Online Vision-Language Preference Learning for VLM Hallucination

Proposes an online preference learning framework to dynamically generate preference pairs during training.

preference learningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted May 21, 2025

preference learning
Authors: Shujun Liu, Siyuan Wang, Zejun LiCorresponding: not specifiedAffiliation: Fudan University; University of Southern California; ByteDance
70

BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models

Employs a bijective maximum likelihood learning approach to suppress hallucinations by optimizing joint probability distributions.

maximum likelihood learningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted May 30, 2025

maximum likelihood learning
Authors: Huu-Thien Tran, Thanh-Dat Truong, Khoa LuuCorresponding: not specifiedAffiliation: CVIU Lab, University of Arkansas
71

Stop learning it all to mitigate visual hallucination, Focus on the hallucination target

Stops learning redundant backgrounds and forces weight updates to focus exclusively on target areas causing hallucinations.

preference optimizationtarget-localized dpomodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jun 13, 2025

target-localized dpo
Authors: Dokyoon Yoon, Youngsook Song, Woomyong ParkCorresponding: not specifiedAffiliation: SIONIC AI
72

Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation

Identifies and purges hallucinated samples from fine-tuning data, using the refined high-quality data for knowledge distillation.

knowledge distillationmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jul 7, 2025

knowledge distillation
Authors: Wenhao Li, Xiu Su, Jingyi WuCorresponding: not specifiedAffiliation: University of Sydney; Central South University; Fudan University; Southeast University; HKUST; Sensetime Research
73

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Proposes a closed-loop framework where the model generates predictions, performs backward verification, and uses errors as training signals.

verificationclosed-loop trainingmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jul 7, 2025

closed-loop training
Authors: Jianjiang Yang, Yanshu li, Ziyan HuangCorresponding: not specifiedAffiliation: Department of Computer Science, University of Bristol; School of Future Technology, South China University of Technology; Brown University
74

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

Analyzes object biases introduced during pre-training and debiases the model using a reverse objective function during fine-tuning.

unlearningbias unlearningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Aug 6, 2025

bias unlearning
Authors: Yifan Li, Kun Zhou, Wayne Xin ZhaoCorresponding: not specifiedAffiliation: Gaoling School of Artificial Intelligence, Renmin University of China; University of California, San Diego; DataCanvas Alaya NeW
75

TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs

Adopts an adaptive MinMax Token preference strategy to dynamically adjust reward allocation.

adaptive preference strategymodel alignment
Mitigation
vision-language

arXiv 2025

First posted Jul 29, 2025

adaptive preference strategy
Authors: Kejia Zhang, Keda Tao, Zhiming LuoCorresponding: not specifiedAffiliation: Xiamen University; Westlake University; DAMO Academy, Alibaba Group; AWS AI Lab, Amazon; Hupan Laboratory
76

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

Integrates the CHAIR metric, utilizing object-level matching scores as the preference reward for DPO.

preference optimizationobject-aware dpomodel alignment
Mitigation
vision-language

arXiv 2025

First posted Aug 27, 2025

object-aware dpo
Authors: Alberto Compagnoni, Davide Caffagni, Nicholas MoratelliCorresponding: not specifiedAffiliation: University of Modena and Reggio Emilia, Modena, Italy
77

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Distinguishes "omission" and "fabrication" hallucination causes and constructs respective fine-tuning data for decoupled training.

decoupled fine-tuningmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Aug 30, 2025

decoupled fine-tuning
Authors: Guangzong Si, Hao Yin, Xianfei LiCorresponding: not specifiedAffiliation: USTC; CowaRobot
78

Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

Proposes semantic curriculum preference optimization, training the model in increasing difficulty to smoothly eliminate hallucinations.

preference optimizationcurriculum preference optimizationmodel alignment
Mitigation
vision-language

arXiv 2025

First posted Sep 29, 2025

curriculum preference optimization
Authors: Yuanshuai Li, Yuping Yan, Junfeng TangCorresponding: not specifiedAffiliation: School of Engineering, Westlake University, Hangzhou, China; School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China
79

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

Proposes a universal training framework across multiple alignment formats (e.g., DPO, PPO) to eliminate vision-language biases.

preference optimizationreinforcement learning
Mitigation
vision-language

arXiv 2025

First posted Nov 21, 2025

unified alignment framework
Authors: Jiaye Qian, Ge Zheng, Yuchen ZhuCorresponding: Sibei YangAffiliation: School of Computer Science and Engineering, Sun Yat-sen University; ShanghaiTech University
80

OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Extracts self-attention matrices to find and penalize attention flows overly reliant on summary tokens.

attention steeringattention penaltyinternal intervention
Mitigation
vision-language

CVPR 2024

First posted Nov 29, 2023

attention penalty
Authors: Qidong Huang, Xiaoyi Dong, Pan ZhangCorresponding: not specifiedAffiliation: University of Science and Technology of China; Shanghai AI Laboratory
81

Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Fuses global and local attention feature matrices internally to enhance fine-grained perception.

attention steeringattention fusioninternal intervention
Mitigation
vision-language

CVPR 2025

First posted Jun 18, 2024

attention fusion
Authors: Wenbin An, Feng Tian, Sicong LengCorresponding: not specifiedAffiliation: Xi'an Jiaotong University; Nanyang Technological University; Lenovo Research; SGIT AI Lab; University of Massachusetts Boston
82

Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs

Adjusts cross-modal attention by forcefully amplifying the weights of visual tokens.

attention steeringattention amplificationinternal intervention
Mitigation
vision-language

ECCV 2024

First posted Jul 31, 2024

attention amplification
Authors: Shi Liu, Kecheng Zheng, Wei ChenCorresponding: not specifiedAffiliation: State Key Lab of CAD&CG, Zhejiang University
83

DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination

Dives into the MLLM to mask specific attention connections overly dominated by text.

attention steeringattention maskinginternal intervention
Mitigation
vision-language

EMNLP 2024

First posted Oct 6, 2024

attention masking
Authors: Xuan Gong, Tianshi Ming, Xinpeng WangCorresponding: not specifiedAffiliation: Department of Computer Science and Technology, Tongji University
84

Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

Uses causal analysis to block specific causal attention paths causing hallucinations during inference.

attention steeringcausal attention disconnectioninternal intervention
Mitigation
vision-language

ICLR 2025

First posted Oct 7, 2024

causal attention disconnection
Authors: Guanyu Zhou, Yibo Yan, Xin ZouCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Tsinghua University
85

Mitigating Object Hallucination via Concentric Causal Attention

Strengthens the causal association between local visual features and global semantics in attention matrices.

attention steeringconcentric causal attentioninternal intervention
Mitigation
vision-language

NIPS 2024

First posted Oct 21, 2024

concentric causal attention
Authors: Yun Xing, Yiheng Li, Ivan LaptevCorresponding: not specifiedAffiliation: Nanyang Technological University; MBZUAI
86

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Projects hidden states into a hallucination space and strips them via orthogonalization to clear hallucinations.

representation editinglatent space orthogonalizationinternal intervention
Mitigation
vision-language

CVPR 2025

First posted Dec 18, 2024

latent space orthogonalization
Authors: Le Yang, Ziwei Zheng, Boxu ChenCorresponding: not specifiedAffiliation: Xi'an Jiaotong University, Xi'an 710049, China
87

VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

Introduces visual-aware sparsification to directly drop irrelevant visual tokens prone to inducing hallucinations.

sparsification droppinginternal intervention
Mitigation
vision-language

CVPR 2025

First posted Jan 11, 2025

sparsification dropping
Authors: Xianwei Zhuang, Zhihong Zhu, Yuxin XieCorresponding: not specifiedAffiliation: SECE of Peking University
88

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

Adaptively constrains inter-layer information flow to block false causal paths causing object hallucinations.

object hallucinationinformation flow constraintinternal intervention
Mitigation
vision-language

AAAI 2025

First posted Feb 28, 2025

information flow constraint
Authors: Jiaqi Bai, Hongcheng Guo, Zhongyuan PengCorresponding: Mohan LiAffiliation: Cyberspace Institute of Advanced Technology, Guangzhou University, China; Huangpu Research School of Guangzhou University, China; CCSE, Beihang University, China; University of the Chinese Academy of Sciences, China
89

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

Amplifies low-level visual signals in multimodal layers to combat dilution by deep language features.

signal enhancementinternal intervention
Mitigation
vision-language

CVPR 2025

First posted Mar 17, 2025

signal enhancement
Authors: Hao Yin, Guangzong Si, Zilei WangCorresponding: not specifiedAffiliation: University of Science and Technology of China
90

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Reconstructs multimodal attention using a causal matrix based on Manhattan distance.

attention steeringmanhattan causal replacementinternal intervention
Mitigation
vision-language

ACMMM 2025

First posted Jul 12, 2025

manhattan causal replacement
Authors: Qiyan Zhao, Xiaofeng Zhang, Yiheng LiCorresponding: not specifiedAffiliation: FKLPRIU, Xiamen University of Technology, China; Shanghai Jiao Tong University, China; Nanyang Technological University, Singapore; Jilin University, China; Monash University, Australia; Zhejiang University, China; Huizhou University, China; Chinese Academy of Sciences, China
91

Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models

Intervenes from the visual processing architecture to fundamentally alter how the model acquires visual information.

pure vision reinforcementinternal intervention
Mitigation
vision-language

EMNLP 2025

First posted Sep 17, 2025

pure vision reinforcement
Authors: Weihang Wang, Xinhao Li, Ziyue WangCorresponding: not specifiedAffiliation: Bilibili; UESTC; University of Virginia
92

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Maps hallucination activations toward trusted attention regions with mixture-Gaussian bridges.

attention steeringrepresentation editing
Mitigation
vision-language

WACV 2026

First posted Feb 10, 2026

attention manifold alignment
Authors: Ziqiang Shi, Rujie Liu, Shanshan YuCorresponding: not specifiedAffiliation: Fujitsu Research & Development Center Co.,LTD., Beijing, China
93

RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Predicts input-specific latent steering vectors with gradient correction in a single forward pass.

attention steeringrectified-flow interventioninternal intervention
Mitigation
vision-language

AAAI 2026

rectified-flow intervention
Authors: Junyu Cheng, Zhibiao Liang, Yidong ChenCorresponding: not specifiedAffiliation: Department of Artificial Intelligence, School of Informatics, Xiamen University, China; School of Computer Science, South China Normal University, China; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China
94

Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models

Filters intermediate attention maps to isolate dominant phantom text tokens and strengthen visual anchor tokens.

attention steeringdata curation
Mitigation
vision-language

AAAI 2026

token-asymmetric filtering
Authors: Shuyi Ouyang, Hongyi Wang, Gongfan FangCorresponding: not specifiedAffiliation: Zhejiang University; National University of Singapore
95

Look Before You Speak: Visual Re-Focusing for Hallucination Mitigation in Large Vision-Language Models

Reorients the model toward image evidence before response generation to reduce visually unsupported claims.

visual re-focusinginternal intervention
Mitigation
vision-language

LNCS 2026

visual re-focusing
Authors: Zheyuan Zhang, Yingmin Liu, Yang ShuCorresponding: Yang ShuAffiliation: Northwestern Polytechnical University, Xi'an, China; Zhejiang University, Hangzhou, China
96

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

Uses an attention lens to locate hallucinated features in middle layers and masks them for correction.

attention steeringmiddle layer maskinginternal intervention
Mitigation
vision-language

arXiv 2024

First posted Nov 23, 2024

middle layer masking
Authors: Zhangqi Jiang, Junkai Chen, Beier ZhuCorresponding: not specifiedAffiliation: National University of Defense Technology; Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and; Its Interdisciplinary Applications (Southeast University), Ministry of Education, China; Nanyang Technological University
97

Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence

Quantifies perception divergence and dynamically masks specific heads that lose visual alignment.

visual groundingattention steering
Mitigation
vision-language

arXiv 2024

First posted Dec 18, 2024

attention head pruning
Authors: Jinghan He, Kuan Zhu, Haiyun GuoCorresponding: not specifiedAffiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; University of Science and Technology of China; Southeast University; National University of Singapore
98

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Reallocates Attention Maps during inference to pull misplaced attention back to visual targets.

attention steeringattention reallocationinternal intervention
Mitigation
vision-language

arXiv 2025

First posted Mar 11, 2025

attention reallocation
Authors: Chongjun Tu, Peng Ye, Dongzhan ZhouCorresponding: Tao ChenAffiliation: Fudan University; The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory; StepFun
99

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Proposes cross-level trusted intervention to block the drift of object representations toward false language priors.

representation editingcross-level blockinginternal intervention
Mitigation
vision-language

arXiv 2024

First posted Nov 22, 2024

cross-level blocking
Authors: Junzhe Chen, Tianshu Zhang, Shiyu HuangCorresponding: not specifiedAffiliation: Tsinghua University; The Hong Kong University of Science and Technology (Guangzhou); Zhipu AI; Chongqing University; Shanghai Jiao Tong University.
100

Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis

Implements causal intervention to cut off spurious feature maps hidden in attention matrices.

attention steeringcausal attention interventioninternal intervention
Mitigation
vision-language

arXiv 2024

First posted Dec 4, 2024

causal attention intervention
Authors: Po-Hsuan Huang, Jeng-Lin Li, Chin-Po ChenCorresponding: not specifiedAffiliation: Inventec Corporation, No. 66, Hougang St., Shihlin Dist., Taipei City; University at Albany - SUNY, Albany, New York
101

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Locates the evolution of object features and corrects visual tokens via early layer interventions.

early layer interventioninternal intervention
Mitigation
vision-language

arXiv 2024

First posted Oct 9, 2024

early layer intervention
Authors: Yuying Shang, Xinyi Zeng, Yutao ZhuCorresponding: not specifiedAffiliation: Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing, China; Gaoling School of Artificial Intelligence, Renmin University of China; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University, Beijing, China; Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University; Kuaishou Technology Inc., Beijing, China
102

Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Accurately locates key attention heads and specifically enhances or suppresses them during forward propagation.

attention steeringhead interventioninternal intervention
Mitigation
vision-language

arXiv 2024

First posted Nov 15, 2024

head intervention
Authors: Xiaofeng Zhang, Yihao Quan, Chaochen GuCorresponding: not specifiedAffiliation: Shanghai Jiao Tong University; Alibaba Group; Beijing Jiaotong University
103

Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Calibrates the weight allocation of attention heads to weaken the implicit coverage of visual inputs by language priors.

uncertaintyattention steering
Mitigation
vision-language

arXiv 2024

First posted May 28, 2024

attention calibration
Authors: Sangmin Woo, Donguk Kim, Jaehyuk JangCorresponding: not specifiedAffiliation: KAIST
104

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

Introduces dual-attention mechanisms to internally rebalance language priors and image feature flows.

attention steeringdual-attention balancinginternal intervention
Mitigation
vision-language

arXiv 2024

First posted Nov 21, 2024

dual-attention balancing
Authors: Haozhe Zhao, Shuzheng Si, Liang ChenCorresponding: Baobao ChangAffiliation: University of Illinois Urbana-Champaign; Peking University; Tsinghua University
105

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

Cuts off implicit semantic biases introduced during cross-lingual transitions inside the model.

attention steeringcross-lingual attention interventioninternal intervention
Mitigation
vision-language

arXiv 2025

First posted Jun 3, 2025

cross-lingual attention intervention
Authors: Zekai Ye, Qiming Li, Xiaocheng FengCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory; Central South University; Huawei Technologies Co., Ltd
106

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

Pinpoints tokens sourcing hallucinations and excises them during forward propagation.

token localization & pruninginternal intervention
Mitigation
vision-language

arXiv 2025

First posted Jun 12, 2025

token localization & pruning
Authors: Eunkyu Park, Minyeong Kim, Gunhee KimCorresponding: not specifiedAffiliation: Seoul National University
107

Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Identifies and edits vision-language feature vectors representing hallucinations in intermediate layers.

representation editinginternal intervention
Mitigation
vision-language

arXiv 2024

First posted Oct 3, 2024

representation editing
Authors: Nick Jiang, Anish Kachinthaya, Suzie PetrykCorresponding: not specifiedAffiliation: University of California, Berkeley
108

Reducing Hallucinations in Vision-Language Models via Latent Space Steering

Injects pre-trained steering vectors into the latent space to guide generation away from hallucinations.

attention steeringlatent space steeringinternal intervention
Mitigation
vision-language

arXiv 2024

First posted Oct 21, 2024

latent space steering
Authors: Sheng Liu, Haotian Ye, Lei XingCorresponding: not specifiedAffiliation: Stanford University
109

Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Penalizes decoding by contrasting the logit output distributions of original and distorted images during inference.

contrastive decodinggeneration control
Mitigation
vision-language

CVPR 2024

First posted Nov 28, 2023

contrastive decoding
Authors: Sicong Leng, Hang Zhang, Guanzheng ChenCorresponding: not specifiedAffiliation: DAMO Academy, Alibaba Group; Nanyang Technological University; Hupan Lab, 310023, Hangzhou, China
110

Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Introduces Classifier-Free Guidance (CFG) during inference to amplify the influence of visual features in the generation process.

inference guidancegeneration control
Mitigation
vision-language

CVPR 2024

First posted Feb 13, 2024

inference guidance
Authors: Linxi Zhao, Yihe Deng, Weitong ZhangCorresponding: Quanquan GuAffiliation: Department of Computer Science, Cornell University, Ithaca, NY, USA; Department of Computer Science, University of California, Los Angeles, CA, USA; School of Data Science and Society, UNC, Chapel Hill, NC, USA
111

IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Penalizes pure language priors by contrasting the logits of inputs with and without images.

image-biased decodinggeneration control
Mitigation
vision-language

ECCV 2024

First posted Feb 28, 2024

image-biased decoding
Authors: Lanyun Zhu, Deyi Ji, Tianrun ChenCorresponding: not specifiedAffiliation: Singapore University of Technology and Design; Alibaba Group; Zhejiang University
112

HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Proposes adaptive focal-contrast decoding to fine-grainedly suppress hallucinations at the local token level.

contrastive decodingfocal contrastive decodinggeneration control
Mitigation
vision-language

ICML 2024

First posted Mar 1, 2024

focal contrastive decoding
Authors: Zhaorun Chen, Zhuokai Zhao, Hongyin LuoCorresponding: Zhaorun Chen, Zhuokai Zhao, Bo Li, Jiawei ZhouAffiliation: University of Chicago, Chicago IL, USA; University of Illinois at Urbana-Champaign, Champaign IL, USA; Massachusetts Institute of Technology, Boston MA, USA; UNC-Chapel Hill, Chapel Hill NC, USA; Toyota Technological Institute at Chicago, Chicago IL, USA
113

Multi-Modal Hallucination Control by Visual Information Grounding

Introduces a visual grounding mechanism during decoding to correct biases by contrasting different information streams.

visual groundingvisual information contrastgeneration control
Mitigation
vision-language

CVPR 2024

First posted Mar 20, 2024

visual information contrast
Authors: Alessandro Favero, Luca Zancato, Matthew TragerCorresponding: not specifiedAffiliation: AWS AI Labs
114

Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Constructs positive and negative instructions, contrasting their logits during inference to guide decoding.

contrastive decodinginstruction contrastive decodinggeneration control
Mitigation
vision-language

ACL 2024

First posted Mar 27, 2024

instruction contrastive decoding
Authors: Xintong Wang, Jingheng Pan, Liang DingCorresponding: not specifiedAffiliation: Department of Informatics, Universität Hamburg; The University of Sydney
115

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

Introduces residual visual connections to add early visual features directly into subsequent decoding computations.

residual decodinggeneration control
Mitigation
vision-language

ACL 2024

First posted Jun 30, 2024

residual decoding
Authors: Weihong Zhong, Xiaocheng Feng, Liang ZhaoCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory
116

ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models

Uses attention visualization to locate hallucination regions and constructs interference images for contrastive decoding.

contrastive decodingattention steering
Mitigation
vision-language

AAAI 2025

First posted Aug 25, 2024

visualized contrastive decoding
Authors: Yeji Park, Deokyeong Lee, Junsuk ChoeCorresponding: Junsuk Choe, Buru ChangAffiliation: Sogang University
117

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Introduces a generative feedback loop to detect and correct the probability distribution of the current token in real time.

self-correctionself-correcting decodinggeneration control
Mitigation
vision-language

ICLR 2025

First posted Feb 10, 2025

self-correcting decoding
Authors: Ce Zhang, Zifu Wan, Zhehan KanCorresponding: not specifiedAffiliation: School of Computer Science, Carnegie Mellon University
118

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Adaptively adjusts the contrastive penalty based on image complexity and local token features.

contrastive decodingdynamic contrastive decodinggeneration control
Mitigation
vision-language

CVPR 2025

First posted Mar 1, 2025

dynamic contrastive decoding
Authors: Wei Suo, Lijun Zhang, Mengyang SunCorresponding: Yanning ZhangAffiliation: School of Computer Science and Ningbo Institute, Northwestern Polytechnical University,China.; School of Cybersecurity, Northwestern Polytechnical University, China.; Computer Science, Swansea University.
119

Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models

Adaptively switches between different sampling algorithms based on attention strength.

attention steeringmixture of decodinggeneration control
Mitigation
vision-language

ACL 2025

First posted May 17, 2025

mixture of decoding
Authors: Xinlong Chen, Yuanxing Zhang, Qiang LiuCorresponding: not specifiedAffiliation: New Laboratory of Pattern Recognition (NLPR); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences; Kuaishou Technology; Nanjing University
120

Med-VCD: Mitigating hallucination for medical large vision language models through visual contrastive decoding

Selects visually informed tokens on the fly, removing redundant tokens while retaining critical medical image context.

contrastive decodingsparse visual contrastive decodinggeneration control
Mitigation
vision-language

Computers in Biology and Medicine 2026

First posted Dec 1, 2025

sparse visual contrastive decoding
Authors: Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman KhalooCorresponding: not specifiedAffiliation: Department of computer science, University of Central Florida, Orlando, USA; Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA; Department of electrical engineering, Columbia university, New York, NY, USA; Technical University of Applied Sciences Regensburg, Regensburg, Germany; Department of Surgery, University of Calgary, Calgary, Alberta, Canada; University College of Nabi Akram, Tabriz, Iran; School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
121

Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)

Constructs an image-free text input as a negative reference and subtracts language prior scores during logit calculation.

contrastive decodinglanguage contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Aug 6, 2024

language contrastive decoding
Authors: Avshalom Manevich, Reut TsarfatyCorresponding: not specifiedAffiliation: Bar Ilan University
122

CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models

Contrasts logit distributions based on the original image versus self-generated descriptions to penalize text-reliant tokens.

contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Jun 4, 2024

contrastive decoding
Authors: Junho Kim, Hyunjun Kim, Yeonju KimCorresponding: Junho KimAffiliation: Integrated Vision and Language Lab, KAIST
123

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Suppresses event hallucinations by contrasting output probabilities between full and truncated videos.

contrastive decodingtemporal contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Sep 25, 2024

temporal contrastive decoding
Authors: Jiacheng Zhang, Yang Jiao, Shaoxiang ChenCorresponding: not specifiedAffiliation: Shanghai Key Lab of Intelligent Information Processing, School of CS, Fudan University; Shanghai Collaborative Innovation Center on Intelligent Visual Computing; Singapore University of Technology and Design; Shanghai Academy of Artificial Intelligence for Science; Meituan
124

Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding

Introduces an external CLIP model to score and guide candidate tokens during decoding to improve image-text consistency.

visual groundingexternal tools
Mitigation
vision-language

arXiv 2024

First posted Feb 23, 2024

external guided decoding
Authors: Ailin Deng, Zhirui Chen, Bryan HooiCorresponding: not specifiedAffiliation: School of Computing, National University of Singapore; Department of Industrial Systems Engineering and Management, National University of Singapore
125

Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding

Dynamically adjusts the logit weight distribution of text and image features during inference.

contrastive decodingre-balancing contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Sep 10, 2024

re-balancing contrastive decoding
Authors: Xiaoyu Liang, Jiayuan Yu, Lianrui MuCorresponding: Haoji HuAffiliation: College of Information Science and Electronic Engineering, Zhejiang University, China
126

Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding

Extracts internally generated facts as references for contrastive decoding to suppress over-extrapolation.

contrastive decodinginternal fact contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2025

First posted Feb 3, 2025

internal fact contrastive decoding
Authors: Chao Wang, Xuancheng Zhou, Weiwei FuCorresponding: Chao Wang, Yang ZhouAffiliation: School of Future Technology, Shanghai University, Shanghai, 200444, China; Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China; School of Mechatronic Engineering and Automation, Shanghai, 200444, China
127

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Introduces a control parameter during decoding to force switching between pure contextual description and parameterized knowledge imagination.

inference interventiongeneration control
Mitigation
vision-language

arXiv 2023

First posted Oct 3, 2023

inference intervention
Authors: Bohan Zhai, Shijia Yang, Chenfeng XuCorresponding: not specifiedAffiliation: ByteDance Inc.; Stanford University; UC Berkeley; UIUC
128

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Introduces language contrastive decoding to weaken misleading priors from user prompts to combat sycophancy-induced hallucinations.

contrastive decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Aug 21, 2024

contrastive decoding
Authors: Yunpu Zhao, Rui Zhang, Junbin XiaoCorresponding: Rui ZhangAffiliation: School of Computer Science and Technology, University of Science and Technology of China; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences; Department of Computer Science, National University of Singapore; University of Illinois Urbana-Champaign; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences
129

Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding

Utilizes image summaries as guidance signals to intervene in output probabilities for contextual consistency.

summary-guided decodinggeneration control
Mitigation
vision-language

arXiv 2024

First posted Oct 17, 2024

summary-guided decoding
Authors: Kyungmin Min, Minbeom Kim, Kang-il LeeCorresponding: not specifiedAffiliation: IPAI, Seoul National University; Dept. of ECE, Seoul National University
130

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

Contrasts logits from the original image and retrieved similar reference images to highlight core facts.

contrastive decodingretrieval
Mitigation
vision-language

arXiv 2025

First posted May 26, 2025

retrieval visual contrastive decoding
Authors: Jihoon Lee, Min SongCorresponding: not specifiedAffiliation: Yonsei University; Onoma AI
131

CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models

Generates multi-granularity hierarchical feedback to dynamically correct outputs.

coarse-to-fine feedback decodinggeneration control
Mitigation
vision-language

arXiv 2025

First posted Dec 29, 2025

coarse-to-fine feedback decoding
Authors: Zongsheng Cao, Yangfan He, Anran LiuCorresponding: Anran Liu, Zepeng WangAffiliation: Researcher; UMN; PCIE
132

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Trains an independent Reviser model specifically to detect and reconstruct hallucinated entities in outputs.

self-correctionpost-hoc revisionexternal grounding
Mitigation
vision-language

ICLR 2023

First posted Oct 1, 2023

post-hoc revision
Authors: Yiyang Zhou, Chenhang Cui, Jaehong YoonCorresponding: not specifiedAffiliation: UNC-Chapel Hill; Rutgers University; Columbia University; Stanford University
133

Woodpecker: Hallucination Correction for Multimodal Large Language Models

Calls external visual foundation models to verify entity facts, finally using an LLM to rewrite and correct.

verificationexternal tools
Mitigation
vision-language

SCIS 2024

First posted Oct 24, 2023

agent correction
Authors: Shukang Yin, Chaoyou Fu, Sirui ZhaoCorresponding: not specifiedAffiliation: School of Data Science, USTC & State Key Laboratory of Cognitive Intelligence; Tencent YouTu Lab
134

Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation

Constructs a pipeline system containing claim extraction, external detection verification, and corrected generation.

verificationexternal tools
Mitigation
vision-language

CVPR 2024

First posted Apr 30, 2024

pipeline agent
Authors: Yunhao Ge, Xiaohui Zeng, Jacob Samuel HuffmanCorresponding: not specifiedAffiliation: NVIDIA
135

Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning

Modifies reference images to expose hallucinating agents and replaces consensus-only debate with evidence-based factual verification.

verificationcounterfactual multi-agent gamingexternal grounding
Mitigation
vision-language

AAAI 2026

counterfactual multi-agent gaming
Authors: Dayong Liang, Xiao-Yong Wei, Changmeng ZhengCorresponding: not specifiedAffiliation: South China University of Technology, Guangzhou, China; Peng Cheng Laboratory, Shenzhen, China; Sichuan University, Chengdu, China; The Hong Kong Polytechnic University, Hong Kong, China
136

Mitigating Object and Relationship Hallucination in Large Vision Language Model with Multi-Agent Guidance

Coordinates specialized agents to identify and correct unsupported object and relationship claims.

verificationmulti-agent verificationexternal grounding
Mitigation
vision-language

ICASSP 2026

multi-agent verification
Authors: Soohyun Kim, Gusang Lee, Kyuhong ShimCorresponding: not specifiedAffiliation: Seoul National University, Department of Electrical and Computer Engineering, Korea; Sungkyunkwan University, Department of Computer Science and Engineering, Korea
137

CEBC: Conformal Evidence-Bounded Control for Low-Hallucination Vision-Language Generation

Calibrates an external detector on held-out data and minimally revises unsupported object mentions under explicit risk bounds.

self-correctionexternal toolsuncertainty
Mitigation
vision-language

ACL 2026

conformal evidence-bounded editing
Authors: Ashish Mishra, Tarun Kumar, Arpit ShahCorresponding: not specifiedAffiliation: Hewlett Packard Labs, Bangalore; Hewlett Packard Labs, USA
138

Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models

Uses external language models to ask questions about generated claims for self-verification.

verificationexternal tools
Mitigation
vision-language

arXiv 2024

First posted Feb 18, 2024

logical loop verification
Authors: Junfei Wu, Qiang Liu, Ding WangCorresponding: not specifiedAffiliation: Center for Research on Intelligent Perception and Computing; State Key Laboratory of Multimodal Artificial Intelligence Systems; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Nanjing University
139

Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning

Utilizes a bottom-up holistic reasoning framework combined with multi-perspective prompt verification.

verificationmultimodal reasoning
Mitigation
vision-language

arXiv 2024

First posted Dec 15, 2024

multi-perspective cross-checking
Authors: Shengqiong Wu, Hao Fei, Liangming PanCorresponding: Hao FeiAffiliation: National University of Singapore, Singapore; University of Arizona, USA; University of California, Santa Barbara, USA; Skywork AI, Singapore; Nanyang Technological University, Singapore
140

Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Uses an external detector to locate relation errors, then applies template-based secondary correction.

external toolsuncertainty
Mitigation
vision-language

arXiv 2024

First posted Aug 18, 2024

external detection calibration
Authors: Kening Zheng, Junkai Chen, Yibo YanCorresponding: Xuming HuAffiliation: Hong Kong University of Science and Technology (Guangzhou); Guangxi Zhuang Autonomous Region Big Data Research Institute; Hong Kong University of Science and Technology
141

Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination

Adopts a "Retrospect-then-Compare" multi-step prompt to make the model review visual details for correction.

multi-step retrospective promptingexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Mar 21, 2024

multi-step retrospective prompting
Authors: Dingchen Yang, Bowen Cao, Guang ChenCorresponding: not specifiedAffiliation: Tongji University; Peking University
142

Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting

Explicitly adds spatial relation logic rules into the prompt to correct spatial confusion.

spatial logic rule promptingexternal grounding
Mitigation
vision-language

arXiv 2025

First posted Feb 12, 2025

spatial logic rule prompting
Authors: Jiarui Wu, Zhuo Liu, Hangfeng HeCorresponding: not specifiedAffiliation: University of Rochester
143

RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models

Applies random geometric transformations to construct multiple views for consistency verification and correction.

verificationmulti-view consistencyexternal grounding
Mitigation
vision-language

arXiv 2024

First posted May 28, 2024

multi-view consistency
Authors: Sangmin Woo, Jaehyuk Jang, Donguk KimCorresponding: not specifiedAffiliation: KAIST
144

Systematic Reward Gap Optimization for Mitigating VLM Hallucinations

The model generates a draft, then independently reviews and self-scores sub-topics for correction.

topic-level self-scoringexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Nov 26, 2024

topic-level self-scoring
Authors: Lehan He, Zeren Chen, Zhelun ShiCorresponding: not specifiedAffiliation: Shanghai AI Laboratory; School of Software, Beihang University; Shanghai Innovation Institute; Tsinghua University
145

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

Deploys a network of introspective and external verification agents for interactive debate.

verificationexternal tools
Mitigation
vision-language

arXiv 2025

First posted Dec 2, 2025

cross-modal introspective debate
Authors: Zhongyu Yang, Yingfang Yuan, Xuanming JiangCorresponding: Xuanming Jiang, Wei PangAffiliation: Xi'an Jiyun Technology Co., Ltd., Xi'an, China; BCML, Heriot-Watt University, Edinburgh, UK
146

Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Evaluates model uncertainty to actively trigger knowledge base retrieval for entity correction.

retrievalexternal toolsuncertainty
Mitigation
vision-language

TOMM 2025

First posted Aug 1, 2024

dynamic rag supplement
Authors: Xiaoye Qu, Qiyuan Chen, Wei WeiCorresponding: not specifiedAffiliation: Huazhong University of Science and Technology; Zhejiang University; Xiamen University; Zhejiang Gongshang University
147

A Unified Hallucination Mitigation Framework for Large Vision-Language Models

Builds a unified cross-modal diagnosis pipeline utilizing external discriminators to intercept and modify outputs.

external toolsexternal discriminator systemexternal grounding
Mitigation
vision-language

TMLR 2024

First posted Sep 24, 2024

external discriminator system
Authors: Yue Chang, Liqiang Jing, Xiaopeng ZhangCorresponding: not specifiedAffiliation: The University of Texas at Dallas
148

Mitigating Object Hallucinations via Sentence-Level Early Intervention

Implements factual truncation prompts in early sentence stages of generation to prevent error propagation.

early sentence truncationexternal grounding
Mitigation
vision-language

ICCV 2025

First posted Jul 16, 2025

early sentence truncation
Authors: Shangpin Peng, Senqiao Yang, Li JiangCorresponding: not specifiedAffiliation: Harbin Institute of Technology, Shenzhen; The Chinese University of Hong Kong; The Chinese University of Hong Kong, Shenzhen
149

Exploring the Transferability of Visual Prompting for Multimodal Large Language Models

Overlays visual prompts (e.g., circles, bounding boxes) on the original image to force model attention.

attention steeringoriginal image overlayexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Apr 17, 2024

original image overlay
Authors: Yichi Zhang, Yinpeng Dong, Siyuan ZhangCorresponding: not specifiedAffiliation: Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center; THBI Lab, BNRist Center, Tsinghua University, Beijing 100084, China; RealAI; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong
150

What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models

Introduces "What if" prompting to compel the model to self-reflect on its initial judgments.

counterfactual promptingexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Mar 20, 2024

counterfactual prompting
Authors: Junho Kim, Yeon Ju Kim, Yong Man RoCorresponding: Yeon Ju KimAffiliation: Integrated Vision and Language Lab KAIST South Korea
151

From Training-Free to Adaptive: Empirical Insights into MLLMs’ Understanding of Detection Information

Feeds information identified by external object detection models directly to the MLLM as context.

external toolsexternal visual guidanceexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Jan 31, 2024

external visual guidance
Authors: Qirui Jiao, Daoyuan Chen, Yilun HuangCorresponding: not specifiedAffiliation: Sun Yat-Sen University, Shenzhen, China; Alibaba Group, Hangzhou, China
152

Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

Pre-supplements missing low-level visual perception details in the input prompt to bridge the gap.

perception detail supplementexternal grounding
Mitigation
vision-language

arXiv 2024

First posted May 24, 2024

perception detail supplement
Authors: Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal KumarCorresponding: not specifiedAffiliation: University of Maryland, College Park, USA; Adobe, USA
153

Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?

Strictly controls the requested detail quantity via prompt to balance richness and hallucination rate.

prompt detail constraintexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Jun 18, 2024

prompt detail constraint
Authors: Mingqian Feng, Yunlong Tang, Zeliang ZhangCorresponding: not specifiedAffiliation: University of Rochester
154

Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning

Adopts multi-view prompts and multi-path reasoning, aggregating and comparing answers.

multimodal reasoningmulti-path votingexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Aug 30, 2024

multi-path voting
Authors: Xiaoye Qu, Jiashuo Sun, Wei WeiCorresponding: not specifiedAffiliation: Huazhong University of Science and Technology; Xiamen University; The Chinese University of Hong Kong
155

Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

Uses prompts to guide the model to actively re-extract visual clues for confirmation during generation.

memory retracing confirmationexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Oct 4, 2024

memory retracing confirmation
Authors: Xin Zou, Yizhou Wang, Yibo YanCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; China University of Geosciences; University of Technology Sydney
156

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Introduces an external Vision Value Model to assist tree search, guiding autoregressive generation.

external toolsexternal scoring searchexternal grounding
Mitigation
vision-language

arXiv 2024

First posted Dec 4, 2024

external scoring search
Authors: Xiyao Wang, Zhengyuan Yang, Linjie LiCorresponding: not specifiedAffiliation: University of Maryland, College Park; Microsoft
157

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

Uses black-box optimization to automatically search for "visual prompt patches" that suppress hallucinations.

external toolsautomated patch searchexternal grounding
Mitigation
vision-language

arXiv 2025

First posted Apr 30, 2025

automated patch search
Authors: Sangmin Woo, Kang Zhou, Yun ZhouCorresponding: Sangmin Woo, Haibo DingAffiliation: Amazon AWS AI; KAIST
158

Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models

Uses uncertainty-aware contrastive calibration to suppress hallucinated actions in vision-language-action generation.

uncertaintyentropy-optimized decodinggeneration control
Mitigation
vision-language

Neurocomputing 2026

entropy-optimized decoding
Authors: Ye Qiu, Zhaoxin Fan, Qingchen YuCorresponding: not specifiedAffiliation:
159

VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models

Uses a reinforcement-learned caption model as an auxiliary visual clue source and applies image-confidence constraints during generation.

uncertaintyvisual-clue-guided decodinggeneration control
Mitigation
vision-language

AAAI 2026

visual-clue-guided decoding
Authors: Guoqing Chen, Fu Zhang, Bingqian LiuCorresponding: not specifiedAffiliation: Northeastern University
160

CTDD: Cumulative Trend Divergence Decoding for Mitigating Hallucination in Large Vision-Language Models

Tracks accumulated divergence in token-distribution trends and calibrates generation when visual and linguistic evidence separate.

uncertaintycumulative trend-divergence decodinggeneration control
Mitigation
vision-language

LNCS 2026

cumulative trend-divergence decoding
Authors: Jiani Hou, Zhixuan You, Siyu LuoCorresponding: Bing GuoAffiliation: College of Software Engineering, Sichuan University, Chengdu, China; International Economics and Trade, Central University of Finance and Economics, Beijing, China; College of Computer Science, Sichuan University, Chengdu, China
161

Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Filters hallucinated words by comparing the confidence of different generation paths via the model's own introspective mechanism.

uncertaintydata curation
Mitigation
vision-language

arXiv 2024

First posted Aug 4, 2024

self-introspective decoding
Authors: Fushuo Huo, Wenchao Xu, Zhong ZhangCorresponding: Wenchao Xu, Peilin ZhaoAffiliation: Department of Computing, The Hong Kong Polytechnic University; Division of Integrative Systems and Design, Hong Kong University of Science and Technology; Tencent AI Lab; Huazhong University of Science and Technology; Tsinghua University
162

HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning

Evaluates synthetic images / vqa with Gen tasks using Acc metrics.

object-levelgenerationsynthetic images vqa
Benchmarks
vision-language

ECCV 2024

First posted Jul 22, 2024

7,748 / Acc
Authors: Zhecan Wang, Garrett Bingham, Adams YuCorresponding: not specifiedAffiliation: Columbia University, New York, NY 10027; Google DeepMind, Mountain View, CA 94043
163

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

Evaluates speech/sound/music with Dis & Gen tasks using Acc / Hallucination Rate metrics.

discriminationgenerationspeechsoundmusic
Benchmarks
vision-language

ACL 2026

5,000+ / Acc / Hallucination Rate
Authors: Feiyu Zhao, Yiming Chen, Wenhuan LuCorresponding: Xianghu YueAffiliation: College of Intelligence and Computing, Tianjin University, China; ASUS Intelligent Cloud Services, Singapore
164

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Evaluates medical context with Dis & Gen tasks using MediHall Score metrics.

object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language

arXiv 2024

First posted Jun 14, 2024

889,125 / MediHall Score
Authors: Jiawei Chen, Dingkang Yang, Tong WuCorresponding: not specifiedAffiliation: Academy for Engineering and Technology, Fudan University, Shanghai, China; Tencent Youtu Lab, Shanghai, China; Cognition and Intelligent Technology Laboratory, Shanghai, China; Engineering Research Center of AI and Robotics, Ministry of Education, Shanghai, China; AI and Unmanned Systems Engineering Research Center of Jilin Province, Changchun, China
165

VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Evaluates temporal/extrinsic with Dis & Gen tasks using Acc metrics.

object-levelrelation-leveldiscriminationgeneration
Benchmarks
vision-language

arXiv 2024

First posted Jun 24, 2024

1,800 / Acc
Authors: Yuxuan Wang, Yueqian Wang, Dongyan ZhaoCorresponding: not specifiedAffiliation: Beijing Institute for General Artificial Intelligence, Beijing China; State Key Laboratory of General Artificial Intelligence, Beijing, China; Wangxuan Institute of Computer Technology, Peking University, Beijing, China; Computer Science and Engineering, University of California, Santa Cruz
166

MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Evaluates medical context with Dis & Gen tasks using Characterization Score metrics.

object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language

arXiv 2024

First posted Jul 3, 2024

2,000 / Characterization Score
Authors: Zishan Gu, Changchang Yin, Fenglin LiuCorresponding: not specifiedAffiliation: The Ohio State University; University of Oxford
167

Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs

Evaluates grounded medical with Gen tasks using Acc / Loc metrics.

object-levelgenerationgrounded medical
Benchmarks
vision-language

arXiv 2025

First posted Apr 30, 2025

67,000 / Acc / Loc
Authors: Dung Nguyen, Minh Khoi Ho, Huy TaCorresponding: not specifiedAffiliation: Hanoi University of Science and Technology; University of Wollongong; Australian Institute for Machine Learning, The University of Adelaide; Griffith University; College of Medicine and Public Health, Flinders University
168

MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models

Evaluates multi-image reasoning with Dis & Gen tasks using Acc metrics.

object-leveldiscriminationgenerationmulti-image reasoning
Benchmarks
vision-language

arXiv 2025

First posted Aug 1, 2025

3,000 / Acc
Authors: Jiale Li, Mingrui Wu, Zixiang JinCorresponding: not specifiedAffiliation: Xiamen University; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China; Zhongguancun Academy
169

The Instinctive Bias: Spurious Images lead to Illusion in MLLMs

Evaluates spurious images with Dis tasks using Acc metrics.

object-leveldiscriminationspurious images
Benchmarks
vision-language

arXiv 2024

First posted Feb 6, 2024

7,308 / Acc
Authors: Tianyang Han, Qing Lian, Rui PanCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology; University of Illinois at Urbana-Champaign; The Hong Kong Polytechnic University
170

Explore the Hallucination on Low-level Perception for MLLMs

Evaluates low-level perception with Gen tasks using Acc metrics.

attribute-levelgenerationlow-level perception
Benchmarks
vision-language

arXiv 2024

First posted Sep 15, 2024

4,989 / Acc
Authors: Yinan Sun, Zicheng Zhang, Haoning WuCorresponding: not specifiedAffiliation: Shanghai Jiao Tong University; Nanyang Technological University
171

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

Evaluates imaginary vqa with Dis & Gen tasks using Acc metrics.

object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language

arXiv 2024

First posted Sep 19, 2024

13,500 / Acc
Authors: Zhecan Wang, Junzhang Liu, Chia-Wei TangCorresponding: not specifiedAffiliation: Columbia University; UCLA; Virginia Tech
172

Detecting and Preventing Hallucinations in Large Vision Language Models

Evaluates fine-grained with Dis tasks using Reward Score metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

AAAI 2024

First posted Aug 11, 2023

16,000 / Reward Score
Authors: Anisha Gunjal, Jihan Yin, Erhan BasCorresponding: not specifiedAffiliation:
173

Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models

Evaluates negative pronoun with Dis tasks using Acc/METEOR metrics.

object-leveldiscriminationnegative pronoun
Benchmarks
vision-language

ACL-W 2024

First posted Oct 9, 2023

29,500 / Acc/METEOR
Authors: Holy Lovenia, Wenliang Dai, Samuel CahyawijayaCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology
174

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

Evaluates free-form generations with Gen tasks using P/R/F metrics.

object-levelgenerationfree-form generations
Benchmarks
vision-language

CVPR 2024

First posted May 8, 2024

5,000 / P/R/F
Authors: Prannay Kaul, Zhizhong Li, Hao YangCorresponding: not specifiedAffiliation: VGG, University of Oxford; AWS AI Labs
175

Quantity Matters: Towards Assessing and Mitigating Number Hallucination in Large Vision-Language Models

Evaluates object counting with Dis tasks using Acc / Consistency metrics.

object-levelattribute-leveldiscriminationobject counting
Benchmarks
vision-language

arXiv 2024

First posted Mar 3, 2024

20,000 / Acc / Consistency
Authors: Huixuan Zhang, Junzhe Zhang, Xiaojun WanCorresponding: not specifiedAffiliation: School of Electronics Engineering and Computer Science, Peking University; Wangxuan Institute of Computer Technology, Peking University
176

CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning

Evaluates multimodal hallucination with Dis tasks using Acc/P/R/F1/Specificity metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

NeurIPS-W 2023

First posted Sep 5, 2023

40,367 / Acc/P/R/F1/Specificity
Authors: Hongyu Hu, Jiyuan Zhang, Minyi ZhaoCorresponding: not specifiedAffiliation: ByteDance Inc, Shanghai
177

Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Evaluates visual grounding with Dis tasks using Accuracy metrics.

object-levelrelation-leveldiscriminationvisual grounding
Benchmarks
vision-language

CVPR 2025

First posted Dec 3, 2023

31,373 / Accuracy
Authors: Andrés Villa, Juan Carlos León Alcázar, Alvaro SotoCorresponding: not specifiedAffiliation: King Abdullah University of Science and Technology (KAUST); Pontificia Universidad Católica de Chile
178

Unified Hallucination Detection for Multimodal Large Language Models

Evaluates t2i with Gen tasks using Acc/P/R/F metrics.

object-levelattribute-levelgenerationt2i
Benchmarks
vision-language

ACL 2024

First posted Feb 5, 2024

1,860 / Acc/P/R/F
Authors: Xiang Chen, Chenxi Wang, Yida XueCorresponding: Huajun ChenAffiliation: College of Computer Science and Technology, Zhejiang University; School of Software Technology, Zhejiang University; Zhejiang University-Ant Group Joint Laboratory of Knowledge Graph; Ant Group
179

Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

Evaluates event hallucination with Dis & Gen tasks using Acc/Score metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

ACMMM 2024

First posted Feb 24, 2024

10,000 / Acc/Score
Authors: Chaoya Jiang, Hongrui Jia, Wei YeCorresponding: not specifiedAffiliation: National Engineering Research Center for Software Engineering, Peking University; DAMO Academy, Alibaba Group
180

VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models

Evaluates multimodal hallucination with Gen tasks using Faithfulness & Coverage metrics.

object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language

ACL 2024

First posted Apr 22, 2024

211 / Faithfulness & Coverage
Authors: Haoyi Qiu, Wenbo Hu, Zi-Yi DouCorresponding: not specifiedAffiliation: University of California, Los Angeles
181

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Evaluates multimodal hallucination with Dis tasks using ASR/MASR/CASR metrics.

object-leveldiscrimination
Benchmarks
vision-language

EMNLP 2024

First posted Jun 16, 2024

5,000 / ASR/MASR/CASR
Authors: Xiyang Wu, Tianrui Guan, Dianqi LiCorresponding: not specifiedAffiliation: University of Maryland, College Park
182

ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models

Evaluates open-set with Dis & Gen tasks using AMBER/Acc metrics.

object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language

CVPR 2025

First posted Sep 14, 2024

8,786 / AMBER/Acc
Authors: Yahan Tu, Rui Hu, Jitao SangCorresponding: not specifiedAffiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
183

Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models

Evaluates ocr/action/counting with Dis & Gen tasks using Hallucination Rate metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

arXiv 2024

First posted Jun 24, 2024

4,000 / Hallucination Rate
Authors: Bei Yan, Jie Zhang, Zheng YuanCorresponding: not specifiedAffiliation: Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Key Laboratory of Al Safety, Chinese Academy of Sciences
184

FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs

Evaluates eval method with Dis & Gen tasks using Acc/P/R/F1 metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

arXiv 2024

First posted Sep 20, 2024

N/A / Acc/P/R/F1
Authors: Bowen Yan, Zhengsong Zhang, Liqiang JingCorresponding: not specifiedAffiliation: University of Texas at Dallas, Richardson, United States
185

TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions

Evaluates unanswerable qs with Dis tasks using Acc metrics.

object-levelrelation-leveldiscriminationunanswerable qs
Benchmarks
vision-language

arXiv 2024

First posted Oct 5, 2024

2,354 / Acc
Authors: Xingwei He, Qianru Zhang, A-Long JinCorresponding: not specifiedAffiliation: The University of Hong Kong; Xi’an Jiaotong-Liverpool University; School of Computer Science and Engineering, Beihang University, Beijing, China; State Key Laboratory of Software, Development Environment; Zhongguancun Laboratory
186

MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models

Evaluates manipulation/ooc/veracity with Dis & Gen tasks using Acc/F1 metrics.

discriminationgenerationmanipulationoocveracity
Benchmarks
vision-language

arXiv 2024

First posted Jun 17, 2024

35,000 / Acc/F1
Authors: Shengkang Wang, Hongzhan Lin, Ziyang LuoCorresponding: not specifiedAffiliation: Beijing University of Posts and Telecommunications; Hong Kong Baptist University; Hong Kong University of Science and Technology
187

Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs

Evaluates perturbed inputs with Dis tasks using Acc metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

arXiv 2024

First posted Aug 2, 2024

1,260 / Acc
Authors: Peng Ding, Jingyu Wu, Jun KuangCorresponding: not specifiedAffiliation: National Key Laboratory for Novel Software Technology, Nanjing University; College of Computer Science and Technology, Zhejiang University; Zhejiang-Singapore Innovation and AI Joint Research Lab, Zhejiang University
188

CAST: Cross-modal Alignment Similarity Test for Vision Language Models

Evaluates self-consistency with Dis & Gen tasks using Similarity metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

arXiv 2024

First posted Sep 17, 2024

15,000 / Similarity
Authors: Gautier Dagan, Olga Loginova, Anil BatraCorresponding: not specifiedAffiliation: University of Edinburgh; University of Trento
189

FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Evaluates obj. counting with Gen tasks using FaithScore metrics.

object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language

EMNLP 2024

First posted Nov 2, 2023

2,000 / FaithScore
Authors: Liqiang Jing, Ruosen Li, Yunmo ChenCorresponding: not specifiedAffiliation: University of Texas at Dallas; Johns Hopkins University; University of Notre Dame
190

ALOHa: A New Measure for Hallucination in Captioning Models

Evaluates captioning metric with Gen tasks using ALOHa metrics.

object-levelgenerationcaptioning metric
Benchmarks
vision-language

arXiv 2024

First posted Apr 3, 2024

N/A / ALOHa
Authors: Suzanne Petryk, David M. Chan, Anish KachinthayaCorresponding: not specifiedAffiliation: University of California, Berkeley
191

Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?

Evaluates metric method with Gen tasks using CHAIR-MEN / FaithScore metrics.

object-levelgenerationmetric method
Benchmarks
vision-language

arXiv 2024

First posted Jun 20, 2024

N/A / CHAIR-MEN / FaithScore
Authors: Gregor Geigle, Radu Timofte, Goran GlavašCorresponding: not specifiedAffiliation: WüNLP; Computer Vision Lab, CAIDAS, University of Würzburg
192

Visual Hallucinations of Multi-modal Large Language Models

Evaluates visual hallucination with Dis & Gen tasks using Acc metrics.

object-levelattribute-leveldiscriminationgeneration
Benchmarks
vision-language

ACL 2024

First posted Feb 22, 2024

1,200 / Acc
Authors: Wen Huang, Hongbin Liu, Minxin GuoCorresponding: not specifiedAffiliation: University of Science & Technology of China; Duke University; The University of Hong Kong
193

PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset

Evaluates sentiment with Dis tasks using PhD Index metrics.

object-levelattribute-levelrelation-leveldiscrimination
Benchmarks
vision-language

CVPR 2025

First posted Mar 17, 2024

102,564 / PhD Index
Authors: Jiazhen Liu, Yuhan Fu, Ruobing XieCorresponding: not specifiedAffiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
194

Automated Multi-level Preference for MLLMs

Evaluates multi-round conv with Gen tasks using Preference Score metrics.

object-levelattribute-levelgenerationmulti-round conv
Benchmarks
vision-language

NeurIPS 2024

First posted May 18, 2024

N/A / Preference Score
Authors: Mengxi Zhang, Wenhao Wu, Yu LuCorresponding: not specifiedAffiliation: Baidu Inc.; Tianjin University; The University of Sydney; University of Technology Sydney; Tsinghua University; Chinese Academy of Science
195

Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models

Evaluates multimodal hallucination with Dis tasks using Acc/P/R/F1 metrics.

object-levelrelation-leveldiscrimination
Benchmarks
vision-language

ICML 2024

First posted Jun 24, 2024

8,030 / Acc/P/R/F1
Authors: Mingrui Wu, Jiayi Ji, Oucheng HuangCorresponding: Jiayi JiAffiliation: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China
196

Multi-Object Hallucination in Vision-Language Models

Evaluates multi-object with Gen tasks using Accuracy metrics.

object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language

NeurIPS 2024

First posted Jul 8, 2024

5,000 / Accuracy
Authors: Xuweiyi Chen, Ziqiao Ma, Xuejun ZhangCorresponding: not specifiedAffiliation: University of Michigan; University of Virginia; New York University
197

BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models

Evaluates before-after changes with Dis tasks using TU/IG/SB/ID metrics.

object-leveldiscriminationbefore-after changes
Benchmarks
vision-language

ECCV 2024

First posted Jul 18, 2024

26,064 / TU/IG/SB/ID
Authors: Moon Ye-Bin, Nam Hyeon-Woo, Wonseok ChoiCorresponding: not specifiedAffiliation: Dept. of EE, POSTECH, Korea; Grad. School of AI, POSTECH, Korea; Institute for Convergence Research and Education in Advanced Technology, Yonsei University, Korea
198

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception

Evaluates cp-bench mitigation with Gen tasks using Acc metrics.

object-levelgenerationcp-bench mitigation
Benchmarks
vision-language

CVPR 2025

First posted Apr 29, 2025

N/A / Acc
Authors: Yuanchen Wu, Lu Zhang, Hang YaoCorresponding: not specifiedAffiliation: School of Computer Engineering & Science, Shanghai University; Tencent YouTu Lab
199

EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflow

Evaluates ophthalmology with Dis & Gen tasks using N/A metrics.

discriminationgenerationophthalmology
Benchmarks
vision-language

Information Fusion 2026

First posted Jul 24, 2025

N/A / N/A
Authors: Xiaoyu Pan, Yang Bai, Ke ZouCorresponding: not specifiedAffiliation: Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR), 1 Fusionopolis Way, Singapore, 138632; Centre for Innovation and Precision Eye Health; and Department of Ophthalmology, NUHS Tower Block, Level 7, 1E Kent Ridge Road, Singapore, 119228; Singapore Eye Research Institute, Singapore National Eye Centre, 20 College Road, Singapore, 169856
200

Evaluation and Analysis of Hallucination in Large Vision-Language Models

Evaluates multimodal hallucination with Gen tasks using LLM Rating metrics.

object-levelgeneration
Benchmarks
vision-language

arXiv 2023

First posted Aug 29, 2023

25,000 / LLM Rating
Authors: Junyang Wang, Yiyang Zhou, Guohai XuCorresponding: not specifiedAffiliation: School of Computer and Information Technology, Beijing Jiaotong University, Beijing, China; School of Software Engineering, Xi'an Jiaotong University, Xi'an, China; School of Software, Shandong University, Jinan, China; MAIS, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China; DAMO Academy, Alibaba Group
201

Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Evaluates model bias with Gen tasks using Human Assessment metrics.

generationmodel bias
Benchmarks
vision-language

arXiv 2023

First posted Nov 6, 2023

370 / Human Assessment
Authors: Chenhang Cui, Yiyang Zhou, Xinyu YangCorresponding: not specifiedAffiliation: UNC-Chapel Hill; Carnegie Mellon University; Stanford University; Rutgers University
202

MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity

Evaluates instruction tuning with Gen tasks using Acc metrics.

object-levelattribute-levelrelation-levelgeneration
Benchmarks
vision-language

arXiv 2024

First posted Jul 22, 2024

973,000 / Acc
Authors: Yangzhou Liu, Yue Cao, Zhangwei GaoCorresponding: not specifiedAffiliation: Shanghai AI Laboratory, Shanghai 200232, China; SenseTime Research, Shanghai 200233, China; Tsinghua University, Beijing 100084, China; Nanjing University, Nanjing 210023, China; Fudan University, Shanghai 200433, China; The Chinese University of Hong Kong, Hong Kong 999077, China; Shanghai Jiao Tong University, Shanghai 200240, China
203

MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Evaluates decoding method with Gen tasks using Acc metrics.

object-levelgenerationdecoding method
Benchmarks
vision-language

arXiv 2024

First posted Oct 15, 2024

N/A / Acc
Authors: Chenxi Wang, Xiang Chen, Ningyu ZhangCorresponding: not specifiedAffiliation: Zhejiang University; National University of Singapore, NUS-NCS Joint Lab, Singapore
204

Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Evaluates unsolvable (upd) with Dis tasks using Acc metrics.

discriminationunsolvable (upd)
Benchmarks
vision-language

arXiv 2024

First posted Mar 29, 2024

2,095 / Acc
Authors: Atsuyuki Miyai, Jingkang Yang, Jingyang ZhangCorresponding: not specifiedAffiliation: The University of Tokyo; S-Lab, Nanyang Technological University; Duke University; University of Wisconsin-Madison; LY Corporation; Tokyo University of Science
205

LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences

Evaluates societal bias/pref with Dis tasks using Alignment/Bias metrics.

discriminationsocietal biaspref
Benchmarks
vision-language

arXiv 2025

First posted Jul 25, 2025

N/A / Alignment/Bias
Authors: Yusuke Hirota, Boyi Li, Ryo HachiumaCorresponding: not specifiedAffiliation: NVIDIA Research; Osaka University; Stanford University
206

How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts

Evaluates deceptive prompts with Gen tasks using Acc metrics.

object-levelattribute-levelgenerationdeceptive prompts
Benchmarks
vision-language

arXiv 2024

First posted Feb 20, 2024

1,000 / Acc
Authors: Yusu Qian, Haotian Zhang, Yinfei YangCorresponding: not specifiedAffiliation: Apple
207

VLind-Bench: Measuring Language Priors in Large Vision-Language Models

Evaluates language priors with Dis tasks using Acc (Y/N) metrics.

object-leveldiscriminationlanguage priors
Benchmarks
vision-language

arXiv 2024

First posted Jun 13, 2024

2,576 / Acc (Y/N)
Authors: Kang-il Lee, Minbeom Kim, Seunghyun YoonCorresponding: not specifiedAffiliation: ECE, SNU; IPAI, SNU
208

MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models

Evaluates relation understanding with Dis & Gen tasks using Acc metrics.

object-levelrelation-leveldiscriminationgeneration
Benchmarks
vision-language

arXiv 2024

First posted Jun 13, 2024

22,500 / Acc
Authors: Jiahao Nie, Gongjie Zhang, Wenbin AnCorresponding: not specifiedAffiliation: IGP, Nanyang Technological University; Nanyang Technological University; Alibaba DAMO Academy; Xi’an Jiaotong University
209

Understanding Multimodal Hallucination with Parameter-Free Representation Alignment

Evaluates alignment method with Dis tasks using Pfram metrics.

object-leveldiscriminationalignment method
Benchmarks
vision-language

arXiv 2024

First posted Sep 2, 2024

N/A / Pfram
Authors: Yueqian Wang, Jianxin Liang, Yuxuan WangCorresponding: not specifiedAffiliation: Wangxuan Institute of Computer Technology, Peking University; Beijing Institute for General Artificial Intelligence; National Key Laboratory of General Artificial Intelligence
210

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

A training-free detector that combines global scene similarity with local visual grounding for object hallucination detection.

global-local similarityVisual Logit Lensobject grounding
Detection
vision-language

NeurIPS 2025

First posted Aug 27, 2025

global-local grounding
Authors: Seongheon Park, Sharon LiCorresponding: not specifiedAffiliation: University of Wisconsin–Madison
211

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

A training-free object hallucination detector that uses calibrated local and instruction-context consistency scores.

instruction embeddingsLogit Lensobject hallucination
Detection
vision-language

ICML 2026

First posted May 12, 2026

instruction-token scoring
Authors: Runhe Lai, Xinhua Lu, Yanqi WuCorresponding: Weijiang Yu, Ruixuan WangAffiliation: Sun Yat-sen University; Peng Cheng Laboratory; Key Laboratory of Machine Intelligence and Advanced Computing, MOE
212

Uncertainty Estimation in Autoregressive Structured Prediction

A general ensemble-based framework for token-level and sequence-level uncertainty estimation in autoregressive structured prediction.

ensemble uncertaintytoken-level estimatessequence-level estimates
Quantification
vision-language

ICLR 2021

First posted Feb 18, 2020

autoregressive uncertainty
Authors: Andrey Malinin, Mark GalesCorresponding: not specifiedAffiliation: Yandex; Higher School of Economics; University of Cambridge
213

Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT

ACT-ViT treats full layer-by-token activation tensors as image-like inputs for cross-LLM hallucination detection.

activation tensorsvision transformercross-model transfer
Detection
vision-language

NeurIPS 2025

First posted Sep 30, 2025

activation-tensor detection
Authors: Guy Bar-Shalom, Fabrizio Frasca, Yaniv GalronCorresponding: not specifiedAffiliation: Technion
214

TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention

A latent-state detector identifies hallucinated object tokens and guides decoding toward a shared truthful direction.

truthful directionlatent subspacepre-intervention
Detection
vision-language

ICCV 2025

First posted Mar 13, 2025

latent-state detection
Authors: Jinhao Duan, Fei Kong, Hao ChengCorresponding: Kaidi XuAffiliation: Drexel University; University of Electronic Science and Technology of China; Hong Kong University of Science and Technology (Guangzhou)
215

Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations

A patch-level detector combines attention dispersion and cross-modal grounding consistency to localize hallucinated tokens.

attention dispersioncross-modal groundingpatch-level detection
Detection
vision-language

CVPR 2026

First posted Apr 6, 2026

token-grounding detection
Authors: Tuan Dung Nguyen, Minh Khoi Ho, Qi ChenCorresponding: Phi Le Nguyen, Vu Minh Hieu PhanAffiliation: Hanoi University of Science and Technology; Australian Institute for Machine Learning, University of Adelaide; Mohamed bin Zayed University of Artificial Intelligence