Skip to content
Home›Mitigation›Overview
🛡️

Mitigation Methods

Methods for reducing multimodal hallucinations.

177Curated entries
6Categories
244Verified links
161Dated papers

Repos, benchmarks, and short notes.

177 curated entries

Method Categories

Curated Methods

1

ViHallu

Visual variations and visual instruction construction improve visual-semantic alignment for LVLM hallucination mitigation.

visual variationsinstruction tuningalignment
Fine-tuning
vision-language

ACM MM 2025

First posted Jul 29, 2025

visual alignment
Authors: Ziyun Dai, Xiaoqiang Li, Shaohua ZhangCorresponding: not specifiedAffiliation: Shanghai University; Shanghai Business School
2

ToR

Token reweighting jointly optimizes perception and reasoning tokens for multimodal RLVR.

RLVRtoken reweightinggrounded reasoning
Perception-grounded RL
vision-language

arXiv 2026

First posted Mar 26, 2026

RLVR / grounding
Authors: Jinda Lu, Junkang Wu, Jinghan LiCorresponding: Jinda LuAffiliation: University of Science and Technology of China
3

PGPO

Policy optimization with token-level visual dependency for grounded multimodal reasoning.

RLVRpolicy optimizationvisual dependency
Perception-grounded RL
vision-language

arXiv 2026

First posted Apr 2, 2026

RLVR / token credit
Authors: Zekai Ye, Qiming Li, Xiaocheng FengCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory
4

VEPO

Vision-anchored token selection combines visual sensitivity with entropy during RL.

RLVRtoken selectionvisual reasoning
Perception-grounded RL
vision-language

arXiv 2026

First posted Jun 2, 2026

RLVR / visual reasoning
Authors: Senjie Jin, Peixin Wang, Boyang LiuCorresponding: Tao GuiAffiliation: Fudan University
5

AOD

Adversarial orthogonal disentanglement with training-free contrastive decoding for LVLM hallucination mitigation.

contrastive decodingdisentanglementtraining-free
Activation Editing
vision-language

arXiv 2026

First posted May 25, 2026

contrastive decoding
Authors: Ruoxi Cheng, Haoxuan Ma, Zhengfei HaiCorresponding: Xingjun MaAffiliation: Fudan University / Tencent; Nanjing University; Southeast University
6

Search-G1

Grounded search agents via representation-based intrinsic rewards.

search agentsintrinsic rewardsgrounding
Tool-augmented
vision-language

arXiv pending

awaiting release
arXiv pending
Author and affiliation metadata pending a public paper record.
7

HIRE

Intermediate representation editing for hallucination mitigation.

representation editingactivation steeringLVLM
Activation Editing
vision-language

arXiv 2026

First posted Mar 31, 2026

representation editing
Authors: Wei Suo, Hanzu Zhang, Lijun ZhangCorresponding: Peng WangAffiliation: Northwestern Polytechnical University
8

NoLan

Suppress language priors dynamically during generation.

language priorsdynamic suppressionobject hallucination
Decoding-time
vision-language

arXiv 2026

First posted Feb 25, 2026

language-prior suppression
Authors: Lingfeng Ren, Weihao Yu, Runpeng YuCorresponding: Weihao Yu, Xinchao WangAffiliation: National University of Singapore; Peking University Shenzhen Graduate School
9

MCoT

Mitigate reasoning-time hallucinations in multimodal CoT.

multimodal CoTreasoningmitigation
Verification-based
vision-language

CVPR 2026

First posted Mar 28, 2026

multimodal CoT
Authors: Ji Ma, Wei Suo, Peng WangCorresponding: Wei SuoAffiliation: Northwestern Polytechnical University
10

R-CoV

Region-aware chain-of-verification for LVLM hallucination.

region-awarechain-of-verificationobject hallucination
Verification-based
vision-language

arXiv 2026

First posted Apr 22, 2026

region verification
Authors: Jiahao Xie, Alessio Tonioni, Nathalie RauschmayrCorresponding: not specifiedAffiliation: Max Planck Institute for Informatics / VIA Research Center; Google
11

LEAD

Latent entropy-aware decoding for multimodal reasoning models.

latent entropydecodingmultimodal reasoning
Uncertainty-aware
vision-language

arXiv 2026

First posted Mar 9, 2026

entropy-aware decoding
Authors: Zhongxing Xu, Zhonghua Wang, Zhe QianCorresponding: not specifiedAffiliation: Monash University
12

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Constructs a robust instruction tuning dataset with positive and negative samples to update the model.

instruction tuningdata curation
Fine-tuning
vision-language

ICLR 2024

First posted Jun 26, 2023

instruction tuning
Authors: Fuxiao Liu, Kevin Lin, Linjie LiCorresponding: not specifiedAffiliation: University of Maryland, College Park; Microsoft Corporation
13

Aligning Large Multimodal Models with Factually Augmented RLHF

Introduces a factually augmented mechanism, using reinforcement learning from human feedback to align the model.

reinforcement learningrlhfmodel alignment
Perception-grounded RL
vision-language

ACL 2024

First posted Sep 25, 2023

rlhf
Authors: Zhiqing Sun, Sheng Shen, Shengcao CaoCorresponding: not specifiedAffiliation: UC Berkeley; CMU; UIUC; UW–Madison; UMass Amherst; Microsoft Research; MIT-IBM Watson AI Lab
14

Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Specifically trains the MLLM to generate self-feedback and revise its output based on it.

self-correctionfine-tuningmodel alignment
Fine-tuning
vision-language

NAACL 2024

First posted Nov 13, 2023

fine-tuning
Authors: Seongyun Lee, Sue Hyun Park, Yongrae JoCorresponding: not specifiedAffiliation: Korea University; KAIST AI; LG AI Research
15

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data

Cleans "hallucinatory toxicity" in existing visual instruction datasets and generates counterfactual data for fine-tuning.

data curationmodel alignment
Fine-tuning
vision-language

CVPR 2024

First posted Nov 22, 2023

data curation
Authors: Qifan Yu, Juncheng Li, Longhui WeiCorresponding: Longhui WeiAffiliation: Zhejiang University; Huawei Cloud; Institute of Computing Technology, Chinese Academy of Sciences
16

Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Constructs hallucination-aware data pairs and applies the DPO algorithm directly to update model weights.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

ICME 2024

First posted Nov 28, 2023

preference optimization
Authors: Zhiyuan Zhao, Bin Wang, Linke OuyangCorresponding: not specifiedAffiliation: Shanghai AI Laboratory
17

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Collects fine-grained, segment-level human correctional feedback for reward modeling and PPO optimization.

reinforcement learningreward modeling
Perception-grounded RL
vision-language

CVPR 2024

First posted Dec 1, 2023

rlhf
Authors: Tianyu Yu, Yuan Yao, Haoye ZhangCorresponding: not specifiedAffiliation: Tsinghua University; National University of Singapore
18

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

Trains a fine-grained visual reward model to guide the model toward better visual alignment.

visual groundingreward modeling
Perception-grounded RL
vision-language

CVPR 2024

First posted Feb 9, 2024

reward modeling
Authors: Siming Yan, Min Bai, Weifeng ChenCorresponding: not specifiedAffiliation: The University of Texas at Austin; AWS AI
19

Visually Dehallucinative Instruction Generation: Know What You Don't Know

Teaches the model to actively output "I don't know" when evidence is insufficient via specialized instruction tuning.

instruction tuningmodel alignment
Fine-tuning
vision-language

ACL 2024

First posted Feb 15, 2024

instruction tuning
Authors: Sungguk Cha, Jusung Lee, Younghyun LeeCorresponding: Sungguk Cha, Cheoljong YangAffiliation: Multimodal AI Lab., NC Research, NCSOFT Corporation
20

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Constructs a preference-based dataset to align the model with high-quality image descriptions via preference fine-tuning.

data curationpreference fine-tuningmodel alignment
Perception-grounded RL
vision-language

NeurIPS 2024

First posted Feb 18, 2024

preference fine-tuning
Authors: Yiyang Zhou, Chenhang Cui, Rafael RafailovCorresponding: Huaxiu YaoAffiliation: UNC-Chapel Hill; Stanford University
21

Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reconstructs the fine-tuning dataset to teach the model to output the End-Of-Sequence (EOS) token early to truncate generation.

instruction tuningdata curation
Fine-tuning
vision-language

ACL 2024

First posted Feb 22, 2024

instruction tuning
Authors: Zihao Yue, Liang Zhang, Qin JinCorresponding: not specifiedAffiliation: Renmin University of China
22

Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning

Introduces adversarial instruction tuning to improve model robustness against misleading user queries.

instruction tuningmodel alignment
Fine-tuning
vision-language

ACL 2024

First posted Mar 15, 2024

instruction tuning
Authors: Dongmin Park, Zhaofang Qian, Guangxing HanCorresponding: not specifiedAffiliation: KRAFTON; UC San Diego; University of Central Florida
23

Calibrated Self-Rewarding Vision Language Models

Proposes a self-calibrating reward mechanism enabling iterative self-generated feedback and preference alignment.

uncertaintyself-rewarding learningmodel alignment
Perception-grounded RL
vision-language

NeurIPS 2024

First posted May 23, 2024

self-rewarding learning
Authors: Yiyang Zhou, Zhiyuan Fan, Dongjie ChengCorresponding: not specifiedAffiliation: UNC-Chapel Hill; University of Chicago; University of Maryland; Rutgers University; Independent Researcher
24

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

Constructs a reflective instruction dataset to teach the model to perform internal visual fact-checking before outputting answers.

instruction tuningdata curation
Fine-tuning
vision-language

ECCV 2024

First posted Jul 16, 2024

instruction tuning
Authors: Jinrui Zhang, Teng Wang, Haigang ZhangCorresponding: Feng ZhengAffiliation: Southern University of Science and Technology; The University of Hong Kong; Shenzhen Polytechnic University; The Cloud Computing and IT Institute of ZTE Corporation; Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory, Shenzhen, China
25

Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models

Cleans and rewrites captions in CLIP pre-training data to reduce object hallucinations at the vision-language source.

object hallucinationdata curation
Fine-tuning
vision-language

EMNLP 2024

First posted Oct 4, 2024

pre-training intervention
Authors: Yufang Liu, Tao Ji, Changzhi SunCorresponding: not specifiedAffiliation: School of Computer Science and Technology, East China Normal University; School of Computer Science, Fudan University; Pazhou Laboratory, Huangpu
26

V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

Uses a vision-guided mechanism to construct high-quality positive and negative preference pairs for DPO.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

EMNLP 2024

First posted Nov 5, 2024

preference optimization
Authors: Yuxi Xie, Guanzhen Li, Xiao XuCorresponding: not specifiedAffiliation: National University of Singapore
27

Hallucination-resistant multimodal content generation through knowledge graph-based reinforcement learning

Grounds multimodal generation in structured knowledge and optimizes graph-based rewards to suppress unsupported content.

reinforcement learningknowledge-graph reinforcement learningmodel alignment
Perception-grounded RL
vision-language

Information Fusion 2026

knowledge-graph reinforcement learning
Authors: Liang Zeng, Xinyi Lin, Shanping YuCorresponding: not specifiedAffiliation:
28

Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training

Selects hallucination-oriented samples, assigns severity-specific loss weights, and modulates localized visual attention during HD-DPO.

preference optimizationattention steering
Perception-grounded RL
vision-language

AAAI 2026

severity-guided dpo
Authors: Yuanyi Xu, Xiangru Zhu, Sihang JiangCorresponding: not specifiedAffiliation: Fudan University; Renmin University of China; Alibaba Group
29

EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation

Builds temporally grounded video preferences from visual and audio evidence and improves supervision with echo-layered keyframe sampling.

preference optimizationaudiovisual preference optimizationmodel alignment
Perception-grounded RL
vision-language

AAAI 2026

audiovisual preference optimization
Authors: Shuai Liu, Da Chen, Yiheng PanCorresponding: not specifiedAffiliation: School of Software Engineering, Xi'an Jiaotong University; ByteDance; School of Cyber Science and Engineering, Xi'an Jiaotong University
30

VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models

Reweights positive words by visual relevance and adapts the negative preference loss using lexical importance differences.

preference optimizationvision-guided lexical dpomodel alignment
Perception-grounded RL
vision-language

TOMM 2026

vision-guided lexical dpo
Authors: Siyuan Li, Feng Wang, Simeng QinCorresponding: not specifiedAffiliation: School of Data Science and Intelligent Media, Communication University of China, Beijing, China; Tianjin University, Tianjin, China; Northeastern University, Shenyang, China; Alibaba Group, Hangzhou, China; Communication University of China, Beijing, China
31

Multimodal Chain-of-Thought Reasoning in Language Models

Adopts a two-stage fine-tuning framework, training the model to generate a CoT reasoning process before the answer.

multimodal reasoningfine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2023

First posted Feb 2, 2023

fine-tuning
Authors: Zhuosheng Zhang, Aston Zhang, Mu LiCorresponding: not specifiedAffiliation: School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; GenAI, Meta; Amazon Web Services; Department of Computer Science and Engineering, Shanghai Jiao Tong University
32

Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

Introduces a hallucination-augmented contrastive learning loss function during the training phase.

contrastive learningmodel alignment
Fine-tuning
vision-language

arXiv 2023

First posted Dec 12, 2023

contrastive learning
Authors: Chaoya Jiang, Haiyang Xu, Mengfan DongCorresponding: not specifiedAffiliation: National Engineering Research Center for Software Engineering, Peking University; Alibaba Group
33

KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

Jointly trains a language model, a vision encoder, and a Graph Neural Network (GNN) to integrate knowledge graphs for reasoning.

external toolsmultimodal reasoning
Fine-tuning
vision-language

arXiv 2024

First posted Jan 23, 2024

joint training
Authors: Debjyoti Mondal, Suraj Modi, Subhadarshi PandaCorresponding: not specifiedAffiliation: Samsung R&D Institute India - Bangalore
34

Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training

Proposes a novel vision-language pre-training (VLP) objective to suppress hallucinations.

pre-trainingmodel alignment
Fine-tuning
vision-language

arXiv 2023

First posted Oct 14, 2022

pre-training
Authors: Wenliang Dai, Zihan Liu, Ziwei JiCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology; NVIDIA
35

Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites

Uses ChatGPT to rewrite image captions for data construction and performs two-stage fine-tuning.

fine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2023

First posted Dec 4, 2023

fine-tuning
Authors: Lei Wang, Jiabang He, Shenshen LiCorresponding: not specifiedAffiliation: Singapore Management University, Beijing Forestry University; University of Electronic Science and Technology of China
36

Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

Proposes hallucination-aware DPO, utilizing fine-grained AI feedback to construct positive and negative pairs for weight updates.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Apr 22, 2024

preference optimization
Authors: Wenyi Xiao, Ziwei Huang, Leilei GanCorresponding: not specifiedAffiliation: Zhejiang University; Alibaba Group; Fudan University
37

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Utilizes fine-grained feedback from open-source AIs to build a preference dataset and updates weights using alignment algorithms.

reinforcement learningdata curation
Perception-grounded RL
vision-language

arXiv 2024

First posted May 27, 2024

rlaif
Authors: Tianyu Yu, Haoye Zhang, Qiming LiCorresponding: not specifiedAffiliation: Department of Computer Science and Technology, Tsinghua University; NExT++ Lab, School of Computing, National University of Singapore
38

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Proposes a conditional preference optimization algorithm to reduce hallucinations without degrading general capabilities.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Jun 17, 2024

preference optimization
Authors: Fei Wang, Wenxuan Zhou, James Y. HuangCorresponding: not specifiedAffiliation: University of Southern California; University of California, Davis; Microsoft Research
39

CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs

Uses a pre-trained CLIP model directly as an AI judge to generate preference pairs for DPO training.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Aug 19, 2024

preference optimization
Authors: Yassine Ouali, Adrian Bulat, Brais MartinezCorresponding: not specifiedAffiliation: Samsung AI Center Cambridge, UK; Technical University of Ia \textcommabelow si, Romania; Queen Mary University of London, UK
40

EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models

Proposes an efficient fine-grained unlearning framework to "forget" the tendency to hallucinate via reverse parameter updates.

unlearningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Feb 15, 2024

unlearning
Authors: Shangyu Xing, Fei Zhao, Zhen WuCorresponding: not specifiedAffiliation: National Key Laboratory for Novel Software Technology, Nanjing University, China
41

Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization

Introduces a hallucination-inducing mechanism during training to construct sample pairs for parameter optimization.

contrastive tuningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted May 24, 2024

contrastive tuning
Authors: Xinyu Lyu, Beitao Chen, Lianli GaoCorresponding: not specifiedAffiliation: Center for Future Media, University of Electronic Science and Technology of China; Shenzhen Institute for Advanced Study, UESTC
42

Mitigating Open-Vocabulary Caption Hallucinations

Trains a reward model based on Natural Language Inference (NLI) and uses RL to fine-tune the MLLM.

reinforcement learningreward modeling
Perception-grounded RL
vision-language

arXiv 2023

First posted Dec 6, 2023

rlhf
Authors: Assaf Ben-Kish, Moran Yanuka, Morris AlperCorresponding: not specifiedAffiliation: Tel-Aviv University
43

Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning

Constructs targeted repair instruction datasets for different hallucination types to perform targeted fine-tuning.

data curationfine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Apr 16, 2024

fine-tuning
Authors: Rui Hu, Yahan Tu, Shuyu WeiCorresponding: not specifiedAffiliation: Beijing Key Lab of Traffic Data Analysis and Mining; Beijing Jiaotong University, China
44

See or Guess: Counterfactually Regularized Image Captioning

Introduces a counterfactual regularization term during training to force distinction between seen visuals and guessed priors.

regularization trainingmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Aug 29, 2024

regularization training
Authors: Qian Cao, Xu Chen, Ruihua SongCorresponding: not specifiedAffiliation: Gaoling School of Artificial Intelligence, Renmin University of China; Tencent AI Lab
45

Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization

Generates targeted preference data specifically for hallucinations and applies DPO to penalize hallucination-prone features.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Nov 15, 2024

preference optimization
Authors: Yuhan Fu, Ruobing Xie, Xingwu SunCorresponding: not specifiedAffiliation: Key Lab of DEKE, Renmin University of China; Machine Learning Platform Department, Tencent
46

Silkie: Preference Distillation for Large Visual Language Models

Uses AI to construct the VLFeedback dataset and distills preferences into the model via DPO.

preference optimizationdata curation
Perception-grounded RL
vision-language

arXiv 2023

First posted Dec 17, 2023

preference distillation
Authors: Lei Li, Zhihui Xie, Mukai LiCorresponding: not specifiedAffiliation: The University of Hong Kong; The Chinese University of Hong Kong (Shenzhen); Peking University
47

Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Constructs hard negative samples via data augmentation and introduces a contrastive loss to strengthen visual grounding.

visual groundingcontrastive tuningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted May 28, 2024

contrastive tuning
Authors: Pritam Sarkar, Sayna Ebrahimi, Ali EtemadCorresponding: not specifiedAffiliation: Queen's University and Vector Institute; Google DeepMind; Queen's University; Google Cloud AI Research
48

Mitigating Hallucination in Visual Language Models with Visual Supervision

Constructs a fine-grained RAI-30k dataset and integrates the SAM model during instruction tuning.

instruction tuningdata curation
Fine-tuning
vision-language

arXiv 2023

First posted Nov 27, 2023

instruction tuning
Authors: Zhiyang Chen, Yousong Zhu, Yufei ZhanCorresponding: not specifiedAffiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Peng Cheng Laboratory; Wuhan AI Research
49

Mitigating Multilingual Hallucination in Large Vision-Language Models

Builds a correctional dataset specifically for multilingual scenarios to mitigate cross-lingual hallucinations caused by language bias.

instruction tuningdata curation
Fine-tuning
vision-language

arXiv 2024

First posted Aug 1, 2024

instruction tuning
Authors: Xiaoye Qu, Mingyang Song, Wei WeiCorresponding: Wei WeiAffiliation: School of Computer Science & Technology, Huazhong University of Science and Technology; School of Computing Science and Technology, Fudan University; Wangxuan Institute of Computer Technology, Peking University; School of Computing Science and Technology, Zhejiang Gongshang University; Department of Computer Science and Engineering, The Chinese University of Hong Kong
50

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

Proposes a modality-fair preference optimization algorithm to balance penalties and prevent language prior dominance.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Oct 20, 2024

preference optimization
Authors: Songtao Jiang, Yan Zhang, Ruizhe ChenCorresponding: not specifiedAffiliation: Zhejiang University; National University of Singapore
51

Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation

Proposes a token-level preference optimization strategy based on self-calibrated visual-anchored rewards.

preference optimizationuncertainty
Perception-grounded RL
vision-language

arXiv 2024

First posted Dec 19, 2024

preference optimization
Authors: Jihao Gu, Yingyao Wang, Meng CaoCorresponding: not specifiedAffiliation: Alibaba Group; Mohamed bin Zayed University of Artificial Intelligence
52

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Introduces Retrieval-Augmented Generation (RAG) into preference data construction to guide DPO training.

preference optimizationretrieval
Perception-grounded RL
vision-language

arXiv 2025

First posted Feb 18, 2025

preference optimization
Authors: Shuo Xing, Peiran Li, Yuping WangCorresponding: not specifiedAffiliation: Texas A&M University; University of Michigan; UIUC; UNC Chapel Hill
53

Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images

Proposes symmetrical visual contrastive optimization using minimal contrastive images for alignment fine-tuning.

contrastive optimizationmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Feb 19, 2025

contrastive optimization
Authors: Shengguang Wu, Fan-Yun Sun, Kaiyue WenCorresponding: not specifiedAffiliation: Stanford University
54

FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback

Uses fine-grained AI feedback to construct preference data and applies DPO for alignment training.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Apr 7, 2024

preference optimization
Authors: Liqiang Jing, Xinya DuCorresponding: not specifiedAffiliation: The University of Texas at Dallas
55

TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Massively scales up text-centric visual instruction tuning data to specifically address text hallucinations in OCR tasks.

instruction tuningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Apr 19, 2024

instruction tuning
Authors: Jingqun Tang, Chunhui Lin, Zhen ZhaoCorresponding: not specifiedAffiliation: ByteDance; East China Normal University; Huazhong University of Science and Technology
56

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Systematically builds datasets covering various alignment strategies to empirically verify their impact on hallucinations.

verificationdata curation
Fine-tuning
vision-language

arXiv 2024

First posted Jul 2, 2024

fine-tuning
Authors: Elmira Amirloo, Jean-Philippe Fauconnier, Christoph RoesmannCorresponding: not specifiedAffiliation: Apple Inc
57

RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data

Optimizes a visual reward model using auxiliary text-only preference data during RLHF to enhance hallucination discrimination.

reinforcement learningreward modeling
Perception-grounded RL
vision-language

arXiv 2024

First posted Aug 22, 2024

reward modeling
Authors: Chenglong Wang, Yang Gan, Yifu HuoCorresponding: Chunliang ZhangAffiliation: School of Computer Science and Engineering, Northeastern University, Shenyang, China; NiuTrans Research, Shenyang, China; CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China
58

HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding

Proposes a fine-tuning framework combining vision-enhanced penalty decoding with hierarchical feedback learning for behavior alignment.

feedback learningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Sep 30, 2024

feedback learning
Authors: Fan Yuan, Chi Qin, Xiaogang XuCorresponding: not specifiedAffiliation: College of Artificial Intelligence; Nanjing University of Aeronautics and Astronautics, Nanjing, China; MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China; The Chinese University of Hong Kong, Hong Kong, China
59

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Proposes entity-centric multimodal preference optimization, focusing penalty granularity precisely on specific visual entity tokens.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Jun 4, 2025

preference optimization
Authors: Jiulong Wu, Zhengliang Shi, Shuaiqiang WangCorresponding: not specifiedAffiliation: Soochow University, Suzhou, China; Baidu Inc., Beijing, China; Shandong University, Qingdao, China
60

Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales

Introduces faithful and concise reasoning rationales as supervision signals during training to enhance logical generation.

multimodal reasoningrationale trainingmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Apr 17, 2024

rationale training
Authors: Minghe Gao, Shuang Chen, Liang PangCorresponding: not specifiedAffiliation: Zhejiang University; Chinese Academy of Sciences; National University of Singapore; Sun Yat-sen University
61

Generating Faithful and Salient Text from Multimodal Data

Constructs a multimodal factual consistency dataset and fine-tunes the model to improve faithfulness in data-to-text generation.

data curationfine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Sep 6, 2024

fine-tuning
Authors: Tahsina Hashem, Weiqing Wang, Derry Tanti WijayaCorresponding: not specifiedAffiliation: Department of Data Science & AI, Monash University, Australia; Department of Data Science, Monash University, Indonesia; Department of CSE, Bangladesh University of Engineering and Technology, Bangladesh
62

Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment

Formulates preference modeling as a next-token prediction task to update the model using fine-grained verifier data.

verificationpreference modelingmodel alignment
Perception-grounded RL
vision-language

arXiv 2024

First posted Oct 18, 2024

preference modeling
Authors: Chenhang Cui, An Zhang, Yiyang ZhouCorresponding: An ZhangAffiliation: National University of Singapore; UNC-Chapel Hill; Chicago University; Nanyang Technological University
63

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

Constructs a dataset targeting verb concept hallucinations and applies specific instruction tuning to correct action biases.

instruction tuningdata curation
Fine-tuning
vision-language

arXiv 2024

First posted Dec 6, 2024

instruction tuning
Authors: Zehao Wang, Xinpeng Liu, Yudonglin ZhangCorresponding: not specifiedAffiliation: Shanghai Jiao Tong University; ARC Lab, Tencent PCG
64

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution

Automatically generates and filters high-quality alignment data through an iterative self-evolution mechanism.

data curationself-evolution learningmodel alignment
Fine-tuning
vision-language

arXiv 2024

First posted Dec 20, 2024

self-evolution learning
Authors: Wentao Tan, Qiong Cao, Yibing ZhanCorresponding: not specifiedAffiliation: South China University of Technology; JD Explore Academy, Beijing; Pazhou Lab, Guangzhou
65

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs

Proposes cross-modal hierarchical DPO to penalize global image-text matching and local object features at different levels.

preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Jan 28, 2025

preference optimization
Authors: Jinlan Fu, Shenzhen Huangfu, Hao FeiCorresponding: not specifiedAffiliation: National University of Singapore; Fudan University; Digital Twin Institute, Eastern Institute of Technology, Ningbo
66

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

Introduces subtle visual perturbations during training forward passes to force the learning of robust visual representations.

representation editingvisual perturbation trainingmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Mar 9, 2025

visual perturbation training
Authors: Cong Chen, Mingyu Liu, Chenchen JingCorresponding: not specifiedAffiliation: Zhejiang University; WeChat Group; Zhejiang University of Technology
67

Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy

Builds a low-level visual hallucination database and applies negative sampling to enhance awareness of low-level features.

negative sample fine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Mar 26, 2025

negative sample fine-tuning
Authors: Yinan Sun, Xiongkuo Min, Zicheng ZhangCorresponding: Xiongkuo MinAffiliation:
68

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning

Adopts rationale-augmented instruction tuning to teach the model to generate critique logic before providing an answer.

instruction tuningmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted May 12, 2025

instruction tuning
Authors: Zexian Yang, Dian Li, Dayan WuCorresponding: Dayan Wu, Gang LiuAffiliation: Institute of Information Engineering, Chinese Academy of Sciences; Foundation Technology Center, Tencent PCG
69

OViP: Online Vision-Language Preference Learning for VLM Hallucination

Proposes an online preference learning framework to dynamically generate preference pairs during training.

preference learningmodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted May 21, 2025

preference learning
Authors: Shujun Liu, Siyuan Wang, Zejun LiCorresponding: not specifiedAffiliation: Fudan University; University of Southern California; ByteDance
70

BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models

Employs a bijective maximum likelihood learning approach to suppress hallucinations by optimizing joint probability distributions.

maximum likelihood learningmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted May 30, 2025

maximum likelihood learning
Authors: Huu-Thien Tran, Thanh-Dat Truong, Khoa LuuCorresponding: not specifiedAffiliation: CVIU Lab, University of Arkansas
71

Stop learning it all to mitigate visual hallucination, Focus on the hallucination target

Stops learning redundant backgrounds and forces weight updates to focus exclusively on target areas causing hallucinations.

preference optimizationtarget-localized dpomodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Jun 13, 2025

target-localized dpo
Authors: Dokyoon Yoon, Youngsook Song, Woomyong ParkCorresponding: not specifiedAffiliation: SIONIC AI
72

Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation

Identifies and purges hallucinated samples from fine-tuning data, using the refined high-quality data for knowledge distillation.

knowledge distillationmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Jul 7, 2025

knowledge distillation
Authors: Wenhao Li, Xiu Su, Jingyi WuCorresponding: not specifiedAffiliation: University of Sydney; Central South University; Fudan University; Southeast University; HKUST; Sensetime Research
73

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Proposes a closed-loop framework where the model generates predictions, performs backward verification, and uses errors as training signals.

verificationclosed-loop trainingmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Jul 7, 2025

closed-loop training
Authors: Jianjiang Yang, Yanshu li, Ziyan HuangCorresponding: not specifiedAffiliation: Department of Computer Science, University of Bristol; School of Future Technology, South China University of Technology; Brown University
74

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

Analyzes object biases introduced during pre-training and debiases the model using a reverse objective function during fine-tuning.

unlearningbias unlearningmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Aug 6, 2025

bias unlearning
Authors: Yifan Li, Kun Zhou, Wayne Xin ZhaoCorresponding: not specifiedAffiliation: Gaoling School of Artificial Intelligence, Renmin University of China; University of California, San Diego; DataCanvas Alaya NeW
75

TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs

Adopts an adaptive MinMax Token preference strategy to dynamically adjust reward allocation.

adaptive preference strategymodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Jul 29, 2025

adaptive preference strategy
Authors: Kejia Zhang, Keda Tao, Zhiming LuoCorresponding: not specifiedAffiliation: Xiamen University; Westlake University; DAMO Academy, Alibaba Group; AWS AI Lab, Amazon; Hupan Laboratory
76

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

Integrates the CHAIR metric, utilizing object-level matching scores as the preference reward for DPO.

preference optimizationobject-aware dpomodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Aug 27, 2025

object-aware dpo
Authors: Alberto Compagnoni, Davide Caffagni, Nicholas MoratelliCorresponding: not specifiedAffiliation: University of Modena and Reggio Emilia, Modena, Italy
77

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Distinguishes "omission" and "fabrication" hallucination causes and constructs respective fine-tuning data for decoupled training.

decoupled fine-tuningmodel alignment
Fine-tuning
vision-language

arXiv 2025

First posted Aug 30, 2025

decoupled fine-tuning
Authors: Guangzong Si, Hao Yin, Xianfei LiCorresponding: not specifiedAffiliation: USTC; CowaRobot
78

Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

Proposes semantic curriculum preference optimization, training the model in increasing difficulty to smoothly eliminate hallucinations.

preference optimizationcurriculum preference optimizationmodel alignment
Perception-grounded RL
vision-language

arXiv 2025

First posted Sep 29, 2025

curriculum preference optimization
Authors: Yuanshuai Li, Yuping Yan, Junfeng TangCorresponding: not specifiedAffiliation: School of Engineering, Westlake University, Hangzhou, China; School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China
79

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

Proposes a universal training framework across multiple alignment formats (e.g., DPO, PPO) to eliminate vision-language biases.

preference optimizationreinforcement learning
Perception-grounded RL
vision-language

arXiv 2025

First posted Nov 21, 2025

unified alignment framework
Authors: Jiaye Qian, Ge Zheng, Yuchen ZhuCorresponding: Sibei YangAffiliation: School of Computer Science and Engineering, Sun Yat-sen University; ShanghaiTech University
80

OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Extracts self-attention matrices to find and penalize attention flows overly reliant on summary tokens.

attention steeringattention penaltyinternal intervention
Activation Editing
vision-language

CVPR 2024

First posted Nov 29, 2023

attention penalty
Authors: Qidong Huang, Xiaoyi Dong, Pan ZhangCorresponding: not specifiedAffiliation: University of Science and Technology of China; Shanghai AI Laboratory
81

Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Fuses global and local attention feature matrices internally to enhance fine-grained perception.

attention steeringattention fusioninternal intervention
Activation Editing
vision-language

CVPR 2025

First posted Jun 18, 2024

attention fusion
Authors: Wenbin An, Feng Tian, Sicong LengCorresponding: not specifiedAffiliation: Xi'an Jiaotong University; Nanyang Technological University; Lenovo Research; SGIT AI Lab; University of Massachusetts Boston
82

Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs

Adjusts cross-modal attention by forcefully amplifying the weights of visual tokens.

attention steeringattention amplificationinternal intervention
Activation Editing
vision-language

ECCV 2024

First posted Jul 31, 2024

attention amplification
Authors: Shi Liu, Kecheng Zheng, Wei ChenCorresponding: not specifiedAffiliation: State Key Lab of CAD&CG, Zhejiang University
83

DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination

Dives into the MLLM to mask specific attention connections overly dominated by text.

attention steeringattention maskinginternal intervention
Activation Editing
vision-language

EMNLP 2024

First posted Oct 6, 2024

attention masking
Authors: Xuan Gong, Tianshi Ming, Xinpeng WangCorresponding: not specifiedAffiliation: Department of Computer Science and Technology, Tongji University
84

Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

Uses causal analysis to block specific causal attention paths causing hallucinations during inference.

attention steeringcausal attention disconnectioninternal intervention
Activation Editing
vision-language

ICLR 2025

First posted Oct 7, 2024

causal attention disconnection
Authors: Guanyu Zhou, Yibo Yan, Xin ZouCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Tsinghua University
85

Mitigating Object Hallucination via Concentric Causal Attention

Strengthens the causal association between local visual features and global semantics in attention matrices.

attention steeringconcentric causal attentioninternal intervention
Activation Editing
vision-language

NIPS 2024

First posted Oct 21, 2024

concentric causal attention
Authors: Yun Xing, Yiheng Li, Ivan LaptevCorresponding: not specifiedAffiliation: Nanyang Technological University; MBZUAI
86

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Projects hidden states into a hallucination space and strips them via orthogonalization to clear hallucinations.

representation editinglatent space orthogonalizationinternal intervention
Activation Editing
vision-language

CVPR 2025

First posted Dec 18, 2024

latent space orthogonalization
Authors: Le Yang, Ziwei Zheng, Boxu ChenCorresponding: not specifiedAffiliation: Xi'an Jiaotong University, Xi'an 710049, China
87

VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

Introduces visual-aware sparsification to directly drop irrelevant visual tokens prone to inducing hallucinations.

sparsification droppinginternal intervention
Activation Editing
vision-language

CVPR 2025

First posted Jan 11, 2025

sparsification dropping
Authors: Xianwei Zhuang, Zhihong Zhu, Yuxin XieCorresponding: not specifiedAffiliation: SECE of Peking University
88

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

Adaptively constrains inter-layer information flow to block false causal paths causing object hallucinations.

object hallucinationinformation flow constraintinternal intervention
Activation Editing
vision-language

AAAI 2025

First posted Feb 28, 2025

information flow constraint
Authors: Jiaqi Bai, Hongcheng Guo, Zhongyuan PengCorresponding: Mohan LiAffiliation: Cyberspace Institute of Advanced Technology, Guangzhou University, China; Huangpu Research School of Guangzhou University, China; CCSE, Beihang University, China; University of the Chinese Academy of Sciences, China
89

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

Amplifies low-level visual signals in multimodal layers to combat dilution by deep language features.

signal enhancementinternal intervention
Activation Editing
vision-language

CVPR 2025

First posted Mar 17, 2025

signal enhancement
Authors: Hao Yin, Guangzong Si, Zilei WangCorresponding: not specifiedAffiliation: University of Science and Technology of China
90

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Reconstructs multimodal attention using a causal matrix based on Manhattan distance.

attention steeringmanhattan causal replacementinternal intervention
Activation Editing
vision-language

ACMMM 2025

First posted Jul 12, 2025

manhattan causal replacement
Authors: Qiyan Zhao, Xiaofeng Zhang, Yiheng LiCorresponding: not specifiedAffiliation: FKLPRIU, Xiamen University of Technology, China; Shanghai Jiao Tong University, China; Nanyang Technological University, Singapore; Jilin University, China; Monash University, Australia; Zhejiang University, China; Huizhou University, China; Chinese Academy of Sciences, China
91

Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models

Intervenes from the visual processing architecture to fundamentally alter how the model acquires visual information.

pure vision reinforcementinternal intervention
Activation Editing
vision-language

EMNLP 2025

First posted Sep 17, 2025

pure vision reinforcement
Authors: Weihang Wang, Xinhao Li, Ziyue WangCorresponding: not specifiedAffiliation: Bilibili; UESTC; University of Virginia
92

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Maps hallucination activations toward trusted attention regions with mixture-Gaussian bridges.

attention steeringrepresentation editing
Activation Editing
vision-language

WACV 2026

First posted Feb 10, 2026

attention manifold alignment
Authors: Ziqiang Shi, Rujie Liu, Shanshan YuCorresponding: not specifiedAffiliation: Fujitsu Research & Development Center Co.,LTD., Beijing, China
93

RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Predicts input-specific latent steering vectors with gradient correction in a single forward pass.

attention steeringrectified-flow interventioninternal intervention
Activation Editing
vision-language

AAAI 2026

rectified-flow intervention
Authors: Junyu Cheng, Zhibiao Liang, Yidong ChenCorresponding: not specifiedAffiliation: Department of Artificial Intelligence, School of Informatics, Xiamen University, China; School of Computer Science, South China Normal University, China; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China
94

Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models

Filters intermediate attention maps to isolate dominant phantom text tokens and strengthen visual anchor tokens.

attention steeringdata curation
Activation Editing
vision-language

AAAI 2026

token-asymmetric filtering
Authors: Shuyi Ouyang, Hongyi Wang, Gongfan FangCorresponding: not specifiedAffiliation: Zhejiang University; National University of Singapore
95

Look Before You Speak: Visual Re-Focusing for Hallucination Mitigation in Large Vision-Language Models

Reorients the model toward image evidence before response generation to reduce visually unsupported claims.

visual re-focusinginternal intervention
Activation Editing
vision-language

LNCS 2026

visual re-focusing
Authors: Zheyuan Zhang, Yingmin Liu, Yang ShuCorresponding: Yang ShuAffiliation: Northwestern Polytechnical University, Xi'an, China; Zhejiang University, Hangzhou, China
96

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

Uses an attention lens to locate hallucinated features in middle layers and masks them for correction.

attention steeringmiddle layer maskinginternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Nov 23, 2024

middle layer masking
Authors: Zhangqi Jiang, Junkai Chen, Beier ZhuCorresponding: not specifiedAffiliation: National University of Defense Technology; Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and; Its Interdisciplinary Applications (Southeast University), Ministry of Education, China; Nanyang Technological University
97

Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence

Quantifies perception divergence and dynamically masks specific heads that lose visual alignment.

visual groundingattention steering
Activation Editing
vision-language

arXiv 2024

First posted Dec 18, 2024

attention head pruning
Authors: Jinghan He, Kuan Zhu, Haiyun GuoCorresponding: not specifiedAffiliation: Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; University of Science and Technology of China; Southeast University; National University of Singapore
98

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Reallocates Attention Maps during inference to pull misplaced attention back to visual targets.

attention steeringattention reallocationinternal intervention
Activation Editing
vision-language

arXiv 2025

First posted Mar 11, 2025

attention reallocation
Authors: Chongjun Tu, Peng Ye, Dongzhan ZhouCorresponding: Tao ChenAffiliation: Fudan University; The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory; StepFun
99

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Proposes cross-level trusted intervention to block the drift of object representations toward false language priors.

representation editingcross-level blockinginternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Nov 22, 2024

cross-level blocking
Authors: Junzhe Chen, Tianshu Zhang, Shiyu HuangCorresponding: not specifiedAffiliation: Tsinghua University; The Hong Kong University of Science and Technology (Guangzhou); Zhipu AI; Chongqing University; Shanghai Jiao Tong University.
100

Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis

Implements causal intervention to cut off spurious feature maps hidden in attention matrices.

attention steeringcausal attention interventioninternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Dec 4, 2024

causal attention intervention
Authors: Po-Hsuan Huang, Jeng-Lin Li, Chin-Po ChenCorresponding: not specifiedAffiliation: Inventec Corporation, No. 66, Hougang St., Shihlin Dist., Taipei City; University at Albany - SUNY, Albany, New York
101

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Locates the evolution of object features and corrects visual tokens via early layer interventions.

early layer interventioninternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Oct 9, 2024

early layer intervention
Authors: Yuying Shang, Xinyi Zeng, Yutao ZhuCorresponding: not specifiedAffiliation: Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing, China; Gaoling School of Artificial Intelligence, Renmin University of China; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University, Beijing, China; Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University; Kuaishou Technology Inc., Beijing, China
102

Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Accurately locates key attention heads and specifically enhances or suppresses them during forward propagation.

attention steeringhead interventioninternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Nov 15, 2024

head intervention
Authors: Xiaofeng Zhang, Yihao Quan, Chaochen GuCorresponding: not specifiedAffiliation: Shanghai Jiao Tong University; Alibaba Group; Beijing Jiaotong University
103

Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Calibrates the weight allocation of attention heads to weaken the implicit coverage of visual inputs by language priors.

uncertaintyattention steering
Activation Editing
vision-language

arXiv 2024

First posted May 28, 2024

attention calibration
Authors: Sangmin Woo, Donguk Kim, Jaehyuk JangCorresponding: not specifiedAffiliation: KAIST
104

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

Introduces dual-attention mechanisms to internally rebalance language priors and image feature flows.

attention steeringdual-attention balancinginternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Nov 21, 2024

dual-attention balancing
Authors: Haozhe Zhao, Shuzheng Si, Liang ChenCorresponding: Baobao ChangAffiliation: University of Illinois Urbana-Champaign; Peking University; Tsinghua University
105

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

Cuts off implicit semantic biases introduced during cross-lingual transitions inside the model.

attention steeringcross-lingual attention interventioninternal intervention
Activation Editing
vision-language

arXiv 2025

First posted Jun 3, 2025

cross-lingual attention intervention
Authors: Zekai Ye, Qiming Li, Xiaocheng FengCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory; Central South University; Huawei Technologies Co., Ltd
106

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

Pinpoints tokens sourcing hallucinations and excises them during forward propagation.

token localization & pruninginternal intervention
Activation Editing
vision-language

arXiv 2025

First posted Jun 12, 2025

token localization & pruning
Authors: Eunkyu Park, Minyeong Kim, Gunhee KimCorresponding: not specifiedAffiliation: Seoul National University
107

Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Identifies and edits vision-language feature vectors representing hallucinations in intermediate layers.

representation editinginternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Oct 3, 2024

representation editing
Authors: Nick Jiang, Anish Kachinthaya, Suzie PetrykCorresponding: not specifiedAffiliation: University of California, Berkeley
108

Reducing Hallucinations in Vision-Language Models via Latent Space Steering

Injects pre-trained steering vectors into the latent space to guide generation away from hallucinations.

attention steeringlatent space steeringinternal intervention
Activation Editing
vision-language

arXiv 2024

First posted Oct 21, 2024

latent space steering
Authors: Sheng Liu, Haotian Ye, Lei XingCorresponding: not specifiedAffiliation: Stanford University
109

Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Penalizes decoding by contrasting the logit output distributions of original and distorted images during inference.

contrastive decodinggeneration control
Decoding-time
vision-language

CVPR 2024

First posted Nov 28, 2023

contrastive decoding
Authors: Sicong Leng, Hang Zhang, Guanzheng ChenCorresponding: not specifiedAffiliation: DAMO Academy, Alibaba Group; Nanyang Technological University; Hupan Lab, 310023, Hangzhou, China
110

Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Introduces Classifier-Free Guidance (CFG) during inference to amplify the influence of visual features in the generation process.

inference guidancegeneration control
Decoding-time
vision-language

CVPR 2024

First posted Feb 13, 2024

inference guidance
Authors: Linxi Zhao, Yihe Deng, Weitong ZhangCorresponding: Quanquan GuAffiliation: Department of Computer Science, Cornell University, Ithaca, NY, USA; Department of Computer Science, University of California, Los Angeles, CA, USA; School of Data Science and Society, UNC, Chapel Hill, NC, USA
111

IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Penalizes pure language priors by contrasting the logits of inputs with and without images.

image-biased decodinggeneration control
Decoding-time
vision-language

ECCV 2024

First posted Feb 28, 2024

image-biased decoding
Authors: Lanyun Zhu, Deyi Ji, Tianrun ChenCorresponding: not specifiedAffiliation: Singapore University of Technology and Design; Alibaba Group; Zhejiang University
112

HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Proposes adaptive focal-contrast decoding to fine-grainedly suppress hallucinations at the local token level.

contrastive decodingfocal contrastive decodinggeneration control
Decoding-time
vision-language

ICML 2024

First posted Mar 1, 2024

focal contrastive decoding
Authors: Zhaorun Chen, Zhuokai Zhao, Hongyin LuoCorresponding: Zhaorun Chen, Zhuokai Zhao, Bo Li, Jiawei ZhouAffiliation: University of Chicago, Chicago IL, USA; University of Illinois at Urbana-Champaign, Champaign IL, USA; Massachusetts Institute of Technology, Boston MA, USA; UNC-Chapel Hill, Chapel Hill NC, USA; Toyota Technological Institute at Chicago, Chicago IL, USA
113

Multi-Modal Hallucination Control by Visual Information Grounding

Introduces a visual grounding mechanism during decoding to correct biases by contrasting different information streams.

visual groundingvisual information contrastgeneration control
Decoding-time
vision-language

CVPR 2024

First posted Mar 20, 2024

visual information contrast
Authors: Alessandro Favero, Luca Zancato, Matthew TragerCorresponding: not specifiedAffiliation: AWS AI Labs
114

Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Constructs positive and negative instructions, contrasting their logits during inference to guide decoding.

contrastive decodinginstruction contrastive decodinggeneration control
Decoding-time
vision-language

ACL 2024

First posted Mar 27, 2024

instruction contrastive decoding
Authors: Xintong Wang, Jingheng Pan, Liang DingCorresponding: not specifiedAffiliation: Department of Informatics, Universität Hamburg; The University of Sydney
115

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

Introduces residual visual connections to add early visual features directly into subsequent decoding computations.

residual decodinggeneration control
Decoding-time
vision-language

ACL 2024

First posted Jun 30, 2024

residual decoding
Authors: Weihong Zhong, Xiaocheng Feng, Liang ZhaoCorresponding: not specifiedAffiliation: Harbin Institute of Technology; Peng Cheng Laboratory
116

ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models

Uses attention visualization to locate hallucination regions and constructs interference images for contrastive decoding.

contrastive decodingattention steering
Decoding-time
vision-language

AAAI 2025

First posted Aug 25, 2024

visualized contrastive decoding
Authors: Yeji Park, Deokyeong Lee, Junsuk ChoeCorresponding: Junsuk Choe, Buru ChangAffiliation: Sogang University
117

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Introduces a generative feedback loop to detect and correct the probability distribution of the current token in real time.

self-correctionself-correcting decodinggeneration control
Decoding-time
vision-language

ICLR 2025

First posted Feb 10, 2025

self-correcting decoding
Authors: Ce Zhang, Zifu Wan, Zhehan KanCorresponding: not specifiedAffiliation: School of Computer Science, Carnegie Mellon University
118

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Adaptively adjusts the contrastive penalty based on image complexity and local token features.

contrastive decodingdynamic contrastive decodinggeneration control
Decoding-time
vision-language

CVPR 2025

First posted Mar 1, 2025

dynamic contrastive decoding
Authors: Wei Suo, Lijun Zhang, Mengyang SunCorresponding: Yanning ZhangAffiliation: School of Computer Science and Ningbo Institute, Northwestern Polytechnical University,China.; School of Cybersecurity, Northwestern Polytechnical University, China.; Computer Science, Swansea University.
119

Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models

Adaptively switches between different sampling algorithms based on attention strength.

attention steeringmixture of decodinggeneration control
Decoding-time
vision-language

ACL 2025

First posted May 17, 2025

mixture of decoding
Authors: Xinlong Chen, Yuanxing Zhang, Qiang LiuCorresponding: not specifiedAffiliation: New Laboratory of Pattern Recognition (NLPR); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences; Kuaishou Technology; Nanjing University
120

Med-VCD: Mitigating hallucination for medical large vision language models through visual contrastive decoding

Selects visually informed tokens on the fly, removing redundant tokens while retaining critical medical image context.

contrastive decodingsparse visual contrastive decodinggeneration control
Decoding-time
vision-language

Computers in Biology and Medicine 2026

First posted Dec 1, 2025

sparse visual contrastive decoding
Authors: Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman KhalooCorresponding: not specifiedAffiliation: Department of computer science, University of Central Florida, Orlando, USA; Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA; Department of electrical engineering, Columbia university, New York, NY, USA; Technical University of Applied Sciences Regensburg, Regensburg, Germany; Department of Surgery, University of Calgary, Calgary, Alberta, Canada; University College of Nabi Akram, Tabriz, Iran; School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
121

Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)

Constructs an image-free text input as a negative reference and subtracts language prior scores during logit calculation.

contrastive decodinglanguage contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Aug 6, 2024

language contrastive decoding
Authors: Avshalom Manevich, Reut TsarfatyCorresponding: not specifiedAffiliation: Bar Ilan University
122

CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models

Contrasts logit distributions based on the original image versus self-generated descriptions to penalize text-reliant tokens.

contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Jun 4, 2024

contrastive decoding
Authors: Junho Kim, Hyunjun Kim, Yeonju KimCorresponding: Junho KimAffiliation: Integrated Vision and Language Lab, KAIST
123

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Suppresses event hallucinations by contrasting output probabilities between full and truncated videos.

contrastive decodingtemporal contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Sep 25, 2024

temporal contrastive decoding
Authors: Jiacheng Zhang, Yang Jiao, Shaoxiang ChenCorresponding: not specifiedAffiliation: Shanghai Key Lab of Intelligent Information Processing, School of CS, Fudan University; Shanghai Collaborative Innovation Center on Intelligent Visual Computing; Singapore University of Technology and Design; Shanghai Academy of Artificial Intelligence for Science; Meituan
124

Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding

Introduces an external CLIP model to score and guide candidate tokens during decoding to improve image-text consistency.

visual groundingexternal tools
Decoding-time
vision-language

arXiv 2024

First posted Feb 23, 2024

external guided decoding
Authors: Ailin Deng, Zhirui Chen, Bryan HooiCorresponding: not specifiedAffiliation: School of Computing, National University of Singapore; Department of Industrial Systems Engineering and Management, National University of Singapore
125

Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding

Dynamically adjusts the logit weight distribution of text and image features during inference.

contrastive decodingre-balancing contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Sep 10, 2024

re-balancing contrastive decoding
Authors: Xiaoyu Liang, Jiayuan Yu, Lianrui MuCorresponding: Haoji HuAffiliation: College of Information Science and Electronic Engineering, Zhejiang University, China
126

Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding

Extracts internally generated facts as references for contrastive decoding to suppress over-extrapolation.

contrastive decodinginternal fact contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2025

First posted Feb 3, 2025

internal fact contrastive decoding
Authors: Chao Wang, Xuancheng Zhou, Weiwei FuCorresponding: Chao Wang, Yang ZhouAffiliation: School of Future Technology, Shanghai University, Shanghai, 200444, China; Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China; School of Mechatronic Engineering and Automation, Shanghai, 200444, China
127

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Introduces a control parameter during decoding to force switching between pure contextual description and parameterized knowledge imagination.

inference interventiongeneration control
Decoding-time
vision-language

arXiv 2023

First posted Oct 3, 2023

inference intervention
Authors: Bohan Zhai, Shijia Yang, Chenfeng XuCorresponding: not specifiedAffiliation: ByteDance Inc.; Stanford University; UC Berkeley; UIUC
128

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Introduces language contrastive decoding to weaken misleading priors from user prompts to combat sycophancy-induced hallucinations.

contrastive decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Aug 21, 2024

contrastive decoding
Authors: Yunpu Zhao, Rui Zhang, Junbin XiaoCorresponding: Rui ZhangAffiliation: School of Computer Science and Technology, University of Science and Technology of China; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences; Department of Computer Science, National University of Singapore; University of Illinois Urbana-Champaign; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences
129

Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding

Utilizes image summaries as guidance signals to intervene in output probabilities for contextual consistency.

summary-guided decodinggeneration control
Decoding-time
vision-language

arXiv 2024

First posted Oct 17, 2024

summary-guided decoding
Authors: Kyungmin Min, Minbeom Kim, Kang-il LeeCorresponding: not specifiedAffiliation: IPAI, Seoul National University; Dept. of ECE, Seoul National University
130

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

Contrasts logits from the original image and retrieved similar reference images to highlight core facts.

contrastive decodingretrieval
Decoding-time
vision-language

arXiv 2025

First posted May 26, 2025

retrieval visual contrastive decoding
Authors: Jihoon Lee, Min SongCorresponding: not specifiedAffiliation: Yonsei University; Onoma AI
131

CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models

Generates multi-granularity hierarchical feedback to dynamically correct outputs.

coarse-to-fine feedback decodinggeneration control
Decoding-time
vision-language

arXiv 2025

First posted Dec 29, 2025

coarse-to-fine feedback decoding
Authors: Zongsheng Cao, Yangfan He, Anran LiuCorresponding: Anran Liu, Zepeng WangAffiliation: Researcher; UMN; PCIE
132

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Trains an independent Reviser model specifically to detect and reconstruct hallucinated entities in outputs.

self-correctionpost-hoc revisionexternal grounding
Verification-based
vision-language

ICLR 2023

First posted Oct 1, 2023

post-hoc revision
Authors: Yiyang Zhou, Chenhang Cui, Jaehong YoonCorresponding: not specifiedAffiliation: UNC-Chapel Hill; Rutgers University; Columbia University; Stanford University
133

Woodpecker: Hallucination Correction for Multimodal Large Language Models

Calls external visual foundation models to verify entity facts, finally using an LLM to rewrite and correct.

verificationexternal tools
Verification-based
vision-language

SCIS 2024

First posted Oct 24, 2023

agent correction
Authors: Shukang Yin, Chaoyou Fu, Sirui ZhaoCorresponding: not specifiedAffiliation: School of Data Science, USTC & State Key Laboratory of Cognitive Intelligence; Tencent YouTu Lab
134

Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation

Constructs a pipeline system containing claim extraction, external detection verification, and corrected generation.

verificationexternal tools
Verification-based
vision-language

CVPR 2024

First posted Apr 30, 2024

pipeline agent
Authors: Yunhao Ge, Xiaohui Zeng, Jacob Samuel HuffmanCorresponding: not specifiedAffiliation: NVIDIA
135

Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning

Modifies reference images to expose hallucinating agents and replaces consensus-only debate with evidence-based factual verification.

verificationcounterfactual multi-agent gamingexternal grounding
Verification-based
vision-language

AAAI 2026

counterfactual multi-agent gaming
Authors: Dayong Liang, Xiao-Yong Wei, Changmeng ZhengCorresponding: not specifiedAffiliation: South China University of Technology, Guangzhou, China; Peng Cheng Laboratory, Shenzhen, China; Sichuan University, Chengdu, China; The Hong Kong Polytechnic University, Hong Kong, China
136

Mitigating Object and Relationship Hallucination in Large Vision Language Model with Multi-Agent Guidance

Coordinates specialized agents to identify and correct unsupported object and relationship claims.

verificationmulti-agent verificationexternal grounding
Verification-based
vision-language

ICASSP 2026

multi-agent verification
Authors: Soohyun Kim, Gusang Lee, Kyuhong ShimCorresponding: not specifiedAffiliation: Seoul National University, Department of Electrical and Computer Engineering, Korea; Sungkyunkwan University, Department of Computer Science and Engineering, Korea
137

CEBC: Conformal Evidence-Bounded Control for Low-Hallucination Vision-Language Generation

Calibrates an external detector on held-out data and minimally revises unsupported object mentions under explicit risk bounds.

self-correctionexternal toolsuncertainty
Verification-based
vision-language

ACL 2026

conformal evidence-bounded editing
Authors: Ashish Mishra, Tarun Kumar, Arpit ShahCorresponding: not specifiedAffiliation: Hewlett Packard Labs, Bangalore; Hewlett Packard Labs, USA
138

Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models

Uses external language models to ask questions about generated claims for self-verification.

verificationexternal tools
Verification-based
vision-language

arXiv 2024

First posted Feb 18, 2024

logical loop verification
Authors: Junfei Wu, Qiang Liu, Ding WangCorresponding: not specifiedAffiliation: Center for Research on Intelligent Perception and Computing; State Key Laboratory of Multimodal Artificial Intelligence Systems; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Nanjing University
139

Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning

Utilizes a bottom-up holistic reasoning framework combined with multi-perspective prompt verification.

verificationmultimodal reasoning
Verification-based
vision-language

arXiv 2024

First posted Dec 15, 2024

multi-perspective cross-checking
Authors: Shengqiong Wu, Hao Fei, Liangming PanCorresponding: Hao FeiAffiliation: National University of Singapore, Singapore; University of Arizona, USA; University of California, Santa Barbara, USA; Skywork AI, Singapore; Nanyang Technological University, Singapore
140

Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Uses an external detector to locate relation errors, then applies template-based secondary correction.

external toolsuncertainty
Verification-based
vision-language

arXiv 2024

First posted Aug 18, 2024

external detection calibration
Authors: Kening Zheng, Junkai Chen, Yibo YanCorresponding: Xuming HuAffiliation: Hong Kong University of Science and Technology (Guangzhou); Guangxi Zhuang Autonomous Region Big Data Research Institute; Hong Kong University of Science and Technology
141

Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination

Adopts a "Retrospect-then-Compare" multi-step prompt to make the model review visual details for correction.

multi-step retrospective promptingexternal grounding
Verification-based
vision-language

arXiv 2024

First posted Mar 21, 2024

multi-step retrospective prompting
Authors: Dingchen Yang, Bowen Cao, Guang ChenCorresponding: not specifiedAffiliation: Tongji University; Peking University
142

Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting

Explicitly adds spatial relation logic rules into the prompt to correct spatial confusion.

spatial logic rule promptingexternal grounding
Verification-based
vision-language

arXiv 2025

First posted Feb 12, 2025

spatial logic rule prompting
Authors: Jiarui Wu, Zhuo Liu, Hangfeng HeCorresponding: not specifiedAffiliation: University of Rochester
143

RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models

Applies random geometric transformations to construct multiple views for consistency verification and correction.

verificationmulti-view consistencyexternal grounding
Verification-based
vision-language

arXiv 2024

First posted May 28, 2024

multi-view consistency
Authors: Sangmin Woo, Jaehyuk Jang, Donguk KimCorresponding: not specifiedAffiliation: KAIST
144

Systematic Reward Gap Optimization for Mitigating VLM Hallucinations

The model generates a draft, then independently reviews and self-scores sub-topics for correction.

topic-level self-scoringexternal grounding
Verification-based
vision-language

arXiv 2024

First posted Nov 26, 2024

topic-level self-scoring
Authors: Lehan He, Zeren Chen, Zhelun ShiCorresponding: not specifiedAffiliation: Shanghai AI Laboratory; School of Software, Beihang University; Shanghai Innovation Institute; Tsinghua University
145

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

Deploys a network of introspective and external verification agents for interactive debate.

verificationexternal tools
Verification-based
vision-language

arXiv 2025

First posted Dec 2, 2025

cross-modal introspective debate
Authors: Zhongyu Yang, Yingfang Yuan, Xuanming JiangCorresponding: Xuanming Jiang, Wei PangAffiliation: Xi'an Jiyun Technology Co., Ltd., Xi'an, China; BCML, Heriot-Watt University, Edinburgh, UK
146

Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Evaluates model uncertainty to actively trigger knowledge base retrieval for entity correction.

retrievalexternal toolsuncertainty
Tool-augmented
vision-language

TOMM 2025

First posted Aug 1, 2024

dynamic rag supplement
Authors: Xiaoye Qu, Qiyuan Chen, Wei WeiCorresponding: not specifiedAffiliation: Huazhong University of Science and Technology; Zhejiang University; Xiamen University; Zhejiang Gongshang University
147

A Unified Hallucination Mitigation Framework for Large Vision-Language Models

Builds a unified cross-modal diagnosis pipeline utilizing external discriminators to intercept and modify outputs.

external toolsexternal discriminator systemexternal grounding
Tool-augmented
vision-language

TMLR 2024

First posted Sep 24, 2024

external discriminator system
Authors: Yue Chang, Liqiang Jing, Xiaopeng ZhangCorresponding: not specifiedAffiliation: The University of Texas at Dallas
148

Mitigating Object Hallucinations via Sentence-Level Early Intervention

Implements factual truncation prompts in early sentence stages of generation to prevent error propagation.

early sentence truncationexternal grounding
Tool-augmented
vision-language

ICCV 2025

First posted Jul 16, 2025

early sentence truncation
Authors: Shangpin Peng, Senqiao Yang, Li JiangCorresponding: not specifiedAffiliation: Harbin Institute of Technology, Shenzhen; The Chinese University of Hong Kong; The Chinese University of Hong Kong, Shenzhen
149

Exploring the Transferability of Visual Prompting for Multimodal Large Language Models

Overlays visual prompts (e.g., circles, bounding boxes) on the original image to force model attention.

attention steeringoriginal image overlayexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Apr 17, 2024

original image overlay
Authors: Yichi Zhang, Yinpeng Dong, Siyuan ZhangCorresponding: not specifiedAffiliation: Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center; THBI Lab, BNRist Center, Tsinghua University, Beijing 100084, China; RealAI; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong
150

What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models

Introduces "What if" prompting to compel the model to self-reflect on its initial judgments.

counterfactual promptingexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Mar 20, 2024

counterfactual prompting
Authors: Junho Kim, Yeon Ju Kim, Yong Man RoCorresponding: Yeon Ju KimAffiliation: Integrated Vision and Language Lab KAIST South Korea
151

From Training-Free to Adaptive: Empirical Insights into MLLMs’ Understanding of Detection Information

Feeds information identified by external object detection models directly to the MLLM as context.

external toolsexternal visual guidanceexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Jan 31, 2024

external visual guidance
Authors: Qirui Jiao, Daoyuan Chen, Yilun HuangCorresponding: not specifiedAffiliation: Sun Yat-Sen University, Shenzhen, China; Alibaba Group, Hangzhou, China
152

Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

Pre-supplements missing low-level visual perception details in the input prompt to bridge the gap.

perception detail supplementexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted May 24, 2024

perception detail supplement
Authors: Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal KumarCorresponding: not specifiedAffiliation: University of Maryland, College Park, USA; Adobe, USA
153

Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?

Strictly controls the requested detail quantity via prompt to balance richness and hallucination rate.

prompt detail constraintexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Jun 18, 2024

prompt detail constraint
Authors: Mingqian Feng, Yunlong Tang, Zeliang ZhangCorresponding: not specifiedAffiliation: University of Rochester
154

Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning

Adopts multi-view prompts and multi-path reasoning, aggregating and comparing answers.

multimodal reasoningmulti-path votingexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Aug 30, 2024

multi-path voting
Authors: Xiaoye Qu, Jiashuo Sun, Wei WeiCorresponding: not specifiedAffiliation: Huazhong University of Science and Technology; Xiamen University; The Chinese University of Hong Kong
155

Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

Uses prompts to guide the model to actively re-extract visual clues for confirmation during generation.

memory retracing confirmationexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Oct 4, 2024

memory retracing confirmation
Authors: Xin Zou, Yizhou Wang, Yibo YanCorresponding: not specifiedAffiliation: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; China University of Geosciences; University of Technology Sydney
156

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Introduces an external Vision Value Model to assist tree search, guiding autoregressive generation.

external toolsexternal scoring searchexternal grounding
Tool-augmented
vision-language

arXiv 2024

First posted Dec 4, 2024

external scoring search
Authors: Xiyao Wang, Zhengyuan Yang, Linjie LiCorresponding: not specifiedAffiliation: University of Maryland, College Park; Microsoft
157

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

Uses black-box optimization to automatically search for "visual prompt patches" that suppress hallucinations.

external toolsautomated patch searchexternal grounding
Tool-augmented
vision-language

arXiv 2025

First posted Apr 30, 2025

automated patch search
Authors: Sangmin Woo, Kang Zhou, Yun ZhouCorresponding: Sangmin Woo, Haibo DingAffiliation: Amazon AWS AI; KAIST
158

Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models

Uses uncertainty-aware contrastive calibration to suppress hallucinated actions in vision-language-action generation.

uncertaintyentropy-optimized decodinggeneration control
Uncertainty-aware
vision-language

Neurocomputing 2026

entropy-optimized decoding
Authors: Ye Qiu, Zhaoxin Fan, Qingchen YuCorresponding: not specifiedAffiliation:
159

VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models

Uses a reinforcement-learned caption model as an auxiliary visual clue source and applies image-confidence constraints during generation.

uncertaintyvisual-clue-guided decodinggeneration control
Uncertainty-aware
vision-language

AAAI 2026

visual-clue-guided decoding
Authors: Guoqing Chen, Fu Zhang, Bingqian LiuCorresponding: not specifiedAffiliation: Northeastern University
160

CTDD: Cumulative Trend Divergence Decoding for Mitigating Hallucination in Large Vision-Language Models

Tracks accumulated divergence in token-distribution trends and calibrates generation when visual and linguistic evidence separate.

uncertaintycumulative trend-divergence decodinggeneration control
Uncertainty-aware
vision-language

LNCS 2026

cumulative trend-divergence decoding
Authors: Jiani Hou, Zhixuan You, Siyu LuoCorresponding: Bing GuoAffiliation: College of Software Engineering, Sichuan University, Chengdu, China; International Economics and Trade, Central University of Finance and Economics, Beijing, China; College of Computer Science, Sichuan University, Chengdu, China
161

Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Filters hallucinated words by comparing the confidence of different generation paths via the model's own introspective mechanism.

uncertaintydata curation
Uncertainty-aware
vision-language

arXiv 2024

First posted Aug 4, 2024

self-introspective decoding
Authors: Fushuo Huo, Wenchao Xu, Zhong ZhangCorresponding: Wenchao Xu, Peilin ZhaoAffiliation: Department of Computing, The Hong Kong Polytechnic University; Division of Integrative Systems and Design, Hong Kong University of Science and Technology; Tencent AI Lab; Huazhong University of Science and Technology; Tsinghua University
162

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

ZINA detects hallucinated spans, classifies six error types, and suggests grounded edits for MLLM outputs.

fine-grained spanserror taxonomyhallucination editing
Detection
vision-language

CVPR 2026

First posted Jun 16, 2025

span-level detection
Authors: Yuiga Wada, Kazuki Matsuda, Komei SugiuraCorresponding: not specifiedAffiliation: Keio AI Research Center; Keio University; Carnegie Mellon University
163

Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

VBackChecker grounds response claims backward to image pixels with rich contextual evidence.

backward visual groundingrich-contextpixel-level grounding
Detection
vision-language

AAAI 2026

First posted Nov 15, 2025

pixel-level grounding
Authors: Pinxue Guo, Chongruo Wu, Xinyu ZhouCorresponding: Wei Zhang, Wenqiang ZhangAffiliation: Fudan University; Independent Researcher
164

MHALO: Evaluating MLLMs as Fine-grained Hallucination Detectors

MHALO evaluates MLLMs on hallucination recognition and token-level localization with fine-grained metrics.

token-level detectionmeta-evaluation benchmarkMLLM as judge
Detection
vision-language

Findings of ACL 2025

fine-grained F1/IoU
Authors: Yishuo Cai, Renjie Gu, Jiaxu LiCorresponding: Xuancheng HuangAffiliation: Central South University; Zhipu AI; Tsinghua University
165

Detecting and Preventing Hallucinations in Large Vision Language Models

M-HalDetect supplies fine-grained multimodal annotations and reward-model signals for hallucination detection.

fine-grained feedbackreward modelingobject attribute relation
Detection
vision-language

AAAI 2024

First posted Aug 11, 2023

fine-grained detection
Authors: Anisha Gunjal, Jihan Yin, Erhan BasCorresponding: Erhan BasAffiliation: Scale AI
166

Hal-Eval: a Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

A universal evaluator spanning object, attribute, relation, and event hallucinations with discriminative and generative modes.

event hallucinationfine-grained evaluationdiscriminative evaluation
Detection
vision-language

ACM MM 2024

First posted Feb 24, 2024

universal hallucination evaluation
Authors: Chaoya Jiang, Hongrui Jia, Mengfan DongCorresponding: Wei YeAffiliation: Peking University
167

TLDR: Token-Level Detective Reward Model for Large Vision Language Models

TLDR assigns token-level rewards to expose hallucinated spans and support visual-language self-correction.

token-level rewardsself-correctionhallucination evaluation
Detection
vision-language

ICLR 2025

First posted Oct 7, 2024

token-level reward
Authors: Deqing Fu, Tong Xiao, Rui WangCorresponding: Deqing Fu, Lawrence ChenAffiliation: Meta; University of Southern California
168

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Med-HallMark and MediHallDetector provide hierarchical medical hallucination evaluation and fine-grained detection.

medical hallucinationhierarchical evaluationmultitask detection
Detection
vision-language

arXiv 2024

First posted Jun 14, 2024

medical hallucination detection
Authors: Jiawei Chen, Dingkang Yang, Tong WuCorresponding: Lihua ZhangAffiliation: Fudan University; Tencent Youtu Lab; Cognition and Intelligent Technology Laboratory
169

Unified Hallucination Detection for Multimodal Large Language Models

UNIHD combines auxiliary tools for claim-level hallucination detection and introduces the MHaluBench benchmark.

tool-augmented verificationclaim-level detectionmulti-category detection
Detection
vision-language

ACL 2024

First posted Feb 5, 2024

unified hallucination detection
Authors: Xiang Chen, Chenxi Wang, Yida XueCorresponding: Ningyu Zhang, Huajun ChenAffiliation: Zhejiang University; Ant Group
170

VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation

VL-Uncertainty estimates response uncertainty under semantic-equivalent perturbations to detect hallucinations.

semantic-equivalent perturbationresponse entropyuncertainty estimation
Detection
vision-language

arXiv 2024

First posted Nov 18, 2024

uncertainty-based detection
Authors: Ruiyang Zhang, Hu Zhang, Zhedong ZhengCorresponding: Zhedong ZhengAffiliation: University of Macau; CSIRO Data61
171

Structural Graph Probing of Vision–Language Models

Graph-based probes model neuron-correlation topology and expose structural signals associated with multimodal behavior and hallucination.

graph probingneuron topologyhallucination classification
Detection
vision-language

arXiv 2026

First posted Mar 28, 2026

structural hallucination probing
Authors: Haoyu He, Yue Zhuo, Yu ZhengCorresponding: not specifiedAffiliation: Northeastern University; Massachusetts Institute of Technology
172

Lyapunov Probes for Hallucination Detection in Large Foundation Models

Stability-constrained perturbation probes distinguish factual representation regions from hallucination-prone boundaries.

stability theoryrepresentation probingperturbation analysis
Detection
vision-language

arXiv 2026

First posted Mar 6, 2026

stability-based detection
Authors: Bozhi Luan, Gen Li, Yalan QinCorresponding: Zhaoxin FanAffiliation: Beihang University; Shanghai University
173

HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token

HALP predicts hallucination risk before decoding by probing visual, vision-token, and query-token representations.

pre-generation detectioninternal representationslightweight probes
Detection
vision-language

EACL 2026

First posted Mar 5, 2026

pre-generation risk prediction
Authors: Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe SunCorresponding: not specifiedAffiliation: Stony Brook University; Toyota Technological Institute at Chicago
174

VADE: Visual Attention Guided Hallucination Detection and Elimination

VADE learns sequential patterns from raw visual attention maps for fine-grained hallucination detection and mitigation.

visual attentionsequence modelingfine-grained detection
Detection
vision-language

Findings of ACL 2025

attention-map detection
Authors: Vishnu Prabhakaran, Purav Aggarwal, Vinay Kumar VermaCorresponding: not specifiedAffiliation: Amazon, India; Amazon, USA
175

HALLUSHIFT++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

HALLUSHIFT++ extends internal distribution-shift analysis to hierarchical category, attribute, and relation hallucinations in MLLMs.

representation shiftshierarchical hallucinationsemantic chunking
Detection
vision-language

arXiv 2025

First posted Dec 8, 2025

hierarchical internal-state detection
Authors: Sujoy Nath, Arkaprabha Basu, Sharanya DasguptaCorresponding: Swagatam DasAffiliation: Netaji Subhash Engineering College; TCG Crest; Indian Statistical Institute
176

PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

PAS measures attention on preliminary output tokens as a training-free, reference-free object hallucination signal.

prelim-token attentiontraining-freeobject hallucination
Detection
vision-language

CVPR 2026

First posted Nov 14, 2025

attention-based object detection
Authors: Nhat Hoang-Xuan, Minh Vu, My T. ThaiCorresponding: not specifiedAffiliation: Los Alamos National Laboratory; University of Florida
177

MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs

MTRE aggregates early multi-token logits with likelihood-ratio evidence to estimate VLM response reliability.

multi-token logitsreliability estimationlikelihood ratios
Detection
vision-language

arXiv 2025

First posted May 16, 2025

multi-token reliability detection
Authors: Geigh Zollicoffer, Minh Vu, Manish BhattaraiCorresponding: not specifiedAffiliation: Los Alamos National Laboratory