Skip to content
Home›Detection Methods›Overview
🔎

Detection Methods

Methods for detecting multimodal hallucinations.

30Curated entries
4Categories
56Verified links
28Dated papers

Links, benchmarks, and short notes.

30 curated entries

Method Categories

Curated Methods

1

HallusionBench

Control-group diagnosis of language hallucination and visual illusion in LVLMs.

control groupsvisual illusionlanguage hallucination
Consistency-based
vision-language

CVPR 2024

First posted Oct 23, 2023

paired visual diagnosis
Authors: Tianrui Guan, Fuxiao Liu, Xiyang WuCorresponding: not specifiedAffiliation: University of Maryland, College Park
2

POPE

Polling-based object probing for detecting object hallucination in LVLMs.

object existenceyes/no probingadversarial negatives
Consistency-based
vision-language

EMNLP 2023

First posted May 17, 2023

object perception probing
Authors: Yifan Li, Yifan Du, Kun ZhouCorresponding: Wayne Xin ZhaoAffiliation: Renmin University of China; Meituan Group
3

FaithScore

Atomic image-fact verification for fine-grained hallucination detection.

atomic factsimage groundingreference-free
Knowledge-based
vision-language

Findings of EMNLP 2024

First posted Nov 2, 2023

atomic fact verification
Authors: Liqiang Jing, Ruosen Li, Yunmo ChenCorresponding: not specifiedAffiliation: University of Texas at Dallas; Johns Hopkins University
4

HaELM

A reproducible local language-model evaluator for LVLM hallucinations.

LLM evaluatorlocal evaluationhallucination scoring
Model-based
vision-language

arXiv 2023

First posted Aug 29, 2023

model-based evaluation
Authors: Junyang Wang, Yiyang Zhou, Guohai XuCorresponding: not specifiedAffiliation: Shandong University; Beijing Jiaotong University; Xi'an Jiaotong University
5

AMBER

An LLM-free pipeline for detecting existence, attribute, and relation hallucinations.

existenceattributesrelations
Knowledge-based
vision-language

arXiv 2023

First posted Nov 13, 2023

multi-dimensional detection
Authors: Junyang Wang, Yuhang Wang, Guohai XuCorresponding: not specifiedAffiliation: Beijing Jiaotong University; Alibaba Group
6

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

A training-free detector that combines global scene similarity with local visual grounding for object hallucination detection.

global-local similarityVisual Logit Lensobject grounding
Model-based
vision-language

NeurIPS 2025

First posted Aug 27, 2025

global-local grounding
Authors: Seongheon Park, Sharon LiCorresponding: not specifiedAffiliation: University of Wisconsin–Madison
7

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

A training-free object hallucination detector that uses calibrated local and instruction-context consistency scores.

instruction embeddingsLogit Lensobject hallucination
Model-based
vision-language

ICML 2026

First posted May 12, 2026

instruction-token scoring
Authors: Runhe Lai, Xinhua Lu, Yanqi WuCorresponding: Weijiang Yu, Ruixuan WangAffiliation: Sun Yat-sen University; Peng Cheng Laboratory; Key Laboratory of Machine Intelligence and Advanced Computing, MOE
8

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

An attention-lens analysis identifies middle-layer signals for object hallucination detection and visual-attention adjustment.

attention lensmiddle layersobject hallucination
Model-based
vision-language

CVPR 2025

First posted Nov 23, 2024

attention-based signals
Authors: Zhangqi Jiang, Junkai Chen, Beier ZhuCorresponding: Tingjin Luo, Xu YangAffiliation: National University of Defense Technology; Southeast University; Nanyang Technological University
9

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

A span-level detector that classifies six hallucination error types and proposes grounded edits for MLLM outputs.

fine-grained spanserror taxonomyhallucination editing
Model-based
vision-language

CVPR 2026

First posted Jun 16, 2025

span-level detection
Authors: Yuiga Wada, Kazuki Matsuda, Komei SugiuraCorresponding: not specifiedAffiliation: Keio AI Research Center; Keio University; Carnegie Mellon University
10

Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

VBackChecker detects hallucinations by grounding response claims backward to image pixels with rich contextual evidence.

backward visual groundingrich-contextpixel-level grounding
Knowledge-based
vision-language

AAAI 2026

First posted Nov 15, 2025

pixel-level grounding
Authors: Pinxue Guo, Chongruo Wu, Xinyu ZhouCorresponding: Wei Zhang, Wenqiang ZhangAffiliation: Fudan University; Independent Researcher
11

Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

A detect-then-rewrite pipeline uses fine-grained object, attribute, and relationship feedback to train HSA-DPO.

fine-grained feedbackseverity-awaredetect-then-rewrite
Model-based
vision-language

AAAI 2025

First posted Apr 22, 2024

sentence-level detection
Authors: Wenyi Xiao, Ziwei Huang, Leilei GanCorresponding: Leilei GanAffiliation: Zhejiang University; Alibaba Group
12

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

HalLoc introduces token-level hallucination localization data and a concurrent low-overhead detector.

token-level localizationgraded confidencehallucination types
Model-based
vision-language

CVPR 2025

First posted Jun 12, 2025

token-level detection
Authors: Eunkyu Park, Minyeong Kim, Gunhee KimCorresponding: Gunhee KimAffiliation: Seoul National University
13

MHALO: Evaluating MLLMs as Fine-grained Hallucination Detectors

A fine-grained hallucination detection benchmark that evaluates MLLMs on recognition and token-level localization.

token-level detectionmeta-evaluation benchmarkMLLM as judge
Model-based
vision-language

Findings of ACL 2025

fine-grained F1/IoU
Authors: Yishuo Cai, Renjie Gu, Jiaxu LiCorresponding: Xuancheng HuangAffiliation: Central South University; Zhipu AI; Tsinghua University
14

Detecting and Preventing Hallucinations in Large Vision Language Models

M-HalDetect provides fine-grained multimodal annotations and reward-model signals for hallucination detection.

fine-grained feedbackreward modelingobject attribute relation
Model-based
vision-language

AAAI 2024

First posted Aug 11, 2023

fine-grained detection
Authors: Anisha Gunjal, Jihan Yin, Erhan BasCorresponding: Erhan BasAffiliation: Scale AI
15

Hal-Eval: a Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

A unified discriminative and generative evaluator covering object, attribute, relation, and event hallucinations.

event hallucinationfine-grained evaluationdiscriminative evaluation
Knowledge-based
vision-language

ACM MM 2024

First posted Feb 24, 2024

universal hallucination evaluation
Authors: Chaoya Jiang, Hongrui Jia, Mengfan DongCorresponding: Wei YeAffiliation: Peking University
16

TLDR: Token-Level Detective Reward Model for Large Vision Language Models

TLDR assigns token-level rewards to expose hallucinated spans and support visual-language self-correction.

token-level rewardsself-correctionhallucination evaluation
Model-based
vision-language

ICLR 2025

First posted Oct 7, 2024

token-level reward
Authors: Deqing Fu, Tong Xiao, Rui WangCorresponding: Deqing Fu, Lawrence ChenAffiliation: Meta; University of Southern California
17

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Med-HallMark and MediHallDetector provide hierarchical medical hallucination evaluation and fine-grained detection.

medical hallucinationhierarchical evaluationmultitask detection
Model-based
vision-language

arXiv 2024

First posted Jun 14, 2024

medical hallucination detection
Authors: Jiawei Chen, Dingkang Yang, Tong WuCorresponding: Lihua ZhangAffiliation: Fudan University; Tencent Youtu Lab; Cognition and Intelligent Technology Laboratory
18

Unified Hallucination Detection for Multimodal Large Language Models

UNIHD unifies claim-level hallucination detection with auxiliary tools and introduces the MHaluBench meta-evaluation benchmark.

tool-augmented verificationclaim-level detectionmulti-category detection
Knowledge-based
vision-language

ACL 2024

First posted Feb 5, 2024

unified hallucination detection
Authors: Xiang Chen, Chenxi Wang, Yida XueCorresponding: Ningyu Zhang, Huajun ChenAffiliation: Zhejiang University; Ant Group
19

VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation

VL-Uncertainty uses semantic-equivalent perturbations and response uncertainty to detect hallucinations without extra labels.

semantic-equivalent perturbationresponse entropyuncertainty estimation
Uncertainty-based
vision-language

arXiv 2024

First posted Nov 18, 2024

uncertainty-based detection
Authors: Ruiyang Zhang, Hu Zhang, Zhedong ZhengCorresponding: Zhedong ZhengAffiliation: University of Macau; CSIRO Data61
20

Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models

LogicCheckGPT probes object-to-attribute and attribute-to-object consistency in a training-free closed loop.

logical consistencyobject hallucinationclosed-loop probing
Consistency-based
vision-language

Findings of ACL 2024

First posted Feb 18, 2024

logical consistency probing
Authors: Junfei Wu, Qiang Liu, Ding WangCorresponding: Shu WuAffiliation: Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Nanjing University
21

Structural Graph Probing of Vision–Language Models

Graph-based probes model neuron-correlation topology to expose structural signals associated with multimodal behavior and hallucination.

graph probingneuron topologyhallucination classification
Model-based
vision-language

arXiv 2026

First posted Mar 28, 2026

structural hallucination probing
Authors: Haoyu He, Yue Zhuo, Yu ZhengCorresponding: not specifiedAffiliation: Northeastern University; Massachusetts Institute of Technology
22

Lyapunov Probes for Hallucination Detection in Large Foundation Models

Lyapunov Probes use stability-constrained perturbation analysis to distinguish stable factual regions from hallucination-prone boundaries.

stability theoryrepresentation probingperturbation analysis
Model-based
vision-language

arXiv 2026

First posted Mar 6, 2026

stability-based detection
Authors: Bozhi Luan, Gen Li, Yalan QinCorresponding: Zhaoxin FanAffiliation: Beihang University; Shanghai University
23

HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token

HALP predicts hallucination risk before decoding by probing visual, vision-token, and query-token representations.

pre-generation detectioninternal representationslightweight probes
Model-based
vision-language

EACL 2026

First posted Mar 5, 2026

pre-generation risk prediction
Authors: Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe SunCorresponding: not specifiedAffiliation: Stony Brook University; Toyota Technological Institute at Chicago
24

VADE: Visual Attention Guided Hallucination Detection and Elimination

VADE learns sequential patterns from raw visual attention maps for fine-grained hallucination detection and mitigation.

visual attentionsequence modelingfine-grained detection
Model-based
vision-language

Findings of ACL 2025

attention-map detection
Authors: Vishnu Prabhakaran, Purav Aggarwal, Vinay Kumar VermaCorresponding: not specifiedAffiliation: Amazon, India; Amazon, USA
25

HALLUSHIFT++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

HALLUSHIFT++ extends internal distribution-shift analysis to hierarchical category, attribute, and relation hallucinations in MLLMs.

representation shiftshierarchical hallucinationsemantic chunking
Model-based
vision-language

arXiv 2025

First posted Dec 8, 2025

hierarchical internal-state detection
Authors: Sujoy Nath, Arkaprabha Basu, Sharanya DasguptaCorresponding: Swagatam DasAffiliation: Netaji Subhash Engineering College; TCG Crest; Indian Statistical Institute
26

PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

PAS measures attention on preliminary output tokens as a training-free, reference-free object hallucination signal.

prelim-token attentiontraining-freeobject hallucination
Model-based
vision-language

CVPR 2026

First posted Nov 14, 2025

attention-based object detection
Authors: Nhat Hoang-Xuan, Minh Vu, My T. ThaiCorresponding: not specifiedAffiliation: Los Alamos National Laboratory; University of Florida
27

MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs

MTRE aggregates early multi-token logits with likelihood-ratio evidence to estimate VLM response reliability.

multi-token logitsreliability estimationlikelihood ratios
Uncertainty-based
vision-language

arXiv 2025

First posted May 16, 2025

multi-token reliability detection
Authors: Geigh Zollicoffer, Minh Vu, Manish BhattaraiCorresponding: not specifiedAffiliation: Los Alamos National Laboratory
28

Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT

ACT-ViT treats full layer-by-token activation tensors as image-like inputs for cross-LLM hallucination detection.

activation tensorsvision transformercross-model transfer
Model-based
vision-language

NeurIPS 2025

First posted Sep 30, 2025

activation-tensor detection
Authors: Guy Bar-Shalom, Fabrizio Frasca, Yaniv GalronCorresponding: not specifiedAffiliation: Technion
29

TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention

A latent-state detector identifies hallucinated object tokens and guides decoding toward a shared truthful direction.

truthful directionlatent subspacepre-intervention
Model-based
vision-language

ICCV 2025

First posted Mar 13, 2025

latent-state detection
Authors: Jinhao Duan, Fei Kong, Hao ChengCorresponding: Kaidi XuAffiliation: Drexel University; University of Electronic Science and Technology of China; Hong Kong University of Science and Technology (Guangzhou)
30

Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations

A patch-level detector combines attention dispersion and cross-modal grounding consistency to localize hallucinated tokens.

attention dispersioncross-modal groundingpatch-level detection
Model-based
vision-language

CVPR 2026

First posted Apr 6, 2026

token-grounding detection
Authors: Tuan Dung Nguyen, Minh Khoi Ho, Qi ChenCorresponding: Phi Le Nguyen, Vu Minh Hieu PhanAffiliation: Hanoi University of Science and Technology; Australian Institute for Machine Learning, University of Adelaide; Mohamed bin Zayed University of Artificial Intelligence