Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

ForgeReason: Multimodal Reasoning with Multi-Agent Collaboration for Generalizable and Explainable Image Forgery Detection

View through CrossRef
Abstract The rapid advancement of generative AI and sophisticated image editing tools has made image forgery increasingly convincing, posing serious threats to media integrity, legal evidence, and public trust. Existing forgery detection methods primarily rely on single-modality analysis and operate as black-box classifiers, suffering from limited generalization to unseen forgery types and a lack of interpretability in their detection decisions. In this paper, we propose ForgeReason, a multimodal reasoning framework that integrates dual-stream feature extraction, reasoning alignment, and multi-agent collaboration for generalizable and explainable image forgery detection. The dual-stream feature extractor simultaneously captures spatial semantic features through a vision transformer and frequency-domain forensic traces through SRM-filtered encoding, with multi-scale attention fusion to enhance manipulation-sensitive representations. The reasoning alignment module bridges visual forensic features and linguistic reasoning spaces via cross-modal contrastive learning, enabling the model to articulate detection evidence in natural language. The multi-agent collaborative decision mechanism coordinates specialized agents for visual analysis, frequency analysis, and reasoning verification, producing comprehensive and reliable detection outcomes with interpretable explanations. Extensive experiments on three benchmark datasets demonstrate that ForgeReason achieves state-of-the-art performance, obtaining 96.8 percent accuracy on CASIA v2, 90.3 percent on IMD2020, and 92.5 percent on AutoSplice, while providing high-quality textual explanations for its detection decisions. Ablation studies confirm that each component contributes positively, with the full framework outperforming all individual and partial configurations.
Title: ForgeReason: Multimodal Reasoning with Multi-Agent Collaboration for Generalizable and Explainable Image Forgery Detection
Description:
Abstract The rapid advancement of generative AI and sophisticated image editing tools has made image forgery increasingly convincing, posing serious threats to media integrity, legal evidence, and public trust.
Existing forgery detection methods primarily rely on single-modality analysis and operate as black-box classifiers, suffering from limited generalization to unseen forgery types and a lack of interpretability in their detection decisions.
In this paper, we propose ForgeReason, a multimodal reasoning framework that integrates dual-stream feature extraction, reasoning alignment, and multi-agent collaboration for generalizable and explainable image forgery detection.
The dual-stream feature extractor simultaneously captures spatial semantic features through a vision transformer and frequency-domain forensic traces through SRM-filtered encoding, with multi-scale attention fusion to enhance manipulation-sensitive representations.
The reasoning alignment module bridges visual forensic features and linguistic reasoning spaces via cross-modal contrastive learning, enabling the model to articulate detection evidence in natural language.
The multi-agent collaborative decision mechanism coordinates specialized agents for visual analysis, frequency analysis, and reasoning verification, producing comprehensive and reliable detection outcomes with interpretable explanations.
Extensive experiments on three benchmark datasets demonstrate that ForgeReason achieves state-of-the-art performance, obtaining 96.
8 percent accuracy on CASIA v2, 90.
3 percent on IMD2020, and 92.
5 percent on AutoSplice, while providing high-quality textual explanations for its detection decisions.
Ablation studies confirm that each component contributes positively, with the full framework outperforming all individual and partial configurations.

Related Results

Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
A Review on Image Forgery Detection Techniques Using Machine Learning
A Review on Image Forgery Detection Techniques Using Machine Learning
Image forgery has evolved into common problem in the digital age, due to the extensive uses of digital image manipulation tools. In a variety of industries, including forensics, jo...
Logical Challenges in Artificial General Intelligence
Logical Challenges in Artificial General Intelligence
The present thesis pertains to the research area of logic for artificial intelligence (AI), and is motivated by the critical role of automated reasoning in AI, particularly by the ...
Explainable Image-Centric Forgery Detection: A Survey
Explainable Image-Centric Forgery Detection: A Survey
The rapid growth of AI-driven image manipulation technologies poses critical challenges for verifying content authenticity. While many forgery detection systems achieve high accura...
VISA-Agent: A Visual Symbolic Agent for Reasoning-Intensive Multimodal Retrieval
VISA-Agent: A Visual Symbolic Agent for Reasoning-Intensive Multimodal Retrieval
Reasoning-intensive multimodal retrieval suffers from a counter-intuitive bottleneck: on MM-BRIGHT multimodal-to-text (Query+Image → Documents), the strongest dense multimodal enco...
Copy-Move Image Forgery Detection Using Deep Learning Approaches: An Abbreviated Survey
Copy-Move Image Forgery Detection Using Deep Learning Approaches: An Abbreviated Survey
Images play a fundamental role in digital media, and altering digital images can present a significant risk since it contributes to disseminating false information. The rapid advan...
Human-centric and Semantics-based Explainable Event Detection: A Survey
Human-centric and Semantics-based Explainable Event Detection: A Survey
Abstract In recent years, there has been a surge in interest in artificial intelligent systems that can provide human-centric explanations for decisions or predictions. No ...

Back to Top