arxiv:2503.15264

LEGION: Learning to Ground and Explain for Synthetic Image Detection

Published on Mar 19

· Submitted by

zichenwen on Mar 20

Upvote

Authors:

Hengrui Kang ,

Zichen Wen ,

Junyan Ye ,

Weijia Li ,

Peilin Feng ,

Baichuan Zhou ,

Dahua Lin ,

Linfeng Zhang ,

Conghui He

Abstract

The rapid advancements in generative technology have emerged as a double-edged sword. While offering powerful tools that enhance convenience, they also pose significant social concerns. As defenders, current synthetic image detection methods often lack artifact-level textual interpretability and are overly focused on image manipulation detection, and current datasets usually suffer from outdated generators and a lack of fine-grained annotations. In this paper, we introduce SynthScars, a high-quality and diverse dataset consisting of 12,236 fully synthetic images with human-expert annotations. It features 4 distinct image content types, 3 categories of artifacts, and fine-grained annotations covering pixel-level segmentation, detailed textual explanations, and artifact category labels. Furthermore, we propose LEGION (LEarning to Ground and explain for Synthetic Image detectiON), a multimodal large language model (MLLM)-based image forgery analysis framework that integrates artifact detection, segmentation, and explanation. Building upon this capability, we further explore LEGION as a controller, integrating it into image refinement pipelines to guide the generation of higher-quality and more realistic images. Extensive experiments show that LEGION outperforms existing methods across multiple benchmarks, particularly surpassing the second-best traditional expert on SynthScars by 3.31% in mIoU and 7.75% in F1 score. Moreover, the refined images generated under its guidance exhibit stronger alignment with human preferences. The code, model, and dataset will be released.

View arXiv page View PDF Project page GitHub repository Add to collection

Community

zichenwen

Paper author Paper submitter about 17 hours ago

We explore the fully synthetic forgery analysis task and introduce SynthScars, a challenging and finely annotated forged image dataset. Additionally, we propose LEGION, a framework supporting three subtasks—artifact localization, explanation generation, and forgery detection. LEGION's detailed feedback further guides image regeneration and inpainting pipelines, promoting higher-quality and more realistic image generation.
Our main contributions are as follows:

We introduce SynthScars, a challenging dataset for synthetic image detection, featuring high-quality synthetic images with diverse content types, as well as fine-grained pixel-level artifact annotations with detailed textual explanations.
We propose LEGION, a comprehensive image forgery analysis framework for artifact localization, explanation generation, and forgery detection, which effectively aids human experts to detect and understand image forgeries.
Extensive experiments demonstrate that LEGION achieves exceptional performance on 4 challenging benchmarks. Comparisons with 19 existing methods show that it achieves state-of-the-art performance on the vast majority of metrics, exhibiting strong robustness and generalization ability.
We position LEGION not only as a defender against ever-evolving generative technologies but also as a controller that guides higher-quality and more realistic image generation. Qualitative and quantitative experiments on image regeneration and inpainting show the great value of LEGION in providing feedbacks for progressive artifact refinement.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Your need to confirm your account before you can post a new comment.

· Sign up or log in to comment

Upvote

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2503.15264 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2503.15264 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2503.15264 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.