Baichuan Zhou

dblp:371/2742 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 36% Trustworthy machine learning · 28% 3D vision · 24%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.722025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Machine learning › Trustworthy machine learning › AI-generated content detection
AI-generated image detection
0.912025
LEGION: Learning to Ground and Explain for Synthetic Image Detection · ICCV 2025
Computer vision › 3D vision › 3d scene understanding › multi-view understanding
cross-view reasoning
0.912025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Computer vision › 3D vision › visual localization
geo-localization
0.912025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
LEGION: Learning to Ground and Explain for Synthetic Image Detection · ICCV 2025
Machine learning › Trustworthy machine learning
synthetic data detection
0.912025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
Machine learning › Trustworthy machine learning
AI-generated content detection
0.312025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
Image and video processing › image enhancement
image refinement
0.312025
LEGION: Learning to Ground and Explain for Synthetic Image Detection · ICCV 2025

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 1.7artifact grounding · 1.7large multimodal model · 1.7cross-view detection-matching · 0.9
YearPublicationVenuePosition
2025 UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
abstract
Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have been limited to evaluating LMMs with basic region-level urban tasks under singular views, leading to incomplete evaluations of LMMs' abilities in urban environments. To address these issues, we present UrBench, a comprehensive benchmark designed for evaluating LMMs in complex multi-view urban scenarios. UrBench contains 11.6K meticulously curated questions at both region-level and role-level that cover 4 task dimensions: Geo-Localization, Scene Reasoning, Scene Understanding, and Object Understanding, totaling 14 task types. In constructing UrBench, we utilize data from existing datasets and additionally collect data from 11 cities, creating new annotations using a cross-view detection-matching method. With these images and annotations, we then integrate LMM-based, rule-based, and human-based methods to construct large-scale high-quality questions. Our evaluations on 21 LMMs show that current LMMs struggle in the urban environments in several aspects. Even the best performing GPT-4o lags behind humans in most tasks, ranging from simple tasks such as counting to complex tasks such as orientation, localization and object attribute recognition, with an average performance gap of 17.4%. Our benchmark also reveals that LMMs exhibit inconsistent behaviors with different urban views, especially with respect to understanding cross-view relations.
Baichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye, Tianyi Bai, Songyang Zhang 0001, Dahua Lin, Conghui He
AAAI1
2025 LEGION: Learning to Ground and Explain for Synthetic Image Detection
abstract
The rapid advancements in generative technology have emerged as a double-edged sword. While offering powerful tools that enhance convenience, they also pose significant social concerns. As defenders, current synthetic image detection methods often lack artifact-level textual interpretability and are overly focused on image manipulation detection, and current datasets usually suffer from outdated generators and a lack of fine-grained annotations. In this paper, we introduce SynthScars, a high-quality and diverse dataset consisting of 12,236 fully synthetic images with human-expert annotations. It features 4 distinct image content types, 3 categories of artifacts, and fine-grained annotations covering pixel-level segmentation, detailed textual explanations, and artifact category labels. Furthermore, we propose LEGION (LEarning to Ground and explain for Synthetic Image detectiON), a multimodal large language model (MLLM)-based image forgery analysis framework that integrates artifact detection, segmentation, and explanation. Building upon this capability, we further explore LEGION as a controller, integrating it into image refinement pipelines to guide the generation of higher-quality and more realistic images. Extensive experiments show that LEGION outperforms existing methods across multiple benchmarks, particularly surpassing the second-best traditional expert on SynthScars by 3.31% in mIoU and 7.75% in F1 score. Moreover, the refined images generated under its guidance exhibit stronger alignment with human preferences. The code, model, and dataset will be released.
Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Peilin Feng, Baichuan Zhou, Bin Wang 0065, Dahua Lin, Linfeng Zhang 0001, Conghui He
ICCV7
2025 LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
abstract
With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large multimodal models (LMMs) in this task has attracted significant interest. LMMs can provide natural language explanations for their authenticity judgments, enhancing the explainability of synthetic content detection. Simultaneously, the task of distinguishing between real and synthetic data effectively tests the perception, knowledge, and reasoning capabilities of LMMs. In response, we introduce LOKI, a novel benchmark designed to evaluate the ability of LMMs to detect synthetic data across multiple modalities. LOKI encompasses video, image, 3D, text, and audio modalities, comprising 18K carefully curated questions across 26 subcategories with clear difficulty levels. The benchmark includes coarse-grained judgment and multiple-choice questions, as well as fine-grained anomaly selection and explanation tasks, allowing for a comprehensive analysis of LMMs. We evaluated 22 open-source LMMs and 6 closed-source models on LOKI, highlighting their potential as synthetic data detectors and also revealing some limitations in the development of LMM capabilities. More information about LOKI can be found at https://opendatalab.github.io/LOKI/.
Junyan Ye, Baichuan Zhou, Junan Zhang, Tianyi Bai, Hengrui Kang, Honglin Lin, Zhizheng Wu 0001, Dahua Lin, Conghui He
ICLR2
2025 BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
abstract
Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating visual input as replaceable context. To address this gap, we introduce BLINK-Twice, a vision-centric reasoning benchmark grounded in challenging perceptual tasks. Instead of relying on external knowledge, our tasks require models to reason from visual content alone, shifting the focus from language-based to image-grounded reasoning. Compared to prior perception benchmarks, it moves beyond shallow perception ("see") and requires fine-grained observation and analytical reasoning ("observe"). BLINK-Twice integrates three core components: seven types of visual challenges for testing visual reasoning, natural adversarial image pairs that enforce reliance on visual content, and annotated reasoning chains for fine-grained evaluation of the reasoning process rather than final answers alone. We evaluate 20 leading MLLMs, including 12 foundation models and 8 reasoning-enhanced models. BLINK-Twice poses a significant challenge to current models. While existing reasoning strategies in the language space—such as chain-of-thought or self-criticism can improve performance, they often result in unstable and redundant reasoning. We observe that repeated image observation improves performance across models, and active visual interaction, as demonstrated by models like o3, highlights the need for a new paradigm for vision reasoning. The dataset is publicly available at https://github.com/PicoTrex/BLINK-Twice.
Junyan Ye, Dongzhi Jiang, Baichuan Zhou, Zhiyuan Yan 0002, Hongsheng Li 0001, Conghui He
NeurIPS4