VLDB 2026 Research / reviewers in the wild / expert
Chengqi Duan
dblp:336/2001
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-5737-019XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 39% Trustworthy machine learning · 27% Vision and language · 24% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
2.0 | 3 | 2025 | T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025 GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025 |
Visual content generation and editing
image editing |
1.7 | 2 | 2025 | GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025 PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025 |
Visual content generation and editing › image generation
text-to-image generation |
1.7 | 2 | 2025 | GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025 PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
compositional text-to-image generation |
0.9 | 1 | 2025 | T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing |
0.9 | 1 | 2025 | GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial defense |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Machine learning › Generative modeling › generative model
generative classifier |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Computer vision › Image recognition and object detection
image classification |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › robust learning
robust classification |
0.8 | 1 | 2024 | Robust Classification via a Single Diffusion Model · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.5semantic-spatial guidance · 1.7multimodal pretraining · 1.7instruction tuning · 1.7chain-of-thought reasoning · 1.7multimodal large language model evaluation · 0.9detection-based metric · 0.9multi-head diffusion · 0.8efficient sampling · 0.8bayes' theorem · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PUMA: Empowering Unified MLLM with Multi-Granular Visual GenerationabstractRecent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content generation. However, existing works have insufficiently addressed the varying granularity demands of different image generation tasks within a unified MLLM paradigm - from the diversity required in text-to-image generation to the precise controllability needed in image manipulation. In this work, we propose PUMA, emPowering Unified MLLM with Multi-grAnular visual generation. PUMA unifies multi-granular visual features as both inputs and outputs of MLLMs, elegantly addressing the different granularity requirements of various image generation tasks within a unified MLLM framework. Following multimodal pretraining and task-specific instruction tuning, PUMA demonstrates proficiency in a wide range of multimodal tasks. This work represents a significant step towards a truly unified MLLM capable of adapting to the granularity demands of various visual tasks. The code and model will be released in https://github.com/rongyaofang/PUMA. Rongyao Fang, Chengqi Duan, Kun Wang 0056, Hao Li 0069, Linjiang Huang, Hao Tian 0006, Xingyu Zeng, Rui Zhao 0001, Jifeng Dai, Hongsheng Li 0001, Xihui Liu |
ICCV | 2 |
| 2025 | GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and EditingabstractCurrent image generation and editing methods primarily process textual prompts as direct inputs without explicit reasoning about visual composition or operational steps. We present Generation Chain-of-Thought (GoT), a novel paradigm that empowers a Multimodal Large Language Model (MLLM) to first generate an explicit, structured reasoning chain in natural language—detailing semantic relationships, object attributes, and, crucially, precise spatial coordinates—before any image synthesis occurs. This intermediate reasoning output directly guides the subsequent visual generation or editing process. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over \textbf{9M} samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. We will release our datasets and models to facilitate future research. Rongyao Fang, Chengqi Duan, Kun Wang 0056, Linjiang Huang, Hao Li 0069, Hao Tian 0006, Shilin Yan, Weihao Yu 0005, Xingyu Zeng, Jifeng Dai, Xihui Liu, Hongsheng Li 0001 |
NeurIPS | 2 |
| 2025 | T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image GenerationabstractDespite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present T2I-CompBench++, an enhanced benchmark for compositional text-to-image generation. T2I-CompBench++ comprises 8,000 compositional text prompts categorized into four primary groups: attribute binding, object relationships, generative numeracy, and complex compositions. These are further divided into eight sub-categories, including newly introduced ones like 3D-spatial relationships and numeracy. In addition to the benchmark, we propose enhanced evaluation metrics designed to assess these diverse compositional challenges. These include a detection-based metric tailored for evaluating 3D-spatial relationships and numeracy, and an analysis leveraging Multimodal Large Language Models (MLLMs), i.e. GPT-4 V, ShareGPT4v as evaluation metrics. Our experiments benchmark 11 text-to-image models, including state-of-the-art models, such as FLUX.1, SD3, DALLE-3, Pixart-$\alpha$α, and SD-XL on T2I-CompBench++. We also conduct comprehensive evaluations to validate the effectiveness of our metrics and explore the potential and limitations of MLLMs. Chengqi Duan, Kaiyue Sun 0001, Enze Xie, Zhenguo Li, Xihui Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Robust Classification via a Single Diffusion ModelabstractDiffusion models have been applied to improve adversarial robustness of image classifiers by purifying the adversarial noises or generating realistic data for adversarial training. However, diffusion-based purification can be evaded by stronger adaptive attacks while adversarial training does not perform well under unseen threats, exhibiting inevitable limitations of these methods. To better harness the expressive power of diffusion models, this paper proposes Robust Diffusion Classifier (RDC), a generative classifier that is constructed from a pre-trained diffusion model to be adversarially robust. RDC first maximizes the data likelihood of a given input and then predicts the class probabilities of the optimized input using the conditional likelihood estimated by the diffusion model through Bayes’ theorem. To further reduce the computational cost, we propose a new diffusion backbone called multi-head diffusion and develop efficient sampling strategies. As RDC does not require training on particular adversarial attacks, we demonstrate that it is more generalizable to defend against multiple unseen threats. In particular, RDC achieves $75.67%$ robust accuracy against various $\ell_\infty$ norm-bounded adaptive attacks with $\epsilon_\infty=8/255$ on CIFAR-10, surpassing the previous state-of-the-art adversarial training models by $+4.77%$. The results highlight the potential of generative classifiers by employing pre-trained diffusion models for adversarial robustness compared with the commonly studied discriminative classifiers. Huanran Chen, Yinpeng Dong, Xiao Yang 0028, Chengqi Duan, Hang Su 0006, Jun Zhu 0001 |
ICML | 5 |