VLDB 2026 Research / reviewers in the wild / expert
Yuke Xing
dblp:383/6940
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal ModelsabstractWith the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticity. Current deepfake detection methods often depend on datasets with limited generation models and content diversity that fail to keep pace with the evolving complexity and increasing realism of the AI-generated content. Large multimodal models (LMMs), widely adopted in various vision tasks, have demonstrated strong zero-shot capabilities, yet their potential in deepfake detection remains largely unexplored. To bridge this gap, we present DFBench, a large-scale DeepFake Benchmark featuring (i) broad diversity, including 540,000 images across real, AI-edited, and AI-generated content, (ii) latest model, the fake images are generated by 12 state-of-the-art generation models, and (iii) bidirectional benchmarking and evaluating for both the detection accuracy of deepfake detectors and the evasion capability of generative models. Based on DFBench, we propose MoA-DF, Mixture of Agents for DeepFake detection, leveraging a combined probability strategy from multiple LMMs. MoA-DF achieves state-of-the-art performance, further proving the effectiveness of leveraging LMMs for deepfake detection. Database and codes are publicly available at https://github.com/IntMeGroup/DFBench. Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 9 |
| 2025 | 3DGS-IEval-15K: A Large-scale Image Quality Evaluation Database for 3D Gaussian-Splatting
Yuke Xing, Peizhi Niu, Guangtao Zhai, Yiling Xu |
ACM Multimedia | 1 |
| 2025 | CompBench: Benchmarking and Comparing Image Generation with Large Multimodal ModelsabstractRecent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench. Huiyu Duan, Yuke Xing, Yiling Xu, Guangtao Zhai, Xiongkuo Min |
MMSP | 3 |
| 2025 | CVBench: Benchmarking and Comparing Video Generation with Large Multimodal ModelsabstractLarge multimodal models (LMMs) have revolutionized both text-to-video (T2V) generation and video-to-text (V2T) interpretation. However, despite these advancements, issues such as imperfect perceptual quality and inconsistent text-video alignment continue to limit the practical deployment of AI-generated videos (AIGVs). Consequently, there is a pressing need for a reliable benchmark and automatic evaluation framework tailored for AIGVs. To this end, we propose CVBench, the largest and most comprehensive dataset for Comparative Video Benchmarking, including 60K video pairs generated by 30 state-of-the-art T2V models and 600K pairwise comparisons annotated with over 1.7 million human judgments from perspectives of both perceptual quality and text-video correspondence. This dataset enables bidirectional benchmarking and evaluation of both T2V generation models and V2T interpretation models. Based on CVBench, we propose VComp, a novel LMM-based evaluation metric that captures fine-grained quality differences from multiple perspectives for pairwise comparison at both the instance level and model level. Extensive experiments show that VComp achieves state-of-the-art alignment with human preferences. Both the CVBench dataset and VComp metric will be available at https://github.com/IntMeGroup/CVBench. Huiyu Duan, Yuke Xing, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
VCIP | 3 |
| 2025 | 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compressionabstract3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual fidelity, but its significant storage demands limit practical deployment, prompting recent methods to integrate compression modules into 3DGS. However, these 3DGS generative compression techniques introduce unique distortions that lack systematic quality assessment research. To this end, we establish 3DGS-VBench, a large-scale Video Quality Assessment (VQA) dataset and benchmark with 660 compressed 3DGS models and video sequences generated from 11 scenes across 6 representative 3DGS compression algorithms with systematically designed parameter levels. With annotations from 50 participants, we obtain MOS scores with outlier removal and validate dataset reliability. We benchmark 6 3DGS compression algorithms on storage efficiency and visual quality, and evaluate 15 quality assessment metrics. Our dataset enables specialized VQA model training for 3DGS. The dataset is available at https://github.com/YukeXing/3DGS-VBench. Yuke Xing, William Gordon, Qi Yang 0003, Kaifa Yang, Yiling Xu |
VCIP | 1 |
| 2024 | A Benchmark for Gaussian Splatting Compression and Quality Assessment Study
Qi Yang 0003, Kaifa Yang, Yuke Xing, Yiling Xu, Zhu Li 0001 |
MMAsia | 3 |
| 2024 | Explicit-NeRF-QA: A Quality Assessment Database for Explicit NeRF Model CompressionabstractIn recent years, Neural Radiance Fields (NeRF) have demonstrated significant advantages in representing and synthesizing 3D scenes. Explicit NeRF models facilitate the practical NeRF applications with faster rendering speed, and also attract considerable attention in NeRF compression due to its huge storage cost. To address the challenge of the NeRF compression study, in this paper, we construct a new dataset, called Explicit-NeRF-QA. We use 22 3D objects with diverse geometries, textures, and material complexities to train four typical explicit NeRF models across five parameter levels. Lossy compression is introduced during the model generation, pivoting the selection of key parameters such as hash table size for InstantNGP and voxel grid resolution for Plenoxels. By rendering NeRF samples to processed video sequences (PVS), a large scale subjective experiment with lab environment is conducted to collect subjective scores from 21 viewers. The diversity of content, accuracy of mean opinion scores (MOS), and characteristics of NeRF distortion are comprehensively presented, establishing the heterogeneity of the proposed dataset. The state-of-the-art objective metrics are tested in the new dataset. Best Pearson correlation, which is around 0.85, is collected from the full-reference objective metric. All tested no-reference metrics report very poor results with 0.4 to 0.6 correlations, demonstrating the need for further development of more robust no-reference metrics. The dataset, including NeRF samples, source 3D objects, multiview images for NeRF generation, PVSs, MOS, is made publicly available at the following location: https://github.com/YukeXing/Explicit-NeRF-QA. Yuke Xing, Qi Yang 0003, Kaifa Yang, Yiling Xu, Zhu Li 0001 |
VCIP | 1 |