Zhaopan Xu

dblp:278/2033 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0009-4985-0528ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2026 MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education.
Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang
AAAI4
2026 Noisy correspondence decomposition for robust composed image retrieval
Zhaopan Xu, Chenrui Zhou, Xi Chen 0110, Wangbo Zhao, Jianning Zhang, Xiaojiang Peng, Hongxun Yao, Kaipeng Zhang
Expert Syst. Appl.1
2026 Dual-prompt-based binary matching for open set domain adaptation
Jidong Yang, Shouxu Jiang, Hongxun Yao, Lingji Xu, Sheng Jin 0002, Huicong Zhang, Zhaopan Xu
Neurocomputing7
2026 Training-Free Noisy Correspondence Rectification With Multimodal Conceptual Knowledge
Zhaopan Xu, Wangbo Zhao, Xi Chen 0110, Sheng Jin 0002, Hongxun Yao
IEEE Signal Process. Lett.1
2026 Learning With Dual Noisy Labels for Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification (TIReID) aims to identify a target person from a given textual description. Although recent work has made significant progress, most of it implicitly assumes that the sample annotations are correct and that the cross-modal correspondence in each image-text pair is well aligned. However, such an assumption requires elaborately annotated datasets, which are expensive and even impossible to obtain in practice. To alleviate this issue, in this letter, we explore a new TIReID setting, termed learning with dual noisy labels, in which the model learns from data with both noisy identity labels and noisy correspondence. We propose a general framework called TDTD (Two stage framework forDual noise ofTIReID) to achieve this. In the first stage, a Noise-Aware Preliminary Learning (NAPL) strategy selects “easy” triplets to train a noise-tolerant initial model. In the second stage, the model leverages reliable representations from NAPL to automatically correct both identity and correspondence errors via soft-label estimation and is then fine-tuned on the entire dataset using a dual noise-robust triplet loss. Extensive experiments on three public benchmarks, CUHK-PEDES, ICFG-PEDES, and RSTPReID, demonstrate the performance and robustness of TDTD, achieving state-of-the-art results under dual noise conditions.
Zhaopan Xu, Wangbo Zhao, Xiaojiang Peng, Hongxun Yao
IEEE Signal Process. Lett.1
2025 Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation
abstract
Visual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on data labeling, domain adaptation offers an alternative solution by adapting models trained on labeled source data to unlabeled target data. Conventional domain adaptation methods require access to source data. However, due to privacy concerns, source emotional data may be inaccessible. To address this issue, we propose an unexplored task: source-free domain adaptation (SFDA) for VER, which does not have access to source data during the adaptation process. To achieve this, we propose a novel framework termed Bridge then Begin Anew (BBA), which consists of two steps: domain-bridged model generation (DMG) and target-related model adaptation (TMA). First, the DMG bridges cross-domain gaps by generating an intermediate model, avoiding direct alignment between two VER datasets with significant differences. Then, the TMA begins training the target model anew to fit the target structure, avoiding the influence of source-specific knowledge. Extensive experiments are conducted on six SFDA settings for VER. The results demonstrate the effectiveness of BBA, which achieves remarkable performance gains compared with state-of-the-art SFDA methods and outperforms representative unsupervised domain adaptation approaches.
Jiankun Zhu, Sicheng Zhao, Wenbo Tang, Zhaopan Xu, Tingting Han 0003, Pengfei Xu 0001, Hongxun Yao
AAAI5
2025 OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
abstract
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to data size and diversity limitations. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.
Xiaopeng Peng 0001, Jiajun Song, Chuanhao Li 0001, Zhaopan Xu, Ziyao Guo, Hao Zhang 0117, Yuqi Lin, Yefei He, Lirui Zhao, Xiaojun Chang, Yu Qiao 0001, Wenqi Shao, Kaipeng Zhang
CVPR5
2025 Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic Segmentation
abstract
Panoramic semantic segmentation has garnered increasing attention due to its ability to provide comprehensive environmental perception. However, it requires a large number of annotated panoramic images to achieve satisfactory performance, which is costly. Recently, Domain Adaptation for Panoramic Semantic Segmentation (DA4PASS) has been proposed to reduce the reliance on annotated data by transferring segmentation models trained on annotated pinhole images to unlabelled panoramic images. Previous DA4PASS methods mainly focus on aligning features between pinhole and panoramic images, overlooking the unique appearance characteristics of panoramic images, particularly object distortion. To address the appearance discrepancies between pinhole and panoramic images, we propose Gaussian Constrained Diffeomorphic Deformation Network (GCDDN), which applies a panoramic deformation transformation obtained by Gaussian kernels to the annotated pinhole images. Specifically, GCDDN predicts multiple Gaussian kernels and performs first-order horizontal/vertical differences to obtain a naturally smooth and reversible panoramic deformation field, which is diffeomorphic. Due to its universality, GCDDN can be integrated into any domain adaptation (DA) method. Extensive experimental results demonstrate that integrating GCDDN leads to substantial improvements in both DA methods for pinhole images and those specifically designed for panoramic images, with a maximum gain of 1.80% in outdoor scenarios. Code is available at https://github.com/jingjiang02/GCDDN.
Jiankun Zhu, Zhaopan Xu, Xi Chen 0110, Sicheng Zhao, Hongxun Yao
ICASSP3
2025 Open-Vocabulary Visual Emotion Adaptation via Prompt Learning
abstract
Visual Emotion Recognition (VER) aims to identify emotions from visual content and has garnered significant attention in recent years due to its wide-ranging applications. Although deep learning-based methods have shown success in VER, they require extensive labeled data, which is costly. Unsupervised Domain Adaptation (UDA) methods can reduce reliance on annotated data by transferring models trained on labeled datasets to unlabeled data. However, these methods assume that the source and target domains share the same label space. In practice, this assumption is often violated due to the inherent ambiguity and subjectivity in emotion labeling. To address this limitation, we propose a novel prompt learning paradigm for open-vocabulary visual emotion UDA, termed Domain-specific Ensemble Prompting (DSEP). DSEP leverages psychological emotion models to unify emotion labels into a common space in an ensemble manner, enhancing the open-vocabulary capabilities of UDA. It then combines ensemble label prompts with domain-specific content prompts to achieve open-vocabulary UDA. To our knowledge, we are the first to explore open-vocabulary adaptation for VER. Extensive experiments demonstrate that DSEP consistently outperforms state-of-the-art methods across four public benchmarks.
Zhaopan Xu, Sicheng Zhao, Xiaojiang Peng, Hongxun Yao
ICASSP1
2025 Learning Class Prototypes for Visual Emotion Recognition
abstract
Visual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1) semantic gap: the large affective gap between visual clues and the emotional expressions; 2) overfitting: the lack of model robustness due to unclear features in the emotional category samples; 3) label ambiguity: the overlap between categories caused by diverse emotional responses. To address these limitations, we present a novel VER method named ProtoEmotion (PoE), exploring discriminative emotional representations by jointly learning prototypes of textual emotional expressions and visual features. Specifically, text prototypes build explicit textual features for each emotion category by extracting prototypes of learnable prompts from multiple aspects, reducing semantic differences. The visual prototypes capture the most defining image features of each category, providing a more robust and discriminative feature representation, while bringing samples closer together to reduce overfitting. In addition, to alleviate the label ambiguity, we propose a label smoothing algorithm based on the prototype distance. Extensive experiments demonstrate the effectiveness of PoE, which outperforms the state-of-the-art by 1.37% on FI and 1.52% on EmotionROI datasets.
Jiankun Zhu, Sicheng Zhao, Zhaopan Xu, Wenbo Tang, Hongxun Yao
ICASSP4
2025 ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges
Jiaxin Ai, Zhaopan Xu, Fanrui Zhang, Zizhen Li, Yukang Feng, Baojin Huang, Zhongyuan Wang 0001, Kaipeng Zhang
ICCV3
2025 DreamAnimate: Temporal Consistency and Detail Preservation for Character Animation
abstract
Character animation aims to generate realistic, high-quality videos from a reference image and target frames. However, existing methods struggle to balance fine-grained detail preservation with temporal consistency. This limitation results in artifacts like flickering and unrealistic deformations, especially in facial and hand regions. To address these challenges, we propose DreamAnimate, a novel framework that synthesizes temporally consistent and detail-rich animations. DreamAnimate integrates three modules: the Progressive Motion Estimation module ensures accurate motion alignment and temporal stability by refining keypoint heatmaps, the Global Affine Transformation module generates dense motion flows to handle complex motions and occlusions, and the Character Animation Fusion module combines intermediate synthesis using a UNet architecture and an Animation Fusion Network to produce high-quality animations. Extensive experiments demonstrate that DreamAnimate outperforms state-of-the-art methods, achieving superior fidelity and effectively capturing intricate facial expressions and hand movements. Code and models will be released at https://github.com/cslltian/DreamAnimate in the near future.
Lulu Tian, Hongxun Yao, Zhaopan Xu, Jiankun Zhu, Xi Chen 0110, Yuxin Hou
ICME3
2025 Sekai: A Video Dataset towards World Exploration
abstract
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications.
Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang
NeurIPS7
2024 Dataset Growth
Ziheng Qin, Zhaopan Xu, Zangwei Zheng, Zebang Cheng, Hao Tang 0005, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang 0036, Yang You 0001
ECCV (9)2
2024 InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data Pruning
abstract
Data pruning aims to obtain lossless performances with less overall cost. A common approach is to filter out samples that make less contribution to the training. This could lead to gradient expectation bias compared to the original data. To solve this problem, we propose InfoBatch, a novel framework aiming to achieve lossless training acceleration by unbiased dynamic data pruning. Specifically, InfoBatch randomly prunes a portion of less informative samples based on the loss distribution and rescales the gradients of the remaining samples to approximate the original gradient. As a plug-and-play and architecture-agnostic framework, InfoBatch consistently obtains lossless training results on classification, semantic segmentation, vision pertaining, and instruction fine-tuning tasks. On CIFAR10/100, ImageNet- 1K, and ADE20K, InfoBatch losslessly saves 40% overall cost. For pertaining MAE and diffusion model, InfoBatch can respectively save 24.8% and 27% cost. For LLaMA instruction fine-tuning, combining InfoBatch and the recent coreset selection method (DQ) can achieve 10 times acceleration. Our results encourage more exploration on the data efficiency aspect of large model training. Code is publicly available at NUS-HPC-AI-Lab/InfoBatch.
Ziheng Qin, Kai Wang 0036, Zangwei Zheng, Jianyang Gu, Zhaopan Xu, Daquan Zhou, Baigui Sun, Xuansong Xie, Yang You 0001
ICLR6
2023 BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency
abstract
As one of the most fundamental techniques in multi-modal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose a general framework called BiCro (Bidirectional Cross-modal similarity consistency), which can be easily integrated into existing cross-modal matching models and improve their robustness against noisy data. Specifically, BiCro aims to estimate soft labels for noisy data pairs to reflect their true correspondence degree. The basic idea of BiCro is motivated by that – taking image-text matching as an example – similar images should have similar textual descriptions and vice versa. Then the consistency of these two similarities can be recast as the estimated soft labels to train the matching model. The experiments on three popular cross-modal matching datasets demonstrate that our method significantly improves the noise-robustness of various matching models, and surpass the state-of-the-art by a clear margin. The code is available at https://github.com/xu5zhao/BiCro.
Shuo Yang 0006, Zhaopan Xu, Kai Wang 0036, Yang You 0001, Hongxun Yao, Tongliang Liu, Min Xu 0001
CVPR2
2020 Stroke controllable style transfer based on dilated convolutions
abstract
Transferring a photo to a stylised image with beautiful texture has become one of the most popular topics in computer vision and the application of image processing. Controlling the stroke size of the texture is one of the challenging problems in this task. Recent representative methods for such problem introduce a pyramid model to regulate receptive fields in the network. Meanwhile, dilated convolutions are proved to be a very efficient way to adjust receptive fields without losing resolution. By combining the advantages of both approaches and making special optimisation for VGG19 model for style transfer tasks, the authors propose to exploit dilated convolutions to extract texture information endowing the network with stroke controllable. Several sets of contrast experiments were conducted and results show that their algorithm can generate more attractive stylisation images and control stroke size flexibly. It demonstrates the superiority of applying dilated convolutions as a texture extraction method for maintaining more texture information and controlling stroke size.
Zhaopan Xu, Yu Zhang 0040, Kang Li 0005, Shengling Geng
IET Comput. Vis.1