VLDB 2026 Research / reviewers in the wild / expert
Yansheng Qiu
dblp:283/7100
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-5619-0902ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 6 |
| 2025 | Does Adding a Modality Really Make Positive Impacts in Incomplete Multi-Modal Brain Tumor Segmentation?abstractPrevious incomplete multi-modal brain tumor segmentation technologies, while effective in integrating diverse modalities, commonly deliver under-expected performance gains. The reason lies in that the new modality may cause confused predictions due to uncertain and inconsistent patterns and quality in some positions, where the direct fusion consequently raises the negative gain for the final decision. In this paper, considering the potentially negative impacts within a modality, we propose multi-modal Positive-Negative impact region Double Calibration pipeline, called PNDC, to mitigate misinformation transfer of modality fusion. Concretely, PNDC involves two elaborate pipelines, Reverse Audit and Forward Checksum. The former is to identify negative regions impacts of each modality. The latter calibrates whether the fusion prediction is reliable in these regions by integrating the positive impacts regions of each modality. Finally, the negative impacts region from each modality and miss-match reliable fusion predictions are utilized to enhance the learning of individual modalities and fusion process. It is noted that PNDC adopts the standard training strategy without specific architectural choices and does not introduce any learning parameters, and thus can be easily plugged into existing network training for incomplete multi-modal brain tumor segmentation. Extensive experiments confirm that our PNDC greatly alleviates the performance degradation of current state-of-the-art incomplete medical multi-modal methods, arising from overlooking the positive/negative impacts regions of the modality. The code is released at PNDC. Yansheng Qiu, Kui Jiang, Hongdou Yao, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | MonOri: Orientation-Guided PnP for Monocular 3-D Object DetectionabstractMonocular 3-D object detection is a challenging task in the field of autonomous driving and has made great progress. However, current monocular image methods tend to incorporate additional information such as pseudolabels to improve algorithm performance while overlooking the geometric relationship between the object's keypoints, resulting in low performance for occluded object detection. To address this issue, we find that introducing the orientation information of objects in the 3-D detection pipeline can help improve the detection performance of occluded objects. An orientation-guided perspective-n-point (PnP) for monocular 3-D object detection method named MonOri is presented in this article, which uses object's orientation to guide keypoints' optimization. Considering the existence of different deformation objects in the scene, we design the feature aggregation detection module (FADM), which consists of the feature focus fusion module (FFFM) and CondConv detection module (CCDM). First, FFFM can highlight signals from irregularly occluded objects, effectively modeling features of elongated and small-sized objects. This module enhances the model's ability to recognize elongated and small-sized objects in complex scenes. Then, the CCDM is designed to improve the network's ability to estimate object keypoints' location regression under occlusion conditions and minimize the network computational overhead. Finally, considering that the unoccluded portions of occluded objects are closely related to the orientation of the objects, an orientation-guided keypoints' selection module (OGKSM) is proposed to enhance the accuracy of objected optimization for keypoint positions and spatial location inference of the object. Experimental results indicate that the MonOri method achieves competitive results; it is also demonstrated that the orientation information is introduced in the PnP algorithm to estimate the object's spatial position that can mitigate the impact of occlusion on object detection, thus improving the recognition rate of occluded objects. Our code is available at https://github.com/DL-YHD/MonOri. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Yansheng Qiu, Xiao Wang 0029, Yimin wang, Xiaoyu Chai, Chenglong Cao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Devil is in Details: Locality-Aware 3D Abdominal CT Volume Generation for Self-Supervised Organ SegmentationabstractIn the realm of medical image analysis, self-supervised learning (SSL) techniques have emerged to alleviate labeling demands, while still facing the challenge of training data scarcity owing to escalating resource requirements and privacy constraints. Numerous efforts employ generative models to generate high-fidelity, unlabeled 3D volumes across diverse modalities and anatomical regions. However, the intricate and indistinguishable anatomical structures within the abdomen pose a unique challenge to abdominal CT volume generation compared to other anatomical regions. To address the overlooked challenge, we introduce the Locality-Aware Diffusion (Lad), a novel method tailored for exquisite 3D abdominal CT volume generation. We design a locality loss to refine crucial anatomical regions and devise a condition extractor to integrate abdominal priori into generation, thereby enabling the generation of large quantities of high-quality abdominal CT volumes essential for SSL tasks without the need for additional data such as labels or radiology reports. Volumes generated through our method demonstrate remarkable fidelity in reproducing abdominal structures, achieving a decrease in FID score from 0.0034 to 0.0002 on AbdomenCT-1K dataset, closely mirroring authentic data and surpassing current methods. Extensive experiments demonstrate the effectiveness of our method in self-supervised organ segmentation tasks, resulting in an improvement in mean Dice scores on two abdominal datasets effectively. These results underscore the potential of synthetic data to advance self-supervised learning in medical image analysis. Yuran Wang 0003, Zhijing Wan, Yansheng Qiu, Zheng Wang 0007 |
ACM Multimedia | 3 |
| 2024 | Occlusion-Aware Plane-Constraints for Monocular 3D Object DetectionabstractThe task of 3D object detection poses a significant challenge for 3D scene understanding and is primarily employed in the fields of robot control and autonomous driving. Monocular-based 3D detection methods are more cost-effective and practical than stereo-based or LiDAR-based methods. Monocular image 3D detection methods have garnered considerable attention from researchers. However, the impact of occlusion scenarios of the objects on the keypoints prediction is often overlooked. To address this issue, the present paper proposes a novel 3D monocular object detection method named MonOAPC, which is equipped with occlusion-aware plane-constraints. This method can adaptively utilize partial keypoints to infer the plane location of the object based on the level of occlusion and highlights that the introduction of plane constraints is advantageous for the 3D detection task. First, the plane information of the object in 3D space is beneficial to optimize the keypoints regression, and considering that each plane holds different significance, an adaptive plane location inference module is proposed to enhance the keypoints location regression. Second, a novel co-depth estimation module is proposed to jointly estimate the object’s spatial location through various depth inference methods, thereby improving the generalization of the object depth estimation. Furthermore, the paper demonstrates that the accuracy of 3D object detection can be indirectly improved by introducing plane information to promote keypoints regression, and that plane information is effective for monocular 3D detection. The experimental outcomes show that the MonOAPC method can attain competitive results. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Scratch Each Other's Back: Incomplete Multi-modal Brain Tumor Segmentation Via Category Aware Group Self-Support LearningabstractAlthough Magnetic Resonance Imaging (MRI) is very helpful for brain tumor segmentation and discovery, it often lacks some modalities in clinical practice. As a result, degradation of prediction performance is inevitable. According to current implementations, different modalities are considered to be independent and non-interfering with each other during the training process of modal feature extraction, however they are complementary. In this paper, considering the sensitivity of different modalities to diverse tumor regions, we propose a Category Aware Group Self-Support Learning framework, called GSS, to make up for the information deficit among the modalities in the individual modal feature extraction phase. Precisely, within each prediction category, predictions of all modalities form a group, where the prediction with the most extraordinary sensitivity is selected as the group leader. Collaborative efforts between group leaders and members identify the communal learning target with high consistency and certainty. As our minor contribution, we introduce a random mask to reduce the possible biases. GSS adopts the standard training strategy without specific architectural choices and thus can be easily plugged into existing incomplete multi-modal brain tumor segmentation. Remarkably, extensive experiments on BraTS2020, BraTS2018, and BraTS2015 datasets demonstrate that GSS can improve the performance of existing SOTA algorithms by 1.27-3.20% in Dice on average. The code is released at https://github.com/qysgithubopen/GSS. Yansheng Qiu, Delin Chen, Hongdou Yao, Yongchao Xu, Zheng Wang 0007 |
ICCV | 1 |
| 2023 | Modal-aware Visual Prompting for Incomplete Multi-modal Brain Tumor SegmentationabstractIn the realm of medical imaging, distinct magnetic resonance imaging (MRI) modalities can provide complementary medical insights. However, it is not uncommon for one or more modalities to be absent due to image corruption, artifacts, acquisition protocols, allergies to contrast agents, or cost constraints, posing a significant challenge for perceiving the modality-absent state in incomplete modality segmentation.In this work, we introduce a novel incomplete multi-modal segmentation framework called Modal-aware Visual Prompting (MAVP), which draws inspiration from the widely used pre-training and prompt adjustment protocol employed in natural language processing (NLP). In contrast to previous prompts that typically use textual network embeddings, we utilize embeddings as the prompts generated by a modality state classifier that focuses on the missing modality states. Additionally, we integrate modality state prompts into both the extraction stage of each modality and the modality fusion stage to facilitate intra/inter-modal adaptation. Our approach achieves state-of-the-art performance in various modality-incomplete scenarios compared to incomplete modality-specific solutions. Yansheng Qiu, Ziyuan Zhao, Hongdou Yao, Delin Chen, Zheng Wang 0007 |
ACM Multimedia | 1 |
| 2023 | Vertex points are not enough: Monocular 3D object detection via intra- and inter-plane constraints
Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
Neural Networks | 6 |
| 2021 | Illuminate Low-Light Image via Coarse-to-fine Multi-level Network
Yansheng Qiu, Jun Chen 0001, Xiao Wang 0029, Kui Jang |
MMM (1) | 1 |
| 2021 | Spatio-Spectral Feature Fusion for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visual quality, which is essential for many downstream computer vision and multimedia tasks. Existing spatial-domain low-light enhancement methods barely focus on the regions containing object boundaries, which take the most informative characteristics. However, solely focusing on enhancing high-frequency details not only causes over-sharpening of an image but also leads to color distortion. In this paper, we propose a novel spatio-spectral feature fusion network (S2F2N), that involves a frequency-feature representation branch (FRB) and a spatial-feature representation branch (SRB) to learn the domain-specific representation individually. Moreover, a spatial-channel mixed attention block (MAB) is introduced to learn the joint representation of spatio-spectral features for final image relighting. Extensive experiments on several benchmark datasets demonstrate that our method can produce high fidelity results for low-light images. Yansheng Qiu, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Chia-Wen Lin |
IEEE Signal Process. Lett. | 1 |