VLDB 2026 Research / reviewers in the wild / expert
Shupan Li
dblp:171/0997
· DBLP profile ↗
24ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-5823-2037ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing Multimodal Domain Generalization via Gradient Modulation and ProjectionabstractMultimodal Domain Generalization (MMDG) leverages the complementary strengths of multiple modalities to enhance model generalization on unseen domains. A central challenge in multimodal learning is optimization imbalance, where modalities converge at different speeds during training. This imbalance leads to unequal gradient contributions, allowing some modalities to dominate the learning process while others lag behind. Existing balancing strategies typically regulate each modality’s gradient contribution based on its classification performance on the source domain to alleviate this issue. However, relying solely on source-domain accuracy neglects a key insight in MMDG: modalities that excel on the source domain may generalize poorly to unseen domains, limiting cross-domain gains. To overcome this limitation, we propose Gradient Modulation Projection (GMP), a unified strategy that promotes balanced optimization in MMDG. GMP first decouples gradients associated with classification and domain-invariance objectives. It then modulates each modality’s gradient based on semantic and domain confidence. Moreover, GMP dynamically adjusts gradient projections by tracking the relative strength of each task, mitigating conflicts between classification and domain-invariant learning within modality-specific encoders. Extensive experiments demonstrate that GMP achieves state-of-the-art performance and integrates flexibly with diverse MMDG methods, significantly improving generalization across multiple benchmarks. Hongzhao Li, Guohao Shen, Shupan Li, Mingliang Xu 0001, Muhammad Haris Khan |
AAAI | 3 |
| 2026 | Congestion-Aware Evidence-Driven Multi-agent Path Finding
Bingqian Chen, Hongzhao Li, Xiangrong Zhong, Shupan Li |
ICIC (2) | 7 |
| 2026 | Transformer and Hypernetwork Enhanced Multi-agent Reinforcement Learning for Multi-depot Vehicle Routing
Xianglong Shen, Tianliang Gao, Chenyang Dong, Zhipeng Xia, Chaochao Li, Hongzhao Li, Chendi Ning, Shupan Li |
ICIC (2) | 8 |
| 2026 | SpecPrompt: Enhancing Few-Shot Generalization in CLIP Prompt Learning via Spectral Priors
Hualei Wan, Chaosen Zhao, Bingqian Chen, Hongzhao Li, Shupan Li |
ICIC (1) | 6 |
| 2026 | Sparse 3D Object Detection via Local Geometric Refinement and Dynamic Context Perception
Bingxi Chen, Xuemeng Li, Mi Guo, Mingyuan Jiu, Shupan Li |
ICPR (9) | 7 |
| 2026 | An adaptive local outlier detection approach in evolving data streams
Shubin Su, Juan Duan, Shupan Li, Rubin Zheng, Xingwang Huang |
Neurocomputing | 5 |
| 2026 | Deep Convolutional Primal-Dual Network for Image DeblurringabstractImage deblurring is a challenging image task, which is regarded as a classical inverse problem. Deep primal-dual proximal network (DeepPDNet) is recently proposed which unrolls the Condat-Vũ primal-dual splitting algorithm as a feed-forward network and it has demonstrated excellent restoration performance. However, the feature patterns in the DeepPDNet are well manually designed and thus the network is not implemented in an efficient convolutional fashion. In this work, we revisit the DeepPDNet and extend it in three respects: i) the convolution and pooling operators as well as their associating adjoint operations are studied in the primal-dual algorithm, and then a deep convolutional primal-dual network (DeepConvPDNet) and its full variant with skips are proposed to preserve the optimization consistence of primal-dual Condat-Vũ algorithm; ii) two (cascade vs parallel) variants of the networks are designed according to the structure of convolutional kernels; iii) rather than that the blur kernels are given as prior knowledge, they can be encoded by a set of convolutional layers and deconvolutional layers for their conjugate, resulting to a full learnable deep convolutional primal-dual neural network.We investigate the proposed networks on the MNIST dataset, the grayscale and color version of BSD dataset and GoPro dataset for image deblurring. Extensive experiments are conducted to validate the performance of the proposed networks, and promising results in term of PSNR and SSIM are obtained in comparison with twelve methods including state-of-the-art methods (e.g. Restormer, DRUNet, and DeblurGAN), which validated its effectiveness. Mingyuan Jiu, Mingjing Peng, Fanfan Zhang, Shupan Li, Hongru Zhao, Rongrong Ji, Mingliang Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Improve QMIX from CBS Intervention Guide and Curiosity Mechanism for Multi-Agent Path Finding
Quanjin Wang, Hanlin Zhu, Bingqian Chen, Shupan Li |
ICIC (20) | 6 |
| 2025 | Multi-Modality and Multi-Grained Transformer for Accurate Radiology Report Generation
Hongzhao Li, Liangzhi Zhang, Xiangrong Zhong, Jingpu Zhang, Shupan Li |
ICIC (27) | 6 |
| 2025 | BiDAFuse: Bimodal Differences-Aware Attentive Network for Infrared and Visible Image Fusion
Keyu Sun, Mingyuan Jiu, Shupan Li, Hongru Zhao |
ICIG (3) | 3 |
| 2025 | AS-Mamba: Multimodal Semantic Segmentation via Mamba with Asymmetric Cross and Spatial PerceptionabstractMultimodal semantic segmentation integrates information from multiple data sources to perform pixel-wise semantic classification of images. By fusing complementary information from modalities such as thermal infrared and depth with traditional RGB, robust and reliable predictions can be achieved. However, existing methods still face challenges such as insufficient utilization of cross-modal information, limitations imposed by local receptive fields, and cross-category semantic confusion. To address these issues, this paper proposes the AS-Mamba framework, which introduces a Mamba structure with linear complexity into the field of image segmentation. Firstly, an Asymmetric Cross Mamba module is constructed to dynamically integrate complementary information from RGB and X modalities through an attention weight allocation strategy and enhanced cross-modal semantic alignment, enabling crossmodal dynamic interaction and global feature fusion. Secondly, a Spatial Perception Mamba module is designed to enhance fine-grained feature representation through a spatial-channel attention mechanism. Experiments on RGB-Thermal and RGBDepth tasks demonstrate the superiority of AS-Mamba. Guohao Shen, Hongzhao Li, Xianglong Shen, Jingya Dong, Shupan Li |
ICPADS | 6 |
| 2025 | SE-D3FNet: A LiDAR-Camera Fusion Network with SE Attention and Dynamic 3D Focal Loss for 3D Object Detectionabstract3D object detection is a critical problem in the field of computer vision, widely applied in autonomous driving, robotic navigation, and other domains. Although modern detectors have achieved success in singlesensor object detection, they remain vulnerable to complex environments due to the limitations of single-sensor modalities. We propose SE-D3FNet, a multi-modal fusion framework for 3D object detection that integrates Squeeze-and-Excitation (SE) channel attention and a dynamic 3D focal loss to significantly improve detection accuracy. We present an enhanced feature extraction network termed SE-ResBlock, which demonstrates superior capability in capturing global contextual information. The loss function is also optimized to more accurately capture targets with poor recognition rates. Experimental results on the KITTI benchmark demonstrate that our proposed 3D object detection algorithm achieves superior performance for the car category compared to existing methods. Mingyuan Jiu, Shupan Li, Hongru Zhao, Mingliang Xu 0001 |
ICPADS | 4 |
| 2025 | PD-YOLOv11s: An End-to-End Paper Surface Detection for Specific Visible Angle DefectabstractSurface defect detection plays a critical role in the paper manufacturing process. However, some defects are only visible from a specific angle, which challenges accurate defect recognition. We propose an innovative video defect dataset and an end-to-end detection method named PD-YOLOv11s to address this issue. We use frame differencing and Gaussian background subtraction in the defect dataset to extract inter-frame information from the video. For PD-YOLOv11s, we improve YOLOv11s by the following: (1) PBottleneck replaces the C3k2 structure to reduce the number of parameters, (2) the DSK attention mechanism is added to the end of each backbone output module to extract features better. PD-YOLOv11s achieves the following performance metrics: 8.7M parameters, 95.2% recall, 95.0% precision, 95.1% F1 score, 98.5% mAP50, and 62.8% mAP50:95. Compared to other methods (SSD, FCOS, Faster-RCNN, etc.), this approach significantly improves both accuracy and parameter efficiency, demonstrating its effectiveness in surface defect detection. Shupan Li, Hanlin Zhu, Xiangrong Zhong, Mingyuan Jiu, Mingliang Xu 0001 |
IJCNN | 1 |
| 2025 | Towards Robust Multimodal Domain Generalization via Modality-Domain Joint Adversarial TrainingabstractMultimodal Domain Generalization (MMDG) aims to enhance the robustness of multimodal models against distribution shifts in unseen target domains. Unlike unimodal domain generalization methods, which primarily focus on mitigating domain bias within individual modalities, MMDG faces unique challenges, notably modality heterogeneity (divergent feature spaces) and stability discrepancy (varying sensitivity to domain shifts). To tackle these challenges, we propose Modality-Domain Joint Adversarial Training, a unified framework that addresses these challenges through two key innovations: (1) a tri-discriminator adversarial module that mitigates domain biases in both modality-specific and multimodal representations, while suppressing modality-heterogeneous patterns in the representation space; and (2) a stability-aware dynamic weighting mechanism that adaptively balances modality contributions based on cross-domain stability, reducing reliance on unstable modalities. Additionally, we provide the first theoretical error bound for MMDG, offering a theoretical foundation that supports the effectiveness of our approach. Our approach achieves state-of-the-art performance on the EPIC-Kitchens and HAC datasets while using 75.2% fewer parameters than previous MMDG methods. The source code is available at https://github.com/lihongzhao99/MMDG-Joint-Adversarial-Training. Hongzhao Li, Hualei Wan, Liangzhi Zhang, Mingyuan Jiu, Shupan Li, Mingliang Xu 0001, Muhammad Haris Khan |
ACM Multimedia | 5 |
| 2025 | LGGFormer: A dual-branch local-guided global self-attention network for surface defect segmentation
Yang Lu 0016, Xiaoheng Jiang, Shaohui Jin, Shupan Li, Mingliang Xu 0001 |
Adv. Eng. Informatics | 5 |
| 2025 | RRGMambaFormer: A hybrid Transformer-Mamba architecture for radiology report generation
Hongzhao Li, Siwei Liu 0001, Xiaoheng Jiang, Mingyuan Jiu, Yang Lu 0016, Shupan Li, Mingliang Xu 0001 |
Expert Syst. Appl. | 8 |
| 2024 | Semi-Adaptive Synergetic Two-Way Pseudoinverse Learning System
Binghong Liu, Shupan Li |
PRCV (4) | 3 |
| 2024 | Hierarchical symmetric cross entropy for distant supervised relation extraction
Xiaoheng Jiang, Pengshuai Lv, Yang Lu 0016, Shupan Li, Kunli Zhang, Mingliang Xu 0001 |
Appl. Intell. | 5 |
| 2022 | ZM-CTC: Covert timing channel construction method based on zigzag matrix
Shupan Li, Shen-Gang Hao, Yuanzhang Li 0001 |
Comput. Commun. | 2 |
| 2020 | Boosting performance of virtualized desktop infrastructure with physical GPU and SPICE
Shupan Li, Chungang Shi, Liequan Che, Changyou Zhang, Yuanzhang Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2019 | Deeper Monocular Depth Prediction via Long and Short Skip ConnectionabstractThis paper presents a fully convolutional neural network to tackle the mapping between single view RGB images and depth maps. To regress the depth maps from monocular images, we leverage deep short skip connections in residual learning for extracting features rather than using hand-crafted features. We further propose long skip connections in up-sampling stage to reuse the feature maps which is proved to enhance the result experimentally. To show the impact of loss functions in monocular depth map predictions, we train our model with kind of loss functions and compare the results qualitatively and quantitatively. The proposed model outperforms all current state-of-the-art results with less training data as well as less than half of training epochs in two standard benchmark data sets without any post-processing procedures or other refinement steps. Zhaokai Wang, Rongbin Xu, Shubin Su, Shupan Li |
IJCNN | 5 |
| 2019 | A novel integrity measurement method based on copy-on-write for region in virtual machine
Shupan Li, Shubin Su |
Future Gener. Comput. Syst. | 1 |
| 2016 | MBFS: a parallel metadata search method based on Bloomfilters using MapReduce for large-scale file systems
Zhisheng Huo, Qiaoling Zhong, Shupan Li, Shouxin Wang, Lihong Fu |
J. Supercomput. | 4 |
| 2015 | A Metadata Cooperative Caching Architecture Based on SSD and DRAM for File Systems
Zhisheng Huo, Qiaoling Zhong, Shupan Li, Shouxin Wang, Lihong Fu |
ICA3PP (2) | 4 |