VLDB 2026 Research / reviewers in the wild / expert
Shiyi Mu
dblp:299/7608
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stereo-Based 3-D Anomaly Object Detection for Autonomous Driving: A New Dataset and Baseline
Shiyi Mu, Zichong Gu, Hanqi Lyu, Shugong Xu |
IEEE Internet Things J. | 1 |
| 2026 | Affective computing in the era of large language models: A survey from the NLP perspective
Xiaocui Yang, Xingle Xu, Zeran Gao, Shiyi Mu, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Kaisong Song, Ge Yu 0001 |
Knowl. Based Syst. | 6 |
| 2026 | StereoDETR: Stereo-Based Transformer for 3D Object DetectionabstractCompared to monocular 3D object detection, stereo-based 3D methods offer significantly higher accuracy but still suffer from high computational overhead and latency. The state-of-the-art stereo 3D detection method achieves twice the accuracy of monocular approaches, yet its inference speed is only half as fast. In this paper, we propose StereoDETR, an efficient stereo 3D object detection framework based on DETR. StereoDETR consists of two branches: a monocular DETR branch and a stereo branch. The DETR branch is built upon 2D DETR with additional channels for predicting object scale, orientation, and sampling points. The stereo branch leverages low-cost multi-scale disparity features to predict object-level depth maps. These two branches are coupled solely through a differentiable depth sampling strategy. To handle occlusion, we introduce a constrained supervision strategy for sampling points without requiring extra annotations. Compared with the existing published monocular and binocular 3D detection methods, StereoDETR breaks the trade-off between speed and accuracy. Through a concise framework, it achieves binocular-level accuracy while maintaining monocular-level inference speed. The code is available at https://github.com/shiyi-mu/StereoDETR-OPEN. Shiyi Mu, Zichong Gu, Zhiqi Ai, Shugong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction-Tuning for Emotion-Cause Pair Extraction
Shiyi Mu, Yongkang Liu 0002, Shi Feng 0001, Xiaocui Yang, Daling Wang, Yifei Zhang 0003 |
CogSci | 1 |
| 2025 | Two-stage model re-optimization and application in face recognition
Jianyu Qian, Shiyi Mu, Hengjie Lu, Shugong Xu |
Neurocomputing | 2 |
| 2025 | Uni-EPM: A Unified Extensible Perception Model Without Labeling EverythingabstractMulti-task perception system to simultaneously perceive various kinds of objects is essential for autonomous driving. Existing perception frameworks always rely on multi-labeled datasets, which encompass labels for all pertinent objects, thereby constraining their adaptability to leverage specialized, task-oriented datasets. This approach hinders the efficient utilization of abundant but focused data. Furthermore, stacking multiple expert networks to address these perception objectives inevitably introduces additional computational overhead. To address this limitation, we propose Uni-EPM (Unified Extensible Perception Model), with a novel training framework for multi-task perception using task prompt selection to decouple tasks, which enables perceiving traffic signs and traffic lights in addition to lane lines and traffic elements from existing task-specific datasets without re-labeling. To the best of our knowledge, Uni-EPM is the first model can do this in the field of autonomous driving. By introducing the parameter-sharing decoder among tasks, we alleviate the problems of stacking task heads, including significant parameter increase, etc. Uni-EPM achieves state-of-the-art results in multi-task algorithms without substantial increase in parameters, which also demonstrates comparable performance to existing standalone models. The efficiency of the design is validated through comprehensive ablation experiments and results. Shiyi Mu, Shugong Xu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | A Learnable Color Correction Matrix for RAW Reconstruction
Shiyi Mu, Shugong Xu |
BMVC | 2 |
| 2024 | Toward Unified End-to-End License Plate Detection and Recognition for Variable Resolution RequirementsabstractIn this paper, we present a new cascade architecture based on a differentiable sample module to satisfy the varied image resolution requirements of license plate detector and recognizer in end-to-end technologies. Based on this module, the network can detect license plates on downsampled low-resolution images and resample them from the original high-definition images to recognize the license plate numbers. Furthermore, since the optimization direction of the detector for the detection boxes and the input requirements of the recognizer are not consistent with each other, we introduce the Bias Detection Head, which decouples the two Bounding Boxes to circumvent this problem. In the meantime, a novel feature fusion module is presented, which simultaneously satisfies the fusion of multi-scale information and the interaction of two Bounding Box features. For the recognizer, we present a unified architecture based on a decoupled attention mechanism for recognizing single and double lines, varying lengths, and tilting on license plates. Shiyi Mu, Shugong Xu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Toward Reliable License Plate Detection in Varied Contexts: Overcoming the Issue of Undersized Plate AnnotationsabstractLicense plate detection and recognition (LPDR) is of paramount importance in the Intelligent Transportation Systems. Most existing license plate (LP) detectors rely on anchors, rendering them vulnerable to multi-scale LPs, especially those of smaller scales, which limited the overall performance of LPDR. Another issue prevalent in LP datasets arises from the inherent ambiguity in manual labeling standards. Owing to this uncertainty, certain small-scale LPs that are distinctly detectable often suffer from annotation omissions. The presence of such noisy data has a detrimental impact on the training of LP detectors. In this paper, we propose ALPD, an anchor-free LP detector along with three key designs, namely the Multi-To-One scale-fusion block (MTO) for cross-scale feature integration, the Multi-Domain Feature Simulation (MDFS) for narrowing the disparities across multiple domains even the unseen ones, and the decoupled heads for better optimizing classification and regression tasks. Besides, ALPD incorporates a semi-supervised training framework using an abstention strategy known as arbitration, wherein a Teacher model and a Student model are trained collaboratively, enabling the supplementation of missing annotations for small license plates. Furthermore, it possesses immunity to model performance degradation when fed with massive quantities of unlabeled or even mislabeled data. The arbitration method along with a penalty factor can effectively guarantee the pseudo-label quality and balance complexity between the Teacher and Student tasks, thus preventing the Student from being constrained by ambiguous pseudo-labels. ALPD outperforms previous state-of-the-art methods on two widely recognized benchmarks and exhibits its robustness and generalizability on the All-round CCPD dataset. Zhongxing Peng, Shiyi Mu, Shugong Xu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | GroupPlate: Toward Multi-Category License Plate RecognitionabstractLicense Plate Detection and Recognition (LPDR) is widely used in Intelligent Transportation Systems (ITS). Although there are typically multiple categories of license plates, the majority of existing research cannot be applied to multi-category plates due to that existing methods are not optimised for multi-category plate scenarios and the scarcity of large-scale multi-category plate datasets. In this paper, we propose a multi-category license plate recognition framework called GroupPlate, which consists of Group Module and Indirect Supervision Module, making full use of the implicit and explicit grouping information of license plate. In addition, the Category Decouple Module is intended to decouple the grouping information from the original features, allowing the decoder to concentrate on character features. Simultaneously, we propose a large-scale All-Category license Plate detection and recognition Dataset (ACPD) for vehicles on the Chinese mainland, which also includes annotations of plates’ categories. Considering the domain gap between synthetic data and real data, we propose a simple but effective strategy called Feature Shift to mitigate the performance degradation caused by this gap. Experiments demonstrate that GroupPlate achieves the comparable performance to the existing methods on single-category license plate dataset and outperforms our baseline on the multi-category license plates dataset. Ablation experiments demonstrate the effectiveness of the modules in GroupPlate. Extensive results demonstrate that the dataset we proposed can mitigate the problem of models trained on a single-category license plate dataset failing to recognize multi-category license plates, and that our model can generalizes well to unseen categories. The work will be available athttps://github.com/YilinGao-SHU/ACPD. Hengjie Lu, Shiyi Mu, Shugong Xu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | IFR: Iterative Fusion Based Recognizer for Low Quality Scene Text Recognition
Zhiwei Jia, Shugong Xu, Shiyi Mu, Yue Tao, Shan Cao 0001 |
PRCV (2) | 3 |