EDBT 2026 Demo / reviewers in the wild / expert
Chunmian Lin
dblp:298/4684
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-9051-8852ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | R-PERL: Resilience-Oriented Mixed Platoon Control Under Traffic Oscillations and Communication Delays
Jianhong Liang, Xuting Duan, Sifan Wu 0004, Jianshan Zhou, Kaige Qu, Chunmian Lin, Ling Wang 0001, Daxin Tian |
IEEE Internet Things J. | 6 |
| 2026 | DCGM: Synergizing Differential Cross-Modal Fusion and Graph-Mamba for Referring Multi-Object Tracking
Zhibiao Xue, Zhanwen Liu, Chunmian Lin, Daniel Jian Sun |
Knowl. Based Syst. | 4 |
| 2026 | SSM-Det: State Space Model-Based Object Detector for Intelligent Transportation SystemabstractThe State Space Model (SSM) has been a growth of interest in computer vision due to its long-term dependency modeling with linear complexity. Despite massive endeavor, it has not been extensively explored in intelligent transportation system (ITS) yet. In this paper, we propose State Space Model-based object Detector (SSM-Det), that is meticulously curated with Direction-aware Visual State Space Encoder (D-VSSE). Specifically, it customizes multi-path pixel exchange and patch re-arrangement via four-direction scanning mechanism, promoting for information communication. To bridge the information bottleneck across high-low level, we further design Split-Fusion (SF) and Skip-Connection (SC) modules for contextual feature propagation before decoding: SF performs multi-channel semantic separation and re-weighting in global-local scope, while SC is responsible for cross-layer feature interaction in a cascaded manner. Empirical studies is conducted on both VisDrone2019-DET and SEU_PML benchmarks, and our proposed SSM-Det reports the state-of-the-art performance against all counterparts by a substantial margin, while maintaining the real-time inference speed. We hope this work contributes to the in-depth investigation of SSM-based detector for intelligent transportation applications. The code is available athttps://buaawjq.github.io/SSM-Det/. Chunmian Lin, Kan Guo, Jiangang Guo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Cooperative Coverage Mission Planning for Multi-UAV Based on the Dual-Ring Dynamic SchedulerabstractUnmanned aerial vehicles (UAVs) have rapidly advanced in applications such as disaster response, infrastructure inspection, and smart city systems. However, cooperative coverage mission planning in dynamic environments poses a persistent challenge in balancing global optimization with real-time responsiveness. To address this, we propose a two-stage task allocation framework based on a dual-ring dynamic scheduler. In the centralized planning phase, an enhanced NSGA-II algorithm is developed, incorporating dual-population initialization, path-exchange crossover, and multi-strategy mutation. Experimental results demonstrate a 49% improvement in the hypervolume indicator over the baseline NSGA-II, with average reductions of 13.33% and 23.71% in total and maximum execution times, respectively. In the dynamic scheduling phase, we design a distributed auction mechanism leveraging a dual-ring communication topology, capable of handling six types of events: UAV failure, UAV addition, node cancel exploration, node urgent exploration, node re-exploration and node in-depth exploration. Through event-driven auctions and group-based bidding, the system maintains a load imbalance under 28% and achieves effective rebalancing even under scenarios with over 50% UAV loss. These results validate the robustness and adaptability of the dual-ring dynamic scheduler in real-time collaborative coverage missions. The proposed method demonstrates significant potential in dynamic, large-scale UAV applications. Source code will be available at: https://github.com/GradualScholar/CoverageMissionPlannin.git Yongzhuo Yu, Xuting Duan, Feiyang Zhao, Jianshan Zhou, Chunmian Lin, Kaige Qu, Daxin Tian |
IEEE Internet Things J. | 5 |
| 2025 | DMP: Difference-Guided Motion Prediction for Vision-Centric Autonomous DrivingabstractVision-centric motion prediction concentrates on accurately determining the instance mask and its future trajectory from surround-view cameras, which manifests inherent merits such as holistic perspective and fully-differentiable spirit. Nonetheless, it is still impeded by sparse bird’s-eye view (BEV) representation and unfavorable temporal context across frames, resulting in a sub-optimal solution to decision-making and vehicle navigation. In this work, we propose a novelDifference-guideMotionPrediction for vision-centric autonomous driving, that is DMP, where it integrates BEV map refinement with spatial-temporal relation modeling in a hierarchical manner. Specifically, a bidirectional view projection strategy is introduced for the complementary BEV feature generation via depth-consistency correction. To promote spatiotemporal context aggregation, we design a difference-guided motion approach by offset approximation to align motion-aware cues between adjacent frames, and a dual-stream pyramid module is further developed for historical information fusion and future instance segmentation during specific durations. Extensive experiments on the large-scale nuScenes dataset demonstrate that it outperforms the baselines by a remarkable margin and delivers competitive motion prediction across diverse scenarios and range settings, suggesting its effectiveness and superiority. The details will be available athttps://github.com/pupu-chenyanyan/DMP-VAD. Chunmian Lin, Xuting Duan, Jianshan Zhou, Kan Guo, Dezong Zhao, Dongpu Cao, Daxin Tian |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | CUDA-X: Unsupervised Domain-Adaptive Vehicle-to-Everything Collaboration via Knowledge Transfer and AlignmentabstractRecently emerged vehicle-to-everything (V2X) perception has revealed great potential to overcome the limitation of single-vehicle intelligence aided by vigorous interaction among on-road agents, while prior endeavors are practically developed on parameter-specific simulation or configuration-dynamic real-world setting, overlooking the transferability across various scenarios. In this article, we propose unsupervised domain-adaptive vehicle-to-everything collaboration framework dubbed CUDA-X, which is built on top of a de facto collective model with key-point information exchange and instance adaptation. Specifically, collaborative knowledge transfer (CKT) is responsible for domain-agnostic feature reconstruction from nearby car or infrastructure by spatial-channel pooling operation in an elementwise manner. To promote the candidate alignment, a brand-new bin-based location correction (BLC) provides an auxiliary supervision for cross-dataset box refinement via residual coordinate encoding (RCE), and category-aware pooling alignment (CPA) is further designed for pulling the category-specific instance closer between source and target samples. We benchmark CUDA-X against the counterparts on four prevalent cooperative perception datasets, i.e., OPV2V, V2X-Sim, V2V4Real, and DAIR-V2X: it establishes the new state-of-the-art vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) performances regardless of simulation or reality. We expect that this appealing attempt would provide an in-depth insight into domain generalization in the context of multiagent perception, and the code is publicly available soon. Daxin Tian, Chunmian Lin, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | CFMMC-Align: Coarse-Fine Multi-Modal Contrastive Alignment Network for Traffic Event Video Question AnsweringabstractTraffic video question answering (TrafficVQA) constitutes a specialized VideoQA task designed to enhance the basic comprehension and intricate reasoning capacities of videos, specifically focusing on traffic events. Recent VideoQA models employ pretrained visual and textual encoder models to bridge the feature space gap between visual and textual data. However, in addressing the unique challenges inherent to the TrafficVQA task, three pivotal issues must be addressed: (i) Dimension Gap: Between the pretrained image (appearance feature) and video (motion feature) models, there exists a conspicuous dimension difference in static and dynamic visual data; (ii) Scene Gap: The common real-world datasets and the traffic event datasets differ in visual scene content; (iii) Modality Gap: A pronounced feature distribution discrepancy emerges between traffic video and text data. To alleviate these challenges, we introduce the coarse-fine multimodal contrastive alignment network (CFMMC-Align). This model leverages sequence-level and token-level multimodal features, grounded in an unsupervised visual multimodal contrastive loss to mitigate dimension and scene gaps and a supervised visual-textual contrastive loss to alleviate modality discrepancies. Finally, the model is validated on the challenging public TrafficVQA dataset SUTD-TrafficQA and outperforms the state-of-the-art method by a substantial margin (50.2%compared to46.0%). The code is available at https://github.com/guokan987/CFMMC-Align. Kan Guo, Daxin Tian, Yongli Hu, Chunmian Lin, Jianshan Zhou, Xuting Duan, Junbin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | V2VFormer++: Multi-Modal Vehicle-to-Vehicle Cooperative Perception via Global-Local TransformerabstractMulti-vehicle cooperative perception has recently emerged for facilitating long-range and large-scale perception ability of connected automated vehicles (CAVs). Nonetheless, enormous efforts formulate collaborative perception as LiDAR-only 3D detection paradigm, neglecting the significance and complementary of dense image. In this work, we construct the first multi-modal vehicle-to-vehicle cooperative perception framework dubbed as V2VFormer++, where individual camera-LiDAR representation is incorporated with dynamic channel fusion (DCF) at bird’s-eye-view (BEV) space and ego-centric BEV maps from adjacent vehicles are aggregated by global-local transformer module. Specifically, channel-token mixer (CTM) with MLP design is developed to capture global response among neighboring CAVs, and position-aware fusion (PAF) further investigate the spatial correlation between each ego-networked map in a local perspective. In this manner, we could strategically determine which CAVs are desirable for collaboration and how to aggregate the foremost information from them. Quantitative and qualitative experiments are conducted on both publicly-available OPV2V and V2X-Sim 2.0 benchmarks, and our proposed V2VFormer++ reports the state-of-the-art cooperative perception performance, demonstrating its effectiveness and advancement. Moreover, ablation study and visualization analysis further suggest the strong robustness against diverse disturbances from real-world scenarios. Daxin Tian, Chunmian Lin, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | DA-RDD: Toward Domain Adaptive Road Damage Detection Across Different CountriesabstractRecent advances on road damage detection relies on a large amount of labeled data, whilst collecting pavement image is labor-intensive and time-consuming. Unsupervised Domain Adaptation (UDA) provides a promising solution to adapt a source domain to the target domain, however, cross-domain crack detection is still an open problem. In this paper, we propose domain adaptive road damage detection termed as DA-RDD, by incorporating image-level with instance-level feature alignment for domain-invariant representation learning in an adversarial manner. Specifically, importance weighting is introduced to evaluate the intermediate samples for image-level alignment between domains, and we aggregate RoI-wise feature with multi-scale contextual information to recover the crack details for progressive domain alignment at instance level. Additionally, a large-scale road damage dataset (based on Road Damage Dataset 2020 (RDD2020)) named as RDD2021 is constructed with$100k$synthetic labeled distress images. Extensive experimental results on damage detection across different countries demonstrate the universality and superiority of DA-RDD, and empirical studies on RDD2021 further claim its effectiveness and advancement. To our best knowledge, it is the first time to investigate domain adaptative pavement crack detection, and we expect the contributions in this work would facilitate the development of generalized road damage detection in the future. Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | 3D-DFM: Anchor-Free Multimodal 3-D Object Detection With Dynamic Fusion Module for Autonomous DrivingabstractRecent advances in cross-modal 3D object detection rely heavily on anchor-based methods, and however, intractable anchor parameter tuning and computationally expensive postprocessing severely impede an embedded system application, such as autonomous driving. In this work, we develop an anchor-free architecture for efficient camera-light detection and ranging (LiDAR) 3D object detection. To highlight the effect of foreground information from different modalities, we propose a dynamic fusion module (DFM) to adaptively interact images with point features via learnable filters. In addition, the 3D distance intersection-over-union (3D-DIoU) loss is explicitly formulated as a supervision signal for 3D-oriented box regression and optimization. We integrate these components into an end-to-end multimodal 3D detector termed 3D-DFM. Comprehensive experimental results on the widely used KITTI dataset demonstrate the superiority and universality of 3D-DFM architecture, with competitive detection accuracy and real-time inference speed. To the best of our knowledge, this is the first work that incorporates an anchor-free pipeline with multimodal 3D object detection. Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | CL3D: Camera-LiDAR 3D Object Detection With Point Feature Enhancement and Point-Guided FusionabstractCamera-LiDAR 3D object detection has been extensively investigated due to its significance for many real-world applications. However, there are still of great challenges to address the intrinsic data difference and perform accurate feature fusion among two modalities. To these ends, we propose a two-stream architecture termed as CL3D, that integrates with point enhancement module, point-guided fusion module with IoU-aware head for cross-modal 3D object detection. Specifically, pseudo LiDAR is firstly generated from RGB image, and point enhancement module (PEM) is then designed to enhance the raw LiDAR with pseudo point. Moreover, point-guided fusion module (PFM) is developed to find image-point correspondence at different resolutions, and incorporate semantic with geometric features in a point-wise manner. We also investigate the inconsistency between localization confidence and classification score in 3D detection, and introduce IoU-aware prediction head (IoU Head) for accurate box regression. Comprehensive experiments are conducted on publicly available KITTI dataset, and CL3D reports the outstanding detection performance compared to both single- and multi-modal 3D detectors, demonstrating its effectiveness and competitiveness. Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | SA-YOLOv3: An Efficient and Accurate Object Detector Using Self-Attention Mechanism for Autonomous DrivingabstractObject detection is becoming increasingly significant for autonomous-driving system. However, poor accuracy or low inference performance limits current object detectors in applying to autonomous driving. In this work, a fast and accurate object detector termed as SA-YOLOv3, is proposed by introducing dilated convolution and self-attention module (SAM) into the architecture of YOLOv3. Furthermore, loss function based on GIoU and focal loss is reconstructed to further optimize detection performance. With an input size of$512\times 512$, our proposed SA-YOLOv3 improves YOLOv3 by 2.58 mAP and 2.63 mAP on KITTI and BDD100K benchmarks, with real-time inference (more than 40 FPS). When compared with other state-of-the-art detectors, it reports better trade-off in terms of detection accuracy and speed, indicating the suitability for autonomous-driving application. To our best knowledge, it is the first method that incorporates YOLOv3 with attention mechanism, and we expect this work would guide for autonomous-driving research in the future. Daxin Tian, Chunmian Lin, Jianshan Zhou, Xuting Duan, Yue Cao 0002, Dezong Zhao, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 2 |