EDBT 2026 Demo / reviewers in the wild / expert
Junwei Li 0009
dblp:06/4732-9
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-6957-3059ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A dual-stream foreground-aware enhancement network with spiralscan-Mamba for vision-based occupancy prediction in autonomous driving
Nannan Liu, Yanyin Guo, Chuiyi Deng, Zhuoyi Zhao, Junwei Li 0009 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | WA-FDNet: A Unified Weight Adaptation Network for Multimodal Image Fusion and Object Detection
Yanyin Guo, Junwei Li 0009, Zhiyuan Zhang 0004 |
CGI (2) | 3 |
| 2025 | Bridging the Modality Gap: Advancing Multimodal Human Pose Estimation with Modality-Adaptive Pose Estimator and Novel Benchmark Datasets
Jiangnan Xia, Zhiyuan Zhang 0004, Yanyin Guo, Jianghan Cheng, Junwei Li 0009 |
CVM (3) | 7 |
| 2025 | Improving Multimodal Human Pose Estimation by Adversarial Modality Enhancement†abstractHuman pose estimation in computer vision predominantly focuses on the visible modality, with limited research on the infrared modality. No existing methods demonstrate robust performance across both modalities, missing their complementary strengths. This gap arises from the lack of a multimodal benchmark and the difficulty of developing robust multimodal capabilities. To address this, we introduce MMPD, a novel visible-infrared multimodal pose benchmark with high-quality annotations for both modalities. Leveraging MMPD, we expose the limitations of state-of-the-art methods due to modality variance. To overcome this challenge, we propose a novel method-agnostic scheme called AMMPE. By employing the Modality Adversarial Enhancement Stage and Modality Interaction Stage, AMMPE easily incorporates multimodal information without additional pose annotations and enhances effective modality interaction. Extensive experiments demonstrate that AMMPE improves performance in both visible and infrared modalities, achieving excellent modality robustness. The code and benchmark is avaible at: https://github.com/ICANDOALLTHINGSSS/Adversarial-Multi-Modality-Pose-Estimation Jiangnan Xia, Yanyin Guo, Jianghan Cheng, Junwei Li 0009, Zhiyuan Zhang 0004 |
ICASSP | 6 |
| 2025 | HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose EstimationabstractWe propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model and refine the objective function design to reduce computational overhead without compromising performance. Evaluation results on the Human3.6M and MPI-INF-3DHP datasets demonstrate that HDiffTG achieves state-of-the-art (SOTA) performance on the MPI-INF-3DHP dataset while excelling in both accuracy and computational efficiency. Additionally, the model exhibits exceptional robustness in noisy and occluded environments. Source codes and models are available at https://github.com/CirceJie/HDiffTG Yajie Fu, Chaorui Huang, Junwei Li 0009, Hui Kong 0001, Yibin Tian, Huakang Li, Zhiyuan Zhang 0004 |
IJCNN | 3 |
| 2025 | TS-Diff: Two-Stage Diffusion Model for Low-Light RAW Image EnhancementabstractThis paper presents a novel Two-Stage Diffusion Model (TS-Diff) for enhancing extremely low-light RAW images. In the pre-training stage, TS-Diff synthesizes noisy images by constructing multiple virtual cameras based on a noise space. Camera Feature Integration (CFI) modules are then designed to enable the model to learn generalizable features across diverse virtual cameras. During the aligning stage, CFIs are averaged to create a target-specific CFIT, which is fine-tuned using a small amount of real RAW data to adapt to the noise characteristics of specific cameras. A structural reparameterization technique further simplifies CFITfor efficient deployment. To address color shifts during the diffusion process, a color corrector is introduced to ensure color consistency by dynamically adjusting global color distributions. Additionally, a novel dataset, QID, is constructed, featuring quantifiable illumination levels and a wide dynamic range, providing a comprehensive benchmark for training and evaluation under extreme low-light conditions. Experimental results demonstrate that TS-Diff achieves state-of-the-art performance on multiple datasets, including QID, SID, and ELD, excelling in denoising, generalization, and color consistency across various cameras and illumination levels. These findings highlight the robustness and versatility of TS-Diff, making it a practical solution for low-light imaging applications. Source codes and models are available at https://github.com/CircccleK/TS-Diff Zhiyuan Zhang 0004, Jiangnan Xia, Jianghan Cheng, Junwei Li 0009, Yibin Tian, Hui Kong 0001 |
IJCNN | 6 |
| 2025 | LSFDNet: A Single-Stage Fusion and Detection Network for Ships Using SWIR and LWIR
Yanyin Guo, Runxuan An, Junwei Li 0009, Zhiyuan Zhang 0004 |
ACM Multimedia | 3 |
| 2025 | GSFANet: Global Spatial-Frequency Attention Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) has progressed significantly in spatial-domain learning. However, single-frame images’ limited spatial semantics impair discrimination between targets and similar noise while complicating integrity detection of large-scale target. To address these, we propose Global Spatial-Frequency Attention Network (GSFANet), which enhances the distribution difference between targets and noise from a frequency-domain perspective while preserving spatial information integrity. The core innovations consist of three modules: 1) Parametric Wavelet Downsampling (PWD), preserving small target details during frequency refinement to prevent feature fragmentation; 2) Hierarchical Gated Kernel Attention (HGKA), capturing cross-level frequency relationships through Cross-channel Kernel Attention (C2K) and maintaining spatial coherence via Cross-spatial Gate Attention (CSG), effectively bridging semantic gaps across layers; 3) Adaptive Frequency-Decoupled Fusion (AdaFD), dynamically fusing target-associated frequency components while suppressing noise. We further develop AdaFL Loss to balance multi-scale target gradients and stabilize training. Experiments on three benchmark datasets demonstrate GSFANet’s superior detection performance and enhanced segmentation robustness in complex scenarios compared to state-of-the-art methods. Our code will be made public at https://github.com/dengfa02/GSFANet_IRSTD. Chuiyi Deng, Zhuoyi Zhao, Xiang Xu 0002, Yixin Xia, Junwei Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography EstimationabstractWe propose Multiscale Correlation searching homogra-phy estimation Network, namely MCNet, an iterative deep homography estimation architecture. Different from previous approaches that achieve iterative refinement by correlation searching within a single scale, MCNet combines the multiscale strategy with correlation searching incur-ring nearly ignored computational overhead. Moreover, MCNet adopts a Fine-Grained Optimization loss function, named FGO loss, to further boost the network training at the convergent stage, which can improve the estimation accuracy without additional computational overhead. Ac-cording to our experiments, using the above two simple strategies can produce significant homography estimation accuracy with considerable efficiency. We show that MC-Net achieves state-of-the-art performance on a variety of datasets, including common scene MSCOCO, cross-modal scene GoogleEarth and GoogleMap, and dynamic scene SPID. Compared to the previous SOTA method, 2-scale RHWF, our MCNet reduces inference time, FLOPs, parameter cost, and memory cost by 78.9%, 73.5%, 34.1%, and 33.2% respectively, while achieving 20.5% (MSCOCO), 43.4% (GoogleEarth), and 41.1% (GoogleMap) mean average corner error (MACE) reduction. Source code is available at https://github.com/zjuzhk/MCNet. Haokai Zhu, Si-Yuan Cao, Jianxin Hu, Sitong Zuo, Beinan Yu, Jiacheng Ying, Junwei Li 0009 |
CVPR | 7 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 7 |
| 2024 | An adaptive network fusing light detection and ranging height-sliced bird's-eye view and vision for place recognition
Zuo Jiang, Yibin Ye, Junwei Li 0009, Zhiyuan Zhang 0004 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerabstractWe propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF. Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009 |
CVPR | 6 |
| 2023 | BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesabstractPlace recognition is a key module for long-term SLAM systems. Current LiDAR-based place recognition methods usually use representations of point clouds such as unordered points or range images. These methods achieve high recall rates of retrieval, but their performance may degrade in the case of view variation or scene changes. In this work, we explore the potential of a different representation in place recognition, i.e. bird’s eye view (BEV) images. We validate that, in scenes of slight viewpoint changes, a simple NetVLAD network trained on BEV images achieves comparable performance to the state-of-the-art place recognition methods. For robustness to view variations, we propose a rotation-invariant network called BEVPlace. We use group convolution to extract rotation-equivariant local features from the images and NetVLAD for global feature aggregation. In addition, we observe that the distance between BEV features is correlated with the geometry distance of point clouds. Based on the observation, we develop a method to estimate the position of the query cloud, extending the usage of place recognition. The experiments conducted on large-scale public datasets show that our method 1) achieves state-of-the-art performance in terms of recall rates, 2) is robust to view changes, 3) shows strong generalization ability, and 4) can estimate the positions of query point clouds. Source codes are publicly available at https://github.com/zjuluolun/BEVPlace. Lun Luo, Shuhang Zheng, Yongzhi Fan, Beinan Yu, Si-Yuan Cao, Junwei Li 0009 |
ICCV | 7 |