EDBT 2026 Demo / reviewers in the wild / expert
Junyin Wang
dblp:291/1477
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Driving with Advice: Large Model as Motion Advisor for Joint PlanningabstractWe address the challenge of integrating high-level semantic reasoning with low-level trajectory planning in end-to-end autonomous driving, where most existing frameworks decouple perception, decision-making, and control, leading to limited interpretability and poor instruction compliance. To bridge this gap, we propose Driving with Advice, a novel closed-loop framework that treats a vision-language model (VLM) as a motion advisor to provide interpretable, language-mediated guidance for trajectory generation. Our approach introduces three key innovations: (1) Semantic-Intentional Pretraining (SIP), which injects driving rationale into a compact VLM via machine-generated question-answering pairs; (2) a discrete action space grounded in directional and speed primitives, enabling structured and interpretable policy learning; and (3) an advice-following diffusion policy refined via Group Relative Policy Optimization under a multi-objective reward that ensures safety, comfort, and alignment with semantic intent. We evaluate our method on the NAVSIM benchmark in a closed-loop setting, achieving a state-of-the-art Predictive Driver Model Score (PDMS) of 91.5, outperforming strong baselines in safety (NC: 99.2). The results demonstrate that leveraging language as a cognitive interface between perception and control enhances both generalization and behavioral transparency, advancing the paradigm of language-conditioned driving. Junyin Wang, Jinlei Yu, Huikai Liu, Wenqian Zhu, Shengwu Xiong 0001 |
AAAI | 1 |
| 2026 | SLOcc: Selective interaction and long-range modelling for occupancy prediction
Junyin Wang, Chenghu Du, Tongao Ge, Shengwu Xiong 0001 |
Expert Syst. Appl. | 1 |
| 2026 | OV-Pro: Enhancing open-vocabulary 3D object detection by prototype contrastive distillation
Tongao Ge, Junyin Wang, Chenghu Du, Hui Li 0010, Huikai Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2026 | D3PD: Dual distillation and dynamic fusion for camera-radar 3D perception
Junyin Wang, Chenghu Du, Tongao Ge, Bingyi Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 1 |
| 2026 | RDNet: Rotate-Groundtruth Augmentation and Decoupled Attention HEAD for 3D Object DetectionabstractLiDAR is one of the most important sensors in the field of autonomous driving and allows for better and more accurate perception of the changes in the surrounding environment. Most of the existing 3D object detection methods use data augmentation and feature fusion enhancement to improve the performance of detection, but the majority of the methods ignore the handling of sample imbalance problems during data augmentation. Also, the designed feature fusion and enhancement methods were not well suited to work with the enhancement methods. To this end, we developed a combined method involving data augmentation and feature enhancement. The designed approach has two main objectives: 1) to address the problem of unbalanced sample distribution in detection scenes through data augmentation, and 2) to enhance feature perception using a special feature enhancement module. Our proposed method solves the problem of class imbalance by directly increasing the number of pedestrian samples in the scene through mixed data augmentation, i.e., RG-Aug. In addition, we introduce the Decoupling and Attention Fusion module (DAF), which combines classification headers with high-level features and prediction branches with low-level features. Leverage data features between different layers of features to get a more robust feature representation. Finally, the multi-scale pyramid attention enhancement module is designed to achieve feature enhancement of multi-scale features by means of attention to improve the detection ability of small objects in the scene, especially the detection ability of pedestrians. Our method can achieve 1.57%, 2.16%, and 2.05% performance improvement on the KITTI dataset for Easy, Mod, and Hard samples, respectively. Furthermore, for the detection of pedestrians, our method has a significant competitive advantage over other state-of-the-art techniques with a mAP of 73.42%. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Junyin Wang, Jianqun Cui, Yanan Chang, Chenghu Du, Jia Wu 0001 |
IEEE Trans. Big Data | 4 |
| 2026 | CrossBEV: enhancing multi-view 3D object detection via spatiotemporal feature cross-enhancement
Jinlei Yu, Junyin Wang, Wenqian Zhu, Huikai Liu, Weidong Yang 0006 |
Vis. Comput. | 2 |
| 2025 | Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationabstractEfficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily dependent on the constraints of inpainting masks. An incorrect mask can mislead the generated results, while overly large mask areas can lose essential original information, thereby hindering the synthesis of high-quality results. To address these problems, we propose a latent diffusion model-based virtual try-on network that achieves fully supervised learning using the concept of cycle consistency and knowledge distillation. Specifically, we divide our approach into pretext and downstream tasks. In the pretext task, we generate a pseudo-label (pseudo-person image) to form paired training samples, which enables the downstream task to achieve fully supervised learning. To prevent the unreliable pseudo-person image from introducing irresponsible prior knowledge, we propose a noise-covering strategy, which aims at fully optimizing the pseudo-label to eliminate the impact of the incorrect inpainting mask as much as possible. Additionally, we propose a skin refinement loss to further enhance the generation of details in the skin region. Extended experiments demonstrate that our proposed method is superior to state-of-the-art methods. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 2 |
| 2025 | GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free ApproachabstractA good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 2 |
| 2025 | ROME: Radar Sparsity Improvement and Omnimodal Enhancement for 3D Object Detection in Bird's Eye ViewsabstractCombining omnimodal feature interaction using LiDAR, surround-view camera, and Radar to form a network has a great guarantee for the safety of autonomous driving, but most of the current omnimodal fusion methods focus on the interaction enhancement of LiDAR and surround-view camera, ignoring the focus on Radar. Enhancing the contextual representation of Radar can ensure better all-weather capability of the perceptual network. To this end, we design the ROME method based on Radar sparsity improvement to better enhance the performance and robustness of the model in terms of alleviating Radar sparsity shortcomings. Firstly, we design the Autocorrelation Point Enhancement (APE) module to improve Radar sparsity leveraging the point-to-point autocorrelation of Radar. Moreover, for omnimodal Bird’s Eye View (BEV) features, an Omnimodal Adaptive Fusion (OAF) module is designed to improve the robustness of BEV features. With the improved Radar modality, the performance of BEV features for the whole driving scene is further improved. Comprehensive experiments on the nuScenes dataset and comparisons with state-of-the-art methods demonstrate the advantages of our proposed method. Yilong Guo, Junyin Wang, Chenghu Du, Shengwu Xiong 0001, Yaxiong Chen |
ICASSP | 2 |
| 2025 | MDC: Modality Distribution Consistent Distillation for Multi-View 3D Object DetectionabstractThe purely visual, multi-view perception approach provides a cost-effective solution for autonomous driving perception. However, vision-based systems struggle to achieve the same precision in object localization as LiDAR due to fundamental differences in their sensing mechanisms. To address this, we introduce MDC, a novel method that integrates LiDAR’s superior spatial information into camera-based systems. Our approach includes three key distillation modules: Distribution Consistency Distillation (DCD), Mask Adaptive Distillation (MAD), and Result Distillation (RD). DCD aligns point cloud and multi-view voxel distributions to boost 3D spatial perception. MAD uses adaptive masking to refine BEV feature alignment with LiDAR. RD ensures consistency in the decoding phase. Experiments conducted on the nuScenes benchmark demonstrate that our method achieves a performance improvement of 2.7% to 3.4% over the student network, highlighting its potential to enhance autonomous driving perception capabilities. Huikai Liu, Junyin Wang, Wenqian Zhu, Shengwu Xiong 0001 |
ICME | 2 |
| 2025 | Mask Does Not Matter: A Unified Latent Diffusion-Enhanced Framework for Mask-Free Virtual Try-OnabstractA good virtual try-on model should introduce minimal redundant conditional information to avoid instability and increase inference efficiency. Existing methods rely on inpainting masks to guide the generation of the object, but the masks, generated by unstable human parsers, often produce unreliable results with fabric residues due to wrong segmentation. Moreover, large mask regions can lose spatial structure and identity information, requiring extra conditional inputs to compensate, which increases model instability and reduces efficiency. To tackle the problem, we present a novel Mask-Free virtual Try-ON (MFTON) framework. Specifically, we propose a mask-free strategy to eliminate all denoising conditions except for clothing and person images, thereby directly extracting spatial structure and identity information from the person image to improve efficiency and reduce instability. Additionally, to optimize the generated clothing regions, we propose a clothing texture-aware attention mechanism to enable the model to focus on texture generation with significant visual differences. We then introduce a geometric detail capture loss to further enable the model to capture more high-frequency information. Finally, we propose an appearance consistency inference method to reduce the initial randomness of the sampling process significantly. Extensive experiments on popular datasets demonstrate that our method outperforms state-of-the-art virtual try-on methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IJCAI | 2 |
| 2025 | Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free FrameworkabstractThis paper investigates the occlusion problems in virtual try-on (VTON) tasks. According to how they affect the try-on results, the occlusion issues of existing VTON methods can be grouped into two categories: (1) Inherent Occlusions, which are the ghosts of the clothing from reference input images that exist in the try-on results. (2) Acquired Occlusions, where the spatial structures of the generated human body parts are disrupted and appear unreasonable. To this end, we analyze the causes of these two types of occlusions, and propose a novel mask-free VTON framework based on our analysis to deal with these occlusions effectively. In this framework, we develop two simple-yet-powerful operations: (1) The background pre-replacement operation prevents the model from confusing the target clothing information with the human body or image background, thereby mitigating inherent occlusions. (2) The covering-and-eliminating operation enhances the model's ability of understanding and modeling human semantic structures, leading to more realistic human body generation and thus reducing acquired occlusions. Moreover, our method is highly generalizable, which can be applied in in-the-wild scenarios, and our proposed operations can also be easily integrated into different generative network architectures (e.g., GANs and diffusion models) in a plug-and-play manner. Extensive experiments on three VTON datasets validate the effectiveness and generalization ability of our method. Both qualitative and quantitative results demonstrate that our method outperforms recently proposed VTON benchmarks. Chenghu Du, Shengwu Xiong 0001, Junyin Wang, Shili Xiong |
NeurIPS | 3 |
| 2025 | Multimodal feature adaptive fusion for anchor-free 3D object detection
Yanli Wu, Junyin Wang, Hui Li 0010, Xiaoxue Ai |
Appl. Intell. | 2 |
| 2025 | GLV: Geometric Correlation Distillation for Latent Diffusion-Enhanced Parser-Free Virtual Try-OnabstractApplying knowledge distillation to virtual try-on tasks is challenging because current methods fail to fully and efficiently exploit responsible teacher knowledge. In other words, existing approaches merely transfer prior knowledge to the student model via pseudo-labels generated by the teacher model, resulting in shallow knowledge representation and low training efficiency. To address these limitations, we propose a novel teacher-student architecture for parser-free virtual try-on, named GLV, which generates high-quality try-on results with realistic body details. Specifically, we propose a deformation-related prior distillation method to effectively leverage the valuable deformation information contained in the teacher warpage model. This enhances the convergence efficiency of the student warpage model, preventing it from getting stuck in a local minima. Moreover, we are the first to propose a geometric correlation distillation, which models the underlying geometric relationship between clothing and the person and transfers this relationship from the teacher to the student. This enables the student warpage model to reduce the entanglement of deformation-irrelevant features, such as color and texture. Finally, we propose a clothing-body retouching method for try-on result synthesis, which refines the denoising process in the latent space of a well-trained diffusion model, thereby preventing catastrophic forgetting. This method seamlessly transforms the parser-based inpainting synthesis paradigm into a parser-free synthesis paradigm and enables efficient convergence of the diffusion model with only fine-tuning. Extensive experiments demonstrate the generality of our approach and highlight its superiority over previous methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | HybridBEV: Hybrid Encode and Distillation for Improved BEV 3D Object DetectionabstractThe development of surround-view cameras is crucial for the advancement of autonomous driving. Utilizing depth information and image features to simulate LiDAR bird’s-eye-view (BEV) features can accomplish efficient 3D object detection tasks. Existing dense BEV generation methods heavily rely on the use of depth features, however, the suboptimal exploitation of these features often results in ambiguity in object location and feature representation during the BEV generation process. To address this, we have designed a hybrid encode and distillation method to enhance 3D object detection performance, termed HybridBEV. Initially, we designed the HybridEncode module, which employs a resampling strategy of depth features in voxel space to obtain BEV features that more accurately reflect the distribution of objects. Subsequently, we introduced multiple distillation methods to supervise the network’s voxel features and BEV feature representations, assisting the student network in learning critical features from the teacher model and ensuring that BEV features can more distinctly represent object distribution. Furthermore, during network training, we loaded pre-trained weights from the teacher network to guide network optimization and accelerate training. Extensive experiments on the nuScenes benchmark demonstrate that HybridBEV can effectively improve the performance of the student network and outperform previous state-of-the-art methods based on surround-view cameras. The code will be published athttps://github.com/wjyxx/HybridBEV Junyin Wang, Chenghu Du, Huikai Liu, Zhenchang Xia, Bingyi Liu, Shengwu Xiong 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | CycleVTON: A Cycle Mapping Framework for Parser-Free Virtual Try-OnabstractImage-based virtual try-on aims to transfer a target clothing onto a specific person. A significant challenge is arbitrarily matched clothing and person lack corresponding ground truth to supervised learning. A recent pioneering work leveraged an improved cycleGAN to enable one network to generate the desired image for another network during training. However, there is no difference in the result distribution before and after the clothing changes. Therefore, using two different networks is unnecessary and may even increase the difficulty of convergence. Furthermore, the introduced human parsing used to provide body structure information in the input also have a negative impact on the try-on result. How to employ a single network for supervised learning while eliminating human parsing? To tackle these issues, we present a Cycle mapping Virtual Try-On Network (CycleVTON), which can produce photo-realistic try-on results by using a cycle mapping framework without the parser. In particular, we introduce a flow constraint loss to achieve supervised learning of arbitrarily matched clothing and person as inputs to the deformer, thus naturally mimicking the interaction between clothing and the human body. Additionally, we design a skin generation strategy that can adapt to the shape of the target clothing by dynamically adjusting the skin region, i.e., by first removing and then filling skin areas. Extensive experiments conducted on challenging benchmarks demonstrate that our proposed method exhibits superior performance compared to state-of-the-art methods. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
AAAI | 2 |
| 2024 | IFNET: Integrating Data Augmentation and Decoupled Attention Fusion for 3D Object DetectionabstractLiDAR is a key sensor for accurately sensing of the environment in autonomous driving. While existing 3D object detection methods generally rely on data augmentation and feature fusion to improve performance, the challenge of dealing with sample imbalance is often overlooked. We design a novel 3D detection network, IFNet, that tackles these issues by introducing mutually reinforcing data augmentation and feature enhancement strategies. It aims to achieve a dual purpose: 1) correcting the category imbalance by directly enhancing pedestrian samples using mixed data augmentation, i.e., RG-Aug; and 2) enhancing feature perception by introducing the decoupling and attention fusion module (DAF). DAF enables robust feature representations across different layers, improving the detection performance, especially for small objects in the scene. Comprehensive experiments on the KITTI dataset and comparisons with state-of-the-art methods demonstrate the superiority of our proposed approach. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Jia Wu 0001, Junyin Wang, Chenghu Du |
ICASSP | 5 |
| 2023 | DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionabstractSurround-view cameras combined with image depth transformation to 3D feature space and fusion with point cloud features are highly regarded. The transformation of 2D features into 3D feature space by means of predefined sampling points and depth distribution happens throughout the scene, and this process generates a large number of redundant features. In addition, multimodal feature fusion unified in 3D space often happens in the previous step of the downstream task, ignoring the interactive fusion between different scales. To this end, we design a new framework, focusing on the design that can give 3D geometric perception information to images and unify them into voxel space to accomplish multi-scale interactive fusion, and we mitigate feature alignment between modal features by geometric relationships between voxel features. The method has two main designs. First, a Segmentation-guided Image View Transformation module is used to accurately transform the pixel region containing the object into a 3D pseudo-point voxel space with the help of a depth distribution. This allows subsequent feature fusion to be performed in a unified voxel feature. Secondly, a Voxel-centric Consistent Fusion module is used to alleviate the errors caused by depth estimation, as well as to achieve better feature fusion between unified modalities. Through extensive experiments on the KITTI and nuScenes datasets, we validate the effectiveness of our camera-LIDAR fusion method. Our proposed approach shows competitive performance on both datasets and outperforms state-of-the-art methods in certain classes of 3D object detection benchmarks. https://github.com/no-Name128/DLFusion [code release] Junyin Wang, Chenghu Du, Hui Li 0010, Shengwu Xiong 0001 |
ACM Multimedia | 1 |
| 2023 | Greatness in Simplicity: Unified Self-Cycle Consistency for Parser-Free Virtual Try-OnabstractImage-based virtual try-on tasks remain challenging, primarily due to inherent complexities associated with non-rigid garment deformation modeling and strong feature entanglement of clothing within human body. Recent groundbreaking formulations, such as in-painting, cycle consistency, and knowledge distillation, have facilitated self-supervised generation of try-on images. However, these paradigms necessitate the disentanglement of garment features within human body features through auxiliary tasks, such as leveraging 'teacher knowledge' and dual generators. The potential presence of irresponsible prior knowledge in the auxiliary task can serve as a significant bottleneck for the main generator (e.g., 'student model') in the downstream task. Moreover, existing garment deformation methods lack the ability to perceive the correlation between the garment and the human body in the real world, leading to unrealistic alignment effects. To tackle these limitations, we present a new parser-free virtual try-on network based on unified self-cycle consistency (USC-PFN), which enables robust translation between different garments using just a single generator, faithfully replicating non-rigid geometric deformation of garments in real-life scenarios. Specifically, we first propose a self-cycle consistency architecture with a circular mode. It utilizes real unpaired garment-person images exclusively as input for training, effectively eliminating the impact of irresponsible prior knowledge at the model input end. Additionally, we formulate a Markov Random Field to simulate a more natural and realistic garment deformation. Furthermore, USC-PFN can leverage a general generator for self-supervised cycle training. Experiments demonstrate that our method achieves state-of-the-art performance on a popular virtual try-on benchmark. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
NeurIPS | 2 |
| 2021 | Efficient and accurate object detection for 3D point clouds in intelligent visual internet of things
Hui Li 0010, Junyin Wang, Lingwei Xu, Ye Tao 0002 |
Multim. Tools Appl. | 2 |
| 2021 | Object detection method based on global feature augmentation and adaptive regression in IoT
Hui Li 0010, Lingwei Xu, Junyin Wang |
Neural Comput. Appl. | 5 |