Wei Li 0111

dblp:64/6025-111 · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
31since 2021 · last 2026
0000-0002-0059-3745ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 21 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Driving
abstract
3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model architecture and computational complexity, especially in high-resolution scenarios. Existing approaches for handling high-resolution scenes typically obtain fine-grained features by grid sampling on low-resolution feature map, resulting in limited sparsity and insufficient feature interaction. This paper presents a framework leveraging SParse representation and SCalable feature interaction to address the aforementioned challenges, called SPSC. Specifically, we maintain sparsity by progressively pruning unoccupied queries during the coarse-to-fine process, thereby reducing the scale of data that the model needs to handle. Subsequently, we introduce query serialization, which transforms queries into an ordered sequence while preserving their spatial structure, This enables fine-grained feature interaction while maintaining linear computational complexity and a larger receptive field. Without complex architectural designs, SPSC significantly outperforms SOTA approaches, relatively enhances the mIoU by 12.0%, 11.0% and 4.8% on nuScenes-Occupancy dataset under the muli-modal, LiDAR and camera settings, respectively.
Qingju Guo, Shuang Li 0008, Binhui Xie, Jing Geng 0002, Wei Li 0111
AAAI5
2026 Real Garment Benchmark (RGBench): A Comprehensive Benchmark for Robotic Garment Manipulation Featuring a High-Fidelity Scalable Simulator
abstract
While there has been significant progress to use simulated data to learn robotic manipulation of rigid objects, applying its success to deformable objects has been hindered by the lack of both deformable object models and realistic non-rigid body simulators. In this paper, we present Real Garment Benchmark (RGBench), a comprehensive benchmark for robotic manipulation of garments. It features a diverse set of over 6000 garment mesh models, a new high-performance simulator, and a comprehensive protocol to evaluate garment simulation quality with carefully measured real garment dynamics. Our experiments demonstrate that our simulator outperforms currently available cloth simulators by a large margin, reducing simulation error by 20% while maintaining a speed of 3 times faster. We will publicly release RGBench to accelerate future research in robotic garment manipulation.
Wenkang Hu, Xincheng Tang, Yanzhi E, Zhengjie Shu, Wei Li 0111, Huamin Wang 0001, Ruigang Yang
AAAI6
2026 FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
abstract
Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods.
Jinxiao Li, Jingnan Wang, Zhiwen Zuo, Jianfeng Dong, Wei Li 0111, Chi Wang 0004, Weiwei Xu 0003, Xun Wang 0007
AAAI6
2026 Cluster-based Pseudo-labeling for Semi-Supervised LiDAR Semantic Segmentation
abstract
The costly annotation process has driven the development of semi-supervised learning (SSL) approaches. Existing semi-supervised LiDAR segmentation methods typically process entire point clouds directly, aiming to assign labels to all points at the scene scale. However, the large number of points, combined with their sparse and irregular nature, makes it challenging to learn scene-level optimization objectives, especially in SSL settings where labeled data are insufficient. This paper presents a Cluster-based pseudo-LAbeling Semi-Supervised technique, called CLASS. CLASS is designed to divide point clouds into several small, pure clusters, thereby decomposing challenging scene-scale segmentation task into more manageable cluster-scale classification and segmentation tasks, enabling the generation of high-quality pseudo labels for unlabeled data. CLASS possesses three key properties. i) Task simplicity: our pseudo-labeling process is based on simpler cluster-scale classification and segmentation tasks, resulting in ease of learning. ii) Labeling effectiveness: CLASS can generate pseudo-labels comparable to ground truth using only approximately 10% labeled data. iii) Universal versatility: CLASS exhibits flexibility regarding LiDAR representations (e.g., BEV, voxel, and range view). Comprehensive experiments on popular LiDAR segmentation benchmarks demonstrate its superiority.
Qingju Guo, Shuang Li 0008, Jing Geng 0002, Binhui Xie, Jiawei Shan, Wei Li 0111
WACV6
2025 Fuel-Optimal Operational Speed Planning for Autonomous Trucking on Highways
abstract
The rapid advancement of autonomous driving technology, particularly in autonomous trucking on highways, shows great value for enhancing efficiency and reducing costs in the logistics industry. In this work, we define the full-trip speed planning problem for autonomous trucks under delivery time and fuel consumption constraints, referred to as the Operational Speed Planning (OSP) problem. To support and accelerate research on the OSP problem, we have developed a comprehensive dataset using a fleet of over 400 trucks. The dataset contains rich, diverse information covering more than 22 million kilometers of real-world highway driving data. In addition to this static dataset, we have developed a closed-loop simulator that allows for the interactive evaluation of OSP solutions, enabling researchers to test speed planning strategies in a realistic environment. Furthermore, we provide an OSP baseline method based on dynamic programming to optimize speed planning, balancing the delivery time requirements and fuel consumption. Our extensive experiments demonstrate both the accuracy of the simulation and the effectiveness of the OSP baseline in planning optimal speeds, proving its capability to meet time constraints while improving fuel efficiency. The dataset, simulator, and baseline will be made publicly available to foster further research and innovation in this area.
Wei Li 0111, Jiahao Xiang, Jiaping Ren, Ruigang Yang
ICRA1
2025 IEMFormer: Internal and External Multi-Fusion Transformer for Indoor RGB-D Semantic Segmentation
abstract
Effectively fusing and complementing RGB and depth modalities while mitigating image noise is a critical challenge in the RGB-D semantic segmentation task. In this paper, we propose a novel Internal and External Multi-fusion Transformer (IEMFormer) to address this issue. IEMFormer incorporates stage-specific fusion strategies to enhance modal complementarity. For internal fusion, we integrate a fusion unit within the traditional Transformer block, combining matching tokens from both modalities on a pixel-by-pixel basis. For external fusion, the proposed External Adaptive Cross-modal Fusion (EACF) module filters dual-modal features across both spatial and channel dimensions, serving the purpose of adaptively weighting complementary channel information and robustly aggregating spatial patterns from both modalities, thereby facilitating the integration of multimodal information. Additionally, the Global Self-attention Guided Fusion (GSGF) module in the decoder refines the fused features from earlier stages, effectively suppressing noise. This is achieved by leveraging high-level semantic features to guide the refinement and incorporating an active noise suppression mechanism to prevent overfitting to dominant, noisy features. Extensive experiments on the NYUv2 and SUN RGB-D datasets demonstrate that IEMFormer achieves highly competitive performance in accurately understanding indoor scenes.
Kaidi Hu, Wei Li 0111, Guangwei Gao, Ruigang Yang
IEEE Signal Process. Lett.2
2025 ABFE-Net: Attention-Based Feature Enhancement Network for Few-Shot Point Cloud Classification
abstract
Few-shot 3D point cloud classification has attracted significant attention due to the challenge of acquiring large-scale labeled data. Existing methods often employ network backbones tailored for fully-supervised learning, which can lead to suboptimal performance in few-shot settings. To tackle these limitations, we propose ABFE-Net, a novel method for point cloud classification with few-shot learning principles. We comprehensively summarize the drawbacks of existing network architectures into four aspects: contextual information loss, channel redundancy, overfitting, and insufficient hidden feature extraction. Accordingly, we design novel modules, such as the Attention-based Dilated Mix-up Module (ADMM) and Attention-based Comprehensive Feature Learning (ACFL), to enhance the network by addressing those issues effectively. Experiments on multiple public datasets demonstrate that ABFE-Net achieves state-of-the-art performance with superior generalization.
Kaidi Hu, Mao Ye 0005, Wei Li 0111, Ruigang Yang
IEEE Signal Process. Lett.4
2024 DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection
abstract
Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X.
Xiang Li 0001, Junbo Yin, Wei Li 0111, Cheng-Zhong Xu 0001, Ruigang Yang, Jianbing Shen
AAAI3
2024 IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
abstract
Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However, objects in the BEV representation typically exhibit small sizes, and the associated point cloud context is inherently sparse, which leads to great challenges for reliable 3D perception. In this paper, we propose IS-Fusion, an innovative multimodal fusion framework that jointly captures the Instance- and Scene-level contextual information. IS-Fusion essentially differs from existing approaches that only focus on the BEV scene-level fusion by explicitly incorporating instance-level multimodal information, thus facilitating the instance-centric tasks like 3D object detection. It comprises a Hierarchical Scene Fusion (HSF) module and an Instance-Guided Fusion (IGF) module. HSF applies Point-to-Grid and Grid-to-Region transformers to capture the multimodal scene context at different granularities. IGF mines instance candidates, explores their relationships, and aggregates the local multimodal context for each instance. These instances then serve as guidance to enhance the scene feature and yield an instance-aware BEV representation. On the challenging nuScenes benchmark, IS-Fusion outperforms all the published multimodal works to date. Code is available at: https://github.com/yinjunbo/IS-Fusion.
Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li 0111, Ruigang Yang, Pascal Frossard, Wenguan Wang
CVPR4
2024 CMD: A Cross Mechanism Domain Adaptation Dataset for 3D Object Detection
Jinhao Deng, Xun Huang 0003, Qiming Xia, Xin Li 0003, Wei Li 0111, Chenglu Wen, Cheng Wang 0003
ECCV (57)8
2024 NPC: Neural Predictive Control for Fuel-Efficient Autonomous Trucks
abstract
Fuel efficiency is a crucial aspect of long-distance cargo transportation by oil-powered trucks that economize on costs and decrease carbon emissions. Current predictive control methods depend on an accurate model of vehicle dynamics and engine, including weight, drag coefficient, and the Brake-specific Fuel Consumption (BSFC) map of the engine. We propose a pure data-driven method, Neural Predictive Control (NPC), which does not use any physical model for the vehicle. After training with over 20,000 km of historical data, the novel proposed NVFormer implicitly models the relationship between vehicle dynamics, road slope, fuel consumption, and control commands using the attention mechanism. Based on the online sampled primitives from the past of the current freight trip and anchor-based future data synthesis, the NVFormer can infer optimal control command for reasonable fuel consumption. The physical model-free NPC outperforms the base PCC method with 2.41% and 3.45% more significant fuel saving in simulation and open-road highway testing, respectively.
Jiaping Ren, Jiahao Xiang, Hongfei Gao, Yiming Ren 0001, Yuexin Ma, Ruigang Yang, Wei Li 0111
ICRA9
2024 ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios
abstract
Emergent-scene safety is the key milestone for fully autonomous driving, and reliable on-time prediction is essential to maintain safety in emergency scenarios. However, these emergency scenarios are long-tailed and hard to collect, which restricts the system from getting reliable predictions. In this paper, we build a new dataset, which aims at the longterm prediction with the inconspicuous state variation in history for the emergency event, named the Extro-Spective Prediction (ESP) problem. Based on the proposed dataset, a flexible feature encoder for ESP is introduced to various prediction methods as a seamless plug-in, and its consistent performance improvement underscores its efficacy. Furthermore, a new metric named clamped temporal error (CTE) is proposed to give a more comprehensive evaluation of prediction performance, especially in time-sensitive emergency events of subseconds. Interestingly, as our ESP features can be described in human-readable language naturally, the application of integrating into ChatGPT also shows huge potential. The ESP-dataset and all benchmarks are released at https://dingrui-wang.github.io/ESP-Dataset/.
Dingrui Wang, Zheyuan Lai, Yuda Li, Yuexin Ma, Johannes Betz, Ruigang Yang, Wei Li 0111
ICRA8
2024 Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation
abstract
Capitalizing on the complementary advantages of generative and discriminative models has always been a compelling vision in machine learning, backed by a growing body of research. This work discloses the hidden semantic structure within score-based generative models, unveiling their potential as effective discriminative priors. Inspired by our theoretical findings, we propose DUSA to exploit the structured semantic priors underlying diffusion score to facilitate the test-time adaptation of image classifiers or dense predictors. Notably, DUSA extracts knowledge from a single timestep of denoising diffusion, lifting the curse of Monte Carlo-based likelihood estimation over timesteps. We demonstrate the efficacy of our DUSA in adapting a wide variety of competitive pre-trained discriminative models on diverse test-time scenarios. Additionally, a thorough ablation study is conducted to dissect the pivotal elements in DUSA. Code is publicly available at https://github.com/BIT-DA/DUSA.
Mingjia Li 0003, Shuang Li 0008, Tongrui Su, Longhui Yuan, Jian Liang 0002, Wei Li 0111
NeurIPS6
2023 SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point Cloud
abstract
LiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations, different cities, and weather conditions. The mainstream approaches tend to solve these challenges by leveraging unsupervised domain adaptation (UDA) techniques. However, these UDA solutions just yield unsatisfactory 3D detection results when there is a severe domain shift, e.g., from Waymo (64-beam) to nuScenes (32-beam). To address this, we present a novel Semi-Supervised Domain Adaptation method for 3D object detection (SSDA3D), where only a few labeled target data is available, yet can significantly improve the adaptation performance. In particular, our SSDA3D includes an Inter-domain Adaptation stage and an Intra-domain Generalization stage. In the first stage, an Inter-domain Point-CutMix module is presented to efficiently align the point cloud distribution across domains. The Point-CutMix generates mixed samples of an intermediate domain, thus encouraging to learn domain-invariant knowledge. Then, in the second stage, we further enhance the model for better generalization on the unlabeled target set. This is achieved by exploring Intra-domain Point-MixUp in semi-supervised learning, which essentially regularizes the pseudo label distribution. Experiments from Waymo to nuScenes show that, with only 10% labeled target data, our SSDA3D can surpass the fully-supervised oracle model with 100% target label. Our code is available at https://github.com/yinjunbo/SSDA3D.
Yan Wang 0116, Junbo Yin, Wei Li 0111, Pascal Frossard, Ruigang Yang, Jianbing Shen
AAAI3
2023 Transformation-Equivariant 3D Object Detection for Autonomous Driving
abstract
3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augmentation are required for robust detection. Recent equivariant networks explicitly model the transformation variations by applying shared networks on multiple transformed point clouds, showing great potential in object geometry modeling. However, it is difficult to apply such networks to 3D object detection in autonomous driving due to its large computation cost and slow reasoning speed. In this work, we present TED, an efficient Transformation-Equivariant 3D Detector to overcome the computation cost and speed issues. TED first applies a sparse convolution backbone to extract multi-channel transformation-equivariant voxel features; and then aligns and aggregates these equivariant features into lightweight and compact representations for high-performance 3D object detection. On the highly competitive KITTI 3D car detection leaderboard, TED ranked 1st among all submissions with competitive efficiency. Code is available at https://github.com/hailanyi/TED.
Chenglu Wen, Wei Li 0111, Xin Li 0003, Ruigang Yang, Cheng Wang 0003
AAAI3
2023 GANet: Goal Area Network for Motion Forecasting
abstract
Predicting the future motion of road participants is crucial for autonomous driving but is extremely challenging due to staggering motion uncertainty. Recently, most motion forecasting methods resort to the goal-based strategy, i.e., predicting endpoints of motion trajectories as conditions to regress the entire trajectories, so that the search space of solution can be reduced. However, accurate goal coordinates are hard to predict and evaluate. In addition, the point representation of the destination limits the utilization of a rich road context, leading to inaccurate prediction results in many cases. Goal area, i.e., the possible destination area, rather than goal coordinate, could provide a more soft constraint for searching potential trajectories by involving more tolerance and guidance. In view of this, we propose a new goal area-based framework, named Goal Area Network (GANet), for motion forecasting, which models goal areas as preconditions for trajectory prediction, performing more robustly and accurately. Specifically, we propose a GoICrop (Goal Area of Interest) operator to effectively aggregate semantic lane features in goal areas and model actors' future interactions as feedback, which benefits a lot for future trajectory estimations. GANet ranks the 1st on the leaderboard of Argoverse Challenge among all public literature (till the paper submission). Code will be available at https://github.com/kingwmk/GANet.
Mingkun Wang, Xinge Zhu, Changqian Yu, Wei Li 0111, Yuexin Ma, Ruochun Jin, Xiaoguang Ren, Dongchun Ren, Wenjing Yang 0002
ICRA4
2023 Bridging Language and Geometric Primitives for Zero-shot Point Cloud Segmentation
abstract
We investigate transductive zero-shot point cloud semantic segmentation, where the network is trained on seen objects and able to segment unseen objects. The 3D geometric elements are essential cues to imply a novel 3D object type. However, previous methods neglect the fine-grained relationship between the language and the 3D geometric elements. To this end, we propose a novel framework to learn the geometric primitives shared in seen and unseen categories' objects and employ a fine-grained alignment between language and the learned geometric primitives. Therefore, guided by language, the network recognizes the novel objects represented with geometric primitives. Specifically, we formulate a novel point visual representation, the similarity vector of the point's feature to the learnable prototypes, where the prototypes automatically encode geometric primitives via back-propagation. Besides, we propose a novel Unknown-aware InfoNCE Loss to fine-grained align the visual representation with language. Extensive experiments show that our method significantly outperforms other state-of-the-art methods in the harmonic mean-intersection-over-union (hIoU), with the improvement of 17.8%, 30.4%, 9.2% and 7.9% on S3DIS, ScanNet, SemanticKITTI and nuScenes datasets, respectively. Codes are available1 https://github.com/runnanchen/Zero-Shot-Point-Cloud-Segmentation.
Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li 0111, Yuexin Ma, Ruigang Yang, Wenping Wang 0001
ACM Multimedia4
2023 Fuel Rate Prediction for Heavy-Duty Trucks
abstract
Fuel cost contributes significantly to the high operation cost of heavy-duty trucks. Developing fuel rate prediction models is the cornerstone of fuel consumption optimization approaches for heavy-duty trucks. However, limited by accurate features directly related to the truck’s fuel consumption, state-of-the-art models show poor performance and are rarely deployed in practice. In this paper, we use the truck’s engine management system (EMS) and Instant Fuel Meter (IFM) to collect a three-month dataset during the period of December 2019 to June 2020. Seven prediction models, including linear regression, polynomial regression, MLP, CNN, LSTM, CNN-LSTM, and AutoML, are investigated and evaluated to predict real-time fuel rate. The evaluation results show that the EMS and IFM dataset help to improve the coefficient of determination of traditional linear/polynomial models from 0.87 to 0.96, while learning-based approach AutoML improves the coefficient of determination to attain 0.99. Besides, we explore the actual deployment of fuel rate prediction with transfer learning and path planning for autonomous driving.
Liangkai Liu, Wei Li 0111, Dawei Wang 0006, Ruigang Yang, Weisong Shi
IEEE Trans. Intell. Transp. Syst.2
2023 A Lightweight and Detector-Free 3D Single Object Tracker on Point Clouds
abstract
Recent works on 3D single object tracking treat the task as a target-specific 3D detection task, where an off-the-shelf 3D detector is commonly employed for the tracking. However, it is non-trivial to perform accurate target-specific detection since the point cloud of objects in raw LiDAR scans is usually sparse and incomplete. In this paper, we address this issue by explicitly leveraging temporal motion cues and propose DMT, a Detector-free Motion-prediction-based 3D Tracking network that completely removes the usage of complicated 3D detectors and is lighter, faster, and more accurate than previous trackers. Specifically, the motion prediction module is first introduced to estimate a potential target center of the current frame in a point-cloud-free manner. Then, an explicit voting module is proposed to directly regress the 3D box from the estimated target center. Extensive experiments on KITTI and NuScenes datasets demonstrate that our DMT can still achieve better performance ($\sim $10% improvement over the NuScenes dataset) and a faster tracking speed (i.e., 72 FPS) than state-of-the-art approaches without applying any complicated 3D detectors. Our code is released athttps://github.com/jimmy-dq/DMT.
Yan Xia 0003, Qiangqiang Wu, Wei Li 0111, Antoni B. Chan, Uwe Stilla
IEEE Trans. Intell. Transp. Syst.3
2022 Making The Best of Both Worlds: A Domain-Oriented Transformer for Unsupervised Domain Adaptation
abstract
Extensive studies on Unsupervised Domain Adaptation (UDA) have propelled the deployment of deep learning from limited experimental datasets into real-world unconstrained domains. Most UDA approaches align features within a common embedding space and apply a shared classifier for target prediction. However, since a perfectly aligned feature space may not exist when the domain discrepancy is large, these methods suffer from two limitations. First, the coercive domain alignment deteriorates target domain discriminability due to lacking target label supervision. Second, the source-supervised classifier is inevitably biased to source data, thus it may underperform in target domain. To alleviate these issues, we propose to simultaneously conduct feature alignment in two individual spaces focusing on different domains, and create for each space a domain-oriented classifier tailored specifically for that domain. Specifically, we design a Domain-Oriented Transformer (DOT) that has two individual classification tokens to learn different domain-oriented representations, and two classifiers to preserve domain-wise discriminability. Theoretical guaranteed contrastive-based alignment and the source-guided pseudo-label refinement strategy are utilized to explore both domain-invariant and specific information. Comprehensive experiments validate that our method achieves state-of-the-art on several benchmarks. Code is released at https://github.com/BIT-DA/Domain-Oriented-Transformer.
Wenxuan Ma 0001, Shuang Li 0008, Chi Harold Liu, Yulin Wang 0002, Wei Li 0111
ACM Multimedia6
2022 Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR-Based Perception
abstract
State-of-the-art methods for driving-scene LiDAR-based perception (including point cloud semantic segmentation, panoptic segmentation and 3D detection, etc.) often project the point clouds to 2D space and then process them via 2D convolution. Although this cooperation shows the competitiveness in the point cloud, it inevitably alters and abandons the 3D topology and geometric relations. A natural remedy is to utilize the 3D voxelization and 3D convolution network. However, we found that in the outdoor point cloud, the improvement obtained in this way is quite limited. An important reason is the property of the outdoor point cloud, namely sparsity and varying density. Motivated by this investigation, we propose a new framework for the outdoor LiDAR segmentation, where cylindrical partition and asymmetrical 3D convolution networks are designed to explore the 3D geometric pattern while maintaining these inherent properties. The proposed model acts as a backbone and the learned features from this model can be used for downstream tasks such as point cloud semantic and panoptic segmentation or 3D detection. In this paper, we benchmark our model on these three tasks. For semantic segmentation, we evaluate the proposed model on several large-scale datasets, i.e., SemanticKITTI, nuScenes and A2D2. Our method achieves the state-of-the-art on the leaderboard of SemanticKITTI (both single-scan and multi-scan challenge), and significantly outperforms existing methods on nuScenes and A2D2 dataset. Furthermore, the proposed 3D framework also shows strong performance and good generalization on LiDAR panoptic segmentation and LiDAR 3D detection.
Xinge Zhu, Hui Zhou 0005, Fangzhou Hong, Wei Li 0111, Yuexin Ma, Hongsheng Li 0001, Ruigang Yang, Dahua Lin
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-Resolution
abstract
Despite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches.
Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang
IEEE Trans. Multim.3
2021 Transferable Semantic Augmentation for Domain Adaptation
abstract
Domain adaptation has been widely explored by transferring the knowledge from a label-rich source domain to a related but unlabeled target domain. Most existing domain adaptation algorithms attend to adapting feature representations across two domains with the guidance of a shared source-supervised classifier. However, such classifier limits the generalization ability towards unlabeled target recognition. To remedy this, we propose a Transferable Semantic Augmentation (TSA) approach to enhance the classifier adaptation ability through implicitly generating source features towards target semantics. Specifically, TSA is inspired by the fact that deep feature transformation towards a certain direction can be represented as meaningful semantic altering in the original input space. Thus, source features can be augmented to effectively equip with target semantics to train a more transferable classifier. To achieve this, for each class, we first use the inter-domain feature mean difference and target intra-class feature covariance to construct a multivariate normal distribution. Then we augment source features with random directions sampled from the distribution class-wisely. Interestingly, such source augmentation is implicitly implemented through an expected transferable cross-entropy loss over the augmented source distribution, where an upper bound of the expected loss is derived and minimized, introducing negligible computational overhead. As a light-weight and general technique, TSA can be easily plugged into various domain adaptation methods, bringing remarkable improvements. Comprehensive experiments on cross-domain benchmarks validate the efficacy of TSA.
Shuang Li 0008, Mixue Xie, Kaixiong Gong, Chi Harold Liu, Yulin Wang 0002, Wei Li 0111
CVPR6
2021 Dynamic Domain Adaptation for Efficient Inference
abstract
Domain adaptation (DA) enables knowledge transfer from a labeled source domain to an unlabeled target domain by reducing the cross-domain distribution discrepancy. Most prior DA approaches leverage complicated and powerful deep neural networks to improve the adaptation capacity and have shown remarkable success. However, they may have a lack of applicability to real-world situations such as real-time interaction, where low target inference latency is an essential requirement under limited computational budget. In this paper, we tackle the problem by proposing a dynamic domain adaptation (DDA) framework, which can simultaneously achieve efficient target inference in low-resource scenarios and inherit the favorable cross-domain generalization brought by DA. In contrast to static models, as a simple yet generic method, DDA can integrate various domain confusion constraints into any typical adaptive network, where multiple intermediate classifiers can be equipped to infer “easier” and “harder” target data dynamically. Moreover, we present two novel strategies to further boost the adaptation performance of multiple prediction exits: 1) a confidence score learning strategy to derive accurate target pseudo labels by fully exploring the prediction consistency of different classifiers; 2) a class-balanced self-training strategy to explicitly adapt multi-stage classifiers from source to target without losing prediction diversity. Extensive experiments on multiple benchmarks are conducted to verify that DDA can consistently improve the adaptation performance and accelerate target inference under domain shift and limited resources scenarios.
Shuang Li 0008, Wenxuan Ma 0001, Chi Harold Liu, Wei Li 0111
CVPR5
2021 Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks
abstract
We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have demonstrated compelling results, most works estimate depth maps of individual video frames independently, without taking into consideration the strong geometric and temporal coherence among the frames. Moreover, current state-of-the-art (SOTA) models mostly adopt a fully 3D convolution network for cost regularization and therefore require high computational cost, thus limiting their deployment in real-world applications. Our method achieves temporally coherent depth estimation results by using a novel Epipolar Spatio-Temporal (EST) transformer to explicitly associate geometric and temporal correlation with multiple estimated depth maps. Furthermore, to reduce the computational cost, inspired by recent Mixture-of-Experts models, we design a compact hybrid network consisting of a 2D context-aware network and a 3D matching network which learn 2D context information and 3D disparity cues separately. Extensive experiments demonstrate that our method achieves higher accuracy in depth estimation and significant speedup than the SOTA methods.
Xiaoxiao Long, Lingjie Liu, Wei Li 0111, Christian Theobalt, Wenping Wang 0001
CVPR3
2021 Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR Segmentation
Xinge Zhu, Hui Zhou 0005, Fangzhou Hong, Yuexin Ma, Wei Li 0111, Hongsheng Li 0001, Dahua Lin
CVPR6
2021 Semantic Concentration for Domain Adaptation
abstract
Domain adaptation (DA) paves the way for label annotation and dataset bias issues by the knowledge transfer from a label-rich source domain to a related but unlabeled target domain. A mainstream of DA methods is to align the feature distributions of the two domains. However, the majority of them focus on the entire image features where irrelevant semantic information, e.g., the messy background, is inevitably embedded. Enforcing feature alignments in such case will negatively influence the correct matching of objects and consequently lead to the semantically negative transfer due to the confusion of irrelevant semantics. To tackle this issue, we propose Semantic Concentration for Domain Adaptation (SCDA), which encourages the model to concentrate on the most principal features via the pair-wise adversarial alignment of prediction distributions. Specifically, we train the classifier to class-wisely maximize the prediction distribution divergence of each sample pair, which enables the model to find the region with large differences among the same class of samples. Meanwhile, the feature extractor attempts to minimize that discrepancy, which suppresses the features of dissimilar regions among the same class of samples and accentuates the features of principal parts. As a general method, SCDA can be easily integrated into various DA methods as a regularizer to further boost their performance. Extensive experiments on the cross-domain benchmarks show the efficacy of SCDA.
Shuang Li 0008, Mixue Xie, Fangrui Lv, Chi Harold Liu, Jian Liang 0002, Chen Qin, Wei Li 0111
ICCV7
2021 Adaptive Surface Normal Constraint for Depth Estimation
abstract
We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimation quality. We therefore introduce a simple yet effective method, named Adaptive Surface Normal (ASN) constraint, to effectively correlate the depth estimation with geometric consistency. Our key idea is to adaptively determine the reliable local geometry from a set of randomly sampled candidates to derive surface normal constraint, for which we measure the consistency of the geometric contextual features. As a result, our method can faithfully reconstruct the 3D geometry and is robust to local shape variations, such as boundaries, sharp corners and noises. We conduct extensive evaluations and comparisons using public datasets. The experimental results demonstrate our method outperforms the state-of-the-art methods and has superior efficiency and robustness. Codes are available at: https://github.com/xxlong0/ASNDepth
Xiaoxiao Long, Cheng Lin 0001, Lingjie Liu, Wei Li 0111, Christian Theobalt, Ruigang Yang, Wenping Wang 0001
ICCV4
2021 CodeVIO: Visual-Inertial Odometry with Learned Optimizable Dense Depth
abstract
In this work, we present a lightweight, tightly-coupled deep depth network and visual-inertial odometry (VIO) system, which can provide accurate state estimates and dense depth maps of the immediate surroundings. Leveraging the proposed lightweight Conditional Variational Autoencoder (CVAE) for depth inference and encoding, we provide the network with previously marginalized sparse features from VIO to increase the accuracy of initial depth prediction and generalization capability. The compact representation of dense depth, termed depth code, can be updated jointly with navigation states in a sliding window estimator in order to provide the dense local scene geometry. We additionally propose a novel method to obtain the CVAE’s Jacobian which is shown to be more than an order of magnitude faster than previous works, and we additionally leverage First-Estimate Jacobian (FEJ) to avoid recalculation. As opposed to previous works that rely on completely dense residuals, we propose to only provide sparse measurements to update the depth code and show through careful experimentation that our choice of sparse measurements and FEJs can still significantly improve the estimated depth maps. Our full system also exhibits state-of-the-art pose estimation accuracy, and we show that it can run in real-time with single-thread execution while utilizing GPU acceleration only for the network and code Jacobian.
Xingxing Zuo 0001, Nathaniel W. Merrill, Wei Li 0111, Yong Liu 0007, Marc Pollefeys, Guoquan Huang 0001
ICRA3
2021 ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion
abstract
We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net. Specifically, the Siamese auto-encoder neural network is adopted to map the partial and complete input point cloud into a shared latent space, which can capture detailed shape prior. Then we design an iterative refinement unit to generate complete shapes with fine-grained details by integrating prior information. Experiments are conducted on the PCN dataset and the Completion3D benchmark, demonstrating the state-of-the-art performance of the proposed ASFM-Net. Our method achieves the 1st place in the leaderboard of Completion3D and outperforms existing methods with a large margin, about 12%. The codes and trained models are released publicly at https://github.com/Yan-Xia/ASFM-Net.
Yaqi Xia, Yan Xia 0003, Wei Li 0111, Rui Song 0003, Kailang Cao, Uwe Stilla
ACM Multimedia3
2021 Image Re-composition via Regional Content-Style Decoupling
abstract
Typical image composition harmonizes regions from different images to a single plausible image. We extend the idea of image composition by introducing the content-style decomposition and combination to form the concept of image re-composition. In other words, our image re-composition could arbitrarily combine those contents and styles decomposed from different images to generate more diverse images in a unified framework. In the decomposition stage, we incorporate the whitening normalization to obtain a more thorough content-style decoupling, which substantially improves the re-composition results. Moreover, to handle the variation of structure and texture of different objects in an image, we design the network to support regional feature representation and achieve region-aware content-style decomposition. Regarding the composition stage, we propose a cycle consistency loss to constrain the network preserving the content and style information during the composition. Our method can produce diverse re-composition results, including content-content, content-style and style-style. Our experimental results demonstrate a large improvement over the current state-of-the-art methods.
Wei Li 0111, Hong Zhang 0009, Ruigang Yang, Weiwei Xu 0003
ACM Multimedia2
2020 AutoRemover: Automatic Object Removal for Autonomous Driving Videos
abstract
Motivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm AutoRemover, designed specifically for generating street-view videos without any moving objects. In our setup we have two challenges: the first is the shadow, shadows are usually unlabeled but tightly coupled with the moving objects. The second is the large ego-motion in the videos. To deal with shadows, we build up an autonomous driving shadow dataset and design a deep neural network to detect shadows automatically. To deal with large ego-motion, we take advantage of the multi-source data, in particular the 3D data, in autonomous driving. More specifically, the geometric relationship between frames is incorporated into an inpainting deep neural network to produce high-quality structurally consistent video output. Experiments show that our method outperforms other state-of-the-art (SOTA) object removal algorithms, reducing the RMSE by over 19%.
Wei Li 0111, Peng Wang 0001, Chenye Guan, Yuhang Song 0003, Baoquan Chen, Weiwei Xu 0003, Ruigang Yang
AAAI2
2020 Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution
abstract
Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what would get from a real depth sensor. In this paper, we argue that DSR models trained under this setting are restrictive and not effective in dealing with realworld DSR tasks. We make two contributions in tackling real-world degradation of different depth sensors. First, we propose to classify the generation of LR depth maps into two types: non-linear downsampling with noise and interval downsampling, for which DSR models are learned correspondingly. Second, we propose a new framework for real-world DSR, which consists of four modules : 1) An iterative residual learning module with deep supervision to learn effective high-frequency components of depth maps in a coarse-to-fine manner; 2) A channel attention strategy to enhance channels with abundant high-frequency components; 3) A multi-stage fusion module to effectively reexploit the results in the coarse-to-fine process; and 4) A depth refinement module to improve the depth map by TGV regularization and input loss. Extensive experiments on benchmarking datasets demonstrate the superiority of our method over current state-of-the-art DSR methods.
Xibin Song, Yuchao Dai, Dingfu Zhou, Liu Liu 0009, Wei Li 0111, Hongdong Li, Ruigang Yang
CVPR5
2020 DVI: Depth Guided Video Inpainting for Autonomous Driving
Miao Liao, Feixiang Lu, Dingfu Zhou, Wei Li 0111, Ruigang Yang
ECCV (21)5
2020 Interactive free-viewpoint video generation
abstract
Free-viewpoint video (FVV) is processed video content in which viewers can freely select the viewing position and angle. FVV delivers an improved visual experience and can also help synthesize special effects and virtual reality content. In this paper, a complete FVV system is proposed to interactively control the viewpoints of video relay programs through multimedia terminals such as computers and tablets. The hardware of the FVV generation system is a set of synchronously controlled cameras, and the software generates videos in novel viewpoints from the captured video using view interpolation. The interactive interface is designed to visualize the generated video in novel viewpoints and enable the viewpoint to be changed interactively. Experiments show that our system can synthesize plausible videos in intermediate viewpoints with a view range of up to 180°.
Hao Zhu 0004, Wei Li 0111, Xun Cao, Ruigang Yang
Virtual Real. Intell. Hardw.4
2019 Fast Texture Mapping Adjustment via Local/Global Optimization
abstract
This paper deals with the texture mapping of a triangular mesh model given a set of calibrated images. Different from the traditional approach of applying projective texture mapping with model parameterizations, we develop an image-space texture optimization scheme that aims to reduce visible seams or misalignment at texture or depth boundaries. Our novel scheme starts with an efficient local (and parallel) texture adjustment scheme at these boundaries, followed by a global correction step to rectify potential texture distortions caused by the local movement. Our phased optimization scheme achieves 50$\sim$∼100 times speed up on GPU (or 6× on CPU) compared to previous state-of-the-art methods. Experiments on a variety of models showed that we achieve this significant speed-up without sacrificing texture quality. Our approach significantly improves resilience to modeling and calibration errors, thereby allowing fast and fully automatic creation of textured models using commodity depth sensors by untrained users.
Wei Li 0111, Huajun Gong, Ruigang Yang
IEEE Trans. Vis. Comput. Graph.1
2018 Inexact descent methods for elastic parameter optimization
abstract
Elastic parameter optimization has revealed its importance in 3D modeling, virtual reality, and additive manufacturing in recent years. Unfortunately, it is known to be computationally expensive, especially if there are many parameters and data samples. To address this challenge, we propose to introduce the inexactness into descent methods, by iteratively solving a forward simulation step and a parameter update step in an inexact manner. The development of such inexact descent methods is centered at two questions: 1) how accurate/inaccurate can the two steps be; and 2) what is the optimal way to implement an inexact descent method. The answers to these questions are in our convergence analysis, which proves the existence of relative error thresholds for the two inexact steps to ensure the convergence. This means we can simply solve each step by a fixed number of iterations, if the iterative solver is at least linearly convergent. While the use of the inexact idea speeds up many descent methods, we specifically favor a GPU-based one powered by state-of-the-art simulation techniques. Based on this method, we study a variety of implementation issues, including backtracking line search, initialization, regularization, and multiple data samples. We demonstrate the use of our inexact method in elasticity measurement and design applications. Our experiment shows the method is fast, reliable, memory-efficient, GPU-friendly, flexible with different elastic models, scalable to a large parameter space, and parallelizable for multiple data samples.
Wei Li 0111, Ruigang Yang, Huamin Wang 0001
ACM Trans. Graph.2