Kewei Wang 0001

dblp:253/2810-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2024
0000-0001-9433-720XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix
abstract
Class-agnostic motion prediction methods aim to comprehend motion within open-world scenarios, holding significance for autonomous driving systems. However, training a high-performance model in a fully-supervised manner always requires substantial amounts of manually annotated data, which can be both expensive and time-consuming to obtain. To address this challenge, our study explores the potential of semi-supervised learning (SSL) for class-agnostic motion prediction. Our SSL framework adopts a consistency-based self-training paradigm, enabling the model to learn from unlabeled data by generating pseudo labels through test-time inference. To improve the quality of pseudo labels, we propose a novel motion selection and re-generation module. This module effectively selects reliable pseudo labels and re-generates unreliable ones. Furthermore, we propose two data augmentation strategies: temporal sampling and BEVMix. These strategies facilitate consistency regularization in SSL. Experiments conducted on nuScenes demonstrate that our SSL method can surpass the self-supervised approach by a large margin by utilizing only a tiny fraction of labeled data. Furthermore, our method exhibits comparable performance to weakly and some fully supervised methods. These results highlight the ability of our method to strike a favorable balance between annotation costs and performance. Code will be available at https://github.com/kwwcv/SSMP.
Kewei Wang 0001, Yizheng Wu, Xingyi Li 0005, Ke Xian, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
AAAI1
2024 S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes
abstract
Current 3D stylization methods often assume static scenes, which violates the dynamic nature of our real world. To address this limitation, we present S-DyRF, a reference-based spatio-temporal stylization method for dynamic neu-ral radiance fields. However, stylizing dynamic 3D scenes is inherently challenging due to the limited availability of stylized reference images along the temporal axis. Our key insight lies in introducing additional temporal cues besides the provided reference. To this end, we generate temporal pseudo-references from the given stylized reference. These pseudo-references facilitate the propagation of style infor-mation from the reference to the entire dynamic 3D scene. For coarse style transfer, we enforce novel views and times to mimic the style details present in pseudo-references at the feature level. To preserve high-frequency details, we create a collection of stylized temporal pseudo-rays from temporal pseudo-references. These pseudo-rays serve as detailed and explicit stylization guidance for achieving fine style trans-fer. Experiments on both synthetic and real-world datasets demonstrate that our method yields plausible stylized re-sults of space-time view synthesis on dynamic 3D scenes.
Xingyi Li 0005, Zhiguo Cao 0001, Yizheng Wu, Kewei Wang 0001, Ke Xian, Zhe Wang 0006, Guosheng Lin
CVPR4
2024 Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations
abstract
The perception of motion behavior in a dynamic environment holds significant importance for autonomous driving systems, wherein class-agnostic motion prediction methods directly predict the motion of the entire point cloud. While most existing methods rely on fully-supervised learning, the manual labeling of point cloud data is laborious and time-consuming. Therefore, several annotation-efficient methods have been proposed to address this challenge. Al-though effective, these methods rely on weak annotations or additional multi-modal data like images, and the potential benefits inherent in the point cloud sequence are still underexplored. To this end, we explore the feasibility of self-supervised motion prediction with only unlabeled Li-DAR point clouds. Initially, we employ an optimal transport solver to establish coarse correspondences between current and future point clouds as the coarse pseudo motion labels. Training models directly using such coarse labels leads to noticeable spatial and temporal prediction in-consistencies. To mitigate these issues, we introduce three simple spatial and temporal regularization losses, which fa-cilitate the self-supervised training process effectively. Experimental results demonstrate the significant superiority of our approach over the state-of-the-art self-supervised methods. Code will be available at https://github.com/kwwcv/SelfMotion.
Kewei Wang 0001, Yizheng Wu, Jun Cen, Xingyi Li 0005, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
CVPR1
2024 iControl3D: An Interactive System for Controllable 3D Scene Generation
abstract
3D content creation has long been a complex and time-consuming process, often requiring specialized skills and resources. While re- cent advancements have allowed for text-guided 3D object and scene generation, they still fall short of providing sufficient control over the generation process, leading to a gap between the user’s creative vision and the generated results. In this paper, we present iControl3D, a novel interactive system that empowers users to gen- erate and render customizable 3D scenes with precise control. To this end, a 3D creator interface has been developed to provide users with fine-grained control over the creation process. Technically, we leverage 3D meshes as an intermediary proxy to iteratively merge individual 2D diffusion-generated images into a cohesive and uni- fied 3D scene representation. To ensure seamless integration of 3D meshes, we propose to perform boundary-aware depth alignment before fusing the newly generated mesh with the existing one in 3D space. Additionally, to effectively manage depth discrepancies between remote content and foreground, we propose to model re- mote content separately with an environment map instead of 3D meshes. Finally, our neural rendering interface enables users to build a radiance field of their scene online and navigate the entire scene. Extensive experiments have been conducted to demonstrate the effectiveness of our system. The code will be made available at https://github.com/xingyi- li/iControl3D.
Xingyi Li 0005, Yizheng Wu, Jun Cen, Juewen Peng, Kewei Wang 0001, Ke Xian, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
ACM Multimedia5
2024 Instance Consistency Regularization for Semi-Supervised 3D Instance Segmentation
abstract
Large-scale datasets with point-wise semantic and instance labels are crucial to 3D instance segmentation but also expensive. To leverage unlabeled data, previous semi-supervised 3D instance segmentation approaches have explored self-training frameworks, which rely on high-quality pseudo labels for consistency regularization. They intuitively utilize both instance and semantic pseudo labels in a joint learning manner. However, semantic pseudo labels contain numerous noise derived from the imbalanced category distribution and natural confusion of similar but distinct categories, which leads to severe collapses in self-training. Motivated by the observation that 3D instances are non-overlapping and spatially separable, we ask whether we can solely rely on instance consistency regularization for improved semi-supervised segmentation. To this end, we propose a novel self-training network InsTeacher3D to explore and exploit pure instance knowledge from unlabeled data. We first build a parallel base 3D instance segmentation model DKNet, which distinguishes each instance from the others via discriminative instance kernels without reliance on semantic segmentation. Based on DKNet, we further design a novel instance consistency regularization framework to generate and leverage high-quality instance pseudo labels. Experimental results on multiple large-scale datasets show that the InsTeacher3D significantly outperforms prior state-of-the-art semi-supervised approaches.
Yizheng Wu, Kewei Wang 0001, Xingyi Li 0005, Jiahao Cui 0002, Liwen Xiao, Guosheng Lin, Zhiguo Cao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Pseudo Label Fusion With Uncertainty Estimation for Semi-Supervised Cropping Box Regression
abstract
Cropping box regression algorithms re-frame the images with predicted cropping boxes for better composition quality, which can save considerable manpower and time for massive image retouching work. Yet, recent learning-based cropping box regression algorithms require expert annotations, which makes the scale of training limited. This consequently incurs a performance bottleneck. To address this issue, previous works seek the help from auxiliary datasets of related tasks,e.g., the composition classification. However, the domain gap between related tasks and the likewise restricted scale of auxiliary datasets are still limiting factors. Hence, our work provides a novel semi-supervised framework that can learn better re-framing knowledge with unlimited unlabeled data. We make use of the unlabeled data via pseudo-labeling, where the model learns from the pseudo labels generated from a temporal ensemble version of itself. To prevent the model learns from its own mistakes,a.k.a. the problem of confirmation bias, we propose to rectify the mistakes by fusing multiple candidate pseudo labels into the better ones. The fusion procedure is based on the uncertainty estimation for each boundary of the candidate cropping boxes. The multiple candidates are from the proposed aesthetic region proposal network. Extensive experimental results explain how the uncertainty-based pseudo label fusion procedure overcomes the confirmation bias and demonstrate the superiority of our semi-supervised cropping box regression framework.
Jiahao Cui 0002, Kewei Wang 0001, Yizheng Wu, Zhiguo Cao 0001
IEEE Trans. Multim.3
2023 BPR-Net: Balancing Precision and Recall for Infrared Small Target Detection
abstract
Most current infrared small target detection methods attempt to fuse local and global information by using single-scale inputs and creating a multi-scale feature pyramid during network feeding forwards. Further to this, our research finds that using high-resolution inputs can improve recall, while low-resolution inputs improve precision. Nevertheless, solely focusing on global or local information can result in missing target and false alarm. To address these issues, we propose the BPR-Net to balance precision and recall via a novel multi-scale attention mechanism, which combines semantic and shallow features of multi-scale inputs. We first scale the input image into multiple images with varying resolutions and feed them into the network. In the encoder, Scale Fusion Module (SFM) fuses features from corresponding images of different resolutions. In the decoder, a Channel Fusion Module (CFM) fuses useful information from multiple channels. Furthermore, a Wavelet Transform cross-layer skip Layer (WTL) is employed to enhance the interaction between decoder layers for more effective multi-scale feature fusion. Experimental results demonstrate that our approach achieves a balance between recall and precision and yields state-of-the-art performance on challenging benchmarks including Sirst, MDvsFA, and SIATD. Notably, our approach achieves an F1 score of 0.9409 on the challenging benchmark SIATD, surpassing the state-of-the-art method by 16.7%.
Shuaiyuan Du, Kewei Wang 0001, Zhiguo Cao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Robust Object Detection with Inaccurate Bounding Boxes
Kewei Wang 0001, Hao Lu 0003, Zhiguo Cao 0001
ECCV (10)2
2022 Interior Attention-Aware Network for Infrared Small Target Detection
abstract
Infrared small target detection plays an important role in target warning, ground monitoring, and flight guidance. Existing methods typically utilize local-contrast information of each pixel to detect infrared small targets, neglecting the interior relation between target pixels or background pixels. The mere use of the local information of one pixel, however, is not sufficient for accurate detection, which may lead to missing detection and false alarms. As a harmonious whole, information between pixels are necessary to determine if a pixel belongs to the target or the background. Motivated by the fact that pixels from targets or backgrounds are correlated with each other, we propose a coarse-to-fine interior attention-aware network (IAANet) for infrared small target detection. Specifically, a region proposal network (RPN) is first applied to obtain coarse target regions and filter out backgrounds. Then, we leverage a transformer encoder to model the attention between pixels in coarse target regions, outputting attention-aware features. Finally, predictions are obtained by feeding attention-aware features to a classification head. Extensive experiments show that our approach is capable of detecting targets precisely, of suppressing a variety of false alarm sources, and works effectively in various background environments and target appearances. We show that our IAANet outperforms the state-of-the-art methods by a large margin. Code will be made available at:https://github.com/kwwcv/iaanet.
Kewei Wang 0001, Shuaiyuan Du, Zhiguo Cao 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 TransView: Inside, Outside, and Across the Cropping View Boundaries
abstract
We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual difference is subtle when a visual element is reserved or removed. Existing methods represent visual elements by extracting region-based convolutional features inside and outside the cropping view boundaries, without probing a fundamental question: why some visual elements are of interest or of discard? In this work, we observe that the relation between different visual elements significantly affects their relative positions to the desired cropping view, and such relation can be characterized by the attraction inside/outside the cropping view boundaries and the repulsion across the boundaries. By instantiating a transformer-based solution that represents visual elements as visual words and that models the dependencies between visual words, we report not only state-of-the-art performance on public benchmarks, but also interesting visualizations that depict the attraction and repulsion between visual elements, which may shed light on what makes for effective cropping view recommendation.
Zhiguo Cao 0001, Kewei Wang 0001, Hao Lu 0003, Weicai Zhong
ICCV3