Zhenyu Wang 0005

dblp:22/1486-5 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0003-3701-8063ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Image recognition and object detection · 42% 3D vision · 23% Trustworthy machine learning · 8%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 24 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
2.232024
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection · NeurIPS 2024
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation · ECCV (47) 2024
Uni3DETR: Unified 3D Detection Transformer · NeurIPS 2023
Computer vision › Image recognition and object detection
object detection
1.932025
UniDetector: Towards Universal Object Detection With Heterogeneous Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification · NeurIPS 2021
Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection · CVPR 2021
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
1.622025
UniDetector: Towards Universal Object Detection With Heterogeneous Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation · ECCV (47) 2024
Computer vision › Image recognition and object detection › object detection
universal object detection
1.522025
UniDetector: Towards Universal Object Detection With Heterogeneous Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Detecting Everything in the Open World: Towards Universal Object Detection · CVPR 2023
Computer vision › 3D vision › 3d object detection
point cloud object detection
1.422024
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection · NeurIPS 2024
Uni3DETR: Unified 3D Detection Transformer · NeurIPS 2023
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection
1.022021
Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification · NeurIPS 2021
Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection · CVPR 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
1.022021
Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification · NeurIPS 2021
Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection · CVPR 2021
Computer vision › Vision and language › cross-modal alignment
image-text alignment
0.912025
UniDetector: Towards Universal Object Detection With Heterogeneous Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › object detection › open-vocabulary object detection
zero-shot object detection
0.912025
UniDetector: Towards Universal Object Detection With Heterogeneous Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision › 3d object detection
open-vocabulary 3d object detection
0.812024
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation · ECCV (47) 2024
Computer vision › Image recognition and object detection › object detection
detection transformer
0.712023
Uni3DETR: Unified 3D Detection Transformer · NeurIPS 2023
Machine learning › Graph learning
graph neural network
0.712023
Uncertainty-aware Consistency Learning for Cold-Start Item Recommendation · SIGIR 2023
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-world detection
0.712023
Detecting Everything in the Open World: Towards Universal Object Detection · CVPR 2023
Machine learning › Graph learning
recommendation
0.712023
Uncertainty-aware Consistency Learning for Cold-Start Item Recommendation · SIGIR 2023
Recommender systems › cold-start recommendation
cold-start item recommendation
0.712023
Uncertainty-aware Consistency Learning for Cold-Start Item Recommendation · SIGIR 2023
Robotics › Robot manipulation › grasping › grasp detection
grasp pose estimation
0.612022
Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022
Robotics › Robot manipulation › grasping
grasp quality evaluation
0.612022
Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022
Computer vision › Segmentation and scene understanding
instance segmentation
0.612022
Noisy Boundaries: Lemon or Lemonade for Semi-supervised Instance Segmentation? · CVPR 2022
Computer vision › Segmentation and scene understanding › instance segmentation
semi-supervised instance segmentation
0.612022
Noisy Boundaries: Lemon or Lemonade for Semi-supervised Instance Segmentation? · CVPR 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertain data
0.512021
Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection · CVPR 2021
Computer vision › 3D vision
3d scene understanding
0.212023
Uni3DETR: Unified 3D Detection Transformer · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212023
Detecting Everything in the Open World: Towards Universal Object Detection · CVPR 2023
Robotics › Robot manipulation › grasping › grasp stability
force-closure grasp
0.212022
Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
noisy pseudo-label learning
0.112021
Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

pseudo-labeling · 1.6probability calibration · 1.5image-text alignment · 0.9heterogeneous supervision · 0.9sparse convolution · 0.8routing mechanism · 0.8cycle-modality propagation · 0.8anchor-free detection · 0.8vision-language alignment · 0.7uncertainty modeling · 0.7teacher-student training · 0.7decoupling training · 0.7consistency learning · 0.7
YearPublicationVenuePosition
2025 UniDetector: Towards Universal Object Detection With Heterogeneous Supervision
abstract
In this paper, we formally address universal object detection, which aims to detect every category in every scene. The dependence on human annotations, the limited visual information, and the novel categories in open world severely restrict the universality of detectors. We propose UniDetector, a universal object detector that recognizes enormous categories in the open world. The critical points for UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces in training through image-text alignment, which guarantees sufficient information for universal representations. 2) it involves heterogeneous supervision training, which alleviates the dependence on the limited fully-labeled images. 3) it generalizes to open world easily while keeping the balance between seen and unseen classes. 4) it further promotes generalizing to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7 k categories, the largest measurable size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot ability on large-vocabulary datasets - it surpasses supervised baselines by more than 5% without seeing any corresponding images. On 13 detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data.
Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
Zhenyu Wang 0005, Yali Li 0001, Taichi Liu, Hengshuang Zhao, Shengjin Wang
ECCV (47)1
2024 One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
abstract
The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however, such multi-domain joint training is highly challenging, because large domain gaps among point clouds from different datasets lead to the severe domain-interference problem. In this paper, we propose OneDet3D, a universal one-for-all model that addresses 3D detection across different domains, including diverse indoor and outdoor scenes, within the same framework and only one set of parameters. We propose the domain-aware partitioning in scatter and context, guided by a routing mechanism, to address the data interference issue, and further incorporate the text modality for a language-guided classification to unify the multi-dataset label spaces and mitigate the category interference issue. The fully sparse structure and anchor-free head further accommodate point clouds with significant scale disparities. Extensive experiments demonstrate the strong universal ability of OneDet3D to utilize only one trained model for addressing almost all 3D object detection tasks (Fig. 1). We will open-source the code for future research and applications.
Zhenyu Wang 0005, Yali Li 0001, Hengshuang Zhao, Shengjin Wang
NeurIPS1
2023 Detecting Everything in the Open World: Towards Universal Object Detection
abstract
In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We propose UniDetector, a universal object detector that has the ability to recognize enormous categories in the open world. The critical points for the universality of UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces for training through the alignment of image and text spaces, which guarantees sufficient information for universal representations. 2) it generalizes to the open world easily while keeping the balance between seen and unseen classes, thanks to abundant information from both vision and language modalities. 3) it further promotes the generalization ability to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7k categories, the largest measurable category size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot generalization ability on largevocabulary datasets - it surpasses the traditional supervised baselines by more than 4% on average without seeing any corresponding images. On 13 public detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data.11Codes are available at https://github.com/zhenyuw16/UniDetector.
Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang
CVPR1
2023 Uni3DETR: Unified 3D Detection Transformer
abstract
Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, there is still a lack of a unified network architecture that can accommodate diverse scenes. In this paper, we propose Uni3DETR, a unified 3D detector that addresses indoor and outdoor 3D detection within the same framework. Specifically, we employ the detection transformer with point-voxel interaction for object prediction, which leverages voxel features and points for cross-attention and behaves resistant to the discrepancies from data. We then propose the mixture of query points, which sufficiently exploits global information for dense small-range indoor scenes and local information for large-range sparse outdoor ones. Furthermore, our proposed decoupled IoU provides an easy-to-optimize training target for localization by disentangling the $xy$ and $z$ space. Extensive experiments validate that Uni3DETR exhibits excellent performance consistently on both indoor and outdoor 3D detection. In contrast to previous specialized detectors, which may perform well on some particular datasets but suffer a substantial degradation on different scenes, Uni3DETR demonstrates the strong generalization ability under heterogeneous conditions (Fig. 1).
Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Hengshuang Zhao, Shengjin Wang
NeurIPS1
2023 Uncertainty-aware Consistency Learning for Cold-Start Item Recommendation
abstract
Graph Neural Network (GNN)-based models have become the mainstream approach for recommender systems. Despite the effectiveness, they are still suffering from the cold-start problem, i.e., recommend for few-interaction items. Existing GNN-based recommendation models to address the cold-start problem mainly focus on utilizing auxiliary features of users and items, leaving the user-item interactions under-utilized. However, embeddings distributions of cold and warm items are still largely different, since cold items' embeddings are learned from lower-popularity interactions, while warm items' embeddings are from higher-popularity interactions. Thus, there is a seesaw phenomenon, where the recommendation performance for the cold and warm items cannot be improved simultaneously. To this end, we proposed a Uncertainty-aware Consistency learning framework for Cold-start item recommendation (shorten as UCC) solely based on user-item interactions. Under this framework, we train the teacher model (generator) and student model (recommender) with consistency learning, to ensure the cold items with additionally generated low-uncertainty interactions can have similar distribution with the warm items. Therefore, the proposed framework improves the recommendation of cold and warm items at the same time, without hurting any one of them. Extensive experiments on benchmark datasets demonstrate that our proposed method significantly outperforms state-of-the-art methods on both warm and cold items, with an average performance improvement of 27.6%.
Taichi Liu, Chen Gao 0001, Zhenyu Wang 0005, Dong Li 0016, Jianye Hao, Depeng Jin, Yong Li 0008
SIGIR3
2022 Noisy Boundaries: Lemon or Lemonade for Semi-supervised Instance Segmentation?
abstract
Current instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employed to boost the performance. We construct a framework for semi-supervised instance segmentation by assigning pixel-level pseudo labels. Under this framework, we point out that noisy boundaries associated with pseudo labels are double-edged. We propose to exploit and resist them in a unified manner simultaneously: 1) To combat the negative effects of noisy boundaries, we propose a noise-tolerant mask head by leveraging low-resolution features. 2) To enhance the positive impacts, we introduce a boundary-preserving map for learning detailed information within boundary-relevant regions. We evaluate our approach by extensive experiments. It behaves extraordinarily, outperforming the supervised baseline by a large margin, more than 6% on Cityscapes, 7% on COCO and 4.5% on BDD100k. On Cityscapes, our method achieves comparable performance by utilizing only 30% labeled images.
Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang
CVPR1
2022 Hybrid Physical Metric For 6-DoF Grasp Pose Detection
abstract
6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric usually generates several discrete levels of grasp confidence scores, which cannot finely distinguish millions of grasp poses and leads to inaccurate prediction results. In this paper, we propose a hybrid physical metric to solve this evaluation insufficiency. First, we define a novel metric is based on the force-closure metric, supplemented by the measurement of the object flatness, gravity and collision. Second, we leverage this hybrid physical metric to generate elaborate confidence scores. Third, to learn the new confidence scores effectively, we design a multi-resolution network called Flatness Gravity Collision GraspNet (FGC-GraspNet). FGC-GraspNet proposes a multi-resolution features learning architecture for multiple tasks and introduces a new joint loss function that enhances the average precision of the grasp detection. The network evaluation and adequate real robot experiments demonstrate the effectiveness of our hybrid physical metric and FGC-GraspNet. Our method achieves 90.5% success rate in real-world cluttered scenes. Our code is available at https://github.com/luyh20IFGC-GraspNet.
Yuhao Lu, Beixing Deng, Zhenyu Wang 0005, Peiyuan Zhi, Yali Li 0001, Shengjin Wang
ICRA3
2021 Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
abstract
In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy labels, thus are deficient in learning different unlabeled knowledge well. To address this issue, we propose a data-uncertainty guided multi-phase learning method for semisupervised object detection. We comprehensively consider divergent types of unlabeled images according to their difficulty levels, utilize them in different phases, and ensemble models from different phases together to generate ultimate results. Image uncertainty guided easy data selection and region uncertainty guided RoI Re-weighting are involved in multi-phase learning and enable the detector to concentrate on more certain knowledge. Through extensive experiments on PASCAL VOC and MS COCO, we demonstrate that our method behaves extraordinarily compared to baseline approaches and outperforms them by a large margin, more than 3% on VOC and 2% on COCO.
Zhenyu Wang 0005, Yali Li 0001, Lu Fang 0001, Shengjin Wang
CVPR1
2021 Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification
abstract
Semi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current methods are easily distracted by noisy regions generated by pseudo labels. To combat the noisy labeling, we propose noise-resistant semi-supervised learning by quantifying the region uncertainty. We first investigate the adverse effects brought by different forms of noise associated with pseudo labels. Then we propose to quantify the uncertainty of regions by identifying the noise-resistant properties of regions over different strengths. By importing the region uncertainty quantification and promoting multi-peak probability distribution output, we introduce uncertainty into training and further achieve noise-resistant learning. Experiments on both PASCAL VOC and MS COCO demonstrate the extraordinary performance of our method.
Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang
NeurIPS1