Jiabao Wang 0005

dblp:37/8167-5 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-4951-3088ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Image recognition and object detection · 52% Efficient and distributed learning · 14% 3D vision · 14%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
2.132025
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
CrossKD: Cross-Head Knowledge Distillation for Object Detection · CVPR 2024
Oriented R-CNN for Object Detection · ICCV 2021
Computer vision › Image recognition and object detection › object detection
oriented object detection
1.322024
Oriented R-CNN and Beyond · Int. J. Comput. Vis. 2024
Oriented R-CNN for Object Detection · ICCV 2021
Machine learning › Deep learning architectures and training
multi-scale representation
0.912025
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Representation and self-supervised learning › representation learning
multi-scale representation learning
0.912025
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › object detection › efficient object detection
real-time object detection
0.912025
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision
3d object detection
0.812024
Towards Stable 3D Object Detection · ECCV (50) 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
CrossKD: Cross-Head Knowledge Distillation for Object Detection · CVPR 2024
Computer vision › Image recognition and object detection › object detection
knowledge distillation for detection
0.812024
CrossKD: Cross-Head Knowledge Distillation for Object Detection · CVPR 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
CrossKD: Cross-Head Knowledge Distillation for Object Detection · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.812024
OPUS: Occupancy Prediction Using a Sparse Set · NeurIPS 2024
Computer vision › Image recognition and object detection › object detection
region-based detection
0.512021
Oriented R-CNN for Object Detection · ICCV 2021
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
0.212024
Towards Stable 3D Object Detection · ECCV (50) 2024

Methods — techniques the papers use, named apart from their topics

multi-branch convolution · 0.9kernel size design · 0.9transformer encoder-decoder · 0.8oriented R-CNN · 0.8nearest neighbor search · 0.8feature imitation · 0.8cross-head prediction mimicking · 0.8coarse-to-fine learning · 0.8chamfer distance loss · 0.83d object detection · 0.8
YearPublicationVenuePosition
2025 YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection
abstract
We aim at providing the object detection community with an efficient and performant object detector, termed YOLO-MS. The core design is based on a series of investigations on how multi-branch features of the basic block and convolutions with different kernel sizes affect the detection performance of objects at different scales. The outcome is a new strategy that can significantly enhance multi-scale feature representations of real-time object detectors. To verify the effectiveness of our work, we train our YOLO-MS on the MS COCO dataset from scratch without relying on any other large-scale datasets, like ImageNet or pre-trained weights. Without bells and whistles, our YOLO-MS outperforms the recent state-of-the-art real-time object detectors, including YOLO-v7, RTMDet, and YOLO-v8. Taking the XS version of YOLO-MS as an example, it can achieve an AP score of 42+% on MS COCO, which is about 2% higher than RTMDet with the same model size. Furthermore, our work can also serve as a plug-and-play module for other YOLO models. Typically, our method significantly advances the APs, APl, and AP of YOLOv8-N from 18%+, 52%+, and 37%+ to 20%+, 55%+, and 40%+, respectively, with even fewer parameters and MACs.
Xinbin Yuan, Jiabao Wang 0005, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 CrossKD: Cross-Head Knowledge Distillation for Object Detection
abstract
Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper, we present a general and effective prediction mimicking distillation scheme, called CrossKD, which delivers the intermediate features of the student's detection head to the teacher's detection head. The resulting crosshead predictions are then forced to mimic the teacher's predictions. This manner relieves the student's head from receiving contradictory supervision signals from the annotations and the teacher's predictions, greatly improving the student's detection performance. Moreover, as mimicking the teacher's predictions is the target of KD, CrossKD offers more task-oriented information in contrast with feature imitation. On MS COCO, with only prediction mimicking losses applied, our CrossKD boosts the average precision of GFL ResNet-50 with 1 × training schedule from 40.2 to 43.7, outperforming all existing KD methods. In addition, our method also works well when distilling detectors with heterogeneous backbones.
Jiabao Wang 0005, Zhaohui Zheng 0003, Xiang Li 0041, Ming-Ming Cheng, Qibin Hou
CVPR1
2024 Towards Stable 3D Object Detection
Jiabao Wang 0005, Guochao Liu, Liujiang Yan, Ming-Ming Cheng, Qibin Hou
ECCV (50)1
2024 OPUS: Occupancy Prediction Using a Sparse Set
abstract
Occupancy prediction, aiming at predicting the occupancy status within voxelized 3D environment, is quickly gaining momentum within the autonomous driving community. Mainstream occupancy prediction works first discretize the 3D environment into voxels, then perform classification on such dense grids. However, inspection on sample data reveals that the vast majority of voxels is unoccupied. Performing classification on these empty voxels demands suboptimal computation resource allocation, and reducing such empty voxels necessitates complex algorithm designs. To this end, we present a novel perspective on the occupancy prediction task: formulating it as a streamlined set prediction paradigm without the need for explicit space modeling or complex sparsification procedures. Our proposed framework, called OPUS, utilizes a transformer encoder-decoder architecture to simultaneously predict occupied locations and classes using a set of learnable queries. Firstly, we employ the Chamfer distance loss to scale the set-to-set comparison problem to unprecedented magnitudes, making training such model end-to-end a reality. Subsequently, semantic classes are adaptively assigned using nearest neighbor search based on the learned locations. In addition, OPUS incorporates a suite of non-trivial strategies to enhance model performance, including coarse-to-fine learning, consistent point sampling, and adaptive re-weighting, etc. Finally, compared with current state-of-the-art methods, our lightest model achieves superior RayIoU on the Occ3D-nuScenes dataset at near 2x FPS, while our heaviest model surpasses previous best results by 6.1 RayIoU.
Jiabao Wang 0005, Zhaojiang Liu, Liujiang Yan, Qibin Hou, Ming-Ming Cheng
NeurIPS1
2024 Oriented R-CNN and Beyond
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xiwen Yao, Junwei Han 0001
Int. J. Comput. Vis.3
2023 Learning Orientation-Aware Distances for Oriented Object Detection
abstract
Oriented object detectors have suffered severely from the discontinuous boundary problem for a long time. In this work, we ingeniously avoid this problem by relating regression outputs to regression target orientations. The core idea of our method is to build a contour function which imports orientations and outputs the corresponding distance predictions. Inspired by Fourier transformations, we assume this function can be represented as a linear combination of trigonometric functions and Fourier series. We replace the final 4D layer in the regression branch of fully convolutional one-stage object detector (FCOS) with a Fourier Series Transformation (FST) module and term this new network FCOSF. By this unique design, the regression outputs in FCOSF can adaptively vary according to the regression target orientations. Thus, the discontinuous boundary has no impact on our FCOSF. More importantly, FCOSF avoids building complicated oriented box representations, which usually cause extra computations and ambiguities. With only flipping augmentation and single-scale training and testing, FCOSF with ResNet-50 achieves 73.64% mAP on the DOTA-v1.0 dataset with up to 23.6 FPS speed, surpassing all one-stage oriented object detectors. On the more challenging DOTA-v2.0 dataset, FCOSF also achieves the highest results of 51.75% mAP among one-stage detectors. More experiments on DIOR-R and HRSC2016 are also conducted to verify the robustness of FCOSF. Code and models will be available at https://github.com/DDGRCF/FCOSF.
Chaofan Rao, Jiabao Wang 0005, Gong Cheng 0003, Xingxing Xie, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Anchor-Free Oriented Proposal Generator for Object Detection
abstract
Oriented object detection is a practical and challenging task in remote sensing image interpretation. Nowadays, oriented detectors mostly use horizontal boxes as intermedium to derive oriented boxes from them. However, the horizontal boxes are inclined to get small Intersection-over-Unions (IoUs) with ground truths, which may have some undesirable effects, such as introducing redundant noise, mismatching with ground truths, detracting from the robustness of detectors, etc. In this paper, we propose a novel Anchor-free Oriented Proposal Generator (AOPG) that abandons horizontal box-related operations from the network architecture. AOPG first produces coarse oriented boxes by a Coarse Location Module (CLM) in an anchor-free manner and then refines them into high-quality oriented proposals. After AOPG, we apply a Fast R-CNN head to produce the final detection results. Furthermore, the shortage of large-scale datasets is also a hindrance to the development of oriented object detection. To alleviate the data insufficiency, we release a new dataset on the basis of our DIOR dataset and name it DIOR-R. Massive experiments demonstrate the effectiveness of AOPG. Particularly, without bells and whistles, we achieve the accuracy of 64.41%, 75.24% and 96.22% mAP on the DIOR-R, DOTA and HRSC2016 datasets respectively. Code and models are available at https://github.com/jbwang1997/AOPG.
Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xingxing Xie, Chunbo Lang, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Dual-Aligned Oriented Detector
abstract
In the past few years, object detection in remote sensing images has achieved remarkable progress. However, the detection of oriented and densely packed objects are still unsatisfactory due to the following spatial and feature misalignments. 1) Most two-stage oriented detectors only introduce an orientation regression branch in the detection head, while still leverage horizontal proposals for classification and regression. This inevitably results in the spatial misalignment problem between horizontal proposals and oriented objects. 2) The features used for classification are in fact extracted from the region proposals which have shifted to the final predictions via the regression branch. This leads to the feature misalignment problem between the classification and the localization tasks. In this article, we present a two-stage oriented object detection method, termed dual-aligned oriented detector (DODet), toward evading the aforementioned problems of spatial and feature misalignments. In DODet, the first stage is an oriented proposal network (OPN), which generates high-quality oriented proposals via a novel representation scheme of oriented objects. The second stage is a localization-guided detection head (LDH) that aims at alleviating the feature misalignment between classification and localization. Comprehensive and extensive evaluations on three benchmarks, including DIOR-R, DOTA, and HRSC2016, indicate that our method could obtain consistent and substantial gains compared with the baseline method. The source code is publicly available athttps://github.com/yanqingyao1994/DODet.
Gong Cheng 0003, Shengyang Li, Ke Li 0005, Xingxing Xie, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 Oriented R-CNN for Object Detection
abstract
Current state-of-the-art two-stage detectors generate oriented proposals through time-consuming schemes. This diminishes the detectors’ speed, thereby becoming the computational bottleneck in advanced oriented object detection systems. This work proposes an effective and simple oriented object detection framework, termed Oriented R-CNN, which is a general two-stage oriented detector with promising accuracy and efficiency. To be specific, in the first stage, we propose an oriented Region Proposal Network (oriented RPN) that directly generates high-quality oriented proposals in a nearly cost-free manner. The second stage is oriented R-CNN head for refining oriented Regions of Interest (oriented RoIs) and recognizing them. Without tricks, oriented R-CNN with ResNet50 achieves state-of-the-art detection accuracy on two commonly-used datasets for oriented object detection including DOTA (75.87% mAP) and HRSC2016 (96.50% mAP), while having a speed of 15.1 FPS with the image size of 1024×1024 on a single RTX 2080Ti. We hope our work could inspire rethinking the design of oriented detectors and serve as a baseline for oriented object detection. Code is available at https://github.com/jbwang1997/OBBDetection.
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001
ICCV3
2021 Object Detection in Optical Remote Sensing Images Based on Positive Sample Reweighting and Feature Decoupling
abstract
The object detection head of Faster R-CNN shared by classification and localization tasks can cause greater harm to training process due to the different features required by the tasks. Besides, Faster R-CNN assigns the same weight to positive proposals, which is unfair to the positive proposal with better regression performance. Aiming at these issues, this paper proposes a novel detection framework, which consists of two important parts, i.e., a Positive Sample Reweighting (PSR) module and a Feature Decoupling (FD) module. Specifically, PSR re-weights each positive proposal according to the positional relationship between each positive proposal and its ground truth. FD employs two parallel branches to extract the weights to decouple features for classification and localization tasks. Experimental results on the DIOR and DOTA datasets show the effectiveness of our proposed method.
Wenqi Yu, Jiabao Wang 0005, Gong Cheng 0003
IGARSS2