Bowen Cheng

dblp:73/8079 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 11 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Simultaneous Suppression of Residual Grating Lobes and Left/Right Ambiguity for Sparse Channel Forward-Looking SAR Imaging
abstract
Multichannel synthetic aperture radar (SAR) has the potential of resolving left/right ambiguity and then achieving high-resolution forward-looking imaging. However, when the sparse channel configuration is adopted, there are always residual grating lobes and left/right ambiguity in the imaging results. In this paper, a scheme of simultaneously suppressing residual grating lobes and left/right ambiguity for sparse channel forward-looking SAR imaging is proposed. Firstly, the formation of the residual grating lobes and left/right ambiguity are analyzed based on the signal model and imaging procedure of multichannel forward-looking SAR. Then, the characteristics of the differences in the position of grating lobes for the imaging results of different snapshots, and the spatiotemporal coupling characteristics of suppression of grating lobes and left/right ambiguity are explained. Next, a space-time steering matrix based on range history information is constructed, which establishes a linear mapping relationship between the space-time echo and the original scene. By solving the linear equation with a regularized iterative adaptive approach, the residual grating lobes and left/right ambiguity are suppressed simultaneously. The extensive simulation and experimental results demonstrate the effectiveness of the proposed method.
Wenchao Li 0002, Rui Chen 0029, Bowen Cheng, Junjie Wu 0001, Jianyu Yang 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Angular Ambiguity Function and Resolution Analysis for Multichannel Radar Forward-Looking Imaging
abstract
With multiple channels in azimuth receiving echoes, multichannel radar has the potential of forward-looking imaging, and various schemes can be formulated. However, due to the different resources utilized by different imaging schemes, the angular resolution will be different. How to analyze the angular resolution and then design appropriate parameters is a key issue in multichannel radar forward-looking imaging. In this paper, based on the echo model of forward-looking imaging, the imaging schemes of synthetic aperture and real aperture are illustrated firstly. Then, based on the ambiguity function theory, the angular ambiguity functions, and the analytical expressions of the angular resolution for different imaging schemes are derived and analyzed. Finally, the simulation results of point targets and extended targets are presented to verify the effectiveness of the theoretical analysis, which would lay a significant foundation for the design of forward-looking imaging schemes of multichannel radar.
Jianyu Yang 0001, Rui Chen 0029, Wenchao Li 0002, Bowen Cheng, Zhongyu Li 0001, Junjie Wu 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud
abstract
The task of 3D semantic scene graph (3D SSG) prediction in the point cloud is challenging since (1) the 3D point cloud only captures geometric structures with limited semantics compared to 2D images, and (2) long-tailed relation distribution inherently hinders the learning of unbiased prediction. Since 2D images provide rich semantics and scene graphs are in nature coped with languages, in this study, we propose Visual-Linguistic Semantics Assisted Training (VL-SAT) scheme that can significantly empower 3DSSG prediction models with discrimination about long-tailed and ambiguous semantic relations. The key idea is to train a powerful multi-modal oracle model to assist the 3D model. This oracle learns reliable structural representations based on semantics from vision, language, and 3D geometry, and its benefits can be heterogeneously passed to the 3D model during the training stage. By effectively utilizing visual-linguistic semantics in training, our VL-SAT can significantly boost common 3DSSG prediction models, such as SGFN and SGGpoint, only with 3D inputs in the inference stage, especially when dealing with tail relation triplets. Comprehensive evaluations and ablation studies on the 3DSSG dataset have validated the effectiveness of the proposed scheme. Code is available at https://github.com/wz7in/CVPR2023-VLSAT.
Ziqin Wang, Bowen Cheng, Lichen Zhao, Dong Xu 0001, Yang Tang 0001, Lu Sheng
CVPR2
2023 Locating Noise is Halfway Denoising for Semi-Supervised Segmentation
abstract
We investigate semi-supervised semantic segmentation with self-training, where a teacher model generates pseudo masks to exploit the benefits of a large amount of unlabeled images. We notice that the noisy label from the generated pseudo masks is the major obstacle to achieving good performance. Previous works all treat the noise in pixel level and ignore the contextual information of the noise. This work shows that locating the patch-wise noisy region is a better way to deal with noise. To be specific, our method, named Uncertainty-aware Patch CutMix (UPC), first estimates the uncertainty of per-pixel prediction for pseudo masks of unlabeled images. Then UPC splits the uncertainty map into patches and calculates patch-wise uncertainty. UPC selects top-k most uncertain patches to generate the uncertain regions. Finally, uncertain regions are replaced with reliable ones from labeled images. We conduct extensive experiments using standard semi-supervised settings on Pascal VOC and Cityscapes. Experiment results show that UPC can significantly boost the performance of the state-of-the-art methods. In addition, we further demonstrate that our UPC is robust to out-of-distribution unlabeled images, e.g., MSCOCO.
Feng Zhu 0005, Bowen Cheng, Luoqi Liu, Yao Zhao 0001, Yunchao Wei
ICCV3
2022 Masked-attention Mask Transformer for Universal Image Segmentation
abstract
Image segmentation groups pixels with different semantics, e.g., category or instance membership. Each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing spe-cialized architectures for each task. We present Masked- attention Mask Transformer (Mask2Former), a new archi-tecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components in-clude masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most no-tably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU onADE20K).
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, Rohit Girdhar
CVPR1
2022 Pointly-Supervised Instance Segmentation
abstract
We propose an embarrassingly simple point annotation scheme to collect weak supervision for instance segmentation. In addition to bounding boxes, we collect binary labels for a set of points uniformly sampled inside each bounding box. We show that the existing instance segmentation models developed for full mask supervision can be seamlessly trained with point-based supervision collected via our scheme. Remarkably, Mask R-CNN trained on COCO, PASCAL VOC, Cityscapes, and LVIS with only 10 annotated random points per object achieves 94%−98% of its fully-supervised performance, setting a strong baseline for weakly-supervised instance segmentation. The new point annotation scheme is approximately 5 times faster than annotating full object masks, making high-quality instance segmentation more accessible in practice. Inspired by the point-based annotation form, we propose a modification to PointRend instance segmentation module. For each object, the new architecture, called Implicit PointRend, generates parameters for a function that makes the final point-level mask prediction. Implicit PointRend is more straightforward and uses a single point-level mask loss. Our experiments show that the new module is more suitable for the point-based supervision.11Project page: https://bowenc0221.github.io/point-sup
Bowen Cheng, Omkar Parkhi, Alexander Kirillov
CVPR1
2022 Bayesian tensor factorization-drive breast cancer subtyping by integrating multi-omics data
Qian Liu 0015, Bowen Cheng, Yongwon Jin, Pingzhao Hu
J. Biomed. Informatics2
2021 Boundary IoU: Improving Object-Centric Image Segmentation Evaluation
abstract
We present Boundary IoU (Intersection-over-Union), a new segmentation evaluation measure focused on boundary quality. We perform an extensive analysis across different error types and object sizes and show that Boundary IoU is significantly more sensitive than the standard Mask IoU measure to boundary errors for large objects and does not over-penalize errors on smaller objects. The new quality measure displays several desirable characteristics like symmetry w.r.t. prediction/ground truth pairs and balanced responsiveness across scales, which makes it more suitable for segmentation evaluation than other boundary-focused measures like Trimap IoU and F-measure. Based on Boundary IoU, we update the standard evaluation protocols for instance and panoptic segmentation tasks by proposing the Boundary AP (Average Precision) and Boundary PQ (Panoptic Quality) metrics, respectively. Our experiments show that the new evaluation metrics track boundary quality improvements that are generally overlooked by current Mask IoU-based evaluation metrics. We hope that the adoption of the new boundary-sensitive evaluation metrics will lead to rapid progress in segmentation methods that improve boundary quality.1
Bowen Cheng, Ross B. Girshick, Piotr Dollár, Alexander C. Berg, Alexander Kirillov
CVPR1
2021 Back-Tracing Representative Points for Voting-Based 3D Object Detection in Point Clouds
abstract
3D object detection in point clouds is a challenging vision task that benefits various applications for understanding the 3D visual world. Lots of recent research focuses on how to exploit end-to-end trainable Hough voting for generating object proposals. However, the current voting strategy can only receive partial votes from the surfaces of potential objects together with severe outlier votes from the cluttered backgrounds, which hampers full utilization of the information from the input point clouds. Inspired by the back-tracing strategy in the conventional Hough voting methods, in this work, we introduce a new 3D object detection method, named as Back-tracing Representative Points Network (BRNet), which generatively back-traces the representative points from the vote centers and also revisits complementary seed points around these generated points, so as to better capture the fine local structural features surrounding the potential objects from the raw point clouds. Therefore, this bottom-up and then top-down strategy in our BRNet enforces mutual consistency between the predicted vote centers and the raw surface points and thus achieves more reliable and flexible object localization and class prediction results. Our BRNet is simple but effective, which significantly outperforms the state-of-the-art methods on two large-scale point cloud datasets, ScanNet V2 (+7.5% in terms of [email protected]) and SUN RGB-D (+4.7% in terms of [email protected]), while it is still lightweight and efficient.
Bowen Cheng, Lu Sheng, Shaoshuai Shi, Dong Xu 0001
CVPR1
2021 An FPGA-based MobileNet Accelerator Considering Network Structure Characteristics
abstract
Convolutional neural networks (CNNs) have been widely deployed in computer vision tasks. However, the computation and resource intensive characteristics of CNN bring obstacles to its application on embedded systems. MobileNet, as a representative of compact models, can reduce the amount of parameters and computation. A high-performance inference accelerator on FPGA for MobileNet is proposed in this paper. With respect to the three types of convolution operations, multiple parallel strategies are exploited and the corresponding hardware structures such as input buffer and configurable adder tree are designed. With respect to the bottleneck block, a dedicated architecture is proposed to reduce data transmission time. In addition, a hardware padding scheme to improve the efficiency of padding is proposed. The accelerator implemented on Virtex-7 FPGA reaches 70.8% Top-1 accuracy under 8-bit quantization. The accelerator achieves 302.3 FPS and 181.8 GOPS, which obtains 22.7x, 3.9x and 1.4x speedup compared to the implementations in Snapdragon 821 CPU, i7-6700HQ CPU and GTX 960M GPU, respectively.
Shun Yan, Zhengyan Liu, Chenglong Zeng, Qiang Liu 0011, Bowen Cheng, Ray C. C. Cheung
FPL6
2021 Per-Pixel Classification is Not All You Need for Semantic Segmentation
abstract
Modern approaches typically formulate semantic segmentation as a per-pixel classification task, while instance-level segmentation is handled with an alternative mask classification. Our key insight: mask classification is sufficiently general to solve both semantic- and instance-level segmentation tasks in a unified manner using the exact same model, loss, and training procedure. Following this observation, we propose MaskFormer, a simple mask classification model which predicts a set of binary masks, each associated with a single global class label prediction. Overall, the proposed mask classification-based method simplifies the landscape of effective approaches to semantic and panoptic segmentation tasks and shows excellent empirical results. In particular, we observe that MaskFormer outperforms per-pixel classification baselines when the number of classes is large. Our mask classification-based method outperforms both current state-of-the-art semantic (55.6 mIoU on ADE20K) and panoptic segmentation (52.7 PQ on COCO) models.
Bowen Cheng, Alexander G. Schwing, Alexander Kirillov
NeurIPS1
2020 Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
abstract
In this work, we introduce Panoptic-DeepLab, a simple, strong, and fast system for panoptic segmentation, aiming to establish a solid baseline for bottom-up methods that can achieve comparable performance of two-stage methods while yielding fast inference speed. In particular, Panoptic-DeepLab adopts the dual-ASPP and dual-decoder structures specific to semantic, and instance segmentation, respectively. The semantic segmentation branch is the same as the typical design of any semantic segmentation model (e.g., DeepLab), while the instance segmentation branch is class-agnostic, involving a simple instance center regression. As a result, our single Panoptic-DeepLab simultaneously ranks first at all three Cityscapes benchmarks, setting the new state-of-art of 84.2% mIoU, 39.0% AP, and 65.5% PQ on test set. Additionally, equipped with MobileNetV3, Panoptic-DeepLab runs nearly in real-time with a single 1025x2049 image (15.8 frames per second), while achieving a competitive performance on Cityscapes (54.1 PQ% on test set). On Mapillary Vistas test set, our ensemble of six models attains 42.7% PQ, outperforming the challenge winner in 2018 by a healthy margin of 1.5%. Finally, our Panoptic-DeepLab also performs on par with several top-down approaches on the challenging COCO dataset. For the first time, we demonstrate a bottom-up approach could deliver state-of-the-art results on panoptic segmentation.
Bowen Cheng, Maxwell D. Collins, Yukun Zhu, Ting Liu 0005, Thomas S. Huang, Hartwig Adam, Liang-Chieh Chen
CVPR1
2020 HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
abstract
Bottom-up human pose estimation methods have difficulties in predicting the correct pose for small persons due to challenges in scale variation. In this paper, we present HigherHRNet: a novel bottom-up human pose estimation method for learning scale-aware representations using high-resolution feature pyramids. Equipped with multi-resolution supervision for training and multi-resolution aggregation for inference, the proposed approach is able to solve the scale variation challenge in bottom-up multi-person pose estimation and localize keypoints more precisely, especially for small person. The feature pyramid in HigherHRNet consists of feature map outputs from HRNet and upsampled higher-resolution outputs through a transposed convolution. HigherHRNet outperforms the previous best bottom-up method by 2.5% AP for medium person on COCO test-dev, showing its effectiveness in handling scale variation. Furthermore, HigherHRNet achieves new state-of-the-art result on COCO test-dev (70.5% AP) without using refinement or other post-processing techniques, surpassing all existing bottom-up methods. HigherHRNet even surpasses all top-down methods on CrowdPose test (67.6% AP), suggesting its robustness in crowded scene.
Bowen Cheng, Bin Xiao 0004, Jingdong Wang 0001, Humphrey Shi, Thomas S. Huang, Lei Zhang 0001
CVPR1
2020 Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
Liang-Chieh Chen, Raphael Gontijo Lopes, Bowen Cheng, Maxwell D. Collins, Ekin Dogus Cubuk, Barret Zoph, Hartwig Adam, Jonathon Shlens
ECCV (9)3
2020 A Target Detection Algorithm of Neural Network Based on Histogram Statistics
abstract
Aiming at the problems of poor adaptability of traditional target detection algorithms and high computational resources of deep learning algorithms, a BP neural network target detection algorithm based on histogram statistics is proposed. It is based on the principle that similar areas have similar histograms. In this algorithm, the two-dimensional image information converts to the one-dimensional histogram information. We establish a three-layer neural network model, and the histogram is used as the input of the BP neural network. Compared to the traditional target detection algorithms, its complexity is low, and its efficiency and accuracy is high. The experimental results show that the fewer classification categories, the higher target detection probability. The computational complexity of the BP neural network is low, so the computational efficiency is quite high. The accuracy of target recognition is higher than 97% with SAR and optical images.
Yalong Pang, Luyuan Wang, Jiyang Yu, Bowen Cheng, Zongling Li
IGARSS5
2019 High Frequency Residual Learning for Multi-Scale Image Classification
Bowen Cheng, Rong Xiao 0003, Thomas S. Huang, Lei Zhang 0001
BMVC1
2019 SPGNet: Semantic Prediction Guidance for Scene Parsing
abstract
Multi-scale context module and single-stage encoder-decoder structure are commonly employed for semantic segmentation. The multi-scale context module refers to the operations to aggregate feature responses from a large spatial extent, while the single-stage encoder-decoder structure encodes the high-level semantic information in the encoder path and recovers the boundary information in the decoder path. In contrast, multi-stage encoder-decoder networks have been widely used in human pose estimation and show superior performance than their single-stage counterpart. However, few efforts have been attempted to bring this effective design to semantic segmentation. In this work, we propose a Semantic Prediction Guidance (SPG) module which learns to re-weight the local features through the guidance from pixel-wise semantic prediction. We find that by carefully re-weighting features across stages, a two-stage encoder-decoder network coupled with our proposed SPG module can significantly outperform its one-stage counterpart with similar parameters and computations. Finally, we report experimental results on the semantic segmentation benchmark Cityscapes, in which our SPGNet attains 81.1% on the test set using only 'fine' annotations.
Bowen Cheng, Liang-Chieh Chen, Yunchao Wei, Yukun Zhu, Jinjun Xiong, Thomas S. Huang, Wen-Mei W. Hwu, Humphrey Shi
ICCV1
2019 Remote Sensing Ship Target Detection and Recognition System Based on Machine Learning
abstract
In this paper, the ship target detection and recognition system of remote sensing imaging based on machine learning is designed according to the sparsity of interest targets in optical remote sensing image, and proposes a method of target detection and recognition based on morphological matching and machine learning. The slices of suspected targets are extracted quickly by visual enhancement technology that the amount of data processed is greatly reduced. The target information of interest is extracted in depth and the false alarm rate of detection is greatly reduced by using machine learning method to classify objects. In the system function and performance verification test, the real-time and accuracy index through 227 targets of 32 scenes GF-2 satellite images are tested what can detect and recognize about 20 objects per second. The recall rate of the system is more than 92%, and the efficiency of target detection method based on traditional morphological matching is less than 60% while the target recognition method base on machine learning improves the precision rate to over 97%.
Zongling Li, Lu-yuan Wang, Ji-Yang Yu, Bowen Cheng, Jianfeng Yin
IGARSS4
2019 Enhance Visual Recognition Under Adverse Conditions via Deep Networks
abstract
Visual recognition under adverse conditions is a very important and challenging problem of high practical value, due to the ubiquitous existence of quality distortions during image acquisition, transmission, or storage. While deep neural networks have been extensively exploited in the techniques of low-quality image restoration and high-quality image recognition tasks respectively, few studies have been done on the important problem of recognition from very low-quality images. This paper proposes a deep learning based framework for improving the performance of image and video recognition models under adverse conditions, using robust adverse pre-training or its aggressive variant. The robust adverse pre-training algorithms leverage the power of pre-training and generalizes conventional unsupervised pre-training and data augmentation methods. We further develop a transfer learning approach to cope with real-world datasets of unknown adverse conditions. The proposed framework is comprehensively evaluated on a number of image and video recognition benchmarks, and obtains significant performance improvements under various single or mixed adverse conditions. Our visualization and analysis further add to the explainability of results.
Ding Liu 0001, Bowen Cheng, Zhangyang Wang, Haichao Zhang 0001, Thomas S. Huang
IEEE Trans. Image Process.2
2018 Visual Recognition in Very Low-Quality Settings: Delving Into the Power of Pre-Training
abstract
Visual recognition from very low-quality images is an extremely challenging task with great practical values. While deep networks have been extensively applied to low-quality image restoration and high-quality image recognition tasks respectively, few works have been done on the important problem of recognition from very low-quality images.This paper presents a degradation-robust pre-training approach on improving deep learning models towards this direction. Extensive experiments on different datasets validate the effectiveness of our proposed method.
Bowen Cheng, Ding Liu 0001, Zhangyang Wang, Haichao Zhang 0001, Thomas S. Huang
AAAI1
2018 Revisiting RCNN: On Awakening the Classification Power of Faster RCNN
Bowen Cheng, Yunchao Wei, Humphrey Shi, Rogério Feris, Jinjun Xiong, Thomas S. Huang
ECCV (15)1
2018 TS ^2 2 C: Tight Box Mining with Surrounding Segmentation Context for Weakly Supervised Object Detection
Yunchao Wei, Bowen Cheng, Humphrey Shi, Jinjun Xiong, Jiashi Feng, Thomas S. Huang
ECCV (11)3
2017 Robust emotion recognition from low quality and low bit rate video: A deep learning approach
abstract
Emotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the transmitted video, which will heavily degrade the recognition reliability. We develop a novel framework to achieve robust emotion recognition from low bit rate video. While video frames are downsampled at the encoder side, the decoder is embedded with a deep network model for joint super-resolution (SR) and recognition. Notably, we propose a novel max-mix training strategy, leading to a single “One-for-All” model that is remarkably robust to a vast range of downsampling factors. That makes our framework well adapted for the varied bandwidths in real transmission scenarios, without hampering scalability or efficiency. The proposed framework is evaluated on the AVEC 2016 benchmark, and demonstrates significantly improved stand-alone recognition performance, as well as rate-distortion (R-D) performance, than either directly recognizing from LR frames, or separating SR and recognition.
Bowen Cheng, Zhangyang Wang, Zhaobin Zhang, Zhu Li 0001, Ding Liu 0001, Jianchao Yang, Shuai Huang 0001, Thomas S. Huang
ACII1