Jinming Cao

dblp:220/4086 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-8614-7366ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Improving Test-Time Efficiency in Source-Free Semantic Segmentation via Multi-Stage Self-Training
abstract
Source-free domain adaptive semantic segmentation aims at adapting a model trained on the source domain to the target domain without requiring access to the source data. Self-training has emerged as a leading approach to address this challenging problem. However, without robust denoising mechanisms to reduce the noise in pseudo labels, it still easily fall into biased estimates. Most existing methods address this issue by introducing novel architectures, but often at the cost of increased model complexity or reliance on additional input modalities. Different from previous studies, this article introduces UniSFDA , a unified multi-stage self-training framework that integrates cross-model transfer learning, uncertainty-aware pseudo label fusion, and intra-domain style augmentation, thereby enhancing both segmentation accuracy and test-time efficiency. Our proposed framework offers exceptional flexibility, with each component being independent and ready to be integrated into any existing self-training framework. Additionally, we investigate the performance of various representative segmentation models, including DeepLabv2, SegFormer, DFormer, and ViT-Adapter, within our framework. It is worth noting that UniSFDA is model-agnostic, allowing both source and target networks to be instantiated with arbitrary segmentation architectures, and thus readily benefiting from future advances in segmentation models. Experiments on the GTA5 \(\rightarrow\) Cityscapes and SYNTHIA \(\rightarrow\) Cityscapes benchmarks demonstrate the effectiveness of our framework. With DeepLabv2 (SegFormer) as the source model, UniSFDA establishes new state-of-the-art performance, achieving mIoU scores of 61.8% (65.4%) and 57.9% (59.6%) on the two benchmarks, respectively.
Yifang Yin, Jinming Cao, Zhenguang Liu, Guanfeng Wang, Shili Xiang, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Class Incremental Learning via Feature Space Calibration
Jinming Cao, Jihie Kim, Roger Zimmermann
Comput. Vis. Media2
2025 BDA: Bi-Directional Attention for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is an important and rapidly growing area of machine learning that aims to recognize new classes without prior training data. Despite its significance, ZSL has faced challenges with overfitting in embedding-based methods and limitations in traditional one-directional attention (ODA) based approaches. To bridge these gaps, this paper proposes the use of bi-directional attention (BDA) to integrate insights from both embedding and attention-based approaches. The proposed BDA system consists of a bi-directional attention network (BDAN) and a synthesized visual embedding network (SVEN) that facilitates visual-semantic interaction for ZSL classification. More specifically, the BDAN employs region self-attention (RSA), semantic synthesis attention (SSA), and visual synthesis attention (VSA) to overcome the overfitting issue in embedding methods and enhance transferability, to associate visual features with semantic property information, and to learn locally improved visual features. Extensive testing on CUB, SUN, and AWA2 datasets confirm the superiority of our proposed method over traditional approaches. Code is available at https://github.com/JunseokLee3/BDA.
Jinming Cao, Yifang Yin, Jihie Kim, Roger Zimmermann
Comput. Vis. Media2
2025 Improving Back-Projection Accuracy for the Semantic Segmentation of Indoor Point Clouds With Fewer & Sparse Image Annotations
abstract
Performing semantic segmentation on point clouds is the primary method by which machines perceive 3D scenes in a fine-grained manner. Deep learning algorithms usually require many pointwise annotations obtained with specialized tools, which is a laborious and inefficient process. To this end, we develop two frameworks for training point cloud semantic segmentation networks, one that utilizes fewer projected image annotations and another that employs sparse scribble image annotations, making the process more flexible and user friendly. However, back-projecting 2D-pixel labels to 3D points during loss calculations always introduces errors. To increase the back-projection accuracy of our approach, we first identify and record potential pixel-point correspondence errors and then develop strategies for constructing an accurate back-projection mapping matrix. Specifically, we filter out occluded and noisy points to avoid incorrect label allocations and permit multiclass assignments to adjust the ambiguity of boundary points. By incorporating an accurate back-projection mechanism into the loss functions of the proposed training frameworks, our networks can perform well with only four projected image annotations or even sparse scribble image annotations for each scene. This results in state-of-the-art performance compared with that of other weakly supervised point cloud semantic segmentation approaches, and the outcomes are even comparable to those produced by fully supervised methods on the S3DIS and ScanNet-v2 datasets.
Peng Jiang 0002, Zhiyi Pan 0001, Jinming Cao, Roger Zimmermann, Changhe Tu
IEEE Trans. Multim.4
2025 ShapeMoiré: Channel-Wise Shape-Guided Network for Image Demoiréing
abstract
Photographing optoelectronic displays often introduces unwanted moiré patterns due to analog signal interference between the pixel grids of the display and the camera sensor arrays. This work identifies two problems that are largely ignored by existing image demoiréing approaches: (1) moiré patterns vary across different channels (RGB); (2) repetitive patterns are constantly observed. However, employing conventional convolutional (CNN) layers cannot address these problems. Instead, this article presents the use of our recently proposed Shape concept. It was originally employed to model consistent features from fragmented regions, particularly when identical or similar objects coexist in an RGB-D image. Interestingly, we find that the Shape information effectively captures the moiré patterns in artifact images. Motivated by this discovery, we propose a new method, ShapeMoiré, for image demoiréing. Beyond modeling shape features at the patch level, we further extend this to the global image level and design a novel Shape-Architecture. Consequently, our proposed method, equipped with both ShapeConv and Shape-Architecture, can be seamlessly integrated into existing approaches without introducing any additional parameters or computation overhead during inference. We conduct extensive experiments on four widely used datasets, and the results demonstrate that our ShapeMoiré achieves state-of-the-art performance, particularly in terms of the PSNR metric. We then apply our method across four popular architectures to showcase its generalization capabilities. Moreover, to further validate its generality beyond the demoiréing task, we apply ShapeMoiré to the image deblurring task, where it continues to deliver consistent performance gains. Finally, experiments on real-world images captured by smartphones confirm the robustness and practical applicability of ShapeMoiré in challenging demoiréing scenarios. We open sourced an implementation of ShapeMoiré in PyTorch at https://github.com/SichengS/ShapeMoire .
Jinming Cao, Sicheng Shen, Qiu Zhou, Yifang Yin, Yangyan Li, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.1
2024 SOGDet: Semantic-Occupancy Guided Multi-View 3D Object Detection
abstract
In the field of autonomous driving, accurate and comprehensive perception of the 3D environment is crucial. Bird's Eye View (BEV) based methods have emerged as a promising solution for 3D object detection using multi-view images as input. However, existing 3D object detection methods often ignore the physical context in the environment, such as sidewalk and vegetation, resulting in sub-optimal performance. In this paper, we propose a novel approach called SOGDet (Semantic-Occupancy Guided Multi-view 3D Object Detection), that leverages a 3D semantic-occupancy branch to improve the accuracy of 3D object detection. In particular, the physical context modeled by semantic occupancy helps the detector to perceive the scenes in a more holistic view. Our SOGDet is flexible to use and can be seamlessly integrated with most existing BEV-based methods. To evaluate its effectiveness, we apply this approach to several state-of-the-art baselines and conduct extensive experiments on the exclusive nuScenes dataset. Our results show that SOGDet consistently enhance the performance of three baseline methods in terms of nuScenes Detection Score (NDS) and mean Average Precision (mAP). This indicates that the combination of 3D object detection and 3D semantic occupancy leads to a more comprehensive perception of the 3D environment, thereby aiding build more robust autonomous driving systems. The codes are available at: https://github.com/zhouqiu/SOGDet.
Qiu Zhou, Jinming Cao, Hanchao Leng, Yifang Yin, Yu Kun, Roger Zimmermann
AAAI2
2022 DO-Conv: Depthwise Over-Parameterized Convolutional Layer
abstract
Convolutional layers are the core building blocks of Convolutional Neural Networks (CNNs). In this paper, we propose to augment a convolutional layer with an additional depthwise convolution, where each input channel is convolved with a different 2D kernel. The composition of the two convolutions constitutes an over-parameterization, since it adds learnable parameters, while the resulting linear operation can be expressed by a single convolution layer. We refer to this depthwise over-parameterized convolutional layer as DO-Conv, which is a novel way of over-parameterization. We show with extensive experiments that the mere replacement of conventional convolutional layers with DO-Conv layers boosts the performance of CNNs on many classical vision tasks, such as image classification, detection, and segmentation. Moreover, in the inference phase, the depthwise convolution is folded into the conventional convolution, reducing the computation to be exactly equivalent to that of a convolutional layer without over-parameterization. As DO-Conv introduces performance gains without incurring any computational complexity increase for inference, we advocate it as an alternative to the conventional convolutional layer. We open sourced an implementation of DO-Conv in Tensorflow, PyTorch and GluonCV at https://github.com/yangyanli/DO-Conv.
Jinming Cao, Yangyan Li, Mingchao Sun, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen, Changhe Tu
IEEE Trans. Image Process.1
2021 ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic Segmentation
abstract
RGB-D semantic segmentation has attracted increasing attention over the past few years. Existing methods mostly employ homogeneous convolution operators to consume the RGB and depth features, ignoring their intrinsic differences. In fact, the RGB values capture the photometric appearance properties in the projected image space, while the depth feature encodes both the shape of a local geometry as well as the base (whereabout) of it in a larger context. Compared with the base, the shape probably is more inherent and has a stronger connection to the semantics, and thus is more critical for segmentation accuracy. Inspired by this observation, we introduce a Shape-aware Convolutional layer (ShapeConv) for processing the depth feature, where the depth feature is firstly decomposed into a shape-component and a base-component, next two learnable weights are introduced to cooperate with them independently, and finally a convolution is applied on the re-weighted combination of these two components. ShapeConv is model-agnostic and can be easily integrated into most CNNs to replace vanilla convolutional layers for semantic segmentation. Extensive experiments on three challenging indoor RGB-D semantic segmentation benchmarks, i.e., NYU-Dv2(-13,-40), SUN RGB-D, and SID, demonstrate the effectiveness of our ShapeConv when employing it over five popular architectures. Moreover, the performance of CNNs with ShapeConv is boosted without introducing any computation and memory increase in the inference phase. The reason is that the learnt weights for balancing the importance between the shape and base components in ShapeConv become constants in the inference phase, and thus can be fused into the following convolution, resulting in a network that is identical to one with vanilla convolutional layers.
Jinming Cao, Hanchao Leng, Dani Lischinski, Daniel Cohen-Or, Changhe Tu, Yangyan Li
ICCV1
2021 APM: Adaptive permutation module for point cloud classification
Yechao Wang, Jinming Cao, Yangyan Li, Changhe Tu
Comput. Graph.2
2021 RGB×D: Learning depth-weighted RGB patches for RGB-D indoor semantic segmentation
Jinming Cao, Hanchao Leng, Daniel Cohen-Or, Dani Lischinski, Changhe Tu, Yangyan Li
Neurocomputing1