Chengliang Zhong

dblp:42/9391 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 IAMTrack: interframe appearance and modality tokens propagation with temporal modeling for RGBT tracking
Huiwei Shi, Xiaodong Mu, Hao He 0005, Chengliang Zhong
Appl. Intell.4
2025 Equivariant Local Reference Frames With Optimization for Robust Non-Rigid Point Cloud Correspondence
abstract
Unsupervised non-rigid point cloud shape correspondence underpins a multitude of 3D vision tasks, yet itself is non-trivial given the exponential complexity stemming from inter-point degree-of-freedom, i.e., pose transformations. Based on the assumption of local rigidity, one solution for reducing complexity is to decompose the overall shape into independent local regions using Local Reference Frames (LRFs) that are equivariant to SE(3) transformations. However, the focus solely on local structure neglects global geometric contexts, resulting in less distinctive LRFs that lack crucial semantic information necessary for effective matching. Furthermore, such complexity introduces out-of-distribution geometric contexts during inference, thus complicating generalization. To this end, we introduce 1) EquiShape, a novel structure tailored to learn pair-wise LRFs with global structural cues for both spatial and semantic consistency, and 2) LRF-Refine, an optimization strategy generally applicable to LRF-based methods, aimed at addressing the generalization challenges. Specifically, for EquiShape, we employ cross-talk within separate equivariant graph neural networks (Cross-GVP) to build long-range dependencies to compensate for the lack of semantic information in local structure modeling, deducing pair-wise independent SE(3)-equivariant LRF vectors for each point. For LRF-Refine, the optimization adjusts LRFs within specific contexts and knowledge, enhancing the geometric and semantic generalizability of point features. Our overall framework surpasses the state-of-the-art methods by a large margin on three benchmarks. Codes are available at https://github.com/2019EPWL/EquiShape.
Runfa Chen, Fuchun Sun 0001, Kai Sun 0014, Chengliang Zhong, Guangyuan Fu, Yikai Wang 0001
IEEE Trans. Image Process.6
2024 MonoOcc: Digging into Monocular Semantic Occupancy Prediction
abstract
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, particularly due to its potential to enhance the 3D perception of autonomous vehicles. However, existing methods rely on a complex cascaded framework with relatively limited information to restore 3D scenes, including a dependency on supervision solely on the whole network’s output, single-frame input, and the utilization of a small backbone. These challenges, in turn, hinder the optimization of the framework and yield inferior prediction results, particularly concerning smaller and long-tailed objects. To address these issues, we propose MonoOcc. In particular, we (i) improve the monocular occupancy prediction framework by proposing an auxiliary semantic loss as supervision to the shallow layers of the framework and an image-conditioned cross-attention module to refine voxel features with visual clues, and (ii) employ a distillation module that transfers temporal information and richer knowledge from a larger image backbone to the monocular semantic occupancy prediction framework with low cost of hardware. With these advantages, our method yields state-of-the-art performance on the camera-based SemanticKITTI Scene Completion benchmark. Codes and models can be accessed at https://github.com/ucaszyp/MonoOcc.
Yupeng Zheng, Xiang Li 0205, Pengfei Li 0007, Yuhang Zheng 0004, Bu Jin, Chengliang Zhong, Xiaoxiao Long, Hao Zhao 0002
ICRA6
2023 3D Implicit Transporter for Temporally Consistent Keypoint Discovery
abstract
Keypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter.
Chengliang Zhong, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Li Yi 0001, Xiaodong Mu, Ling Wang 0001, Pengfei Li 0007, Guyue Zhou, Chao Yang 0026, Jian Zhao 0006
ICCV1
2023 STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth Estimation
abstract
Self-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capa-bilities of self-driving vehicles. However, it intrinsically relies upon the photometric consistency assumption, which hardly holds during nighttime. Although various supervised night-time image enhancement methods have been proposed, their generalization performance in challenging driving scenarios is not satisfactory. To this end, we propose the first method that jointly learns a nighttime image enhancer and a depth estimator, without using ground truth for either task. Our method tightly entangles two self-supervised tasks using a newly proposed uncertain pixel masking strategy. This strategy originates from the observation that nighttime images not only suffer from underexposed regions but also from overexposed regions. By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally. We benchmark the method on two established datasets: nuScenes and RobotCar and demonstrate state-of-the-art performance on both of them. Detailed ablations also reveal the mechanism of our proposal. Last but not least, to mitigate the problem of sparse ground truth of existing datasets, we provide a new photo-realistically enhanced nighttime dataset based upon CARLA. It brings meaningful new challenges to the community. Codes, data, and models are available at https://github.com/ucaszyp/STEPS.
Yupeng Zheng, Chengliang Zhong, Pengfei Li 0007, Huan-ang Gao, Yuhang Zheng 0004, Bu Jin, Ling Wang 0001, Hao Zhao 0002, Guyue Zhou, Dongbin Zhao
ICRA2
2022 Sim2Real Object-Centric Keypoint Detection and Description
abstract
Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, requires further identifying which object each interest point belongs to. With such fine-grained information, our framework enables more downstream potentials, such as object-level matching and pose estimation in a clustered environment. To get around the difficulty of label collection in the real world, we develop a sim2real contrastive learning mechanism that can generalize the model trained in simulation to real-world applications. The novelties of our training method are three-fold: (i) we integrate the uncertainty into the learning framework to improve feature description of hard cases, e.g., less-textured or symmetric patches; (ii) we decouple the object descriptor into two independent branches, intra-object salience and inter-object distinctness, resulting in a better pixel-wise description; (iii) we enforce cross-view semantic consistency for enhanced robustness in representation learning. Comprehensive experiments on image matching and 6D pose estimation verify the encouraging generalization ability of our method. Particularly for 6D pose estimation, our method significantly outperforms typical unsupervised/sim2real methods, achieving a closer gap with the fully supervised counterpart.
Chengliang Zhong, Chao Yang 0026, Fuchun Sun 0001, Jinshan Qi, Xiaodong Mu, Huaping Liu 0001, Wenbing Huang 0001
AAAI1
2022 SNAKE: Shape-aware Neural 3D Keypoint Field
abstract
Detecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoints that are invariant to transformation. Nevertheless, the idea of incorporating shape reconstruction into 3D keypoint detection is under-explored. We argue that this is restricted by former problem formulations. To this end, a novel unsupervised paradigm named SNAKE is proposed, which is short for shape-aware neural 3D keypoint field. Similar to recent coordinate-based radiance or distance field, our network takes 3D coordinates as inputs and predicts implicit shape indicators and keypoint saliency simultaneously, thus naturally entangling 3D keypoint detection and shape reconstruction. We achieve superior performance on various public benchmarks, including standalone object datasets ModelNet40, KeypointNet, SMPL meshes and scene-level datasets 3DMatch and Redwood. Intrinsic shape awareness brings several advantages as follows. (1) SNAKE generates 3D keypoints consistent with human semantic annotation, even without such supervision. (2) SNAKE outperforms counterparts in terms of repeatability, especially when the input point clouds are down-sampled. (3) the generated keypoints allow accurate geometric registration, notably in a zero-shot setting. Codes and models are available at https://github.com/zhongcl-thu/SNAKE.
Chengliang Zhong, Peixing You, Xiaoxue Chen, Hao Zhao 0002, Fuchun Sun 0001, Guyue Zhou, Xiaodong Mu, Chuang Gan 0001, Wenbing Huang 0001
NeurIPS1
2022 Hierarchical attention and feature projection for click-through rate prediction
Chengliang Zhong, Shouxiang Fan, Xiaodong Mu, Zhen Ni
Appl. Intell.2
2022 Combining feature importance and neighbor node interactions for cold start recommendation
Chenhui Ma, Chengliang Zhong, Xiaodong Mu
Eng. Appl. Artif. Intell.3
2022 Multi-scale and multi-channel neural network for click-through rate prediction
Chenhui Ma, Chengliang Zhong, Xiaodong Mu
Neurocomputing3
2021 MBPI: Mixed behaviors and preference interaction for session-based recommendation
Chenhui Ma, Chengliang Zhong, Xiaodong Mu, Lizhi Wang 0009
Appl. Intell.3
2021 Recurrent convolutional neural network for session-based recommendation
Chenhui Ma, Xiaodong Mu, Chengliang Zhong, A. Ruhan
Neurocomputing5
2019 SAR Target Image Classification Based on Transfer Learning and Model Compression
abstract
When convolutional neural networks (CNNs) are applied to the synthetic aperture radar (SAR) image classification, they are prone to overfitting due to scarce SAR image data, and CNNs require a large amount of storage and long computing time, so it is difficult to deploy them on resource constrained devices. This letter proposes a simple and feasible approach that can effectively solve these problems. First, the convolutional layers of the pretrained model on the ImageNet data set are transferred, and a new convolutional layer and global pooling layer are added afterward. Then, fine-tuning is performed on the new network from the SAR image data set. Finally, a filterbased pruning method is used on the convolutional layers to obtain a compact network. Compared with the all-convolutional network (A-ConvNets) which is the state-of-the-art method on the moving and stationary target acquisition and recognition data set, our method achieves about 3.6× speedup during forward propagation and 3.7× compression of the parameters, with only a 1.42% decrease in the accuracy.
Chengliang Zhong, Xiaodong Mu, Xiangchen He
IEEE Geosci. Remote. Sens. Lett.1
2018 Classification for SAR Scene Matching Areas Based on Convolutional Neural Networks
abstract
The selection of scene matching areas is a difficult problem in the field of matching guidance. Compared with the traditional methods of matching feature extraction and pattern classification, this letter applies convolutional neural networks (CNN) to the extraction of synthetic aperture radar (SAR) scene matching regions for the first time. First of all, we match the SAR images of the same land taken by satellites from different angles and in different phases, and then automatically label the matching suitability of the images as the output of the network according to the matching results. Next, the digital elevation model data reflecting the elevation information and the SAR image grayscale information are fused as the input to the network. Finally, CNN is used to automatically extract the matching features and classify the suitability of the SAR images. The proposed method avoids the steps of extracting features manually and improves the classification performance of SAR scene matching area. Compared with the support vector machine method, the classification accuracy increases from 86.1% to 93.3%.
Chengliang Zhong, Xiaodong Mu, Xiangchen He, Bichao Zhan
IEEE Geosci. Remote. Sens. Lett.1