VLDB 2026 Research / reviewers in the wild / expert
Zhongyao Cheng
dblp:247/1787
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2025
0009-0006-2363-450XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Global-Aware Monocular Semantic Scene Completion with State Space Models
Shijie Li 0006, Zhongyao Cheng, Juergen Gall, Xun Xu 0002, Xulei Yang |
ICCV | 2 |
| 2025 | AIC3DOD: Advancing Indoor Class-Incremental 3D Object Detection with Point Transformer Architecture and Room Layout ConstraintsabstractOver the recent years, there has been a growing interest in class-incremental 3D object detection based on point clouds. However, the current state-of-the-art (SOTA) methods still fall short of practical adoption, mainly due to two key observations. Firstly, existing SOTA methods are limited by the capability of feature representation from the object detection model. Secondly, these methods overlook the importance of incorporating prior information or geometry constraints, which are crucial elements for 3D point cloud tasks. In this study, we strive to enhance the performance of class-incremental 3D object detection for indoor scenes by proposing AIC3DOD - Advancing Indoor Classincremental 3D Object Detection using the point transformer architecture with room layout constraints. Our approach employs a transformer architecture in our detection model and optimizes the class incremental step in the transformer architecture. Besides, AIC3DOD incorporates additional prior information, namely room layout, to impose physical constraints on detected objects, thereby enhancing overall object detection performance. Extensive experimental results on the ScanNet dataset demonstrate the effectiveness of our approach, showcasing our superior performance compared to other SOTA methods in the class-incremental 3D object detection task. Zhongyao Cheng, Fang Wu 0009, Peisheng Qian, Ziyuan Zhao, Xulei Yang |
WACV | 1 |
| 2024 | Hybrid Explainable Network Intrusion Detection Framework Based on Shapley Additive ExplanationsabstractWith the rapid advancements in network technology and automation processes, the threats posed by cyberattacks have become increasingly significant. To address these threats, numerous researchers have developed various network intrusion detection systems (NIDS) to monitor network traffic. However, with the continuous complexification of 5G networks and the exponential increase in network traffic, the emergence of new attacks alongside the lack of interpretability in NIDS posed challenges to the performance and efficiency of network intrusion detection. To tackle these issues, this paper proposes a hybrid explainable network intrusion detection framework that combines the strengths of supervised and unsupervised learning, enabling effective detection of emerging attacks within the network. Specifically, we utilize a Light Gradient Boosting Machine (LightGBM) model for supervised learning, followed by the SHapley Additive exPlanations (SHAP) method for model explanation and feature selection. Additionally, we employ a network utilizing Convolutional Neural Network and Long short term memory for encoding and decoding (ACLNet) purposes in unsupervised learning. Finally, the results from these two learning processes are integrated for anomaly detection. The simulation experimental results on the NSL-KDD dataset verify that our approach generates detection performance on par with state-of-the-art methods while offering significantly enhanced interpretability and improving detection efficiency. Sijin Chen, Jing Liu 0032, Cen Chen 0002, Songyu Xie, Zhongyao Cheng |
ISPA | 5 |
| 2024 | On-the-fly Point Feature Representation for Point Clouds AnalysisabstractPoint cloud analysis is challenging due to its unique characteristics of unorderness, sparsity and irregularity. Prior works attempt to capture local relationships by convolution operations or attention mechanisms, exploiting geometric information from coordinates implicitly. These methods, however, are insufficient to describe the explicit local geometry, e.g., curvature and orientation. In this paper, we propose On-the-fly Point Feature Representation (OPFR), which captures abundant geometric information explicitly through Curve Feature Generator module. This is inspired by Point Feature Histogram (PFH) from computer vision community. However, the utilization of vanilla PFH encounters great difficulties when applied to large datasets and dense point clouds, as it demands considerable time for feature generation. In contrast, we introduce the Local Reference Constructor module, which approximates the local coordinate systems based on triangle sets. Owing to this, our OPFR only requires extra 1.56ms for inference (65X faster than vanilla PFH) and 0.012M more parameters, and it can serve as a versatile plug-and-play module for various backbones, particularly MLP-based and Transformer-based backbones examined in this study. Additionally, we introduce the novel Hierarchical Sampling module aimed at enhancing the quality of triangle sets, thereby ensuring robustness of the obtained geometric features. Our proposed method improves overall accuracy (OA) on ModelNet40 from 90.7% to 94.5% (+3.8%) for classification, and OA on S3DIS Area-5 from 86.4% to 90.0% (+3.6%) for semantic segmentation, respectively, building upon PointNet++ backbone. When integrated with Point Transformer backbone, we achieve state-of-the-art results on both tasks: 94.8% OA on ModelNet40 and 91.7% OA on S3DIS Area-5. Jiangyi Wang, Zhongyao Cheng, Na Zhao 0004, Jun Cheng 0003, Xulei Yang |
ACM Multimedia | 2 |
| 2024 | HEN: a novel hybrid explainable neural network based framework for robust network intrusion detection
Wei Wei 0006, Sijin Chen, Cen Chen 0002, Heshi Wang, Jing Liu 0032, Zhongyao Cheng, Xiaofeng Zou |
Sci. China Inf. Sci. | 6 |
| 2023 | COCO-TEACH: A Contrastive Co-Teaching Network For Incremental 3D Object DetectionabstractDeep learning (DL) models for 3D object detection from point clouds have shown remarkable progress in various autonomous perception scenarios. However, the issue of catastrophic forgetting seriously hinders the deployment of these models in real-world applications where new classes are encountered over time. In order to address this issue, we present the Contrastive Co-Teaching Network (COCO-TEACH) framework for class-incremental 3D object detection. Our proposed framework consists of two teacher networks: a primary teacher network that detects old class objects in new data and provides them with pseudo-labels and an auxiliary teacher network that leverages the unlabelled objects in new data. The two teacher models transfer their learned knowledge to the target student model through a class-aware consistency loss. To enhance this transfer, a supervised contrastive loss is further incorporated into the loss function. We evaluate the performance of our proposed method against baseline methods through extensive experiments on two benchmark datasets. The results show that our proposed framework achieves state-of-the-art performance on incremental 3D object detection. Zhongyao Cheng, Cen Chen 0002, Ziyuan Zhao, Peisheng Qian, Xiaoli Li 0001, Xulei Yang |
ICIP | 1 |
| 2023 | Controlling Facial Attribute Synthesis by Disentangling Attribute Feature Axes in Latent SpaceabstractIn this study, we propose a novel approach to synthesize high-resolution and hyper-realistic face images with controlled attributes. Firstly, by training an attribute classifier to assign attribute labels to given synthesized face images, we build the links between latent vectors and face attributes. Secondly, we adapt the regression method to match the distributions of latent vectors with the corresponding face attributes, to control the attribute synthesis in the face images. Finally, we use the Gram-Schmidt orthogonalization algorithm to disentangle the attribute feature axes in latent space, such that a change in one attribute will not cause any changes in other attributes. Extensive experiments demonstrate the effectiveness of the proposed approach for high-quality face image synthesis with controlled attributes. Qiyu Wei, Zhongyao Cheng, Zeng Zeng, Xulei Yang |
ICIP | 4 |
| 2023 | Spatial and Temporal Dual-Scale Adaptive Pruning for Point Cloud VideosabstractPoint clouds, characterized by irregularity and disorder, are widely utilized in the domain of the Internet of Things, including applications such as terrain exploration and autonomous driving. 3D action recognition and semantic segmentation widely employ them. However, point cloud videos often exhibit massive data volume and contain substantial data redundancy. This situation is highly detrimental to the real-time applications of point clouds such as autonomous driving and virtual reality. Network pruning is imperative to mitigate the redundancy. This paper proposes a spatial and temporal dual-scale adaptive pruning (STDAP) method for point cloud videos based on the attention mechanism to reduce redundancy. This method operates in both temporal and spatial dimensions, enabling adaptive pruning of point cloud videos and cutting down the inference time and memory usage. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques regarding accuracy and inference time. It achieves notable acceleration while maintaining efficacy. Songyu Xie, Jing Liu 0032, Cen Chen 0002, Zhongyao Cheng |
ICPADS | 4 |
| 2023 | Cascade Graph Neural Networks for Few-Shot Learning on Point CloudsabstractPoint cloud data, a flexible 3D object representation, is critical for various applications such as autonomous driving, robotics and remote sensing. Despite the recent success of deep neural networks (DNNs) on supervised point cloud analysis tasks, they still rely on tedious manual annotation of point clouds and cannot make predictions for new classes. Unlike few-shot learning for 2D images with the advantages of large-scale datasets and high-quality deep pre-trained models like ResNet, for 3D few-shot learning, obtaining discriminative representations of unseen classes with high intra-class similarity and inter-class difference is very challenging. To address this issue, this work proposes a novel cascade graph neural network for few-shot learning on point clouds, termed as CGNN, in which two cascade GNNs are adopted to extract the intra-object topological information and learn the inter-object relations respectively. To further increase the discriminability of point cloud features, we first design a novel discriminative edge label to model the intra-class similarity and inter-class dissimilarity based on channel-wise feature variance and class consistency. Second, we propose a novel few-shot circle loss which classifies the nodes into two subsets, i.e., support to support pairs and support to query pairs, and optimizes the pair-wise similarity on two subsets independently. Extensive experiments on benchmark CAD and real LiDAR point cloud datasets have demonstrated that CGNN improves accuracy by 5.98% over the state-of-the-art GNN-based few-shot classification methods. Yangfan Li 0001, Cen Chen 0002, Weiquan Yan, Zhongyao Cheng, Hui Li Tan, Wenjie Zhang 0004 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Latent Vector Prototypes Guided Conditional Face SynthesisabstractRecent advances in deep neural networks, especially in generative adversarial networks (GAN), have shown remarkable progress in face image generations. However, most of the existing face image generators can only synthesize random face images, but are not able to control the attributes of the generated face images. Though conditional GAN based methods can manipulate the attributes to some extent, but can only generate low-resolution face images up to 256 × 256. In this study, based on StyleGAN, one of the state-of-the-art image generators for synthesizing high-quality face images, we propose a simple but efficient approach to generate high-resolution and hyper-realistic face images with any desired attribute. By training an attribute classifier to assign attribute labels to given synthesized face images, we build the links between latent vectors and face attributes. In such a way, the latent vectors can be grouped into different clusters, one cluster corresponding to one face attribute, respectively. We then extract the prototypes for the clusters, which are used to control the attribute of the generated face image. Extensive experiments demonstrate the effectiveness of the proposed approach for high-quality face image generation with predefined attributes. Qiyu Wei, Xulei Yang, Tong Sang, Huijiao Wang, Xiaofeng Zou, Zhongyao Cheng, Ziyuan Zhao, Zeng Zeng |
ICIP | 6 |
| 2022 | Exploring Structural Knowledge for Automated Visual Inspection of Moving TrainsabstractDeep learning methods are becoming the de-facto standard for generic visual recognition in the literature. However, their adaptations to industrial scenarios, such as visual recognition for machines, product streamlines, etc., which consist of countless components, have not been investigated well yet. Compared with the generic object detection, there is some strong structural knowledge in these scenarios (e.g., fixed relative positions of components, component relationships, etc.). A case worth exploring could be automated visual inspection for trains, where there are various correlated components. However, the dominant object detection paradigm is limited by treating the visual features of each object region separately without considering common sense knowledge among objects. In this article, we propose a novel automated visual inspection framework for trains exploring structural knowledge for train component detection, which is called SKTCD. SKTCD is an end-to-end trainable framework, in which the visual features of train components and structural knowledge (including hierarchical scene contexts and spatial-aware component relationships) are jointly exploited for train component detection. We propose novel residual multiple gated recurrent units (Res-MGRUs) that can optimally fuse the visual features of train components and messages from the structural knowledge in a weighted-recurrent way. In order to verify the feasibility of SKTCD, a dataset that contains high-resolution images captured from moving trains has been collected, in which 18 590 critical train components are manually annotated. Extensive experiments on this dataset and on the PASCAL VOC dataset have demonstrated that SKTCD outperforms the existing challenging baselines significantly. The dataset as well as the source code can be downloaded online (https://github.com/smartprobe/SKCD). Cen Chen 0002, Xiaofeng Zou, Zeng Zeng, Zhongyao Cheng, Le Zhang 0001, Steven C. H. Hoi |
IEEE Trans. Cybern. | 4 |
| 2022 | A Hybrid Deep Learning Based Framework for Component Defect Detection of Moving TrainsabstractDefect detection of trains is of great significance for operation safety and maintenance efficiency for railway maintenance. Nowadays, China railway system utilizes high-speed line scan cameras to capture images of critical parts of moving trains. The visual inspection on the images still heavily relies on manual interpretation. To reduce the labor requirements, we propose a novel two-stage deep learning based framework for component defect detection of moving trains. The proposed framework is composed of two major successive stages: detecting train components by using our proposed hierarchical object detection scheme (HOD), and detecting component defects based on multiple neural networks and image processing methods. Our proposed HOD can effectively detect and localize train components from large to small in a hierarchical way. Furthermore, a gated feature fusion method that can extract and combine the hierarchical contextual features and spatial contexts is also proposed to improve the performance. To the best of our knowledge, it is the first time in the literature that component defect detection of moving trains is systematically analyzed. Extensive experiments on real images from China railway system have demonstrated that our framework outperforms the state-of-the-art baselines significantly. Cen Chen 0002, Kenli Li 0001, Zhongyao Cheng, Francesco Piccialli, Steven C. H. Hoi, Zeng Zeng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Hierarchical Semantic Graph Reasoning for Train Component DetectionabstractRecently, deep learning-based approaches have achieved superior performance on object detection applications. However, object detection for industrial scenarios, where the objects may also have some structures and the structured patterns are normally presented in a hierarchical way, is not well investigated yet. In this work, we propose a novel deep learning-based method, hierarchical graphical reasoning (HGR), which utilizes the hierarchical structures of trains for train component detection. HGR contains multiple graphical reasoning branches, each of which is utilized to conduct graphical reasoning for one cluster of train components based on their sizes. In each branch, the visual appearances and structures of train components are considered jointly with our proposed novel densely connected dual-gated recurrent units (Dense-DGRUs). To the best of our knowledge, HGR is the first kind of framework that explores hierarchical structures among objects for object detection. We have collected a data set of 1130 images captured from moving trains, in which 17 334 train components are manually annotated with bounding boxes. Based on this data set, we carry out extensive experiments that have demonstrated our proposed HGR outperforms the existing state-of-the-art baselines significantly. The data set and the source code can be downloaded online at https://github.com/ChengZY/HGR. Cen Chen 0002, Kenli Li 0001, Xiaofeng Zou, Zhongyao Cheng, Wei Wei 0006, Qi Tian 0001, Zeng Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Multiple convolutional neural networks for multivariate time series prediction
Kenli Li 0001, Liqian Zhou, Yikun Hu 0001, Zhongyao Cheng, Jing Liu 0032, Cen Chen 0002 |
Neurocomputing | 5 |