VLDB 2026 Research / reviewers in the wild / expert
Kai Zeng 0010
dblp:80/1651-10
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-2745-1253ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Expert Imitation With Purifying Latent Feature for Generalization in Visual Reinforcement LearningabstractThe generalization ability of visual reinforcement learning, which allows the policy trained in the source domain to guide agents in similar unknown target environments, is one of the cores applied to visual navigation and autonomous driving. Recently, methods such as data augmentation techniques, self-supervised learning methods, and the generative adversarial network were employed to enhance the generalization capability of policy neural networks in visual reinforcement learning. However, current state-of-the-art methods, after utilizing domain-general latent features to train the RL policy, result in the loss of certain state-specific features, leading to diminished policy performance following generalization. To tackle these challenges, we designed a technical framework called self-expert imitation with purifying latent features, which enables the trained policy to effectively guide agents in scenarios similar to the training environment, without compromising the performance of the policy-guided agent in task completion. Additionally, a novel method was developed for separating domain-general and domain-specific latent vectors based on a variational autoencoder, enabling the domain-general component to exhibit strong and stable zero-shot generalization performance in unseen visually similar domains. Extensive experiments on the CarRacing game demonstrated that our approach achieves strong and stable generalization performance in unseen environments, without compromising the performance of the policy in guiding agents to complete tasks. Lin Chen 0034, Yang Mo, Yaonan Wang 0001, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Zhen Zhou 0003, Sifei Wang, Danwei Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power SceneabstractIn the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line InspectionabstractReal-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Deep spatial and discriminative feature enhancement network for stereo matching
Guowei An, Yaonan Wang 0001, Kai Zeng 0010, Qing Zhu 0003, Xiaofang Yuan |
Vis. Comput. | 3 |
| 2024 | Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent FeatureabstractGeneralizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techniques, self-supervised learning methods, and the generative adversarial network framework to train policy neural networks with enhanced generalizability. However, current methods, upon extracting domain-general latent features, further utilize these features to train the reinforcement learning policy, resulting in a decline in the performance of the learned policy guiding the agent to accomplish tasks. To tackle these challenges, a framework of self-expert imitation with purifying latent features was devised, empowering the policy to achieve robust and stable zero-shot generalization performance in visually similar domains previously unseen, without diminishing the performance of guiding the agent to accomplish tasks. The extraction method of domain-general latent features is proposed to enhance their quality based on the variational autoencoder. Extensive experiments have shown that our policy, compared with state-of-the-art counterparts, does not diminish the performance of the policy guiding the agent to accomplish tasks after generalization. Lin Chen 0034, Jianan Huang 0002, Zhen Zhou 0003, Yaonan Wang 0001, Yang Mo, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Danwei Wang |
IROS | 7 |
| 2024 | Deep Stereo Network With MRF-Based Cost AggregationabstractDespite the remarkable progress made in learning-based stereo-matching algorithms, it is an open challenge for stereo-matching in disparity discontinuities and textureless regions. In this paper, we propose the deep Markov Random Field based cost aggregation network (DMCA-Net) for stereo matching, which is an end-to-end model-driven network architecture. This architecture introduces an efficient feature extraction network to extract richer textual and contextual feature information for stereo feature similarity representation at multi-stages and levels. Furthermore, with the aim of alleviating the edge-fattening phenomenon at disparity discontinuities and generating accurate disparities in textureless regions, we proposed the differentiable Markov Random Field model for cost aggregation, where the model’s data term utilizes image detail information, such as boundary and contour features, to guide matching cost aggregation, and the model’s smoothness term penalizes the adjacency similarity of the cost between the four-nearest neighboring pixel pairs to predict the disparity in textureless regions. The detailed experiment demonstrates that DMCA network achieves competitive performance on the SceneFlow, KITTI 2012, KITTI 2015, and Middlebury 2014 datasets. Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Deep Stereo Matching with Superpixel Based Feature and Cost
Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
PRCV (2) | 1 |
| 2023 | SSAPN: Spectral-Spatial Anomaly Perception Network for Unsupervised Vaccine DetectionabstractVaccines are the most significant and effective way to prevent disease and safeguard human health. However, it is easy to produce or mix foreign matters during the manufacturing process. Moreover, foreign matters are extremely faint that it is difficult to obtain images and detect them accurately. To tackle imaging challenges, in this article, we built a hyperspectral imaging system to construct a first-of-its-kind HSI dataset with pixel-level annotation for vaccine anomaly detection, where the vaccine comes from the actual pharmaceutical company. To address the problem of low detection accuracy, we propose a spectral–spatial anomaly perception module joint with an unsupervised autoencoder network (SSAPN), in which nonlinear features learned from the encoder are divided into nonoverlapping patches and mapped to efficiently encode spectral and spatial feature information. The spectral–spatial multilayer perceptrons (MLP) module consists of continuous and alternating spectral MLP with spatial MLP, which achieves spectral with spatial perception in the global receptive field, captures long-range dependencies, and extracts the most discriminative spectral–spatial features. Experimental results show that our SSAPN model outperforms other state-of-the-art anomaly detection methods in terms of both detection and generalization performance. This work will help speed up the production process in the vaccine pharmaceutical industry and ensure vaccine quality. Ating Yin, Yaonan Wang 0001, Yurong Chen 0003, Kai Zeng 0010, Hui Zhang 0023, Jianxu Mao |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Deep Confidence Propagation Stereo NetworkabstractStereo matching depth estimation based on rectified image pairs is of great importance to many computer vision tasks such as vehicle navigation and autonomous driving. Confidence measures are typically used to refine stereo matching results, which provides robustness and efficiency for disparity estimation. However, previous learning-based confidence methods for stereo matching usually use the middle results or composition as a post-processing step to refine the stereo matching results. This cannot be optimized end-to-end and the performance is limited by the quality of the tri-modal output. To handle this issue, in this paper, we pursue an end-to-end hierarchical architecture and propose a differentiable confidence propagation (DCP) model of a cost aggregation network for stereo matching. The DCP model is integrated into an end-to-end neural network hierarchical architecture to guide matching cost volume aggregation. More specifically, to better represent the similarity of left and right feature maps, we extract unary context feature maps with an effective attention mechanism for matching cost construction. Moreover, we aggregate the cost volume with the multiple stacked DCP cost aggregation (DCPCA) networks to generate a more reliable and finer cost volume. This network suppresses multi-level disparity maps. Each output disparity is supervised with different training weights to learn in a coarse-to-fine way. Our method outperforms previous methods on the Sceneflow dataset by achieving the$0.6735px$EPE error, achieving 1.53% D1-all metric of Non-occluded pixels regions and 0.72% Non-occluded pixels of$5px$metric on KITTI 2015 and 2012 dataset. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our DCPCA-Net can significantly minimize the trade-off between accuracy and efficiency for stereo matching. Kai Zeng 0010, Yaonan Wang 0001, Wei Wang 0025, Hui Zhang 0023, Jianxu Mao, Qing Zhu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Deep Stereo Matching With Hysteresis Attention and Supervised Cost Volume ConstructionabstractStereo matching disparity prediction for rectified image pairs is of great importance to many vision tasks such as depth sensing and autonomous driving. Previous work on the end-to-end unary trained networks follows the pipeline of feature extraction, cost volume construction, matching cost aggregation, and disparity regression. In this paper, we propose a deep neural network architecture for stereo matching aiming at improving the first and second stages of the matching pipeline. Specifically, we show a network design inspired by hysteresis comparator in the circuit as our attention mechanism. Our attention module is multiple-block and generates an attentive feature directly from the input. The cost volume is constructed in a supervised way. We try to use data-driven to find a good balance between informativeness and compactness of extracted feature maps. The proposed approach is evaluated on several benchmark datasets. Experimental results demonstrate that our method outperforms previous methods on SceneFlow, KITTI 2012, and KITTI 2015 datasets. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Caiping Liu, Weixing Peng, Yin Yang 0004 |
IEEE Trans. Image Process. | 1 |
| 2022 | Deep Progressive Fusion Stereo NetworkabstractStereo matching depth estimation for rectified image pairs is of great importance to many compute vision tasks, specifically in autonomous driving. With the flourishing of convolution neural networks, responsible depth estimation of stereo matching with artificial intelligence is the most severe challenge for autonomous driving in recent years. Previous research on end-to-end trainable stereo matching networks has usually used cascading convolution blocks with down-sampling or pooling operations to extract the unary features required for matching cost construction. Such approaches lack a reconstruction stage for increasing feature map pixel-wise alignment and strength, factors which play an important role in representing the similarity between stereo image pairs. To address this issue, in this paper, we propose the progressive fusion stereo matching network (PFSM-Net). We exploit an encoder-decoder feature extraction network architecture for multi-stage and -scale dynamic feature extraction. Moreover, we propose a group-wise concatenation method to construct the cost volume, which provides a more efficient cost volume for cost aggregation. Furthermore, we propose the use of multi-scale cost aggregation networks with a progressive fusion strategy. The aggregated cost volume is progressively fused with the multi-stage and -scale cost volume as the size of the cost volume increases. Multi-stage and -scale outputs are supervised with and learned in a coarse-to-fine manner. Experimental results demonstrate that our method outperforms previous methods on the SceneFlow, KITTI 2012, and KITTI 2015 datasets. Kai Zeng 0010, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Hui Zhang 0023 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Deep residual deconvolutional networks for defocus blur detectionabstractAbstract Accurate defocus blur detection has instigated wide research interest for the last few years. However, it is still a meaningful yet challenging machine vision task, and most methods rely on prior knowledge. Convolutional neural networks have proved the huge success for different tasks within the computer vision, and machine learning flew. A simple yet effective method of defocus blur detection was proposed in this paper, which by applying the deep residual convolutional encoder‐decoder network. The aims of DRDN is to automatically generate pixel‐level predictions for defocus blur images, and reconstruct output detection results of the same size as the input, which by performing several deconvolution operations at multiple scales through the transposed convolution, and skip connection. Afterwards, we used the slide window detection strategy and traversed the input image with a certain stride. Experiments on challenging benchmarks of defocus blur detection show that our algorithm achieved state‐of‐the‐art performance, and powerfully balanced the detection accuracy, and detection time. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Xianen Zhou |
IET Image Process. | 1 |
| 2019 | A Local Metric for Defocus Blur Detection Based on CNN Feature LearningabstractDefocus blur detection is an important and challenging task in computer vision and digital imaging fields. Previous work on defocus blur detection has put a lot of effort into designing local sharpness metric maps. This paper presents a simple yet effective method to automatically obtain the local metric map for defocus blur detection, which based on the feature learning of multiple convolutional neural networks (ConvNets). The ConvNets automatically learn the most locally relevant features at the super-pixel level of the image in a supervised manner. By extracting convolution kernels from the trained neural network structures and processing it with principal component analysis, we can automatically obtain the local sharpness metric by reshaping the principal component vector. Meanwhile, an effective iterative updating mechanism is proposed to refine the defocus blur detection result from coarse to fine by exploiting the intrinsic peculiarity of the hyperbolic tangent function. The experimental results demonstrate that our proposed method consistently performed better than the previous state-of-the-art methods. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Junyang Liu, Weixing Peng, Nankai Chen |
IEEE Trans. Image Process. | 1 |