VLDB 2026 Research / reviewers in the wild / expert
Ke Zhang 0005
dblp:20/4152-5
· DBLP profile ↗
21ranked-venue papers
9as first author
14since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Auto-metric network for open-set recognition
Shuoshi Li, Yuan Zhou 0006, Ke Zhang 0005, Sun-Yuan Kung |
Expert Syst. Appl. | 3 |
| 2026 | A Segmentation-Driven Editing Method for Bolt Defect Augmentation and DetectionabstractBolt defect detection is critical to ensure the safety of transmission lines. However, the scarcity of defect images and imbalanced data distributions significantly limit detection performance. To address this problem, we propose a segmentationdriven bolt defect editing method (SBDE) to augment the dataset. First, a bolt attribute segmentation model (Bolt-SAM) is proposed, which enhances the segmentation of complex bolt attributes through the CLAHE-FFT Adapter (CFA) and Multipart- Aware Mask Decoder (MAMD), generating high-quality masks for subsequent editing tasks. Second, a mask optimization module (MOD) is designed and integrated with the image inpainting model (LaMa) to construct the bolt defect attribute editing model (MOD-LaMa), which converts normal bolts into defective ones through attribute editing. Finally, an editing recovery augmentation (ERA) strategy is proposed to recover and put the edited defect bolts back into the original inspection scenes and expand the defect detection dataset. We constructed multiple bolt datasets and conducted extensive experiments. Experimental results demonstrate that the bolt defect images generated by SBDE significantly outperform state-of-the-art image editing models, and effectively improve the performance of bolt defect detection, which fully verifies the effectiveness and application potential of the proposed method. The code of the project is available at https://github.com/Jay-xyj/SBDE. Yangjie Xiao, Ke Zhang 0005, Jiacun Wang 0001, Xin Sheng 0001, Yurong Guo 0001, Meijuan Chen, Zehua Ren, Zhaoye Zheng, Zhenbing Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | A ground-based cloud image classification method for photovoltaic power prediction based on Convolutional Neural Networks and Vision Transformer
Chaojun Shi, Hongyin Xiang, Ke Zhang 0005, Sihao Ju, Leile Han |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Anti-vibration hammer defect detection based on structural knowledge representationabstractAnti-vibration hammer defects pose a significant risk to the safe operation of power transmission lines. Addressing the issues posed by the various manifestations of identical category defects and the similarities among different defects, this paper introduces a defect detection algorithm for anti-vibration hammers based on structural knowledge representation. Initially, a Structural Knowledge Enhancement (StKE) Module is proposed to conduct a statistical analysis of the aspect ratios of different defects in anti-vibration hammers, effectively extracting the structural features of the anti-vibration hammers and their corresponding defects. Subsequently, a Structural Knowledge Representation (StKR) Module is introduced, which bolsters the model’s ability to precisely locate defects that disrupt structural symmetry. The model incorporates the Coordinate Attention (CA) mechanism to acquire contextual information, thereby enhancing detection accuracy. Experimental results demonstrate that the improved model achieves a 6.1% increase in mean detection precision over the baseline model, and a notable improvement in detection accuracy compared to other advanced algorithms. Zhenbing Zhao, Guangxue Guo, Yitian Pan, Ke Zhang 0005, Yongjie Zhai |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | A decoupled scene-equipment fusion method for power substation equipment detection
Yurong Guo 0001, Ke Zhang 0005, Tiefeng Zhang, Zinian Wang, Haiqi Chen |
Knowl. Based Syst. | 3 |
| 2024 | Learning Conditional Prompt for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) strives to learn attributes and objects from seen compositions and transfer the acquired knowledge to unseen compositions. Existing methods either learn primitive concepts in an entangled manner, leading to the model relying on spurious correlations between attributes and objects. Alternatively, they adopt a decoupled approach, causing the model to overlook relationships between attributes and objects. In this paper, we propose a conditional prompting (CoP) method to enhance the performance of vision-language models (e.g., CLIP) in CZSL. Specifically, we utilize two image-to-word mapping networks to learn pseudo attribute and object word embeddings that can represent the corresponding semantics of input images. Subsequently, the model recognizes one concept based on another generated pseudo word embeddings, enabling the recognition of individual sub-concepts while leveraging the image-specific correlations between attributes and objects. The experimental results on three CZSL benchmarks indicate that the proposed method achieves competitive performance compared to previous state-of-the-art methods. Tian Zhang 0029, Kongming Liang, Ke Zhang 0005, Zhanyu Ma |
ICME | 3 |
| 2024 | CloudFU-Net: A Fine-Grained Segmentation Method for Ground-Based Cloud Images Based on an Improved Encoder-Decoder StructureabstractThe segmentation of ground-based cloud image is a crucial aspect of ground-based cloud observation, with significant implications for meteorological forecasting, photovoltaic power prediction, and other related tasks. At present, the proposed method of ground-based cloud image segmentation only separates cloud from the sky background without further classifying the cloud categories. Clouds have rich fine-grained semantic features, and different types of clouds have different effects on solar irradiance, which in turn has different effects on photovoltaic power. In this paper, a fine-grained segmentation method for ground-based cloud images is proposed, which is based on an improved encoder-decoder structure named CloudFU-Net. Firstly, a ground-based cloud image fine-grained segmentation data set for Photovoltaic power prediction is constructed, and the clouds are divided into five categories with different colors under the guidance of meteorologists; Secondly, Selective Kernel (SK) is introduced in CloudFU-Net encoder to better capture cloud of different sizes. Then, Parallel dilated convolution model (PDCM) is proposed to segment small target clouds more accurately. Finally, a Content-Aware ReAssembly of Features(CARAFE) is introduced into CloudFU-Net decoder to replace the original interpolating upsampling to better recover fine-grained semantic features. Lastly, the experimental results show that the proposed CloudFU-Net has the best segmentation performance compared with other segmentation models, with Miou reaching 61.9%, which can efficiently segment different cloud genera and lay a solid foundation for accurate prediction of photovoltaic power. Chaojun Shi, Zibo Su, Ke Zhang 0005, Xiongbin Xie, Xian Zheng, Qiaochu Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | LSKF-YOLO: Large Selective Kernel Feature Fusion Network for Power Tower Detection in High-Resolution Satellite Remote Sensing ImagesabstractWith the rapid development of high-resolution satellite remote sensing observation technology, power tower detection based on satellite remote sensing images has become a key research focus for power intelligent inspection. However, the performance of power tower detection in satellite remote sensing images needs improvement due to complex back-grounds, small and non-uniform target sizes. To address this, this paper first constructs a multi-scene high-resolution satellite remote sensing power tower dataset, and then proposes the LSKF-YOLO network for high-resolution satellite remote sensing images. This network primarily consists of a large spatial kernel selective attention fusion module and a multi-scale feature alignment fusion structure. The large spatial selective kernel mechanism is improved by using the attentional feature fusion module, provides richer feature information for accurately locating the position of the power tower. The multi-scale feature alignment fusion structure effectively utilizes low-level semantic information, mitigates feature ambiguity in deeper network layers, and enables multi-scale feature fusion of power towers within complex backgrounds. Additionally, the introduction of MPDIoU enhances CIoU, further improving the model’s performance. The results demonstrate that the F1 score and mAP0.5 of the LSKF-YOLO network reach 0.764 and 77.47%, respectively. Compared with other deep learning-based satellite remote sensing power tower inspection methods, the LSKF-YOLO network significantly enhances detection accuracy and provides crucial technical support for intelligent inspection of power lines via satellite remote sensing. Chaojun Shi, Xian Zheng, Zhenbing Zhao, Ke Zhang 0005, Zibo Su, Qiaochu Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Transmission Line Component Defect Detection Based on UAV Patrol Images: A Self-Supervised HC-ViT MethodabstractThe unmanned aerial vehicle (UAV) patrol inspection has become an efficient method to ensure the operation condition of transmission lines. The detection of key components with defects in transmission lines is a critical task in maintaining a power system’s stability. However, the complex inspection environment and the imbalance between the number of normal component samples and that of defect samples significantly affect the detection accuracy. In this article, we present a novel method for defect detection in UAV patrol images, based on a hierarchical convolutional vision transformer (HC-ViT) and a simple contrastive masked autoencoder (SC-MAE). The HC-ViT backbone integrates the advantages of vision transformer and convolution, while the SC-MAE is a self-supervised learning method that extracts useful features from normal samples. By introducing the normal features into the backbone, we enhance the performance of the defect detection task. We demonstrate the effectiveness of our method through experiments, and show that it can leverage a large amount of unlabeled normal images, reducing the need for manual annotation. Our method offers a new way to exploit the potential features of patrol inspection images. Ke Zhang 0005, Ruiheng Zhou, Jiacun Wang 0001, Yangjie Xiao, Xiwang Guo 0001, Chaojun Shi |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Multi-Head Uncertainty Inference for Adversarial Attack DetectionabstractDeep neural networks (DNNs) are sensitive and susceptible to tiny perturbations by adversarial attacks which cause erroneous predictions. Various methods, including adversarial defense and uncertainty inference (UI), have been developed to overcome adversarial attacks in recent years. In this paper, we propose a multi-head uncertainty inference (MH-UI) framework for detecting adversarial attack examples. We adopt a multi-head architecture with multiple prediction heads (i.e., classifiers) to obtain predictions from different depths in the DNNs and introduce shallow information for the UI. Using independent heads at different depths, the normalized predictions are assumed to follow the same Dirichlet distribution, and we estimate the distribution parameter of it by moment matching. Cognitive uncertainty brought by the adversarial attacks will be reflected and amplified in the distribution. Experimental results show that the proposed MH-UI framework has good performance in different settings of adversarial attack detection tasks. Songyun Yang, Jiyang Xie 0001, Zhongwei Si, Ke Zhang 0005, Kongming Liang |
ICASSP | 6 |
| 2023 | Multi-task multi-scale attention learning-based facial age estimationabstractAbstract Face‐based age estimation strongly depends on deep residual networks (ResNets), used as the backbone in the relevant research. However, ResNet‐based methods ignore the importance of some large‐scale facial information and other facial age attributes. Inspired by the attention mechanism, a multi‐task learning framework for face‐based age estimation called multi‐task multi‐scale attention is proposed. First, the authors embed the alternative strategy structure of dilated convolution into ResNet34 to construct a multi‐scale attention module (MSA) to improve the network's receptive field, which extracts local age‐sensitive information while obtaining multi‐scale features. The MSA can have a larger receptive field to extract both large‐scale and local detailed feature information. Second, multi‐task learning network structures are built to predict gender and race, which can share rigid network parameters to improve age estimation and improve the accuracy of age estimation by other age‐related parameters. Finally, the Kullback‐Leibler divergence loss is adopted between a Dirac delta label and a Gaussian prediction to guide the training. The numerical tests on the MORPH Album II and Adience datasets prove the superiority of the proposed method over other state‐of‐the‐art ones. Chaojun Shi, Ke Zhang 0005, Xiaohan Feng |
IET Signal Process. | 3 |
| 2022 | Real-time facial expression recognition based on iterative transfer learning and efficient attention networkabstractAbstract Real‐time facial expression recognition is the basis for computers to understand human emotions and detect abnormalities in time. To effectively solve the problems of server overload and privacy information leakage, a real‐time facial expression recognition method based on iterative transfer learning and efficient attention network (EAN) for edge resource‐constrained scenes is proposed in this paper. Firstly, an EAN is designed with its parameter number and computation amount strictly limited by depth separable convolution and local channel attention mechanism. Then, the soft labels of facial expression data were obtained by EAN based on the idea of knowledge distillation, so as to provide more supervision information for the training process. Finally, an iterative transfer learning method of teacher‐student (T‐S) network was proposed; it refines the soft labels of the teacher network and further improves the recognition accuracy of the student network. The tests on the public datasets, FER2013 and RAF‐DB, show that this method can significantly reduce the model complexity and achieve high recognition accuracy. Compared with other advanced methods, the proposed method strikes a good balance between complexity and accuracy, and well meets the real‐time deployment requirements of facial expression recognition technology for edge resource‐constrained scenes. Yinghui Kong, Shuaitong Zhang, Ke Zhang 0005, Qiang Ni, Jungong Han |
IET Image Process. | 3 |
| 2021 | HRM-CenterNet: A High-Resolution Real-time Fittings Detection MethodabstractMost successful fittings detectors are anchor-based, which is challenging to meet the lightweight and real-time requirements of the edge computing system. We propose a high-resolution real-time network HRM-CenterNet. Firstly, the lightweight MobileNetV3 is used to extract multi-level features from images. Then, to improve the resolution of the feature maps and reduce the spatial semantic information loss during the image downsampling process, a high-resolution feature fusion network based on iterative aggregation is introduced. Finally, we conduct experiments on the PASCAL VOC dataset and fittings dataset. The results show that HRM-CenterNet improves accuracy as well as robustness, and meets the performance requirements of real-time edge detection. Ke Zhang 0005, Xiwang Guo 0001, Xiaohan Feng, Ying Tang 0001 |
SMC | 1 |
| 2021 | Competing ratio loss for discriminative multi-class image classification
Ke Zhang 0005, Yurong Guo 0001, Dongliang Chang, Zhenbing Zhao, Zhanyu Ma, Tony X. Han |
Neurocomputing | 1 |
| 2020 | Fine-Grained Age Estimation in the Wild With Attention LSTM NetworksabstractAge estimation from a single face image has been an essential task in the field of human-computer interaction and computer vision, which has a wide range of practical application values. Accuracy of age estimation of face images in the wild is relatively low for existing methods, because they only take into account the global features, while neglecting the fine-grained features of age-sensitive areas. We propose a novel method based on our attention long short-term memory (AL) network for fine-grained age estimation in the wild, inspired by the fine-grained categories and the visual attention mechanism. This method combines the residual networks (ResNets) or the residual network of residual network (RoR) models with LSTM units to construct AL-ResNets or AL-RoR networks to extract local features of age-sensitive regions, which effectively improves the age estimation accuracy. First, a ResNets or a RoR model pretrained on ImageNet dataset is selected as the basic model, which is then fine-tuned on the IMDB-WIKI-101 dataset for age estimation. Then, we fine-tune the ResNets or the RoR on the target age datasets to extract the global features of face images. To extract the local features of age-sensitive regions, the LSTM unit is then presented to obtain the coordinates of the age-sensitive region automatically. Finally, the age group classification is conducted directly on the Adience dataset, and age-regression experiments are performed by the Deep EXpectation algorithm (DEX) on MORPH Album 2, FG-NET and 15/16LAP datasets. By combining the global and the local features, we obtain our final prediction results. Experimental results illustrate the effectiveness and robustness of the proposed AL-ResNets or AL-RoR for age estimation in the wild, where it achieves better state-of-the-art performance than all other convolutional neural network (CNN) methods on the Adience, MORPH Album 2, FG-NET and 15/16LAP datasets. Ke Zhang 0005, Xingfang Yuan, Xinyao Guo, Ce Gao, Zhenbing Zhao, Zhanyu Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Channel-Wise and Feature-Points Reweights Densenet for Image ClassificationabstractRecent network research has demonstrated that the performance of convolutional neural networks can be improved by introducing a learning block that capture spatial correlations and channel-wise correlations . In this work, we propose a novel Channel-wise and Feature-points Reweights DenseNet (CAPR-DenseNet) architecture. The CAPR-DenseNet improves the representation power of the DenseNet by adaptively recalibrating the channel-wise feature responses and explicitly modeling the interdependencies between feature-points. First, in order to perform dynamic channel-wise feature recalibration, we construct the Channel-wise Feature Reweight DenseNet (CFR-DenseNet) by introducing the Squeeze-and-Excitation Module (SEM) to DenseNet. Then, we present a novel CAPR-DenseNet by adding a Feature-points Reweight Module (FPRM) to the CFR-DenseNet. Through massive experiments, we demonstrate that by recalibrating the channel-wise feature and the feature-points responses. Our CAPR-DenseNet performs better than DenseNet across challenging datasets CIFAR-10 and CIFAR-100. Ke Zhang 0005, Yurong Guo 0001, Jinsha Yuan, Zhanyu Ma, Zhenbing Zhao |
ICIP | 1 |
| 2019 | Competing Ratio Loss for Multi-class image ClassificationabstractCross-entropy loss function (CEL) is widely used for training a multi-class classification deep convolutional neural network (DCNN). While CEL has been successfully implemented in image classification tasks, it only focuses on the posterior probability of correct class when the labels of training images are one-hot. It cannot be discriminated against the classes not belong to correct class (wrong classes) directly. Negative Log Likelihood Ratio Loss (NLLR) is proposed to better discriminate the correct class from competing wrong classes. But optimization of the loss function is normally presented as a minimization problem. In training DCNN, the value of NLLR is not constantly positive or negative, which affects the convergence of NLLR adversely. So, we propose competing ratio loss (CRL), which calculates the posterior probability ratio between the correct class and competing wrong classes to better widen the difference between the probability of the correct class and the probabilities of wrong classes, which also assures the value of CRL is constantly positive. Through massive experiments, we demonstrate the effectiveness and robustness of CRL on deep convolutional neural networks, our CRL outperforms CEL and NLLR on CIFAR-10/100 datasets. Ke Zhang 0005, Yurong Guo 0001, Zhenbing Zhao, Zhanyu Ma |
VCIP | 1 |
| 2018 | Optimization Method of Residual Networks of Residual Networks for Image Classification
Long Lin, Liru Guo, Yingqun Kuang, Ke Zhang 0005 |
ICIC (3) | 5 |
| 2018 | Fine-Grained Age Group Classification in the wildabstractAge estimation from a single face image has been an essential task in the field of human-computer interaction and computer vision which has a wide range of practical application value. Concerning the problem that accuracy of age estimation of face images under unconstrained conditions are relatively low for existing methods, we propose a method based on Attention LSTM network for Fine-Grained age group classification in the wild based on the idea of Fine-Grained categories and visual attention. This method combines ResNets models with LSTM unit to construct AL-ResNets networks to extract age-sensitive local regions, which effectively improves age estimation accuracy. Firstly, ResNets model pre-trained on ImageNet data set is selected as the basic model, which is then fine-tuned on the IMDB-WIKI-101 data set for age estimation. Then, we fine-tune ResNets on the Adience data set to extract the global features of face images. To extract the local characteristics of age-sensitive areas, the LSTM unit is then presented to obtain the coordinates of the age-sensitive region automatically. Finally, by combining the global and local features, we got our final prediction results. Our experiments illustrate the effectiveness of AL-ResNets for age group classification in the wild, where it achieves new state-of-the-art performance than all other CNN methods on the Adience data set. Ke Zhang 0005, Xingfang Yuan, Xinyao Guo, Ce Gao, Zhenbing Zhao |
ICPR | 1 |
| 2018 | Residual Networks of Residual Networks: Multilevel Residual NetworksabstractA residual networks family with hundreds or even thousands of layers dominates major image recognition tasks, but building a network by simply stacking residual blocks inevitably limits its optimization ability. This paper proposes a novel residual network architecture, residual networks of residual networks (RoR), to dig the optimization ability of residual networks. RoR substitutes optimizing residual mapping of residual mapping for optimizing original residual mapping. In particular, RoR adds levelwise shortcut connections upon original residual networks to promote the learning capability of residual networks. More importantly, RoR can be applied to various kinds of residual networks (ResNets, Pre-ResNets, and WRN) and significantly boost their performance. Our experiments demonstrate the effectiveness and versatility of RoR, where it achieves the best performance in all residual-network-like structures. Our RoR-3-WRN58-4 + SD models achieve new state-of-the-art results on CIFAR-10, CIFAR-100, and SVHN, with the test errors of 3.77%, 19.73%, and 1.59%, respectively. RoR-3 models also achieve state-of-the-art results compared with ResNets on the ImageNet data set. Ke Zhang 0005, Tony X. Han, Xingfang Yuan, Liru Guo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Age group classification in the wild with deep RoR architectureabstractAutomatically predicting age group from face images acquired in unconstrained conditions is an important and challenging task in many real-world applications. Nevertheless, the conventional methods with manually-designed features on in-the-wild benchmarks are unsatisfactory because of incompetency to tackle large variations in unconstrained images. In this paper, we propose a new CNN based method for age group classification leveraging Residual Networks of Residual Networks (RoR), which exhibits better optimization ability for age group classification than other CNN architectures. Moreover, two modest mechanisms based on observation of the characteristics of age group are presented to further improve the performance of age estimation. Our experiments illustrate the effectiveness of RoR method for age estimation in the wild, where it achieves better performance than other CNN methods. Finally, the Pre-RoR-58+SD with two mechanisms achieves new state-of-the-art results on Adience benchmark. Ke Zhang 0005, Liru Guo, Ce Gao, Zhenbing Zhao, Xingfang Yuan |
ICIP | 1 |