EDBT 2026 Demo / reviewers in the wild / expert
Keyang Cheng
dblp:67/1926
· DBLP profile ↗
41ranked-venue papers
22as first author
29since 2021 · last 2026
0000-0001-5240-1605ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Infrared small UAV target detection via depthwise separable residual dense attention network
Keyang Cheng, Changsheng Peng |
J. Vis. Commun. Image Represent. | 1 |
| 2026 | DGPT: Domain-guided prompt tuning for improved generalization in vision-language models
Keyang Cheng, Haoming Hou, Liutao Wei, Akira David Diaz Campaña |
Knowl. Based Syst. | 1 |
| 2026 | Interpreting networks via semantic probes with correction and counterfactuals
Keyang Cheng, Ligang He, Maozhen Li |
Pattern Recognit. | 1 |
| 2026 | Adaptive proximal regularization for image smoothing
Yang Yang 0046, Shunli Ji, Lanling Zeng, Keyang Cheng |
Pattern Recognit. | 4 |
| 2026 | Part-Based Feature Complementary Denoising for Unsupervised Person Re-Identification
Qing Tian 0001, Bin Wang 0062, Jiashuo Shen, Keyang Cheng, Weihua Ou, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Inter-image Token Relation Learning for weakly supervised semantic segmentation
Jingfeng Tang, Keyang Cheng, Liutao Wei, Yongzhao Zhan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Constraint embedding for prompt tuning in vision-language pre-trained model
Keyang Cheng, Liutao Wei, Jingfeng Tang, Yongzhao Zhan 0001 |
Multim. Syst. | 1 |
| 2025 | Enhancing Open-Set Domain Adaptation through Optimal Transport and Adversarial Learning
Qing Tian 0001, Keyang Cheng, Tinghuai Ma |
Neural Networks | 3 |
| 2024 | BAMG: Text-Based Person Re-identification via Bottlenecks Attention and Masked Graph Modeling
Keyang Cheng, Wenxuan Zou, Hongjian Gu, Anxiang Ouyang |
ACCV (8) | 1 |
| 2024 | Parallel Interpretation Network via Semantic Visual Probe and Counterfactual Verification
Keyang Cheng |
ICONIP (1) | 2 |
| 2024 | Interpreting Convolutional Neural Network Decision via Pixel-Wise Interaction Hierarchy Graph
Keyang Cheng |
ICPR (7) | 1 |
| 2024 | Advancements in Photorealistic Style Translation with a Hybrid Generative Adversarial Network
Keyang Cheng, Rabia Tahir |
PRCV (4) | 1 |
| 2024 | Person re-identification via deep compound eye network and pose repair moduleabstractAbstract Person re‐identification is aimed at searching for specific target pedestrians from non‐intersecting cameras. However, in real complex scenes, pedestrians are easily obscured, which makes the target pedestrian search task time‐consuming and challenging. To address the problem of pedestrians' susceptibility to occlusion, a person re‐identification via deep compound eye network (CEN) and pose repair module is proposed, which includes (1) A deep CEN based on multi‐camera logical topology is proposed, which adopts graph convolution and a Gated Recurrent Unit to capture the temporal and spatial information of pedestrian walking and finally carries out pedestrian global matching through the Siamese network; (2) An integrated spatial‐temporal information aggregation network is designed to facilitate pose repair. The target pedestrian features under the multi‐level logic topology camera are utilised as auxiliary information to repair the occluded target pedestrian image, so as to reduce the impact of pedestrian mismatch due to pose changes; (3) A joint optimisation mechanism of CEN and pose repair network is introduced, where multi‐camera logical topology inference provides auxiliary information and retrieval order for the pose repair network. The authors conducted experiments on multiple datasets, including Occluded‐DukeMTMC, CUHK‐SYSU, PRW, SLP, and UJS‐reID. The results indicate that the authors’ method achieved significant performance across these datasets. Specifically, on the CUHK‐SYSU dataset, the authors’ model achieved a top‐1 accuracy of 89.1% and a mean Average Precision accuracy of 83.1% in the recognition of occluded individuals. Hongjian Gu, Wenxuan Zou, Keyang Cheng, Humaira abdul Ghafoor, Yongzhao Zhan 0001 |
IET Comput. Vis. | 3 |
| 2024 | β-CLVAE: a semantic disentangled generative model
Keyang Cheng, Chunyun Meng, Guojian Ma, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 1 |
| 2024 | Tiny Object Detection via Regional Cross Self-Attention NetworkabstractAs vision sensor technology continues to evolve, the requirements for detecting targets of interest in the images captured by the sensors are increasing. Considering fast detection and high accuracy, the industry favors geometric key point-based solutions. However, there are a large number of small and fuzzy objects in the real world. Geometric key point detectors do not effectively utilize the contextual features of the region of interest, leading to excessive false positive and false negative results. In this work, a simple, effective, and interpretable tiny object detection method called Regional Cross Self-Attention Object Detection Network (RCSANet) is proposed. It adopts Region Proposal Networks and transformers to capture regional background relations and uses regional background relations to generate key point sequences. The regional cross self-attention mechanism is introduced to curtail computation redundancy and minimize the interference of redundant information to the target region. Additionally, a position coding called dynamic implicit position coding is proposed to cooperate with regional cross self-attentiveness. Dynamic implicit location coding can encode arbitrarily long input sequences. The computational cost of RCSANet is significantly lower than that of state-of-the-art object detection solutions. Moreover, RCSANet improves the performance on the four benchmark datasets, of MSCOCO, Tinyperson, DOTA, and AI-TOD, by about 3.0%AP. Keyang Cheng, Honggang Cui, Humaira abdul Ghafoor, Qirong Mao, Yongzhao Zhan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Cross-Block Sparse Class Token Contrast for Weakly Supervised Semantic SegmentationabstractMost existing Vision Transformer-based frameworks for weakly supervised semantic segmentation utilize class activation maps to generate pseudo masks. Although it mitigates the class-agnostic issue, this approach still suffers from misclassification and noise in segmentation results. To overcome these limitations, we propose an attention-based framework named Cross-block Sparse Class Token Contrast (CB-SCTC), which incorporates Dynamic Sparse Attention module (DSA) and Cross-block Class Token Contrast scheme (CB-CTC). Specifically, the proposed Cross-block Class Token Contrast scheme forces diversity between the final class tokens by learning from the lower similarity of the class tokens in the relatively shallower blocks. Moreover, the Dynamic Sparse Attention module is designed to post-process the output from the softmax function in the attention mechanism to reduce noise. Extensive experiments prove the proposed framework is a valid alternative to class activation maps. Our framework demonstrates competitive mIoU scores on the PASCAL VOC 2012(val:75.5%, test:75.2%) and MS COCO 2014 dataset(val:46.9%). Our code is available athttps://github.com/Jingfeng-Tang/CB-SCTC. Keyang Cheng, Jingfeng Tang, Hongjian Gu, Maozhen Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | FlowMFD: Characterisation and classification of tor traffic using MFD chromatographic features and spatial-temporal modellingabstractAbstract Tor traffic tracking is valuable for combating cybercrime as it provides insights into the traffic active on the Tor network. Tor‐based application traffic classification is one of the tracking methods, which can effectively classify Tor application services. However, it is not effective in classifying specific applications due to more complicated traffic patterns in the spatial and temporal dimensions. As a solution, the authors propose FlowMFD, a novel Tor‐based application traffic classification approach using amount‐frequency‐direction (MFD) chromatographic features and spatial‐temporal modelling. Expressly, FlowMFD mines the interaction pattern between Tor applications and servers by analysing the time series features (TSFs) of different size packets. Then MFD chromatographic features (MFDCF) are designed to represent the pattern. Those features integrate multiple low‐dimensional TSFs into a single plane and retain most pattern information. In addition, FlowMFD utilises a cascaded model with a two‐dimensional convolutional neural network (2D‐CNN) and a bidirectional gated recurrent unit to capture spatial‐temporal dependencies between MFDCF. The authors evaluate FlowMFD under the public ISCXTor2016 dataset and the self‐collected dataset, where we achieve an accuracy of 92.1% (4.2%↑) and 88.3% (4.5%↑), respectively, outperforming state‐of‐the‐art comparison methods. Liukun He, Liangmin Wang 0001, Keyang Cheng |
IET Inf. Secur. | 3 |
| 2023 | Sonar image garbage detection via global despeckling and dynamic attention graph optimization
Keyang Cheng, Liuyang Yan, Yi Ding 0001, Maozhen Li 0001, Humaira abdul Ghafoor |
Neurocomputing | 1 |
| 2023 | Logical Topology Inference via CPGCN Joint Optimizing With Pedestrian Re-IdabstractWith the rise of artificial intelligence, deep learning has become the main research method of pedestrian recognition re-identification (re-id). However, most of the existing researches usually just determine the retrieval order based on the geographical location of cameras, which ignore the spatio-temporal logic characteristics of pedestrian flow. Furthermore, most of these methods rely on common object detection to detect and match pedestrians directly, which will separate the logical connection between videos from different cameras. In this research, a novel pedestrian re-identification model assisted by logical topological inference is proposed, which includes: 1) a joint optimization mechanism of pedestrian re-identification and multicamera logical topology inference, which makes the multicamera logical topology provides the retrieval order and the confidence for re-identification. And meanwhile, the results of pedestrian re-identification as a feedback modify logical topological inference; 2) a dynamic spatio-temporal information driving logical topology inference method via conditional probability graph convolution network (CPGCN) with random forest-based transition activation mechanism (RF-TAM) is proposed, which focuses on the pedestrian's walking direction at different moments; and 3) a pedestrian group cluster graph convolution network (GC-GCN) is designed to measure the correlation between embedded pedestrian features. Some experimental analyses and real scene experiments on datasets CUHK-SYSU, PRW, SLP, and UJS-reID indicate that the designed model can achieve a better logical topology inference with an accuracy of 87.3% and achieve the top-1 accuracy of 77.4% and the mAP accuracy of 74.3% for pedestrian re-identification. Keyang Cheng, Qing Liu 0015, Rabia Tahir, Liangmin Wang 0001, Maozhen Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Improving Interpretability by Information Bottleneck Saliency Guided Localization
Keyang Cheng, Yu Si, Liuyang Yan |
BMVC | 2 |
| 2022 | MMDV: Interpreting DNNs via Building Evaluation Metrics, Manual Manipulation and Decision VisualizationabstractThe unexplainability and untrustworthiness of deep neural networks hinder their application in various high-risk fields. The existing methods lack solid evaluation metrics, interpretable models, and controllable manual manipulation. This paper presents Manual Manipulation and Decision Visualization (MMDV) which makes Human-in-the-loop improve the interpretability of deep neural networks. The MMDV offers three unique benefits: 1) The Expert-drawn CAM (Draw CAM) is presented to manipulate the key feature map and update the convolutional layer parameters, which makes the model focus on and learn the important parts by making a mask of the input image from the CAM drawn by the expert; 2) A hierarchical learning structure with sequential decision trees is proposed to provide a decision path and give strong interpretability for the fully connected layer of DNNs; 3) A novel metric, Data-Model-Result interpretable evaluation(DMR metric), is proposed to assess the interpretability of data, model and the results. Comprehensive experiments are conducted on the pre-trained models and public datasets. The results of the DMR metric are 0.4943, 0.5280, 0.5445 and 0.5108. These data quantifications represent the interpretability of the model and results. The attention force ratio is about 6.5% higher than the state-of-the-art methods. The Average Drop rate achieves 26.2% and the Average Increase rate achieves 36.6%. We observed that MMDV is better than other explainable methods by attention force ratio under the positioning evaluation. Furthermore, the manual manipulation disturbance experiments show that MMDV correctly locates the most responsive region in the target item and explains the model's internal decision-making basis. The MMDV not only achieves easily understandable interpretability but also makes it possible for people to be in the loop. Keyang Cheng, Yu Si, Rabia Tahir |
ACM Multimedia | 1 |
| 2022 | Dual Attention-Guided Network for Anchor-Free Apple Instance Segmentation in Complex Environments
Yunshen Pei, Yi Ding 0001, Xuesen Zhu, Liuyang Yan, Keyang Cheng |
PRCV (4) | 5 |
| 2022 | Spatial-temporal correlations learning and action-background jointed attention for weakly-supervised temporal action localization
Huifen Xia 0001, Yongzhao Zhan 0001, Keyang Cheng |
Multim. Syst. | 3 |
| 2022 | Reliable Sensing Data Fusion Through Robust Multiview Prototype LearningabstractDue to emerging development of intelligent sensing technologies in Internet of Things, multisensor cooperation has been widely deployed in applications. Although multisensor information fusion can be addressed by multiview learning, its performance tends to degrade if any one sensor is disturbed with annoying noises by the environment or other factors. Therefore, fusing these cross-sensor data in a reliable and secure manner while removing those noises is crucial. Although there have been outlier-against multiview works proposed, most of them suffer from redundant parameters or performance degradation. Even worse, few of them have considered the complementary information across the sensors. In this article, we argue that in multiview information fusion, not only the clean data, but also those outliers share the same prototypes in a common space, except that the outliers are disturbed with noises. To this end, we propose a type of robust multiview prototype (RMVP) learning to fuse the sensing data while removing the noises automatically in the learning process. Specifically, in RMVP, projection matrices are designed for each sensing view to sketch the data prototypes. In addition, one auxiliary margin matrix is modeled for each sensing view to capture its data noises through penalizing a sparsity regularization on it. Afterwards, an alternating algorithm is presented to solve the proposed model. Finally, extensive experiments on intelligent sensing data sets are conducted to testify the effectiveness of the proposed method. Qing Tian 0001, Shiyu Xia, Meng Cao 0005, Keyang Cheng |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | SMA: SRv6-Based Multidomain Integrated Architecture for Industrial InternetabstractWith the increasing requirements of industrial production efficiency, the Industrial Internet has played a very important role in the fourth industrial revolution. However, the current Industrial Internet still has many drawbacks, especially in terms of network systems, such as low network expansion, inconvenient troubleshooting, and low data transmission efficiency. For this motivation, a novel SRv6-based multidomain integrated architecture (SMA) for the Industrial Internet has been proposed. Multilayer controllers are deployed in the SMA, and a software-defined network controller that generates the transmission path is replaced by SMA nodes, which realizes the high network scalability and efficient data transmission of the Industrial Internet. The faulty node in the SMA can be quickly and accurately identified through the periodic detection actively sent by the controller node in the domain and the passive feedback of the SMA nodes, and the generated SMA node trusted set (SNTS) can be used for forwarding path generation. A Bellman–Ford algorithm with a hop count constraint based on the total number of SNTS nodes is proposed, which effectively avoids long-path forwarding and improves network resource utilization. Through theoretical analysis, the safety and scalability of the SMA have been fully verified. The simulation results of the SMA on the experimental platform show that the SMA is superior to the existing Industrial Internet network structure in terms of troubleshooting efficiency of faulty nodes, network throughput, and data communication overhead. In the Industrial Internet, when the proportion of SMA nodes reaches 30%, the SMA controller can control nearly 80% of the traffic. In addition, the maximum link utilization rate will be greatly reduced, which means better adjustment of network load balance. Liangmin Wang 0001, Fan Wen, Keyang Cheng, Xia Feng, Hao Shentu |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Multi-Camera Logical Topology Inference via Conditional Probability Graph Convolution NetworkabstractIn order to improve the efficiency of pedestrian retrieval and re-identification with numerous surveillance cameras, a novel multi-camera dynamic logical topology inference method is proposed, which includes:(1) A conditional probability graph convolution network(CPG) is designed, which samples and aggregates information of multi-order neighbor nodes according to the conditional probability. The CPG is employed to aggregate the influence of all nodes relative to the target node and calculate the global correlation between each node.(2) A dynamic spatio-temporal information aggregation model(STIA) in a multi-camera system is proposed. The dynamic logical topology of the multi-camera system is inferred based on the pedestrian’s walking direction and the spatio-temporal factors.(3) A novel correlation indicator in multi-camera system is proposed. This indicator detects and quantifies temporal and causal relationships within and across camera views by a designed time-delayed Jensen-Shannon divergence(TDJS). It can be used to measure the causal correlation between two camera nodes over a long time delay. Some ablation studies and simulation experiments are performed on a dataset collected from real scenes show that our method can efficiently infer the logical topology of multiple cameras. Keyang Cheng, Qing Liu 0015, Rabia Tahir, Lubamba Kasangu Eric, Ligang He |
ICME | 1 |
| 2021 | Anti-occluded Person Re-identification via Pose Restoration and Dual Channel Feature Distance Measurement
Keyang Cheng, Chunyun Meng, Sai Liang |
PRCV (4) | 2 |
| 2021 | Nonlinear dimensionality reduction in robot vision for industrial monitoring process via deep three dimensional Spearman correlation analysis (D3D-SCA)
Keyang Cheng, Muhammad Saddam Khokhar, Misbah Ayoub, Zakria Jamali |
Multim. Tools Appl. | 1 |
| 2021 | A motor imagery EEG signal classification algorithm based on recurrence plot convolution neural network
Xianjia Meng, Shi Qiu 0002, Shaohua Wan 0001, Keyang Cheng |
Pattern Recognit. Lett. | 4 |
| 2020 | Person Search via Anchor-Free Detection and Part-Based Group Feature Similarity Estimation
Qing Liu 0015, Keyang Cheng |
PRCV (2) | 2 |
| 2020 | A Contribution Algorithm from LDRI to HDRIabstractHigh dynamic range image (HDRI) which is combined with low dynamic range image (LDRI) needs to be mapped to a low dynamic area to display. In the process of mapping, it is impossible to determine the contribution of low dynamic image sequences in the display images, so that it results in a problem that the low dynamic images cannot be accurately selected. In this paper, for the first time, a contribution algorithm from LDRI to HDRI according to the corresponding response curve of the camera is proposed. Junsong Luo, Shi Qiu 0002, Yizhang Jiang, Keyang Cheng, Huping Ye, Mingjin Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2020 | An analysis of generative adversarial networks and variants for image synthesis on MNIST dataset
Keyang Cheng, Rabia Tahir, Lubamba Kasangu Eric, Maozhen Li 0001 |
Multim. Tools Appl. | 1 |
| 2020 | Hierarchical attributes learning for pedestrian re-identification via parallel stochastic gradient descent combined with momentum correction and adaptive learning rate
Keyang Cheng, Yongzhao Zhan 0001, Maozhen Li 0001, Kenli Li 0001 |
Neural Comput. Appl. | 1 |
| 2019 | Automatic Image Annotation and Deep Learning for Tooth CT Image Segmentation
Miao Gou, Yunbo Rao, Minglu Zhang, Jianxun Sun, Keyang Cheng |
ICIG (2) | 5 |
| 2018 | A Stochastic Parallel Gradient Descent Algorithm for Person Re-identification*abstractPedestrian re-identification is a hot topic in computer vision. Convolutional neural network(CNN) has achieved good performance in pedestrian re-identification. However, CNN is computationally intensive because of vast pedestrian data and depth of CNN training. As the requirement of higher accuracy, the training always takes days and even weeks. In this paper, we propose a parallel stochastic gradient descent(SGD) algorithm, where five-hierarchy parallel structure sets up blocks based on pedestrian attributes. Moreover, the interval for updating parameters is analyzed to optimize parameter selections. Momentum-combined adaptive learning rate is also adopted. Our results show that this method successfully speeds up the training process by five times and surpasses state-of-the-art in accuracy as well. Keyang Cheng |
VCIP | 1 |
| 2018 | Data-driven pedestrian re-identification based on hierarchical semantic representationabstractSummary Limited number of labeled data of surveillance video causes the training of supervised model for pedestrian re‐identification to be a difficult task. Besides, applications of pedestrian re‐identification in pedestrian retrieving and criminal tracking are limited because of the lack of semantic representation. In this paper, a data‐driven pedestrian re‐identification model based on hierarchical semantic representation is proposed, extracting essential features with unsupervised deep learning model and enhancing the semantic representation of features with hierarchical mid‐level ‘attributes’. Firstly, CNNs, well‐trained with the training process of CAEs, is used to extract features of horizontal blocks segmented from unlabeled pedestrian images. Then, these features are input into corresponding attribute classifiers to judge whether the pedestrian has the attributes. Lastly, with a table of ‘attributes‐classes mapping relations’, final result can be calculated. Under the premise of improving the accuracy of attribute classifier, our qualitative results show its clear advantages over the CHUK02, VIPeR, and i‐LIDS data set. Our proposed method is proved to effectively solve the problem of dependency on labeled data and lack of semantic expression, and it also significantly outperforms the state‐of‐the‐art in terms of accuracy and semanteme. Keyang Cheng, Fangjie Xu, Man Qi, Maozhen Li 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | AL-DDCNN: a distributed crossing semantic gap learning for person re-identificationabstractSummary By the reason of the variability of light and pedestrians' appearance, it is hard for a camera to obtain a clear human figure. Person re‐identification with different cameras is a difficult visual recognition task. In this paper, a novel approach called attribute learning based on distributed deep convolutional neural network model is proposed to address person re‐identification task. It shows how attributes, namely the mid‐level medium between classes and features, are obtained automatically and how they are employed to re‐identify person with semantics when an author‐topic model is used to mapping category. Besides, considering the ability to operate on raw pixel input without the need to design special features, deep convolutional neural network is employed to generate features without supervision for attributes learning model. To overcome the model's weakness in computing speed, parallelized implementations such as distributed parameter manipulation and attributes learning are employed in attribute learning based on distributed deep convolutional neural network model. Experiments show that the proposed approach achieves state‐of‐the‐art recognition performance in the VIPeR data set and is with a good semantic explanation. Copyright © 2016 John Wiley & Sons, Ltd. Keyang Cheng, Yongzhao Zhan 0001, Man Qi |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Sparse representations based distributed attribute learning for person re-identification
Keyang Cheng, Kaifa Hui, Yongzhao Zhan 0001, Maozhen Li 0001 |
Multim. Tools Appl. | 1 |
| 2016 | Discriminative sparsity preserving graph embeddingabstractIn this paper, we propose a new dimensionality reduction method called discriminative sparsity preserving graph embedding (DSPGE). Unlike many existing graph embedding methods such as locality preserving projections (LPP) and sparsity preserving projections (SPP), the aim of DSPGE is to preserve the sparse reconstructive relationships of data while simultaneously capture the geometric and discriminant structure of data in the embedding space. Through the sparse reconstruction and class-specific adjacent graphs, DSPGE characterizes the intra-class and inter-class sparsity preserving scatters, seeking to achieve the optimal projections that simultaneously maximize the inter-class sparsity preserving scatter and minimize intra-class sparsity preserving scatter. The effectiveness of the proposed DSPGE is demonstrated on two popular face databases, compared to up-to-date methods. The experimental results show that DSPGE outperforms the competing methods with the satisfactory classification performance. Jianping Gou, Lan Du 0002, Keyang Cheng, Yingfeng Cai |
CEC | 3 |
| 2014 | Sparse representations based attribute learning for flower classification
Keyang Cheng, Xiaoyang Tan |
Neurocomputing | 1 |
| 2010 | A New Classifier for Facial Expression Recognition: Fuzzy Buried Markov Model
Yongzhao Zhan 0001, Keyang Cheng, Ya-Bi Chen, Chuan-Jun Wen |
J. Comput. Sci. Technol. | 2 |