VLDB 2026 Research / reviewers in the wild / expert
Xiao Ke
dblp:78/9040
· DBLP profile ↗
61ranked-venue papers
36as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 17 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 7 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentabstractMultimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities are frequently unavailable at the inference stage in reality. The absence of any modality often renders existing multimodal models inoperable. Furthermore, it triggers catastrophic performance degradation due to interruptions in cross-modal interactions. To address this issue, we propose a novel Missing Completion Framework with Mixture of Experts (MCMoE) that unifies unimodal and joint representation learning in single-stage training. Specifically, we propose an adaptive gated modality generator that dynamically fuses available information to reconstruct missing modalities. We then design modality experts to learn unimodal knowledge and dynamically mix the knowledge of all experts to extract cross-modal joint representations. With a mixture of experts, missing modalities are further refined and complemented. Finally, in the training phase, we mine the complete multimodal features and unimodal expert knowledge to guide modality generation and generation-based joint representation extraction. Extensive experiments demonstrate that our MCMoE achieves state-of-the-art results in both complete and incomplete multimodal learning on three public AQA benchmarks. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Rui Xu 0028, Jinglin Xu |
AAAI | 3 |
| 2026 | SFCE-Det: Sub-Feature Fusion and Cross-Layer Perceptual Enhancement DetectorabstractEdge devices face a pressing demand for low-cost object detection networks. However, because of limited computational resources, lightweight detectors often suffer significant performance degradation. In this paper, we propose SFCE-Det, an efficient object detector that achieves remarkable performance with remarkably few parameters and GFLOPs. The key contribution of our work lies in the novel subfeature fusion and cross-layer perceptual enhancement block (SFCE-Block), which effectively extracts feature information from images at a very low computational cost. SFCE-Block can be seamlessly integrated into existing convolutional neural networks and serves as a plug-and-play component for lightweight upgrades to the network. SFCE-Block can not only be used to upgrade classic models but also has excellent lightweight effects on state-of-the-art models (e.g. YoLOv8). Additionally, we propose a dynamic label assignment strategy that leverages global label correlation to further enhance the performance of SFCE-Det. Experimental results demonstrate that SFCE-Det surpasses many state-of-the-art lightweight object detectors, on multiple public datasets while maintaining an extremely low cost. For example, SFCE-Det-D2 achieves an impressive mAP of 83.4% on the PASCAL VOC dataset, comparable to YOLOv8-S. However, SFCE-Det-D2 requires only 26% of the parameters and 35% of the GFLOPs, which are 2.96M parameters and 9.9 GFLOPs, respectively. Xiao Ke, Wenyao Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | CoRe: An End-to-End Collaborative Refinement Network for Medical Image SegmentationabstractThe anatomical information obtained from medical image segmentation will provide a crucial decision-making basis for clinical diagnosis and treatment. Deep networks with encoder-decoder architecture proposed recently have achieved impressive results. However, these existing deep networks have some inherent flaws, e.g., network depth and downsampling operators jointly determine the loss of spatial detail information of deep features. We find that it is the lack of targeted solutions to these inherent flaws that make it difficult to further improve the segmentation performance. Therefore, based on these findings, we propose an end-to-end collaborative refinement method (CoRe). Specifically, we first design to generate an Error-Prone Region (EPR) by predicting uncertainty map and foreground boundary map to simulate the error region, and after locating pixels with high error proneness, we propose a feature refinement module (FRM) based on neighborhood-aware features and foreground-boundary-enhanced features to refine the upsampling features of the decoder, so as to better reconstruct the lost spatial detail information. In addition, a segmentation refinement module (SRM) is proposed to refine coarse segmentation prediction by establishing highly representative global class centers that comprehensively contain the intrinsic properties of each segmentation target. Finally, we conduct extensive experiments on five datasets with different modalities and segmentation targets. The results show that our method achieves significant improvements and competes favorably with current state-of-the-art methods. Xiao Ke, Wenzhong Guo |
IEEE J. Biomed. Health Informatics | 1 |
| 2026 | MDANet: A Lightweight Multi-Task Dynamic Adaptive Network for Real-Time Visual Perception in Autonomous Driving
Xiao Ke, Jingyi Fang, Chaoying Chen, Huanqi Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseabstractThe fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix proposes a bidirectional spatial-temporal context optical flow correction (BOFC) method. This approach leverages the consistency and complementarity of motion information between two modalities: optical flow, which excels at pixel capture, and lightweight skeleton data. It enables the extraction of pixel-level motion changes and the correction of abnormal skeleton data. Furthermore, we propose a part-level dance dataset (Dancer Parts) and part-level motion feature extraction based on task decoupling (PETD). This aims to decouple complex whole-body parts tracking into fine-grained limb-level motion extraction, enhancing the confidence of temporal information and the accuracy of correction for abnormal data. Finally, we present the DNV dataset, which simulates fully neat group dance scenes and provides reliable labels and validation methods for the newly introduced group dance neatness assessment (GDNA). To the best of our knowledge, this is the first work to develop quantitative criteria for assessing limb and joint neatness in group dance. We conduct experiments on DNV and video-based public JHMDB datasets. Our method effectively corrects abnormal skeleton points, flexibly embeds, and improves the accuracy of existing pose estimation algorithms. Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Peirong Xu, Wenzhong Guo |
AAAI | 2 |
| 2025 | Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentabstractLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large number of model parameters to learn potential associations between actions and music. To address this issue, we propose a language-guided audio-visual learning (MLAVL) framework that models "audio-action-visual" correlations guided by low-cost language modality. In our framework, multidimensional domain-based actions form action knowledge graphs, motivating audio-visual modalities to focus on task-relevant actions. We further design a shared-specific context encoder to integrate deep multimodal semantics, and an audio-visual cross-modal fusion module to evaluate action-music consistency. To match the sport’s rules, we then propose a dual-branch prompt-guided grading module to weigh both visual and audio-visual performance. Extensive experiments demonstrate that our approach achieves state-of-the-art on four public long-term sports benchmarks while maintaining low parameters.1 Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Wenzhong Guo |
CVPR | 2 |
| 2025 | Achieving Zero-Glance Unlearning with Data-Free Inversion and Selective Parameters SuppressionabstractIn machine learning (ML), data deletion involves more than just removing data from a dataset. Machine unlearning enables ML models to eliminate the effects of specific data that needs to be deleted. Under zero-glance settings, we may lack the right to utilize the data slated for removal during unlearning, thereby heightening the complexity. To address this challenge, we propose UISPS, which employs data-free inversion to generate replacement data for unavailable forgotten data. Utilizing the generated data, we propose selective parameter suppression to address the issue of catastrophic forgetting during unlearning effectively. Its interpretability improves the reliability of unlearning under zero-glance conditions. The experiments demonstrate that UISPS performs forgotten tasks with commendable results. Meanwhile, UISPS maintains higher accuracy on retained data, improving it by up to 3.34% while reducing the attack success rate of membership inference attacks by 25.85% ∼ 39.54% compared to the state-of-the-art. Puwei Lian, Xiao Ke, Zhou Tan, Ximeng Liu |
ICME | 2 |
| 2025 | Progressive Modality-Adaptive Interactive Network for Multi-Modality Image FusionabstractMulti-modality image fusion (MMIF) integrates features from distinct modalities to enhance visual quality and improve downstream task performance. However, existing methods often overlook the sparsity variations and dynamic correlations between infrared and visible images, potentially limiting the utilization of both modalities. To address these challenges, we propose the Progressive Modality-Adaptive Interactive Network (PoMAI), a novel framework that not only dynamically adapts to the sparsity and structural disparities of each modality but also enhances inter-modal correlations, thereby optimizing fusion quality. The training process consists of two stages: in the first stage, the Neighbor-Group Matching Model (NGMM) models the high sparsity of infrared features, while the Context-Aware Modeling Network (CAMN) captures rich structural details in visible features, jointly refining modality-specific characteristics for fusion. In the second stage, the Modality-Interactive Compensation Module (MICM) refines inter-modal correlations via dynamic compensation mechanism, while freezing the first-stage modules to focus MICM solely on the compensation task. Extensive experiments on benchmark datasets demonstrate that PoMAI surpasses state-of-the-art methods in fusion quality and excels in downstream tasks. Chaowei Huang, Yaru Su, Huangbiao Xu, Xiao Ke |
IJCAI | 4 |
| 2025 | Projection, Interaction and Fusion: A Progressive Difference Fusion Network for Salient Object DetectionabstractIn recent years, deep learning-based Salient Object Detection (SOD) methods have made tremendous progress; however, their performance in complex scenarios has reached a bottleneck. In this paper, we propose a novel Progressive Difference Fusion Network (PDFNet) based on fine-grained feature fusion. First, to address the scale variability of salient objects, we introduce a Self-Guided Module (SGM) with dynamic receptive fields. Second, to tackle the shape variability of salient objects, we design a Feature Aggregation Module (FAM) incorporating cross convolutions and a feedback loop. Finally, to alleviate the issue of confusion between global and detail information during multi-scale feature fusion in existing models, we develop a Progressive Difference Fusion Unit (PDFU) to project multi-scale features into fine-grained nodes and enhance them through node interaction based on difference features. Additionally, we propose a Conditional Random Field Based on Patch (CRFbp), which focuses on handling discrete points, further improving the model’s performance. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on five benchmark datasets. Code is available at: https://github.com/pdfnet2025/PDFNet.git. Xiao Ke, Yuzhen Niu |
IJCAI | 1 |
| 2025 | The Devil in the Stego Image: Far from Being Usable in Real-World ScenariosabstractDigital images, serving as the primary carrier of information, have been wildly spread on the Internet. Image steganography is a technology that employs images as the carrier for information hiding. While current deep image steganography demonstrated impressive encoding abilities across various media, two serious problems have been overlooked in deep image-to-image steganography and hinder its application under real-world scenarios, which we define as the problem of Pixel Value Overflow and Gap of Precision. In this paper, we explore the cause of those problems and introduce a plug-and-play Universal Suppressor to solve the application problems of deep image-to-image steganography in real-world scenarios, which can be flexibly applied to various models with different structures. Experiments demonstrate that our Universal Suppressor performs well in existing state-of-the-art (SOTA) models and confers them with intrinsic robustness for real-world deployment. The code will be released at https://github.com/aoli-gei/USP. Huanqi Wu 0001, Huangbiao Xu, Xiao Ke |
ACM Multimedia | 3 |
| 2025 | WSDNet: A Dual-Branch Wavelet-Structured Network for Image Forgery DetectionabstractImage Manipulation Detection and Localization (IML) is a critical task in multimedia forensics, aiming to identify whether an image has been manipulated and to accurately localize the manipulated regions. While high-frequency components often reveal local anomalies such as edge disruptions and texture inconsistencies, low-frequency components encode global structural and semantic information that is equally vital for reliable detection. However, most existing methods primarily emphasize high-frequency clues, overlooking the complementary role of low-frequency guidance in enhancing contextual understanding.To address this limitation, we propose a Wavelet-Structured Dual-branch Network (WSDNet) that explicitly leverages low-frequency information to guide high-frequency modulation. The network integrates a frequency-guided modulation branch that performs progressive wavelet decomposition across multiple scales. At each scale, the low-frequency subbands generate attention maps to modulate corresponding high-frequency components, enhancing sensitivity to structural inconsistencies and boundary artifacts.Experimental results on multiple benchmark datasets, including CASIA and Coverage, demonstrate that WSDNet achieves state-of-the-art performance, with notable improvements in both AUC and F1 metrics. The proposed low-frequency guided modulation also improves robustness under complex backgrounds and subtle manipulation patterns, validating the effectiveness of the frequencyspatial collaborative design. Xiao Ke, Song Yun |
TrustCom | 1 |
| 2025 | Modality-specific adaptive scaling and attention network for cross-modal retrieval
Xiao Ke, Baitao Chen, Yuhang Cai, Hao Liu 0076, Wenzhong Guo |
Neurocomputing | 1 |
| 2025 | Zero-shot 3D anomaly detection via online voter mechanism
Wukun Zheng, Xiao Ke, Wenzhong Guo |
Neural Networks | 2 |
| 2025 | Multi-granularity interaction and feature recombination network for fine-grained visual classification
Xiao Ke, Yuhang Cai, Baitao Chen, Hao Liu 0076, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2025 | Cross-modal independent matching network for image-text retrieval
Xiao Ke, Baitao Chen, Yuhang Cai, Hao Liu 0076, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2025 | Language-Image Consistency Augmentation and Distillation Network for visual grounding
Xiao Ke, Peirong Xu, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2025 | MSP: Multimodal Self-Attention Prompt LearningabstractMultimodal prompt learning has emerged as an effective strategy for adapting vision-language models such as CLIP to downstream tasks. However, conventional approaches typically operate at the input level, forcing learned prompts to propagate through a sequence of frozen Transformer layers. This indirect adaptation introduces cumulative geometric distortions, a limitation that we formalize as the indirect learning dilemma (ILD), leading to overfitting of the base class and reduced generalization to novel classes. To overcome this challenge, we propose the Multimodal Self-Attention Prompt (MSP) framework, which shifts adaptation into the semantic core of the model by injecting learnable prompts directly into the key and value sequences of attention blocks. This direct modulation preserves the pretrained embedding geometry while enabling more precise downstream adaptation. MSP further incorporates distance-aware optimization to maintain semantic consistency with CLIP's original representation space, and partial prompt learning via stochastic dimension masking to improve robustness and prevent over-specialization. Extensive evaluations across 11 benchmarks demonstrate the effectiveness of MSP. It achieves a state-of-the-art harmonic mean accuracy of 80.67%, with 77.32% accuracy on novel classes-representing a 2.18% absolute improvement over prior methods-while requiring only 0.11M learnable parameters. Notably, MSP surpasses CLIP's zero-shot performance on 10 out of 11 datasets, establishing a new paradigm for efficient and generalizable prompt-based adaptation. Our implementation is available at https://github.com/laixinyi023/Multimodal-Self-Attention-Prompt. Xinyi Lai, Xiao Ke, Huangbiao Xu, Shanghui Wu, Wenzhong Guo |
IEEE Trans. Image Process. | 2 |
| 2025 | Quality-Guided Vision-Language Learning for Long-Term Action Quality AssessmentabstractLong-term action quality assessment poses a challenging visual task since it requires assessing technical actions at different skill levels in a long video. Recent state-of-the-art methods incorporate additional modality information to aid in understanding action semantics, which incurs extra annotation costs and imposes higher constraints on action scenes and datasets. To address this issue, we propose a Quality-Guided Vision-Language Learning (QGVL) method to map visual features into appropriate fine-grained intervals of quality scores. Specifically, we use a set of quality-related textual prompts as quality prototypes to guide the discrimination and aggregation of specific visual actions. To avoid fuzzy rule mapping, we further propose a progressive semantic learning strategy with a Granularity-Adaptive Semantic Learning Module (GSLM) that refines accurate score intervals from coarse to fine at clip, grade, and score levels. The quality-related semantics we designed are universal to all types of action scenarios without any additional annotations. Extensive experiments show that our approach outperforms previous work by a significant margin and establishes new state-of-the-art on four public AQA benchmarks: Rhythmic Gymnastics, Fis-V, FS1000, and FineFS. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Yuezhou Li, Rui Xu 0028, Wenzhong Guo |
IEEE Trans. Multim. | 3 |
| 2024 | StegFormer: Rebuilding the Glory of Autoencoder-Based SteganographyabstractImage hiding aims to conceal one or more secret images within a cover image of the same resolution. Due to strict capacity requirements, image hiding is commonly called large-capacity steganography. In this paper, we propose StegFormer, a novel autoencoder-based image-hiding model. StegFormer can conceal one or multiple secret images within a cover image of the same resolution while preserving the high visual quality of the stego image. In addition, to mitigate the limitations of current steganographic models in real-world scenarios, we propose a normalizing training strategy and a restrict loss to improve the reliability of the steganographic models under realistic conditions. Furthermore, we propose an efficient steganographic capacity expansion method to increase the capacity of steganography and enhance the efficiency of secret communication. Through this approach, we can increase the relative payload of StegFormer to 96 bits per pixel without any training strategy modifications. Experiments demonstrate that our StegFormer outperforms existing state-of-the-art (SOTA) models. In the case of single-image steganography, there is an improvement of more than 3 dB and 5 dB in PSNR for secret/recovery image pairs and cover/stego image pairs. Xiao Ke, Huanqi Wu 0001, Wenzhong Guo |
AAAI | 1 |
| 2024 | Surpassing Immediate Spatio-temporal Metaphors: The Enduring Impact of Language and Visuospatial Experience on Temporal Cognition in Native and Near-Native Mandarin Speakers
Chunying Wu, Lingyue Kong, Xiao Ke |
CogSci | 3 |
| 2024 | Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment
Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu 0028, Huanqi Wu 0001, Wenzhong Guo |
ECCV (42) | 2 |
| 2024 | Cutransnet: Transformers to Make Strong Encoders for Multi-Task Vision Perception of Autonomous DrivingabstractIn autonomous driving, perception plays a critical role as it serves as a fundamental requirement for both planning and control. Currently, most perception tasks are processed independently, which requires designing multiple models and networks to handle multiple tasks. This division leads to multiple sub-tasks in real-world environments, making it difficult to ensure real-time performance and communication between tasks. The paper introduces an innovative neural network architecture termed CUTransNet, which presents a unified approach capable of simultaneously detecting drivable areas, lane markings, and traffic objects, achieving multi-task processing with a single model. To build a more robust feature encoder, we proposed the CUT module that combines the global context ability of Transformers. The module is integrated into the convolutional neural network’s backbone to compensate for the low-level visual clues lost by Transformers and achieve higher detection accuracy. Experimental results on BDD100K demonstrate that CUT model surpasses traditional multi-task networks in both task accuracy and computational efficiency while maintaining high real-time performance. Xiao Ke, JinCheng Wan, Guozhen Tan |
ICASSP | 2 |
| 2024 | IF-Font: Ideographic Description Sequence-Following Font GenerationabstractFew-shot font generation (FFG) aims to learn the target style from a limited number of reference glyphs and generate the remaining glyphs in the target font. Previous works focus on disentangling the content and style features of glyphs, combining the content features of the source glyph with the style features of the reference glyph to generate new glyphs. However, the disentanglement is challenging due to the complexity of glyphs, often resulting in glyphs that are influenced by the style of the source glyph and prone to artifacts. We propose IF-Font, a novel paradigm which incorporates Ideographic Description Sequence (IDS) instead of the source glyph to control the semantics of generated glyphs. To achieve this, we quantize the reference glyphs into tokens, and model the token distribution of target glyphs using corresponding IDS and reference tokens. The proposed method excels in synthesizing glyphs with neat and correct strokes, and enables the creation of new glyphs based on provided IDS. Extensive experiments demonstrate that our method greatly outperforms state-of-the-art methods in both one-shot and few-shot settings, particularly when the target styles differ significantly from the training font styles. The code is available at [https://github.com/Stareven233/IF-Font](https://github.com/Stareven233/IF-Font). Xinping Chen, Xiao Ke, Wenzhong Guo |
NeurIPS | 2 |
| 2024 | Two-path target-aware contrastive regression for action quality assessment
Xiao Ke, Huangbiao Xu, Wenzhong Guo |
Inf. Sci. | 1 |
| 2024 | No-reference stereoscopic image quality assessment based on binocular collaboration
Hanling Wang, Xiao Ke, Wenzhong Guo, Wukun Zheng |
Neural Networks | 2 |
| 2024 | Text-based person search via cross-modal alignment learning
Xiao Ke, Hao Liu 0076, Peirong Xu, Xinru Lin, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2024 | U-Transformer-based multi-levels refinement for weakly supervised action segmentation
Xiao Ke, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2024 | GFENet: Generalization Feature Extraction Network for Few-Shot Object DetectionabstractFew-shot object detection achieves rapid detection of novel-class objects by training detectors with a minimal number of novel-class annotated instances. Transfer learning-based few-shot object detection methods have shown better performance compared to other methods such as meta-learning. However, when training with base-class data, the model may gradually bias towards learning the characteristics of each category in the base-class data, which could result in a decrease in learning ability during fine-tuning on novel classes, and further overfitting due to data scarcity. In this paper, we first find that the generalization performance of the base-class model has a significant impact on novel class detection performance and proposes a generalization feature extraction network framework to address this issue. This framework perturbs the base model during training to encourage it to learn generalization features and solves the impact of changes in object shape and size on overall detection performance, improving the generalization performance of the base model. Additionally, we propose a feature-level data augmentation method based on self-distillation to further enhance the overall generalization ability of the model. Our method achieves state-of-the-art results on both the COCO and PASCAL VOC datasets, with a 6.94% improvement on the PASCAL VOC 10-shot dataset. Xiao Ke, Qiuqin Chen, Hao Liu 0076, Wenzhong Guo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Object detection based on knowledge graph network
Guozhen Tan, Xiao Ke, Huaiwei Si, Yanfei Peng |
Appl. Intell. | 3 |
| 2023 | S-pad: self-learning padding mechanism
Xiao Ke, Xinru Lin, Wenzhong Guo |
Mach. Vis. Appl. | 1 |
| 2023 | HPFace: a high speed and accuracy face detector
Xiao Ke, Wenzhong Guo |
Neural Comput. Appl. | 1 |
| 2023 | Granularity-aware distillation and structure modeling region proposal network for fine-grained image classification
Xiao Ke, Yuhang Cai, Baitao Chen, Hao Liu 0076, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2023 | An Ultra-Fast Automatic License Plate Recognition Approach for Unconstrained ScenariosabstractRecently, with the development of deep learning, Automatic License Plate Recognition (ALPR) has made great progress, However, there are still many challenges to accomplish license plate (LP) recognition under various traffic scenarios. One of them is the detection speed and recognition speed, and the other is the difficulty to recognize the low resolution and highly tilted LP images. In this paper, we present a two-stage ALPR framework to achieve efficient LP detection and recognition in unconstrained scenarios. Our LP detector is based on improved Yolov3-tiny, and we also propose a lightweight recognition network MRNet based on multi-scale features. In order to improve the inference speed, we abandoned the rectification of LP images, and RNNs that are difficult to compute in parallel. Additionally, we also propose a license plate data augmentation method, which achieves more effective augmentation and improves the generalization ability of the network through secondary random and hyperparametric search. As an additional contribution, we provide a challenging dataset collected from real-world driving recorders. The dataset is for multiple LPs in a single image, which makes up for the lack of public datasets in multiple LP scenarios. We evaluated the results on five datasets and showed that we achieved the best performance on almost all test sets, achieving 99.8% accuracy on CCPD with more than 180,000 license plate test sets. In terms of speed, the inference speed for detecting license plates reaches 751 FPS, and the fastest inference time for recognizing a single license plate takes only 2.9 ms. Xiao Ke, Ganxiong Zeng, Wenzhong Guo |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Progressive Multitask Learning Network for Online Chinese Signature Segmentation and Recognition
Xunhui Qin, Hanyue Zhang, Xiao Ke, Zhonghao Shen, Songmao Qi |
ICFHR | 3 |
| 2022 | Sar Ship Detection Based on Swin Transformer and Feature Enhancement Feature Pyramid NetworkabstractWith the booming of Convolutional Neural Networks (CNNs), CNNs such as VGG-16 and ResNet-50 widely serve as backbone in SAR ship detection. However, CNN based backbone is hard to model long-range dependencies, and causes the lack of enough high-quality semantic infor-mation in feature maps of shallow layers, which leads to poor detection performance in complicated background and small-sized ships cases. To address these problems, we pro-pose a SAR ship detection method based on Swin Trans-former and Feature Enhancement Feature Pyramid Network (FEFPN). Swin Transformer serves as backbone to model long-range dependencies and generates hierarchical features maps. FEFPN is proposed to further improve the quality of feature maps by gradually enhencing the semantic infor-mation of feature maps at all levels, especially feature maps in shallow layers. Experiments conducted on SAR ship de-tection dataset (SSDD) reveal the advantage of our pro-posed methods. Xiao Ke, Xiaoling Zhang 0002, Tianwen Zhang, Jun Shi 0002, Shunjun Wei |
IGARSS | 1 |
| 2022 | A deep learning based bank card detection and recognition method in complex scenes
Hanyang Lin, Xiao Ke |
Appl. Intell. | 4 |
| 2022 | Weakly supervised fine-grained image classification via two-level attention activation model
Xiao Ke, Yanyan Huang, Wenzhong Guo |
Comput. Vis. Image Underst. | 1 |
| 2022 | LocalFace: Learning significant local features for deep face recognition
Xiao Ke, BingHui Lin, Wenzhong Guo |
Image Vis. Comput. | 1 |
| 2022 | Learning deep convolutional descriptor aggregation for efficient visual tracking
Xiao Ke, Yuezhou Li, Wenzhong Guo, Yanyan Huang |
Neural Comput. Appl. | 1 |
| 2022 | 100+ FPS detector of personal protective equipment for worker safety: A deep learning approach for green edge computing
Xiao Ke, Wenyao Chen, Wenzhong Guo |
Peer-to-Peer Netw. Appl. | 1 |
| 2022 | Joint Sample Enhancement and Instance-Sensitive Feature Learning for Efficient Person SearchabstractPerson search, consisting of jointly or separately trained person detection stage and person Re-ID stage, suffers from significant challenges such as inefficiency and difficulty in acquiring discriminative features. However, certain work has either turned to the end-to-end framework whose performance is limited by task conflicts or has consistently attempted to obtain more accurate bounding boxes (Bboxes). Few studies have focused on the impact of sample-specificity in person search datasets for training a fine-grain Re-ID model, and few have considered obtaining discriminative Re-ID features from Bboxes in a more efficient way. In this paper, a novel sample-enhanced and instance-sensitive (SEIE) framework is designed to boost performance. By analyzing the structure of person search framework, our method refines the two stages separately. For the detection stage, we re-design the usage of Bbox and a sample enhancement combination is proposed to further enhance the quality and quantity of Bboxes. SEC can suppress false positive detection results and randomly generate high-quality positive samples. For the Re-ID stage, we contribute an instance similarity loss to exploit the similarity between classless instances, and an Omni-scale Re-ID backbone is employed to learn more discriminative features. We obtain a more efficient and discriminative person search framework by concatenating the two stages. Extensive experiments demonstrate that our method achieves state-of-the-art performance with a high speed, and significantly outperforms other existing methods. Xiao Ke, Hao Liu 0076, Wenzhong Guo, Baitao Chen, Yuhang Cai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | HOG-ShipCLSNet: A Novel Deep Learning Network With HOG Feature Fusion for SAR Ship ClassificationabstractShip classification in synthetic aperture radar (SAR) images is a fundamental and significant step in ocean surveillance. Recently, with the rise of deep learning (DL), modern abstract features from convolutional neural networks (CNNs) have hugely improved SAR ship classification accuracy. However, most existing CNN-based SAR ship classifiers overly rely on abstract features, but uncritically abandon traditional mature hand-crafted features, which may incur some challenges for further improving accuracy. Hence, this article proposes a novel DL network with histogram of oriented gradient (HOG) feature fusion (HOG-ShipCLSNet) for preferable SAR ship classification. In HOG-ShipCLSNet, four mechanisms are proposed to ensure superior classification accuracy, that is, 1) a multiscale classification mechanism (MS-CLS-Mechanism); 2) a global self-attention mechanism (GS-ATT-Mechanism); 3) a fully connected balance mechanism (FC-BAL-Mechanism); and 4) an HOG feature fusion mechanism (HOG-FF-Mechanism). We perform sufficient ablation studies to confirm the effectiveness of these four mechanisms. Finally, our experimental results on two open SAR ship datasets (OpenSARShip and FUSAR-Ship) jointly reveal that HOG-ShipCLSNet dramatically outperforms both modern CNN-based methods and traditional hand-crafted feature methods. Tianwen Zhang, Xiaoling Zhang 0002, Xiao Ke, Xiaowo Xu, Xu Zhan, Chen Wang 0041, Yue Zhou 0005, Dece Pan, Jun Shi 0002, Shunjun Wei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SAR Ship Detection Based on an Improved Faster R-CNN Using Deformable ConvolutionabstractWith the rise of Deep Learning (DL), numerous DL-based SAR ship detectors, represented by Faster R-CNN, is constantly breaking the record of detection accuracy. However, these detectors still face huge challenges in modeling the geometric transformation of shape-changeable ships, due to their used conventional convolution kernels whose structure is fixed. Therefore, to address this problem, we propose an improved Faster R-CNN by using deformable convolution kernels for SAR ship detection. We substitute some conventional shape-changeless convolution kernels in Faster R-CNN with deformable convolution ones that can adaptively learn additional 2-D offsets of the raw convolution kernels, to better model the geometric transformation of shape-changeable ships. Finally, the experimental results on the open SAR Ship Detection Dataset (SSDD) reveal that our improved Faster R-CNN achieves a 2.02% mean Average Precision (mAP) improvement than the raw Faster R-CNN. Xiao Ke, Xiaoling Zhang 0002, Tianwen Zhang, Jun Shi 0002, Shunjun Wei |
IGARSS | 1 |
| 2021 | U-FPNDet: A one-shot traffic object detector based on U-shaped feature pyramid moduleabstractAbstract In the field of automatic driving, identifying vehicles and pedestrians is the starting point of other automatic driving techniques. Using the information collected by the camera to detect traffic targets is particularly important. The main bottleneck of traffic object detection is due to the same category of targets, which may have different scales. For example, the pixel‐level of cars may range from 30 to 300 px, which will cause instability of positioning and classification. In this paper, a multi‐dimension feature pyramid is constructed in order to solve the multi‐scale problem. The feature pyramid is built by developing a U‐shaped module and using a cascade‐method. In order to verify the effectiveness of the U‐shaped module, we also designed a new one‐shot detector U‐FPNDet. The model first extracts the basic feature map by using the basic network and constructs the multi‐dimension feature pyramid. Next, a pyramid pooling module is used to get more context information from the scene. Finally, the detection network is run on each level of the pyramid to obtain the final result by NMS. By using this method, a state‐of‐the‐art performance is achieved on both detection and classification on commonly used benchmarks. Xiao Ke |
IET Image Process. | 1 |
| 2021 | Lightweight convolutional neural network-based pedestrian detection and re-identification in multiple scenarios
Xiao Ke, Xinru Lin, Liyun Qin |
Mach. Vis. Appl. | 1 |
| 2021 | Template Enhancement and Mask Generation for Siamese TrackingabstractSiamese tracking methods have become the focus of visual tracking in recent years. Advanced Siamese trackers perform well on certain benchmarks, but there are still some limitations. First, most Siamese trackers adopt the initial frame as a single template, which leads to underfitting and reduces the ability to predict instances. Second, mainstream trackers report a rectangular bounding box as a prediction, resulting in poor accuracy of non-rigid objects. Therefore, we propose the template enhancement and mask generation for Siamese tracking. Given that the essence of Siamese trackers is instance learning, we propose constructing an alternative template explicitly to address the underfitting of the instance space. Moreover, in order to improve the tracking accuracy, we obtain the descriptor aggregation to transform the semantic segmentation outputs for mask prediction. Finally, we propose the SiamEM through the fusion of the above approaches. Comprehensive experiments show that template enhancement and mask generation significantly improve Siamese trackers on benchmarks. Xiao Ke, Yuezhou Li, Yu Ye 0004, Wenzhong Guo |
IEEE Signal Process. Lett. | 1 |
| 2020 | Fine-grained vehicle type detection and recognition based on dense attention network
Xiao Ke |
Neurocomputing | 1 |
| 2019 | Dense small face detection based on regional cascade multi-scale methodabstractIn the field of object detection, the research on the problem of detecting small face is the most extensive, but when there are objects with obvious scale differences in the image, the detection performance is not obvious, which is due to the scale invariance properties of the deep convolutional neural networks. Although in recent years, there have been some methods proposed to solve this problem such as FPN and SNIP, which is based on feature pyramid. However, they have not fundamentally solved the problem. A regional cascade multi‐scale detection method has been proposed. First, a global detector and several local detectors have been trained, respectively. The global detector is trained by the original training set, while the local detector is trained by the sub‐training set generated by the original training set. Second, the global detector can detect object roughly and the local detectors can produce more detailed results that improve the performance of global detector. Finally, to integrate the detection results of global detector and local detectors as the output, non‐maximum suppression methods are used. The method can be carried in any depth model of object detection, has good scalability, and is more suitable for dense face detection. Xiao Ke, Wenzhong Guo |
IET Image Process. | 1 |
| 2019 | A Robust Moving Object Detection in Multi-Scenario Big Data for Video SurveillanceabstractAdvanced wireless imaging sensors and cloud data storage contribute to video surveillance by enabling the generation of large amounts of video footage every second. Consequently, surveillance videos have become one of the largest sources of unstructured data. Because multi-scenario surveillance videos are often continuously produced, using these videos to detect moving objects is challenging for the conventional moving object detection methods. This paper presents a novel model that harnesses both sparsity and low-rankness with contextual regularization to detect moving objects in multi-scenario surveillance data. In the proposed model, we consider moving objects as a contiguous outlier detection problem through the use of low-rank constraint with contextual regularization, and we construct dedicated backgrounds for multiple scenarios using dictionary learning-based sparse representation, which ensures that our model can be effectively applied to multi-scenario videos. Quantitative and qualitative assessments indicate that the proposed model outperforms existing methods and achieves substantially more robust performance than the other state-of-the-art methods. Xiao Ke |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Multi-Dimensional Traffic Congestion Detection Based on Fusion of Visual Features and Convolutional Neural NetworkabstractIn intelligent transportation systems, there are many tasks that rely on the detection of road congestion, such as traffic signal scheduling and traffic accident detection. As traditional methods for traffic congestion detection are difficult to use, expensive, and may cause damage to the road surface, this paper presents a method for road congestion detection that is based on multidimensional visual features and a convolutional neural network (CNN). This method first detects the density of foreground objects by using a gray-level co-occurrence matrix; second, the speed of moving objects is detected by using the Lucas-Kanade optical flow with pyramid implementation. Third, a Gaussian mixture model is used to model the background, and the CNN is then used to accurately detect the final foreground from the candidate foregrounds. Finally, the proposed method performs road congestion detection in terms of a multidimensional feature space, including traffic density, traffic velocity, road occupancy, and traffic flow. Furthermore, we propose an information entropy method using a histogram of optical flow to enhance the accuracy and reliability of road congestion detection. Simulation results via quantitative and qualitative assessment indicate that the proposed method is able to significantly outperform the state-of-the-art road-traffic congestion detection methods due to the fusion of multidimensional features using the CNN. Xiao Ke, Wenzhong Guo, Dewang Chen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | End-to-End Automatic Image Annotation Based on Deep CNN and Multi-Label Data AugmentationabstractAutomatic image annotation is a key step in image retrieval and image understanding. In this paper, we present an end-to-end automatic image annotation method based on a deep convolutional neural network (CNN) and multi-label data augmentation. Different from traditional annotation models that usually perform feature extraction and annotation as two independent tasks, we propose an end-to-end automatic image annotation model based on deep CNN (E2E-DCNN). E2E-DCNN transforms the image annotation problem into a multi-label learning problem. It uses a deep CNN structure to carry out the adaptive feature learning before constructing the end-to-end annotation structure using multiple cross-entropy loss functions for training. It is difficult to train a deep CNN model using small-scale datasets or scale up multi-label datasets using traditional data augmentation methods; hence, we propose a multi-label data augmentation method based on Wasserstein generative adversarial networks (ML-WGAN). The ML-WGAN generator can approximate the data distribution of a single multi-label image. The images generated by ML-WGAN can assist in the reduction of the over-fitting problem of training a deep CNN model and enhance the generalization ability of the trained CNN model. We optimize the network structure by using deformable convolution and spatial pyramid pooling. We experiment the proposed E2E-DCNN model with data augmentation by the proposed ML-WGAN on several public datasets. The experimental results demonstrate that the proposed model outperforms the state-of-the-art automatic image annotation models. Xiao Ke, Jiawei Zou, Yuzhen Niu |
IEEE Trans. Multim. | 1 |
| 2018 | Stereoscopic Image Quality Assessment Based on both Distortion and DisparityabstractUnderstanding the characteristics of high-quality stereoscopic 3D (S3D) images has great significance for S3D image classification, quality assessment, and quality enhancement. Existing works assess the quality of an S3D image from a single perspective and use databases with subjective opinion scores obtained by conducting subjective experiments in labs. So the performance of these works in real applications is unclear and questionable. In this paper, we propose an S3D image quality assessment index based on two important factors, namely DisTortion and DisParity (DTDP). Distortion reflects the extent of distortion of each view of an S3D image, respectively. Disparity is a distinguishing factor for an S3D image as compared with a monocular image. We design some features to represent these two factors and use a random forest regression (RFR) to learn the mapping between the features and subjective opinion scores. The database we choose is the NVIDIA 3D VISION LIVE Highest Rated database which is from real application. The images and the subjective opinion scores are directly come from the website viewers all over the world. Experimental results demonstrate a superior performance of the proposed DTDP index as compared to the existing methods. Yuzhen Niu, Yini Zhong, Xiao Ke, Yiqing Shi |
VCIP | 3 |
| 2018 | CF-based optimisation for saliency detectionabstractIn view of the observation that saliency maps generated by saliency detection algorithms usually show similarity imperfection against the ground truth, the authors propose an optimisation algorithm based on clustering and fitting (CF) for saliency detection. The algorithm uses a fitting model to represent the quantitative relationship between ground truth and algorithm‐generated saliency maps. The authors use the K ‐means method to cluster the images into k clusters according to the similarities among images. Image similarity is measured in terms of scene and colour by using the GIST and colour histogram features, after which the fitting model for each cluster is calculated. The saliency map of a new image is optimised by using one of the fitting models which correspond to the cluster to which the image belongs. Experimental results show that their CF‐based optimisation algorithm improves the performance of various single image saliency detection algorithms. Moreover, the improvement achieved by their algorithm when using both CF strategies is greater than the improvement achieved by the same algorithm when not using the clustering strategy. In addition, their proposed optimisation algorithm can also effectively optimise co‐saliency detection algorithms which already consider multiple similar images simultaneously to improve saliency of single images. Yuzhen Niu, Wenqi Lin, Xiao Ke |
IET Comput. Vis. | 3 |
| 2018 | Multi-modality weakly labeled sentiment learning based on Explicit Emotion Signal for Chinese microblog
Dazhen Lin, Donglin Cao, Yanping Lv, Xiao Ke |
Neurocomputing | 5 |
| 2017 | Fitting-based optimisation for image visual salient object detectionabstractTo overcome some major problems with traditional saliency evaluation metrics, full‐reference image quality assessment (IQA) metrics, which have similar but stricter objectives, are used. Inspired by the root mean absolute error, the authors propose a fitting‐based optimisation method for salient object detection algorithms. Their algorithm analyses the quantitative relationship between saliency and ground truth values, and uses the derived relationship to fit the saliency values to the original saliency maps. This ensures that the resulting images, which are composed of fitted values, are closer to the ground truth. The proposed algorithm first computes the statistics of the ground truth and saliency maps computed by each salient object detection algorithm. These statistics are used to compute the parameters of four fitting models, which generally agree with the characteristics of the statistical data. For a new saliency map, they use the fitting model with the computed parameters to obtain the fitted saliency values, which are confined to the range [0, 255]. Finally, they evaluate their saliency optimisation algorithm using traditional evaluation metrics, IQA metrics, and a content‐based image retrieval application. The results show that the proposed approach improves the quality of the optimised saliency maps. Yuzhen Niu, Wenqi Lin, Xiao Ke, Lingling Ke |
IET Comput. Vis. | 3 |
| 2017 | Construction of uniform designs via an adjusted threshold accepting algorithm
Kai-Tai Fang, Xiao Ke, Ahmed M. Elsawah |
J. Complex. | 2 |
| 2017 | Data equilibrium based automatic image annotation by fusing deep model and semantic propagation
Xiao Ke, Ming-Ke Zhou, Yuzhen Niu, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2017 | Sparse Representation-Based Semi-Supervised Regression for People CountingabstractLabel imbalance and the insufficiency of labeled training samples are major obstacles in most methods for counting people in images or videos. In this work, a sparse representation-based semi-supervised regression method is proposed to count people in images with limited data. The basic idea is to predict the unlabeled training data, select reliable samples to expand the labeled training set, and retrain the regression model. In the algorithm, the initial regression model, which is learned from the labeled training data, is used to predict the number of people in the unlabeled training dataset. Then, the unlabeled training samples are regarded as an over-complete dictionary. Each feature of the labeled training data can be expressed as a sparse linear approximation of the unlabeled data. In turn, the labels of the labeled training data can be estimated based on a sparse reconstruction in feature space. The label confidence in labeling an unlabeled sample is estimated by calculating the reconstruction error. The training set is updated by selecting unlabeled samples with minimal reconstruction errors, and the regression model is retrained on the new training set. A co-training style method is applied during the training process. The experimental results demonstrate that the proposed method has a low mean square error and mean absolute error compared with those of state-of-the-art people-counting benchmarks. Hongbo Zhang 0002, Bineng Zhong 0001, Jixiang Du, Jialin Peng, Duansheng Chen, Xiao Ke |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2016 | Multi-scale salient region and relevant visual keywords based model for automatic image annotation
Xiao Ke, Wenzhong Guo |
Multim. Tools Appl. | 1 |
| 2015 | Two- and three-level lower bounds for mixture L2-discrepancy and construction of uniform designs by threshold accepting
Xiao Ke, Hua-Jun Ye |
J. Complex. | 1 |
| 2012 | A two-level model for automatic image annotation
Xiao Ke, Shaozi Li, Donglin Cao |
Multim. Tools Appl. | 1 |