Shizhuo Deng

dblp:183/2132 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-6863-8516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 BASE: A boundary-aware adaptive semantic evidential framework for open-set skeleton-based action recognition
Dongyue Chen 0001, Dingyu Xue, Shizhuo Deng, Tong Jia 0001
Expert Syst. Appl.4
2026 WeCo-OSAR: Weighted Contrastive Learning for Open-Set Skeleton-based Action Recognition with pseudo-OOD samples
Dongyue Chen 0001, Dingyu Xue, Shizhuo Deng, Tong Jia 0001
Knowl. Based Syst.4
2026 DCART: A dual contrastive alignment residual transformer model for visual grounding
Dongyue Chen 0001, Hao Wang 0073, Tong Jia 0001, Shizhuo Deng
Pattern Recognit.5
2026 Contextual Style Coherence Network for X-Ray Prohibited Item Image Synthesis
abstract
Prohibited item detection in X-Ray baggage images plays a crucial role for preventing the social security and stability. Well annotated X-Ray prohibited item training samples show necessity in achieving high detection performance for X-Ray inspection system. While collection of massive samples is extremely laborious and costly, especially for those X-Ray images, which need professional inspection machine. Synthesizing X-Ray images through Threat Image Projection (TIP) is a promising solution to overcome the data insufficient limitation in prohibited item detection. However, TIP based methods rarely consider the contextual style coherence between the foreground prohibited items and background images, resulting in generating low realistic X-Ray security images. For improving image quality and diversity, we propose a Contextual Style Coherence Network for X-Ray Prohibited item Image Synthesis. Specifically, we first propose a style fusion module to guarantee the style coherence and consistency between the foreground prohibited items and background images. We transfer the threat image projection from image space to feature space, and an affine transformation matrix is applied to uniformly sample the location, ratio and scale of the prohibited items to improve the sample diversity. We further normalize the features of the foreground prohibited item by implementing the style transfer through Gram matrix. Then, a mask partial convolution is designed for inpainting the non-object regions of the foreground prohibited items to achieve a better style transition, especially for the boundary parts. The whole network follows the adversarial training pipeline in an unsupervised manner guided by the incorporation of adversarial loss and total variation regularization. We evaluate the synthetic images generated by our method from different evaluating metrics including image quality and object detection performance on various prohibited item detection datasets. The results verify that our method can effectively generate realistic X-Ray prohibited item images and improve the detection performance.
Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng
IEEE Trans. Image Process.4
2026 Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection
abstract
Recently, with the development of the Vision-Language Model (VLM), adopting such VLM (e.g., CLIP) into object detection framework has gradually become a promising and attractive research direction, and the resulted open vocabulary object detection methods can effectively alleviate the limitations in those close-set ones, making the detectors perceive the unseen world. The core issue in open vocabulary object detection is to design an effective and efficient alignment between the visual (e.g., image) and textual (e.g., caption) features in the semantic space, so that the detectors can capture more information around the open-set scene. Current approaches deploy extra uncurated image-text pairs to pre-train a detector for obtaining a better visual-textual alignment in the feature space. Besides, knowledge distillation technology is also adopted to design an appropriate information transferring flow for aligning the visual-textual knowledge. However, large-scale image-text pairs are not always available to obtain, and the pre-training process will inevitable introduce much more computation overhead. While knowledge distillation methods focus on aligning between the local region visual feature in RoI and the textual features of VLM, neglecting the global information alignment between the image and text. For addressing the dilemmas in these alignment manners, we propose a Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection in this paper. Specifically, our proposed method integrates global image-caption and local region-prompt alignments into a unified learning paradigm. The global alignment takes the whole image and caption as the visual and textual inputs, respectively, and matches the image and caption representations from the detector and the text encoder in CLIP by contrastive learning from the overall perspective. Different from global alignment, the local one concentrates on the accordance between regions and prompts from the aspect of portion description. It extracts and aligns the embeddings for the visual patch RoIs from the image encoder in CLIP and discriminating textual token prompts from the text encoder. Moreover, we also design a prompt tuning strategy, which contains global and local components corresponding to the alignment procedure, for better adapting CLIP to downstream task object detection in a parameter-efficient learning manner. By implementation on Faster R-CNN, we conduct experiments on open vocabulary benchmarks OV-COCO and OV-LVIS, respectively. The results verify that our proposed method can achieve clear improvement over counterparts on novel categories, while performing favorably against state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Shizhuo Deng, Dongyue Chen 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Image Process.3
2025 Efficient Indoor Depth Completion Network Using Mask-adaptive Gated Convolution
abstract
Most indoor depth completion tasks rely on convolutional auto-encoders to reconstruct depth images, especially in areas with significant missing values. While traditional convolution treats valid and missing pixels equally, Partial Convolution (PConv) has mitigated this limitation. However, PConv fails to distinguish the varying degree of invalidity across different missing areas, which highlights the need for a more refined strategy. To solve this problem, we propose a novel system for indoor depth completion tasks that leverages Mask-adaptive Gated Convolution (MagaConv). MagaConv utilizes gated signals to selectively apply convolution kernels based on the characteristics of missing depth data. These gating signals are generated using shared convolution kernels that jointly process depth features and corresponding masks, ensuring coherent weight optimization. Additionally, the mask undergoes iterative updates according to predefined rules. To improve the fusion of depth and color information, we introduce a Bi-directional Aligning Projection (Bid-AP) module, which utilizes a bi-directional projection scheme with global spatial-channel attention mechanisms to filter out depth-irrelevant features from other modalities. Extensive experiments on popular benchmarks, including NYU-Depth V2, DIML, and SUN RGB-D, demonstrate that our model outperforms state-of-the-art methods in both accuracy and efficiency.
Tingxuan Huang, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001
AAAI3
2025 DUPL: Domain-agnostic Unknown-aware Prompt Learning for Threshold-free Open-set Domain Generalization
abstract
Open-set domain generalization (OSDG) aims to recognize known categories in unseen target domains without fine-tuning, while rejecting unknown categories. Existing methods show limited practicality for the requirement of extra generated or collected unknown samples, or for the assumption that the known categories appear in all source domains. Additionally, they need to determine an optimal threshold for distinguishing between known and unknown samples during testing, which is impractical in OSDG, as the target domain is unavailable during training. Besides, they are usually studied with conventional CNNs and shows unsatisfactory generalizability. To address these issues, we harness the transferable property of the pre-trained vision-language model CLIP, and propose domain-agnostic unknown-aware prompt learning (DUPL) framework to achieve threshold-free OSDG. Specifically, we train unknown tokens (UT) to enable threshold-free unknown rejection through unknown-aware prompt learning (UAPL) with only known data, and then introduce a Fourier-based data augmentation (FDA) strategy to obtain domain-agnostic prompts via domain-agnostic semantic consistency (DASC) regularization. Extensive experiments show that our method achieves state-of-the-art OSDG performance. Code is available at https://github.com/X-funbean/DUPL.
Fangbin Xu, Dongyue Chen 0001, Shizhuo Deng, Tong Jia 0001, Hao Wang 0073
ICME3
2025 Ensemble CLIPs: Effective Zero-shot Classification with Hundreds of Multi-modal CLIPs
Shizhuo Deng, Zehua Gan, Da Teng, Dongyue Chen 0001, Tong Jia 0001
ICMR2
2025 Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning
abstract
Image captioning aims to create natural language descriptions of images. Recent advancements in image captioning have explored text-only training methods that eliminate the need for image annotations. However, these methods are prone to generate descriptions that include objects that do not actually appear in the image, but are instead drawn from the retrieval texts or hard prompts-resulting in object hallucinations. To address this issue, we propose synergistic prompting mechanism called NASCap (Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning). Our method further improves the accuracy of generated captions by designing a fusion model with mask attention that isolate integrates retrieved captions with input features. The hard prompt mixed with negative entities is designed to further improve model's robust to wrong information. Additionally, we introduce a training-free multi-granularity fusion strategy that dynamically perceive and enhance salient regions into global representation. Extensive experiments demonstrate that NASCap sets a new state-of-the art cross-domain (transferable) captioning and performs Through extensive experiments, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in image captioning compared to zero-shot captioning based on text-only training.
Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng
ACM Multimedia5
2025 Weighted Evidential Continual Learning with Logits-Angle Knowledge Distillation
Dingyu Xue, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001
PRCV (1)3
2024 Self-Supervised Federated Learning for Personalized Human Activity Recognition
abstract
Personalized Human Activity Recognition (PHAR) based on wearable sensors is crucial in the medical, sports, industrial and other fields. PHAR faces challenges of privacy leakage and a shortage of labeled data. Therefore, we propose a framework called self-supervised federated learning for personalized human activity recognition (SSF-HAR) to implement private PHAR. To protect user privacy, our framework integrates federated learning (FL) to achieve the transmission of only model parameters between the cloud and clients, rather than user data. Besides, we propose a strategy of weighted aggregation to update the cloud model with the client models. To overcome the lack of labeled data, our framework introduces self-supervised learning (SSL) tasks to pretrain a feature extractor in the cloud. The proxy task of SSL transforms data and provides pseudo-labels in three forms. We test the performance on the benchmark datasets MotionSense and WIDSM. The experiments show that SSF-HAR outperforms other FL frameworks for PHAR.
Shizhuo Deng, Da Teng, Zhubao Guo, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073
ICME1
2024 SpikMamba: When SNN meets Mamba in Event-based Human Action Recognition
Yan Yang 0011, Shizhuo Deng, Da Teng, Liyuan Pan
MMAsia3
2024 APPN: An Attention-based Pseudo-label Propagation Network for few-shot learning with noisy labels
Shizhuo Deng, Da Teng, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073
Neurocomputing2
2024 A lightweight Transformer-based visual question answering network with Weight-Sharing Hybrid Attention
Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng
Neurocomputing4
2024 Ensembling disentangled domain-specific prompts for domain generalization
Fangbin Xu, Shizhuo Deng, Tong Jia 0001, Xiaosheng Yu 0001, Dongyue Chen 0001
Knowl. Based Syst.2
2024 Conjoined triple deep network for video anomaly detection
Xingya Chang, Yunhe Wu, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001
Multim. Tools Appl.3
2024 Delving Into Cluttered Prohibited Item Detection for Security Inspection System
abstract
Prohibited item detection in X-Ray baggage images can efficiently prevent the social security and stability. With the development of deep learning, applying specific methods in computer vision tasks to prohibited item detection has shown promising perspectives, and the resulted intelligent security inspection system can effectively address the limitations of human inspection. Despite several deep learning based methods have been proposed to flourish this researching field, there are still two issues which have not been fully explored. First, most of them suffer from the shortage of dataset and collection of massive and well-annotated samples is extremely laborious and costly. Second, the designation of backbone modules for current prohibited item detection methods shares similar idea with the ones in object detection. While little consideration has been paid to the properties of prohibited item, where most of them are over-lapped and cluttered. To overcome limitations, this paper proposes a cluttered prohibited item detection method for security inspection system. Specifically, our method first generates synthetic X-Ray images through cut-and-paste strategy from the training samples in each training mini-batch, where the strategy effectively and efficiently augments the dataset and quality of the synthetic samples can be guaranteed. Then, a high-order dilated convolution module is developed for enriching the representation ability of the feature, further promoting the localization ability for over-lapped and cluttered prohibited items. Experiments show that our proposed method can be well generalized to various datasets, and achieve clear improvement over state-of-the-art methods.
Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng
IEEE Trans. Ind. Informatics5
2024 LHAR: Lightweight Human Activity Recognition on Knowledge Distillation
abstract
Sensor-based Human Activity Recognition (HAR) is widely used in daily life and is the basic-level bridge to virtual healthcare in the metaverse. The current challenge is the low recognition accuracy for personalized users on smart wearable devices. The limited resource cannot support large deep learning models updated locally. Besides, integrating and transmitting sensor data to the cloud would reduce the efficiency. Considering the tradeoff between performance and complexity, we propose a Lightweight Human Activity Recognition (LHAR) framework. In LHAR, we combine the cross-people HAR task with the lightweight model task. LHAR framework is designed on the teacher-student architecture and the student network consists of multiple depthwise separable convolution layers to achieve fewer parameters. The dark knowledge distilled from the complex teacher model enhances the generalization ability of LHAR. To achieve effective knowledge distillation, we propose two optimization methods. Firstly, we train the teacher model by ensemble learning to promote teacher performance. Secondly, a multi-channel data augmentation method is proposed for the diversity of the dataset, which is a plug-in operation for the ensemble teacher model. In the experiments, we compare LHAR with state-of-art models in comparison evaluation, ablation study and the hyperparameter analysis, which proves the better performance of LHAR in efficiency and effectiveness.
Shizhuo Deng, Da Teng, Chuangui Yang, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073
IEEE J. Biomed. Health Informatics1
2024 CBDMoE: Consistent-but-Diverse Mixture of Experts for Domain Generalization
abstract
Machine learning models often suffer from severe performance degradation due to distributional shifts between testing and training data. To address this issue, researchers have focused on domain generalization (DG), which aims to generalize a model trained on source domains to arbitrary unseen target domains. Recently, ensemble learning has emerged as a popular strategy for addressing the DG problem, and domain-specific experts are typically involved. However, the existing methods do not sufficiently consider the generalizability of individual experts or leverage the consistency and diversity among them, thus limiting the generalizability of the constructed models. In this paper, we propose a consistent-but-diverse mixture of experts (CBDMoE) algorithm, which is an improved MoE framework that effectively harnesses ensemble learning for solving the DG problem. Specifically, we introduce individual expert learning (IEL), which incorporates a novel domain-class-balanced subset division (DCBSD)-based sampling strategy to facilitate a generalizable expert learning process. Additionally, we present consistent-but-diverse learning (CBDL), which employs two regularizing losses to encourage consistency and diversity in the predictions of the experts. Our proposed strategy significantly enhances the generalizability of the MoE framework. Extensive experiments conducted on three popular DG benchmark datasets demonstrate that our method outperforms the state-of-the-art approaches.
Fangbin Xu, Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng, Hao Wang 0073
IEEE Trans. Multim.4
2023 AGG-Net: Attention Guided Gated-convolutional Network for Depth Image Completion
abstract
Recently, stereo vision based on lightweight RGBD cameras has been widely used in various fields. However, limited by the imaging principles, the commonly used RGB-D cameras based on TOF, structured light, or binocular vision acquire some invalid data inevitably, such as weak reflection, boundary shadows, and artifacts, which may bring adverse impacts to the follow-up work. In this paper, we propose a new model for depth image completion based on the Attention Guided Gated-convolutional Network (AGG-Net), through which more accurate and reliable depth images can be obtained from the raw depth maps and the corresponding RGB images. Our model employs a UNet-like architecture which consists of two parallel branches of depth and color features. In the encoding stage, an Attention Guided Gated-Convolution (AG-GConv) module is proposed to realize the fusion of depth and color features at different scales, which can effectively reduce the negative impacts of invalid depth data on the reconstruction. In the decoding stage, an Attention Guided Skip Connection (AG-SC) module is presented to avoid introducing too many depth-irrelevant features to the reconstruction. The experimental results demonstrate that our method outperforms the state-of-the-art methods on the popular benchmarks NYU-Depth V2, DIML, and SUN RGB-D. https://github.com/htx0601/AGG-Net
Dongyue Chen 0001, Tingxuan Huang, Zhimin Song, Shizhuo Deng, Tong Jia 0001
ICCV4
2023 Fast Personalized Human Activity Recognition on Heuristic Parameter Estimation
abstract
Personalized human activity recognition (HAR) on wearable sensor data is crucial for healthcare, industry and sport. Personalized HAR has two challenges: very few high-quality labeled personalized data and limited terminal computing resource. Most previous domain adaptation models ignore the transfer efficiency on the terminal, which is also one criterion of Artificial Intelligence of Things. Therefore, we propose a fast transfer HAR, FastTrans, to improve the efficiency and get a trade-off with recognition effectiveness. A fusion feature extraction module is designed to learn multi-scale features on the improved hybrid loss. The proposed heuristic parameter estimation method learns the approximate solutions of the classification weights in FastTrans by scanning the adaptation data only once. Besides, an efficient time series data augmentation is proposed as a plugin for dataset variety. The results on benchmark datasets show the dramatically competitive performance of FastTrans on transferring efficiency with close accuracy to other models.
Shizhuo Deng, Chuangui Yang, Zhubao Guo, Boqian Lin, Dongyue Chen 0001, Tong Jia 0001
ICME1
2023 Prior Knowledge Guided Network for Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) involves detecting anomalous events in videos, presenting a significant and intricate task within intelligent video surveillance. Existing studies often concentrate solely on features acquired from limited normal data, disregarding the latent prior knowledge present in extensive natural image datasets. To address this constraint, we propose a Prior Knowledge Guided Network(PKG-Net) for the VAD task. First, an auto-encoder network is incorporated into a teacher-student architecture to learn two designated proxy tasks: future frame prediction and teacher network imitation, which can provide better generalization ability on unknown samples. Second, knowledge distillation on proper feature blocks is also proposed to increase the multi-scale detection ability of the model. In addition, prediction error and teacher-student feature inconsistency are combined to evaluate anomaly scores of inference samples more comprehensively. Experimental results on three public benchmarks validate the effectiveness and accuracy of our method, which surpasses recent state-of-the-arts.
Zhewen Deng, Dongyue Chen 0001, Shizhuo Deng
MMAsia3
2023 Self-relation attention networks for weakly supervised few-shot activity recognition
Shizhuo Deng, Zhubao Guo, Da Teng, Boqian Lin, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073
Knowl. Based Syst.1
2021 Reliable Recommendation with Review-level Explanations
abstract
The quality of user-generated reviews is significant for users to understand recommendation results and make online purchasing decisions correctly. However, the reliability of a review, which captures the likelihood that a review is benign, is ignored by many studies. The low reliability reviews cause a recommendation system's unsatisfying performance. Especially the fake reviews written by fraudulent users mislead the system into generating error recommendation results and explanations, which confuse customers and deprive customers of confidence in the system. In this paper, we propose a model, Reliable Recommendation with Review-level Explanations (RRRE), which detects reliable reviews and improves the performance of the explainable recommendation system as well. Recognizing the textual content of reviews, user-item interactions are valuable features for both rating prediction and reliability prediction. RRRE builds a uniform framework to predict rating scores and reliability scores simultaneously. Firstly, RRRE embeds user preferences and item profiles, which are extracted from textual and interactive features, into the representation of the review. Secondly, the supervised information of two subtasks is jointly combined. It makes the optimization of RRRE faster and better. Finally, the reviews with both high reliability scores and rating scores are given to customers as reliable explanations. To the best of our knowledge, we are the first to consider the reliability of reviews for improving explainable recommender system. And the experimental results confirm this idea and show that our model outperforms other baseline methods on Yelp and Amazon datasets.
Yanzhang Lyu, Hongzhi Yin, Jun Liu 0002, Mengyue Liu, Huan Liu 0012, Shizhuo Deng
ICDE6
2020 Few-Shot Human Activity Recognition on Noisy Wearable Sensor Data
Shizhuo Deng, Wen Hua, Guoren Wang, Xiaofang Zhou 0001
DASFAA (2)1
2020 Self-Adaptive Framework for Efficient Stream Data Classification on Storm
abstract
In this era of big data, stream data classification which is one of typical data stream applications has become more and more significant and challengeable. In these applications, it is obvious that data classification is much more frequent than model training. The ratio of stream data to be classified is rapid and time-varying, so it is an important problem to classify the stream data efficiently with high throughput. In this paper, we first analyze and categorize the current data stream machine learning algorithms according to their data structures. Then, we propose stream data classification topology (SDC-Topology) on Storm. For the classification algorithms based on the matrix, we propose self-adaptive stream data classification framework (SASDC-Framework) for efficient stream data classification on Storm. In SASDC-Framework, all the data sets arriving at the same unit time are partitioned into subsets with the nearly best partition size and processed in parallel. To select the nearly best partition size for the stream data sets efficiently, we adopt bisection method strategy and inverse distance weighted strategy. Extreme learning machine, which is a fast and accurate machine learning method based on matrix calculating, is used to test the efficiency of our proposals. According to evaluation results, the throughputs based on SASDC-Framework are 8-35 times higher than those based on SDC-Topology and the best throughput is more than 40000 prediction requests per second in our environment.
Shizhuo Deng, Shan Huang 0007, Chuncheng Yue, Jianpeng Zhou, Guoren Wang
IEEE Trans. Syst. Man Cybern. Syst.1
2017 CHAR-HMM: An Improved Continuous Human Activity Recognition Algorithm Based on Hidden Markov Model
Chuangui Yang, Shizhuo Deng, Guangxin Liu, Yuru Kang, Huichao Men
MSN4