Shunzhi Zhu

dblp:35/952 · DBLP profile ↗
← Back
78ranked-venue papers
5as first author
54since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 2Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Diverse feature generation for zero-shot Chinese character recognition
Song-Liang Pan, Kunchi Li, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
Expert Syst. Appl.6
2026 MoKA-HP: Motion-aware KAdaptation with historical prompts for efficient and robust RGB-T tracking
Zhixi Wu, Si Chen 0002, Dahan Wang, Shunzhi Zhu
Neurocomputing4
2026 Learning relationship-guided vision-language transformer for facial attribute recognition
Si Chen 0002, Mingxuan Lei, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu
Pattern Recognit.6
2026 HOH-Net: High-Order Hierarchical Middle-Feature Learning Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match images of the same person across visible (VIS) and infrared (IR) modalities. Existing VI-ReID methods ignore high-order structure information of features and struggle to learn a reliable common feature space due to the modality discrepancy between VIS and IR images. To alleviate the above issues, we propose a novel high-order hierarchical middle-feature learning network (HOH-Net) for VI-ReID. We introduce a high-order structure learning (HSL) module to explore the high-order relationships of short- and long-range feature nodes, for significantly mitigating model collapse and effectively obtaining discriminative features. We further develop a fine-coarse graph attention alignment (FCGA) module, which efficiently aligns multi-modality feature nodes from node-level and region-level perspectives, ensuring reliable middle-feature representations. Moreover, we exploit a hierarchical middle-feature agent learning (HMAL) loss to hierarchically reduce the modality discrepancy at each stage of the network by using the agents of middle features. The proposed HMAL loss also exchanges detailed and semantic information between low- and high-stage networks. Finally, we introduce a modality-range identity-center contrastive (MRIC) loss to minimize the distances between VIS, IR, and middle features. Extensive experiments demonstrate that the proposed HOH-Net yields state-of-the-art performance on the image-based and video-based VI-ReID datasets. The code is available at: https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Amurep: Adaptively Multi-layer Region Partitioning for Next POI Recommendation
abstract
In the secnario of next POI (point of interest) recommendation, interactions between users and POIs are important to represent users’ requirements, so graph neural networks(GNNs) are often used to model User-POI bipartite graphs based on their interactions. Due to the large number of users and POIs, the sparsity of corresponding interaction graph is very high. Adding intermediate layers for aggregating several POIs into a POI group is considered to improve recommendation performance. Existing methods often use static methods based on priori knowledge, such as geographic proximity of POIs. However, those methods are not adaptable for dynamical requirements of users. In this paper, a method Amurep(Adaptively Multi-layer Region Partitioning) is presented, for adaptively aggregating POIs into multi-layer POI groups to improves recommendation performance. Specifically, POI groups in multi-scales are learned to constrain the range of recommended next POI. In addition, for better capturing users’ dynamic requirements, high-order information, such as multi-hop relations of POIs within user check-in trajectories, are used. In the end, extensive experiments are conducted on two real-world datasets to show that Amurep outperforms the state-of-the-art methods to the best of our knowledge and the learned POI groups are also interpretable.
Weiqiang Jiang, Shunzhi Zhu
IJCNN4
2025 Dual-Branch Residual Wavelet Attention Network for Colorectal Cancer Magnifying Endoscopy Image Classification
Linrui He, Yun Wu 0001, Dahan Wang, Shunzhi Zhu, Xuyao Zhang
PRCV (14)5
2025 Simultaneously local and global contrastive learning of graph representations
Binsheng Hong, Zhaori Guo, Shunzhi Zhu, Kaibiao Lin, Fan Yang 0010
Eng. Appl. Artif. Intell.4
2025 Low-light image enhancement with quality-oriented pseudo labels via semi-supervised contrastive learning
Nanfeng Jiang, Yiwen Cao, Xu-Yao Zhang, Dahan Wang, Chiming Wang, Shunzhi Zhu
Expert Syst. Appl.7
2025 AMST: Object tracking based on collaborative framework with adaptive multi-strategy
Rui Xu 0028, Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
Inf. Sci.5
2025 ADR-Net: Attention-oriented detail recovery network for document image shadow removal
Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Yun Wu 0001, Shunzhi Zhu
Knowl. Based Syst.6
2025 Joint radical embedding and detection for zero-shot Chinese character recognition
Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
Pattern Recognit.5
2025 Stain-adaptive self-supervised learning for histopathology image analysis
Haili Ye, Shunzhi Zhu, Dahan Wang, Xu-Yao Zhang, Heguang Huang
Pattern Recognit.3
2025 Toward High-Quality Spatiotemporal Recommendation: Trajectory Recovery Based on Spatial and Temporal Dependencies
abstract
The rapid advancement of location and information technologies has generated a significant volume of human mobility data, which has been extensively utilized in spatiotemporal recommendation systems, including personalized point-of-interest recommendation, route recommendation, and location-aware event recommendation. Achieving high-quality recommendation results necessitates excellent quality of input trajectory data. However, trajectories obtained from GPS-enabled devices often contain missing and erroneous data that is unevenly distributed over time and highly sparse, which significantly hampers the effectiveness spatiotemporal data analytics. Therefore, trajectory recovery plays an important role in spatiotemporal recommendation systems. The objective of trajectory recovery is to utilize historical trajectories to restore missing locations, providing high-quality data for spatiotemporal recommendation systems. The development of an effective trajectory recovery mechanism faces three major challenges: 1) Complex and multi-granularity transition patterns among different locations; 2) Difficulty in discovering spatio-temporal dependencies; and 3) Data sparsity and noise. To address these challenges, we propose an attentional model with spatio-temporal recurrent neural networks, ARMove, to recover human mobility from long and sparse trajectories. In ARMove, we first design a spatio-temporal weighted recurrent neural network to capture users' long-term preferences. Next, we introduce a multi-granularity trajectory encoder to model complex transition patterns and multi-level periodicity of human mobility. An attention-based history aggregation module is proposed to leverage historical mobility information. Extensive evaluation results reveal that our model outperforms the state-of-the-art models, demonstrating its ability to reconstruct high-quality and fine-grained human mobility trajectories.
Chenhao Wang 0007, Shunzhi Zhu, Lisi Chen 0001
IEEE Trans. Big Data4
2025 DHLA: Dynamic Hybrid Label Assignment for End-to-End Object Detection
abstract
The recent one-to-one label assignment plays a crucial role in removing the last non-differentiable component, i.e., Non-Maximum Suppression (NMS), used in the post-processing step of the one-to-many label assignment, thus building an efficient end-to-end detection system. However, due to the limited number of foreground samples, the one-to-one label assignment often suffers from insufficient representation learning, and its performance is inferior to that of traditional detectors trained using the one-to-many label assignment. To solve these problems, we introduce a novel Dynamic Hybrid Label Assignment (DHLA) method, including a Hybrid Sample Selection (HSS) strategy and a Stage-aware Soft-label Adjustment (SSA) mechanism. In order to enhance the ability of representation learning of the one-to-one label assignment, the HSS strategy subtly integrates the one-to-many and the one-to-one label assignment rules to form a simple and effective hybrid assignment rule, where high-quality samples are selected for training according to an effective task consistency metric. Moreover, the SSA mechanism dynamically adjusts the contributions of different foreground samples at different training stages, thus effectively achieving the transition from one-to-many to one-to-one label assignment. In addition, we leverage a ranking loss function to widen the score gaps between the highest scoring position and surrounding areas for effectively removing duplicate bounding boxes. As a result, our method not only learns robust feature representations during training but also performs efficient end-to-end detection during inference. Extensive experiments demonstrate our method achieves competitive performance compared to state-of-the-art detectors on the challenging COCO and CrowdHuman datasets.
Zhi-Liang Hu, Si Chen 0002, Yang Hua 0001, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual Tracking
abstract
In recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods.
Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.6
2024 High-Order Structure Based Middle-Feature Learning for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same persons captured by visible (VIS) and infrared (IR) cameras. Existing VI-ReID methods ignore high-order structure information of features while being relatively difficult to learn a reasonable common feature space due to the large modality discrepancy between VIS and IR images. To address the above problems, we propose a novel high-order structure based middle-feature learning network (HOS-Net) for effective VI-ReID. Specifically, we first leverage a short- and long-range feature extraction (SLE) module to effectively exploit both short-range and long-range features. Then, we propose a high-order structure learning (HSL) module to successfully model the high-order relationship across different local features of each person image based on a whitened hypergraph network. This greatly alleviates model collapse and enhances feature representations. Finally, we develop a common feature space learning (CFL) module to learn a discriminative and reasonable common feature space based on middle features generated by aligning features from different modalities and ranges. In particular, a modality-range identity-center contrastive (MRIC) loss is proposed to reduce the distances between the VIS, IR, and middle features, smoothing the training process. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that our HOS-Net achieves superior state-of-the-art performance. Our code is available at https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Yan Yan 0001, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu
AAAI6
2024 Learning Explicit Radical Representations for Zero-Shot Chinese Character Recognition
Song-Liang Pan, Dahan Wang, Nanfeng Jiang, Xu-Yao Zhang, Shunzhi Zhu
ICPR (31)5
2024 DocHFormer: Document Image Dewarping via Harmonized Modeling of Hierarchical Priors
Xinyue Zhou, Guanting Li, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
ICPR (31)6
2024 RCFormer : Interactive Image Segmentation via Reconstructing Click Vision Transformers
abstract
Click-based interactive image segmentation intends to segment an object from the background under user click guidance. Recently, Vision Transformer has made significant strides in interactive image segmentation. However, the previous studies 1) overlook the importance of different clicks in terms of their contribution to the segmentation results; and 2) suffer from inconsistency across different feature scales in the multi-scale structure. In this paper, we propose a new interactive segmentation framework, named RCFormer, with two novel components: reconstruct click patch embedding (RCPE) for encoding the importance of clicks, and multi-scale adaptive fusion (MSAF) for the adaptive fusion of feature maps across different scales. RCPE enhances the effectiveness of click interactions by spatially distinguishing the importance of clicks. MSAF adaptively fuses useful spatial information and filters the redundant feature at multi-scales. The experiments on several benchmarks show that our proposed approach achieves state-of-the-art performance. Notably, our method achieves 2.31 NoC@90 on the Berkeley dataset, improving by 8.6% over the previous best results.
PanPan Chen, Dahan Wang, Yun Wu 0001, Xu-Yao Zhang, Shunzhi Zhu
IJCNN7
2024 Enhancing Lightweight Remote Sensing Semantic Segmentation via Weak Consistency Regularization
abstract
Remote sensing image semantic segmentation has widespread applications in urban planning and land monitoring. In recent years, U-Net and its variant networks have almost dominated the research in the field of semantic segmentation. However, many models pay less attention to computational efficiency, rendering them ineffective in scenarios with computational resource and timeliness constraints, such as autonomous driving and disaster monitoring. To address this issue, we propose the USA-Net (UNet-like with Shifted Axial), a lightweight hybrid model based on convolution and MLP (Multi-Layer Perceptron). Specifically, we design the ST Block (Shift Tokenized Block), which introduces local features into global operations in MLP through spatial shift, and then use ELCM (Efficient Large-kernel Convolution Module) to enlarge the model’s receptive field and learn the shape features of objects. Additionally, we propose a new semi-supervised learning framework to further improve the model’s generalization performance. On the ISPRS Vaihingen and ISPRS Potsdam datasets, USA-Net significantly outperforms most state-of-the-art methods in terms of segmentation accuracy and efficiency.
Zirong Chen, Shunxin Xiao, Wang Man, Dahan Wang, Shunzhi Zhu
IJCNN6
2024 Character Relationship Refinement Network for Handwritten Mathematical Expression Recognition
abstract
Most current Handwritten Mathematical Expression Recognition (HMER) methods employ an attention-based encoder-decoder framework, which generates LaTeX sequences from the given images, following the paradigm of predicting "one-by-one". However, this paradigm may have some challenges: 1) without considering the connectivity between characters, the prior information in the prediction process will be ignored inadvertently, especially implicit information, such as " " and " ˆ ". 2) Some characters of high similarities, such as "6/b" and "o/O", will have negative effects on prediction results. To solve these issues, we propose a simple but effective Character Relationship Refinement Network (CRRN), which consists of Joint Character Learning (JCL) and Character Refinement Mask (CRM). Specifically, JCL calculates the relationship probability between characters and uses them to improve prediction accuracy. CRM takes the character confidence coefficient in a coarse-to-fine way that can reassign the weights of all characters to improve model discriminability on easily confused characters. With the collaboration of both modules, our proposed CRRN can outperform the state-of-the-art on popular datasets.
LiWei Jiang, Nanfeng Jiang, Yun Wu 0001, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
IJCNN6
2024 Bridge the Gap of Semantic Context: A Boundary-Guided Context Fusion UNet for Medical Image Segmentation
Dahan Wang, Shunzhi Zhu
PRCV (15)5
2024 VATBoost-Net: Integrating Enhanced Feature Perturbation and Detail Enhancement for Medical Image Segmentation
Baichen Liu, Shunzhi Zhu
PRCV (14)3
2024 FedAFR: Enhancing Federated Learning with adaptive feature reconstruction
abstract
Federated learning is a distributed machine learning method where clients train models on local data to ensure that data will not be transmitted to a central server, providing unique advantages in privacy protection. However, in real-world scenarios, data between different clients may be non-Independently and Identically Distributed (non-IID) and imbalanced, leading to discrepancies among local models and impacting the efficacy of global model aggregation. To tackle this issue, this paper proposes a novel framework, FedARF, designed to improve Federated Learning performance by adaptively reconstructing local features during training. FedARF offers a simple reconstruction module for aligning feature representations from various clients, thereby enhancing the generalization capability of cross-client aggregated models. Additionally, to better adapt the model to each client’s data distribution , FedARF employs an adaptive feature fusion strategy for a more effective blending of global and local model information, augmenting the model’s accuracy and generalization performance . Experimental results demonstrate that our proposed Federated Learning method significantly outperforms existing methods in variety image classification tasks, achieving faster model convergence and superior performance when dealing with non-IID data distributions.
Youxin Huang, Shunzhi Zhu, Zhicai Huang
Comput. Commun.2
2024 An IoT and blockchain based logistics application of UAV
Chin-Ling Chen, Yong-Yuan Deng, Shunzhi Zhu, Woei-Jiunn Tsaur, Wei Weng 0002
Multim. Tools Appl.3
2024 GCAT: graph calibration attention transformer for robust object tracking
Si Chen 0002, Xinxin Hu, Dahan Wang, Yan Yan 0001, Shunzhi Zhu
Neural Comput. Appl.5
2024 HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011.
Si Chen 0002, Hui Da, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu
IEEE Trans. Circuits Syst. Video Technol.6
2024 Developing Deep LSTMs With Later Temporal Attention for Predicting COVID-19 Severity, Clinical Outcome, and Antibody Level by Screening Serological Indicators Over Time
abstract
OBJECTIVE: The clinical course of COVID-19, as well as the immunological reaction, is notable for its extreme variability. Identifying the main associated factors might help understand the disease progression and physiological status of COVID-19 patients. The dynamic changes of the antibody against Spike protein are crucial for understanding the immune response. This work explores a temporal attention (TA) mechanism of deep learning to predict COVID-19 disease severity, clinical outcomes, and Spike antibody levels by screening serological indicators over time. METHODS: We use feature selection techniques to filter feature subsets that are highly correlated with the target. The specific deep Long Short-Term Memory (LSTM) models are employed to capture the dynamic changes of disease severity, clinical outcome, and Spike antibody level. We also propose deep LSTMs with a TA mechanism to emphasize the later blood test records because later records often attract more attention from doctors. RESULTS: Risk factors highly correlated with COVID-19 are revealed. LSTM achieves the highest classification accuracy for disease severity prediction. Temporal Attention Long Short-Term Memory (TA-LSTM) achieves the best performance for clinical outcome prediction. For Spike antibody level prediction, LSTM achieves the best permanence. CONCLUSION: The experimental results demonstrate the effectiveness of the proposed models. The proposed models can provide a computer-aided medical diagnostics system by simply using time series of serological indicators.
Yang Li 0168, Baichen Liu, Zhixi Wu, Shengjun Zhu, Qiliang Chen, Hongyan Hou, Zhibin Guo, Hewei Jiang, Shujuan Guo, Feng Wang 0061, Shengjing Huang, Shunzhi Zhu, Xionglin Fan, Sheng-ce Tao
IEEE J. Biomed. Health Informatics14
2024 Multi-Branch Enhanced Discriminative Network for Vehicle Re-Identification
abstract
Vehicle re-identification (ReID) is the task of identifying the same vehicle across numerous cameras. This is a complex classification task, and the fine-grained information and strong discrimination features have proven to be effective in handling the re-identification classification task. However, most existing methods focuses on extracting local area features in combination with global features, while exploring subtle distinguishing features, which is a difficult task, remains an open problem and unsolved. In this paper, we propose a multi-branch enhanced discriminative network (MED) to better extract subtle distinguishing features that have high discriminative power to improve the ReID performance. In the proposed MED method, each feature map obtained by convolutional neural network (CNN) is divided into 4 spatial sub-maps, on each of which, the vertical and the horizontal branches are used to extract the subtle distinguishing features intrinsically contained in sub-areas. The vertical and the horizontal branches are combined with the global branch to perform the ReID task. Moreover, our proposed method is capable of extracting rich fine-grained features without the need of extra manual annotation while maintaining a simple design structure. We conducted extensive experiments on the vehicle ReID datasets (VehicleID and VeRi-776), showing that the proposed MED method outperforms most existing methods. Further, we directly apply the MED method to the pedestrian ReID problem on the Market-1501, DUKEMTMC, and MSMT17 datasets, achieving the state-of-the-art (SOTA) performance as well. This demonstrates that the proposed method has good generality and can be flexibly applied to the ReID tasks.
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.4
2023 A Novel Clustering Model TEC for Station Classification
Shunzhi Zhu
AINA (1)3
2023 Multi-Zone Transformer Based on Self-Distillation for Facial Attribute Recognition
abstract
Recently, transformers have shown great promising performance in various computer vision tasks. However, the current transformer based methods ignore the information exchanges between transformer blocks, and they have not been applied in the facial attribute recognition task. In this paper, we propose a multi-zone transformer based on self-distillation for FAR, termed MZTS, to predict the facial attributes. A multi-zone transformer encoder is firstly presented to achieve the interactions of the different transformer encoder blocks, thus avoiding forgetting the effective information between the transformer encoder block groups during the iteration process. Furthermore, we introduce a new self-distillation mechanism based on class tokens, which distills the class tokens obtained from the last transformer encoder block group to the other shallow groups by interacting with the significant information between the different transformer blocks through attention. Extensive experiments on the challenging CelebA and LFWA datasets have demonstrated the excellent performance of the proposed method for FAR.
Si Chen 0002, Xueyan Zhu, Dahan Wang, Shunzhi Zhu, Yun Wu 0001
FG4
2023 A Shallow Graph Neural Network with Innovative Node Updating for Online Handwritten Stroke Classification
Yan-Rong Wang, Dahan Wang, Xiao-Long Yun, Shunzhi Zhu
ICDAR (4)6
2023 UAM-Net: An Attention-Based Multi-level Feature Fusion UNet for Remote Sensing Image Segmentation
Yiwen Cao, Nanfeng Jiang, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
PRCV (4)5
2023 MCKIE: Multi-class Key Information Extraction from Complex Documents Based on Graph Convolutional Network
Zhicai Huang, Shunxin Xiao, Dahan Wang, Shunzhi Zhu
PRCV (7)4
2023 Pseudo Labels Refinement with Stable Cluster Reconstruction for Unsupervised Re-identification
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu, Dewu Ge
PRCV (4)6
2023 GridIIS: Grid Based Interactive Image Segmentation
Pengqi Zhu, Dahan Wang, Shunzhi Zhu
PRCV (11)3
2023 Self-supervised contrastive representation learning for large-scale trajectories
Shuzhe Li, Bingqi Yan, Shunzhi Zhu, Yanwei Yu
Future Gener. Comput. Syst.5
2023 Continuous trajectory similarity search with result diversification
Shunzhi Zhu, Yongjun Ren
Future Gener. Comput. Syst.2
2023 SiamCCF: Siamese visual tracking via cross-layer calibration fusion
abstract
Abstract Siamese networks have attracted wide attention in visual tracking due to their competitive accuracy and speed. However, the existing Siamese trackers usually leverage a fixed linear aggregation of feature maps, which does not effectively fuse the different layers of features with attention. Besides, most of Siamese trackers calculate the similarity between the template and the search region through a cross‐correlation operation between the features of the last blocks from the two branches, which might introduce the redundant noise information. In order to solve these problems, this study proposes a novel Siamese visual tracking method via cross‐layer calibration fusion, termed SiamCCF. An attention‐based feature fusion module is employed using local attention and non‐local attention to fuse the features from the deep and shallow layers, so as to capture both local details and high‐level semantic information. Moreover, a cross‐layer calibration module can use the fused features to calibrate the features of the last network blocks and build the cross‐layer long‐range spatial and inter‐channel dependencies around each spatial location. Extensive experiments demonstrate that the proposed method has achieved competitive tracking performance compared with state‐of‐the‐art trackers on challenging benchmarks, including OTB100, OTB2013, UAV123, UAV20L, and LaSOT.
Si Chen 0002, Shunzhi Zhu, Huarong Xu, Dahan Wang
IET Comput. Vis.3
2023 Learning an attention-aware parallel sharing network for facial attribute recognition
Si Chen 0002, Xinyu Lai, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
J. Vis. Commun. Image Represent.5
2023 MTNet: Mutual tri-training network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Liuxiang Qiu, Zimin Tian, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
J. Vis. Commun. Image Represent.6
2023 Self-information of radicals: A new clue for zero-shot Chinese character recognition
Dahan Wang, Xia Du, Huayi Yin, Xu-Yao Zhang, Shunzhi Zhu
Pattern Recognit.6
2023 Identity-Aware Contrastive Knowledge Distillation for Facial Attribute Recognition
abstract
Facial attribute recognition (FAR) is an important and yet challenging multi-label learning task in computer vision. Existing FAR methods have achieved promising performance with the development of deep learning. However, they usually suffer from prohibitive computational and memory costs. In this paper, we propose an identity-aware contrastive knowledge distillation method, termed ICKD, to compress the FAR model. A nonlinear weight-sharing mapping (NWSM) mechanism is firstly designed to avoid the difficulty of directly matching features of the teacher and student networks due to the lower representation ability of the student network. Furthermore, an identity-aware contrastive distillation (ICD) loss is employed to guide the student network to effectively learn the mutual relations between samples with multiple attributes. In addition, an adjustable ladder distillation (ALD) loss is developed to automatically adjust the importance of different distillation points with the progress of training. Extensive experiments demonstrate that our method can significantly improve the performance of student networks and outperforms the existing FAR methods on the public challenging datasets.
Si Chen 0002, Xueyan Zhu, Yan Yan 0001, Shunzhi Zhu, Shaozi Li, Dahan Wang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Continuous spatial keyword search with query result diversifications
Shunzhi Zhu
World Wide Web (WWW)3
2022 Graph Attention Transformer Network for Robust Visual Tracking
Si Chen 0002, Dahan Wang, Shunzhi Zhu
ICONIP (4)5
2022 Critical Radical Analysis Network for Chinese Character Recognition
abstract
Zero-shot learning is a challenging problem in many tasks due to the lack of training samples of the unseen classes. The radical-based zero-shot Chinese character recognition methods treat Chinese characters as a combination of radicals and structures, and recognize Chinese characters by identifying the radicals and structures contained in them. Current approaches generally treat the contribution of all radicals to Chinese character recognition as the same, and the recognition results rely on the network’s ability to recognize radicals and their corresponding position information, ignoring the potential value of radicals themselves in eliminating the uncertainty of Chinese characters. In this paper, we model the problem of radical-based Chinese character recognition as an uncertainty elimination problem and propose a Critical Radical Analysis Network (CRAN) to explore the Ideographic Description Sequence (IDS) information for zero-shot Chinese character recognition. Specifically, we propose a novel method to compute the critical values of radicals based on information theory using the predefined Chinese character IDS dictionary. In recognition, we use an iterative approach to translate the predicted radical sequence to target Chinese characters. That is, the radicals of the predicted sequence are sorted in descending order of the critical value, and then the radicals are continuously selected in this order as the information obtained to eliminate the uncertainty of the Chinese character until the character is recognized. We conduct experiments on the CTW, CASIA-AHCDB, and CASIA-HWDB datasets. The experimental results show that the proposed method improves the ability of recognizing unseen Chinese characters, demonstrating the effectiveness of the proposed method.
Huayi Yin, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
ICPR5
2022 Exploiting Robust Memory Features for Unsupervised Reidentification
Jiawei Lian, Dahan Wang, Xia Du, Yun Wu 0001, Shunzhi Zhu
PRCV (2)5
2022 Semantic-Aware Non-local Network for Handwritten Mathematical Expression Recognition
Xiang-Hao Liu, Dahan Wang, Xia Du, Shunzhi Zhu
PRCV (3)4
2022 Joint Pixel-Level and Feature-Level Unsupervised Domain Adaptation for Surveillance Face Recognition
Huangkai Zhu, Huayi Yin, Du Xia, Dahan Wang, Xianghao Liu, Shunzhi Zhu
PRCV (3)6
2022 Fuzzy granular convolutional classifiers
Yumin Chen 0002, Shunzhi Zhu, Wei Li 0069, Nan Qin
Fuzzy Sets Syst.2
2022 Learning meta-adversarial features via multi-stage adaptation network for robust visual object tracking
Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
Neurocomputing6
2021 LIDUSA - A Learned Index Structure for Dynamical Uneven Spatial Data
Zejian Zhang, Shunzhi Zhu
ICA3PP (3)3
2021 Detecting urban hot regions by using massive geo-tagged image data
Dahan Wang, Shunzhi Zhu
Neurocomputing3
2021 Few-labeled visual recognition for self-driving using multi-view visual-semantic representation
Dahan Wang, Shunzhi Zhu
Neurocomputing3
2020 An efficient algorithm for spatio-textual location matching
Jianping Zeng 0004, Shunzhi Zhu
Distributed Parallel Databases4
2020 Privacy-preserving spatial keyword location-to-trajectory matching
Jianping Zeng 0004, Wenxing Hong, Shunzhi Zhu
Distributed Parallel Databases4
2020 Multisource Aggregation Search and Scheduling for Remote Sensing Data Cluster
abstract
Multisource aggregation (MSA) is an important function in the remote sensing data processing. In this letter, we propose and study a novel MSA search and scheduling problem to improve the performance of remote sensing data cluster. Given a static spatial network, a set of query points $Q$ , and a set of data locations $O$ (candidates for cluster centers), the MSA function retrieves the data location with the minimum aggregation distance (the sum of the distances to all query points). We believe that such function plays an important role in remote sensing data cluster and classification. The MSA problem faces two challenges: 1) how to prune the search space effectively and retrieve the results of MSA in real time and 2) how to schedule multiple query sources during search processing. To overcome these challenges, we make the following contributions. First, upper and lower bounds on aggregation distance are defined to prune the search space effectively. Second, each query source is given a priority label, and a best-first scheduling strategy is developed to further enhance the query performance. Finally, we conduct extensive experiments on real and synthetic data sets to verify the high performance of the developed algorithms.
Hao Wang 0013, Shunzhi Zhu
IEEE Geosci. Remote. Sens. Lett.2
2020 Inferring region significance by using multi-source spatial data
Shunzhi Zhu, Dahan Wang, Danhuai Guo
Neural Comput. Appl.1
2019 Collaborative Cross-Domain k NN Search for Remote Sensing Image Processing
abstract
kNN search is a fundamental function in image processing, which is useful in many real applications, including image cluster, image classification, and image understanding and analysis in general. In this light, we propose and study a novel collaborative cross-domain kNN search (CD-kNN) in multidomain space. Given a query location q in a multidomain space (e.g., spatial domain, temporal domain, textual domain, and so on), the CD-kNN finds top-k data points with the minimum distance to q. This problem is challenging due to two reasons. First, how to define practical distance measures to evaluate the distance in multidomain space. Second, how to prune the search space efficiently in multiple domains. To address the challenges, we define a linear combination method-based distance measure for multidomain space. Based on the distance measure, a collaborative search method is developed to constrain the CD search space in a comparable smaller range. A pair of upper and lower bounds is defined to prune the search space in multiple domains effectively. Finally, we conduct extensive experiments to verify that the developed methods can achieve a high performance.
Wei Weng 0002, Shunzhi Zhu
IEEE Geosci. Remote. Sens. Lett.4
2019 Discovery of accessible locations using region-based geo-social data
Shunzhi Zhu, Danhuai Guo, Shuo Shang
World Wide Web4
2018 Aggregate location recommendation in dynamic transportation networks
Danhuai Guo, Shunzhi Zhu
World Wide Web5
2017 Embedding Factorization Models for Jointly Recommending Items and User Generated Lists
abstract
Existing recommender algorithms mainly focused on recommending individual items by utilizing user-item interactions. However, little attention has been paid to recommend user generated lists (e.g., playlists and booklists). On one hand, user generated lists contain rich signal about item co-occurrence, as items within a list are usually gathered based on a specific theme. On the other hand, a user's preference over a list also indicate her preference over items within the list. We believe that 1) if the rich relevance signal within user generated lists can be properly leveraged, an enhanced recommendation for individual items can be provided, and 2) if user-item and user-list interactions are properly utilized, and the relationship between a list and its contained items is discovered, the performance of user-item and user-list recommendations can be mutually reinforced.
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Shunzhi Zhu, Tat-Seng Chua
SIGIR5
2017 Location-Based Top-k Term Querying over Sliding Window
Lisi Chen 0001, Bin Yao 0002, Shuo Shang, Shunzhi Zhu, Kai Zheng 0001
WISE (1)5
2017 Probabilistic routing using multimodal data
Shunzhi Zhu, Shuo Shang, Jiye Wang
Neurocomputing1
2017 Image feature detection algorithm based on the spread of Hessian source
Shunzhi Zhu, Lizhao Liu, Si Chen 0002
Multim. Syst.1
2017 Discovery of probabilistic nearest neighbors in traffic-aware spatial networks
Shuo Shang, Shunzhi Zhu, Danhuai Guo, Minhua Lu
World Wide Web2
2016 Probabilistic Nearest Neighbor Query in Traffic-Aware Spatial Networks
Shuo Shang, Zhewei Wei, Ji-Rong Wen, Shunzhi Zhu
APWeb (1)4
2016 Discriminative local collaborative representation for online object tracking
Si Chen 0002, Shaozi Li, Rongrong Ji, Yan Yan 0001, Shunzhi Zhu
Knowl. Based Syst.5
2016 Robust visual tracking via online semi-supervised co-boosting
Si Chen 0002, Shunzhi Zhu, Yan Yan 0001
Multim. Syst.2
2015 Anti-counterfeiting digital watermarking algorithm for printed QR barcode
Rongsheng Xie, Shunzhi Zhu, Dapeng Tao
Neurocomputing3
2015 k-jump: A strategy to design publicly-known algorithms for privacy preserving micro-data disclosure
abstract
Abstract Data owners are expected to disclose micro-data for research, analysis, and various other purposes. In disclosing micro-data with sensitive attributes, the goal is usually two fold. First, the data utility of disclosed data should be maximized for analysis purposes. Second, the private information contained in such data must be to an acceptable level. Typically, a disclosure algorithm evaluates potential generalization functions in a predetermined order, and then discloses the first generalization that satisfies the desired privacy property. Recent studies show that adversarial inferences using knowledge about such disclosure algorithms can usually render the algorithm unsafe. In this paper, we show that an existing unsafe algorithm can be transformed into a large family of safe algorithms, namely, k-jump algorithms. We then prove that the data utility of different k-jump algorithms is generally incomparable. The comparison of data utility is independent of utility measures and syntactic privacy models. Finally, we analyze the computational complexity of k-jump algorithms, and confirm the necessity of safe algorithms even when a secret choice is made among algorithms.
Wen Ming Liu, Lingyu Wang 0001, Lei Zhang 0004, Shunzhi Zhu
J. Comput. Secur.4
2014 Combining the requirement information for software defect estimation in design time
Shunzhi Zhu, Ke Qin, Guangchun Luo
Inf. Process. Lett.2
2014 PPTP: Privacy-Preserving Traffic Padding in Web-Based Applications
abstract
Web-based applications are gaining popularity as they require less client-side resources, and are easier to deliver and maintain. On the other hand, web applications also pose new security and privacy challenges. In particular, recent research revealed that many high profile web applications might cause sensitive user inputs to be leaked from encrypted traffic due to side-channel attacks exploiting unique patterns in packet sizes and timing. Moreover, existing solutions, such as random padding and packet-size rounding, were shown to incur prohibitive overhead while still failing to guarantee sufficient privacy protection. In this paper, we first observe an interesting similarity between this privacy-preserving traffic padding (PPTP) issue and another well studied problem, privacy-preserving data publishing (PPDP). Based on such a similarity, we present a formal PPTP model encompassing the privacy requirements, padding costs, and padding methods. We then formulate PPTP problems under different application scenarios, analyze their complexity, and design efficient heuristic algorithms. Finally, we confirm the effectiveness and efficiency of our algorithms by comparing them to existing solutions through experiments using real-world web applications.
Wen Ming Liu, Lingyu Wang 0001, Pengsu Cheng, Kui Ren 0001, Shunzhi Zhu, Mourad Debbabi
IEEE Trans. Dependable Secur. Comput.5
2013 Searching similar segments over textual event sequences
abstract
Sequential data is prevalent in many scientific and commercial applications such as bioinformatics, system security and networking. Similarity search has been widely studied for symbolic and time series data in which each data object is a symbol or numeric value. Textual event sequences are sequences of events, where each object is a message describing an event. For example, system logs are typical textual event sequences and each event is a textual message recording internal system operations, statuses, configuration modifications or execution errors. Similar segments of an event sequence reveals similar system behaviors in the past which are helpful for system administrators to diagnose system problems. Existing search indexing for textual data only focus on unordered data. Substring matching methods are able to efficiently find matched segments over a sequence, however, their sequences are single values rather than texts. In this paper, we propose a method, suffix matrix, for efficiently searching similar segments over textual event sequences. It provides an integration of two disparate techniques: locality-sensitive hashing and suffix arrays. This method also supports the k-dissimilar segment search. A k-dissimilar segment is a segment that has at most k dissimilar events to the query sequence. By using random sequence mask proposed in this paper, this method can have a high probability to reach all k-dissimilar segments without increasing much search cost. We conduct experiments on real system log data and the experimental results show that our proposed method outperforms alternative methods using existing techniques.
Tao Li 0001, Shu-Ching Chen, Shunzhi Zhu
CIKM4
2013 Finding multiple global linear correlations in sparse and noisy data sets
Shunzhi Zhu, Tao Li 0001
Knowl. Based Syst.1
2011 Personalized News Recommendation: A Review and an Experimental Investigation
Lei Li 0001, Dingding Wang 0001, Shunzhi Zhu, Tao Li 0001
J. Comput. Sci. Technol.3
2010 Data clustering with size constraints
Shunzhi Zhu, Dingding Wang 0001, Tao Li 0001
Knowl. Based Syst.1
2006 Scheduling of Re-entrant Lines with Neuro-Dynamic Programming Based on a New Evaluating Criterion
Huiyu Jin, Shunzhi Zhu, Maoqing Li
ISNN (2)3