VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Shi 0002
dblp:240/7215-2
· DBLP profile ↗
34ranked-venue papers
5as first author
28since 2021 · last 2025
0000-0001-7858-5369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 21 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 10 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person RetrievalabstractThe aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Models (LVLMs), the advanced capabilities of contemporary models in image understanding have led to the generation of highly accurate captions. Therefore, this paper explores the potential of employing Large Vision-Language Models for unsupervised text-based pedestrian image retrieval and proposes a Multi-grained Uncertainty Modeling and Alignment framework (MUMA). Initially, multiple Large Vision-Language Models are employed to generate diverse and hierarchically structured pedestrian descriptions across different styles and granularities. However, the generated captions inevitably introduce noise. To address this issue, an uncertainty-guided sample filtration module is proposed to estimate and filter out unreliable image-text pairs. Additionally, to simulate the diversity of styles and granularities in captions, a multi-grained uncertainty modeling approach is applied to model the distributions of captions, with each caption represented as a multivariate Gaussian distribution. Finally, a multi-level consistency distillation loss is employed to integrate and align the multi-grained captions, aiming to transfer knowledge across different granularities. Experimental evaluations conducted on three widely-used datasets demonstrate the significant advancements achieved by our approach. Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Shijuan Huang, Linnan Tu, Fei Shen 0004 |
AAAI | 3 |
| 2025 | TSAD: Temporal-spatial association differences-based unsupervised anomaly detection for multivariate time-series
Hanbing Zhu, Zongyi Li, Yuxuan Shi 0002, Chuang Zhao 0001, Hongxu Ji, Ping Li 0021 |
Neurocomputing | 5 |
| 2024 | Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature ExplorationabstractPerson search is a unified system aimed at jointly localizing and identifying a person of interest from a gallery of whole scene images. Due to the inherent properties of the person search, it faces significant challenges of large-scale variations, inaccurate detection boxes, and crowded scenes. To address these issues, we proposed an uncertainty-guided framework coupled with auxiliary shallow feature exploration, which includes a shallow feature fusion module and an uncertainty-guided module. Firstly, considering the scales of the person are varied due to various scenes and their relative positions to the camera, a shallow feature fusion module is designed to extract multi-scale features to assist the re-id sub-task. Additionally, a self-distillation loss is proposed to align features across different scales. Furthermore, to alleviate the problem that the model can be easily affected by coarse samples resulting from crowded scenes and inaccurate detection boxes, we introduce an uncertainty guidance module to reduce the negative impact of these coarse targets. The experimental results demonstrate the effectiveness of our proposed methods on two benchmarks (i.e., CUHK-SYSU, and PRW). Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Ping Li 0021 |
ICASSP | 3 |
| 2024 | Improving Visual Quality and Transferability of Adversarial Attacks on Face Recognition Simultaneously with Adversarial RestorationabstractAdversarial face examples possess two critical properties: Visual Quality and Transferability. However, existing approaches rarely address these properties simultaneously, leading to subpar results. To address this issue, we propose a novel adversarial attack technique known as Adversarial Restoration (AdvRestore), which enhances both visual quality and transferability of adversarial face examples by leveraging a face restoration prior. In our approach, we initially train a Restoration Latent Diffusion Model (RLDM) designed for face restoration. Subsequently, we employ the inference process of RLDM to generate adversarial face examples. The adversarial perturbations are applied to the intermediate features of RLDM. Additionally, by treating RLDM face restoration as a sibling task, the transferability of the generated adversarial face examples is further improved. Our experimental results validate the effectiveness of the proposed attack method. Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021 |
ICASSP | 3 |
| 2024 | Cross-modal Generation and Alignment via Attribute-guided Prompt for Unsupervised Text-based Person Retrieval
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Shijuan Huang |
IJCAI | 3 |
| 2024 | Importance-Aware Spatial-Temporal representation Learning for Gait RecognitionabstractAs an important biometric technology, gait recognition methods are used to identify individuals based on their unique walking patterns. Empirically, different local regions of a pedestrian contribute variably to identification due to variations in walking conditions. Besides, the appearance of silhouette frames tends to be imperfect and variable, as segmentation errors frequently occur during pedestrian segmentation and pedestrians present different postures in walking. From a perspective of intuition, it is pivotal to measure the spatial importance of each local body region and the temporal importance of each frame during the gait representation learning. Based on this analysis, we heuristically propose Importance-Aware Spatial-Temporal representation Learning (IAGait), a novel gait recognition framework which utilize spatial and temporal importance to aggregate discriminative gait representations, assessing the contributions of local spatial regions and different temporal frames. IAGait is composed of two main modules: the Spatial-Importance-Aware Module (SIAM) which focuses on learning the importance of spatial horizontal parts, and the Temporal-Importance-Aware Module (TIAM) which learns temporal importance using the self-attention mechanism. In detail, SIAM introduces a weak supervision strategy for the spatial importance learning and Pseudo-Label Assignment (PLA) for part importance with metric learning hardness. And TIAM adopts a View-Information-Erasing Module (VIEM) to eliminate view information and obtain view-invariant representations, since the view variation can disrupt temporal importance awareness. Extensive experiments demonstrate the effectiveness of the proposed modules, with our method achieving state-of-the-art performance on popular datasets. We plan to release the source code for further research and development. Bohao Wei, Yuxuan Shi 0002 |
IJCNN | 3 |
| 2024 | DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions LocalizationabstractEmbedding invisible hyperlinks or hidden codes in images to replace QR codes has become a hot topic recently. This technology requires first localizing the embedded region in the captured photos before decoding. Existing methods that train models to find the invisible embedded region struggle to obtain accurate localization results, leading to degraded decoding accuracy. This limitation is primarily because the CNN network is sensitive to low-frequency signals, while the embedded signal is typically in the high-frequency form. Based on this, this paper proposes a Dual-Branch Dual-Head (DBDH) neural network tailored for the precise localization of invisible embedded regions. Specifically, DBDH uses a low-level texture branch containing 62 high-pass filters to capture the high-frequency signals induced by embedding. A high-level context branch is used to extract discriminative features between the embedded and normal regions. DBDH employs a detection head to directly detect the four vertices of the embedding region. In addition, we introduce an extra segmentation head to segment the mask of the embedding region during training. The segmentation head provides pixel-level supervision for model learning, facilitating better learning of the embedded signals. Based on two state-of-the-art invisible offline-to-online messaging methods, we construct two datasets and augmentation strategies for training and testing localization models. Extensive experiments demonstrate the superior performance of the proposed DBDH over existing methods. Chengxin Zhao, Sijing Xie, Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen |
IJCNN | 6 |
| 2024 | Improve Deep Hashing with Language Guidance for Unsupervised Image RetrievalabstractHashing method is widely used in multimedia retrieval systems because of its outstanding retrieval efficiency and low storage cost. Most existing unsupervised hashing methods learn binary hash codes through similarity structure preserving or contrastive learning of hash codes. However, these methods usually use the visual similarity of images to guide hash learning, which does not fully utilize the high-level semantic concept information contained in images, resulting in limited retrieval performance. To tackle this problem, we propose a novel deep unsupervised hashing method called Language Guidance Hashing (LGH). Specifically, LGH utilizes a language model to mine high-level semantic concept information in images and construct a language-based similarity structure, which is used to guide hash learning. By introducing features of textual modality, higher information gain can be brought. In addition, we also propose a language-guided contrastive learning method for learning high-quality binary hash codes. Extensive experimental results show that LGH significantly outperforms state-of-the-art unsupervised hashing methods on three benchmark image datasets. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021 |
ICMR | 4 |
| 2024 | Improving the Transferability of Adversarial Attacks on Face Recognition With Beneficial Perturbation Feature AugmentationabstractFace recognition (FR) models can be easily fooled by adversarial examples, which are crafted by adding imperceptible perturbations on benign face images. The existence of adversarial face examples poses a great threat to the security of society. To build a more sustainable digital nation, in this article, we improve the transferability of adversarial face examples to expose more blind spots of the existing FR models. Though generating hard samples has shown its effectiveness in improving the generalization of models in training tasks, the effectiveness of using this idea to improve the transferability of adversarial face examples remains unexplored. To this end, based on the property of hard samples and the symmetry between training tasks and adversarial attack tasks, we propose the concept of hard models, which have similar effects as hard samples for adversarial attack tasks. Using the concept of hard models, we propose a novel attack method called beneficial perturbation feature augmentation attack (BPFA), which reduces the overfitting of adversarial examples to surrogate FR models by constantly generating new hard models to craft the adversarial examples. Specifically, in the backpropagation, BPFA records the gradients on preselected feature maps and uses the gradient on the input image to craft the adversarial example. In the next forward propagation, BPFA leverages the recorded gradients to add beneficial perturbations on their corresponding feature maps to increase the loss. Extensive experiments demonstrate that BPFA can significantly boost the transferability of adversarial attacks on FR. Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Zongyi Li, Ping Li 0021 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Knowledge Consistency Distillation for Weakly Supervised One Step Person SearchabstractWeakly supervised person search targets to detect and identify a person with only bounding box annotations. Recent approaches have focused on learning person relations in a single model, ignoring the conflicts between the detection and Re-ID heads, along with the influence of background elements, which may lead to noisy pseudo labels and inaccurate Re-ID features. To address this challenge, we introduce a novel framework named Knowledge Consistency Distillation (KCD) for weakly supervised person search, which explores the capabilities of an advanced unsupervised person re-identification (Re-ID) model to mitigate the conflicts and background influences. We propose hierarchical consistency alignments, including feature-level, cluster-level, and instance-level consistency alignment, to synchronize the knowledge from the state-of-the-art unsupervised Re-ID model. Specifically, the feature-level consistency aligns the feature through both context and relation alignment. The cluster-level consistency aligns the teacher cluster information by reusing its OIM module. To tackle the inconsistency problem between student instances and teacher cluster centroids, we incorporate pseudo-label refinement to assist the student model in comprehending the teacher’s knowledge at cluster-level while mitigating the negative effects of noisy labels. Finally, an instance-level consistency loss weighted by the similarity between the instance and its corresponding cluster is proposed to align the positive instance correlations. Our approach aims to train a one-step weakly supervised model for person search by exploiting the characteristics of unsupervised person Re-ID. Extensive experiments illustrate that our method achieves state-of-the-art performance on two widely-used person search datasets, CUHK-SYSU and PRW. Our code will be available on GitHub athttps://github.com/zongyi1999/KCD. Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Chengxin Zhao, Qian Wang 0001, Shijuan Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Gait Recognition With Multi-Level Skeleton-Guided RefinementabstractExisting methods combining skeleton and silhouette representations demonstrate explicit effectiveness for gait recognition. However, current related methods simply combine the video-level representations of model-based skeleton data and gait silhouettes for retrieval. Therefore, diverse skeleton information is not fully exploited in existing related works: Firstly, the position and movement of bones are not clear from individual silhouettes. This indicates that the frame-level interaction between features of skeletons and silhouettes is critical, which is ignored by previous methods. Secondly, diverse part-level skeleton-guided gait features are not fully captured in existing related approaches. To solve the above issues, we present a novel framework with multi-level skeleton-guided refinement, including frame-level, part-level, and video-level skeleton-guided refinement, for comprehensive skeleton-aided gait representation learning. First, two modules are proposed for frame-level skeleton-guided refinement. Specifically, Visual Skeleton Enhanced Backbone (VSEB) is proposed to visually highlight the global and part-level skeleton regions for the feature of each silhouette frame. Moreover, Cross-Visual-Model Frame-level Interaction (CVMFI) is proposed to further transfer the model-based skeleton information to features of the visual modalities. Secondly, part-level visual and model-based skeleton features are utilized to refine the final gait representation. Concretely, in VSEB, Part Skeleton Enhance Network (PSEN) is proposed to visually enhance the position and movement of part-level skeletons. In addition, Semantic Part Pooling (SPP) is proposed for capturing the model-based skeleton features of different semantic parts. Finally, as the video-level skeleton-guided refinement, multimodal video-level features are combined to boost the final recognition performance. Extensive experimental results on prevailing datasets demonstrate that our approach outperforms most existing methods, including the skeleton-aided multi-modal methods. With the multi-level refinement guided by the skeleton modalities, the framework is expected to provide a deeper understanding of skeleton-aided gait recognition. Runsheng Wang, Yuxuan Shi 0002, Zongyi Li, Chengxin Zhao, Bohao Wei, He Li 0052, Ping Li 0021 |
IEEE Trans. Multim. | 2 |
| 2024 | Viewpoint Disentangling and Generation for Unsupervised Object Re-IDabstractUnsupervised object Re-ID aims to learn discriminative identity features from a fully unlabeled dataset to solve the open-class re-identification problem. Satisfying results have been achieved in existing unsupervised Re-ID methods, primarily trained with pseudo-labels created by feature clustering. However, the viewpoint variation of objects is the key challenge, introducing noisy labels in the clustering process. To address this problem, a novel viewpoint disentangling and generation framework (VDG) is proposed to learn viewpoint-invariant ID features, including a disentangling and generation module, as well as a contrastive learning module. First, we design an ID encoder to map the viewpoint and identity features into the latent space. Second, a generator is used to disentangle view features and synthesize images with different orientations. Especially, the well-trained encoder serves as a pre-trained feature extractor in the contrastive learning module. Third, a viewpoint-aware loss and a class-level loss are integrated to facilitate contrastive learning between original and novel views. The generation of novel view images and the application of viewpoint-aware contrastive loss mutually assist model learning viewpoint-invariant ID features. Extensive experiments on Market-1501, DukeMTMC, MSMT17, and VeRi-776 demonstrate the effectiveness of the proposed VDG framework, as well as its superiority over the existing state-of-the-art approaches. The VDG model also demonstrates high quality in the image generation tasks. Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Boyuan Liu, Runsheng Wang, Chengxin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Deep Unsupervised Hashing with Hyperbolic Multi-Structure LearningabstractUnsupervised hashing aims to learn a compact binary hash code to represent complex image content without label information. Existing deep unsupervised hashing methods typically first employ extracted image embeddings to construct semantic similarity structures and then map the images into compact hash codes while preserving the semantic similarity structure. However, the limited representation power of embeddings in Euclidean space and the inadequate exploration of the similarity structure in current methods often result in poorly discriminative hash codes. In this paper, we propose a novel method called Hyperbolic Multi-Structure Hashing (HMSH) to address these issues. Specifically, to increase the representation power of embeddings, we propose to map embeddings from Euclidean space to hyperbolic space and use the similarity structure constructed in hyperbolic space to guide hash learning. Meanwhile, to fully explore the structural information, we investigate four kinds of data structures, including local neighborhood structure, global clustering structure, inter/intra-class variation and variation under perturbation. Different data structures can complement each other, which is beneficial for hash learning. Extensive experimental results on three benchmark image datasets show that HMSH significantly outperforms state-of-the-art unsupervised hashing methods for image retrieval. Chuang Zhao 0001, Yuxuan Shi 0002, Jiazhong Chen |
ECAI | 3 |
| 2023 | Mutual Relative Position Learning Transformer for Cross-View Geo-LocalizationabstractCross-view geo-localization refers to matching ground images with geo-tagged satellite imagery. Existing methods are mainly two-stage, applying a polar transform to roughly eliminate the gap between these two domains, but this might introduce distortions and reduce the discriminativeness of features. In this work, we propose a transformer-based one-stage approach, which unifies gap elimination and feature extraction. The relative position among objects provides critical clues for this task and has strong spatial correspondences between the two views. Firstly, we form the relative position by selecting representative tokens from different regions. Then the relative positions of the two views predict each other and eliminate the gap through mutual learning. Finally, we introduce a novel consistency loss to enhance feature learning by mutual transfer of relational knowledge among samples. Extensive experiments demonstrate that our method achieves state-of-the-art results on both standard and fine-grained datasets.1 Yuxuan Shi 0002, Zongyi Li, Chuang Zhao 0001, Ping Li 0021 |
ICIP | 3 |
| 2023 | MEGL: Multi-Experts Guided Learning Network for Single Camera Training Person Re-IdentificationabstractThe time-saving single-camera training(SCT) person re-identification aims to learn camera-invariant information without cross-camera pedestrian annotations. To address this challenging task, we propose a novel approach called Multi-Experts Guided Learning Network (MEGL-Net) for SCT-ReID that can obtain features not influenced by camera views at the global and local levels under the guidance of multi-camera experts. Firstly, to obtain camera-invariant features, an adaptive feature integration module (AFI) is introduced to adaptively integrate expert-guided features from different camera branches. Then, the proposed camera-local interactive module (CLI) facilitates interaction between the local branch and the camera experts branch for automatically extracting discriminative, domain-invariant features at a fine-grained level. Finally, our framework aggregates expert-guided features with global features and enhanced local features in the testing stage for pedestrian retrieval. Under the Market-SCT and Duke-SCT datasets, experimental results demonstrate that our approach significantly improves ReID performance and outperforms existing state-of-the-art (SOTA) methods. He Li 0052, Yuxuan Shi 0002, Zongyi Li, Runsheng Wang, Chengxin Zhao, Ping Li 0021 |
ICIP | 2 |
| 2023 | Deep Unsupervised Hashing with Semantic Consistency LearningabstractHashing method has attracted more attention in recent years because of its low storage consumption and high retrieval performance. Most unsupervised hashing methods first construct local similarity structure in high-dimensional feature space, and then learn binary hash codes which maintain similarity structure information. However, this local structure based on pairwise distance will bring false guidance and misguide the hashing model. Besides, previous methods rarely consider the robustness of the hashing model, resulting in the unstable hash codes generated under perturbation. Toward these issues, we propose a novel Semantic Consistency Hashing (SCH). Specifically, to avoid misguidance caused by local similarity structure, SCH converts the similarity structure into the probability distribution and preserves semantic information from the perspective of global data distribution. In addition, to improve the robustness of hash codes, we introduce transformation consistency learning to maximize the similarity of hash codes under different transformations of the same image. Experiments on three popular datasets show that SCH outperforms the state-of-the-art methods. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Ping Li 0021 |
ICIP | 4 |
| 2023 | Improve Unsupervised Deep Hashing Via Masked Contrastive LearningabstractUnsupervised hashing method aims to generate compact binary hash codes for images without label supervision. Existing unsupervised hashing methods usually learn binary hash codes by reconstructing input data or preserving similarity structures. However, these methods will either force the hash code to retain a large amount of redundant information or will learn a similarity structure with noise due to biased prior knowledge, resulting in poor retrieval performance. In this paper, we introduce a novel unsupervised hashing method called Masked Contrastive Hashing (MCH). Specifically, to maximally preserve meaningful semantic information into the binary hash code, MCH adopts an encoder-decoder structure and extracts the binary representation from the random masked image to reconstruct the original image. Furthermore, MCH maximizes the consistency of the enhanced views of the same image while minimizing the consistency of different images to establish the similarity relationship between images, which is helpful to generate hash codes that are more suitable for retrieval tasks. Extensive experiments show that the proposed MCH significantly outperforms existing state-of-the-art methods on several benchmark datasets. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Ping Li 0021 |
ICIP | 4 |
| 2023 | Unsupervised Deep Hashing With Deep Semantic DistillationabstractMany existing unsupervised hashing methods attempt to preserve as much semantic information as possible by reconstructing the input data. However, this approach can result in the hash code preserving a lot of redundant information. Besides, previous works usually adopt local structures to guide hashing learning, which will mislead hashing model due to a large amount of noise existing in the local structure. In this paper, we propose a novel Deep Semantic Distillation Hashing (DSDH) to solve the above problems. Specifically, to ensure that the hashing model focuses on preserving more discriminative information rather than background noise, we use random masked images as input for feature extraction. We then apply empirical Maximum Mean Discrepancy to match the output feature distribution with that of the original image. Additionally, to avoid misleading, we propose to constrain the consistency of the similarity structures of the two spaces from the perspective of global distribution, thus transferring the knowledge of the feature space to Hamming space. Experiments conducted on three benchmarks show the superiority of DSDH. Chuang Zhao 0001, Yuxuan Shi 0002, Shijie Lu, Ping Li 0021 |
ICIP | 3 |
| 2023 | Deep Unsupervised Hashing with Selective Semantic MiningabstractMost of the existing unsupervised hashing methods usually construct semantic similarity structure to guide hashing learning. However, due to the lack of filtering of useless information, some wrong guiding information in the similarity structure may damage the retrieval performance. Besides, some works adopt the framework of contrastive learning to preserve the discriminative semantic information that is more important for the hashing task. But such a training strategy may incorrectly embed some semantically similar samples far away due to the absence of manual label supervision, thus producing sub-optimal hash codes. To solve the aforementioned problems, we propose a novel method named Deep Selective Semantic Mining Hashing (DSSMH). Specifically, with the prior knowledge obtained by clustering, we select semantically correct image pairs with high confidence to alleviate the guidance of wrong information and correct sampling bias in contrastive learning. Extensive experiments demonstrate that DSSMH outperforms existing state-of-the-art methods. Chuang Zhao 0001, Yuxuan Shi 0002, Chengxin Zhao, Jiazhong Chen |
ICME | 3 |
| 2022 | Reliability Exploration with Self-Ensemble Learning for Domain Adaptive Person Re-identificationabstractPerson re-identifcation (Re-ID) based on unsupervised domain adaptation (UDA) aims to transfer the pre-trained model from one labeled source domain to an unlabeled target domain. Existing methods tackle this problem by using clustering methods to generate pseudo labels. However, pseudo labels produced by these techniques may be unstable and noisy, substantially deteriorating models’ performance. In this paper, we propose a Reliability Exploration with Self-ensemble Learning (RESL) framework for domain adaptive person ReID. First, to increase the feature diversity, multiple branches are presented to extract features from different data augmentations. Taking the temporally average model as a mean teacher model, online label refning is conducted by using its dynamic ensemble predictions from different branches as soft labels. Second, to combat the adverse effects of unreliable samples in clusters, sample reliability is estimated by evaluating the consistency of different clusters’ results, followed by selecting reliable instances for training and re-weighting sample contribution within Re-ID losses. A contrastive loss is also utilized with cluster-level memory features which are updated by the mean feature. The experiments demonstrate that our method can signifcantly surpass the state-of-the-art performance on the unsupervised domain adaptive person ReID. Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Qian Wang 0001, Fengfan Zhou |
AAAI | 2 |
| 2022 | A Triplet Deviation Network Framework: Boosting Weakly-supervised Anomaly Detection By Ensemble LearningabstractWeakly-supervised anomaly detection refers to extracting the information to identify anomalies from the limited labeled anomalies. Existing work has used the deviation network to learn the anomaly score end-to-end, which treats unlabeled data as normal data, and makes the output score of labeled anomalies deviate greatly from the normal data. However, due to the diversity of anomaly types, the model trained by limited labeled anomaly is one-sided, and can not be generalized to identify other types of anomalies. Simply treating unlabeled data as normal data can not extract features from unlabeled data to improve the generalization and accuracy of the model. In this paper, a triplet deviation network framework(TDNF) is proposed. Compared with the original deviation network, it adds a potential anomalies filtering module and a prior anomaly score generation module. The potential anomalies filtering module ensemble multiple unsupervised methods to evaluate data and filter potential anomalies. Labeled anomalies, potential anomalies, unlabeled data compose multiple triplets, and input to deviation network to improve the ability of the model to identify different types of anomalies. The prior anomaly score generation module uses one unsupervised method to generate normalized prior anomaly scores. The prior anomaly scores as prior knowledge of deviation network, which help to fine-tune the model's learning of unlabeled data to optimize the model's ability of anomaly ranking for unlabeled data. We give a triplet deviation network instance TDN-IHSC and carry out extensive experiments on multiple real-world datasets. The results show that our method is effective and performs better than the other four advanced competitive methods. Shuhui Pan, Yuxuan Shi 0002, Ping Li 0021 |
IJCNN | 3 |
| 2022 | Attribute disentanglement and registration for occluded person re-identification
Yuxuan Shi 0002, Lei Wu 0010, Baiyan Zhang, Ping Li 0021 |
Neurocomputing | 1 |
| 2022 | Spatial-wise and channel-wise feature uncertainty for occluded person re-identification
Yuxuan Shi 0002, Weiyi Tian, Zongyi Li, Ping Li 0021 |
Neurocomputing | 1 |
| 2022 | Instance Correlation Graph for Unsupervised Domain AdaptationabstractIn recent years, deep neural networks have emerged as a dominant machine learning tool for a wide variety of application fields. Due to the expensive cost of manual labeling efforts, it is important to transfer knowledge from a label-rich source domain to an unlabeled target domain. The core problem is how to learn a domain-invariant representation to address the domain shift challenge, in which the training and test samples come from different distributions. First, considering the geometry of space probability distributions, we introduce an effective Hellinger Distance to match the source and target distributions on statistical manifold. Second, the data samples are not isolated individuals, and they are interrelated. The correlation information of data samples should not be neglected for domain adaptation. Distinguished from previous works, we pay attention to the correlation distributions over data samples. We design elaborately a Residual Graph Convolutional Network to construct the Instance Correlation Graph (ICG). The correlation information of data samples is exploited to reduce the domain shift. Therefore, a novel Instance Correlation Graph for Unsupervised Domain Adaptation is proposed, which is trained end-to-end by jointly optimizing three types of losses, i.e., Supervised Classification loss for source domain, Centroid Alignment loss to measure the centroid difference between source and target domain, ICG Alignment loss to match Instance Correlation Graph over two related domains. Extensive experiments are conducted on several hard transfer tasks to learn domain-invariant representations on three benchmarks: Office-31, Office-Home, and VisDA2017. Compared with other state-of-the-art techniques, our method achieves superior performance. Lei Wu 0010, Yuxuan Shi 0002, Baiyan Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Hands-On Guidance for Distilling Object DetectorsabstractKnowledge distillation can lead to deploy-friendly networks against the plagued computational complexity problem, but previous methods neglect the feature hierarchy in detectors. Motivated by this, we propose a general framework for detection distillation. Our method, called Hands-on Guidance Distillation, distills the latent knowledge of all stage features for imposing more comprehensive supervision, and focuses on the essence simultaneously for promoting more intense knowledge absorption. Specifically, a series of novel mechanisms are designed elaborately, including correspondence establishment for consistency, hands-on imitation loss measure and re-weighted optimization from both micro and macro perspectives. We conduct extensive evaluations with different distillation configurations over VOC and COCO datasets, which show better performance on accuracy and speed trade-offs. Meanwhile, feasibility experiments on different structural networks further prove the robustness of our HGD. Yangyang Qin, Zhenghai He, Yuxuan Shi 0002, Lei Wu 0010 |
ICME | 4 |
| 2021 | Re-ranking person re-identification using distance aggregation of k-nearest neighbors hierarchical tree
Muhammad Hanif 0004, Weiyi Tian, Yuxuan Shi 0002, Mudassar Rauf |
Multim. Tools Appl. | 4 |
| 2021 | Person Retrieval in Surveillance Videos Via Deep Attribute Mining and ReasoningabstractPerson retrieval largely relies on the appearance features of pedestrians. This task is rather more difficult in surveillance videos due to the limitations of extracting robust appearance features brought by the cross-view and cross-camera data with lower image resolution, motion blur, occlusion and other kinds of image degradation. To build up a more reliable person retrieval system, recent works introduced appearance attribute models to describe and distinguish different persons with high-level semantic concepts. Despite the progress of previous works, the value of utilizing appearance attributes is still under-explored. On one hand, existing methods lack for concise and precise attribute representations that are specific for each attribute category and, in the meantime, are able to filter noisy information in irrelevant spatial locations and useless patterns. On the other hand, correlation and reasoning between different attributes are neglected, which could generate more useful information and add more robustness to the retrieval system. In this paper, we propose an Attribute Mining and Reasoning (AMR) framework which is capable to handle the issues in question. The AMR makes better use of appearance attributes with two main components. First, the AMR disentangles the representations of different attributes by localizing their spatial positions and identifying their effective patterns in a weakly supervised manner. To achieve more reliable localization, we propose the Attribute Localization Ensemble (ALE) module that is consisted of multiple localization heads and a voting mechanism. Second, we introduce the Attribute Reasoning (AR) module to correlate different attributes together with the global appearance features and discover their latent relations to generate more comprehensive descriptions of pedestrians. Extensive experiments on DukeMTMC-ReID and Market-1501 datasets demonstrate the effectiveness of the proposed AMR framework as well as its superiority over the existing state-of-the-art methods. The AMR model also shows great generalization ability on the unseen CUHK03 dataset when it is only trained on Market-1501 dataset. Yuxuan Shi 0002, Zhen Wei 0001, Jialie Shen 0001, Ping Li 0021 |
IEEE Trans. Multim. | 1 |
| 2021 | Adaptive and Robust Partition Learning for Person Retrieval With Policy GradientabstractPerson retrieval aims at effectively matching the pedestrian images over an extensive database given a specified identity. As extracting effective features is crucial in a high-performance retrieval system, recent significant progress was achieved by part-based models that have constructed robust local representations on top of vertically striped part features. However, this kind of models use predefined partitioning strategies, making the number and size of each partition identical even when input images vary a lot. This unchangeable setting usually leads to less flexibility and robustness in capturing visual variance. The primary reason for such a negative effect is that a fixed partitioning strategy is unable to deal with (a) the significant variance from pose, illumination and viewpoint which is common in a pedestrian image dataset, and (b) also the inference error and misalignment of human bodies introduced by the prepositive pedestrian detection module or human pose estimation module. In this paper, we tackle this problem via introducing the novel Adaptive Partition Network (APN). The APN utilizes deep reinforcement learning and applies an agent to generate optimal partitioning strategies dynamically for different input images. The agent inside the APN is optimized with the policy gradient algorithm and maximizes the reward of choosing the best partition setting. By leveraging the supervision cues from the objective partitioning strategies that are generated on a set of held-out training images, the agent is trained jointly with other parts of APN, which ensures the APN's robustness and generalization ability. Extensive experimental results on multiple datasets, including CUHK03, DukeMTMC and Market-1501, demonstrate the superiority of APN over the state-of-the-art models. Yuxuan Shi 0002, Zhen Wei 0001, Pengfei Zhu 0001, Jialie Shen 0001, Ping Li 0021 |
IEEE Trans. Multim. | 1 |
| 2020 | Selective Convolutional Network: An Efficient Object Detector with Ignoring BackgroundabstractIt is well known that attention mechanisms can effectively improve the performance of many CNNs including object detectors. Instead of refining feature maps prevalently, we reduce the prohibitive computational complexity by a novel attempt at attention. Therefore, we introduce an efficient object detector called Selective Convolutional Network (SCN), which selectively calculates only on the locations that contain meaningful and conducive information. The basic idea is to exclude the insignificant background areas, which effectively reduces the computational cost especially during the feature extraction. To solve it, we design an elaborate structure with negligible overheads to guide the network where to look next. It’s end-to-end trainable and easy-embedding. Without additional segmentation datasets, we explores two different train strategies including direct supervision and indirect supervision. Extensive experiments assess the performance on PASCAL VOC2007 and MS COCO detection datasets. Results show that SSD and Pelee integrated with our method averagely reduce the calculations in a range of 1/5 and 1/3 with slight loss of accuracy, demonstrating the feasibility of SCN. Yangyang Qin, Yuxuan Shi 0002, Ping Li 0021 |
ICASSP | 4 |
| 2020 | Lightweight Action Recognition with Sequence-Specific Global ContextabstractWith the emergence of a large number of video resources, video action recognition is attracting much attention. Recently, realizing the outstanding performance of three-dimensional (3D) convolutional neural networks (CNNs), many works have began to apply them for action recognition and obtained satisfactory results. However, high computational over-heads greatly reduce the efficiency of 3D CNNs. To make up for the shortcoming, in this paper, we first propose two innovations - the Xwise Separable Convolution and the SS block, both of which are lightweight. Then we build an efficient 3D CNN called the XwiseNet based on our innovations. Our work aims to make 3D CNNs lightweight without reducing the recognition accuracy. The key idea of the Xwise Separable Convolution is extremely decoupling the 3D convolution in channel, spatial, and temporal dimensions. The SS block can capture temporal long-range dependencies via aggregating sequence-specific global context to each sequence feature. Experiments have verified that our XwiseNet achieves competitive performance with the least computational overhead. Jiazhong Chen, Lei Wu 0010, Yuxuan Shi 0002 |
IJCNN | 5 |
| 2020 | Learning refined attribute-aligned network with attribute selection for person re-identification
Yuxuan Shi 0002, Lei Wu 0010, Jialie Shen 0001, Ping Li 0021 |
Neurocomputing | 1 |
| 2020 | XwiseNet: action recognition with Xwise separable convolutions
Jiazhong Chen, Lei Wu 0010, Yuxuan Shi 0002 |
Multim. Tools Appl. | 5 |
| 2019 | BMNet: A Reconstructed Network for Lightweight Object Detection via Branch Merging
Yangyang Qin, Yuxuan Shi 0002, Lei Wu 0010, Jiazhong Chen, Baiyan Zhang |
BMVC | 4 |
| 2019 | Improving person re-identification by multi-task learning
Ping Li 0021, Yuxuan Shi 0002, Jiazhong Chen, Fuhao Zou |
Neurocomputing | 4 |