VLDB 2026 Research / reviewers in the wild / expert
Ping Li 0021
dblp:62/5860-21
· DBLP profile ↗
36ranked-venue papers
1as first author
22since 2021 · last 2025
0000-0002-9534-816XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TSAD: Temporal-spatial association differences-based unsupervised anomaly detection for multivariate time-series
Hanbing Zhu, Zongyi Li, Yuxuan Shi 0002, Chuang Zhao 0001, Hongxu Ji, Ping Li 0021 |
Neurocomputing | 8 |
| 2025 | GSyncCode: Geometry Synchronous Hidden Code for One-step Photography DecodingabstractInvisible hyperlinks and hidden barcodes have recently emerged as a hot topic in offline-to-online messaging, where an invisible message or barcode is embedded in an image and can be decoded via camera shooting. Current schemes involve a two-step decoding process: starting with vertex localization of the embedded region to correct the perspective distortion introduced by shooting, followed by decoding the message from the corrected region. However, vertex localization can be complex and time-consuming, which affects the efficiency and accuracy of message decoding. To address this issue, this article proposes a geometry synchronous decoding scheme called GSyncCode, allowing for one-step extraction of a Data Matrix code from the photograph. Instead of correction before decoding, GSyncCode directly decodes a geometry-transformed Data Matrix that is synchronized with the embedded region. A barcode scanner is then used to efficiently retrieve messages. We design a Haar transform-based encoder HaarUNet and a HaarLoss visual function to select the key component of the Data Matrix for embedding. They improve the visual quality of the embedded image by reducing redundant embedding signals. Extensive simulated and real-world experiments demonstrate the superiority of GSyncCode in both decoding efficiency and accuracy. Our codes are published at: https://github.com/zcx-language/GSyncCode . Chengxin Zhao, Jialie Shen 0001, Han Fang 0004, Sijing Xie, Yaokun Fang, Zongyi Li, Ping Li 0021 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature ExplorationabstractPerson search is a unified system aimed at jointly localizing and identifying a person of interest from a gallery of whole scene images. Due to the inherent properties of the person search, it faces significant challenges of large-scale variations, inaccurate detection boxes, and crowded scenes. To address these issues, we proposed an uncertainty-guided framework coupled with auxiliary shallow feature exploration, which includes a shallow feature fusion module and an uncertainty-guided module. Firstly, considering the scales of the person are varied due to various scenes and their relative positions to the camera, a shallow feature fusion module is designed to extract multi-scale features to assist the re-id sub-task. Additionally, a self-distillation loss is proposed to align features across different scales. Furthermore, to alleviate the problem that the model can be easily affected by coarse samples resulting from crowded scenes and inaccurate detection boxes, we introduce an uncertainty guidance module to reduce the negative impact of these coarse targets. The experimental results demonstrate the effectiveness of our proposed methods on two benchmarks (i.e., CUHK-SYSU, and PRW). Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Ping Li 0021 |
ICASSP | 7 |
| 2024 | Improving Visual Quality and Transferability of Adversarial Attacks on Face Recognition Simultaneously with Adversarial RestorationabstractAdversarial face examples possess two critical properties: Visual Quality and Transferability. However, existing approaches rarely address these properties simultaneously, leading to subpar results. To address this issue, we propose a novel adversarial attack technique known as Adversarial Restoration (AdvRestore), which enhances both visual quality and transferability of adversarial face examples by leveraging a face restoration prior. In our approach, we initially train a Restoration Latent Diffusion Model (RLDM) designed for face restoration. Subsequently, we employ the inference process of RLDM to generate adversarial face examples. The adversarial perturbations are applied to the intermediate features of RLDM. Additionally, by treating RLDM face restoration as a sibling task, the transferability of the generated adversarial face examples is further improved. Our experimental results validate the effectiveness of the proposed attack method. Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021 |
ICASSP | 5 |
| 2024 | Improve Deep Hashing with Language Guidance for Unsupervised Image RetrievalabstractHashing method is widely used in multimedia retrieval systems because of its outstanding retrieval efficiency and low storage cost. Most existing unsupervised hashing methods learn binary hash codes through similarity structure preserving or contrastive learning of hash codes. However, these methods usually use the visual similarity of images to guide hash learning, which does not fully utilize the high-level semantic concept information contained in images, resulting in limited retrieval performance. To tackle this problem, we propose a novel deep unsupervised hashing method called Language Guidance Hashing (LGH). Specifically, LGH utilizes a language model to mine high-level semantic concept information in images and construct a language-based similarity structure, which is used to guide hash learning. By introducing features of textual modality, higher information gain can be brought. In addition, we also propose a language-guided contrastive learning method for learning high-quality binary hash codes. Extensive experimental results show that LGH significantly outperforms state-of-the-art unsupervised hashing methods on three benchmark image datasets. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021 |
ICMR | 6 |
| 2024 | Improving the Transferability of Adversarial Attacks on Face Recognition With Beneficial Perturbation Feature AugmentationabstractFace recognition (FR) models can be easily fooled by adversarial examples, which are crafted by adding imperceptible perturbations on benign face images. The existence of adversarial face examples poses a great threat to the security of society. To build a more sustainable digital nation, in this article, we improve the transferability of adversarial face examples to expose more blind spots of the existing FR models. Though generating hard samples has shown its effectiveness in improving the generalization of models in training tasks, the effectiveness of using this idea to improve the transferability of adversarial face examples remains unexplored. To this end, based on the property of hard samples and the symmetry between training tasks and adversarial attack tasks, we propose the concept of hard models, which have similar effects as hard samples for adversarial attack tasks. Using the concept of hard models, we propose a novel attack method called beneficial perturbation feature augmentation attack (BPFA), which reduces the overfitting of adversarial examples to surrogate FR models by constantly generating new hard models to craft the adversarial examples. Specifically, in the backpropagation, BPFA records the gradients on preselected feature maps and uses the gradient on the input image to craft the adversarial example. In the next forward propagation, BPFA leverages the recorded gradients to add beneficial perturbations on their corresponding feature maps to increase the loss. Extensive experiments demonstrate that BPFA can significantly boost the transferability of adversarial attacks on FR. Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Zongyi Li, Ping Li 0021 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Gait Recognition With Multi-Level Skeleton-Guided RefinementabstractExisting methods combining skeleton and silhouette representations demonstrate explicit effectiveness for gait recognition. However, current related methods simply combine the video-level representations of model-based skeleton data and gait silhouettes for retrieval. Therefore, diverse skeleton information is not fully exploited in existing related works: Firstly, the position and movement of bones are not clear from individual silhouettes. This indicates that the frame-level interaction between features of skeletons and silhouettes is critical, which is ignored by previous methods. Secondly, diverse part-level skeleton-guided gait features are not fully captured in existing related approaches. To solve the above issues, we present a novel framework with multi-level skeleton-guided refinement, including frame-level, part-level, and video-level skeleton-guided refinement, for comprehensive skeleton-aided gait representation learning. First, two modules are proposed for frame-level skeleton-guided refinement. Specifically, Visual Skeleton Enhanced Backbone (VSEB) is proposed to visually highlight the global and part-level skeleton regions for the feature of each silhouette frame. Moreover, Cross-Visual-Model Frame-level Interaction (CVMFI) is proposed to further transfer the model-based skeleton information to features of the visual modalities. Secondly, part-level visual and model-based skeleton features are utilized to refine the final gait representation. Concretely, in VSEB, Part Skeleton Enhance Network (PSEN) is proposed to visually enhance the position and movement of part-level skeletons. In addition, Semantic Part Pooling (SPP) is proposed for capturing the model-based skeleton features of different semantic parts. Finally, as the video-level skeleton-guided refinement, multimodal video-level features are combined to boost the final recognition performance. Extensive experimental results on prevailing datasets demonstrate that our approach outperforms most existing methods, including the skeleton-aided multi-modal methods. With the multi-level refinement guided by the skeleton modalities, the framework is expected to provide a deeper understanding of skeleton-aided gait recognition. Runsheng Wang, Yuxuan Shi 0002, Zongyi Li, Chengxin Zhao, Bohao Wei, He Li 0052, Ping Li 0021 |
IEEE Trans. Multim. | 8 |
| 2023 | Mutual Relative Position Learning Transformer for Cross-View Geo-LocalizationabstractCross-view geo-localization refers to matching ground images with geo-tagged satellite imagery. Existing methods are mainly two-stage, applying a polar transform to roughly eliminate the gap between these two domains, but this might introduce distortions and reduce the discriminativeness of features. In this work, we propose a transformer-based one-stage approach, which unifies gap elimination and feature extraction. The relative position among objects provides critical clues for this task and has strong spatial correspondences between the two views. Firstly, we form the relative position by selecting representative tokens from different regions. Then the relative positions of the two views predict each other and eliminate the gap through mutual learning. Finally, we introduce a novel consistency loss to enhance feature learning by mutual transfer of relational knowledge among samples. Extensive experiments demonstrate that our method achieves state-of-the-art results on both standard and fine-grained datasets.1 Yuxuan Shi 0002, Zongyi Li, Chuang Zhao 0001, Ping Li 0021 |
ICIP | 6 |
| 2023 | MEGL: Multi-Experts Guided Learning Network for Single Camera Training Person Re-IdentificationabstractThe time-saving single-camera training(SCT) person re-identification aims to learn camera-invariant information without cross-camera pedestrian annotations. To address this challenging task, we propose a novel approach called Multi-Experts Guided Learning Network (MEGL-Net) for SCT-ReID that can obtain features not influenced by camera views at the global and local levels under the guidance of multi-camera experts. Firstly, to obtain camera-invariant features, an adaptive feature integration module (AFI) is introduced to adaptively integrate expert-guided features from different camera branches. Then, the proposed camera-local interactive module (CLI) facilitates interaction between the local branch and the camera experts branch for automatically extracting discriminative, domain-invariant features at a fine-grained level. Finally, our framework aggregates expert-guided features with global features and enhanced local features in the testing stage for pedestrian retrieval. Under the Market-SCT and Duke-SCT datasets, experimental results demonstrate that our approach significantly improves ReID performance and outperforms existing state-of-the-art (SOTA) methods. He Li 0052, Yuxuan Shi 0002, Zongyi Li, Runsheng Wang, Chengxin Zhao, Ping Li 0021 |
ICIP | 7 |
| 2023 | Deep Unsupervised Hashing with Semantic Consistency LearningabstractHashing method has attracted more attention in recent years because of its low storage consumption and high retrieval performance. Most unsupervised hashing methods first construct local similarity structure in high-dimensional feature space, and then learn binary hash codes which maintain similarity structure information. However, this local structure based on pairwise distance will bring false guidance and misguide the hashing model. Besides, previous methods rarely consider the robustness of the hashing model, resulting in the unstable hash codes generated under perturbation. Toward these issues, we propose a novel Semantic Consistency Hashing (SCH). Specifically, to avoid misguidance caused by local similarity structure, SCH converts the similarity structure into the probability distribution and preserves semantic information from the perspective of global data distribution. In addition, to improve the robustness of hash codes, we introduce transformation consistency learning to maximize the similarity of hash codes under different transformations of the same image. Experiments on three popular datasets show that SCH outperforms the state-of-the-art methods. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Ping Li 0021 |
ICIP | 6 |
| 2023 | Improve Unsupervised Deep Hashing Via Masked Contrastive LearningabstractUnsupervised hashing method aims to generate compact binary hash codes for images without label supervision. Existing unsupervised hashing methods usually learn binary hash codes by reconstructing input data or preserving similarity structures. However, these methods will either force the hash code to retain a large amount of redundant information or will learn a similarity structure with noise due to biased prior knowledge, resulting in poor retrieval performance. In this paper, we introduce a novel unsupervised hashing method called Masked Contrastive Hashing (MCH). Specifically, to maximally preserve meaningful semantic information into the binary hash code, MCH adopts an encoder-decoder structure and extracts the binary representation from the random masked image to reconstruct the original image. Furthermore, MCH maximizes the consistency of the enhanced views of the same image while minimizing the consistency of different images to establish the similarity relationship between images, which is helpful to generate hash codes that are more suitable for retrieval tasks. Extensive experiments show that the proposed MCH significantly outperforms existing state-of-the-art methods on several benchmark datasets. Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Ping Li 0021 |
ICIP | 5 |
| 2023 | Unsupervised Deep Hashing With Deep Semantic DistillationabstractMany existing unsupervised hashing methods attempt to preserve as much semantic information as possible by reconstructing the input data. However, this approach can result in the hash code preserving a lot of redundant information. Besides, previous works usually adopt local structures to guide hashing learning, which will mislead hashing model due to a large amount of noise existing in the local structure. In this paper, we propose a novel Deep Semantic Distillation Hashing (DSDH) to solve the above problems. Specifically, to ensure that the hashing model focuses on preserving more discriminative information rather than background noise, we use random masked images as input for feature extraction. We then apply empirical Maximum Mean Discrepancy to match the output feature distribution with that of the original image. Additionally, to avoid misleading, we propose to constrain the consistency of the similarity structures of the two spaces from the perspective of global distribution, thus transferring the knowledge of the feature space to Hamming space. Experiments conducted on three benchmarks show the superiority of DSDH. Chuang Zhao 0001, Yuxuan Shi 0002, Shijie Lu, Ping Li 0021 |
ICIP | 6 |
| 2023 | Detecting Adversarial Faces Using Only Real Face Self-PerturbationsabstractAdversarial attacks aim to disturb the functionality of a target system by adding specific noise to the input samples, bringing potential threats to security and robustness when applied to facial recognition systems. Although existing defense techniques achieve high accuracy in detecting some specific adversarial faces (adv-faces), new attack methods especially GAN-based attacks with completely different noise patterns circumvent them and reach a higher attack success rate. Even worse, existing techniques require attack data before implementing the defense, making it impractical to defend newly emerging attacks that are unseen to defenders. In this paper, we investigate the intrinsic generality of adv-faces and propose to generate pseudo adv-faces by perturbing real faces with three heuristically designed noise patterns. We are the first to train an adv-face detector using only real faces and their self-perturbations, agnostic to victim facial recognition systems, and agnostic to unseen attacks. By regarding adv-faces as out-of-distribution data, we then naturally introduce a novel cascaded system for adv-face detection, which consists of training data self-perturbations, decision boundary regularization, and a max-pooling-based binary classifier focusing on abnormal local color aberrations. Experiments conducted on LFW and CelebA-HQ datasets with eight gradient-based and two GAN-based attacks validate that our method generalizes to a variety of unseen adversarial attacks. Qian Wang 0001, Yongqin Xian, Xiaorui Lin, Ping Li 0021, Jiazhong Chen, Ning Yu 0006 |
IJCAI | 6 |
| 2023 | General Adversarial Perturbation Simulating: Protect Unknown System by Detecting Unknown Adversarial FacesabstractBenefitting from the development of convolutional neural networks (CNNs), face recognition systems (FRSs) play a key role in many security-critical systems. However, FRSs have been proved to be vulnerable to adversarial faces (adv-faces). Adv-faces aim to change classification results by adding a subtle perturbation on real faces. The existence of adv-faces poses a significant threat to financial and privacy security. Previous detection methods require either training on pre-computed adv-faces or accessing to protected victim FRSs, bringing a dilemma in practical using. In this work, we heuristically propose an adversarial face detection method called General Adversarial Perturbation Simulating (GAPS) which is blind to both adversarial attacks and FRSs. Simulating noise patterns of several gradient-based adversarial perturbations, GAPS is able to generate simulated adversarial faces (sadv-faces) guiding detectors to learn general adversarial perturbation features and focus on classifying sensitive regions. Extensive experiments on LFW and CASIA-WebFace show that our method outperforms 9 state-of-the-art baseline methods and demonstrate the effectiveness of GAPS. Feiran Sun, Xiaorui Lin, Jiazhong Chen, Ping Li 0021, Qian Wang 0001 |
IJCNN | 6 |
| 2022 | A Triplet Deviation Network Framework: Boosting Weakly-supervised Anomaly Detection By Ensemble LearningabstractWeakly-supervised anomaly detection refers to extracting the information to identify anomalies from the limited labeled anomalies. Existing work has used the deviation network to learn the anomaly score end-to-end, which treats unlabeled data as normal data, and makes the output score of labeled anomalies deviate greatly from the normal data. However, due to the diversity of anomaly types, the model trained by limited labeled anomaly is one-sided, and can not be generalized to identify other types of anomalies. Simply treating unlabeled data as normal data can not extract features from unlabeled data to improve the generalization and accuracy of the model. In this paper, a triplet deviation network framework(TDNF) is proposed. Compared with the original deviation network, it adds a potential anomalies filtering module and a prior anomaly score generation module. The potential anomalies filtering module ensemble multiple unsupervised methods to evaluate data and filter potential anomalies. Labeled anomalies, potential anomalies, unlabeled data compose multiple triplets, and input to deviation network to improve the ability of the model to identify different types of anomalies. The prior anomaly score generation module uses one unsupervised method to generate normalized prior anomaly scores. The prior anomaly scores as prior knowledge of deviation network, which help to fine-tune the model's learning of unlabeled data to optimize the model's ability of anomaly ranking for unlabeled data. We give a triplet deviation network instance TDN-IHSC and carry out extensive experiments on multiple real-world datasets. The results show that our method is effective and performs better than the other four advanced competitive methods. Shuhui Pan, Yuxuan Shi 0002, Ping Li 0021 |
IJCNN | 5 |
| 2022 | Attribute disentanglement and registration for occluded person re-identification
Yuxuan Shi 0002, Lei Wu 0010, Baiyan Zhang, Ping Li 0021 |
Neurocomputing | 5 |
| 2022 | Spatial-wise and channel-wise feature uncertainty for occluded person re-identification
Yuxuan Shi 0002, Weiyi Tian, Zongyi Li, Ping Li 0021 |
Neurocomputing | 5 |
| 2021 | Selective Adversarial Adaptation Learning via Exclusive Regularization for Partial Domain AdaptationabstractIn consideration of the suitability for the application scenario, partial domain adaptation is more significant and more valuable than traditional domain adaptation. Most existing partial domain adaptation methods adopt weighting mechanism to avoid negative migration which is caused by outlier classes samples. However, these methods give the equal consideration of each category in the source domain and determine the classes weight by classifier or discriminator, and they do not consider the possible misprediction of the similar samples from classes which are difficult to distinguish in the source domain. This situation may cause the misalignment of the outlier source classes and target classes, and the wrong alignment of the discriminators. In this work, we propose a selective adversarial adaptation learning method via exclusive regularization for partial domain adaptation (ERPDA) to solve these problems. Specifically, we utilize the exclusive regularization to extend the distance between samples of different classes in source domain to learn an inter-class separable discriminant representation to avoid negative transfer. Meanwhile, the positive transfer is performed by Joint Maximum Mean Discrepancy (JMMD) based on selective adaptation adversarial learning via multi-discriminator. Extensive experiments show that ERPDA achieves state-of-the-art results on several partial domain adaptation benchmark datasets. Ping Li 0021, LinLin Shen, Lei Wu 0010, Qian Wang 0001, Chuang Zhao 0001 |
IJCNN | 1 |
| 2021 | Learning Diverse Local Patterns for Deepfake Detection with Image-level SupervisionabstractTo prevent the Deepfake-like videos from spreading, researchers have proposed many anti-forgery methods. However, most approaches require pixel-wise annotation, which conflicts with the real Deepfake detection scenario. To make full use of the image-level label, we propose a Local-Prediction framework that indirectly allows the image-level label to supervise local regions. To further enrich the local feature, we introduced the Local-Diversity concept in the Deepfake detection field for the first time. We proposed the Local-Diversity Loss based on the motivation that regional pattern differences can provide semi-supervised information during training. Compared to the previous method, our approach limits each classification unit's receptive field and enriches the feature diversity. In the experiment, our method is evaluated on three benchmark datasets of four widely-used manipulation types. The result shows that the Local-Prediction framework is beneficial to different CNN backbones and achieved significant performance. The proposed LD loss enriches the learned patterns of binary classifiers. Furthermore, we provide visualization and ablation studies to understand the mechanism. Junrui Huang, Chengxin Zhao, Yutong Yao, Jiazhong Chen, Ping Li 0021 |
IJCNN | 6 |
| 2021 | Attention-Based Convolutional Neural Network for ASV Spoofing Detection
Leichao Huang, Junrui Huang, Baiyan Zhang, Ping Li 0021 |
Interspeech | 5 |
| 2021 | Person Retrieval in Surveillance Videos Via Deep Attribute Mining and ReasoningabstractPerson retrieval largely relies on the appearance features of pedestrians. This task is rather more difficult in surveillance videos due to the limitations of extracting robust appearance features brought by the cross-view and cross-camera data with lower image resolution, motion blur, occlusion and other kinds of image degradation. To build up a more reliable person retrieval system, recent works introduced appearance attribute models to describe and distinguish different persons with high-level semantic concepts. Despite the progress of previous works, the value of utilizing appearance attributes is still under-explored. On one hand, existing methods lack for concise and precise attribute representations that are specific for each attribute category and, in the meantime, are able to filter noisy information in irrelevant spatial locations and useless patterns. On the other hand, correlation and reasoning between different attributes are neglected, which could generate more useful information and add more robustness to the retrieval system. In this paper, we propose an Attribute Mining and Reasoning (AMR) framework which is capable to handle the issues in question. The AMR makes better use of appearance attributes with two main components. First, the AMR disentangles the representations of different attributes by localizing their spatial positions and identifying their effective patterns in a weakly supervised manner. To achieve more reliable localization, we propose the Attribute Localization Ensemble (ALE) module that is consisted of multiple localization heads and a voting mechanism. Second, we introduce the Attribute Reasoning (AR) module to correlate different attributes together with the global appearance features and discover their latent relations to generate more comprehensive descriptions of pedestrians. Extensive experiments on DukeMTMC-ReID and Market-1501 datasets demonstrate the effectiveness of the proposed AMR framework as well as its superiority over the existing state-of-the-art methods. The AMR model also shows great generalization ability on the unseen CUHK03 dataset when it is only trained on Market-1501 dataset. Yuxuan Shi 0002, Zhen Wei 0001, Jialie Shen 0001, Ping Li 0021 |
IEEE Trans. Multim. | 6 |
| 2021 | Adaptive and Robust Partition Learning for Person Retrieval With Policy GradientabstractPerson retrieval aims at effectively matching the pedestrian images over an extensive database given a specified identity. As extracting effective features is crucial in a high-performance retrieval system, recent significant progress was achieved by part-based models that have constructed robust local representations on top of vertically striped part features. However, this kind of models use predefined partitioning strategies, making the number and size of each partition identical even when input images vary a lot. This unchangeable setting usually leads to less flexibility and robustness in capturing visual variance. The primary reason for such a negative effect is that a fixed partitioning strategy is unable to deal with (a) the significant variance from pose, illumination and viewpoint which is common in a pedestrian image dataset, and (b) also the inference error and misalignment of human bodies introduced by the prepositive pedestrian detection module or human pose estimation module. In this paper, we tackle this problem via introducing the novel Adaptive Partition Network (APN). The APN utilizes deep reinforcement learning and applies an agent to generate optimal partitioning strategies dynamically for different input images. The agent inside the APN is optimized with the policy gradient algorithm and maximizes the reward of choosing the best partition setting. By leveraging the supervision cues from the objective partitioning strategies that are generated on a set of held-out training images, the agent is trained jointly with other parts of APN, which ensures the APN's robustness and generalization ability. Extensive experimental results on multiple datasets, including CUHK03, DukeMTMC and Market-1501, demonstrate the superiority of APN over the state-of-the-art models. Yuxuan Shi 0002, Zhen Wei 0001, Pengfei Zhu 0001, Jialie Shen 0001, Ping Li 0021 |
IEEE Trans. Multim. | 7 |
| 2020 | Selective Convolutional Network: An Efficient Object Detector with Ignoring BackgroundabstractIt is well known that attention mechanisms can effectively improve the performance of many CNNs including object detectors. Instead of refining feature maps prevalently, we reduce the prohibitive computational complexity by a novel attempt at attention. Therefore, we introduce an efficient object detector called Selective Convolutional Network (SCN), which selectively calculates only on the locations that contain meaningful and conducive information. The basic idea is to exclude the insignificant background areas, which effectively reduces the computational cost especially during the feature extraction. To solve it, we design an elaborate structure with negligible overheads to guide the network where to look next. It’s end-to-end trainable and easy-embedding. Without additional segmentation datasets, we explores two different train strategies including direct supervision and indirect supervision. Extensive experiments assess the performance on PASCAL VOC2007 and MS COCO detection datasets. Results show that SSD and Pelee integrated with our method averagely reduce the calculations in a range of 1/5 and 1/3 with slight loss of accuracy, demonstrating the feasibility of SCN. Yangyang Qin, Yuxuan Shi 0002, Ping Li 0021 |
ICASSP | 5 |
| 2020 | Multi-Object Tracking Via Multi-AttentionabstractData association plays a crucial role in Multi-Object Tracking(MOT), but it is usually suppressed by occlusion. In this paper, we propose an online MOT approach via multiple attention mechanism(Multi-Attention) to handle the frequent interactions between targets. Specifically, the proposed Multi-Attention consists of spatial-attention, channel-attention, and temporal-attention three modules. The spatial-attention module lets the network focus on visible local areas by generating a visibility map, and the channel-attention module combines texture information and context information adaptively to build a recognizable object descriptor, then the temporal-attention module pays different attention to objects in the same trajectory avoiding the suppress caused by contaminated samples. Besides, a multiple branch convolutional block called receptive filed module(RFModule) is introduced to learn multiple levels of information for Multi-Attention. The experimental results on MOTChallenging benchmarks demonstrate the effectiveness of the proposed MOT algorithm against both online and offline trackers. Xianrui Wang, Jiazhong Chen, Ping Li 0021 |
IJCNN | 4 |
| 2020 | Learning refined attribute-aligned network with attribute selection for person re-identification
Yuxuan Shi 0002, Lei Wu 0010, Jialie Shen 0001, Ping Li 0021 |
Neurocomputing | 5 |
| 2020 | Attention-based convolutional neural network for deep face recognition
Jiyang Wu, Junrui Huang, Jiazhong Chen, Ping Li 0021 |
Multim. Tools Appl. | 5 |
| 2019 | Saliency Detection via Topological Feature Modulated Deep LearningabstractThe topological feature in an image, such as connectivity and adjacency, plays an important role in eye fixation detection. However, due to the scalar and additive nature of neurons to aggregate the node values in a local neighborhood, it is hard for convolutional neural networks (CNNs) to directly obtain and model the topological feature. Thus we adopt the topological feature which is pre-trained in accordance with the relationship between the figure and ground. Then we add a topological feature modulated convolutional layer into CNNs. By this way, the topological feature is automatically modulated by the deep features and well modeled by the CNNs. Experimental results show the proposed method surpasses the state-of-the-art by a big margin. Jiazhong Chen, Lei Wu 0010, Baiyan Zhang, Ping Li 0021 |
ICIP | 7 |
| 2019 | Robust Mutual Learning HashingabstractWith the advances in deep learning, deep hashing methods have achieved promising results in recent years. However, tackling the distribution gap between train data and test data still remains unsolved. In this paper, motivated by Spatial Transformer Networks (STN) and mutual learning, we propose a novel robust hashing method (RMLH) for effective image retrieval. Specifically, the network learns more flexible transformation and makes itself generalize better to test data by plugging STN module. Then the mutual learning strategy is introduced to stabilize the training process. Furthermore, we relax the binary variables into continuous variables to avoid introducing any auxiliary variable. Finally, the experimental results show that our method has achieved the state-of-the-art performance on benchmark datasets. Lei Wu 0010, Jiazhong Chen, Ping Li 0021 |
ICIP | 5 |
| 2019 | Improving person re-identification by multi-task learning
Ping Li 0021, Yuxuan Shi 0002, Jiazhong Chen, Fuhao Zou |
Neurocomputing | 3 |
| 2019 | Saliency prediction by Mahalanobis distance of topological feature on deep color components
Jiazhong Chen, Ping Li 0021, Lei Wu 0010 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Objectness Region Enhancement Networks for Scene Parsing
Xin-Yu Ou, Ping Li 0021, Si Liu 0001, Tianjiang Wang, Dan Li 0012 |
J. Comput. Sci. Technol. | 2 |
| 2017 | Adult Image and Video Recognition by a Deep Multicontext Network and Fine-to-Coarse StrategyabstractAdult image and video recognition is an important and challenging problem in the real world. Low-level feature cues do not produce good enough information, especially when the dataset is very large and has various data distributions. This issue raises a serious problem for conventional approaches. In this article, we tackle this problem by proposing a deep multicontext network with fine-to-coarse strategy for adult image and video recognition. We employ a deep convolution networks to model fusion features of sensitive objects in images. Global contexts and local contexts are both taken into consideration and are jointly modeled in a unified multicontext deep learning framework. To make the model more discriminative for diverse target objects, we investigate a novel hierarchical method, and a task-specific fine-to-coarse strategy is designed to make the multicontext modeling more suitable for adult object recognition. Furthermore, some recently proposed deep models are investigated. Our approach is extensively evaluated on four different datasets. One dataset is used for ablation experiments, whereas others are used for generalization experiments. Results show significant and consistent improvements over the state-of-the-art methods. Xinyu Ou, Ping Li 0021, Fuhao Zou, Si Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2013 | Nonnegative sparse coding induced hashing for image copy detection
Fuhao Zou, Hui Feng 0002, Cong Liu 0008, Lingyu Yan, Ping Li 0021, Dan Li 0012 |
Neurocomputing | 6 |
| 2011 | A novel image copy detection scheme based on the local multi-resolution histogram descriptor
Zhihua Xu, Fuhao Zou, Zhengding Lu, Ping Li 0021 |
Multim. Tools Appl. | 5 |
| 2011 | Robust video watermarking based on affine invariant regions in the compressed domain
Liyun Wang, Fuhao Zou, Zhengding Lu, Ping Li 0021 |
Signal Process. | 5 |
| 2009 | Fast and robust video copy detection scheme using full DCT coefficientsabstractIn this paper, a fast and robust video copy detection scheme is proposed, which is suitable for the DCT-coded video sequences. To address the efficiency and effectiveness issue, we extract the video signature directly from the compressed domain. The video sequence clusters are constructed with a fixed length. Each cluster consists of several fictional key-frames. For each key-frame, some low-middle frequency full DCT coefficients are obtained directly from block DCT coefficients, and their ordinal measure is computed and acts as video signature. A rotation compensational strategy is further employed to resist the rotation attacks. The experimental results show that the proposed scheme can be resilient to various types of video transformations, including scaling, rotation, speed change, text insertion, and subsequence insertion/deletion etc.. The most important thing is that the proposed approach not only handles geometric distortion perfectly, but also reduces the computation costs substantially. Zhihua Xu, Fuhao Zou, Zhengding Lu, Ping Li 0021, Tianjiang Wang |
ICME | 5 |