Shiming Ge

dblp:93/8104 · DBLP profile ↗
← Back
92ranked-venue papers
19as first author
52since 2021 · last 2026
0000-0001-5293-310XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 68 · 16 first-author · 37 since 2021Artificial intelligence and machine learning · 35 · 6 first-author · 23 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Masked Face Recognition Under Multi-Representation Alignment Distillation
Chu-Yuan Zhang, Jun-Zheng Zhang, Han-Song Zhang, Yun-Ze Wang, Shiming Ge
ICIC (17)5
2026 Few-shot Class-Incremental Learning via Generative Co-Memory Regularization
Kexin Bao, Dan Zeng 0001, Shiming Ge
Int. J. Comput. Vis.4
2026 Privacy-Preserving Model Transcription With Differentially Private Synthetic Distillation
abstract
While many deep learning models trained on private datasets have been deployed in various practical tasks, they may pose a privacy leakage risk as attackers could recover informative data or label knowledge from models. In this work, we present privacy-preserving model transcription, a data-free model-to-model conversion solution to facilitate model deployment with a privacy guarantee. To this end, we propose a cooperative-competitive learning approach termed differentially private synthetic distillation that learns to convert a pretrained model (teacher) into its privacy-preserving counterpart (student) via a trainable generator without access to private data. The learning collaborates with three players in a unified framework and performs alternate optimization: i) the generator is learned to generate synthetic data, ii) the teacher and student accept the synthetic data and compute differential private labels by flexible data or label noisy perturbation, and iii) the student is updated with noisy labels and the generator is updated by taking the student as a discriminator for adversarial training. We theoretically prove that our approach can guarantee differential privacy and convergence. The transcribed student has good performance and privacy protection, while the resulting generator can generate private synthetic data for downstream tasks. Extensive experiments clearly demonstrate that our approach outperforms 26 state-of-the-arts.
Bochao Liu, Shiming Ge, Shikun Li, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
abstract
Very low-resolution face recognition is challenging due to the serious loss of informative facial details in resolution degradation. Recent approaches based on knowledge distillation provide an effective solution by distilling knowledge from a well-trained teacher for high-resolution face recognition and transferring it to a student for low-resolution face recognition. In general, the existing approaches usually take a discriminative model as teacher, where the teacher knowledge is trained in an abstract manner and provides poor transfer efficiency to compensate for the missing knowledge in low-resolution faces. To make more complete knowledge transfer, we propose a generative-discriminative representation distillation approach that combines generative representation with cross-resolution aligned knowledge distillation. This approach facilitates very low-resolution face recognition by jointly distilling generative and discriminative models via two distillation modules. Firstly, the generative representation distillation takes the encoder of a diffusion model pretrained for face super-resolution as the generative teacher to supervise the learning of the student backbone via feature regression, and then freezes the student backbone. After that, the discriminative representation distillation further considers a pretrained face recognizer as the discriminative teacher to supervise the learning of the student head via cross-resolution relational contrastive distillation. In this way, the general backbone representation can be transformed into discriminative head representation, leading to a robust and discriminative student model for very low-resolution face recognition. Our approach improves the recovery of the missing details in very low-resolution faces and achieves better knowledge transfer. Extensive experiments on face datasets demonstrate that our approach enhances the recognition accuracy of very low-resolution faces, showcasing its effectiveness and adaptability.
Junzheng Zhang, Weijia Guo, Bochao Liu, Ruixin Shi, Shiming Ge
ICASSP6
2025 CD^2: Constrained Dataset Distillation for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) receives significant attention from the public to perform classification continuously with a few training samples, which suffers from the key catastrophic forgetting problem. Existing methods usually employ an external memory to store previous knowledge and treat it with incremental classes equally, which cannot properly preserve previous essential knowledge. To solve this problem and inspired by recent distillation works on knowledge transfer, we propose a framework termed Constrained Dataset Distillation (CD^2) to facilitate FSCIL, which includes a dataset distillation module (DDM) and a distillation constraint module (DCM). Specifically, the DDM synthesizes highly condensed samples guided by the classifier, forcing the model to learn compacted essential class-related clues from a few incremental samples. The DCM introduces a designed loss to constrain the previously learned class distribution, which can preserve distilled knowledge more sufficiently. Extensive experiments on three public datasets show the superiority of our method against other state-of-the-art competitors.
Kexin Bao, Daichi Zhang, Hansong Zhang 0003, Yutao Yue, Shiming Ge
IJCAI6
2025 Divide and Conquer: Static-Dynamic Collaboration for Few-Shot Class-Incremental Learning
abstract
Continual learning systems suffer from catastrophic forgetting, where updates for new tasks destructively interfere with previously acquired knowledge. Recent empirical advances—including flatness-based optimization, static–dynamic architectural decomposition, and probabilistic reg- ularization— have demonstrated strong mitigation of forgetting. However, a unified structural explanation for why these methods succeed remains underdeveloped. This paper proposes a constraint geometry perspective on representation updates in continual learning. We argue that catastrophic forgetting can be interpreted as a curvature-induced vio- lation of constraint-preserving update dynamics. Under this view, successful continual learning methods implicitly regulate update directions in high-curvature regions of the loss landscape. Rather than introducing a new algorithm, this work provides a structural interpretation that clarifies why diverse empirical strategies succeed. Identifying and preserving geometric constraints during gradient-based updates may serve as a guiding principle for future continual learning research.
Kexin Bao, Daichi Zhang, Dan Zeng 0001, Shiming Ge
ICMR5
2025 PKI: Prior knowledge-infused neural network for few-shot class-incremental learning
Kexin Bao, Fanzhao Lin, Dan Zeng 0001, Shiming Ge
Neural Networks6
2025 DSENet++: A Coarse-to-Fine Framework for Enhanced Sub-Region Detection in Aerial Images
Xiangjie Wang, Liang Chen 0004, Junjie Zhang 0002, Jian Zhang 0002, Shiming Ge, Dan Zeng 0001
IEEE Trans. Multim.6
2024 Coupled Confusion Correction: Learning from Crowds with Sparse Annotations
abstract
As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.
Hansong Zhang 0003, Shikun Li, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
AAAI5
2024 M3D: Dataset Condensation by Minimizing Maximum Mean Discrepancy
abstract
Training state-of-the-art (SOTA) deep models often requires extensive data, resulting in substantial training and storage costs. To address these challenges, dataset condensation has been developed to learn a small synthetic set that preserves essential information from the original large-scale dataset. Nowadays, optimization-oriented methods have been the primary method in the field of dataset condensation for achieving SOTA results. However, the bi-level optimization process hinders the practical application of such methods to realistic and larger datasets. To enhance condensation efficiency, previous works proposed Distribution-Matching (DM) as an alternative, which significantly reduces the condensation cost. Nonetheless, current DM-based methods still yield less comparable results to SOTA optimization-oriented methods. In this paper, we argue that existing DM-based methods overlook the higher-order alignment of the distributions, which may lead to sub-optimal matching results. Inspired by this, we present a novel DM-based method named M3D for dataset condensation by Minimizing the Maximum Mean Discrepancy between feature representations of the synthetic and real images. By embedding their distributions in a reproducing kernel Hilbert space, we align all orders of moments of the distributions of real and synthetic images, resulting in a more generalized condensed set. Notably, our method even surpasses the SOTA optimization-oriented method IDC on the high-resolution ImageNet dataset. Extensive analysis is conducted to verify the effectiveness of the proposed method. Source codes are available at https://github.com/Hansong-Zhang/M3D.
Hansong Zhang 0003, Shikun Li, Dan Zeng 0001, Shiming Ge
AAAI5
2024 Learning Differentially Private Diffusion Models via Stochastic Adversarial Distillation
Bochao Liu, Shiming Ge
ECCV (7)3
2024 Learning Natural Consistency Representation for Face Forgery Video Detection
Daichi Zhang, Zihao Xiao 0002, Shikun Li, Fanzhao Lin, Shiming Ge
ECCV (83)6
2024 Masked Face Recognition with Generative-to-Discriminative Representations
abstract
Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified deep network to learn generative-to-discriminative representations for facilitating masked face recognition. To this end, we split the network into three modules and learn them on synthetic masked faces in a greedy module-wise pretraining manner. First, we leverage a generative encoder pretrained for face inpainting and finetune it to represent masked faces into category-aware descriptors. Attribute to the generative encoder’s ability in recovering context information, the resulting descriptors can provide occlusion-robust representations for masked faces, mitigating the effect of diverse masks. Then, we incorporate a multi-layer convolutional network as a discriminative reformer and learn it to convert the category-aware descriptors into identity-aware vectors, where the learning is effectively supervised by distilling relation knowledge from off-the-shelf face recognition model. In this way, the discriminative reformer together with the generative encoder serves as the pretrained backbone, providing general and discriminative representations towards masked faces. Finally, we cascade one fully-connected layer following by one softmax layer into a feature classifier and finetune it to identify the reformed identity-aware vectors. Extensive experiments on synthetic and realistic datasets demonstrate the effectiveness of our approach in recognizing masked faces.
Shiming Ge, Weijia Guo, Chenyu Li 0001, Junzheng Zhang, Dan Zeng 0001
ICML1
2024 DANCE: Dual-View Distribution Alignment for Dataset Condensation
Hansong Zhang 0003, Shikun Li, Fanzhao Lin, Weiping Wang 0005, Zhenxing Qian, Shiming Ge
IJCAI6
2024 Unsupervised Video Face Super-Resolution via Untrained Neural Network Priors
abstract
The goal of video face super-resolution is to reliably reconstruct clear face sequences from low-resolution input videos. Recent approaches either apply a single face super-resolution model directly or train a specific designed model from scratch. Compared with single face super-resolution, video face super-resolution should not only achieve visually plausible results, but also maintain coherency in both space and time. Existing deep-learning based approaches usually train a model using a large amount of external videos which usually are difficult to collect. In this paper, we propose an unsupervised video face super-resolution approach (UVFSR), leveraging the idea of untrained neural network prior to avoid the need for supervised learning. Specifically, we use a single generative convolutional neural network with random initialization to control the exploration of video generation in an online optimization manner. Besides, we explore this unsupervised paradigm with a new iteratively training strategy to make the optimization process more efficient and suitable for video. Moreover, we introduce a consistency-aware loss based on joint face image and face flow prediction to facilitate the temporal consistency. In this way, the super-resolution results can be plausible not only in the appearance domain but also in the motion domain. Experiments validate the effectiveness of our approach in terms of quantitative metrics and visual quality.
Ruixin Shi, Weijia Guo, Shiming Ge
IJCNN3
2024 Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
abstract
Low-resolution face recognition is a challenging task due to the missing of informative details. Recent approaches based on knowledge distillation have proven that high-resolution clues can well guide low-resolution face recognition via proper knowledge transfer. However, due to the distribution difference between training and testing faces, the learned models often suffer from poor adaptability. To address that, we split the knowledge transfer process into distillation and adaptation steps, and propose an adaptable instance-relation distillation approach to facilitate low-resolution face recognition. In the approach, the student distills knowledge from high-resolution teacher in both instance level and relation level, providing sufficient cross-resolution knowledge transfer. Then, the learned student can be adaptable to recognize low-resolution faces with adaptive batch normalization in inference. In this manner, the capability of recovering missing details of familiar low-resolution faces can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on low-resolution face recognition clearly demonstrate the effectiveness and adaptability of our approach.
Ruixin Shi, Weijia Guo, Shiming Ge
IJCNN3
2024 Fusion of Current and Historical Knowledge for Personalized Federated Learning
abstract
Data heterogeneity poses a significant challenge in the realm of federated learning. Personalized federated learning has emerged as a crucial solution to mitigate this challenge. In these methods, the local model is initially replaced with a global model, followed by personalized processing to adapt to the local data. However, a substantial disparity in knowledge representation between the global and local models can lead to the loss of previously acquired knowledge by the local model, which is commonly referred to as catastrophic forgetting. To tackle this issue, we propose a method named FedKML, which utilizes knowledge mutual learning to achieve the fusion of current and historical knowledge. Our method involves preserving the trained local model as a historical local model for each client, thereby retaining valuable personalized knowledge from the past. By facilitating mutual learning between the current and historical local models, the local model can effectively integrate both the current generalized knowledge and the historical personalized knowledge. Extensive experiments showcase the superiority of our method compared to state-of-the-art methods.
Bochao Liu, Weijia Guo, Shiming Ge
IJCNN5
2024 Private Gradient Estimation is Useful for Generative Modeling
abstract
While generative models have proved successful in many domains, they may pose a privacy leakage risk in practical deployment. To address this issue, differentially private generative model learning has emerged as a solution to train private generative models for different downstream tasks. However, existing private generative modeling approaches face significant challenges in generating high-dimensional data due to the inherent complexity involved in modeling such data. In this work, we present a new private generative modeling approach where samples are generated via Hamiltonian dynamics with gradients of the private dataset estimated by a well-trained network. In the approach, we achieve differential privacy by perturbing the projection vectors in the estimation of gradients with sliced score matching. In addition, we enhance the reconstruction ability of the model by incorporating a residual enhancement module during the score matching. For sampling, we perform Hamiltonian dynamics with gradients estimated by the well-trained network, allowing the sampled data close to the private dataset's manifold step by step. In this way, our model is able to generate data with a resolution of 256×256. Extensive experiments and analysis clearly demonstrate the effectiveness and rationality of the proposed approach.
Bochao Liu, Weijia Guo, Liansheng Zhuang, Weiping Wang 0005, Shiming Ge
ACM Multimedia7
2024 DB-FSCIL: Few-Shot Class-Incremental Learning Using Dual Bridges
Kexin Bao, Fanzhao Lin, Ruyue Liu, Shiming Ge
PRICAI (3)4
2024 Towards Personalized Federated Learning via Comprehensive Knowledge Distillation
abstract
Federated learning is a distributed machine learning paradigm designed to protect data privacy. However, data heterogeneity across various clients results in catastrophic forget-ting, where the model rapidly forgets previous knowledge while acquiring new knowledge. To address this challenge, personalized federated learning has emerged to customize a personalized model for each client. However, the inherent limitation of this mechanism is its excessive focus on personalization, potentially hindering the generalization of those models. In this paper, we present a novel personalized federated learning method that uses global and historical models as teachers and the local model as the student to facilitate comprehensive knowledge distillation. The historical model represents the local model from the last round of client training, containing historical personalized knowledge, while the global model represents the aggregated model from the last round of server aggregation, containing global generalized knowledge. By applying knowledge distillation, we effectively transfer global generalized knowledge and historical personalized knowledge to the local model, thus mitigating catastrophic forgetting and enhancing the general performance of personalized models. Extensive experimental results demonstrate the significant advantages of our method.
Bochao Liu, Weijia Guo, Shiming Ge
SMC5
2024 Transferring Annotator- and Instance-Dependent Transition Matrix for Learning From Crowds
abstract
Learning from crowds describes that the annotations of training data are obtained with crowd-sourcing services. Multiple annotators each complete their own small part of the annotations, where labeling mistakes that depend on annotators occur frequently. Modeling the label-noise generation process by the noise transition matrix is a powerful tool to tackle the label noise. In real-world crowd-sourcing scenarios, noise transition matrices are both annotator- and instance-dependent. However, due to the high complexity of annotator- and instance-dependent transition matrices (AIDTM), annotation sparsity, which means each annotator only labels a tiny part of instances, makes modeling AIDTM very challenging. Without prior knowledge, existing works simplify the problem by assuming the transition matrix is instance-independent or using simple parametric ways, which lose modeling generality. Motivated by this, we target a more realistic problem, estimating general AIDTM in practice. Without losing modeling generality, we parameterize AIDTM with deep neural networks. To alleviate the modeling challenge, we suppose every annotator shares its noise pattern with similar annotators, and estimate AIDTM via knowledge transfer. We hence first model the mixture of noise patterns by all annotators, and then transfer this modeling to individual annotators. Furthermore, considering that the transfer from the mixture of noise patterns to individuals may cause two annotators with highly different noise generations to perturb each other, we employ the knowledge transfer between identified neighboring annotators to calibrate the modeling. Theoretical analyses are derived to demonstrate that both the knowledge transfer from global to individuals and the knowledge transfer between neighboring individuals can effectively help mitigate the challenge of modeling general AIDTM. Experiments confirm the superiority of the proposed approach on synthetic and real-world crowd-sourcing data.
Shikun Li, Xiaobo Xia, Jiankang Deng, Shiming Ge, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Low-Resolution Object Recognition With Cross-Resolution Relational Contrastive Distillation
abstract
Recognizing objects in low-resolution images is a challenging task due to the lack of informative details. Recent studies have shown that knowledge distillation approaches can effectively transfer knowledge from a high-resolution teacher model to a low-resolution student model by aligning cross-resolution representations. However, these approaches still face limitations in adapting to the situation where the recognized objects exhibit significant representation discrepancies between training and testing images. In this study, we propose a cross-resolution relational contrastive distillation approach to facilitate low-resolution object recognition. Our approach enables the student model to mimic the behavior of a well-trained teacher model which delivers high accuracy in identifying high-resolution objects. To extract sufficient knowledge, the student learning is supervised with contrastive relational distillation loss, which preserves the similarities in various relational structures in contrastive representation space. In this manner, the capability of recovering missing details of familiar low-resolution objects can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on low-resolution object classification and low-resolution face recognition clearly demonstrate the effectiveness and adaptability of our approach.
Kangkai Zhang, Shiming Ge, Ruixin Shi, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learning Contrast-Enhanced Shape-Biased Representations for Infrared Small Target Detection
abstract
Detecting infrared small targets under cluttered background is mainly challenged by dim textures, low contrast and varying shapes. This paper proposes an approach to facilitate infrared small target detection by learning contrast-enhanced shape-biased representations. The approach cascades a contrast-shape encoder and a shape-reconstructable decoder to learn discriminative representations that can effectively identify target objects. The contrast-shape encoder applies a stem of central difference convolutions and a few large-kernel convolutions to extract shape-preserving features from input infrared images. This specific design in convolutions can effectively overcome the challenges of low contrast and varying shapes in a unified way. Meanwhile, the shape-reconstructable decoder accepts the edge map of input infrared image and is learned by simultaneously optimizing two shape-related consistencies: the internal one decodes the encoder representations by upsampling reconstruction and constraints segmentation consistency, whilst the external one cascades three gated ResNet blocks to hierarchically fuse edge maps and decoder representations and constrains contour consistency. This decoding way can bypass the challenge of dim texture and varying shapes. In our approach, the encoder and decoder are learned in an end-to-end manner, and the resulting shape-biased encoder representations are suitable for identifying infrared small targets. Extensive experimental evaluations are conducted on public benchmarks and the results demonstrate the effectiveness of our approach.
Fanzhao Lin, Kexin Bao, Dan Zeng 0001, Shiming Ge
IEEE Trans. Image Process.5
2024 Multimodal Composition Example Mining for Composed Query Image Retrieval
abstract
Composed query image retrieval task aims to retrieve the target image in the database by a query that composes two different modalities: a reference image and a sentence declaring that some details of the reference image need to be modified and replaced by new elements. Tackling this task needs to learn a multimodal embedding space, which can make semantically similar targets and queries close but dissimilar targets and queries as far away as possible. Most of the existing methods start from the perspective of model structure and design some clever interactive modules to promote the better fusion and embedding of different modalities. However, their learning objectives use conventional query-level examples as negatives while neglecting the composed query's multimodal characteristics, leading to the inadequate utilization of the training data and suboptimal construction of metric space. To this end, in this paper, we propose to improve the learning objective by constructing and mining hard negative examples from the perspective of multimodal fusion. Specifically, we compose the reference image and its logically unpaired sentences rather than paired ones to create component-level negative examples to better use data and enhance the optimization of metric space. In addition, we further propose a new sentence augmentation method to generate more indistinguishable multimodal negative examples from the element level and help the model learn a better metric space. Massive comparison experiments on four real-world datasets confirm the effectiveness of the proposed method.
Gangjian Zhang, Shikun Li, Shikui Wei, Shiming Ge, Na Cai, Yao Zhao 0001
IEEE Trans. Image Process.4
2024 Learning Shape-Biased Representations for Infrared Small Target Detection
abstract
Typically, infrared small target detection aims to accurately localize objects from complex backgrounds where the object textures are often dim and the object shapes are varying. A feasible solution is learning discriminative representations with deep convolutional neural networks (CNNs). However, the representations learned by traditional deep CNNs often suffer from low shape bias. In this work, we propose a unified framework to learn shape-biased representations for facilitating infrared small target detection by explicitly incorporating shape information into model learning. The framework cascades a large-kernel encoder and a shape-guided decoder to learn discriminative shape-biased representations in an end-to-end manner. The large-kernel encoder describes infrared images into shape-preserving representations by using a few convolutions whose kernel size is as large as$9\times 9$, in contrast to commonly used$3\times 3$. The shape-guided decoder simultaneously addresses two tasks: decodes the encoder representations via upsampling reconstruction to reconstruct the segmentation, and hierarchically fuses the decoder representations and edge information via cascaded gated ResNet blocks to reconstruct the contour. In this way, the learned shape-biased representations are effective for identifying infrared small targets. Extensive experiments show our approach outperforms 18 state-of-the-arts.
Fanzhao Lin, Shiming Ge, Kexin Bao, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Multim.2
2023 Bootstrapping Multi-View Representations for Fake News Detection
abstract
Previous researches on multimedia fake news detection include a series of complex feature extraction and fusion networks to gather useful information from the news. However, how cross-modal consistency relates to the fidelity of news and how features from different modalities affect the decision-making are still open questions. This paper presents a novel scheme of Bootstrapping Multi-view Representations (BMR) for fake news detection. Given a multi-modal news, we extract representations respectively from the views of the text, the image pattern and the image semantics. Improved Multi-gate Mixture-of-Expert networks (iMMoE) are proposed for feature refinement and fusion. Representations from each view are separately used to coarsely predict the fidelity of the whole news, and the multimodal representations are able to predict the cross-modal consistency. With the prediction scores, we reweigh each view of the representations and bootstrap them for fake news detection. Extensive experiments conducted on typical fake news detection datasets prove that BMR outperforms state-of-the-art schemes.
Qichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian, Dan Zeng 0001, Shiming Ge
AAAI6
2023 Model Conversion via Differentially Private Data-Free Distillation
abstract
While massive valuable deep models trained on large-scale data have been released to facilitate the artificial intelligence community, they may encounter attacks in deployment which leads to privacy leakage of training data. In this work, we propose a learning approach termed differentially private data-free distillation (DPDFD) for model conversion that can convert a pretrained model (teacher) into its privacy-preserving counterpart (student) via an intermediate generator without access to training data. The learning collaborates three parties in a unified way. First, massive synthetic data are generated with the generator. Then, they are fed into the teacher and student to compute differentially private gradients by normalizing the gradients and adding noise before performing descent. Finally, the student is updated with these differentially private gradients and the generator is updated by taking the student as a fixed discriminator in an alternate manner. In addition to a privacy-preserving student, the generator can generate synthetic data in a differentially private way for other down-stream tasks. We theoretically prove that our approach can guarantee differential privacy and well convergence. Extensive experiments that significantly outperform other differentially private generative approaches demonstrate the effectiveness of our approach.
Bochao Liu, Shikun Li, Dan Zeng 0001, Shiming Ge
IJCAI5
2023 Learning Generalized Representations for Open-Set Temporal Action Localization
abstract
Open-set Temporal Action Localization (OSTAL) is a critical and challenging task that aims to recognize and temporally localize human actions in untrimmed videos in open word scenarios. The main challenge in this task is the knowledge transfer from known actions to unknown actions. However, existing methods utilize limited training data and overparameterized deep neural network, which have poor generalization. This paper proposes a novel Generalized OSTAL model (namely GOTAL) to learn generalized representations of actions. GOTAL utilizes a Transformer network to model actions and a open-set detection head to perform action localization and recognition. Benefitting from Transformer's temporal modeling capabilities, GOTAL facilitates the extraction of human motion information from videos to mitigate the effects of irrelevant background data. Furthermore, a sharpness minimization algorithm is used to learn the network parameters of GOTAL, which facilitates the convergence of network parameters towards flatter minima by simultaneously minimizing the training loss value and sharpness of the loss plane. The collaboration of the above components significantly enhances the generalization of the representation. Experimental results demonstrate that GOTAL achieves the state-of-the-art performance on THUMOS14 and ActivityNet1.3 benchmarks, confirming the effectiveness of our proposed method.
Junshan Hu, Liansheng Zhuang, Weisong Dong, Shiming Ge, Shafei Wang
ACM Multimedia4
2023 Federated Learning with Label-Masking Distillation
abstract
Federated learning provides a privacy-preserving manner to collaboratively train models on data distributed over multiple local clients via the coordination of a global server. In this paper, we focus on label distribution skew in federated learning, where due to the different user behavior of the client, label distributions between different clients are significantly different. When faced with such cases, most existing methods will lead to a suboptimal optimization due to the inadequate utilization of label distribution information in clients. Inspired by this, we propose a label-masking distillation approach termed FedLMD to facilitate federated learning via perceiving the various label distributions of each client. We classify the labels into majority and minority labels based on the number of examples per class during training. The client model learns the knowledge of majority labels from local data. The process of distillation masks out the predictions of majority labels from the global model, so that it can focus more on preserving the minority label knowledge of the client. A series of experiments show that the proposed approach can achieve state-of-the-art performance in various cases. Moreover, considering the limited resources of the clients, we propose a variant FedLMD-Tf that does not require an additional teacher, which outperforms previous lightweight approaches without increasing computational costs. Our code is available at https://github.com/wnma3mz/FedLMD.
Jianghu Lu, Shikun Li, Kexin Bao, Zhenxing Qian, Shiming Ge
ACM Multimedia6
2023 Personalized Federated Learning via Backbone Self-Distillation
abstract
In practical scenarios, federated learning frequently necessitates training personalized models for each client using heterogeneous data. This paper proposes a backbone self-distillation approach to facilitate personalized federated learning. In this approach, each client trains its local model and only sends the backbone weights to the server. These weights are then aggregated to create a global backbone, which is returned to each client for updating. However, the client’s local backbone lacks personalization because of the common representation. To solve this problem, each client further performs backbone self-distillation by using the global backbone as a teacher and transferring knowledge to update the local backbone. This process involves learning two components: the shared backbone for common representation and the private head for local personalization, which enables effective global knowledge transfer. Extensive experiments and comparisons with 12 state-of-the-art approaches demonstrate the effectiveness of our approach.
Bochao Liu, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
MMAsia5
2023 WebUAV-3M: A Benchmark for Unveiling the Power of Million-Scale Deep UAV Tracking
abstract
Unmanned aerial vehicle (UAV) tracking is of great significance for a wide range of applications, such as delivery and agriculture. Previous benchmarks in this area mainly focused on small-scale tracking problems while ignoring the amounts of data, types of data modalities, diversities of target categories and scenarios, and evaluation protocols involved, greatly hiding the massive power of deep UAV tracking. In this article, we propose WebUAV-3M, the largest public UAV tracking benchmark to date, to facilitate both the development and evaluation of deep UAV trackers. WebUAV-3M contains over 3.3 million frames across 4,500 videos and offers 223 highly diverse target categories. Each video is densely annotated with bounding boxes by an efficient and scalable semi-automatic target annotation (SATA) pipeline. Importantly, to take advantage of the complementary superiority of language and audio, we enrich WebUAV-3M by innovatively providing both natural language specifications and audio descriptions. We believe that such additions will greatly boost future research in terms of exploring language features and audio cues for multi-modal UAV tracking. In addition, a fine-grained UAV tracking-under-scenario constraint (UTUSC) evaluation protocol and seven challenging scenario subtest sets are constructed to enable the community to develop, adapt and evaluate various types of advanced trackers. We provide extensive evaluations and detailed analyses of 43 representative trackers and envision future research directions in the field of deep UAV tracking and beyond. The dataset, toolkits, and baseline results are available at https://github.com/983632847/WebUAV-3M.
Chunhui Zhang 0001, Guanjie Huang, Li Liu 0036, Shiming Ge, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Learning Privacy-Preserving Student Networks via Discriminative-Generative Distillation
abstract
While deep models have proved successful in learning rich knowledge from massive well-annotated data, they may pose a privacy leakage risk in practical deployment. It is necessary to find an effective trade-off between high utility and strong privacy. In this work, we propose a discriminative-generative distillation approach to learn privacy-preserving deep models. Our key idea is taking models as bridge to distill knowledge from private data and then transfer it to learn a student network via two streams. First, discriminative stream trains a baseline classifier on private data and an ensemble of teachers on multiple disjoint private subsets, respectively. Then, generative stream takes the classifier as a fixed discriminator and trains a generator in a data-free manner. After that, the generator is used to generate massive synthetic data which are further applied to train a variational autoencoder (VAE). Among these synthetic data, a few of them are fed into the teacher ensemble to query labels via differentially private aggregation, while most of them are embedded to the trained VAE for reconstructing synthetic data. Finally, a semi-supervised student learning is performed to simultaneously handle two tasks: knowledge transfer from the teachers with distillation on few privately labeled synthetic data, and knowledge enhancement with tangent-normal adversarial regularization on many triples of reconstructed synthetic data. In this way, our approach can control query cost over private data and mitigate accuracy degradation in a unified manner, leading to a privacy-preserving student model. Extensive experiments and analysis clearly show the effectiveness of the proposed approach.
Shiming Ge, Bochao Liu, Dan Zeng 0001
IEEE Trans. Image Process.1
2023 Learning Patch-Channel Correspondence for Interpretable Face Forgery Detection
abstract
Beyond high accuracy, good interpretability is very critical to deploy a face forgery detection model for visual content analysis. In this paper, we propose learning patch-channel correspondence to facilitate interpretable face forgery detection. Patch-channel correspondence aims to transform the latent features of a facial image into multi-channel interpretable features where each channel mainly encoders a corresponding facial patch. Towards this end, our approach embeds a feature reorganization layer into a deep neural network and simultaneously optimizes classification task and correspondence task via alternate optimization. The correspondence task accepts multiple zero-padding facial patch images and represents them into channel-aware interpretable representations. The task is solved by step-wisely learning channel-wise decorrelation and patch-channel alignment. Channel-wise decorrelation decouples latent features for class-specific discriminative channels to reduce feature complexity and channel correlation, while patch-channel alignment then models the pairwise correspondence between feature channels and facial patches. In this way, the learned model can automatically discover corresponding salient features associated to potential forgery regions during inference, providing discriminative localization of visualized evidences for face forgery detection while maintaining high detection accuracy. Extensive experiments on popular benchmarks clearly demonstrate the effectiveness of the proposed approach in interpreting face forgery detection without sacrificing accuracy. The source code is available at https://github.com/Jae35/IFFD.
Yingying Hua, Ruixin Shi, Shiming Ge
IEEE Trans. Image Process.4
2023 Trustable Co-Label Learning From Multiple Noisy Annotators
abstract
Supervised deep learning depends on massive accurately annotated examples, which is usually impractical in many real-world scenarios. A typical alternative is learning from multiple noisy annotators. Numerous earlier works assume that all labels are noisy, while it is usually the case that a few trusted samples with clean labels are available. This raises the following important question: how can we effectively use a small amount of trusted data to facilitate robust classifier learning from multiple annotators? This paper proposes a data-efficient approach, calledTrustable Co-label Learning(TCL), to learn deep classifiers from multiple noisy annotators when a small set of trusted data is available. This approach follows the coupled-view learning manner, which jointly learns the data classifier and the label aggregator. It effectively uses trusted data as a guide to generate trustable soft labels (termed co-labels). A co-label learning can then be performed by alternately reannotating the pseudo labels and refining the classifiers. In addition, we further improve TCL for a special complete data case, where each instance is labeled by all annotators and the label aggregator is represented by multilayer neural networks to enhance model capacity. Extensive experiments on synthetic and real datasets clearly demonstrate the effectiveness and robustness of the proposed approach. Source code is available athttps://github.com/ShikunLi/TCL.
Shikun Li, Tongliang Liu, Jiyong Tan, Dan Zeng 0001, Shiming Ge
IEEE Trans. Multim.5
2022 Selective-Supervised Contrastive Learning with Noisy Labels
abstract
Deep networks have strong capacities of embedding data into latent representations and finishing following tasks. However, the capacities largely come from high-quality annotated labels, which are expensive to collect. Noisy labels are more affordable, but result in corrupted representations, leading to poor generalization performance. To learn robust representations and handle noisy labels, we propose selective-supervised contrastive learning (Sel-CL) in this paper. Specifically, Sel-CL extend supervised contrastive learning (Sup-CL), which is powerful in representation learning, but is degraded when there are noisy labels. Sel-CL tackles the direct cause of the problem of Sup-CL. That is, as Sup-CL works in a pair-wise manner, noisy pairs built by noisy labels mislead representation learning. To alleviate the issue, we select confident pairs out of noisy ones for Sup-CL without knowing noise rates. In the selection process, by measuring the agreement between learned representations and given labels, we first identify confident examples that are exploited to build confident pairs. Then, the representation similarity distribution in the built confident pairs is exploited to identify more confident pairs out of noisy pairs. All obtained confident pairs are finally used for Sup-CL to enhance representations. Experiments on multiple noisy datasets demonstrate the robustness of the learned representations by our method, following the state-of-the-art performance. Source codes are available at https://github.com/ShikunLi/Sel-Cl.
Shikun Li, Xiaobo Xia, Shiming Ge, Tongliang Liu
CVPR3
2022 Regularized Latent Space Exploration for Discriminative Face Super-Resolution
abstract
Learning face super-resolution models is challenged in many practical scenarios where high-resolution and low-resolution face pairs usually are difficult to collect for training examples. Recent self-supervised approach provides a feasible solution by using low-resolution faces to guide the generation of the corresponding high-resolution ones with a pretrained generator. In this paper, we propose a regularized latent space exploration approach to facilitate self-supervised face super-resolution. In the approach, a pretrained generative adversarial network (GAN) is fully used to control the exploration of high-resolution face generation in an iterative optimization manner for a low-resolution face. During the iteration, super-resolution faces are continually generated from a feasible latent space by the generator and evaluated by the discriminator, while the generator is online finetuned. The generation is evaluated by measuring the semantic loss as well as pixel loss between ground-truth low-resolution faces and the corresponding downsampled super-resolution faces. In this way, the generated faces can be appearance natural and semantic discriminative. Experiments validate the effectiveness of our approach in terms of quantitative metrics and visual quality.
Ruixin Shi, Junzheng Zhang, Shiming Ge
ICASSP4
2022 Robust Weight Perturbation for Adversarial Training
abstract
Overfitting widely exists in adversarial robust training of deep networks. An effective remedy is adversarial weight perturbation, which injects the worst-case weight perturbation during network training by maximizing the classification loss on adversarial examples. Adversarial weight perturbation helps reduce the robust generalization gap; however, it also undermines the robustness improvement. A criterion that regulates the weight perturbation is therefore crucial for adversarial training. In this paper, we propose such a criterion, namely Loss Stationary Condition (LSC) for constrained perturbation. With LSC, we find that it is essential to conduct weight perturbation on adversarial data with small classification loss to eliminate robust overfitting. Weight perturbation on adversarial data with large classification loss is not necessary and may even lead to poor robustness. Based on these observations, we propose a robust perturbation strategy to constrain the extent of weight perturbation. The perturbation strategy prevents deep networks from overfitting while avoiding the side effect of excessive weight perturbation, significantly improving the robustness of adversarial training. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-art adversarial training methods.
Chaojian Yu, Bo Han 0003, Mingming Gong, Li Shen 0008, Shiming Ge, Bo Du 0001, Tongliang Liu
IJCAI5
2022 Deepfake Video Detection with Spatiotemporal Dropout Transformer
abstract
While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches often focus on single frame and ignore the spatiotemporal cues hidden in deepfake videos, resulting in poor generalization and robustness. The key of a video-level detector is to fully exploit the spatiotemporal inconsistency distributed in local facial regions across different frames in deepfake videos. Inspired by that, this paper proposes a simple yet effective patch-level approach to facilitate deepfake video detection via spatiotemporal dropout transformer. The approach reorganizes each input video into bag of patches that is then fed into a vision transformer to achieve robust representation. Specifically, a spatiotemporal dropout operation is proposed to fully explore patch-level spatiotemporal cues and serve as effective data augmentation to further enhance model's robustness and generalization ability. The operation is flexible and can be easily plugged into existing vision transformers. Extensive experiments demonstrate the effectiveness of our approach against 25 state-of-the-arts with impressive robustness, generalizability, and representation ability.
Daichi Zhang, Fanzhao Lin, Yingying Hua, Dan Zeng 0001, Shiming Ge
ACM Multimedia6
2022 Privacy-Preserving Student Learning with Differentially Private Data-Free Distillation
abstract
Deep learning models can achieve high inference accuracy by extracting rich knowledge from massive well-annotated data, but may pose the risk of data privacy leakage in practical deployment. In this paper, we present an effective teacher-student learning approach to train privacy-preserving deep learning models via differentially private data-free distillation. The main idea is generating synthetic data to learn a student that can mimic the ability of a teacher well-trained on private data. In the approach, a generator is first pretrained in a data-free manner by incorporating the teacher as a fixed discriminator. With the generator, massive synthetic data can be generated for model training without exposing data privacy. Then, the synthetic data is fed into the teacher to generate private labels. Towards this end, we propose a label differential privacy algorithm termed selective randomized response to protect the label information. Finally, a student is trained on the synthetic data with the supervision of private labels. In this way, both data privacy and label privacy are well protected in a unified framework, leading to privacy-preserving models. Extensive experiments and analysis clearly demonstrate the effectiveness of our approach.
Bochao Liu, Jianghu Lu, Junjie Zhang 0002, Dan Zeng 0001, Zhenxing Qian, Shiming Ge
MMSP7
2022 Estimating Noise Transition Matrix with Label Correlations for Noisy Multi-Label Learning
abstract
In label-noise learning, the noise transition matrix, bridging the class posterior for noisy and clean data, has been widely exploited to learn statistically consistent classifiers. The effectiveness of these algorithms relies heavily on estimating the transition matrix. Recently, the problem of label-noise learning in multi-label classification has received increasing attention, and these consistent algorithms can be applied in multi-label cases. However, the estimation of transition matrices in noisy multi-label learning has not been studied and remains challenging, since most of the existing estimators in noisy multi-class learning depend on the existence of anchor points and the accurate fitting of noisy class posterior. To address this problem, in this paper, we first study the identifiability problem of the class-dependent transition matrix in noisy multi-label learning, and then inspired by the identifiability results, we propose a new estimator by exploiting label correlations without neither anchor points nor accurate fitting of noisy class posterior. Specifically, we estimate the occurrence probability of two noisy labels to get noisy label correlations. Then, we perform sample selection to further extract information that implies clean label correlations, which is used to estimate the occurrence probability of one noisy label when a certain clean label appears. By utilizing the mismatch of label correlations implied in these occurrence probabilities, the transition matrix is identifiable, and can then be acquired by solving a simple bilinear decomposition problem. Empirical results demonstrate the effectiveness of our estimator to estimate the transition matrix with label correlations, leading to better classification performance. Source codes are available at https://github.com/tmllab/Multi-Label-T.
Shikun Li, Xiaobo Xia, Hansong Zhang 0003, Yibing Zhan, Shiming Ge, Tongliang Liu
NeurIPS5
2022 Structure injected weight normalization for training deep networks
Xu Yuan 0008, Xiangjun Shen, Sumet Mehta, Teng Li 0001, Shiming Ge, Zhengjun Zha
Multim. Syst.5
2022 Efficient Video Grounding With Which-Where Reading Comprehension
abstract
Video grounding aims at localizing the temporal moment related to the given language description, which is very helpful to many cross-modal content understanding applications like visual question answering and sentence-video search. Existing approaches usually directly regress the temporal boundaries of an event described by a query sentence in the video sequence. This direct regression manner often encounters a large decision space due to diverse target events and variable video durations, leading to inaccurate localization as well as inefficient grounding. This paper presents an efficient framework termed from which to where to facilitate video grounding. The core idea is imitating the reading comprehension process to gradually narrow the decision space, in what we decompose the direct regression into two steps. The “which” step first roughly selects a candidate area by evaluating which video segment in the predefined set is closest to the ground truth. To this end, we formulate this step into a multi-choice reading comprehension problem and propose a criterion to select the best-matched segment. In this way, the excessive decision space is effectively reduced. The “where” step aims to precisely regress the temporal boundary of the selected video segment from the shrunk decision space. We thus introduce a triple-span representation for each candidate video segment to use the regional context for better boundary regression. The “which” and “where” steps can be combined into a unified framework and learned end-to-end, leading to an efficient video grounding system. Extensive experiments on Charades-STA, ActivityNet-Captions, and TACoS benchmarks clearly demonstrate the effectiveness of our framework.
Jialin Gao, Xin Sun 0020, Bernard Ghanem, Xi Zhou 0001, Shiming Ge
IEEE Trans. Circuits Syst. Video Technol.5
2022 Student Network Learning via Evolutionary Knowledge Distillation
abstract
Knowledge distillation provides an effective way to transfer knowledge via teacher-student learning, where most existing distillation approaches apply a fixed pre-trained model as teacher to supervise the learning of student network. This manner usually brings in a big capability gap between teacher and student networks during learning. Recent researches have observed that a small teacher-student capability gap can facilitate knowledge transfer. Inspired by that, we propose an evolutionary knowledge distillation approach to improve the transfer effectiveness of teacher knowledge. Instead of a fixed pre-trained teacher, an evolutionary teacher is learned online and consistently transfers intermediate knowledge to supervise student network learning on-the-fly. To enhance intermediate knowledge representation and mimicking, several simple guided modules are introduced between corresponding teacher-student blocks. In this way, the student can simultaneously obtain rich internal knowledge and capture its growth process, leading to effective student network learning. Extensive experiments clearly demonstrate the effectiveness of our approach as well as good adaptability in the low-resolution and few-sample scenarios.
Kangkai Zhang, Chunhui Zhang 0001, Shikun Li, Dan Zeng 0001, Shiming Ge
IEEE Trans. Circuits Syst. Video Technol.5
2022 Deepfake Video Detection via Predictive Representation Learning
abstract
Increasingly advanced deepfake approaches have made the detection of deepfake videos very challenging. We observe that the general deepfake videos often exhibit appearance-level temporal inconsistencies in some facial components between frames, resulting in discriminative spatiotemporal latent patterns among semantic-level feature maps. Inspired by this finding, we propose a predictive representative learning approach termed Latent Pattern Sensing to capture these semantic change characteristics for deepfake video detection. The approach cascades a Convolution Neural Network-based encoder, a ConvGRU-based aggregator, and a single-layer binary classifier. The encoder and aggregator are pretrained in a self-supervised manner to form the representative spatiotemporal context features. Then, the classifier is trained to classify the context features, distinguishing fake videos from real ones. Finally, we propose a selective self-distillation fine-tuning method to further improve the robustness and performance of the detector. In this manner, the extracted features can simultaneously describe the latent patterns of videos across frames spatially and temporally in a unified way, leading to an effective and robust deepfake video detector. Extensive experiments and comprehensive analysis prove the effectiveness of our approach, e.g., achieving a very highest Area Under Curve (AUC) score of 99.94% on FaceForensics++ benchmark and surpassing 12 states of the art at least 7.90%@AUC and 8.69%@AUC on challenging DFDC and Celeb-DF(v2) benchmarks, respectively.
Shiming Ge, Fanzhao Lin, Chenyu Li 0001, Daichi Zhang, Weiping Wang 0005, Dan Zeng 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Double-Blinded Finder: a two-side secure children face recognition system
Xin Jin 0015, Jicheng Lei, Shiming Ge, Chenggen Song, Chuanqiang Wu
Wirel. Networks3
2021 Interpret The Predictions Of Deep Networks Via Re-Label Distillation
abstract
Interpreting the predictions of a black-box deep network can facilitate the reliability of its deployment. In this work, we propose a re-label distillation approach to learn a direct map from the input to the prediction in a self-supervision manner. The image is projected into a VAE subspace to generate some synthetic images by randomly perturbing its latent vector. Then, these synthetic images can be annotated into one of two classes by identifying whether their labels shift. After that, using the labels annotated by the deep network as teacher, a linear student model is trained to approximate the annotations by mapping these synthetic images to the classes. In this manner, these re-labeled synthetic images can well describe the local classification mechanism of the deep network, and the learned student can provide a more intuitive explanation towards the predictions. Extensive experiments verify the effectiveness of our approach qualitatively and quantitatively.
Yingying Hua, Shiming Ge, Daichi Zhang
ICME2
2021 Detecting Deepfake Videos with Temporal Dropout 3DCNN
abstract
While the abuse of deepfake technology has brought about a serious impact on human society, the detection of deepfake videos is still very challenging due to their highly photorealistic synthesis on each frame. To address that, this paper aims to leverage the possible inconsistent cues among video frames and proposes a Temporal Dropout 3-Dimensional Convolutional Neural Network (TD-3DCNN) to detect deepfake videos. In the approach, the fixed-length frame volumes sampled from a video are fed into a 3-Dimensional Convolutional Neural Network (3DCNN) to extract features across different scales and identified whether they are real or fake. Especially, a temporal dropout operation is introduced to randomly sample frames in each batch. It serves as a simple yet effective data augmentation and can enhance the representation and generalization ability, avoiding model overfitting and improving detecting accuracy. In this way, the resulting video-level classifier is accurate and effective to identify deepfake videos. Extensive experiments on benchmarks including Celeb-DF(v2) and DFDC clearly demonstrate the effectiveness and generalization capacity of our approach.
Daichi Zhang, Chenyu Li 0001, Fanzhao Lin, Dan Zeng 0001, Shiming Ge
IJCAI5
2021 Latent Pattern Sensing: Deepfake Video Detection via Predictive Representation Learning
abstract
Increasingly advanced deepfake approaches have made the detection of deepfake videos very challenging. We observe that the general deepfake videos often exhibit appearance-level temporal inconsistencies in some facial components between frames, resulting in discriminable spatiotemporal latent patterns among semantic-level feature maps. Inspired by this finding, we propose a predictive representative learning approach termed Latent Pattern Sensing to capture these semantic change characteristics for deepfake video detection. The approach cascades a CNN-based encoder, a ConvGRU-based aggregator and a single-layer binary classifier. The encoder and aggregator are pre-trained in a self-supervised manner to form the representative spatiotemporal context features. Finally, the classifier is trained to classify the context features, distinguishing fake videos from real ones. In this manner, the extracted features can simultaneously describe the latent patterns of videos across frames spatially and temporally in a unified way, leading to an effective deepfake video detector. Extensive experiments prove our approach’s effectiveness, e.g., surpassing 10 state-of-the-arts at least 7.92%@AUC on challenging Celeb-DF(v2) benchmark.
Shiming Ge, Fanzhao Lin, Chenyu Li 0001, Daichi Zhang, Jiyong Tan, Weiping Wang 0005, Dan Zeng 0001
MMAsia1
2021 Differentially Private Learning with Grouped Gradient Clipping
abstract
While deep learning has proved success in many critical tasks by training models from large-scale data, some private information within can be recovered from the released models, leading to the leakage of privacy. To address this problem, this paper presents a differentially private deep learning paradigm to train private models. In the approach, we propose and incorporate a simple operation termed grouped gradient clipping to modulate the gradient weights. We also incorporated the smooth sensitivity mechanism into differentially private deep learning paradigm, which bounds the adding Gaussian noise. In this way, the resulting model can simultaneously provide with strong privacy protection and avoid accuracy degradation, providing a good trade-off between privacy and performance. The theoretic advantages of grouped gradient clipping are well analyzed. Extensive evaluations on popular benchmarks and comparisons with 11 state-of-the-arts clearly demonstrate the effectiveness and genearalizability of our approach.
Chenyu Li 0001, Bochao Liu, Shiming Ge, Weiping Wang 0005
MMAsia5
2021 Skeleton-Based Action Recognition With Focusing-Diffusion Graph Convolutional Networks
abstract
Graph Convolutional Networks have been successfully applied in skeleton-based action recognition. The key is fully exploring the spatial-temporal context. This letter proposes a Focusing-Diffusion Graph Convolutional Network (FDGCN) to address this issue. Each skeleton frame is first decomposed into two opposite-direction graphs for subsequent focusing and diffusion processes. Next, the focusing process generates a spatial-level representation for each frame individually by an attention module. This representation is regarded as a supernode to aggregate the feature from each joint node in each frame for spatial context extraction. After generating supernodes for the entire sequence, a transformer encoder layer is proposed to capture the temporal context further. Finally, these supernodes pass the embedded spatial-temporal context back to the spatial joints through the diffusion graph in the diffusing process. Extensive experiments on the NTU RGB+D and Skeleton-Kinetics benchmarks demonstrate the effectiveness of our approach.
Jialin Gao, Tong He 0002, Xi Zhou 0001, Shiming Ge
IEEE Signal Process. Lett.4
2021 Self-Guided Body Part Alignment With Relation Transformers for Occluded Person Re-Identification
abstract
Person re-identification in the wild is often challenged by occlusion. Existing methods mainly rely on learned external cues like pose or parsing to ease occlusion distraction. This knowledge highly related to body semantics may introduce alignment effects, leading to additional requirements for dedicated training data and inference computation. We propose the Self-guided Body Part Alignment method that learns cue-free semantic-aligned local prediction for feature representations to avoid high-cost dependence on external cues. First, scale-wise global spatial attention is utilized to determine essential body parts automatically. A relation transformer network is then employed to predict semantic-aligned local parts, guided with anchored global information by constraint loss. Similarity metrics for all parts are merged with threshold conditions to filter invisible body parts comprehensively. Experimental results on occluded and holistic person reID benchmarks show the proposed method outperforms other cue-relied and cue-free methods. As far as we know, this is the first method that applies transformer networks on local predictions for occluded reID tasks.
Guanshuo Wang, Jialin Gao, Xi Zhou 0001, Shiming Ge
IEEE Signal Process. Lett.5
2021 Cascaded Correlation Refinement for Robust Deep Tracking
abstract
Recent deep trackers have shown superior performance in visual tracking. In this article, we propose a cascaded correlation refinement approach to facilitate the robustness of deep tracking. The core idea is to address accurate target localization and reliable model update in a collaborative way. To this end, our approach cascades multiple stages of correlation refinement to progressively refine target localization. Thus, the localized object could be used to learn an accurate on-the-fly model for improving the reliability of model update. Meanwhile, we introduce an explicit measure to identify the tracking failure and then leverage a simple yet effective look-back scheme to adaptively incorporate the initial model and on-the-fly model to update the tracking model. As a result, the tracking model can be used to localize the target more accurately. Extensive experiments on OTB2013, OTB2015, VOT2016, VOT2018, UAV123, and GOT-10k demonstrate that the proposed tracker achieves the best robustness against the state of the arts.
Shiming Ge, Chunhui Zhang 0001, Shikun Li, Dan Zeng 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2020 Ultrafast Video Attention Prediction with Coupled Knowledge Distillation
abstract
Large convolutional neural network models have recently demonstrated impressive performance on video attention prediction. Conventionally, these models are with intensive computation and large memory. To address these issues, we design an extremely light-weight network with ultrafast speed, named UVA-Net. The network is constructed based on depth-wise convolutions and takes low-resolution images as input. However, this straight-forward acceleration method will decrease performance dramatically. To this end, we propose a coupled knowledge distillation strategy to augment and train the network effectively. With this strategy, the model can further automatically discover and emphasize implicit useful cues contained in the data. Both spatial and temporal knowledge learned by the high-resolution complex teacher networks also can be distilled and transferred into the proposed low-resolution light-weight spatiotemporal network. Experimental results show that the performance of our model is comparable to 11 state-of-the-art models in video attention prediction, while it costs only 0.68 MB memory footprint, runs about 10,106 FPS on GPU and 404 FPS on CPU, which is 206 times faster than previous models.
Kui Fu, Peipei Shi, Yafei Song 0002, Shiming Ge, Xiangju Lu, Jia Li 0003
AAAI4
2020 Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid Network
abstract
Accurate temporal action proposals play an important role in detecting actions from untrimmed videos. The existing approaches have difficulties in capturing global contextual information and simultaneously localizing actions with different durations. To this end, we propose a Relation-aware pyramid Network (RapNet) to generate highly accurate temporal action proposals. In RapNet, a novel relation-aware module is introduced to exploit bi-directional long-range relations between local features for context distilling. This embedded module enhances the RapNet in terms of its multi-granularity temporal proposal generation ability, given predefined anchor boxes. We further introduce a two-stage adjustment scheme to refine the proposal boundaries and measure their confidence in containing an action with snippet-level actionness. Extensive experiments on the challenging ActivityNet and THUMOS14 benchmarks demonstrate our RapNet generates superior accurate proposals over the existing state-of-the-art methods.
Jialin Gao, Zhixiang Shi, Guanshuo Wang, Yufeng Yuan, Shiming Ge, Xi Zhou 0001
AAAI6
2020 Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
abstract
In spite of great success in many image recognition tasks achieved by recent deep models, directly applying them to recognize low-resolution images may suffer from low accuracy due to the missing of informative details during resolution degradation. However, these images are still recognizable for subjects who are familiar with the corresponding high-resolution ones. Inspired by that, we propose a teacher-student learning approach to facilitate low-resolution image recognition via hybrid order relational knowledge distillation. The approach refers to three streams: the teacher stream is pretrained to recognize high-resolution images in high accuracy, the student stream is learned to identify low-resolution images by mimicking the teacher's behaviors, and the extra assistant stream is introduced as bridge to help knowledge transfer across the teacher to the student. To extract sufficient knowledge for reducing the loss in accuracy, the learning of student is supervised with multiple losses, which preserves the similarities in various order relational structures. In this way, the capability of recovering missing details of familiar low-resolution images can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on metric learning, low-resolution image classification and low-resolution face recognition tasks show the effectiveness of our approach, while taking reduced models.
Shiming Ge, Kangkai Zhang, Yingying Hua, Shengwei Zhao, Xin Jin 0015
AAAI1
2020 Coupled-View Deep Classifier Learning from Multiple Noisy Annotators
abstract
Typically, learning a deep classifier from massive cleanly annotated instances is effective but impractical in many real-world scenarios. An alternative is collecting and aggregating multiple noisy annotations for each instance to train the classifier. Inspired by that, this paper proposes to learn deep classifier from multiple noisy annotators via a coupled-view learning approach, where the learning view from data is represented by deep neural networks for data classification and the learning view from labels is described by a Naive Bayes classifier for label aggregation. Such coupled-view learning is converted to a supervised learning problem under the mutual supervision of the aggregated and predicted labels, and can be solved via alternate optimization to update labels and refine the classifiers. To alleviate the propagation of incorrect labels, small-loss metric is proposed to select reliable instances in both views. A co-teaching strategy with class-weighted loss is further leveraged in the deep classifier learning, which uses two networks with different learning abilities to teach each other, and the diverse errors introduced by noisy labels can be filtered out by peer networks. By these strategies, our approach can finally learn a robust data classifier which less overfits to label noise. Experimental results on synthetic and real data demonstrate the effectiveness and robustness of the proposed approach.
Shikun Li, Shiming Ge, Yingying Hua, Chunhui Zhang 0001, Tengfei Liu 0007, Weiqiang Wang 0002
AAAI2
2020 Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
abstract
Many real-world applications today like video surveillance and urban governance need to address the recognition of masked faces, where content replacement by diverse masks often brings in incomplete appearance and ambiguous representation, leading to a sharp drop in accuracy. Inspired by recent progress on amodal perception, we propose to migrate the mechanism of amodal completion for the task of masked face recognition with an end-to-end de-occlusion distillation framework, which consists of two modules. The de-occlusion module applies a generative adversarial network to perform face completion, which recovers the content under the mask and eliminates appearance ambiguity. The distillation module takes a pre-trained general face recognition model as the teacher and transfers its knowledge to train a student for completed faces using massive online synthesized face pairs. Especially, the teacher knowledge is represented with structural relations among instances in multiple orders, which serves as a posterior regularization to enable the adaptation. In this way, the knowledge can be fully distilled and transferred to identify masked faces. Experiments on synthetic and realistic datasets show the efficacy of the proposed approach.
Chenyu Li 0001, Shiming Ge, Daichi Zhang, Jia Li 0003
ACM Multimedia2
2020 Talking Face Generation with Expression-Tailored Generative Adversarial Network
abstract
A key of automatically generating vivid talking faces is to synthesize identity-preserving natural facial expressions beyond audio-lip synchronization, which usually need to disentangle the informative features from multiple modals and then fuse them together. In this paper, we propose an end-to-end Expression-Tailored Generative Adversarial Network (ET-GAN) to generate an expression enriched talking face video of arbitrary identity. Different from talking face generation based on identity image and audio, an expressional video of arbitrary identity serves as the expression source in our approach. Expression encoder is proposed to disentangle expression-tailored representation from the guiding expressional video, while audio encoder disentangles audio-lip representation. Instead of using single image as identity input, multi-image identity encoder is proposed by learning different views of faces and merging a unified representation. Multiple discriminators are exploited to keep both image-aware and the video-aware realistic details, including a spatial-temporal discriminator for visual continuity of expression synthesis and facial movements. We conduct extensive experimental evaluations on quantitative metrics, expression retention quality and audio-visual synchronization. The results show the effectiveness of our ET-GAN in generating high quality expressional talking face videos against existing state-of-the-arts.
Dan Zeng 0001, Shiming Ge
ACM Multimedia4
2020 Accurate UAV Tracking with Distance-Injected Overlap Maximization
abstract
UAV tracking is usually challenged by the dual-dynamic disturbances that arise from not only diverse moving target but also motion camera, leading to a more serious model drift issue than traditional visual tracking. In this work, we propose to alleviate this issue with distance-injected overlap maximization. Our idea is improving the accuracy of target localization by deriving a conceptually simple target localization loss and a global feature recalibration scheme in a mutual reinforced way. In particular, the target localization loss is designed by simply incorporating the normalized distance of target offset and generic semantic IoU loss, resulting in the distance-injected semantic IoU loss, and its minimal solution can alleviate the drift problem caused by camera motion. Moreover, the deep feature extractor is reconstructed and alternated with a feature recalibration network, which can leverage the global information to recalibrate significant features and suppress negligible features. Following by multi-scale feature concat, the proposed tracker can improve the discriminative capability of feature representation for UAV targets on the fly. Extensive experimental results on four benchmarks, i.e. UAV123, UAVDT, DTB70, and VisDrone, demonstrate the superiority of the proposed tracker against existing state-of-the-arts on UAV tracking.
Chunhui Zhang 0001, Shiming Ge, Kangkai Zhang, Dan Zeng 0001
ACM Multimedia2
2020 View-specific subspace learning and re-ranking for semi-supervised person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Shiming Ge
Pattern Recognit.5
2020 Occluded Face Recognition in the Wild by Identity-Diversity Inpainting
abstract
Face recognition has achieved advanced development by using convolutional neural network (CNN) based recognizers. Existing recognizers typically demonstrate powerful capacity in recognizing un-occluded faces, but often suffer from accuracy degradation when directly identifying occluded faces. This is mainly due to insufficient visual and identity cues caused by occlusions. On the other hand, generative adversarial network (GAN) is particularly suitable when it needs to reconstruct visually plausible occlusions by face inpainting. Motivated by these observations, this paper proposes identity-diversity inpainting to facilitate occluded face recognition. The core idea is integrating GAN with an optimized pre-trained CNN recognizer which serves as the third player to compete with the generator by distinguishing diversity within the same identity class. To this end, a collect of identity-centered features is applied in the recognizer as supervision to enable the inpainted faces clustering towards their identity centers. In this way, our approach can benefit from GAN for reconstruction and CNN for representation, and simultaneously addresses two challenging tasks, face inpainting and face recognition. Experimental results compared with 4 state-of-the-arts prove the efficacy of the proposed approach.
Shiming Ge, Chenyu Li 0001, Shengwei Zhao, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Distilling Channels for Efficient Deep Tracking
abstract
Deep trackers have proven success in visual tracking. Typically, these trackers employ optimally pre-trained deep networks to represent all diverse objects with multi-channel features from some fixed layers. The deep networks employed are usually trained to extract rich knowledge from massive data used in object classification and so they are capable to represent generic objects very well. However, these networks are too complex to represent a specific moving object, leading to poor generalization as well as high computational and memory costs. This paper presents a novel and general framework termed channel distillation to facilitate deep trackers. To validate the effectiveness of channel distillation, we take discriminative correlation filter (DCF) and ECO for example. We demonstrate that an integrated formulation can turn feature compression, response map generation, and model update into a unified energy minimization problem to adaptively select informative feature channels that improve the efficacy of tracking moving objects on the fly. Channel distillation can accurately extract good channels, alleviating the influence of noisy channels and generally reducing the number of channels, as well as adaptively generalizing to different channels and networks. The resulting deep tracker is accurate, fast, and has low memory requirements. Extensive experimental evaluations on popular benchmarks clearly demonstrate the effectiveness and generalizability of our framework.
Shiming Ge, Zhao Luo, Chunhui Zhang 0001, Yingying Hua, Dacheng Tao
IEEE Trans. Image Process.1
2020 Efficient Low-Resolution Face Recognition via Bridge Distillation
abstract
Face recognition in the wild is now advancing towards light-weight models, fast inference speed and resolution-adapted capability. In this paper, we propose a bridge distillation approach to turn a complex face model pretrained on private high-resolution faces into a light-weight one for low-resolution face recognition. In our approach, such a cross-dataset resolution-adapted knowledge transfer problem is solved via two-step distillation. In the first step, we conduct cross-dataset distillation to transfer the prior knowledge from private high-resolution faces to public high-resolution faces and generate compact and discriminative features. In the second step, the resolution-adapted distillation is conducted to further transfer the prior knowledge to synthetic low-resolution faces via multi-task learning. By learning low-resolution face representations and mimicking the adapted high-resolution knowledge, a light-weight student model can be constructed with high efficiency and promising accuracy in recognizing low-resolution faces. Experimental results show that the student model performs impressively in recognizing low-resolution faces with only 0.21M parameters and 0.057MB memory. Meanwhile, its speed reaches up to 14,705, 934 and 763 faces per second on GPU, CPU and mobile phone, respectively.
Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Yu Zhang 0035, Jia Li 0003
IEEE Trans. Image Process.1
2020 Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
abstract
The performance of video saliency estimation techniques has achieved significant advances along with the rapid development of Convolutional Neural Networks (CNNs). However, devices like cameras and drones may have limited computational capability and storage space so that the direct deployment of complex deep saliency models becomes infeasible. To address this problem, this paper proposes a dynamic saliency estimation approach for aerial videos via spatiotemporal knowledge distillation. In this approach, five components are involved, including two teachers, two students and the desired spatiotemporal model. The knowledge of spatial and temporal saliency is first separately transferred from the two complex and redundant teachers to their simple and compact students, while the input scenes are also degraded from high-resolution to low-resolution to remove the probable data redundancy so as to greatly speed up the feature extraction process. After that, the desired spatiotemporal model is further trained by distilling and encoding the spatial and temporal saliency knowledge of two students into a unified network. In this manner, the inter-model redundancy can be removed for the effective estimation of dynamic saliency on aerial videos. Experimental results show that the proposed approach is comparable to 11 state-of-the-art models in estimating visual saliency on aerial videos, while its speed reaches up to 28,738 FPS and 1,490.5 FPS on the GPU and CPU platforms, respectively.
Jia Li 0003, Kui Fu, Shengwei Zhao, Shiming Ge
IEEE Trans. Image Process.4
2020 Receptive Multi-Granularity Representation for Person Re-Identification
abstract
A key for person re-identification is achieving consistent local details for discriminative representation across variable environments. Current stripe-based feature learning approaches have delivered impressive accuracy, but do not make a proper trade-off between diversity, locality, and robustness, which easily suffers from part semantic inconsistency for the conflict between rigid partition and misalignment. This paper proposes a receptive multi-granularity learning approach to facilitate stripe-based feature learning. This approach performs local partition on the intermediate representations to operate receptive region ranges, rather than current approaches on input images or output features, thus can enhance the representation of locality while remaining proper local association. Toward this end, the local partitions are adaptively pooled by using significance-balanced activations for uniform stripes. Random shifting augmentation is further introduced for a higher variance of person appearing regions within bounding boxes to ease misalignment. By twobranch network architecture, different scales of discriminative identity representation can be learned. In this way, our model can provide a more comprehensive and efficient feature representation without larger model storage costs. Extensive experiments on intra-dataset and cross-dataset evaluations demonstrate the effectiveness of the proposed approach. Especially, our approach achieves a state-of-the-art accuracy of 96.2%@Rank-1 or 90.0%@mAP on the challenging Market-1501 benchmark.
Guanshuo Wang, Yufeng Yuan, Shiming Ge, Xi Zhou 0001
IEEE Trans. Image Process.4
2019 Robust Deep Tracking with Two-step Augmentation Discriminative Correlation Filters
abstract
Recently, deep trackers have proven success in visual tracking due to their powerful feature representation. Among them, discriminative correlation filter (DCF) paradigm is widely used. However, these trackers are still difficult to learn an adaptive appearance model of the object due to the limited data available. To address that, this paper proposes a two-step augmentation discriminative correlation filters (TADCF) approach to improve robustness. Firstly, we propose an online frame augmentation scheme to obtain rich and robust deep features which can effectively alleviate background distractors, leading to better generalization and adaptation of the learned model. Secondly, an object augmentation mechanism is implemented by exploiting rotation continuity restriction, which simultaneously models target appearance changes from rotation and scale variations. Extensive experiments on four benchmarks illustrate that the proposed approach performs favorably against state-of-the-art trackers.
Chunhui Zhang 0001, Shiming Ge, Yingying Hua, Dan Zeng 0001
ICME2
2019 Fewer-Shots and Lower-Resolutions: Towards Ultrafast Face Recognition in the Wild
abstract
Is it possible to train an effective face recognition model with fewer shots that works efficiently on low-resolution faces in the wild? To answer this question, this paper proposes a few-shot knowledge distillation approach to learn an ultrafast face recognizer via two steps. In the first step, we initialize a simple yet effective face recognition model on synthetic low-resolution faces by distilling knowledge from an existing complex model. By removing the redundancies in both face images and the model structure, the initial model can provide an ultrafast speed with impressive recognition accuracy. To further adapt this model into the wild scenarios with fewer faces per person, the second step refines the model via few-shot learning by incorporating a relation module that compares low-resolution query faces with faces in the support set. In this manner, the performance of the model can be further enhanced with only fewer low-resolution faces in the wild. Experimental results show that the proposed approach performs favorably against state-of-the-arts in recognizing low-resolution faces with an extremely low memory of 30KB and runs at an ultrafast speed of 1,460 faces per second on CPU or 21,598 faces per second on GPU.
Shiming Ge, Shengwei Zhao, Xindi Gao, Jia Li 0003
ACM Multimedia1
2019 Defending Against Adversarial Examples via Soft Decision Trees Embedding
abstract
Convolutional neural networks (CNNs) have shown vulnerable to adversarial examples which contain imperceptible perturbations. In this paper, we propose an approach to defend against adversarial examples with soft decision trees embedding. Firstly, we extract the semantic features of adversarial examples with a feature extraction network. Then, a specific soft decision tree is trained and embedded to select the key semantic features for each feature map from convolutional layers and the selected features are fed to a light-weight classification network. To this end, we use the probability distributions of each tree node to quantify the semantic features. In this way, some small perturbations can be effectively removed and the selected features are more discriminative in identifying adversarial examples. Moreover, the influence of adversarial perturbations on classification can be reduced by migrating the interpretability of soft decision trees into the black-box neural networks. We conduct experiments to defend the state-of-the-art adversarial attacks. The experimental results demonstrate that our proposed approach can effectively defend against these attacks and improve the robustness of deep neural networks.
Yingying Hua, Shiming Ge, Xindi Gao, Xin Jin 0015, Dan Zeng 0001
ACM Multimedia2
2019 Aesthetic Attributes Assessment of Images
abstract
Image aesthetic quality assessment has been a relatively hot topic during the last decade. Most recently, comments type assessment (aesthetic captions) has been proposed to describe the general aesthetic impression of an image using text. In this paper, we propose Aesthetic Attributes Assessment of Images, which means the aesthetic attributes captioning. This is a new formula of image aesthetic assessment, which predicts aesthetic attributes captions together with the aesthetic score of each attribute. We introduce a new dataset named DPC-Captions which contains comments of up to 5 aesthetic attributes of one image through knowledge transfer from a full-annotated small-scale dataset. Then, we propose Aesthetic Multi-Attribute Network (AMAN), which is trained on a mixture of fully-annotated small-scale PCCD dataset and weakly-annotated large-scale DPC-Captions dataset. Our AMAN makes full use of transfer learning and attention model in a single framework. The experimental results on our DPC-Captions and PCCD dataset reveal that our method can predict captions of 5 aesthetic attributes together with numerical score assessment of each attribute. We use the evaluation criteria used in image captions to prove that our specially designed AMAN model outperforms traditional CNN-LSTM model and modern SCA-CNN model of image captions.
Xin Jin 0015, Geng Zhao 0001, Xiaodong Li 0013, Xiaokun Zhang 0002, Shiming Ge, Dongqing Zou, Xinghui Zhou
ACM Multimedia6
2019 ILGNet: inception modules with connected local and global features for efficient image aesthetic quality classification using domain adaptation
abstract
In this study, the authors address a challenging problem of aesthetic image classification, which is to label an input image as high‐ or low‐aesthetic quality. We take both the local and global features of images into consideration. A novel deep convolutional neural network named ILGNet is proposed, which combines both the inception modules and a connected layer of both local and global features. The ILGnet is based on GoogLeNet. Thus, it is easy to use a pre‐trained GoogLeNet for large‐scale image classification problem and fine tune their connected layers on a large‐scale database of aesthetic‐related images: AVA, i.e. domain adaptation . The experiments reveal that their model achieves the state of the arts in AVA database. Both the training and testing speeds of their model are higher than those of the original GoogLeNet.
Xin Jin 0015, Xiaodong Li 0013, Xiaokun Zhang 0002, Jingying Chi, Siwei Peng, Shiming Ge, Geng Zhao 0001
IET Comput. Vis.7
2019 Proposal pyramid networks for fast face detection
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002, Zhijiang Zhang
Inf. Sci.4
2019 Fast cascade face detection with pyramid network
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002
Pattern Recognit. Lett.3
2019 Secure face retrieval for group mobile users
Xin Jin 0015, Yujie Li 0001, Shiming Ge, Chenggen Song, Xinghui Zhou
Soft Comput.3
2019 Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation
abstract
Typically, the deployment of face recognition models in the wild needs to identify low-resolution faces with extremely low computational cost. To address this problem, a feasible solution is compressing a complex face model to achieve higher speed and lower memory at the cost of minimal performance drop. Inspired by that, this paper proposes a learning approach to recognize low-resolution faces via selective knowledge distillation. In this approach, a two-stream convolutional neural network (CNN) is first initialized to recognize high-resolution faces and resolution-degraded faces with a teacher stream and a student stream, respectively. The teacher stream is represented by a complex CNN for high-accuracy recognition, and the student stream is represented by a much simpler CNN for low-complexity recognition. To avoid significant performance drop at the student stream, we then selectively distil the most informative facial features from the teacher stream by solving a sparse graph optimization problem, which are then used to regularize the finetuning process of the student stream. In this way, the student stream is actually trained by simultaneously handling two tasks with limited computational resources: approximating the most informative facial cues via feature regression, and recovering the missing facial cues via low-resolution face classification. Experimental results show that the student stream performs impressively in recognizing low-resolution faces and costs only 0.15MB memory and runs at 418 faces per second on CPU and 9; 433 faces per second on GPU.
Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Jia Li 0003
IEEE Trans. Image Process.1
2018 Predicting Aesthetic Score Distribution Through Cumulative Jensen-Shannon Divergence
abstract
Aesthetic quality prediction is a challenging task in the computer vision community because of the complex interplay with semantic contents and photographic technologies. Recent studies on the powerful deep learning based aesthetic quality assessment usually use a binary high-low label or a numerical score to represent the aesthetic quality. However the scalar representation cannot describe well the underlying varieties of the human perception of aesthetics. In this work, we propose to predict the aesthetic score distribution (i.e., a score distribution vector of the ordinal basic human ratings) using Deep Convolutional Neural Network (DCNN). Conventional DCNNs which aim to minimize the difference between the predicted scalar numbers or vectors and the ground truth cannot be directly used for the ordinal basic rating distribution. Thus, a novel CNN based on the Cumulative distribution with Jensen-Shannon divergence (CJS-CNN) is presented to predict the aesthetic score distribution of human ratings, with a new reliability-sensitive learning method based on the kurtosis of the score distribution, which eliminates the requirement of the original full data of human ratings (without normalization). Experimental results on large scale aesthetic dataset demonstrate the effectiveness of our introduced CJS-CNN in this task.
Xin Jin 0015, Xiaodong Li 0013, Siwei Peng, Jingying Chi, Shiming Ge, Chenggen Song, Geng Zhao 0001
AAAI7
2018 Predicting Aesthetic Radar Map Using a Hierarchical Multi-task Network
Xin Jin 0015, Xinghui Zhou, Geng Zhao 0001, Xiaokun Zhang 0002, Xiaodong Li 0013, Shiming Ge
PRCV (2)7
2018 Image editing by object-aware optimal boundary searching and mixed-domain composition
abstract
When combining very different images which often contain complex objects and backgrounds, producing consistent compositions is a challenging problem requiring seamless image editing. In this paper, we propose a general approach, called object-aware image editing , to obtain consistency in structure, color, and texture in a unified way. Our approach improves upon previous gradient-domain composition in three ways. Firstly, we introduce an iterative optimization algorithm to minimize mismatches on the boundaries when the target region contains multiple objects of interest. Secondly, we propose a mixed-domain consistency metric for measuring gradients and colors, and formulate composition as a unified minimization problem that can be solved with a sparse linear system. In particular, we encode texture consistency using a patch-based approach without searching and matching. Thirdly, we adopt an object-aware approach to separately manipulate the guidance gradient fields for objects of interest and backgrounds of interest, which facilitates a variety of seamless image editing applications. Our unified method outperforms previous state-of-the-art methods in preserving global texture consistency in addition to local structure continuity.
Shiming Ge, Xin Jin 0015, Qiting Ye, Zhao Luo, Qiang Li 0007
Comput. Vis. Media1
2018 Color image encryption in non-RGB color spaces
Xin Jin 0015, Sui Yin, Ningning Liu, Xiaodong Li 0013, Geng Zhao 0001, Shiming Ge
Multim. Tools Appl.6
2018 Enhancing heterogeneous similarity estimation via neighborhood reversibility
Shikui Wei, Yao Zhao 0001, Shiming Ge
Multim. Tools Appl.5
2017 Detecting Masked Faces in the Wild with LLE-CNNs
abstract
Detecting faces with occlusions is a challenging task due to two main reasons: 1) the absence of large datasets of masked faces, and 2) the absence of facial cues from the masked regions. To address these two issues, this paper first introduces a dataset, denoted as MAFA, with 30, 811 Internet images and 35, 806 masked faces. Faces in the dataset have various orientations and occlusion degrees, while at least one part of each face is occluded by mask. Based on this dataset, we further propose LLE-CNNs for masked face detection, which consist of three major modules. The Proposal module first combines two pre-trained CNNs to extract candidate facial regions from the input image and represent them with high dimensional descriptors. After that, the Embedding module is incorporated to turn such descriptors into a similarity-based descriptor by using locally linear embedding (LLE) algorithm and the dictionaries trained on a large pool of synthesized normal faces, masked faces and non-faces. In this manner, many missing facial cues can be largely recovered and the influences of noisy cues introduced by diversified masks can be greatly alleviated. Finally, the Verification module is incorporated to identify candidate facial regions and refine their positions by jointly performing the classification and regression tasks within a unified CNN. Experimental results on the MAFA dataset show that the proposed approach remarkably outperforms 6 state-of-the-arts by at least 15.6%.
Shiming Ge, Jia Li 0003, Qiting Ye, Zhao Luo
CVPR1
2017 Compressing deep neural networks for efficient visual inference
abstract
The deployments of deep neural network models on mobile or embedded devices have been challenged due to two main reasons: 1) the large model size for storage, and 2) the large memory bandwidth for inference. To address these issues, this paper develops a deep neural network compression framework to reduce the resource usage for efficient visual inference. By reviewing the trained deep model, we propose a hybrid model compression algorithm via four major modules. Approximation module reduces the number of weights in each fully connected layer with low rank approximation. Then, quantization module analyzes weight distribution in each layer and represents them with low precision fixed point, which reduces the bits for storing each weight. After that, pruning module suppresses small weights to further reduce the number of parameters. Finally, coding module joint optimizes the representation and encoding of the sparse structure of the pruned weights with relative index by Huffman coding. Beyond the compression of model size, we propose an adaptive fixed point memory allocation algorithm to reduce memory footprint in inference. The proposed framework, along with the model compression and memory allocation algorithms, can provide 20-30x compression rate with negligible accuracy loss. We conduct an evaluation on two representative models, AlexNet and VGG-16, for object recognition and face verification tasks, which demonstrate the effectiveness of our proposed compression framework.
Shiming Ge, Zhao Luo, Shengwei Zhao, Xin Jin 0015, Xiaoyu Zhang 0002
ICME1
2017 Efficient privacy preserving Viola-Jones type object detection via random base image representation
abstract
A cloud server spent a lot of time, energy and money to train a Viola-Jones type object detector [1] with high accuracy. Clients can upload their photos to the cloud server to find objects. However, the client does not want the leakage of the content of his/her photos. In the meanwhile, the cloud server is also reluctant to leak any parameters of the trained object detectors. 10 years ago, Avidan & Butman introduced Blind Vision, which is a method for securely evaluating a ViolaJones type object detector. Blind Vision uses standard cryptographic tools and is painfully slow to compute, taking a couple of hours to scan a single image. The purpose of this work is to explore an efficient method that can speed up the process. We propose the Random Base Image (RBI) Representation. The original image is divided into random base images. Only the base images are submitted randomly to the cloud server. Thus, the content of the image can not be leaked. In the meanwhile, a random vector and the secure Millionaire protocol are leveraged to protect the parameters of the trained object detector. The RBI makes the integral-image enable again for the great acceleration. The experimental results reveal that our method can retain the detection accuracy of that of the plain vision algorithm and is significantly faster than the traditional blind vision, with only a very low probability of the information leakage theoretically.
Xin Jin 0015, Xiaodong Li 0013, Chenggen Song, Shiming Ge, Geng Zhao 0001, Yingya Chen
ICME5
2017 3D textured model encryption via 3D Lu chaotic mapping
Xin Jin 0015, Shuyun Zhu, Chaoen Xiao, Xiaodong Li 0013, Geng Zhao 0001, Shiming Ge
Sci. China Inf. Sci.7
2016 Learning multi-channel correlation filter bank for eye localization
Shiming Ge, Kaixuan Xie, Hongsong Zhu, Shuixian Chen
Neurocomputing1
2016 Global image completion with joint sparse patch selection and optimal seam synthesis
Shiming Ge, Kaixuan Xie
Signal Process.1
2015 Target Domain Adaptation for Face Detection in a Smart Camera Network with Peer-to-Peer Communications
abstract
With the fast advance of mobile chips technologies, a node in a smart camera network can afford sophisticated processing via on-board multicore CPUs and GPUs, e.g., face detection. The performance of a general purpose face detector, however, may degrade seriously under specific situations with unexpected challenges such as facial coverage or bad illumination. This degradation is due to the difference of probability distributions between training data and testing data, known as source data domain and target data domain, respectively. To better adapt a smart camera network to a specific situation, some form of target domain adaptation is needed, which usually requires both the source domain data and as much as possible target domain data at each node, which may strain storage capacity and bandwidth. In this paper, we propose an adaptation method which fuses the source specific hypotheses (SSHs) and target specific hypotheses (TSHs) - requiring only a pre-trained face detector and a few target data to be shared by peer-to-peer communications, thus relieving the storage and bandwidth constraints. The method uses the "accuracy-regularization" objective as the adaptation model, to fuse SSHs and TSHs, and tries to minimize the misclassification error on target data. With an existing frontal face detector, we conduct experiments to verify our algorithm, covering cases of video surveillance, extreme pose challenge, and different illumination spectra. Significant performance gains are observed with only dozens of target data in all the experiments, demonstrating the effectiveness of the proposed adaptation model. Therefore, the proposed adaption can be applied to a smart camera network with peer-to-peer communications to improve the network's overall face detection performance.
Shuixian Chen, Xiang Lu 0004, Limin Sun 0001, Shiming Ge
GLOBECOM4
2015 Abnormal event detection via adaptive cascade dictionary learning
abstract
Detecting abnormal events plays an essential role in video content analysis and has received increasing attention in surveillance system. One of the major problems in abnormal event detection is the imbalanced classification issue due to the rare abnormal samples. Another problem is the difficulty of detecting anomalies within a reasonable amount of computation time. To address these problems, we propose an adaptive cascade dictionary learning framework for detecting the anomalies. The framework considers anomaly detection as an one-class classification problem with a cascade of dictionaries. Each stage of the cascade constructs an adaptive dictionary to detect the anomalies with costless least square optimization solution. The experiments on benchmark datasets demonstrate that the proposed method has a better performance while comparing with several state-of-the-art methods.
Hui Wen 0001, Shiming Ge, Shuixian Chen, Hongtao Wang 0002, Limin Sun 0001
ICIP2
2014 Eye localization based on correlation filter bank
abstract
Eye localization is a key step in many face analysis related applications. In this paper, we present a novel eye localization method based on a group of trained filters called correlation filter bank (CFB). We formulate the eye localization problem as an optimization problem with a well-defined cost function based on CFB. The CFB is trained with an EM-like adaptive clustering approach. The trained filter bank includes several discriminative filter templates, each of them suits to a different face condition from the others, thus can provide accurate eye localization ability for variable poses, appearances and illuminations. Simulation comparisons with cascade classifier-based method [1], traditional single correlation filter based methods [2][3] and pictorial structure model based method [4] demonstrates the superiority of the proposed method both in detection ratio and localization accuracy.
Shiming Ge, Hui Wen 0001, Shuixian Chen, Limin Sun 0001
ICME1
2014 Image Completion Using Global Patch Matching and Optimal Seam Synthesis
abstract
This paper presented a global exemplar-based image completion method for filling large missing or damaged regions in an image. Based on three proposed completion rules, the image completion problem is formulated as a global discrete optimization problem with a well-defined energy function. The energy function can evaluate image consistency globally and is minimized with an expectation-maximization (EM) like algorithm, which considers patch matching and patch synthesis in a unified way. In the algorithm, M step and E step are achieved by fast coherent searching and optimal seam synthesis respectively. Moreover, E step combines image patch synthesis and coherent correction simultaneously. We analyzed our global energy function and optimization method in theory. Simulation comparisons with other state-of-the-art methods show the superiority of our proposed method in ensuring global coherent and avoiding image blurring.
Shiming Ge, Kaixuan Xie, Zhiqiang Shi
ICPR1
2014 Poster: Crowdsourcing for video traffic surveillance
abstract
No abstract available.
Hui Wen 0001, Qiang Li 0007, Qi Han 0001, Shiming Ge, Limin Sun 0001
MobiSys4
2009 A unified gradient domain method for seamless image processing
abstract
Seamless image processing concerns stitching parts of images in a visually natural manner. In this paper, we present a seamless processing method which unifies Featuring, Optimal seam method and some recent gradient domain methods including Poisson image editing, Drag-and-drop pasting and GIST. To evaluate the processing quality visually, a variational energy function is proposed to compute both the similarity of the stitched image to each of the input images and the visibility of the seam between the stitched images. The minimum of the energy function gives a globally consistent composition in both geometrical and photometric structures. Based on 3 propositions, we study the energy function and compare it with other methods theoretically. The experimental results show the benefits of our stitching method.
Shiming Ge, Kongqiao Wang
ICIP1
2009 A fast and effective outlier detection method for matching uncalibrated images
abstract
Many image analysis tasks require an outlier detection procedure to identify the false matches. In this paper, a fast and effective outlier detection method is presented to match images in the uncalibrated case. This method employs a hypothesis test on the consistency of dominant orientations of the feature points to significantly increase the detection speed. Moreover, it can also effectively find the outliers that can not be identified by traditional RANSAC-based methods using epipolar constraint. Note that our method does not require the prior knowledge of camera parameters or the percentage of outliers. The experimental results show that our method outperforms the classical RANSAC-based methods both in speed and in accuracy of the results.
Xiujuan Chai, Shiming Ge
ICIP4