Ming Shao

dblp:49/5671 · DBLP profile ↗
← Back
107ranked-venue papers
16as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 8 first-author · 18 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust defense strategies for multimodal contrastive learning: efficient fine-tuning against backdoor attacks
Afia Sajeeda, Neeresh Kumar Perla, Ming Shao
Multim. Tools Appl.4
2026 A survey of recent advances in adversarial attack and defense on vision-language models
Neeresh Kumar Perla, Afia Sajeeda, Si-Yu Xia, Ming Shao
Neural Networks5
2025 Supportive Negatives Spectral Augmentation for Source-Free Cross-Domain Segmentation
abstract
Source-free domain adaptation (SFDA) aims to transfer knowledge from the well-trained source model and optimize it to adapt target data distribution. SFDA methods are suitable for medical image segmentation task due to its data-privacy protection and achieve promising performances. However, cross-domain distribution shift makes it difficult for the adapted model to provide accurate decisions on several hard instances and negatively affects model generalization. To overcome this limitation, a novel method `supportive negatives spectral augmentation' (SNSA) is presented in this work. Concretely, SNSA includes the instance selection mechanism to automatically discover a few hard samples for which source model produces incorrect predictions. And, active learning strategy is adopted to re-calibrate their predictive masks. Moreover, SNSA deploys the spectral augmentation between hard instances and others to encourage source model to gradually capture and adapt the attributions of target distribution. Considerable experimental studies demonstrate that annotating merely 4%~5% of negative instances from the target domain significantly improves segmentation performance over previous methods.
Kexin Zheng, Haifeng Xia, Si-Yu Xia, Ming Shao, Zhengming Ding
AAAI4
2025 IPNet: Interpretable Prototype Network for Multi-Source Domain Adaptation
abstract
Multi-source domain adaptation (MSDA) borrows intrinsic knowledge from well-annotated source domains to identify target visual signals. The main challenges are effectively mitigating cross-domain shift and extracting discriminative target features via the suitable source semantics. To overcome them, this paper proposes a novel Interpretable Prototype Network (IPNet) with channel-wise augmentation and multi-domain prototype mechanism. Specifically, IPNet explores the parameterized channel fusion paradigm across multiple source domains and target one to generate intermediate instances and achieve beneficial alignment. Moreover, IPNet analyzes contributions of source domains with interpretable learning approach and adjusts their effects on representations of target signals. Extensive experiments on three MSDA benchmark datasets suggest the advantages of our IPNet over others and exhibit the path of knowledge transfer.
Haifeng Xia, Si-Yu Xia, Ming Shao, Zhengming Ding
ICASSP4
2025 QuantBEVFusion: A Fully Quantized Framework for LiDAR-camera 3D Object Detection
abstract
3D object detection is essential for robust environmental perception in autonomous driving and robotics. While LiDAR-camera fusion methods offer high accuracy, their computational complexity hinders deployment on resource-constrained edge devices. To address this, we introduce QuantBEVFusion, a fully quantized 3D object detection framework that prioritizes both quantization and operator optimization. Our approach tackles the inherent asymmetry between LiDAR and camera data in pillar-based models by incorporating a novel pillar bird’s-eye view (BEV) encoder, significantly boosting performance. Furthermore, we introduce 1) an optimized LiDAR input processing method that filters noise and enables per-tensor quantization; 2) an improved sparse feature quantization process with log-histogram balancing, adaptive bin widths, and distillation loss for enhanced accuracy; and 3) a deployment-friendly 3D-to-2D transformation operator facilitating fixed-point implementation. Extensive experiments demonstrate that QuantBEVFusion achieves state-of-the-art quantization performance while maintaining accuracy suitable for real-time applications on edge devices.
Xubin Wen, Ming Shao, Libo Sun 0001, Wenhu Qin, Si-Yu Xia
IJCNN2
2025 Are Exemplar-Based Class Incremental Learning Models Victim of Black-Box Poison Attacks?
abstract
Class Incremental Learning (CIL) models are designed to continuously learn new classes without forgetting previously learned ones, often relying on an exemplar set to retain a portion of knowledge from previously learned classes. However, their vulnerability to adversarial attacks under novel and unexplored conditions remains unstudied. In this work, we are the first to evaluate the robustness of exemplar-based CIL models using a non-overlapping dataset, where the dataset is independent of the training and test sets of the target model. We propose and implement a novel black-box attack framework targeting the exemplar set of class incremental learning models using zero-overlapping data. Specifically, we focus on scenarios where the target model provides only hard-label predictions with-out interactive access. Our experimental evaluation covers a range of exemplar-based incremental learning algorithms, different surrogate models, and black-box attack options. Our findings reveal significant vulnerabilities in exemplar-based CIL models to poisoning-based attacks using a non-overlapping dataset.
Neeresh Kumar Perla, Afia Sajeeda, Ming Shao
WACV4
2025 SR-BigGAN: lightweight image super-resolution with priors
Deepak Kumar 0008, Harshitha Srinivas Rao, Chetan Kumar, Ming Shao
Mach. Vis. Appl.4
2025 ACL-SAR: model agnostic adversarial contrastive learning for robust skeleton-based action recognition
Jiaxuan Zhu, Ming Shao, Libo Sun 0001, Si-Yu Xia
Vis. Comput.2
2024 Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action Recognition
abstract
Human action video recognition has recently attracted more attention in applications such as video security and sports posture correction. Popular solutions, including graph convolutional networks (GCNs) that model the human skeleton as a spatiotemporal graph, have proven very effective. GCNs-based methods with stacked blocks usually utilize top-layer semantics for classification/annotation purposes. Although the global features learned through the procedure are suitable for the general classification, they have difficulty capturing fine-grained action change across adjacent frames - decisive factors in sports actions. In this paper, we propose a novel “Cross-block Fine-grained Semantic Cascade (CFSC)” module to overcome this challenge. In summary, the proposed CFSC progressively integrates shallow visual knowledge into high-level blocks to allow networks to focus on action details. In particular, the CFSC module utilizes the GCN feature maps produced at different levels, as well as aggregated features from proceeding levels to consolidate fine-grained features. In addition, a dedicated temporal convolution is applied at each level to learn short-term temporal features, which will be carried over from shallow to deep layers to maximize the leverage of low-level details. This cross-block feature aggregation methodology, capable of mitigating the loss of fine-grained information, has resulted in improved performance. Last, FD-7, a new action recognition dataset for fencing sports, was collected and will be made publicly available. Experimental results and empirical analysis on public benchmarks (FSD-10) and self-collected (FD-7) demonstrate the advantage of our CFSC module on learning discriminative patterns for action classification over others.
Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia
FG5
2024 Autonomous Generative Feature Replay for Non-Exemplar Class-Incremental Learning
abstract
Deep neural networks have been successfully applied in many computer vision tasks. However, these models suffer catastrophic forgetting when learning new knowledge incrementally. To overcome the stability-plasticity dilemma, class incremental learning (CIL) has been widely discussed recently. The state-of-the-art CIL methods mainly leverage additional exemplar sets, thus memory costly and may raise privacy issues. To that end, we propose an autonomous generative feature replay (AGFR) framework without using exemplar sets. It consists of three modules: the feature extractor module, the feature generator module, and the unified classification module. First, to stabilize features over tasks, robust feature extractors are learned in a self-supervised manner and thus generalize well to unseen data. Second, instead of using exemplar sets or producing raw images, we propose an autonomous generative feature replay scheme to constantly update unified classifier in CIL without saving any image data. This strategy avoids overwhelming memory usage or poor quality of the generated raw images. Experiments demonstrate that our method achieves state-of-the-art performance in terms of average classification accuracy.⋆
Yinjie Zhang, Ming Shao, Wenlong Shi, Haifeng Xia, Si-Yu Xia
ICASSP2
2024 Task-aware Disentanglement for Object Detection
abstract
Sibling-head structure is widely used to alleviate the feature conflict between classification and regression tasks in most object detectors. However, as the two branches of the sibling head are trained with exactly the same positive samples and lack explicit feature disentanglement in the forward propagation, the classification-sensitive features and localization-sensitive features are still somewhat coupled. As a result, the feature conflict between the two tasks still remains, which seriously hurts the performance of the classifier and regressor in the testing phase. In this paper, we propose a Task-Aware Disentangled object Detector (TDD) that explicitly disentangles the classification and regression from the aspect of feature disentanglement and sampling strategy. In terms of feature disentanglement, we design a task-aware activation head driven by a reconstruction-activation mechanism to explicitly activate corresponding sensitive features for classification and localization in the forward propagation. Furthermore, we explore a novel task-aware sampling strategy that explicitly assigns the task-adaptive samples for classification and regression tasks according to their quality distributions. Extensive experiments on MS COCO show that our TDD consistently surpasses the baseline by ~2.0 AP with different backbones. Moreover, our best model achieves 55.1 AP, outperforming most state-of-the-art detectors.
Keyang Wang, Fei Wu 0001, Ming Shao
IJCNN4
2024 Sketch3D: Style-Consistent Guidance for Sketch-to-3D Generation
Wangguandong Zheng, Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia, Zhengming Ding
ACM Multimedia5
2024 Few-shot Shape Recognition by Learning Deep Shape-aware Features
abstract
Traditional shape descriptors have been gradually replaced by convolutional neural networks due to their superior performance in feature extraction and classification. The state-of-the-art methods recognize object shapes via image reconstruction or pixel classification. However, these methods are biased toward texture information and overlook the essential shape descriptions, thus, they fail to generalize to unseen shapes. We are the first to propose a few-shot shape descriptor (FSSD) to recognize object shapes given only one or a few samples. We employ an embedding module for FSSD to extract transformation-invariant shape features. Secondly, we develop a dual attention mechanism to decompose and reconstruct the shape features via learnable shape primitives. In this way, any shape can be formed through a finite set basis, and the learned representation model is highly interpretable and extendable to unseen shapes. Thirdly, we propose a decoding module to include the supervision of shape masks and edges and align the original and reconstructed shape features, enforcing the learned features to be more shape-aware. Lastly, all the proposed modules are assembled into a few-shot shape recognition scheme. Experiments on five datasets show that our FSSD significantly improves the shape classification compared to the state-of-the-art under the few-shot setting.
Wenlong Shi, Changsheng Lu, Ming Shao, Yinjie Zhang, Si-Yu Xia, Piotr Koniusz
WACV3
2024 fNIRSNET: A multi-view spatio-temporal convolutional neural network fusion for functional near-infrared spectroscopy-based auditory event classification
Pankaj Pandey, John McLinden, Neela Rahimi, Chetan Kumar, Ming Shao, Kevin M. Spencer, Sarah Ostadabbas, Yalda Shahriari
Eng. Appl. Artif. Intell.5
2024 Utilizing Inherent Bias for Memory Efficient Continual Learning: A Simple and Robust Baseline
Neela Rahimi, Ming Shao
Image Vis. Comput.2
2023 Graphrpe: Relative Position Encoding Graph Transformer for 3d Human Pose Estimation
abstract
The graph neural network has been playing increasingly important roles in 2D-3D lifting based single-frame human pose estimation. However, it still suffers from inferior modeling of local and global associations between 2D nodes. To this end, we introduce two relative position encoding approaches and propose a novel graph-transformer structure "GraphRPE." The model consists of two components: graph relative position encoding (GRPE) and universal relative position encoding (URPE). GRPE embeds graph structure prior information into the attention map to correct the self-attention weights and prompt appropriate interactions between 2D nodes, both locally and globally. On the other hand, URPE introduces the Toeplitz matrix to address the limited representation capability of the transformer. In addition, we investigate several graph edge types and their impacts on the results. Extensive experimental results demonstrate that our proposed method achieves SOTA performance on the Human3.6M dataset.
Junjie Zou, Ming Shao, Si-Yu Xia
ICIP2
2023 The forecast of power consumption and freshwater generation in a solar-assisted seawater greenhouse system using a multi-layer perceptron neural network
Ming Shao
Expert Syst. Appl.2
2023 High-precision skeleton-based human repetitive action counting
abstract
Abstract A novel counting model is presented by the authors to estimate the number of repetitive actions in temporal 3D skeleton data. As per the authors’ knowledge, this is the first work of this kind using skeleton data for high‐precision repetitive action counting. Different from existing works on RGB video data, the authors’ model follows a bottom‐up pipeline to clip the sub‐action first followed by robust aggregation in inference. First, novel counting loss functions and robust inference with backtracking is proposed to pursue precise per‐frame count as well as overall count with boundary frames. Second, an efficient synthetic approach is proposed to augment skeleton data in training and thus avoid time‐consuming repetitive action data collection work. Finally, a challenging human repetitive action counting dataset named VSRep is collected with various types of action to evaluate the proposed model. Experiments demonstrate that the proposed counting model outperforms existing video‐based methods by a large margin in terms of accuracy in real‐time inference.
Chengxian Li, Ming Shao, Si-Yu Xia
IET Comput. Vis.2
2023 LRPRNet: Lightweight Deep Network by Low-Rank Pointwise Residual Convolution
abstract
Deep learning has become popular in recent years primarily due to powerful computing devices such as graphics processing units (GPUs). However, it is challenging to deploy these deep models to multimedia devices, smartphones, or embedded systems with limited resources. To reduce the computation and memory costs, we propose a novel lightweight deep learning module by low-rank pointwise residual (LRPR) convolution, called LRPRNet. Essentially, LRPR aims at using a low-rank approximation in pointwise convolution to further reduce the module size while keeping depthwise convolutions as the residual module to rectify the LRPR module. This is critical when the low-rankness undermines the convolution process. Moreover, our LRPR is quite general and can be directly applied to many existing network architectures such as MobileNetv1, ShuffleNetv2, MixNet, and so on. Experiments on visual recognition tasks, including image classification and face alignment on popular benchmarks, show that our LRPRNet achieves competitive performance but with a significant reduction of Flops and memory cost compared to the state-of-the-art deep lightweight models.
Bin Sun 0002, Jun Li 0027, Ming Shao, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 ElDet: An Anchor-Free General Ellipse Object Detector
Tian Wang 0001, Changsheng Lu, Ming Shao, Si-Yu Xia
ACCV (3)3
2022 Critic-over-Actor-Critic Modeling: Finding Optimal Strategy in ICU Environments
abstract
Reinforcement learning (RL) is mechanized to learn from experience. It solves the problem in sequential decisions by optimizing reward-punishment through experimentation of the distinct actions in an environment. Unlike supervised learning models, RL lacks static input-output mappings and the objective of minimization of a vector error. However, to find out an optimal strategy, it is crucial to learn both continuous feedback from training data and the offline rules of the experiences with no explicit dependence on online samples. In this paper, we present a study of a multi-agent RL framework which involves a Critic in semi-offline mode criticizing over an online Actor-Critic network, namely, Critic-over-Actor-Critic (CoAC) model, in finding optimal treatment plan of ICU patients as well as optimal strategy in a combative battle game. For further validation, we also examine the model in the adversarial assignment.
Riazat Ryan, Ming Shao
IEEE Big Data2
2022 TMCR: A Twin Matching Networks for Chinese Scene Text Retrieval
Zhiheng Peng, Ming Shao, Si-Yu Xia
PRCV (3)2
2022 Generative Adversarial Attack on Ensemble Clustering
abstract
Adversarial attack on learning tasks has attracted substantial attention in recent years; however, most existing works focus on supervised learning. Recently, research has shown that unsupervised learning, such as clustering, tends to be vulnerable due to adversarial attack. In this paper, we focus on a clustering algorithm widely used in the real-world environment, namely, ensemble clustering (EC). EC algorithms usually leverage basic partition (BP) and ensemble techniques to improve the clustering performance collaboratively. Each BP may stem from one trial of clustering, feature segment, or part of data stored on the cloud. We have observed that the attack tends to be less perceivable when only a few BPs are compromised. To explore plausible attack strategies, we propose a novel generative adversarial attack (GA2) model for EC, titled GA2EC. First, we show that not all BPs are equally important, and some of them are more vulnerable under adversarial attack. Second, we develop a generative adversarial model to mimic the attack on EC. In particular, the generative model will simulate behaviors of both clean BPs and perturbed key BPs, and their derived graphs, and thus can launch effective attacks with less attention. We have conducted extensive experiments on eleven clustering benchmarks and have demonstrated that our approach is effective in attacking EC under both transductive and inductive settings.
Chetan Kumar, Deepak Kumar 0008, Ming Shao
WACV3
2022 Survey on the Analysis and Modeling of Visual Kinship: A Decade in the Making
abstract
Kinship recognition is a challenging problem with many practical applications. With much progress and milestones having been reached after ten years, we are now able to survey the research and create new milestones. We review the public resources and data challenges that enabled and inspired many to hone-in on the views of automatic kinship recognition in the visual domain. The different tasks are described in technical terms and syntax consistent across the problem domain and the practical value of each discussed and measured. State-of-the-art methods for visual kinship recognition problems, whether to discriminate between or generate from, are examined. As part of such, we review systems proposed as part of a recent data challenge held in conjunction with the 2020 IEEE Conference on Automatic Face and Gesture Recognition. We establish a stronghold for the state of progress for the different problems in a consistent manner. This survey will serve as the central resource for the work of the next decade to build upon. For the tenth anniversary, the demo code is provided for the various kin-based tasks. Detecting relatives with visual recognition and classifying the relationship is an area with high potential for impact in research and practice.
Joseph P. Robinson, Ming Shao, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Multiview Video-Based 3-D Pose Estimation of Patients in Computer-Assisted Rehabilitation Environment (CAREN)
abstract
The computer-assisted rehabilitation environment (CAREN) system plays an important role in the training of rehabilitation patients, where the capture of the patient's 3-D pose and gait is critical for assessing the patient's requirements for effective training. Vision-based methods are highly effective for this task due to their low cost, high speed, and noninterference. Although various general vision-based pose estimation methods were developed recently, their performance is limited in the CAREN system due to the specific environment. To address these problems, we propose an improved framework for accurate 2-D and 3-D pose estimation for the CAREN system through using multiview videos. First, for 2-D pose estimation, we propose a coarse-to-fine heatmap shrinking (CFHS) strategy that gradually reduces the kernel size of the heatmap of joints during training to improve the performance. Second, to further obtain 3-D pose estimations, we propose a novel spatial-temporal perception network that fuses the 2-D results from multiple views and multiple moments; multiview early fusion uses complementary spatial information from different views, and multimoment late fusion leverages temporal information from the sequential input for higher accuracy. The experimental results, based on CAREN videos of 225 orthopedic patients, showed that the accuracy of 2-D human pose estimations with the CFHS training strategy reached 99.85% [email protected]. For 3-D results, the mean per joint position error was 25.22 mm, and the 3DPCK reached 98.71%, which outperformed existing general video-based methods. The results showed that the proposed system is capable of estimating human poses with high accuracy for clinical applications.
Wei Xu 0046, Donghai Xiang, Guotai Wang, Ruisong Liao, Ming Shao, Kang Li 0004
IEEE Trans. Hum. Mach. Syst.5
2022 Families in Wild Multimedia: A Multimodal Database for Recognizing Kinship
abstract
Kinship, a soft biometric detectable in media, is fundamental for a myriad of use-cases. Despite the difficulty of detecting kinship, annual data challenges using still-images have consistently improved performances and attracted new researchers. Now, systems reach performance levels unforeseeable a decade ago, closing in on performances acceptable to deploy in practice. Like other biometric tasks, we expect systems can receive help from other modalities. We hypothesize that adding modalities toFamilies In the Wild(FIW), which has only still-images, will improve performance. Thus, to narrow the gap between research and reality and enhance the power of kinship recognition systems, we extend FIW with multimedia (MM) data (i.e., video, audio, and text captions). Specifically, we introduce the first publicly available multi-task MM kinship dataset. To buildFIW in Multimedia(FIW MM), we developed machinery to automatically collect, annotate, and prepare the data, requiring minimal human input and no financial cost. The proposed MM corpus allows the problem statements to be more realistic template-based protocols. We show significant improvements in all benchmarks with the added modalities. The results highlight edge cases to inspire future research with different areas of improvement. FIW MM supplies the data needed to increase the potential of automated systems to detect kinship in MM. It also allows experts from diverse fields to collaborate in novel ways.
Joseph P. Robinson, Zaid Khan 0001, Yu Yin 0001, Ming Shao, Yun Fu 0001
IEEE Trans. Multim.4
2022 Music2Dance: DanceNet for Music-Driven Dance Generation
abstract
Synthesize human motions from music (i.e., music to dance) is appealing and has attracted lots of research interests in recent years. It is challenging because of the requirement for realistic and complex human motions for dance, but more importantly, the synthesized motions should be consistent with the style, rhythm, and melody of the music. In this article, we propose a novel autoregressive generative model, DanceNet, to take the style, rhythm, and melody of music as the control signals to generate 3D dance motions with high realism and diversity. Due to the high long-term spatio-temporal complexity of dance, we propose the dilated convolution to improve the receptive field, and adopt the gated activation unit as well as separable convolution to enhance the fusion of motion features and control signals. To boost the performance of our proposed model, we capture several synchronized music-dance pairs by professional dancers and build a high-quality music-dance pair dataset. Experiments have demonstrated that the proposed method can achieve state-of-the-art results.
Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang 0001, Ming Shao, Si-Yu Xia
ACM Trans. Multim. Comput. Commun. Appl.5
2021 The 5th Recognizing Families in the Wild Data Challenge: Predicting Kinship from Faces
abstract
Recognizing Families In the Wild (RFIW), held as a data challenge in conjunction with the 16thIEEE International Conference on Automatic Face and Gesture Recognition (FG), is a large-scale, multi-track visual kinship recognition evaluation. For the fifth edition of RFIW, we continue to attract scholars, bring together professionals, publish new work, and discuss prospects. In this paper, we summarize submissions for the three tasks of this year's RFIW: specifically, we review the results for kinship verification, tri-subject verification, and family member search and retrieval. We look at the RFIW problem, share current efforts, and make recommendations for promising future directions.
Joseph P. Robinson, Can Qin, Ming Shao, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001
FG3
2021 Adversarial Attacks on Kinship Verification using Transformer
abstract
Visual kinship verification is one of the key research problems in computer vision with significant progress made in the past decade. Meanwhile, the harnessing of visual kinship models may lead to personal privacy leaking and raise people's concerns, especially the heavy social media users. One promising countermeasure is to overlay additional noise on images through the adversarial attack on kinship verification models to protect personal privacy. Motivated by the recent success of Transformer models in visual tasks, we propose a novel Transformer-based adversarial attack method named “Kinship-advTransGAN” towards the attack on kinship verification model. Essentially, Kinship-advTransGAN replaces the well-established CNN structure in conventional advGAN by TransGAN to generate adversarial samples with more sparse noise but comparable successful attacking rate. We verify our proposed method on a few open benchmarks, including FIW datasets and Kaggle Kinship Verification Challenges. Among these challenging tasks, we achieved surprisingly good performance: over 90% successful attacking rate on FIW datasets and 76.86% successful rate of attacks on Kaggle Kinship Verification Challenge, but with less visually perceivable noise on face images.
Jiaxuan Zhu, Ming Shao, Hong Pan 0001, Si-Yu Xia
FG2
2021 DNA-Net: Age and Gender Aware Kin Face Synthesizer
abstract
Visual kinship verification aims to detect blood relatives in facial images. Its practical application have motivated many researchers to focus on the topic as of recent. In this paper, we focus on a new view of visual kinship technology: kin-based face generation. Specifically, we propose a two-stage kin-face generation model to predict the appearance of a child given a pair of parents. The first stage includes a deep generative adversarial auto-encoder conditioned on ages and genders to map between facial appearance and high-level features. The second stage is our proposed DNA-Net, which serves as a transformation between the deep and genetic features based on a random selection process to fuse genes of a parent pair to form the genes of a child. We demonstrate the effectiveness of the proposed method quantitatively and qualitatively. Experiments validate that the proposed model synthesizes convincing kin-faces using both subjective and objective standards.
Pengyu Gao, Joseph P. Robinson, Jiaxuan Zhu, Ming Shao, Si-Yu Xia
ICME5
2021 Collaborative knowledge distillation for incomplete multi-view action prediction
Deepak Kumar 0008, Chetan Kumar, Ming Shao
Image Vis. Comput.3
2021 Cost-sensitive selection of variables by ensemble of model sequences
Donghui Yan, Songxiang Gu, Haiping Xu, Ming Shao
Knowl. Inf. Syst.5
2020 Adversary for Social Good: Protecting Familial Privacy through Joint Adversarial Attacks
abstract
Social media has been widely used among billions of people with dramatical participation of new users every day. Among them, social networks maintain the basic social characters and host huge amount of personal data. While protecting user sensitive data is obvious and demanding, information leakage due to adversarial attacks is somehow unavoidable, yet hard to detect. For example, implicit social relation such as family information may be simply exposed by network structure and hosted face images through off-the-shelf graph neural networks (GNN), which will be empirically proved in this paper. To address this issue, in this paper, we propose a novel adversarial attack algorithm for social good. First, we start from conventional visual family understanding problem, and demonstrate that familial information can easily be exposed to attackers by connecting sneak shots to social networks. Second, to protect family privacy on social networks, we propose a novel adversarial attack algorithm that produces both adversarial features and graph under a given budget. Specifically, both features on the node and edges between nodes will be perturbed gradually such that the probe images and its family information can not be identified correctly through conventional GNN. Extensive experiments on a popular visual social dataset have demonstrated that our defense strategy can significantly mitigate the impacts of family information leakage.
Chetan Kumar, Riazat Ryan, Ming Shao
AAAI3
2020 Localin Reshuffle Net: Toward Naturally and Efficiently Facial Image Blending
Chengyao Zheng, Si-Yu Xia, Joseph P. Robinson, Changsheng Lu, Wayne Wu, Chen Qian 0006, Ming Shao
ACCV (5)7
2020 Recognizing Families In the Wild (RFIW): The 4th Edition
abstract
Recognizing Families In the Wild (RFIW)- an annual large-scale, multi-track automatic kinship recognition evaluation- supports various visual kin-based problems on scales much higher than ever before. Organized in conjunction with the as a Challenge, RFIW provides a platform for publishing original work and the gathering of experts for a discussion of the next steps. This paper summarizes the supported tasks (i.e., kinship verification, tri-subject verification, and search & retrieval of missing children) in the evaluation protocols, which include the practical motivation, technical background, data splits, metrics, and benchmark results. Furthermore, top submissions (i.e., leader-board stats) are listed and reviewed as a high-level analysis on the state of the problem. In the end, the purpose of this paper is to describe the 2020 RFIW challenge, end-to-end, along with forecasts in promising future directions.
Joseph P. Robinson, Yu Yin 0001, Zaid Khan 0001, Ming Shao, Si-Yu Xia, Michael Stopa, Samson Timoner, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001
FG4
2020 Finding Achilles' Heel: Adversarial Attack on Multi-modal Action Recognition
abstract
Neural network-based models are notoriously known for their adversarial vulnerability. Recent adversarial machine learning mainly focused on images, where a small perturbation can be simply added to fool the learning model. Very recently, this practice has been explored in human action video attacks by adding perturbation to key frames. Unfortunately, frame selection is usually computationally expensive in run-time, and adding noises to all frames is unrealistic, either. In this paper, we present a novel yet efficient approach to address this issue. Multi-modal video data such as RGB, depth and skeleton data have been widely used for human action modeling, and they have been demonstrated with superior performance than a single modality. Interestingly, we observed that the skeleton data is more "vulnerable" under adversarial attack, and we propose to leverage this "Achilles' Heel" to attack multi-modal video data. In particular, first, an adversarial learning paradigm is designed to perturb skeleton data for a specific action under a black box setting, which highlights how body joints and key segments in videos are subject to attack. Second, we propose a graph attention model to explore the semantics between segments from different modalities and within a modality. Third, the attack will be launched in run-time on all modalities through the learned semantics. The proposed method has been extensively evaluated on multi-modal visual action datasets, including PKU-MMD and NTU-RGB+D to validate its effectiveness.
Deepak Kumar 0008, Chetan Kumar, Chun-Wei Seah, Si-Yu Xia, Ming Shao
ACM Multimedia5
2020 Arc-Support Line Segments Revisited: An Efficient High-Quality Ellipse Detection
abstract
Over the years many ellipse detection algorithms spring up and are studied broadly, while the critical issue of detecting ellipses accurately and efficiently in real-world images remains a challenge. In this paper, we propose a valuable industry-oriented ellipse detector by arc-support line segments, which simultaneously reaches high detection accuracy and efficiency. To simplify the complicated curves in an image while retaining the general properties including convexity and polarity, the arc-support line segments are extracted, which grounds the successful detection of ellipses. The arc-support groups are formed by iteratively and robustly linking the arc-support line segments that latently belong to a common ellipse. Afterward, two complementary approaches, namely, locally selecting the arc-support group with higher saliency and globally searching all the valid paired groups, are adopted to fit the initial ellipses in a fast way. Then, the ellipse candidate set can be formulated by hierarchical clustering of 5D parameter space of initial ellipses. Finally, the salient ellipse candidates are selected and refined as detections subject to the stringent and effective verification. Extensive experiments on three public datasets are implemented and our method achieves the best F-measure scores compared to the state-of-the-art methods. The source code is available at https://github.com/AlanLuSun/High-quality-ellipse-detection.
Changsheng Lu, Si-Yu Xia, Ming Shao, Yun Fu 0001
IEEE Trans. Image Process.3
2019 Weakly Supervised Vitiligo Segmentation in Skin Image through Saliency Propagation
abstract
Vitiligo is a skin disorder where pale or white patches develop due to the lack or absence of melanocytes. Vitiligo affects around 0.5% to 1% of the world's population, and it may have a profound psychological impact on patients' quality of life. In this paper, we present a novel weakly supervised framework to segment vitiligo regions with high quality, which is a fundamental task for the assessment of vitiligo. The proposed framework starts with pre-training a classification network using only image-level labels. Then we observed that the activation map obtained from the image classification network could be further exploited and introduced into the saliency propagation process as useful information. Finally, the saliency propagation process is performed on the graph built on superpixels to obtain a meaningful saliency map. These three steps lead to a compelling yet elegant method. Moreover, we propose a new large vitiligo image dataset named Vit2019. To the best of our knowledge, this is currently the first dataset for image segmentation of vitiligo diseases. Experimental results demonstrate the superiority of the proposed model over state-of-the-arts.
Zhangxing Bian, Si-Yu Xia, Ming Shao
BIBM4
2019 CTC-Attention based Non-Parametric Inference Modeling for Clinical State Progression
abstract
Predictive modeling of patient state to state medical conditions in ICU is a critical yet challenging task in health informatics and machine learning. Prior critical stages from the same ICU admission may contribute differently to the next stages. That said, stages are interdependent, and disease progression is a multi-step temporal observation. In this paper, we formally name this problem as “Clinical State Progression Prediction (CSPP).” Conventional temporal modeling may fit well to predictions of fixed size observations and number of stages, but have troubles and less flexibility when addressing CSPP. To that end, we cast this problem as multi-label learning on time series data in which each stage is marked by a label. The implementation of entire framework includes two phases. In learning, an RNN based Encoder-Decoder deep model is developed for basic temporal modeling. In addition, Attention mechanism and Connectionist Temporal Classification (CTC) are integrated to explicitly model the temporal dependency as well as monotonic relation between input time series and output label space. In inference, based on the observed multi-stage labels, a non-parametric retrieval is carried out first to build up the reference patient records. Then, based on CTC-Attention learning model, consistent progressions are computed and ranked to contribute to the prediction of the clinical state progression in the next few hours. Extensive experiments on MIMIC III and Parkinson datasets demonstrate that the proposed predictive modeling for CSPP outperforms state-of-the-art works on Sepsis, Kidney-Sepsis-Mortality, Heart-Sepsis-Mortality, and Parkinson Progression.
Riazat Ryan, Handong Zhao, Ming Shao
IEEE BigData3
2019 Fast Facial Image Analogy with Spatial Guidance
abstract
This paper proposes a novel method for fast facial image analogy with spatial guidance. Given a facial image A and another one B in a different style (color, tone, or texture), we are allowed to render the face in A with style B to output a stylized facial image A' in a timely manner. For traditional image analogy, such a process could take an unbearable time and considerable computing resources. In this paper, for the first time, we make it possible to do image analogy at a speed much faster than the state-of-the-art method. Specifically, we first extract deep image features using a VGG-19 encoder, and implement patch-match with the guidance of facial landmarks. Then Procrustes analysis is applied to accelerate the program as a coarse-to-fine strategy. We repeat the above process at each layer of VGG-19 and decode image A' from bottom to top. Experimental results show that our method not only spends much less time but also provides high-quality image analogy results. Moreover, our method can be naturally extended to many face-related applications including but not limited to face swapping, makeup and style to photo.
Chengyao Zheng, Si-Yu Xia, Ming Shao, Yun Fu 0001
FG3
2019 Generative Zero-Shot Learning via Low-Rank Embedded Semantic Dictionary
abstract
Zero-shot learning for visual recognition, which approaches identifying unseen categories through a shared visual-semantic function learned on the seen categories and is expected to well adapt to unseen categories, has received considerable research attention most recently. However, the semantic gap between discriminant visual features and their underlying semantics is still the biggest obstacle, because there usually exists domain disparity across the seen and unseen classes. To deal with this challenge, we design two-stage generative adversarial networks to enhance the generalizability of semantic dictionary through low-rank embedding for zero-shot learning. In detail, we formulate a novel framework to simultaneously seek a two-stage generative model and a semantic dictionary to connect visual features with their semantics under a low-rank embedding. Our first-stage generative model is able to augment more semantic features for the unseen classes, which are then used to generate more discriminant visual features in the second stage, to expand the seen visual feature space. Therefore, we will be able to seek a better semantic dictionary to constitute the latent basis for the unseen classes based on the augmented semantic and visual data. Finally, our approach could capture a variety of visual characteristics from seen classes that are "ready-to-use" for new classes. Extensive experiments on four zero-shot benchmarks demonstrate that our proposed algorithm outperforms the state-of-the-art zero-shot algorithms.
Zhengming Ding, Ming Shao, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Robust Discriminative Metric Learning for Image Representation
abstract
Metric learning has attracted significant attention in the past decades, because of its appealing advances in various real-world tasks, e.g., person re-identification and face recognition. Traditional supervised metric learning attempts to seek a discriminative metric, which could minimize the pairwise distance of within-class data samples, while maximizing the pairwise distance of data samples from various classes. However, it is still a challenge to build a robust and discriminative metric, especially for corrupted data in the real-world application. In this paper, we propose a Robust Discriminative Metric Learning algorithm through fast low-rank representation and denoising strategy. To be specific, the metric learning problem is guided by a discriminative regularization by incorporating the pair-wise or class-wise information. Moreover, the low-rank basis learning is jointly optimized with the metric to better uncover the global data structure and remove noise. Furthermore, the fast low-rank representation is implemented to mitigate the computational burden and ensure the scalability on large-scale datasets. Finally, we evaluate our learned metric on several challenging tasks, e.g., face recognition/verification, object recognition, image clustering, and person re-identification. The experimental results verify the effectiveness of our proposed algorithm in comparison to many metric learning algorithms, even deep learning ones.
Zhengming Ding, Ming Shao, Wonjun Hwang, Sungjoo Suh, Jae-Joon Han, Changkyu Choi, Yun Fu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Structure-Preserved Unsupervised Domain Adaptation
abstract
Domain adaptation has been a primal approach to addressing the issues by lack of labels in many data mining tasks. Although considerable efforts have been devoted to domain adaptation with promising results, most existing work learns a classifier on a source domain and then predicts the labels for target data, where only the instances near the boundary determine the hyperplane and the whole structure information is ignored. Moreover, little work has been done regarding to multi-source domain adaptation. To that end, we develop a novel unsupervised domain adaptation framework, which ensures the whole structure of source domains is preserved to guide the target structure learning in a semi-supervised clustering fashion. To our knowledge, this is the first time when the domain adaptation problem is re-formulated as a semi-supervised clustering problem with target labels as missing values. Furthermore, by introducing an augmented matrix, a non-trivial solution is designed, which can be exactly mapped into a K-means-like optimization problem with modified distance function and update rule for centroids in an efficient way. Extensive experiments on several widely-used databases show the substantial improvements of our proposed approach over the state-of-the-art methods.
Hongfu Liu 0001, Ming Shao, Zhengming Ding, Yun Fu 0001
IEEE Trans. Knowl. Data Eng.2
2019 Feature Selection with Unsupervised Consensus Guidance
abstract
Most of the unsupervised feature selection methods employ pseudo labels generated by clustering to guide the feature selection; however, noisy and irrelevant features degrade the cluster structure, which is ineffective to supervise feature selection. In light of this, we propose the Consensus Guided Unsupervised Feature Selection (CGUFS) framework, which introduces consensus clustering to generate pseudo labels for feature selection. Generally speaking, multiple diverse basic partitions are generated from the data and the consensus clustering is employed to provide the high-quality and robust partition to guide the feature selection in a one-step framework. In addition, complex constraints such as non-negative are removed due to the crisp indicators of consensus clustering. Based on the CGUFS framework, two formulations are put forward by using the utility function and co-association matrix, respectively, and we propose the (weighted) K-means-like optimization solution for efficient solutions with theoretical supports. Moreover, we extend the CGUFS framework to handle multi-view data feature selection. Extensive experiments on several singleview and multi-view data mining data sets in different domains demonstrate that our methods outperform the most recent state-ofthe-art work in terms of effectiveness and efficiency. Some important impact factors and model parameters within CGUFS are thoroughly discussed for practical use.
Hongfu Liu 0001, Ming Shao, Yun Fu 0001
IEEE Trans. Knowl. Data Eng.2
2018 Album to Family Tree: A Graph Based Method for Family Relationship Recognition
Si-Yu Xia, Ming Shao, Yun Fu 0001
ACCV (2)3
2018 Divide-and-Conquer Kronecker Product Decomposition for Memory-Efficient Graph Approximation
abstract
Graphs are a widely used data structure for modeling objects in several domains, ranging from social media analytics to molecular biology. Recently, with the surge of big data, finding compact representations of large graphs has become an integral part of large-scale data analysis. To that end, we explore the effectiveness of the Kronecker Product SVD (KPSVD) for scalable sparse graph approximation, in a Divide-and-Conquer fashion. The KPSVD seeks to represent a graph as the sum of the Kronecker Product (KP) of several smaller factor matrices. In our method, we first partition the graph into inter-graph and intra-graph, and then use the Van Loan-Pitsianis (VLP) SVD-based algorithm to find the low-rank KPSVD of each subgraph to approximate the intra-graph. We use both the inter-graph and the intra-graph to find the approximation for the whole graph. We perform experiments on small-scale to large-scale real-world datasets to test the effectiveness of our method in terms of approximation error and spectral clustering results. The experiments demonstrate that our approach can provide better or competitive performance in terms of approximation error and clustering results while saving memory, compared to other state-of-the-art algorithms.
Venkata Suhas Maringanti, Ming Shao
IEEE BigData2
2018 Deep Evolutionary 3D Diffusion Heat Maps for Large-pose Face Alignment
Bin Sun 0002, Ming Shao, Si-Yu Xia, Yun Fu 0001
BMVC2
2018 Graph Adaptive Knowledge Transfer for Unsupervised Domain Adaptation
Zhengming Ding, Sheng Li 0001, Ming Shao, Yun Fu 0001
ECCV (2)3
2018 IEEE Std P1838's flexible parallel port and its specification with Google's protocol buffers
abstract
IEEE Std P1838 is the DfT standard-under-development for 3D test access into dies meant to be used in 3D multi-die stack assemblies. P1838 is the first DfT standard to include a flexible parallel port (FPP): an optional, scalable multi-bit ('parallel') test access mechanism, offering higher test access bandwidth compared to the mandatory one-bit ('serial') port. In this paper, we describe P1838's FPP and propose a formal FPP specification language based on Google's Protocol Buffers (PBs), that potentially could become part of the standard. For a realistic example FPP, we provide its formal specification. Finally, we report on a demonstrator software tool, developed by using PBs-generated data access routines, that converts an FPP specification into a corresponding Verilog netlist.
Yu Li 0007, Ming Shao, Hailong Jiao, Adam Cron, Sandeep Bhatia, Erik Jan Marinissen
ETS2
2018 Robust Multi-view Representation: A Unified Perspective from Multi-view Learning to Domain Adaption
abstract
Multi-view data are extensively accessible nowadays thanks to various types of features, different view-points and sensors which tend to facilitate better representation in many key applications. This survey covers the topic of robust multi-view data representation, centered around several major visual applications. First of all, we formulate a unified learning framework which is able to model most existing multi-view learning and domain adaptation in this line. Following this, we conduct a comprehensive discussion across these two problems by reviewing the algorithms along these two topics, including multi-view clustering, multi-view classification, zero-shot learning, and domain adaption. We further present more practical challenges in multi-view data analysis. Finally, we discuss future research including incomplete, unbalance, large-scale multi-view learning. This would benefit AI community from literature review to future direction.
Zhengming Ding, Ming Shao, Yun Fu 0001
IJCAI2
2018 To Recognize Families In the Wild: A Machine Vision Tutorial
abstract
Automatic kinship recognition has relevance in an abundance of applications. For starters, aiding forensic investigations, as kinship is a powerful cue that could narrow the search space (e.g., knowledge that the 'Boston Bombers' were brothers could have helped identify the suspects sooner). In short, there are many beneficiaries that could result from such technologies: whether the consumer (e.g., automatic photo library management), scholar (e.g., historic lineage & genealogical studies), data analyzer (e.g., social-media- based analysis), investigator (e.g., cases of missing children and human trafficking. For instance, it is unlikely that a missing child found online would be in any database, however, more than likely a family member would be), or even refugees. Besides application- based problems, and as already hinted, kinship is a powerful cue that could serve as a face attribute capable of greatly reducing the search space in more general face-recognition problems. In this tutorial, we will introduce the background information, progress leading us up to these points, several current state-of-the-art algorithms spanning various views of the kinship recognition problem (e.g., verification, classification, tri-subject). We will then cover our large-scale Families In the Wild (FIW) image collection, several challenge competitions it as been used in, along with the top per- forming deep learning approaches. The tutorial will end with a discussion about future research directions and practical use-cases.
Joseph P. Robinson, Ming Shao, Yun Fu 0001
ACM Multimedia2
2018 Graph Based Family Relationship Recognition from a Single Image
Si-Yu Xia, Ming Shao
PRICAI (1)5
2018 Infinite ensemble clustering
Hongfu Liu 0001, Ming Shao, Sheng Li 0001, Yun Fu 0001
Data Min. Knowl. Discov.2
2018 Person Re-Identification by Cross-View Multi-Level Dictionary Learning
abstract
Person re-identification plays an important role in many safety-critical applications. Existing works mainly focus on extracting patch-level features or learning distance metrics. However, the representation power of extracted features might be limited, due to the various viewing conditions of pedestrian images in complex real-world scenarios. To improve the representation power of features, we learn discriminative and robust representations via dictionary learning in this paper. First, we propose a Cross-view Dictionary Learning (CDL) model, which is a general solution to the multi-view learning problem. Inspired by the dictionary learning based domain adaptation, CDL learns a pair of dictionaries from two views. In particular, CDL adopts a projective learning strategy, which is more efficient than the optimization in traditional dictionary learning. Second, we propose a Cross-view Multi-level Dictionary Learning (CMDL) approach based on CDL. CMDL contains dictionary learning models at different representation levels, including image-level, horizontal part-level, and patch-level. The proposed models take advantages of the view-consistency information, and adaptively learn pairs of dictionaries to generate robust and compact representations for pedestrian images. Third, we incorporate a discriminative regularization term to CMDL, and propose a CMDL-Dis approach which learns pairs of discriminative dictionaries in image-level and part-level. We devise efficient optimization algorithms to solve the proposed models. Finally, a fusion strategy is utilized to generate the similarity scores for test images. Experiments on the public VIPeR, CUHK Campus, iLIDS, GRID and PRID450S datasets show that our approach achieves the state-of-the-art performance.
Sheng Li 0001, Ming Shao, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Learning Consensus Representation for Weak Style Classification
abstract
Style classification (e.g., Baroque and Gothic architecture style) is grabbing increasing attention in many fields such as fashion, architecture, and manga. Most existing methods focus on extracting discriminative features from local patches or patterns. However, the spread out phenomenon in style classification has not been recognized yet. It means that visually less representative images in a style class are usually very diverse and easily getting misclassified. We name them weak style images. Another issue when employing multiple visual features towards effective weak style classification is lack of consensus among different features. That is, weights for different visual features in the local patch should have been allocated similar values. To address these issues, we propose a Consensus Style Centralizing Auto-Encoder (CSCAE) for learning robust style features representation, especially for weak style classification. First, we propose a Style Centralizing Auto-Encoder (SCAE) which centralizes weak style features in a progressive way. Then, based on SCAE, we propose both the non-linear and linear version CSCAE which adaptively allocate weights for different features during the progressive centralization process. Consensus constraints are added based on the assumption that the weights of different features of the same patch should be similar. Specifically, the proposed linear counterpart of CSCAE motivated by the "shared weights" idea as well as group sparsity improves both efficacy and efficiency. For evaluations, we experiment extensively on fashion, manga and architecture style classification problems. In addition, we collect a new dataset-Online Shopping, for fashion style classification, which will be publicly available for vision based fashion style research. Experiments demonstrate the effectiveness of the SCAE and CSCAE on both public and newly collected datasets when compared with the most recent state-of-the-art works.
Shuhui Jiang, Ming Shao, Chengcheng Jia, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Visual Kinship Recognition of Families in the Wild
abstract
We present the largest database for visual kinship recognition, Families In the Wild (FIW), with over 13,000 family photos of 1,000 family trees with 4-to-38 members. It took only a small team to build FIW with efficient labeling tools and work-flow. To extend FIW, we further improved upon this process with a novel semi-automatic labeling scheme that used annotated faces and unlabeled text metadata to discover labels, which were then used, along with existing FIW data, for the proposed clustering algorithm that generated label proposals for all newly added data-both processes are shared and compared in depth, showing great savings in time and human input required. Essentially, the clustering algorithm proposed is semi-supervised and uses labeled data to produce more accurate clusters. We statistically compare FIW to related datasets, which unarguably shows enormous gains in overall size and amount of information encapsulated in the labels. We benchmark two tasks, kinship verification and family classification, at scales incomparably larger than ever before. Pre-trained CNN models fine-tuned on FIW outscores other conventional methods and achieved state-of-the art on the renowned KinWild datasets. We also measure human performance on kinship recognition and compare to a fine-tuned CNN.
Joseph P. Robinson, Ming Shao, Yue Wu 0008, Hongfu Liu 0001, Timothy Gillis, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Stacked Denoising Tensor Auto-Encoder for Action Recognition With Spatiotemporal Corruptions
abstract
Spatially or temporally corrupted action videos are impractical for recognition via vision or learning models. It usually happens when streaming data are captured from unintended moving cameras, which bring occlusion or camera vibration and accordingly result in arbitrary loss of spatiotemporal information. In reality, it is intractable to deal with both spatial and temporal corruptions at the same time. In this paper, we propose a coupled stacked denoising tensor auto-encoder (CSDTAE) model, which approaches this corruption problem in a divide-and-conquer fashion by jointing both the spatial and temporal schemes together. In particular, each scheme is a SDTAE designed to handle either spatial or temporal corruption, respectively. SDTAE is composed of several blocks, each of which is a denoising tensor auto-encoder (DTAE). Therefore, CSDTAE is designed based on several DTAE building blocks to solve the spatiotemporal corruption problem simultaneously. In one DTAE, the video features are represented as a high-order tensor to preserve the spatiotemporal structure of data, where the temporal and spatial information are processed separately in different hidden layers via tensor unfolding. In summary, DTAE explores the spatial and temporal structure of the tensor representation, and SDTAE handles different corrupted ratios progressively to extract more discriminative features. CSDTAE couples the temporal and spatial corruptions of the same data through a thorough step-by-step procedure based on canonical correlation analysis, which integrates the two sub-problems into one problem. The key point is solving the spatiotemporal corruption in one model by considering them as noises in either spatial or temporal direction. Extensive experiments on three action data sets demonstrate the effectiveness of our model, especially when large volumes of corruption in the video.
Chengcheng Jia, Ming Shao, Sheng Li 0001, Handong Zhao, Yun Fu 0001
IEEE Trans. Image Process.2
2018 Multi-View Low-Rank Analysis with Applications to Outlier Detection
abstract
Detecting outliers or anomalies is a fundamental problem in various machine learning and data mining applications. Conventional outlier detection algorithms are mainly designed for single-view data. Nowadays, data can be easily collected from multiple views, and many learning tasks such as clustering and classification have benefited from multi-view data. However, outlier detection from multi-view data is still a very challenging problem, as the data in multiple views usually have more complicated distributions and exhibit inconsistent behaviors. To address this problem, we propose a multi-view low-rank analysis (MLRA) framework for outlier detection in this article. MLRA pursuits outliers from a new perspective, robust data representation. It contains two major components. First, the cross-view low-rank coding is performed to reveal the intrinsic structures of data. In particular, we formulate a regularized rank-minimization problem, which is solved by an efficient optimization algorithm. Second, the outliers are identified through an outlier score estimation procedure. Different from the existing multi-view outlier detection methods, MLRA is able to detect two different types of outliers from multiple views simultaneously. To this end, we design a criterion to estimate the outlier scores by analyzing the obtained representation coefficients. Moreover, we extend MLRA to tackle the multi-view group outlier detection problem. Extensive evaluations on seven UCI datasets, the MovieLens, the USPS-MNIST, and the WebKB datasets demon strate that our approach outperforms several state-of-the-art outlier detection methods.
Sheng Li 0001, Ming Shao, Yun Fu 0001
ACM Trans. Knowl. Discov. Data2
2018 Incomplete Multisource Transfer Learning
abstract
Transfer learning is generally exploited to adapt well-established source knowledge for learning tasks in weakly labeled or unlabeled target domain. Nowadays, it is common to see multiple sources available for knowledge transfer, each of which, however, may not include complete classes information of the target domain. Naively merging multiple sources together would lead to inferior results due to the large divergence among multiple sources. In this paper, we attempt to utilize incomplete multiple sources for effective knowledge transfer to facilitate the learning task in target domain. To this end, we propose an incomplete multisource transfer learning through two directional knowledge transfer, i.e., cross-domain transfer from each source to target, and cross-source transfer. In particular, in cross-domain direction, we deploy latent low-rank transfer learning guided by iterative structure learning to transfer knowledge from each single source to target domain. This practice reinforces to compensate for any missing data in each source by the complete target data. While in cross-source direction, unsupervised manifold regularizer and effective multisource alignment are explored to jointly compensate for missing data from one portion of source to another. In this way, both marginal and conditional distribution discrepancy in two directions would be mitigated. Experimental results on standard cross-domain benchmarks and synthetic data sets demonstrate the effectiveness of our proposed model in knowledge transfer from incomplete multiple sources.
Zhengming Ding, Ming Shao, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Probabilistic Low-Rank Multitask Learning
abstract
In this paper, we consider the problem of learning multiple related tasks simultaneously with the goal of improving the generalization performance of individual tasks. The key challenge is to effectively exploit the shared information across multiple tasks as well as preserve the discriminative information for each individual task. To address this, we propose a novel probabilistic model for multitask learning (MTL) that can automatically balance between low-rank and sparsity constraints. The former assumes a low-rank structure of the underlying predictive hypothesis space to explicitly capture the relationship of different tasks and the latter learns the incoherent sparse patterns private to each task. We derive and perform inference via variational Bayesian methods. Experimental results on both regression and classification tasks on real-world applications demonstrate the effectiveness of the proposed method in dealing with the MTL problems.
Yu Kong 0001, Ming Shao, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Collaborative Random Faces-Guided Encoders for Pose-Invariant Face Representation Learning
abstract
Learning discriminant face representation for pose-invariant face recognition has been identified as a critical issue in visual learning systems. The challenge lies in the drastic changes of facial appearances between the test face and the registered face. To that end, we propose a high-level feature learning framework called "collaborative random faces (RFs)-guided encoders" toward this problem. The contributions of this paper are three fold. First, we propose a novel supervised autoencoder that is able to capture the high-level identity feature despite of pose variations. Second, we enrich the identity features by replacing the target values of conventional autoencoders with random signals (RFs in this paper), which are unique for each subject under different poses. Third, we further improve the performance of the framework by incorporating deep convolutional neural network facial descriptors and linking discriminative identity features from different RFs for the augmented identity features. Finally, we conduct face identification experiments on Multi-PIE database, and face verification experiments on labeled faces in the wild and YouTube Face databases, where face recognition rate and verification accuracy with Receiver Operating Characteristic curves are rendered. In addition, discussions of model parameters and connections with the existing methods are provided. These experiments demonstrate that our learning system works fairly well on handling pose variations.
Ming Shao, Yizhe Zhang 0001, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2017 Cross-database mammographic image analysis through unsupervised domain adaptation
abstract
World Health Organization report shows 519,000 deaths due to breast cancer in 2014 and it was much more in 2008. Therefore, it is required to take early steps in detection and diagnosis of breast cancer to decrease the associated death rate. Computer Aided Diagnosis (CAD) is useful in mass screening of breast cancer datasets. Data mining and machine learning technologies have already achieved significant success in many knowledge engineering areas including classification, regression and clustering, and most recently, have been employed to assist the diagnosis of cancers with promising outcomes. Traditional machine learning models are characterized by training and testing data with the same input feature space and data distribution. But when distribution changes, most machine learning models need to be modified or rebuilt from scratch to work on newly collected data. In many real world applications, it is expensive or impossible to recollect the needed data and rebuild the models. Therefore, there is a need to create high-performance learners trained with more easily obtained data from different domains. This methodology is referred as Transfer Learning. In this paper, we explore the usage of transfer learning, specially, unsupervised domain adaptation for breast cancer diagnosis to address the issues of fewer training data on target image dataset. On the strength of recent developed deep descriptors, we are able to adapt recent transfer learning methodologies, e.g., TCA (Transfer Component Analysis), CORAL (Correlation Alignment), BDA(Balanced Distribution Adaptation) to breast cancer diagnosis across multiple mammographic image databases including CBIS-DDSM, InBreast, MIAS, etc, and evaluate their performance. Experiments demonstrate that, without any labels in the target database, transfer learning is able to help improve the classification accuracy.
Deepak Kumar 0008, Chetan Kumar, Ming Shao
IEEE BigData3
2017 Low-Rank Embedded Ensemble Semantic Dictionary for Zero-Shot Learning
abstract
Zero-shot learning for visual recognition has received much interest in the most recent years. However, the semantic gap across visual features and their underlying semantics is still the biggest obstacle in zero-shot learning. To fight off this hurdle, we propose an effective Low-rank Embedded Semantic Dictionary learning (LESD) through ensemble strategy. Specifically, we formulate a novel framework to jointly seek a low-rank embedding and semantic dictionary to link visual features with their semantic representations, which manages to capture shared features across different observed classes. Moreover, ensemble strategy is adopted to learn multiple semantic dictionaries to constitute the latent basis for the unseen classes. Consequently, our model could extract a variety of visual characteristics within objects, which can be well generalized to unknown categories. Extensive experiments on several zero-shot benchmarks verify that the proposed model can outperform the state-of-the-art approaches.
Zhengming Ding, Ming Shao, Yun Fu 0001
CVPR2
2017 Circle detection by arc-support line segments
abstract
Circle detection is fundamental in both object detection and high accuracy localization in visual control systems. We propose a novel method for circle detection by analysing and refining arc-support line segments. The key idea is to use line segment detector to extract the arc-support line segments which are likely to make up the circle, instead of all line segments. Each couple of line segments is analyzed to form a valid pair and followed by generating initial circle set. Through the mean shift clustering, the circle candidates are generated and verified based on the geometric attributes of circle edge. Finally, twice circle fitting is applied to increase the accuracy for circle locating and radius measuring. The experimental results demonstrate that the proposed method performs better than other well known approaches on circles that are incomplete, occluded, blurry and over-illumination. Moreover, our method shows significant improvement in accuracy, robustness and efficiency on the industrial Printed Circuit Board (PCB) images as well as the synthesized, natural and complicated images.
Changsheng Lu, Si-Yu Xia, Wanming Huang, Ming Shao, Yun Fu 0001
ICIP4
2017 Family Photo Recognition via Multiple Instance Learning
abstract
Family photo recognition is an important task in social media analytics. Previous methods use singleton global features and conventional binary classifiers to distinguish family group photos from non-family ones. Different from them, we propose a novel family recognition approach with three dedicated local representations under Multiple Instance Learning framework, where geometry, kinship and semantic features are integrated to overcome issues in the previous work. Experimental results show that our method achieves the state-of-the-art result among global-feature models.
Junkang Zhang, Si-Yu Xia, Ming Shao, Yun Fu 0001
ICMR3
2017 RFIW: Large-Scale Kinship Recognition Challenge
abstract
Recognizing Families In the Wild (RFIW) is organized as a Data Challenge Workshop in conjunction with ACM MM 2017. The workshop is scheduled for the afternoon of October 27th. RFIW is the 1st large-scale kinship recognition challenge and is made up of 2 tracks, kinship verification and family classification. In total, 12 final submissions were made. This big data challenge was achieved with our FIW dataset which is, by far, the largest image collection of its kind. Potential next steps for FIW are abundant.
Joseph P. Robinson, Ming Shao, Handong Zhao, Yue Wu 0008, Timothy Gillis, Yun Fu 0001
ACM Multimedia2
2017 Sparse Canonical Temporal Alignment With Deep Tensor Decomposition for Action Recognition
abstract
In this paper we solve three problems in action recognition: sub-action, multi-subject, and multi-modality, by reducing the diversity of intra-class samples. The main stage contains canonical temporal alignment and key frames selection. As we know, temporal alignment aims to reduce the diversity of intra-class samples, however, dense frames may yield misalignment or overlapped alignment and decrease recognition performance. To overcome this problem, we propose a Sparse Canonical Temporal Alignment (SCTA) method which selects and aligns key frames from two sequences to reduce diversity. To extract better features from the key frames, we propose a Deep Non-negative Tensor Factorization (DNTF) method to find a tensor subspace integrated with SCTA scheme. First we model an action sequence as a third-order tensor with spatiotemporal structure. Then we design a DNTF scheme to find a tensor subspace in both spatial and temporal directions. Particularly, in the first layer the original tensor is decomposed into two lowrank tensors by Non-negative Tensor Factorization (NTF), and in the second layer each low-rank tensor is further decomposed by Tensor-Train (TT) for time efficiency. Finally, our framework composed of SCTA and DNTF could solve the three problems and extract effective features for action recognition. Experiments on synthetic data, MSRDailyActivity3D and MSRActionPairs datasets show that our method works better than competitive methods in terms of accuracy.
Chengcheng Jia, Ming Shao, Yun Fu 0001
IEEE Trans. Image Process.2
2017 Cross-Modality Feature Learning Through Generic Hierarchical Hyperlingual-Words
abstract
Recognizing facial images captured under visible light has long been discussed in the past decades. However, there are many impact factors that hinder its successful application in real-world, e.g., illumination, pose variations. Recent work has concentrated on different spectrals, i.e., near infrared, that can only be perceived by specifically designed device to avoid the illumination problem. However, this inevitably introduces a new problem, namely, cross-modality classification. In brief, images registered in the system are in one modality, while images that captured momentarily used as the tests are in another modality. In addition, there could be many within-modality variations-pose and expression-leading to a more complicated problem for the researchers. To address this problem, we propose a novel framework called hierarchical hyperlingual-words (Hwords) in this paper. First, we design a novel structure, called generic Hwords, to capture the high-level semantics across different modalities and within each modality in weakly supervised fashion, meaning only modality pair and variations information are needed in the training. Second, to improve the discriminative power of Hwords, we propose a novel distance metric through the hierarchical structure of Hwords. Extensive experiments on multimodality face databases demonstrate the superiority of our method compared with the state-of-the-art works on face recognition tasks subject to pose and expression variations.
Ming Shao, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2016 Consensus Style Centralizing Auto-Encoder for Weak Style Classification
abstract
Style classification (e.g., architectural, music, fashion) attracts an increasing attention in both research and industrial fields. Most existing works focused on low-level visual features composition for style representation. However, little effort has been devoted to automatic mid-level or high-level style features learning by reorganizing low-level descriptors. Moreover, styles are usually spread out and not easy to differentiate from one to another. In this paper, we call these less representative images as weak style images. To address these issues, we propose a consensus style centralizing auto-encoder (CSCAE) to extract robust style features to facilitate weak style classification. CSCAE is the ensemble of several style centralizing auto-encoders (SCAEs) with consensus constraint. Each SCAE centralizes each feature of certain category in a progressive way. We apply our method in fashion style classification and manga style classification as two example applications. In addition, we collect a new dataset, Online Shopping, for fashion style classification evaluation, which will be publicly available for vision based fashion style research. Experiments demonstrate the effectiveness of SCAE and CSCAE on both public and newly collected datasets when compared with the most recent state-of-the-art works.
Shuhui Jiang, Ming Shao, Chengcheng Jia, Yun Fu 0001
AAAI2
2016 Consensus Guided Unsupervised Feature Selection
abstract
Feature selection has been widely recognized as one of the key problems in data mining and machine learning community, especially for high-dimensional data with redundant information, partial noises and outliers. Recently, unsupervised feature selection attracts substantial research attentions since data acquisition is rather cheap today but labeling work is still expensive and time consuming. This is specifically useful for effective feature selection of clustering tasks. Recent works using sparse projection with pre-learned pseudo labels achieve appealing results; however, they generate pseudo labels with all features so that noisy and ineffective features degrade the cluster structure and further harm the performance of feature selection; besides, these methods suffer from complex composition of multiple constraints and computational inefficiency, e.g., eigen-decomposition. Differently, in this work we introduce consensus clustering for pseudo labeling, which gets rid of expensive eigen-decomposition and provides better clustering accuracy with high robustness. In addition, complex constraints such as non-negative are removed due to the crisp indicators of consensus clustering. Specifically, we propose one efficient formulation for our unsupervised feature selection by using the utility function and provide theoretical analysis on optimization rules and model convergence. Extensive experiments on several popular data sets demonstrate that our methods are superior to the most recent state-of-the-art works in terms of NMI.
Hongfu Liu 0001, Ming Shao, Yun Fu 0001
AAAI2
2016 Spectral Bisection Tree Guided Deep Adaptive Exemplar Autoencoder for Unsupervised Domain Adaptation
abstract
Learning with limited labeled data is always a challenge in AI problems, and one of promising ways is transferring well-established source domain knowledge to the target domain, i.e., domain adaptation. In this paper, we extend the deep representation learning to domain adaptation scenario, and propose a novel deep model called ``Deep Adaptive Exemplar AutoEncoder (DAE$^2$)''. Different from conventional denoising autoencoders using corrupted inputs, we assign semantics to the input-output pairs of the autoencoders, which allow us to gradually extract discriminant features layer by layer. To this end, first, we build a spectral bisection tree to generate source-target data compositions as the training pairs fed to autoencoders. Second, a low-rank coding regularizer is imposed to ensure the transferability of the learned hidden layer. Finally, a supervised layer is added on top to transform learned representations into discriminant features. The problem above can be solved iteratively in an EM fashion of learning. Extensive experiments on domain adaptation tasks including object, handwritten digits, and text data classifications demonstrate the effectiveness of the proposed method.
Ming Shao, Zhengming Ding, Handong Zhao, Yun Fu 0001
AAAI1
2016 A Multi-stream Bi-directional Recurrent Neural Network for Fine-Grained Action Detection
abstract
We present a multi-stream bi-directional recurrent neural network for fine-grained action detection. Recently, twostream convolutional neural networks (CNNs) trained on stacked optical flow and image frames have been successful for action recognition in videos. Our system uses a tracking algorithm to locate a bounding box around the person, which provides a frame of reference for appearance and motion and also suppresses background noise that is not within the bounding box. We train two additional streams on motion and appearance cropped to the tracked bounding box, along with full-frame streams. Our motion streams use pixel trajectories of a frame as raw features, in which the displacement values corresponding to a moving scene point are at the same spatial position across several frames. To model long-term temporal dynamics within and between actions, the multi-stream CNN is followed by a bi-directional Long Short-Term Memory (LSTM) layer. We show that our bi-directional LSTM network utilizes about 8 seconds of the video sequence to predict an action label. We test on two action detection datasets: the MPII Cooking 2 Dataset, and a new MERL Shopping Dataset that we introduce and make available to the community with this paper. The results demonstrate that our method significantly outperforms state-of-the-art action detection methods on both datasets.
Tim K. Marks, Michael J. Jones 0001, Oncel Tuzel, Ming Shao
CVPR5
2016 Deep Robust Encoder Through Locality Preserving Low-Rank Dictionary
Zhengming Ding, Ming Shao, Yun Fu 0001
ECCV (6)2
2016 Structure-Preserved Multi-source Domain Adaptation
abstract
Domain adaptation has achieved promising results in many areas, such as image classification and object recognition. Although a lot of algorithms have been proposed to solve the task with different domain distributions, it remains a challenge for multi-source unsupervised domain adaptation. In addition, most of the existing algorithms learn a classifier on the source domain and predict the labels for the target data, which indicates that only the knowledge derived from the hyperplane is transferred to the target domain and the structure information is ignored. In light of this, we propose a novel algorithm for multi-source unsupervised domain adaptation. Generally speaking, we aim to preserve the whole structure from source domains and transfer it to serve the task on the target domain. The source and target data are put together for clustering, which simultaneously explores the structures of the source and target domains. The structure-preserved information from source domain further guides the clustering process on the target domain. Extensive experiments on two widely used databases on object recognition and face identification show the substantial improvement of our proposed approach over several state-of-the-art methods. Especially, our algorithm can take use of multi-source domains and achieve robust and better performance compared with the single source domain adaptation methods.
Hongfu Liu 0001, Ming Shao, Yun Fu 0001
ICDM2
2016 Transfer learning for image classification with incomplete multiple sources
abstract
Transfer learning plays a powerful role in mitigating the discrepancy between test data (target) and auxiliary data (source). There is often the case that multiple sources are available in transfer learning. However, naively combining multiple sources does not lead to valid results, since they will introduce negative transfer as well. Furthermore, each single source from multiple sources may not cover all the labels of the target data. In this paper, we consider the problem that how to better utilize multiple incomplete sources for effective knowledge transfer. To this end, we propose a Bi-directional Low-Rank Transfer learning framework (BLRT). First, we adapt the conventional low-rank transfer learning to multiple sources knowledge transfer scenario. Second, an iterative structure learning is proposed to better use prior knowledge for transfer learning coefficients matrix. Third, a cross-source regularizer is added to couple the same labels from multiple incomplete sources, so that they could jointly compensate missing data from other sources. Experimental results on three groups of databases including face and object images have demonstrated that our method can successfully inherit knowledge from incomplete multiple sources and adapt to the target data successfully.
Zhengming Ding, Ming Shao, Yun Fu 0001
IJCNN2
2016 Sparse alignment for video analysis in discriminant Tensor space
abstract
RGB-D action streams have aroused impressive attentions for recognition task, for its geometric characteristic and less influence of illumination. However, there exists large divergences of intra-class actions performed between sub-action, multi-subject and multi-modality, which may affect the result of action recognition. In order to solve these three problems, we propose a Sparse alignment guided Non-negative Tensor Factorization (SaNTF) framework with non-negative tensor factorization (NTF) for subspace learning. SaNTF selects the key frames from two intra-class action sequences by sparse regression, and aligns them to mitigate the diversity in the new tensor subspace. In this paper, the high-dimensional RGB-D action sequence is represented as a third-order tensor to preserve the original spatiotemporal structure, and we employ NTF to find a common tensor subspace for realistic action recognition. The experiments on MSRDailyActivity3D action and MSRPair3D action datasets show the higher accuracy compared with the state-of-the-art temporal alignment methods.
Chengcheng Jia, Ming Shao, Yun Fu 0001
IJCNN2
2016 Infinite Ensemble for Image Clustering
abstract
Image clustering has been a critical preprocessing step for vision tasks, e.g., visual concept discovery, content-based image retrieval. Conventional image clustering methods use handcraft visual descriptors as basic features via K-means, or build the graph within spectral clustering. Recently, representation learning with deep structure shows appealing performance in unsupervised feature pre-treatment. However, few studies have discussed how to deploy deep representation learning to image clustering problems, especially the unified framework which integrates both representation learning and ensemble clustering for efficient image clustering still remains void. In addition, even though it is widely recognized that with the increasing number of basic partitions, ensemble clustering gets better performance and lower variances, the best number of basic partitions for a given data set is a pending problem. In light of this, we propose the Infinite Ensemble Clustering (IEC), which incorporates the power of deep representation and ensemble clustering in a one-step framework to fuse infinite basic partitions. Generally speaking, a set of basic partitions is firstly generated from the image data, then by converting the basic partitions to the 1-of-K codings, we link the marginalized auto-encoder to the infinite ensemble clustering with i.i.d. basic partitions, which can be approached by the closed-form solutions, finally we follow the layer-wise training procedure and feed the concatenated deep features to K-means for final clustering. Extensive experiments on diverse vision data sets with different levels of visual descriptors demonstrate both the time efficiency and superior performance of IEC compared to the state-of-the-art ensemble clustering and deep clustering methods.
Hongfu Liu 0001, Ming Shao, Sheng Li 0001, Yun Fu 0001
KDD2
2016 Families in the Wild (FIW): Large-Scale Kinship Image Database and Benchmarks
abstract
We present the largest kinship recognition dataset to date, Families in the Wild (FIW). Motivated by the lack of a single, unified dataset for kinship recognition, we aim to provide a dataset that captivates the interest of the research community. With only a small team, we were able to collect, organize, and label over 10,000 family photos of 1,000 families with our annotation tool designed to mark complex hierarchical relationships and local label information in a quick and efficient manner. We include several benchmarks for two image-based tasks, kinship verification and family recognition. For this, we incorporate several visual features and metric learning methods as baselines. Also, we demonstrate that a pre-trained Convolutional Neural Network (CNN) as an off-the-shelf feature extractor outperforms the other feature types. Then, results were further boosted by fine-tuning two deep CNNs on FIW data: (1) for kinship verification, a triplet loss function was learned on top of the network of pre-train weights; (2) for family recognition, a family-specific softmax classifier was added to the network.
Joseph P. Robinson, Ming Shao, Yue Wu 0008, Yun Fu 0001
ACM Multimedia2
2016 Scalable Nearest Neighbor Sparse Graph Approximation by Exploiting Graph Structure
abstract
We consider exploiting graph structure for sparse graph approximations. Graphs play increasingly important roles in learning problems: manifold learning, kernel learning, and spectral clustering. Specifically, in this paper, we concentrate on nearest neighbor sparse graphs which are widely adopted in learning problems due to its spatial efficiency. Nonetheless, we raise an even challenging problem: can we save more memory space while keep competitive performance for the sparse graph? To this end, first, we propose to partition the entire graph into intra- and inter-graphs by exploring the graph structure, and use both of them for graph approximation. Therefore, neighborhood similarities within each cluster, and between different clusters are well preserved. Second, we improve the space use of the entire approximation algorithm. Specially, a novel sparse inter-graph approximation algorithm is proposed, and corresponding approximation error bound is provided. Third, extensive experiments are conducted on 11 real-world datasets, ranging from small to large-scales, demonstrating that when using less space, our approximate graph can provide comparable or even better performance, in terms of approximation error, and clustering accuracy. In large-scale test, we use less than 1/100 memory of comparable algorithms, but achieve very appealing results.
Ming Shao, Xindong Wu 0001, Yun Fu 0001
IEEE Trans. Big Data1
2015 Part-Level Regularized Semi-Nonnegative Coding for Semi-Supervised Learning
abstract
Graph-based semi-supervised learning method has been influential in the data mining and machine learning fields. The key is to construct an effective graph to capture the intrinsic data structure, which further benefits for propagating the unlabeled data over the graph. The existing methods have shown the effectiveness of a graph regularization term on measuring the similarities among samples, which further uncovers the data structure. However, all the existing graph-based methods are on the sample-level, i.e. calculate the similarity based on sample-level representation coefficients, inevitably overlooking the underlying part-level structure within sample. Inspired by the strong interpretability of Non-negative Matrix Factorization (NMF) method, we design a more robust and discriminative graph, by integrating low-rank factorization and graph regularizer into a unified framework. Specifically, a novel low-rank factorization through Semi-Non-negative Matrix Factorization (SNMF) is proposed to extract the semantically part-level representation. Moreover, instead of incorporating a graph regularization on sample-level, we propose a sparse graph regularization term built on the decomposed part-level representation. This practice results in a more accurate measurement among samples, generating a more discriminative graph for semi-supervised learning. As a non-trivial contribution, we also provide an optimization solution to the proposed method. Comprehensive experimental evaluations show that our proposed method is able to achieve superior performance compared with the state-of-the-art semi-supervised classification baselines in both transductive and inductive scenarios.
Handong Zhao, Zhengming Ding, Ming Shao, Yun Fu 0001
ICDM3
2015 Deep Low-Rank Coding for Transfer Learning
Zhengming Ding, Ming Shao, Yun Fu 0001
IJCAI2
2015 Cross-View Projective Dictionary Learning for Person Re-Identification
Sheng Li 0001, Ming Shao, Yun Fu 0001
IJCAI2
2015 Deep Linear Coding for Fast Graph Clustering
Ming Shao, Sheng Li 0001, Zhengming Ding, Yun Fu 0001
IJCAI1
2015 Multi-View Low-Rank Analysis for Outlier Detection
abstract
Outlier detection is a fundamental problem in data mining. Unlike most existing methods that are designed for single-view data, we propose a multi-view outlier detection approach in this paper. Multi-view data can provide plentiful information of samples, however, detecting outliers from multi-view data is still a challenging problem due to the complicated distribution and inconsistent behavior of samples across different views. We address this problem through robust data representation, by building a Multi-view Low-Rank Analysis (MLRA) framework. Our framework contains two major components. First, it performs cross-view low-rank analysis for revealing the intrinsic structures of data. Second, it identifies outliers by estimating the outlier score for each test sample. Specifically, we formulate the cross-view low-rank analysis as a constrained rank-minimization problem, and present an efficient optimization algorithm to solve it. Different from the existing multi-view outlier detection methods, our framework is able to detect two different types of outliers from multiple views simultaneously. To this end, we design a criterion to estimate the outlier scores by analyzing the obtained representation coefficients. Experimental results on seven UCI datasets and the USPS-MNIST dataset demonstrate that our approach outperforms several state-of-the-art single-view and multi-view outlier detection methods in most cases.
Sheng Li 0001, Ming Shao, Yun Fu 0001
SDM2
2015 Missing Modality Transfer Learning via Latent Low-Rank Constraint
abstract
Transfer learning is usually exploited to leverage previously well-learned source domain for evaluating the unknown target domain; however, it may fail if no target data are available in the training stage. This problem arises when the data are multi-modal. For example, the target domain is in one modality, while the source domain is in another. To overcome this, we first borrow an auxiliary database with complete modalities, then consider knowledge transfer across databases and across modalities within databases simultaneously in a unified framework. The contributions are threefold: 1) a latent factor is introduced to uncover the underlying structure of the missing modality from the known data; 2) transfer learning in two directions allows the data alignment between both modalities and databases, giving rise to a very promising recovery; and 3) an efficient solution with theoretical guarantees to the proposed latent low-rank transfer learning algorithm. Comprehensive experiments on multi-modal knowledge transfer with missing target modality verify that our method can successfully inherit knowledge from both auxiliary database and source modality, and therefore significantly improve the recognition performance even when test modality is inaccessible in the training stage.
Zhengming Ding, Ming Shao, Yun Fu 0001
IEEE Trans. Image Process.2
2014 Latent Low-Rank Transfer Subspace Learning for Missing Modality Recognition
abstract
We consider an interesting problem in this paper that uses transfer learning in two directions to compensate missing knowledge from the target domain. Transfer learning tends to be exploited as a powerful tool that mitigates the discrepancy between different databases used for knowledge transfer. It can also be used for knowledge transfer between different modalities within one database. However, in either case, transfer learning will fail if the target data are missing. To overcome this, we consider knowledge transfer between different databases and modalities simultaneously in a single framework, where missing target data from one database are recovered to facilitate recognition task. We referred to this framework as Latent Low-rank Transfer Subspace Learning method (L2TSL). We first propose to use a low-rank constraint as well as dictionary learning in a learned subspace to guide the knowledge transfer between and within different databases. We then introduce a latent factor to uncover the underlying structure of the missing target data. Next, transfer learning in two directions is proposed to integrate auxiliary database for transfer learning with missing target data. Experimental results of multi-modalities knowledge transfer with missing target data demonstrate that our method can successfully inherit knowledge from the auxiliary database to complete the target domain, and therefore enhance the performance when recognizing data from the modality without any training data.
Zhengming Ding, Ming Shao, Yun Fu 0001
AAAI2
2014 Guided Fast Local Search for speeding up a financial forecasting algorithm
abstract
Guided Local Search is a powerful meta-heuristic algorithm that has been applied to a successful Genetic Programming Financial Forecasting tool called EDDIE. Although previous research has shown that it has significantly improved the performance of EDDIE, it also increased its computational cost to a high extent. This paper presents an attempt to deal with this issue by combining Guided Local Search with Fast Local Search, an algorithm that has shown in the past to be able to significantly reduce the computational cost of Guided Local Search. Results show that EDDIE's computational cost has been reduced by an impressive 77%, while at the same time there is no cost to the predictive performance of the algorithm.
Ming Shao, Dafni Smonou, Michael Kampouridis, Edward P. K. Tsang
CIFEr1
2014 Learning relative features through adaptive pooling for image classification
abstract
Bag-of-Feature (BoF) representations and spatial constraints have been popular in image classification research. One of the most successful methods uses sparse coding and spatial pooling to build discriminative features. However, minimizing the reconstruction error by sparse coding only considers the similarity between the input and codebooks. In contrast, this paper describes a novel feature learning approach for image classification by considering the dissimilarity between inputs and prototype images, or what we called reference basis (RB). First, we learn the feature representation by max-margin criterion between the input and the RB. The learned hyperplane is stored as the relative feature. Second, we propose an adaptive pooling technique to assemble multiple relative features generated by different RBs under the SVM framework, where the classifier and the pooling weights are jointly learned. Experiments based on three challenging datasets: Caltech-101, Scene 15 and Willow-Actions, demonstrate the effectiveness and generality of our framework.
Ming Shao, Sheng Li 0001, Tongliang Liu, Dacheng Tao, Thomas S. Huang, Yun Fu 0001
ICME1
2014 Locality linear fitting one-class SVM with low-rank constraints for outlier detection
abstract
We propose a novel outlier detection approach in this paper, which learns the most accurate hyperspheres for the normal data through a top-down procedure. Conventional one-class support vector machine (SVM) based approaches aim to find nonlinear global solutions for all the normal data, with the benefit of kernel trick. However, those methods are intractable when data are in large-scale and inaccurate when data are under complex distributions. It's observed that high dimensional data, e.g., features of texts or images, are always sparse, and linear classifier usually performs well. A specific class of data seldom lie in one single subspace. In this paper, we propose to learn multiple discriminative hyper-spheres locally based on the data distributions, and fit them globally to formulate a more discriminative boundary for the normal data. By far, neural mechanisms used by human brain-mind for outlier detection are not known, however, the top-down strategy proposed in this paper would inspire understanding of the human neural mechanisms. The benefits of our model are two-folds. First, the distribution of each local cluster is much simpler than that in a global view, which makes the fitting processing for each individual cluster much easier and insensitive to the choice of kernel. In particular, we adopt low-rank constraints to find multiple clusters automatically. Secondly, the proposed approach trains the model linearly which tackles the large-scale problem, substantially reducing training time and memory space. Extensive experimental results on three image databases demonstrate that our approach outperforms several related methods.
Sheng Li 0001, Ming Shao, Yun Fu 0001
IJCNN2
2014 Attractive or Not?: Beauty Prediction with Attractiveness-Aware Encoders and Robust Late Fusion
abstract
Facial attractiveness is an ever-lasting issue in art and social science. It also draws considerable attention from multimedia community recently. In this paper, we develop a framework highlighting attractiveness-aware feature extracted from a pair of auto-encoders to learn human-like assessment of facial beauty. Our work is fully-automatic that does not require any landmark and puts no restrictions on the faces' pose, expressions, and lighting conditions and therefore is applicable on a larger and more diverse dataset. To this end, first, a pair of auto-encoders is built respectively with beauty images and non-beauty images, which can be used to extract attractiveness-aware features by putting test images into both encoders. Second, we further enhance the performance using an efficient robust low-rank fusion framework to integrate the predicted confidence scores which are obtained based on certain kinds of features. We show that our attractiveness-aware model with multiple layers of auto-encoders produces appealing results and performs better than previous appearance-based approaches.
Ming Shao, Yun Fu 0001
ACM Multimedia2
2014 Generalized Transfer Subspace Learning Through Low-Rank Constraint
Ming Shao, Dmitry Kit, Yun Fu 0001
Int. J. Comput. Vis.1
2013 What Do You Do? Occupation Recognition in a Photo via Social Context
abstract
In this paper, we investigate the problem of recognizing occupations of multiple people with arbitrary poses in a photo. Previous work utilizing single person's nearly frontal clothing information and fore/background context preliminarily proves that occupation recognition is computationally feasible in computer vision. However, in practice, multiple people with arbitrary poses are common in a photo, and recognizing their occupations is even more challenging. We argue that with appropriately built visual attributes, co-occurrence, and spatial configuration model that is learned through structure SVM, we can recognize multiple people's occupations in a photo simultaneously. To evaluate our method's performance, we conduct extensive experiments on a new well-labeled occupation database with 14 representative occupations and over 7K images. Results on this database validate our method's effectiveness and show that occupation recognition is solvable in a more general case.
Ming Shao, Liangyue Li, Yun Fu 0001
ICCV1
2013 Random Faces Guided Sparse Many-to-One Encoder for Pose-Invariant Face Recognition
abstract
One of the most challenging task in face recognition is to identify people with varied poses. Namely, the test faces have significantly different poses compared with the registered faces. In this paper, we propose a high-level feature learning scheme to extract pose-invariant identity feature for face recognition. First, we build a single-hidden-layer neural network with sparse constraint, to extract pose-invariant feature in a supervised fashion. Second, we further enhance the discriminative capability of the proposed feature by using multiple random faces as the target values for multiple encoders. By enforcing the target values to be unique for input faces over different poses, the learned high-level feature that is represented by the neurons in the hidden layer is pose free and only relevant to the identity information. Finally, we conduct face identification on CMU Multi-PIE, and verification on Labeled Faces in the Wild (LFW) databases, where identification rank-1 accuracy and face verification accuracy with ROC curve are reported. These experiments demonstrate that our model is superior to other state-of-the-art approaches on handling pose variations.
Yizhe Zhang 0001, Ming Shao, Edward K. Wong, Yun Fu 0001
ICCV2
2012 Low-Rank Transfer Subspace Learning
abstract
One of the most important challenges in machine learning is performing effective learning when there are limited training data available. However, there is an important case when there are sufficient training data coming from other domains (source). Transfer learning aims at finding ways to transfer knowledge learned from a source domain to a target domain by handling the subtle differences between the source and target. In this paper, we propose a novel framework to solve the aforementioned knowledge transfer problem via low-rank representation constraints. This is achieved by finding an optimal subspace where each datum in the target domain can be linearly represented by the corresponding subspace in the source domain. Extensive experiments on several databases, i.e., Yale B, CMU PIE, UB Kin Face databases validate the effectiveness of the proposed approach and show the superiority to the existing, well-established methods.
Ming Shao, Carlos Castillo 0002, Zhenghong Gu, Yun Fu 0001
ICDM1
2012 Discriminative metric: Schatten norm vs. vector norm
Zhenghong Gu, Ming Shao, Liangyue Li, Yun Fu 0001
ICPR2
2012 Toward kinship verification using visual attributes
Si-Yu Xia, Ming Shao, Yun Fu 0001
ICPR2
2012 Understanding Kin Relationships in a Photo
abstract
There is an urgent need to organize and manage images of people automatically due to the recent explosion of such data on the Web in general and in social media in particular. Beyond face detection and face recognition, which have been extensively studied over the past decade, perhaps the most interesting aspect related to human-centered images is the relationship of people in the image. In this work, we focus on a novel solution to the latter problem, in particular the kin relationships. To this end, we constructed two databases: the first one named UB KinFace Ver2.0, which consists of images of children, their young parents and old parents, and the second one named FamilyFace. Next, we develop a transfer subspace learning based algorithm in order to reduce the significant differences in the appearance distributions between children and old parents facial images. Moreover, by exploring the semantic relevance of the associated metadata, we propose an algorithm to predict the most likely kin relationships embedded in an image. In addition, human subjects are used in a baseline study on both databases. Experimental results have shown that the proposed algorithms can effectively annotate the kin relationships among people in an image and semantic context can further improve the accuracy.
Si-Yu Xia, Ming Shao, Jiebo Luo 0001, Yun Fu 0001
IEEE Trans. Multim.2
2011 Kinship Verification through Transfer Learning
Si-Yu Xia, Ming Shao, Yun Fu 0001
IJCAI2
2010 A BEMD based normalization method for face recognition under variable illuminations
abstract
Face recognition remains challenging in computer vision due to variations on face, especially for illuminations. In this paper, a novel face illumination normalization method is proposed. By using Bidimensional Empirical Mode Decomposition (BEMD), a series of normalization images (BIMF) from one subject can be extracted with different spatial scales, each of which possesses a high recognition rate compared with former representative methods, i.e., SQI, LOG-DCT and LTV. What's more, canonical correlation analysis (CCA) is adopted in this paper to combine images generated from one input to form more discrimative features. Experiments on Yale B, Extended Yale B and CMU PIE show that the proposed method, though simple, is very effective when dealing with face recognition under variable lighting conditions.
Ming Shao, Yunhong Wang 0001, Xue Ling
ICASSP1
2009 Face Relighting Based on Multi-spectral Quotient Image and Illumination Tensorfaces
Ming Shao, Yunhong Wang 0001, Peijiang Liu
ACCV (3)1
2009 Recovering Facial Intrinsic Images from a Single Input
Ming Shao, Yunhong Wang 0001
ICIC (1)1
2009 Joint Features for Face Recognition under Variable Illuminations
abstract
In this paper, we propose a new method using joint features extracted from four efficient face illumination normalization approaches to deal with the face recognition problems under variable lighting conditions. These four methods (Logarithm Total Variation, Generic Intrinsic Illumination Subspace, Self-Quotient Image and Discrete Cosine Transform in Logarithm Domain) can indeed improve recognition rates solely when testing on face database, i.e. Yale B, Extended Yale B and CMU PIE. However, in this paper, we argue that single feature extracted from one method is useful but not adequate to high-accuracy face recognition system. Joint features generated by canonical correlation analysis (CCA) from more than one method can enhance the performance of existing algorithms. It is also suggested that CCA can project different features to the direction that maximize the correlation between them thus leading to an optimized joint feature. Experiments show that our method is not only simple but also effective on promoting face recognition rates.
Ming Shao, Yunhong Wang 0001
ICIG1
2009 A super-resolution based method to synthesize visual images from near infrared
abstract
In this paper, we propose a new method to enhance the quality of near infrared face image using tensorface, super-resolution and image fusion. Given a single model of near infrared face image which is not suitable for human to recognize or verify and its low-resolution sample, we can synthesize an image under visible light environment by building multiple factors training tensors and super-resolving its high-resolution visible light reconstructions across different modalities. The training tensor space consists of near infrared and visible light face images pairs of different people. Fusion is performed between the reconstructions and the original near infrared images. Experiments show promising results of synthesized visible light face images.
Ming Shao, Yunhong Wang 0001
ICIP1
2008 An adaptive counter propagation network based on soft competition
Yi-hong Dong, Ming Shao, Xiaoying Tai
Pattern Recognit. Lett.2
2005 Formal Verification Techniques Based on Boolean Satisfiability Problem
Xiaowei Li 0001, Guanghui Li 0001, Ming Shao
J. Comput. Sci. Technol.3
2003 Design Error Diagnosis Based on Verification Techniques
abstract
Error diagnosis is becoming more difficult in VLSI circuit designs due to the increasing complexity. In this paper, we present an algorithm based on verification for improving the accuracy of design error diagnosis. This algorithm integrates three-valued logic simulation and Boolean satisfiability (SAT). It uses test patterns generated by a gate level stuck-at fault ATPG tool for parallel pattern simulation, and uses SAT-based Boolean comparison to enhance the three-valued simulation, in which universally quantified conjunction normal formulas (CNF) represent the unknown constraints in the implementation with black boxes, and does not need circuit structural transformation. Our approach can quickly and efficiently eliminate many false candidates, experimental results on ISCAS'85 circuits show the accuracy and the speed of this approach.
Guanghui Li 0001, Ming Shao, Xiaowei Li 0001
Asian Test Symposium2
2003 SAT-Based Algorithm of Verification for Port Order Fault
abstract
In verification of embedded core-based design, the port order fault (POF) model focuses on the errors in connections between the ports of the cores and the surrounding circuits, thus considerably reduces the verification complexity and time. This paper investigated the automatic verification pattern generation for POF and developed an effective algorithm of verification for POF using SAT instead of BDD. The problem of detecting POF was transformed into SAT, which was efficiently solved by a state-of-the-art efficient SAT solver.
Ming Shao, Guanghui Li 0001, Xiaowei Li 0001
Asian Test Symposium1