Zhongfei Zhang

dblp:z/ZhongfeiMarkZhang · also Zhongfei (Mark) Zhang · DBLP profile ↗
← Back
231ranked-venue papers
17as first author
44since 2021 · last 2026
0000-0001-5098-2506ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 141 · 10 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 98 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 59 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 first-author · 1 since 2021Computer networks · 3 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Emphasizing Domain Differences Through Interactive-Augmented Prompts in Continual Audio-Visual Speech Recognition
abstract
Audio-Visual Speech Recognition (AVSR) has been studied for a long time in the literature. By leveraging the complementary information from both acoustic and visual modalities, this approach offers a promising solution for robust speech transcription. While recent AVSR models have achieved impressive performance on large-scale, uniformly distributed datasets, they often overlook the challenges posed by real-world scenarios-where data is collected across multiple sessions and environments, leading to significant domain shifts and heterogeneous distributions. Such heterogeneity can result in catastrophic forgetting and hinder the generalization ability of the conventional models. To bridge this gap, we introduce the Continual Audio-Visual Speech Recognition (CL-AVSR) problem, which formulates AVSR as a continual learning task. We establish a dedicated benchmark for CL-AVSR by designing three experimental scenarios that reflect real-world challenges: introducing varying background noise for the audio stream, degrading video quality for the visual stream, and dividing tasks by speaker characteristics to jointly affect both modalities. These scenarios systematically evaluate the model's ability to adapt and retain knowledge across dynamic and non-stationary data streams. To address the unique challenges of CL-AVSR, we propose the Interaction-enhanced Multimodal Prompt learning (IMP) framework. IMP builds upon a pre-trained AV-HuBERT backbone and integrates task-relevant soft prompts with cross-modal and cross-task interactions, enabling efficient knowledge transfer from high-quality source domains to typical low-quality target domains with minimal parameter overhead. The interactive prompts facilitate fine-grained alignment and adaptation between modalities and tasks, while contrastive regularization further mitigates catastrophic forgetting. Furthermore, we devise a multi-modal prompt selection strategy that leverages clustering-based feature analysis, empowering the model to dynamically select optimal prompts for unseen data distributions during inference. Extensive experiments on the LRS2 dataset demonstrate that IMP achieves substantial improvements over strong baselines, setting new state-of-the-art performance in all CL-AVSR scenarios. Our results highlight the effectiveness of IMP in enhancing continual learning capabilities for AVSR, paving the way for more robust and adaptable multi-modal speech recognition systems in real-world applications.
Xize Cheng, Jingyuan Chen 0003, Tao Jin 0004, Zhongfei Zhang
IEEE Trans. Image Process.5
2026 PRISM: Link Prediction in Attributed Networks With Uncertain Modalities
abstract
Link prediction for attributed graphs has garnered significant attention due to its ability to enhance predictive performance by leveraging multi-modal node attributes. However, real-world challenges such as privacy concerns, content restrictions, and attribute constraints often result in nodes facing varying degrees of missing modalities in their attributes, significantly limiting the effectiveness of existing approaches. Building on this fact, we propose a model for linkPRediction in attrIbuted networkSwith uncertainModalities (PRISM), which learns the shared representations across various scenarios of missing modalities through dual-level adversarial training.PRISMcomprises four modules,i.e.,a GCN extractor, an adversarial extractor, an attentive fusion, and an adaptive aggregator. The GCN extractor leverages graph convolutional networks (GCN) to extract fundamental representations from the network topology. The adversarial extractor employs dual-level adversarial training to acquire the shared representations across various multi-modal scenarios at the node-level and link-level, respectively. The attentive fusion applies the multi-head attention mechanism to integrate the shared representations and the fundamental representations. The adaptive aggregator comprehensively considers both node-level and link-level representations to predict the existence of links. Experimental evaluation using real-world datasets demonstrates thatPRISMsignificantly outperforms existing state-of-the-art link prediction methods for multi-modal attributed graphs under missing modalities by improving the Recall@50 metric (R@50) by up to 38.79%.
Muhammad Asif Ali, Huan Wang 0005, Zhongfei Zhang, Junyang Chen 0001, Di Wang 0015
IEEE Trans. Knowl. Data Eng.4
2025 Cloud-fog-edge based computing architechture and a hierarchical decision approach for distributed synchronized manufacturing systems
Zhicong Hong, Ting Qu 0002, Yongheng Zhang 0004, Zhongfei Zhang, George Q. Huang
Adv. Eng. Informatics4
2025 Recognize-and-tell: Generating video captions with textual cue in scene
Tao Jin 0004, Hao Jiang 0062, Jian Wang 0119, Jingyuan Chen 0003, Zhou Zhao 0001, Zhongfei Zhang
Expert Syst. Appl.9
2025 Balancing Feature Alignment and Uniformity for Few-Shot Classification
abstract
In Few-Shot Learning (FSL), the objective is to correctly recognize new samples from novel classes with only a few available samples per class. Existing methods in FSL primarily focus on learning transferable knowledge from base classes by maximizing the information between feature representations and their corresponding labels. However, this approach may suffer from the "supervision collapse" issue, which arises due to a bias towards the base classes. In this paper, we propose a solution to address this issue by preserving the intrinsic structure of the data and enabling the learning of a generalized model for the novel classes. Following the InfoMax principle, our approach maximizes two types of mutual information (MI): between the samples and their feature representations, and between the feature representations and their class labels. This allows us to strike a balance between discrimination (capturing class-specific information) and generalization (capturing common characteristics across different classes) in the feature representations. To achieve this, we adopt a unified framework that perturbs the feature embedding space using two low-bias estimators. The first estimator maximizes the MI between a pair of intra-class samples, while the second estimator maximizes the MI between a sample and its augmented views. This framework effectively combines knowledge distillation between class-wise pairs and enlarges the diversity in feature representations. By conducting extensive experiments on popular FSL benchmarks, our proposed approach achieves comparable performances with state-of-the-art competitors. For example, we achieved an accuracy of 69.53% on the miniImageNet dataset and 77.06% on the CIFAR-FS dataset for the 5-way 1-shot task.
Yunlong Yu 0001, Dingyi Zhang, Zhong Ji, Xi Li 0001, Jungong Han, Zhongfei Zhang
IEEE Trans. Image Process.6
2024 Enhancing trusted synchronization in open production logistics: A platform framework integrating blockchain and digital twin under social manufacturing
Zhongfei Zhang, Ting Qu 0002, Kuo Zhao, Yongheng Zhang 0004, Wenyou Guo
Adv. Eng. Informatics1
2024 Uncertainty-aware enhanced dark experience replay for continual learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang
Appl. Intell.4
2024 Dense affinity matching for Few-Shot Segmentation
Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Yingming Li, Jungong Han, Zhongfei Zhang
Neurocomputing7
2024 Dual-DIANet: A sharing-learnable multi-task network based on dense information aggregation
Kejie Lyu, Yingming Li, Zhongfei Zhang
Neurocomputing3
2024 Tolerant Self-Distillation for image classification
Mushui Liu, Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang
Neural Networks5
2024 Multi-Content Interaction Network for Few-Shot Segmentation
abstract
Few-Shot Segmentation (FSS) poses significant challenges due to limited support images and large intra-class appearance discrepancies. Most existing approaches focus on aligning the support-query correlations from the same layer of the frozen backbone while neglecting the bias between different tasks and different layers. In this article, we propose a Multi-Content Interaction Network (MCINet) to remedy these issues by fully exploiting and interacting with the different contextual information contained in distinct branches. Specifically, MCINet improves FSS from three perspectives: (1) boosting the query representations through incorporating the independent information from another learnable branch into the features from the frozen backbone, (2) enhancing the support-query correlations by exploiting both the same-layer and adjacent-layer features, and (3) refining the predicted results with a multi-scale mask prediction strategy. Experiments on three benchmarks demonstrate that our approach reaches state-of-the-art performances and outperforms the best competitors with many desirable advantages, especially on the challenging COCO dataset. Code will be released on GitHub ( https://github.com/chenhao-zju/mcinet ).
Hao Chen 0107, Yunlong Yu 0001, Yonghan Dong, Zheming Lu 0001, Yingming Li, Zhongfei Zhang
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Multi-Relational Contrastive Learning Graph Neural Network for Drug-Drug Interaction Event Prediction
abstract
Drug-drug interactions (DDIs) could lead to various unexpected adverse consequences, so-called DDI events. Predicting DDI events can reduce the potential risk of combinatorial therapy and improve the safety of medication use, and has attracted much attention in the deep learning community. Recently, graph neural network (GNN)-based models have aroused broad interest and achieved satisfactory results in the DDI event prediction. Most existing GNN-based models ignore either drug structural information or drug interactive information, but both aspects of information are important for DDI event prediction. Furthermore, accurately predicting rare DDI events is hindered by their inadequate labeled instances. In this paper, we propose a new method, Multi-Relational Contrastive learning Graph Neural Network, MRCGNN for brevity, to predict DDI events. Specifically, MRCGNN integrates the two aspects of information by deploying a GNN on the multi-relational DDI event graph attributed with the drug features extracted from drug molecular graphs. Moreover, we implement a multi-relational graph contrastive learning with a designed dual-view negative counterpart augmentation strategy, to capture implicit information about rare DDI events. Extensive experiments on two datasets show that MRCGNN outperforms the state-of-the-art methods. Besides, we observe that MRCGNN achieves satisfactory performance when predicting rare DDI events.
Zhankun Xiong, Shichao Liu 0002, Feng Huang 0004, Xuan Liu 0010, Zhongfei Zhang, Wen Zhang 0008
AAAI6
2023 Multi-Scale Similarity Aggregation for Dynamic Metric Learning
abstract
In this paper, we propose a new multi-scale similarity aggregation method (MSA) for dynamic metric learning (DyML), which adopts a pretraining-finetuning scheme and efficiently learns the similarity relationship for each semantic level. In particular, building upon the framework of self-supervised pretraining, the output embedding layer is divided into three learners to learn the similarity relations in each level individually. Then for training these learners, the hierarchical prior information is fully considered. Specifically, in light of the class hierarchy that each class in a coarse level corresponds to a set of subclasses in a finer level, multi-proxy learning is employed to facilitate the single-level similarity learning of each learner. On the other hand, following the hierarchical consistency property, a cross-level similarity constraint is further presented to encourage the estimated similarities of the three learners to be hierarchically consistent. Extensive experiments on three DyML datasets show that MSA significantly outperforms the existing state-of-the-art methods and allows for a better generalization for different semantic scales.
Dingyi Zhang, Yingming Li, Zhongfei Zhang
ACM Multimedia3
2023 HimGNN: a novel hierarchical molecular graph representation learning framework for property prediction
abstract
Accurate prediction of molecular properties is an important topic in drug discovery. Recent works have developed various representation schemes for molecular structures to capture different chemical information in molecules. The atom and motif can be viewed as hierarchical molecular structures that are widely used for learning molecular representations to predict chemical properties. Previous works have attempted to exploit both atom and motif to address the problem of information loss in single representation learning for various tasks. To further fuse such hierarchical information, the correspondence between learned chemical features from different molecular structures should be considered. Herein, we propose a novel framework for molecular property prediction, called hierarchical molecular graph neural networks (HimGNN). HimGNN learns hierarchical topology representations by applying graph neural networks on atom- and motif-based graphs. In order to boost the representational power of the motif feature, we design a Transformer-based local augmentation module to enrich motif features by introducing heterogeneous atom information in motif representation learning. Besides, we focus on the molecular hierarchical relationship and propose a simple yet effective rescaling module, called contextual self-rescaling, that adaptively recalibrates molecular representations by explicitly modelling interdependencies between atom and motif features. Extensive computational experiments demonstrate that HimGNN can achieve promising performances over state-of-the-art baselines on both classification and regression tasks in molecular property prediction.
Shen Han, Haitao Fu, Yuyang Wu, Ganglan Zhao, Feng Huang 0004, Zhongfei Zhang, Shichao Liu 0002, Wen Zhang 0008
Briefings Bioinform.7
2023 Zero-shot classification with unseen prototype learning
Zhong Ji, Biying Cui, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Neural Comput. Appl.5
2023 COREN: Multi-Modal Co-Occurrence Transformer Reasoning Network for Image-Text Retrieval
Zhong Ji, Yanwei Pang, Zhongfei Zhang
Neural Process. Lett.5
2023 Enhancing Drug-Drug Interaction Prediction Using Deep Attention Neural Networks
abstract
Drug-drug interactions are one of the main concerns in drug discovery. Accurate prediction of drug-drug interactions plays a key role in increasing the efficiency of drug research and safety when multiple drugs are co-prescribed. With various data sources that describe the relationships and properties between drugs, the comprehensive approach that integrates multiple data sources would be considerably effective in making high-accuracy prediction. In this paper, we propose a Deep Attention Neural Network based Drug-Drug Interaction prediction framework, abbreviated as DANN-DDI, to predict unobserved drug-drug interactions. First, we construct multiple drug feature networks and learn drug representations from these networks using the graph embedding method; then, we concatenate the learned drug embeddings and design an attention neural network to learn representations of drug-drug pairs; finally, we adopt a deep neural network to accurately predict drug-drug interactions. The experimental results demonstrate that our model DANN-DDI has improved prediction performance compared with state-of-the-art methods. Moreover, the proposed model can predict novel drug-drug interactions and drug-drug interaction-associated events.
Shichao Liu 0002, Yang Zhang 0123, Zhongfei Zhang, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 Complementary Calibration: Boosting General Continual Learning With Collaborative Distillation and Self-Supervision
abstract
General Continual Learning (GCL) aims at learning from non independent and identically distributed stream data without catastrophic forgetting of the old tasks that don't rely on task boundaries during both training and testing stages. We reveal that the relation and feature deviations are crucial problems for catastrophic forgetting, in which relation deviation refers to the deficiency of the relationship among all classes in knowledge distillation, and feature deviation refers to indiscriminative feature representations. To this end, we propose a Complementary Calibration (CoCa) framework by mining the complementary model's outputs and features to alleviate the two deviations in the process of GCL. Specifically, we propose a new collaborative distillation approach for addressing the relation deviation. It distills model's outputs by utilizing ensemble dark knowledge of new model's outputs and reserved outputs, which maintains the performance of old tasks as well as balancing the relationship among all classes. Furthermore, we explore a collaborative self-supervision idea to leverage pretext tasks and supervised contrastive learning for addressing the feature deviation problem by learning complete and discriminative features for all classes. Extensive experiments on six popular datasets show that our CoCa framework achieves superior performance against state-of-the-art methods. Code is available at https://github.com/lijincm/CoCa.
Zhong Ji, Jin Li 0054, Qiang Wang 0056, Zhongfei Zhang
IEEE Trans. Image Process.4
2023 Knowledge Distillation Classifier Generation Network for Zero-Shot Learning
abstract
In this article, we present a conceptually simple but effective framework called knowledge distillation classifier generation network (KDCGN) for zero-shot learning (ZSL), where the learning agent requires recognizing unseen classes that have no visual data for training. Different from the existing generative approaches that synthesize visual features for unseen classifiers' learning, the proposed framework directly generates classifiers for unseen classes conditioned on the corresponding class-level semantics. To ensure the generated classifiers to be discriminative to the visual features, we borrow the knowledge distillation idea to both supervise the classifier generation and distill the knowledge with, respectively, the visual classifiers and soft targets trained from a traditional classification network. Under this framework, we develop two, respectively, strategies, i.e., class augmentation and semantics guidance, to facilitate the supervision process from the perspectives of improving visual classifiers. Specifically, the class augmentation strategy incorporates some additional categories to train the visual classifiers, which regularizes the visual classifier weights to be compact, under supervision of which the generated classifiers will be more discriminative. The semantics-guidance strategy encodes the class semantics into the visual classifiers, which would facilitate the supervision process by minimizing the differences between the generated and the real-visual classifiers. To evaluate the effectiveness of the proposed framework, we have conducted extensive experiments on five datasets in image classification, i.e., AwA1, AwA2, CUB, FLO, and APY. Experimental results show that the proposed approach performs best in the traditional ZSL task and achieves a significant performance improvement on four out of the five datasets in the generalized ZSL task.
Yunlong Yu 0001, Bin Li 0038, Zhong Ji, Jungong Han, Zhongfei Zhang
IEEE Trans. Neural Networks Learn. Syst.5
2022 Reducing Flipping Errors in Deep Neural Networks
abstract
Deep neural networks (DNNs) have been widely applied in various domains in artificial intelligence including computer vision and natural language processing. A DNN is typically trained for many epochs and then a validation dataset is used to select the DNN in an epoch (we simply call this epoch ``the last epoch") as the final model for making predictions on unseen samples, while it usually cannot achieve a perfect accuracy on unseen samples. An interesting question is ``how many test (unseen) samples that a DNN misclassifies in the last epoch were ever correctly classified by the DNN before the last epoch?". In this paper, we empirically study this question and find on several benchmark datasets that the vast majority of the misclassified samples in the last epoch were ever classified correctly before the last epoch, which means that the predictions for these samples were flipped from ``correct" to ``wrong". Motivated by this observation, we propose to restrict the behavior changes of a DNN on the correctly-classified samples so that the correct local boundaries can be maintained and the flipping error on unseen samples can be largely reduced. Extensive experiments on different benchmark datasets with different modern network architectures demonstrate that the proposed flipping error reduction (FER) approach can substantially improve the generalization, the robustness, and the transferability of DNNs without introducing any additional network parameters or inference cost, only with a negligible training overhead.
Xiang Deng 0002, Bo Long, Zhongfei Zhang
AAAI4
2022 Deep Hybrid Models for Out-of-Distribution Detection
abstract
We propose a principled and practical method for out-of-distribution (OoD) detection with deep hybrid models (DHMs), which model the joint density p(x, y) of features and labels with a single forward pass. By factorizing the joint density p(x, y) into three sources of uncertainty, we show that our approach has the ability to identify samples semantically different from the training data. To ensure computational scalability, we add a weight normalization step during training, which enables us to plug in state-of-the-art (SoTA) deep neural network (DNN) architectures for approximately modeling and inferring expressive probability distributions. Our method provides an efficient, general, and flexible framework for predictive uncertainty estimation with promising results and theoretical support. To our knowledge, this is the first work to reach 100% in OoD detection tasks on both vision and language datasets, especially on notably difficult dataset pairs such as CIFAR -10 vs. SVHN and CIFAR-100 vs. CIFAR-10. This work is a step towards enabling DNNs in real-world deployment for safety-critical applications.
Senqi Cao, Zhongfei Zhang
CVPR2
2022 Personalized Education: Blind Knowledge Distillation
Xiang Deng 0002, Zhongfei Zhang
ECCV (34)3
2022 Video-Grounded Dialogues with Joint Video and Image Training
abstract
In this paper, we propose a multi-modal transformer model for end-to-end training of video-grounded dialogue generation. In particular, LayerScale regularized spatio-temporal self-attention blocks are first introduced to enable us to flexibly train end-to-end from both video and image data, without extracting offline visual features. Further, a pre-trained generative language architecture BART is employed to encode different modalities and perform dialogue generation. Extensive experiments on Audio-Visual Scene-Aware Dialog (AVSD) dataset demonstrate its effectiveness and superiority to the state-of-the-art methods.
Yingming Li, Zhongfei Zhang
ICIP3
2022 Deep Causal Metric Learning
abstract
Deep metric learning aims to learn distance metrics that measure similarities and dissimilarities between samples. The existing approaches typically focus on designing different hard sample mining or distance margin strategies and then minimize a pair/triplet-based or proxy-based loss over the training data. However, this can lead the model to recklessly learn all the correlated distances found in training data including the spurious distance (e.g., background differences) that is not the distance of interest and can harm the generalization of the learned metric. To address this issue, we study metric learning from a causality perspective and accordingly propose deep causal metric learning (DCML) that pursues the true causality of the distance between samples. DCML is achieved through explicitly learning environment-invariant attention and task-invariant embedding based on causal inference. Extensive experiments on several benchmark datasets demonstrate the superiority of DCML over the existing methods.
Xiang Deng 0002, Zhongfei Zhang
ICML2
2022 Multi-Proxy Learning from an Entropy Optimization Perspective
abstract
Deep Metric Learning, a task that learns a feature embedding space where semantically similar samples are located closer than dissimilar samples, is a cornerstone of many computer vision applications. Most of the existing proxy-based approaches usually exploit the global context via learning a single proxy for each training class, which struggles in capturing the complex non-uniform data distribution with different patterns. In this work, we present an easy-to-implement framework to effectively capture the local neighbor relationships via learning multiple proxies for each class that collectively approximate the intra-class distribution. In the context of large intra-class visual diversity, we revisit the entropy learning under the multi-proxy learning framework and provide a training routine that both minimizes the entropy of intra-class probability distribution and maximizes the entropy of inter-class probability distribution. In this way, our model is able to better capture the intra-class variations and smooth the inter-class differences and thus facilitates to extract more semantic feature representations for the downstream tasks. Extensive experimental results demonstrate that the proposed approach achieves competitive performances. Codes and an appendix are provided.
Yunlong Yu 0001, Dingyi Zhang, Yingming Li, Zhongfei Zhang
IJCAI4
2022 Training a Lightweight ViT Network for Image Retrieval
Yunlong Yu 0001, Yingming Li, Zhongfei Zhang
PRICAI (3)4
2022 META-DDIE: predicting drug-drug interaction events with few-shot learning
abstract
Drug-drug interactions (DDIs) are one of the major concerns in pharmaceutical research, and a number of computational methods have been developed to predict whether two drugs interact or not. Recently, more attention has been paid to events caused by the DDIs, which is more useful for investigating the mechanism hidden behind the combined drug usage or adverse reactions. However, some rare events may only have few examples, hindering them from being precisely predicted. To address the above issues, we present a few-shot computational method named META-DDIE, which consists of a representation module and a comparing module, to predict DDI events. We collect drug chemical structures and DDIs from DrugBank, and categorize DDI events into hundreds of types using a standard pipeline. META-DDIE uses the structures of drugs as input and learns the interpretable representations of DDIs through the representation module. Then, the model uses the comparing module to predict whether two representations are similar, and finally predicts DDI events with few labeled examples. In the computational experiments, META-DDIE outperforms several baseline methods and especially enhances the predictive capability for rare events. Moreover, META-DDIE helps to identify the key factors that may cause DDI events and reveal the relationship among different events.
Xinran Xu, Shichao Liu 0002, Zhongfei Zhang, Shanfeng Zhu, Wen Zhang 0008
Briefings Bioinform.5
2022 JSENet: A deep convolutional neural network for joint image super-resolution and enhancement
Kejie Lyu, Sicheng Pan, Yingming Li, Zhongfei Zhang
Neurocomputing4
2022 Local spatial alignment network for few-shot learning
Yunlong Yu 0001, Dingyi Zhang, Sidi Wang, Zhong Ji, Zhongfei Zhang
Neurocomputing5
2022 Hierarchical Correlations Replay for Continual Learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang
Knowl. Based Syst.5
2022 Sparsity-control ternary weight networks
Xiang Deng 0002, Zhongfei Zhang
Neural Networks2
2022 Efficient Relational Sentence Ordering Network
abstract
In this paper, we propose a novel deep Efficient Relational Sentence Ordering Network (referred to as ERSON) by leveraging pre-trained language model in both encoder and decoder architectures to strengthen the coherence modeling of the entire model. Specifically, we first introduce a divide-and-fuse BERT (referred to as DF-BERT), a new refactor of BERT network, where lower layers in the improved model encode each sentence in the paragraph independently, which are shared by different sentence pairs, and the higher layers learn the cross-attention between sentence pairs jointly. It enables us to capture the semantic concepts and contextual information between the sentences of the paragraph, while significantly reducing the runtime and memory consumption without sacrificing the model performance. Besides, a Relational Pointer Decoder (referred to as RPD) is developed, which utilizes the pre-trained Next Sentence Prediction (NSP) task of BERT to capture the useful relative ordering information between sentences to enhance the order predictions. In addition, a variety of knowledge distillation based losses are added as auxiliary supervision to further improve the ordering performance. The extensive evaluations on Sentence Ordering, Order Discrimination, and Multi-Document Summarization tasks show the superiority of ERSON to the state-of-the-art ordering methods.
Yingming Li, Baiyun Cui, Zhongfei Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Semantic-Guided Class-Imbalance Learning Model for Zero-Shot Image Classification
abstract
In this article, we focus on the task of zero-shot image classification (ZSIC) that equips a learning system with the ability to recognize visual images from unseen classes. In contrast to the traditional image classification, ZSIC more easily suffers from the class-imbalance issue since it is more concerned with the class-level knowledge transferring capability. In the real world, the sample numbers of different categories generally follow a long-tailed distribution, and the discriminative information in the sample-scarce seen classes is hard to transfer to the related unseen classes in the traditional batch-based training manner, which degrades the overall generalization ability a lot. To alleviate the class-imbalance issue in ZSIC, we propose a sample-balanced training process to encourage all training classes to contribute equally to the learned model. Specifically, we randomly select the same number of images from each class across all training classes to form a training batch to ensure that the sample-scarce classes contribute equally as those classes with sufficient samples during each iteration. Considering that the instances from the same class differ in class representativeness, we further develop an efficient semantic-guided feature fusion model to obtain the discriminative class visual prototype for the following visual-semantic interaction process via distributing different weights to the selected samples based on their class representativeness. Extensive experiments on three imbalanced ZSIC benchmark datasets for both traditional ZSIC and generalized ZSIC tasks demonstrate that our approach achieves promising results, especially for the unseen categories that are closely related to the sample-scarce seen categories. Besides, the experimental results on two class-balanced datasets show that the proposed approach also improves the classification performance against the baseline model.
Zhong Ji, Xuejie Yu, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
IEEE Trans. Cybern.5
2022 Task-Oriented High-Order Context Graph Networks for Few-Shot Human-Object Interaction Recognition
abstract
Few-shot human-object interaction (FS-HOI) recognition aims at inferring new interactions between human actions and surrounding objects merely with a few available instances. It is beneficial to alleviate the long-tail and combinatorial explosion problems in human-object interaction (HOI). Nevertheless, the existing FS-HOI methods only focus on modeling the relationships between labeled samples and unlabeled samples in the Euclidean domain, which neglects the rich relational structures of the visual information among labeled samples and between human actions and objects. Accordingly, we tackle the few-shot HOI task in the non-Euclidean domain and present a graph-based model, namely, task-oriented high-order context graph network (THCG-Net). It contains a task attention module (TA-Module) and a high-order context graph module (HG-Module). In TA-Module, an attention mechanism is designed by utilizing task information to build a task-oriented space, in which the discriminative information for the current task (episode) is captured by embedding the visual features into the task-oriented space. The HG-Module is proposed to construct a task-level graph and takes the context information as high-order knowledge, which provides discriminative guidance for propagating visual information. It captures the discriminability among different categories while highlights the commonality of related categories adaptively, which effectively transfers knowledge to related categories. Extensive experimental results on two benchmark datasets, HICO-FS and TUHOI-FS, are provided. It demonstrates that our THCG-Net significantly outperforms the state-of-the-art approaches, which proves its impressive effectiveness in recognizing various human actions and surrounding objects in few-shot scenarios.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Ling Shao 0001, Zhongfei Zhang
IEEE Trans. Syst. Man Cybern. Syst.6
2021 Learning with Retrospection
abstract
Deep neural networks have been successfully deployed in various domains of artificial intelligence, including computer vision and natural language processing. We observe that the current standard procedure for training DNNs discards all the learned information in the past epochs except the current learned weights. An interesting question is: is this discarded information indeed useless? We argue that the discarded information can benefit the subsequent training. In this paper, we propose learning with retrospection (LWR) which makes use of the learned information in the past epochs to guide the subsequent training. LWR is a simple yet effective training framework to improve accuracies, calibration, and robustness of DNNs without introducing any additional network parameters or inference cost, but only with a negligible training overhead. Extensive experiments on several benchmark datasets demonstrate the superiority of LWR for training DNNs.
Xiang Deng 0002, Zhongfei Zhang
AAAI2
2021 CoCosNet v2: Full-Resolution Correspondence Learning for Image Translation
abstract
We present the full-resolution correspondence learning for cross-domain images, which aids image translation. We adopt a hierarchical strategy that uses the correspondence from coarse level to guide the fine levels. At each hierarchy, the correspondence can be efficiently computed via PatchMatch that iteratively leverages the matchings from the neighborhood. Within each PatchMatch iteration, the ConvGRU module is employed to refine the current correspondence considering not only the matchings of larger context but also the historic estimates. The proposed Co-CosNet v2, a GRU-assisted PatchMatch approach, is fully differentiable and highly efficient. When jointly trained with image translation, full-resolution semantic correspondence can be established in an unsupervised manner, which in turn facilitates the exemplar-based image translation. Experiments on diverse translation tasks show that CoCosNet v2 performs considerably better than state-of-the-art literature on producing high-resolution images.
Xingran Zhou, Bo Zhang 0025, Ting Zhang 0002, Pan Zhang 0003, Jianmin Bao, Dong Chen 0003, Zhongfei Zhang, Fang Wen 0001
CVPR7
2021 Graph-Free Knowledge Distillation for Graph Neural Networks
abstract
Knowledge distillation (KD) transfers knowledge from a teacher network to a student by enforcing the student to mimic the outputs of the pretrained teacher on training data. However, data samples are not always accessible in many cases due to large data sizes, privacy, or confidentiality. Many efforts have been made on addressing this problem for convolutional neural networks (CNNs) whose inputs lie in a grid domain within a continuous space such as images and videos, but largely overlook graph neural networks (GNNs) that handle non-grid data with different topology structures within a discrete space. The inherent differences between their inputs make these CNN-based approaches not applicable to GNNs. In this paper, we propose to our best knowledge the first dedicated approach to distilling knowledge from a GNN without graph data. The proposed graph-free KD (GFKD) learns graph topology structures for knowledge transfer by modeling them with multinomial distribution. We then introduce a gradient estimator to optimize this framework. Essentially, the gradients w.r.t. graph structures are obtained by only using GNN forward-propagation without back-propagation, which means that GFKD is compatible with modern GNN libraries such as DGL and Geometric. Moreover, we provide the strategies for handling different types of prior knowledge in the graph data or the GNNs. Extensive experiments demonstrate that GFKD achieves the state-of-the-art performance for distilling knowledge from GNNs without training data.
Xiang Deng 0002, Zhongfei Zhang
IJCAI2
2021 Comprehensive Knowledge Distillation with Causal Intervention
abstract
Knowledge distillation (KD) addresses model compression by distilling knowledge from a large model (teacher) to a smaller one (student). The existing distillation approaches mainly focus on using different criteria to align the sample representations learned by the student and the teacher, while they fail to transfer the class representations. Good class representations can benefit the sample representation learning by shaping the sample representation distribution. On the other hand, the existing approaches enforce the student to fully imitate the teacher while ignoring the fact that the teacher is typically not perfect. Although the teacher has learned rich and powerful representations, it also contains unignorable bias knowledge which is usually induced by the context prior (e.g., background) in the training data. To address these two issues, in this paper, we propose comprehensive, interventional distillation (CID) that captures both sample and class representations from the teacher while removing the bias with causal intervention. Different from the existing literature that uses the softened logits of the teacher as the training targets, CID considers the softened logits as the context information of an image, which is further used to remove the biased knowledge based on causal inference. Keeping the good representations while removing the bad bias enables CID to have a better generalization ability on test data and a better transferability across different datasets against the existing state-of-the-art approaches, which is demonstrated by extensive experiments on several benchmark datasets.
Xiang Deng 0002, Zhongfei Zhang
NeurIPS2
2021 Joint structured pruning and dense knowledge distillation for efficient transformer model compression
Baiyun Cui, Yingming Li, Zhongfei Zhang
Neurocomputing3
2021 Reweighting and information-guidance networks for Few-Shot Learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Zhongfei Zhang
Neurocomputing4
2021 Coordinating Experience Replay: A Harmonious Experience Retention approach for Continual Learning
Zhong Ji, Qiang Wang 0056, Zhongfei Zhang
Knowl. Based Syst.4
2021 Efficient Person Search via Expert-Guided Knowledge Distillation
abstract
The person search problem aims to find the target person in the scene images, which presents high demands for both effectiveness and efficiency. In this paper, we present a unified person search framework which jointly handles the two demands for real-world applications. We explore the technique of knowledge distillation (KD), which allows the student network to share capabilities of the deep expert networks with much fewer parameters and less computing time. To achieve this, we describe an efficient person search network and a set of deep and well-engineered expert networks, to build a tiny and compact model that can approximate the representations of the expert networks in a multitask learning manner. We present extensive experiments on three customized student networks with different scales of networks and show strong performance compared to the state-of-the-art methods on both mean average precision and top-1 accuracies. We further demonstrate the efficiency of the proposed network at 120 frames/s in the feedforward time with only a little sacrifice on the accuracy.
Xi Li 0001, Zhongfei Zhang
IEEE Trans. Cybern.3
2021 Multitask Non-Autoregressive Model for Human Motion Prediction
abstract
Human motion prediction, which aims at predicting future human skeletons given the past ones, is a typical sequence-to-sequence problem. Therefore, extensive efforts have been devoted to exploring different RNN-based encoder-decoder architectures. However, by generating target poses conditioned on the previously generated ones, these models are prone to bringing issues such as error accumulation problem. In this paper, we argue that such issue is mainly caused by adopting autoregressive manner. Hence, a novel Non-AuToregressive model (NAT) is proposed with a complete non-autoregressive decoding scheme, as well as a context encoder and a positional encoding module. More specifically, the context encoder embeds the given poses from temporal and spatial perspectives. The frame decoder is responsible for predicting each future pose independently. The positional encoding module injects positional signal into the model to indicate the temporal order. Besides, a multitask training paradigm is presented for both low-level human skeleton prediction and high-level human action recognition, resulting in the considerable improvement for the prediction task. Our approach is evaluated on Human3.6M and CMU-Mocap benchmarks and outperforms state-of-the-art autoregressive methods.
Bin Li 0038, Zhongfei Zhang, Hailin Feng, Xi Li 0001
IEEE Trans. Image Process.3
2021 Local-Global Memory Neural Network for Medication Prediction
abstract
Electronic medical records (EMRs) play an important role in medical data mining and sequential data learning. In this article, we propose to use a sequential neural network with dynamic content-based memories to predict future medications, given EMRs. The local-global memory neural network contains two layers of memories: the local memory and the global memory. Particularly, our method learns the hidden knowledge within EMRs by locally remembering individual patterns of a patient (via local memory) and globally remembering group evidence of disease (via global memory). In addition, we show how our model can be modified to classify the hidden states of EMRs from different patients at each time step into different phases that indicate the progressions of medications in terms of a specific disease, in an unsupervised manner. Experimental results on real EMRs data sets show that, by learning EMRs with external local and global memories, with regard to a given disease, our model improves the prediction performance compared with several alternative methods.
Jun Song 0004, Siliang Tang, Yin Zhang 0006, Zhigang Chen 0003, Zhongfei Zhang, Tong Zhang 0001, Fei Wu 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 Episode-Based Prototype Generating Network for Zero-Shot Learning
abstract
We introduce a simple yet effective episode-based training framework for zero-shot learning (ZSL), where the learning system requires to recognize unseen classes given only the corresponding class semantics. During training, the model is trained within a collection of episodes, each of which is designed to simulate a zero-shot classification task. Through training multiple episodes, the model progressively accumulates ensemble experiences on predicting the mimetic unseen classes, which will generalize well on the real unseen classes. Based on this training framework, we propose a novel generative model that synthesizes visual prototypes conditioned on the class semantic prototypes. The proposed model aligns the visual-semantic interactions by formulating both the visual prototype generation and the class semantic inference into an adversarial framework paired with a parameter-economic Multi-modal Cross-Entropy Loss to capture the discriminative information. Extensive experiments on four datasets under both traditional ZSL and generalized ZSL tasks show that our model outperforms the state-of-the-art approaches by large margins.
Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang
CVPR4
2020 BERT-enhanced Relational Sentence Ordering Network
abstract
In this paper, we introduce a novel BERTenhanced Relational Sentence Ordering Network (referred to as BERSON) by leveraging BERT for capturing a better dependency relationship among sentences to enhance the coherence modeling for the entire paragraph.In particular, we develop a new Relational Pointer Decoder (referred as RPD) by incorporating the relative ordering information into the pointer network with a Deep Relational Module (referred as DRM), which utilizes BERT to exploit the deep semantic connection and relative ordering between sentences.This enables us to strengthen both local and global dependencies among sentences.Extensive evaluations are conducted on six public datasets.The experimental results demonstrate the effectiveness and promise of BERSON, showing a significant improvement over the state-of-the-art by a wide margin.
Baiyun Cui, Yingming Li, Zhongfei Zhang
EMNLP (1)3
2020 Stacked Pooling for Boosting Scale Invariance of Crowd Counting
abstract
In this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance. Existing deep crowd counting approaches mainly focus on the multi-scale techniques over convolutional layers to capture scale-adaptive features, resulting in high computing costs. In this paper, we propose simple but effective pooling variants, i.e., multi-kernel pooling and stacked pooling, to take place of the vanilla pooling layers in convolutional neural networks (CNNs) for boosting the scale invariance. Our proposed pooling modules do not introduce extra parameters and can be easily implemented in practice. Empirical studies on two benchmark crowd counting datasets show that the proposed pooling modules beat the vanilla pooling layer in most experimental cases.
Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001
ICASSP4
2020 Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
abstract
Substantial efforts have been made on improving the generalization abilities of deep neural networks (DNNs) in order to obtain better performances without introducing more parameters. On the other hand, meta-learning approaches exhibit powerful generalization on new tasks in few-shot learning. Intuitively, few- shot learning is more challenging than the standard supervised learning as each target class only has a very few or no training samples. The natural question that arises is whether the meta-learning idea can be used for improving the generalization of DNNs on the standard supervised learning. In this paper, we propose a novel meta-learning based training procedure (MLTP) for DNNs and demonstrate that the meta-learning idea can indeed improve the generalization abilities of DNNs. MLTP simulates the meta-training process by considering a batch of training samples as a task. The key idea is that the gradient descent step for improving the current task performance should also improve a new task performance, which is ignored by the current standard procedure for training neural networks. MLTP also benefits from all the existing training techniques such as dropout, weight decay, and batch normalization. We evaluate MLTP by training a variety of small and large neural networks on three benchmark datasets, i.e., CIFAR-10, CIFAR-100, and Tiny ImageNet. The experimental results show a consistently improved generalization performance on all the DNNs with different sizes, which verifies the promise of MLTP and demonstrates that the meta-learning idea is indeed able to improve the generalization of DNNs on the standard supervised learning.
Xiang Deng 0002, Zhongfei Zhang
ICPR2
2020 SBAT: Video Captioning with Sparse Boundary-Aware Transformer
abstract
In this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation task such as machine translation. However, video captioning is a multimodal learning problem, and the video features have much redundancy between different time steps. Based on these concerns, we propose a novel method called sparse boundary-aware transformer (SBAT) to reduce the redundancy in video representation. SBAT employs boundary-aware pooling operation for scores from multihead attention and selects diverse features from different scenarios. Also, SBAT includes a local correlation scheme to compensate for the local information loss brought by sparse operation. Based on SBAT, we further propose an aligned cross-modal encoding scheme to boost the multimodal interaction. Experimental results on two benchmark datasets show that SBAT outperforms the state-of-the-art methods under most of the metrics.
Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang
IJCAI5
2020 Deep Metric Learning with Spherical Embedding
abstract
Deep metric learning has attracted much attention in recent years, due to seamlessly combining the distance metric learning and deep neural network. Many endeavors are devoted to design different pair-based angular loss functions, which decouple the magnitude and direction information for embedding vectors and ensure the training and testing measure consistency. However, these traditional angular losses cannot guarantee that all the sample embeddings are on the surface of the same hypersphere during the training stage, which would result in unstable gradient in batch optimization and may influence the quick convergence of the embedding learning. In this paper, we first investigate the effect of the embedding norm for deep metric learning with angular distance, and then propose a spherical embedding constraint (SEC) to regularize the distribution of the norms. SEC adaptively adjusts the embeddings to fall on the same hypersphere and performs more balanced direction update. Extensive experiments on deep metric learning, face recognition, and contrastive self-supervised learning show that the SEC-based angular space learning strategy significantly improves the performance of the state-of-the-art.
Dingyi Zhang, Yingming Li, Zhongfei Zhang
NeurIPS3
2020 Multi-modal generative adversarial network for zero-shot learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Zhongfei Zhang
Knowl. Based Syst.5
2020 Improved prototypical networks for few-Shot learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Pattern Recognit. Lett.5
2020 Attention-Aware Multi-Task Convolutional Neural Networks
abstract
Multi-task deep learning methods learn multiple tasks simultaneously and share representations amongst them, so information from related tasks improves learning within one task. The generalization capabilities of the produced models are substantially enhanced. Typical multi-task deep learning models usually share representations of different tasks in lower layers of the network, and separate representations of different tasks in higher layers. However, different groups of tasks always have different requirements for sharing representations, so the required design criterion does not necessarily guarantee that the obtained network architecture is optimal. In addition, most existing methods ignore the redundancy problem and lack the pre-screening process for representations before they are shared. Here, we propose a model called Attention-aware Multi-task Convolutional Neural Network, which automatically learns appropriate sharing through end-to-end training. The attention mechanism is introduced into our architecture to suppress redundant contents contained in the representations. The shortcut connection is adopted to preserve useful information. We evaluate our model by carrying out experiments on different task groups and different datasets. Our model demonstrates an improvement over existing techniques in many experiments, indicating the effectiveness and the robustness of the model. We also demonstrate the importance of attention mechanism and shortcut connection in our model.
Kejie Lyu, Yingming Li, Zhongfei Zhang
IEEE Trans. Image Process.3
2019 Spatio-Temporal Graph Routing for Skeleton-Based Action Recognition
abstract
With the representation effectiveness, skeleton-based human action recognition has received considerable research attention, and has a wide range of real applications. In this area, many existing methods typically rely on fixed physicalconnectivity skeleton structure for recognition, which is incapable of well capturing the intrinsic high-order correlations among skeleton joints. In this paper, we propose a novel spatio-temporal graph routing (STGR) scheme for skeletonbased action recognition, which adaptively learns the intrinsic high-order connectivity relationships for physicallyapart skeleton joints. Specifically, the scheme is composed of two components: spatial graph router (SGR) and temporal graph router (TGR). The SGR aims to discover the connectivity relationships among the joints based on sub-group clustering along the spatial dimension, while the TGR explores the structural information by measuring the correlation degrees between temporal joint node trajectories. The proposed scheme is naturally and seamlessly incorporated into the framework of graph convolutional networks (GCNs) to produce a set of skeleton-joint-connectivity graphs, which are further fed into the classification networks. Moreover, an insightful analysis on receptive field of graph node is provided to explain the necessity of our method. Experimental results on two benchmark datasets (NTU-RGB+D and Kinetics) demonstrate the effectiveness against the state-of-the-art.
Bin Li 0038, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001
AAAI3
2019 Cross-Relation Cross-Bag Attention for Distantly-Supervised Relation Extraction
abstract
Distant supervision leverages knowledge bases to automatically label instances, thus allowing us to train relation extractor without human annotations. However, the generated training data typically contain massive noise, and may result in poor performances with the vanilla supervised learning. In this paper, we propose to conduct multi-instance learning with a novel Cross-relation Cross-bag Selective Attention (C2SA), which leads to noise-robust training for distant supervised relation extractor. Specifically, we employ the sentence-level selective attention to reduce the effect of noisy or mismatched sentences, while the correlation among relations were captured to improve the quality of attention weights. Moreover, instead of treating all entity-pairs equally, we try to pay more attention to entity-pairs with a higher quality. Similarly, we adopt the selective attention mechanism to achieve this goal. Experiments with two types of relation extractor demonstrate the superiority of the proposed approach over the state-of-the-art, while further ablation studies verify our intuitions and demonstrate the effectiveness of our proposed two techniques.
Yujin Yuan, Siliang Tang, Zhongfei Zhang, Yueting Zhuang, Shiliang Pu, Fei Wu 0001, Xiang Ren 0001
AAAI4
2019 Learning a Key-Value Memory Co-Attention Matching Network for Person Re-Identification
abstract
Person re-identification (Re-ID) is typically cast as the problem of semantic representation and alignment, which requires precisely discovering and modeling the inherent spatial structure information on person images. Motivated by this observation, we propose a Key-Value Memory Matching Network (KVM-MN) model that consists of key-value memory representation and key-value co-attention matching. The proposed KVM-MN model is capable of building an effective local-position-aware person representation that encodes the spatial feature information in the form of multi-head key-value memory. Furthermore, the proposed KVM-MN model makes use of multi-head co-attention to automatically learn a number of cross-person-matching patterns, resulting in more robust and interpretable matching results. Finally, we build a setwise learning mechanism that implements a more generalized query-to-gallery-image-set learning procedure. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art.
Xi Li 0001, Zhongfei Zhang
AAAI3
2019 Text Guided Person Image Synthesis
abstract
This paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled down to two stages: 1) text guided pose generation and 2) visual appearance transferred image synthesis. In the first stage, our method infers a reasonable target human pose based on the text. In the second stage, our method synthesizes a realistic and appearance transferred person image according to the text in conjunction with the target pose. Our method extracts sufficient information from the text and establishes a mapping between the image space and the language space, making generating and editing images corresponding to the description possible. We conduct extensive experiments to reveal the effectiveness of our method, as well as using the VQA Perceptual Score as a metric for evaluating the method. It shows for the first time that we can automatically edit the person image from the natural language descriptions.
Xingran Zhou, Siyu Huang, Bin Li 0038, Yingming Li, Zhongfei Zhang
CVPR6
2019 Fine-tune BERT with Sparse Self-Attention Mechanism
abstract
Baiyun Cui, Yingming Li, Ming Chen, Zhongfei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Baiyun Cui, Yingming Li, Zhongfei Zhang
EMNLP/IJCNLP (1)4
2019 Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning
abstract
Tao Jin, Siyu Huang, Yingming Li, Zhongfei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang
EMNLP/IJCNLP (1)4
2019 Efficient Network Representations Learning: An Edge-Centric Perspective
Shichao Liu 0002, Shuangfei Zhai, Lida Zhu, Fuxi Zhu, Zhongfei Zhang, Wen Zhang 0008
KSEM (2)5
2019 Recurrent convolutional video captioning with global and local attention
Tao Jin 0004, Yingming Li, Zhongfei Zhang
Neurocomputing3
2019 A Bilinear Ranking SVM for Knowledge Based Relation Prediction and Classification
abstract
As an important and challenging problem, knowledge representation and inference are typically carried out in a knowledge embedding framework over a multi-relational knowledge graph, and thus have a wide range of applications such as semantic retrieval and question answering. In this paper, we propose a bilinear learning framework which performs cross-entity knowledge relation analysis in the continuous vector space (derived from knowledge embedding). In the framework, we effectively model the intrinsic correlations among different types of knowledge relations within a max-margin multi-relational ranking scheme, which jointly optimizes the tasks of entity embedding and cross-entity relation prediction in terms of multi-relational structures of the knowledge graph. Specifically, we devise a bilinear scoring function that aims to evaluate the confidence degree of semantic relation prediction for entity pairs through a multi-relational learning-to-rank pipeline. In essence, the pipeline formulates the problem of relation prediction for entity pairs as that of learning relation-specific ranking functions by max-margin optimization. Experimental results demonstrate the effectiveness of the proposed framework on two common benchmark datasets.
Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Jingdong Wang 0001, Yueting Zhuang, Xuelong Li 0001
IEEE Trans. Big Data4
2019 Zero-Shot Learning via Latent Space Encoding
abstract
Zero-shot learning (ZSL) is typically achieved by resorting to a class semantic embedding space to transfer the knowledge from the seen classes to unseen ones. Capturing the common semantic characteristics between the visual modality and the class semantic modality (e.g., attributes or word vector) is a key to the success of ZSL. In this paper, we propose a novel encoder-decoder approach, namely latent space encoding (LSE), to connect the semantic relations of different modalities. Instead of requiring a projection function to transfer information across different modalities like most previous work, LSE performs the interactions of different modalities via a feature aware latent space, which is learned in an implicit way. Specifically, different modalities are modeled separately but optimized jointly. For each modality, an encoder-decoder framework is performed to learn a feature aware latent space via jointly maximizing the recoverability of the original space from the latent space and the predictability of the latent space from the original space. To relate different modalities together, their features referring to the same concept are enforced to share the same latent codings. In this way, the common semantic characteristics of different modalities are generalized with the latent representations. Another property of the proposed approach is that it is easily extended to more modalities. Extensive experimental results on four benchmark datasets [animal with attribute, Caltech UCSD birds, aPY, and ImageNet] clearly demonstrate the superiority of the proposed approach on several ZSL tasks, including traditional ZSL, generalized ZSL, and zero-shot retrieval.
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Zhongfei Zhang
IEEE Trans. Cybern.4
2019 User-Ranking Video Summarization With Multi-Stage Spatio-Temporal Representation
abstract
Video summarization is a challenging task, mainly due to the difficulties in learning complicated semantic structural relations between videos and summaries. In this paper, we present a novel supervised video summarization scheme based on threestage deep neural networks. The scheme takes a divide-andconquer strategy to resolve the complicated task of 3D video summarization into a set of easy and flexible computational subtasks, and then to sequentially perform 2D CNNs, 1D CNNs, and LSTM to address the subtasks in an hierarchical fashion. The hierarchical modeling of spatio-temporal structure leads to high performance and efficiency. In addition, we propose a simple but effective user-ranking method to cope with the labeling subjectivity problem of user-created video summarization, leading to the labeling quality refinement for robust supervised learning. Experimental results show that our approach outperforms the state-of-the-art video summarization methods on two benchmark datasets.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Junwei Han 0001
IEEE Trans. Image Process.3
2019 A Survey of Multi-View Representation Learning
abstract
Recently, multi-view representation learning has become a rapidly growing direction in machine learning and data mining areas. This paper introduces two categories for multi-view representation learning: multi-view representation alignment and multi-view representation fusion. Consequently, we first review the representative methods and theories of multi-view representation learning based on the perspective of alignment, such as correlation-based alignment. Representative examples are canonical correlation analysis (CCA) and its several extensions. Then, from the perspective of representation fusion, we investigate the advancement of multi-view representation learning that ranges from generative methods including multi-modal topic learning, multi-view sparse coding, and multi-view latent space Markov networks, to neural network-based methods including multi-modal autoencoders, multi-view convolutional neural networks, and multi-modal recurrent neural networks. Further, we also investigate several important applications of multi-view representation learning. Overall, this survey aims to provide an insightful overview of theoretical foundation and state-of-the-art developments in the field of multi-view representation learning and to help researchers find the most appropriate tools for particular applications.
Yingming Li, Ming Yang 0012, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.3
2018 FR-ANet: A Face Recognition Guided Facial Attribute Classification Network
abstract
In this paper, we study the problem of facial attribute learning. In particular, we propose a Face Recognition guided facial Attribute classification Network, called FR-ANet. All the attributes share low-level features, while high-level features are specially learned for attribute groups. Further, to utilize the identity information, high-level features are merged to perform face identity recognition. The experimental results on CelebA and LFWA datasets demonstrate the promise of the FR-ANet.
Jiajiong Cao, Yingming Li, Xi Li 0001, Zhongfei Zhang
AAAI4
2018 Learning With Incomplete Labels
Yingming Li, Zenglin Xu, Zhongfei Zhang
AAAI3
2018 Multi-Channel Pyramid Person Matching Network for Person Re-Identification
abstract
In this work, we present a Multi-Channel deep convolutional Pyramid Person Matching Network (MC-PPMN) based on the combination of the semantic-components and the color-texture distributions to address the problem of person re-identification. In particular, we learn separate deep representations for semantic-components and color-texture distributions from two person images and then employ pyramid person matching network (PPMN) to obtain correspondence representations. These correspondence representations are fused to perform the re-identification task. Further, the proposed framework is optimized via a unified end-to-end deep learning scheme. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our approach against the state-of-the-art literature, especially on the rank-1 recognition rate.
Chaojie Mao, Yingming Li, Zhongfei Zhang, Xi Li 0001
AAAI4
2018 TVT: Two-View Transformer Network for Video Captioning
abstract
Video captioning is a task of automatically generating the natural text description of a given video. There are two main challenges in video captioning under the context of an encoder-decoder framework: 1) How to model the sequential information; 2) How to combine the modalities including video and text. For challenge 1), the recurrent neural networks (RNNs) based methods are currently the most common approaches for learning temporal representations of videos, while they suffer from a high computational cost. For challenge 2), the features of different modalities are often roughly concatenated together without insightful discussion. In this paper, we introduce a novel video captioning framework, i.e., Two-View Transformer (TVT). TVT comprises of a backbone of Transformer network for sequential representation and two types of fusion blocks in decoder layers for combining different modalities effectively. Empirical study shows that our TVT model outperforms the state-of-the-art methods on the MSVD dataset and achieves a competitive performance on the MSR-VTT dataset under four common metrics.
Yingming Li, Zhongfei Zhang, Siyu Huang
ACML3
2018 Relative Attribute Learning with Deep Attentive Cross-image Representation
abstract
In this paper, we study the relative attribute learning problem, which refers to comparing the strengths of a specific attribute between image pairs, with a new perspective of cross-image representation learning. In particular, we introduce a deep attentive cross-image representation learning (DACRL) model, which first extracts single-image representation with one shared subnetwork, and then learns attentive cross-image representation through considering the channel-wise attention of concatenated single-image feature maps. Taking a pair of images as input, DACRL outputs a posterior probability indicating whether the first image in the pair has a stronger presence of attribute than the second image. The whole network is jointly optimized via a unified end-to-end deep learning scheme. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our approach against the state-of-the-art methods.
Zeshang Zhang, Yingming Li, Zhongfei Zhang
ACML3
2018 Partially Shared Multi-Task Convolutional Neural Network With Local Constraint for Face Attribute Learning
abstract
In this paper, we study the face attribute learning problem by considering the identity information and attribute relationships simultaneously. In particular, we first introduce a Partially Shared Multi-task Convolutional Neural Network (PS-MCNN), in which four Task Specific Networks (TSNets) and one Shared Network (SNet) are connected by Partially Shared (PS) structures to learn better shared and task specific representations. To utilize identity information to further boost the performance, we introduce a local learning constraint which minimizes the difference between the representations of each sample and its local geometric neighbours with the same identity. Consequently, we present a local constraint regularized multitask network, called Partially Shared Multi-task Convolutional Neural Network with Local Constraint (PS-MCNN-LC), where PS structure and local constraint are integrated together to help the framework learn better attribute representations. The experimental results on CelebA and LFWA demonstrate the promise of the proposed methods.
Jiajiong Cao, Yingming Li, Zhongfei Zhang
CVPR3
2018 Deep Attentive Sentence Ordering Network
abstract
In this paper, we propose a novel deep attentive sentence ordering network (referred as ATTOrderNet) which integrates self-attention mechanism with LSTMs in the encoding of input sentences.It enables us to capture global dependencies among sentences regardless of their input order and obtains a reliable representation of the sentence set.With this representation, a pointer network is exploited to generate an ordered sequence.The proposed model is evaluated on Sentence Ordering and Order Discrimination tasks.The extensive experimental results demonstrate its effectiveness and superiority to the state-ofthe-art methods.
Baiyun Cui, Yingming Li, Zhongfei Zhang
EMNLP4
2018 Celeb-500K: A Large Training Dataset for Face Recognition
abstract
In this paper, we propose a large training dataset named Celeb-500K for face recognition, which contains 50M images from 500K persons. To better facilitate academic research, we clean Celeb-500K to obtain Celeb-500K-2R, which contains 25M aligned face images from 365K persons. Based on the developed dataset, we achieve state-of-the-art face recognition performance and reveal two important observations on face recognition study. First, metric learning methods have limited performance gain when the training dataset contains a large number of identities. Second, in order to develop an efficient training dataset, the number of identities is more important than the average image number of each identity from the perspective of face recognition performance. Extensive experimental results show the superiority of Celeb-500K and provide a strong support to the two observations.
Jiajiong Cao, Yingming Li, Zhongfei Zhang
ICIP3
2018 Image Annotation Retrieval with Text-Domain Label Denoising
abstract
This work explores the problem of making user-generated text data, in the form of noisy tags, usable for tasks such as automatic image annotation and image retrieval by denoising the data. Earlier work in this area has focused on filtering out noisy, sparse, or incorrect tags by representing an image by the accumulation of the tags of its nearest neighbors in the visual space. However, this imposes an expensive preprocessing step that must be performed for each new set of images and tags and relies on assumptions about the way the images have been labelled that we find do not always hold. We instead propose a technique for calculating a set of probabilities for the relevance of each tag for a given image relying soley on information in the text domain, namely through widely-available pretrained continous word embeddings. By first clustering the word embeddings for the tags, we calculate a set of weights representing the probability that each tag is meaningful to the image content. Given the set of tags denoised in this way, we use kernel canonical correlation analysis (KCCA) to learn a semantic space which we can project into to retrieve relevant tags for unseen images or to retrieve images for unseen tags. This work also explores the deficiencies of the use of continuous word embeddings for automatic image annotation in the existing KCCA literature and introduces a new method for constructing textual kernel matrices using these word vectors that improves tag retrieval results for both user-generated tags as well as expert labels.
Zachary Seymour, Zhongfei Zhang
ICMR2
2018 Multi-label Triplet Embeddings for Image Annotation from User-Generated Tags
abstract
This work studies the representational embedding of images and their corresponding annotations--in the form of tag metadata--such that, given a piece of the raw data in one modality, the corresponding semantic description can be retrieved in terms of the raw data in another. While convolutional neural networks (CNNs) have been widely and successfully applied in this domain with regards to detecting semantically simple scenes or categories (even though many such objects may be simultaneously present in an image), this work approaches the task of dealing with image annotations in the context of noisy, user-generated, and semantically complex multi-labels, widely available from social media sites. In this case, the labels for an image are diverse, noisy, and often not specifically related to an object, but rather descriptive or user-specific. Furthermore, the existing deep image annotation literature using this type of data typically utilizes the so-called CNN-RNN framework, combining convolutional and recurrent neural networks. We offer a discussion of why RNNs may not be the best choice in this case, though they have been shown to perform well on the similar captioning tasks. Our model exploits the latent image-text space through the use of a triplet loss framework to learn a joint embedding space for the images and their tags, in the presence of multiple, potentially positive exemplar classes. We present state-of-the-art results of the representational properties of these embeddings on several image annotation datasets to show the promise of this approach.
Zachary Seymour, Zhongfei Zhang
ICMR2
2018 GNAS: A Greedy Neural Architecture Search Method for Multi-Attribute Learning
abstract
A key problem in deep multi-attribute learning is to effectively discover the inter-attribute correlation structures. Typically, the conventional deep multi-attribute learning approaches follow the pipeline of manually designing the network architectures based on task-specific expertise prior knowledge and careful network tunings, leading to the inflexibility for various complicated scenarios in practice. Motivated by addressing this problem, we propose an efficient greedy neural architecture search approach (GNAS) to automatically discover the optimal tree-like deep architecture for multi-attribute learning. In a greedy manner, GNAS divides the optimization of global architecture into the optimizations of individual connections step by step. By iteratively updating the local architectures, the global tree-like architecture gets converged where the bottom layers are shared across relevant attributes and the branches in top layers more encode attribute-specific features. Experiments on three benchmark multi-attribute datasets show the effectiveness and compactness of neural architectures derived by GNAS, and also demonstrate the efficiency of GNAS in searching neural architectures.
Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001
ACM Multimedia4
2018 Stacked Semantics-Guided Attention Model for Fine-Grained Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) is generally achieved via aligning the semantic relationships between the visual features and the corresponding class semantic descriptions. However, using the global features to represent fine-grained images may lead to sub-optimal results since they neglect the discriminative differences of local regions. Besides, different regions contain distinct discriminative information. The important regions should contribute more to the prediction. To this end, we propose a novel stacked semantics-guided attention (S2GA) model to obtain semantic relevant features by using individual class semantic features to progressively guide the visual features to generate an attention map for weighting the importance of different local regions. Feeding both the integrated visual features and the class semantic features into a multi-class classification architecture, the proposed framework can be trained end-to-end. Extensive experimental results on CUB and NABird datasets show that the proposed approach has a consistent improvement on both fine-grained zero-shot classification and retrieval tasks.
Yunlong Yu 0001, Zhong Ji, Yanwei Fu 0001, Jichang Guo, Yanwei Pang, Zhongfei Zhang
NeurIPS6
2018 Multimodal Deep Embedding via Hierarchical Grounded Compositional Semantics
abstract
For a number of important problems, isolated semantic representations of individual syntactic words or visual objects do not suffice, but instead a compositional semantic representation is required; for example, a literal phrase or a set of spatially concurrent objects. In this paper, we aim to harness the existing image-sentence databases to exploit the compositional nature of image-sentence data for multimodal deep embedding. In particular, we propose an approach called hierarchical-alike (bottom-up two layers) multimodal grounded compositional semantics (hiMoCS) learning. The proposed hiMoCS systemically captures the compositional semantic connotation of multimodal data in the setting of hierarchical-alike deep learning by modeling the inherent correlations between two modalities of collaboratively grounded semantics, such as the textual entity (with its describing attribute) and visual object, the phrase (e.g., subject-verb-object triplet), and spatially concurrent objects. We argue that hiMoCS is more appropriate to reflect the multimodal compositional semantics of the image and its narrative textual sentence, which are strongly coupled. We evaluate hiMoCS on the several benchmark data sets and show that the utilization of the hiMoCS (textual entities and visual objects, textual phrase, and spatially concurrent objects) achieves a much better performance than only using the flat grounded compositional semantics.
Yueting Zhuang, Jun Song 0004, Fei Wu 0001, Xi Li 0001, Zhongfei Zhang, Yong Rui
IEEE Trans. Circuits Syst. Video Technol.5
2018 Transductive Zero-Shot Learning With a Self-Training Dictionary Approach
abstract
As an important and challenging problem in computer vision, zero-shot learning (ZSL) aims at automatically recognizing the instances from unseen object classes without training data. To address this problem, ZSL is usually carried out in the following two aspects: 1) capturing the domain distribution connections between seen classes data and unseen classes data and 2) modeling the semantic interactions between the image feature space and the label embedding space. Motivated by these observations, we propose a bidirectional mapping-based semantic relationship modeling scheme that seeks for cross-modal knowledge transfer by simultaneously projecting the image features and label embeddings into a common latent space. Namely, we have a bidirectional connection relationship that takes place from the image feature space to the latent space as well as from the label embedding space to the latent space. To deal with the domain shift problem, we further present a transductive learning approach that formulates the class prediction problem in an iterative refining process, where the object classification capacity is progressively reinforced through bootstrapping-based model updating over highly reliable instances. Experimental results on four benchmark datasets (animal with attribute, Caltech-UCSD Bird2011, aPascal-aYahoo, and SUN) demonstrate the effectiveness of the proposed approach against the state-of-the-art approaches.
Yunlong Yu 0001, Zhong Ji, Xi Li 0001, Jichang Guo, Zhongfei Zhang, Haibin Ling, Fei Wu 0001
IEEE Trans. Cybern.5
2018 Body Structure Aware Deep Crowd Counting
abstract
Crowd counting is a challenging task, mainly due to the severe occlusions among dense crowds. This paper aims to take a broader view to address crowd counting from the perspective of semantic modeling. In essence, crowd counting is a task of pedestrian semantic analysis involving three key factors: pedestrians, heads, and their context structure. The information of different body parts is an important cue to help us judge whether there exists a person at a certain position. Existing methods usually perform crowd counting from the perspective of directly modeling the visual properties of either the whole body or the heads only, without explicitly capturing the composite body-part semantic structure information that is crucial for crowd counting. In our approach, we first formulate the key factors of crowd counting as semantic scene models. Then, we convert the crowd counting problem into a multi-task learning problem, such that the semantic scene models are turned into different sub-tasks. Finally, the deep convolutional neural networks are used to learn the sub-tasks in a unified scheme. Our approach encodes the semantic nature of crowd counting and provides a novel solution in terms of pedestrian semantic analysis. In experiments, our approach outperforms the state-of-the-art methods on four benchmark crowd counting data sets. The semantic structure information is demonstrated to be an effective cue in scene of crowd counting.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Shenghua Gao, Rongrong Ji, Junwei Han 0001
IEEE Trans. Image Process.3
2018 Deep Air Learning: Interpolation, Prediction, and Feature Analysis of Fine-Grained Air Quality
abstract
The interpolation, prediction, and feature analysis of fine-gained air quality are three important topics in the area of urban air computing. The solutions to these topics can provide extremely useful information to support air pollution control, and consequently generate great societal and technical impacts. Most of the existing work solves the three problems separately by different models. In this paper, we propose a general and effective approach to solve the three problems in one model called the Deep Air Learning (DAL). The main idea of DAL lies in embedding feature selection and semi-supervised learning in different layers of the deep learning network. The proposed approach utilizes the information pertaining to the unlabeled spatio-temporal data to improve the performance of the interpolation and the prediction, and performs feature selection and association analysis to reveal the main relevant features to the variation of the air quality. We evaluate our approach with extensive experiments based on real data sources obtained in Beijing, China. Experiments show that DAL is superior to the peer models from the recent literature when solving the topics of interpolation, prediction, and feature analysis of fine-gained air quality.
Zhongang Qi, Tianchun Wang, Guojie Song, Weisong Hu, Xi Li 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.6
2018 Scalable Distributed Nonnegative Matrix Factorization with Block-Wise Updates
abstract
Nonnegative Matrix Factorization (NMF) has been applied with great success on a wide range of applications. As NMF is increasingly applied to massive datasets such as web-scale dyadic data, it is desirable to leverage a cluster of machines to store those datasets and to speed up the factorization process. However, it is challenging to efficiently implement NMF in a distributed environment. In this paper, we show that by leveraging a new form of update functions, we can perform local aggregation and fully explore parallelism. Therefore, the new form is much more efficient than the traditional form in distributed implementations. Moreover, under the new form of update functions, we can perform frequent updates and lazy updates, which aim to use the most recently updated data whenever possible and avoid unnecessary computations. As a result, frequent updates and lazy updates are more efficient than their traditional concurrent counterparts. Through a series of experiments on a local cluster as well as the Amazon EC2 cloud, we demonstrate that our implementations with frequent updates or lazy updates are up to two orders of magnitude faster than the existing implementation with the traditional form of update functions.
Jiangtao Yin, Lixin Gao 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.3
2017 Learning with Feature Network and Label Network Simultaneously
Yingming Li, Ming Yang 0012, Zenglin Xu, Zhongfei Zhang
AAAI4
2017 Structural Correspondence Learning for Cross-Lingual Sentiment Classification with One-to-Many Mappings
abstract
Structural correspondence learning (SCL) is an effective method for cross-lingual sentiment classification. This approach uses unlabeled documents along with a word translation oracle to automatically induce task specific, cross-lingual correspondences. It transfers knowledge through identifying important features, i.e., pivot features. For simplicity, however, it assumes that the word translation oracle maps each pivot feature in source language to exactly only one word in target language. This one-to-one mapping between words in different languages is too strict. Also the context is not considered at all. In this paper, we propose a cross-lingual SCL based on distributed representation of words; it can learn meaningful one-to-many mappings for pivot words using large amounts of monolingual data and a small dictionary. We conduct experiments on NLP&CC 2013 cross-lingual sentiment analysis dataset, employing English as source language, and Chinese as target language. Our method does not rely on the parallel corpora and the experimental results show that our approach is more competitive than the state-of-the-art methods in cross-lingual sentiment classification.
Shuangfei Zhai, Zhongfei Zhang, Boying Liu
AAAI3
2017 Pyramid Person Matching Network for Person Re-identification
abstract
In this work, we present a deep convolutional pyramid person matching network (PPMN) with specially designed Pyramid Matching Module to address the problem of person re-identification. The architecture takes a pair of RGB images as input, and outputs a similiarity value indicating whether the two input images represent the same person or not. Based on deep convolutional neural networks, our approach first learns the discriminative semantic representation with the semantic-component-aware features for persons and then employs the Pyramid Matching Module to match the common semantic-components of persons, which is robust to the variation of spatial scales and misalignment of locations posed by viewpoint changes. The above two processes are jointly optimized via a unified end-to-end deep learning scheme. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our approach against the state-of-the-art approaches, especially on the rank-1 recognition rate.
Chaojie Mao, Yingming Li, Zhongfei Zhang, Xi Li 0001
ACML3
2017 Text Coherence Analysis Based on Deep Neural Network
abstract
In this paper, we propose a novel deep coherence model (DCM) using a convolutional neural network architecture to capture the text coherence. The text coherence problem is investigated with a new perspective of learning sentence distributional representation and text coherence modeling simultaneously. In particular, the model captures the interactions between sentences by computing the similarities of their distributional representations. Further, it can be easily trained in an end-to-end fashion. The proposed model is evaluated on a standard Sentence Ordering task. The experimental results demonstrate its effectiveness and promise in coherence assessment showing a significant improvement over the state-of-the-art by a wide margin.
Baiyun Cui, Yingming Li, Zhongfei Zhang
CIKM4
2017 S3Pool: Pooling with Stochastic Spatial Sampling
Shuangfei Zhai, Hui Wu 0009, Abhishek Kumar 0001, Yu Cheng 0001, Yongxi Lu, Zhongfei Zhang, Rogério Feris
CVPR6
2017 Boosted Zero-Shot Learning with Semantic Correlation Regularization
abstract
We study zero-shot learning (ZSL) as a transfer learning problem, and focus on the two key aspects of ZSL, model effectiveness and model adaptation. For effective modeling, we adopt the boosting strategy to learn a zero-shot classifier from weak models to a strong model. For adaptable knowledge transfer, we devise a Semantic Correlation Regularization (SCR) approach to regularize the boosted model to be consistent with the inter-class semantic correlations. With SCR embedded in the boosting objective, and with a self-controlled sample selection for learning robustness, we propose a unified framework, Boosted Zero-shot classification with Semantic Correlation Regularization (BZ-SCR). By balancing the SCR-regularized boosted model selection and the self-controlled sample selection, BZ-SCR is capable of capturing both discriminative and adaptable feature-to-class semantic alignments, while ensuring the reliability and adaptability of the learned samples. The experiments on two ZSL datasets show the superiority of BZ-SCR over the state-of-the-arts.
Te Pi, Xi Li 0001, Zhongfei Zhang
IJCAI3
2017 Manifold regularized cross-modal embedding for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jichang Guo, Zhongfei Zhang
Inf. Sci.5
2017 Joint entity-relation knowledge embedding via cost-sensitive learning
abstract
As a joint-optimization problem which simultaneously fulfills two different but correlated embedding tasks (i.e., entity embedding and relation embedding), knowledge embedding problem is solved in a joint embedding scheme. In this embedding scheme, we design a joint compatibility scoring function to quantitatively evaluate the relational facts with respect to entities and relations, and further incorporate the scoring function into the max-margin structure learning process that explicitly learns the embedding vectors of entities and relations using the context information of the knowledge base. By optimizing the joint problem, our design is capable of effectively capturing the intrinsic topological structures in the learned embedding spaces. Experimental results demonstrate the effectiveness of our embedding scheme in characterizing the semantic correlations among different relation units, and in relation prediction for knowledge inference.
Shengkang Yu, Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
Frontiers Inf. Technol. Electron. Eng.4
2017 Zero-shot learning with Multi-Battery Factor Analysis
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Signal Process.5
2017 Data-Dependent Label Distribution Learning for Age Estimation
abstract
As an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples with respect to several ordinal age labels, which have intrinsically cross-age correlations across adjacent age dimensions. As a result, such correlations usually lead to the age label ambiguities of the face samples. Namely, each face sample is associated with a latent label distribution that encodes the cross-age correlation information on label ambiguities. Motivated by this observation, we propose a totally data-driven label distribution learning approach to adaptively learn the latent label distributions. The proposed approach is capable of effectively discovering the intrinsic age distribution patterns for cross-age correlation analysis on the basis of the local context structures of face samples. Without any prior assumptions on the forms of label distribution learning, our approach is able to flexibly model the sample-specific context aware label distribution properties by solving a multi-task problem, which jointly optimizes the tasks of age-label distribution learning and age prediction for individuals. Experimental results demonstrate the effectiveness of our approach.
Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Xin Geng 0001, Ming-Hsuan Yang 0001, Yueting Zhuang
IEEE Trans. Image Process.3
2017 Learning Bregman Distance Functions for Structural Learning to Rank
abstract
We study content-based learning to rank from the perspective of learning distance functions. Standardly, the two key issues of learning to rank, feature mappings and score functions, are usually modeled separately, and the learning is usually restricted to modeling a linear distance function such as the Mahalanobis distance. However, the modeling of feature mappings and score functions are mutually interacted, and the patterns underlying the data are probably complicated and nonlinear. Thus, as a general nonlinear distance family, the Bregman distance is a suitable distance function for learning to rank, due to its strong generalization ability for distance functions, and its nonlinearity for exploring the general patterns of data distributions. In this paper, we study learning to rank as a structural learning problem, and devise a Bregman distance function to build the ranking model based on structural SVM. To improve the model robustness to outliers, we develop a robust structural learning framework for the ranking model. The proposed model Robust Structural Bregman distance functions Learning to Rank (RSBLR) is a general and unified framework for learning distance functions to rank. The experiments of data ranking on real-world datasets show the superiority of this method to the state-of-the-art literature, as well as its robustness to the noisily labeled outliers.
Xi Li 0001, Te Pi, Zhongfei Zhang, Xueyi Zhao, Meng Wang 0001, Xuelong Li 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2017 Hierarchical Contextual Attention Recurrent Neural Network for Map Query Suggestion
abstract
The query logs from an on-line map query system provide rich cues to understand the behaviors of human crowds. With the growing ability of collecting large scale query logs, the query suggestion has been a topic of recent interest. In general, query suggestion aims at recommending a list of relevant queries w.r.t. users’ inputs via an appropriate learning of crowds’ query logs. In this paper, we are particularly interested in map query suggestions (e.g., the predictions of location-related queries) and propose a novel modelHierarchical Contextual Attention Recurrent Neural Network(HCAR-NN) for map query suggestion in an encoding-decoding manner. Given crowds map query logs, our proposed HCAR-NN not only learns the local temporal correlation among map queries in a query session (e.g., queries in a short-term interval are relevant to accomplish a search mission), but also captures the global longer range contextual dependencies among map query sessions in query logs (e.g., how a sequence of queries within a short-term interval has an influence on another sequence of queries). We evaluate our approach over millions of queries from a commercial search engine (i.e.,Baidu Map). Experimental results show that the proposed approach provides significant performance improvements over the competitive existing methods in terms of classical metrics (i.e.,Recall@KandMRR) as well as the prediction of crowds’ search missions.
Jun Song 0004, Jun Xiao 0001, Fei Wu 0001, Haishan Wu, Tong Zhang 0001, Zhongfei Zhang, Wenwu Zhu 0001
IEEE Trans. Knowl. Data Eng.6
2017 Bag-of-Discriminative-Words (BoDW) Representation via Topic Modeling
abstract
Many of the words in a given document either deliver facts (objective) or express opinions (subjective), respectively, depending on the topics they are involved in. For example, given a bunch of documents, the word “bug” assigned to the topic “order Hemiptera” apparently remarks one object (i.e., one kind of insects), while the same word assigned to the topic “software” probably conveys a negative opinion. Motivated by the intuitive assumption that different words have varying degrees of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as discriminatively objective-subjective LDA (dosLDA) is proposed in this paper. The essential idea underlying the proposed dosLDA is that a pair of objective and subjective selection variables are explicitly employed to encode the interplay between topics and discriminative power for the words in documents in a supervised manner. As a result, each document is appropriately represented as “bag-of-discriminativewords” (BoDW). The experiments reported on documents and images demonstrate that dosLDA not only performs competitively over traditional approaches in terms of topic modeling and document classification, but also has the ability to discern the discriminative power of each word in terms of its objective or subjective sense with respect to its assigned topic.
Yueting Zhuang, Hanqi Wang, Jun Xiao 0001, Fei Wu 0001, Yi Yang 0001, Weiming Lu 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.7
2016 Learning with Marginalized Corrupted Features and Labels Together
abstract
Tagging has become increasingly important in many real-world applications noticeably including web applications, such as web blogs and resource sharing systems. Despite this importance, tagging methods often face difficult challenges such as limited training samples and incomplete labels, which usually lead to degenerated performance on tag prediction. To improve the generalization performance, in this paper, we propose Regularized Marginalized Cross-View learning (RMCV) by jointly modeling on attribute noise and label noise. In more details, the proposed model constructs infinite training examples with attribute noises from known exponential-family distributions and exploits label noise via marginalized denoising autoencoder. Therefore, the model benefits from its robustness and alleviates the problem of tag sparsity. While RMCV is a general method for learning tagging, in the evaluations we focus on the specific application of multi-label text tagging. Extensive evaluations on three benchmark data sets demonstrate that RMCV outstands with a superior performance in comparison with state-of-the-art methods.
Yingming Li, Ming Yang 0012, Zenglin Xu, Zhongfei Zhang
AAAI4
2016 Semisupervised Autoencoder for Sentiment Analysis
abstract
In this paper, we investigate the usage of autoencoders in modeling textual data. Traditional autoencoders suffer from at least two aspects: scalability with the high dimensionality of vocabulary size and dealing with task-irrelevant words. We address this problem by introducing supervision via the loss function of autoencoders. In particular, we first train a linear classifier on the labeled data, then define a loss for the autoencoder with the weights learned from the linear classifier. To reduce the bias brought by one single classifier, we define a posterior probability distribution on the weights of the classifier, and derive the marginalized loss of the autoencoder with Laplace approximation. We show that our choice of loss function can be rationalized from the perspective of Bregman Divergence, which justifies the soundness of our model. We evaluate the effectiveness of our model on six sentiment analysis datasets, and show that our model significantly outperforms all the competing methods with respect to classification accuracy. We also show that our model is able to take advantage of unlabeled dataset and get improved performance. We further show that our model successfully learns highly discriminative feature maps, which explains its superior performance.
Shuangfei Zhai, Zhongfei Zhang
AAAI2
2016 Deep Structured Energy Based Models for Anomaly Detection
abstract
In this paper, we attack the anomaly detection problem by directly modeling the data distribution with deep architectures. We hence propose deep structured energy based models (DSEBMs), where the energy function is the output of a deterministic deep neural network with structure. We develop novel model architectures to integrate EBMs with different types of data such as static data, sequential data, and spatial data, and apply appropriate model architectures to adapt to the data structure. Our training algorithm is built upon the recent development of score matching (Hyvarinen, 2005), which connects an EBM with a regularized autoencoder, eliminating the need for complicated sampling method. Statistically sound decision criterion can be derived for anomaly detection purpose from the perspective of the energy landscape of the data distribution. We investigate two decision criteria for performing anomaly detection: the energy score and the reconstruction error. Extensive empirical studies on benchmark anomaly detection tasks demonstrate that our proposed model consistently matches or outperforms all the competing methods.
Shuangfei Zhai, Yu Cheng 0001, Weining Lu, Zhongfei Zhang
ICML4
2016 Multi-View Learning with Limited and Noisy Tagging
Yingming Li, Ming Yang 0012, Zenglin Xu, Zhongfei Zhang
IJCAI4
2016 Self-Paced Boost Learning for Classification
Te Pi, Xi Li 0001, Zhongfei Zhang, Deyu Meng, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang
IJCAI3
2016 Semantics-Aware Deep Correspondence Structure Learning for Robust Person Re-Identification
Xi Li 0001, Zhongfei Zhang
IJCAI4
2016 DeepIntent: Learning Attentions for Online Advertising with Recurrent Neural Networks
abstract
In this paper, we investigate the use of recurrent neural networks (RNNs) in the context of search-based online advertising. We use RNNs to map both queries and ads to real valued vectors, with which the relevance of a given (query, ad) pair can be easily computed. On top of the RNN, we propose a novel attention network, which learns to assign attention scores to different word locations according to their intent importance (hence the name DeepIntent). The vector output of a sequence is thus computed by a weighted sum of the hidden states of the RNN at each word according their attention scores. We perform end-to-end training of both the RNN and attention network under the guidance of user click logs, which are sampled from a commercial search engine. We show that in most cases the attention network improves the quality of learned vector representations, evaluated by AUC on a manually labeled dataset. Moreover, we highlight the effectiveness of the learned attention scores from two aspects: query rewriting and a modified BM25 metric. We show that using the learned attention scores, one is able to produce sub-queries that are of better qualities than those of the state-of-the-art methods. Also, by modifying the term frequency with the attention scores in a standard BM25 formula, one is able to improve its performance evaluated by AUC.
Shuangfei Zhai, Keng-hao Chang, Ruofei Zhang, Zhongfei Zhang
KDD4
2016 Doubly Convolutional Neural Networks
abstract
Building large models with parameter sharing accounts for most of the success of deep convolutional neural networks (CNNs). In this paper, we propose doubly convolutional neural networks (DCNNs), which significantly improve the performance of CNNs by further exploring this idea. In stead of allocating a set of convolutional filters that are independently learned, a DCNN maintains groups of filters where filters within each group are translated versions of each other. Practically, a DCNN can be easily implemented by a two-step convolution procedure, which is supported by most modern deep learning libraries. We perform extensive experiments on three image classification benchmarks: CIFAR-10, CIFAR-100 and ImageNet, and show that DCNNs consistently outperform other competing architectures. We have also verified that replacing a convolutional layer with a doubly convolutional layer at any depth of a CNN can improve its performance. Moreover, various design choices of DCNNs are demonstrated, which shows that DCNN can serve the dual purpose of building more accurate models and/or reducing the memory footprint without sacrificing the accuracy.
Shuangfei Zhai, Yu Cheng 0001, Zhongfei Zhang, Weining Lu
NIPS3
2016 LSTM-in-LSTM for generating long descriptions of images
abstract
In this paper, we propose an approach for generating rich fine-grained textual descriptions of images. In particular, we use an LSTM-in-LSTM (long short-term memory) architecture, which consists of an inner LSTM and an outer LSTM. The inner LSTM effectively encodes the long-range implicit contextual interaction between visual cues (i.e., the spatiallyconcurrent visual objects), while the outer LSTM generally captures the explicit multi-modal relationship between sentences and images (i.e., the correspondence of sentences and images). This architecture is capable of producing a long description by predicting one word at every time step conditioned on the previously generated word, a hidden vector (via the outer LSTM), and a context vector of fine-grained visual cues (via the inner LSTM). Our model outperforms state-of-theart methods on several benchmark datasets (Flickr8k, Flickr30k, MSCOCO) when used to generate long rich fine-grained descriptions of given images in terms of four different metrics (BLEU, CIDEr, ROUGE-L, and METEOR).
Jun Song 0004, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Zhongfei Zhang
Comput. Vis. Media5
2016 Visual search reranking with RElevant Local Discriminant Analysis
Peiguang Jing, Zhong Ji, Yunlong Yu 0001, Zhongfei Zhang
Neurocomputing4
2016 Online Metric-Weighted Linear Representations for Robust Visual Tracking
abstract
In this paper, we propose a visual tracker based on a metric-weighted linear representation of appearance. In order to capture the interdependence of different feature dimensions, we develop two online distance metric learning methods using proximity comparison information and structured output learning. The learned metric is then incorporated into a linear representation of appearance. We show that online distance metric learning significantly improves the robustness of the tracker, especially on those sequences exhibiting drastic appearance changes. In order to bound growth in the number of training samples, we design a time-weighted reservoir sampling method. Moreover, we enable our tracker to automatically perform object identification during the process of object tracking, by introducing a collection of static template samples belonging to several object classes of interest. Object identification results for an entire video sequence are achieved by systematically combining the tracking information and visual recognition at each frame. Experimental results on challenging video sequences demonstrate the effectiveness of the method for both inter-frame tracking and object identification.
Xi Li 0001, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Structure-Aware Slow Feature Analysis for Age Estimation
abstract
As an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples. However, most existing efforts to age estimation usually cope with the face samples individually, which do not take full advantage of the temporal structure and contextual structure of the face samples. In this letter, we propose an age estimation approach named structure-aware slow feature analysis, which is capable of effectively capturing the structure of human faces in the aspects of time-related smoothness for progressive age variation as well as face-related attribute constraints for face age consistency. As a result, we present an iterative optimization scheme to effectively learn the slowly varying feature transformation. Experimental results demonstrate the effectiveness of our approach on the Morph dataset.
Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Jun Xiao 0001
IEEE Signal Process. Lett.3
2016 Deep Learning Driven Visual Path Prediction From a Single Image
abstract
Capabilities of inference and prediction are the significant components of visual systems. Visual path prediction is an important and challenging task among them, with the goal to infer the future path of a visual object in a static scene. This task is complicated as it needs high-level semantic understandings of both the scenes and underlying motion patterns in video sequences. In practice, cluttered situations have also raised higher demands on the effectiveness and robustness of models. Motivated by these observations, we propose a deep learning framework, which simultaneously performs deep feature learning for visual representation in conjunction with spatiotemporal context modeling. After that, a unified path-planning scheme is proposed to make accurate path prediction based on the analytic results returned by the deep context models. The highly effective visual representation and deep context models ensure that our framework makes a deep semantic understanding of the scenes and motion patterns, consequently improving the performance on visual path prediction task. In experiments, we extensively evaluate the model's performance by constructing two large benchmark datasets from the adaptation of video tracking datasets. The qualitative and quantitative experimental results show that our approach outperforms the state-of-the-art approaches and owns a better generalization capability.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Zhouzhou He, Fei Wu 0001, Wei Liu 0005, Jinhui Tang 0001, Yueting Zhuang
IEEE Trans. Image Process.3
2016 Joint Multilabel Classification With Community-Aware Label Graph Learning
abstract
As an important and challenging problem in machine learning and computer vision, multilabel classification is typically implemented in a max-margin multilabel learning framework, where the inter-label separability is characterized by the sample-specific classification margins between labels. However, the conventional multilabel classification approaches are usually incapable of effectively exploring the intrinsic inter-label correlations as well as jointly modeling the interactions between inter-label correlations and multilabel classification. To address this issue, we propose a multilabel classification framework based on a joint learning approach called label graph learning (LGL) driven weighted Support Vector Machine (SVM). In principle, the joint learning approach explicitly models the inter-label correlations by LGL, which is jointly optimized with multilabel classification in a unified learning scheme. As a result, the learned label correlation graph well fits the multilabel classification task while effectively reflecting the underlying topological structures among labels. Moreover, the inter-label interactions are also influenced by label-specific sample communities (each community for the samples sharing a common label). Namely, if two labels have similar label-specific sample communities, they are likely to be correlated. Based on this observation, LGL is further regularized by the label Hypergraph Laplacian. Experimental results have demonstrated the effectiveness of our approach over several benchmark data sets.
Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Yueting Zhuang, Jingdong Wang 0001, Xuelong Li 0001
IEEE Trans. Image Process.3
2016 Learning of Multimodal Representations With Random Walks on the Click Graph
abstract
In multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. With the click data collected from the users' searching behavior, existing approaches take either one-to-one paired data (text-image pairs) or ranking examples (text-query-image and/or image-query-text ranking lists) as training examples, which do not make full use of the click data, particularly the implicit connections among the data objects. In this paper, we treat the click data as a large click graph, in which vertices are images/text queries and edges indicate the clicks between an image and a query. We consider learning a multimodal representation from the perspective of encoding the explicit/implicit relevance relationship between the vertices in the click graph. By minimizing both the truncated random walk loss as well as the distance between the learned representation of vertices and their corresponding deep neural network output, the proposed model which is named multimodal random walk neural network (MRW-NN) can be applied to not only learn robust representation of the existing multimodal data in the click graph, but also deal with the unseen queries and images to support cross-modal retrieval. We evaluate the latent representation learned by MRW-NN on a public large-scale click log data set Clickture and further show that MRW-NN achieves much better cross-modal retrieval performance on the unseen queries/images than the other state-of-the-art methods.
Fei Wu 0001, Jun Song 0004, Shuicheng Yan, Zhongfei Zhang, Yong Rui, Yueting Zhuang
IEEE Trans. Image Process.5
2016 Multimodal Data Mining in a Multimedia Database Based on Structured Max Margin Learning
abstract
Mining knowledge from a multimedia database has received increasing attentions recently since huge repositories are made available by the development of the Internet. In this article, we exploit the relations among different modalities in a multimedia database and present a framework for general multimodal data mining problem where image annotation and image retrieval are considered as the special cases. Specifically, the multimodal data mining problem can be formulated as a structured prediction problem where we learn the mapping from an input to the structured and interdependent output variables. In addition, in order to reduce the demanding computation, we propose a new max margin structure learning approach called Enhanced Max Margin Learning (EMML) framework, which is much more efficient with a much faster convergence rate than the existing max margin learning methods, as verified through empirical evaluations. Furthermore, we apply EMML framework to develop an effective and efficient solution to the multimodal data mining problem that is highly scalable in the sense that the query response time is independent of the database scale. The EMML framework allows an efficient multimodal data mining query in a very large scale multimedia database, and excels many existing multimodal data mining methods in the literature that do not scale up at all. The performance comparison with a state-of-the-art multimodal data mining method is reported for the real-world image databases.
Zhongfei Zhang, Eric P. Xing, Christos Faloutsos
ACM Trans. Knowl. Discov. Data2
2016 Bayesian Multi-Task Relationship Learning with Link Structure
abstract
In this paper, we study the multi-task learning problem with a new perspective of considering the link structure of data and task relationship modeling simultaneously. In particular, we first introduce the Matrix Gaussian (MG) distribution and Matrix Generalized Inverse Gaussian (MGIG) distribution, then define a Matrix Gaussian Matrix Generalized Inverse Gaussian (MG-MGIG) prior. Based on this prior, we propose a novel multi-task learning algorithm, the Bayesian Multi-task Relationship Learning (BMTRL) algorithm. To incorporate the link structure into the framework of BMTRL, we propose link constraints between samples. Through combining the BMTRL algorithm with the link constraints, we propose the Bayesian Multi-task Relationship Learning with Link Constraints (BMTRL-LC) algorithm. Further, we apply the manifold theory to provide an extension of BMTRL-LC to data with no link structure. Specifically, BMTRL-LC is effective for multi-task learning with only limited training samples, which is not addressed in the existing literature. To make the computation tractable, we simultaneously use a convex optimization method and sampling techniques. In particular, we adopt two stochastic EM algorithms for BMTRL and BMTRL-LC, respectively. The experimental results on three real datasets demonstrate the promise of the proposed algorithms.
Yingming Li, Ming Yang 0012, Zhongang Qi, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.4
2016 Scalable Linear Visual Feature Learning via Online Parallel Nonnegative Matrix Factorization
abstract
Visual feature learning, which aims to construct an effective feature representation for visual data, has a wide range of applications in computer vision. It is often posed as a problem of nonnegative matrix factorization (NMF), which constructs a linear representation for the data. Although NMF is typically parallelized for efficiency, traditional parallelization methods suffer from either an expensive computation or a high runtime memory usage. To alleviate this problem, we propose a parallel NMF method called alternating least square block decomposition (ALSD), which efficiently solves a set of conditionally independent optimization subproblems based on a highly parallelized fine-grained grid-based blockwise matrix decomposition. By assigning each block optimization subproblem to an individual computing node, ALSD can be effectively implemented in a MapReduce-based Hadoop framework. In order to cope with dynamically varying visual data, we further present an incremental version of ALSD, which is able to incrementally update the NMF solution with a low computational cost. Experimental results demonstrate the efficiency and scalability of the proposed methods as well as their applications to image clustering and image retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang, Chunhua Shen, Yueting Zhuang, Lixin Gao 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2015 Structured Embedding via Pairwise Relations and Long-Range Interactions in Knowledge Base
abstract
We consider the problem of embedding entities and relations of knowledge bases into low-dimensional continuous vector spaces (distributed representations). Unlike most existing approaches, which are primarily efficient for modelling pairwise relations between entities, we attempt to explicitly model both pairwise relations and long-range interactions between entities, by interpreting them as linear operators on the low-dimensional embeddings of the entities. Therefore, in this paper we introduces Path-Ranking to capture the long-range interactions of knowledge graph and at the same time preserve the pairwise relations of knowledge graph; we call it 'structured embedding via pairwise relation and long-range interactions' (referred to as SePLi). Comparing with the-state-of-the-art models, SePLi achieves better performances of embeddings.
Fei Wu 0001, Jun Song 0004, Yi Yang 0001, Xi Li 0001, Zhongfei Zhang, Yueting Zhuang
AAAI5
2015 Sensor network partitioning based on homogeneity
abstract
Discovering communities from sensor networks is an important problem in many real-world applications. The problem desiderata requires to partition the network according to the homogeneity of the sensor measurement and at the same time the partitioning result to be as independent as possible of the typical local change in the network topology. This poses new challenges to the current graph partitioning methodologies. We develop a graph partitioning approach from the perspective of homogeneity. In order to avoid the inherent incompatibility of the current graph partitioning methodologies, an objective functional is designed to be asymptotically close to the Mumford-Shah functional [1]. An variational algorithm ZERO-CUT and a fast approximation algorithm GREEDY-SACK are developed to solve the objective functional. We evaluate the performance of the proposed algorithms on both synthetic and real-world graphs in different applications to demonstrate its advantage and promise in solving the problem of discovering communities from sensor networks.
Zhongfei Zhang, Philip S. Yu
DSAA2
2015 Convex Approximation to the Integral Mixture Models Using Step Functions
abstract
The parameter estimation to mixture models has been shown as a local optimal solution for decades. In this paper, we propose a functional estimation to mixture models using step functions. We show that the proposed functional inference yields a convex formulation and consequently the mixture models are feasible for a global optimum inference. The proposed approach further unifies the existing isolated exemplar-based clustering techniques at a higher level of generality, e.g. it provides a theoretical justification for the heuristics of the clustering by affinity propagation Frey & Dueck (2007), it reproduces Lashkari & Golland (2007)'s's convex formulation as a special case under this step function construction. Empirical studies also verify the theoretic justifications.
Zhongfei Zhang, Philip S. Yu
ICDM3
2015 Sketch the Storyline with CHARCOAL: A Non-Parametric Approach
Siliang Tang, Fei Wu 0001, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang
IJCAI5
2015 Dropout Training of Matrix Factorization and Autoencoder for Link Prediction in Sparse Graphs
abstract
Matrix factorization (MF) and Autoencoder (AE) are among the most successful approaches of unsupervised learning. While MF based models have been extensively exploited in the graph modeling and link prediction literature, the AE family has not gained much attention. In this paper we investigate both MF and AE's application to the link prediction problem in sparse graphs. We show the connection between AE andMF from the perspective of multiview learning, and further propose MF+AE: a model training MF and AE jointly with shared parameters. We apply dropout to training both the MF and AE parts, and show that it can significantly prevent overfitting by acting as an adaptive regularization. We conduct experiments on six real world sparse graph datasets, and show that MF+AE consistently outperforms the competing methods, especially on datasets that demonstrate strong non-cohesive structures.
Shuangfei Zhai, Zhongfei Zhang
SDM2
2015 Tracking news article evolution by dense subgraph learning
Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001
Neurocomputing4
2015 Deep learning driven blockwise moving object detection with binary scene modeling
Xi Li 0001, Zhongfei Zhang, Fei Wu 0001
Neurocomputing3
2015 E-commerce business model mining and prediction
abstract
We study the problem of business model mining and prediction in the e-commerce context. Unlike most existing approaches where this is typically formulated as a regression problem or a time-series prediction problem, we take a different formulation to this problem by noting that these existing approaches fail to consider the potential relationships both among the consumers (consumer influence) and among the shops (competitions or collaborations). Taking this observation into consideration, we propose a new method for e-commerce business model mining and prediction, called EBMM, which combines regression with community analysis. The challenge is that the links in the network are typically not directly observed, which is addressed by applying information diffusion theory through the consumer-shop network. Extensive evaluations using Alibaba Group e-commerce data demonstrate the promise and superiority of EBMM to the state-of-the-art methods in terms of business model mining and prediction.
Zhouzhou He, Zhongfei Zhang, Chun-Ming Chen, Zheng-Gang Wang
Frontiers Inf. Technol. Electron. Eng.2
2015 Multimedia Retrieval via Deep Learning to Rank
abstract
Many existing learning-to-rank approaches are incapable of effectively modeling the intrinsic interaction relationships between the feature-level and ranking-level components of a ranking model. To address this problem, we propose a novel joint learning-to-rank approach called Deep Latent Structural SVM (DL-SSVM), which jointly learns deep neural networks and latent structural SVM (connected by a set of latent feature grouping variables) to effectively model the interaction relationships at two levels (i.e., feature-level and ranking-level). To make the joint learning problem easier to optimize, we present an effective auxiliary variable-based alternating optimization approach with respect to deep neural network learning and structural latent SVM learning. Experimental results on several challenging datasets have demonstrated the effectiveness of the proposed learning to rank approach in real-world information retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
IEEE Signal Process. Lett.3
2015 Weakly Semi-Supervised Deep Learning for Multi-Label Image Annotation
abstract
In this paper, we study leveraging both weakly labeled images and unlabeled images for multi-label image annotation. Motivated by the recent advance in deep learning, we propose an approach called weakly semi-supervised deep learning for multi-label image annotation (WeSed). In WeSed, a novel weakly weighted pairwise ranking loss is effectively utilized to handle weakly labeled images, while a triplet similarity loss is employed to harness unlabeled images. WeSed enables us to train deep convolutional neural network (CNN) with images from social networks where images are either only weakly labeled with several labels or unlabeled. We also design an efficient algorithm to sample high-quality image triplets from large image datasets to fine-tune the CNN. WeSed is evaluated on benchmark datasets for multi-label annotation. The experiments demonstrate the effectiveness of our proposed approach and show that the leverage of the weakly labeled images and unlabeled images leads to a significantly better performance.
Fei Wu 0001, Zhuhao Wang, Zhongfei Zhang, Yi Yang 0001, Jiebo Luo 0001, Wenwu Zhu 0001, Yueting Zhuang
IEEE Trans. Big Data3
2015 Cross-Modal Learning to Rank via Latent Joint Representation
abstract
Cross-modal ranking is a research topic that is imperative to many applications involving multimodal data. Discovering a joint representation for multimodal data and learning a ranking function are essential in order to boost the cross-media retrieval (i.e., image-query-text or text-query-image). In this paper, we propose an approach to discover the latent joint representation of pairs of multimodal data (e.g., pairs of an image query and a text document) via a conditional random field and structural learning in a listwise ranking manner. We call this approach cross-modal learning to rank via latent joint representation (CML²R). In CML²R, the correlations between multimodal data are captured in terms of their sharing hidden variables (e.g., topics), and a hidden-topic-driven discriminative ranking function is learned in a listwise ranking manner. The experiments show that the proposed approach achieves a good performance in cross-media retrieval and meanwhile has the capability to learn the discriminative representation of multimodal data.
Fei Wu 0001, Xinyang Jiang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Image Process.6
2015 Joint Structural Learning to Rank with Deep Linear Feature Learning
abstract
Multimedia information retrieval usually involves two key modules including effective feature representation and ranking model construction. Most existing approaches are incapable of well modeling the inherent correlations and interactions between them, resulting in the loss of the latent consensus structure information. To alleviate this problem, we propose a learning to rank approach that simultaneously obtains a set of deep linear features and constructs structure-aware ranking models in a joint learning framework. Specifically, the deep linear feature learning corresponds to a series of matrix factorization tasks in a hierarchical manner, while the learning-to-rank part concentrates on building a ranking model that effectively encodes the intrinsic ranking information by structural SVM learning. Through a joint learning mechanism, the two parts are mutually reinforced in our approach, and meanwhile their underlying interaction relationships are implicitly reflected by solving an alternating optimization problem. Due to the intrinsic correlations among different queries (i.e., similar queries for similar ranking lists), we further formulate the learning-to-rank problem as a multi-task problem, which is associated with a set of mutually related query-specific learning-to-rank subproblems. For computational efficiency and scalability, we design a MapReduce-based parallelization approach to speed up the learning processes. Experimental results demonstrate the efficiency, effectiveness, and scalability of the proposed approach in multimedia information retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.3
2015 Probabilistic Word Selection via Topic Modeling
abstract
We propose selective supervised Latent Dirichlet Allocation (ssLDA) to boost the prediction performance of the widely studied supervised probabilistic topic models. We introduce a Bernoulli distribution for each word in one given document to selectthis word as a strongly or weakly discriminative one with respect to its assigned topic. The Bernoulli distribution is parameterized by the discrimination power of the word for its assigned topic. As a result, the document is represented as a “bag-of-selective-words” instead of the probabilistic “bag-of-topics” in the topic modeling domain or the flat “bag-of-words” in the traditional natural language processing domain to form a new perspective. Inheriting the general framework of supervised LDA (sLDA), ssLDA can also predict many types of response specified by a Gaussian Linear Model (GLM). Focusing on the utilization of this word selection mechanism for singe-label document classification in this paper, we conduct the variational inference for approximating the intractable posterior and derive a maximum-likelihood estimation of parameters in ssLDA. The experiments reported on textual documents show that ssLDA not only performs competitively over “state-of-the-art” classification approaches based on both the flat “bag-of-words” and probabilistic “bag-of-topics” representation in terms of classification performance, but also has the ability to discover the discrimination power of the words specified in the topics (compatible with our rational knowledge).
Yueting Zhuang, Haidong Gao, Fei Wu 0001, Siliang Tang, Yin Zhang 0006, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.6
2014 Structural Bregman Distance Functions Learning to Rank with Self-Reinforcement
abstract
Learning to rank is an important task for many data mining applications. Essentially, the goal of learning to rank is to learn an appropriate similarity or distance metric to determine the relevance relationships among data points. However, most of the existing approaches for distance metric learning are limited in three aspects. First, they often assume a fixed form of distance metric for the entire input space. Second, the assumed distance functions are often computationally expensive or even intractable to learn for high dimensional data, such as Mahalanobis distance. Third, most of these approaches lack robustness to noisily labeled data, which is pervasive in many real-world applications. In this paper, we study learning to rank as a problem of distance metric learning to address the above three problems. We choose Bregman distance as the target distance function, due to its general functional form as a generalization of a wide class of distance functions, and its capacity of exploiting complicated nonlinear patterns underlying the data. Under the framework of structural SVM, we formulate the problem of learning Bregman distance functions for ranking as a QP problem by a nonparametric approach, and present an effective algorithm. Furthermore, we propose a self-reinforcement scheme that adaptively differentiates each data point in the role of learning to secure the robustness. We emphasize that the proposed method SBLR-S (Structural Bregman distance functions Learning to Rank with Self-reinforcement) is more general than the conventional distance metric learning approaches, and is able to handle high dimensional data as well as noisily labeled data. The experiments of data ranking on real-world datasets show the superiority of this method to the state-of-the-art literature.
Te Pi, Xi Li 0001, Zhongfei Zhang
ICDM3
2014 Distributed Binary Subspace Learning on large-scale cross media data
abstract
Due to the ubiquitous existence of large-scale data in today's real-world applications including learning on cross media data, we propose a semi-supervised learning method named Multiple Binary Subspace Regression (MBSR) for cross media data classification. In order to mine the common features among the data with multiple modalities, we project the original cross-media data into the same low-rank representation simultaneously by mapping to the corresponding subspaces for dimension reduction. All the subspaces are set to be binary, which only involve the addition operations and omit the multiplication operations in the subsequent computation owing to the good property of the binary values. The dimension reduction to a binary subspace and the classification on this subspace are also optimized simultaneously leading to a semi-supervised model. For dealing with large-scale data, our learning method is easily implemented to run in a MapReduce-based Hadoop system. Empirical studies demonstrate its competitive performance on convergence, efficiency, and scalability in comparison with the state-of-the-art literature.
Xueyi Zhao, Chenyi Zhang 0002, Zhongfei Zhang
ICME3
2014 Cross Domain Shared Subspace Learning for Unsupervised Transfer Classification
abstract
Transfer learning aims to address the problem where we lack the labeled data for training in one domain while utilizing the sufficient training data from other relevant domains. The problem becomes even more challenging when there are no labeled data in the target domain to build the association between two domains, which is more common in real-world scenarios. In this paper, we tackle with the challenge through learning the shared subspace across domains. The subspace is able to capture the intrinsic domain invariant innate characteristics for feature representations. Meanwhile in the learning procedure we train the classifiers in the source domain and predict the labels in the target domain simultaneously. We also incorporate the inherent data structure in the predicted labels to enhance the robustness against the misclassification. Extensive experimental evaluations on the public datasets demonstrate the effectiveness and promise of our method compared with the state-of-the-art transfer learning methods.
Zheng Fang 0013, Zhongfei Zhang
ICPR2
2014 Scalable Nonnegative Matrix Factorization with Block-wise Updates
Jiangtao Yin, Lixin Gao 0001, Zhongfei Zhang
ECML/PKDD (3)3
2014 Scientific articles recommendation with topic regression and relational matrix factorization
abstract
In this paper we study the problem of recommending scientific articles to users in an online community with a new perspective of considering topic regression modeling and articles relational structure analysis simultaneously. First, we present a novel topic regression model, the topic regression matrix factorization (tr-MF), to solve the problem. The main idea of tr-MF lies in extending the matrix factorization with a probabilistic topic modeling. In particular, tr-MF introduces a regression model to regularize user factors through probabilistic topic modeling under the basic hypothesis that users share similar preferences if they rate similar sets of items. Consequently, tr-MF provides interpretable latent factors for users and items, and makes accurate predictions for community users. To incorporate the relational structure into the framework of tr-MF, we introduce relational matrix factorization. Through combining tr-MF with the relational matrix factorization, we propose the topic regression collective matrix factorization (tr-CMF) model. In addition, we also present the collaborative topic regression model with relational matrix factorization (CTR-RMF) model, which combines the existing collaborative topic regression (CTR) model and relational matrix factorization (RMF). From this point of view, CTR-RMF can be considered as an appropriate baseline for tr-CMF. Further, we demonstrate the efficacy of the proposed models on a large subset of the data from CiteULike, a bibliography sharing service dataset. The proposed models outperform the state-of-the-art matrix factorization models with a significant margin. Specifically, the proposed models are effective in making predictions for users with only few ratings or even no ratings, and support tasks that are specific to a certain field, neither of which has been addressed in the existing literature.
Ming Yang 0012, Yingming Li, Zhongfei Zhang
J. Zhejiang Univ. Sci. C3
2014 A Regularized Approach for Geodesic-Based Semisupervised Multimanifold Learning
abstract
Geodesic distance, as an essential measurement for data dissimilarity, has been successfully used in manifold learning. However, most geodesic distance-based manifold learning algorithms have two limitations when applied to classification: 1) class information is rarely used in computing the geodesic distances between data points on manifolds and 2) little attention has been paid to building an explicit dimension reduction mapping for extracting the discriminative information hidden in the geodesic distances. In this paper, we regard geodesic distance as a kind of kernel, which maps data from linearly inseparable space to linear separable distance space. In doing this, a new semisupervised manifold learning algorithm, namely regularized geodesic feature learning algorithm, is proposed. The method consists of three techniques: a semisupervised graph construction method, replacement of original data points with feature vectors which are built by geodesic distances, and a new semisupervised dimension reduction method for feature vectors. Experiments on the MNIST, USPS handwritten digit data sets, MIT CBCL face versus nonface data set, and an intelligent traffic data set show the effectiveness of the proposed algorithm.
Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao
IEEE Trans. Image Process.4
2014 A Two-Level Topic Model Towards Knowledge Discovery from Citation Networks
abstract
Knowledge discovery from scientific articles has received increasing attention recently since huge repositories are made available by the development of the Internet and digital databases. In a corpus of scientific articles such as a digital library, documents are connected by citations and one document plays two different roles in the corpus: document itself and a citation of other documents. In the existing topic models, little effort is made to differentiate these two roles. We believe that the topic distributions of these two roles are different and related in a certain way. In this paper, we propose a Bernoulli process topic (BPT) model which considers the corpus at two levels: document level and citation level. In the BPT model, each document has two different representations in the latent topic space associated with its roles. Moreover, the multi-level hierarchical structure of citation network is captured by a generative process involving a Bernoulli process. The distribution parameters of the BPT model are estimated by a variational approximation approach. An efficient computation algorithm is proposed to overcome the difficulty of matrix inverse operation. In addition to conducting the experimental evaluations on the document modeling and document clustering tasks, we also apply the BPT model to well known corpora to discover the latent topics, recommend important citations, detect the trends of various research areas in computer science between 1991 and 1998, and to investigate the interactions among the research areas. The comparisons against state-of-the-art methods demonstrate a very promising performance. The implementations and the data sets are available online .
Zhongfei Zhang, Shenghuo Zhu, Yun Chi, Yihong Gong
IEEE Trans. Knowl. Data Eng.2
2014 Context-Aware Hypergraph Construction for Robust Spectral Clustering
abstract
Spectral clustering is a powerful tool for unsupervised data analysis. In this paper, we propose a context-aware hypergraph similarity measure (CAHSM), which leads to robust spectral clustering in the case of noisy data. We construct three types of hypergraphs-the pairwise hypergraph, the k-nearest-neighbor (kNN) hypergraph, and the high-order over-clustering hypergraph. The pairwise hypergraph captures the pairwise similarity of data points; the kNNhypergraph captures the neighborhood of each point; and the clustering hypergraph encodes high-order contexts within the dataset. By combining the affinity information from these three hypergraphs, the CAHSM algorithm is able to explore the intrinsic topological information of the dataset. Therefore, data clustering using CAHSM tends to be more robust. Considering the intra-cluster compactness and the inter-cluster separability of vertices, we further design a discriminative hypergraph partitioning criterion (DHPC). Using both CAHSM and DHPC, a robust spectral clustering algorithm is developed. Theoretical analysis and experimental evaluation demonstrate the effectiveness and robustness of the proposed algorithm.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.5
2013 Discriminative feature selection for multi-view cross-domain learning
abstract
In many data mining applications, we often face the problem of cross-domain learning, i.e., to transfer the already learned knowledge from a source domain to a target domain. In particular, this problem becomes very challenging when there is no or little labeled training data available in the target domain, which is not an uncommon scenario as it is expensive and in certain cases even impossible to obtain any labeled training data in the target domain in many real world applications. In the literature, though few efforts are reported to attempt to solve this challenging problem, the solutions are all rather limited making this problem still open and challenging. On the other hand, as it is not uncommon to face this problem in many applications, an effective solution to this problem shall generate substantial societal impacts. In this paper, we address this problem and propose a new framework, called DISMUTE, taking advantage of the typically available multiple views of the data in domains. Consequently, DISMUTE is based on discriminative feature selection for multi-view cross-domain learning. Theoretic analysis and extensive evaluations in the specific application of object identification and image classification against several state-of-the-art methods demonstrate the outstanding superiority of DISMUTE.
Zheng Fang 0013, Zhongfei Zhang
CIKM2
2013 Scientific articles recommendation
abstract
We study the problem of recommending scientific articles to users in an online community and present a novel matrix factorization model, the topic regression Matrix Factorization (tr-MF), to solve the problem. The main idea of tr-MF lies in extending the matrix factorization with a probabilistic topic modeling. Instead of regularizing item factors through the probabilistic topic modeling as in the framework of the CTR model, tr-MF introduces a regression model to regularize user factors through the probabilistic topic modeling under the basic hypothesis that users share the similar preferences if they rate similar sets of items. Consequently, tr-MF provides interpretable latent factors for users and items, and makes accurate predictions for community users. Specifically, it is effective in making predictions for users with only few ratings or even no ratings, and supports tasks that are specific to a certain field, neither of which is addressed in the existing literature. Further, we demonstrate the efficacy of tr-MF on a large subset of the data from CiteULike, a bibliography sharing service dataset. The proposed model outperforms the state-of-the-art matrix factorization models with a significant margin.
Yingming Li, Ming Yang 0012, Zhongfei Zhang
CIKM3
2013 Bayesian Multi-Task Relationship Learning with Link Structure
abstract
In this paper, we study the multi-task learning problem with a new perspective of considering the link structure of data and task relationship modeling simultaneously. In particular, we first introduce the Matrix Generalized Inverse Gaussian (MGIG) distribution and define a Matrix Gaussian Matrix Generalized Inverse Gaussian (MG-MGIG) prior. Based on this prior, we propose a novel multi-task learning algorithm, the Bayesian Multi-task Relationship Learning (BMTRL) algorithm. To incorporate the link structure into the framework of BMTRL, we propose link constraints between samples. Through combining the BMTRL algorithm with the link constraints, we propose the Bayesian Multi-task Relationship Learning with Link Constraints (BMTRL-LC) algorithm. To make the computation tractable, we simultaneously use a convex optimization method and sampling techniques. In particular, we adopt two stochastic EM algorithms for BMTRL and BMTRL-LC, respectively. The experimental results on Cora dataset demonstrate the promise of the proposed algorithms.
Yingming Li, Ming Yang 0012, Zhongang Qi, Zhongfei Zhang
ICDM4
2013 Multi-Task Learning with Gaussian Matrix Generalized Inverse Gaussian Model
abstract
In this paper, we study the multi-task learning problem with a new perspective of considering the structure of the residue error matrix and the low-rank approximation to the task covariance matrix simultaneously. In particular, we first introduce the Matrix Generalized Inverse Gaussian (MGIG) prior and define a Gaussian Matrix Generalized Inverse Gaussian (GMGIG) model for low-rank approximation to the task covariance matrix. Through combining the GMGIG model with the residual error structure assumption, we propose the GMGIG regression model for multi-task learning. To make the computation tractable, we simultaneously use variational inference and sampling techniques. In particular, we propose two sampling strategies for computing the statistics of the MGIG distribution. Experiments show that this model is superior to the peer methods in regression and prediction.
Ming Yang 0012, Yingming Li, Zhongfei Zhang
ICML (3)3
2013 Learning with limited and noisy tagging
abstract
With the rapid development of social networks, tagging has become an important means responsible for such rapid development. A robust tagging method must have the capability to meet the two challenging requirements: limited labeled training samples and noisy labeled training samples. In this paper, we investigate this challenging problem of learning with limited and noisy tagging and propose a discriminative model, called SpSVM-MC, that exploits both labeled and unlabeled data through a semi-parametric regularization and takes advantage of the multi-label constraints into the optimization. While SpSVM-MC is a general method for learning with limited and noisy tagging, in the evaluations we focus on the specific application of noisy image tagging with limited labeled training samples on a benchmark dataset. Theoretical analysis and extensive evaluations in comparison with state-of-the-art literature demonstrate that SpSVM-MC outstands with a superior performance.
Yingming Li, Zhongang Qi, Zhongfei Zhang, Ming Yang 0012
ACM Multimedia3
2013 Cross-media semantic representation via bi-directional learning to rank
abstract
In multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. Existing approaches take either one-to-one paired data or uni-directional ranking examples (i.e., utilizing only text-query-image ranking examples or image-query-text ranking examples) as training examples, which do not make full use of bi-directional ranking examples (bi-directional ranking means that both text-query-image and image-query-text ranking examples are utilized in the training period) to achieve a better performance. In this paper, we consider learning a cross-media representation model from the perspective of optimizing a listwise ranking problem while taking advantage of bi-directional ranking examples. We propose a general cross-media ranking algorithm to optimize the bi-directional listwise ranking loss with a latent space embedding, which we call Bi-directional Cross-Media Semantic Representation Model (Bi-CMSRM). The latent space embedding is discriminatively learned by the structural large margin learning for optimization with certain ranking criteria (mean average precision in this paper) directly. We evaluate Bi-CMSRM on the Wikipedia and NUS-WIDE datasets and show that the utilization of the bi-directional ranking examples achieves a much better performance than only using the uni-directional ranking examples.
Fei Wu 0001, Zhongfei Zhang, Shuicheng Yan, Yong Rui, Yueting Zhuang
ACM Multimedia3
2013 Discriminative Transfer Learning on Manifold
abstract
Collective matrix factorization has achieved a remarkable success in document classification in the literature of transfer learning. However, the learned latent factors still suffer from the divergence between different domains and thus are usually not discriminative for an appropriate assignment of category labels. Based on these observations, we impose a discriminative regression model over the latent factors to enhance the capability of label prediction. Moreover, we propose to minimize the Maximum Mean Discrepancy in the latent manifold subspace, as opposed to typically in the original data space, to bridge the gap between different domains. Specifically, we formulate these objectives into a joint optimization framework with two matrix tri-factorizations for the source and target domains simultaneously. An iterative algorithm DTLM is developed and the theoretical analysis of its convergence is discussed. Empirical study on benchmark datasets validates that DTLM improves the classification accuracy consistently compared with the state-of-the-art transfer learning methods.
Zheng Fang 0013, Zhongfei Zhang
SDM2
2013 A low rank structural large margin method for cross-modal ranking
abstract
Cross-modal retrieval is a classic research topic in multimedia information retrieval. The traditional approaches study the problem as a pairwise similarity function problem. In this paper, we consider this problem from a new perspective as a listwise ranking problem and propose a general cross-modal ranking algorithm to optimize the listwise ranking loss with a low rank embedding, which we call Latent Semantic Cross-Modal Ranking (LSCMR). The latent low-rank embedding space is discriminatively learned by structural large margin learning to optimize for certain ranking criteria directly. We evaluate LSCMR on the Wikipedia and NUS-WIDE dataset. Experimental results show that this method obtains significant improvements over the state-of-the-art methods.
Fei Wu 0001, Siliang Tang, Zhongfei Zhang, Xiaofei He 0001, Yueting Zhuang
SIGIR4
2013 An Incremental DPMM-Based Method for Trajectory Clustering, Modeling, and Retrieval
abstract
Trajectory analysis is the basis for many applications, such as indexing of motion events in videos, activity recognition, and surveillance. In this paper, the Dirichlet process mixture model (DPMM) is applied to trajectory clustering, modeling, and retrieval. We propose an incremental version of a DPMM-based clustering algorithm and apply it to cluster trajectories. An appropriate number of trajectory clusters is determined automatically. When trajectories belonging to new clusters arrive, the new clusters can be identified online and added to the model without any retraining using the previous data. A time-sensitive Dirichlet process mixture model (tDPMM) is applied to each trajectory cluster for learning the trajectory pattern which represents the time-series characteristics of the trajectories in the cluster. Then, a parameterized index is constructed for each cluster. A novel likelihood estimation algorithm for the tDPMM is proposed, and a trajectory-based video retrieval model is developed. The tDPMM-based probabilistic matching method and the DPMM-based model growing method are combined to make the retrieval model scalable and adaptable. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our algorithm.
Weiming Hu 0004, Xi Li 0001, Guodong Tian, Stephen J. Maybank, Zhongfei Zhang
IEEE Trans. Pattern Anal. Mach. Intell.5
2013 Visual Tracking With Spatio-Temporal Dempster-Shafer Information Fusion
abstract
A key problem in visual tracking is how to effectively combine spatio-temporal visual information from throughout a video to accurately estimate the state of an object. We address this problem by incorporating Dempster-Shafer (DS) information fusion into the tracking approach. To implement this fusion task, the entire image sequence is partitioned into spatially and temporally adjacent subsequences. A support vector machine (SVM) classifier is trained for object/nonobject classification on each of these subsequences, the outputs of which act as separate data sources. To combine the discriminative information from these classifiers, we further present a spatio-temporal weighted DS (STWDS) scheme. In addition, temporally adjacent sources are likely to share discriminative information on object/nonobject classification. To use such information, an adaptive SVM learning scheme is designed to transfer discriminative information across sources. Finally, the corresponding DS belief function of the STWDS scheme is embedded into a Bayesian tracking model. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracking approach.
Xi Li 0001, Anthony R. Dick, Chunhua Shen, Zhongfei Zhang, Anton van den Hengel, Hanzi Wang
IEEE Trans. Image Process.4
2013 A survey of appearance models in visual object tracking
abstract
Visual object tracking is a significant computer vision task which can be applied to many domains, such as visual surveillance, human computer interaction, and video compression. Despite extensive research on this topic, it still suffers from difficulties in handling complex object appearance changes caused by factors such as illumination variation, partial occlusion, shape deformation, and camera motion. Therefore, effective modeling of the 2D appearance of tracked objects is a key issue for the success of a visual tracker. In the literature, researchers have proposed a variety of 2D appearance models. To help readers swiftly learn the recent advances in 2D appearance models for visual object tracking, we contribute this survey, which provides a detailed review of the existing 2D appearance models. In particular, this survey takes a module-based architecture that enables readers to easily grasp the key points of visual object tracking. In this survey, we first decompose the problem of appearance modeling into two different processing stages: visual representation and statistical modeling. Then, different 2D appearance models are categorized and discussed with respect to their composition modules. Finally, we address several issues of interest as well as the remaining challenges for future research on this topic. The contributions of this survey are fourfold. First, we review the literature of visual representations according to their feature-construction mechanisms (i.e., local and global). Second, the existing statistical modeling schemes for tracking-by-detection are reviewed according to their model-construction mechanisms: generative, discriminative, and hybrid generative-discriminative. Third, each type of visual representations or statistical modeling techniques is analyzed and discussed from a theoretical or practical viewpoint. Fourth, the existing benchmark resources (e.g., source codes and video datasets) are examined in this survey.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Zhongfei Zhang, Anthony R. Dick, Anton van den Hengel
ACM Trans. Intell. Syst. Technol.4
2012 Mining noisy tagging from multi-label space
abstract
In this paper we study the problem of mining noisy tagging. Most of the existing discriminative classification methods to this problem only consider one tag at a time as the classification target, and completely ignore the rest of the given tags at the same time. In this paper we argue that all the given multiple tags can be utilized simultaneously as an additional feature and the information contained in the multi-label space can be taken advantage of to improve the performance of the classification. We first propose a novel distance measure to compute the distance between instances in the multi-label space. Then we propose several novel methods to incorporate the information of the multi-label space into the discriminative classification methods in one view learning or in two views learning to solve a general multi-label classification problem and to mitigate the influence of the noise in the classification. We apply the proposed solutions to the problem with a more specific context - noisy image annotation, and evaluate the proposed methods on a standard dataset from the related literature. Experiments show that they are superior to the peer methods in the existing literature on solving the problem of mining noisy tagging.
Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang
CIKM3
2012 Geodesic Based Semi-supervised Multi-manifold Feature Extraction
abstract
Manifold learning is an important feature extraction approach in data mining. This paper presents a new semi-supervised manifold learning algorithm, called Multi-Manifold Discriminative Analysis (Multi-MDA). The proposed method is designed to explore the discriminative information hidden in geodesic distances. The main contributions of the proposed method are: 1) we propose a semi-supervised graph construction method which can effectively capture the multiple manifolds structure of the data, 2) each data point is replaced with an associated feature vector whose elements are the graph distances from it to the other data points. Information of the nonlinear structure is contained in the feature vectors which are helpful for classification, 3) we propose a new semi-supervised linear dimension reduction method for feature vectors which introduces the class information into the manifold learning process and establishes an explicit dimension reduction mapping. Experiments on benchmark data sets are conducted to show the effectiveness of the proposed method.
Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao
ICDM4
2012 Simultaneously Combining Multi-view Multi-label Learning with Maximum Margin Classification
abstract
Multiple feature views arise in various important data classification scenarios. However, finding a consensus feature view from multiple feature views for a classifier is still a challenging task. We present a new classification framework using the multi-label correlation information to address the problem of simultaneously combining multiple feature views and maximum margin classification. Under this framework, we propose a novel algorithm that iteratively computes the multiple view feature mapping matrices, the consensus feature view representation, and the coefficients of the classifier. Extensive experimental evaluations demonstrate the effectiveness and promise of this framework as well as the algorithm for discovering a consensus view from multiple feature views.
Zheng Fang 0013, Zhongfei Zhang
ICDM2
2012 Multi-view learning from imperfect tagging
abstract
In many real-world applications, tagging is imperfect: incomplete, inconsistent, and error-prone. Solutions to this problem will generate societal and technical impacts. In this paper, we investigate this arguably new problem: learning from imperfect tagging. We propose a general and effective learning scheme called the Multi-view Imperfect Tagging Learning (MITL) to this problem. The main idea of MITL lies in extracting the information of the imperfectly tagged training dataset from multiple views to differentiate the data points in the role of classification. Further, a novel discriminative classification method is proposed under the framework of MITL, which explicitly makes use of the given multiple labels simultaneously as an additional feature to deliver a more effective classification performance than the existing literature where one label is considered at a time as the classification target while the rest of the given labels are completely ignored at the same time. The proposed methods can not only complete the incomplete tagging but also denoise the noisy tagging through an inductive learning. We apply the general solution to the problem with a more specific context - imperfect image annotation, and evaluate the proposed methods on a standard dataset from the related literature. Experiments show that they are superior to the peer methods on solving the problem of learning from imperfect tagging in cross-media.
Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang
ACM Multimedia3
2012 Unsupervised Ensemble Learning for Mining Top-n Outliers
Weiming Hu 0004, Zhongfei Zhang, Ou Wu 0001
PAKDD (1)3
2012 Overlapping community detection combining content and link
abstract
In classic community detection, it is assumed that communities are exclusive, in the sense of either soft clustering or hard clustering. It has come to attention in the recent literature that many real-world problems violate this assumption, and thus overlapping community detection has become a hot research topic. The existing work on this topic uses either content or link information, but not both of them. In this paper, we deal with the issue of overlapping community detection by combining content and link information. We develop an effective solution called subgraph overlapping clustering (SOC) and evaluate this new approach in comparison with several peer methods in the literature that use either content or link information. The evaluations demonstrate the effectiveness and promise of SOC in dealing with large scale real datasets.
Zhouzhou He, Zhongfei Zhang, Philip S. Yu
J. Zhejiang Univ. Sci. C2
2012 Societally connected multimedia across cultures
abstract
The advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity.
Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu
J. Zhejiang Univ. Sci. C1
2012 Single and Multiple Object Tracking Using Log-Euclidean Riemannian Subspace and Block-Division Appearance Model
abstract
Object appearance modeling is crucial for tracking objects, especially in videos captured by nonstationary cameras and for reasoning about occlusions between multiple moving objects. Based on the log-euclidean Riemannian metric on symmetric positive definite matrices, we propose an incremental log-euclidean Riemannian subspace learning algorithm in which covariance matrices of image features are mapped into a vector space with the log-euclidean Riemannian metric. Based on the subspace learning algorithm, we develop a log-euclidean block-division appearance model which captures both the global and local spatial layout information about object appearances. Single object tracking and multi-object tracking with occlusion reasoning are then achieved by particle filtering-based Bayesian state inference. During tracking, incremental updating of the log-euclidean block-division appearance model captures changes in object appearance. For multi-object tracking, the appearance models of the objects can be updated even in the presence of occlusions. Experimental results demonstrate that the proposed tracking algorithm obtains more accurate results than six state-of-the-art tracking algorithms.
Weiming Hu 0004, Xi Li 0001, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank, Zhongfei Zhang
IEEE Trans. Pattern Anal. Mach. Intell.6
2012 Generative Models for Evolutionary Clustering
abstract
This article studies evolutionary clustering, a recently emerged hot topic with many important applications, noticeably in dynamic social network analysis. In this article, based on the recent literature on nonparametric Bayesian models, we have developed two generative models: DPChain and HDP-HTM. DPChain is derived from the Dirichlet process mixture (DPM) model, with an exponential decaying component along with the time. HDP-HTM combines the hierarchical dirichlet process (HDP) with a hierarchical transition matrix (HTM) based on the proposed Infinite hierarchical Markov state model (iHMS). Both models substantially advance the literature on evolutionary clustering, in the sense that not only do they both perform better than those in the existing literature, but more importantly, they are capable of automatically learning the cluster numbers and explicitly addressing the corresponding issues. Extensive evaluations have demonstrated the effectiveness and the promise of these two solutions compared to the state-of-the-art literature.
Tianbing Xu, Zhongfei Zhang, Philip S. Yu, Bo Long
ACM Trans. Knowl. Discov. Data2
2011 Pattern change discovery between high dimensional data sets
abstract
This paper investigates the general problem of pattern change discovery between high-dimensional data sets. Current methods either mainly focus on magnitude change detection of low-dimensional data sets or are under supervised frameworks. In this paper, the notion of the principal angles between the subspaces is introduced to measure the subspace difference between two high-dimensional data sets. Principal angles bear a property to isolate subspace change from the magnitude change. To address the challenge of directly computing the principal angles, we elect to use matrix factorization to serve as a statistical framework and develop the principle of the dominant subspace mapping to transfer the principal angle based detection to a matrix factorization problem. We show how matrix factorization can be naturally embedded into the likelihood ratio test based on the linear models. The proposed method is of an unsupervised nature and addresses the statistical significance of the pattern changes between high-dimensional data sets. We have showcased the different applications of this solution in several specific real-world applications to demonstrate the power and effectiveness of this method.
Zhongfei Zhang, Philip S. Yu, Bo Long
CIKM2
2011 Mining partially annotated images
abstract
In this paper, we study the problem of mining partially annotated images. We first define what the problem of mining partially annotated images is, and argue that in many real-world applications annotated images are typically partially annotated and thus that the problem of mining partially annotated images exists in many situations. We then propose an effective solution to this problem based on a statistical model we have developed called the Semi-Supervised Correspondence Hierarchical Dirichlet Process (SSCHDP). The main idea of this model lies in exploiting the information pertaining to partially annotated images or even unannotated images to achieve semi-supervised learning under the HDP structure. We apply this model to completing the annotations appropriately for partially annotated images in the training data and then to predicting the annotations appropriately and completely for all the unannotated images either in the training data or in any unseen data beyond the training process. Experiments show that SSC-HDP is superior to the peer models from the recent literature when they are applied to solving the problem of mining partially annotated images.
Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang
KDD3
2011 RKOF: Robust Kernel-Based Local Outlier Detection
Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Ou Wu 0001
PAKDD (2)3
2011 Incremental Tensor Subspace Learning and Its Applications to Foreground Segmentation and Tracking
abstract
Appearance modeling is very important for background modeling and object tracking. Subspace learning-based algorithms have been used to model the appearances of objects or scenes. Current vector subspace-based algorithms cannot effectively represent spatial correlations between pixel values. Current tensor subspace-based algorithms construct an offline representation of image ensembles, and current online tensor subspace learning algorithms cannot be applied to background modeling and object tracking. In this paper, we propose an online tensor subspace learning algorithm which models appearance changes by incrementally learning a tensor subspace representation through adaptively updating the sample mean and an eigenbasis for each unfolding matrix of the tensor. The proposed incremental tensor subspace learning algorithm is applied to foreground segmentation and object tracking for grayscale and color image sequences. The new background models capture the intrinsic spatiotemporal characteristics of scenes. The new tracking algorithm captures the appearance characteristics of an object during tracking and uses a particle filter to estimate the optimal object state. Experimental evaluations against state-of-the-art algorithms demonstrate the promise and effectiveness of the proposed incremental tensor subspace learning algorithm, and its applications to foreground segmentation and object tracking.
Weiming Hu 0004, Xi Li 0001, Xiaoqin Zhang 0002, Xinchu Shi, Stephen J. Maybank, Zhongfei Zhang
Int. J. Comput. Vis.6
2011 Recognition of adult images, videos, and web page bags
abstract
In this article, we develop an integrated adult-content recognition system which can detect adult images, adult videos, and adult Web page bags, where a Web page bag consists of a Web page and a predefined number of Web pages linked to it through hyperlinks. In our adult image-recognition algorithm, we model skin patches rather than skin pixels, resulting in better results than state-of-the-art algorithms which model skin pixels. In our adult video-recognition algorithm, information from the accompanying audio section around an image in an adult video is used to obtain a prior classification of the image. The algorithm achieves a better performance than the ones which use image information alone or audio information alone. The adult Web page bag recognition is carried out using multi-instance learning based on the combination of classifying texts, images and videos in Web pages. Both the speed and the accuracy for recognizing the Web adult content are increased, in contrast to recognizing Web pages one-by-one.
Weiming Hu 0004, Haiqiang Zuo, Ou Wu 0001, Yunfei Chen 0002, Zhongfei Zhang, David Suter
ACM Trans. Multim. Comput. Commun. Appl.5
2010 A Topic Model for Linked Documents and Update Rules for its Estimation
abstract
The latent topic model plays an important role in the unsupervised learning from a corpus, which provides a probabilistic interpretation of the corpus in terms of the latent topic space. An underpinning assumption which most of the topic models are based on is that the documents are assumed to be independent of each other. However, this assumption does not hold true in reality and the relations among the documents are available in different ways, such as the citation relations among the research papers. To address this limitation, in this paper we present a Bernoulli Process Topic (BPT) model, where the interdependence among the documents is modeled by a random Bernoulli process. In the BPT model a document is modeled as a distribution over topics that is a mixture of the distributions associated with the related documents. Although BPT aims at obtaining a better document modeling by incorporating the relations among the documents, it could also be applied to many applications including detecting the topics from corpora and clustering the documents. We apply the BPT model to several document collections and the experimental comparisons against several state-of-the-art approaches demonstrate the promising performance.
Shenghuo Zhu, Zhongfei Zhang, Yun Chi, Yihong Gong
AAAI3
2010 Local Outlier Detection Based on Kernel Regression
abstract
Outlier detection keeps an important and attractive task of the knowledge discovery in databases. In this paper, a novel approach named Multi-scale Local Kernel Regression is proposed. It transfers the unsupervised learning of outlier detection to the classic non-parameter regression learning. Through preprocessing the original data by the basic local density-based method, it adopts the local kernel regression estimator in the multiple scale neighborhoods to determine outliers. Experiments on several real life data sets demonstrate that this approach is promising in detection performance.
Weiming Hu 0004, Wei Li 0034, Zhongfei Zhang, Ou Wu 0001
ICPR4
2010 Unsupervised Learning from Linked Documents
abstract
Documents in many corpora, such as digital libraries and webpages, contain both content and link information. In a traditional topic model which plays an important role in the unsupervised learning, the link information is either totally ignored or treated as a feature similar to content. We believe that neither approach is capable of accurately capturing the relations represented by links. To address the limitation of traditional topic models, in this paper we propose a citation-topic (CT) model that explicitly considers the document relations represented by links. In the CT model, instead of being treated as yet another feature, links are used to form the structure of the generative model. As a result, in the CT model a given document is modeled as a mixture of a set of topic distributions, each of which is borrowed (cited) from a document that is related to the given document. We apply the CT model to several document collections and the experimental comparisons against state-of-the-art approaches demonstrate very promising performances.
Shenghuo Zhu, Yun Chi, Zhongfei Zhang, Yihong Gong
ICPR4
2010 Overview of ACM international workshop on connected multimedia
abstract
Following the very first international workshop on connected multimedia held in Hangzhou, China, in October of 2009 jointly sponsored by US National Science Foundation and Zhejiang University of China, this is the very first ACM International Workshop on Connected Multimedia in conjunction with ACM International Conference on Multimedia held in Florence, Italy, in October of 2010. In this workshop overview, we first define what we mean by connected multimedia, and then briefly overview the program of this workshop.
Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang
ACM Multimedia1
2010 Linear discriminant analysis using rotational invariant L1 norm
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang
Neurocomputing4
2010 Robust object tracking using a spatial pyramid heat kernel structural information representation
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang
Neurocomputing4
2010 A general framework for relation graph clustering
Bo Long, Zhongfei Zhang, Philip S. Yu
Knowl. Inf. Syst.2
2010 Heat Kernel Based Local Binary Pattern for Face Representation
abstract
Face classification has recently become a very hot research topic in computer vision and multimedia information processing. It has many potential applications, in which face representation is the most fundamental task. Most existing face representation methods perform poorly in capturing the intrinsic structural information of face appearance. To address this problem, we propose a novel multiscale heat kernel based face representation, for heat kernels perform well in characterizing the topological structural information of face appearance. Further, the local binary pattern (LBP) descriptor is incorporated into the multiscale heat kernel face representation for the purpose of capturing texture information of face appearance. As a result, we have the heat kernel based local binary pattern (HKLBP) descriptor. Finally, a Support Vector Machine (SVM) classifier is learned in theHKLBPfeature space for face classification. Experimental results demonstrate the effectiveness and superiority of our face classification framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Hanzi Wang
IEEE Signal Process. Lett.3
2009 Spectral Graph Partitioning Based on a Random Walk Diffusion Similarity Measure
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Yang Liu 0020
ACCV (2)3
2009 Knowledge Discovery from Citation Networks
abstract
Knowledge discovery from scientific articles has received increasing attentions recently since huge repositories are made available by the development of the Internet and digital databases. In a corpus of scientific articles such as a digital library, documents are connected by citations and one document plays two different roles in the corpus: \emph{document itself} and \emph{a citation of other documents}. In the existing topic models, little effort is made to differentiate these two roles. We believe that the topic distributions of these two roles are different and related in a certain way. In this paper we propose a \emph{Bernoulli Process Topic}~(BPT) model which models the corpus at two levels: \emph{document level} and \emph{citation level}. In the BPT model, each document has two different representations in the latent topic space associated with its roles. Moreover, the multi-level hierarchical structure of the citation network is captured by a generative process involving a Bernoulli process. The distribution parameters of the BPT model are estimated by a variational approximation approach. In addition to conducting the experimental evaluations on the document modeling task, we also apply the BPT model to a well known scientific corpus to discover the latent topics. The comparisons against state-of-the-art methods demonstrate a very promising performance.
Zhongfei Zhang, Shenghuo Zhu, Yun Chi, Yihong Gong
ICDM2
2009 Towards developing a unified multimodal image retrieval framework
abstract
Built upon the previous work on automatic image annotation and multimodal image retrieval, in this paper we present a unified multimodal image retrieval framework, called UPMIR. The contributions of this paper include: (1) the development of the UPMIR framework; and (2) extensive evaluations of UPMIR including evaluations against a state-of-the-art image retrieval system in a large scale, visually and semantically diverse database crawled from the Web to demonstrate that UPMIR not only has more effective and efficient retrieval performance, but also facilitates more enhanced retrieval modalities.
Zhongfei Zhang, Ruofei Zhang
ICME1
2009 A latent topic model for linked documents
abstract
Documents in many corpora, such as digital libraries and webpages, contain both content and link information. To explicitly consider the document relations represented by links, in this paper we propose a citation-topic (CT) model which assumes a probabilistic generative process for corpora. In the CT model a given document is modeled as a mixture of a set of topic distributions, each of which is borrowed (cited) from a document that is related to the given document. Moreover, the CT model contains a random process for selecting the related documents according to the structure of the generative model determined by links and therefore, the transitivity of the relations among documents is captured. We apply the CT model on the document clustering task and the experimental comparisons against several state-of-the-art approaches demonstrate very promising performances.
Shenghuo Zhu, Yun Chi, Zhongfei Zhang, Yihong Gong
SIGIR4
2008 Clustering on Complex Graphs
Bo Long, Zhongfei Zhang, Philip S. Yu, Tianbing Xu
AAAI2
2008 Trajectory-Based Video Retrieval Using Dirichlet Process Mixture Models
abstract
In this paper, we present a trajectory-based video retrieval framework using Dirichlet process mixture models. The main contribution of this framework is four-fold. (1) We apply a Dirichlet process mixture model (DPMM) to unsupervised trajectory learning. DPMM is a countably infinite mixture model with its components growing by itself. (2) We employ a time-sensitive Dirichlet process mixture model (tDPMM) to learn trajectories ’ time-series characteristics. Furthermore, a novel likelihood estimation algorithm for tDPMM is proposed for the first time. (3) We develop a tDPMM-based probabilistic model matching scheme, which is empirically shown to be more error-tolerating and is able to deliver higher retrieval accuracy than the peer methods in the literature. (4) The framework has a nice scalability and adaptability in the sense that when new cluster data are presented, the framework automatically identifies the new cluster information without having to redo the training. Theoretic analysis and experimental evaluations against the state-of-the-art methods demonstrate the promise and effectiveness of the framework. 1
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
BMVC3
2008 Visual tracking via incremental Log-Euclidean Riemannian subspace learning
abstract
Recently, a novel Log-Euclidean Riemannian metric is proposed for statistics on symmetric positive definite (SPD) matrices. Under this metric, distances and Riemannian means take a much simpler form than the widely used affine-invariant Riemannian metric. Based on the Log-Euclidean Riemannian metric, we develop a tracking framework in this paper. In the framework, the covariance matrices of image features in the five modes are used to represent object appearance. Since a nonsingular covariance matrix is a SPD matrix lying on a connected Riemannian manifold, the Log-Euclidean Riemannian metric is used for statistics on the covariance matrices of image features. Further, we present an effective online Log-Euclidean Riemannian subspace learning algorithm which models the appearance changes of an object by incrementally learning a low-order Log-Euclidean eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking is then led by the Bayesian state inference framework in which a particle filter is used for propagating sample distributions over the time. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of the proposed framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Mingliang Zhu, Jian Cheng 0002
CVPR3
2008 Robust Visual Tracking Based on an Effective Appearance Model
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002
ECCV (4)3
2008 Dirichlet Process Based Evolutionary Clustering
abstract
Evolutionary Clustering has emerged as an important research topic in recent literature of data mining, and solutions to this problem have found a wide spectrum of applications, particularly in social network analysis. In this paper, based on the recent literature on Dirichlet processes, we have developed two different and specific models as solutions to this problem: DPChain and HDP-EVO. Both models substantially advance the literature on evolutionary clustering in the sense that not only they both perform better than the existing literature, but more importantly they are capable of automatically learning the cluster numbers and structures during the evolution. Extensive evaluations have demonstrated the effectiveness and promise of these models against the state-of-the-art literature.
Tianbing Xu, Zhongfei Zhang, Philip S. Yu, Bo Long
ICDM2
2008 Evolutionary Clustering by Hierarchical Dirichlet Process with Hidden Markov State
abstract
This paper studies evolutionary clustering, which is a recently hot topic with many important applications, noticeably in social network analysis. In this paper, based on the recent literature on Hierarchical Dirichlet Process (HDP) and Hidden Markov Model (HMM), we have developed a statistical model HDP-HTM that combines HDP with a Hierarchical Transition Matrix (HTM) based on the proposed Infinite Hierarchical Hidden Markov State model (iH2MS) as an effective solution to this problem. The HDP-HTM model substantially advances the literature on evolutionary clustering in the sense that not only it performs better than the existing literature, but more importantly it is capable of automatically learning the cluster numbers and structures and at the same time explicitly addresses the correspondence issue during the evolution. Extensive evaluations have demonstrated the effectiveness and promise of this solution against the state-of-the-art literature.
Tianbing Xu, Zhongfei Zhang, Philip S. Yu, Bo Long
ICDM2
2008 Multiclass spectral clustering based on discriminant analysis
abstract
Many existing spectral clustering algorithms share a conventional graph partitioning criterion: normalized cuts (NC). However, one problem with NC is that it poorly captures the graph¿s local marginal information which is very important to graph-based clustering. In this paper, we present a discriminant analysis based graph partitioning criterion (DAC), which is designed to effectively capture the graph¿s local marginal information characterized by the intra-class compactness and the inter-class separability. DAC preserves the intrinsic topological structures of the similarity graph on data points by constructing a k-nearest neighboring subgraph for each data point. Consequently, the clustering results generated by the DAC-based clustering algorithm (DACA) are robust to the outlier disturbance. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of DACA.
Xi Li 0001, Zhongfei Zhang, Yanguo Wang, Weiming Hu 0004
ICPR2
2008 Supervised Learning for Head Pose Estimation Using SVD and Gabor Wavelets
abstract
This paper presents the use of a template based method in order to make a head pose estimation. As an image classification problem the aim of this kind of techniques is to convert the input head image into a feature vector. The feature vectors of different persons taken at the same pose will serve to learn a head pose classifier. The aim of this work is to estimate the head pose of people looking at a target scene in order to extract the location of their gaze in the scene.
Adel Lablack, Zhongfei Zhang, Chaabane Djeraba
ISM2
2008 Mining Bulletin Board Systems Using Community Generation
Ming Li 0005, Zhongfei Zhang, Zhi-Hua Zhou
PAKDD2
2008 Semi-Supervised Learning Based on Semiparametric Regularization
abstract
Semi-supervised learning plays an important role in the recent literature on machine learning and data mining and the developed semisupervised learning techniques have led to many data mining applications in recent years. This paper addresses the semi-supervised learning problem by developing a semiparametric regularization based approach, which attempts to discover the marginal distribution of the data to learn the parametric function through exploiting the geometric distribution of the data. This learned parametric function can then be incorporated into the supervised learning on the available labeled data as the prior knowledge. Specifically, our contributions are: (1) We present a semi-supervised learning approach which incorporates the unlabeled data into the supervised learning by a parametric function learned from the whole data including the labeled and unlabeled data. The parametric function reflects the geometric structure of the marginal distribution of the data. Furthermore, the proposed approach which naturally extends to the out-of-sample data is an inductive learning method in nature. (2) This approach allows a family of algorithms to be developed based on various choices of the original RKHS and the loss function. (3) We provide experimental comparisons showing that the proposed approach leads the state-of-the-art performance on a variety of classification tasks. In particular, we demonstrate that this approach can be used successfully in both transductive and semisupervised settings. 1
Zhongfei Zhang, Eric P. Xing, Christos Faloutsos
SDM2
2008 A General Model for Multiple View Unsupervised Learning
abstract
Multiple view data, which have multiple representations from different feature spaces or graph spaces, arise in various data mining applications such as information retrieval, bioinformatics and social network analysis. Since different representations could have very different statistical properties, how to learn a consensus pattern from multiple representations is a challenging problem. In this paper, we propose a general model for multiple view unsupervised learning. The proposed model introduces the concept of mapping function to make the different patterns from different pattern spaces comparable and hence an optimal pattern can be learned from the multiple patterns of multiple representations. Under this model, we formulate two specific models for two important cases of unsupervised learning, clustering and spectral dimensionality reduction; we derive an iterating algorithm for multiple view clustering, and a simple algorithm providing a global optimum to multiple spectral dimensionality reduction. We also extend the proposed model and algorithms to evolutionary clustering and unsupervised learning with side information. Empirical evaluations on both synthetic and real data sets demonstrate the effectiveness of the proposed model and algorithms.
Bo Long, Philip S. Yu, Zhongfei Zhang
SDM3
2008 Automatic medical image annotation and retrieval
Jian Yao 0003, Zhongfei Zhang, Sameer K. Antani, L. Rodney Long, George R. Thoma
Neurocomputing2
2008 Editorial: Introduction to the Special Issue on Multimedia Data Mining
abstract
The twelve papers in this special issue focus on multimedia data mining. The special issue evolved from a successful workshop organized in conjunction with the 2006 ACM KDD conference, but the special issue was open to the whole community.
Zhongfei Zhang, Florent Masseglia, Ramesh Jain 0001, Alberto Del Bimbo
IEEE Trans. Multim.1
2008 Improving Web search using image snippets
abstract
The Web has become the largest information repository in the world; thus, effectively and efficiently searching the Web becomes a key challenge. Interactive Web search divides the search process into several rounds, and for each round the search engine interacts with the user for more knowledge of the user's information requirement. Previous research mainly uses the text information on Web pages, while little attention is paid to other modalities. This article shows that Web search performance can be significantly improved if imagery is considered in interactive Web search. Compared with text, imagery has its own advantage: the time for “reading” an image is as little as that for reading one or two words, while the information brought by an image is as much as that conveyed by a whole passage of text. In order to exploit the advantages of imagery, a novel interactive Web search framework is proposed, where image snippets are first extracted from Web pages and then provided, along with the text snippets, to the user for result presentation and relevance feedback, as well as being presented alone to the user for image suggestion. User studies show that it is more convenient for the user to identify the Web pages he or she expects and to reformulate the initial query. Further experiments demonstrate the promise of introducing multimodal techniques into the proposed interactive Web search framework.
Xiao-Bing Xue, Zhi-Hua Zhou, Zhongfei Zhang
ACM Trans. Internet Techn.3
2007 Graph Partitioning Based on Link Distributions
Bo Long, Zhongfei Zhang, Philip S. Yu
AAAI2
2007 Robust Visual Tracking Based on Incremental Tensor Subspace Learning
abstract
Most existing subspace analysis-based tracking algorithms utilize a flattened vector to represent a target, resulting in a high dimensional data learning problem. Recently, subspace analysis is incorporated into the multilinear framework which offline constructs a representation of image ensembles using high-order tensors. This reduces spatio-temporal redundancies substantially, whereas the computational and memory cost is high. In this paper, we present an effective online tensor subspace learning algorithm which models the appearance changes of a target by incrementally learning a low-order tensor eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking then is led by the state inference within the framework in which a particle filter is used for propagating sample distributions over the time. A novel likelihood function, based on the tensor reconstruction error norm, is developed to measure the similarity between the test image and the learned tensor subspace model during the tracking. Theoretic analysis and experimental evaluations against a state-of-the-art method demonstrate the promise and effectiveness of this algorithm.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
ICCV3
2007 Community Learning by Graph Approximation
abstract
Learning communities from a graph is an important problem in many domains. Different types of communities can be generalized as link-pattern based communities. In this paper, we propose a general model based on graph approximation to learn link-pattern based community structures from a graph. The model generalizes the traditional graph partitioning approaches and is applicable to learning various community structures. Under this model, we derive a family of algorithms which are flexible to learn various community structures and easy to incorporate the prior knowledge of the community structures. Experimental evaluation and theoretical analysis show the effectiveness and great potential of the proposed model and algorithms.
Bo Long, Zhongfei Zhang, Philip S. Yu
ICDM3
2007 Corner Detection of Contour Images using Spectral Clustering
abstract
Corner detection plays an important role in object recognition and motion analysis. In this paper, we propose a hierarchical corner detection framework based on spectral clustering (SC). The framework consists of three stages: contour smoothing, corner cell extraction and corner localization. In the contour smoothing stage, wavelet decomposition is imposed on the raw contour to reduce noise. In the corner cell extraction stage, several atomic corner cells are obtained by SC. In the corner localization stage, the corner points of each corner cell are located by the corner locator based on the kernel-weighted cosine curvature measure. Experimental results demonstrate the superiority of our framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang
ICIP (3)3
2007 A Max Margin Framework on Image Annotation and Multimodal Image Retrieval
abstract
This paper presents a max margin framework on image annotation and multimodal image retrieval as a structured prediction model. Following the max margin approach the image retrieval problem is formulated as a quadratic programming problem. By properly selecting joint feature representation between different modalities, our framework captures the dependency information between different modalities and avoids retraining the model from scratch when database undergoes dynamic updates. While this framework is a general approach which can be applied to multimodal information retrieval in any domains, we apply this approach to the Berkeley Drosophila embryo image database for the evaluation purpose. Experimental results show significant performance improvements over a state-of-the-art method.
Zhongfei Zhang, Eric P. Xing, Christos Faloutsos
ICME2
2007 Relational clustering by symmetric convex coding
abstract
Relational data appear frequently in many machine learning applications. Relational data consist of the pairwise relations (similarities or dissimilarities) between each pair of implicit objects, and are usually stored in relation matrices and typically no other knowledge is available. Although relational clustering can be formulated as graph partitioning in some applications, this formulation is not adequate for general relational data. In this paper, we propose a general model for relational clustering based on symmetric convex coding. The model is applicable to all types of relational data and unifies the existing graph partitioning formulation. Under this model, we derive two alternative bound optimization algorithms to solve the symmetric convex coding under two popular distance functions, Euclidean distance and generalized I-divergence. Experimental evaluation and theoretical analysis show the effectiveness and great potential of the proposed model and algorithms.
Bo Long, Zhongfei Zhang, Xiaoyun Wu, Philip S. Yu
ICML2
2007 Enhanced max margin learning on multimodal data mining in a multimedia database
abstract
The problem of multimodal data mining in a multimedia database can be addressed as a structured prediction problem where we learn the mapping from an input to the structured and interdependent output variables. In this paper, built upon the existing literature on the max margin based learning, we develop a new max margin learning approach called Enhanced Max Margin Learning (EMML) framework. In addition, we apply EMML framework to developing an effective and efficient solution to the multimodal data mining problem in a multimedia database. The main contributions include: (1) we have developed a new max margin learning approach - the enhanced max margin learning framework that is much more efficient in learning with a much faster convergence rate, which is verified in empirical evaluations; (2) we have applied this EMML approach to developing an effective and efficient solution to the multimodal data mining problem that is highly scalable in the sense that the query response time is independent of the database scale, allowing facilitating a multimodal data mining querying to a very large scale multimedia database,and excelling many existing multimodal data mining methods in the literature that do not scale up at all; this advantage is also supported through the complexity analysis as well as empirical evaluations against a state-of-the-art multimodal data mining method from the literature. While EMML is a general framework, for the evaluation purpose, we apply it to the Berkeley Drosophila embryo image database, and report the performance comparison with a state-of-the-art multimodal data mining method.
Zhongfei Zhang, Eric P. Xing, Christos Faloutsos
KDD2
2007 A probabilistic framework for relational clustering
abstract
Relational clustering has attracted more and more attention due to its phenomenal impact in various important applications which involve multi-type interrelated data objects, such as Web mining, search marketing, bioinformatics, citation analysis, and epidemiology. In this paper, we propose a probabilistic model for relational clustering, which also provides a principal framework to unify various important clustering tasks including traditional attributes-based clustering, semi-supervised clustering, co-clustering and graph clustering. The proposed model seeks to identify cluster structures for each type of data objects and interaction patterns between different types of objects. Under this model, we propose parametric hard and soft relational clustering algorithms under a large number of exponential family distributions. The algorithms are applicable to relational data of various structures and at the same time unifies a number of stat-of-the-art clustering algorithms: co-clustering algorithms, the k-partite graph clustering, Bregman k-means, and semi-supervised clustering based on hidden Markov random fields.
Bo Long, Zhongfei Zhang, Philip S. Yu
KDD2
2007 Effective Image Retrieval Based on Hidden Concept Discovery in Image Database
abstract
This paper addresses content-based image retrieval in general, and in particular, focuses on developing a hidden semantic concept discovery methodology to address effective semantics-intensive image retrieval. In our approach, each image in the database is segmented into regions associated with homogenous color, texture, and shape features. By exploiting regional statistical information in each image and employing a vector quantization method, a uniform and sparse region-based representation is achieved. With this representation, a probabilistic model based on statistical-hidden-class assumptions of the image database is obtained, to which the expectation-maximization technique is applied to analyze semantic concepts hidden in the database. An elaborated retrieval algorithm is designed to support the probabilistic model. The semantic similarity is measured through integrating the posterior probabilities of the transformed query image, as well as a constructed negative example, to the discovered semantic concepts. The proposed approach has a solid statistical foundation; the experimental evaluations on a database of 10000 general-purposed images demonstrate its promise and effectiveness.
Ruofei Zhang, Zhongfei Zhang
IEEE Trans. Image Process.2
2006 Improve Web Search Using Image Snippets
Xiao-Bing Xue, Zhi-Hua Zhou, Zhongfei Zhang
AAAI3
2006 Automatic Medical Image Annotation and Retrieval Using SECC
abstract
The demand for automatically annotating and retrieving medical images is growing faster than ever. In this paper, we present a novel medical image annotation method based on the proposed Semantic Error-Correcting output Codes (SECC). With this annotation method, we present a new semantic image retrieval method, which exploits the high level semantic similarity. For example, a user may query the system using an image of arm while he/she expects images of hand. This cannot be realized by traditional retrieval methods. The experimental results on the IMAGECLEF 2005 annotation data set clearly show the strength and the promise of the presented methods.
Jian Yao 0003, Sameer K. Antani, L. Rodney Long, George R. Thoma, Zhongfei Zhang
CBMS5
2006 Automatic Medical Image Annotation and Retrieval using SEMI-SECC
abstract
The demand for automatically annotating and retrieving medical images is growing faster than ever. In this paper, we present a novel medical image retrieval method based on SEMI-supervised Semantic Error-Correcting output Codes (SEMI-SEC). The experimental results on IMAGECLEF 2005 [1] annotation data set clearly show the strength and the promise of the presented methods.
Jian Yao 0003, Zhongfei Zhang, Sameer K. Antani, L. Rodney Long, George R. Thoma
ICME2
2006 Spectral clustering for multi-type relational data
abstract
Clustering on multi-type relational data has attracted more and more attention in recent years due to its high impact on various important applications, such as Web mining, e-commerce and bioinformatics. However, the research on general multi-type relational data clustering is still limited and preliminary. The contribution of the paper is three-fold. First, we propose a general model, the collective factorization on related matrices, for multi-type relational data clustering. The model is applicable to relational data with various structures. Second, under this model, we derive a novel algorithm, the spectral relational clustering, to cluster multi-type interrelated data objects simultaneously. The algorithm iteratively embeds each type of data objects into low dimensional spaces and benefits from the interactions among the hidden structures of different types of data objects. Extensive experiments demonstrate the promise and effectiveness of the proposed algorithm. Third, we show that the existing spectral clustering algorithms can be considered as the special cases of the proposed model and algorithm. This demonstrates the good theoretic generality of the proposed model and algorithm.
Bo Long, Zhongfei Zhang, Xiaoyun Wu, Philip S. Yu
ICML2
2006 Unsupervised learning on k-partite graphs
abstract
Various data mining applications involve data objects of multiple types that are related to each other, which can be naturally formulated as a k-partite graph. However, the research on mining the hidden structures from a k-partite graph is still limited and preliminary. In this paper, we propose a general model, the relation summary network, to find the hidden structures (the local cluster structures and the global community structures) from a k-partite graph. The model provides a principal framework for unsupervised learning on k-partite graphs of various structures. Under this model, we derive a novel algorithm to identify the hidden structures of a k-partite graph by constructing a relation summary network to approximate the original k-partite graph under a broad range of distortion measures. Experiments on both synthetic and real data sets demonstrate the promise and effectiveness of the proposed model and algorithm. We also establish the connections between existing clustering approaches and the proposed model to provide a unified view to the clustering approaches.
Bo Long, Xiaoyun Wu, Zhongfei Zhang, Philip S. Yu
KDD3
2006 Hierarchical shadow detection for color aerial images
Jian Yao 0003, Zhongfei Zhang
Comput. Vis. Image Underst.2
2006 BALAS: Empirical Bayesian learning in the relevance feedback for image retrieval
Ruofei Zhang, Zhongfei Zhang
Image Vis. Comput.2
2006 Guest Editors' introduction to the special issue: machine learning approaches to multimedia information retrieval
Arif Ghafoor, Zhongfei Zhang, Michael S. Lew, Zhi-Hua Zhou
Multim. Syst.2
2006 A probabilistic semantic model for image annotation and multi-modal image retrieval
Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang
Multim. Syst.2
2005 Semi-Supervised Learning Based Object Detection in Aerial Imagery
abstract
Object detection in aerial imagery has been well studied in computer vision for years. However, given the complexity of large variations of the appearance of the object and the background in a typical aerial image, a robust and efficient detection is still considered as an open and challenging problem. In this paper, we have developed a theoretic foundation for aerial imagery object detection using semi-supervised learning algorithms. Based on this theory, we have proposed a context-based object detection methodology. Both theoretic analyses and experimental evaluations have successfully demonstrated the great promise of the developed theory and the related detection methodology.
Jian Yao 0003, Zhongfei Zhang
CVPR (1)2
2005 Object Detection in Aerial Imagery Based on Enhanced Semi-Supervised Learning
abstract
Object detection in aerial imagery has been well studied in computer vision for years. However, given the complexity of large variations of the appearance of the object and the background in a typical aerial image, a robust and efficient detection is still considered as an open and challenging problem. In this paper, we present the enhanced semi-supervised learning (ESL) framework and apply this framework to revising an object detection methodology we have developed in a previous effort. Theoretic analysis and experimental evaluation using the UCI machine learning repository clearly indicate the superiority of the ESL framework. The performance evaluations of the revised object detection methodology against the original one clearly demonstrate the promise and superiority of this approach.
Jian Yao 0003, Zhongfei Zhang
ICCV2
2005 A Probabilistic Semantic Model for Image Annotation and Multi-Modal Image Retrieva
abstract
This paper addresses automatic image annotation problem and its application to multi-modal image retrieval. The contribution of our work is three-fold. (1) We propose a probabilistic semantic model in which the visual features and the textual words are connected via a hidden layer which constitutes the semantic concepts to be discovered to explicitly exploit the synergy among the modalities. (2) The association of visual features and textual words is determined in a Bayesian framework such that the confidence of the association can be provided. (3) Extensive evaluation on a large-scale, visually and semantically diverse image collection crawled from Web is reported to evaluate the prototype system based on the model. In the proposed probabilistic model, a hidden concept layer which connects the visual feature and the word layer is discovered by fitting a generative model to the training image and annotation words through an Expectation-Maximization (EM) based iterative learning procedure. The evaluation of the prototype system on 17,000 images and 7,736 automatically extracted annotation words from crawled Web pages for multi-modal image retrieval has indicated that the proposed semantic model and the developed Bayesian framework are superior to a state-of-the-art peer system in the literature.
Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang
ICCV2
2005 Combining Multiple Clusterings by Soft Correspondence
abstract
Combining multiple clusterings arises in various important data mining scenarios. However, finding a consensus clustering from multiple clusterings is a challenging task because there is no explicit correspondence between the classes from different clusterings. We present a new framework based on soft correspondence to directly address the correspondence problem in combining multiple clusterings. Under this framework, we propose a novel algorithm that iteratively computes the consensus clustering and correspondence matrices using multiplicative updating rules. This algorithm provides a final consensus clustering as well as correspondence matrices that gives intuitive interpretation of the relations between the consensus clustering and each clustering from clustering ensembles. Extensive experimental evaluations also demonstrate the effectiveness and potential of this framework as well as the algorithm for discovering a consensus clustering from multiple clusterings.
Bo Long, Zhongfei Zhang, Philip S. Yu
ICDM2
2005 Image Database Classification based on Concept Vector Model
abstract
Automatic semantic classification of image databases is very useful for users' searching and browsing, but it is at the same time a very challenging research problem as well. In this paper, we develop a hidden semantic concept discovery methodology to address effective semantics-intensive image database classification. Each image is segmented into regions and then a uniform and sparse region-based representation is obtained. With this representation a probabilistic model based on statistical-hidden-class assumptions of the image database is proposed, to which the Expectation-Maximization (EM) technique is applied to analyze semantic concepts hidden in the database. Two methods are proposed to make use of the semantic concepts discovered from the probabilistic model for unsupervised and supervised image database classifications, respectively, based on the automatically learned concept vectors. It is shown that the concept vectors are more reliable and robust and thus promising than the low level features through the theoretic analysis and the experimental evaluations on a database of 10,000 general-purpose images.
Ruofei Zhang, Zhongfei Zhang
ICME2
2005 Co-clustering by block value decomposition
abstract
Dyadic data matrices, such as co-occurrence matrix, rating matrix, and proximity matrix, arise frequently in various important applications. A fundamental problem in dyadic data analysis is to find the hidden block structure of the data matrix. In this paper, we present a new coclustering framework, block value decomposition(BVD), for dyadic data, which factorizes the dyadic data matrix into three components, the row-coefficient matrix R, the block value matrix B, and the column-coefficient matrix C. Under this framework, we focus on a special yet very popular case – non-negative dyadic data, and propose a specific novel co-clustering algorithm that iteratively computes the three decomposition matrices based on the multiplicative updating rules. Extensive experimental evaluations also demonstrate the effectiveness and potential of this framework as well as the specific algorithms for co-clustering, and in particular, for discovering the hidden block structure in the dyadic data.
Bo Long, Zhongfei Zhang, Philip S. Yu
KDD2
2005 FAST: Toward more effective and efficient image retrieval
Ruofei Zhang, Zhongfei Zhang
Multim. Syst.2
2004 Hidden Semantic Concept Discovery in Region Based Image Retrieval
Ruofei Zhang, Zhongfei Zhang
CVPR (2)2
2004 Stretching Bayesian Learning in the Relevance Feedback of Image Retrieval
Ruofei Zhang, Zhongfei Zhang
ECCV (3)2
2004 Cognitive bridge between haptic impressions and texture images for subjective image retrieval
abstract
As a step towards subjective image retrieval, This work reports an on-going collaboration project between Waseda University and SUNY, Binghamton, on relating texture images to haptic impressions. To capture the surface height variations, texture images are taken under different illuminations and viewing conditions. Our method applies a new frequency analysis method to the texture images. We evaluate the performances of our feature and other typical conventional features by checking whether texture images are correctly classified into "soft" or "hard" by the SVM (support vector machine) method, where the training data for the SVM are collected by subjective tests. Experimental results show that our texture feature can classify "soft" or "hard" better than the other features.
Yuichi Kobayashi, Jun Ohya, Zhongfei Zhang
ICME3
2004 Exploiting the cognitive synergy between different media modalities in multimodal information retrieval
abstract
This is a position paper reporting an on-going collaboration project between SUNY Binghamton, USA, and Waseda University, Japan, on multimodal information retrieval through exploiting the cognitive synergy across the different modalities of the information, to facilitate an effective retrieval. Specifically, we focus on image retrieval in the applications where imagery data appear along with collateral text. It is noted that these applications are ubiquitous. We have proposed the synergistic indexing scheme (SIS) to explicitly exploit the synergy between the information of imagery and text modalities. Since the synergy we have exploited between the information of imagery and text modalities is subjective and depends on specific cognitive context, we call this type of synergy as cognitive synergy. We have reported part of the empirical evaluation and are in the process of fully implementing the SIS prototype for an extensive evaluation.
Zhongfei Zhang, Ruofei Zhang, Jun Ohya
ICME1
2004 Semantic repository modeling in image database
abstract
This work is about content based image database retrieval, focusing on developing a classification based methodology to address semantics-intensive image retrieval. With self organization map based image feature grouping, a visual dictionary is created for color, texture, and shape feature attributes, respectively. Labeling each training image with the keywords in the visual dictionary, a classification tree is built. Based on the statistical properties of the feature space we define a structure, called /spl alpha/-semantics graph, to discover the hidden semantic relationships among the semantic repositories embodied in the image database. With the /spl alpha/-semantics graph, each semantic repository is modeled as a unique fuzzy set to explicitly address the semantic uncertainty and the semantic overlap existing among the repositories in the feature space. A retrieval algorithm combining the classification tree with the fuzzy set models to deliver semantically relevant image retrieval is provided. The experimental evaluations have demonstrated that the proposed approach models the semantic relationships effectively and outperforms a state-of-the-art content based image retrieval system in the literature both in effectiveness and efficiency.
Ruofei Zhang, Zhongfei Zhang, Zhongyuan Qin
ICME2
2004 A data mining approach to modeling relationships among categories in image collection
abstract
This paper proposes a data mining approach to modeling relationships among categories in image collection. In our approach, with image feature grouping, a visual dictionary is created for color, texture, and shape feature attributes respectively. Labeling each training image with the keywords in the visual dictionary, a classification tree is built. Based on the statistical properties of the feature space we define a structure, called α-Semantics Graph, to discover the hidden semantic relationships among the semantic categories embodied in the image collection. With the α-Semantics Graph, each semantic category is modeled as a unique fuzzy set to explicitly address the semantic uncertainty and semantic overlap among the categories in the feature space. The model is utilized in the semantics-intensive image retrieval application. An algorithm using the classification accuracy measures is developed to combine the built classification tree with the fuzzy set modeling method to deliver semantically relevant image retrieval for a given query image. The experimental evaluations have demonstrated that the proposed approach models the semantic relationships effectively and the image retrieval prototype system utilizing the derived model is promising both in effectiveness and efficiency.
Ruofei Zhang, Zhongfei Zhang, Sandeep Khanzode
KDD2
2003 "Automatic" Multimodal Medical Image Fusion
Zhongfei Zhang, Jian Yao 0003, Saeed Bajwa, Thomas Gudas
CBMS1
2003 Applying data mining in investigating money laundering crimes
abstract
In this paper, we study the problem of applying data mining to facilitate the investigation of money laundering crimes (MLCs). We have identified a new paradigm of problems --- that of automatic community generation based on uni-party data, the data in which there is no direct or explicit link information available. Consequently, we have proposed a new methodology for Link Discovery based on Correlation Analysis (LDCA). We have used MLC group model generation as an exemplary application of this problem paradigm, and have focused on this application to develop a specific method of automatic MLC group model generation based on timeline analysis using the LDCA methodology, called CORAL. A prototype of CORAL method has been implemented, and preliminary testing and evaluations based on a real MLC case data are reported. The contributions of this work are: (1) identification of the uni-party data community generation problem paradigm, (2) proposal of a new methodology LDCA to solve for problems in this paradigm, (3) formulation of the MLC group model generation problem as an example of this paradigm, (4) application of the LDCA methodology in developing a specific solution (CORAL) to the MLC group model generation problem, and (5) development, evaluation, and testing of the CORAL prototype in a real MLC case data.
Zhongfei Zhang, John J. Salerno, Philip S. Yu
KDD1
2003 A Unified Fuzzy Feature Indexing Scheme for Region Based Online Image Querying
abstract
This paper describes a novel indexing and retrieval methodology integrating color, texture and shape information for content-based image retrieval in online image databases. This methodology, called PicSearcher, applies unsupervised image segmentation to partition an image into a set of regions, then fuzzy color histogram as well as fuzzy texture and shape properties of each region is calculated to be part of their signatures. The fuzzification procedures resolve the recognition uncertainty stemming from color quantization and human perception of colors. At the same time, this unified fuzzy scheme incorporates the segmentation-related uncertainties into the retrieval algorithm. Then an adaptive and effective measure for the overall similarity between images is developed by integrating properties of all the regions in the image. An implemented prototype system of PicSearcher has demonstrated a promising retrieval performance for an online test database containing 10,000 general-purpose color images, as compared with its peer systems in the literature. 1.
Ruofei Zhang, Zhongfei Zhang, Jian Yao 0003
Web Intelligence2
2002 Mining Surveillance Video for Independent Motion Detection
abstract
This paper addresses the special applications of data mining techniques in homeland defense. The problem targeted, which is frequently encountered in military/intelligence surveillance, is to mine a massive surveillance video database automatically collected to retrieve the shots containing independently moving targets. A novel solution to this problem is presented in this paper, which offers a completely qualitative approach to solving for the automatic independent motion detection problem directly from the compressed surveillance video in a faster than real-time mining performance. This approach is based on the linear system consistency analysis, and consequently is called QLS. Since the QLS approach only focuses on what exactly is necessary to compute a solution, it saves the computation to a minimum and achieves the efficacy to the maximum. Evaluations from real data show that QLS delivers effective mining performance at the achieved efficiency.
Zhongfei Zhang
ICDM1
2002 A Clustering Based Approach to Efficient Image Retrieval
abstract
This paper addresses the issue of effective and efficient content based image retrieval by presenting a novel indexing and retrieval methodology that integrates color, texture, and shape information for the indexing and retrieval, and applies these features in regions obtained through unsupervised segmentation, as opposed to applying them to the whole image domain. In order to address the typical color feature "inaccuracy" problem in the literature, fuzzy logic is applied to the traditional color histogram to solve for the problem to a certain degree. The similarity is defined through a balanced combination between global and regional similarity measures incorporating all the features. In order to further improve the retrieval efficiency, a secondary clustering technique is developed and employed to significantly save query processing time without compromising the retrieval precision. An implemented prototype system has demonstrated a promising retrieval performance for a test database containing 2000 general-purpose color images, as compared with its peer systems in the literature.
Ruofei Zhang, Zhongfei Zhang
ICTAI2
2002 Subspace morphing theory for appearance based object identification
Zhongfei Zhang, Rohini K. Srihari
Pattern Recognit.1
2001 A Web-Based Multimedia Diabetes Mellitus Education Tool for School Nurses
abstract
Presents a Web-based diabetes mellitus education program for school nurses, called WebDiamen. Compared with existing systems for medical education using computer-assisted technologies, WebDiamen exhibits the advantages of a powerful browsing capability, a friendly GUI, an individualized presentation schedule tailored to each user based on their special browsing interests, powerful online processing and annotation capabilities, and powerful querying capabilities. The system is being evaluated for school nurses in Onondaga County, NY, USA.
Zhongfei Zhang, Paul E. Knudson, Ruth S. Weinstock, Suzanne Meyer
CBMS1
2001 IBMAS: An Internet Based Medical Archive System
abstract
Presents IBMAS (Internet-Based Medical Archive System). IBMAS is a prototype system of a proposed research paradigm aiming at building up a distributed archival system over the Internet to facilitate medical practitioners to maintain, share, update, search and process medical information conveniently and consistently. IBMAS differs from existing systems in the sense that it not only offers a full spectrum of online communications, processing and annotation tools, but also provides powerful multi-modal search functionalities to the users. In addition, the database is always kept in a "live" mode such that information contributed by users is periodically indexed. A preliminary version of IBMAS containing over 2,000 medical images in different modalities along with associated annotation text for each of the images has been fully implemented and is under evaluation.
Zhongfei Zhang, Guangbiao Pu, Andrzej Król
CBMS1
2000 Intelligent Indexing and Semantic Retrieval of Multimodal Documents
Rohini K. Srihari, Zhongfei Zhang, Aibing Rao
Inf. Retr.2
1999 Spatial Color Histograms for Content-Based Image Retrieval
abstract
The color histogram is an important technique for color image database indexing and retrieval. In this paper, the traditional color histogram is modified to capture the spatial layout information of each color, and three types of spatial color histograms are introduced: annular, angular and hybrid color histograms. Experiments show that, with a proper trade-off between the granularity in the color and spatial dimensions, these histograms outperform both the traditional color histogram and some existing histogram refinements such as the color coherent vector.
Aibing Rao, Rohini K. Srihari, Zhongfei Zhang
ICTAI3
1999 Face Detection and Its Applications in Intelligent and Focused Image Retrieval
abstract
This paper presents a face detection technique and its applications in image retrieval. Even though this face detection method has relatively high false positives and a low detection rate (as compared with the dedicated face detection systems in the literature of image understanding), because of its simple and fast nature, it has been shown that this system may be well applied in image retrieval in certain focused application domains. Two application examples are given: one combining face detection with indexed collateral text for image retrieval regarding human beings, and the other combining face detection with conventional similarity matching techniques for image retrieval with similar background. Experimental results show that our proposed approaches have significantly improved image retrieval precision over existing search engines in these focused application domains.
Zhongfei Zhang, Rohini K. Srihari, Aibing Rao
ICTAI1
1998 Identifying human faces in general appearances
abstract
Identifying people in general appearances is a very practical, widely applicable, but extremely difficult problem, because it requires solutions to both automatic detection and recognition of human faces in arbitrary appearances and complex backgrounds. General appearances are referred to arbitrary variations in terms of face pose, orientation, expression, scale, background, as well as image contrast. A streamlined solution to face detection and face recognition in general appearances is proposed, which has the advantage of no requirement on normalizing image scales either in the stage of building up the face image library for recognition or in the stage of recognition per se. This is valuable in practice for it allows the application of this technique in many general scenarios in which general face appearances are expected under complex backgrounds. Experimental evaluations show that this method is a vital approach to many real applications.
Zhongfei Zhang
SMC1
1997 Obstacle Detection Based on Qualitative and Quantitative 3D Reconstruction
abstract
Three different algorithms for mobile robot vision-based obstacle detection are presented in this paper each based on different assumptions The first two algorithms are qualitative in that they return only yes/no answers regarding the presence of obstacles in the field of view; no 3D reconstruction is performed. They have the advantage of fast determination of the existence of obstacles in a scene based on the solvability of a linear system. The first algorithm uses information about the ground plane, while the second only assumes that the ground is planar. The third algorithm is quantitative in that it continuously estimates the ground plane and reconstructs partial 3D structures by determining the height above the ground plane of each point in the scene. Experimental results are presented for real and simulated data, and the performance of the three algorithms under different noise levels is compared in simulation. We conclude that in terms of the robustness of performance, the third algorithm is superior to the other two.
Zhongfei Zhang, Richard Weiss 0001, Allen R. Hanson
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Qualitative obstacle detection
abstract
Three different algorithms for qualitative obstacle detection are presented in this paper. Each one is based on different assumptions. The first two algorithms are aimed at yes/no obstacle detection without indicating which points are obstacles. They have the advantage of fast determination of the existence of obstacles in a scene based on the solvability of a linear system. The first algorithm uses information about the ground plane, while the second algorithm only assumes that the ground is planar. The third algorithm continuously estimates the ground plane, and based on that determines the height of each matched point in the scene. Experimental results are presented for real and simulated data, and performances of the three algorithms under different noise levels are compared in simulation. We conclude that in terms of the robustness of performance, the third one works best.>
Zhongfei Zhang, Richard Weiss 0001, Allen R. Hanson
CVPR1
1991 Feature matching in 360° waveforms for robot navigation
abstract
A target location is represented by compressing a 360 degrees image from a spherical mirror into a circular waveform. Navigation tasks are specified as a sequence of homing tasks on target locations. This is accomplished by matching landmarks from the current view with the next target view. A method is provided for matching using qualitative geometric features of each of the waveforms. In addition, a geometric analysis allows one to do 3-D reasoning about the environment, which is sufficient for computing the rotation and translation between the current location and the target location. It eventually may allow acquisition of a 3-D model of the environment.>
Zhongfei Zhang, Richard Weiss 0001, Edward M. Riseman
CVPR1