VLDB 2026 Research / reviewers in the wild / expert
Zhigang Ma
dblp:71/9924
· DBLP profile ↗
44ranked-venue papers
15as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 12 first-authorArtificial intelligence and machine learning · 22 · 5 first-authorComputer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Efficient and distributed learning · 56% Video understanding and tracking · 21% Transfer learning and domain adaptation · 7% | |
| Databases, data mining, and information retrieval
8 papers |
Data mining · 59% Recommender systems · 27% Information retrieval · 9% | |
| Computer graphics and multimedia
12 papers |
Multimedia analysis and retrieval · 92% Image and video processing · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Smart cities and intelligent transportation · 100% | |
| Computer networks
1 paper |
Edge and fog computing · 100% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI Recommendations · ACM Trans. Inf. Syst. 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI Recommendations · ACM Trans. Inf. Syst. 2026 |
Machine learning › Efficient and distributed learning › model compression › quantization
product quantization |
1.0 | 1 | 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI Recommendations · ACM Trans. Inf. Syst. 2026 |
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation |
1.0 | 1 | 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI Recommendations · ACM Trans. Inf. Syst. 2026 |
Smart cities and intelligent transportation
traffic prediction |
1.0 | 1 | 2026 | Autonomous Spatiotemporal Graph Learning for Proactive Edge-Based Traffic Forecasting · INFOCOM 2026 |
Recommender systems
point-of-interest recommendation |
1.0 | 1 | 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI Recommendations · ACM Trans. Inf. Syst. 2026 |
Multimedia analysis and retrieval › event detection
video event detection |
0.6 | 4 | 2013 | Multimedia Event Detection Using A Classifier-Specific Intermediate Representation · IEEE Trans. Multim. 2013 We are not equally negative: fine-grained labeling for multimedia event detection · ACM Multimedia 2013 Complex Event Detection via Multi-source Video Attributes · CVPR 2013 |
Data mining › dimensionality reduction
feature selection |
0.6 | 4 | 2013 | GLocal structural feature selection with sparsity for multimedia data understanding · ACM Multimedia 2013 Discriminating Joint Feature Analysis for Multimedia Data Understanding · IEEE Trans. Multim. 2012 Web Image Annotation Via Subspace-Sparsity Collaborated Feature Selection · IEEE Trans. Multim. 2012 |
Data mining
clustering |
0.5 | 2 | 2018 | Balanced Clustering via Exclusive Lasso: A Pragmatic Approach · AAAI 2018 A Convex Formulation for Spectral Shrunk Clustering · AAAI 2015 |
Computer vision › Video understanding and tracking
action recognition |
0.5 | 3 | 2014 | Semi-Supervised Multiple Feature Analysis for Action Recognition · IEEE Trans. Multim. 2014 Harnessing Lab Knowledge for Real-World Action Recognition · Int. J. Comput. Vis. 2014 Action recognition by exploring data distribution and feature correlation · CVPR 2012 |
Computer vision › Video understanding and tracking › event recognition
complex event detection |
0.5 | 2 | 2017 | The Many Shades of Negativity · IEEE Trans. Multim. 2017 How Related Exemplars Help Complex Event Detection in Web Videos? · ICCV 2013 |
Multimedia analysis and retrieval
video analysis |
0.4 | 2 | 2014 | Multiple Features But Few Labels?: A Symbiotic Solution Exemplified for Video Analysis · ACM Multimedia 2014 Complex Event Detection via Multi-source Video Attributes · CVPR 2013 |
Multimedia analysis and retrieval
event detection |
0.3 | 2 | 2014 | Event Detection Using Multi-level Relevance Labels and Multiple Features · CVPR 2014 Robust cross-media transfer for visual event detection · ACM Multimedia 2012 |
Data mining › clustering › constrained clustering
balanced clustering |
0.3 | 1 | 2018 | Balanced Clustering via Exclusive Lasso: A Pragmatic Approach · AAAI 2018 |
Data mining › clustering
k-means clustering |
0.3 | 1 | 2018 | Balanced Clustering via Exclusive Lasso: A Pragmatic Approach · AAAI 2018 |
Data mining › clustering › graph clustering
min-cut clustering |
0.3 | 1 | 2018 | Balanced Clustering via Exclusive Lasso: A Pragmatic Approach · AAAI 2018 |
Recommender systems › click-through rate prediction
feature interaction |
0.3 | 1 | 2017 | Feature Interaction Augmented Sparse Learning for Fast Kinect Motion Detection · IEEE Trans. Image Process. 2017 |
Machine learning and data management
sparse learning |
0.3 | 1 | 2017 | Feature Interaction Augmented Sparse Learning for Fast Kinect Motion Detection · IEEE Trans. Image Process. 2017 |
Machine learning › Learning paradigms
semi-supervised learning |
0.3 | 3 | 2014 | Semi-Supervised Multiple Feature Analysis for Action Recognition · IEEE Trans. Multim. 2014 Multiple Features But Few Labels?: A Symbiotic Solution Exemplified for Video Analysis · ACM Multimedia 2014 Exploiting the entire feature space with sparsity for automatic image annotation · ACM Multimedia 2011 |
Machine learning › Efficient and distributed learning
active learning |
0.2 | 1 | 2015 | Multi-Class Active Learning by Uncertainty Sampling with Diversity Maximization · Int. J. Comput. Vis. 2015 |
Machine learning › Efficient and distributed learning › active learning
uncertainty sampling |
0.2 | 1 | 2015 | Multi-Class Active Learning by Uncertainty Sampling with Diversity Maximization · Int. J. Comput. Vis. 2015 |
Data mining › dimensionality reduction
manifold learning |
0.2 | 1 | 2015 | A Convex Formulation for Spectral Shrunk Clustering · AAAI 2015 |
Data mining › clustering
spectral clustering |
0.2 | 1 | 2015 | A Convex Formulation for Spectral Shrunk Clustering · AAAI 2015 |
Computer vision › Video understanding and tracking › action recognition
cross-domain action recognition |
0.2 | 1 | 2014 | Harnessing Lab Knowledge for Real-World Action Recognition · Int. J. Comput. Vis. 2014 |
Computer vision › Video understanding and tracking › event recognition
multimedia event detection |
0.2 | 1 | 2014 | Knowledge Adaptation with PartiallyShared Features for Event DetectionUsing Few Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2014 |
Multimedia analysis and retrieval
multimodal fusion |
0.2 | 1 | 2014 | Event Detection Using Multi-level Relevance Labels and Multiple Features · CVPR 2014 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2013 | Thinking of Images as What They Are: Compound Matrix Regression for Image Classification · IJCAI 2013 |
Machine learning › Learning paradigms › multi-task learning › multi-task feature learning
multi-task feature selection |
0.2 | 1 | 2013 | Feature Selection for Multimedia Analysis by Sharing Information Among Multiple Tasks · IEEE Trans. Multim. 2013 |
Computer vision › Video understanding and tracking › event recognition
video event detection |
0.2 | 1 | 2013 | Multimedia Event Detection Using A Classifier-Specific Intermediate Representation · IEEE Trans. Multim. 2013 |
Multimedia analysis and retrieval › event detection
complex event detection |
0.2 | 1 | 2013 | Complex Event Detection via Multi-source Video Attributes · CVPR 2013 |
Methods — techniques the papers use, named apart from their topics
product quantization · 2.0mixture of experts · 2.0knowledge distillation · 2.0graph learning · 2.0spatiotemporal learning · 1.0spatio-temporal learning · 1.0multimodal representation · 1.0multi-modal representation · 1.0semi-supervised learning · 0.8manifold learning · 0.5joint optimization · 0.5knowledge transfer · 0.3exclusive lasso · 0.3sparse modeling · 0.3schatten-p norm · 0.3linear classification · 0.3convolutional neural network · 0.3spectral clustering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Autonomous Spatiotemporal Graph Learning for Proactive Edge-Based Traffic Forecasting
Ruifeng Pan, Nan Zhao 0006, Zhigang Ma |
INFOCOM | 4 |
| 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI RecommendationsabstractIn large-scale Point-of-Interest (POI) recommendation, the conflict between accuracy and computational efficiency intensifies as POI catalogs grow. Traditional deep models struggle to balance quality with efficiency. To address this challenge, we propose a knowledge-distilled product quantization framework EffiPOI for efficient POI recommendation. EffiPOI jointly optimizes accuracy and efficiency by integrating product quantization with multi-modal knowledge distillation. Specifically, we first construct service-oriented multi-modal POI representations, which comprehensively capture each POI’s spatial coverage, temporal activity patterns, and semantic attributes. Based on these representations, we design a teacher-student distillation paradigm. The teacher model adopts a Mixture-of-Experts architecture to generate discriminative and semantically expressive POI representations, which serve as high-quality supervision signals for guiding the student model through knowledge distillation. The student model leverages product quantization to encode POIs into compact and computation-friendly representations, achieving a favorable tradeoff between representational compactness and predictive accuracy. To alleviate the performance degradation due to quantization, we develop a hybrid knowledge distillation strategy that transfers both response-aware and feature-aware knowledge from the teacher model to the student model. Experimental results on three real-world datasets show that the proposed method achieves 4.6%–12.1% improvements in accuracy and over 10× speedup in inference efficiency, outperforming existing POI recommendation models. Code is available at: https://github.com/pcm1217/EffiPOI . Chengmei Peng, Yang Xu 0025, Lei Zhu 0002, Fengling Li 0001, Huaxiang Zhang 0001, Zhigang Ma |
ACM Trans. Inf. Syst. | 6 |
| 2024 | Self-Supervised Transfer Learning for Remote Wear Evaluation in Machine Tool Elements With Imaging Transmission AttenuationabstractThe Industrial Internet of Things (IIoT) has significantly advanced traditional industrial systems, especially in facilitating remote monitoring and predictive maintenance for Computer Numerical Control (CNC) machines. Ball Screw Drives (BSDs), crucial in CNC machining, require regular upkeep, often challenged by environmental influences and the limitations of wired sensor-based diagnostic procedures. Wireless sensors offer a cost-effective solution but struggle with data integrity during transmission, impacting remote wear evaluation of BSDs. Current recovery methods are not always adequate, often relying on extensive historical data and suffering from accumulating errors. Addressing these limitations, a novel Self-supervised Transfer Learning (SSTL) model is proposed for remote wear assessment of machine tool components. This model integrates an image capture module into CNC surveillance systems and employs a LocalMIM module, which is pre-trained and fine-tuned to adapt to various domains, especially for image quality deterioration during data transmission. The SSTL model is designed to function effectively despite significant pixel data loss, eliminating the need for historical data for image restoration. This innovation is particularly adept at evaluating wear in environments with compromised image transmission, providing a robust predictive maintenance strategy for BSDs that is less affected by data loss, aiding a pressing industrial requirement for consistent and dependable remote monitoring solutions. Zhigang Ma, Chaojun Xu, Yaqiang Jin, Chengning Zhou |
IEEE Internet Things J. | 2 |
| 2019 | Adaptive Structure Discovery for Multimedia Analysis Using Multiple FeaturesabstractMultifeature learning has been a fundamental research problem in multimedia analysis. Most existing multifeature learning methods exploit graph, which must be computed beforehand, as input to uncover data distribution. These methods have two major problems confronted. First, graph construction requires calculating similarity based on nearby data pairs by a fixed function, e.g., the RBF kernel, but the intrinsic correlation among different data pairs varies constantly. Therefore, feature learning based on such predefined graphs may degrade, especially when there is dramatic correlation variation between nearby data pairs. Second, in most existing algorithms, each single-feature graph is computed independently and then combine them for learning, which ignores the correlation between multiple features. In this paper, a new unsupervised multifeature learning method is proposed to make the best utilization of the correlation among different features by jointly optimizing data correlation from multiple features in an adaptive way. As opposed to computing the affinity weight of data pairs by a fixed function, the weight of affinity graph is learned by a well-designed optimization problem. Additionally, the affinity graph of data pairs from different features is optimized in a global level to better leverage the correlation among different channels. In this way, the adaptive approach correlates the features of all features for a better learning process. Experimental results on real-world datasets demonstrate that our approach outperforms the state-of-the-art algorithms on leveraging multiple features for multimedia analysis. Kun Zhan, Xiaojun Chang, Junpeng Guan, Ling Chen 0006, Zhigang Ma, Yi Yang 0001 |
IEEE Trans. Cybern. | 5 |
| 2018 | Balanced Clustering via Exclusive Lasso: A Pragmatic ApproachabstractClustering is an effective technique in data mining to generate groups that are the matter of interest.Among various clustering approaches, the family of k-means algorithms and min-cut algorithms gain most popularity due to their simplicity and efficacy. The classical k-means algorithm partitions a number of data points into several subsets by iteratively updating the clustering centers and the associated data points. By contrast, a weighted undirected graph is constructed in min-cut algorithms which partition the vertices of the graph into two sets. However, existing clustering algorithms tend to cluster minority of data points into a subset, which shall be avoided when the target dataset is balanced. To achieve more accurate clustering for balanced dataset, we propose to leverage exclusive lasso on k-means and min-cut to regulate the balance degree of the clustering results. By optimizing our objective functions that build atop the exclusive lasso, we can make the clustering result as much balanced as possible. Extensive experiments on several large-scale datasets validate the advantage of the proposed algorithms compared to the state-of-the-art clustering algorithms. Zhihui Li 0001, Feiping Nie 0001, Xiaojun Chang, Zhigang Ma, Yi Yang 0001 |
AAAI | 4 |
| 2018 | Joint Concept Correlation and Feature-Concept Relevance Learning for Multilabel ClassificationabstractIn recent years, multilabel classification has attracted significant attention in multimedia annotation. However, most of the multilabel classification methods focus only on the inherent correlations existing among multiple labels and concepts and ignore the relevance between features and the target concepts. To obtain more robust multilabel classification results, we propose a new multilabel classification method aiming to capture the correlations among multiple concepts by leveraging hypergraph that is proved to be beneficial for relational learning. Moreover, we consider mining feature-concept relevance, which is often overlooked by many multilabel learning algorithms. To better show the feature-concept relevance, we impose a sparsity constraint on the proposed method. We compare the proposed method with several other multilabel classification methods and evaluate the classification performance by mean average precision on several data sets. The experimental results show that the proposed method outperforms the state-of-the-art methods. Zhigang Ma |
Neural Comput. | 2 |
| 2018 | Joint Attributes and Event Analysis for Multimedia Event DetectionabstractSemantic attributes have been increasingly used the past few years for multimedia event detection (MED) with promising results. The motivation is that multimedia events generally consist of lower level components such as objects, scenes, and actions. By characterizing multimedia event videos with semantic attributes, one could exploit more informative cues for improved detection results. Much existing work obtains semantic attributes from images, which may be suboptimal for video analysis since these image-inferred attributes do not carry dynamic information that is essential for videos. To address this issue, we propose to learn semantic attributes from external videos using their semantic labels. We name them video attributes in this paper. In contrast with multimedia event videos, these external videos depict lower level contents such as objects, scenes, and actions. To harness video attributes, we propose an algorithm established on a correlation vector that correlates them to a target event. Consequently, we could incorporate video attributes latently as extra information into the event detector learnt from multimedia event videos in a joint framework. To validate our method, we perform experiments on the real-world large-scale TRECVID MED 2013 and 2014 data sets and compare our method with several state-of-the-art algorithms. The experiments show that our method is advantageous for MED. Zhigang Ma, Xiaojun Chang, Zhongwen Xu, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Bi-Level Semantic Representation Analysis for Multimedia Event DetectionabstractMultimedia event detection has been one of the major endeavors in video event analysis. A variety of approaches have been proposed recently to tackle this problem. Among others, using semantic representation has been accredited for its promising performance and desirable ability for human-understandable reasoning. To generate semantic representation, we usually utilize several external image/video archives and apply the concept detectors trained on them to the event videos. Due to the intrinsic difference of these archives, the resulted representation is presumable to have different predicting capabilities for a certain event. Notwithstanding, not much work is available for assessing the efficacy of semantic representation from the source-level. On the other hand, it is plausible to perceive that some concepts are noisy for detecting a specific event. Motivated by these two shortcomings, we propose a bi-level semantic representation analyzing method. Regarding source-level, our method learns weights of semantic representation attained from different multimedia archives. Meanwhile, it restrains the negative influence of noisy or irrelevant concepts in the overall concept-level. In addition, we particularly focus on efficient multimedia event detection with few positive examples, which is highly appreciated in the real-world scenario. We perform extensive experiments on the challenging TRECVID MED 2013 and 2014 datasets with encouraging results that validate the efficacy of our proposed approach. Xiaojun Chang, Zhigang Ma, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Cybern. | 2 |
| 2017 | Feature Interaction Augmented Sparse Learning for Fast Kinect Motion DetectionabstractThe Kinect sensing devices have been widely used in current Human-Computer Interaction entertainment. A fundamental issue involved is to detect users' motions accurately and quickly. In this paper, we tackle it by proposing a linear algorithm, which is augmented by feature interaction. The linear property guarantees its speed whereas feature interaction captures the higher order effect from the data to enhance its accuracy. The Schatten-p norm is leveraged to integrate the main linear effect and the higher order nonlinear effect by mining the correlation between them. The resulted classification model is a desirable combination of speed and accuracy. We propose a novel solution to solve our objective function. Experiments are performed on three public Kinect-based entertainment data sets related to fitness and gaming. The results show that our method has its advantage for motion detection in a real-time Kinect entertaining environment. Xiaojun Chang, Zhigang Ma, Ming Lin 0002, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | The Many Shades of NegativityabstractComplex event detection has been progressively researched in recent years for the broad interest of video indexing and retrieval. To fulfill the purpose of event detection, one needs to train a classifier using both positive and negative examples. Current classifier training treats the negative videos as equally negative. However, we notice that many negative videos resemble the positive videos in different degrees. Intuitively, we may capture more informative cues from the negative videos if we assign them fine-grained labels, thus benefiting the classifier learning. Aiming for this, we use a statistical method on both the positive and negative examples to get the decisive attributes of a specific event. Based on these decisive attributes, we assign the fine-grained labels to negative examples to treat them differently for more effective exploitation. The resulting fine-grained labels may be not optimal to capture the discriminative cues from the negative videos. Hence, we propose to jointly optimize the fine-grained labels with the classifier learning, which brings mutual reciprocality. Meanwhile, the labels of positive examples are supposed to remain unchanged. We thus additionally introduce a constraint for this purpose. On the other hand, the state-of-the-art deep convolutional neural network features are leveraged in our approach for event detection to further boost the performance. Extensive experiments on the challenging TRECVID MED 2014 dataset have validated the efficacy of our proposed approach. Zhigang Ma, Xiaojun Chang, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Special Issue on Event-based Media Processing and Analysis
Bogdan Ionescu, Giulia Boato, Zhigang Ma, Ioannis Kompatsiaris, Nicu Sebe, Shuicheng Yan |
Image Vis. Comput. | 3 |
| 2016 | Event-based media processing and analysis: A survey of the literatureabstractResearch on event-based processing and analysis of media is receiving an increasing attention from the scientific community due to its relevance for an abundance of applications, from consumer video management and video surveillance to lifelogging and social media.Events have the ability to semantically encode relationships of different informational modalities, such as visual-audio-text, time, involved agents and objects, with the spatio-temporal component of events being a key feature for contextual analysis.This unveils an enormous potential for exploiting new information sources and opening new research directions.In this paper, we survey the existing literature in this field.We extensively review the employed conceptualization of the notion of event in multimedia, the techniques for event representation and modeling, the feature representation and event inference approaches for the problems of event detection in audio, visual, and textual content.Furthermore, we review some key event-based multimedia applications, and various benchmarking activities that provide solid frameworks for measuring the performance of different event processing and analysis systems.We provide an in-depth discussion of the insights obtained from reviewing the literature and identify future directions and challenges. Christos Tzelepis, Zhigang Ma, Vasileios Mezaris, Bogdan Ionescu, Ioannis Kompatsiaris, Giulia Boato, Nicu Sebe, Shuicheng Yan |
Image Vis. Comput. | 2 |
| 2016 | Outlier correction method of telemetry data based on wavelet transformation and Wright criterion
Zhigang Ma |
Multim. Tools Appl. | 1 |
| 2016 | Guest Editorial: Representation Learning for Multimedia Data Understanding
Yan Yan 0002, Zhigang Ma, Bingbing Ni |
Multim. Tools Appl. | 2 |
| 2015 | A Convex Formulation for Spectral Shrunk ClusteringabstractSpectral clustering is a fundamental technique in the field of data mining and information processing. Most existing spectral clustering algorithms integrate dimensionality reduction into the clustering process assisted by manifold learning in the original space. However, the manifold in reduced-dimensional subspace is likely to exhibit altered properties in contrast with the original space. Thus, applying manifold information obtained from the original space to the clustering process in a low-dimensional subspace is prone to inferior performance. Aiming to address this issue, we propose a novel convex algorithm that mines the manifold structure in the low-dimensional subspace. In addition, our unified learning process makes the manifold learning particularly tailored for the clustering. Compared with other related methods, the proposed algorithm results in more structured clustering result. To validate the efficacy of the proposed algorithm, we perform extensive experiments on several benchmark datasets in comparison with some state-of-the-art clustering approaches. The experimental results demonstrate that the proposed algorithm has quite promising clustering performance. Xiaojun Chang, Feiping Nie 0001, Zhigang Ma, Yi Yang 0001, Xiaofang Zhou 0001 |
AAAI | 3 |
| 2015 | Multi-Class Active Learning by Uncertainty Sampling with Diversity Maximization
Yi Yang 0001, Zhigang Ma, Feiping Nie 0001, Xiaojun Chang, Alex Hauptmann 0001 |
Int. J. Comput. Vis. | 2 |
| 2015 | Multitask Spectral Clustering by Exploring Intertask CorrelationabstractClustering, as one of the most classical research problems in pattern recognition and data mining, has been widely explored and applied to various applications. Due to the rapid evolution of data on the Web, more emerging challenges have been posed on traditional clustering techniques: 1) correlations among related clustering tasks and/or within individual task are not well captured; 2) the problem of clustering out-of-sample data is seldom considered; and 3) the discriminative property of cluster label matrix is not well explored. In this paper, we propose a novel clustering model, namely multitask spectral clustering (MTSC), to cope with the above challenges. Specifically, two types of correlations are well considered: 1) intertask clustering correlation, which refers the relations among different clustering tasks and 2) intratask learning correlation, which enables the processes of learning cluster labels and learning mapping function to reinforce each other. We incorporate a novel l2,p -norm regularizer to control the coherence of all the tasks based on an assumption that related tasks should share a common low-dimensional representation. Moreover, for each individual task, an explicit mapping function is simultaneously learnt for predicting cluster labels by mapping features to the cluster label matrix. Meanwhile, we show that the learning process can naturally incorporate discriminative information to further improve clustering performance. We explore and discuss the relationships between our proposed model and several representative clustering techniques, including spectral clustering, k -means and discriminative k -means. Extensive experiments on various real-world datasets illustrate the advantage of the proposed MTSC model compared to state-of-the-art clustering approaches. Yang Yang 0002, Zhigang Ma, Yi Yang 0001, Feiping Nie 0001, Heng Tao Shen |
IEEE Trans. Cybern. | 2 |
| 2015 | Semisupervised Feature Selection via Spline Regression for Video Semantic RecognitionabstractTo improve both the efficiency and accuracy of video semantic recognition, we can perform feature selection on the extracted video features to select a subset of features from the high-dimensional feature set for a compact and accurate video data representation. Provided the number of labeled videos is small, supervised feature selection could fail to identify the relevant features that are discriminative to target classes. In many applications, abundant unlabeled videos are easily accessible. This motivates us to develop semisupervised feature selection algorithms to better identify the relevant video features, which are discriminative to target classes by effectively exploiting the information underlying the huge amount of unlabeled video data. In this paper, we propose a framework of video semantic recognition by semisupervised feature selection via spline regression (S(2)FS(2)R) . Two scatter matrices are combined to capture both the discriminative information and the local geometry structure of labeled and unlabeled training videos: A within-class scatter matrix encoding discriminative information of labeled training videos and a spline scatter output from a local spline regression encoding data distribution. An l2,1 -norm is imposed as a regularization term on the transformation matrix to ensure it is sparse in rows, making it particularly suitable for feature selection. To efficiently solve S(2)FS(2)R , we develop an iterative algorithm and prove its convergency. In the experiments, three typical tasks of video semantic recognition, such as video concept detection, video classification, and human action recognition, are used to demonstrate that the proposed S(2)FS(2)R achieves better performance compared with the state-of-the-art methods. Yahong Han, Yi Yang 0001, Yan Yan 0006, Zhigang Ma, Nicu Sebe, Xiaofang Zhou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Event Detection Using Multi-level Relevance Labels and Multiple FeaturesabstractWe address the challenging problem of utilizing related exemplars for complex event detection while multiple features are available. Related exemplars share certain positive elements of the event, but have no uniform pattern due to the huge variance of relevance levels among different related exemplars. None of the existing multiple feature fusion methods can deal with the related exemplars. In this paper, we propose an algorithm which adaptively utilizes the related exemplars by cross-feature learning. Ordinal labels are used to represent the multiple relevance levels of the related videos. Label candidates of related exemplars are generated by exploring the possible relevance levels of each related exemplar via a cross-feature voting strategy. Maximum margin criterion is then applied in our framework to discriminate the positive and negative exemplars, as well as the related exemplars from different relevance levels. We test our algorithm using the large scale TRECVID 2011 dataset and it gains promising performance. Zhongwen Xu, Ivor W. Tsang, Yi Yang 0001, Zhigang Ma, Alex Hauptmann 0001 |
CVPR | 4 |
| 2014 | Multiple Features But Few Labels?: A Symbiotic Solution Exemplified for Video AnalysisabstractVideo analysis has been attracting increasing research due to the proliferation of internet videos. In this paper, we investigate how to improve the performance on internet quality video analysis. Particularly, we work on the scenario of few labeled training videos being provided, which is less focused in multimedia. To being with, we consider how to more effectively harness the evidences from the low-level features. Researchers have developed several promising features to represent videos to capture the semantic information. However, as videos usually characterize rich semantic contents, the analysis performance by using one single feature is potentially limited. Simply combining multiple features through early fusion or late fusion to incorporate more informative cues is doable but not optimal due to the heterogeneity and different predicting capability of these features. For better exploitation of multiple features, we propose to mine the importance of different features and cast it into the learning of the classification model. Our method is based on multiple graphs from different features and uses the Riemannian metric to evaluate the feature importance. On the other hand, to be able to use limited labeled training videos for a respectable accuracy we formulate our method in a semi-supervised way. The main contribution of this paper is a novel scheme of evaluating the feature importance that is further casted into a unified framework of harnessing multiple weighted features with limited labeled training videos. We perform extensive experiments on video action recognition and multimedia event recognition and the comparison to other state-of-the-art multi-feature learning algorithms has validated the efficacy of our framework. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2014 | GLocal tells you more: Coupling GLocal structural for feature selection with sparsity for image and video classification
Yan Yan 0002, Haoquan Shen, Gaowen Liu, Zhigang Ma, Chenqiang Gao, Nicu Sebe |
Comput. Vis. Image Underst. | 4 |
| 2014 | Harnessing Lab Knowledge for Real-World Action Recognition
Zhigang Ma, Yi Yang 0001, Feiping Nie 0001, Nicu Sebe, Shuicheng Yan, Alex Hauptmann 0001 |
Int. J. Comput. Vis. | 1 |
| 2014 | E-LAMP: integration of innovative ideas for multimedia event detection
Yi Yang 0001, Lu Jiang 0004, Shoou-I Yu, Zhen-Zhong Lan, Zhigang Ma, Waito Sze, Ehsan Younessian, Alex Hauptmann 0001 |
Mach. Vis. Appl. | 6 |
| 2014 | Knowledge Adaptation with PartiallyShared Features for Event DetectionUsing Few ExemplarsabstractMultimedia event detection (MED) is an emerging area of research. Previous work mainly focuses on simple event detection in sports and news videos, or abnormality detection in surveillance videos. In contrast, we focus on detecting more complicated and generic events that gain more users' interest, and we explore an effective solution for MED. Moreover, our solution only uses few positive examples since precisely labeled multimedia content is scarce in the real world. As the information from these few positive examples is limited, we propose using knowledge adaptation to facilitate event detection. Different from the state of the art, our algorithm is able to adapt knowledge from another source for MED even if the features of the source and the target are partially different, but overlapping. Avoiding the requirement that the two domains are consistent in feature types is desirable as data collection platforms change or augment their capabilities and we should be able to respond to this with little or no effort. We perform extensive experiments on real-world multimedia archives consisting of several challenging events. The results show that our approach outperforms several other state-of-the-art detection algorithms. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Image Attribute AdaptationabstractVisual attributes can be considered as a middle-level semantic cue that bridges the gap between low-level image features and high-level object classes. Thus, attributes have the advantage of transcending specific semantic categories or describing objects across categories. Since attributes are often human-nameable and domain specific, much work constructs attribute annotations ad hoc or take them from an application-dependent ontology. To facilitate other applications with attributes, it is necessary to develop methods which can adapt a well-defined set of attributes to novel images. In this paper, we propose a framework for image attribute adaptation. The goal is to automatically adapt the knowledge of attributes from a well-defined auxiliary image set to a target image set, thus assisting in predicting appropriate attributes for target images. In the proposed framework, we use a non-linear mapping function corresponding to multiple base kernels to map each training images of both the auxiliary and the target sets to a Reproducing Kernel Hilbert Space (RKHS), where we reduce the mismatch of data distributions between auxiliary and target images. In order to make use of un-labeled images, we incorporate a semi-supervised learning process. We also introduce a robust loss function into our framework to remove the shared irrelevance and noise of training images. Experiments on two couples of auxiliary-target image sets demonstrate that the proposed framework has better performance of predicting attributes for target testing images, compared to three baselines and two state-of-the-art domain adaptation methods. Yahong Han, Yi Yang 0001, Zhigang Ma, Haoquan Shen, Nicu Sebe, Xiaofang Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Semi-Supervised Multiple Feature Analysis for Action RecognitionabstractThis paper presents a semi-supervised method for categorizing human actions using multiple visual features. The proposed algorithm simultaneously learns multiple features from a small number of labeled videos, and automatically utilizes data distributions between labeled and unlabeled data to boost the recognition performance. Shared structural analysis is applied in our approach to discover a common subspace shared by each type of feature. In the subspace, the proposed algorithm is able to characterize more discriminative information of each feature type. Additionally, data distribution information of each type of feature has been preserved. The aforementioned attributes make our algorithm robust for action recognition, especially when only limited labeled training samples are provided. Extensive experiments have been conducted on both the choreographed and the realistic video datasets, including KTH, Youtube action and UCF50. Experimental results show that our method outperforms several state-of-the-art algorithms. Most notably, much better performances have been achieved when there are only a few labeled training samples. Sen Wang 0001, Zhigang Ma, Yi Yang 0001, Xue Li 0001, Chaoyi Pang, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | Complex Event Detection via Multi-source Video AttributesabstractComplex events essentially include human, scenes, objects and actions that can be summarized by visual attributes, so leveraging relevant attributes properly could be helpful for event detection. Many works have exploited attributes at image level for various applications. However, attributes at image level are possibly insufficient for complex event detection in videos due to their limited capability in characterizing the dynamic properties of video data. Hence, we propose to leverage attributes at video level (named as video attributes in this work), i.e., the semantic labels of external videos are used as attributes. Compared to complex event videos, these external videos contain simple contents such as objects, scenes and actions which are the basic elements of complex events. Specifically, building upon a correlation vector which correlates the attributes and the complex event, we incorporate video attributes latently as extra informative cues into the event detector learnt from complex event videos. Extensive experiments on a real-world large-scale dataset validate the efficacy of the proposed approach. Zhigang Ma, Yi Yang 0001, Zhongwen Xu, Shuicheng Yan, Nicu Sebe, Alex Hauptmann 0001 |
CVPR | 1 |
| 2013 | How Related Exemplars Help Complex Event Detection in Web Videos?abstractCompared to visual concepts such as actions, scenes and objects, complex event is a higher level abstraction of longer video sequences. For example, a "marriage proposal" event is described by multiple objects (e.g., ring, faces), scenes (e.g., in a restaurant, outdoor) and actions (e.g., kneeling down). The positive exemplars which exactly convey the precise semantic of an event are hard to obtain. It would be beneficial to utilize the related exemplars for complex event detection. However, the semantic correlations between related exemplars and the target event vary substantially as relatedness assessment is subjective. Two related exemplars can be about completely different events, e.g., in the TRECVID MED dataset, both bicycle riding and equestrianism are labeled as related to "attempting a bike trick" event. To tackle the subjectiveness of human assessment, our algorithm automatically evaluates how positive the related exemplars are for the detection of an event and uses them on an exemplar-specific basis. Experiments demonstrate that our algorithm is able to utilize related exemplars adaptively, and the algorithm gains good performance for complex event detection. Yi Yang 0001, Zhigang Ma, Zhongwen Xu, Shuicheng Yan, Alex Hauptmann 0001 |
ICCV | 2 |
| 2013 | Thinking of Images as What They Are: Compound Matrix Regression for Image Classification
Zhigang Ma, Yi Yang 0001, Feiping Nie 0001, Nicu Sebe |
IJCAI | 1 |
| 2013 | We are not equally negative: fine-grained labeling for multimedia event detectionabstractMultimedia event detection (MED) is an effective technique for video indexing and retrieval. Current classifier training for MED treats the negative videos equally. However, many negative videos may resemble the positive videos in different degrees. Intuitively, we may capture more informative cues from the negative videos if we assign them fine-grained labels, thus benefiting the classifier learning. Aiming for this, we use a statistical method on both the positive and negative examples to get the decisive attributes of a specific event. Based on these decisive attributes, we assign the fine-grained labels to negative examples to treat them differently for more effective exploitation. The resulting fine-grained labels may be not accurate enough to characterize the negative videos. Hence, we propose to jointly optimize the fine-grained labels with the knowledge from the visual features and the attributes representations, which brings mutual reciprocality. Our model obtains two kinds of classifiers, one from the attributes and one from the features, which incorporate the informative cues from the fine-grained labels. The outputs of both classifiers on the testing videos are fused for detection. Extensive experiments on the challenging TRECVID MED 2012 development set have validated the efficacy of our proposed approach. Zhigang Ma, Yi Yang 0001, Zhongwen Xu, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2013 | GLocal structural feature selection with sparsity for multimedia data understandingabstractThe selection of discriminative features is an important and effective technique for many multimedia tasks. Using irrelevant features in classification or clustering tasks could deteriorate the performance. Thus, designing efficient feature selection algorithms to remove the irrelevant features is a possible way to improve the classification or clustering performance. With the successful usage of sparse models in image and video classification and understanding, imposing structural sparsity in \emph{feature selection} has been widely investigated during the past years. Motivated by the merit of sparse models, we propose a novel feature selection method using a sparse model in this paper. Different from the state of the art, our method is built upon $\ell _{2,p}$-norm and simultaneously considers both the global and local (GLocal) structures of data distribution. Our method is more flexible in selecting the discriminating features as it is able to control the degree of sparseness. Moreover, considering both global and local structures of data distribution makes our feature selection process more effective. An efficient algorithm is proposed to solve the $\ell_{2,p}$-norm sparsity optimization problem in this paper. Experimental results performed on real-world image and video datasets show the effectiveness of our feature selection method compared to several state-of-the-art methods. Yan Yan 0002, Zhongwen Xu, Gaowen Liu, Zhigang Ma, Nicu Sebe |
ACM Multimedia | 4 |
| 2013 | Fusing inherent and external knowledge with nonlinear learning for cross-media retrieval
Yun Liu 0021, Zhigang Ma |
Neurocomputing | 3 |
| 2013 | Image classification with manifold learning for out-of-sample data
Yahong Han, Zhongwen Xu, Zhigang Ma, Zi Huang |
Signal Process. | 3 |
| 2013 | Multimedia Event Detection Using A Classifier-Specific Intermediate RepresentationabstractMultimedia event detection (MED) plays an important role in many applications such as video indexing and retrieval. Current event detection works mainly focus on sports and news event detection or abnormality detection in surveillance videos. Differently, our research aims to detect more complicated and generic events within a longer video sequence. In the past, researchers have proposed using intermediate concept classifiers with concept lexica to help understand the videos. Yet it is difficult to judge how many and what concepts would be sufficient for the particular video analysis task. Additionally, obtaining robust semantic concept classifiers requires a large number of positive training examples, which in turn has high human annotation cost. In this paper, we propose an approach that exploits the external concepts-based videos and event-based videos simultaneously to learn an intermediate representation from video features. Our algorithm integrates the classifier inference and latent intermediate representation into a joint framework. The joint optimization of the intermediate representation and the classifier makes them mutually beneficial and reciprocal. Effectively, the intermediate representation and the classifier are tightly correlated. The classifier dependent intermediate representation not only accurately reflects the task semantics but is also more suitable for the specific classifier. Thus we have created a discriminative semantic analysis framework based on a tightly coupled intermediate representation. Extensive experiments on multimedia event detection using real-world videos demonstrate the effectiveness of the proposed approach. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Kai Zheng 0001, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | Feature Selection for Multimedia Analysis by Sharing Information Among Multiple TasksabstractWhile much progress has been made to multi-task classification and subspace learning, multi-task feature selection has long been largely unaddressed. In this paper, we propose a new multi-task feature selection algorithm and apply it to multimedia (e.g., video and image) analysis. Instead of evaluating the importance of each feature individually, our algorithm selects features in a batch mode, by which the feature correlation is considered. While feature selection has received much research attention, less effort has been made on improving the performance of feature selection by leveraging the shared knowledge from multiple related tasks. Our algorithm builds upon the assumption that different related tasks have common structures. Multiple feature selection functions of different tasks are simultaneously learned in a joint framework, which enables our algorithm to utilize the common knowledge of multiple tasks as supplementary information to facilitate decision making. An efficient iterative algorithm is proposed to optimize it, whose convergence is guaranteed. Experiments on different databases have demonstrated the effectiveness of the proposed algorithm. Yi Yang 0001, Zhigang Ma, Alex Hauptmann 0001, Nicu Sebe |
IEEE Trans. Multim. | 2 |
| 2013 | Multi-Feature Fusion via Hierarchical Regression for Multimedia AnalysisabstractMultimedia data are usually represented by multiple features. In this paper, we propose a new algorithm, namely Multi-feature Learning via Hierarchical Regression for multimedia semantics understanding, where two issues are considered. First, labeling large amount of training data is labor-intensive. It is meaningful to effectively leverage unlabeled data to facilitate multimedia semantics understanding. Second, given that multimedia data can be represented by multiple features, it is advantageous to develop an algorithm which combines evidence obtained from different features to infer reliable multimedia semantic concept classifiers. We design a hierarchical regression model to exploit the information derived from each type of feature, which is then collaboratively fused to obtain a multimedia semantic concept classifier. Both label information and data distribution of different features representing multimedia data are considered. The algorithm can be applied to a wide range of multimedia applications and experiments are conducted on video data for video concept annotation and action recognition. Using Trecvid and CareMedia video datasets, the experimental results show that it is beneficial to combine multiple features. The performance of the proposed algorithm is remarkable when only a small amount of labeled training data are available. Yi Yang 0001, Jingkuan Song, Zi Huang, Zhigang Ma, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 4 |
| 2012 | Action recognition by exploring data distribution and feature correlationabstractHuman action recognition in videos draws strong research interest in computer vision because of its promising applications for video surveillance, video annotation, interactive gaming, etc. However, the amount of video data containing human actions is increasing exponentially, which makes the management of these resources a challenging task. Given a database with huge volumes of unlabeled videos, it is prohibitive to manually assign specific action types to these videos. Considering that it is much easier to obtain a small number of labeled videos, a practical solution for organizing them is to build a mechanism which is able to conduct action annotation automatically by leveraging the limited labeled videos. Motivated by this intuition, we propose an automatic video annotation algorithm by integrating semi-supervised learning and shared structure analysis into a joint framework for human action recognition. We apply our algorithm on both synthetic and realistic video datasets, including KTH [20], CareMedia dataset [1], Youtube action [12] and its extended version, UCF50 [2]. Extensive experiments demonstrate that the proposed algorithm outperforms the compared algorithms for action recognition. Most notably, our method has a very distinct advantage over other compared algorithms when we have only a few labeled samples. Sen Wang 0001, Yi Yang 0001, Zhigang Ma, Xue Li 0001, Chaoyi Pang, Alex Hauptmann 0001 |
CVPR | 3 |
| 2012 | Classifier-specific intermediate representation for multimedia tasksabstractVideo annotation and multimedia classification play important roles in many applications such as video indexing and retrieval. To improve video annotation and event detection, researchers have proposed using intermediate concept classifiers with concept lexica to help understand the videos. Yet it is difficult to judge how many and what concepts would be sufficient for the particular video analysis task. Additionally, obtaining robust semantic concept classifiers requires a large number of positive training examples, which in turn has high human annotation cost. In this paper, we propose an approach that is able to automatically learn an intermediate representation from video features together with a classifier. The joint optimization of the two components makes them mutually beneficial and reciprocal. Effectively, the intermediate representation and the classifier are tightly correlated. The classifier dependent intermediate representation not only accurately reflects the task semantics but is also more suitable for the specific classifier. Thus we have created a discriminative semantic analysis framework based on a tightly-coupled intermediate representation. Several experiments on video annotation and multimedia event detection using real-world videos demonstrate the effectiveness of the proposed approach. Zhigang Ma, Alex Hauptmann 0001, Yi Yang 0001, Nicu Sebe |
ICMR | 1 |
| 2012 | Knowledge adaptation for ad hoc multimedia event detection with few exemplarsabstractMultimedia event detection (MED) has a significant impact on many applications. Though video concept annotation has received much research effort, video event detection remains largely unaddressed. Current research mainly focuses on sports and news event detection or abnormality detection in surveillance videos. Our research on this topic is capable of detecting more complicated and generic events. Moreover, the curse of reality, i.e., precisely labeled multimedia content is scarce, necessitates the study on how to attain respectable detection performance using only limited positive examples. Research addressing these two aforementioned issues is still in its infancy. In light of this, we explore Ad Hoc MED, which aims to detect complicated and generic events by using few positive examples. To the best of our knowledge, our work makes the first attempt on this topic. As the information from these few positive examples is limited, we propose to infer knowledge from other multimedia resources to facilitate event detection. Experiments are performed on real-world multimedia archives consisting of several challenging events. The results show that our approach outperforms several other detection algorithms. Most notably, our algorithm outperforms SVM by 43% and 14% comparatively in Average Precision when using Gaussian and Χ2 kernel respectively. Zhigang Ma, Yi Yang 0001, Yang Cai 0002, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2012 | Robust cross-media transfer for visual event detectionabstractIn this paper, we present a novel approach, named Robust Cross-Media Transfer (RCMT), for visual event detection in social multimedia environments. Different from most existing methods, the proposed method can directly take different types of noisy social multimedia data as input and conduct robust event detection. More specifically, we build a robust model by employing an l2,1-norm regression model featuring noise tolerance, and also manage to integrate different types of social multimedia data by minimizing the distribution difference among them. Experimental results on real-life Flickr image dataset and YouTube video dataset demonstrate the effectiveness of our proposal, compared to state-of-the-art algorithms. Yang Yang 0002, Yi Yang 0001, Zi Huang, Jiajun Liu 0004, Zhigang Ma |
ACM Multimedia | 5 |
| 2012 | Web Image Annotation Via Subspace-Sparsity Collaborated Feature SelectionabstractThe number of web images has been explosively growing due to the development of network and storage technology. These images make up a large amount of current multimedia data and are closely related to our daily life. To efficiently browse, retrieve and organize the web images, numerous approaches have been proposed. Since the semantic concepts of the images can be indicated by label information, automatic image annotation becomes one effective technique for image management tasks. Most existing annotation methods use image features that are often noisy and redundant. Hence, feature selection can be exploited for a more precise and compact representation of the images, thus improving the annotation performance. In this paper, we propose a novel feature selection method and apply it to automatic image annotation. There are two appealing properties of our method. First, it can jointly select the most relevant features from all the data points by using a sparsity-based model. Second, it can uncover the shared subspace of original features, which is beneficial for multi-label learning. To solve the objective function of our method, we propose an efficient iterative algorithm. Extensive experiments are performed on large image databases that are collected from the web. The experimental results together with the theoretical analysis have validated the effectiveness of our method for feature selection, thus demonstrating its feasibility of being applied to web image annotation. Zhigang Ma, Feiping Nie 0001, Yi Yang 0001, Jasper R. R. Uijlings, Nicu Sebe |
IEEE Trans. Multim. | 1 |
| 2012 | Discriminating Joint Feature Analysis for Multimedia Data UnderstandingabstractIn this paper, we propose a novel semi-supervised feature analyzing framework for multimedia data understanding and apply it to three different applications: image annotation, video concept detection and 3-D motion data analysis. Our method is built upon two advancements of the state of the art: (1)l2, 1-norm regularized feature selection which can jointly select the most relevant features from all the data points. This feature selection approach was shown to be robust and efficient in literature as it considers the correlation between different features jointly when conducting feature selection; (2) manifold learning which analyzes the feature space by exploiting both labeled and unlabeled data. It is a widely used technique to extend many algorithms to semi-supervised scenarios for its capability of leveraging the manifold structure of multimedia data. The proposed method is able to learn a classifier for different applications by selecting the discriminating features closely related to the semantic concepts. The objective function of our method is non-smooth and difficult to solve, so we design an efficient iterative algorithm with fast convergence, thus making it applicable to practical applications. Extensive experiments on image annotation, video concept detection and 3-D motion data analysis are performed on different real-world data sets to demonstrate the effectiveness of our algorithm. Zhigang Ma, Feiping Nie 0001, Yi Yang 0001, Jasper R. R. Uijlings, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | l2, 1-Norm Regularized Discriminative Feature Selection for Unsupervised LearningabstractCompared with supervised learning for feature selection, it is much more difficult to select the discriminative features in unsupervised learning due to the lack of label information. Traditional unsupervised feature selection algorithms usually select the features which best preserve the data distribution, e.g., manifold structure, of the whole feature set. Under the assumption that the class label of input data can be predicted by a linear classifier, we incorporate discriminative analysis and ℓ2,1-norm minimization into a joint framework for unsupervised feature selection. Different from existing unsupervised feature selection algorithms, our algorithm selects the most discriminative feature subset from the whole feature set in batch mode. Extensive experiment on different data types demonstrates the effectiveness of our algorithm. Yi Yang 0001, Heng Tao Shen, Zhigang Ma, Zi Huang, Xiaofang Zhou 0001 |
IJCAI | 3 |
| 2011 | Exploiting the entire feature space with sparsity for automatic image annotationabstractThe explosive growth of digital images requires effective methods to manage these images. Among various existing methods, automatic image annotation has proved to be an important technique for image management tasks, e.g., image retrieval over large-scale image databases. Automatic image annotation has been widely studied during recent years and a considerable number of approaches have been proposed. However, the performance of these methods is yet to be satisfactory, thus demanding more effort on research of image annotation. In this paper, we propose a novel semi supervised framework built upon feature selection for automatic image annotation. Our method aims to jointly select the most relevant features from all the data points by using a sparsity-based model and exploiting both labeled and unlabeled data to learn the manifold structure. Our framework is able to simultaneously learn a robust classifier for image annotation by selecting the discriminating features related to the semantic concepts. To solve the objective function of our framework, we propose an efficient iterative algorithm. Extensive experiments are performed on different real-world image datasets with the results demonstrating the promising performance of our framework for automatic image annotation. Zhigang Ma, Yi Yang 0001, Feiping Nie 0001, Jasper R. R. Uijlings, Nicu Sebe |
ACM Multimedia | 1 |