Senjian An

dblp:17/3718 · DBLP profile ↗
← Back
52ranked-venue papers
16as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 14 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 8 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Video understanding and tracking · 26% Image recognition and object detection · 23% Learning theory · 11%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › multi-graph learning
multi-view graph learning
0.712023
Multi-View Diffusion Process for Spectral Clustering and Image Retrieval · IEEE Trans. Image Process. 2023
Data mining
clustering
0.712023
Multi-View Diffusion Process for Spectral Clustering and Image Retrieval · IEEE Trans. Image Process. 2023
Data mining › clustering
spectral clustering
0.712023
Multi-View Diffusion Process for Spectral Clustering and Image Retrieval · IEEE Trans. Image Process. 2023
Multimedia analysis and retrieval
image retrieval
0.712023
Multi-View Diffusion Process for Spectral Clustering and Image Retrieval · IEEE Trans. Image Process. 2023
Computer vision › Face, body and person analysis
face recognition
0.652015
Deep Reconstruction Models for Image Set Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Learning Non-linear Reconstruction Models for Image Set Classification · CVPR 2014
Exploiting side information in locality preserving projection · CVPR 2008
Computer vision › Video understanding and tracking
action recognition
0.622018
Learning Clip Representations for Skeleton-Based 3D Action Recognition · IEEE Trans. Image Process. 2018
A New Representation of Skeleton Sequences for 3D Action Recognition · CVPR 2017
Computer vision › Image recognition and object detection › image classification
image set classification
0.632015
Deep Reconstruction Models for Image Set Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Reverse Training: An Efficient Approach for Image Set Classification · ECCV (6) 2014
Learning Non-linear Reconstruction Models for Image Set Classification · CVPR 2014
Machine learning › Deep learning architectures and training
convolutional neural network
0.532018
A New Representation of Skeleton Sequences for 3D Action Recognition · CVPR 2017
Learning Clip Representations for Skeleton-Based 3D Action Recognition · IEEE Trans. Image Process. 2018
A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification · IEEE Trans. Image Process. 2016
Computer vision › Video understanding and tracking
action anticipation
0.412020
Learning Latent Global Network for Skeleton-Based Action Prediction · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking › action recognition › skeleton-based action recognition
3d skeleton-based action recognition
0.312018
Learning Clip Representations for Skeleton-Based 3D Action Recognition · IEEE Trans. Image Process. 2018
Computer vision › Video understanding and tracking › video analytics › behavior analysis › human behavior analysis
human interaction prediction
0.312018
Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction · IEEE Trans. Multim. 2018
Computer vision › Image recognition and object detection
object detection
0.332011
Efficient subwindow search with submodular score functions · CVPR 2011
Exploiting Monge structures in optimum subwindow search · CVPR 2010
Efficient algorithms for subwindow search in object detection and localization · CVPR 2009
Machine learning › Learning paradigms
multi-task learning
0.312017
A New Representation of Skeleton Sequences for 3D Action Recognition · CVPR 2017
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.312017
A New Representation of Skeleton Sequences for 3D Action Recognition · CVPR 2017
Computer vision › Image recognition and object detection › scene recognition
indoor scene recognition
0.212016
A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification · IEEE Trans. Image Process. 2016
Computer vision › Image recognition and object detection
scene recognition
0.212016
A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification · IEEE Trans. Image Process. 2016
Computer vision › Image recognition and object detection › object detection
subwindow search
0.222011
Efficient subwindow search with submodular score functions · CVPR 2011
Efficient algorithms for subwindow search in object detection and localization · CVPR 2009
Computer vision › Face, body and person analysis › face recognition › face matching
image set-based face recognition
0.212015
Deep Reconstruction Models for Image Set Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Machine learning › Learning theory
linear separability
0.212015
How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances? · ICML 2015
Machine learning › Learning theory
margin maximization
0.212015
Contractive Rectifier Networks for Nonlinear Maximum Margin Classification · ICCV 2015
Machine learning › Kernel, tree and ensemble methods › large margin methods
maximum margin classifiers
0.212015
Contractive Rectifier Networks for Nonlinear Maximum Margin Classification · ICCV 2015
Machine learning › Learning theory › classification
neural network classifier
0.212015
Contractive Rectifier Networks for Nonlinear Maximum Margin Classification · ICCV 2015
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.112020
Learning Latent Global Network for Skeleton-Based Action Prediction · IEEE Trans. Image Process. 2020
Mathematical optimization
submodular optimization
0.112011
Efficient subwindow search with submodular score functions · CVPR 2011
Human-robot interaction
interaction anticipation
0.112018
Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction · IEEE Trans. Multim. 2018
Computer vision › Image recognition and object detection › object localization
efficient subwindow search
0.112009
Efficient algorithms for subwindow search in object detection and localization · CVPR 2009
Computer vision › Image recognition and object detection
object localization
0.112009
Efficient algorithms for subwindow search in object detection and localization · CVPR 2009
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal feature learning
0.112017
A New Representation of Skeleton Sequences for 3D Action Recognition · CVPR 2017
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.112008
Exploiting side information in locality preserving projection · CVPR 2008
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
locality preserving projection
0.112008
Exploiting side information in locality preserving projection · CVPR 2008

Methods — techniques the papers use, named apart from their topics

weight learning · 2.0tensor product graph · 2.0diffusion process · 2.0alternating optimization · 2.0long short-term memory · 0.7latent global network · 0.4adversarial learning · 0.4gaussian restricted boltzmann machines · 0.4ranking score fusion · 0.3optical flow · 0.3multi-task convolutional neural network · 0.3CLIP representation · 0.3branch-and-bound · 0.2submodular bound function · 0.1monge property · 0.1monge structure · 0.1alternating column and row search · 0.1
YearPublicationVenuePosition
2024 Combining Raft-Based Stereo Disparity and Optical Flow Models For Scene Flow Estimation
abstract
In stereo-based scene flow estimation, two problems are often addressed in tandem—optical flow and stereo depth estimation. Both of these problems require dense point matching within distinct search domains. Despite their similarities, more investigation has yet to be conducted on improving optical flow performance through stereo datasets and vice versa. This paper introduces a conjoined network for optical flow and stereo disparity estimation based on the Recurrent All-Pairs Field Transforms (RAFT) architecture. The network features a shared backbone, reducing parameters and facilitating joint training on stereo disparity and optical flow data. Our joint model surpasses the baseline RAFT and RAFT-Stereo models, demonstrating that the two dense matching tasks can be effectively addressed using the same encoded features. Experiments show that training on data for each task improves the model’s performance on the other task. The joint model offers the advantage of training with more data to improve the encoder.
Huizhu Pan, Ling Li 0006, Senjian An, Hui Xie 0002
ICIP3
2024 EdgeConvFormer: An Unsupervised Anomaly Detection Method for Multivariate Time Series
Jie Liu 0076, Senjian An, Bradley Ezard, Ling Li 0006
ICPR (4)3
2024 Invertible Residual Blocks in Deep Learning Networks
abstract
Residual blocks have been widely used in deep learning networks. However, information may be lost in residual blocks due to the relinquishment of information in rectifier linear units (ReLUs). To address this issue, invertible residual networks have been proposed recently but are generally under strict restrictions which limit their applications. In this brief, we investigate the conditions under which a residual block is invertible. A sufficient and necessary condition is presented for the invertibility of residual blocks with one layer of ReLU inside the block. In particular, for widely used residual blocks with convolutions, we show that such residual blocks are invertible under weak conditions if the convolution is implemented with certain zero-padding methods. Inverse algorithms are also proposed, and experiments are conducted to show the effectiveness of the proposed inverse algorithms and prove the correctness of the theoretical results.
Ruhua Wang, Senjian An, Wanquan Liu, Ling Li 0006
IEEE Trans. Neural Networks Learn. Syst.2
2023 Multi-View Diffusion Process for Spectral Clustering and Image Retrieval
abstract
This paper presents a novel approach to multi-view graph learning that combines weight learning and graph learning in an alternating optimization framework. Multi-view graph learning refers to the problem of constructing a unified affinity graph using heterogeneous sources of data representation, which is a popular technique in many learning systems where no prior knowledge of data distribution is available. Our approach is based on a fusion-and-diffusion strategy, in which multiple affinity graphs are fused together via a weight learning scheme based on the unsupervised graph smoothness and utilised as a consensus prior to the diffusion. We propose a novel multi-view diffusion process that learns a manifold-aware affinity graph by propagating affinities on tensor product graphs, leveraging high-order contextual information to enhance pairwise affinities. In contrast to existing multi-view graph learning approaches, our approach is not limited by the quality of initial graphs or the assumption of a latent common subspace among multiple views. Instead, our approach is able to identify the consistency among views and fuse multiple graphs adaptively. We formulate both weight learning and diffusion-based affinity learning in a unified framework and propose an alternating optimization solver that is guaranteed to converge. The proposed approach is applied to image retrieval and clustering tasks on 16 real-world datasets. Extensive experimental results demonstrate that our approach outperforms state-of-the-art methods for both retrieval and clustering on 13 out of 16 datasets.
Senjian An, Ling Li 0006, Wanquan Liu, Yanda Shao
IEEE Trans. Image Process.2
2022 Unsupervised learning of multi-task deep variational model
Ling Li 0006, Wanquan Liu, Senjian An, Kylie Munyard
J. Vis. Commun. Image Represent.4
2021 Semisupervised Learning on Graphs With an Alternating Diffusion Process
abstract
Graph-based semisupervised learning is of great importance in many effective learning systems, particularly in agnostic settings where no parametric information or other prior knowledge about the data distribution is available. It leverages the graph structure to propagate labels from labeled nodes to unlabeled ones. Two separate stages are usually involved: constructing an affinity graph and propagating labels on the graph for transductive inference. It is suboptimal to manage them independently, as the correlation between the affinity graph and the labels would not be fully exploited. In this article, we integrate these two stages into one unified framework by formulating the graph construction as a regularized function estimation problem, similar to label propagation. We then propose an alternating diffusion process to solve them alternately, which allows us to learn the graph and unknown labels in an iterative fashion. With the proposed framework, we can construct a dynamic graph adapted to the given and predicted labels iteratively, resulting in more accurate and robust label propagation performance. Extensive experiments on synthetic data and various real-world data have demonstrated the superiority of the proposed method compared with other state-of-the-art methods.
Senjian An, Wanquan Liu, Ling Li 0006
IEEE Trans. Neural Networks Learn. Syst.2
2020 Haze pollution causality mining and prediction based on multi-dimensional time series with PS-FCM
Wanquan Liu, Senjian An
Inf. Sci.3
2020 ResFeats: Residual network based features for underwater image classification
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
Image Vis. Comput.3
2020 Learning Latent Global Network for Skeleton-Based Action Prediction
abstract
Human actions represented with 3D skeleton sequences are robust to clustered backgrounds and illumination changes. In this paper, we investigate skeleton-based action prediction, which aims to recognize an action from a partial skeleton sequence that contains incomplete action information. We propose a new Latent Global Network based on adversarial learning for action prediction. We demonstrate that the proposed network provides latent long-term global information that is complementary to the local action information of the partial sequences and helps improve action prediction. We show that action prediction can be improved by combining the latent global information with the local action information. We test the proposed method on three challenging skeleton datasets and report state-of-the-art performance.
Qiuhong Ke, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Image Process.4
2019 Efficient Gaussian Distance Transforms for Image Processing
Senjian An, Wanquan Liu, Ling Li 0006
ADMA1
2019 Improved Algorithms for Zero Shot Image Super-Resolution with Parametric Rectifiers
Senjian An, Wanquan Liu, Ling Li 0006
ADMA2
2019 An Improved Approach to Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with image-level labels is of great significance since it alleviates the dependency on dense annotations. However, it is a challenging task as it aims to achieve a mapping from high-level semantics to low-level features. In this work, we propose a three-step method to bridge this gap. First, we rely on the interpretable ability of deep neural networks to generate attention maps with class localization information by back-propagating gradients. Secondly, we employ an off-the-shelf object saliency detector with an iterative erasing strategy to obtain saliency maps with spatial extent information of objects. Finally, we combine these two complementary maps to generate pseudo ground-truth images for the training of the segmentation network. With the help of the pre-trained model on the MS-COCO dataset and a multi-scale fusion method, we obtained mIoU of 62.1% and 63.3% on PASCAL VOC 2012 val and test sets, respectively, achieving new state-of-the-art results for the weakly supervised semantic segmentation task.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel
ICASSP4
2019 Coral Classification Using DenseNet and Cross-modality Transfer Learning
abstract
Coral classification is a challenging task due to the complex morphology and ambiguous boundaries of corals. This paper investigates the benefits of Densely connected convolutional network (DenseNet) and multi-modal image translation techniques in boosting image classification performance by synthesizing missing fluorescence information. To this end, an imageconditional Generative Adversarial Network (GAN) based image translator is trained to model the relationship between reflectance and fluorescence images. Through this image translator, fluorescence images can be generated from the available reflectance images to provide complementary information. During the classification phase, reflectance and translated fluorescence images are combined to obtain more discriminative representations and produce improved classification performance. We present results on the EFC and MLC datasets and report state-of-the-art coral classification performance.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel
IJCNN4
2018 Global Regularizer and Temporal-Aware Cross-Entropy for Skeleton-Based Early Action Recognition
Qiuhong Ke, Jun Liu 0036, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd
ACCV (4)5
2018 Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network Representations
abstract
Coral species, with complex morphology and ambiguous boundaries, pose a great challenge for automated classification. CNN activations, which are extracted from fully connected layers of deep networks (FC features), have been successfully used as powerful universal representations in many visual tasks. In this paper, we investigate the transferability and combined performance of FC features and CONY features (extracted from convolutional layers) in the coral classification of two image modalities (reflectance and fluorescence), using a typical deep network (e.g. VGGNet). We exploit vector of locally aggregated descriptors (VLAD) encoding and principal component analysis (PCA) to compress dense CONY features into a compact representation. Experimental results demonstrate that encoded CONV3 features achieve superior performances on reflectance and fluorescence coral images, compared to FC features. The combination of these two features further improves the overall accuracy and achieves state-of-the-art performance on the challenging EFC dataset.
Lian Xu, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
ICASSP3
2018 Exploiting layerwise convexity of rectifier networks with sign constrained weights
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel
Neural Networks1
2018 Response to "Ghost Numbers"
abstract
This note clarifies the experimental settings of [1] and shows that the issue raised by [2] is due to a lack of details in [1] which resulted in a misinterpretation of the experimental settings.
Munawar Hayat, Mohammed Bennamoun, Senjian An
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Learning Clip Representations for Skeleton-Based 3D Action Recognition
abstract
This paper presents a new representation of skeleton sequences for 3D action recognition. Existing methods based on hand-crafted features or recurrent neural networks cannot adequately capture the complex spatial structures and the long-term temporal dynamics of the skeleton sequences, which are very important to recognize the actions. In this paper, we propose to transform each channel of the 3D coordinates of a skeleton sequence into a clip. Each frame of the generated clip represents the temporal information of the entire skeleton sequence and one particular spatial relationship between the skeleton joints. The entire clip incorporates multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We also propose a multitask convolutional neural network (MTCNN) to learn the generated clips for action recognition. The proposed MTCNN processes all the frames of the generated clips in parallel to explore the spatial and temporal information of the skeleton sequences. The proposed method has been extensively tested on six challenging benchmark datasets. Experimental results consistently demonstrate the superiority of the proposed clip representation and the feature learning method for 3D action recognition compared to the existing techniques.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Image Process.3
2018 Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction
abstract
Predicting an interaction before it is fully executed is very important in applications, such as human-robot interaction and video surveillance. In a two-human interaction scenario, there are often contextual dependency structures between the global interaction context of the two humans and the local context of the different body parts of each human. In this paper, we propose to learn the structure of the interaction contexts and combine it with the spatial and temporal information of a video sequence to better predict the interaction class. The structural models, including the spatial and the temporal models, are learned with long short term memory (LSTM) networks to capture the dependency of the global and local contexts of each RGB frame and each optical flow image, respectively. LSTM networks are also capable of detecting the key information from global and local interaction contexts. Moreover, to effectively combine the structural models with the spatial and temporal models for interaction prediction, a ranking score fusion method is introduced to automatically compute the optimal weight of each model for score fusion. Experimental results on the BIT-Interaction Dataset and the UT-Interaction Dataset clearly demonstrate the benefits of the proposed method.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Multim.3
2017 A New Representation of Skeleton Sequences for 3D Action Recognition
abstract
This paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several frames for spatial temporal feature learning using deep neural networks. Each clip is generated from one channel of the cylindrical coordinates of the skeleton sequence. Each frame of the generated clips represents the temporal information of the entire skeleton sequence, and incorporates one particular spatial relationship between the joints. The entire clips include multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We propose to use deep convolutional neural networks to learn long-term temporal information of the skeleton sequence from the frames of the generated clips, and then use a Multi-Task Learning Network (MTLN) to jointly process all frames of the clips in parallel to incorporate spatial structural information for action recognition. Experimental results clearly show the effectiveness of the proposed new representation and feature learning method for 3D action recognition.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
CVPR3
2017 Resfeats: Residual network based features for image classification
abstract
Deep residual networks have recently emerged as the state-of-the-art architecture in image classification and object detection. In this paper, we propose new image features (called ResFeats) extracted from the last convolutional layer of the deep residual networks pre-trained on ImageNet. We propose to use ResFeats for diverse image classification tasks namely, object classification, scene classification and coral classification and show that ResFeats consistently perform better than their CNN counterparts on these classification tasks. Since the ResFeats are large feature vectors, we explore dimensionality reduction methods. Experimental results are provided to show the effectiveness of ResFeats with state-of-the-art classification accuracies on Caltech-101, Caltech-256 and MLC datasets and a significant performance improvement on MIT-67 dataset compared to the widely used CNN features.
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel
ICIP3
2017 SkeletonNet: Mining Deep Part Features for 3-D Action Recognition
abstract
This letter presents SkeletonNet, a deep learning framework for skeleton-based 3-D action recognition. Given a skeleton sequence, the spatial structure of the skeleton joints in each frame and the temporal information between multiple frames are two important factors for action recognition. We first extract body-part-based features from each frame of the skeleton sequence. Compared to the original coordinates of the skeleton joints, the proposed features are translation, rotation, and scale invariant. To learn robust temporal information, instead of treating the features of all frames as a time series, we transform the features into images and feed them to the proposed deep learning network, which contains two parts: one to extract general features from the input images, while the other to generate a discriminative and compact representation for action recognition. The proposed method is tested on the SBU kinect interaction dataset, the CMU dataset, and the large-scale NTU RGB+D dataset and achieves state-of-the-art performance.
Qiuhong Ke, Senjian An, Mohammed Bennamoun, Ferdous Sohel, Farid Boussaïd
IEEE Signal Process. Lett.2
2016 Coral classification with hybrid feature representations
abstract
Coral reefs exhibit significant within-class variations, complex between-class boundaries and inconsistent image clarity. This makes coral classification a challenging task. In this paper, we report the application of generic CNN representations combined with hand-crafted features for coral reef classification to take advantage of the complementary strengths of these representation types. We extract CNN based features from patches centred at labelled pixels at multiple scales. We use texture and color based hand-crafted features extracted from the same patches to complement the CNN features. Our proposed method achieves a classification accuracy that is higher than the state-of-art methods on the MLC benchmark dataset for corals.
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd, Renae Hovey, Gary A. Kendrick, Robert B. Fisher
ICIP3
2016 A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification
abstract
Unlike standard object classification, where the image to be classified contains one or multiple instances of the same object, indoor scene classification is quite different since the image consists of multiple distinct objects. Furthermore, these objects can be of varying sizes and are present across numerous spatial locations in different layouts. For automatic indoor scene categorization, large-scale spatial layout deformations and scale variations are therefore two major challenges and the design of rich feature descriptors which are robust to these challenges is still an open problem. This paper introduces a new learnable feature descriptor called “spatial layout and scale invariant convolutional activations” to deal with these challenges. For this purpose, a new convolutional neural network architecture is designed which incorporates a novel “spatially unstructured” layer to introduce robustness against spatial layout deformations. To achieve scale invariance, we present a pyramidal image representation. For feasible training of the proposed network for images of indoor scenes, this paper proposes a methodology, which efficiently adapts a trained network model (on a large-scale data) for our task with only a limited amount of available training data. The efficacy of the proposed approach is demonstrated through extensive experiments on a number of data sets, including MIT-67, Scene-15, Sports-8, Graz-02, and NYU data sets.
Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Senjian An
IEEE Trans. Image Process.4
2015 Contractive Rectifier Networks for Nonlinear Maximum Margin Classification
abstract
To find the optimal nonlinear separating boundary with maximum margin in the input data space, this paper proposes Contractive Rectifier Networks (CRNs), wherein the hidden-layer transformations are restricted to be contraction mappings. The contractive constraints ensure that the achieved separating margin in the input space is larger than or equal to the separating margin in the output layer. The training of the proposed CRNs is formulated as a linear support vector machine (SVM) in the output layer, combined with two or more contractive hidden layers. Effective algorithms have been proposed to address the optimization challenges arising from contraction constraints. Experimental results on MNIST, CIFAR-10, CIFAR-100 and MIT-67 datasets demonstrate that the proposed contractive rectifier networks consistently outperform their conventional unconstrained rectifier network counterparts.
Senjian An, Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel
ICCV1
2015 How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances?
abstract
This paper investigates how hidden layers of deep rectifier networks are capable of transforming two or more pattern sets to be linearly separable while preserving the distances with a guaranteed degree, and proves the universal classification power of such distance preserving rectifier networks. Through the nearly isometric nonlinear transformation in the hidden layers, the margin of the linear separating plane in the output layer and the margin of the nonlinear separating boundary in the original data space can be closely related so that the maximum margin classification in the input data space can be achieved approximately via the maximum margin linear classifiers in the output layer. The generalization performance of such distance preserving deep rectifier neural networks can be well justified by the distance-preserving properties of their hidden layers and the maximum margin property of the linear classifiers in the output layer.
Senjian An, Farid Boussaïd, Mohammed Bennamoun
ICML1
2015 Sign Constrained Rectifier Networks with Applications to Pattern Decompositions
Senjian An, Qiuhong Ke, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel
ECML/PKDD (1)1
2015 Deep Reconstruction Models for Image Set Classification
abstract
Image set classification finds its applications in a number of real-life scenarios such as classification from surveillance videos, multi-view camera networks and personal albums. Compared with single image based classification, it offers more promises and has therefore attracted significant research attention in recent years. Unlike many existing methods which assume images of a set to lie on a certain geometric surface, this paper introduces a deep learning framework which makes no such prior assumptions and can automatically discover the underlying geometric structure. Specifically, a Template Deep Reconstruction Model (TDRM) is defined whose parameters are initialized by performing unsupervised pre-training in a layer-wise fashion using Gaussian Restricted Boltzmann Machines (GRBMs). The initialized TDRM is then separately trained for images of each class and class-specific DRMs are learnt. Based on the minimum reconstruction errors from the learnt class-specific models, three different voting strategies are devised for classification. Extensive experiments are performed to demonstrate the efficacy of the proposed framework for the tasks of face and object recognition from image sets. Experimental results show that the proposed method consistently outperforms the existing state of the art methods.
Munawar Hayat, Mohammed Bennamoun, Senjian An
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Quantitative Error Analysis of Bilateral Filtering
abstract
One of the fastest acceleration techniques for bilateral image filtering is the real time O(1) quantization method proposed by Yang 2009, which first computes some Principal Bilateral Filtered Image Components (PBFICs) and then applies linear interpolation to estimate the filtered output images. There is a trade-off between accuracy and efficiency in selecting the number of PBFICs: the more PBFICs are used, the higher the accuracy, and the higher the computational cost. A question arises: how many PBFICs are required to achieve a certain level of accuracy? In this letter, we address this question by investigating the properties of bilateral filtering and deriving the linear interpolation error bounds when only a subset of PBFICs is used. The provided theoretical analysis indicates that the necessary number of PBFICs for user-provided precision depends on the range kernel and, for typical Gaussian range kernels, a small percentage (typically less than 4%) of the PBFICs are enough for good approximations.
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel
IEEE Signal Process. Lett.1
2014 Learning Non-linear Reconstruction Models for Image Set Classification
abstract
We propose a deep learning framework for image set classification with application to face recognition. An Adaptive Deep Network Template (ADNT) is defined whose parameters are initialized by performing unsupervised pre-training in a layer-wise fashion using Gaussian Restricted Boltzmann Machines (GRBMs). The pre-initialized ADNT is then separately trained for images of each class and class-specific models are learnt. Based on the minimum reconstruction error from the learnt class-specific models, a majority voting strategy is used for classification. The proposed framework is extensively evaluated for the task of image set classification based face recognition on Honda/UCSD, CMU Mobo, YouTube Celebrities and a Kinect dataset. Our experimental results and comparisons with existing state-of-the-art methods show that the proposed method consistently achieves the best performance on all these datasets.
Munawar Hayat, Mohammed Bennamoun, Senjian An
CVPR3
2014 Reverse Training: An Efficient Approach for Image Set Classification
Munawar Hayat, Mohammed Bennamoun, Senjian An
ECCV (6)3
2014 Robust Face Recognition by Utilizing Color Information and Sparse Representation
abstract
In this paper, we consider the problem of robust face recognition using color information. In this context, sparse representation-based algorithms are the state-of-the-art solutions for gray facial images. We will integrate the existing sparse representation-based algorithms with color information and this integration can improve the previous performances significantly. Furthermore, we propose a new performance metric, namely the discriminativeness (DIS) to describe the recognition effectiveness for sparse representation algorithms. We find out that the richer information in color space can be used to increase the DIS, i.e. enhancing the robustness in face recognition. Extensive experiments have been conducted under different conditions, including various feature extractors, random pixel corruptions and occlusions on AR and GT databases, to demonstrate the advantages of using color information in robust face recognition. Detailed analysis is also included for each experiment to explain why and how color improve the robustness of different sparse representation-based methods.
Billy Y. L. Li, Wanquan Liu, Senjian An, Aneesh Krishna
Int. J. Pattern Recognit. Artif. Intell.3
2012 Tensor based robust color face recognition
Billy Y. L. Li, Wanquan Liu, Senjian An, Aneesh Krishna
ICPR3
2012 Face recognition using various scales of discriminant color space transform
Billy Y. L. Li, Wanquan Liu, Senjian An, Aneesh Krishna, Tianwei Xu
Neurocomputing3
2011 Efficient subwindow search with submodular score functions
abstract
Subwindow search aims to find the optimal subimage which maximizes the score function of an object to be detected. After the development of the branch and bound (B&B) method called Efficient Subwindow Search (ESS), several algorithms (IESS, AESS, ARCS) have been proposed to improve the performance of ESS. For n×n images, IESS's time complexity is bounded by O(n3) which is better than ESS, but only applicable to linear score functions. Other work shows that Monge properties can hold in subwindow search and can be used to speed up the search to O(n3), but only applies to certain types of score functions. In this paper we explore the connection between submodular functions and the Monge property, and prove that sub-modular score functions can be used to achieve O(n3) time complexity for object detection. The time complexity can be further improved to be sub-cubic by applying B&B methods on row interval only, when the score function has a multivariate submodular bound function. Conditions for sub-modularity of common non-linear score functions and multivariate submodularity of their bound functions are also provided, and experiments are provided to compare the proposed approach against ESS and ARCS for object detection with some nonlinear score functions.
Senjian An, Patrick Peursum, Wanquan Liu, Svetha Venkatesh
CVPR1
2011 The MCF Model: Utilizing Multiple Colors for Face Recognition
abstract
Finding a good color space is one of the main research goals for color face recognition. Existing research shows that RGB can improve over gray-scale, while some other color spaces (YQCr for instance) can improve over RGB. However, all developed color models consist of only three color components transformed linearly from RGB. Since three colors may not capture sufficient information for solving complex face recognition problems and non-linear transformed colors usually encode very different information, this paper investigates the feasibility and effectiveness of using more than three colors including some non-linear color spaces. We propose a novel color combination algorithm namely the Multiple Color Fusion (MCF) model to utilize multiple colors. Experiment 4 on FRGC2 is conducted to demonstrate the effectiveness of MCF. In particular, MCF outperforms any existing three-color based methods by at least 3% and improves over RGB by 8%.
Billy Y. L. Li, Senjian An, Wanquan Liu, Aneesh Krishna
ICIG2
2011 Margin Preserving Projection for Image Set Based Face Recognition
Wanquan Liu, Senjian An
ICONIP (2)3
2011 Unified formulation of linear discriminant analysis methods and optimal parameter selection
Senjian An, Wanquan Liu, Svetha Venkatesh, Hong Yan 0001
Pattern Recognit.1
2010 Exploiting Monge structures in optimum subwindow search
abstract
Optimum subwindow search for object detection aims to find a subwindow so that the contained subimage is most similar to the query object. This problem can be formulated as a four dimensional (4D) maximum entry search problem wherein each entry corresponds to the quality score of the subimage contained in a subwindow. For n × n images, a naive exhaustive search requires O(n4) sequential computations of the quality scores for all subwindows. To reduce the time complexity, we prove that, for some typical similarity functions like Euclidian metric, χ2metric on image histograms, the associated 4D array carries some Monge structures and we utilise these properties to speed up the optimum subwindow search and the time complexity is reduced to O(n3). Furthermore, we propose a locally optimal alternating column and row search method with typical quadratic time complexity O(n2). Experiments on PASCAL VOC 2006 demonstrate that the alternating method is significantly faster than the well known efficient subwindow search (ESS) method whilst the performance loss due to local maxima problem is negligible.
Senjian An, Patrick Peursum, Wanquan Liu, Svetha Venkatesh
CVPR1
2009 Efficient algorithms for subwindow search in object detection and localization
abstract
Recently, a simple yet powerful branch-and-bound method called Efficient Subwindow Search (ESS) was developed to speed up sliding window search in object detection. A major drawback of ESS is that its computational complexity varies widely from O(n2) to O(n4) for n × n matrices. Our experimental experience shows that the ESS's performance is highly related to the optimal confidence levels which indicate the probability of the object's presence. In particular, when the object is not in the image, the optimal subwindow scores low and ESS may take a large amount of iterations to converge to the optimal solution and so perform very slow. Addressing this problem, we present two significantly faster methods based on the linear-time Kadane's Algorithm for 1D maximum subarray search. The first algorithm is a novel, computationally superior branch-and-bound method where the worst case complexity is reduced to O(n3). Experiments on the PASCAL VOC 2006 data set demonstrate that this method is significantly and consistently faster (approximately 30 times faster on average) than the original ESS. Our second algorithm is an approximate algorithm based on alternating search, whose computational complexity is typically O(n2). Experiments shows that (on average) it is 30 times faster again than our first algorithm, or 900 times faster than ESS. It is thus well-suited for real time object detection.
Senjian An, Patrick Peursum, Wanquan Liu, Svetha Venkatesh
CVPR1
2008 Exploiting side information in locality preserving projection
abstract
Even if the class label information is unknown, side information represents some equivalence constraints between pairs of patterns, indicating whether pairs originate from the same class. Exploiting side information, we develop algorithms to preserve both the intra-class and inter-class local structures. This new type of locality preserving projection (LPP), called LPP with side information (LPPSI), preserves the data’s local structure in the sense that the close, similar training patterns will be kept close, whilst the close but dissimilar ones are separated. Our algorithms balance these conflicting requirements, and we further improve this technique using kernel methods. Experiments conducted on popular face databases demonstrate that the proposed algorithm significantly outperforms LPP. Further, we show that the performance of our algorithm with partial side information (that is, using only small amount of pair-wise similarity/dissimilarity information during training) is comparable with that when using full side information. We conclude that exploiting side information by preserving both similar and dissimilar local structures of the data significantly improves performance.
Senjian An, Wanquan Liu, Svetha Venkatesh
CVPR1
2008 Double Sides 2DPCA for Face Recognition
Chong Lu, Wanquan Liu, Xiaodong Liu 0001, Senjian An
ICIC (1)4
2008 A simplified GLRAM algorithm for face recognition
Chong Lu, Wanquan Liu, Senjian An
Neurocomputing3
2007 Face Recognition Using Kernel Ridge Regression
abstract
In this paper, we present novel ridge regression (RR) and kernel ridge regression (KRR) techniques for multivariate labels and apply the methods to the problem efface recognition. Motivated by the fact that the regular simplex vertices are separate points with highest degree of symmetry, we choose such vertices as the targets for the distinct individuals in recognition and apply RR or KRR to map the training face images into a face subspace where the training images from each individual will locate near their individual targets. We identify the new face image by mapping it into this face subspace and comparing its distance to all individual targets. An efficient cross-validation algorithm is also provided for selecting the regularization and kernel parameters. Experiments were conducted on two face databases and the results demonstrate that the proposed algorithm significantly outperforms the three popular linear face recognition techniques (Eigenfaces, Fisher faces and Laplacian faces) and also performs comparably with the recently developed Orthogonal Laplacian faces with the advantage of computational speed. Experimental results also demonstrate that KRR outperforms RR as expected since KRR can utilize the nonlinear structure of the face images. Although we concentrate on face recognition in this paper, the proposed method is general and may be applied for general multi-category classification problems.
Senjian An, Wanquan Liu, Svetha Venkatesh
CVPR1
2007 Face Recognition via the Overlapping Energy Histogram
Ronny Tjahyadi, Wanquan Liu, Senjian An, Svetha Venkatesh
IJCAI3
2007 Fast cross-validation algorithms for least squares support vector machine and kernel ridge regression
Senjian An, Wanquan Liu, Svetha Venkatesh
Pattern Recognit.1
2006 A Fast Feature-based Dimension Reduction Algorithm for Kernel Classifiers
Senjian An, Wanquan Liu, Svetha Venkatesh, Ronny Tjahyadi
Neural Process. Lett.1
2005 Fast cross-validation of kernel Fisher discriminant classifiers
abstract
Given n training examples, the training of a kernel Fisher discriminant (KFD) classifier corresponds to solving a linear system of dimension n. In cross-validating KFD, the training examples are split into 2 distinct subsets for a number of times (L) wherein a subset of m examples is used for validation and the other subset of (n - m) examples is used for training the classifier. In this case L linear systems of dimension (n - m) need to be solved. We propose a novel method for cross-validation of KFD in which instead of solving L linear systems of dimension (n - m), we compute the inverse of an n /spl times/ n matrix and solve L linear systems of dimension 2m, thereby reducing the complexity when L is large and/or m is small. For typical 10-fold and leave-one-out cross-validations, the proposed algorithm is approximately 4 and ( 4/9 n ) times respectively as efficient as the naive implementations. Simulations are provided to demonstrate the efficiency of the proposed algorithms.
Senjian An, Wanquan Liu, Svetha Venkatesh
ICMLA1
2003 Blind identification of FIR MIMO channels by group decorrelation
abstract
We present a new method for identification of FIR MIMO channels driven by unknown, uncorrelated and colored sources. This method, belonging to the BID (blind identification by decorrelation) family, makes use of the mutual uncorrelation of the unknown sources by first decorrelating the observed signals into two uncorrelated groups. The two decorrelators are then used to estimate the channel matrix (i.e., MIMO channel transfer function matrix) up to a constant matrix. This constant matrix is finally determined using a BID method for instantaneous MIMO channels. This new method, named BID-G, is shown to be much more robust than the subspace method that requires the channel matrix to be irreducible and column-reduced.
Senjian An, Yingbo Hua, Jonathan H. Manton
ICASSP (5)1
2002 Separating colorred signals distorted by convolutive channels using diagonal constrained decorrelation
abstract
We consider the problem of separating colorred signals mixed by unknown convolutive channels. We introduce a variation of a previous approach of separating signals via decorrelation. This variation minimizes the mutual correlation among the output signals of a decorrelation matrix subject to a diagonal constraint. The diagonal constraint ensures a better quality of separation. The diagonal constraint is shown to be essential when the number of original signals is less than the number of distorted signals
Yingbo Hua, Senjian An, Alex Acero
ICASSP3
2001 Blind identification and equalization of FIR MIMO channels by BIDS
abstract
This paper presents an algorithm of blind identification and equalization of finite-impulse-response and multiple-input and multiple-output (FIR MIMO) channels driven by colored signals. This algorithm is an improved realization of a concept referred to as blind identification via decorrelating subchannels (BIDS). This BIDS algorithm first constructs a set of decorrelators which decorrelate the output signals of subchannels, and then estimates the channel matrix using the transfer functions of the decorrelators, and finally recovers the input signals using the estimated channel matrix. This BIDS algorithm in general assumes that the channel matrix is irreducible and the input signals are mutually uncorrelated and of sufficiently diverse power spectra. However, for channel matrix identification, this BIDS algorithm only requires the channel matrix to be nonsingular (ie, full rank almost everywhere as opposed to everywhere) and column-wise coprime. Such a channel matrix may have zeros and be of non-minimum phase.
Yingbo Hua, Senjian An
ICASSP2
2001 Experimental investigation of delayed instantaneous demixer for speech enhancement
abstract
This paper presents a delayed instantaneous demixer (DID) for speech signal separation from real recordings. Based on the fact that the original signals are colored and mutually uncorrelated, a simple algorithm is derived to estimate the parameters of the demixer. This algorithm consists of two parts: a grid searching method to estimate time delays and an alternating projection method to estimate gain coefficients. Experimental result demonstrates the performance of the model and the algorithm.
Yingbo Hua, Senjian An, Alex Acero
ICASSP3