Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Vijayan K. Asari

dblp:71/3518 · also K. Vijayan Asari · DBLP profile ↗
← Back
70ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-3751-5492ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8Systems, architecture and hardware · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 46% Deep learning architectures and training · 22% Face, body and person analysis · 22%
Computer graphics and multimedia
1 paper
Image and video processing · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 94% Electronic design automation · 6%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d human pose estimation
0.922021
Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions · Int. J. Comput. Vis. 2021
Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction · CVPR 2020
Computer vision › 3D vision
event-based vision
0.712023
Time-Ordered Recent Event (TORE) Volumes for Event Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › event-based vision
event camera
0.712023
Time-Ordered Recent Event (TORE) Volumes for Event Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › reasoning about action and change
event representation
0.712023
Time-Ordered Recent Event (TORE) Volumes for Event Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Face, body and person analysis
human pose estimation
0.512021
Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions · Int. J. Comput. Vis. 2021
Computer vision › 3D vision › 3d human pose estimation
video-based 3d pose estimation
0.512021
Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions · Int. J. Comput. Vis. 2021
Machine learning › Deep learning architectures and training
attention mechanism
0.412020
Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction · CVPR 2020
Machine learning › Deep learning architectures and training
convolutional neural network
0.412020
Event Probability Mask (EPM) and Event Denoising Convolutional Neural Network (EDnCNN) for Neuromorphic Cameras · CVPR 2020
Machine learning › Deep learning architectures and training › attention mechanism
temporal attention
0.412020
Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction · CVPR 2020
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation
0.412020
Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction · CVPR 2020
Emerging computing paradigms
neuromorphic computing
0.412020
Event Probability Mask (EPM) and Event Denoising Convolutional Neural Network (EDnCNN) for Neuromorphic Cameras · CVPR 2020
Image and video processing › image restoration
image deraining
0.212015
Utilizing Local Phase Information to Remove Rain from Video · Int. J. Comput. Vis. 2015
Image and video processing
video restoration
0.212015
Utilizing Local Phase Information to Remove Rain from Video · Int. J. Comput. Vis. 2015
Computer vision › Face, body and person analysis
head pose estimation
0.212013
A Two-Layer Framework for Piecewise Linear Manifold-Based Head Pose Estimation · Int. J. Comput. Vis. 2013
Computer vision › Face, body and person analysis
face recognition
0.112009
Facial Recognition Using Multisensor Images Based on Localized Kernel Eigen Spaces · IEEE Trans. Image Process. 2009
Computer vision › Face, body and person analysis › face recognition › infrared face recognition
thermal face recognition
0.112009
Facial Recognition Using Multisensor Images Based on Localized Kernel Eigen Spaces · IEEE Trans. Image Process. 2009
Image and video processing
video enhancement
0.112015
Utilizing Local Phase Information to Remove Rain from Video · Int. J. Comput. Vis. 2015
Electronic design automation
logic synthesis
0.011994
An Optimization Technique for the Design of Multiple Valued PLA's · IEEE Trans. Computers 1994
Electronic design automation › logic synthesis
multiple-valued logic synthesis
0.011994
An Optimization Technique for the Design of Multiple Valued PLA's · IEEE Trans. Computers 1994
Electronic design automation › logic synthesis › programmable logic array
programmable logic array design
0.011994
An Optimization Technique for the Design of Multiple Valued PLA's · IEEE Trans. Computers 1994

Methods — techniques the papers use, named apart from their topics

dilated convolution · 0.9event probability mask · 0.9denoising convolutional neural network · 0.9attention-based neural network · 0.5local phase information · 0.2image decomposition · 0.2phase congruency · 0.1kernel methods · 0.1decision-level fusion · 0.1output encoding · 0.0literal circuit minimization · 0.0
YearPublicationVenuePosition
2026 PoseGaussian: Pose-Driven Novel View Synthesis for Robust 3D Human Reconstruction
abstract
We propose PoseGaussian, a pose-guided Gaussian Splatting framework for high-fidelity human novel view synthesis. Human body pose serves a dual purpose in our design: as a structural prior, it is fused with a color encoder to refine depth estimation; as a temporal cue, it is processed by a dedicated pose encoder to enhance temporal consistency across frames. These components are integrated into a fully differentiable, end-to-end trainable pipeline. Unlike prior works that use pose only as a condition or for warping, PoseGaussian embeds pose signals into both geometric and temporal stages to improve robustness and generalization. It is specifically designed to address challenges inherent in dynamic human scenes, such as articulated motion and severe self-occlusion. Notably, our framework achieves real-time rendering at 100 FPS, maintaining the efficiency of standard Gaussian Splatting pipelines. We validate our approach on ZJU-MoCap, THuman2.0, and in-house datasets, demonstrating state-of-the-art performance in perceptual quality and structural accuracy (PSNR 30.86, SSIM 0.979, LPIPS 0.028).
Ju Shen, Chen Chen 0001, Tam V. Nguyen 0002, Vijayan K. Asari
WACV4
2025 Extrapolation Convolution for Data Prediction on a 2-D Grid: Bridging Spatial and Frequency Domains With Applications in Image Outpainting and Compressed Sensing
abstract
Extrapolation plays a critical role in machine/deep learning (ML/DL), enabling models to predict data points beyond their training constraints, particularly useful in scenarios deviating significantly from training conditions. This article addresses the limitations of current convolutional neural networks (CNNs) in extrapolation tasks within image restoration and compressed sensing (CS). While CNNs show potential in tasks such as image outpainting and CS, traditional convolutions are limited by their reliance on interpolation, failing to fully capture the dependencies needed for predicting values outside the known data. This work proposes an extrapolation convolution (EC) framework that models missing data prediction as an extrapolation problem using linear prediction within DL architectures. The approach is applied in two domains: first, image outpainting, where EC in encoder-decoder (EnDec) networks replaces conventional interpolation methods to reduce artifacts and enhance fine detail representation; second, Fourier-based CS-magnetic resonance imaging (CS-MRI), where it predicts high-frequency signal values from undersampled measurements in the frequency domain, improving reconstruction quality and preserving subtle structural details at high acceleration factors. Comparative experiments demonstrate that the proposed EC-DecNet and FDRN outperform traditional CNN-based models, achieving high-quality image reconstruction with finer details, as shown by improved peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and kernel inception distance (KID)/Frechet inception distance (FID) scores. Ablation studies and analysis highlight the effectiveness of larger kernel sizes and multilevel semi-supervised learning in FDRN for enhancing extrapolation accuracy in the frequency domain.
Vazim Ibrahim, Faouzi Alaya Cheikh, Vijayan K. Asari, Joseph Suresh Paul
IEEE Trans. Neural Networks Learn. Syst.3
2024 Pyramid Point: A Multilevel Focusing Network for Revisiting Feature Layers
abstract
We present a mcthod to lcarn a diverse group of object categories from an unordcrcd point set. We propose our Pyramid Point network, which uses a dense pyramid structure instead of thc traditional ’U’ shape, typically seen in semantic segmentation networks. This pyramid structure gives a second look, allowing thc network to revisit different layers creating from various leveis on the network, allowing for feature propagation. We introduce a Focused Kemel Point convolution (FKP Conv), which expands on the traditional point convolutions by adding an attention mcchanism to the kemel outputs. This FKP Conv increases our feature quality and allows us to weigh thc kemel outputs dynamically. These FKP Convs are the central part of our Recurrent FKP Bottlcneck block, which makes up thc backbonc of our encoder. With this distinct network, we demonstrate competitive performance on threc benchmark data sets.
Nina M. Varney, Vijayan K. Asari
IEEE Geosci. Remote. Sens. Lett.2
2023 Time-Ordered Recent Event (TORE) Volumes for Event Cameras
abstract
Event cameras are an exciting, new sensor modality enabling high-speed imaging with extremely low-latency and wide dynamic range. Unfortunately, most machine learning architectures are not designed to directly handle sparse data, like that generated from event cameras. Many state-of-the-art algorithms for event cameras rely on interpolated event representations-obscuring crucial timing information, increasing the data volume, and limiting overall network performance. This paper details an event representation called Time-Ordered Recent Event (TORE) volumes. TORE volumes are designed to compactly store raw spike timing information with minimal information loss. This bio-inspired design is memory efficient, computationally fast, avoids time-blocking (i.e., fixed and predefined frame rates), and contains "local memory" from past data. The design is evaluated on a wide range of challenging tasks (e.g., event denoising, image reconstruction, classification, and human pose estimation) and is shown to dramatically improve state-of-the-art performance. TORE volumes are an easy-to-implement replacement for any algorithm currently utilizing event representations.
Raymond Baldwin, Ruixu Liu, Mohammed Almatrafi, Vijayan K. Asari, Keigo Hirakawa
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions
Ruixu Liu, Ju Shen, Chen Chen 0001, Sen-Ching S. Cheung, Vijayan K. Asari
Int. J. Comput. Vis.6
2021 Inception recurrent convolutional neural network for object recognition
Md. Zahangir Alom, Mahmudul Hasan 0003, Chris Yakopcic, Tarek M. Taha, Vijayan K. Asari
Mach. Vis. Appl.5
2020 Event Probability Mask (EPM) and Event Denoising Convolutional Neural Network (EDnCNN) for Neuromorphic Cameras
abstract
This paper presents a novel method for labeling real-world neuromorphic camera sensor data by calculating the likelihood of generating an event at each pixel within a short time window, which we refer to as “event probability mask” or EPM. Its applications include (i) objective benchmarking of event denoising performance, (ii) training convolutional neural networks for noise removal called “event denoising convolutional neural network” (EDnCNN), and (iii) estimating internal neuromorphic camera parameters. We provide the first dataset (DVSNOISE20) of real-world labeled neuromorphic camera events for noise removal.
Raymond Baldwin, Mohammed Almatrafi, Vijayan K. Asari, Keigo Hirakawa
CVPR3
2020 Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction
abstract
We propose a novel attention-based framework for 3D human pose estimation from a monocular video. Despite the general success of end-to-end deep learning paradigms, our approach is based on two key observations: (1) temporal incoherence and jitter are often yielded from a single frame prediction; (2) error rate can be remarkably reduced by increasing the receptive field in a video. Therefore, we design an attentional mechanism to adaptively identify significant frames and tensor outputs from each deep neural net layer, leading to a more optimal estimation. To achieve large temporal receptive fields, multi-scale dilated convolutions are employed to model long-range dependencies among frames. The architecture is straightforward to implement and can be flexibly adopted for real-time applications. Any off-the-shelf 2D pose estimation system, e.g. Mocap libraries, can be easily integrated in an ad-hoc fashion. We both quantitatively and qualitatively evaluate our method on various standard benchmark datasets (e.g. Human3.6M, HumanEva). Our method considerably outperforms all the state-of-the-art algorithms up to 8% error reduction (average mean per joint position error: 34.7) as compared to the best-reported results. Code is available at: (https://github.com/lrxjason/Attention3DHumanPose)
Ruixu Liu, Ju Shen, Chen Chen 0001, Sen-Ching S. Cheung, Vijayan K. Asari
CVPR6
2020 Improved inception-residual convolutional neural network for object recognition
Md. Zahangir Alom, Mahmudul Hasan 0003, Chris Yakopcic, Tarek M. Taha, Vijayan K. Asari
Neural Comput. Appl.5
2019 Active Recall Networks for Multiperspectivity Learning through Shared Latent Space Optimization
abstract
Given that there are numerous amounts of unlabeled data available for usage in training neural networks, it is desirable to implement a neural network architecture and training paradigm to maximize the ability of the latent space representation. Through multiple perspectives of the latent space using adversarial learning and autoencoding, data requirements can be reduced, which improves learning ability across domains. The entire goal of the proposed work is not to train exhaustively, but to train with multiperspectivity. We propose a new neural network architecture called Active Recall Network (ARN) for learning with less labels by optimizing the latent space. This neural network architecture learns latent space features of unlabeled data by using a fusion framework of an autoencoder and a generative adversarial network. Variations in the latent space representations will be captured and modeled by generation, discrimination, and reconstruction strategies in the network using both unlabeled and labeled data. Performance evaluations conducted on the proposed ARN architectures with two popular datasets demonstrated promising results in terms of generative capabilities and latent space effectiveness. Through the multiple perspectives that are embedded in ARN, we envision that this architecture will be incredibly versatile in every application that requires learning with less labels.
Theus H. Aspiras, Ruixu Liu, Vijayan K. Asari
IJCCI3
2018 Long Short Working Memory (LSWM) Integration with Polynomial Connectivity for Object Tracking in Wide Area Motion Imagery
abstract
High-value-target tracking in full-motion-video is a difficult surveillance task due to model drift while online training. We propose a tracker that adaptively fuses detections from multiple target models trained using three memory types to overcome model drift and other challenges. The short-term memory uses correlation and histogram features to detect the target from its recent appearance. The working memory uses a deep extreme learning network with polynomial connectivity that is trained online using a set of target appearances and background from the recent past. The long-term memory uses an offline trained polynomial CNN as a vehicle detector. Occlusion detection along with image registration and motion estimation stages aid in tracking the target through occlusions. A motion detection module reduces the effects of model drift and the scale change detector keeps the boundary accurate to the target. The proposed tracker is evaluated against state-of-the-art tracking methods on a vehicle dataset that includes challenging scenarios such as tree canopy occlusion, shadow, and erratic vehicle motion.
Evan Krieger, Theus H. Aspiras, Vijayan K. Asari, Yakov Diskin
AVSS3
2018 Adaptive Trigonometric Transformation Function With Image Contrast and Color Enhancement: Application to Unmanned Aerial System Imagery
abstract
An unmanned aerial system (UAS)-based imaging technology has gained great interests in modern photogrammetry and remote sensing. However, due to the limitations of UAS imaging devices, image enhancement (IE) has become a necessary process for improving the visual appearance of UAS images. Although a great amount of effort has been focused on improving image quality from different aspects, the major obstacles are from computational efficiency and complexity, such as manually adjusting the associated algorithmic parameters that account for various image luminance. To overcome these drawbacks, we propose a new adaptive yet highly efficient luminance enhancement method, namely, adaptive trigonometric transformation function (ATTF), for enhancing the visual quality of digital color images captured by a UAS. The ATTF is derived from a tangent-based transformation function whose characteristics adaptively change with respect to the variation of the image luminance. By combining ATTF with a Laplacian operator and a color restoration process, a well-balanced color image is obtained. The effectiveness of the proposed technique is evaluated on various UAS-based images and compared with other IE techniques.
Sidike Paheding, Vasit Sagan, Maher B. Qumsiyeh, Maitiniyazi Maimaitijiang, Almabrok Essa, Vijayan K. Asari
IEEE Geosci. Remote. Sens. Lett.6
2017 Visual perception based adaptive feature fusion for visual object tracking
abstract
To overcome visual object tracking challenges, various feature-based object trackers use feature combination. Each feature component is developed to overcome certain tracking challenges, but the interaction between the components may cause tracking errors. We propose a tracking solution based on human vision principles to reduce combination errors by adaptively fusing each feature using its previous performance. An adaptive fusion technique is developed to determine feature quality using feature likelihood map variance ratios. The proposed method is completely modular, while reducing the risk of tracker failure. Experimental results on the Visual Object Tracking database show the proposed tracker's robustness and its advantage over state-of-the-art trackers.
Evan Krieger, Vijayan K. Asari
SMC2
2017 Hierarchical Autoassociative Polynimial Network (HAP Net) for pattern recognition
Theus H. Aspiras, Vijayan K. Asari
Neurocomputing2
2017 Volumetric Directional Pattern for Spatial Feature Extraction in Hyperspectral Imagery
abstract
In this letter, we propose to use an enhanced version of volumetric directional pattern to efficiently extract rich spatial context information in the hyperspectral imagery (HSI). The proposed technique fuses the texture information from three consecutive bands in the input HSI. The extracted local image texture features for each pixel of interest are then fed into an extreme learning machine classifier to assign object category. The experimental results on three standard hyperspectral data sets demonstrate the effectiveness of the proposed method for HSI classification compared with that of a set of state-of-the-art spatial extraction methods.
Almabrok Essa, Sidike Paheding, Vijayan K. Asari
IEEE Geosci. Remote. Sens. Lett.3
2017 State Preserving Extreme Learning Machine: A Monotonically Increasing Learning Approach
Md. Zahangir Alom, Sidike Paheding, Tarek M. Taha, Vijayan K. Asari
Neural Process. Lett.4
2016 Automatic building change detection through adaptive local textural features and sequential background removal
abstract
In this paper, we present a new framework for building change detection from monocular aerial imagery that automatically predicts building candidates based on adaptive local textural features with successive background removal. An adaptive local entropy feature is developed based on quadratic regression and Random Sample Consensus (RANSAC) for extracting potential building candidates. Then a majority voting aggregation strategy is employed to accurately estimate shadow direction associated with building objects to aid in reducing false positives in the detected building candidates. A ground plane estimation method is proposed to distinguish building and non-building objects that share similar textural features. Finally, to better estimate changes in building area, a double convex hull based morphological merging technique is introduced. The evaluation results of the proposed framework performed on three-band (RGB) aerial images indicate its capability to successfully detect building changes in urban, suburban, and rural areas.
Sidike Paheding, Daniel Prince, Almabrok Essa, Vijayan K. Asari
IGARSS4
2016 Multiclass Object Detection With Single Query in Hyperspectral Imagery Using Class-Associative Spectral Fringe-Adjusted Joint Transform Correlation
abstract
We present a deterministic object detection algorithm capable of detecting multiclass objects in hyperspectral imagery (HSI) without any training or preprocessing. The proposed method, which is named class-associative spectral fringe-adjusted joint transform correlation (CSFJTC), is based on joint transform correlation (JTC) between object and nonobject spectral signatures to search for a similar match, which only requires one query (training-free) from the object's spectral signature. Our method utilizes class-associative filtering, modified Fourier plane image subtraction, and fringe-adjusted JTC techniques in spectral correlation domain to perform the object detection task. The output of CSFJTC yields a pair of sharp correlation peaks for a matched target and negligible or no correlation peaks for a mismatch. Experimental results, in terms of receiver operating characteristic (ROC) curves and area-under-ROC (AUROC), on three popular real-world hyperspectral data sets demonstrate the superiority of the proposed CSFJTC technique over other well-known hyperspectral object detection approaches.
Sidike Paheding, Vijayan K. Asari, Mohammad S. Alam
IEEE Trans. Geosci. Remote. Sens.2
2015 A self-organizing lattice Boltzmann active contour (SOLBAC) approach for fast and robust object region segmentation
abstract
In this paper, we propose a self-organized learning based active contour model with a lattice Boltzmann convergence criteria for fast and effective segmentation preserving the precise details of the object's region of interest. A dual self-organizing map approach is being used to learn the object of interest and the background independently in order to guide the active contour to extract the target region. The lattice Boltzmann method is utilized to evolve the level-set function faster and terminate the evolution of the curve at the most optimum region, which segments objects in cluttered environments. Experiments performed on a challenging dataset (PSCAL 2011) show promising results in terms of time and quality of the segmentation and that our method is more than 53% faster than other state-of-the-art learning-based active contour model approaches.
Fatema A. Albalooshi, Vijayan K. Asari
ICIP2
2015 Directional ringlet intensity feature transform for tracking
abstract
The challenges existing for current intensity-based histogram feature tracking methods in wide area motion imagery include object structural information distortions and background variations, such as different pavement or ground types. All of these challenges need to be met in order to have a robust object tracker, while attaining to be computed at an appropriate speed for real-time processing. To achieve this we propose a novel method, Directional Ringlet Intensity Feature Transform (DRIFT), that employs Kirsch kernel filtering and Gaussian ringlet feature mapping. We evaluated the DRIFT on two challenging datasets, namely Columbus Large Image Format (CLIF) and Large Area Image Recorder (LAIR), to evaluate its robustness and efficiency. Experimental results show that the proposed approach yields the highest accuracy compared to state-of-the-art object tracking methods.
Evan Krieger, Sidike Paheding, Theus H. Aspiras, Vijayan K. Asari
ICIP4
2015 State Preserving Extreme Learning Machine for face recognition
abstract
Extreme Learning Machine (ELM) has been introduced as a new algorithm for training single hidden layer feed-forward neural networks (SLFNs) instead of the classical gradient-based algorithms. Based on the consistency property of data, which enforce similar samples to share similar properties, ELM is a biologically inspired learning algorithm with SLFNs that learns much faster with good generalization and performs well in classification applications. However, the random generation of the weight matrix in current ELM based techniques leads to the possibility of unstable outputs in the learning and testing phases. Therefore, we present a novel approach for computing the weight matrix in ELM which forms a State Preserving Extreme Leaning Machine (SPELM). The SPELM stabilizes ELM training and testing outputs while monotonically increases its accuracy by preserving state variables. Furthermore, three popular feature extraction techniques, namely Gabor, Pyramid Histogram of Oriented Gradients (PHOG) and Local Binary Pattern (LBP) are incorporated with the SPELM for performance evaluation. Experimental results show that our proposed algorithm yields the best performance on the widely used face datasets such as Yale, CMU and ORL compared to state-of-the-art ELM based classifiers.
Md. Zahangir Alom, Sidike Paheding, Vijayan K. Asari, Tarek M. Taha
IJCNN3
2015 Utilizing Local Phase Information to Remove Rain from Video
Varun Santhaseelan, Vijayan K. Asari
Int. J. Comput. Vis.2
2015 No-Reference Video Quality Assessment Based on Artifact Measurement and Statistical Analysis
abstract
A discrete cosine transform (DCT)-based no-reference video quality prediction model is proposed that measures artifacts and analyzes the statistics of compressed natural videos. The model has two stages: 1) distortion measurement and 2) nonlinear mapping. In the first stage, an unsigned ac band, three frequency bands, and two orientation bands are generated from the DCT coefficients of each decoded frame in a video sequence. Six efficient frame-level features are then extracted to quantify the distortion of natural scenes. In the second stage, each frame-level feature of all frames is transformed to a corresponding video-level feature via a temporal pooling, then a trained multilayer neural network takes all video-level features as inputs and outputs, a score as the predicted quality of the video sequence. The proposed method was tested on videos with various compression types, content, and resolution in four databases. We compared our model with a linear model, a support-vector-regression-based model, a state-of-the-art training-based model, and a four popular full-reference metrics. Detailed experimental results demonstrate that the results of the proposed method are highly correlated with the subjective assessments.
Kongfeng Zhu, Chengqing Li, Vijayan K. Asari, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.3
2014 Gaussian ringlet intensity distribution (GRID) features for rotation-invariant object detection in wide area motion imagery
abstract
Most detection algorithms are established by using well defined features. Since wide area imagery is low resolution and has features that are not well defined, a local intensity distribution based methodology seems a likely candidate. We propose a new methodology, Gaussian Ringlet Intensity Distribution (GRID), which is a derivative of the ring-partitioned histograms for local intensity distribution based object tracking in low-resolution environments, which deals with the issue of rotation invariance. We observed that the proposed algorithm produces the highest accuracy among other state of the art methodologies and provides robust features for rotationally invariant detection and tracking in wide area motion imagery.
Theus H. Aspiras, Vijayan K. Asari, Juan R. Vasquez
ICIP2
2013 A no-reference video quality assessment based on Laplacian pyramids
abstract
This paper presents an approach to predict the quality of compressed videos with content of natural scenes. The method is focused on measuring the distortion of compressed video without reference. There are two main steps of the proposed method: measuring distortion and predicting video quality. Each frame of the distorted video sequence is first decomposed to an N-subband Laplacian pyramid, then their intra-subband and inter-subband statistical features are fully exploited. Three intra-subband features and three inter-subband features are taken as inputs of the prediction model. Its output is a single score as the predicted video quality. The performance of the proposed method is evaluated on the LIVE video database and the LIVE mobile video database. Results show that the predicted quality scores are well correlated with the mean opinion scores associated to the subjective assessment.
Kongfeng Zhu, Keigo Hirakawa, Vijayan K. Asari, Dietmar Saupe
ICIP3
2013 Local Difference of Gaussian Binary Pattern: Robust Features for Face Sketch Recognition
abstract
Automatic recognition of face sketches is a challenging problem with application in criminal investigations. We propose a method that allows face sketch recognition across modalities called Local Difference of Gaussian Binary Pattern (LDoGBP). LDoGBP is based on the fact that the sketches are similar to their corresponding photos even though they are prone to shape distoration. This similarity between sketch and photo is captured and used for recognition across modalities. In this method, the face image characteristics are captured in the Difference of Gaussian (DoG) representation of the image patches. The Local Binary Pattern(LBP) corresponding to the DoG representation is then generated. These histograms are concatenated to generate the feature vector corresponding to input image. These feature vectors are compared using Earth Mover's Distance for recognition. Experiments on the CUFS(Chinese University of Hong Kong (CUHK) Face Sketch Database) and CUFSF (CUHK Face Sketch FERET Database) datesets prove the effectiveness of this feature in Face Sketch Recognition.
Ann Theja Alex, Vijayan K. Asari, Alex Mathew
SMC2
2013 Learning Multi-level Local Phase Relationship for Single Image Resolution Enhancement
abstract
In this paper, a novel approach for image spatial resolution enhancement based on multi-level local Fourier phase features is proposed. This method uses adaptive kernel regression technique based on multi-level local covariance to estimate the high resolution image from a low resolution input. However, this concept is similar to other regression and covariance based methods, our method uses multi-level Fourier image features to learn the local covariance from geometric similarity between low resolution image and its corresponding high resolution image. For each local region, four weighted integrated directional variances are estimated to adapt the interpolated pixels. This method is tested on various natural and aerial images at higher resolution scales. The results confirm that the proposed technique performs better especially at high resolution scales in comparison with other state of art techniques.
Saibabu Arigela, Vijayan K. Asari, Maher B. Qumsiyeh
SMC2
2013 3D Scene Reconstruction for Aiding Unmanned Vehicle Navigation
abstract
We present a 3D reconstruction algorithm designed to support various autonomous vehicle navigation applications. The algorithm presented focuses on the 3D reconstruction of a scene using only a single moving camera. Utilizing video frames captured at different points in time allows us to determine the relative depths in a scene. The original reconstruction process resulting in a point cloud was computed based on feature matching and depth triangulation analysis. In an improved version of the algorithm, we utilized optical flow features to create an extremely dense representation model. Although dense, this model is hindered due to its low disparity resolution. With the third algorithmic modification, we introduce the addition of the preprocessing step of nonlinear super resolution. With this addition, the accuracy and quantity of features is significantly increased since the number of features is directly proportional to the resolution and high frequencies of the input images. Our final contribution of additional pre and post processing steps are designed to filter noise points and mismatched features, completing the presentation of our Dense Point-cloud Representation (DPR) technique. We measure the success of DPR by evaluating the visual appeal, density, usability and computational expense of the reconstruction technique and compare with two state-of-the-art techniques.
Yakov Diskin, Vijayan K. Asari
SMC2
2013 Regression Based Learning of Human Actions from Video Using HOF-LBP Flow Patterns
abstract
A human action recognition framework is proposed which models motion variations corresponding to a particular class of actions without the need for sequence length normalization. The motion descriptors used in this framework are based on the optical flow vectors computed at every point on the silhouette of the human body. Histogram of flow(HOF) is computed from the optical flow vectors and these give the motion orientation in a local neighborhood. To get a relationship between the motion vectors at a particular instant, the magnitude and direction of the optical flow vector are coded with local binary patterns(LBP). The concatenation of these histograms(HOF-LBP) are considered as the action feature set to be used in the proposed framework. We illustrate that this motion descriptor is suitable for classifying various human actions when used in conjunction with the proposed action recognition framework which models the motion variations in time for each class using regression based techniques. The feature vectors extracted from the training set are suitably mapped to a lower dimensional space using Empirical Orthogonal Functional Analysis. A regression based technique such as Generalized Regression Neural Networks(GRNN), are used to compute the functional mapping from the action feature vectors to its reduced Eigenspace representation for each class, thereby obtaining separate action manifolds. The feature set obtained from a test sequence are compared with each of the action manifolds by comparing the test coefficients with the ones corresponding to the manifold (as estimated by GRNN) to determine the class using Mahalanobis distance.
Binu M. Nair, Vijayan K. Asari
SMC2
2013 Whale blow detection in infrared video using fractal analysis as tool for representing dynamic shape variation
abstract
In this paper we propose a new method to detect characteristic shape variations in infrared video using the concept of fractals. The proof of the concept is presented in terms of an application to detect whale blows in infrared video. This application will be of immense use to researchers who study whale behavior like its migration patterns. Whale blows appear as characteristic patterns with relatively higher intensity when compared to its surroundings. The pattern of whale blow formation is detected using fractal analysis. We have been able to experimentally prove that the fractal dimension can be used as an efficient tool to detect the presence of whale blows subject to some thresholding constraints. The development of thresholding constraints based on the variation of the coefficient of variation is also explained in the paper. We have been able to detect whale blows in infrared video with a very high degree of accuracy irrespective of the size of the whale blow and its distance from the camera.
Varun Santhaseelan, Vijayan K. Asari
WACV2
2013 A Two-Layer Framework for Piecewise Linear Manifold-Based Head Pose Estimation
Jacob Foytik, Vijayan K. Asari
Int. J. Comput. Vis.2
2013 Human action recognition using hull convexity defect features with multi-modality setups
M. M. Youssef, Vijayan K. Asari
Pattern Recognit. Lett.2
2012 Time Invariant Gesture Recognition by Modelling Body Posture Space
Binu M. Nair, Vijayan K. Asari
IEA/AIE2
2012 Gradient feature matching for expression invariant face recognition using single reference image
abstract
Automatic recognition of human faces irrespective of the expression variations is a challenging problem. In this paper, we propose a novel method for face recognition based on `edge-strings'. Experimental studies on face perception have shown the significance of edge features in visual perception and learning. In the proposed technique, the edges of a face are identified, and a feature string is created from edge pixels. This forms a symbolic descriptor corresponding to the edge image referred to as `edge-string'. The `edge-strings' are then compared using the Smith-Waterman algorithm to match them. The class corresponding to each image is identified based on the number of string primitives that match. Local string alignment algorithm is more robust to noise than global alignment algorithm; it gives better performance even if the input image is noisy. In addition, this method needs only a single training image per class. The proposed technique is a good solution for expression invariant face recognition. The effectiveness of the proposed method is compared with state-of-the-art algorithms on the Yale Face database, the Japanese Female Face Expression database (JAFFE) and CMU AMP Face EXpression database.
Ann Theja Alex, Vijayan K. Asari, Alex Mathew
SMC2
2012 Local region statistical distance measure for tracking in Wide Area Motion Imagery
abstract
In this paper we propose a novel tracking method in Wide Area Motion Imagery (WAMI) data based on local region histogram feature and a statistical distance measure. The aspects that make tracking particularly challenging are global camera motion, large movement of targets, poor gradient and texture information and absence of color information. Global camera motion is reduced or eliminated by registering the images from frame to frame employing SURF (Speeded Up Robust Feature). The proposed method is based on a variant of intensity histogram that encodes both spatial and intensity information. The method is evaluated on aerial WAMI data. The robustness of the feature eliminates the need for background subtraction in videos. A performance comparison of our feature descriptor with other descriptors such as HOG (Histogram of Gradients), SURF and SIFT (Scale Invariant Feature Transform) shows the effectiveness of the proposed method. We also show a comparison of our method with mean-shift tracking to show its effectiveness in tracking on WAMI data.
Alex Mathew, Vijayan K. Asari
SMC2
2010 A second order polynomial based subspace projection method for dimensionality reduction
abstract
A novel feature extraction method that utilizes nonlinear mapping from the original data space to the feature space is presented in this paper. For most practical systems, the meaningful features of a pattern class lie in a low dimensional nonlinear constraint region (manifold) within the high dimensional data space. A learning algorithm to model this nonlinear region and to project patterns to this feature space is developed. Least squares estimation approach that utilizes interdependency between points in training patterns is used to form the nonlinear region. A feature space encompassing multiple pattern classes can be trained by modeling a separate constraint region for each pattern class and obtaining a mean constraint region by averaging all the individual regions. Unlike most other nonlinear techniques, the proposed method provides an easy intuitive way to place new points onto a nonlinear region in the feature space. Classification accuracy is further improved by introducing the concepts of modularity and discriminant analysis into the proposed method.
Praveen Sankaran, Vijayan K. Asari
ICIP2
2010 Learning as a nonlinear line of attraction in a recurrent neural network
Ming-Jung Seow, Vijayan K. Asari, Adam R. Livingston
Neural Comput. Appl.2
2009 Towards representation of a perceptual color manifold using associative memory for color constancy
Ming-Jung Seow, Vijayan K. Asari
Neural Networks2
2009 Facial Recognition Using Multisensor Images Based on Localized Kernel Eigen Spaces
abstract
A feature selection technique along with an information fusion procedure for improving the recognition accuracy of a visual and thermal image-based facial recognition system is presented in this paper. A novel modular kernel eigenspaces approach is developed and implemented on the phase congruency feature maps extracted from the visual and thermal images individually. Smaller sub-regions from a predefined neighborhood within the phase congruency images of the training samples are merged to obtain a large set of features. These features are then projected into higher dimensional spaces using kernel methods. The proposed localized nonlinear feature selection procedure helps to overcome the bottlenecks of illumination variations, partial occlusions, expression variations and variations due to temperature changes that affect the visual and thermal face recognition techniques. AR and Equinox databases are used for experimentation and evaluation of the proposed technique. The proposed feature selection procedure has greatly improved the recognition accuracy for both the visual and thermal images when compared to conventional techniques. Also, a decision level fusion methodology is presented which along with the feature selection procedure has outperformed various other face recognition techniques in terms of recognition accuracy.
Satyanadh Gundimada, Vijayan K. Asari
IEEE Trans. Image Process.2
2008 Image enhancement for improving face detection under non-uniform lighting conditions
abstract
A new wavelet-based image enhancement algorithm is proposed to improve performance of face detection in non-uniform lighting environment with high dynamic range. Wavelet transform is used for dimension reduction so that dynamic range compression with local contrast enhancement algorithm is applied only to the approximation coefficients. The normalized approximation coefficients are transformed using a hyperbolic sine curve which achieves dynamic range compression. Contrast enhancement is realized by tuning the magnitude of each coefficient with respect to its surroundings. The detail coefficients are also modified to prevent the edge deformation. Experimental results on the proposed algorithm show improvement on the performance of the Viola-Jones face detector when compared to other prominent enhancement techniques.
Numan Unaldi, Praveen Sankaran, Vijayan K. Asari, Zia-ur Rahman 0001
ICIP3
2008 Design of a systolic-pipelined architecture for real-time enhancement of color video stream based on an illuminance-reflectance model
Hau T. Ngo, Vijayan K. Asari, Ming Z. Zhang
Integr.2
2007 Design and Implementation of an Efficient and Power-Aware Architecture for Skin Segmentation in Color Video Stream
abstract
In this paper, an efficient design for the high performance, power-aware architecture to extract skinlike regions in the video stream is presented. Skin segmentation is an important step in many image processing and computer vision applications such as face detection and hand gesture recognition. The design utilizes the high correlation and similarity of neighboring pixels in video streams to reduce switching activity (hence reducing dynamic power dissipation) in the arithmetic unit. The proposed design is implemented and fitted in the Altera's Cyclone II FPGA which is available in the DE2 development and educational board. The pipelined system is capable of performing the skin segmentation procedure in real-time with a processing rate of 654 frames per second for video frames with standard size of 640*480. It is observed that the proposed design helps to reduce operations and switching activities in the processing unit up to 42 percent which results in lower dynamic power dissipation with low hardware overhead.
Hau T. Ngo, Satyanadh Gundimada, Vijayan K. Asari
ASAP3
2007 A New Framework for Automatic Feature Selection for Tracking
abstract
A new framework of recurrent neural network is proposed in this paper for automatic feature selection for tracking. The network is not designed particularly for conventional applications such as pattern classification, association, and recognition; instead, it captures parts of those ingredients for identification of unique features from given sets of data. The architecture extracts different types of textures defined by natural importance to the datasets. These textural layers are then fused into single layer feature where the neurons compete and converge with few iterations based on the criteria of uniqueness of the textually maximized features. The automatically selected features by winning neurons, if any, are determined once and applied for subsequent feature tracking within the same architecture. Experiments performed on video sequence showed that the framework for feature selection and tracking is acceptable to gradual in-plane rotation and some degree of scale and out-of-plane rotation.
Ming Z. Zhang, Vijayan K. Asari
IJCNN2
2007 An efficient multiplier-less architecture for 2-D convolution with quadrant symmetric kernels
Ming Z. Zhang, Vijayan K. Asari
Integr.2
2006 An Adaptive Weight Assignment Scheme in Linear Subspace Approaches for Face Recognition
Satyanadh Gundimada, Vijayan K. Asari
ACCV (2)2
2006 On the Divergence Dynamics of the Nonlinear Line of Attraction
abstract
In designing a recurrent neural network, it is usually of prime importance to guarantee the convergence in the dynamics of the network. We propose to modify this picture: if the brain remembers by converging to the state representing familiar patterns, it should also diverge from such states when presented with an unknown encoded representation of a visual image. We propose to capture this behavior using a nonlinear line attractor network. This model encapsulates attractive fixed points scattered in the state space representing patterns with similar characteristics as an attractive curved line. The dynamics of the nonlinear line attractor network is designed such that when the network is able to reach equilibrium (stable), the input is considered as one of the stored patterns. Conversely, when the network is unable to reach equilibrium (unstable), the input is considered to be dissimilar to the stored patterns and therefore is considered as pattern of another class. Several experiments on benchmark problems have shown that the proposed model can be very useful for discriminating patterns.
Ming-Jung Seow, Vijayan K. Asari
IJCNN2
2006 A Hardware Architecture for Color Image Enhancement Using a Machine Learning Approach with Adaptive Parameterization
abstract
A novel architecture for performing color image enhancement using a machine learning algorithm called Ratio Rule is proposed in this paper. The threshold width of the activation function is automatically determined from the image characteristics. The approach promotes log-domain computation to eliminate all multiplications and divisions, utilizing approximation techniques for efficient estimation of the log2and inverse-log2. The design incorporates the dynamic thresholds of the activation functions and update rate. The improved quadrant symmetric architecture is also presented to provide very high throughput rate for homomorphic filters which is part of the pixel intensity enhancement across RGB components in the system. The pipelined design of the filter features the flexibility in reloading a wide range of kernels for different frequency responses. A new approach for the design of the uniform filters is also presented to reduce the processing element arrays (PEAs) from W PEAs to 2 PEAs for W × W window. This new concept is applied to assist in training the synaptic weights of the neural network for color balancing to restore the intensity enhanced image to its natural color existed in the original image. The concept of uniform filter is further extended to design max/min filters. It is observed that the performance of the system with parallel and pipelined architectures is able to achieve 139.3 million outputs per second (MOPS), or equivalently 54.7 billion operations per second on Xilinx's Virtex II XC2V2000-4ff896 FPGA at a clock frequency of 139.3 MHz.
Ming Z. Zhang, Ming-Jung Seow, Vijayan K. Asari
IJCNN3
2006 Robust Learning by Self-organization of Nonlinear Lines of Attractions
Ming-Jung Seow, Vijayan K. Asari
ISNN (1)2
2006 Ratio rule and homomorphic filter for enhancement of digital colour image
Ming-Jung Seow, Vijayan K. Asari
Neurocomputing2
2006 Recurrent neural network as a linear attractor for pattern association
abstract
We propose a linear attractor network based on the observation that similar patterns form a pipeline in the state space, which can be used for pattern association. To model the pipeline in the state space, we present a learning algorithm using a recurrent neural network. A least-squares estimation approach utilizing the interdependency between neurons defines the dynamics of the network. The region of convergence around the line of attraction is defined based on the statistical characteristics of the input patterns. Performance of the learning algorithm is evaluated by conducting several experiments in benchmark problems, and it is observed that the new technique is suitable for multiple-valued pattern association.
Ming-Jung Seow, Vijayan K. Asari
IEEE Trans. Neural Networks2
2005 On using an associative memory for improving digital color images: color characterization, enhancement, and color balancing
abstract
Actual observed scenes usually produce a wide dynamic range. Currently available visual systems use low dynamic range light detector that provide only 8 bits of brightness information at each pixel. This greatly limits what visual systems can do for surveillance applications such as face detection and face recognition. A color image enhancement procedure based on the concept of color characterization, enhancement, and color balancing is proposed in this paper. The enhancement technique directly operates on pixels using a hyperbolic tangent function to increase the dynamic range of the pixel. The global and local statistics of the image is used to control the curvature of the hyperbolic tangent function. The color characterization and color balancing processes are based on a new nonlinear line attractor network to create a color manifold to restore the relationship of red, green, and blue components of the pixels. The proposed enhancement approach greatly improves the dynamic range compression and color rendition of an image.
Ming-Jung Seow, Vijayan K. Asari
IJCNN2
2005 Associative Memory Using Nonlinear Line Attractor Network for Multi-valued Pattern Association
Ming-Jung Seow, Vijayan K. Asari
ISNN (1)2
2005 Color Characterization and Balancing by a Nonlinear Line Attractor Network for Image Enhancement
Ming-Jung Seow, Vijayan K. Asari
Neural Process. Lett.2
2005 A pipelined architecture for real-time correction of barrel distortion in wide-angle camera images
abstract
An efficient pipelined architecture for the real-time correction of barrel distortion in wide-angle camera images is presented in this paper. The distortion correction model is based on least-squares estimation to correct the nonlinear distortion in images. The model parameters include the expanded/corrected image size, the back-mapping coefficients, distortion center, and corrected center. The coordinate rotation digital computer (CORDIC) based hardware design is suitable for an input image size of 1028/spl times/1028 pixels and is pipelined to operate at a clock frequency of 40 MHz. The VLSI system will facilitate the use of a dedicated hardware that could be mounted along with the camera unit.
Hau T. Ngo, Vijayan K. Asari
IEEE Trans. Circuits Syst. Video Technol.2
2004 Learning skin distribution using a sparse map
abstract
We present a new skin modeling technique based on SNoW (sparse network of Winnows) for accurate and robust skin region detection. A skin distribution map (SDM) representing the sparse network is trained with skin pixels to learn their distribution in a color space. We then train the SDM with non-skin pixels to unlearn the distribution of the non-skin pixels, which overlap with the skin pixels in the color space. This skin model can be used for skin detection on any color space. We have found the accuracy of skin detection using SDM to be slightly better than that using the skin probability map (SPM) method. The main advantage of using the SDM method over the SPM method is that the complexity, memory requirements and time for skin detection are reduced significantly.
Rajkiran Gottumukkal, Vijayan K. Asari
ICIP2
2004 Face detection technique based on intesity and skin color distribution
abstract
A rotation invariant human face detection system in color images based on human skin color distribution and intensity is proposed in this paper. Skin color distribution typical to a human face is used as a feature along with the intensity variations to classify the candidate regions into faces and nonfaces. The detection process is carried out in YCbCr color space. Sparse Network of Winnows architecture is used to train three networks one for intensity and two for the color distributions for classification of candidate regions. Rotation invariance in detection of faces is achieved by training multiple classifiers, each to detect faces at a particular orientation. The detection process also implements a non linear luminance based lighting compensation method which is very efficient in enhancing and restoring the natural colors into the images which are taken in darker and varying lighting conditions. Experimental results show that the new face detection technique is highly efficient in terms of speed and accuracy in detecting frontal view faces at different orientations in complex environments.
Satyanadh Gundimada, Vijayan K. Asari
ICIP3
2004 Homomorphic processing system and ratio rule for color image enhancement
abstract
Homomorphic filter is an illumination-reflectance model that can be used to develop a frequency domain procedure for improving the appearance of an image by simultaneous gray-level range compression and contrast enhancement. Many previously reported methods on homomorphic filter for color images shows that the homomorphic filter consistently provides excellent dynamic range compression but is lacking final color rendition. We present a novel color image enhancement process to overcome this limitation. The color image enhancement process involved using a neural network algorithm, namely ratio rule, to pre-process and post-process the color image in the homomorphic system. This method improves the appearance of images as perceived by the human eye, and/or to render these images more suitable for computer analysis. That is, both color rendition and dynamic range compression are achieved using this method.
Ming-Jung Seow, Vijayan K. Asari
IJCNN2
2004 Recurrent Network as a Nonlinear Line Attractor for Skin Color Association
Ming-Jung Seow, Vijayan K. Asari
ISNN (1)2
2004 Design of an efficient VLSI architecture for non-linear spatial warping of wide-angle camera images
Vijayan K. Asari
J. Syst. Archit.1
2004 An improved face recognition technique based on modular PCA approach
Rajkiran Gottumukkal, Vijayan K. Asari
Pattern Recognit. Lett.2
2004 Learning using distance based training algorithm for pattern recognition,
Ming-Jung Seow, Vijayan K. Asari
Pattern Recognit. Lett.2
2003 System level design of real time face recognition architecture based on composite PCA
abstract
Design and implementation of a fast parallel architecture based on an improved principal component analysis (PCA) method called Composite PCA suitable for real-time face recognition is presented in this paper. The proposed architecture performs the tasks of both feature extraction and classification. Composite PCA takes in to consideration the local features of face images, which do not vary widely between face images of the same person taken under varying expression, illumination and pose. Hence it leads to a better recognition rate than PCA. Composite PCA has more parallelism than conventional PCA and this parallelism is utilized to design an efficient architecture capable of performing real-time face recognition. The face recognition system is implemented in an FPGA environment and tested using standard databases. The system is able to recognize a person from a database of 110 images of 10 individuals in approximately 4 ms.
Rajkiran Gottumukkal, Vijayan K. Asari
ACM Great Lakes Symposium on VLSI2
2003 High performance associative memory with distance based training algorithm for character recognition
abstract
The consequence of reducing the impact of the synaptic weights from neurons farther away from the neuron under consideration on a modular two-dimensional Hopfield network using Hebbian learning rule is examined for image processing applications. A generalized modular architecture is developed by defining each module with a group of neighboring neurons and all modules communicating with each other. A spatially decaying distance factor is introduced into the Hebbian rule to reduce the effect of neurons from farther modules. A biologically inspired visual perception concept has been adopted for defining the variation of the distance factor. The performance of the new technique is evaluated by conducting several experiments on character images and it is observed that the proposed method increases the learning ability and convergence rate of the network. The nature of the distance factor helps the removal of several synaptic weights farther away from a particular neuron and this leads to the reduction of complexity of the network in terms of both software and hardware implementation.
Ming-Jung Seow, Vijayan K. Asari
IJCNN2
2003 Associative memory using ratio rule for multi-valued pattern association
abstract
A novel learning algorithm, named ratio rule, for association of multi-valued patterns in a recurrent neural network is proposed in this paper. The learning is performed based on the degree of similarity between the relative magnitudes of the output of each neuron with respect to that of all other neurons. The dynamics of the neural network functions as a line attractor as opposed to the common concept of point attractor. The limit of the convergence region around the line of attraction is defined based on the statistical characteristics of the input patterns. Theoretical analysis of the associativity of the network with the ratio rule confirms the authenticity of its learning ability. The performance of the ratio rule on associativity and convergence of the recurrent network is evaluated by conducting several experiments on face images. It is observed that the ratio rule is suitable for retention, reconstruction, and restoration of learned patterns with varying face expressions.
Ming-Jung Seow, Vijayan K. Asari
IJCNN2
2003 L2-norm approximation based learning in recurrent neural networks for expression invariant face recognition
abstract
A new learning algorithm based on L2-norm approximation to define the relationship between two neurons in a recurrent neural network is proposed in this paper. The learning process utilizes the statistical relationship between each component of the input pattern with respect to every other component. The activation function of a neuron is a rectangular function whose position changes adaptively with respect to the input pattern and its left and right wings are decided by the mean of maximum variations of the training signals to that neuron. The new training algorithm is applied for recognition of faces images with varying expressions. 975 face images of 13 persons from the Carnegie Mellon University (CMU) face expression variant database are used for evaluating the performance of the network. The network has been trained with 5 images and tested with the remaining 70 images of each person. The recurrent neural network with the new learning algorithm recognized all the 13 persons in this database without error.
Ming-Jung Seow, Deepthi Valaparla, Vijayan K. Asari
SMC3
2002 Unbiased frequency estimation of narrowband signals using Procrustes type subspace rotation
abstract
We consider an improved least square algorithm based on subspace rotations using Procrustes approximation. Robust estimation of the poles of a narrowband signal is achieved by reducing the bias of the estimates. The key idea is to consider the signal as a vector in multi-dimensional space, and to separate the “pure signal” and the noise into two mutually orthogonal components lying in different subspaces. Though the quality of the estimated signal is not much improved as compared to the conventional least square based subspace approaches, the advantage of using the proposed algorithm is the asymptotically unbiased nature of the estimates.
Joseph S. Paul, Chirag B. Patel, Vijayan K. Asari, David L. Sherman
ICASSP3
2002 Segmenting endoscopic images using adaptive progressive thresholding: a hardware perspective
Vijayan K. Asari, Thambipillai Srikanthan
J. Syst. Archit.1
2001 Training of a feedforward multiple-valued neural network by error backpropagation with a multilevel threshold function
abstract
A technique for the training of multiple-valued neural networks based on a backpropagation learning algorithm employing a multilevel threshold function is proposed. The optimum threshold width of the multilevel function and the range of the learning parameter to be chosen for convergence are derived. Trials performed on a benchmark problem demonstrate the convergence of the network within the specified range of parameters.
Vijayan K. Asari
IEEE Trans. Neural Networks1
1999 A New Approach for Nonlinear Distortion Correction in Endoscopic Images Based on Least Squares Estimation
abstract
Images captured with a typical endoscope show spatial distortion, which necessitates distortion correction for subsequent analysis. In this paper, a new methodology based on least squares estimation is proposed to correct the nonlinear distortion in the endoscopic images. A mathematical model based on polynomial mapping is used to map the images from distorted image space onto the corrected image space. The model parameters include the polynomial coefficients, distortion center, and corrected center. The proposed method utilizes a line search approach of global convergence for the iterative procedure to obtain the optimum expansion coefficients. A new technique to find the distortion center of the image based on curvature criterion is presented. A dual-step approach comprising token matching and integrated neighborhood search is also proposed for accurate extraction of the centers of the dots contained in a rectangular grid, used for the model parameter estimation. The model parameters were verified with different grid patterns. The distortion-correction model is applied to several gastrointestinal images and the results are presented. The proposed technique provides high-speed response and forms a key step toward online camera calibration, which is required for accurate quantitative analysis of the images.
Vijayan K. Asari, Sanjiv Kumar, D. Radhakrishnan
IEEE Trans. Medical Imaging1
1994 An Optimization Technique for the Design of Multiple Valued PLA's
abstract
An optimization technique for the design of two types of multiple-valued PLAs is described. In a type-I PLA, the multiple-valued function is realized directly, whereas in a type-II PLA, output encoding is used to encode the binary output of the PLA. In both types, multiple function literal circuits are used for the purpose of minimization. It is shown that the proposed technique leads to a considerably reduced size of PLA when compared to the earlier techniques.>
Vijayan K. Asari, C. Eswaran
IEEE Trans. Computers1