Parham Aarabi

dblp:a/ParhamAarabi · DBLP profile ↗
← Back
67ranked-venue papers
13as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2024 SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation
abstract
Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task, existing state-of-the-art (SOTA) methods (1) extract coarse features from backbones that are trained with instance-level self-supervised learning (SSL) paradigms, which neglect the dense prediction nature of the task, (2) aggregate them into memory-intensive hypercolumn formations, and (3) su-pervise lightweight projector networks to naï vely establish full local correspondences among all pairs of spatial features. In this paper, we introduce SCE-MAE, a framework that (1) leverages the MAE [14], a region-level SSL method that naturally better suits the landmark prediction task, (2) operates on the vanilla feature map instead of on expen-sive hypercolumns, and (3) employs a Correspondence Ap-proximation and Refinement Block (CARB) that utilizes a simple density peak clustering algorithm and our proposed Locality-Constrained Repellence Loss to directly hone only select local correspondences. We demonstrate through extensive experiments that SCE-MAE is highly effective and robust, outperforming existing SOTA methods by large mar-gins of ~20%-44% on the landmark matching and ~9%-15% on the landmark detection tasks.
Kejia Yin, Varshanth S. Rao, Ruowei Jiang, Parham Aarabi, David B. Lindell
CVPR5
2023 Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers
abstract
Vision Transformers (ViT) have shown competitive advantages in terms of performance compared to convolutional neural networks (CNNs), though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearby tokens to accelerate the ViT's multi-head self-attention (MHSA) operations. However, such structured attention patterns limit the token-to-token connections to their spatial relevance, which disregards learned semantic connections from a full attention mask. In this work, we propose an approach to learn instance-dependent attention patterns, by devising a lightweight connectivity predictor module that estimates the connectivity score of each pair of tokens. Intuitively, two tokens have high connectivity scores if the features are considered relevant either spatially or semantically. As each token only attends to a small number of other tokens, the binarized connectivity masks are often very sparse by nature and therefore provide the opportunity to reduce network FLOPs via sparse computations. Equipped with the learned unstructured attention pattern, sparse attention ViT (Sparsifiner) produces a superior Pareto frontier between FLOPs and top-1 accuracy on ImageNet compared to token sparsity. Our method reduces 48% ~ 69% FLOPs of MHSA while the accuracy drop is within 0.4%. We also show that combining attention and token sparsity reduces ViT FLOPs by over 60%.
Cong Wei 0001, Brendan Duke, Ruowei Jiang, Parham Aarabi, Graham W. Taylor, Florian Shkurti
CVPR4
2022 Exploring Gradient-Based Multi-directional Controls in GANs
Ruowei Jiang, Brendan Duke, Han Zhao 0002, Parham Aarabi
ECCV (23)5
2022 Synthesizing ultraviolet skin images via GAN with Gaussian weighted patch blending
abstract
In this work, we explore a novel application of synthesizing ultraviolet skin images from RGB images using an unpaired training framework for image-to-image translation. To synthesize high resolution outputs, we propose a novel Gaussian-based patch blending technique that is designed following the characteristics of GANs. Specifically, we weigh the pixels at the same coordinates among multiple generated patches based on their distance to the center point. Our proposed method is performant, taking 0.72s for a whole face image at resolution 960x720 at inference time and generates realistic-looking ultraviolet images. We also show high correspondence of our synthesized images with the true ultraviolet images qualitatively. Finally, our novel blending approach achieves significant improvements compared with other blending methods.
Ruowei Jiang, Brendan Duke, Frédéric Flament, Parham Aarabi
ISM4
2021 SSTVOS: Sparse Spatiotemporal Transformers for Video Object Segmentation
abstract
In this paper we introduce a Transformer-based approach to video object segmentation (VOS). To address compounding error and scalability issues of prior work, we propose a scalable, end-to-end method for VOS called Sparse Spatiotemporal Transformers (SST). SST extracts per-pixel representations for each object in a video using sparse attention over spatiotemporal features. Our attention-based formulation for VOS allows a model to learn to attend over a history of multiple frames and provides suitable inductive bias for performing correspondence-like computations necessary for solving motion segmentation. We demonstrate the effectiveness of attention-based over recurrent networks in the spatiotemporal domain. Our method achieves competitive results on YouTube-VOS and DAVIS 2017 with improved scalability and robustness to occlusions compared with the state of the art. Code is available at https://github.com/dukebw/SSTVOS.
Brendan Duke, Abdalla Ahmed, Christian Wolf 0001, Parham Aarabi, Graham W. Taylor
CVPR4
2021 Continuous Face Aging via Self-Estimated Residual Age Embedding
abstract
Face synthesis, including face aging, in particular, has been one of the major topics that witnessed a substantial improvement in image fidelity by using generative adversarial networks (GANs). Most existing face aging approaches divide the dataset into several age groups and leverage group-based training strategies, which lacks the ability to provide fine-controlled continuous aging synthesis in nature. In this work, we propose a unified network structure that embeds a linear age estimator into a GAN-based model, where the embedded age estimator is trained jointly with the encoder and decoder to estimate the age of a face image and provide a personalized target age embedding for age progression/regression. The personalized target age embedding is synthesized by incorporating both personalized residual age embedding of the current age and exemplar-face aging basis of the target age, where all preceding aging bases are derived from the learned weights of the linear age estimator. This formulation brings the unified perspective of estimating the age and generating personalized aged face, where self-estimated age embeddings can be learned for every single age. The qualitative and quantitative evaluations on different datasets further demonstrate the significant improvement in the continuous face aging aspect over the state-of-the-art.
Zeqi Li, Ruowei Jiang, Parham Aarabi
CVPR3
2021 LOHO: Latent Optimization of Hairstyles via Orthogonalization
abstract
Hairstyle transfer is challenging due to hair structure differences in the source and target hair. Therefore, we propose Latent Optimization of Hairstyles via Orthogonalization (LOHO), an optimization-based approach using GAN inversion to infill missing hair structure details in latent space during hairstyle transfer. Our approach decomposes hair into three attributes: perceptual structure, appearance, and style, and includes tailored losses to model each of these attributes independently. Furthermore, we propose two-stage optimization and gradient orthogonalization to enable disentangled latent space optimization of our hair attributes. Using LOHO for latent space manipulation, users can synthesize novel photorealistic images by manipulating hair attributes either individually or jointly, transferring the desired attributes from reference hairstyles. LOHO achieves a superior FID compared with the current state-of-the-art (SOTA) for hairstyle transfer. Additionally, LOHO preserves the subject’s identity comparably well according to PSNR and SSIM when compared to SOTA image embedding pipelines. Code is available at https://github.com/dukebw/LOHO.
Rohit Saha, Brendan Duke, Florian Shkurti, Graham W. Taylor, Parham Aarabi
CVPR5
2020 Semantic Relation Preserving Knowledge Distillation for Image-to-Image Translation
Zeqi Li, Ruowei Jiang, Parham Aarabi
ECCV (26)3
2020 Dynamic Memory Regeneration
abstract
In this paper, we examine a practical implementation of dynamic memory regeneration for the purpose of treating memory loss. Inspired by the refresh cycles in computer Dynamic Random Access Memory (DRAM) circuits, we propose a condensed audiovisual summary that refreshes a subject's memory within a specific amount of time. We explore the impact of the audiovisual refresh cycle time on the overall probability of memory loss, and examine how this refresh time is related to parameters such as refresh cycle duration, and the intrinsic memory loss time. Based on several simplifying assumptions, we propose a model for an ideal memory refresh duration that would minimize the probability of memory loss.
Pegah Aarabi, Parham Aarabi
SMC2
2019 Quantitative Measurement of VR Stereoscopic Video Recording Quality Based on Visual Acuity Loss
abstract
A quality of experience measure for virtual reality is important for both understanding the limitations of existing technology and to provide a focus for specific improvements necessary for virtual reality to achieve full realism. In this paper, we provide a mechanism for the quantitative measurement of scene reconstruction in virtual reality based on visual acuity loss (VAL) as measured by both a virtual and real LogMAR eye chart tests. We illustrate that this test is viable for measuring the entire stack of visual recording and playback in virtual reality, and provide quantitative results for popular virtual reality recording and playback device combinations. We specifically show that visual acuity loss appears to be primarily based on the recording quality, rather than the playback device.
Parham Aarabi, Tzu-An Chen, Vladislav Il'govskiy, Anastasia Kolesnikov, Nathaniel Xu, Benzakhar Manashirov
MMSP1
2019 Virtual Fakes: DeepFakes for Virtual Reality
abstract
The proliferation of data and computational resources has led into many advancements in computer vision for facial data including easily replacing a face in one video with another one, the so called DeepFake. In this paper, we apply techniques to generate DeepFakes for virtual reality applications. We empirically validate our method by generating, for the first time, Deep Fake videos in virtual reality.
Joey Bose, Parham Aarabi
MMSP2
2018 Adversarial Attacks on Face Detectors Using Neural Net Based Constrained Optimization
abstract
Adversarial attacks involve adding, small, often imperceptible, perturbations to inputs with the goal of getting a machine learning model to misclassifying them. While many different adversarial attack strategies have been proposed on image classification models, object detection pipelines have been much harder to break. In this paper, we propose a novel strategy to craft adversarial examples by solving a constrained optimization problem using an adversarial generator network. Our approach is fast and scalable, requiring only a forward pass through our trained generator network to craft an adversarial sample. Unlike in many attack strategies we show that the same trained generator is capable of attacking new images without explicitly optimizing on them. We evaluate our attack on a trained Faster R-CNN face detector on the cropped 300-W face dataset where we manage to reduce the number of detected faces to 0.5% of all originally detected faces. In a different experiment, also on 300-W, we demonstrate the robustness of our attack to a JPEG compression based defense typical JPEG compression level of 75% reduces the effectiveness of our attack from only 0.5% of detected faces to a modest 5.0%.
Joey Bose, Parham Aarabi
MMSP2
2018 Hybrid eye center localization using cascaded regression and hand-crafted model fitting
Alex Levinshtein, Edmund Phung, Parham Aarabi
Image Vis. Comput.3
2018 Hair Segmentation Using Heuristically-Trained Neural Networks
abstract
We present a method for binary classification using neural networks (NNs) that performs training and classification on the same data using the help of a pretraining heuristic classifier. The heuristic classifier is initially used to segment data into three clusters of high-confidence positives, high-confidence negatives, and low-confidence sets. The high-confidence sets are used to train an NN, which is then used to classify the low-confidence set. Applying this method to the binary classification of hair versus nonhair patches, we obtain a 2.2% performance increase using the heuristically trained NN over the current state-of-the-art hair segmentation method.
Wenzhangzhi Guo, Parham Aarabi
IEEE Trans. Neural Networks Learn. Syst.2
2017 A convolutional neural network for search term detection
abstract
Pathfinding in hospitals is challenging for patients, visitors, and even employees. Many people have experienced getting lost due to lack of clear guidance, large footprint of hospitals, and confusing array of hospital wings. In this paper, we propose Halo; An indoor navigation application based on voice-user interaction to help provide directions for users without assistance of a localization system. The main challenge is accurate detection of origin and destination search terms. A custom convolutional neural network (CNN) is proposed to detect origin and destination search terms from transcription of a submitted speech query. The CNN is trained based on a set of queries tailored specifically for hospital and clinic environments. Performance of the proposed model is studied and compared with Levenshtein distance-based word matching.
Hojjat Salehinejad, Joseph Barfett, Parham Aarabi, Shahrokh Valaee, Errol Colak, Bruce Gray, Timothy Dowdell
PIMRC3
2016 7 surprising lessons learned from teaching iOS programming to 30, 000+ MOOC students
abstract
In this paper, we experimentally explore the impact of different teaching paradigms on teaching a large-scale iOS programming MOOC consisting of 30,162 students. Our initial approach utilizes methods from our in-person lecturing experience. After launching the initial version of the course, we analyzed our performance and student feedback based on which we recreated and re-launched the entire course with a particular focus on clarity, video-quality, and packaging the topics in 1-2 minute micro-modules. Based on feedback from 650 students, we observed that the overall lecture positive feedback increased from 65% before our course adjustment to 83% after the adjustment. In this paper we will provide a detailed overview of the lessons learned and the impact of our new teaching methods on each section of the course. We also noticed a slight increase in course completion rates, from 5.4% before the adjustment to 5.8% after the adjustment.
Parham Aarabi, Narges Norouzi, Jack Wu, Michael Spears
FIE1
2016 Hierarchical differential image filters for skin analysis
abstract
In this paper we present a framework for analyzing skin parameters from portrait images and videos. Using a series of Hierarchical Differential Image Filters (HDIF), it becomes possible to detect different skin features such as wrinkles, spots, and roughness. These detected features are used to compute skin ratings that are compared to actual ratings by dermatologists. Analyzing a database of 49 images with ratings by a panel of dermatologists, the proposed HDIF method is able to detect skin roughness, dark spots, and deep wrinkles with an average rating error of 11.3%, 17.6%, and 15.6%, respectively, as compared to individual dermatologist rating errors of 8.2%, 7.4%, and 6.5%. Although dermatologist ratings are more accurate than the proposed HDIF method, the ratings are close enough that the HDIF ratings can be a viable solution where dermatologist ratings are not readily available.
Parham Aarabi
MMSP2
2015 Automatic Segmentation of Hair in Images
abstract
Based on a multi-step process, an automatic hair segmentation method is created and tested on a database of 115 manually segmented hair images. By extracting various information components from an image, including background color, face position, hair color, skin color, and skin mask, a heuristic-based method is created for the detection and segmentation of hair that can detect hair with an accuracy of approximately 75% and with a false hair overestimation error of 34%. Furthermore, it is shown that downsampling the image down to a face width of 25px results in a 73% reduction in computation time with insignificant change in detection accuracy.
Parham Aarabi
ISM1
2015 Precise Skin-Tone and Under-Tone Estimation by Large Photo Set Information Fusion
abstract
This paper proposes a novel method for the estimation of a person's skin-tone and under-tone by analyzing a large collection of photos of that person. By excluding badly lit images, and analyzing well-lit skin pixels, it becomes possible to compute an overall skin-tone estimate which is in-line with the person's true skin shade, and based on this, to determine a person's under-tone. Based on a study involving 15,590 user sessions and 104,366 photos, it was found that the proposed methodology can detect the normalized RGB of the person's skin-tone with 2.3% RMSE, or based on the CIE76 color difference measure, obtain an average Delta E color difference of 3.15 in L*a*b* color space.
Parham Aarabi, Benzakhar Manashirov, Edmund Phung, Kyung Moon Lee
ISM1
2015 Quantitative Evaluation of Hair Texture
abstract
In this paper, we quantitatively evaluate the role of texture in hair patches, with a primary motivation of understanding what can be learned and applied by machine learning systems for texture-based hair detection. We evaluate the distribution of gradient directions in hair patches, and explore the relation between proximity to the face and the angle of the gradients for 2,870,000 hair patches selected from 100 manually silhouetted hairstyles.
Wenzhangzhi Guo, Parham Aarabi
ISM2
2014 Skin lens: Skin assessment video filters
abstract
This paper presents a framework for real-time skin analysis. Our system consists of two filters for skin redness analysis and pigmented skin lesions analysis. In skin redness analysis, we search for skin regions in video frames and assess size and degree of redness in skin regions. In order to analyze pigmented skin lesions we localize potential moles of different sizes utilizing multiple difference of Gaussian filters and apply trained Support Vector Machine and K-Nearest Neighbors to eliminate non-mole candidates from the previous stage. After finding mole candidates we present information about the symmetry and border irregularity of candidate moles. To test real-time performance of our system a prototype application on iPhone 4S and iPhone 5S was implemented.
Inaz Alaei-Novin, Parham Aarabi
SMC2
2014 Extracting deep social relationships from photos
abstract
Hidden within the relative location of tags in images is a relational model that can identify how close two individuals are, or, the affinity of a person to an object or a brand. Based on this model we can 1) better understand the relationship between users/tags, 2) find photos where a user is pictured but not tagged in, and 3) enable searching “inside” images by clicking on any location within an image to start a search. This paper proposes a method of modeling the relationship between objects based on their spatial arrangement in a set of tagged images. Based on the relative coordinates of each object tag, we compute a joint relativity between each tag pair, generate a social relationship graph and propose an efficient image search method using the joint Relativity graph. We evaluated our approach with real world data from Facebook, showing a direct relationship between the number of tagged photos and the amount of information obtained from these photos, and an average correlation coefficient of 0.8 between user-generated relativity scores and those obtained by our algorithm.
Yelei Lu, Parham Aarabi
SMC2
2013 Extended touch mobile user interfaces through sensor fusion
Tusi Chowdhury, Parham Aarabi, Weijian Zhou, Yuan Zhonglin
FUSION2
2013 Fusion of spatial and visual information for object tracking on iPhone
Amin Heidari, Inaz Alaei-Novin, Parham Aarabi
FUSION3
2013 Extended touch user interfaces
abstract
This paper explores an extended touch-based input interface for mobile devices by inferring the user-tapped location on any neighboring surface. We propose the feasibility of achieving a portable solution by using only the device microphone, without needing multiple sensors or highly specialized single piezoelectric sensor as has been done in past research. Training and classification of discrete tap locations is done using cross-correlation analysis on the audio feature vector. We first explore a binary classifier that performs pair-wise detection of taps, which results in a 92% detection rate across different surfaces. The solution is extended to multi-tap detection using a K-Nearest Neighbor (kNN) algorithm trained on 17 output classes, resulting in 92% knuckle tap and 63% soft tap detection.
Tusi Chowdhury, Parham Aarabi, Weijian Zhou, Yuan Zhonglin
ICME2
2013 Relational Social Image Search
abstract
This paper proposes a method of finding the relationship between objects based on their spatial arrangement in a set of tagged images. Based on the relative coordinates of each object tag, we compute a joint Relativity between each tag pair. We then propose an efficient image search method using the joint Relativity graphs and provide simple examples where the proposed Relational Social Image (RSI) search produces more relevant and intuitive results than simple search.
Parham Aarabi
ISM1
2011 Tiny Videos: A Large Data Set for Nonparametric Video Retrieval and Frame Classification
abstract
In this paper, we present a large database of over 50,000 user-labeled videos collected from YouTube. We develop a compact representation called "tiny videos" that achieves high video compression rates while retaining the overall visual appearance of the video as it varies over time. We show that frame sampling using affinity propagation-an exemplar-based clustering algorithm-achieves the best trade-off between compression and video recall. We use this large collection of user-labeled videos in conjunction with simple data mining techniques to perform related video retrieval, as well as classification of images and video frames. The classification results achieved by tiny videos are compared with the tiny images framework [24] for a variety of recognition tasks. The tiny images data set consists of 80 million images collected from the Internet. These are the largest labeled research data sets of videos and images available to date. We show that tiny videos are better suited for classifying scenery and sports activities, while tiny images perform better at recognizing objects. Furthermore, we demonstrate that combining the tiny images and tiny videos data sets improves classification precision in a wider range of categories.
Alexandre Karpenko, Parham Aarabi
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Intelligent ad resizing
abstract
No abstract available.
Anthony P. Badali, Parham Aarabi, Ron D. Appel
WWW2
2009 FLoSS: Facility location for subspace segmentation
abstract
Subspace segmentation is the task of segmenting data lying on multiple linear subspaces. Its applications in computer vision include motion segmentation in video, structure-from-motion, and image clustering. In this work, we describe a novel approach for subspace segmentation that uses probabilistic inference via a message-passing algorithm. We cast the subspace segmentation problem as that of choosing the best subset of linear subspaces from a set of candidate subspaces constructed from the data. Under this formulation, subspace segmentation corresponds to facility location, a well studied operational research problem. Approximate solutions to this NP-hard optimization problem can be found by performing maximum-a-posteriori (MAP) inference in a probabilistic graphical model. We describe the graphical model and a message-passing inference algorithm. We demonstrate the performance of Facility Location for Subspace Segmentation, or FLoSS, on synthetic data as well as on 3D multi-body video motion segmentation from point correspondences.
Nevena Lazic, Inmar E. Givoni, Brendan J. Frey, Parham Aarabi
ICCV4
2009 Fourier-based Rotation Invariant image features
abstract
Fourier Coefficients have long been used to achieve invariance to signal transformations. For the purposes of image processing, the magnitude of the Fourier transform has been used in conjunction with other transforms to achieve invariance to rotation. In this paper we propose a Rotation Invariant Descriptor for matching images based on features derived from the Discrete Fourier Transform (DFT). The features combine both the phase and the magnitude information to achieve invariance. Experiments are conducted to show the robustness of these features under changes of scale and compression of images.
Sam Mavandadi, Parham Aarabi, Konstantinos N. Plataniotis
ICIP2
2009 Learning from 1, 000, 000 user-uploaded faces
abstract
This paper combines two face detection algorithms to create a hybrid method which is more accurate and robust than either of the original methods. The face detectors are compared using a database of 1,000 user uploaded photos, which is a small subset of a much larger 1 million photo database that was generated through a popular online application that enables people to upload facial photos and to accurately locate the face in each photo.
Benzakhar Manashirov, Parham Aarabi
ICME2
2009 Evaluating real-time audio localization algorithms for artificial audition in robotics
abstract
Although research on localization of sound sources using microphone arrays has been carried out for years, providing such capabilities on robots is rather new. Artificial audition systems on robots currently exist, but no evaluation of the methods used to localize sound sources has yet been conducted. This paper presents an evaluation of various real-time audio localization algorithms using a medium-sized microphone array which is suitable for applications in robotics. The techniques studied here are implementations and enhancements of steered response power - phase transform beamformers, which represent the most popular methods for time difference of arrival audio localization. In addition, two different grid topologies for implementing source direction search are also compared. Results show that a direction refinement procedure can be used to improve localization accuracy and that more efficient and accurate direction searches can be performed using a uniform triangular element grid rather than the typical rectangular element grid.
Anthony P. Badali, Jean-Marc Valin, François Michaud, Parham Aarabi
IROS4
2009 Tiny Videos: A Large Dataset for Image and Video Frame Categorization
abstract
This paper presents a new method for video and image categorization based on a database of over 50,000 videos collected from YouTube and down-sampled to tiny size. The categorization results achieved by tiny videos are compared with the tiny images framework for a variety of recognition tasks. The tiny images dataset consists of 80 million images collected from the Internet. These are the largest labeled research datasets of videos and images available to date. We show that tiny videos are better suited for classifying sports activities and scenery, while tiny images perform better at recognizing objects. Furthermore, we demonstrate that combining the tiny images and tiny videos datasets improves categorization precision in a wider range of categories.
Alexandre Karpenko, Parham Aarabi
ISM2
2008 Tiny Videos: Non-parametric Content-Based Video Retrieval and Recognition
abstract
This work extends the tiny images techniques developed by Torralba et al. to videos. A dataset of 6,612 videos was collected from YouTube in the Sports and News sections. We present a method for compressing the temporal dimension nonuniformly using affinity propagation. We show that nonuniform sampling using affinity propagation outperforms temporal sampling at uniform intervals, because it covers a greater range of visual appearances in the video for the same number of samples. We examine two main applications for the tiny video dataset: duplicate video detection and related video retrieval. We also show that the scope of text-based searches on YouTube can be significantly increased by incorporating visual similarity.
Alexandre Karpenko, Parham Aarabi
ISM2
2008 Spoken Term Detection Using Visual Spectrogram Matching
abstract
This work proposes a novel spoken term detection technique, where the query is in audio format. Detection and retrieval are performed by matching the spectrograms of the spoken document and query as visual images, using ideas from computer vision. Local descriptors are computed on a dense grid over each spectrogram, and the query term is detected using deformable template matching of grids. Detection experiments are performed on an hour-long newscast recording, involving 10 query terms of length 2-3 words. When the query term comes from the document, nearly all other instances of the term in the document are detected; performance degrades when the query is recorded by the user.
Nevena Lazic, Parham Aarabi
ISM2
2008 A Cyclic Interface for the Presentation of Multiple Music Files
abstract
This paper proposes a novel cyclic interface for browsing through a song database. The method, which sums multiple audio streams on a server and broadcasts only a single summed stream, allows the user to hear different parts of each audio stream by cycling through all available streams. Songs are summed into a single stream based on a combination of spectral entropy and local power of each song's waveform. Perceptual parameters of the system are determined based on experiments conducted on 20 users, for three, four, and five songs. Results illustrate that the proposed methodology requires less listening time as compared to traditional list-based interfaces when the desired audio clip is among one of the audio streams. Applications of this methodology include any search system which returns multiple audio search results, including music query by example. The proposed methodology can be used for real-time searching with an ordinary internet browser.
Sarah Ali, Parham Aarabi
IEEE Trans. Multim.2
2007 Face detection using information fusion
abstract
The fundamental point of this paper is that the fusion of several simple, somewhat unreliable, and somewhat inefficient frontal face detectors results in an efficient and reliable frontal face detector which, without any training, performs similarly to a state-ofthe- art neural network based face detector trained on 60,000 images. The simple detectors used include a skin detector, symmetry detectors, as well as structural face detectors. On a test set of 30 color images containing frontal faces, the fused face detector had an accuracy of 93% with a RMSE of 4.96 pixels, as compared to an accuracy of 87% and a RMSE of 8.00 pixels for the neural network based face detector. On the Caltech Face Database, the fused face detector had a 90% detection rate which is on par with state-of-the-art face detection methods that utilize extensive prior training, including the neural network approach which achieves a detection rate of 94%.
Parham Aarabi, Jerry Chi-Ling Lam, Arezou Keshavarz
FUSION1
2007 Importance of Feature Locations in Bag-of-Words Image Classification
abstract
The impact of image feature locations in the bag-of-words model for object classification is examined. It is demonstrated that a simple variance-based method works well and offers advantages over several other methods. In essence, the feature locations are selected intelligently, decreasing the redundancy and cost sometimes associated with feature extraction on dense grids. Classification results on two databases are presented, using a support vector machine classifier.
Nevena Lazic, Parham Aarabi
ICASSP (1)2
2007 Rotation Invariance in Images
abstract
Rotation is one of the most basic transformations that can relate two images. Two determine if two images are rotated versions of each other, one can either exhaustively rotate them in order to find out if the two match up at some angle, or alternatively extract features from the images that can then be compared to make the same decision. In this paper, we will propose a novel method for extracting the components of an image that are invariant to rotation based on the Fourier transform. We will compare the performance of the algorithm to the exhaustive search method and show that this is a much faster technique, that is also accurate in matching rotated images.
Sam Mavandadi, Parham Aarabi
ICASSP (1)2
2007 Variational Probabilistic Speech Separation Using Microphone Arrays
abstract
Separating multiple speech sources using a limited number of noisy sensor measurements presents a difficult problem, but one that is of great practical interest. Although previously introduced source separation methods [such as independent component analysis (ICA)] can be made to work in many situations, most of these methods fail when the sensors are very noisy or when the number of sources exceeds the number of sensors. Our approach to this problem is to combine the multiple sensor likelihoods [obtained using time-delay-of-arrival (TDOA) information] with a generative probability model of the sources. This model accounts for the power spectrum of each source using a mixture model, and accounts for the phase of each source using one discretized hidden phase variable for each frequency. Source separation is achieved by identifying the source vector configuration of maximum a posteriori probability, given all available information. An exhaustive search for the MAP configuration is computationally intractable, but we present an efficient variational technique that performs approximate probabilistic inference. For the problem of separating delayed additive noise corrupted speech mixtures, the algorithm is able to improve upon the signal-to-noise ratio (SNR) gain performance of existing state-of-the-art probabilistic and TDOA-based speech separation algorithms by over 10 dB. This significant performance improvement is obtained by combining the information utilized by these approaches intelligently under a representative probabilistic description of the speech production and mixing process. The method is capable of recovering high fidelity estimates of the underlying speech sources even when there are more sources than microphone observations
Steven J. Rennie, Parham Aarabi, Brendan J. Frey
IEEE Trans. Speech Audio Process.2
2007 Phase-Based Dual-Microphone Speech Enhancement Using A Prior Speech Model
abstract
This paper proposes a phase-based dual-microphone speech enhancement technique that utilizes a prior speech model. Recently, it has been shown that phase-based dual-microphone filters can result in significant noise reduction in low signal-to-noise ratio [(SNR) less than 10 dB] conditions and negligible distortion at high SNRs (greater than 10 dB), as long as a correct filter parameter is chosen at each SNR. While prior work utilizes a constant parameter for all SNRs, we present an SNR-adaptive filter parameter estimation algorithm that maximizes the likelihood of the enhanced speech features based on a prior speech model. Experimental results using the CARVUI database show significant speech recognition accuracy rate improvement over alternative techniques in low SNR situations (e.g., an improvement of 11% in word error rate (WER) over postfiltering and 23% over delay-and-sum beamforming at 0 dB) and negligible distortion at high SNRs. The proposed adaptive approach also significantly outperforms the original phase-based filter with a constant parameter. Furthermore, it improves the filter's robustness when there are errors in time delay estimation
Guangji Shi, Parham Aarabi, Hui Jiang 0001
IEEE Trans. Speech Audio Process.2
2006 Face Fusion: An Automatic Method For Virtual Plastic Surgery
abstract
This paper describes a system that replaces an individual's facial features with corresponding features of another individual -possibly of different skin color- and fuses the replaced features with the original face, such that the resulting face looks natural. The final face resulting from the fusion of the original face with exogenous features, lacks the characteristic discontinuities that would have been expected if only a replacement operation was performed. The proposed system could be used to simulate and predict the outcome of aesthetic maxillofacial plastic surgeries. To achieve its task, the system uses five modules: face detection, feature detection, replacement, shifting and blending. While these modules are designed to address the problem of face fusion, some of the novel algorithms and techniques introduced in this paper could be useful in other image processing and fusion applications
Alireza Seyed Rabi, Parham Aarabi
FUSION2
2006 A Novel Interface for Audio Search
abstract
In this paper a novel cyclic interface for searching through a song database is proposed. The method, which merges multiple audio streams on a server and broadcasts only a single merged stream, allows the user to hear different parts of each audio stream by cycling through all available streams. Experimental results on 21 users illustrate that the proposed interface requires less listening time as compared to traditional list-based interfaces when the desired song/audio clip is among one of the audio streams. The average search time for the proposed interface was 7.3 seconds, compared to 12.1 seconds for the traditional list-based interface when searching for a song which is included among the audio streams
Sarah Ali, Parham Aarabi
ICME2
2006 Predictive Dynamic User Interfaces for Interactive Visual Search
abstract
This paper proposes a method for designing user interfaces based on ideas rooted in data communication theory. It suggests that a visual user interface should be treated as a multitransmitter, single-receiver communication system, where the total available bandwidth for transmission is limited. The proposed design entails the scaling of visual components that are displayed according to their degree of relevance to the user, or in other words, their probability of selection by the user
Sam Mavandadi, Parham Aarabi, Azadeh Khaleghi, Ron D. Appel
ICME2
2006 Sound Localization-Based Navigational User Interfaces
abstract
In this paper, we propose and compare three navigational user interfaces that are based on acoustic information. The user can navigate through a Web page or change slides in a presentation by changing its spatial location. To this end, we employ an array of 24 microphones to record speech signals and localize the speaker. The proposed interfaces are: 1) The parallax system, which behaves as if the user is looking out of a window, 2) The corner system, which resizes each image based on the proximity of the speaker to the corners of the environment, and 3) The tiled system, which divides the environment into a number of tiles and loads the image corresponding to the tile on which the user is standing on. A set of experiments were performed on 15 participants, which showed that the interaction time required for the tiled system and the corner system is far less than that required for the parallax system. The participants were also asked to rank the ease of use of each of the interfaces; it was observed that the users are more comfortable with an interface that requires minimal interaction time
Arezou Keshavarz, Parham Aarabi
ISM2
2006 On the importance of phase in human speech recognition
abstract
In this paper, we analyze the effects of uncertainty in the phase of speech signals on the word recognition error rate of human listeners. The motivating goal is to get a quantitative measure on the importance of phase in automatic speech recognition by studying the effects of phase uncertainty on human perception. Listening tests were conducted for 18 listeners under different phase uncertainty and signal-to-noise ratio (SNR) conditions. These results indicate that a small amount of phase error or uncertainty does not affect the recognition rate, but a large amount of phase uncertainty significantly affects the recognition rate. The degree of the importance of phase also seems to be an SNR-dependent one, such that at lower SNRs the effects of phase uncertainty are more pronounced than at higher SNRs. For example, at an SNR of -10 dB, having random phases at all frequencies results in a word error rate (WER) of 63% compared to 24% if the phase was unaltered. In comparison, at 0 dB, random phase results in a 25% WER as compared to 11% for the unaltered phase case. Listening tests were also conducted for the case of reconstructed phase based on the least square error estimation approach. The results indicate that the recognition rate for the reconstructed phase case is very close to that of the perfect phase case (a WER difference of 4% on average).
Guangji Shi, Maryam Modir Shanechi, Parham Aarabi
IEEE Trans. Speech Audio Process.3
2006 Communication Over an Acoustic Channel Using Data Hiding Techniques
abstract
This work proposes an audio data hiding system that hides information into signals not known beforehand. The system can hide data into live music, or ambient sounds in general, and can be used to communicate information acoustically from one device to another. An important benefit of such a system is backward-compatibility, as the transmitter is a speaker, and the receiver is a microphone, both of which are already present in numerous devices and environments. The highest data rate achieved was 213 bits/s, while keeping the error rate under 10%. The developed technique was applied in a simple navigation system, where acoustic data embedded into background music indicates the location of the receiver
Nevena Lazic, Parham Aarabi
IEEE Trans. Multim.2
2006 Real-time face detection and lip feature extraction using field-programmable gate arrays
abstract
This paper proposes a new technique for face detection and lip feature extraction. A real-time field-programmable gate array (FPGA) implementation of the two proposed techniques is also presented. Face detection is based on a naive Bayes classifier that classifies an edge-extracted representation of an image. Using edge representation significantly reduces the model's size to only 5184 B, which is 2417 times smaller than a comparable statistical modeling technique, while achieving an 86.6% correct detection rate under various lighting conditions. Lip feature extraction uses the contrast around the lip contour to extract the height and width of the mouth, metrics that are useful for speech filtering. The proposed FPGA system occupies only 15050 logic cells, or about six times less than a current comparable FPGA face detection system.
Duy Cuong Nguyen, David Halupka, Parham Aarabi, Ali Sheikholeslami
IEEE Trans. Syst. Man Cybern. Part B3
2005 Real-time dual-microphone speech enhancement using field programmable gate arrays
abstract
The paper discusses an implementation of a dual-microphone phase-based speech enhancement technique. By using the phases of the incoming sound signals, we mask frequencies with low signal-to-noise ratio (SNR) between the two microphones. Phase-based filtering can achieve high SNR gains with just two microphones, making it ideal for hand-held devices. However, these devices have a limited battery life and lack the processing power needed for a software based implementation. The paper presents a field programmable gate array (FPGA) implementation that was designed specifically for low-power operation. The FPGA based implementation is compared, with respect to processing capabilities and power utilization, with an off-the-shelf low-power digital signal processor (DSP) implementation.
David Halupka, Alireza Seyed Rabi, Parham Aarabi, Ali Sheikholeslami
ICASSP (5)3
2004 Integrated displacement tracking and sound localization
abstract
An algorithm is proposed for the robust localization of a vehicle using both displacement tracking and sound localization. The displacement tracking is performed by optical encoders that enable the turn angle and the movement distance of the vehicle to be estimated. The sound localization utilizes a speaker mounted on the vehicle and an array of 24 microphones deployed in the environment. The two modalities are integrated by modeling the displacement tracking uncertainty by a Gaussian mixture model (GMM) and combining it with the probability distribution obtained from the sound localization system. It is shown that the proposed integrated system results in an average localization error (at best 11 cm) that is better than either modality alone.
Parham Aarabi, QingHua Wang, Maoud Yeganegi
ICASSP (5)1
2004 Distributed spectrum estimation in sensor networks
abstract
This paper considers the problem of fusing the statistical information gained by a distributed network of sensors. We formulate the problem as a convex feasibility problem, that is, finding a point in the intersection of finitely many closed convex sets. We then present a distributed optimization algorithm to solve the problem.
Omid S. Jahromi, Parham Aarabi
ICASSP (3)2
2004 Multiple-microphone time-varying filters for robust speech recognition
abstract
A multiple microphone time varying filter that is an extension of the dual-microphone speech enhancement technique of P. Aarabi et al. (see Proceedings of the IEEE Conference on Multimedia and Expo, Baltimore, Maryland, July 2003) is proposed and experimentally analyzed. The technique utilizes information regarding the locations of the speech source of interest and the microphones to compute a time varying filter that results in substantial noise reduction over other speech enhancement techniques such as delay-and-sum beamforming and superdirective beamforming. For example, digit recognition results in an environment with two speakers and a reverberation time of 0.1s show a recognition accuracy rate increase of 25.2% over delay-and-sum beamforming and an increase of 26.5% over superdirective beamforming using six microphones.
Calvin Yiu-Kit Lai, Parham Aarabi
ICASSP (1)2
2004 Localization-based sensor validation using the Kullback-Leibler divergence
abstract
A sensor validation criteria based on the sensor's object localization accuracy is proposed. Assuming that the true probability distribution of an object or event in space f(x) is known and a spatial likelihood function (SLF) psi(x) for the same object or event in space is obtained from a sensor, then the expected value of the SLF E[psi(x)] is proposed as a suitable validity metric for the sensor, where the expectation is performed over the distribution f(x). It is shown that for the class of increasing linear log likelihood SLFs, the proposed validity metric is equivalent to the Kullback-Leibler distance between f(x) and the unknown sensor-based distribution g(x) where the SLF psi(x) is an observable increasing function of the unobservable g(x). The proposed technique is illustrated through several simulated and experimental examples.
Parham Aarabi
IEEE Trans. Syst. Man Cybern. Part B1
2004 Phase-based dual-microphone robust speech enhancement
abstract
A dual-microphone speech-signal enhancement algorithm, utilizing phase-error based filters that depend only on the phase of the signals, is proposed. This algorithm involves obtaining time-varying, or alternatively, time-frequency (TF), phase-error filters based on prior knowledge regarding the time difference of arrival (TDOA) of the speech source of interest and the phases of the signals recorded by the microphones. It is shown that by masking the TF representation of the speech signals, the noise components are distorted beyond recognition while the speech source of interest maintains its perceptual quality. This is supported by digit recognition experiments which show a substantial recognition accuracy rate improvement over prior multimicrophone speech enhancement algorithms. For example, for a case with two speakers with a 0.1 s reverberation time, the phase-error based technique results in a 28.9% recognition rate gain over the single channel noisy signal, a gain of 22.0% over superdirective beamforming, and a gain of 8.5% over postfiltering.
Parham Aarabi, Guangji Shi
IEEE Trans. Syst. Man Cybern. Part B1
2004 Enhanced sound localization
abstract
A new approach to sound localization, known as enhanced sound localization, is introduced, offering two major benefits over state-of-the-art algorithms. First, higher localization accuracy can be achieved compared to existing methods. Second, an estimate of the source orientation is obtained jointly, as a consequence of the proposed sound localization technique. The orientation estimates and improved localizations are a result of explicitly modeling the various factors that affect a microphone's level of access to different spatial positions and orientations in an acoustic environment. Three primary factors are accounted for, namely the source directivity, microphone directivity, and source-microphone distances. Using this model of the acoustic environment, several different enhanced sound localization algorithms are derived. Experiments are carried out in a real environment whose reverberation time is 0.1 seconds, with the average microphone SNR ranging between 10-20 dB. Using a 24-element microphone array, a weighted version of the SRP-PHAT algorithm is found to give an average localization error of 13.7 cm with 3.7% anomalies, compared to 14.7 cm and 7.8% anomalies with the standard SRP-PHAT technique.
Bobji Mungamuru, Parham Aarabi
IEEE Trans. Syst. Man Cybern. Part B2
2003 Time delay estimation and signal reconstruction using multi-rate measurements
abstract
This paper considers the problem of fusing two low-rate sensors (e.g., microphones) for reconstructing one high-resolution signal when time delay of arrival (TDOA) is present as well. We show that under certain conditions the phase of the cross-spectrum-density of low-rate measurements becomes independent of the signal in the high-rate front end of the system. We then utilize this fact to demonstrate that it is possible to extend a class of TDOA estimation techniques known as the generalized cross correlation technique to linear-phase multi-rate sensor systems. Finally, we illustrate how the combination of the theory of linear-phase multirate filter banks and TDOA estimation can result in a practical, multi-sensor signal reconstruction system.
Omid S. Jahromi, Parham Aarabi
ICASSP (6)2
2003 Real-time sound localization using field-programmable gate arrays
abstract
This paper presents a single FPGA implementation of a real-time sound localization system using two microphones. The implementation, utilizing a cross-correlation technique based on a modified version of the phase transform, successfully localizes sound sources in noisy environments with as low an SNR as 10 dB. Using the same algorithm and similar hardware architecture, it is shown that up to 5 parallel systems (using 10 microphones), all real-time, can be implemented on a single FPGA while only utilizing an estimated 77mW-108mW per microphone.
Duy Cuong Nguyen, Parham Aarabi, Ali Sheikholeslami
ICASSP (2)2
2003 Robust variational speech separation using fewer microphones than speakers
abstract
A variational inference algorithm for robust speech separation, capable of recovering the underlying speech sources even in the case of more sources than microphone observations, is presented. The algorithm is based upon a generative probabilistic model that fuses time-delay of arrival (TDOA) information with prior information about the speakers and application, to produce an optimal estimate of the underlying speech sources. Simulation results are presented for the case of two, three and four underlying sources and two microphone observations corrupted by noise. The resulting SNR gains (32 dB with two sources, 23 dB with three sources, and 16 dB with four sources) are significantly higher than previous speech separation techniques.
Steven J. Rennie, Parham Aarabi, Trausti T. Kristjansson, Brendan J. Frey, Kannan Achan
ICASSP (1)2
2003 Robust digit recognition using phase-dependent time-frequency masking
abstract
A technique using the time-frequency phase information of two microphones is proposed to estimate an ideal time-frequency mask using time-delay-of-arrival (TDOA) of the signal of interest. At a signal-to-noise ratio (SNR) of 0 dB, the proposed technique using two microphones achieves a digit recognition rate (average over 5 speakers, each speaking 20-30 digits) of 71%. In contrast, delay-and-sum beamforming only achieves a 40% recognition rate with two microphones and 60% with four microphones. Superdirective beamforming achieves a 44% recognition rate with two microphones and 65% with four microphones.
Guangji Shi, Parham Aarabi
ICASSP (1)2
2003 Scene reconstruction using distributed microphone arrays
abstract
A method for the joint localization and orientation estimation of a directional sound source using distributed microphones is presented. By modeling the signal attenuation due to the microphone directivity, the source directivity, and the source-microphone distance, a multi-dimensional search over all possible sound source scene reconstruction algorithm is presented in the context of an experiment with 24 microphones and a dynamic speech source. At a signal-to-noise ratio of 20 dB and with a reverberation time of approximately 0.1 s, accurate location estimates (20 cm error) and orientation estimates (less than 10/spl deg/ average error) are obtained.
Parham Aarabi, Bobji Mungamuru
ICME1
2003 Robust speech separation using time-frequency masking
abstract
A multi-microphone time-frequency speech masking technique is proposed. This technique utilizes both the time-frequency magnitude and phase information in order to estimate the signal-to-noise ratio (SNR) maximizing masking coefficients for each time-frequency block given that the direction (or alternatively, the time-delay of arrival) of the speaker of interest is known. Using this masking algorithm, speech features (such as formants) from the direction of interest are preserved while features from other directions are severely degraded. Digit recognition experiments indicate that the proposed technique can result in a substantial increase in the digit recognition accuracy rate. At 0 dB, for example, the proposed technique results in a digit recognition accuracy rate improvement of 26% over the single microphone case and an improvement of 12% over the two microphone superdirective beamforming case.
Parham Aarabi, Guangji Shi, Omid S. Jahromi
ICME1
2003 Time delay estimation and signal reconstruction using multi-rate measurements
abstract
This paper considers the problem of fusing two low-rate sensors (e.g., microphones) for reconstructing one high-resolution signal when time delay of arrival (TDOA) is present as well. We show that under certain conditions the phase of the cross-spectrum-density of low-rate measurements becomes independent of the signal in the high-rate front end of the system. We then utilize this fact to demonstrate that it is possible to extend a class of TDOA estimation techniques known as the generalized cross correlation technique to linear-phase multi-rate sensor systems. Finally, we illustrate how the combination of the theory of linear-phase multi-rate filter banks and TDOA estimation can result in a practical, multi-sensor signal reconstruction system.
Omid S. Jahromi, Parham Aarabi
ICME2
2003 Real-time sound localization using field-programmable gate arrays
abstract
This paper presents a single FPGA implementation of a real-time sound localization system using two microphones. The implementation, utilizing a cross-correlation technique based on a modified version of the phase transform, successfully localizes sound sources in noisy environments with as low an SNR as 10 dB. Using the same algorithm and similar hardware architecture, it is shown that up to 5 parallel systems (using 10 microphones), all real-time, can be implemented on a single FPGA while only utilizing an estimated 77 mW-108 mW per microphone.
Duy Cuong Nguyen, Parham Aarabi, Ali Sheikholeslami
ICME2
2003 Robust variational speech separation using fewer microphones than speakers
abstract
A variational inference algorithm for robust speech separation, capable of recovering the underlying speech sources even in the case of more sources than microphone observations, is presented. The algorithm is based upon an generative probabilistic model that fuses time-delay of arrival (TDOA) information with prior information about the speakers and application, to produce an optimal estimate of the underlying speech sources. Simulation results are presented for the case of two, three and four underlying sources and two microphones observations corrupted by noise. The resulting SNR gains (32 dB with two sources, 23 dB with three sources, and 16 dB with four sources) are significantly higher than previous speech separation techniques.
Steven J. Rennie, Parham Aarabi, Trausti T. Kristjansson, Brendan J. Frey, Kannan Achan
ICME2
2003 Robust digit recognition using phase-dependent time-frequency masking
abstract
A technique using the time-frequency phase information of two microphones is proposed to estimate an ideal time-frequency mask using time-delay-of-arrival (TDOA) of the signal of interest. At a signal-to-noise ratio (SNR) of 0dB, the proposed technique using two microphones achieves a digit recognition rate (average over 5 speakers, each speaking 20-30 digits) of 71%. In contrast, delay-and-sum beamforming only achieves a 40% recognition rate with two microphones and 60% with four microphones. Superdirective beamforming achieves a 44% recognition rate with two microphones and 65% with four microphones.
Guangji Shi, Parham Aarabi
ICME2
2002 The relation between speech segment selectivity and source localization accuracy
abstract
An experimental analysis of the relation between speech signal segment power and the source direction-of-arrival-estimation accuracy is conducted. A total of 10 different speakers, including both male and female speakers, totaling to approximately 2 hours of speech are used to analyze the performance of the Phase Transform, the Maximum Likelihood, and the Unfiltered Cross Correlation time-delay estimation techniques. For female speakers, it is determined that the Phase Transform technique has a lower percentage of anomalies and a lower direction-of-arrival root mean-square error (DOA RMSE). Conversely, for male speakers, it is determined that the Unfiltered Cross Correlation has a lower percentage of anomalies although the Phase Transform has a lower DOA RMSE. The spatial distribution of the errors as well as the speech segment power relation to the errors are also presented.
Parham Aarabi, Albarz Mahdavi
ICASSP1
2001 The automatic measurement of facial beauty
abstract
We develop an automatic facial beauty scoring system based on ratios between facial features. After isolating the face, eyes, eyebrows and mouth in a portrait photograph, we represent a face abstractly as an 8-element vector of ratios between these features. We use a variant of the K-nearest neighbor algorithm, in the context of a parameterized metric space optimized using a genetic algorithm, to learn a beauty assignment function from a training set of photographs rated by humans. We assess performance on a test set of photographs, concluding that when facial ratios are accurately extracted in the computer vision phase, the results of the program are highly correlated with median-human ratings of beauty.
Parham Aarabi, Dominic Hughes, Keyvan Mohajer, Majid Emami
SMC1