Harry Wechsler

dblp:w/HarryWechsler · DBLP profile ↗
← Back
139ranked-venue papers
12as first author
0since 2021 · last 2019
0000-0002-5454-8819ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 112 · 10 first-authorGraphics, computer vision, multimedia, augmented reality and games · 54 · 2 first-authorSecurity and privacy · 7Human-computer interaction and ubiquitous computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
22 papers
Face, body and person analysis · 36% Learning theory · 17% Representation and self-supervised learning · 13%
Computer graphics and multimedia
7 papers
Image and video processing · 98% Multimedia analysis and retrieval · 2%
Databases, data mining, and information retrieval
4 papers
Data mining · 81% Information retrieval · 19%
Network and information security
2 papers
Biometric security · 100%

Topics — the 30 heaviest of 50, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
face recognition
0.3102005
Open Set Face Recognition Using Transduction · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition · IEEE Trans. Image Process. 2002
A shape- and texture-based enhanced Fisher classifier for face recognition · IEEE Trans. Image Process. 2001
Machine learning › Learning theory › hypothesis testing
exchangeability testing
0.112010
A Martingale Framework for Detecting Changes in Data Streams by Testing Exchangeability · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Data mining
anomaly detection
0.112010
A Martingale Framework for Detecting Changes in Data Streams by Testing Exchangeability · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Efficient and distributed learning
active learning
0.112008
Query by Transduction · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Machine learning › Transfer learning and domain adaptation
transductive inference
0.112008
Query by Transduction · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › discriminant analysis
fisher discriminant analysis
0.132002
Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition · IEEE Trans. Image Process. 2002
A shape- and texture-based enhanced Fisher classifier for face recognition · IEEE Trans. Image Process. 2001
Face Recognition Using Shape and Texture · CVPR 1999
Biometric security
open-set recognition
0.112005
Open Set Face Recognition Using Transduction · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Image and video processing
motion estimation
0.012004
Motion Estimation Using Statistical Learning Theory · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Machine learning › Learning theory › statistical learning theory
complexity control
0.012003
Controlling Model Complexity in Flow Estimation · ICCV 2003
Image and video processing › motion estimation
flow estimation
0.012003
Controlling Model Complexity in Flow Estimation · ICCV 2003
Image and video processing › motion estimation
optical flow
0.012003
Controlling Model Complexity in Flow Estimation · ICCV 2003
User interface design and tools
adaptive user interfaces
0.012002
Integrating perceptual and cognitive modeling for adaptive and intelligent human-computer interaction · Proc. IEEE 2002
Machine learning › Representation and self-supervised learning › representation learning
feature extraction
0.012001
A Gabor Feature Classifier for Face Recognition · ICCV 2001
Computer vision › Face, body and person analysis › face recognition › video-based face recognition
surveillance face recognition
0.011998
Face Surveillance · ICCV 1998
Computer vision › Image recognition and object detection
object recognition
0.031990
Selective and Focused Invariant Recognition Using Distributed Associative Memories (DAM) · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Distributed Associative Memory (DAM) for Bin-Picking · IEEE Trans. Pattern Anal. Mach. Intell. 1989
2-D Invariant Object Recognition Using Distributed Associative Memory · IEEE Trans. Pattern Anal. Mach. Intell. 1988
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation
0.012005
Open Set Face Recognition Using Transduction · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Image and video processing › motion analysis
motion tracking
0.012004
Motion Estimation Using Statistical Learning Theory · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Machine learning › Kernel, tree and ensemble methods
decision tree learning
0.011995
Hybrid Learning Using Genetic Algorithms and Decision Trees for Pattern Classification · IJCAI (1) 1995
Machine learning › Reinforcement learning › population-based learning
evolutionary learning
0.011995
Hybrid Learning Using Genetic Algorithms and Decision Trees for Pattern Classification · IJCAI (1) 1995
Machine learning › Optimization for machine learning › evolutionary computation
genetic algorithms
0.011995
Hybrid Learning Using Genetic Algorithms and Decision Trees for Pattern Classification · IJCAI (1) 1995
Computer vision › Image recognition and object detection › region localization
region of interest detection
0.011994
Integration of bottom-up and top-down cues for visual attention using non-linear relaxation · CVPR 1994
Machine learning › Deep learning architectures and training › attention mechanism
visual attention
0.011994
Integration of bottom-up and top-down cues for visual attention using non-linear relaxation · CVPR 1994
Computer vision › Image recognition and object detection › object recognition
invariant object recognition
0.021990
Selective and Focused Invariant Recognition Using Distributed Associative Memories (DAM) · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Invariant Object Recognition Using a Distributed Associative Memory · NIPS 1987
Robotics › Robot navigation and mapping
active perception
0.011992
Active Perception Using DAM and Estimation Techniques · ECCV 1992
Image and video processing › image representation
spatial-frequency representation
0.021990
Segmentation of Textured Images and Gestalt Organization Using Spatial/Spatial-Frequency Representations · IEEE Trans. Pattern Anal. Mach. Intell. 1990
A Theory for Invariant Object Recognition in the Frontoparallel Plane · IEEE Trans. Pattern Anal. Mach. Intell. 1984
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.011999
Face Recognition Using Shape and Texture · CVPR 1999
Image and video processing
low-level vision
0.011990
Segmentation of Textured Images and Gestalt Organization Using Spatial/Spatial-Frequency Representations · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Image and video processing
perceptual grouping
0.011990
Segmentation of Textured Images and Gestalt Organization Using Spatial/Spatial-Frequency Representations · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Image and video processing › image segmentation
texture segmentation
0.011990
Segmentation of Textured Images and Gestalt Organization Using Spatial/Spatial-Frequency Representations · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian classification
bayes classifier
0.011998
Probabilistic Reasoning Models for Face Recognition · CVPR 1998

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.3sliding window · 0.2martingale · 0.2kolmogorov complexity · 0.1k-nearest neighbors · 0.1fisher linear discriminant · 0.1principal component analysis · 0.1p-values · 0.1kullback-leibler divergence · 0.1enhanced fisher linear discriminant model · 0.1transductive inference · 0.1statistical learning theory · 0.0ridge regression · 0.0VC theory · 0.0perceptual modeling · 0.0nonverbal behavior parsing · 0.0cognitive modeling · 0.0decision tree · 0.0
YearPublicationVenuePosition
2019 Question action relevance and editing for visual question answering
Andeep S. Toor, Harry Wechsler, Michele Nappi
Multim. Tools Appl.2
2019 Biometric surveillance using visual question answering
Andeep S. Toor, Harry Wechsler, Michele Nappi
Pattern Recognit. Lett.2
2019 Modern art challenges face detection
Harry Wechsler, Andeep S. Toor
Pattern Recognit. Lett.1
2018 Visual Question Authentication Protocol (VQAP)
Andeep S. Toor, Harry Wechsler, Michele Nappi, Kim-Kwang Raymond Choo
Comput. Secur.2
2018 Biometrics and forensics integration using deep multi-modal semantic alignment and joint embedding
Andeep S. Toor, Harry Wechsler
Pattern Recognit. Lett.2
2017 Leveraging implicit demographic information for face recognition using a multi-expert system
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler
Multim. Tools Appl.4
2016 Eye movement analysis for human authentication: a critical survey
Chiara Galdi, Michele Nappi, Daniel Riccio, Harry Wechsler
Pattern Recognit. Lett.4
2016 Towards demographic categorization using gaze analysis
Chiara Galdi, Harry Wechsler, Virginio Cantoni, Marco Porta, Michele Nappi
Pattern Recognit. Lett.2
2015 Using Bipartite Graphs for 3D Cardiac Model Retrieval
abstract
Three-dimensional models have been used to aid medical diagnoses, using images generated by modalities like Magnetic Resonance Imaging. They can provide a more complete vision of objects since their depth is taken into account. Content-based Image Retrieval (CBIR) has also been used to aid the diagnosis. One important step in Three-dimensional CBIR (Model Retrieva) systems is the comparison between two models by using a set of features extracted and stored in a database. In this paper we present a novel method to compare two models, using the Bipartite graphs technique, with the aim to improve the retrieval precision. This technique retrieves 3D medical models of the left ventricle in order to aid the diagnosis of Congestive Heart Failure. Results showed that the novel method improved the precision by 10% when compared to the Similarity Function of Euclidean and Manhattan distance. These results confirmed that bipartite graph techniques can be used to improve the accuracy of Model Retrieval systems.
Leila C. C. Bergamasco, Hellyan Oliveira, Helton H. Biscaro, Harry Wechsler, Fátima L. S. Nunes
CBMS4
2015 Content-Based Image Retrieval of 3D Cardiac Models to Aid the Diagnosis of Congestive Heart Failure by Using Spectral Clustering
abstract
This paper describes a novel application of Content-Based Image Retrieval (CBIR) to search a medical database consisting of 3D models for diagnosis purposes. The 3D models, which are generated using Magnetic Resonance Imaging and include depth information, are used to search for similarity a database of 3D annotated medical cases using their pairwise feature similarity. The 3D models consist of both local and global feature descriptors that consider the surface of the 3D model and the overall geometry of the medical artifact. The models are then matched using spectral clustering that embeds the Euclidean distance for affinity and partitions the models into two groups, Congestive Heart Failure (CHF) and non-CHF. This suffices to demarcate using pairwise similarity the existence of CHF for the left ventricle. Experimental results using thirty 3D models show the utility of the new 3D method compared to existing methods. In particular, the novel method yields 83% overall accuracy.
Leila C. C. Bergamasco, Rafael Alves Paes de Oliveira, Harry Wechsler, Caina Dajuda, Márcio Eduardo Delamaro, Fátima L. S. Nunes
CBMS3
2015 Robust face recognition after plastic surgery using region-based approaches
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler
Pattern Recognit.4
2015 Mobile Iris Challenge Evaluation (MICHE)-I, biometric iris dataset and protocols
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler
Pattern Recognit. Lett.4
2014 Robust Face Recognition for Uncontrolled Settings
Harry Wechsler
ICPRAM1
2014 Identifying users with application-specific command streams
abstract
This paper proposes and describes an active authentication model based on user profiles built from user-issued commands when interacting with GUI-based application. Previous behavioral models derived from user issued commands were limited to analyzing the user's interaction with the *Nix (Linux or Unix) command shell program. Human-computer interaction (HCI) research has explored the idea of building users profiles based on their behavioral patterns when interacting with such graphical interfaces. It did so by analyzing the user's keystroke and/or mouse dynamics. However, none had explored the idea of creating profiles by capturing users' usage characteristics when interacting with a specific application beyond how a user strikes the keyboard or moves the mouse across the screen. We obtain and utilize a dataset of user command streams collected from working with Microsoft (MS) Word to serve as a test bed. User profiles are first built using MS Word commands and identification takes place using machine learning algorithms. Best performance in terms of both accuracy and Area under the Curve (AUC) for Receiver Operating Characteristic (ROC) curve is reported using Random Forests (RF) and AdaBoost with random forests.
Ala'a El Masri, Harry Wechsler, Peter Likarish, Brent ByungHoon Kang
PST2
2013 Adversarial Spam Detection Using the Randomized Hough Transform-Support Vector Machine
abstract
In public e-mail systems, it is possible to solicit annotation help from users to train spam detection models. For example, we can occasionally ask a selected user to annotate whether a randomly selected message destined for their inbox is spam or not spam. Unfortunately, it is also possible that the user being solicited is an internal threat and has malicious intent. Similar to an adversary, such a user may want to introduce noise: to confuse the spam classifier into believing a spam message is not spam (to ensure delivery of similar messages), or to confuse the spam classifier into believing a non-spam message is spam (to prevent delivery of similar messages). Inspired by the Randomized Hough Transform (RHT), a set of Support Vector Machines (SVMs) is trained from randomly chosen data subsets to vote to identify training examples that have been mislabeled. The labels for messages which on the average appear on the wrong side of the decision boundary are flipped and a final SVM model is trained using the modified labels. Two data sets are used for evaluating the proposed RHT-SVM method: the TREC 2007 Spam Track data and the CEAS 2008 Spam data. To preserve the time ordered nature of the data stream, for each of the data sets, the first 10% of the messages are used for training, and the remaining 90% of the messages are used for evaluation. Separate adversarial experiments are conducted for flipping spam labels and non-spam labels. For 10 iterations, labels are flipped for a randomly selected subset of 5% of the training data and the final RHT-SVM is evaluated on the test set. Performance of the RHT-SVM is compared to the performance of the state of the art Reject On Negative Impact (RONI) algorithm. RHT-SVM shows an average 9.3% increase in the F measure compared to RONI (99.0% versus 90.6%), as well as significant improvements in other evaluation metrics. The flip sensitivity for RHT-SVM is 95.9% and the flip specificity is 99.0%. It also takes over 90% less time to complete the RHT-SVM experiments compared to the RONI experiments (20 minutes per experiment instead of 360 minutes).
David DeBarr, Harry Wechsler
ICMLA (1)3
2013 Real-Time Covert Timing Channel Detection in Networked Virtual Environments
Anyi Liu, Jim X. Chen, Harry Wechsler
IFIP Int. Conf. Digital Forensics3
2013 Phishing detection using traffic behavior, spectral clustering, and random forests
abstract
Phishing is an attempt to steal a user's identity. This is typically accomplished by sending an email message to a user, with a link directing the user to a web site used to collect personal information. Phishing detection systems typically rely on content filtering techniques, such as Latent Dirichlet Allocation (LDA), to identify phishing messages. In the case of spear phishing, however, this may be ineffective because messages from a trusted source may contain little content. In order to handle such emerging spear phishing behavior, we propose as a first step the use of Spectral Clustering to analyze messages based on traffic behavior. In particular, Spectral Clustering analyzes the links between URL substrings for web sites found in the message contents. Cluster membership is then used to construct a Random Forest classifier for phishing. Data from the Phishing Email Corpus and the Spam Assassin Email Corpus are used to evaluate this approach. Performance evaluation metrics include the Area Under the receiver operating characteristic Curve (AUC), as well as accuracy, precision, recall, and the (harmonic mean) F measure. Performance of the integrated Spectral Clustering and Random Forest approach is found to provide significant improvements in all the metrics listed, compared to a content filtering technique such as LDA coupled with text message deletion done randomly or in an adaptive fashion using adversarial learning. The Spectral Clustering approach is robust against the absence of content. In particular, we show that Spectral Clustering yields (99.8%, 97.8%) for (AUC, F measure) compared to LDA that yields (94.6%, 89.4%) and (79.6%, 57.9%) when the content of the messages is reduced to 10% of their original size using random and adversarial deletion, respectively. The difference is most striking at low False Positive (FP) rates.
David DeBarr, Venkatesh Ramanathan, Harry Wechsler
ISI3
2013 Phishing detection and impersonated entity discovery using Conditional Random Field and Latent Dirichlet Allocation
Venkatesh Ramanathan, Harry Wechsler
Comput. Secur.2
2013 Robust Face Recognition for Uncontrolled Pose and Illumination Changes
abstract
Face recognition has made significant advances in the last decade, but robust commercial applications are still lacking. Current authentication/identification applications are limited to controlled settings, e.g., limited pose and illumination changes, with the user usually aware of being screened and collaborating in the process. Among others, pose and illumination changes are limited. To address challenges from looser restrictions, this paper proposes a novel framework for real-world face recognition in uncontrolled settings named Face Analysis for Commercial Entities (FACE). Its robustness comes from normalization (“correction”) strategies to address pose and illumination variations. In addition, two separate image quality indices quantitatively assess pose and illumination changes for each biometric query, before submitting it to the classifier. Samples with poor quality are possibly discarded or undergo a manual classification or, when possible, trigger a new capture. After such filter, template similarity for matching purposes is measured using a localized version of the image correlation index. Finally, FACE adopts reliability indices, which estimate the “acceptability” of the final identification decision made by the classifier. Experimental results show that the accuracy of FACE (in terms of recognition rate) compares favorably, and in some cases by significant margins, against popular face recognition methods. In particular, FACE is compared against SVM, incremental SVM, principal component analysis, incremental LDA, ICA, and hierarchical multiscale local binary pattern. Testing exploits data from different data sets: CelebrityDB, Labeled Faces in the Wild, SCface, and FERET. The face images used present variations in pose, expression, illumination, image quality, and resolution. Our experiments show the benefits of using image quality and reliability indices to enhance overall accuracy, on one side, and to provide for individualized processing of biometric probes for better decision-making purposes, on the other side. Both kinds of indices, owing to the way they are defined, can be easily integrated within different frameworks and off-the-shelf biometric applications for the following: 1) data fusion; 2) online identity management; and 3) interoperability. The results obtained by FACE witness a significant increase in accuracy when compared with the results produced by the other algorithms considered.
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler
IEEE Trans. Syst. Man Cybern. Syst.4
2012 Phishing website detection using Latent Dirichlet Allocation and AdaBoost
abstract
One of the ways criminals steal identity in the cyberspace is using phishing. Attackers host phishing websites that resemble a legitimate website and entice users to click on hyperlinks which directs them to these fake websites. Attackers use these fake sites to capture personal information such as login, passwords and social security numbers from innocent victims, which they later use to commit crimes. We propose here a robust methodology to detect phishing websites that employs for semantic analysis a topic modeling technique, Latent Dirichlet Allocation, and for classification, AdaBoost. The methodology developed is a content driven approach that is device independent and language neutral. The website content of mobile and desktop clients are collected by employing an intelligent web crawler. The website contents that are not in English are translated to English using Google's language translator. Topic model is built using the translated contents of desktop and mobile clients. The phishing website classifier is built using (i) distribution probabilities for the topics found as features using Latent Dirichlet Allocation and (ii) AdaBoost voting technique. Experiments were conducted using one of the large public corpus of website data containing 47500 phishing websites and 52500 good websites. Results show that our method achieves a F-measure of 99%.
Venkatesh Ramanathan, Harry Wechsler
ISI2
2012 Musical keys and chords recognition using unsupervised learning with infinite Gaussian mixture
abstract
This paper presents a Bayesian-based methodology that determines keys and chords of symbolic music using unsupervised learning guided by constraints. An infinite Gaussian mixture, a type of Dirichlet Process mixture, is constructed to model the generative processes of keys and chords. Keys and chords are recognized iteratively by generated key and chord samples using the same algorithm. The model only uses very simple profiles for keys and chords to guide the learning process without the use of training data or rules. The technique is also capable of recognizing key modulations. We demonstrate the performance of the proposed approach by comparing it against existing methods using 159 songs from the Beatles MIDI collection. The results show that the proposed method performs well in both keys and chords finding.
Harry Wechsler
ICMR2
2012 phishGILLNET-phishing detection using probabilistic latent semantic analysis
abstract
Identity theft is one of the most profitable crimes committed by felons. In the cyber space, this is commonly achieved using phishing. We propose here robust server side methodology to detect phishing attacks, called phishGILLNET, which incorporates the power of natural language processing and machine learning techniques. phishGILLNET is a multi-layered approach to detect phishing attacks. The first layer (phishGILLNET1) employs Probabilistic Latent Semantic Analysis (PLSA) to build a topic model. The topic model handles synonym (multiple words with similar meaning), polysemy (words with multiple meanings), and other linguistic variations found in phishing. Intentional misspelled words found in phishing are handled using Levenshtein editing and Google APIs for correction. Based on term document frequency matrix as input PLSA finds phishing and non-phishing topics using tempered expectation maximization. The performance of phishGILLNET1 is evaluated using PLSA fold in technique and the classification is achieved using Fisher similarity. The second layer of phishGILLNET (phishGILLNET2) employs AdaBoost to build a robust classifier. Using probability distributions of the best PLSA topics as features the classifier is built using AdaBoost. The third layer (phishGILLNET3) further expands phishGILLNET2 by building a classifier from labeled and unlabeled examples by employing Co-Training. Experiments were conducted using one of the largest public corpus of email data containing 400,000 emails. Results show that phishGILLNET3 outperforms state of the art phishing detection methods and achieves F -measure of 100%. Moreover, phishGILLNET3 requires only a small percentage (10%) of data be annotated thus saving significant time, labor, and avoiding errors incurred in human annotation.
Venkatesh Ramanathan, Harry Wechsler
EURASIP J. Inf. Secur.2
2012 Spam detection using Random Boost
David DeBarr, Harry Wechsler
Pattern Recognit. Lett.2
2012 Novel pattern recognition-based methods for re-identification in biometric context
Mislav Grgic, Michele Nappi, Harry Wechsler
Pattern Recognit. Lett.3
2012 Robust re-identification using randomness and statistical learning: Quo vadis
Michele Nappi, Harry Wechsler
Pattern Recognit. Lett.2
2010 Adaptive biometric authentication using nonlinear mappings on quality measures and verification scores
abstract
Three methods to improve the performance of biometric matchers based on vectors of quality measures associated with biometric samples are described. The first two methods select samples and matching scores based on predicted values of Quality of Sample (QS) index (defined here as d-prime) and Confidence in matching Scores (CS), respectively. The third method treats quality measures as weak but useful features for discrimination between genuine and imposter matching scores. The unifying theme for the three methods consists of a nonlinear mapping between quality measures and the predicted values of QS, CS, and combined quality measures and matching scores, respectively. The proposed methodology is generic and is suitable for any biometric modality. The experimental results reported show significant performance improvements for all the three methods when applied to iris biometrics.
Jinyu Zuo, Francesco Nicolo, Natalia A. Schmid, Harry Wechsler
ICIP4
2010 A Martingale Framework for Detecting Changes in Data Streams by Testing Exchangeability
abstract
In a data streaming setting, data points are observed sequentially. The data generating model may change as the data are streaming. In this paper, we propose detecting this change in data streams by testing the exchangeability property of the observed data. Our martingale approach is an efficient, nonparametric, one-pass algorithm that is effective on the classification, cluster, and regression data generating models. Experimental results show the feasibility and effectiveness of the martingale methodology in detecting changes in the data generating model for time-varying data streams. Moreover, we also show that: 1) An adaptive support vector machine (SVM) utilizing the martingale methodology compares favorably against an adaptive SVM utilizing a sliding window, and 2) a multiple martingale video-shot change detector compares favorably against standard shot-change detection algorithms.
Shen-Shyang Ho, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Robust human authentication using appearance and holistic anthropometric features
Venkatesh Ramanathan, Harry Wechsler
Pattern Recognit. Lett.2
2009 Editorial
Mislav Grgic, Shiguang Shan, Rastislav Lukac, Harry Wechsler, Marian Stewart Bartlett
Int. J. Pattern Recognit. Artif. Intell.4
2009 Face Authentication Using Recognition-by-Parts, Boosting and Transduction
abstract
The paper describes an integrated recognition-by-parts architecture for reliable and robust face recognition. Reliability and robustness are characteristic of the ability to deploy full-fledged and operational biometric engines, and handling adverse image conditions that include among others uncooperative subjects, occlusion, and temporal variability, respectively. The architecture proposed is model-free and non-parametric. The conceptual framework draws support from discriminative methods using likelihood ratios. At the conceptual level it links forensics and biometrics, while at the implementation level it links the Bayesian framework and statistical learning theory (SLT). Layered categorization starts with face detection using implicit rather than explicit segmentation. It proceeds with face authentication that involves feature selection of local patch instances including dimensionality reduction, exemplar-based clustering of patches into parts, and data fusion for matching using boosting driven by parts that play the role of weak-learners. Face authentication shares the same implementation with face detection. The implementation, driven by transduction, employs proximity and typicality (ranking) realized using strangeness and p-values, respectively. The feasibility and reliability of the proposed architecture are illustrated using FRGC data. The paper concludes with suggestions for augmenting and enhancing the scope and utility of the proposed architecture.
Fayin Li, Harry Wechsler
Int. J. Pattern Recognit. Artif. Intell.2
2008 Gait Analysis using Independent Components of image motion
abstract
We propose a novel approach to gait analysis using the independent components of motion. Our motion representation uses the amount of translation in small image patches. For each subject in the training set, several short image sequences are selected at random. Spatiotemporal independent component analysis (stICA) estimates independent components (ICs) of the training sequences. Given an unknown subject, we compute stIC coefficients for all image subsequences. These coefficients are compared with the training set using the cosine similarity measure. We demonstrate feasibility by using the nearest neighbor to classify the gender of a subject.
Wallace E. Lawson, Zoran Duric, Harry Wechsler
FG3
2008 Robust fusion using boosting and transduction for component-based face recognition
abstract
Face recognition performance depends upon the input variability as encountered during biometric data capture including occlusion and disguise. The challenge met in this paper is to expand the scope and utility of biometrics by discarding unwarranted assumptions regarding the completeness and quality of the data captured. Towards that end we propose a model-free and non-parametric component-based face recognition strategy with robust decisions for data fusion that are driven by transduction and boosting. The conceptual framework draws support throughout from discriminative methods using likelihood ratios. It links at the conceptual level forensics and biometrics, while at the implementation level it links the Bayesian framework and statistical learning theory (SLT). Feature selection of local patch instances and their corresponding high-order combinations, exemplar-based clustering (of patches) as components including the sharing (of exemplars) among components, and finally decision-making regarding authentication using boosting driven by components that play the role of weak-learners, are implemented in a similar fashion using transduction driven by a strangeness measure akin to typicality. The feasibility, reliability, and utility of the proposed open set face recognition architecture vis-à-vis adverse image capture conditions are illustrated using FRGC data. The potential for future developments concludes the paper.
Fayin Li, Harry Wechsler, Massimo Tistarelli
ICARCV2
2008 Reliable face recognition using adaptive and robust correlation filters
Hung Lai, Venkatesh Ramanathan, Harry Wechsler
Comput. Vis. Image Underst.3
2008 Query by Transduction
abstract
There has been recently a growing interest in the use of transductive inference for learning. We expand here the scope of transductive inference to active learning in a stream-based setting. Towards that end this paper proposes Query-by-Transduction (QBT) as a novel active learning algorithm. QBT queries the label of an example based on the p-values obtained using transduction. We show that QBT is closely related to Query-by-Committee (QBC) using relations between transduction, Bayesian statistical testing, Kullback-Leibler divergence, and Shannon information. The feasibility and utility of QBT is shown on both binary and multi-class classification tasks using SVM as the choice classifier. Our experimental results show that QBT compares favorably, in terms of mean generalization, against random sampling, committee-based active learning, margin-based active learning, and QBC in the stream-based setting.
Shen-Shyang Ho, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Detecting Changes in Unlabeled Data Streams Using Martingale
Shen-Shyang Ho, Harry Wechsler
IJCAI2
2005 Adaptive Support Vector Machine for Time-Varying Data Streams Using Martingale
Shen-Shyang Ho, Harry Wechsler
IJCAI2
2005 On the Detection of Concept Changes in Time-Varying Data Stream by Testing Exchangeability
Shen-Shyang Ho, Harry Wechsler
UAI2
2005 Special issue: eye detection and tracking
Harry Wechsler, Andrew T. Duchowski, Myron Flickner
Comput. Vis. Image Underst.2
2005 Open Set Face Recognition Using Transduction
abstract
This paper motivates and describes a novel realization of transductive inference that can address the Open Set face recognition task. Open Set operates under the assumption that not all the test probes have mates in the gallery. It either detects the presence of some biometric signature within the gallery and finds its identity or rejects it, i.e., it provides for the "none of the above" answer. The main contribution of the paper is Open Set TCM-kNN (Transduction Confidence Machine-k Nearest Neighbors), which is suitable for multiclass authentication operational scenarios that have to include a rejection option for classes never enrolled in the gallery. Open Set TCM-kNN, driven by the relation between transduction and Kolmogorov complexity, provides a local estimation of the likelihood ratio needed for detection tasks. We provide extensive experimental data to show the feasibility, robustness, and comparative advantages of Open Set TCM-kNN on Open Set identification and watch list (surveillance) tasks using challenging FERET data. Last, we analyze the error structure driven by the fact that most of the errors in identification are due to a relatively small number of face patterns. Open Set TCM-kNN is shown to be suitable for PSEI (pattern specific error inhomogeneities) error analysis in order to identify difficult to recognize faces. PSEI analysis improves biometric performance by removing a small number of those difficult to recognize faces responsible for much of the original error in performance and/or by using data fusion.
Fayin Li, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Multiple regression estimation for motion analysis and segmentation
abstract
This paper describes multiple model estimation for motion analysis and segmentation (aka spatial partitioning), from point correspondences in two successive images. In motion analysis applications, available (training) data is generated by several unknown models (motions). However, the correspondence between data samples and different models (motions) is unknown. Hence, the goal of learning (motion estimation) is two-fold, i.e. estimation (learning) of unknown motions (models) and separation (segmentation) of available data into several subsets corresponding to different motions. We present the mathematical formulation for multiple motion estimation, as a problem of learning several (regression) mappings, from a single data set, and then show a constructive (SVM-based) learning algorithm developed for this setting. Experimental results show potential advantages of the proposed method.
Vladimir Cherkassky, Yunqian Ma, Harry Wechsler
IJCNN3
2004 Motion Estimation Using Statistical Learning Theory
abstract
This paper describes a novel application of Statistical Learning Theory (SLT) to single motion estimation and tracking. The problem of motion estimation can be related to statistical model selection, where the goal is to select one (correct) motion model from several possible motion models, given finite noisy samples. SLT, also known as Vapnik-Chervonenkis (VC), theory provides analytic generalization bounds for model selection, which have been used successfully for practical model selection. This paper describes a successful application of an SLT-based model selection approach to the challenging problem of estimating optimal motion models from small data sets of image measurements (flow). We present results of experiments on both synthetic and real image sequences for motion interpolation and extrapolation; these results demonstrate the feasibility and strength of our approach. Our experimental results show that for motion estimation applications, SLT-based model selection compares favorably against alternative model selection methods, such as the Akaike's fpe, Schwartz' criterion (sc), Generalized Cross-Validation (gcv), and Shibata's Model Selector (sms). The paper also shows how to address the aperture problem using SLT-based model selection for penalized linear (ridge regression) formulation.
Harry Wechsler, Zoran Duric, Fayin Li, Vladimir Cherkassky
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Partial Faces for Face Recognition: Left vs Right Half
Srinivas Gutta, Harry Wechsler
CAIP2
2003 Controlling Model Complexity in Flow Estimation
Zoran Duric, Fayin Li, Harry Wechsler, Vladimir Cherkassky
ICCV3
2003 A Distributed Reinforcement Learning Approach to Pattern Inference in Go
Myriam Abramson, Harry Wechsler
ICMLA2
2003 Competitive reinforcement learning in continuous control tasks
abstract
This paper describes a novel hybrid reinforcement learning algorithm, Sarsa Learning Vector Quantization (SLVQ), that leaves the reinforcement part intact but employs a more effective representation of the policy function using a piecewise constant function based upon "policy prototypes". The prototypes correspond to the pattern classes induced by the Voronoi tessellation generated by self-organizing methods like Learning Vector Quantization (LVQ). The determination of the optimal policy function can be now viewed as a pattern recognition problem in the sense that the assignment of an action to a point in the phase space is similar to the assignment of a pattern class to a point in phase space. The distributed LVQ representation of the policy function automatically generates a piecewise constant tessellation of the state space and yields in a major simplification of the learning task relative to the standard reinforcement learning algorithms for whom a discontinuous table look function has to be learned. The feasibility and comparative advantages of the new algorithm is shown on the cart centering and mountain car problems, two control problems of increased difficulty.
Myriam Abramson, Peter W. Pachowicz, Harry Wechsler
IJCNN3
2003 Tabu search exploration for on-policy reinforcement learning
abstract
On-policy reinforcement learning provides online adaptation, a characteristic of intelligent systems and lifelong learning. Unlike dynamic programming, an exhaustive sweep of the search space is not necessary for convergence in reinforcement learning with an efficient exploration strategy. For efficient and "believable" online performance, an exploration strategy also has to avoid cycling through previous solutions and know when to stop without getting stuck in a local optimum. This paper addresses the above problem with tabu search (TS) exploration. Several strategies for reinforcement learning are introduced. Experimental results are presented in the game of Go, a deterministic, perfect-information two-player game, and Sarsa learning vector quantization (SLVQ), an on-policy reinforcement learning algorithm.
Myriam Abramson, Harry Wechsler
IJCNN2
2003 Transductive confidence machine for active learning
abstract
This paper describes a novel active learning strategy using universal p-value measures of confidence based on algorithmic randomness, and transconductive inference. The early stopping criterion for active learning is based on the bias-variance tradeoff for classification. This corresponds to that learning instance when the boundary bias becomes positive, and requires one to switch from active to random selection of learning examples. The sign for the boundary and the increase in the classification error are two manifestations of the same phenomena, i.e., over-training. The experimental results presented show the feasibility and usefulness of our novel approach using a non-separable two-class classification problem. Our hybrid learning strategy achieves competitive performance against standard nearest neighbor methods using much fewer training examples.
Shen-Shyang Ho, Harry Wechsler
IJCNN2
2003 Independent component analysis of Gabor features for face recognition
abstract
We present an independent Gabor features (IGFs) method and its application to face recognition. The novelty of the IGF method comes from 1) the derivation of independent Gabor features in the feature extraction stage and 2) the development of an IGF features-based probabilistic reasoning model (PRM) classification method in the pattern recognition stage. In particular, the IGF method first derives a Gabor feature vector from a set of downsampled Gabor wavelet representations of face images, then reduces the dimensionality of the vector by means of principal component analysis, and finally defines the independent Gabor features based on the independent component analysis (ICA). The independence property of these Gabor features facilitates the application of the PRM method for classification. The rationale behind integrating the Gabor wavelets and the ICA is twofold. On the one hand, the Gabor transformed face images exhibit strong characteristics of spatial locality, scale, and orientation selectivity. These images can, thus, produce salient local features that are most suitable for face recognition. On the other hand, ICA would further reduce redundancy and represent independent features explicitly. These independent features are most useful for subsequent pattern discrimination and associative recall. Experiments on face recognition using the FacE REcognition Technology (FERET) and the ORL datasets, where the images vary in illumination, expression, pose, and scale, show the feasibility of the IGF method. In particular, the IGF method achieves 98.5% correct face recognition accuracy when using 180 features for the FERET dataset, and 100% accuracy for the ORL dataset using 88 features.
Chengjun Liu, Harry Wechsler
IEEE Trans. Neural Networks2
2002 Integrating perceptual and cognitive modeling for adaptive and intelligent human-computer interaction
abstract
This paper describes technology and tools for intelligent human-computer interaction (IHCI) in which human cognitive, perceptual, motor and affective factors are modeled and used to adapt the H-C interface. IHCI emphasizes that human behavior encompasses both apparent human behavior and the hidden mental state behind behavioral performance. IHCI expands on the interpretation of human activities, known as W4 (what, where, when, who). While W4 only addresses the apparent perceptual aspect of human behavior the W5+ technology for IHCI described in this paper addresses also the why and how questions, whose solution requires recognizing specific cognitive states. IHCI integrates parsing and interpretation of nonverbal information with a computational cognitive model of the user which, in turn, feeds into processes that adapt the interface to enhance operator performance and provide for rational decision-making. The technology proposed is based on a general four-stage interactive framework, which moves from parsing the raw sensory-motor input, to interpreting the user's motions and emotions, to building an understanding of the user's current cognitive state. It then diagnoses various problems in the situation and adapts the interface appropriately. The interactive component of the system improves processing at each stage. Examples of perceptual, behavioral, and cognitive tools are described throughout the paper Adaptive and intelligent HCI are important for novel applications of computing, including ubiquitous and human-centered computing.
Zoran Duric, Wayne D. Gray, Ric Heishman, Fayin Li, Azriel Rosenfeld, Michael J. Schoelles, Christian D. Schunn, Harry Wechsler
Proc. IEEE8
2002 Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition
abstract
This paper introduces a novel Gabor-Fisher (1936) classifier (GFC) for face recognition. The GFC method, which is robust to changes in illumination and facial expression, applies the enhanced Fisher linear discriminant model (EFM) to an augmented Gabor feature vector derived from the Gabor wavelet representation of face images. The novelty of this paper comes from 1) the derivation of an augmented Gabor feature vector, whose dimensionality is further reduced using the EFM by considering both data compression and recognition (generalization) performance; 2) the development of a Gabor-Fisher classifier for multi-class problems; and 3) extensive performance evaluation studies. In particular, we performed comparative studies of different similarity measures applied to various classifiers. We also performed comparative experimental studies of various face recognition schemes, including our novel GFC method, the Gabor wavelet method, the eigenfaces method, the Fisherfaces method, the EFM method, the combination of Gabor and the eigenfaces method, and the combination of Gabor and the Fisherfaces method. The feasibility of the new GFC method has been successfully tested on face recognition using 600 FERET frontal face images corresponding to 200 subjects, which were acquired under variable illumination and facial expressions. The novel GFC method achieves 100% accuracy on face recognition using only 62 features.
Chengjun Liu, Harry Wechsler
IEEE Trans. Image Process.2
2001 A Gabor Feature Classifier for Face Recognition
abstract
This paper describes a novel Gabor feature classifier (GFC) method for face recognition. The GFC method employs an enhanced Fisher discrimination model on an augmented Gabor feature vector, which is derived from the Gabor wavelet transformation of face images. The Gabor wavelets, whose kernels are similar to the 2D receptive field profiles of the mammalian cortical simple cells, exhibit desirable characteristics of spatial locality and orientation selectivity. As a result, the Gabor transformed face images produce salient local and discriminating features that are suitable for face recognition. The feasibility of the new GFC method has been successfully tested on face recognition using 600 FERET frontal face images, which involve different illumination and varied facial expressions of 200 subjects. The effectiveness of the novel GFC method is shown in terms of both absolute performance indices and comparative performance against some popular face recognition schemes such as the eigenfaces method and some other Gabor wavelet based classification methods. In particular, the novel GFC method achieves 100% recognition accuracy using only 62 features.
Chengjun Liu, Harry Wechsler
ICCV2
2001 A shape- and texture-based enhanced Fisher classifier for face recognition
abstract
This paper introduces a new face coding and recognition method, the enhanced Fisher classifier (EFC), which employs the enhanced Fisher linear discriminant model (EFM) on integrated shape and texture features. Shape encodes the feature geometry of a face while texture provides a normalized shape-free image. The dimensionalities of the shape and the texture spaces are first reduced using principal component analysis, constrained by the EFM for enhanced generalization. The corresponding reduced shape and texture features are then combined through a normalization procedure to form the integrated features that are processed by the EFM for face recognition. Experimental results, using 600 face images corresponding to 200 subjects of varying illumination and facial expressions, show that (1) the integrated shape and texture features carry the most discriminating information followed in order by textures, masked images, and shape images, and (2) the new coding and face recognition method, EFC, performs the best among the eigenfaces method using L(1) or L(2) distance measure, and the Mahalanobis distance classifiers using a common covariance matrix for all classes or a pooled within-class covariance matrix. In particular, EFC achieves 98.5% recognition accuracy using only 25 features.
Chengjun Liu, Harry Wechsler
IEEE Trans. Image Process.2
2000 Tracking Interacting People
abstract
A computer vision system for tracking multiple people in relatively unconstrained environments is described. Tracking is performed at three levels of abstraction: regions, people and groups. A novel, adaptive background subtraction method that combines colour and gradient information is used to cope with shadows and unreliable colour cues. People are tracked through mutual occlusions as they form groups and part from one another. Strong use is made of colour information to disambiguate occlusions and to provide qualitative estimates of depth ordering and position during occlusion. Some simple interactions with objects can also be detected. The system is tested using indoor and outdoor sequences. It is robust and should provide a useful mechanism for bootstrapping and reinitialisation of tracking using more-specific but less-robust human models.
Stephen J. McKenna, Sumer Jabri, Zoran Duric, Harry Wechsler
FG4
2000 Detection and Location of People in Video Images Using Adaptive Fusion of Color and Edge Information
abstract
A new method of finding people in video images is presented. The detection is based on a novel background modeling and subtraction approach which uses both color and edge information. We introduce confidence maps gray-scale images whose intensity is a function of confidence that a pixel has changed - to fuse intermediate results and represent the results of background subtraction. The latter is used to delineate a person's body by guiding contour collection to segment the person from the background. The method is tolerant to scene clutter, slow illumination changes, and camera noise, and runs in near real time on a standard platform.
Sumer Jabri, Zoran Duric, Harry Wechsler, Azriel Rosenfeld
ICPR3
2000 Learning the Face Space - Representation and Recognition
abstract
This paper advances an integrated learning and evolutionary computation methodology for approaching the task of learning the face space. The methodology is geared to provide a framework whereby enhanced and robust face coding and classification schemes can be derived and evaluated using both machine and human benchmark studies. In particular we take an interdisciplinary approach, drawing from the accumulated and vast knowledge of both the computer vision and psychology communities, and describe how evolutionary computation and statistical learning can engage in mutually beneficial relationships in order to define an exemplar (absolute)-based coding of multidimensional face space representation for successfully coping with changing population (face) types, and to leverage past experience for incremental face space definition.
Chengjun Liu, Harry Wechsler
ICPR2
2000 Automatic View Based Caricaturing
abstract
This paper describes a new method for the automatic view based generation of caricatures using image operators and lending itself to parallel implementation. This work is motivated by the need for a realistic and automatic caricaturing process for both human studies related to understanding why caricatures are recognized better than real images, and automatic face recognition. The proposed method determines how the valleys of an image should be distorted in order to obtain a caricature. The propagation of this sparse transformation to the full image provides a dense and smooth deformation field that, when it is properly applied, yields a realistic face caricature. The new caricaturing method does not require manual annotation and yields more realistic images due to the dense and smooth deformation field. A new shape and texture descriptor for representing faces based on the new caricatured image is also proposed for face recognition.
Albert Pujol, Juan José Villanueva, Harry Wechsler
ICPR3
2000 Network Ensembles for Facial Analysis Tasks
abstract
We propose an approach for combining the outputs of multiple neural network classifiers to reach a unified decision with improved performance in terms of higher recognition/classification rates. Our architecture consists of an ensemble of connectionist networks-radial basis functions (RBF)-and inductive decision trees (DT). The specific characteristics of our architecture include (a) query by consensus as provided by ensembles of networks for coping with the inherent variability of the image formation and data acquisition process, (b) categorical classifications using decision trees, (c) flexible and adaptive thresholds as opposed to ad hoc and hard thresholds. Experiments proving the feasibility of our architecture were performed on face recognition using 900 images, gender and ethnic classification tasks using 3000 images from the FERET facial database. Specifically, we observe that a small number of networks (two or three) were sufficient for yielding a much improved classification rate as opposed to using a single RBF network or an ensemble of RBF networks employing a voting scheme.
Srinivas Gutta, Harry Wechsler
IJCNN (3)2
2000 Tracking Groups of People
Stephen J. McKenna, Sumer Jabri, Zoran Duric, Azriel Rosenfeld, Harry Wechsler
Comput. Vis. Image Underst.5
2000 Evolutionary Pursuit and Its Application to Face Recognition
abstract
Introduces evolutionary pursuit (EP) as an adaptive representation method for image encoding and classification. In analogy to projection pursuit, EP seeks to learn an optimal basis for the dual purpose of data compression and pattern classification. It should increase the generalization ability of the learning machine as a result of seeking the trade-off between minimizing the empirical risk encountered during training and narrowing the confidence interval for reducing the guaranteed risk during testing. It therefore implements strategies characteristic of GA for searching the space of possible solutions to determine the optimal basis. It projects the original data into a lower dimensional whitened principal component analysis (PCA) space. Directed random rotations of the basis vectors in this space are searched by GA where evolution is driven by a fitness function defined by performance accuracy (empirical risk) and class separation (confidence interval). Accuracy indicates the extent to which learning has been successful, while separation gives an indication of expected fitness. The method has been tested on face recognition using a greedy search algorithm. To assess both accuracy and generalization capability, the data includes for each subject images acquired at different times or under different illumination conditions. EP has better recognition performance than PCA (eigenfaces) and better generalization abilities than the Fisher linear discriminant (Fisherfaces).
Chengjun Liu, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Color image compression using PCA and backpropagation learning
Clifford Clausen, Harry Wechsler
Pattern Recognit.2
2000 Visual routines for eye location using learning and evolution
abstract
Eye location is used as a test bed for developing navigation routines implemented as visual routines within the framework of adaptive behavior-based AI. The adaptive eye location approach seeks first where salient objects are, and then what their identity is. Specifically, eye location involves: 1) the derivation of the saliency attention map, and 2) the possible classification of salient locations as eve regions. The saliency ("where") map is derived using a consensus between navigation routines encoded as finite-state automata exploring the facial landscape and evolved using genetic algorithms (GAs). The classification ("what") stage is concerned with the optimal selection of features, and the derivation of decision trees, using GAs, to possibly classify salient locations as eyes. The experimental results, using facial image data, show the feasibility of our method, and suggest a novel approach for the adaptive development of task-driven active perception and navigational mechanisms.
Jeffrey Huang, Harry Wechsler
IEEE Trans. Evol. Comput.2
2000 Robust coding schemes for indexing and retrieval from large face databases
abstract
This paper introduces two new coding schemes, probabilistic reasoning models (PRM) and enhanced FLD (Fisher linear discriminant) models (EFM), for indexing and retrieval of large image databases with applications to face recognition. The unifying theme of the new schemes is that of lowering the space dimension ("data compression") subject to increased fitness for the discrimination index.
Chengjun Liu, Harry Wechsler
IEEE Trans. Image Process.2
2000 Quad-Q-learning
abstract
This paper develops the theory of quad-Q-learning which is a new learning algorithm that evolved from Q-learning. Quad-Q-learning is applicable to problems that can be solved by "divide and conquer" techniques. Quad-Q-learning concerns an autonomous agent that learns without supervision to act optimally to achieve specified goals. The learning agent acts in an environment that can be characterized by a state. Although Q-learning and quad-Q-learning are in many respects similar, they differ in their notion of state transitions. In the Q-learning environment, when an action is taken, a reward is received and a single new state results. The objective of Q-learning is to learn a policy function that maps states to actions so as to maximize a function of the rewards such as the sum of rewards. However, with respect to quad-Q-learning, when an action is taken from a state either an immediate reward and no new state results, or no reward is received and four new states result from taking that action. If four new states result, then each new state is treated independently as if four new environments resulted. If no new state results, no further action is taken in the associated environment. The environment in which quad-Q-learning operates can thus be viewed as a hierarchy of states where lower level states are the children of higher level states. The hierarchical aspect of quad-Q-learning leads to a bottom up view of learning that improves the efficiency of learning at higher levels in the hierarchy. The objective of quad-Q-learning is to maximize the sum of rewards obtained from each of the environments that result as actions are taken. Quad-Q-learning can be readily generalized to a family of n-Q-learning algorithms which are applicable to learning how to partition large intractable problem domains into smaller problems that can be solved independently. Two versions of quad-Q-learning are discussed; these are discrete state and mixed discrete and continuous state quad-Q-learning. The discrete state version is only applicable to problems with small numbers of states. The problem with the discrete case is that learning for one state does not improve learning with respect to other nearby states nor does it generalize to previously unseen states. Scaling up to problems with practical numbers of states requires a continuous state learning method. Continuous state learning can be accomplished using functional approximation methods. Application of quad-Q-learning to image compression is briefly described.
Clifford Clausen, Harry Wechsler
IEEE Trans. Neural Networks Learn. Syst.2
2000 Mixture of experts for classification of gender, ethnic origin, and pose of human faces
abstract
In this paper we describe the application of mixtures of experts on gender and ethnic classification of human faces, and pose classification, and show their feasibility on the FERET database of facial images. The FERET database allows us to demonstrate performance on hundreds or thousands of images. The mixture of experts is implemented using the "divide and conquer" modularity principle with respect to the granularity and/or the locality of information. The mixture of experts consists of ensembles of radial basis functions (RBFs). Inductive decision trees (DTs) and support vector machines (SVMs) implement the "gating network" components for deciding which of the experts should be used to determine the classification output and to restrict the support of the input space. Both the ensemble of RBF's (ERBF) and SVM use the RBF kernel ("expert") for gating the inputs. Our experimental results yield an average accuracy rate of 96% on gender classification and 92% on ethnic classification using the ERBF/DT approach from frontal face images, while the SVM yield 100% on pose classification.
Srinivas Gutta, Jeffrey Huang, P. Jonathon Phillips, Harry Wechsler
IEEE Trans. Neural Networks Learn. Syst.4
1999 Face Recognition Using Shape and Texture
abstract
We introduce in this paper a new face coding and recognition method which employs the Enhanced FLD (Fisher Linear Discrimimant) Model (EFM) on integrated shape (vector) and texture ('shape-free' image) information. Shape encodes the feature geometry of a face while texture provides a normalized shape-free image by warping the original face image to the mean shape, i.e., the average of aligned shapes. The dimensionalities of the shape and the texture spaces are first reduced using Principal Component Analysis (PCA). The corresponding but reduced shape find texture features are then integrated through a normalization procedure to form augmented features. The dimensionality reduction procedure, constrained by EFM for enhanced generalization, maintains a proper balance between the spectral energy needs of PCA for adequate representation, and the FLD discrimination requirements, that the eigenvalues of the within-class scatter matrix should not include small trailing values after the dimensionality reduction procedure as they appear in the denominator.
Chengjun Liu, Harry Wechsler
CVPR2
1999 Reinforcement Algorithms Using Functional Approximation for Generalization and their Application to Cart Centering and Fractal Compression
Clifford Claussen, Srinivas Gutta, Harry Wechsler
IJCAI3
1999 Gender and ethnic classification of human faces using hybrid classifiers
abstract
This paper considers hybrid classification architectures for gender and ethnic classification of human faces and shows their feasibility using a collection of 3006 face images corresponding to 1009 subjects from the FERET database. The hybrid approach consists of an ensemble of RBF networks and inductive decision trees (DT). Experimental cross validation (CV) results yield an average accuracy rate of (a) 96% on the gender classification task and (b) 94% on the ethnic classification task. The benefits of our hybrid architecture include: (i) robustness via query by consensus provided by the ensembles of RBF networks, and (ii) flexible and adaptive thresholds as opposed to ad hoc and hard thresholds provided by using only DT.
Srinivas Gutta, Harry Wechsler
IJCNN2
1999 An integrated shape and intensity coding scheme for face recognition
abstract
This paper introduces a new face coding scheme which employs an enhanced Fisher classifier (EFC) operating on integrated shape and intensity features. The dimensionalities of the shape and the intensity image spaces are first reduced using the principal component analysis, constrained by the EFC for enhanced generalization. The reduced shape and the intensity features are then integrated through a normalization procedure to form integrated features. Experiments using 600 face images from the FERET database of varying illumination and corresponding to 200 subjects, whose facial expression can vary, show the feasibility of the new face coding scheme. In particular, the EFC achieves 98.5% recognition rate using only 25 features. Our experiments also show that the integrated shape and intensity features carry the most discriminating information followed in order by textures, shape vectors, masked images and shape images.
Chengjun Liu, Harry Wechsler
IJCNN2
1999 Eye Detection Using Optimal Wavelet Packets and Radial Basis Functions (RBFs)
abstract
The eyes are important facial landmarks, both for image normalization due to their relatively constant interocular distance, and for post processing due to the anchoring on model-based schemes. This paper introduces a novel approach for the eye detection task using optimal wavelet packets for eye representation and Radial Basis Functions (RBFs) for subsequent classification ("labeling") of facial areas as eye versus non-eye regions. Entropy minimization is the driving force behind the derivation of optimal wavelet packets. It decreases the degree of data dispersion and it thus facilitates clustering ("prototyping") and capturing the most significant characteristics of the underlying (eye regions) data. Entropy minimization is thus functionally compatible with the first operational stage of the RBF classifier, that of clustering, and this explains the improved RBF performance on eye detection. Our experiments on the eye detection task prove the merit of this approach as they show that eye images compressed using optimal wavelet packets lead to improved and robust performance of the RBF classifier compared to the case where original raw images are used by the RBF classifier.
Jeffrey Huang, Harry Wechsler
Int. J. Pattern Recognit. Artif. Intell.2
1998 Probabilistic Reasoning Models for Face Recognition
abstract
We introduce in this paper two probabilistic reasoning models (PPM-1 and PRM-2) which combine the Principal Component Analysis (PCA) technique and the Bayes classifier and show their feasibility on the face recognition problem. The conditional probability density function for each class is modeled using the within class scatter and the Maximum A Posteriori (MAP) classification rule is implemented in the reduced PCA subspace. Experiments carried out using 1107 facial images corresponding to 369 subjects (with 169 subjects having duplicate images) from the FERET database show that the PRM approach compares favorably against the two well-known methods for face recognition-the Eigenfaces and Fisherfaces.
Chengjun Liu, Harry Wechsler
CVPR2
1998 Face Recognition Using Evolutionary Pursuit
Chengjun Liu, Harry Wechsler
ECCV (2)2
1998 Facial image retrieval using sequential classifiers
Srinivas Gutta, Harry Wechsler
ESANN2
1998 Gender and Ethnic Classification of Face Images
Srinivas Gutta, Harry Wechsler, P. Jonathon Phillips
FG2
1998 Evolution of Optimal Projection Axes (OPA) for Face Recognition
Chengjun Liu, Harry Wechsler
FG2
1998 Face Recognition Using Binary Image Metrics
Bálint Takács, Harry Wechsler
FG2
1998 Face Surveillance
abstract
Most of the research on face recognition addresses the MATCH problem and it assumes a closed universe where there is no need for a REJECT ('false positive') option. The SURVEILLANCE problem is addressed indirectly, if at all, through the MATCH problem, where the size of the gallery rather than that of the probe set is very large. This paper addresses the proper surveillance problem where the size of the probe ('unknown image') set vs. gallery ('known image') set is 450 vs. 50 frontal images. We developed robust face ID verification ('classification') and retrieval schemes based on hybrid classifiers and showed their feasibility using the FERET face data base. The hybrid classifier architecture consists of an ensemble of connectionist networks-Radial Basis Functions (RBF) and inductive decision trees (DT). Experimental results prove the feasibility of our approach and yield 97% accuracy using the probe and gallery sets specified above.
Srinivas Gutta, Jeffrey Huang, Vishal Kakkad, Harry Wechsler
ICCV4
1998 A Unified Bayesian Framework for Face Recognition
abstract
This paper introduces a Bayesian framework for face recognition which unifies popular methods such as the eigenfaces and Fisherfaces and can generate two novel probabilistic reasoning models (PRM) with enhanced performance. The Bayesian framework first applies principal component analysis (PCA) for dimensionality reduction with the resulting image representation enjoying noise reduction and enhanced generalization abilities for classification tasks. Following data compression, the Bayes classifier which yields the minimum error when the underlying probability density functions (PDF) are known, carries out the recognition in the reduced PCA subspace using the maximum a posteriori (MAP) rule, which is the optimal criterion for classification because it measures class separability. The PRM models are described within this unified Bayesian framework and shown to yield better performance against both the eigenfaces and Fisherfaces methods.
Chengjun Liu, Harry Wechsler
ICIP (1)2
1998 Face pose discrimination using support vector machines (SVM)
abstract
This paper describes an approach for the problem of face pose discrimination using support vector machines (SVM). Face pose discrimination means that one can label the face image as one of several known poses. Face images are drawn from the standard FERET database. The training set consists of 150 images equally distributed among frontal, approximately 33.75/spl deg/ rotated left and right poses, respectively, and the test set consists of 450 images again equally distributed among the three different types of poses. SVM achieved perfect accuracy-100%-discriminating between the three possible face poses on unseen test data, using either polynomials of degree 3 or radial basis functions (RBF) as kernel approximation functions.
Jeffrey Huang, Xuhui Shao, Harry Wechsler
ICPR3
1998 Enhanced Fisher linear discriminant models for face recognition
abstract
We introduce two enhanced Fisher linear discriminant (FLD) models (EFM) in order to improve the generalization ability of the standard FLD based classifiers such as Fisherfaces. Similar to Fisherfaces, both EFM models apply first principal component analysis (PCA) for dimensionality reduction before proceeding with FLD type of analysis. EFM-1 implements the dimensionality reduction with the goal to balance between the need that the selected eigenvalues account for most of the spectral energy of the raw data and the requirement that the eigenvalues of the within-class scatter matrix in the reduced PCA subspace are not too small. EFM-2 implements the dimensionality reduction as Fisherfaces do. It proceeds with the whitening of the within-class scatter matrix in the reduced PCA subspace and then chooses a small set of features (corresponding to the eigenvectors of the within-class scatter matrix) so that the smaller trailing eigenvalues are not included in further computation of the between-class scatter matrix. Experimental data using a large set of faces-1,107 images drawn from 369 subjects and including duplicates acquired at a later time under different illumination-from the FERET database shows that the EFM models outperform the standard FLD based methods.
Chengjun Liu, Harry Wechsler
ICPR2
1998 Optimal training set design for 3D object recognition
abstract
We describe a general approach for the representation and recognition of 3D objects. The method is based on a novel view selection mechanism that develops "visual filters" responsive to specific object classes to encode the complete viewing sphere with a small number of prototypical examples. The optimal set of visual filters is found via a cross-validation-like data reduction algorithm used to train banks of back propagation (BP) neural networks. Experimental results on real-world imagery demonstrate the feasibility of our approach.
Barnabás Takács, Lev Sadovnik, Harry Wechsler
ICPR3
1998 Fast searching of digital face libraries using binary image metrics
abstract
We describe a shape comparison method applicable to fast screening of large facial databases. The proposed technique derives holistic similarity measures without the explicit need of point-to-point correspondence thus delivering speed and tolerance to local non-rigid distortions. Specifically, we developed a face similarity measure derived as a variant of the Hausdorff distance by introducing the notion of a neighborhood function and associated penalties. Binary edge representation is used to provide robustness to changes in illumination. Experimental results on a large facial data set demonstrate that our approach produces excellent search results even when less than 1% of the original grey-scale face image information is stored in the face database.
Barnabás Takács, Harry Wechsler
ICPR2
1998 A Dynamic and Multiresolution Model of Visual Attention and Its Application to Facial Landmark Detection
Barnabás Takács, Harry Wechsler
Comput. Vis. Image Underst.2
1998 The FERET database and evaluation procedure for face-recognition algorithms
P. Jonathon Phillips, Harry Wechsler, Jeffrey Huang, Patrick J. Rauss
Image Vis. Comput.2
1998 A discrete dynamics model for synchronization of pulse-coupled oscillators
abstract
Biological information processing systems employ a variety of feature types. It has been postulated that oscillator synchronization is the mechanism for binding these features together to realize coherent perception. A discrete dynamic model of a coupled system of oscillators is presented. The network of oscillators converges to a state where subpopulations of cells become phase synchronized. It has potential applications to describing biological perception as well as for the construction of multifeature pattern recognition systems. It is shown that this model can be used to detect the presence of short line segments in the boundary contour of an object. The Hough transform, which is the standard method for detecting curve segments of a specified shape in an image was found not to be effective for this application. Implementation of the discrete dynamics model of oscillator synchronization is much easier than the differential equation models that have appeared in the literature. A systematic numerical investigation of the convergence properties of the model has been performed and it is shown that the discrete dynamics model can scale up to large number of oscillators.
Abraham Schultz, Harry Wechsler
IEEE Trans. Neural Networks2
1997 Hand Gesture Recognition Using Ensembles of Radial Basis Function (RBF) Networks and Decision Trees
abstract
Hand gestures are the natural form of communication among people, yet human-computer interaction is still limited to mice movements. The use of hand gestures in the field of human-computer interaction has attracted renewed interest in the past several years. Special glove-based devices have been developed to analyze finger and hand motion and use them to manipulate and explore virtual worlds. To further enrich the naturalness of the interaction, different computer vision-based techniques have been developed. At the same time the need for more efficient systems has resulted in new gesture recognition approaches. In this paper we present an hybrid intelligent system for hand gesture recognition. The hybrid approach consists of an ensemble of connectionist networks — radial basis functions (RBF) — and inductive decision trees (AQDT). Cross Validation (CV) experimental results yield a false negative rate of 1.7% and a false positive rate of 1% while the evaluation takes place on a data base including 150 images corresponding to 15 gestures of 5 subjects. In order to assess the robustness of the system, the vocabulary of the gestures has been increased from 15 to 25 and the size of the database from 150 to 750 images corresponding now to 15 subjects. Cross Validation (CV) experimental results yield a false negative rate of 3.6% and a false positive rate of 1.8% respectively. The benefits of our hybrid architecture include (i) robustness via query by consensus as provided by ensembles of networks when facing the inherent variability of the image formation and data acquisition process, (ii) classifications made using decision trees, (iii) flexible and adaptive thresholds as opposed to ad hoc and hard thresholds and (iv) interpretability of the way classification and retrieval is eventually achieved.
Srinivas Gutta, Ibrahim F. Imam, Harry Wechsler
Int. J. Pattern Recognit. Artif. Intell.3
1997 Face recognition using hybrid classifiers
Srinivas Gutta, Harry Wechsler
Pattern Recognit.2
1997 Detection of faces and facial landmarks using iconic filter banks
Barnabás Takács, Harry Wechsler
Pattern Recognit.2
1996 Face and Hand Gesture Recognition Using Hybrid Classifiers
abstract
This paper advances the methodology of hybrid classification architectures for face and hand gesture recognition tasks and shows their feasibility through experimental studies using the FERET data base and gesture images. The hybrid architecture, consisting of an ensemble of connectionist networks-radial basis functions (RBF)-and inductive decision trees (DT), combines the merits of 'holistic' template matching with those of 'abstractive' matching using discrete features and subject to both positive and negative learning. The hybrid architecture, quite general as it applies to both face and hand gesture recognition, derives its robustness from (i) consensus using ensembles of RBF network;, and (ii) flexible matching using categorical classification via decision trees. The experimental results, proving the feasibility of our approach, yield (i) 93% accuracy, using cross validation, for contents-based image retrieval (CBIR) subject to correct ID matching tasks, such as 'find Joe Smith with/without glasses', on a data bate of 200 images, and (ii) 96% accuracy using cross validation, for forensic verification on a data base consisting of 102 images corresponding to 350 subjects (of whom 102 are duplicates). Cross validation results on the hand gesture recognition task yield a false negative rate of 3.6% and a false positive rate of 1.8%, using a data base of 750 images corresponding to 25 hand gestures.
Srinivas Gutta, Jeffrey Huang, Ibrahim F. Imam, Harry Wechsler
FG4
1996 Detection of human faces using decision trees
abstract
The paper proposes a novel algorithm for face detection using decision trees (DT) and shows its generality and feasibility using a database consisting of 2340 face images from the FERET database (corresponding to 817 subjects and including 190 sets of duplicates) over a semi-uniform background. The approach used for face detection involves three main stages, those of location, cropping, and post-processing. The first stage finds a rough approximation for the possible location of the face box, the second stage will refine it, and the last stage decider whether a face is present in the image and if the answer is positive would normalize the face image. The algorithm does not require multiple (scale) templates and the accuracy achieved is 96%. Accuracy is based on the visual observation that the face box includes both eyes, nose, and mouth, and that the top side of the box is below the hairline. Experiments were also performed to assess the accuracy of the algorithm in rejecting images where no face is present. Using a small database of 25 images of various but complex backgrounds the algorithm failed on two images for an overall accuracy rate of 92%.
Jeffrey Huang, Srinivas Gutta, Harry Wechsler
FG3
1996 Visual Filters for Face Recognition
Barnabás Takács, Harry Wechsler
FG2
1996 Visual routine for eye detection using hybrid genetic architectures
abstract
We address the problem of crafting visual routines for detection tasks. Emphasis is placed on both competition and learning to help with specific visual tasks involved in localization and identification. Crafting of visual routines presents difficult optimization problems and leads to evolutionary computation using a hybrid genetic architecture consisting of natural selection, learning, and their beneficial interactions. Base features representations and visual routines for detection represented as decision trees are evolved. The visual routine considered is that of eye detection. The experimental results reported herein prove the feasibility of our approach in terms of feature selection (data compression) and the corresponding eye detection (pattern recognition).
Jerzy W. Bala, Kenneth A. De Jong, Jeffrey Huang, Haleh Vafaie, Harry Wechsler
ICPR5
1996 Face recognition using ensembles of networks
abstract
We describe a novel approach for fully automated face recognition and show its feasibility on a large database of facial images (FERET). Our approach, based on a hybrid architecture consisting of an ensemble of radial basis function (RBF) neural networks and inductive decision trees, combines the merits of "abstractive" features with those of "holistic" template matching. The benefits of our architecture include: 1) robust detection of facial landmarks using decision trees, and 2) robust face recognition using consensus methods over ensembles of RBF networks. Experiments carried out using k-fold cross validation on a large database consisting of 748 images corresponding to 374 subjects, among them 11 duplicates, yield on the average 87% correct match, and 99% correct surveillance ("verification").
Srinivas Gutta, Jeffrey Huang, Barnabás Takács, Harry Wechsler
ICPR4
1996 Automated borders detection and adaptive segmentation for binary document images
abstract
This paper describes two new and effective algorithms: one for detecting the page borders for documents available as binary images, and the other an adaptive segmentation algorithm using a bottom-up approach for segmenting binary images into blocks. The borders detection algorithm relies upon the classification of blank/textual/non-textual rows and columns, objects segmentation, and an analysis of projection profiles and crossing counts. Segmentation, done by an adaptive smearing technique, is different from all previous bottom-up approaches because any decisions on merging and/or separating are based on the estimated font information in binary document images.
Daniel X. Le, George R. Thoma, Harry Wechsler
ICPR3
1996 Attention and pattern detection using sensory and reactive control mechanisms
abstract
We introduce a biologically motivated low level model of visual attention and saccade generation based on data-driven dynamic processes governing foveation and recognition of object primitives. The approach consists of two major processing pathways, magno- (M) and parvocellular (P), and it employs: 1) retinal sampling, 2) active foveation, and 3) low-level ("coarse") recognition mechanisms. The M ("where") channel, responsible for object localization and corresponding reflexive saccades, feeds the P channel with salient locations for pattern detection. The P ("what") channel matches the image locations ("sensory") channel against previously interpreted and possibly labelled them. The P ("reactive") channel also generates the conditional saccades needed to collect additional information as it might be appropriate for full pattern interpretation. Simulation results, in the context of face recognition and using a large data set of 200 subjects, demonstrate the feasibility of our approach.
Barnabás Takács, Harry Wechsler
ICPR2
1996 Learning with Noise in Engineering Domains
Jerzy W. Bala, Peter W. Pachowicz, Harry Wechsler
ISMIS3
1996 Using Learning to Facilitate the Evolution of Features for Recognizing Visual Concepts
abstract
This paper describes a hybrid methodology that integrates genetic algorithms (GAs) and decision tree learning in order to evolve useful subsets of discriminatory features for recognizing complex visual concepts. A GA is used to search the space of all possible subsets of a large set of candidate discrimination features. Candidate feature subsets are evaluated by using C4.5, a decision tree learning algorithm, to produce a decision tree based on the given features using a limited amount of training data. The classification performance of the resulting decision tree on unseen testing data is used as the fitness of the underlying feature subset. Experimental results are presented to show how increasing the amount of learning significantly improves feature set evolution for difficult visual recognition problems involving satellite and facial image data. In addition, we also report on the extent to which other more subtle aspects of the Baldwin effect are exhibited by the system.
Jerzy W. Bala, Kenneth A. De Jong, Jeffrey Huang, Haleh Vafaie, Harry Wechsler
Evol. Comput.5
1996 Shape analysis using hybrid learning
Jerzy W. Bala, Harry Wechsler
Pattern Recognit.2
1996 Detection and localization of objects in time-varying imagery using attention, representation and memory pyramids
Vicente Concepcion, Harry Wechsler
Pattern Recognit.2
1995 Document image analysis using integrated image and neural processing
abstract
In this paper we present robust algorithms for detecting the page orientation (portrait/landscape) and the degree of skew for binary document images, and a method for classification of binary document images into textual or non-textual data blocks using neural network models. The performance of four neural network models are compared in terms of training times, memory requirements, and classification accuracy, and it was found that the radial basis functions performed best. The experiments show the feasibility of building an integrated document analysis system for page orientation and skew angle detection, and textual block classification.
Daniel X. Le, George R. Thoma, Harry Wechsler
ICDAR3
1995 Hybrid Learning Using Genetic Algorithms and Decision Trees for Pattern Classification
Jerzy W. Bala, Jeffrey Huang, Haleh Vafaie, Kenneth A. De Jong, Harry Wechsler
IJCAI (1)5
1995 Classification of binary document images into textual or nontextual data blocks using neural network models
Daniel X. Le, George R. Thoma, Harry Wechsler
Mach. Vis. Appl.3
1994 Integration of bottom-up and top-down cues for visual attention using non-linear relaxation
abstract
Active and selective perception seeks regions of interest in an image in order to reduce the computational complexity associated with time-consuming processes such as object recognition. We describe in this paper a visual attention system that extracts regions of interest by integrating multiple image cues. Bottom-up cues are detected by decomposing the image into a number: of feature and conspicuity maps, while a-priori knowledge (i.e. models) about objects is used to generate top-down attention cues. Bottom-up and top-down information is combined through a non-linear relaxation process using energy minimization-like procedures. The functionality of the attention system is expanded by the introduction of an alerting (motion-based) system able to explore and avoid obstacles. Experimental results are reported, using cluttered and noisy scenes.>
Ruggero Milanese, Harry Wechsler, Sylvia Gil, Jean-Marc Bost, Thierry Pun
CVPR2
1994 Detection of human speech in structured noise
abstract
This paper describes research to develop an efficient system that provides a binary decision as to the presence of speech in a short (one to three second) time sample of an acoustic signal. A method which is efficient and reliably detects human speech in the presence of structured noise (such as wind, music, traffic sounds, etc.) is described. Two separate algorithms were developed. The first algorithm detects the presence of speech by testing for concave and/or convex formant shapes. The second algorithm is a statistical pattern classifier utilizing radial basis function (RBF) networks with mel-cepstra feature vectors. Classification errors are not consistent across these two different methods. As a consequence, we plan to reduce our error rate by fusion of these methods.>
John D. Hoyt, Harry Wechsler
ICASSP (2)2
1994 Multiresolution attention and associative memory systems for time-varying imagery
abstract
This paper describes a vision system that integrates attention and recognition systems using active (where and what to sense) and selective (how to control the stream of computation) perception. Both time-varying imagery input and recognition memory are organized as pyramids and a uniform indexing and classification interface is established. The tools discussed in this paper include the use of: (1) scale-space representation using wavelet pyramids; (2) hierarchical, localized, and robust indexing and classification of the feature models using distributed associative memory; and (3) adaptive saccade and zoom strategies guided by saliency in determining focus of attention to locate and identify target objects.
Vicente Concepcion, Harry Wechsler
ICPR (1)2
1994 Detection of human speech using hybrid recognition models
abstract
This paper describes the research to develop an efficient system that provides a binary decision as to the presence of speech in a short time sample of an acoustic signal. A method which is efficient and reliably detects human speech in the presence of structured noise (such as wind, music, traffic sounds, etc.) is described. There are methods which work well to detect speech in a communications environment, but previous methods can not distinguish speech from a quasi-periodic signal that have a spectral power density similar to speech (such as music). Two separate feature sets are evaluated, reliable detection is obtained down to signal to noise ratios (SNR) as low as 0 dB. The algorithm utilized is a statistical pattern classifier with radial basis function networks. Mel-cepstra and wavelet feature vectors are compared. A method of obtaining the temporal feature information is also described.
John D. Hoyt, Harry Wechsler
ICPR (2)2
1994 Locating facial features using SOFM
abstract
We describe a novel and general approach for the detection of facial features such as the eyes. The approach is based on biologically motivated processing and classification schemes. The processing involves retinal sampling along P-type lattices and micro saccades, while classification is done using the self-organizing feature map (SOFM). The optimal set of eye templates is found by an enhanced SOFM approach using cross-validation training. Experimental results are presented to prove the feasibility of our approach.
Barnabás Takács, Harry Wechsler
ICPR (2)2
1994 Automated page orientation and skew angle detection for binary document images
Daniel X. Le, George R. Thoma, Harry Wechsler
Pattern Recognit.3
1993 Shape analysis using genetic algorithms
Jerzy W. Bala, Harry Wechsler
Pattern Recognit. Lett.2
1992 Active Perception Using DAM and Estimation Techniques
Wolfgang Pölzleitner, Harry Wechsler
ECCV2
1992 Robust shape analysis using multistrategy learning
abstract
This paper describes how to integrate subsymbolic and symbolic processes in order to create high-performance shape analysis systems. The specific methodology introduced integrates morphological processing and machine learning techniques such as genetic algorithms (GAs) and empirical inductive generalization. The optimal operators (defined as variable morphological structuring elements) evolved by GAs are used to derive discriminant feature vectors, which are then used by empirical inductive learning to generate rule-based class description in disjunctive normal form. The rule-based descriptions are finally optimized by removing small disjuncts in order to enhance the robustness of the shape analysis system. Experimental results are presented to illustrate the feasibility of the methodology for discriminating among classes of arbitrarily shaped objects, for learning the concepts of convexity and concavity, and for building robust recognition methods.>
Jerzy W. Bala, Harry Wechsler
ICPR (2)2
1992 The wavelet transform-a CMOS VLSI ASIC implementation
abstract
Describes a complementary metal oxide semiconductor (CMOS) very large scale integration (VLSI) application specific integrated circuit (ASIC) implementation of the wavelet transform. This custom hardware can provide higher speed, in the order of 5-50 million data samples/sec (needed for video applications) and/or lower power (needed for battery powered applications) than software implementations of either the wavelet or windowed Fourier transform on a conventional Von Neumann programmable computer.>
John D. Hoyt, Harry Wechsler
ICPR (4)2
1992 Selective and robust perception using multiresolution estimation techniques
abstract
Data and model-driven processes are basic to image understanding systems and correspond preattentive and focal attentive processes, respectively. The authors describe how to develop preattentive and focal attentive processes primed by known object models complementary to efforts on data driven preattentive processes. The techniques used to implement attentional mechanisms are shown to be robust with respect to noise given as changes in illumination, able to reduce crosstalk and able to count outliers.>
Wolfgang Pölzleitner, Harry Wechsler
ICPR (2)2
1992 An Overview of Parallel Hardware Architectures for Computer Vision
abstract
The sheer complexity of the visual task and the need for robust behavior provide a unifying theme for computer vision — what can be accomplished is constrained by relatively limited computational resources, but the resulting performance must be robust. As a consequence parallel computation over space and time becomes essential for machine vision systems. Parallel computation is all encompassing and includes both image representations and processing. We give an overview in this paper of a taxonomy of parallel hardware architectures which includes pipelining, Single instruction multiple data (SIMD), multiple instruction multiple data (MIMD), data-flow, and neurocomputing. Specific applications and the relevance of parallel hardware architectures for computer vision tasks are discussed as well.
Harry Wechsler
Int. J. Pattern Recognit. Artif. Intell.1
1991 Shape analysis using morphological processing and genetic algorithms
abstract
A novel way of combining morphological processing and genetic algorithms (GAs) to generate high-performance shape discrimination operators is presented. GAs can evolve operators that discriminate among classes comprising different shapes. The operators are defined as variable structuring elements and can be sequenced as program forms. The population of such operators, evaluated according to an index of performance corresponding to shape discrimination ability, evolves into an optimal set of operators using genetic search. Experimental results are presented to illustrate the feasibility of the approach for shape discrimination.>
Jerzy W. Bala, Harry Wechsler
ICTAI2
1991 Fault-tolerant database using distributed associative memories
Vladimir Cherkassky, Malathi Rao, Harry Wechsler
Inf. Sci.3
1991 Spatial/spatial-frequency representations for image segmentation and grouping
Todd R. Reed, Harry Wechsler
Image Vis. Comput.2
1990 An examination of the application of multi-layer neural networks to audio signal processing
abstract
A method is described for utilizing a multilayer neural network for audio signal processing. Normal finite and infinite impulse response (FIR and IIR) filters are of little use in separating the signal from noise because the cases considered are those where the noise source has statistical properties similar to the signal. Artificial neural systems (ANSs) offer potential alternatives to signal-processing and pattern-recognition techniques. ANSs were utilized using back propagation in an adaptive filter, and how single-layer and multilayer networks compare for adaptive filtering was explored. The experimental results indicate that the multilayer ANS did provide better performance
John D. Hoyt, Harry Wechsler
IJCNN2
1990 Simultaneous fitting of several planes to point sets using neural networks
Behrooz Kamgar-Parsi, Behzad Kamgar-Parsi, Harry Wechsler
Comput. Vis. Graph. Image Process.3
1990 Selective and Focused Invariant Recognition Using Distributed Associative Memories (DAM)
abstract
A method of 2-D object recognition based on the Moore-Penrose distributed associative memory is presented. Using known relationships between DAMs and regression analysis, the selectivity of the association weights is improved in an iterative way be discarding from further consideration response vectors deemed to be insignificant. Such selectivity allow the system to focus on the significant associations and to reduce crosstalk effects. The same formalism that allows the significance of the association weights to be computed also provides for a reject option. Experiments incorporating the proposed method onto an invariant recognition system prove the feasibility and benefits of the recognition scheme.>
Wolfgang Pölzleitner, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Segmentation of Textured Images and Gestalt Organization Using Spatial/Spatial-Frequency Representations
abstract
The generic issue of clustering/grouping is addressed. Recent research, both in computer and human vision, suggests the use of joint spatial/spatial-frequency (s/sf) representations. The spectrogram, the difference of Gaussians representation, the Gabor representation, and the Wigner distribution are discussed and compared. It is noted that the Wigner distribution gives superior joint resolution. Experimental results in the area of texture segmentation and Gestalt grouping using the Wigner distribution are presented, proving the feasibility of using s/sf representations for low-level (early, preattentive) vision.>
Todd R. Reed, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Texture segmentation using a diffusion region growing technique
Todd R. Reed, Harry Wechsler, Michael Werman
Pattern Recognit.2
1989 Distributed Associative Memory (DAM) for Bin-Picking
abstract
The feasibility of using a distributed associative memory as the recognition component for a bin-picking system is established. The system displays invariance to metric distortions and a robust response in the presence of noise, occlusions, and faults. Although the system is primarily concerned with two-dimensional problems, eight extensions to the system allow the three-dimensional bin-picking problem to be addressed. It is noted that there are implicit weaknesses in the neural network model chosen for the heart of the recognition system. The distributed associative memory used is linear, and as a result there are certain desirable properties that cannot be exhibited by the computer vision system.>
Harry Wechsler, George Lee Zimmerman
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Edge detection by associative mapping
Peter Meer, Harry Wechsler
Pattern Recognit.3
1988 2-D Invariant Object Recognition Using Distributed Associative Memory
abstract
Complex-log conformal mapping is combined with a distributed associative memory to create a system that recognizes objects regardless of changes in rotation or scale. Information recalled from the memorized database is used to classify an object, reconstruct the memorized version of the object, and estimate the magnitude of changes in scale or rotation. The system response is resistant to moderate amounts of noise and occlusion. Several experiments using real gray-scale images are presented to show the feasibility of the approach.>
Harry Wechsler, George Lee Zimmerman
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Distributed and Fault-Tolerant Computation for Retrieval Tasks Using Distributed Associative Memories
abstract
The distributed associative memory (DAM) model is proposed for distributed and fault-tolerant computation related to retrieval tasks. The fault tolerance is with respect to noise in the input key data and/or local and global failures in the memory itself. Working models for fault-tolerant image reconfiguration and database information retrieval have been developed and backed up by experimental results that show the feasibility of such an approach.>
Jois Malathi Char, Vladimir Cherkassky, Harry Wechsler, George Lee Zimmerman
IEEE Trans. Computers3
1987 Invariant Object Recognition Using a Distributed Associative Memory
Harry Wechsler, George Lee Zimmerman
NIPS1
1987 Derivation of optical flow using a spatiotemporal-Frequency approach
Lowell Jacobson, Harry Wechsler
Comput. Vis. Graph. Image Process.2
1984 An image reconstruction algorithm for time varying X-ray projection image sequences
abstract
A new method to improve the spatial resolution in reconstructed images of moving objects using high-speed CT is described. Improvement is obtained by providing the reconstruction algorithm with increased number of views. This is done by mathematically estimating projections around the moving object corresponding to a given instant of time with the aid of Kalman-type filters. The variations of the projection data as a function of time is modeled using a discrete time state space model. A discrete transformation is used to map the signal dependent nature of the noise into a signal independent one. Results of several computer simulations using the algorithm are presented.
R. S. Acharya, Richard A. Robb, Harry Wechsler
ICASSP3
1984 Invariant image representation: A path toward solving the bin-picking problem
abstract
We review the definition of the bin-picking problem and briefly summarize past approaches toward its solution, arguing that a major problem with conventional approaches is their reliance upon impoverished image representations. We then describe a recently developed image representation that uniquely encodes the information in a grey-scale image, decouples the effects of illumination, reflectance, and angle of incidence, and is invariant, within a linear shift, to perspective, position, orientation, and size of all planar forms. This representation is composed of discrete approximations to the complex-logarithmic conformally mapped Wigner distributions of multiple deprojection images. Such a representation could open the frontiers to the development of general-purpose robot vision systems which routinely perform size, orientation, and perspective invariant part recognition while also specifying part location and attitude in 3-D space.
Lowell Jacobson, Harry Wechsler
ICRA2
1984 A Theory for Invariant Object Recognition in the Frontoparallel Plane
abstract
We suggest a new computational theory for position, size, and orientation invariant recognition of planar gray-scale forms in frontoparallel view. The theory employs a spatial-frequency representation which approximates the complex-logarithmic conformally mapped Wigner distribution of a gray-scale image. Our discussion emphasizes the theoretical significance, method of computation, and practical utility of this powerful new representation for image pattern recognition.
Lowell Jacobson, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
1984 Invariant analogical image representation and pattern recognition
Lowell Jacobson, Harry Wechsler
Pattern Recognit. Lett.2
1983 The composite pseudo Wigner distribution (CPWD): A computable and versatile approximation to the Wigner distribution (WD)
abstract
The Wigner distribution (WD) has recently received much attention from researchers interested in dual time/frequency or space/spatial-frequency signal representation. However, an adequate computable approximation to the WD has been lacking. In order to fill this void, we introduce in this paper a new and highly versatile approximation known as the composite pseudo Wigner distribution (CPWD).
Lowell Jacobson, Harry Wechsler
ICASSP2
1982 On the Difficulties Involved in the Segmentation of Pictures
abstract
The problem of picture segmentation is widely believed to be practically unsolvable in its more general case formulation. However, as far as the authors know no formal argumentation has yet been given to this belief related to the segmentation problem. This correspondence provides some theoretical aspects of this phenomenon.
Eitan M. Gurari, Harry Wechsler
IEEE Trans. Pattern Anal. Mach. Intell.2
1982 A paradigm for invariant object recognition of brightness, optical flow and binocular disparity images
Lowell Jacobson, Harry Wechsler
Pattern Recognit. Lett.2
1980 Feature extraction for texture classification
Harry Wechsler, Todd K. Citron
Pattern Recognit.1
1979 A Random Walk Procedure for Texture Discrimination
abstract
We consider the problem of texture discrimination. Random walks are performed in a plain domain D bounded by an absorbing boundary ¿ and the absorption distribution is calculated. Measurements derived from such distributions are the features used for discrimination. Both problems of texture discrimination and edge segment detection can be solved using the same random walk approach. The border distributions and their differences with respect to a homogeneous image can classify two different images as having similar or dissimilar textures. The existence of an edge segment is concluded if the boundary distribution for a given window (subimage) differs significantly from the boundary distribution for a homogeneous (uniform grey level) window. The random walk procedure has been implemented and results of texture discrimination are shown. A comparison is made between results obtained using the random walk approach and the first-or second-order statistics, respectively. The random walk procedure is intended mainly for the texture discrimination problem, and its possible application to the edge detection problem (as shown in this paper) is just a by-product.
Harry Wechsler, Masatsugu Kidode
IEEE Trans. Pattern Anal. Mach. Intell.1
1977 Finding the rib cage in chest radiographs
Harry Wechsler, Jack Sklansky
Pattern Recognit.1
1977 A New Edge Detection Technique and Its Implementation
abstract
A new edge detection technique for digital image processing is presented. The technique is viewed as a low-level operator in the context of the more complex problem, that of scene analysis, and a conceptual comparison with some previous edge detectors is done. The new edge detection technique has been implemented and its results are compared with two other edge detection techniques using the same kind of pictures as input data. The software implementation of this new edge detector written in machine language takes 32 s for an image of size 128 by 128 picture elements. The edge detector can be hardware-implemented and such an implementation, for which the estimated processing time will be about half a second, is given in the Appendix.
Harry Wechsler, Masatsugu Kidode
IEEE Trans. Syst. Man Cybern.1
1975 Automatic Detection Of Rib Contours in Chest Radiographs
Harry Wechsler, Jack Sklansky
IJCAI1