Shumeet Baluja

dblp:78/4447 · DBLP profile ↗
← Back
66ranked-venue papers
42as first author
5since 2021 · last 2023
0000-0002-8696-8711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 30 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 14 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2023 The infinite doodler: expanding textures within tightly constrained manifolds
Shumeet Baluja
Vis. Comput.1
2022 Adding Non-Linear Context to Deep Networks
abstract
Enormous success has been achieved with deep neural networks consisting of standard linear-convolutions followed by simple non-linear mapping functions. In this paper, we add easily-computed non-linear local and global statistics to deep architectures, augmenting the information available at each layer. This additional information is then used in an identical manner to current processing. The summary statistics, which can be as simple as calculating within-channel variance, introduces little run-time computational overhead and can be instantiated with few extra parameters. All standard training procedures can be used without modification for training these augmented networks. We show, through extensive testing with ResNet on ImageNet, performance improvements across a wide range of network sizes. Additionally, we provide a detailed study of where within the deep networks these statistics are most effective.
Michele Covell, David Marwood, Shumeet Baluja
ICIP3
2022 A natural representation of colors with textures
Shumeet Baluja
Vis. Comput.1
2021 Interpretable Actions: Controlling Experts with Understandable Commands
Shumeet Baluja, David Marwood, Michele Covell
AAAI1
2021 Contextual Convolution Blocks
David Marwood, Shumeet Baluja
BMVC2
2020 Hiding Images within Images
abstract
We present a system to hide a full color image inside another of the same size with minimal quality loss to either image. Deep neural networks are simultaneously trained to create the hiding and revealing processes and are designed to specifically work as a pair. The system is trained on images drawn randomly from the ImageNet database, and works well on natural images from a wide variety of sources. Beyond demonstrating the successful application of deep learning to hiding images, we examine how the result is achieved and apply numerous transformations to analyze if image quality in the host and hidden image can be maintained. These transformation range from simple image manipulations to sophisticated machine learning-based adversaries. Two extensions to the basic system are presented that mitigate the possibility of discovering the content of the hidden image. With these extensions, not only can the hidden information be kept secure, but the system can be used to hide even more than a single image. Applications for this technology include image authentication, digital watermarks, finding exact regions of image manipulation, and storing meta-information about image rendering and content.
Shumeet Baluja
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Learning to Render Better Image Previews
abstract
A rapidly increasing portion of Internet traffic is dominated by requests from mobile devices with limited and metered bandwidth constraints. To satisfy these requests, it has become standard practice for websites to transmit small and extremely compressed image previews as part of the initial page-load process. Recent work, based on an adaptive triangulation of the target image, has performed well at extreme compression rates: 200 bytes or less. Gains have been shown, in terms of PSNR and SSIM, over both JPEG and WebP standards. However, qualitative assessments and preservation of semantic content can be less favorable. We present a novel method to significantly improve the reconstruction quality of the original image that requires no changes to the encoded information. Our neural-based decoding triples the amount of semantic-level content preservation while also improving both SSIM and PSNR scores. In addition, by keeping the same encoding stream, our solution is completely inter-operable with the original, and remains suitable for small-device deployment.
Shumeet Baluja, David Marwood, Nicholas Johnston, Michele Covell
ICIP1
2018 Learning to Attack: Adversarial Transformation Networks
abstract
With the rapidly increasing popularity of deep neural networks for image recognition tasks, a parallel interest in generating adversarial examples to attack the trained models has arisen. To date, these approaches have involved either directly computing gradients with respect to the image pixels or directly solving an optimization on the image pixels. We generalize this pursuit in a novel direction: can a separate network be trained to efficiently attack another fully trained network? We demonstrate that it is possible, and that the generated attacks yield startling insights into the weaknesses of the target network. We call such a network an Adversarial Transformation Network (ATN). ATNs transform any input into an adversarial attack on the target network, while being minimally perturbing to the original inputs and the target network's outputs. Further, we show that ATNs are capable of not only causing the target network to make an error, but can be constructed to explicitly control the type of misclassification made. We demonstrate ATNs on both simple MNIST-digit classifiers and state-of-the-art ImageNet classifiers deployed by Google, Inc.: Inception ResNet-v2.
Shumeet Baluja, Ian Fischer
AAAI1
2018 Representing Images in 200 Bytes: Compression via Triangulation
abstract
A rapidly increasing portion of internet traffic is dominated by requests from mobile devices with limited and metered bandwidth constraints. To satisfy these requests, it has become standard practice for websites to transmit small and extremely compressed image previews as part of the initial page load process to improve responsiveness. Increasing thumbnail compression beyond the capabilities of existing codecs is therefore an active research direction. In this work, we concentrate on extreme compression rates, where the size of the image is typically 200 bytes or less. First, we propose a novel approach for image compression that, unlike commonly used methods, does not rely on block-based statistics. We use an approach based on an adaptive triangulation of the target image, devoting more triangles to high entropy regions of the image. Second, we present a novel algorithm for encoding the triangles. The results show favorable statistics, in terms of PSNR and SSIM, over both the JPEG and the WebP standards.
David Marwood, Pascal Massimino, Michele Covell, Shumeet Baluja
ICIP4
2017 Hiding Images in Plain Sight: Deep Steganography
abstract
Steganography is the practice of concealing a secret message within another, ordinary, message. Commonly, steganography is used to unobtrusively hide a small message within the noisy regions of a larger image. In this study, we attempt to place a full size color image within another image of the same size. Deep neural networks are simultaneously trained to create the hiding and revealing processes and are designed to specifically work as a pair. The system is trained on images drawn randomly from the ImageNet database, and works well on natural images from a wide variety of sources. Beyond demonstrating the successful application of deep learning to hiding images, we carefully examine how the result is achieved and explore extensions. Unlike many popular steganographic methods that encode the secret message within the least significant bits of the carrier image, our approach compresses and distributes the secret image's representation across all of the available bits.
Shumeet Baluja
NIPS1
2017 Learning typographic style: from discrimination to synthesis
Shumeet Baluja
Mach. Vis. Appl.1
2016 Labeling the Features Not the Samples: Efficient Video Classification with Minimal Supervision
abstract
Feature selection is essential for effective visual recognition. We propose an efficient joint classifier learning and feature selection method that discovers sparse, compact representations of input features from a vast sea of candidates, with an almost unsupervised formulation. Our method requires only the following knowledge, which we call the feature sign - whether or not a particular feature has on average stronger values over positive samples than over negatives. We show how this can be estimated using as few as a single labeled training sample per class. Then, using these feature signs, we extend an initial supervised learning problem into an (almost) unsupervised clustering formulation that can incorporate new data without requiring ground truth labels. Our method works both as a feature selection mechanism and as a fully competitive classifier. It has important properties, low computational cost annd excellent accuracy, especially in difficult cases of very limited training data. We experiment on large-scale recognition in video and show superior speed and performance to established feature selection approaches such as AdaBoost, Lasso, greedy forward-backward selection, and powerful classifiers such as SVM.
Marius Leordeanu, Alexandra Radu, Shumeet Baluja, Rahul Sukthankar
AAAI3
2015 The Virtues of Peer Pressure: A Simple Method for Discovering High-Value Mistakes
Shumeet Baluja, Michele Covell, Rahul Sukthankar
CAIP (2)1
2013 Point representation for local optimization
abstract
In the context of stochastic search, once regions of high performance are found, having the property that small changes in the candidate solution correspond to searching nearby neighborhoods provides the ability to perform effective local optimization. To achieve this, Gray Codes are often employed for encoding ordinal points or discretized real numbers. In this paper, we present a method to label similar and/or close points within arbitrary graphs with small Hamming distances. The resultant point labels can be viewed as an approximate high-dimensional variant of Gray Codes. The labeling procedure is useful for any task in which the solution requires the search algorithm to select a small subset of items out of many. A large number of empirical results using these encodings with a combination of genetic algorithms and hill-climbing are presented.
Shumeet Baluja, Michele Covell
IEEE Congress on Evolutionary Computation1
2010 Beyond "Near Duplicates": Learning Hash Codes for Efficient Similar-Image Retrieval
abstract
Finding similar images in a large database is an important, but often computationally expensive, task. In this paper, we present a two-tier similar-image retrieval system with the efficiency characteristics found in simpler systems designed to recognize near-duplicates. We compare the efficiency of lookups based on random projections and learned hashes to 100-times-more-frequent exemplar sampling. Both approaches significantly improve on the results from exemplar sampling, despite having significantly lower computational costs. Learned-hash keys provide the best result, in terms of both recall and efficiency.
Shumeet Baluja, Michele Covell
ICPR1
2009 LSH banding for large-scale retrieval with memory and recall constraints
abstract
Locality Sensitive Hashing (LSH) is widely used for efficient retrieval of candidate matches in very large audio, video, and image systems. However, extremely large reference databases necessitate a guaranteed limit on the memory used by the table lookup itself, no matter how the entries crowd different parts of the signature space, a guarantee that LSH does not give. In this paper, we provide such guaranteed limits, primarily through the design of the LSH bands. When combined with data-adaptive bin splitting (needed on only 0.04% of the occupied bins) this approach provides the required guarantee on memory usage. At the same time, it avoids the reduced recall that more extensive use of bin splitting would give.
Michele Covell, Shumeet Baluja
ICASSP2
2009 Finding Images and Line-Drawings in Document-Scanning Systems
abstract
The system presented in this paper finds images and line-drawings in scanned pages; it is a crucial processing step in the creation of a large-scale system to detect and index images found in books and historic documents. Within the scanned pages that contain both text and images, the images are found through the use of SIFT-based local-features applied to the complete scanned-page. This is followed by a novel learning system to categorize the found SIFT features into either text or image. The discrimination is based on using multiple classifiers trained via AdaBoost. Through the use of this system, we improve image detection by finding more line-drawings, graphics, and photographs, as well as by reducing the number of spurious detections due to misclassified text, discolorations, and scanning artifacts.
Shumeet Baluja, Michele Covell
ICDAR1
2009 What's up CAPTCHA?: a CAPTCHA based on image orientation
abstract
We present a new CAPTCHA which is based on identifying an image's upright orientation. This task requires analysis of the often complex contents of an image, a task which humans usually perform well and machines generally do not. Given a large repository of images, such as those from a web search result, we use a suite of automated orientation detectors to prune those images that can be automatically set upright easily. We then apply a social feedback mechanism to verify that the remaining images have a human-recognizable upright orientation. The main advantages of our CAPTCHA technique over the traditional text recognition techniques are that it is language-independent, does not require text-entry (e.g. for a mobile device), and employs another domain for CAPTCHA generation beyond character obfuscation. This CAPTCHA lends itself to rapid implementation and has an almost limitless supply of images. We conducted extensive experiments to measure the viability of this technique.
Rich Gossweiler, Maryam Kamvar, Shumeet Baluja
WWW3
2008 Query suggestions for mobile search: understanding usage patterns
abstract
Entering search terms on mobile phones is a time consuming and cumbersome task. In this paper, we explore the usage patterns of query entry interfaces that display suggestions. Our primary goal is to build a usage model of query suggestions in order to provide user interface guidelines for mobile text prediction interfaces. We find that users who were asked to enter queries on a search interface with query suggestions rated their workload lower and their enjoyment higher. They also saved, on average, approximately half of the key presses compared to users who were not shown suggestions, despite no associated decrease in time to enter a query. Surprisingly, users also accepted suggestions when the process of doing so resulted in an increase in the number of total key presses.
Maryam Kamvar, Shumeet Baluja
CHI2
2008 Permutation grouping: intelligent Hash function design for audio & image retrieval
abstract
The combination of MinHash-based signatures and locality- sensitive hashing (LSH) schemes has been effectively used for finding approximate matches in very large audio and image retrieval systems. In this study, we introduce the idea of permutation-grouping to intelligently design the hash functions that are used to index the LSH tables. This helps to overcome the inefficiencies introduced by hashing real-world data that is noisy, structured, and most importantly is not independently and identically distributed. Through extensive tests, we find that permutation-grouping dramatically increases the efficiency of the overall retrieval system by lowering the number of low-probability candidates that must be examined by 30-50%.
Shumeet Baluja, Michele Covell, Sergey Ioffe
ICASSP1
2008 Video suggestion and discovery for youtube: taking random walks through the view graph
abstract
The rapid growth of the number of videos in YouTube provides enormous potential for users to find content of interest to them. Unfortunately, given the difficulty of searching videos, the size of the video repository also makes the discovery of new content a daunting task. In this paper, we present a novel method based upon the analysis of the entire user-video graph to provide personalized video suggestions for users. The resulting algorithm, termed Adsorption, provides a simple method to efficiently propagate preference information through a variety of graphs. We extensively test the results of the recommendations on a three month snapshot of live data from YouTube.
Shumeet Baluja, Rohan Seth, Yushi Jing, Jay Yagnik, Shankar Kumar, Deepak Ravichandran, Mohamed Aly 0002
WWW1
2008 Pagerank for product image search
abstract
In this paper, we cast the image-ranking problem into the task of identifying "authority" nodes on an inferred visual similarity graph and propose an algorithm to analyze the visual link structure that can be created among a group of images. Through an iterative procedure based on the PageRank computation, a numerical weight is assigned to each image; this measures its relative importance to the other images being considered. The incorporation of visual signals in this process differs from the majority of large-scale commercial-search engines in use today. Commercial search-engines often solely rely on the text clues of the pages in which images are embedded to rank images, and often entirely ignore the content of the images themselves as a ranking signal. To quantify the performance of our approach in a real-world system, we conducted a series of experiments based on the task of retrieving images for 2000 of the most popular products queries. Our experimental results show significant improvement, in terms of user satisfaction and relevancy, in comparison to the most recent Google Image Search results.
Yushi Jing, Shumeet Baluja
WWW2
2008 Learning to hash: forgiving hash functions and applications
Shumeet Baluja, Michele Covell
Data Min. Knowl. Discov.1
2008 Mass personalization: social and interactive applications using sound-track identification
Michael Fink 0002, Michele Covell, Shumeet Baluja
Multim. Tools Appl.3
2008 VisualRank: Applying PageRank to Large-Scale Image Search
abstract
Because of the relative ease in understanding and processing text, commercial image-search systems often rely on techniques that are largely indistinguishable from text-search. Recently, academic studies have demonstrated the effectiveness of employing image-based features to provide alternative or additional signals. However, it remains uncertain whether such techniques will generalize to a large number of popular web queries, and whether the potential improvement to search quality warrants the additional computational cost. In this work, we cast the image-ranking problem into the task of identifying "authority" nodes on an inferred visual similarity graph and propose VisualRank to analyze the visual link structures among images. The images found to be "authorities" are chosen as those that answer the image-queries well. To understand the performance of such an approach in a real system, we conducted a series of large-scale experiments based on the task of retrieving images for 2000 of the most popular products queries. Our experimental results show significant improvement, in terms of user satisfaction and relevancy, in comparison to the most recent Google Image Search results. Maintaining modest computational cost is vital to ensuring that this procedure can be used in practice; we describe the techniques required to make this system practical for large scale deployment in commercial search engines.
Yushi Jing, Shumeet Baluja
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Waveprint: Efficient wavelet-based audio fingerprinting
Shumeet Baluja, Michele Covell
Pattern Recognit.1
2007 Audio Fingerprinting: Combining Computer Vision & Data Stream Processing
abstract
In this paper, we present waveprint, a novel system for audio identification. Waveprint uses a combination of computer-vision techniques and large-scale-data-stream processing algorithms to create compact fingerprints of audio data that can be efficiently matched. The resulting system has excellent identification capabilities for small snippets of audio that have been degraded in a variety of manners, including competing noise, poor recording quality, and cell-phone playback. We measure the tradeoffs between performance, memory usage, and computation through extensive experimentation. The system is more efficient in terms of memory usage and computation, while being more accurate, when compared with previous state of the art systems.
Shumeet Baluja, Michele Covell
ICASSP (2)1
2007 Known-Audio Detection using Waveprint: Spectrogram Fingerprinting by Wavelet Hashing
abstract
In this paper, we present a novel system for detecting known audio. We start with Waveprint, an audio identification system that, given a probe snippet, efficiently provides reliable forced-choice ranking of entries from an audio database. For open-set detection, we can re-examine the best-ranked matches from waveprint using simple temporal-ordering-based processing. The resulting system has excellent detection capabilities for small snippets of audio that have been degraded in a variety of manners, including competing noise, poor recording quality, and cell-phone playback. The system is more accurate than the previous state-of-the-art system while being more efficient and flexible in memory usage and computation.
Michele Covell, Shumeet Baluja
ICASSP (1)2
2007 Learning "Forgiving" Hash Functions: Algorithms and Large Scale Tests
Shumeet Baluja, Michele Covell
IJCAI1
2007 The role of context in query input: using contextual signals to complete queries on mobile devices
abstract
The difficulty of entering queries from impoverished keyboards impedes the use of web search on mobile devices. On average, it takes a mobile user approximately 60 seconds to enter a query from a 9-key keypad [1]. In this paper, we explore the use of contextual signals to facilitate query entry on mobile phones. We present a query prediction system which offers automatically generated word completions as the user is typing her query. The query prediction system redefines the prediction dictionary after considering contextual signals such as the application being used (e.g. search vs. general text messaging), the inferred location of the user, the time of day and day of week. We demonstrate a 46.4% improvement in query entry, measured by number of key presses needed to enter queries. We found that the two contextual signals that make the largest impact are knowledge of the application being used and the location of the user.
Maryam Kamvar, Shumeet Baluja
Mobile HCI2
2007 Boosting Sex Identification Performance
Shumeet Baluja, Henry A. Rowley
Int. J. Comput. Vis.1
2007 Automated image-orientation detection: a scalable boosting approach
Shumeet Baluja
Pattern Anal. Appl.1
2006 A large scale study of wireless search behavior: Google mobile search
abstract
We present a large scale study of search patterns on Google's mobile search interface. Our goal is to understand the current state of wireless search by analyzing over 1 Million hits to Google's mobile search sites. Our study also includes the examination of search queries and the general categories under which they fall. We follow users throughout multiple interactions to determine search behavior; we estimate how long they spend inputting a query, viewing the search results, and how often they click on a search result. We also compare and contrast search patterns between 12-key keypad phones (cellphones), phones with QWERTY keyboards (PDAs) and conventional computers.
Maryam Kamvar, Shumeet Baluja
CHI2
2006 Advertisement Detection and Replacement using Acoustic and Visual Repetition
abstract
In this paper, we propose a method for detecting and precisely segmenting repeated sections of broadcast streams. This method allows advertisements to be removed and replaced with new ads in redistributed television material. The detection stage starts from acoustic matches and validates the hypothesized matches using the visual channel. Finally, the precise segmentation uses fine-grain acoustic match profiles to determine start and end-points. The approach is both efficient and robust to broadcast noise and differences in broadcaster signals. Our final result is nearly perfect, with better than 99% precision, at a recall rate of 95% for repeated advertisements
Michele Covell, Shumeet Baluja, Michael Fink 0002
MMSP2
2006 Browsing on small screens: recasting web-page segmentation into an efficient machine learning framework
abstract
Fitting enough information from webpages to make browsing on small screens compelling is a challenging task. One approach is to present the user with a thumbnail image of the full web page and allow the user to simply press a single key to zoom into a region (which may then be transcoded into wml/xhtml, summarized, etc). However, if regions for zooming are presented naively, this yields a frustrating experience because of the number of coherent regions, sentences, images, and words that may be inadvertently separated. Here, we cast the web page segmentation problem into a machine learning framework, where we re-examine this task through the lens of entropy reduction and decision tree learning. This yields an efficient and effective page segmentation algorithm. We demonstrate how simple techniques from computer vision can be used to fine-tune the results. The resulting segmentation keeps coherent regions together when tested on a broad set of complex webpages.
Shumeet Baluja
WWW1
2005 Boosting Sex Identification Performance
Shumeet Baluja, Henry A. Rowley
AAAI1
2005 Large scale performance measurement of content-based automated image-orientation detection
abstract
With the proliferation of digital cameras and self-publishing of photos, automatic detection of image orientation will become an important part of photo management systems. In this paper, we perform a large scale empirical test to determine whether the common techniques to automatically determine a photo's orientation are robust enough to handle the breadth of real-world images. We use a wide variety of features and color-spaces to address this problem. We use test photos gathered from the Web and photo collections, including photos that are in color and black and white, realistic and abstract, and outdoor and indoor. Results show that current methods give satisfactory results on only a small subset of these images.
Shumeet Baluja, Henry A. Rowley
ICIP (2)1
2004 Efficient face orientation discrimination
abstract
The paper presents efficient methods to address the problem of discriminating between live facial orientations. We present the most efficient methods for this task to date, which can accurately discriminate between five facial orientations with approximately 92% accuracy using fewer than 30 pixel comparisons and greater than 99% accuracy using 150 pixel comparisons. We achieve these rates by using a boosting method to select from a large set of extremely simple features. Comparisons to other methods are given.
Shumeet Baluja, Mehran Sahami, Henry A. Rowley
ICIP1
2004 The Happy Searcher: Challenges in Web Information Retrieval
Mehran Sahami, Vibhu O. Mittal, Shumeet Baluja, Henry A. Rowley
PRICAI3
2002 Using a priori knowledge to create probabilistic models for optimization
Shumeet Baluja
Int. J. Approx. Reason.1
2000 Memory-Based Face Recognition for Visitor Identification
abstract
We show that a simple, memory-based technique for appearance-based face recognition, motivated by the real-world task of visitor identification, can outperform more sophisticated algorithms that use principal components analysis (PCA) and neural networks. This technique is closely related to correlation templates; however, we show that the use of novel similarity measures greatly improves performance. We also show that augmenting the memory base with additional, synthetic face images results in further improvements in performance. Results of extensive empirical testing on two standard face recognition datasets are presented, and direct comparisons with published work show that our algorithm achieves comparable (or superior) results. Our system is incorporated into an automated visitor identification system that has been operating successfully in an outdoor environment since January 1999.
Terence Sim, Rahul Sukthankar, Matthew D. Mullin, Shumeet Baluja
FG4
2000 Applying Machine Learning for High-Performance Named-Entity Extraction
abstract
This paper describes a machine learning approach to building an efficient and accurate name spotting system. Finding names in free text is an important task in many text‐based applications. Most previous approaches were based on hand‐crafted modules encoding language and genre‐specific knowledge. These approaches had at least two shortcomings: They required large amounts of time and expertise to develop and were not easily portable to new languages and genres. This paper describes an extensible system that automatically combines weak evidence from different, easily available sources: parts‐of‐speech tags, dictionaries, and surface‐level syntactic information such as capitalization and punctuation. Individually, each piece of evidence is insufficient for robust name detection. However, the combination of evidence, through standard machine learning techniques, yields a system that achieves performance equivalent to the best existing hand‐crafted approaches.
Shumeet Baluja, Vibhu O. Mittal, Rahul Sukthankar
Comput. Intell.1
2000 Using Labeled and Unlabeled Data for Probabilistic Modeling of Face Orientation
abstract
This paper describes probabilistic modeling methods to solve the problem of discriminating between five facial orientations with very little labeled data. Three models are explored. The first model maintains no inter-pixel dependencies, the second model is capable of modeling a set of arbitrary pair-wise dependencies, and the last model allows dependencies only between neighboring pixels. We show that for all three of these models, the accuracy of the learned models can be greatly improved by augmenting a small number of labeled training images with a large set of unlabeled images using Expectation–Maximization. This is important because it is often difficult to obtain image labels, while many unlabeled images are readily available. Through a large set of empirical tests, we examine the benefits of unlabeled data for each of the models. By using only two randomly selected labeled examples per class, we can discriminate between the five facial orientations with an accuracy of 94%; with six labeled examples, we achieve an accuracy of 98%.
Shumeet Baluja
Int. J. Pattern Recognit. Artif. Intell.1
1998 Rotation Invariant Neural Network-Based Face Detection
abstract
In this paper, we present a neural network-based face detection system. Unlike similar systems which are limited to detecting upright, frontal faces, this system detects faces at any degree of rotation in the image plane. The system employs multiple networks; a "router" network first processes each input window to determine its orientation and then uses this information to prepare the window for one or more "detector" networks. We present the training methods for both types of networks. We also perform sensitivity analysis on the networks, and present empirical results on a large test set. Finally, we present preliminary results for detecting faces rotated out of the image plane, such as profiles and semi-profiles.
Henry A. Rowley, Shumeet Baluja, Takeo Kanade
CVPR2
1998 Rotation Invariant Neural Network-Based Face Detection
abstract
In this paper, we present a neural network-based face detection system. Unlike similar systems which are limited to detecting upright, frontal faces, this system detects faces at any degree of rotation in the image plane. The system employs multiple networks; a "router" network first processes each input window to determine its orientation and then uses this information to prepare the window for one or more "detector" networks. We present the training methods for both types of networks. We also perform sensitivity analysis on the networks, and present empirical results on a large test set. Finally, we present preliminary results for detecting faces rotated out of the image plane, such as profiles and semi-profiles. 1. Introduction In our observations of face detector demonstrations, we have found that users expect faces to be detected at any angle, as shown in Figure 1. In this paper, we present a neural network-based algorithm to detect faces in gray-scale images. Unlike similar pre...
Henry A. Rowley, Shumeet Baluja, Takeo Kanade
CVPR2
1998 Making Templates Rotationally Invariant. An Application to Rotated Digit Recognition
Shumeet Baluja
NIPS1
1998 Probabilistic Modeling for Face Orientation Discrimination: Learning from Labeled and Unlabeled Data
Shumeet Baluja
NIPS1
1998 Finding Regions of Uncertainty in Learned Models: An Application to Face Detection
Shumeet Baluja
PPSN1
1998 Evolution-Based Methods for Selecting Point Data for Object Localization: Applications to Computer-Assisted Surgery
Shumeet Baluja, David Simon
Appl. Intell.1
1998 Multiple Adaptive Agents for Tactical Driving
Rahul Sukthankar, Shumeet Baluja, John A. Hancock
Appl. Intell.2
1998 Neural Network-Based Face Detection
abstract
We present a neural network-based upright frontal face detection system. A retinally connected neural network examines small windows of an image and decides whether each window contains a face. The system arbitrates between multiple networks to improve performance over a single network. We present a straightforward procedure for aligning positive face examples for training. To collect negative examples, we use a bootstrap algorithm, which adds false detections into the training set as training progresses. This eliminates the difficult task of manually selecting nonface training examples, which must be chosen to span the entire space of nonface images. Simple heuristics, such as using the fact that faces rarely overlap in images, can further improve the accuracy. Comparisons with several other state-of-the-art face detection systems are presented, showing that our system has comparable performance in terms of detection and false-positive rates.
Henry A. Rowley, Shumeet Baluja, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Using Optimal Dependency-Trees for Combinational Optimization
Shumeet Baluja, Scott Davies
ICML1
1997 Evolving an intelligent vehicle for tactical reasoning in traffic
abstract
Recent research in automated highway systems has ranged from low-level vision-based controllers to high-level route-guidance software. However there is currently no system for tactical-level reasoning. Such a system should address tasks such as passing cars, making exits on time, and merging into a traffic stream. Our approach to this intermediate-level planning combines a distributed reasoning system (PolySAPIENT) with a novel evolutionary optimization strategy (PBIL). PBIL automatically tunes PolySAPIENT module parameters in simulation by evaluating candidate modules on various traffic scenarios. Since the control interface to the simulated vehicles is identical to that on the Carnegie Mellon Navlab vehicles, modules developed using this process can be directly ported to existing hardware. This method is currently being applied to the automated highway system domain; it also generalizes to many complex robotics tasks where multiple interacting modules must simultaneously be configured without individual module feedback.
Rahul Sukthankar, Shumeet Baluja, John A. Hancock
ICRA2
1997 Using Expectation to Guide Processing: A Study of Three Real-World Applications
Shumeet Baluja
NIPS1
1997 Dynamic Relevance: Vision-Based Focus of Attention Using Artificial Neural Networks. (Technical Note)
Shumeet Baluja, Dean Pomerleau
Artif. Intell.1
1996 Neural Network-Based Face Detection
abstract
We present a neural network-based face detection system. A retinally connected neural network examines small windows of an image and decides whether each window contains a face. The system arbitrates between multiple networks to improve performance over a single network. We use a bootstrap algorithm for training the networks, which adds false detections into the training set as training progresses. This eliminates the difficult task of manually selecting non-face training examples, which must be chosen to span the entire space of non-face images. Comparisons with other state-of-the-art face detection systems are presented; our system has better performance in terms of detection and false-positive rates.
Henry A. Rowley, Shumeet Baluja, Takeo Kanade
CVPR2
1996 Genetic Algorithms and Explicit Search Statistics
Shumeet Baluja
NIPS1
1996 Evolution of an artificial neural network based autonomous land vehicle controller
abstract
This paper presents an evolutionary method for creating an artificial neural network based autonomous land vehicle controller. The evolved controllers perform better in unseen situations than those trained with an error backpropagation learning algorithm designed for this task. In this paper, an overview of the previous connectionist based approaches to this task is given, and the evolutionary algorithms used in this study are described in detail. Methods for reducing the high computational costs of training artificial neural networks with evolutionary algorithms are explored. Error metrics specific to the task of autonomous vehicle control are introduced; the evolutionary algorithms guided by these error metrics reveal improved performance over those guided by the standard sum-squared error metric. Finally, techniques for integrating evolutionary search and error backpropagation are presented. The evolved networks are designed to control Carnegie Mellon University's NAVLAB vehicles in road following tasks.
Shumeet Baluja
IEEE Trans. Syst. Man Cybern. Part B1
1995 Removing the Genetics from the Standard Genetic Algorithm
Shumeet Baluja, Rich Caruana
ICML1
1995 Using the Representation in a Neural Network's Hidden Layer for Task-Specific Focus of Attention
Shumeet Baluja, Dean Pomerleau
IJCAI1
1995 Using the Future to Sort Out the Present: Rankprop and Multitask Learning for Medical Risk Evaluation
Rich Caruana, Shumeet Baluja, Tom M. Mitchell
NIPS2
1995 Human Face Detection in Visual Scenes
Henry A. Rowley, Shumeet Baluja, Takeo Kanade
NIPS2
1994 Using a Saliency Map for Active Spatial Selective Attention: Implementation & Initial Results
abstract
In many vision based tasks, the ability to focus attention on the important portions of a scene is crucial for good performance on the tasks. In this paper we present a simple method of achieving spatial selective attention through the use of a saliency map. The saliency map indicates which regions of the input retina are important for performing the task. The saliency map is cre(cid:173) ated through predictive auto-encoding. The performance of this method is demonstrated on two simple tasks which have multiple very strong distract(cid:173) ing features in the input retina. Architectural extensions and application directions for this model are presented.
Shumeet Baluja, Dean Pomerleau
NIPS1
1994 Towards Automated Artificial Evolution for Computer-generated Images
abstract
In 1991, Karl Sims presented work on artificial evolution in which he used genetic algorithms to evolve complex structures for use in computer-generated images and animations. The evolution of the computer-generated images progressed from simple, randomly generated shapes to interesting images which the users created interactively. The evolution advanced under the constant guidance and supervision of the user. This paper describes attempts to automate the process of image evolution through the use of artificial neural networks. The central objective of this study is to learn the user's preferences, and to apply this knowledge to evolve aesthetically pleasing images which are similar to those evolved through interactive sessions with the user. This paper presents a detailed performance analysis of both the successes and shortcomings encountered in the use of five artificial neural network architectures. Further possibilities for improving the performance of a fully automated system are also discussed.
Shumeet Baluja, Dean Pomerleau, Todd Jochem
Connect. Sci.1
1993 The Evolution of Gennetic Algorithms: Towards Massive Parallelism
Shumeet Baluja
ICML1
1993 Non-Intrusive Gaze Tracking Using Artificial Neural Networks
Shumeet Baluja, Dean Pomerleau
NIPS1