Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Thomas K. Leung

dblp:39/2593 · also Thomas Leung · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 9 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Trustworthy machine learning · 21% Generative modeling · 20% Image recognition and object detection · 15%
Computer graphics and multimedia
8 papers
Visual content generation and editing · 74% Image and video processing · 14% Computational photography and imaging · 9%

Topics — the 30 heaviest of 55, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › AI-generated content detection
AI-generated image detection
0.812024
FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable Diffusion · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.812024
Directed Diffusion: Direct Control of Object Placement through Attention Guidance · AAAI 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Directed Diffusion: Direct Control of Object Placement through Attention Guidance · AAAI 2024
Visual content generation and editing
image generation
0.812024
Directed Diffusion: Direct Control of Object Placement through Attention Guidance · AAAI 2024
Machine learning › Trustworthy machine learning
robustness
0.622018
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels · ICML 2018
Improving the Robustness of Deep Neural Networks via Stability Training · CVPR 2016
Computer vision › Video understanding and tracking
visual summarization
0.612022
NewsStories: Illustrating Articles with Visual Summaries · ECCV (36) 2022
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.522018
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels · ICML 2018
Handling label noise in video classification via multiple instance learning · ICCV 2011
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.422017
No Fuss Distance Metric Learning Using Proxies · ICCV 2017
MatchNet: Unifying feature and metric learning for patch-based matching · CVPR 2015
Computer vision › Image recognition and object detection
image classification
0.322018
Improving the Robustness of Deep Neural Networks via Stability Training · CVPR 2016
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels · ICML 2018
Machine learning › Learning paradigms
curriculum learning
0.312018
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels · ICML 2018
Machine learning › Learning paradigms › curriculum learning
data-driven curriculum
0.312018
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels · ICML 2018
Computer vision › Video understanding and tracking
video classification
0.322014
Large-Scale Video Classification with Convolutional Neural Networks · CVPR 2014
Handling label noise in video classification via multiple instance learning · ICCV 2011
Machine learning › Representation and self-supervised learning › representation learning › metric learning › deep metric learning
triplet loss
0.312017
No Fuss Distance Metric Learning Using Proxies · ICCV 2017
Computer vision › Image recognition and object detection › image classification
robust image classification
0.212016
Improving the Robustness of Deep Neural Networks via Stability Training · CVPR 2016
Computer vision › 3D vision › local feature descriptor
local descriptor learning
0.212015
MatchNet: Unifying feature and metric learning for patch-based matching · CVPR 2015
Computer vision › Video understanding and tracking
action recognition
0.212014
Large-Scale Video Classification with Convolutional Neural Networks · CVPR 2014
Machine learning › Deep learning architectures and training
deep ranking
0.212014
Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014
Computer vision › Image recognition and object detection
image retrieval
0.212014
Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014
Computer vision › Image recognition and object detection
image similarity
0.212014
Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal CNN
0.212014
Large-Scale Video Classification with Convolutional Neural Networks · CVPR 2014
Machine learning › Learning paradigms
multiple instance learning
0.112011
Handling label noise in video classification via multiple instance learning · ICCV 2011
Computer vision › Image recognition and object detection
weakly-labeled data
0.112011
Handling label noise in video classification via multiple instance learning · ICCV 2011
Computer vision › Image recognition and object detection
texture classification
0.122004
Texton Correlation for Recognition · ECCV (1) 2004
Recognizing Surfaces using Three-Dimensional Textons · ICCV 1999
Computer vision › Face, body and person analysis
person identification
0.112006
Context-Aided Human Recognition - Clustering · ECCV (3) 2006
Image and video processing
image segmentation
0.132001
Contour and Texture Analysis for Image Segmentation · Int. J. Comput. Vis. 2001
Contour Continuity in Region Based Image Segmentation · ECCV (1) 1998
Detecting, localizing and grouping repeated scene elements from an image · ECCV (1) 1996
Computer vision › Image recognition and object detection › texture classification
material recognition
0.122001
Representing and Recognizing the Visual Appearance of Materials using Three-dimensional Textons · Int. J. Comput. Vis. 2001
Recognizing Surfaces using Three-Dimensional Textons · ICCV 1999
Rendering
appearance modeling
0.012001
Representing and Recognizing the Visual Appearance of Materials using Three-dimensional Textons · Int. J. Comput. Vis. 2001
Computational photography and imaging
illumination modeling
0.012001
Statistics of Real-World Illumination · CVPR (2) 2001
Computational photography and imaging
projector-camera systems
0.012001
Self-Calibrating Camera Projector Systems for Interactive Displays and Presentations · ICCV 2001
Computational photography and imaging › camera calibration
self-calibration
0.012001
Self-Calibrating Camera Projector Systems for Interactive Displays and Presentations · ICCV 2001

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.5cross-attention guidance · 1.5reverse image search · 0.8diffusion model inversion · 0.8student network · 0.3sample weighting · 0.3mentor network · 0.3triplet mining · 0.3proxy-based loss · 0.3deep neural network · 0.2camera calibration · 0.1wavelet coefficient analysis · 0.0three-dimensional textons · 0.0high dynamic range imaging · 0.0harmonic spectra analysis · 0.0non-parametric sampling · 0.0markov random field · 0.0contour continuity · 0.0
YearPublicationVenuePosition
2024 Directed Diffusion: Direct Control of Object Placement through Attention Guidance
abstract
Text-guided diffusion models such as DALLE-2, Imagen, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality. However, these models often struggle to compose scenes containing several key objects such as characters in specified positional relationships. The missing capability to ``direct'' the placement of characters and objects both within and across images is crucial in storytelling, as recognized in the literature on film and animation theory. In this work, we take a particularly straightforward approach to providing the needed direction. Drawing on the observation that the cross-attention maps for prompt words reflect the spatial layout of objects denoted by those words, we introduce an optimization objective that produces ``activation'' at desired positions in these cross-attention maps. The resulting approach is a step toward generalizing the applicability of text-guided diffusion models beyond single images to collections of related images, as in storybooks. Directed Diffusion provides easy high-level positional control over multiple objects, while making use of an existing pre-trained model and maintaining a coherent blend between the positioned objects and the background. Moreover, it requires only a few lines to implement.
Wan-Duo Kurt Ma, Avisek Lahiri, John P. Lewis, Thomas K. Leung, W. Bastiaan Kleijn
AAAI4
2024 FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable Diffusion
abstract
Due to the high potential for abuse of GenAl systems, the task of detecting synthetic images has recently become of great interest to the research community. Unfortunately, ex-isting image-space detectors quickly become obsolete as new high-fidelity text-to-image models are developed at blinding speed. In this work, we propose a new synthetic image detector that uses features obtained by inverting an open-source pre-trained Stable Diffusion model. We show that these inversion features enable our detector to generalize well to unseen generators of high visual fidelity (e.g., DALL.E 3) even when the detector is trained only on lower fidelity fake images generated via Stable Diffusion. This detector achieves new state-of-the-art across multiple training and evaluation se-tups. Moreover, we introduce a new challenging evaluation protocol that uses reverse image search to mitigate stylistic and thematic biases in the detector evaluation. We show that the resulting evaluation scores align well with detectors' in-the-wild performance, and release these datasets as public benchmarks for future research.
George Cazenavette, Avneesh Sud, Thomas K. Leung, Ben Usman
CVPR3
2022 NewsStories: Illustrating Articles with Visual Summaries
Reuben Tan, Bryan A. Plummer, Kate Saenko, John P. Lewis, Avneesh Sud, Thomas K. Leung
ECCV (36)6
2018 Towards A Semantic Perceptual Image Metric
abstract
We present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 140k unique images annotated with ground truth by human raters who received minimal instruction. The resulting metric shows competitive performance on TID 2013, a database widely used to assess image quality assessments methods. More interestingly, it shows strong responses to objects potentially carrying semantic relevance such as faces and text, which we demonstrate using a visualization technique and ablation experiments. In effect, the metric appears to model a higher influence of semantic context on judgments, which we observe particularly in untrained raters. As the vast majority of users of image processing systems are unfamiliar with Image Quality Assessment (IQA) tasks, these findings may have significant impact on real-world applications of perceptual metrics.
Troy T. Chinen, Jona Ballé, Chunhui Gu, Sung Jin Hwang, Sergey Ioffe, Nicholas Johnston, Thomas K. Leung, David Minnen, Sean M. O'Malley, Charles Rosenberg 0001, George Toderici
ICIP7
2018 MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
abstract
Recent deep networks are capable of memorizing the entire data even when the labels are completely random. To overcome the overfitting on corrupted labels, we propose a novel technique of learning another neural network, called MentorNet, to supervise the training of the base deep networks, namely, StudentNet. During training, MentorNet provides a curriculum (sample weighting scheme) for StudentNet to focus on the sample the label of which is probably correct. Unlike the existing curriculum that is usually predefined by human experts, MentorNet learns a data-driven curriculum dynamically with StudentNet. Experimental results demonstrate that our approach can significantly improve the generalization performance of deep networks trained on corrupted training data. Notably, to the best of our knowledge, we achieve the best-published result on WebVision, a large benchmark containing 2.2 million images of real-world noisy labels.
Lu Jiang 0004, Zhengyuan Zhou, Thomas K. Leung, Li-Jia Li 0001, Li Fei-Fei 0001
ICML3
2017 No Fuss Distance Metric Learning Using Proxies
abstract
We address the problem of distance metric learning (DML), defined as learning a distance consistent with a notion of semantic similarity. Traditionally, for this problem supervision is expressed in the form of sets of points that follow an ordinal relationship - an anchor point x is similar to a set of positive points Y , and dissimilar to a set of negative points Z, and a loss defined over these distances is minimized. While the specifics of the optimization differ, in this work we collectively call this type of supervision Triplets and all methods that follow this pattern Triplet-Based methods. These methods are challenging to optimize. A main issue is the need for finding informative triplets, which is usually achieved by a variety of tricks such as increasing the batch size, hard or semi-hard triplet mining, etc. Even with these tricks, the convergence rate of such methods is slow. In this paper we propose to optimize the triplet loss on a different space of triplets, consisting of an anchor data point and similar and dissimilar proxy points which are learned as well. These proxies approximate the original data points, so that a triplet loss over the proxies is a tight upper bound of the original loss. This proxy-based loss is empirically better behaved. As a result, the proxy-loss improves on state-of-art results for three standard zero-shot learning datasets, by up to 15% points, while converging three times as fast as other triplet-based losses.
Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung, Sergey Ioffe
ICCV3
2016 Improving the Robustness of Deep Neural Networks via Stability Training
abstract
In this paper we address the issue of output instability of deep neural networks: small perturbations in the visual input can significantly distort the feature embeddings and output of a neural network. Such instability affects many deep architectures with state-of-the-art performance on a wide range of computer vision tasks. We present a general stability training method to stabilize deep networks against small input distortions that result from various types of common image processing, such as compression, rescaling, and cropping. We validate our method by stabilizing the state of-the-art Inception architecture [11] against these types of distortions. In addition, we demonstrate that our stabilized model gives robust state-of-the-art performance on largescale near-duplicate detection, similar-image ranking, and classification on noisy datasets.
Stephan Zheng, Yang Song 0009, Thomas K. Leung, Ian J. Goodfellow
CVPR3
2015 MatchNet: Unifying feature and metric learning for patch-based matching
abstract
Motivated by recent successes on learning feature representations and on learning feature comparison functions, we propose a unified approach to combining both for training a patch matching system. Our system, dubbed Match-Net, consists of a deep convolutional network that extracts features from patches and a network of three fully connected layers that computes a similarity between the extracted features. To ensure experimental repeatability, we train MatchNet on standard datasets and employ an input sampler to augment the training set with synthetic exemplar pairs that reduce overfitting. Once trained, we achieve better computational efficiency during matching by disassembling MatchNet and separately applying the feature computation and similarity networks in two sequential stages. We perform a comprehensive set of experiments on standard datasets to carefully study the contributions of each aspect of MatchNet, with direct comparisons to established methods. Our results confirm that our unified approach improves accuracy over previous state-of-the-art results on patch matching datasets, while reducing the storage requirement for descriptors. We make pre-trained MatchNet publicly available.
Xufeng Han, Thomas K. Leung, Yangqing Jia, Rahul Sukthankar, Alexander C. Berg
CVPR2
2014 Large-Scale Video Classification with Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have been established as a powerful class of models for image recognition problems. Encouraged by these results, we provide an extensive empirical evaluation of CNNs on large-scale video classification using a new dataset of 1 million YouTube videos belonging to 487 classes. We study multiple approaches for extending the connectivity of a CNN in time domain to take advantage of local spatio-temporal information and suggest a multiresolution, foveated architecture as a promising way of speeding up the training. Our best spatio-temporal networks display significant performance improvements compared to strong feature-based baselines (55.3% to 63.9%), but only a surprisingly modest improvement compared to single-frame models (59.3% to 60.9%). We further study the generalization performance of our best model by retraining the top layers on the UCF-101 Action Recognition dataset and observe significant performance improvements compared to the UCF-101 baseline model (63.3% up from 43.9%).
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas K. Leung, Rahul Sukthankar, Li Fei-Fei 0001
CVPR4
2014 Learning Fine-Grained Image Similarity with Deep Ranking
abstract
Learning fine-grained image similarity is a challenging task. It needs to capture between-class and within-class image differences. This paper proposes a deep ranking model that employs deep learning techniques to learn similarity metric directly from images. It has higher learning capability than models based on hand-crafted features. A novel multiscale network structure has been developed to describe the images effectively. An efficient triplet sampling algorithm is also proposed to learn the model with distributed asynchronized stochastic gradient. Extensive experiments show that the proposed algorithm outperforms models based on hand-crafted visual features and deep classification models.
Jiang Wang 0001, Yang Song 0009, Thomas K. Leung, Charles Rosenberg 0001, Jingbin Wang, James Philbin, Bo Chen 0019, Ying Wu 0001
CVPR3
2011 Handling label noise in video classification via multiple instance learning
abstract
In many classification tasks, the use of expert-labeled data for training is often prohibitively expensive. The use of weakly-labeled data is an attractive solution but raises the problem of label noise. Multiple instance learning, whereby training samples are “bagged” instead of treated as singletons, offers a possible approach to mitigating the effects of label noise. In this paper, we propose the use of MILBoost [28] in a large-scale video taxonomic classification system comprised of hundreds of binary classifiers to handle noisy training data. We test on data with both artificial and real-world noise and compare against the state-of-the-art classifiers based on AdaBoost. We also explore the effects of different bag sizes on different levels of noise on the final classifier performance. Experiments show that when training classifiers with noisy data, MILBoost provides an improvement in performance.
Thomas K. Leung, Yang Song 0009
ICCV1
2006 Context-Aided Human Recognition - Clustering
Yang Song 0009, Thomas K. Leung
ECCV (3)2
2004 Texton Correlation for Recognition
Thomas K. Leung
ECCV (1)1
2001 Statistics of Real-World Illumination
abstract
While computer vision systems often assume simple illumination models, real-world illumination is highly complex, consisting of reflected light from every direction as well as distributed and localized primary light sources. One can capture the illumination incident at a point in the real world from every direction photographically using a spherical illumination map. This paper illustrates, through analysis of photographically-acquired, high dynamic range illumination maps, that real-world illumination shares many of the statistical properties of natural images. In particular, the marginal and joint wavelet coefficient distributions, directional derivative distributions, and harmonic spectra of illumination maps resemble those documented in the natural image statistics literature. However, illumination maps differ from standard photographs in that illumination maps are statistically non-stationary and may contain localized light sources that dominate their power spectra. Our work provides a foundation for statistical models of real-world illumination that may facilitate robust estimation of shape, reflectance, and illumination from images.
Ron O. Dror, Thomas K. Leung, Edward H. Adelson, Alan S. Willsky
CVPR (2)2
2001 Self-Calibrating Camera Projector Systems for Interactive Displays and Presentations
abstract
The authors demonstrate a self-calibrating system that employs uncalibrated cameras and microportable projectors to create novel interactive displays and presentations. Three benefits of ther system are detailed.
Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar, James M. Rehg, David Hsu, Thomas K. Leung
ICCV6
2001 Representing and Recognizing the Visual Appearance of Materials using Three-dimensional Textons
Thomas K. Leung, Jitendra Malik
Int. J. Comput. Vis.1
2001 Contour and Texture Analysis for Image Segmentation
Jitendra Malik, Serge J. Belongie, Thomas K. Leung, Jianbo Shi
Int. J. Comput. Vis.3
1999 Texture Synthesis by Non-parametric Sampling
abstract
A non-parametric method for texture synthesis is proposed. The texture synthesis process grows a new image outward from an initial seed, one pixel at a time. A Markov random field model is assumed, and the conditional distribution of a pixel given all its neighbors synthesized so far is estimated by querying the sample image and finding all similar neighborhoods. The degree of randomness is controlled by a single perceptually intuitive parameter. The method aims at preserving as much local structure as possible and produces good results for a wide variety of synthetic and real-world textures.
Alexei A. Efros, Thomas K. Leung
ICCV2
1999 Recognizing Surfaces using Three-Dimensional Textons
abstract
We study the recognition of surfaces made from different materials such as concrete, rug, marble or leather on the basis of their textural appearance. Such natural textures arise from spatial variation of two surface attributes: (1) reflectance and (2) surface normal. In this paper, we provide a unified model to address both these aspects of natural texture. The main idea is to construct a vocabulary of prototype tiny surface patches with associated local geometric and photometric properties. We call these 3D textons. Examples might be ridges, grooves, spots or stripes or combinations thereof Associated with each texton is an appearance vector, which characterizes the local irradiance distribution, represented as a set of linear Gaussian derivative filter outputs, under different lighting and viewing conditions. Given a large collection of images of different materials, a clustering approach is used to acquire a small (on the order of 100) 3D texton vocabulary. Given a few (1 to 4) images of any material, it can be characterized using these textons. We demonstrate the application of this representation for recognition of the material viewed under novel lighting and viewing conditions.
Thomas K. Leung, Jitendra Malik
ICCV1
1999 Textons, Contours and Regions: Cue Integration in Image Segmentation
abstract
The paper makes two contributions: it provides (1) an operational definition of textons, the putative elementary units of texture perception, and (2) an algorithm for partitioning the image into disjoint regions of coherent brightness and texture, where boundaries of regions are defined by peaks in contour orientation energy and differences in texton densities across the contour. B. Julesz (1981) introduced the term texton, analogous to a phoneme in speech recognition, but did not provide an operational definition for gray-level images. We re-invent textons as frequently co-occurring combinations of oriented linear filter outputs. These can be learned using a K-means approach. By mapping each pixel to its nearest texton, the image can be analyzed into texton channels, each of which is a point set where discrete techniques such as Voronoi diagrams become applicable. Local histograms of texton frequencies can be used with a /spl chi//sup 2/ test for significant differences to find texture boundaries. Natural images contain both textured and untextured regions, so we combine this cue with that of the presence of peaks of contour energy derived from outputs of odd- and even-symmetric oriented Gaussian derivative filters. Each of these cues has a domain of applicability, so to facilitate cue combination we introduce a gating operator based on a statistical test for isotropy of Delaunay neighbors. Having obtained a local measure of how likely two nearby pixels are to belong to the same region, we use the spectral graph theoretic framework of normalized cuts to find partitions of the image into regions of coherent texture and brightness. Experimental results on a wide range of images are shown.
Jitendra Malik, Serge J. Belongie, Jianbo Shi, Thomas K. Leung
ICCV4
1998 Probablistic Affine Invariants for Recognition
abstract
Under a weak perspective camera model, the image plane coordinates in different views of a planar object are related by an affine transformation. Because of this property, researchers have attempted to use affine invariants for recognition. However, there are two problems with this approach: (1) objects or object classes with inherent variability cannot be adequately treated using invariants; and (2) in practice the calculated affine invariants can be quite sensitive to errors in the image plane measurements. In this paper we use probability distributions to address both of these difficulties. Under the assumption that the feature positions of a planar object can be modeled using a jointly Gaussian density, we have derived the joint density over the corresponding set of affine coordinates. Even when the assumptions of a planar object and a weak perspective camera model do not strictly hold, the results are useful because deviations from the ideal can be treated as deformability in the underlying object model.
Thomas K. Leung, Michael C. Burl, Pietro Perona
CVPR1
1998 Contour Continuity in Region Based Image Segmentation
Thomas K. Leung, Jitendra Malik
ECCV (1)1
1998 Image and Video Segmentation: The Normalized Cut Framework
abstract
We propose a segmentation system based on the normalized cut framework proposed by Shi and Malik (see Proc. IEEE Conf. Computer Vision and Pattern Recognition, San Juan, Puerto Rico, p.731-7, 1997). The goal is to partition the image from a big picture point of view. Perceptually significant groups are detected first while small variations and details are treated later. Different image features-intensity, color, texture, contour continuity, motion and stereo disparity are treated in one uniform framework.
Jianbo Shi, Serge J. Belongie, Thomas K. Leung, Jitendra Malik
ICIP (1)3
1997 On Perpendicular Texture: Why do we see more flowers in the distance?
abstract
Almost all work on texture in the computer vision and graphics communities has modeled the texture as tangential, i.e. lying in the tangent plane to the surface. This is equivalent to thinking of the texture as a pattern painted on the surface. Three-dimensional textures, where the elements may point out of the surface, have largely been ignored. We study a special class of 3D textures, perpendicular textures where we can model the elements as being normal to the surface. The perspective projection of perpendicularly textured surfaces results in several interesting phenomena, which do not occur in the much-studied tangential texture case. These include occlusion, foreshortening and illumination. In this paper, we study the geometry of the problem, modeling the locations of the elements of the texture as being a realization of a spatial point process. Relations between slant and tilt of the surface, density and height of elements and occlusions are derived. Occlusions can now be used as a cue to infer shape, instead of being treated as a source of error.
Thomas K. Leung, Jitendra Malik
CVPR1
1996 Detecting, localizing and grouping repeated scene elements from an image
Thomas K. Leung, Jitendra Malik
ECCV (1)1
1996 Finding objects in image databases by grouping
abstract
Retrieving images from very large collections, using image content as a key, is becoming an important problem. Finding objects in image databases is a big challenge in the field. The paper describes our approach to object recognition, which is distinguished by: a rich involvement of early visual primitives, including color and texture; hierarchical grouping and learning strategies in the classification process; the ability to deal with rather general objects in uncontrolled configurations and contexts. We illustrate these properties with three case studies: one demonstrating the use of color and texture descriptors; one learning scenery concepts using grouped features; and one demonstrating a possible application domain in detecting naked people in a scene.
Jitendra Malik, David A. Forsyth, Margaret M. Fleck, Hayit Greenspan, Thomas K. Leung, Chad Carson, Serge J. Belongie, Christoph Bregler
ICIP (2)5
1995 Finding Faces in Cluttered Scenes Using Labeled Random Graph Matching
abstract
An algorithm for locating quasi-frontal views of human faces in cluttered scenes is presented. The algorithm works by coupling a set of local feature detectors with a statistical model of the mutual distances between facial features it is invariant with respect to translation, rotation (in the plane), and scale and can handle partial occlusions of the face. On a challenging database with complicated and varied backgrounds, the algorithm achieved a correct localization rate of 95% in images where the face appeared quasi-frontally.>
Thomas K. Leung, Michael C. Burl, Pietro Perona
ICCV1