Simon Baker

dblp:85/2539 · DBLP profile ↗
← Back
86ranked-venue papers
24as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 22 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 11 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
41 papers
Face, body and person analysis · 26% 3D vision · 21% Question answering and dialogue systems · 14%
Computer graphics and multimedia
26 papers
Image and video processing · 45% Computational photography and imaging · 18% Geometric modeling and processing · 15%

Topics — the 30 heaviest of 120, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
biomedical text mining
2.752026
BioTriplex : a full-text annotated corpus for fine-tuning language models in gene-disease relation extraction tasks · Bioinform. 2026
Text mining for contexts and relationships in cancer genomics literature · Bioinform. 2024
LION LBD: a literature-based discovery system for cancer biology · Bioinform. 2019
Bioinformatics and computational biology › biomedical text mining
named entity recognition
1.122024
Text mining for contexts and relationships in cancer genomics literature · Bioinform. 2024
LION LBD: a literature-based discovery system for cancer biology · Bioinform. 2019
Bioinformatics and computational biology › biomedical text mining
gene-disease relation extraction
1.012026
BioTriplex : a full-text annotated corpus for fine-tuning language models in gene-disease relation extraction tasks · Bioinform. 2026
Bioinformatics and computational biology
cancer genomics
0.812024
Text mining for contexts and relationships in cancer genomics literature · Bioinform. 2024
Bioinformatics and computational biology › biomedical text mining
relation extraction
0.812024
Text mining for contexts and relationships in cancer genomics literature · Bioinform. 2024
Computer vision › Face, body and person analysis
face recognition
0.692012
Memory constrained face recognition · CVPR 2012
Which faces to tag: Adding prior constraints into active learning · ICCV 2009
Simultaneous super-resolution and feature extraction for recognition of low-resolution faces · CVPR 2008
Machine learning › Learning paradigms
curriculum learning
0.512021
Dialogue Response Selection with Hierarchical Curriculum Learning · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems
response selection
0.512021
Dialogue Response Selection with Hierarchical Curriculum Learning · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › dialogue generation
stylized dialogue generation
0.512021
PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval Memory · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Bioinformatics and computational biology
cancer biology
0.422019
Cancer Hallmarks Analytics Tool (CHAT): a text mining approach to organize and evaluate scientific literature on cancer · Bioinform. 2017
LION LBD: a literature-based discovery system for cancer biology · Bioinform. 2019
Bioinformatics and computational biology › biomedical text mining
literature-based discovery
0.412019
LION LBD: a literature-based discovery system for cancer biology · Bioinform. 2019
Computer vision › Face, body and person analysis › face modeling
active appearance model
0.352008
Multi-View AAM Fitting and Construction · Int. J. Comput. Vis. 2008
Increasing the density of Active Appearance Models · CVPR 2008
Multi-View AAM Fitting and Camera Calibration · ICCV 2005
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312026
BioTriplex : a full-text annotated corpus for fine-tuning language models in gene-disease relation extraction tasks · Bioinform. 2026
Computer vision › 3D vision › motion estimation
optical flow
0.242011
A Database and Evaluation Methodology for Optical Flow · Int. J. Comput. Vis. 2011
A Database and Evaluation Methodology for Optical Flow · ICCV 2007
Three-Dimensional Scene Flow · ICCV 1999
Machine learning › Representation and self-supervised learning › representation learning › metric learning
local distance function
0.222011
Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Local distance functions: A taxonomy, new algorithms, and an evaluation · ICCV 2009
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.222011
Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Local distance functions: A taxonomy, new algorithms, and an evaluation · ICCV 2009
Image and video processing
super-resolution
0.242008
Simultaneous super-resolution and feature extraction for recognition of low-resolution faces · CVPR 2008
Resolution-Aware Fitting of Active Appearance Models to Low Resolution Images · ECCV (2) 2006
Limits on Super-Resolution and How to Break Them · IEEE Trans. Pattern Anal. Mach. Intell. 2002
Natural language and speech › Information extraction and text analysis
lexical semantics
0.212014
An Unsupervised Model for Instance Level Subcategorization Acquisition · EMNLP 2014
Natural language and speech › Information extraction and text analysis › lexical resources › lexical resource construction
subcategorization frame acquisition
0.212014
An Unsupervised Model for Instance Level Subcategorization Acquisition · EMNLP 2014
Natural language and speech › Information extraction and text analysis › lexical semantics
verb similarity
0.212014
An Unsupervised Model for Instance Level Subcategorization Acquisition · EMNLP 2014
Information retrieval › search interfaces
search result presentation
0.212014
Mining text snippets for images on the web · KDD 2014
Multimedia analysis and retrieval › cross-modal alignment
text-image alignment
0.212014
Mining text snippets for images on the web · KDD 2014
Computer vision › Image recognition and object detection
object recognition
0.232011
Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Local distance functions: A taxonomy, new algorithms, and an evaluation · ICCV 2009
Pattern Rejection · CVPR 1996
Computer vision › 3D vision › 3d reconstruction
shape from silhouette
0.232007
Coplanar Shadowgrams for Acquiring Visual Hulls of Intricate Objects · ICCV 2007
Shape-From-Silhouette Across Time Part I: Theory and Algorithms · Int. J. Comput. Vis. 2004
Shape-From-Silhouette of Articulated Objects and its Use for Human Body Kinematics Estimation and Motion Capture · CVPR (1) 2003
Computer vision › Face, body and person analysis
face modeling
0.132008
2D vs. 3D Deformable Face Models: Representational Power, Construction, and Real-Time Fitting · Int. J. Comput. Vis. 2007
Active Appearance Models Revisited · Int. J. Comput. Vis. 2004
Multi-View AAM Fitting and Construction · Int. J. Comput. Vis. 2008
Computer vision › Face, body and person analysis
face tracking
0.122007
The Asymmetry of Image Registration and Its Application to Face Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Leveragingarchivalvideo for building face datasets · ICCV 2007
Computer vision › Video understanding and tracking › action recognition › human action recognition
human interaction recognition
0.112012
Recognizing proxemics in personal photos · CVPR 2012
Machine learning › Efficient and distributed learning
resource-constrained learning
0.112012
Memory constrained face recognition · CVPR 2012
Image and video processing
image restoration
0.122009
Filter flow · ICCV 2009
Limits on Super-Resolution and How to Break Them · IEEE Trans. Pattern Anal. Mach. Intell. 2002
Computer vision › 3D vision › motion estimation › optical flow
optical flow evaluation
0.112011
A Database and Evaluation Methodology for Optical Flow · Int. J. Comput. Vis. 2011

Methods — techniques the papers use, named apart from their topics

large language model fine-tuning · 2.0LLaMA 3.1 · 2.0text mining · 1.3taxonomy construction · 0.6supervised machine learning · 0.5style-aware learning · 0.5information retrieval · 0.5hierarchical curriculum learning · 0.5denoising · 0.5biomedical text mining · 0.5active appearance model · 0.4ontology grounding · 0.4natural language processing · 0.4named entity recognition · 0.4co-occurrence metrics · 0.4machine-learned relevance scoring · 0.4combinatorial subset selection · 0.4interpolation · 0.2
YearPublicationVenuePosition
2026 BioTriplex : a full-text annotated corpus for fine-tuning language models in gene-disease relation extraction tasks
abstract
MOTIVATION: Automatic information extraction from biomedical texts requires machine learning methodology that can recognize biomedical entities, characterize inter-entity relationships, and relate extracted information to specific research topics. Large language models (LLMs) excel in general tasks but perform less reliably in the biomedical domain, where texts are characterized by extensive technical terminology and semantic variations from general literature. There is an unmet need for annotated full-text datasets that can be used to fine-tune language models for significant biomedical applications. Here, we focus on extraction of the complex relationships between genes and diseases. RESULTS: We present BioTriplex, a corpus of 100 full-length biomedical research articles (comprising 604 subsection texts) manually annotated with disease names, genes, and 21 subtypes of disease-gene relationships. We employ BioTriplex to train the LLaMA 3.1 8B language model in gene-disease relation extraction. Our fine-tuned model outperforms zero-shot and few-shot approaches, both within the LLaMA 3.1 architecture and across the larger state-of-the-art LLMs GPT-4 and Claude Sonnet 3.7, and classifies gene-disease relation types with broader scope and greater granularity than previously described. These results validate BioTriplex as a useful full-text data resource and underscore the value of specialized datasets in fine-tuning language models for important biomedical tasks. AVAILABILITY AND IMPLEMENTATION: https://github.com/PanagiotisFytas/BioTriplex.
Charlotte Collins, Panagiotis Fytas, Ilknur Karadeniz, Huiyuan Zheng, Simon Baker, Ulla Stenius, Anna Korhonen
Bioinform.5
2024 Text mining for contexts and relationships in cancer genomics literature
abstract
MOTIVATION: Scientific advances build on the findings of existing research. The 2001 publication of the human genome has led to the production of huge volumes of literature exploring the context-specific functions and interactions of genes. Technology is needed to perform large-scale text mining of research papers to extract the reported actions of genes in specific experimental contexts and cell states, such as cancer, thereby facilitating the design of new therapeutic strategies. RESULTS: We present a new corpus and Text Mining methodology that can accurately identify and extract the most important details of cancer genomics experiments from biomedical texts. We build a Named Entity Recognition model that accurately extracts relevant experiment details from PubMed abstract text, and a second model that identifies the relationships between them. This system outperforms earlier models and enables the analysis of gene function in diverse and dynamically evolving experimental contexts. AVAILABILITY AND IMPLEMENTATION: Code and data are available here: https://github.com/cambridgeltl/functional-genomics-ie.
Charlotte Collins, Simon Baker, Huiyuan Zheng, Adelyne Chan, Ulla Stenius, Masashi Narita, Anna Korhonen
Bioinform.2
2021 Dialogue Response Selection with Hierarchical Curriculum Learning
abstract
Yixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060
ACL/IJCNLP (1)5
2021 Non-Autoregressive Text Generation with Pre-trained Language Models
abstract
Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, Nigel Collier. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Yixuan Su, Deng Cai 0002, Yan Wang 0060, David Vandyke, Simon Baker, Piji Li, Nigel Collier
EACL5
2021 PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval Memory
abstract
The ability of dialogue systems to express pre-specified style during conversations has a direct, positive impact on their usability and user satisfaction. While it has attracted much research interest, existing methods often generate stylistic responses at the cost of content quality. In this work, we introduce a prototype-to-style (PS) framework to tackle the challenge of stylistic dialogue generation. The proposed framework first exploits an Information Retrieval (IR) system and extracts a response prototype from the retrieved response. A stylistic response generator then takes the response prototype and the desired style as input to produce a high-quality and stylistic response. To effectively train the proposed model and imitate the real testing environment, we introduce a new style-aware learning objective and a denoising learning strategy. Results on three benchmark datasets (gender, emotion, and sentiment) from two languages demonstrate that the proposed approach significantly outperforms existing baselines both in terms of in-domain and cross-domain evaluations.
Yixuan Su, Yan Wang 0060, Deng Cai 0002, Simon Baker, Anna Korhonen, Nigel Collier
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity
abstract
We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each language data set is annotated for the lexical relation of semantic similarity and contains 1,888 semantically aligned concept pairs, providing a representative coverage of word classes (nouns, verbs, adjectives, adverbs), frequency ranks, similarity intervals, lexical fields, and concreteness levels. Additionally, owing to the alignment of concepts across languages, we provide a suite of 66 crosslingual semantic similarity data sets. Because of its extensive size and language coverage, Multi-SimLex provides entirely novel opportunities for experimental evaluation and analysis. On its monolingual and crosslingual benchmarks, we evaluate and analyze a wide array of recent state-of-the-art monolingual and crosslingual representation models, including static and contextualized word embeddings (such as fastText, monolingual and multilingual BERT, XLM), externally informed lexical representations, as well as fully unsupervised and (weakly) supervised crosslingual word embeddings. We also present a step-by-step data set creation protocol for creating consistent, Multi-Simlex–style resources for additional languages. We make these contributions—the public release of Multi-SimLex data sets, their creation protocol, strong baseline results, and in-depth analyses which can be helpful in guiding future developments in multilingual lexical semantics and representation learning—available via a Web site that will encourage community effort in further expansion of Multi-Simlex to many more languages. Such a large-scale semantic resource could inspire significant further advances in NLP across languages.
Ivan Vulic, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, Anna Korhonen
Comput. Linguistics2
2020 A systematic literature review of automatic Alzheimer's disease detection from speech and language
abstract
OBJECTIVE: In recent years numerous studies have achieved promising results in Alzheimer's Disease (AD) detection using automatic language processing. We systematically review these articles to understand the effectiveness of this approach, identify any issues and report the main findings that can guide further research. MATERIALS AND METHODS: We searched PubMed, Ovid, and Web of Science for articles published in English between 2013 and 2019. We performed a systematic literature review to answer 5 key questions: (1) What were the characteristics of participant groups? (2) What language data were collected? (3) What features of speech and language were the most informative? (4) What methods were used to classify between groups? (5) What classification performance was achieved? RESULTS AND DISCUSSION: We identified 33 eligible studies and 5 main findings: participants' demographic variables (especially age ) were often unbalanced between AD and control group; spontaneous speech data were collected most often; informative language features were related to word retrieval and semantic, syntactic, and acoustic impairment; neural nets, support vector machines, and decision trees performed well in AD detection, and support vector machines and decision trees performed well in decline detection; and average classification accuracy was 89% in AD and 82% in mild cognitive impairment detection versus healthy control groups. CONCLUSION: The systematic literature review supported the argument that language and speech could successfully be used to detect dementia automatically. Future studies should aim for larger and more balanced datasets, combine data collection methods and the type of information analyzed, focus on the early stages of the disease, and report performance using standardized metrics.
Ulla Petti, Simon Baker, Anna Korhonen
J. Am. Medical Informatics Assoc.2
2019 LION LBD: a literature-based discovery system for cancer biology
abstract
MOTIVATION: The overwhelming size and rapid growth of the biomedical literature make it impossible for scientists to read all studies related to their work, potentially leading to missed connections and wasted time and resources. Literature-based discovery (LBD) aims to alleviate these issues by identifying implicit links between disjoint parts of the literature. While LBD has been studied in depth since its introduction three decades ago, there has been limited work making use of recent advances in biomedical text processing methods in LBD. RESULTS: We present LION LBD, a literature-based discovery system that enables researchers to navigate published information and supports hypothesis generation and testing. The system is built with a particular focus on the molecular biology of cancer using state-of-the-art machine learning and natural language processing methods, including named entity recognition and grounding to domain ontologies covering a wide range of entity types and a novel approach to detecting references to the hallmarks of cancer in text. LION LBD implements a broad selection of co-occurrence based metrics for analyzing the strength of entity associations, and its design allows real-time search to discover indirect associations between entities in a database of tens of millions of publications while preserving the ability of users to explore each mention in its original context in the literature. Evaluations of the system demonstrate its ability to identify undiscovered links and rank relevant concepts highly among potential connections. AVAILABILITY AND IMPLEMENTATION: The LION LBD system is available via a web-based user interface and a programmable API, and all components of the system are made available under open licenses from the project home page http://lbd.lionproject.net. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sampo Pyysalo, Simon Baker, Stefan Haselwimmer, Tejas Shah, Johan Högberg, Ulla Stenius, Masashi Narita, Anna Korhonen
Bioinform.2
2018 Variable Typing: Assigning Meaning to Variables in Mathematical Text
abstract
Yiannos Stathopoulos, Simon Baker, Marek Rei, Simone Teufel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yiannos Stathopoulos, Simon Baker, Marek Rei, Simone Teufel
NAACL-HLT2
2018 Towards improving decision making and estimating the value of decisions in value-based software engineering: the VALUE framework
abstract
To sustain growth, maintain competitive advantage and to innovate, companies must make a paradigm shift in which both short- and long-term value aspects are employed to guide their decision-making. Such need is clearly pressing in innovative industries, such as ICT, and is also the core of Value-based Software Engineering (VBSE). The goal of this paper is to detail a framework called VALUE—improving decision-making relating to software-intensive products and services development—and to show its application in practice to a large ICT company in Finland. The VALUE framework includes a mixed-methods approach, as follows: to elicit key stakeholders’ tacit knowledge regarding factors used during a decision-making process, either transcripts from interviews with key stakeholders are analysed and validated in focus group meetings or focus-group meeting(s) are directly applied. These value factors are later used as input to a Web-based tool (Value tool) employed to support decision making. This tool was co-created with four industrial partners in this research via a design science approach that includes several case studies and focus-group meetings. Later, data on key stakeholders’ decisions gathered using the Value tool, plus additional input from key stakeholders, are used, in combination with the Expert-based Knowledge Engineering of Bayesian Network (EKEBN) process, coupled with the weighed sum algorithm (WSA) method, to build and validate a company-specific value estimation model. The application of our proposed framework to a real case, as part of an ongoing collaboration with a large software company (company A), is presented herein. Further, we also provide a detailed example, partially using real data on decisions, of a value estimation Bayesian network (BN) model for company A. This paper presents some empirical results from applying the VALUE Framework to a large ICT company; those relate to eliciting key stakeholders’ tacit knowledge, which is later used as input to a pilot study where these stakeholders employ the Value tool to select features for one of their company’s chief products. The data on decisions obtained from this pilot study is later applied to a detailed example on building a value estimation BN model for company A. We detail a framework—VALUE framework—to be used to help companies improve their value-based decisions and to go a step further and also estimate the overall value of each decision.
Emilia Mendes, Pilar Rodríguez 0002, Vitor Freitas, Simon Baker, Mohamed Amine Atoui
Softw. Qual. J.4
2018 Correction to: Towards improving decision making and estimating the value of decisions in value-based software engineering: the VALUE framework
Emilia Mendes, Pilar Rodríguez 0002, Vitor Freitas, Simon Baker, Mohamed Amine Atoui
Softw. Qual. J.4
2017 Cancer Hallmarks Analytics Tool (CHAT): a text mining approach to organize and evaluate scientific literature on cancer
abstract
MOTIVATION: To understand the molecular mechanisms involved in cancer development, significant efforts are being invested in cancer research. This has resulted in millions of scientific articles. An efficient and thorough review of the existing literature is crucially important to drive new research. This time-demanding task can be supported by emerging computational approaches based on text mining which offer a great opportunity to organize and retrieve the desired information efficiently from sizable databases. One way to organize existing knowledge on cancer is to utilize the widely accepted framework of the Hallmarks of Cancer. These hallmarks refer to the alterations in cell behaviour that characterize the cancer cell. RESULTS: We created an extensive Hallmarks of Cancer taxonomy and developed automatic text mining methodology and a tool (CHAT) capable of retrieving and organizing millions of cancer-related references from PubMed into the taxonomy. The efficiency and accuracy of the tool was evaluated intrinsically as well as extrinsically by case studies. The correlations identified by the tool show that it offers a great potential to organize and correctly classify cancer-related literature. Furthermore, the tool can be useful, for example, in identifying hallmarks associated with extrinsic factors, biomarkers and therapeutics targets. AVAILABILITY AND IMPLEMENTATION: CHAT can be accessed at: http://chat.lionproject.net. The corpus of hallmark-annotated PubMed abstracts and the software are available at: http://chat.lionproject.net/about. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Simon Baker, Ilona Silins, Sampo Pyysalo, Johan Högberg, Ulla Stenius, Anna Korhonen
Bioinform.1
2016 Robust Text Classification for Sparsely Labelled Data Using Multi-level Embeddings
abstract
The conventional solution for handling sparsely labelled data is extensive feature engineering. This is time consuming and task and domain specific. We present a novel approach for learning embedded features that aims to alleviate this problem. Our approach jointly learns embeddings at different levels of granularity (word, sentence and document) along with the class labels. The intuition is that topic semantics represented by embeddings at multiple levels results in better classification. We evaluate this approach in unsupervised and semi-supervised settings on two sparsely labelled classification tasks, outperforming the handcrafted models and several embedding baselines.
Simon Baker, Douwe Kiela, Anna Korhonen
COLING1
2016 Automatic semantic classification of scientific literature according to the hallmarks of cancer
abstract
MOTIVATION: The hallmarks of cancer have become highly influential in cancer research. They reduce the complexity of cancer into 10 principles (e.g. resisting cell death and sustaining proliferative signaling) that explain the biological capabilities acquired during the development of human tumors. Since new research depends crucially on existing knowledge, technology for semantic classification of scientific literature according to the hallmarks of cancer could greatly support literature review, knowledge discovery and applications in cancer research. RESULTS: We present the first step toward the development of such technology. We introduce a corpus of 1499 PubMed abstracts annotated according to the scientific evidence they provide for the 10 currently known hallmarks of cancer. We use this corpus to train a system that classifies PubMed literature according to the hallmarks. The system uses supervised machine learning and rich features largely based on biomedical text mining. We report good performance in both intrinsic and extrinsic evaluations, demonstrating both the accuracy of the methodology and its potential in supporting practical cancer research. We discuss how this approach could be developed and applied further in the future. AVAILABILITY AND IMPLEMENTATION: The corpus of hallmark-annotated PubMed abstracts and the software for classification are available at: http://www.cl.cam.ac.uk/∼sb895/HoC.html. CONTACT: [email protected].
Simon Baker, Ilona Silins, Johan Högberg, Ulla Stenius, Anna Korhonen
Bioinform.1
2014 An Unsupervised Model for Instance Level Subcategorization Acquisition
abstract
Most existing systems for subcategorization frame (SCF) acquisition rely on supervised parsing and infer SCF distributions at type, rather than instance level.These systems suffer from poor portability across domains and their benefit for NLP tasks that involve sentence-level processing is limited.We propose a new unsupervised, Markov Random Field-based model for SCF acquisition which is designed to address these problems.The system relies on supervised POS tagging rather than parsing, and is capable of learning SCFs at instance level.We perform evaluation against gold standard data which shows that our system outperforms several supervised and type-level SCF baselines.We also conduct task-based evaluation in the context of verb similarity prediction, demonstrating that a vector space model based on our SCFs substantially outperforms a lexical model and a model based on a supervised parser 1 .
Simon Baker, Roi Reichart, Anna Korhonen
EMNLP1
2014 Mining text snippets for images on the web
abstract
Images are often used to convey many different concepts or illustrate many different stories. We propose an algorithm to mine multiple diverse, relevant, and interesting text snippets for images on the web. Our algorithm scales to all images on the web. For each image, all webpages that contain it are considered. The top-K text snippet selection problem is posed as combinatorial subset selection with the goal of choosing an optimal set of snippets that maximizes a combination of relevancy, interestingness, and diversity. The relevancy and interestingness are scored by machine learned models. Our algorithm is run at scale on the entire image index of a major search engine resulting in the construction of a database of images with their corresponding text snippets. We validate the quality of the database through a large-scale comparative study. We showcase the utility of the database through two web-scale applications: (a) augmentation of images on the web as webpages are browsed and (b)~an image browsing experience (similar in spirit to web browsing) that is enabled by interconnecting semantically related images (which may not be visually related) through shared concepts in their corresponding text snippets.
Anitha Kannan, Simon Baker, Krishnan Ramnath, Juliet Fiss, Dahua Lin, Lucy Vanderwende, Rizwan Ansary, Ashish Kapoor, Qifa Ke, Matthew Uyttendaele, Xin-Jing Wang, Lei Zhang 0001
KDD2
2014 Car make and model recognition using 3D curve alignment
abstract
We present a new approach for recognizing the make and model of a car from a single image. While most previous methods are restricted to fixed or limited viewpoints, our system is able to verify a car's make and model from an arbitrary view. Our model consists of 3D space curves obtained by backprojecting image curves onto silhouette-based visual hulls and then refining them using three-view curve matching. We also build an appearance model of taillights which is used as an additional cue. Our approach is able to verify the exact make and model of a car over a wide range of viewpoints and background clutter.
Edward Hsiao, Sudipta N. Sinha, Krishnan Ramnath, Simon Baker, C. Lawrence Zitnick, Richard Szeliski
WACV4
2014 AutoCaption: Automatic caption generation for personal photos
abstract
AutoCaption is a system that helps a smartphone user generate a caption for their photos. It operates by uploading the photo to a cloud service where a number of parallel modules are applied to recognize a variety of entities and relations. The outputs of the modules are combined to generate a large set of candidate captions, which are returned to the phone. The phone client includes a convenient user interface that allows users to select their favorite caption, reorder, add, or delete words to obtain the grammatical style they prefer. The user can also select from multiple candidates returned by the recognition modules.
Krishnan Ramnath, Simon Baker, Lucy Vanderwende, Motaz Ahmad El-Saban, Sudipta N. Sinha, Anitha Kannan, Noran Hassan, Michel Galley, Yi Yang 0007, Deva Ramanan, Alessandro Bergamo, Lorenzo Torresani
WACV2
2012 Memory constrained face recognition
abstract
Real-time recognition may be limited by scarce memory and computing resources for performing classification. Although, prior research has addressed the problem of training classifiers with limited data and computation, few efforts have tackled the problem of memory constraints on recognition. We explore methods that can guide the allocation of limited storage resources for classifying streaming data so as to maximize discriminatory power. We focus on computation of the expected value of information with nearest neighbor classifiers for online face recognition. Experiments on real-world datasets show the effectiveness and power of the approach. The methods provide a principled approach to vision under bounded resources, and have immediate application to enhancing recognition capabilities in consumer devices with limited memory.
Ashish Kapoor, Simon Baker, Sumit Basu, Eric Horvitz
CVPR2
2012 Recognizing proxemics in personal photos
abstract
Proxemics is the study of how people interact. We present a computational formulation of visual proxemics by attempting to label each pair of people in an image with a subset of physically based “touch codes.” A baseline approach would be to first perform pose estimation and then detect the touch codes based on the estimated joint locations. We found that this sequential approach does not perform well because pose estimation step is too unreliable for images of interacting people, due to difficulties with occlusion and limb ambiguities. Instead, we propose a direct approach where we build an articulated model tuned for each touch code. Each such model contains two people, connected in an appropriate manner for the touch code in question. We fit this model to the image and then base classification on the fitting error. Experiments show that this approach significantly outperforms the sequential baseline as well as other related approches.
Yi Yang 0007, Simon Baker, Anitha Kannan, Deva Ramanan
CVPR2
2011 A Database and Evaluation Methodology for Optical Flow
abstract
The quantitative evaluation of optical flow algorithms by Barron et al. (1994) led to significant advances in performance. The challenges for optical flow algorithms today go beyond the datasets and evaluation methods proposed in that paper. Instead, they center on problems associated with complex natural scenes, including nonrigid motion, real sensor noise, and motion discontinuities. We propose a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: (1) sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture, (2) realistic synthetic sequences, (3) high frame-rate video used to study interpolation error, and (4) modified stereo sequences of static scenes. In addition to the average angular error used by Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and results at motion discontinuities and in textureless regions. In October 2007, we published the performance of several well-known methods on a preliminary version of our data to establish the current state of the art. We also made the data freely available on the web at http://vision.middlebury.edu/flow/ . Subsequently a number of researchers have uploaded their results to our website and published papers using the data. A significant improvement in performance has already been achieved. In this paper we analyze the results obtained to date and draw a large number of conclusions from them.
Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski
Int. J. Comput. Vis.1
2011 Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation
abstract
We present a taxonomy for local distance functions where most existing algorithms can be regarded as approximations of the geodesic distance defined by a metric tensor. We categorize existing algorithms by how, where, and when they estimate the metric tensor. We also extend the taxonomy along each axis. How: We introduce hybrid algorithms that use a combination of techniques to ameliorate overfitting. Where: We present an exact polynomial-time algorithm to integrate the metric tensor along the lines between the test and training points under the assumption that the metric tensor is piecewise constant. When: We propose an interpolation algorithm where the metric tensor is sampled at a number of references points during the offline phase. The reference points are then interpolated during the online classification phase. We also present a comprehensive evaluation on tasks in face recognition, object recognition, and digit recognition.
Deva Ramanan, Simon Baker
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Removing rolling shutter wobble
abstract
We present an algorithm to remove wobble artifacts from a video captured with a rolling shutter camera undergoing large accelerations or jitter. We show how estimating the rapid motion of the camera can be posed as a temporal super-resolution problem. The low-frequency measurements are the motions of pixels from one frame to the next. These measurements are modeled as temporal integrals of the underlying high-frequency jitter of the camera. The estimated high-frequency motion of the camera is then used to re-render the sequence as though all the pixels in each frame were imaged at the same time. We also present an auto-calibration algorithm that can estimate the time between the capture of subsequent rows in the camera.
Simon Baker, Eric P. Bennett, Sing Bing Kang, Richard Szeliski
CVPR1
2010 Joint People, Event, and Location Recognition in Personal Photo Collections Using Cross-Domain Context
Dahua Lin, Ashish Kapoor, Gang Hua 0001, Simon Baker
ECCV (1)4
2010 Evaluating the Weighted Sum Algorithm for Estimating Conditional Probabilities in Bayesian Networks
Simon Baker, Emilia Mendes
SEKE1
2010 Multi-PIE
Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker
Image Vis. Comput.5
2010 Internet Vision
abstract
The ten papers in this special issue focus on Internet vision. The goal is to provide the reader with a general sense of the research being conducted in Internet vision, defined as the intersection of computer vision and the Internet. This issue includes coverage of a number of significant advances in this field.
Shai Avidan, Simon Baker, Ying Shan
Proc. IEEE2
2009 Clustering Videos by Location
abstract
We propose an algorithm to cluster video shots by the location in which they were captured. Each shot is represented as a set of keyframes and each keyframe is represented by a histogram of textons. Clustering is performed using an energy-based formulation. We propose an energy function for the clusters that matches the expected distribution of viewpoints in any one location and use the chi-squared distance to measure the similarity of two shots. We also add a temporal prior to model the fact that temporally neighboring shots are more likely to have been captured in the same location. We test our algorithm on both home videos and professionally edited footage (sitcoms). Quantitative results are presented to justify each choice made in the design of our algorithm, as well as comparisons with k-means, connected components, and spectral clustering. 1
Florian Schroff, C. Lawrence Zitnick, Simon Baker
BMVC3
2009 Which faces to tag: Adding prior constraints into active learning
abstract
We introduce an algorithm that guides the user to tag faces in the best possible order during a face recognition assisted tagging scenario. In particular, we extend the active learning paradigm to take advantage of constraints known a priori. For example, in the context of personal photo collections, if two faces come from the same source photograph, we know that they must be of different people. Similarly, in the context of video, we know that the faces from a single track must be of the same person. Given a set of unlabeled images and constraints, we use a probabilistic discriminative model that models the posterior distributions by propagating label information using a message passing scheme. The uncertainty estimate provided by the model naturally allows for active learning paradigms where the user is consulted after each iteration to tag additional faces. Our experiments show that performing active learning while incorporating a priori constraints provides a significant boost in many real-world face recognition tasks.
Ashish Kapoor, Gang Hua 0001, Amir Akbarzadeh, Simon Baker
ICCV4
2009 Local distance functions: A taxonomy, new algorithms, and an evaluation
abstract
We present a taxonomy for local distance functions where most existing algorithms can be regarded as approximations of the geodesic distance defined by a metric tensor. We categorize existing algorithms by how, where and when they estimate the metric tensor. We also extend the taxonomy along each axis. How: We introduce hybrid algorithms that use a combination of dimensionality reduction and metric learning to ameliorate over-fitting. Where: We present an exact polynomial time algorithm to integrate the metric tensor along the lines between the test and training points under the assumption that the metric tensor is piecewise constant. When: We propose an interpolation algorithm where the metric tensor is sampled at a number of references points during the offline phase, which are then interpolated during online classification. We also present a comprehensive evaluation of all the algorithms on tasks in face recognition, object recognition, and digit recognition.
Deva Ramanan, Simon Baker
ICCV2
2009 Filter flow
abstract
The filter flow problem is to compute a space-variant linear filter that transforms one image into another. This framework encompasses a broad range of transformations including stereo, optical flow, lighting changes, blur, and combinations of these effects. Parametric models such as affine motion, vignetting, and radial distortion can also be modeled within the same framework. All such transformations are modeled by selecting a number of constraints and objectives on the filter entries from a catalog which we enumerate. Most of the constraints are linear, leading to globally optimal solutions (via linear programming) for affine transformations, depth-from-defocus, and other problems. Adding a (non-convex) compactness objective enables solutions for optical flow with illumination changes, space-variant defocus, and higher-order smoothness.
Steven M. Seitz, Simon Baker
ICCV2
2009 Robust low-resolution face identification and verification using high-resolution features
abstract
In this work, we elaborate on a rather intuitive hypothesis: face recognition of low-resolution faces can be improved if the processes of reconstruction and recognition are considered simultaneously, instead of sequentially, without feedback or any interaction. Given a high-resolution training set, matching low-resolution probe images with good accuracy is an open problem. We have recently introduced [Hennings-Yeomans, Baker, and Kumar, CVPR, June 2008] a new framework for low-resolution face recognition that uses models from an image formation process, super-resolution priors and face feature extraction methods. By measuring how well an intermediate super-resolution reconstruction of the probe image fits into the models used in the process, the proposed matching algorithm extracts new features for recognition. In this paper, we present results for an improved design of these new features. We show that the proposed algorithm improves performance in both, identification and verification tasks on a large database of 337 subjects that also captures illumination variations.
Pablo H. Hennings-Yeomans, B. V. K. Vijaya Kumar, Simon Baker
ICIP3
2009 The Theory and Practice of Coplanar Shadowgram Imaging for Acquiring Visual Hulls of Intricate Objects
Shuntaro Yamazaki, Srinivasa G. Narasimhan, Simon Baker, Takeo Kanade
Int. J. Comput. Vis.3
2009 An MRF-Based DeInterlacing Algorithm With Exemplar-Based Refinement
abstract
In this paper, we propose an MRF-based deinterlacing algorithm that combines the benefits of rule-based algorithms such as motion-adaptation, edge-directed interpolation, and motion compensation, with those of an MRF formulation. MRF-based interpolation and enhancement algorithms are typically formulated as an optimization over pixel intensities or colors, which can make them relatively slow. In comparison, our MRF-based deinterlacing algorithm uses interpolation functions as labels.We use seven interpolants (three spatial, three temporal, and one for motion compensation). The core dynamic programming algorithm is, therefore, sped up greatly over the direct use of intensity as labels. We also show how an exemplar-based learning algorithm can be used to refine the output of our MRF-based algorithm. The training set can be augmented with exemplars from static regions of the same video, as a form of "self-learning."
Shengyang Dai, Simon Baker, Sing Bing Kang
IEEE Trans. Image Process.2
2008 Model-Based De-Identification of Facial Images
Ralph Gross, Latanya Sweeney, Jeffrey F. Cohn, Fernando De la Torre, Simon Baker
AMIA5
2008 Semi-supervised learning of multi-factor models for face de-identification
abstract
With the emergence of new applications centered around the sharing of image data, questions concerning the protection of the privacy of people visible in the scene arise. Recently, formal methods for the de-identification of images have been proposed which would benefit from multi-factor coding to separate identity and non-identity related factors. However, existing multi-factor models require complete labels during training which are often not available in practice. In this paper we propose a new multi-factor framework which unifies linear, bilinear, and quadratic models. We describe a new fitting algorithm which jointly estimates all model parameters and show that it outperforms the standard alternating algorithm. We furthermore describe how to avoid overfitting the model and how to train the model in a semi-supervised manner. In experiments on a large expression-variant face database we show that data coded using our multi-factor model leads to improved data utility while providing the same privacy protection.
Ralph Gross, Latanya Sweeney, Fernando De la Torre, Simon Baker
CVPR4
2008 Simultaneous super-resolution and feature extraction for recognition of low-resolution faces
abstract
Face recognition degrades when faces are of very low resolution since many details about the difference between one person and another can only be captured in images of sufficient resolution. In this work, we propose a new procedure for recognition of low-resolution faces, when there is a high-resolution training set available. Most previous super-resolution approaches are aimed at reconstruction, with recognition only as an after-thought. In contrast, in the proposed method, face features, as they would be extracted for a face recognition algorithm (e.g., eigenfaces, Fisher-faces, etc.), are included in a super-resolution method as prior information. This approach simultaneously provides measures of fit of the super-resolution result, from both reconstruction and recognition perspectives. This is different from the conventional paradigms of matching in a low-resolution domain, or, alternatively, applying a super-resolution algorithm to a low-resolution face and then classifying the super-resolution result. We show, for example, that recognition of faces of as low as 6 times 6 pixel size is considerably improved compared to matching using a super-resolution reconstruction followed by classification, and to matching with a low-resolution training set.
Pablo H. Hennings-Yeomans, Simon Baker, B. V. K. Vijaya Kumar
CVPR2
2008 Increasing the density of Active Appearance Models
abstract
Active appearance models (AAMs) typically only use 50-100 mesh vertices because they are usually constructed from a set of training images with the vertices hand-labeled on them. In this paper, we propose an algorithm to increase the density of an AAM. Our algorithm operates by iteratively building the AAM, refitting the AAM to the training data, and refining the AAM.We compare our algorithm with the state of the art in optical flow algorithms and find it to be significantly more accurate. We also show that dense AAMs can be fit more robustly than sparse ones. Finally, we show how our algorithm can be used to construct AAMs automatically, starting with a single affine model that is subsequently refined to model non-planarity and non-rigidity.
Krishnan Ramnath, Simon Baker, Iain A. Matthews, Deva Ramanan
CVPR2
2008 Multi-PIE
abstract
A close relationship exists between the advancement of face recognition algorithms and the availability of face databases varying factors that affect facial appearance in a controlled manner. The CMU PIE database has been very influential in advancing research in face recognition across pose and illumination. Despite its success the PIE database has several shortcomings: a limited number of subjects, a single recording session and only few expressions captured. To address these issues we collected the CMU Multi-PIE database. It contains 337 subjects, imaged under 15 view points and 19 illumination conditions in up to four recording sessions. In this paper we introduce the database and describe the recording procedure. We furthermore present results from baseline experiments using PCA and LDA classifiers to highlight similarities and differences between PIE and Multi-PIE.
Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker
FG5
2008 Multi-View AAM Fitting and Construction
Krishnan Ramnath, Seth Koterba, Jing Xiao 0006, Changbo Hu, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade
Int. J. Comput. Vis.6
2008 Guest Editors' Introduction to the Special Section on CVPR Papers
abstract
The four papers in this special section are extended versions of award-winning papers from the 2007 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2007).
Simon Baker, Jiri Matas, Ramin Zabih
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 A Database and Evaluation Methodology for Optical Flow
abstract
The quantitative evaluation of optical flow algorithms by Barron et al. led to significant advances in the performance of optical flow methods. The challenges for optical flow today go beyond the datasets and evaluation methods proposed in that paper and center on problems associated with nonrigid motion, real sensor noise, complex natural scenes, and motion discontinuities. Our goal is to establish a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture; realistic synthetic sequences; high frame-rate video used to study interpolation error; and modified stereo sequences of static scenes. In addition to the average angular error used in Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and flow accuracy at motion boundaries and in textureless regions. We evaluate the performance of several well-known methods on this data to establish the current state of the art. Our database is freely available on the web together with scripts for scoring and publication of the results at http://vision.middlebury.edu/flow/.
Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski
ICCV1
2007 Leveragingarchivalvideo for building face datasets
abstract
We introduce a semi-supervised method for building large, labeled datasets effaces by leveraging archival video. Specifically, we have implemented a system for labeling 11 years worth of archival footage from a television show. We have compiled a dataset of 611,770 faces, orders of magnitude larger than existing collections. It includes variation in appearance due to age, weight gain, changes in hairstyles, and other factors difficult to observe in smaller-scale collections. Face recognition in an uncontrolled setting can be difficult. We argue (and demonstrate) that there is much structure at varying timescales in the video data that make recognition much easier. At local time scales, one can use motion and tracking to group face images together - we may not know the identity, but we know a single label applies to all faces in a track. At medium time scales (say, within a scene), one can use appearance features such as hair and clothing to group tracks across shot boundaries. However, at longer timescales (say, across episodes), one can no longer use clothing as a cue. This suggests that one needs to carefully encode representations of appearance, depending on the timescale at which one intends to match. We assemble our final dataset by classifying groups of tracks in a nearest-neighbors framework. We use a face library obtained by labeling track clusters in a reference episode. We show that this classification is significantly easier when exploiting the hierarchical structure naturally present in the video sequences. From a data-collection point of view, tracking is vital because it adds non-frontal poses to our face collection. This is important because we know of no other method for collecting images of non-frontal faces "in the wild".
Deva Ramanan, Simon Baker, Sham M. Kakade
ICCV2
2007 Coplanar Shadowgrams for Acquiring Visual Hulls of Intricate Objects
abstract
Acquiring 3D models of intricate objects (like tree branches, bicycles and insects) is a hard problem due to severe self-occlusions, repeated thin structures and surface discontinuities. In theory, a shape-from-silhouettes (SFS) approach can overcome these difficulties and use many views to reconstruct visual hulls that are close to the actual shapes. In practice, however, SFS is highly sensitive to errors in silhouette contours and the calibration of the imaging system, and therefore not suitable for obtaining reliable shapes with a large number of views. We present a practical approach to SFS using a novel technique called coplanar shadowgram imaging, that allows us to use dozens to even hundreds of views for visual hull reconstruction. Here, a point light source is moved around an object and the shadows (silhouettes) cast onto a single background plane are observed. We characterize this imaging system in terms of image projection, reconstruction ambiguity, epipolar geometry, and shape and source recovery. The coplanarity of the shadowgrams yields novel geometric properties that are not possible in traditional multi-view camera- based imaging systems. These properties allow us to derive a robust and automatic algorithm to recover the visual hull of an object and the 3D positions of light source simultaneously, regardless of the complexity of the object. We demonstrate the acquisition of several intricate shapes with severe occlusions and thin structures, using 50 to 120 views.
Shuntaro Yamazaki, Srinivasa G. Narasimhan, Simon Baker, Takeo Kanade
ICCV3
2007 2D vs. 3D Deformable Face Models: Representational Power, Construction, and Real-Time Fitting
Iain A. Matthews, Jing Xiao 0006, Simon Baker
Int. J. Comput. Vis.3
2007 The Asymmetry of Image Registration and Its Application to Face Tracking
abstract
Most image registration problems are formulated in an asymmetric fashion. Given a pair of images, one is implicitly or explicitly regarded as a template and warped onto the other to match as well as possible. In this paper, we focus on this seemingly arbitrary choice of the roles and reveal how it may lead to biased warp estimates in the presence of relative scaling. We present a principled way of selecting the template and explain why only the correct asymmetric form, with the potential inclusion of a blurring step, can yield an unbiased estimator. We validate our analysis in the domain of model-based face tracking. We show how the usual Active Appearance Model (AAM) formulation overlooks the asymmetry issue, causing the fitting accuracy to degrade quickly when the observed objects are smaller than their model. We formulate a novel, "resolution-aware fitting" (RAF) algorithm that respects the asymmetry and incorporates an explicit model of the blur caused by the camera's sensing elements into the fitting formulation. We compare the RAF algorithm against a state-of-the-art tracker across a variety of resolutions and AAM complexity levels. Experimental results show that RAF significantly improves the estimation accuracy of both shape and appearance parameters when fitting to low-resolution data. Recognizing and accounting for the asymmetry of image registration leads to tangible accuracy improvements in analyzing low-resolution imagery.
Göksel Dedeoglu, Takeo Kanade, Simon Baker
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Resolution-Aware Fitting of Active Appearance Models to Low Resolution Images
Göksel Dedeoglu, Simon Baker, Takeo Kanade
ECCV (2)2
2006 Active appearance models with occlusion
Ralph Gross, Iain A. Matthews, Simon Baker
Image Vis. Comput.3
2005 Representational Oriented Component Analysis (ROCA) for Face Recognition with One Sample Image per Training Class
abstract
Subspace methods such as PCA, LDA, ICA have become a standard tool to perform visual learning and recognition. In this paper we propose representational oriented component analysis (ROCA), an extension of OCA, to perform face recognition when just one sample per training class is available. Several novelties are introduced in order to improve generalization and efficiency: (1) combining several OCA classifiers based on different image representations of the unique training sample is shown to greatly improve the recognition performance. (2) To improve generalization and to account for small misregistration effect, a learned subspace is added to constrain the OCA solution, (3) a stable/efficient generalized eigenvector algorithm that solves the small size sample problem and avoids overfitting. Preliminary experiments in the FRGC Ver 1.0 dataset show that ROCA outperforms existing linear techniques (PCA, OCA) and some commercial systems.
Fernando De la Torre, Ralph Gross, Simon Baker, B. V. K. Vijaya Kumar
CVPR (2)3
2005 Multi-View AAM Fitting and Camera Calibration
abstract
In this paper, we study the relationship between multi-view active appearance model (AAM) fitting and camera calibration. In the first part of the paper we propose an algorithm to calibrate the relative orientation of a set of N > 1 cameras by fitting an AAM to sets of N images. In essence, we use the human face as a (non-rigid) calibration grid. Our algorithm calibrates a set of 2 /spl times/ 3 weak-perspective camera projection matrices, protections of the world coordinate system origin into the images, depths of the world coordinate system origin, and focal lengths. We demonstrate that the performance of this algorithm is comparable to a standard algorithm using a calibration grid. In the second part of the paper, we show how calibrating the cameras improves tile performance of multi-view AAM fitting.
Seth Koterba, Simon Baker, Iain A. Matthews, Changbo Hu, Jing Xiao 0006, Jeffrey F. Cohn, Takeo Kanade
ICCV2
2005 Shape-From-Silhouette Across Time Part II: Applications to Human Modeling and Markerless Motion Tracking
German K. M. Cheung, Simon Baker, Takeo Kanade
Int. J. Comput. Vis.2
2005 Generic vs. person specific active appearance models
Ralph Gross, Iain A. Matthews, Simon Baker
Image Vis. Comput.3
2005 Three-Dimensional Scene Flow
abstract
Just as optical flow is the two-dimensional motion of points in an image, scene flow is the three-dimensional motion of points in the world. The fundamental difficulty with optical flow is that only the normal flow can be computed directly from the image measurements, without some form of smoothing or regularization. In this paper, we begin by showing that the same fundamental limitation applies to scene flow; however, many cameras are used to image the scene. There are then two choices when computing scene flow: 1) perform the regularization in the images or 2) perform the regularization on the surface of the object in the scene. In this paper, we choose to compute scene flow using regularization in the images. We describe three algorithms, the first two for computing scene flow from optical flows and the third for constraining scene tructure from the inconsistencies in multiple optical flows.
Sundar Vedula, Simon Baker, Peter Rander, Robert T. Collins, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Image-based spatio-temporal modeling and view interpolation of dynamic events
abstract
We present an approach for modeling and rendering a dynamic, real-world event from an arbitrary viewpoint, and at any time, using images captured from multiple video cameras. The event is modeled as a nonrigidly varying dynamic scene, captured by many images from different viewpoints, at discrete times. First, the spatio-temporal geometric properties (shape and instantaneous motion) are computed. The view synthesis problem is then solved using a reverse mapping algorithm, ray-casting across space and time, to compute a novel image from any viewpoint in the 4D space of position and time. Results are shown on real-world events captured in the CMU 3D Room, by creating synthetic renderings of the event from novel, arbitrary positions in space and time. Multiple such recreated renderings can be put together to create retimed fly-by movies of the event, with the resulting visual experience richer than that of a regular video clip, or switching between images from multiple cameras.
Sundar Vedula, Simon Baker, Takeo Kanade
ACM Trans. Graph.2
2004 Generic vs. Person Specific Active Appearance Models
abstract
Active Appearance Models (AAMs) are generative parametric models that have been successfully used in the past to model faces. Anecdotal evidence, however, suggests that the performance of an AAM built to model the variation in appearance of a single person across pose, illumination, and expression (a Person Specific AAM) is substantially better than the performance of an AAM built to model the variation in appearance of many faces, including unseen subjects not in the training set (a Generic AAM). In this paper, we present an empirical evaluation that shows that Person Specific AAMs are, as expected, both easier to build and more robust to fit than Generic AAMs. Moreover, we show that: (1) building a generic shape model is far easier than building a generic appearance model, and (2) the shape component is the main cause of the reduced fitting robustness of Generic AAMs. We then proceed to describe two refinements to Generic AAMs to improve their performance: (1) a refitting procedure to improve the quality of the ground-truth data used to build the AAM and (2) a new fitting algorithm. For both refinements we demonstrate dramatically improved fitting performance. Finally, we evaluate the effect of these improvements on a combined model construction and fitting task.
Ralph Gross, Iain A. Matthews, Simon Baker
BMVC3
2004 Fitting a Single Active Appearance Model Simultaneously to Multiple Images
abstract
Active Appearance Models (AAMs) are a well studied 2D deformable model. One recently proposed extension of AAMs to multiple images is the Coupled-View AAM. Coupled-View AAMs model the 2D shape and appearance of a face in two or more views simultaneously. The major limitation of Coupled-View AAMs, however, is that they are specific to a particular set of cameras, both in geometry and the photometric responses. In this paper, we describe how a single AAM can be fit to multiple images, captured simultaneously by cameras with arbitrary geometry and response functions. Our algorithm retains the major benefits of Coupled-View AAMs: the integration of information from multiple images into a single model, and improved fitting robustness. 1
Changbo Hu, Jing Xiao 0006, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade
BMVC4
2004 Real-Time Combined 2D+3D Active Appearance Models
Jing Xiao 0006, Simon Baker, Iain A. Matthews, Takeo Kanade
CVPR (2)2
2004 Lucas-Kanade 20 Years On: A Unifying Framework
Simon Baker, Iain A. Matthews
Int. J. Comput. Vis.1
2004 Shape-From-Silhouette Across Time Part I: Theory and Algorithms
German K. M. Cheung, Simon Baker, Takeo Kanade
Int. J. Comput. Vis.2
2004 Active Appearance Models Revisited
Iain A. Matthews, Simon Baker
Int. J. Comput. Vis.2
2004 Automatic Construction of Active Appearance Models as an Image Coding Problem
abstract
The automatic construction of Active Appearance Models (AAMs) is usually posed as finding the location of the base mesh vertices in the input training images. In this paper, we repose the problem as an energy-minimizing image coding problem and propose an efficient gradient-descent algorithm to solve it.
Simon Baker, Iain A. Matthews, Jeff G. Schneider
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Appearance-Based Face Recognition and Light-Fields
abstract
Arguably the most important decision to be made when developing an object recognition algorithm is selecting the scene measurements or features on which to base the algorithm. In appearance-based object recognition, the features are chosen to be the pixel intensity values in an image of the object. These pixel intensities correspond directly to the radiance of light emitted from the object along certain rays in space. The set of all such radiance values over all possible rays is known as the plenoptic function or light-field. In this paper, we develop a theory of appearance-based object recognition from light-fields. This theory leads directly to an algorithm for face recognition across pose that uses as many images of the face as are available, from one upwards. All of the pixels, whichever image they come from, are treated equally and used to estimate the (eigen) light-field of the object. The eigen light-field is then used as the set of features on which to base recognition, analogously to how the pixel intensities are used in appearance-based face and object recognition.
Ralph Gross, Iain A. Matthews, Simon Baker
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 The Template Update Problem
abstract
Template tracking dates back to the 1981 Lucas-Kanade algorithm. One question that has received very little attention, however, is how to update the template so that it remains a good model of the tracked object. We propose a template update algorithm that avoids the "drifting" inherent in the naive algorithm.
Iain A. Matthews, Takahiro Ishikawa, Simon Baker
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 The Template Update Problem
abstract
Template tracking dates back to the 1981 Lucas-Kanade algorithm. One question that has received very little attention, however, is how to update the template so that it remains a good model of the tracked object. We propose a template update algorithm that avoids the "drifting" inherent in the naive algorithm.
Iain A. Matthews, Takahiro Ishikawa, Simon Baker
BMVC3
2003 Shape-From-Silhouette of Articulated Objects and its Use for Human Body Kinematics Estimation and Motion Capture
abstract
Shape-from-silhouette (SFS), also known as visual hull (VH) construction, is a popular 3D reconstruction method, which estimates the shape of an object from multiple silhouette images. The original SFS formulation assumes that the entire silhouette images are captured either at the same time or while the object is static. This assumption is violated when the object moves or changes shape. Hence the use of SFS with moving objects has been restricted to treating each time instant sequentially and independently. Recently we have successfully extended the traditional SFS formulation to refine the shape of a rigidly moving object over time. We further extend SFS to apply to dynamic articulated objects. Given silhouettes of a moving articulated object, the process of recovering the shape and motion requires two steps: (1) correctly segmenting (points on the boundary of) the silhouettes to each articulated part of the object, (2) estimating the motion of each individual part using the segmented silhouette. In this paper, we propose an iterative algorithm to solve this simultaneous assignment and alignment problem. Once we have estimated the shape and motion of each part of the object, the articulation points between each pair of rigid parts are obtained by solving a simple motion constraint between the connected parts. To validate our algorithm, we first apply it to segment the different body parts and estimate the joint positions of a person. The acquired kinematic (shape and joint) information is then used to track the motion of the person in new video sequences.
German K. M. Cheung, Simon Baker, Takeo Kanade
CVPR (1)2
2003 Visual Hull Alignment and Refinement Across Time: A 3D Reconstruction Algorithm Combining Shape-From-Silhouette with Stereo
abstract
Visual hull (VH) construction from silhouette images is a popular method of shape estimation. The method, also known as shape-from-silhouette (SFS), is used in many applications such as non-invasive 3D model acquisition, obstacle avoidance, and more recently human motion tracking and analysis. One of the limitations of SFS, however, is that the approximated shape can be very coarse when there are only a few cameras. In this paper, we propose an algorithm to improve the shape approximation by combining multiple silhouette images captured across time. The improvement is achieved by first estimating the rigid motion between the visual hulls formed at different time instants (visual hull alignment) and then combining them (visual hull refinement) to get a tighter bound on the object's shape. Our algorithm first constructs a representation of the VHs called the bounding edge representation. Utilizing a fundamental property of visual hulls, which states that each bounding edge must touch the object at at least one point, we use multi-view stereo to extract points called colored surface points (CSP) on the surface of the object. These CSPs are then used in a 3D image alignment algorithm to find the 6 DOF rigid motion between two visual hulls. Once the rigid motion across time is known, all of the silhouette images are treated as being captured at the same time instant and the shape of the object is refined. We validate our algorithm on both synthetic and real data and compare it with space carving.
German K. M. Cheung, Simon Baker, Takeo Kanade
CVPR (2)2
2003 Tele-Graffiti: A Camera-Projector Based Remote Sketching System with Hand-Based User Interface and Automatic Session Summarization
Naoya Takao, Jianbo Shi, Simon Baker
Int. J. Comput. Vis.3
2003 When Is the Shape of a Scene Unique Given Its Light-Field: A Fundamental Theorem of 3D Vision?
abstract
The complete set of measurements that could ever be used by a passive 3D vision algorithm is the plenoptic function or light-field. We give a concise characterization of when the light-field of a Lambertian scene uniquely determines its shape and, conversely, when the shape is inherently ambiguous. In particular, we show that stereo computed from the light-field is ambiguous if and only if the scene is radiating light of a constant intensity (and color, etc.) over an extended region.
Simon Baker, Terence Sim, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 The CMU Pose, Illumination, and Expression Database
abstract
In the Fall of 2000, we collected a database of more than 40,000 facial images of 68 people. Using the Carnegie Mellon University 3D Room, we imaged each person across 13 different poses, under 43 different illumination conditions, and with four different expressions. We call this the CMU pose, illumination, and expression (PIE) database. We describe the imaging hardware, the collection procedure, the organization of the images, several possible uses, and how to obtain the database.
Terence Sim, Simon Baker, Maan Bsat
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Limits on Super-Resolution and How to Break Them
abstract
Nearly all super-resolution algorithms are based on the fundamental constraints that the super-resolution image should generate low resolution input images when appropriately warped and down-sampled to model the image formation process. (These reconstruction constraints are normally combined with some form of smoothness prior to regularize their solution.) We derive a sequence of analytical results which show that the reconstruction constraints provide less and less useful information as the magnification factor increases. We also validate these results empirically and show that, for large enough magnification factors, any smoothness prior leads to overly smooth results with very little high-frequency content. Next, we propose a super-resolution algorithm that uses a different kind of constraint in addition to the reconstruction constraints. The algorithm attempts to recognize local features in the low-resolution images and then enhances their resolution in an appropriate manner. We call such a super-resolution algorithm a hallucination or reconstruction algorithm. We tried our hallucination algorithm on two different data sets, frontal images of faces and printed Roman text. We obtained significantly better results than existing reconstruction-based algorithms, both qualitatively and in terms of RMS pixel error.
Simon Baker, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Equivalence and Efficiency of Image Alignment Algorithms
abstract
There are two major formulations of image alignment using gradient descent. The first estimates an additive increment to the parameters (the additive approach), the second an incremental warp (the compositional approach). We first prove that these two formulations are equivalent. A very efficient algorithm was proposed by Hager and Belhumeur (1998) using the additive approach that unfortunately can only be applied to a very restricted class of warps. We show that using the compositional approach an equally efficient algorithm (the inverse compositional algorithm) can be derived that can be applied to any set of warps which form a group. While most warps used in computer vision form groups, there are a certain warps that do not. Perhaps most notable is the set of piecewise affine warps used in flexible appearance models (FAMs). We end this paper by extending the inverse compositional algorithm to apply to FAMs.
Simon Baker, Iain A. Matthews
CVPR (1)1
2001 A Characterization of Inherent Stereo Ambiguities
Simon Baker, Terence Sim, Takeo Kanade
ICCV1
2001 Tele-Graffiti: A Pen and Paper-Based Remote Sketching System
abstract
Tele-Graffiti is a system allowing two.or more users to communicate remotely via hand-drawn sketches. What one person writes at one site is captured using a video camera, transmitted to the other site(s), and displayed there using an LCD projector. The advantage of our system over other intelligent desktops and white-boards is that the users are free to move the pieces of paper on which they are writing. In Tele-Graffiti, paper detection and tracking is based on real-time paper boundary detection.
Naoya Takao, Jianbo Shi, Simon Baker, Iain A. Matthews, Bart C. Nabbe
ICCV3
2000 Limits on Super-Resolution and How to Break Them
abstract
We analyze the super-resolution reconstruction constraints. In particular we derive a sequence of results which all show that the constraints provide far less useful information as the magnification factor increases. It is well established that the use of a smoothness prior may help somewhat, however for large enough magnification factors any smoothness prior leads to overly smooth results. We therefore propose an algorithm that learns recognition-based priors for specific classes of scenes, the use of which gives far better super-resolution results for both faces and text.
Simon Baker, Takeo Kanade
CVPR1
2000 Shape and Motion Carving in 6D
abstract
The motion of a non-rigid scene over time imposes more constraints on its structure than those derived from images at a single time instant alone. An algorithm is presented for simultaneously recovering dense scene shape and scene flow (i.e. the instantaneous 3D motion at every point in the scene). The algorithm operates by carving away hexels, or points in the 6D space of all possible shapes and flows that are inconsistent with the images captures at either time instant, or across time. The recovered shape is demonstrated to be more accurate than that recovered using images at a single time instant. Applications of the combined scene shape and flow include motion capture for animation, retiming of videos, and non-rigid motion analysis.
Sundar Vedula, Simon Baker, Steven M. Seitz, Takeo Kanade
CVPR2
2000 Hallucinating Faces
abstract
Faces often appear very small in surveillance imagery because of the wide fields of view that are typically used and the relatively large distance between the cameras and the scene. For tasks such as face recognition, resolution enhancement techniques are therefore generally needed. Although numerous resolution enhancement algorithms have been proposed in the literature, most of them are limited by the fact that they make weak, if any, assumptions about the scene. We propose an algorithm to learn a prior on the spatial distribution of the image gradient for frontal images of faces. We proceed to show how such a prior can be incorporated into a resolution enhancement algorithm to yield 4- to 8-fold improvements in resolution (i.e., 16 to 64 times as many pixels). The additional pixels are, in effect, hallucinated.
Simon Baker, Takeo Kanade
FG1
1999 Global Measures of Coherence for Edge Detector Evaluation
abstract
We propose a class of benchmarks for edge detector evaluation that require no ground truth. Each benchmark consists of a large number of images of a carefully designed scene for which we enforce a constraint on the edges, for example, that they are co-linear. We sample the space of edge appearances as densely as possible by capturing the images under widely varying imaging conditions. Not only do we change the viewing geometry and the illumination direction, but we also vary the camera parameters and the physical properties of the objects in the scene. We show that the degrees to which the constraints hold in the output edge-maps can be used as highly discriminating measures of edge detector performance. The code, images, and results which form our benchmarks are all available from the website http://www.cs.columbia.edu/CAVE/. The code and images enable a user to compare any new detector against several previous ones with minimal effort.
Simon Baker, Shree K. Nayar
CVPR1
1999 Three-Dimensional Scene Flow
abstract
Scene flow is the three-dimensional motion field of points in the world, just as optical flow is the two-dimensional motion field of points in an image. Any optical flow is simply the projection of the scene flow onto the image plane of a camera. We present a framework for the computation of dense, non-rigid scene flow from optical flow. Our approach leads to straightforward linear algorithms and a classification of the task into three major scenarios: complete instantaneous knowledge of the scene structure; knowledge only of correspondence information; and no knowledge of the scene structure. We also show that multiple estimates of the normal flow cannot be used to estimate dense scene flow directly without some form of smoothing or regularization.
Sundar Vedula, Simon Baker, Peter Rander, Robert T. Collins, Takeo Kanade
ICCV2
1999 A Theory of Single-Viewpoint Catadioptric Image Formation
Simon Baker, Shree K. Nayar
Int. J. Comput. Vis.1
1998 A Layered Approach to Stereo Reconstruction
abstract
We propose a framework for extracting structure from stereo which represents the scene as a collection of approximately planar layers. Each layer consists of an explicit 3D plane equation, a colored image with per-pixel opacity (a sprite), and a per-pixel depth offset relative to the plane. Initial estimates of the layers are recovered using techniques taken from parametric motion estimation. These initial estimates are then refined using a re-synthesis algorithm which takes into account both occlusions and mixed pixels. Reasoning about such effects allows the recovery of depth and color information with high accuracy even in partially occluded regions. Another important benefit of our framework is that the output consists of a collection of approximately planar regions, a representation which is far more appropriate than a dense depth map for many applications such as rendering and video parsing.
Simon Baker, Richard Szeliski, P. Anandan 0001
CVPR1
1998 A Theory of Catadioptric Image Formation
abstract
Conventional video cameras have limited fields of view which make them restrictive for certain applications in computational vision. A catadioptric sensor uses a combination of lenses and mirrors placed in a carefully arranged configuration to capture a much wider field of view. When designing a catadioptric sensor, the shape of the mirror(s) should ideally be selected to ensure that the complete catadioptric system has a single effective viewpoint. In this paper, we derive the complete class of single-lens single-mirror catadioptric sensors which have a single viewpoint and an expression for the spatial resolution of a catadioptric sensor in terms of the resolution of the camera used to construct it. We also include a preliminary analysis of the defocus blur caused by the use of a curved mirror.
Simon Baker, Shree K. Nayar
ICCV1
1998 Interactive 3D modeling from multiple images using scene regularities
abstract
Due to the complexity of real scenes and the fragility of fully automated vision techniques, results from many automated modeling systems are disappointing. Automated techniques often require manual clean-up and postprocessing to segment the scene into coherent objects and surfaces, or to triangulate sparse point matches. They may also be required to enforce geometric constraints such as known orientations of surfaces. For instance, building interiors and exteriors provide vertical and horizontal lines and parallel and perpendicular planes. In this paper, we attack the 3D modeling problem from the other side: we specify some geometric knowledge ahead of time (e.g., known orientations of lines, co-planarity of points, initial scene segmentations), and use these constraints to guide our matching and reconstruction algorithms. We present two interactive (semi-automated) systems for recovering 3D models of large-scale environments from multiple images.
Harry Shum, Richard Szeliski, Simon Baker, P. Anandan 0001
WACV3
1998 Parametric Feature Detection
Simon Baker, Shree K. Nayar, Hiroshi Murase
Int. J. Comput. Vis.1
1996 Pattern Rejection
abstract
The efficiency of pattern recognition is particularly crucial in two scenarios; whenever there are a large number of classes to discriminate, and, whenever recognition must be performed a large number of times. We propose a single technique, namely, pattern rejection, that greatly enhances efficiency in both cases. A rejector is a generalization of a classifier, that quickly eliminates a large fraction of the candidate classes or inputs. This allows a recognition algorithm to dedicate its efforts to a much smaller number of possibilities. Importantly, a collection of rejectors may be combined to form a composite rejector, which is shown to be far more effective than any of its individual components. A simple algorithm is proposed for the construction of each of the component rejectors. Its generality is established through close relationships with the Karhunen-Loeve expansion and Fisher's discriminant analysis. Composite rejectors were constructed for two representative applications, namely, appearance matching based object recognition and local feature detection. The results demonstrate substantial efficiency improvements over existing approaches, most notably Fisher's discriminant analysis.
Simon Baker, Shree K. Nayar
CVPR1
1996 Parametric Feature Detection
abstract
We propose an algorithm to automatically construct feature detectors for arbitrary parametric features. To obtain a high level of robustness we advocate the use of realistic multi-parameter feature models and incorporate optical and sensing effects. Each feature is represented as a densely sampled parametric manifold in a low dimensional subspace of a Hilbert space. During detection, the brightness distribution around each image pixel is projected into the subspace. If the projection lies sufficiently close to the feature manifold, the feature is detected and the location of the closest manifold point yields the feature parameters. The concepts of parameter reduction by normalization, dimension reduction, pattern rejection, and heuristic search are all employed to achieve the required efficiency. By applying the algorithm to appropriate parametric feature models, detectors have been constructed for five features, namely, step edge, roof edge, line, corner, and circular disc. Detailed experiments are reported on the robustness of detection and the accuracy of parameter estimation.
Shree K. Nayar, Simon Baker, Hiroshi Murase
CVPR2
1996 Algorithms for pattern rejection
abstract
The efficiency of pattern recognition is particularly crucial in two situations; whenever there are a large number of classes to discriminate, and, whenever recognition must be performed a large number of times. We develop a number of algorithms to cope with the demands of these difficult conditions. The algorithms achieve high efficiency by using pattern rejectors. A pattern rejector is a generalization of a classifier that quickly eliminates a large fraction of the candidate classes or inputs. After applying a rejector the recognition algorithms can concentrate their computational efforts on verifying the small number of remaining possibilities. The generality of our algorithms is established through a close relationship with the Karhunen-Loeve expansion. We experimented on two representative applications, namely, object recognition and feature detection. The results demonstrate substantial efficiency improvements over existing approaches, most notably Fisher's discriminant analysis (1939).
Simon Baker, Shree K. Nayar
ICPR1