Michael Isard

dblp:17/5751 · DBLP profile ↗
← Back
54ranked-venue papers
14as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 9 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-authorSoftware engineering, systems software and programming languages · 15 · 2 first-authorSystems, architecture and hardware · 8 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
24 papers
Language models and text generation · 28% Efficient and distributed learning · 14% Deep learning architectures and training · 10%
Computer architecture, parallel and distributed computing, and storage systems
11 papers
Parallel and multicore computing · 33% Distributed systems · 27% Memory systems · 15%
Software engineering, system software, and programming languages
5 papers
Operating systems · 40% Concurrent programming · 26% Programming languages and type systems · 22%
Databases, data mining, and information retrieval
7 papers
Information retrieval · 68% Query processing and optimization · 14% Distributed and cloud data management · 9%

Topics — the 30 heaviest of 100, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
instruction following
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
large language model
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
scaling laws
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning
distributed training
0.522023
Dynamic control flow in large-scale machine learning · EuroSys 2018
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Operating systems › resource management
memory management
0.412020
Learning-based Memory Allocation for C++ Server Workloads · ASPLOS 2020
Memory systems › memory management
memory allocation
0.412020
Learning-based Memory Allocation for C++ Server Workloads · ASPLOS 2020
Parallel and multicore computing
data-parallel programming
0.332013
Naiad: a timely dataflow system · SOSP 2013
Distributed data-parallel computing using a high-level programming language · SIGMOD Conference 2009
DryadLINQ: A System for General-Purpose Distributed Data-Parallel Computing Using a High-Level Language · OSDI 2008
Distributed systems
distributed machine learning
0.312018
Dynamic control flow in large-scale machine learning · EuroSys 2018
Parallel and multicore computing › data-parallel programming
distributed data-parallel execution
0.332013
Optimus: a dynamic rewriting framework for data-parallel execution plans · EuroSys 2013
DryadLINQ: A System for General-Purpose Distributed Data-Parallel Computing Using a High-Level Language · OSDI 2008
Dryad: distributed data-parallel programs from sequential building blocks · EuroSys 2007
Machine learning › Efficient and distributed learning
large-scale learning
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Distributed systems
large-scale machine learning systems
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Concurrent programming
transactional memory
0.222011
Semantics of transactional memory and automatic mutual exclusion · ACM Trans. Program. Lang. Syst. 2011
Semantics of transactional memory and automatic mutual exclusion · POPL 2008
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.212023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Representation and self-supervised learning › multimodal representation learning
multimodal embedding
0.212014
A Multi-View Embedding Space for Modeling Internet Images, Tags, and Their Semantics · Int. J. Comput. Vis. 2014
Information retrieval
image retrieval
0.222009
Bundling features for large scale partial-duplicate web image search · CVPR 2009
Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval · ICCV 2007
Distributed systems › distributed data processing
dataflow systems
0.212013
Naiad: a timely dataflow system · SOSP 2013
Processor architecture and microarchitecture
dynamic optimization
0.212013
Optimus: a dynamic rewriting framework for data-parallel execution plans · EuroSys 2013
Robotics › Robot navigation and mapping
object search
0.222008
Lost in quantization: Improving particular object retrieval in large scale image databases · CVPR 2008
Object retrieval with large vocabularies and fast spatial matching · CVPR 2007
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference
0.112012
Loose-limbed People: Estimating 3D Human Pose and Motion Using Non-parametric Belief Propagation · Int. J. Comput. Vis. 2012
Computer vision › 3D vision › pose estimation
human pose and motion estimation
0.112012
Loose-limbed People: Estimating 3D Human Pose and Motion Using Non-parametric Belief Propagation · Int. J. Comput. Vis. 2012
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › belief propagation
nonparametric belief propagation
0.112012
Loose-limbed People: Estimating 3D Human Pose and Motion Using Non-parametric Belief Propagation · Int. J. Comput. Vis. 2012
Query processing and optimization › query optimization
declarative query optimization
0.112011
Steno: automatic optimization of declarative queries · PLDI 2011
Programming languages and type systems
language semantics
0.112011
Semantics of transactional memory and automatic mutual exclusion · ACM Trans. Program. Lang. Syst. 2011
Compilers and program optimization › domain-specific compilation
query compilation
0.112011
Steno: automatic optimization of declarative queries · PLDI 2011
Storage systems › file systems
distributed file system
0.112011
TidyFS: A Simple and Small Distributed File System · USENIX ATC 2011
Cloud and datacenter computing
cluster resource management and scheduling
0.122009
Quincy: fair scheduling for distributed computing clusters · SOSP 2009
Dryad: distributed data-parallel programs from sequential building blocks · EuroSys 2007
Computer vision › 3D vision › local feature descriptor
descriptor learning
0.112010
Descriptor Learning for Efficient Retrieval · ECCV (3) 2010
Computer vision › Image recognition and object detection
image retrieval
0.112010
Descriptor Learning for Efficient Retrieval · ECCV (3) 2010

Methods — techniques the papers use, named apart from their topics

learning-based allocation · 0.9huge pages · 0.9transformer · 0.7pathways · 0.7data flow graph · 0.7data flow graphs · 0.6iterative computation · 0.3incremental computation · 0.3particle filtering · 0.3nonparametric belief propagation · 0.2dynamic rewriting · 0.2particle message passing · 0.1type system · 0.1operational semantics · 0.1minhash · 0.1min-hash · 0.1descriptor learning · 0.1locality-aware scheduling · 0.1
YearPublicationVenuePosition
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.26
2021 Falkirk Wheel: Rollback Recovery for Dataflow Systems
abstract
Data processing applications often combine computations with disparate fault-tolerance requirements. For example, batch computations prioritize throughput over recovery latency, and can tolerate recovery delays of up to several minutes, while streaming computations expect recovery latencies of at most a few seconds. However, state-of-the-art data systems each offer a single fault-tolerance regime, so complex applications either: (i) suffer performance degradation in steady state and during recovery due to the poor fit of the fault-tolerance regime for parts of the applications, or (ii) are difficult to maintain because they are developed using fragile combinations of batch and streaming systems that provide different APIs and schedulers, and evolve independently.
Ionel Gog, Michael Isard, Martín Abadi
SoCC2
2020 Learning-based Memory Allocation for C++ Server Workloads
abstract
Modern C++ servers have memory footprints that vary widely over time, causing persistent heap fragmentation of up to 2x from long-lived objects allocated during peak memory usage. This fragmentation is exacerbated by the use of huge (2MB) pages, a requirement for high performance on large heap sizes. Reducing fragmentation automatically is challenging because C++ memory managers cannot move objects.
Martin Maas 0001, David G. Andersen, Michael Isard, Mohammad Mahdi Javanmard, Kathryn S. McKinley, Colin Raffel
ASPLOS3
2019 Machine Learning Systems are Stuck in a Rut
abstract
In this paper we argue that systems for numerical computing are stuck in a local basin of performance and programmability. Systems researchers are doing an excellent job improving the performance of 5-year-old benchmarks, but gradually making it harder to explore innovative machine learning research ideas.
Paul Barham 0001, Michael Isard
HotOS2
2018 Dynamic control flow in large-scale machine learning
abstract
Many recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments.
Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng
EuroSys11
2016 TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng
OSDI10
2015 Timely Dataflow: A Model
Martín Abadi, Michael Isard
FORTE2
2015 Broom: Sweeping Out Garbage Collection from Big Data Systems
Ionel Gog, Jana Giceva, Malte Schwarzkopf, Kapil Vaswani, Dimitrios Vytiniotis, G. Ramalingam, Manuel Costa, Derek Gordon Murray, Steven Hand 0001, Michael Isard
HotOS10
2015 Scalability! But at what COST?
Frank McSherry, Michael Isard, Derek Gordon Murray
HotOS2
2014 A Multi-View Embedding Space for Modeling Internet Images, Tags, and Their Semantics
Yunchao Gong, Qifa Ke, Michael Isard, Svetlana Lazebnik
Int. J. Comput. Vis.3
2013 Differential Dataflow
Frank McSherry, Derek Gordon Murray, Rebecca Isaacs, Michael Isard
CIDR4
2013 Optimus: a dynamic rewriting framework for data-parallel execution plans
abstract
In distributed data-parallel computing, a user program is compiled into an execution plan graph (EPG), typically a directed acyclic graph. This EPG is the core data structure used by modern distributed execution engines for task distribution, job management, and fault tolerance. Once submitted for execution, the EPG remains largely unchanged at runtime except for some limited modifications. This makes it difficult to employ dynamic optimization techniques that could substantially improve the distributed execution based on runtime information.
Qifa Ke, Michael Isard
EuroSys2
2013 Naiad: a timely dataflow system
abstract
Naiad is a distributed system for executing data parallel, cyclic dataflow programs. It offers the high throughput of batch processors, the low latency of stream processors, and the ability to perform iterative and incremental computations. Although existing systems offer some of these features, applications that require all three have relied on multiple platforms, at the expense of efficiency, maintainability, and simplicity. Naiad resolves the complexities of combining these features in one framework.
Derek Gordon Murray, Frank McSherry, Rebecca Isaacs, Michael Isard, Paul Barham 0001, Martín Abadi
SOSP4
2012 Loose-limbed People: Estimating 3D Human Pose and Motion Using Non-parametric Belief Propagation
abstract
We formulate the problem of 3D human pose estimation and tracking as one of inference in a graphical model. Unlike traditional kinematic tree representations, our model of the body is a collection of loosely-connected body-parts. In particular, we model the body using an undirected graphical model in which nodes correspond to parts and edges to kinematic, penetration, and temporal constraints imposed by the joints and the world. These constraints are encoded using pair-wise statistical distributions, that are learned from motion-capture training data. Human pose and motion estimation is formulated as inference in this graphical model and is solved using Particle Message Passing ( PaMPas ). PaMPas is a form of non-parametric belief propagation that uses a variation of particle filtering that can be applied over a general graphical model with loops. The loose-limbed model and decentralized graph structure allow us to incorporate information from “bottom-up” visual cues, such as limb and head detectors, into the inference process. These detectors enable automatic initialization and aid recovery from transient tracking failures. We illustrate the method by automatically tracking people in multi-view imagery using a set of calibrated cameras and present quantitative evaluation using the HumanEva dataset.
Leonid Sigal, Michael Isard, Horst W. Haussecker, Michael J. Black
Int. J. Comput. Vis.2
2011 Steno: automatic optimization of declarative queries
abstract
Declarative queries enable programmers to write data manipulation code without being aware of the underlying data structure implementation. By increasing the level of abstraction over imperative code, they improve program readability and, crucially, create opportunities for automatic parallelization and optimization. For example, the Language Integrated Query (LINQ) extensions to C# allow the same declarative query to process in-memory collections, and datasets that are distributed across a compute cluster. However, our experiments show that the serial performance of declarative code is several times slower than the equivalent hand-optimized code, because it is implemented using run-time abstractions---such as iterators---that incur overhead due to virtual function calls and superfluous instructions.
Derek Gordon Murray, Michael Isard
PLDI2
2011 TidyFS: A Simple and Small Distributed File System
Dennis Fetterly, Maya Haridasan, Michael Isard, Swaminathan Sundararaman
USENIX ATC3
2011 Semantics of transactional memory and automatic mutual exclusion
abstract
Software Transactional Memory (STM) is an attractive basis for the development of language features for concurrent programming. However, the semantics of these features can be delicate and problematic. In this article we explore the trade-offs semantic simplicity, the viability of efficient implementation strategies, and the flexibility of language constructs. Specifically, we develop semantics and type systems for the constructs of the Automatic Mutual Exclusion (AME) programming model; our results apply also to other constructs, such as atomic blocks. With this semantics as a point of reference, we study several implementation strategies. We model STM systems that use in-place update, optimistic concurrency, lazy conflict detection, and rollback. These strategies are correct only under nontrivial assumptions that we identify and analyze. One important source of errors is that some efficient implementations create dangerous “zombie” computations where a transaction keeps running after experiencing a conflict; the assumptions confine the effects of these computations.
Martín Abadi, Andrew Birrell, Tim Harris 0001, Michael Isard
ACM Trans. Program. Lang. Syst.4
2010 Partition Min-Hash for Partial Duplicate Image Discovery
David C. Lee, Qifa Ke, Michael Isard
ECCV (1)3
2010 Descriptor Learning for Efficient Retrieval
James Philbin, Michael Isard, Josef Sivic, Andrew Zisserman
ECCV (3)2
2009 Implementation and Use of Transactional Memory with Dynamic Separation
Martín Abadi, Andrew Birrell, Tim Harris 0001, Johnson Hsieh, Michael Isard
CC5
2009 Bundling features for large scale partial-duplicate web image search
abstract
In state-of-the-art image retrieval systems, an image is represented by a bag of visual words obtained by quantizing high-dimensional local image descriptors, and scalable schemes inspired by text retrieval are then applied for large scale image indexing and retrieval. Bag-of-words representations, however: 1) reduce the discriminative power of image features due to feature quantization; and 2) ignore geometric relationships among visual words. Exploiting such geometric constraints, by estimating a 2D affine transformation between a query image and each candidate image, has been shown to greatly improve retrieval precision but at high computational cost. In this paper we present a novel scheme where image features are bundled into local groups. Each group of bundled features becomes much more discriminative than a single feature, and within each group simple and robust geometric constraints can be efficiently enforced. Experiments in Web image search, with a database of more than one million images, show that our scheme achieves a 49% improvement in average precision over the baseline bag-of-words approach. Retrieval performance is comparable to existing full geometric verification approaches while being much less computationally expensive. When combined with full geometric verification we achieve a 77% precision improvement over the baseline bag-of-words approach, and a 24% improvement over full geometric verification alone.
Qifa Ke, Michael Isard, Jian Sun 0001
CVPR3
2009 Distributed data-parallel computing using a high-level programming language
abstract
The Dryad and DryadLINQ systems offer a new programming model for large scale data-parallel computing. They generalize previous execution environments such as SQL and MapReduce in three ways: by providing a general-purpose distributed execution engine for data-parallel applications; by adopting an expressive data model of strongly typed .NET objects; and by supporting general-purpose imperative and declarative operations on datasets within a traditional high-level programming language.
Michael Isard
SIGMOD Conference1
2009 Quincy: fair scheduling for distributed computing clusters
abstract
This paper addresses the problem of scheduling concurrent jobs on clusters where application data is stored on the computing nodes. This setting, in which scheduling computations close to their data is crucial for performance, is increasingly common and arises in systems such as MapReduce, Hadoop, and Dryad as well as many grid-computing environments. We argue that data-intensive computation benefits from a fine-grain resource sharing model that differs from the coarser semi-static resource allocations implemented by most existing cluster computing architectures. The problem of scheduling with locality and fairness constraints has not previously been extensively studied under this resource-sharing model.
Michael Isard, Vijayan Prabhakaran, Jon Currey, Udi Wieder, Kunal Talwar, Andrew V. Goldberg
SOSP1
2009 Distributed aggregation for data-parallel computing: interfaces and implementations
abstract
Data-intensive applications are increasingly designed to execute on large computing clusters. Grouped aggregation is a core primitive of many distributed programming models, and it is often the most efficient available mechanism for computations such as matrix multiplication and graph traversal. Such algorithms typically require non-standard aggregations that are more sophisticated than traditional built-in database functions such as Sum and Max. As a result, the ease of programming user-defined aggregations, and the efficiency of their implementation, is of great current interest.
Pradeep Kumar Gunda, Michael Isard
SOSP3
2008 Lost in quantization: Improving particular object retrieval in large scale image databases
abstract
The state of the art in visual object retrieval from large databases is achieved by systems that are inspired by text retrieval. A key component of these approaches is that local regions of images are characterized using high-dimensional descriptors which are then mapped to ldquovisual wordsrdquo selected from a discrete vocabulary.This paper explores techniques to map each visual region to a weighted set of words, allowing the inclusion of features which were lost in the quantization stage of previous systems. The set of visual words is obtained by selecting words based on proximity in descriptor space. We describe how this representation may be incorporated into a standard tf-idf architecture, and how spatial verification is modified in the case of this soft-assignment. We evaluate our method on the standard Oxford Buildings dataset, and introduce a new dataset for evaluation. Our results exceed the current state of the art retrieval performance on these datasets, particularly on queries with poor initial recall where techniques like query expansion suffer. Overall we show that soft-assignment is always beneficial for retrieval with large vocabularies, at a cost of increased storage requirements for the index.
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman
CVPR3
2008 Continuously-adaptive discretization for message-passing algorithms
abstract
Continuously-Adaptive Discretization for Message-Passing (CAD-MP) is a new message-passing algorithm employing adaptive discretization. Most previous message-passing algorithms approximated arbitrary continuous probability distributions using either: a family of continuous distributions such as the exponential family; a particle-set of discrete samples; or a fixed, uniform discretization. In contrast, CAD-MP uses a discretization that is (i) non-uniform, and (ii) adaptive. The non-uniformity allows CAD-MP to localize interesting features (such as sharp peaks) in the marginal belief distributions with time complexity that scales logarithmically with precision, as opposed to uniform discretization which scales at best linearly. We give a principled method for altering the non-uniform discretization according to information-based measures. CAD-MP is shown in experiments on simulated data to estimate marginal beliefs much more precisely than competing approaches for the same computational expense.
Michael Isard, John MacCormick, Kannan Achan
NIPS1
2008 DryadLINQ: A System for General-Purpose Distributed Data-Parallel Computing Using a High-Level Language
Michael Isard, Dennis Fetterly, Mihai Budiu, Úlfar Erlingsson, Pradeep Kumar Gunda, Jon Currey
OSDI2
2008 Semantics of transactional memory and automatic mutual exclusion
abstract
Software Transactional Memory (STM) is an attractive basis for the development of language features for concurrent programming. However, the semantics of these features can be delicate and problematic. In this paper we explore the tradeoffs between semantic simplicity, the viability of efficient implementation strategies, and the flexibilityof language constructs. Specifically, we develop semantics and type systems for the constructs of the Automatic Mutual Exclusion (AME) programming model; our results apply also to other constructs, such as atomic blocks. With this semantics as a point of reference, we study several implementation strategies. We model STM systems that use in-place update, optimistic concurrency, lazy conflict detection, and roll-back. These strategies are correct only under non-trivial assumptions that we identify and analyze. One important source of errors is that some efficient implementations create dangerous 'zombie' computations where a transaction keeps running after experiencing a conflict; the assumptions confine the effects of these computations.
Martín Abadi, Andrew Birrell, Tim Harris 0001, Michael Isard
POPL4
2007 Object retrieval with large vocabularies and fast spatial matching
abstract
In this paper, we present a large-scale object retrieval system. The user supplies a query object by selecting a region of a query image, and the system returns a ranked list of images that contain the same object, retrieved from a large corpus. We demonstrate the scalability and performance of our system on a dataset of over 1 million images crawled from the photo-sharing site, Flickr [3], using Oxford landmarks as queries. Building an image-feature vocabulary is a major time and performance bottleneck, due to the size of our dataset. To address this problem we compare different scalable methods for building a vocabulary and introduce a novel quantization method based on randomized trees which we show outperforms the current state-of-the-art on an extensive ground-truth. Our experiments show that the quantization has a major effect on retrieval quality. To further improve query performance, we add an efficient spatial verification stage to re-rank the results returned from our bag-of-words model and show that this consistently improves search quality, though by less of a margin when the visual vocabulary is large. We view this work as a promising step towards much larger, "web-scale" image corpora.
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman
CVPR3
2007 Dryad: distributed data-parallel programs from sequential building blocks
abstract
Dryad is a general-purpose distributed execution engine for coarse-grain data-parallel applications. A Dryad application combines computational "vertices" with communication "channels" to form a dataflow graph. Dryad runs the application by executing the vertices of this graph on a set of available computers, communicating as appropriate through flies, TCP pipes, and shared-memory FIFOs.
Michael Isard, Mihai Budiu, Andrew Birrell, Dennis Fetterly
EuroSys1
2007 Automatic Mutual Exclusion
Michael Isard, Andrew Birrell
HotOS1
2007 Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval
abstract
Given a query image of an object, our objective is to retrieve all instances of that object in a large (1M+) image database. We adopt the bag-of-visual-words architecture which has proven successful in achieving high precision at low recall. Unfortunately, feature detection and quantization are noisy processes and this can result in variation in the particular visual words that appear in different images of the same object, leading to missed results. In the text retrieval literature a standard method for improving performance is query expansion. A number of the highly ranked documents from the original query are reissued as a new query. In this way, additional relevant terms can be added to the query. This is a form of blind relevance feedback and it can fail if `outlier' (false positive) documents are included in the reissued query. In this paper we bring query expansion into the visual domain via two novel contributions. Firstly, strong spatial constraints between the query image and each result allow us to accurately verify each return, suppressing the false positives which typically ruin text-based query expansion. Secondly, the verified images can be used to learn a latent feature model to enable the controlled construction of expanded queries. We illustrate these ideas on the 5000 annotated image Oxford building database together with more than 1M Flickr images. We show that the precision is substantially boosted, achieving total recall in many cases.
Ondrej Chum, James Philbin, Josef Sivic, Michael Isard, Andrew Zisserman
ICCV4
2006 Dense Motion and Disparity Estimation Via Loopy Belief Propagation
Michael Isard, John MacCormick
ACCV (2)1
2005 Estimating Disparity and Occlusions in Stereo Video Sequences
abstract
We propose an algorithm for estimating disparity and occlusion in stereo video sequences. The algorithm defines a prior on sequences of disparity maps using a 3D Markov random field, and approximately computes the MAP estimate for the disparity sequence using loopy belief propagation. In contrast to previous work on temporal stereo, the algorithm (i) correctly models half-occlusions - scene points visible in one camera but not the other - and (ii) enforces the so-called "monotonicity constraint" on the boundary of half-occluded regions. The algorithm is also able to exploit temporal coherence more appropriately than many previous approaches to temporal stereo, by employing additional states in the Markov random field. These additional states permit rudimentary motion estimation to be performed as part of the belief propagation, thus improving the quality of temporal inference. Parameters of the algorithm are learned from the ground truth disparities of a real stereo sequence. Qualitative results are shown on real sequences, including comparisons with competing approaches, and the performance of the algorithm is assessed quantitatively using the ground truth data.
Oliver Williams, Michael Isard, John MacCormick
CVPR (2)2
2004 Tracking Loose-Limbed People
Leonid Sigal, Sidharth Bhatia, Stefan Roth 0001, Michael J. Black, Michael Isard
CVPR (1)5
2003 PAMPAS: Real-Valued Graphical Models for Computer Vision
abstract
Probabilistic models have been adopted for many computer vision applications, however inference in high-dimensional spaces remains problematic. As the state-space of a model grows, the dependencies between the dimensions lead to an exponential growth in computation when performing inference. Many common computer vision problems naturally map onto the graphical model framework; the representation is a graph where each node contains a portion of the state-space and there is an edge between two nodes only if they are not independent conditional on the other nodes in the graph. When this graph is sparsely connected, belief propagation algorithms can turn an exponential inference computation into one, which is linear in the size of the graph. However belief propagation is only applicable when the variables in the nodes are discrete-valued or jointly represented by a single multivariate Gaussian distribution, and this rules out many computer vision applications. This paper combines belief propagation with ideas from particle filtering; the resulting algorithm performs inference on graphs containing both cycles and continuous-valued latent variables with general conditional probability distributions. Such graphical models have wide applicability in the computer vision domain and we test the algorithm on example problems of low-level edge linking and locating jointed structures in clutter.
Michael Isard
CVPR (1)1
2003 Attractive People: Assembling Loose-Limbed Models using Non-parametric Belief Propagation
abstract
The detection and pose estimation of people in images and video is made challenging by the variability of human appearance, the complexity of natural scenes, and the high dimensionality of articulated body mod- els. To cope with these problems we represent the 3D human body as a graphical model in which the relationships between the body parts are represented by conditional probability distributions. We formulate the pose estimation problem as one of probabilistic inference over a graphi- cal model where the random variables correspond to the individual limb parameters (position and orientation). Because the limbs are described by 6-dimensional vectors encoding pose in 3-space, discretization is im- practical and the random variables in our model must be continuous- valued. To approximate belief propagation in such a graph we exploit a recently introduced generalization of the particle filter. This framework facilitates the automatic initialization of the body-model from low level cues and is robust to occlusion of body parts and scene clutter.
Leonid Sigal, Michael Isard, Benjamin H. Sigelman, Michael J. Black
NIPS2
2003 A Cooperative Internet Backup Scheme
Mark Lillibridge, Sameh Elnikety, Andrew Birrell, Michael Burrows, Michael Isard
USENIX ATC, General Track5
2003 Distributed rendering of interactive soft shadows
Michael Isard, Mark Shand, Alan Heirich
Parallel Comput.1
2002 Automatic Camera Calibration from a Single Manhattan Image
Jonathan Deutscher, Michael Isard, John MacCormick
ECCV (4)2
2001 BraMBLe: A Bayesian Multiple-Blob Tracker
abstract
Blob trackers have become increasingly powerful in recent years largely due to the adoption of statistical appearance models which allow effective background subtraction and robust tracking of deforming foreground objects. It has been standard, however, to treat background and foreground modelling as separate processes-background subtraction is followed by blob detection and tracking-which prevents a principled computation of image likelihoods. This paper presents two theoretical advances which address this limitation and lead to a robust multiple-person tracking system suitable for single-camera real-time surveillance applications. The first innovation is a multi-blob likelihood function which assigns directly comparable likelihoods to hypotheses containing different numbers of objects. This likelihood function has a rigorous mathematical basis: it is adapted from the theory of Bayesian correlation, but uses the assumption of a static camera to create a more specific background model while retaining a unified approach to background and foreground modelling. Second we introduce a Bayesian filter for tracking multiple objects when the number of objects present is unknown and varies over time. We show how a particle filter can be used to perform joint inference on both the number of objects present and their configurations. Finally we demonstrate that our system runs comfortably in real time on a modest workstation when the number of blobs in the scene is small.
Michael Isard, John MacCormick
ICCV1
2001 Bayesian Object Localisation in Images
Josephine Sullivan, Andrew Blake 0001, Michael Isard, John MacCormick
Int. J. Comput. Vis.3
2000 Partitioned Sampling, Articulated Objects, and Interface-Quality Hand Tracking
John MacCormick, Michael Isard
ECCV (2)2
2000 Learning and Classification of Complex Dynamics
abstract
Standard, exact techniques based on likelihood maximization are available for learning auto-regressive process models of dynamical processes. The uncertainty of observations obtained from real sensors means that dynamics can be observed only approximately. Learning can still be achieved via "EM-K"-expectation-maximization (EM) based on Kalman filtering. This cannot handle more complex dynamics, however, involving multiple classes of motion. A problem arises also in the case of dynamical processes observed visually: background clutter arising for example, in camouflage, produces non-Gaussian observation noise. Even with a single dynamical class, non-Gaussian observations put the learning problem beyond the scope of EM-K. For those cases, we show here how "EM-C"-based on the CONDENSATION algorithm which propagates random "particle-sets," can solve the learning problem. Here, learning in clutter is studied experimentally using visual observations of a hand moving over a desktop. The resulting learned dynamical model is shown to have considerable predictive value: when used as a prior for estimation of motion, the burden of computation in visual observation is significantly reduced. Multiclass dynamics are studied via visually observed juggling; plausible dynamical models have been found to emerge from the learning process, and accurate classification of motion has resulted. In practice, EM-C learning is computationally burdensome and the paper concludes with some discussion of computational complexity.
Ben North, Andrew Blake 0001, Michael Isard, Jens Rittscher
IEEE Trans. Pattern Anal. Mach. Intell.3
1999 Object Localization by Bayesian Correlation
abstract
Maximisation of cross correlation is a commonly used principle for intensity based object localization that gives a single estimate of location. However, to facilitate sequential inference (e.g. over time or scale) and to allow the representation of ambiguity, it is desirable to represent an entire probability distribution for object location. Although the cross correlation itself (or some function of it) has sometimes been treated as a probability distribution, this is not generally justifiable. Bayesian correlation achieves a consistent probabilistic treatment by combining several developments. The first is the interpretation of correlation matching functions in probabilistic terms, as observation likelihoods. Second, probability distributions of filter bank responses are learned from training examples. Inescapably, response learning also demands statistical modelling of background intensities, and there are links here with image coding and Independent Component Analysis. Lastly, multi scale processing is achieved in a Bayesian context by means of a new algorithm, layered sampling, for which asymptotic properties are derived.
Josephine Sullivan, Andrew Blake 0001, Michael Isard, John MacCormick
ICCV3
1998 A Smoothing Filter for CONDENSATION
Michael Isard, Andrew Blake 0001
ECCV (1)1
1998 ICONDENSATION: Unifying Low-Level and High-Level Tracking in a Stochastic Framework
Michael Isard, Andrew Blake 0001
ECCV (1)1
1998 A Mixed-State CONDENSATION Tracker with Automatic Model-Switching
abstract
There is considerable interest in the computer vision community in representing and modelling motion. Motion models are used as predictors to increase the robustness and accuracy of visual trackers, and as classifiers for gesture recognition. This paper presents a significant development of random sampling methods to allow automatic switching between multiple motion models as a natural extension of the tracking process. The Bayesian mixed-state framework is described in its generality, and the example of a bouncing ball is used to demonstrate that a mixed-state model can significantly improve tracking performance in heavy clutter. The relevance of the approach to the problem of gesture recognition is then investigated using a tracker which is able to follow the natural drawing action of a hand holding a pen, and switches state according to the hand's motion.
Michael Isard, Andrew Blake 0001
ICCV1
1998 Learning Multi-Class Dynamics
Andrew Blake 0001, Ben North, Michael Isard
NIPS3
1998 CONDENSATION - Conditional Density Propagation for Visual Tracking
Michael Isard, Andrew Blake 0001
Int. J. Comput. Vis.1
1996 Contour Tracking by Stochastic Propagation of Conditional Density
Michael Isard, Andrew Blake 0001
ECCV (1)1
1996 The CONDENSATION Algorithm - Conditional Density Propagation and Applications to Visual Tracking
Andrew Blake 0001, Michael Isard
NIPS2
1995 Learning to Track the Visual Motion of Contours
Andrew Blake 0001, Michael Isard, David Reynard
Artif. Intell.2
1994 3D position, attitude and shape input using video tracking of hands and lips
abstract
Recent developments in video-tracking allow the outlines of moving, natural objects in a video-camera input stream to be tracked live, at full video-rate. Previous systems have been available to do this for specially illuminated objects or for naturally illuminated but polyhedral objects. Other systems have been able to track nonpolyhedral objects in motion, in some cases from live video, but following only centroids or key-points rather than tracking whole curves. The system described here can track accurately the curved silhouettes of moving non-polyhedral objects at frame-rate, for example hands, lips, legs, vehicles, fruit, and without any special hardware beyond a desktop workstation and a video-camera and framestore.
Andrew Blake 0001, Michael Isard
SIGGRAPH2