VLDB 2026 Research / reviewers in the wild / expert
James Philbin
dblp:48/6239
· DBLP profile ↗
19ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Image recognition and object detection · 17% Autonomous driving · 16% Face, body and person analysis · 13% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 94% Indexing and storage engines · 6% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 42% Parallel and multicore computing · 33% Interconnection networks and networks-on-chip · 17% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Autonomous driving
behavior prediction |
0.4 | 1 | 2019 | Rules of the Road: Predicting Driving Behavior With a Convolutional Model of Semantic Interactions · CVPR 2019 |
Robotics › Autonomous driving › behavior prediction
driving behavior prediction |
0.4 | 1 | 2019 | Rules of the Road: Predicting Driving Behavior With a Convolutional Model of Semantic Interactions · CVPR 2019 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.4 | 1 | 2019 | Rules of the Road: Predicting Driving Behavior With a Convolutional Model of Semantic Interactions · CVPR 2019 |
Computer vision › Image recognition and object detection
image retrieval |
0.3 | 2 | 2014 | Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014 Descriptor Learning for Efficient Retrieval · ECCV (3) 2010 |
Computer vision › 3D vision › stereo vision › stereo matching
deep stereo matching |
0.2 | 1 | 2016 | Deep Stereo: Learning to Predict New Views from the World's Imagery · CVPR 2016 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end learning |
0.2 | 1 | 2016 | Deep Stereo: Learning to Predict New Views from the World's Imagery · CVPR 2016 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.2 | 1 | 2016 | The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition · ECCV (3) 2016 |
Computer vision › 3D vision › visual localization › geo-localization
image geo-localization |
0.2 | 1 | 2016 | PlaNet - Photo Geolocation with Convolutional Neural Networks · ECCV (8) 2016 |
Machine learning › Trustworthy machine learning › robustness › noisy data
noisy training data |
0.2 | 1 | 2016 | The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition · ECCV (3) 2016 |
Rendering
novel view synthesis |
0.2 | 1 | 2016 | Deep Stereo: Learning to Predict New Views from the World's Imagery · CVPR 2016 |
Computer vision › Face, body and person analysis
face clustering |
0.2 | 1 | 2015 | FaceNet: A unified embedding for face recognition and clustering · CVPR 2015 |
Computer vision › Face, body and person analysis
face recognition |
0.2 | 1 | 2015 | FaceNet: A unified embedding for face recognition and clustering · CVPR 2015 |
Computer vision › Face, body and person analysis › face recognition
face verification |
0.2 | 1 | 2015 | FaceNet: A unified embedding for face recognition and clustering · CVPR 2015 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.2 | 1 | 2015 | FaceNet: A unified embedding for face recognition and clustering · CVPR 2015 |
Machine learning › Representation and self-supervised learning › representation learning › metric learning
triplet embedding |
0.2 | 1 | 2015 | FaceNet: A unified embedding for face recognition and clustering · CVPR 2015 |
Information retrieval
image retrieval |
0.2 | 2 | 2011 | Geometric Latent Dirichlet Allocation on a Matching Graph for Large-scale Image Datasets · Int. J. Comput. Vis. 2011 Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval · ICCV 2007 |
Machine learning › Deep learning architectures and training
deep ranking |
0.2 | 1 | 2014 | Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014 |
Computer vision › Image recognition and object detection
image similarity |
0.2 | 1 | 2014 | Learning Fine-Grained Image Similarity with Deep Ranking · CVPR 2014 |
Robotics › Robot navigation and mapping
object search |
0.2 | 2 | 2008 | Lost in quantization: Improving particular object retrieval in large scale image databases · CVPR 2008 Object retrieval with large vocabularies and fast spatial matching · CVPR 2007 |
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation |
0.1 | 1 | 2011 | Geometric Latent Dirichlet Allocation on a Matching Graph for Large-scale Image Datasets · Int. J. Comput. Vis. 2011 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2011 | Geometric Latent Dirichlet Allocation on a Matching Graph for Large-scale Image Datasets · Int. J. Comput. Vis. 2011 |
Information retrieval
multimedia analysis and retrieval |
0.1 | 1 | 2011 | Geometric Latent Dirichlet Allocation on a Matching Graph for Large-scale Image Datasets · Int. J. Comput. Vis. 2011 |
Computer vision › 3D vision › local feature descriptor
descriptor learning |
0.1 | 1 | 2010 | Descriptor Learning for Efficient Retrieval · ECCV (3) 2010 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
soft labels |
0.1 | 1 | 2008 | Lost in quantization: Improving particular object retrieval in large scale image databases · CVPR 2008 |
Computer vision › Image recognition and object detection
visual word quantization |
0.1 | 1 | 2008 | Lost in quantization: Improving particular object retrieval in large scale image databases · CVPR 2008 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise |
0.1 | 1 | 2016 | The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition · ECCV (3) 2016 |
Information retrieval › query reformulation › query expansion
automatic query expansion |
0.1 | 1 | 2007 | Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval · ICCV 2007 |
Information retrieval › image retrieval
object retrieval |
0.1 | 1 | 2007 | Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval · ICCV 2007 |
Information retrieval › query reformulation
query expansion |
0.1 | 1 | 2007 | Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval · ICCV 2007 |
Parallel and multicore computing › parallel computing › parallel communication
user-level communication |
0.1 | 2 | 2002 | Experiences with VI Communication for Database Storage · ISCA 2002 User-Space Communication: A Quantitative Study · SC 1998 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 0.6end-to-end training · 0.5deep neural network · 0.5supervised learning · 0.4latent dirichlet allocation · 0.2graph-based matching · 0.2online triplet mining · 0.2deep convolutional network · 0.2triplet sampling · 0.2distributed asynchronous stochastic gradient descent · 0.2spatial constraints · 0.1relevance feedback · 0.1microbenchmarking · 0.1generative feature model · 0.1multi-queue replacement · 0.1zero-copy protocol · 0.0quantitative comparison · 0.0thread creation hints · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | QuadroNet: Multi-Task Learning for Real-Time Semantic Depth Aware Instance SegmentationabstractVision for autonomous driving is a uniquely challenging problem: the number of tasks required for full scene understanding is large and diverse; the quality requirements on each task are stringent due to the safety-critical nature of the application; and the latency budget is limited, requiring real-time solutions. In this work we address these challenges with QuadroNet, a one-shot network that jointly produces four outputs: 2D detections, instance segmentation, semantic segmentation, and monocular depth estimates in real-time (>60fps) on consumer-grade GPU hardware. On a challenging real-world autonomous driving dataset, we demonstrate an increase of+2.4% mAP for detection, +3.15% mIoU for semantic segmentation, +5.05% [email protected] for instance segmentation and +1.36% in δ <; 1.25 for depth prediction over a baseline approach. We also compare our work against other multi-task learning approaches on Cityscapes and demonstrate state-of-the-art results. Kratarth Goel, Praveen Srinivasan, Sarah Tariq, James Philbin |
WACV | 4 |
| 2019 | Rules of the Road: Predicting Driving Behavior With a Convolutional Model of Semantic InteractionsabstractWe focus on the problem of predicting future states of entities in complex, real-world driving scenarios. Previous research has approached this problem via low-level signals to predict short time horizons, and has not addressed how to leverage key assets relied upon heavily by industry self-driving systems: (1) large 3D perception efforts which provide highly accurate 3D states of agents with rich attributes, and (2) detailed and accurate semantic maps of the environment (lanes, traffic lights, crosswalks, etc). We present a unified representation which encodes such high-level semantic information in a spatial grid, allowing the use of deep convolutional models to fuse complex scene context. This enables learning entity-entity and entity-environment interactions with simple, feed-forward computations in each timestep within an overall temporal model of an agent's behavior. We propose different ways of modelling the future as a {\em distribution} over future states using standard supervised learning. We introduce a novel dataset providing industry-grade rich perception and semantic inputs, and empirically show we can effectively learn fundamentals of driving behavior. Joey Hong, Benjamin Sapp, James Philbin |
CVPR | 3 |
| 2016 | Deep Stereo: Learning to Predict New Views from the World's ImageryabstractDeep networks have recently enjoyed enormous success when applied to recognition and classification problems in computer vision [22, 33], but their use in graphics problems has been limited ([23, 7] are notable recent exceptions). In this work, we present a novel deep architecture that performs new view synthesis directly from pixels, trained from a large number of posed image sets. In contrast to traditional approaches, which consist of multiple complex stages of processing, each of which requires careful tuning and can fail in unexpected ways, our system is trained end-to-end. The pixels from neighboring views of a scene are presented to the network, which then directly produces the pixels of the unseen view. The benefits of our approach include generality (we only require posed image sets and can easily apply our method to different domains), and high quality results on traditionally difficult scenes. We believe this is due to the end-to-end nature of our system, which is able to plausibly generate pixels according to color, depth, and texture priors learnt automatically from the training data. We show view interpolation results on imagery from the KITTI dataset [12], from data from [1] as well as on Google Street View images. To our knowledge, our work is the first to apply deep learning to the problem of new view synthesis from sets of real-world, natural imagery. John Flynn, Ivan Neulander, James Philbin, Noah Snavely |
CVPR | 3 |
| 2016 | The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition
Jonathan Krause, Benjamin Sapp, Howard Zhou, Alexander Toshev, Tom Duerig, James Philbin, Li Fei-Fei 0001 |
ECCV (3) | 7 |
| 2016 | PlaNet - Photo Geolocation with Convolutional Neural Networks
Tobias Weyand, Ilya Kostrikov, James Philbin |
ECCV (8) | 3 |
| 2015 | FaceNet: A unified embedding for face recognition and clusteringabstractDespite significant recent advances in the field of face recognition [10, 14, 15, 17], implementing face verification and recognition efficiently at scale presents serious challenges to current approaches. In this paper we present a system, called FaceNet, that directly learns a mapping from face images to a compact Euclidean space where distances directly correspond to a measure offace similarity. Once this space has been produced, tasks such as face recognition, verification and clustering can be easily implemented using standard techniques with FaceNet embeddings asfeature vectors. Our method uses a deep convolutional network trained to directly optimize the embedding itself, rather than an intermediate bottleneck layer as in previous deep learning approaches. To train, we use triplets of roughly aligned matching / non-matching face patches generated using a novel online triplet mining method. The benefit of our approach is much greater representational efficiency: we achieve state-of-the-artface recognition performance using only 128-bytes perface. On the widely used Labeled Faces in the Wild (LFW) dataset, our system achieves a new record accuracy of 99.63%. On YouTube Faces DB it achieves 95.12%. Our system cuts the error rate in comparison to the best published result [15] by 30% on both datasets. Florian Schroff, Dmitry Kalenichenko, James Philbin |
CVPR | 3 |
| 2014 | Learning Fine-Grained Image Similarity with Deep RankingabstractLearning fine-grained image similarity is a challenging task. It needs to capture between-class and within-class image differences. This paper proposes a deep ranking model that employs deep learning techniques to learn similarity metric directly from images. It has higher learning capability than models based on hand-crafted features. A novel multiscale network structure has been developed to describe the images effectively. An efficient triplet sampling algorithm is also proposed to learn the model with distributed asynchronized stochastic gradient. Extensive experiments show that the proposed algorithm outperforms models based on hand-crafted visual features and deep classification models. Jiang Wang 0001, Yang Song 0009, Thomas K. Leung, Charles Rosenberg 0001, Jingbin Wang, James Philbin, Bo Chen 0019, Ying Wu 0001 |
CVPR | 6 |
| 2011 | Geometric Latent Dirichlet Allocation on a Matching Graph for Large-scale Image Datasets
James Philbin, Josef Sivic, Andrew Zisserman |
Int. J. Comput. Vis. | 1 |
| 2010 | Descriptor Learning for Efficient Retrieval
James Philbin, Michael Isard, Josef Sivic, Andrew Zisserman |
ECCV (3) | 1 |
| 2008 | Near Duplicate Image Detection: min-Hash and tf-idf Weighting
Ondrej Chum, James Philbin, Andrew Zisserman |
BMVC | 2 |
| 2008 | Geometric LDA: A Generative Model for Particular Object DiscoveryabstractAutomatically organizing collections of images presents serious challenges to the current state-of-the art methods in image data mining. Often, what is required is that images taken in the same place, of the same thing, or of the same person be conceptually grouped together. To achieve this, we introduce the Geometric Latent Dirichlet Allocation (gLDA) model for unsupervised particular object discovery in unordered image collections. This explicitly represents documents as mixtures of particular objects or facades, and builds rich latent topic models which incorporate the identity and locations of visual words specific to the topic in a geometrically consistent way. Applying standard inference techniques to this model enables images likely to contain the same object to be probabilistically grouped and ranked. We demonstrate the model on a publicly available dataset of Oxford images, and show examples of spatially consistent groupings. 1 James Philbin, Josef Sivic, Andrew Zisserman |
BMVC | 1 |
| 2008 | Lost in quantization: Improving particular object retrieval in large scale image databasesabstractThe state of the art in visual object retrieval from large databases is achieved by systems that are inspired by text retrieval. A key component of these approaches is that local regions of images are characterized using high-dimensional descriptors which are then mapped to ldquovisual wordsrdquo selected from a discrete vocabulary.This paper explores techniques to map each visual region to a weighted set of words, allowing the inclusion of features which were lost in the quantization stage of previous systems. The set of visual words is obtained by selecting words based on proximity in descriptor space. We describe how this representation may be incorporated into a standard tf-idf architecture, and how spatial verification is modified in the case of this soft-assignment. We evaluate our method on the standard Oxford Buildings dataset, and introduce a new dataset for evaluation. Our results exceed the current state of the art retrieval performance on these datasets, particularly on queries with poor initial recall where techniques like query expansion suffer. Overall we show that soft-assignment is always beneficial for retrieval with large vocabularies, at a cost of increased storage requirements for the index. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman |
CVPR | 1 |
| 2007 | Object retrieval with large vocabularies and fast spatial matchingabstractIn this paper, we present a large-scale object retrieval system. The user supplies a query object by selecting a region of a query image, and the system returns a ranked list of images that contain the same object, retrieved from a large corpus. We demonstrate the scalability and performance of our system on a dataset of over 1 million images crawled from the photo-sharing site, Flickr [3], using Oxford landmarks as queries. Building an image-feature vocabulary is a major time and performance bottleneck, due to the size of our dataset. To address this problem we compare different scalable methods for building a vocabulary and introduce a novel quantization method based on randomized trees which we show outperforms the current state-of-the-art on an extensive ground-truth. Our experiments show that the quantization has a major effect on retrieval quality. To further improve query performance, we add an efficient spatial verification stage to re-rank the results returned from our bag-of-words model and show that this consistently improves search quality, though by less of a margin when the visual vocabulary is large. We view this work as a promising step towards much larger, "web-scale" image corpora. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman |
CVPR | 1 |
| 2007 | Total Recall: Automatic Query Expansion with a Generative Feature Model for Object RetrievalabstractGiven a query image of an object, our objective is to retrieve all instances of that object in a large (1M+) image database. We adopt the bag-of-visual-words architecture which has proven successful in achieving high precision at low recall. Unfortunately, feature detection and quantization are noisy processes and this can result in variation in the particular visual words that appear in different images of the same object, leading to missed results. In the text retrieval literature a standard method for improving performance is query expansion. A number of the highly ranked documents from the original query are reissued as a new query. In this way, additional relevant terms can be added to the query. This is a form of blind relevance feedback and it can fail if `outlier' (false positive) documents are included in the reissued query. In this paper we bring query expansion into the visual domain via two novel contributions. Firstly, strong spatial constraints between the query image and each result allow us to accurately verify each return, suppressing the false positives which typically ruin text-based query expansion. Secondly, the verified images can be used to learn a latent feature model to enable the controlled construction of expanded queries. We illustrate these ideas on the 5000 annotated image Oxford building database together with more than 1M Flickr images. We show that the precision is substantially boosted, achieving total recall in many cases. Ondrej Chum, James Philbin, Josef Sivic, Michael Isard, Andrew Zisserman |
ICCV | 2 |
| 2002 | Experiences with VI Communication for Database StorageabstractThis paper examines how VI-based interconnects can be used to improve I/O path performance between a database server and the storage subsystem. We design and implement a software layer, DSA, that is layered between the application and VI. DSA takes advantage of specific VI features and deals with many of its shortcomings. We provide and evaluate one kernel-level and two user-level implementations of DSA. These implementations trade transparency and generality for performance at different degrees, and unlike research prototypes are designed to be suitable for realworld deployment. We present detailed measurements using a commercial database management system with both micro-benchmarks and industrial database workloads on a mid-size, 4 CPU, and a large, 32 CPU, database server. Our results show that VI-based interconnects and user-level communication can improve all aspects of the I/O path between the database system and the storage back-end. We also find that to make effective use of VI in I/O intensive environments we need to provide substantial additional functionality than what is currently provided by VI. Finally, new storage APIs that help minimize kernel involvement in the I/O path are needed to fully exploit the benefits of VI-based communication. Yuanyuan Zhou 0001, Kai Li 0001, Angelos Bilas, Suresh Jagannathan, Cezary Dubnicki, James Philbin |
ISCA | 6 |
| 2001 | The Multi-Queue Replacement Algorithm for Second Level Buffer Caches
Yuanyuan Zhou 0001, James Philbin, Kai Li 0001 |
USENIX ATC, General Track | 2 |
| 1998 | User-Space Communication: A Quantitative StudyabstractPowerful commodity systems and networks offer a promising direction for high performance computing because they are inexpensive and they closely track technology progress. However, high, raw-hardware performance is rarely delivered to the end user. Previous work has shown that the bottleneck in these architectures is the overheads imposed by the software communication layer. To reduce these overheads, researchers have proposed a number of user-space communication models. The common feature of these models is that applications have direct access to the network, bypassing the operating system in the common case and thus avoiding the cost of send/receive system calls. In this paper we examine five user-space communication layers, that represent different points in the configuration space: Generic AM, BIP-0.92, FM-2.02, PM-1.2, and VMMC-2. Although these systems support different communication paradigms and employ a variety of different implementation tradeoffs, we are able to quantitatively compare them on a single testbed consisting of a cluster of high-end PCs connected by a Myrinet network. We find that all five communication systems have very low latency for small messages, in the range of 5 to 17 s. Not surprisingly, this range is strongly influenced by the functionality offered by each system. We are encouraged, however, to find that features such as protected and reliable communication at user level and multiprogramming can be provided at very low cost. Bandwidth, however, depends primarily on how data is transferred between host memory and the network. Most of the investigated libraries support zero-copy protocols for certain types of data transfers, but differ significantly in the bandwidth delivered to end users. The highest bandwidth, between 95 and 125 MBytes/s for long message transfers, is delivered by libraries that use DMA on both send and receive sides and avoid all data copies. Libraries that perform additional data copies or use programmed I/O to send data to the network achieve lower maximum bandwidth, in the range of 60-70 MBytes/s. Soichiro Araki, Angelos Bilas, Cezary Dubnicki, Jan Edler, Koichi Konishi, James Philbin |
SC | 6 |
| 1996 | Thread Scheduling for Cache LocalityabstractThis paper describes a method to improve the cache locality of sequential programs by scheduling fine-grained threads. The algorithm relies upon hints provided at the time of thread creation to determine a thread execution order likely to reduce cache misses. This technique may be particularly valuable when compiler-directed tiling is not feasible. Experiments with several application programs, on two systems with different cache structures, show that our thread scheduling method can improve program performance by reducing second-level cache misses. James Philbin, Jan Edler, Otto J. Anshus, Craig C. Douglas, Kai Li 0001 |
ASPLOS | 1 |
| 1992 | A Customizable Substrate for Concurrent LanguagesabstractWe describe an approach to implementing a wide-range of concurrency paradigms in high-level (symbolic) programming languages. The focus of our discussion is STING, a dialect of Scheme, that supports lightweight threads of control and virtual processors as first-class objects. Given the significant degree to which the behavior of these objects may be customized, we can easily express a variety of concurrency paradigms and linguistic structures within a common framework without loss of efficiency. Suresh Jagannathan, James Philbin |
PLDI | 2 |