EDBT 2026 Demo / reviewers in the wild / expert
Omkar M. Parkhi
dblp:38/10771
· DBLP profile ↗
15ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0001-8959-3284ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 41% Video understanding and tracking · 22% Segmentation and scene understanding · 14% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 74% Virtual and augmented reality · 26% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
activity recognition |
0.9 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Computer vision › 3D vision
egocentric vision |
0.9 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Wearable and physiological sensing › wearable camera › egocentric vision
egocentric sensing |
0.9 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Computer vision › 3D vision › 3d object detection
3d object detection and tracking |
0.7 | 1 | 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023 |
Computer vision › 3D vision
3d scene reconstruction |
0.7 | 1 | 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023 |
Computer vision › 3D vision › 3d scene understanding
egocentric 3d perception |
0.7 | 1 | 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023 |
Computer vision › Video understanding and tracking › multi-object tracking
object detection and tracking |
0.7 | 1 | 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.7 | 1 | 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023 |
Computer vision › Face, body and person analysis
face recognition |
0.6 | 2 | 2020 | Automated Video Face Labelling for Films and TV Material · IEEE Trans. Pattern Anal. Mach. Intell. 2020 A Compact and Discriminative Face Track Descriptor · CVPR 2014 |
Visual content generation and editing
image editing |
0.6 | 1 | 2022 | End-to-End Visual Editing with a Generatively Pre-trained Artist · ECCV (15) 2022 |
Wearable and physiological sensing › wearable display
smart glasses |
0.3 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Computer vision › Face, body and person analysis › face recognition
face verification |
0.2 | 1 | 2014 | A Compact and Discriminative Face Track Descriptor · CVPR 2014 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.2 | 2 | 2012 | The truth about cats and dogs · ICCV 2011 Cats and dogs · CVPR 2012 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.1 | 1 | 2012 | Cats and dogs · CVPR 2012 |
Computer vision › Face, body and person analysis
face tracking |
0.1 | 1 | 2020 | Automated Video Face Labelling for Films and TV Material · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Computer vision › Image recognition and object detection › object detection
deformable object detection |
0.1 | 1 | 2011 | The truth about cats and dogs · ICCV 2011 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2011 | The truth about cats and dogs · ICCV 2011 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.1 | 1 | 2011 | The truth about cats and dogs · ICCV 2011 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2014 | A Compact and Discriminative Face Track Descriptor · CVPR 2014 |
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model |
0.0 | 1 | 2011 | The truth about cats and dogs · ICCV 2011 |
Computer vision › Image recognition and object detection › object detection
template-based detection |
0.0 | 1 | 2011 | The truth about cats and dogs · ICCV 2011 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.7head pose · 1.7eye gaze · 1.7sim-to-real learning · 1.3image translation · 1.3weak supervision from aligned transcripts · 0.4linear programming · 0.4convnet face features · 0.4bag-of-words · 0.3binarization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reading Recognition in the WildabstractTo enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public. Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004 |
NeurIPS | 10 |
| 2024 | ICDAR 2024 Competition on Reading Documents Through Aria Glasses
Soumya Jahagirdar, Ajoy Mondal, Yuheng (Carl) Ren, Omkar M. Parkhi, C. V. Jawahar |
ICDAR (6) | 4 |
| 2023 | Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine PerceptionabstractWe introduce the Aria Digital Twin (ADT)1- an egocentric dataset captured using Aria glasses with extensive object, environment, and human level ground truth. This ADT release contains 200 sequences of real-world activities conducted by Aria wearers in two real indoor scenes with 398 object instances (344 stationary and 74 dynamic). Each sequence consists of: a) raw data of two monochrome camera streams, one RGB camera stream, two IMU streams; b) complete sensor calibration; c) ground truth data including continuous 6-degree-of-freedom (6DoF) poses of the Aria devices, object 6DoF poses, 3D eye gaze vectors, 3D human poses, 2D image segmentations, image depth maps; and d) photo-realistic synthetic renderings. To the best of our knowledge, there is no existing egocentric dataset with a level of accuracy, photo-realism and comprehensiveness comparable to ADT. By contributing ADT to the research community, our mission is to set a new standard for evaluation in the egocentric machine perception domain, which includes very challenging research problems such as 3D object detection and tracking, scene reconstruction and understanding, sim-to-real learning, human pose prediction - while also inspiring new machine perception tasks for augmented reality (AR) applications. To kick start exploration of the ADT research use cases, we evaluated several existing state-of-the-art methods for object detection, segmentation and image translation tasks that demonstrate the usefulness of ADT as a benchmarking dataset. Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, Carl Yuheng Ren |
ICCV | 7 |
| 2022 | End-to-End Visual Editing with a Generatively Pre-trained Artist
Andrew Brown 0006, Cheng-Yang Fu, Omkar M. Parkhi, Tamara L. Berg, Andrea Vedaldi |
ECCV (15) | 3 |
| 2020 | Automated Video Face Labelling for Films and TV MaterialabstractThe objective of this work is automatic labelling of characters in TV video and movies, given weak supervisory information provided by an aligned transcript. We make five contributions: (i) a new strategy for obtaining stronger supervisory information from aligned transcripts; (ii) an explicit model for classifying background characters, based on their face-tracks; (iii) employing new ConvNet based face features, and (iv) a novel approach for labelling all face tracks jointly using linear programming. Each of these contributions delivers a boost in performance, and we demonstrate this on standard benchmarks using tracks provided by authors of prior work. As a fifth contribution, we also investigate the generalisation and strength of the features and classifiers by applying them "in the raw" on new video material where no supervisory information is used. In particular, to provide high quality tracks on those material, we propose efficient track classifiers to remove false positive tracks by the face tracker. Overall we achieve a dramatic improvement over the state of the art on both TV series and film datasets, and almost saturate performance on some benchmarks. Omkar M. Parkhi, Esa Rahtu, Qiong Cao, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | VGGFace2: A Dataset for Recognising Faces across Pose and AgeabstractIn this paper, we introduce a new large-scale face dataset named VGGFace2. The dataset contains 3.31 million images of 9131 subjects, with an average of 362.6 images for each subject. Images are downloaded from Google Image Search and have large variations in pose, age, illumination, ethnicity and profession (e.g. actors, athletes, politicians). The dataset was collected with three goals in mind: (i) to have both a large number of identities and also a large number of images for each identity; (ii) to cover a large range of pose, age and ethnicity; and (iii) to minimise the label noise. We describe how the dataset was collected, in particular the automated and manual filtering stages to ensure a high accuracy for the images of each identity. To assess face recognition performance using the new dataset, we train ResNet-50 (with and without Squeeze-and-Excitation blocks) Convolutional Neural Networks on VGGFace2, on MS-Celeb-1M, and on their union, and show that training on VGGFace2 leads to improved recognition performance over pose and age. Finally, using the models trained on these datasets, we demonstrate state-of-the-art performance on the IJB-A and IJB-B face recognition benchmarks, exceeding the previous state-of-the-art by a large margin. The dataset and models are publicly available. Qiong Cao, Li Shen 0005, Weidi Xie, Omkar M. Parkhi, Andrew Zisserman |
FG | 4 |
| 2018 | Template adaptation for face verification and identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
Image Vis. Comput. | 4 |
| 2017 | Template Adaptation for Face Verification and Identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
FG | 4 |
| 2015 | Face Painting: querying art with photosabstractWe study the problem of matching photos of a person to paintings of that person, in order to retrieve similar paintings given a query photo. This is challenging as paintings span many media (oil, ink, watercolor) and can vary tremendously in style (caricature, pop art, minimalist). We make the following contributions: (i) we show that, depending on the face representation used, performance can be improved substantially by learning – either by a linear projection matrix common across identities, or by a per-identity classifier. We compare Fisher Vector and Convolutional Neural Network representations for this task; (ii) we introduce new datasets for learning and evaluating this problem; (iii) we also consider the reverse problem of retrieving photos from a large corpus given a painting; and finally, (iv) using the learnt descriptors, we show that, given a photo of a person, we are able to find their doppelganger in a large dataset of oil paintings, and how this result can be varied by modifying attributes (e.g. frowning, old looking). Elliot Crowley, Omkar M. Parkhi, Andrew Zisserman |
BMVC | 2 |
| 2015 | Deep Face RecognitionabstractThe goal of this paper is face recognition – from either a single photograph or from a set of faces tracked in a video. Recent progress in this area has been due to two factors: (i) end to end learning for the task using a convolutional neural network (CNN), and (ii) the availability of very large scale training datasets. We make two contributions: first, we show how a very large scale dataset (2.6M images, over 2.6K people) can be assembled by a combination of automation and human in the loop, and discuss the trade off between data purity and time; second, we traverse through the complexities of deep network training and face recognition to present methods and procedures to achieve comparable state of the art results on the standard LFW and YTF face benchmarks. Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman |
BMVC | 1 |
| 2014 | A Compact and Discriminative Face Track DescriptorabstractOur goal is to learn a compact, discriminative vector representation of a face track, suitable for the face recognition tasks of verification and classification. To this end, we propose a novel face track descriptor, based on the Fisher Vector representation, and demonstrate that it has a number of favourable properties. First, the descriptor is suitable for tracks of both frontal and profile faces, and is insensitive to their pose. Second, the descriptor is compact due to discriminative dimensionality reduction, and it can be further compressed using binarization. Third, the descriptor can be computed quickly (using hard quantization) and its compact size and fast computation render it very suitable for large scale visual repositories. Finally, the descriptor demonstrates good generalization when trained on one dataset and tested on another, reflecting its tolerance to the dataset bias. In the experiments we show that the descriptor exceeds the state of the art on both face verification task (YouTube Faces without outside training data, and INRIA-Buffy benchmarks), and face classification task (using the Oxford-Buffy dataset). Omkar M. Parkhi, Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
CVPR | 1 |
| 2013 | Fisher Vector Faces in the WildabstractSeveral recent papers on automatic face verification have significantly raised the performance bar by developing novel, specialised representations that outperform standard features such as SIFT for this problem. This paper makes two contributions: first, and somewhat surprisingly, we show that Fisher vectors on densely sampled SIFT features, i.e. an off-the-shelf object recognition representation, are capable of achieving state-of-the-art face verification performance on the challenging “Labeled Faces in the Wild” benchmark; second, since Fisher vectors are very high dimensional, we show that a compact descriptor can be learnt from them using discriminative metric learning. This compact descriptor has a better recognition accuracy and is very well suited to large scale identification tasks. Karen Simonyan, Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman |
BMVC | 2 |
| 2013 | The AXES PRO video search systemabstractWe demonstrate a multimedia content information retrieval engine developed for audiovisual digital libraries targeted at media professionals. It is the first of three multimedia IR systems being developed by the AXES project. The system brings together traditional text IR and state-of-the-art content indexing and retrieval technologies to allow users to search and browse digital libraries in novel ways. Key features include: metadata and ASR search and filtering, on-the-fly visual concept classification (categories, faces, places, and logos), and similarity search (instances and faces). Kevin McGuinness, Noel E. O'Connor, Robin Aly, Franciska de Jong, Ken Chatfield, Omkar M. Parkhi, Relja Arandjelovic, Andrew Zisserman, Matthijs Douze, Cordelia Schmid |
ICMR | 6 |
| 2012 | Cats and dogsabstractWe investigate the fine grained object categorization problem of determining the breed of animal from an image. To this end we introduce a new annotated dataset of pets covering 37 different breeds of cats and dogs. The visual problem is very challenging as these animals, particularly cats, are very deformable and there can be quite subtle differences between the breeds. We make a number of contributions: first, we introduce a model to classify a pet breed automatically from an image. The model combines shape, captured by a deformable part model detecting the pet face, and appearance, captured by a bag-of-words model that describes the pet fur. Fitting the model involves automatically segmenting the animal in the image. Second, we compare two classification approaches: a hierarchical one, in which a pet is first assigned to the cat or dog family and then to a breed, and a flat one, in which the breed is obtained directly. We also investigate a number of animal and image orientated spatial layouts. These models are very good: they beat all previously published results on the challenging ASIRRA test (cat vs dog discrimination). When applied to the task of discriminating the 37 different breeds of pets, the models obtain an average accuracy of about 59%, a very encouraging result considering the difficulty of the problem. Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, C. V. Jawahar |
CVPR | 1 |
| 2011 | The truth about cats and dogsabstractTemplate-based object detectors such as the deformable parts model of Felzenszwalb et al. [11] achieve state-of-the-art performance for a variety of object categories, but are still outperformed by simpler bag-of-words models for highly flexible objects such as cats and dogs. In these cases we propose to use the template-based model to detect a distinctive part for the class, followed by detecting the rest of the object via segmentation on image specific information learnt from that part. This approach is motivated by two observations: (i) many object classes contain distinctive parts that can be detected very reliably by template-based detectors, whilst the entire object cannot; (ii) many classes (e.g. animals) have fairly homogeneous coloring and texture that can be used to segment the object once a sample is provided in an image. We show quantitatively that our method substantially outperforms whole-body template-based detectors for these highly deformable object categories, and indeed achieves accuracy comparable to the state-of-the-art on the PASCAL VOC competition, which includes other models such as bag-of-words. Omkar M. Parkhi, Andrea Vedaldi, C. V. Jawahar, Andrew Zisserman |
ICCV | 1 |