VLDB 2026 Research / reviewers in the wild / expert
Saptarshi Sinha
dblp:151/6360
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HD-EPIC: A Highly-Detailed Egocentric Video DatasetabstractWe present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HD-EPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments.We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to recognise recipes, ingredients, nutrition, fine-grained actions, 3D perception, object motion, and gaze direction. The powerful long-context Gemini Pro only achieves 37.6% on this benchmark, showcasing its difficulty and highlighting shortcomings in current VLMs. We additionally assess action recognition, sound recognition, and long-term video-object segmentation on HD-EPIC.HD-EPIC is 41 hours of video in 9 kitchens with digital twins of 413 kitchen fixtures, capturing 69 recipes, 59K fine-grained actions, 51K audio events, 20K object movements and 37K object masks lifted to 3D. On average, we have 263 annotations per minute of our unscripted videos. Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu 0001, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu 0006, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen |
CVPR | 3 |
| 2024 | Every Shot Counts: Using Exemplars for Repetition Counting in Videos
Saptarshi Sinha, Alexandros Stergiou, Dima Damen |
ACCV (3) | 1 |
| 2023 | MILA: Memory-Based Instance-Level Adaptation for Cross-Domain Object Detection
Onkar Krishna, Hiroki Ohashi, Saptarshi Sinha |
BMVC | 3 |
| 2023 | Use Your Head: Improving Long-Tail Video RecognitionabstractThis paper presents an investigation into long-tail video recognition. We demonstrate that, unlike naturally-collected video datasets and existing long-tail image benchmarks, current video benchmarks fall short on multiple long-tailed properties. Most critically, they lack few-shot classes in their tails. In response, we propose new video benchmarks that better assess long-tail recognition, by sampling subsets from two datasets: SSv2 and VideoLT. We then propose a method, Long-Tail Mixed Reconstruction (LMR), which reduces overfitting to instances from few-shot classes by reconstructing them as weighted combinations of samples from head classes. LMR then employs label mixing to learn robust decision boundaries. It achieves state-of-the-art average class accuracy on EPIC-KITCHENS and the proposed SSv2-LT and VideoLT-LT. Benchmarks and code at: github.com/tobyperrett/lmr Toby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi, Dima Damen |
CVPR | 2 |
| 2023 | Difficulty-Net: Learning to Predict Difficulty for Long-Tailed RecognitionabstractLong-tailed datasets, where head classes comprise much more training samples than tail classes, cause recognition models to get biased towards the head classes. Weighted loss is one of the most popular ways of mitigating this issue, and a recent work has suggested that class-difficulty might be a better clue than conventionally used class-frequency to decide the distribution of weights. A heuristic formulation was used in the previous work for quantifying the difficulty, but we empirically find that the optimal formulation varies depending on the characteristics of datasets. Therefore, we propose Difficulty-Net, which learns to predict the difficulty of classes using the model’s performance in a meta-learning framework. To make it learn reasonable difficulty of a class within the context of other classes, we newly introduce two key concepts, namely the relative difficulty and the driver loss. The former helps Difficulty-Net take other classes into account when calculating difficulty of a class, while the latter is indispensable for guiding the learning to a meaningful direction. Extensive experiments on popular long-tailed datasets demonstrated the effectiveness of the proposed method, and it achieved state-of-the-art performance on multiple long-tailed datasets. Code is available at https://github.com/hitachi-rd-cv/Difficulty_Net. Saptarshi Sinha, Hiroki Ohashi |
WACV | 1 |
| 2022 | Class-Difficulty Based Methods for Long-Tailed Visual RecognitionabstractAbstract Long-tailed datasets are very frequently encountered in real-world use cases where few classes or categories (known as majority or head classes) have higher number of data samples compared to the other classes (known as minority or tail classes). Training deep neural networks on such datasets gives results biased towards the head classes. So far, researchers have come up with multiple weighted loss and data re-sampling techniques in efforts to reduce the bias. However, most of such techniques assume that the tail classes are always the most difficult classes to learn and therefore need more weightage or attention. Here, we argue that the assumption might not always hold true. Therefore, we propose a novel approach to dynamically measure the instantaneous difficulty of each class during the training phase of the model. Further, we use the difficulty measures of each class to design a novel weighted loss technique called ‘class-wise difficulty based weighted (CDB-W) loss’ and a novel data sampling technique called ‘class-wise difficulty based sampling (CDB-S)’. To verify the wide-scale usability of our CDB methods, we conducted extensive experiments on multiple tasks such as image classification, object detection, instance segmentation and video-action classification. Results verified that CDB-W loss and CDB-S could achieve state-of-the-art results on many class-imbalanced datasets such as ImageNet-LT, LVIS and EGTEA, that resemble real-world use cases. Saptarshi Sinha, Hiroki Ohashi, Katsuyuki Nakamura |
Int. J. Comput. Vis. | 1 |
| 2021 | Network approach to mutagenesis sheds insight on phage resistance in mycobacteriaabstractMOTIVATION: A rigorous yet general mathematical approach to mutagenesis, especially one capable of delivering systems-level perspectives would be invaluable. Such systems-level understanding of phage resistance is also highly desirable for phage-bacteria interactions and phage therapy research. Independently, the ability to distinguish between two graphs with a set of common or identical nodes and identify the implications thereof, is important in network science. RESULTS: Herein, we propose a measure called shortest path alteration fraction (SPAF) to compare any two networks by shortest paths, using sets. When SPAF is one, it can identify node pairs connected by at least one shortest path, which are present in either network but not both. Similarly, SPAF equalling zero identifies identical shortest paths, which are simultaneously present between a node pair in both networks. We study the utility of our measure theoretically in five diverse microbial species, to capture reported effects of well-studied mutations and predict new ones. We also scrutinize the effectiveness of our procedure through theoretical and experimental tests on Mycobacterium smegmatis mc2155 and by generating a mutant of mc2155, which is resistant to mycobacteriophage D29. This mutant of mc2155, which is resistant to D29 exhibits significant phenotypic alterations. Whole-genome sequencing identifies mutations, which cannot readily explain the observed phenotypes. Exhaustive analyses of protein-protein interaction network of the mutant and wild-type, using the machinery of topological metrics and differential networks does not yield a clear picture. However, SPAF coherently identifies pairs of proteins at the end of a subset of shortest paths, from amongst hundreds of thousands of viable shortest paths in the networks. The altered functions associated with the protein pairs are strongly correlated with the observed phenotypes. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Saptarshi Sinha, Sourabh Samaddar, Sujoy K. Das Gupta, Soumen Roy |
Bioinform. | 1 |
| 2020 | Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
Saptarshi Sinha, Hiroki Ohashi, Katsuyuki Nakamura |
ACCV (6) | 1 |