EDBT 2026 Demo / reviewers in the wild / expert
Sameer Dharur
dblp:277/0795
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-7131-4539ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 18% Question answering and dialogue systems · 18% Segmentation and scene understanding · 18% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › multimodal question answering
egocentric question answering |
0.6 | 1 | 2022 | Episodic Memory Question Answering · CVPR 2022 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.6 | 1 | 2022 | Episodic Memory Question Answering · CVPR 2022 |
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | Episodic Memory Question Answering · CVPR 2022 |
Robotics › Robot navigation and mapping
embodied AI simulation |
0.5 | 1 | 2021 | Habitat 2.0: Training Home Assistants to Rearrange their Habitat · NeurIPS 2021 |
Robotics › Robot manipulation
mobile manipulation |
0.5 | 1 | 2021 | Habitat 2.0: Training Home Assistants to Rearrange their Habitat · NeurIPS 2021 |
Computer vision › Video understanding and tracking
egocentric video understanding |
0.2 | 1 | 2022 | Episodic Memory Question Answering · CVPR 2022 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.1 | 1 | 2021 | Habitat 2.0: Training Home Assistants to Rearrange their Habitat · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
question grounding · 0.6allocentric top-down semantic feature map · 0.6hierarchical RL · 0.5deep reinforcement learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Modality Drop-Out for Multimodal Device Directed Speech Detection Using Verbal and Non-Verbal FeaturesabstractDevice-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background speech. State-of-the-art DDSD systems use verbal cues, e.g acoustic, text and/or automatic speech recognition system (ASR) features, to classify speech as device-directed or otherwise, and often have to contend with one or more of these modalities being unavailable when deployed in real-world settings. In this paper, we investigate fusion schemes for DDSD systems that can be made more robust to missing modalities. Concurrently, we study the use of non-verbal cues, specifically prosody features, in addition to verbal cues for DDSD. We present different approaches to combine scores and embeddings from prosody with the corresponding verbal cues, finding that prosody improves DDSD performance by upto 8.5% in terms of false acceptance rate (FA) at a given fixed operating point via non-linear intermediate fusion, while our use of modality dropout techniques improves the performance of these models by 7.4% in terms of FA when evaluated with missing modalities during inference time. Gautam Krishna, Sameer Dharur, Ognjen Rudovic, Pranay Dighe, Saurabh Adya, Ahmed Hussen Abdelaziz, Ahmed H. Tewfik |
ICASSP | 2 |
| 2024 | Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Ognjen Rudovic, Sameer Dharur, Florian Pesce, Gautam Krishna, Aswin Sivaraman, Jack Berkowitz, Ahmed Hussen Abdelaziz, Saurabh Adya, Ahmed H. Tewfik |
INTERSPEECH | 3 |
| 2022 | Episodic Memory Question AnsweringabstractEgocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human communicates with an AI agent powering such a device by asking questions (e.g., “where did you last see my keys?”). In order to succeed at this task, the egocentric AI assistant must (1) construct semantically rich and efficient scene memories that encode spatio-temporal infor-mation about objects seen during the tour and (2) possess the ability to understand the question and ground its answer into the semantic memory representation. Towards that end, we introduce (1) a new task - Episodic Memory Question Answering (EMQA) wherein an egocentric AI assistant is provided with a video sequence (the tour) and a question as an input and is asked to localize its answer to the question within the tour, (2) a dataset of grounded questions designed to probe the agent's spatio-temporal understanding of the tour, and (3) a model for the task that encodes the scene as an allocentric, top-down semantic feature map and grounds the question into the map to localize the answer. We show that our choice of episodic scene memory outperforms naive, off-the-shelf solutions for the task as well as a host of very competitive baselines and is robust to noise in depth, pose as well as camera jitter. Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai, Mukul Khanna, Dhruv Batra, Devi Parikh |
CVPR | 2 |
| 2021 | SOrT-ing VQA Models : Contrastive Gradient Learning for Improved ConsistencyabstractSameer Dharur, Purva Tendulkar, Dhruv Batra, Devi Parikh, Ramprasaath R. Selvaraju. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Sameer Dharur, Purva Tendulkar, Dhruv Batra, Devi Parikh, Ramprasaath R. Selvaraju |
NAACL-HLT | 1 |
| 2021 | Habitat 2.0: Training Home Assistants to Rearrange their HabitatabstractWe introduce Habitat 2.0 (H2.0), a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios. We make comprehensive contributions to all levels of the embodied AI stack – data, simulation, and benchmark tasks. Specifically, we present: (i) ReplicaCAD: an artist-authored, annotated, reconfigurable 3D dataset of apartments (matching real spaces) with articulated objects (e.g. cabinets and drawers that can open/close); (ii) H2.0: a high-performance physics-enabled 3D simulator with speeds exceeding 25,000 simulation steps per second (850x real-time) on an 8-GPU node, representing 100x speed-ups over prior work; and, (iii) Home Assistant Benchmark (HAB): a suite of common tasks for assistive robots (tidy the house, stock groceries, set the table) that test a range of mobile manipulation capabilities. These large-scale engineering contributions allow us to systematically compare deep reinforcement learning (RL) at scale and classical sense-plan-act (SPA) pipelines in long-horizon structured tasks, with an emphasis on generalization to new objects, receptacles, and layouts. We find that (1) flat RL policies struggle on HAB compared to hierarchical ones; (2) a hierarchy with independent skills suffers from ‘hand-off problems’, and (3) SPA pipelines are more brittle than RL policies. Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, John M. Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel X. Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, Dhruv Batra |
NeurIPS | 13 |