EDBT 2026 Demo / reviewers in the wild / expert
Rohan Choudhury
dblp:234/6273
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0004-8307-8395ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 28% Efficient and distributed learning · 21% Face, body and person analysis · 18% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-robot interaction · 79% Immersive interaction · 21% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
video question answering |
1.8 | 2 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 Video Question Answering with Procedural Programs · ECCV (38) 2024 |
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.8 | 1 | 2024 | Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
token pruning |
0.8 | 1 | 2024 | Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
token reduction |
0.8 | 1 | 2024 | Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › transformer › temporal transformer
video transformer |
0.8 | 1 | 2024 | Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024 |
Computer vision › Face, body and person analysis
human pose estimation |
0.7 | 1 | 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023 |
Computer vision › Face, body and person analysis › human pose estimation
human pose forecasting |
0.7 | 1 | 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023 |
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking |
0.7 | 1 | 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023 |
Computer vision › 3D vision › pose estimation › multi-view pose estimation
multi-view 3d pose estimation |
0.7 | 1 | 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023 |
Human-robot interaction
robot learning |
0.4 | 1 | 2019 | On the Utility of Model Learning in HRI · HRI 2019 |
Human-robot interaction › cognitive human-robot interaction
theory of mind |
0.4 | 1 | 2019 | On the Utility of Model Learning in HRI · HRI 2019 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Video understanding and tracking › deep video understanding
video reasoning |
0.2 | 1 | 2024 | Video Question Answering with Procedural Programs · ECCV (38) 2024 |
Immersive interaction › virtual reality
virtual reality simulation |
0.2 | 1 | 2024 | JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions · ICRA 2024 |
Human-robot interaction
human modeling |
0.1 | 1 | 2019 | On the Utility of Model Learning in HRI · HRI 2019 |
Methods — techniques the papers use, named apart from their topics
human-in-the-loop simulation · 1.5multimodal benchmark · 1.0human baseline · 1.0run-length encoding · 0.8procedural programs · 0.8positional encoding · 0.8neural module networks · 0.8spatiotemporal representation learning · 0.7recurrent neural network · 0.7reinforcement learning · 0.4model-free learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXabstractWe introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below. Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni |
AAAI | 11 |
| 2024 | Video Question Answering with Procedural Programs
Rohan Choudhury, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
ECCV (38) | 1 |
| 2024 | JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle InteractionsabstractDeveloping autonomous vehicles that can safely interact with pedestrians requires large amounts of pedestrian and vehicle data in order to learn accurate pedestrian-vehicle interaction models. However, gathering data that include crucial but rare scenarios - such as pedestrians jaywalking into heavy traffic - can be costly and unsafe to collect. We propose a virtual reality human-in-the-loop simulator, JaywalkerVR, to obtain vehicle-pedestrian interaction data to address these challenges. Our system enables efficient, affordable, and safe collection of long-tail pedestrian-vehicle interaction data. Using our proposed simulator, we create a high-quality dataset with vehicle-pedestrian interaction data from safety critical scenarios called CARLA-VR. The CARLA-VR dataset addresses the lack of long-tail data samples in commonly used real world autonomous driving datasets. We demonstrate that models trained with CARLA-VR improve displacement error and collision rate by 10.7% and 4.9%, respectively, and are more robust in rare vehicle-pedestrian scenarios. Kenta Mukoya, Erica Weng, Rohan Choudhury, Kris Makoto Kitani |
ICRA | 3 |
| 2024 | Don't Look Twice: Faster Video Transformers with Run-Length TokenizationabstractVideo transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We present Run-Length Tokenization (RLT), a simple approach to speed up video transformers inspired by run-length encoding for data compression. RLT efficiently finds and removes `runs' of patches that are repeated over time before model inference, then replaces them with a single patch and a positional encoding to represent the resulting token's new length.
Our method is content-aware, requiring no tuning for different datasets, and fast, incurring negligible overhead.
RLT yields a large speedup in training, reducing the wall-clock time to fine-tune a video transformer by 30% while matching baseline model performance. RLT also works without training, increasing model throughput by 35% with only 0.1% drop in accuracy.
RLT speeds up training at 30 FPS by more than 100%, and on longer video datasets, can reduce the token count by up to 80\%. Our project page is at rccchoudhury.github.io/projects/rlt. Rohan Choudhury, Guanglei Zhu, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
NeurIPS | 1 |
| 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and ForecastingabstractExisting volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a robust spatiotemporal representation, improving pose accuracy while also tracking and forecasting human pose. We significantly reduce computation compared to the state-of-the-art by recurrently computing per-person 2D pose features, fusing both spatial and temporal information into a single representation. In doing so, our model is able to use spatiotemporal context to predict more accurate human poses without sacrificing efficiency. We further use this representation to track human poses over time as well as predict future poses. Finally, we demonstrate that our model is able to generalize across datasets without scene-specific fine-tuning. TEMPO achieves 10% better MPJPE with a 33× improvement in FPS compared to TesseTrack on the challenging CMU Panoptic Studio dataset. Our code and demos are available at https://rccchoudhury.github.io/tempo2023/. Rohan Choudhury, Kris Makoto Kitani, László A. Jeni |
ICCV | 1 |
| 2019 | On the Utility of Model Learning in HRIabstractFundamental to robotics is the debate between model-based and model-free learning: should the robot build an explicit model of the world, or learn a policy directly? In the context of HRI, part of the world to be modeled is the human. One option is for the robot to treat the human as a black box and learn a policy for how they act directly. But it can also model the human as an agent, and rely on a “theory of mind” to guide or bias the learning (grey box). We contribute a characterization of the performance of these methods under the optimistic case of having an ideal theory of mind, as well as under different scenarios in which the assumptions behind the robot's theory of mind for the human are wrong, as they inevitably will be in practice. We find that there is a significant sample complexity advantage to theory of mind methods and that they are more robust to covariate shift, but that when enough interaction data is available, black box approaches eventually dominate. Rohan Choudhury, Dylan Hadfield-Menell, Anca D. Dragan |
HRI | 1 |