Rohan Choudhury

dblp:234/6273 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0004-8307-8395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 28% Efficient and distributed learning · 21% Face, body and person analysis · 18%
Human-computer interaction and pervasive computing
2 papers
Human-robot interaction · 79% Immersive interaction · 21%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video question answering
1.822026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Video Question Answering with Procedural Programs · ECCV (38) 2024
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding
1.012026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
token pruning
0.812024
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024
Machine learning › Efficient and distributed learning
token reduction
0.812024
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer › temporal transformer
video transformer
0.812024
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization · NeurIPS 2024
Computer vision › Face, body and person analysis
human pose estimation
0.712023
TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023
Computer vision › Face, body and person analysis › human pose estimation
human pose forecasting
0.712023
TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking
0.712023
TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023
Computer vision › 3D vision › pose estimation › multi-view pose estimation
multi-view 3d pose estimation
0.712023
TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting · ICCV 2023
Human-robot interaction
robot learning
0.412019
On the Utility of Model Learning in HRI · HRI 2019
Human-robot interaction › cognitive human-robot interaction
theory of mind
0.412019
On the Utility of Model Learning in HRI · HRI 2019
Natural language and speech › Language models and text generation
large language model evaluation
0.312026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Computer vision › Video understanding and tracking › deep video understanding
video reasoning
0.212024
Video Question Answering with Procedural Programs · ECCV (38) 2024
Immersive interaction › virtual reality
virtual reality simulation
0.212024
JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions · ICRA 2024
Human-robot interaction
human modeling
0.112019
On the Utility of Model Learning in HRI · HRI 2019

Methods — techniques the papers use, named apart from their topics

human-in-the-loop simulation · 1.5multimodal benchmark · 1.0human baseline · 1.0run-length encoding · 0.8procedural programs · 0.8positional encoding · 0.8neural module networks · 0.8spatiotemporal representation learning · 0.7recurrent neural network · 0.7reinforcement learning · 0.4model-free learning · 0.4
YearPublicationVenuePosition
2026 MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
abstract
We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below.
Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni
AAAI11
2024 Video Question Answering with Procedural Programs
Rohan Choudhury, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni
ECCV (38)1
2024 JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions
abstract
Developing autonomous vehicles that can safely interact with pedestrians requires large amounts of pedestrian and vehicle data in order to learn accurate pedestrian-vehicle interaction models. However, gathering data that include crucial but rare scenarios - such as pedestrians jaywalking into heavy traffic - can be costly and unsafe to collect. We propose a virtual reality human-in-the-loop simulator, JaywalkerVR, to obtain vehicle-pedestrian interaction data to address these challenges. Our system enables efficient, affordable, and safe collection of long-tail pedestrian-vehicle interaction data. Using our proposed simulator, we create a high-quality dataset with vehicle-pedestrian interaction data from safety critical scenarios called CARLA-VR. The CARLA-VR dataset addresses the lack of long-tail data samples in commonly used real world autonomous driving datasets. We demonstrate that models trained with CARLA-VR improve displacement error and collision rate by 10.7% and 4.9%, respectively, and are more robust in rare vehicle-pedestrian scenarios.
Kenta Mukoya, Erica Weng, Rohan Choudhury, Kris Makoto Kitani
ICRA3
2024 Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
abstract
Video transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We present Run-Length Tokenization (RLT), a simple approach to speed up video transformers inspired by run-length encoding for data compression. RLT efficiently finds and removes `runs' of patches that are repeated over time before model inference, then replaces them with a single patch and a positional encoding to represent the resulting token's new length. Our method is content-aware, requiring no tuning for different datasets, and fast, incurring negligible overhead. RLT yields a large speedup in training, reducing the wall-clock time to fine-tune a video transformer by 30% while matching baseline model performance. RLT also works without training, increasing model throughput by 35% with only 0.1% drop in accuracy. RLT speeds up training at 30 FPS by more than 100%, and on longer video datasets, can reduce the token count by up to 80\%. Our project page is at rccchoudhury.github.io/projects/rlt.
Rohan Choudhury, Guanglei Zhu, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni
NeurIPS1
2023 TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting
abstract
Existing volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a robust spatiotemporal representation, improving pose accuracy while also tracking and forecasting human pose. We significantly reduce computation compared to the state-of-the-art by recurrently computing per-person 2D pose features, fusing both spatial and temporal information into a single representation. In doing so, our model is able to use spatiotemporal context to predict more accurate human poses without sacrificing efficiency. We further use this representation to track human poses over time as well as predict future poses. Finally, we demonstrate that our model is able to generalize across datasets without scene-specific fine-tuning. TEMPO achieves 10% better MPJPE with a 33× improvement in FPS compared to TesseTrack on the challenging CMU Panoptic Studio dataset. Our code and demos are available at https://rccchoudhury.github.io/tempo2023/.
Rohan Choudhury, Kris Makoto Kitani, László A. Jeni
ICCV1
2019 On the Utility of Model Learning in HRI
abstract
Fundamental to robotics is the debate between model-based and model-free learning: should the robot build an explicit model of the world, or learn a policy directly? In the context of HRI, part of the world to be modeled is the human. One option is for the robot to treat the human as a black box and learn a policy for how they act directly. But it can also model the human as an agent, and rely on a “theory of mind” to guide or bias the learning (grey box). We contribute a characterization of the performance of these methods under the optimistic case of having an ideal theory of mind, as well as under different scenarios in which the assumptions behind the robot's theory of mind for the human are wrong, as they inevitably will be in practice. We find that there is a significant sample complexity advantage to theory of mind methods and that they are more robust to covariate shift, but that when enough interaction data is available, black box approaches eventually dominate.
Rohan Choudhury, Dylan Hadfield-Menell, Anca D. Dragan
HRI1