Pooyan Fazli

dblp:35/3763 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-2625-8216ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
abstract
The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing efficiency methods mainly target inference via token reduction or merging, offering limited benefits during training. We introduce ReGATE (Reference-Guided Adaptive Token Elision), an adaptive token pruning method for accelerating MLLM training. ReGATE adopts a teacher-student framework, in which a frozen teacher LLM provides per-token guidance losses that are fused with an exponential moving average of the student's difficulty estimates. This adaptive scoring mechanism dynamically selects informative tokens while skipping redundant ones in the forward pass, substantially reducing computation without altering the model architecture. Across three representative MLLMs, ReGATE matches the peak accuracy of standard training on MVBench up to 2$\times$ faster, using only 38% of the tokens. With extended training, it even surpasses the baseline across multiple multimodal benchmarks, cutting total token usage by over 41%.
Chaoyu Li, Yogesh Kulkarni, Pooyan Fazli
ACL (1)3
2026 ChartQA-X: Generating Explanations for Visual Chart Reasoning
abstract
The ability to explain complex information from chart images is vital for effective data-driven decision-making. In this work, we address the challenge of generating detailed explanations alongside answering questions about charts. We present ChartQA-X, a comprehensive dataset comprising 30,799 chart samples across four chart types, each paired with contextually relevant questions, answers, and explanations. Explanations are generated and selected based on metrics such as faithfulness, informativeness, coherence, and perplexity. Our human evaluation with 245 participants shows that model-generated explanations in ChartQA-X surpass human-written explanations in accuracy and logic and are comparable in terms of clarity and overall quality. Moreover, models fine-tuned on ChartQA-X show substantial improvements across various metrics, including absolute gains of up to 24.57 points in explanation quality, 18.96 percentage points in question-answering accuracy, and 14.75 percentage points on unseen benchmarks for the same task. By integrating explanatory narratives with answers, our approach enables agents to convey complex visual information more effectively, improving comprehension and greater trust in the generated responses.
Shamanthak Hegde, Pooyan Fazli, Hasti Seifi
WACV2
2025 Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
abstract
Audio descriptions (AD) make videos accessible for blind and low vision (BLV) users by describing visual elements that cannot be understood from the main audio track. AD created by professionals or novice describers is time-consuming and offers little customization or control to BLV viewers on description length and content and when they receive it. To address this gap, we explore user-driven AI-generated descriptions, enabling BLV viewers to control both the timing and level of detail of the descriptions they receive. In a study, 20 BLV participants activated audio descriptions for seven different video genres with two levels of detail: concise and detailed. Our findings reveal differences in the preferred frequency and level of detail of ADs for different videos, participants' sense of control with this style of AD delivery, and its limitations. We discuss the implications of these findings for the development of future AD tools for BLV users.
Maryam Cheema, Hasti Seifi, Pooyan Fazli
Conference on Designing Interactive Systems3
2025 DescribePro: Collaborative Audio Description with Human-AI Interaction
abstract
with 18 describers (9 professionals and 9 novices) using quantitative and qualitative methods. Results show that AI support reduces repetitive work while helping professionals preserve their stylistic choices and easing the cognitive load for novices. Collaborative tags and variations show potential for providing customizations, version control, and training new describers. These findings highlight the potential of collaborative, AI-assisted tools to enhance and scale AD authorship.
Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, Hasti Seifi
ASSETS4
2025 VideoA11y: Method and Dataset for Accessible Video Description
abstract
Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human annotations within training datasets, resulting in descriptions that do not fully meet BLV users' needs. To address this gap, we introduce VideoA11y, an approach that leverages multimodal large language models (MLLMs) and video accessibility guidelines to generate descriptions tailored for BLV individuals. Using this method, we have curated VideoA11y-40K, the largest and most comprehensive dataset of 40,000 videos described for BLV users. Rigorous experiments across 15 video categories, involving 347 sighted participants, 40 BLV participants, and seven professional describers, showed that VideoA11y descriptions outperform novice human annotations and are comparable to trained human annotations in clarity, accuracy, objectivity, descriptiveness, and user satisfaction. We evaluated models on VideoA11y-40K using both standard and custom metrics, demonstrating that MLLMs fine-tuned on this dataset produce high-quality accessible descriptions. Code and dataset are available at https://people-robots.github.io/VideoA11y/.
Chaoyu Li, Sid Padmanabhuni, Maryam Cheema, Hasti Seifi, Pooyan Fazli
CHI5
2025 VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
abstract
The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multi-modal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessment of their ability to understand video compositions, the nuanced interpretation of how visual elements combine and interact within highly compiled video contexts. We introduce VidComposition, a new benchmark specifically designed to evaluate the video composition understanding capabilities of MLLMs using carefully curated compiled videos and cinematic-level annotations. VidComposition includes 982 videos with 1706 multiple-choice questions, covering various compositional aspects such as camera movement, angle, shot size, narrative structure, character actions and emotions, etc. Our comprehensive evaluation of 33 open-source and proprietary MLLMs reveals a significant performance gap between human and model capabilities. This highlights the limitations of current MLLMs in understanding complex, compiled video compositions and offers insights into areas for further improvement. Our benchmark is publicly available at https://yunlong10.github.io/VidComposition/.
Yunlong Tang 0002, Junjia Guo, Hang Hua, Susan Liang, Mingqian Feng, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Zeliang Zhang 0001, Pooyan Fazli, Chenliang Xu
CVPR11
2025 VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
abstract
Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate inaccurate or misleading content, remains underexplored in the video domain. Building on the observation that MLLM visual encoders often fail to distinguish visually different yet semantically similar video pairs, we introduce VIDHALLUC, the largest benchmark designed to examine hallucinations in MLLMs for video understanding. It consists of 5,002 videos, paired to highlight cases prone to hallucinations. VIDHALLUC assesses hallucinations across three critical dimensions: (1) action, (2) temporal sequence, and (3) scene transition. Comprehensive testing shows that most MLLMs are vulnerable to hallucinations across these dimensions. Furthermore, we propose DINO-HEAL, a trainingfree method that reduces hallucinations by incorporating spatial saliency from DINOv2 to reweight visual features during inference. Our results show that DINO-HEAL consistently improves performance on VIDHALLUC, achieving an average improvement of 3.02% in mitigating hallucinations across all tasks. Both the VIDHALLUC benchmark and DINO-HEAL code are available at https://people-robots.github.io/vidhalluc.
Chaoyu Li, Eun Woo Im, Pooyan Fazli
CVPR3
2025 VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
abstract
Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity.To address these limitations, we introduce VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries), a framework that enhances Video-LLMs through targeted preference optimization.VideoPASTA trains models to distinguish accurate video representations from carefully crafted adversarial examples that deliberately violate spatial, temporal, or cross-frame relationships.With only 7,020 preference pairs and Direct Preference Optimization, VideoPASTA enables models to learn robust representations that capture finegrained spatial details and long-range temporal dynamics.Experiments demonstrate that VideoPASTA is model agnostic and significantly improves performance, for example, achieving gains of up to +3.8 percentage points on LongVideoBench, +4.1 on VideoMME, and +4.0 on MVBench, when applied to various state-of-the-art Video-LLMs.These results demonstrate that targeted alignment, rather than massive pretraining or architectural modifications, effectively addresses core videolanguage challenges.Notably, VideoPASTA achieves these improvements without any human annotation or captioning, relying solely on 32-frame sampling.This efficiency makes our approach a scalable plug-and-play solution that seamlessly integrates with existing models while preserving their original capabilities.
Yogesh Kulkarni, Pooyan Fazli
EMNLP2
2025 RCareGen: An Interface for Scene and Task Generation in RCareWorld
abstract
This late-breaking report presents RCareGen, a graphical interface that integrates natural language commands with RCare World, a physics simulator for robotic caregiving scenarios. RCareGen has three core modules: (1) a front-end web interface, (2) an LLM-based code generator, and (3) the RCareWorld simulation backend. The front-end web interface enables novice users to input natural language, which is translated into Python code by an LLM-based code generator. This generated code interacts with RCareWorld APIs to run the simulation backend, facilitating scene setup, modifications, simple movements, and human-robot interaction tasks. Additionally, the system supports iterative feedback, allowing users to refine scenes and tasks interactively. By simplifying simulation setup and enhancing task diversity, RCareGen introduces a novel interface that democratizes robot simulation and programming across diverse domains.
Shuaixing Chen, Ruolin Ye, Saurabh Dingwani, Pooyan Fazli, Hasti Seifi, Tapomayukh Bhattacharjee
HRI4
2023 First-Hand Impressions: Charting and Predicting User Impressions of Robot Hands
abstract
Designing robotic hands has been an active area of research and innovation in the last decade. However, little is known about how people perceive robot hands and react to being touched by them. To inform hand design for social robots, we created a database of 73 robot hands and ran two user studies. In the first study, 160 online users rated the hands in our database. Variations in user ratings mostly centered on the perceived Comfortableness , Interestingness , and Industrialness of the hands. In a second lab-based study, users evaluated seven physical hands and had similar ratings to results from the online study. Furthermore, we did not find a significant difference in user ratings before and after the users were touched by the hands. We provide regression models that can predict user ratings from the hand features (e.g., number of fingers) and an online interface for using our robot hand database and predictive models.
Hasti Seifi, Steven A. Vasquez, Hyunyoung Kim 0001, Pooyan Fazli
ACM Trans. Hum. Robot Interact.4
2020 Human-in-the-Loop Machine Learning to Increase Video Accessibility for Visually Impaired and Blind Users
abstract
Video accessibility is crucial for blind and visually impaired individuals for education, employment, and entertainment purposes. However, professional video descriptions are costly and time-consuming. Volunteer-created video descriptions could be a promising alternative, however, they can vary in quality and can be intimidating for novice describers. We developed a Human-in-the-Loop Machine Learning (HILML) approach to video description by automating video text generation and scene segmentation and allowing humans to edit the output. The HILML approach facilitates human-machine collaboration to produce high quality video descriptions while keeping a low barrier to entry for volunteer describers. Our HILML system was significantly faster and easier to use for first-time video describers compared to a human-only control condition with no machine learning assistance. The quality of the video descriptions and understanding of the topic created by the HILML system compared to the human-only condition were rated as being significantly higher by blind and visually impaired users.
Beste F. Yuksel, Pooyan Fazli, Umang Mathur 0002, Vaishali Bisht, Soo Jung Kim 0001, Joshua Junhee Lee, Seung Jung Jin, Yue-Ting Siu, Joshua A. Miele, Ilmi Yoon
Conference on Designing Interactive Systems2
2020 Multi-Robot Task Allocation with Time Window and Ordering Constraints
abstract
The multi-robot task allocation problem comprises task assignment, coalition formation, task scheduling, and routing. We extend the distributed constraint optimization problem (DCOP) formalism to allocate tasks to a team of robots. The tasks have time window and ordering constraints. Each robot creates a simple temporal network to maintain the tasks in its schedule. The proposed layered framework, called L-DCOP, forms efficient coalitions among robots to accomplish the tasks more efficiently as a result of their collective abilities. We conduct extensive experiments to assess the performance of the proposed algorithm and compare it against a benchmark auction-based approach. The results show that L-DCOP increases the task completion rate and task completion frequency by 1.7% and 10.1%, respectively, and reduces the task execution time by 52.5% on average.
Elina Suslova, Pooyan Fazli
IROS2
2019 DeepMoTIon: Learning to Navigate Like Humans
abstract
We present a novel human-aware navigation approach, where the robot learns to mimic humans to navigate safely in crowds. The presented model, referred to as Deep-MoTIon, is trained with pedestrian surveillance data to predict human velocity in the environment. The robot processes LiDAR scans via the trained network to navigate to the target location. We conduct extensive experiments to assess the components of our network and prove their necessity to imitate humans. Our experiments show that DeepMoTIion outperforms all the benchmarks in terms of human imitation, achieving a 24% reduction in time series-based path deviation over the next best approach. In addition, while many other approaches often failed to reach the target, our method reached the target in 100% of the test cases while complying with social norms and ensuring human safety.
Mahmoud Hamandi, Mike D'Arcy, Pooyan Fazli
RO-MAN3
2010 On Multi-Robot Area Coverage
Pooyan Fazli
AAAI1
2010 Complete and robust cooperative robot area coverage with limited range
abstract
We address the problem of multi-robot area coverage and present a new approach in the case where the map of the area and its static obstacles are known and the robots have a limited visibility range. The proposed method starts by locating a set of static guards on the map of the target area and then builds a graph called Reduced-CDT, a new environment representation method based on theConstrained Delaunay Triangulation(CDT).Multi-Prim'sis used to decompose the graph into a forest ofpartial spanning trees(PSTs). Each PST is then modified through a mechanism calledConstrained Spanning Tour(CST) to build a cycle which is then assigned to an individual robot. Subsequently, robots start navigating the cycles and consequently cover the whole area. We show that the proposed approach is complete and robust with respect to robot failure.
Pooyan Fazli, Alireza Davoodi, Philippe Pasquier, Alan K. Mackworth
IROS1
2008 Unsupervised Categorization (Filtering) of Google Images Based on Visual Consistency
Pooyan Fazli, Ara Bedrosian
AAAI1
2007 The UBC Semantic Robot Vision System
Scott Helmer, David Meger, Per-Erik Forssén, Tristram Southey, Sancho McCann, Pooyan Fazli, James J. Little, David G. Lowe
AAAI6