VLDB 2026 Research / reviewers in the wild / expert
Ali Vosoughi
dblp:348/7127
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-1014-2937ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 48% Trustworthy machine learning · 16% 3D vision · 12% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.6 | 2 | 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025 EAGLE: Egocentric AGgregated Language-video Engine · ACM Multimedia 2024 |
Computer vision › Vision and language
video captioning |
1.0 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Computer vision › Vision and language › vision-language model › prompt learning
vision-language model prompting |
1.0 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Computer vision › 3D vision › multi-view geometry › camera geometry
perspective geometry |
0.9 | 1 | 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.9 | 1 | 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
0.8 | 1 | 2024 | Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQA · IEEE Trans. Multim. 2024 |
Computer vision › Video understanding and tracking
egocentric video understanding |
0.8 | 1 | 2024 | EAGLE: Egocentric AGgregated Language-video Engine · ACM Multimedia 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQA · IEEE Trans. Multim. 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 2 | 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025 Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQA · IEEE Trans. Multim. 2024 |
Computer vision › Video understanding and tracking
video object segmentation |
0.3 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Computer vision › 3D vision
spatial consistency |
0.3 | 1 | 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.2 | 1 | 2024 | EAGLE: Egocentric AGgregated Language-video Engine · ACM Multimedia 2024 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation |
0.2 | 1 | 2024 | Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQA · IEEE Trans. Multim. 2024 |
Methods — techniques the papers use, named apart from their topics
chain-of-thought prompting · 1.9temporal analysis · 1.0segment anything · 1.0multimodal prompting · 1.0benchmark construction · 0.9multimodal large language model · 0.8instruction tuning · 0.8counterfactual inference · 0.8causal inference · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal PromptingabstractIn this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data. Yunlong Tang 0002, Jing Bi 0002, Chao Huang 0033, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Pinxin Liu, Mingqian Feng, Junjia Guo, Luchuan Song, Ali Vosoughi, Jinxi He, Zeliang Zhang 0001, Jiebo Luo 0001, Chenliang Xu |
AAAI | 14 |
| 2026 | Video Understanding With Large Language Models: A SurveyabstractWith the rapid growth of online video platforms and the escalating volume of video content, the need for proficient video understanding tools has increased significantly. Given the remarkable capabilities of large language models (LLMs) in language and multimodal tasks, this survey provides a detailed overview of recent advances in video understanding that harness the power of LLMs (Vid-LLMs). The emergent capabilities of Vid-LLMs are surprisingly advanced, particularly their ability for open-ended multi-granularity (abstract, temporal, and spatiotemporal) reasoning combined with common-sense knowledge, suggesting a promising path for future video understanding. We examine the unique characteristics and capabilities of Vid-LLMs, categorizing the approaches into three main types:Video Analyzer × LLM, Video Embedder × LLM, and (Analyzer + Embedder) × LLM. We identify five subtypes based on the functions of LLMs in Vid-LLMs:LLMas Summarizer,LLMas Manager,LLMas Text Decoder,LLMas Regressor, andLLMas Hidden Layer. This survey also presents a comprehensive study of the tasks, datasets, benchmarks, and evaluation methods for Vid-LLMs. Additionally, it explores the extensive applications of Vid-LLMs in various domains, highlighting their remarkable scalability and versatility in real-world video understanding challenges. Additionally, it summarizes the limitations of existing Vid-LLMs and outlines directions for future research. For more information, readers are encouraged to visit the repository at https://github.com/yunlong10/Awesome-LLMs-for-Video-Understanding. Yunlong Tang 0002, Jing Bi 0002, Siting Xu, Luchuan Song, Susan Liang, Teng Wang 0007, Daoan Zhang, Jie An 0002, Rongyi Zhu, Ali Vosoughi, Chao Huang 0033, Zeliang Zhang 0001, Pinxin Liu, Mingqian Feng, Feng Zheng 0001, Jianguo Zhang 0001, Ping Luo 0002, Jiebo Luo 0001, Chenliang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and RobustnessabstractUnderstanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources are available at https://yunlong10.github.io/MMPerspective/ Yunlong Tang 0002, Pinxin Liu, Mingqian Feng, Zhangyun Tan, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Yunzhong Xiao, Susan Liang, Hang Hua, Ali Vosoughi, Luchuan Song, Zeliang Zhang 0001, Chenliang Xu |
NeurIPS | 11 |
| 2024 | Learning Audio Concepts from Counterfactual Natural LanguageabstractConventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from raw audio-text pairs describing audio in natural language. Despite recent advancements, there is little exploration of systematic methods to train models for recognizing sound events and sources in alternative scenarios, such as distinguishing fireworks from gunshots at outdoor events in similar situations. This study introduces causal reasoning and counterfactual analysis in the audio domain. We use counterfactual instances and include them in our model across different aspects. Our model considers acoustic characteristics and sound source information from human-annotated reference texts. To validate the effectiveness of our model, we conducted pre-training utilizing multiple audio captioning datasets. We then evaluate with several common downstream tasks, demonstrating the merits of the proposed method as one of the first works leveraging counterfactual information in audio domain. Specifically, the top-1 accuracy in open-ended language-based audio retrieval task increased by more than 43%. Ali Vosoughi, Luca Bondi, Ho-Hsiang Wu, Chenliang Xu |
ICASSP | 1 |
| 2024 | EAGLE: Egocentric AGgregated Language-video EngineabstractThe rapid evolution of egocentric video analysis brings new insights into understanding human activities and intentions from a first-person perspective. Despite this progress, the fragmentation in tasks like action recognition, procedure learning, and moment retrieval, \etc, coupled with inconsistent annotations and isolated model development, hinders a holistic interpretation of video content. In response, we introduce the EAGLE (Egocentric AGgregated Language-video Engine) model and the EAGLE-400K dataset to provide a unified framework that integrates various egocentric video understanding tasks. EAGLE-400K, the \textit{first} large-scale instruction-tuning dataset tailored for egocentric video, features 400K diverse samples to enhance a broad spectrum of tasks from activity recognition to procedure knowledge learning. Moreover, EAGLE, a strong video multimodal large language model (MLLM), is designed to effectively capture both spatial and temporal information. In addition, we propose a set of evaluation metrics designed to facilitate a thorough assessment of MLLM for egocentric video understanding. Our extensive experiments demonstrate EAGLE's superior performance over existing models, highlighting its ability to balance task-specific understanding with holistic video interpretation. With EAGLE, we aim to pave the way for research opportunities and practical applications in real-world scenarios. Jing Bi 0002, Yunlong Tang 0002, Luchuan Song, Ali Vosoughi, Chenliang Xu |
ACM Multimedia | 4 |
| 2024 | Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQAabstractTo increase the generalization capability of VQA systems, many recent studies have tried to de-bias spurious language or vision associations that shortcut the question or image to the answer. Despite these efforts, the literature fails to address the confounding effect of vision and language simultaneously. As a result, when they reduce bias learned from one modality, they usually increase bias from the other. In this paper, we first model a confounding effect that causes language and vision bias simultaneously, then propose a counterfactual inference to remove the influence of this effect. The model trained in this strategy can concurrently and efficiently reduce vision and language bias. To the best of our knowledge, this is the first work to reduce biases resulting from confounding effects of vision and language in VQA, leveraging causal explain-away relations. We accompany our method with an explain-away strategy, pushing the accuracy of the questions with numerical answers results compared to existing methods that have been an open problem. The proposed method outperforms the state-of-the-art methods in VQA-CP v2 datasets. R2: Providing brief insights into the experimental setup and results would add valuable context for readers. In response to R2, we released the code and documentation for the implementation as follows. Our codes are available at https://github.com/ali-vosoughi/PW-VQA. Ali Vosoughi, Shijian Deng, Songyang Zhang 0004, Yapeng Tian, Chenliang Xu, Jiebo Luo 0001 |
IEEE Trans. Multim. | 1 |