EDBT 2026 Demo / reviewers in the wild / expert
Chao Huang 0033
dblp:18/4087-33
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-1469-1020ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal PromptingabstractIn this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data. Yunlong Tang 0002, Jing Bi 0002, Chao Huang 0033, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Pinxin Liu, Mingqian Feng, Junjia Guo, Luchuan Song, Ali Vosoughi, Jinxi He, Zeliang Zhang 0001, Jiebo Luo 0001, Chenliang Xu |
AAAI | 3 |
| 2026 | High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
Chao Huang 0033, Susan Liang, Yapeng Tian, Chenliang Xu |
Int. J. Comput. Vis. | 1 |
| 2026 | Video Understanding With Large Language Models: A SurveyabstractWith the rapid growth of online video platforms and the escalating volume of video content, the need for proficient video understanding tools has increased significantly. Given the remarkable capabilities of large language models (LLMs) in language and multimodal tasks, this survey provides a detailed overview of recent advances in video understanding that harness the power of LLMs (Vid-LLMs). The emergent capabilities of Vid-LLMs are surprisingly advanced, particularly their ability for open-ended multi-granularity (abstract, temporal, and spatiotemporal) reasoning combined with common-sense knowledge, suggesting a promising path for future video understanding. We examine the unique characteristics and capabilities of Vid-LLMs, categorizing the approaches into three main types:Video Analyzer × LLM, Video Embedder × LLM, and (Analyzer + Embedder) × LLM. We identify five subtypes based on the functions of LLMs in Vid-LLMs:LLMas Summarizer,LLMas Manager,LLMas Text Decoder,LLMas Regressor, andLLMas Hidden Layer. This survey also presents a comprehensive study of the tasks, datasets, benchmarks, and evaluation methods for Vid-LLMs. Additionally, it explores the extensive applications of Vid-LLMs in various domains, highlighting their remarkable scalability and versatility in real-world video understanding challenges. Additionally, it summarizes the limitations of existing Vid-LLMs and outlines directions for future research. For more information, readers are encouraged to visit the repository at https://github.com/yunlong10/Awesome-LLMs-for-Video-Understanding. Yunlong Tang 0002, Jing Bi 0002, Siting Xu, Luchuan Song, Susan Liang, Teng Wang 0007, Daoan Zhang, Jie An 0002, Rongyi Zhu, Ali Vosoughi, Chao Huang 0033, Zeliang Zhang 0001, Pinxin Liu, Mingqian Feng, Feng Zheng 0001, Jianguo Zhang 0001, Ping Luo 0002, Jiebo Luo 0001, Chenliang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 12 |
| 2025 | VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?abstractThe advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multi-modal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessment of their ability to understand video compositions, the nuanced interpretation of how visual elements combine and interact within highly compiled video contexts. We introduce VidComposition, a new benchmark specifically designed to evaluate the video composition understanding capabilities of MLLMs using carefully curated compiled videos and cinematic-level annotations. VidComposition includes 982 videos with 1706 multiple-choice questions, covering various compositional aspects such as camera movement, angle, shot size, narrative structure, character actions and emotions, etc. Our comprehensive evaluation of 33 open-source and proprietary MLLMs reveals a significant performance gap between human and model capabilities. This highlights the limitations of current MLLMs in understanding complex, compiled video compositions and offers insights into areas for further improvement. Our benchmark is publicly available at https://yunlong10.github.io/VidComposition/. Yunlong Tang 0002, Junjia Guo, Hang Hua, Susan Liang, Mingqian Feng, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Zeliang Zhang 0001, Pooyan Fazli, Chenliang Xu |
CVPR | 8 |
| 2025 | Learning to Highlight Audio by Watching MoviesabstractRecent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset—THE MUDDY MIX DATASET, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process—separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/. Chao Huang 0033, Ruohan Gao, J. M. F. Tsang, Jan Kurcius, Cagdas Bilen, Chenliang Xu, Sanjeel Parekh |
CVPR | 1 |
| 2025 | $\pi$-AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?
Susan Liang, Chao Huang 0033, Yunlong Tang 0002, Zeliang Zhang 0001, Chenliang Xu |
ICCV | 2 |
| 2025 | ZeroSep: Separate Anything in Audio with Zero TrainingabstractAudio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the immense variability and open-set nature of real-world acoustic scenes. Inspired by the success of generative foundation models, we investigate whether pre-trained text-guided audio diffusion models can overcome these limitations. We make a surprising discovery: zero-shot source separation can be achieved purely through a pre-trained text-guided audio diffusion model under the right configuration. Our method, named ZeroSep, works by inverting the mixed audio into the diffusion model's latent space and then using text conditioning to guide the denoising process to recover individual sources. Without any task-specific training or fine-tuning, ZeroSep repurposes the generative diffusion model for a discriminative separation task and inherently supports open-set scenarios through its rich textual priors. ZeroSep is compatible with a variety of pre-trained text-guided audio diffusion backbones and delivers strong separation performance on multiple separation benchmarks, surpassing even supervised methods. Chao Huang 0033, Yuesheng Ma, Junxuan Huang, Susan Liang, Yunlong Tang 0002, Jing Bi 0002, Nima Mesgarani, Chenliang Xu |
NeurIPS | 1 |
| 2025 | Harnessing the Computation Redundancy in ViTs to Boost Adversarial TransferabilityabstractVision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. Many previous studies have observed that adversarial examples crafted on ViTs exhibit higher transferability than those crafted on CNNs, indicating that ViTs contain structural characteristics favorable for transferable attacks. In this work, we take a further step to deeply investigate the role of computational redundancy brought by its unique characteristics in ViTs and its impact on adversarial transferability. Specifically, we identify two forms of redundancy, including the data-level and model-level, that can be harnessed to amplify attack effectiveness. Building on this insight, we design a suite of techniques, including attention sparsity manipulation, attention head permutation, clean token regularization, ghost MoE diversification, and learn to robustify before the attack. A dynamic online learning strategy is also proposed to fully leverage these operations to enhance the adversarial transferability. Extensive experiments on the ImageNet-1k dataset validate the effectiveness of our approach, showing that our methods significantly outperform existing baselines in both transferability and generality across diverse model architectures, including different variants of ViTs and mainstream Vision Large Language Models (VLLMs). Jiani Liu 0013, Zeliang Zhang 0001, Chao Huang 0033, Susan Liang, Yunlong Tang 0002, Chenliang Xu |
NeurIPS | 4 |
| 2025 | MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and RobustnessabstractUnderstanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources are available at https://yunlong10.github.io/MMPerspective/ Yunlong Tang 0002, Pinxin Liu, Mingqian Feng, Zhangyun Tan, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Yunzhong Xiao, Susan Liang, Hang Hua, Ali Vosoughi, Luchuan Song, Zeliang Zhang 0001, Chenliang Xu |
NeurIPS | 6 |
| 2024 | High-Quality Visually-Guided Sound Separation from Diverse Categories
Chao Huang 0033, Susan Liang, Yapeng Tian, Chenliang Xu |
ACCV (6) | 1 |
| 2024 | Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
Susan Liang, Chao Huang 0033, Yapeng Tian, Chenliang Xu |
ACCV (6) | 2 |
| 2024 | Modeling and Driving Human Body Soundfields Through Acoustic Primitives
Chao Huang 0033, Dejan Markovic, Chenliang Xu, Alexander Richard |
ECCV (10) | 1 |
| 2023 | Egocentric Audio-Visual Object LocalizationabstractHumans naturally perceive surrounding scenes by unifying sound and sight from a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric perspective. In this paper, we explore the challenging egocentric audio-visual object localization task and observe that 1) egomotion commonly exists in first-person recordings, even within a short duration; 2) The out-of-view sound components can be created when wearers shift their attention. To address the first problem, we propose a geometryaware temporal aggregation module that handles the egomotion explicitly. The effect of egomotion is mitigated by estimating the temporal geometry transformation and exploiting it to update visual representations. Moreover, we propose a cascaded feature enhancement module to over-come the second issue. It improves cross-modal localization robustness by disentangling visually-indicated audio representation. During training, we take advantage of the naturally occurring audio-visual temporal synchronization as the “free” self-supervision to avoid costly labeling. We also annotate and create the Epic Sounding Object dataset for evaluation purposes. Extensive experiments show that our method achieves state-of-the-art localization performance in egocentric videos and can be generalized to diverse audio-visual scenes. Code is available at https://github.com/WikiChao/Ego-AV-Loc. Chao Huang 0033, Yapeng Tian, Chenliang Xu |
CVPR | 1 |
| 2023 | AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisabstractCan machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task---real-world audio-visual scene synthesis---and a first-of-its-kind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset. Notably, we refer readers to view our demo videos for convincing comparisons. Susan Liang, Chao Huang 0033, Yapeng Tian, Chenliang Xu |
NeurIPS | 2 |
| 2022 | How to Prepare for the Next Pandemic - Investigation of Correlation Between Food Prices and COVID-19 From Global and Local PerspectivesabstractThe coronavirus disease (COVID-19) has caused enormous disruptions to not only the United States, but also the global economy. Due to the pandemic, issues in the supply chain and concerns about food shortage drove up the food prices. According to the U.S. Bureau of Labor Statistics, the prices for food increased 4.1% and 3.7% over the year ended in August 2020 and August 2021, respectively, while the amount of annual increase in the food prices prior to the COVID-19 pandemic is less than 2.0%. Previous studies show that such kinds of exogenous disasters, including the 2011 Tohoku Earthquake, 9/11 terrorist attacks, and major infectious diseases, and the resulted unusual food prices often led to subsequent changes in people’s consumption behaviors. We hypothesize that the COVID-19 pandemic causes food price changes and the price changes alter people’s grocery shopping behaviors as well. To thoroughly explore this, we formulate our analysis from two different perspectives, by collecting data both globally, from China, Japan, United Kingdom, and United States, and locally, from different groups of people inside the US. In particular, we analyze the trends between food prices and COVID-19 as well as between food prices and spending, aiming to find out their correlations and the lessons for preparing the next pandemic. Chao Huang 0033, Jiebo Luo 0001 |
IEEE Big Data | 2 |