Xiulong Liu 0002

dblp:136/3336-2 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-8697-1489ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 21% Representation and self-supervised learning · 20% Video understanding and tracking · 11%
Computer graphics and multimedia
4 papers
Audio and music processing · 87% Multimedia analysis and retrieval · 13%

Topics — the 26 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › sound synthesis
video-to-audio generation
1.222024
Tell What You Hear From What You See - Video to Audio Generation Through Text · NeurIPS 2024
Audeo: Audio Generation for a Silent Performance Video · NeurIPS 2020
Multimedia analysis and retrieval › multimodal learning
cross-modal generation
0.922021
How Does it Sound? · NeurIPS 2021
Audeo: Audio Generation for a Silent Performance Video · NeurIPS 2020
Audio and music processing
music generation
0.922021
How Does it Sound? · NeurIPS 2021
Audeo: Audio Generation for a Silent Performance Video · NeurIPS 2020
Audio and music processing › music generation
video-to-music generation
0.922021
How Does it Sound? · NeurIPS 2021
Audeo: Audio Generation for a Silent Performance Video · NeurIPS 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning
0.912025
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
audio-visual large language model
0.912025
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025
Audio and music processing › room acoustics
room impulse response estimation
0.912025
Hearing Anywhere in Any Environment · CVPR 2025
Audio and music processing
spatial audio
0.912025
Hearing Anywhere in Any Environment · CVPR 2025
Machine learning › Generative modeling
audio generation
0.812024
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation · ICML 2024
Robotics › Robot navigation and mapping › mobile robot navigation › sensor-based navigation
audio-visual navigation
0.812024
CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments · AAAI 2024
Machine learning › Representation and self-supervised learning › multimodal representation learning › cross-modal representation learning
audio-visual representation learning
0.812024
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation · ICML 2024
Machine learning › Generative modeling
multimodal generation
0.812024
Tell What You Hear From What You See - Video to Audio Generation Through Text · NeurIPS 2024
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.812024
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation · ICML 2024
Machine learning › Generative modeling › audio generation
video-to-audio generation
0.812024
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation · ICML 2024
Audio and music processing › audio-language processing
audio captioning
0.812024
Tell What You Hear From What You See - Video to Audio Generation Through Text · NeurIPS 2024
Audio and music processing
sound synthesis
0.812024
Tell What You Hear From What You See - Video to Audio Generation Through Text · NeurIPS 2024
Computer vision › Video understanding and tracking
action recognition
0.412020
PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition · CVPR 2020
Machine learning › Probabilistic and Bayesian machine learning › clustering
sequence clustering
0.412020
PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition · CVPR 2020
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.412020
PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition · CVPR 2020
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412020
PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition · CVPR 2020
Computer vision › 3D vision
3d scene understanding
0.312025
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025
Computer vision › 3D vision
depth estimation
0.312025
Hearing Anywhere in Any Environment · CVPR 2025
Robotics › Robot navigation and mapping › robot mapping
global map building
0.312025
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025
Computer vision › Video understanding and tracking
object tracking
0.312025
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025
Computer vision › 3D vision › depth estimation › panoramic depth estimation
panoramic indoor depth
0.312025
Hearing Anywhere in Any Environment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked token prediction
0.212024
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation · ICML 2024

Methods — techniques the papers use, named apart from their topics

sim-to-real transfer · 1.7geometric feature extraction · 1.7RIR encoding · 1.7transformer · 1.3egocentric spatial track estimation · 0.9coordinate transformation · 0.9trajectory forecasting network · 0.8neural audio codec · 0.8natural language question generation · 0.8multimodal large language model · 0.8large language model · 0.8iterative decoding · 0.8audio tokenizer · 0.8u-net · 0.5skeleton keypoint extraction · 0.5piano-roll representation · 0.4MIDI synthesis · 0.4
YearPublicationVenuePosition
2025 Hearing Anywhere in Any Environment
abstract
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present xRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce AcousticRooms, a new dataset featuring high-fidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four real-world environments, demonstrating the generalizability of our approach and the realism of our dataset.
Xiulong Liu 0002, Anurag Kumar 0003, Paul Calamia, Sebastià Vicenc Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip W. Robinson, Eli Shlizerman, Vamsi K. Ithapu, Ruohan Gao
CVPR1
2025 SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
abstract
3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of carefully curated question–answer pairs probing both directional and distance relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.
Mingfei Chen, Zijun Cui, Xiulong Liu 0002, Jinlin Xiang, Eli Shlizerman
NeurIPS3
2024 CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments
abstract
Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied Navigation framework in which the agent may interact with a human/oracle for solving the task of navigating to an audio goal. Specifically, CAVEN is modeled as a budget-aware partially observable semi-Markov decision process that implicitly learns the uncertainty in the audio-based navigation policy to decide when and how the agent may interact with the oracle. Our CAVEN agent can engage in fully-bidirectional natural language conversations by producing relevant questions and interpret free-form, potentially noisy responses from the oracle based on the audio-visual context. To enable such a capability, CAVEN is equipped with: i) a trajectory forecasting network that is grounded in audio-visual cues to produce a potential trajectory to the estimated goal, and (ii) a natural language based question generation and reasoning network to pose an interactive question to the oracle or interpret the oracle's response to produce navigation instructions. To train the interactive modules, we present a large scale dataset: AVN-Instruct, based on the Landmark-RxR dataset. To substantiate the usefulness of conversations, we present experiments on the benchmark audio-goal task using the SoundSpaces simulator under various noisy settings. Our results reveal that our fully-conversational approach leads to nearly an order-of-magnitude improvement in success rate, especially in localizing new sound sources and against methods that use only uni-directional interaction.
Xiulong Liu 0002, Sudipta Paul 0007, Moitreya Chatterjee, Anoop Cherian
AAAI1
2024 From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
abstract
Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and visual elements. Previous studies of audio-visual modalities primarily focused on either audio-visual representation learning or generative modeling of a modality conditioned on the other, creating a disconnect between these two branches. A unified framework that learns representation and generates modalities has not been developed yet. In this work, we introduce a novel framework called Vision to Audio and Beyond (VAB) to bridge the gap between audio-visual representation learning and vision-to-audio generation. The key approach of VAB is that rather than working with raw video frames and audio data, VAB performs representation learning and generative modeling within latent spaces. In particular, VAB uses a pre-trained audio tokenizer and an image encoder to obtain audio tokens and visual features, respectively. It then performs the pre-training task of visual-conditioned masked audio token prediction. This training strategy enables the model to engage in contextual learning and simultaneous video-to-audio generation. After the pre-training phase, VAB employs the iterative-decoding approach to rapidly generate audio tokens conditioned on visual features. Since VAB is a unified model, its backbone can be fine-tuned for various audio-visual downstream tasks. Our experiments showcase the efficiency of VAB in producing high-quality audio from video, and its capability to acquire semantic audio-visual features, leading to competitive results in audio-visual retrieval and classification.
Xiulong Liu 0002, Eli Shlizerman
ICML2
2024 Tell What You Hear From What You See - Video to Audio Generation Through Text
abstract
The content of visual and audio scenes is multi-faceted such that a video stream can be paired with various audio streams and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the generated audio. While Video-to-Audio generation is a well-established generative task, existing methods lack such controllability. In this work, we propose VATT, a multi-modal generative framework that takes a video and an optional text prompt as input, and generates audio and optional textual description (caption) of the audio. Such a framework has two unique advantages: i) Video-to-Audio generation process can be refined and controlled via text which complements the context of the visual information, and ii) The model can suggest what audio to generate for the video by generating audio captions. VATT consists of two key modules: VATT Converter, which is an LLM that has been fine-tuned for instructions and includes a projection layer that maps video features to the LLM vector space, and VATT Audio, a bi-directional transformer that generates audio tokens from visual frames and from optional text prompt using iterative parallel decoding. The audio tokens and the text prompt are used by a pretrained neural codec to convert them into a waveform. Our experiments show that when VATT is compared to existing video-to-audio generation methods in objective metrics, such as VGGSound audiovisual dataset, it achieves competitive performance when the audio caption is not provided. When the audio caption is provided as a prompt, VATT achieves even more refined performance (with lowest KLD score of 1.41). Furthermore, subjective studies asking participants to choose the most compatible generated audio for a given silent video, show that VATT Audio has been chosen on average as a preferred generated audio than the audio generated by existing methods. VATT enables controllable video-to-audio generation through text as well as suggesting text prompts for videos through audio captions, unlocking novel applications such as text-guided video-to-audio generation and video-to-audio captioning.
Xiulong Liu 0002, Eli Shlizerman
NeurIPS1
2024 Tackling Data Bias in MUSIC-AVQA: Crafting a Balanced Dataset for Unbiased Question-Answering
abstract
In recent years, there has been a growing emphasis on the intersection of audio, vision, and text modalities, driving forward the advancements in multimodal research. However, strong bias that exists in any modality can lead to the model neglecting the others. Consequently, the model’s ability to effectively reason across these diverse modalities is compromised, impeding further advancement.In this paper, we meticulously review each question type from the original dataset, selecting those with pronounced answer biases. To counter these biases, we gather complementary videos and questions, ensuring that no answers have outstanding skewed distribution. In particular, for binary questions, we strive to ensure that both answers are almost uniformly spread within each question category. As a result, we construct a new dataset, named MUSIC-AVQA v2.0, which is more challenging and we believe could better foster the progress of AVQA task. Furthermore, we present a novel baseline model that delves deeper into the audiovisual-text interrelation. On MUSIC-AVQA v2.0, this model surpasses all the existing benchmarks, improving accuracy by 2% on MUSIC-AVQA v2.0, setting a new state-of-theart performance. Dataset: https://github.com/DragonLiu1995/MUSIC-AVQA-v2.0/
Xiulong Liu 0002, Zhikang Dong
WACV1
2024 Let the Beat Follow You - Creating Interactive Drum Sounds From Body Rhythm
abstract
It is often the case that human body movements include rhythmic patterns. A video camera system that captures these patterns and responds to them with rhythmic sounds or music, as these happen, could create a unique interactive experience. Creating such an experience requires a real-time translation of related visual cues into in-rhythm sounds and warrants novel real-time methods. In this work, we propose a novel learning-based system, called ‘InteractiveBeat’, which generates an evolving interactive soundtrack for a camera input that captures person’s movements. InteractiveBeat infers body skeleton keypoints and translates them into drum rhythms using a series of sequence models. It then implements a conditional drum generation network for generating polyphonic drum sounds based on the rhythms. To guarantee real-time function, these models are integrated into a time-evolving pipeline with rules for updates. InteractiveBeat is trained and evaluated on a well-annotated large-scale dance database (AIST), and in addition, we collected a dataset of in-the-wild videos with people performing movements of various activities that correspond to background music. Furthermore, we develop a ‘live’ demo prototype of the system. Our evaluation results show that the system can generate interactive rhythmic drums more accurately than existing methods and achieves a non-cumulative latency of 34ms (approx. 30 fps). This allows InteractiveBeat to be synchronized with the video stream and react to real-time movements.
Xiulong Liu 0002, Eli Shlizerman
WACV1
2021 How Does it Sound?
abstract
One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with the rhythmic features of these activities. How would this soundtrack sound? Such a problem is challenging since little is known about capturing the rhythmic nature of free body movements. In this work, we explore this problem and propose a novel system, called `RhythmicNet', which takes as an input a video which includes human movements and generates a soundtrack for it. RhythmicNet works directly with human movements by extracting skeleton keypoints and implements a sequence of models which translate the keypoints to rhythmic sounds.RhythmicNet follows the natural process of music improvisation which includes the prescription of streams of the beat, the rhythm and the melody. In particular, RhythmicNet first infers the music beat and the style pattern from body keypoints per each frame to produce rhythm. Next, it implements a transformer-based model to generate the hits of drum instruments and implements a U-net based model to generate the velocity and the offsets of the instruments. Additional types of instruments are added to the soundtrack by further conditioning on the generated drum sounds. We evaluate RhythmicNet on large scale datasets of videos that include body movements with inherit sound association, such as dance, as well as 'in the wild' internet videos of various movements and actions. We show that the method can generate plausible music that aligns well with different types of human movements.
Xiulong Liu 0002, Eli Shlizerman
NeurIPS2
2020 PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition
abstract
We propose a novel system for unsupervised skeleton-based action recognition. Given inputs of body-keypoints sequences obtained during various movements, our system associates the sequences with actions. Our system is based on an encoder-decoder recurrent neural network, where the encoder learns a separable feature representation within its hidden states formed by training the model to perform the prediction task. We show that according to such unsupervised training, the decoder and the encoder self-organize their hidden states into a feature space which clusters similar movements into the same cluster and distinct movements into distant clusters. Current state-of-the-art methods for action recognition are strongly supervised, i.e., rely on providing labels for training. Unsupervised methods have been proposed, however, they require camera and depth inputs (RGB+D) at each time step. In contrast, our system is fully unsupervised, does not require action labels at any stage and can operate with body-keypoints input only. Furthermore, the method can perform on various dimensions of body-keypoints (2D or 3D) and can include additional cues describing movements. We evaluate our system on three action recognition benchmarks with different numbers of actions and examples. Our results outperform prior unsupervised skeleton-based methods, unsupervised RGB+D based methods on cross-view tests and while being unsupervised have similar performance to supervised skeleton-based action recognition.
Xiulong Liu 0002, Eli Shlizerman
CVPR2
2020 Audeo: Audio Generation for a Silent Performance Video
abstract
We present a novel system that gets as an input, video frames of a musician playing the piano, and generates the music for that video. The generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named 'Audeo' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. In the last step, we implement Midi synthesizers to generate realistic music. Audeo converts video to audio smoothly and clearly with only a few setup constraints. We evaluate Audeo on piano performance videos collected from Youtube and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.
Xiulong Liu 0002, Eli Shlizerman
NeurIPS2