Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hankun Xu

dblp:419/7273 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 61% Multimedia analysis and retrieval · 30% Virtual and augmented reality · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › multimedia dataset construction
multimodal dataset
0.912025
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations · NeurIPS 2025
Audio and music processing
sound source localization
0.912025
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations · NeurIPS 2025
Audio and music processing
spatial audio
0.912025
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations · NeurIPS 2025
Virtual and augmented reality
immersive audio
0.312025
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

binaural audio · 0.9ambisonic audio · 0.9
YearPublicationVenuePosition
2025 MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
abstract
Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address these challenges, we introduce MRSAudio, a large-scale multimodal spatial audio dataset designed to advance research in spatial audio understanding and generation. MRSAudio spans four distinct components: MRSLife, MRSSpeech, MRSMusic, and MRSSing, covering diverse real-world scenarios. The dataset includes synchronized binaural and ambisonic audio, exocentric and egocentric video, motion trajectories, and fine-grained annotations such as transcripts, phoneme boundaries, lyrics, scores, and prompts.To demonstrate the utility and versatility of MRSAudio, we establish five foundational tasks: audio spatialization, and spatial text to speech, spatial singing voice synthesis, spatial music generation and sound event localization and detection. Results show that MRSAudio enables high-quality spatial modeling and supports a broad range of spatial audio research.Demos and dataset access are available at https://mrsaudio.github.io.
Wenxiang Guo, Changhao Pan, Xintong Hu, Yu Zhang 0126, Han Wang 0019, Zongbao Zhang, Hankun Xu, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu 0001, Zhou Zhao 0001
NeurIPS12