VLDB 2026 Research / reviewers in the wild / expert
Zeyue Tian
dblp:335/0981
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-7278-3708ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VMChill: A Dataset for Fine-Grained Visual-Musical SynergyabstractMassive multi-modality datasets are fundamental to the success of large video-language models. However, existing datasets often focus on providing textual descriptions for visual content, treating audio, particularly music, as weakly related information. This overlooks the inherent semantic correlation between visual narratives and musical scores, limiting the development of models for fine-grained cross-modal understanding and generation. To address this gap, we introduce VMChill, a large-scale, fine-grained multimodal video dataset. We leverage trailers as our data source, as they are professionally edited to create a strong synergy between visual pacing, scene transitions, and background music for narrative and emotional impact. Our dataset comprises over 20 million video clips derived from more than 27.1k hours of high-resolution trailer videos. To annotate this data, we propose a systematic multimodal captioning framework. This framework first employs specialized unimodal models to extract descriptive features from multiple perspectives, including visual content, motion dynamics, and musical attributes (e.g., genre, instruments, mood). Subsequently, a large language model (LLM) is utilized to adaptively fuse these diverse descriptions into a single, coherent, and rich multimodal caption. This process yields VMChill-2M, a high-quality subset of 2 million clips with detailed multimodal annotations, and VMChill-Test, a manually refined test set for evaluation. We conduct extensive experiments on downstream tasks, including video understanding and generation, to establish benchmarks and demonstrate the dataset's quality. The results validate that VMChill effectively enhances model performance, highlighting its potential to facilitate future research in fine-grained multimodal learning. We will release the dataset, annotation codebase, and processing pipelines to support community research. Xiaowei Chi, Zeyue Tian, Wei Xue 0002 |
AAAI | 2 |
| 2026 | Inference-time Scaling for Diffusion-based Audio Super-resolutionabstractDiffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for achieving superior audio quality. However, while existing diffusion approaches typically increase sampling steps to improve quality, the performance remains fundamentally limited by the stochastic nature of the sampling process, leading to high-variance and quality-limited outputs. Here, rather than simply increasing the number of sampling steps, we propose a different paradigm through inference-time scaling for SR, which explores multiple solution trajectories during the sampling process. Different task-specific verifiers are developed, and two search algorithms, including the random search and zero-order search for SR, are introduced. By actively guiding the exploration of the high-dimensional solution space through verifier-algorithm combinations, we enable more robust and higher-quality outputs. Through extensive validation across diverse audio domains (speech, music, sound effects) and frequency ranges, we demonstrate consistent performance gains, achieving improvements of up to 9.70% in aesthetics, 5.88% in speaker similarity, 15.20% in word error rate, and 46.98% in spectral distance for speech SR from 4 kHz to 24 kHz, showcasing the effectiveness of our approach. Yizhu Jin, Zhen Ye 0006, Zeyue Tian, Haohe Liu, Qiuqiang Kong, Yike Guo, Wei Xue 0002 |
AAAI | 3 |
| 2025 | VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term ModelingabstractIn this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset by collecting 360K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framework for generating music aligned with video inputs. VidMuse stands out by producing high-fidelity music that is both acoustically and semantically aligned with the video. By incorporating local and global visual cues, VidMuse enables the creation of coherent music tracks that consistently match the video content through Long-Short-Term modeling. Through extensive experiments, VidMuse outperforms existing models in terms of audio quality, diversity, and audio-visual alignment. The code and datasets are available at https://vidmuse.github.io/ Zeyue Tian, Zhaoyang Liu 0001, Ruibin Yuan, Xu Tan 0003, Qifeng Chen 0001, Wei Xue 0002, Yike Guo |
CVPR | 1 |
| 2024 | Multitarget Device-Free Localization via Cross-Domain Wi-Fi RSS Training Data and Attentional Prior FusionabstractDevice-free localization (DFL) using easily-obtained Wi-Fi received signal strength (RSS) has wide real-world applications for not requiring people to carry trackable devices. However, accurate multitarget DFL remains challenging due to the unknown number of targets, multipath interference (MPI), especially between nearby targets, and limited real-world data. In this study, we pioneeringly propose a transformer-based learning method with Wi-Fi RSS as input, and an attentional prior fusion module, to simultaneously locate an unknown number of people at random positions. To overcome the multitarget data collection challenges, we contribute a large-scale cross-domain real-simulation-augmentation training dataset with one and two real-world nearby non-person objects at limited positions and up to five simulated and augmented randomly distributed targets. Experimental results demonstrate our method's improved accuracy, generalization ability, and robustness with fewer Wi-Fi nodes than previous methods. Na Fan 0002, Zeyue Tian, Amartansh Dubey, Samruddhi Deshmukh, Ross Murch, Qifeng Chen 0001 |
AAAI | 2 |
| 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersabstractVideo and audio content creation serves as the core technique for the movie industry and professional users. Re-cently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from academia to industry. In this work, we aim at filling the gap, with a carefully designed optimization-based framework for cross-visual-audio and joint-visual-audio generation. We observe the powerful generation abil-ity of off-the-shelf video or audio generation models. Thus, instead of training the giant models from scratch, we pro-pose to bridge the existing strong models with a shared la-tent representation space. Specifically, we propose a mul-timodality latent aligner with the pre-trained ImageBind model. Our latent aligner shares a similar core as the clas-sifier guidance that guides the diffusion denoising process during inference time. Through carefully designed opti-mization strategy and loss functions, we show the superior performance of our method on joint video-audio generation, visual-steered audio generation, and audio-steered vi-sual generation tasks. The project website can be found at https://yzxing87.github.io/Seeing-and-Hearing/. Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 0002, Qifeng Chen 0001 |
CVPR | 3 |
| 2023 | Mixed Neural Voxels for Fast Multi-view Video SynthesisabstractSynthesizing high-fidelity videos from real-world multi-view input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such models on real-world scenes is time-consuming, usually taking days or weeks. In this paper, we present a novel method named MixVoxels to efficiently represent dynamic scenes, enabling fast training and rendering speed. The proposed MixVoxels represents the 4D dynamic scenes as a mixture of static and dynamic voxels and processes them with different networks. In this way, the computation of the required modalities for static voxels can be processed by a lightweight model, which essentially reduces the amount of computation as many daily dynamic scenes are dominated by static backgrounds. To distinguish the two kinds of voxels, we propose a novel variation field to estimate the temporal variance of each voxel. For the dynamic representations, we design an inner product time query method to efficiently query multiple time steps, which is essential to recover the high-dynamic movements. As a result, with 15 minutes of training for dynamic scenes with inputs of 300-frame videos, MixVoxels achieves better PSNR than previous methods. For rendering, MixVoxels can render a novel view video with 1K resolution at 37 fps. Codes and trained models are available at https://github.com/fengres/mixvoxels. Feng Wang 0034, Sinan Tan, Xinghang Li, Zeyue Tian, Huaping Liu 0001 |
ICCV | 4 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 10 |
| 2023 | Deep Cascade Gradient RBF Networks With Output-Relevant Feature Extraction and Adaptation for Nonlinear and Nonstationary ProcessesabstractThe main challenge for industrial predictive models is how to effectively deal with big data from high-dimensional processes with nonstationary characteristics. Although deep networks, such as the stacked autoencoder (SAE), can learn useful features from massive data with multilevel architecture, it is difficult to adapt them online to track fast time-varying process dynamics. To integrate feature learning and online adaptation, this article proposes a deep cascade gradient radial basis function (GRBF) network for online modeling and prediction of nonlinear and nonstationary processes. The proposed deep learning method consists of three modules. First, a preliminary prediction result is generated by a GRBF weak predictor, which is further combined with raw input data for feature extraction. By incorporating the prior weak prediction information, deep output-relevant features are extracted using a SAE. Online prediction is finally produced upon the extracted features with a GRBF predictor, whose weights and structure are updated online to capture fast time-varying process characteristics. Three real-world industrial case studies demonstrate that the proposed deep cascade GRBF network outperforms existing state-of-the-art online modeling approaches as well as deep networks, in terms of both online prediction accuracy and computational complexity. Tong Liu 0014, Zeyue Tian, Sheng Chen 0001, Kai Wang 0003, Christopher J. Harris 0001 |
IEEE Trans. Cybern. | 2 |