EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Xie
dblp:271/8450
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2025
0009-0001-9546-3301ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PicoAudio: Enabling Precise Temporal Controllability in Text-to-Audio GenerationabstractRecently, audio generation tasks have attracted considerable research interests. Despite rapid advancements in generating high-fidelity audio that is coarsely aligned with the text description, precise temporal controllability is still a challenge, which is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio generation framework, PicoAudio. It leverages data crawling, segmentation and filtering to simulate fine-grained temporally-aligned audio-text data. Furthermore, PicoAudio integrates temporal information to guide audio generation through tailored model design. With the effective text processing capabilities from large language models, PicoAudio can take natural language input and generate audio that aligns well with the temporal description in the input. Both subjective and objective evaluation demonstrate that PicoAudio dramatically surpasses current state-of-the-art generation models in terms of timestamp and occurrence frequency controllability. Generation samples are available at the $PicoAudio - Demo$. Zeyu Xie, Xuenan Xu, Zhizheng Wu 0001, Mengyue Wu |
ICASSP | 1 |
| 2025 | AudioTime: A Temporally-aligned Audio-text Benchmark DatasetabstractRecent advances in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relation, a critical feature for audio content, is currently underrepresented in mainstream models, resulting in an imprecise temporal controllability. Specifically, users cannot accurately control the timestamps of sound events using free-form text. One significant challenge is the absence of a high-quality, temporally-aligned audio-text dataset, which is essential for training models with temporal control. The more temporally-aligned the annotations, the better the models can understand the precise relationship between audio outputs and temporal textual prompts. Therefore, we propose a temporally-aligned audio-text dataset, AudioTime. It provides text annotations rich in temporal information such as timestamps, duration, frequency, and ordering, covering almost all aspects of temporal control. Additionally, we offer a comprehensive test set and evaluation metric to assess the temporal control performance of text-to-audio generation models. Examples are available on the $AudioTime - Demo$. Zeyu Xie, Xuenan Xu, Zhizheng Wu 0001, Mengyue Wu |
ICASSP | 1 |
| 2025 | Robust Visual Food Recognition for Enriching Nutrition Knowledge BasesabstractAcquiring nutrition information and health-related knowledge about food is a common need among individuals. However, using conventional food names as search queries often fails to yield accurate matches to entries within food nutrition knowledge bases (FoodnKB), which frequently utilize scientific or product names. In this study, we present a method for enriching FoodnKB entries with imagery and facilitating visual access to food-related knowledge through image recognition. We start with an official food nutrition database and propose a consensus-based approach using Large Language Models to identify visually discernible and directly edible foods, expanding food synonyms and harnessing diverse web-based food images for comprehensive visual representation. To minimize manual annotation of noisy web images, we introduce a cyclic training-based area under the margin metric (cAUM) approach that effectively distinguishes appropriate images, including rare instances, from noisy ones. Additionally, we design a generic accuracy gap (AccGap) algorithm to automatically estimate the noise ratio of the web-harnessed data. Our integrated cAUM and AccGap method demonstrates superior performance in noise detection and enhancement of image recognition accuracy compared to existing noise-robust frameworks. Furthermore, we successfully apply the visually enriched FoodnKB and food recognition capabilities within a smart nutritionist mobile application. Zhaoyan Ming, Zeyu Xie, Kui Su, Changzheng Yuan, Tat-Seng Chua |
IEEE Trans. Multim. | 2 |
| 2024 | Enhancing Audio Generation Diversity with Visual InformationabstractAudio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited research on enhancing the diversity of generated audio, particularly when it comes to audio generation within specific categories. Current models tend to produce homogeneous audio samples within a category. This work aims to address this limitation by improving the diversity of generated audio with visual information. We propose a clustering-based method, leveraging visual information to guide the model in generating distinct audio content within each category. Results on seven categories indicate that extra visual input can largely enhance audio generation diversity. Audio samples are available at DemoWeb. Zeyu Xie, Baihan Li, Xuenan Xu, Mengyue Wu, Kai Yu 0004 |
ICASSP | 1 |
| 2024 | A Detailed Audio-Text Data Simulation Pipeline Using Single-Event SoundsabstractRecently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such descriptions are significantly limited. In this paper, we first analyze the detailed information that human descriptions of audio may contain beyond sound event labels. Based on the analysis, we propose an automatic pipeline for curating audio-text pairs with rich details1. Leveraging the property that sounds can be mixed and concatenated in the time domain, we control details in four aspects: temporal relationship, loudness, speaker identity, and occurrence number, in simulating audio mixtures. Corresponding details are transformed into captions by large language models. Audio-text pairs with rich details in text descriptions are thereby obtained. We validate the effectiveness of our pipeline with a small amount of simulated data, demonstrating that the simulated data enables models to learn detailed audio captioning. Xuenan Xu, Xiaohang Xu 0004, Zeyu Xie, Pingyue Zhang, Mengyue Wu, Kai Yu 0004 |
ICASSP | 3 |
| 2024 | DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
Baihan Li, Zeyu Xie, Xuenan Xu, Ming Yan 0008, Ji Zhang 0011, Kai Yu 0004, Mengyue Wu |
INTERSPEECH | 2 |
| 2024 | FakeSound: Deepfake General Audio Detection
Zeyu Xie, Baihan Li, Xuenan Xu, Mengyue Wu |
INTERSPEECH | 1 |
| 2024 | Beyond the Status Quo: A Contemporary Survey of Advances and Challenges in Audio CaptioningabstractAutomated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress over the last few years. AAC requires recognizing contents such as the environment, sound events and the temporal relationships between sound events and describing these elements with a fluent sentence. Currently, an encoder-decoder-based deep learning framework is the standard approach to tackle this problem. Plenty of works have proposed novel network architectures and training schemes, including extra guidance, reinforcement learning, audio-text self-supervised learning and diverse or controllable captioning. Effective data augmentation techniques, especially based on large language models are explored. Benchmark datasets and AAC-oriented evaluation metrics also accelerate the improvement of this field. This article situates itself as a comprehensive survey covering the comparison between AAC and its related tasks, the existing deep learning techniques, datasets, and the evaluation metrics in AAC, with insights provided to guide potential future research directions. Xuenan Xu, Zeyu Xie, Mengyue Wu, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Enhance Temporal Relations in Audio Captioning with Sound Event DetectionabstractAutomated audio captioning aims at generating natural language descriptions for given audio clips, not only detecting and classifying sounds, but also summarizing the relationships between audio events.Recent research advances in audio captioning have introduced additional guidance to improve the accuracy of audio events in generated sentences.However, temporal relations between audio events have received little attention while revealing complex relations is a key component in summarizing audio content.Therefore, this paper aims to better capture temporal relationships in caption generation with sound event detection (SED), a task that locates events' timestamps.We investigate the best approach to integrate temporal information in a captioning model and propose a temporal tag system to transform the timestamps into comprehensible relations.Results evaluated by the proposed temporal metrics suggest that great improvement is achieved in terms of temporal relation generation 1 . Zeyu Xie, Xuenan Xu, Mengyue Wu, Kai Yu 0004 |
INTERSPEECH | 1 |
| 2023 | BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic DataabstractCompared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text pre-training, which inevitably induces data noise. In this paper, we propose to utilize audio captioning to generate text directly from audio, without the aid of the visual modality so that potential noise from modality mismatch is eliminated. Furthermore, we propose caption generation under the guidance of AudioSet tags, leading to more accurate captions. With the above two improvements, we curate high-quality, large-scale parallel audio-text data, based on which we perform audio-text pre-training. We comprehensively demonstrate the performance of the pre-trained model on a series of downstream audio-related tasks, including single-modality tasks like audio classification and tagging, as well as cross-modal tasks consisting of audio-text retrieval and audio-based text generation. Experimental results indicate that our approach achieves state-of-the-art zero-shot classification performance on most datasets, suggesting the effectiveness of our synthetic data. The audio encoder also serves as an efficient pattern recognition model by fine-tuning it on audio-related tasks. Synthetic data and pre-trained models are available online1 The code, checkpoints and data are available at https://github.com/wsntxxn/BLAT and https://zenodo.org/record/8218696/. Xuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang, Zeyu Xie, Mengyue Wu, Kenny Q. Zhu |
ACM Multimedia | 5 |
| 2022 | Can Audio Captions Be Evaluated With Image Caption Metrics?abstractAutomated audio captioning aims at generating textual descriptions for an audio clip. To evaluate the quality of generated audio captions, previous works directly adopt image captioning metrics like SPICE and CIDEr, without justifying their suitability in this new domain, which may mislead the development of advanced models. This problem is still unstudied due to the lack of human judgment datasets on caption quality. Therefore, we first construct two evaluation benchmarks, AudioCaps-Eval and Clotho-Eval. They are established with pairwise comparison instead of absolute rating to achieve better inter-annotator agreement. Current metrics are found in poor correlation with human annotations on these datasets. To overcome their limitations, we propose a metric named FENSE, where we combine the strength of Sentence-BERT in capturing similarity, and a novel Error Detector to penalize erroneous sentences for robustness. On the newly established benchmarks, FENSE outperforms current metrics by 14-25% accuracy.1 Zelin Zhou, Zhiling Zhang, Xuenan Xu, Zeyu Xie, Mengyue Wu, Kenny Q. Zhu |
ICASSP | 4 |
| 2021 | Investigating Local and Global Information for Automated Audio Captioning with Transfer LearningabstractAutomated audio captioning (AAC) aims at generating summarizing descriptions for audio clips. Multitudinous concepts are described in an audio caption, ranging from local information such as sound events to global information like acoustic scenery. Currently, the mainstream paradigm for AAC is the end-to-end encoder-decoder architecture, expecting the encoder to learn all levels of concepts embedded in the audio automatically. This paper first proposes a topic model for audio descriptions, comprehensively analyzing the hierarchical audio topics that are commonly covered. We then explore a transfer learning scheme to access local and global information. Two source tasks are identified to respectively represent local and global information, being Audio Tagging (AT) and Acoustic Scene Classification (ASC). Experiments are conducted on the AAC benchmark dataset Clotho and Audiocaps, amounting to a vast increase in all eight metrics with topic transfer learning. Further, it is discovered that local information and abstract representation learning are more crucial to AAC than global information and temporal relationship learning. Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Zeyu Xie, Kai Yu 0004 |
ICASSP | 4 |
| 2020 | Medical image fusion using the PCNN based on IQPSO in NSST domainabstractIn this study, an improved quantum‐behaved particle swarm optimisation based pulse‐coupled neural network (IQPSO‐PCNN) is proposed in the non‐subsampled shearlet transform (NSST) domain for medical image fusion. First, NSST tool is used to decompose the source image into low‐frequency and high‐frequency subbands. Then, for low‐frequency subbands, the fusion rules of two different functions are presented, which simultaneously addresses two key issues of energy preservation and detail extraction. For high‐frequency subbands, unlike conventional PCNN‐based methods, parameters are manually set based on experience, and the decomposed high‐frequency subbands share a set of parameters. The IQPSO‐PCNN model can obtain the optimal parameters for each high‐frequency subband adaptively according to its own information. Finally, the fused low‐frequency subband and high‐frequency subbands are inversely transformed by NSST to acquire the final fused image. The proposed algorithm uses >90 pairs of images with four different modalities. In addition, fusion experiments are performed on different sequences of the three modes. The experimental results demonstrate that the proposed method is superior to existing state‐of‐art methods in subjective visual performance and objective evaluation. Di Gai, Xuanjing Shen, Haipeng Chen 0002, Zeyu Xie, Pengxiang Su |
IET Image Process. | 4 |