EDBT 2026 Demo / reviewers in the wild / expert
Zhiyong Yan
dblp:80/5482
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | APSys: Disaggregated LLM Serving System with an Adaptive Parallel Strategy
Zhaochen Li, Huaxia Wang, Anfa Zhang, Zhiyong Yan, Runhuai Huang |
IEEE Big Data | 4 |
| 2025 | Joint Prediction and Matching for Computing Resource Exchange PlatformsabstractThe rapid growth of deep learning has created unprecedented demand for computing resources, while many small and enterprise-level clusters remain underutilized. Computing resource exchange platforms offer a solution by aggregating these idle resources. However, effective cluster-task matching depends on accurate performance prediction. Existing approaches, which decouple prediction from matching, often lead to suboptimal decisions due to misaligned objectives. We propose a Matching-Focused Cluster Performance Predictor (MFCP), an end-to-end framework that integrates performance prediction with task matching to improve decision accuracy and resource utilization. Unlike existing methods that prioritize prediction accuracy, MFCP minimizes decision regret by aligning the predictor’s loss with optimal matching objectives. To handle non-differentiable matching optimization, we use continuous relaxation and incorporate constraints via an interior-point method, ensuring meaningful gradients for training. For non-convex optimization, we approximate optimal decisions with gradient descent and estimate gradients using zeroth-order perturbation. Experiments show that MFCP consistently outperforms existing methods across different cluster environments and scales, achieving lower matching regret and higher resource utilization. Da Huo 0002, Zhenzhe Zheng 0001, Xiaoyao Huang, Hao Chen 0181, Jianfeng Hu 0003, Zhiyong Yan, Fan Wu 0006, Jie Wu 0001 |
ICPP | 6 |
| 2024 | CED: Consistent Ensemble Distillation for Audio TaggingabstractAugmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the widely recognized Audioset (AS) benchmark. Although both techniques are effective individually, their combined use, called consistent teaching, hasn’t been explored before. This paper proposes CED, a simple training framework that distils student models from large teacher ensembles with consistent teaching. To achieve this, CED efficiently stores logits as well as the augmentation methods on disk, making it scalable to large-scale datasets. Central to CED’s efficacy is its label-free nature, meaning that only the stored logits are used for the optimization of a student model only requiring 0.3% additional disk space for AS. The study trains various transformer-based models, including a 10M parameter model achieving a 49.0 mean average precision (mAP) on AS. Pretrained models and code are available online. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 3 |
| 2024 | Scaling up masked audio encoder learning for general audio classification
Heinrich Dinkel, Zhiyong Yan, Bin Wang 0004 |
INTERSPEECH | 2 |
| 2024 | Streaming Audio Transformers for Online Audio Tagging
Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 2 |
| 2024 | Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Jizhong Liu, Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 6 |
| 2024 | Bridging Language Gaps in Audio-Text Retrieval
Zhiyong Yan, Heinrich Dinkel, Jizhong Liu |
INTERSPEECH | 1 |
| 2023 | Unified Keyword Spotting and Audio Tagging on Mobile Devices with TransformersabstractKeyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that adds additional capabilities in the form of audio tagging (AT) to a KWS model. However, previous work did not consider the real-world deployment of a UniKW-AT model, where factors such as model size and inference speed are more important than performance alone. This work introduces three mobile-device deployable models named Unified Transformers (UiT). Our best model achieves an mAP of 34.09 on Audioset, and an accuracy of 97.76 on the public Google Speech Commands V1 dataset. Further, we benchmark our proposed approaches on four mobile platforms, revealing that the proposed UiT models can achieve a speedup of 2 - 6 times against a competitive MobileNetV2. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 3 |
| 2023 | Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker ExtractionabstractVisual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio and visual. AV-SepFormer splits the audio feature into a number of chunks, equivalent to the length of the visual feature. Then self- and cross-attention are employed to model the multi-modal features. Furthermore, we use a novel 2D positional encoding, that introduces the positional information between and within chunks and provides significant gains over the traditional positional encoding. Our model has two key advantages: the time granularity of audio chunked feature is synchronized to the visual feature, which alleviates the harm caused by the inconsistency of audio and video sampling rate; by combining self- and cross-attention, feature fusion and speech extraction processes are unified within an attention paradigm. The experimental results show that AV-SepFormer significantly outperforms other existing methods. Jiuxin Lin, Xinyu Cai, Heinrich Dinkel, Jun Chen 0024, Zhiyong Yan, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 5 |
| 2023 | Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
Jiuxin Lin, Heinrich Dinkel, Jun Chen 0024, Zhiyong Wu 0001, Zhiyong Yan |
INTERSPEECH | 6 |
| 2022 | Pseudo Strong Labels for Large Scale Weakly Supervised Audio TaggingabstractLarge-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual labeling cost. This work proposes pseudo strong labels (PSL), a simple label augmentation framework that enhances the supervision quality for large-scale weakly supervised audio tagging. A machine annotator is first trained on a large weakly supervised dataset, which then provides finer supervision for a student model. Using PSL we achieve an mAP of 35.95 balanced train subsets of Audioset using a MobileNetV2 backend, significantly outperforming approaches without PSL. An analysis is provided which reveals that PSL mitigates missing labels. Lastly, we show that models trained with PSL are also superior at generalizing to the Freesound datasets (FSD) than their weakly trained counterparts. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 2 |
| 2022 | UniKW-AT: Unified Keyword Spotting and Audio TaggingabstractWithin the audio research community and the industry, keyword spotting (KWS) and audio tagging (AT) are seen as two distinct tasks and research fields. However, from a technical point of view, both of these tasks are identical: they predict a label (keyword in KWS, sound event in AT) for some fixed-sized input audio segment. This work proposes UniKW-AT: An initial approach for jointly training both KWS and AT. UniKW-AT enhances the noise-robustness for KWS, while also being able to predict specific sound events and enabling conditional wake-ups on sound events. Our approach extends the AT pipeline with additional labels describing the presence of a keyword. Experiments are conducted on the Google Speech Commands V1 (GSCV1) and the balanced Audioset (AS) datasets. The proposed MobileNetV2 model achieves an accuracy of 97.53% on the GSCV1 dataset and an mAP of 33.4 on the AS evaluation set. Further, we show that significant noise-robustness gains can be observed on a real-world KWS dataset, greatly outperforming standard KWS approaches. Our study shows that KWS and AT can be merged into a single framework without significant performance degradation. Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 3 |
| 2021 | GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10, 000 Hours of Transcribed AudioabstractThis paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training.Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription.For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%.The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.Baseline systems are provided for popular speech recognition toolkits, namely Athena, ESPnet, Kaldi and Pika. Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Weiqiang Zhang 0001, Chao Weng, Dan Su 0002, Daniel Povey, Jan Trmal, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe 0001, Shuaijiang Zhao, Xiangang Li, Xuchen Yao, Zhao You, Zhiyong Yan |
Interspeech | 20 |
| 2021 | speechocean762: An Open-Source Non-Native English Speech Corpus for Pronunciation AssessmentabstractThis paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children.Five experts annotated each of the utterances at sentence-level, wordlevel and phoneme-level.A baseline system is released in open source to illustrate the phoneme-level pronunciation assessment workflow on this corpus.This corpus is allowed to be used freely for commercial and non-commercial purposes.It is available for free download from OpenSLR, and the corresponding baseline system is published in the Kaldi speech recognition toolkit. Zhiyong Yan, Qiong Song, Ke Li 0018, Daniel Povey |
Interspeech | 4 |
| 2011 | Improving naive Bayes classifier by dividing its decision regionsabstractClassification can be regarded as dividing the data space into decision regions separated by decision boundaries. In this paper we analyze decision tree algorithms and the NBTree algorithm from this perspective. Thus, a decision tree can be regarded as a classifier tree, in which each classifier on a non-root node is trained in decision regions of the classifier on the parent node. Meanwhile, the NBTree algorithm, which generates a classifier tree with the C4.5 algorithm and the naive Bayes classifier as the root and leaf classifiers respectively, can also be regarded as training naive Bayes classifiers in decision regions of the C4.5 algorithm. We propose a second division (SD) algorithm and three soft second division (SD-soft) algorithms to train classifiers in decision regions of the naive Bayes classifier. These four novel algorithms all generate two-level classifier trees with the naive Bayes classifier as root classifiers. The SD and three SD-soft algorithms can make good use of both the information contained in instances near decision boundaries, and those that may be ignored by the naive Bayes classifier. Finally, we conduct experiments on 30 data sets from the UC Irvine (UCI) repository. Experiment results show that the SD algorithm can obtain better generalization abilities than the NBTree and the averaged one-dependence estimators (AODE) algorithms when using the C4.5 algorithm and support vector machine (SVM) as leaf classifiers. Further experiments indicate that our three SD-soft algorithms can achieve better generalization abilities than the SD algorithm when argument values are selected appropriately. Zhiyong Yan, Congfu Xu, Yunhe Pan |
J. Zhejiang Univ. Sci. C | 1 |