EDBT 2026 Demo / reviewers in the wild / expert
Zili Huang
dblp:210/0905
· DBLP profile ↗
21ranked-venue papers
8as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Factorized RVQ-GAN For Disentangled Speech TokenizationabstractInternational audience Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, François G. Germain, Gordon Wichern, Jonathan Le Roux |
INTERSPEECH | 8 |
| 2025 | VSLAM-BA: Algorithm and Hardware Co-Design for High Performance and Energy-Efficient Visual SLAM Backend Hardware AcceleratorabstractVisual Simultaneous Localization and Mapping (VSLAM) is a key localization technology for emerging applications such as autonomous driving and uncrewed aerial vehicles (UAVs). Compared with VSLAM frontend, VSLAM backend plays a more important role as it is employed to improve the localization accuracy. However, the VSLAM backend usually uses Bundle Adjustment (BA) as its core optimization method which is well-known for its large scale of problem construction, high computational complexity and high serialization of data processing, making it difficult to achieve high performance and energy efficiency on platforms such as CPUs or GPUs. Although there are some VSLAM backend accelerators proposed recently for addressing the above issues, they did not well exploit the data regularity and computational characteristics, resulting in limited performance/energy efficiency improvements or degraded accuracy. In this work, we propose VSLAM-BA which is a high performance and energy-efficient VSLAM backend accelerator with algorithm-hardware co-design. On the algorithm level, a keyframe-split-based Schur elimination scheme is proposed to reduce latency, power consumption and memory storage while maintaining accuracy. On the hardware level, a column-folding-based computing architecture is proposed to boost performance and energy efficiency. A loading-sensitive matrix-computing technique with an adaptive task scheduler is proposed to reduce the latency and energy consumption. Further, a recyclable computing technique with point-aware solver is proposed to reduce the memory and energy consumption. The experimental results show that the proposed VSLAM-BA achieves the highest performance (380 fps) and the highest energy efficiency (0.51 mJ per frame) with high accuracy and low memory storage, compared with the SOTA designs. The proposed accelerator can work with different VSLAM frontend for backend optimization of localization accuracy. Ye Liu 0011, Xiuyuan Qi, Shuang Hao 0005, Zili Huang, Neng Zhao, Ruixin Mao, Sixu Li, Ang Hu, Yu Long 0005, Shanshan Liu 0001, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | An Ultra-High Performance and Scalable Optical Flow Hardware Accelerator Based on FPGA for Autonomous DrivingabstractOptical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. Nevertheless, the incorporation of the image pyramid elevates the computational complexity. Moreover, the data relationships between pyramid layers result in strong data dependencies, which renders it difficult to accelerate via parallel processing and makes it challenging to fulfill real-time demands in practical scenarios. To address this issue, in this paper, we propose an ultra-high performance and scalable optical flow hardware accelerator based on FPGA with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 1.02) compared with SOTA hardware accelerators. Ye Liu 0011, Shuang Hao 0005, Xiuyuan Qi, Zili Huang, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech ProcessingabstractThe speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. In response to the diversity of microphone configurations in use, we introduce the UniX-Encoder, a universal encoder for multi-channel speech recordings. The UniX-Encoder is versatile, catering to a variety of speech tasks, and seamlessly integrates with any microphone array, whether in single-talker or multi-talker environments. Our research enhances previous multi-channel speech processing efforts in four aspects: 1) Adaptability: Contrasting traditional models constrained to certain microphone array configurations, our encoder is universally compatible. 2) Multi-Task Capability: Contrasting previous systems that were designed for single-task applications, the UniX-Encoder serves as a versatile upstream model, capable of extracting features for diverse speech tasks. 3) Self-Supervised Training: The UniX-Encoder is pretrained without the need for labeled multi-channel data. 4) End-to-End Integration: In contrast to models that first beamform then process single-channels, our encoder offers an end-to-end solution, bypassing explicit beamforming or separation. To validate its effectiveness, we tested the UniX-Encoder on a synthetic multi-channel dataset from the LibriSpeech corpus. Across various tasks, including ASR and speaker diarization, our encoder consistently outperformed combinations such as the WavLM model with the BeamformIt frontend. Zili Huang, Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 1 |
| 2024 | An FPGA-based Ultra-High Performance and Scalable Optical Flow Hardware Accelerator for Autonomous DrivingabstractOptical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. However, the introduction of the image pyramid increases the computational complexity. Additionally, the data relationships between pyramid layers result in strong data dependencies, making it difficult to accelerate using parallel processing, and it is challenging to meet real-time requirements in practical scenarios. To address this issue, in this paper, we propose an FPGA-based ultra-high performance and scalable optical flow hardware accelerator with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 0.64). Ye Liu 0011, Shuang Hao 0005, Zili Huang, Xiuyuan Qi, Yu Long 0005, Jun Zhou 0017 |
ISCAS | 5 |
| 2024 | A Large-Scale Evaluation of Speech Foundation ModelsabstractThe foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific data collection and modeling. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. To bridge this gap, we establish the Speech processing Universal PERformance Benchmark (SUPERB). SUPERB represents an ecosystem designed to evaluate foundation models across a wide range of speech processing tasks, facilitating the sharing of results on an online leaderboard and fostering collaboration through a community-driven benchmark database that aids in new development cycles. We present a unified learning framework for solving the speech processing tasks in SUPERB with the frozen foundation model followed by task-specialized lightweight prediction heads. Combining our results with community submissions, we verify that the framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models and the statistical significance and robustness of the benchmark. Shu-Wen Yang, Heng-Jui Chang, Zili Huang, Andy T. Liu, Cheng-I Lai, Jiatong Shi, Xuankai Chang, Hsiang-Sheng Tsai, Wen-Chin Huang, Tzu-hsun Feng, Po-Han Chi, Yist Y. Lin, Yung-Sung Chuang, Tzu-Hsien Huang, Wei-Cheng Tseng, Kushal Lakhotia, Shang-Wen Li 0001, Abdel-rahman Mohamed, Shinji Watanabe 0001, Hung-yi Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech RecognitionabstractSelf-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However, most of the previous studies focused on offline single-talker applications, with limited investigations in multi-talker cases, especially for streaming scenarios. In this paper, we investigate SSL for streaming multi-talker speech recognition, which generates transcriptions of overlapping speakers in a streaming fashion. Firstly, we observe that conventional SSL techniques do not work well on this task due to the poor representation of overlapping speech. We then propose a novel SSL training objective, referred to as bi-label masked speech prediction, which explicitly preserves representations of all speakers in overlapping speech. We investigate various aspects of the proposed system, including data configuration and quantizer selection. The proposed SSL setup achieves substantially better word error rates on the LibriSpeechMix dataset. Zili Huang, Zhuo Chen 0006, Naoyuki Kanda, Jian Wu 0027, Jinyu Li 0001, Takuya Yoshioka, Xiaofei Wang 0009 |
ICASSP | 1 |
| 2023 | Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker EmbeddingsabstractSelf-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to the domain mismatch — which severely limits their use for such applications. In this paper, we investigate the adaptation of upstream SSL models to the multi-talker automatic speech recognition (ASR) task under two conditions. First, when segmented utterances are given, we show that adding a target speaker extraction (TSE) module based on enrollment embeddings is complementary to mixture-aware pre-training. Second, for unsegmented mixtures, we propose a novel joint speaker modeling (JSM) approach, which aggregates information from all speakers in the mixture through their embeddings. With controlled experiments on Libri2Mix, we show that using speaker embeddings provides relative WER improvements of 9.1% and 42.1% over strong baselines for the segmented and unsegmented cases, respectively. We also demonstrate the effectiveness of our models for real conversational mixtures through experiments on the AMI dataset. Our code and models are open-sourced on https://github.com/HuangZiliAndy/SSL_for_multitalker. Zili Huang, Desh Raj, L. Paola García-Perera, Sanjeev Khudanpur |
ICASSP | 1 |
| 2022 | SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative CapabilitiesabstractHsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-wen Yang, Shuyan Dong, Andy Liu, Cheng-I Lai, Jiatong Shi, Xuankai Chang, Phil Hall, Hsuan-Jui Chen, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, Hung-yi Lee. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-Wen Yang, Shuyan Dong, Andy T. Liu, Cheng-I Lai, Jiatong Shi, Xuankai Chang, Phil Hall, Hsuan-Jui Chen, Shang-Wen Li 0001, Shinji Watanabe 0001, Abdel-rahman Mohamed, Hung-yi Lee |
ACL (1) | 4 |
| 2022 | Investigating Self-Supervised Learning for Speech Enhancement and SeparationabstractSpeech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a great number of supervised learning-based enhancement and separation methods having been proposed and achieving good performance, studies on applying self-supervised learning (SSL) to enhancement and separation are limited. In this paper, we evaluate 13 SSL upstream methods on speech enhancement and separation downstream tasks. Our experimental results on Voicebank-DEMAND and Libri2Mix show that some SSL representations consistently outperform baseline features including the short-time Fourier transform (STFT) magnitude and log Mel filterbank (FBANK). Furthermore, we analyze the factors that make existing SSL frameworks difficult to apply to speech enhancement and separation and discuss the representation properties desired for both tasks. Our study is included as the official speech enhancement and separation downstreams for SUPERB. Zili Huang, Shinji Watanabe 0001, Shu-Wen Yang, L. Paola García-Perera, Sanjeev Khudanpur |
ICASSP | 1 |
| 2022 | Superb @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation LearningabstractWe present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge builds upon the SUPERB benchmark and implements metrics to measure the computation requirements of self-supervised learning (SSL) representation and to evaluate its generalizability and performance across the diverse SUPERB tasks. The SUPERB benchmark provides comprehensive coverage of popular speech processing tasks, from speech and speaker recognition to audio generation and semantic understanding. As SSL has gained interest in the speech community and showed promising outcomes, we envision the challenge to uplevel the impact of SSL techniques by motivating more practical designs of techniques beyond task performance. We summarize the results of 14 submitted models in this paper. We also discuss the main findings from those submissions and the future directions of SSL research. Tzu-hsun Feng, Shuyan Dong, Ching-Feng Yeh, Shu-Wen Yang, Tzu-Quan Lin, Jiatong Shi, Kai-Wei Chang 0001, Zili Huang, Xuankai Chang, Shinji Watanabe 0001, Abdel-rahman Mohamed, Shang-Wen Li 0001, Hung-yi Lee |
SLT | 8 |
| 2022 | Joint speaker diarization and speech recognition based on region proposal networks
Zili Huang, Marc Delcroix, L. Paola García-Perera, Shinji Watanabe 0001, Desh Raj, Sanjeev Khudanpur |
Comput. Speech Lang. | 1 |
| 2021 | Target-Speaker Voice Activity Detection with Improved i-Vector Estimation for Unknown Number of SpeakerabstractTarget-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fixed (and known) number of speakers, which limits its application to real conversations. In this paper, we extend TS-VAD to speaker diarization with unknown numbers of speakers. This is achieved by two steps: first, an initial diarization system is applied for speaker number estimation, followed by TS-VAD network output masking according to this estimate. We further investigate different diarization methods, including clustering-based and region proposal networks, for estimating the initial i-vectors. Since these systems have complementary strengths, we propose a fusion-based method to combine frame-level decisions from the systems for an improved initialization. We demonstrate through experiments on variants of the LibriCSS meeting corpus that our proposed approach can improve the DER by up to 50\% relative across varying numbers of speakers. This improvement also results in better downstream ASR performance approaching that using oracle segments. Maokui He, Desh Raj, Zili Huang, Jun Du 0002, Zhuo Chen 0006, Shinji Watanabe 0001 |
Interspeech | 3 |
| 2021 | SUPERB: Speech Processing Universal PERformance BenchmarkabstractSelf-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various tasks with minimal adaptation.However, the speech processing community lacks a similar setup to systematically explore the paradigm.To bridge this gap, we introduce Speech processing Universal PERformance Benchmark (SUPERB).SUPERB is a leaderboard to benchmark the performance of a shared model across a wide range of speech processing tasks with minimal architecture changes and labeled data.Among multiple usages of the shared model, we especially focus on extracting the representation learned from SSL for its preferable re-usability.We present a simple framework to solve SUPERB tasks by learning task-specialized lightweight prediction heads on top of the frozen shared model.Our results demonstrate that the framework is promising as SSL representations show competitive generalizability and accessibility across SUPERB tasks.We release SUPERB as a challenge with a leaderboard 1 and a benchmark toolkit 2 to fuel the research in representation learning and general speech processing. Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li 0001, Shinji Watanabe 0001, Abdel-rahman Mohamed, Hung-yi Lee |
Interspeech | 15 |
| 2021 | Integration of Speech Separation, Diarization, and Recognition for Multi-Speaker Meetings: System Description, Comparison, and AnalysisabstractMulti-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, speaker diarization, and automatic speech recognition (ASR) in the last decade, it has become possible to build pipelines that achieve reasonable error rates on this task. In this paper, we propose an end-to-end modular system for the LibriCSS meeting data, which combines independently trained separation, diarization, and recognition components, in that order. We study the effect of different state-of-the-art methods at each stage of the pipeline, and report results using task-specific metrics like SDR and DER, as well as downstream WER. Experiments indicate that the problem of overlapping speech for diarization and ASR can be effectively mitigated with the presence of a well-trained separation module. Our best system achieves a speaker-attributed WER of 12.7%, which is close to that of a non-overlapping ASR. Desh Raj, Pavel Denisov, Zhuo Chen 0006, Hakan Erdogan, Zili Huang, Maokui He, Shinji Watanabe 0001, Jun Du 0002, Takuya Yoshioka, Yi Luo 0004, Naoyuki Kanda, Jinyu Li 0001, Scott Wisdom, John R. Hershey |
SLT | 5 |
| 2021 | DOVER-Lap: A Method for Combining Overlap-Aware Diarization OutputsabstractSeveral advances have been made recently towards handling overlapping speech for speaker diarization. Since speech and natural language tasks often benefit from ensemble techniques, we propose an algorithm for combining outputs from such diarization systems through majority voting. Our method, DOVER-Lap, is inspired from the recently proposed DOVER algorithm, but is designed to handle overlapping segments in diarization outputs. We also modify the pair-wise incremental label mapping strategy used in DOVER, and propose an approximation algorithm based on weighted k-partite graph matching, which performs this mapping using a global cost tensor. We demonstrate the strength of our method by combining outputs from diverse systems - clustering-based, region proposal networks, and target-speaker voice activity detection - on AMI and LibriCSS datasets, where it consistently outperforms the single best system. Additionally, we show that DOVER-Lap can be used for late fusion in multichannel diarization, and compares favorably with early fusion methods like beamforming. Desh Raj, L. Paola García-Perera, Zili Huang, Shinji Watanabe 0001, Daniel Povey, Andreas Stolcke, Sanjeev Khudanpur |
SLT | 3 |
| 2021 | Multi-Class Spectral Clustering with Overlaps for Speaker DiarizationabstractThis paper describes a method for overlap-aware speaker diarization. Given an overlap detector and a speaker embedding extractor, our method performs spectral clustering of segments informed by the output of the overlap detector. This is achieved by transforming the discrete clustering problem into a convex optimization problem which is solved by eigen-decomposition. Thereafter, we discretize the solution by alternatively using singular value decomposition and a modified version of non-maximal suppression which is constrained by the output of the overlap detector. Furthermore, we detail an HMM-DNN based overlap detector which performs frame-level classification and enforces duration constraints through HMM state transitions. Our method achieves a test diarization error rate (DER) of 24.0% on the mixed-headset setting of the AMI meeting corpus, which is a relative improvement of 15.2% over a strong agglomerative hierarchical clustering baseline, and compares favorably with other overlap-aware diarization methods. Further analysis on the LibriCSS data demonstrates the effectiveness of the proposed method in high overlap conditions. Desh Raj, Zili Huang, Sanjeev Khudanpur |
SLT | 2 |
| 2020 | Speaker Diarization with Region Proposal NetworkabstractSpeaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline. Zili Huang, Shinji Watanabe 0001, Yusuke Fujita, L. Paola García-Perera, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 1 |
| 2019 | Discriminative Neural Embedding Learning for Short-Duration Text-Independent Speaker VerificationabstractShort duration text-independent speaker verification remains a hot research topic in recent years, and deep neural network based embeddings have shown impressive results in such conditions. Good speaker embeddings require the property of both small intra-class variation and large inter-class difference, which is critical for the ability of discrimination and generalization. Current embedding learning strategies can be grouped into two frameworks: “Cascade embedding learning” with multiple stages and “direct embedding learning” from spectral feature directly. We propose new approaches to achieve more discriminant speaker embeddings. Within the cascade framework, a neural network based deep discriminant analysis (DDA) is proposed to project i-vector to more discriminative embeddings. Within the direct embedding framework, a deep model with more advanced center loss and A-softmax loss is used, the focal loss is also investigated in this framework. Moreover, the traditional i-vector and neural embeddings are finally combined with neural network based DDA to achieve further gain. Main experiments are carried out on a short-duration text-independent speaker verification dataset generated from the SRE corpus. The results show that the newly proposed method is promising for short-duration text-independent speaker verification, and it is consistently better than traditional i-vector and neural embedding baselines. The best embeddings achieve roughly 30% relative EER reduction compared to the i-vector baseline, which could be further enhanced when combined with the i-vector system. Shuai Wang 0016, Zili Huang, Yanmin Qian, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Joint I-Vector with End-to-End System for Short Duration Text-Independent Speaker VerificationabstractFactor analysis based i-vector has been the state-of-the-art method for speaker verification. Recently, researchers propose to build DNN based end-to-end speaker verification systems and achieve comparable performance withi-vector. Since these two methods possess their own property and differ from each other significantly, we explore a framework to integrate these two paradigms together to utilize their complementarity. More specifically, in this paper we develop and compare four methodologies to integrate traditionali-vector into end-to-end systems, including score fusion, embeddings concatenation, transformed concatenation and joint learning. All these approaches achieve significant gains. Moreover, the hard trial selection is performed on the end-to-end architecture which further improves the performance. Experimental results on a text-independent short-duration dataset generated from SRE 2010 reveal that the newly proposed method reduces the EER by relative 31.0% and 28.2% compared to the i-vector and end-to-end baselines respectively. Zili Huang, Shuai Wang 0016, Yanmin Qian |
ICASSP | 1 |
| 2018 | Angular Softmax for Short-Duration Text-independent Speaker Verification
Zili Huang, Shuai Wang 0016, Kai Yu 0004 |
INTERSPEECH | 1 |