VLDB 2026 Research / reviewers in the wild / expert
Hongbin Suo
dblp:99/5346
· DBLP profile ↗
20ranked-venue papers
0as first author
12since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 since 2021Artificial intelligence and machine learning · 14 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 4 |
| 2023 | Multi-channel multi-speaker transformer for speech recognitionabstractWith the development of teleconferencing and in-vehicle voice assistants, far-field multi-speaker speech recognition has become a hot research topic. Recently, a multi-channel transformer (MCT) has been proposed, which demonstrates the ability of the transformer to model far-field acoustic environments. However, MCT cannot encode high-dimensional acoustic features for each speaker from mixed input audio because of the interference between speakers. Based on these, we propose the multi-channel multi-speaker transformer (M2Former) for far-field multi-speaker ASR in this paper. Experiments on the SMS-WSJ benchmark show that the M2Former outperforms the neural beamformer, MCT, dual-path RNN with transform-average-concatenate and multi-channel deep clustering based end-to-end systems by 9.2%, 14.3%, 24.9%, and 52.2% respectively, in terms of relative word error rate reduction. Hongbin Suo, Yulong Wan |
INTERSPEECH | 3 |
| 2023 | Task-Agnostic Structured Pruning of Speech Representation Models
Haoyu Wang 0014, Siyuan Wang 0002, Weiqiang Zhang 0001, Hongbin Suo, Yulong Wan |
INTERSPEECH | 4 |
| 2023 | Robust Audio Anti-spoofing Countermeasure with Joint Training of Front-end and Back-end Models
Xingming Wang, Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 3 |
| 2023 | SEF-Net: Speaker Embedding Free Target Speaker Extraction Network
Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 2 |
| 2023 | Outlier-aware Inlier Modeling and Multi-scale Scoring for Anomalous Sound Detection via Multitask LearningabstractThis paper proposes an approach for anomalous sound detection that incorporates outlier exposure and inlier modeling within a unified framework by multitask learning. While outlier exposure-based methods can extract features efficiently, it is not robust. Inlier modeling is good at generating robust features, but the features are not very effective. Recently, serial approaches are proposed to combine these two methods, but it still requires a separate training step for normal data modeling. To overcome these limitations, we use multitask learning to train a conformer-based encoder for outlier-aware inlier modeling. Moreover, our approach provides multi-scale scores for detecting anomalies. Experimental results on the MIMII and DCASE 2020 task 2 datasets show that our approach outperforms state-of-the-art single-model systems and achieves comparable results with top-ranked multi-system ensembles. Yucong Zhang, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 2 |
| 2022 | Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting DataabstractUnsupervised clustering on speakers is becoming increasingly important for its potential uses in semi-supervised learning. In reality, we are often presented with enormous amounts of unlabeled data from multi-party meetings and discussions. An effective unsupervised clustering approach would allow us to significantly increase the amount of training data without additional costs for annotations. Recently, methods based on graph convolutional networks (GCN) have received growing attention for unsupervised clustering, as these methods exploit the connectivity patterns between nodes to improve learning performance. In this work, we present a GCN-based approach for semi-supervised learning. Given a pre-trained embedding extractor, a graph convolutional network is trained on the labeled data and clusters unlabeled data with "pseudo-labels". We present a self-correcting training mechanism that iteratively runs the cluster-train-correct process on pseudo-labels. We show that this proposed approach effectively uses unlabeled data and improves speaker recognition accuracy. Fuchuan Tong, Yafeng Chen, Hongbin Suo, Qingyang Hong, Lin Li 0032 |
ICASSP | 5 |
| 2022 | Reformulating Speaker Diarization As Community Detection With Emphasis On Topological StructureabstractClustering-based speaker diarization has stood firm as one of the major approaches in reality, despite recent development in end-to-end diarization. However, clustering methods have not been explored extensively for speaker diarization. Commonly-used methods such as k-means, spectral clustering, and agglomerative hierarchical clustering only take into account properties such as proximity and relative densities. In this paper we propose to view clustering-based diarization as a community detection problem. By doing so the topological structure is considered. This work has four major contributions. First it is shown that Leiden community detection algorithm significantly outperforms the previous methods on the clustering of speaker-segments. Second, we propose to use uniform manifold approximation to reduce dimension while retaining global and local topological structure. Third, a masked filtering approach is introduced to extract "clean" speaker embeddings. Finally, the community structure is applied to an end-to-end post-processing network to obtain diarization results. The final system presents a relative DER reduction of up to 70 percent. The breakdown contribution of each component is analyzed. Hongbin Suo |
ICASSP | 2 |
| 2022 | PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker VerificationabstractSpeaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to biased estimations, especially when handling shorter utterances. In this paper we propose to represent a speaker utterance as "floating" vector whose state is indeterminate without knowing the context. The state of a speaker representation is jointly determined by itself, other speech from the same speaker, as well as other speakers it is being compared to. The content of the speech also contributes to determining the final state of a speaker representation. We pre-train an indeterminate speaker representation model that estimates the state of an utterance based on the context. The pre-trained model can be fine-tuned for downstream tasks such as speaker verification, speaker clustering, and speaker diarization. Substantial improvements are observed across all downstream tasks. Hongbin Suo, Qian Chen 0003 |
INTERSPEECH | 2 |
| 2021 | Cam: Context-Aware Masking for Robust Speaker VerificationabstractPerformance degradation caused by noise has been a long-standing challenge for speaker verification. Previous methods usually involve applying a denoising transformation to speaker embeddings or enhancing input features. Nevertheless, these methods are lossy and inefficient for speaker embedding. In this paper, we propose context- aware masking (CAM), a novel method to extract robust speaker embedding. CAM enables the speaker embedding network to "focus" on the speaker of interest and "blur" unrelated noise. The threshold of masking is dynamically controlled by an auxiliary context embedding that captures speaker and noise characteristics. Moreover, models adopting CAM can be trained in an end-to-end manner without using synthesized noisy-clean speech pairs. Our results show that CAM improves speaker verification performance in the wild by a large margin, compared to the baselines. Ya-Qi Yu, Hongbin Suo, Yun Lei, Wu-Jun Li |
ICASSP | 3 |
| 2021 | A Real-Time Speaker Diarization System Based on Spatial SpectrumabstractIn this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in speaker diarization tasks: (1) to segment and separate overlapping speech from two speakers; (2) to estimate the number of speakers when participants may enter or leave the conversation at any time; (3) to provide accurate speaker identification on short text-independent utterances; (4) to track down speakers movement during the conversation; (5) to detect speaker change incidence real-time. First, a differential directional microphone array-based approach is exploited to capture the target speakers’ voice in far-field adverse environment. Second, an online speaker-location joint clustering approach is proposed to keep track of speaker location. Third, an instant speaker number detector is developed to trigger the mechanism that separates overlapped speech. The results suggest that our system effectively incorporates spatial information and achieves significant gains. Weilong Huang, Xianliang Wang, Hongbin Suo, Jinwei Feng, Zhijie Yan |
ICASSP | 4 |
| 2021 | Investigation of Spatial-Acoustic Features for Overlapping Speech Detection in Multiparty Meetings
Shiliang Zhang, Weilong Huang, Hongbin Suo, Jinwei Feng, Zhijie Yan |
Interspeech | 5 |
| 2020 | Phonetically-Aware Coupled Network For Short Duration Text-Independent Speaker Verification
Yun Lei, Hongbin Suo |
INTERSPEECH | 3 |
| 2019 | Towards a Fault-Tolerant Speaker Verification System: A Regularization Approach to Reduce the Condition Number
Hongbin Suo, Yun Lei |
INTERSPEECH | 3 |
| 2019 | Autoencoder-Based Semi-Supervised Curriculum Learning for Out-of-Domain Speaker Verification
Hongbin Suo, Yun Lei |
INTERSPEECH | 3 |
| 2012 | Factor analysis of Laplacian approach for speaker recognitionabstractIn this study, we introduce a new factor analysis of Laplacian approach to speaker recognition under the support vector machine (SVM) framework. The Laplacian-projected supervector from our proposed Laplacian approach, which finds an embedding that preserves local information by locality preserving projections (LPP), is believed to contain speaker dependent information. The proposed method was compared with the state-of-the-art total variability approach on 2010 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) corpus. According to the compared results, our proposed method is effective. Jinchao Yang, Chunyan Liang, Lin Yang 0016, Hongbin Suo, Yonghong Yan 0002 |
ICASSP | 4 |
| 2012 | Maximum A Posteriori Linear Regression for language recognition
Jinchao Yang, Xiang Zhang 0014, Hongbin Suo, Jianping Zhang 0001, Yonghong Yan 0002 |
Expert Syst. Appl. | 3 |
| 2010 | Speaker recognition using the resynthesized speech via spectrum modeling
Xiang Zhang 0014, Chuan Cao, Lin Yang 0016, Hongbin Suo, Jianping Zhang 0001, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2009 | A Novel Fuzzy-Based Automatic Speaker Clustering Algorithm
Xiang Zhang 0014, Hongbin Suo, Qingwei Zhao, Yonghong Yan 0002 |
ISNN (2) | 3 |
| 2007 | Spoken language identification using score vector modeling and support vector machineabstractThe support vector machine (SVM) framework based on generalized linear discriminate sequence (GLDS) kernel has been shown effective and widely used in language identifica-tion tasks. In this paper, in order to compensate the distortions due to inter-speaker variability within the same language and solve the practical limitation of computer memory requested by large database training, multiple speaker group based discrim-inative classifiers are employed to map the cepstral features of speech utterances into discriminative language characterization score vectors (DLCSV). Furthermore, backend SVM classifiers are used to model the probability distribution of each target language in the DLCSV space and the output scores of back-end classifiers are calibrated as the final language recognition scores by a pair-wise posterior probability estimation algorithm. The proposed SVM framework is evaluated on 2003 NIST Lan-guage Recognition Evaluation databases, achieving an equal er-ror rate of 4.0 % in 30-second tasks, which outperformed the state-of-art SVM system by more than 30 % relative error re-duction. Index Terms: spoken language identification, support vector machine, score vector modeling Ming Li 0026, Hongbin Suo, Ping Lu 0009, Yonghong Yan 0002 |
INTERSPEECH | 2 |