VLDB 2026 Research / reviewers in the wild / expert
Binghuai Lin
dblp:146/2946
· DBLP profile ↗
33ranked-venue papers
17as first author
28since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 17 first-author · 20 since 2021Artificial intelligence and machine learning · 23 · 8 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HeadMap: Locating and Enhancing Knowledge Circuits in LLMsabstractLarge language models (LLMs), through pretraining on extensive corpora, encompass rich semantic knowledge and exhibit the potential for efficient adaptation to diverse downstream tasks. However, the intrinsic mechanisms underlying LLMs remain unexplored, limiting the efficacy of applying these models to downstream tasks. In this paper, we explore the intrinsic mechanisms of LLMs from the perspective of knowledge circuits. Specifically, considering layer dependencies, we propose a layer-conditioned locating algorithm to identify a series of attention heads, which is a knowledge circuit of some tasks. Experiments demonstrate that simply masking a small portion of attention heads in the knowledge circuit can significantly reduce the model's ability to make correct predictions. This suggests that the knowledge flow within the knowledge circuit plays a critical role when the model makes a correct prediction. Inspired by this observation, we propose a novel parameter-efficient fine-tuning method called HeadMap, which maps the activations of these critical heads in the located knowledge circuit to the residual stream by two linear layers, thus enhancing knowledge flow from the knowledge circuit in the residual stream. Extensive experiments conducted on diverse datasets demonstrate the efficiency and efficacy of the proposed method. Our code is available at https://github.com/XuehaoWangFi/HeadMap. Xuehao Wang, Binghuai Lin |
ICLR | 3 |
| 2024 | Large Language Models are not Fair EvaluatorsabstractPeiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 6 |
| 2024 | PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQLabstractLarge Language Models (LLMs) have emerged as powerful tools for Text-to-SQL tasks, exhibiting remarkable reasoning capabilities.Different from tasks such as math word problems and commonsense reasoning, SQL solutions have a relatively fixed pattern.This facilitates the investigation of whether LLMs can benefit from categorical thinking, mirroring how humans acquire knowledge through inductive reasoning based on comparable examples.In this study, we propose that employing query group partitioning allows LLMs to focus on learning the thought processes specific to a single problem type, consequently enhancing their reasoning abilities across diverse difficulty levels and problem categories.Our experiments reveal that multiple advanced LLMs, when equipped with PTD-SQL, can either surpass or match previous state-of-theart (SOTA) methods on the Spider and BIRD datasets.Intriguingly, models with varying initial performances have exhibited significant improvements, mainly at the boundary of their capabilities after targeted drilling, suggesting a parallel with human progress.Code is available at https://github.com/lrlbbzl/PTD-SQL. Ruilin Luo, Binghuai Lin, Zicheng Lin, Yujiu Yang 0001 |
EMNLP | 3 |
| 2024 | A Study of Mispronunciation Detection and Diagnosis Based on Meta-LearningabstractThe majority of the current mispronunciation detection and diagnosis (MD&D) methods rely on manually annotated data for model training. However, annotating mispronunciations produced by second language (L2) learners is costly. Consequently, data scarcity emerges as a significant challenge in MD&D tasks. In this paper, we employ model-agnostic meta-learning (MAML) to train a phoneme recognition model for MD&D. We conduct experiments using varied meta-learning task partitioning and training strategies to endow the model’s ability to rapidly adapt to unfamiliar speakers. Our best-performing method achieves an F-measure of 61.45%, surpassing both the method using fine-tuned pre-trained model wav2vec2.0 and the approach of incorporating reference text during training. These related works also aim to address the challenge of data scarcity in MD&D. Notably, with few-shot fine-tuning, our model still yielded some remarkable results on F-measure, which suggest that in MD&D tasks, meta-learning is indeed effective. Yukai Wan, Yuqi Shi, Binghuai Lin, Yanlu Xie |
ICASSP | 3 |
| 2024 | DialogVCS: Robust Natural Language Understanding in Dialogue System UpgradeabstractZefan Cai, Xin Zheng, Tianyu Liu, Haoran Meng, Jiaqi Han, Gang Yuan, Binghuai Lin, Baobao Chang, Yunbo Cao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zefan Cai, Tianyu Liu 0001, Haoran Meng, Binghuai Lin, Baobao Chang, Yunbo Cao |
NAACL-HLT | 7 |
| 2023 | Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionabstractShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui |
ACL (1) | 5 |
| 2023 | Soft Language Clustering for Multilingual Model Pre-trainingabstractJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016 |
ACL (1) | 6 |
| 2023 | Multi-Lingual Pronunciation Assessment with Unified Phoneme Set and Language-Specific EmbeddingsabstractAutomatic pronunciation assessment is commonly trained and applied for a specific language, which is not practical in multi-lingual or low-resource scenarios. In this paper, we propose a unified method to take advantage of multi-lingual data for multi-lingual pronunciation assessment. To this end, we first construct a concise unified phoneme set for multi-lingual phoneme recognition based on a pre-trained acoustic model. In this way we can not only share language-independent knowledge but also try to discriminate language-specific information for pronunciation assessment. Second, we employ language-specific embeddings for different languages, which act like language-specific assessment criteria to adaptively adjust the feature weights based on an attention mechanism. The whole network is optimized in a unified framework. Experimental results based on multi-lingual datasets demonstrate its superiority to different baselines in Pearson correlation coefficient (PCC). We also illustrate the generalizability of the proposed method for both seen and unseen data. Binghuai Lin |
ICASSP | 1 |
| 2023 | Robust multi-modal speech emotion recognition with ASR error adaptationabstractMulti-modal speech emotion recognition (SER) brings performance improvement compared with single-modal systems. However, the ASR errors from text modality may deteriorate the SER performance. This paper proposes an SER method robust to ASR errors. We explore complementary semantic in-formation from the audio to reduce the impact of ASR errors, which is done by an attention mechanism to calculate weighted acoustic representations. This information is fused with the text representations of ASR hypotheses utilizing an adaptive weight, which is determined by an auxiliary ASR error detection task. Finally, the fused text representations are concatenated with acoustic representations to perform SER. Results based on the public Emotional Dyadic Motion Capture (IEMOCAP) dataset show when using ASR hypotheses with high word error rate (WER), the proposed method is proved to be robust with very slight performance drops compared to traditional multi-modal models. We further demonstrate its robustness using ASR hypotheses with different WERs. Binghuai Lin |
ICASSP | 1 |
| 2023 | Multi-modal ASR error correction with joint ASR error detectionabstractTo tackle the recognition error problem for Automatic speech recognition (ASR), one common approach is to utilize text-based ASR error correction methods focusing on text error pat-terns. To include the audio information for better error correction, we propose a sequence-to-sequence multi-modal ASR error correction model. The multi-modal representations from pre-trained audio and text encoders are fused and aligned based on an attention mechanism. The decoder then searches for the corresponding correction results based on the fused representations. To better explore the correlations between different modalities, an additional ASR error detection task is applied on top of the fused representations. We optimize the network by a multi-task learning method combining both ASR error detection and correction tasks. Experimental results based on a 200-hour dataset recorded by Chinese English-as-second-language (ESL) learners show the proposed correction model can achieve significant improvement compared to the baselines with or without other ASR error correction methods. Binghuai Lin |
ICASSP | 1 |
| 2022 | Learning Robust Representations for Continual Relation Extraction via Adversarial Class AugmentationabstractContinual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks. Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui |
EMNLP | 4 |
| 2022 | HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy.Recently, the pretrained language models (PLM) have been widely adopted in HTC through a finetuning paradigm.However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pretraining tasks of PLMs and thus the potential of PLMs cannot be fully tapped.To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective.Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM.Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations. Peiyi Wang, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui, Houfeng Wang |
EMNLP | 4 |
| 2022 | Phoneme Mispronunciation Detection By Jointly Learning To AlignabstractPhoneme mispronunciation detection plays an important role in Computer-Assisted Pronunciation Training. Traditional methods either rely on phone recognition, which has the limitation of incapability of detecting mispronounced phonemes out of the dictionary, or the need of external phoneme alignment for extracting acoustic features. In this paper, we propose a method for phoneme mispronunciation detection by jointly learning to align. Specifically, we first obtain acoustic and canonical phoneme representations utilizing acoustic and phoneme encoders. Second, we utilize an attention mechanism to fuse acoustic features of each frame and phoneme representations. Finally, a convolutional neural network (CNN)-based layer following the fused representations is utilized for better exploring local context. The network is jointly optimized for phoneme mispronunciation and phoneme alignment based on a multi-task learning framework. Experimental results based on a public dataset L2-ARCTIC show the state-of-the-art (SOTA) performance with an F1-score of 63.04%. It is also found that optimizing the phoneme alignment can further improve the performance of phoneme mispronunciation detection. Binghuai Lin |
ICASSP | 1 |
| 2022 | Fast Task-Specific Adaptation in Spoken Language Assessment with Meta-LearningabstractAutomatic spoken language assessment plays an important role in assessing English proficiency of non-native learners, which involves tasks ranging from restricted tasks such as Repeat Sentences to more open-ended tasks such as unconstrained spontaneous speech. Traditional methods typically focus on specific task types and rely on a significant amount of human-labelled data. In this paper, we propose a fast adaptation framework with meta-learning for various task types in spoken language assessment under low-resource settings. To better adapt to tasks with different grading criteria, we incorporate a memory network acting as an external memory for these criteria. Experimental results based on data from different spoken language tests demonstrate the superiority of the proposed method to the baselines in Pearson correlation coefficient and accuracy when adapted to various task types, especially in low-resource settings. Binghuai Lin |
ICASSP | 1 |
| 2022 | Learning Acoustic Frame Labeling for Phoneme Segmentation with Regularized Attention MechanismabstractPhoneme segmentation plays an important role in various speech processing applications such as keyword spotting, automatic pronunciation assessment, and automatic speech recognition. In this paper, we propose a method for phoneme segmentation based on a regularized attention mechanism. Specifically, the representations of speech utterance for each frame are extracted from a pre-trained acoustic encoder and combined with presumed phoneme sequences based on the attention mechanism. By fusing acoustic representations with these aligned phoneme representations, we learn phoneme labeling for each frame to obtain final segmentation. For better alignment between the pronounced phoneme sequence and utterance, we regularize the attention matrix utilizing an extra attention loss. The whole network is optimized by a multi-task learning framework (MTL). Experimental results based on the TIMIT and Buckeye corpora show the proposed method is superior to the previous baselines and reaches the state-of-the-art (SOTA) performance in F1 score and R-value. Binghuai Lin |
ICASSP | 1 |
| 2022 | Exploiting Information From Native Data for Non-Native Automatic Pronunciation AssessmentabstractThis paper proposes an end-to-end pronunciation assessment method to exploit the adequate native data and reduce the need for non-native data costly to label. To obtain discriminative acoustic representations at the phoneme level, the pretrained wav2vec 2.0 is re-trained with connectionist temporal classification (CTC) loss for phoneme recognition using native data. These acoustic representations are fused with phoneme representations derived from a phoneme encoder to obtain final pronunciation scores. An efficient fusion mechanism aligns each phoneme with acoustic frames based on attention, where all blank frames recognized by the CTC-based phoneme recognition are masked. Finally, the whole network is optimized by a multi-task learning framework combining CTC loss and mean square error loss between predicted and human scores. Extensive experiments demonstrate that it outperforms previous baselines in the Pearson correlation coefficient even with much fewer labeled non-native data. Binghuai Lin |
SLT | 1 |
| 2021 | Uncertainty-Aware Pseudo-Labeling for Spoken Language AssessmentabstractAutomatic spoken language assessment has gained popularity in computer-assisted language learning (CALL). Normally building these systems relies heavily on labor-intensive human-labeled data. In this paper, we adopt the pseudo-labeling (PL) method to make better use of the massive amount of unlabeled data. Traditional pseudo-labeling mechanism selects predictions of high confidence from the unlabeled pool to augment the labeled data, which may be over-confident and less informative. To select more informative, less noisy data, we optimize the pseudo-labeling process by modeling both data and knowledge uncertainty. Based on a self-adaptive learning mechanism, we sequentially and iteratively select unlabeled samples to augment the training sets during the multi-step self-training procedure until the stop criterion is satisfied. Experimental results based on data from the spoken English tests demonstrate superior performance compared to the baselines and other semi-supervised learning (SSL) methods in Pearson correlation coefficient (PCC) and accuracy. Binghuai Lin |
ASRU | 1 |
| 2021 | Attention-Based Multi-Encoder Automatic Pronunciation AssessmentabstractAutomatic pronunciation assessment plays an important role in Computer-Assisted Pronunciation Training (CAPT). Traditional methods for pronunciation assessment of reading aloud tasks utilize features derived from automatic speech recognition (ASR) and thus are sensitive to the accuracy of ASR and the effectiveness of features. Moreover, the representation capability of the features is also affected by the inconsistent optimization goals between the ASR and scoring tasks. In this paper we propose an end-to-end (E2E) pronunciation scoring network based on attention mechanism and multi-encoder consisting of audio and text encoders. The network optimized by a multi-task learning (MTL) framework can provide scoring at sentence-level as well as detailed scoring at word-level. Due to data scarcity for pronunciation scoring, we utilize ASR data and synthetic data to pre-train the network in two steps, and then fine-tune the network using the limited high-quality scoring data. Experimental results based on the dataset recorded by Chinese English-as-second-language (ESL) learners and labeled by three experts demonstrate that the proposed model outperforms the baseline in Pearson correlation coefficient (PCC). Binghuai Lin |
ICASSP | 1 |
| 2021 | F0 Patterns of L2 English Speech by Mandarin Chinese Learners
Binghuai Lin |
Interspeech | 2 |
| 2021 | A Noise Robust Method for Word-Level Pronunciation Assessment
Binghuai Lin |
Interspeech | 1 |
| 2021 | End to End Transformer-Based Contextual Speech Recognition Based on Pointer Network
Binghuai Lin |
Interspeech | 1 |
| 2021 | A Neural Network-Based Noise Compensation Method for Pronunciation Assessment
Binghuai Lin |
Interspeech | 1 |
| 2021 | Deep Feature Transfer Learning for Automatic Pronunciation Assessment
Binghuai Lin |
Interspeech | 1 |
| 2021 | A Study on Fine-Tuning wav2vec2.0 Model for the Task of Mispronunciation Detection and Diagnosis
Linkai Peng, Kaiqi Fu, Binghuai Lin, Dengfeng Ke, Jinsong Zhang 0001 |
Interspeech | 3 |
| 2021 | Explore wav2vec 2.0 for Mispronunciation Detection
Xiaoshuo Xu, Yueteng Kang, Songjun Cao, Binghuai Lin |
Interspeech | 4 |
| 2021 | A Preliminary Study on Discourse Prosody Encoding in L1 and L2 English Spontaneous Narratives
Yuqing Zhang 0003, Binghuai Lin, Jinsong Zhang 0001 |
Interspeech | 3 |
| 2021 | Relationships Between Perceptual Distinctiveness, Articulatory Complexity and Functional Load in Speech Communication
Yuqing Zhang 0003, Yanlu Xie, Binghuai Lin, Jinsong Zhang 0001 |
Interspeech | 5 |
| 2021 | Improving L2 English Rhythm Evaluation with Automatic Sentence Stress DetectionabstractEnglish is a stress-timed language, for which sentence stress or prosodic stress plays an important role. It's then difficult for Chinese who are used to the syllable-timed rhythm to learn the rhythm of English [1]. In this paper, we investigate how to improve the rhythm evaluation based on the sentence stress for Chinese who learn English as a second language (ESL). Particularly, we explore some rhythm measures to quantify rhythmic differences among second language (L2) learners based on sentence stress. To relieve the dependency on labeled data of sentence stress, we predict sentence stress automatically utilizing a hierarchical network with bidirectional Long Short-Term Memory (BLSTM) [2]. We evaluate the proposed method based on the corpus consisting of 3,500 sentences recorded by 100 Chinese speakers aging from 10 to 20 years old, which was marked with the sentence stress labels and scored by three experts. Experimental results show the proposed sentence stress measure is well correlated with labeled prosody scores with a correlation coefficient of -0.73 and the automatic labeling method achieves comparable results with the method with gold labels. Binghuai Lin, Xiaoli Feng |
SLT | 1 |
| 2020 | Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating MechanismabstractFormant tracking is one of the most fundamental problems in speech processing.Traditionally, formants are estimated using signal processing methods.Recent studies showed that generic convolutional architectures can outperform recurrent networks on temporal tasks such as speech synthesis and machine translation.In this paper, we explored the use of Temporal Convolutional Network (TCN) for formant tracking.In addition to the conventional implementation, we modified the architecture from three aspects.First, we turned off the "causal" mode of dilated convolution, making the dilated convolution see the future speech frames.Second, each hidden layer reused the output information from all the previous layers through dense connection.Third, we also adopted a gating mechanism to alleviate the problem of gradient disappearance by selectively forgetting unimportant information.The model was validated on the open access formant database VTR.The experiment showed that our proposed model was easy to converge and achieved an overall mean absolute percent error (MAPE) of 8.2% on speech-labeled frames, compared to three competitive baselines of 9.4% (LSTM), 9.1% (Bi-LSTM) and 8.9% (TCN). Wang Dai, Jinsong Zhang 0001, Yingming Gao, Dengfeng Ke, Binghuai Lin, Yanlu Xie |
INTERSPEECH | 6 |
| 2020 | A Comparison of English Rhythm Produced by Native American Speakers and Mandarin ESL Primary School Learners
Binghuai Lin, Ruomei Fang |
INTERSPEECH | 2 |
| 2020 | Joint Prediction of Punctuation and Disfluency in Speech Transcripts
Binghuai Lin |
INTERSPEECH | 1 |
| 2020 | Automatic Scoring at Multi-Granularity for L2 Pronunciation
Binghuai Lin, Xiaoli Feng, Jinsong Zhang 0001 |
INTERSPEECH | 1 |
| 2020 | Joint Detection of Sentence Stress and Phrase Boundary for Prosody
Binghuai Lin, Xiaoli Feng, Jinsong Zhang 0001 |
INTERSPEECH | 1 |