Chang Su 0001

dblp:07/2757-1 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang
Frontiers Comput. Sci.11
2025 Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
abstract
Log analysis represents a critical sub-domain within AI applications that facilitates automatic approaches to fault and error management of large-scaled software systems, saving labors of traditional manual methods. While existing solutions using large language models (LLMs) show promise, they are limited by a significant domain gap between natural and log languages (the latter contains rich domain-specific tokens such as status codes, IP addresses, resource pathes), which restricts their effectiveness in real-world applications. However, directly adapting general-purpose LLMs to log analysis using raw logs may degrade their performance due to inconsistent token distribution. In this paper, we present a domain adaptation approach that addresses these limitations by integrating interpretable domain knowledge into open-source LLMs through continual pre-training (CPT), which bridges this domain gap by adapting LLMs on interpretable natural texts with log knowledge (instead of raw logs) to reduce distribution discrepancy. To achieve this, we developed NLPLog, a comprehensive dataset containing over 250,000 question-answer pairs on log-related knowledge. Our resulting model, SuperLog, achieves the best performance across four log analysis tasks, with an average accuracy improvement of 12.01% over the second-best model. Ablation study also suggests advantages of domain adaption using interpretable log knowledge over using raw logs.
Yuhe Ji, Yilun Liu 0001, Feiyu Yao, Minggui He, Shimin Tao, Chang Su 0001, Xinhua Yang, Weibin Meng, Yuming Xie, Boxing Chen, Shenglin Zhang, Yongqian Sun
CIKM7
2025 SuperFC: Selective Data Utilization for a Sustainable and Effective Function-Calling Agent
abstract
The function-calling agent is obtained by performing agent tuning to the large language model (LLM) on function-calling dataset. However, even state-of-the-art datasets (e.g., xlam-function-calling-60k datasets) still contain numerous misleading examples of low-quality data, wasting significant computational resources and result in an unnecessary carbon footprint. Furthermore, such inductive bad data negatively impacts the performance of the agent. In this paper, we propose a set of scoring criteria specifically tailored to evaluate function-calling data and use these criteria to develop a data filtering framework. By applying this framework to filter out low-quality data, we fine-tuned SuperFC, which demonstrates substantial improvements in both sustainability and performance. The SuperFC-7B training process reduced training time from 455 minutes to 85 minutes, resulting in a 80.02% reduction in carbon footprint. Simultaneously, fine-tuning on high-quality data subsets led to performance improvements of up to 3.68%. Additionally, we provide an in-depth analysis of the causes behind the low quality of synthetic function-calling data, offering valuable insights for future data synthesis in this domain. We have also released a high-quality function-calling dataset, available at: https://github.com/Zire-Young/SuperFC
Xinhua Yang, Yilun Liu 0001, Shimin Tao, Chunguang Zhao, Weibin Meng, Minggui He, Chang Su 0001, Hongxia Ma, Jingzhou Du, Hao Yang 0006, Boxing Chen, Chuanwen Li
IJCNN7
2024 CB-Whisper: Contextual Biasing Whisper Using Open-Vocabulary Keyword-Spotting
abstract
End-to-end automatic speech recognition (ASR) systems often struggle to recognize rare name entities, such as personal names, organizations and terminologies that are not frequently encountered in the training data. This paper presents Contextual Biasing Whisper (CB-Whisper), a novel ASR system based on OpenAI’s Whisper model that can recognize user-defined name entities by performing open-vocabulary keyword-spotting (KWS) before the decoder. The KWS module leverages text-to-speech (TTS) techniques and a convolutional neural network (CNN) classifier to match the features between the entities and the utterances. To integrate the recognized entities into the Whipser decoder and avoid hallucinations, we carefully crafted multiple prompts with spoken form hints. Experiments show that the KWS module based on Whisper encoder’s features can recognize unseen user-defined keywords effectively. More importantly, the proposed CB-Whisper substantially improves the mixed-error-rate (MER) and entity recall compared to the original Whisper model on three internal datasets and two publicly available datasets including Aishell and ACL datasets that cover English-only, Chinese-only, and code-switching scenarios.
Yuang Li, Yinglu Li, Min Zhang 0042, Chang Su 0001, Mengyao Piao, Xiaosong Qiao, Miaomiao Ma, Hao Yang 0006
LREC/COLING4
2024 CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
abstract
Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set11https://github.com/lunyiliu/CoachLM.
Yilun Liu 0001, Shimin Tao, Ming Zhu 0010, Wenbing Ma, Chang Su 0001, Yutai Hou, Min Zhang 0042, Hongxia Ma, Hao Yang 0006, Yanfei Jiang
ICDE7
2024 From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
abstract
Machine Translation Quality Estimation (MTQE) is the task of estimating the quality of machine-translated text in real time without the need for reference translations, which is of great importance for the development of MT. After two decades of evolution, QE has yielded a wealth of results. This article provides a comprehensive overview of QE datasets, annotation methods, shared tasks, methodologies, challenges, and future research directions. It begins with an introduction to the background and significance of QE, followed by an explanation of the concepts and evaluation metrics for word-level QE, sentence-level QE, document-level QE, and explainable QE. The paper categorizes the methods developed throughout the history of QE into those based on handcrafted features, deep learning, and Large Language Models (LLMs), with a further division of deep learning-based methods into classic deep learning and those incorporating pre-trained language models (LMs). Additionally, the article details the advantages and limitations of each method and offers a straightforward comparison of different approaches. Finally, the paper discusses the current challenges in QE research and provides an outlook on future research directions.
Haofei Zhao, Yilun Liu 0001, Shimin Tao, Weibin Meng, Xiang Geng, Chang Su 0001, Min Zhang 0042, Hao Yang 0006
IJCNN7
2024 A Multitask Training Approach to Enhance Whisper with Open-Vocabulary Keyword Spotting
abstract
The recognition of rare named entities, such as personal names and terminologies, is challenging for automatic speech recognition (ASR) systems, especially when they are not frequently observed in the training data.In this paper, we introduce keyword spotting enhanced Whisper (KWS-Whisper), a novel ASR system that leverages the Whisper model and performs openvocabulary keyword spotting (OV-KWS) on the hidden states of the Whisper encoder to recognize user-defined named entities.These entities serve as prompts for the Whisper decoder.To optimize the model, we propose a multitask training approach that learns OV-KWS and contextual-ASR tasks.We evaluate our approach on Chinese Aishell hot word subsets and two internal code-switching test sets and show that it significantly improves the entity recall compared to the original Whisper model.Moreover, we demonstrate that the OV-KWS can be a plug-andplay module to enhance the ASR error correction methods and frozen Whisper models.
Yuang Li, Min Zhang 0042, Chang Su 0001, Yinglu Li, Xiaosong Qiao, Mengxin Ren, Miaomiao Ma, Daimeng Wei, Shimin Tao, Hao Yang 0006
INTERSPEECH3
2023 Knowledge Prompt for Whisper: An ASR Entity Correction Approach with Knowledge Base
abstract
Entity correction is crucial in Automatic Speech TABLE I Recognition (ASR), since erroneous entities seriously affect our understanding of ASR results. In this paper, in order to correct entity errors, we propose a knowledge prompt approach for Whisper (a recent ASR model trained with a corpus containing 680k hours of labeled speech recorded in various conditions). For a given audio, our approach consists of three steps: (1) obtaining its ASR result by Whisper; (2) fuzzy matching the ASR result with a knowledge base to obtain candidate entities; (3) using the candidate entities as a prompt to obtain the final ASR result by Whisper again. We conduct experiments on the test dataset of open-source Chinese speech corpus AISHELLNER. Experimental results show that our approach not only significantly improves the entity recall rate in ASR results (from 70.97% to 84.82%), but also reduces the overall Character Error Rate (CER).
Min Zhang 0042, Xiaosong Qiao, Chang Su 0001, Yinglu Li, Yuang Li, Ming Zhu 0010, Mengyao Piao, Shimin Tao, Hao Yang 0006, Yanfei Jiang
IEEE Big Data4
2023 DA-Parser: A Pre-trained Domain-aware Parsing Framework for Heterogeneous Log Analysis
abstract
Automated log analysis is widely applied in modern software-intensive systems to ensure resilience and sustainability, where log parsing is a vital initial step, converting unstructured logs into structured data for downstream analysis. However, traditional log parsing algorithms are designed to process logs within a single domain. As cross-domain dependencies and interactions between sub-modules of software systems increase, these algorithms struggle to handle the challenges posed by multi-domain log inputs, which results in a significant decline in parsing accuracy when facing heterogeneous logs. Additionally, current solutions for heterogeneous log parsing require extensive manual labeling efforts. In this paper, we propose Domain-aware Parser (DA-Parser), a framework that consists of a domain-aware head to identify the source domains of heterogeneous logs and then converts the multi-domain log parsing problem into a series of single-domain parsing problems. The domain-aware head is pretrained using a corpus of logs from 16 domains, which allows for the classification of the source domains of most heterogeneous log set without additional human labeling. Source domain tags predicted by the domain-aware head serve as a constraint to limit the template extraction process to logs from the same domain. Empirical evaluation is conducted on a multi-domain dataset containing logs from 7 domains. DA-Parser can be integrated with existing single-domain algorithms and are compatible with them, achieving superior parsing accuracy with an average of 9.26% improvement compared with single-domain algorithms.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Jingyu Wang 0001, Chang Su 0001, Weinan Tian, Min Zhang 0042, Hao Yang 0006, Xun Chen 0001
COMPSAC6
2023 UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
abstract
Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets.
Minghan Wang, Xiaosong Qiao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP9
2023 Zephyr: Zero-Shot Punctuation Restoration
abstract
Punctuation restoration can be crucial for the cascade speech translation system. Traditional approaches typically treat it as a sequential tagging problem, predicting which punctuation mark should follow a given word. However, this often requires significant computational and storage resources for full-stage training or fine-tuning. Our argument is that pre-trained language models (PLMs) can directly leverage their learned knowledge for punctuation generation, making additional training unnecessary. In this paper, we propose the Zephyr algorithm, which utilizes PLMs to perform zero-shot and few-shot punctuation restoration for both offline and streaming scenarios. Our experimental results demonstrate that, in comparison to fine-tuning-based baselines, Zephyr achieves competitive performance while requiring little to no training cost and exhibiting better generalizability in zeroshot and few-shot settings.
Minghan Wang, Yinglu Li, Xiaosong Qiao, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP5
2023 Biglog: Unsupervised Large-scale Pre-training for a Unified Log Representation
abstract
Automated log analysis has been widely applied in modern data-center network, performing critical tasks such as log parsing, log anomaly detection and log-based failure prediction. However, existing approaches rely on hand-crafted features or domain-specific vectors to represent logs, which are either laborious in manual efforts or ineffective facing multiple domains in a system. Furthermore, general-purpose word embeddings are not optimized for log data, thus are data-inefficient in handling complex log analysis tasks. In this paper, we present a pre-training phase for language models to understand both in-sentence and cross-sentence features of logs, resulting in a unified representation of logs that is well-suited for various downstream analysis tasks. The pre-training phase is unsupervised, utilizing 0.45 billion logs from 16 diverse domains. Experiments on 12 publicly available evaluation datasets across 3 tasks indicate superiority of our approach against existing approaches, especially in online scenarios with limited historical logs. Our approach also exhibits remarkable few-shot learning ability and domain-adaptiveness, which not only outperforms existing approaches using only 0.0025% of their required training data, but also adapts into new domains via only a few in-domain logs. We release our code and pre-trained model.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Zuomin Ren, Hao Yang 0006, Xun Chen 0001, Yuming Xie, Chang Su 0001, Xiaosong Oiao, Weinan Tian, Yichen Zhu 0001
IWQoS9
2023 HWCGEC:HW-TSC's 2023 Submission for the NLPCC2023's Chinese Grammatical Error Correction Task
Chang Su 0001, Xiaosong Qiao, Min Zhang 0042, Hao Yang 0006, Ming Zhu 0010, Wenbing Ma
NLPCC (3)1
2022 Diformer: Directional Transformer for Neural Machine Translation
abstract
Autoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current combination frameworks focus more on the integration of multiple decoding paradigms with a unified generative model, e.g. Masked Language Model. However, the generalization can be harmful on the performance due to the gap between training objective and inference. In this paper, we aim to close the gap by preserving the original objective of AR and NAR under a unified framework. Specifically, we propose the Directional Transformer (Diformer) by jointly modelling AR and NAR into three generation directions (left-to-right, right-to-left and straight) with a newly introduced direction variable, which works by controlling the prediction of each token to have specific dependencies under that direction. The unification achieved by direction successfully preserves the original dependency assumption used in AR and NAR, retaining both generalization and performance. Experiments on 4 WMT benchmarks demonstrate that Diformer outperforms current united-modelling works with more than 1.5 BLEU points for both AR and NAR decoding, and is also competitive to the state-of-the-art independent AR and NAR models.
Minghan Wang, Yuxia Wang 0003, Daimeng Wei, Hengchao Shang, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
EAMT7
2021 HI-CMLM: Improve CMLM with Hybrid Decoder Input
abstract
Mask-predict CMLM (Ghazvininejad et al., 2019) has achieved stunning performance among non-autoregressive NMT models, but we find that the mechanism of predicting all of the target words only depending on the hidden state of [MASK] is not effective and efficient in initial iterations of refinement, resulting in ungrammatical repetitions and slow convergence.In this work, we mitigate this problem by combining copied source with embeddings of [MASK] in decoder.Notably.it's not a straightforward copying that is shown to be useless, but a novel heuristic hybrid strategy -fence-mask.Experimental results show that it gains consistent boosts on both WMT14 En↔De and WMT16 En↔Ro corpus by 0.5 BLEU on average, and 1 BLEU for lessinformative short sentences.This reveals that incorporating additional information by proper strategies is beneficial to improve CMLM, particularly translation quality of short texts and speeding up early-stage convergence.
Minghan Wang, Yuxia Wang 0003, Chang Su 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
INLG5
2021 Make the Blind Translator See The World: A Novel Transfer Learning Solution for Multimodal Machine Translation
abstract
Based on large-scale pretrained networks and the liability to be easily overfitting with limited labelled training data of multimodal translation (MMT) is a critical issue in MMT. To this end and we propose a transfer learning solution. Specifically and 1) A vanilla Transformer is pre-trained on massive bilingual text-only corpus to obtain prior knowledge; 2) A multimodal Transformer named VLTransformer is proposed with several components incorporated visual contexts; and 3) The parameters of VLTransformer are initialized with the pre-trained vanilla Transformer and then being fine-tuned on MMT tasks with a newly proposed method named cross-modal masking which forces the model to learn from both modalities. We evaluated on the Multi30k en-de and en-fr dataset and improving up to 8% BLEU score compared with the SOTA performance. The experimental result demonstrates that performing transfer learning with monomodal pre-trained NMT model on multimodal NMT tasks can obtain considerable boosts.
Minghan Wang, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
MTSummit (1)4