Shimin Tao

dblp:188/1236 · DBLP profile ↗
← Back
58ranked-venue papers
2as first author
53since 2021 · last 2026
0000-0002-2795-6921ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 since 2021Databases, data management, data science and information retrieval · 9 · 9 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLM
abstract
As large language models (LLMs) are increasingly deployed in high-stakes domains such as education, healthcare, and law, accurately evaluating their nuanced reasoning process becomes essential to ensure their safety, reliability, and trustworthiness. However, most existing benchmarks evaluate LLMs at a coarse granularity. Current benchmarks lack a unified framework and rely on single‐task datasets, overlooking the intermediate steps of complex reasoning. This results in redundant overlap across benchmarks, poor generalization to multifaceted real-world tasks, and underutilizes the rich reasoning traces generated by advanced LLMs.
Cui Danxin, Sihang Jiang 0001, Zhiyi Duan, Yanghua Xiao, Bi Yude, Jiaqing Liang, Minggui He, Shimin Tao, Yilun Liu 0001
AAAI9
2026 MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis
abstract
Despite doubts on data quality, instruction synthesis has been widely applied into instruction tuning (IT) of LLMs as an economic and rapid alternative. Recent endeavors focus on improving data quality for synthesized instruction pairs in English and have facilitated IT of English-centric LLMs. However, data quality issues in multilingual synthesized instruction pairs are even more severe, since the common synthesizing practice is to translate English synthesized data into other languages using machine translation (MT). Besides the known content errors in these English synthesized data, multilingual synthesized instruction data are further exposed to defects introduced by MT and face insufficient localization of the target languages, leading to cultural inequality in trained LLMs. In this paper, we propose MIDB, a Multilingual Instruction Data Booster to automatically address the quality issues in multilingual synthesized data. MIDB is trained on around 36.8k revision examples across 16 languages by human linguistic experts, thereby can boost the low-quality data by addressing content errors and MT defects, and improving localization in these synthesized data. Both automatic and human evaluation indicate that not only MIDB steadily improved instruction data quality in 16 languages, but also the instruction-following and cultural-understanding abilities of multilingual LLMs fine-tuned on MIDB-boosted data were significantly enhanced, suggesting an improved linguistic and cultural equality.
Yilun Liu 0001, Chunguang Zhao, Xinhua Yang, Hongyong Zeng, Shimin Tao, Weibin Meng, Minggui He, Hongxia Ma, Daimeng Wei, Boxing Chen
AAAI5
2026 ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph Reconstruction
abstract
Pairwise evaluation of large language models (LLMs) has become the dominant paradigm for benchmarking open-ended tasks, yet non-transitive preferences—where evaluators prefer A over B, B over C, but C over A—fundamentally undermine ranking reliability. We show that this critical issue stems largely from low-quality data that contains inherently ambiguous preference pairs. To address this challenge, we propose ELSPR, a principled graph-theoretic framework that models pairwise preferences as tournament graphs and systematically identifies problematic training data. ELSPR quantifies non-transitivity through strongly connected components (SCCs) analysis and measures overall preference clarity using a novel normalized directed graph structural entropy metric. Our filtering methodology selectively removes preference data that induce non-transitivity while preserving transitive preferences. Extensive experiments on the AlpacaEval benchmark demonstrate that models fine-tuned on ELSPR-filtered data achieve substantial improvements: a 13.8% reduction in non-transitivity, a 0.088 decrease in structural entropy, and significantly enhanced discriminative power in real-world evaluation systems. Human validation confirms that discarded data exhibit dramatically lower inter-annotator agreement (34.4% vs. 52.6%) and model-human consistency (51.2% vs. 80.6%) compared to cleaned data. These findings establish ELSPR as an effective data self-purification approach for developing more robust, consistent, and human-aligned LLM evaluation systems.
Yilun Liu 0001, Minggui He, Shimin Tao, Weibin Meng, Xinhua Yang, Hongxia Ma, Dengye Li, Daimeng Wei, Boxing Chen, Fuliang Li
AAAI4
2026 The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
abstract
Ying He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui HE, Shimin Tao, Mahongxia, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ying He 0010, Sihang Jiang 0001, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui He, Shimin Tao, Hongxia Ma, Yanghua Xiao
ACL (1)7
2026 DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning
abstract
Rongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang, Min Zhang, Shimin Tao, Daimeng Wei, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rongqing Jiang, Xuebo Liu 0002, Shengxin Liu, Min Zhang 0005, Shimin Tao, Daimeng Wei
ACL (1)6
2026 The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
abstract
Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao
ACL (1)5
2026 M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets
abstract
Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curated multilingual IFT datasets remain scarce. To address this gap, we propose M-DaQ (Multilingual Diversity and Quality), a diversity-aware sampling framework that jointly optimizes instruction-response quality and cross-lingual semantic diversity. M-DaQ leverages a fine-tuned Quality Scoring Model alongside a maximal marginal relevance-inspired selection strategy to construct balanced, high-fidelity training data. Furthermore, we present the first systematic investigation of the Superficial Alignment Hypothesis in multilingual settings. Extensive evaluations across 18 languages demonstrate that models trained on M-DaQ-curated data achieve average win rates exceeding 60% against strong baselines on Alpaca-Eval and MT-Bench. Complementary human evaluations corroborate these gains, highlighting significant improvements in cultural relevance, contextual appropriateness, and instruction-following capability. The code are publicly released to facilitate reproducibility and future research.
Chunguang Zhao, Yilun Liu 0001, Pufan Zeng, Yuanchang Luo, Shimin Tao, Minggui He, Weibin Meng, Hongxia Ma, Boxing Chen, Daimeng Wei
SIGIR5
2026 Chart specification: Structural representations for incentivizing VLM reasoning in chart-to-code generation
abstract
Vision-Language Models (VLMs) have shown promise in generating plotting code from chart images, yet achieving structural fidelity remains challenging. Existing approaches largely rely on supervised fine-tuning, encouraging surface-level token imitation rather than faithful modeling of chart structure, which often leads to hallucinated or semantically inconsistent outputs. We propose Chart Specification, a canonical structural representation that shifts training from mimicking training-code patterns to structure-grounded learning. By normalizing plotting code into structure-equivalent specifications, it enables (i) the construction of a structurally balanced training set, and (ii) a Spec-Align Reward that provides fine-grained, verifiable feedback on structural correctness for reinforcement learning. Under an explicit reasoning-to-code generation paradigm, this reward encourages structure-aware reasoning that produces constraint-consistent plotting code. Experiments on three public benchmarks show that our method consistently outperforms prior approaches. With only 3K training samples, we achieve strong data efficiency, surpassing leading baselines by up to 61.7% on complex benchmarks, and scaling to 4K samples establishes new state-of-the-art results across all evaluated metrics. Overall, our results demonstrate that precise structural supervision offers an efficient pathway to high-fidelity chart-to-code generation. Code and dataset are available at: https://github.com/Mighten/chart-specification-paper .
Minggui He, Mingchen Dai, Yilun Liu 0001, Shimin Tao, Pufan Zeng, Osamu Yoshie, Yuya Ieiri
Neurocomputing5
2026 Multiphase and Multitask Prompt Tuning for LLM-Based Context-Aware Machine Translation
abstract
Large language models (LLMs) are typically adapted for context-aware machine translation (MT) by combining both the source sentence and its surrounding sentences into a single input. This unified input is then processed in one go, with the model producing the target translation step by step. However, this method treats the intrasentence and intersentence contexts similarly, even though they play distinct roles. In this study, we present a novel strategy called multiphase prompt tuning (MPT) to address this issue by enabling LLMs to treat these two context types differently. MPT divides the context-aware translation task into three phases: encoding the intersentence context, encoding the source sentence, and the final decoding phase. Each phase incorporates distinct continuous prompts that help the model focus on the appropriate task for each type of context. We also introduce a multitask fine-tuning approach to emphasize the distinction between intersentence and intrasentence contexts and enhance intersentence dependencies. This includes two auxiliary tasks: context-agnostic translation and cross-lingual next sentence generation, which help extract additional information and improve the model's handling of discourse-related challenges.
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE Trans. Neural Networks Learn. Syst.5
2025 SRDC: Semantics-based Ransomware Detection and Classification with LLM-assisted Pre-training
abstract
In recent years, ransomware has emerged as a formidable data security threat, causing significant data privacy breaches that inflict substantial financial, reputational, and operational damages on society. Many studies employ dynamic feature analysis for ransomware detection. However, these methods utilize neither the internal semantic information (semantic information inherent in the features), nor external semantics (the wealth of existing knowledge and expert experience with regard to ransomware detection). Moreover, conventional methods rely on training data from known ransomware families, while zero-day ransomware often has unknown data distribution patterns, posing detection challenges. In this paper, we propose a Semantics-based Ransomware Detection and family Classification (SRDC) framework that can utilize both internal and external semantics of software. To bolster semantic analysis in zero-day attacks, we also design a procedure called LLM-assisted task-adaptive pre-training (LATAP). In LATAP, ransomware semantics from human experts and LLMs are employed to pre-train the detection model (GPT-2). By fully utilizing semantics, the proposed SRDC framework outperforms the SOTA methods by 12.15% for ransomware family classification tasks, and by 4.03% for zero-day ransomware detection tasks. SRDC also exhibits excellent data efficiency, requiring only two ransom families for training, which is only 35% of the data required by existing methods, to achieve a 90%+ accuracy of zero-day ransomware detection in nine unseen ransom families.
Ce Zhou, Yilun Liu 0001, Weibin Meng, Shimin Tao, Weinan Tian, Feiyu Yao, Boxing Chen, Hao Yang 0006
AAAI4
2025 Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
abstract
Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub.
Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ACL (1)6
2025 Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
abstract
Log analysis represents a critical sub-domain within AI applications that facilitates automatic approaches to fault and error management of large-scaled software systems, saving labors of traditional manual methods. While existing solutions using large language models (LLMs) show promise, they are limited by a significant domain gap between natural and log languages (the latter contains rich domain-specific tokens such as status codes, IP addresses, resource pathes), which restricts their effectiveness in real-world applications. However, directly adapting general-purpose LLMs to log analysis using raw logs may degrade their performance due to inconsistent token distribution. In this paper, we present a domain adaptation approach that addresses these limitations by integrating interpretable domain knowledge into open-source LLMs through continual pre-training (CPT), which bridges this domain gap by adapting LLMs on interpretable natural texts with log knowledge (instead of raw logs) to reduce distribution discrepancy. To achieve this, we developed NLPLog, a comprehensive dataset containing over 250,000 question-answer pairs on log-related knowledge. Our resulting model, SuperLog, achieves the best performance across four log analysis tasks, with an average accuracy improvement of 12.01% over the second-best model. Ablation study also suggests advantages of domain adaption using interpretable log knowledge over using raw logs.
Yuhe Ji, Yilun Liu 0001, Feiyu Yao, Minggui He, Shimin Tao, Chang Su 0001, Xinhua Yang, Weibin Meng, Yuming Xie, Boxing Chen, Shenglin Zhang, Yongqian Sun
CIKM5
2025 Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance
abstract
Yilun Liu, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Jian Gao, Zhang Li, Hao Yang, Boxing Chen, Osamu Yoshie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yilun Liu 0001, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Hao Yang 0006, Boxing Chen, Osamu Yoshie
EMNLP5
2025 SuperFC: Selective Data Utilization for a Sustainable and Effective Function-Calling Agent
abstract
The function-calling agent is obtained by performing agent tuning to the large language model (LLM) on function-calling dataset. However, even state-of-the-art datasets (e.g., xlam-function-calling-60k datasets) still contain numerous misleading examples of low-quality data, wasting significant computational resources and result in an unnecessary carbon footprint. Furthermore, such inductive bad data negatively impacts the performance of the agent. In this paper, we propose a set of scoring criteria specifically tailored to evaluate function-calling data and use these criteria to develop a data filtering framework. By applying this framework to filter out low-quality data, we fine-tuned SuperFC, which demonstrates substantial improvements in both sustainability and performance. The SuperFC-7B training process reduced training time from 455 minutes to 85 minutes, resulting in a 80.02% reduction in carbon footprint. Simultaneously, fine-tuning on high-quality data subsets led to performance improvements of up to 3.68%. Additionally, we provide an in-depth analysis of the causes behind the low quality of synthetic function-calling data, offering valuable insights for future data synthesis in this domain. We have also released a high-quality function-calling dataset, available at: https://github.com/Zire-Young/SuperFC
Xinhua Yang, Yilun Liu 0001, Shimin Tao, Chunguang Zhao, Weibin Meng, Minggui He, Chang Su 0001, Hongxia Ma, Jingzhou Du, Hao Yang 0006, Boxing Chen, Chuanwen Li
IJCNN3
2025 Rethinking Diffusion Bridge Model with Dual Alignments for Medical Image Synthesis
abstract
Medical image synthesis is crucial in clinical workflows, enabling the generation of missing modalities from available imaging data. While recent diffusion-based models show promise in medical image synthesis, they face two key limitations: progressive distribution drift from coarse intermediate samples and structural granularity loss due to missing high-frequency constraints. To address these challenges, we propose Dual Diffusion Bridge (DualDB), a framework integrating implicit distribution alignment and explicit structural constraints within a unified diffusion bridge paradigm. First, implicit distribution alignment employs optimal transport-guided adversarial learning to minimize statistical discrepancies between intermediate and target distributions, mitigating global distribution drift. Second, explicit structural alignment applies gradient-driven constraints to preserve high-frequency anatomical features, preventing structural degradation during reverse diffusion. This complementary design ensures both global statistical consistency and local anatomical precision in the synthesized results. Extensive experiments on multi-contrast MRI and MRI-CT translation show that DualDB outperforms state-of-the-art methods in quantitative performance and visual fidelity, maintaining superior anatomical accuracy even under noisy conditions.
Jinbao Wei, Shimin Tao, Aiping Liu, Xun Chen 0001
ACM Multimedia5
2025 Improving LLM-Based Document-Level MT with Multi-Knowledge Fusion
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
NLPCC (3)6
2025 LogEval: A comprehensive benchmark suite for LLMs in log analysis
Tianyu Cui, Shiyu Ma, Tong Xiao 0002, Shimin Tao, Yilun Liu 0001, Shenglin Zhang, Duoming Lin, Changchang Liu, Yuzhe Cai, Weibin Meng, Yongqian Sun, Dan Pei
Empir. Softw. Eng.6
2025 Degradation-Aware Prompted Transformer for Unified Medical Image Restoration
abstract
Medical image restoration (MedIR) aims to recover high-quality images from degraded inputs, yet faces unique challenges from physics-driven degradations and multi-modal task interference. While existing all-in-one methods handle natural image degradations well, they struggle with medical scenarios due to limited degradation perception and suboptimal multi-task optimization. In response, we introduce DaPT, a Degradation-aware Prompted Transformer, which integrates dynamic prompt learning and modular expert mining for unified MedIR. First, DaPT introduces spatially compact prompts with optimal transport regularization, amplifying inter-prompt differences to capture diverse degradation patterns. Second, a mixture of experts dynamically routes inputs to specialized modules via prompt guidance, resolving task conflicts while reducing computational overhead. The synergy of prompt learning and expert mining further enables robust restoration across multi-modal medical data, offering a practical solution for clinical imaging. Extensive experiments across multiple modalities (MRI, CT, PET) and diverse degradations, covering both in-distribution and out-of-distribution scenarios, demonstrate that DaPT consistently outperforms state-of-the-art methods and generalizes reliably to unseen settings, underscoring its robustness, effectiveness, and clinical practicality. The source code will be released at https://github.com/weijinbao1998/DaPT.
Jinbao Wei, Shimin Tao, Aiping Liu, Xun Chen 0001
IEEE Trans. Image Process.4
2024 Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models
abstract
To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal translations often miss the intended meaning. Traditional methods, which replace idioms using existing knowledge bases (KBs), often lack scale and context-awareness. Addressing these challenges, our approach prioritizes context-awareness and scalability, allowing for offline storage of idioms in a manageable KB size. This ensures efficient serving with smaller models and provides a more comprehensive understanding of idiomatic expressions. We introduce a multilingual idiom KB (IdiomKB) developed using large LMs to address this. This KB facilitates better translation by smaller models, such as BLOOMZ (7.1B), Alpaca (7B), and InstructGPT (6.7B), by retrieving idioms' figurative meanings. We present a novel, GPT-4-powered metric for human-aligned evaluation, demonstrating that IdiomKB considerably boosts model performance. Human evaluations further validate our KB's quality.
Jiangjie Chen, Hao Yang 0006, Shimin Tao, Yanghua Xiao
AAAI6
2024 Evaluation Dataset for Lexical Translation Consistency in Chinese-to-English Document-level Translation
abstract
Lexical translation consistency is one of the most common discourse phenomena in Chinese-to-English document-level translation. To better evaluate the performance of lexical translation consistency, previous researches assumes that all repeated source words should be translated consistently. However, constraining translations of repeated source words to be consistent will hurt word diversity and human translators tend to use different words in translation. Therefore, in this paper we construct a test set of 310 bilingual news articles to properly evaluate lexical translation consistency. We manually differentiate those repeated source words whose translations are consistent into two types: true consistency and false consistency. Then based on the constructed test set, we evaluate the performance of lexical translation consistency for several typical NMT systems.
Xiangyu Lei, Junhui Li 0001, Shimin Tao, Hao Yang 0006
LREC/COLING3
2024 Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation
abstract
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Mahong Xia, Zhang Li, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yuan Ge 0001, Yilun Liu 0001, Chi Hu, Weibin Meng, Shimin Tao, Mahong Xia, Boxing Chen, Hao Yang 0006, Tong Xiao 0001
EMNLP5
2024 DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators
abstract
Generally, the decoder-only large language models (LLMs) are adapted to context-aware neural machine translation (NMT) in a concatenating way, where LLMs take the concatenation of the source sentence (i.e., intrasentence context) and the inter-sentence context as the input, and then to generate the target tokens sequentially.This adaptation strategy, i.e., concatenation mode, considers intrasentence and inter-sentence contexts with the same priority, despite an apparent difference between the two kinds of contexts.In this paper, we propose an alternative adaptation approach, named Decoding-enhanced Multiphase Prompt Tuning (DeMPT), to make LLMs discriminately model and utilize the inter-and intra-sentence context and more effectively adapt LLMs to context-aware NMT.First, DeMPT divides the context-aware NMT process into three separate phases.During each phase, different continuous prompts are introduced to make LLMs discriminately model various information.Second, DeMPT employs a heuristic way to further discriminately enhance the utilization of the source-side interand intra-sentence information at the final decoding phase.Experiments show that our approach significantly outperforms the concatenation method, and further improves the performance of LLMs in discourse modeling.
Xinglin Lyu, Junhui Li 0001, Min Zhang 0042, Daimeng Wei, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP6
2024 CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
abstract
Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set11https://github.com/lunyiliu/CoachLM.
Yilun Liu 0001, Shimin Tao, Ming Zhu 0010, Wenbing Ma, Chang Su 0001, Yutai Hou, Min Zhang 0042, Hongxia Ma, Hao Yang 0006, Yanfei Jiang
ICDE2
2024 From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
abstract
Machine Translation Quality Estimation (MTQE) is the task of estimating the quality of machine-translated text in real time without the need for reference translations, which is of great importance for the development of MT. After two decades of evolution, QE has yielded a wealth of results. This article provides a comprehensive overview of QE datasets, annotation methods, shared tasks, methodologies, challenges, and future research directions. It begins with an introduction to the background and significance of QE, followed by an explanation of the concepts and evaluation metrics for word-level QE, sentence-level QE, document-level QE, and explainable QE. The paper categorizes the methods developed throughout the history of QE into those based on handcrafted features, deep learning, and Large Language Models (LLMs), with a further division of deep learning-based methods into classic deep learning and those incorporating pre-trained language models (LMs). Additionally, the article details the advantages and limitations of each method and offers a straightforward comparison of different approaches. Finally, the paper discusses the current challenges in QE research and provides an outlook on future research directions.
Haofei Zhao, Yilun Liu 0001, Shimin Tao, Weibin Meng, Xiang Geng, Chang Su 0001, Min Zhang 0042, Hao Yang 0006
IJCNN3
2024 Using Large Language Model for End-to-End Chinese ASR and NER
Yuang Li, Min Zhang 0042, Mengxin Ren, Shimin Tao, Jinsong Su, Hao Yang 0006
INTERSPEECH7
2024 A Multitask Training Approach to Enhance Whisper with Open-Vocabulary Keyword Spotting
abstract
The recognition of rare named entities, such as personal names and terminologies, is challenging for automatic speech recognition (ASR) systems, especially when they are not frequently observed in the training data.In this paper, we introduce keyword spotting enhanced Whisper (KWS-Whisper), a novel ASR system that leverages the Whisper model and performs openvocabulary keyword spotting (OV-KWS) on the hidden states of the Whisper encoder to recognize user-defined named entities.These entities serve as prompts for the Whisper decoder.To optimize the model, we propose a multitask training approach that learns OV-KWS and contextual-ASR tasks.We evaluate our approach on Chinese Aishell hot word subsets and two internal code-switching test sets and show that it significantly improves the entity recall compared to the original Whisper model.Moreover, we demonstrate that the OV-KWS can be a plug-andplay module to enhance the ASR error correction methods and frozen Whisper models.
Yuang Li, Min Zhang 0042, Chang Su 0001, Yinglu Li, Xiaosong Qiao, Mengxin Ren, Miaomiao Ma, Daimeng Wei, Shimin Tao, Hao Yang 0006
INTERSPEECH9
2024 Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies
abstract
Automated log analysis is crucial in modern software-intensive systems for facilitating program comprehension throughout software maintenance and engineering life cycles. Existing methods perform tasks such as log parsing and log anomaly detection by providing a single prediction value without interpretation. However, given the increasing volume of system events, the limited interpretability of analysis results hinders analysts' comprehension of program status and their ability to take appropriate actions. Moreover, these methods require substantial in-domain training data, and their performance declines sharply (by up to 62.5%) in online scenarios involving unseen logs from new domains, a common occurrence due to rapid software updates. In this paper, we propose LogPrompt, a novel interpretable log analysis approach for online scenarios. LogPrompt employs large language models (LLMs) to perform online log analysis tasks via a suite of advanced prompt strategies tailored for log tasks, which enhances LLMs' performance by up to 380.7% compared with simple prompts. Experiments on nine publicly available evaluation datasets across two tasks demonstrate that LogPrompt, despite requiring no in-domain training, outperforms existing approaches trained on thousands of logs by up to 55.9%. We also conduct a human evaluation of LogPrompt's interpretability, with six practitioners possessing over 10 years of experience, who highly rated the generated content in terms of usefulness and readability (averagely 4.42/5). LogPrompt also exhibits remarkable compatibility with open-source and smaller-scale LLMs, making it flexible for practical deployment. Code of LogPrompt is available at https://github.com/lunyiliu/LogPrompt.
Yilun Liu 0001, Shimin Tao, Weibin Meng, Jingyu Wang 0001, Wenbing Ma, Hao Yang 0006, Yanfei Jiang
ICPC2
2024 Multi-Source Log Parsing With Pre-Trained Domain Classifier
abstract
Automated log analysis with AI technologies is commonly used in network, system, and service operation and maintenance to ensure reliability and quality assurance. Log parsing serves as an essential primary stage in log analysis, where unstructured logs are transformed into structured data to facilitate subsequent downstream analysis. However, traditional log parsing algorithms designed for single-domain processing struggle to handle the challenges posed by multi-source log inputs, leading to a decline in parsing accuracy. Adapting these algorithms to multi-source logs often requires extensive manual labeling efforts. To address this, we propose Domain-aware Parser (DA-Parser), a framework that includes a domain classifier to identify the source domains of multi-source logs. This enables the conversion of the multi-source log parsing problem into a series of single-source parsing problems. The classifier is pre-trained on a corpus of logs from 16 domains, eliminating the need for additional human labeling. The predicted source domain tags serve as constraints, limiting the template extraction process to logs from the same domain. Empirical evaluation on a multi-domain dataset demonstrates that DA-Parser outperforms the existing SOTA algorithm by 21.6% in terms of parsing accuracy. The proposed approach also shows potential efficiency improvements, requiring only 6.67% of the time consumed by existing parsers, while maintaining robustness against minor domain classification errors.
Yilun Liu 0001, Shimin Tao, Weibin Meng, Jingyu Wang 0001, Hao Yang 0006, Yanfei Jiang
IEEE Trans. Netw. Serv. Manag.2
2023 Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning
abstract
Quality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains QE models with pseudo QE data and obtains remarkable performance. However, there tends to be inevitable noise in the pseudo data, hindering models from learning QE accurately. Our study shows that the noise mainly comes from the differences between pseudo and real translation outputs. To handle this problem, we propose CLQE, a denoising pre-training framework for QE based on curriculum learning. More specifically, we propose to measure the degree of noise in the pseudo QE data with some metrics based on statistical or distributional features. With the guidance of these metrics, CLQE gradually pre-trains the QE model using data from cleaner to noisier. Experiments on various benchmarks reveal that CLQE outperforms DirectQE and other strong baselines. We also show that with our framework, pre-training converges faster than directly using the pseudo data. We make our CLQE code available (https://github.com/NJUNLP/njuqe).
Xiang Geng, Jiahuan Li, Shujian Huang, Hao Yang 0006, Shimin Tao, Jiajun Chen 0001
AAAI6
2023 Knowledge Prompt for Whisper: An ASR Entity Correction Approach with Knowledge Base
abstract
Entity correction is crucial in Automatic Speech TABLE I Recognition (ASR), since erroneous entities seriously affect our understanding of ASR results. In this paper, in order to correct entity errors, we propose a knowledge prompt approach for Whisper (a recent ASR model trained with a corpus containing 680k hours of labeled speech recorded in various conditions). For a given audio, our approach consists of three steps: (1) obtaining its ASR result by Whisper; (2) fuzzy matching the ASR result with a knowledge base to obtain candidate entities; (3) using the candidate entities as a prompt to obtain the final ASR result by Whisper again. We conduct experiments on the test dataset of open-source Chinese speech corpus AISHELLNER. Experimental results show that our approach not only significantly improves the entity recall rate in ASR results (from 70.97% to 84.82%), but also reduces the overall Character Error Rate (CER).
Min Zhang 0042, Xiaosong Qiao, Chang Su 0001, Yinglu Li, Yuang Li, Ming Zhu 0010, Mengyao Piao, Shimin Tao, Hao Yang 0006, Yanfei Jiang
IEEE Big Data10
2023 DA-Parser: A Pre-trained Domain-aware Parsing Framework for Heterogeneous Log Analysis
abstract
Automated log analysis is widely applied in modern software-intensive systems to ensure resilience and sustainability, where log parsing is a vital initial step, converting unstructured logs into structured data for downstream analysis. However, traditional log parsing algorithms are designed to process logs within a single domain. As cross-domain dependencies and interactions between sub-modules of software systems increase, these algorithms struggle to handle the challenges posed by multi-domain log inputs, which results in a significant decline in parsing accuracy when facing heterogeneous logs. Additionally, current solutions for heterogeneous log parsing require extensive manual labeling efforts. In this paper, we propose Domain-aware Parser (DA-Parser), a framework that consists of a domain-aware head to identify the source domains of heterogeneous logs and then converts the multi-domain log parsing problem into a series of single-domain parsing problems. The domain-aware head is pretrained using a corpus of logs from 16 domains, which allows for the classification of the source domains of most heterogeneous log set without additional human labeling. Source domain tags predicted by the domain-aware head serve as a constraint to limit the template extraction process to logs from the same domain. Empirical evaluation is conducted on a multi-domain dataset containing logs from 7 domains. DA-Parser can be integrated with existing single-domain algorithms and are compatible with them, achieving superior parsing accuracy with an average of 9.26% improvement compared with single-domain algorithms.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Jingyu Wang 0001, Chang Su 0001, Weinan Tian, Min Zhang 0042, Hao Yang 0006, Xun Chen 0001
COMPSAC1
2023 Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam Search
abstract
Xiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She, Wei Zou, Shimin Tao, Hao Yang, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Xiang Geng, Zhejian Lai, Shuaijie She, Shimin Tao, Hao Yang 0006, Jiajun Chen 0001, Shujian Huang
EMNLP6
2023 UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
abstract
Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets.
Minghan Wang, Xiaosong Qiao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP11
2023 Zephyr: Zero-Shot Punctuation Restoration
abstract
Punctuation restoration can be crucial for the cascade speech translation system. Traditional approaches typically treat it as a sequential tagging problem, predicting which punctuation mark should follow a given word. However, this often requires significant computational and storage resources for full-stage training or fine-tuning. Our argument is that pre-trained language models (PLMs) can directly leverage their learned knowledge for punctuation generation, making additional training unnecessary. In this paper, we propose the Zephyr algorithm, which utilizes PLMs to perform zero-shot and few-shot punctuation restoration for both offline and streaming scenarios. Our experimental results demonstrate that, in comparison to fine-tuning-based baselines, Zephyr achieves competitive performance while requiring little to no training cost and exhibiting better generalizability in zeroshot and few-shot settings.
Minghan Wang, Yinglu Li, Xiaosong Qiao, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP7
2023 CONFPILOT: A Pilot for Faster Configuration by Learning from Device Manuals
abstract
The command line interface (CLI) is widely used to configure and manage network devices. However, as heterogeneous devices are introduced into the network, the CLI-based method is becoming time-consuming and inefficient because much effort is required to learn proprietary configuration languages of different vendors or consult online documents. In this work, we present CONFPILOT, an assistant system that can accelerate configuration by automatically converting natural language intents into commands. Our solution is based on a retrieval-augmented generation framework that features a unified parser that parses device manuals into a searchable configuration library, a vendor-agnostic retriever that finds the most relevant$k$syntaxes through a two-stage coarse-to-fine process, and a reliable generator that predicts syntactically correct commands via a pointer-generator network and syntax-guided decoding. In a nutshell, CONFPILOT frees engineers from most time-consuming efforts by learning directly from device manuals to generate configuration commands. Our evaluation and user study show, CONFPILOT can speed up the configuration process by 60x compared to manual lookup while maintaining acceptable exact match accuracy. Furthermore, CONFPILOT can quickly adapt to new vendors and devices with little human effort and time cost.
Jinyu Zhao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Shimin Tao, Jianxin Liao
ICDCS6
2023 WhiSLU: End-to-End Spoken Language Understanding with Whisper
Minghan Wang, Yinglu Li, Xiaosong Qiao, Hengchao Shang, Daimeng Wei, Shimin Tao, Min Zhang 0042, Hao Yang 0006
INTERSPEECH8
2023 Biglog: Unsupervised Large-scale Pre-training for a Unified Log Representation
abstract
Automated log analysis has been widely applied in modern data-center network, performing critical tasks such as log parsing, log anomaly detection and log-based failure prediction. However, existing approaches rely on hand-crafted features or domain-specific vectors to represent logs, which are either laborious in manual efforts or ineffective facing multiple domains in a system. Furthermore, general-purpose word embeddings are not optimized for log data, thus are data-inefficient in handling complex log analysis tasks. In this paper, we present a pre-training phase for language models to understand both in-sentence and cross-sentence features of logs, resulting in a unified representation of logs that is well-suited for various downstream analysis tasks. The pre-training phase is unsupervised, utilizing 0.45 billion logs from 16 diverse domains. Experiments on 12 publicly available evaluation datasets across 3 tasks indicate superiority of our approach against existing approaches, especially in online scenarios with limited historical logs. Our approach also exhibits remarkable few-shot learning ability and domain-adaptiveness, which not only outperforms existing approaches using only 0.0025% of their required training data, but also adapts into new domains via only a few in-domain logs. We release our code and pre-trained model.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Zuomin Ren, Hao Yang 0006, Xun Chen 0001, Yuming Xie, Chang Su 0001, Xiaosong Oiao, Weinan Tian, Yichen Zhu 0001
IWQoS1
2023 Multi-order Matched Neighborhood Consistent Graph Alignment in a Union Vector Space
abstract
In this paper, we study the unsupervised plain graph alignment problem, which aims to find node correspondences across two graphs without any side information. The majority of previous works addressed UPGA based on structural information, which will inevitably lead to subgraph isomorphism issues. That is, unaligned nodes could take similar local structural information. To mitigate this issue, we present the Multi-order Matched Neighborhood Consistent (MMNC) which tries to match nodes by aligning the learned node embeddings with only a small number of pseudo alignment seeds. In particular, we extend matched neighborhood consistency (MNC) to vector space and further develop embedding-based MNC (EMNC). By minimizing the EMNC-based loss function, we can utilize the limited pseudo alignment seeds to approximate the orthogonal transformation matrix between two groups of node embeddings with high efficiency and accuracy. Through extensive experiments on public benchmarks, we show that the proposed methods achieve a good balance between alignment accuracy and speed over multiple datasets compared with existing methods.
Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039, Hao Yang 0006, Shimin Tao
SIGIR7
2023 Weakly Supervised Entity Alignment with Positional Inspiration
abstract
The current success of entity alignment (EA) is still mainly based on large-scale labeled anchor links. However, the refined annotation of anchor links still consumes a lot of manpower and material resources. As a result, an increasing number of works based on active learning, few-shot learning, or other deep network learning techniques have been developed to address the performance bottleneck caused by a lack of labeled data. These works focus either on the strategy of choosing more informative labeled data or on the strategy of model training, while it remains opaque why existing popular EA models (e.g., GNN-based models) fail the EA task with limited labeled data. To overcome this issue, this paper analyzes the problem of weakly supervised EA from the perspective of model design and proposes a novel weakly supervised learning framework, Position Enhanced Entity Alignment (PEEA). Besides absorbing structural and relational information, PEEA aims to increase the connections between far-away entities and labeled ones by incorporating positional information into the representation learning with a Position Attention Layer (PAL). To fully utilize the limited anchor links, we further introduce a novel position encoding method that considers both anchor links and relational information from a global view. The proposed position encoding will be fed into PEEA as additional entity features. Extensive experiments on public datasets demonstrate the effectiveness of PEEA.
Wei Tang 0013, Fenglong Su, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Shimin Tao, Hao Yang 0006
WSDM6
2023 Collective Human Opinions in Semantic Textual Similarity
abstract
Abstract Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as gold standard. Averaging masks the true distribution of human opinions on examples of low agreement, and prevents models from capturing the semantic vagueness that the individual ratings represent. In this work, we introduce USTS, the first Uncertainty-aware STS dataset with ∼15,000 Chinese sentence pairs and 150,000 labels, to study collective human opinions in STS. Analysis reveals that neither a scalar nor a single Gaussian fits a set of observed judgments adequately. We further show that current STS models cannot capture the variance caused by human disagreement on individual instances, but rather reflect the predictive confidence over the aggregate dataset.
Yuxia Wang 0003, Shimin Tao, Hao Yang 0006, Timothy Baldwin, Karin Verspoor
Trans. Assoc. Comput. Linguistics2
2023 P-Transformer: Towards Better Document-to-Document Neural Machine Translation
abstract
Directly training a document-to-document (Doc2Doc) neural machine translation (NMT) via Transformer from scratch, especially on small datasets, usually fails to converge. Our dedicated probing tasks show that 1) both the absolute position and relative position information gets gradually weakened or even vanished once it reaches the upper encoder layers, and 2) the vanishing of absolute position information in encoder output causes the training failure of Doc2Doc NMT. To alleviate this problem, we propose a position-aware Transformer (P-Transformer) to enhance both the absolute and relative position information in both self-attention and cross-attention. Specifically, we integrate absolute positional information, i.e., position embeddings, into the query-key pairs both in self-attention and cross-attention through a simple yet effective addition operation. Moreover, we also integrate relative position encoding in self-attention. The proposed P-Transformer utilizes sinusoidal position encoding and does not require any task-specified position embedding, segment embedding, or attention mechanism. Through the above methods, we build a Doc2Doc NMT model with P-Transformer, which ingests the source document and completely generates the target document in a sequence-to-sequence (seq2seq) way. In addition, P-Transformer can be applied to seq2seq-based document-to-sentence (Doc2Sent) and sentence-to-sentence (Sent2Sent) translations. Extensive experimental results of Doc2Doc NMT show that P-Transformer significantly outperforms strong baselines on the widely-used 9 document-level datasets in 7 language pairs, covering small-, middle-, and large-scales, and achieves a new state-of-the-art. Experimentation on discourse phenomena shows that our Doc2Doc NMT models improve the translation quality in both BLEU and discourse coherence. We make our code available on Github.
Yachao Li 0003, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Exploiting Spatial-Temporal Behavior Patterns for Fraud Detection in Telecom Networks
abstract
Fraud detection in telecom network is a crucial problem that threatens users’ privacy and property security. In recent years, fraudsters adopt more advanced camouflage strategies to avoid being detected by traditional algorithms. To deal with these new types of fraud, it is necessary to analyze the integrated spatial-temporal features, which are rarely involved in existing literature. In this article, we propose a novel fraud detection model based on the intertwined spatial-temporal patterns of user behaviors. Specifically, we first introduce the extension of statistical and interactive features to dynamic call patterns, and build a probabilistic model to simulate users’ call behaviors. Then the sequential patterns reflecting users’ own behaviors are obtained by the mixture Hidden Markov Models, and the structural patterns reflecting the collaboration between users in the telecom network are obtained by the attention-based Graph-SAGE model. Finally, our model outputs a fraud score for each user to detect potential fraudsters. We conduct extensive experiments on a real-world telecom dataset. The experimental results demonstrate that our intertwined spatial-temporal call patterns can effectively represent user behavior and improve the accuracy of fraud detection compared with state-of-the-art methods. The results also validate the efficiency and the interpretability of our model.
Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006, Jianxin Liao, Zhu Han 0001
IEEE Trans. Dependable Secur. Comput.5
2023 LogSummary: Unstructured Log Summarization for Software Systems
abstract
We propose LogSummary, an automatic, unsupervised end-to-end log summarization framework for software system maintenance in this work. LogSummary obtains the summarized triples of necessary logs for a given log sequence. It integrates a novel information extraction method that considers semantic information and domain knowledge with a new triple-ranking approach using the global knowledge learned from all logs. Given the lack of a publicly-available gold standard for log summarization, we have manually labeled the summaries of four open-source log datasets and made them publicly available. The evaluation of these datasets and the case studies on real-world logs demonstrate that LogSummary produces highly representative (average ROUGE F1 score of 0.741) summaries efficiently. We have packaged LogSummary into an open-source toolkit and hope it can be a standard baseline and benefit future log summarization works.
Weibin Meng, Federico Zaiter, Ying Liu 0024, Shenglin Zhang, Shimin Tao, Yichen Zhu 0001, En Wang, Dan Pei
IEEE Trans. Netw. Serv. Manag.6
2022 EntityRank: Unsupervised Mining of Bilingual Named Entity Pairs from Parallel Corpora for Neural Machine Translation
abstract
As Neural Machine Translation (NMT) heavily relies on training data, finding an effective method to help NMT make better use of limited data is of great significance. In this paper, with the motivation of the famous Google’s PageRank algorithm, we propose a novel unsupervised method EntityRank for mining bilingual named entity pairs from parallel corpora, which involves three critical components (Generator, Scorer and Filter). To apply the pairs mined by EntityRank to NMT, we design a data augmentation strategy for the state-of-the-art (SOTA) model Transformer. From the experimental results on the CCMT20 English-Chinese and WMT14 English-German news parallel corpora, it can be seen that the unsupervised method EntityRank could obtain relatively high quality bilingual named entity pairs; and with the designed data augmentation strategy, the mined pairs could not only significantly improve the translation quality of their covered data, but also benefit the translation quality of the overall data.
Min Zhang 0042, Hao Yang 0006, Xiaosong Qiao, Shimin Tao, Yanfei Jiang
IEEE Big Data7
2022 Diformer: Directional Transformer for Neural Machine Translation
abstract
Autoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current combination frameworks focus more on the integration of multiple decoding paradigms with a unified generative model, e.g. Masked Language Model. However, the generalization can be harmful on the performance due to the gap between training objective and inference. In this paper, we aim to close the gap by preserving the original objective of AR and NAR under a unified framework. Specifically, we propose the Directional Transformer (Diformer) by jointly modelling AR and NAR into three generation directions (left-to-right, right-to-left and straight) with a newly introduced direction variable, which works by controlling the prediction of each token to have specific dependencies under that direction. The unification achieved by direction successfully preserves the original dependency assumption used in AR and NAR, retaining both generalization and performance. Experiments on 4 WMT benchmarks demonstrate that Diformer outperforms current united-modelling works with more than 1.5 BLEU points for both AR and NAR decoding, and is also competitive to the state-of-the-art independent AR and NAR models.
Minghan Wang, Yuxia Wang 0003, Daimeng Wei, Hengchao Shang, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
EAMT10
2022 Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine Translation
abstract
In this paper we aim to relieve the issue of lexical translation inconsistency for documentlevel neural machine translation (NMT) by modeling consistency preference for lexical chains which consist of repeated words in a source-side document and provide a representation of the lexical consistency structure of the document.Specifically, we first propose lexical-consistency attention to capture consistency context among words in the same lexical chains.Then for each lexical chain we define and learn a consistency-tailored latent variable, which will guide the translation of corresponding sentences to enhance lexical translation consistency.Experimental results on Chinese→English and French→English document-level translation tasks show that our approach not only significantly improves translation performance in BLEU, but also substantially alleviates the problem of the lexical translation inconsistency.
Xinglin Lyu, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP3
2022 CCDC: A Chinese-Centric Cross Domain Contrastive Learning Framework
Hao Yang 0006, Shimin Tao, Minghan Wang, Min Zhang 0042, Daimeng Wei, Shuai Zhao 0001, Miaomiao Ma
KSEM (2)2
2022 Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical Constraints
abstract
Chun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu, Hao Yang, Qin Ying, Shimin Tao, Yanghua Xiao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu 0026, Hao Yang 0006, Shimin Tao, Yanghua Xiao
NAACL-HLT7
2022 WRS: Workflow Retrieval System for Cloud Automatic Remediation
abstract
Remediation systems in modern information infrastructure gradually take over the increasing workload of unexpected events from operators. The core logic in such systems — remediation rule— still needs operators to fill in manually, which is a tedious and error-prone process. We propose Workflow Retrieval System (a.k.a. WRS), which helps the operator to build remediation rules.WRS is a recommendation system that recommends structures from existing rules to operators when they are creating new ones. WRS formalizes the workflows in remediation rules as trees and extracts representative atomic structures from them. Then WRS organizes atomic structures in two ways to accelerate the later retrieval — a two-level hierarchy which helps searching similar structures, and a keyword indexing structure which helps searching by words. With the two retrieval structures, we build two applications: one is real-time workflow auto completion application which recommends remaining structures based on the existing partial structure, and another is workflow recommendation where WRS recommends atomic structures based on the description of the remediation symptoms.We prototype WRS and evaluate it with legacy device vendors’ operation manual. Our evaluation shows that WRS reasonably extracts and organizes atomic structures in remediation workflows, and the two applications atop it show fast execution time and high accuracy.
Hongyi Huang, Wenfei Wu, Shimin Tao
NOMS3
2021 Prefix-Graph: A Versatile Log Parsing Approach Merging Prefix Tree with Probabilistic Graph
abstract
Logs play an important part in analyzing system behavior and diagnosing system failures. As the basic step of log analysis, log parsing converts raw log messages into structured log templates. However, existing log parsing approaches are not adaptive and versatile enough to ensure their high accuracy on all types of datasets. In particular, it is required to design regular expressions or fine-tune the hyper-parameters manually for the best performance. In this paper, we propose Prefix-Graph, an online versatile log parsing approach. Prefix-Graph is a probabilistic graph structure extended from prefix tree. It iteratively merges together two branches which have high similarity in probability distribution, and represents log templates as the combination of cut-edges in root-to-leaf paths of the graph. Since no domain knowledge is used and all the parameters are fixed, Prefix-Graph can be easily applied to different log datasets without any additional manual work. We evaluate our approach on 10 real-world datasets and 117GB log messages obtained from Huawei. The experimental results demonstrate that Prefix-Graph achieves the highest average accuracy of 0.975 and the smallest standard deviation of 0.037. Our approach is superior to baseline methods in terms of adaptability and versatility.
Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Jianxin Liao
ICDE5
2021 HI-CMLM: Improve CMLM with Hybrid Decoder Input
abstract
Mask-predict CMLM (Ghazvininejad et al., 2019) has achieved stunning performance among non-autoregressive NMT models, but we find that the mechanism of predicting all of the target words only depending on the hidden state of [MASK] is not effective and efficient in initial iterations of refinement, resulting in ungrammatical repetitions and slow convergence.In this work, we mitigate this problem by combining copied source with embeddings of [MASK] in decoder.Notably.it's not a straightforward copying that is shown to be useless, but a novel heuristic hybrid strategy -fence-mask.Experimental results show that it gains consistent boosts on both WMT14 En↔De and WMT16 En↔Ro corpus by 0.5 BLEU on average, and 1 BLEU for lessinformative short sentences.This reveals that incorporating additional information by proper strategies is beneficial to improve CMLM, particularly translation quality of short texts and speeding up early-stage convergence.
Minghan Wang, Yuxia Wang 0003, Chang Su 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
INLG8
2021 Make the Blind Translator See The World: A Novel Transfer Learning Solution for Multimodal Machine Translation
abstract
Based on large-scale pretrained networks and the liability to be easily overfitting with limited labelled training data of multimodal translation (MMT) is a critical issue in MMT. To this end and we propose a transfer learning solution. Specifically and 1) A vanilla Transformer is pre-trained on massive bilingual text-only corpus to obtain prior knowledge; 2) A multimodal Transformer named VLTransformer is proposed with several components incorporated visual contexts; and 3) The parameters of VLTransformer are initialized with the pre-trained vanilla Transformer and then being fine-tuned on MMT tasks with a newly proposed method named cross-modal masking which forces the model to learn from both modalities. We evaluated on the Multi30k en-de and en-fr dataset and improving up to 8% BLEU score compared with the SOTA performance. The experimental result demonstrates that performing transfer learning with monomodal pre-trained NMT model on multimodal NMT tasks can obtain considerable boosts.
Minghan Wang, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
MTSummit (1)6
2021 Deep graph alignment network
Wei Tang 0013, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006
Neurocomputing5
2020 LogParse: Making Log Parsing Adaptive through Word Classification
abstract
Logs are one of the most valuable data sources for large-scale service (e.g., social network, search engine) maintenance. Log parsing serves as the the first step towards automated log analysis. However, the current log parsing methods are not adaptive. Without intra-service adaptiveness, log parsing cannot handle software/firmware upgrade because learned templates cannot match new type of logs. In addition, without cross-service adaptiveness, the logs of a new type of service cannot be accurately parsed when this service is newly deployed. We propose LogParse, an adaptive log parsing framework, to support intra-service and cross-service incremental template learning and update. LogParse turns the template generation problem into a word classification problem and learns the features of template words and variable words. We evaluate LogParse on four public production log datasets. The results demonstrate that LogParse supports accurate adaptive template update (increased from 0.559 to nearly 1.0 parsing accuracy), and a trained LogParse is adaptive for a brand new service’s log parsing. Because of LogParse’s adaptiveness, we also apply LogParse to an interesting application, log compression and deployed log compression in a top cloud service provider. We package LogParse into an open-source toolkit.
Weibin Meng, Ying Liu 0024, Federico Zaiter, Shenglin Zhang, Yichen Zhu 0001, En Wang, Shimin Tao, Dian Yang, Dan Pei
ICCCN10
2019 LogAnomaly: Unsupervised Detection of Sequential and Quantitative Anomalies in Unstructured Logs
abstract
Recording runtime status via logs is common for almost every computer system, and detecting anomalies in logs is crucial for timely identifying malfunctions of systems. However, manually detecting anomalies for logs is time-consuming, error-prone, and infeasible. Existing automatic log anomaly detection approaches, using indexes rather than semantics of log templates, tend to cause false alarms. In this work, we propose LogAnomaly, a framework to model unstructured a log stream as a natural language sequence. Empowered by template2vec, a novel, simple yet effective method to extract the semantic information hidden in log templates, LogAnomaly can detect both sequential and quantitive log anomalies simultaneously, which were not done by any previous work. Moreover, LogAnomaly can avoid the false alarms caused by the newly appearing log templates between periodic model retrainings. Our evaluation on two public production log datasets show that LogAnomaly outperforms existing log-based anomaly detection methods.
Weibin Meng, Ying Liu 0024, Yichen Zhu 0001, Shenglin Zhang, Dan Pei, Shimin Tao
IJCAI9
2018 FUNNEL: Assessing Software Changes in Web-Based Services
abstract
The detection of performance changes in software change roll-outs in Internet-based services is crucial for an operations team, because it allows timely roll-back of a software change when performance degrades unexpectedly. However, it is infeasible to manually investigate millions of performance measurements of many roll-outs. In this paper, we present an automated tool, FUNNEL, for rapid and robust impact assessment of software changes in large Internet-based services. FUNNEL automatically collects the related performance measurements for each software change. To detect significant performance behavior changes, FUNNEL adopts singular spectrum transform (SST) algorithm as the core algorithm, uses various techniques to improve its robustness and reduce its computational cost, and applies a difference-in-difference (DiD) method to differentiate the true causality from the random correlations between the performance change and the software change. Evaluation through historical data in real-word services shows that FUNNEL achieves accuracy of more than 99.7 percent. Compared with previous methods, FUNNEL's detection delay is 38.02 to 64.99 percent shorter, and its computation speed is 4.59-7,098 times faster. In real deployment, FUNNEL achieves a 98.21 percent precision, high robustness, fast detection speed, and shows its capability in detecting unexpected behavior changes.
Shenglin Zhang, Ying Liu 0024, Dan Pei, Xianping Qu, Shimin Tao, Zhi Zang, Xiaowei Jing, Mei Feng
IEEE Trans. Serv. Comput.6
2017 Segmentation of Time Series Based on Kinetic Characteristics for Storage Consumption Prediction
abstract
The Internet services generate huge amount of data, which require large space for storage. Determining device purchase plan turns out to be very important for the service providers. Under-purchasing might lead to data loss, while over-purchasing would result in waste. In this paper, we propose a linear regression based approach to predict the storage demand according to the time series of the storage consumption. We partitioned the storage con-sumption time series into several linear segments, and perform prediction on the last segment using linear regression. Since the position of turning points between adjacent segments and the total number of the segments are both unknown, how to achieve the online segmentation becomes a big challenge. Aiming to solve this problem, we carried out the Kalman-Anova segmentation method. Experiment results show that our method has good accuracy in precision, recall and F-measure values. Moreover, the method is able to segment nonlinear time series as well, suggesting a potential wider application. The proposed method has been deployed in Baidu Inc. and saves about 45 thousand dollars in one of its device purchase program.
Beibei Miao, Xue-bo Jin 0001, Xianping Qu, Shimin Tao, Zhi Zang
ICDCS6
2015 Rapid and robust impact assessment of software changes in large internet-based services
abstract
The detection of performance changes in software change roll-outs in Internet-based services is crucial for an operations team, because it allows timely roll-back of a software change when performance degrades unexpectedly. However, it is infeasible to manually investigate millions of performance measurements of many roll-outs.
Shenglin Zhang, Ying Liu 0024, Dan Pei, Xianping Qu, Shimin Tao, Zhi Zang
CoNEXT6