Hao Yang 0006

dblp:54/4089-6 · DBLP profile ↗
← Back
75ranked-venue papers
5as first author
67since 2021 · last 2026
0000-0001-8861-7010ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 3 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 19 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang
Frontiers Comput. Sci.16
2026 Multiphase and Multitask Prompt Tuning for LLM-Based Context-Aware Machine Translation
abstract
Large language models (LLMs) are typically adapted for context-aware machine translation (MT) by combining both the source sentence and its surrounding sentences into a single input. This unified input is then processed in one go, with the model producing the target translation step by step. However, this method treats the intrasentence and intersentence contexts similarly, even though they play distinct roles. In this study, we present a novel strategy called multiphase prompt tuning (MPT) to address this issue by enabling LLMs to treat these two context types differently. MPT divides the context-aware translation task into three phases: encoding the intersentence context, encoding the source sentence, and the final decoding phase. Each phase incorporates distinct continuous prompts that help the model focus on the appropriate task for each type of context. We also introduce a multitask fine-tuning approach to emphasize the distinction between intersentence and intrasentence contexts and enhance intersentence dependencies. This includes two auxiliary tasks: context-agnostic translation and cross-lingual next sentence generation, which help extract additional information and improve the model's handling of discourse-related challenges.
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE Trans. Neural Networks Learn. Syst.6
2025 SRDC: Semantics-based Ransomware Detection and Classification with LLM-assisted Pre-training
abstract
In recent years, ransomware has emerged as a formidable data security threat, causing significant data privacy breaches that inflict substantial financial, reputational, and operational damages on society. Many studies employ dynamic feature analysis for ransomware detection. However, these methods utilize neither the internal semantic information (semantic information inherent in the features), nor external semantics (the wealth of existing knowledge and expert experience with regard to ransomware detection). Moreover, conventional methods rely on training data from known ransomware families, while zero-day ransomware often has unknown data distribution patterns, posing detection challenges. In this paper, we propose a Semantics-based Ransomware Detection and family Classification (SRDC) framework that can utilize both internal and external semantics of software. To bolster semantic analysis in zero-day attacks, we also design a procedure called LLM-assisted task-adaptive pre-training (LATAP). In LATAP, ransomware semantics from human experts and LLMs are employed to pre-train the detection model (GPT-2). By fully utilizing semantics, the proposed SRDC framework outperforms the SOTA methods by 12.15% for ransomware family classification tasks, and by 4.03% for zero-day ransomware detection tasks. SRDC also exhibits excellent data efficiency, requiring only two ransom families for training, which is only 35% of the data required by existing methods, to achieve a 90%+ accuracy of zero-day ransomware detection in nine unseen ransom families.
Ce Zhou, Yilun Liu 0001, Weibin Meng, Shimin Tao, Weinan Tian, Feiyu Yao, Boxing Chen, Hao Yang 0006
AAAI10
2025 Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
abstract
Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub.
Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ACL (1)7
2025 Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality Estimation
abstract
Quality Estimation (QE) models evaluate the quality of machine translations without reference translations, serving as the reward models for the translation task.Due to the data scarcity, synthetic data generation has emerged as a promising solution.However, synthetic QE data often suffers from distribution shift, which can manifest as discrepancies between pseudo and real translations, or in pseudo labels that do not align with human preferences.To tackle this issue, we introduce DCSQE, a novel framework for alleviating distribution shift in synthetic QE data.To reduce the difference between pseudo and real translations, we employ the constrained beam search algorithm and enhance translation diversity through the use of distinct generation models.DCSQE uses references—i.e., translation supervision signals—to guide both the generation and annotation processes, enhancing the quality of token-level labels.DCSQE further identifies the shortest phrase covering consecutive error tokens, mimicking human annotation behavior, to assign the final phrase-level labels.Specially, we underscore that the translation model can not annotate translations of itself accurately.Extensive experiments demonstrate that DCSQE outperforms SOTA baselines like CometKiwi in both supervised and unsupervised settings.Further analysis offers insights into synthetic data generation that could benefit reward models for other tasks.The code is available at https://github.com/NJUNLP/njuqe.
Xiang Geng, Zhejian Lai, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang
ACL (1)4
2025 Basic Reading Distillation
abstract
Large language models (LLMs) have demonstrated remarkable abilities in various natural language processing areas, but they demand high computation resources which limits their deployment in real-world.Distillation is one technique to solve this problem through either knowledge distillation or task distillation.Both distillation approaches train small models to imitate specific features of LLMs, but they all neglect basic reading education for small models on generic texts that are unrelated to downstream tasks.In this paper, we propose basic reading distillation (BRD) which educates a small model to imitate LLMs basic reading behaviors, such as named entity recognition, question raising and answering, on each sentence.After such basic education, we apply the small model on various tasks including language inference benchmarks and BIG-bench tasks.It shows that the small model can outperform or perform comparable to over 20x bigger LLMs.Analysis reveals that BRD effectively influences the probability distribution of the small model, and has orthogonality to either knowledge distillation or task distillation.
Sirui Miao, Xiangyu Duan, Hao Yang 0006, Min Zhang 0005
ACL (1)4
2025 Enhancing Large Language Models for Document-Level Translation Post-Editing Using Monolingual Data
abstract
The translation capabilities of neural machine translation (NMT) models based on the encoder-decoder framework are extremely potent. Although Large Language Models (LLMs) have achieved remarkable results in many tasks, they have not reached state-of-the-art performance in NMT. However, traditional NMT still faces significant challenges in areas of document translation such as context consistency, tense, and pronoun resolution, where LLMs inherently possess substantial advantages. Instead of directly using LLMs for translation, employing them for Automatic Post-Editing (APE) to post-edit NMT outputs proves to be a viable option. However, document-level bilingual data is extremely scarce. This paper proposes a method that can effectively leverage the capabilities of LLMs to optimize document translation using only monolingual data. By employing two NMT models in opposite directions (Source-to-Target and Target-to-Source), we generate pseudo-document training data for the training of APE. We have identified and resolved the issue between training and inference mode inconsistency brought about by the pseudo-document training data. The final experimental results demonstrate that by using only document-level monolingual data, we can significantly improve the quality of NMT and greatly enhance issues such as reference and contextual consistency in NMT.
Zhiqiang Rao, Hengchao Shang, Daimeng Wei, Hao Yang 0006
COLING7
2025 Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance
abstract
Yilun Liu, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Jian Gao, Zhang Li, Hao Yang, Boxing Chen, Osamu Yoshie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yilun Liu 0001, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Hao Yang 0006, Boxing Chen, Osamu Yoshie
EMNLP10
2025 Generative Annotation for ASR Named Entity Correction
abstract
Yuanchang Luo, Daimeng Wei, Shaojun Li, Hengchao Shang, Jiaxin Guo, Zongyao Li, Zhanglin Wu, Xiaoyu Chen, Zhiqiang Rao, Jinlong Yang, Hao Yang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yuanchang Luo, Daimeng Wei, Hengchao Shang, Zhanglin Wu, Xiaoyu Chen 0004, Zhiqiang Rao, Hao Yang 0006
EMNLP11
2025 Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
abstract
Weiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo, Chen Xu, Xiaofeng Zhao, Long Meng, Yunfei Lu, Min Zhang, Hao Yang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weiqiao Shan, Yuang Li, Yingfeng Luo, Chen Xu 0008, Long Meng, Yunfei Lu, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001
EMNLP10
2025 Investigating Numerical Translation with Large Language Models
abstract
The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explored. This study focuses on evaluating the reliability of LLM-based machine translation systems when handling numerical data. In order to systematically test the numerical translation capabilities of currently open source LLMs, we have constructed a numerical translation dataset between Chinese and English based on real business data, encompassing ten types of numerical translation. Experiments on the dataset indicate that errors in numerical translation are a common issue, with most open-source LLMs faltering when faced with our test scenarios. Especially when it comes to numerical types involving large units like "million", "billion", and "亿" , even the latest llama3.1 8b model can have error rates as high as 20%. Finally, we introduce three potential strategies to mitigate the numerical mistranslations for large units.
Wei Tang 0013, Yuang Li, Min Zhang 0042, Hao Yang 0006
ICASSP8
2025 Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
abstract
Large language models (LLMs) can enhance automatic speech recognition (ASR) systems through generative error correction (GEC). In this paper, we propose Pinyin-enhanced GEC (PY-GEC), which leverages Pinyin—the phonetic representation of Mandarin Chinese—as supplementary information to improve Chinese ASR error correction. Our approach only utilizes synthetic errors for training and employs the one-best hypothesis during inference. Additionally, we introduce a multitask training approach involving conversion tasks between Pinyin and text to align their feature spaces. Experiments on the Aishell-1 and the Common Voice datasets demonstrate that our approach consistently outperforms GEC with text-only input. More importantly, we provide intuitive explanations for the effectiveness of PY-GEC and multitask training from two aspects: 1) increased attention weight on Pinyin features; and 2) aligned feature space between Pinyin and text hidden states.
Yuang Li, Xiaosong Qiao, Wei Tang 0013, Min Zhang 0042, Hao Yang 0006
ICASSP7
2025 Optimizing Speech Multi-View Feature Fusion through Conditional Computation
abstract
Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.
Weiqiao Shan, Yuchen Han 0001, Yuang Li, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001
ICASSP8
2025 "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities
abstract
Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly when dealing with previously unseen named entities. To tackle this challenge, we propose a method for generating Spoken NER data based on a named entity dictionary (NED) to reduce costs. Specifically, we first use a large language model (LLM) to generate sentences from the sampled named entities and then use a text-to-speech (TTS) system to generate the speech. Furthermore, we introduce a noise metric to filter out noisy data. To evaluate our approach, we release a novel Spoken NER benchmark along with a corresponding NED containing 8,853 entities. Experiment results show that our method achieves state-of-the-art (SOTA) performance in the in-domain, zero-shot domain adaptation, and fully zero-shot settings. Our data will be available at https://github.com/DeepLearnXMU/HeardU.
Xiang Geng, Yuang Li, Mengxin Ren, Wei Tang 0013, Jiahuan Li, Zhibin Lan, Min Zhang 0042, Hao Yang 0006, Shujian Huang, Jinsong Su
ICASSP9
2025 A Hybrid Graph Neural Network for Enhanced EEG-Based Depression Detection
abstract
Graph neural networks (GNNs) are gaining increasing popularity for EEG-based depression detection. However, previous GNN-based methods inadequately consider the characteristics of depression, which limit their performance. First, neuroscience studies indicate that patients with depression exhibit both common and individualized brain abnormalities. Previous GNN-based approaches typically focus either on common graph connections to capture common brain abnormalities or on individualized connections to capture individualized patterns, which is insufficient for depression detection. Second, brain network exhibits a hierarchical structure, ranging from channel-level graphs to region-level graphs. This hierarchical structure varies across individuals and contains significant information relevant to detecting depression. However, previous GNN-based methods overlook this individualized hierarchical information. To address these issues, we propose a Hybrid GNN (HybGNN) that combines a Common Graph Neural Network (CGNN) branch using common connections and an Individualized Graph Neural Network (IGNN) branch employing individualized connections. The two branches capture common and individualized depression patterns, respectively, complementing each other. Furthermore, we enhance the HybGNN with a Cross-Branch Hierarchical Information Extractor (CB-HIE) to extract more task-relevant individualized hierarchical information. Extensive experiments on the MODMA and HUSM datasets demonstrate that the proposed HybGNN achieves state-of-the-art performance.
Yiye Wang, Wenming Zheng, Yang Li 0019, Hao Yang 0006
IJCNN4
2025 SuperFC: Selective Data Utilization for a Sustainable and Effective Function-Calling Agent
abstract
The function-calling agent is obtained by performing agent tuning to the large language model (LLM) on function-calling dataset. However, even state-of-the-art datasets (e.g., xlam-function-calling-60k datasets) still contain numerous misleading examples of low-quality data, wasting significant computational resources and result in an unnecessary carbon footprint. Furthermore, such inductive bad data negatively impacts the performance of the agent. In this paper, we propose a set of scoring criteria specifically tailored to evaluate function-calling data and use these criteria to develop a data filtering framework. By applying this framework to filter out low-quality data, we fine-tuned SuperFC, which demonstrates substantial improvements in both sustainability and performance. The SuperFC-7B training process reduced training time from 455 minutes to 85 minutes, resulting in a 80.02% reduction in carbon footprint. Simultaneously, fine-tuning on high-quality data subsets led to performance improvements of up to 3.68%. Additionally, we provide an in-depth analysis of the causes behind the low quality of synthetic function-calling data, offering valuable insights for future data synthesis in this domain. We have also released a high-quality function-calling dataset, available at: https://github.com/Zire-Young/SuperFC
Xinhua Yang, Yilun Liu 0001, Shimin Tao, Chunguang Zhao, Weibin Meng, Minggui He, Chang Su 0001, Hongxia Ma, Jingzhou Du, Hao Yang 0006, Boxing Chen, Chuanwen Li
IJCNN14
2025 PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Tengfei Song, Hao Yang 0006, Zhanglin Wu, Wenming Zheng
INTERSPEECH5
2025 Graph Alignment Using Seed-Oriented Subgraph Matching
abstract
This paper addresses the challenge of unsupervised plain graph alignment, specifically in scenarios where auxiliary information, such as node attributes, is unavailable. Existing alignment algorithms primarily fall into two categories: spectral methods and representation learning-based methods. Spectral methods typically leverage alignment consistency principles, employing heuristic strategies to iteratively infer the alignment matrix. In contrast, representation learning methods focus on encoding the geometric structural features of nodes to generate node representations, thereby transforming the node matching task into a similarity computation based on these representations. While both approaches demonstrate robust performance in the graph alignment domain, their time complexity poses significant concerns. To mitigate this issue, we propose a novel, efficient algorithm grounded in seed-oriented subgraph matching. Our method begins by extracting a limited number of reliable pseudo alignment seeds derived from graph geometric features. Subsequently, we extract the corresponding K-hop seed-oriented subgraphs, allowing us to reformulate the graph alignment problem into a series of subgraph matching tasks. The final alignment matrix is then constructed by aggregating the results of these subgraph matches. Experimental evaluations conducted on public datasets reveal that our method not only improves efficiency but also outperforms current state-of-the-art techniques in terms of accuracy.
Wei Tang 0013, Xinglin Lv, Yuang Li, Min Zhang 0042, Hao Yang 0006
ICMR5
2025 Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided Memory for Document-Level Machine Translation
Yuanchang Luo, Daimeng Wei, Hengchao Shang, Zhiqiang Rao, Zhanglin Wu, Hao Yang 0006
NLPCC (3)11
2025 Improving LLM-Based Document-Level MT with Multi-Knowledge Fusion
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
NLPCC (3)7
2024 Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models
abstract
To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal translations often miss the intended meaning. Traditional methods, which replace idioms using existing knowledge bases (KBs), often lack scale and context-awareness. Addressing these challenges, our approach prioritizes context-awareness and scalability, allowing for offline storage of idioms in a manageable KB size. This ensures efficient serving with smaller models and provides a more comprehensive understanding of idiomatic expressions. We introduce a multilingual idiom KB (IdiomKB) developed using large LMs to address this. This KB facilitates better translation by smaller models, such as BLOOMZ (7.1B), Alpaca (7B), and InstructGPT (6.7B), by retrieving idioms' figurative meanings. We present a novel, GPT-4-powered metric for human-aligned evaluation, demonstrating that IdiomKB considerably boosts model performance. Human evaluations further validate our KB's quality.
Jiangjie Chen, Hao Yang 0006, Shimin Tao, Yanghua Xiao
AAAI5
2024 Submodular-based In-context Example Selection for LLMs-based Machine Translation
abstract
Large Language Models (LLMs) have demonstrated impressive performances across various NLP tasks with just a few prompts via in-context learning. Previous studies have emphasized the pivotal role of well-chosen examples in in-context learning, as opposed to randomly selected instances that exhibits unstable results.A successful example selection scheme depends on multiple factors, while in the context of LLMs-based machine translation, the common selection algorithms only consider the single factor, i.e., the similarity between the example source sentence and the input sentence.In this paper, we introduce a novel approach to use multiple translational factors for in-context example selection by using monotone submodular function maximization.The factors include surface/semantic similarity between examples and inputs on both source and target sides, as well as the diversity within examples.Importantly, our framework mathematically guarantees the coordination between these factors, which are different and challenging to reconcile.Additionally, our research uncovers a previously unexamined dimension: unlike other NLP tasks, the translation part of an example is also crucial, a facet disregarded in prior studies.Experiments conducted on BLOOMZ-7.1B and LLAMA2-13B, demonstrate that our approach significantly outperforms random selection and robust single-factor baselines across various machine translation tasks.
Baijun Ji, Xiangyu Duan, Zhenyu Qiu, Junhui Li 0001, Hao Yang 0006, Min Zhang 0005
LREC/COLING6
2024 Evaluation Dataset for Lexical Translation Consistency in Chinese-to-English Document-level Translation
abstract
Lexical translation consistency is one of the most common discourse phenomena in Chinese-to-English document-level translation. To better evaluate the performance of lexical translation consistency, previous researches assumes that all repeated source words should be translated consistently. However, constraining translations of repeated source words to be consistent will hurt word diversity and human translators tend to use different words in translation. Therefore, in this paper we construct a test set of 310 bilingual news articles to properly evaluate lexical translation consistency. We manually differentiate those repeated source words whose translations are consistent into two types: true consistency and false consistency. Then based on the constructed test set, we evaluate the performance of lexical translation consistency for several typical NMT systems.
Xiangyu Lei, Junhui Li 0001, Shimin Tao, Hao Yang 0006
LREC/COLING4
2024 CB-Whisper: Contextual Biasing Whisper Using Open-Vocabulary Keyword-Spotting
abstract
End-to-end automatic speech recognition (ASR) systems often struggle to recognize rare name entities, such as personal names, organizations and terminologies that are not frequently encountered in the training data. This paper presents Contextual Biasing Whisper (CB-Whisper), a novel ASR system based on OpenAI’s Whisper model that can recognize user-defined name entities by performing open-vocabulary keyword-spotting (KWS) before the decoder. The KWS module leverages text-to-speech (TTS) techniques and a convolutional neural network (CNN) classifier to match the features between the entities and the utterances. To integrate the recognized entities into the Whipser decoder and avoid hallucinations, we carefully crafted multiple prompts with spoken form hints. Experiments show that the KWS module based on Whisper encoder’s features can recognize unseen user-defined keywords effectively. More importantly, the proposed CB-Whisper substantially improves the mixed-error-rate (MER) and entity recall compared to the original Whisper model on three internal datasets and two publicly available datasets including Aishell and ACL datasets that cover English-only, Chinese-only, and code-switching scenarios.
Yuang Li, Yinglu Li, Min Zhang 0042, Chang Su 0001, Mengyao Piao, Xiaosong Qiao, Miaomiao Ma, Hao Yang 0006
LREC/COLING10
2024 Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation
abstract
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Mahong Xia, Zhang Li, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yuan Ge 0001, Yilun Liu 0001, Chi Hu, Weibin Meng, Shimin Tao, Mahong Xia, Boxing Chen, Hao Yang 0006, Tong Xiao 0001
EMNLP10
2024 Cross-Domain Audio Deepfake Detection: Dataset and Analysis
abstract
Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy.Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance.However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models.In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zeroshot TTS models.To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets.Experiments show that, through novel attackaugmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1% and 6.5% respectively.Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data.Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.Our dataset is publicly available 1 .
Yuang Li, Min Zhang 0042, Mengxin Ren, Xiaosong Qiao, Miaomiao Ma, Daimeng Wei, Hao Yang 0006
EMNLP7
2024 DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators
abstract
Generally, the decoder-only large language models (LLMs) are adapted to context-aware neural machine translation (NMT) in a concatenating way, where LLMs take the concatenation of the source sentence (i.e., intrasentence context) and the inter-sentence context as the input, and then to generate the target tokens sequentially.This adaptation strategy, i.e., concatenation mode, considers intrasentence and inter-sentence contexts with the same priority, despite an apparent difference between the two kinds of contexts.In this paper, we propose an alternative adaptation approach, named Decoding-enhanced Multiphase Prompt Tuning (DeMPT), to make LLMs discriminately model and utilize the inter-and intra-sentence context and more effectively adapt LLMs to context-aware NMT.First, DeMPT divides the context-aware NMT process into three separate phases.During each phase, different continuous prompts are introduced to make LLMs discriminately model various information.Second, DeMPT employs a heuristic way to further discriminately enhance the utilization of the source-side interand intra-sentence information at the final decoding phase.Experiments show that our approach significantly outperforms the concatenation method, and further improves the performance of LLMs in discourse modeling.
Xinglin Lyu, Junhui Li 0001, Min Zhang 0042, Daimeng Wei, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP7
2024 CSNet: Contrastive Siamese Network for Robust SLU
abstract
Automatic speech recognition (ASR) results based on clean references are much more accurate than those based on ASR transcripts in spoken language understanding (SLU). Effective utilization of manually-checked clean transcripts is key to improving SLU performance. This paper proposes a siamese network with contrastive learning to enhance SLU effects. A siamese network on sentence pairs that are composed of ASR transcripts and clean transcripts is used for the SLU task. During training, contrastive learning brings closer the sentence-level semantic representations of ASR transcripts and clean transcripts. During inference, k-nearest neighbors (KNN) semantic search via the siamese network first finds the pseudo clean transcript, then forms a sentence pair based on the ASR transcript and pseudo clean transcript for prediction. Experiments on three benchmark datasets prove the effectiveness of our proposed approach, which improves the Intent Classification (IC) performance by over 1.3% on the SLURP dataset.
Hao Yang 0006, Min Zhang 0042, Daimeng Wei
ICASSP1
2024 CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
abstract
Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set11https://github.com/lunyiliu/CoachLM.
Yilun Liu 0001, Shimin Tao, Ming Zhu 0010, Wenbing Ma, Chang Su 0001, Yutai Hou, Min Zhang 0042, Hongxia Ma, Hao Yang 0006, Yanfei Jiang
ICDE13
2024 From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
abstract
Machine Translation Quality Estimation (MTQE) is the task of estimating the quality of machine-translated text in real time without the need for reference translations, which is of great importance for the development of MT. After two decades of evolution, QE has yielded a wealth of results. This article provides a comprehensive overview of QE datasets, annotation methods, shared tasks, methodologies, challenges, and future research directions. It begins with an introduction to the background and significance of QE, followed by an explanation of the concepts and evaluation metrics for word-level QE, sentence-level QE, document-level QE, and explainable QE. The paper categorizes the methods developed throughout the history of QE into those based on handcrafted features, deep learning, and Large Language Models (LLMs), with a further division of deep learning-based methods into classic deep learning and those incorporating pre-trained language models (LMs). Additionally, the article details the advantages and limitations of each method and offers a straightforward comparison of different approaches. Finally, the paper discusses the current challenges in QE research and provides an outlook on future research directions.
Haofei Zhao, Yilun Liu 0001, Shimin Tao, Weibin Meng, Xiang Geng, Chang Su 0001, Min Zhang 0042, Hao Yang 0006
IJCNN9
2024 Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
Daimeng Wei, Hengchao Shang, Zhanglin Wu, Zhiqiang Rao, Yuanchang Luo, Xianghui He, Hao Yang 0006
INTERSPEECH10
2024 Using Large Language Model for End-to-End Chinese ASR and NER
Yuang Li, Min Zhang 0042, Mengxin Ren, Shimin Tao, Jinsong Su, Hao Yang 0006
INTERSPEECH9
2024 A Multitask Training Approach to Enhance Whisper with Open-Vocabulary Keyword Spotting
abstract
The recognition of rare named entities, such as personal names and terminologies, is challenging for automatic speech recognition (ASR) systems, especially when they are not frequently observed in the training data.In this paper, we introduce keyword spotting enhanced Whisper (KWS-Whisper), a novel ASR system that leverages the Whisper model and performs openvocabulary keyword spotting (OV-KWS) on the hidden states of the Whisper encoder to recognize user-defined named entities.These entities serve as prompts for the Whisper decoder.To optimize the model, we propose a multitask training approach that learns OV-KWS and contextual-ASR tasks.We evaluate our approach on Chinese Aishell hot word subsets and two internal code-switching test sets and show that it significantly improves the entity recall compared to the original Whisper model.Moreover, we demonstrate that the OV-KWS can be a plug-andplay module to enhance the ASR error correction methods and frozen Whisper models.
Yuang Li, Min Zhang 0042, Chang Su 0001, Yinglu Li, Xiaosong Qiao, Mengxin Ren, Miaomiao Ma, Daimeng Wei, Shimin Tao, Hao Yang 0006
INTERSPEECH10
2024 RASU: Retrieval Augmented Speech Understanding through Generative Modeling
Hao Yang 0006, Min Zhang 0042, Minghan Wang
INTERSPEECH1
2024 Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies
abstract
Automated log analysis is crucial in modern software-intensive systems for facilitating program comprehension throughout software maintenance and engineering life cycles. Existing methods perform tasks such as log parsing and log anomaly detection by providing a single prediction value without interpretation. However, given the increasing volume of system events, the limited interpretability of analysis results hinders analysts' comprehension of program status and their ability to take appropriate actions. Moreover, these methods require substantial in-domain training data, and their performance declines sharply (by up to 62.5%) in online scenarios involving unseen logs from new domains, a common occurrence due to rapid software updates. In this paper, we propose LogPrompt, a novel interpretable log analysis approach for online scenarios. LogPrompt employs large language models (LLMs) to perform online log analysis tasks via a suite of advanced prompt strategies tailored for log tasks, which enhances LLMs' performance by up to 380.7% compared with simple prompts. Experiments on nine publicly available evaluation datasets across two tasks demonstrate that LogPrompt, despite requiring no in-domain training, outperforms existing approaches trained on thousands of logs by up to 55.9%. We also conduct a human evaluation of LogPrompt's interpretability, with six practitioners possessing over 10 years of experience, who highly rated the generated content in terms of usefulness and readability (averagely 4.42/5). LogPrompt also exhibits remarkable compatibility with open-source and smaller-scale LLMs, making it flexible for practical deployment. Code of LogPrompt is available at https://github.com/lunyiliu/LogPrompt.
Yilun Liu 0001, Shimin Tao, Weibin Meng, Jingyu Wang 0001, Wenbing Ma, Hao Yang 0006, Yanfei Jiang
ICPC8
2024 Multi-Source Log Parsing With Pre-Trained Domain Classifier
abstract
Automated log analysis with AI technologies is commonly used in network, system, and service operation and maintenance to ensure reliability and quality assurance. Log parsing serves as an essential primary stage in log analysis, where unstructured logs are transformed into structured data to facilitate subsequent downstream analysis. However, traditional log parsing algorithms designed for single-domain processing struggle to handle the challenges posed by multi-source log inputs, leading to a decline in parsing accuracy. Adapting these algorithms to multi-source logs often requires extensive manual labeling efforts. To address this, we propose Domain-aware Parser (DA-Parser), a framework that includes a domain classifier to identify the source domains of multi-source logs. This enables the conversion of the multi-source log parsing problem into a series of single-source parsing problems. The classifier is pre-trained on a corpus of logs from 16 domains, eliminating the need for additional human labeling. The predicted source domain tags serve as constraints, limiting the template extraction process to logs from the same domain. Empirical evaluation on a multi-domain dataset demonstrates that DA-Parser outperforms the existing SOTA algorithm by 21.6% in terms of parsing accuracy. The proposed approach also shows potential efficiency improvements, requiring only 6.67% of the time consumed by existing parsers, while maintaining robustness against minor domain classification errors.
Yilun Liu 0001, Shimin Tao, Weibin Meng, Jingyu Wang 0001, Hao Yang 0006, Yanfei Jiang
IEEE Trans. Netw. Serv. Manag.5
2023 Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning
abstract
Quality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains QE models with pseudo QE data and obtains remarkable performance. However, there tends to be inevitable noise in the pseudo data, hindering models from learning QE accurately. Our study shows that the noise mainly comes from the differences between pseudo and real translation outputs. To handle this problem, we propose CLQE, a denoising pre-training framework for QE based on curriculum learning. More specifically, we propose to measure the degree of noise in the pseudo QE data with some metrics based on statistical or distributional features. With the guidance of these metrics, CLQE gradually pre-trains the QE model using data from cleaner to noisier. Experiments on various benchmarks reveal that CLQE outperforms DirectQE and other strong baselines. We also show that with our framework, pre-training converges faster than directly using the pseudo data. We make our CLQE code available (https://github.com/NJUNLP/njuqe).
Xiang Geng, Jiahuan Li, Shujian Huang, Hao Yang 0006, Shimin Tao, Jiajun Chen 0001
AAAI5
2023 Text Style Transfer Back-Translation
abstract
Daimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li, Minghan Wang, Jiaxin Guo, Xiaoyu Chen, Zhengzhe Yu, Hao Yang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Daimeng Wei, Zhanglin Wu, Hengchao Shang, Minghan Wang, Xiaoyu Chen 0004, Zhengzhe Yu, Hao Yang 0006
ACL (1)9
2023 Knowledge Prompt for Whisper: An ASR Entity Correction Approach with Knowledge Base
abstract
Entity correction is crucial in Automatic Speech TABLE I Recognition (ASR), since erroneous entities seriously affect our understanding of ASR results. In this paper, in order to correct entity errors, we propose a knowledge prompt approach for Whisper (a recent ASR model trained with a corpus containing 680k hours of labeled speech recorded in various conditions). For a given audio, our approach consists of three steps: (1) obtaining its ASR result by Whisper; (2) fuzzy matching the ASR result with a knowledge base to obtain candidate entities; (3) using the candidate entities as a prompt to obtain the final ASR result by Whisper again. We conduct experiments on the test dataset of open-source Chinese speech corpus AISHELLNER. Experimental results show that our approach not only significantly improves the entity recall rate in ASR results (from 70.97% to 84.82%), but also reduces the overall Character Error Rate (CER).
Min Zhang 0042, Xiaosong Qiao, Chang Su 0001, Yinglu Li, Yuang Li, Ming Zhu 0010, Mengyao Piao, Shimin Tao, Hao Yang 0006, Yanfei Jiang
IEEE Big Data11
2023 DA-Parser: A Pre-trained Domain-aware Parsing Framework for Heterogeneous Log Analysis
abstract
Automated log analysis is widely applied in modern software-intensive systems to ensure resilience and sustainability, where log parsing is a vital initial step, converting unstructured logs into structured data for downstream analysis. However, traditional log parsing algorithms are designed to process logs within a single domain. As cross-domain dependencies and interactions between sub-modules of software systems increase, these algorithms struggle to handle the challenges posed by multi-domain log inputs, which results in a significant decline in parsing accuracy when facing heterogeneous logs. Additionally, current solutions for heterogeneous log parsing require extensive manual labeling efforts. In this paper, we propose Domain-aware Parser (DA-Parser), a framework that consists of a domain-aware head to identify the source domains of heterogeneous logs and then converts the multi-domain log parsing problem into a series of single-domain parsing problems. The domain-aware head is pretrained using a corpus of logs from 16 domains, which allows for the classification of the source domains of most heterogeneous log set without additional human labeling. Source domain tags predicted by the domain-aware head serve as a constraint to limit the template extraction process to logs from the same domain. Empirical evaluation is conducted on a multi-domain dataset containing logs from 7 domains. DA-Parser can be integrated with existing single-domain algorithms and are compatible with them, achieving superior parsing accuracy with an average of 9.26% improvement compared with single-domain algorithms.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Jingyu Wang 0001, Chang Su 0001, Weinan Tian, Min Zhang 0042, Hao Yang 0006, Xun Chen 0001
COMPSAC9
2023 Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam Search
abstract
Xiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She, Wei Zou, Shimin Tao, Hao Yang, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Xiang Geng, Zhejian Lai, Shuaijie She, Shimin Tao, Hao Yang 0006, Jiajun Chen 0001, Shujian Huang
EMNLP7
2023 UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
abstract
Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets.
Minghan Wang, Xiaosong Qiao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP12
2023 Zephyr: Zero-Shot Punctuation Restoration
abstract
Punctuation restoration can be crucial for the cascade speech translation system. Traditional approaches typically treat it as a sequential tagging problem, predicting which punctuation mark should follow a given word. However, this often requires significant computational and storage resources for full-stage training or fine-tuning. Our argument is that pre-trained language models (PLMs) can directly leverage their learned knowledge for punctuation generation, making additional training unnecessary. In this paper, we propose the Zephyr algorithm, which utilizes PLMs to perform zero-shot and few-shot punctuation restoration for both offline and streaming scenarios. Our experimental results demonstrate that, in comparison to fine-tuning-based baselines, Zephyr achieves competitive performance while requiring little to no training cost and exhibiting better generalizability in zeroshot and few-shot settings.
Minghan Wang, Yinglu Li, Xiaosong Qiao, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ICASSP8
2023 WhiSLU: End-to-End Spoken Language Understanding with Whisper
Minghan Wang, Yinglu Li, Xiaosong Qiao, Hengchao Shang, Daimeng Wei, Shimin Tao, Min Zhang 0042, Hao Yang 0006
INTERSPEECH10
2023 Biglog: Unsupervised Large-scale Pre-training for a Unified Log Representation
abstract
Automated log analysis has been widely applied in modern data-center network, performing critical tasks such as log parsing, log anomaly detection and log-based failure prediction. However, existing approaches rely on hand-crafted features or domain-specific vectors to represent logs, which are either laborious in manual efforts or ineffective facing multiple domains in a system. Furthermore, general-purpose word embeddings are not optimized for log data, thus are data-inefficient in handling complex log analysis tasks. In this paper, we present a pre-training phase for language models to understand both in-sentence and cross-sentence features of logs, resulting in a unified representation of logs that is well-suited for various downstream analysis tasks. The pre-training phase is unsupervised, utilizing 0.45 billion logs from 16 diverse domains. Experiments on 12 publicly available evaluation datasets across 3 tasks indicate superiority of our approach against existing approaches, especially in online scenarios with limited historical logs. Our approach also exhibits remarkable few-shot learning ability and domain-adaptiveness, which not only outperforms existing approaches using only 0.0025% of their required training data, but also adapts into new domains via only a few in-domain logs. We release our code and pre-trained model.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Zuomin Ren, Hao Yang 0006, Xun Chen 0001, Yuming Xie, Chang Su 0001, Xiaosong Oiao, Weinan Tian, Yichen Zhu 0001
IWQoS5
2023 Twin Graph Attention Network with Evolution Pattern Learner for Few-Shot Temporal Knowledge Graph Completion
Shuai Zhao 0001, Bo Cheng 0001, Hao Yang 0006
KSEM (1)4
2023 HWCGEC:HW-TSC's 2023 Submission for the NLPCC2023's Chinese Grammatical Error Correction Task
Chang Su 0001, Xiaosong Qiao, Min Zhang 0042, Hao Yang 0006, Ming Zhu 0010, Wenbing Ma
NLPCC (3)5
2023 Multi-order Matched Neighborhood Consistent Graph Alignment in a Union Vector Space
abstract
In this paper, we study the unsupervised plain graph alignment problem, which aims to find node correspondences across two graphs without any side information. The majority of previous works addressed UPGA based on structural information, which will inevitably lead to subgraph isomorphism issues. That is, unaligned nodes could take similar local structural information. To mitigate this issue, we present the Multi-order Matched Neighborhood Consistent (MMNC) which tries to match nodes by aligning the learned node embeddings with only a small number of pseudo alignment seeds. In particular, we extend matched neighborhood consistency (MNC) to vector space and further develop embedding-based MNC (EMNC). By minimizing the EMNC-based loss function, we can utilize the limited pseudo alignment seeds to approximate the orthogonal transformation matrix between two groups of node embeddings with high efficiency and accuracy. Through extensive experiments on public benchmarks, we show that the proposed methods achieve a good balance between alignment accuracy and speed over multiple datasets compared with existing methods.
Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039, Hao Yang 0006, Shimin Tao
SIGIR6
2023 Weakly Supervised Entity Alignment with Positional Inspiration
abstract
The current success of entity alignment (EA) is still mainly based on large-scale labeled anchor links. However, the refined annotation of anchor links still consumes a lot of manpower and material resources. As a result, an increasing number of works based on active learning, few-shot learning, or other deep network learning techniques have been developed to address the performance bottleneck caused by a lack of labeled data. These works focus either on the strategy of choosing more informative labeled data or on the strategy of model training, while it remains opaque why existing popular EA models (e.g., GNN-based models) fail the EA task with limited labeled data. To overcome this issue, this paper analyzes the problem of weakly supervised EA from the perspective of model design and proposes a novel weakly supervised learning framework, Position Enhanced Entity Alignment (PEEA). Besides absorbing structural and relational information, PEEA aims to increase the connections between far-away entities and labeled ones by incorporating positional information into the representation learning with a Position Attention Layer (PAL). To fully utilize the limited anchor links, we further introduce a novel position encoding method that considers both anchor links and relational information from a global view. The proposed position encoding will be fed into PEEA as additional entity features. Extensive experiments on public datasets demonstrate the effectiveness of PEEA.
Wei Tang 0013, Fenglong Su, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Shimin Tao, Hao Yang 0006
WSDM7
2023 TransAM: Transformer appending matcher for few-shot knowledge graph completion
Shuai Zhao 0001, Bo Cheng 0001, Hao Yang 0006
Neurocomputing4
2023 Collective Human Opinions in Semantic Textual Similarity
abstract
Abstract Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as gold standard. Averaging masks the true distribution of human opinions on examples of low agreement, and prevents models from capturing the semantic vagueness that the individual ratings represent. In this work, we introduce USTS, the first Uncertainty-aware STS dataset with ∼15,000 Chinese sentence pairs and 150,000 labels, to study collective human opinions in STS. Analysis reveals that neither a scalar nor a single Gaussian fits a set of observed judgments adequately. We further show that current STS models cannot capture the variance caused by human disagreement on individual instances, but rather reflect the predictive confidence over the aggregate dataset.
Yuxia Wang 0003, Shimin Tao, Hao Yang 0006, Timothy Baldwin, Karin Verspoor
Trans. Assoc. Comput. Linguistics4
2023 P-Transformer: Towards Better Document-to-Document Neural Machine Translation
abstract
Directly training a document-to-document (Doc2Doc) neural machine translation (NMT) via Transformer from scratch, especially on small datasets, usually fails to converge. Our dedicated probing tasks show that 1) both the absolute position and relative position information gets gradually weakened or even vanished once it reaches the upper encoder layers, and 2) the vanishing of absolute position information in encoder output causes the training failure of Doc2Doc NMT. To alleviate this problem, we propose a position-aware Transformer (P-Transformer) to enhance both the absolute and relative position information in both self-attention and cross-attention. Specifically, we integrate absolute positional information, i.e., position embeddings, into the query-key pairs both in self-attention and cross-attention through a simple yet effective addition operation. Moreover, we also integrate relative position encoding in self-attention. The proposed P-Transformer utilizes sinusoidal position encoding and does not require any task-specified position embedding, segment embedding, or attention mechanism. Through the above methods, we build a Doc2Doc NMT model with P-Transformer, which ingests the source document and completely generates the target document in a sequence-to-sequence (seq2seq) way. In addition, P-Transformer can be applied to seq2seq-based document-to-sentence (Doc2Sent) and sentence-to-sentence (Sent2Sent) translations. Extensive experimental results of Doc2Doc NMT show that P-Transformer significantly outperforms strong baselines on the widely-used 9 document-level datasets in 7 language pairs, covering small-, middle-, and large-scales, and achieves a new state-of-the-art. Experimentation on discourse phenomena shows that our Doc2Doc NMT models improve the translation quality in both BLEU and discourse coherence. We make our code available on Github.
Yachao Li 0003, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Exploiting Spatial-Temporal Behavior Patterns for Fraud Detection in Telecom Networks
abstract
Fraud detection in telecom network is a crucial problem that threatens users’ privacy and property security. In recent years, fraudsters adopt more advanced camouflage strategies to avoid being detected by traditional algorithms. To deal with these new types of fraud, it is necessary to analyze the integrated spatial-temporal features, which are rarely involved in existing literature. In this article, we propose a novel fraud detection model based on the intertwined spatial-temporal patterns of user behaviors. Specifically, we first introduce the extension of statistical and interactive features to dynamic call patterns, and build a probabilistic model to simulate users’ call behaviors. Then the sequential patterns reflecting users’ own behaviors are obtained by the mixture Hidden Markov Models, and the structural patterns reflecting the collaboration between users in the telecom network are obtained by the attention-based Graph-SAGE model. Finally, our model outputs a fraud score for each user to detect potential fraudsters. We conduct extensive experiments on a real-world telecom dataset. The experimental results demonstrate that our intertwined spatial-temporal call patterns can effectively represent user behavior and improve the accuracy of fraud detection compared with state-of-the-art methods. The results also validate the efficiency and the interpretability of our model.
Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006, Jianxin Liao, Zhu Han 0001
IEEE Trans. Dependable Secur. Comput.6
2022 Exploring Entity Interactions for Few-Shot Relation Learning (Student Abstract)
abstract
Few-shot relation learning refers to infer facts for relations with a few observed triples. Existing metric-learning methods mostly neglect entity interactions within and between triples. In this paper, we explore this kind of fine-grained semantic meaning and propose our model TransAM. Specifically, we serialize reference entities and query entities into sequence and apply transformer structure with local-global attention to capture intra- and inter-triple entity interactions. Experiments on two public datasets with 1-shot setting prove the effectiveness of TransAM.
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
AAAI5
2022 EntityRank: Unsupervised Mining of Bilingual Named Entity Pairs from Parallel Corpora for Neural Machine Translation
abstract
As Neural Machine Translation (NMT) heavily relies on training data, finding an effective method to help NMT make better use of limited data is of great significance. In this paper, with the motivation of the famous Google’s PageRank algorithm, we propose a novel unsupervised method EntityRank for mining bilingual named entity pairs from parallel corpora, which involves three critical components (Generator, Scorer and Filter). To apply the pairs mined by EntityRank to NMT, we design a data augmentation strategy for the state-of-the-art (SOTA) model Transformer. From the experimental results on the CCMT20 English-Chinese and WMT14 English-German news parallel corpora, it can be seen that the unsupervised method EntityRank could obtain relatively high quality bilingual named entity pairs; and with the designed data augmentation strategy, the mined pairs could not only significantly improve the translation quality of their covered data, but also benefit the translation quality of the overall data.
Min Zhang 0042, Hao Yang 0006, Xiaosong Qiao, Shimin Tao, Yanfei Jiang
IEEE Big Data3
2022 Diformer: Directional Transformer for Neural Machine Translation
abstract
Autoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current combination frameworks focus more on the integration of multiple decoding paradigms with a unified generative model, e.g. Masked Language Model. However, the generalization can be harmful on the performance due to the gap between training objective and inference. In this paper, we aim to close the gap by preserving the original objective of AR and NAR under a unified framework. Specifically, we propose the Directional Transformer (Diformer) by jointly modelling AR and NAR into three generation directions (left-to-right, right-to-left and straight) with a newly introduced direction variable, which works by controlling the prediction of each token to have specific dependencies under that direction. The unification achieved by direction successfully preserves the original dependency assumption used in AR and NAR, retaining both generalization and performance. Experiments on 4 WMT benchmarks demonstrate that Diformer outperforms current united-modelling works with more than 1.5 BLEU points for both AR and NAR decoding, and is also competitive to the state-of-the-art independent AR and NAR models.
Minghan Wang, Yuxia Wang 0003, Daimeng Wei, Hengchao Shang, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
EAMT11
2022 Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine Translation
abstract
In this paper we aim to relieve the issue of lexical translation inconsistency for documentlevel neural machine translation (NMT) by modeling consistency preference for lexical chains which consist of repeated words in a source-side document and provide a representation of the lexical consistency structure of the document.Specifically, we first propose lexical-consistency attention to capture consistency context among words in the same lexical chains.Then for each lexical chain we define and learn a consistency-tailored latent variable, which will guide the translation of corresponding sentences to enhance lexical translation consistency.Experimental results on Chinese→English and French→English document-level translation tasks show that our approach not only significantly improves translation performance in BLEU, but also substantially alleviates the problem of the lexical translation inconsistency.
Xinglin Lyu, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP4
2022 Tackling Solitary Entities for Few-Shot Knowledge Graph Completion
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
KSEM (1)5
2022 CCDC: A Chinese-Centric Cross Domain Contrastive Learning Framework
Hao Yang 0006, Shimin Tao, Minghan Wang, Min Zhang 0042, Daimeng Wei, Shuai Zhao 0001, Miaomiao Ma
KSEM (2)1
2022 Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical Constraints
abstract
Chun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu, Hao Yang, Qin Ying, Shimin Tao, Yanghua Xiao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu 0026, Hao Yang 0006, Shimin Tao, Yanghua Xiao
NAACL-HLT5
2022 Augmented Topic-Specific Summarization for Domain Dialogue Text
Zhiqiang Rao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Zhanglin Wu, Lizhi Lei, Hao Yang 0006
NLPCC (2)10
2022 Explore Modeling Relation Information and Direction Information in KBQA
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
Neurocomputing5
2021 Integrating Subgraph-Aware Relation and Direction Reasoning for Question Answering
abstract
Question Answering (QA) models over Knowledge Bases (KBs) are capable of providing more precise answers by utilizing relation information among entities. Although effective, most of these models solely rely on fixed relation representations to obtain answers for different question-related KB subgraphs. Hence, the rich structured information of these subgraphs may be overlooked by the relation representation vectors. Meanwhile, the direction information of reasoning, which has been proven effective for the answer prediction on graphs, has not been fully explored in existing work. To address these challenges, we propose a novel neural model, Relation-updated Direction-guided Answer Selector (RDAS), which converts relations in each subgraph to additional nodes to learn structure information. Additionally, we utilize direction information to enhance the reasoning ability. Experimental results show that our model yields substantial improvements on two widely used datasets.
Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Ivan Sekulic, Guoshun Nan
ICASSP6
2021 On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 0002, Hao Yang 0006, Qun Liu 0001, Jakob Grue Simonsen
ICLR5
2021 HI-CMLM: Improve CMLM with Hybrid Decoder Input
abstract
Mask-predict CMLM (Ghazvininejad et al., 2019) has achieved stunning performance among non-autoregressive NMT models, but we find that the mechanism of predicting all of the target words only depending on the hidden state of [MASK] is not effective and efficient in initial iterations of refinement, resulting in ungrammatical repetitions and slow convergence.In this work, we mitigate this problem by combining copied source with embeddings of [MASK] in decoder.Notably.it's not a straightforward copying that is shown to be useless, but a novel heuristic hybrid strategy -fence-mask.Experimental results show that it gains consistent boosts on both WMT14 En↔De and WMT16 En↔Ro corpus by 0.5 BLEU on average, and 1 BLEU for lessinformative short sentences.This reveals that incorporating additional information by proper strategies is beneficial to improve CMLM, particularly translation quality of short texts and speeding up early-stage convergence.
Minghan Wang, Yuxia Wang 0003, Chang Su 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
INLG9
2021 Make the Blind Translator See The World: A Novel Transfer Learning Solution for Multimodal Machine Translation
abstract
Based on large-scale pretrained networks and the liability to be easily overfitting with limited labelled training data of multimodal translation (MMT) is a critical issue in MMT. To this end and we propose a transfer learning solution. Specifically and 1) A vanilla Transformer is pre-trained on massive bilingual text-only corpus to obtain prior knowledge; 2) A multimodal Transformer named VLTransformer is proposed with several components incorporated visual contexts; and 3) The parameters of VLTransformer are initialized with the pre-trained vanilla Transformer and then being fine-tuned on MMT tasks with a newly proposed method named cross-modal masking which forces the model to learn from both modalities. We evaluated on the Multi30k en-de and en-fr dataset and improving up to 8% BLEU score compared with the SOTA performance. The experimental result demonstrates that performing transfer learning with monomodal pre-trained NMT model on multimodal NMT tasks can obtain considerable boosts.
Minghan Wang, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006
MTSummit (1)7
2021 Deep graph alignment network
Wei Tang 0013, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006
Neurocomputing6
2020 HGMAN: Multi-Hop and Multi-Answer Question Answering Based on Heterogeneous Knowledge Graph (Student Abstract)
abstract
Multi-hop question answering models based on knowledge graph have been extensively studied. Most existing models predict a single answer with the highest probability by ranking candidate answers. However, they are stuck in predicting all the right answers caused by the ranking method. In this paper, we propose a novel model that converts the ranking of candidate answers into individual predictions for each candidate, named heterogeneous knowledge graph based multi-hop and multi-answer model (HGMAN). HGMAN is capable of capturing more informative representations for relations assisted by our heterogeneous graph, which consists of multiple entity nodes and relation nodes. We rely on graph convolutional network for multi-hop reasoning and then binary classification for each node to get multiple answers. Experimental results on MetaQA dataset show the performance of our proposed model over all baselines.
Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Guoshun Nan
AAAI6
2020 Deep Spatio-Temporal Multiple Domain Fusion Network for Urban Anomalies Detection
abstract
Multiple domain fusion has been widely used for urban anomalies forecasting problem, as urban anomalies such as traffic accidents or illegal assembly are usually caused by many complex factors and they would affect many fields. Although many efforts have been devoted to fusing multiple datasets for anomalies detection, most of the work is to extract the spatio-temporal features one by one from multiple datasets and then fuse to get the result or anomaly score. However, the correlation between data from multiple domains at each moment is ignored, which is especially important when detecting anomalies by analyzing the impacts from multiple datasets. In this paper, we propose a novel end-to-end deep learning based framework, namely deep spatio-temporal multiple domain fusion network to collect the impacts of urban anomalies on multiple datasets and detect anomalies in each region of the city at next time interval in turn. We formulate the problem on a weighted graph and obtain spatiotemporal features with adaptive graph convolution and temporal convolution. In addition, a cross-domain convolution network is applied to fully obtain connection between multiple domains. We evaluate our method with real-world dataset collected in New York City and experiments on our model show the advantages nearly 10% beyond the state-of-the-art urban anomalies detection methods.
Ruiqiang Liu, Shuai Zhao 0001, Bo Cheng 0001, Hao Yang 0006, Haina Tang, Taoyu Li
CIKM4
2020 Modelling Long-distance Node Relations for KBQA with Global Dynamic Graph
abstract
The structural information of Knowledge Bases (KBs) has proven effective to Question Answering (QA).Previous studies rely on deep graph neural networks (GNNs) to capture rich structural information, which may not model node relations in particularly long distance due to oversmoothing issue.To address this challenge, we propose a novel framework GlobalGraph, which models long-distance node relations from two views: 1) Node type similarity: GlobalGraph assigns each node a global type label and models long-distance node relations through the global type label similarity; 2) Correlation between nodes and questions: we learn similarity scores between nodes and the question, and model long-distance node relations through the sum score of two nodes.We conduct extensive experiments on two widely used multi-hop KBQA datasets to prove the effectiveness of our method.
Shuai Zhao 0001, Jiale Han 0001, Bo Cheng 0001, Hao Yang 0006, Jianchang Ao, Zhenzi Li
COLING5
2020 Unified Humor Detection Based on Sentence-pair Augmentation and Transfer Learning
abstract
We propose a unified multilingual model for humor detection which can be trained under a transfer learning framework. 1) The model is built based on pre-trained multilingual BERT, thereby is able to make predictions on Chinese, Russian and Spanish corpora. 2) We step out from single sentence classification and propose sequence-pair prediction which considers the inter-sentence relationship. 3) We propose the Sentence Discrepancy Prediction (SDP) loss, aiming to measure the semantic discrepancy of the sequence-pair, which often appears in the setup and punchline of a joke. Our method achieves two SoTA and a second-place on three humor detection corpora in three languages (Russian, Spanish and Chinese), and also improves F1-score by 4%-6%, which demonstrates the effectiveness of it in humor detection tasks.
Minghan Wang, Hao Yang 0006, Shiliang Sun
EAMT2
2020 Efficient Transfer Learning for Quality Estimation with Bottleneck Adapter Layer
abstract
The Predictor-Estimator framework for quality estimation (QE) is commonly used for its strong performance. Where the predictor and estimator works on feature extraction and quality evaluation, respectively. However, training the predictor from scratch is computationally expensive. In this paper, we propose an efficient transfer learning framework to transfer knowledge from NMT dataset into QE models. A Predictor-Estimator alike model named BAL-QE is also proposed, aiming to extract high quality features with pre-trained NMT model, and make classification with a fine-tuned Bottleneck Adapter Layer (BAL). The experiment shows that BAL-QE achieves 97% of the SOTA performance in WMT19 En-De and En-Ru QE tasks by only training 3% of parameters within 4 hours on 4 Titan XP GPUs. Compared with the commonly used NuQE baseline, BAL-QE achieves 47% (En-Ru) and 75% (En-De) of performance promotions.
Hao Yang 0006, Minghan Wang
EAMT1
2020 ST-MFM: A Spatiotemporal Multi-Modal Fusion Model for Urban Anomalies Prediction
abstract
Urban anomaly prediction is of great importance for urban management and public safety. Accurate anomaly prediction can avoid much unnecessary loss. Urban anomalies are usually caused by many complex factors, such as festivals, demonstrations and market promotions. It is not possible to predict anomalies from the perspective of reason, thus, most of the previous work analyzes the impacts of anomalies from multiple crowd flow datasets and observes the shift to ordinary distribution when they occur. Most existing models use observation-based methods to extract relevant spatiotemporal features, which are difficult to fully extract hidden relationships and eventually lead to low accuracy and low recall. In this paper, we propose an end-to-end deep learning based approach, called spatiotemporal multi-modal fusion model to collect the impacts of urban anomalies on multiple crowd flow datasets and predict anomalies in each region of the city for next time interval in turn. More specifically, we model the city into a graph and regard each region as a node. We use graph convolution network to obtain its spatial features and use gate recurrent units to obtain its temporal features. The features of those multiple modalities are further aggregated with points of interest in a two-stage-fusion method for assigning different weights to different functional regions. We evaluate our method using five datasets associated with New York City: 311 complaints, taxicab data, bike rental data, points of interest and road network dataset. Results show the advantages nearly 10% beyond the-state-of-the-art urban anomalies prediction methods.
Ruiqiang Liu, Shuai Zhao 0001, Bo Cheng 0001, Hao Yang 0006, Haina Tang, Fangfang Yang
ECAI4
2020 DVKCM: Knowledge-guided Conversation Generation with Dynamic Vocabulary
abstract
Knowledge-guided conversation models, whose inputs are current input sentence with its background knowledge, make the generation of responses more informative and meaningful. Existing methods assume that words in responses come from the vocabulary of the whole corpus. However, for specific input and knowledge, only a small vocabulary is useful in prediction and other words lead to uncorrelated noise. In this paper, we propose a Dynamic Vocabulary based Knowledge-guided Conversation Model (DVKCM). Inspired by dynamic vocabulary mechanism, DVKCM adopts the vocabulary construction module to allocate the sentence-level vocabulary which relates to the input sentence and background knowledge, and then only uses the small vocabulary to execute the decoding part. Through the sentence-level vocabulary mechanism, we reduce the generation of noise effectively. Experiments on both automatic and human evaluation verify the performance of our model compared with previous models. Moreover, we find that dynamic vocabulary can be applied to other conversation models to improve their performance.
Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Xiangsheng Wei, Hao Yang 0006
IJCNN7
2008 A Dynamic Agent-Based Web Service Invocation Infrastructure
abstract
Web services have led a revolution of Internet technology architecture by their platform-independence, language- independence and other characters. But traditional Web service architecture is based on "Client/Server" model, where server is always providing service reactively. Software agents are now increasingly used in commercial applications to solve complex engineering problems, for their autonomous, proactive and social capabilities. And these applications often make use of Web services. As such, this paper presents a Web service invocation infrastructure based on software agents. The infrastructure is a hybrid peer-to-peer model, using agents to describe service providers and service customers. This invocation model is more flexible than traditional Web service model, for (l)agents can invoke services in a proactive manner no matter whether they act like service customers or providers, and(2)agents can also act as multi-role actors in service domain. And with inspiration from Aspect-Oriented programming, web services are mapped as aspects, while agents are mapped as node. In this way, Web service policies in an agent can be considered to be form an filter chain, either incoming filter chain or outgoing filter chain, which is used to describe agent's request or response filter policies. When service contractor satisfies both the incoming filter chain and the outgoing filter chain at the same time, the corresponding service can be invoked dynamically. And experiments show that dynamic service invocation can be achieved in our infrastructure.
Hao Yang 0006, Junliang Chen 0001, Xiangwu Meng
ACHI1