EDBT 2026 Demo / reviewers in the wild / expert
Gongshen Liu
dblp:02/2134
· DBLP profile ↗
100ranked-venue papers
0as first author
57since 2021 · last 2026
0000-0001-5194-1570ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 11 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving AgentsabstractRecent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks.With the paradigm shift toward self-evolving agentic learning, models are increasingly expected to learn from trajectories by synthesizing tools or accumulating explicit experiences.However, prevailing methods typically rely on large-scale LLMs or multi-agent frameworks, which hinder their deployment in resource-constrained environments.The inherent sparsity of outcome-based rewards also poses a substantial challenge, as agents typically receive feedback only upon completion of tasks.To address these limitations, we introduce a Tool-Memory based self-evolving agentic framework SEARL.Unlike approaches that directly utilize interaction experiences, our method constructs a structured experience memory that integrates planning with execution.This provides a novel state abstraction that facilitates generalization across analogous contexts, such as tool reuse.Consequently, agents extract explicit knowledge from historical data while leveraging inter-trajectory correlations to densify reward signals.We evaluate our framework on knowledge reasoning and mathematics tasks, demonstrating its effectiveness in achieving more practical and efficient learning 1 . Xinshun Feng, Xinhao Song, Gongshen Liu |
ACL (1) | 4 |
| 2026 | Domain Adaptation of MLLM-Based GUI Agents in Documentation-Rich Environments with Standard Operating Procedures
Lingzhong Dong, Pengzhou Cheng, Zongru Wu, Zheng Wu 0005, Gongshen Liu, Zhuosheng Zhang 0001 |
KSEM (3) | 6 |
| 2026 | OpenImplicit: Benchmarking Implicit Reasoning in MLLMs via Open-Ended Evaluation
Jidong Li, Xiaofei Yin, Shuheng Zhou 0001, Haodong Zhao, Sufeng Duan, Gongshen Liu, Huijia Zhu |
ICMR | 8 |
| 2026 | A robust multi-source transfer classification method based on belief functions for cross-domain pattern recognition
Linqing Huang, Jinfu Fan, Gongshen Liu, Shi-Lin Wang |
Int. J. Approx. Reason. | 3 |
| 2025 | ALIS: Aligned LLM Instruction Security Strategy for Unsafe Input PromptabstractIn large language models, existing instruction tuning methods may fail to balance the performance with robustness against attacks from user input like prompt injection and jailbreaking. Inspired by computer hardware and operating systems, we propose an instruction tuning paradigm named Aligned LLM Instruction Security Strategy (ALIS) to enhance model performance by decomposing user inputs into irreducible atomic instructions and organizing them into instruction streams which will guide the response generation of model. ALIS is a hierarchical structure, in which user inputs and system prompts are treated as user and kernel mode instructions respectively. Based on ALIS, the model can maintain security constraints by ignoring or rejecting the input instructions when user mode instructions attempt to conflict with kernel mode instructions. To build ALIS, we also develop an automatic instruction generation method for training ALIS, and give one instruction decomposition task and respective datasets. Notably, the ALIS framework with a small model to generate instruction streams still improve the resilience of LLM to attacks substantially without any lose on general capabilities. Xinhao Song, Sufeng Duan, Gongshen Liu |
COLING | 3 |
| 2025 | Gracefully Filtering Backdoor Samples for Generative Large Language Models without RetrainingabstractBackdoor attacks remain significant security threats to generative large language models (LLMs). Since generative LLMs output sequences of high-dimensional token logits instead of low-dimensional classification logits, most existing backdoor defense methods designed for discriminative models like BERT are ineffective for generative LLMs. Inspired by the observed differences in learning behavior between backdoor and clean mapping in the frequency space, we transform gradients of each training sample, directly influencing parameter updates, into the frequency space. Our findings reveal a distinct separation between the gradients of backdoor and clean samples in the frequency space. Based on this phenomenon, we propose Gradient Clustering in the Frequency Space for Backdoor Sample Filtering (GraCeFul), which leverages sample-wise gradients in the frequency space to effectively identify backdoor samples without requiring retraining LLMs. Experimental results show that GraCeFul outperforms baselines significantly. Notably, GraCeFul exhibits remarkable computational efficiency, achieving nearly 100% recall and F1 scores in identifying backdoor samples, reducing the average success rate of various backdoor attacks to 0% with negligible drops in clean accuracy across multiple free-style question answering datasets. Additionally, GraCeFul generalizes to Llama-2 and Vicuna. The codes are publicly available at https://github.com/ZrW00/GraceFul. Zongru Wu, Pengzhou Cheng, Lingyong Fang, Zhuosheng Zhang 0001, Gongshen Liu |
COLING | 5 |
| 2025 | Towards A Distribution Alignment Framework for Incomplete Data ClassificationabstractMissing attribute values frequently affect data classification, reducing accuracy as most models rely on complete datasets. Imputing missing values is typically used to restore data completeness, which is essential for building models. The effectiveness of imputation significantly impacts the classification accuracy. Therefore, improving imputed values’ quality is crucial for better classification outcomes. Here, we introduce a new distribution alignment framework (DAF) to address classification issues with complete training data but incomplete test data. Initially, DAF imputes missing test data values using mean vectors from complete training data, minimizing the first-order distributional discrepancies. Next, it aligns the second-order statistical distributions, specifically covariance matrices, of both training and imputed test data to derive a feature transformation matrix. This matrix generates new feature representations for the incomplete test data. The classifier trained on the complete training data then classifies the imputed test data under this new feature representation. The experiments on several benchmark datasets show that DAF usually outperforms many advanced methods, achieving the higher classification performance. Linqing Huang, Jinfu Fan, Shi-Lin Wang, Gongshen Liu, Shouxuan Liu |
ICASSP | 4 |
| 2025 | Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge Editing
Lingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu, Huijia Zhu, Zhuosheng Zhang 0001, Gongshen Liu |
ICCV | 8 |
| 2025 | New Multi-Source Distributed Transfer Learning FrameworkabstractIn pattern recognition, where the labeled data is scarce, transfer learning (also called domain adaptation in some cases) methods frequently come into play to transfer knowledge from the source domains to bolster the construction of classification models within the target domain. The judicious fusion of information from multiple source domains typically enhances classification precision. In light of this, we introduce a new Multi-source Distributed Transfer Learning (MDTL) framework designed to adeptly integrate complementary information across various source domains through the application of belief functions. In this approach, the distributions of each source and target domain are aligned independently. Subsequently, the resultant soft classification outcomes, facilitated by different source domains, are amalgamated using belief functions. This integration incorporates novel weighting factors that consider both distribution discrepancies and classifier effectiveness. The effectiveness of MDTL was assessed against a range of related methods, and the experimental findings confirm that it markedly improves classification accuracy in the target domain. Linqing Huang, Yumei Hu, Shi-Lin Wang, Gongshen Liu, Jinfu Fan |
ICIP | 5 |
| 2025 | KMoP: Knowledge-injected Mixture-of-Prefix for Joint Multimodal Aspect-Based Sentiment AnalysisabstractMultimodal Aspect-based Sentiment Analysis (MABSA) aims to extract aspect-sentiment pairs from a combination of text and images. However, images often contain content that is either irrelevant to the textual information or not related to the sentiment prediction, which can adversely affect the accuracy of the model predictions. Furthermore, existing models neglect precise regional information beyond global image features, which could also assist in enhancing aspect-based sentiment prediction, but may also introduce noise that damages the model. To address these issues, this study proposes Knowledge-injected Mixture-of-Prefix (KMoP) to inject various types of knowledge into the language model and reduce external noise. Specifically, external knowledge is injected in the form of prefixes into the language model, which minimizes catastrophic forgetting issue and generate noise-insensitive representations. Additionally, to allow different layers of the language model to automatically select the required knowledge, we differentiate the aggregation of prefixes from different knowledge sources for each layer through Mixture-of-Prefix. This paper simultaneously divides the training process into two parts, with the first phase training on the original clean dataset and the second phase fine-tuning on the original dataset with added noise. KMoP achieves state-of-the-art performance on MABSA task, with extensive supplementary experiments demonstrating its enhanced robustness to noise. Xinzhong Wang, Lingyong Fang, Jidong Li, Gongshen Liu |
ICME | 5 |
| 2025 | Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language ModelsabstractMulti-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently memorize private content that is entirely irrelevant to the training tasks. In this paper, we investigate how randomly generated task-irrelevant private content can become spuriously correlated with downstream objectives due to partial mini-batch training dynamics, thus causing inadvertent memorization. Concretely, we randomly generate task-irrelevant watermarks into VQA fine-tuning images at varying probabilities and propose a novel probing framework to determine whether MLLMs have inadvertently encoded such content. Our experiments reveal that MLLMs exhibit notably different training behaviors in partial mini-batch settings with task-irrelevant watermarks embedded. Furthermore, through layer-wise probing, we demonstrate that MLLMs trigger distinct representational patterns when encountering previously seen task-irrelevant knowledge, even if this knowledge does not influence their output during prompting. Our code is available at https://github.com/illusionhi/ProbingPrivacy. Tianjie Ju, Hao Fei 0001, Zhenyu Shao, Yubin Zheng, Haodong Zhao, Mong-Li Lee, Wynne Hsu, Zhuosheng Zhang 0001, Gongshen Liu |
ICML | 10 |
| 2025 | GTRNet: Graph Topology-aware Refinement Network for User Role Classification in Social NetworksabstractWith the rapid development of Internet technology, social networks have become essential platforms for communication, and in the process, they generate vast amounts of user data that reflect interests, relationships, and habits. User role classification in social networks is critical for personalized recommendations, precision marketing, and security. In this work, we investigate Graph Neural Networks (GNNs) for user role classification in social networks, and a novel GNN architecture, termed Graph Topology-aware Refinement Network (GTRNet), is designed. GTRNet comprises two modules: Graph Topology Encoding (GTE) and Node Representation Refinement (NRR). They combine the node features and network topological information via convolutional layers to enhance the use role classification performance. The experimental results demonstrate that GTRNet usually outperforms the state-of-the-art methods on the Facebook, Cora, and Citeseer datasets, with micro-F1 score improvements of 0.021, 0.009, and 0.014, respectively. It verifies GTRNet’s effectiveness in addressing social network user role classification task. Tingxuan Gu, Linqing Huang, Gongshen Liu |
SMC | 3 |
| 2025 | Improving semi-autoregressive machine translation with the guidance of syntactic dependency parsing structure
Sufeng Duan, Gongshen Liu |
Neurocomputing | 3 |
| 2025 | Transferable and Robust Dynamic Adversarial Attack Against Object Detection ModelsabstractObject detection models have been widely deployed in physical world applications, and they are vulnerable to adversarial attacks. However, most adversarial attacks are implemented in a glass box setting, and under ideal shooting conditions, such as fixed distances and angles, and thus have limited attack success rate (ASR) in practice. In this article, we present a transferable and robust dynamic adversarial attack where the adversarial patches can be printed on or attached to nonrigid objects, such as clothes. We develop a cascade module with a momentum-based technique to optimize adversarial patches against various object detection models, achieving better transferability of the patches in a closed box setting. We also develop a strategy of distance-adaptive patch generation and employ perspective transformation to enhance the robustness of patches. To evaluate the attack performance, we conduct extensive experiments on seven mainstream object detection models at different distances and angles. The results show that our method can achieve an average ASR of 69.85%, which is 3.27 times that of the baseline method at 3 m. Jing Chen 0003, Zijun Zhang 0003, Kun He 0008, Zongru Wu, Ruiying Du, Gongshen Liu |
IEEE Internet Things J. | 8 |
| 2025 | Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security ReviewabstractLanguage models (LMs) are becoming increasingly popular in real-world applications. Outsourcing model training and data hosting to third-party platforms has become a standard method for reducing costs. In such a situation, the attacker can manipulate the training process or data to inject a backdoor into models. Backdoor attacks are a serious threat where malicious behavior is activated when triggers are present; otherwise, the model operates normally. However, there is still no systematic and comprehensive review of LMs from the attacker's capabilities and purposes on different backdoor attack surfaces. Moreover, there is a shortage of analysis and comparison of the diverse emerging backdoor countermeasures. Therefore, this work aims to provide the natural language processing (NLP) community with a timely review of backdoor attacks and countermeasures. According to the attackers' capability and affected stage of the LMs, the attack surfaces are formalized into four categorizations: attacking the pretrained model with fine-tuning (APMF) or parameter-efficient fine-tuning (PEFT), attacking the final model with training (AFMT), and attacking large language model (ALLM). Thus, attacks under each categorization are combed. The countermeasures are categorized into two general classes: sample inspection and model inspection. Thus, we review countermeasures and analyze their advantages and disadvantages. Also, we summarize the benchmark datasets and provide comparable evaluations for representative attacks and defenses. Drawing the insights from the review, we point out the crucial areas for future research on the backdoor, especially soliciting more efficient and practical countermeasures. Pengzhou Cheng, Zongru Wu, Haodong Zhao, Wei Lu 0011, Gongshen Liu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | FedSTS: A Stratified Client Selection Framework for Consistently Fast Federated LearningabstractIn this article, we investigate random client selection in the context of horizontal federated learning (FL), whereby only a randomly selected subset of clients transmit their model updates to the server instead of yielding all clients involved. Many researchers have demonstrated that clustering-based client selection constitutes a simple yet efficacious approach to the identification of those clients possessing representative gradient information. Despite the extensive body of research on modified selection methodologies, the majority of prior work is predicated upon the assumption of consistently effective clustering. However, raw gradient-based clustering methods are subject to several challenges: 1) poor effectiveness, the raw high-dimensional gradient of a client is too complex to serve as an appropriate feature for grouping, resulting in large intra-cluster distances and 2) fluctuating effectiveness, due to inherent limitations in clustering, the effectiveness can vary significantly, leading to clusters with diverse levels of heterogeneity. In practice, suboptimal and inconsistent clustering effects can result in clusters with low intra-cluster similarity among clients. The selection of clients from such clusters may impede the overall convergence of training. In this article, we propose FedSTS, a novel client selection scheme to accelerate the FL convergence by variance reduction. The main idea of FedSTS is to stratify a compressed model update in order to ensure an excellent grouping effect, and at the same time reduce the cross-client variance by re-allocating the sample chance among different groups based on their diverse heterogeneity. It strikes this convergence acceleration by paying more attention to those client groups with relatively low similarity and then improving the representativeness of the selected subset as much as possible. Theoretically, we demonstrate the critical improvement of the proposed scheme in variance reduction and present equivalence conditions among different client selection methods. We also present the tighter convergence guarantee of the proposed method thanks to the variance reduction. Experimental results confirm the exceeded efficiency of our approach compared to alternatives. Dehong Gao, Duanxiao Song, Guangyuan Shen, Xiaoyan Cai, Libin Yang, Gongshen Liu, Zhen Wang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language ModelsabstractTianjie Ju, Yijin Chen, Xinwei Yuan, Zhuosheng Zhang, Wei Du, Yubin Zheng, Gongshen Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tianjie Ju, Xinwei Yuan, Zhuosheng Zhang 0001, Yubin Zheng, Gongshen Liu |
ACL (1) | 7 |
| 2024 | Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency SpaceabstractDespite the notable success of language models (LMs) in various natural language processing (NLP) tasks, the reliability of LMs is susceptible to backdoor attacks.Prior research attempts to mitigate backdoor learning while training the LMs on the poisoned dataset, yet struggles against complex backdoor attacks in real-world scenarios.In this paper, we investigate the learning mechanisms of backdoor LMs in the frequency space by Fourier analysis.Our findings indicate that the backdoor mapping presented on the poisoned datasets exhibits a more discernible inclination towards lower frequency compared to clean mapping, resulting in the faster convergence of backdoor mapping.To alleviate this dilemma, we propose Multi-Scale Low-Rank Adaptation (MuScleLoRA), which deploys multiple radial scalings in the frequency space with lowrank adaptation to the target model and further aligns the gradients when updating parameters.Through downscaling in the frequency space, MuScleLoRA encourages the model to prioritize the learning of relatively highfrequency clean mapping, consequently mitigating backdoor learning.Experimental results demonstrate that MuScleLoRA outperforms baselines significantly.Notably, MuScleLoRA reduces the average success rate of diverse backdoor attacks to below 15% across multiple datasets and generalizes to various backbone LMs, including BERT, RoBERTa, GPT2-XL, and Llama2.The codes are publicly available at https://github.com/ZrW00/MuScleLoRA. Zongru Wu, Zhuosheng Zhang 0001, Pengzhou Cheng, Gongshen Liu |
ACL (1) | 4 |
| 2024 | Backdoor NLP Models via AI-Generated TextabstractBackdoor attacks pose a critical security threat to natural language processing (NLP) models by establishing covert associations between trigger patterns and target labels without affecting normal accuracy. Existing attacks usually disregard fluency and semantic fidelity of poisoned text, rendering the malicious data easily detectable. However, text generation models can produce coherent and content-relevant text given prompts. Moreover, potential differences between human-written and AI-generated text may be captured by NLP models while being imperceptible to humans. More insidious threats could arise if attackers leverage latent features of AI-generated text as trigger patterns. We comprehensively investigate backdoor attacks on NLP models using AI-generated poisoned text obtained via continued writing or paraphrasing, exploring three attack scenarios: data, model and pre-training. For data poisoning, we fine-tune generators with attribute control to enhance the attack performance. For model poisoning, we leverage downstream tasks to derive specialized generators. For pre-training poisoning, we train multiple attribute-based generators and align their generated text with pre-defined vectors, enabling task-agnostic migration attacks. Experiments demonstrate that our method achieves effective attacks while maintaining fluency and semantic similarity across all scenarios. We hope this work can raise awareness of the security risks hidden in AI-generated text. Tianjie Ju, Gaolei Li, Gongshen Liu |
LREC/COLING | 5 |
| 2024 | How Large Language Models Encode Context Knowledge? A Layer-Wise Probing StudyabstractPrevious work has showcased the intriguing capability of large language models (LLMs) in retrieving facts and processing context knowledge. However, only limited research exists on the layer-wise capability of LLMs to encode knowledge, which challenges our understanding of their internal mechanisms. In this paper, we devote the first attempt to investigate the layer-wise capability of LLMs through probing tasks. We leverage the powerful generative capability of ChatGPT to construct probing datasets, providing diverse and coherent evidence corresponding to various facts. We employ \mathcal V-usable information as the validation metric to better reflect the capability in encoding context knowledge across different layers. Our experiments on conflicting and newly acquired knowledge show that LLMs: (1) prefer to encode more context knowledge in the upper layers; (2) primarily encode context knowledge within knowledge-related entity tokens at lower layers while progressively expanding more knowledge within other tokens at upper layers; and (3) gradually forget the earlier context knowledge retained within the intermediate layers when provided with irrelevant evidence. Code is publicly available at https://github.com/Jometeorie/probing_llama. Tianjie Ju, Weiwei Sun 0001, Xinwei Yuan, Zhaochun Ren, Gongshen Liu |
LREC/COLING | 6 |
| 2024 | FashionGPT: A Large Vision-Language Model for Enhancing Fashion Understanding
Duanxiao Song, Dehong Gao, Gongshen Liu |
ICANN (5) | 3 |
| 2024 | NWS: Natural Textual Backdoor Attacks Via Word SubstitutionabstractBackdoor attacks pose a serious security threat for natural language processing (NLP). Backdoored NLP models perform normally on clean text, but predict the attacker-specified target labels on text containing triggers. Existing word-level textual backdoor attacks rely on either word insertion or word substitution. Word-insertion backdoor attacks can be easily detected by simple backdoor defenses. Meanwhile, word-substitution backdoor attacks tend to substantially degrade the fluency and semantic consistency of the poisoned text. In this paper, we propose a more natural word substitution method to implement covert textual backdoor attacks. Specifically, we combine three different ways to construct a diverse synonym thesaurus for clean text. We then train a learnable word selector for producing poisoned text using a composite loss function of poison and fidelity terms. This enables automated selection of minimal critical word substitutions necessary to induce the backdoor. Experiments demonstrate our method achieves high attack performance with less impact on fluency and semantics. We hope this work can raise awareness regarding the threat of subtle, fluent word substitution attacks. Tongxin Yuan, Haodong Zhao, Gongshen Liu |
ICASSP | 4 |
| 2024 | Multi-Grained Multimodal Interaction Network for Sentiment AnalysisabstractMultimodal sentiment analysis aims to utilize different modalities including language, visual, and audio to identify human emotions in videos. Multimodal interaciton mechanism is the key challenge. Previous works lack modeling of multimodal interaction at different grain levels, and does not suppress redundant information in multimodal interaction. This leads to incomplete multimodal representation with noisy information. To address these issues, we propose Multi-grained Multimodal Interaction Network (MMIN) to provide a more complete view of multimodal representation. Coarse-grained Interaction Network (CIN) exploits the unique characteristics of different modalities at a coarse-grained level and adversarial learning is used to reduce redundancy. Fine-grained Interaction Network (FIN) employ sparse-attention mechanism to capture fine-grained interactions between multimodal sequences across distinct time steps and reduce irrelevant fine-grained multimodal interaction. Experimental results on two public datasets demonstrate the effectiveness of our model in multimodal sentiment analysis. Lingyong Fang, Gongshen Liu, Ru Zhang 0002 |
ICASSP | 2 |
| 2024 | Improving Non-autoregressive Machine Translation with Error Exposure and Consistency Regularization
Sufeng Duan, Gongshen Liu |
NLPCC (3) | 3 |
| 2024 | Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference List
Yun Lin 0001, Xiwen Teoh, Gongshen Liu, Zhiyong Huang 0010, Jin Song Dong 0001 |
USENIX Security Symposium | 4 |
| 2024 | LSF-IDM: Deep learning-based lightweight semantic fusion intrusion detection model for automotive
Pengzhou Cheng, Haobin Jiang, Gongshen Liu |
Peer Peer Netw. Appl. | 4 |
| 2023 | PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language ModelsabstractThe huge training overhead, considerable commercial value, and various potential security risks make it urgent to protect the intellectual property (IP) of Deep Neural Networks (DNNs). DNN watermarking has become a plausible method to meet this need. However, most of the existing watermarking schemes focus on image classification tasks. The schemes designed for the textual domain lack security and reliability. Moreover, how to protect the IP of widely-used pre-trained language models (PLMs) remains a blank. To fill these gaps, we propose PLMmark, the first secure and robust black-box watermarking framework for PLMs. It consists of three phases: (1) In order to generate watermarks that contain owners’ identity information, we propose a novel encoding method to establish a strong link between a digital signature and trigger words by leveraging the original vocabulary tables of PLMs. Combining this with public key cryptography ensures the security of our scheme. (2) To embed robust, task-agnostic, and highly transferable watermarks in PLMs, we introduce a supervised contrastive loss to deviate the output representations of trigger sets from that of clean samples. In this way, the watermarked models will respond to the trigger sets anomaly and thus can identify the ownership. (3) To make the model ownership verification results reliable, we perform double verification, which guarantees the unforgeability of ownership. Extensive experiments on text classification tasks demonstrate that the embedded watermark can transfer to all the downstream tasks and can be effectively extracted and verified. The watermarking scheme is robust to watermark removing attacks (fine-pruning and re-initializing) and is secure enough to resist forgery attacks. Pengzhou Cheng, Fangqi Li 0001, Haodong Zhao, Gongshen Liu |
AAAI | 6 |
| 2023 | Semantic Information Mining and Fusion Method for Bot Detection
Lijia Liang, Xinzhong Wang, Gongshen Liu |
ICANN (9) | 3 |
| 2023 | FedPrompt: Communication-Efficient and Privacy-Preserving Prompt Tuning in Federated LearningabstractFederated learning (FL) has enabled global model training on decentralized data in a privacy-preserving way. However, for tasks that utilize pre-trained language models (PLMs) with massive parameters, there are considerable communication costs. Prompt tuning, which tunes soft prompts without modifying PLMs, has achieved excellent performance as a new learning paradigm. In this paper, we want to combine these methods and explore the effect of prompt tuning under FL. We propose "FedPrompt" studying prompt tuning in a model split aggregation way using FL, and prove that split aggregation greatly reduces the communication cost, only 0.01% of the PLMs’ parameters, with little decrease on accuracy both on IID and Non-IID data distribution. We further conduct backdoor attacks by data poisoning on FedPrompt. Experiments show that attack achieve a quite low attack success rate and can not inject backdoor effectively, proving the robustness of FedPrompt. Haodong Zhao, Fangqi Li 0001, Gongshen Liu |
ICASSP | 5 |
| 2023 | SDPSAT: Syntactic Dependency Parsing Structure-Guided Semi-Autoregressive Machine Translation
Yuran Zhao, Jianming Guo, Sufeng Duan, Gongshen Liu |
ICONIP (8) | 5 |
| 2023 | Neural Linguistic Steganography with Controllable SecurityabstractInformation hiding is an art and science with a long history and is widely used in covert communication. There are many ways to hide secret data in image, audio, and video. However, relatively few systems can hide information in text. Generative text steganography is a promising topic in natural language text infor-mation hiding. Previous generative text steganography methods use a fixed candidate pool generation rule, and they cannot effec-tively control the security of the generated text. The perceptual-imperceptibility and statistical-imperceptibility conflict effect also causes the poor quality of the steganographic text generated by previous generative text steganography methods. Moreover, pre-vious generative text steganography approaches barely discuss the robustness of steganographic text. This paper proposes a security controllable text steganography method that can generate natural-looking steganographic text with a statistical distribution that matches the natural language distribution. The proposed method combines the metrics of per-ceptual-imperceptibility and statistical-imperceptibility to calcu-late the combined distortion. It selects the tokens with the smallest combined distortion to construct a candidate pool at each time step. Moreover, the maximum combined distortion threshold is set when embedding secret messages to ensure controllable security. We conducted several experiments to evaluate the proposed model from the perspectives of embedding rate, perceptual-impercepti-bility, statistical-imperceptibility, and anti-attack ability. The ex-perimental results show that the proposed method can generate smooth and readable steganographic sentences with good re-sistance to steganalysis and high robustness. Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju |
IJCNN | 2 |
| 2023 | Robust Secret Data Hiding for Transformer-based Neural Machine TranslationabstractHiding secret information in text is a research area of significant importance and a great challenge. In recent years, there have been huge developments and exciting advances in generation-based text information hiding techniques. Current generative text information hiding methods mainly establish correspondence between token and secret bits based on probability distributions given by language models. However, the semantic control of such methods is weak, and their robustness is not discussed. In this paper, we investigate an end-to-end generation-based text information hiding scheme. The proposed method uses a sequence-to-sequence model with adversarial training as a machine translation model. It converts the secret information into an embedding vector to be added to each position of the hidden state representation of the source language text, which in turn allows the model to automatically learn to produce translation results with the embedded secret information without using fixed rules. The semantics of the text with embedded secret messages obtained by translation can be controlled by the meaning of the source language text. Our experiments show that the proposed method can embed the secret message into the translation results with little loss of the translation quality and is robust to active attacks such as word deletion or synonym substitution. Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju |
IJCNN | 2 |
| 2023 | SCS-VAE: Generate Style Headlines Via Novel DisentanglementabstractCurrent headline generation models only focus on consistent headlines while lacking attention to headline styles. However, headlines of different styles could have different effects on producers and viewers. Capturing salient content while following a unique style is challenging when generating text. In this paper, we propose a Semantic Content-Style VAE(SCS-VAE), combining the Variational Auto-Encoder(VAE) and the dictionary learning method to solve the content-related and stylized problems of generating headlines simultaneously. Specifically, we disentangle the semantic information in content and style space, allowing us to control the generation of headline content and styles. Moreover, we design interpretable loss functions better to supervise the semantic information into two different feature spaces. Experiment results show that SCS-VAE has SOTA performance in headline generation method and makes the style of the headlines more diversified. Zhaoqian Zhu, Gongshen Liu, Bo Su 0003, Tianhe Lu |
IJCNN | 3 |
| 2023 | DESC-IDS: Towards an efficient real-time automotive intrusion detection system based on deep evolving stream clustering
Pengzhou Cheng, Mu Han, Gongshen Liu |
Future Gener. Comput. Syst. | 3 |
| 2022 | Few-shot Table-to-text Generation with Prefix-Controlled GeneratorabstractNeural table-to-text generation approaches are data-hungry, limiting their adaption for low-resource real-world applications. Previous works mostly resort to Pre-trained Language Models (PLMs) to generate fluent summaries of a table. However, they often contain hallucinated contents due to the uncontrolled nature of PLMs. Moreover, the topological differences between tables and sequences are rarely studied. Last but not least, fine-tuning on PLMs with a handful of instances may lead to over-fitting and catastrophic forgetting. To alleviate these problems, we propose a prompt-based approach, Prefix-Controlled Generator (i.e., PCG), for few-shot table-to-text generation. We prepend a task-specific prefix for a PLM to make the table structure better fit the pre-trained input. In addition, we generate an input-specific prefix to control the factual contents and word order of the generated text. Both automatic and human evaluations on different domains (humans, books and songs) of the Wikibio dataset prove the effectiveness of our approach. Yutao Luo, Menghua Lu, Gongshen Liu, Shi-Lin Wang |
COLING | 3 |
| 2022 | A Multi-Task Dual-Tree Network for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims at extracting triplets from a given sentence, where each triplet includes an aspect, its sentiment polarity, and a corresponding opinion explaining the polarity. Existing methods are poor at detecting complicated relations between aspects and opinions as well as classifying multiple sentiment polarities in a sentence. Detecting unclear boundaries of multi-word aspects and opinions is also a challenge. In this paper, we propose a Multi-Task Dual-Tree Network (MTDTN) to address these issues. We employ a constituency tree and a modified dependency tree in two sub-tasks of Aspect Opinion Co-Extraction (AOCE) and ASTE, respectively. To enhance the information interaction between the two sub-tasks, we further design a Transition-Based Inference Strategy (TBIS) that transfers the boundary information from tags of AOCE to ASTE through a transition matrix. Extensive experiments are conducted on four popular datasets, and the results show the effectiveness of our model. Yichun Zhao, Gongshen Liu, Jintao Du, Huijia Zhu |
COLING | 3 |
| 2022 | Parallel Relationship Graph to Improve Multi-Document Summarization
Menghua Lu, Lijia Liang, Gongshen Liu |
ICANN (2) | 3 |
| 2022 | PPT: Backdoor Attacks on Pre-trained Models via Poisoned Prompt TuningabstractRecently, prompt tuning has shown remarkable performance as a new learning paradigm, which freezes pre-trained language models (PLMs) and only tunes some soft prompts. A fixed PLM only needs to be loaded with different prompts to adapt different downstream tasks. However, the prompts associated with PLMs may be added with some malicious behaviors, such as backdoors. The victim model will be implanted with a backdoor by using the poisoned prompt. In this paper, we propose to obtain the poisoned prompt for PLMs and corresponding downstream tasks by prompt tuning. We name this Poisoned Prompt Tuning method "PPT". The poisoned prompt can lead a shortcut between the specific trigger word and the target label word to be created for the PLM. So the attacker can simply manipulate the prediction of the entire model by just a small prompt. Our experiments on various text classification tasks show that PPT can achieve a 99% attack success rate with almost no accuracy sacrificed on original task. We hope this work can raise the awareness of the possible security threats hidden in the prompt. Yichun Zhao, Boqun Li, Gongshen Liu, Shi-Lin Wang |
IJCAI | 4 |
| 2022 | Sense-aware BERT and Multi-task Fine-tuning for Multimodal Sentiment AnalysisabstractHumans convey emotions through verbal and non-verbal signals when communicating face-to-face. Pre-trained language model such as BERT can be fine-tuned to improve the performance of various downstream tasks including sentiment analysis. However, most prior works about BERT fine-tuning contains only textual unimodal data and lacks information from sense organs, such as audio and visual signals, which are crucial for sentiment analysis. In this paper, we propose Sense-aware BERT (SenBERT) which allows sense information integrated with BERT during fine-tuning. In particular, we exploit multimodal multi-head attention to capture the interaction between unaligned multimodal data. Additionally, due to the variable information richness of different modalities, multimodal network may be dominated by some modalities during training process, so we propose unimodal sentiment analysis auxiliary tasks for multi-task learning which forces the model to focus on all modalities. We conduct experiments on CMU-MOSI and CMU-MOSEI datasets for multimodal sentiment analysis. The results show the superior performance of SenBERT on all the metrics over previous baselines. Lingyong Fang, Gongshen Liu, Ru Zhang 0002 |
IJCNN | 2 |
| 2022 | Context Modeling with Hierarchical Shallow Attention Structure for Document-Level NMTabstractIt is acknowledged that neural machine translation(NMT) can be improved by considering context information. Nevertheless, the progress of context-aware NMT has encountered some challenges. Firstly, effectively utilizing valuable information contained in context is still challenging. Moreover, as the number of sentences increases, the parameters of context-aware NMT models will surge, which costs computing power and prevents them from transferring to other translation tasks. Therefore, we propose a hierarchical shallow attention structure for document-level NMT to tackle the problems above. We employ hierarchical encoders to extract both sentence and context information. Then the integration of hierarchical attention is incorporated with the self-attention of the target sentences in the decoding phase. Moreover, we employ shallow attention to reduce model complexity. Assessments of several linguistic phenomena demonstrate that the proposed approach can balance model complexity and translation performance while getting SOTA BLEU scores in several translation tasks. Jianming Guo, Weijie Yuan 0003, Jianshen Zhang, Gongshen Liu |
IJCNN | 6 |
| 2022 | TIMS: A Novel Approach for Incrementally Few-Shot Text Instance Selection via Model SimilarityabstractLarge-scale pre-trained models' demand for high-quality instances forces people to consider how to select instances for annotation with limited resources. Nonetheless, little attention has been paid to the scenario where the number of instances that ultimately need to be annotated is agnostic. Meanwhile, the anisotropy of the sentence vector output by pre-trained models makes it hard to represent the instance itself well. Faced with the two challenges, we propose an incrementally few-shot instance selection approach (TIMS) based on model similarity and outlier detection, which suits the starting step of active learning well and serves as a better benchmark for few-shot learning. Specifically, TIMS determines the representative candidate set by calculating the similarity between changes in model parameters caused by each instance and by the full dataset. Meanwhile, Isolation Forest is adopted to select instances from the candidate set for annotation, which prevents selected instances from being too similar. Comprehensive experiments on WikiLingua & SQuAD show that TIMS outperforms other algorithms across almost every circumstance. It inspires us that the proper implementation of model similarity detection and outlier detection is of great help to select representative instances incrementally. Tianjie Ju, Han Liao, Gongshen Liu |
IJCNN | 3 |
| 2022 | KLAttack: Towards Adversarial Attack and Defense on Neural Dependency Parsing ModelsabstractAlthough neural language models achieve great performance on many Natural Language Processing tasks, they suffer from various adversarial attacks. Previous works mainly focus on semantic adversarial examples, which have similar semantics to the original sentences, while syntactic adversarial attacks against the dependency parsing task are still in an early stage of research. In this paper, we propose a novel method KLAttack, crafting word-level adversarial examples to attack neural-network-based dependency parsing models. Specifically, we retrieve the class probabilities from the victim dependency parsing model and compute the KL divergence by masking every word in a sentence. Then we use pre-trained language models and reference parsers to generate candidates for substitution. Experiments on the English Penn Treebank (PTB) dataset show that our method improves the attack success rate against Deep Biaffine Parser by up to 13.04% compared with previous related studies. Based on KLAttack, we further propose Syntax-Aware Transformer for Input Reconstruction, a denoiser to recover the original sentences from the adversarial examples. Trained adversarially with successfully attacked sentences from KLAttack, we enhance the robustness of the dependency parsing models by concatenating the denoiser ahead of the victim models. Yutao Luo, Menghua Lu, Chaoqi Yang, Gongshen Liu, Shi-Lin Wang |
IJCNN | 4 |
| 2022 | SECT: A Successively Conditional Transformer for Controllable Paraphrase GenerationabstractParaphrase generation has consistently been a challenging area in the field of NLP. Despite the considerable achievements made by previous work, existing methods lack a flexible way to include multiple controllable attributes to enhance the diversity of paraphrased sentences. To overcome this challenge, we propose a Successively Conditional Transformer (SECT) to tackle this task. SECT is based on a combination of Conditional Variational AutoEncoder (CVAE) and Transformer framework to generate diversified words. More specifically, our SECT deploys multi-head attention and memory gate mechanism to keep the interaction between each of the attributes and the corresponding encoder layer hidden state. To address the problem of absorbing flexible attributes, we apply a successive structure to our SECT, which enables the framework to couple the CVAE latent variables with the encoder layer hidden states progressively. In addition, our SECT is trained by minimizing a tailor-designed loss for producing paraphrased sentences as required. Finally, we conduct extensive experiments to substantiate the validity and effectiveness of our proposed model. The results show that SECT significantly outperforms the existing state-of-the-art approaches and generates more diverse paraphrased sentences. Tang Xue, Yuran Zhao, Chaoqi Yang, Gongshen Liu |
IJCNN | 4 |
| 2022 | Label-Dividing Gated Graph Neural Network for Hierarchical Text ClassificationabstractMulti-label hierarchical text classification (MLHTC) is an essential yet challenging task of natural language processing (NLP). Existing methods lack attention to predicting sibling labels. In addition, methods based on graph convolutional networks (GCN) meet the problem of over-smoothing, which further deepens the difficulty of distinguishing sibling labels. In this paper, we propose a label-dividing gated graph neural network (LD-GGNN), which can better distinguish sibling labels and achieve adaptive interaction between text and labels. We optimize gated graph neural network (GGNN) to accurately capture structural features of label hierarchy and deeply explore label dependence. Stronger nonlinear characteristics of GGNN are used to solve the problem of over-smoothing. Furthermore, we propose a dynamic label dividing mechanism (DLDM), which can guide the model to distinguish sibling labels by introducing a dividing bias. Compared with previous works, LD-GGNN achieves significant and consistent improvements on both Micro-F1 and Macro-F1 score on multiple datasets. Jie Zhou 0013, Gongshen Liu |
IJCNN | 4 |
| 2022 | A Universal Identity Backdoor Attack against Speaker Verification based on Siamese NetworkabstractSpeaker verification has been widely used in many authentication scenarios.However, training models for speaker verification requires large amounts of data and computing power, so users often use untrustworthy third-party data or deploy thirdparty models directly, which may create security risks.In this paper, we propose a backdoor attack for the above scenario.Specifically, for the Siamese network in the speaker verification system, we try to implant a universal identity in the model that can simulate any enrolled speaker and pass the verification.So the attacker does not need to know the victim, which makes the attack more flexible and stealthy.In addition, we design and compare three ways of selecting attacker utterances and two ways of poisoned training for the GE2E loss function in different scenarios.The results on the TIMIT and Voxceleb1 datasets show that our approach can achieve a high attack success rate while guaranteeing the normal verification accuracy.Our work reveals the vulnerability of the speaker verification system and provides a new perspective to further improve the robustness of the system. Haodong Zhao, Junjie Guo, Gongshen Liu |
INTERSPEECH | 4 |
| 2022 | Improving Constituent Representation with Hypertree Neural NetworksabstractMany natural language processing tasks involve text spans and thus high-quality span representations are needed to enhance neural approaches to these tasks.Most existing methods of span representation are based on simple derivations (such as max-pooling) from word representations and do not utilize compositional structures of natural language.In this paper, we aim to improve representations of constituent spans using a novel hypertree neural networks (HTNN) that is structured with constituency parse trees.Each node in the HTNN represents a constituent of the input sentence and each hyperedge represents a composition of smaller child constituents into a larger parent constituent.In each update iteration of the HTNN, the representation of each constituent is computed based on all the hyperedges connected to it, thus incorporating both bottom-up and top-down compositional information.We conduct comprehensive experiments to evaluate HTNNs against other span representation models and the results show the effectiveness of HTNN. Hao Zhou 0044, Gongshen Liu, Kewei Tu |
NAACL-HLT | 2 |
| 2022 | An Adversarial Approach for Unsupervised Syntax-Guided Paraphrase Generation
Tang Xue, Yuran Zhao, Gongshen Liu |
NLPCC (1) | 3 |
| 2022 | Multi-Objective Actor-Critics for Real-Time Bidding in Display Advertising
Haolin Zhou, Chaoqi Yang, Xiaofeng Gao 0001, Gongshen Liu, Guihai Chen |
ECML/PKDD (4) | 5 |
| 2021 | Dynamic Tuning and Weighting of Meta-learning for NMT Domain Adaptation
Ziyue Song, Kaiyue Qi, Gongshen Liu |
ICANN (5) | 4 |
| 2021 | Speaker Verification with Disentangled Self-attention
Junjie Guo, Haodong Zhao, Gongshen Liu |
ICONIP (1) | 4 |
| 2021 | A Hierarchical Graph-Based Neural Network for Malware Classification
Yuran Zhao, Gongshen Liu, Bo Su 0003 |
ICONIP (4) | 3 |
| 2021 | A Multi-Channel Graph Attention Network for Chinese NER
Yichun Zhao, Gongshen Liu |
ICONIP (1) | 3 |
| 2021 | A Multi -Role Graph Attention Network for Knowledge Graph AlignmentabstractEntity alignment aims to automatically align entities which are equivalent in the real world from different knowledge graphs (KGs). The task is challenged by the diversity of KGs, the multi-relational graph structures and limited seed entity pairs. Most existing GNN-based models do not distinguish the types or the directions of relations while the existing translation-based models merely rely on local semantics. To tackle the challenges, we propose a Multi-Role Graph Attention network (MRGA), which incorporates local semantics in triples with global structural information in graphs and exploits rich relational information effectively. MRGA devises a GAT-based layer which regards each entity as a combination of roles that it plays in different triples. Multifaceted entity semantics (i.e. roles), associated with different triples, are aggregated to enrich entity embeddings. An additional attention mechanism is applied to capture the correlation in relations of different KGs. The model is evaluated on three publicly available datasets. The results show the superior performance of MRGA compared with baseline methods and demonstrate the effectiveness of each component in MRGA via ablation study. Linyi Ding, Weijie Yuan 0003, Gongshen Liu |
IJCNN | 4 |
| 2021 | Bridging the Gap of Dimensions in Distillation: Understanding the knowledge transfer between different-dimensional semantic spacesabstractIn recent years, knowledge distillation has been widely used in the field of deep learning in order to reduce the model size and save time and space. The student-teacher paradigm is a framework for knowledge distillation, and knowledge distillation proposed to minimize the KL divergence between the probabilistic outputs of a teacher and student network. However, apart from the probabilistic outputs, there are much valuable information contained in the middle layers of the teacher network. As for NLP tasks, the hidden vectors from different layers of a model have different semantic information, but the vectors' dimension of the student network is different from that of the teacher network in many cases, which makes hidden layer distillation hard to be performed directly. We propose to simply use a transition matrix to project the student's vector to a space of the same dimension as the teacher's vector, and we theoretically prove the effectiveness of this method. Our analysis shows how the transition matrix preserve important semantic information, which is closely related to the vector's characteristic in Euclidean space. We provide a geometric method for the interpretability of shared knowledge space for student-teacher architectures. Our experiments show that this method can significantly improve the performance of a small model in different tasks with different models. Ziyue Song, Haodong Zhao, Gongshen Liu |
IJCNN | 5 |
| 2021 | ASM: Augmentation-based Semantic Mechanism on Abstractive SummarizationabstractMany transformer-based encoder-decoder models have made significant progress on summary generating tasks. And the availability of pre-trained models further improves its performance with self-supervised objectives on large text corpora. However, most models' architectures and their training criteria pay more attention to the lexical and syntactic structure rather than semantic similarity. In this paper, we augment training data in semantic space and propose Augmentation-based Semantic Mechanism (ASM) for encoder feedback with corresponding criterion to capture global semantic meanings. Notably, we enhance the encoder's comprehension of summaries in semantic space, and facilitate the integration of global semantics and local syntax during generating summaries. By leveraging pre-trained language models, we have driven our results to a new level (45.11 on CNN/DailyMail, 45.35 on XSum in ROUGE-1). Additionally, the human evaluation and further experiments also validate the effectiveness of our proposed method for generating abstractive summaries. Our augmented data and source code for summarization will be made public. Weidong Ren, Hao Zhou 0044, Gongshen Liu, Fei Huan |
IJCNN | 3 |
| 2021 | Text Generation with Syntax - Enhanced Variational AutoencoderabstractText generation is one of the essential yet challenging tasks in natural language processing. However, the input text alone is usually hard to provide enough information to generate the desired output. Previous work attempts to incorporate syntactic information into the generative models based on variational autoencoder(VAE). But these methods have difficulty in adequately modeling the tree structure of syntactic data. In this paper, we formulate the syntactic structure as a graph and introduce a syntax encoder based on graph neural network(GNN) to model the syntactic information of sentences. Based on the syntax encoder, we propose a novel syntax-enhanced variational autoencoder(SEVAE) with two variants. The variant SEVAE-m merges sentence information and syntactic information into one latent space to enrich the fine-grained syntactic information of latent representations. And the variant SEVAE-s with two separate latent spaces allows the sentence decoder to dynamically attend to semantic and syntactic information from two latent variables. Experiments on two benchmark datasets show that our methods achieve significant and consistent improvements compared with previous work. Weijie Yuan 0003, Linyi Ding, Gongshen Liu |
IJCNN | 4 |
| 2021 | Large-Scale Malicious Software Classification With Fuzzified Features and Boosted Fuzzy Random ForestabstractClassification of malicious software, especially in a very large dataset, is a challenging task for machine intelligence. Malware can have highly diversified features, each of which has highly heterogeneous distributions. These factors increase the difficulties for traditional data analytic approaches to deal with them. Although deep learning based methods have reported good classification performance, the deep models usually lack interpretability and are fragile under adversarial attacks. To solve these problems, fuzzy systems have become a competitive candidate in malware analysis. In this article, a new fuzzy-based approach is proposed for malware classification. We focused on portable executable files in the Windows platform and analyzed the distributions of static features and content-oriented features. Fuzzification was used to reduce the ubiquitous impact of noise and outliers in a very large dataset. Finally, a novel boosted classifier consisted of fuzzy decision trees and support vector machine is proposed to perform the malware classification. By using fuzzy decision trees, the inner structure of the classifier can be readily interpreted as discriminative rules, whereas the novel boosting strategy provides state-of-the-art classification performance. Extensive experimental results showed that our method significantly outperformed several state-of-the-art classifiers. Fangqi Li 0001, Shi-Lin Wang, Alan Wee-Chung Liew, Weiping Ding 0001, Gongshen Liu |
IEEE Trans. Fuzzy Syst. | 5 |
| 2020 | A Character-Centric Neural Model for Automated Story GenerationabstractAutomated story generation is a challenging task which aims to automatically generate convincing stories composed of successive plots correlated with consistent characters. Most recent generation models are built upon advanced neural networks, e.g., variational autoencoder, generative adversarial network, convolutional sequence to sequence model. Although these models have achieved prompting results on learning linguistic patterns, very few methods consider the attributes and prior knowledge of the story genre, especially from the perspectives of explainability and consistency. To fill this gap, we propose a character-centric neural storytelling model, where a story is created encircling the given character, i.e., each part of a story is conditioned on a given character and corresponded context environment. In this way, we explicitly capture the character information and the relations between plots and characters to improve explainability and consistency. Experimental results on open dataset indicate that our model yields meaningful improvements over several strong baselines on both human and automatic evaluations. Juntao Li 0005, Meng-Hsuan Yu, Ziming Huang, Gongshen Liu, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 5 |
| 2020 | Hierarchy-Aware Global Model for Hierarchical Text ClassificationabstractHierarchical text classification is an essential yet challenging subtask of multi-label text classification with a taxonomic hierarchy.Existing methods have difficulties in modeling the hierarchical label structure in a global view.Furthermore, they cannot make full use of the mutual interactions between the text feature space and the label space.In this paper, we formulate the hierarchy as a directed graph and introduce hierarchy-aware structure encoders for modeling label dependencies.Based on the hierarchy encoder, we propose a novel end-to-end hierarchy-aware global model (Hi-AGM) with two variants.A multi-label attention variant (HiAGM-LA) learns hierarchyaware label embeddings through the hierarchy encoder and conducts inductive fusion of labelaware text features.A text feature propagation model (HiAGM-TP) is proposed as the deductive variant that directly feeds text features into hierarchy encoders.Compared with previous works, both HiAGM-LA and HiAGM-TP achieve significant and consistent improvements on three benchmark datasets. Jie Zhou 0013, Chunping Ma, Dingkun Long, Ning Ding 0002, Pengjun Xie, Gongshen Liu |
ACL | 8 |
| 2020 | ASAP-Net: Attention and Structure Aware Point Cloud Sequence Segmentation
Yongyi Lu, Bo Pang 0003, Cewu Lu, Alan L. Yuille, Gongshen Liu |
BMVC | 6 |
| 2020 | A Memory-Based Sentence Split and Rephrase Model with Multi-task Training
Xiaoning Fan, Gongshen Liu, Bo Su 0003 |
ICONIP (1) | 3 |
| 2020 | Hierarchical Sentiment Estimation Model for Potential Topics of Individual Tweets
Qian Ji, Yilin Dai, Yinghua Ma, Gongshen Liu, Quanhai Zhang |
ICONIP (4) | 4 |
| 2020 | Learning Interactions at Multiple Levels for Abstractive Multi-document Summarization
Xiaoning Fan, Jie Zhou 0013, Gongshen Liu |
ICONIP (4) | 4 |
| 2020 | Sparse Hierarchical Modeling of Deep Contextual Attention for Document-Level Neural Machine Translation
Jianshen Zhang, YongAn Li, Gongshen Liu |
ICONIP (1) | 4 |
| 2020 | Neural Machine Translation with Soft Reordering Knowledge
Leiying Zhou, Jie Zhou 0013, Gongshen Liu |
ICONIP (4) | 5 |
| 2020 | Multi-paragraph Reading Comprehension with Token-level Dynamic Reader and Hybrid VerifierabstractMulti-paragraph reading comprehension requires the model to infer answers of arbitrary user-generated questions by reasoning cross-passage information. Previous work usually generates answer by directly employing a pointer network to predict the start and end position of the answer. However, span-level reading is insufficient since intermediate words may matter more. In this paper, we propose a novel unified network that includes a selector, a Token-level dynamic reader, and a Hybrid verifier (TH-Net). The core of token-level dynamic reader is a gate mechanism which dynamically selects important intermediate words according to boundary words. We decide the reader score from each token being both the boundary and the content. Moreover, we adopt a hybrid network verifier considering semantic answer-answer and entailment question-answer relationships to robust the model in case of being fooled by adversarial answers. Our experiments on SQuAD-document, SQuAD-open, and Trivia-wiki datasets show significant and consistent improvement as compared to other baselines and achieve the state-of-the-art performance on two of them. Yilin Dai, Qian Ji, Gongshen Liu, Bo Su 0003 |
IJCNN | 3 |
| 2020 | Unleashing the Potential of Attention Model for News Headline GenerationabstractHeadline generation is a special summarization generation task and the difficulty lies in requiring the generated headline to be concise, fluent and informative. Limited by the ability of commonly used encoder and decoder modules to capture long-term dependencies in seq2seq tasks, previous work rarely researched headline generation by end-to-end methods. However, the recent success of Transformer model and its subsequent improved versions demonstrate their remarkable performance on seq2seq tasks, which provide us with a feasible solution. In this paper, we propose a novel model Transformer(XL)-CC to generate headline from the perspective of understanding the whole text, the segment-level recurrence mechanism and relative positional encoding make our model learn ultra-long-term dependencies. In addition, we combine the copy and coverage mechanisms to generate more readable titles. Experimental results on the NYT and Chinese LSCC news datasets also confirm that our method significantly achieves better performance on the headline generation task. Jianshen Zhang, Gongshen Liu |
IJCNN | 4 |
| 2020 | Challenge Training to Simulate Inference in Machine TranslationabstractDespite much success has been achieved, neural machine translation (NMT) suffers from exposure bias and evaluation discrepancy. To be specific, the generation inconsistency between the training and inference process further causes error accumulation and distribution disparity. Furthermore, NMT models are generally optimized on word-level cross-entropy loss function but evaluated by sentence-level metrics. This evaluation-level mismatch may mislead the promotion of translation performance. To address these two drawbacks, we propose to challenge training to gradually simulate inference. Namely, the decoder is fed with inferred words rather than ground truth words during training with a dynamic probability. To ensure accuracy and integrity, we adopt alignment and tailoring on the inferred words. Therefore, these words can leverage inferred information to help improve the training process. As for the dynamic simulation, we define a novel loss-sensitive probability that can sense the converge of training and finetune itself in turn. Experimental results on IWSLT 2016 German-English and WMT 2019 English-Chinese datasets demonstrate that our methodology can significantly improve translation quality. The approach of alignment and tailoring outperforms previous works. Meanwhile, the proposed loss-sensitive sampling is also useful for other state-of-the-art scheduled sampling methods to achieve further promotion. Jie Zhou 0013, Leiying Zhou, Gongshen Liu, Quanhai Zhang |
IJCNN | 4 |
| 2020 | A Submodular Optimization-Based VAE-Transformer Framework for Paraphrase Generation
Xiaoning Fan, Xuejian Wang, Gongshen Liu, Bo Su 0003 |
NLPCC (1) | 5 |
| 2020 | Learning to Consider Relevance and Redundancy Dynamically for Abstractive Multi-document Summarization
Xiaoning Fan, Jie Zhou 0013, Gongshen Liu |
NLPCC (1) | 5 |
| 2020 | Incorporating Named Entity Information into Neural Machine Translation
Leiying Zhou, Jie Zhou 0013, Gongshen Liu |
NLPCC (1) | 5 |
| 2020 | Depth-Wise Separable Convolutions and Multi-Level Pooling for an Efficient Spatial CNN-Based SteganalysisabstractFor steganalysis, many studies showed that convolutional neural network (CNN) has better performances than the two-part structure of traditional machine learning methods. Existing CNN architectures use various tricks to improve the performance of steganalysis, such as fixed convolutional kernels, the absolute value layer, data augmentation and the domain knowledge. However, some designing of the network structure were not extensively studied so far, such as different convolutions (inception, xception, etc.) and variety ways of pooling(spatial pyramid pooling, etc.). In this paper, we focus on designing a new CNN network structure to improve detection accuracy of spatial-domain steganography. First, we use$3\times 3$kernels instead of the traditional$5\times 5$kernels and optimize convolution kernels in the preprocessing layer. The smaller convolution kernels are used to reduce the number of parameters and model the features in a small local region. Next, we use separable convolutions to utilize channel correlation of the residuals, compress the image content and increase the signal-to-noise ratio (between the stego signal and the image signal). Then, we use spatial pyramid pooling (SPP) to aggregate the local features and enhance the representation ability of features by multi-level pooling. Finally, data augmentation is adopted to further improve network performance. The experimental results show that the proposed CNN structure is significantly better than other five methods such as SRM, Ye-Net, Xu-Net, Yedroudj-Net and SRNet, when it is used to detect three spatial algorithms such as WOW, S-UNIWARD and HILL with a wide variety of datasets and payloads. Ru Zhang 0002, Gongshen Liu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Lip Image Segmentation in Mobile Devices Based on Alternative Knowledge DistillationabstractLip image segmentation, as the first step in many lip-related tasks (e.g. automatic lipreading), is of vital significance for the subsequent procedures. Nowadays, with the increasing computational power of the mobile devices, mobile applications become more and more popular. In this paper, a new approach is proposed, which is able to segment the lip region in natural scenes and is of acceptable computational complexity to be implemented in mobile devices. Two networks including a complex teacher network and a compact student network with the same structure are employed. With the proposed remedy loss and the alternative knowledge distillation scheme, the student network can learn useful knowledge from the teacher network effectively and efficiently, and even rectify some of its segmentation errors. A dataset containing 49 people captured under natural scenes by various cellphone cameras is adopted for evaluation and the experiment results have demonstrated that the proposed student network even outperforms the teacher network with much less computational cost. Cheng Guan, Shi-Lin Wang, Gongshen Liu, Alan Wee-Chung Liew |
ICIP | 3 |
| 2019 | Viewpoint Estimation in Images by a Key-Point Based Deep Neural NetworkabstractViewpoint estimation in a 2D image is a challenging task due to the great variations in the object's shape, appearance, visible parts, etc. To overcome the above difficulties, a new deep neural network is proposed, which employs the key-points of the object as a regularization term and a semantic bridge connecting the raw pixels with the object's viewpoint. A series of Hourglass structures are adopted for key-point extraction. With the extracted key-points, an LSTM based network is designed to model both the intrinsic relationship among the key-points and the underlying connections between the key-points and the object's viewpoint. A multitasks learning scheme is designed to optimize the key-point detection and the viewpoint estimation performance simultaneously. The experiment results on the PASCAL 3D+ dataset have demonstrated the effectiveness of the proposed approach. Jiana Yang, Shi-Lin Wang, Gongshen Liu |
ICIP | 3 |
| 2019 | Transformer-DW: A Transformer Network with Dynamic and Weighted Head
Ruxin Tan, Bo Su 0003, Gongshen Liu |
ICONIP (2) | 4 |
| 2019 | Hie-Transformer: A Hierarchical Hybrid Transformer for Abstractive Article Summarization
Xuewen Zhang, Gongshen Liu |
ICONIP (3) | 3 |
| 2019 | Paragraph-Level Hierarchical Neural Machine Translation
Gongshen Liu |
ICONIP (3) | 3 |
| 2019 | Hierarchical Attention CNN and Entity-Aware for Relation Extraction
Gongshen Liu, Bo Su 0003 |
ICONIP (4) | 2 |
| 2019 | Adaptive User Modeling with Long and Short-Term Preferences for Personalized RecommendationabstractUser modeling is an essential task for online recommender systems. In the past few decades, collaborative filtering (CF) techniques have been well studied to model users' long term preferences. Recently, recurrent neural networks (RNN) have shown a great advantage in modeling users' short term preference. A natural way to improve the recommender is to combine both long-term and short-term modeling. Previous approaches neglect the importance of dynamically integrating these two user modeling paradigms. Moreover, users' behaviors are much more complex than sentences in language modeling or images in visual computing, thus the classical structures of RNN such as Long Short-Term Memory (LSTM) need to be upgraded for better user modeling. In this paper, we improve the traditional RNN structure by proposing a time-aware controller and a content-aware controller, so that contextual information can be well considered to control the state transition. We further propose an attention-based framework to combine users' long-term and short-term preferences, thus users' representation can be generated adaptively according to the specific context. We conduct extensive experiments on both public and industrial datasets. The results demonstrate that our proposed method outperforms several state-of-art methods consistently. Zeping Yu, Jianxun Lian, Ahmad Mahmoody, Gongshen Liu, Xing Xie 0001 |
IJCAI | 4 |
| 2019 | A Transformer-Based Variational Autoencoder for Sentence GenerationabstractThe variational autoencoder(VAE) has been proved to be a most efficient generative model, but its applications in natural language tasks have not been fully developed. A novel variational autoencoder for natural texts generation is presented in this paper. Compared to the previously introduced variational autoencoder for natural text where both the encoder and decoder are RNN-based, we propose a new transformer-based architecture and augment the decoder with an LSTM language model layer to fully exploit information of latent variables. We also propose some methods to deal with problems during training time, such as KL divergency collapsing and model degradation. In the experiment, we use random sampling and linear interpolation to test our model. Results show that the generated sentences by our approach are more meaningful and the semantics are more coherent in the latent space. Gongshen Liu |
IJCNN | 2 |
| 2019 | Extending the Transformer with Context and Multi-dimensional Mechanism for Dialogue Response Generation
Ruxin Tan, Bo Su 0003, Gongshen Liu |
NLPCC (2) | 4 |
| 2019 | Dynamic Label Correction for Distant Supervision Relation Extraction via Semantic Similarity
Gongshen Liu, Bo Su 0003, Jan P. Nees |
NLPCC (2) | 2 |
| 2019 | A Bayesian Possibilistic C-Means clustering approach for cervical cancer screening
Fangqi Li 0001, Shi-Lin Wang, Gongshen Liu |
Inf. Sci. | 3 |
| 2019 | A global and local context integration DCNN for adult image classification
Shi-Lin Wang, Alan Wee-Chung Liew, Gongshen Liu |
Pattern Recognit. | 5 |
| 2018 | A Joint Selective Mechanism for Abstractive Sentence SummarizationabstractSequence-to-sequence (Seq2Seq) learning framework has been widely used in many natural language processing (NLP) tasks, including abstractive summarization and machine translation (MT). However, abstractive summarization generates the output in a lossy manner, in comparison with MT which is almost loss-less. We model this by introducing a joint selective mechanism: (i) A selective gate is added after encoding phase of the Seq2Seq learning framework, which learns to tailor the original input information and generates a selected input representation. (ii) A selection loss function is also added to help our selective gate function well, which is computed by looking at the input and the output jointly. Experimental results show that our proposed model outperforms most of the baseline models and is comparable to the state-of-the-art model in automatic evaluations. Junjie Fu, Gongshen Liu |
ACML | 2 |
| 2018 | Sliced Recurrent Neural NetworksabstractRecurrent neural networks have achieved great success in many NLP tasks. However, they have difficulty in parallelization because of the recurrent structure, so it takes much time to train RNNs. In this paper, we introduce sliced recurrent neural networks (SRNNs), which could be parallelized by slicing the sequences into many subsequences. SRNNs have the ability to obtain high-level information through multiple layers with few extra parameters. We prove that the standard RNN is a special case of the SRNN when we use linear activation functions. Without changing the recurrent units, SRNNs are 136 times as fast as standard RNNs and could be even faster when we train longer sequences. Experiments on six large-scale sentiment analysis datasets show that SRNNs achieve better performance than standard RNNs. Zeping Yu, Gongshen Liu |
COLING | 2 |
| 2018 | Modeling Multi-turn Conversation with Deep Utterance AggregationabstractMulti-turn conversation understanding is a major challenge for building intelligent dialogue systems. This work focuses on retrieval-based response matching for multi-turn conversation whose related work simply concatenates the conversation utterances, ignoring the interactions among previous utterances for context modeling. In this paper, we formulate previous utterances into context using a proposed deep utterance aggregation model to form a fine-grained context representation. In detail, a self-matching attention is first introduced to route the vital information in each utterance. Then the model matches a response with each refined utterance and the final matching score is obtained after attentive turns aggregation. Experimental results show our model outperforms the state-of-the-art methods on three multi-turn conversation benchmarks, including a newly introduced e-commerce dialogue corpus. Zhuosheng Zhang 0001, Jiangtong Li, Pengfei Zhu 0003, Hai Zhao 0001, Gongshen Liu |
COLING | 5 |
| 2018 | A Unified Syntax-aware Framework for Semantic Role LabelingabstractSemantic role labeling (SRL) aims to recognize the predicate-argument structure of a sentence.Syntactic information has been paid a great attention over the role of enhancing SRL.However, the latest advance shows that syntax would not be so important for SRL with the emerging much smaller gap between syntax-aware and syntax-agnostic SRL.To comprehensively explore the role of syntax for SRL task, we extend existing models and propose a unified framework to investigate more effective and more diverse ways of incorporating syntax into sequential neural networks.Exploring the effect of syntactic input quality on SRL performance, we confirm that high-quality syntactic parse could still effectively enhance syntactically-driven SRL.Using empirically optimized integration strategy, we even enlarge the gap between syntax-aware and syntax-agnostic SRL.Our framework achieves state-of-the-art results on CoNLL-2009 benchmarks both for English and Chinese, substantially outperforming all previous models. Zuchao Li, Shexia He, Jiaxun Cai, Zhuosheng Zhang 0001, Hai Zhao 0001, Gongshen Liu, Linlin Li 0001, Luo Si |
EMNLP | 6 |
| 2018 | Detecting Double Jpeg Compression with Same Quantization Matrix Based on Dense Cnn FeatureabstractDetection of double JPEG compression with same quantization matrix has been regarded as a challenging task in digital image forensics because there are very few modification cues in the tampered images especially when the compression quality factor is low. In order to solve this problem, a comprehensive feature representation based on the dense CNN framework is proposed, which is sensitive to the artifacts caused by double JPEG compression and is not related to the image content. With the appropriate network design and contributing to the characteristics of average pooling, dense connection and transition, the proposed network can differentiate double JPEG compression artifacts accurately. Experiment results on the two datasets have demonstrated that the proposed feature outperforms several state-of-the-art approaches investigated. Xiaosa Huang, Shi-Lin Wang, Gongshen Liu |
ICIP | 3 |
| 2018 | 3D Convolutional Neural Networks Based Speaker Identification and AuthenticationabstractResearch shows that human lips can be used as a new kind of biometrics in personal identification and authentication. In this letter, a novel end-to-end method based on 3D convolutional neural network (3DCNN) is proposed to extract discriminative spatiotemporal features from raw lip video streams. In our approach, the lip video is first divided into a series of overlapping clips. For each clip, the lip-characteristics network is proposed to characterize the minutiae of the lip region and its movement. Finally, the entire lip video is represented by a set of sub-features corresponding to each clip in it. Experiments have been performed on a dataset with 200 speakers and the proposed method achieves high identification accuracy of 99.18% and very low authentication error (HTER of 0.15%). Compared with several state-of-the-art methods, our approach achieves better performance and higher robustness against variations caused by different speaker's pose and position. Jianguo Liao, Shi-Lin Wang, Xingxuan Zhang, Gongshen Liu |
ICIP | 4 |
| 2018 | Adult Image Classification by a Local-Context Aware NetworkabstractTo build a healthy online environment, adult image recognition is a crucial and challenging task. Recent deep learning based methods have brought great advances to this task. However, the recognition accuracy and generalization ability need to be further improved. In this paper, a local-context aware network is proposed to improve the recognition accuracy and a corresponding curriculum learning strategy is proposed to guarantee a good generalization ability. The main idea is to integrate the global classification and the local sensitive region detection into one network and optimize them simulatenously. Such strategy helps the classification networks focus more on suspicious regions and thus provide better recognition performance. Two datasets containing over 150,000 images have been collected to evaluate the performance of the proposed approach. From the experiment results, it is observed that our approach can always achieve the best classification accuracy compared with several state-of-the-art approaches investigated. Shi-Lin Wang, Huanrong Sun, Gongshen Liu |
ICIP | 5 |
| 2018 | Leveraging Inner-Connection of Message Sequence for Traffic Classification: A Deep Learning ApproachabstractClassifying traffic flows into source applications is of great value for intelligent network management, which can help to detect malicious attacks, monitor the network, optimize network behaviors and then improve user experience, etc. However, to achieve high-accuracy traffic classification, especially in real time, is very challenging due to very complicated behaviors of traffic flows where network applications could often transmit traffics with encryption at randomized port numbers under highly dynamic network conditions. In this paper, by collecting extensive application traffic flows at the exit router of Shanghai Maritime University (the traffic rate can reach up to 7 GB/s at peak time), we identify that there is a very distinct characteristic in inner-connection of message (grouped by single or multiple consecutive TCP packets) sequence for different application flows. We then propose our traffic classification algorithm, which essentially adopts a Long Short-Term Memory (LSTM) neural network to output a classifier with message sequence vector (not necessarily covering all messages) of a traffic flow as the training input, to conduct online traffic flow classification. Extensive simulations are conduced considering varied training data size and diverse source applications, and an average about 97 % accuracy on per-flow classification can be achieved. Renjie Jin, Guangtao Xue, Feng Lyu 0001, Hao Sheng 0001, Gongshen Liu, Minglu Li 0001 |
ICPADS | 5 |
| 2018 | Five-Stroke Based CNN-BiRNN-CRF Network for Chinese Named Entity Recognition
Jianhu Zhang, Gongshen Liu, Jie Zhou 0013, Huanrong Sun |
NLPCC (1) | 3 |
| 2018 | LM Enhanced BiRNN-CRF for Joint Chinese Word Segmentation and POS Tagging
Jianhu Zhang, Gongshen Liu, Jie Zhou 0013, Huanrong Sun |
NLPCC (2) | 2 |
| 2018 | Paraphrase Identification Based on Weighted URAE, Unit Similarity and Context Correlation Feature
Jie Zhou 0013, Gongshen Liu, Huanrong Sun |
NLPCC (2) | 2 |
| 2017 | Ada-copy: An Adaptive Memory Copy Strategy for Virtual Machine Live MigrationabstractIn the cloud computing architecture, virtual machine (VM) live migration is a fundamental research topic which has drawn extensive attention from communities of industry and academy. It is critical to transfer the VM memory pages that contain essential state information to resume the VM on another host during virtual machine (VM) live migration. There are many memory copy methods, such as pre-copy and post-copy. However, these methods have two limitations: application generality and performance imbalance. In this paper, we propose Ada-copy (adaptive copy), an adaptive memory copy strategy for VM live migration. The basic idea of Ada-copy is that the memory copy method of a VM should be determined by its workload characteristics. Specifically, based on the variation of current dirty page rate of memory, Ada-copy can adaptively select the most appropriate migration method to copy memory pages, thus addressing the two limitations of existing memory copy methods. To evaluate the effectiveness of our proposed strategy, we experiment with the Ada-copy on a variety of migration tasks with different dirty page rate and diverse memory usage workloads. Evaluation results show, compared with traditional methods, Ada-copy can significantly reduce the total migration time by 26%, the VM downtime by 42% and the amount of pages transferred by 35% in average. Zhong Wang 0013, Guangtao Xue, Shiyou Qian, Gongshen Liu, Minglu Li 0001, Jian Cao 0001, Jiadi Yu |
ICPADS | 4 |
| 2016 | Uncovering fuzzy communities in networks with structural similarity
Xiaofeng Wang 0004, Gongshen Liu, Li Pan 0002, Jianhua Li 0001 |
Neurocomputing | 2 |
| 2014 | An modularity-based overlapping community structure detecting algorithmabstractMany algorithms have been designed to detect community structure in social networks. However, most algorithms can only detect disjoint communities effectively. A new overlapping community structure detecting algorithm is proposed in this paper, which adopts modularity to community clustering. In order to evaluate the algorithm, Modularity by Newman and the NMI (Normalized Mutual Information) by Lancichinetti are used as the evaluation metrics. It is approved by the experiments that the proposed method works well to the real overlapping communities. Gongshen Liu, Jianhua Li 0001 |
ASONAM | 2 |
| 2014 | A detecting community method in complex networks with fuzzy clusteringabstractDetection of community structure in complex networks is a significant aspect in social network analysis. A novel fuzzy clustering method is proposed in this paper, by which the community structure can be divided. In contrast to previous studies, the proposed method processes similarity of connecting vertices with fuzzy relation. In our method, we globally consider the fuzzy relation between vertices and the similarity in network topology to divide vertices into communities. In addition, smaller grained communities can be detected by adjusting fuzzy parameter. In order to avoid subjectivity in the selection of cluster number, a new modularity is introduced to evaluate the effectiveness of the clustering analysis. It's proved by experiments that the method is efficient in detecting both good communities and appropriate number of clusters. Xiaofeng Wang 0004, Gongshen Liu, Jianhua Li 0001 |
DSAA | 2 |
| 2008 | A cross-language state mapping approach to bilingual (Mandarin-English) TTSabstractWe propose a cross-language state mapping approach to HMM-based bilingual TTS. Two language-dependent decision trees are built first with a bilingual speech database recorded by a single speaker. A state mapping for every leaf node in the decision tree of a target language is created by finding the nearest leaf node in the tree of a source language. Kullback-Leibler divergence between two distributions is used to find the nearest leaf node. To synthesize target language speech by a monolingual, (source language) speaker's voice, we find HMM parameters trained by the monolingual (source language) speaker in the mapped leaf nodes. Similar mappings can be constructed by reversing the source and target languages. With these bi-directional cross-lingual mappings, we can synthesize bilingual or mixed-code speech by HMMs trained by any monolingual speaker. High voice (speaker) similarity is preserved in synthesized speech of the target language. Two perceptual tests on synthesized Mandarin speech confirms high intelligibility with a Chinese character transcription accuracy of 92.1% and an MOS score of 3.08. Yao Qian, Frank K. Soong, Gongshen Liu |
ICASSP | 4 |