VLDB 2026 Research / reviewers in the wild / expert
Degen Huang
dblp:67/5547
· DBLP profile ↗
60ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-8860-7805ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 2 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 14Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkabstractCurrent large language models (LLMs) often exhibit performance imbalances between dominant languages (e.g., English) and nondominant ones due to the skewed distribution of pretraining data.A common strategy to address this issue is to enhance cross-lingual alignment, thereby facilitating non-dominant language processing.However, existing methods typically rely on additional training objectives or language-specific parameters, which increase training complexity and cost.In this work, we propose a selective bidirectional language projection framework that enables efficient multilingual alignment and language shift using the intrinsic parameters.Specifically, we first identify the layers most sensitive to language projection between non-dominant and dominant languages through neuron activation analysis.We then perform sequential language projection within the selected layers by mapping non-dominant representations into the dominant language space and reverting them before generation.The bidirectional projection benefits the subsequent instruction tuning in non-dominant languages.Experiments on seven benchmarks demonstrate that our method remarkably enhances the performance of nondominant languages.Further analyses indicate that our method learns better internal representations and exhibits strong generalization capabilities. Junpeng Liu 0002, Jiuyi Li, Bo Jin 0001, Degen Huang, Hui Xiong 0001 |
ACL (1) | 5 |
| 2026 | One Pair Suffices: Unlocking Universal Zero-Shot Translation via Cross-Architecture AlignmentabstractHao Zong, Cong Hu Yuan, Chao Bei, Wentao Chen, Huan Liu, Kaiyu Huang, Degen Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Zong, Conghu Yuan, Chao Bei, Degen Huang |
ACL (1) | 7 |
| 2026 | AlignExperts: Knowledge alignment experts for summarization faithfulness evaluation
Jiuyi Li, Degen Huang, Junpeng Liu 0002, Jinshuang Zhao |
Expert Syst. Appl. | 2 |
| 2026 | CrossSeg-GvT: multi-view graph vision transformers with context-aware memory and meta prompting for cross-domain few-shot semantic segmentation
Anil Ahmed, Degen Huang, Salahuddin Unar, Mobeen Nazar |
Neurocomputing | 2 |
| 2025 | Continual Learning for Multilingual Neural Machine Translation via Meta-Contrastive Memory Replay
Degen Huang, Fengran Mo |
NLPCC (3) | 2 |
| 2025 | Bridging Semantic and Emotional Gaps in Speech Emotion Recognition: A Novel LightGBM-Stacking Architecture
Zhilong Duan, Degen Huang, Xuewen Shi 0001 |
PAKDD (3) | 3 |
| 2025 | Best of Both Worlds? A Glance at Efficient Reasoning for LLM-Based Machine Translation
Hao Zong, Chao Bei, Conghu Yuan, Degen Huang |
PRICAI | 7 |
| 2025 | Towards better text image machine translation with multimodal codebook and multi-stage training
Zhibin Lan, Junfeng Yao, Degen Huang, Jinsong Su |
Neural Networks | 5 |
| 2025 | Urdu Word Sense Disambiguation: Leveraging Contextual Stacked Embedding, Siamese Transformer Encoder 1DCNN-BiLSTM, and Gloss Data AugmentationabstractWord Sense Disambiguation (WSD) in Natural Language Processing (NLP) is crucial for discerning the correct meaning of words with multiple senses in various contexts. Recent advancements in this field, particularly Deep Learning (DL) and sophisticated language models such as BERT and GPT, have significantly improved WSD performance. However, challenges persist, especially with languages like Urdu, which are known for their linguistic complexity and limited digital resources compared to English. This study addresses the challenge of advancing WSD in Urdu by developing and applying tailored Data Augmentation (DA) techniques. We introduce an innovative approach, Prompt Engineering with Retrieval Augmented Generation (RAG), leveraging GPT-3.5-turbo to generate context-sensitive Gloss Definitions (GD). Additionally, we employ sentence-level and word-level DA techniques, including Back Translation (BT) and Masked Word Prediction (MWP). To enhance sentence understanding, we combine three BERT embedding models: mBERT, mDistilBERT, and Roberta_Urdu, facilitating a more nuanced comprehension of sentences and improving word disambiguation in complex linguistic contexts. Furthermore, we propose a novel network architecture merging Transformer Encoder (TE)-CNN and TE-BiLSTM models with Multi-Head Self-Attention (MHSA), One-Dimensional Convolutional Neural Network (1DCNN), and Bidirectional Long Short-Term Memory (BiLSTM). This architecture is tailored to address polysemy and capture short- and long-range dependencies critical for effective WSD in Urdu. Empirical evaluations on Lexical Sample (LS) and All Word (AW) tasks demonstrate the effectiveness of our approach, achieving an 88.9% F1 score on the LS and a 79.2% F1 score on AW tasks. These results underscore the importance of language-specific approaches and the potential of DA and advanced modeling techniques in overcoming challenges associated with WSD in languages with limited resources. Anil Ahmed, Degen Huang, Syed Yasser Arafat, Khawaja Iftekhar Rashid |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | How to Leverage Personal Textual Knowledge for Personalized Conversational Information RetrievalabstractPersonalized conversational information retrieval (CIR) combines conversational and personalizable elements to satisfy various users' complex information needs through multi-turn interaction based on their backgrounds. The key promise is that the personal textual knowledge base (PTKB) can improve the CIR effectiveness because the retrieval results can be more related to the user's background. However, PTKB is noisy: not every piece of knowledge in PTKB is relevant to the specific query at hand. In this paper, we explore and test several ways to select knowledge from PTKB and use it for query reformulation by using a large language model (LLM). The experimental results show the PTKB might not always improve the search results when used alone, but LLM can help generate a more appropriate personalized query when high-quality guidance is provided. Fengran Mo, Longxiang Zhao, Yue Dong 0002, Degen Huang, Jian-Yun Nie |
CIKM | 5 |
| 2024 | Context-Aware Non-Autoregressive Document-Level Translation with Sentence-Aligned Connectionist Temporal ClassificationabstractPrevious studies employ the autoregressive translation (AT) paradigm in the document-to-document neural machine translation. These methods extend the translation unit from a single sentence to a pseudo-document and encodes the full pseudo-document, avoiding the redundant computation problem in context. However, the AT methods cannot parallelize decoding and struggle with error accumulation, especially when the length of sentences increases. In this work, we propose a context-aware non-autoregressive framework with the sentence-aligned connectionist temporal classification (SA-CTC) loss for document-level neural machine translation. In particular, the SA-CTC loss reduces the search space of the decoding path by fixing the positions of the beginning and end tokens for each sentence in the document. Meanwhile, the context-aware architecture introduces preset nodes to represent sentence-level information and utilizes a hierarchical attention structure to regulate the attention hypothesis space. Experimental results show that our proposed method can achieve competitive performance compared with several strong baselines. Our method implements non-autoregressive modeling in Doc-to-Doc translation manner, achieving an average 46X decoding speedup compared to the document-level AT baselines on three benchmarks. Hao Yu 0019, Junpeng Liu 0002, Degen Huang |
LREC/COLING | 5 |
| 2024 | Enriching Urdu NER with BERT Embedding, Data Augmentation, and Hybrid Encoder-CNN ArchitectureabstractNamed Entity Recognition (NER) is an indispensable component of Natural Language Processing (NLP), which aims to identify and classify entities within text data. While Deep Learning (DL) models have excelled in NER for well-resourced languages such as English, Spanish, and Chinese, they face significant hurdles when dealing with low-resource languages such as Urdu. These challenges stem from the intricate linguistic characteristics of Urdu, including morphological diversity, a context-dependent lexicon, and the scarcity of training data. This study addresses these issues by focusing on Urdu Named Entity Recognition (U-NER) and introducing three key contributions. First, various pre-trained embedding methods are employed, encompassing Word2vec (W2V), GloVe, FastText, Bidirectional Encoder Representations from Transformers (BERT), and Embeddings from language models (ELMo). In particular, fine-tuning is performed on BERT BASE and ELMo using Urdu Wikipedia and news articles. Second, a novel generative Data Augmentation (DA) technique replaces Named Entities (NEs) with mask tokens, employing pre-trained masked language models to predict masked tokens, effectively expanding the training dataset. Finally, the study introduces a novel hybrid model combining a Transformer Encoder with a Convolutional Neural Network (CNN) to capture the intricate morphology of Urdu. These modules enable the model to handle polysemy, extract short- and long-range dependencies, and enhance learning capacity. Empirical experiments demonstrate that the proposed model, incorporating BERT embeddings and an innovative DA approach, attains the highest F1-score of 93.99%, highlighting its efficacy for the U-NER task. Anil Ahmed, Degen Huang, Syed Yasser Arafat, Imran Hameed |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Boundary-Aware Abstractive Summarization with Entity-Augmented Attention for Enhancing FaithfulnessabstractWith the successful application of deep learning, document summarization systems can produce more readable results. However, abstractive summarization still suffers from unfaithful outputs and factual errors, especially in named entities. Current approaches tend to employ external knowledge to improve model performance while neglecting the boundary information and the semantics of the entities. In this article, we propose an entity-augmented method (EAM) to encourage the model to make full use of the entity boundary information and pay more attention to the critical entities. Experimental results on three Chinese and English summarization datasets show that our method outperforms several strong baselines and achieves state-of-the-art performance on the CLTS dataset. Our method can also improve the faithfulness of the summary and generalize well to different pre-trained language models. Moreover, we propose a method to evaluate the integrity of generated entities. Besides, we adapt the data augmentation method in the FactCC model according to the difference between Chinese and English in grammar and train a new evaluation model for factual consistency evaluation in Chinese summarization. Jiuyi Li, Junpeng Liu 0002, Degen Huang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Exploring Better Text Image Translation with Multimodal CodebookabstractZhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang, Jian Luan, Bin Wang, Degen Huang, Jinsong Su. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhibin Lan, Xiang Li 0104, Wen Zhang 0015, Jian Luan 0001, Bin Wang 0004, Degen Huang, Jinsong Su |
ACL (1) | 7 |
| 2023 | Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionabstractA persistent goal of multilingual neural machine translation (MNMT) is to continually adapt the model to support new language pairs or improve some current language pairs without accessing the previous training data.To achieve this, the existing methods primarily focus on preventing catastrophic forgetting by making compromises between the original and new language pairs, leading to sub-optimal performance on both translation tasks.To mitigate this problem, we propose a dual importancebased model division method to divide the model parameters into two parts and separately model the translation of the original and new tasks.Specifically, we first remove the parameters that are negligible to the original tasks but essential to the new tasks to obtain a pruned model, which is responsible for the original translation tasks.Then we expand the pruned model with external parameters and fine-tune the newly added parameters with new training data.The whole fine-tuned model will be used for the new translation tasks.Experimental results show that our method can efficiently adapt the original model to various new translation tasks while retaining the performance of the original tasks.Further analyses demonstrate that our method consistently outperforms several strong baselines under different incremental translation scenarios. Junpeng Liu 0002, Hao Yu 0019, Jiuyi Li, Jinsong Su, Degen Huang |
EMNLP | 6 |
| 2023 | Multi-modal graph contrastive encoding for neural machine translation
Yongjing Yin, Jiali Zeng, Jinsong Su, Chulun Zhou, Fandong Meng, Jie Zhou 0016, Degen Huang, Jiebo Luo 0001 |
Artif. Intell. | 7 |
| 2023 | Multigranularity Pruning Model for Subject Recognition Task under Knowledge Base Question Answering When General Models FailabstractIn general knowledge base question answering (KBQA) models, subject recognition (SR) is usually a precondition of finding an answer, and it is a common way to employ a general named entity recognition (NER) model such as BERT‐CRF to recognize the subject. However, in previous researches, the difference between a NER task and a SR task is usually ignored, and a wrong entity recognized by the NER model will certainly lead to a wrong answer in the KBQA task, which is one bottleneck for KBQA performance. In this paper, a multigranularity pruning model (MGPM) is proposed to answer a question when general models fail to recognize a subject. In MGPM, the set of all possible subjects in the Knowledge Base (KB) is pruned by 4 multigranularity pruning submodels successively based on the constraint of relation (domain and tuple), string similarity, and semantic similarity. Experimental results show that our model is compatible with various KBQA models for both single‐relation and complex questions answering. The integrated MGPM model (with the BERT‐CRF model) achieves a SR accuracy of 94.4% on the SimpleQuestions dataset, 68.6% on the WebQuestionsSP dataset, and 63.7% on the WebQuestions dataset, which outperforms the original model by a margin of 3.6%, 8.6%, and 5.3%, respectively. Xirong Xu, Xiaopeng Wei, Degen Huang |
Int. J. Intell. Syst. | 6 |
| 2023 | Exploiting Japanese-Chinese Cognates with Shared Private Representations for NMTabstractNeural machine translation has achieved remarkable progress over the past several years; however, little attention has been paid to machine translation (MT) between Japanese and Chinese, which share a large proportion of cognate words that can be utilized as additional linguistic knowledge to enhance translation performance. In this article, we seek to strengthen the semantic correlation between Japanese and Chinese by leveraging cognate words that share common Chinese characters. Specifically, we experiment with three strategies: (1) a shared vocabulary with cognate lexicon induction, which models the commonality between source and target cognates; (2) a shared private representation with a dynamic gating mechanism, which models the language-specific features on the source side; and (3) an embedding shortcut, which enables the decoder to access the shared private representation with shortest distance and aids the training process. The experiments and analysis presented in this article demonstrate that our proposed approaches can significantly improve the performance of both Japanese-to-Chinese and Chinese-to-Japanese translations and verify the effectiveness of exploiting Japanese–Chinese cognates for MT. Fuji Ren, Xiao Sun 0003, Degen Huang, Piao Shi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Multilingual BERT-based Word Alignment By Incorporating Common Chinese CharactersabstractWord alignment is an important task of detecting translation equivalents between a sentence pair. Although word alignment is no longer necessarily needed for neural machine translation, it’s still useful in a wealth of applications, e.g., bilingual lexicon induction, constraint decoding, and so on. However, the most well-known word aligners are still Giza++ and fastAlign, both of which are implementations of traditional IBM models. To keep pace with the advance in NMT, there has been a surge of interest in replacing the IBM models with neural models. We follow this trend but aim to boost performance of word alignment between Japanese and Chinese, which share a large portion of Chinese characters. Our key idea is to leverage these common Chinese characters in both languages as an indicator for inferring alignment; i.e., the source and target words with the common Chinese characters should be most likely aligned. Following this idea, we propose three methods that leverage common Chinese characters to boost the mBERT-based word alignment, including reward factor, representation alignment, and contrastive training. Furthermore, we annotate and release a golden dataset for Japanese-Chinese word alignment. Experiments on the dataset show that our methods outperform several strong baselines in terms of AER score and verify the effectiveness of exploiting common Chinese characters. Xiao Sun 0003, Fuji Ren, Degen Huang, Piao Shi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Multi-task Label-wise Transformer for Chinese Named Entity RecognitionabstractBenefiting from the improvement of positional encoding and the introduction of lexical knowledge, Transformer has achieved superior performance than the prevailing BiLSTM-based models in named entity recognition (NER) task. However, existing Transformer-based models for Chinese NER pay less attention to the information captured by the bottom layers of Transformer and the significance of representation subspace where each head of Transformer is projected. In this article, we propose M ulti- T ask L abel- W ise T ransformer (MTLWT). From a global perspective, we assign entity boundary prediction (EBP) and entity type prediction (ETP) tasks to the first two layers. In this way, we stimulate lower layers to participate more in constructing character representation. Besides, in each multi-head self-attention (MHSA) layer, we provide a specific focus for each individual head, making the head project into a significant subspace. Experiments on four datasets from different domains show that our proposed model achieves comparable performance with other state-of-the-art models. In particular, MTLWT outperforms the other frameworks without external knowledge on all the datasets. Xirong Xu, Degen Huang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | Towards Robust k-Nearest-Neighbor Machine Translationabstractk-Nearest-Neighbor Machine Translation (kNN-MT) becomes an important research direction of NMT in recent years.Its main idea is to retrieve useful key-value pairs from an additional datastore to modify translations without updating the NMT model.However, the underlying retrieved noisy pairs will dramatically deteriorate the model performance.In this paper, we conduct a preliminary study and find that this problem results from not fully exploiting the prediction of the NMT model.To alleviate the impact of noise, we propose a confidence-enhanced kNN-MT model with robust training.Concretely, we introduce the NMT confidence to refine the modeling of two important components of kNN-MT: kNN distribution and the interpolation weight.Meanwhile we inject two types of perturbations into the retrieved pairs for robust training.Experimental results on four benchmark datasets demonstrate that our model not only achieves significant improvements over current kNN-MT models, but also exhibits better robustness.Our code is available at https://github.com/ DeepLearnXMU/Robust-knn-mt. Ziyao Lu, Fandong Meng, Chulun Zhou, Jie Zhou 0016, Degen Huang, Jinsong Su |
EMNLP | 6 |
| 2022 | Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation aims to translate multiple language pairs in a single model and has shown great success thanks to the knowledge transfer across languages with the shared parameters.Despite promising, this share-all paradigm suffers from insufficient ability to capture language-specific features.Currently, the common practice is to insert or search language-specific networks to balance the shared and specific features.However, those two types of features are not sufficient enough to model the complex commonality and divergence across languages, such as the locally shared features among similar languages, which leads to sub-optimal transfer, especially in massively multilingual translation.In this paper, we propose a novel token-level feature mixing method that enables the model to capture different features and dynamically determine the feature sharing across languages.Based on the observation that the tokens in the multilingual model are usually shared by different languages, we insert a feature mixing layer into each Transformer sublayer and model each token representation as a mix of different features, with a proportion indicating its feature preference.In this way, we can perform fine-grained feature sharing and achieve better multilingual transfer.Experimental results on multilingual datasets show that our method outperforms various strong baselines and can be extended to zero-shot translation.Further analyses reveal that our method can capture different linguistic features and bridge the representation gap across languages. Junpeng Liu 0002, Jiuyi Li, Jinsong Su, Degen Huang |
EMNLP | 6 |
| 2022 | SSMFRP: Semantic Similarity Model for Relation Prediction in KBQA Based on Pre-trained Models
Xirong Xu, Xinzi Li, Xiaopeng Wei, Degen Huang |
ICANN (2) | 6 |
| 2022 | Domain-Aware Word Segmentation for Chinese Language: A Document-Level Context-Aware ModelabstractWord segmentation is an essential and challenging task in natural language processing, especially for the Chinese language due to its high linguistic complexity. Existing methods for Chinese word segmentation, including statistical machine learning methods and neural network methods, usually have good performance in specific knowledge domains. Given the increasing importance of interdisciplinary and cross-domain studies, one of the challenges in cross-domain word segmentation is to handle the out-of-vocabulary (OOV) words. Existing methods show unsatisfactory performance to meet the practical standard. To this end, we propose a document-level context-aware model that can automatically perceive and identify OOV words from different domains. Our method jointly implements a word-based and a character-based model and then processes the results with a newly proposed reconstruction model. We evaluate the new method by designing and conducting comprehensive experiments on two real-world datasets (e.g., news from different domains). The results demonstrate the superiority of our method over the state-of-the-art models in handling texts from different domains. Importantly, when doing the word segmentation under the cross-domain scenario, our proposed method can improve the performance of OOV words recognition. Keli Xiao, Fengran Mo, Bo Jin 0001, Zhuang Liu 0001, Degen Huang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2021 | Exploring Dynamic Selection of Branch Expansion Orders for Code GenerationabstractHui Jiang, Chulun Zhou, Fandong Meng, Biao Zhang, Jie Zhou, Degen Huang, Qingqiang Wu, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chulun Zhou, Fandong Meng, Biao Zhang 0002, Jie Zhou 0016, Degen Huang, Qingqiang Wu 0001, Jinsong Su |
ACL/IJCNLP (1) | 6 |
| 2021 | Towards User-Driven Neural Machine TranslationabstractHuan Lin, Liang Yao, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Degen Huang, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Weihua Luo, Degen Huang, Jinsong Su |
ACL/IJCNLP (1) | 7 |
| 2021 | Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise OrderingsabstractShaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou 0016, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su |
EMNLP (1) | 8 |
| 2021 | Adaptive Transformer for Multilingual Neural Machine Translation
Junpeng Liu 0002, Jiuyi Li, Degen Huang |
NLPCC (1) | 5 |
| 2020 | Dual Head-wise Coattention Network for Machine Comprehension with Multiple-Choice QuestionsabstractMultiple-choice Machine Comprehension (MC) is an important and challenging nature language processing (NLP) task where the machine is required to make the best answer from candidate answer set given particular passage and question. Existing approaches either only utilize the powerful pre-trained language models or only rely on an over complicated matching network that is design supposed to capture the relationship effectively among the triplet of passage, question and candidate answers. In this paper, we present a novel architecture, Dual Head-wise Coattention network (called DHC), which is a simple and efficient attention neural network designed to perform multiple-choice MC task. Our proposed DHC not only support a powerful pre-trained language model as encoder, but also models the MC relationship as attention mechanism straightforwardly, by head-wise matching and aggregating method on multiple layers, which better model relationships sufficiently between question and passage, and cooperate with large pre-trained language models more efficiently. To evaluate the performance, we test our proposed model on five challenging and well-known datasets for multiple-choice MC: RACE, DREAM, SemEval-2018 Task 11, OpenBookQA, and TOEFL. Extensive experimental results demonstrate that our proposal can achieve a significant increase in accuracy comparing existing models based on all five datasets, and it consistently outperforms all tested baselines including the state-of-the-arts techniques. More remarkably, our proposal is a pluggable and more flexible model, and it thus can be plugged into any pre-trained Language Models based on BERT. Ablation studies demonstrate its state-of-the-art performance and generalization. Zhuang Liu 0001, Degen Huang, Zhuang Liu 0005, Jun Zhao 0022 |
CIKM | 3 |
| 2020 | Semantics-Reinforced Networks for Question Generation
Zhuang Liu 0001, Degen Huang, Jun Zhao 0022 |
ECAI | 3 |
| 2020 | A Joint Multiple Criteria Model in Transfer Learning for Cross-domain Chinese Word SegmentationabstractWord-level information is important in natural language processing (NLP), especially for the Chinese language due to its high linguistic complexity. Chinese word segmentation (CWS) is an essential task for Chinese downstream NLP tasks. Existing methods have already achieved a competitive performance for CWS on large-scale annotated corpora. However, the accuracy of the method will drop dramatically when it handles an unsegmented text with lots of out-of-vocabulary (OOV) words. In addition, there are many different segmentation criteria for addressing different requirements of downstream NLP tasks. Excessive amounts of models with saving different criteria will generate the explosive growth of the total parameters. To this end, we propose a joint multiple criteria model that shares all parameters to integrate different segmentation criteria into one model. Besides, we utilize a transfer learning method to improve the performance of OOV words. Our proposed method is evaluated by designing comprehensive experiments on multiple benchmark datasets (e.g., Bakeoff 2005, Bakeoff 2008 and SIGHAN 2010). Our method achieves the state-of-the-art performances on all datasets. Importantly, our method also shows a competitive practicability and generalization ability for the CWS task. Degen Huang, Zhuang Liu 0001, Fengran Mo |
EMNLP (1) | 2 |
| 2020 | FinBERT: A Pre-trained Financial Language Representation Model for Financial Text MiningabstractThere is growing interest in the tasks of financial text mining. Over the past few years, the progress of Natural Language Processing (NLP) based on deep learning advanced rapidly. Significant progress has been made with deep learning showing promising results on financial text mining models. However, as NLP models require large amounts of labeled training data, applying deep learning to financial text mining is often unsuccessful due to the lack of labeled training data in financial fields. To address this issue, we present FinBERT (BERT for Financial Text Mining) that is a domain specific language model pre-trained on large-scale financial corpora. In FinBERT, different from BERT, we construct six pre-training tasks covering more knowledge, simultaneously trained on general corpora and financial domain corpora, which can enable FinBERT model better to capture language knowledge and semantic information. The results show that our FinBERT outperforms all current state-of-the-art models. Extensive experimental results demonstrate the effectiveness and robustness of FinBERT. The source code and pre-trained models of FinBERT are available online. Zhuang Liu 0001, Degen Huang, Jun Zhao 0022 |
IJCAI | 2 |
| 2020 | Chinese Syntax Parsing Based on Sliding Match of Semantic StringabstractDifferent from the current syntax parsing based on deep learning, we present a novel Chinese parsing method, which is based on Sliding Match of Semantic String (SMOSS). (1) Training stage: In a treebank, headwords of tree nodes are represented by semantic codes given in the Synonym Dictionary (Tongyici Cilin). N-gram semantic templates are extracted from every layer of a syntax tree by means of sliding window to establish one N-gram semantic template library. (2) Parsing stage: Words of a sentence, including headwords of chunks, are represented by the semantic codes from Tongyici Cilin. With the sliding window method, N-gram semantic code strings are extracted to match with the templates in the N-gram semantic template library; subsequently, the mapping information of the matched templates is employed to guide the chunking of semantic code strings. The Chinese syntax parsing is completed through continuous matching and chunking. On the same training scale, N-gram semantic template can create favorable conditions for flexible matching and improve the syntax parsing performance. With train and test sets from the Tsinghua Chinese Treebank (TCT), the results are F1-score 99.71% (closed test) and F1-score 70.43% (open test), respectively. Wei Wang 0412, Degen Huang, Jingxiang Cao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | Unified Generative Adversarial Networks for Multiple-Choice Oriented Machine ComprehensionabstractIn this article, we address the multiple-choice machine comprehension (MC) problem in natural language processing. Existing approaches for MC are usually designed for general cases; however, we specially develop a novel method for solving the multiple-choice MC problem. We take the inspiration generative adversarial networks (GANs) and first propose an adversarial framework for multiple-choice oriented MC, named McGAN . Specifically, our approach is designed as a GAN-based method that unifies both generative and discriminative MC models. Working together, the generative model focuses on predicting relevant answer given a passage (text) and a question; the discriminative model focuses on predicting their relevancy given an answer-passage-question set. Based on the competition via adversarial training in a minimize-maximize game, the proposed method takes advantages from both models. To evaluate the performance, we test our McGAN model on three well-known datasets for multiple-choice MC. Our results show that McGAN can achieve a significant increase in accuracy compared to existing models based on all three datasets, and it consistently outperforms all tested baselines, including state-of-the-art techniques. Zhuang Liu 0001, Keli Xiao, Bo Jin 0001, Degen Huang, Yunxia Zhang |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Combining Context and Knowledge Representations for Chemical-Disease Relation ExtractionabstractAutomatically extracting the relationships between chemicals and diseases is significantly important to various areas of biomedical research and health care. Biomedical experts have built many large-scale knowledge bases (KBs) to advance the development of biomedical research. KBs contain huge amounts of structured information about entities and relationships, therefore plays a pivotal role in chemical-disease relation (CDR) extraction. However, previous researches pay less attention to the prior knowledge existing in KBs. This paper proposes a neural network-based attention model (NAM) for CDR extraction, which makes full use of context information in documents and prior knowledge in KBs. For a pair of entities in a document, an attention mechanism is employed to select important context words with respect to the relation representations learned from KBs. Experiments on the BioCreative V CDR dataset show that combining context and knowledge representations through the attention mechanism, could significantly improve the CDR extraction performance while achieve comparable results with state-of-the-art systems. Huiwei Zhou, Shixian Ning, Zhuang Liu 0001, Chengkun Lang, Yingyu Lin, Degen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2018 | Incorporating Prior Knowledge into Word Embedding for Chinese Word Similarity MeasurementabstractWord embedding-based methods have received increasing attention for their flexibility and effectiveness in many natural language-processing (NLP) tasks, including Word Similarity (WS). However, these approaches rely on high-quality corpus and neglect prior knowledge. Lexicon-based methods concentrate on human’s intelligence contained in semantic resources, e.g., Tongyici Cilin, HowNet, and Chinese WordNet, but they have the drawback of being unable to deal with unknown words. This article proposes a three-stage framework for measuring the Chinese word similarity by incorporating prior knowledge obtained from lexicons and statistics into word embedding: in the first stage, we utilize retrieval techniques to crawl the contexts of word pairs from web resources to extend context corpus. In the next stage, we investigate three types of single similarity measurements, including lexicon similarities, statistical similarities, and embedding-based similarities. Finally, we exploit simple combination strategies with math operations and the counter-fitting combination strategy using optimization method. To demonstrate our system’s efficiency, comparable experiments are conducted on the PKU-500 dataset. Our final results are 0.561/0.516 of Spearman/Pearson rank correlation coefficient, which outperform the state-of-the-art performance to the best of our knowledge. Experiment results on Chinese MC-30 and SemEval-2012 datasets show that our system also performs well on other Chinese datasets, which proves its transferability. Besides, our system is not language-specific and can be applied to other languages, e.g., English. Degen Huang, Jiahuan Pei |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2017 | A Hierarchical Iterative Attention Model for Machine Comprehension
Zhuang Liu 0001, Degen Huang, Yunxia Zhang |
NLDB | 2 |
| 2017 | Biomedical Domain-Oriented Word Embeddings via Small Background Texts for Biomedical Text Mining Tasks
Lishuang Li, Degen Huang |
NLPCC | 3 |
| 2017 | Drug-drug interaction extraction from biomedical literature using support vector machine and long short term memory networks
Degen Huang, Zhenchao Jiang, Lishuang Li |
Inf. Sci. | 1 |
| 2016 | Biomedical event extraction via Long Short Term Memory networks along dynamic extended treeabstractExtracting knowledge from unstructured text is one of the most important goals of Natural Language Processing, especially in biomedical event extraction domain. In this paper, we describe a system for extracting biomedical events among biotope and bacteria from biomedical literature, using the corpus from the BioNLP'16 Shared Task on Bacteria Biotope task. The current mainstream methods for event extraction are based on shallow machine learning methods. However, these methods mainly rely on domain experience and need enormous manual efforts to select features. Therefore, we propose a novel Long Short Term Memory (LSTM) Networks framework DETBLSTM for event extraction. In our framework, a dynamic extended tree is introduced as the input instead of the original sentences, which utilizes the syntactic information. Furthermore, the POS and distance embeddings are added to enrich input information and thus the complex feature extraction can be skipped. In final, we construct a bidirectional LSTM model to extract biomedical events and achieve 57.14% F-score in the test set. Our model obtains a better F-score than all official submissions to BioNLP-ST 2016, which is 1.34% higher than the best system. Lishuang Li, Jieqiong Zheng, Degen Huang, Xiaohui Lin 0002 |
BIBM | 4 |
| 2016 | An Unsupervised Graph Based Continuous Word Representation Method for Biomedical Text MiningabstractIn biomedical text mining tasks, distributed word representation has succeeded in capturing semantic regularities, but most of them are shallow-window based models, which are not sufficient for expressing the meaning of words. To represent words using deeper information, we make explicit the semantic regularity to emerge in word relations, including dependency relations and context relations, and propose a novel architecture for computing continuous vector representation by leveraging those relations. The performance of our model is measured on word analogy task and Protein-Protein Interaction Extraction (PPIE) task. Experimental results show that our method performs overall better than other word representation models on word analogy task and have many advantages on biomedical text mining. Zhenchao Jiang, Lishuang Li, Degen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Extracting Biomedical Event with Dual Decomposition Integrating Word EmbeddingsabstractExtracting biomedical event from literatures has attracted much attention recently. By now, most of the state-of-the-art systems have been based on pipelines which suffer from cascading errors, and the words encoded by one-hot are unable to represent the semantic information. Joint inference with dual decomposition and novel word embeddings are adopted to address the two problems, respectively, in this work. Word embeddings are learnt from large scale unlabeled texts and integrated as an unsupervised feature into other rich features based on dependency parse graphs to detect triggers and arguments. The proposed system consists of four components: trigger detector, argument detector, jointly inference with dual decomposition, and rule-based semantic post-processing, and outperforms the state-of-the-art systems. On the development set of BioNLP'09, the F-score is 59.77 percent on the primary task, which is 0.96 percent higher than the best system. On the test set of BioNLP'11, the F-score is 56.09 and 0.89 percent higher than the best published result that do not adopt additional techniques. On the test set of BioNLP'13, the F-score reaches 53.19 percent which is 2.22 percent higher than the best result. Lishuang Li, Meiyue Qin, Degen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2015 | Learning Bilingual Sentiment Word Embeddings for Cross-language Sentiment ClassificationabstractHuiWei Zhou, Long Chen, Fulin Shi, Degen Huang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Huiwei Zhou, Long Chen 0019, Fulin Shi, Degen Huang |
ACL (1) | 4 |
| 2015 | Training word embeddings for deep learning in biomedical text mining tasksabstractMost word embedding methods are proposed with general purpose which take a word as a basic unit and learn embeddings according to words' external contexts. However, in biomedical text mining, there are many biomedical entities and syntactic chunks which contain rich domain information, and the semantic meaning of a word is also strongly related to those information. Hence, we present a biomedical domain-specific word embedding model by incorporating stem, chunk and entity to train word embeddings. We also present two deep learning architectures respectively for two biomedical text mining tasks, by which we evaluate our word embeddings and compare them with other models. Experimental results show that our biomedical domain-specific word embeddings overall outperform other general-purpose word embeddings in these deep learning methods for biomedical text mining tasks. Zhenchao Jiang, Lishuang Li, Degen Huang, Liuke Jin |
BIBM | 3 |
| 2015 | Biomedical named entity recognition based on extended Recurrent Neural NetworksabstractBiomedical named entity recognition (bio-NER), which extracts important entities such as genes and proteins, has become one of the most fundamental tasks in biomedical knowledge acquisition. However, the performance of traditional NER systems is always limited to the construction of complex hand-designed features which are derived from various linguistic analyses and maybe only adapted to specified area. In this paper we mainly focus on building a simple and efficient system for bio-NER with the extended Recurrent Neural Network (RNN) which considers the predicted information from the prior node and external context information (topical information & clustering information). Extracting complex hand-designed features is skipped and replaced with word embeddings. The experiments conducted on the BioCreative II GM data set demonstrate RNN models outperform CRF model and deep neural networks (DNN); furthermore, the extended RNN model performs better than the original RNN model. Lishuang Li, Liuke Jin, Zhenchao Jiang, Dingxin Song, Degen Huang |
BIBM | 5 |
| 2015 | Protein-protein interaction extraction based on actively transfer learningabstractIn this paper, we present an actively transfer learning framework to extract PPI. Experimental results show that the proposed ActTrAdaBoost method performs much better than the baseline SVM and the original transfer learning method. In PPIE transfer learning task, our ActTrAdaBoost method presents better performance. Lishuang Li, Jieqiong Zheng, Dingxin Song, Degen Huang |
BIBM | 5 |
| 2015 | A distributed meta-learning system for Chinese entity relation extraction
Lishuang Li, Jing Zhang 0028, Liuke Jin, Degen Huang |
Neurocomputing | 5 |
| 2014 | Improving Kernel-based protein-protein interaction extraction by unsupervised word representationabstractAs an important branch of biomedical information extraction, Protein-Protein Interaction extraction (PPIe) from biomedical literatures has been widely researched, and machine learning methods have achieved great success for this task. However, the word feature generally adopted in the existing methods suffers badly from vocabulary gap and data sparseness, weakening the classification performance. In this paper, the unsupervised word representation approach is introduced to address these problems. Three word representation methods are adopted to improve the performance of PPIe: distributed representation, vector clustering and Brown clusters representation. Experimental results show that our method outperforms the state-of-the-art methods on five publicly available corpora. Lishuang Li, Zhenchao Jiang, Degen Huang |
BIBM | 4 |
| 2014 | A general instance representation architecture for protein-protein interaction extractionabstractPrevious researches have shown that supervised Protein-Protein Interaction Extraction (PPIE) can get high accuracies with elaborately selected features and kernels. However, most features and kernels rest upon domain knowledge and natural language analysis, which makes the supervised model expensive, heavy and brittle. Moreover, the one-hot encoding, a commonly used representation technique, fails to capture the semantic similarity between words. To reduce the manual labor and overcome the shortage of one-hot encoding, we put forward a general instance representation architecture for PPIE, which integrates word representation and vector composition. Our method obtains F-scores of 69.4%, 78.8%, 76.0%, 74.0% and 81.1% on AIMed, BioInfer, HPRD50, IEPA and LLL respectively. Lishuang Li, Zhenchao Jiang, Degen Huang |
BIBM | 3 |
| 2014 | Coreference resolution in biomedical textsabstractCoreference resolution recently plays a more and more important role for many natural language processing tasks. In this paper, we propose two methods for the biomedical coreference resolution. One is the single machine learning method (SVM ranker-learning algorithm) which selects appropriate features for the pronoun and noun phrase coreference resolution respectively. The other one is the hybrid method which adopts the rule-based method or the machine learning method for relative pronouns, non-relative pronouns and noun phrases coreference resolution respectively. Experiments are carried out on Biomedical Natural Language Process Shared Task (BioNLP-ST)12011 coreference resolution corpus. In the first method (the single machine learning method), the F-score is 49.36%, higher than that using the same method with the features in the Reconcile system by 10.06%. In the second method (the hybrid method), the F-score is 68.61%, higher than that of the currently best system by 1.21%. Lishuang Li, Liuke Jin, Zhenchao Jiang, Jing Zhang 0028, Degen Huang |
BIBM | 5 |
| 2014 | The Protein-Protein Interaction extraction based on full textsabstractProtein-Protein Interaction (PPI) extraction from literatures is becoming a more and more significant task in the biomedical information extraction. Though many methods for PPI extraction have achieved promising results, they all concentrated on the abstracts of literatures rather than full texts. In this paper, we append full-text features, namely Location and Co-occurrence to extract PPIs from full texts. Location describes where the protein pair appears in the article. Co-occurrence is the frequency of each protein pair occurring in the article. In addition, syntactic patterns are extracted as features, and then feature selection is applied to improve the performance and reduce the dimension of feature vectors in SVM. Finally, the selected features are combined with two-level DET tree kernel. Experimental results show that the presented approach can achieve an F-score of 74.46% and an AUC of 78.50%. Lishuang Li, Liuke Jin, Jieqiong Zheng, Degen Huang |
BIBM | 5 |
| 2014 | Cross-Lingual Sentiment Classification Based on Denoising Autoencoder
Huiwei Zhou, Long Chen 0019, Degen Huang |
NLPCC | 3 |
| 2013 | A Two-Phase Bio-NER System Based on Integrated Classifiers and Multiagent StrategyabstractBiomedical named entity recognition (Bio-NER) is a fundamental step in biomedical text mining. This paper presents a two-phase Bio-NER model targeting at JNLPBA task. Our two-phase method divides the task into two subtasks: named entity detection (NED) and named entity classification (NEC). The NED subtask is accomplished based on the two-layer stacking method in the first phase, where named entities (NEs) are distinguished from nonnamed-entities (NNEs) in biomedical literatures without identifying their types. Then six classifiers are constructed by four toolkits (CRF++, YamCha, maximum entropy, Mallet) with different training methods and integrated based on the two-layer stacking method. In the second phase for the NEC subtask, the multiagent strategy is introduced to determine the correct entity type for entities identified in the first phase. The experiment results show that the presented approach can achieve an F-score of 76.06 percent, which outperforms most of the state-of-the-art systems. Lishuang Li, Wenting Fan, Degen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2012 | Boosting performance of gene mention tagging system by hybrid methods
Lishuang Li, Wenting Fan, Degen Huang, Yanzhong Dang |
J. Biomed. Informatics | 3 |
| 2011 | POS Tagging of English Particles for Machine Translation
Degen Huang, Wenfeng Sheng |
MTSummit | 2 |
| 2011 | Chinese New Word Identification: A Latent Discriminative Model with Global Features
Xiao Sun 0003, Degen Huang, Fuji Ren |
J. Comput. Sci. Technol. | 2 |
| 2011 | Mining English-Chinese Named Entity Pairs from Comparable CorporaabstractBilingual Named Entity (NE) pairs are valuable resources for many NLP applications. Since comparable corpora are more accessible, abundant and up-to-date, recent researches have concentrated on mining bilingual lexicons using comparable corpora. Leveraging comparable corpora, this research presents a novel approach to mining English-Chinese NE translations by combining multi-dimension features from various information sources for every possible NE pair, which include the transliteration model, English-Chinese matching, Chinese-English matching, translation model, length, and context vector. These features are integrated into one model with linear combination and minimum sample risk (MSR) algorithm. As for the high type-dependence of NE translation, we integrate different features according to different NE types. We experiment with the above individual feature or integrated features to mine person NE (PN) pairs, location NE (LN) pairs and organization NE (ON) pairs. When using transliteration and length to mine PN pairs, we achieve the best performance of 84.9% ( F -score). The LN pairs can be mined with the features of transliteration model, length, translation model, English-Chinese matching and Chinese-English matching. And the best performance is 83.4% ( F -score). The ON pairs can be mined with the features of English-Chinese matching and Chinese-English matching. It reaches the best performance with 84.1% ( F -score). Lishuang Li, Degen Huang, Lian Zhao |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2008 | HMM and CRF Based Hybrid Model for Chinese Lexical Analysis
Degen Huang, Xiao Sun 0003, Shidou Jiao, Lishuang Li, Zhuoye Ding, Ru Wan |
IJCNLP | 1 |
| 2005 | Chinese Syntactic Category Disambiguation Using Support Vector Machines
Lishuang Li, Lihua Li 0006, Degen Huang, Heping Song |
ISNN (2) | 3 |
| 2004 | Identifying Pronunciation-Translated Names from Chinese Texts Based on Support Vector Machines
Lishuang Li, Chunrong Chen, Degen Huang, Yuansheng Yang |
ISNN (1) | 3 |