Qingcai Chen

dblp:15/1052 · DBLP profile ↗
← Back
117ranked-venue papers
3as first author
53since 2021 · last 2026
0000-0001-8473-7293ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 2 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 FSE: Continual learning for named entity recognition by fast-slow experts
Yunan Zhang 0003, Xiangping Wu 0001, Qingcai Chen
Pattern Recognit. Lett.5
2025 Reasoning Graph Enhanced Exemplars Retrieval for In-Context Learning
abstract
Large language models (LLMs) have exhibited remarkable few-shot learning capabilities and unified the paradigm of NLP tasks through the in-context learning (ICL) technique. Despite the success of ICL, the quality of the exemplar demonstrations can significantly influence the LLM’s performance. Existing exemplar selection methods mainly focus on the semantic similarity between queries and candidate exemplars. On the other hand, the logical connections between reasoning steps can also be beneficial to depict the problem-solving process. This paper proposes a novel method named Reasoning Graph-enhanced Exemplar Retrieval (RGER). RGER first queries LLM to generate an initial response and then expresses intermediate problem-solving steps to a graph structure. After that, it employs a graph kernel to select exemplars with semantic and structural similarity. Extensive experiments demonstrate the structural relationship is helpful to the alignment of queries and candidate exemplars. The efficacy of RGER on mathematics and logical reasoning tasks showcases its superiority over state-of-the-art retrieval-based approaches.
Yukang Lin, Bingchen Zhong, Shuoran Jiang, Joanna Siebert, Qingcai Chen
COLING5
2025 STPformer: Mutation-Aware Spatial-Temporal Pivotal Attention Networks for Transformer-Based Traffic Forecasting
Hongyang Su, Chenyun Yu, Qingcai Chen, Beibei Kong, Lei Cheng 0005, Chengxiang Zhuo, Zang Li, Xiaolong Wang 0001
DASFAA (1)3
2025 Local and Global Aware Document Image Enhancement with Residual Denoising Diffusion Model
abstract
In document image enhancement scenarios, due to the limitations of high computational complexity caused by high-resolution input images, current methods often process these original degraded images by cropping them into patches of specified sizes. However, previous approaches that solely rely on cropped patches or merely use the document enhancement result of the global image as a reference are difficult to fully utilize global image information. This limitation often results in inconsistent enhancement effects across different regions of the same image. In this paper, we introduce LGA-Doc, a novel two-stage local-global information aware generative framework for document image enhancement. Our approach employs a context-aware image feature fusion module that facilitates feature interaction between local document patches and the global image, enabling deep integration of multi-granularity information. The experimental results demonstrate that our method achieves state-of-the-art performance on both the deblurring dataset and the binarization evaluation dataset. Ablation studies further validate the effectiveness of our local-global information aware module.
Hongrui Tie, Heng Li 0014, Xiangping Wu 0001, Qingcai Chen
ICMR4
2025 CLSurCoder: LLMs Based Cross-Lingual Transfer Learning for Low-Resource Language Surgical Records Coding
Dawen Chu, Fen Yang, Bingchen Zhong, Xiangping Wu 0001, Shuoran Jiang, Qingcai Chen
NLPCC (2)9
2024 Discriminative Forests Improve Generative Diversity for Generative Adversarial Networks
abstract
Improving the diversity of Artificial Intelligence Generated Content (AIGC) is one of the fundamental problems in the theory of generative models such as generative adversarial networks (GANs). Previous studies have demonstrated that the discriminator in GANs should have high capacity and robustness to achieve the diversity of generated data. However, a discriminator with high capacity tends to overfit and guide the generator toward collapsed equilibrium. In this study, we propose a novel discriminative forest GAN, named Forest-GAN, that replaces the discriminator to improve the capacity and robustness for modeling statistics in real-world data distribution. A discriminative forest is composed of multiple independent discriminators built on bootstrapped data. We prove that a discriminative forest has a generalization error bound, which is determined by the strength of individual discriminators and the correlations among them. Hence, a discriminative forest can provide very large capacity without any risk of overfitting, which subsequently improves the generative diversity. With the discriminative forest framework, we significantly improved the performance of AutoGAN with a new record FID of 19.27 from 30.71 on STL10 and improved the performance of StyleGAN2-ADA with a new record FID of 6.87 from 9.22 on LSUN-cat.
Junjie Chen 0004, Qingcai Chen, Hongchang Gao, Wendy Hui Wang, Zenglin Xu, Xinghua Shi
AAAI5
2024 TDeLTA: A Light-Weight and Robust Table Detection Method Based on Learning Text Arrangement
abstract
The diversity of tables makes table detection a great challenge, leading to existing models becoming more tedious and complex. Despite achieving high performance, they often overfit to the table style in training set, and suffer from significant performance degradation when encountering out-of-distribution tables in other domains. To tackle this problem, we start from the essence of the table, which is a set of text arranged in rows and columns. Based on this, we propose a novel, light-weighted and robust Table Detection method based on Learning Text Arrangement, namely TDeLTA. TDeLTA takes the text blocks as input, and then models the arrangement of them with a sequential encoder and an attention module. To locate the tables precisely, we design a text-classification task, classifying the text blocks into 4 categories according to their semantic roles in the tables. Experiments are conducted on both the text blocks parsed from PDF and extracted by open-source OCR tools, respectively. Compared to several state-of-the-art methods, TDeLTA achieves competitive results with only 3.1M model parameters on the large-scale public datasets. Moreover, when faced with the cross-domain data under the 0-shot setting, TDeLTA outperforms baselines by a large margin of nearly 7%, which shows the strong robustness and transferability of the proposed model.
Xiangping Wu 0001, Qingcai Chen, Heng Li 0014, Zhixiang Cai, Qitian Wu
AAAI3
2024 ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order Optimization
abstract
Lowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants.
Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang 0003, Yukang Lin, Xiangping Wu 0001, Chuanyi Liu, Xiaobao Song
AAAI2
2024 Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language Models
abstract
Ensuring robustness is especially important when AI is deployed in responsible or safety-critical environments. ChatGPT can perform brilliantly in both adversarial and out-of-distribution (OOD) robustness, while other popular large language models (LLMs), like LLaMA-2, ERNIE and ChatGLM, do not perform satisfactorily in this regard. Therefore, it is valuable to study what efforts play essential roles in ChatGPT, and how to transfer these efforts to other LLMs. This paper experimentally finds that linguistic rule induction is the foundation for identifying the cause-effect relationships in LLMs. For LLMs, accurately processing the cause-effect relationships improves its adversarial and OOD robustness. Furthermore, we explore a low-cost way for aligning LLMs with linguistic rules. Specifically, we constructed a linguistic rule instruction dataset to fine-tune LLMs. To further energize LLMs for reasoning step-by-step with the linguistic rule, we construct the task-relevant LingR-based chain-of-thoughts. Experiments showed that LingR-induced LLaMA-13B achieves comparable or better results with GPT-3.5 and GPT-4 on various adversarial and OOD robustness evaluations.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Yukang Lin
LREC/COLING2
2024 PVEIN: A Pretrained Vertex Embedding Infer Network for Open-Domain Question Answer Scoring
Kai Chen 0020, Yingping Deng, Qingcai Chen
ICONIP (9)3
2024 Attention based adaptive spatial-temporal hypergraph convolutional networks for stock price trend prediction
Hongyang Su, Xiaolong Wang 0001, Yang Qin 0001, Qingcai Chen
Expert Syst. Appl.4
2024 Confounder balancing in adversarial domain adaptation for pre-trained large models fine-tuning
abstract
The excellent generalization, contextual learning, and emergence abilities in the pre-trained large models (PLMs) handle specific tasks without direct training data, making them the better foundation models in the adversarial domain adaptation (ADA) methods to transfer knowledge learned from the source domain to target domains. However, existing ADA methods fail to account for the confounder properly, which is the root cause of the source data distribution that differs from the target domains. This study proposes a confounder balancing method in adversarial domain adaptation for PLMs fine-tuning (CadaFT), which includes a PLM as the foundation model for a feature extractor, a domain classifier and a confounder classifier, and they are jointly trained with an adversarial loss. This loss is designed to improve the domain-invariant representation learning by diluting the discrimination in the domain classifier. At the same time, the adversarial loss also balances the confounder distribution among source and unmeasured domains in training. Compared to newest ADA methods, CadaFT can correctly identify confounders in domain-invariant features, thereby eliminating the confounder biases in the extracted features from PLMs. The confounder classifier in CadaFT is designed as a plug-and-play and can be applied in the confounder measurable, unmeasurable, or partially measurable environments. Empirical results on natural language processing and computer vision downstream tasks show that CadaFT outperforms the newest GPT-4, LLaMA2, ViT and ADA methods.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001, Yukang Lin
Neural Networks2
2024 BaSFormer: A Balanced Sparsity Regularized Attention Network for Transformer
abstract
Attention networks often make decisions relying solely on a few pieces of tokens, even if those reliances are not truly indicative of the underlying meaning or intention of the full context. This can lead to over-fitting in transformers and hinder their ability to generalize. Attention regularization and sparsity-based methods have been used to overcome this issue. However, these methods cannot guarantee that all tokens have sufficient receptive fields for global information inference. Thus, the impact of individual biases cannot be effectively reduced. As a result, the generalization of these approaches improved slightly from the training data to new data. To address these limitations, we propose a balanced sparsity (BaS) regularized attention network on top of the transformers, called BaSFormer. BaS regularization introduces the K-regular graph constraint on self-attention connections, which replaces SoftMax with SparseMax in the attention transformation. In BaS-regularized self-attention, SparseMax assigns zero attention scores to low-scoring connections, highlighting influential and meaningful contexts. The K-regular graph constraint ensures that all tokens have an equal-sized receptive field to aggregate information, which facilitates the involvement of global tokens in the feature update of each layer and reduces the impact of individual biases. Given that there is no continuous loss can be used for the K-regular graph regularization, we propose an exponential extremum loss with an augmented Lagrangian function. The experimental results showed that BaSFormer improved the effectiveness of debiasing compared to that of the newest LLMs, such as the GPT-3.5, GPT-4 and LLaMA. In addition, BaSFormer achieves new state-of-the-art (SOTA) results in text generation tasks. Interestingly, this work also shows that BaSFormer can learn hierarchical linguistic dependencies in gradient attributions, which improves interpretability and adversarial robustness.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Learning to Improve Out-of-Distribution Generalization via Self-Adaptive Language Masking
abstract
Although the pre-trained Transformers learned general linguistic knowledge from large-scale corpus, they still over-fit on the lexical biases when fine-tuning on specific datasets. This problem limits the generalizability of pre-trained models, particularly when learning over out-of-distribution (OOD) data. To address this issue, this paper proposes a self-adaptive language masking (AdaLMask) paradigm to fine-tune the pre-trained Transformers. AdaLMask obviates lexical biases by eliminating the dependence on semantically inessential words. Specifically, AdaLMask learns a Gumbel-Softmax distribution to determine the desired masking positions, and the distribution parameters are optimized via a representation-invariant (RInv) objective to ensure the masked positions are semantically lossless. Four natural language processing tasks are chosen to evaluate the effectiveness of the proposed method on the robustness of lexical biases and OOD generalization. All empirical results demonstrate that the AdaLMask paradigm substantially improves the OOD generalization of pre-trained Transformers.
Shuoran Jiang, Youcheng Pan, Qingcai Chen, Yang Xiang 0003, Xiangping Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 BioPRO: Context-Infused Prompt Learning for Biomedical Entity Linking
abstract
Recent research tends to address the biomedical entity linking problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing thevarietychallenge of the heterogeneous naming of biomedical concepts. Yet, theambiguitychallenge that the same word under different contexts can be used to refer to distinct concepts is usually ignored. To address this challenge, we propose BioPRO, a two-stage entity linking algorithm to enhance the biomedical entity representations based on context-infused prompt learning. The first stage includes a coarse-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity's surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a fine-grained encoder based on prompt-tuning that sufficiently stimulates knowledge in contextual information of mentions and entities. Furthermore, the trained fine-grained encoder can be utilized to generate deep representations of bio-entities and boost candidate retrieval in the first stage. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art (SOTA) techniques on 4 biomedical corpora. We also observe by cases that the proposed context-infused prompt-tuning strategy is effective in solving both thevarietyandambiguitychallenges in the linking task.
Tiantian Zhu 0002, Yang Qin 0001, Ming Feng, Qingcai Chen, Baotian Hu, Yang Xiang 0003
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 A Neural Span-Based Continual Named Entity Recognition Model
abstract
Named Entity Recognition (NER) models capable of Continual Learning (CL) are realistically valuable in areas where entity types continuously increase (e.g., personal assistants). Meanwhile the learning paradigm of NER advances to new patterns such as the span-based methods. However, its potential to CL has not been fully explored. In this paper, we propose SpanKL, a simple yet effective Span-based model with Knowledge distillation (KD) to preserve memories and multi-Label prediction to prevent conflicts in CL-NER. Unlike prior sequence labeling approaches, the inherently independent modeling in span and entity level with the designed coherent optimization on SpanKL promotes its learning at each incremental step and mitigates the forgetting. Experiments on synthetic CL datasets derived from OntoNotes and Few-NERD show that SpanKL significantly outperforms previous SoTA in many aspects, and obtains the smallest gap from CL to the upper bound revealing its high practiced value. The code is available at https://github.com/Qznan/SpanKL.
Yunan Zhang 0003, Qingcai Chen
AAAI2
2023 FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-Training
abstract
Fashion vision-language pre-training models have shown efficacy for a wide range of downstream tasks. However, general vision-language pre-training models pay less attention to fine-grained domain features, while these features are important in distinguishing the specific domain tasks from general tasks. We propose a method for fine-grained fashion vision-language pre-training based on fashion Symbols and Attributes Prompt (FashionSAP) to model fine-grained multi-modalities fashion attributes and characteristics. Firstly, we propose the fashion symbols, a novel abstract fashion concept layer, to represent different fashion items and to generalize various kinds of fine- grained fashion features, making modelling fine-grained attributes more effective. Secondly, the attributes prompt method is proposed to make the model learn specific attributes of fashion items explicitly. We design proper prompt templates according to the format of fashion data. Comprehensive experiments are conducted on two public fashion benchmarks, i.e., FashionGen and FashionIQ, and FashionSAP gets SOTA performances for four popular fashion tasks. The ablation study also shows the proposed abstract fashion symbols, and the attribute prompt method enables the model to acquire fine-grained semantics in the fashion domain effectively. The obvious performance gains from FashionSAP provide a new baseline for future fashion task research.11The source code is available at https://github.com/hssip/FashionSAP
Yunpeng Han, Lisai Zhang, Qingcai Chen, Jianxin Yang, Zhao Cao
CVPR3
2023 Controllable Contrastive Generation for Multilingual Biomedical Entity Linking
abstract
Multilingual biomedical entity linking (MBEL) aims to map language-specific mentions in the biomedical text to standardized concepts in a multilingual knowledge base (KB) such as Unified Medical Language System (UMLS).In this paper, we propose Con2GEN, a prompt-based controllable contrastive generation framework for MBEL, which summarizes multidimensional information of the UMLS concept mentioned in biomedical text into a natural sentence following a predefined template.Instead of tackling the MBEL problem with a discriminative classifier, we formulate it as a sequence-tosequence generation task, which better exploits the shared dependencies between source mentions and target entities.Moreover, Con2GEN matches against UMLS concepts in as many languages and types as possible, hence facilitating cross-information disambiguation.Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the XL-BEL and the Mantra GSC datasets spanning 12 typologically diverse languages.
Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Xin Mu, Changlong Yu, Yang Xiang 0003
EMNLP3
2023 Foreground and Text-lines Aware Document Image Rectification
abstract
This paper aims at the distorted document image rectification problem, the objective to eliminate the geometric distortion in the document images and realize document intelligence. Improving the readability of distorted documents is crucial to effectively extract information from deformed images. According to our observations, the foreground and text-line of the original warped image can represent the deformation tendency. However, previous distorted image rectification methods pay little attention to the readability of the warped paper. In this paper, we focus on the foreground and text-line regions of distorted paper and proposes a global and local fusion method to improve the rectification effect of distorted images and enhance the readability of document images. We introduce cross attention to capture the features of the foreground and text-lines in the warped document and effectively fuse them. The proposed method is evaluated quantitatively and qualitatively on the public DocUNet benchmark and DIR300 Dataset, which achieve state-of-the-art performances. Experimental analysis shows the proposed method can well perform overall geometric rectification of distorted images and effectively improve document readability (using the metrics of Character Error Rate and Edit Distance). The code is available at https://github.com/xiaomore/Document-Image-Dewarping.
Heng Li 0014, Xiangping Wu 0001, Qingcai Chen, Qianjin Xiang
ICCV3
2023 Efficient Adaptive Spatial-Temporal Attention Network for Traffic Flow Forecasting
Hongyang Su, Xiaolong Wang 0001, Qingcai Chen, Yang Qin 0001
ECML/PKDD (5)3
2023 Learning to generate complex question with intent prediction from long passage
Youcheng Pan, Baotian Hu, Shiyue Wang, Xiaolong Wang 0001, Qingcai Chen, Zenglin Xu, Min Zhang 0005
Appl. Intell.5
2023 VGbel: An exploration of ensemble learning incorporating non-Euclidean structural representation for time series classification
Shaocong Wu, Mengxia Liang, Xiaolong Wang 0001, Qingcai Chen
Expert Syst. Appl.4
2023 Fine-grained biomedical knowledge negation detection via contrastive learning
Tiantian Zhu 0002, Yang Xiang 0003, Qingcai Chen, Yang Qin 0001, Baotian Hu, Wentai Zhang 0003
Knowl. Based Syst.3
2023 Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network
abstract
Deep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely.
Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002
IEEE Trans. Multim.3
2022 Diaformer: Automatic Diagnosis via Symptoms Sequence Generation
abstract
Automatic diagnosis has attracted increasing attention but remains challenging due to multi-step reasoning. Recent works usually address it by reinforcement learning methods. However, these methods show low efficiency and require task-specific reward functions. Considering the conversation between doctor and patient allows doctors to probe for symptoms and make diagnoses, the diagnosis process can be naturally seen as the generation of a sequence including symptoms and diagnoses. Inspired by this, we reformulate automatic diagnosis as a symptoms Sequence Generation (SG) task and propose a simple but effective automatic Diagnosis model based on Transformer (Diaformer). We firstly design the symptom attention framework to learn the generation of symptom inquiry and the disease diagnosis. To alleviate the discrepancy between sequential generation and disorder of implicit symptoms, we further design three orderless training mechanisms. Experiments on three public datasets show that our model outperforms baselines on disease diagnosis by 1%, 6% and 11.5% with the highest training efficiency. Detailed analysis on symptom inquiry prediction demonstrates that the potential of applying symptoms sequence generation for automatic diagnosis.
Dongfang Li 0002, Qingcai Chen, Wenxiu Zhou, Xin Liu 0054
AAAI3
2022 Unifying Model Explainability and Robustness for Joint Text Classification and Rationale Extraction
abstract
Recent works have shown explainability and robustness are two crucial ingredients of trustworthy and reliable text classification. However, previous works usually address one of two aspects: i) how to extract accurate rationales for explainability while being beneficial to prediction; ii) how to make the predictive model robust to different types of adversarial attacks. Intuitively, a model that produces helpful explanations should be more robust against adversarial attacks, because we cannot trust the model that outputs explanations but changes its prediction under small perturbations. To this end, we propose a joint classification and rationale extraction model named AT-BMC. It includes two key mechanisms: mixed Adversarial Training (AT) is designed to use various perturbations in discrete and embedding space to improve the model’s robustness, and Boundary Match Constraint (BMC) helps to locate rationales more precisely with the guidance of boundary information. Performances on benchmark datasets demonstrate that the proposed AT-BMC outperforms baselines on both classification and rationale extraction by a large margin. Robustness analysis shows that the proposed AT-BMC decreases the attack success rate effectively by up to 69%. The results indicate that there are connections between robust models and better explanations.
Dongfang Li 0002, Baotian Hu, Qingcai Chen, Tujie Xu, Jingcong Tao, Yunan Zhang 0003
AAAI3
2022 CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
abstract
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ningyu Zhang 0001, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li 0040, Xin Shang, Kangping Yin, Chuanqi Tan, Fei Huang 0002, Luo Si, Yuan Ni, Guo Tong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan 0002, Jun Yan 0010, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen
ACL (1)23
2022 Contrastive Label Correlation Enhanced Unified Hashing Encoder for Cross-modal Retrieval
abstract
Cross-modal hashing (CMH) has been widely used in multimedia retrieval applications for its low storage cost and fast indexing speed. Thanks to the success of deep learning, cross-modal hashing has made significant progress with high-quality deep features. However, the modal gap is still a crucial bottleneck for existing cross-modal hashing methods: the commonly used convolutional neural network and bag-of-words encoders are customized for single modal prior, limiting the models to learn semantics representation in a cross-modal space. To overcome modality heterogeneity, we propose a shared transformer encoder (UniHash) to unify the cross-modal hashing into the same semantic space. A contrastive label correlation learning (CLC) loss using the category labels as modality bridge is designed together to improve the representation quality. Moreover, we take advantage of the multi-hot label space and propose a negative label generation (NegLG) strategy to get richer and uniformly distributed negative labels for contrast. Extensive experiments on three benchmarks verify the advantage of our proposed method. Besides, the proposed UniHash outperforms state-of-the-art cross-modal hashing methods significantly, establishing a new important baseline for the cross-modal hashing research. Codes are released github.com/idealwhite/Unihash.
Hongfa Wu, Lisai Zhang, Qingcai Chen, Yimeng Deng, Joanna Siebert, Yunpeng Han, Dejiang Kong, Zhao Cao
CIKM3
2022 Prompt-based Text Entailment for Low-Resource Named Entity Recognition
abstract
Pre-trained Language Models (PLMs) have been applied in NLP tasks and achieve promising results. Nevertheless, the fine-tuning procedure needs labeled data of the target domain, making it difficult to learn in low-resource and non-trivial labeled scenarios. To address these challenges, we propose Prompt-based Text Entailment (PTE) for low-resource named entity recognition, which better leverages knowledge in the PLMs. We first reformulate named entity recognition as the text entailment task. The original sentence with entity type-specific prompts is fed into PLMs to get entailment scores for each candidate. The entity type with the top score is then selected as final label. Then, we inject tagging labels into prompts and treat words as basic units instead of n-gram spans to reduce time complexity in generating candidates by n-grams enumeration. Experimental results demonstrate that the proposed method PTE achieves competitive performance on the CoNLL03 dataset, and better than fine-tuned counterparts on the MIT Movie and Few-NERD dataset in low-resource settings.
Dongfang Li 0002, Baotian Hu, Qingcai Chen
COLING3
2022 Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence Estimates
abstract
Calibration strengthens the trustworthiness of black-box models by producing better accurate confidence estimates on given examples.However, little is known about if model explanations can help confidence calibration.Intuitively, humans look at important features attributions and decide whether the model is trustworthy.Similarly, the explanations can tell us when the model may or may not know.Inspired by this, we propose a method named CME that leverages model explanations to make the model less confident with non-inductive attributions.The idea is that when the model is not highly confident, it is difficult to identify strong indications of any class, and the tokens accordingly do not have high attribution scores for any class and vice versa.We conduct extensive experiments on six datasets with two popular pre-trained language models in the in-domain and out-of-domain settings.The results show that CME improves calibration performance in all settings.The expected calibration errors are further reduced when combined with temperature scaling.Our findings highlight that model explanations can help calibrate posterior estimates.
Dongfang Li 0002, Baotian Hu, Qingcai Chen
EMNLP3
2022 End-to-End ASR-Enhanced Neural Network for Alzheimer's Disease Diagnosis
abstract
This paper presents an approach to Alzheimer’s disease (AD) diagnosis from spontaneous speech using an end-to-end ASR-enhanced neural network. Under the condition that only audio data are provided and accurate transcripts are unavailable, this paper proposes a system that can analyze utterances to differentiate between AD patients, healthy controls, and individuals with mild cognitive impairment. The ASR-enhanced model comprises automatic speech recognition (ASR) with an encoder-decoder structure and the encoder followed by an AD classification network. The encoder takes a Mel spectrogram as input and transforms it into high-level acoustic features that correlate with AD. The classification network then maps intermediate acoustic features to three categories. In the training phase, the AD classification and speech recognition tasks are optimized simultaneously. Experimental results obtained from an AD recognition dataset of Chinese spontaneous speech1illustrate the effectiveness of integrating ASR into AD diagnosis in an end-to-end manner. Further, our model has low dependency on accurate ASR transcripts. This work achieved accuracy scores of 89.1% and 82.6% for long and short utterance tracks, respectively.
Jiancheng Gui, Kai Chen 0020, Joanna Siebert, Qingcai Chen
ICASSP5
2022 Enhancing Entity Representations with Prompt Learning for Biomedical Entity Linking
abstract
Biomedical entity linking aims to map mentions in biomedical text to standardized concepts or entities in a curated knowledge base (KB) such as Unified Medical Language System (UMLS). The latest research tends to solve this problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing the variety challenge of the heterogeneous naming of biomedical concepts. Yet, the ambiguity challenge that the same word under different contexts may refer to distinct entities is usually ignored. To address this challenge, we propose a two-stage linking algorithm to enhance the entity representations based on prompt learning. The first stage includes a coarser-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity’s surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a finer-grained encoder based on prompt-tuning that utilizes the contextual information. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the largest biomedical public dataset MedMentions and the NCBI disease corpus. We also observe by cases that the proposed prompt-tuning strategy is effective in solving both the variety and ambiguity challenges in the linking task.
Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Baotian Hu, Yang Xiang 0003
IJCAI3
2022 Medical Dialogue Response Generation with Pivotal Information Recalling
abstract
Medical dialogue generation is an important yet challenging task. Most previous works rely on the attention mechanism and large-scale pretrained language models. However, these methods often fail to acquire pivotal information from the long dialogue history to yield an accurate and informative response, due to the fact that the medical entities usually scatters throughout multiple utterances along with the complex relationships between them. To mitigate this problem, we propose a medical response generation model with Pivotal Information Recalling (MedPIR), which is built on two components, i.e., knowledge-aware dialogue graph encoder and recall-enhanced generator. The knowledge-aware dialogue graph encoder constructs a dialogue graph by exploiting the knowledge relationships between entities in the utterances, and encodes it with a graph attention network. Then, the recall-enhanced generator strengthens the usage of these pivotal information by generating a summary of the dialogue before producing the actual response. Experimental results on two large-scale medical dialogue datasets show that MedPIR outperforms the strong baselines in BLEU scores and medical entities F1 measure.
Yu Zhao 0043, Yunxin Li, Yuxiang Wu, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001, Min Zhang 0005
KDD5
2022 PreRBP-TL: prediction of species-specific RNA-binding proteins based on transfer learning
abstract
MOTIVATION: RNA-binding proteins (RBPs) play crucial roles in post-transcriptional regulation. Accurate identification of RBPs helps to understand gene expression, regulation, etc. In recent years, some computational methods were proposed to identify RBPs. However, these methods fail to accurately identify RBPs from some specific species with limited data, such as bacteria. RESULTS: In this study, we introduce a computational method called PreRBP-TL for identifying species-specific RBPs based on transfer learning. The weights of the prediction model were initialized by pretraining with the large general RBP dataset and then fine-tuned with the small species-specific RPB dataset by using transfer learning. The experimental results show that the PreRBP-TL achieves better performance for identifying the species-specific RBPs from Human, Arabidopsis, Escherichia coli and Salmonella, outperforming eight state-of-the-art computational methods. It is anticipated PreRBP-TL will become a useful method for identifying RBPs. AVAILABILITY AND IMPLEMENTATION: For the convenience of researchers to identify RBPs, the web server of PreRBP-TL was established, freely available at http://bliulab.net/PreRBP-TL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Zhang 0078, Ke Yan 0003, Qingcai Chen, Bin Liu 0014
Bioinform.3
2022 Biomedical relation extraction via knowledge-enhanced reading comprehension
abstract
BACKGROUND: In biomedical research, chemical and disease relation extraction from unstructured biomedical literature is an essential task. Effective context understanding and knowledge integration are two main research problems in this task. Most work of relation extraction focuses on classification for entity mention pairs. Inspired by the effectiveness of machine reading comprehension (RC) in the respect of context understanding, solving biomedical relation extraction with the RC framework at both intra-sentential and inter-sentential levels is a new topic worthy to be explored. Except for the unstructured biomedical text, many structured knowledge bases (KBs) provide valuable guidance for biomedical relation extraction. Utilizing knowledge in the RC framework is also worthy to be investigated. We propose a knowledge-enhanced reading comprehension (KRC) framework to leverage reading comprehension and prior knowledge for biomedical relation extraction. First, we generate questions for each relation, which reformulates the relation extraction task to a question answering task. Second, based on the RC framework, we integrate knowledge representation through an efficient knowledge-enhanced attention interaction mechanism to guide the biomedical relation extraction. RESULTS: The proposed model was evaluated on the BioCreative V CDR dataset and CHR dataset. Experiments show that our model achieved a competitive document-level F1 of 71.18% and 93.3%, respectively, compared with other methods. CONCLUSION: Result analysis reveals that open-domain reading comprehension data and knowledge representation can help improve biomedical relation extraction in our proposed KRC framework. Our work can encourage more research on bridging reading comprehension and biomedical relation extraction and promote the biomedical relation extraction.
Baotian Hu, Weihua Peng, Qingcai Chen, Buzhou Tang
BMC Bioinform.4
2022 A stock time series forecasting approach incorporating candlestick patterns and sequence similarity
abstract
This article aims to implement trend forecasting of stock time series based on candlestick patterns and sequence similarity. Financial time series forecasting plays a central role in hedging market risks and optimizing investment portfolios. This is a challenging task, as financial engineering requires the proposed approach to be interpretable, robust, and compatible. It is noted that many published research studies are based on multi-modal data, which makes the prediction approaches increasingly complex, difficult to interpret, and does not allow the migration across different data. Given this situation, it is believed that a candlestick data-based approach is promising. It is already recognized by the technical analyses, prevalent in financial markets, more readily available, and has better interpretability. In this paper, the forecasting approach is divided into two steps. In the first step, sequential pattern mining is used to obtain candlestick patterns from multidimensional candlestick data, and the correlation between different patterns and the corresponding future trends are calculated. In the second step, a new sequence similarity is proposed to match the diverse candlestick sequences with the existing patterns. The method is validated on real data from 800 stocks in the Chinese stock market, which are divided into two groups of experiments, and the average accuracy achieved by the proposed method is 56.04% and 55.56%, which is higher than the SVM model (50.83% and 51.32%) and the LSTM model (50.71% and 50.68%) used for comparison, proving that our work is more stable and accurate. This work is instructive for further research around candlestick data to follow.
Mengxia Liang, Shaocong Wu, Xiaolong Wang 0001, Qingcai Chen
Expert Syst. Appl.4
2022 Leveraging Multi-source knowledge for Chinese clinical named entity recognition via relational graph convolutional network
Yang Xiang 0003, Ka-Chun Wong, Qingcai Chen, Jun Yan 0010, Buzhou Tang
J. Biomed. Informatics5
2022 VLDeformer: Vision-Language Decomposed Transformer for fast cross-modal retrieval
Lisai Zhang, Hongfa Wu, Qingcai Chen, Yimeng Deng, Joanna Siebert, Yunpeng Han, Dejiang Kong, Zhao Cao
Knowl. Based Syst.3
2022 A Unified Machine Reading Comprehension Framework for Cohort Selection
abstract
Cohort selection is an essential prerequisite for clinical research, determining whether an individual satisfies given selection criteria. Previous works for cohort selection usually treated each selection criterion independently and ignored not only the meaning of each selection criterion but the relations among cohort selection criteria. To solve the problems above, we propose a novel unified machine reading comprehension (MRC) framework. In this MRC framework, we design simple rules to generate questions for each criterion from cohort selection guidelines and treat clues extracted by trigger words from patients' medical records as passages. A series of state-of-the-art MRC models based on BiDAF, BIMPM, BERT, BioBERT, NCBI-BERT, and RoBERTa are deployed to determine which question and passage pairs match. We also introduce a cross-criterion attention mechanism on representations of question and passage pairs to model relations among cohort selection criteria. Results on two datasets, that is, the dataset of the 2018 National NLP Clinical Challenge (N2C2) for cohort selection and a dataset from the MIMIC-III dataset, show that our NCBI-BERT MRC model with cross-criterion attention mechanism achieves the highest micro-averaged F1-score of 0.9070 on the N2C2 dataset and 0.8353 on the MIMIC-III dataset. It is competitive to the best system that relies on a large number of rules defined by medical experts on the N2C2 dataset. Comparing these two models, we find that the NCBI-BERT MRC model mainly performs worse on mathematical logic criteria. When using rules instead of the NCBI-BERT MRC model on some criteria regarding mathematical logic on the N2C2 dataset, we obtain a new benchmark with an F1-score of 0.9163, indicating that it is easy to integrate rules into MRC models for improvement.
Weihua Peng, Qingcai Chen, Zhengxing Huang, Buzhou Tang
IEEE J. Biomed. Health Informatics3
2021 Multi-hop Graph Convolutional Network with High-order Chebyshev Approximation for Text Reasoning
abstract
Shuoran Jiang, Qingcai Chen, Xin Liu, Baotian Hu, Lisai Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shuoran Jiang, Qingcai Chen, Xin Liu 0054, Baotian Hu, Lisai Zhang
ACL/IJCNLP (1)2
2021 Leveraging Capsule Routing to Associate Knowledge with Medical Literature Hierarchically
abstract
Integrating knowledge into text is a promising way to enrich text representation, especially in the medical field.However, undifferentiated knowledge not only confuses the text representation but also imports unexpected noises.In this paper, to alleviate this problem, we propose leveraging capsule routing to associate knowledge with medical literature hierarchically (called HiCapsRKL).Firstly, HiCapsRKL extracts two empirically designed text fragments from medical literature and encodes them into fragment representations respectively.Secondly, the capsule routing algorithm is applied to two fragment representations.Through the capsule computing and dynamic routing, each representation is processed into a new representation (denoted as caps-representation), and we integrate the caps-representations as information gain to associate knowledge with medical literature hierarchically.Finally, HiCapsRKL are validated on relevance prediction and medical literature retrieval test sets.The experimental results and analyses show that HiCapsRKL can more accurately associate knowledge with medical literature than the mainstream methods.In summary, HiCapsRKL can efficiently help selecting the most relevant knowledge to the medical literature, which may be an alternative attempt to improve knowledge-based text representation.Source code is released on GitHub 1 .
Xin Liu 0054, Qingcai Chen, Wenxiu Zhou, Tingyu Liu, Xinlan Yang, Weihua Peng
EMNLP (1)2
2021 A Large-Scale Chinese Long-Text Extractive Summarization Corpus
abstract
Recently, large-scale datasets have vastly facilitated the development in nearly domains of Natural Language Processing. However, lacking large scale Chinese corpus is still a critical bottleneck for further research on deep text summarization methods. In this paper, we publish a large-scale Chinese Long-text Extractive Summarization corpus named CLES. The CLES contains about 104Kpairs, which is originally collected from Sina Weibo1. To verify the quality of the corpus, we also manually tagged the relevance score of 5,000pairs. Our benchmark models on the proposed corpus include conventional deep learning based extractive models and several pre-trained Bert-based algorithms. Their performances are reported and briefly analyzed to facilitate further research on the corpus. We will release this corpus for further research2.
Kai Chen 0020, Guanyu Fu, Qingcai Chen, Baotian Hu
ICASSP3
2021 HCAG: A Hierarchical Context-Aware Graph Attention Model for Depression Detection
abstract
Depression is one of the most common mental health disorders, it’s crucial to design an effective and robust model for automatic depression detection (ADD). Although current approaches rely on extra topic models or manually topic-selection procedures which is time-consuming, they still haven’t thoroughly explored the sufficient context information among clinical interviews. In this paper, we propose HCAG, a novel Hierarchical Context-Aware Graph attention model for ADD. Our model mirrors the hierarchical structure of depression assessment and leverages the Graph Attention Network (GAT) to grasp relational contextual information of text/audio modality. Experiments on the DAIC-WOZ dataset show a great performance improvement, with the Fl-score of 0.92, a Mean Absolute Error (MAE) of 2.94, and a Root Mean Square Error (RMSE) of 3.80. To the best of our knowledge, our model outperforms the existing state-of-the-art methods.
Meng Niu, Kai Chen 0020, Qingcai Chen, Lufeng Yang
ICASSP3
2021 Semi-supervised Visual Feature Integration for Language Models through Sentence Visualization
abstract
Integrating visual features has been proved useful for natural language understanding tasks. Nevertheless, most existing multimodal language models highly rely on training on aligned image and text data. In this paper, we propose a novel semi-supervised visual integration framework for pre-trained language models. In the framework, the visual features are obtained through a sentence visualization and vision-language fusion mechanism. The uniqueness includes: 1) the integration is conducted via a semi-supervised framework and does not require aligned images for the processed sentences. 2) the framework works as an auxiliary component, and will not affect the language processing ability of the integrated language model. Experimental results on both natural language inference and reading comprehension tasks demonstrate that our framework improves the strong baseline language models. Considering that our framework only requires an image database, and does not require aligned images for the processed texts, it provides a feasible way for multimodal language learning.
Lisai Zhang, Qingcai Chen, Joanna Siebert, Buzhou Tang
ICMI2
2021 FHTC: Few-Shot Hierarchical Text Classification in Financial Domain
Qingcai Chen, Dongfang Li 0002
ICONIP (2)2
2021 NCBRPred: predicting nucleic acid binding residues in proteins based on multilabel learning
abstract
The interactions between proteins and nucleic acid sequences play many important roles in gene expression and some cellular activities. Accurate prediction of the nucleic acid binding residues in proteins will facilitate the research of the protein functions, gene expression, drug design, etc. In this regard, several computational methods have been proposed to predict the nucleic acid binding residues in proteins. However, these methods cannot satisfactorily measure the global interactions among the residues along protein. Furthermore, these methods are suffering cross-prediction problem, new strategies should be explored to solve this problem. In this study, a new computational method called NCBRPred was proposed to predict the nucleic acid binding residues based on the multilabel sequence labeling model. NCBRPred used the bidirectional Gated Recurrent Units (BiGRUs) to capture the global interactions among the residues, and treats this task as a multilabel learning task. Experimental results on three widely used benchmark datasets and an independent dataset showed that NCBRPred achieved higher predictive results with lower cross-prediction, outperforming 10 existing state-of-the-art predictors. The web-server and a stand-alone package of NCBRPred are freely available at http://bliulab.net/NCBRPred. It is anticipated that NCBRPred will become a very useful tool for identifying nucleic acid binding residues.
Jun Zhang 0078, Qingcai Chen, Bin Liu 0014
Briefings Bioinform.2
2021 Improving deep learning method for biomedical named entity recognition by using entity definition information
abstract
BACKGROUND: Biomedical named entity recognition (NER) is a fundamental task of biomedical text mining that finds the boundaries of entity mentions in biomedical text and determines their entity type. To accelerate the development of biomedical NER techniques in Spanish, the PharmaCoNER organizers launched a competition to recognize pharmacological substances, compounds, and proteins. Biomedical NER is usually recognized as a sequence labeling task, and almost all state-of-the-art sequence labeling methods ignore the meaning of different entity types. In this paper, we investigate some methods to introduce the meaning of entity types in deep learning methods for biomedical NER and apply them to the PharmaCoNER 2019 challenge. The meaning of each entity type is represented by its definition information. MATERIAL AND METHOD: We investigate how to use entity definition information in the following two methods: (1) SQuad-style machine reading comprehension (MRC) methods that treat entity definition information as query and biomedical text as context and predict answer spans as entities. (2) Span-level one-pass (SOne) methods that predict entity spans of one type by one type and introduce entity type meaning, which is represented by entity definition information. All models are trained and tested on the PharmaCoNER 2019 corpus, and their performance is evaluated by strict micro-average precision, recall, and F1-score. RESULTS: Entity definition information brings improvements to both SQuad-style MRC and SOne methods by about 0.003 in micro-averaged F1-score. The SQuad-style MRC model using entity definition information as query achieves the best performance with a micro-averaged precision of 0.9225, a recall of 0.9050, and an F1-score of 0.9137, respectively. It outperforms the best model of the PharmaCoNER 2019 challenge by 0.0032 in F1-score. Compared with the state-of-the-art model without using manually-crafted features, our model obtains a 1% improvement in F1-score, which is significant. These results indicate that entity definition information is useful for deep learning methods on biomedical NER. CONCLUSION: Our entity definition information enhanced models achieve the state-of-the-art micro-average F1 score of 0.9137, which implies that entity definition information has a positive impact on biomedical NER detection. In the future, we will explore more entity definition information from knowledge graph.
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010, Yi Zhou 0005
BMC Bioinform.4
2021 Distantly supervised biomedical relation extraction using piecewise attentive convolutional neural network and reinforcement learning
abstract
OBJECTIVE: There have been various methods to deal with the erroneous training data in distantly supervised relation extraction (RE), however, their performance is still far from satisfaction. We aimed to deal with the insufficient modeling problem on instance-label correlations for predicting biomedical relations using deep learning and reinforcement learning. MATERIALS AND METHODS: In this study, a new computational model called piecewise attentive convolutional neural network and reinforcement learning (PACNN+RL) was proposed to perform RE on distantly supervised data generated from Unified Medical Language System with MEDLINE abstracts and benchmark datasets. In PACNN+RL, PACNN was introduced to encode semantic information of biomedical text, and the RL method with memory backtracking mechanism was leveraged to alleviate the erroneous data issue. Extensive experiments were conducted on 4 biomedical RE tasks. RESULTS: The proposed PACNN+RL model achieved competitive performance on 8 biomedical corpora, outperforming most baseline systems. Specifically, PACNN+RL outperformed all baseline methods with the F1-score of 0.5592 on the may-prevent dataset, 0.6666 on the may-treat dataset, and 0.3838 on the DDI corpus, 2011. For the protein-protein interaction RE task, we obtained new state-of-the-art performance on 4 out of 5 benchmark datasets. CONCLUSIONS: The performance on many distantly supervised biomedical RE tasks was substantially improved, primarily owing to the denoising effect of the proposed model. It is anticipated that PACNN+RL will become a useful tool for large-scale RE and other downstream tasks to facilitate biomedical knowledge acquisition. We also made the demonstration program and source code publicly available at http://112.74.48.115:9000/.
Tiantian Zhu 0002, Yang Qin 0001, Yang Xiang 0003, Baotian Hu, Qingcai Chen, Weihua Peng
J. Am. Medical Informatics Assoc.5
2021 Neural data-to-text generation with dynamic content planning
Kai Chen 0020, Fayuan Li, Baotian Hu, Weihua Peng, Qingcai Chen, Hong Yu 0001, Yang Xiang 0003
Knowl. Based Syst.5
2021 Attentive capsule network for click-through rate and conversion rate prediction in online advertising
Dongfang Li 0002, Baotian Hu, Qingcai Chen, Quanchang Qi, Liubin Wang, Haishan Liu
Knowl. Based Syst.3
2021 Decomposing word embedding with the capsule network
Xin Liu 0054, Qingcai Chen, Yan Liu 0004, Joanna Siebert, Baotian Hu, Xiangping Wu 0001, Buzhou Tang
Knowl. Based Syst.2
2021 DeepDRBP-2L: A New Genome Annotation Predictor for Identifying DNA-Binding Proteins and RNA-Binding Proteins Using Convolutional Neural Network and Long Short-Term Memory
abstract
DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) are two kinds of crucial proteins, which are associated with various cellule activities and some important diseases. Accurate identification of DBPs and RBPs facilitate both theoretical research and real world application. Existing sequence-based DBP predictors can accurately identify DBPs but incorrectly predict many RBPs as DBPs, and vice versa, resulting in low prediction precision. Moreover, some proteins (DRBPs) interacting with both DNA and RNA play important roles in gene expression and cannot be identified by existing computational methods. In this study, a two-level predictor named DeepDRBP-2L was proposed by combining Convolutional Neural Network (CNN) and the Long Short-Term Memory (LSTM). It is the first computational method that is able to identify DBPs, RBPs and DRBPs. Rigorous cross-validations and independent tests showed that DeepDRBP-2L is able to overcome the shortcoming of the existing methods and can go one further step to identify DRBPs. Application of DeepDRBP-2L to tomato genome further demonstrated its performance. The webserver of DeepDRBP-2L is freely available at http://bliulab.net/DeepDRBP-2L.
Jun Zhang 0078, Qingcai Chen, Bin Liu 0014
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 LCSegNet: An Efficient Semantic Segmentation Network for Large-Scale Complex Chinese Character Recognition
abstract
Complex scene character recognition is a challenging yet important task in machine learning, especially for languages with large character sets, such as Chinese, which is composed of hieroglyphics with large-scale categories and similar glyphs. Recently, state-of-the-art methods based on semantic segmentation have achieved great success in scene parsing and have been applied in scene text recognition. However, because of limitations in terms of memory and computation, they are only applied in the small category recognition tasks, such as tasks involving English alphabets and digits. In this paper, we propose an efficient semantic segmentation model based on label coding (LC), called LCSegNet, to recognize large-scale Chinese characters. First, to reduce the number of labels, we design a new label coding method based on the Wubi Chinese characters code, called Wubi-CRF. In this method, glyphs and structure information of Chinese characters are encoded into 140-bit labels. Second, we employ an efficient semantic segmentation model for pixel-wise prediction and utilize a conditional random field (CRF) module to learn the constraint rules of Wubi-like coding. Finally, experiments are conducted on three benchmarks: a large Chinese text dataset in the wild (CTW), ICDAR2019-ReCTS, and HIT-OR3C dataset. Results show that the proposed method achieves state-of-the-art performances in both complex scene and handwritten character recognition tasks.
Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Xin Liu 0054, Baotian Hu
IEEE Trans. Multim.2
2020 KEoG: A knowledge-aware edge-oriented graph neural network for document-level relation extraction
abstract
Document-level relation extraction (RE) has attracted more and more attentions recently. Edge-oriented graph neural network (EoG) is a new neural network exhibiting greater potential than previous node-oriented graph neural networks for document-level RE. In this paper, we propose a novel EoG, called knowledge-aware edge-oriented GNN (KEoG) for document-level RE. In KEoG, we further introduce not only two types of nodes to represent documents and external knowledge respectively, but also soft F-Measure loss function to solve the inherent class imbalance problem in document-level RE. Experiments conducted on two document-level datasets show that KEoG outperforms other state-of-the-art methods for comparison on both intra-sentence and inter-sentence relation extractions, indicating that KEoG is an effective extension of EoG.
Weihua Peng, Qingcai Chen, Xiaolong Wang 0001, Buzhou Tang
BIBM3
2020 MedWriter: Knowledge-Aware Medical Text Generation
abstract
To exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical.However, most of recent works only consider the information of unstructured text rather than structured information of the knowledge graph.In this paper, we focus on the medical topic-to-text generation task and adapt a knowledge-aware text generation model to the medical domain, named MedWriter, which not only introduces the specific knowledge from the external MKG but also is capable of learning graph-level representation.We conduct experiments on a medical literature dataset collected from medical journals, each of which has a set of topic words, an abstract of medical literature and a corresponding knowledge graph from CMeKG.Experimental results demonstrate incorporating knowledge graph into generation model can improve the quality of the generated text and has robust superiority over the competitor methods.
Youcheng Pan, Qingcai Chen, Weihua Peng, Xiaolong Wang 0001, Baotian Hu, Xin Liu 0054, Wenxiu Zhou
COLING2
2020 Towards Medical Machine Reading Comprehension with Structural Knowledge and Plain Text
abstract
Machine reading comprehension (MRC) has achieved significant progress on the open domain in recent years, mainly due to large-scale pre-trained language models.However, it performs much worse in specific domains such as the medical field due to the lack of extensive training data and professional structural knowledge neglect.As an effort, we first collect a large scale medical multi-choice question dataset (more than 21k instances) for the National Licensed Pharmacist Examination in China.It is a challenging medical examination with a passing rate of less than 14.2% in 2018.Then we propose a novel reading comprehension model KMQA, which can fully exploit the structural medical knowledge (i.e., medical knowledge graph) and the reference medical plain text (i.e., text snippets retrieved from reference books).The experimental results indicate that the KMQA outperforms existing competitive models with a large margin and passes the exam with 61.8% accuracy rate on the test set.
Dongfang Li 0002, Baotian Hu, Qingcai Chen, Weihua Peng
EMNLP (1)3
2020 SED-MDD: Towards Sentence Dependent End-To-End Mispronunciation Detection and Diagnosis
abstract
A mispronunciation detection and diagnosis (MD&D) system typically consists of multiple stages, such as an acoustic model, a language model and a Viterbi decoder. In order to integrate these stages, we propose SED-MDD, an end-to-end model for sentence dependent mispronunciation detection and diagnosis (MD&D) . Our proposed model takes mel-spectrogram and characters as inputs and outputs the corresponding phone sequence. Our experiments prove that SED-MDD can implicitly learn the phonological rules in both acoustic and linguistic features directly from the phonological annotation and transcription in the training data. To the best of our knowledge, SED-MDD is the first model of its kind and it achieves an accuracy of 86.35% and a correctness of 88.61% on L2-ARCTIC which significantly outperforms the existing end-to-end mispronunciation detection and diagnosis (MD&D) model CNN-RNN-CTC.
Yiqing Feng, Guanyu Fu, Qingcai Chen, Kai Chen 0020
ICASSP3
2020 Learning to Generate Diverse Questions from Keywords
abstract
Diverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., generating a series of questions for each set of keywords (one-to-many). As an effort towards this, we propose a novel neural generative model, which incorporates context information and control signal to produce multiple diverse questions from a given fixed set of keywords. The control signal is designed to increase the diversity of questions by capturing the diverse patterns from the entire dataset. The context information is used to guarantee the generated questions are highly related to the given keywords. To evaluate the effectiveness of the proposed model, we collect a dataset which contains 62835 questions with respect to 12567 sets of keywords.1To the best of our knowledge, it's the first Chinese financial dataset for diverse question generation. The experimental results show that our model outperforms the competitor methods in terms of BLEU and Distinct. The qualitative evaluation indicates that our model is able to generate diverse and meaningful questions.
Youcheng Pan, Baotian Hu, Qingcai Chen, Yang Xiang 0003, Xiaolong Wang 0001
ICASSP3
2020 AdaHGNN: Adaptive Hypergraph Neural Networks for Multi-Label Image Classification
abstract
Multi-label image classification is an important and challenging task in computer vision and multimedia fields. Most of the recent works only capture the pair-wise dependencies among multiple labels through statistical co-occurrence information, which cannot model the high-order semantic relations automatically. In this paper, we propose a high-order semantic learning model based on adaptive hypergraph neural networks (AdaHGNN) to boost multi-label classification performance. Firstly, an adaptive hypergraph is constructed by using label embeddings automatically. Secondly, image features are decoupled into feature vectors corresponding to each label, and hypergraph neural networks (HGNN) are employed to correlate these vectors and explore the high-order semantic interactions. In addition, multi-scale learning is used to reduce sensitivity to object size inconsistencies. Experiments are conducted on four benchmarks: MS-COCO, NUS-WIDE, Visual Genome, and Pascal VOC 2007, which cover large, medium, and small-scale categories. State-of-the-art performances are achieved on three of them. Results and analysis demonstrate that the proposed method has the ability to capture high-order semantic dependencies.
Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Baotian Hu
ACM Multimedia2
2020 Text-Guided Neural Image Inpainting
abstract
Image inpainting task requires filling the corrupted image with contents coherent with the context. This research field has achieved promising progress by using neural image inpainting methods. Nevertheless, there is still a critical challenge in guessing the missed content with only the context pixels. The goal of this paper is to fill the semantic information in corrupted images according to the provided descriptive text. Unique from existing text-guided image generation works, the inpainting models are required to compare the semantic content of the given text and the remaining part of the image, then find out the semantic content that should be filled for missing part. To fulfill such a task, we propose a novel inpainting model named Text-Guided Dual Attention Inpainting Network (TDANet). Firstly, a dual multimodal attention mechanism is designed to extract the explicit semantic information about the corrupted regions, which is done by comparing the descriptive text and complementary image areas through reciprocal attention. Secondly, an image-text matching loss is applied to maximize the semantic similarity of the generated image and the text. Experiments are conducted on two open datasets. Results show that the proposed TDANet model reaches new state-of-the-art on both quantitative and qualitative measures. Result analysis suggests that the generated images are consistent with the guidance text, enabling the generation of various results by providing different descriptions. Codes are available at https://github.com/idealwhite/TDANet
Lisai Zhang, Qingcai Chen, Baotian Hu, Shuoran Jiang
ACM Multimedia2
2020 Gated Semantic Difference Based Sentence Semantic Equivalence Identification
abstract
This article proposes a novel sentence semantic equivalence identification (SSEI) method by using the semantic difference features between sentences. The lexical differences of a sentence pair are first extracted, and the bidirectional long short term memory (BiLSTM) network is then applied on them to generate the semantic difference representations. Finally, an efficient gate mechanism is proposed to integrate the semantic differences with existing models (called base model) to enhance their encoding capability in the SSEI task. Exhaustive experiments conducted on the standard Quora corpus, and the Large-scale Chinese Question Matching Corpus (LCQMC) show that the proposed gated semantic difference (GSD) method brings significant improvement for different existing state-of-the-art models. When the bidirectional encoder representations from transformers model (BERT) is used as the base model, the accuracy for SSEI on Quora is improved from 90.63% to 91.98%, and the F1 score on the LCQMC is improved from 87.0% to 87.7%, which outperforms the best-published results.
Xin Liu 0054, Qingcai Chen, Xiangping Wu 0001, Yang Hua 0004, Dongfang Li 0002, Buzhou Tang, Xiaolong Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Stroke Sequence-Dependent Deep Convolutional Neural Network for Online Handwritten Chinese Character Recognition
abstract
We propose a novel model, called stroke sequence-dependent deep convolutional neural network (SSDCNN), which uses the stroke sequence information and eight-directional features of Chinese characters for online handwritten Chinese character recognition (OLHCCR). SSDCNN learns the representation of OLHCCs by incorporating the natural sequence information of the strokes. Furthermore, it naturally incorporates the eight-directional features. First, SSDCNN inputs the stroke sequence and transforms it into stacks of feature maps following the writing order of the strokes. Second, the fixed-length, stroke sequence-dependent representations of OLHCC are derived through convolutional, residual, and max-pooling operations. Third, the stroke sequence-dependent representation is combined with the eight-directional features via a number of fully connected neural network layers. Finally, the Chinese characters are recognized using a softmax classifier. The SSDCNN is trained in two stages: 1) the whole architecture is pretrained using the training data until the performance converges to an acceptable degree. 2) The stroke sequence-dependent representation is combined with the eight-directional features by a fully connected neural network and a softmax layer for further training. The model was experimentally evaluated on the OLHCCR competition tasks of International Conference on Document Analysis and Recognition (ICDAR) 2013. The recognition error was a maximum 58.28% lower in SSDCNN than in a model using the eight-directional features alone (5.13% versus 2.14%). Owing to its high accuracy (97.86%), the proposed SSDCNN reduced the recognition error by approximately 18.0% as compared with that of the winning system in the ICDAR 2013 competition. SSDCNN integrated with an adaptation mechanism, called the SSDCNN+Adapt model, and reached a new state-of-the-art (SOTA) standard with an accuracy of 97.94%. The SSDCNN exploits the stroke sequence information to learn high-quality OLHCC representations. Moreover, the learned representation and the classical eight-directional features complement each other within the SSDCNN architecture.
Xin Liu 0054, Baotian Hu, Qingcai Chen, Xiangping Wu 0001, Jinghan You
IEEE Trans. Neural Networks Learn. Syst.3
2019 De-identification of Clinical Text via Bi-LSTM-CRF with Neural Language Models
Buzhou Tang, Dehuan Jiang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010, Ying Shen 0001
AMIA3
2019 A Study on Automatic Generation of Chinese Discharge Summary
abstract
Discharge summary, which summarizes a patient's health information during hospitalization, is very important for transferring information between the hospitalist and primary care physician. Discharge summary writing is a necessary but time-consuming job for physicians. Automatically generating discharge summaries using information technology is helpful to physicians, but challenging as discharge summaries are typically long and contain amounts of information to outline patient's reason for admission, labtests, examinations, diagnostic findings, treatments, and medication care plan. In this study, we propose a framework based on deep learning methods for automatic discharge summary generation. In the framework, we split information of discharge summary into two parts: 1) history information such as reason for admission, admission diagnosis; 2) outcome information such as discharge diagnosis and medication care plan, and deploy different neural networks to generate them separately. Hierarchical sequence labeling methods are proposed to select key sentences from existing documents, e.g., admission notes, progress notes, examination reports as the history information, and multi-task learning methods to predict the outcomes, e.g., diagnosis and medication care plan. Experiments on a Chinese corpus show that our approach has the ability to generate effective discharge summaries.
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010
BIBM3
2019 Multi-strategies Method for Cold-Start Stage Question Matching of rQA Task
Dongfang Li 0002, Qingcai Chen, Songjian Chen, Xin Liu 0054, Buzhou Tang, Ben Tan
NLPCC (1)2
2019 Extracting entities with attributes in clinical text via joint deep learning
abstract
OBJECTIVE: Extracting clinical entities and their attributes is a fundamental task of natural language processing (NLP) in the medical domain. This task is typically recognized as 2 sequential subtasks in a pipeline, clinical entity or attribute recognition followed by entity-attribute relation extraction. One problem of pipeline methods is that errors from entity recognition are unavoidably passed to relation extraction. We propose a novel joint deep learning method to recognize clinical entities or attributes and extract entity-attribute relations simultaneously. MATERIALS AND METHODS: The proposed method integrates 2 state-of-the-art methods for named entity recognition and relation extraction, namely bidirectional long short-term memory with conditional random field and bidirectional long short-term memory, into a unified framework. In this method, relation constraints between clinical entities and attributes and weights of the 2 subtasks are also considered simultaneously. We compare the method with other related methods (ie, pipeline methods and other joint deep learning methods) on an existing English corpus from SemEval-2015 and a newly developed Chinese corpus. RESULTS: Our proposed method achieves the best F1 of 74.46% on entity recognition and the best F1 of 50.21% on relation extraction on the English corpus, and 89.32% and 88.13% on the Chinese corpora, respectively, which outperform the other methods on both tasks. CONCLUSIONS: The joint deep learning-based method could improve both entity recognition and relation extraction from clinical text in both English and Chinese, indicating that the approach is promising.
Xue Shi, Yingping Yi, Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Zongcheng Ji, Yaoyun Zhang, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2019 Cohort selection for clinical trials using hierarchical neural network
abstract
OBJECTIVE: Cohort selection for clinical trials is a key step for clinical research. We proposed a hierarchical neural network to determine whether a patient satisfied selection criteria or not. MATERIALS AND METHODS: We designed a hierarchical neural network (denoted as CNN-Highway-LSTM or LSTM-Highway-LSTM) for the track 1 of the national natural language processing (NLP) clinical challenge (n2c2) on cohort selection for clinical trials in 2018. The neural network is composed of 5 components: (1) sentence representation using convolutional neural network (CNN) or long short-term memory (LSTM) network; (2) a highway network to adjust information flow; (3) a self-attention neural network to reweight sentences; (4) document representation using LSTM, which takes sentence representations in chronological order as input; (5) a fully connected neural network to determine whether each criterion is met or not. We compared the proposed method with its variants, including the methods only using the first component to represent documents directly and the fully connected neural network for classification (denoted as CNN-only or LSTM-only) and the methods without using the highway network (denoted as CNN-LSTM or LSTM-LSTM). The performance of all methods was measured by micro-averaged precision, recall, and F1 score. RESULTS: The micro-averaged F1 scores of CNN-only, LSTM-only, CNN-LSTM, LSTM-LSTM, CNN-Highway-LSTM, and LSTM-Highway-LSTM were 85.24%, 84.25%, 87.27%, 88.68%, 88.48%, and 90.21%, respectively. The highest micro-averaged F1 score is higher than our submitted 1 of 88.55%, which is 1 of the top-ranked results in the challenge. The results indicate that the proposed method is effective for cohort selection for clinical trials. DISCUSSION: Although the proposed method achieved promising results, some mistakes were caused by word ambiguity, negation, number analysis and incomplete dictionary. Moreover, imbalanced data was another challenge that needs to be tackled in the future. CONCLUSION: In this article, we proposed a hierarchical neural network for cohort selection. Experimental results show that this method is good at selecting cohort.
Xue Shi, Dehuan Jiang, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Jun Yan 0010
J. Am. Medical Informatics Assoc.7
2019 Unconstrained Offline Handwritten Word Recognition by Position Embedding Integrated ResNets Model
abstract
The state-of-the-art methods usually integrate with linguistic knowledge in the recognizer, which makes models more complicated and hard for resource-lacking languages. This letter proposes a new method for unconstrained offline handwritten word recognition by combining position embeddings with residual networks (ResNets) and bidirectional long short-term memory (BiLSTM) networks. At first, ResNets are used to extract abundant features from the input image. Then, position embeddings are used as indices of the character sequence corresponding to a word. By combining the ResNets features with each position embedding, the model generates different inputs for the BiLSTM networks. Finally, the state sequence of the BiLSTM is used to recognize corresponding characters. Without additional language resource, the proposed model achieved the best result on two public corpora, i.e., the 2017 ICDAR word-level information extraction in historical handwritten records competition and the RIMES public dataset on character error rate.
Xiangping Wu 0001, Qingcai Chen, Jinghan You, Yulun Xiao
IEEE Signal Process. Lett.2
2018 LCQMC: A Large-scale Chinese Question Matching Corpus
abstract
The lack of large-scale question matching corpora greatly limits the development of matching methods in question answering (QA) system, especially for non-English languages. To ameliorate this situation, in this paper, we introduce a large-scale Chinese question matching corpus (named LCQMC), which is released to the public1. LCQMC is more general than paraphrase corpus as it focuses on intent matching rather than paraphrase. How to collect a large number of question pairs in variant linguistic forms, which may present the same intent, is the key point for such corpus construction. In this paper, we first use a search engine to collect large-scale question pairs related to high-frequency words from various domains, then filter irrelevant pairs by the Wasserstein distance, and finally recruit three annotators to manually check the left pairs. After this process, a question matching corpus that contains 260,068 question pairs is constructed. In order to verify the LCQMC corpus, we split it into three parts, i.e., a training set containing 238,766 question pairs, a development set with 8,802 question pairs, and a test set with 12,500 question pairs, and test several well-known sentence matching methods on it. The experimental results not only demonstrate the good quality of LCQMC but also provide solid baseline performance for further researches on this corpus.
Xin Liu 0054, Qingcai Chen, Chong Deng, Hua-Jun Zeng, Dongfang Li 0002, Buzhou Tang
COLING2
2018 The BQ Corpus: A Large-scale Domain-specific Chinese Corpus For Sentence Semantic Equivalence Identification
abstract
This paper introduces the Bank Question (BQ) corpus, a Chinese corpus for sentence semantic equivalence identification (SSEI). The BQ corpus contains 120,000 question pairs from 1-year online bank custom service logs. To efficiently process and annotate questions from such a large scale of logs, this paper proposes a clustering based annotation method to achieve questions with the same intent. First, the deduplicated questions with the same answer are clustered into stacks by the Word Mover’s Distance (WMD) based Affinity Propagation (AP) algorithm. Then, the annotators are asked to assign the clustered questions into different intent categories. Finally, the positive and negative question pairs for SSEI are selected in the same intent category and between different intent categories respectively. We also present six SSEI benchmark performance on our corpus, including state-of-the-art algorithms. As the largest manually annotated public Chinese SSEI corpus in the bank domain, the BQ corpus is not only useful for Chinese question semantic matching research, but also a significant resource for cross-lingual and cross-domain SSEI research. The corpus is available in public.
Qingcai Chen, Xin Liu 0054, Daohe Lu, Buzhou Tang
EMNLP2
2018 Recurrent convolutional neural network for answer selection in community question answering
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001
Neurocomputing3
2018 Recognizing Continuous and Discontinuous Adverse Drug Reaction Mentions from Social Media Using LSTM-CRF
abstract
Social media in medicine, where patients can express their personal treatment experiences by personal computers and mobile devices, usually contains plenty of useful medical information, such as adverse drug reactions (ADRs); mining this useful medical information from social media has attracted more and more attention from researchers. In this study, we propose a deep neural network (called LSTM‐CRF) combining long short‐term memory (LSTM) neural networks (a type of recurrent neural networks) and conditional random fields (CRFs) to recognize ADR mentions from social media in medicine and investigate the effects of three factors on ADR mention recognition. The three factors are as follows: (1) representation for continuous and discontinuous ADR mentions: two novel representations, that is, “BIOHD” and “Multilabel,” are compared; (2) subject of posts: each post has a subject (i.e., drug here); and (3) external knowledge bases. Experiments conducted on a benchmark corpus, that is, CADEC, show that LSTM‐CRF achieves better F ‐score than CRF; “Multilabel” is better in representing continuous and discontinuous ADR mentions than “BIOHD”; both subjects of comments and external knowledge bases are individually beneficial to ADR mention recognition. To the best of our knowledge, this is the first time to investigate deep neural networks to mine continuous and discontinuous ADRs from social media.
Buzhou Tang, Jianglu Hu, Xiaolong Wang 0001, Qingcai Chen
Wirel. Commun. Mob. Comput.4
2017 Chemical-induced disease extraction via convolutional neural networks with attention
abstract
Extracting relationships between chemicals and diseases from unstructured literature is very important for many biomedical applications such as pharmacovigilance and drug repositioning. Automatic chemical-induced disease extraction is usually recognized as a classification task, and several systems have been proposed for this task recently due to some annotated corpora publicly available. Most of the systems are based on machine learning methods with many manually-crafted features. In recent years, deep learning that does not only can avoid verbose feature engineering but also shows competitive performance has been widely used in various types of tasks, including classification task. Therefore, deep learning has great potential on chemical-induced disease extraction. In this paper, we proposed an architecture of convolutional neural networks (CNN) with attention mechanism for chemical-induced disease extraction, which does not only avoid verbose feature engineering but also integrates domain knowledge in a simple way. Experiments on a benchmark dataset demonstrate that the proposed CNN-based chemical-induced disease extraction system is competitive with other state-of-the-art systems.
Haodi Li, Qingcai Chen, Buzhou Tang, Xiaolong Wang 0001
BIBM2
2017 Answer Selection in Community Question Answering by Normalizing Support Answers
Zhihui Zheng, Daohe Lu, Qingcai Chen, Yang Xiang 0003, Youcheng Pan
NLPCC3
2017 CNN-based ranking for biomedical entity normalization
abstract
BACKGROUND: Most state-of-the-art biomedical entity normalization systems, such as rule-based systems, merely rely on morphological information of entity mentions, but rarely consider their semantic information. In this paper, we introduce a novel convolutional neural network (CNN) architecture that regards biomedical entity normalization as a ranking problem and benefits from semantic information of biomedical entities. RESULTS: The CNN-based ranking method first generates candidates using handcrafted rules, and then ranks the candidates according to their semantic information modeled by CNN as well as their morphological information. Experiments on two benchmark datasets for biomedical entity normalization show that our proposed CNN-based ranking method outperforms traditional rule-based method with state-of-the-art performance. CONCLUSIONS: We propose a CNN architecture that regards biomedical entity normalization as a ranking problem. Comparison results show that semantic information is beneficial to biomedical entity normalization and can be well combined with morphological information in our CNN architecture for further improvement.
Haodi Li, Qingcai Chen, Buzhou Tang, Xiaolong Wang 0001, Hua Xu 0001
BMC Bioinform.2
2017 Answer Selection in Community Question Answering via Attentive Neural Networks
abstract
Answer selection in community question answering (cQA) is a challenging task in natural language processing. The difficulty lies in that it not only needs the consideration of semantic matching between question answer pairs but also requires a serious modeling of contextual factors. In this letter, we propose an attentive deep neural network architecture so as to learn the deterministic information for answer selection. The architecture can support various input formats through the organization of convolutional neural networks, attention-based long short-term memory, and conditional random fields. Experiments are carried out on the SemEval-2015 cQA dataset. We attain 58.35% on macroaveraged F1, which outperforms the Top-1 system in the shared task by 1.16% and improves the state-of-the-art deep-neural-network-based method by 2.21%.
Yang Xiang 0003, Qingcai Chen, Xiaolong Wang 0001, Yang Qin 0001
IEEE Signal Process. Lett.2
2016 CMedTEX: A Rule-based Temporal Expression Extraction and Normalization System for Chinese Clinical Notes
Zengjian Liu, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Haodi Li, Junzhao Bu, Jingzhi Jiang, Qiwen Deng, Suisong Zhu
AMIA4
2016 Dependency-based convolutional neural network for drug-drug interaction extraction
abstract
Drug-drug interactions (DDIs) are crucial for healthcare. Besides DDIs reported in medical knowledge bases such as DrugBank, a large number of latest DDI findings are also reported in unstructured biomedical literature. Extracting DDIs from unstructured biomedical literature is a worthy addition to the existing knowledge bases. Currently, convolutional neural network (CNN) is a state-of-the-art method for DDI extraction. One limitation of CNN is that it neglects long distance dependencies between words in candidate DDI instances, which may be helpful for DDI extraction. In order to incorporate the long distance dependencies between words in candidate DDI instances, in this work, we propose a dependency-based convolutional neural network (DCNN) for DDI extraction. Experiments conducted on the DDIExtraction 2013 corpus show that DCNN using a public state-of-the-art dependency parser achieves an F-score of 70.19%, outperforming CNN by 0.44%. By analyzing errors of DCNN, we find that errors from dependency parsers are propagated into DCNN and affect the performance of DCNN. To reduce error propagation, we design a simple rule to combine CNN with DCNN, that is, using DCNN to extract DDIs in short sentences and CNN to extract DDIs in long distances as most dependency parsers work well for short sentences but bad for long sentences. Finally, our system that combines CNN and DCNN achieves an F-score of 70.81%, outperforming CNN by 1.06% and DNN by 0.62% on the DDIExtraction 2013 corpus.
Kai Chen 0020, Qingcai Chen, Buzhou Tang
BIBM3
2016 Incorporating Label Dependency for Answer Quality Tagging in Community Question Answering via CNN-LSTM-CRF
abstract
In community question answering (cQA), the quality of answers are determined by the matching degree between question-answer pairs and the correlation among the answers. In this paper, we show that the dependency between the answer quality labels also plays a pivotal role. To validate the effectiveness of label dependency, we propose two neural network-based models, with different combination modes of Convolutional Neural Net-works, Long Short Term Memory and Conditional Random Fields. Extensive experi-ments are taken on the dataset released by the SemEval-2015 cQA shared task. The first model is a stacked ensemble of the networks. It achieves 58.96% on macro averaged F1, which improves the state-of-the-art neural network-based method by 2.82% and outper-forms the Top-1 system in the shared task by 1.77%. The second is a simple attention-based model whose input is the connection of the question and its corresponding answers. It produces promising results with 58.29% on overall F1 and gains the best performance on the Good and Bad categories.
Yang Xiang 0003, Xiaoqiang Zhou, Qingcai Chen, Zhihui Zheng, Buzhou Tang, Xiaolong Wang 0001, Yang Qin 0001
COLING3
2016 A novel word embedding learning model using the dissociation between nouns and verbs
Baotian Hu, Buzhou Tang, Qingcai Chen, Longbiao Kang
Neurocomputing3
2015 Recognizing Disjoint Clinical Concepts in Clinical Text Using Machine Learning-based Methods
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001
AMIA2
2015 LCSTS: A Large Scale Chinese Short Text Summarization Dataset
abstract
Automatic text summarization is widely regarded as the highly difficult problem, partially because of the lack of large text summarization data set.Due to the great challenge of constructing the large scale summaries for full text, in this paper, we introduce a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public 1 .This corpus consists of over 2 million real Chinese short texts with short summaries given by the author of each text.We also manually tagged the relevance of 10,666 short summaries with their corresponding short texts.Based on the corpus, we introduce recurrent neural network for the summary generation and achieve promising results, which not only shows the usefulness of the proposed corpus for short text summarization research, but also provides a baseline for further research on this topic.
Baotian Hu, Qingcai Chen, Fangze Zhu
EMNLP2
2015 Structural Regularity Exploration in Multidimensional Networks
Yi Chen 0019, Xiaolong Wang 0001, Buzhou Tang, Junzhao Bu, Qingcai Chen
ICONIP (3)5
2015 An Auto-Encoder for Learning Conversation Representation Using LSTM
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001
ICONIP (1)3
2015 Visual orientation inhomogeneity based scale-invariant feature transform
Shenghua Zhong, Yan Liu 0004, Qingcai Chen
Expert Syst. Appl.3
2015 An automatic system to identify heart disease risk factors in clinical texts over time
abstract
Despite recent progress in prediction and prevention, heart disease remains a leading cause of death. One preliminary step in heart disease prediction and prevention is risk factor identification. Many studies have been proposed to identify risk factors associated with heart disease; however, none have attempted to identify all risk factors. In 2014, the National Center of Informatics for Integrating Biology and Beside (i2b2) issued a clinical natural language processing (NLP) challenge that involved a track (track 2) for identifying heart disease risk factors in clinical texts over time. This track aimed to identify medically relevant information related to heart disease risk and track the progression over sets of longitudinal patient medical records. Identification of tags and attributes associated with disease presence and progression, risk factors, and medications in patient medical history were required. Our participation led to development of a hybrid pipeline system based on both machine learning-based and rule-based approaches. Evaluation using the challenge corpus revealed that our system achieved an F1-score of 92.68%, making it the top-ranked system (without additional annotations) of the 2014 i2b2 clinical NLP challenge.
Qingcai Chen, Haodi Li, Buzhou Tang, Xiaolong Wang 0001, Xin Liu 0054, Zengjian Liu, Weida Wang, Qiwen Deng, Suisong Zhu, Yangxin Chen
J. Biomed. Informatics1
2015 Automatic de-identification of electronic medical records using token-level and character-level conditional random fields
abstract
De-identification, identifying and removing all protected health information (PHI) present in clinical data including electronic medical records (EMRs), is a critical step in making clinical data publicly available. The 2014 i2b2 (Center of Informatics for Integrating Biology and Bedside) clinical natural language processing (NLP) challenge sets up a track for de-identification (track 1). In this study, we propose a hybrid system based on both machine learning and rule approaches for the de-identification track. In our system, PHI instances are first identified by two (token-level and character-level) conditional random fields (CRFs) and a rule-based classifier, and then are merged by some rules. Experiments conducted on the i2b2 corpus show that our system submitted for the challenge achieves the highest micro F-scores of 94.64%, 91.24% and 91.63% under the "token", "strict" and "relaxed" criteria respectively, which is among top-ranked systems of the 2014 i2b2 challenge. After integrating some refined localization dictionaries, our system is further improved with F-scores of 94.83%, 91.57% and 91.95% under the "token", "strict" and "relaxed" criteria respectively.
Zengjian Liu, Yangxin Chen, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Haodi Li, Qiwen Deng, Suisong Zhu
J. Biomed. Informatics5
2014 Hybrid Deep Belief Networks for Semi-supervised Sentiment Classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
COLING2
2014 Convolutional Neural Network Architectures for Matching Natural Language Sentences
Baotian Hu, Zhengdong Lu, Hang Li 0001, Qingcai Chen
NIPS4
2014 A Short Texts Matching Method Using Shallow Features and Deep Features
Longbiao Kang, Baotian Hu, Xiangping Wu 0001, Qingcai Chen
NLPCC4
2014 Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection
abstract
MOTIVATION: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. RESULTS: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/~binliu/remote/
Bin Liu 0014, Deyuan Zhang, Ruifeng Xu 0001, Jinghao Xu, Xiaolong Wang 0001, Qingcai Chen, Qiwen Dong, Kuo-Chen Chou
Bioinform.6
2014 Using distances between Top-n-gram and residue pairs for protein remote homology detection
abstract
BACKGROUND: Protein remote homology detection is one of the central problems in bioinformatics, which is important for both basic research and practical application. Currently, discriminative methods based on Support Vector Machines (SVMs) achieve the state-of-the-art performance. Exploring feature vectors incorporating the position information of amino acids or other protein building blocks is a key step to improve the performance of the SVM-based methods. RESULTS: Two new methods for protein remote homology detection were proposed, called SVM-DR and SVM-DT. SVM-DR is a sequence-based method, in which the feature vector representation for protein is based on the distances between residue pairs. SVM-DT is a profile-based method, which considers the distances between Top-n-gram pairs. Top-n-gram can be viewed as a profile-based building block of proteins, which is calculated from the frequency profiles. These two methods are position dependent approaches incorporating the sequence-order information of protein sequences. Various experiments were conducted on a benchmark dataset containing 54 families and 23 superfamilies. Experimental results showed that these two new methods are very promising. Compared with the position independent methods, the performance improvement is obvious. Furthermore, the proposed methods can also provide useful insights for studying the features of protein families. CONCLUSION: The better performance of the proposed methods demonstrates that the position dependant approaches are efficient for protein remote homology detection. Another advantage of our methods arises from the explicit feature space representation, which can be used to analyze the characteristic features of protein families. The source code of SVM-DT and SVM-DR is available at http://bioinformatics.hitsz.edu.cn/DistanceSVM/index.jsp.
Bin Liu 0014, Jinghao Xu, Quan Zou 0001, Ruifeng Xu 0001, Xiaolong Wang 0001, Qingcai Chen
BMC Bioinform.6
2014 Fuzzy deep belief networks for semi-supervised sentiment classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
Neurocomputing2
2014 Handwritten Chinese text editing and recognition system
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
Multim. Tools Appl.2
2014 Multilinear Sparse Principal Component Analysis
abstract
In this brief, multilinear sparse principal component analysis (MSPCA) is proposed for feature extraction from the tensor data. MSPCA can be viewed as a further extension of the classical principal component analysis (PCA), sparse PCA (SPCA) and the recently proposed multilinear PCA (MPCA). The key operation of MSPCA is to rewrite the MPCA into multilinear regression forms and relax it for sparse regression. Differing from the recently proposed MPCA, MSPCA inherits the sparsity from the SPCA and iteratively learns a series of sparse projections that capture most of the variation of the tensor data. Each nonzero element in the sparse projections is selected from the most important variables/factors using the elastic net. Extensive experiments on Yale, Face Recognition Technology face databases, and COIL-20 object database encoded the object images as second-order tensors, and Weizmann action database as third-order tensors demonstrate that the proposed MSPCA algorithm has the potential to outperform the existing PCA-based subspace learning algorithms.
Zhihui Lai 0001, Yong Xu 0001, Qingcai Chen, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2013 Automatic Corpora Construction for Text Classification
Qingcai Chen, Xiaolong Wang 0001, Bingyang Yu
IJCNLP2
2013 Active deep learning method for semi-supervised sentiment classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
Neurocomputing2
2013 Convolutional Deep Networks for Visual Data Classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
Neural Process. Lett.2
2012 An Empirical Evaluation on Online Chinese Handwriting Databases
abstract
Several online Chinese handwriting databases have been proposed recently. Though they have been introduced in detail, to date, no one has ever evaluated these databases with experimental comparison. To help the researchers use the corresponding database properly for algorithm evaluation and real application, we compare the property of the handwriting characters in these databases, and evaluate them with the same experimental setup and handwriting recognizer. Moreover, we analyze the connection between the property and the corresponding recognition accuracy for the handwriting characters in different databases. These empirical evaluation results can help the researchers choose the right database for different algorithms and applications.
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001, Zou Chen, Suqin Ao
Document Analysis Systems2
2011 An Empirical Evaluation on HIT-OR3C Database
abstract
Recently, we have proposed a handwriting Chinese character database HIT-OR3C. Though it has been introduced in detail, to date, it has not been evaluated by any handwriting recognition method. To help the researchers use this database for algorithm evaluation, we propose the structure of HIT-OR3C database. Moreover, we evaluate the OR3C database with a series of experiments using state-of-the-art handwriting recognizer. These experiment results on the different subsets can be a benchmark for the researchers who will use the database. The low average recognition rate confirms that the HIT-OR3C database is challenging.
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
ICDAR2
2011 Macro Features Based Text Categorization
Qingcai Chen, Xiaolong Wang 0001, Buzhou Tang
ICONIP (2)2
2011 Dynamic Template Based Online Event Detection
Qingcai Chen, Xiaolong Wang 0001, Jiacai Weng
ICONIP (3)2
2011 Deep Belief Networks for Automatic Music Genre Classification
Xiaohong Yang, Qingcai Chen, Shusen Zhou, Xiaolong Wang 0001
INTERSPEECH2
2011 Tolerance Rough Set Based Attribute Extraction Approach for Multiple Semantic Knowledge Base Integration
abstract
In the integration of multiple semantic knowledge bases (SKBs), the inconsistence of the items or their attributes appeared in different SKBs is still an opening challenge for researchers. To address this issue, this paper presents an innovative approach which bases on extracting common class attributes and establishing unified category-attribute templates. Since the natural properties of uncertainty and vagueness of semantic analysis involved in selecting a specific attribute from numerous candidates, the tolerance rough set (TRS) techniques are applied in constructing class-attribute templates from online SKBs. The extraction of attribute is fulfilled by statistical techniques and is integrated into the TRS framework. Finally, experiments are conducted on random selected categories. Experimental results show the effectiveness of the proposed approach.
Hongzhi Guo 0007, Qingcai Chen, Xiaolong Wang 0001
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2011 Discriminative deep belief networks for visual data classification
Yan Liu 0004, Shusen Zhou, Qingcai Chen
Pattern Recognit.3
2010 HIT-OR3C: an opening recognition corpus for Chinese characters
abstract
This paper proposes an opening recognition corpus, HIT-OR3C, and its construction toolkit to facilitate the unconstrained online Chinese handwriting text recognition. The characters of HIT-OR3C are collected through handwriting pad and are recorded and labeled automatically via the proposed handwriting document collection software OR3C Toolkit. HIT-OR3C consists of 5 subsets, namely GB1, GB2, Letter, Digit and Document. The first 4 corpora contain 6,825 categories produced by 122 persons and 832,650 samples in total. The document corpus is corresponding to 10 news articles that contain 2,442 categories produced by 20 persons and 77,168 samples in total. HIT-OR3C can be used for training and evaluation of character recognition algorithms. The OR3C Toolkit provides an efficient, device-independent, and unconstrained platform for the building of large scale handwriting corpus.
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
Document Analysis Systems2
2010 Discriminative Deep Belief Networks for image classification
abstract
This paper presents a novel semi-supervised learning algorithm called Discriminative Deep Belief Networks (DDBN), to address the image classification problem with limited labeled data. We first construct a new deep architecture for classification using a set of Restricted Boltzmann Machines (RBM). The parameter space of the deep architecture is initially determined using labeled data together with abundant of unlabeled data, by greedy layer-wise unsupervised learning. Then, we fine-tune the whole deep networks using an exponential loss function to maximize the separability of the labeled data, by gradient-descent based supervised learning. Experiments on the artificial dataset and real image datasets show that DDBN outperforms most semi-supervised algorithm and deep learning techniques, especially for the hard classification tasks.
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
ICIP2
2010 Reranking for Stacking Ensemble Learning
Buzhou Tang, Qingcai Chen, Xuan Wang 0002, Xiaolong Wang 0001
ICONIP (1)2
2010 Deep Quantum Networks for Classification
abstract
This paper introduces a new type of deep learning method named Deep Quantum Network (DQN) for classification. DQN inherits the capability of modeling the structure of a feature space by fuzzy sets. At first, we propose the architecture of DQN, which consists of quantum neuron and sigmoid neuron and can guide the embedding of samples divisible in new Euclidean space. The parameter of DQN is initialized through greedy layer-wise unsupervised learning. Then, the parameter space of the deep architecture and quantum representation are refined by supervised learning based on the global gradient-descent procedure. An exponential loss function is introduced in this paper to guide the supervised learning procedure. Experiments conducted on standard datasets show that DQN outperforms other feed forward neural networks and neuro-fuzzy classifiers.
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001
ICPR2
2009 STRank: A SiteRank Algorithm Using Semantic Relevance and Time Frequency
abstract
Most of the researches on Web information processing are concentrated on the Web pages and the hyperlinks among them. One of the important facts that a Web page is just one building block of the whole Website had been ignored. But the situation is gradually changed in recent years for the needs of Website reputation calculation, the high level Website structure mining etc. It causes the Website ranking become one of the hot research topics and various site ranking algorithms, such as SiteRank, AggregateRank etc., had been proposed. But most of existing Website ranking algorithm just take use of Website link graphs and the content of Websites are usually not put into consideration. It is obviously not enough for a reliable ranking of Websites. To address this issue, this paper introduces two content based features, i.e., semantic relevance and time frequency and proposes a new STRank algorithm based on these two features. We firstly conduct a series of experiments to verify the feasibility of these two factors in site ranking task. Then the semantic relevance is applied in the calculation of transition probability, and the updating frequency of sites is combined into the ranking task. Since traditional Kendall's ¿ distance and Spearman's footrule distance is not appropriate for the evaluation of site ranking, we make some modifications accordingly to evaluate Website ranking algorithms. Finally, our experiments show that the STRank algorithm outperforms existing approaches on both effectiveness and efficiency.
Hongzhi Guo 0007, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001
SMC2
2009 Text Clustering Approach Based on Maximal Frequent Term Sets
abstract
Classical text clustering algorithms are usually based on vector space model or its variants. Because of the high computing complexity and the difficulty of controlling clustering results, this kind of approaches are hard to be applied for the purpose of the large scale text clustering. Clustering algorithms based on frequent term sets make use of relationship among documents and their shared frequent term sets to achieve high accuracy and effectiveness in clustering. But since the number of frequent terms is usually too large to reach the efficiency requirement for large collection texts clustering, this paper proposes a novel text clustering approach based on maximal frequent term sets (MFTSC). This approach firstly mines maximal frequent term sets from text set and then clusters texts by following steps: at first, the maximal frequent term sets are clustered based on the criterion of k-mismatch; then texts are clustered according to term sets clustering results; finally, we categorize the left texts uncovered in previous step into produced text clusters Be compared with existing approaches, our experimental results show an average gain of 10% on F-Measure score with better performance on scalability and efficiency.
Chong Su, Qingcai Chen, Xiaolong Wang 0001, Xianjun Meng
SMC2
2008 Adaptive filter based prosody modification approach
Qingcai Chen, Shusen Zhou, Xiaohong Yang
INTERSPEECH1
2008 Semantic feature reduction in chinese document clustering
abstract
Text clustering techniques were usually used to structure the text documents into topic related groups which can facilitate users to get a comprehensive understanding on corpus or results from information retrieval system. Most of existing text clustering algorithm which derived from traditional formatted data clustering heavily rely on term analysis methods and adopted vector space model (VSM) as their document representation. But because of the essential characteristic underlying text such as high dimensionality features vector space, the problem of sparseness has a strong impact on the clustering algorithm. So feature reduction is an important preprocess step for improving the efficiency and accuracy of clustering algorithm by removing redundant and irrelevant terms from corpus. Even the clustering is considered as an unsupervised learning method, but in text, there is still some priori knowledge we can use from NLP analysis based approach. In this paper, we propose a semantic analysis based feature reduction method which used in Chinese text clustering. Our method bases on a dedicated Part-of-Speech tags selection and synonyms consolidation and can reduce the feature space of documents more effectively compared with traditional feature reduction method tfidf and stopwords removal; meanwhile it preserves or sometimes even improves the accuracy of clustering algorithm. In our experiment, we tested our feature reduction method using bisecting k-means algorithm which was proved be efficient in text clustering. The results show that our method can reduce the feature space significantly, and meanwhile have a better clustering accuracy in terms of the purity.
Xianjun Meng, Qingcai Chen, Xiaolong Wang 0001
SMC2
2008 Basic semantic units based web page content extraction
abstract
Web page content extraction can be achieved by node-based and segmentation-based algorithms respectively on top of the document object model (DOM). However, the node-based algorithm often removes content embedded as anchor text; while the segmentation-based way can not distinguish irrelevant text from content text when they are divided into the same segment. The two kinds of algorithms don't keep the paragraph information of the original page either. In this paper, a new basic semantic unit (BSU) with granularity between nodes in the DOM tree and content block is defined. Two different methods based on BSU, using clustering and heuristic rules are developed to extract page content. The clustering method gets the best precision 96.88%; while the heuristic rules obtain the best F1-value 95.28%. Compared with the baseline method which uses text blocks segmented byandas Web page content, the F1-values are enhanced by 8.92% and 9.42% respectively.
Qingcai Chen, Xiaolong Wang 0001, Hongzhi Guo 0007
SMC2
2008 Auto Adapted English Pronunciation Evaluation: a Fuzzy Integral Approach
abstract
To evaluate the pronunciation skills of spoken English is one of the key tasks for computer-aided spoken language learning (CALL). While most of the researchers focus on improving the speech recognition techniques to build a reliable evaluation system, another important aspect of this task has been ignored, i.e. the pronunciation evaluation model that integrates both the reliabilities of existing speech processing systems and the learner's pronunciation personalities. To take this aspect into consideration, a Sugeno integral-based evaluation model is introduced in this paper. At first, the English phonemes that are hard to be distinguished (HDP) for Chinese language learners are grouped into different HDP sets. Then, the system reliabilities for distinguishing the phonemes within a HDP set are computed from the standard speech corpus and are integrated with the phoneme recognition results under the Sugeno integral framework. The fuzzy measures are given for each subset of speech segments that contains n occurrences of phonemes within a HDP set. Rather than providing a quantity of scores, the linguistic descriptions of evaluation results are given by the model, which is more helpful for the users to improve their spoken language skills. To get a better performance, generic algorithm (GA)-based parameter optimization is also applied to optimize the model parameters. Experiments are conducted on the Sphinx-4 speech recognition platform. They show that, with 84.7% of average recognition rate of the SR system on standard speech corpus, our pronunciation evaluation model has got reasonable and reliable results for three kinds of test corpora.
Qingcai Chen, Xiaolong Wang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2007 Improving web search ranking by incorporating summarization
abstract
Though link analysis based page ranking approaches have reached great success in commercial search engines (SE), the content based relevance computing approaches also play a very important role in the ranking of information retrieval results. Since most of existing relevance computing algorithms are running on the full text of a web page, this paper is focused on the relevance computing between user’s query and the auto-generated text summarization of each webpage. The first part of this paper provides a brief introduction of the state of art of relevance computing in SE. The inference network approach is especially concerned in this paper since it is the baseline method in our experiment SE system. Then the auto text summarization method based on multi-source integration is introduced, and the full text of each web page is replaced by its auto-generated abstract to compute the relevance between the webpage and user query. To evaluate the effect of the condensation representation of full text on the relevance based page rank of a system, several experiments are conducted in the last part of this paper, which include the method remarked above with different compress ratio, and the full text based ranking. In addition to the efficiency gain of the SE system, the experiment results also shows that the ranking results based on the summary generated by our text summarization system with 30% compress ratio can also get 11.29% of the precision improvement for the SE system.
Xianjun Meng, Qingcai Chen, Xiaolong Wang 0001, Xiao-Hong Yang
SMC2
2004 Mining Pinyin-to-character conversion rules from large-scale corpus: a rough set approach
abstract
This paper introduces a rough set technique for solving the problem of mining Pinyin-to-character (PTC) conversion rules. It first presents a text-structuring method by constructing a language information table from a corpus for each pinyin, which it will then apply to a free-form textual corpus. Data generalization and rule extraction algorithms can then be used to eliminate redundant information and extract consistent PTC conversion rules. The design of our model also addresses a number of important issues such as the long-distance dependency problem, the storage requirements of the rule base, and the consistency of the extracted rules, while the performance of the extracted rules as well as the effects of different model parameters are evaluated experimentally. These results show that by the smoothing method, high precision conversion (0.947) and recall rates (0.84) can be achieved even for rules represented directly by pinyin rather than words. A comparison with the baseline tri-gram model also shows good complement between our method and the tri-gram language model.
Xiaolong Wang 0001, Qingcai Chen, Daniel S. Yeung
IEEE Trans. Syst. Man Cybern. Part B2