EDBT 2026 Demo / reviewers in the wild / expert
Xiaolong Wang 0001
dblp:91/952-1
· DBLP profile ↗
154ranked-venue papers
2as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 86 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 51 · 5 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 13 · 2 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Computer networks · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STPformer: Mutation-Aware Spatial-Temporal Pivotal Attention Networks for Transformer-Based Traffic Forecasting
Hongyang Su, Chenyun Yu, Qingcai Chen, Beibei Kong, Lei Cheng 0005, Chengxiang Zhuo, Zang Li, Xiaolong Wang 0001 |
DASFAA (1) | 8 |
| 2024 | DKINet: Medication Recommendation via Domain Knowledge Informed Deep LearningabstractMedication recommendation is a fundamental yet crucial branch of healthcare that presents opportunities to assist physicians in making more accurate medication prescriptions for patients with complex health conditions. Previous studies have primarily concentrated on deriving patient representations from electronic health records (EHRs) to recommend medications, often overlooking the effective integration of domain-specific prior knowledge. However, integrating domain knowledge with the patient’s clinical manifestations can be challenging, particularly when dealing with complex clinical manifestations. Therefore, in this paper, we first identify comprehensive domain-specific prior knowledge, namely the Unified Medical Language System (UMLS), which is a comprehensive repository of biomedical vocabularies and standards, for knowledge extraction. Subsequently, we propose a knowledge injection module that addresses the effective integration of domain knowledge with complex clinical manifestations, enabling an effective characterization of the health conditions of the patient. Moreover, acknowledging the influence of historical medications on patients’ current treatments, we propose a historical medication-aware patient representation module to capture the longitudinal influence of historical medication information on the representation of current patients. Extensive experiments on three publicly benchmark datasets verify the superiority of our proposed method, which outperformed other methods by a significant margin. The code is available at: https://github.com/sherry6247/DKINet. Sicen Liu, Xiaolong Wang 0001, Xianbing Zhao, Hao Chen 0011 |
BIBM | 2 |
| 2024 | Attention based adaptive spatial-temporal hypergraph convolutional networks for stock price trend prediction
Hongyang Su, Xiaolong Wang 0001, Yang Qin 0001, Qingcai Chen |
Expert Syst. Appl. | 2 |
| 2023 | Efficient Adaptive Spatial-Temporal Attention Network for Traffic Flow Forecasting
Hongyang Su, Xiaolong Wang 0001, Qingcai Chen, Yang Qin 0001 |
ECML/PKDD (5) | 2 |
| 2023 | Learning to generate complex question with intent prediction from long passage
Youcheng Pan, Baotian Hu, Shiyue Wang, Xiaolong Wang 0001, Qingcai Chen, Zenglin Xu, Min Zhang 0005 |
Appl. Intell. | 4 |
| 2023 | VGbel: An exploration of ensemble learning incorporating non-Euclidean structural representation for time series classification
Shaocong Wu, Mengxia Liang, Xiaolong Wang 0001, Qingcai Chen |
Expert Syst. Appl. | 3 |
| 2023 | Improving stock trend prediction through financial time series classification and temporal correlation analysis based on aligning change point
Mengxia Liang, Xiaolong Wang 0001, Shaocong Wu |
Soft Comput. | 2 |
| 2023 | SHAPE: A Sample-Adaptive Hierarchical Prediction Network for Medication RecommendationabstractEffectively medication recommendation with complex multimorbidity conditions is a critical yet challenging task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the encoding format of intra-visit medical events are serialized and information transmitted patterns of learning longitudinal sequence data are stable. However, the following conditions may have been ignored: 1) A more compact encoder for intra-relationship in the intra-visit medical event is urgent; 2) Strategies for learning accurate representations of the variable longitudinal sequences of patients are different. In this article, we proposed a novel Sample-adaptive Hierarchical medicAtion Prediction nEtwork, termed SHAPE, to tackle the above challenges in the medication recommendation task. Specifically, we design a compact intra-visit set encoder to encode the relationship in the medical event for obtaining visit-level representation and then develop an inter-visit longitudinal encoder to learn the patient-level longitudinal representation efficiently. To endow the model with the capability of modeling the variable visit length, we introduce a soft curriculum learning method to assign the difficulty of each sample automatically by the visit length. Extensive experiments on a benchmark dataset verify the superiority of our model compared with several state-of-the-art baselines. Sicen Liu, Xiaolong Wang 0001, Jingcheng Du, Yongshuai Hou, Xianbing Zhao, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Multimodal Data Matters: Language Model Pre-Training Over Structured and Unstructured Electronic Health RecordsabstractAs two important textual modalities in electronic health records (EHR), both structured data (clinical codes) and unstructured data (clinical narratives) have recently been increasingly applied to the healthcare domain. Most existing EHR-oriented studies, however, either focus on a particular modality or integrate data from different modalities in a straightforward manner, which usually treats structured and unstructured data as two independent sources of information about patient admission and ignore the intrinsic interactions between them. In fact, the two modalities are documented during the same encounter where structured data inform the documentation of unstructured data and vice versa. In this paper, we proposed a Medical Multimodal Pre-trained Language Model, named MedM-PLM, to learn enhanced EHR representations over structured and unstructured data and explore the interaction of two modalities. In MedM-PLM, two Transformer-based neural network components are firstly adopted to learn representative characteristics from each modality. A cross-modal module is then introduced to model their interactions. We pre-trained MedM-PLM on the MIMIC-III dataset and verified the effectiveness of the model on three downstream clinical tasks, i.e., medication recommendation, 30-day readmission prediction and ICD coding. Extensive experiments demonstrate the power of MedM-PLM compared with state-of-the-art methods. Further analyses and visualizations show the robustness of our model, which could potentially provide more comprehensive interpretations for clinical decision-making. Sicen Liu, Xiaolong Wang 0001, Yongshuai Hou, Ge Li 0002, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion NetworkabstractDeep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely. Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002 |
IEEE Trans. Multim. | 5 |
| 2022 | Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match OptimizationabstractExtracting arguments for the pre-defined roles is a crucial step for event extraction. Recently, there are some insightful works that view it as a machine reading comprehension problem and achieve significant progress. However, most of them need multi-turns to extract the arguments of each role independently, which ignores the relationships among roles in the same event. To alleviate this problem, we propose a novel Multi-Role Argument Extraction method named MRAE which can exploit the relationship of event roles by extracting all arguments for an event simultaneously. To force MRAE to locate more arguments accurately, we propose an argument match optimization loss based on the minimum risk training to exploit sentence-level F1 score. We conduct experiments on the widely used ACE2005 dataset. The experimental results demonstrate that MRAE outperforms the competitor methods by at least +1.2% F1 score on argument extraction, and also shows superiority on data scarce scenarios. Jingcong Tao, Youcheng Pan, Baotian Hu, Weihua Peng, Cuiyun Han, Xiaolong Wang 0001 |
ICASSP | 7 |
| 2022 | Medical Dialogue Response Generation with Pivotal Information RecallingabstractMedical dialogue generation is an important yet challenging task. Most previous works rely on the attention mechanism and large-scale pretrained language models. However, these methods often fail to acquire pivotal information from the long dialogue history to yield an accurate and informative response, due to the fact that the medical entities usually scatters throughout multiple utterances along with the complex relationships between them. To mitigate this problem, we propose a medical response generation model with Pivotal Information Recalling (MedPIR), which is built on two components, i.e., knowledge-aware dialogue graph encoder and recall-enhanced generator. The knowledge-aware dialogue graph encoder constructs a dialogue graph by exploiting the knowledge relationships between entities in the utterances, and encodes it with a graph attention network. Then, the recall-enhanced generator strengthens the usage of these pivotal information by generating a summary of the dialogue before producing the actual response. Experimental results on two large-scale medical dialogue datasets show that MedPIR outperforms the strong baselines in BLEU scores and medical entities F1 measure. Yu Zhao 0043, Yunxin Li, Yuxiang Wu, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001, Min Zhang 0005 |
KDD | 6 |
| 2022 | Statistical analysis of the community lockdown for COVID-19 pandemicabstractAs the global pandemic of the COVID-19 continues, the statistical modeling and analysis of the spreading process of COVID-19 have attracted widespread attention. Various propagation simulation models have been proposed to predict the spread of the epidemic and the effectiveness of related control measures. These models play an indispensable role in understanding the complex dynamic situation of the epidemic. Most existing work studies the spread of epidemic at two levels including population and agent. However, there is no comprehensive statistical analysis of community lockdown measures and corresponding control effects. This paper performs a statistical analysis of the effectiveness of community lockdown based on the Agent-Level Pandemic Simulation (ALPS) model. We propose a statistical model to analyze multiple variables affecting the COVID-19 pandemic, which include the timings of implementing and lifting lockdown, the crowd mobility, and other factors. Specifically, a motion model followed by ALPS and related basic assumptions is discussed first. Then the model has been evaluated using the real data of COVID-19. The simulation study and comparison with real data have validated the effectiveness of our model. Shaocong Wu, Xiaolong Wang 0001, Jingyong Su |
Appl. Intell. | 2 |
| 2022 | CATNet: Cross-event attention-based time-aware network for medical event prediction
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
Artif. Intell. Medicine | 2 |
| 2022 | A stock time series forecasting approach incorporating candlestick patterns and sequence similarityabstractThis article aims to implement trend forecasting of stock time series based on candlestick patterns and sequence similarity. Financial time series forecasting plays a central role in hedging market risks and optimizing investment portfolios. This is a challenging task, as financial engineering requires the proposed approach to be interpretable, robust, and compatible. It is noted that many published research studies are based on multi-modal data, which makes the prediction approaches increasingly complex, difficult to interpret, and does not allow the migration across different data. Given this situation, it is believed that a candlestick data-based approach is promising. It is already recognized by the technical analyses, prevalent in financial markets, more readily available, and has better interpretability. In this paper, the forecasting approach is divided into two steps. In the first step, sequential pattern mining is used to obtain candlestick patterns from multidimensional candlestick data, and the correlation between different patterns and the corresponding future trends are calculated. In the second step, a new sequence similarity is proposed to match the diverse candlestick sequences with the existing patterns. The method is validated on real data from 800 stocks in the Chinese stock market, which are divided into two groups of experiments, and the average accuracy achieved by the proposed method is 56.04% and 55.56%, which is higher than the SVM model (50.83% and 51.32%) and the LSTM model (50.71% and 50.68%) used for comparison, proving that our work is more stable and accurate. This work is instructive for further research around candlestick data to follow. Mengxia Liang, Shaocong Wu, Xiaolong Wang 0001, Qingcai Chen |
Expert Syst. Appl. | 3 |
| 2022 | Jointly modeling transfer learning of industrial chain information and deep learning for stock predictionabstractThe prediction of stock price has always been a main challenge. The time series of stock price tends to exhibit very strong nonlinear characteristics. In recent years, with the rapid development of deep learning, the ability to automatically extract nonlinear features has significantly attracted scholars’ attention. However, the majority of the relevant studies have concentrated on prediction of the changes of stock market based on the data of the specific stock (e.g., transaction data, financial data, etc.), while those studies ignored the interaction between stocks of different industries, especially the interaction between the stocks of upstream enterprises and downstream enterprises in the industrial chain. This paper aims to propose a combination of transfer learning of industrial chain information and deep learning models, including multilayer perceptron (MLP), recurrent neural network (RNN), long short-term memory (LSTM), and gated recurrent unit (GRU), for stock market prediction. These models are used to predict the trend of the 379 stock market indices by industry in China, and the DM test was employed for validation of the prediction results. It can be concluded that RNNs are not necessarily such an optimal choice for the prediction when dealing with specific time series data, and it could be justified by using the local interpretable model-agnostic explanations (LIME) algorithm. Hence, the MLP was selected to effectively improve the accuracy of the prediction of the stock market indices based on the transfer learning of industrial chain information. The investment strategy is constructed according to the prediction results, and the yield of maturity exceeds that of the buy-and-hold strategy. Dingming Wu 0004, Xiaolong Wang 0001, Shaocong Wu |
Expert Syst. Appl. | 2 |
| 2022 | A hybrid framework based on extreme learning machine, discrete wavelet transform, and autoencoder with feature penalty for stock predictionabstractAccurate prediction of the stock market trend can assist efficient portfolio and risk management. In recent years, with the rapid development of deep learning, it can make the classifiers more robust, which can be used for solving nonlinear problems. In our previous research, we proposed a numerical model for predicting the stock market via combination of discrete wavelet transform (DWT) denoising and extreme learning machine (ELM), and promising outcomes are achieved. The current research presents a hybrid framework using DWT, ELM, and autoencoder (AE) with feature penalty. Firstly, the backpropagation of the AE with feature penalty was deduced theoretically. Then, the raw data were denoised by DWT. The denoised data were used to train the AE with feature penalty after feature preprocessing and utilization of the labeling method. Afterward, the encoder part of the well-trained AE was utilized as the feature extraction model to train ELM model, and the hybrid framework named DAELM (DWT-AE-ELM) could be successfully developed. We also carried out experiments on the corresponding dataset of 400 stocks, and the prediction accuracy of the current study was higher than that of our previous research. According to the predicted labels, we presented an investment strategy, and the yield-to-maturity of 400 stocks was significantly higher than that of the buy-and-hold (BAH) strategy. The results confirmed the superiority of the proposed hybrid framework. Dingming Wu 0004, Xiaolong Wang 0001, Shaocong Wu |
Expert Syst. Appl. | 2 |
| 2022 | Multi-channel fusion LSTM for medical event prediction using EHRs
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
J. Biomed. Informatics | 2 |
| 2022 | Construction of stock portfolios based on k-means clustering of continuous trend features
Dingming Wu 0004, Xiaolong Wang 0001, Shaocong Wu |
Knowl. Based Syst. | 2 |
| 2021 | AGCNT: Adaptive Graph Convolutional Network for Transformer-based Long Sequence Time-Series ForecastingabstractLong sequence time-series forecasting(LSTF) plays an important role in a variety of real-world application scenarios, such as electricity forecasting, weather forecasting, and traffic flow forecasting. It has previously been observed that transformer-based models have achieved outstanding results on LSTF tasks, which can reduce the complexity of the model and maintain stable prediction accuracy. Nevertheless, there are still some issues that limit the performance of transformer-based models for LSTF tasks: (i) the potential correlation between sequences is not considered; (ii) the inherent structure of encoder-decoder is difficult to expand after being optimized from the aspect of complexity. In order to solve these two problems, we propose a transformer-based model, named AGCNT, which is efficient and can capture the correlation between the sequences in the multivariate LSTF task without causing the memory bottleneck. Specifically, AGCNT has several characteristics: (i) a probsparse adaptive graph self-attention, which maps long sequences into a low-dimensional dense graph structure with an adaptive graph generation and captures the relationships between sequences with an adaptive graph convolution; (ii) the stacked encoder with distilling probsparse graph self-attention integrates the graph attention mechanism and retains the dominant attention of the cascade layer, which preserves the correlation between sparse queries from long sequences; (iii) the stacked decoder with generative inference generates all prediction values in one forward operation, which can improve the inference speed of long-term predictions. Experimental results on 4 large-scale datasets demonstrate the AGCNT outperforms state-of-the-art baselines. Hongyang Su, Xiaolong Wang 0001, Yang Qin 0001 |
CIKM | 2 |
| 2021 | Enriching BERT With Knowledge Graph Embedding For Industry Classification
Shiyue Wang, Youcheng Pan, Zhenran Xu, Baotian Hu, Xiaolong Wang 0001 |
ICONIP (6) | 5 |
| 2021 | MSDF: A General Open-Domain Multi-skill Dialog Framework
Yu Zhao 0043, Xinshuo Hu, Yunxin Li, Baotian Hu, Dongfang Li 0002, Sichao Chen, Xiaolong Wang 0001 |
NLPCC (2) | 7 |
| 2021 | Improving deep learning method for biomedical named entity recognition by using entity definition informationabstractBACKGROUND: Biomedical named entity recognition (NER) is a fundamental task of biomedical text mining that finds the boundaries of entity mentions in biomedical text and determines their entity type. To accelerate the development of biomedical NER techniques in Spanish, the PharmaCoNER organizers launched a competition to recognize pharmacological substances, compounds, and proteins. Biomedical NER is usually recognized as a sequence labeling task, and almost all state-of-the-art sequence labeling methods ignore the meaning of different entity types. In this paper, we investigate some methods to introduce the meaning of entity types in deep learning methods for biomedical NER and apply them to the PharmaCoNER 2019 challenge. The meaning of each entity type is represented by its definition information. MATERIAL AND METHOD: We investigate how to use entity definition information in the following two methods: (1) SQuad-style machine reading comprehension (MRC) methods that treat entity definition information as query and biomedical text as context and predict answer spans as entities. (2) Span-level one-pass (SOne) methods that predict entity spans of one type by one type and introduce entity type meaning, which is represented by entity definition information. All models are trained and tested on the PharmaCoNER 2019 corpus, and their performance is evaluated by strict micro-average precision, recall, and F1-score. RESULTS: Entity definition information brings improvements to both SQuad-style MRC and SOne methods by about 0.003 in micro-averaged F1-score. The SQuad-style MRC model using entity definition information as query achieves the best performance with a micro-averaged precision of 0.9225, a recall of 0.9050, and an F1-score of 0.9137, respectively. It outperforms the best model of the PharmaCoNER 2019 challenge by 0.0032 in F1-score. Compared with the state-of-the-art model without using manually-crafted features, our model obtains a 1% improvement in F1-score, which is significant. These results indicate that entity definition information is useful for deep learning methods on biomedical NER. CONCLUSION: Our entity definition information enhanced models achieve the state-of-the-art micro-average F1 score of 0.9137, which implies that entity definition information has a positive impact on biomedical NER detection. In the future, we will explore more entity definition information from knowledge graph. Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010, Yi Zhou 0005 |
BMC Bioinform. | 5 |
| 2020 | KEoG: A knowledge-aware edge-oriented graph neural network for document-level relation extractionabstractDocument-level relation extraction (RE) has attracted more and more attentions recently. Edge-oriented graph neural network (EoG) is a new neural network exhibiting greater potential than previous node-oriented graph neural networks for document-level RE. In this paper, we propose a novel EoG, called knowledge-aware edge-oriented GNN (KEoG) for document-level RE. In KEoG, we further introduce not only two types of nodes to represent documents and external knowledge respectively, but also soft F-Measure loss function to solve the inherent class imbalance problem in document-level RE. Experiments conducted on two document-level datasets show that KEoG outperforms other state-of-the-art methods for comparison on both intra-sentence and inter-sentence relation extractions, indicating that KEoG is an effective extension of EoG. Weihua Peng, Qingcai Chen, Xiaolong Wang 0001, Buzhou Tang |
BIBM | 4 |
| 2020 | MedWriter: Knowledge-Aware Medical Text GenerationabstractTo exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical.However, most of recent works only consider the information of unstructured text rather than structured information of the knowledge graph.In this paper, we focus on the medical topic-to-text generation task and adapt a knowledge-aware text generation model to the medical domain, named MedWriter, which not only introduces the specific knowledge from the external MKG but also is capable of learning graph-level representation.We conduct experiments on a medical literature dataset collected from medical journals, each of which has a set of topic words, an abstract of medical literature and a corresponding knowledge graph from CMeKG.Experimental results demonstrate incorporating knowledge graph into generation model can improve the quality of the generated text and has robust superiority over the competitor methods. Youcheng Pan, Qingcai Chen, Weihua Peng, Xiaolong Wang 0001, Baotian Hu, Xin Liu 0054, Wenxiu Zhou |
COLING | 4 |
| 2020 | Learning to Generate Diverse Questions from KeywordsabstractDiverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., generating a series of questions for each set of keywords (one-to-many). As an effort towards this, we propose a novel neural generative model, which incorporates context information and control signal to produce multiple diverse questions from a given fixed set of keywords. The control signal is designed to increase the diversity of questions by capturing the diverse patterns from the entire dataset. The context information is used to guarantee the generated questions are highly related to the given keywords. To evaluate the effectiveness of the proposed model, we collect a dataset which contains 62835 questions with respect to 12567 sets of keywords.1To the best of our knowledge, it's the first Chinese financial dataset for diverse question generation. The experimental results show that our model outperforms the competitor methods in terms of BLEU and Distinct. The qualitative evaluation indicates that our model is able to generate diverse and meaningful questions. Youcheng Pan, Baotian Hu, Qingcai Chen, Yang Xiang 0003, Xiaolong Wang 0001 |
ICASSP | 5 |
| 2020 | Gated Semantic Difference Based Sentence Semantic Equivalence IdentificationabstractThis article proposes a novel sentence semantic equivalence identification (SSEI) method by using the semantic difference features between sentences. The lexical differences of a sentence pair are first extracted, and the bidirectional long short term memory (BiLSTM) network is then applied on them to generate the semantic difference representations. Finally, an efficient gate mechanism is proposed to integrate the semantic differences with existing models (called base model) to enhance their encoding capability in the SSEI task. Exhaustive experiments conducted on the standard Quora corpus, and the Large-scale Chinese Question Matching Corpus (LCQMC) show that the proposed gated semantic difference (GSD) method brings significant improvement for different existing state-of-the-art models. When the bidirectional encoder representations from transformers model (BERT) is used as the base model, the accuracy for SSEI on Quora is improved from 90.63% to 91.98%, and the F1 score on the LCQMC is improved from 87.0% to 87.7%, which outperforms the best-published results. Xin Liu 0054, Qingcai Chen, Xiangping Wu 0001, Yang Hua 0004, Dongfang Li 0002, Buzhou Tang, Xiaolong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2019 | De-identification of Clinical Text via Bi-LSTM-CRF with Neural Language Models
Buzhou Tang, Dehuan Jiang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010, Ying Shen 0001 |
AMIA | 4 |
| 2019 | A Study on Automatic Generation of Chinese Discharge SummaryabstractDischarge summary, which summarizes a patient's health information during hospitalization, is very important for transferring information between the hospitalist and primary care physician. Discharge summary writing is a necessary but time-consuming job for physicians. Automatically generating discharge summaries using information technology is helpful to physicians, but challenging as discharge summaries are typically long and contain amounts of information to outline patient's reason for admission, labtests, examinations, diagnostic findings, treatments, and medication care plan. In this study, we propose a framework based on deep learning methods for automatic discharge summary generation. In the framework, we split information of discharge summary into two parts: 1) history information such as reason for admission, admission diagnosis; 2) outcome information such as discharge diagnosis and medication care plan, and deploy different neural networks to generate them separately. Hierarchical sequence labeling methods are proposed to select key sentences from existing documents, e.g., admission notes, progress notes, examination reports as the history information, and multi-task learning methods to predict the outcomes, e.g., diagnosis and medication care plan. Experiments on a Chinese corpus show that our approach has the ability to generate effective discharge summaries. Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Jun Yan 0010 |
BIBM | 4 |
| 2019 | A comprehensive review and comparison of existing computational methods for intrinsically disordered protein and region predictionabstractIntrinsically disordered proteins and regions are widely distributed in proteins, which are associated with many biological processes and diseases. Accurate prediction of intrinsically disordered proteins and regions is critical for both basic research (such as protein structure and function prediction) and practical applications (such as drug development). During the past decades, many computational approaches have been proposed, which have greatly facilitated the development of this important field. Therefore, a comprehensive and updated review is highly required. In this regard, we give a review on the computational methods for intrinsically disordered protein and region prediction, especially focusing on the recent development in this field. These computational approaches are divided into four categories based on their methodologies, including physicochemical-based method, machine-learning-based method, template-based method and meta method. Furthermore, their advantages and disadvantages are also discussed. The performance of 40 state-of-the-art predictors is directly compared on the target proteins in the task of disordered region prediction in the 10th Critical Assessment of protein Structure Prediction. A more comprehensive performance comparison of 45 different predictors is conducted based on seven widely used benchmark data sets. Finally, some open problems and perspectives are discussed. Xiaolong Wang 0001, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2019 | Extracting entities with attributes in clinical text via joint deep learningabstractOBJECTIVE: Extracting clinical entities and their attributes is a fundamental task of natural language processing (NLP) in the medical domain. This task is typically recognized as 2 sequential subtasks in a pipeline, clinical entity or attribute recognition followed by entity-attribute relation extraction. One problem of pipeline methods is that errors from entity recognition are unavoidably passed to relation extraction. We propose a novel joint deep learning method to recognize clinical entities or attributes and extract entity-attribute relations simultaneously. MATERIALS AND METHODS: The proposed method integrates 2 state-of-the-art methods for named entity recognition and relation extraction, namely bidirectional long short-term memory with conditional random field and bidirectional long short-term memory, into a unified framework. In this method, relation constraints between clinical entities and attributes and weights of the 2 subtasks are also considered simultaneously. We compare the method with other related methods (ie, pipeline methods and other joint deep learning methods) on an existing English corpus from SemEval-2015 and a newly developed Chinese corpus. RESULTS: Our proposed method achieves the best F1 of 74.46% on entity recognition and the best F1 of 50.21% on relation extraction on the English corpus, and 89.32% and 88.13% on the Chinese corpora, respectively, which outperform the other methods on both tasks. CONCLUSIONS: The joint deep learning-based method could improve both entity recognition and relation extraction from clinical text in both English and Chinese, indicating that the approach is promising. Xue Shi, Yingping Yi, Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Zongcheng Ji, Yaoyun Zhang, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Cohort selection for clinical trials using hierarchical neural networkabstractOBJECTIVE: Cohort selection for clinical trials is a key step for clinical research. We proposed a hierarchical neural network to determine whether a patient satisfied selection criteria or not. MATERIALS AND METHODS: We designed a hierarchical neural network (denoted as CNN-Highway-LSTM or LSTM-Highway-LSTM) for the track 1 of the national natural language processing (NLP) clinical challenge (n2c2) on cohort selection for clinical trials in 2018. The neural network is composed of 5 components: (1) sentence representation using convolutional neural network (CNN) or long short-term memory (LSTM) network; (2) a highway network to adjust information flow; (3) a self-attention neural network to reweight sentences; (4) document representation using LSTM, which takes sentence representations in chronological order as input; (5) a fully connected neural network to determine whether each criterion is met or not. We compared the proposed method with its variants, including the methods only using the first component to represent documents directly and the fully connected neural network for classification (denoted as CNN-only or LSTM-only) and the methods without using the highway network (denoted as CNN-LSTM or LSTM-LSTM). The performance of all methods was measured by micro-averaged precision, recall, and F1 score. RESULTS: The micro-averaged F1 scores of CNN-only, LSTM-only, CNN-LSTM, LSTM-LSTM, CNN-Highway-LSTM, and LSTM-Highway-LSTM were 85.24%, 84.25%, 87.27%, 88.68%, 88.48%, and 90.21%, respectively. The highest micro-averaged F1 score is higher than our submitted 1 of 88.55%, which is 1 of the top-ranked results in the challenge. The results indicate that the proposed method is effective for cohort selection for clinical trials. DISCUSSION: Although the proposed method achieved promising results, some mistakes were caused by word ambiguity, negation, number analysis and incomplete dictionary. Moreover, imbalanced data was another challenge that needs to be tackled in the future. CONCLUSION: In this article, we proposed a hierarchical neural network for cohort selection. Experimental results show that this method is good at selecting cohort. Xue Shi, Dehuan Jiang, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Jun Yan 0010 |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Dynamic Working Memory for Context-Aware Response GenerationabstractIn human-to-human conversations, the context generally provides several backgrounds and strategic points for the following response. Therefore, many response generation approaches have explored the methodologies to incorporate the context into the encoder-decoder architecture, to generate context-aware responses that are remarkably relevant and cohesive to the given context. However, most approaches pay less attention to semantic interactions implicitly existing within contextual utterances, which are of great importance to capture semantic clues of the given dialog context, indeed. This paper proposes a dynamic working memory mechanism to model long-term semantic hints in the conversation context, by performing semantic interactions between utterances and updating context representation dynamically. Then, the outputs of the dynamic working memory are employed to provide helpful clues for the encoder-decoder architecture to generate responses to the given dialog. We have evaluated the proposed approach on Twitter Customer Service Corpus and OpenSubtitles Corpus, with several automatic evaluation metrics and the human evaluation, and the empirical results show the effectiveness of the proposed method. Zhen Xu 0003, Chengjie Sun, Yinong Long, Bingquan Liu, Baoxun Wang, Mingjiang Wang, Min Zhang 0005, Xiaolong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2019 | Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency MatrixabstractProtein remote homology detection and fold recognition are two critical tasks for the studies of protein structures and functions. Currently, the profile-based methods achieve the state-of-the-art performance in these fields. However, the widely used sequence profiles, like position-specific frequency matrix (PSFM) and position-specific scoring matrix (PSSM), ignore the sequence-order effects along protein sequence. In this study, we have proposed a novel profile, called sequence-order frequency matrix (SOFM), to extract the sequence-order information of neighboring residues from multiple sequence alignment (MSA). Combined with two profile feature extraction approaches, top-n-grams and the Smith-Waterman algorithm, the SOFMs are applied to protein remote homology detection and fold recognition, and two predictors called SOFM-Top and SOFM-SW are proposed. Experimental results show that SOFM contains more information content than other profiles, and these two predictors outperform other state-of-the-art methods. It is anticipated that SOFM will become a very useful profile in the studies of protein structures and functions. Bin Liu 0014, Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | LSDSCC: a Large Scale Domain-Specific Conversational Corpus for Response Generation with Diversity Oriented Evaluation MetricsabstractZhen Xu, Nan Jiang, Bingquan Liu, Wenge Rong, Bowen Wu, Baoxun Wang, Zhuoran Wang, Xiaolong Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Zhen Xu 0003, Nan Jiang 0010, Bingquan Liu, Wenge Rong, Bowen Wu 0001, Baoxun Wang, Xiaolong Wang 0001 |
NAACL-HLT | 8 |
| 2018 | A comprehensive review and comparison of different computational methods for protein remote homology detectionabstractProtein remote homology detection is one of the most fundamental and central problems for the studies of protein structures and functions, aiming to detect the distantly evolutionary relationships among proteins via computational methods. During the past decades, many computational approaches have been proposed to solve this important task. These methods have made a substantial contribution to protein remote homology detection. Therefore, it is necessary to give a comprehensive review and comparison on these computational methods. In this article, we divide these computational approaches into three categories, including alignment methods, discriminative methods and ranking methods. Their advantages and disadvantages are discussed in a comprehensive perspective, and their performance is compared on widely used benchmark data sets. Finally, some open questions in this field are further explored and discussed. Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2018 | Entity disambiguation with memory network
Yaming Sun, Zhenzhou Ji, Lei Lin 0001, Xiaolong Wang 0001, Duyu Tang |
Neurocomputing | 4 |
| 2018 | Recurrent convolutional neural network for answer selection in community question answering
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001 |
Neurocomputing | 4 |
| 2018 | Structural regularity exploration in multidimensional networks via Bayesian inference
Yi Chen 0019, Xiaolong Wang 0001, Buzhou Tang |
Neural Comput. Appl. | 2 |
| 2018 | Learning to recognize opinion targets using recurrent neural networks
Yuanchao Liu, Xiaolong Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Content-Oriented User Modeling for Personalized Response Ranking in ChatbotsabstractAutomatic chatbots (also known as chat-agents) have attracted much attention from both researching and industrial fields. Generally, the semantic relevance between users' queries and the corresponding responses is considered as the essential element for conversation modeling in both generation and ranking based chat systems. By contrast, it is a nontrivial task to adopt the users' information, such as preference, social role, etc., into conversational models reasonably, while users' profiles play a significant role in the procedure of conversations by providing the implicit contexts. This paper aims to address the personalized response ranking task by incorporating user profiles into the conversation model. In our approach, users' personalized representations are latently learned from the contents posted by them via a two-branch neural network. After that, a deep neural network architecture is further presented to learn the fusion representation of posts, responses, and personal information. In this way, the proposed model could understand conversations from the users' perspective; hence, the more appropriate responses are selected for a specified person. The experimental results on two datasets from social network services demonstrate that our approach is hopeful to represent users' personal information implicitly based on user generated contents, and it is promising to perform as an important component in chatbots to select the personalized responses for each user. Bingquan Liu, Zhen Xu 0003, Chengjie Sun, Baoxun Wang, Xiaolong Wang 0001, Derek F. Wong, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Recognizing Continuous and Discontinuous Adverse Drug Reaction Mentions from Social Media Using LSTM-CRFabstractSocial media in medicine, where patients can express their personal treatment experiences by personal computers and mobile devices, usually contains plenty of useful medical information, such as adverse drug reactions (ADRs); mining this useful medical information from social media has attracted more and more attention from researchers. In this study, we propose a deep neural network (called LSTM‐CRF) combining long short‐term memory (LSTM) neural networks (a type of recurrent neural networks) and conditional random fields (CRFs) to recognize ADR mentions from social media in medicine and investigate the effects of three factors on ADR mention recognition. The three factors are as follows: (1) representation for continuous and discontinuous ADR mentions: two novel representations, that is, “BIOHD” and “Multilabel,” are compared; (2) subject of posts: each post has a subject (i.e., drug here); and (3) external knowledge bases. Experiments conducted on a benchmark corpus, that is, CADEC, show that LSTM‐CRF achieves better F ‐score than CRF; “Multilabel” is better in representing continuous and discontinuous ADR mentions than “BIOHD”; both subjects of comments and external knowledge bases are individually beneficial to ADR mention recognition. To the best of our knowledge, this is the first time to investigate deep neural networks to mine continuous and discontinuous ADRs from social media. Buzhou Tang, Jianglu Hu, Xiaolong Wang 0001, Qingcai Chen |
Wirel. Commun. Mob. Comput. | 3 |
| 2017 | Chemical-induced disease extraction via convolutional neural networks with attentionabstractExtracting relationships between chemicals and diseases from unstructured literature is very important for many biomedical applications such as pharmacovigilance and drug repositioning. Automatic chemical-induced disease extraction is usually recognized as a classification task, and several systems have been proposed for this task recently due to some annotated corpora publicly available. Most of the systems are based on machine learning methods with many manually-crafted features. In recent years, deep learning that does not only can avoid verbose feature engineering but also shows competitive performance has been widely used in various types of tasks, including classification task. Therefore, deep learning has great potential on chemical-induced disease extraction. In this paper, we proposed an architecture of convolutional neural networks (CNN) with attention mechanism for chemical-induced disease extraction, which does not only avoid verbose feature engineering but also integrates domain knowledge in a simple way. Experiments on a benchmark dataset demonstrate that the proposed CNN-based chemical-induced disease extraction system is competitive with other state-of-the-art systems. Haodi Li, Qingcai Chen, Buzhou Tang, Xiaolong Wang 0001 |
BIBM | 4 |
| 2017 | Neural Response Generation via GAN with an Approximate Embedding LayerabstractThis paper presents a Generative Adversarial Network (GAN) to model singleturn short-text conversations, which trains a sequence-to-sequence (Seq2Seq) network for response generation simultaneously with a discriminative classifier that measures the differences between human-produced responses and machinegenerated ones.In addition, the proposed method introduces an approximate embedding layer to solve the non-differentiable problem caused by the sampling-based output decoding procedure in the Seq2Seq generative model.The GAN setup provides an effective way to avoid noninformative responses (a.k.a "safe responses"), which are frequently observed in traditional neural response generators.The experimental results show that the proposed approach significantly outperforms existing neural response generation models in diversity metrics, with slight increases in relevance scores as well, when evaluated on both a Mandarin corpus and an English corpus. Zhen Xu 0003, Bingquan Liu, Baoxun Wang, Chengjie Sun, Xiaolong Wang 0001 |
EMNLP | 5 |
| 2017 | SOFM-Top: Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency Matrix
Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014 |
ICIC (2) | 3 |
| 2017 | Predicting Users' Negative Feedbacks in Multi-Turn Human-Computer DialoguesabstractUser experience is essential for human-computer dialogue systems. However, it is impractical to ask users to provide explicit feedbacks when the agents’ responses displease them. Therefore, in this paper, we explore to predict users’ imminent dissatisfactions caused by intelligent agents by analysing the existing utterances in the dialogue sessions. To our knowledge, this is the first work focusing on this task. Several possible factors that trigger negative emotions are modelled. A relation sequence model (RSM) is proposed to encode the sequence of appropriateness of current response with respect to the earlier utterances. The experimental results show that the proposed structure is effective in modelling emotional risk (possibility of negative feedback) than existing conversation modelling approaches. Besides, strategies of obtaining distance supervision data for pre-training are also discussed in this work. Balanced sampling with respect to the last response in the distance supervision data are shown to be reliable for data augmentation. Xin Wang 0017, Yuanchao Liu, Xiaolong Wang 0001, Baoxun Wang |
IJCNLP(1) | 4 |
| 2017 | Incorporating loose-structured knowledge into conversation modeling via recall-gate LSTMabstractIt is critical for automatic chat-bots to gain the ability of conversation comprehension, which is the essence to provide context-aware responses to conduct smooth dialogues with human beings. As the basis of this task, conversation modeling will notably benefit from the background knowledge, since such knowledge indeed implicates semantic hints that help to further clarify the relationships between sentences within a conversation. In this paper, a deep neural network is proposed to incorporate background knowledge for conversation modeling. Through a recall mechanism with a specially designed recall-gate, background knowledge as global memory can be motivated to cooperate with local cell memory of Long Short-Term Memory (LSTM), so as to enrich the ability of LSTM to capture the implicit semantic clues in conversations. In addition, this paper introduces the loose-structured domain knowledge as background knowledge, which can be built with slight amount of manual work and easily adopted by the recall-gate. Our model is evaluated on the context-oriented response selecting task, and experimental results on two datasets have shown that our approach is promising for modeling conversations and building key components of automatic chat systems. Zhen Xu 0003, Bingquan Liu, Baoxun Wang, Chengjie Sun, Xiaolong Wang 0001 |
IJCNN | 5 |
| 2017 | CNN-based ranking for biomedical entity normalizationabstractBACKGROUND: Most state-of-the-art biomedical entity normalization systems, such as rule-based systems, merely rely on morphological information of entity mentions, but rarely consider their semantic information. In this paper, we introduce a novel convolutional neural network (CNN) architecture that regards biomedical entity normalization as a ranking problem and benefits from semantic information of biomedical entities. RESULTS: The CNN-based ranking method first generates candidates using handcrafted rules, and then ranks the candidates according to their semantic information modeled by CNN as well as their morphological information. Experiments on two benchmark datasets for biomedical entity normalization show that our proposed CNN-based ranking method outperforms traditional rule-based method with state-of-the-art performance. CONCLUSIONS: We propose a CNN architecture that regards biomedical entity normalization as a ranking problem. Comparison results show that semantic information is beneficial to biomedical entity normalization and can be well combined with morphological information in our CNN architecture for further improvement. Haodi Li, Qingcai Chen, Buzhou Tang, Xiaolong Wang 0001, Hua Xu 0001 |
BMC Bioinform. | 4 |
| 2017 | Answer Selection in Community Question Answering via Attentive Neural NetworksabstractAnswer selection in community question answering (cQA) is a challenging task in natural language processing. The difficulty lies in that it not only needs the consideration of semantic matching between question answer pairs but also requires a serious modeling of contextual factors. In this letter, we propose an attentive deep neural network architecture so as to learn the deterministic information for answer selection. The architecture can support various input formats through the organization of convolutional neural networks, attention-based long short-term memory, and conditional random fields. Experiments are carried out on the SemEval-2015 cQA dataset. We attain 58.35% on macroaveraged F1, which outperforms the Top-1 system in the shared task by 1.16% and improves the state-of-the-art deep-neural-network-based method by 2.21%. Yang Xiang 0003, Qingcai Chen, Xiaolong Wang 0001, Yang Qin 0001 |
IEEE Signal Process. Lett. | 3 |
| 2016 | Write-righter: An Academic Writing Assistant SystemabstractWriting academic articles in English is a challenging task for non-native speakers, as more effort has to be spent to enhance their language expressions. This paper presents an academic writing assistant system called Write-righter, which can provide real-time hint and recommendation by analyzing the input context. To achieve this goal, some novel strategies, e.g., semantic extension based sentence retrieval and LDA based sentence structure identification have been proposed. Write-righter is expected to help people express their ideas correctly by recommending top N most possible expressions. Yuanchao Liu, Xin Wang 0017, Ming Liu 0004, Xiaolong Wang 0001 |
AAAI | 4 |
| 2016 | CMedTEX: A Rule-based Temporal Expression Extraction and Normalization System for Chinese Clinical Notes
Zengjian Liu, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Haodi Li, Junzhao Bu, Jingzhi Jiang, Qiwen Deng, Suisong Zhu |
AMIA | 3 |
| 2016 | Incorporating Label Dependency for Answer Quality Tagging in Community Question Answering via CNN-LSTM-CRFabstractIn community question answering (cQA), the quality of answers are determined by the matching degree between question-answer pairs and the correlation among the answers. In this paper, we show that the dependency between the answer quality labels also plays a pivotal role. To validate the effectiveness of label dependency, we propose two neural network-based models, with different combination modes of Convolutional Neural Net-works, Long Short Term Memory and Conditional Random Fields. Extensive experi-ments are taken on the dataset released by the SemEval-2015 cQA shared task. The first model is a stacked ensemble of the networks. It achieves 58.96% on macro averaged F1, which improves the state-of-the-art neural network-based method by 2.82% and outper-forms the Top-1 system in the shared task by 1.77%. The second is a simple attention-based model whose input is the connection of the question and its corresponding answers. It produces promising results with 58.29% on overall F1 and gains the best performance on the Good and Bad categories. Yang Xiang 0003, Xiaoqiang Zhou, Qingcai Chen, Zhihui Zheng, Buzhou Tang, Xiaolong Wang 0001, Yang Qin 0001 |
COLING | 6 |
| 2016 | App relationship calculation: An iterative processabstractToday, plenty of apps are released to help users make the best use of their mobile phones. Facing the large amount of apps, app retrieval and app recommendation are extensively adopted to help users obtain their favorite apps. To acquire the high-quality retrieval or recommending results, it needs to obtain the accurate app relationship calculating results in advance. Unfortunately, recent methods are conducted mostly depending on user's log or app's contexts, which can only detect whether two apps are downloaded, installed meanwhile or provide similar functions or not. In fact, apps contain many deep relationships other than similarity, e.g., one app needs another app to cooperate to fulfill its work. Obviously, app's reviews contain user's viewpoint. They are useful to help dig deep relationship between apps. Therefore, to calculate relationship between apps via reviews, we propose an iterative process by combining review similarity calculation and app relationship calculation together. Ming Liu 0004, Chong Wu 0001, Xiang-Nan Zhao, Chin-Yew Lin, Xiaolong Wang 0001 |
ICDE | 5 |
| 2016 | Extended Dependency-Based Word Embeddings for Aspect Extraction
Xin Wang 0017, Yuanchao Liu, Chengjie Sun, Ming Liu 0004, Xiaolong Wang 0001 |
ICONIP (4) | 5 |
| 2015 | Predicting Polarities of Tweets by Composing Word Embeddings with Long Short-Term MemoryabstractXin Wang, Yuanchao Liu, Chengjie Sun, Baoxun Wang, Xiaolong Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Xin Wang 0017, Yuanchao Liu, Chengjie Sun, Baoxun Wang, Xiaolong Wang 0001 |
ACL (1) | 5 |
| 2015 | Recognizing Disjoint Clinical Concepts in Clinical Text Using Machine Learning-based Methods
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001 |
AMIA | 3 |
| 2015 | Structural Regularity Exploration in Multidimensional Networks
Yi Chen 0019, Xiaolong Wang 0001, Buzhou Tang, Junzhao Bu, Qingcai Chen |
ICONIP (3) | 2 |
| 2015 | User Recommendation Based on Network Structure in Social Networks
Yi Chen 0019, Xiaolong Wang 0001, Buzhou Tang, Junzhao Bu |
ICONIP (3) | 2 |
| 2015 | Multimodal Deep Belief Network Based Link Prediction and User Comment Generation
Feng Liu 0041, Bingquan Liu, Chengjie Sun, Ming Liu 0004, Xiaolong Wang 0001 |
ICONIP (4) | 5 |
| 2015 | Distant Supervision for Relation Extraction via Group Selection
Yang Xiang 0003, Xiaolong Wang 0001, Yaoyun Zhang, Yang Qin 0001, Shixi Fan |
ICONIP (2) | 2 |
| 2015 | An Auto-Encoder for Learning Conversation Representation Using LSTM
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001 |
ICONIP (1) | 4 |
| 2015 | VRCA: A Clustering Algorithm for Massive Amount of Texts
Ming Liu 0004, Lei Chen 0072, Bingquan Liu, Xiaolong Wang 0001 |
IJCAI | 4 |
| 2015 | Modeling Mention, Context and Entity with Neural Networks for Entity Disambiguation
Yaming Sun, Lei Lin 0001, Duyu Tang, Nan Yang 0002, Zhenzhou Ji, Xiaolong Wang 0001 |
IJCAI | 6 |
| 2015 | Multimodal Learning Based Approaches for Link Prediction in Social NetworksabstractThe link prediction problem in social networks is to estimate the value of the link that can represent relationship between social members. Researchers have proposed several methods for solving link prediction and a number of features have been used. Most of these models are learned with only considering the features from one kind of data. In this paper, by considering the data from link network structure and user comment, both of which could imply the concept of link value, we propose multimodal learning based approaches to predict the link values. The experiment results done on dataset from typical social networks show that our model could learn the joint representation of these datas properly, and the method MDBN outperforms other state-of-art link prediction methods. Feng Liu 0041, Bingquan Liu, Chengjie Sun, Ming Liu 0004, Xiaolong Wang 0001 |
NLPCC | 5 |
| 2015 | Convolutional Neural Networks for Correcting English Article ErrorsabstractIn this paper, convolutional neural networks are employed for English article error correction. Instead of employing features relying on human ingenuity and prior natural language processing knowledge, the words surrounding the context of the article are taken as features. Our approach could be trained both on an error annotated corpus and an error non-annotated corpus. Experiments are conducted on CoNLL-2013 data set. Our approach achieves 38.10 % in F1, and outperforms the best system (33.40 %) that participates in the task. Experimental results demonstrate the effectiveness of our proposed approach. Chengjie Sun, Xiaoqiang Jin, Lei Lin 0001, Xiaolong Wang 0001 |
NLPCC | 5 |
| 2015 | Computing Semantic Text Similarity Using Rich Features
Yang Liu 0054, Chengjie Sun, Lei Lin 0001, Xiaolong Wang 0001 |
PACLIC | 4 |
| 2015 | Application of learning to rank to protein remote homology detectionabstractMOTIVATION: Protein remote homology detection is one of the fundamental problems in computational biology, aiming to find protein sequences in a database of known structures that are evolutionarily related to a given query protein. Some computational methods treat this problem as a ranking problem and achieve the state-of-the-art performance, such as PSI-BLAST, HHblits and ProtEmbed. This raises the possibility to combine these methods to improve the predictive performance. In this regard, we are to propose a new computational method called ProtDec-LTR for protein remote homology detection, which is able to combine various ranking methods in a supervised manner via using the Learning to Rank (LTR) algorithm derived from natural language processing. RESULTS: Experimental results on a widely used benchmark dataset showed that ProtDec-LTR can achieve an ROC1 score of 0.8442 and an ROC50 score of 0.9023 outperforming all the individual predictors and some state-of-the-art methods. These results indicate that it is correct to treat protein remote homology detection as a ranking problem, and predictive performance improvement can be achieved by combining different ranking approaches in a supervised manner via using LTR. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the software tools of three basic ranking predictors and Learning to Rank algorithm were provided at http://bioinformatics.hitsz.edu.cn/ProtDec-LTR/home/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Junjie Chen 0004, Xiaolong Wang 0001 |
Bioinform. | 3 |
| 2015 | repDNA: a Python package to generate various modes of feature vectors for DNA sequences by incorporating user-defined physicochemical properties and sequence-order effectsabstractUNLABELLED: In order to develop powerful computational predictors for identifying the biological features or attributes of DNAs, one of the most challenging problems is to find a suitable approach to effectively represent the DNA sequences. To facilitate the studies of DNAs and nucleotides, we developed a Python package called representations of DNAs (repDNA) for generating the widely used features reflecting the physicochemical properties and sequence-order effects of DNAs and nucleotides. There are three feature groups composed of 15 features. The first group calculates three nucleic acid composition features describing the local sequence information by means of kmers; the second group calculates six autocorrelation features describing the level of correlation between two oligonucleotides along a DNA sequence in terms of their specific physicochemical properties; the third group calculates six pseudo nucleotide composition features, which can be used to represent a DNA sequence with a discrete model or vector yet still keep considerable sequence-order information via the physicochemical properties of its constituent oligonucleotides. In addition, these features can be easily calculated based on both the built-in and user-defined properties via using repDNA. AVAILABILITY AND IMPLEMENTATION: The repDNA Python package is freely accessible to the public at http://bioinformatics.hitsz.edu.cn/repDNA/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Fule Liu, Longyun Fang, Xiaolong Wang 0001, Kuo-Chen Chou |
Bioinform. | 4 |
| 2015 | An automatic system to identify heart disease risk factors in clinical texts over timeabstractDespite recent progress in prediction and prevention, heart disease remains a leading cause of death. One preliminary step in heart disease prediction and prevention is risk factor identification. Many studies have been proposed to identify risk factors associated with heart disease; however, none have attempted to identify all risk factors. In 2014, the National Center of Informatics for Integrating Biology and Beside (i2b2) issued a clinical natural language processing (NLP) challenge that involved a track (track 2) for identifying heart disease risk factors in clinical texts over time. This track aimed to identify medically relevant information related to heart disease risk and track the progression over sets of longitudinal patient medical records. Identification of tags and attributes associated with disease presence and progression, risk factors, and medications in patient medical history were required. Our participation led to development of a hybrid pipeline system based on both machine learning-based and rule-based approaches. Evaluation using the challenge corpus revealed that our system achieved an F1-score of 92.68%, making it the top-ranked system (without additional annotations) of the 2014 i2b2 clinical NLP challenge. Qingcai Chen, Haodi Li, Buzhou Tang, Xiaolong Wang 0001, Xin Liu 0054, Zengjian Liu, Weida Wang, Qiwen Deng, Suisong Zhu, Yangxin Chen |
J. Biomed. Informatics | 4 |
| 2015 | Automatic de-identification of electronic medical records using token-level and character-level conditional random fieldsabstractDe-identification, identifying and removing all protected health information (PHI) present in clinical data including electronic medical records (EMRs), is a critical step in making clinical data publicly available. The 2014 i2b2 (Center of Informatics for Integrating Biology and Bedside) clinical natural language processing (NLP) challenge sets up a track for de-identification (track 1). In this study, we propose a hybrid system based on both machine learning and rule approaches for the de-identification track. In our system, PHI instances are first identified by two (token-level and character-level) conditional random fields (CRFs) and a rule-based classifier, and then are merged by some rules. Experiments conducted on the i2b2 corpus show that our system submitted for the challenge achieves the highest micro F-scores of 94.64%, 91.24% and 91.63% under the "token", "strict" and "relaxed" criteria respectively, which is among top-ranked systems of the 2014 i2b2 challenge. After integrating some refined localization dictionaries, our system is further improved with F-scores of 94.83%, 91.57% and 91.95% under the "token", "strict" and "relaxed" criteria respectively. Zengjian Liu, Yangxin Chen, Buzhou Tang, Xiaolong Wang 0001, Qingcai Chen, Haodi Li, Qiwen Deng, Suisong Zhu |
J. Biomed. Informatics | 4 |
| 2015 | Predicting the quality of user-generated answers using co-training in community-based question answering portals
Bingquan Liu, Ming Liu 0004, Haifeng Hu 0002, Xiaolong Wang 0001 |
Pattern Recognit. Lett. | 5 |
| 2015 | APP Relationship Calculation: An Iterative ProcessabstractToday, plenty of apps are released to enable users to make the best use of their cell phones. Facing the large amount of apps, app retrieval and app recommendation become important, since users can easily use them to acquire their desired apps. To obtain high-quality retrieval and recommending results, it needs to obtain the precise app relationship calculating results. Unfortunately, the recent methods are conducted mostly relying on user's log or app's description, which can only detect whether two apps are downloaded, installed meanwhile or provide similar functions or not. In fact, apps contain many general relationships other than similarity, such as one app needs another app as its tool. These relationships cannot be dug via user's log or app's description. Reviews contain user's viewpoint and judgment to apps, thus they can be used to calculate relationship between apps. To use reviews, this paper proposes an iterative process by combining review similarity and app relationship together. Experimental results demonstrate that via this iterative process, relationship between apps can be calculated exactly. Furthermore, this process is improved in two aspects. One is to obtain excellent results even with weak initialization. The other is to apply matrix product to reduce running time. Ming Liu 0004, Chong Wu 0001, Xiang-Nan Zhao, Chin-Yew Lin, Xiaolong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Hybrid Deep Belief Networks for Semi-supervised Sentiment Classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
COLING | 3 |
| 2014 | CTR Prediction for DSP with Improved Cube Factorization Model from Historical Bidding Log
Lili Shan, Lei Lin 0001, Di Shao, Xiaolong Wang 0001 |
ICONIP (3) | 4 |
| 2014 | Radical-Enhanced Chinese Character Embedding
Yaming Sun, Lei Lin 0001, Nan Yang 0002, Zhenzhou Ji, Xiaolong Wang 0001 |
ICONIP (2) | 5 |
| 2014 | Computing Semantic Relatedness Using a Word-Text Mutual Guidance Model
Bingquan Liu, Ming Liu 0004, Feng Liu 0041, Xiaolong Wang 0001 |
NLPCC | 5 |
| 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detectionabstractMOTIVATION: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. RESULTS: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/~binliu/remote/ Bin Liu 0014, Deyuan Zhang, Ruifeng Xu 0001, Jinghao Xu, Xiaolong Wang 0001, Qingcai Chen, Qiwen Dong, Kuo-Chen Chou |
Bioinform. | 5 |
| 2014 | Using distances between Top-n-gram and residue pairs for protein remote homology detectionabstractBACKGROUND: Protein remote homology detection is one of the central problems in bioinformatics, which is important for both basic research and practical application. Currently, discriminative methods based on Support Vector Machines (SVMs) achieve the state-of-the-art performance. Exploring feature vectors incorporating the position information of amino acids or other protein building blocks is a key step to improve the performance of the SVM-based methods. RESULTS: Two new methods for protein remote homology detection were proposed, called SVM-DR and SVM-DT. SVM-DR is a sequence-based method, in which the feature vector representation for protein is based on the distances between residue pairs. SVM-DT is a profile-based method, which considers the distances between Top-n-gram pairs. Top-n-gram can be viewed as a profile-based building block of proteins, which is calculated from the frequency profiles. These two methods are position dependent approaches incorporating the sequence-order information of protein sequences. Various experiments were conducted on a benchmark dataset containing 54 families and 23 superfamilies. Experimental results showed that these two new methods are very promising. Compared with the position independent methods, the performance improvement is obvious. Furthermore, the proposed methods can also provide useful insights for studying the features of protein families. CONCLUSION: The better performance of the proposed methods demonstrates that the position dependant approaches are efficient for protein remote homology detection. Another advantage of our methods arises from the explicit feature space representation, which can be used to analyze the characteristic features of protein families. The source code of SVM-DT and SVM-DR is available at http://bioinformatics.hitsz.edu.cn/DistanceSVM/index.jsp. Bin Liu 0014, Jinghao Xu, Quan Zou 0001, Ruifeng Xu 0001, Xiaolong Wang 0001, Qingcai Chen |
BMC Bioinform. | 5 |
| 2014 | Fuzzy deep belief networks for semi-supervised sentiment classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
Neurocomputing | 3 |
| 2014 | Handwritten Chinese text editing and recognition system
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
Multim. Tools Appl. | 3 |
| 2013 | Chinese Emotion Lexicon Developing via Multi-lingual Lexical Resources Integration
Jun Xu 0007, Ruifeng Xu 0001, Yanzhen Zheng, Qin Lu 0001, Kam-Fai Wong, Xiaolong Wang 0001 |
CICLing (2) | 6 |
| 2013 | Deep Learning Approaches for Link Prediction in Social Network Services
Feng Liu 0041, Bingquan Liu, Chengjie Sun, Ming Liu 0004, Xiaolong Wang 0001 |
ICONIP (2) | 5 |
| 2013 | Automatic Corpora Construction for Text Classification
Qingcai Chen, Xiaolong Wang 0001, Bingyang Yu |
IJCNLP | 3 |
| 2013 | Grammatical Error Correction Using Feature Selection and Confidence Tuning
Yang Xiang 0003, Yaoyun Zhang, Xiaolong Wang 0001, Chongqiang Wei, Xiaoqiang Zhou, Yuxiu Hu, Yang Qin 0001 |
IJCNLP | 3 |
| 2013 | Active deep learning method for semi-supervised sentiment classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
Neurocomputing | 3 |
| 2013 | A support vector machine based MSM model for financial short-term volatility forecasting
Hejiao Huang, Xiaolong Wang 0001 |
Neural Comput. Appl. | 3 |
| 2013 | Convolutional Deep Networks for Visual Data Classification
Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
Neural Process. Lett. | 3 |
| 2012 | Coarse-to-fine sentence-level emotion classification based on the intra-sentence features and sentential contextabstractThis paper proposes a novel approach using a coarse-to-fine analysis strategy for sentence-level emotion classification which takes into consideration of similarities to sentences in training set as well as adjacent sentences in the context. First, we use intra-sentence based features to determine the emotion label set of a target sentence coarsely through the statistical information gained from the label sets of the k most similar sentences in the training data. Then, we use the emotion transfer probabilities between neighboring sentences to refine the emotion labels of the target sentences. Such iterative refinements terminate when the emotion classification converges. The proposed algorithm is evaluated on Ren-CECps, a Chinese blog emotion corpus. Experimental results show that the coarse-to-fine emotion classification algorithm improves the sentence-level emotion classification by 19.11% on the average precision metric, which outperforms the baseline methods. Jun Xu 0007, Ruifeng Xu 0001, Qin Lu 0001, Xiaolong Wang 0001 |
CIKM | 4 |
| 2012 | An Empirical Evaluation on Online Chinese Handwriting DatabasesabstractSeveral online Chinese handwriting databases have been proposed recently. Though they have been introduced in detail, to date, no one has ever evaluated these databases with experimental comparison. To help the researchers use the corresponding database properly for algorithm evaluation and real application, we compare the property of the handwriting characters in these databases, and evaluate them with the same experimental setup and handwriting recognizer. Moreover, we analyze the connection between the property and the corresponding recognition accuracy for the handwriting characters in different databases. These empirical evaluation results can help the researchers choose the right database for different algorithms and applications. Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001, Zou Chen, Suqin Ao |
Document Analysis Systems | 3 |
| 2012 | Features for link prediction in social networks: A comprehensive studyabstractWith the development of social media websites, more and more users start to show their attitudes and emotions to each other. Some of these interactions can be represented as links with sign values(positive or negative). In this paper, a unified method is proposed for link prediction and feature analysis. This paper focuses on the data from social media websites and tries to find the features that determine the sign value mostly. Based on the features extracted from the users' self statuses and from their relationships with neighbors, our method can predict the links' values with high accuracy. By analyzing the models generated over different datasets, our experiments find out the common determining features for link prediction. Based on our results, advices on how to predict links' values and get more positive links in future are given to users. Feng Liu 0041, Bingquan Liu, Xiaolong Wang 0001, Ming Liu 0004, Baoxun Wang |
SMC | 3 |
| 2012 | A novel text mining approach to financial time series forecasting
Hejiao Huang, Xiaolong Wang 0001 |
Neurocomputing | 3 |
| 2011 | Partially Supervised Text Classification with Multi-Level ExamplesabstractPartially supervised text classification has received great research attention since it only uses positive and unlabeled examples as training data. This problem can be solved by automatically labeling some negative (and more positive) examples from unlabeled examples before training a text classifier. But it is difficult to guarantee both high quality and quantity of the new labeled examples. In this paper, a multi-level example based learning method for partially supervised text classification is proposed, which can make full use of all unlabeled examples. A heuristic method is proposed to assign possible labels to unlabeled examples and partition them into multiple levels according to their labeling confidence. A text classifier is trained on these multi-level examples using weighted support vector machines. Experiments show that the multi-level example based learning method is effective for partially supervised text classification, and outperforms the existing popular methods such as Biased-SVM, ROC-SVM, S-EM and WL. Tao Liu 0001, Xiaoyong Du 0001, Yong-Dong Xu, Xiaolong Wang 0001 |
AAAI | 5 |
| 2011 | An Empirical Evaluation on HIT-OR3C DatabaseabstractRecently, we have proposed a handwriting Chinese character database HIT-OR3C. Though it has been introduced in detail, to date, it has not been evaluated by any handwriting recognition method. To help the researchers use this database for algorithm evaluation, we propose the structure of HIT-OR3C database. Moreover, we evaluate the OR3C database with a series of experiments using state-of-the-art handwriting recognizer. These experiment results on the different subsets can be a benchmark for the researchers who will use the database. The low average recognition rate confirms that the HIT-OR3C database is challenging. Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
ICDAR | 3 |
| 2011 | Using Hybrid Kernel Method for Question Classification in CQA
Shixi Fan, Xiaolong Wang 0001, Xuan Wang 0002, Xiaohong Yang |
ICONIP (3) | 2 |
| 2011 | Macro Features Based Text Categorization
Qingcai Chen, Xiaolong Wang 0001, Buzhou Tang |
ICONIP (2) | 3 |
| 2011 | Dynamic Template Based Online Event Detection
Qingcai Chen, Xiaolong Wang 0001, Jiacai Weng |
ICONIP (3) | 3 |
| 2011 | Complex Detection Based on Integrated Properties
Lei Lin 0001, Chengjie Sun, Xiaolong Wang 0001, Xuan Wang 0002 |
ICONIP (1) | 4 |
| 2011 | Making Image to Class Distance Comparable
Deyuan Zhang, Bingquan Liu, Chengjie Sun, Xiaolong Wang 0001 |
ICONIP (2) | 4 |
| 2011 | Diversifying Question Recommendations in Community-Based Question Answering
Yaoyun Zhang, Xiaolong Wang 0001, Xuan Wang 0002, Ruifeng Xu 0001, Buzhou Tang |
ICONIP (3) | 2 |
| 2011 | Diversifying Information Needs in Results of Question Retrieval
Yaoyun Zhang, Xiaolong Wang 0001, Xuan Wang 0002, Ruifeng Xu 0001, Jun Xu 0007, Shixi Fan |
IJCNLP | 2 |
| 2011 | Deep Belief Networks for Automatic Music Genre Classification
Xiaohong Yang, Qingcai Chen, Shusen Zhou, Xiaolong Wang 0001 |
INTERSPEECH | 4 |
| 2011 | A language model approach for tag recommendation
Ke Sun 0007, Xiaolong Wang 0001, Chengjie Sun, Lei Lin 0001 |
Expert Syst. Appl. | 2 |
| 2011 | Tolerance Rough Set Based Attribute Extraction Approach for Multiple Semantic Knowledge Base IntegrationabstractIn the integration of multiple semantic knowledge bases (SKBs), the inconsistence of the items or their attributes appeared in different SKBs is still an opening challenge for researchers. To address this issue, this paper presents an innovative approach which bases on extracting common class attributes and establishing unified category-attribute templates. Since the natural properties of uncertainty and vagueness of semantic analysis involved in selecting a specific attribute from numerous candidates, the tolerance rough set (TRS) techniques are applied in constructing class-attribute templates from online SKBs. The extraction of attribute is fulfilled by statistical techniques and is integrated into the TRS framework. Finally, experiments are conducted on random selected categories. Experimental results show the effectiveness of the proposed approach. Hongzhi Guo 0007, Qingcai Chen, Xiaolong Wang 0001 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 3 |
| 2011 | Analysis of the Petri net model of parallel manufacturing processes with shared resources
Farooq Ahmad, Hejiao Huang, Xiaolong Wang 0001 |
Inf. Sci. | 3 |
| 2011 | Deep Learning Approaches to Semantic Relevance Modeling for Chinese Question-Answer PairsabstractThe human-generated question-answer pairs in the Web social communities are of great value for the research of automatic question-answering technique. Due to the large amount of noise information involved in such corpora, it is still a problem to detect the answers even though the questions are exactly located. Quantifying the semantic relevance between questions and their candidate answers is essential to answer detection in social media corpora. Since both the questions and their answers usually contain a small number of sentences, the relevance modeling methods have to overcome the problem of word feature sparsity. In this article, the deep learning principle is introduced to address the semantic relevance modeling task. Two deep belief networks with different architectures are proposed by us to model the semantic relevance for the question-answer pairs. According to the investigation of the textual similarity between the community-driven question-answering (cQA) dataset and the forum dataset, a learning strategy is adopted to promote our models’ performance on the social community corpora without hand-annotating work. The experimental results show that our method outperforms the traditional approaches on both the cQA and the forum corpora. Baoxun Wang, Bingquan Liu, Xiaolong Wang 0001, Chengjie Sun, Deyuan Zhang |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2010 | Modeling Semantic Relevance for Question-Answer Pairs in Web Social Communities
Baoxun Wang, Xiaolong Wang 0001, Chengjie Sun, Bingquan Liu, Lin Sun 0010 |
ACL | 2 |
| 2010 | HIT-OR3C: an opening recognition corpus for Chinese charactersabstractThis paper proposes an opening recognition corpus, HIT-OR3C, and its construction toolkit to facilitate the unconstrained online Chinese handwriting text recognition. The characters of HIT-OR3C are collected through handwriting pad and are recorded and labeled automatically via the proposed handwriting document collection software OR3C Toolkit. HIT-OR3C consists of 5 subsets, namely GB1, GB2, Letter, Digit and Document. The first 4 corpora contain 6,825 categories produced by 122 persons and 832,650 samples in total. The document corpus is corresponding to 10 news articles that contain 2,442 categories produced by 20 persons and 77,168 samples in total. HIT-OR3C can be used for training and evaluation of character recognition algorithms. The OR3C Toolkit provides an efficient, device-independent, and unconstrained platform for the building of large scale handwriting corpus. Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
Document Analysis Systems | 3 |
| 2010 | Discriminative Deep Belief Networks for image classificationabstractThis paper presents a novel semi-supervised learning algorithm called Discriminative Deep Belief Networks (DDBN), to address the image classification problem with limited labeled data. We first construct a new deep architecture for classification using a set of Restricted Boltzmann Machines (RBM). The parameter space of the deep architecture is initially determined using labeled data together with abundant of unlabeled data, by greedy layer-wise unsupervised learning. Then, we fine-tune the whole deep networks using an exponential loss function to maximize the separability of the labeled data, by gradient-descent based supervised learning. Experiments on the artificial dataset and real image datasets show that DDBN outperforms most semi-supervised algorithm and deep learning techniques, especially for the hard classification tasks. Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
ICIP | 3 |
| 2010 | Reranking for Stacking Ensemble Learning
Buzhou Tang, Qingcai Chen, Xuan Wang 0002, Xiaolong Wang 0001 |
ICONIP (1) | 4 |
| 2010 | Learning the Kernel Combination for Object CategorizationabstractAlthough Support Vector Machines(SVM) succeed in classifying several image databases using image descriptors proposed in the literature, no single descriptor can be optimal for general object categorization. This paper describes a novel framework to learn the optimal combination of kernels corresponding to multiple image descriptors before SVM training, leading to solve a quadratic programming problem efficiently. Our framework takes into account the variation of kernel matrix and imbalanced dataset, which are common in real world image categorization tasks. Experimental results on Graz-01 and Caltech-101 image databases show the effectiveness and robustness of our algorithm. Deyuan Zhang, Xiaolong Wang 0001, Bingquan Liu |
ICPR | 2 |
| 2010 | Deep Quantum Networks for ClassificationabstractThis paper introduces a new type of deep learning method named Deep Quantum Network (DQN) for classification. DQN inherits the capability of modeling the structure of a feature space by fuzzy sets. At first, we propose the architecture of DQN, which consists of quantum neuron and sigmoid neuron and can guide the embedding of samples divisible in new Euclidean space. The parameter of DQN is initialized through greedy layer-wise unsupervised learning. Then, the parameter space of the deep architecture and quantum representation are refined by supervised learning based on the global gradient-descent procedure. An exponential loss function is introduced in this paper to guide the supervised learning procedure. Experiments conducted on standard datasets show that DQN outperforms other feed forward neural networks and neuro-fuzzy classifiers. Shusen Zhou, Qingcai Chen, Xiaolong Wang 0001 |
ICPR | 3 |
| 2010 | Petri net modeling and deadlock analysis of parallel manufacturing processes with shared-resources
Farooq Ahmad, Hejiao Huang, Xiaolong Wang 0001 |
J. Syst. Softw. | 3 |
| 2009 | Protein Long Disordered Region Prediction Based on Profile-Level Disorder Propensities and Position-Specific Scoring MatrixesabstractIdentification of long disordered regions in protein sequence is important for understanding protein function. In this work, a class of novel propensities at profile level is presented, namely, the order profile disorder propensities, which use the evolutionary information of profile for protein long disorder prediction. These propensities, combined with position-specific scoring matrices, are inputted to the logistic regression (LR) for the prediction of protein long disordered regions. In 5-fold cross-validation test, our method can achieve an area of 97.5% under the ROC cure. Testing on a blind-test set, our method is significantly more accurate than several existing disorder predictors. Bin Liu 0014, Lei Lin 0001, Xiaolong Wang 0001, Xuan Wang 0002 |
BIBM | 3 |
| 2009 | Learning to recommend questions based on user ratingsabstractAt community question answering services, users are usually encouraged to rate questions by votes. The questions with the most votes are then recommended and ranked on the top when users browse questions by category. As users are not obligated to rate questions, usually only a small proportion of questions eventually gets rating. Thus, in this paper, we are concerned with learning to recommend questions from user ratings of a limited size. To overcome the data sparsity, we propose to utilize questions without users rating as well. Further, as there exist certain noises within user ratings (the preference of some users expressed in their ratings diverges from that of the majority of users), we design a new algorithm called 'majority-based perceptron algorithm' which can avoid the influence of noisy instances by emphasizing its learning over data instances from the majority users. Experimental results from a large collection of real questions confirm the effectiveness of our proposals. Ke Sun 0007, Yunbo Cao, Xinying Song, Young-In Song, Xiaolong Wang 0001, Chin-Yew Lin |
CIKM | 5 |
| 2009 | STRank: A SiteRank Algorithm Using Semantic Relevance and Time FrequencyabstractMost of the researches on Web information processing are concentrated on the Web pages and the hyperlinks among them. One of the important facts that a Web page is just one building block of the whole Website had been ignored. But the situation is gradually changed in recent years for the needs of Website reputation calculation, the high level Website structure mining etc. It causes the Website ranking become one of the hot research topics and various site ranking algorithms, such as SiteRank, AggregateRank etc., had been proposed. But most of existing Website ranking algorithm just take use of Website link graphs and the content of Websites are usually not put into consideration. It is obviously not enough for a reliable ranking of Websites. To address this issue, this paper introduces two content based features, i.e., semantic relevance and time frequency and proposes a new STRank algorithm based on these two features. We firstly conduct a series of experiments to verify the feasibility of these two factors in site ranking task. Then the semantic relevance is applied in the calculation of transition probability, and the updating frequency of sites is combined into the ranking task. Since traditional Kendall's ¿ distance and Spearman's footrule distance is not appropriate for the evaluation of site ranking, we make some modifications accordingly to evaluate Website ranking algorithms. Finally, our experiments show that the STRank algorithm outperforms existing approaches on both effectiveness and efficiency. Hongzhi Guo 0007, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001 |
SMC | 3 |
| 2009 | Text Clustering Approach Based on Maximal Frequent Term SetsabstractClassical text clustering algorithms are usually based on vector space model or its variants. Because of the high computing complexity and the difficulty of controlling clustering results, this kind of approaches are hard to be applied for the purpose of the large scale text clustering. Clustering algorithms based on frequent term sets make use of relationship among documents and their shared frequent term sets to achieve high accuracy and effectiveness in clustering. But since the number of frequent terms is usually too large to reach the efficiency requirement for large collection texts clustering, this paper proposes a novel text clustering approach based on maximal frequent term sets (MFTSC). This approach firstly mines maximal frequent term sets from text set and then clusters texts by following steps: at first, the maximal frequent term sets are clustered based on the criterion of k-mismatch; then texts are clustered according to term sets clustering results; finally, we categorize the left texts uncovered in previous step into produced text clusters Be compared with existing approaches, our experimental results show an average gain of 10% on F-Measure score with better performance on scalability and efficiency. Chong Su, Qingcai Chen, Xiaolong Wang 0001, Xianjun Meng |
SMC | 3 |
| 2009 | Study on Feature Selection in Finance Text CategorizationabstractDocument genre information is one of the most distinguishing features in information retrieval, which brings order to the search results. What the genre classification concerned is not the topic but the genre of document. In this paper, two different feature sets were employed: bag of words which are derived by feature selection method and structural features which are selected manually and subjectively. And a comparative study on feature selection in genre classification of Chinese finance text is presented. In empirical results with classifiers on the real world corpora, we find that those manual labeled features can improve the performance clearly. Changqiu Sun, Xiaolong Wang 0001, Jun Xu 0007 |
SMC | 2 |
| 2009 | Extracting Chinese Question-Answer Pairs from Online ForumsabstractExtracting question-answer pairs from online forums is a meaningful work due to the huge amount of valuable user generated resource contained in forums. In this paper we consider the problem of extracting Chinese question-answer pairs for the first time. We present a strategy to detect Chinese questions and their answers. We propose a sequential rule based method to find questions in a forum thread, then we adopt non-textual features based on forum structure to improve the performance of answer detecting in the same thread. Experimental results show that our techniques are very effective. Baoxun Wang, Bingquan Liu, Chengjie Sun, Xiaolong Wang 0001, Lin Sun 0010 |
SMC | 4 |
| 2009 | CRF-based Active Learning for Chinese Named Entity RecognitionabstractConditional Random Fields (CRFs) have been used for many sequence labeling tasks and got excellent results. Further, the supervised model strongly depends on the huge training data. Active learning is a different way rather than relying on a large amount random sampling. However, random sampling constructively participates in the optimal choosing training examples. Based on different query strategies, active learning can combine with other machine learning methods to reduce the annotation cost while maintaining the accuracy. This paper proposes a new active learning strategy based on Information Density (ID) integrated with CRFs for Chinese Named Entity Recognition (NER). On Sighan bakeoff 2006 MSRA NER corpus, an F1 score of 77.2% is achieved by using only 10,000 labeled training sentences chosen by the proposed active learning strategy. Lin Yao 0004, Chengjie Sun, Xiaolong Wang 0001, Xuan Wang 0002 |
SMC | 4 |
| 2009 | Using Question Classification to Model User Intentions of Different LevelsabstractUser information need detection is a fundamental issue in automatic question answering systems. Based on real questions collected from on-line question answering communities, this paper proposes a three-level question type taxonomy to model user information need. The three levels are based on interrogative patterns, hidden user intentions and specific answer expectations. One question can have multiple types in level 2&3. Question type assignment of level 2&3 is subjective-orientated, and may vary between different users. Shallow lexical, syntactic and semantic features are used to model the inherent subjectivity of user intentions. Classification experiments are conducted on a corpus of real questions collected from the web. Different machine learning methods are employed. Experimental results are promising. This indicates the capability of modeling user information need and subjectivity statistically, and that strong correlations exist between question types of the same level. Yaoyun Zhang, Xuan Wang 0002, Xiaolong Wang 0001, Shixi Fan, Daoxu Zhang |
SMC | 3 |
| 2009 | Prediction of protein binding sites in protein structures using hidden Markov support vector machineabstractBACKGROUND: Predicting the binding sites between two interacting proteins provides important clues to the function of a protein. Recent research on protein binding site prediction has been mainly based on widely known machine learning techniques, such as artificial neural networks, support vector machines, conditional random field, etc. However, the prediction performance is still too low to be used in practice. It is necessary to explore new algorithms, theories and features to further improve the performance. RESULTS: In this study, we introduce a novel machine learning model hidden Markov support vector machine for protein binding site prediction. The model treats the protein binding site prediction as a sequential labelling task based on the maximum margin criterion. Common features derived from protein sequences and structures, including protein sequence profile and residue accessible surface area, are used to train hidden Markov support vector machine. When tested on six data sets, the method based on hidden Markov support vector machine shows better performance than some state-of-the-art methods, including artificial neural networks, support vector machines and conditional random field. Furthermore, its running time is several orders of magnitude shorter than that of the compared methods. CONCLUSION: The improved prediction performance and computational efficiency of the method based on hidden Markov support vector machine can be attributed to the following three factors. Firstly, the relation between labels of neighbouring residues is useful for protein binding site prediction. Secondly, the kernel trick is very advantageous to this field. Thirdly, the complexity of the training step for hidden Markov support vector machine is linear with the number of training samples by using the cutting-plane algorithm. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Buzhou Tang, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 2 |
| 2009 | Channel assignment using block design in wireless mesh networks
Hejiao Huang, Xiaolu Cao, Xiaohua Jia, Xiaolong Wang 0001 |
Comput. Commun. | 4 |
| 2008 | Discriminative Learning of Syntactic and Semantic Dependencies
Shixi Fan, Xuan Wang 0002, Xiaolong Wang 0001 |
CoNLL | 4 |
| 2008 | A Study of Chinese Lexical Analysis Based on Discriminative Models
Guang-Lu Sun, Chengjie Sun, Ke Sun 0007, Xiaolong Wang 0001 |
IJCNLP | 4 |
| 2008 | Name Origin Recognition Using Maximum Entropy Model and Diverse Features
Min Zhang 0005, Chengjie Sun, Haizhou Li 0001, AiTi Aw, Chew Lim Tan, Xiaolong Wang 0001 |
IJCNLP | 6 |
| 2008 | Chunking with Max-Margin Markov Networks
Buzhou Tang, Xuan Wang 0002, Xiaolong Wang 0001 |
PACLIC | 3 |
| 2008 | A technique for generating the reduced reachability graph of petri net modelsabstractThe reachability graph technology is the basic and important method of analysis and verification of a system but it suffers from state explosion problem. To cope with this problem, this paper introduces transition vectors which can detect all the enabled transitions at given state of system and identify them as dependent or independent. Transition vectors have been found useful and functional for simultaneous execution of concurrently enabled transitions. An efficient algorithm based on transition vectors has been presented to generate the reduced reachability graph and compared with the reduced reachability graph constructed by stubborn set method. Farooq Ahmad, Hejiao Huang, Xiaolong Wang 0001, Waqas Anwer |
SMC | 3 |
| 2008 | Semantic Chunk Annotation for questions using Maximum Entropyabstractwe present a ME (Maximum Entropy) model for Semantic Chunk Annotation in a Chinese Question and Answer (Q&A) system. The model was derived from a corpus of real world questions, which are collected from some discussion groups on the Internet. The questions are supposed to be answered by other people, so the questions are very complex. The semantic chunks were introduced. Feature for the model was described and MI (Mutual Information) was adopted for feature selection. The training data consists of 14000 sentences and the test data consists of 4000 sentences. The result: F-score is 90.68%. Shixi Fan, Yaoyun Zhang, Wing W. Y. Ng, Xuan Wang 0002, Xiaolong Wang 0001 |
SMC | 5 |
| 2008 | Topology simplification and channel assignment in multi-radio Wireless Mesh NetworksabstractAs the price of hardware decreases, Wireless Mesh Networks (WMNs) will become practical commodity in a few years. By equipping each node with multiple network interface cards, multiple channels can be applied to improve the performance of WMNs. With the advantages of both wireless LAN and Ad Hoc, WMN also has challenges. Connectivity of the communication network and the interferences between the links of the network are two most important points to be considered carefully. This paper presents some efficient technologies to handle the two points. It firstly simplifies the topology of the communication network and gets a bi-connected network; then for the resulted network, it greedily assigns channels such that the interference of the network is as small as possible. The simulation results show that, by applying the technologies presented in this paper, a bi-connected network with minimum interference is created. Hejiao Huang, Xiaolu Cao, Xiaolong Wang 0001 |
SMC | 3 |
| 2008 | Semantic feature reduction in chinese document clusteringabstractText clustering techniques were usually used to structure the text documents into topic related groups which can facilitate users to get a comprehensive understanding on corpus or results from information retrieval system. Most of existing text clustering algorithm which derived from traditional formatted data clustering heavily rely on term analysis methods and adopted vector space model (VSM) as their document representation. But because of the essential characteristic underlying text such as high dimensionality features vector space, the problem of sparseness has a strong impact on the clustering algorithm. So feature reduction is an important preprocess step for improving the efficiency and accuracy of clustering algorithm by removing redundant and irrelevant terms from corpus. Even the clustering is considered as an unsupervised learning method, but in text, there is still some priori knowledge we can use from NLP analysis based approach. In this paper, we propose a semantic analysis based feature reduction method which used in Chinese text clustering. Our method bases on a dedicated Part-of-Speech tags selection and synonyms consolidation and can reduce the feature space of documents more effectively compared with traditional feature reduction method tfidf and stopwords removal; meanwhile it preserves or sometimes even improves the accuracy of clustering algorithm. In our experiment, we tested our feature reduction method using bisecting k-means algorithm which was proved be efficient in text clustering. The results show that our method can reduce the feature space significantly, and meanwhile have a better clustering accuracy in terms of the purity. Xianjun Meng, Qingcai Chen, Xiaolong Wang 0001 |
SMC | 3 |
| 2008 | Basic semantic units based web page content extractionabstractWeb page content extraction can be achieved by node-based and segmentation-based algorithms respectively on top of the document object model (DOM). However, the node-based algorithm often removes content embedded as anchor text; while the segmentation-based way can not distinguish irrelevant text from content text when they are divided into the same segment. The two kinds of algorithms don't keep the paragraph information of the original page either. In this paper, a new basic semantic unit (BSU) with granularity between nodes in the DOM tree and content block is defined. Two different methods based on BSU, using clustering and heuristic rules are developed to extract page content. The clustering method gets the best precision 96.88%; while the heuristic rules obtain the best F1-value 95.28%. Compared with the baseline method which uses text blocks segmented byandas Web page content, the F1-values are enhanced by 8.92% and 9.42% respectively. Qingcai Chen, Xiaolong Wang 0001, Hongzhi Guo 0007 |
SMC | 3 |
| 2008 | Genre identification of Chinese finance text using machine learning methodabstractDocument genre information is one of the most distinguishing features in information retrieval, which brings order to the search results. What the genre classification concerned is not the topic but the genre of document. In this paper, we examine the effectiveness of using machine learning techniques to solve genre classification of Chinese text with the same topic, viz. finance. Based on the likelihood ratio test, we present a new method for selecting feature terms, which can improve the performance clearly and perform better than others with up to 80% terms removal. In empirical results with SVMs classifier on the real world corpora, we find that this method can gain a better selecting effect and likelihood ratio is a reliable measure for selecting informative features. Jun Xu 0007, Xiaolong Wang 0001, Yonghui Wu 0001 |
SMC | 3 |
| 2008 | A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysisabstractBACKGROUND: Protein remote homology detection and fold recognition are central problems in bioinformatics. Currently, discriminative methods based on support vector machine (SVM) are the most effective and accurate methods for solving these problems. A key step to improve the performance of the SVM-based methods is to find a suitable representation of protein sequences. RESULTS: In this paper, a novel building block of proteins called Top-n-grams is presented, which contains the evolutionary information extracted from the protein sequence frequency profiles. The protein sequence frequency profiles are calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into Top-n-grams. The protein sequences are transformed into fixed-dimension feature vectors by the occurrence times of each Top-n-gram. The training vectors are evaluated by SVM to train classifiers which are then used to classify the test protein sequences. We demonstrate that the prediction performance of remote homology detection and fold recognition can be improved by combining Top-n-grams and latent semantic analysis (LSA), which is an efficient feature extraction technique from natural language processing. When tested on superfamily and fold benchmarks, the method combining Top-n-grams and LSA gives significantly better results compared to related methods. CONCLUSION: The method based on Top-n-grams significantly outperforms the methods based on many other building blocks including N-grams, patterns, motifs and binary profiles. Therefore, Top-n-gram is a good building block of the protein sequences and can be widely used in many tasks of the computational biology, such as the sequence alignment, the prediction of domain boundary, the designation of knowledge-based potentials and the prediction of protein binding sites. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 2 |
| 2008 | ConSOM: A conceptional self-organizing map model for text clustering
Yuanchao Liu, Xiaolong Wang 0001, Chong Wu 0001 |
Neurocomputing | 2 |
| 2008 | Auto Adapted English Pronunciation Evaluation: a Fuzzy Integral ApproachabstractTo evaluate the pronunciation skills of spoken English is one of the key tasks for computer-aided spoken language learning (CALL). While most of the researchers focus on improving the speech recognition techniques to build a reliable evaluation system, another important aspect of this task has been ignored, i.e. the pronunciation evaluation model that integrates both the reliabilities of existing speech processing systems and the learner's pronunciation personalities. To take this aspect into consideration, a Sugeno integral-based evaluation model is introduced in this paper. At first, the English phonemes that are hard to be distinguished (HDP) for Chinese language learners are grouped into different HDP sets. Then, the system reliabilities for distinguishing the phonemes within a HDP set are computed from the standard speech corpus and are integrated with the phoneme recognition results under the Sugeno integral framework. The fuzzy measures are given for each subset of speech segments that contains n occurrences of phonemes within a HDP set. Rather than providing a quantity of scores, the linguistic descriptions of evaluation results are given by the model, which is more helpful for the users to improve their spoken language skills. To get a better performance, generic algorithm (GA)-based parameter optimization is also applied to optimize the model parameters. Experiments are conducted on the Sphinx-4 speech recognition platform. They show that, with 84.7% of average recognition rate of the SR system on standard speech corpus, our pronunciation evaluation model has got reasonable and reliable results for three kinds of test corpora. Qingcai Chen, Xiaolong Wang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2008 | A New Measurement of Systematic SimilarityabstractThe relationship of similarity may be the most universal relationship that exists between every two objects in either the material world or the mental world. Although similarity modeling has been the focus of cognitive science for decades, many theoretical and realistic issues are still under controversy. In this paper, a new theoretical framework that conforms to the nature of similarity and incorporates the current similarity models into a universal model is presented. The new model, i.e., the systematic similarity model, which is inspired by the contrast model of similarity and structure mapping theory in cognitive psychology, is the universal similarity measurement that has many potential applications in text, image, or video retrieval. The text relevance ranking experiments undertaken in this research tentatively show the validity of the new model. Yi Guan, Xiaolong Wang 0001, Qiang Wang 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2007 | Using Maximum Entropy Model to Extract Protein-Protein Interaction Information from Biomedical Literature
Chengjie Sun, Lei Lin 0001, Xiaolong Wang 0001, Yi Guan |
ICIC (1) | 3 |
| 2007 | Model fusion of conditional random fieldsabstractThis paper introduces two model fusion methods on a series of sub-models of Conditional Random Fields (CRFs): majority voting and feature fusion. The former performs on the results of each participant without any consideration about the underlying details of each sub-model, and the latter takes place on feature level to produce modified feature weights of CRFs to merge all sub-models into a single one. Experiments on syntactic data and part-of-speech tagging problem shows that by dividing training corpus into small parts and using model fusion techniques, comparable results will be achieved. Xuan Wang 0002, Yanbing Yu, Xiaolong Wang 0001 |
SMC | 4 |
| 2007 | Extracting domain-specific terms from unlabeled web documents by bootstrapping and term classifiersabstractDomain-specific term extraction contributes to all domain-oriented natural language processing tasks. Given a small set of domain-specific terms as seed terms, new terms from unlabeled corpora can be extracted by bootstrapping a term classifier to discover the association between seed terms and new terms. Traditional term representation method for domain-specific term extraction represents a term in a feature space of documents, which depicts association of terms which share common documents. This representation can't depict the inner-document information of terms and requires extracted terms to occur in multiple documents. A new term representation method in global contextual space is proposed for domain-specific term extraction in this paper. This representation mechanism depicts the association of terms which share common global contexts. The information of terms within certain document and among corpora is depicted by global contexts. Experiments on Chinese web corpus show that the proposed domain-specific term extraction method with global contextual representation outperforms traditional method with representation mechanism in documents space. The improvement for low frequency terms is much higher for the proposed method. Tao Liu 0001, Xiaolong Wang 0001, Bingquan Liu, Yuanchao Liu |
SMC | 2 |
| 2007 | Improving web search ranking by incorporating summarizationabstractThough link analysis based page ranking approaches have reached great success in commercial search engines (SE), the content based relevance computing approaches also play a very important role in the ranking of information retrieval results. Since most of existing relevance computing algorithms are running on the full text of a web page, this paper is focused on the relevance computing between user’s query and the auto-generated text summarization of each webpage. The first part of this paper provides a brief introduction of the state of art of relevance computing in SE. The inference network approach is especially concerned in this paper since it is the baseline method in our experiment SE system. Then the auto text summarization method based on multi-source integration is introduced, and the full text of each web page is replaced by its auto-generated abstract to compute the relevance between the webpage and user query. To evaluate the effect of the condensation representation of full text on the relevance based page rank of a system, several experiments are conducted in the last part of this paper, which include the method remarked above with different compress ratio, and the full text based ranking. In addition to the efficiency gain of the SE system, the experiment results also shows that the ranking results based on the summary generated by our text summarization system with 30% compress ratio can also get 11.29% of the precision improvement for the SE system. Xianjun Meng, Qingcai Chen, Xiaolong Wang 0001, Xiao-Hong Yang |
SMC | 3 |
| 2007 | Intelligent chinese text input technology for mobile computingabstractA sentence-based Chinese text input method system is proposed in this paper, which is implemented on both Symbian S60 and Windows Mobile platform with such characters as easy-to-use, efficient and smart. The whole system is compacted within 150k, and can be integrated with cell phone, PDA and remoter. Xuan Wang 0002, Lin Yao 0004, Xiaolong Wang 0001 |
SMC | 3 |
| 2007 | Multi-document summarization based on rhetorical structure: Sentence extraction and evaluationabstractA Multi-document Rhetorical Structure (MRS) is proposed for multi-document automatic summarization task. This structure can represent interrelationship between text units at different levels of granularity and can describe simultaneously the happen and change of various events. MRS simplify traditional multi-document representation in cross structure theory and supplement change and distribution information of events topics which cannot be obtained in information fusion theory. Concretely, a series of algorithms including building MRS, multi-document information fusion based MRS and summarization generation are proposed. The capability of concurrently fuse multiple knowledge sources of MRS strategies is testified by sets of experiments and shows good result. Yong-Dong Xu, Xiaolong Wang 0001, Tao Liu 0001, Zhi-Ming Xu |
SMC | 2 |
| 2007 | Protein-protein interaction site prediction based on conditional random fieldsabstractMOTIVATION: We are motivated by the fast-growing number of protein structures in the Protein Data Bank with necessary information for prediction of protein-protein interaction sites to develop methods for identification of residues participating in protein-protein interactions. We would like to compare conditional random fields (CRFs)-based method with conventional classification-based methods that omit the relation between two labels of neighboring residues to show the advantages of CRFs-based method in predicting protein-protein interaction sites. RESULTS: The prediction of protein-protein interaction sites is solved as a sequential labeling problem by applying CRFs with features including protein sequence profile and residue accessible surface area. The CRFs-based method can achieve a comparable performance with state-of-the-art methods, when 1276 nonredundant hetero-complex protein chains are used as training and test set. Experimental result shows that CRFs-based method is a powerful and robust protein-protein interaction site prediction method and can be used to guide biologists to make specific experiments on proteins. AVAILABILITY: http://www.insun.hit.edu.cn/~mhli/site_CRFs/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Lin 0001, Xiaolong Wang 0001, Tao Liu 0001 |
Bioinform. | 3 |
| 2007 | Exploiting residue-level and profile-level interface propensities for usage in binding sites prediction of proteinsabstractBACKGROUND: Recognition of binding sites in proteins is a direct computational approach to the characterization of proteins in terms of biological and biochemical function. Residue preferences have been widely used in many studies but the results are often not satisfactory. Although different amino acid compositions among the interaction sites of different complexes have been observed, such differences have not been integrated into the prediction process. Furthermore, the evolution information has not been exploited to achieve a more powerful propensity. RESULT: In this study, the residue interface propensities of four kinds of complexes (homo-permanent complexes, homo-transient complexes, hetero-permanent complexes and hetero-transient complexes) are investigated. These propensities, combined with sequence profiles and accessible surface areas, are inputted to the support vector machine for the prediction of protein binding sites. Such propensities are further improved by taking evolutional information into consideration, which results in a class of novel propensities at the profile level, i.e. the binary profiles interface propensities. Experiment is performed on the 1139 non-redundant protein chains. Although different residue interface propensities among different complexes are observed, the improvement of the classifier with residue interface propensities can be negligible in comparison with that without propensities. The binary profile interface propensities can significantly improve the performance of binding sites prediction by about ten percent in term of both precision and recall. CONCLUSION: Although there are minor differences among the four kinds of complexes, the residue interface propensities cannot provide efficient discrimination for the complicated interfaces of proteins. The binary profile interface propensities can significantly improve the performance of binding sites prediction of protein, which indicates that the propensities at the profile level are more accurate than those at the residue level. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001, Yi Guan |
BMC Bioinform. | 2 |
| 2007 | The study of a nonstationary maximum entropy Markov model and its application on the pos-tagging taskabstractSequence labeling is a core task in natural language processing. The maximum entropy Markov model (MEMM) is a powerful tool in performing this task. This article enhances the traditional MEMM by exploiting the positional information of language elements. The stationary hypothesis is relaxed in MEMM, and the nonstationary MEMM (NS-MEMM) is proposed. Several related issues are discussed in detail, including the representation of positional information, NS-MEMM implementation, smoothing techniques, and the space complexity issue. Furthermore, the asymmetric NS-MEMM presents a more flexible way to exploit positional information. In the experiments, NS-MEMM is evaluated on both the Chinese and the English pos-tagging tasks. According to the experimental results, NS-MEMM yields effective improvements over MEMM by exploiting positional information. The smoothing techniques in this article effectively solve the NS-MEMM data-sparseness problem; the asymmetric NS-MEMM is also an improvement by exploiting positional information in a more flexible way. JingHui Xiao, Xiaolong Wang 0001, Bingquan Liu |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2006 | Conditional Random Fields Based Label Sequence and Information Feedback
Wei Jiang 0036, Yi Guan, Xiaolong Wang 0001 |
ICIC (2) | 3 |
| 2006 | Application of latent semantic analysis to protein remote homology detectionabstractMOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001 |
Bioinform. | 2 |
| 2006 | Novel knowledge-based mean force potential at the profile levelabstractBACKGROUND: The development and testing of functions for the modeling of protein energetics is an important part of current research aimed at understanding protein structure and function. Knowledge-based mean force potentials are derived from statistical analyses of interacting groups in experimentally determined protein structures. Current knowledge-based mean force potentials are developed at the atom or amino acid level. The evolutionary information contained in the profiles is not investigated. Based on these observations, a class of novel knowledge-based mean force potentials at the profile level has been presented, which uses the evolutionary information of profiles for developing more powerful statistical potentials. RESULTS: The frequency profiles are directly calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into binary profiles with a probability threshold. As a result, the protein sequences are represented as sequences of binary profiles rather than sequences of amino acids. Similar to the knowledge-based potentials at the residue level, a class of novel potentials at the profile level is introduced. We develop four types of profile-level statistical potentials including distance-dependent, contact, Phi/Psi dihedral angle and accessible surface statistical potentials. These potentials are first evaluated by the fold assessment between the correct and incorrect models generated by comparative modeling from our own and other groups. They are then used to recognize the native structures from well-constructed decoy sets. Experimental results show that all the knowledge-base mean force potentials at the profile level outperform those at the residue level. Significant improvements are obtained for the distance-dependent and accessible surface potentials (5-6%). The contact and Phi/Psi dihedral angle potential only get a slight improvement (1-2%). Decoy set evaluation results show that the distance-dependent profile-level potentials even outperform other atom-level potentials. We also demonstrate that profile-level statistical potentials can improve the performance of threading. CONCLUSION: The knowledge-base mean force potentials at the profile level can provide better discriminatory ability than those at the residue level, so they will be useful for protein structure prediction and model refinement. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001 |
BMC Bioinform. | 2 |
| 2005 | Principles of Non-stationary Hidden Markov Model and Its Applications to Sequence Labeling Task
JingHui Xiao, Bingquan Liu, Xiaolong Wang 0001 |
IJCNLP | 3 |
| 2005 | Combining multiple classifiers based on a statistical method for handwritten Chinese character recognitionabstractCombining multiple classifiers is a new method that achieves a substantial gain in performance in many areas of pattern recognition. This paper demonstrates a novel method (based on statistics) of combining multiple classifiers to address the task of recognizing handwritten Chinese characters. Fusion strategies are discussed to provide a basis for the architecture of the combined classifiers. The weights of these fusion strategies are assigned via a genetic algorithm (GA). These fusion strategies are then tested using our online system for handwritten Chinese character recognition. In addition, different combinatory approaches are tested for comparison purposes. These include the conventional approach that is based on the Bayesian principle and the improved weighted combination, employing shared and distinct representations. Our experimental results demonstrate the effectiveness of these combinatory approaches. Lei Lin 0001, Xiaolong Wang 0001, Daniel S. Yeung |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | A Hybrid Language Model Based On Statistics And Linguistic RulesabstractLanguage modeling is a current research topic in many domains including speech recognition, optical character recognition, handwriting recognition, machine translation and spelling correction. There are two main types of language models, the mathematical and the linguistic. The most widely used mathematical language model is the n-gram model inferred from statistics. This model has three problems: long distance restriction, recursive nature and partial language understanding. Language models based on linguistics present many difficulties when applied to large scale real texts. We present here a new hybrid language model that combines the advantages of the n-gram statistical language model with those of a linguistic language model which makes use of grammatical or semantic rules. Using suitable rules, this hybrid model can solve problems such as long distance restriction, recursive nature and partial language understanding. The new language model has been effective in experiments and has been incorporated in Chinese sentence input products for Windows and Macintosh OS. Xiaolong Wang 0001, Daniel S. Yeung, James Nga-Kwok Liu, Robert Wing Pong Luk, Xuan Wang 0002 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | A Study of Semi-discrete Matrix Decomposition for LSI in Automated Text Categorization
Qiang Wang 0001, Xiaolong Wang 0001, Yi Guan |
IJCNLP | 2 |
| 2004 | Mining Pinyin-to-character conversion rules from large-scale corpus: a rough set approachabstractThis paper introduces a rough set technique for solving the problem of mining Pinyin-to-character (PTC) conversion rules. It first presents a text-structuring method by constructing a language information table from a corpus for each pinyin, which it will then apply to a free-form textual corpus. Data generalization and rule extraction algorithms can then be used to eliminate redundant information and extract consistent PTC conversion rules. The design of our model also addresses a number of important issues such as the long-distance dependency problem, the storage requirements of the rule base, and the consistency of the extracted rules, while the performance of the extracted rules as well as the effects of different model parameters are evaluated experimentally. These results show that by the smoothing method, high precision conversion (0.947) and recall rates (0.84) can be achieved even for rules represented directly by pinyin rather than words. A comparison with the baseline tri-gram model also shows good complement between our method and the tri-gram language model. Xiaolong Wang 0001, Qingcai Chen, Daniel S. Yeung |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | An Approach To Natural Stroke Extraction For Off-Line Loosely-Constrained Handwritten Chinese CharactersabstractThis paper proposes a new approach to extracting natural strokes from the skeletons of loosely-constrained, off-line handwritten Chinese characters. It admits the output substrokes from a previously proposed fuzzy substroke extractor as its inputs. By identifying a number of expected ambiguities which include mutual similarities, unstable touches and joint/cross distortions, fuzzy stroke models are constructed and a "hit-all" fuzzy stroke matching strategy is pursued. Fuzzy partitioning technique is used to generate a ranked list of consistent stroke sets from the set of fuzzy strokes being identified. With this approach, a maximum of 20 distinct natural stroke classes can be extracted from each input character, together with an estimate on the actual count of strokes which compose the character. Our system offers a number of performance tuning capabilities such as the computation of the fuzzy scores of each extracted stroke, the adjustment on the fuzzy stroke model parameters, and the potential of incorporating one's personal writing styles into our methodology. Daniel S. Yeung, Hank-Shun Fong, Eric C. C. Tsang, Wenhao Shu, Xiaolong Wang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |