VLDB 2026 Research / reviewers in the wild / expert
Xinglong Wang
dblp:05/3134
· DBLP profile ↗
31ranked-venue papers
13as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-authorSystems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion model assisted designing self-assembling collagen mimetic peptides as biocompatible materialsabstractCollagen self-assembly supports its mechanical function, but controlling collagen mimetic peptides (CMPs) to self-assemble into higher-order oligomers with numerous functions remains challenging due to the vast potential amino acid sequence space. Herein, we developed a diffusion model to learn features from different types of human collagens and generate CMPs; obtaining 66% of synthetic CMPs could self-assemble into triple helices. Triple-helical and untwisting states were probed by melting temperature (Tm); hence, we developed a model to predict collagen Tm, achieving a state-of-art Pearson's correlation (PC) of 0.95 by cross-validation and a PC of 0.8 for predicting Tm values of synthetic CMPs. Our chemically synthesized short CMPs and recombinantly expressed long CMPs could self-assemble, with the lowest requirement for hydrogel formation at a concentration of 0.08% (w/v). Five CMPs could promote osteoblast differentiation. Our results demonstrated the potential for using computer-aided methods to design functional self-assembling CMPs. Xinglong Wang, Kangjie Xu, Ruoxi Sun 0010, Ruiyan Wang, Wenwen Tao, Kai Linghu, Shuyao Yu |
Briefings Bioinform. | 1 |
| 2025 | GRLGRN: graph representation-based learning to infer gene regulatory networks from single-cell RNA-seq dataabstractBACKGROUND: A gene regulatory network (GRN) is a graph-level representation that describes the regulatory relationships between transcription factors and target genes in cells. The reconstruction of GRNs can help investigate cellular dynamics, drug design, and metabolic systems, and the rapid development of single-cell RNA sequencing (scRNA-seq) technology provides important opportunities while posing significant challenges for reconstructing GRNs. A number of methods for inferring GRNs have been proposed in recent years based on traditional machine learning and deep learning algorithms. However, inferring the GRN from scRNA-seq data remains challenging owing to cellular heterogeneity, measurement noise, and data dropout. RESULTS: In this study, we propose a deep learning model called graph representational learning GRN (GRLGRN) to infer the latent regulatory dependencies between genes based on a prior GRN and data on the profiles of single-cell gene expressions. GRLGRN uses a graph transformer network to extract implicit links from the prior GRN, and encodes the features of genes by using both an adjacency matrix of implicit links and a matrix of the profile of gene expression. Moreover, it uses attention mechanisms to improve feature extraction, and feeds the refined gene embeddings into an output module to infer gene regulatory relationships. To evaluate the performance of GRLGRN, we compared it with prevalent models and performed ablation experiments on seven cell-line datasets with three ground-truth networks. The results showed that GRLGRN achieved the best predictions in AUROC and AUPRC on 78.6% and 80.9% of the datasets, and achieved an average improvement of 7.3% in AUROC and 30.7% in AUPRC. The interpretation discussion and the network visualization were conducted. CONCLUSIONS: The experimental results and case studies illustrate the considerable performance of GRLGRN in predicting gene interactions and provide interpretability for the prediction tasks, such as identifying hub genes in the network and uncovering implicit links. Kai Wang 0017, Fei Liu 0001, Xiaoli Luan, Xinglong Wang |
BMC Bioinform. | 5 |
| 2025 | PLMAM-PLA: A Method Using Pretrained Language Models and Attention Mechanisms for Protein-Ligand Binding Affinity PredictionabstractProtein-ligand binding affinity measures the strength of interactions between proteins and ligands. Accurately predicting this value is crucial for drug discovery and estimating enzyme kinetic parameters. In recent years, various computational models based on deep learning algorithms have been developed for predicting protein-ligand binding affinity. Most of these require data on protein structure or pockets in addition to protein sequences and ligand SMILES strings. Although integrating structural or pocket information can enhance prediction performances, sequence-based affinity prediction methods using only protein sequences and ligand SMILES strings are more convenient and efficient in practice. We have developed a novel sequence-based deep learning model, called PLMAM-PLA, to predict protein-ligand binding affinity. This model simultaneously extracts global and local features from both protein sequences and ligand SMILES by leveraging pretrained language models (ESM-2 and MolFormer) and dilated convolutional neural networks. The features are enhanced by SKNets and SENets and are further fused by successively using cross-attention and self-attention mechanisms. The output module provides the final affinity prediction value. Ablation studies emphasize the important contributions of the different modules, while visualization experiments demonstrate the efficacy of PLMAM-PLA in capturing meaningful feature representations. Additionally, case studies highlight the powerful generalization capabilities of the model, while comparisons with state-of-the-art models confirm its superior performance in predicting protein-ligand binding affinities. Kai Wang 0017, Aijie Song, Fei Liu 0001, Xiaoli Luan, Xinglong Wang |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Machine learning-assisted substrate binding pocket engineering based on structural informationabstractEngineering enzyme-substrate binding pockets is the most efficient approach for modifying catalytic activity, but is limited if the substrate binding sites are indistinct. Here, we developed a 3D convolutional neural network for predicting protein-ligand binding sites. The network was integrated by DenseNet, UNet, and self-attention for extracting features and recovering sample size. We attempted to enlarge the dataset by data augmentation, and the model achieved success rates of 48.4%, 35.5%, and 43.6% at a precision of ≥50% and 52%, 47.6%, and 58.1%. The distance of predicted and real center is ≤4 Å, which is based on SC6K, COACH420, and BU48 validation datasets. The substrate binding sites of Klebsiella variicola acid phosphatase (KvAP) and Bacillus anthracis proline 4-hydroxylase (BaP4H) were predicted using DUnet, showing high competitive performance of 53.8% and 56% of the predicted binding sites that critically affected the catalysis of KvAP and BaP4H. Virtual saturation mutagenesis was applied based on the predicted binding sites of KvAP, and the top-ranked 10 single mutations contributed to stronger enzyme-substrate binding varied while the predicted sites were different. The advantage of DUnet for predicting key residues responsible for enzyme activity further promoted the success rate of virtual mutagenesis. This study highlighted the significance of correctly predicting key binding sites for enzyme engineering. Xinglong Wang, Kangjie Xu, Kai Linghu, Beichen Zhao, Shangyang Yu, Shuyao Yu, Weizhu Zeng, Kai Wang 0017 |
Briefings Bioinform. | 1 |
| 2024 | BERT-TFBS: a novel BERT-based model for predicting transcription factor binding sites by transfer learningabstractTranscription factors (TFs) are proteins essential for regulating genetic transcriptions by binding to transcription factor binding sites (TFBSs) in DNA sequences. Accurate predictions of TFBSs can contribute to the design and construction of metabolic regulatory systems based on TFs. Although various deep-learning algorithms have been developed for predicting TFBSs, the prediction performance needs to be improved. This paper proposes a bidirectional encoder representations from transformers (BERT)-based model, called BERT-TFBS, to predict TFBSs solely based on DNA sequences. The model consists of a pre-trained BERT module (DNABERT-2), a convolutional neural network (CNN) module, a convolutional block attention module (CBAM) and an output module. The BERT-TFBS model utilizes the pre-trained DNABERT-2 module to acquire the complex long-term dependencies in DNA sequences through a transfer learning approach, and applies the CNN module and the CBAM to extract high-order local features. The proposed model is trained and tested based on 165 ENCODE ChIP-seq datasets. We conducted experiments with model variants, cross-cell-line validations and comparisons with other models. The experimental results demonstrate the effectiveness and generalization capability of BERT-TFBS in predicting TFBSs, and they show that the proposed model outperforms other deep-learning models. The source code for BERT-TFBS is available at https://github.com/ZX1998-12/BERT-TFBS. Kai Wang 0017, Fei Liu 0001, Xiaoli Luan, Xinglong Wang |
Briefings Bioinform. | 6 |
| 2018 | Reachability for airline networks: fast algorithm for shortest path problem with time windows
Xiaofeng Gao 0001, Yueyang Xianzang, Xiaotian You, Yaru Dang, Guihai Chen, Xinglong Wang |
Theor. Comput. Sci. | 6 |
| 2017 | Research on Arrival Integration Method for Point Merge System in Tactical Operation
Yannan Qi, Xinglong Wang |
COCOA (1) | 2 |
| 2017 | FTRS: A mechanism for reducing flow table entries in software defined networks
Bing Leng, Liusheng Huang, Chunming Qiao, Hongli Xu 0001, Xinglong Wang |
Comput. Networks | 5 |
| 2017 | Auction-based resource allocation for cooperative cognitive radio networks
Xinglong Wang, Liusheng Huang, Hongli Xu 0001, He Huang 0001 |
Comput. Commun. | 1 |
| 2017 | FAST: truthful auction with access flexibility for cooperative communicationsabstractCooperative communication has great potential to enhance the performance of wireless networks by exploiting relay nodes' spatial diversity. Relay assignment plays a vital role in making the most of this potential. However, most previous studies investigate relay assignment in a rather static manner which may lead to the under‐utilisation of relay nodes, especially in the scenarios where source nodes' demands are time varying. Hence, in this study, the authors consider the relay assignment problem for cooperative networks using auction model with access flexibility, i.e. providing sufficient flexibility for source nodes to access relay nodes in time domain. They divide relay nodes' access time into multiple smaller access units and permit each source node to bid for a bundle of desired access units in the auction. This model is formulated as a combinational auction whose winner determination problem is NP hard. They thus present an auction mechanism with an elaborated greed strategy which not only achieves adorable properties such as truthfulness, individual rationality and computational efficiency, but also guarantees near‐optimal social welfare. Finally, they conduct extensive evaluations to verify the performance of their mechanism. Xinglong Wang, Liusheng Huang, Hongli Xu 0001, He Huang 0001 |
IET Commun. | 1 |
| 2017 | Price-based resource allocation for revenue maximization with cooperative communication
Hongli Xu 0001, Shaojie Tang 0001, Xinglong Wang, Long Chen 0006, Liusheng Huang |
Wirel. Networks | 3 |
| 2016 | A Comprehensive Reachability Evaluation for Airline Networks with Multi-constraints
Xiaotian You, Xiaofeng Gao 0001, Yaru Dang, Guihai Chen, Xinglong Wang |
COCOA | 5 |
| 2016 | Social Welfare Maximization Auction for Secondary Spectrum Markets: A Long-Term PerspectiveabstractDynamic secondary spectrum markets have gained tremendous attentions recently, which can provide significant flexibility for trading spectrum by conducting auctions periodically. The high price of spectrum necessitates a thorough consideration on secondary users' budget constraint, i.e. the total money they could pay. However, previous studies rarely deal with the practical issue of budget constraint, and concentrate on maximizing social welfare greedily in a single round auction, which may not guarantee their performance after multiple rounds of auctions. In this paper, we investigate the techniques for budget constrained periodic spectrum auction, which ensures approximate social welfare maximization from a long- term perspective. With the celebrated primal-dual method, we present a Periodic Spectrum Auction framework (PSA) that runs a tailored One Round Spectrum Auction (ORSA) in each round. In the ORSA, we achieve critical properties such as truthfulness, individual rationality and computational efficiency. Due to the dual fitting technique, ORSA not only achieves approximate social welfare maximization in one round, but also guarantees only a small loss of approximation ratio when runs in multiple rounds under the PSA framework. Finally, we conduct extensive simulations to demonstrate the performance of our schemes. Xinglong Wang, Liusheng Huang, Hongli Xu 0001, He Huang 0001 |
SECON | 1 |
| 2016 | A self-adaptive reconfiguration scheme for throughput maximization in municipal WMNs
Bing Leng, Liusheng Huang, Hongli Xu 0001, Chenkai Yang, Xinglong Wang |
J. Parallel Distributed Comput. | 5 |
| 2016 | Shared Relay Assignment (SRA) for Many-to-One Traffic in Cooperative NetworksabstractRelay assignment significantly affects the performance of the cooperative communication, which is an emerging technology for the future mobile system. Previous studies in this area have mostly focused on assigning a dedicated relay to each source-destination pair for one-to-one (121) traffic. However, many-to-one (M21) traffic, which is also common in many situations (for example, several users associate with one access point in a wireless access network such as a WLAN), hasn't been well studied. This paper addresses the shared relay assignment (SRA) problem for M21 traffic. We formulate two new optimization problems: one is to maximize the minimum throughput among all the sources (hereafter called M21-SRA-MMT), and the other is to maximize the total throughput over all the sources while maintaining some degree of fairness (hereafter called M21-SRA-MTT). As the optimal solutions to the two problems are hard to find, we propose two approximation algorithms whose performance factors are 5.828 and 3, respectively, based on the rounding mechanism. Extensive simulation results show that our algorithms for M21-SRA-MMTcan significantly improve the minimum throughput compared with existing algorithms, while our algorithm for M21-SRA-MTTcan achieve the close-to-optimal performance. Hongli Xu 0001, Liusheng Huang, Chunming Qiao, Xinglong Wang, Shan Lin 0001, Yu-e Sun |
IEEE Trans. Mob. Comput. | 4 |
| 2015 | A mechanism for reducing flow tables in software defined networkabstractThe software defined network (SDN) has been developing tremendously in recent years. Numerous studies are proposed on the performance of SDN in academia. The idea of programmable network which is the foundation of SDN ensures dynamic network management by separating the control plane away from the network switches. The focus of this paper is the flow table's size in OpenFlow switches which is a significant bottleneck in SDN's practical application. We propose a mechanism named “Flow Table Reduction Scheme”(FTRS) for reducing the flow table size and maintaining the omnipotent controller's functions at the same time. We test the performance of FTRS both in simulation and experiment and the results show that FTRS is able to reduce 98% at most of the size of flow table with no impact on network's normal functions. Bing Leng, Liusheng Huang, Xinglong Wang, Hongli Xu 0001 |
ICC | 3 |
| 2015 | Primary Secrecy Is Achievable: Optimal Secrecy Rate in Overlay CRNs with an Energy Harvesting Secondary TransmitterabstractTo tackle the challenging secrecy communication problem in energy harvesting cognitive radio networks, this paper considers an overlay system with one energy harvesting secondary user (SU) to assist primary transmission under the assumption that the primary channel at primary receiver is worse than the eavesdropper. Under such scenario, we optimize the secrecy rate of the PU transmitter by jointly investigating energy harvesting slot, cooperative transmission slot and so on. Given the transmission rate requirement between SUs, the optimization problem is formulated as a mixed integer non-linear (MINLP) program. Due to the special features, we design a polynomial time algorithm SRMA to optimally solve this problem. The algorithm computes the lower bound and upper bound of the transmission power in a secondary transmitter, which are relative with the QoS requirement and energy harvesting parameters. Then SRMA determines its optimal transmission power by iteratively searching between two bounds. Numerical results demonstrate that the primary secrecy rate grows with the increasing energy save ratio and optimal energy save ratio is inversely proportional to the energy harvesting rate. Long Chen 0006, Liusheng Huang, Hongli Xu 0001, Chenkai Yang, Zehao Sun, Xinglong Wang |
ICCCN | 6 |
| 2015 | Truthful Auction for Resource Allocation in Cooperative Cognitive Radio NetworksabstractCooperative cognitive radio network (CCRN) is a promising paradigm to increase spectrum utilization and exploit spatial diversity. The allocation of two related resources, i.e. spectrum and relay nodes, plays a fundamental role in the performance of CCRNs. However, previous works either lack of incentives for both primary users (PUs) and relay nodes to participate in or consider spectrum auction and relay auction separately. In this paper, we consider a static cooperative cognitive radio network scenario with several PUs and multiple secondary user coteries, each of which consists of a set of secondary users who are interested in sharing the same secondary relay node. We model the problem of joint spectrum allocation and relay allocation as a hierarchical auction and propose TERA, which is the first Truthful auction mechanism for Efficient Resource Allocation in CCRNs. We show that TERA satisfies critical economic properties such as truthful, individual rationality, budget balance, supply limits and computational efficiency. Furthermore, we theoretically prove TERA can achieve near-optimal revenue with high probability. Finally, extensive simulation results show that TERA is efficient and able to improve the utility of PUs and relay nodes significantly up to 125% and 151% respectively. Xinglong Wang, Liusheng Huang, Hongli Xu 0001, He Huang 0001 |
ICCCN | 1 |
| 2015 | A self-adaptive reconfiguration scheme for throughput maximization in municipal WMNsabstractWireless mesh networks (WMNs) are being used and deployed widely all around the world for various reasons such as public safety, environmental monitoring and city-wide wireless Internet services [4], [12]. Although faced with some difficulties, wireless mesh network is thought to be a preferred municipal Internet service provider [5] with huge potential due to its automatic connection, ease of installation, dynamic route discovery, flexibility and other advantages. Since plenty of cities such as Singapore and the city of Cambridge have put providing ubiquitous Internet access on the agenda [3], the research on providing QoS-guaranteed municipal WMNs is in urgent need. Bing Leng, Liusheng Huang, Chenkai Yang, Hongli Xu 0001, Xinglong Wang |
IWQoS | 5 |
| 2014 | Replica placement in content delivery networks with stochastic demands and M/M/1 serversabstractContent Delivery Network (CDN) is proposed for replicating data objects at multiple locations in the network and encounters vast potential for future development, as a result of which, a number of replica placement techniques have been proposed over the last decade. However, most of the existing works on replica placement (RP) ignore the statistical property of the demands and the restricted service rate of the servers. In this paper, we investigate the techniques of replica placement in CDNs with stochastic demands and M/M/1 servers to optimize the overall performance in the network. We first model the demands and the servers as independent Poisson streams and simple M/M/1 queueing systems, respectively. Then, a formal definition and formalization of RP problem will be given. We show that RP problem is NP-complete and propose two heuristic algorithms: Greedy Dropping (GD) and Tabu Search (TS). We conduct abundant simulation experiments to evaluate the performance of our proposed algorithms. According to our simulation results, both of the two algorithms are efficient in finding a feasible solution with high probability. Especially, the TS decreases the average delay of the demands about 50% on average. Chenkai Yang, Liusheng Huang, Bing Leng, Hongli Xu 0001, Xinglong Wang |
IPCCC | 5 |
| 2014 | Shared relay assignment (SRA) for many-to-one traffic in cooperative wireless transmissionsabstractRelay assignment significantly affects the performance of cooperative communications. Previous studies in this area have mostly focused on assigning a dedicated relay to each source-destination pair for one-to-one (121) traffic. On the other hand, many-to-one (M21) traffic, which is also common in many situations (for example, several users associate with one access point in a wireless access network such as a WLAN), hasn't been well studied. This paper addresses the shared relay assignment (SRA) problem for M21 traffic. We formulate two new optimization problems: one is to maximize the minimum throughput among all the sources (hereafter called M21-SRA-MMT), and the other is to maximize the total throughput over all the sources while maintaining some degree of fairness (hereafter called M21-SRA-MTT). As both of these problems are NP-hard, we propose two approximation algorithms whose performance factors are 5.828 and 3, respectively, based on the rounding mechanism. Extensive simulation results show that our algorithm for M21-SRA-MMT can significantly improve the minimum throughput compared with existing algorithms, while our algorithm for M21-SRA-MTT can achieve the close-to-optimal performance. Hongli Xu 0001, Liusheng Huang, Chunming Qiao, Xinglong Wang, Yu-e Sun |
IWQoS | 4 |
| 2013 | Topology Control with vMIMO Communication in Wireless Sensor NetworksabstractVirtual multi-input-multi-output (or vMIMO) communication is a promising technology to improve the spatial diversity of wireless networks. Using this mechanism, multiple single-antenna nodes can coordinate their transmissions and receptions so as to reduce power consumption. This paper studies the problem of constructing an energy-efficient topology in wireless sensor networks using vMIMO communication. We first define the problem involving joint optimization of vMIMO, partner selection and topology control. As this problem is NP-Complete, a distributed and heuristic algorithm, called vMIMO topology control (VMTC), is proposed to solve this problem. The algorithm uses an improved binary searching method to obtain an initial power assignment. A local competition method is then adopted to implement the partner selection. At last, we reduce the power consumption of each node by using efficient vMIMO modes. Our theoretical analysis show that this algorithm can achieve an approximate performance of O(1). Our simulations show that VMTC helps to decrease the power consumptions by about 32% compared to the existing algorithms. Hongli Xu 0001, Liusheng Huang, Chunming Qiao, Xinglong Wang, Yu-e Sun |
IEEE Trans. Wirel. Commun. | 4 |
| 2011 | Automatic extraction of angiogenesis bioprocess from textabstractMOTIVATION: Understanding key biological processes (bioprocesses) and their relationships with constituent biological entities and pharmaceutical agents is crucial for drug design and discovery. One way to harvest such information is searching the literature. However, bioprocesses are difficult to capture because they may occur in text in a variety of textual expressions. Moreover, a bioprocess is often composed of a series of bioevents, where a bioevent denotes changes to one or a group of cells involved in the bioprocess. Such bioevents are often used to refer to bioprocesses in text, which current techniques, relying solely on specialized lexicons, struggle to find. RESULTS: This article presents a range of methods for finding bioprocess terms and events. To facilitate the study, we built a gold standard corpus in which terms and events related to angiogenesis, a key biological process of the growth of new blood vessels, were annotated. Statistics of the annotated corpus revealed that over 36% of the text expressions that referred to angiogenesis appeared as events. The proposed methods respectively employed domain-specific vocabularies, a manually annotated corpus and unstructured domain-specific documents. Evaluation results showed that, while a supervised machine-learning model yielded the best precision, recall and F1 scores, the other methods achieved reasonable performance and less cost to develop. AVAILABILITY: The angiogenesis vocabularies, gold standard corpus, annotation guidelines and software described in this article are available at http://text0.mib.man.ac.uk/~mbassxw2/angiogenesis/ CONTACT: [email protected]. Xinglong Wang, Iain McKendrick, Ian P. Barrett, Ian Dix, Tim French 0003, Jun'ichi Tsujii, Sophia Ananiadou |
Bioinform. | 1 |
| 2011 | The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full textabstractBACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows. Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu |
BMC Bioinform. | 26 |
| 2011 | Detecting experimental techniques and selecting relevant documents for protein-protein interactions from biomedical literatureabstractBACKGROUND: The selection of relevant articles for curation, and linking those articles to experimental techniques confirming the findings became one of the primary subjects of the recent BioCreative III contest. The contest's Protein-Protein Interaction (PPI) task consisted of two sub-tasks: Article Classification Task (ACT) and Interaction Method Task (IMT). ACT aimed to automatically select relevant documents for PPI curation, whereas the goal of IMT was to recognise the methods used in experiments for identifying the interactions in full-text articles. RESULTS: We proposed and compared several classification-based methods for both tasks, employing rich contextual features as well as features extracted from external knowledge sources. For IMT, a new method that classifies pair-wise relations between every text phrase and candidate interaction method obtained promising results with an F1 score of 64.49%, as tested on the task's development dataset. We also explored ways to combine this new approach and more conventional, multi-label document classification methods. For ACT, our classifiers exploited automatically detected named entities and other linguistic information. The evaluation results on the BioCreative III PPI test datasets showed that our systems were very competitive: one of our IMT methods yielded the best performance among all participants, as measured by F1 score, Matthew's Correlation Coefficient and AUC iP/R; whereas for ACT, our best classifier was ranked second as measured by AUC iP/R, and also competitive according to other metrics. CONCLUSIONS: Our novel approach that converts the multi-class, multi-label classification problem to a binary classification problem showed much promise in IMT. Nevertheless, on the test dataset the best performance was achieved by taking the union of the output of this method and that of a multi-class, multi-label document classifier, which indicates that the two types of systems complement each other in terms of recall. For ACT, our system exploited a rich set of features and also obtained encouraging results. We examined the features with respect to their contributions to the classification results, and concluded that contextual words surrounding named entities, as well as the MeSH headings associated with the documents were among the main contributors to the performance. Xinglong Wang, Rafal Rak, Angelo C. Restificar, Chikashi Nobata, C. J. Rupp, Riza Theresa Batista-Navarro, Raheel Nawaz, Sophia Ananiadou |
BMC Bioinform. | 1 |
| 2011 | Extracting Secondary Bio-Event Arguments with Extraction ConstraintsabstractThis paper describes our bio-event extraction system developed for the BioNLP 2009 Shared Task 2, with focus on its capability of extracting secondary biological event arguments from literature. Shared Task 2 is particularly interesting because when browsing literature, biologists often need to understand conditions surrounding biological events, which are usually expressed by secondary event arguments (e.g., binding sites). To achieve our goal, we take an approach that extracts n-ary relations from text using event extraction constraints automatically generated from a training corpus. Event constraints consist of sequences of trigger words and semantic roles which we automatically identify using Conditional Random Fields (CRFs). Unlike most other systems participating in this shared task, our system is light-weight and relies on neither external resources (e.g., Ontologies and dictionaries) nor natural language processing software (e.g., POS taggers and parsers). The official test results show that our approach performed well on extracting secondary arguments in Task 2, yielding the highest precision at 76.62% and the second highest F-measure at 43.22%. Yutaka Sasaki, Xinglong Wang, Sophia Ananiadou |
Comput. Intell. | 2 |
| 2010 | Disambiguating the species of biomedical named entities using natural language parsersabstractMOTIVATION: Text mining technologies have been shown to reduce the laborious work involved in organizing the vast amount of information hidden in the literature. One challenge in text mining is linking ambiguous word forms to unambiguous biological concepts. This article reports on a comprehensive study on resolving the ambiguity in mentions of biomedical named entities with respect to model organisms and presents an array of approaches, with focus on methods utilizing natural language parsers. RESULTS: We build a corpus for organism disambiguation where every occurrence of protein/gene entity is manually tagged with a species ID, and evaluate a number of methods on it. Promising results are obtained by training a machine learning model on syntactic parse trees, which is then used to decide whether an entity belongs to the model organism denoted by a neighbouring species-indicating word (e.g. yeast). The parser-based approaches are also compared with a supervised classification method and results indicate that the former are a more favorable choice when domain portability is of concern. The best overall performance is obtained by combining the strengths of syntactic features and supervised classification. AVAILABILITY: The corpus and demo are available at http://www.nactem.ac.uk/deca_details/start.cgi, and the software is freely available as U-Compare components (Kano et al., 2009): NaCTeM Species Word Detector and NaCTeM Species Disambiguator. U-Compare is available at http://-compare.org/ Xinglong Wang, Jun'ichi Tsujii, Sophia Ananiadou |
Bioinform. | 1 |
| 2009 | Classifying Relations for Biomedical Named Entity Disambiguation
Xinglong Wang, Jun'ichi Tsujii, Sophia Ananiadou |
EMNLP | 1 |
| 2008 | Learning the Species of Biomedical Named Entities from Annotated Corpora
Xinglong Wang, Claire Grover |
LREC | 1 |
| 2008 | Distinguishing the species of biomedical named entities for term identificationabstractBACKGROUND: Term identification is the task of grounding ambiguous mentions of biomedical named entities in text to unique database identifiers. Previous work on term identification has focused on studying species-specific documents. However, full-length articles often describe entities across a number of species, in which case resolving the ambiguity of model organisms in entities is critical to achieving accurate term identification. RESULTS: We developed and compared a number of rule-based and machine-learning based approaches to resolving species ambiguity in mentions of biomedical named entities, and demonstrated that a hybrid method achieved the best overall accuracy at 71.7%, as tested on the gold-standard ITI-TXM corpora. By utilising the species information predicted by the hybrid tagger, our rule-based term identification system was improved significantly by up to 11.6%. CONCLUSION: This paper shows that, in the context of identifying terms involving multiple model organisms, integration of an accurate species disambiguation system can significantly improve the performance of term identification systems. Xinglong Wang, Michael Matthews |
BMC Bioinform. | 1 |
| 2007 | Rule-Based Protein Term Identification with Help from Automatic Species Tagging
Xinglong Wang |
CICLing | 1 |