VLDB 2026 Research / reviewers in the wild / expert
Goran Nenadic
dblp:60/6717
· DBLP profile ↗
67ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-0795-5363ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware RefinementabstractHao Li, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Batista-Navarro, Goran Nenadic. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Li 0074, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Theresa Batista-Navarro, Goran Nenadic |
ACL (1) | 6 |
| 2025 | Evaluating Differentially Private Generation of Domain-Specific TextabstractGenerative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To address this, differentially private synthetic data generation has emerged as a promising alternative. In this work, we introduce a unified benchmark to systematically evaluate the utility and fidelity of text datasets generated under formal Differential Privacy (DP) guarantees. Our benchmark addresses key challenges in domain-specific benchmarking, including choice of representative data and realistic privacy budgets, accounting for pre-training and a variety of evaluation metrics. We assess state-of-the-art privacy-preserving generation methods across five domain-specific datasets, revealing significant utility and fidelity degradation compared to real data, especially under strict privacy constraints. These findings underscore the limitations of current approaches, outline the need for advanced privacy-preserving data sharing methods and set a precedent regarding their evaluation in realistic scenarios. Viktor Schlegel, Srinivasan Nandakumar, Iqra Zahid, Yuping Wu 0001, Warren Del-Pinto, Goran Nenadic, Siew-Kei Lam, Jie Zhang 0073, Anil A. Bharath |
CIKM | 7 |
| 2025 | DocDiscNER: Enhanced Document-Level Discontinuous NER via Coordination Ellipses Resolution and Self-Consistency DecodingabstractIdentifying entities in medical text often involves dealing with discontinuous word sequences or entities sharing a common head, which pose significant challenges for traditional Named Entity Recognition (NER) systems. Current state-of-the-art discontinuous NER models typically process each sentence in isolation, overlooking valuable intra-sentence context. However, recent studies have shown that large language models (LLMs) perform exceptionally well when provided such context. In this work, we introduce DocDiscNER, a novel approach to discontinuous NER, which features (i) a context-aware document chunking method that provides contextually related segments as input for LLM-based NER models; (ii) a dataset and approach for coordination ellipses resolution, to address candidate spans sharing common heads and (iii) a self-consistency decoding strategy that uses self-ensembling and a majority voting mechanism to select the most consistent predictions as entity spans. We demonstrate the effectiveness and generalisability of our method on three discontinuous NER benchmarks, achieving new state-of-the-art (SOTA) performance on two of them–CADEC and ShARe-14 (2.48 and 2.2 absolute F1 points gain, respectively); while achieving competitive results on ShARe-13. In addition, our method surpasses previous SOTA performance specifically in recognising discontinuous mentions. A deeper analysis unveils that incorporating semantically relevant context significantly enhances overall NER performance compared to using individual sentences as input. Areej Alhassan, Viktor Schlegel, Rina Carines Cabral, Riza Theresa Batista-Navarro, Soyeon Caren Han, Josiah Poon, Goran Nenadic |
ECAI | 7 |
| 2025 | BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion ModelingabstractTime-series Generation (TSG) is a prominent research area with broad applications in simulations, data augmentation, and counterfactual analysis. While existing methods have shown promise in unconditional single-domain TSG, real-world applications demand for cross-domain approaches capable of controlled generation tailored to domain-specific constraints and instance-level requirements. In this paper, we argue that text can provide semantic insights, domain information and instance-specific temporal patterns, to guide and improve TSG. We introduce “Text-Controlled TSG”, a task focused on generating realistic time series by incorporating textual descriptions. To address data scarcity in this setting, we propose a novel LLM-based Multi-Agent framework that synthesizes diverse, realistic text-to-TS datasets. Furthermore, we introduce Bridge, a hybrid text-controlled TSG framework that integrates semantic prototypes with text description for supporting domain-level guidance. This approach achieves state-of-the-art generation fidelity on 11 of 12 datasets, and improves controllability by up to 12% on MSE and 6% MAE compared to no text input generation, highlighting its potential for generating tailored time-series data. Hao Li 0074, Yu-Hao Huang 0002, Chang Xu 0008, Viktor Schlegel, Renhe Jiang, Riza Theresa Batista-Navarro, Goran Nenadic, Jiang Bian 0002 |
ICML | 7 |
| 2025 | CAST: Corpus-Aware Self-similarity Enhanced Topic modellingabstractYanan Ma, Chenghao Xiao, Chenhan Yuan, Sabine N Van Der Veer, Lamiece Hassan, Chenghua Lin, Goran Nenadic. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chenghao Xiao, Chenhan Yuan, Sabine van der Veer, Lamiece Hassan, Goran Nenadic |
NAACL (Long Papers) | 7 |
| 2025 | TriG-NER: Triplet-Grid Framework for Discontinuous Named Entity RecognitionabstractDiscontinuous Named Entity Recognition (DNER) presents a challenging problem where entities may be scattered across multiple non-adjacent tokens, making traditional sequence labelling approaches inadequate. Existing methods predominantly rely on custom tagging schemes to handle these discontinuous entities, resulting in models tightly coupled to specific tagging strategies and lacking generalisability across diverse datasets. To address these challenges, we propose TriG-NER, a novel Triplet-Grid Framework that introduces a generalisable approach to learning robust token-level representations for discontinuous entity extraction. Our framework applies triplet loss at the token level, where similarity is defined by word pairs existing within the same entity, effectively pulling together similar and pushing apart dissimilar ones. This approach enhances entity boundary detection and reduces the dependency on specific tagging schemes by focusing on word-pair relationships within a flexible grid structure. We evaluate TriG-NER on three benchmark DNER datasets and demonstrate significant improvements over existing grid-based architectures. These results underscore our framework's effectiveness in capturing complex entity structures and its adaptability to various tagging schemes, setting a new benchmark for discontinuous entity extraction. Rina Carines Cabral, Soyeon Caren Han, Areej Alhassan, Riza Theresa Batista-Navarro, Goran Nenadic, Josiah Poon |
WWW | 5 |
| 2025 | Discontinuous named entities in clinical text: A systematic literature reviewabstractOBJECTIVE: Extracting named entities from clinical free-text presents unique challenges, particularly when dealing with discontinuous entities-mentions that are separated by unrelated words. Traditional NER methods often struggle to accurately identify these entities, prompting the development of specialised computational solutions. This paper systematically reviews and presents the methodologies developed for Discontinuous Named Entity Recognition in clinical texts, highlighting their effectiveness and the challenges they face. METHOD: We conducted a systematic literature review focused on discontinuous named entities, using structured searches across four Computer Science-related and one medical-related electronic database. A combination of search terms, grouped into three synonym categories-problem, entity/approach, and task-yielded 2,442 articles. Guided by our research objectives, we identified five key dimensions to systematically annotate and normalise the data for comprehensive analysis. RESULT: The review included 44 studies which were coded across several key dimensions: the chronological development of approaches, the corpora used, the downstream tasks affected by discontinuous named entities, the methodological approaches proposed to address the issue, and the reported performance outcomes. The discussion section examines the challenges encountered in this area and suggests potential directions for future research. CONCLUSION: Significant progress has been made in discontinuous named entity recognition; however, there remains a need for more adaptable, generalisable solutions that are independent of custom annotation schemes. Exploring various configurations of generative language models presents a promising avenue for advancing this area. Additionally, future research should investigate the impact of precise versus imprecise recognition of discontinuous entities on clinical downstream tasks to better understand its practical implications in healthcare applications. Areej Alhassan, Viktor Schlegel, Monira Aloud, Riza Theresa Batista-Navarro, Goran Nenadic |
J. Biomed. Informatics | 5 |
| 2024 | MTUncertainty: Assessing the Need for Post-editing of Machine Translation Outputs by Fine-tuning OpenAI LLMsabstractTranslation Quality Evaluation (TQE) is an essential step of the modern translation production process. TQE is critical in assessing both machine translation (MT) and human translation (HT) quality without reference translations. The ability to evaluate or even simply estimate the quality of translation automatically may open significant efficiency gains through process optimisation.This work examines whether the state-of-the-art large language models (LLMs) can be used for this uncertainty estimation of MT output quality. We take OpenAI models as an example technology and approach TQE as a binary classification task.On eight language pairs including English to Italian, German, French, Japanese, Dutch, Portuguese, Turkish, and Chinese, our experimental results show that fine-tuned gpt3.5 can demonstrate good performance on translation quality prediction tasks, i.e. whether the translation needs to be edited.Another finding is that simply increasing the sizes of LLMs does not lead to apparent better performances on this task by comparing the performance of three different versions of OpenAI models: curie, davinci, and gpt3.5 with 13B, 175B, and 175B parameters, respectively. Serge Gladkoff, Lifeng Han, Gleb Erofeev, Irina Sorokina, Goran Nenadic |
EAMT (1) | 5 |
| 2024 | CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models using Real and Synthetic Back-Translation DataabstractNeural Machine Translation (NMT) for low-resource languages remains a challenge for many NLP researchers. In this work, we deploy a standard data augmentation methodology by back-translation to a new language translation direction, i.e., Cantonese-to-English. We present the models we fine-tuned using the limited amount of real data and the synthetic data we generated using back-translation by three models: OpusMT, NLLB, and mBART.We carried out automatic evaluation using a range of different metrics including those that are lexical-based and embedding-based.Furthermore, we create a user-friendly interface for the models we included in this project, CantonMT, and make it available to facilitate Cantonese-to-English MT research. Researchers can add more models to this platform via our open-source CantonMT toolkit, available at https://github.com/kenrickkung/CantoneseTranslation. Kung Yin Hong, Lifeng Han, Riza Theresa Batista-Navarro, Goran Nenadic |
EAMT (1) | 4 |
| 2024 | De-identification of clinical free text using natural language processing: A systematic review of current approaches
Aleksandar Kovacevic, Bojana Basaragin, Nikola Milosevic, Goran Nenadic |
Artif. Intell. Medicine | 4 |
| 2023 | Do You Hear The People Sing? Key Point Analysis via Iterative Clustering and Abstractive SummarisationabstractArgument summarisation is a promising but currently under-explored field.Recent work has aimed to provide textual summaries in the form of concise and salient short texts, i.e., key points (KPs), in a task known as Key Point Analysis (KPA).One of the main challenges in KPA is finding high-quality key point candidates from dozens of arguments even in a small corpus.Furthermore, evaluating key points is crucial in ensuring that the automatically generated summaries are useful.Although automatic methods for evaluating summarisation have considerably advanced over the years, they mainly focus on sentence-level comparison, making it difficult to measure the quality of a summary (a set of KPs) as a whole.Aggravating this problem is the fact that human evaluation is costly and unreproducible.To address the above issues, we propose a two-step abstractive summarisation framework based on neural topic modelling with an iterative clustering procedure, to generate key points which are aligned with how humans identify key points.Our experiments show that our framework advances the state of the art in KPA, with performance improvement of up to 14 (absolute) percentage points, in terms of both ROUGE and our own proposed evaluation metrics 1 .Furthermore, we evaluate the generated summaries using a novel set-based evaluation toolkit.Our quantitative analysis demonstrates the effectiveness of our proposed evaluation metrics in assessing the quality of generated KPs.Human evaluation further demonstrates the advantages of our approach and validates that our proposed evaluation metric is more consistent with human judgment than ROUGE scores.Children express themselves through the clothes they wear and should be able to do this at school School uniform is harming the student's self expression School uniforms are expensive [...]Children should be able to dress as they wish, within reason, at school rather than being restricted from expressing themselves through their clothes.Children should be allowed to express themselves School uniform is unaffordable for many single parents and should be abandoned.School uniforms are an expense that many families can't afford.there are plenty of ways to get very cheap clothing, but discounted uniforms are more difficult to obtain.School uniforms are expensive and puts an undue burden on the parents of the students.School uniforms are expensive for the school and take money from other important programs.School unforms stifle freedom of expression.they can be costly and make circumstances difficult for those on a budget [... Hao Li 0074, Viktor Schlegel, Riza Theresa Batista-Navarro, Goran Nenadic |
ACL (1) | 4 |
| 2023 | Exploring the Value of Pre-trained Language Models for Clinical Named Entity RecognitionabstractThe practice of fine-tuning Pre-trained Language Models (PLMs) from general or domain-specific data to a specific task with limited resources, has gained popularity within the field of natural language processing (NLP). In this work, we re-visit this assumption and carry out an investigation in clinical NLP, specifically Named Entity Recognition (NER) on drugs and their related attributes. We compare Transformer models that are trained from scratch to fine-tuned BERT-based Large Language Models (LLMs) namely BERT, BioBERT, and ClinicalBERT. Furthermore, we examine the impact of an additional Conditional Random Field (CRF) layer on such models to encourage contextual learning. We use n2c2-2018 shared task data for model development and evaluations. The experimental outcomes show that 1) CRF layers improved all language models; 2) referring to BIO-strict span level evaluation using macro-average F1 score, although the fine-tuned LLMs achieved 0.83+ scores, the TransformerCRF model trained from scratch achieved 0.78+, demonstrating comparable performances with much lower cost, e.g. with 39.80% less training parameters; 3) referring to BIO-strict span-level evaluation using weighted-average F1 score, ClinicalBERT-CRF, BERT-CRF, and TransformerCRF exhibited lower score differences, with 97.59%/97.44%/96.84% respectively. 4) applying efficient training by down-sampling for better data distribution further reduced the training cost and need for data, while maintaining similar scores -i.e. around 0.02 points lower compared to using the full dataset. This This TRANSFORMERCRF project is hosted at https://github.com/HECTA-UoM/TransformerCRF Samuel Belkadi, Lifeng Han, Yuping Wu 0001, Goran Nenadic |
IEEE Big Data | 4 |
| 2023 | Extraction of Medication and Temporal Relation from Clinical Text using Neural Language ModelsabstractClinical texts, represented in electronic medical records (EMRs), contain rich medical information and are essential for disease prediction, personalised information recommendation, clinical decision support, and medication pattern mining and measurement. Relation extractions between medication mentions and temporal information can further help clinicians better understand the patients’ treatment history. To evaluate the performances of deep learning (DL) and large language models (LLMs) in medication extraction and temporal relations classification, we carry out an empirical investigation of MEDTEM project using several advanced learning structures including BiLSTM-CRF and CNN-BiLSTM for a clinical domain named entity recognition (NER), and BERT-CNN for temporal relation extraction (RE), in addition to the exploration of different word embedding techniques. Furthermore, we also designed a set of post-processing roles to generate structured output on medications and the temporal relation. Our experiments show that CNN-BiLSTM slightly wins the BiLSTM-CRF model on the i2b2-2009 clinical NER task yielding 75.67, 77.83, and 78.17 for precision, recall, and F1 scores using Macro Average. BERT-CNN model also produced reasonable evaluation scores 64.48, 67.17, and 65.03 for P/R/F1 using Macro Avg on the temporal relation extraction test set from i2b2-2012 challenges. Code and Tools from MEDTEM will be hosted at https://github.com/HECTA-UoM/MedTem Hangyu Tu, Lifeng Han, Goran Nenadic |
IEEE Big Data | 3 |
| 2023 | MC-DRE: Multi-Aspect Cross Integration for Drug Event/Entity ExtractionabstractExtracting meaningful drug-related information chunks, such as adverse drug events (ADE), is crucial for preventing morbidity and saving many lives. Most ADEs are reported via an unstructured conversation with the medical context, so applying a general entity recognition approach is not sufficient enough. In this paper, we propose a new multi-aspect cross-integration framework for drug entity/event detection by capturing and aligning different context/language/knowledge properties from drug-related documents. We first construct multi-aspect encoders to describe semantic, syntactic, and medical document contextual information by conducting those slot tagging tasks, main drug entity/event detection, part-of-speech tagging, and general medical named entity recognition. Then, each encoder conducts cross-integration with other contextual information in three ways: the key-value cross, attention cross, and feedforward cross, so the multi-encoders are integrated in depth. Our model outperforms all SOTA on two widely used tasks, flat entity detection and discontinuous event extraction. Soyeon Caren Han, Siqu Long, Josiah Poon, Goran Nenadic |
CIKM | 5 |
| 2023 | A survey of methods for revealing and overcoming weaknesses of data-driven Natural Language UnderstandingabstractAbstract Recent years have seen a growing number of publications that analyse Natural Language Understanding (NLU) datasets for superficial cues, whether they undermine the complexity of the tasks underlying those datasets and how they impact those models that are optimised and evaluated on this data. This structured survey provides an overview of the evolving research area by categorising reported weaknesses in models and datasets and the methods proposed to reveal and alleviate those weaknesses for the English language. We summarise and discuss the findings and conclude with a set of recommendations for possible future research directions. We hope that it will be a useful resource for researchers who propose new datasets to assess the suitability and quality of their data to evaluate various phenomena of interest, as well as those who propose novel NLU approaches, to further understand the implications of their improvements with respect to their model’s acquired capabilities. Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro |
Nat. Lang. Eng. | 2 |
| 2022 | Pandemic Planning using Text Analytics on Hospital Outpatient Letters: a Case Study on Covid-19 Shielding for Rheumatology Patients
Meghna Jani, Ghada Alfattni, Maksim Belousov, Michael Cheng, Andrew S. Kanter, William G. Dixon, Goran Nenadic |
AMIA | 8 |
| 2021 | Semantics Altering Modifications for Evaluating Comprehension in Machine ReadingabstractAdvances in NLP have yielded impressive results for the task of machine reading comprehension (MRC), with approaches having been reported to achieve performance comparable to that of humans. In this paper, we investigate whether state-of-the-art MRC models are able to correctly process Semantics Altering Modifications (SAM): linguistically-motivated phenomena that alter the semantics of a sentence while preserving most of its lexical surface form. We present a method to automatically generate and align challenge sets featuring original and altered examples. We further propose a novel evaluation methodology to correctly assess the capability of MRC systems to process these examples independent of the data they were optimised on, by discounting for effects introduced by domain shift. In a large-scale empirical study, we apply the methodology in order to evaluate extractive MRC models with regard to their capability to correctly process SAM-enriched data. We comprehensively cover 12 different state-of-the-art neural architecture configurations and four training datasets and find that -- despite their well-known remarkable performance -- optimised models consistently struggle to correctly process semantically altered data. Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro |
AAAI | 2 |
| 2021 | Mining a stroke knowledge graph from literatureabstractBACKGROUND: Stroke has an acute onset and a high mortality rate, making it one of the most fatal diseases worldwide. Its underlying biology and treatments have been widely studied both in the "Western" biomedicine and the Traditional Chinese Medicine (TCM). However, these two approaches are often studied and reported in insolation, both in the literature and associated databases. RESULTS: To aid research in finding effective prevention methods and treatments, we integrated knowledge from the literature and a number of databases (e.g. CID, TCMID, ETCM). We employed a suite of biomedical text mining (i.e. named-entity) approaches to identify mentions of genes, diseases, drugs, chemicals, symptoms, Chinese herbs and patent medicines, etc. in a large set of stroke papers from both biomedical and TCM domains. Then, using a combination of a rule-based approach with a pre-trained BioBERT model, we extracted and classified links and relationships among stroke-related entities as expressed in the literature. We construct StrokeKG, a knowledge graph includes almost 46 k nodes of nine types, and 157 k links of 30 types, connecting diseases, genes, symptoms, drugs, pathways, herbs, chemical, ingredients and patent medicine. CONCLUSIONS: Our Stroke-KG can provide practical and reliable stroke-related knowledge to help with stroke-related research like exploring new directions for stroke research and ideas for drug repurposing and discovery. We make StrokeKG freely available at http://114.115.208.144:7474/browser/ (Please click "Connect" directly) and the source structured data for stroke at https://github.com/yangxi1016/Stroke. Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001 |
BMC Bioinform. | 3 |
| 2021 | Correction to: Mining a stroke knowledge graph from literature
Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001 |
BMC Bioinform. | 3 |
| 2021 | Corrigendum to "Extraction of temporal relations from clinical free text: A systematic review of current approaches" [J. Biomed. Inf. 108 (2020) 103488]
Ghada Alfattni, Niels Peek, Goran Nenadic |
J. Biomed. Informatics | 3 |
| 2021 | Attention-based bidirectional long short-term memory networks for extracting temporal relationships from clinical discharge summaries
Ghada Alfattni, Niels Peek, Goran Nenadic |
J. Biomed. Informatics | 3 |
| 2020 | A Framework for Evaluation of Machine Reading Comprehension Gold StandardsabstractMachine Reading Comprehension (MRC) is the task of answering a question over a paragraph of text. While neural MRC systems gain popularity and achieve noticeable performance, issues are being raised with the methodology used to establish their performance, particularly concerning the data design of gold standards that are used to evaluate them. There is but a limited understanding of the challenges present in this data, which makes it hard to draw comparisons and formulate reliable hypotheses. As a first step towards alleviating the problem, this paper proposes a unifying framework to systematically investigate the present linguistic features, required reasoning and background knowledge and factual correctness on one hand, and the presence of lexical cues as a lower bound for the requirement of understanding on the other hand. We propose a qualitative annotation schema for the first and a set of approximative metrics for the latter. In a first application of the framework, we analyse modern MRC gold standards and present our findings: the absence of features that contribute towards lexical ambiguity, the varying factual correctness of the expected answers and the presence of lexical cues, all of which potentially lower the reading comprehension complexity and quality of the evaluation data. Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, Riza Theresa Batista-Navarro |
LREC | 4 |
| 2020 | Extraction of temporal relations from clinical free text: A systematic review of current approaches
Ghada Alfattni, Niels Peek, Goran Nenadic |
J. Biomed. Informatics | 3 |
| 2019 | Using Social Media to Study Mental Health Conditions - Challenges and Opportunities
Vasa Curcin, Elizabeth Ford, Jyotishman Pathak, Goran Nenadic |
AMIA | 4 |
| 2019 | Wind Turbine operational state prediction: towards featureless, end-to-end predictive maintenanceabstractTraditionally, predictive maintenance of wind turbines has relied on experts to perform time consuming feature pre-processing using statistical, time and frequency domain analysis. Recent advancements in Convolutional Neural Networks have opened the potential for using featureless approaches that learn the discriminating patterns from big data sets without expert intervention. Given multi-dimensional time series data representing sensed electric currents, in this paper we explore the optimal window length that can be used to accurately determine its speed and load using Convolutional and Residual Networks. Choosing an optimal window length is a trade-off between accuracy and number of predictions per second. Fast predictions for operating parameters are useful in maintenance strategies but they come at a cost of decreased accuracy (as there is less data available in a shorter time interval). We show how fusing multiple signals can achieve high accuracy with nine speed and load cases varying from 375rpm with 0% load to 1500rpm and 100% load. Using Class Activation Maps we investigate features in the time domain picked up by the network to make its classification decisions. Further, we train the networks to predict the drive loads in a regression setting where the models are able to generalize well on unseen cases. Adrian Stetco, Anees Mohammed, Sinisa Djurovic, Goran Nenadic, John A. Keane |
IEEE BigData | 4 |
| 2019 | From Web Crawled Text to Project Descriptions: Automatic Summarizing of Social Innovation Projects
Nikola Milosevic, Dimitar Marinov, Abdullah Gök, Goran Nenadic |
NLDB | 4 |
| 2019 | A framework for information extraction from tables in biomedical literatureabstractThe scientific literature is growing exponentially, and professionals are no more able to cope with the current amount of publications. Text mining provided in the past methods to retrieve and extract information from text; however, most of these approaches ignored tables and figures. The research done in mining table data still does not have an integrated approach for mining that would consider all complexities and challenges of a table. Our research is examining the methods for extracting numerical (number of patients, age, gender distribution) and textual (adverse reactions) information from tables in the clinical literature. We present a requirement analysis template and an integral methodology for information extraction from tables in clinical domain that contains 7 steps: (1) table detection, (2) functional processing, (3) structural processing, (4) semantic tagging, (5) pragmatic processing, (6) cell selection and (7) syntactic processing and extraction. Our approach performed with the F -measure ranged between 82 and 92%, depending on the variable, task and its complexity. Nikola Milosevic, Cassie Gregson, Robert Hernandez, Goran Nenadic |
Int. J. Document Anal. Recognit. | 4 |
| 2018 | Classification of Intangible Social Innovation Concepts
Nikola Milosevic, Abdullah Gök, Goran Nenadic |
NLDB | 3 |
| 2018 | Using set theory to reduce redundancy in pathway setsabstractBACKGROUND: The consolidation of pathway databases, such as KEGG, Reactome and ConsensusPathDB, has generated widespread biological interest, however the issue of pathway redundancy impedes the use of these consolidated datasets. Attempts to reduce this redundancy have focused on visualizing pathway overlap or merging pathways, but the resulting pathways may be of heterogeneous sizes and cover multiple biological functions. Efforts have also been made to deal with redundancy in pathway data by consolidating enriched pathways into a number of clusters or concepts. We present an alternative approach, which generates pathway subsets capable of covering all of genes presented within either pathway databases or enrichment results, generating substantial reductions in redundancy. RESULTS: We propose a method that uses set cover to reduce pathway redundancy, without merging pathways. The proposed approach considers three objectives: removal of pathway redundancy, controlling pathway size and coverage of the gene set. By applying set cover to the ConsensusPathDB dataset we were able to produce a reduced set of pathways, representing 100% of the genes in the original data set with 74% less redundancy, or 95% of the genes with 88% less redundancy. We also developed an algorithm to simplify enrichment data and applied it to a set of enriched osteoarthritis pathways, revealing that within the top ten pathways, five were redundant subsets of more enriched pathways. Applying set cover to the enrichment results removed these redundant pathways allowing more informative pathways to take their place. CONCLUSION: Our method provides an alternative approach for handling pathway redundancy, while ensuring that the pathways are of homogeneous size and gene coverage is maximised. Pathways are not altered from their original form, allowing biological knowledge regarding the data set to be directly applicable. We demonstrate the ability of the algorithms to prioritise redundancy reduction, pathway size control or gene set coverage. The application of set cover to pathway enrichment results produces an optimised summary of the pathways that best represent the differentially regulated gene set. Ruth Alexandra Stoney, Jean-Marc Schwartz, David L. Robertson, Goran Nenadic |
BMC Bioinform. | 4 |
| 2018 | Extracting useful software development information from mobile application reviews: A survey of intelligent mining techniques and toolsabstractMobile application (app) websites such as Google Play and AppStore allow users to review their downloaded apps. Such reviews can be useful for app users, as they may help users make an informed decision; such reviews can also be potentially useful for app developers, if they contain valuable information concerning user needs and requirements. However, in order to unleash the value of app reviews for mobile app development, intelligent mining tools that can help discern relevant reviews from irrelevant ones must be provided. This paper surveys the state of the art in the development of such tools and techniques behind them. To gain insight into the maturity of the current support mining tools, the paper will also find out what app development information these tools have discovered and what challenges they are facing. The results of this survey can inform the development of more effective and intelligent app review mining techniques and tools. Mohammadali Tavakoli, Liping Zhao 0001, Atefeh Heydari, Goran Nenadic |
Expert Syst. Appl. | 4 |
| 2018 | Data and systems for medication-related text classification and concept normalization from Twitter: insights from the Social Media Mining for Health (SMM4H)-2017 shared taskabstractObjective: We executed the Social Media Mining for Health (SMM4H) 2017 shared tasks to enable the community-driven development and large-scale evaluation of automatic text processing methods for the classification and normalization of health-related text from social media. An additional objective was to publicly release manually annotated data. Materials and Methods: We organized 3 independent subtasks: automatic classification of self-reports of 1) adverse drug reactions (ADRs) and 2) medication consumption, from medication-mentioning tweets, and 3) normalization of ADR expressions. Training data consisted of 15 717 annotated tweets for (1), 10 260 for (2), and 6650 ADR phrases and identifiers for (3); and exhibited typical properties of social-media-based health-related texts. Systems were evaluated using 9961, 7513, and 2500 instances for the 3 subtasks, respectively. We evaluated performances of classes of methods and ensembles of system combinations following the shared tasks. Results: Among 55 system runs, the best system scores for the 3 subtasks were 0.435 (ADR class F1-score) for subtask-1, 0.693 (micro-averaged F1-score over two classes) for subtask-2, and 88.5% (accuracy) for subtask-3. Ensembles of system combinations obtained best scores of 0.476, 0.702, and 88.7%, outperforming individual systems. Discussion: Among individual systems, support vector machines and convolutional neural networks showed high performance. Performance gains achieved by ensembles of system combinations suggest that such strategies may be suitable for operational systems relying on difficult text classification tasks (eg, subtask-1). Conclusions: Data imbalance and lack of context remain challenges for natural language processing of social media text. Annotated data from the shared task have been made available as reference standards for future studies (http://dx.doi.org/10.17632/rxwfb3tysd.1). Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kiritchenko, Farrokh Mehryary, Sifei Han, Tung Tran 0001, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, Graciela Gonzalez-Hernandez |
J. Am. Medical Informatics Assoc. | 15 |
| 2016 | Identification of Occupation Mentions in Clinical Narratives
Azad Dehghan, Tom Liptrot, Daniel Tibble, Matthew Barker-Hewitt, Goran Nenadic |
NLDB | 5 |
| 2016 | Disentangling the Structure of Tables in Scientific Literature
Nikola Milosevic, Cassie Gregson, Robert Hernandez, Goran Nenadic |
NLDB | 4 |
| 2016 | Inferring Methodological Meta-knowledge from Large Biomedical Corpora
Goran Nenadic |
PACLIC | 1 |
| 2015 | Contextualisation of Biomedical Knowledge Through Large-Scale Processing of Literature, Clinical Narratives and Social Media
Goran Nenadic |
AIME | 1 |
| 2015 | Temporal expression extraction with extensive feature type selection and a posteriori label adjustmentabstractThe automatic extraction of temporal information from written texts is pivotal for many Natural Language Processing applications such as question answering, text summarisation and information retrieval. It allows to filter information and infer temporal flows of events. This paper presents ManTIME, a general domain temporal expression identification and normalisation system, and systematically explores the impact of different features and training corpora on the performance. The identification phase combines the use of conditional random fields along with a post-processing pipeline, whereas the normalisation phase is carried out using NorMA, an open-source rule-based temporal normaliser. We investigate the performance variation with respect to different feature types. Specifically, we show that the use of WordNet-based features in the identification task negatively affects the overall performance, and that there is no statistically significant difference in the results based on gazetteers, shallow parsing and propositional noun phrases labels on top of the morpho-lexical features. We also show that the use of silver data (alone or in addition to the human-annotated ones) does not improve the performance. We evaluate six combinations of training data and post-processing pipeline with respect to the TempEval-3 benchmark test set. The best run achieved 0.95 (precision), 0.85 (recall) and 0.90 (Fβ=1) in the identification phase. Normalisation accuracies are 0.86 (for type attribute) and 0.77 (for value attribute). The proposed approach ranked 3rd in the TempEval-3 challenge (task A) as the best performing machine learning-based system among 21 participants. Michele Filannino, Goran Nenadic |
Data Knowl. Eng. | 2 |
| 2015 | Combining knowledge- and data-driven methods for de-identification of clinical narrativesabstractA recent promise to access unstructured clinical data from electronic health records on large-scale has revitalized the interest in automated de-identification of clinical notes, which includes the identification of mentions of Protected Health Information (PHI). We describe the methods developed and evaluated as part of the i2b2/UTHealth 2014 challenge to identify PHI defined by 25 entity types in longitudinal clinical narratives. Our approach combines knowledge-driven (dictionaries and rules) and data-driven (machine learning) methods with a large range of features to address de-identification of specific named entities. In addition, we have devised a two-pass recognition approach that creates a patient-specific run-time dictionary from the PHI entities identified in the first step with high confidence, which is then used in the second pass to identify mentions that lack specific clues. The proposed method achieved the overall micro F1-measures of 91% on strict and 95% on token-level evaluation on the test dataset (514 narratives). Whilst most PHI entities can be reliably identified, particularly challenging were mentions of Organizations and Professions. Still, the overall results suggest that automated text mining methods can be used to reliably process clinical notes to identify personal information and thus providing a crucial step in large-scale de-identification of unstructured data for further clinical and epidemiological studies. Azad Dehghan, Aleksandar Kovacevic, George Karystianis, John A. Keane, Goran Nenadic |
J. Biomed. Informatics | 5 |
| 2015 | Using local lexicalized rules to identify heart disease risk factors in clinical notesabstractHeart disease is the leading cause of death globally and a significant part of the human population lives with it. A number of risk factors have been recognized as contributing to the disease, including obesity, coronary artery disease (CAD), hypertension, hyperlipidemia, diabetes, smoking, and family history of premature CAD. This paper describes and evaluates a methodology to extract mentions of such risk factors from diabetic clinical notes, which was a task of the i2b2/UTHealth 2014 Challenge in Natural Language Processing for Clinical Data. The methodology is knowledge-driven and the system implements local lexicalized rules (based on syntactical patterns observed in notes) combined with manually constructed dictionaries that characterize the domain. A part of the task was also to detect the time interval in which the risk factors were present in a patient. The system was applied to an evaluation set of 514 unseen notes and achieved a micro-average F-score of 88% (with 86% precision and 90% recall). While the identification of CAD family history, medication and some of the related disease factors (e.g. hypertension, diabetes, hyperlipidemia) showed quite good results, the identification of CAD-specific indicators proved to be more challenging (F-score of 74%). Overall, the results are encouraging and suggested that automated text mining methods can be used to process clinical notes to identify risk factors and monitor progression of heart disease on a large-scale, providing necessary data for clinical and epidemiological studies. George Karystianis, Azad Dehghan, Aleksandar Kovacevic, John A. Keane, Goran Nenadic |
J. Biomed. Informatics | 5 |
| 2014 | Identification of Multi-Focal Questions in Question and Answer Reports
Mona Mohamed Zaki Ali, Goran Nenadic, Babis Theodoulidis |
NLDB | 2 |
| 2014 | Extracting patterns of database and software usage from the bioinformatics literatureabstractMOTIVATION: As a natural consequence of being a computer-based discipline, bioinformatics has a strong focus on database and software development, but the volume and variety of resources are growing at unprecedented rates. An audit of database and software usage patterns could help provide an overview of developments in bioinformatics and community common practice, and comparing the links between resources through time could demonstrate both the persistence of existing software and the emergence of new tools. RESULTS: We study the connections between bioinformatics resources and construct networks of database and software usage patterns, based on resource co-occurrence, that correspond to snapshots of common practice in the bioinformatics community. We apply our approach to pairings of phylogenetics software reported in the literature and argue that these could provide a stepping stone into the identification of scientific best practice. AVAILABILITY AND IMPLEMENTATION: The extracted resource data, the scripts used for network generation and the resulting networks are available at http://bionerds.sourceforge.net/networks/. Geraint Duck, Goran Nenadic, Andy Brass, David L. Robertson, Robert Stevens 0001 |
Bioinform. | 2 |
| 2013 | Challenges in Clinical Named Entity Recognition for Decision SupportabstractIn addition to structured data, electronic health records contain unstructured clinical notes and narratives. The identification and classification of mentions of relevant clinical concepts is a crucial preprocessing step in designing and developing clinical decision support systems. While this task has gained significant attention in recent years, there are still a number of issues that need further investigation. This paper explores a variety of common challenges faced by clinical named entity recognition and classification methods as well as current approaches to handling them. Azad Dehghan, John A. Keane, Goran Nenadic |
SMC | 3 |
| 2013 | A Bayesian Association Rule Mining AlgorithmabstractThis paper proposes a Bayesian association rule mining algorithm (BAR) which combines the Apriori association rule mining algorithm with Bayesian networks. Two interesting-ness measures of association rules: Bayesian confidence (BC) and Bayesian lift (BL) which measure conditional dependence and independence relationships between items are defined based on the joint probabilities represented by the Bayesian networks of association rules. BAR outputs best rules according to BC and BL. BAR is evaluated for its performance using two anonymized clinical phenotype datasets from the UCI Repository: Thyroid disease and Diabetes. The results show that BAR is capable of finding the best rules which have the highest BC, BL and very high support, confidence and lift. David Tian, Ann Gledson, Athos Antoniades, Aristos Aristodimou, Dimitrios Ntalaperas, Ratnesh Sahay, Jianxin Pan, Stavros Stivaros, Goran Nenadic, Xiaojun Zeng, John A. Keane |
SMC | 9 |
| 2013 | bioNerDS: exploring bioinformatics' database and software use through literature miningabstractBACKGROUND: Biology-focused databases and software define bioinformatics and their use is central to computational biology. In such a complex and dynamic field, it is of interest to understand what resources are available, which are used, how much they are used, and for what they are used. While scholarly literature surveys can provide some insights, large-scale computer-based approaches to identify mentions of bioinformatics databases and software from primary literature would automate systematic cataloguing, facilitate the monitoring of usage, and provide the foundations for the recovery of computational methods for analysing biological data, with the long-term aim of identifying best/common practice in different areas of biology. RESULTS: We have developed bioNerDS, a named entity recogniser for the recovery of bioinformatics databases and software from primary literature. We identify such entities with an F-measure ranging from 63% to 91% at the mention level and 63-78% at the document level, depending on corpus. Not attaining a higher F-measure is mostly due to high ambiguity in resource naming, which is compounded by the on-going introduction of new resources. To demonstrate the software, we applied bioNerDS to full-text articles from BMC Bioinformatics and Genome Biology. General mention patterns reflect the remit of these journals, highlighting BMC Bioinformatics's emphasis on new tools and Genome Biology's greater emphasis on data analysis. The data also illustrates some shifts in resource usage: for example, the past decade has seen R and the Gene Ontology join BLAST and GenBank as the main components in bioinformatics processing. ABSTRACT: Conclusions We demonstrate the feasibility of automatically identifying resource names on a large-scale from the scientific literature and show that the generated data can be used for exploration of bioinformatics database and software usage. For example, our results help to investigate the rate of change in resource usage and corroborate the suspicion that a vast majority of resources are created, but rarely (if ever) used thereafter. bioNerDS is available at http://bionerds.sourceforge.net/. Geraint Duck, Goran Nenadic, Andy Brass, David L. Robertson, Robert Stevens 0001 |
BMC Bioinform. | 2 |
| 2013 | Combining rules and machine learning for extraction of temporal expressions and events from clinical narrativesabstractOBJECTIVE: Identification of clinical events (eg, problems, tests, treatments) and associated temporal expressions (eg, dates and times) are key tasks in extracting and managing data from electronic health records. As part of the i2b2 2012 Natural Language Processing for Clinical Data challenge, we developed and evaluated a system to automatically extract temporal expressions and events from clinical narratives. The extracted temporal expressions were additionally normalized by assigning type, value, and modifier. MATERIALS AND METHODS: The system combines rule-based and machine learning approaches that rely on morphological, lexical, syntactic, semantic, and domain-specific features. Rule-based components were designed to handle the recognition and normalization of temporal expressions, while conditional random fields models were trained for event and temporal recognition. RESULTS: The system achieved micro F scores of 90% for the extraction of temporal expressions and 87% for clinical event extraction. The normalization component for temporal expressions achieved accuracies of 84.73% (expression's type), 70.44% (value), and 82.75% (modifier). DISCUSSION: Compared to the initial agreement between human annotators (87-89%), the system provided comparable performance for both event and temporal expression mining. While (lenient) identification of such mentions is achievable, finding the exact boundaries proved challenging. CONCLUSIONS: The system provides a state-of-the-art method that can be used to support automated identification of mentions of clinical events and temporal expressions in narratives either to support the manual review process or as a part of a large-scale processing of electronic health databases. Aleksandar Kovacevic, Azad Dehghan, Michele Filannino, John A. Keane, Goran Nenadic |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | BioContext: an integrated text mining system for large-scale extraction and contextualization of biomolecular eventsabstractMOTIVATION: Although the amount of data in biology is rapidly increasing, critical information for understanding biological events like phosphorylation or gene expression remains locked in the biomedical literature. Most current text mining (TM) approaches to extract information about biological events are focused on either limited-scale studies and/or abstracts, with data extracted lacking context and rarely available to support further research. RESULTS: Here we present BioContext, an integrated TM system which extracts, extends and integrates results from a number of tools performing entity recognition, biomolecular event extraction and contextualization. Application of our system to 10.9 million MEDLINE abstracts and 234 000 open-access full-text articles from PubMed Central yielded over 36 million mentions representing 11.4 million distinct events. Event participants included over 290 000 distinct genes/proteins that are mentioned more than 80 million times and linked where possible to Entrez Gene identifiers. Over a third of events contain contextual information such as the anatomical location of the event occurrence or whether the event is reported as negated or speculative. AVAILABILITY: The BioContext pipeline is available for download (under the BSD license) at http://www.biocontext.org, along with the extracted data which is also available for online browsing. Martin Gerner, Farzaneh Sarafraz, Casey M. Bergman, Goran Nenadic |
Bioinform. | 4 |
| 2012 | Mining methodologies from NLP publications: A case study in automatic terminology recognition
Aleksandar Kovacevic, Zora Konjovic, Branko Milosavljevic, Goran Nenadic |
Comput. Speech Lang. | 4 |
| 2011 | The GNAT library for local and remote gene mention normalizationabstractSUMMARY: Identifying mentions of named entities, such as genes or diseases, and normalizing them to database identifiers have become an important step in many text and data mining pipelines. Despite this need, very few entity normalization systems are publicly available as source code or web services for biomedical text mining. Here we present the Gnat Java library for text retrieval, named entity recognition, and normalization of gene and protein mentions in biomedical text. The library can be used as a component to be integrated with other text-mining systems, as a framework to add user-specific extensions, and as an efficient stand-alone application for the identification of gene and protein names for data analysis. On the BioCreative III test data, the current version of Gnat achieves a Tap-20 score of 0.1987. AVAILABILITY: The library and web services are implemented in Java and the sources are available from http://gnat.sourceforge.net. CONTACT: [email protected]. Jörg Hakenberg, Martin Gerner, Maximilian Haeussler, Illés Solt, Conrad Plake, Michael Schroeder 0001, Graciela Gonzalez-Hernandez, Goran Nenadic, Casey M. Bergman |
Bioinform. | 8 |
| 2010 | LINNAEUS: A species name identification system for biomedical literatureabstractBACKGROUND: The task of recognizing and identifying species names in biomedical literature has recently been regarded as critical for a number of applications in text and data mining, including gene name recognition, species-specific document retrieval, and semantic enrichment of biomedical articles. RESULTS: In this paper we describe an open-source species name recognition and normalization software system, LINNAEUS, and evaluate its performance relative to several automatically generated biomedical corpora, as well as a novel corpus of full-text documents manually annotated for species mentions. LINNAEUS uses a dictionary-based approach (implemented as an efficient deterministic finite-state automaton) to identify species names and a set of heuristics to resolve ambiguous mentions. When compared against our manually annotated corpus, LINNAEUS performs with 94% recall and 97% precision at the mention level, and 98% recall and 90% precision at the document level. Our system successfully solves the problem of disambiguating uncertain species mentions, with 97% of all mentions in PubMed Central full-text documents resolved to unambiguous NCBI taxonomy identifiers. CONCLUSIONS: LINNAEUS is an open source, stand-alone software system capable of recognizing and normalizing species name mentions with speed and accuracy, and can therefore be integrated into a range of bioinformatics and text-mining applications. The software and manually annotated corpus can be downloaded freely at http://linnaeus.sourceforge.net/. Martin Gerner, Goran Nenadic, Casey M. Bergman |
BMC Bioinform. | 2 |
| 2010 | Medication information extraction with linguistic pattern matching and semantic rulesabstractOBJECTIVE: This study presents a system developed for the 2009 i2b2 Challenge in Natural Language Processing for Clinical Data, whose aim was to automatically extract certain information about medications used by a patient from his/her medical report. The aim was to extract the following information for each medication: name, dosage, mode/route, frequency, duration and reason. DESIGN: The system implements a rule-based methodology, which exploits typical morphological, lexical, syntactic and semantic features of the targeted information. These features were acquired from the training dataset and public resources such as the UMLS and relevant web pages. Information extracted by pattern matching was combined together using context-sensitive heuristic rules. MEASUREMENTS: The system was applied to a set of 547 previously unseen discharge summaries, and the extracted information was evaluated against a manually prepared gold standard consisting of 251 documents. The overall ranking of the participating teams was obtained using the micro-averaged F-measure as the primary evaluation metric. RESULTS: The implemented method achieved the micro-averaged F-measure of 81% (with 86% precision and 77% recall), which ranked this system third in the challenge. The significance tests revealed the system's performance to be not significantly different from that of the second ranked system. Relative to other systems, this system achieved the best F-measure for the extraction of duration (53%) and reason (46%). CONCLUSION: Based on the F-measure, the performance achieved (81%) was in line with the initial agreement between human annotators (82%), indicating that such a system may greatly facilitate the process of extracting relevant information from medical records by providing a solid basis for a manual review process. Irena Spasic, Farzaneh Sarafraz, John A. Keane, Goran Nenadic |
J. Am. Medical Informatics Assoc. | 4 |
| 2009 | Mining Semantic Descriptions of Bioinformatics Web Resources from the Literature
Hammad Afzal, Robert Stevens 0001, Goran Nenadic |
ESWC | 3 |
| 2009 | Research Paper: A Text Mining Approach to the Prediction of Disease Status from Clinical Discharge SummariesabstractOBJECTIVE The authors present a system developed for the Challenge in Natural Language Processing for Clinical Data-the i2b2 obesity challenge, whose aim was to automatically identify the status of obesity and 15 related co-morbidities in patients using their clinical discharge summaries. The challenge consisted of two tasks, textual and intuitive. The textual task was to identify explicit references to the diseases, whereas the intuitive task focused on the prediction of the disease status when the evidence was not explicitly asserted. DESIGN The authors assembled a set of resources to lexically and semantically profile the diseases and their associated symptoms, treatments, etc. These features were explored in a hybrid text mining approach, which combined dictionary look-up, rule-based, and machine-learning methods. MEASUREMENTS The methods were applied on a set of 507 previously unseen discharge summaries, and the predictions were evaluated against a manually prepared gold standard. The overall ranking of the participating teams was primarily based on the macro-averaged F-measure. RESULTS The implemented method achieved the macro-averaged F-measure of 81% for the textual task (which was the highest achieved in the challenge) and 63% for the intuitive task (ranked 7(th) out of 28 teams-the highest was 66%). The micro-averaged F-measure showed an average accuracy of 97% for textual and 96% for intuitive annotations. CONCLUSIONS The performance achieved was in line with the agreement between human annotators, indicating the potential of text mining for accurate and efficient prediction of disease statuses from clinical discharge summaries. Hui Yang 0004, Irena Spasic, John A. Keane, Goran Nenadic |
J. Am. Medical Informatics Assoc. | 4 |
| 2009 | Assigning roles to protein mentions: The case of transcription factors
Hui Yang 0004, John A. Keane, Casey M. Bergman, Goran Nenadic |
J. Biomed. Informatics | 4 |
| 2008 | Identification of transcription factor contexts in literature using machine learning approachesabstractBACKGROUND: Availability of information about transcription factors (TFs) is crucial for genome biology, as TFs play a central role in the regulation of gene expression. While manual literature curation is expensive and labour intensive, the development of semi-automated text mining support is hindered by unavailability of training data. There have been no studies on how existing data sources (e.g. TF-related data from the MeSH thesaurus and GO ontology) or potentially noisy example data (e.g. protein-protein interaction, PPI) could be used to provide training data for identification of TF-contexts in literature. RESULTS: In this paper we describe a text-classification system designed to automatically recognise contexts related to transcription factors in literature. A learning model is based on a set of biological features (e.g. protein and gene names, interaction words, other biological terms) that are deemed relevant for the task. We have exploited background knowledge from existing biological resources (MeSH and GO) to engineer such features. Weak and noisy training datasets have been collected from descriptions of TF-related concepts in MeSH and GO, PPI data and data representing non-protein-function descriptions. Three machine-learning methods are investigated, along with a vote-based merging of individual approaches and/or different training datasets. The system achieved highly encouraging results, with most classifiers achieving an F-measure above 90%. CONCLUSIONS: The experimental results have shown that the proposed model can be used for identification of TF-related contexts (i.e. sentences) with high accuracy, with a significantly reduced set of features when compared to traditional bag-of-words approach. The results of considering existing PPI data suggest that there is not as high similarity between TF and PPI contexts as we have expected. We have also shown that existing knowledge sources are useful both for feature engineering and for obtaining noisy positive training data. Hui Yang 0004, Goran Nenadic, John A. Keane |
BMC Bioinform. | 2 |
| 2006 | Towards a terminological resource for biomedical text mining
Goran Nenadic, Naoaki Okazaki, Sophia Ananiadou |
LREC | 1 |
| 2006 | Mining semantically related terms from biomedical literatureabstractDiscovering links and relationships is one of the main challenges in biomedical research, as scientists are interested in uncovering entities that have similar functions, take part in the same processes, or are coregulated. This article discusses the extraction of such semantically related entities (represented by domain terms) from biomedical literature. The method combines various text-based aspects, such as lexical, syntactic, and contextual similarities between terms. Lexical similarities are based on the level of sharing of word constituents. Syntactic similarities rely on expressions (such as term enumerations and conjunctions) in which a sequence of terms appears as a single syntactic unit. Finally, contextual similarities are based on automatic discovery of relevant contexts shared among terms. The approach is evaluated using the Genia resources, and the results of experiments are presented. Lexical and syntactic links have shown high precision and low recall, while contextual similarities have resulted in significantly higher recall with moderate precision. By combining the three metrics, we achieved F measures of 68% for semantically related terms and 37% for highly related entities. Goran Nenadic, Sophia Ananiadou |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2005 | Mining protein function from text using term-based support vector machinesabstractBACKGROUND: Text mining has spurred huge interest in the domain of biology. The goal of the BioCreAtIvE exercise was to evaluate the performance of current text mining systems. We participated in Task 2, which addressed assigning Gene Ontology terms to human proteins and selecting relevant evidence from full-text documents. We approached it as a modified form of the document classification task. We used a supervised machine-learning approach (based on support vector machines) to assign protein function and select passages that support the assignments. As classification features, we used a protein's co-occurring terms that were automatically extracted from documents. RESULTS: The results evaluated by curators were modest, and quite variable for different problems: in many cases we have relatively good assignment of GO terms to proteins, but the selected supporting text was typically non-relevant (precision spanning from 3% to 50%). The method appears to work best when a substantial set of relevant documents is obtained, while it works poorly on single documents and/or short passages. The initial results suggest that our approach can also mine annotations from text even when an explicit statement relating a protein to a GO term is absent. CONCLUSION: A machine learning approach to mining protein function predictions from text can yield good performance only if sufficient training data is available, and significant amount of supporting data is used for prediction. The most promising results are for combined document retrieval and GO term assignment, which calls for the integration of methods developed in BioCreAtIvE Task 1 and Task 2. Simon B. Rice, Goran Nenadic, Benjamin J. Stapley |
BMC Bioinform. | 2 |
| 2004 | Enhancing automatic term recognition through recognition of variation
Goran Nenadic, Sophia Ananiadou, John McNaught |
COLING | 1 |
| 2004 | Learning to Classify Biomedical Terms Through Literature Mining and Genetic Algorithms
Irena Spasic, Goran Nenadic, Sophia Ananiadou |
IDEAL | 2 |
| 2004 | Mining Biomedical Abstracts: What's in a Term?
Goran Nenadic, Irena Spasic, Sophia Ananiadou |
IJCNLP | 1 |
| 2004 | Exploring Balkanet Shared Ontology for Multilingual Conceptual Indexing
Sofia Stamou, Goran Nenadic, Dimitris Christodoulakis |
LREC | 2 |
| 2004 | Term identification in the biomedical literature
Michael Krauthammer, Goran Nenadic |
J. Biomed. Informatics | 2 |
| 2003 | An Integrated Term-Based Corpus Query System
Kostas Manios, Goran Nenadic, Irena Spasic, Sophia Ananiadou |
EACL | 2 |
| 2003 | Terminology-driven Mining of Biomedical LiteratureabstractAbstract Motivation: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective literature mining techniques that can help biologists to gather and make use of the knowledge encoded in text documents. Although the knowledge is organized around sets of domain-specific terms, few literature mining systems incorporate deep and dynamic terminology processing. Results: In this paper, we present an overview of an integrated framework for terminology-driven mining from biomedical literature. The framework integrates the following components: automatic term recognition, term variation handling, acronym acquisition, automatic discovery of term similarities and term clustering. The term variant recognition is incorporated into terminology recognition process by taking into account orthographical, morphological, syntactic, lexico-semantic and pragmatic term variations. In particular, we address acronyms as a common way of introducing term variants in biomedical papers. Term clustering is based on the automatic discovery of term similarities. We use a hybrid similarity measure, where terms are compared by using both internal and external evidence. The measure combines lexical, syntactical and contextual similarity. Experiments on terminology recognition and clustering performed on a corpus of Medline abstracts recorded the precision of 98 and 71% respectively. Availability: Software for the terminology management is available upon request. Contact: [email protected] * To whom correspondence should be addressed at: Department of Computation/Department of BioMolecular Sciences, UMIST, Manchester, UK. Goran Nenadic, Irena Spasic, Sophia Ananiadou |
Bioinform. | 1 |
| 2002 | A Methodology for Terminology-based Knowledge Acquisition and Integration
Hideki Mima, Sophia Ananiadou, Goran Nenadic, Jun'ichi Tsujii |
COLING | 3 |
| 2002 | Supervised Learning of Term Similarities
Irena Spasic, Goran Nenadic, Kostas Manios, Sophia Ananiadou |
IDEAL | 2 |
| 2002 | Automatic Acronym Acquisition and Term Variation Management within Domain-Specific Texts
Goran Nenadic, Irena Spasic, Sophia Ananiadou |
LREC | 1 |
| 2002 | Tuning Context Features with Genetic Algorithms
Irena Spasic, Goran Nenadic, Sophia Ananiadou |
LREC | 2 |