EDBT 2026 Demo / reviewers in the wild / expert
Nada Lavrac
dblp:l/NadaLavrac
· DBLP profile ↗
146ranked-venue papers
23as first author
11since 2021 · last 2025
0000-0002-9995-7093ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 15 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 58 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 30 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6 · 2 first-authorTheory of computation · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 4Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Make Literature-Based Discovery Great Again Through Reproducible Pipelines
Bojan Cestnik, Andrej Kastrin, Boshko Koloski, Nada Lavrac |
IDA | 4 |
| 2025 | HorNets: learning from discrete and continuous signals with routing neural networksabstractAbstract Construction of neural network architectures suitable for learning from both continuous and discrete tabular data is challenging, as contemporary high-dimensional tabular data sets are often characterized by a relatively small set of instances and the request for efficient learning. We propose HorNets (Horn Networks), a neural network architecture with state-of-the-art performance on synthetic and real-life data sets from scarce-data tabular domains. HorNets are based on a clipped polynomial-like activation function, extended by a custom discrete-continuous routing mechanism that decides which part of the neural network to optimize based on the input’s cardinality. By explicitly modeling parts of the feature combination space or combining whole space in a linear attention-like manner, HorNets dynamically decide which mode of operation is the most suitable for a given piece of data with no explicit supervision. This architecture is one of the few approaches that reliably retrieves logical clauses (including noisy XNOR) and achieves state-of-the-art classification performance on 14 real-life biomedical high-dimensional data sets. HorNets are made freely available under a permissive license alongside a synthetic generator of categorical benchmarks. Boshko Koloski, Nada Lavrac, Blaz Skrlj |
Mach. Learn. | 2 |
| 2024 | AHAM: Adapt, Help, Ask, Model Harvesting LLMs for Literature Mining
Boshko Koloski, Nada Lavrac, Bojan Cestnik, Senja Pollak, Blaz Skrlj, Andrej Kastrin |
IDA (1) | 2 |
| 2022 | Deep node ranking for neuro-symbolic structural node embedding and classificationabstractNetwork node embedding is an active research subfield of complex network analysis. This paper contributes a novel approach to learning network node embeddings and direct node classification using a node ranking scheme coupled with an autoencoder-based neural network architecture. The main advantages of the proposed Deep Node Ranking (DNR) algorithm are competitive or better classification performance, significantly higher learning speed and lower space requirements when compared to state-of-the-art approaches on 15 real-life node classification benchmarks. Furthermore, it enables exploration of the relationship between symbolic and the derived sub-symbolic node representations, offering insights into the learned node space structure. To avoid the space complexity bottleneck in a direct node classification setting, DNR computes stationary distributions of personalized random walks from given nodes in mini-batches, scaling seamlessly to larger networks. The scaling laws associated with DNR were also investigated on 1488 synthetic Erd\H{o}s-R\'enyi networks, demonstrating its scalability to tens of millions of links. Blaz Skrlj, Jan Kralj, Janez Konc, Marko Robnik-Sikonja, Nada Lavrac |
Int. J. Intell. Syst. | 5 |
| 2022 | ReliefE: feature ranking in high-dimensional spaces via manifold embeddingsabstractAbstract Feature ranking has been widely adopted in machine learning applications such as high-throughput biology and social sciences. The approaches of the popular Relief family of algorithms assign importances to features by iteratively accounting for nearest relevant and irrelevant instances. Despite their high utility, these algorithms can be computationally expensive and not-well suited for high-dimensional sparse input spaces. In contrast, recent embedding-based methods learn compact, low-dimensional representations, potentially facilitating down-stream learning capabilities of conventional learners. This paper explores how the Relief branch of algorithms can be adapted to benefit from (Riemannian) manifold-based embeddings of instance and target spaces, where a given embedding’s dimensionality is intrinsic to the dimensionality of the considered data set. The developed ReliefE algorithm is faster and can result in better feature rankings, as shown by our evaluation on 20 real-life data sets for multi-class and multi-label classification tasks. The utility of ReliefE for high-dimensional data sets is ensured by its implementation that utilizes sparse matrix algebraic operations. Finally, the relation of ReliefE to other ranking algorithms is studied via the Fuzzy Jaccard Index. Blaz Skrlj, Saso Dzeroski, Nada Lavrac, Matej Petkovic |
Mach. Learn. | 3 |
| 2021 | Stratification of Parkinson's Disease Patients via Multi-view Clustering
Anita Valmarska, Nada Lavrac, Marko Robnik-Sikonja |
AIME | 2 |
| 2021 | Prioritization of COVID-19-Related Literature via Unsupervised Keyphrase Extraction and Document Representation Learning
Blaz Skrlj, Marko Jukic, Nika Erzen, Senja Pollak, Nada Lavrac |
DS | 5 |
| 2021 | Scientific Question Generation: Pattern-Based and Graph-Based RoboCHAIR Methods
Senja Pollak, Vid Podpecan, Janez Kranjc, Borut Lesjak, Nada Lavrac |
ICCC | 5 |
| 2021 | CaNDis: a web server for investigation of causal relationships between diseases, drugs and drug targetsabstractMOTIVATION: Causal biological interaction networks represent cellular regulatory pathways. Their fusion with other biological data enables insights into disease mechanisms and novel opportunities for drug discovery. RESULTS: We developed Causal Network of Diseases (CaNDis), a web server for the exploration of a human causal interaction network, which we expanded with data on diseases and FDA-approved drugs, on the basis of which we constructed a disease-disease network in which the links represent the similarity between diseases. We show how CaNDis can be used to identify candidate genes with known and novel roles in disease co-occurrence and drug-drug interactions. AVAILABILITYAND IMPLEMENTATION: CaNDis is freely available to academic users at http://candis.ijs.si and http://candis.insilab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Blaz Skrlj, Nika Erzen, Nada Lavrac, Tanja Kunej, Janez Konc |
Bioinform. | 3 |
| 2021 | tax2vec: Constructing Interpretable Features from Taxonomies for Short Text ClassificationabstractThe use of background knowledge is largely unexploited in text classification tasks. This paper explores word taxonomies as means for constructing new semantic features, which may improve the performance and robustness of the learned classifiers. We propose tax2vec, a parallel algorithm for constructing taxonomy-based features, and demonstrate its use on six short text classification problems: prediction of gender, personality type, age, news topics, drug side effects and drug effectiveness. The constructed semantic features, in combination with fast linear classifiers, tested against strong baselines such as hierarchical attention neural networks, achieves comparable classification results on short text documents. The algorithm's performance is also tested in a few-shot learning setting, indicating that the inclusion of semantic features can improve the performance in data-scarce situations. The tax2vec capability to extract corpus-specific semantic keywords is also demonstrated. Finally, we investigate the semantic space of potential features, where we observe a similarity with the well known Zipf's law. Blaz Skrlj, Matej Martinc, Jan Kralj, Nada Lavrac, Senja Pollak |
Comput. Speech Lang. | 4 |
| 2021 | autoBOT: evolving neuro-symbolic representations for explainable low resource text classificationabstractLearning from texts has been widely adopted throughout industry and science. While state-of-the-art neural language models have shown very promising results for text classification, they are expensive to (pre-)train, require large amounts of data and tuning of hundreds of millions or more parameters. This paper explores how automatically evolved text representations can serve as a basis for explainable, low-resource branch of models with competitive performance that are subject to automated hyperparameter tuning. We present autoBOT (automatic Bags-Of-Tokens), an autoML approach suitable for low resource learning scenarios, where both the hardware and the amount of data required for training are limited. The proposed approach consists of an evolutionary algorithm that jointly optimizes various sparse representations of a given text (including word, subword, POS tag, keyword-based, knowledge graph-based and relational features) and two types of document embeddings (non-sparse representations). The key idea of autoBOT is that, instead of evolving at the learner level, evolution is conducted at the representation level. The proposed method offers competitive classification performance on fourteen real-world classification tasks when compared against a competitive autoML approach that evolves ensemble models, as well as state-of-the-art neural language models such as BERT and RoBERTa. Moreover, the approach is explainable, as the importance of the parts of the input space is part of the final solution yielded by the proposed optimization procedure, offering potential for meta-transfer learning. Blaz Skrlj, Matej Martinc, Nada Lavrac, Senja Pollak |
Mach. Learn. | 3 |
| 2020 | Multi-view Clustering with mvReliefF for Parkinson's Disease Patients Subgroup Detection
Anita Valmarska, Dragana Miljkovic, Nada Lavrac, Marko Robnik-Sikonja |
AIME | 3 |
| 2020 | COVID-19 Therapy Target Discovery with Context-Aware Literature Mining
Matej Martinc, Blaz Skrlj, Sergej Pirkmajer, Nada Lavrac, Bojan Cestnik, Martin Marzidovsek, Senja Pollak |
DS | 4 |
| 2020 | Feature Importance Estimation with Self-Attention NetworksabstractBlack-box neural network models are widely used in industry and science, yet are hard to understand and interpret. Recently, the attention mechanism was introduced, offering insights into the inner workings of neural language models. This paper explores the use of attention-based neural networks mechanism for estimating feature importance, as means for explaining the models learned from propositional (tabular) data. Feature importance estimates, assessed by the proposed Self-Attention Network (SAN) architecture, are compared with the established ReliefF, Mutual Information and Random Forest-based estimates, which are widely used in practice for model interpretation. For the first time we conduct scale-free comparisons of feature importance estimates across algorithms on ten real and synthetic data sets to study the similarities and differences of the resulting feature importance estimates, showing that SANs identify similar high-ranked features as the other methods. We demonstrate that SANs identify feature interactions which in some cases yield better predictive performance than the baselines, suggesting that attention extends beyond interactions of just a few key features and detects larger feature subsets relevant for the considered learning task. Blaz Skrlj, Saso Dzeroski, Nada Lavrac, Matej Petkovic |
ECAI | 3 |
| 2020 | Bisociative Literature-Based Discovery: Lessons Learned and New Prospects
Nada Lavrac, Matej Martinc, Senja Pollak, Bojan Cestnik |
ICCC | 1 |
| 2020 | Propositionalization and embeddings: two sides of the same coinabstractData preprocessing is an important component of machine learning pipelines, which requires ample time and resources. An integral part of preprocessing is data transformation into the format required by a given learning algorithm. This paper outlines some of the modern data processing techniques used in relational learning that enable data fusion from different input data types and formats into a single table data representation, focusing on the propositionalization and embedding data transformation approaches. While both approaches aim at transforming data into tabular data format, they use different terminology and task definitions, are perceived to address different goals, and are used in different contexts. This paper contributes a unifying framework that allows for improved understanding of these two data transformation techniques by presenting their unified definitions, and by explaining the similarities and differences between the two approaches as variants of a unified complex data transformation task. In addition to the unifying framework, the novelty of this paper is a unifying methodology combining propositionalization and embeddings, which benefits from the advantages of both in solving complex data transformation and learning tasks. We present two efficient implementations of the unifying methodology: an instance-based PropDRM approach, and a feature-based PropStar approach to data transformation and learning, together with their empirical evaluation on several relational problems. The results show that the new algorithms can outperform existing relational learners and can solve much larger problems. Nada Lavrac, Blaz Skrlj, Marko Robnik-Sikonja |
Mach. Learn. | 1 |
| 2020 | Embedding-based Silhouette community detectionabstractMining complex data in the form of networks is of increasing interest in many scientific disciplines. Network communities correspond to densely connected subnetworks, and often represent key functional parts of real-world systems. This paper proposes the embedding-based Silhouette community detection (SCD), an approach for detecting communities, based on clustering of network node embeddings, i.e. real valued representations of nodes derived from their neighborhoods. We investigate the performance of the proposed SCD approach on 234 synthetic networks, as well as on a real-life social network. Even though SCD is not based on any form of modularity optimization, it performs comparably or better than state-of-the-art community detection algorithms, such as the InfoMap and Louvain. Further, we demonstrate that SCD's outputs can be used along with domain ontologies in semantic subgroup discovery, yielding human-understandable explanations of communities detected in a real-life protein interaction network. Being embedding-based, SCD is widely applicable and can be tested out-of-the-box as part of many existing network learning and exploration pipelines. Blaz Skrlj, Jan Kralj, Nada Lavrac |
Mach. Learn. | 3 |
| 2019 | Connection Between the Parkinson's Disease Subtypes and Patients' Symptoms Progression
Anita Valmarska, Dragana Miljkovic, Marko Robnik-Sikonja, Nada Lavrac |
AIME | 4 |
| 2019 | Symbolic Graph Embedding Using Frequent Pattern Mining
Blaz Skrlj, Nada Lavrac, Jan Kralj |
DS | 2 |
| 2019 | Interactive exploration of heterogeneous biological networks with Biomine ExplorerabstractSUMMARY: Biomine Explorer is a web application that enables interactive exploration of large heterogeneous biological networks constructed from selected publicly available biological knowledge sources. It is built on top of Biomine, a system which integrates cross-references from several biological databases into a large heterogeneous probabilistic network. Biomine Explorer offers user-friendly interfaces for search, visualization, exploration and manipulation as well as public and private storage of discovered subnetworks with permanent links suitable for inclusion into scientific publications. A JSON-based web API for network search queries is also available for advanced users. AVAILABILITY AND IMPLEMENTATION: Biomine Explorer is implemented as a web application, which is publicly available at https://biomine.ijs.si. Registration is not required but registered users can benefit from additional features such as private network repositories. Vid Podpecan, Ziva Ramsak, Kristina Gruden, Hannu Toivonen, Nada Lavrac |
Bioinform. | 5 |
| 2019 | CBSSD: community-based semantic subgroup discoveryabstractModern data mining algorithms frequently need to address the task of learning from heterogeneous data, including various sources of background knowledge. A data mining task where ontologies are used as background knowledge in data analysis is referred to as semantic data mining. A specific semantic data mining task is semantic subgroup discovery: a rule learning approach enabling ontology terms to be used in subgroup descriptions learned from class labeled data. This paper presents Community-Based Semantic Subgroup Discovery (CBSSD), a novel approach that advances ontology-based subgroup identification by exploiting the structural properties of induced complex networks related to the studied phenomenon. Following the idea of multi-view learning, using different sources of information to obtain better models, the CBSSD approach can leverage different types of nodes of the induced complex network, simultaneously using information from multiple levels of a biological system. The approach was tested on ten data sets consisting of genes related to complex diseases, as well as core metabolic processes. The experimental results demonstrate that the CBSSD approach is scalable, applicable to large complex networks, and that it can be used to identify significant combinations of terms, which can not be uncovered by contemporary term enrichment analysis approaches. Blaz Skrlj, Jan Kralj, Nada Lavrac |
J. Intell. Inf. Syst. | 3 |
| 2019 | NetSDM: Semantic Data Mining with Network AnalysisabstractSemantic data mining (SDM) is a form of relational data mining that uses annotated data together with complex semantic background knowledge to learn rules that can be easily interpreted. The drawback of SDM is a high computational complexity of existing SDM algorithms, resulting in long run times even when applied to relatively small data sets. This paper proposes an effective SDM approach, named NetSDM, which first transforms the available semantic background knowledge into a network format, followed by network analysis based node ranking and pruning to significantly reduce the size of the original background knowledge. The experimental evaluation of the NetSDM methodology on acute lymphoblastic leukemia and breast cancer data demonstrates that NetSDM achieves radical time efficiency improvements and that learned rules are comparable or better than the rules obtained by the original SDM algorithms. Jan Kralj, Marko Robnik-Sikonja, Nada Lavrac |
J. Mach. Learn. Res. | 3 |
| 2018 | Visualization and Analysis of Parkinson's Disease Status and Therapy Patterns
Anita Valmarska, Dragana Miljkovic, Marko Robnik-Sikonja, Nada Lavrac |
DS | 4 |
| 2018 | Conceptualising Computational Creativity: Towards automated historiography of a research field
Geraint A. Wiggins, Nada Lavrac, Vid Podpecan, Senja Pollak |
ICCC | 2 |
| 2018 | Targeted End-to-End Knowledge Graph Decomposition
Blaz Skrlj, Jan Kralj, Nada Lavrac |
ILP | 3 |
| 2018 | Symptoms and medications change patterns for Parkinson's disease patients stratification
Anita Valmarska, Dragana Miljkovic, Spiros Konitsiotis, Dimitrios A. Gatsios, Nada Lavrac, Marko Robnik-Sikonja |
Artif. Intell. Medicine | 5 |
| 2018 | HINMINE: heterogeneous information network mining with information retrieval heuristics
Jan Kralj, Marko Robnik-Sikonja, Nada Lavrac |
J. Intell. Inf. Syst. | 3 |
| 2018 | Redescription mining augmented with random forest of multi-target predictive clustering trees
Matej Mihelcic, Saso Dzeroski, Nada Lavrac, Tomislav Smuc |
J. Intell. Inf. Syst. | 3 |
| 2018 | Analysis of medications change in Parkinson's disease progression data
Anita Valmarska, Dragana Miljkovic, Nada Lavrac, Marko Robnik-Sikonja |
J. Intell. Inf. Syst. | 3 |
| 2018 | Correction to: Analysis of medications change in Parkinson's disease progression data
Anita Valmarska, Dragana Miljkovic, Nada Lavrac, Marko Robnik-Sikonja |
J. Intell. Inf. Syst. | 3 |
| 2017 | Combining Multitask Learning and Short Time Series Analysis in Parkinson's Disease Patients Stratification
Anita Valmarska, Dragana Miljkovic, Spiros Konitsiotis, Dimitrios A. Gatsios, Nada Lavrac, Marko Robnik-Sikonja |
AIME | 5 |
| 2017 | Outlier based literature exploration for cross-domain linking of Alzheimer's disease and gut microbiota
Donatella Gubiani, Elsa Fabbretti, Bojan Cestnik, Nada Lavrac, Tanja Urbancic |
Expert Syst. Appl. | 4 |
| 2017 | A framework for redescription set construction
Matej Mihelcic, Saso Dzeroski, Nada Lavrac, Tomislav Smuc |
Expert Syst. Appl. | 3 |
| 2017 | Refinement and selection heuristics in subgroup discovery and classification rule learning
Anita Valmarska, Nada Lavrac, Johannes Fürnkranz, Marko Robnik-Sikonja |
Expert Syst. Appl. | 2 |
| 2017 | ClowdFlows: Online workflows for distributed big data mining
Janez Kranjc, Roman Orac, Vid Podpecan, Nada Lavrac, Marko Robnik-Sikonja |
Future Gener. Comput. Syst. | 4 |
| 2016 | Detection of dependencies between literature domains through relation extraction and copulasabstractKnowledge discovery, especially in the field of literature mining, is often involved in searching for interconnecting concepts between two different literature domains, which might bring new understanding of the two domains. This paper presents a new approach to discovering dependencies between different biological domains based on copula analysis of literature mining results. In the use case of the domains of plant defence response and redox potential literatures we have first performed relation extraction with Bio3graph, which is a rule-based natural language processing tool for extracting relations in the form (subject, predicate, object) triplets. The results of triplets extraction were analysed by using copula-based approach, which showed that dependencies exist between the two domains indicating a potential for cross-domain literature exploration. Both Bio3graph and software for Clayton and Frank fully nested copulas, which were used in this work, are publicly available. Dragana Miljkovic, Nada Lavrac, Marko Bohanec, Biljana Mileva-Boshkoska |
CoDIT | 2 |
| 2016 | Computational Creativity Conceptualisation Grounded on ICCC Papers
Senja Pollak, Biljana Mileva-Boshkoska, Dragana Miljkovic, Geraint A. Wiggins, Nada Lavrac |
ICCC | 5 |
| 2016 | Computational Creativity Infrastructure for Online Software Composition: A Conceptual Blending Use Case
Martin Znidarsic, Amílcar Cardoso, Pablo Gervás, Pedro Martins 0003, Raquel Hervás, Ana Alves 0001, Hugo Gonçalo Oliveira, Ping Xiao, Simo Linkola, Hannu Toivonen, Janez Kranjc, Nada Lavrac |
ICCC | 12 |
| 2016 | Explaining mixture models through semantic pattern mining and banded matrix visualization
Prem Raj Adhikari, Anze Vavpetic, Jan Kralj, Nada Lavrac, Jaakko Hollmén |
Mach. Learn. | 4 |
| 2016 | TextFlows: A visual programming platform for text mining and natural language processing
Matic Perovsek, Janez Kranjc, Tomaz Erjavec, Bojan Cestnik, Nada Lavrac |
Sci. Comput. Program. | 5 |
| 2015 | The Good, the Bad, and the AHA! Blends
Pedro Martins 0003, Tanja Urbancic, Senja Pollak, Nada Lavrac, Amílcar Cardoso |
ICCC | 4 |
| 2015 | Relational and Semantic Data Mining - - Invited Talk -
Nada Lavrac, Anze Vavpetic |
LPNMR | 1 |
| 2015 | Mining Text Enriched Heterogeneous Citation Networks
Jan Kralj, Anita Valmarska, Marko Robnik-Sikonja, Nada Lavrac |
PAKDD (1) | 4 |
| 2015 | Wordification: Propositionalization by unfolding relational data into bags of words
Matic Perovsek, Anze Vavpetic, Janez Kranjc, Bojan Cestnik, Nada Lavrac |
Expert Syst. Appl. | 5 |
| 2015 | Relating ensemble diversity and performance: A study in class noise detection
Borut Sluban, Nada Lavrac |
Neurocomputing | 2 |
| 2015 | Active learning for sentiment analysis on data streams: Methodology and workflow implementation in the ClowdFlows platform
Janez Kranjc, Jasmina Smailovic, Vid Podpecan, Miha Grcar, Martin Znidarsic, Nada Lavrac |
Inf. Process. Manag. | 6 |
| 2014 | Explaining Mixture Models through Semantic Pattern Mining and Banded Matrix Visualization
Prem Raj Adhikari, Anze Vavpetic, Jan Kralj, Nada Lavrac, Jaakko Hollmén |
Discovery Science | 4 |
| 2014 | Multilayer Clustering: A Discovery Experiment on Country Level Trading Data
Dragan Gamberger, Matej Mihelcic, Nada Lavrac |
Discovery Science | 3 |
| 2014 | Baseline Methods for Automated Fictional Ideation
Maria Teresa Llano, Rose Hepworth, Simon Colton, Jeremy Gow, John William Charnley, Nada Lavrac, Martin Znidarsic, Matic Perovsek, Mark Granroth-Wilding, Stephen Clark |
ICCC | 6 |
| 2014 | Propositionalization Online
Nada Lavrac, Matic Perovsek, Anze Vavpetic |
ECML/PKDD (3) | 1 |
| 2014 | GMOseek: a user friendly tool for optimized GMO testingabstractBACKGROUND: With the increasing pace of new Genetically Modified Organisms (GMOs) authorized or in pipeline for commercialization worldwide, the task of the laboratories in charge to test the compliance of food, feed or seed samples with their relevant regulations became difficult and costly. Many of them have already adopted the so called "matrix approach" to rationalize the resources and efforts used to increase their efficiency within a limited budget. Most of the time, the "matrix approach" is implemented using limited information and some proprietary (if any) computational tool to efficiently use the available data. RESULTS: The developed GMOseek software is designed to support decision making in all the phases of routine GMO laboratory testing, including the interpretation of wet-lab results. The tool makes use of a tabulated matrix of GM events and their genetic elements, of the laboratory analysis history and the available information about the sample at hand. The tool uses an optimization approach to suggest the most suited screening assays for the given sample. The practical GMOseek user interface allows the user to customize the search for a cost-efficient combination of screening assays to be employed on a given sample. It further guides the user to select appropriate analyses to determine the presence of individual GM events in the analyzed sample, and it helps taking a final decision regarding the GMO composition in the sample. GMOseek can also be used to evaluate new, previously unused GMO screening targets and to estimate the profitability of developing new GMO screening methods. CONCLUSION: The presented freely available software tool offers the GMO testing laboratories the possibility to select combinations of assays (e.g. quantitative real-time PCR tests) needed for their task, by allowing the expert to express his/her preferences in terms of multiplexing and cost. The utility of GMOseek is exemplified by analyzing selected food, feed and seed samples from a national reference laboratory for GMO testing and by comparing its performance to existing tools which use the matrix approach. GMOseek proves superior when tested on real samples in terms of GMO coverage and cost efficiency of its screening strategies, including its capacity of simple interpretation of the testing results. Dany Morisset, Petra Kralj Novak, Darko Zupanic, Kristina Gruden, Nada Lavrac, Jana Zel |
BMC Bioinform. | 5 |
| 2014 | Ensemble-based noise detection: noise ranking and visual performance evaluation
Borut Sluban, Dragan Gamberger, Nada Lavrac |
Data Min. Knowl. Discov. | 3 |
| 2014 | Stream-based active learning for sentiment analysis in the financial domain
Jasmina Smailovic, Miha Grcar, Nada Lavrac, Martin Znidarsic |
Inf. Sci. | 3 |
| 2014 | Semantic subgroup explanations
Anze Vavpetic, Vid Podpecan, Nada Lavrac |
J. Intell. Inf. Syst. | 3 |
| 2013 | Integrating semantic transcriptomic data analysis and knowledge extraction from biological literatureabstractThe paper presents an approach to the holistic analysis of transcriptomic data which integrates two state-of-the-art methodologies into a coherent framework. The aim of the proposed approach is to give insight into the discovered patterns, help explaining the observed phenomena, enable the creation of new research hypotheses and assist in design of new experiments. We have integrated a methodology for semantic analysis of transcriptomic data, a system for automated extraction of biological relations from the literature, and a number of supporting components. The approach is demonstrated and evaluated on a publicly available dataset from a clinical trial in acute lymphoblastic leukaemia and a document corpus of full-text articles from the PubMed Open Access Subset. Vid Podpecan, Dragana Miljkovic, Marko Petek, Tjasa Stare, Kristina Gruden, Igor Mozetic, Nada Lavrac |
BIBM | 7 |
| 2013 | Real-time data analysis in ClowdFlowsabstractClowdFlows is an open cloud based platform for composition, execution, and sharing of interactive data mining workflows. In this paper we extend the ClowdFlows platform with the ability to mine real-time data streams. This functionality was implemented by creating a specialized type of workflow component and a stream mining daemon that delegates the execution of workflows in real-time. In this way, we have transformed a batch data processing platform into a real-time stream mining platform with an intuitive user interface. The real-time analytics aspect of the platform is demonstrated in a Twitter sentiment analysis use case where the sentiment of tweets about whistleblower Edward Snowden was monitored for approximately one month. Janez Kranjc, Vid Podpecan, Nada Lavrac |
IEEE BigData | 3 |
| 2013 | A Wordification Approach to Relational Data Mining
Matic Perovsek, Anze Vavpetic, Bojan Cestnik, Nada Lavrac |
Discovery Science | 4 |
| 2013 | Semantic Data Mining of Financial News Articles
Anze Vavpetic, Petra Kralj Novak, Miha Grcar, Igor Mozetic, Nada Lavrac |
Discovery Science | 5 |
| 2013 | Towards Narrative Ideation via Cross-Context Link Discovery Using Banded Matrices
Matic Perovsek, Bojan Cestnik, Tanja Urbancic, Simon Colton, Nada Lavrac |
IDA | 5 |
| 2013 | ViperCharts: Visual Performance Evaluation Platform
Borut Sluban, Nada Lavrac |
ECML/PKDD (3) | 2 |
| 2013 | A Methodology for Mining Document-Enriched Heterogeneous Information NetworksabstractThe paper presents a new methodology for mining heterogeneous information networks, motivated by the fact that, in many real-life scenarios, documents are available in heterogeneous information networks, such as interlinked multimedia objects containing titles, descriptions and subtitles. The methodology consists of transforming documents into bag-of-words vectors, decomposing the corresponding heterogeneous network into separate graphs, computing structural-context feature vectors with PageRank, and finally, constructing a common feature vector space in which knowledge discovery is performed. We exploit this feature vector construction process to devise an efficient centroid-based classification algorithm. We demonstrate the approach by applying it to the task of categorizing video lectures. We show that our approach exhibits low time and space complexity without compromising the classification accuracy. In addition, we provide a qualitative analysis of the results by employing a data visualization technique. Miha Grcar, Nejc Trdin, Nada Lavrac |
Comput. J. | 3 |
| 2013 | Contrasting Subgroup DiscoveryabstractSubgroup discovery methods find interesting subsets of objects of a given class. Motivated by an application in bioinformatics, we first define a generalized subgroup discovery problem. In this setting, a subgroup is interesting if its members are characteristic for their class, even if the classes are not identical. Then we further refine this setting for the case where subsets of objects, for example, subsets of objects that represent different time points or different phenotypes, are contrasted. We show that this allows finding subgroups of objects that could not be found with classical subgroup discovery. To find such subgroups, we propose an approach that consists of two subgroup discovery steps and an intermediate, contrast set definition step. This approach is applicable in various application areas. An example is biology, where interesting subgroups of genes are searched by using gene expression data. We address the problem of finding enriched gene sets that are specific for virus infected samples for a specific time point or a specific phenotype. We report on experimental results on a time series data set for virus infected Solanum tuberosum (potato) plants. The results on S. tuberosum’s response to virus infection revealed new research hypotheses for plant biologists. Laura Langohr, Vid Podpecan, Marko Petek, Igor Mozetic, Kristina Gruden, Nada Lavrac, Hannu Toivonen |
Comput. J. | 6 |
| 2013 | Semantic Subgroup Discovery Systems and Workflows in the SDM-ToolkitabstractThis paper addresses semantic data mining, a new data mining paradigm in which ontologies are exploited in the process of data mining and knowledge discovery. This paradigm is introduced together with new semantic subgroup discovery systems SDM-search for enriched gene sets (SEGS) and SDM-Aleph. These systems are made publicly available in the new SDM-Toolkit for semantic data mining. The toolkit is implemented in the Orange4WS data mining platform that supports knowledge discovery workflow construction from local and distributed data mining services. On the basis of the experimental evaluation of semantic subgroup discovery systems on two publicly available biomedical datasets, the paper results in a thorough quantitative and qualitative evaluation of SDM-SEGS and SDM-Aleph and their comparison with SEGS, a system for enriched gene set discovery from microarray data. Anze Vavpetic, Nada Lavrac |
Comput. J. | 2 |
| 2012 | Advances in Data Mining for Biomedical ResearchabstractThis CBMS-2012 keynote first outlines standard approaches to data mining, with the emphasis on subgroup discovery which proves to be an effective tool for data analysis in biomedical applications. The core of this paper is devoted to inductive logic programming and relational data mining which also have a great potential for biomedical research, with a focus on recently developed approaches to semantic data mining which enable the use of domain ontologies as background knowledge in data analysis. The use of described techniques and tools is illustrated on selected biomedical applications. Nada Lavrac |
CBMS | 1 |
| 2012 | Cross-domain literature mining: Finding bridging concepts with CrossBee
Matjaz Jursic, Bojan Cestnik, Tanja Urbancic, Nada Lavrac |
ICCC | 4 |
| 2012 | Irregularity Detection in Categorized Document Corpora
Borut Sluban, Senja Pollak, Roel Coesemans, Nada Lavrac |
LREC | 4 |
| 2012 | ClowdFlows: A Cloud Based Scientific Workflow Platform
Janez Kranjc, Vid Podpecan, Nada Lavrac |
ECML/PKDD (2) | 3 |
| 2012 | Explaining Subgroups through Ontologies
Anze Vavpetic, Vid Podpecan, Stijn Meganck, Nada Lavrac |
PRICAI | 4 |
| 2012 | Outlier Detection in Cross-Context Link Discovery for Creative Literature MiningabstractThis paper investigates the role of outliers in literature-based knowledge discovery. It shows that detecting interesting outliers which appear in the literature on a given phenomenon can help the expert to find implicit relationships among concepts of different domains. The underlying assumption is that while the majority of articles in the given scientific domain describe matters related to a common understanding of the domain, the exploration of outliers may lead to the detection of scientifically interesting bridging concepts among disjoint sets of scientific articles. The proposed approach contributes to cross-context link discovery by proving the utility of outlier detection for finding bisociative links in the process of autism literature exploration, as well as by uncovering implicit relationships in the articles from the migraine domain. Ingrid Petric, Bojan Cestnik, Nada Lavrac, Tanja Urbancic |
Comput. J. | 3 |
| 2012 | Orange4WS Environment for Service-Oriented Data MiningabstractNovel data-mining tasks in e-science involve mining of distributed, highly heterogeneous data and knowledge sources. However, standard data mining platforms, such as Weka and Orange, involve only their own data mining algorithms in the process of knowledge discovery from local data sources. In contrast, next generation data mining technologies should enable processing of distributed data sources, the use of data mining algorithms implemented as web services, as well as the use of formal descriptions of data sources and knowledge discovery tools in the form of ontologies, enabling automated composition of complex knowledge discovery workflows for a given data mining task. This paper proposes a novel Service-oriented Knowledge Discovery framework and its implementation in a service-oriented data mining environment Orange4WS (Orange for Web Services), based on the existing Orange data mining toolbox and its visual programming environment, which enables manual composition of data mining workflows. The new service-oriented data mining environment Orange4WS includes the following new features: simple use of web services as remote components that can be included into a data mining workflow; simple incorporation of relational data mining algorithms; a knowledge discovery ontology to describe workflow components (data, knowledge and data mining services) in an abstract and machine-interpretable way, and its use by a planner that enables automated composition of data mining workflows. These new features are showcased in three real-world scenarios. Vid Podpecan, Monika Zemenova, Nada Lavrac |
Comput. J. | 3 |
| 2011 | Evaluating Outliers for Cross-Context Link Discovery
Borut Sluban, Matjaz Jursic, Bojan Cestnik, Nada Lavrac |
AIME | 4 |
| 2011 | A Methodology for Mining Document-Enriched Heterogeneous Information Networks
Miha Grcar, Nada Lavrac |
Discovery Science | 2 |
| 2011 | Using Ontologies in Semantic Data Mining with SEGS and g-SEGS
Nada Lavrac, Anze Vavpetic, Larisa N. Soldatova, Igor Trajkovski, Petra Kralj Novak |
Discovery Science | 1 |
| 2011 | SegMine workflows for semantic microarray data analysis in Orange4WSabstractBACKGROUND: In experimental data analysis, bioinformatics researchers increasingly rely on tools that enable the composition and reuse of scientific workflows. The utility of current bioinformatics workflow environments can be significantly increased by offering advanced data mining services as workflow components. Such services can support, for instance, knowledge discovery from diverse distributed data and knowledge sources (such as GO, KEGG, PubMed, and experimental databases). Specifically, cutting-edge data analysis approaches, such as semantic data mining, link discovery, and visualization, have not yet been made available to researchers investigating complex biological datasets. RESULTS: We present a new methodology, SegMine, for semantic analysis of microarray data by exploiting general biological knowledge, and a new workflow environment, Orange4WS, with integrated support for web services in which the SegMine methodology is implemented. The SegMine methodology consists of two main steps. First, the semantic subgroup discovery algorithm is used to construct elaborate rules that identify enriched gene sets. Then, a link discovery service is used for the creation and visualization of new biological hypotheses. The utility of SegMine, implemented as a set of workflows in Orange4WS, is demonstrated in two microarray data analysis applications. In the analysis of senescence in human stem cells, the use of SegMine resulted in three novel research hypotheses that could improve understanding of the underlying mechanisms of senescence and identification of candidate marker genes. CONCLUSIONS: Compared to the available data analysis systems, SegMine offers improved hypothesis generation and data interpretation for bioinformatics in an easy-to-use integrated workflow environment. Vid Podpecan, Nada Lavrac, Igor Mozetic, Petra Kralj Novak, Igor Trajkovski, Laura Langohr, Kimmo Kulovesi, Hannu Toivonen, Marko Petek, Helena Motaln, Kristina Gruden |
BMC Bioinform. | 2 |
| 2011 | Automating Knowledge Discovery Workflow Composition Through Ontology-Based PlanningabstractThe problem addressed in this paper is the challenge of automated construction of knowledge discovery workflows, given the types of inputs and the required outputs of the knowledge discovery process. Our methodology consists of two main ingredients. The first one is defining a formal conceptualization of knowledge types and data mining algorithms by means of knowledge discovery ontology. The second one is workflow composition formalized as a planning task using the ontology of domain and task descriptions. Two versions of a forward chaining planning algorithm were developed. The baseline version demonstrates suitability of the knowledge discovery ontology for planning and uses Planning Domain Definition Language (PDDL) descriptions of algorithms; to this end, a procedure for converting data mining algorithm descriptions into PDDL was developed. The second directly queries the ontology using a reasoner. The proposed approach was tested in two use cases, one from scientific discovery in genomics and another from advanced engineering. The results show the feasibility of automated workflow construction achieved by tight integration of planning and ontological reasoning. Monika Záková, Petr Kremen, Filip Zelezný, Nada Lavrac |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2010 | Efficient Visualization of Document Streams
Miha Grcar, Vid Podpecan, Matjaz Jursic, Nada Lavrac |
Discovery Science | 4 |
| 2010 | Advances in Class Noise DetectionabstractNoise filtering is usually used in data preprocessing to improve the accuracy of induced classifiers. Our goal is different: we aim at detecting noisy instances to be inspected by the domain expert in the phase of data understanding. Consequently, our noise detection algorithms should have high precision of class noise detection, where the precision-recall trade-off is modeled using the F-measure. New variants of class noise detection algorithms have been developed, including the high agreement random forest filter which ensures very high precision of identified erroneous data instances. Borut Sluban, Dragan Gamberger, Nada Lavrac |
ECAI | 3 |
| 2010 | Bisociative Knowledge Discovery for Microarray Data Analysis
Igor Mozetic, Nada Lavrac, Vid Podpecan, Petra Kralj Novak, Helena Motaln, Marko Petek, Kristina Gruden, Hannu Toivonen, Kimmo Kulovesi |
ICCC | 2 |
| 2010 | Workflow Construction for Service-Oriented Knowledge Discovery
Vid Podpecan, Monika Záková, Nada Lavrac |
ISoLA (1) | 3 |
| 2010 | Semi-supervised Constrained Clustering: An Expert-Guided Data Analysis Methodology
Vid Podpecan, Miha Grcar, Nada Lavrac |
PRICAI | 3 |
| 2010 | Identification of concepts bridging diverse biomedical domainsabstractBackground In biology and medicine, experts are challenged daily with linking information from various highly specialized subfields. The individual subfields can be considered habitually different domains since experts usually master only one of them. However, many novel discoveries are achieved by gaining new insights and knowledge via fusing two or more diverse fields. In this work we propose a method that reveals key concepts which are the most informative and promising to pursue when bridging Matjaz Jursic, Igor Mozetic, Miha Grcar, Bojan Cestnik, Nada Lavrac |
BMC Bioinform. | 5 |
| 2010 | Supporting the search for cross-context links by outlier detection methodsabstractBackground and relation to previous work Outliers in data can either present noise in the data, which has harmful effects on knowledge discovery (and should therefore best be eliminated), or correct data instances that belong to a specific subconcept of the main domain concept (and can potentially carry new interesting insights). Several outlier detection methods have been developed in data and text mining, mainly used for noise filtering and error detection purposes. Except for [1], outlier detection in text mining has not yet been used for exploratory purposes. Our work focuses on using noise/outlier detection methods for a novel task of cross-context link discovery. Borut Sluban, Nada Lavrac |
BMC Bioinform. | 2 |
| 2009 | CSM-SD: Methodology for contrast set mining through subgroup discovery
Petra Kralj Novak, Nada Lavrac, Dragan Gamberger, Antonija Krstacic |
J. Biomed. Informatics | 2 |
| 2009 | Supervised Descriptive Rule Discovery: A Unifying Survey of Contrast Set, Emerging Pattern and Subgroup Mining
Petra Kralj Novak, Nada Lavrac, Geoffrey I. Webb |
J. Mach. Learn. Res. | 2 |
| 2009 | Guest editors' introduction: Special issue on Inductive Logic Programming (ILP-2008)
Filip Zelezný, Nada Lavrac |
Mach. Learn. | 2 |
| 2008 | On the Design of Knowledge Discovery Services Design Patterns and Their Application in a Use Case Implementation
Jeroen S. de Bruin, Joost N. Kok, Nada Lavrac, Igor Trajkovski |
ISoLA | 3 |
| 2008 | Advancing Topic Ontology Learning through Term Extraction
Blaz Fortuna, Nada Lavrac, Paola Velardi |
PRICAI | 2 |
| 2008 | Handling Unknown and Imprecise Attribute Values in Propositional Rule Learning: A Feature-Based Approach
Dragan Gamberger, Nada Lavrac, Johannes Fürnkranz |
PRICAI | 2 |
| 2008 | SEGS: Search for enriched gene sets in microarray data
Igor Trajkovski, Nada Lavrac, Jakub Tolar |
J. Biomed. Informatics | 2 |
| 2008 | Closed Sets for Labeled Data
Gemma C. Garriga, Petra Kralj Novak, Nada Lavrac |
J. Mach. Learn. Res. | 3 |
| 2008 | Learning Relational Descriptions of Differentially Expressed Gene GroupsabstractThis paper presents a method that uses gene ontologies (GOs), together with the paradigm of relational subgroup discovery, to find compactly described groups of genes differentially expressed in specific cancers. The groups are described by means of relational logic features, extracted from publicly available GO information, and are straightforwardly interpretable by medical experts. We applied the proposed method to three gene expression data sets with the following respective sets of sample classes: 1) acute lymphoblastic leukemia (ALL) versus acute myeloid leukemia (AML); 2) seven subtypes of ALL; and 3) 14 different types of cancers. Significant number of discovered groups of genes had a description that highlighted the underlying biological process responsible for distinguishing one class from the other classes. The quality of the discovered descriptions was also verified by cross validation. We believe that the presented approach will significantly contribute to the application of relational machine learning to gene expression analysis, given the expected increase in both the quality and quantity of gene/protein annotations in the near future. Igor Trajkovski, Filip Zelezný, Nada Lavrac, Jakub Tolar |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2007 | Supporting Factors in Descriptive Analysis of Brain Ischaemia
Dragan Gamberger, Nada Lavrac |
AIME | 2 |
| 2007 | Contrast Set Mining for Distinguishing Between Similar Diseases
Petra Kralj Novak, Nada Lavrac, Dragan Gamberger, Antonija Krstacic |
AIME | 2 |
| 2007 | Monitoring Human Resources of a Public Health-Care System Through Intelligent Data Analysis and Visualization
Aleksander Pur, Marko Bohanec, Nada Lavrac, Bojan Cestnik, Marko Debeljak, Anton Gradisek |
AIME | 3 |
| 2007 | Interpreting Gene Expression Data by Searching for Enriched Gene Sets
Igor Trajkovski, Nada Lavrac |
AIME | 2 |
| 2007 | Efficient Generation of Biologically Relevant Enriched Gene Sets
Igor Trajkovski, Nada Lavrac |
ISBRA | 2 |
| 2007 | Contrast Set Mining Through Subgroup Discovery Applied to Brain Ischaemina Data
Petra Kralj Novak, Nada Lavrac, Dragan Gamberger, Antonija Krstacic |
PAKDD | 2 |
| 2007 | Clinical data analysis based on iterative subgroup discovery: experiments in brain ischaemia data analysis
Dragan Gamberger, Nada Lavrac, Antonija Krstacic, Goran Krstacic |
Appl. Intell. | 2 |
| 2007 | Data mining and visualization for decision support and modeling of public health-care resources
Nada Lavrac, Marko Bohanec, Aleksander Pur, Bojan Cestnik, Marko Debeljak, Andrej Kobler |
J. Biomed. Informatics | 1 |
| 2007 | Trust Modeling for Networked Organizations Using Reputation and Collaboration EstimatesabstractThe main motivation for organizations and individuals to collaborate is to enable knowledge and resource sharing in order to effectively fulfill a joint business opportunity. This correspondence focuses on virtual organizations (VOs) and virtual teams (VTs), whose strengths lie in the range of competencies of their members, offered jointly through collaboration. One of the difficulties in VO and VT creation is partner selection using partners' mutual trust as one of the selection criteria. This correspondence provides an analysis of trust relationships based on the principal–agent theory, and proposes an approach to hierarchical multiattribute decision-support-based trust estimation applied to a network of collaborating organizations (VO) and a network of collaborating individuals (VT). The correspondence presents two case studies, one using a questionnaire-based approach and the other using automated reputation and collaboration estimation from data gathered by Web crawling. Nada Lavrac, Peter Ljubic, Tanja Urbani, Gregor Papa, Mitja Jermol, Stefan Bollhalter |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2007 | An Ontology for Virtual Organization Breeding EnvironmentsabstractCompanies and individuals connect into networks to share their resources with the purpose of achieving a common goal. The field of collaborative network organizations covers various types of organizational structures. Knowledge, which is stored in such networks, can be separated into two different levels. First, there is a common knowledge about the organizational structure itself that can be used and reused in any of such networks. The second level represents the domain-specific knowledge, which such networks cover and use to function. In this paper, we address both levels, first, by proposing an ontology representing the common vocabulary and identifying the actors and relationships in a specific type of network, namely virtual organization breeding environment (VBE), and second, by proposing a methodology for extracting network-specific knowledge related to competencies. The instantiation of the proposed VBE ontology and the developed approach to semiautomated construction of competencies have been applied to real problem scenarios of Virtuelle Fabrik, a Swiss-German cluster of companies in mechanical engineering. Joël Plisson, Peter Ljubic, Igor Mozetic, Nada Lavrac |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2006 | Relational Data Mining Applied to Virtual Engineering of Product Designs
Monika Záková, Filip Zelezný, Javier A. García-Sedano, Cyril Masia Tissot, Nada Lavrac, Petr Kremen, Javier Molina |
ILP | 5 |
| 2006 | Closed Sets for Labeled Data
Gemma C. Garriga, Petra Kralj Novak, Nada Lavrac |
PKDD | 3 |
| 2006 | Propositionalization-based relational subgroup discovery with RSD
Filip Zelezný, Nada Lavrac |
Mach. Learn. | 2 |
| 2005 | Resource Modeling and Analysis of Regional Public Health Care Data by Means of Knowledge Technologies
Nada Lavrac, Marko Bohanec, Aleksander Pur, Bojan Cestnik, Mitja Jermol, Tanja Urbancic, Marko Debeljak, Branko Kavsek, Tadeja Kopac |
AIME | 1 |
| 2005 | A Decision Support Approach to Modeling Trust in Networked Organizations
Nada Lavrac, Peter Ljubic, Mitja Jermol, Gregor Papa |
IEA/AIE | 1 |
| 2005 | Data Mining for Decision Support: An Application in Public Health Care
Aleksander Pur, Marko Bohanec, Bojan Cestnik, Nada Lavrac, Marko Debeljak, Tadeja Kopac |
IEA/AIE | 4 |
| 2005 | Hierarchical Multi-Attribute Decision Support Approach to Virtual Organization Creation
Toni Jarimo, Peter Ljubic, Iiro Salkari, Marko Bohanec, Nada Lavrac, Martin Znidarsic, Stefan Bollhalter, Jirí Hodík |
PRO-VE | 5 |
| 2005 | A Decision Support Approach to Trust Modeling in Networked Organizations
Nada Lavrac, Peter Ljubic, Mitja Jermol, Stefan Bollhalter |
PRO-VE | 1 |
| 2005 | Subgroup Discovery Techniques and Applications
Nada Lavrac |
PAKDD | 1 |
| 2004 | Avoiding Data Overfitting in Scientific Discovery: Experiments in Functional Genomics
Dragan Gamberger, Nada Lavrac |
ECAI | 2 |
| 2004 | Induction of comprehensible models for gene expression datasets by subgroup discovery methodology
Dragan Gamberger, Nada Lavrac, Filip Zelezný, Jakub Tolar |
J. Biomed. Informatics | 2 |
| 2004 | Subgroup Discovery with CN2-SD
Nada Lavrac, Branko Kavsek, Peter A. Flach, Ljupco Todorovski |
J. Mach. Learn. Res. | 1 |
| 2004 | Decision Support Through Subgroup Discovery: Three Case Studies and the Lessons Learned
Nada Lavrac, Bojan Cestnik, Dragan Gamberger, Peter A. Flach |
Mach. Learn. | 1 |
| 2004 | Editorial: Data Mining Lessons Learned
Nada Lavrac, Hiroshi Motoda, Tom Fawcett |
Mach. Learn. | 1 |
| 2004 | Introduction: Lessons Learned from Data Mining Applications and Collaborative Problem Solving
Nada Lavrac, Hiroshi Motoda, Tom Fawcett, Robert C. Holte, Pat Langley, Pieter W. Adriaans |
Mach. Learn. | 1 |
| 2003 | Analysis of Gene Expression Data by the Logic Minimization Approach
Dragan Gamberger, Nada Lavrac |
AIME | 2 |
| 2003 | APRIORI-SD: Adapting Association Rule Learning to Subgroup Discovery
Branko Kavsek, Nada Lavrac, Viktor Jovanoski |
IDA | 2 |
| 2003 | Comparative Evaluation of Approaches to Propositionalization
Mark-A. Krogel, Simon Alan Rawles, Filip Zelezný, Peter A. Flach, Nada Lavrac, Stefan Wrobel |
ILP | 5 |
| 2003 | Active subgroup mining: a case study in coronary heart disease risk group detection
Dragan Gamberger, Nada Lavrac |
Artif. Intell. Medicine | 2 |
| 2002 | Adapting classification rule induction to subgroup discoveryabstractRule learning is typically used for solving classification and prediction tasks. However learning of classification rules can be adapted also to subgroup discovery. This paper shows how this can be achieved by modifying the covering algorithm and the search heuristic, performing probabilistic classification of instances, and using an appropriate measure for evaluating the results of subgroup discovery. Experimental evaluation of the CN2-SD subgroup discovery algorithm on 17 UCI data sets demonstrates substantial reduction of the number of induced rules, increased rule coverage and rule significance, as well as slight improvements in terms of the area under the ROC curve. Nada Lavrac, Peter A. Flach, Branko Kavsek, Ljupco Todorovski |
ICDM | 1 |
| 2002 | Descriptive Induction through Subgroup Discovery: A Case Study in a Medical Domain
Dragan Gamberger, Nada Lavrac |
ICML | 2 |
| 2002 | Virtual Enterprise for Data Mining and Decision Support
Nada Lavrac |
PRO-VE | 1 |
| 2002 | RSD: Relational Subgroup Discovery through First-Order Feature Construction
Nada Lavrac, Filip Zelezný, Peter A. Flach |
ILP | 1 |
| 2002 | Generating Actionable Knowledge by Expert-Guided Subgroup Discovery
Dragan Gamberger, Nada Lavrac |
PKDD | 2 |
| 2002 | Expert-Guided Subgroup Discovery: Methodology and ApplicationabstractThis paper presents an approach to expert-guided subgroup discovery. The main step of the subgroup discovery process, the induction of subgroup descriptions, is performed by a heuristic beam search algorithm, using a novel parametrized definition of rule quality which is analyzed in detail. The other important steps of the proposed subgroup discovery process are the detection of statistically significant properties of selected subgroups and subgroup visualization: statistically significant properties are used to enrich the descriptions of induced subgroups, while the visualization shows subgroup properties in the form of distributions of the numbers of examples in the subgroups. The approach is illustrated by the results obtained for a medical problem of early detection of patient risk groups. Dragan Gamberger, Nada Lavrac |
J. Artif. Intell. Res. | 2 |
| 2001 | Consensus Decision Trees: Using Consensus Hierarchical Clustering for Data Relabelling and Reduction
Branko Kavsek, Nada Lavrac, Anuska Ferligoj |
ECML | 2 |
| 2001 | AIM portraits: tracing the evolution of artificial intelligence in medicine and predicting its future in the new millennium
Elpida T. Keravnou, Nada Lavrac |
Artif. Intell. Medicine | 2 |
| 2001 | An extended transformation approach to inductive logic programmingabstractInductive logic programming (ILP) is concerned with learning relational descriptions that typically have the form of logic programs. In a transformation approach, an ILP task is transformed into an equivalent learning task in a different representation formalism. Propositionalization is a particular transformation method, in which the ILP task is compiled to an attribute-value learning task. The main restriction of propositionalization methods such as LINUS is that they are unable to deal with nondeterminate local variables in the body of hypothesis clauses. In this paper we show how this limitation can be overcome., by systematic first-order feature construction using a particular individual-centered feature bias. The approach can be applied in any domain where there is a clear notion of individual. We also show how to improve upon exhaustive first-order feature construction by using a relevancy filter. The proposed approach is illustrated on the “trains” and “mutagenesis” ILP domains. Nada Lavrac, Peter A. Flach |
ACM Trans. Comput. Log. | 1 |
| 2000 | Confirmation Rule Sets
Dragan Gamberger, Nada Lavrac |
PKDD | 2 |
| 2000 | Predictive Performance of Weghted Relative Accuracy
Ljupco Todorovski, Peter A. Flach, Nada Lavrac |
PKDD | 3 |
| 1999 | Experiments with Noise Filtering in a Medical Domain
Dragan Gamberger, Nada Lavrac, Ciril Groselj |
ICML | 2 |
| 1999 | Selected techniques for data mining in medicine
Nada Lavrac |
Artif. Intell. Medicine | 1 |
| 1999 | Editorial
Saso Dzeroski, Nada Lavrac |
Data Min. Knowl. Discov. | 2 |
| 1997 | Machine Learning Applied to Diagnosis of Sport Injuries
Igor Zelic, Igor Kononenko 0001, Nada Lavrac, Vanja Vuga |
AIME | 3 |
| 1997 | Using machine learning for outcome prediction of patients with severe head injuryabstractThe paper presents an application of decision tree induction to the problem of the prediction of outcome after a severe head injury. The study shows that induced decision trees are useful for the analysis of the importance of clinical parameters and of their combinations for the evaluation of the severity of brain injury and for outcome prediction. Iztok A. Pilih, Dunja Mladenic, Nada Lavrac, Tine S. Prevec |
CBMS | 3 |
| 1997 | Diagnosis of sport injuries with machine learning: explanation of induced decisionsabstractMachine learning techniques can be used to extract knowledge from data stored in medical databases. In our application, various machine learning algorithms were used to extract diagnostic knowledge to support decisions in the diagnosis of sport injuries. Igor Zelic, Igor Kononenko 0001, Nada Lavrac, Vanja Vuga |
CBMS | 3 |
| 1997 | Conditions for Occam's Razor Applicability and Noise Elimination
Dragan Gamberger, Nada Lavrac |
ECML | 2 |
| 1996 | A Reply to Pazzani's Book Review of "Inductive Logic Programming: Techniques and Applications"
Nada Lavrac, Saso Dzeroski |
Mach. Learn. | 1 |
| 1994 | Weakening the language bias in LINUSabstractThe two main limitations of propositional inductive learning algorithms are the limited capability of taking into account available background knowledge and the limited expressiveness of the knowledge representation formalism used for describing examples, background knowledge and concepts. The paper presents a method for using background knowledge effectively in learning both propositional and relational descriptions. The method, implemented in the system LINUS, uses propositional learners in a more expressive logic programming framework. This allows for learning of logic programs in the form of constrained deductive hierarchical database clauses and determinate deductive database clauses. Nada Lavrac, Saso Dzeroski |
J. Exp. Theor. Artif. Intell. | 1 |
| 1993 | Multiple Predicate Learning
Luc De Raedt, Nada Lavrac, Saso Dzeroski |
IJCAI | 2 |
| 1993 | The Many Faces of Inductive Logic Programming
Luc De Raedt, Nada Lavrac |
ISMIS | 2 |
| 1993 | Inductive Learning in Deductive DatabasesabstractMost current applications of inductive learning in databases take place in the context of a single extensional relation. The authors place inductive learning in the context of a set of relations defined either extensionally or intentionally in the framework of deductive databases. LINUS, an inductive logic programming system that induces virtual relations from example positive and negative tuples and already defined relations in a deductive database, is presented. Based on the idea of transforming the problem of learning relations to attribute-value form, several attribute-value learning systems are incorporated. As the latter handle noisy data successfully, LINUS is able to learn relations from real-life noisy databases. The use of LINUS for learning virtual relations is illustrated, and a study of its performance on noisy data is presented.> Saso Dzeroski, Nada Lavrac |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1992 | Stochastic Search in Inductive Logic Programming
Matevz Kovacic, Nada Lavrac, Marko Grobelnik, Darko Zupanic, Dunja Mladenic |
ECAI | 2 |
| 1991 | Learning Relations from Noisy Examples: An Empirical Comparison of LINUS and FOIL
Saso Dzeroski, Nada Lavrac |
ML | 2 |
| 1986 | The Multi-Purpose Incremental Learning System AQ15 and Its Testing Application to Three Medical Domains
Ryszard S. Michalski, Igor Mozetic, Jiarong Hong, Nada Lavrac |
AAAI | 4 |