Blaz Zupan

dblp:z/BlazZupan · DBLP profile ↗
← Back
71ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0002-5864-7056ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 27 · 9 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorComputer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Inferring Chronic Treatment Onset from ePrescription Data: A Renewal Process Approach
Pavlin Gregor Policar, Dalibor Stanimirovic, Blaz Zupan
AIME (2)3
2026 Online tutorial on survival analysis for biomarker discovery
abstract
In biomedicine, survival analysis addresses time-to-event data to study outcomes like patient survival and treatment response, and supports biomarker discovery. Yet, teaching this analysis is often hindered by mathematical and programming barriers. We present a structured, hands-on tutorial that goes beyond a typical online guide-offering integrated video lectures, literature, quizzes, and practical exercises. Built around Orange Data Mining, an open and free no-code visual analytics platform, the tutorial covers key concepts such as censoring, Kaplan-Meier curves, group comparisons, and biomarker discovery through real-world datasets. Organized in four pedagogical units, it progresses from basic survival data analysis to gene and gene-set biomarker discovery. Designed for 2-3 hours of learning, it supports both individual study and classroom use, and was successfully tested with over 120 participants.
Jaka Kokosar, Ela Praznik, Martin Spendl, Nancy P. Moreno, Alana Newell, Gad Shaulsky, Blaz Zupan
PLoS Comput. Biol.7
2026 Ten simple rules for coordinating a large digital health project: Perspectives from EU and implications for global contexts
abstract
Coordinating a large-scale digital health project requires a unique mix of scientific leadership, administrative skill, and human sensitivity. Drawing from our experience leading CAPABLE, a European Horizon 2020 project aimed at improving the quality of life of cancer patients through AI and telemedicine, we present ten practical rules for navigating the complex landscape of multi-partner biomedical research. These rules address challenges such as building balanced consortia, managing timelines and regulatory requirements, ensuring cultural alignment, and promoting long-term impact through dissemination and exploitation. The paper specifically addresses international research projects at the intersection of healthcare and IT and their peculiar challenges, typically connected to the interplay of different actors such as academics, healthcare personnel, and industry partners located in different countries, each from diverse backgrounds and different working practices. Our goal is to provide researchers and project coordinators with concrete guidance to increase the likelihood of success in future large digital health initiatives.
Lucia Sacchi, Blaz Zupan, Silvana Quaglini
PLoS Comput. Biol.2
2025 Large scale gene set ranking for survival-related gene sets
abstract
Disease progression is closely linked to shifts in the expression levels of specific genes within molecular pathways. While gene set enrichment analysis is a widely employed method for identifying key disease markers, it has been underutilized in survival analysis. Here, we introduce a novel computational approach that adapts gene set enrichment analysis for survival analysis. The proposed approach considers a gene set, computes a single-sample gene set enrichment score, and, based on this score, splits the samples into cohorts. It then scores the gene sets by evaluating the differences in survival rates between the resulting cohorts. We aim to find gene sets that can lead to cohorts with significantly different survival probabilities. Utilizing gene expression data from The Cancer Genome Atlas and gene sets from the Molecular Signature Database, our results demonstrate that existing empirical research consistently supports the top gene sets our approach associates with survival prognosis. The proposed method broadens gene set enrichment analysis applications to include information on survival, bridging the gap between alterations in molecular pathways and their implications on survival.
Martin Spendl, Jaka Kokosar, Ela Praznik, Luka Ausec, Miha Stajdohar, Blaz Zupan
Artif. Intell. Medicine6
2025 Automated assignment grading with large language models: insights from a bioinformatics course
abstract
MOTIVATION: Providing students with individualized feedback through assignments is a cornerstone of education that supports their learning and development. Studies have shown that timely, high-quality feedback plays a critical role in improving learning outcomes. However, providing personalized feedback on a large scale in classes with large numbers of students is often impractical due to the significant time and effort required. Recent advances in natural language processing and large language models (LLMs) offer a promising solution by enabling the efficient delivery of personalized feedback. These technologies can reduce the workload of course staff while improving student satisfaction and learning outcomes. Their successful implementation, however, requires thorough evaluation and validation in real classrooms. RESULTS: We present the results of a practical evaluation of LLM-based graders for written assignments in the 2024/25 iteration of the Introduction to Bioinformatics course at the University of Ljubljana. Over the course of the semester, more than 100 students answered 36 text-based questions, most of which were automatically graded using LLMs. In a blind study, students received feedback from both LLMs and human teaching assistants (TAs) without knowing the source, and later rated the quality of the feedback. We conducted a systematic evaluation of six commercial and open-source LLMs and compared their grading performance with human TAs. Our results show that with well-designed prompts, LLMs can achieve grading accuracy and feedback quality comparable to human graders. Our results also suggest that open-source LLMs perform as well as commercial LLMs, allowing schools to implement their own grading systems while maintaining privacy.
Pavlin Gregor Policar, Martin Spendl, Tomaz Curk, Blaz Zupan
Bioinform.4
2025 Uncovering temporal patterns in visualizations of high-dimensional data
abstract
Abstract With the increasing availability of high-dimensional data, analysts often rely on exploratory data analysis to understand complex data sets. A key approach to exploring such data is dimensionality reduction, which embeds high-dimensional data in two dimensions to enable visual exploration. However, popular embedding techniques, such as t-SNE and UMAP, typically assume that data points are independent. When this assumption is violated, as in time-series data, the resulting visualizations may fail to reveal important temporal patterns and trends. To address this, we propose a formal extension to existing dimensionality reduction methods that incorporates two temporal loss terms that explicitly highlight temporal progression in the embedded visualizations. Through a series of experiments on both synthetic and real-world datasets, we demonstrate that our approach effectively uncovers temporal patterns and improves the interpretability of the visualizations. Furthermore, the method improves temporal coherence while preserving the fidelity of the embeddings, providing a robust tool for dynamic data analysis.
Pavlin Gregor Policar, Blaz Zupan
Mach. Learn.2
2024 Latent Embedding Based on a Transcription-Decay Decomposition of mRNA Dynamics Using Self-supervised CoxPH
Martin Spendl, Tomaz Curk, Blaz Zupan
DS (1)3
2024 Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students
abstract
MOTIVATION: We learn more effectively through experience and reflection than through passive reception of information. Bioinformatics offers an excellent opportunity for project-based learning. Molecular data are abundant and accessible in open repositories, and important concepts in biology can be rediscovered by reanalyzing the data. RESULTS: In the manuscript, we report on five hands-on assignments we designed for master's computer science students to train them in bioinformatics for genomics. These assignments are the cornerstones of our introductory bioinformatics course and are centered around the study of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). They assume no prior knowledge of molecular biology but do require programming skills. Through these assignments, students learn about genomes and genes, discover their composition and function, relate SARS-CoV-2 to other viruses, and learn about the body's response to infection. Student evaluation of the assignments confirms their usefulness and value, their appropriate mastery-level difficulty, and their interesting and motivating storyline. AVAILABILITY AND IMPLEMENTATION: The course materials are freely available on GitHub at https://github.com/IB-ULFRI.
Pavlin Gregor Policar, Martin Spendl, Tomaz Curk, Blaz Zupan
Bioinform.4
2024 Hands-on training about data clustering with orange data mining toolbox
abstract
Data clustering is a core data science approach widely used and referenced in the scientific literature. Its algorithms are often intuitive and can lead to exciting, insightful results that are easy to interpret. For these reasons, data clustering techniques could be the first method encountered in data science training. This paper proposes a hands-on approach to data clustering training suitable for introductory courses. The education approach features problem-based training that starts with the data and gradually introduces various data processing and analysis methods, illustrating them through visual representations of data and models. The proposed training is suitable for a general audience, does not require a background in statistics, mathematics, or computer science, and aims to engage the audience through practical examples, an exploratory approach to data analysis with visual analysis, experimentation, and a gentle learning curve. The manuscript details the pedagogical units of the training, motivates them through the sequence of methods introduced, and proposes data sets and data analysis workflows to be explored in the class.
Janez Demsar, Blaz Zupan
PLoS Comput. Biol.2
2023 Nation-Wide ePrescription Data Reveals Landscape of Physicians and Their Drug Prescribing Patterns in Slovenia
Pavlin Gregor Policar, Dalibor Stanimirovic, Blaz Zupan
AIME3
2023 Ranking of Survival-Related Gene Sets Through Integration of Single-Sample Gene Set Enrichment and Survival Analysis
Martin Spendl, Jaka Kokosar, Ela Praznik, Luka Ausec, Blaz Zupan
AIME5
2023 Gene Interactions in Survival Data Analysis: A Data-Driven Approach Using Restricted Mean Survival Time and Literature Mining
abstract
Abstract Unveiling gene interactions is crucial for comprehending biological processes, particularly their combined impact on phenotypes. Computational methodologies for gene interaction discovery have been extensively studied, but their application to censored data has yet to be thoroughly explored. Our work introduces a data-driven approach to identifying gene interactions that profoundly influence survival rates through the use of survival analysis. Our approach calculates the restricted mean survival time (RMST) for gene pairs and compares it against their individual expressions. If the interaction’s RMST exceeds that of the individual gene expressions, it suggests a potential functional association. We focused on L1000 landmark genes using TCGA na METABRIC data sets. Our findings demonstrate numerous additive and competing interactions and a scarcity of XOR-type interactions. We substantiated our results by cross-referencing with existing interactions in STRING and BioGRID databases and using large language models to summarize complex biological data. Although many potential gene interactions were hypothesized, only a fraction have been experimentally explored. This novel approach enables biologists to initiate a further investigation based on our ranked gene pairs and the generated literature summaries, thus offering a comprehensive, data-driven approach to understanding gene interactions affecting survival rates.
Jaka Kokosar, Martin Spendl, Blaz Zupan
DS3
2023 Refining Temporal Visualizations Using the Directional Coherence Loss
abstract
Abstract Many real-world data sets contain a temporal component or include transitions from state to state. For exploratory data analysis, we can present these high-dimensional data sets in two-dimensional maps, using embeddings of data objects under exploration and representing their temporal relations with directed edges. Most existing dimensionality reduction techniques, such as t-SNE and UMAP, disregard the temporal or relational nature of the data during embedding construction, leading to cluttered visualizations obscuring potentially interesting temporal patterns. To address this issue, we introduce Directional Coherence Loss (DCL), a differentiable loss function that we can incorporate into existing dimensionality reduction techniques. We have designed DCL to highlight the temporal aspects of the data, revealing temporal patterns that might otherwise remain unnoticed. By encouraging local directional coherence of the directed edges, the DCL produces more temporally-meaningful and less-cluttered visualizations. We demonstrate the effectiveness of our approach on a real-world multivariate time-series data set tracking the progression of the COVID-19 pandemic in Slovenia. We show that incorporating the DCL into the t-SNE algorithm elucidates the time progression of the pandemic in the embedding and reveals interesting cyclical patterns otherwise hidden in standard embeddings.
Pavlin Gregor Policar, Blaz Zupan
DS2
2023 Embedding to reference t-SNE space addresses batch effects in single-cell classification
abstract
Abstract Dimensionality reduction techniques, such as t-SNE, can construct informative visualizations of high-dimensional data. When jointly visualising multiple data sets, a straightforward application of these methods often fails; instead of revealing underlying classes, the resulting visualizations expose dataset-specific clusters. To circumvent these batch effects, we propose an embedding procedure that uses a t-SNE visualization constructed on a reference data set as a scaffold for embedding new data points. Each data instance from a new, unseen, secondary data is embedded independently and does not change the reference embedding. This prevents any interactions between instances in the secondary data and implicitly mitigates batch effects. We demonstrate the utility of this approach by analyzing six recently published single-cell gene expression data sets with up to tens of thousands of cells and thousands of genes. The batch effects in our studies are particularly strong as the data comes from different institutions using different experimental protocols. The visualizations constructed by our proposed approach are clear of batch effects, and the cells from secondary data sets correctly co-cluster with cells of the same type from the primary data. We also show the predictive power of our simple, visual classification approach in t-SNE space matches the accuracy of specialized machine learning techniques that consider the entire compendium of features that profile single cells.
Pavlin Gregor Policar, Martin Strazar, Blaz Zupan
Mach. Learn.3
2021 Hands-on training about overfitting
abstract
Overfitting is one of the critical problems in developing models by machine learning. With machine learning becoming an essential technology in computational biology, we must include training about overfitting in all courses that introduce this technology to students and practitioners. We here propose a hands-on training for overfitting that is suitable for introductory level courses and can be carried out on its own or embedded within any data science course. We use workflow-based design of machine learning pipelines, experimentation-based teaching, and hands-on approach that focuses on concepts rather than underlying mathematics. We here detail the data analysis workflows we use in training and motivate them from the viewpoint of teaching goals. Our proposed approach relies on Orange, an open-source data science toolbox that combines data visualization and machine learning, and that is tailored for education in machine learning and explorative data analysis.
Janez Demsar, Blaz Zupan
PLoS Comput. Biol.2
2020 RICERCANDO: Data mining toolkit for mobile broadband measurements
Veljko Pejovic, Ivan Majhen, Miha Janez, Blaz Zupan
Comput. Networks4
2019 Embedding to Reference t-SNE Space Addresses Batch Effects in Single-Cell Classification
Pavlin Gregor Policar, Martin Strazar, Blaz Zupan
DS3
2019 scOrange - a tool for hands-on training of concepts from single-cell data analytics
abstract
MOTIVATION: Single-cell RNA sequencing allows us to simultaneously profile the transcriptomes of thousands of cells and to indulge in exploring cell diversity, development and discovery of new molecular mechanisms. Analysis of scRNA data involves a combination of non-trivial steps from statistics, data visualization, bioinformatics and machine learning. Training molecular biologists in single-cell data analysis and empowering them to review and analyze their data can be challenging, both because of the complexity of the methods and the steep learning curve. RESULTS: We propose a workshop-style training in single-cell data analytics that relies on an explorative data analysis toolbox and a hands-on teaching style. The training relies on scOrange, a newly developed extension of a data mining framework that features workflow design through visual programming and interactive visualizations. Workshops with scOrange can proceed much faster than similar training methods that rely on computer programming and analysis through scripting in R or Python, allowing the trainer to cover more ground in the same time-frame. We here review the design principles of the scOrange toolbox that support such workshops and propose a syllabus for the course. We also provide examples of data analysis workflows that instructors can use during the training. AVAILABILITY AND IMPLEMENTATION: scOrange is an open-source software. The software, documentation and an emerging set of educational videos are available at http://singlecell.biolab.si.
Martin Strazar, Lan Zagar, Jaka Kokosar, Vesna Tanko, Ales Erjavec, Pavlin Gregor Policar, Anze Staric, Janez Demsar, Gad Shaulsky, Vilas Menon, Andrew Lemire, Anup Parikh, Blaz Zupan
Bioinform.13
2017 dictyExpress: a web-based platform for sequence data management and analytics in Dictyostelium and beyond
abstract
BACKGROUND: Dictyostelium discoideum, a soil-dwelling social amoeba, is a model for the study of numerous biological processes. Research in the field has benefited mightily from the adoption of next-generation sequencing for genomics and transcriptomics. Dictyostelium biologists now face the widespread challenges of analyzing and exploring high dimensional data sets to generate hypotheses and discovering novel insights. RESULTS: We present dictyExpress (2.0), a web application designed for exploratory analysis of gene expression data, as well as data from related experiments such as Chromatin Immunoprecipitation sequencing (ChIP-Seq). The application features visualization modules that include time course expression profiles, clustering, gene ontology enrichment analysis, differential expression analysis and comparison of experiments. All visualizations are interactive and interconnected, such that the selection of genes in one module propagates instantly to visualizations in other modules. dictyExpress currently stores the data from over 800 Dictyostelium experiments and is embedded within a general-purpose software framework for management of next-generation sequencing data. dictyExpress allows users to explore their data in a broader context by reciprocal linking with dictyBase-a repository of Dictyostelium genomic data. In addition, we introduce a companion application called GenBoard, an intuitive graphic user interface for data management and bioinformatics analysis. CONCLUSIONS: dictyExpress and GenBoard enable broad adoption of next generation sequencing based inquiries by the Dictyostelium research community. Labs without the means to undertake deep sequencing projects can mine the data available to the public. The entire information flow, from raw sequence data to hypothesis testing, can be accomplished in an efficient workspace. The software framework is generalizable and represents a useful approach for any research community. To encourage more wide usage, the backend is open-source, available for extension and further development by bioinformaticians and data scientists.
Miha Stajdohar, Rafael D. Rosengarten, Janez Kokosar, Luka Jeran, Domen Blenkus, Gad Shaulsky, Blaz Zupan
BMC Bioinform.7
2016 Orthogonal matrix factorization enables integrative analysis of multiple RNA binding proteins
abstract
MOTIVATION: RNA binding proteins (RBPs) play important roles in post-transcriptional control of gene expression, including splicing, transport, polyadenylation and RNA stability. To model protein-RNA interactions by considering all available sources of information, it is necessary to integrate the rapidly growing RBP experimental data with the latest genome annotation, gene function, RNA sequence and structure. Such integration is possible by matrix factorization, where current approaches have an undesired tendency to identify only a small number of the strongest patterns with overlapping features. Because protein-RNA interactions are orchestrated by multiple factors, methods that identify discriminative patterns of varying strengths are needed. RESULTS: We have developed an integrative orthogonality-regularized nonnegative matrix factorization (iONMF) to integrate multiple data sources and discover non-overlapping, class-specific RNA binding patterns of varying strengths. The orthogonality constraint halves the effective size of the factor model and outperforms other NMF models in predicting RBP interaction sites on RNA. We have integrated the largest data compendium to date, which includes 31 CLIP experiments on 19 RBPs involved in splicing (such as hnRNPs, U2AF2, ELAVL1, TDP-43 and FUS) and processing of 3'UTR (Ago, IGF2BP). We show that the integration of multiple data sources improves the predictive accuracy of retrieval of RNA binding sites. In our study the key predictive factors of protein-RNA interactions were the position of RNA structure and sequence motifs, RBP co-binding and gene region type. We report on a number of protein-specific patterns, many of which are consistent with experimentally determined properties of RBPs. AVAILABILITY AND IMPLEMENTATION: The iONMF implementation and example datasets are available at https://github.com/mstrazar/ionmf CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Strazar, Marinka Zitnik, Blaz Zupan, Jernej Ule, Tomaz Curk
Bioinform.3
2016 Jumping across biomedical contexts using compressive data fusion
abstract
MOTIVATION: The rapid growth of diverse biological data allows us to consider interactions between a variety of objects, such as genes, chemicals, molecular signatures, diseases, pathways and environmental exposures. Often, any pair of objects-such as a gene and a disease-can be related in different ways, for example, directly via gene-disease associations or indirectly via functional annotations, chemicals and pathways. Different ways of relating these objects carry different semantic meanings However, traditional methods disregard these semantics and thus cannot fully exploit their value in data modeling. RESULTS: We present Medusa, an approach to detect size-k modules of objects that, taken together, appear most significant to another set of objects. Medusa operates on large-scale collections of heterogeneous datasets and explicitly distinguishes between diverse data semantics. It advances research along two dimensions: it builds on collective matrix factorization to derive different semantics, and it formulates the growing of the modules as a submodular optimization program. Medusa is flexible in choosing or combining semantic meanings and provides theoretical guarantees about detection quality. In a systematic study on 310 complex diseases, we show the effectiveness of Medusa in associating genes with diseases and detecting disease modules. We demonstrate that in predicting gene-disease associations Medusa compares favorably to methods that ignore diverse semantic meanings. We find that the utility of different semantics depends on disease categories and that, overall, Medusa recovers disease modules more accurately when combining different semantics. AVAILABILITY AND IMPLEMENTATION: Source code is at http://github.com/marinkaz/medusa CONTACT: [email protected], [email protected].
Marinka Zitnik, Blaz Zupan
Bioinform.2
2015 Gene network inference by fusing data from diverse distributions
abstract
MOTIVATION: Markov networks are undirected graphical models that are widely used to infer relations between genes from experimental data. Their state-of-the-art inference procedures assume the data arise from a Gaussian distribution. High-throughput omics data, such as that from next generation sequencing, often violates this assumption. Furthermore, when collected data arise from multiple related but otherwise nonidentical distributions, their underlying networks are likely to have common features. New principled statistical approaches are needed that can deal with different data distributions and jointly consider collections of datasets. RESULTS: We present FuseNet, a Markov network formulation that infers networks from a collection of nonidentically distributed datasets. Our approach is computationally efficient and general: given any number of distributions from an exponential family, FuseNet represents model parameters through shared latent factors that define neighborhoods of network nodes. In a simulation study, we demonstrate good predictive performance of FuseNet in comparison to several popular graphical models. We show its effectiveness in an application to breast cancer RNA-sequencing and somatic mutation data, a novel application of graphical models. Fusion of datasets offers substantial gains relative to inference of separate networks for each dataset. Our results demonstrate that network inference methods for non-Gaussian data can help in accurate modeling of the data generated by emergent high-throughput technologies. AVAILABILITY AND IMPLEMENTATION: Source code is at https://github.com/marinkaz/fusenet.
Marinka Zitnik, Blaz Zupan
Bioinform.2
2015 Sieve-based relation extraction of gene regulatory networks from biological literature
abstract
BACKGROUND: Relation extraction is an essential procedure in literature mining. It focuses on extracting semantic relations between parts of text, called mentions. Biomedical literature includes an enormous amount of textual descriptions of biological entities, their interactions and results of related experiments. To extract them in an explicit, computer readable format, these relations were at first extracted manually from databases. Manual curation was later replaced with automatic or semi-automatic tools with natural language processing capabilities. The current challenge is the development of information extraction procedures that can directly infer more complex relational structures, such as gene regulatory networks. RESULTS: We develop a computational approach for extraction of gene regulatory networks from textual data. Our method is designed as a sieve-based system and uses linear-chain conditional random fields and rules for relation extraction. With this method we successfully extracted the sporulation gene regulation network in the bacterium Bacillus subtilis for the information extraction challenge at the BioNLP 2013 conference. To enable extraction of distant relations using first-order models, we transform the data into skip-mention sequences. We infer multiple models, each of which is able to extract different relationship types. Following the shared task, we conducted additional analysis using different system settings that resulted in reducing the reconstruction error of bacterial sporulation network from 0.73 to 0.68, measured as the slot error rate between the predicted and the reference network. We observe that all relation extraction sieves contribute to the predictive performance of the proposed approach. Also, features constructed by considering mention words and their prefixes and suffixes are the most important features for higher accuracy of extraction. Analysis of distances between different mention types in the text shows that our choice of transforming data into skip-mention sequences is appropriate for detecting relations between distant mentions. CONCLUSIONS: Linear-chain conditional random fields, along with appropriate data transformations, can be efficiently used to extract relations. The sieve-based architecture simplifies the system as new sieves can be easily added or removed and each sieve can utilize the results of previous ones. Furthermore, sieves with conditional random fields can be trained on arbitrary text data and hence are applicable to broad range of relation extraction tasks and data domains.
Slavko Zitnik, Marinka Zitnik, Blaz Zupan, Marko Bajec
BMC Bioinform.3
2015 Data Fusion by Matrix Factorization
abstract
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly associated data together with contextual data and data about system's constraints. In the paper we describe a data fusion approach with penalized matrix tri-factorization (DFMF) that simultaneously factorizes data matrices to reveal hidden associations. The approach can directly consider any data that can be expressed in a matrix, including those from feature-based representations, ontologies, associations and networks. We demonstrate the utility of DFMF for gene function prediction task with eleven different data sources and for prediction of pharmacologic actions by fusing six data sources. Our data fusion algorithm compares favorably to alternative data integration approaches and achieves higher accuracy than can be obtained from any single data source alone.
Marinka Zitnik, Blaz Zupan
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Gene Prioritization by Compressive Data Fusion and Chaining
abstract
Data integration procedures combine heterogeneous data sets into predictive models, but they are limited to data explicitly related to the target object type, such as genes. Collage is a new data fusion approach to gene prioritization. It considers data sets of various association levels with the prediction task, utilizes collective matrix factorization to compress the data, and chaining to relate different object types contained in a data compendium. Collage prioritizes genes based on their similarity to several seed genes. We tested Collage by prioritizing bacterial response genes in Dictyostelium as a novel model system for prokaryote-eukaryote interactions. Using 4 seed genes and 14 data sets, only one of which was directly related to the bacterial response, Collage proposed 8 candidate genes that were readily validated as necessary for the response of Dictyostelium to Gram-negative bacteria. These findings establish Collage as a method for inferring biological knowledge from the integration of heterogeneous and coarsely related data sets.
Marinka Zitnik, Edward A. Nam, Christopher Dinh, Adam Kuspa, Gad Shaulsky, Blaz Zupan
PLoS Comput. Biol.6
2015 Integrative Clustering by Nonnegative Matrix Factorization Can Reveal Coherent Functional Groups From Gene Profile Data
abstract
Recent developments in molecular biology and techniques for genome-wide data acquisition have resulted in abundance of data to profile genes and predict their function. These datasets may come from diverse sources and it is an open question how to commonly address them and fuse them into a joint prediction model. A prevailing technique to identify groups of related genes that exhibit similar profiles is profile-based clustering. Cluster inference may benefit from consensus across different clustering models. In this paper, we propose a technique that develops separate gene clusters from each of available data sources and then fuses them by means of nonnegative matrix factorization. We use gene profile data on the budding yeast S. cerevisiae to demonstrate that this approach can successfully integrate heterogeneous datasets and yield high-quality clusters that could otherwise not be inferred by simply merging the gene profiles prior to clustering.
Sanja Brdar, Vladimir S. Crnojevic, Blaz Zupan
IEEE J. Biomed. Health Informatics3
2014 Imputation of Quantitative Genetic Interactions in Epistatic MAPs by Interaction Propagation Matrix Completion
Marinka Zitnik, Blaz Zupan
RECOMB2
2014 Gene network inference by probabilistic scoring of relationships from a factorized model of interactions
abstract
MOTIVATION: Epistasis analysis is an essential tool of classical genetics for inferring the order of function of genes in a common pathway. Typically, it considers single and double mutant phenotypes and for a pair of genes observes whether a change in the first gene masks the effects of the mutation in the second gene. Despite the recent emergence of biotechnology techniques that can provide gene interaction data on a large, possibly genomic scale, few methods are available for quantitative epistasis analysis and epistasis-based network reconstruction. RESULTS: We here propose a conceptually new probabilistic approach to gene network inference from quantitative interaction data. The approach is founded on epistasis analysis. Its features are joint treatment of the mutant phenotype data with a factorized model and probabilistic scoring of pairwise gene relationships that are inferred from the latent gene representation. The resulting gene network is assembled from scored pairwise relationships. In an experimental study, we show that the proposed approach can accurately reconstruct several known pathways and that it surpasses the accuracy of current approaches. AVAILABILITY AND IMPLEMENTATION: Source code is available at http://github.com/biolab/red.
Marinka Zitnik, Blaz Zupan
Bioinform.2
2014 Heterogeneous computing architecture for fast detection of SNP-SNP interactions
abstract
BACKGROUND: The extent of data in a typical genome-wide association study (GWAS) poses considerable computational challenges to software tools for gene-gene interaction discovery. Exhaustive evaluation of all interactions among hundreds of thousands to millions of single nucleotide polymorphisms (SNPs) may require weeks or even months of computation. Massively parallel hardware within a modern Graphic Processing Unit (GPU) and Many Integrated Core (MIC) coprocessors can shorten the run time considerably. While the utility of GPU-based implementations in bioinformatics has been well studied, MIC architecture has been introduced only recently and may provide a number of comparative advantages that have yet to be explored and tested. RESULTS: We have developed a heterogeneous, GPU and Intel MIC-accelerated software module for SNP-SNP interaction discovery to replace the previously single-threaded computational core in the interactive web-based data exploration program SNPsyn. We report on differences between these two modern massively parallel architectures and their software environments. Their utility resulted in an order of magnitude shorter execution times when compared to the single-threaded CPU implementation. GPU implementation on a single Nvidia Tesla K20 runs twice as fast as that for the MIC architecture-based Xeon Phi P5110 coprocessor, but also requires considerably more programming effort. CONCLUSIONS: General purpose GPUs are a mature platform with large amounts of computing power capable of tackling inherently parallel problems, but can prove demanding for the programmer. On the other hand the new MIC architecture, albeit lacking in performance reduces the programming effort and makes it up with a more general architecture suitable for a wider range of problems.
Davor Sluga, Tomaz Curk, Blaz Zupan, Uros Lotric
BMC Bioinform.3
2013 Orange: data mining toolbox in python
Janez Demsar, Tomaz Curk, Ales Erjavec, Crtomir Gorup, Tomaz Hocevar, Mitar Milutinovic, Martin Mozina, Matija Polajnar, Marko Toplak, Anze Staric, Miha Stajdohar, Lan Umek, Lan Zagar, Jure Zbontar, Marinka Zitnik, Blaz Zupan
J. Mach. Learn. Res.16
2012 NIMFA: A Python Library for Nonnegative Matrix Factorization
Marinka Zitnik, Blaz Zupan
J. Mach. Learn. Res.2
2011 A Data Mining Library for miRNA Annotation and Analysis
Angelo Nuzzo, Riccardo Beretta, Francesca Mulas, Valerie Roobrouck, Catherine M. Verfaillie, Blaz Zupan, Riccardo Bellazzi
AIME6
2011 Ranking and 1-Dimensional Projection of Cell Development Transcription Profiles
Lan Zagar, Francesca Mulas, Riccardo Bellazzi, Blaz Zupan
AIME4
2011 Stage prediction of embryonic stem cell differentiation from genome-wide expression data
abstract
MOTIVATION: The developmental stage of a cell can be determined by cellular morphology or various other observable indicators. Such classical markers could be complemented with modern surrogates, like whole-genome transcription profiles, that can encode the state of the entire organism and provide increased quantitative resolution. Recent findings suggest that such profiles provide sufficient information to reliably predict the cell's developmental stage. RESULTS: We use whole-genome transcription data and several data projection methods to infer differentiation stage prediction models for embryonic cells. Given a transcription profile of an uncharacterized cell, these models can then predict its developmental stage. In a series of experiments comprising 14 datasets from the Gene Expression Omnibus, we demonstrate that the approach is robust and has excellent prediction ability both within a specific cell line and across different cell lines. AVAILABILITY: Model inference and computational evaluation procedures in the form of Python scripts and accompanying datasets are available at http://www.biolab.si/supp/stagerank. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lan Zagar, Francesca Mulas, Silvia Garagna, Maurizio Zuccotti, Riccardo Bellazzi, Blaz Zupan
Bioinform.6
2011 Subgroup discovery in data sets with multi-dimensional responses
abstract
Most of the present subgroup discovery approaches aim at finding subsets of attribute-value data with unusual distribution of a single output variable. In general, real-life problems may be described with richer, multi-dimensional descriptions of the outcome. The discovery task in such domains is t o find subsets of data instances with similar outcome description that are separable from the rest of the instances in the input space. We have developed a technique that directly addresses this problem and uses a combination of agglomerative clustering to find subgroup candidates in the space of output attributes, and predictive modeling to score and describe these candidates in the input attribute space. Experiments with the proposed method on a set of synthetic and on a real social survey data set demonstrate its ability to discover relevant and interesting subgroups from the data with multi-dimensional fesponses.
Lan Umek, Blaz Zupan
Intell. Data Anal.2
2010 New components of the Dictyostelium PKA pathway revealed by Bayesian analysis of expression data
abstract
BACKGROUND: Identifying candidate genes in genetic networks is important for understanding regulation and biological function. Large gene expression datasets contain relevant information about genetic networks, but mining the data is not a trivial task. Algorithms that infer Bayesian networks from expression data are powerful tools for learning complex genetic networks, since they can incorporate prior knowledge and uncover higher-order dependencies among genes. However, these algorithms are computationally demanding, so novel techniques that allow targeted exploration for discovering new members of known pathways are essential. RESULTS: Here we describe a Bayesian network approach that addresses a specific network within a large dataset to discover new components. Our algorithm draws individual genes from a large gene-expression repository, and ranks them as potential members of a known pathway. We apply this method to discover new components of the cAMP-dependent protein kinase (PKA) pathway, a central regulator of Dictyostelium discoideum development. The PKA network is well studied in D. discoideum but the transcriptional networks that regulate PKA activity and the transcriptional outcomes of PKA function are largely unknown. Most of the genes highly ranked by our method encode either known components of the PKA pathway or are good candidates. We tested 5 uncharacterized highly ranked genes by creating mutant strains and identified a candidate cAMP-response element-binding protein, yet undiscovered in D. discoideum, and a histidine kinase, a candidate upstream regulator of PKA activity. CONCLUSIONS: The single-gene expansion method is useful in identifying new components of known pathways. The method takes advantage of the Bayesian framework to incorporate prior biological knowledge and discovers higher-order dependencies among genes while greatly reducing the computational resources required to process high-throughput datasets.
Anup Parikh, Eryong Huang, Christopher Dinh, Blaz Zupan, Adam Kuspa, Devika Subramanian, Gad Shaulsky
BMC Bioinform.4
2010 FragViz: visualization of fragmented networks
abstract
BACKGROUND: Researchers in systems biology use network visualization to summarize the results of their analysis. Such networks often include unconnected components, which popular network alignment algorithms place arbitrarily with respect to the rest of the network. This can lead to misinterpretations due to the proximity of otherwise unrelated elements. RESULTS: We propose a new network layout optimization technique called FragViz which can incorporate additional information on relations between unconnected network components. It uses a two-step approach by first arranging the nodes within each of the components and then placing the components so that their proximity in the network corresponds to their relatedness. In the experimental study with the leukemia gene networks we demonstrate that FragViz can obtain network layouts which are more interpretable and hold additional information that could not be exposed using classical network layout optimization algorithms. CONCLUSIONS: Network visualization relies on computational techniques for proper placement of objects under consideration. These algorithms need to be fast so that they can be incorporated in responsive interfaces required by the explorative data analysis environments. Our layout optimization technique FragViz meets these requirements and specifically addresses the visualization of fragmented networks, for which standard algorithms do not consider similarities between unconnected components. The experiments confirmed the claims on speed and accuracy of the proposed solution.
Miha Stajdohar, Minca Mramor, Blaz Zupan, Janez Demsar
BMC Bioinform.3
2009 On Quality of Different Annotation Sources for Gene Expression Analysis
Francesca Mulas, Tomaz Curk, Riccardo Bellazzi, Blaz Zupan
AIME4
2009 Subgroup Discovery in Data Sets with Multi-dimensional Responses: A Method and a Case Study in Traumatology
Lan Umek, Blaz Zupan, Marko Toplak, Annie Morin, Jean-Hugues Chauchat, Gregor Makovec, Dragica Smrke
AIME2
2009 dictyExpress: a Dictyostelium discoideum gene expression database with an explorative data analysis web-based interface
abstract
BACKGROUND: Bioinformatics often leverages on recent advancements in computer science to support biologists in their scientific discovery process. Such efforts include the development of easy-to-use web interfaces to biomedical databases. Recent advancements in interactive web technologies require us to rethink the standard submit-and-wait paradigm, and craft bioinformatics web applications that share analytical and interactive power with their desktop relatives, while retaining simplicity and availability. RESULTS: We have developed dictyExpress, a web application that features a graphical, highly interactive explorative interface to our database that consists of more than 1000 Dictyostelium discoideum gene expression experiments. In dictyExpress, the user can select experiments and genes, perform gene clustering, view gene expression profiles across time, view gene co-expression networks, perform analyses of Gene Ontology term enrichment, and simultaneously display expression profiles for a selected gene in various experiments. Most importantly, these tasks are achieved through web applications whose components are seamlessly interlinked and immediately respond to events triggered by the user, thus providing a powerful explorative data analysis environment. CONCLUSION: dictyExpress is a precursor for a new generation of web-based bioinformatics applications with simple but powerful interactive interfaces that resemble that of the modern desktop. While dictyExpress serves mainly the Dictyostelium research community, it is relatively easy to adapt it to other datasets. We propose that the design ideas behind dictyExpress will influence the development of similar applications for other model organisms.
Gregor Rot, Anup Parikh, Tomaz Curk, Adam Kuspa, Gad Shaulsky, Blaz Zupan
BMC Bioinform.6
2009 Textual features for corpus visualization using correspondence analysis
abstract
Explorative data analysis in text mining essentially relies on effective visualization techniques which can expose hidden relationships among documents and reveal correspondence between documents and their features. In text mining, the documents are
Sasa Petrovic, Bojana Dalbelo Basic, Annie Morin, Blaz Zupan, Jean-Hugues Chauchat
Intell. Data Anal.4
2007 Visualization-based cancer microarray data classification analysis
abstract
MOTIVATION: Methods for analyzing cancer microarray data often face two distinct challenges: the models they infer need to perform well when classifying new tissue samples while at the same time providing an insight into the patterns and gene interactions hidden in the data. State-of-the-art supervised data mining methods often cover well only one of these aspects, motivating the development of methods where predictive models with a solid classification performance would be easily communicated to the domain expert. RESULTS: Data visualization may provide for an excellent approach to knowledge discovery and analysis of class-labeled data. We have previously developed an approach called VizRank that can score and rank point-based visualizations according to degree of separation of data instances of different class. We here extend VizRank with techniques to uncover outliers, score features (genes) and perform classification, as well as to demonstrate that the proposed approach is well suited for cancer microarray analysis. Using VizRank and radviz visualization on a set of previously published cancer microarray data sets, we were able to find simple, interpretable data projections that include only a small subset of genes yet do clearly differentiate among different cancer types. We also report that our approach to classification through visualization achieves performance that is comparable to state-of-the-art supervised data mining techniques. AVAILABILITY: VizRank and radviz are implemented as part of the Orange data mining suite (http://www.ailab.si/orange). SUPPLEMENTARY INFORMATION: Supplementary data are available from http://www.ailab.si/supp/bi-cancer.
Minca Mramor, Gregor Leban, Janez Demsar, Blaz Zupan
Bioinform.4
2007 Polyketide synthase genes and the natural products potential of Dictyostelium discoideum
abstract
MOTIVATION: The genome of the social amoeba Dictyostelium discoideum contains an unusually large number of polyketide synthase (PKS) genes. An analysis of the genes is a first step towards understanding the biological roles of their products and exploiting novel products. RESULTS: A total of 45 Type I iterative PKS genes were found, 5 of which are probably pseudogenes. Catalytic domains that are homologous with known PKS sequences as well as possible novel domains were identified. The genes often occurred in clusters of 2-5 genes, where members of the cluster had very similar sequences. The D.discoideum PKS genes formed a clade distinct from fungal and bacterial genes. All nine genes examined by RT-PCR were expressed, although at different developmental stages. The promoters of PKS genes were much more divergent than the structural genes, although we have identified motifs that are unique to some PKS gene promoters.
Jurica Zucko, N. Skunca, Tomaz Curk, Blaz Zupan, Paul F. Long, John Cullum, R. H. Kessin, Daslav Hranueli
Bioinform.4
2007 Towards knowledge-based gene expression data mining
Riccardo Bellazzi, Blaz Zupan
J. Biomed. Informatics2
2007 FreeViz - An intelligent multivariate visualization approach to explorative analysis of biomedical data
Janez Demsar, Gregor Leban, Blaz Zupan
J. Biomed. Informatics3
2006 Knowledge-based data analysis and interpretation
Blaz Zupan, John H. Holmes, Riccardo Bellazzi
Artif. Intell. Medicine1
2006 VizRank: Data Visualization Guided by Machine Learning
Gregor Leban, Blaz Zupan, Gaj Vidmar, Ivan Bratko
Data Min. Knowl. Discov.2
2006 Spam Filtering Using Statistical Data Compression Models
abstract
Spam filtering poses a special problem in text categorization, of which the defining characteristic is that filters face an active adversary, which constantly attempts to evade filtering. Since spam evolves continuously and most practical applications are based on online user feedback, the task calls for fast, incremental and robust learning algorithms. In this paper, we investigate a novel approach to spam filtering based on adaptive statistical data compression models. The nature of these models allows them to be employed as probabilistic text classifiers based on character-level or binary sequences. By modeling messages as sequences, tokenization and other error-prone preprocessing steps are omitted altogether, resulting in a method that is very robust. The models are also fast to construct and incrementally updateable. We evaluate the filtering performance of two different compression algorithms; dynamic Markov compression and prediction by partial matching. The results of our empirical evaluation indicate that compression models outperform currently established spam filters, as well as a number of methods proposed in previous studies.
Andrej Bratko, Gordon V. Cormack, Bogdan Filipic, Thomas R. Lynam, Blaz Zupan
J. Mach. Learn. Res.5
2005 Conquering the Curse of Dimensionality in Gene Expression Cancer Diagnosis: Tough Problem, Simple Models
Minca Mramor, Gregor Leban, Janez Demsar, Blaz Zupan
AIME4
2005 Nomograms for visualizing support vector machines
abstract
We propose a simple yet potentially very effective way of visualizing trained support vector machines. Nomograms are an established model visualization technique that can graphically encode the complete model on a single page. The dimensionality of the visualization does not depend on the number of attributes, but merely on the properties of the kernel. To represent the effect of each predictive feature on the log odds ratio scale as required for the nomograms, we employ logistic regression to convert the distance from the separating hyperplane into a probability. Case studies on selected data sets show that for a technique thought to be a black-box, nomograms can clearly expose its internal structure. By providing an easy-to-interpret visualization the analysts can gain insight and study the effects of predictive factors.
Aleks Jakulin, Martin Mozina, Janez Demsar, Ivan Bratko, Blaz Zupan
KDD5
2005 Simple and effective visual models for gene expression cancer diagnostics
abstract
In the paper we show that diagnostic classes in cancer gene expression data sets, which most often include thousands of features (genes), may be effectively separated with simple two-dimensional plots such as scatterplot and radviz graph. The principal innovation proposed in the paper is a method called VizRank, which is able to score and identify the best among possibly millions of candidate projections for visualizations. Compared to recently much applied techniques in the field of cancer genomics that include neural networks, support vector machines and various ensemble-based approaches, VizRank is fast and finds visualization models that can be easily examined and interpreted by domain experts. Our experiments on a number of gene expression data sets show that VizRank was always able to find data visualizations with a small number of (two to seven) genes and excellent class separation. In addition to providing grounds for gene expression cancer diagnosis, VizRank and its visualizations also identify small sets of relevant genes, uncover interesting gene interactions and point to outliers and potential misclassifications in cancer data sets.
Gregor Leban, Minca Mramor, Ivan Bratko, Blaz Zupan
KDD4
2005 Microarray data mining with visual programming
abstract
UNLABELLED: Visual programming offers an intuitive means of combining known analysis and visualization methods into powerful applications. The system presented here enables users who are not programmers to manage microarray and genomic data flow and to customize their analyses by combining common data analysis tools to fit their needs. AVAILABILITY: http://www.ailab.si/supp/bi-visprog SUPPLEMENTARY INFORMATION: http://www.ailab.si/supp/bi-visprog.
Tomaz Curk, Janez Demsar, Qikai Xu, Gregor Leban, Uros Petrovic, Ivan Bratko, Gad Shaulsky, Blaz Zupan
Bioinform.8
2005 VizRank: finding informative data projections in functional genomics by machine learning
abstract
UNLABELLED: VizRank is a tool that finds interesting two-dimensional projections of class-labeled data. When applied to multi-dimensional functional genomics datasets, VizRank can systematically find relevant biological patterns. AVAILABILITY: http://www.ailab.si/supp/bi-vizrank SUPPLEMENTARY INFORMATION: http://www.ailab.si/supp/bi-vizrank.
Gregor Leban, Ivan Bratko, Uros Petrovic, Tomaz Curk, Blaz Zupan
Bioinform.5
2004 Orange: From Experimental Machine Learning to Interactive Data Mining
Janez Demsar, Blaz Zupan, Gregor Leban, Tomaz Curk
PKDD2
2004 Nomograms for Visualization of Naive Bayesian Classifier
Martin Mozina, Janez Demsar, Michael W. Kattan, Blaz Zupan
PKDD4
2004 A function-decomposition method for development of hierarchical multi-attribute decision models
Marko Bohanec, Blaz Zupan
Decis. Support Syst.2
2003 Attribute Interactions in Medical Data Analysis
Aleks Jakulin, Ivan Bratko, Dragica Smrke, Janez Demsar, Blaz Zupan
AIME5
2003 GenePath: a system for inference of genetic networks and proposal of genetic experiments
Blaz Zupan, Ivan Bratko, Janez Demsar, Peter Juvan, Tomaz Curk, Urban Borstnik, J. Robert Beck, John A. Halter, Adam Kuspa, Gad Shaulsky
Artif. Intell. Medicine1
2003 GenePath: a system for automated construction of genetic networks from mutant data
abstract
MOTIVATION: Genetic networks are often used in the analysis of biological phenomena. In classical genetics, they are constructed manually from experimental data on mutants. The field lacks formalism to guide such analysis, and accounting for all the data becomes complicated when large amounts of data are considered. RESULTS: We have developed GenePath, an intelligent assistant that automates the analysis of genetic data. GenePath employs expert-defined patterns to uncover gene relations from the data, and uses these relations as constraints in the search for a plausible genetic network. GenePath formalizes genetic data analysis, facilitates the consideration of all the available data in a consistent manner, and the examination of the large number of possible consequences of planned experiments. It also provides an explanation mechanism that traces every finding to the pertinent data. AVAILABILITY: GenePath can be accessed at http://genepath.org. SUPPLEMENTARY INFORMATION: Supplementary material is available at http://genepath.org/bi-.supp.
Blaz Zupan, Janez Demsar, Ivan Bratko, Peter Juvan, John A. Halter, Adam Kuspa, Gad Shaulsky
Bioinform.1
2002 Prognostic Model for Terminally Ill Cancer Patients on Laboratory Data
Kenji Hira, Noriaki Aoki, Akitoshi Hayashi, Janez Demsar, Blaz Zupan, Kim Dunn, William Jack Schull, Tsuguya Fukui
AMIA5
2001 Abductive Inference of Genetic Networks
Blaz Zupan, Ivan Bratko, Janez Demsar, J. Robert Beck, Adam Kuspa, Gad Shaulsky
AIME1
2000 Induction of Concept Hierarchies from Noisy Data
Blaz Zupan, Ivan Bratko, Marko Bohanec, Janez Demsar
ICML1
2000 Machine learning for survival analysis: a case study on recurrence of prostate cancer
Blaz Zupan, Janez Demsar, Michael W. Kattan, J. Robert Beck, Ivan Bratko
Artif. Intell. Medicine1
1999 Learning by Discovering Concept Hierarchies
Blaz Zupan, Marko Bohanec, Janez Demsar, Ivan Bratko
Artif. Intell.1
1998 Acquiring background knowledge for machine learning using function decomposition: a case study in rheumatology
Blaz Zupan, Saso Dzeroski
Artif. Intell. Medicine1
1998 Optimization of Rule-Based Systems Using State Space Graphs
abstract
Embedded rule-based expert systems must satisfy stringent timing constraints when applied to real-time environments. The paper describes a novel approach to reduce the response time of rule-based expert systems. The optimization method is based on a construction of the reduced cycle-free finite state space graph. In contrast with traditional state space graph derivation, the optimization algorithm starts from the final states (fixed points) and gradually expands the state space graph until all of the states with a reachable fixed point are found. The new and optimized system is then synthesized from the constructed state space graph. The authors present several algorithms implementing the optimization method. They vary in complexity as well as in the usage of concurrency and state-equivalency-both targeted toward minimizing the size of the optimized state space graph. Though depending on the algorithm used, optimized rule-based systems: (1) in general have better response time in that they require fewer rule firings to reach the fixed point; (2) are stable, i.e., have no cycles that would result in the instability of execution; and (3) have no redundant rules. They also address the issue of deterministic execution and propose optimization algorithms that generate the rule-bases with single corresponding fixed points for every initial state. The synthesis method also determines the tight response time bound of the new system and can identify unstable states in the original rule-base.
Blaz Zupan, Albert Mo Kim Cheng
IEEE Trans. Knowl. Data Eng.1
1997 Acquiring and Validating Background Knowledge for Machine Learning Using Function Decomposition
Blaz Zupan, Saso Dzeroski
AIME1
1997 Relating clinical and neurophysiological assessment of spasticity by machine learning
abstract
Spasticity following spinal cord injury (SCI) is most often assessed clinically using a five point Ashworth Score (AS). A more objective assessment of altered motor control may be achieved by using a comprehensive protocol based on a surface electromyographic (sEMG) activity recorded from thigh and leg muscles. However, the relation between clinical and neurophysiological assessments is still unknown. We employ three different classification methods to investigate this relationship. The experimental results indicate that if the appropriate set of sEMG features is used, the neurophysiological assessment is related to clinical findings and can be used to predict the AS. A comprehensive and objective sEMG assessment may be proven useful for the assessment of interventions and follow up of SCI patients.
Blaz Zupan, Dobrivoje S. Stokic, Marko Bohanec, Michael M. Priebe, Arthur M. Sherwood
CBMS1
1997 Constructing Intermediate Concepts by Decomposition of Real Functions
Janez Demsar, Blaz Zupan, Marko Bohanec, Ivan Bratko
ECML2
1997 Machine Learning by Function Decomposition
Blaz Zupan, Marko Bohanec, Ivan Bratko, Janez Demsar
ICML1
1997 A Dataset Decomposition Approach to Data Mining and Machine Discovery
Blaz Zupan, Marko Bohanec, Ivan Bratko, Bojan Cestnik
KDD1