Tomaz Curk

dblp:69/3710 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-4888-7256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Computing education · 48% Bioinformatics and computational biology · 33% Computational science and engineering · 19%
Databases, data mining, and information retrieval
1 paper
Data mining · 77% Data integration and cleaning · 23%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education › STEM education
bioinformatics education
1.022025
Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students · Bioinform. 2024
Automated assignment grading with large language models: insights from a bioinformatics course · Bioinform. 2025
Computing education
automated assessment
0.912025
Automated assignment grading with large language models: insights from a bioinformatics course · Bioinform. 2025
Computational science and engineering › active learning
project-based learning
0.812024
Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students · Bioinform. 2024
Bioinformatics and computational biology › sequence analysis › RNA sequence analysis
RNA-binding site prediction
0.212016
Orthogonal matrix factorization enables integrative analysis of multiple RNA binding proteins · Bioinform. 2016
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA-protein interaction prediction
0.212016
Orthogonal matrix factorization enables integrative analysis of multiple RNA binding proteins · Bioinform. 2016
Bioinformatics and computational biology
genomics
0.212024
Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students · Bioinform. 2024
Bioinformatics and computational biology › genomics › viral genomics
SARS-CoV-2 genome analysis
0.212024
Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students · Bioinform. 2024
Bioinformatics and computational biology › computational microbiology
biosynthetic gene cluster analysis
0.112007
Polyketide synthase genes and the natural products potential of Dictyostelium discoideum · Bioinform. 2007
Bioinformatics and computational biology › genomics
genome analysis
0.112007
Polyketide synthase genes and the natural products potential of Dictyostelium discoideum · Bioinform. 2007
Bioinformatics and computational biology › drug discovery
natural product discovery
0.112007
Polyketide synthase genes and the natural products potential of Dictyostelium discoideum · Bioinform. 2007
Bioinformatics and computational biology
functional genomics
0.112005
VizRank: finding informative data projections in functional genomics by machine learning · Bioinform. 2005
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.112005
Microarray data mining with visual programming · Bioinform. 2005
Data integration and cleaning
data preprocessing
0.012013
Orange: data mining toolbox in python · J. Mach. Learn. Res. 2013
Visualization and visual analytics
dimensionality reduction
0.012005
VizRank: finding informative data projections in functional genomics by machine learning · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

prompt engineering · 0.9large language model · 0.9project-based learning · 0.8orthogonality regularization · 0.2non-negative matrix factorization · 0.2sequence homology analysis · 0.1phylogenetic analysis · 0.1RT-PCR · 0.1visualization · 0.1projection ranking · 0.1machine learning · 0.1data flow · 0.1
YearPublicationVenuePosition
2025 Automated assignment grading with large language models: insights from a bioinformatics course
abstract
MOTIVATION: Providing students with individualized feedback through assignments is a cornerstone of education that supports their learning and development. Studies have shown that timely, high-quality feedback plays a critical role in improving learning outcomes. However, providing personalized feedback on a large scale in classes with large numbers of students is often impractical due to the significant time and effort required. Recent advances in natural language processing and large language models (LLMs) offer a promising solution by enabling the efficient delivery of personalized feedback. These technologies can reduce the workload of course staff while improving student satisfaction and learning outcomes. Their successful implementation, however, requires thorough evaluation and validation in real classrooms. RESULTS: We present the results of a practical evaluation of LLM-based graders for written assignments in the 2024/25 iteration of the Introduction to Bioinformatics course at the University of Ljubljana. Over the course of the semester, more than 100 students answered 36 text-based questions, most of which were automatically graded using LLMs. In a blind study, students received feedback from both LLMs and human teaching assistants (TAs) without knowing the source, and later rated the quality of the feedback. We conducted a systematic evaluation of six commercial and open-source LLMs and compared their grading performance with human TAs. Our results show that with well-designed prompts, LLMs can achieve grading accuracy and feedback quality comparable to human graders. Our results also suggest that open-source LLMs perform as well as commercial LLMs, allowing schools to implement their own grading systems while maintaining privacy.
Pavlin Gregor Policar, Martin Spendl, Tomaz Curk, Blaz Zupan
Bioinform.3
2024 Latent Embedding Based on a Transcription-Decay Decomposition of mRNA Dynamics Using Self-supervised CoxPH
Martin Spendl, Tomaz Curk, Blaz Zupan
DS (1)2
2024 Teaching bioinformatics through the analysis of SARS-CoV-2: project-based training for computer science students
abstract
MOTIVATION: We learn more effectively through experience and reflection than through passive reception of information. Bioinformatics offers an excellent opportunity for project-based learning. Molecular data are abundant and accessible in open repositories, and important concepts in biology can be rediscovered by reanalyzing the data. RESULTS: In the manuscript, we report on five hands-on assignments we designed for master's computer science students to train them in bioinformatics for genomics. These assignments are the cornerstones of our introductory bioinformatics course and are centered around the study of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). They assume no prior knowledge of molecular biology but do require programming skills. Through these assignments, students learn about genomes and genes, discover their composition and function, relate SARS-CoV-2 to other viruses, and learn about the body's response to infection. Student evaluation of the assignments confirms their usefulness and value, their appropriate mastery-level difficulty, and their interesting and motivating storyline. AVAILABILITY AND IMPLEMENTATION: The course materials are freely available on GitHub at https://github.com/IB-ULFRI.
Pavlin Gregor Policar, Martin Spendl, Tomaz Curk, Blaz Zupan
Bioinform.3
2021 Sparse data embedding and prediction by tropical matrix factorization
abstract
BACKGROUND: Matrix factorization methods are linear models, with limited capability to model complex relations. In our work, we use tropical semiring to introduce non-linearity into matrix factorization models. We propose a method called Sparse Tropical Matrix Factorization (STMF) for the estimation of missing (unknown) values in sparse data. RESULTS: We evaluate the efficiency of the STMF method on both synthetic data and biological data in the form of gene expression measurements downloaded from The Cancer Genome Atlas (TCGA) database. Tests on unique synthetic data showed that STMF approximation achieves a higher correlation than non-negative matrix factorization (NMF), which is unable to recover patterns effectively. On real data, STMF outperforms NMF on six out of nine gene expression datasets. While NMF assumes normal distribution and tends toward the mean value, STMF can better fit to extreme values and distributions. CONCLUSION: STMF is the first work that uses tropical semiring on sparse data. We show that in certain cases semirings are useful because they consider the structure, which is different and simpler to understand than it is with standard linear algebra.
Amra Omanovic, Hilal Kazan, Polona Oblak, Tomaz Curk
BMC Bioinform.4
2020 Relation chaining in binary positive-only recommender systems
Rok Gomiscek, Tomaz Curk
Expert Syst. Appl.2
2020 Simultaneous incremental matrix factorization for streaming recommender systems
Martin Jakomin, Zoran Bosnic, Tomaz Curk
Expert Syst. Appl.3
2019 Approximate multiple kernel learning with least-angle regression
abstract
Kernel methods provide a principled way for general data representations. Multiple kernel learning and kernel approximation are often treated as separate tasks, with considerable savings in time and memory expected if the two are performed simultaneously. Our proposed Mklaren algorithm selectively approximates multiple kernel matrices in regression. It uses Incomplete Cholesky Decomposition and Least-angle regression (LAR) to select basis functions, achieving linear complexity both in the number of data points and kernels. Since it approximates kernel matrices rather than functions, it allows to combine an arbitrary set of kernels. Compared to single kernel-based approximations, it selectively approximates different kernels in different regions of the input spaces. The LAR criterion provides a robust selection of inducing points in noisy settings, and an accurate modelling of regression functions in continuous and discrete input spaces. Among general kernel matrix decompositions, Mklaren achieves minimal approximation rank required for performance comparable to using the exact kernel matrix, at a cost lower than 1% of required operations. Finally, we demonstrate the scalability and interpretability in settings with millions of data points and thousands of kernels.
Martin Strazar, Tomaz Curk
Neurocomputing2
2016 Orthogonal matrix factorization enables integrative analysis of multiple RNA binding proteins
abstract
MOTIVATION: RNA binding proteins (RBPs) play important roles in post-transcriptional control of gene expression, including splicing, transport, polyadenylation and RNA stability. To model protein-RNA interactions by considering all available sources of information, it is necessary to integrate the rapidly growing RBP experimental data with the latest genome annotation, gene function, RNA sequence and structure. Such integration is possible by matrix factorization, where current approaches have an undesired tendency to identify only a small number of the strongest patterns with overlapping features. Because protein-RNA interactions are orchestrated by multiple factors, methods that identify discriminative patterns of varying strengths are needed. RESULTS: We have developed an integrative orthogonality-regularized nonnegative matrix factorization (iONMF) to integrate multiple data sources and discover non-overlapping, class-specific RNA binding patterns of varying strengths. The orthogonality constraint halves the effective size of the factor model and outperforms other NMF models in predicting RBP interaction sites on RNA. We have integrated the largest data compendium to date, which includes 31 CLIP experiments on 19 RBPs involved in splicing (such as hnRNPs, U2AF2, ELAVL1, TDP-43 and FUS) and processing of 3'UTR (Ago, IGF2BP). We show that the integration of multiple data sources improves the predictive accuracy of retrieval of RNA binding sites. In our study the key predictive factors of protein-RNA interactions were the position of RNA structure and sequence motifs, RBP co-binding and gene region type. We report on a number of protein-specific patterns, many of which are consistent with experimentally determined properties of RBPs. AVAILABILITY AND IMPLEMENTATION: The iONMF implementation and example datasets are available at https://github.com/mstrazar/ionmf CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Strazar, Marinka Zitnik, Blaz Zupan, Jernej Ule, Tomaz Curk
Bioinform.5
2014 Heterogeneous computing architecture for fast detection of SNP-SNP interactions
abstract
BACKGROUND: The extent of data in a typical genome-wide association study (GWAS) poses considerable computational challenges to software tools for gene-gene interaction discovery. Exhaustive evaluation of all interactions among hundreds of thousands to millions of single nucleotide polymorphisms (SNPs) may require weeks or even months of computation. Massively parallel hardware within a modern Graphic Processing Unit (GPU) and Many Integrated Core (MIC) coprocessors can shorten the run time considerably. While the utility of GPU-based implementations in bioinformatics has been well studied, MIC architecture has been introduced only recently and may provide a number of comparative advantages that have yet to be explored and tested. RESULTS: We have developed a heterogeneous, GPU and Intel MIC-accelerated software module for SNP-SNP interaction discovery to replace the previously single-threaded computational core in the interactive web-based data exploration program SNPsyn. We report on differences between these two modern massively parallel architectures and their software environments. Their utility resulted in an order of magnitude shorter execution times when compared to the single-threaded CPU implementation. GPU implementation on a single Nvidia Tesla K20 runs twice as fast as that for the MIC architecture-based Xeon Phi P5110 coprocessor, but also requires considerably more programming effort. CONCLUSIONS: General purpose GPUs are a mature platform with large amounts of computing power capable of tackling inherently parallel problems, but can prove demanding for the programmer. On the other hand the new MIC architecture, albeit lacking in performance reduces the programming effort and makes it up with a more general architecture suitable for a wider range of problems.
Davor Sluga, Tomaz Curk, Blaz Zupan, Uros Lotric
BMC Bioinform.2
2013 Orange: data mining toolbox in python
Janez Demsar, Tomaz Curk, Ales Erjavec, Crtomir Gorup, Tomaz Hocevar, Mitar Milutinovic, Martin Mozina, Matija Polajnar, Marko Toplak, Anze Staric, Miha Stajdohar, Lan Umek, Lan Zagar, Jure Zbontar, Marinka Zitnik, Blaz Zupan
J. Mach. Learn. Res.2
2009 On Quality of Different Annotation Sources for Gene Expression Analysis
Francesca Mulas, Tomaz Curk, Riccardo Bellazzi, Blaz Zupan
AIME2
2009 dictyExpress: a Dictyostelium discoideum gene expression database with an explorative data analysis web-based interface
abstract
BACKGROUND: Bioinformatics often leverages on recent advancements in computer science to support biologists in their scientific discovery process. Such efforts include the development of easy-to-use web interfaces to biomedical databases. Recent advancements in interactive web technologies require us to rethink the standard submit-and-wait paradigm, and craft bioinformatics web applications that share analytical and interactive power with their desktop relatives, while retaining simplicity and availability. RESULTS: We have developed dictyExpress, a web application that features a graphical, highly interactive explorative interface to our database that consists of more than 1000 Dictyostelium discoideum gene expression experiments. In dictyExpress, the user can select experiments and genes, perform gene clustering, view gene expression profiles across time, view gene co-expression networks, perform analyses of Gene Ontology term enrichment, and simultaneously display expression profiles for a selected gene in various experiments. Most importantly, these tasks are achieved through web applications whose components are seamlessly interlinked and immediately respond to events triggered by the user, thus providing a powerful explorative data analysis environment. CONCLUSION: dictyExpress is a precursor for a new generation of web-based bioinformatics applications with simple but powerful interactive interfaces that resemble that of the modern desktop. While dictyExpress serves mainly the Dictyostelium research community, it is relatively easy to adapt it to other datasets. We propose that the design ideas behind dictyExpress will influence the development of similar applications for other model organisms.
Gregor Rot, Anup Parikh, Tomaz Curk, Adam Kuspa, Gad Shaulsky, Blaz Zupan
BMC Bioinform.3
2007 Polyketide synthase genes and the natural products potential of Dictyostelium discoideum
abstract
MOTIVATION: The genome of the social amoeba Dictyostelium discoideum contains an unusually large number of polyketide synthase (PKS) genes. An analysis of the genes is a first step towards understanding the biological roles of their products and exploiting novel products. RESULTS: A total of 45 Type I iterative PKS genes were found, 5 of which are probably pseudogenes. Catalytic domains that are homologous with known PKS sequences as well as possible novel domains were identified. The genes often occurred in clusters of 2-5 genes, where members of the cluster had very similar sequences. The D.discoideum PKS genes formed a clade distinct from fungal and bacterial genes. All nine genes examined by RT-PCR were expressed, although at different developmental stages. The promoters of PKS genes were much more divergent than the structural genes, although we have identified motifs that are unique to some PKS gene promoters.
Jurica Zucko, N. Skunca, Tomaz Curk, Blaz Zupan, Paul F. Long, John Cullum, R. H. Kessin, Daslav Hranueli
Bioinform.3
2005 Microarray data mining with visual programming
abstract
UNLABELLED: Visual programming offers an intuitive means of combining known analysis and visualization methods into powerful applications. The system presented here enables users who are not programmers to manage microarray and genomic data flow and to customize their analyses by combining common data analysis tools to fit their needs. AVAILABILITY: http://www.ailab.si/supp/bi-visprog SUPPLEMENTARY INFORMATION: http://www.ailab.si/supp/bi-visprog.
Tomaz Curk, Janez Demsar, Qikai Xu, Gregor Leban, Uros Petrovic, Ivan Bratko, Gad Shaulsky, Blaz Zupan
Bioinform.1
2005 VizRank: finding informative data projections in functional genomics by machine learning
abstract
UNLABELLED: VizRank is a tool that finds interesting two-dimensional projections of class-labeled data. When applied to multi-dimensional functional genomics datasets, VizRank can systematically find relevant biological patterns. AVAILABILITY: http://www.ailab.si/supp/bi-vizrank SUPPLEMENTARY INFORMATION: http://www.ailab.si/supp/bi-vizrank.
Gregor Leban, Ivan Bratko, Uros Petrovic, Tomaz Curk, Blaz Zupan
Bioinform.4
2004 Orange: From Experimental Machine Learning to Interactive Data Mining
Janez Demsar, Blaz Zupan, Gregor Leban, Tomaz Curk
PKDD4
2003 GenePath: a system for inference of genetic networks and proposal of genetic experiments
Blaz Zupan, Ivan Bratko, Janez Demsar, Peter Juvan, Tomaz Curk, Urban Borstnik, J. Robert Beck, John A. Halter, Adam Kuspa, Gad Shaulsky
Artif. Intell. Medicine5