Peter Dawyndt

dblp:95/1229 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-1623-9070ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Source Code Plagiarism Detection as a Service with Dolos
abstract
Source code similarity detection tools are crucial for preventing and identifying plagiarism in programming courses. While tools like JPlag and Moss are effective, their complexity often hinders widespread adoption. To address this, we developed Dolos, a user-friendly source code similarity detection pipeline that enhances the user experience with interactive dashboards. The Dolos ecosystem includes software libraries, a CLI tool, a web UI, a web server, and API server enabling seamless integration with educational environments. The Dolos web server allows instructors to perform plagiarism detection directly within their browsers, eliminating the need for additional software installation. A publicly available instance and the option for private instances further enhance accessibility. Dolos' API facilitates integration as a microservice within programming exercise platforms like Dodona, A+ and Radar, and Codio, showcasing its versatility and effectiveness. At Ghent University, Dolos is integral to the plagiarism detection strategy, with visualisations serving as both a detection tool and a deterrent. The focus on user experience, flexibility, and comprehensive documentation has attracted scholars to use Dolos for innovative applications, even outside of the educational field. We invite instructors to explore Dolos and integrate it within their educational platforms for programming assignments.
Rien Maertens, Peter Dawyndt, Bart Mesuere
ITiCSE (2)2
2025 Are LLMs Good at Answering Student Questions in CS1 Courses?
abstract
Educators often spend a significant amount of time answering student coding questions, which can lead to rushed or incomplete responses. Additionally, generative AI tools are starting to and will play an increasing role in students' careers. These tools are easy to use and can provide heaps of information almost instantly. However, they often generate answers that provide full assignment solutions rather than guiding students towards the correct solution, which can be detrimental to their learning process. To address this issue, we are exploring the potential of large language models (LLMs) in improving student support by generating draft responses. These drafts are designed to provide students with meaningful guidance without giving away direct solutions. To evaluate these drafts, an LLM-as-a-judge is employed to compare the generated answers from various LLMs, different prompts, and best available human answers against a ground truth dataset. We present the evaluation process using an LLM-as-a-judge based benchmark, discuss the results obtained by different models and prompts, and compare them to the best available human responses. These evaluations give an indication of how LLMs can aid in computer science education, by reducing the time needed to answer questions and increasing both the accuracy and effectiveness of responses.
Thomas Van Mullem, Bart Mesuere, Peter Dawyndt
ITiCSE (2)3
2025 Direct construction of sparse suffix arrays with Libsais
abstract
BACKGROUND: Pattern matching is a fundamental challenge in bioinformatics, especially in the fields of genomics, transcriptomics and proteomics. Efficient indexing structures, such as suffix arrays, are critical for searching large datasets. A sparse suffix array (SSA) retains only suffixes at every k-th position in the text, where k is the sparseness factor. While sparse suffix arrays offer significant memory savings compared to full suffix arrays, they typically still require the construction of a full suffix array prior to a sampling step, resulting in substantial memory overhead during the construction phase. RESULTS: We present an alternative method to directly construct the sparse suffix array using a simple, yet powerful text encoding. This encoding reduces the input text length by grouping characters, thereby enabling direct SSA construction by extending the widely used Libsais library. This approach bypasses the need to construct a full suffix array, reducing memory usage and construction time by 50 to 75% when building a sparse suffix array with sparseness factor 3 or 4 for various nucleotide and amino acid datasets. Depending on the alphabet size, similar gains can be achieved for sparseness factors up to 8. For higher sparseness factors, comparable performance improvements can be obtained by constructing the SSA using a suitable divisor of the desired sparseness factor, followed by a subsampling step. The method is particularly effective for applications with small alphabets, such as a nucleotide or amino acid alphabet. An open-source implementation of this method is available on GitHub, enabling easy adoption for large-scale bioinformatics applications. CONCLUSIONS: We introduce an efficient method for the construction of sparse suffix arrays for large datasets. Central to this approach is the introduction of a simple text transformation, which then serves as input to Libsais. This method reduces the length of both the input text and the resulting suffix array by a factor of k, which improves execution time and memory usage significantly.
Simon Van de Vyver, Tibo Vande Moortele, Peter Dawyndt, Bart Mesuere, Pieter Verschaffelt
BMC Bioinform.3
2023 Dolos 2.0: Towards Seamless Source Code Plagiarism Detection in Online Learning Environments
abstract
With the increasing demand for programming skills comes a trend towards more online programming courses and assessments. While this allows educators to teach larger groups of students, it also opens the door to dishonest student behaviour, such as copying code from other students. When teachers use assignments where all students write code for the same problem, source code similarity tools can help to combat plagiarism. Unfortunately, teachers often do not use these tools to prevent such behaviour.
Rien Maertens, Peter Dawyndt, Bart Mesuere
ITiCSE (2)2
2023 Dodona: Learn to Code with a Virtual Co-teacher that Supports Active Learning
abstract
Dodona (dodona.ugent.be) is an intelligent tutoring system for learning computer programming, statistics and data science. It bridges the gap between assessment and learning by providing real-time data and feedback to help students learn better, teachers teach better and educational technology become more effective.
Charlotte Van Petegem, Peter Dawyndt, Bart Mesuere
ITiCSE (2)2
2023 Blink: An Educational Software Debugger for Scratch
abstract
Debugging is an important aspect of programming. Most programming languages have some features and tools to facilitate debugging. As the debugging process is also frustrating, it requires good scaffolding, in which a debugger can be a useful tool [3]. Scratch is a visual block-based programming language that is commonly used to teach programming to children, aged 10--14 [4]. It comes with its own integrated development environment (IDE), where children can edit and run their code. This IDE misses some of the tools that are available in traditional IDEs, such as a debugger. In response to this challenge, we developed Blink. Blink is a debugger for Scratch with the aim of being usable to the young audience that typically uses Scratch.
Niko Strijbol, Christophe Scholliers, Peter Dawyndt
ITiCSE (2)3
2022 Unipept Visualizations: an interactive visualization library for biological data
abstract
SUMMARY: The Unipept Visualizations library is a JavaScript package to generate interactive visualizations of both hierarchical and non-hierarchical quantitative data. It provides four different visualizations: a sunburst, a treemap, a treeview and a heatmap. Every visualization is fully configurable, supports TypeScript and uses the excellent D3.js library. AVAILABILITY AND IMPLEMENTATION: The Unipept Visualizations library is available for download on NPM: https://npmjs.com/unipept-visualizations. All source code is freely available from GitHub under the MIT license: https://github.com/unipept/unipept-visualizations.
Pieter Verschaffelt, James H. Collier, Alexander Botzki, Lennart Martens, Peter Dawyndt, Bart Mesuere
Bioinform.5
2022 FragGeneScanRs: faster gene prediction for short reads
abstract
BACKGROUND: FragGeneScan is currently the most accurate and popular tool for gene prediction in short and error-prone reads, but its execution speed is insufficient for use on larger data sets. The parallelization which should have addressed this is inefficient. Its alternative implementation FragGeneScan+ is faster, but introduced a number of bugs related to memory management, race conditions and even output accuracy. RESULTS: This paper introduces FragGeneScanRs, a faster Rust implementation of the FragGeneScan gene prediction model. Its command line interface is backward compatible and adds extra features for more flexible usage. Its output is equivalent to the original FragGeneScan implementation. CONCLUSIONS: Compared to the current C implementation, shotgun metagenomic reads are processed up to 22 times faster using a single thread, with better scaling for multithreaded execution. The Rust code of FragGeneScanRs is freely available from GitHub under the GPL-3.0 license with instructions for installation, usage and other documentation ( https://github.com/unipept/FragGeneScanRs ).
Felix Van der Jeugt, Peter Dawyndt, Bart Mesuere
BMC Bioinform.2
2020 Unipept CLI 2.0: adding support for visualizations and functional annotations
abstract
SUMMARY: Unipept is an ecosystem of tools developed for fast metaproteomics data-analysis consisting of a web application, a set of web services (application programming interface, API) and a command-line interface (CLI). After the successful introduction of version 4 of the Unipept web application, we here introduce version 2.0 of the API and CLI. Next to the existing taxonomic analysis, version 2.0 of the API and CLI provides access to Unipept's powerful functional analysis for metaproteomics samples. The functional analysis pipeline supports retrieval of Enzyme Commission numbers, Gene Ontology terms and InterPro entries for the individual peptides in a metaproteomics sample. This paves the way for other applications and developers to integrate these new information sources into their data processing pipelines, which greatly increases insight into the functions performed by the organisms in a specific environment. Both the API and CLI have also been expanded with the ability to render interactive visualizations from a list of taxon ids. These visualizations are automatically made available on a dedicated website and can easily be shared by users. AVAILABILITY AND IMPLEMENTATION: The API is available at http://api.unipept.ugent.be. Information regarding the CLI can be found at https://unipept.ugent.be/clidocs. Both interfaces are freely available and open-source under the MIT license. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pieter Verschaffelt, Philippe Van Thienen, Tim Van Den Bossche, Felix Van der Jeugt, Caroline De Tender, Lennart Martens, Peter Dawyndt, Bart Mesuere
Bioinform.7
2016 Unipept web services for metaproteomics analysis
abstract
UNLABELLED: Unipept is an open source web application that is designed for metaproteomics analysis with a focus on interactive datavisualization. It is underpinned by a fast index built from UniProtKB and the NCBI taxonomy that enables quick retrieval of all UniProt entries in which a given tryptic peptide occurs. Unipept version 2.4 introduced web services that provide programmatic access to the metaproteomics analysis features. This enables integration of Unipept functionality in custom applications and data processing pipelines. AVAILABILITY AND IMPLEMENTATION: The web services are freely available at http://api.unipept.ugent.be and are open sourced under the MIT license. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bart Mesuere, Toon Willems, Felix Van der Jeugt, Bart Devreese, Peter Vandamme, Peter Dawyndt
Bioinform.6
2015 A Long Fragment Aligner called ALFALFA
abstract
BACKGROUND: Rapid evolutions in sequencing technology force read mappers into flexible adaptation to longer reads, changing error models, memory barriers and novel applications. RESULTS: ALFALFA achieves a high performance in accurately mapping long single-end and paired-end reads to gigabase-scale reference genomes, while remaining competitive for mapping shorter reads. Its seed-and-extend workflow is underpinned by fast retrieval of super-maximal exact matches from an enhanced sparse suffix array, with flexible parameter tuning to balance performance, memory footprint and accuracy. CONCLUSIONS: ALFALFA is open source and available at http://alfalfa.ugent.be .
Michaël Vyverman, Bernard De Baets, Veerle Fack, Peter Dawyndt
BMC Bioinform.4
2013 essaMEM: finding maximal exact matches using enhanced sparse suffix arrays
abstract
Abstract Summary: We have developed essaMEM, a tool for finding maximal exact matches that can be used in genome comparison and read mapping. essaMEM enhances an existing sparse suffix array implementation with a sparse child array. Tests indicate that the enhanced algorithm for finding maximal exact matches is much faster, while maintaining the same memory footprint. In this way, sparse suffix arrays remain competitive with the more complex compressed suffix arrays. Availability: Source code is freely available at https://github.ugent.be/ComputationalBiology/essaMEM. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Michaël Vyverman, Bernard De Baets, Veerle Fack, Peter Dawyndt
Bioinform.4
2010 From learning taxonomies to phylogenetic learning: Integration of 16S rRNA gene data into FAME-based bacterial classification
abstract
BACKGROUND: Machine learning techniques have shown to improve bacterial species classification based on fatty acid methyl ester (FAME) data. Nonetheless, FAME analysis has a limited resolution for discrimination of bacteria at the species level. In this paper, we approach the species classification problem from a taxonomic point of view. Such a taxonomy or tree is typically obtained by applying clustering algorithms on FAME data or on 16S rRNA gene data. The knowledge gained from the tree can then be used to evaluate FAME-based classifiers, resulting in a novel framework for bacterial species classification. RESULTS: In view of learning in a taxonomic framework, we consider two types of trees. First, a FAME tree is constructed with a supervised divisive clustering algorithm. Subsequently, based on 16S rRNA gene sequence analysis, phylogenetic trees are inferred by the NJ and UPGMA methods. In this second approach, the species classification problem is based on the combination of two different types of data. Herein, 16S rRNA gene sequence data is used for phylogenetic tree inference and the corresponding binary tree splits are learned based on FAME data. We call this learning approach 'phylogenetic learning'. Supervised Random Forest models are developed to train the classification tasks in a stratified cross-validation setting. In this way, better classification results are obtained for species that are typically hard to distinguish by a single or flat multi-class classification model. CONCLUSIONS: FAME-based bacterial species classification is successfully evaluated in a taxonomic framework. Although the proposed approach does not improve the overall accuracy compared to flat multi-class classification, it has some distinct advantages. First, it has better capabilities for distinguishing species on which flat multi-class classification fails. Secondly, the hierarchical classification structure allows to easily evaluate and visualize the resolution of FAME data for the discrimination of bacterial species. Summarized, by phylogenetic learning we are able to situate and evaluate FAME-based bacterial species classification in a more informative context.
Bram Slabbinck, Willem Waegeman, Peter Dawyndt, Paul De Vos, Bernard De Baets
BMC Bioinform.3
2010 Semantic integration of isolation habitat and location in StrainInfo
abstract
Table 1 Example isolation habitat and location data of a Pichia guilliermondii strain, as listed by different BRCs.For each column, we want to calculate a consensus value for the complete strain Strain number of culture Isolation habitat Isolation location CECT 1456 insect frass on Ulmus americana (elm tree) n/a CLIB
Bert Verslyppe, Wim De Smet, Paul De Vos, Bernard De Baets, Peter Dawyndt
BMC Bioinform.5
2009 Bayesian Clustering of Fuzzy Feature Vectors Using a Quasi-Likelihood Approach
abstract
Bayesian model-based classifiers, both unsupervised and supervised, have been studied extensively and their value and versatility have been demonstrated on a wide spectrum of applications within science and engineering. A majority of the classifiers are built on the assumption of intrinsic discreteness of the considered data features or on the discretization of them prior to the modeling. On the other hand, Gaussian mixture classifiers have also been utilized to a large extent for continuous features in the Bayesian framework. Often the primary reason for discretization in the classification context is the simplification of the analytical and numerical properties of the models. However, the discretization can be problematic due to its \textit{ad hoc} nature and the decreased statistical power to detect the correct classes in the resulting procedure. We introduce an unsupervised classification approach for fuzzy feature vectors that utilizes a discrete model structure while preserving the continuous characteristics of data. This is achieved by replacing the ordinary likelihood by a binomial quasi-likelihood to yield an analytical expression for the posterior probability of a given clustering solution. The resulting model can be justified from an information-theoretic perspective. Our method is shown to yield highly accurate clusterings for challenging synthetic and empirical data sets.
Pekka Marttinen, Jing Tang 0002, Bernard De Baets, Peter Dawyndt, Jukka Corander
IEEE Trans. Pattern Anal. Mach. Intell.4
2008 StrainInfo.net Web Services: Enabling Microbiologic Workflows Such as Phylogenetic Tree Building and Biomarker Comparison
abstract
In this paper we present novel web services offered by the StrainInfo.net bioportal. This portal integrates information in the domain of microbiology and offers a uniform web interface to a multitude of data providers. By providing web services, the integration results of StrainInfo.net become available for automated processing. Several classes of web services are implemented and some interesting uses are discussed in more detail. Combined with third-party services, the StrainInfo.net services can be integrated into workflows. We describe two example workflows: one basic workflow for the construction of a phylogenetic tree based on 16S rRNA gene sequences retrieved from the species of a given genus and a more advanced workflow to collect data of several biomarkers, to calculate the corresponding distance matrices, and to visualize the intra- and inter-species variation among the different biomarkers using the TaxonGap tool. Hereby, the tedious and manual work of collecting and analyzing data, and of visualizing the analysis results has become automated.
Bert Verslyppe, Bram Slabbinck, Wim De Smet, Paul De Vos, Bernard De Baets, Peter Dawyndt
eScience6
2008 TaxonGap: a visualization tool for intra- and inter-species variation among individual biomarkers
abstract
UNLABELLED: Selection of optimal biomarkers for the identification of different operational taxonomic units (OTUs) may be a hard and tedious task, especially when phylogenetic trees for multiple genes need to be compared. With TaxonGap we present a novel and easy-to-handle software tool that allows visual comparison of the discriminative power of multiple biomarkers for a set of OTUs. The compact graphical output allows for easy comparison and selection of individual biomarkers. AVAILABILITY: Graphical User Interface; Executable JAVA archive file, source code, supplementary information and sample files can be downloaded from the website: http://www.kermit.ugent.be/taxongap
Bram Slabbinck, Peter Dawyndt, M. Martens, Paul De Vos, Bernard De Baets
Bioinform.2
2006 UPGMA clustering revisited: A weight-driven approach to transitive approximation
Peter Dawyndt, Hans E. De Meyer, Bernard De Baets
Int. J. Approx. Reason.1
2005 Improving interoperability between microbial information and sequence databases
abstract
BACKGROUND: Biological resources are essential tools for biomedical research. Their availability is promoted through on-line catalogues. Common Access to Biological Resources and Information (CABRI) is a service for distribution of biological resources and related data collected by 28 European culture collections. Linking this information to bioinformatics databanks can make the collections' holdings more visible after a search in molecular biology databanks and vice-versa. Identification of links to sequence databases can be useful, but annotation and indexing problems, together with compilation errors, immediately arise. In this paper, we present our efforts for the identification of cross-references between CABRI catalogues and the EMBL Data Library and related results. RESULTS: An SRS site with both EMBL and CABRI catalogues has been set up. Ad-hoc changes in indexing scripts allowed to achieve homogeneous index keys and SRS link features have been used to identify links between databases. After manual checking and comparison with an alternative procedure, about 67,500 valid cross-references were identified, added to the EMBL Data Library and are now distributed with it. HTML links can be established from EMBL to CABRI network service. Procedures can be executed whenever needed. CONCLUSION: Links between EMBL and CABRI catalogues constitute an improved access to micro-organisms of certified quality and can produce positive effects on biomedical research. Further links between CABRI catalogues and other bioinformatics databases can now easily be defined by using these cross-references. Linking genetic information onto natural resources information may stand model for the integration of other databases containing empirical data on these materials.
Paolo Romano 0001, Peter Dawyndt, Francesca Piersigilli, Jean Swings
BMC Bioinform.2
2005 The complete linkage clustering algorithm revisited
Peter Dawyndt, Hans E. De Meyer, Bernard De Baets
Soft Comput.1
2005 Knowledge Accumulation and Resolution of Data Inconsistencies during the Integration of Microbial Information Sources
abstract
The Internet has emerged as an ever-increasing environment of multiple heterogeneous and autonomous data sources that contain relevant but overlapping information on microorganisms. Microbiologists might therefore seriously benefit from the design of intelligent software agents that assist in the navigation through this information-rich environment, together with the development of data mining tools that can aid in the discovery of new information. These applications heavily depend upon well-conditioned data samples that are correlated with multiple information sources, hence, accurate database merging operations are desirable. Information systems designed for joining the related knowledge provided by different microbial data sources are hampered by the labeling mechanism for referencing microbial strains and cultures that suffers from syntactical variation in the practical usage of the labels, whereas, additionally, synonymy and homonymy are also known to exist amongst the labels. This situation is even complicated by the observation that the label equivalence knowledge is itself fragmentarily recorded over several data sources which can be suspected of providing information that might be both incomplete and incorrect. This paper presents how extraction and integration of label equivalence information from several distributed data sources has led to the construction of a so-called integrated strain database, which helps to resolve most of the above problems. Given the fact that information retrieved from autonomous resources might be overlapping, incomplete, and incorrect, much energy was spent into the completion of missing information, the discovery of new associations between information objects, and the development and application of tools for error detection and correction. Through a thorough evaluation of the different levels of incompleteness and incorrectness encountered within the incorporated data sources, we have finally given proof of the added value of the integrated strain database as a necessary service provider for the seamless integration of microbial information sources.
Peter Dawyndt, Marc Vancanneyt, Hans E. De Meyer, Jean Swings
IEEE Trans. Knowl. Data Eng.1
2004 On the min-transitive approximation of symmetric fuzzy relations
abstract
Two new algorithms are proposed for generating a min-transitive approximation of a given reflexive and symmetric fuzzy relation which, in general, deviates less from the given fuzzy relation than its min-transitive closure, and which is guaranteed to be still reflexive and symmetric. Since the new algorithms are weight-driven, they can be used to generate layer by layer the partition tree associated to the corresponding min-transitive approximation. We report on numerical tests that have been carried out on synthetic data to compare the approximations generated by the new algorithms to the min-transitive closure and the min-transitive approximation delivered by the UPGMA clustering algorithm.
Peter Dawyndt, Hans E. De Meyer, Bernard De Baets
FUZZ-IEEE1