Marco Spruit

dblp:98/217 · also Marco R. Spruit, Marco René Spruit · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-9237-221XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Design and evaluation of semantically-valid negative samples integration techniques for scalable semi-automated drug repurposing prediction pipelines in rare disease research
abstract
BACKGROUND: Computational approaches involving complex data structures (e.g. machine learning, knowledge graphs) have been more prominent in biological studies for the last two decades. Due to increasingly larger amounts of data collected with modern omics techniques, there is a need for methods that can process such data quickly and thoroughly. In addition, those techniques can be applied to extrapolate results from a limited number of observations. Rare disease research benefits particularly from those new computational approaches as each rare disease affects a small percentage of the population. Nevertheless, finding effective treatments benefits a wide portion of the world’s individuals if measured in absolute numbers: 10% of the whole world population is affected by rare diseases as a whole. In the context of rare diseases, drug repurposing (i.e. testing existing approved drugs against other diseases) stands as a viable alternative to traditional drug discovery—thus reducing costs compared to novel drug discovery. RESULTS: We introduce a novel approach for initial candidate drugs selection which is based on a knowledge graph of biological associations between genes involved in the disease and drugs from experimental and clinical databases. Additionally, our approach generates semantically valid negative samples to further improve the selection of candidate drugs. We tested it on Huntington’s disease, a model condition for rare disease research. CONCLUSIONS: Our main contribution is that the approach we introduce in this paper does not require human-curated datasets, resulting in a scalable drug repurposing workflow that leverages information on known and missing associations between gene and drugs to predict candidate repurposed drugs—while implementing strategies that limit hardware resource consumptions, hence reducing computing time.
Niccolò Bianchi, Armel E. J. L. Lefebvre, Katherine J. Wolstencroft, Marco Spruit
BMC Bioinform.4
2024 Federated Learning Analytics: Investigating the Privacy-Performance Trade-Off in Machine Learning for Educational Analytics
Max van Haastrecht, Matthieu J. S. Brinkhuis, Marco Spruit
AIED (2)3
2024 Modelling Pragmatic Inference in Children's Use of Perception Verbs with Language Models
Bram van Dijk, Max J. van Duijn, Li Kloostra, Marco Spruit, Barend Beekhuizen
CogSci4
2024 Attend All Options at Once: Full Context Input for Multi-choice Reading Comprehension
Runda Wang, Suzan Verberne, Marco Spruit
ECIR (1)3
2024 Exploring the potential of federated learning in mental health research: a systematic literature review
abstract
Abstract The rapid advancement of technology has created new opportunities to improve the accuracy and efficiency of medical diagnoses, treatments, and overall patient care in several medical domains, including mental health. One promising novel approach is federated learning, a machine learning approach that allows multiple devices to train a shared model without exchanging raw data. Instead of centralizing the data in one location, each device or machine holds a portion of the data and collaborates with other devices to update the shared model. In this way, federated learning enables training on more extensive and diverse datasets than would be possible with centralized training while preserving the privacy and security of individual data. In the mental health domain, federated learning has the potential to improve mental disorders’ detection, diagnosis, and treatment. By pooling data from multiple sources while maintaining patient privacy by keeping data secure and ensuring that they are not used for unauthorized purposes. This literature survey reviews recent studies that have exploited federated learning in the psychiatric domain, covering multiple data resources and different machine-learning techniques. Furthermore, we formulate the gap in the current methodologies and propose new research directions.
Samar Samir Khalil, Noha S. Tawfik, Marco Spruit
Appl. Intell.3
2023 ChiSCor: A Corpus of Freely-Told Fantasy Stories by Dutch Children for Computational Linguistics and Cognitive Science
abstract
In this resource paper we release ChiSCor, a new corpus containing 619 fantasy stories, told freely by 442 Dutch children aged 4-12. ChiSCor was compiled for studying how children render character perspectives, and unravelling language and cognition in development, with computational tools. Unlike existing resources, ChiSCor’s stories were produced in natural contexts, in line with recent calls for more ecologically valid datasets. ChiSCor hosts text, audio, and annotations for character complexity and linguistic complexity. Additional metadata (e.g. education of caregivers) is available for one third of the Dutch children. ChiSCor also includes a small set of 62 English stories. This paper details how ChiSCor was compiled and shows its potential for future work with three brief case studies: i) we show that the syntactic complexity of stories is strikingly stable across children’s ages; ii) we extend work on Zipfian distributions in free speech and show that ChiSCor obeys Zipf’s law closely, reflecting its social context; iii) we show that even though ChiSCor is relatively small, the corpus is rich enough to train informative lemma vectors that allow us to analyse children’s language use. We end with a reflection on the value of narrative datasets in computational linguistics.
Bram van Dijk, Max J. van Duijn, Suzan Verberne, Marco Spruit
CoNLL4
2023 Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests
abstract
To what degree should we ascribe cognitive capacities to Large Language Models (LLMs), such as the ability to reason about intentions and beliefs known as Theory of Mind (ToM)? Here we add to this emerging debate by (i) testing 11 base- and instruction-tuned LLMs on capabilities relevant to ToM beyond the dominant false-belief paradigm, including non-literal language usage and recursive intentionality; (ii) using newly rewritten versions of standardized tests to gauge LLMs' robustness; (iii) prompting and scoring for open besides closed questions; and (iv) benchmarking LLM performance against that of children aged 7-10 on the same tasks. We find that instruction-tuned LLMs from the GPT family outperform other models, and often also children. Base-LLMs are mostly unable to solve ToM tasks, even with specialized prompting. We suggest that the interlinked evolution and development of language and ToM may help explain what instruction-tuning adds: rewarding cooperative communication that takes into account interlocutor and context. We conclude by arguing for a nuanced perspective on ToM in LLMs.
Max J. van Duijn, Bram van Dijk, Tom Kouwenhoven, Werner de Valk, Marco Spruit, Peter van der Putten
CoNLL5
2023 Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding
abstract
Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.LLMs are appearing rapidly, and debates on LLM capacities have taken off, but reflection is lagging behind.Thus, in this position paper, we first zoom in on the debate and critically assess three points recurring in critiques of LLM capacities: i) that LLMs only parrot statistical patterns in the training data; ii) that LLMs master formal but not functional language competence; and iii) that language learning in LLMs cannot inform human language learning.Drawing on empirical and theoretical arguments, we show that these points need more nuance.Second, we outline a pragmatic perspective on the issue of 'real' understanding and intentionality in LLMs.Understanding and intentionality pertain to unobservable mental states we attribute to other humans because they have pragmatic value: they allow us to abstract away from complex underlying mechanics and predict behaviour effectively.We reflect on the circumstances under which it would make sense for humans to similarly attribute mental states to LLMs, thereby outlining a pragmatic philosophical context for LLMs as an increasingly prominent technology in society.
Bram van Dijk, Tom Kouwenhoven, Marco Spruit, Max J. van Duijn
EMNLP3
2023 Embracing Trustworthiness and Authenticity in the Validation of Learning Analytics Systems
abstract
Learning analytics sits in the middle space between learning theory and data analytics. The inherent diversity of learning analytics manifests itself in an epistemology that strikes a balance between positivism and interpretivism, and knowledge that is sourced from theory and practice. In this paper, we argue that validation approaches for learning analytics systems should be cognisant of these diverse foundations. Through a systematic review of learning analytics validation research, we find that there is currently an over-reliance on positivistic validity criteria. Researchers tend to ignore interpretivistic criteria such as trustworthiness and authenticity. In the 38 papers we analysed, researchers covered positivistic validity criteria 221 times, whereas interpretivistic criteria were mentioned 37 times. We motivate that learning analytics can only move forward with holistic validation strategies that incorporate “thick descriptions” of educational experiences. We conclude by outlining a planned validation study using argument-based validation, which we believe will yield meaningful insights by considering a diverse spectrum of validity criteria.
Max van Haastrecht, Matthieu J. S. Brinkhuis, Jessica Peichl, Bernd Remmele, Marco Spruit
LAK5
2023 Adaptable Security Maturity Assessment and Standardization for Digital SMEs
abstract
Small and Medium-sized Enterprises (SMEs) constitute a very large part of every country’s economy and play an essential role in economic growth and social development. SMEs are frequent targets of cyberattacks. Unlike large enterprises, SMEs generally have limited capabilities regarding cybersecurity practices. Assessment and improvement of cybersecurity capabilities are crucial for SMEs to survive and sustain their operations. Despite the availability of maturity assessment models and standards to assess and improve cybersecurity capabilities, SMEs’ specific requirements and roles in the digital ecosystem are often neglected. This paper presents high-level SME requirements regarding cybersecurity maturity assessment and standardization and translates them into an Adaptable Security Maturity Assessment and Standardization (ASMAS) framework to address this gap. The framework is demonstrated by a web-based software prototype. In the evaluation study conducted with SMEs, we obtained positive results for perceived usefulness, perceived ease of use of the framework, and intention to use it.
Bilge Yigit Ozkan, Marco Spruit
J. Comput. Inf. Syst.2
2022 FuzzyTM: a Software Package for Fuzzy Topic Modeling
abstract
Unstructured text data is collected daily in large amounts by many organizations. Analyzing all this data is time intensive and too costly in many cases. One technique to systematically analyze large corpora of texts is topic modeling, which returns the latent topics present in a corpus. Recently, several fuzzy topic modeling algorithms have been proposed and have shown superior results over the existing algorithms. Although various Python libraries offer topic modeling algorithms, none includes fuzzy topic models. Therefore, we present FuzzyTM, a Python library for training fuzzy topic models and creating topic embeddings for downstream tasks. The user-friendly pipelines with default values allow practitioners to train a topic model with minimal effort. Meanwhile, its modular design allows researchers to modify each software element and for future methods to be added.
Emil Rijcken, Pablo Mosteiro, Kalliopi Zervanou, Marco Spruit, Floor Scheepers, Uzay Kaymak
FUZZ-IEEE4
2022 Exploring Embedding Spaces for more Coherent Topic Modeling in Electronic Health Records
abstract
The written notes in the Electronic Health Records contain a vast amount of information about patients. Implementing automated approaches for text classification tasks requires the automated methods to be well-interpretable, and topic models can be used for this goal as they can indicate what topics in a text are relevant to making a decision. We propose a new topic modeling algorithm, FLSA-E, and compare it with another state-of-the-art algorithm FLSA-W. In FLSA-E, topics are found by fuzzy clustering in a word embedding space. Since we use word embeddings as the basis for our clustering, we extend our evaluation with word-embeddings-based evaluation metrics. We find that different evaluation metrics favour different algorithms. Based on the results, there is evidence that FLSA-E has fewer outliers in its topics, a desirable property, given that within-topic words need to be semantically related.
Emil Rijcken, Kalliopi Zervanou, Marco Spruit, Pablo Mosteiro, Floor Scheepers, Uzay Kaymak
SMC3
2022 Federated learning for violence incident prediction in a simulated cross-institutional psychiatric setting
abstract
Inpatient violence is a common and severe problem within psychiatry. Knowing who might become violent can influence staffing levels and mitigate severity. Predictive machine learning models can assess each patient’s likelihood of becoming violent based on clinical notes. Yet, while machine learning models benefit from having more data, data availability is limited as hospitals typically do not share their data for privacy preservation. Federated Learning (FL) can overcome the problem of data limitation by training models in a decentralised manner, without disclosing data between collaborators. However, although several FL approaches exist, none of these train Natural Language Processing models on clinical notes. In this work, we investigate the application of Federated Learning to clinical Natural Language Processing, applied to the task of Violence Risk Assessment by simulating a cross-institutional psychiatric setting. We train and compare four models: two local models, a federated model and a data-centralised model. Our results indicate that the federated model outperforms the local models and has similar performance as the data-centralised model. These findings suggest that Federated Learning can be used successfully in a cross-institutional setting and is a step towards new applications of Federated Learning based on clinical notes.
Thomas Borger, Pablo Mosteiro, Heysem Kaya, Emil Rijcken, Albert Ali Salah, Floor Scheepers, Marco Spruit
Expert Syst. Appl.7
2021 A Threat-Based Cybersecurity Risk Assessment Approach Addressing SME Needs
abstract
Cybersecurity incidents are commonplace nowadays, and Small- and Medium-Sized Enterprises (SMEs) are exceptionally vulnerable targets. The lack of cybersecurity resources available to SMEs implies that they are less capable of dealing with cyber-attacks. Motivation to improve cybersecurity is often low, as the prerequisite knowledge and awareness to drive motivation is generally absent at SMEs. A solution that aims to help SMEs manage their cybersecurity risks should therefore not only offer a correct assessment but should also motivate SME users. From Self-Determination Theory (SDT), we know that by promoting perceived autonomy, competence, and relatedness, people can be motivated to take action. In this paper, we explain how a threat-based cybersecurity risk assessment approach can help to address the needs outlined in SDT. We propose such an approach for SMEs and outline the data requirements that facilitate automation. We present a practical application covering various user interfaces, showing how our threat-based cybersecurity risk assessment approach turns SME data into prioritised, actionable recommendations.
Max van Haastrecht, Injy Sarhan, Alireza Shojaifar, Louis Baumgartner, Wissam Mallouli, Marco Spruit
ARES6
2021 Open-CyKG: An Open Cyber Threat Intelligence Knowledge Graph
abstract
Instant analysis of cybersecurity reports is a fundamental challenge for security experts as an immeasurable amount of cyber information is generated on a daily basis, which necessitates automated information extraction tools to facilitate querying and retrieval of data. Hence, we present Open-CyKG: an Open Cyber Threat Intelligence (CTI) Knowledge Graph (KG) framework that is constructed using an attention-based neural Open Information Extraction (OIE) model to extract valuable cyber threat information from unstructured Advanced Persistent Threat (APT) reports. More specifically, we first identify relevant entities by developing a neural cybersecurity Named Entity Recognizer (NER) that aids in labeling relation triples generated by the OIE model. Afterwards, the extracted structured data is canonicalized to build the KG by employing fusion techniques using word embeddings. As a result, security professionals can execute queries to retrieve valuable information from the Open-CyKG framework. Experimental results demonstrate that our proposed components that build up Open-CyKG outperform state-of-the-art models.1
Injy Sarhan, Marco Spruit
Knowl. Based Syst.2
2020 Evaluating sentence representations for biomedical text: Methods and experimental results
Noha S. Tawfik, Marco Spruit
J. Biomed. Informatics2
2019 Using Cluster Ensembles to Identify Psychiatric Patient Subgroups
Vincent Jorn Menger, Marco Spruit, Wouter van der Klift, Floor Scheepers
AIME2
2019 Contextualized Word Embeddings in a Neural Open Information Extraction Model
Injy Sarhan, Marco Spruit
NLDB2
2019 PreMedOnto: A Computer Assisted Ontology for Precision Medicine
Noha S. Tawfik, Marco Spruit
NLDB2
2019 Towards Recognition of Textual Entailment in the Biomedical Domain
Noha S. Tawfik, Marco Spruit
NLDB2
2017 Risk Mediation in Association Rules - The Case of Decision Support in Medication Review
Michiel Meulendijk, Marco Spruit, Sjaak Brinkkemper
AIME2
2017 Full-Text or Abstract? Examining Topic Coherence Scores Using Latent Dirichlet Allocation
abstract
This paper assesses topic coherence and human topic ranking of uncovered latent topics from scientific publications when utilizing the topic model latent Dirichlet allocation (LDA) on abstract and full-text data. The coherence of a topic, used as a proxy for topic quality, is based on the distributional hypothesis that states that words with similar meaning tend to co-occur within a similar context. Although LDA has gained much attention from machine-learning researchers, most notably with its adaptations and extensions, little is known about the effects of different types of textual data on generated topics. Our research is the first to explore these practical effects and shows that document frequency, document word length, and vocabulary size have mixed practical effects on topic coherence and human topic ranking of LDA topics. We furthermore show that large document collections are less affected by incorrect or noise terms being part of the topic-word distributions, causing topics to be more coherent and ranked higher. Differences between abstract and full-text data are more apparent within small document collections, with differences as large as 90% high-quality topics for full-text data, compared to 50% high-quality topics for abstract data.
Shaheen Syed, Marco Spruit
DSAA2
2016 Organizational Characteristics Influencing SME Information Security Maturity
abstract
In the current business environment, many organizations use popular standards such as the ISO 27000x series, COBIT, and related frameworks to protect themselves against security incidents. However, these standards and frameworks are overly complicated for small to medium-sized enterprises, leaving these organizations with no easy to understand toolkit to address their security needs. This research builds upon the recent Information Security Focus Area Maturity (ISFAM) model for SME information security as a cornerstone in the development of an assessment tool for tailor-made, fast, and easy-to-use information security advice for SMEs. By performing an extensive literature review and evaluating the results with security experts, we propose the Characterizing Organizations’ Information Security for SMEs (CHOISS) model to relate measurable organizational characteristics in four categories through 47 parameters to help SMEs distinguish and prioritize which risks to mitigate.
Frederik Mijnhardt, Thijs Baars, Marco Spruit
J. Comput. Inf. Syst.3
2013 Mobile Business Intelligence: Key Considerations for Implementations Projects
abstract
The new generation of mobile devices, such as smartphones and tablets, is enabling employees to access business insights anytime, anywhere. This trend in Business Intelligence (BI) is popularized under the term mobile BI. Various studies indicate a strong increase in the adoption of this technology. However, mobile BI implementations remain unexplored and unsupported by implementation methods. By devising a Mobile BI Implementation (MOBII) framework, this study aims to fill in this research gap. A systematic literature review revealed the following major implementation themes: (1) value creation, (2) application deployment, (3) information security, (4) workforce mobilization, (5) information delivery and (6) device management. Moreover, expert interviews revealed twenty key considerations, which are also included in the framework. Using a single case study the MOBII framework was successfully evaluated and its practical applicability was demonstrated by adapting an actively used BI implementation method.
Kim Verkooij, Marco Spruit
J. Comput. Inf. Syst.2
2012 Designing a Secure Cloud Architecture: The SeCA Model
abstract
Security issues are paramount when considering adoption of any cloud technology. This article proposes the Secure Cloud Architecture (SeCA) model on the basis of data classifications which defines a properly secure cloud architecture by testing the cloud environment on eight attributes. The SeCA model is developed using a literature review and a Delphi study with seventeen experts, consisting of three rounds. The authors integrate the CI3A—an extension on the CIA-triad—to create a basic framework for testing the classification inputted. The data classification is then tested on regional, geo-spatial, delivery, deployment, governance and compliance, network, premise and encryption attributes. After this testing has been executed, a specification for a secure cloud architecture is outputted.
Thijs Baars, Marco Spruit
Int. J. Inf. Secur. Priv.2
2012 CITS: The Cost of IT Security Framework
abstract
Organizations know that investing in security measures is an important requirement for doing business. But how much should they invest and how should those investments be directed? Many organizations have turned to a risk management approach to identify the largest threats and the control measures that could help mitigate those threats. This research presents the Cost of IT Security (CITS) Framework to support analysis of the costs and benefits of those control measures. This analysis can be performed by using either quantification methods or by using a qualitative approach. Based on a study of five distinct security areas–Identity Management, Network Access Control, Intrusion Detection Systems, Business Continuity Management and Data Loss Prevention–nine cost factors are identified for IT security, and for only five of those nine a quantitative approach is feasible for the cost factor. This study finds that even though quantification methods are useful, organizations that wish to use those should do this together with more qualitative approaches in the decision-making process for security measures.
Marco Spruit, Wouter de Bruijn
Int. J. Inf. Secur. Priv.1
2010 A Framework for Process Improvement in Software Product Management
Willem Bekkers, Inge van de Weerd, Marco Spruit, Sjaak Brinkkemper
EuroSPI3