Sophia Tsoka

dblp:33/5439 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-8403-1282ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge Transfer
abstract
Generalized few-shot semantic segmentation (GFSS) aims to segment objects of both base and novel classes, using sufficient samples of base classes and few samples of novel classes. Representative GFSS approaches typically employ a two-phase training scheme, involving base class pre-training followed by novel class fine-tuning, to learn the classifiers for base and novel classes respectively. Nevertheless, distribution gap exists between base and novel classes in this process. To narrow this gap, we exploit effective knowledge transfer from base to novel classes. First, a novel prototype modulation module is designed to modulate novel class prototypes by exploiting the correlations between base and novel classes. Second, a novel classifier calibration module is proposed to calibrate the weight distribution of the novel classifier according to that of the base classifier. Furthermore, existing GFSS approaches suffer from a lack of contextual information for novel classes due to their limited samples, we thereby introduce a context consistency learning scheme to transfer the contextual knowledge from base to novel classes. Extensive experiments on PASCAL-5i and COCO-20i demonstrate that our approach significantly enhances the state of the art in the GFSS setting.
Xinyue Chen 0009, Miaojing Shi, Zijian Zhou 0002, Lianghua He, Sophia Tsoka
AAAI5
2025 Interpretable optimisation-based approach for hyper-box classification
abstract
Data classification is considered a fundamental research subject within the machine learning community. Researchers seek the improvement of machine learning algorithms in not only accuracy, but also interpretability. Interpretable algorithms allow humans to easily understand the decisions that a machine learning model makes, which is challenging for black box models. Mathematical programming-based classification algorithms have attracted considerable attention due to their ability to effectively compete with leading-edge algorithms in terms of both accuracy and interpretability. Meanwhile, the training of a hyper-box classifier can be mathematically formulated as a Mixed Integer Linear Programming (MILP) model and the predictions combine accuracy and interpretability. In this work, an optimisation-based approach is proposed for multi-class data classification using a hyper-box representation, thus facilitating the extraction of compact IF-THEN rules. The key novelty of our approach lies in the minimisation of the number and length of the generated rules for enhanced interpretability. Through a number of real-world datasets, it is demonstrated that the algorithm exhibits favorable performance when compared to well-known alternatives in terms of prediction accuracy and rule set simplicity.
Georgios I. Liapis, Sophia Tsoka, Lazaros G. Papageorgiou
Mach. Learn.2
2024 Pathway Activity Autoencoders for Enhanced Omics Analysis and Clinical Interpretability
abstract
Large-scale analyses of omics data are crucial for advancing precision medicine and personalised treatments. Current methods to unravel cancer development and prognosis exhibit a dichotomy between interpretability and representational power. Most either rely on post-hoc interpretability techniques with black-box models or use linear models that might miss complex biological interactions. We propose a novel configurable prior-knowledge-based deep auto-encoding framework called PAAE and its generative variant PAVAE, for analyzing cancer RNA-seq data. Our method constrains its learned internal representation with biological pathways, providing interpretability without sacrificing predictive power. Our model is tested on 3 different downstream tasks: cancer subtype classification, survival analysis and unsupervised clustering of the learned features. Our models outperform baselines while having orders of magnitude less parameters than naive models. Extensive interpretability analyses, including task-relevant feature identification demonstrate our model’s effectiveness at identifying underlying biological signals in an unsupervised fashion. Visualisations are used to highlight the intrinsic interpretability of our models. The source code of this study is available at github.com/phcavelar/pathwayae
Pedro H. C. Avelar, Le Ou-Yang, Min Wu 0008, Sophia Tsoka
BIBM4
2024 Optimisation-based modelling for explainable lead discovery in malaria
abstract
BACKGROUND: The search for new antimalarial treatments is urgent due to growing resistance to existing therapies. The Open Source Malaria (OSM) project offers a promising starting point, having extensively screened various compounds for their effectiveness. Further analysis of the chemical space surrounding these compounds could provide the means for innovative drugs. METHODS: We report an optimisation-based method for quantitative structure-activity relationship (QSAR) modelling that provides explainable modelling of ligand activity through a mathematical programming formulation. The methodology is based on piecewise regression principles and offers optimal detection of breakpoint features, efficient allocation of samples into distinct sub-groups based on breakpoint feature values, and insightful regression coefficients. Analysis of OSM antimalarial compounds yields interpretable results through rules generated by the model that reflect the contribution of individual fingerprint fragments in ligand activity prediction. Using knowledge of fragment prioritisation and screening of commercially available compound libraries, potential lead compounds for antimalarials are identified and evaluated experimentally via a Plasmodium falciparum asexual growth inhibition assay (PfGIA) and a human cell cytotoxicity assay. CONCLUSIONS: Three compounds are identified as potential leads for antimalarials using the methodology described above. This work illustrates how explainable predictive models based on mathematical optimisation can pave the way towards more efficient fragment-based lead discovery as applied in malaria.
Yutong Li 0006, Jonathan Cardoso-Silva, John M. Kelly, Michael J. Delves, Nicholas Furnham, Lazaros G. Papageorgiou, Sophia Tsoka
Artif. Intell. Medicine7
2024 CLigOpt: controllable ligand design through target-specific optimization
abstract
MOTIVATION: A key challenge in deep generative models for molecular design is to navigate random sampling of the vast molecular space, and produce promising molecules that strike a balance across multiple chemical criteria. Fragment-based drug design (FBDD), using fragments as starting points, is an effective way to constrain chemical space and improve generation of biologically active molecules. Furthermore, optimization approaches are often implemented with generative models to search through chemical space, and identify promising samples which satisfy specific properties. Controllable FBDD has promising potential in efficient target-specific ligand design. RESULTS: We propose a controllable FBDD model, CLigOpt, which can generate molecules with desired properties from a given fragment pair. CLigOpt is a variational autoencoder-based model which utilizes co-embeddings of node and edge features to fully mine information from molecular graphs, as well as a multi-objective Controllable Generation Module to generate molecules under property controls. CLigOpt achieves consistently strong performance in generating structurally and chemically valid molecules, as evaluated across six metrics. Applicability is illustrated through ligand candidates for hDHFR and it is shown that the proportion of feasible active molecules from the generated set is increased by 10%. Molecular docking and synthesizability prediction tasks are conducted to prioritize generated molecules to derive potential lead compounds. AVAILABILITY AND IMPLEMENTATION: The source code is available via https://github.com/yutongLi1997/CLigOpt-Controllable-Ligand-Design-through-Target-Specific-Optimisation.
Yutong Li 0006, Pedro H. C. Avelar, Xinyue Chen 0009, Min Wu 0008, Sophia Tsoka
Bioinform.6
2023 Drug repurposing and prediction of multiple interaction types via graph embedding
abstract
BACKGROUND: Finding drugs that can interact with a specific target to induce a desired therapeutic outcome is key deliverable in drug discovery for targeted treatment. Therefore, both identifying new drug-target links, as well as delineating the type of drug interaction, are important in drug repurposing studies. RESULTS: A computational drug repurposing approach was proposed to predict novel drug-target interactions (DTIs), as well as to predict the type of interaction induced. The methodology is based on mining a heterogeneous graph that integrates drug-drug and protein-protein similarity networks, together with verified drug-disease and protein-disease associations. In order to extract appropriate features, the three-layer heterogeneous graph was mapped to low dimensional vectors using node embedding principles. The DTI prediction problem was formulated as a multi-label, multi-class classification task, aiming to determine drug modes of action. DTIs were defined by concatenating pairs of drug and target vectors extracted from graph embedding, which were used as input to classification via gradient boosted trees, where a model is trained to predict the type of interaction. After validating the prediction ability of DT2Vec+, a comprehensive analysis of all unknown DTIs was conducted to predict the degree and type of interaction. Finally, the model was applied to propose potential approved drugs to target cancer-specific biomarkers. CONCLUSION: DT2Vec+ showed promising results in predicting type of DTI, which was achieved via integrating and mapping triplet drug-target-disease association graphs into low-dimensional dense vectors. To our knowledge, this is the first approach that addresses prediction between drugs and targets across six interaction types.
E. Amiri Souri, A. Chenoweth, S. N. Karagiannis, Sophia Tsoka
BMC Bioinform.4
2022 Novel drug-target interactions via link prediction and network embedding
abstract
BACKGROUND: As many interactions between the chemical and genomic space remain undiscovered, computational methods able to identify potential drug-target interactions (DTIs) are employed to accelerate drug discovery and reduce the required cost. Predicting new DTIs can leverage drug repurposing by identifying new targets for approved drugs. However, developing an accurate computational framework that can efficiently incorporate chemical and genomic spaces remains extremely demanding. A key issue is that most DTI predictions suffer from the lack of experimentally validated negative interactions or limited availability of target 3D structures. RESULTS: We report DT2Vec, a pipeline for DTI prediction based on graph embedding and gradient boosted tree classification. It maps drug-drug and protein-protein similarity networks to low-dimensional features and the DTI prediction is formulated as binary classification based on a strategy of concatenating the drug and target embedding vectors as input features. DT2Vec was compared with three top-performing graph similarity-based algorithms on a standard benchmark dataset and achieved competitive results. In order to explore credible novel DTIs, the model was applied to data from the ChEMBL repository that contain experimentally validated positive and negative interactions which yield a strong predictive model. Then, the developed model was applied to all possible unknown DTIs to predict new interactions. The applicability of DT2Vec as an effective method for drug repurposing is discussed through case studies and evaluation of some novel DTI predictions is undertaken using molecular docking. CONCLUSIONS: The proposed method was able to integrate and map chemical and genomic space into low-dimensional dense vectors and showed promising results in predicting novel DTIs.
E. Amiri Souri, Roman Laddach, S. N. Karagiannis, Lazaros G. Papageorgiou, Sophia Tsoka
BMC Bioinform.5
2018 Reactome Pengine: a web-logic API to the Homo sapiens reactome
abstract
Summary: Existing ways of accessing data from the Reactome database are limited. Either a researcher is restricted to particular queries defined by a web application programming interface (API) or they have to download the whole database. Reactome Pengine is a web service providing a logic programming-based API to the human reactome. This gives researchers greater flexibility in data access than existing APIs, as users can send their own small programs (alongside queries) to Reactome Pengine. Availability and implementation: The server and an example notebook can be found at https://apps.nms.kcl.ac.uk/reactome-pengine. Source code is available at https://github.com/samwalrus/reactome-pengine and a Docker image is available at https://hub.docker.com/r/samneaves/rp4/. Supplementary information: Supplementary data are available at Bioinformatics online.
Samuel R. Neaves, Sophia Tsoka, Louise A. C. Millard
Bioinform.2
2017 A regression tree approach using mathematical programming
abstract
Regression analysis is a machine learning approach that aims to accurately predict the value of continuous output variables from certain independent input variables, via automatic estimation of their latent relationship from data. Tree-based regression models are popular in literature due to their flexibility to model higher order non-linearity and great interpretability. Conventionally, regression tree models are trained in a two-stage procedure, i.e. recursive binary partitioning is employed to produce a tree structure, followed by a pruning process of removing insignificant leaves, with the possibility of assigning multivariate functions to terminal leaves to improve generalisation. This work introduces a novel methodology of node partitioning which, in a single optimisation model, simultaneously performs the two tasks of identifying the break-point of a binary split and assignment of multivariate functions to either leaf, thus leading to an efficient regression tree model. Using six real world benchmark problems, we demonstrate that the proposed method consistently outperforms a number of state-of-the-art regression tree models and methods based on other techniques, with an average improvement of 7–60% on the mean absolute errors (MAE) of the predictions.
Lingjian Yang, Sophia Tsoka, Lazaros G. Papageorgiou
Expert Syst. Appl.3
2016 Mathematical programming for piecewise linear regression analysis
Lingjian Yang, Sophia Tsoka, Lazaros G. Papageorgiou
Expert Syst. Appl.3
2015 Using ILP to Identify Pathway Activation Patterns in Systems Biology
Samuel R. Neaves, Louise A. C. Millard, Sophia Tsoka
ILP3
2014 Pathway activity inference for multiclass disease classification through a mathematical programming optimisation framework
abstract
BACKGROUND: Applying machine learning methods on microarray gene expression profiles for disease classification problems is a popular method to derive biomarkers, i.e. sets of genes that can predict disease state or outcome. Traditional approaches where expression of genes were treated independently suffer from low prediction accuracy and difficulty of biological interpretation. Current research efforts focus on integrating information on protein interactions through biochemical pathway datasets with expression profiles to propose pathway-based classifiers that can enhance disease diagnosis and prognosis. As most of the pathway activity inference methods in literature are either unsupervised or applied on two-class datasets, there is good scope to address such limitations by proposing novel methodologies. RESULTS: A supervised multiclass pathway activity inference method using optimisation techniques is reported. For each pathway expression dataset, patterns of its constituent genes are summarised into one composite feature, termed pathway activity, and a novel mathematical programming model is proposed to infer this feature as a weighted linear summation of expression of its constituent genes. Gene weights are determined by the optimisation model, in a way that the resulting pathway activity has the optimal discriminative power with regards to disease phenotypes. Classification is then performed on the resulting low-dimensional pathway activity profile. CONCLUSIONS: The model was evaluated through a variety of published gene expression profiles that cover different types of disease. We show that not only does it improve classification accuracy, but it can also perform well in multiclass disease datasets, a limitation of other approaches from the literature. Desirable features of the model include the ability to control the maximum number of genes that may participate in determining pathway activity, which may be pre-specified by the user. Overall, this work highlights the potential of building pathway-based multi-phenotype classifiers for accurate disease diagnosis and prognosis problems.
Lingjian Yang, Chrysanthi Ainali, Sophia Tsoka, Lazaros G. Papageorgiou
BMC Bioinform.3
2010 Promoter Complexity and Tissue-Specific Expression of Stress Response Components in Mytilus galloprovincialis, a Sessile Marine Invertebrate Species
abstract
The mechanisms of stress tolerance in sessile animals, such as molluscs, can offer fundamental insights into the adaptation of organisms for a wide range of environmental challenges. One of the best studied processes at the molecular level relevant to stress tolerance is the heat shock response in the genus Mytilus. We focus on the upstream region of Mytilus galloprovincialis Hsp90 genes and their structural and functional associations, using comparative genomics and network inference. Sequence comparison of this region provides novel evidence that the transcription of Hsp90 is regulated via a dense region of transcription factor binding sites, also containing a region with similarity to the Gamera family of LINE-like repetitive sequences and a genus-specific element of unknown function. Furthermore, we infer a set of gene networks from tissue-specific expression data, and specifically extract an Hsp class-associated network, with 174 genes and 2,226 associations, exhibiting a complex pattern of expression across multiple tissue types. Our results (i) suggest that the heat shock response in the genus Mytilus is regulated by an unexpectedly complex upstream region, and (ii) provide new directions for the use of the heat shock process as a biosensor system for environmental monitoring.
Chrysa Pantzartzi, Elena Drosopoulou, Minas Yiangou, Ignat Drozdov, Sophia Tsoka, Christos A. Ouzounis, Zacharias G. Scouras
PLoS Comput. Biol.5
2010 A Systems Model for Immune Cell Interactions Unravels the Mechanism of Inflammation in Human Skin
abstract
Inflammation is characterized by altered cytokine levels produced by cell populations in a highly interdependent manner. To elucidate the mechanism of an inflammatory reaction, we have developed a mathematical model for immune cell interactions via the specific, dose-dependent cytokine production rates of cell populations. The model describes the criteria required for normal and pathological immune system responses and suggests that alterations in the cytokine production rates can lead to various stable levels which manifest themselves in different disease phenotypes. The model predicts that pairs of interacting immune cell populations can maintain homeostatic and elevated extracellular cytokine concentration levels, enabling them to operate as an immune system switch. The concept described here is developed in the context of psoriasis, an immune-mediated disease, but it can also offer mechanistic insights into other inflammatory pathologies as it explains how interactions between immune cell populations can lead to disease phenotypes.
Najl V. Valeyev, Christian Hundhausen, Yoshinori Umezawa, Nikolai V. Kotov, Gareth Williams, Alex Clop, Chrysanthi Ainali, Christos A. Ouzounis, Sophia Tsoka, Frank O. Nestle
PLoS Comput. Biol.9
2009 Stratification of co-evolving genomic groups using ranked phylogenetic profiles
abstract
BACKGROUND: Previous methods of detecting the taxonomic origins of arbitrary sequence collections, with a significant impact to genome analysis and in particular metagenomics, have primarily focused on compositional features of genomes. The evolutionary patterns of phylogenetic distribution of genes or proteins, represented by phylogenetic profiles, provide an alternative approach for the detection of taxonomic origins, but typically suffer from low accuracy. Herein, we present rank-BLAST, a novel approach for the assignment of protein sequences into genomic groups of the same taxonomic origin, based on the ranking order of phylogenetic profiles of target genes or proteins across the reference database. RESULTS: The rank-BLAST approach is validated by computing the phylogenetic profiles of all sequences for five distinct microbial species of varying degrees of phylogenetic proximity, against a reference database of 243 fully sequenced genomes. The approach - a combination of sequence searches, statistical estimation and clustering - analyses the degree of sequence divergence between sets of protein sequences and allows the classification of protein sequences according to the species of origin with high accuracy, allowing taxonomic classification of 64% of the proteins studied. In most cases, a main cluster is detected, representing the corresponding species. Secondary, functionally distinct and species-specific clusters exhibit different patterns of phylogenetic distribution, thus flagging gene groups of interest. Detailed analyses of such cases are provided as examples. CONCLUSION: Our results indicate that the rank-BLAST approach can capture the taxonomic origins of sequence collections in an accurate and efficient manner. The approach can be useful both for the analysis of genome evolution and the detection of species groups in metagenomics samples.
Shiri Freilich, Leon Goldovsky, Assaf Gottlieb, Eric Blanc, Sophia Tsoka, Christos A. Ouzounis
BMC Bioinform.5
2005 CoGenT++: an extensive and extensible data environment for computational genomics
abstract
MOTIVATION: CoGenT++ is a data environment for computational research in comparative and functional genomics, designed to address issues of consistency, reproducibility, scalability and accessibility. DESCRIPTION: CoGenT++ facilitates the re-distribution of all fully sequenced and published genomes, storing information about species, gene names and protein sequences. We describe our scalable implementation of ProXSim, a continually updated all-against-all similarity database, which stores pairwise relationships between all genome sequences. Based on these similarities, derived databases are generated for gene fusions--AllFuse, putative orthologs--OFAM, protein families--TRIBES, phylogenetic profiles--ProfUse and phylogenetic trees. Extensions based on the CoGenT++ environment include disease gene prediction, pattern discovery, automated domain detection, genome annotation and ancestral reconstruction. CONCLUSION: CoGenT++ provides a comprehensive environment for computational genomics, accessible primarily for large-scale analyses as well as manual browsing.
Leon Goldovsky, Paul J. Janssen, Dag G. Ahrén, Benjamin Audit, Ildefonso Cases, Nikos Darzentas, Anton J. Enright, Núria López-Bigas, José M. Peregrín-Alvarez, Mike Smith 0001, Sophia Tsoka, Victor Kunin, Christos A. Ouzounis
Bioinform.11
2003 Evaluation of annotation strategies using an entire genome sequence
abstract
Abstract Motivation: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. Results: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences. Availability and supplementary information: http://www.ebi.ac.uk/research/cgg/annotation/cteval/ Contact: [email protected] * To whom correspondence should be addressed. † INA-EKETA, GR-57001 Thessaloniki, Greece ‡ Computational Biology Center, Memorial Sloan-Kettering Cancer Center, New York, NY 10021, USA § Aetion Technologies LLC, Worthington, OH 43085, USA ¶ Institut Curie, F-75248 Paris, France ∥ CNRS, UMR6543, F-06108 Nice, France ** Alma Bioinformatics, E-28760 Madrid, Spain †† Cap Gemini Ernst & Young, London SW1X 7LX, UK ‡‡ Univ. of Rome ‘La Sapienza’, I-00185 Rome, Italy §§ MWG-Biotech AG, Ebersberg, D-85560 Berlin, Germany ¶¶ Wellcome Trust Biocentre, Univ. of Dundee, Dundee DD1 5HN, UK
Sophia Tsoka, Miguel A. Andrade-Navarro, Anton J. Enright, Mark Carroll, Patrick Poullet, Vasilis J. Promponas, Theodore Liakopoulos, Giorgos Palaios, Claude Pasquier, Stavros J. Hamodrakas, Javier Tamames, Asutosh T. Yagnik, Anna Tramontano, Damien Devos, Christian Blaschke, Alfonso Valencia, David Brett, David M. A. Martin, Christophe Leroy, Isidore Rigoutsos, Chris Sander, Christos A. Ouzounis
Bioinform.2
2002 Modeling the percolation of annotation errors in a database of protein sequences
abstract
Public sequence databases contain information on the sequence, structure and function of proteins. Genome sequencing projects have led to a rapid increase in protein sequence information, but reliable, experimentally verified, information on protein function lags a long way behind. To address this deficit, functional annotation in protein databases is often inferred by sequence similarity to homologous, annotated proteins, with the attendant possibility of error. Now, the functional annotation in these homologous proteins may itself have been acquired through sequence similarity to yet other proteins, and it is generally not possible to determine how the functional annotation of any given protein has been acquired. Thus the possibility of chains of misannotation arises, a process we term 'error percolation'. With some simple assumptions, we develop a dynamical probabilistic model for these misannotation chains. By exploring the consequences of the model for annotation quality it is evident that this iterative approach leads to a systematic deterioration of database quality.
Walter R. Gilks, Benjamin Audit, Daniela De Angelis, Sophia Tsoka, Christos A. Ouzounis
Bioinform.4
2000 CAST: an iterative algorithm for the complexity analysis of sequence tracts
abstract
MOTIVATION: Sensitive detection and masking of low-complexity regions in protein sequences. Filtered sequences can be used in sequence comparison without the risk of matching compositionally biased regions. The main advantage of the method over similar approaches is the selective masking of single residue types without affecting other, possibly important, regions. RESULTS: A novel algorithm for low-complexity region detection and selective masking. The algorithm is based on multiple-pass Smith-Waterman comparison of the query sequence against twenty homopolymers with infinite gap penalties. The output of the algorithm is both the masked query sequence for further analysis, e.g. database searches, as well as the regions of low complexity. The detection of low-complexity regions is highly specific for single residue types. It is shown that this approach is sufficient for masking database query sequences without generating false positives. The algorithm is benchmarked against widely available algorithms using the 210 genes of Plasmodium falciparum chromosome 2, a dataset known to contain a large number of low-complexity regions. AVAILABILITY: CAST (version 1.0) executable binaries are available to academic users free of charge under license. Web site entry point, server and additional material: http://www.ebi.ac.uk/research/cgg/services/cast/
Vasilis J. Promponas, Anton J. Enright, Sophia Tsoka, David P. Kreil, Christophe Leroy, Stavros J. Hamodrakas, Chris Sander, Christos A. Ouzounis
Bioinform.3