Guy-Bart Stan

dblp:52/7139 · also Guy-Bart Vincent Stan · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-5560-902XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 86% Medical and health informatics · 14%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genome annotation
0.912025
Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual Annotation · NeurIPS 2025
Bioinformatics and computational biology › genomics › machine learning for genomics
genomic foundation models
0.912025
Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual Annotation · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.912025
Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual Annotation · NeurIPS 2025
Machine learning › Generative modeling
autoregressive model
0.812024
Absorb & Escape: Overcoming Single Model Limitations in Generating Heterogeneous Genomic Sequences · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
Absorb & Escape: Overcoming Single Model Limitations in Generating Heterogeneous Genomic Sequences · NeurIPS 2024
Bioinformatics and computational biology › clinical bioinformatics
genetic variant classification
0.812024
GV-Rep: A Large-Scale Dataset for Genetic Variant Representation Learning · NeurIPS 2024
Bioinformatics and computational biology › clinical bioinformatics
genetic variant interpretation
0.812024
GV-Rep: A Large-Scale Dataset for Genetic Variant Representation Learning · NeurIPS 2024
Medical and health informatics › clinical informatics
laboratory data management
0.712023
Parsley: a web app for parsing data from plate readers · Bioinform. 2023

Methods — techniques the papers use, named apart from their topics

transformer · 0.9language model pretraining · 0.9fine-tuning · 0.9representation learning · 0.8post-training sampling · 0.8genomic foundation model · 0.8diffusion model · 0.8autoregressive model · 0.8MCMC · 0.8web application · 0.7
YearPublicationVenuePosition
2025 Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual Annotation
abstract
The interpretation of genomic sequences is crucial for understanding biological processes. To handle the growing volume of DNA sequence data, Genomic Foundation Models (GFMs) have been developed by adapting architectures and training paradigms from Large Language Models (LLMs). Despite their remarkable performance in DNA sequence classification tasks, there remains a lack of systematic understanding regarding the training and task-adaptation processes of GFMs. Moreover, existing GFMs cannot achieve state-of-the-art performance on both short and long-context tasks and lacks multimodal abilities. By revisiting pre-training architectures and post-training techniques, we propose **Omni-DNA**, a family of models spanning 20M to 1.1B parameters that supports sequence understanding, long-context genomic reasoning, and natural-language annotation. **Omni-DNA** establishes new state-of-the-art results on 18 of 26 evaluations drawn from Nucleotide Transformer and Genomic Benchmarks. When jointly fine-tuning on biologically related tasks, **Omni-DNA** consistently outperform existing models and demonstrate multi-tasking abilities. To enable processing of arbitrary sequence lengths, we introduce **SEQPACK**—an adaptive compression operator that packs historical tokens into a learned synopsis using a position-aware learnable sampling mechanism, enabling transformer-based models to process ultra-long sequences with minimal memory and computational requirements. Our approach demonstrates superior performance on enhancer-target interaction tasks, capturing distant regulatory interactions at the 450kbp range more effectively than existing models. Finally, we present a new dataset termed **seq2func**, enabling Omni-DNA to generate accurate and functionally meaningful interpretations of DNA sequences, unlocking new possibilities for genomic analysis and discovery.
Vallijah Subasri, Yifei Shen 0004, Dongsheng Li 0002, Wentao Gu, Guy-Bart Stan
NeurIPS6
2024 Enhancing Node Representations for Real-World Complex Networks with Topological Augmentation
abstract
Graph augmentation methods play a crucial role in improving the performance and enhancing generalisation capabilities in Graph Neural Networks (GNNs). Existing graph augmentation methods mainly perturb the graph structures, and are usually limited to pairwise node relations. These methods cannot fully address the complexities of real-world large-scale networks, which often involve higher-order node relations beyond only being pairwise. Meanwhile, real-world graph datasets are predominantly modelled as simple graphs, due to the scarcity of data that can be used to form higher-order edges. Therefore, reconfiguring the higher-order edges as an integration into graph augmentation strategies lights up a promising research path to address the aforementioned issues. In this paper, we present Topological Augmentation (TopoAug), a novel graph augmentation method that builds a combinatorial complex from the original graph by constructing virtual hyperedges directly from the raw data. TopoAug then produces auxiliary node features by extracting information from the combinatorial complex, which are used for enhancing GNN performances on downstream tasks. We design three diverse virtual hyperedge construction strategies to accompany the construction of combinatorial complexes: (1) via graph statistics, (2) from multiple data perspectives, and (3) utilising multi-modality. Furthermore, to facilitate TopoAug evaluation, we provide 23 novel real-world graph datasets across various domains including social media, biology, and e-commerce. Our empirical study shows that TopoAug consistently and significantly outperforms GNN baselines and other graph augmentation methods, across a variety of application contexts, which clearly indicates that it can effectively incorporate higher-order node relations into the graph augmentation for real-world complex networks.
Mingzhu Shen, Guy-Bart Stan, Pietro Liò
ECAI4
2024 Absorb & Escape: Overcoming Single Model Limitations in Generating Heterogeneous Genomic Sequences
abstract
Recent advances in immunology and synthetic biology have accelerated the development of deep generative methods for DNA sequence design. Two dominant approaches in this field are AutoRegressive (AR) models and Diffusion Models (DMs). However, genomic sequences are functionally heterogeneous, consisting of multiple connected regions (e.g., Promoter Regions, Exons, and Introns) where elements within each region come from the same probability distribution, but the overall sequence is non-homogeneous. This heterogeneous nature presents challenges for a single model to accurately generate genomic sequences. In this paper, we analyze the properties of AR models and DMs in heterogeneous genomic sequence generation, pointing out crucial limitations in both methods: (i) AR models capture the underlying distribution of data by factorizing and learning the transition probability but fail to capture the global property of DNA sequences. (ii) DMs learn to recover the global distribution but tend to produce errors at the base pair level. To overcome the limitations of both approaches, we propose a post-training sampling method, termed Absorb & Escape (A&E) to perform compositional generation from AR models and DMs. This approach starts with samples generated by DMs and refines the sample quality using an AR model through the alternation of the Absorb and Escape steps. To assess the quality of generated sequences, we conduct extensive experiments on 15 species for conditional and unconditional DNA generation. The experiment results from motif distribution, diversity checks, and genome integration tests unequivocally show that A&E outperforms state-of-the-art AR models and DMs in genomic sequence generation. A&E does not suffer from the slowness of traditional MCMC to sample from composed distributions with Energy-Based Models whilst it obtains higher quality samples than single models. Our research sheds light on the limitations of current single-model approaches in DNA generation and provides a simple but effective solution for heterogeneous sequence generation. Code is available at the [Github Repo](https://github.com/Zehui127/Absorb-Escape).
Yuhao Ni, Guoxuan Xia, William A. V. Beardall, Akashaditya Das, Guy-Bart Stan
NeurIPS6
2024 GV-Rep: A Large-Scale Dataset for Genetic Variant Representation Learning
abstract
Genetic variants (GVs) are defined as differences in the DNA sequences among individuals and play a crucial role in diagnosing and treating genetic diseases. The rapid decrease in next generation sequencing cost, analogous to Moore’s Law, has led to an exponential increase in the availability of patient-level GV data. This growth poses a challenge for clinicians who must efficiently prioritize patient-specific GVs and integrate them with existing genomic databases to inform patient management. To addressing the interpretation of GVs, genomic foundation models (GFMs) have emerged. However, these models lack standardized performance assessments, leading to considerable variability in model evaluations. This poses the question: *How effectively do deep learning methods classify unknown GVs and align them with clinically-verified GVs?* We argue that representation learning, which transforms raw data into meaningful feature spaces, is an effective approach for addressing both indexing and classification challenges. We introduce a large-scale Genetic Variant dataset, named $\textsf{GV-Rep}$, featuring variable-length contexts and detailed annotations, designed for deep learning models to learn GV representations across various traits, diseases, tissue types, and experimental contexts. Our contributions are three-fold: (i) $\textbf{Construction}$ of a comprehensive dataset with 7 million records, each labeled with characteristics of the corresponding variants, alongside additional data from 17,548 gene knockout tests across 1,107 cell types, 1,808 variant combinations, and 156 unique clinically-verified GVs from real-world patients. (ii) $\textbf{Analysis}$ of the structure and properties of the dataset. (iii) $\textbf{Experimentation}$ of the dataset with pre-trained genomic foundation models (GFMs). The results highlight a significant disparity between the current capabilities of GFMs and the accurate representation of GVs. We hope this dataset will advance genomic deep learning to bridge this gap.
Vallijah Subasri, Guy-Bart Stan
NeurIPS3
2023 Parsley: a web app for parsing data from plate readers
abstract
SUMMARY: As demand for the automation of biological assays has increased over recent years, the range of measurement types implemented by multiwell plate readers has broadened and the list of published software packages that caters to their analysis has grown. However, most plate readers export data in esoteric formats with little or no metadata, while most analytical software packages are built to work with tidy data accompanied by associated metadata. 'Parser' functions are therefore required to prepare raw data for analysis. Such functions are instrument- and data type-specific, and to date, no generic tool exists that can parse data from multiple data types or multiple plate readers, despite the potential for such a tool to speed up access to analysed data and remove an important barrier for less confident coders. We have developed the interactive web application, Parsley, to bridge this gap. Unlike conventional programmatic parser functions, Parsley makes few assumptions about exported data, instead employing user inputs to identify and extract data from data files. In doing so, it is designed to enable any user to parse plate reader data and can handle a wide variety of instruments (10+) and data types (53+). Parsley is freely available via a web interface, enabling access to its unique plate reader data parsing functionality, without the need to install software or write code. AVAILABILITY AND IMPLEMENTATION: The Parsley web application can be accessed at: https://gbstan.shinyapps.io/parsleyapp/. The source code is available at: https://github.com/ec363/parsleyapp and is archived on Zenodo: https://zenodo.org/records/10011752.
Eszter Csibra, Guy-Bart Stan
Bioinform.2