VLDB 2026 Research / reviewers in the wild / expert
Susan Walsh
dblp:244/3356
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-7064-1589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimized phenotyping of complex morphological traits: enhancing discovery of common and rare genetic variantsabstractGenotype-phenotype (G-P) analyses for complex morphological traits typically utilize simple, predetermined anatomical measures or features derived via unsupervised dimension reduction techniques (e.g. principal component analysis (PCA) or eigen-shapes). Despite the popularity of these approaches, they do not necessarily reveal axes of phenotypic variation that are genetically relevant. Therefore, we introduce a framework to optimize phenotyping for G-P analyses, such as genome-wide association studies (GWAS) of common variants or rare variant association studies (RVAS) of rare variants. Our strategy is two-fold: (i) we construct a multidimensional feature space spanning a wide range of phenotypic variation, and (ii) within this feature space, we use an optimization algorithm to search for directions or feature combinations that are genetically enriched. To test our approach, we examine human facial shape in the context of GWAS and RVAS. In GWAS, we optimize for phenotypes exhibiting high heritability, estimated from either family data or genomic relatedness measured in unrelated individuals. In RVAS, we optimize for the skewness of phenotype distributions, aiming to detect commingled distributions that suggest single or few genomic loci with major effects. We compare our approach with eigen-shapes as baseline in GWAS involving 8246 individuals of European ancestry and in gene-based tests of rare variants with a subset of 1906 individuals. After applying linkage disequilibrium score regression to our GWAS results, heritability-enriched phenotypes yielded the highest SNP heritability, followed by eigen-shapes, while commingling-based traits displayed the lowest SNP heritability. Heritability-enriched phenotypes also exhibited higher discovery rates, identifying the same number of independent genomic loci as eigen-shapes with a smaller effective number of traits. For RVAS, commingling-based traits resulted in more genes passing the exome-wide significance threshold than eigen-shapes, while heritability-enriched phenotypes lead to only a few associations. Overall, our results demonstrate that optimized phenotyping allows for the extraction of genetically relevant traits that can specifically enhance discovery efforts of common and rare variants, as evidenced by their increased power in facial GWAS and RVAS. Seppe Goovaerts, Myoung K. Lee, Jay Devine, Stephen Richmond, Susan Walsh, Mark D. Shriver, John R. Shaffer, Mary L. Marazita, Hilde Peeters, Seth M. Weinberg, Peter Claes |
Briefings Bioinform. | 6 |
| 2025 | Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging dataabstractMOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD. Zuqi Li, Sam F. L. Windels, Noël Malod-Dognin, Seth M. Weinberg, Mary L. Marazita, Susan Walsh, Mark D. Shriver, David W. Fardo, Peter Claes, Natasa Przulj, Kristel Van Steen |
Bioinform. | 6 |
| 2024 | Mapping genes for human face shape: Exploration of univariate phenotyping strategiesabstractHuman facial shape, while strongly heritable, involves both genetic and structural complexity, necessitating precise phenotyping for accurate assessment. Common phenotyping strategies include simplifying 3D facial features into univariate traits such as anthropometric measurements (e.g., inter-landmark distances), unsupervised dimensionality reductions (e.g., principal component analysis (PCA) and auto-encoder (AE) approaches), and assessing resemblance to particular facial gestalts (e.g., syndromic facial archetypes). This study provides a comparative assessment of these strategies in genome-wide association studies (GWASs) of 3D facial shape. Specifically, we investigated inter-landmark distances, PCA and AE-derived latent dimensions, and facial resemblance to random, extreme, and syndromic gestalts within a GWAS of 8,426 individuals of recent European ancestry. Inter-landmark distances exhibit the highest SNP-based heritability as estimated via LD score regression, followed by AE dimensions. Conversely, resemblance scores to extreme and syndromic facial gestalts display the lowest heritability, in line with expectations. Notably, the aggregation of multiple GWASs on facial resemblance to random gestalts reveals the highest number of independent genetic loci. This novel, easy-to-implement phenotyping approach holds significant promise for capturing genetically relevant morphological traits derived from complex biomedical imaging datasets, and its applications extend beyond faces. Nevertheless, these different phenotyping strategies capture different genetic influences on craniofacial shape. Thus, it remains valuable to explore these strategies individually and in combination to gain a more comprehensive understanding of the genetic factors underlying craniofacial shape and related traits. Seppe Goovaerts, Michiel Vanneste, Harold S. Matthews, Hanne Hoskens, Stephen Richmond, Ophir D. Klein, Richard A. Spritz, Benedikt Hallgrímsson, Susan Walsh, Mark D. Shriver, John R. Shaffer, Seth M. Weinberg, Hilde Peeters, Peter Claes |
PLoS Comput. Biol. | 10 |
| 2023 | Data-driven trait heritability-based extraction of human facial phenotypesabstractA genome-wide association study (GWAS) of a complex, multi-dimensional morphological trait, such as the human face, typically relies on predefined and simplified phenotypic measurements, such as inter-landmark distances and angles. These measures are predominantly designed by human experts based on perceived biological or clinical knowledge. To avoid use handcrafted phenotypes (i.e., a priori expert-identified phenotypes), alternative automatically extracted phenotypic descriptors, such as features derived from dimension reduction techniques (e.g., principal component analysis), are employed. While the features generated by such computational algorithms capture the geometric variations of the biological shape, they are not necessarily genetically relevant. Therefore, genetically informed data-driven phenotyping is desirable. Here, we propose an approach where phenotyping is done through a data-driven optimization of trait heritability, defined as the degree of variation in a phenotypic trait in a population that is due to genetic variation. The resulting phenotyping process consists of two steps: 1) constructing a feature space that models shape variations using dimension reduction techniques, and 2) searching for directions in the feature space exhibiting high trait heritability using a genetic search algorithm (i.e., heuristic inspired by natural selection). We show that the phenotypes resulting from the proposed trait heritability-optimized training differ from those of principal components in the following aspects: 1) higher trait heritability, 2) higher SNP heritability, and 3) identification of the same number of independent genetic loci with a smaller number of effective traits. Our results demonstrate that data-driven trait heritability-based optimization enables the automatic extraction of genetically relevant phenotypes, as shown by their increased power in genome-wide association scans. Seppe Goovaerts, Hanne Hoskens, Stephen Richmond, Susan Walsh, Mark D. Shriver, John R. Shaffer, Mary L. Marazita, Seth M. Weinberg, Hilde Peeters, Peter Claes |
BIBM | 5 |
| 2023 | ILIAD: a suite of automated Snakemake workflows for processing genomic data for downstream applicationsabstractBACKGROUND: Processing raw genomic data for downstream applications such as imputation, association studies, and modeling requires numerous third-party bioinformatics software tools. It is highly time-consuming and resource-intensive with computational demands and storage limitations that pose significant challenges that increase cost. The use of software tools independent of one another, in a disjointed stepwise fashion, increases the difficulty and sets forth higher error rates because of fragmented job executions in alignment, variant calling, and/or build conversion complications. As sequencing data availability grows, the ability for biologists to process it using stable, automated, and reproducible workflows is paramount as it significantly reduces the time to generate clean and reliable data. RESULTS: The Iliad suite of genomic data workflows was developed to provide users with seamless file transitions from raw genomic data to a quality-controlled variant call format (VCF) file for downstream applications. Iliad benefits from the efficiency of the Snakemake best practices framework coupled with Singularity and Docker containers for repeatability, portability, and ease of installation. This feat is accomplished from the onset with download acquisitions of any raw data type (FASTQ, CRAM, IDAT) straight through to the generation of a clean merged data file that can combine any user-preferred datasets using robust programs such as BWA, Samtools, and BCFtools. Users can customize and direct their workflow with one straightforward configuration file. Iliad is compatible with Linux, MacOS, and Windows platforms and scalable from a local machine to a high-performance computing cluster. CONCLUSION: Iliad offers automated workflows with optimized time and resource management that are comparable to other workflows available but generates analysis-ready VCF files from the most common datatypes using a single command. The storage footprint challenge of genomic data is overcome by utilizing temporary intermediate files before the final VCF is generated. This file is ready for use in imputation, genome-wide association study (GWAS) pipelines, high-throughput population genetics studies, select gene candidate studies, and more. Iliad was developed to be portable, compatible, scalable, robust, and repeatable with a simplistic setup, so biologists that are less familiar with programming can manage their own big data with this open-source suite of workflows. Noah Herrick, Susan Walsh |
BMC Bioinform. | 2 |
| 2020 | 3D Facial Matching by Spiral Convolutional Metric Learning and a Biometric Fusion-Net of Demographic PropertiesabstractFace recognition is a widely accepted biometric verification tool, as the face contains a lot of information about the identity of a person. In this study, a 2-step neural-based pipeline is presented for matching 3D facial shape to multiple DNA-related properties (sex, age, BMI and genomic background). The first step consists of a triplet loss-based metric learner that compresses facial shape into a lower dimensional embedding while preserving information about the property of interest. Most studies in the field of metric learning have only focused on 2D Euclidean data. In this work, geometric deep learning is employed to learn directly from 3D facial meshes. To this end, spiral convolutions are used along with a novel mesh-sampling scheme that retains uniformly sampled 3D points at different levels of resolution. The second step is a multi-biometric fusion by a fully connected neural network. The network takes an ensemble of embeddings and property labels as input and returns genuine and imposter scores. Since embeddings are accepted as an input, there is no need to train classifiers for the different properties and available data can be used more efficiently. Results obtained by a to-fold cross-validation for biometric verification show that combining multiple properties leads to stronger biometric systems. Furthermore, the proposed neural-based pipeline outperforms a linear baseline, which consists of principal component analysis, followed by classification with linear support vector machines and a Naïve Bayes-based score-fuser. Soha Sadat Mahdi, Nele Nauwelaers, Philip Joris, Giorgos Bouritsas, Shunwang Gong, Sergiy Bokhnyak, Susan Walsh, Mark D. Shriver, Michael M. Bronstein, Peter Claes |
ICPR | 7 |
| 2019 | Automatic Landmark Placement for Large 3D Facial Image DatasetabstractFacial landmark placement is a key step in many biomedical and biometrics applications. This paper presents a computational method that efficiently performs automatic 3D facial landmark placement based on training images containing manually placed anthropological facial landmarks. After 3D face registration by an iterative closest point (ICP) technique, a visual analytics approach is taken to generate local geometric patterns for individual landmark points. These individualized local geometric patterns are derived interactively by a user's initial visual pattern detection. They are used to guide the refinement process for landmark points projected from a template face to achieve accurate landmark placement. Compared to traditional methods, this technique is simple, robust, and does not require a large number of training samples (e.g. in machine learning based methods) or complex 3D image analysis procedures. This technique and the associated software tool are being used in a 3D biometrics project that aims to identify links between human facial phenotypes and their genetic association. Jerry Wang, Shiaofen Fang, Meie Fang, Jeremy Wilson, Noah Herrick, Susan Walsh |
IEEE BigData | 6 |
| 2019 | Odyssey: a semi-automated pipeline for phasing, imputation, and analysis of genome-wide genetic dataabstractBACKGROUND: Genome imputation, admixture resolution and genome-wide association analyses are timely and computationally intensive processes with many composite and requisite steps. Analysis time increases further when building and installing the run programs required for these analyses. For scientists that may not be as versed in programing language, but want to perform these operations hands on, there is a lengthy learning curve to utilize the vast number of programs available for these analyses. RESULTS: In an effort to streamline the entire process with easy-to-use steps for scientists working with big data, the Odyssey pipeline was developed. Odyssey is a simplified, efficient, semi-automated genome-wide imputation and analysis pipeline, which prepares raw genetic data, performs pre-imputation quality control, phasing, imputation, post-imputation quality control, population stratification analysis, and genome-wide association with statistical data analysis, including result visualization. Odyssey is a pipeline that integrates programs such as PLINK, SHAPEIT, Eagle, IMPUTE, Minimac, and several R packages, to create a seamless, easy-to-use, and modular workflow controlled via a single user-friendly configuration file. Odyssey was built with compatibility in mind, and thus utilizes the Singularity container solution, which can be run on Linux, MacOS, and Windows platforms. It is also easily scalable from a simple desktop to a High-Performance System (HPS). CONCLUSION: Odyssey facilitates efficient and fast genome-wide association analysis automation and can go from raw genetic data to genome: phenome association visualization and analyses results in 3-8 h on average, depending on the input data, choice of programs within the pipeline and available computer resources. Odyssey was built to be flexible, portable, compatible, scalable, and easy to setup. Biologists less familiar with programing can now work hands on with their own big data using this easy-to-use pipeline. Ryan J. Eller, Sarath Chandra Janga, Susan Walsh |
BMC Bioinform. | 3 |