Miguel Graça

dblp:205/9057 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0001-5852-2851ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale Datasets
abstract
DNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need for a high-performant, portable and efficient solution for variant calling, which can further improve our understanding of genomic structure and genetic basis for complex diseases. In its most common formulation, the Pair Hidden Markov Model (PairHMM) algorithm for variant calling stands as the main bottleneck in the pipeline, accounting for up to 70% of the execution time in large-scale genomic datasets. The state-of-the-art approaches for accelerating PairHMM in CPUs and GPUs do not scale to long DNA sequences and only explore very limited anti-diagonal data parallelism, which yields poor performance. In this work, Endeavor is proposed as a new parallelization strategy for PairHMM that redefines its traditional formulation to explore row-level fine-grained parallelism without loss in solution accuracy. Based on this, a novel and portable SIMD-based approach is derived for efficient and high-performance processing of short and long sequences in CPUs and GPUs, leveraging novel levels of parallelism and synchronization to achieve high throughput in sequences up to 100k basepairs for the first time. Evaluation on Intel and AMD CPUs shows that Endeavor outperforms GKL up to 2.14x in peak throughput and GATK HaplotypeCaller by at least 2x in real-world datasets, while NVIDIA and AMD GPUs achieve up to 2.05x speedups in genome-scale datasets when compared to state-of-the-art GPU-based methods.
Miguel Graça, Aleksandar Ilic
HPDC1
2026 TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
Miguel Graça, Aleksandar Ilic
IPDPS1
2025 EPIClear: Exploiting Domain-Specific Features for Epistasis Detection Acceleration on Tensor Cores
abstract
High-order epistasis detection is challenging, making it important to efficiently leverage today's supercomputers.The fastest approaches are those relying on binary precision tensorized operations on modern GPUs.This paper presents a novel approach that significantly surpasses the state-of-theart in high-order epistasis detection by leveraging previously unexplored domain-specific features on the genotype distribution patterns in the dataset.It accelerates time-to-solution with a computational step that reduces the volume of data that needs to be processed to count genotypes.The proposed approach achieves 4× higher performance on a A100 GPU than the previously fastest approach when processing balanced genotype distributions.Evaluation on datasets with unbalanced genotype distributions, which is something that is bound to happen in real datasets, results in significantly higher performance.The proposed accelerating scheme exhibits high scalability.Epistasis detection searches on the MeluXina supercomputer with 32 A100 GPUs resulted in a speedup of up to 30× in comparison to a single GPU, and in achieving a performance scaled to sample size of up to 13 Peta SNP combinations per second for the genotype distribution most unfavorable to the proposed accelerating scheme.
Ricardo Nobre, Miguel Graça, Leonel Sousa, Aleksandar Ilic
ICS2
2025 Bridging Portability and Performance in Sparse Tensor Computations Using SYCL
abstract
ABSTRACT Sparse tensors have become prevalent data structures in multiple applications, such as medical imaging and machine learning, making operations that decompose them, that is, creating smaller structures that retain most of the original information, essential. Two of the most commonly used tensor decomposition methods are the Canonical Polyadic and Tucker Decomposition, with the most time‐consuming operations being the MTTKRP and TTM‐chain, respectively. Modern computing platforms combine multiple devices with different architectures to achieve unprecedented levels of performance, creating an environment where portability is as important as performance. To tackle this challenge, this work proposes SYCL‐based MTTKRP and TTM‐chain approaches for sparse tensors, which are portable to any CPU or GPU, extending previous literature by handling mode‐4 and mode‐5 tensors and tackling the TTM‐chain operation as a whole, allowing for further optimizations. The experimental results show that the proposed approaches present linear to superlinear scalability as the problem size grows and outperform the portable state‐of‐the‐art by 4.9× on average.
Daniel Pacheco, Miguel Graça, Filipe Borralho, Leonel Sousa, Aleksandar Ilic
Concurr. Comput. Pract. Exp.2
2020 When and Why is Unsupervised Neural Machine Translation Useless?
abstract
This paper studies the practicality of the current state-of-the-art unsupervised methods in neural machine translation (NMT). In ten translation tasks with various data settings, we analyze the conditions under which the unsupervised methods fail to produce reasonable translations. We show that their performance is severely affected by linguistic dissimilarity and domain mismatch between source and target monolingual data. Such conditions are common for low-resource language pairs, where unsupervised learning works poorly. In all of our experiments, supervised and semi-supervised baselines with 50k-sentence bilingual data outperform the best unsupervised results. Our analyses pinpoint the limits of the current unsupervised NMT and also suggest immediate research directions.
Yunsu Kim 0001, Miguel Graça, Hermann Ney
EAMT2