EDBT 2026 Demo / reviewers in the wild / expert
Aleksandar Ilic
dblp:83/1727
· DBLP profile ↗
60ranked-venue papers
15as first author
18since 2021 · last 2026
0000-0002-8594-3539ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 4 first-author · 15 since 2021Theory of computation · 11 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale DatasetsabstractDNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need for a high-performant, portable and efficient solution for variant calling, which can further improve our understanding of genomic structure and genetic basis for complex diseases. In its most common formulation, the Pair Hidden Markov Model (PairHMM) algorithm for variant calling stands as the main bottleneck in the pipeline, accounting for up to 70% of the execution time in large-scale genomic datasets. The state-of-the-art approaches for accelerating PairHMM in CPUs and GPUs do not scale to long DNA sequences and only explore very limited anti-diagonal data parallelism, which yields poor performance. In this work, Endeavor is proposed as a new parallelization strategy for PairHMM that redefines its traditional formulation to explore row-level fine-grained parallelism without loss in solution accuracy. Based on this, a novel and portable SIMD-based approach is derived for efficient and high-performance processing of short and long sequences in CPUs and GPUs, leveraging novel levels of parallelism and synchronization to achieve high throughput in sequences up to 100k basepairs for the first time. Evaluation on Intel and AMD CPUs shows that Endeavor outperforms GKL up to 2.14x in peak throughput and GATK HaplotypeCaller by at least 2x in real-world datasets, while NVIDIA and AMD GPUs achieve up to 2.05x speedups in genome-scale datasets when compared to state-of-the-art GPU-based methods. Miguel Graça, Aleksandar Ilic |
HPDC | 2 |
| 2026 | TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
Miguel Graça, Aleksandar Ilic |
IPDPS | 2 |
| 2025 | Thievory: Graph Processing with Multi-GPU Memory StealingabstractGraph workloads have gained widespread attention across diverse applications. These workloads often display large memory footprints while offering significant parallelism opportunities in graph traversals. To exploit this parallelism, the state-of-the-art approaches focus on single-GPU environments and/or out-of-memory processing. However, these approaches are individually insufficient to address the challenges of modern GPU-accelerated graph processing. First, the ever-increasing graph size can span well beyond the available GPU memory. Second, modern computing platforms are de facto multi-GPU, thus bringing a new set of challenges to be taken into account for taming the highly irregular nature of modern large graphs. Novel approaches must be devised to achieve efficient multi-device graph processing by exploiting the fast cross-device interconnects and seamless peer-to-peer communication. This work proposes Thievory – the first out-of-memory graph processing framework that explores novel execution patterns in multi-GPU environments. In particular, we explore the use of GPUs as data buffers, effectively using spare neighbor GPUs’ memory to boost out-of-memory graph traversals, thus increasing resource efficiency in these multi-GPU environments. Moreover, we also analyse a rich set of data transfer optimizations and caching mechanisms. Thievory shows up to 10.5× speedups over existing state-of-the-art approaches when evaluated with a representative set of graphs. João Brotas, Ricardo Nobre, Aleksandar Ilic |
ICPP | 3 |
| 2025 | EPIClear: Exploiting Domain-Specific Features for Epistasis Detection Acceleration on Tensor CoresabstractHigh-order epistasis detection is challenging, making it important to efficiently leverage today's supercomputers.The fastest approaches are those relying on binary precision tensorized operations on modern GPUs.This paper presents a novel approach that significantly surpasses the state-of-theart in high-order epistasis detection by leveraging previously unexplored domain-specific features on the genotype distribution patterns in the dataset.It accelerates time-to-solution with a computational step that reduces the volume of data that needs to be processed to count genotypes.The proposed approach achieves 4× higher performance on a A100 GPU than the previously fastest approach when processing balanced genotype distributions.Evaluation on datasets with unbalanced genotype distributions, which is something that is bound to happen in real datasets, results in significantly higher performance.The proposed accelerating scheme exhibits high scalability.Epistasis detection searches on the MeluXina supercomputer with 32 A100 GPUs resulted in a speedup of up to 30× in comparison to a single GPU, and in achieving a performance scaled to sample size of up to 13 Peta SNP combinations per second for the genotype distribution most unfavorable to the proposed accelerating scheme. Ricardo Nobre, Miguel Graça, Leonel Sousa, Aleksandar Ilic |
ICS | 4 |
| 2025 | Bridging Portability and Performance in Sparse Tensor Computations Using SYCLabstractABSTRACT Sparse tensors have become prevalent data structures in multiple applications, such as medical imaging and machine learning, making operations that decompose them, that is, creating smaller structures that retain most of the original information, essential. Two of the most commonly used tensor decomposition methods are the Canonical Polyadic and Tucker Decomposition, with the most time‐consuming operations being the MTTKRP and TTM‐chain, respectively. Modern computing platforms combine multiple devices with different architectures to achieve unprecedented levels of performance, creating an environment where portability is as important as performance. To tackle this challenge, this work proposes SYCL‐based MTTKRP and TTM‐chain approaches for sparse tensors, which are portable to any CPU or GPU, extending previous literature by handling mode‐4 and mode‐5 tensors and tackling the TTM‐chain operation as a whole, allowing for further optimizations. The experimental results show that the proposed approaches present linear to superlinear scalability as the problem size grows and outperform the portable state‐of‐the‐art by 4.9× on average. Daniel Pacheco, Miguel Graça, Filipe Borralho, Leonel Sousa, Aleksandar Ilic |
Concurr. Comput. Pract. Exp. | 5 |
| 2025 | SpEpistasis: A sparse approach for three-way epistasis detection
Diogo Marques 0003, Leonel Sousa, Aleksandar Ilic |
J. Parallel Distributed Comput. | 3 |
| 2024 | External Memory Protection on FPGA-Based Embedded SystemsabstractWith the proliferation and increased capabilities of embedded systems, they become more exposed and easier targets to attacks. This includes the external components such as DRAM, particularly vulnerable to attacks, especially regarding unauthorized access to the stored data. With the goal of increasing storage security, this paper proposes a memory bridge evaluation platform and an improvement to the authenticated-encryption of off-chip memory, minimizing critical memory access latency. Most state-of-the-art works either use custom high-latency solutions, or frequently recur exclusively to AES-GCM standard for Authenticated Encryption with Associated Data. Besides AES-GCM, this work explores and implements different protection solutions, including NOEKEON-GCM and AEGIS-128L. This work also explores how much of the algorithms can be pre-computed between memory transmissions, by removing the address and data from critical path computations, placing it in the AAD field. The presented prototypes were evaluated on a Xilinx Zynq-7100, showing that the proposed platform allows to analyse and assess different approaches. The obtained experimental results suggest that with a careful selection of algorithms and implementations, memory access latency improvements up to 56% can be achieved in regard to equivalent AES-GCM designs. João Carlos Resende, Aleksandar Ilic, Ricardo Chaves |
DSD | 2 |
| 2024 | IPU-EpiDet: Identifying Gene Interactions on Massively Parallel Graph-Based AI AcceleratorsabstractEpistasis detection is a bioinformatics application that searches for associations between sets of single nucleotide polymorphisms (SNPs) and a given trait of a population. Epistasis detection is a computationally complex problem, especially when tackling interaction orders above two. This paper presents a pioneering approach for performing third-order epistasis searches that has been devised around the bulk-synchronous parallel (BSP) model of execution, which is used in processor designs that target artificial intelligence (AI) workloads, such as Graphcore’s Intelligence Processing Unit (IPU). We propose a parallelization approach and a set of optimizations to efficiently exploit the computation and communication resources of these novel AI processors to perform precise bioinformatics searches. The proposed approach achieves a performance of 3.9 Tera SNP combinations evaluated per second, scaled to sample size, on an IPU-M2000 AI accelerator, while close to a linear speedup (up to 15.01×) is achieved with 16 IPU-M2000 accelerators. Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa |
IPDPS | 2 |
| 2024 | Maximal diameter of integral circulant graphs
Milan Basic, Aleksandar Ilic, Aleksandar Stamenkovic |
Inf. Comput. | 2 |
| 2023 | Bringing Order to Sparsity: A Sparse Matrix Reordering Study on Multicore CPUsabstractMany real-world computations involve sparse data structures in the form of sparse matrices. A common strategy for optimizing sparse matrix operations is to reorder a matrix to improve data locality. However, it's not always clear whether reordering will provide benefits over the unordered matrix, as its effectiveness depends on several factors, such as structural features of the matrix, the reordering algorithm and the hardware that is used. This paper aims to establish the relationship between matrix reordering algorithms and the performance of sparse matrix operations. We thoroughly evaluate six different matrix reordering algorithms on 490 matrices across eight multicore architectures, focusing on the commonly used sparse matrix-vector multiplication (SpMV) kernel. We find that reordering based on graph partitioning provides better SpMV performance than the alternatives for a large majority of matrices, and that the resulting performance is explained through a combination of data locality and load balancing concerns. James D. Trotter, Sinan Ekmekçibasi, Johannes Langguth, Tugba Torun, Emre Düzakin, Aleksandar Ilic, Didem Unat |
SC | 6 |
| 2023 | Special issue: 20th international workshop on algorithms, models and tools for parallel computing on heterogeneous platforms (HeteroPar'22)abstractHeterogeneity is emerging as one of the most profound and challenging characteristics of parallel environments. From the macro level, where distributed systems are built around heterogeneous networks connecting multiple nodes with computing devices of diverse architectures, to the micro level, where ever-deeper memory hierarchies and specialized accelerators are increasingly common, the impact of heterogeneity on parallel processing is rapidly increasing. Traditional parallel algorithms, programming environments and tools designed for legacy homogeneous multiprocessors achieve at best a small fraction of the efficiency and performance expected from highly heterogeneous parallel computing systems. Therefore, innovative models, algorithms, programming environments and tools are required to efficiently tackle the challenges and fully exploit the resources of modern parallel and heterogeneous platforms. The international workshop on algorithms, models and tools for parallel computing on heterogeneous platforms (HeteroPar) is a premium forum for researchers on algorithms, programming languages, tools, and theoretical models for efficiently solving complex problems on heterogeneous parallel platforms. The 14th edition of HeteroPar (2022) took place in Glasgow, Scotland, co-located with the Euro-Par annual international conference. The workshop includes one keynote and 11 technical presentations. The selected papers cover a good spectrum of research topics in heterogeneous computing showing the challenges present on these modern platforms, hopefully indicating to interested readers possible directions for further research in this field. On behalf of every reader of this Special Issue, the Guest Editors would like to thank all the authors who submitted their papers and worked hard to respond to Reviewers' requests in due time, all the anonymous reviewers who participated in the review process providing helpful suggestions, as well as the Editor in Chief and the entire staff of Wiley's Concurrency and Computation: Practice and Experience who oversaw the whole process. Aleksandar Ilic, Leonel Sousa |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | Tensor-Accelerated Fourth-Order Epistasis Detection on GPUsabstractThe improved accessibility of gene sequencing technologies has led to creation of huge datasets, i.e. patient records related to certain human diseases (phenotypes). Hence, deriving fast and accurate algorithms for efficiently processing these datasets is a paramount concern to enable some key healthcare scenarios, such as personalizing treatments, explaining the occurrence of and/or susceptibility to complex conditions and reducing the spread of infectious diseases. This is especially true for high-order epistasis detection, one of the most computationally challenging problems in bioinformatics, where associations between a given phenotype and single nucleotide polymorphisms (SNPs) of a population can often only be uncovered through evaluation of a large number of SNP combinations. To tackle this challenge, we propose a novel fourth-order epistasis detection algorithm that leverages tensor processing capabilities of two distinct accelerator architectures by efficiently mapping core computations related to processing quads of SNPs to binary tensor-accelerated matrix operations. Experimental results show that the proposed approach delivers very high performance even in single-GPU environments, e.g., 27.8 and 90.9 tera quads of SNPs per second, scaled to the sample size, were processed on Titan RTX (Turing) and A100 (Ampere) PCIe GPUs, respectively. Being the first approach that exploits tensor cores for accelerating searches with interaction order above three, the proposed method achieved a performance of up to 835.4 tera quads of SNPs per second on the 8-GPU HGX A100 server, which represents performance two or more orders of magnitude higher than that of related art. Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa |
ICPP | 2 |
| 2022 | Unlocking Personalized Healthcare on Modern CPUs/GPUs: Three-way Gene Interaction StudyabstractDevelopments in Genome-Wide Association Studies have led to the increasing notion that future healthcare techniques will be personalized to the patient, by relying on genetic tests to determine the risk of developing a disease. To this end, the detection of gene interactions that cause complex diseases constitutes an important application. Similarly to many applications in this field, extensive data sets containing genetic information for a series of patients are used (such as Single-Nucleotide Polymorphisms), leading to high computational complexity and memory utilization, thus constituting a major challenge when targeting high-performance execution in modern computing systems. To close this gap, this work proposes several novel approaches for the detection of three-way gene interactions in modern CPUs and GPUs, making use of different optimizations to fully exploit the target architectures. Crucial insights from the Cache-Aware Roofline Model are used to ensure the suitability of the applications to the computing devices. An extensive study of the architectural features of 13 CPU and GPU devices from all main vendors is also presented, allowing to understand the features relevant to obtain high-performance in this bioinformatics domain. To the best of our knowledge, this study is the first to perform such evaluation for epistasis detection. The proposed approaches are able to surpass the performance of state-of-the-art works in the tested platforms, achieving an average speedup of 3.9× (7.3× on CPUs and 2.8× on GPUs) and maximum speedup of 10.6× on Intel UHD P630 GPU. Diogo Marques 0003, Rafael Campos, Sergio Santander-Jiménez, Zakhar Matveev, Leonel Sousa, Aleksandar Ilic |
IPDPS | 6 |
| 2021 | HEDAcc: FPGA-based Accelerator for High-order Epistasis DetectionabstractThe manifestation of important genetic diseases is often a consequence of the interactions between Single Nucleotide Polymorphisms (SNPs), also known as epistasis. Detecting epistasis for high-order interactions results in a huge computational complexity, as the number of SNP combinations to be evaluated exponentially grows with the interaction order. To address this challenge, state-of-the-art exhaustive search-based methods for epistasis detection rely on GPUs and FPGAs to provide high-performance solutions tailored for a specific order (second and rarely third-order interactions) and/or specific data-set sizes. In this paper, a novel parameterizable architecture is proposed that enables the deployment of FPGA-based accelerators targeting any order of interactions and any data-set size. By relying on a set of algorithmic and architecture optimizations, the proposed accelerator showed to outperform current FPGA accelerators for second and third-order interactions by as much as 4.6× and 9.5×, respectively. The proposed solution also showed comparable performance to current GPGPU third-order implementations, while consuming up to 8.2× less energy. Finally, the proposed architecture allowed for the implementation of a fourth-order epistasis detection accelerator in an FPGA platform. Gaspar Ribeiro, Nuno Neves 0002, Sergio Santander-Jiménez, Aleksandar Ilic |
FCCM | 4 |
| 2021 | Fourth-Order Exhaustive Epistasis Detection for the xPU EraabstractThe investigation of highly-efficient parallel algorithms targeting modern heterogeneous systems can provide bioinformaticians new mechanisms to find relations between genetics, phenotype and environment. This paper proposes an approach for fourth-order exhaustive, i.e. as precise as possible, epistasis detection targeting modern heterogeneous systems. Being implemented in Data Parallel C++ / SYCL, the proposed approach relies on technologies and tools built around open standards, making it able to target different types of architectures and devices. As a means to show interoperability with hardware from different sources, we have included performance results obtained from execution on different systems. Scaled to the number of samples, the proposed approach achieved a performance per GPU stream core of up to 236, 472 or 487 mega quads of SNPs processed per second, on GPUs with the Gen9.5, Gen12 and Turing architectures. This metric is reported for different implementation variants combining different considered optimizations. The proposed approach is able to target a wide range of CPU and GPU devices, enabling more users to access high-throughput epistasis detection software. Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa |
ICPP | 2 |
| 2021 | On conjectures of network distance measures by using graph spectra
Aleksandar Ilic, Modjtaba Ghorbani, Seyran Azizi, Matthias Dehmer |
Discret. Appl. Math. | 1 |
| 2021 | Comparing Wiener complexity with eccentric complexity
Kexiang Xu, Aleksandar Ilic, Vesna Irsic Chenoweth, Sandi Klavzar |
Discret. Appl. Math. | 2 |
| 2021 | Retargeting Tensor Accelerators for Epistasis DetectionabstractThe substitution of nucleotides at specific positions in the genome of a population, known as single-nucleotide polymorphisms (SNPs), has been correlated with a number of important diseases. Complex conditions such as Alzheimer's disease or Crohn's disease are significantly linked to genetics when the impact of multiple SNPs is considered. SNPs often interact in an epistatic manner, where the joint effect of multiple SNPs may not be simply mapped to a linear additive combination of individual effects. Genome-wide association studies considering epistasis are computationally challenging, especially when performing triplet searches is required. Some contemporary computer architectures support fused XOR and population count as the highest throughput operations as part of tensor operations. This article presents a new approach for efficiently repurposing this capability to accelerate 2-way (pairs) and 3-way (triplets) epistasis detection searches. Experimental evaluation targeting the Turing GPU architecture resulted in previously unattainable levels of performance, with the proposal being able to evaluate up to 108.1 and 54.5 tera unique sets of SNPs per second, scaled to the sample size, in 2-way and 3-way searches, respectively. Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Heterogeneous CPU+iGPU Processing for Efficient Epistasis Detection
Rafael Campos, Diogo Marques 0003, Sergio Santander-Jiménez, Leonel Sousa, Aleksandar Ilic |
Euro-Par | 5 |
| 2020 | Performance Optimization and Scalability Analysis of the MGB Hydrological ModelabstractHydrological models are extensively used in applications such as water resources, climate change, land use, and forecast systems. The focus of this paper is performance optimization of the MGB hydrological model, which is widely employed to simulate water flows in large-scale watersheds. The optimization strategies that we selected include AVX-512 vectorization, thread-parallelism on multi-core CPUs (OpenMP), and data-parallelism on many-core GPUs (CUDA). We conducted experiments for real-world input datasets on state-of-the-art HPC systems based on Intel's Skylake CPUs and NVIDIA GPUs. In addition, a Roofline model characterization for these datasets confirmed performance improvements of up to 37.5x on the most time-consuming part of the code and 8.6x on the full MGB model. The work proposed herein shows that careful optimizations are needed for hydrological models to achieve a significant fraction of the performance potential in modern processors. Henrique Rennó de Azeredo Freitas, Celso L. Mendes, Aleksandar Ilic |
HiPC | 3 |
| 2020 | Exploring the Binary Precision Capabilities of Tensor Cores for Epistasis DetectionabstractGenome-wide association studies are performed to correlate a number of diseases and other physical or even psychological conditions (phenotype) with substitutions of nucleotides at specific positions in the human genome, mainly single-nucleotide polymorphisms (SNPs). Some conditions, possibly because of the complexity of the mechanisms that give rise to them, have been identified to be more statistically correlated with genotype when multiple SNPs are jointly taken into account. However, the discovery of new associations between genotype and phenotype is exponentially slowed down by the increase of computational power required when epistasis, i.e., interactions between SNPs, is considered. This paper proposes a novel graphics processing unit (GPU)-based approach for epistasis detection that combines the use of modern tensor cores with native support for processing binarized inputs with algorithmic and target-focused optimizations. Using only a single mid-range Turing-based GPU, the proposed approach is able to evaluate 64.8×1012and 25.4×1012sets of SNPs per second, normalized to the number of patients, when considering 2-way and 3-way epistasis detection, respectively. This proposal is able to surpass the state-of-the-art approach by 6× and 8.2× in terms of the number of pairs and triplets of SNP allelic patient data evaluated per unit of time per GPU. Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa |
IPDPS | 2 |
| 2020 | Accelerating 3-Way Epistasis Detection with CPU+GPU Processing
Ricardo Nobre, Sergio Santander-Jiménez, Leonel Sousa, Aleksandar Ilic |
JSSPP | 4 |
| 2020 | Application-driven Cache-Aware Roofline Model
Diogo Marques 0003, Aleksandar Ilic, Zakhar Matveev, Leonel Sousa |
Future Gener. Comput. Syst. | 2 |
| 2019 | HeTM: Transactional Memory for Heterogeneous SystemsabstractModern heterogeneous computing architectures, which couple multi-core CPUs with discrete many-core GPUs (or other specialized hardware accelerators), enable unprecedented peak performance and energy efficiency levels. However, developing applications that can take full advantage of the potential of heterogeneous systems is a notoriously hard task. This work takes a step towards reducing the complexity of programming heterogeneous systems by introducing the abstraction of Heterogeneous Transactional Memory (HeTM). HeTM provides programmers with the illusion of a single memory region, shared among the CPUs and the (discrete) GPU(s) of a heterogeneous system, with support for atomic transactions. Besides introducing the abstract semantics and programming model of HeTM, we present the design and evaluation of a concrete implementation of the proposed abstraction, referred herein as Speculative HeTM (SHeTM). SHeTM makes use of a novel design that leverages speculative techniques, which aims at hiding the inherently large communication latency between CPUs and discrete GPUs and at minimizing inter-device synchronization overhead. We demonstrate the efficiency of the SHeTM via an extensive quantitative study based both on synthetic benchmarks and on a popular object caching system. Daniel Castro 0004, Paolo Romano 0002, Aleksandar Ilic, Amin M. Khan |
PACT | 3 |
| 2019 | Fast block distributed CUDA implementation of the Hungarian algorithm
Paulo Alexandre Crisóstomo Lopes, Satyendra Singh Yadav, Aleksandar Ilic, Sarat Kumar Patra |
J. Parallel Distributed Comput. | 3 |
| 2019 | DVFS-aware application classification to improve GPGPUs energy efficiency
João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás |
Parallel Comput. | 2 |
| 2019 | Modeling Non-Uniform Memory Access on Large Compute Nodes with the Cache-Aware Roofline ModelabstractNUMA platforms, emerging memory architectures with on-package high bandwidth memories bring new opportunities and challenges to bridge the gap between computing power and memory performance. Heterogeneous memory machines feature several performance trade-offs, depending on the kind of memory used, when writing or reading it. Finding memory performance upper-bounds subject to such trade-offs aligns with the numerous interests of measuring computing system performance. In particular, representing applications performance with respect to the platform performance bounds has been addressed in the state-of-the-art Cache-Aware Roofline Model (CARM) to troubleshoot performance issues. In this paper, we present a Locality-Aware extension (LARM) of the CARM to model NUMA platforms bottlenecks, such as contention and remote access. On top of this, the new contribution of this paper is the design and validation of a novel hybrid memory bandwidth model. This new hybrid model quantifies the achievable bandwidth upper-bound under above-described trade-offs with less than 3 percent error. Hence, when comparing applications performance with the maximum attainable performance, software designers can now rely on more accurate information. Nicolas Denoyelle, Brice Goglin, Aleksandar Ilic, Emmanuel Jeannot, Leonel Sousa |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Modeling and Decoupling the GPU Power Consumption for Cross-Domain DVFSabstractDynamic voltage and frequency scaling (DVFS) is a popular technique to improve the energy-efficiency of high-performance computing systems. It allows placing the devices into lower performance states when the computational demands are lower, opening the possibility for significant power/energy savings. This work presents a GPU power consumption model, used to predict the GPU power consumption of any application at different frequency levels. To obtain this model, an estimation algorithm is proposed, relying on careful benchmarking of the GPU architecture. The model can estimate the contribution of twelve different GPU components (FP32-ADD/MUL/FMA, FP64-ADD/MUL/FMA, INT, SF, CF units, shared memory, L2-cache, and DRAM) to the GPU power consumption. Different model use cases are evaluated (fixed-frequency, DVFS, and scaling-factors), which can obtain both the total or the per-component GPU power consumption. A technique to export models to a distinct GPU from the one it was estimated on is also proposed. These approaches were extensively validated on five different GPUs from the three most recent microarchitectures with a set of 42 standard benchmarks, achieving very accurate predictions. In particular, the scaling-factor power model achieves an average prediction error of 3.5 percent (Titan Xp), 4.6 percent (GTX Titan X), 3.1 percent (GTX 980) and 2.4 percent (Tesla K40c). João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | GPGPU Power Modeling for Multi-domain Voltage-Frequency ScalingabstractDynamic Voltage and Frequency Scaling (DVFS) on Graphics Processing Units (GPUs) components is one of the most promising power management strategies, due to its potential for significant power and energy savings. However, there is still a lack of simple and reliable models for the estimation of the GPU power consumption under a set of different voltage and frequency levels. Accordingly, a novel GPU power estimation model with both core and memory frequency scaling is herein proposed. This model combines information from both the GPU architecture and the executing GPU application and also takes into account the non-linear changes in the GPU voltage when the core and memory frequencies are scaled. The model parameters are estimated using a collection of 83 microbenchmarks carefully crafted to stress the main GPU components. Based on the hardware performance events gathered during the execution of GPU applications on a single frequency configuration, the proposed model allows to predict the power consumption of the application over a wide range of frequency configurations, as well as to decompose the contribution of different parts of the GPU pipeline to the overall power consumption. Validated on 3 GPU devices from the most recent NVIDIA microarchitectures (Pascal, Maxwell and Kepler), by using a collection of 26 standard benchmarks, the proposed model is able to achieve accurate results (7%, 6% and 12% mean absolute error) for the target GPUs (Titan Xp, GTX Titan X and Tesla K40c). João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás |
HPCA | 2 |
| 2018 | Highly parallel HEVC decoding for heterogeneous systems with CPU and GPUabstractThe High Efficiency Video Coding HEVC standard provides a higher compression efficiency than other video coding standards but at the cost of an increased computational load, which makes hard to achieve real-time encoding/decoding for ultra high-resolution and high-quality video sequences. Graphics Processing Units GPU are known to provide massive processing capability for highly parallel and regular computing kernels, but not all HEVC decoding procedures are suited for GPU execution. Furthermore, if HEVC decoding is accelerated by GPUs, energy efficiency is another concern for heterogeneous CPU+GPU decoding. In this paper, a highly parallel HEVC decoder for heterogeneous CPU+GPU system is proposed. It exploits available parallelism in HEVC decoding on the CPU, GPU, and between the CPU and GPU devices simultaneously. On top of that, different workload balancing schemes can be selected according to the devoted CPU and GPU computing resources. Furthermore, an energy optimized solution is proposed by tuning GPU clock rates. Results show that the proposed decoder achieves better performance than the state-of-the-art CPU decoder, and the best performance among the workload balancing schemes depends on the available CPU and GPU computing resources. In particular, with an NVIDIA Titan X Maxwell GPU and an Intel Xeon E5-2699v3 CPU, the proposed decoder delivers 167 frames per second (fps) for Ultra HD 4K videos, when four CPU cores are used. Compared to the state-of-the-art CPU decoder using four CPU cores, the proposed decoder gains a speedup factor of 2.2×. When decoding performance is bounded by the CPU, a system wise energy reduction up to 36% is achieved by using fixed (and lower) GPU clocks, compared to the default dynamic clock settings on the GPU. Biao Wang 0001, Diego F. de Souza, Mauricio Alvarez-Mesa, Chi Ching Chi, Ben H. H. Juurlink, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
Signal Process. Image Commun. | 6 |
| 2017 | Exploring GPU performance, power and energy-efficiency bounds with Cache-aware Roofline ModelingabstractOptimization, portability and development of GPGPU applications are not trivial tasks, since the capabilities and organization of GPU processing elements and memory subsystem greatly differ from the traditional CPU concepts, as well as among different GPU architectures. This work goes a step further in aiding this process by delivering a set of visual models that can be used by GPU programmers to analyze and improve application performance and energy-efficiency across a range of different GPU devices. For the first time in this paper, the state-of-the-art Cache-aware Roofline Modeling principles are applied for insightful modeling of GPU upper-bounds for performance, power consumption and energy-efficiency. The proposed models are developed by relying on extensive GPU micro-benchmarking aimed at fully exercising the capabilities of GPU functional units and memory hierarchy levels. The models are experimentally validated across 8 GPU devices from 3 different NVIDIA generations, and their benefits are explored when characterizing the behavior of 23 real-world applications from 5 different benchmark suites. Furthermore, the DVFS effects on GPU performance upper-bounds are also analyzed by scaling both core and memory frequencies. Andre Lopes, Frederico Pratas, Leonel Sousa, Aleksandar Ilic |
ISPASS | 4 |
| 2017 | Energy-aware mechanism for stencil-based MPDATA algorithm with constraintsabstractSummary In this paper, we propose an energy‐aware task management mechanism designed for the forward‐in‐time algorithms running on multicore central processing units (CPUs), where the multidimensional positive definite advection transport algorithm stencil‐based algorithm is one of the representative examples. This mechanism is based on the dynamic voltage and frequency scaling technique and allows the reduction of energy consumption for an existing algorithm (or application) such that the predefined execution time is respected, without requiring any modifications in the algorithm itself. This paper also provides the formulation of a method for minimizing the energy consumption with time constraints, which is based on the adaptive scheduling with online modeling. Finally, using the autotuning technique, we provide the automation of the process for creation and determination of the best energy profile at runtime, even in the presence of additional CPU workloads. The experimental results on a 6‐core computing platform show that the proposed mechanism provides the energy savings of up to 1.43x when compared to the default Linux scaling governor. Also, we confirm the effectiveness of the self‐adaptive feature of the proposed mechanism, by showing its ability to maintain the requested execution time in spite of additional CPU workloads imposed by other applications. Krzysztof Rojek, Aleksandar Ilic, Roman Wyrzykowski, Leonel Sousa |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Accelerating the phylogenetic parsimony function on heterogeneous systemsabstractSummary The availability of heterogeneous CPU+GPU systems has opened the door to new opportunities for the development of parallel solutions to tackle complex biological problems. The reconstruction of evolutionary histories among species represents a grand computational challenge, which can be addressed by exploiting this kind of hardware designs. In this research, we study the application of heterogeneous computing with OpenCL to accelerate one of the most well‐known objective functions for inferring phylogenies, the phylogenetic parsimony function. For this purpose, we undertake the design of CPU and GPU kernel implementations of this relevant function, proposing a heterogeneous CPU+GPU multidevice approach that distributes multiple parsimony evaluations among processing devices. Experiments on 6 real nucleotide data sets and comparisons with other parallel implementations give account of the benefits of the proposal in this paper, obtaining significant parallel results by combining CPU and GPU capabilities in accordance with the characteristics of the input data. Sergio Santander-Jiménez, Aleksandar Ilic, Leonel Sousa, Miguel A. Vega-Rodríguez |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Beyond the Roofline: Cache-Aware Power and Energy-Efficiency Modeling for Multi-CoresabstractTo foster the energy-efficiency in current and future multi-core processors, the benefits and trade-offs of a large set of optimization solutions must be evaluated. For this purpose, it is often crucial to consider how key micro-architecture aspects, such as accessing different memory levels and functional units, affect the attainable power and energy consumption. To ease this process, we propose a set of insightful cache-aware models to characterize the upper-bounds for power, energy and energy-efficiency of modern multi-cores in three different domains of the processor chip: cores, uncore and package. The practical importance of the proposed models is illustrated when optimizing matrix multiplication and deriving a set of power envelopes and energy-efficiency ranges of the micro-architecture for different operating frequencies. The proposed models are experimentally validated on a computing platform with a quad-core Intel 3770K processor by using hardware counters, on-chip power monitoring facilities and assembly micro-benchmarks. Aleksandar Ilic, Frederico Pratas, Leonel Sousa |
IEEE Trans. Computers | 1 |
| 2017 | GHEVC: An Efficient HEVC Decoder for Graphics Processing UnitsabstractThe high compression efficiency that is provided by the high efficiency video coding (HEVC) standard comes at the cost of a significant increase of the computational load at the decoder. Such an increased burden is a limiting factor to accomplish real-time decoding, specially for high definition video sequences (e.g., Ultra HD 4K). In this scenario, a highly parallel HEVC decoder for the state-of-the-art graphics processor units (GPUs) is presented, i.e., GHEVC. Contrasting to our previous contributions, the data-parallel GHEVC decoder integrates the whole decompression pipeline (except for the entropy decoding), both for intra- and interframes. Furthermore, its processing efficiency was highly optimized by keeping the decompressed frames in the GPU memory for subsequent inter frame prediction. The proposed GHEVC decoder is fully compliant with the HEVC standard, where explicit synchronization points ensure the correct HEVC module execution order. Moreover, the GPU-based HEVC decoder is experimentally evaluated for different GPU devices, an extensive range of recommended HEVC configurations and video sequences, where an average frame rate of 145, 318, and 605 frames per second for Ultra HD 4K, WQXGA, and Full HD, respectively, was obtained in the Random Access configuration with the NVIDIA GeForce GTX TITAN X GPU. Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
IEEE Trans. Multim. | 2 |
| 2016 | SET response of a SEL protection switch for 130 and 250 nm CMOS technologiesabstractThis paper analyzes the single event transient (SET) response of a single event latchup (SEL) protection switch (SPS) designed in the 130 and 250 nm bulk CMOS technologies. The analysis has been conducted through the SPICE simulations, using the standard double exponential current source as the SET model. It has been confirmed that the 130 nm SPS cell is more susceptible to SETs than the 250 nm version, i.e. the 130 nm SPS cell has exhibited significantly lower critical charge. Based on the simulation results, an analytical model for estimating the critical charge in terms of the transistor size, number of load cells, and duration of the SET current pulse, has been derived. Use of the proposed critical charge model simplifies the analysis of the SPS cell's susceptibility to SETs for custom designs. Marko S. Andjelkovic, Aleksandar Ilic, Vladimir Petrovic, Miljana Nenadovic, Zoran Stamenkovic, Goran S. Ristic |
IOLTS | 2 |
| 2016 | Efficient HEVC decoder for heterogeneous CPU with GPU systemsabstractThe High Efficiency Video Coding (HEVC) standard provides higher compression efficiency than other video coding standards but at the cost of increased computational load, which makes it hard to achieve real-time encoding/decoding of high-resolution, high-quality video sequences. In this paper, we investigate how Graphics Processing Units (GPUs) can be employed to accelerate HEVC decoding. GPUs are known to provide massive processing capability for throughput computing kernels, but the HEVC entropy decoding kernel cannot be executed efficiently on GPUs. We therefore propose a complete HEVC decoding solution for heterogeneous CPU+GPU systems, in which the entropy decoder is executed on the CPU and the remaining kernels on the GPU. Furthermore, the decoder is pipelined such that the CPU and the GPU can decode different frames in parallel. The proposed CPU+GPU decoder achieves an average frame rate of 150 frames per second for Ultra HD 4K video sequences when four CPU cores are used with an NVIDIA GeForce Titan X GPU. Biao Wang 0001, Mauricio Alvarez-Mesa, Chi Ching Chi, Ben H. H. Juurlink, Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
MMSP | 6 |
| 2016 | A proof of the conjecture regarding the sum of domination number and average eccentricity
Zhibin Du, Aleksandar Ilic |
Discret. Appl. Math. | 2 |
| 2016 | On the extremal values of general degree-based graph entropies
Aleksandar Ilic |
Inf. Sci. | 1 |
| 2016 | A Framework for Application-Guided Task Management on Heterogeneous Embedded SystemsabstractIn this article, we propose a general framework for fine-grain application-aware task management in heterogeneous embedded platforms, which allows integration of different mechanisms for an efficient resource utilization, frequency scaling, and task migration. The proposed framework incorporates several components for accurate runtime monitoring by relying on the OS facilities and performance self-reporting for parallel and iterative applications. The framework efficiency is experimentally evaluated on a real hardware platform, where significant power and energy savings are attained for SPEC CPU2006 and PARSEC benchmarks, by guiding frequency scaling and intercluster migrations according to the runtime application behavior and predefined performance targets. Francisco Gaspar, Luís Taniça, Pedro Tomás, Aleksandar Ilic, Leonel Sousa |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | Adaptive Scheduling Framework for Real-Time Video Encoding on Heterogeneous SystemsabstractTo challenge real-time encoding of high-definition video sequences on heterogeneous desktop systems, a collaborative central processing units (CPU) + graphics processing unit (GPU) framework for interloop video encoding is proposed herein. The proposed framework considers the overall complexity of the collaborative interloop encoding as a unified optimization problem. Several functional blocks are integrated for simultaneous execution control, automatic data access management, performance characterization, and adaptive scheduling and load balancing. These blocks aim at fully exploiting the performance of heterogeneous devices, asymmetric bandwidth of communication links, and several levels of concurrency between computation and communication. To support a wide range of CPU and GPU architectures, a specific encoding library is developed with highly optimized algorithms for all interloop modules. The experimental results show that the proposed framework allows achieving a real-time encoding of full high-definition sequences in several CPU + GPU systems. It also delivers performance improvements of up to 61.2% over the state-of-the-art solution, while outperforming individual GPU and quad-core CPU executions by more than 2 and 5 times, respectively. Aleksandar Ilic, Svetislav Momcilovic, Nuno Roma, Leonel Sousa |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Towards GPU HEVC intra decoding: Seizing fine-grain parallelismabstractTo satisfy the growing demands on real-time video decoders for high frame resolutions, novel GPU parallel algorithms are proposed herein for fully compliant HEVC de-quantization, inverse transform and intra prediction. The proposed algorithms are designed to fully exploit and leverage the fine grain parallelism within these computationally demanding and highly data dependent modules. Moreover, the proposed approaches allow the efficient utilization of the GPU computational resources, while carefully managing the data accesses in the complex GPU memory hierarchy. The experimental results show that the real-time processing is achieved for all tested sequences and the most demanding QP, while delivering average fps of 118.6, 89.2 and 49.7 for Full HD, 2160p and Ultra HD 4K sequences, respectively. Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
ICME | 2 |
| 2015 | Multi-kernel Auto-Tuning on GPUs: Performance and Energy-Aware OptimizationabstractPrompted by their very high computational capabilities and memory bandwidth, Graphics Processing Units (GPUs) are already widely used to accelerate the execution of many scientific applications. However, programmers are still required to have a very detailed knowledge of the GPU internal architecture when tuning the kernels, in order to improve either performance or energy-efficiency. Moreover, different GPU devices have different characteristics, moving a kernel to a different GPU typically requires re-tuning the kernel execution, in order to efficiently exploit the underlying hardware. The procedure proposed in this work is based on real-time kernel profiling and GPU monitoring and it automatically tunes parameters from several concurrent kernels to maximize the performance or minimize the energy consumption. Experimental results on NVIDIA GPU devices with up to 4 concurrent kernels show that the proposed solution achieves near optimal configurations. Furthermore, significant energy savings can be achieved by using the proposed energy-efficiency auto-tuning procedure. João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás |
PDP | 2 |
| 2015 | On the variable common due date, minimal tardy jobs bicriteria two-machine flow shop problem with ordered machines
Aleksandar Ilic |
Theor. Comput. Sci. | 1 |
| 2015 | Finding Critical Regions and Region-Disjoint Paths in a NetworkabstractDue to their importance to society, communication networks should be built and operated to withstand failures. However, cost considerations make network providers less inclined to take robustness measures against failures that are unlikely to manifest, like several failures coinciding simultaneously in different geographic regions of their network. Considering networks embedded in a two-dimensional plane, we study the problem of finding a critical region-a part of the network that can be enclosed by a given elementary figure of predetermined size-whose destruction would lead to the highest network disruption. We determine that only a polynomial, in the input, number of nontrivial positions for such a figure needs to be considered and propose a corresponding polynomial-time algorithm. In addition, we consider region-aware network augmentation to decrease the impact of a regional failure. We subsequently address the region-disjoint paths problem, which asks for two paths with minimum total weight between a source (s) and a destination (d) that cannot both be cut by a single regional failure of diameter D (unless that failure includes s or d). We prove that deciding whether region-disjoint paths exist is NP-hard and propose a heuristic region-disjoint paths algorithm. Stojan Trajanovski, Fernando A. Kuipers, Aleksandar Ilic, Jon Crowcroft, Piet Van Mieghem |
IEEE/ACM Trans. Netw. | 3 |
| 2014 | Collaborative inter-prediction on CPU+GPU systemsabstractIn this paper we propose an efficient method for collaborative H.264/AVC inter-prediction in heterogeneous CPU+GPU systems. In order to minimize the overall encoding time, the proposed method provides stable and balanced load distribution of the most computationally demanding video encoding modules, by relying on accurate and dynamically built functional performance models. In an extensive RD analysis, an efficient temporary dependent prediction of the search area center is proposed, which allows dependency-aware workload partitioning and efficient GPU parallelization, while preserving high compression efficiency. The proposed method also introduces efficient communication-aware techniques, which maximize data reusing, and decrease the overhead of expensive data transfers in collaborative video encoding. The experimental results show that the proposed method is able of achieving real-time video encoding for very demanding video coding parameters, i.e. full HD video format, 64×64 pixels search area and the exhaustive motion estimation. Svetislav Momcilovic, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
ICIP | 2 |
| 2014 | FEVES: Framework for Efficient Parallel Video Encoding on Heterogeneous SystemsabstractLead by high performance computing potential of modern heterogeneous desktop systems and predominance of video content in general applications, we propose herein an autonomous unified video encoding framework for hybrid multi-core CPU and multi-GPU platforms. To fully exploit the capabilities of these platforms, the proposed framework integrates simultaneous execution control, automatic data access management, and adaptive scheduling and load balancing strategies to deal with the overall complexity of the video encoding procedure. These strategies consider the collaborative inter-loop encoding as a unified optimization problem to efficiently exploit several levels of concurrency between computation and communication. To support a wide range of CPU and GPU architectures, a specific encoding library is developed with highly optimized algorithms for all inter-loop modules. The obtained experimental results show that the proposed framework allows achieving a real-time encoding of full high-definition sequences in the state-of-the-art CPU+GPU systems, by outperforming individual GPU and quad-core CPU executions for more than 2 and 5 times, respectively. Aleksandar Ilic, Svetislav Momcilovic, Nuno Roma, Leonel Sousa |
ICPP | 1 |
| 2014 | Performance-Aware Task Management and Frequency Scaling in Embedded SystemsabstractDue to the dissemination of smartphones and tablets, a constant complexity growth can be observed for both embedded systems and mobile applications. However, this results in an increase in energy consumption. To guarantee longer battery life cycles, it is fundamental to develop system level strategies that allow guaranteeing the applications' required quality of service by managing the available system resources. In this paper a new task management framework is proposed that controls, in real-time, the execution of multi-threaded applications in order to meet their performance targets. For this, we amend the Linux CFS scheduler decisions to efficiently control the shared resource utilization of parallel applications. The proposed framework relies on runtime performance modelling of both the underlying architecture and the running applications to scale the system resource allocation and frequency. As a result, efficient application execution is achieved not only in terms of performance, but also in energy consumption. Experimental results show that the proposed approach satisfies the applications required performance level by decreasing the relative performance error from 2.801 to 0.168, while achieving 49 % energy savings. Francisco Gaspar, Aleksandar Ilic, Pedro Tomás, Leonel Sousa |
SBAC-PAD | 2 |
| 2014 | Dynamic Load Balancing for Real-Time Video Encoding on Heterogeneous CPU+GPU SystemsabstractThe high computational demands and overall encoding complexity make the processing of high definition video sequences hard to be achieved in real-time. In this manuscript, we target an efficient parallelization and RD performance analysis of H.264/AVC inter-loop modules and their collaborative execution in hybrid multi-core CPU and multi-GPU systems. The proposed dynamic load balancing algorithm allows efficient and concurrent video encoding across several heterogeneous devices by relying on realistic run-time performance modeling and module-device execution affinities when distributing the computations. Due to an online adjustment of load balancing decisions, this approach is also self-adaptable to different execution scenarios. Experimental results show the proposed algorithm's ability to achieve real-time encoding for different resolutions of high-definition sequences in various heterogeneous platforms. Speed-up values of up to 2.6 were obtained when compared to the video inter-loop encoding on a single GPU device, and up to 8.5 when compared to a highly optimized multi-core CPU execution. Moreover, the proposed algorithm also provides an automatic tuning of the encoding parameters, in order to meet strict encoding constraints. Svetislav Momcilovic, Aleksandar Ilic, Nuno Roma, Leonel Sousa |
IEEE Trans. Multim. | 2 |
| 2013 | Critical regions and region-disjoint paths in a network
Stojan Trajanovski, Fernando A. Kuipers, Piet Van Mieghem, Aleksandar Ilic, Jon Crowcroft |
Networking | 4 |
| 2013 | Efficient algorithm for the vertex connectivity of trapezoid graphs
Aleksandar Ilic |
Inf. Process. Lett. | 1 |
| 2012 | Hierarchical Partitioning Algorithm for Scientific Computing on Highly Heterogeneous CPU + GPU Clusters
David Clarke, Aleksandar Ilic, Alexey L. Lastovetsky, Leonel Sousa |
Euro-Par | 2 |
| 2012 | Simultaneous Multi-Level Divisible Load Balancing for Heterogeneous Desktop SystemsabstractIn this paper, we propose an algorithm for efficient divisible load balancing across all processing devices available in a heterogeneous desktop system. The proposed algorithm allows to achieve simultaneous load balancing at different execution levels, namely between execution subdomains defined with several processing devices, and between devices in each subdomain. Moreover, the algorithm builds partial performance models for each execution subdomain, using the minimal set of approximation points determined during the algorithm run, while converging towards the optimal multi level load distributions. The proposed approach was experimentally evaluated in a real desktop system with a quad core CPU and two GPUs, for matrix multiplication. Experimental results show the ability of the algorithm to provide significant performance improvements with very low scheduling overhead when compared to similar scheduling approaches. Aleksandar Ilic, Leonel Sousa |
ISPA | 1 |
| 2012 | On Realistic Divisible Load Scheduling in Highly Heterogeneous Distributed SystemsabstractThis paper investigates the problem of scheduling discretely divisible applications in highly heterogeneous distributed platforms which deploy modern desktop systems with limited memory as computing nodes. We propose an algorithm for hierarchical load balancing at both inter- and intra-node platform levels which relies on realistic performance models of computation and communication resources. An iterative procedure, based on the proposed algorithm, is also presented for building accurate performance models during the application run-time. The presented approach was evaluated for a 2D FFT batch application executed on a distributed system with four CPU+GPU nodes. The experimental results show the advantages of using the proposed approach by outperforming the "optimal" implementation by at least 4 times on GPU devices. Aleksandar Ilic, Leonel Sousa |
PDP | 1 |
| 2012 | On reformulated Zagreb indices
Aleksandar Ilic, Bo Zhou 0007 |
Discret. Appl. Math. | 1 |
| 2012 | Ballot matrix as Catalan matrix power and related identities
Stefan Stanimirovic, Predrag S. Stanimirovic, Aleksandar Ilic |
Discret. Appl. Math. | 3 |
| 2012 | The index of a binary word
Aleksandar Ilic, Sandi Klavzar, Yoomi Rho |
Theor. Comput. Sci. | 1 |
| 2011 | Degree distance of unicyclic and bicyclic graphs
Aleksandar Ilic, Dragan Stevanovic, Lihua Feng, Guihai Yu, Peter Dankelmann |
Discret. Appl. Math. | 1 |
| 2010 | Distance spectral radius of trees with given matching number
Aleksandar Ilic |
Discret. Appl. Math. | 1 |
| 2008 | Distributed Web-based Platform for Computer Architecture SimulationabstractComputer architecture simulation and modeling require a huge amount of time and resources, not only for the simulation itself but also regarding the configuration and submission procedures. A quite common simulation toolset (SimpleScalar) has been used to model a variety of platforms ranging from simple unpipelined processors to detailed dynamically scheduled microarchitectures with multiple-level memory hierarchies. In this paper we propose a platform for automatically executing a massive number of simulations in parallel, by exploiting a distributed computing approach. We developed a Web-based simulation system consisting in a front-end user interface and a back-end part supported on a grid system. The front-end is responsible for configuring the simulation and parsing the results, while the back-end distributes the workload by using Condor scheduler. Experimental results show that it is very easy to use the system, even when dealing with a huge number of simulations, and also it provides results in a very suitable format. Moreover, it has been concluded that a significant speedup can be achieved, by exploiting parallelism at the benchmark levels or also by sampling each benchmark with the SimPoint tool. Aleksandar Ilic, Frederico Pratas, Leonel Sousa |
ISPDC | 1 |