Claude Tadonki

dblp:75/1189 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-1194-6400ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Leveraging cutting-edge high performance computing for large-scale applications
Claude Tadonki, Gabriele Mencagli, Leonel Sousa
Future Gener. Comput. Syst.1
2025 A Framework for Analytical Performance and Energy Prediction of DL Training on GPUs
abstract
The rapid scaling of deep learning (DL) models raises the need for accurate and understandable performance/energy prediction tools to support efficient resource management and sustainable AI development. Existing modeling approaches often lack both sufficient granularity to capture nuanced hardware-software interactions and suitable flexibility to adapt to diverse modern architectures. This paper introduces an analytical framework for time/energy prediction of DL training workloads on GPU. Our framework integrates detailed workload characterization that includes FLOPs, memory access, kernel activities, and novel structural features to derive an architecture-aware efficiency model, which considers a saturation-based function to capture dimensional scaling effects on hardware utilization. We propose an iterative refinement methodology, which incorporates model-specific scalars to address particular architectures like ALBERT and precision-specific calibrations for BF16 operations. Our benchmark with six advanced DL models (including CNNs, BERT-style Transformers, and LLMs like TinyLlama) on NVIDIA A100 GPUs under various configurations (1/4 GPUs, FP32/TF32/mixed BF16) shows that our approach achieves a high predictive accuracy, with an overall relative error of 4.14% (3.05% for time, 5.78% for power). The framework is intended to provide valuable insights for HPC-AI co-design, energy-aware scheduling, and performance optimization.
Roblex Nana Tchakoute, Claude Tadonki, Petr Dokládal, Youssef Mesri
SBAC-PAD2
2023 Dynamic data replication and placement strategy in geographically distributed data centers
abstract
Abstract With the evolution of geographically distributed data centers in the Cloud Computing landscape along with the amount of data being processed in these data centers, which is growing at an exponential rate, processing massive data applications become an important topic. Since a given task may require many datasets for its execution and the datasets are spread over several different data centers, finding an efficient way to manage the datasets storage across nodes of a Cloud system is a difficult problem. In fact, the execution time of a task might be influenced by the cost of data transfers, which mainly depends on two criterias. The first one is the initial placement of the input datasets during the build‐time phase, while the second is the replication of the datasets during the runtime phase. The replication is explicitly considered when datasets are being migrated over the data centers in order to make them locally available wherever needed. Data placement and data replication are important challenges in Cloud Computing. Nevertheless, many studies focus on data placement or data replication exclusively. In this paper, a combination of a data placement strategy followed by a dynamic data replication management strategy is proposed, with the purpose of reducing the associated cost of all data transfers between the (distant) data centers. Our proposed data placement approach considers the main characteristics of a data center such asstorage capacityandread/write speedsto efficiently store the datasets, while our dynamic data replication management approach considers three parameters: thenumber of replicasin the system, thedependency between datasetsand tasks and thestorage capacityof data centers. The decision of when and whether to keep or to delete replicas is determined by the fulfillment of those three parameters. Our approach estimates the total execution time of the tasks as well as the monetary cost, considering the data transfers activity. Our experiments are conducted using Cloudsim simulator. The obtained results show that our proposed strategies produce an efficient data management by reducing the overheads of the data transfers, compared to both a data placement without replication (by 76%) and the selected data replication approach from Kouidri et al. (by 52%), and by improving the financial cost.
Laila Bouhouch, Mostapha Zbakh, Claude Tadonki
Concurr. Comput. Pract. Exp.3
2023 Optimizing computational costs of Spark for SARS-CoV-2 sequences comparisons on a commercial cloud
abstract
Summary Cloud computing is currently one of the prime choices in the computing infrastructure landscape. In addition to advantages such as the pay‐per‐use bill model and resource elasticity, there are technical benefits regarding heterogeneity and large‐scale configuration. Alongside the classical need for performance, for example, time, space, and energy, there is an interest in the financial cost that might come from budget constraints. Based on scalability considerations and the pricing model of traditional public clouds, a reasonable optimization strategy output could be the most suitable configuration of virtual machines to run a specific workload. From the perspective of runtime and monetary cost optimizations, we provide the adaptation of a Hadoop applications execution cost model extracted from the literature aiming at Spark applications modeled with the MapReduce paradigm. We evaluate our optimizer model executing an improved version of the Diff Sequences Spark application to perform SARS‐CoV‐2 coronavirus pairwise sequence comparisons using the AWS EC2's virtual machine instances. The experimental results with our model outperformed 80% of the random resource selection scenarios. By only employing spot worker nodes exposed to revocation scenarios rather than on‐demand workers, we obtained an average monetary cost reduction of 35.66% with a slight runtime increase of 3.36%.
Alan L. Nunes, Alba Cristina Magalhaes Alves de Melo, Claude Tadonki, Cristina Boeres, Daniel de Oliveira 0001, Lúcia M. A. Drummond
Concurr. Comput. Pract. Exp.3
2020 Performance Analysis and Optimization of the Vector-Kronecker Product Multiplication
abstract
The Kronecker product, also called tensor product, is a fundamental matrix algebra operation, used to model complex systems using structured descriptions. This operation needs to be computed efficiently, since it is a critical kernel for iterative algorithms. In this work, we focus on the vector-kronecker product operation, where we present an in-depth performance analysis of a sequential and a parallel algorithm previously proposed. Based on this analysis, we proposed three optimizations: changing the memory access pattern, reducing load imbalance and manually vectorizing some portions of the code with Intel SSE4.2 intrinsics. The obtained results show better cache usage and load balance, thus improving the performance, especially for larger matrices.
Alexandre Azevedo, Cristiana Bentes, Maria Clicia Stelling de Castro, Claude Tadonki
SBAC-PAD4
2018 Meta-programming for cross-domain tensor optimizations
abstract
Many modern application domains crucially rely on tensor operations. The optimization of programs that operate on tensors poses difficulties that are not adequately addressed by existing languages and tools. Frameworks such as TensorFlow offer good abstractions for tensor operations, but target a specific domain, i.e. machine learning, and their optimization strategies cannot easily be adjusted to other domains. General-purpose optimization tools such as Pluto and existing meta-languages offer more flexibility in applying optimizations but lack abstractions for tensors. This work closes the gap between domain-specific tensor languages and general-purpose optimization tools by proposing the Tensor optimizations Meta-Language (TeML). TeML offers high-level abstractions for both tensor operations and loop transformations, and enables flexible composition of transformations into effective optimization paths. This compositionality is built into TeML's design, as our formal language specification will reveal. We also show that TeML can express tensor computations as comfortably as TensorFlow and that it can reproduce Pluto's optimization paths. Thus, optimized programs generated by TeML execute at least as fast as the corresponding Pluto programs. In addition, TeML enables optimization paths that often allow outperforming Pluto.
Adilla Susungi, Norman A. Rink, Albert Cohen 0001, Jerónimo Castrillón, Claude Tadonki
GPCE5
2018 Evaluation of an OPENMP Parallelization of Lucas-Kanade on a NUMA-Manycore
abstract
Lucas-Kanade algorithm is a well-known optical flow estimator widely used in image processing for motion detection and object tracking. As a typical image processing algorithm, the procedure is a series of convolution masks followed by 2×2 linear systems for the optical flow vectors. Since we are dealing with a stencil computation for each stage of the algorithm, the overhead from memory accesses is expected to stand as a serious scalability bottleneck, especially on a NUMA manycore configuration. The objective of this study is therefore to investigate an openMP parallelization of Lucas-kanade algorithm on a NUMA manycore, including the performance impact of NUMA-aware settings at runtime. Experimental results on a dual-socket INTEL Broadwell-EIEP is provided together with the corresponding technical discussions.
Olfa Haggui, Claude Tadonki, Fatma Sayadi, Bouraoui Ouni
SBAC-PAD2
2018 Performance comparison between Hadoop and Spark frameworks using HiBench benchmarks
abstract
Summary Big Data has become one of the major areas of research for cloud service providers due to a large amount of data produced every day and the inefficiency of traditional algorithms and technologies to handle these large amounts of data. Big Data with its characteristics such as volume, variety, and veracity (3V) requires efficient technologies to process in real time. To solve this problem and to process and analyze this vast amount of data, there are many powerful tools like Hadoop and Spark, which are mainly used in the context of Big Data. They work following the principles of parallel computing. The challenge is to specify which Big Data's tool is better depending on the processing context. In this paper, we present and discuss a performance comparison between two popular Big Data frameworks deployed on virtual machines. Hadoop MapReduce and Apache Spark are used to efficiently process a vast amount of data in parallel and distributed mode on large clusters, and both of them suit for Big Data processing. We also present the execution results of Apache Hadoop in Amazon EC2, a major cloud computing environment. To compare the performance of these two frameworks, we use HiBench benchmark suite, which is an experimental approach for measuring the effectiveness of any computer system. The comparison is made based on three criteria: execution time, throughput, and speedup. We test Wordcount workload with different data sizes for more accurate results. Our experimental results show that the performance of these frameworks varies significantly based on the use case implementation. Furthermore, from our results we draw the conclusion that Spark is more efficient than Hadoop to deal with a large amount of data in major cases. However, Spark requires higher memory allocation, since it loads the data to be processed into memory and keeps them in caches for a while, just like standard databases. So the choice depends on performance level and memory constraints.
Yassir Samadi, Mostapha Zbakh, Claude Tadonki
Concurr. Comput. Pract. Exp.3
2018 Harris corner detection on a NUMA manycore
Olfa Haggui, Claude Tadonki, Lionel Lacassagne, Fatma Sayadi, Bouraoui Ouni
Future Gener. Comput. Syst.2
2017 Towards compositional and generative tensor optimizations
abstract
Many numerical algorithms are naturally expressed as operations on tensors (i.e. multi-dimensional arrays). Hence, tensor expressions occur in a wide range of application domains, e.g. quantum chemistry and physics; big data analysis and machine learning; and computational fluid dynamics. Each domain, typically, has developed its own strategies for efficiently generating optimized code, supported by tools such as domain-specific languages, compilers, and libraries. However, strategies and tools are rarely portable between domains, and generic solutions typically act as ''black boxes'' that offer little control over code generation and optimization. As a consequence, there are application domains without adequate support for easily generating optimized code, e.g. computational fluid dynamics. In this paper we propose a generic and easily extensible intermediate language for expressing tensor computations and code transformations in a modular and generative fashion. Beyond being an intermediate language, our solution also offers meta-programming capabilities for experts in code optimization. While applications from the domain of computational fluid dynamics serve to illustrate our proposed solution, we believe that our general approach can help unify research in tensor optimizations and make solutions more portable between domains.
Adilla Susungi, Norman A. Rink, Jerónimo Castrillón, Immo Huismann, Albert Cohen 0001, Claude Tadonki, Jörg Stiller, Jochen Fröhlich
GPCE6
2016 Power-aware server consolidation for federated clouds
abstract
Summary Cloud computing has evolved to provide computing resources on‐demand through a virtualized infrastructure, letting applications, computing power, data storage, and network resources to be provisioned and managed over private networks or over the Internet. Cloud services normally run on large data centers and demand a huge amount of electricity. Consequently, the electricity cost represents one of the major concerns of data centers, because it is sometimes nonlinear with the capacity of the data centers, and it is also associated with a high amount of carbon emission (CO2). However, energy‐saving schemes that result in too much degradation of the system performance or in violations of service‐level agreement (SLA) parameters would eventually cause the users to move to another cloud provider. Thus, there is a need to reach a balance between energy savings and the costs incurred by these savings in the execution of the applications. Therefore, in this paper, we propose and evaluate a power and SLA‐aware application consolidation solution for cloud federations. It comprises a multi‐agent system for server consolidation, taking into account SLA, power consumption, and carbon footprint. Different for similar solutions available in the literature, in our solution, when a cloud is overloaded, its data center needs to negotiate with other data centers before migrating the workload to another cloud. Simulation results show that our approach can reduce up to 46% of the power consumption while trying to meet performance requirements. Furthermore, we show that federated clouds can provide an adequate solution to deal with power consumption in the clouds. Copyright © 2016 John Wiley & Sons, Ltd.
Alessandro Ferreira Leite, Azzedine Boukerche, Alba Cristina Magalhaes Alves de Melo, Christine Eisenbeis, Claude Tadonki, Célia Ghedini Ralha
Concurr. Comput. Pract. Exp.5
2015 Automating Resource Selection and Configuration in Inter-clouds through a Software Product Line Method
abstract
Nowadays, cloud users face three important problems: (a) choosing one or more appropriate cloud provider(s) to run their application(s), (b) selecting appropriate cloud resources, which implies having enough information about the available resources, including their characteristics and constraints, and (c) configuring the cloud resources. These problems are mostly due to the wide range of resources. These resources usually have distinct dependencies, and they are offered at various clouds' layers. In this complex scenario, the users often have to handle cloud resources and their dependencies manually. This is an error-prone and time-consuming activity, even for skilled cloud users and system administrators. In this context, this paper proposes a software product line engineering (SPLE) method and a tool to deal with these issues. Our SPL-based engineering method enables a declarative and goal-oriented strategy. Furthermore, it allows resource selection and configuration in inter-cloud environments. In our proposal, the cloud users specify their applications and requirements, and our tool automatically selects and configures a suitable computing environment, taking into account temporal and functional dependencies. Experimental results on Amazon EC2 and Google Compute Engine (GCE) show that our approach enables unskilled users to have access to advanced inter-cloud computing configurations, without being concerned with the characteristics of each cloud.
Alessandro Ferreira Leite, Vander Alves, Genaína Nunes Rodrigues, Claude Tadonki, Christine Eisenbeis, Alba Cristina Magalhaes Alves de Melo
CLOUD4
2012 3D shape Retrieval Using Bag-of-Feature Method Basing on Local Codebooks
Elwardani Dadi, El Mostafa Daoudi, Claude Tadonki
ICISP3
2012 Accelerator-Based implementation of the Harris Algorithm
Claude Tadonki, Lionel Lacassagne, Elwardani Dadi, El Mostafa Daoudi
ICISP1
2009 Algorithmic Skeletons within an Embedded Domain Specific Language for the CELL Processor
abstract
Efficiently using the hardware capabilities of the Cell processor, a heterogeneous chip multiprocessor that uses several levels of parallelism to deliver high performance, and being able to reuse legacy code are real challenges for application developers. We propose to use Generative Programming and more precisely template meta-programming to design an domain specific embedded language using algorithmic skeletons to generate applications based on a high-level mapping description. The method is easy to use by developers and delivers performance close to the performance of optimized hand-written code, as shown on various benchmarks ranging from simple BLAS kernels to image processing applications.
Tarik Saidani, Joël Falcou, Claude Tadonki, Lionel Lacassagne, Daniel Etiemble
PACT3
2003 Combinatorial Techniques for Memory Power State Scheduling in Energy-Constrained Systems
Claude Tadonki, Mitali Singh, José D. P. Rolim, Viktor Prasanna 0001
WAOA1
1999 The Algebraic Path Problem Revisited
Sanjay V. Rajopadhye, Claude Tadonki, Tanguy Risset
Euro-Par2