EDBT 2026 Demo / reviewers in the wild / expert
Alessandro Cilardo
dblp:18/2951
· DBLP profile ↗
56ranked-venue papers
34as first author
10since 2021 · last 2026
0000-0002-1685-8736ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 28 first-author · 10 since 2021Software engineering, systems software and programming languages · 17 · 10 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Formally Verified Secure Caching Mechanism on TrustZone-enabled MicrocontrollersabstractTrusted Execution Environments (TEEs) on resource-constrained microcontrollers are an emerging area of interest, yet they present unique security challenges, particularly in managing encrypted code execution through limited secure memory. This paper presents a formal verification approach for Umbra, a TEE framework for ARM TrustZone-M, currently under development, that implements secure caching mechanisms to execute encrypted enclaves from flash memory. We employ model checking techniques to formally analyze critical security properties, including data isolation between secure and non-secure worlds, integrity of the Enclave Flash Block Cache (EFBC), and resilience against identified threats such as Direct Memory Access (DMA) handover attacks and timing-based side channels. Our threat model considers privileged attackers in the non-secure world and compromised host operating systems, analyzing vulnerabilities in DMA reconfiguration windows and context switch dependencies. Through formal modeling, we identify replay and timing side-channel attacks; by introducing countermeasures, these guarantees are restored in the model. Salvatore Bramante, Matteo Busi 0001, Alessandro Cilardo, Riccardo Focardi, Flaminia L. Luccio, Stefano Mercogliano |
DATE | 3 |
| 2026 | A hardware/software architecture for multi-threaded offloading of erasure codes in distributed file systemsabstractBig Data analytics and cloud computing impose an ever-growing demand for data-center providers in terms of computational requirements, latency, and storage. Distributed file systems offer the strategic advantage of scaling-out computing and storage resources, hence allowing for notable speed-ups with massively parallel and distributed computing paradigms. On the other hand, such distributed clusters are constantly challenged with storage failures. Data replication is often deployed to ensure fault tolerance and business continuity, typically in a 3x configuration. This results in expensive 200 % overheads in storage space, write propagation, and energy costs. Erasure codes offer an alternative approach for fault tolerance by allowing reconstruction of erased data chunks, while reducing storage overhead down to 30 %. However, a considerable share of CPU cycles and energy is spent computing such codes, effectively reducing the cluster’s efficiency and starving other user and system tasks. Offloading on a custom accelerator is a non-trivial issue, due to the highly multi-threaded nature of such tasks and the lack of robust multi-threading support in conventional accelerator runtimes. In this work, we present a heterogeneous hardware/software architectural design for large-scale and multi-threaded acceleration of distributed erasure codes on PCIe accelerators, and a new abstraction and integration model for distributed accelerators in fault-tolerant storage systems. We enable safe and seamless deployment of multi-threaded SYCL-based IP cores through a hardware thread proxying layer providing software thread-isolation, and integration with cluster-level middlewares. In addition, our design allows for heterogeneous cluster configurations, with full compatibility and transparent integration of heterogeneously-accelerated and CPU-only nodes. We systematically evaluate the individual layers of our architecture and validate design’s integration in a container-based HDFS cluster, comparing performance against the state-of-the-art AVX-512-accelerated ISA-L library and other SYCL substrates, such as GPUs and single-threaded FPGAs. Vincenzo Maisto, Alessandro Cilardo, Emilio Billi, Chuck Fader |
Future Gener. Comput. Syst. | 2 |
| 2026 | Distilling knowledge for low-energy AIoTabstractThe Artificial Intelligence of Things (AIoT) empowers IoT devices to leverage the advantages of AI near data-sources, reducing data movement, latency, and mitigating privacy issues. However, AI workloads are notoriously energy-intensive, posing significant challenges for energy-constrained IoT devices. Since such devices are often deployed in thousands of instances, even minor inefficiencies can significantly increase carbon emissions and energy consumptions. Model compression techniques have been employed to enable AI inference in resource-constrained environments. For example, Knowledge Distillation (KD) is an elaborate approach targeting low-footprint and high-accuracy models, although introducing further complexity during training due to inefficient grid searches of additional hyperparameters. The emerging wave of AIoT, however, calls for prioritizing energy-awareness both in inference and training. To address this shortcoming, this work proposes a three-stage design workflow for low-energy AIoT applications, driven primarily by an input energy budget characterizing the target IoT scenario. Given a specific CNN architecture and IoT platform, our workflow identifies the most effective student under the imposed energy constrained and derives an efficient configuration of the KD hyperparameters that maximizes student accuracy, while avoiding inefficient and expensive grid-search. Hence, this approach enable energy-efficient CNN inference while substantially reducing overall training costs. We validate our workflow with a systematic experimental campaign using ResNets and DenseNets on CIFAR-10, CIFAR-100,and Tiny Imagenet datasets, on an AMD Xilinx Zynq Ultrascale+ ZCU102 MPSoC. Our proposal maintains high accuracy while lowering energy consumption by up to 80%, highlighting the potential of our flow for real-world AIoT applications. Franca Rocco di Torrepadula, Vincenzo Maisto, Alessandro Cilardo, Nicola Mazzocca |
J. Syst. Archit. | 3 |
| 2026 | The Simply-V Framework: An Extensible RISC-V Reconfigurable Soft-SoC for Open Research and Fast PrototypingabstractThe recent rise of open hardware, mainly driven by the momentum of the RISC-V ecosystem, has sparked significant innovation in the development of open-source CPUs and SoCs. This movement has enabled broad exploration across academia and industry, fostering collaboration and reuse. However, the diversity and openness that empower this space also introduce challenges: academic projects often fall short of industry-grade robustness, and meaningful comparison across hardware platforms remains difficult due to ad hoc infrastructures, lack of standardization, and simulation limitations. To ease the work of researchers some key challenges must be faced in open hardware development: platforms’ reconfigurability, ease of integration of third-party IPs, and support for technological heterogeneity. A core problem lies in validating and comparing CPUs and SoC components across varying protocols, toolchains, and design languages, especially in real hardware settings. To address these issues, we present Simply-V, a flexible FPGA-based soft-SoC platform designed for rapid prototyping and open hardware research. Simply-V enables plug-and-play support for multiple CPUs, IPs and accelerators, offers structured configurability across embedded and high-performance profiles, and supports the integration of both RTL and HLS-based components. Capabilities such as a high-level configuration flow, frequency scaling, and cross-device portability make our platform a powerful tool to simplify open hardware research. We demonstrate the SoC generator’s capabilities through multi-task FreeRTOS examples, platform-fair CPU benchmarking and the iterative development of HLS-designed convolutional accelerators. Moreover, we validate multi-accelerator and multi-CPU scalability and compare with the state-of-the-art SoC generators. Our platform showcases simplified fast prototyping, configurability, scalability and heterogeneous IP support on real hardware. Simply-V is openly available at https://github.com/HiSA-Team/Simply-V . Vincenzo Maisto, Stefano Mercogliano, Manuel Maddaluno, Alessandro Cilardo |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | Umbra: An Efficient Framework for Trusted Execution on Modern TrustZone-Enabled MicrocontrollersabstractThe rise of microcontrollers in critical systems demands robust security measures beyond traditional methods like Memory Protection Units. ARM's TrustZone-M offers enhanced protection for secure applications, yet its potential for deploying Trusted Execution Environments often remains untapped, leaving room for innovation in managing security on resource-constrained devices. This paper presents Umbra, a Rust-based framework that isolates mutually distrustful applications and integrates with untrusted embedded OSes. Leveraging modern security hardware, Umbra features an efficient secure caching mechanism that encrypts all code exposed to attackers, decrypting and validating only necessary blocks during execution, achieving practical Trusted Execution Environments on modern microcontrollers. Stefano Mercogliano, Alessandro Cilardo |
DATE | 2 |
| 2024 | Lightweight and Predictable Memory Virtualization on Medium-Size MicrocontrollersabstractNowadays industry research is heading towards the consolidation of multiple real-time applications and execution environments on single microcontrollers, with the aim of optimizing area, power, and cost while keeping an eye on protection and flexibility. To this end, virtualization seems an attractive solution, but it must be redesigned according to the specific requirements of microcontroller tasks, different than traditional application processor workloads. This paper examines two possible hardware-based models to support virtual machines on medium-size microcontrollers providing an extensive and reproducible analysis over a RISC-V processor. Stefano Mercogliano, Daniele Ottaviano, Alessandro Cilardo, Marcello Cinque |
DATE | 3 |
| 2024 | An Approach to the Systematic Characterization of Multitask Accelerated CNN Inference in Edge MPSoCsabstractDeep Learning is ubiquitous today and is increasingly moving from the cloud down to the edge of networked infrastructures, where it enables embedded applications to perform complex inference tasks close to the data sources, reducing long-distance data movement and alleviating the need for a powerful cloud infrastructure. Edge-class multi-processor system on chip (MPSoC) devices featuring an on-chip FPGA fabric offer key advantages for Deep Learning inference tasks, especially for complex applications where multiple models may be run concurrently in the same platform. In this work, we propose an approach and a practical framework for the systematic characterization of multithreaded Deep Learning inference on edge FPGA MPSoCs. We instantiate the framework into a real-world MPSoC platform, targeting Xilinx Vitis-AI as a representative example of a commercial Deep Learning acceleration toolkit for edge environments. We design a comprehensive experimental campaign and apply it to the platform for several convolutional neural networks, each trained on three different datasets. We show that our approach can be used for both hardware- and software-level analysis of a target system. Among other findings, the analysis revealed a suboptimal behavior of the underlying toolkit runtime, involving the utilization of the accelerator cores and the uneven software latency of the support library, influenced by the shapes of the input tensors. Alessandro Cilardo, Vincenzo Maisto, Nicola Mazzocca, Franca Rocco di Torrepadula |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | A Pluggable Vector Unit for RISC-V Vector ExtensionabstractVector extensions have become increasingly important for accelerating data-parallel applications in areas like multimedia, data-streaming, and Machine Learning. This interactive presentation in-troduces a microarchitectural design of a vector unit compliant with the RISC- V vector extension v1.0. While we targeted a specific core for demonstration, CVA6, our architecture is designed so as to ensure extensibility, maintainability, and re-usability in other cores. Furthermore, as a distinctive feature, we support speculative execution and precise vector traps. The paper provides an overview of the main motivation, design choices, and implementation details, followed by a qualitative and quantitative discussion of the results collected from the synthesis of the extended CVA6 RISC-V core. Vincenzo Maisto, Alessandro Cilardo |
DATE | 2 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 6 |
| 2021 | FPGA-based real-time monitoring support for CAN applicationsabstractThis technical contribution deals with monitoring support for CAN, a popular protocol in automotive and robotics applications with various levels of criticality, therefore requiring strict reliability and performance guarantees. While software implementations for CAN-based monitoring applications are very flexible, they may face prohibitive overheads in terms of latency and responsiveness. We present a customizable hardware-based CAN filter designed to enable real-time monitoring and anomaly detection, which can be employed in critical systems with stringent response time requirements. As shown in the paper, an advanced CAN monitor relying on the customizable FPGA-based filter can bridge the limitations of software solutions by drastically reducing latency –around 10X compared to software– showing that the adoption of FPGA technologies in a critical industrial environment can bring key benefits in terms of real-time features and flexibility. Alessandro Cilardo, Stefano Mercogliano |
DSD | 1 |
| 2019 | Lightweight hardware support for selective coherence in heterogeneous manycore acceleratorsabstractShared memory coherence is a key feature in many-core accelerators, ensuring programmability and application portability. Most established solutions for coherence in homogeneous systems cannot be simply reused because of the special requirements of accelerator architectures. This paper introduces a low-overhead hardware coherence system for heterogeneous accelerators, with customizable granularity and noncoherent region support. The coherence system has been demonstrated in operation in a full manycore accelerator, exhibiting significant improvements in terms of network load, execution time, and power consumption. Alessandro Cilardo, Mirko Gagliardi, Vincenzo Scotti 0002 |
DATE | 1 |
| 2019 | Challenges in Deeply Heterogeneous High Performance SystemsabstractRECIPE (REliable power and time-ConstraInts-aware Predictive management of heterogeneous Exascale systems) is a recently started project funded within the H2020 FETHPC programme, which is expressly targeted at exploring new High-Performance Computing (HPC) technologies. RECIPE aims at introducing a hierarchical runtime resource management infrastructure to optimize energy efficiency and minimize the occurrence of thermal hotspots, while enforcing the time constraints imposed by the applications and ensuring reliability for both time-critical and throughput-oriented computation that run on deeply heterogeneous accelerator-based systems. This paper presents a detailed overview of RECIPE, identifying the fundamental challenges as well as the key innovations addressed by the project, which span run-time management, heterogeneous computing architectures, HPC memory/interconnection infrastructures, thermal modelling, reliability, programming models, and timing analysis. For each of these areas, the paper describes the relevant state of the art as well as the specific actions that the project will take to effectively address the identified technological challenges. Giovanni Agosta, William Fornaciari, David Atienza 0001, Ramon Canal, Alessandro Cilardo, José Flich, Carles Hernández 0001, Michal Kulczewski, Giuseppe Massari, Rafael Tornero, Marina Zapater |
DSD | 5 |
| 2018 | Understanding turn models for adaptive routing: The modular approachabstractRouting algorithms were extensively studied first in multi-computer systems, then in multi- and many-core architectures. Among the commonly used routing techniques, the turn model seems the most promising solution when targeting adaptiveness. Based on the turn model, several alternative approaches with different turn prohibition schemes were proposed. This paper gives a new theoretical background for designing deadlock-free partially adaptive logic-based distributed routing algorithms that are based on the turn model. Two properties are presented, including a necessary and sufficient condition to prove that a routing algorithm is deadlock-free as long as turn restrictions follow a modular distribution. Existing approaches can be considered a subset of the solution space identified by this work. Finally, we propose a novel routing algorithm exhibiting encouraging performance improvements over state-of-the-art approaches. Edoardo Fusella, Alessandro Cilardo |
DATE | 2 |
| 2018 | Reducing Power Consumption of Lasers in Photonic NoCs through Application-Specific MappingabstractTo face the complex communication problems that arise as the number of on-chip components grows up, photonic networks-on-chip (NoCs) have been recently proposed to replace electronic interconnects. However, photonic NoCs lack efficient laser sources, possibly resulting in an inefficient or inoperable architecture. In this article, we introduce a methodology for the design space exploration of optical NoC mapping solutions, which automatically assigns IPs/cores to the network tiles such that the laser power consumption is minimized. The experimental evaluation shows average reductions of 34.7% and 27.3% in the power consumption compared to, respectively, application-oblivious and randomly mapped photonic NoCs, allowing improved energy efficiency. Edoardo Fusella, Alessandro Cilardo |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | Lattice-Based Turn Model for Adaptive RoutingabstractThis paper presents a model for designing partially adaptive logic-based distributed routing algorithms. Unlike the previous methods, the Lattice-based Turn Model associates turn prohibitions with the points of different full-rank integer lattices. Due to the generality of the proposed model, existing approaches can be considered a subset of the solution space identified by this work. Morover, we propose three theorems that are instrumental to the design of lattice-based routing algorithms. In particular, the second theorem gives a necessary and sufficient condition to prove that a lattice-based routing algorithm is deadlock-free as long as the lattice basis meets certain requirements. Based on the proposed model, a novel routing algorithm, called lattice-based routing algorithm (LBRA), is presented. Simulation results exhibit encouraging performance improvements over state-of-the-art approaches when considering real and synthetic benchmarks. For instance, average 71 and 18 percent latency reductions are observed under transpose1 traffic compared to, respectively, Odd-Even and Repetitive Turn Model routing algorithms. Furthermore, LBRA achieves up to 38 and 17 percent performance improvement under real traffic as compared to Odd-Even and Repetitive Turn Model routing algorithms. Edoardo Fusella, Alessandro Cilardo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | MANGO: Exploring Manycore Architectures for Next-GeneratiOn HPC SystemsabstractThe Horizon 2020 MANGO project aims at exploring deeply heterogeneous accelerators for use in High-Performance Computing systems running multiple applications with different Quality of Service (QoS) levels. The main goal of the project is to exploit customization to adapt computing resources to reach the desired QoS. For this purpose, it explores different but interrelated mechanisms across the architecture and system software. In particular, in this paper we focus on the runtime resource management, the thermal management, and support provided for parallel programming, as well as introducing three applications on which the project foreground will be validated. José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Etienne Cappe, Alessandro Cilardo, Leon Dragic, Alexandre Dray, Alen Duspara, William Fornaciari, Gerald Guillaume, Ynse Hoornenborg, Arman Iranfar, Mario Kovac, Simone Libutti, Bruno Maitre, José Maria Martínez, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Tomás Picornell, Igor Piljic, Anna Pupykina, Federico Reghenzani, Isabelle Staub, Rafael Tornero, Marina Zapater, Davide Zoni |
DSD | 7 |
| 2017 | Path Setup for Hybrid NoC Architectures Exploiting Flooding and StandbyabstractFuture many-core systems will require energy-efficient, high-throughput and low-latency communication architectures. Silicon Photonics appears today a promising solution towards these goals. The inability of photonics networks to perform inflight buffering and logic computation suggests the use of hybrid photonic-electronic architectures. In order to exploit the full potential of photonics, it is essential to carefully design the path-setup architecture, which is a primary source of performance degradation and power consumption. In this paper, we propose a new path-setup approach which can put allocated circuits in a stand-by state, rapidly restoring them when needed. Path-setup messages are sent using a flooding routing strategy to enhance the possibility of finding free optical paths. We compare the proposed approach with a commonly used path-setup strategy as well as some other alternatives available. The results exhibit encouraging improvements in terms of both performance and energy consumption. Edoardo Fusella, José Flich, Alessandro Cilardo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | H2ONoC: A Hybrid Optical-Electronic NoC Based on Hybrid TopologyabstractNext-generation chip multiprocessors will require communication performance levels that cannot be achieved by traditional electronic ON-chip interconnects. Silicon photonics has recently emerged as a promising alternative to handle future communication needs thanks to the ultrahigh bandwidth and low power consumption. Optical networks-on-chip (ONoCs) are affected by insertion loss and crosstalk noise effects, which constrain the network scalability and impact the power consumption. This paper proposes a hybrid electronic/photonic, hybrid-topology ONoC (H2ONoC), based on a novel architecture aimed at mitigating the above effects. This paper provides a thorough description of the H2ONoC architectures as well as an experimental evaluation based on both synthetic benchmarks and real-world applications. Compared with hybrid mesh- and torus-based network-on-chip architectures, H2ONoC achieves, respectively, 13% and 18% less insertion loss, 32% and 8% less energy consumption under synthetic traffic, 74% and 14% less energy consumption with real applications, as well as better SNR when the system size scales up. Edoardo Fusella, Alessandro Cilardo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Securing the cloud with reconfigurable computing: An FPGA accelerator for homomorphic encryption
Alessandro Cilardo, Domenico Argenziano |
DATE | 1 |
| 2016 | Enabling HPC for QoS-sensitive applications: The MANGO approach
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Alessandro Cilardo, William Fornaciari, Ynse Hoornenborg, Mario Kovac, Bruno Maitre, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Fabrice Roudet, Rafael Tornero, Davide Zoni |
DATE | 6 |
| 2016 | PhoNoCMap: An application mapping tool for photonic networks-on-chip
Edoardo Fusella, Alessandro Cilardo |
DATE | 2 |
| 2016 | Design automation for application-specific on-chip interconnects: A survey
Alessandro Cilardo, Edoardo Fusella |
Integr. | 1 |
| 2016 | Crosstalk-Aware Automated Mapping for Optical Networks-on-Chip
Edoardo Fusella, Alessandro Cilardo |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Interplay of loop unrolling and multidimensional memory partitioning in HLS
Alessandro Cilardo, Luca Gallo |
DATE | 1 |
| 2015 | Variable-latency signed addition on FPGAsabstractVariable-latency, or speculative, addition is an effective technique to implement fast adders working on very long operands. Most approaches to speculative addition are either based on the assumption that operands have equiprobable independent bits, which is rarely the case in real applications due to sign-extension, or they can handle the case of signed numbers at the price of a considerable area overhead. Furthermore, many existing approaches require ad-hoc schemes preventing the reuse of standard adders typically available as optimized library components in many technologies, most notably Field-Programmable Gate Arrays. This paper introduces an innovative scheme for speculative addition that effectively addresses both problems, yielding fast and low-area circuits able to handle sign-extended numbers speculatively and only made of optimized carry-propagation adders based on fast carry circuitry as basic building blocks. Alessandro Cilardo |
FPL | 1 |
| 2015 | Scheduling-aware interconnect synthesis for FPGA-based Multi-Processor Systems-on-ChipabstractMulti-Processor System-on-Chip (MPSoC) applications can rely today on a very large spectrum of interconnection architectures determining various trade-offs between cost and performance. An automated methodology for optimizing FPGA-based MPSoC interconnect architectures is summarized in this poster paper. Based on the application communication requirements, the methodology concurrently defines the structure of the interconnect and the communication task scheduling, taking into account possible dependencies between tasks under given area constraints. The resulting architecture improves the level of communication parallelism while containing area and power costs. Edoardo Fusella, Alessandro Cilardo, Antonino Mazzeo |
FPL | 2 |
| 2015 | Exploiting Concurrency for the Automated Synthesis of MPSoC InterconnectsabstractMultiprocessor Systems-on-Chip (MPSoC) applications can rely today on a very large spectrum of interconnection topologies potentially meeting given communication requirements, determining various trade-offs between cost and performance. Building interconnects that enable concurrent communication tasks introduces decisive opportunities for reducing the overall communication latency. This work identifies three levels of parallelism at the interconnect level: global parallelism across different independent domains; local or intradomain parallelism, relying on inherently concurrent interconnect components such as crossbars; and interdomain parallelism, where multiple concurrent paths across different local domains are exploited. We propose an automated methodology to search the design space, aimed at maximizing the exploitation of these forms of parallelism. The approach also takes into consideration possible dependencies between communication tasks, which further constrains the design space, making the identification of a feasible solution more challenging. By jointly solving a scheduling and interconnect synthesis problem, the methodology turns the description of the application communication requirements, including data dependencies, into an on-chip synthesizable interconnection structure along with a communication schedule satisfying given area constraints. The article thoroughly describes the formalisms and the methodology used to derive such optimized heterogeneous topologies. It also discusses some case studies emphasizing the impact of the proposed approach and highlighting the essential differences with a few other solutions presented in the technical literature. Alessandro Cilardo, Edoardo Fusella, Luca Gallo, Antonino Mazzeo |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | New Techniques and Tools for Application-Dependent Testing of FPGA-Based ComponentsabstractField programmable gate array (FPGA) devices are increasingly being deployed in industrial environments, making reconfigurable hardware testing and reliability an active area of investigation. While FPGA devices can be tested exhaustively, the so-called application-dependent test (ADT) has emerged as an effective approach ensuring reduced testing efforts and improving the manufacturing yield since it can selectively exclude a subset of faults not affecting a given design. In addition to manufacturing, ADT can be used online, providing a solution for fast runtime fault detection and diagnostics. This paper identifies a number of issues in existing ADT techniques which limit their applicability and proposes new approaches improving the range of covered faults, with special emphasis on feedback bridging faults, as well as new algorithms for generating ADT test configurations. Furthermore, the work introduces a software environment addressing the current lack of tools, either academic or commercial, supporting ADT techniques. The architecture of the environment is highly modular and extensively based on a plug-in approach. To demonstrate the potential of the toolset, we developed a complete suite of plug-ins, based on both state-of-the-art ADT techniques and the novel approaches introduced here. The experimental results presented at the end of the paper confirm the impact of the proposed techniques. Alessandro Cilardo |
IEEE Trans. Ind. Informatics | 1 |
| 2014 | Joint communication scheduling and interconnect synthesis for FPGA-based many-core systemsabstractThis work proposes an automated methodology for optimizing FPGA-based many-core interconnect architectures. Based on the application communication requirements, the methodology concurrently defines the structure of the interconnect and the communication task scheduling, taking into account possible dependencies between tasks under given area constraints. The resulting architecture improves the level of communication parallelism that can be exploited while keeping area costs low. The paper thoroughly describes the proposed approach and discusses a few case-studies showing the impact of the proposed technique. Alessandro Cilardo, Edoardo Fusella, Luca Gallo, Antonino Mazzeo |
DATE | 1 |
| 2014 | Generating On-Chip Heterogeneous Systems from High-Level Parallel CodeabstractThis work addresses the generation of parallel on-chip heterogeneous systems starting from high-level code with explicit parallelism, based on a custom compiler and a high-level synthesis flow. Blending parallel software programming paradigms with high-level synthesis introduces a range of challenges at both the architectural level and the programming paradigm level, particularly involving the mismatches between the semantics of general high-level parallel code and the coding style imposed by high-level synthesis for FPGAs. We addressed these challenges by introducing some transformations of the source code to adapt it to the underlying hardware synthesis process, and devising a few innovative architectural solutions for supporting OpenMP functionalities. We developed a prototypical toolchain for the generation of heterogeneous hardware/software systems and used it to perform an experimental evaluation of the proposed approach. The results collected from the generated systems exhibit limited performance overheads and high application scalability, confirming the potential impact for the automated synthesis of highly parallel software applications. Alessandro Cilardo, Luca Gallo |
DSD | 1 |
| 2014 | Area implications of memory partitioning for high-level synthesis on FPGAsabstractFPGAs normally have numerous independent memory banks that can be accessed simultaneously, potentially offering a very large memory bandwidth. Adopting a suitable application-based memory partitioning strategy is thus vital to take full advantage of the memory architecture. In addition to improving the potential memory bandwidth, partitioning also affects the area complexity of the generated system because the required steering logic depends on the partitioning scheme. This work describes the area implications of a lattice-based memory partitioning technique in the context of high-level synthesis for FPGAs. Experimental results with a commercial HLS tool show that the proposed partitioning technique improves area efficiency compared to alternative approaches. Luca Gallo, Alessandro Cilardo, David B. Thomas, Samuel Bayliss, George A. Constantinides |
FPL | 2 |
| 2014 | ASP-based optimized mapping in a simulink-to-MPSoC design flow
Alessandro Cilardo, Dario Socci, Nicola Mazzocca |
J. Syst. Archit. | 1 |
| 2014 | Improving Multibank Memory Access Parallelism with Lattice-Based PartitioningabstractEmerging architectures, such as reconfigurable hardware platforms, provide the unprecedented opportunity of customizing the memory infrastructure based on application access patterns. This work addresses the problem of automated memory partitioning for such architectures, taking into account potentially parallel data accesses to physically independent banks. Targeted at affine static control parts (SCoPs), the technique relies on the Z-polyhedral model for program analysis and adopts a partitioning scheme based on integer lattices. The approach enables the definition of a solution space including previous works as particular cases. The problem of minimizing the total amount of memory required across the partitioned banks, referred to as storage minimization throughout the article, is tackled by an optimal approach yielding asymptotically zero memory waste or, as an alternative, an efficient approach ensuring arbitrarily small waste. The article also presents a prototype toolchain and a detailed step-by-step case study demonstrating the impact of the proposed technique along with extensive comparisons with alternative approaches in the literature. Alessandro Cilardo, Luca Gallo |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Efficient and scalable OpenMP-based system-level designabstractIn this work we present an experimental environment for electronic system-level design based on the OpenMP programming paradigm. Fully compliant with the OpenMP standard, the environment allows the generation of heterogeneous hardware/software systems exhibiting good scalability with respect to the number of threads and limited performance overheads. Based on well-established OpenMP benchmarks, the paper also presents some comparisons with high-performance software implementations as well as with previous proposals oriented to pure hardware translation. The results confirm that the proposed approach achieves improved results in terms of both efficiency and scalability. Alessandro Cilardo, Luca Gallo, Antonino Mazzeo, Nicola Mazzocca |
DATE | 1 |
| 2013 | Automated synthesis of FPGA-based heterogeneous interconnect topologiesabstractThe choice of the communication topology in many systems is of vital importance because it affects the entire inter-component data traffic and impacts significantly the overall system performance and cost. On the other hand, there is a very large spectrum of interconnection topologies that potentially meet given communication requirements, determining various trade-offs between cost and performance. This work proposes an automated methodology to choose among all of these possibilities, avoiding a manual and time consuming design space search process. The methodology takes as input the description of the application communication requirements, and gives as output an on-chip synthesizable interconnection structure satisfying given area constraints. Targeted at FPGA technologies, the approach generates an interconnection structure combining crossbars and shared buses, connected through bridges, yielding a scalable, efficient structure. To the best of the authors' knowledge, it provides the first method to automatically generate FPGA-based communication architectures where heterogeneous communication elements, such as shared buses and crossbar switches, coexist in a network inherently supporting multiple communication paths. The resulting architecture improves the level of communication parallelism that can be exploited, while keeping area requirements low. The paper thoroughly describes the formalisms and the methodology used to derive such optimized heterogeneous topologies. It also discusses a couple of case-study applications emphasizing the impact of the proposed approach and highlighting the essential differences with a few other solutions in the literature. Alessandro Cilardo, Edoardo Fusella, Luca Gallo, Antonino Mazzeo |
FPL | 1 |
| 2013 | Heterogeneous Computing vs. Big Data: The Case of Cryptanalytical Applications
Alessandro Cilardo |
ICA3PP (2) | 1 |
| 2013 | Design space exploration for high-level synthesis of multi-threaded applications
Alessandro Cilardo, Luca Gallo, Nicola Mazzocca |
J. Syst. Archit. | 1 |
| 2013 | Fast Parallel GF(2^m) Polynomial Multiplication for All DegreesabstractNumerous works have addressed efficient parallel GF(2m) multiplication based on polynomial basis or some of its variants. For those field degrees where neither irreducible trinomials nor Equally Spaced Polynomials (EPSs) exist, the best area/time performance has been achieved for special-type irreducible pentanomials, which however do not exist for all degrees. In other words, no multiplier architecture has been proposed so far achieving the best performance and, at the same time, being general enough to support any field degrees. In this paper, we propose a new representation, based on what we called Generalized Polynomial Bases (GPBs), covering polynomial bases and the so-called Shifted Polynomial Bases (SPBs) as special cases. In order to study the new representation, we introduce a novel formulation for polynomial basis and its variants, which is able to express concisely all implementation aspects of interest, i.e., gate count, subexpression sharing, and time delay. The methodology enabled by the new formulation is completely general and repetitive in its application, allowing the development of an ad-hoc software tool to derive proofs for area complexity and time delays automatically. As the central contribution of this paper, we introduce some new types of irreducible pentanomials and an associated GPB. Based on the above formulation, we prove that carefully chosen GPBs yield multiplier architectures matching, or even outperforming, the best special-type pentanomials from both the area and time point of view. Most importantly, the proposed GPB architectures require pentanomials existing for all degrees of practical interest. A list of suitable irreducible pentanomials for all degrees less than 1,000 is given in the appendix (Fig. 5 and Tables 4-11 are provided in a separate file containing the body of Appendix, which can be found on the Computer Society Digital Library at >http://doi.ieeecomputersociety.org/10.1109/TC.2012.63). Alessandro Cilardo |
IEEE Trans. Computers | 1 |
| 2013 | Exploiting Vulnerabilities in Cryptographic Hash Functions Based on Reconfigurable HardwareabstractCryptanalysis, i.e., the study of methods for breaking cryptographic algorithms, can greatly benefit from hardware acceleration as a key aspect enabling high-performance attacks. This work investigates the new opportunities inherently provided by a particular class of hardware technologies, i.e., reconfigurable hardware devices, addressing the cryptanalysis of the SHA-1 hash function as a case study. We show how hardware reconfiguration enables some unexplored approaches such as algorithm and architecture exploration, as well as on-the-fly system specialization relying on hardware programmability. We also identify some new cryptanalysis methods, including two novel techniques for SHA-1 cryptanalysis called interbit constraints and constraint relaxation. Relying on the proposed approaches, we designed an FPGA-based platform targeting 71- and 75-round versions of SHA-1. Under the same cost budget, the estimated times for a collision achieved by the platform are at least one order of magnitude lower than other solutions based on high-end supercomputing facilities, reaching the highest performance/cost ratio for SHA-1 collision search and providing a striking confirmation of the impact of hardware reconfigurability. Alessandro Cilardo, Nicola Mazzocca |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | The potential of reconfigurable hardware for HPC cryptanalysis of SHA-1abstractModern reconfigurable technologies can have a number of inherent advantages for cryptanalytic applications. Aimed at the cryptanalysis of the SHA-1 hash function, this work explores this potential showing new approaches inherently based on hardware reconfigurability, enabling algorithm and architecture exploration, input-dependent system specialization, and low-level optimizations based on static/dynamic reconfiguration. As a result of this approach, we identified a number of new techniques, at both the algorithmic and architectural level, to effectively improve the attacks against SHA-1. We also defined the architecture of a high-performance FPGA-based cluster, that turns out to be the solution with the highest speed/cost ratio for SHA-1 collision search currently available. A small-scale prototype of the cluster enabled us to reach a real collision for a 72-round version of the hash function. Alessandro Cilardo |
DATE | 1 |
| 2011 | Revisiting Application-Dependent Test for FPGA DevicesabstractFPGA testing poses a number of challenges related to both the complexity of the device under test and the opportunities introduced by its support to hardware reconfiguration. Application-dependent testing (ADT) provides an effective answer to these challenges. The study presented in this paper identifies some limitations of state-of-the-art ADT approaches, which prevent a complete coverage for bridging faults and the practical applicability of the algorithms for test configuration generation. The work also introduces a set of new techniques that enabled us to overcome these limitations and effectively extend previous methodologies for ADT. Alessandro Cilardo, Carmelo Lofiego, Antonino Mazzeo, Nicola Mazzocca |
ETS | 1 |
| 2011 | Exploring the Potential of Threshold Logic for Cryptography-Related OperationsabstractMotivated by the emerging interest in new VLSI processes and technologies, such as Resonant Tunneling Diodes (RTDs), Single-Electron Tunneling (SET), Quantum Cellular Automata (QCA), and Tunneling Phase Logic (TPL), this paper explores the application of the non-Boolean computational paradigms enabled by such new technologies. In particular, we consider Threshold Logic functions, directly implementable as primitive gates in the above-mentioned technologies, and study their application to the domain of cryptographic computing. From a theoretical perspective, we present a study on the computational power of linear threshold functions related to modular reduction and multiplication, the central operations in many cryptosystems such as RSA and Elliptic Curve Cryptography. We establish an optimal bound to the delay of a threshold logic circuit implementing Montgomery modular reduction and multiplication. In particular, we show that fixed-modulus Montgomery reduction can be implemented as a polynomial-size depth-2 threshold circuit, while Montgomery multiplication can be implemented as a depth-3 circuit. We also propose an architecture for Montgomery modular reduction and multiplication, which ensures feasible O(n2) area requirements, preserving the properties of constant latency and a low architectural critical path independent of the input size n. We compare this result with existing polynomial-size solutions based on the Boolean computational model, showing that the presented approach has intrinsically better architectural delay and latency, both O(1). Alessandro Cilardo |
IEEE Trans. Computers | 1 |
| 2010 | Early Prediction of Hardware Complexity in HLL-to-HDL TranslationabstractEarly prediction of hardware complexity is essential in driving hardware/software partitioning and the automatic generation of HDL descriptions from high-level code. In fact, early prediction helps estimate the “hardware cost” of a given high-level code segment before the actual synthesis, dramatically reducing the time required for an exhaustive exploration of different design choices. Clearly, this early estimation is inherently influenced by the specific toolchain for HLL-to-HDL translation. As a consequence, suitable early prediction metrics should be studied and carefully selected for each given toolchain. In this paper, we propose a general framework for the systematic study of such metrics. Unlike some previous works, the proposed framework is not specific to a given toolchain as it lets designers plug their own synthesis tool and characterize its behaviour in order to identify the most effective metrics to be used during the design space exploration. The framework is developed on top of the LLVM compiler infrastructure along with the R statistical package used to perform regression analysis. For a specific HLL-to-HDL compiler chosen for tests, we collected extensive experimental results on a large base of benchmarks, which show interesting accuracy improvements over some related work previously presented and confirm the effectiveness of the framework in deriving a characterization of the underlying hardware compiler. Alessandro Cilardo, Paolo Durante, Carmelo Lofiego, Antonino Mazzeo |
FPL | 1 |
| 2010 | A CellBE-based HPC Application for the Analysis of Vulnerabilities in Cryptographic Hash FunctionsabstractAfter some recent breaks presented in the technical literature, it has become of paramount importance to gain a deeper understanding of the robustness and weaknesses of cryptographic hash functions. In particular, in the light of the recent attacks to the MD5 hash function, SHA-1 remains currently the only function that can be used in practice, since it is the only alternative to MD5 in many security standards. This work presents a study of vulnerabilities in the SHA family, namely the SHA-0 and SHA-1 hash functions, based on a high-performance computing application run on the MariCel cluster available at the Barcelona Supercomputing Center. The effectiveness of the different optimizations and search strategies that have been used is validated by a comprehensive set of quantitative evaluations, presented in the paper. Most importantly, at the conclusion of our study, we were able to identify an actual collision for a 71-round version of SHA-1, the first ever found so far. Alessandro Cilardo, Luigi Esposito, Antonio Veniero, Antonino Mazzeo, Vicenç Beltran 0001, Eduard Ayguadé |
HPCC | 1 |
| 2009 | A new speculative addition architecture suitable for two's complement operationsabstractExisting architectures for speculative addition are all based on the assumption that operands have uniformly distributed bits, which is rarely verified in real applications. As a consequence, they may be disadvantageous for real-world workloads, although in principle faster than standard adders. To address this limitation, we introduce a new architecture based on an innovative technique for speculative global carry evaluation. The proposed architecture solves the main drawback of existing schemes and, evaluated on real-world benchmarks, it exhibits an interesting performance improvement with respect to both standard adders and alternative architectures for speculative addition. Alessandro Cilardo |
DATE | 1 |
| 2009 | Efficient Bit-Parallel GF(2^m) Multiplier for a Large Class of Irreducible PentanomialsabstractThis work studies efficient bit-parallel multiplication in GF(2m) for irreducible pentanomials, based on the so-called shifted polynomial bases (SPBs). We derive a closed expression of the reduced SPB product for a class of polynomials xm+ xks+ xks-1+ hellip + xk-1+ 1, with ks- k1les m+1/ 2. Then, we apply the above formulation to the case of pentanomials. The resulting multiplier outperforms, or is as efficient as the best proposals in the technical literature, but it is suitable for a much larger class of pentanomials than those studied so far. Unlike previous works, this property enables the choice of pentanomials optimizing different field operations (for example, inversion), yet preserving an optimal implementation of field multiplication, as discussed and quantitatively proved in the last part of the paper. Alessandro Cilardo |
IEEE Trans. Computers | 1 |
| 2008 | Virtual Scan Chains for Online Testing of FPGA-based Embedded SystemsabstractWhile techniques for offline testing of FPGAs, either manufacturing-oriented or application-oriented, are today relatively mature, in critical applications such as avionics, space, and even numerous commercial products it is often necessary to perform online testing. In this paper, we present a technique for online testing of digital designs implemented on an FPGA. The approach enables application-oriented testing, in that it covers the subset of the FPGA which is actually used for the implemented design, and considers scenarios where the FPGA component is a part of a larger embedded system. The proposed approach is in fact based on a software framework, which acts as an abstraction layer for reconfigurable hardware resources. Essentially, the framework exposes to software applications a Register-Transfer Level view of the underlying hardware, allowing test procedures to be implemented as software programs. Our approach is especially advantageous when memory is a constraint, the case of many embedded systems. As proved by experimental results, in fact, test procedures turn out to be very compact and much more memory-efficient than conventional approaches relying on static sets of FPGA testing configurations to be stored in system memory. Alessandro Cilardo, Nicola Mazzocca, Luigi Coppolino |
DSD | 1 |
| 2007 | Adaptable Parsing of Real-Time Data StreamsabstractToday's business processes are rarely accomplished inside the companies domains. More often they involve entities geographically distributed which interact in a loosely coupled cooperation. While cooperating, these entities generate transactional data streams, such as sequences of stock-market buy/sell orders, credit-card purchase records, Web server entries, and electronic fund transfer orders. Such streams are often collections of events stored and processed locally, and they thus have typically ad-hoc, heterogeneous formats. On the other hand, elements in such data streams usually share a common semantics and indeed they can be profitably mined in order to obtain combined global events. In this paper, we present an approach to the parsing of heterogeneous data streams based on the definition of format-dependent grammars and automatic production of ad-hoc parsers. The stream-dependent parsers can be obtained dynamically in a totally automatic way, provided that the appropriate grammar, written in a common format, is fed into the system. We also present a fully working implementation, that has been successfully integrated into a telecommunication environment for real-time processing of billing information flows Ferdinando Campanile, Alessandro Cilardo, Luigi Coppolino, Luigi Romano |
PDP | 2 |
| 2007 | Combining Programmable Hardware and Web Services Technologies for Delivering High-Performance and Interoperable SecurityabstractInformation security is a key requirement in emerging networked scenarios, which typically involve a large variety of heterogeneous, often resource-constrained devices. Providing security to this emerging class of distributed applications raises a number of new challenges. This paper discusses such challenges with respect to two key security services, namely public key certification and digital timestamping, and presents a multi-tier architecture which combines a hardware-accelerated back-end and a Web Services based Web tier to for achieving interoperability while boosting performance. The paper describes the organization of the multi-tier architecture, provides a detailed description of individual components, and presents the results of a thorough experimental campaign Alessandro Cilardo, Luigi Coppolino, Antonino Mazzeo, Luigi Romano |
PDP | 1 |
| 2007 | Performance Evaluation of Security Services: An Experimental ApproachabstractRecent advances in wireless technologies have enabled pervasive connectivity to Internet scale systems which include heterogeneous mobile devices, such as mobile phones and personal digital assistants, a trend which is generally referred to as ubiquitous computing. This leads to the need for providing security functions to applications which are partially deployed over wireless devices. Delivering security services to mobile devices raises a number of challenging issues, mostly related to the limited amount of computing power which is typically available on the target plat-forms. Some promising solutions rely on multi-tier architectures, which are based on the emerging Web services technology. In this scenario, understanding the impact of architectural characteristics of specific platforms is a key issue for practitioners who have to develop and deploy efficient security-enabled applications on mobile devices. This paper provides an experimental study of the impact that specific characteristics of individual mobile device platforms have on the final performance of security applications. Focus is on performance and resource utilization, which are key aspects when one develops applications on mobile devices. The case study is a Web services based solution for delivering public key infrastructure (PKI) services to mobile devices. Experiments have been conducted on three different mobile terminals, which span a large range of characteristics in the class of resource-constrained devices. Results show that: i) performance figures are not uniform in spite of similar underlying hardware characteristics, and ii) security and performance are often conflicting requirements Alessandro Cilardo, Luigi Coppolino, Antonino Mazzeo, Luigi Romano |
PDP | 1 |
| 2006 | Elliptic Curve Cryptography EngineeringabstractIn recent years, elliptic curve cryptography (ECC) has gained widespread exposure and acceptance, and has already been included in many security standards. Engineering of ECC is a complex, interdisciplinary research field encompassing such fields as mathematics, computer science, and electrical engineering. In this paper, we survey ECC implementation issues as a prominent case study for the relatively new discipline of cryptographic engineering. In particular,we show that the requirements of efficiency and security considered at the implementation stage affect not only mere low-level, technological aspects but also, significantly, higher level choices, ranging from finite field arithmetic up to curve mathematics and protocols. Alessandro Cilardo, Luigi Coppolino, Nicola Mazzocca, Luigi Romano |
Proc. IEEE | 1 |
| 2005 | A Novel Unified Architecture for Public-Key CryptographyabstractWe propose a fully-parallel, bit-sliced unified architecture designed to perform modular multiplication/exponentiation and GF(2/sup M/) multiplication as the core operations of RSA and EC cryptography. The architecture uses a radix-2 Montgomery technique for modular arithmetic, and a radix-4 MSD-first approach for GF(2/sup M/) multiplication. To the best of our knowledge, it is the first unified proposal based on such a hybrid approach. The architecture structure is bit-sliced and is highly regular, modular, and scalable, as virtually any datapath length can be obtained at a linear cost in terms of hardware resources and no costs in terms of critical path. Our proposal outperforms all similar unified architectures found in the technical literature in terms of clock count and critical path. The architecture has been implemented on a field-programmable gate array (FPGA) device. A highly compact and efficient design was obtained taking advantage of the architectural characteristics. Alessandro Cilardo, Antonino Mazzeo, Nicola Mazzocca, Luigi Romano |
DATE | 1 |
| 2005 | High-Performance and Interoperable Security Services for Mobile Environments
Alessandro Cilardo, Luigi Coppolino, Antonino Mazzeo, Luigi Romano |
HPCC | 1 |
| 2005 | Reconfigurable systems self-healing using mobile hardware agentsabstractTechnology constantly follows the needs of the marketplace. The current trend of producing digitally-aware environments keeps growing, and it is very likely that in the future we will be surrounded by many heterogeneous devices which communicate each other through distributed and wireless interfaces. We will see a new society of digital systems, whose individuals will be ubiquitous and heterogeneous systems interacting together and providing high productivity and great flexibility. Two technologies seem to be emerging to support this new paradigm: mobile data agents to handle the complexity and heterogeneity of networked infrastructures, and real-time reconfigurable systems to implement flexible, adaptable, and high performance individuals. This paper analyzes how the innovative aspects of these new technologies could be exploited to implement innovative and efficient test and repair strategies Alfredo Benso, Alessandro Cilardo, Nicola Mazzocca, Liviu Miclea, Paolo Prinetto, Szilárd Enyedi |
ITC | 2 |
| 2004 | Carry-Save Montgomery Modular Exponentiation on Reconfigurable HardwareabstractIn this paper we present a hardware implementation of the RSA algorithm for public-key cryptography. Basically, the RSA algorithm entails a modular exponentiation operation on large integers, which is considerably time-consuming to implement. To this end, we adopted a novel algorithm combining the Montgomery's technique and the carry-save representation of numbers. A highly modular, bit-slice based architecture has been designed for executing the algorithm in hardware. We also propose an FPGA-based implementation of the architecture developed. The characteristics of the algorithm, the regularity of the architecture, and the data-flow aware placement of the FPGA resources resulted in a considerable performance improvement, as compared to other implementations presented in the literature. Alessandro Cilardo, Antonino Mazzeo, Luigi Romano, Giacinto Paolo Saggese |
DATE | 1 |
| 2004 | A Web Services Based Architecture for Digital Time Stamping
Alessandro Cilardo, Antonino Mazzeo, Luigi Romano, Giacinto Paolo Saggese, Giuseppe Cattaneo |
J. Web Eng. | 1 |