EDBT 2026 Demo / reviewers in the wild / expert
Alberto García Ortiz
dblp:80/3384 · also Alberto García-Ortiz
· DBLP profile ↗
51ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-6461-3864ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 2 since 2021Computer networks · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: COIN-3D - Collaborative Innovation in 3D VLSI ReliabilityabstractAs semiconductor manufacturing advances from the 3-nm process toward the sub-nanometer regime and transitions from FinFETs to gate-all-around field-effect transistors (GAAFETs), the resulting complexity and manufacturing challenges continue to increase. In this context, 3D chiplet-based approaches have emerged as key enablers to address these limitations while exploiting the expanded design space. Specifically, chiplets help address the lower yields typically associated with large monolithic designs. This paradigm enables the modular design of heterogeneous systems consisting of multiple chiplets (e.g., CPUs, GPUs, memory) fabricated using different technology nodes and processes. Consequently, it offers a capable and cost-effective strategy for designing heterogeneous systems.This paper introduces the Horizon Europe Twinning project COIN-3D (Collaborative Innovation in 3D VLSI Reliability), which aims to strengthen research excellence in 2.5D/3D VLSI systems reliability through collaboration between leading European institutions. More specifically, our primary scientific goal is the provision of novel open-source Electronic Design Automation (EDA) tools for reliability assessment of 3D systems, integrating advanced algorithms for physical- and system-level reliability analysis. George Rafael Gourdoumanis, Fotoini Oikonomou, Maria Pantazi-Kypraiou, Pavlos Stoikos, Olympia Axelou, Athanasios Tziouvaras, Georgios Karakonstantis, Tahani Aladwani, Christos Anagnostopoulos 0001, Yixian Shen, Anuj Pathania, Alberto García Ortiz, George Floros 0002 |
DATE | 12 |
| 2025 | 3DPX - An Open-Source Methodology for 3D Physical Design ExplorationabstractArchitectural exploration of novel technological options, such as 3D integration, requires close interaction with physical implementation. However, the lack of information regarding the technical aspects of the 3D flavor, and lack of open source tools are the key obstacles inhibiting the wide use of physicalaware 3D architectural exploration. To palliate this problem, we present 3DPX, an open-source methodology for 3D physical design exploration based on the OpenROAD framework. By leveraging open standards and tools, our methodology enables evaluation of the impact of 3D stacking on performance, power, and area (PPA). Experimental results on a set of RISC-V benchmark circuits show the expected 50% area reduction, while allowing users to analyze and optimize timing and power just as they would in a traditional 2D flow. George Rafael Goudroumanis, Maria Pantazi-Kypraiou, George Floros 0002, Athanasios Tziouvaras, Georgios I. Stamoulis, Alberto García Ortiz |
ICCD | 6 |
| 2024 | GNN-assisted Back-side Clock Routing Methodology for Advance TechnologiesabstractThe back-side metal layers exhibit lower parasitics compared to the front-side layers in advanced technologies, making them suitable for clock-net distribution. In this study, we explore the advantages of using back-side metal layers for clock routing, which is shared with a power delivery network. Our Graph Neural Network (GNN) based framework, effectively distributes the clock-tree between the front and back sides. We address the back-side clock nets' creation by incorporating back-side buffers. Our results demonstrate better clock and full-chip metrics represented by an increase of up to 13% in the effective frequency with equivalent power consumption, using 3 nm technology. Nesara Eranna Bethur, Pruek Vanna-Iampikul, Odysseas Zografos, Lingjun Zhu, Giuliano Sisto, Dragomir Milojevic, Alberto García Ortiz, Geert Hellings, Julien Ryckaert, Francky Catthoor, Sung Kyu Lim |
DAC | 7 |
| 2024 | ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration
Zeqi Zhu, Alberto García Ortiz, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira |
ECCV (11) | 2 |
| 2024 | Hier-3D: A Methodology for Physical Hierarchy Exploration of 3-D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. Complemented with an automated floorplanning solution, the flow allows for system-level physical and architectural exploration of 3D designs. As a result, we significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 -2.2× compared with 2D, where all metrics are improved simultaneously, including up to 20% power savings. Nesara Eranna Bethur, Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | MemPool-3D: Boosting Performance and Efficiency of Shared-L1 Memory Many-Core Clusters with 3D IntegrationabstractThree-dimensional integrated circuits promise power, performance, and footprint gains compared to their 2D counter-parts, thanks to drastic reductions in the interconnects' length through their smaller form factor. We can leverage the potential of 3D integration by enhancing MemPool, an open-source many-core design with 256 cores and a shared pool of L1 scratchpad memory connected with a low-latency interconnect. MemPool's baseline 2D design is severely limited by routing congestion and wire propagation delay, making the design ideal for 3D integration. In architectural terms, we increase MemPool's scratchpad memory capacity beyond the sweet spot for 2D designs, improving performance in a common digital signal processing kernel. We propose a 3D MemPool design that leverages a smart partitioning of the memory resources across two layers to balance the size and utilization of the stacked dies. In this paper, we explore the architectural and the technology parameter spaces by analyzing the power, performance, area, and energy efficiency of MemPool instances in 2D and 3D with 1 MiB, 2 MiB, 4 MiB, and 8 MiB of scratchpad memory in a commercial 28 nm technology node. We observe a performance gain of 9.1% when running a matrix multiplication on MemPool-3D with 4 MiB of scratchpad memory compared to the MemPool 2D counterpart. In terms of energy efficiency, we can implement the MemPool-3D instance with 4 MiB of L1 memory on an energy budget 15 % smaller than its 2D counterpart, and 3.7 % smaller than the MemPool-2D instance with a fourth of the L1 scratchpad memory capacity. Matheus A. Cavalcante, Anthony Agnesina, Samuel Riedel, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Sung Kyu Lim, Luca Benini |
DATE | 5 |
| 2022 | Hier-3D: A Hierarchical Physical Design Methodology for Face-to-Face-Bonded 3D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. We significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 to 2.2 × compared with 2D, where all metrics are improved simultaneously, including up to power savings. Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
ISLPED | 3 |
| 2022 | DiNS: Nature Disaster in Network SimulationsabstractWireless sensor networks (WSNs) is a promising solution for disaster management because of its scalability and low cost operations. However, testing the effectiveness of WSNs in real-world disasters is time consuming, costly and in some cases even infeasible. In this paper, we propose a Disaster in Network Simulations (DiNS) framework based on OMNeT ++ to replicate a WSN deployed in a disaster in a simulated environment. DiNS allows researchers to observe how the disaster influences the sensor network and how the network can respond in return. A generic coupling interface is developed to support different disaster types. Moreover, we develop a verification tool for functional debugging and verification during new functions de-veloping of sensor nodes, and an optimization tool to support the mathematical optimization of the network operations in response to the disaster. The functionality of DiNS is demonstrated with a case study using wildfire disaster. It provides an easy way for validating and optimizing the disaster management with WSNs. Nisal Hemadasa, Wanli Yu, Yanqiu Huang, Leonardo Sarmiento, Amila Wickramasinghe, Alberto García Ortiz |
MSN | 6 |
| 2021 | Bridging the Frequency Gap in Heterogeneous 3D SoCs through Technology-Specific NoC Router ArchitecturesabstractIn heterogeneous 3D System-on-Chips (SoCs), NoCs with uniform properties suffer one major limitation; the clock frequency of routers varies due to different manufacturing technologies. For example, digital nodes allow for a higher clock frequency of routers than mixed-signal nodes. This large frequency gap is commonly tackled by complex and expensive pseudo-mesochronous or asynchronous router architectures. Here, a more efficient approach is chosen to bridge the frequency gap. We propose to use a heterogeneous network architecture. We show that reducing the number of VCs allows to bridge a frequency gap of up to 2x. We achieve a system-level latency improvement of up to 47% for uniform random traffic and up to 59% for PARSEC benchmarks, a maximum throughput increase of 50%, up to 68% reduced area and 38% reduced power in an exemplary setting combining 15-nm digital and 30-nm mixed-signal nodes and comparing against a homogeneous synchronous network architecture. Versus asynchronous and pseudo-mesochronous router architectures, the proposed optimization consistently performs better in area, in power and the average flit latency improvement can be larger than 51%. Jan Moritz Joseph, Lennart Bamberg, Geonhwa Jeong, Ruei-Ting Chien, Rainer Leupers, Alberto García Ortiz, Tushar Krishna, Thilo Pionteck |
ASP-DAC | 6 |
| 2021 | Power, Performance, Area and Cost Analysis of Memory-on-Logic Face-to-Face Bonded 3D Processor DesignsabstractIn this paper, we present a power, performance, area and cost (PPAC) analysis for large-scale 3D processor designs based on wafer-to-wafer bonding. From the evaluation of our cost model, we investigate a typically disregarded opportunity in 3D that is area savings due to buffer savings and better routability, offering unexpected cost savings. We explore the viability of this factor with the feedback of a state-of-the-art 3D memory-on-logic implementation flow. We show how this affects the PPAC of full-chip GDS implementations of a large-scale manycore processor design. Experiments show that our memory-on-logic 3D implementation offers 7% silicon area savings, resulting in 53.5% footprint reduction. We also obtain a 40% power-performance-cost improvement compared with 2D counterparts Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Manu Perumkunnil Komalan, Sung Kyu Lim |
ISLPED | 4 |
| 2021 | Combination of Task Allocation and Approximate Computing for Fog-Architecture-Based IoTabstractAchieving energy efficiency is always a primary concern for fog-architecture-based Internet of Things (IoT) applications. As the IoT devices are typically of small sizes and powered by battery energy, it is essential to address the energy consumption at all levels from the circuit to the system. Two of the promising solutions at circuit and system levels are approximate computing and energy-aware task allocation, respectively. However, the existing task allocation approaches are designed without considering the aspect of approximate computing. In this work, we fill this gap and aim to maximize the network lifetime subject to the accuracy requirements of the applications. By considering both the approximate computing and task allocation simultaneously, a nonlinear problem is obtained to allocate the tasks for the devices (fog nodes and IoT end devices) and to select the corresponding execution modes (tasks in approximate or exact modes). To efficiently solve this problem, a centralized algorithm is first proposed by transferring the above nonlinear problem as a linear programming (LP) problem. As executing the centralized algorithm is a challenge for the resource-limited IoT devices, this work further proposes an optimal distributed algorithm based on Dantzig-Wolfe decomposition to solve the problem of tasks distribution and execution modes selection. The centralized large-scaled LP problem is decomposed into small-scaled subproblems, which can be efficiently solved by each IoT device. The proposed algorithms are tested by extensive simulations. The results demonstrate that the distributed algorithm achieves the same results as the centralized algorithm, and both of them significantly outperform the previous approaches. Wanli Yu, Ardalan Najafi, Yanqiu Huang, Alberto García Ortiz |
IEEE Internet Things J. | 4 |
| 2021 | High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A ProcessorsabstractMonolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively. Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2020 | Macro-3D: A Physical Design Methodology for Face-to-Face-Stacked Heterogeneous 3D ICsabstractMemory-on-logic and sensor-on-logic face-to-face stacking are emerging design approaches that promise a significant increase in the performance of modern systems-on-chip at reasonable costs. In this work, a netlist-to-layout design flow for such heterogeneous 3D systems is proposed. The proposed technique overcomes the severe limitations of existing 3D physical design methodologies. A RISC-V-based multi-core system, implemented in a commercial technology, is used as a case study to evaluate the proposed design flow. The case study is performed for modern/large and small cache sizes to show the superiority of the proposed methodology for a broad set of systems. While previous 3D design flows do not show to optimize performance against 2D baseline designs for processor systems with a significant memory area occupation, the proposed flow shows a performance and power improvement by 20.4-28.2% and 3.2-3.8%, respectively. Lennart Bamberg, Alberto García Ortiz, Lingjun Zhu, Sai Pentapati, Da Eun Shim, Sung Kyu Lim |
DATE | 2 |
| 2020 | Stochastic Mixed-PR: A Stochastically-Tunable Low-Error AdderabstractApproximate computing is a promising paradigm in low-power design to trade off power efficiency by accuracy. However, its usability in real applications is restricted due to the lack of a dynamic configuration of the error characteristics. Most of the existing approximate adders have a fixed level of accuracy. On the other hand, the approximate adders which are configurable can only switch between an exact mode and a fixed level of approximation. This brief presents the mathematical stochastic error analysis of the adders, for the first time. Furthermore, based on the analysis, we propose to use a mixed adder as a reconfigurable architecture. This adder outperforms the state-of-the-art configurable adder. Moreover, it reduces the energy-delay product considerably in comparison with its conventional counterpart. Ardalan Najafi, Alberto García Ortiz |
ISCAS | 2 |
| 2020 | TAAC: Task Allocation Meets Approximate Computing for Internet of ThingsabstractUltra-low-power operation, as required by Internet of Things (IoT) systems, requires to address energy consumption at all levels from circuit to system. Two of the promising solutions at circuit and system levels are approximate computing and energy aware task allocation, respectively. However, the existing task allocation approaches are designed without considering the aspect of approximate computing. For the first time, this work proposes an optimal Task Allocation algorithm taking the Approximate Computing into account (TAAC) to fill this gap. The problem of tasks assignment and executing modes selection (approximate or exact modes of the tasks) can be efficiently solved by formulating it as a linear programming problem. The extensive simulation results show that the proposed TAAC algorithm significantly outperforms the previous approaches. Wanli Yu, Ardalan Najafi, Yarib Nevarez, Yanqiu Huang, Alberto García Ortiz |
ISCAS | 5 |
| 2020 | Misalignment-aware energy modeling of narrow buses for data encoding schemes
Amir Najafi 0001, Lennart Bamberg, Alberto García Ortiz |
Integr. | 3 |
| 2019 | Symbolic Circuit Analysis under an Arc Based Timing ModelabstractTools for Automatic Test Pattern Generation (ATPG) typically abstract timing. When a more detailed timing model is needed, either simulation or statistic timing analysis is usually applied. Our symbolic engine based on Satisfiability Modulo Theories can reason over a model using pin-to-pin timing arcs as available after synthesis or after place and route. We study the differences of a simple fixed gate delay versus the arc-based timing model. Görschwin Fey, Alberto García Ortiz |
ETS | 2 |
| 2019 | Poster: Event-triggered State Estimation Meets Duty Cycling Protocol
Yanqiu Huang, Wanli Yu, Alex Leong, Alberto García Ortiz |
EWSN | 4 |
| 2019 | System-Level Optimization of Network-on-Chips for Heterogeneous 3D System-on-ChipsabstractFor a system-level design of Networks-on-Chip for 3D heterogeneous System-on-Chip (SoC), the locations of components, routers and vertical links are determined from an application model and technology parameters. In conventional methods, the two inputs are accounted for separately; here, we define an integrated problem that considers both application model and technology parameters. We show that this problem does not allow for exact solution in reasonable time, as common for many design problems. Therefore, we contribute a heuristic by proposing design steps, which are based on separation of intralayer and interlayer communication. The advantage is that this new problem can be solved with well-known methods. We use 3D Vision SoC case studies to quantify the advantages and the practical usability of the proposed optimization approach. We achieve up to 18.8% reduced white space and up to 12.4% better network performance in comparison to conventional approaches. Jan Moritz Joseph, Dominik Ermel, Lennart Bamberg, Alberto García Ortiz, Thilo Pionteck |
ICCD | 4 |
| 2019 | HotAging - Impact of Power Dissipation on Hardware DegradationabstractSafety and dependability are of utmost importance for many integrated systems. Hence, it must be guaranteed throughout the whole system's lifetime that no ambient and internal influences can affect the system's integrity. Under this scope and having in mind the side-effects of today's nanoscale technologies, hardware degradation is of rising concern. However, related studies should not solely focus on aging effect itself, but also consider its relation to any accelerating factors, especially temperature. Towards this end, this work presents a study on how the power dissipation of a circuit, and thus, its temperature, can expedite wear-out effects. Therefore, three different analysis are performed-aging without and with consideration of temperature and the study on how guard-banding strategies are affected. In order to distinguish random and, maliciously intended or accidentally produced, worst case scenarios, we implemented an algorithm that determines a combination of input vectors that forces high aging states and high power dissipation. Results indicate that aging under consideration of temperature can increase circuit delay by more than 26% (random case) and by nearly 40% (worst case). That means, if a maximum acceptable delay degradation is defined, designs can enter malfunction states already in a period of weeks (worst case) or months (random case). These results underline the importance of considering power dissipation, and thus temperature, when doing aging analysis and aging verification. Frank Sill, Alberto García Ortiz, Rolf Drechsler |
ISCAS | 2 |
| 2019 | Crosstalk optimization for through-silicon vias by exploiting temporal signal misalignment
Lennart Bamberg, Jan Moritz Joseph, Thilo Pionteck, Alberto García Ortiz |
Integr. | 4 |
| 2019 | Edge effect aware low-power crosstalk avoidance technique for 3D integration
Lennart Bamberg, Amir Najafi 0001, Alberto García Ortiz |
Integr. | 3 |
| 2019 | Simulation environment for link energy estimation in networks-on-chip with virtual channels
Jan Moritz Joseph, Lennart Bamberg, Imad Hajjar, Robert Schmidt 0003, Thilo Pionteck, Alberto García Ortiz |
Integr. | 6 |
| 2019 | FLINT+: A runtime-configurable emulation-based stochastic timing analysis framework
Moritz Weißbrich, Lukas Gerlach 0001, Holger Blume, Ardalan Najafi, Alberto García Ortiz, Guillermo Payá-Vayá |
Integr. | 5 |
| 2019 | EPKF: Energy Efficient Communication Schemes Based on Kalman Filter for IoTabstractThe Internet of Things (IoT) has been recognized as the next technological revolution. It faces two challenges: 1) how to achieve energy efficient communication for the battery constrained devices and 2) how to connect a very large number of devices to the Internet with low latency, high efficiency, and reliability. To address these problems, this paper proposes two methods based on Kalman filter (KF), termed as extensions of predicable KF (EPKF). They locally reduce the unnecessary transmission (access) of end devices to the network (Internet) utilizing the spatial and temporal correlations with low algorithmic overhead. Each transmitting device (TD) independently controls its transmission using the temporal correlation; and the receiving device (RD) exploits the spatial correlation among the TDs to further improve the reconstruction quality. The reconstruction problem in the RD is nonlinear. To reduce the computation complexity, an in-depth analysis of the local estimate error is conducted and the approximated linear solutions are thereupon obtained. They are fundamental methods applicable to any IoT monitored/controlled physical system that can be modeled as a linear state space representation. The pedestrian-position application is used as a case study to demonstrate the efficiency in the simulation. Remarkably, the EPKF methods using the linear combinations of the local estimates from multiple TDs reduce the transmission rate to 10%, while achieving the same reconstruction quality as using KF in the traditional manner. Yanqiu Huang, Wanli Yu, Enjie Ding, Alberto García Ortiz |
IEEE Internet Things J. | 4 |
| 2019 | Comparing vertical and horizontal SIMD vector processor architectures for accelerated image feature extraction
Moritz Weißbrich, Alberto García Ortiz, Guillermo Payá-Vayá |
J. Syst. Archit. | 2 |
| 2019 | Coding-Based Low-Power Through-Silicon-Via Redundancy Schemes for Heterogeneous 3-D SoCsabstractThree-dimensional integration, employing through-silicon vias (TSVs), improves the system-on-chip (SoC) performance. However, redundancy schemes are required to cope with the relatively poor TSV manufacturing yield. Existing redundancy schemes do not exploit technological heterogeneity between the dies. Hardware costs can differ for the individual dies. This demands asymmetrical schemes with low complexity in costly mixed signal or RF dies. Furthermore, redundant TSVs are only used in the case of a defect. In the most probable case of correct manufacturing, they are unused. Another emerging technique using redundant lines is low-power coding (LPC). This paper presents a hybrid TSV redundancy technique based on coding, which can be used for LPC and for yield enhancement. Furthermore, the approach is strongly asymmetric. In case of a fault, a configuration is only required for the encoder or decoder located in the cheaper die, while in the costly die, a minimal set of XOR gates is sufficient. A case study for an existing heterogeneous SoC shows that the proposed technique decreases area overhead and power consumption compared to the best previous technique by over 69 % and 33%, respectively. Lennart Bamberg, Alberto García Ortiz |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Coding approach for low-power 3D interconnectsabstractThrough-silicon vias (TSVs) in 3D ICs show a significant power consumption, which can be reduced using coding techniques. This work presents an approach which reduces the TSV power consumption by a signal-aware bit assignment which includes inversions to exploit the MOS effect. The approach causes no overhead and results in a guaranteed reduction of the overall power consumption. An analysis of our technique shows a reduction in the TSV power consumption by up to 48 % for real correlated data streams (e.g. image sensor), and 11 % for low-power encoded random data streams. Lennart Bamberg, Robert Schmidt 0003, Alberto García Ortiz |
DAC | 3 |
| 2018 | Reliability Improvements for Multiprocessor Systems by Health-Aware Task SchedulingabstractThe probability that a particular device is operational for a given duration, or reliability, is a dependability attribute and key metric for systems in critical applications. For example, systems for long-term autonomous exploration missions have to be operational during their complete mission. Other critical applications like banking, medical automotive or aerospace face similar reliability requirements that are only met by dependable systems. Traditional dependable systems, compared to their non-dependable counterparts, have three key issues: They are more expensive, consume more power, and provide less performance. Robert Schmidt 0003, Rehab Massoud, Jaan Raik, Alberto García Ortiz, Rolf Drechsler |
IOLTS | 4 |
| 2018 | A comprehensive survey on wireless sensor node hardware platforms
Fatma Karray, Mohamed Wassim Jmal, Alberto García Ortiz, Mohamed Abid, Abdulfattah Mohammad Obeid |
Comput. Networks | 3 |
| 2018 | Edge effects on the TSV array capacitances and their performance influence
Lennart Bamberg, Amir Najafi 0001, Alberto García Ortiz |
Integr. | 3 |
| 2018 | Systematic Design of an Approximate Adder: The Optimized Lower Part Constant-OR Adder
Ayad Dalloo, Ardalan Najafi, Alberto García Ortiz |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Temporal redundancy latch-based architecture for soft error mitigationabstractCurrent transients caused by energetic particle strikes are a serious threat for digital circuits in aerospace applications. Such single-event transients (SETs) can corrupt the circuit state, with possibly devastating consequences. Although it is possible to protect circuits with spatial redundancy techniques, the area and power overhead is high. Therefore aerospace circuits would benefit from adopting temporal redundancy instead, but existing solutions prioritize performance over reliability. Our proposed temporal redundancy latch-based architecture (TRLA) is a standard cell, static CMOS temporal redundancy technique, with area savings of 26%, power savings of 46%, and 14% faster circuit operation compared to triple modular redundancy (TMR). Robert Schmidt 0003, Alberto García Ortiz, Görschwin Fey |
IOLTS | 2 |
| 2017 | High-Level Energy Estimation for Submicrometric TSV ArraysabstractThe 3-D integration using through silicon vias (TSVs) is one of the most promising approaches to overcome the interconnect delay problem of current CMOS technologies. Nevertheless, the TSV energy consumption is not negligible due to the high capacitive coupling. This paper presents an abstract and yet accurate model to estimate the pattern-dependent energy consumption in arrays of TSVs; it is the first high-level model including the effects of the voltage-dependent metal-oxide-semiconductor (MOS) capacitances surrounding each TSV and a possible temporal misalignment between the input signals. We propose a regression method to estimate the dynamic size of the coupling capacitances as a function of the bit probabilities. Experimental results for real and synthetic data streams, a submicrometer 9-bit TSV array and a 65-nm technology show that the presented TSV energy model exhibits a maximum error of 5.53%, while the traditional high-level model shows errors of up to 79.77%. Furthermore, the new insights provided by our model reveal a possibility to easily boost the efficiency of existing low-power codes for TSV structures by over 10% without affecting the coding efficiency for the planar metal wires or the encoder complexity. Lennart Bamberg, Alberto García Ortiz |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Synthesis of approximate coders for on-chip interconnects using reversible logic
Robert Wille, Oliver Keszöcze, Stefan Hillmich, Marcel Walter, Alberto García Ortiz |
DATE | 5 |
| 2016 | Modeling Optimal Dynamic Scheduling for Energy-Aware Workload Distribution in Wireless Sensor NetworksabstractEnergy-aware workload distribution becomes crucial for extending the lifetime of wireless sensor networks (WSNs) in complex applications as those in Internet-of-Things or in-network DSP processing scenarios. Today static workload schedules are well understood, while dynamic schedules (i.e., with multiple partitions) remain unexplored. This paper models the dynamic scheduling by considering both the communication and computation energy consumption. It formulates a series of (integer) linear programming problems to characterize the optimal scheduling strategies. Surprisingly, even 2-partition scheduling can provide the maximum gains. Besides the interest to evaluate the optimality of on-line heuristics for dynamic scheduling, the reported off-line strategies can be immediately applied to WSN applications. Wanli Yu, Yanqiu Huang, Alberto García Ortiz |
DCOSS | 3 |
| 2016 | PKF-ST: A Communication Cost Reduction Scheme Using Spatial and Temporal Correlation for Wireless Sensor Networks
Yanqiu Huang, Wanli Yu, Alberto García Ortiz |
EWSN | 3 |
| 2016 | EARNPIPE: A Testbed for Smart Water Pipeline Monitoring Using Wireless Sensor NetworkabstractLarge quantities of water are wasted daily due to leakages in pipelines. In order to decrease this waste and preserve water, advanced systems could be used. In this context, a Wireless Sensor Network (WSN) is increasingly required to optimize the reliability of the inspection and improve the accuracy of the water pipeline monitoring. A WSN solution is proposed in this paper with a view to detecting and locating leaks for long distance pipelines. It combines powerful leak detection and localization algorithms and an efficient wireless sensor node System on Chip (SoC) architecture. In fact, a novel hybrid Water Pipeline Monitoring (WPM) method has been proposed using Leak detection Predictive Kalman Filter (LPKF) and Modified Time Difference of Arrival (TDOA) method based on pressure measurements. The data collected from sensors are filtered, analyzed and compressed with the same Kalman Filter (KF) based algorithm instead of using various algorithms that deeply damage the battery of the node. The local low power pre-processing is efficient to save the power of the sensor nodes. Moreover, a laboratory testbed has been constructed using plumbing components and validated by an ARM-based prototyping platform with pressure sensors. Fatma Karray, Alberto García Ortiz, Mohamed Wassim Jmal, Abdulfattah Mohammad Obeid, Mohamed Abid |
KES | 2 |
| 2016 | Analysis of PKF: A Communication Cost Reduction Scheme for Wireless Sensor NetworksabstractEnergy efficiency is a primary concern for wireless sensor networks (WSNs). One of its most energy-intensive processes is the radio communication. This work uses a predictor combined with a Kalman filter (KF) to reduce the communication energy cost for cluster-based WSNs. The technique, called PKF, is suitable for typical WSN applications with adjustable data quality and tens of picojoule computation cost. However, it is challenging to precisely quantify its underlying process from a mathematical point of view. Through an in-depth mathematical analysis, we formulate the tradeoff between energy efficiency and reconstruction quality of PKF. One of our prominent results for that is the explicit expression for the covariance of the doubly truncated multivariate normal distribution; it improves the previous methods and has generality. The validity and accuracy of the analysis are verified with both artificial and real signals. The simulation results, using real temperature values, demonstrate the efficiency of PKF: without additional data degradation, it reduces the communication cost by more than 88%. Compared to previous works based on KF, PKF requires less computational effort while improving the reconstruction quality; compared with the techniques without KF, the advantages of PKF are even more significant. It reduces the transmission rate of them by at least 29%. Besides, it can be integrated into network level techniques to further extend the whole network lifetime. Yanqiu Huang, Wanli Yu, Christof Osewold, Alberto García Ortiz |
IEEE Trans. Wirel. Commun. | 4 |
| 2015 | A rapid prototyping framework for nano-photonic acceleratorsabstractA rapid-prototyping framework for nano-photonic accelerators embedded in digital multiprocessor systems is highly desired. Nano-photonic technologies advance at a fast pace. They provide a promising technological substrate for implementing on-chip optical computing. Currently, numerous competing physical implementations and architectural strategies are under active investigation; however, the functionality, performance, and error characteristics of these technologies differ substantially from those of standard CMOS technologies, making system exploration hard. This work presents a technology-agnostic rapid prototyping framework for nano-photonic accelerators focusing on optical analog processing and digital optical gates. It allows to explore the implementation-space of optical accelerators and to analyse the influence of optical effects on the overall system at an early stage. Wolfgang Büter, Alberto García Ortiz, S. Mahmood, S. Arefin, V. V. Parsi Sreenivas, R. B. Bergman |
FPL | 2 |
| 2015 | An altruistic compression-scheduling scheme for cluster-based wireless sensor networksabstractData compression using temporal and/or spatial correlations has been extensively studied to prolong the lifetime of wireless sensor networks (WSNs). In order to maximize the gain of these techniques, this work proposes an off-line altruistic compression-scheduling (ACS) scheme for cluster-based WSNs. It schedules when and how (i.e. without compression, with either of temporal and spatial compression, or with both of them) each sensor node executes compression techniques. The optimum scheduling solution to maximize the network lifetime is obtained by solving a linear program, whose computational complexity and runtime are efficiently reduced by a grouping and filtering algorithm. In addition, we propose a transmission power increment (TPI) method for WSNs with isolated nodes to improve the spatial compression possibilities. It can be used in ACS to further extend the lifetime of such networks. The efficiency of ACS is demonstrated by extensive simulation results using realistic models. It increases the network lifetime by a factor of 1.67 to 12.49 depending on different temporal correlations for typical WSNs. Compared with previous scheduling methods, it further extends the network lifetime with even less runtime. The combination of TPI and ACS significantly prolongs the lifetime of networks with isolated nodes. Wanli Yu, Yanqiu Huang, Alberto García Ortiz |
SECON | 3 |
| 2014 | DCM: An IP for the autonomous control of optical and electrical reconfigurable NoCsabstractThe increasing requirements for bandwidth and quality-of-service motivate the use of parallel interconnect architectures with several degrees of reconfiguration. This paper presents an IP, called Distributed Channel Management (DCM), to extend existing packet-switched NoCs with a reconfigurable point-to-point network seamlessly, i.e., without the need for any modification on the routers. The configuration of the reconfigurable network takes place dynamically and autonomously, so that the topology can be changed at run time. Furthermore, the architecture is scalable due to the autonomous decentralized administration of the links. The Paper reports a thorough experimental analysis of the overhead of the approach at the gate level that considers different network parameters such as flit size and timing constraints. Wolfgang Büter, Christof Osewold, Daniel Gregorek, Alberto García Ortiz |
DATE | 4 |
| 2013 | A Scalable Hardware Implementation of a Best-Effort Scheduler for Multicore ProcessorsabstractThe trend for multicore processor architectures indicates an ongoing increase in computing cores per chip. The resulting challenges demand for a revision of the applicability of existing hardware operating systems. We propose a scalable best-effort task scheduler implemented in hardware, which services a homogeneous multiprocessor architecture. The hardware scheduler realizes a master/slave system to maximize available parallelism. Experimental results show the scalability of the hardware scheduler in terms of performance, area and power. A design pattern to generate a hierarchical communication architecture for task management is prospected. Daniel Gregorek, Christof Osewold, Alberto García Ortiz |
DSD | 3 |
| 2012 | Automatic design of low-power encoders using reversible circuit synthesisabstractThe application of coding strategies is an established methodology to improve the characteristics of on-chip interconnect architectures. Therefore, design methods are required which realize the corresponding encoders and decoders with as small as possible overhead in terms of power and delay. In the past, conventional design methods have been applied for this purpose. Robert Wille, Rolf Drechsler, Christof Osewold, Alberto García Ortiz |
DATE | 4 |
| 2006 | A high-level compact pattern-dependent delay model for high-speed point-to-point interconnectsabstractThis work introduces an extended linear pattern-dependent model for high-level signal delay estimation in high-speed very deep submicron point-to-point interconnects. The proposed model accurately predicts the delay in both inductively and capacitively coupled lines for the complete set of the switching patterns and not only for capacitively coupled lines or worst-case delay as in previous works. We also consider process variations in the formulation of the model and propose a moment-based approach for the inclusion of variations. The accuracy of the model has been assessed by means of extensive experiments. Moreover, we show how the model can be applied at high levels of abstraction in order to explore coding-based alternatives to improve throughput. Tudor Murgan, Massoud Momeni, Alberto García Ortiz, Manfred Glesner |
ICCAD | 3 |
| 2003 | Evaluation and Run-Time Optimization of On-chip Communication Structures in Reconfigurable Architectures
Tudor Murgan, Mihail Petrov, Alberto García Ortiz, Ralf Ludewig, Peter Zipf, Thomas Hollstein, Manfred Glesner, Bernard Ölkrug, Jörg Brakensiek |
FPL | 3 |
| 2003 | Arbitrary function approximation in HDLs with application to the N-body problemabstractA module generator is described that allows for the generation of synthesizable VHDL modules which implement arbitrary functions in fixed point precision using the Symmetric Table Addition Method (STAM). This module generator was interfaced to a high level synthesis tool "fly" which automatically generates fully-pipelined circuits from a Perl-like language. The resulting system was applied to the N-body problem and results are presented. It was found that a function generator module is a very useful addition to a hardware description language. Chun Hok Ho, Kuen Hung Tsoi, Jackson H. C. Yeung, Yuet Ming Lam, Kin-Hong Lee, Philip H. W. Leong, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPT | 9 |
| 2003 | Moment-Based Power Estimation in Very Deep Submicron Technologies
Alberto García Ortiz, Lukusa D. Kabulepa, Tudor Murgan, Manfred Glesner |
ICCAD | 1 |
| 2002 | Estimation of Power Consumption in Encoded Data BusesabstractBecause of the increasing importance of cross coupled capacitances in deep submicron technologies, it is of great interest to extend the existing high-level power estimation techniques by considering the spatial correlation between adjacent lines. This work addresses the modeling and estimation of power dissipation in on-chip buses based on the statistical properties of data sequences. Using the derived models, a power estimation technique is proposed and evaluated for various coding schemes. For different DSP applications, our results depict less than 5% discrepancy with precise bit level estimations. Alberto García Ortiz, Lukusa D. Kabulepa, Manfred Glesner |
DATE | 1 |
| 2002 | Fly - A Modifiable Hardware Compiler
Chun Hok Ho, Philip H. W. Leong, Kuen Hung Tsoi, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPL | 6 |
| 2002 | Efficient estimation of signal transition activity in MAC architecturesabstractBecause of the increasing demand of portable digital systems, it is of great interest to extend the existing high-level power estimation techniques to handle architectures with non linear components, as they appear in relevant practical applications. In this paper we focus on the estimation of the transition activity in MAC structures implementing FIR filters. Based on a divide and conquer approach, an accurate yet efficient estimation procedure is developed. The technique has been evaluated for different synthetic and real data sets. In all cases, our results depict only very slight discrepancies with respect to precise bit level simulations. Alberto García Ortiz, Lukusa D. Kabulepa, Manfred Glesner |
ISLPED | 1 |