VLDB 2026 Research / reviewers in the wild / expert
Christian Weis
dblp:51/10052
· DBLP profile ↗
37ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-4152-0200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 5 first-author · 16 since 2021Software engineering, systems software and programming languages · 18 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi Partner Project: STRATUM, co-creation protocol and advanced smart GUI for a 3D neurosurgery supporting toolabstractSTRATUM is a Horizon Europe multi-partner project developing a clinically validated, real-time 3D decision support tool for brain tumour surgery. The system integrates Hyperspectral Imaging (HSI), AI-based multimodal data fusion, and heterogeneous High-Performance Computing (HPC) architectures combining Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Processing-In-Memory (PIM) technologies. A touchless augmented reality interface facilitates safe and intuitive intraoperative interaction. The distinguishing characteristic of STRATUM is its end-to-end co-designed approach, which integrates advanced computing, state-of-the-art imaging and clinical expertise into a unified Point-of-Care (PoC) platform. Utilising a structured co-creation methodology involving surgeons, engineers, and social scientists, the project ensures usability, safety and regulatory compliance from its early design stages to its clinical validation. The usability of STRATUM will be tested in three hospitals located in different European regions with diverse conditions and regulations. This will allow to collect advice and remarks from surgical staff in a continuous co-creation and co-tuning protocol. Beyond its clinical objectives, STRATUM contributes to the advancement of heterogeneous computing for real-time diagnostics, AI acceleration in critical medical environments and energy-efficient system integration. Furthermore, it delivers open datasets, validated AI pipelines, and performance benchmarks with a view to fostering future research and industrial innovation in digital surgery. The STRATUM project establishes a replicable model for intelligent, human-centred computing integrating microelectronics, AI and medicine.The paper presents an overview of the project in terms of aims, concepts and technologies and the description of the state of the work when approaching the end of the second of the five years planned. Specifically, the outcomes of the steps related to the collaboration with surgeons and medical staff (co-creation process) and the intelligent Graphical User Interface (GUI) development will be described. The latter allows for contactless interaction of the surgeon with several functions that have already been developed in the system. Emanuele Torti, Himar Fabelo, Elisa Marenzi, Maria Luisa Alvarez-Male, Chrysanthi Bairaktari, Beatriz Noriega-Ortega, Raquel León, Santiago Marco, Asaf Badouh, Max Verbers, Javier Santana-Nunez, Yolanda Ramallo-Fariña, Christian Weis, Ana M. Wägner, Eduardo Juárez Martínez, Claudio Rial, Alfonso Lagares, Gustav Burström, Luis Jimenez-Roldan, Teresa Cervero, Miquel Moretó, Giovanni Danese, Svitlana Zinger, Francesca Manni, Miguel A. García-Bello, Lidia García, Jesús Morera, Juan F. Piñeiro, Bernardino Clavo, Francesco Leporati, Gustavo M. Callicó |
DATE | 14 |
| 2025 | Design of a Low-Power 4.3 Gb/s Transceiver Using Pre-computed Lookup TablesabstractHigh-speed memory interfaces require the design of optimised low-power and robust analog circuits. This poses significant challenges in sizing due to the inability of using complex modern transistor models, such as the Berkeley Shortchannel IGFET Model (BSIM), to calculate the sizing of the circuits based on hand-analysis. This forces designers to perform iterative simulations, which is time-intensive and error-prone. To address these issues, this work presents an automated sizing approach using pre-computed lookup tables (LUTs) for a 4.3 Gb/s LPDDR4X transceiver in 12nm FinFET technology. The receiver is designed based on gm/IDsizing methodology where look-up tables are used to compute a set of matrices representing the possible design points based on the circuit topology. The design space is constrained by the biasing level, gain, and bandwidth to find an optimised design point in terms of operation region, speed and power. The driver circuit design is automated based on a new algorithm which computes the estimated ON-resistance from the look up table and finds an optimum sizing which fulfills the required impedance range. The calculated sizing of the proposed design approach was used as an input for spectre simulator to compare the simulation results with the input specifications, showing an error margin of less than 1%. Furthermore, the receiver power consumption was evaluated to be 70% less than the work in literature. The proposed driver topology, which uses low voltage swing terminated logic (LVSTL) and near-ground signaling (NGS), provides a38−120Ωimpedance range at all process, voltage and temperature variations (PVT) for the postlayout results and consumes relatively low power when compared to literature. Hussien Abdo, Jan Lappas, Mohammadreza Esmaeilpour, Christian Weis, Norbert Wehn |
DDECS | 4 |
| 2025 | Analysis and Mitigation of Radiation Effects in SRAM-based Register Files
Surendra Hemaram, Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
ETS | 4 |
| 2024 | Timing Analysis beyond Complementary CMOS Logic StylesabstractWith scaling unabated, device density continues to increase, but power and thermal budgets prevent the full use of all available devices. This leads to the exploration of alternative circuit styles beyond traditional CMOS, especially dynamic data-dependent styles, but the excessive pessimism inherent in conventional static timing analysis tools presents a barrier to adoption. One such circuit family is Pass-Transistor Logic (PTL), which holds significant promise but behaves differently from CMOS in that traditional CMOS-oriented EDA tools cannot produce sufficiently accurate performance estimates. In this work, we revisit timing analysis and its premises and show a significantly improved methodology of a more generalized dynamic timing engine that accurately predicts timing performance for traditional CMOS as well as PTL with an accuracy of 4.0% compared to SPICE and with a run-time comparable to traditional gate-level simulation. The run-time improvement compared with SPICE is four orders of magnitude. Jan Lappas, Mohamed Amine Riahi, Christian Weis, Norbert Wehn, Sani R. Nassif |
ASPDAC | 3 |
| 2024 | 3D Decision Support Tool for Brain Tumour Surgery: The STRATUM ProjectabstractIntegrated digital diagnostics can support complex surgical procedures in many anatomical sites, brain tumour surgery being the most complex. STRATUM is a 5-year Horizon Europe funded project with the goal of developing an innovative 3D decision support tool for brain tumour surgeries, based on real-time multimodal data processing using artificial intelligence algorithms. The proposed tool is envisioned as an energy-efficient Point-of-Care computing system to be integrated within neurosurgical workflows to aid surgeons to make informed, efficient, and accurate decisions during surgical procedures. The expected long-term impact of STRATUM is to reduce the duration of surgical procedures, thus decreasing patients' risks, but also optimising the resources of European health care systems. Himar Fabelo, Raquel León, Emanuele Torti, Santiago Marco, Max Verbers, Yann Falevoz, Yolanda Ramallo-Fariña, Christian Weis, Ana M. Wägner, Eduardo Juárez Martínez, Claudio Rial, Alfonso Lagares, Gustav Burström, Francesco Leporati, Elisa Marenzi, Teresa Cervero, Miquel Moretó, Giovanni Danese, Svitlana Zinger, Francesca Manni, Maria Luisa Alvarez-Male, Jesús Morera, Bernardino Clavo, Gustavo M. Callicó |
DSD | 8 |
| 2024 | Do Radiation and Aging Impact DVFS? TCAD-based Analysis on 22 nm FDSOI Latches
Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif |
IOLTS | 2 |
| 2024 | Testing for aging in advanced SRAM: From front end of the line transistors to back end of the line interconnectsabstractThe long-term reliability of Static Random Access Memory (SRAM) is crucial for safety-critical applications, such as those in the automotive industry. In the front-end-of-line (FEoL), the transistor elements are susceptible to negative bias temperature instability (NBTI), while in the back-end-of-line (BEoL) the interconnects are susceptible to electromigration (EM), especially in scaled technology nodes. To meet safety-critical standards, it is essential to investigate the combined aging mechanisms within the SRAM array and to develop effective testing methodologies during the operational lifetime of the system. Such methodologies are also crucial for enabling the early detection of in-field failures. In this paper, a precise aging model is presented that extends the Technology Computer-Aided Design (TCAD) transistor model with a detailed NBTI model and includes physical modeling for EM. This approach provides insights into the combined effects of NBTI and EM on the degradation of SRAM writability, considering the entire SRAM subarray, including the bit-cell array and peripheral circuits in Fin Field-Effect Transistors (FinFET) technology. Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
ITC | 3 |
| 2024 | A Low-Power Linear Phase Interpolation-Based Delay Line in 12nm FinFET TechnologyabstractA novel low-power high-linear phase interpolation-based delay line in 12nm FinFET technology is detailed in this paper. The proposed delay line exhibits 50% improvement in terms of power consumption compared to the previous work. In addition, the presented architecture to the best of our knowledge is the most efficient delay line for fine tuning in advanced technology nodes due to the low complexity and complete controllability over resolution, delay range and target frequency. The analysis in this paper indicates that the input slew rate plays an indispensable role in the linearity of the delay line. Consequently, two identical resistors are added to the input of the phase interpolator unit to decrease the slew rate. This approach significantly improves the linearity over a wide frequency range. The proposed delay line dissipates 0.56 mW from a 0.8 V supply voltage and 5 GHz operating frequency. Mohammadreza Esmaeilpour, Jan Lappas, Christian Weis, Norbert Wehn |
VLSI-SoC | 3 |
| 2024 | Addressing the Combined Effect of Transistor and Interconnect Aging in SRAM towards Silicon Lifecycle ManagementabstractThe long-term reliability of the Static Random Access Memory (SRAM) module, as an important component of computing architectures, is crucial for safety-critical applications such as automotive. In the front end of the line (FEoL), the transistor elements are vulnerable to negative bias temperature instability (NBTI), while the back end of the line (BEoL) interconnect is prone to electromigration (EM). Complying with safety-critical standards as part of silicon lifecycle management (SLM) infrastructure requires an understanding of the combined aging mechanisms of transistors and interconnects in SRAM. Moreover, a precise aging model is a prerequisite for effective aging testing and mitigation strategies. For this aim, we augment the Technology Computer-Aided Design (TCAD) transistor model with a detailed NBTI model at the FEoL, and use measurement-calibrated physical modeling of EM at the BEoL, to create an integrated analysis that can provide deeper insights into the individual and combined effects of NBTI and EM for SRAM operation. Our findings reveal the mutual acceleration of delay faults and hard stuck-at faults caused by NBTI and EM in SRAM, offering a precise methodology for estimating the time to failure under these conditions. Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
VTS | 3 |
| 2023 | ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous DrivingabstractModern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned. Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001 |
DATE | 10 |
| 2023 | A Learning-Based Approach for Single Event Transient Analysis in Pass Transistor LogicabstractPass transistor logic (PTL) has emerged recently in advanced high-speed optical communication system due to its higher speed and lower power consumption compared to traditional CMOS logic. However, the sensitivity to radiation-induced soft errors of PTL implementations is significant different from CMOS circuitry, which emphasizes the need for understanding the mechanism of soft error propagation in PTL. Due to the non-conventional logic structure in PTL, previous approaches of pulse width modelling in CMOS logic are no more applicable since they are not always measurable. Hence, in this paper, we propose a learning-based structural regression modeling approach to explore the soft error propagation mechanism in PTL at transistor level. Our models can be easily mapped onto higher level to analyze soft error propagation in any complex PTL designs. The experimental results on a 4-bit ripple carry adder demonstrate that our models can achieve high accuracy compared with SPICE simulation. Zhihang Wu, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori |
IOLTS | 3 |
| 2022 | Revisiting Pass-Transistor Logic Styles in a 12nm FinFET Technology NodeabstractWith the slow-down of Moore's law and the increasing requirements on energy efficiency, alternative logic styles compared to complementary static CMOS have to be revisited for digital circuit implementations. Pass Transistor Logic (PTL) gained much attention in the '90s, however, only a limited number of recent investigations and publications regarding PTL exist that use advanced technology nodes. This paper compares key performance metrics of 22 different PTL based 1-bit full adder designs to a complementary static CMOS logic reference, using a recent 12nm FinFET technology. The figures of merit are the propagation delay, the energy consumption, and the energy-delay-product (EDP). Our investigations show that PTL based adder circuits can have an up to 49% decreased delay and a 48% and 63% reduced energy consumption and EDP, respectively, compared to a state-of-the-art complementary CMOS logic reference. In addition, we analyzed the impact of PVT variations on the delay for selected PTL full adder designs. Jan Lappas, André Lucas Chinazzo, Christian Weis, Chenyang Xia, Zhihang Wu, Leibin Ni, Norbert Wehn |
DATE | 3 |
| 2022 | Machine learning based soft error rate estimation of pass transistor logic in high-speed communicationabstractRecent advanced high-speed communication systems, such as optical systems, require highest reliability at lowest possible power consumption. Thus, Pass Transistor Logic (PTL) is gaining lots of interest in these communication systems due to its power saving potential compared to traditional CMOS logic. However, due to the non-conventional logic structure, its susceptibility to radiation-induced soft errors is different from CMOS circuitry. Due to the unique generation and propagation of Single Event Transients (SETs) in PTL, different approaches for PTL soft error rate (SER) estimation are required. In this paper we propose a machine learning (ML) approach for SET propagation in PTL logic. Multi-layer feed-forward neural network together with support vector classifier (SVC) are used to build the SET pulse width and pulse amplitude models. Bayesian optimization using Gaussian Processes is utilized to tune the hyperparameters of neural network. The experimental results on full adder (FA), which is the key component in many large cirucits such as ALU, and comparison with Monte Carlo (MC) spectre simulations confirm the accuracy and speed of the proposed method. Jan Lappas, André Lucas Chinazzo, Christian Weis, Zhihang Wu, Leibin Ni, Norbert Wehn, Mehdi Baradaran Tahoori |
ETS | 4 |
| 2022 | Optimization of DRAM based PIM Architecture for Energy-Efficient Deep Neural Network TrainingabstractDeep Neural Network (DNN) training consumes high-energy. On the other hand, DNNs deployed on edge devices demand very high-energy efficiency. In this context, Processing-in-Memory (PIM) is an emerging compute paradigm that bridges the memory-computation gap to improve the energy-efficiency. DRAMs are one such memory type employed for designing energy-efficient PIM architectures for DNN training. One of the major issues of DRAM-PIM architectures designed for DNN training is the high number of internal data accesses within a bank between the memory arrays and the PIM computation units (e.g. 51% more than inference). These internal data accesses in the state-of-the-art DRAM PIM architectures consume very high energy compared to computation units. Hence, it is important to reduce the internal data access energy within the DRAM bank for further improving the energy efficiency of DRAMPIM architectures. We present three novel optimizations that together reduce the internal data access energy up to 81.54%. Our first optimization modifies the bank data access circuit to enable partial accesses of data instead of the conventional fixed granularity accesses, thereby exploiting the available sparsity during training. The second optimization is to have a dedicated low-energy region within the DRAM bank that has low capacitive load of global wires and shorter data movement. Finally, we propose a 12-bit high dynamic range floating-point format called TinyFloat that reduces the total number of data access energy by 20% compared to IEEE 754 half and single precision. Chirag Sudarshan, Mohammad Hassani Sadi, Christian Weis, Norbert Wehn |
ISCAS | 3 |
| 2022 | FeFET versus DRAM based PIM Architectures: A Comparative StudyabstractThe throughput and energy efficiency of compute-centric architectures for memory intensive Deep Neural Networks (DNN) applications are limited by memory bound issues like high data-access energy, long latencies, and limited bandwidth. Processing-in-Memory (PIM) is a very promising approach to address these challenges and bridge the memory-computation gap. PIM places computational logic inside the memory to exploit minimum data movement and massive internal data parallelism. There are currently two PIM trends: 1) Use of emerging non-volatile memories to perform highly parallel analog computation of MAC operations and implicit storage of weights within the memory arrays, and 2) exploiting mature memory technologies that are enhanced by additional logic to enable efficient computation of MAC operations near the memory arrays. In this paper, we will compare both trends from an architectural perspective. Our study mainly emphasizes on FeFET memories (an emerging memory candidate) and DRAM memories (a mature memory candidate). We will highlight the major architectural constraints of these memory candidates that impact the PIM designs and their overall performance. Finally, we will assess feasible choice of candidate for different computations or DNN task types. Chirag Sudarshan, Taha Soliman, Thomas Kämpfe, Christian Weis, Norbert Wehn |
VLSI-SoC | 4 |
| 2021 | A Novel DRAM-Based Process-in-Memory Architecture and its Implementation for CNNsabstractProcessing-in-Memory (PIM) is an emerging approach to bridge the memory-computation gap. One of the key challenges of PIM architectures in the scope of neural network inference is the deployment of traditional area-intensive arithmetic multipliers in memory technology, especially for DRAM-based PIM architectures. Hence, existing DRAM PIM architectures are either confined to binary networks or exploit the analog property of the sub-array bitlines to perform bulk bit-wise logic operations. The former reduces the accuracy of predictions, i.e. Quality-of-results, while the latter increases overall latency and power consumption. Chirag Sudarshan, Taha Soliman, Cecilia De la Parra, Christian Weis, Leonardo Ecco, Matthias Jung 0001, Norbert Wehn, Andre Guntoro |
ASP-DAC | 4 |
| 2020 | Access-Aware Per-Bank DRAM Refresh for Reduced DRAM Refresh OverheadabstractThe performance and energy penalties of DRAM refresh have increased in successive generations of higher capacity DRAM devices. This trend is likely to continue in future systems where the internal DRAM refresh cycle is opaque to the memory controller and the memory device oblivious of context. This paper presents Access-Aware Per-bank DRAM Refresh, a refresh control method that mitigates the negative impacts of refreshes, and its memory controller architecture. Novel capabilities are introduced in the memory controller. An access-aware refresh control unit analyses the short-term history of memory accesses translating row activations into refresh masks. Refresh masks are used either to skip rows, shortening the refresh cycle, or to completely omit refresh operations. The set of DRAM commands is extended with two new per-bank refresh commands that provide the memory controller not only an ability to omit refreshes, but also a context-rich fine-grained control of refresh operations. A proof of concept model of our architecture is implemented in a virtual platform where a set of applications is used to exercise the memory subsystem. Evaluations show that for the workloads considered the proposed architecture and refresh control method improve, either by reducing the latency or by completely omitting, up to 19% of the refresh operations. Éder Zulian, Christian Weis, Norbert Wehn |
ISCAS | 2 |
| 2019 | An In-DRAM Neural Network Processing EngineabstractMany advanced neural network inference engines are bounded by the available memory bandwidth. The conventional approach to address this issue is to employ high bandwidth memory devices or to adapt data compression techniques (reduced precision, sparse weight matrices). Alternatively, an emerging approach to bridge the memory-computation gap and to exploit extreme data parallelism is Processing in Memory (PIM). The close proximity of the computation units to the memory cells reduces the amount of external data transactions and it increases the overall energy efficiency of the memory system. In this work, we present a novel PIM based Binary Weighted Network (BWN) inference accelerator design that is inline with the commodity Dynamic Random Access Memory (DRAM) design and process. In order to exploit data parallelism and minimize energy, the proposed architecture integrates the basic BWN computation units at the output of the Primary Sense Amplifiers (PSAs) and the rest of the substantial logic near the Secondary Sense Amplifiers (SSAs). The power and area values are obtained at sub-array (SA) level using exhaustive circuit level simulations and full-custom layout. The proposed architecture results in an area overhead of 25 % compared to a commodity 8 Gb DRAM and delivers a throughput of 63.59 FPS (Frames per Second) for AlexNet. We also demonstrate that our architecture is extremely energy efficient, 7.25× higher FPS/W, as compared to previous works. Chirag Sudarshan, Jan Lappas, Muhammad Mohsin Ghaffar, Vladimir Rybalkin, Christian Weis, Matthias Jung 0001, Norbert Wehn |
ISCAS | 5 |
| 2018 | Improving the error behavior of DRAM by exploiting its Z-channel propertyabstractIn this paper, we present a new communication theoretic channel model for Dynamic Random Access Memory (DRAM) retention errors, that relies on the fully asymmetric retention error behavior of DRAM cells. This new model shows that the traditional approach is over pessimistic and we confirm this with real measurements of DDR3 and DDR4 DRAM devices. Together with an exploitation of the vendor specific true- and anti-cell structure, a low complexity bit-flipping approach is presented, that can largely increase DRAM's reliability with minimum overhead. Kira Kraft, Chirag Sudarshan, Deepak M. Mathew, Christian Weis, Norbert Wehn, Matthias Jung 0001 |
DATE | 4 |
| 2018 | An analysis on retention error behavior and power consumption of recent DDR4 DRAMsabstractDRAM technology is scaling aggressively that results in high leakage power, worse data retention time behavior, and large process variations. Due to these process variations, vendors provide large guard bands on various DRAM currents and timing specifications that are over pessimistic. Detailed knowledge on the DRAM retention behavior and currents for the average case allow to improve memory system performance and energy efficiency of specific applications by moving away from worst case behavior. In this paper, we present an advanced measurement platform to investigate off-the-shelf DDR4 DRAMs' retention behavior, and to precisely measure various DRAM currents (IDDs and IPPs) at a wide range of operating temperatures. Error Checking and Correction (ECC) schemes are popular in correcting randomly scattered single bit errors. Since retention failures also occur randomly, ECCs can be used to improve DRAM retention behavior. Therefore, for the first time, we show the influence of ECC on the retention behavior of recent DDR4 DRAMs, and how it varies across various DRAM architectures considering detailed structure of the DRAM (true-cell devices/mixed-cell devices). Deepak M. Mathew, Martin Schultheis, Carl Christian Rheinländer, Chirag Sudarshan, Christian Weis, Norbert Wehn, Matthias Jung 0001 |
DATE | 5 |
| 2018 | The Role of Memories in Transprecision ComputingabstractComputing paradigms largely evolved over the last decades mainly driven by continuously increasing performance requirements, energy efficiency and power density challenges. Heterogeneous highly parallel architectures enhanced with dedicated accelerators tuned to specific applications, near-threshold computing, and recently approximate computing are examples of these new approaches. In this context the memory part was relatively untouched. However, memories play a central role in any computing system, are a major source of power consumption, and limit in many applications the overall compute performance. In this paper, we focus mainly on Dynamic Random Access Memories (DRAMs), which are today's most prominent external memories. In transprecision computing we address the DRAM memory challenge by several new approaches that are strongly related to the new techniques known on the compute side. In particular, these are the concept of approximate DRAM, advanced power-down modes and the integration of application knowledge into the memory system. Christian Weis, Matthias Jung 0001, Éder Zulian, Chirag Sudarshan, Deepak M. Mathew, Norbert Wehn |
ISCAS | 1 |
| 2017 | An advanced embedded architecture for connected component analysis in industrial applicationsabstractIn recent years, connected component analysis (CCA) has become one of the vital image/video processing algorithms due to its wide-range applicability in the field of computer vision. Numerous applications such as pattern recognition, object detection and image segmentation involve connected component analysis. In the context of camera-based inspection systems, CCA plays an important role for quality assurance. State-of-the-art hardware architectures offer high performance implementations of CCA using field programmable gate arrays (FPGAs). However, due to their high memory-demand, most of these implementations inhibit a large resource utilization. In this paper, we propose a hybrid software-hardware architecture of CCA for an industrial application using Xilinx Zynq-7000 All Programmable System on Chip (SoC). By offloading the most resource consuming part of the algorithm to the embedded CPU, we achieved high performance, while reducing the required resources on the FPGA. Our proposed architecture saves more than 30% of on-chip memory (Block RAMs) compared to state-of-the-art hardware architectures without affecting the throughput. Furthermore, due to the embedded CPU, our system provides a versatile and highly flexible feature extraction at run-time without the necessity to reconfigure the FPGA. Menbere Tekleyohannes, MohammadSadegh Sadri, Christian Weis, Norbert Wehn, Martin Klein 0005, Michael Siegrist |
DATE | 3 |
| 2016 | Efficient reliability management in SoCs - an approximate DRAM perspectiveabstractIn today's computing systems Dynamic Random Access Memories (DRAMs) have a large influence on performance and contribute significantly to the total power consumption. Thus, recent research activities bring the idea of approximate DRAM into focus to save power and improve performance by lowering the refresh rate or disabling refresh completely. Hence, fast and accurate models are required for a thoroughly exploration of approximate DRAM for error resilient applications. In this paper we present a holistic simulation environment for investigations on approximate DRAM and show the impact on error resilient applications. Matthias Jung 0001, Deepak M. Mathew, Christian Weis, Norbert Wehn |
ASP-DAC | 3 |
| 2016 | Invited - Approximate computing with partially unreliable dynamic random access memory - approximate DRAMabstractIn the context of approximate computing, Approximate Dynamic Random Access Memory (ADRAM) enables the tradeoff between energy efficiency, performance and reliability. The inherent error resilience of applications allows sacrificing data storage robustness and stability by lowering the refresh rate or disabling refresh in DRAMs completely. Consequently, it is important to know exactly the statistical DRAM behavior with respect to retention time, process variation and temperature to manage this trade-off and thereby deliberately exploiting the error resilience of different target applications. Matthias Jung 0001, Deepak M. Mathew, Christian Weis, Norbert Wehn |
DAC | 3 |
| 2016 | Error resilience and energy efficiency: An LDPC decoder design study
Philipp Schläfer, Chu-Hsiang Huang, Clayton Schoeny, Christian Weis, Yao Li 0007, Norbert Wehn, Lara Dolecek |
DATE | 4 |
| 2015 | Retention time measurements and modelling of bit error rates of WIDE I/O DRAM in MPSoCs
Christian Weis, Matthias Jung 0001, Peter Ehses, Cristiano Santos, Pascal Vivet, Sven Goossens, Martijn Koedam, Norbert Wehn |
DATE | 1 |
| 2014 | Exploiting expendable process-margins in DRAMs for run-time performance optimizationabstractManufacturing-time process (P) variations and runtime voltage (V) and temperature (T) variations can affect a DRAM's performance severely. To counter these effects, DRAM vendors provide substantial design-time PVT timing margins to guarantee correct DRAM functionality under worst-case operating conditions. Unfortunately, with technology scaling these timing margins have become large and very pessimistic for a majority of the manufactured DRAMs. While run-time variations are specific to operating conditions and as a result, their margins difficult to optimize, process variations are manufacturing-time effects and excessive process-margins can be reduced at run-time, on a per-device basis, if properly identified. In this paper, we propose a generic post-manufacturing performance characterization methodology for DRAMs that identifies this excess in process-margins for any given DRAM device at runtime, while retaining the requisite margins for voltage (noise) and temperature variations. By doing so, the methodology ascertains the actual impact of process-variations on the particular DRAM device and optimizes its access latencies (timings), thereby improving its overall performance. We evaluate this methodology on 48 DDR3 devices (from 12 DIMMs) and verify the derived timings under worst-case operating conditions, showing up to 33.3% and 25.9% reduction in DRAM read and write latencies, respectively. Karthik Chandrasekar 0001, Sven Goossens, Christian Weis, Martijn Koedam, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 3 |
| 2014 | Hybrid memory architecture for voltage scaling in ultra-low power multi-core biomedical processorsabstractTechnology scaling enables today the design of sensor-based ultra-low cost chips well suited for emerging applications such as wireless body sensor networks, urban life and environment monitoring. Energy consumption is the key limiting factor of this up-coming revolution and memories are often the energy bottleneck mainly due to leakage power. This paper proposes an ultra-low power multi-core architecture targeting eHealth monitoring systems, where applications involve collection of sequences of slow biomedical signals and highly parallel computations at very low voltage. We propose a hybrid memory architecture that combines 6T-SRAM and 8T-SRAM operating in the same voltage domain and capable of dispatching at high voltage a normal operation and at low voltage a fully reliable small memory partition (8T) while the rest of the memory (6T) is state-retentive. Our architecture offers significant energy savings with a low area overhead in typical eHealth Compressed Sensing-based applications. Daniele Bortolotti, Andrea Bartolini, Christian Weis, Davide Rossi 0001, Luca Benini |
DATE | 3 |
| 2014 | Energy optimization in 3D MPSoCs with Wide-I/O DRAM using temperature variation aware bank-wise refreshabstractHeterogeneous 3D integrated systems with Wide-I/O DRAMs are a promising solution to squeeze more functionality and storage bits into an ever decreasing volume. Unfortunately, with 3D stacking, the challenges of high power densities and thermal dissipation are exacerbated. We improve DRAM refresh power by considering the lateral and vertical temperature variations in the 3D structure and adapting the per-DRAM-bank refresh period accordingly. In order to provide proof of our concepts we develop an advanced virtual platform which models the performance, power, and thermal behavior of a 3D-integrated MPSoC with Wide-I/O DRAMs in detail. On this platform we run the Android OS with real-world benchmarks to quantify the advantages of our ideas. We show improvements of 16% in DRAM refresh power due to temperature variation aware bank-wise refresh. Furthermore, two solutions are investigated to speedup system simulations: (1) Adaptive tuning of sampling intervals based on the estimated chip thermal profile, which results in speedups of 2X. (2) Hardware acceleration of thermal simulations using the Maxeler engine, which shows possible speedups of 12X. MohammadSadegh Sadri, Matthias Jung 0001, Christian Weis, Norbert Wehn, Luca Benini |
DATE | 3 |
| 2014 | Optimized active and power-down mode refresh control in 3D-DRAMsabstract3D stacked systems with Wide-I/O DRAMs are the future density optimized mobile computing platforms. Unfortunately, with 3D integration, the power densities and thermal dissipation are increased dramatically. In this paper, we investigate the effectiveness of power-down mode policies (using precharge power down, active power-down and self-refresh) and bank-wise refresh in active mode. We run real-life benchmarks to quantify the impact of each power-down mode setting. We derive a power-down mode policy which shows up to 10% energy reduction in high activity periods and up to 13% in idle phases. Further, we improve DRAM refresh power by considering the lateral and vertical temperature variations in the 3D structure and adapting the per-DRAM-bank refresh period accordingly. To achieve this, a per DRAM array hotspot detector, designed with DRAM cells and circuits, is used to acquire temperature and refresh information directly from the DRAM array. We show 16% improvements in DRAM refresh power due to hotspot detectors inside the DRAM enabling temperature variation aware bank-wise refresh. For all the above mentioned investigations a detailed DRAM controller model with accurate functionality, timing, and power estimation in SystemC TLM-2.0 (Transaction Level Modeling) and a highly sophisticated virtual hardware platform are mandatory to achieve a through analysis. Matthias Jung 0001, Christian Weis, Norbert Wehn, MohammadSadegh Sadri, Luca Benini |
VLSI-SoC | 2 |
| 2013 | Towards variation-aware system-level power estimation of DRAMs: an empirical approachabstractDRAM vendors provide pessimistic current measures in memory datasheets to account for worst-case impact of process variations and to improve their production yield, leading to unrealistic power consumption estimates. In this paper, we first demonstrate the possible effects of process variations on DRAM performance and power consumption by performing Monte-Carlo simulations on a detailed DRAM cross-section. We then propose a methodology to empirically determine the actual impact for any given DRAM memory by assessing its performance characteristics during the DRAM calibration phase at system boot-time, thereby enabling its optimal use at run-time. We further employ our analysis on Micron's 2Gb DDR3-1600-x16 memory and show considerable over-estimation in the datasheet measures and the energy estimates (up to 28%), by using realistic current measures for a set of MediaBench applications. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DAC | 2 |
| 2013 | System and circuit level power modeling of energy-efficient 3D-stacked wide I/O DRAMsabstractJEDEC recently introduced its new standard for 3D-stacked Wide I/O DRAM memories, which defines their architecture, design, features and timing behavior. With improved performance/power trade-offs over previous generation DRAMs, Wide I/O DRAMs provide an extremely energy-efficient green memory solution required for next-generation embedded and high-performance computing systems. With both industry and academia pushing to evaluate and employ these highly anticipated memories, there is an urgent need for an accurate power model targeting Wide I/O DRAMs that enables their efficient integration and energy management in DRAM stacked SoC architectures. In this paper, we present the first system-level power model of 3D-stacked Wide I/O DRAM memories that is almost as accurate as detailed circuit-level power models of 3D-DRAMs. To verify its accuracy, we experimentally compare its power and energy estimates for different memory workloads and operations against those of a circuit-level 3D-DRAM power model and show less than 2% difference between the two sets of estimates. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 2 |
| 2013 | Exploration and Optimization of 3-D Integrated DRAM SubsystemsabstractEnergy efficiency is the major optimization criterion for systems-on-chip (SoCs) for mobile devices (smartphones and tablets). Through silicon via (TSV) technology enables 3-D integration of dies and the heterogeneous stacking of multiple memory or logic layers, allowing increased bandwidth and lower energy consumption of the memory interface compared to traditional approaches. In this paper, we explore the 3-D-DRAM architecture design space. The result is an optimized 2 Gb 3-D-DRAM, which shows a 83% lower energy/bit than a 2 Gb device. Furthermore, we propose a highly energy-efficient DRAM subsystem for next-generation 3-D-integrated SoCs, consisting of a SDR/DDR 3-D-DRAM controller and an attached 3-D-DRAM cube with fine-grained access and a flexible (WIDE-IO) interface. We assess the energy efficiency using a synthesizable model of the SDR/DDR 3-D-DRAM channel controller (CC) as well as functional models of the 3-D-stacked DRAM, including an accurate power estimation engine. We also investigate different DRAM families (WIDE IO SDR/DDR, LPDDR, and LPDDR2) and densities from 256 Mb to 4 Gb per channel. The implementation results of the proposed 3-D-DRAM subsystem show that energy optimized accesses to the 3-D-DRAM enable up to 50% energy savings compared to standard accesses. To the best of our knowledge this is the first design space exploration for 3-D-stacked DRAM considering different technologies based on real-world physical data and the first design of a 3-D-DRAM CC and 3-D-DRAM model featuring co-optimization of memory and controller architecture. Christian Weis, Igor Loi, Luca Benini, Norbert Wehn |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | DRAM selection and configuration for real-time mobile systemsabstractThe performance and power consumption of mobile DRAMs (LPDDRs) depend on the configuration of system-level parameters, such as operating frequency, interface width, request size, and memory map. In mobile systems running both real-time and non-real-time applications, the memory configuration must satisfy bandwidth requirements of real-time applications, meet the power consumption budget, and offer the best average-case execution time to the non-real-time applications. There is currently no well-defined methodology for selecting a suitable memory configuration for real-time mobile systems. The worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across generations have furthermore not been investigated. This paper has two main contributions. 1) We analyze the worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across three generations: LPDDR, LPDDR2 and Wide-IO-based 3D-stacked DRAM. 2) Based on our analysis, we propose a methodology for selecting memory configurations in real-time mobile systems.We show that LPDDR (32-bit IO), LPDDR2 (32-bit IO) and 3D-DRAM (128-bit IO) provide worst-case bandwidth up to 0.75 GB/s, 1.6 GB/s and 3.1 GB/s, respectively. We furthermore show for an H.263 decoder that LPDDR2 and 3D-DRAM reduce power consumption with up to 25% and 67%, respectively, compared to LPDDR, and reduce the execution time with up to 18% and 25%. Manil Dev Gomony, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 2 |
| 2012 | An energy efficient DRAM subsystem for 3D integrated SoCsabstractEnergy efficiency is the key driver for the design optimization of System-on-Chips for mobile terminals (smartphones and tablets). 3D integration of heterogeneous dies based on TSV (through silicon via) technology enables stacking of multiple memory or logic layers and has the advantage of higher bandwidth at lower energy consumption for the memory interface. In this work we propose a highly energy efficient DRAM subsystem for next-generation 3D integrated SoCs, which will consist of a SDR/DDR 3D-DRAM controller and an attached 3D-DRAM cube with a fine-grained access and a very flexible (WIDE-IO) interface. We implemented a synthesizable model of the SDR/DDR 3D-DRAM channel controller and a functional model of the 3D-stacked DRAM which embeds an accurate power estimation engine. We investigated different DRAM families (WIDE IO DDR/SDR, LPDDR and LPDDR2) and densities that range from 256Mb to 4Gb per channel. The implementation results of the proposed 3D-DRAM subsystem show that energy optimized accesses to the 3D-DRAM enable an overall average of 37% power savings as compared to standard accesses. To the best of our knowledge this is the first design of a 3D-DRAM channel controller and 3D-DRAM model featuring co-optimization of memory and controller architecture. Christian Weis, Igor Loi, Luca Benini, Norbert Wehn |
DATE | 1 |
| 2011 | Design space exploration for 3D-stacked DRAMsabstract3D integration based on TSV (through silicon via) technology enables stacking of multiple memory layers and has the advantage of higher bandwidth at lower energy consumption for the memory interface. As in mobile applications energy efficiency is key, 3D integration is especially here a strategic technology. In this paper we focus on the design space exploration of 3D-stacked DRAMs with respect to performance, energy and area efficiency for densities from 256Mbit to 4Gbit per 3D-DRAM channel. We investigate four different technology nodes from 75nm down to 45nm and show the optimal design point for the currently most common commodity DRAM density of 1Gbit. Multiple channels can be combined for main memory sizes of up to 32GB. We present a functional SystemC model for the 3D-stacked DRAM which is coupled with a SDR/DDR 3D-DRAM channel controller. Parameters for this model were derived from detailed circuit level simulations. The exploration demonstrates that an optimized 1Gbit 3D-DRAM stack is 15× more energy efficient compared to a commodity Low-Power DDR SDRAM part without IO drivers and pads. To the best of our knowledge this is the first design space exploration for 3D-stacked DRAM considering different technologies and real world physical commodity DRAM data. Christian Weis, Norbert Wehn, Igor Loi, Luca Benini |
DATE | 1 |
| 2011 | Bringing C++ productivity to VHDL world: From language definition to a case study
Ivan Shcherbakov, Christian Weis, Norbert Wehn |
FDL | 2 |