Mehdi Saligane

dblp:160/5074 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-0237-4690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FLARE: Finetuning ReLU With FIRE for Efficient Long-Context Inference
abstract
Deploying large language models (LLMs) on resource-constrained edge devices, such as mobile phones or IoT devices, is highly desirable for enabling secure, personalized on-device AI. However, there are significant challenges due to these models’ high computational and memory demands. A key bottleneck lies in the Transformer’s attention block, especially when handling long contexts. Techniques like model architectures with Rectified Linear Unit (ReLU) activations for Softmax and FIRE positional encoding (a resource-efficient, automatic context-length-scaling alternative to Rotary Positional Embedding (RoPE)) have each independently shown promise in reducing the computational complexity of the attention block, but the proper alchemy for combining their benefits remains underexplored. In this paper, we show a method for combining FIRE and ReLU that maintains low-validation loss at long contexts. We also introduce FLARE, a new algorithm that further improves efficiency by removing operations from the learned relative position encoding in FIRE. Our approach leads to faster inference on long sequences, robust generalization to varying context lengths, and lower validation loss compared to baseline models. FLARE achieves a significant reduction in power and area consumption. On custom hardware, it achieves a 6× higher operating frequency than Softmax, while occupying 57× less silicon area (measured under different throughput settings) and consuming 600× less energy. Our results indicate that FLARE represents a significant step towards deploying powerful LLMs efficiently on resource-limited devices. Code and hardware designs are publicly available at: https://github.com/ReaLLMASIC/nanoGPT.
Michael Moffatt, Junyi Luo, Xinting Jiang, Guanchen Tao, Shiwei Liu 0002, Kauna Lei, Gregory Kielian, Mehdi Saligane
DATE10
2026 Pinball: A Cryogenic Predecoder for Quantum Error Correction Decoding Under Circuit-Level Noise
abstract
Scaling fault tolerant quantum computers, especially cryogenic systems based on the surface code, to millions of qubits is very challenging due to poorly-scaling data processing and power consumption overheads. One key challenge is the design of decoders for real-time quantum error correction (QEC), which demands high data rates for error processing; this is particularly apparent in systems with cryogenic qubits and room temperature (RT) decoders. In response, cryogenic predecoding using lightweight logic has been proposed to handle common, sparse errors within the cryogenic domain. However, prior work only accounts for a subset of the error sources present in real-world quantum systems with limited accuracy, often degrading performance below a useful level in practical scenarios. Furthermore, prior reliance on SFQ logic precludes detailed architecture-technology co-optimization. To address these shortcomings, this paper introduces Pinball<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>1Source code available at: https://github.com/aknapen/Pinball, a comprehensive design in cryogenic CMOS of a QEC predecoder for the surface code, tailored to realistic, circuit-level noise. By accounting for error generation and propagation through QEC circuits, our design achieves higher predecoding accuracy, outperforming logical error rates of the current state-of-theart cryogenic predecoder by nearly six orders of magnitude. Remarkably, despite operating under much stricter power and area constraints, Pinball also reduces logical error rates by 32.58× and 5×, respectively, compared to the state-of-the-art RT predecoder and an RT ensemble configuration. By increasing cryogenic coverage, we also reduce syndrome bandwidth up to 3780.72×. Through co-design with 4 K -characterized 22 nm FDSOI technology, we achieve a peak power consumption under 0.56 mW. Voltage/frequency scaling and body biasing enable 22.2× lower typical power consumption, yielding up to 67.4× total energy savings. Assuming a 4 K power budget of 1.5 W, our predecoder can support up to <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{2, 6 6 8}$</tex> logical qubits at <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$d=21$</tex>.
Alexander Knapen, Guanchen Tao, Jacob Mack, Tomas Bruno, Mehdi Saligane, Dennis Sylvester, Qirui Zhang 0001, Gokul Subramanian Ravi
HPCA5
2026 Toward Generative Silicon: The Next Frontier in Open-Source and AI-Driven Analog Design
Mehdi Saligane, Anhang Li, Boris Murmann
ISCAS1
2026 Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh Environments
abstract
In harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM.
Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto
IEEE Trans. Computers11
2024 Reinforcement Learning-Enhanced Cloud-Based Open Source Analog Circuit Generator for Standard and Cryogenic Temperatures in 130-nm and 180-nm OpenPDKs
abstract
This work introduces an open-source, Process Technology-agnostic framework for hierarchical circuit netlist, layout, and Reinforcement Learning (RL) optimization. The layout, netlist, and optimization python API is fully modular and publicly installable. It features a bottom-up hierarchical construction, which allows for complete design reuse across provided PDKs. The modular hierarchy also facilitates parallel circuit design iterations on cloud platforms. To illustrate its capabilities, a two-stage OpAmp with a 5T first-stage, common-source second-stage, and miller compensation is implemented. We instantiate the OpAmp in two different open-source process design kits (OpenPDKs) using both room-temperature models and cryogenic (4K) models. With a human designed version as the baseline, we leveraged the parameterization capabilities of the framework and applied the RL optimizer to adapt to the power consumption limits suitable for cryogenic applications while maintaining gain and bandwidth performance. Using the modular RL optimization framework we achieve a 6x reduction in power consumption compared to manually designed circuits while maintaining gain to within 2%.
Ali Hammoud, Anhang Li, Ayushman Tripathi, Harsh Khandeparkar, Ryan Wans, Gregory Kielian, Boris Murmann, Dennis Sylvester, Mehdi Saligane
ICCAD10
2024 ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
abstract
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvement, achieving real-time LLM inference on silicon remains challenging due to the extensive use of Softmax in self-attention. In addition to the non-linearity, the low arithmetic intensity significantly limits processing parallelism, especially when working with longer contexts. To address this challenge, we propose Constant Softmax (ConSmax), a software-hardware co-design that serves as an efficient alternative to Softmax. ConSmax utilizes differentiable normalization parameters to eliminate the need for maximum searching and denominator summation in Softmax. This approach enables extensive parallelization while still executing the essential functions of Softmax. Moreover, a scalable ConSmax hardware design with a bitwidth-split look-up table (LUT) can achieve lossless non-linear operations and support mixed-precision computing. Experimental results show that ConSmax achieves a minuscule power consumption of 0.2mW and an area of 0.0008mm2 at 1250MHz working frequency in 16nm FinFET technology. For open-source contribution, we further implement our design with the OpenROAD toolchain under SkyWater's 130nm CMOS technology. The corresponding power is 2.69mW and the area is 0.007mm2. ConSmax achieves 3.35× power savings and 2.75× area savings in 16nm technology, and 3.15× power savings and 4.14× area savings with the open-source EDA toolchain. In the meantime, it also maintains comparable accuracy on the GPT-2 model and the WikiText103 dataset. The project is available at https://github.com/ReaLLMASIC/ConSmax.
Shiwei Liu 0002, Guanchen Tao, Yifei Zou, Derek Chow, Zichen Fan, Kauna Lei, Bangfei Pan, Dennis Sylvester, Gregory Kielian, Mehdi Saligane
ICCAD10
2023 An Open Source Compatible Framework to Fully Autonomous Digital LDO Generation
abstract
This work presents an open-source methodology to automate the design and layout of a low dropout (LDO) regulator from high-level performance specifications. LDO designs with this methodology have been demonstrated in commercial 65nm, 12nm, 130nm processes and the open-source 130nm Skywater PDK. The tool currently supports LDO designs with 50mV/100mV dropout for an input voltage range of 0.6V-1.3V (130nm and 65nm), 0.6V-0.9V (12nm), 1.8V-3.3V (Skywater 130nm) and a maximum load current ranging from 0.5mA-25mA (130nm and 65nm), 1mA-20mA (12nm), 0.5mA-50mA (Skywater 130nm). Cell-based design approach is adopted using an auxiliary cell library to enable mixed-signal design synthesis. A port to a new technology only requires a one-time manual layout for auxiliary library generation. The design automation includes a technology-agnostic modeling step and generates the LDO layout automatically. A bi-directional shift register based DLDO with a 1-bit comparator and a stochastic flash ADC (achieving 15x faster settling time) for error detection has been generated using the tool and validated using silicon measurements from 65nm process. LDO designs with different switch types and load configurations have been fabricated in open-source Skywater 130nm.
Yaswanth K. Cherivirala, Mehdi Saligane, David D. Wentzloff
ISCAS2
2020 The Missing Pieces of Open Design Enablement: A Recent History of Google Efforts : lnvited Paper
abstract
In an initiative to advance the open-source electronic design automation (EDA) and hardware design community, Google has been spearheading a global collaborative effort involving investigators from academia, start-ups as well as foundries. Open-source silicon being the end goal, multiple blossoming projects are supported to drive the renewed open-source wave to break down the barriers of EDA tooling and ultimately hardware design. This push toward the democratization of hardware also aims to develop and release an open-source platform of silicon-proven analog and digital IP blocks to serve as a foundation for rapid design of complex, secure systems-on-chip (SoCs) at bleeding edge technology nodes. This paper details the different efforts that constitute the missing pieces standing in the way of open design enablement such as: OpenROAD for EDA tooling, digital libraries such as OpenRAM and standard cells, and finally analog and mixed-signal (AMS) building blocks for SoCs such as: BAG and FASoC.
Tim Ansell, Mehdi Saligane
ICCAD2
2020 Bridging Academic Open-Source EDA to Real-World Usability
abstract
Several academic EDA tools have been released; however, few are used in real tapeouts, even by other academics. Robust open-source tools require feedback and direction from users. To this end, OpenROAD employs end-users as internal design advisors who bring with them the experience of multiple tapeouts and EDA tool flow development. This paper discusses the OpenROAD design advisors' ongoing work to bring OpenROAD from a collection of tools to an end-to-end autonomous design flow. We discuss our work to fill in the gaps for a full RTL-to-GDS design flow, assemble a full-flow test suite reflective of real tapeouts, debug flow-level issues between tools, and bridge the gap between OpenROAD developers and others in the open-source community. Lastly, we discuss OpenROAD's long-term goal to become fully autonomous, and what that means from a user's perspective.
Austin Rovinski, Tutu Ajayi, Guanru Wang, Mehdi Saligane
ICCAD5
2020 An Open-source Framework for Autonomous SoC Design with Analog Block Generation
abstract
We present the world's first autonomous mixed-signal SoC framework, driven entirely by user constraints, along with a suite of automated generators for analog blocks. The process-agnostic framework takes high-level user intent as inputs to generate optimized and fully verified analog blocks using a cell-based design methodology. Our approach is highly scalable and silicon-proven by an SoC prototype which includes 2 PLLs, 3 LDOs, 1 SRAM, and 2 temperature sensors fully integrated with a processor in a 65nm CMOS process. The physical design of all blocks, including analog, is achieved using optimized synthesis and APR flows in commercially available tools. The framework is portable across different processes and requires no-human-in-the-Ioop, dramatically accelerating design time.
Tutu Ajayi, Sumanth Kamineni, Yaswanth K. Cherivirala, Morteza Fayazi, Kyumin Kwon, Mehdi Saligane, Shourya Gupta, Chien-Hen Chen, Dennis Sylvester, David T. Blaauw, Ronald G. Dreslinski, Benton H. Calhoun, David D. Wentzloff
VLSI-SOC6
2019 Toward an Open-Source Digital Flow: First Learnings from the OpenROAD Project
abstract
We describe the planned Alpha release of OpenROAD, an open-source end-to-end silicon compiler. OpenROAD will help realize the goal of "democratization of hardware design", by reducing cost, expertise, schedule and risk barriers that confront system designers today. The development of open-source, self-driving design tools is in and of itself a "moon shot" with numerous technical and cultural challenges. The open-source flow incorporates a compatible open-source set of tools that span logic synthesis, floorplanning, placement, clock tree synthesis, global routing and detailed routing. The flow also incorporates analysis and support tools for static timing analysis, parasitic extraction, power integrity analysis, and cloud deployment. We also note several observed challenges, or "lessons learned", with respect to development of open-source EDA tools and flows.
Tutu Ajayi, Vidya A. Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B. Kahng, Jeongsup Lee, Uday Mallappa, Marina Neseem, Geraldo Pradipta, Sherief Reda, Mehdi Saligane, Sachin S. Sapatnekar, Carl Sechen, Mohamed Shalan, William Swartz, Lutong Wang, Zhehong Wang, Mingyu Woo, Bangqi Xu
DAC13