VLDB 2026 Research / reviewers in the wild / expert
Borivoje Nikolic
dblp:40/6998
· DBLP profile ↗
77ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0003-2324-1715ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 1 first-author · 22 since 2021Computer networks · 20 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 14 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simulator-Driven Deep Reinforcement Learning for Analog Circuit DesignabstractThis work addresses the use of reinforcement learning in the design of analog and mixed-signal (AMS) circuits. With recent advanced angstrom-technology-nodes adding new complexities, this highly manual process has grown increasingly challenging and less aligned with conventional design intuition. The presented approach modifies circuit topologies at the transistor-level to meet design requirements. We present, for the first time, a deep reinforcement learning (RL) framework capable of generating novel circuit topologies by using graph encodings for targeted specifications, starting from a minimal expert design and a user-specified testbench. To highlight the capabilities of the approach, we demonstrate the topological modification and expansion of incomplete sub-circuits to satisfy user-provided performance for three different types of circuits: 1) a ring oscillator, 2) a comparator, and 3) an operational transconductance amplifier. Our results demonstrate that our method is capable of generating previously unseen topologies that reach user-defined performance targets. In each design case, 100% of generated circuit netlists are correct by construction and over 90% of generated circuits demonstrate intended functionality and targeted performance when simulated with commercial tools. Felicia B. Guo, Ken T. Ho, Andrei Vladimirescu, Borivoje Nikolic |
DATE | 4 |
| 2026 | Substrate: A Statically Typed Framework for Designing Highly Configurable Analog and Mixed-Signal Circuit GeneratorsabstractAnalo and mixed-signal (AMS) integrated circuit design is often a time-consuming and costly process, due in part to manual design flows and long layout iterations. A number of tools have been developed aiming to automate the process of creating AMS designs. However, existing tools are often difficult to use due to unclear application programming interfaces (APIs), limited levels of abstraction, or insufficient control over generated collateral. We introduce Substrate, an open-source, statically typed framework for creating highly configurable schematic and layout generators using the Rust programming language. Substrate provides multiple levels of abstraction, allowing designers to navigate the tradeoff between fine-grained control over a design and increased automation. We also describe algorithms for programmatically creating and modifying circuit layouts, including two methods for automatically adjusting the aspect ratio of a layout. We use Substrate to design generators for a StrongARM comparator and a programmable resistor bank in Skywater 130nm and Intel 16nm, demonstrating 90 degree rotation, array folding, and the ability to change the aspect ratio by a factor of over 10 in both processes. These generators highlight Substrate’s ability to facilitate design reuse, process portability, and performance and area optimization. Borivoje Nikolic |
DATE | 3 |
| 2025 | SuperNoVA: Algorithm-Hardware Co-Design for Resource-Aware SLAM
Seah Kim, Roger Hsiao, Borivoje Nikolic, James Demmel, Sophia Shao |
ASPLOS (1) | 3 |
| 2025 | Taping Out Three Class Chips Per Semester in Intel 16 TechnologyabstractWe present an agile methodology based on the open source Chipyard framework used for designing and validating manufacturable and performant heterogeneous RISC-V SoCs within the constraints of 15-week semesters by classes composed primarily of undergraduate students. Chipyard integrates configurable, generator-based IP blocks and flows, including the modular VLSI flow, Hammer, developed over a decade of tapeouts in different technologies. Students iterate their custom RTL and AMS blocks through integration, verification, and place-and-route, then write full-stack applications, characterizing performance. One recent semester’s class chips in FinFET are described: COSMIC (FFT, convolution, DMA accelerators), MELLIS (sparse‑matrix, convolution, quantized transformer engines, near‑memory MAC), and SCμM‑V (low-power crystal-free transceiver, general-purpose AFE, on-chip power management, clock generation). For example, COSMIC reaches 1.25 GHz, accelerates compute 2-12× with energy savings, and runs live demos. New documentation and infrastructure, such as the new bring-up platform, Baremetal, make Chipyard even more accessible. Lucy Revina, Ethan Gao, Ken Ho, Daniel Lovell, Kristofer S. J. Pister, Borivoje Nikolic |
HCS | 6 |
| 2025 | Design Techniques for a Multi-Phase Injection- Based Eight-Phase 17-GHz Clock Generator for Multi-Phase Wireline ReceiversabstractClock generation for high-speed wireline receivers must provide multiple clock phases with high-resolution rotation. To address this, an 8-phase 17 GHz clock generation circuit with built-in 6b rotation is presented. Multi-phase injection is used to perform reference-side phase rotation to efficiently generate and rotate eight clock phases. The injection method is analyzed with a model to study the introduced nonlinearity, and the effect of the injection strength is discussed. Designed by using BAG3++ for layout-aware design optimization, the proposed circuit achieves 98 fs RMS jitter and a measured DNLpp and INLpp of 1.26 and 4.05 LSB respectively, while consuming 33 mW. Bob L. Zhou, Borivoje Nikolic |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | RTL-Repair: Fast Symbolic Repair of Hardware Design CodeabstractWe present RTL-Repair, a semantics-based repair tool for register transfer level circuit descriptions. Compared to the previous state-of-the-art tool, RTL-Repair generates more correct repairs within seconds instead of minutes or even hours. We imagine that RTL-Repair could thus be integrated into an IDE to give developers repair suggestions promptly. Our new SMT-based one-step fault localization and repair algorithm for digital hardware designs uses optimization to generate minimal changes that a user can easily understand. A novel adaptive windowing approach allows us to avoid scalability issues by focusing the repair search on the parts of the test that matter the most. RTL-Repair provides repairs that pass their testbench for 9 out of 12 real bugs collected from open-source hardware projects. Two repairs fully match the ground truth, one partially, four more repairs change the correct expression but in a way that overfits the testbench, and only three repairs differ strongly from the ground truth. Kevin Laeufer, Brandon Fajardo, Abhik Ahuja, Vighnesh Iyer, Borivoje Nikolic, Koushik Sen |
ASPLOS (3) | 5 |
| 2024 | NeCTAr and RASoC: Tale of Two Class SoCs for Language Model Interference and Robotics in Intel 16abstractThis paper introduces NeCTAr (Near-Cache Transformer Accelerator), a 16nm heterogeneous multicore RISC-V SoC for sparse and dense machine learning kernels with both near-core and near-memory accelerators. A prototype chip runs at 400MHz at 0.85V and performs matrix-vector multiplications with 109 GOPs/W. The effectiveness of the design is demonstrated by running inference on a sparse language model, ReLU-Llama. Viansa Schmulbach, Ethan Gao, Nikhil Jha, Ethan Wu, Oliver Yu, Ben Oliveau, Brendan Roberts, Connor McMahon, Lixiang Yin, Vamber Yang, Brendan Brenner, George Moujaes, Boyu Hao, Lucy Revina, Bryan Ngo, Yufeng Chi, Hongyi Huang, Reza Sajadiany, Raghav Gupta 0001, Ella Schwarz, Jennifer Zhou, Ken Ho, Jerry Zhao, Anita Flynn, Borivoje Nikolic |
HCS | 28 |
| 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsabstractPre-silicon validation and end-to-end system evaluation are integral parts of hardware development as they provide architects with insights about the complex interactions between various hardware components, system software, and application code. Although this process can be accelerated using FPGAs as a simulation host, existing platforms fall short when the resource requirements of a custom hardware design exceed a single FPGA. We present FireAxe, an open-source FPGA-accelerated RTL simulation platform that supports push-button user-guided partitioning across multiple FPGAs, using a compiler called FireRipper. Given a partition point, FireRipper automatically maps a monolithic RTL design onto multiple FPGAs while providing hardware designers quick feedback about the partition interface and expected simulation performance. Furthermore, FireRipper enables users to choose between an exact-mode which provides cycle-exact results with RTL-level fidelity, or a fast-mode that improves simulation rate while sacrificing fidelity only at the partition boundary. Built on FireSim, FireAxe preserves the ability to elastically scale simulations from on-premises FPGAs to cloud FPGAs. For example, pulling out a core from a systemon-chip (SoC) onto a separate FPGA, we achieve simulation rates of 1.6 MHz using on-premises FPGAs connected by direct-attach cables and 1 MHz on AWS F1 FPGAs using peer-to-peer PCIe. To show FireAxe’s ability to enable pre-silicon performance validation at unprecedented scale, we show several case studies. First, we replicate full-stack system-level effects such as latency spikes from garbage collection in a Golang application on an SoC containing 4 out-of-order (OoO) cores. We also boot Linux on, to our knowledge, the largest OoO core ever cycle-exactly simulated in academia. Lastly, we simulate a system-on-chip containing 24 OoO cores mapped onto five datacenter-class FPGAs. We discover an RTL bug when trying to run Linux user-space applications that did not appear with less substantial software stacks. This was discovered in less than 2 hours using FireAxe and would have taken weeks in a commercial software RTL simulator. Joonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Abraham Gonzalez, Raghav Gupta 0001, Nivedha Krishnakumar, Sagar Karandikar, Borivoje Nikolic, Sophia Shao, Krste Asanovic |
ISCA | 9 |
| 2024 | Design Approach for Die-to-Die Interfaces to Enable Energy-Efficient Chiplet SystemsabstractHeterogeneous chiplet integration and advanced packaging have given a new lease to scaling of compute in a post-Moore era. A critical aspect of designing chiplet systems is die-to-die interfaces for aggregation of smaller disaggregated chiplets. In recent years, interconnects and packaging have made a huge leap forward, thereby, enabling high bandwidth and energy-efficient parallel die-to-die (D2D) interfaces. Instead of bespoke solutions, building a standardized D2D interface provides a mechanism for interoperability between heterogeneous chiplets and facilitates low power and energy-efficient design by prescribing implementation strategy. In this paper, we discuss some of the recent efforts in the standardization of die-to-die interfaces. Starting with an overview of Advanced Interface Bus (AIB) PHY, we emphasize its simplicity and high energy efficiency. Followed by a case study that demonstrates an effective chiplet integration employing AIB. A more recent open standard for chiplet integration is the Universal Chiplet Interconnect Express (UCIe). We discuss the key aspects of UCIe, including its electrical and packaging characteristics, as well as low-power features that lead to a 10x power reduction compared to typical off-package I/O. Finally, we discuss the UCIe-lite controller, an effort to democratize the chiplet infrastructure by providing a simplified open-source RTL generator of the D2D interface. The generator is highly parameterizable and lightweight, enabling chiplet systems for energy-efficient applications. Vikram Jain, Wei Tang 0010, Zuoguo Wu, Viansa Schmulbach, Sophia Shao, Zhengya Zhang, Borivoje Nikolic |
ISLPED | 7 |
| 2023 | Simulator Independent Coverage for RTL Hardware LanguagesabstractWe demonstrate a new approach to implementing automated coverage metrics including line, toggle, and finite state machine coverage. Each metric is implemented through a compiler pass with a report generator. They are decoupled from the backend simulation, emulation, or formal verification tool through a simple API designed around a single new cover primitive. Our prototype for the Chisel hardware construction language demonstrates support across three software simulators, the FPGA-accelerated FireSim simulator and a formal tool. We demonstrate collecting line coverage while booting Linux with FireSim at a target frequency of 65MHz. By construction, coverage can be trivially merged across backends. Kevin Laeufer, Vighnesh Iyer, David Biancolin, Jonathan Bachrach, Borivoje Nikolic, Koushik Sen |
ASPLOS (3) | 5 |
| 2023 | A Heterogeneous SoC for Bluetooth LE in 28nmabstractOsciBear is a system-on-chip (SoC) featuring a RISC-V 32-bit 5-stage in-order scalar processor, AES accelerator, BLE 1M baseband-modem, and a 2.4 GHz radio front end (RFE) transceiver. It was designed in TSMC's 28nm process with a total die area of 1 mm2during the course of a 14-week semester by 18 students - 4 Ph.D students, 6 masters students, and 8 undergraduates - enrolled in UC Berkeley's special topics course “28nm SoC for loT” in Spring 2021. Additionally, a PCB was designed with off-chip reference clocks, bring-up tooling, as well as power amplifiers, RF switch, and an antenna to complete the radio front-end. The CPU has been demonstrated to run up to 30 MHz in typical operating conditions. The BLE 1M-compliant PHY layer packet assembly and disassembly has been verified in-hardware through “loopback” testing. Adherence to BLE's PHY FM specifications has also been verified with a commercial BLE receiver. In total, the chip consumes 8.43 mW of static power. Felicia Guo, Nayiri Krzysztofowicz, Alex Moreno, Jeffrey Ni, Daniel Lovell, Yufeng Chi, Kareem Ahmad, Sherwin Afshar, Josh Alexander, Dylan Brater, Daniel Fan, Ryan Lund, Jackson Paddock, Griffin Prechter, Troy Sheldon, Shreesha Sreedhara, Anson Tsai, Eric Wu, Kerry Yu, Daniel Fritchman, Aviral Pandey, Ali M. Niknejad, Kristofer S. J. Pister, Borivoje Nikolic |
HCS | 25 |
| 2023 | MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural NetworksabstractDriven by the wide adoption of deep neural networks (DNNs) across different application domains, multi-tenancy execution, where multiple DNNs are deployed simultaneously on the same hardware, has been proposed to satisfy the latency requirements of different applications while improving the overall system utilization. However, multi-tenancy execution could lead to undesired system-level resource contention, causing quality-of-service (QoS) degradation for latency-critical applications.To address this challenge, we propose MoCA1, an adaptive multi-tenancy system for DNN accelerators. Unlike existing solutions that focus on compute resource partition, MoCA dynamically manages shared memory resources of co-located applications to meet their QoS targets. Specifically, MoCA leverages the regularities in both DNN operators and accelerators to dynamically modulate memory access rates based on their latency targets and user-defined priorities so that co-located applications get the resources they demand without significantly starving their co-runners. We demonstrate that MoCA improves the satisfaction rate of the service level agreement (SLA) up to 3.9× (1.8× average), system throughput by 2.3× (1.7× average), and fairness by 1.3× (1.2× average), compared to prior work. Seah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic, Borivoje Nikolic, Sophia Shao |
HPCA | 5 |
| 2023 | CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale SystemsabstractGeneral-purpose lossless data compression and decompression ("(de)compression") are used widely in hyperscale systems and are key "datacenter taxes". However, designing optimal hardware compression and decompression processing units ("CDPUs") is challenging due to the variety of algorithms deployed, input data characteristics, and evolving costs of CPU cycles, network bandwidth, and memory/storage capacities. Sagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo, Jerry Zhao, Svilen Kanev, Edwin Lim, Jyrki Alakuijala, Vrishab Madduri, Sophia Shao, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan |
ISCA | 11 |
| 2023 | RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC EvaluationabstractRobotic systems, such as autonomous unmanned aerial vehicles (UAVs) and self-driving cars, have been widely deployed in many scenarios and have the potential to revolutionize the future generation of computing. To improve the performance and energy efficiency of robotic platforms, significant research efforts are being devoted to developing hardware accelerators for workloads that form bottlenecks in the robotics software pipeline. Although domain-specific accelerators can offer improved efficiency over general-purpose processors on isolated robotics benchmarks, system-level constraints such as data movement and contention over shared resources can significantly impact the achievable end-to-end acceleration. In addition, the closed-loop nature of robotic systems, where there is a tight interaction across different deployed environments, software stacks, and hardware architecture, further exacerbates the difficulties of evaluating robotics SoCs. Dima Nikiforov, Shengjun Chris Dong, Chengyi Lux Zhang, Seah Kim, Borivoje Nikolic, Sophia Shao |
ISCA | 5 |
| 2023 | AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsabstractWith the widespread adoption of deep neural networks (DNNs) across applications, there is a growing demand for DNN deployment solutions that can seamlessly support multi-tenant execution. This involves simultaneously running multiple DNN workloads on heterogeneous architectures with domain-specific accelerators. However, existing accelerator interfaces directly bind the accelerator’s physical resources to user threads, without an efficient mechanism to adaptively re-partition available resources. This leads to high programming complexities and performance overheads due to sub-optimal resource allocation, making scalable many-accelerator deployment impractical. Seah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic, Sophia Shao |
MICRO | 4 |
| 2022 | Hammer: a modular and reusable physical design flow tool: invitedabstractProcess technology scaling and hardware architecture specialization have vastly increased the need for chip design space exploration, while optimizing for power, performance, and area. Hammer is an open-source, reusable physical design (PD) flow generator that reduces design effort and increases portability by enforcing a separation among design-, tool-, and process technology-specific concerns with a modular software architecture. In this work, we outline Hammer's structure and highlight recent extensions that support both physical chip designers and hardware architects evaluating the merit and feasibility of their proposed designs. This is accomplished through the integration of more tools and process technologies---some open-source---and the designer-driven development of flow step generators. An evaluation of chip designs in process technologies ranging from 130nm down to 12nm across a series of RISC-V-based chips shows how Hammer-generated flows are reusable and enable efficient optimization for diverse applications. Harrison Liew, Daniel Grubb, Colin Schmidt 0001, Nayiri Krzysztofowicz, Adam M. Izraelevitz, Krste Asanovic, Jonathan Bachrach, Borivoje Nikolic |
DAC | 10 |
| 2022 | Automated Design of Analog Circuits Using Reinforcement LearningabstractAnalog and mixed-signal (AMS) blocks are often a crucial and time-consuming part of System-on-Chip (SoC) design, primarily due to a manual circuit and layout iterations. Existing automated solutions for selecting circuit parameters for a given target specification are often not efficient, accurate, or reliable. In order for an automated sizing tool to be practical, we posit that it must: 1) return valid results for a large range of target specifications; 2) understand where and why it is unable to meet certain specifications; 3) consider true layout parasitic simulations for complete end-to-end design; and 4) be automated, allowing most of the design effort to fall on the tool. In this article, we address these critical points by establishing an automated reinforcement learning framework, AutoCkt, by 1) successfully deploying it on a complex two-stage transimpedance amplifier and two-stage folded cascode with biasing in the 16-nm FinFet technology; 2) implementing a new combined distribution deployment algorithm to improve efficiency; 3) analyzing in-depth the efficacy of the trained agent; and 4) demonstrating the functionality of this tool when considering a topology that is highly sensitive to layout parasitics. Our algorithm not only successfully reaches unique, valid, and practical performances, but also does so in state-of-the-art run time, up to 38X more efficient than prior work. In addition, our tool averages just four parasitic simulations obtained by using the Berkeley Analog Generator, to achieve a target specification post-layout for the folded cascode. AutoCkt successfully generates LVS-passed designs with validation in process corner variation results. Keertana Settaluri, Zhaokai Liu, Rishubh Khurana, S. Arash Mirhaj, Rajeev Jain, Borivoje Nikolic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationabstractDNN accelerators are often developed and evaluated in isolation without considering the cross-stack, system-level effects in real-world environments. This makes it difficult to appreciate the impact of Systemon-Chip (SoC) resource contention, OS overheads, and programming-stack inefficiencies on overall performance/energy-efficiency. To address this challenge, we present Gemmini, an open-source, full-stack DNN accelerator generator. Gemmini generates a wide design-space of efficient ASIC accelerators from a flexible architectural template, together with flexible programming stacks and full SoCs with shared resources that capture system-level effects. Gemmini-generated accelerators have also been fabricated, delivering up to three orders-of-magnitude speedups over high-performance CPUs on various DNN benchmarks. Hasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali, Vighnesh Iyer, Pranav Prakash, Jerry Zhao, Daniel Grubb, Harrison Liew, Howard Mao, Albert J. Ou, Colin Schmidt 0001, Samuel Steffl, John Charles Wright, Ion Stoica, Jonathan Ragan-Kelley, Krste Asanovic, Borivoje Nikolic, Sophia Shao |
DAC | 18 |
| 2021 | An Automated and Process-Portable Generator for Phase-Locked LoopabstractWe present a bang-bang phase-locked loop (PLL) generator that encapsulates design methodologies for its circuit blocks and the complete PLL system. The generator is fully automated and parameterized, producing the layout and schematic based on process characterization and top-level specifications. Three 14GHz PLLs are instantiated in TSMC 16nm, GF 14nm and Intel 22nm technologies, demonstrating the process portability. The rapid generation time of less than four days enables fast PLL design and technology porting. The PLL design fabricated in TSMC 16nm shows RMS jitter of 565.4fs and power of 6.64mW from a 0.9V supply. Zhongkai Wang, Minsoo Choi 0002, Eric Chang, John Charles Wright, Wooham Bae, Sijun Du, Zhaokai Liu, Nathan Narevsky, Colin Schmidt 0001, Ayan Biswas 0004, Borivoje Nikolic, Elad Alon |
DAC | 11 |
| 2021 | A Scalable Massive MIMO Uplink Baseband Processing GeneratorabstractThis paper describes a scalable, highly portable, and power-efficient generator for massive multiple-input multiple-output (MIMO) uplink baseband processing. This generator is written in Chisel, and produces hardware instances for the distributed processing in a scalable massive MIMO system. The generator is parameterized in both the MIMO system and hardware datapath elements. The performance of several generator instances with different parameter values are validated by emulation on a field-programmable gate array (FPGA), demonstrating both functionality and scalability, and operation up to 6.4Gb/s data throughput. Greg LaCaille, Harrison Liew, James Dunn 0003, Borivoje Nikolic |
ICC | 5 |
| 2021 | Vertically Integrated Computing Labs Using Open-Source Hardware Generators and Cloud-Hosted FPGAsabstractThe design of computing systems has changed dramatically over the past decade, but most courses in advanced computer architecture remain unchanged. Computer architecture education lies at the intersection between computer science and electrical engineering, with practical exercises in classes based on appropriate levels of abstraction in the computing system design stack. Hardware-centric lab exercises often require broad infrastructure resources and tend to navigate around tedious practical implementation concepts, while software-centric exercises leave a gap between modeling and system implementation implications that students later need to overcome in professional settings. Vertical integration trends in domain-specific compute systems, as well as software-hardware co-design, are often covered in classroom lectures, but are not reflected in laboratory exercises due to complex tooling and simulation infrastructure. We describe our experiences with a joint hardware-software approach to exploring computer architecture concepts in class exercises, by using open- source processor hardware implementations, generator-based hardware design methodologies, and cloud-hosted FPGAs. This approach further enables scaling course enrollment, remote learning and a cross-class collaborative lab ecosystem, creating a connecting thread between computer science and electrical engineering experience-based curricula. Alon Amid, Albert J. Ou, Krste Asanovic, Sophia Shao, Borivoje Nikolic |
ISCAS | 5 |
| 2021 | Memory-Efficient Hardware Performance Counters with Approximate-Counting AlgorithmsabstractHardware performance counters are special registers on processors that track the hardware activities. While the performance counter data are useful for many applications, there are challenges in efficiently collecting many event statistics simultaneously, due to the limited number of performance counters on chip. We propose an efficient hardware performance counter design that uses approximate-counting algorithms to improve the number of events tracked on-chip without incurring significant memory overhead. These counters are more memory efficient because they increment counts according to a dynamic probability and approximate the exact counts. Compared with multiplexed hardware performance counters, our approximate hardware counters have a statistically provable memory-accuracy trade-off and are entirely managed in hardware. Sehoon Kim 0001, Borivoje Nikolic, Sophia Shao |
ISPASS | 3 |
| 2021 | A Hardware Accelerator for Protocol BuffersabstractSerialization frameworks are a fundamental component of scale-out systems, but introduce significant compute overheads. However, they are amenable to acceleration with specialized hardware. To understand the trade-offs involved in architecting such an accelerator, we present the first in-depth study of serialization framework usage at scale by profiling Protocol Buffers (“protobuf”) usage across Google’s datacenter fleet. We use this data to build HyperProtoBench, an open-source benchmark representative of key serialization-framework user services at scale. In doing so, we identify key insights that challenge prevailing assumptions about serialization framework usage. Sagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao, Dinesh Parimi, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan |
MICRO | 6 |
| 2021 | LAYGO: A Template-and-Grid-Based Layout Generation Engine for Advanced CMOS TechnologiesabstractLAYout with Gridded Objects (LAYGO), a Python-based layout-generation engine for enhancing the design productivity of custom circuit layouts in advanced CMOS processes, is presented and verified by implementing a time-interleaved SAR (TI-SAR) ADC instance in a 16 nm CMOS FinFET technology. LAYGO supports rapid generation by placing customized templates on process-specific placement grids, thereby encapsulating the design rules and process-specific structures. The templates can be located based on their relative positional information, which further enhances the description capability and portability. Interconnecting wires are routed on the grids for design rule abstractions, with additional customizations and support for multi-patterning. The functions for the on-grid placement and routing use advanced indexing and slicing with multi-dimensional object containers to improve the description and parameterization capabilities. Multiple TI-SAR ADC layouts are generated using LAYGO in 28-16 nm CMOS technologies. One instance is fabricated in a 16 nm CMOS FinFET process and measured, achieving a 38.2 dB signal-to-noise-and-distortion ratio (SNDR) at 7 GS/s after digital calibration and consuming 45.2 mW. Owing to its high customization capability, the design achieved the highest sampling rate (7 GS/s) among the generated ADCs. Jaeduk Han, Woo-Rham Bae, Eric Chang, Zhongkai Wang, Borivoje Nikolic, Elad Alon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignabstractAchieving high-performance when developing specialized hardware/software systems requires understanding and improving not only core compute kernels, but also intricate and elusive system-level bottlenecks. Profiling these bottlenecks requires both high-fidelity introspection and the ability to run sufficiently many cycles to execute complex software stacks, a challenging combination. In this work, we enable agile full-system performance optimization for hardware/software systems with FirePerf, a set of novel out-of-band system-level performance profiling capabilities integrated into the open-source FireSim FPGA-accelerated hardware simulation platform. Using out-of-band call stack reconstruction and automatic performance counter insertion, FirePerf enables introspecting into hardware and software at appropriate abstraction levels to rapidly identify opportunities for software optimization and hardware specialization, without disrupting end-to-end system behavior like traditional profiling tools. We demonstrate the capabilities of FirePerf with a case study that optimizes the hardware/software stack of an open-source RISC-V SoC with an Ethernet NIC to achieve 8x end-to-end improvement in achievable bandwidth for networking applications running on Linux. We also deploy a RISC-V Linux kernel optimization discovered with FirePerf on commercial RISC-V silicon, resulting in up to 1.72x improvement in network performance. Sagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao, Randy H. Katz, Borivoje Nikolic, Krste Asanovic |
ASPLOS | 6 |
| 2020 | Invited: Chipyard - An Integrated SoC Research and Implementation EnvironmentabstractContinued improvement in computing efficiency requires functional specialization of hardware designs. We present an agile design flow for custom SoCs using the Chipyard framework, an integrated SoC research and implementation environment for custom systems. Chipyard includes configurable, composable, open-source, generator-based designs that can be used across multiple stages of the hardware development flow while maintaining design intent and integration consistency. Through cloud FPGA simulation and rapid ASIC implementation, we demonstrate an iterative agile hardware design cycle which enables continuous validation of physically-realizable customized systems. Alon Amid, David Biancolin, Abraham Gonzalez, Daniel Grubb, Sagar Karandikar, Harrison Liew, Albert Magyar, Howard Mao, Albert J. Ou, Nathan Pemberton, Paul Rigge, Colin Schmidt 0001, John Charles Wright, Jerry Zhao, Jonathan Bachrach, Sophia Shao, Borivoje Nikolic, Krste Asanovic |
DAC | 17 |
| 2020 | AutoCkt: Deep Reinforcement Learning of Analog Circuit DesignsabstractDomain specialization under energy constraints in deeply-scaled CMOS has been driving the need for agile development of Systems on a Chip (SoCs). While digital subsystems have design flows that are conducive to rapid iterations from specification to layout, analog and mixed-signal modules face the challenge of a long human-in-the-middle iteration loop that requires expert intuition to verify that post-layout circuit parameters meet the original design specification. Existing automated solutions that optimize circuit parameters for a given target design specification have limitations of being schematic-only, inaccurate, sample-inefficient or not generalizable. This work presents AutoCkt, a machine learning optimization framework trained using deep reinforcement learning that not only finds post-layout circuit parameters for a given target specification, but also gains knowledge about the entire design space through a sparse subsampling technique. Our results show that for multiple circuit topologies, AutoCkt is able to converge and meet all target specifications on at least 96.3% of tested design goals in schematic simulation, on average 40× faster than a traditional genetic algorithm. Using the Berkeley Analog Generator, AutoCkt is able to design 40 LVS passed operational amplifiers in 68 hours, 9.6× faster than the state-of-the-art when considering layout parasitics. Keertana Settaluri, Ameer Haj-Ali, Qijing Huang 0001, Kourosh Hakhamaneshi, Borivoje Nikolic |
DATE | 5 |
| 2020 | Wireless Channel Dynamics for Relay Selection under Ultra-Reliable Low-Latency CommunicationabstractUltra-reliable, low-latency communication (URLLC) is being developed to support critical control applications over wireless networks. Exploiting spatial diversity through relays is a promising technique for achieving the stringent requirements of URLLC, but coordinating relays reliably and with low overhead is a challenge. Adaptive relay selection techniques have been proposed as a way to simplify implementation while still achieving the requirements of URLLC. Identifying good relays with low overhead and high confidence is critical for such adaptive relay selection techniques.Channel dynamics must be taken into account by adaptive relay selection algorithms because channel quality may degrade in the time it takes to estimate the relay's channel and schedule a transmission. Spatial channel dynamics are well studied in many settings such as RADAR and the fast-fading wireless channels, but less so in the URLLC context where rare events neglected in other models may be important. In this work, we perform measurements to validate channel models in the slow fading regime of interest. We compare measurements to Jakes's model and discuss the appropriateness of Jakes's model for URLLC relay selection. This is further applied to demonstrate that easily implementable relay selection techniques perform well in practical settings.Polynomial interpolation and neural-net-based algorithms were evaluated as channel prediction algorithms. These techniques perform orders of magnitude better than relay selection on average (nominal) SNR. Paul Rigge, Vasuki Narasimha Swamy, Christian Nelson, Fredrik Tufvesson, Anant Sahai, Borivoje Nikolic |
PIMRC | 6 |
| 2020 | A Dual-Core RISC-V Vector Processor With On-Chip Fine-Grain Power Management in 28-nm FD-SOIabstractThis work demonstrates a dual-core RISC-V system-on-chip (SoC) with integrated fine-grain power management. The 28-nm fully depleted silicon-on-insulator (FD-SOI) SoC integrates switched-capacitor voltage converters and 4-Gb/s off-chip serial links. The SoC runs applications with operating system support on dual RISC-V Rocket cores with vector accelerators. Runtime monitoring of microarchitectural counters allows prediction of future compute intensity, enabling the voltage state of the managed core to be adjusted quickly to optimize energy efficiency without sacrificing overall performance. John Charles Wright, Colin Schmidt 0001, Ben Keller, Daniel Palmer Dabbelt, Jaehwa Kwak, Vighnesh Iyer, Nandish Mehta, Pi-Feng Chiu, Stevo Bailey, Krste Asanovic, Borivoje Nikolic |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2019 | Open-Source EDA Tools and IP, A View from the TrenchesabstractWe describe our experience developing and promoting a set of open-source tools and IP over the last 9 years, including the Chisel hardware construction language, the Rocket Chip SoC generator, and the BAG analog layout generator. Elad Alon, Krste Asanovic, Jonathan Bachrach, Borivoje Nikolic |
DAC | 4 |
| 2019 | RTL bug localization through LTL specification mining (WIP)abstractAs the complexity of contemporary hardware designs continues to grow, functional verification demands more effort and resources in the design cycle than ever. As a result, manually debugging RTL designs is extremely challenging even with full signal traces after detecting errors in chip-level software simulation or FPGA emulation. Therefore, it is necessary to reduce the burden of verification by automating RTL debugging processes. Vighnesh Iyer, Borivoje Nikolic, Sanjit A. Seshia |
MEMOCODE | 3 |
| 2019 | Wireless Channel Dynamics and Robustness for Ultra-Reliable Low-Latency CommunicationsabstractInteractive, immersive, and other timing-critical applications demand ultra-reliable low-latency communication (URLLC). To build wireless communication systems that can support these applications, understanding the relevant characteristics of the wireless medium is paramount. Although wireless channel characteristics and dynamics have been extensively studied, it is important to revisit these concepts in the context of the strict demands of low-latency and ultra-high reliability. In this paper, we bring a modeling approach from robust control to wireless communication-the wireless channel characteristics are given a nominal model around which we allow for some quantified uncertainty. We propose certain key URLLC-relevant parameters along which the model uncertainty is to be bounded. To validate the nominal model of the spatially independent quasi-static Rayleigh fading, we take an in-depth look at the spatial and temporal correlations based on Jakes' model. We find that although the Rayleigh fading process is band-limited, the quasi-static assumption is not safe for relay selection even well within a single coherence time. We also find that under reasonable conditions, the spatial correlation of channels provide a fading distribution that is not too far off from an independent spatial fading model. In addition, we look at the impact of these channel models on cooperative communication-based systems. We find that while spatial-diversity-based techniques are necessary to combat the adverse effects of fading, time-diversity-based techniques are necessary to be robust against unmodeled errors. Robust URLLC systems need to operate with both an adequate SNR margin and a time margin through repetitions. Vasuki Narasimha Swamy, Paul Rigge, Gireeja Ranade, Borivoje Nikolic, Anant Sahai |
IEEE J. Sel. Areas Commun. | 4 |
| 2018 | ACED: a hardware library for generating DSP systemsabstractDesigners translate DSP algorithms into application-specific hardware via primitives composed in various ways for different architectural realizations. Despite sharing underlying algorithms and hardware constructs, designs are often difficult to reuse, leading to redeveloping/reverifying conceptually similar instances. Hardware generators are attractive solutions for effectively balancing fine-grained control of implementation details with simple, retargetable hardware descriptions. This work presents ACED, a hardware library for generating DSP systems. It extends the Chisel hardware construction language and FIRRTL compiler and operates on three principles: zero-cost abstraction, unobtrusive downstream optimization/specialization promoting generator reusability, and unified, portable systems modeling and verification. Angie Wang, Paul Rigge, Adam M. Izraelevitz, Chick Markley, Jonathan Bachrach, Borivoje Nikolic |
DAC | 6 |
| 2018 | FireSim: FPGA-Accelerated Cycle-Exact Scale-Out System Simulation in the Public CloudabstractWe present FireSim, an open-source simulation platform that enables cycle-exact microarchitectural simulation of large scale-out clusters by combining FPGA-accelerated simulation of silicon-proven RTL designs with a scalable, distributed network simulation. Unlike prior FPGA-accelerated simulation tools, FireSim runs on Amazon EC2 F1, a public cloud FPGA platform, which greatly improves usability, provides elasticity, and lowers the cost of large-scale FPGA-based experiments. We describe the design and implementation of FireSim and show how it can provide sufficient performance to run modern applications at scale, to enable true hardware-software co-design. As an example, we demonstrate automatically generating and deploying a target cluster of 1,024 3.2 GHz quad-core server nodes, each with 16 GB of DRAM, interconnected by a 200 Gbit/s network with 2 microsecond latency, which simulates at a 3.4 MHz processor clock rate (less than 1,000x slowdown over real-time). In aggregate, this FireSim instantiation simulates 4,096 cores and 16 TB of memory, runs ~14 billion instructions per second, and harnesses 12.8 million dollars worth of FPGAs—at a total cost of only ~$100 per simulation hour to the user. We present several examples to show how FireSim can be used to explore various research directions in warehouse-scale machine design, including modeling networks with high-bandwidth and low-latency, integrating arbitrary RTL designs for a variety of commodity and specialized datacenter nodes, and modeling a variety of datacenter organizations, as well as reusing the scale-out FireSim infrastructure to enable fast, massively parallel cycle-exact single-node microarchitectural experimentation. Sagar Karandikar, Howard Mao, David Biancolin, Alon Amid, Dayeol Lee, Nathan Pemberton, Emmanuel Amaro, Colin Schmidt 0001, Aditya Chopra, Qijing Huang 0001, Kyle Kovacs, Borivoje Nikolic, Randy H. Katz, Jonathan Bachrach, Krste Asanovic |
ISCA | 13 |
| 2018 | Predicting Wireless Channels for Ultra-Reliable Low-Latency CommunicationsabstractUltra-reliable, low-latency wireless communication is essential to enable critical and interactive applications. The cooperative communication schemes for such ultra-reliable communication must harvest multi-user diversity to achieve their specifications. The underlying low-latency space-time codes for a large number of users (> 10) place burdens on practical implementations due to the large number of simultaneous relays they must use. To address this, we propose an adaptive relay selection technique that selects a small set of good relays, instead of using every available radio to relay. Using our simple relay-selection schemes, we can support a network with 30 nodes requiring system failure probability under 10 -9and 2ms latency with only 3 simultaneously active relays per message. In contrast, in the absence of adaptive relay selection, we must rely on 13 relays to achieve the same reliability. To arrive at such relay selection schemes, we revisit the fading dynamics of wireless channels in the context of ultra-high reliability. Contrary to what has been claimed in the literature, we find that standard Rayleigh fading processes are not bandlimited. However, these fading processes are fairly predictable on the short time scales of the regime of interest. Vasuki Narasimha Swamy, Paul Rigge, Gireeja Ranade, Borivoje Nikolic, Anant Sahai |
ISIT | 4 |
| 2017 | Use of Phase Delay Analysis for Evaluating Wideband Circuits: An Alternative to Group Delay AnalysisabstractA phase delay analysis is proposed against pursuing a flat group delay response for the design of wideband circuits. While it is believed that a large group delay variation introduces a large data-dependent jitter, this brief reconsiders the effectiveness of the group delay analysis in the evaluation of wideband circuits. Because of its own differentiating nature, the group delay provides a good insight on the delay variation at a vicinity of a certain frequency. However, for certain kind of wideband circuits, the group delay analysis cannot provide a sufficient insight since much of useful information is lost or distorted during the differentiating operation. In this brief, the effectiveness of the phase delay analysis is investigated and comparison with a traditional group delay analysis is presented with a theoretical approach and through a few circuit examples. Woo-Rham Bae, Borivoje Nikolic, Deog-Kyoon Jeong |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Real-Time Cooperative Communication for Automation Over WirelessabstractHigh-performance industrial automation systems rely on tens of simultaneously active sensors and actuators and have stringent communication latency and reliability requirements. Current wireless technologies, such as Wi-Fi, Bluetooth, and LTE are unable to meet these requirements, forcing the use of wired communication in industrial control systems. This paper introduces a wireless communication protocol that capitalizes on multiuser diversity and cooperative communication to achieve the ultra-reliability with a low-latency constraint. Our protocol is analyzed using the communication-theoretic delay-limitedcapacity framework and compared with baseline schemes that primarily exploit frequency diversity. For a scenario inspired by an industrial printing application with 30 nodes in the control loop, 20-B messages transmitted between pairs of nodes and a cycle time of 2 ms, an idealized protocol can achieve a cycle failure probability (probability that any packet in a cycle is not successfully delivered) lower than 10-9with nominal SNR below 5 dB in a 20-MHz wide channel. Vasuki Narasimha Swamy, Sahaana Suri, Paul Rigge, Matthew Weiner, Gireeja Ranade, Anant Sahai, Borivoje Nikolic |
IEEE Trans. Wirel. Commun. | 7 |
| 2016 | A generator of memory-based, runtime-reconfigurable 2N3M5K FFT enginesabstractRuntime-reconfigurable, mixed-radix FFT/IFFT engines are essential for modern wireless communication systems. To comply with varying standards requirements, these engines are customized for each modem. The Chisel hardware construction language has been used in this work to create a generator of runtime-reconfigurable 2n3m5k FFT engines targeting software-defined radios (SDR) for modern communications, but with flexibility to support a wide range of applications. The generator uses a conflict-free, in-place, multi-bank SRAM design, and exploits the duality of decimation-in-frequency (DIF) and decimation-in-time (DIT) FFTs to support continuous data flow with only 2N memory blocks. DFT decomposition using the prime-factor algorithm (PFA) followed by the Cooley-Tukey algorithm (CTA) reduces twiddle ROM sizes. A programmable Winograd's Fourier Transform (WFTA) butterfly supporting radix-2/3/4/5/7 operations reuses radix-7 hardware to support reconfigurability with minimal area penalty. The generated FFTs use 50% less memory than iterative FFTs from Spiral. The twiddle ROM size of the generated LTE/WiFi FFT engine is 16% smaller than that of a 2048-pt Spiral design. Angie Wang, Jonathan Bachrach, Borivoje Nikolic |
ICASSP | 3 |
| 2016 | Phase noise scaling and tracking in OFDM multi-user beamforming arraysabstractMany-element antenna arrays, used for multi-user MIMO, are expected to be one of the cornerstone technologies for 5G wireless systems. Large arrays also offer the opportunity to average out some of the transceivers' analog imperfections, potentially enabling a lower-power implementation. In this paper we study the effect of local oscillator phase noise on beamforming MU-MIMO-OFDM systems. We show that the array does average out uncorrelated phase noise at each element. Exploiting this, we propose scaling the per-element phase noise specification proportionally to the array size, thereby maintaining constant array-level performance with lower power consumption. However, if the phase noise is entirely uncorrelated, this scaling causes a substantial degradation in the recovered signal energy. If, instead, some correlated low-frequency phase noise is introduced at each element, we show that phase noise scaling incurs no performance loss. In fact, under these conditions, a single, global pilot tracking loop can replace carrier recovery at each element. Additionally, this level of phase noise correlation eliminates the phase noise-induced channel aging effect. This type of correlation can be achieved by distributing a common reference and optimizing the bandwidth of the PLL. Antonio Puglielli, Greg LaCaille, Ali M. Niknejad, Gregory Wright, Borivoje Nikolic, Elad Alon |
ICC | 5 |
| 2016 | Network coding for high-reliability low-latency wireless controlabstractThe Internet of Things (IoT) envisions simultaneous sensing and actuation of numerous wirelessly connected devices. Emerging human-in-the-loop applications demand low-latency high-reliability communication protocols, paralleling the requirements for high-performance industrial control. This paper introduces a wireless communication protocol based on network coding that in conjunction with cooperative communication techniques builds the necessary diversity to achieve the target reliability. The proposed protocol, XOR-CoW, is analyzed by using a communication theoretic delay-limited-capacity framework and compared to different realizations of previously proposed protocols without network coding. The results show that as the network size or payload increases, XOR-CoW gains advantage in minimum SNR to achieve the target latency. For a scenario inspired by an industrial printing application with 30 nodes in the control loop, total information throughput of 4.8 Mb/s, 20MHz of bandwidth and cycle time under 2 ms, the protocol can robustly achieve a system probability of error better than 10-9with a nominal SNR less than 2 dB with Rayleigh fading. Vasuki Narasimha Swamy, Paul Rigge, Gireeja Ranade, Anant Sahai, Borivoje Nikolic |
WCNC | 5 |
| 2016 | Design of Energy- and Cost-Efficient Massive MIMO ArraysabstractLarge arrays of radios have been exploited for beamforming and null steering in both radar and communication applications, but cost and form factor limitations have precluded their use in commercial systems. This paper discusses how to build arrays that enable multiuser massive multiple-input-multiple-output (MIMO) and aggressive spatial multiplexing with many users sharing the same spectrum. The focus of the paper is the energy- and cost-efficient realization of these arrays in order to enable new applications. Distributed algorithms for beamforming are proposed, and the optimum array size is considered as a function of the performance of the receiver, transmitter, frequency synthesizer, and signal distribution within the array. The effects of errors such as phase noise and synchronization skew across the array are analyzed. The paper discusses both RF frequencies below 10 GHz, where fully digital techniques are preferred, and operation at millimeter (mm)-wave bands where a combination of digital and analog techniques are needed to keep cost and power low. Antonio Puglielli, Andrew Townley, Greg LaCaille, Vladimir M. Milovanovic, Pengpeng Lu, Konstantin Trotskovsky, Amy Whitcombe, Nathan Narevsky, Gregory Wright, Thomas A. Courtade, Elad Alon, Borivoje Nikolic, Ali M. Niknejad |
Proc. IEEE | 12 |
| 2015 | Raven: A 28nm RISC-V vector processor with integrated switched-capacitor DC-DC converters and adaptive clocking
Yunsup Lee, Brian Zimmer, Andrew Waterman, Alberto Puggelli, Jaehwa Kwak, Ruzica Jevtic, Ben Keller, Stevo Bailey, Milovan Blagojevic, Pi-Feng Chiu, Henry Cook, Rimas Avizienis, Brian C. Richards, Elad Alon, Borivoje Nikolic, Krste Asanovic |
Hot Chips Symposium | 15 |
| 2015 | Cooperative communication for high-reliability low-latency wireless controlabstractThe Internet of Things envisions not only sensing but also actuation of numerous wirelessly connected devices. Seamless control with humans in the loop requires latencies on the order of a millisecond with very high reliabilities, paralleling the requirements for high-performance industrial control. Today's practical wireless systems cannot meet these reliability and latency requirements, forcing the use of wired systems. This paper introduces a wireless communication protocol, dubbed “Occupy CoW,” based on cooperative communication among nodes in the network to build the diversity necessary for the target reliability. Simultaneous retransmission by many relays achieves this without significantly decreasing throughput or increasing latency. The protocol is analyzed using the communication theoretic delay-limited-capacity framework and compared to baseline schemes that primarily exploit frequency diversity. In particular, we develop a novel “diversity meter” designed to measure “effective diversity” in the non-asymptotic regime. For a scenario inspired by an industrial printing application with 30 nodes in the control loop, total information throughput of 4.8 Mb/s, and cycle time under 2 ms, the protocol can robustly achieve a system probability of error better than 10−9with nominal SNR below 5 dB. Vasuki Narasimha Swamy, Sahaana Suri, Paul Rigge, Matthew Weiner, Gireeja Ranade, Anant Sahai, Borivoje Nikolic |
ICC | 7 |
| 2015 | Per-Core DVFS With Switched-Capacitor Converters for Energy Efficiency in Manycore ProcessorsabstractIntegrating multiple power converters on-chip improves energy efficiency of manycore architectures. Switched-capacitor (SC) dc-dc converters are compatible with conventional CMOS processes, but traditional implementations suffer from limited conversion efficiency. We propose a dynamic voltage and frequency scaling scheme with SC converters that achieves high converter efficiency by allowing the output voltage to ripple and having the processor core frequency track the ripple. Minimum core energy is achieved by hopping between different converter modes and tuning body-bias voltages. A multicore processor model based on a 28-nm technology shows conversion efficiencies of 90% along with over 25% improvement in the overall chip energy efficiency. Ruzica Jevtic, Hanh-Phuc Le, Milovan Blagojevic, Stevo Bailey, Krste Asanovic, Elad Alon, Borivoje Nikolic |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2014 | Design of a low-latency, high-reliability wireless communication system for control applicationsabstractHigh-performance industrial control systems with tens to hundreds of sensors and actuators use wired connections between all of their components because they require low-latency, high-reliability links to maintain stability; however, the wires cause many mechanical problems that moving to wireless links would solve. No existing or proposed wireless system can achieve the latency and reliability required by the control algorithms because they are designed for either high-throughput or low-power communication between a pair or a small number of terminals. A preliminary wireless system architecture is proposed that focuses on low-latency operation through the use of reliable broadcasting, semi-fixed resource allocation, and low-rate coding. For an industrial printer application with 30 nodes in the control loop and a moderate information throughput of 4.8Mb/s, the system can achieve latencies under 2ms for SNRs above 7dB. Matthew Weiner, Milos Jorgovanovic, Anant Sahai, Borivoje Nikolic |
ICC | 4 |
| 2013 | Relay scheduling and interference cancellation for quantize-map-and-forward cooperative relayingabstractThis paper presents system design aspects of a multi-relay half-duplex QMF cooperative system. We propose two simple algorithms that address the design of relay scheduling and inter-relay interference schemes. Proposed linear-complexity scheduling algorithm is proven to be optimal for a multi-relay diamond network under specific channel conditions. We demonstrate through simulations that for typical channel conditions, the achievable QMF rate of a five-relay cooperative system is up to 3 times higher compared to a system without cooperation. Milos Jorgovanovic, Matthew Weiner, David Tse, Borivoje Nikolic, I-Hsiang Wang, Vinayak Nagpal |
ISIT | 4 |
| 2013 | Coding and System Design for Quantize-Map-and-Forward RelayingabstractIn this paper we develop a low-complexity coding scheme and system design framework for the half duplex relay channel based on the Quantize-Map-and-Forward (QMF) relaying scheme. The proposed framework allows linear complexity operations at all network terminals. We propose the use of binary LDPC codes for encoding at the source and LDGM codes for mapping at the relay. We express joint decoding at the destination as a belief propagation algorithm over a factor graph. This graph has the LDPC and LDGM codes as subgraphs connected via probabilistic constraints that model the QMF relay operations. We show that this coding framework extends naturally to the high SNR regime using bit interleaved coded modulation (BICM). We develop density evolution analysis tools for this factor graph and demonstrate the design of practical codes for the half-duplex relay channel that perform within 1dB of information theoretic QMF threshold. Vinayak Nagpal, I-Hsiang Wang, Milos Jorgovanovic, David Tse, Borivoje Nikolic |
IEEE J. Sel. Areas Commun. | 5 |
| 2012 | A 15 MHz to 600 MHz, 20 mW, 0.38 mm2 Split-Control, Fast Coarse Locking Digital DLL in 0.13 µ m CMOSabstractA digital delay-locked loop (DLL) suitable for generation of multiphase clocks in applications such as time-interleaved and pipelined analog-to-digital converters (ADCs) locks in a very wide (40×) frequency range. The DLL provides 12 uniformly delayed phases, free of false harmonic locking. A two-stage digital split-control loop is implemented: a fast-locking coarse acquisition is achieved in four cycles using binary search; a fine linear loop achieves low jitter (9 ps rms @ 600 MHz) and tracks process, voltage, and temperature (PVT) variations. The false harmonic locking detector, the frequency range and the jitter performance among other design considerations are analyzed in detail. The DLL consumes 20 mW and occupies a 470 μm × 800 μm in 0.13 μm CMOS. Sebastian Hoyos, Cheongyuen W. Tsang, Johan P. Vanderhaegen, Yun Chiu, Yasutoshi Aibara, Haideh Khorramabadi, Borivoje Nikolic |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2011 | LDPC decoder architecture for high-data rate personal-area networksabstractEmerging standards for wireless communications in the 60GHz band, such as WiGig, IEEE 802.11ad, and IEEE 802.15.3c, require throughputs between 1.5 and 6Gb/s and use rate adaptive low-density parity-check (LDPC) codes as the main form of forward error correction. State-of-the-art flexible LDPC decoders cannot simultaneously achieve the high throughput mandated by these standards and the low power needed for mobile applications. This work develops a flexible, fully pipelined architecture for the IEEE 802.11ad standard capable of achieving both goals. Based on a decoder synthesized in a low-power 65nm CMOS technology, the decoder dissipates 42mW at the 1.5Gb/s throughput and 84mW at the 3Gb/s throughput for the worst-case matrix in the standard. Matthew Weiner, Borivoje Nikolic, Zhengya Zhang |
ISCAS | 2 |
| 2010 | SRAM design in fully-depleted SOI technologyabstractContinued increase in variability is a challenge for SRAM scaling into sub-22 nm nodes, and presents an opportunity for the introduction of alternate technologies. In this work, the performance and threshold-voltage variability of vertical SOI finFETs are compared against those of planar fully depleted (FD) SOI MOSFETs with thin buried oxide, and are presented as an alternative to planar bulk CMOS. Analytical modeling derived from 3D device simulations is used to estimate six-transistor SRAM cell performance and yield metrics. Borivoje Nikolic, Changhwan Shin, Min Hee Cho, Tsu-Jae King Liu, Bich-Yen Nguyen |
ISCAS | 1 |
| 2010 | Analysis of absorbing sets and fully absorbing sets of array-based LDPC codesabstractThe class of low-density parity-check (LDPC) codes is attractive, since such codes can be decoded using practical message-passing algorithms, and their performance is known to approach the Shannon limits for suitably large block lengths. For the intermediate block lengths relevant in applications, however, many LDPC codes exhibit a so-called “error floor,” corresponding to a significant flattening in the curve that relates signal-to-noise ratio (SNR) to the bit-error rate (BER) level. Previous work has linked this behavior to combinatorial substructures within the Tanner graph associated with an LDPC code, known as (fully) absorbing sets. These fully absorbing sets correspond to a particular type of near-codewords or trapping sets that are stable under bit-flipping operations, and exert the dominant effect on the low BER behavior of structured LDPC codes. This paper provides a detailed theoretical analysis of these (fully) absorbing sets for the class of$C_{p, \gamma}$array-based LDPC codes, including the characterization of all minimal (fully) absorbing sets for the array-based LDPC codes for$\gamma = 2,3,4$, and moreover, it provides the development of techniques to enumerate them exactly. Theoretical results of this type provide a foundation for predicting and extrapolating the error floor behavior of LDPC codes. Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Martin J. Wainwright, Borivoje Nikolic |
IEEE Trans. Inf. Theory | 5 |
| 2010 | SRAM Read/Write Margin Enhancements Using FinFETsabstractProcess-induced variations and sub-threshold leakage in bulk-Si technology limit the scaling of SRAM into sub-32 nm nodes. New device architectures are being considered to improve$V_{T}$control and reduce short channel effects. Among the likely candidates, FinFETs are the most attractive option because of their good scalability and possibilities for further SRAM performance and yield enhancement through independent gating. The enhancements to read/write margins and yield are investigated in detail for two cell designs employing independently gated FinFETs. It is shown that FinFET-based 6-T SRAM cells designed with pass-gate feedback (PGFB) achieve significant improvements in the cell read stability without area penalty. The write-ability of the cell can be improved through the use of pull-up write gating (PUWG) with a separate write word line (WWL). The benefits of these two approaches are complementary and additive, allowing for simultaneous read and write yield enhancements when the PGFB and PUWG designs are used in combination. Andrew Carlson, Zheng Guo 0004, Sriram Balasubramanian, Radu Zlatanovici, Tsu-Jae King Liu, Borivoje Nikolic |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2009 | Cooperative multiplexing in the multiple antenna half duplex relay channelabstractCooperation between terminals has been proposed to improve the reliability and throughput of wireless communication. While recent work has shown that relay cooperation provides increased diversity, increased multiplexing gain over that offered by direct link has largely been unexplored. In this work we show that cooperative multiplexing gain can be achieved by using a half duplex relay. We capture relative distances between terminals in the high SNR diversity multiplexing tradeoff (DMT) framework. The DMT performance is then characterized for a network having a single antenna half-duplex relay between a single-antenna source and two-antenna destination. Our results show that the achievable multiplexing gain using cooperation can be greater than that of the direct link and is a function of the relative distance between source and relay compared to the destination. Moreover, for multiplexing gains less than 1, a simple scheme of the relay listening 1/3 of the time and transmitting 2/3 of the time can achieve the 2 by 2 MIMO DMT. Vinayak Nagpal, Sameer Pawar, David Tse, Borivoje Nikolic |
ISIT | 4 |
| 2009 | Predicting error floors of structured LDPC codes: deterministic bounds and estimatesabstractThe error-correcting performance of low-density parity check (LDPC) codes, when decoded using practical iterative decoding algorithms, is known to be close to Shannon limits for codes with suitably large blocklengths. A substantial limitation to the use of finite-length LDPC codes is the presence of an error floor in the low frame error rate (FER) region. This paper develops a deterministic method of predicting error floors, based on high signal-to-noise ratio (SNR) asymptotics, applied to absorbing sets within structured LDPC codes. The approach is illustrated using a class of array-based LDPC codes, taken as exemplars of high-performance structured LDPC codes. The results are in very good agreement with a stochastic method based on importance sampling which, in turn, matches the hardware-based experimental results. The importance sampling scheme uses a mean-shifted version of the original Gaussian density, appropriately centered between a codeword and a dominant absorbing set, to produce an unbiased estimator of the FER with substantial computational savings over a standard Monte Carlo estimator. Our deterministic estimates are guaranteed to be a lower bound to the error probability in the high SNR regime, and extend the prediction of the error probability to as low as 10-30. By adopting a channel-independent viewpoint, the usefulness of these results is demonstrated for both the standard Gaussian channel and a channel with mixture noise. Lara Dolecek, Pamela Lee, Zhengya Zhang, Venkat Anantharam, Borivoje Nikolic, Martin J. Wainwright |
IEEE J. Sel. Areas Commun. | 5 |
| 2009 | Design of LDPC decoders for improved low error rate performance: quantization and algorithm choicesabstractMany classes of high-performance low-density parity-check (LDPC) codes are based on parity check matrices composed of permutation submatrices. We describe the design of a parallel-serial decoder architecture that can be used to map any LDPC code with such a structure to a hardware emulation platform. High-throughput emulation allows for the exploration of the low bit-error rate (BER) region and provides statistics of the error traces, which illuminate the causes of the error floors of the (2048, 1723) Reed-Solomon based LDPC (RS-LDPC) code and the (2209, 1978) array-based LDPC code. Two classes of error events are observed: oscillatory behavior and convergence to a class of non-codewords, termed absorbing sets. The influence of absorbing sets can be exacerbated by message quantization and decoder implementation. In particular, quantization and the log-tanh function approximation in sum-product decoders Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
IEEE Trans. Commun. | 3 |
| 2008 | Lowering LDPC Error Floors by PostprocessingabstractA class of combinatorial structures, called absorbing sets, strongly influences the performance of low-density parity-check (LDPC) decoders at low error rates. Past experiments have shown that a class of (8,8) absorbing sets determines the error floor performance of the (2048,1723) Reed-Solomon based LDPC code (RS-LDPC). A postprocessing approach is formulated to exploit the structure of the absorbing set by biasing the reliabilities of selected messages in a message-passing decoder. The approach converges quickly and can be efficiently implemented with minimal overhead. Hardware emulation of the decoder with postprocessing shows more than two orders of magnitude improvement in the very low bit error rate performance and error- floor-free operation below a BER of 10-12. Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
GLOBECOM | 3 |
| 2008 | Error floors in LDPC codes: Fast simulation, bounds and hardware emulationabstractAbstract — The error-correcting performance of low-density parity check (LDPC) codes, when decoded using practical iterative decoding, is known to approach Shannon limits in the asymptotic limit of large blocklengths. A substantial limitation to the use of finite-length LDPC codes is the presence of an error floor in the low frame error rate (FER) region. This paper develops a method, based on importance sampling and high SNR asymptotics as applied to suitably defined absorbing structures within the LDPC code, to predict error floors. Our results are in very close agreement with hardware-based experimental results, and moreover extend the prediction of the error probability to even lower regions. We compute both importance sampling estimates of error probabilities and deterministic estimates that are guaranteed to lower bound the error probability in the high SNR regime. I. Pamela Lee, Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Borivoje Nikolic, Martin J. Wainwright |
ISIT | 5 |
| 2007 | Analysis of Absorbing Sets for Array-Based LDPC CodesabstractLow density parity check codes (LDPC) are known to perform very well under iterative decoding. However, these codes also exhibit a change in the slope of the bit error rate (BER) vs. signal to noise ratio (SNR) curve in the very low BER region. In our earlier work using hardware emulation in this deep BER regime we argue that this behavior can be attributed to specific structures within the Tanner graph associated with an LDPC code, called absorbing sets. In this paper we provide a detailed theoretical analysis of absorbing sets for array-based LDPC codes Cp.gamma. Specifically, we identify and enumerate all the smallest absorbing sets for these array-based LDPC codes with gamma = 2,3,4 with standard parity check matrix. Experiments carried out on the emulation platform show excellent agreement with our theoretical results. Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Martin J. Wainwright, Borivoje Nikolic |
ICC | 5 |
| 2007 | Quantization Effects in Low-Density Parity-Check DecodersabstractA. class of combinatorial structures, called absorbing sets, strongly influences the performance of low-density parity- check (LDPC) decoders. In particular, the quantization scheme strongly affects which absorbing sets dominate in the error-floor region. Absorbing sets may be characterized as weak or strong. They are a characteristic of the parity check matrix of a code. Conventional quantization schemes applied to a (2209,1978) array-based LDPC code can induce low-weight weak absorbing sets and, as a result, elevate the error floor. Adaptive quantization schemes alleviate the effects of weak absorbing sets, and, as a result, only the strong ones dominate the error floor of an optimized decoder implementation. Another benefit of an adaptive quantization scheme is that it performs well even in very few iterations. Zhengya Zhang, Lara Dolecek, Martin J. Wainwright, Venkat Anantharam, Borivoje Nikolic |
ICC | 5 |
| 2006 | Investigation of Error Floors of Structured Low-Density Parity-Check Codes by Hardware EmulationabstractSeveral high performance LDPC codes have parity-check matrices composed of permutation submatrices. We design a parallel-serial architecture to map the decoder of any structured LDPC code in this large family to a hardware emulation platform. A peak throughput of 240 Mb/s is achieved in decoding the (2048,1723) Reed-Solomon based LDPC (RS-LDPC) code. Experiments in the low bit error rate (BER) region provide statistics of the error traces, which are used to investigate the causes of the error floor. In a low precision implementation, the error floors are dominated by the fixed-point decoding effects, whereas in a higher precision implementation the errors are attributed to special configurations within the code, whose effect is exacerbated in a fixed-point decoder. This new characterization leads to an improved decoding strategy and higher performance. Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
GLOBECOM | 3 |
| 2006 | Power and Area Efficient VLSI Architectures for Communication Signal ProcessingabstractA methodology for VLSI realization of signal processing algorithms for wireless communications is presented that optimizes architecture for reduced power and area. When power is limited, optimal architecture represents a point on the best power-area tradeoff curve that is obtained by balancing the algorithm throughput with the power-performance tradeoff of the underlying building blocks. Architectural optimization is done in the graphical Matlab/Simulink environment, which is also used for algorithm verification. Hardware description language produced by Simulink enables algorithm emulation on the FPGA and also serves as design entry for the chip realization. This is illustrated on complex multi-dimensional algorithms such as wideband MIMO channel decoupling through singular value decomposition (SVD) using 16 sub-carriers. Dejan Markovic, Borivoje Nikolic, Robert W. Brodersen |
ICC | 2 |
| 2006 | L. Embedding Mixed-Signal Design in Systems-on-ChipabstractWith semiconductor technology feature size scaling below 100 nm, mixed-signal design faces some important challenges, caused among others by reduced supply voltages, process variation, and declining intrinsic device gains. Addressing these challenges requires innovative solutions, at the technology, circuit, architecture, and design-methodology level. We present some of these solutions, including a structured platform-based design methodology to enable a meaningful exploration of the broad design space and to classify potential solutions in terms of the relevant metrics. Jan M. Rabaey, Fernando De Bernardinis, Ali M. Niknejad, Borivoje Nikolic, Alberto L. Sangiovanni-Vincentelli |
Proc. IEEE | 4 |
| 2005 | FinFET-based SRAM designabstractIntrinsic variations and challenging leakage control in today's bulk-Si MOSFETs limit the scaling of SRAM. Design tradeoffs in six-transistor (6-T) and four-transistor (4-T) SRAM cells are presented in this work. It is found that 6-T and 4-T FinFET-based SRAM cells designed with built-in feedback achieve significant improvements in the cell static noise margin (SNM) without area penalty. Up to 2x improvement in SNM can be achieved in 6-T FinFET-based SRAM cells. A 4-T FinFET-based SRAM cell with built-in feedback can achieve sub-100pA per-cell standby current and offer the similar improvements in SNM as the 6-T cell with feedback, making them attractive for low-power, low-voltage applications Zheng Guo 0004, Sriram Balasubramanian, Radu Zlatanovici, Tsu-Jae King Liu, Borivoje Nikolic |
ISLPED | 5 |
| 2004 | Low-density parity-check code constructions for hardware implementationabstractWe present several hardware architectures to implement low-density parity-check (LDPC) decoders for codes constructed with a hierarchical structure. The proposed hierarchical formulation of the LDPC code allows a structured hardware realization of the decoder. For a fully-parallel implementation, there is a reduced routing congestion that allows implementations for blocks sizes up to 1024 bits in 0.13/spl mu/m technology. Partially and fully serial implementations benefits greatly from the structure of the code as well, leading to several flexible, efficient architectures. In a general purpose 0.13/spl mu/m technology, the approximate area required by a 1024-bit fully-parallel LDPC decoder is found to be 12.5 mm/sup 2/ while a serial decoder can be implemented in an area of 0.15 mm/sup 2/. Edward Liao, Engling Yeo, Borivoje Nikolic |
ICC | 3 |
| 2004 | Level conversion for dual-supply systemsabstractDual-supply voltage design using a clustered voltage scaling (CVS) scheme is an effective approach to reduce chip power. The optimal CVS design relies on a level converter implemented in a flip-flop to minimize energy, delay, and area penalties due to level conversion. Additionally, circuit robustness against supply bounce is a key property that differentiates good level converter design. Novel flip-flops presented in this paper incorporate a half-latch level converter and a precharged level converter. These flip-flops are optimized in the energy-delay design space to achieve over 30% reduction of energy-delay product and about 10% savings of total power in a CVS design as compared to the conventional flip-flop. These benefits are accompanied by 24% flip-flop robustness improvement leading to 13% delay spread reduction in a CVS critical path. The proposed flip-flops also show 18% layout area reduction. Advantages of level conversion in a flip-flop over asynchronous level conversion in combinational logic are also discussed in terms of delay penalty and its sensitivity to supply bounce. Fujio Ishihara, Farhana Sheikh, Borivoje Nikolic |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Level conversion for dual-supply systemsabstractDual-supply voltage design using a clustered voltage scaling (CVS) scheme is an effective approach to reduce chip power. The optimal CVS design relies on a level converter (LC) implemented in a flip-flop to minimize energy, delay, and area penalties due to level conversion. Novel flip-flops presented in this paper incorporate a half-latch LC and a precharged LC. These flip-flops are optimized in the energy-delay design space to achieve over 30% reduction of energy-delay product and about 10% savings of total power in a CVS design as compared to the conventional flip-flop. These benefits are accompanied by 24% robustness improvement and 18% layout area reduction. Fujio Ishihara, Farhana Sheikh, Borivoje Nikolic |
ISLPED | 3 |
| 2002 | Methods for true power minimizationabstractThis paper presents methods for efficient power minimization at circuit and micro-architectural levels. The potential energy savings are strongly related to the energy profile of a circuit. These savings are obtained by using gate sizing, supply voltage, and threshold voltage optimization, to minimize energy consumption subject to a delay constraint. The true power minimization is achieved when the energy reduction potentials of all tuning variables are balanced. We derive the sensitivity of energy to delay for each of the tuning variables connecting its energy saving potential to the physical properties of the circuit. This helps to develop understanding of optimization performance and identify the most efficient techniques for energy reduction. The optimizations are applied to some examples that span typical circuit topologies including inverter chains, SRAM decoders, and adders. At a delay of 20% larger than the minimum, energy savings of 40% to 70% are possible, indicating that achieving peak performance is expensive in terms of energy. Energy savings of about 50% can be achieved without delay penalty with the balancing of sizes, supplies, and thresholds. Robert W. Brodersen, Mark Horowitz, Dejan Markovic, Borivoje Nikolic, Vladimir Stojanovic |
ICCAD | 4 |
| 2002 | Combining Dual-Supply, Dual-Threshold and Transistor Sizing for Power ReductionabstractMultiple supply voltages, multiple transistor thresholds and transistor sizing could be used to reduce the power dissipation of digital blocks. This paper presents a framework for evaluating the effectiveness of each of these approaches independently and in conjunction with each other. Results show the advantages of multiple supply, transistor sizing, and multiple threshold can be compounded to maximize power reduction. The order of application of these techniques determines the final savings in active and leakage power. Stephanie Augsburger, Borivoje Nikolic |
ICCD | 2 |
| 2001 | Achieving 550Mhz in an ASIC MethodologyabstractTypically, good automated ASIC designs may be two to five times slower than handcrafted custom designs. At last year's DAC this was examined and causes of the speed gap between custom circuits and ASICs were identified. In particular, faster custom speeds are achieved by a combination of factors: good architecture with well-balanced pipelines; compact logic design; timing overhead minimization; careful floorplanning, partitioning and placement; dynamic logic; post-layout transistor and wire sizing; and speed binning of chips. Closing the speed gap requires improving these same factors in ASICs, as far as possible. In this paper we examine a practical example of how these factors may be improved in ASICs. In particular we show how techniques commonly found in custom design were applied to design a high-speed 550 MHz disk drive read channel in an ASIC design flow. David G. Chinnery, Borivoje Nikolic, Kurt Keutzer |
DAC | 2 |
| 2001 | Design methodology for PicoRadio networksabstractOne of the most compelling challenges of the next decade is the "last-meter" problem, extending the expanding data network into end-user data-collection and monitoring devices. PicoRadio supports the assembly of an ad hoc wireless network of self-contained mesoscale, low-cost, low-energy sensor and monitor nodes. While technology advances have made it conceivable to deploy wireless networks of heterogeneous nodes, the design of a low-power, low-cost, adaptive node in a reduced time to market is still a challenge. We present a design methodology for PicoRadio Networks, from system conception and optimization to silicon platform implementation. For each phase of the design, we demonstrate the applicability of our methodology through promising experimental results. Julio Leao da Silva Jr., J. Shamberger, M. Josie Ammer, Chunlong Guo, Suet-Fei Li, Rahul C. Shah, Tim Tuan, Michael Sheets, Jan M. Rabaey, Borivoje Nikolic, Alberto L. Sangiovanni-Vincentelli, Paul K. Wright |
DATE | 10 |
| 2001 | List Viterbi decoding with continuous error detection for magnetic recordingabstractThe list Viterbi algorithm (LVA) with an arithmetic coding based continuous error detection (CED) scheme is applied to high-order partial-response magnetic recording channels. Commonly used magnetic recording systems employ distance-enhancing codes along with parity-check post-processors to correct most dominant error events. The system presented here inserts a CED encoder between the outer Reed-Solomon (RS) code and the channel, and utilizes LVA along with CED's error detection capabilities to improve the performance of the Viterbi decoder. Simulations show that this system results in a 2 dB improvement over MTR encoded EEPRML decoding at a BER of 2/spl times/10/sup -6/ in additive white Gaussian noise and localizes error occurrences to the end of the sector. Dragan Petrovic, Borivoje Nikolic, Kannan Ramchandran |
GLOBECOM | 2 |
| 2001 | High throughput low-density parity-check decoder architecturesabstractTwo decoding schedules and the corresponding serialized architectures for low-density parity-check (LDPC) decoders are presented. They are applied to codes with parity-check matrices generated either randomly or using geometric properties of elements in Galois fields. Both decoding schedules have low computational requirements. The original concurrent decoding schedule has a large storage requirement that is dependent on the total number of edges in the underlying bipartite graph, while a new, staggered decoding schedule which uses an approximation of the belief propagation, has a reduced memory requirement that is dependent only on the number of bits in the block. The performance of these decoding schedules is evaluated through simulations on a magnetic recording channel. Engling Yeo, Payam Pakzad, Borivoje Nikolic, Venkat Anantharam |
GLOBECOM | 3 |
| 2001 | Analysis and design of low-energy flip-flopsabstractThis paper develops a methodology for selecting and optimizing flip-flops for low-energy systems with constant throughput. Characterization metrics, relevant to low-energy systems are discussed, providing insight into timing and energy parameters at both the circuit and system levels. Transistor sizes are optimized for minimal delay under constrained energy consumption. This methodology is applied to characterization of various flip-flop styles and their comparison in 0.25µm CMOS technology under scaled supply voltages. A transmission-gate master-slave latchpair has the largest internal race margin, lowest energy consumption, and has energy-delay product comparable to much faster pulse-triggered latches. Dejan Markovic, Borivoje Nikolic, Robert W. Brodersen |
ISLPED | 2 |
| 2001 | Reduced complexity sequence detection for high-order partial response channelsabstractDetector hardware complexity of high-order partial response magnetic read channels is a major obstacle to high data rate operation and reduced area and power consumption. The method presented here reduces the complexity of single-step and two-step implementations of the Viterbi detector by applying a distance-enhancing code that eliminates some states from the code trellis. The complexity of the detector is further reduced by eliminating less-probable branches from the trellis. This is accomplished by a simple control mechanism that uses the signs of the consecutive input samples. The reduced set of add-compare-select (ACS) units is dynamically assigned to the detector states, decreasing the complexity of the Viterbi detector by roughly 50%. This method is demonstrated on high-order partial response systems with the E/sup 2/PR4 target and an 11-level/32-state target. The simulation results show negligible bit error rate (BER) degradation for signal-to-noise ratios In the range of operation of contemporary disk drive read channels. Michael Leung 0001, Borivoje Nikolic, Leo Ki-Chun Fu, Taehyun Jeon |
IEEE J. Sel. Areas Commun. | 2 |
| 2000 | Clocked CMOS adiabatic logic with integrated single-phase power-clock supplyabstractThe design and experimental evaluation of a clocked adiabatic logic (GAL) is described in this paper. CAL is a dual-rail logic that operates from a single-phase AC power-clock supply. This new low-energy logic makes it possible to integrate all power control circuitry on the chip, resulting in better system efficiency, lower cost, and simpler power distribution. CAL can also be operated from a DC power supply in a nonenergy-recovery mode compatible with standard CMOS logic. In the adiabatic mode, the power-clock supply waveform is generated using an on-chip switching transistor and a small external inductor between the chip and a low-voltage DC supply. Circuit operation and performance are evaluated using a chain of inverters realized in a 1.2 /spl mu/m CMOS technology. Experimental results show that energy savings are achieved at clock frequencies up to about 40 MHz as compared to the nonadiabatic mode. Since CAL can operate both in adiabatic and nonadiabatic modes, power management strategies may be based upon switching between modes when necessary. Dragan Maksimovic, Vojin G. Oklobdzija, Borivoje Nikolic, K. Wayne Current |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | A rate 8/9 sliding block trellis code with stationary detector for magnetic recordingabstractA new trellis code for partial-response magnetic recording channels, that eliminates most frequent errors while keeping the code trellis invariant in time, is proposed. Previously published codes either had lower code density or resulted in time varying trellises with a period of 9, thus requiring a higher complexity of detectors. The new code introduces dependency between codewords to achieve the same coding constraints as an 8/9 code, with the same code density, resulting in a trellis that has a period of two. The new trellis code eliminates two states in every second step of the E/sup 2/PR4 trellis, requiring a 14-state two-step sequence detector. Borivoje Nikolic, Michael Leung 0001, Leo Ki-Chun Fu |
ICC | 1 |
| 1997 | Clocked CMOS adiabatic logic with integrated single-phase power-clock supply: experimental resultsabstractIn this paper we describe the design and experimental evaluation of a clocked CMOS adiabatic logic (CAL). CAL is a dual-rail logic that operates from a single-phase AC power-clock supply in the 'adiabatic' mode, or from a DC power supply in the 'non-adiabatic' mode. In the adiabatic mode, the power-clock supply waveform is generated using an on-chip switching transistor and a small external inductor between the chip and a low-voltage DC supply. Circuit operation and performance are evaluated using a chain of inverters realized in 1.2 /spl mu/m technology. Experimental results show energy savings in the adiabatic mode versus the non-adiabatic mode at clock frequencies up to about 40 MHz. Dragan Maksimovic, Vojin G. Oklobdzija, Borivoje Nikolic, K. Wayne Current |
ISLPED | 3 |