EDBT 2026 Demo / reviewers in the wild / expert
H.-S. Philip Wong
dblp:48/6697 · also Hon-Sum Philip Wong
· DBLP profile ↗
60ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0002-0096-1472ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 6Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultrafast Generative AI by Ultradense 3D Integration: A Case Study on LLM-based Edge InferenceabstractGenerative AI (GenAI) is one of the most critical applications today, continually challenging the limits of semiconductor technology. We introduce a very fine-grained 3D memory-on-logic architecture along with a novel data mapping strategy to support Large Language Model (LLM)-based GenAI, including both prefill and generation stages. Our conceptual analysis shows how ultradense 3D connectivity can enhance text generation speed and energy-efficiency well-beyond current limits. Preliminary findings from a basic analytical model indicate that the single batch autoregressive generation rate for Llama 3.2 1B could surpass 5K tokens/sec by maximizing weight locality and enhancing memory bandwidth through massively parallel 3D links between Multiply-Accumulate (MAC) units in the logic tier and their dedicated memory partitions in the 3D stack. We also explore the impact of advanced logic nodes and quantify their benefits in reducing prefill latency. Finally, we examine the challenges associated with memory access power and power density under extreme bandwidth conditions and present pipelined access strategies to address them. Kerem Akarvardar, Xiaoyu Sun 0001, Brian Crafton, Xiaochen Peng, Haruki Mori, Abhiroop Bhattacharjee, Hidehiro Fujiwara, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2024 | Bitwise Adaptive Early Termination in Hyperdimensional Computing InferenceabstractHyperdimensional computing (HDC), a powerful paradigm for cognitive tasks, often demands hypervectors of high dimensions (e.g., 10,000) to achieve competitive accuracy. However, processing such large-dimensional data poses challenges for performance and energy efficiency, particularly on resource-constrained devices. In this paper, We present a framework to terminate bit-serial HDC inference early when sufficient confidence is attained in the prediction. This approach integrates a Naive Bayes model to replace the conventional associative memory in HDC. This transformation allows for a probabilistic interpretation of the model outputs, steering away from mere similarity measures. We reduce more than 70% of bits that need to be processed while maintaining comparable accuracy across diverse benchmarks. In addition, We show the adaptability of our early termination algorithm during on-the-fly learning scenarios. Wei-Chen Chen, H.-S. Philip Wong, Sara Achour |
DAC | 2 |
| 2024 | Efficient Open Modification Spectral Library Searching in High-Dimensional Space with Multi-Level-Cell MemoryabstractOpen Modification Search (OMS) is a promising algorithm for mass spectrometry analysis that enables the discovery of modified peptides. However, OMS encounters challenges as it exponentially extends the search scope. Existing OMS accelerators either have limited parallelism or struggle to scale effectively with growing data volumes. In this work, we introduce an OMS accelerator utilizing multi-level-cell (MLC) RRAM memory to enhance storage capacity by 3x. Through in-memory computing, we achieve up to 77x faster data processing with two to three orders of magnitude better energy efficiency. Testing was done on a fabricated MLC RRAM chip. We leverage hyperdimensional computing to tolerate up to 10% memory errors while delivering massive parallelism in hardware. Keming Fan, Wei-Chen Chen, Sumukh Pinge, H.-S. Philip Wong, Tajana Rosing |
DAC | 4 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 14 |
| 2023 | Technology Prospects for Data-Intensive ComputingabstractFor many decades, progress in computing hardware has been closely associated with CMOS logic density, performance, and cost. As such, slowdown in 2-D scaling, frequency saturation in CPUs, and increased cost of design and chip fabrication for advanced technology nodes since the early 2000s have led to concerns about how semiconductor technology may evolve in the future. However, the last two decades have also witnessed a parallel development in the application landscape: the advent of big data and consequent rise of data-intensive computing, using techniques such as machine learning. In this article, we advance the idea that data-intensive computing would further cement semiconductor technology as a foundational technology with multidimensional pathways for growth. Continued progress of semiconductor technology in this new context would require the adoption of a system-centric perspective to holistically harness logic, memory, and packaging resources. After examining the performance metrics for data-intensive computing, we present the historical trends for general-purpose graphics processing unit (GPGPU) as a representative data-intensive computing hardware. Thereon, we estimate the values of the key data-intensive computing parameters for the next decade, and our projections may serve as a precursor for a dedicated technology roadmap. By analyzing the compiled data, we identify and discuss specific opportunities and challenges for data-intensive computing hardware technology. Kerem Akarvardar, H.-S. Philip Wong |
Proc. IEEE | 2 |
| 2023 | Micro/Nano Circuits and Systems Design and Design Automation: Challenges and OpportunitiesabstractThe field of design and design automation of micro-/nano-circuits and systems has played a pivotal role in advancing information technologies that are an inseparable part of all our lives. Without the fundamental principles and tools created in this field, modern-day electronic systems that form the foundations of today's information age would not be a reality. Though the field has achieved tremendous success in the past few decades, it is now facing some unprecedented challenges, stemming from foundational technologies all the way to new applications. Business-as-usual approaches are plateauing. New, fundamental research and innovation are needed to sustain the demanded growth. This paper aims to summarize the key challenges and future research directions in the field of micro/nano circuits and systems design and design automation. Gert Cauwenberghs, Jason Cong, Xiaobo Sharon Hu, Siddharth Joshi 0001, Subhasish Mitra, Wolfgang Porod, H.-S. Philip Wong |
Proc. IEEE | 7 |
| 2023 | Neural Network Compression for Noisy Storage DevicesabstractCompression and efficient storage of neural network (NN) parameters is critical for applications that run on resource-constrained devices. Despite the significant progress in NN model compression, there has been considerably less investigation in the actual physical storage of NN parameters. Conventionally, model compression and physical storage are decoupled, as digital storage media with error-correcting codes (ECCs) provide robust error-free storage. However, this decoupled approach is inefficient as it ignores the overparameterization present in most NNs and forces the memory device to allocate the same amount of resources to every bit of information regardless of its importance. In this work, we investigate analog memory devices as an alternative to digital media – one that naturally provides a way to add more protection for significant bits unlike its counterpart, but is noisy and may compromise the stored model’s performance if used naively. We develop a variety of robust coding strategies for NN weight storage on analog devices, and propose an approach to jointly optimize model compression and memory resource allocation. We then demonstrate the efficacy of our approach on models trained on MNIST, CIFAR-10, and ImageNet datasets for existing compression techniques. Compared to conventional error-free digital storage, our method reduces the memory footprint by up to one order of magnitude, without significantly compromising the stored model’s accuracy. Berivan Isik, Kristy Choi, Xin Zheng 0013, Tsachy Weissman, Stefano Ermon, H.-S. Philip Wong, Armin Alaghi |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2020 | A Density Metric for Semiconductor Technology [Point of View]abstractSince its inception, the semiconductor industry has used a physical dimension (the minimum gate length of a transistor) as a means to gauge continuous technology advancement. This metric is all but obsolete today. As a replacement, we propose a density metric, which aims to capture how advances in semiconductor device technologies enable system-level benefits. The proposed metric can be used to gauge advances in future generations of semi-conductor technologies in a holistic way, by accounting for the progress in logic, memory, and packaging/integration technologies simultaneously. H.-S. Philip Wong, Kerem Akarvardar, Dimitri A. Antoniadis, Jeffrey Bokor, Chenming Hu, Tsu-Jae King Liu, Subhasish Mitra, James D. Plummer, Sayeef S. Salahuddin |
Proc. IEEE | 1 |
| 2020 | Scanning the IssueabstractThis month’s issue offers insight into efficient compression and execution of DNNs, the challenge of connecting rural areas, and the clique problem in wireless communication. which H.-S. Philip Wong, Kerem Akarvardar, Dimitri A. Antoniadis, Jeffrey Bokor, Chenming Hu, Tsu-Jae King Liu, Subhasish Mitra, James D. Plummer, Sayeef S. Salahuddin, Lei Deng 0003, Song Han 0003, Luping Shi, Yuan Xie 0001, Elias Yaacoub, Mohamed-Slim Alouini, Ahmed Douik, Hayssam Dahrouj, Tareq Y. Al-Naffouri |
Proc. IEEE | 1 |
| 2019 | On-Chip Memory Technology Design Space Explorations for Mobile Deep Neural Network AcceleratorsabstractDeep neural network (DNN) inference tasks have become ubiquitous workloads on mobile SoCs and demand energy-efficient hardware accelerators. Mobile DNN accelerators are heavily area-constrained, with only minimal on-chip SRAM, which results in heavy use of inefficient off-chip DRAM. With diminishing returns from conventional silicon technology scaling, emerging memory technologies that offer better area density than SRAM can boost accelerator efficiency by minimizing costly off-chip DRAM accesses. This paper presents a detailed design space exploration (DSE) of technology-system co-design for systolic-array accelerators. We focus on practical/mature on-chip memory technologies, including SRAM, eDRAM, MRAM, and 3D vertical RRAM (VRRAM). The DSE employs state-of-the-art optimizations (e.g., model compression and optimized buffer scheduling), and evaluates results on important models including ResNet-50, MobileNet, and Faster-RCNN. Compared to an SRAM/DRAM baseline, MRAM-based accelerators show up to 4.68× energy benefits (57% area overhead), while a 3D VRRAM-based design achieves 2.22× energy benefits (33% area reduction). Haitong Li, Mudit Bhargava, Paul N. Whatmough, H.-S. Philip Wong |
DAC | 4 |
| 2019 | IC Technology - What Will the Next Node Offer Us?abstractThis article consists of a collection of slides from the author's conference presentation. H.-S. Philip Wong |
Hot Chips Symposium | 1 |
| 2019 | The N3XT Approach to Energy-Efficient Abundant-Data ComputingabstractThe world's appetite for analyzing massive amounts of structured and unstructured data has grown dramatically. The computational demands of these abundant-data applications, such as deep learning, far exceed the capabilities of today's computing systems and are unlikely to be met with isolated improvements in transistor or memory technologies, or integrated circuit architectures alone. To achieve unprecedented functionality, speed, and energy efficiency, one must create transformative nanosystems whose architectures are based on the salient properties of the underlying nanotechnologies. Our Nano-Engineered Computing Systems Technology (N3XT) approach makes such nanosystems possible through new computing system architectures leveraging emerging device (logic and memory) nanotechnologies and their dense 3-D integration with fine-grained connectivity to immerse computing in memory and new logic devices (such as carbon nanotube field-effect transistors for implementing high-speed and low-energy logic circuits) as well as high-density nonvolatile memory (such as resistive memory), and amenable to ultradense (monolithic) 3-D integration of thin layers of logic and memory devices that are fabricated at low temperature. In addition, we explore the use of several device and integration technologies in the N3XT beyond the specific ones mentioned earlier that are also used in our main nanosystem prototypes. We also present an efficient resiliency technique to overcome endurance challenges in certain resistive memory technologies. N3XT hardware prototypes demonstrate the practicality of our architectures. We evaluate the benefits of the N3XT using a simulation framework calibrated using experimental measurements. System-level energy-delay product of common implementations of abundant-data workloads improves by three orders of magnitude in the N3XT compared with conventional architectures. These improvements impact a broad range of application workloads and architecture configurations, from embedded systems to the cloud. Mohamed M. Sabry, Tony F. Wu, Andrew Bartolo, Yash H. Malviya, William Hwang, Gage Hills, Igor L. Markov, Mary Wootters, Max M. Shulaker, H.-S. Philip Wong, Subhasish Mitra |
Proc. IEEE | 10 |
| 2018 | TRIG: hardware accelerator for inference-based applications and experimental demonstration using carbon nanotube FETsabstractThe energy efficiency demands of future abundant-data applications, e.g., those which use inference-based techniques to classify large amounts of data, exceed the capabilities of digital systems today. Field-effect transistors (FETs) built using nanotechnologies, such as carbon nanotubes (CNTs), can improve energy efficiency significantly. However, carbon nanotube FETs (CNFETs) are subject to process variations inherent to CNTs: variations in CNT type (semiconductor or metallic), CNT density, or CNT diameter, to name a few. These CNT variations can degrade CNFET benefits at advanced technology nodes. One path to overcome CNT variations is to co-optimize CNT processing and CNFET circuit design; however, the required CNT process advancements have not been achieved experimentally. We present a new design approach (TRIG, Technique for Reducing errors using Iterative Gray code) to overcome process variations in hardware accelerators targeting inference-based applications that use serial matrix operations (serial: accumulated over at least 2 clock cycles). We demonstrate that TRIG can retain the major energy efficiency benefits (quantified using Energy Delay Product or EDP) of CNFETs despite CNT variations that exist in today's CNFET fabrication - without requiring further CNT processing improvements to overcome CNT variations. As a case study, we analyze the effectiveness of TRIG for a binary neural network hardware accelerator that classifies images. Despite CNT variations that exist today, TRIG can maintain 99% (90%) of projected EDP benefits of CNFET digital circuits for 90% (99%) image classification accuracy target. We also demonstrate experimentally fabricated CNFET circuits to compute scalar product (a common matrix operation, also called dot product), with and without TRIG: TRIG reduces the mean difference between the expected result (no errors) and the experimentally computed result by 30× in the presence of CNT variations, shown experimentally. Gage Hills, Daniel Bankman, Bert Moons, Lita Yang, Jake Hillard, Alex Kahng, Rebecca Park, Marian Verhelst, Boris Murmann, Max M. Shulaker, H.-S. Philip Wong, Subhasish Mitra |
DAC | 11 |
| 2018 | Joint Source-Channel Coding with Neural Networks for Analog Data Compression and StorageabstractWe provide an encoding and decoding strategy for efficient storage of analog data onto an array of Phase-Change Memory (PCM) devices. The PCM array is treated as an analog channel, with the stochastic relationship between write voltage and read resistance for each device determining its theoretical capacity. The encoder and decoder are implemented as neural networks with parameters that are trained end-to-end to minimize distortion for a fixed number of devices. To minimize distortion, the encoder and decoder must adapt jointly to the statistics of images and the statistics of the channel. Similar to Balle et al. (2017), we find that incorporating divisive normalization in the encoder, paired with de-normalization in the decoder, improves model performance. We show that the autoencoder achieves a rate-distortion performance above that achieved by a separate JPEG source coding and binary channel coding scheme. These results demonstrate the feasibility of exploiting the full analog dynamic range of PCM or other emerging memory devices for efficient storage of analog image data. Ryan Zarcone, Dylan M. Paiton, Alex Anderson, Jesse H. Engel, H.-S. Philip Wong, Bruno A. Olshausen |
DCC | 5 |
| 2018 | Coming Up N3XT, After 2D Scaling of Si CMOSabstractAs two-dimensional scaling of Si CMOS crosses the nanometer threshold, from 7 nm, 5 nm, 3 nm, toward 1 nm technology nodes, will it continue to provide the energy efficiency required of future computing systems? A scalable, fast, and energy-efficient computation platform that may provide another 1,000× in computing energy efficiency (energy-execution time product) will have massive on-chip memory co-located with highly energy-efficient computing logic, enabled by 3D integration (e.g., monolithic) with ultra-dense and fine-grained connectivity. There will be multiple layers of memories interleaved with computing logic, sensors, and application-specific devices. We call this technology platform N3XT, Nano-engineered Computing Systems Technology. In this paper, we give an overview of the nanoscale memory and logic technologies that enable N3XT. William Hwang, Weier Wan, Subhasish Mitra, H.-S. Philip Wong |
ISCAS | 4 |
| 2017 | In Quest of the Next Information Processing Substrate: Extended Abstract: InvitedabstractConventional CMOS scaling and the Moore's law have been the cornerstone of progress in computing hardware technology. However, with dimensional scaling expected to end soon, there is a pressing need to find the next information processing hardware that can continue to support the technology revolution. Will this hardware solution be an enhanced or an augmented version of MOSFET or a switch based on a radically new switching mechanism. Ultimately, do we require a complete deviation from the Boolean paradigm itself? In this invited paper, we will review some of the actively pursued future logic, merged logic-memory and related concepts. Suman Datta, Alan C. Seabaugh, Michael T. Niemier, Arijit Raychowdhury, Darrell Schlom, Debdeep Jena, Huili Grace Xing, H.-S. Philip Wong, Eric Pop, Sayeef S. Salahuddin, Sumeet Kumar Gupta, Supratik Guha |
DAC | 8 |
| 2017 | A Systems Approach to Computing in Beyond CMOS Fabrics: InvitedabstractNo abstract available. Ameya Patil 0001, Naresh R. Shanbhag, Lav R. Varshney, Eric Pop, H.-S. Philip Wong, Subhasish Mitra, Jan M. Rabaey, Jeffrey A. Weldon, Lawrence T. Pileggi, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young |
DAC | 5 |
| 2016 | TPAD: Hardware Trojan Prevention and Detection for Trusted Integrated CircuitsabstractThere are increasing concerns about possible malicious modifications of integrated circuits (ICs) used in critical applications. Such attacks are often referred to as hardware Trojans. While many techniques focus on hardware Trojan detection during IC testing, it is still possible for attacks to go undetected. Using a combination of new design techniques and new memory technologies, we present a new approach that detects a wide variety of hardware Trojans during IC testing and also during system operation in the field. Our approach can also prevent a wide variety of attacks during synthesis, place-and-route, and fabrication of ICs. It can be applied to any digital system, and can be tuned for both traditional and split-manufacturing methods. We demonstrate its applicability for both application-specified integrated circuits and field-programmable gate arrays. Using fabricated test chips with Trojan emulation capabilities and also using simulations, we demonstrate: 1) the area and power costs of our approach can range between 7.4%-165% and 7%-60%, respectively, depending on the design and the attacks targeted; 2) the speed impact can be minimal (close to 0%); 3) our approach can detect 99.998% of Trojans (emulated using test chips) that do not require detailed knowledge of the design being attacked; 4) our approach can prevent 99.98% of specific attacks (simulated) that utilize detailed knowledge of the design being attacked (e.g., through reverse engineering); and 5) our approach never produces any false positives, i.e., it does not report attacks when the IC operates correctly. Tony F. Wu, Karthik Ganesan 0001, Yunqing Alexander Hu, H.-S. Philip Wong, S. Simon Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Modeling and design optimization of ReRAMabstractResistive switching memories (ReRAM) have been widely studied for applications in next-generation data storage and neurormorphic computing systems. To enable device-circuit-system co-design and optimization, a SPICE model of ReRAM that can reproduce the device characteristics in circuit simulations is needed. In this paper, we present a novel tool for ReRAM design including a physics-based SPICE model, the model parameters extraction strategy, as well as the system assessment method. This physics-based SPICE model can capture all the essential features of HfOx-based ReRAM including the DC/AC and multi-level switching behaviors, switching reliability, and intrinsic device variations. A strategy is developed to extract the critical model parameters from the fabricated ReRAM devices. A variety of electrical measurements on various ReRAMs are performed to verify and calibrate the model. The assessment method based on the experimentally verified SPICE model can be applied to explore a wide range of applications including: 1) variation-aware and reliability-emphasized system design; 2) system performance evaluation; 3) array architecture optimization. This verified design tool not only enables system design but also enables system optimization that capitalizes on device/circuit interaction for both data storage and neuromorphic computing applications. Jinfeng Kang, Haitong Li, Peng Huang 0004, Bin Gao 0006, Zizhen Jiang, H.-S. Philip Wong |
ASP-DAC | 8 |
| 2015 | Contact pitch and location prediction for Directed Self-Assembly template verificationabstractDirected Self-Assembly (DSA) is a promising technique for contacts/vias patterning in 7 nm technology nodes. In DSA process, groups of contact holes/vias are generated by the self-assembly process guided by the `guiding templates'. The guiding templates are patterned by conventional optical lithography process such as 193 nm immersion lithography. As a result, the patterning fidelity and variation in the template shapes is very likely to affect the final contact holes/vias. While feasible in principle, rigorous DSA process simulation is unacceptably slow for full chip verification in practice. This paper proposes a machine learning based verification that can predict the pitch size of the contact holes and the hole centers. Given a set of training data that consists of simulated template and contact hole patterns, our method is able to learn a highly accurate predictive model for pitch size and hole location. To build a statistical model for prediction, we utilize computer vision techniques to extract various geometric and image features. We conduct extensive experiments to explore the effectiveness of the proposed features, and compare several machine learning algorithms to achieve an effective and efficient prediction. The experimental results show that compared to the minutes or even hours of simulation time in rigorous methods, our best prediction model achieves very promising results (RMSE = 0.135 pitch grid) with less than one second of training and predicting runtime overhead. Zigang Xiao, Yuelin Du, Martin D. F. Wong, He Yi, H.-S. Philip Wong, Hongbo Zhang 0001 |
ASP-DAC | 5 |
| 2015 | Layout optimization and template pattern verification for directed self-assembly (DSA)abstractRecently, block copolymer directed self-assembly (DSA) has demonstrated great advantages in patterning contacts/vias for the 7 nm technology node and beyond. The high throughput and low process cost of DSA makes it the most promising candidate in patterning tight pitched dense patterns for the next generation lithography. Since DSA is very sensitive to the shapes and distributions of the guiding templates, it is necessary to develop new EDA algorithms and tools to address the patterning rules and constraints of the process. This paper presents a set of DSA-aware optimization techniques targeting the most urgent problems for DSA technology, including layout optimization and template pattern verification. Zigang Xiao, Daifeng Guo, Martin D. F. Wong, He Yi, Maryann C. Tung, H.-S. Philip Wong |
DAC | 6 |
| 2015 | Variation-aware, reliability-emphasized design and optimization of RRAM using SPICE model
Haitong Li, Zizhen Jiang, Peng Huang 0004, Hong-Yu Chen, Bin Gao 0006, Jinfeng Kang, H.-S. Philip Wong |
DATE | 9 |
| 2015 | Monolithic 3D integration: a path from concept to reality
Max M. Shulaker, Tony F. Wu, Mohamed M. Sabry, Hai Wei, H.-S. Philip Wong, Subhasish Mitra |
DATE | 5 |
| 2015 | Time-based sensor interface circuits in carbon nanotube technologyabstractCarbon nanotube technology is a promising technology to further reduce the energy consumption in electronics, as it is projected to achieve an order of magnitude improvement in energy-delay product compared to Silicon CMOS at highly-scaled technology nodes. In addition, CNTs are excellent candidates to be functionalized as sensors, and can potentially improve the energy efficiency of sensors and sensor interfaces for future autonomy-demanding applications. This paper presents an overview of time-based sensor interfaces implemented in a CNT technology. Time-based sensor interfaces yield highly-digital architectures, allowing for scalable and robust designs. All of the presented CNFET-based sensor interface circuits have been fabricated in a VLSI-compatible manner and have been validated through measurements. Georges Gielen, Jelle Van Rethy, Max M. Shulaker, Gage Hills, H.-S. Philip Wong, Subhasish Mitra |
ISCAS | 5 |
| 2015 | Physical Layout Design of Directed Self-Assembly Guiding Alphabet for IC Contact Hole/via PatterningabstractThe continued scaling of feature size has brought increasingly significant challenges to conventional optical lithography.[1-3] The rising cost and limited resolution of current lithography technologies have opened up opportunities for alternative patterning approaches. Among the emerging patterning approaches, block copolymer self-assembly for device fabrication has been envisioned for over a decade. Block copolymer DSA is a result of spontaneous microphase separation of block copolymer films, forming periodic microdomains including cylinders, spheres, and lamellae, in the same way that snowflakes and clamshells are formed in nature - by self-assembly due to forces of nature (Fig. 1a). DSA can generate closely packed and well controlled sub-20 nm features with low cost and high throughput, therefore stands out among other emerging lithographic solutions, including extreme ultraviolet lithography (EUV), electron beam lithography (e-beam), and multiple patterning lithography (MPL).[2;6] H.-S. Philip Wong, He Yi, Maryann C. Tung, Kye Okabe |
ISPD | 1 |
| 2015 | Rapid Co-Optimization of Processing and Circuit Design to Overcome Carbon Nanotube VariationsabstractCarbon nanotube field-effect transistors (CNFETs) are promising candidates for building energy-efficient digital systems at highly scaled technology nodes. However, carbon nanotubes (CNTs) are inherently subject to variations that reduce circuit yield, increase susceptibility to noise, and severely degrade their anticipated energy and speed benefits. Joint exploration and optimization of CNT processing options and CNFET circuit design are required to overcome this outstanding challenge. Unfortunately, existing approaches for such exploration and optimization are computationally expensive, and mostly rely on trial-and-error-based ad hoc techniques. In this paper, we present a framework that quickly evaluates the impact of CNT variations on circuit delay and noise margin, and systematically explores the large space of CNT processing options to derive optimized CNT processing and CNFET circuit design guidelines. We demonstrate that our framework: 1) runs over 100× faster than existing approaches and 2) accurately identifies the most important CNT processing parameters, together with CNFET circuit design parameters (e.g., for CNFET sizing and standard cell layouts), to minimize the impact of CNT variations on CNFET circuit speed with ≤5% energy cost, while simultaneously meeting circuit-level noise margin and yield constraints. Gage Hills, Jie Zhang 0007, Max M. Shulaker, Hai Wei, Chi-Shuen Lee, Arjun Balasingam, H.-S. Philip Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2014 | Directed Self-Assembly (DSA) Template Pattern VerificationabstractDirected Self-Assembly (DSA) is a promising technique for contacts/vias patterning, where groups of contacts/vias are patterned by guiding templates. As the templates are patterned by traditional lithography, their shapes may vary due to the process variations, which will ultimately affect the contacts/vias even for the same type of template. Due to the complexity of the DSA process, rigorous process simulation is unacceptably slow for full chip verification. This paper formulate several critical problems in DSA verification, and proposes a design automation methodology that consists of a data preparation and a model learning stage. We present a novel DSA model with Point Correspondence and Segment Distance features for robust learning. Following the methodology, we propose an effective machine learning (ML) based method for DSA hotspot detection. The results of our initial experiments have already demonstrated the high-efficiency of our ML-based approach with over 85% detection accuracy. Compared to the minutes or even hours of simulation time in rigorous method, the methodology in this paper validates the research potential along this direction. Zigang Xiao, Yuelin Du, Haitong Tian, Martin D. F. Wong, He Yi, H.-S. Philip Wong, Hongbo Zhang 0001 |
DAC | 6 |
| 2014 | Video analytics using beyond CMOS devicesabstractThe human vision system understands and interprets complex scenes for a variety of visual tasks in real-time while consuming less than 20 Watts of power. The holistic design of artificial vision systems that will approach and eventually exceed the capabilities of human vision systems is a grand challenge. The design of such a system needs advances in multiple disciplines. This paper focuses on advances needed in the computational fabric and provides an overview of a new-genre of architectures inspired by advances in both the understanding of the visual cortex and the emergence of devices with new mechanisms for state computations. Narayanan Vijaykrishnan, Suman Datta, Gert Cauwenberghs, Donald M. Chiarulli, Steven P. Levitan, H.-S. Philip Wong |
DATE | 6 |
| 2014 | Scaling and operation characteristics of HfOx based vertical RRAM for 3D cross-point architectureabstractStacked HfOxbased vertical RRAM with interface engineering for 3D cross-point architecture is fabricated using a cost-effective fabrication process. The excellent performances such as low reset current, fast switching speed, high switching endurance and disturbance immunity, good retention and self-selectivity are demonstrated in the fabricated HfOxbased vertical RRAM devices. The scaling limit and the functionality along with a viable write/read scheme of the presented vertical RRAM are investigated. The experiments show that the pillar electrode thickness and the plane electrode thickness of the vertical RRAM can be scaled down to 3nm and 5nm without significant performance degradation, respectively. Jinfeng Kang, Bin Gao 0006, Peng Huang 0004, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong, Shimeng Yu |
ISCAS | 9 |
| 2014 | Design guidelines for 3D RRAM cross-point architectureabstractDesign guidelines were proposed to evaluate and optimize the 3D RRAM cross-point architecture by a full-size 3D circuit simulation in SPICE. The performance metrics that were evaluated include the write/read margin, access latency, energy consumption per programming, and the density per bit. Different 3D cross-point architecture including the horizontally stacked or the vertically stacked structure were compared in terms of these metrics, revealing the advantages of the vertical RRAM structure. Then the scaling trend of the vertical RRAM based 3D array with respect to the scaling of lateral feature size, vertical electrode thickness and vertical isolation layer thickness were evaluated. The design parameters that affect the scaling trend include the metal interconnect resistance, RRAM on-state cell resistance (or the nonlinearity of the I-V). The design trade-offs are discussed considering those parameters constraints. Shimeng Yu, Yexin Deng, Bin Gao 0006, Peng Huang 0004, Jinfeng Kang, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong |
ISCAS | 10 |
| 2014 | Design considerations of synaptic device for neuromorphic computingabstractHardware implementation of neuromorphic computing is attractive as a computing paradigm beyond the conventional digital Boolean computing. Recently, two-terminal emerging memory devices that show electrically-triggered resistance modulation have been proposed as synaptic devices for neuromorphic computing. The synaptic device candidates include phase change memory (PCM), resistive RAM (RRAM) and conductive bridge RAM (CBRAM), etc. In this paper, we discuss the general design considerations of synaptic devices for plasticity and learning. As a rule of thumb for performance metrics assessment, an ideal synaptic device should have characteristics such as dimension, energy consumption, operation frequency, dynamic range, etc. that are scalable to biological systems with comparable complexity. Shimeng Yu, Duygu Kuzum, H.-S. Philip Wong |
ISCAS | 3 |
| 2014 | System Level Benchmarking with Yield-Enhanced Standard Cell Library for Carbon Nanotube VLSI CircuitsabstractThe quest for technologies with superior device characteristics has showcased Carbon-Nanotube Field-Effect Transistors (CNFET) into limelight. In this work we present physical design techniques to improve the yield of CNFET circuits in the presence of Carbon Nanotube (CNT) imperfections. Various layout schemes are studied for enhancing the yield of CNFET standard cell library. With the help of existing ASIC design flow, we perform system-level benchmarking of CNFET circuits and compare them to CMOS circuits at various technology nodes. With CNFET technology, we observe maximum performance gains for circuits with gate-dominated delays. Averaged across various benchmarks at 16 nm, we report 8× improvement in Energy-Delay-Product (EDP) with CNFET circuits when compared to CMOS counterpart. We also study the performance of a complete OpenRISC processor, where we see 1.5× improvement in EDP over CMOS at 16 nm technology node. Voltage scaling enabled by CNFETs can be explored in the future for further performance benefits. Shashikanth Bobba, Jie Zhang 0007, Pierre-Emmanuel Gaillardon, H.-S. Philip Wong, Subhasish Mitra, Giovanni De Micheli |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2013 | Rapid exploration of processing and design guidelines to overcome carbon nanotube variationsabstractCarbon nanotube field-effect transistors (CNFETs) are promising candidates for building energy-efficient digital systems at highly-scaled technology nodes. However, carbon nanotubes (CNTs) are inherently subject to variations that reduce circuit yield, increase susceptibility to noise, and severely degrade their anticipated energy and speed benefits. Joint exploration and optimization of CNT processing options and CNFET circuit design are required to overcome this outstanding challenge. Unfortunately, existing approaches for such exploration and optimization are computationally expensive, and mostly rely on trial-and-error-based ad-hoc techniques. In this paper, we present a systematic framework which quickly evaluates the impact of CNT variations on circuit delay and noise margin, and automatically explores the large space of CNT processing options to derive optimized CNT processing and CNFET circuit design guidelines. We demonstrate that: 1. Our new framework runs over 100X faster than existing approaches. 2. It accurately identifies the most important CNT processing parameters, together with CNFET circuit sizing, to minimize the impact of CNT variations while meeting circuit-level noise margin constraints. Gage Hills, Jie Zhang 0007, Charles Mackin, Max M. Shulaker, Hai Wei, H.-S. Philip Wong, Subhasish Mitra |
DAC | 6 |
| 2013 | Sacha: the Stanford carbon nanotube controlled handshaking robotabstractLow-power applications, such as sensing, are becoming increasingly important and demanding in terms of minimizing energy consumption, driving the search for new and innovative interface architectures and technologies. Carbon Nanotube FETs (CNFETs) are excellent candidates for further energy reduction, as CNFET-based digital circuits are projected to potentially achieve an order of magnitude improvement in energy-delay product at highly scaled technology nodes. This paper presents an overview of the first demonstration of a complete sub-system, a sensor interface circuit, implemented entirely using CNFETs. The demonstrated sub-system is an all-digital capacitive sensor to digital converter. The CNFET sensor interface is demonstrated by using the CNFET circuitry to interface with a sensor used to control a handshaking robot. Max M. Shulaker, Jelle Van Rethy, Gage Hills, Hong-Yu Chen, Georges Gielen, H.-S. Philip Wong, Subhasish Mitra |
DAC | 6 |
| 2013 | Carbon nanotube circuits: opportunities and challengesabstractCarbon Nanotube Field-Effect Transistors (CNFETs) are excellent candidates for building highly energy-efficient digital systems. However, imperfections inherent in carbon nanotubes (CNTs) pose significant hurdles to realizing practical CNFET circuits. In order to achieve CNFET VLSI systems in the presence of these inherent imperfections, careful orchestration of design and processing is required: from device processing and circuit integration, all the way to large-scale system design and optimization. In this paper, we summarize the key ideas that enabled the first experimental demonstration of CNFET arithmetic and storage elements. We also present an overview of a probabilistic framework to analyze the impact of various CNFET circuit design techniques and CNT processing options on system-level energy and delay metrics. We demonstrate how this framework can be used to improve the energy-delay-product (EDP) of CNFET-based digital systems. Hai Wei, Max M. Shulaker, Gage Hills, Hong-Yu Chen, Chi-Shuen Lee, Luckshitha Liyanage, Jie Zhang 0007, H.-S. Philip Wong, Subhasish Mitra |
DATE | 8 |
| 2013 | Block copolymer directed self-assembly (DSA) aware contact layer optimization for 10 nm 1D standard cell libraryabstractAt the 10 nm technology node, the contact layers of integrated circuits (IC) designs are too dense to be printed by single exposure using 193 nm immersion (193i) lithography. Among all the emerging patterning approaches, block copolymer directed self-assembly (DSA) is a promising candidate with high throughput and low cost for sub-20 nm features. Traditionally, the study of DSA has focused on achieving periodic regular patterns over large area. Realizing that long range order is not needed for patterning irregularly distributed contact holes, we use topographical guiding templates to alter the natural symmetry of block copolymer and achieve controlled irregular DSA patterns. However, DSA patterning must satisfy the overlay accuracy requirements while the guiding templates also need to be printable by conventional lithography. This presents a unique opportunity of DSA patterning and layout design co-optimization for improving the manufacturability of DSA. This paper discusses the DSA-aware contact layer optimization problem for 10 nm 1D standard cell library. For the first time we propose a cost function for each DSA template based on its overlay accuracy performance. Then given a standard cell library, we simultaneously optimize the layouts of every cell, such that the contact layer of any cell in the library can be fully patterned by a set of guiding templates, and the total cost of the templates is minimal. This optimization problem is first proved to be NP-hard and formulated as a Weighted Partial Maximum Satisfiability (MAXSAT) problem, which can be optimally solved with a public SAT solver. Then we propose a bounded approximation algorithm that solves the problem much more efficiently. The experimental results demonstrate that our approach is remarkably promising in practice and validate the proposed optimization problem. Yuelin Du, Daifeng Guo, Martin D. F. Wong, He Yi, H.-S. Philip Wong, Hongbo Zhang 0001, Qiang Ma 0002 |
ICCAD | 5 |
| 2013 | Effect of Wordline/Bitline Scaling on the Performance, Energy Consumption, and Reliability of Cross-Point Memory ArrayabstractThe impact of wordline/bitline metal wire scaling on the write/read performance, energy consumption, speed, and reliability of the cross-point memory array is quantitatively studied for technology nodes down to single-digit nm. The impending resistivity increase in the Cu wires is found to cause significant decrease of both write and read window margins at the regime when electron surface scattering and grain boundary scattering are substantial. At deeply-scaled device dimensions, the wire energy dissipation and wire latency become comparable to or even exceed the intrinsic values of memory cells. The large current density flowing through the wordlines/bitlines raises additional reliability concerns for the cross-point memory array. All these issues are exacerbated at smaller memory resistance values and larger memory array sizes. They thereby impose strict constraints on the memory device design and preclude the realization of large-scale cross-point memory array with minimum feature sizes beyond the 10 nm node. A rethink in the design methodology of cross-point memory to incorporate and mitigate the scaling effects of wordline/bitline is necessary. Possible solutions include the use of memory wires with better conductivity and scalability, memory arrays with smaller partition sizes, and memory elements with larger resistance values and resistance ratios. Jiale Liang, Chih-Wei Stanley Yeh, S. Simon Wong, H.-S. Philip Wong |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2013 | Combinational Logic Design Using Six-Terminal NEM RelaysabstractThis paper presents techniques for designing nanoelectromechanical relay-based logic circuits using six-terminal relays that behave as universal logic gates. With proper biasing, a compact 2-to-1 multiplexer can be implemented using a single six-terminal relay. Arbitrary combinational logic functions can then be implemented using well-known binary decision diagram (BDD) techniques. Compared to a CMOS-style implementation using four-terminal relays, the BDD-based implementation can result in lower area without major impact on performance metrics such as delay, and energy (when the relays are scaled to small dimensions). Although it is possible to implement any combinational circuit with a single mechanical delay, the relay count can be significantly reduced for complex logic functions by allowing multiple mechanical delays. Daesung Lee 0002, W. Scott Lee, Chen Chen 0018, Farzan Fallah, J. Provine, Soogine Chong, John Watkins, Roger T. Howe, H.-S. Philip Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2013 | Impact of III-V and Ge Devices on Circuit PerformanceabstractIII-V and germanium (Ge) field-effect transistors (FETs) have been studied as candidates for post Si CMOS. In this paper, the performance of various digital blocks and static random access memory (SRAM) with different combinations of Si, III-V and Ge devices are studied. SPICE-compatible III-V n-channel FET (nFET) and Ge p-channel FET (pFET) models are developed for the analysis. The delay and energy of the different combinations are estimated and compared. In typical digital design, the driving capability of the nFET and pFET should be matched for optimum noise margin and performance. The combination of III-V nFET with low input capacitance and Ge pFET achieves the best energy-delay performance for many digital logic circuits. The read margin of SRAM is maximized with a Si pass-gate, and an inverter of III-V nFET and Ge pFET. Jeongha Park, Saeroonter Oh, H.-S. Philip Wong, S. Simon Wong |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Nano-Electro-Mechanical (NEM) relays and their application to FPGA routingabstractNano-Electro-Mechanical (NEM) relays are nano-scale switches that can be mechanically actuated by an electrical signal. Unlike conventional CMOS transistors, NEM relays exhibit zero off-state leakage and very sharp on-off transitions. As a result, NEM relays can be potentially used to design highly energy-efficient digital systems. NEM relays are also excellent candidates for programmable routing switches in Field Programmable Gate Arrays (FPGAs) due to their potentially low on-state resistances despite their long mechanical delays. Low-temperature fabrication of NEM relays creates opportunities for their integration on top of silicon CMOS circuits. Hysteresis properties of NEM relays can enable their use as FPGA programmable routing switches without requiring additional routing SRAM cells. In this talk, we will present an overview of NEM relays and their use in digital system design, and discuss design considerations for hybrid CMOS-NEM FPGAs. Chen Chen 0018, W. Scott Lee, J. Provine, Soogine Chong, Roozbeh Parsa, Daesung Lee 0002, Roger T. Howe, H.-S. Philip Wong, Subhasish Mitra |
ASP-DAC | 8 |
| 2012 | Nano-Electro-Mechanical relays for FPGA routing: Experimental demonstration and a design techniqueabstractNano-Electro-Mechanical (NEM) relays are excellent candidates for programmable routing in Field Programmable Gate Arrays (FPGAs). FPGAs that combine CMOS circuits with NEM relays are referred to as CMOS-NEM FPGAs. In this paper, we experimentally demonstrate, for the first time, correct functional operation of NEM relays as programmable routing switches in FPGAs, and their programmability by utilizing hysteresis properties of NEM relays. In addition, we present a technique that utilizes electrical properties of NEM relays and selectively removes or downsizes routing buffers for designing energy-efficient CMOS-NEM FPGAs. Simulation results indicate that such CMOS-NEM FPGAs can achieve 10-fold reduction in leakage power, 2-fold reduction in dynamic power, and 2-fold reduction in area, simultaneously, without application speed penalty when compared to a 22nm CMOS-only FPGA. Chen Chen 0018, W. Scott Lee, Roozbeh Parsa, Soogine Chong, J. Provine, Jeff Watt, Roger T. Howe, H.-S. Philip Wong, Subhasish Mitra |
DATE | 8 |
| 2012 | Metal-Oxide RRAMabstractIn this paper, recent progress of binary metal–oxide resistive switching random access memory (RRAM) is reviewed. The physical mechanism, material properties, and electrical characteristics of a variety of binary metal–oxide RRAM are discussed, with a focus on the use of RRAM for nonvolatile memory application. A review of recent development of large-scale RRAM arrays is given. Issues such as uniformity, endurance, retention, multibit operation, and scaling trends are discussed. H.-S. Philip Wong, Heng-Yuan Lee, Shimeng Yu, Yu-Sheng Chen, Yi Wu 0016, Pang-Shiu Chen, Byoungil Lee, Frederick T. Chen, Ming-Jinn Tsai |
Proc. IEEE | 1 |
| 2012 | Carbon Nanotube Robust Digital VLSIabstractCarbon nanotube field-effect transistors (CNFETs) are excellent candidates for building highly energy-efficient electronic systems of the future. Fundamental limitations inherent to carbon nanotubes (CNTs) pose major obstacles to the realization of robust CNFET digital very large-scale integration (VLSI): 1) it is nearly impossible to guarantee perfect alignment and positioning of all CNTs despite near-perfect CNT alignment achieved in recent years; 2) CNTs can be metallic or semiconducting depending on chirality; and 3) CNFET circuits can suffer from large performance variations, reduced yield, and increased susceptibility to noise. Today's CNT process improvements alone are inadequate to overcome these challenges. This paper presents an overview of: 1) imperfections and variations inherent to CNTs; 2) design and processing techniques, together with a probabilistic analysis framework, for robust CNFET digital VLSI circuits immune to inherent CNT imperfections and variations; and 3) recent experimental demonstration of CNFET digital circuits that are immune to CNT imperfections. Significant advances in design tools can enable robust and scalable CNFET circuits that overcome the challenges of the CNFET technology while retaining its energy-efficiency benefits. Jie Zhang 0007, Albert Lin 0002, Nishant Patil, Hai Wei, H.-S. Philip Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2011 | Carbon nanotube imperfection-immune digital VLSI: Frequently asked questions updatedabstractCarbon Nanotube Field-Effect Transistors (CNFETs) are excellent candidates for designing highly energy-efficient future digital systems. However, carbon nanotubes (CNTs) are inherently highly subject to imperfections that pose major obstacles to robust CNFET digital VLSI. This paper summarizes commonly raised questions and concerns about CNFET technology through a series of frequently asked questions. The specific questions addressed in this paper are motivated by recent advances in the field since the publication of our earlier paper on frequently asked questions in the Proceedings of the 2009 Design Automation Conference. Hai Wei, Jie Zhang 0007, Nishant Patil, Albert Lin 0002, Max M. Shulaker, Hong-Yu Chen, H.-S. Philip Wong, Subhasish Mitra |
ICCAD | 8 |
| 2011 | Characterization and Design of Logic Circuits in the Presence of Carbon Nanotube Density VariationsabstractVariations in the spatial density of carbon nanotubes (CNTs), resulting from the lack of precise control over CNT positioning during chemical synthesis, is a major hurdle to the scalability of carbon nanotube field effect transistor (CNFET) circuits. Such CNT density variations can lead to non-functional CNFET circuits. This paper presents a probabilistic framework for modeling the CNT count distribution contained in a CNFET of given width, and establishes the accuracy of the model using experimental data obtained from CNT growth. Using this model, we estimate the impact of CNT density variations on the yield of CNFET very large-scale integrated circuits. Our estimation results demonstrate that CNT density variations can significantly degrade the yield of CNFETs, and can be a major concern for scaled CNFET circuits. Finally, we analyze the impact of CNT correlation (i.e., correlation of CNT count between CNFETs) that exists in CNT growth, and demonstrate how the yield of a CNFET storage circuit (primarily limited by its noise immunity) can be significantly improved by taking advantage of such correlation. Jie Zhang 0007, Nishant Patil, Arash Hazeghi, H.-S. Philip Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Detachable nano-carbon chip with ultra low powerabstractThis paper describes ultra-low-power chip design using nano-scale electro-mechanical switches (NEMS) with graphene. This chip is attachable and detachable onto the top of other chips due to remarkable stickiness of carbon-nanotube interconnects. New 3D-IC can be thus constructed for reconfigurable system-on-chips. Furthermore, due to a floating gate built in NEMS, their logic performance is much superior to that of NEMS-based logic in previous works, and even better than that of conventional CMOS. Shinobu Fujita, Shinichi Yasuda, Daesung Lee 0002, Deji Akinwande, H.-S. Philip Wong |
DAC | 6 |
| 2010 | Carbon nanotube correlation: promising opportunity for CNFET circuit yield enhancementabstractCarbon Nanotubes (CNTs) are grown using chemical synthesis, and the exact positioning and chirality of CNTs are very difficult to control. As a result, "small-width" Carbon Nanotube Field-Effect Transistors (CNFETs) can have a high probability of containing no semiconducting CNTs, resulting in CNFET failures. Upsizing these vulnerable small-width CNFETs is an expensive design choice since it can result in substantial area/power penalties. This paper introduces a processing/design co-optimization approach to reduce probability of CNFET failures at the chip-level. Large degree of spatial correlation observed in directional CNT growth presents a unique opportunity for such optimization. Maximum benefits from such correlation can be realized by enforcing the active regions of CNFETs to be aligned with each other. This approach relaxes the device-level failure probability requirement by 350X at the 45nm technology node, leading to significantly reduced costs associated with upsizing the small-width CNFETs. Jie Zhang 0007, Shashikanth Bobba, Nishant Patil, Albert Lin 0002, H.-S. Philip Wong, Giovanni De Micheli, Subhasish Mitra |
DAC | 5 |
| 2010 | Carbon nanotube circuits: Living with imperfections and variationsabstractCarbon Nanotube Field-Effect Transistors (CNFETs) can potentially provide significant energy-delay-product benefits compared to silicon CMOS. However, CNFET circuits are subject to several sources of imperfections. These imperfections lead to incorrect logic functionality and substantial circuit performance variations. Processing techniques alone are inadequate to overcome the challenges resulting from these imperfections. An imperfection-immune design methodology is required. We present an overview of imperfection-immune design techniques to overcome two major sources of CNFET imperfections: metallic Carbon Nanotubes (CNTs) and CNT density variations. Jie Zhang 0007, Nishant Patil, Albert Lin 0002, H.-S. Philip Wong, Subhasish Mitra |
DATE | 4 |
| 2010 | Efficient FPGAs using nanoelectromechanical relaysabstractNanoelectromechanical (NEM) relays are promising candidates for programmable routing in Field-Programmable-Gate Arrays (FPGAs). This is due to their zero leakage and potentially low on-resistance. Moreover, NEM relays can be fabricated using a low-temperature process and, hence, may be monolithically integrated on top of CMOS circuits. Hysteresis characteristics of NEM relays can be utilized for designing programmable routing switches in FPGAs without requiring corresponding routing SRAM cells. Our simulation results demonstrate that the use of NEM relays for programmable routing in FPGAs can simultaneously provide 43.6% footprint area reduction, 37% leakage power reduction, and up to 28% critical path delay reduction compared to traditional SRAM-based CMOS FPGAs at the 22nm technology node. Chen Chen 0018, Roozbeh Parsa, Nishant Patil, Soogine Chong, Kerem Akarvardar, J. Provine, Jeff Watt, Roger T. Howe, H.-S. Philip Wong, Subhasish Mitra |
FPGA | 10 |
| 2010 | Phase Change MemoryabstractIn this paper, recent progress of phase change memory (PCM) is reviewed. The electrical and thermal properties of phase change materials are surveyed with a focus on the scalability of the materials and their impact on device design. Innovations in the device structure, memory cell selector, and strategies for achieving multibit operation and 3-D, multilayer high-density memory arrays are described. The scaling properties of PCM are illustrated with recent experimental results using special device test structures and novel material synthesis. Factors affecting the reliability of PCM are discussed. H.-S. Philip Wong, Simone Raoux, SangBum Kim, Jiale Liang, John P. Reifenberg, Bipin Rajendran, Mehdi Asheghi, Kenneth E. Goodson |
Proc. IEEE | 1 |
| 2009 | Digital VLSI logic technology using Carbon Nanotube FETs: frequently asked questionsabstractCarbon Nanotube Field-Effect Transistors (CNFETs) show promise as extensions to silicon-CMOS. Ideal CNFET circuits can potentially provide 20X Energy-Delay-Product benefits over silicon-CMOS at the 16 nm technology node. However, several challenges must be overcome before such performance benefits can be experimentally realized. In this paper, we present a brief overview of CNFET technology, and address commonly raised concerns through a series of Frequently Asked Questions (FAQs). We also provide a CNFET technology outlook which includes a survey of challenges as well as existing and potential solutions to these challenges. Nishant Patil, Albert Lin 0002, Jie Zhang 0007, H.-S. Philip Wong, Subhasish Mitra |
DAC | 4 |
| 2009 | Nanoelectromechanical (NEM) relays integrated with CMOS SRAM for improved stability and low leakageabstractWe present a hybrid nanoelectromechanical (NEM)/CMOS static random access memory (SRAM) cell, in which the two pull-down transistors of a conventional CMOS six transistor (6T) SRAM cell are replaced with NEM relays. This SRAM cell utilizes the infinite subthreshold slope and hysteretic properties of NEM relays to dramatically increase the cell stability compared to the conventional CMOS 6T SRAM cells. It also utilizes the zero off-state leakage of NEM relays to significantly decrease static power dissipation. The structure is designed so that the relatively long mechanical delay of the NEM relays does not result in performance degradation. Circuit simulations are performed using a VerilogA model of a NEM relay. Compared to a 65nm CMOS 6T SRAM cell, when 10nm-gap NEM relays (pull-in voltage = 0.8V, pull-out voltage = 0.2V, on resistance = 1kΩ) are integrated, hold and read static noise margin (SNM) improve by ~110% and ~250%, respectively. In addition, static power dissipation decreases by ~85%. The write delay decreases by ~60%, while read delay decreases by ~10%. The advantages in SNM and static power dissipation are expected to increase with scaling. Soogine Chong, Kerem Akarvardar, Roozbeh Parsa, Jun-Bo Yoon, Roger T. Howe, Subhasish Mitra, H.-S. Philip Wong |
ICCAD | 7 |
| 2009 | Fabrication and Characterization of Emerging Nanoscale MemoryabstractConventional solid state memory technologies such as flash memory, DRAM, and SRAM are facing scaling challenges due to fundamental limitations. Therefore, various new memory technologies are being widely researched and evaluated to continue the cost/performance improvement trend of solid state memory devices. To assess the potential scalability of emerging nanoscale memory beyond conventional limits, it is essential to characterize and understand how differently they perform at the nanoscale compared to known properties in the microscale. New nanoscale fabrication methods and new memory technologies offer a great opportunity for future memory device research. In this regard, we evaluated characteristics of nanoscale phase change memory and Ni oxide memory using nanofabrication technologies such as nanowire growth, nanocrystal synthesis, diblock copolymer patterning, and ebeam lithography. Evaluated characteristics include not only their device performance but also key material properties that might affect the ultimate device performance. The nanofabrication method for each memory material is also discussed due to its potential to overcome the difficulties of conventional semiconductor fabrication process. SangBum Kim, Byoungil Lee, Marissa Caldwell, H.-S. Philip Wong |
ISCAS | 5 |
| 2008 | Carbon nanotube transistor compact model for circuit design and performance optimizationabstractIn this paper, we describe the development of the Stanford University Carbon Nanotube FET (CNFET) Compact Model. The CNFET Model is a circuit-compatible, compact model which describes enhancement-mode, CMOS-like CNFETs. It can be used to simulate both functionality and performance of large-scale circuits with hundreds of CNFETs. To produce realistic and relevant results, the model accounts for several practical non-idealities such as scattering in the near-ballistic channel, effects of the source/drain extension region, and charge-screening for multiple-nanotube CNFETs. The model also includes a full transcapacitance network for more accurate transient and AC results. The Stanford University CNFET Model is implemented in both HSPICE macro language and VerilogA. The VerilogA implementation shows speedups of roughly 7x∼15x over HSPICE. Applications of the model suggest that n- and p-CNFETs will have 6x and 13x speed advantage over Si n- and p-MOSFETs respectively at the 32nm node, and that a CNT density of 250 CNTs/um is ideal for multiple-nanotube gates. Such a compact CNFET model will be absolutely essential in ushering in the Design Era of CNFET circuits as carbon nanotube technology outgrows its “science discovery” phase. Albert Lin 0002, Gordon C. Wan, H.-S. Philip Wong |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2008 | Design Methods for Misaligned and Mispositioned Carbon-Nanotube Immune CircuitsabstractCarbon-nanotube (CNT) field-effect transistors (CNFETs) are promising extensions to silicon CMOS. Simulations show that CNFET inverters fabricated with a perfect CNFET technology have 13 times better energy delay product compared with 32-nm silicon CMOS inverters. The following two fundamental challenges prevent the fabrication of CNFET circuits with the aforementioned advantages: 1) misaligned and mispositioned CNTs and 2) metallic CNTs. Misaligned and mispositioned CNTs can cause incorrect functionality. This paper presents a technique for designing arbitrary logic functions using CNFET circuits that are guaranteed to implement correct functions even in the presence of a large number of misaligned and mispositioned CNTs. Experimental demonstration of misaligned and mispositioned CNT-immune logic structures is also presented. Nishant Patil, Albert Lin 0002, H.-S. Philip Wong, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Automated Design of Misaligned-Carbon-Nanotube-Immune CircuitsabstractCarbon Nanotube Field-Effect Transistors (CNFETs) are promising candidates as extensions to Silicon CMOS due to excellent CV/I device performance. An ideal CNFET inverter fabricated using a perfect CNFET technology can have 5.1 times faster F04 delay and 2.6 times lower energy per cycle compared to a 32nm Silicon CMOS inverter. Two fundamental challenges prevent us from creating CNFET-based logic designs with the advantages quoted above: 1. Misaligned Carbon Nanotubes (CNTs), and 2. Metallic CNTs. Misaligned CNTs can result in incorrect logic function implementations. This paper presents a technique for designing CNFET-based arbitrary logic functions that are guaranteed to be correct even in the presence of a large number of misaligned CNTs. Nishant Patil, H.-S. Philip Wong, Subhasish Mitra |
DAC | 3 |
| 2006 | Carbon nanotube transistor circuits: models and tools for design and performance optimizationabstractIn this paper, we describe the development of device models and tools for the design of new transistors such as the carbon nanotube transistor. An HSPICE model for enhancement mode nanotube transistor has been developed. It can be used for design of nanotube transistor circuits as well as to study performance benefits of the new transistor. A model of the carbon nanotube transistor with Schottky barrier is presented. The model enables device design and performance optimization. H.-S. Philip Wong, Arash Hazeghi, Tejas Krishnamohan, Gordon C. Wan |
ICCAD | 1 |
| 2001 | Device scaling limits of Si MOSFETs and their application dependenciesabstractThis paper presents the current state of understanding of the factors that limit the continued scaling of Si complementary metal-oxide-semiconductor (CMOS) technology and provides an analysis of the ways in which application-related considerations enter into the determination of these limits. The physical origins of these limits are primarily in the tunneling currents, which leak through the various barriers in a MOS field-effect transistor (MOSFET) when it becomes very small, and in the thermally generated subthreshold currents. The dependence of these leakages on MOSFET geometry and structure is discussed along with design criteria for minimizing short-channel effects and other issues related to scaling. Scaling limits due to these leakage currents arise from application constraints related to power consumption and circuit functionality. We describe how these constraints work out for some of the most important application classes: dynamic random access memory (DRAM), static random access memory (SRAM), low-power portable devices, and moderate and high-performance CMOS logic. As a summary, we provide a table of our estimates of the scaling limits for various applications and device types. The end result is that there is no single end point for scaling, but that instead there are many end points, each optimally adapted to its particular applications. David J. Frank, Robert H. Dennard, Edward J. Nowak, Paul M. Solomon, Yuan Taur, H.-S. Philip Wong |
Proc. IEEE | 6 |
| 1999 | Nanoscale CMOSabstractThis paper examines the apparent limits, possible extensions, and applications of CMOS technology in the nanometer regime. Starting from device scaling theory and current industry projections, we analyze the achievable performance and possible limits of CMOS technology from the point of view of device physics, device technology, and power consumption. Various possible extensions to the basic logic and memory devices are reviewed, with emphasis on novel devices that are structurally distinct front conventional bulk CMOS logic and memory devices. Possible applications of nanoscale CMOS are examined, with a view to better defining the likely capabilities of future microelectronic systems. This analysis covers both data processing applications and nondata processing applications such as RF and imaging. Finally, we speculate on the future of CMOS for the coming 15-20 years. H.-S. Philip Wong, David J. Frank, Paul M. Solomon, Clement H. J. Wann, Jefferey J. Welser |
Proc. IEEE | 1 |
| 1997 | CMOS scaling into the nanometer regimeabstractStarting with a brief review on 0.1-/spl mu/m (100 nm) CMOS status, this paper addresses the key challenges in further scaling of CMOS technology into the nanometer (sub-100 nm) regime in light of fundamental physical effects and practical considerations. Among the issues discussed are: lithography, power supply and threshold voltage, short-channel effect, gate oxide, high-field effects, dopant number fluctuations and interconnect delays. The last part of the paper discusses several alternative or unconventional device structures, including silicon-on-insulator (SOI), SiGe MOSFET's, low-temperature CMOS, and double-gate MOSFET's, which may lead to the outermost limits of silicon scaling. Yuan Taur, Douglas A. Buchanan, David J. Frank, Khalid E. Ismail, Shih-Hsien Lo, George A. Sai-Halasz, Raman G. Viswanathan, Hsing-Jen C. Wann, Shalom J. Wind, H.-S. Philip Wong |
Proc. IEEE | 11 |