Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Leland Chang

dblp:45/4122 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Hardware accelerators and domain-specific architectures · 67% Emerging computing paradigms · 18% Energy-efficient computing · 10%
Artificial intelligence
4 papers
Efficient and distributed learning · 96% Deep learning architectures and training · 4%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.642021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks · DAC 2019
Machine learning › Efficient and distributed learning
model compression
0.832022
Deep Compression of Pre-trained Transformer Models · NeurIPS 2022
BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks · DAC 2019
Compensated-DNN: energy efficient low-precision deep neural networks by compensating quantization errors · DAC 2018
Machine learning › Efficient and distributed learning › model compression
quantization
0.832022
Deep Compression of Pre-trained Transformer Models · NeurIPS 2022
BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks · DAC 2019
Compensated-DNN: energy efficient low-precision deep neural networks by compensating quantization errors · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
low-precision DNN inference
0.722019
BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks · DAC 2019
Compensated-DNN: energy efficient low-precision deep neural networks by compensating quantization errors · DAC 2018
Machine learning › Efficient and distributed learning › model compression › quantization › low-bit quantization
4-bit quantization
0.612022
Deep Compression of Pre-trained Transformer Models · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression › sparsity
structured sparsity
0.612022
Deep Compression of Pre-trained Transformer Models · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
transformer compression
0.612022
Deep Compression of Pre-trained Transformer Models · NeurIPS 2022
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.512021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Emerging computing paradigms
approximate computing
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Hardware accelerators and domain-specific architectures
approximate computing accelerator
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Emerging computing paradigms › approximate computing
cross-layer approximate computing
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Energy-efficient computing › voltage scaling
near-threshold computing
0.112012
Near-threshold operation for power-efficient computing?: it depends · DAC 2012
Machine learning › Deep learning architectures and training
neural network inference
0.112020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Energy-efficient computing
power-performance tradeoff
0.112010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
Energy-efficient computing
voltage scaling
0.112010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
GPUs and heterogeneous computing
heterogeneous architecture
0.012012
Near-threshold operation for power-efficient computing?: it depends · DAC 2012
Integrated circuit design › semiconductor devices › multi-gate devices
FinFET
0.012003
Extremely scaled silicon nano-CMOS devices · Proc. IEEE 2003
Integrated circuit design › semiconductor devices
multi-gate devices
0.012003
Extremely scaled silicon nano-CMOS devices · Proc. IEEE 2003
Memory systems
cache
0.012010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
Integrated circuit design › semiconductor devices
silicon-on-insulator
0.012003
Extremely scaled silicon nano-CMOS devices · Proc. IEEE 2003

Methods — techniques the papers use, named apart from their topics

quantization · 0.9pruning · 0.9mixed-precision arithmetic · 0.9custom number representation · 0.9two-scale-factor fixed-point quantization · 0.8quantization error compensation · 0.7fixed-point quantization · 0.7sparsity-aware fine-tuning · 0.6scheduled dropout · 0.6quantization-aware fine-tuning · 0.6performance modeling · 0.5variability mitigation · 0.1power delivery · 0.1
YearPublicationVenuePosition
2022 Deep Compression of Pre-trained Transformer Models
abstract
Pre-trained transformer models have achieved remarkable success in natural language processing (NLP) and have recently become competitive alternatives to Convolution Neural Networks (CNN) and Recurrent Neural Networks (RNN) in vision and speech tasks, respectively. Due to excellent computational efficiency and scalability, transformer models can be trained on exceedingly large amounts of data; however, model sizes can grow tremendously. As high performance, large-scale, and pre-trained transformer models become available for users to download and fine-tune for customized downstream tasks, the deployment of these models becomes challenging due to the vast amount of operations and large memory footprint. To address this challenge, we introduce methods to deeply compress pre-trained transformer models across three major application domains: NLP, speech, and vision. Specifically, we quantize transformer backbones down to 4-bit and further achieve 50% fine-grained structural sparsity on pre-trained BERT, Wav2vec2.0 and Vision Transformer (ViT) models to achieve 16x compression while maintaining model accuracy. This is achieved by identifying the critical initialization for quantization/sparsity aware fine-tuning, as well as novel techniques including quantizers with zero-preserving format and scheduled dropout. These hardware-friendly techniques need only to be applied in the fine-tuning phase for downstream tasks; hence, are especially suitable for acceleration and deployment of pre-trained transformer models.
Naigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen, Chia-Yu Chen, Kaoutar El Maghraoui, Vijayalakshmi Srinivasan, Leland Chang
NeurIPS8
2021 RaPiD: AI Accelerator for Ultra-low Precision Training and Inference
abstract
The growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications.
Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan
ISCA53
2020 Efficient AI System Design With Cross-Layer Approximate Computing
abstract
Advances in deep neural networks (DNNs) and the availability of massive real-world data have enabled superhuman levels of accuracy on many AI tasks and ushered the explosive growth of AI workloads across the spectrum of computing devices. However, their superior accuracy comes at a high computational cost, which necessitates approaches beyond traditional computing paradigms to improve their operational efficiency. Leveraging the application-level insight of error resilience, we demonstrate how approximate computing (AxC) can significantly boost the efficiency of AI platforms and play a pivotal role in the broader adoption of AI-based applications and services. To this end, we present RaPiD, a multi-tera operations per second (TOPS) AI hardware accelerator core (fabricated at 14-nm technology) that we built from the ground-up using AxC techniques across the stack including algorithms, architecture, programmability, and hardware. We highlight the workload-guided systematic explorations of AxC techniques for AI, including custom number representations, quantization/pruning methodologies, mixed-precision architecture design, instruction sets, and compiler technologies with quality programmability, employed in the RaPiD accelerator.
Swagath Venkataramani, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Mingu Kang, Ankur Agarwal, Jinwook Oh, Shubham Jain 0004, Tina Babinsky, Nianzheng Cao, Thomas W. Fox, Bruce M. Fleischer, George Gristede, Michael Guillorn, Howard Haynie, Hiroshi Inoue, Kazuaki Ishizaki, Michael J. Klaiber, Shih-Hsien Lo, Gary W. Maier, Silvia M. Müller, Michael Scheuermann, Eri Ogawa, Marcel Schaal, Mauricio J. Serrano, Joel Silberman, Christos Vezyrtzis, Wei Wang 0333, Fanchieh Yee, Matthew M. Ziegler, Ching Zhou, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Vijayalakshmi Srinivasan, Leland Chang, Kailash Gopalakrishnan
Proc. IEEE39
2019 BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks
abstract
Fixed-point implementations (FxP) are prominently used to realize Deep Neural Networks (DNNs) efficiently on energy-constrained platforms. The choice of bit-width is often constrained by the ability of FxP to represent the entire range of numbers in the datastructure with sufficient resolution. At low bit-widths (< 8 bits), state-of-the-art DNNs invariably suffer a loss in classification accuracy due to quantization/saturation errors.
Shubham Jain 0004, Swagath Venkataramani, Vijayalakshmi Srinivasan, Jungwook Choi, Kailash Gopalakrishnan, Leland Chang
DAC6
2019 Memory and Interconnect Optimizations for Peta-Scale Deep Learning Systems
abstract
Hardware accelerators are a promising solution to the stringent computational requirements of Deep Neural Networks (DNNs). Ranging from low-power IP cores to server class systems, various accelerator architectures with high TOPS/W peak processing efficiencies and flexibility to execute different DNN topologies have been proposed. Prior efforts improve core utilization through better data-flows and computation sequencing, but little effort has thus far been devoted to systematically programming DNN accelerators to extract best possible system utilization, particularly for DNN training, which can be parallelized across peta-scale systems. In this work, we address the hitherto open challenge of systematically mapping computations onto Peta-scale accelerator systems, comprising many (thousands of) processing cores spanning many chips, while maximizing overall system performance. We achieve this by characterizing the design space of possible mapping configurations, building a detailed performance model that incorporates every computation and data-transfer involved in DNN training, and using a design space exploration tool called DEEPSPATIALMATRIX to identify the performance optimal configuration. We highlight 4 key optimizations built within DEEPSPATIALMATRIX - hybrid data-model parallelism, inter-layer memory reuse, time-step pipelining, and dynamic spatial minibatching - each of which improve system utilization by carefully managing the available memory capacity and interconnect bandwidth to balance the compute vs. communication costs. On a 8-peta-FLOP accelerator system, we demonstrate 1.36×-32× improvement in training performance through our design space exploration and optimizations across image recognition (VGG16, ResNet50) and machine translation (GNMT) DNN models.
Swagath Venkataramani, Vijayalakshmi Srinivasan, Jungwook Choi, Philip Heidelberger, Leland Chang, Kailash Gopalakrishnan
HiPC5
2018 Compensated-DNN: energy efficient low-precision deep neural networks by compensating quantization errors
abstract
Deep Neural Networks (DNNs) represent the state-of-the-art in many Artificial Intelligence (AI) tasks involving images, videos, text, and natural language. Their ubiquitous adoption is limited by the high computation and storage requirements of DNNs, especially for energy-constrained inference tasks at the edge using wearable and IoT devices. One promising approach to alleviate the computational challenges is implementing DNNs using low-precision fixed point (<16 bits) representation. However, the quantization error inherent in any Fixed Point (FxP) implementation limits the choice of bit-widths to maintain application-level accuracy. Prior efforts recommend increasing the network size and/or re-training the DNN to minimize loss due to quantization, albeit with limited success.
Shubham Jain 0004, Swagath Venkataramani, Vijayalakshmi Srinivasan, Jungwook Choi, Pierce Chuang, Leland Chang
DAC6
2018 Across the Stack Opportunities for Deep Learning Acceleration
abstract
The combination of growth in compute capabilities and availability of large datasets has led to a re-birth of deep learning. Deep Neural Networks (DNNs) have become state-of-the-art in a variety of machine learning tasks spanning domains across vision, speech, and machine translation. Deep Learning (DL) achieves high accuracy in these tasks at the expense of 100s of ExaOps of computation; posing significant challenges to efficient large-scale deployment in both resource-constrained environments and data centers.
Vijayalakshmi Srinivasan, Bruce M. Fleischer, Sunil Shukla, Matthew M. Ziegler, Joel Silberman, Jinwook Oh, Jungwook Choi, Silvia M. Müller, Ankur Agrawal, Tina Babinsky, Nianzheng Cao, Chia-Yu Chen, Pierce Chuang, Thomas W. Fox, George Gristede, Michael Guillorn, Howard Haynie, Michael J. Klaiber, Dongsoo Lee, Shih-Hsien Lo, Gary W. Maier, Michael Scheuermann, Swagath Venkataramani, Christos Vezyrtzis, Naigang Wang, Fanchieh Yee, Ching Zhou, Pong-Fei Lu, Brian W. Curran, Leland Chang, Kailash Gopalakrishnan
ISLPED30
2018 Taming the beast: Programming Peta-FLOP class Deep Learning Systems
abstract
No abstract available.
Swagath Venkataramani, Vijayalakshmi Srinivasan, Jungwook Choi, Kailash Gopalakrishnan, Leland Chang
ISLPED5
2017 POSTER: Design Space Exploration for Performance Optimization of Deep Neural Networks on Shared Memory Accelerators
abstract
The growing prominence and computational challenges imposed by Deep Neural Networks (DNNs) has fueled the design of specialized accelerator architectures and associated dataflows to improve their implementation efficiency. Each of these solutions serve as a datapoint on the throughput vs. energy trade-offs for a given DNN and a set of architectural constraints. In this paper, we set out to explore whether it is possible to systematically explore the design space so as to estimate a given DNN's (both inference and training) performance on an shared memory architecture specification using a variety of data-flows. To this end, we have developed a framework, DEEPMATRIX, which given a description of a DNN and a hardware architecture, automatically identifies how the computations of the DNN's layers need to partitioned and mapped on to the architecture such that the overall performance is maximized, while meeting the constraints imposed by the hardware (processing power, memory capacity, bandwidth etc.) We demonstrate DEEPMATRIX's effectiveness for the VGG DNN benchmark, showing the trade-offs and sensitivity of utilization based on different architecture constraints.
Swagath Venkataramani, Jungwook Choi, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan, Leland Chang
PACT5
2017 Cognitive Data-Centric Systems
abstract
With rapid growth in the availability of massive amounts of data and the development of new machine learning and deep learning techniques, significant opportunities exist in the application of computing to learn from data, build models, and discover insights -- cognitive tasks that can augment human expertise in a broad range of industries. Computing systems must evolve to efficiently meet these needs by leveraging innovation in heterogeneous systems infrastructure and information technology consumption models that are increasingly driven by cloud-based delivery. These new systems must be designed to accommodate the entirety of the overall workflow, including not just machine learning and analytics tasks, but also data management and manipulation. In a convergence with systems for classical modeling and simulation (HPC and technical computing), cognitive workloads can benefit dramatically from hardware acceleration. As decades of sustained CMOS technology scaling begins to slow, the specificity and optimality of hardware accelerators will be a key enabler for system-level performance while simultaneously presenting challenges in composing systems that seamlessly integrate traditional CPUs, multiple accelerators, and different memories. This talk will discuss cognitive data-centric systems for the next era of computing, in which balanced heterogeneous systems are delivered through the cloud.
Leland Chang
ACM Great Lakes Symposium on VLSI1
2016 Synthesis design strategies for energy-efficient microprocessors
abstract
A detailed synthesis study has been performed on a functional unit from a recent IBM microprocessor to explore the voltage-frequency space for energy-efficient design points across a wide performance spectrum ranging from 625 MHz at 0.48V to 5.6 GHz at 0.95V. It is found that the optimal operating voltage depends strongly on frequency for an energy-efficient design. Circuit characteristics, as represented by the combination of the average gate width, effective VT, and buffering scheme, differ significantly between designs optimized for low voltage-frequency and for high voltage-frequency operations and suggest a distinct application dependence in the selection of standard cell images and optimal design points. In particular, for optimal energy efficiency at a given frequency, low voltage designs should utilize smaller gate width and lower VT. Though a design energy-optimized near 1V is more scalable over a wide frequency range when operating at low voltages, designs optimized at a lower voltage-frequency point can be leveraged to offer better solutions in both performance and energy efficiency within a narrow frequency range near the design point.
Ching Zhou, Yu-Shiang Lin, Pong-Fei Lu, Bruce M. Fleischer, David J. Frank, Leland Chang
ICCD6
2013 Low-Power Circuit Analysis and Design Based on Heterojunction Tunneling Transistors (HETTs)
abstract
The theoretical lower limit of subthreshold swing in mosfets (60 mV/decade) significantly restricts low-voltage operation since it results in a low ON -to- OFF current ratio at low supply voltages. This paper investigates extremely low-power circuits based on new Si/SiGe heterojunction tunneling transistors (HETTs) that have a subthreshold swing of . Device characteristics, as determined through technology computer aided design tools, are used to develop a Verilog-A device model to simulate and evaluate a range of HETT-based circuits. We show that an HETT-based ring oscillator (RO) shows a 9-19 times reduction in dynamic power compared to a CMOS RO. We also explore two key differences between HETTs and traditional mosfets, namely, asymmetric current flow and increased Miller capacitance, analyze their effect on circuit behavior, and propose methods to address them. HETT characteristics have the most dramatic impact on static random access memory (SRAM) operation and we propose a novel seven-transistor HETT-based SRAM cell topology to overcome, and take advantage of, the asymmetric current flow. This new HETT SRAM design achieves 7-37 times reduction in leakage power compared to CMOS.
Yoonmyung Lee, Jin Cai, Isaac Lauer, Leland Chang, Steven J. Koester, David T. Blaauw, Dennis Sylvester
IEEE Trans. Very Large Scale Integr. Syst.5
2012 Near-threshold operation for power-efficient computing?: it depends
abstract
While it has long been argued that near-threshold (~0.5V) operation of CMOS technologies can dramatically improve power efficiency, widespread application of such low voltage operation to VLSI systems has yet to materialize. This is due in part to practical system workload demands, in which single-thread performance needs can limit strategies to improve parallelizeable throughput performance, but also due to barriers in the ability of supporting hardware to counter variability and reliability concerns while maintaining power efficiency throughout the system. This paper describes the issues on which the realization of near-threshold computing depends to explain why this strategy is not yet pervasive today. However, recent advancements across the spectrum of system design--including heterogeneous architectures, transistor and memory technologies, power delivery, packaging, and I/O--suggest that as the market for throughput performance grows, hardware technologies may soon become available to practically harness the promise of near-threshold operation.
Leland Chang, Wilfried Haensch
DAC1
2010 Practical Strategies for Power-Efficient Computing Technologies
abstract
After decades of continuous scaling, further advancement of silicon microelectronics across the entire spectrum of computing applications is today limited by power dissipation. While the trade-off between power and performance is well-recognized, most recent studies focus on the extreme ends of this balance. By concentrating instead on an intermediate range, an ~ 8× improvement in power efficiency can be attained without system performance loss in parallelizable applications-those in which such efficiency is most critical. It is argued that power-efficient hardware is fundamentally limited by voltage scaling, which can be achieved only by blurring the boundaries between devices, circuits, and systems and cannot be realized by addressing any one area alone. By simultaneously considering all three perspectives, the major issues involved in improving power efficiency in light of performance and area constraints are identified. Solutions for the critical elements of a practical computing system are discussed, including the underlying logic device, associated cache memory, off-chip interconnect, and power delivery system. The IBM Blue Gene system is then presented as a case study to exemplify several proposed directions. Going forward, further power reduction may demand radical changes in device technologies and computer architecture; hence, a few such promising methods are briefly considered.
Leland Chang, David J. Frank, Robert K. Montoye, Steven J. Koester, Brian L. Ji, Paul Coteus, Robert H. Dennard, Wilfried Haensch
Proc. IEEE1
2009 Low power circuit design based on heterojunction tunneling transistors (HETTs)
abstract
The theoretical lower limit of subthreshold swing in MOSFETs (60 mV/decade) significantly restricts low voltage operation since it results in a low ON to OFF current ratio at low supply voltages. This paper investigates extremely-low power circuits based on new Si/SiGe HEterojunction Tunneling Transistors (HETTs) that have subthreshold swing < 60 mV/decade. Device characteristics as determined through Technology Computer Aided Design (TCAD) tools are used to develop a Verilog-A device model to simulate and evaluate a range of HETT-based circuits. We show that a HETT-based ring oscillator (RO) shows a 9−19X reduction in dynamic power compared to a CMOS RO. We also explore two key differences between HETTs and traditional MOSFETs, namely asymmetric current flow and increased Miller capacitance, analyzing their effect on circuit behavior and proposing methods to address them. Finally, HETT characteristics have the most dramatic impact on SRAM operation and hence we propose a novel 7-transistor HETT-based SRAM cell topology to overcome, and take advantage of, the asymmetric current flow. This new HETT SRAM design achieves 7−37X reduction in leakage power compared to CMOS.
Yoonmyung Lee, Jin Cai, Isaac Lauer, Leland Chang, Steven J. Koester, Dennis Sylvester, David T. Blaauw
ISLPED5
2003 Extremely scaled silicon nano-CMOS devices
abstract
Silicon-based CMOS technology can be scaled well into the nanometer regime. High-performance, planar, ultrathin-body devices fabricated on silicon-on-insulator substrates have been demonstrated down to 15-nm gate lengths. We have also introduced the FinFET, a double-gate device structure that is relatively simple to fabricate and can be scaled to gate lengths below 10 nm. In this paper, some of the key elements of these technologies are described, including sublithographic patterning, the effects of crystal orientation and roughness on carrier mobility, gate work function engineering, circuit performance, and sensitivity to process-induced variations.
Leland Chang, Yang-Kyu Choi, Daewon Ha, Pushkar Ranade, Shiying Xiong, Jeffrey Bokor, Chenming Hu, Tsu-Jae King Liu
Proc. IEEE1