Tianjian Li

dblp:38/3553 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
8since 2021 · last 2026
0009-0004-0985-8684ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Hardware accelerators and domain-specific architectures · 20% Electronic design automation · 16% Memory systems · 12%
Artificial intelligence
4 papers
Language models and text generation · 31% Reinforcement learning · 23% Trustworthy machine learning · 18%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL · ASPLOS (2) 2026
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models
1.012026
History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL · ASPLOS (2) 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.922021
ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor Cores · DAC 2020
Natural language and speech › Language models and text generation
alignment
0.912025
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning · ICML 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning · ICML 2025
Machine learning › Reinforcement learning
preference learning
0.912025
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning · ICML 2025
High-performance computing › large-scale training
large language model training
0.912025
ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention · NeurIPS 2025
Parallel and multicore computing › parallel algorithms › parallel algorithm design
sequence parallelism
0.912025
ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness
noisy data
0.812024
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
robust learning
0.812024
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models · ICLR 2024
Natural language and speech › Machine translation
robust machine translation
0.812024
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models · ICLR 2024
Electronic design automation › hardware verification and test
memory testing
0.522017
Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique · DAC 2017
A Novel Test Method for Metallic CNTs in CNFET-Based SRAMs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Emerging computing paradigms › quantum computing › quantum machine learning
noise-aware training
0.512021
ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Memory systems › in-memory computing
ReRAM crossbar accelerator
0.512021
ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
direct convolution
0.412020
GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor Cores · DAC 2020
GPUs and heterogeneous computing › GPU computing
tensor cores
0.412020
GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor Cores · DAC 2020
Electronic design automation
hardware verification and test
0.322017
A Novel Test Method for Metallic CNTs in CNFET-Based SRAMs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique · DAC 2017
Hardware accelerators and domain-specific architectures
approximate computing accelerator
0.312018
A FPGA Friendly Approximate Computing Framework with Hybrid Neural Networks: (Abstract Only) · FPGA 2018
Reconfigurable computing and FPGAs › FPGA-based heterogeneous computing
CPU-FPGA platform
0.312018
A FPGA Friendly Approximate Computing Framework with Hybrid Neural Networks: (Abstract Only) · FPGA 2018
Reconfigurable computing and FPGAs
FPGA accelerator
0.312018
A FPGA Friendly Approximate Computing Framework with Hybrid Neural Networks: (Abstract Only) · FPGA 2018
Processor architecture and microarchitecture
SIMD
0.312018
CNFET-Based High Throughput SIMD Architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Natural language and speech › Language models and text generation
large language model inference
0.312026
History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL · ASPLOS (2) 2026
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.312026
History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL · ASPLOS (2) 2026
Electronic design automation › hardware verification and test
fault diagnosis
0.312017
Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique · DAC 2017
Memory systems › non-volatile memory
resistive memory
0.312017
Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique · DAC 2017
Memory systems › emerging memory technologies
RRAM crossbar
0.312017
Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique · DAC 2017
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention
0.312025
ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention · NeurIPS 2025
Electronic design automation › hardware verification and test
fault modeling
0.212016
A Novel Test Method for Metallic CNTs in CNFET-Based SRAMs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Natural language and speech › Language models and text generation
text summarization
0.212024
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models · ICLR 2024
Emerging computing paradigms
approximate and stochastic computing
0.112021
ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021

Methods — techniques the papers use, named apart from their topics

sequence parallelism · 1.7all-scan collective communication · 1.7speculative decoding · 1.0scheduling · 1.0direct preference optimization · 0.9data mixing · 0.9negative log-likelihood · 0.8error norm truncation · 0.8software-hardware co-design · 0.5on-device retraining · 0.5dynamic adjustment · 0.5stripe-mined convolution · 0.4multi-precision dataflow · 0.4pipelined datapath · 0.3multi-class classifier · 0.3co-training · 0.3
YearPublicationVenuePosition
2026 History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL
abstract
With the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlike traditional pre-training approaches, RL encompasses multiple stages: rollout, reward, and training, which necessitates collaboration among various worker types. However, current RL systems continue to grapple with substantial GPU underutilization, due to two primary factors: (1) The rollout stage dominates the overall RL process due to test-time scaling; (2) Imbalances in rollout lengths (within the same batch) result in GPU bubbles. While prior solutions like asynchronous execution and truncation offer partial relief, they may compromise training accuracy for efficiency. Our key insight stems from a previously overlooked observation: rollout responses exhibit remarkable similarity across adjacent training epochs. Based on the insight, we introduce RhymeRL, an LLM RL system designed to accelerate RL training with two key innovations. First, to enhance rollout generation, we present HistoSpec, a speculative decoding inference engine that utilizes the similarity of historical rollout token sequences to obtain accurate drafts. Second, to tackle rollout bubbles, we introduce HistoPipe, a two-tier scheduling strategy that leverages the similarity of historical rollout distributions to balance workload among rollout workers. Experimental results demonstrate that RhymeRL achieves up to a 2.6x performance improvement over existing methods, without compromising accuracy or modifying the RL paradigm.
Jingkai He, Tianjian Li, Erhu Feng, Dong Du 0003, Qian Liu 0033, Yubin Xia, Haibo Chen 0001
ASPLOS (2)2
2025 SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
abstract
Aligning language models with human preferences relies on pairwise preference datasets. While some studies suggest that on-policy data consistently outperforms off-policy data for preference learning, others indicate that the advantages of on-policy data are task-dependent, highlighting the need for a systematic exploration of their interplay. In this work, we show that on-policy and off-policy data offer complementary strengths: on-policy data is particularly effective for reasoning tasks like math and coding, while off-policy data performs better on subjective tasks such as creative writing and making personal recommendations. Guided by these findings, we introduce SimpleMix, an approach to combine the complementary strengths of on-policy and off-policy preference learning by simply mixing these two data sources. Our empirical results across diverse tasks and benchmarks demonstrate that SimpleMix substantially improves language model alignment. Specifically, SimpleMix improves upon on-policy DPO and off-policy DPO by an average of 6.03 on Alpaca Eval 2.0. Moreover, it surpasses prior approaches that are much more complex in combining on- and off-policy data, such as HyPO and DPO-Mix-P, by an average of 3.05. These findings validate the effectiveness and efficiency of SimpleMix for enhancing preference-based alignment.
Tianjian Li, Daniel Khashabi
ICML1
2025 Upsample or Upweight? Balanced Training on Heavily Imbalanced Datasets
abstract
Tianjian Li, Haoran Xu, Weiting Tan, Kenton Murray, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tianjian Li, Weiting Tan, Kenton Murray, Daniel Khashabi
NAACL (Long Papers)1
2025 Benchmarking Language Model Creativity: A Case Study on Code Generation
abstract
Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang 0001, Daniel Khashabi
NAACL (Long Papers)3
2025 Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
abstract
Jingyu Zhang, Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi
NAACL (Long Papers)3
2025 ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention
abstract
Linear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these workloads across devices, become the primary performance bottleneck due to substantial communication overhead. In this paper, we introduce ZeCO (Zero Communication Overhead) sequence parallelism for linear attention models, a new SP method designed to overcome these limitations and achieve practically end-to-end near-linear scalability for long sequence training. For example, training a model with a 1M sequence length across 64 devices using ZeCO takes roughly the same time as training with an 16k sequence on a single device. At the heart of ZeCO lies All-Scan, a novel collective communication primitive. All-Scan provides each SP rank with precisely the initial operator state it requires while maintaining a minimal communication footprint, effectively eliminating communication overhead. Theoretically, we prove the optimaity of ZeCO, showing that it introduces only negligible time and space overhead. Empirically, we compare the communication costs of different sequence parallelism strategies and demonstrate that All-Scan achieves the fastest communication in SP scenarios. Specifically, on 256 GPUs with an 8M sequence length, ZeCO achieves a 60\% speedup compared to the current state-of-the-art (SOTA) SP method. We believe ZeCO establishes a clear path toward efficiently training next-generation LLMs on previously intractable sequence lengths.
Yuhong Chou, Rui-Jie Zhu 0003, Tianjian Li, Congying Chu, Qian Liu 0033, Jibin Wu, Zejun Ma 0001
NeurIPS5
2024 Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
abstract
Text generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50\% of noise is added to the data.
Tianjian Li, Philipp Koehn, Daniel Khashabi, Kenton Murray
ICLR1
2021 ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator
abstract
Deep neural networks (DNNs) have gained a strong momentum among various applications. The enormous matrix-multiplication exhibited in the above DNNs is computation and memory intensive. Resistive random-access memory crossbar (RRAM-crossbar) consisting of memristor cells can naturally carry out the matrix-vector multiplication. RRAM-crossbar-based accelerator, therefore, has two orders of magnitude of higher energy-efficiency than conventional accelerators. The imperfect fabrication process of RRAM-crossbars, however, causes various defects and process variations. These fabrication imperfections not only result in significant yield loss but also degrade the accuracy of DNNs executed on the RRAM-crossbars. In this article, we first propose an accelerator-friendly neural-network training method, by leveraging the inherent self-healing capability of the neural network, to prevent the large-weight synapses from being mapped to the imperfect memristors. Next, we propose a dynamic adjustment mechanism to extend the above method for DNNs, such as multilayer perceptrons (MLPs), wherein the imperfect-memristor induced errors can accumulate and magnify through multiple layers. Such off-device training method is a pure software solution, and it is unable to provide enough accuracy for convolutional neural networks (CNNs). Several works propose error-tolerable hardware design by allowing the retraining of CNNs on the RRAM-crossbar. Although this hardware-based on-device training method is effective, the frequent write operation on RRAM-crossbar hurt the endurance of RRAM-crossbars. Consequently, we propose a software and hardware co-design methodology to effectively preserve the classification accuracy of CNN with few on-device training iterations. The experimental results show that the proposed method can guarantee ≤1.1% loss of accuracy for resistance variations in MLP and CNN. Moreover, the proposed method can guarantee ≤1% loss of accuracy even when stuck-at-faults (SAFs) rate = 20%.
Zhuoran Song, Yanan Sun 0003, Lerong Chen, Tianjian Li, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor Cores
abstract
To tailor for DNN (Deep Neural Network) acceleration, GPU has migrated to new architectures such as NVIDIA Volta and Turing that incorporate dedicated Tensor Cores. Although good at GEMM (generic matrix-matrix multiplication), Tensor Cores still have inefficiency facing convolutions with certain layer structures. This paper proposes a GPNPU (General-Purpose Neural-network Processing Unit) architecture, which offers another option of direct convolution in GPU. It stitches the direct convolution dataflow into the Tensor Cores with little hardware support, and resorts to regulated data layout with stripe-mined convolution execution to achieve higher performance and power efficiency, while retaining the general programability as GPU. We further apply a unified core design to support varied operand types and precision for higher computing throughput. The evaluation shows that GPNPU can outperform Tensor Cores on typical DNNs by 1.4X for inference (FP16) and 1.2X for training with much reduced power. The INT8 performance even increases to 2.4X. Our study demonstrates that it is possible and appealing to refine the Tensor Cores for greater DNN acceleration, while conforming to GPU architecture for the programmability necessary in future DNN evolution.
Zhuoran Song, Tianjian Li, Li Jiang 0002, Jing Ke, Xiaoyao Liang, Naifeng Jing
DAC3
2019 HUBPA: high utilization bidirectional pipeline architecture for neuromorphic computing
abstract
Training Convolutional Neural Networks(CNNs) is both memory-and computation-intensive. The resistive random access memory (ReRAM) has shown its advantage to accelerate such tasks with high energy-efficiency. However, the ReRAM-based pipeline architecture suffers from the low utilization of computing resource, caused by the imbalanced data throughput in different pipeline stages because of the inherent down-sampling effect in CNNs and the inflexible usage of ReRAM cells. In this paper, we propose a novel ReRAM-based bidirectional pipeline architecture, named HUBPA, to accelerate the training with higher utilization of the computing resource. Two stages of the CNN training, forward and backward propagations, are scheduled in HUBPA dynamically to share the computing resource. We design an accessory control scheme for the context switch of these two tasks. We also propose an efficient algorithm to allocate computing resource for each neural network layer. Our experiment results show that, compared with state-of-the-art ReRAM pipeline architecture, HUBPA improves the performance by 1.7X and reduces the energy consumption by 1.5X, based on the current benchmarks.
Houxiang Ji, Li Jiang 0002, Tianjian Li, Naifeng Jing, Jing Ke, Xiaoyao Liang
ASP-DAC3
2018 In-growth test for monolithic 3D integrated SRAM
abstract
Monolithic three-dimensional integration (M3I) directly fabricates tiers of integrated circuits upon each other and provides millions of vertical interconnections with inter-layer vias (ILVs). It thus brings higher integration density and communication capability compared with three-dimensional stacked integration (3D-SI). However, the Known-Good-Die problem haunting 3D-SI-a faulty tier causes the failure of the entire stack-also occurs in M3I. Lack of efficient test methodologies such as the pre-bond testing in 3D-SI, M3I may have a more significant yield drop and thus its cost may be unacceptable for main-stream adoption. This paper introduces a novel In-growth test method for M3I SRAM. We propose a novel Design-for-Test (DfT) methodology to enable the proposed In-growth test on cell-level partitioned incomplete SRAM cells. We also build a statistical model of cost and discover a prospective judgement to determine whether or not to stop the fabrication, in order to prevent from raising the cost of fabricating more tiers upon the irreparable tiers. We find that a “sweet point” exists in the judgement, which can minimize the overall cost. Experimental results show the effectiveness of our proposed test methodology.
Pu Pang, Yixun Zhang, Tianjian Li, Sung Kyu Lim, Quan Chen 0002, Xiaoyao Liang, Li Jiang 0002
DATE3
2018 A FPGA Friendly Approximate Computing Framework with Hybrid Neural Networks: (Abstract Only)
abstract
Neural approximate computing is promising to gain energy-efficiency at the cost of tolerable quality loss. The architecture contains two neural networks: the approximate accelerator generates approximate results while the classifier determines whether input data can be safely approximated. However, they are not compatible to a heterogeneous computing platform, due to the large communication overhead between the approximate accelerator and accurate cores, and the large speed gap between them. This paper proposes a software-hardware co-design strategy. With deep exploration of data distributions in the feature space, we first propose a novel approximate computing architecture containing a multi-class classifier and multiple approximate accelerator; this architecture, derived by the existing iterative co-training methods, can shift more data from accurate computation (in CPU) to approximate accelerator (in FPGA); the increased invocation of the approximate accelerator thus can yield higher utilization of the FPGA-based accelerator, resulting in the enhanced the performance. Moreover, much less input data is redistributed, by the classifier (also in FPGA), back to CPU, which can minimize the CPU-FPGA communication. Second, we design a pipelined data-path with batched input/output for the proposed hybrid architecture to efficiently hide the communication latency. A mask technique is proposed to decouple the synchronization between CPU and FPGA, in order to minimize the frequency of communication.
Haiyue Song, Tianjian Li, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002
FPGA3
2018 CNFET-Based High Throughput SIMD Architecture
abstract
Carbon nanotube field effect transistor (CNFET), using the carbon nanotubes (CNTs) as the material for conducting, is a promising alternative of CMOS technology to overcome the “power wall” issue. Recently, a microprocessor solely based on CNFETs was fabricated and demonstrated, which is a big step forward to the industrial practice. However, CNFETs are inherently subject to much larger process variation or manufacturing defects; thereby it may cause significant design cost to build high performance processors. This is exacerbated in the large register file (RF) architectures widely used in single instruction multiple data (SIMD) architectures, e.g., general public utilities style processors, where the number of critical paths are multiplied by the SIMD width and thread count. In this paper, we seek cost-effective approaches to address the issues by judiciously exploiting the strong asymmetric spatial correlation in the variation unique to the CNFET fabrication process. This paper presents a microarchitectural model to characterize CNFET delay variation and malfunction, under which we show that the RF organizations coupled with the architectural schemes are critical to the performance and power consumption of the SIMD processor. Therefore, we propose several architectural techniques to mitigate the performance degradation and the impact of CNT metallization, leveraging the distinctive CNFET characteristics and the unique features in the SIMD processors. Experimental results verify the effectiveness of the proposed techniques and demonstrate the great opportunity offered by this new device technology.
Li Jiang 0002, Tianjian Li, Naifeng Jing, Nam Sung Kim, Minyi Guo, Xiaoyao Liang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Sneak-Path Based Test and Diagnosis for 1R RRAM Crossbar Using Voltage Bias Technique
abstract
Metal-oxide resistive random access memories with a single memristor device at the crosspoint (1R RRAM) is a promising alternative to next generation storage technology due to their high density, scalability, non-volatility and low power consumption. However, the imperfect fabrication process introduces high defect rates of the nanoscale memristor devices and leads to yield degradation. In addition, sneak-paths occur in 1R RRAM crossbar that can jeaperdize the normal read/write operation. Previous work proposes voltage bias technique to eliminate the sneak-paths. Instead, in the paper, we leverage voltage bias to manipulate various distribution of sneak-paths that can screen one or multiple faults out of a 4 x 4 region of memristors at once, and consequently diagnose the exact location of each faulty memristor within three write-read operations. The SPICE simulation results highlight the effectiveness and efficiency of the proposed test method.
Tianjian Li, Xiangyu Bi, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002
DAC1
2017 Fault clustering technique for 3D memory BISR
abstract
Three Dimensional (3D) memory has gained a great momentum because of its large storage capacity, bandwidth and etc. A critical challenge for 3D memory is the significant yield loss due to the disruptive integration process: any memory die that cannot be successfully repaired leads to the failure of the whole stack. The repair ratio of each die must be as high as possible to guarantee the overall yield. Existing memory repair methods, however, follow the traditional way of using redundancies: a redundant row/column replaces a row/column containing few or even one faulty cell. We propose a novel technique specifically in 3D memory that can overcome this limitation. It can cluster faulty cells across layers to the same row/column in the same memory array so that each redundant row/column can repair more “faults”. Moreover, it can be applied to the existing repair algorithms. We design the BIST and BISR modules to implement the proposed repair technique. Experimental results show more than 71% enhancement of the repair ratio over the global 3D GESP solution and 80% redundancy-cost reduction, respectively.
Tianjian Li, Xiaoyao Liang, Hsien-Hsin S. Lee, Li Jiang 0002
DATE1
2016 CNFET-based high throughput register file architecture
abstract
A Carbon Nanotube field-effect transistor (CNFET) is a promising alternative to a traditional metal-oxide-semiconductor field-effect transistor (MOSFET) to overcome the “Power Wall” challenge. However, CNFETs are inherently subject to much larger process variation and thereby they can incur a significant design cost to build high-performance processors. Particularly, the large register files (RF) of SIMD GPU-style processors suffer more from such process variations because the number of critical paths are multiplied by the SIMD width and thread count. In this paper, we first show that RF organizations coupled with architectural techniques are critical to RF performance under CNFET-specific variations. Second, we propose several architectural techniques to mitigate the performance degradation, leveraging distinctive characteristics of CNFETs and unique features of SIMD processors. Our experiments demonstrate that the average RF performance is 53% higher than the worst design under variation and only 7% lower than the design with no variation.
Tianjian Li, Li Jiang 0002, Naifeng Jing, Nam Sung Kim, Xiaoyao Liang
ICCD1
2016 Defect tolerance for CNFET-based SRAMs
abstract
SRAMs based on carbon nanotube field-effect transistors (CNFETs) offer a promising alternative to conventional SRAMs due to their high energy efficiency and low leakage. However, the imperfect CNT fabrication process introduces high defect rates and a unique defect distribution; these problems may offset the power/performance benefits of CNFET-based SRAMs and lead to yield degradation. We propose a redundancy architecture with asymmetrically partitioned column blocks and the sharing of spares among column blocks. We also present a analytical model to characterize the distribution of faults, which can guide the design exploration of the proposed redundancy architecture. Simulation results highlight the accuracy of the proposed model, as well as the efficiency and effectiveness of the redundancy architecture.
Tianjian Li, Li Jiang 0002, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty
ITC1
2016 A Novel Test Method for Metallic CNTs in CNFET-Based SRAMs
abstract
Static random access memories (SRAMs) built on carbon nanotube field effect transistors (CNFETs) are promising alternatives to conventional CMOS-based SRAMs, due to their advantages in terms of power consumption and noise immunity. However, the nonideal carbon nanotube (CNT) fabrication process generates metallic-CNTs (m-CNTs) along with semiconductor-CNTs, leading to correlated faulty cells along the growth direction of the m-CNTs. In this paper, we propose a novel low-cost test solution to detect such faults. Instead of using conventional March test to test each and every SRAM cell, we selectively test certain SRAM cells and judiciously skip testing other SRAM cells between the selected cells. To ensure high fault coverage, we propose three jump test algorithms for different CNFET-SRAM layouts. Moreover, we model m-CNT-induced SRAM faults and characterize their distribution in the SRAM array. Experimental results show that the proposed solutions are able to achieve high fault coverage with low test cost.
Tianjian Li, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty, Naifeng Jing, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 Approximating optimal survivable scheduled service provisioning in WDM optical networks with Shared Risk Link Groups
abstract
Survivable service provisioning design has been an important issue in communication networks. In this work, we study survivable service provisioning using shared path based protection under a scheduled traffic model in wavelength convertible WDM optical mesh networks with Shared Risk Link Groups (SRLGs). In the scheduled traffic model, a set of demands is given, and the setup time and teardown time of a demand are known in advance. The objective is to minimize the total network resources (e.g., the number of wavelength-links) used by working paths and protection paths of the given set of demands while 100% restorability is guaranteed against any single SRLG failure. This problem is known to be NP-hard. We therefore study a time efficient approach to approximating the optimal solution to the problem. Our proposed approach is based on an iterative survivable routing scheme that utilizes a capacity provision matrix and processes demands sequentially. Our simulation results indicate that the proposed ISR-SRLG algorithm achieves excellent performance in terms of the total network resources used.
Tianjian Li, Bin Wang 0002
BROADNETS1
2006 Approximating Optimal Survivable Scheduled Service Provisioning in WDM Optical Networks with Iterative Survivable Routing
abstract
Survivable service provisioning design has emerged as one of the most important issues in communication networks in recent years. In this work, we study survivable service provisioning with shared protection under a scheduled traffic model in wavelength convertible WDM optical mesh networks. In this model, a set of demands is given, and the setup time and teardown time of a demand are known in advance. Based on different protection schemes used, this problem has been formulated as integer linear programs with different optimization objectives and constraints in our previous work [8, 9]. The problem is shown to be NP- hard. We therefore study time efficient approaches to approximating the optimal solution to the problem. Our proposed approach is based on an iterative survivable routing (ISR) scheme that utilizes a capacity provision matrix and processes demands sequentially using different demand scheduling policies. The objective is to minimize the total network resources (e.g., number of wavelength-links) used by working paths and protection paths of a given set of demands while 100% restorability is guaranteed against any single failure. The proposed algorithm is evaluated against solutions obtained by integer linear programming. Our simulation results indicate that the proposed ISR algorithm is extremely time efficient while achieving excellent performance in terms of total network resources used. The impact of demand scheduling policies on the ISR algorithm is also studied.
Tianjian Li, Bin Wang 0002
BROADNETS1
2005 Routing and wavelength assignment under a scheduled traffic model in reconfigurable WDM optical networks
abstract
In this paper, we propose a general scheduled traffic model, sliding scheduled traffic model. In this model, the setup time t/sub s/ of a demand whose holding time is T time units is not known in advance. Rather t/sub s/ is allowed to begin in a pre-specified time window [l,T] subject to the constraint that l/spl les/t/sub s//spl les/r-T. We then consider two problems: (1) how to properly place a demand within its associated time window to reduce overlapping in time among a set of demands; and (2) route and assign wavelengths (RWA) to a set of demands under the proposed sliding scheduled traffic model in mesh reconfigurable WDM optical networks without wavelength conversion. In addition, we consider how to rearrange a demand by negotiating a new setup time that minimizes the demand schedule change in case that the demand is blocked. To maximize temporal resource reuse, we propose a demand time conflict reduction algorithm to solve the first problem. Two algorithms, window based RWA algorithm and traffic matrix based RWA algorithm, are then proposed for the second problem. We compare the proposed RWA algorithms against a customized tabu search scheme. Simulation results show that the proposed demand time conflict reduction algorithm can resolve well over 50% of time conflicts and the space-time RWA algorithms are effective in satisfying demand requirements and minimizing total network resources used, d.
Bin Wang 0002, Tianjian Li, Chunsheng Xin
BROADNETS2
2005 On survivable service provisioning in WDM optical networks under a scheduled traffic model
abstract
We study survivable service provisioning under a scheduled traffic model in wavelength convertible WDM optical mesh networks. In this model, a set of demands is given, and the setup time and teardown time of each demand are known in advance. We formulate the problem as integer linear programs that maximally exploit network resource reuse in both space and time. The objective is to minimize the total number of wavelength-links used by working paths and protection paths of all traffic demands while 100% restorability is guaranteed against any single failures. Our simulation results indicate that joint optimization of resource sharing in space and time enabled by our connection holding-time aware protection schemes can achieve significantly better resource utilization than schemes that are holding-time unaware
Tianjian Li, Bin Wang 0002, Chunsheng Xin, Xinhui Zhang
GLOBECOM1
2004 Optimal configuration of p-cycles in WDM optical networks with sparse wavelength conversion
abstract
In this paper, we study the optimal configuration of p-cycles in survivable WDM optical mesh networks with sparse wavelength conversion while 100% restorability is guaranteed against any single failures. We formulate the problem as an integer linear program. In our optimization model, working paths are known before protection configuration is done. p-cycles and wavelength converters are then optimally determined subject to the constraint that only a given number of nodes have wavelength conversion capability. The objective is to minimize the cost of link capacity used by all p-cycles to accommodate a set of traffic demands. In the proposed p-cycle configuration architecture, we take into account converter sharing: (a) when converters are used for accessing p-cycles and for wavelength conversion between two adjacent on-cycle spans; (b) when converters are used among disjoint straddling spans incident to the same node for accessing p-cycles. Converter sharing enables the network to require as few converters as possible to attain a satisfactory level of performance. Our simulation results indicate that the proposed approach significantly outperforms the approach for WP networks in terms of protection cost and can obtain the optimal performance as achieved by the approach for VWP networks, but requires fewer wavelength conversion sites and fewer wavelength converters.
Tianjian Li, Bin Wang 0002
GLOBECOM1