Tianmu Li

dblp:199/8757 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-1078-6743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learned Approximate Computing: Algorithm Hardware Co-Optimization
abstract
Approximate hardware trades acceptable error for improved performance and previous literature focuses on optimizing this tradeoff in the hardware. We show in this article that the application and the hardware can be co-optimized to achieve the best-quality-performance tradeoff. We propose LAC: learned approximate computing to optimize the algorithm and approximate hardware at the same time to maximize quality of output. Our approach allows automatic selection of approximate computing hardware while achieving similar quality as dedicated training for a single hardware configuration. Our improved training algorithm allows simultaneous hardware selection and application optimization without additional runtime overhead. Multihardware setup chooses a separate approximate hardware for each part of an application which allows for more hardware configurations and further improves quality.
Egor Glukhov, Tianmu Li, Vaibhav Gupta, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 DRDebug: Automated Design Rule Debugging
abstract
Design rule checking (DRC) is an important step in the physical design flow that checks if a design meets the manufacturing constraints or design rules imposed by the process technology. It allows the foundry to ensure high acceptable manufacturing yield. Design rule verification is one of the most challenging steps because of the sheer size of the rule decks and the lack of standardization among these design rule manuals (DRMs). One way of efficiently discovering missed rule checks is by comparing the rule deck with that of another mature or well-established process. In this work, we develop two complementary techniques for comparing process design rule decks and automatically establishing a one-to-one correspondence between rules from two different process design kits (PDKs). The first approach, random layout generation (RLG), creates random layouts of different shapes and sizes. The generated layout is checked using both rule decks. The rules are then matched based on the violations generated. The second approach, based on rule language processing (RLP), matches rules based on the similarity between rule commands. Rules are directly matched based on the layer names and keywords present in the DRC commands. The two approaches are complementary and together they can correctly match more than 80% of the rules in two DRMs.
Irina Alam, Tianmu Li, Sean Brock, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 REX-SC: Range-Extended Stochastic Computing Accumulation for Neural Network Acceleration
abstract
Deep learning has grown in capability and size in recent years, prompting research on alternative computing methods to cope with the increased compute cost. Stochastic computing (SC) promises higher compute efficiency with its compact compute units, but accuracy issues have prevented wide adoption, and accuracy-improving techniques have sacrificed runtime or training performance. In this work, we propose extended range SC—Range-Extended SC Accumulation to deal with the accuracy issues of SC. By modifying the functionality of OR-based SC accumulation, we increase SC computation accuracy without sacrificing the performance benefits. Our approach achieves a$2\times $reduction in stream length for the same accuracy compared to SC with OR-based accumulation and an up to$3.6\times $improvement in energy compared to SC with binary addition. With proper modeling, our approach improves training performance for SC-based neural networks and makes training SC models practical for large datasets like ImageNet.
Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 LAC: Learned Approximate Computing
abstract
Approximate hardware trades acceptable error for improved performance and previous literature focuses on optimizing this trade-off in the hardware. We show in this paper that the application (i.e., the software) can be optimized for better accuracy without losing any performance benefits of the approximate hardware. We propose LAC: learned approximate computing as a method of tuning the application parameters to compensate for hardware errors. Our approach showed improvements across a variety of standard signal/image processing applications delivering an average improvement of 5.82db in PSNR and 0.23 in SSIM of the outputs. This translates to up to 87% power reduction and 83% area reduction for similar application quality. LAC allows the same approximate hardware to be used for multiple applications.
Vaibhav Gupta, Tianmu Li, Puneet Gupta 0001
DATE2
2022 Detecting deepfake videos based on spatiotemporal attention and convolutional LSTM
Beijing Chen, Tianmu Li, Weiping Ding 0001
Inf. Sci.2
2022 SASCHA - Sparsity-Aware Stochastic Computing Hardware Architecture for Neural Network Acceleration
abstract
Stochastic computing (SC) has recently emerged as a promising method for efficient machine learning acceleration. Its high compute density, affinity with dense linear algebra primitives, and approximation properties have an uncanny level of synergy with the deep neural network computational requirements. However, there is a conspicuous lack of works trying to integrate SC hardware with sparsity awareness, which has brought significant performance improvements to conventional architectures. In this work, we identify why common sparsity-exploiting techniques are not easily applicable to SC accelerators and propose a new architecture—SASCHA—sparsity-aware SC hardware architecture for the neural network acceleration that addresses those issues. SASCHA encompasses a set of techniques that make utilizing sparsity in inference practical for different types of SC computation. At 90% weight sparsity, SASCHA can be up to$6.5\times $faster and$5.5\times $more energy-efficient than comparable dense SC accelerators with a similar area without sacrificing the dense network throughput. SASCHA also outperforms sparse fixed-point accelerators by up to$4\times $in terms of latency. To the best of our knowledge, SASCHA is the first SC accelerator architecture oriented around sparsity.
Wojciech Romaszkan, Tianmu Li, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 GEO: Generation and Execution Optimized Stochastic Computing Accelerator for Neural Networks
abstract
Stochastic computing (SC) has seen a renaissance in recent years as a means for machine learning acceleration due to its compact arithmetic and approximation properties. Still, SC accuracy remains an issue, with prior works either not fully utilizing the computational density or suffering from significant accuracy losses. In this work, we propose GEO - Generation and Execution Optimized Stochastic Computing Accelerator for Neural Networks, which optimizes stream generation and execution components of SC, and bridges the accuracy gap between stochastic computing and fixed-point neural networks. It improves accuracy by coupling controlled stream sharing with training and balancing OR and binary accumulations. GEO further optimizes the SC execution through progressive shadow buffering and architectural optimizations. GEO can improve accuracy compared to state-of-the-art SC by 2.2-4.0% points while being up to 4.4X faster and 5.3X more energy efficient. GEO eliminates the accuracy gap between SC and fixed-point architectures while delivering up to 5.6X higher throughput and 2.6X lower energy.
Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001
DATE1
2020 ACOUSTIC: Accelerating Convolutional Neural Networks through Or-Unipolar Skipped Stochastic Computing
abstract
As privacy and latency requirements force a move towards edge Machine Learning inference, resource constrained devices are struggling to cope with large and computationally complex models. For Convolutional Neural Networks, those limitations can be overcome by taking advantage of enormous data reuse opportunities and amenability to reduced precision. To do that however, a level of compute density unattainable for conventional binary arithmetic is required. Stochastic Computing can deliver such density, but it has not lived up to its full potential because of multiple underlying precision issues. We present ACOUSTIC: Accelerating Convolutions through Or-Unipolar Skipped sTochastIc Computing, an accelerator framework that enables fully stochastic, high-density CNN inference. Leveraging split-unipolar representation, OR-based accumulation and novel computation-skipping approach, ACOUSTIC delivers server-class parallelism within a mobile area and power budget - a 12mm2accelerator can be as much as 38.7x more energy efficient and 72.5x faster than conventional fixed-point accelerators. It can also be up to 79.6x more energy efficient than state-of-the-art stochastic accelerators. At the lower-end ACOUSTIC achieves 8x-120X inference throughput improvement with similar energy and area when compared to recent mixed-signal/neuromorphic accelerators.
Wojciech Romaszkan, Tianmu Li, Tristan Melton, Sudhakar Pamarti, Puneet Gupta 0001
DATE2
2020 3PXNet: Pruned-Permuted-Packed XNOR Networks for Edge Machine Learning
abstract
As the adoption of Neural Networks continues to proliferate different classes of applications and systems, edge devices have been left behind. Their strict energy and storage limitations make them unable to cope with the sizes of common network models. While many compression methods such as precision reduction and sparsity have been proposed to alleviate this, they don’t go quite far enough. To push size reduction to its absolute limits, we combine binarization with sparsity in Pruned-Permuted-Packed XNOR Networks (3PXNet), which can be efficiently implemented on even the smallest of embedded microcontrollers. 3PXNets can reduce model sizes by up to 38X and reduce runtime by up to 3X compared with already compact conventional binarized implementations with less than 3% accuracy reduction. We have created the first software implementation of sparse-binarized Neural Networks, released as open source library targeting edge devices. Our library is complete with training methodology and model generating scripts, making it easy and fast to deploy.
Wojciech Romaszkan, Tianmu Li, Puneet Gupta 0001
ACM Trans. Embed. Comput. Syst.2
2017 Hybrid VC-MTJ/CMOS non-volatile stochastic logic for efficient computing
abstract
In this paper, we propose a non-volatile stochastic computing (SC) scheme using voltage-controlled magnetic tunnel junction (VC-MTJ) and negative differential resistance (NDR). The proposed design includes a VC-MTJ based true stochastic bit stream generator and VC-MTJ and NDR based stochastic adder, multiplier, register, which are experimentally demonstrated using 60nm VC-MTJ and CMOS NDR connected on die. These components are then used to realize FIR filter and AdaBoost (machine-learning algorithm). 3X–37X energy advantage is shown for the proposed SC compared with CMOS binary arithmetic ASIC and SC designs.
Shaodi Wang, Saptadeep Pal, Tianmu Li, Andrew Pan, Cecile Grezes, Pedram Khalili Amiri, Kang L. Wang, Puneet Gupta 0001
DATE3