VLDB 2026 Research / reviewers in the wild / expert
Tom Glint
dblp:282/4948
· DBLP profile ↗
9ranked-venue papers
9as first author
8since 2021 · last 2024
0000-0002-4053-0247ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CANSim: When to Utilize Synchronous and Asynchronous Routers in Large and Complex NoCsabstractAsynchronous routers offer benefits in latency, energy, and area over synchronous routers, but their large system design adoption is limited due to inadequate tool support for performance quantification. We introduce CANSim, a fast and accurate simulator for complex asynchronous and synchronous NoCs. Verified against synthesis models, CANSim models any graph-representable topology, asymmetric links, data-dependent delays, meta-stability, and supports synthetic and real-world traffic. It also accommodates power gating, hierarchical networks, and GALS-like networks. Our comparisons reveal that asynchronous NoCs provide lower packet latency at no-load conditions, but synchronous NoCs may outperform near saturation. On average, asynchronous NoCs offer up to 36% latency and 52% power benefits, with power-gating potentially reducing static power by 90%. Tom Glint, Manu Awasthi, Joycee Mekie |
ASPDAC | 1 |
| 2024 | Hardware-Software Co-Design of a Collaborative DNN Accelerator for 3D Stacked Memories with Multi-Channel DataabstractHardware accelerators are preferred over general-purpose processors for processing Deep Neural Networks (DNN) as the later suffer from power and memory walls. However, hardware accelerators designed as a separate logic chip from the memory still suffer from memory wall. Processing-in-memory accelerators, which try to overcome this memory wall by developing the compute elements as part of the memory structures, are highly constrained due to the memory manufacturing process. Near-data-processing (NDP) based hardware accelerator design is an alternative paradigm that could combine the benefit of high bandwidth, low access energy of processing-in-memory, and design flexibility of separate logic chip. However, NDP has area, data flow and thermal constraints, hindering high throughput designs. In this work, we propose an HBM3-based NDP accelerator that tackles the constraints of NDP with a hardware-software co-design approach. The proposed design takes only 50% area, delivers a speed-up of $3\times$, and is about $6 \times$ more energy efficient than state-of-the-art NDP hardware accelerator for inferencing workloads such as AlexNet, MobileNet, ResNet, and VGG without loss of accuracy. Tom Glint, Manu Awasthi, Joycee Mekie |
ASPDAC | 1 |
| 2024 | DeepFrack: A Comprehensive Framework for Layer Fusion, Face Tiling, and Efficient Mapping in DNN Hardware AcceleratorsabstractDeepFrack is a novel framework developed for enhancing energy efficiency and reducing latency in deep learning workloads executed on hardware accelerators. By optimally fusing layers and implementing an asymmetric tiling strategy, DeepFrack addresses the limitations of traditional layer-by-layer scheduling. The computational efficiency of our method is underscored by significant performance improvements seen across various deep neural network architectures such as AlexNet, VGG, and ResNets when run on Eyeriss and Simba accelerators. The reduction in latency (30 % to 40 %) and energy consumption (30 % to 50 %) are further enhanced by the efficient usage of the on-chip buffer and reduction of external memory bandwidth bottleneck. This work contributes to the ongoing efforts in designing more efficient hardware accelerators for machine learning workloads. Tom Glint, Mithil Pechimuthu, Joycee Mekie |
DATE | 1 |
| 2023 | Hardware-Software Codesign of DNN Accelerators Using Approximate Posit MultipliersabstractEmerging data intensive AI/ML workloads encounter memory and power wall when run on general-purpose compute cores. This has led to the development of a myriad of techniques to deal with such workloads, among which DNN accelerator architectures have found a prominent place. In this work, we propose a hardware-software co-design approach to achieve system-level benefits. We propose a quantized data-aware POSIT number representation that leads to a highly optimized DNN accelerator. We demonstrate this work on SOTA SIMBA architecture, extendable to any other accelerator. Our proposal reduces the buffer/storage requirements within the architecture and reduces the data transfer cost between the main memory and the DNN accelerator. We have investigated the impact of using integer, IEEE floating point, and posit multipliers for LeNet, ResNet and VGG NNs trained and tested on MNIST, CIFAR10 and ImageNet datasets, respectively. Our system-level analysis shows that the proposed approximate-fixed-posit multiplier when implemented on SIMBA architecture, achieves on average ~2.2× speed up, consumes ~3.1× less energy and requires ~3.2× less area, respectively, against the baseline SOTA architecture, without loss of accuracy (~±1%) Tom Glint, Kailash Prasad, Jinay Dagli, Krishil Gandhi, Vrajesh Patel, Joycee Mekie |
ASP-DAC | 1 |
| 2023 | REDRAW: Fast and Efficient Hardware Accelerator with Reduced Reads And Writes for 3D UNetabstractHardware Accelerators (HAs) proposed so far have been designed with a focus on 2D Convolutional Neural Networks (CNNs) and 3D CNNs using temporal data. To the best of our knowledge, there is no existing HA for 3D CNNs using spatial data. 3D UNet is a 3D CNN with significant applications in the medical domain. However, the total on-chip buffer size (>20 MB) required for the complete stationery approach of processing 3D UNet is cost prohibitive. In this work, we analyze the 3D UNet workload and propose a HA with an optimized memory hierarchy with a total on-chip buffer of less than 4 MB while conceding near theoretical minimum memory accesses required for processing 3D UNet. We demonstrate the efficiency of the proposed HA by comparing it with SOTA Simba architecture with the same number of MAC Units and show a 1.3× increase in TOPSwatt for an ISO-area design. Further, we revise the proposed architecture to increase the ratio of compute operations to memory operations and to meet the latency requirement of 3D UNet-based embedded applications. The revised architecture, compared against a dual instance of Simba, has similar latency. Against the dual instance of Simba, the proposed architecture achieves a 1.8 × increase in TOPS/watt in a similar area. Tom Glint, Manu Awasthi, Joycee Mekie |
DATE | 1 |
| 2023 | Analysis of Quantization Across DNN Accelerator Architecture ParadigmsabstractQuantization techniques promise to significantly reduce the latency, energy, and area associated with multiplier hardware. This work, to the best of our knowledge, for the first time, shows the system-level impact of quantization on SOTA DNN accelerators from different digital accelerator paradigms. Based on the placement of data and compute site, we identify SOTA designs from Conventional Hardware Accelerators (CHA), Near Data Processors (NDP), and Processing-in-Memory (PIM) paradigms and show the impact of quantization when inferencing CNN and Fully Connected Layer (FCL) workloads. We show that the 32-bit implementation of SOTA from PIM consumes less energy than the 8-bit implementation of SOTA from CHA for FCL, while the trend reverses for CNN workloads. Further, PIM has stable latency while scaling the word size while CHA and NDP suffer 20% to$2\times$slow down for doubling word size. Tom Glint, Chandan Kumar Jha 0001, Manu Awasthi, Joycee Mekie |
DATE | 1 |
| 2023 | Impact of Optimal Design Point on Performance Metrics of DNN accelerators in FPGAabstractDue to their flexibility, FPGAs are used to deploy Deep Neural Network Accelerators (DA) at various compute locations such as servers and edge computes. In this work, we show the impact of choosing architectural parameters on performance metrics for SOTA DA implemented on FPGA, and the optimal design point for maximum compute throughput. Tom Glint, Daniel Giftson, Gaurav Shah, Vrajesh Patel, Ruchit Chudasama, Sukanya More, Joycee Mekie |
ISPASS | 1 |
| 2023 | Analysis of Conventional, Near-Memory, and In-Memory DNN AcceleratorsabstractVarious DNN accelerators based on Conventional compute Hardware Accelerator (CHA), Near-Data-Processing (NDP) and Processing-in-Memory (PIM) paradigms have been proposed to meet the challenges of inferencing Deep Neural Networks (DNNs). To the best of our knowledge, this work aims to perform the first quantitative as well as qualitative comparison among the state-of-the-art accelerators from each digital DNN accelerator paradigm. Our study provides insights into selecting the best architecture for a given DNN workload. We have used workloads of the MLPerf Inference benchmark. We observe that for Fully Connected Layer (FCL) DNNs, PIM-based accelerator is 21× and 3× faster than CHA and NDP-based accelerator respectively. However, NDP is 9× and 2.5× more energy efficient than CHA and PIM for FCL. For Convolutional Neural Network (CNN) workloads, CHA is 10% and 5× faster than NDP and PIM-based accelerator respectively. Further, CHA is 1.5× and 6× more energy efficient than NDP and PIM-based accelerators respectively. Tom Glint, Chandan Kumar Jha 0001, Manu Awasthi, Joycee Mekie |
ISPASS | 1 |
| 2020 | ANSim: A Fast and Versatile Asynchronous Network-On-Chip SimulatorabstractIn a large chip, an asynchronous Network-on-Chip (NoC) is a suitable candidate for establishing an interconnection network between varied components. Architectural level simulation is an accepted methodology for evaluating such systems. In this paper, we propose a fast and versatile Asynchronous Network-on-Chip (NoC) Simulator - ANSim, which brings down the simulation time by 25 ×, compared to the state of the art simulators. It can model and analyze all synchronous, asynchronous, and mixed synchronous-asynchronous system of cores connected through NoC. ANSim can model routers with different delays, routers with asynchronous arbitration, connected in a wide range of topologies. ANSim supports individual routers modeled to have varying timing constraints. Further, it supports synthetic and real-workloads, and produces system-level latency, throughput, power, power-gating, and arbitration reports. ANSim has been verified against RTL models of NoCs, and other RTL verified simulators. An open-source synchronous NoC router, TNoC, and its asynchronous derivative are used to demonstrate ANSim's usefulness and features. Tom Glint, Jitesh Sah, Manu Awasthi, Joycee Mekie |
ICCD | 1 |