EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Wang 0005
dblp:27/5665-5
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-5883-7327ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Eidetic: An In-Memory Matrix Multiplication Accelerator for Neural NetworksabstractThis paper presents theEideticarchitecture, which is an SRAM-based ASIC neural network accelerator that eliminates the need to continuously load weights from off-chip, while also minimizing the need to go off chip for intermediate results. Using in-situ arithmetic in the SRAM arrays, this architecture can supports a variety of precision types allowing for effective inference. We also present different data mapping policies for matrix-vector based networks (RNN and MLP) on theEideticarchitecture and describe the tradeoffs involved. With this architecture, multiple layers of a network can be concurrently mapped, storing both the layer weights and intermediate results on-chip, removing the energy and latency penalty of off-chip memory accesses. We evaluateEideticon Google's Neural Machine Translation System (GNMT) encoder and demonstrate a 17.20× increase in throughput and 7.77× reduction in average latency over a single TPUv2 chip. Charles Eckert, Arun Subramaniyan 0001, Xiaowei Wang 0005, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
IEEE Trans. Computers | 3 |
| 2021 | Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAsabstractThe density of FPGA on-chip memory has been continuously increasing with modern FPGAs having thousands of block RAMs (BRAMs) distributed across their reconfigurable fabric. These distributed BRAMs can provide a tremendous amount of on-chip bandwidth for efficient acceleration of data-intensive applications. In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our proposed architectural change results in 1.6× and 2.3× increase in the peak multiply-accumulate throughput of a large Stratix 10 FPGA, at a minimal cost of only 1.8% increase in the FPGA die size and no change to the BRAM's interface to the programmable routing. Then, we present RIMA, a reconfigurable in-memory accelerator architecture for deep learning (DL) inference. RIMA exploits the proposed compute-capable BRAMs and the FPGA's reconfigurability to achieve 1.25× and 3× higher performance compared to the state-of-the-art Brainwave DL soft processor for 8-bit integer and block floating-point precisions, respectively. In addition, RIMA implemented on a Stratix 10 FPGA enhanced with compute-capable BRAMs can achieve an order of magnitude higher performance compared to a same-generation GPU. Xiaowei Wang 0005, Vidushi Goyal, Jiecao Yu, Valeria Bertacco, Andrew Boutros, Eriko Nurvitadhi, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
FCCM | 1 |
| 2021 | Cache Compression with Efficient in-SRAM Data ComparisonabstractWe present a novel cache compression method that leverages the fine-grained data duplication across cache lines. We leverage the XOR operation of the in-SRAM bit-line computing peripherals, to search for compressible data over a wide range of data locations on cache, reducing the data movement requirements. To reduce the decompression latency, we design specialized compression schemes by fetching the data with the same parallelism as the original cache, according to the architecture of the last-level cache slice. The proposed compression method achieves a 2.05× compression ratio on average (up to 67×), and 4.73% of speedup on average (up to 29%), over the SPEC2006 benchmarks. Xiaowei Wang 0005, Charles Augustine, Eriko Nurvitadhi, Ravi R. Iyer 0001, Li Zhao 0002, Reetuparna Das |
NAS | 1 |
| 2020 | High Throughput CNN Inference and Training with In-Cache ComputationabstractWe present an architecture for CNN training and batched inference with in-cache computing. Specifically targeting the high throughput requirements of the training and inference workloads, we propose novel resource partitioning and work scheduling strategies to balance the in-place computing and data storage requirements on the last level cache. Further, we propose a compression mechanism to reduce the data movement between the cache and the main memory. For ResNet-50, the proposed architecture achieves 2.2× better throughput, compared to a state-of-the-art CNN inference engine with in-cache computing [1] with a baseline scheduling policy, and a training throughput of 65.4 images per second. Xiaowei Wang 0005 |
ICCD | 1 |
| 2020 | Neksus: An Interconnect for Heterogeneous System-In-Package ArchitecturesabstractIn the embedded systems industry today, skyrocketing design and manufacturing costs of Systems-on-Chip (SoCs) are key limiting factors for growth. Emerging 2.5D-based System-In-Package (SiP) architectures show potential to lower these costs by enabling the reuse of hard core units and providing higher manufacturing yields due to small chiplet sizes.In this paper, we present Neksus, a novel architecture designed to lower SiP manufacturing costs, support modular "plug-and-play" chiplet integration, and leverage the unique properties of interposers. Key to Neksus is a new dedicated interconnect chiplet that addresses the limitations of SiP packaging technology by leveraging direct communication over a mini-chain IP-connection topology. In addition to satisfying SiP technology constraints, because our mini-chain design provides high-bandwidth IP-to-IP communication, it is particularly well-suited for bandwidth-intensive mobile applications. Our evaluation shows Neksus provides up to 28% performance improvement and 31% energy savings over recent SiP architecture. Vidushi Goyal, Xiaowei Wang 0005, Valeria Bertacco, Reetuparna Das |
IPDPS | 2 |
| 2019 | Bit Prudent In-Cache Acceleration of Deep Convolutional Neural NetworksabstractWe propose Bit Prudent In-Cache Acceleration of Deep Convolutional Neural Networks - an in-SRAM architecture for accelerating Convolutional Neural Network (CNN) inference by leveraging network redundancy and massive parallelism. The network redundancy is exploited in two ways. First, we prune and fine-tune the trained network model and develop two distinct methods - coalescing and overlapping - to run inferences efficiently with sparse models. Second, we propose an architecture for network models with a reduced bit width by leveraging bit-serial computation. Our proposed architecture achieves a 17.7×/3.7× speedup over server class CPU/GPU, and a 1.6× speedup compared to the relevant in-cache accelerator, with 2% area overhead each processor die, and no loss on top-1 accuracy for AlexNet. With a relaxed accuracy limit, our tunable architecture achieves higher speedups. Xiaowei Wang 0005, Jiecao Yu, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
HPCA | 1 |
| 2018 | Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural NetworksabstractThis paper presents the Neural Cache architecture, which re-purposes cache structures to transform them into massively parallel compute units capable of running inferences for Deep Neural Networks. Techniques to do in-situ arithmetic in SRAM arrays, create efficient data mapping and reducing data movement are proposed. The Neural Cache architecture is capable of fully executing convolutional, fully connected, and pooling layers in-cache. The proposed architecture also supports quantization in-cache. Our experimental results show that the proposed architecture can improve inference latency by 8.3× over state-of-art multi-core CPU (Xeon E5), 7.7× over server class GPU (Titan Xp), for Inception v3 model. Neural Cache improves inference throughput by 12.4× over CPU (2.2× over GPU), while reducing power consumption by 50% over CPU (53% over GPU). Charles Eckert, Xiaowei Wang 0005, Arun Subramaniyan 0001, Ravi R. Iyer 0001, Dennis Sylvester, David T. Blaauw, Reetuparna Das |
ISCA | 2 |