EDBT 2026 Demo / reviewers in the wild / expert
Yongming Shen 0001
dblp:156/6476
· DBLP profile ↗
7ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 42% Processor architecture and microarchitecture · 32% Cloud and datacenter computing · 21% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.6 | 2 | 2017 | Maximizing CNN Accelerator Efficiency Through Resource Partitioning · ISCA 2017 Storage-Efficient Batching for Minimizing Bandwidth of Fully-Connected Neural Network Layers (Abstract Only) · FPGA 2017 |
Cloud and datacenter computing › resource allocation › resource allocation policy
resource partitioning |
0.3 | 1 | 2017 | Maximizing CNN Accelerator Efficiency Through Resource Partitioning · ISCA 2017 |
Processor architecture and microarchitecture
speculative execution |
0.2 | 1 | 2015 | Architectural Support for Dynamic Linking · ASPLOS 2015 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.1 | 1 | 2017 | Storage-Efficient Batching for Minimizing Bandwidth of Fully-Connected Neural Network Layers (Abstract Only) · FPGA 2017 |
Methods — techniques the papers use, named apart from their topics
trampoline elimination · 0.4hardware speculation · 0.4batching · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Medusa: A Scalable Interconnect for Many-Port DNN Accelerators and Wide DRAM Controller InterfacesabstractTo cope with the increasing demand and computational intensity of deep neural networks (DNNs), industry and academia have turned to accelerator technologies. In particular, FPGAs have been shown to provide a good balance between performance and energy efficiency for accelerating DNNs. While significant research has focused on how to build efficient layer processors, the computational building blocks of DNN accelerators, relatively little attention has been paid to the on-chip interconnects that sit between the layer processors and the FPGA's DRAM controller. We observe a disparity between DNN accelerator interfaces, which tend to comprise many narrow ports, and FPGA DRAM controller interfaces, which tend to be wide buses. This mismatch causes traditional interconnects to consume significant FPGA resources. To address this problem, we designed Medusa: an optimized FPGA memory interconnect which transposes data in the interconnect fabric, tailoring the interconnect to the needs of DNN layer processors. Compared to a traditional FPGA interconnect, our design can reduce LUT and FF use by 4.7x and 6.0x, and improves frequency by 1.8x. Yongming Shen 0001, Tianchu Ji, Michael Ferdman, Peter A. Milder |
FPL | 1 |
| 2017 | Escher: A CNN Accelerator with Flexible Buffering to Minimize Off-Chip TransferabstractConvolutional neural networks (CNNs) are used to solve many challenging machine learning problems. Interest in CNNs has led to the design of CNN accelerators to improve CNN evaluation throughput and efficiency. Importantly, the bandwidth demand from weight data transfer for modern large CNNs causes CNN accelerators to be severely bandwidth bottlenecked, prompting the need for processing images in batches to increase weight reuse. However, existing CNN accelerator designs limit the choice of batch sizes and lack support for batch processing of convolutional layers. We observe that, for a given storage budget, choosing the best batch size requires balancing the input and weight transfer. We propose Escher, a CNN accelerator with a flexible data buffering scheme that ensures a balance between the input and weight transfer bandwidth, significantly reducing overall bandwidth requirements. For example, compared to the state-of-the-art CNN accelerator designs targeting a Virtex-7 690T FPGA, Escher reduces the accelerator peak bandwidth requirements by 2.4x across both fully-connected and convolutional layers on fixed-point AlexNet, and reduces convolutional layer bandwidth by up to 10.5x on fixed-point GoogleNet. Yongming Shen 0001, Michael Ferdman, Peter A. Milder |
FCCM | 1 |
| 2017 | Storage-Efficient Batching for Minimizing Bandwidth of Fully-Connected Neural Network Layers (Abstract Only)
Yongming Shen 0001, Michael Ferdman, Peter A. Milder |
FPGA | 1 |
| 2017 | Maximizing CNN Accelerator Efficiency Through Resource Partitioning
Yongming Shen 0001, Michael Ferdman, Peter A. Milder |
ISCA | 1 |
| 2016 | Overcoming resource underutilization in spatial CNN acceleratorsabstractConvolutional neural networks (CNNs) are revolutionizing a variety of machine learning tasks, but they present significant computational challenges. Recently, FPGA-based accelerators have been proposed to improve the speed and efficiency of CNNs. Current approaches construct an accelerator optimized to maximize the overall throughput of iteratively computing the CNN layers. However, this approach leads to dynamic resource underutilization because the same accelerator is used to compute CNN layers of radically varying dimensions. We present a new CNN accelerator design that improves the dynamic resource utilization. Using the same FPGA resources, we build multiple accelerators, each specialized for specific CNN layers. Our design achieves 1.3× higher throughput than the state of the art when evaluating the convolutional layers of the popular AlexNet CNN on a Xilinx Virtex-7 FPGA. Yongming Shen 0001, Michael Ferdman, Peter A. Milder |
FPL | 1 |
| 2016 | Demystifying cloud benchmarkingabstractThe popularity of online services has grown exponentially, spurring great interest in improving server hardware and software. However, conducting research on servers has traditionally been challenging due to the complexity of setting up representative server configurations and measuring their performance. Recent work has eased the effort of benchmarking servers by making benchmarking software and benchmarking instructions readily available to the research community. Unfortunately, the existing benchmarks are a black box; their users are expected to trust the design decisions made in the construction of these benchmarks with little justification and few cited sources. In this work, we have attempted to overcome this problem by building new server benchmarks for three popular network-intensive workloads: video streaming, web serving, and object caching. This paper documents the benchmark construction process, describes the software, and provides the resources we used to justify the design decisions that make our benchmarks representative for system-level studies. Tapti Palit, Yongming Shen 0001, Michael Ferdman |
ISPASS | 2 |
| 2015 | Architectural Support for Dynamic LinkingabstractAll software in use today relies on libraries, including standard libraries (e.g., C, C++) and application-specific libraries (e.g., libxml, libpng). Most libraries are loaded in memory and dynamically linked when programs are launched, resolving symbol addresses across the applications and libraries. Dynamic linking has many benefits: It allows code to be reused between applications, conserves memory (because only one copy of a library is kept in memory for all the applications that share it), and allows libraries to be patched and updated without modifying programs, among numerous other benefits. However, these benefits come at the cost of performance. For every call made to a function in a dynamically linked library, a trampoline is used to read the function address from a lookup table and branch to the function, incurring memory load and branch operations. Static linking avoids this performance penalty, but loses all the benefits of dynamic linking. Given its myriad benefits, dynamic linking is the predominant choice today, despite the performance cost. In this work, we propose a speculative hardware mechanism to optimize dynamic linking by avoiding executing the trampolines for library function calls, providing the benefits of dynamic linking with the performance of static linking. Speculatively skipping the memory load and branch operations of the library call trampolines improves performance by reducing the number of executed instructions and gains additional performance by reducing pressure on the instruction and data caches, TLBs, and branch predictors. Because the indirect targets of library call trampolines do not change during program execution, our speculative mechanism never misspeculates in practice. We evaluate our technique on real hardware with production software and observe up to 4% speedup using only 1.5KB of on-chip storage. Varun Agrawal, Abhiroop Dabral, Tapti Palit, Yongming Shen 0001, Michael Ferdman |
ASPLOS | 4 |