EDBT 2026 Demo / reviewers in the wild / expert
Yu Yang 0020
dblp:16/4505-20
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2396-3590ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modeling and Scheduling of Composable Instruction SetabstractState-of-the-art hardware accelerators with custom instruction set architectures (ISAs) are widely used in AI/ML applications. Despite the outstanding progress of modern accelerators, their performance potential is limited by their ISA. Recently, research around composable instruction sets (CIS) has emerged with the aim of improving the efficiency of modern accelerator ISAs. However, while CIS significantly outperforms existing ISAs with near-optimal PE utilization, no existing instruction scheduling approach supports the temporal composability requirements to correctly schedule CIS programs. The goal of this work is to propose a scalable scheduling algorithm that can meet CIS requirements. We propose a novel timing model structure built upon five fundamental concepts - operations, events, transformations, anchors, and constraints - that accurately capture CIS timing behavior and constraints. Using this model, we automatically formulate scheduling problems that can be solved using constraint programming (CP) solvers. In addition, we propose a methodology to synchronize scheduled instructions and generate functional CIS assembly code. Finally, we design experiments to study the scalability and quality of our scheduling approach. Across multiple dimensions of complexity, our approach scales linearly and produces near-optimal scheduling solutions. Yu Yang 0020, Paul Delestrac, Ahmed Hemani |
DSD | 1 |
| 2024 | FPGA-Based HPC for Associative Memory SystemabstractAssociative memory plays a crucial role in the cognitive capabilities of the human brain. The Bayesian Confidence Propagation Neural Network (BCPNN) is a cortex model capable of emulating brain-like cognitive capabilities, particularly associative memory. However, the existing GPU-based approach for BCPNN simulations faces challenges in terms of time overhead and power efficiency. In this paper, we propose a novel FPGA-based high performance computing (HPC) design for the BCPNN-based associative memory system. Our design endeavors to maximize the spatial and timing utilization of FPGA while adhering to the constraints of the available hardware resources. By incorporating optimization techniques including shared parallel computing units, hybrid-precision computing for a hybrid update mechanism, and the globally asynchronous and locally synchronous (GALS) strategy, we achieve a maximum network size of $150 \times 10$ and a peak working frequency of 100 MHz for the BCPNN-based associative memory system on the Xilinx Alveo U200 Card. The tradeoff between performance and hardware overhead of the design is explored and evaluated. Compared with the GPU counterpart, the FPGA-based implementation demonstrates significant improvements in both performance and energy efficiency, achieving a maximum latency reduction of $33.25 \times$, and a power reduction of over $6.9 \times$, all while maintaining the same network configuration. Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani, Anders Lansner, Jiawei Xu 0002, Lirong Zheng 0001, Zhuo Zou |
ASPDAC | 3 |
| 2024 | Exploration of Custom Floating-Point Formats: A Systematic ApproachabstractThe remarkable advancements in AI algorithms over the past three decades have been paralleled by an exponential growth in their complexity, with parameter counts soaring from 60,000 in LeNet during the late 1980s to a staggering 175 billion in ChatGPT 3.0. To mitigate this surge in memory footprint, approximate computing has emerged as a promising strategy, focusing on deploying the minimal resolution necessary to maintain acceptable accuracy. Yet, current practices are hindered by two major challenges: a) the process of identifying the optimal resolution and representation format for each tensor remains a manual, ad hoc task, and b) the representation, typically in floating point (FP) format, is confined to standardized norms predominantly supported by commercial-off-the-shelf (COTS) products like GPUs. This paper tackles these issues by introducing a systematic approach to exploring the FP representation design space to find the ideal FP format for each tensor, thereby leveraging the full potential of FP quantization techniques. It is designed for custom hardware, enabling access to arbitrary FP formats, but also allows users to limit their exploration to standard FP formats, making it compatible with COTS. Additionally, the proposed method explores the Block Floating-Point (BFP) and automatically decides on the size of the blocks. A heuristic-based search method is proposed to handle the large design space. The proposed approach is general, and the heuristic is not biased towards any specific category of algorithms. We apply this method to a Self-Organizing Map (SOM) for bacterial genome identification and LeNet-5 neural network, demonstrating a significant reduction in memory footprint by around 94% and 96%, respectively, compared to the conventional 32-bit FP baseline. Saba Yousefzadeh, Yu Yang 0020, Astile Peter, Dimitrios Stathis 0001, Ahmed Hemani |
DSD | 2 |
| 2024 | Multi-objective preference-free exact design space exploration of static DSP on multicore platformsabstractA challenge in designing resource-constrained embedded systems for digital signal processing (DSP) is their complexity due to their vast design spaces, where only a fraction of implementations are feasible or optimal. A crucial tool to aid in this challenge is automated design space exploration (DSE). However, no exact, multi-objective, and preference-free DSE approach exists for DSP applications on resource-constrained embedded platforms.We propose a novel DSE solution with these ideal characteristics to perform DSE of analyzable DSP applications for tile-based multiprocessing embedded platforms. Our proposal harmonizes the exactness of constraint programming (CP) and the exploration efficiency of genetic algorithms (GA). Through this synergy, no single-objective reduction strategy or a priori objective preferences is required.We evaluate the proposal through state-of-the-art single-objective case studies and multi-objective case studies inspired by these. The evaluations show that our proposal improves the single-objective state-of-the-art and finds high-quality approximate Pareto-frontiers for the multi-objective case study. Therefore, our proposal is a more performant single-objective DSE solution than the state-of-the-art, and it is the first exact, multi-objective, and preference-free DSE approach for the problem addressed. Rodolfo Jordão, Fahimeh Bahrami, Yu Yang 0020, Matthias Becker 0004, Ingo Sander, Kathrin Rosvall |
FDL | 3 |
| 2022 | Reducing the Configuration Overhead of the Distributed Two-level Control SystemabstractWith the growing demand for more efficient hardware accelerators for streaming applications, a novel Coarse-Grained Reconfigurable Architecture (CGRA) that uses a Dis-tributed Two-Level Control (D2LC) system has been proposed in the literature. Even though the highly distributed and parallel structure makes it fast and energy-efficient, the single-issue instruction channel between the level-l and level-2 controller in each D2LC cell becomes the bottleneck of its performance. In this paper, we improve its design to mimic a multi-issued architecture by inserting shadow instruction buffers between the level-l and level-2 controllers. Together with a zero-overhead hardware loop, the improved D2LC architecture can enable efficient overlap between loop iterations. We also propose a complete constraint programming based instruction scheduling algorithm to support the above hardware features. The experiment result shows that the improved D2LC architecture can achieve up to 25% of reduction on the instruction execution cycles and 35% reduction on the energy-delay product. Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani |
DATE | 1 |
| 2021 | Approximate computation of post-synaptic spikes reduces bandwidth to synaptic storage in a model of cortex
Dimitrios Stathis 0001, Yu Yang 0020, Ahmed Hemani, Anders Lansner |
DATE | 2 |
| 2021 | Scheduling Persistent and Fully Cooperative InstructionsabstractParallel, distributed two-level control system has been adopted in streaming application accelerators that implement atomic vector operations. Each instruction of such architecture deals with one aspect (arithmetic, interconnect, storage, etc.) of an atomic vector operation. Such instructions are persistent and fully cooperative. Their lifetimes vary because of the vector size and the degree of parallelism. More complex constraints are also required to express the cooperation among these instructions. The conventional instruction behavior models are no longer suitable for such instructions. Therefore, we develop a novel instruction behavior model to address the scheduling aspect of the instruction set required by such architecture. Based on the behavior model, we formally define the scheduling problem and formulate it as a constraint satisfaction optimization problem (CSOP). However, the naive CSOP formulation quickly becomes unscalable. Thus a heuristic enhanced scheduling algorithm is introduced to make the CSOP approach scalable. The enhanced algorithm’s scalability is validated by a large set of experiments varying in problem size. Yu Yang 0020, Ahmed Hemani, Kolin Paul |
DSD | 1 |
| 2021 | Scheduling Persistent and Fully Cooperative InstructionsabstractThe distributed two-level control (D2LC) system has been adopted in streaming application accelerators that implement atomic vector operations. Each instruction of such architecture deals with one aspect (arithmetic, interconnect, storage, etc.) of an atomic vector operation. The D2LC architecture is different from traditional computer architecture such as MMX or VLIW. The D2LC architecture consists of many cells interconnected via a NoC. In each cell, there is a two-level controller. The level-1 controller sends instructions to configure their level-2 controllers. Each level-2 controller, once configured, works as an independent finite state machine (FSM). It manages a datapath for specific functionality, including computation, interconnection, as well as storage. We can say that the level-1 controllers implement threads while the level-2 controllers implement micro-threads. For generality, these microthreads, even though distributed in different cells, can be grouped by the interconnection units and orchestrated to implement a larger functionality. The compiler of D2LC architecture needs to schedule these micro-threads correctly so that the larger functionality is reached. Yu Yang 0020, Ahmed Hemani, Kolin Paul |
FCCM | 1 |