EDBT 2026 Demo / reviewers in the wild / expert
Fuyu Wang 0001
dblp:88/8346-1
· DBLP profile ↗
10ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-3165-873XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsabstractSpatial accelerators are promising to satiate the growing demands for performance and energy efficiency in deep neural networks (DNNs). Due to the speed gap between onchip compute cores and off-chip memory bandwidth, common DNNs suffer from poor operational intensity and are increasingly memory-bound. While operator fusion has shown potential in alleviating this bottleneck, existing approaches suffer from two key limitations. They rely on predefined fusion templates before tensor mapping and impose tile constraints during mapping. As a result, they overlook the potential of fusing more operators and lead to sub-optimal performance. In this paper, we propose a segment fusion dataflow optimization framework called SFD. Central to this framework is the dataflow abstraction that enables template-free operator fusion after mapping and supports tile constraint relaxation through tile scheduling. Based on this abstraction, we first introduce a memory-centric mapper, which defines a design space and incorporates an algorithm to facilitate design space exploration (DSE). Then we propose an analytical network segmenter, which leverages mapping results to analyze tensor lifetimes and on-chip memory usage, fusing operators into variable-length segments. Finally, we introduce a dependency-aware tile scheduler, which develops a priority queue for each segment to ensure correct execution order. Extensive experiments with different DNNs demonstrate SFD achieves$1.4 \times$to$2.2 \times$speedup for spatial accelerators over state-of-the-art fusion frameworks. Fuyu Wang 0001, Minghua Shen, Yufei Ding 0001, Nong Xiao 0001, Yutong Lu |
HPCA | 1 |
| 2025 | Poros: One-Level Architecture-Mapping Co-Exploration for Tensor AlgorithmsabstractTensor algorithms increasingly rely on specialized accelerators to meet growing performance and efficiency demands. Given the rapid evolution of these algorithms and the high cost of designing accelerators, automated solutions for jointly optimizing both architectures and mappings have gained attention. However, the joint design space is non-convex and non-smooth, hindering the finding of optimal or near-optimal designs. Moreover, prior work conducts two-level exploration, resulting in a combinatorial explosion. In this paper, we propose Poros, a one-level architecture-mapping co-exploration framework. Poros directly explores a batch of architecture-mapping configurations and evaluates their performance. It then exploits reinforcement learning to perform gradient-based search in the non-smooth joint design space. By sampling from the policy, Poros keeps exploring new actions to address non-convexity. Experimental results demonstrate that Poros achieves up to 5.32 × and 2.15 × better EDP compared with hand-designed accelerators and state-of-the-art automatic approaches respectively. Through one-level exploration scheme, Poros also converges at least 20% faster than other approaches. Fuyu Wang 0001, Minghua Shen |
DATE | 1 |
| 2025 | Ceiba: An Efficient and Scalable DNN Scheduler for Spatial AcceleratorsabstractSpatial accelerators are domain-specific architectures to elevate performance and energy efficiency for deep neural networks (DNNs). They also bring a large number of schedule parameters to determine computation and data movement patterns of DNNs. Previous works formulate the schedule problem as design space exploration or integer linear programming. However, these advanced techniques face the challenge of efficiency or scalability. In this article, we propose Ceiba, which is a deep reinforcement learning-based DNN scheduler for spatial accelerators. Ceiba observes the running DNN computation as well as the spatial architecture to make schedule decisions. Then, Ceiba receives a reward to learn and produce the best-fit policy. To provide efficient and scalable scheduling, Ceiba constructs a DNN-architecture-specific action space. It is defined by upper and lower bounds to exclude invalid and sub-optimal schedule candidates. Extensive experiments demonstrate that Ceiba generally provides better performance for spatial accelerators under a fixed number of searching steps or a fixed amount of time. Specifically, Ceiba achieves an average 2.2× speedup for the Simba accelerator, compared with the state-of-the-art scheduler. When scaling the batch size and the hardware architecture up by 64×, the performance gains of Ceiba are 1.8× and 1.2× on average, respectively. Moreover, Ceiba exhibits better scalability for the Eyeriss accelerator. Fuyu Wang 0001, Minghua Shen, Yutong Lu, Nong Xiao 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | TileMap: Mapping Multi-Head Attention on Spatial Accelerators with Tile-based AnalysisabstractMulti-head attention demonstrates enormous potential but incurs high costs (e.g., memory access) when deploying transformer models on spatial accelerators. Operator fusion is prevalent in conventional deep learning mappers to reduce off-chip memory access. However, designing operator-fusion mapping for multi-head attention is challenging. There exist strict data dependency between operators and hardware resource constraints of accelerators. In this paper, we propose TileMap, an operator-fusion mapping framework for multi-head attention on spatial accelerators. Central to this framework is tile-based analysis that can automatically satisfy both data dependency and resource constraints. Based on this analysis, we construct a mapping design space, which significantly prunes invalid candidates. We then propose an RL-based searching algorithm to explore the mapping space. The RL agent regards the mapping space as action space and further optimizes it to preserve all constraints. Experiments show TileMap achieves$1.8\times$to$3.1\times$speedup on different spatial accelerators, relative to state-of-the-art mapping approaches. Fuyu Wang 0001, Minghua Shen |
ICCD | 1 |
| 2024 | Soter: Analytical Tensor-Architecture Modeling and Automatic Tensor Program Tuning for Spatial AcceleratorsabstractSpatial accelerator is a specialized hardware to provide noticeable performance speedup for tensor computations. It also brings a challenge to map tensor computations on spatial accelerators. Auto-tuning compiler is one of the most promising directions for tensor mapping. However, existing auto-tuning compilers suffer from either numerous invalid and inefficient programs or inaccurate evaluation of incomplete programs, leading to sub-optimal performance.In this paper, we propose Soter, a novel auto-tuning tensor compilation framework for spatial accelerators. The key is to perform exploration in a both valid and efficient program design space and perform optimization according to accurate evaluation of complete programs. First, we design an analytical model to generate a high-quality program design space, which excludes invalid and inefficient programs. Second, we design an automatic program tuner to efficiently explore the program space and avoid evaluating incomplete programs. Finally, we coordinate the model and the tuner to further improve the quality of program space. The program space is identified by the model and is updated during the exploration of tuner. On average, Soter achieves 2.1× to 3.5× speedup over the state-of-the-art tensor compilers. Moreover, Soter shows better scalability for larger-scale tensor computations and spatial architectures. Fuyu Wang 0001, Minghua Shen, Yufei Ding 0001, Nong Xiao 0001 |
ISCA | 1 |
| 2024 | TensorMap: A Deep RL-Based Tensor Mapping Framework for Spatial AcceleratorsabstractThe mapping of tensor computation is a complex and important process for spatial accelerators. Today's mapping works depend on hand-tuned kernel libraries or search-based heuristics from human experts. The former is time-intensive while the latter easily leads to sub-optimal performance. In this paper, we propose TensorMap, a deep reinforcement learning (RL)-based mapping framework for tensor computations on spatial accelerators. We propose a sequential generation mode for mapping optimization and construct a coarse-grained action space to reduce the complexity of the mapping search space. An efficient policy network is devised to optimize mapping primitives in the RL-based search. We then propose a stop signal that is sampled fromBernoullidistribution to facilitate multi-level loop unrolling for spatial accelerators. Finally, a genetic algorithm is employed to further refine the optimized mappings. In the experiments, we demonstrate TensorMap's ability for different spatial accelerators with various tensor computations. On TPU, TensorMap provides 2.6$\times$, 2.7$\times$, and 2.4$\times$better energy-delay product (EDP) on average compared with FlexTensor, Ansor, and AMOS respectively. On Eyeriss, TensorMap provides 2.1$\times$, 1.8$\times$, and 1.7$\times$better EDP on average compared with FlexTensor, Ansor, and AMOS respectively. Fuyu Wang 0001, Minghua Shen, Yutong Lu, Nong Xiao 0001 |
IEEE Trans. Computers | 1 |
| 2023 | Automatic Kernel Generation for Large Language Models on Deep Learning AcceleratorsabstractLarge language model (LLM) is a promising trend to sustain accuracy growth with billions of parameters for various application domains. Deep learning (DL) accelerators that are typically designed in spatial architecture are potential platforms to handle the substantial computational demands of LLMs. Exploiting DL accelerators for LLMs require high-performance kernels, which are manually optimized or automatically generated. While prior automatic works reduce development costs, they fail to find optimal or near-optimal kernels. This is because the kernel design space is large containing many invalid kernels, and non-convex containing many local minima. In this paper, we propose an automatic kernel generation framework for large language models on deep learning accelerators. The key idea is the reinforcement learning (RL) formulation to generate a kernel with multi-step decision-making. We first develop a high-quality action space to satisfy architectural constraints of accelerator. Then, we provide a practical RL implementation by devising a policy network with Transformer and variance reduction techniques for gradients. Experimental results show our framework achieves average 3.5× speedup on TensorCore compared with exploration-based Ansor; 2.6× speedup on Simba compared with solver-based CoSA. Also, our framework achieves better energy efficiency compared to the state-of-the-art works. Fuyu Wang 0001, Minghua Shen |
ICCAD | 1 |
| 2022 | Unifying Relational Sentence Generation and Retrieval for Medical Image Report CompositionabstractBeyond generating long and topic-coherent paragraphs in traditional captioning tasks, the medical image report composition task poses more task-oriented challenges by requiring both the highly accurate medical term diagnosis and multiple heterogeneous forms of information, including impression and findings. Current methods often generate the most common sentences due to dataset bias for the individual case, regardless of whether the sentences properly capture key entities and relationships. Such limitations severely hinder their applicability and generalization capability in medical report composition, where the most critical sentences lie in the descriptions of abnormal diseases that are relatively rare. Moreover, some medical terms appearing in one report are often entangled with each other and co-occurred, for example, symptoms associated with a specific disease. To enforce the semantic consistency of medical terms to be incorporated into the final reports and encourage the sentence generation for rare abnormal descriptions, we propose a novel framework that unifies template retrieval and sentence generation to handle both common and rare abnormality while ensuring the semantic coherency among the detected medical terms. Specifically, our approach exploits hybrid-knowledge co-reasoning: 1) explicit relationships among all abnormal medical terms to induce the visual attention learning and topic representation encoding for better topic-oriented symptoms descriptions and 2) adaptive generation mode that changes between the template retrieval and sentence generation according to a contextual topic encoder. The experimental results on two medical report benchmarks demonstrate the superiority of the proposed framework in terms of both human and metrics evaluation. Fuyu Wang 0001, Xiaodan Liang, Liang Lin 0004 |
IEEE Trans. Cybern. | 1 |
| 2021 | Medical-VLBERT: Medical Visual Language BERT for COVID-19 CT Report Generation With Alternate LearningabstractMedical imaging technologies, including computed tomography (CT) or chest X-Ray (CXR), are largely employed to facilitate the diagnosis of the COVID-19. Since manual report writing is usually too time-consuming, a more intelligent auxiliary medical system that could generate medical reports automatically and immediately is urgently needed. In this article, we propose to use the medical visual language BERT (Medical-VLBERT) model to identify the abnormality on the COVID-19 scans and generate the medical report automatically based on the detected lesion regions. To produce more accurate medical reports and minimize the visual-and-linguistic differences, this model adopts an alternate learning strategy with two procedures that are knowledge pretraining and transferring. To be more precise, the knowledge pretraining procedure is to memorize the knowledge from medical texts, while the transferring procedure is to utilize the acquired knowledge for professional medical sentences generations through observations of medical images. In practice, for automatic medical report generation on the COVID-19 cases, we constructed a dataset of 368 medical findings in Chinese and 1104 chest CT scans from The First Affiliated Hospital of Jinan University, Guangzhou, China, and The Fifth Affiliated Hospital of Sun Yat-sen University, Zhuhai, China. Besides, to alleviate the insufficiency of the COVID-19 training samples, our model was first trained on the large-scale Chinese CX-CHR dataset and then transferred to the COVID-19 CT dataset for further fine-tuning. The experimental results showed that Medical-VLBERT achieved state-of-the-art performances on terminology prediction and report generation with the Chinese COVID-19 CT dataset and the CX-CHR dataset. The Chinese COVID-19 CT dataset is available at https://covid19ct.github.io/. Guangyi Liu 0005, Yinghong Liao, Fuyu Wang 0001, Lu Zhang 0051, Xiaodan Liang, Shaolin Li, Zhen Li 0026, Shuixing Zhang, Shuguang Cui |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Simultaneous Lung Field Detection and Segmentation for Pediatric Chest Radiographs
Guanbin Li, Fuyu Wang 0001, Longjiang E, Yizhou Yu, Liang Lin 0004, Huiying Liang |
MICCAI (6) | 3 |