Zimeng Fan 0001

dblp:260/3655-1 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-8281-0458ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AGF: Adaptive GNN Framework via Subgraph and Model Partitioning for Edge Analysis
abstract
AI has found extensive application in edge analytics, with Deep Neural Networks effectively processing and analyzing data. Some applications need to deal with unstructured data like graphs. This complexity and diversity of unstructured features pose challenges for traditional DNN models. Therefore, Graph Neural Networks have emerged as a significant solution. However, GNNs encounter two primary challenges in edge analytics scenarios: scalability and efficiency. These challenges arise primarily from significant structural differences among various graphs and models. To address these issues, we propose AGF, an Adaptive GNN Framework. AGF consists of three components: subgraph partitioning and model partitioning, scalable hardware architecture, and design space exploration. Subgraph partitioning and model partitioning transform complex and diverse GNNs into a uniform computing flow and subgraphs. These can be directly mapped onto the hardware kernel, leveraging GNN characteristics to achieve independent parallelism. The hardware architecture is based on a multi-level processing element structure, which can efficiently parallelize subgraphs and vertices based on the partitioning. The design space exploration selects kernel allocation strategies and setups based on the hardware, dataset, and model. We validated AGF using six datasets and three GNN models. The results demonstrate that our framework achieves an acceleration ranging from 4.54 to 53.17× compared to the GPU-based GNN framework. When compared to other FPGA-based GNN accelerators, AGF achieves a latency reduction ranging from 1.16 to 33.6× in model execution and from 1.09 to 2.37× in end-to-end scenarios.
Zimeng Fan 0001, Min Peng 0002
ACM Trans. Internet Techn.1
2025 GNNmap: A Scalable Framework for GNN Deployment through Co-Optimized Graph Partitioning and Mapping
abstract
Graph Neural Networks (GNNs) have become pivotal for analyzing relational data in embedded intelligent systems such as IOT devices. However, their deployment on resource-constrained devices faces critical barriers: traditional graph partitioning methods induce unbalanced computational loads due to rigid granularity, while hardware mapping strategies cause inefficient resource utilization under dynamic graph structures. These limitations conflict with the requirements of embedded systems for resource efficiency and scalability. To address this, we present GNNmap, a hardware-software co-design framework that synergizes multi-granular graph partitioning with topology-aware GNN mapping. The framework first reconstructs input graphs into balanced kernel groups comprising cohesive supernodes (corresponding to parallelizable subgraphs). By combining coarse-grained partitioning with fine-grained optimization, GNNmap ensures load balance while dramatically reducing cross-subgraph communication. Concurrently, a subgraph-PE mapping based on coarse-grained reconfigurable architectures (CGRAs) enables efficient graph-to-hardware matching through the joint modeling of graph topological features and hardware resource constraints. By dynamically coordinating graph reorganization and hardware resource allocation, GNNmap resolves the intrinsic mismatch between irregular graph computations and static hardware configurations. Experimental results demonstrate that GNNmap achieves improvements over existing works, improving inference performance by 1.47× to 62.8×, resource efficiency by 1.15× to 3.06×, and energy efficiency by 1.34× to 3.50×.
Zimeng Fan 0001, Min Peng 0002
ACM Trans. Embed. Comput. Syst.1
2025 DGMF: A Unified Dynamic Mapping Framework for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have seen considerable advancements across various applications. However, the computationally and storage-intensive nature of GNNs presents unique challenges for hardware design. Existing GNN accelerators frequently grapple with low execution efficiency and workload imbalance. Therefore, this article introduces DGMF, a unified dynamic mapping approach that optimizes dataflow and computation mode based on GNN models and datasets. To support this dynamic mapping approach, we also present a dedicated hardware architecture. This architecture dynamically adjusts resource allocation and parallelism based on models and datasets, thereby ensuring optimal utilization. Additionally, we propose a Design Space Exploration (DSE) algorithm that traverses the defined design space to identify the optimal design solution, effectively managing diverse design constraints and objectives. To validate our approach, we developed an accelerator based on three datasets and five types of GNN models. Experimental results demonstrate that DGMF achieves performance by 10.5×–2,548×, and 1.01×–51.7× compared to GPUs and existing accelerators. It also improves resource efficiency, energy efficiency, and DRAM access by 0.87×–1.47×, 0.92×–2.18×, and 1.01×–47.25×, respectively, compared to other works.
Zimeng Fan 0001, Min Peng 0002
ACM Trans. Reconfigurable Technol. Syst.1
2024 A Hardware Design Framework for Computer Vision Models Based on Reconfigurable Devices
abstract
In computer vision, the joint development of the algorithm and computing dimensions cannot be separated. Models and algorithms are constantly evolving, while hardware designs must adapt to new or updated algorithms. Reconfigurable devices are recognized as important platforms for computer vision applications because of their reconfigurability. There are two typical design approaches: customized and overlay design. However, existing work is unable to achieve both efficient performance and scalability to adapt to a wide range of models. To address both considerations, we propose a design framework based on reconfigurable devices to provide unified support for computer vision models. It provides software-programmable modules while leaving unit design space for problem-specific algorithms. Based on the proposed framework, we design a model mapping method and a hardware architecture with two processor arrays to enable dynamic and static reconfiguration, thereby relieving redesign pressure. In addition, resource consumption and efficiency can be balanced by adjusting the hyperparameter. In experiments on CNN, vision Transformer, and vision MLP models, our work’s throughput is improved by 18.8x–33.6x and 1.4x–2.0x compared to CPU and GPU. Compared to others on the same platform, accelerators based on our framework can better balance resource consumption and efficiency.
Zimeng Fan 0001, Wei Hu 0001, Fang Liu 0031, Dian Xu, Hong Guo 0005, Yanxiang He, Min Peng 0002
ACM Trans. Reconfigurable Technol. Syst.1
2022 Software and Hardware Fusion Multi-Head Attention
Wei Hu 0001, Dian Xu, Fang Liu 0031, Zimeng Fan 0001
KSEM (3)4
2022 Classification of Heads in Multi-head Attention Mechanisms
Feihu Huang 0003, Min Jiang 0015, Fang Liu 0031, Dian Xu, Zimeng Fan 0001, Yonghao Wang
KSEM (3)5
2021 Hardware and Algorithm Co-Optimization for pointwise convolution and channel shuffle in ShuffleNet V2
abstract
Since the convolutional neural network (CNN) was proposed, it has achieved remarkable results in image recognition, and its accuracy rate surpasses other methods. In recent years, with the continuous improvement of computational complexity and the development of research on the deployment of neural networks on terminal devices, new neural network models suitable for terminal devices such as MobileNet, SqueezeNet and ShuffleNet have also appeared in CNN model. These network models have special structures or units to reduce the amount of calculation.the lightweight features of the model cannot take advantage of due to the lack of customized design for them when GPU is used for acceleration. The characteristics of FPGA such as customization and parallelism just meet the needs of research. However, there is a little research on implementing and optimizing new lightweight networks on FPGAs. In this paper, we optimize each part of the model based on FPGA, achieve 8bit quantification, and redesign depthwise separable convolution operation and channel shuffle module, which make the module can be operated in a hardware-friendly way. Under the premise of maintaining the original model structure and performance, we make full use of the abundant LUT, FF and other resources on the FPGA. we use HLS (high-level synthesis) to analyze the structure of the ShuffleNet V2 model on xilinxzynqxc-7Z045. Compared with the performance of the original model in CPU, GPU and FPGA, our design achieves the acceleration performance of 8.7x, 1.3x and 3.3x respectively, and the accuracy is only reduced by 0.8%. Compared with other designs, it also achieved a result of 3.09x 60x in terms of latency.
Zimeng Fan 0001, Wei Hu 0001, Hong Guo 0005, Fang Liu 0031, Dian Xu
SMC1