EDBT 2026 Demo / reviewers in the wild / expert
Fei Wen 0003
dblp:21/8350-3
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8789-8495ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RIFT: A Single-Bitstream, Runtime-Adaptive FPGA-Based Accelerator for Multimodal AIabstractMultimodal models spanning ViTs, CNNs, GNNs, and NLP stress embedded systems because their heterogeneous compute and memory behaviors complicate resource allocation, load balancing, and real-time inference. We present RIFT, a single-bitstream FPGA accelerator and compiler for end-to-end multimodal inference. RIFT unifies layers as DDMM/SDDMM/SpMM kernels executed on a runtime mode-switchable engine that morphs among weight-/output-stationary systolic, 1×CSSIMD, and a routable adder tree (RADT) on a shared datapath. A two-stage hardware top-k unit, width-matched to the array, performs in-stream token pruning with minimal buffering, and dependency-aware scheduling overlaps independent kernels across multiple RPUs—achieving adaptation without bitstream reconfiguration. On Alveo U50 and ZCU104, RIFT reduces latency by up to 22.57× versus an RTX 4090 and 6.86× versus a Jetson Orin Nano at ∼20–21W; pruning alone yields up to 7.8× on ViT-heavy workloads. Hyunwoo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Suyeon Jang, Behnam Khaleghi, Fei Wen 0003, Mohsen Imani |
DATE | 7 |
| 2025 | A Multimodal AI Acceleration with Dynamic Pruning and Run-Time ConfigurationabstractThe computational diversity of multimodal AI workloads-spanning vision transformers (ViTs), graph neural networks (GNNs), CNNs, and transformer-based NLP—poses a fundamental challenge to embedded acceleration platforms. We propose a fully integrated FPGA-based acceleration framework that addresses this heterogeneity via compile-time and run-time configurability. Our system introduces a reconfigurable processing unit (RPU) capable of executing dense and sparse matrix operations (DDMM, SpMM, SDDMM), a scalable top-k pruning engine for ViTs, and a domain-specific compiler for hardware-software co-design. The architecture supports real-time configuration without reloading bitstreams, enabling unified deployment across tasks. Implementations on Xilinx U50 and ZCU104 demonstrate up to 22.57× and 6.86× latency reductions versus RTX 4090 and Jetson Orin Nano, respectively, validating the design's efficiency for real-time, resource-limited environments. Hyun Woo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Behnam Khaleghi, Fei Wen 0003, Mohsen Imani |
FCCM | 6 |
| 2025 | Revisiting Reconfigurable Acceleration of Vision Transformer with Patch PruningabstractVision Transformers (ViTs) have become the backbone of numerous cutting-edge vision applications. The attention modules within ViTs play a crucial role in modeling spatial relationships between pixels. Although these attention modules enhance the accuracy of ViT models, they also increase computational demands, limiting the deployment of ViTs in edge computing environments. To address this issue, prior research has focused on optimizing ViTs from both software and hardware perspectives. A notable software optimization technique is reducing the image patches involved in attention computations. Two common methods to achieve this are window attention and patch pruning. However, they introduce new challenges for existing hardware platforms regarding attention computation. Therefore, it is essential to develop new hardware modules to simultaneously support pruned attention computations and efficient window shifts. In this study, we introduce an FPGA-based token reduction vision transformer accelerator called TRFPA. Experiments conducted on the Xilinx ZCU104 and Alveo U50 demonstrate that TRFPA outperforms previous FPGA-based ViT accelerators, achieving a 7× speedup and a 3× improvement in energy efficiency. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Tamoghno Das, Fei Wen 0003, Mohsen Imani |
ISLPED | 6 |
| 2025 | VLTP: Vision-Language Guided Token Pruning for Task-Oriented SegmentationabstractVision Transformers (ViTs) have emerged as the backbone of many segmentation models, consistently achieving state-of-the-art (SOTA) performance. However, their success comes at a significant computational cost. Image token pruning is one of the most effective strategies to address this complexity. However, previous approaches fall short when applied to more complex task-oriented segmentation (TOS), where the class of each image patch is not predefined but dependent on the specific input task. This work introduces the Vision Language Guided Token Pruning (VLTP), a novel token pruning mechanism that can accelerate ViT-based segmentation models, particularly for TOS guided by multi-modal large language model (MLLM). We argue that ViT does not need to process every image token through all of its layers—only the tokens related to reasoning tasks are necessary. We design a new pruning decoder to take both image tokens and vision-language guidance as input to predict the relevance of each image token to the task. Only image tokens with high relevance are passed to deeper layers of the ViT. Experiments show that the VLTP framework reduces the computational costs of ViT by approximately 25% without performance degradation and by around 40% with only a 1% performance drop. The code associated with this study can be found at this URL. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Yezi Liu, Sungheon Jeong 0001, Fei Wen 0003, Nathaniel D. Bastian, Hugo Latapie, Mohsen Imani |
WACV | 6 |
| 2025 | Recoverable Anonymization for Pose Estimation: A Privacy-Enhancing ApproachabstractHuman pose estimation (HPE) is crucial for various applications. However, deploying HPE algorithms in surveillance contexts raises significant privacy concerns due to the potential leakage of sensitive personal information (SPI) such as facial features, and ethnicity. Existing privacy-enhancing methods often compromise either privacy or performance, or they require costly additional modalities. We propose a novel privacy-enhancing system that generates privacy-enhanced portraits while maintaining high HPE performance. Our key innovations include the reversible recovery of SPI for authorized personnel and the preservation of contextual information. By Jointly optimizing a privacy-enhancing module, a privacy recovery module, and a pose estimator, our system ensures robust privacy protection, efficient SPI recovery, and high-performance HPE. Experimental results demonstrate the system's robust performance in privacy enhancement, SPI recovery, and HPE. The code associated with this study can be found at this URL. Wenjun Huang 0001, Yang Ni 0001, Arghavan Rezvani, Sungheon Jeong 0001, Hanning Chen, Yezi Liu, Fei Wen 0003, Mohsen Imani |
WACV | 7 |
| 2024 | High-Performance Reconfigurable Accelerator for Knowledge Graph ReasoningabstractIn recent times, a plethora of hardware accelerators has emerged, catering to graph learning applications. However, the focus has primarily been on accelerating graph analysis, graph clustering, and graph mining, with a lack of attention to knowledge graph reasoning. Graph reasoning requires a more complex model to handle the complicated knowledge graph compared to other graph learning tasks. A primary knowledge graph reasoning task is to find the implicit relations between entities of a given knowledge graph, which demands a significantly longer training time than traditional graph learning algorithms due to the model complexity. Therefore, it is essential to develop an acceleration method to mitigate the training cost for the practical deployment of this task. Prior work in this field has solely considered using a single GPU or distributed GPU cluster to accelerate translational embedding models. However, as demonstrated in this paper, such general-purpose GPUs don't provide satisfactory results for more complex reinforcement learning-based models. Hence, it becomes necessary to design customized domain-specific accelerators. This work proposes GraFlex, the first domain specific accelerator for reinforcement learning-based knowledge graph reasoning, implemented on FPGA. We first develop a compression method for knowledge graphs. Then, we explore FPGAs of different sizes, analyze their on-chip resources, and suggest a mechanism to achieve high-speed training on devices with insufficient resources using the aforementioned compression method. Hanning Chen, Ali Zakeri, Yang Ni 0001, Fei Wen 0003, Behnam Khaleghi, Hugo Latapie, Mohsen Imani |
FCCM | 4 |
| 2023 | HyperGRAF: Hyperdimensional Graph-Based Reasoning Acceleration on FPGAabstractThe latest hardware accelerators proposed for graph applications primarily focus on graph neural networks (GNNs) and graph mining. High-level graph reasoning tasks, such as graph memorization and neighborhood reconstruction, have barely been addressed. Compared to low-level learning applications like node classification and clustering, high-level reasoning typically requires a more complex model to mimic human brain functionalities. Brain-inspired Hyper-Dimensional Computing (HDC) has recently introduced a promising lightweight and efficient machine learning solution, particularly for symbolic representation. General-purpose computing platforms (CPU/GPU) have been revealed to be inefficient for HDC applications. Therefore, it becomes essential to design a domain-specific accelerator targeting HDC-based graph reasoning algorithms. In this work, we propose the first domain-specific accelerator for HDC-based graph reasoning, HyperGRAF. We first develop a scheduler to balance the sparse matrix computation workloads, before parallelizing the hypervector calculations on two levels for the graph memorization task. Finally, we design a pipelinestyle matrix multiplication accelerator for the neighborhood reconstruction task. We evaluate our design under a wide range of generated graphs with different sizes and sparsity. The results show that HyperGRAF achieves over 100× improvement in both speedup and energy efficiency of graph reasoning compared to NVIDIA Jetson Orin. Hanning Chen, Ali Zakeri, Fei Wen 0003, Hamza Errahmouni Barkam, Mohsen Imani |
FPL | 3 |
| 2023 | KVRangeDB: Range Queries for a Hash-based Key-Value DeviceabstractKey–value (KV) software has proven useful to a wide variety of applications including analytics, time-series databases, and distributed file systems. To satisfy the requirements of diverse workloads, KV stores have been carefully tailored to best match the performance characteristics of underlying solid-state block devices. Emerging KV storage device is a promising technology for both simplifying the KV software stack and improving the performance of persistent storage-based applications. However, while providing fast, predictable put and get operations, existing KV storage devices do not natively support range queries that are critical to all three types of applications described above. In this article, we present KVRangeDB, a software layer that enables processing range queries for existing hash-based KV solid-state disks (KVSSDs). As an effort to adapt to the performance characteristics of emerging KVSSDs, KVRangeDB implements log-structured merge tree key index that reduces compaction I/O, merges keys when possible, and provides separate caches for indexes and values. We evaluated the KVRangeDB under a set of representative workloads, and compared its performance with two existing database solutions: a Rocksdb variant ported to work with the KVSSD, and Wisckey, a key–value database that is carefully tuned for conventional block devices. On filesystem aging workloads, KVRangeDB outperforms Wisckey by 23.7× in terms of throughput and reduce CPU usage and external write amplifications by 14.3× and 9.8×, respectively. Qing Zheng, Jason Lee 0004, Bradley W. Settlemyer, Fei Wen 0003, A. L. Narasimha Reddy, Paul Gratz |
ACM Trans. Storage | 5 |
| 2022 | SIMD-Matcher: A SIMD-based Arbitrary Matching FrameworkabstractPacket classification methods rely upon matching packet content/header against pre-defined rules, which are generated by network applications and their configurations. With the rapid development of network technology and the fast-growing network applications, users seek more enhanced, secure, and diverse network services. Hence it becomes critical to improve the performance of arbitrary matching operations. This article presents SIMD-Matcher, an efficient Single Instruction Multiple Data (SIMD) and cache-friendly arbitrary matching framework. To further improve the arbitrary matching performance, SIMD-Matcher adopts a trie node with a fixed high fanout and a varying span for each node depending on the data distribution. The trie node layout leverages cache and modern processor features such as SIMD instructions. To support arbitrary matching, we first interpret arbitrary rules into three fields: value, mask, and priority. Second, to support insertion of randomly positioned wildcards to arbitrary rules, we propose the SIMD-Matcher extraction algorithm to process the wildcard bits. Third, we add an array of wildcard entries to the leaf entries, which store the wildcard rules and guarantee the correctness of matching results. Experiments show that SIMD-Matcher outperforms GenMatcher under large-scale ruleset and key set, in terms of search time, insert time, and memory cost. Specifically with 5M rules, our method achieves a 2.7X speedup on search time, and the insertion time takes \( ~\sim \!\! 7.3 \) seconds, gaining a 1.38X speedup; meanwhile, the memory cost reduction is up to 6.17X. Ping Wang 0043, Fei Wen 0003, Paul Gratz, Alexander Sprintson |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | Software Hint-Driven Data Management for Hybrid Memory in Mobile SystemsabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. Recently emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of distinct memory classes render a new challenge for memory system design. Ideally, pages should be placed or migrated between the two types of memories according to the data objects’ access properties. Prior system software approaches exploit the program information from OS but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recent memory requests, due to the limited on-chip resources. In this work, we propose OpenMem: a hardware-software cooperative approach that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. First, we built a hardware-based memory manager unit (HMMU) that can learn the short-term access patterns by online profiling, and execute data migration efficiently. Then, we built a heap memory manager for the heterogeneous memory systems that allows the programmer to directly customize each data object’s allocation to a favorable memory device within the presumed object life cycle. With the programmer’s hints guiding the data placement at allocation time, data objects with similar properties will be congregated to reduce unnecessary page migrations. We implemented the whole system on the FPGA board with embedded ARM processors. In testing under a set of benchmark applications from SPEC 2017 and PARSEC, experimental results show that OpenMem reduces 44.6% energy consumption with only a 16% performance degradation compared to the all-DRAM memory system. The amount of writes to the NVM is reduced by 14% versus the HMMU-only, extending the NVM device lifetime. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | OpenMem: Hardware/Software Cooperative Management for Mobile Memory SystemabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. NVM technologies have higher capacity density, minimal static power consumption, but longer access latency and limited write endurance compared to DRAM. The different characteristics of these two memory classes, however, pose new challenges for memory system design. Ideally, pages shall be placed or migrated between the two types of memories according to the data objects’ access properties. Prior works use the OS for placement and migration in these systems, but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recently memory request, due to the limited on-chip resources.In this work, we propose OpenMem: a hardware-software cooperative approach to address placement and migration within hybrid memory systems, that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. We emulate OpenMem on an FPGA board with embedded ARM CPU, and run a set of benchmark applications from SPEC 2017 and PARSEC. Experimental results show that OpenMem reduces energy consumption by 44.6% with only a 16% performance degradation compared to an all-DRAM memory system. Further, writes to the NVM are reduced by 14% versus a hardware-only approach, extending the NVM device lifetime. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
DAC | 1 |
| 2021 | An FPGA-based Hybrid Memory Emulation SystemabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of applications. Emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of two memory classes point towards the design of hybrid memory systems containing multiple classes of main memory.In the iterative and incremental development of new architectures, the timeliness of simulation completion is critical to project progression. Hence, a highly efficient simulation method is needed to evaluate the performance of different hybrid memory system designs. Design exploration for hybrid memory systems is challenging, because it requires emulation of the full system stack, including the OS, memory controller, and interconnect. Moreover, benchmark applications for memory performance tests typically have much larger working sets, thus taking an even longer simulation warm-up period.In this paper, we propose an FPGA-based hybrid memory system emulation platform. We target the mobile computing system, which is sensitive to energy consumption and is likely to adopt NVM for its power efficiency. The focus of our platform is on the design of hybrid memory system, so we leverage the on-board hard IP ARM processors to enhance simulation performance while improving the accuracy of results. Thus, users can implement their data placement/migration policies with the FPGA logic elements and evaluate new designs quickly and effectively. Results show that our emulation platform provides a speedup of 9280x in simulation time compared to the software counterpart gem5. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
FPL | 1 |
| 2020 | Hardware Memory Management for Future Mobile Hybrid Memory SystemsabstractThe current mobile applications have rapidly growing memory footprints, posing a great challenge for memory system design. Insufficient DRAM main memory will incur frequent data swaps between memory and storage, a process that hurts performance, consumes energy, and deteriorates the write endurance of typical flash storage devices. Alternately, a larger DRAM has higher leakage power and drains the battery faster. Furthermore, DRAM scaling trends make further growth of DRAM in the mobile space prohibitive due to cost. Emerging nonvolatile memory (NVM) has the potential to alleviate these issues due to its higher capacity per cost than DRAM and minimal static power. Recently, a wide spectrum of NVM technologies, including phase-change memories (PCMs), memristor, and 3-D XPoint has emerged. Despite the mentioned advantages, NVM has longer access latency compared to DRAM and NVM writes can incur higher latencies and wear costs. Therefore, the integration of these new memory technologies in the memory hierarchy requires a fundamental rearchitecting of traditional system designs. In this work, we propose a hardware-accelerated memory manager (HMMU) that addresses in a flat address space, with a small partition of the DRAM reserved for subpage block-level management. We design a set of data placement and data migration policies within this memory manager such that we may exploit the advantages of each memory technology. By augmenting the system with this HMMU, we reduce the overall memory latency while also reducing writes to the NVM. The experimental results show that our design achieves a 39% reduction in energy consumption with only a 12% performance degradation versus an all-DRAM baseline that is likely untenable in the future. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |