VLDB 2026 Research / reviewers in the wild / expert
Yimin Wang 0001
dblp:07/4113-1
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0008-7292-0670ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRIMAL: Processing-In-Memory based Low-Rank Adaptation for LLM Inference AcceleratorabstractThis paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM processing elements (PEs), interconnected by 2D-mesh inter-PE computational network (IPCN). A novel SRAM reprogramming and power gating (SRPG) scheme enables pipelined LoRA updates and sub-linear power scaling by overlapping reconfiguration with computation and gating idle resources. PRIMAL employs optimized spatial mapping and dataflow orchestration to minimize communication overhead, and achieves $1.5\times$ throughput and $25\times$ energy efficiency over NVIDIA H100 with LoRA rank 8 (Q,V) on Llama-13B. Yue Jiet Chong, Yimin Wang 0001, Xuanyao Fong |
ISCAS | 2 |
| 2026 | Topology-Mapping Co-Design for Scalable LLM Accelerators with Hierarchical Mesh-of-Trees NoC
Yimin Wang 0001, Yue Jiet Chong, Xuanyao Fong |
ISCAS | 2 |
| 2026 | Cross-Layer Evaluation for On-Chip Training of BEOL FeTFT-based CIM M3D Accelerator
Jianze Wang, Yimin Wang 0001, Xuanyao Fong |
ISCAS | 3 |
| 2025 | LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained ParallelismabstractLarge language model (LLM) inference has been a prevalent demand in daily life and industries. The large tensor sizes and computing complexities in LLMs have brought challenges to memory, computing, and databus. This paper proposes a computation/memory/communication co-designed non-von Neumann accelerator by aggregating processing-in-memory (PIM) and computational network-on-chip (NoC), termed LEAP. The matrix multiplications in LLMs are assigned to PIM or NoC based on the data dynamicity to maximize data locality. Model partition and mapping are optimized by heuristic design space exploration. Dedicated fine-grained parallelism and tiling techniques enable high-throughput dataflow across the distributed resources in PIM and NoC. The architecture is evaluated on Llama 1B/8B/13B models and shows ~2.55× throughput (tokens/sec) improvement and ~71.94× energy efficiency (tokens/Joule) boost compared to the A100 GPU. Yimin Wang 0001, Yue Jiet Chong, Xuanyao Fong |
ICCAD | 1 |
| 2024 | Design Framework for Ising Machines with Bistable Latch-Based Spins and All-to-All Resistive CouplingabstractMemristive crossbars have been proposed as a promising pathway to enabling all-to-all connections for a variety of applications. One such application is the latch-based dynamical Ising machine, which have been proposed for solving combinatorial optimization problems. However, the impact of the design of the memristive crossbar on the performance of the latch-based Ising machine remains unclear. We present a SPICE-level model and simulation framework for analyzing and evaluating the latch-based dynamical Ising machine that use memristive crossbars to achieve all-to-all coupling between Ising spins implemented as latches. Our design space exploration reveals that the solution quality is highly sensitive to the circuit and device design parameters and sizes of the problems. An effective metric that can capture the effect of several design parameters on the statistical results is then proposed to quantify the system functionality. Finally, optimization techniques based on our proposed metric that models the effect of several design parameters on the solution quality of the Ising machine are proposed and demonstrated on MaxCut solving. Our evaluation results on a 128×128 crossbar array-based design show >98% solution quality across a wide range of problem sizes is achievable. Yimin Wang 0001, Yunuo Cen, Xuanyao Fong |
ISCAS | 1 |
| 2024 | Energy-Efficient Ising Machines Using Capacitance-Coupled Latches for MaxCut SolvingabstractLatch-based Ising machines (LIMs) have aroused research interest due to their speed and area efficiency for solving combinatorial optimization problems (COPs). However, existing LIMs based on resistive coupling are sensitive to many design parameters and suffer from solution quality degradation. This paper explores a capacitive-coupling approach that improves the robustness of the LIM to variations in the system design parameters. Evaluation results show that the LIMs using capacitive coupling can solve the MaxCut problem with high solution quality in a ∼1000× wider range of coupling strength compared to the resistive coupling approach. It also achieves an energy-time product of 313.5 nJ•s, with an 84.3% and 86.4% reduction compared to the state-of-the-art resistance-coupled and capacitance-coupled counterparts respectively. Yimin Wang 0001, Xuanyao Fong |
ISCAS | 1 |
| 2024 | Transposable Memory Based on the Ferroelectric Field-Effect TransistorabstractNon-volatile ferroelectric field-effect transistor (FeFET) technology is a promising CMOS process compatible solution for fast, energy efficient on-chip memories that are needed to enable ubiquitous deployment of AI. Existing memories are linear by design and the bandwidth may not be sufficient to support transformers that underlie the now-popular Large Language Models (LLMs) for generative AI, which require fast accesses to matrices and their transposes. In this work, two designs of transposable FeFET memories are proposed to address this challenge. We evaluate our designs using compact models for the FeFET devices, which we have calibrated to experimentally measured device characterization data. We demonstrate that 1 ns read latencies are possible, which underscores the potential of FeFET technology for AI accelerators. The area efficiency of our proposed dual-gated transposable memory and the speed of transpose operations are studied, showing that it occupies 30.02 F2cell area and is capable of maximum 82.24× faster transpose operation compared to non-transposable memory. Jianze Wang, Yimin Wang 0001, Leming Jiao, Xiaolin Wang 0007, Xuanyao Fong |
ISCAS | 4 |
| 2021 | Design Framework for SRAM-Based Computing-In-Memory Edge CNN AcceleratorsabstractThis paper presents an architectural framework and an evaluation model for Static Random Access Memory (SRAM)-based Computing-in-Memory (CIM) edge Convolutional Neural Network (CNN) accelerators. To provide a baseline for system-level design perspectives, an architectural framework for SRAM-CIM design concerning the key design points in state-of-the-art works is proposed. Furthermore, a configurable evaluation model featuring top-down design flow based on the proposed framework is established to investigate design space explorations. Case studies validated the framework and evaluation model using LeNet-5, AlexNet and VGG-16 to achieve energy-aware optimizations. The optimized memory scale for LeNet-5 is "16 PEs and 120 tiles" with the minimal estimated inference energy of 0.0018J, while for AlexNet and VGG-16, "16 PEs and 120 tiles" is better achieving minimal energy consumption of 0.1733mJ and 0.6825mJ respectively. Estimation results highlight tradeoffs among data represent- tation parameters and memory partitioning parameters. This work provides specific SRAM-CIM design guidelines from a system-level perspective. Yimin Wang 0001, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 1 |