EDBT 2026 Demo / reviewers in the wild / expert
Yihan Fu
dblp:335/2717
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaRS: A Multi-modality Very-high-resolution Remote Sensing Foundation Model with Cross-Granularity Meta-Modality LearningabstractThe multi-modality remote sensing foundation model (MM-RSFM) has made notable progress recently. However, most existing approaches remain limited to medium-resolution, single-modality, restricting their performance in fine-grained downstream applications such as disaster response and urban planning. In this work, MaRS is proposed, a multi-modality very-high-resolution (VHR) remote sensing foundation model designed for cross-modality granularity interpretation of complex scenes. To achieve this, a multi-modality VHR SAR-optical dataset, MaRS-16M, is constructed through large-scale collection and semi-automated processing, comprising over 16 million paired samples. Unlike previous work, MaRS tackles two fundamental challenges in VHR SAR-optical self-supervised learning (SSL) techniques. Cross-granularity contrastive learning (CGCL) is introduced to alleviate alignment inconsistencies caused by imaging differences, and meta-modality attention (MMA) is designed to unify heterogeneous physical characteristics across modalities. Compared to existing remote sensing foundation models (RSFMs) and general vision foundation models (VFMs), MaRS performs better as a pre-trained backbone across nine multi-modality VHR downstream tasks. Ruoyu Yang, Yinhe Liu, Heng Yan, Yiheng Zhou, Yihan Fu, Yanfei Zhong |
AAAI | 5 |
| 2026 | M3DKV: Monolithic 3D Gain Cell Memory Enabled Efficient KV Cache & ProcessingabstractTransformer-based generative large language models (LLMs) have revolutionized natural language processing, yet their quadratic growth in computational complexity in context length creates severe inference bottlenecks. While LLM keyvalue cache (KV cache) enhances decoding efficiency, prolonged contexts infer frequent KV cache reloads that exacerbate memory bandwidth constraints. To address this hardware challenge, we propose M3DKV-a monolithic three-dimensional (3D) gain cell near-memory computing accelerator featuring back-end-of-line (BEOL) cache layers for in-situ KV matrix buffering and computation and a front-end-of-line (FEOL) base layer for full selfattention operations. Through optimized 3D data organization, inter-layer dataflow management, and intelligent computation scheduling, our design achieves $0.29 \mathrm{~TB} / \mathrm{s} /$ core on-die bandwidth while demonstrating $97.03 \times / 268.01 \times$ speedup over GPU/CPU in the decoding stage and $1.72 \times-262.16 \times$ better area efficiency per parameter versus state-of-the-art accelerators. Jiaqi Yang 0009, Yanbo Su, Yihan Fu, Jianshi Tang, Bonan Yan |
ASP-DAC | 3 |
| 2026 | Reconfigurable Computing Challenge: FPGA-Gym-v2: FPGA-Based RL Environment Acceleration with LLM-Assisted OnboardingabstractReinforcement learning (RL) repeatedly executes environment rollouts computation, where environment stepping may dominate training time when many environments run in parallel. Building on our prior FPGA-Gym/PEARL acceleration backbone, this work introduces FPGA-Gym-v2, a upgraded framework that keeps RL inference, training, and replay-buffer management on the host CPU/GPU, but offloads massively parallel environment stepping to an FPGA. It reduces host–FPGA traffic with compact encoding, on-chip state residency, and compute–communication overlap. Across representative Gymnasium workloads, FPGA-Gym-v2 achieves 4.36×–972.6× throughput speedup over EnvPool and up to 4699.3× over VectorEnv, while reducing DQN/PPO training time by 7%–15% over EnvPool and 18%–61% over VectorEnv. Further, FPGA-Gym-v2 adds an large-language-model-assisted (LLM-assisted) onboarding flow that generates environment-specific specifications, wrappers, Verilog HDL skeletons, and golden tests under explicit validation gates. The code is available at https://github.com/Selinaee/FPGA_Gym. Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Bonan Yan |
FCCM | 4 |
| 2026 | Reconfigurable Computing Challenge: FPGA-Based WebAssembly Stack Co-ProcessorabstractLarge language models suffer from hallucinations when performing scientific computing, motivating the use of AI agents such as IronClaw that offload computation to specialized tools. IronClaw invokes tools implemented as WebAssembly (Wasm) plugins for security and extensibility, but the stack-based Wasm bytecode is mismatched with register-based processors (x86, ARM), causing runtime overhead. We propose PAWS, a native Wasm coprocessor that directly executes Wasm bytecode in hardware. PAWS features: (1) full support for all five Wasm instruction types; (2) dual digital stack circuits (operand stack and control stack) replacing register files to minimize memory access latency; (3) dedicated control logic for block-based branching; and (4) a sliding-window instruction fetch unit that decodes variable-length Wasm instructions. Evaluated on the PolyBench suite, PAWS achieves average execution latencies 28.6× lower than an Intel Xeon processor and 40.6× lower than an Nvidia Jetson TX2, making it highly suitable for IronClaw’s compute-intensive scientific applications. The design is available at https://github.com/Iris-WQP/PAWS_FPGA_softcore. Qiuping Wu, Mugeng Liu 0001, Hongxiao Zhao, Yihan Fu, Gang Huang 0001, Yun Ma 0002, Bonan Yan |
FCCM | 5 |
| 2026 | Fusion-driven graph representation enhancement for predicting interactions of new drugsabstractAccurate prediction of drug–drug interactions (DDIs) for newly synthesized compounds enables early, in-silico safety screening in drug discovery and formulary review. We target the cold-start regime, where (i) new compounds are topologically isolated on external biomedical knowledge graphs (KGs) and on the DDI graph, and (ii) sparse supervision hampers the learning of discriminative representations. We propose an early-fusion method (LINCS-DDI) that inserts shared substructure nodes to connect a molecular-fingerprint knowledge graph with the DDI graph, turning structural similarity into topological links, and providing two-hop connectivity directly from the Simplified Molecular Input Line Entry System (SMILES) without prior inclusion in external KGs. Building on this substrate, we introduce Native Dual-View Contrastive Learning (NDV-CL): within a single pass of a flow-based graph neural network (GNN), forward and reverse message-passing representations of the same drug pair are treated as deterministic positives, while label-guided negatives (screened using only training-split interaction labels) are mined within the induced subgraph, improving representation quality without stochastic augmentations. Under strict cold-start settings on two open-source datasets, LINCS-DDI improves macro-F1 by up to 4.1% over the best baseline and reduces contrastive overhead by up to 66%. These properties make the approach suitable for routine, large-scale preclinical DDI triage, prioritizing high-risk combinations for wet-lab validation, and informing pharmacovigilance pipelines. Yihan Fu, Lin Wang 0023 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | PEARL: FPGA-Based Reinforcement Learning Acceleration with Pipelined Parallel EnvironmentsabstractReinforcement learning (RL) is an effective machine learning approach that enables artificial intelligence agents to perform complex tasks and make decisions in dynamic situations. Training an RL agent demands its repetitive interaction with the environment to learn optimal policies. To efficiently collect training data, parallelizing environments is a widely used technique by enabling simultaneous interactions between multiple agents and environments. However, existing CPU-based RL software frameworks face a key challenge of slow multi-environmental update computation. To solve this problem, we present a novel FPGA-based RL accelerating framework-PEARL. PEARL instantiates multiple parallel environments and accelerates them with a carefully designed pipeline scheme to hide data transfer latency within the computation time. We evaluate PEARL on respective RL environments and achieve 4.36 x to 972.6 x speedup over the existing fastest software-based framework for parallel environment execution. When scaling the number of environments from 1024 to 43008 (42x) in CliffWalking benchmark, the power consumption increases marginally by 3%, while LUT and flip-flops utilization rise by 2.24 x and 3.08 x, respectively. This demonstrates efficient resource usage and power management in PEARL. Further, PEARL allows users to define and add their environments within the framework flexibly. We have established an open-source repository for users to utilize and expand. We also implement PEARL with the existing RL algorithm and save 7% -15% training time. All the source code is available online https://github.com/Selinaee/FPGA_Gym. Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Yuchao Yang 0001, Bonan Yan |
DATE | 4 |
| 2025 | PROCA: Programmable Probabilistic Processing Unit Architecture with Accept/Reject Prediction & Multicore Pipelining for Causal InferenceabstractCausal inference is an important field in data science and cognitive artificial intelligence. It requires the construction of complex probabilistic models to describe the causal relationships between random variables. Probabilistic models rely on probabilistic programming as a flexible framework. However, the computing speed of probabilistic programming is often hindered by the extensive use of Markov chain Monte Carlo (MCMC) algorithms, even though they are powerful in Bayesian inference. To accelerate MCMC, this work presents PROCA, a programmable MCMC-based probabilistic processing unit architecture. PROCA exploits processing-in-memory function units to generate new samples of Markov chains. PROCA is programmable to execute the computation for arbitrary forms of posterior distribution formulas that software probabilistic programming frameworks support. We develop a novel accept/reject prediction methodology to accelerate the sequential MCMC computation, thereby introducing efficient multi-core pipelining methods. We implement and validate the PROCA architecture with commercial process development kits. The implementation is evaluated based on 9 representative benchmarks, covering PyMC official tutorial probabilistic problems, single-variable probabilistic problems, and real-world causal inference problems. Our comprehensive experiments demonstrate that PROCA achieves a speedup of 172~4871 $\times$ compared to Intel Xeon Gold CPU, $42 \sim 1058 \times$ compared to NVIDIA A100 GPU, and $1.765 \times$ over state-of-the-art MCMC accelerators, respectively. PROCA achieves comparable statistical robustness to the software probabilistic programming frameworks. Compared with state-of-the-art MCMC domain-specific accelerators, our design boosts the energy efficiency by $9.47 \times$. Yihan Fu, Anjunyi Fan, Wenshuo Yue, Hongxiao Zhao, Daijing Shi, Qiuping Wu, Yaoyu Tao, Yuchao Yang 0001, Bonan Yan |
HPCA | 1 |
| 2024 | Probabilistic Compute-in-Memory Design for Efficient Markov Chain Monte Carlo SamplingabstractMarkov chain Monte Carlo (MCMC) is a widely used sampling method in modern artificial intelligence and probabilistic computing systems. It involves repetitive random number generations and thus often dominates the latency of probabilistic model computing. Hence, we propose a compute-in-memory (CIM) based MCMC design as a hardware acceleration solution. This work investigates SRAM bitcell stochasticity and proposes a novel “pseudo-read” operation, based on which we offer a block-wise random number generation circuit scheme for fast random number generation. Moreover, this work proposes a novel multi-stage exclusive-OR gate (MSXOR) design method to generate strictly uniformly distributed random numbers. The probability error deviating from a uniform distribution is suppressed under$10^{-6}$. Also, this work presents a novel in-memory copy circuit scheme to realize data copy inside a CIM sub-array, significantly reducing the use of R/W circuits for power saving. Evaluated in a commercial 28-nm process development kit, this CIM-based MCMC design generates 4-bit$\sim$32-bit samples with an energy efficiency of 0.53 pJ/sample and high throughput of up to 1066.7M samples/s. Compared to conventional processors, the overall energy efficiency improves$2.12\times10^{9}$to$9.58\times10^{9}$times. Yihan Fu, Daijing Shi, Anjunyi Fan, Wenshuo Yue, Yuchao Yang 0001, Ru Huang 0001, Bonan Yan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |