VLDB 2026 Research / reviewers in the wild / expert
Yiheng Wu
dblp:224/4714
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OpenACM: An Open-Source SRAM-Based Approximate CiM CompilerabstractThe rise of data-intensive AI workloads has exacerbated the "memory wall" bottleneck. Digital Compute-in-Memory (DCiM) using SRAM offers a scalable solution, but its vast design space makes manual design impractical, creating a need for automated compilers. A key opportunity lies in approximate computing, which leverages the error tolerance of AI applications for significant energy savings. However, existing DCiM compilers focus on exact arithmetic, failing to exploit this optimization. This paper introduces OpenACM, the first open-source, accuracy-aware compiler for SRAM-based approximate DCiM architectures. OpenACM bridges the gap between application error tolerance and hardware automation. Its key contribution is an integrated library of accuracy-configurable multipliers (exact, tunable approximate, and logarithmic), enabling designers to make fine-grained accuracy-energy trade-offs. The compiler automates the generation of the DCiM architecture, integrating a transistor-level customizable SRAM macro with variation-aware characterization into a complete, open-source physical design flow based on OpenROAD and the FreePDK45 library. This ensures full reproducibility and accessibility, removing dependencies on proprietary tools. Experimental results on representative convolutional neural networks (CNNs) demonstrate that OpenACM achieves energy savings of up to 64% with negligible loss in application accuracy. The framework is available on OpenACM:URL. JunHao Ma, Xingyang Li, Yule Sheng, Bochang Wang, Yiheng Wu, Shan Shen, Daying Sun |
DATE | 8 |
| 2026 | EpiGator: An Event-based Surveillance System for Infectious Disease OutbreaksabstractPeer reviewed Yiheng Wu, Sathianpong Trangcasanchai, Lidia Pivovarova, Roman Yangarber |
LREC | 1 |
| 2026 | Thinking Bidirectionally: A Reasoning and Self-Correction Approach for Text-Based Event Prediction with Large Language ModelsabstractUsing Web Mining and Content Analysis to find and understand clues from the massive amount of unstructured text online is very important for predicting future events and providing early risk warnings in important fields like finance and public safety. While Large Language Models (LLMs) exhibit potential in processing and understanding text, current text-based event prediction faces two primary challenges: first, an insufficient utilization of potential information within the text, such as causal relationships and latent associations, and second, limited predictive reliability constrained by issues like the LLM's own ability and hallucinations. To address these challenges, we propose a novel event prediction framework, Bidirectional Reasoning with Self-Correction (BRSC). BRSC comprises two complementary reasoning dimensions: temporal deductive reasoning, which analyzes the trajectory of historical events along the timeline to enable accurate trend extrapolation, and synchronic associative reasoning, which deeply mines details and latent connections from documents within a specific time window to extended semantic information. In addition, we use a self-correction mechanism that identifies and rectifies potential hallucinations and errors during the reasoning process. Extensive experiments on international relations event prediction demonstrate that BRSC achieves significant improvements over several leading LLM-based methods. Liwei Qian, Hang Zhang 0008, Yiheng Wu, Yanmin Li, Mengna Zhu, Lihua Liu 0002, Jibing Wu |
WWW | 3 |
| 2026 | An Extended Full GKS Formulation for High-Efficiency and Low-Memory Two-Phase Flow SimulationabstractTwo-phase flows are ubiquitous in nature, exhibiting complex fluid-fluid interactions that challenges numerical simulators. To accurately and efficiently solve two-phase flows, grid-based methods have been widely adopted. Navier-Stokes (NS) solvers consume a small memory footprint, but simultaneously achieving both high performance and low numerical dissipation remains a significant challenge. In contrast, lattice Boltzmann solvers are efficient and have low numerical dissipation, yet they remain memoryintensive, even with state-of-the-art moment-encoding schemes. To date, the simultaneous attainment of high accuracy, exceptional efficiency, and a low memory footprint remains a major challenge in the field. In this paper, we propose a novel two-phase flow solver that achieves this objective. Our work is motivated by extending gas-kinetic scheme (GKS), which is adapted to handle nearly incompressible flows. To allow stable and accurate two-phase flow simulations, we systematically derive a coupled formulation of the GKS method and the phase-field model, incorporating novel mathematical constructs. Combined with robust boundary treatments and specialized techniques for handling turbulent flows, this results a unified framework capable of efficiently simulating two-phase flows, even those with large density contrasts and high Reynolds numbers. Since our formulation is explicit, it achieves exceptional performance when optimized on GPU, making it the fastest kinetic two-phase flow solver to date. Additionally, as it is derived from GKS, it obviates the need to store distribution functions. Thus, it has a small memory footprint, competitive with, or even lower than, that of many NS solvers. As a result, our solver can efficiently simulate complex two-phase flow dynamics at high resolutions using a single commodity GPU. We validate the accuracy of our solver via several benchmark tests, compare its performance with recent methods in various aspects, and demonstrate its capability to replicate a broad range of two-phase flow phenomena, encompassing both typical and large-scale scenarios. Yiheng Wu, Kai Bai, Xiaopei Liu |
ACM Trans. Graph. | 1 |
| 2025 | Can Large Language Models Tackle Graph Partitioning?abstractLarge language models (LLMs) demonstrate remarkable capabilities in understanding complex tasks and have achieved commendable performance in graph-related tasks, such as node classification, link prediction, and subgraph classification.These tasks primarily depend on the local reasoning capabilities of the graph structure.However, research has yet to address the graph partitioning task that requires global perception abilities.Our preliminary findings reveal that vanilla LLMs can only handle graph partitioning on extremely small-scale graphs.To overcome this limitation, we propose a three-phase pipeline to empower LLMs for large-scale graph partitioning: coarsening, reasoning, and refining.The coarsening phase reduces graph complexity.The reasoning phase captures both global and local patterns to generate a coarse partition.The refining phase ensures topological consistency by projecting the coarse-grained partitioning results back to the original graph structure.Extensive experiments demonstrate that our framework enables LLMs to perform graph partitioning across varying graph scales, validating both the effectiveness of LLMs for partitioning tasks and the practical utility of our proposed methodology. Yiheng Wu, Ningchao Ge, Yanmin Li, Liwei Qian, Mengna Zhu, Haiwen Chen, Jibing Wu |
EMNLP | 1 |
| 2025 | OpenYield: An Open-Source SRAM Yield Analysis and Optimization Benchmark SuiteabstractStatic Random-Access Memory (SRAM) yield analysis is essential for semiconductor innovation, yet research progress faces a critical challenge: the large gap between simplified academic models and the complexities observed in practice. The lack of open, higher-fidelity benchmarks has hindered reproducibility and transferability, as promising academic techniques often fail to carry over to more realistic settings. We present OpenYield, an open-source ecosystem that aims to narrow this gap through three contributions: (i) An SRAM circuit generator that explicitly incorporates second-order effects (interconnect/line parasitics, inter-cell leakage coupling, and peripheralcircuit variations) that are commonly omitted in academic studies. (ii) A standardized evaluation platform with a simple interface and baseline yield-analysis implementations to enable fair comparisons and reproducible research on these higherfidelity circuits. (iii) An optimization platform for transistor-level sizing under these models, supporting reproducible studies of robustness/efficiency trade-offs. OpenYield aims to foster more reproducible and transferable progress in SRAM-yield research. The framework is publicly available at OpenYield:URL. Shan Shen, Xingyang Li, Zhuohua Liu, Junhao Ma, Yiheng Wu, Yuquan Sun, Wei W. Xing |
ICCD | 6 |
| 2025 | Simulating Two-Phase Fluid-Rigid Interactions With an Overset-Grid Kinetic SolverabstractSimulating the coupled dynamics between rigid bodies and two-phase fluids, especially those with a large density ratio and a high Reynolds number, is computationally demanding but visually compelling with a broad range of applications. Traditional approaches that directly solve the Navier-Stokes equations often struggle to reproduce these flow phenomena due to stronger numerical diffusion, resulting in lower accuracy. While recent advancements in kinetic lattice Boltzmann methods for two-phase flows have notably enhanced efficiency and accuracy, challenges remain in correctly managing fluid-rigid boundaries, resulting in physically inconsistent results. In this article, we propose a novel kinetic framework for fluid-rigid interaction involving two fluid phases. Our approach leverages the idea of an overset grid, and proposes a novel formulation in the two-phase flow context with multiple improvements to handle complex scenarios and support moving multi-resolution domains with boundary layer control. These new contributions successfully resolve many issues inherent in previous methods and enable physically more consistent simulations of two-phase flow phenomena. We have conducted both quantitative and qualitative evaluations, compared our method to previous techniques, and validated its physical consistency through real-world experiments. Additionally, we demonstrate the versatility of our method across various scenarios. Xiaoyu Xiao, Ding Lin, Yiheng Wu, Kai Bai, Xiaopei Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | YOLO-DCNet: A Semantic-Based Novel Flexible Lightweight Human Detection AlgorithmabstractEnhanced processors empower edge devices like smartphones for human detection, yet their application is constrained by algorithmic efficiency and precision. This paper introduces YOLO-DCNet, a lightweight neural network detector built upon YOLOv7-tiny. Incorporating a dynamic multi-head structural re-parameterization (DMSR) module within its backbone network enables effective processing of the features utilized in the model. To improve multi-scale feature aggregation, the model integrates a channel information compression and linear mapping (CLM) module into its feature pyramid architecture. Moreover, the optimization of training and inference performance is achieved by employing RepVGG blocks between the main computational modules of the model. Experimental data reveal that the enhanced YOLOv7-tiny model achieves a 31.7% faster inference speed and marginal gains of 0.7% in [email protected] and 0.5% in [email protected]:0.95 over the original. This underscores the model's improved performance and applicability for real-time human detection on edge devices across diverse applications. Yiheng Wu, Jiaqiang Dong |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2023 | A Lightweight Real-Time System for Object Detection in Enterprise Information Systems for Frequency-Based Feature SeparationabstractIn the domain of target detection in mobile and embedded devices, neural network model inference speed is a crucial metric. This paper introduces YOLO-FLNet, a lightweight algorithm for detecting people in open scenes. The model utilizes the DFEM structure to capture and process high-frequency and low-frequency information in the feature map. Additionally, the VoV-DFEM structure, based on the concept of one-shot aggregation, enhances feature aggregation from different scales and frequencies in the backbone network. To validate its performance, experiments were conducted using publicly available datasets on a computer with dedicated GPUs. As a result, compared to YOLOv7-tiny, YOLO-FLNet achieved a 0.3% [email protected] improvement, reduced parameter size by 52.9%, and increased inference speed by 30.2%. These characteristics make it valuable for person detection in engineering domains, providing theoretical guidance for lightweight models in edge computing. Yiheng Wu |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2023 | Building a Virtual Weakly-Compressible Wind Tunnel Testing FacilityabstractVirtual wind tunnel testing is a key ingredient in the engineering design process for the automotive and aeronautical industries as well as for urban planning: through visualization and analysis of the simulation data, it helps optimize lift and drag coefficients, increase peak speed, detect high pressure zones, and reduce wind noise at low cost prior to manufacturing. In this paper, we develop an efficient and accurate virtual wind tunnel system based on recent contributions from both computer graphics and computational fluid dynamics in high-performance kinetic solvers. Running on one or multiple GPUs, our massively-parallel lattice Boltzmann model meets industry standards for accuracy and consistency while exceeding current mainstream industrial solutions in terms of efficiency --- especially for unsteady turbulent flow simulation at very high Reynolds number (on the order of 10 7 ) --- due to key contributions in improved collision modeling and boundary treatment, automatic construction of multiresolution grids for complex models, as well as performance optimization. We demonstrate the efficacy and reliability of our virtual wind tunnel testing facility through comparisons of our results to multiple benchmark tests, showing an increase in both accuracy and efficiency compared to state-of-the-art industrial solutions. We also illustrate the fine turbulence structures that our system can capture, indicating the relevance of our solver for both VFX and industrial product design. Chaoyang Lyu, Kai Bai, Yiheng Wu, Mathieu Desbrun, Changxi Zheng, Xiaopei Liu |
ACM Trans. Graph. | 3 |
| 2020 | A 1036-F2/Bit High Reliability Temperature Compensated Cross-Coupled Comparator-Based PUFabstractIn this article, a compact physical unclonable function (PUF) based on cross-coupled comparator is presented. Featuring a positive feedback response generation mechanism, the mismatch in analog signals between the cross-coupled transistor pair is quickly amplified to prevent its polarity from flipping by the temporal noise. The rapid enlargement of noise margin by the sense amplifier also contributes to stabilizing the response against supply voltage variations. To improve its temperature stability, the counteracting effect of complementary-to-absolutetemperature (CTAT) and proportional-to-absolute-temperature (PTAT) drives are considered in sizing the bit cell transistors. The proposed design is fabricated in a standard 65-nm CMOS process. The bit cell occupies an area of only 4.38 μm2(i.e., 1036 F2), and the overall PUF chip consumes 2.98 pJ/bit at the throughput of 8 Mb/s, of which only 1.61 pJ/bit is due to the PUF's core. With the uniqueness measured to be 49.53%, the unpredictability of the fabricated PUF chips is validated by autocorrelation function and NIST randomness tests. Compared with the state-of-the-art implementations, the proposed PUF has the lowest native response instability of 1.46% with 500 repeated PUF readouts at 27 °C and 1.2 V. By varying the operating temperature from -50 °C to 150 °C in a step size of 10 °C and the supply voltage from 1.0 to 1.4 V in a step size of 0.1 V simultaneously, the average reliability of the proposed PUF obtained from the 2-D plot of all operating conditions is found to be 96.87% without correction and 99.31% with spatial majority voting (SMV). Yiheng Wu, Xiaojin Zhao, Yuan Cao 0003, Chip-Hong Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |