EDBT 2026 Demo / reviewers in the wild / expert
Weihang Li
dblp:296/1000
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PF-LLM: Large Language Model Hinted Hardware PrefetchingabstractHardware data prefetching is a critical technique for mitigating memory latency in modern processors. While sophisticated hardware prefetching algorithms exist, their exclusive reliance on runtime information limits their ability to adapt quickly and comprehend broader program context. Our key insight is that the optimal prefetching strategy for a load instruction is often discernible from its static code context -- a task at which experienced developers excel. This motivates our central question: can a Large Language Model (LLM) be trained to perform this analysis automatically? We introduce PF-LLM, an LLM fine-tuned to analyze the assembly context surrounding a load instruction and generate prefetching hints. These offline-generated hints are consumed at runtime by LMHint Prefetcher, a lightweight hardware prefetcher ensemble designed to leverage this static guidance. Our approach boosts the performance of the on-chip hardware prefetcher by moving the hard ''when, how, and how aggressively to prefetch'' decisions out of the runtime hardware and into an offline LLM-powered analysis. This turns the on-chip prefetcher into a zero-latency, oracle-level system that always follows the best prefetching policy for every single load instruction. Our evaluation shows that our approach achieves a 9.8% instruction-per-cycle (IPC) improvement on average for memory-intensive SPEC 2017 benchmarks over state-of-the-art hardware prefetching baselines and 18.9% improvement on average over state-of-the-art ensemble methods, demonstrating the significant potential of leveraging LLMs to guide microarchitectural decisions. Ceyu Xu, Xiangfeng Sun, Weihang Li, Bangyan Wang, Mengming Li, Zhiyao Xie, Yuan Xie 0001 |
ASPLOS (2) | 3 |
| 2026 | DynSUP: Dynamic Gaussian Splatting From an Unposed Image PairabstractRecent advances in 3D Gaussian Splatting have shown promising results. Existing methods typically assume static scenes and/or multiple images with prior poses. Dynamics, sparse views, and unknown poses significantly increase the problem complexity due to insufficient geometric constraints. To overcome this challenge, we propose a method that can use only two images without prior poses to fit Gaussians in dynamic environments. To achieve this, we introduce two technical contributions. First, we propose an object-level two-view bundle adjustment. This strategy decomposes dynamic scenes into piece-wise rigid components, and jointly estimates the relative camera motion and dynamic object motions for dynamic Gaussian initialization. Second, we design an SE(3) field-driven Gaussian training method. It enables fine-grained motion modeling through learnable per-Gaussian transformations. Our method leads to high-fidelity novel view synthesis of dynamic scenes while accurately preserving temporal consistency and object motion. Experiments on both synthetic and real-world datasets demonstrate that our method significantly outperforms state-of-the-art approaches designed for the cases of static environments, multiple images, and/or known poses. Our project page is available at https://colin-de.github.io/DynSUP/. Weihang Li, Shenhan Qian, Benjamin Busam, Daniel Cremers, Haoang Li |
IEEE Trans. Image Process. | 1 |
| 2025 | GCE-Pose: Global Context Enhancement for Category-level Object Pose EstimationabstractA key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Recent approaches leverage foundational features to capture semantic and geometry cues from data. However, these approaches fail under partial visibility. We overcome this with a first-complete-then-aggregate strategy for feature extraction utilizing class priors. In this paper, we present GCE-Pose, a method that enhances pose estimation for novel instances by integrating category-level global context prior. GCE-Pose performs semantic shape reconstruction with a proposed Semantic Shape Reconstruction (SSR) module. Given an unseen partial RGB-D object instance, our SSR module reconstructs the instance’s global geometry and semantics by deforming category-specific 3D semantic prototypes through a learned deep Linear Shape Model. We further introduce a Global Context Enhanced (GCE) feature fusion module that effectively fuses features from partial RGB-D observations and the reconstructed global context. Extensive experiments validate the impact of our global context prior and the effectiveness of the GCE fusion module, demonstrating that GCE-Pose significantly outperforms existing methods on challenging real-world datasets House-Cat6D and NOCS-REAL275. Our project page is available at https://colin-de.github.io/GCE-Pose/. Weihang Li, Junwen Huang 0001, Peter KT Yu, Nassir Navab, Benjamin Busam |
CVPR | 1 |
| 2024 | Knowledge-based Programming by Demonstration using semantic action models for industrial assemblyabstractIn this paper, we introduce a knowledge-based Programming by Demonstration (kb-PbD) paradigm to facilitate robot programming in small and medium-sized enterprises (SMEs). PbD in production scenarios requires the recognition of product-specific actions but faces challenges in the lack of suitable and comprehensive datasets, due to the large variety of involved hand actions across different production scenarios. To address this issue, we utilize standardized grasp types as the fundamental feature to recognize basic hand movements, where a Long Short-Term Memory (LSTM) network is employed to recognize grasp types from hand landmarks. The product-specific actions, aggregated from the basic hand movements, are formally modeled in a semantic description language based on the Web Ontology Language (OWL). Description Logic (DL) is used to define the actions with their characteristic properties, which enables the efficient classification of new action instances by an OWL reasoner.The semantic models of hand actions, robot tasks, and work-cell resources are interconnected and stored in a Knowledge Base (KB), which enables the efficient pair-wise translation between hand actions and robot tasks. For the reproduction of human assembly processes, actions are converted to robot tasks via skill descriptions, while reusing the action parameters of involved objects to ensure product integrity. We showcase and evaluate our method in an industrial production setting for control cabinet assembly. Demonstration video available at: https://kb-pbd.github.io/. Junsheng Ding, Haifan Zhang, Weihang Li, Liangwei Zhou, Alexander Clifford Perzylo |
IROS | 3 |
| 2024 | Determining the Minimum Number of Virtual Networks for Different Coherence ProtocolsabstractWe revisit the question of how many virtual networks (VNs) are required to provably avoid deadlock in a cache coherence protocol. The textbook way of reasoning about VNs says that the number of VNs depends on the longest chain of message dependencies in the protocol. We show that this conventional wisdom is incorrect and results in a number of virtual networks that is neither necessary nor sufficient for the general system model of an arbitrary interconnection network (ICN) topology and multiple directories. We have created a formalism for modeling coherence protocols and their interactions with ICN queueing. Using that formalism, we have developed an algorithm that (a) determines the minimum number of virtual networks required to avoid deadlock and (b) generates the mappings from message types to virtual networks. Weihang Li, Andres Goens, Nicolai Oswald, Vijay Nagarajan, Daniel J. Sorin |
ISCA | 1 |
| 2024 | SCRREAM : SCan, Register, REnder And Map: A Framework for Annotating Accurate and Dense 3D Indoor Scenes with a BenchmarkabstractTraditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be problematic as the meshes of the dataset are often incomplete and may produce wrong ground truth to evaluate the details. In this paper, we propose SCRREAM, a dataset annotation framework that allows annotation of fully dense meshes of objects in the scene and registers camera poses on the real image sequence, which can produce accurate ground truth for both sparse 3D as well as dense 3D tasks. We show the details of the dataset annotation pipeline and showcase four possible variants of datasets that can be obtained from our framework with example scenes, such as indoor reconstruction and SLAM, scene editing & object removal, human reconstruction and 6d pose estimation. Recent pipelines for indoor reconstruction and SLAM serve as new benchmarks. In contrast to previous indoor dataset, our design allows to evaluate dense geometry tasks on eleven sample scenes against accurately rendered ground truth depth maps. Weihang Li, William Bittner, Nikolas Brasch, Jifei Song, Eduardo Pérez-Pellitero, Zhensong Zhang, Arthur Moreau, Nassir Navab, Benjamin Busam |
NeurIPS | 2 |
| 2023 | TOD: Trend-Oriented Delay-Based Congestion Control in Lossless Datacenter NetworkabstractIn current high-speed data center networks, congestion control is crucial for ensuring consistent high performance. Over the past decade, researchers and developers have explored several congestion signals such as ECN, RTT, and INT. However, most of the existing congestion control algorithms suffer from either imprecise congestion detection due to ambiguous signals or excessive bandwidth loss due to aggressive rate decrease. This paper proposes a novel congestion control mechanism called TOD, which is a trend-oriented delay-based approach designed for lossless data center networks. TOD leverages the change in RTT to learn the congestion trend and adjusts the sending rate accordingly. By analyzing the congestion trend, the sender reacts by adjusting the sending rate to a reasonable level, while still maintaining high bandwidth utilization to dismiss congestion. The sender uses a reference rate, which is calculated by the receiver and communicated back to the sender, to achieve this target. Therefore, TOD is a sender-receiver cooperative congestion control mechanism. We evaluate TOD extensively in NS-3 simulations using both microbenchmark and macrobenchmark. Our experiments demonstrate that TOD outperforms DCQCN and Timely in terms of FCT and convergence speed. Kaixin Huang, Weihang Li, Lang Cheng |
APNet | 3 |
| 2023 | Deep Unrolling Shrinkage Network for Dynamic MR ImagingabstractDeep unrolling networks that utilize sparsity priors have achieved great success in dynamic magnetic resonance (MR) imaging. The convolutional neural network (CNN) is usually utilized to extract the transformed domain, and then the soft thresholding (ST) operator is applied to the CNN-transformed data to enforce the sparsity priors. However, the ST operator is usually constrained to be the same across all channels of the CNN-transformed data. In this paper, we propose a novel operator, called soft thresholding with channel attention (AST), that learns the threshold for each channel. In particular, we put forward a novel deep unrolling shrinkage network (DUS-Net) by unrolling the alternating direction method of multipliers (ADMM) for optimizing the transformed l1norm dynamic MR reconstruction model. Experimental results on an open-access dynamic cine MR dataset demonstrate that the proposed DUS-Net outperforms the state-of-the-art methods. The source code is available at https://github.com/yhao-z/DUS-Net. Xiaodi Li 0003, Weihang Li, Yue Hu 0003 |
ICIP | 3 |
| 2022 | ReverSearch: Search-based energy-efficient Processing-in-Memory ArchitectureabstractRecent development of the processing-in-memory (PIM) architecture has demonstrated high efficiency by reducing data movements. However, the performance of the conventional PIM architecture is limited by several issues, including frequent bit-line operations, complicated control of data flow, and massive inter-macro data movements. In addition, both analog- and digital-PIM solutions have obstacles to meet requirement of high-precision computation. In this work, we explore the tradeoff between data movement and energy efficiency of PIM architecture. We develop a PIM architecture, namely ReverSearch, to accelerate multiple-and-accumulate operation, equipped with reverse searching engine and look up table operations. Also, the corresponding data mapping and data flow methods are provided to improve the performance of the ReverSearch architecture. Based on our evaluation, ReverSearch improves the energy efficiency by 17.26 × and 3.68 ×, compared to the baseline of LUT-Cache [1] and LAcc [2]. Weihang Li, Liang Chang 0002, Jiajing Fan, Xin Zhao 0044, Hengtan Zhang, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 1 |
| 2021 | Energy-Efficient Spin-Orbit Torque MRAM Operations for Neural Network ProcessorabstractEmerging energy-efficient neural network processor is a promising hardware design to accelerate neural network algorithms with high performance and low power consumption. Typically, static random-access memory (SRAM) is employed to develop large buffers using in the processor. The bit cell of SRAM contains six transistors, leading to low density and large leakage current. In particular, several AI processors need multiple port and transfer-based SRAMs, which decrease the density and increase the power consumption. Recently, emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) becomes a possible solution to replace the SRAM as working memory. However, more operations should be supported by the SOT- MRAM to provide sufficient functions, such as multiple-port memory, transpose memory, data-streaming operations. In this paper, we develop the working memory of neural network processor with SOT-MRAM to build the design library including the transpose operations, multiple-port memory, and data-streaming based buffer arrays. Equiped with those operations provided by SOT-MRAM, we can build high performance and energy-efficient neural network processors. Liang Chang 0002, Zixuan Zhu 0001, Zhen Zhu 0005, Siqi Yang 0002, Weihang Li, Jun Zhou 0017 |
ISCAS | 5 |
| 2021 | Energy-efficient computing-in-memory architecture for AI processor: device, circuit, architecture perspective
Liang Chang 0002, Zhaomin Zhang, Jianbiao Xiao, Zhen Zhu 0005, Weihang Li, Zixuan Zhu 0001, Siqi Yang 0002, Jun Zhou 0017 |
Sci. China Inf. Sci. | 7 |