Mugeng Liu 0001

dblp:375/3225-1 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0002-7625-8721ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reconfigurable Computing Challenge: FPGA-Based WebAssembly Stack Co-Processor
abstract
Large language models suffer from hallucinations when performing scientific computing, motivating the use of AI agents such as IronClaw that offload computation to specialized tools. IronClaw invokes tools implemented as WebAssembly (Wasm) plugins for security and extensibility, but the stack-based Wasm bytecode is mismatched with register-based processors (x86, ARM), causing runtime overhead. We propose PAWS, a native Wasm coprocessor that directly executes Wasm bytecode in hardware. PAWS features: (1) full support for all five Wasm instruction types; (2) dual digital stack circuits (operand stack and control stack) replacing register files to minimize memory access latency; (3) dedicated control logic for block-based branching; and (4) a sliding-window instruction fetch unit that decodes variable-length Wasm instructions. Evaluated on the PolyBench suite, PAWS achieves average execution latencies 28.6× lower than an Intel Xeon processor and 40.6× lower than an Nvidia Jetson TX2, making it highly suitable for IronClaw’s compute-intensive scientific applications. The design is available at https://github.com/Iris-WQP/PAWS_FPGA_softcore.
Qiuping Wu, Mugeng Liu 0001, Hongxiao Zhao, Yihan Fu, Gang Huang 0001, Yun Ma 0002, Bonan Yan
FCCM2
2026 LaTune: Lightweight and Adaptive Configuration Tuning for LLM Inference on Edge Devices
abstract
Large Language Models (LLMs) are increasingly deployed on edge devices to address privacy and latency concerns in modern Web applications. While numerous studies focus on inference frameworks, the critical problem of tuning runtime configurations remains largely underexplored. This endeavor is particularly challenging on edge devices due to severe budget limitations and the dynamic variability of system resources.
Siqi Zhong, Mugeng Liu 0001, Haiyang Shen, Chongyang Pan, Yun Ma 0002
WWW2
2025 SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection
abstract
Spreadsheets are critical to data-centric tasks, with rich, structured layouts that enable efficient information transmission.Given the time and expertise required for manual spreadsheet layout design, there is an urgent need for automated solutions.However, existing automated layout models are ill-suited to spreadsheets, as they often (1) treat components as axis-aligned rectangles with continuous coordinates, overlooking the inherently discrete, gridbased structure of spreadsheets; and (2) neglect interrelated semantics, such as data dependencies and contextual links, unique to spreadsheets.In this paper, we first formalize the spreadsheet layout generation task, supported by a seven-criterion evaluation protocol and a dataset of 3,326 spreadsheets.We then introduce SheetDesigner, a zero-shot and trainingfree framework using Multimodal Large Language Models (MLLMs) that combines rule and vision reflection for component placement and content population.SheetDesigner outperforms five baselines by at least 22.6%.We further find that through vision modality, MLLMs handle overlap and balance well but struggle with alignment, necessitates hybrid rule and visual reflection strategies.Our codes and data is available at Github.
Yuanyi Ren, Xiaojun Ma 0001, Mugeng Liu 0001, Shi Han, Dongmei Zhang 0001
EMNLP4
2025 WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers
abstract
Approximate nearest neighbor search (ANNS) has become vital to modern AI infrastructure, particularly in retrieval-augmented generation (RAG) applications. Numerous in-browser ANNS engines have emerged to seamlessly integrate with popular LLM-based web applications, while addressing privacy protection and challenges of heterogeneous device deployments. However, web browsers present unique challenges for ANNS, including computational limitations, external storage access issues, and memory utilization constraints, which state-of-the-art (SOTA) solutions fail to address comprehensively.
Mugeng Liu 0001, Siqi Zhong, Yudong Han 0001, Xuanzhe Liu, Yun Ma 0002
SIGIR1
2025 WeInfer: Unleashing the Power of WebGPU on LLM Inference in Web Browsers
abstract
Web-based large language model (LLM) has garnered significant attention from both academia and industry as it combines the benefits of on-device computation with the accessibility and portability of Web applications. The advent of WebGPU, a modern browser API that enables Web applications to utilize a device's GPU, has opened up new possibilities for GPU-accelerated LLM inference within browsers. However, our experiment reveals that existing Web-based LLM inference frameworks exhibit inefficiencies in GPU utilization, limiting the inference speed. These inefficiencies primarily arise from underutilizing the full capabilities of WebGPU, particularly in resource management and execution synchronization. To address these limitations, we present WeInfer, an efficient Web-based LLM inference framework specifically designed to unleash the power of WebGPU. WeInfer incorporates two key innovations: 1) buffer reuse strategies that reduce the overhead associated with resource preparation, optimizing the lifecycle management of WebGPU buffers, and 2) an asynchronous pipeline that decouples resource preparation from GPU execution, enabling parallelized computation and deferred result fetching to improve overall efficiency. We conduct extensive evaluations across 9 different LLMs and 5 heterogeneous devices, covering a broad spectrum of model architectures and hardware configurations. The results demonstrate that WeInfer delivers substantial improvements in decoding speed, achieving up to a 3.76× performance boost compared with WebLLM, the state-of-the-art Web-based LLM inference framework.
Yun Ma 0002, Haiyang Shen, Mugeng Liu 0001
WWW4
2025 WebAssembly for Container Runtime: Are We There Yet?
abstract
To pursue more efficient software deployment with containers, WebAssembly (abbreviated as Wasm) has long been regarded as a promising alternative to native container runtime (such as Docker container) due to its features of secure memory sandbox, lightweight isolation, portability, and multi-language support. However, it remains unknown whether and how much Wasm indeed brings benefits for containerized software applications. To fill the knowledge gap, this paper presents the first measurement study on Wasm-based container runtime (i.e., Wasm container) by comparison with the Docker container and native standalone Wasm runtime for execution performance in terms of the startup, computation, system interface access, and resource consumption. Surprisingly, we find that the Wasm container does not achieve better performance versus the Docker container as expected and introduces significant overhead compared to the standalone Wasm runtime. Through comparison, we identify the main causes of performance degradation for Wasm containers. Some stem from the heavy containerization overhead similar to Docker containers, while others are inherently caused by Wasm VMs and the WASI interface. Our findings can help software developers, Wasm container developers and the Wasm community improve the efficiency of utilizing Wasm-based container runtime, ultimately optimizing software performance.
Mugeng Liu 0001, Haiyang Shen, Hong Mei 0001, Yun Ma 0002
ACM Trans. Softw. Eng. Methodol.1
2025 Research on WebAssembly Runtimes: A Survey
abstract
WebAssembly (abbreviated as Wasm) was initially introduced for the Web and quickly extended its reach into various domains beyond the Web. To create Wasm applications, developers can compile high-level programming languages into Wasm binaries or manually write the textual format of Wasm and translate it into Wasm binaries by the toolchain. Regardless of whether it is utilized within or outside the Web, the execution of Wasm binaries is supported by the Wasm runtime. Such a runtime provides a secure, memory-efficient, and sandboxed execution environment to execute Wasm binaries. This article provides a comprehensive survey of research on Wasm runtimes with 103 collected research papers related to Wasm runtimes following the traditional systematic literature review process. It characterizes existing studies from two different angles, including the internal research of Wasm runtimes (Wasm runtime design, testing, and analysis) and the external research (applying Wasm runtimes to various domains). This article also proposes future research directions about Wasm runtimes.
Mugeng Liu 0001, Haoyu Wang 0001, Yun Ma 0002, Gang Huang 0001, Xuanzhe Liu
ACM Trans. Softw. Eng. Methodol.2
2024 Research artifacts in software engineering publications: Status and trends
Mugeng Liu 0001, Yibing Xie, Jie Zhang 0050, Xiang Jing, Zhenpeng Chen 0001, Yun Ma 0002
J. Syst. Softw.1