Haiyang Shen

dblp:239/3363 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0000-4599-3198ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MCP-Focus: Leveraging Function-Oriented Document Enhancement for MCP Server Retrieval
abstract
Model Context Protocol (MCP) has emerged as a practical standard for connecting LLM-based agents with external tools and services through MCP servers. Driven by the open-source community, the MCP ecosystem is rapidly expanding, resulting in a large and growing collection of third-party MCP servers. Accurately selecting MCP servers that satisfy functional requirements from many candidates, therefore, becomes an increasingly important problem. However, MCP server documents are often unstructured and exhibit ambiguous function semantics, making it difficult to align user requirements with server capabilities during retrieval. To address this issue, we propose MCP-Focus, a function-oriented document enhancement framework that produces retrieval-ready MCP server documentation via a multi-stage agentic pipeline for white-box code analysis and document generation. Specifically, MCP-Focus first extracts a comprehensive tool inventory with metadata, then refines tool-level descriptions grounded in each extracted tool's implementation, and finally aggregates the refined tool descriptions into a structured server-level overview as the retrieval document. To better evaluate MCP server retrieval, we construct a benchmark comprising 3k+ open-source MCP servers and human-guided queries that vary in semantic ambiguity, input-output specificity, and the number of involved function points. Experiments across multiple dense retrievers show that fine-tuning with MCP-Focus-enhanced documents consistently improves retrieval effectiveness over baseline document methods on multiple benchmarks. Code and data: https://github.com/JingWC/MCP-Focus.
Wenchun Jing, Haiyang Shen, Qi Liu 0071, Ningyuan Li 0005, Chaoran Luo, Yun Ma 0002
SIGIR2
2026 LaTune: Lightweight and Adaptive Configuration Tuning for LLM Inference on Edge Devices
abstract
Large Language Models (LLMs) are increasingly deployed on edge devices to address privacy and latency concerns in modern Web applications. While numerous studies focus on inference frameworks, the critical problem of tuning runtime configurations remains largely underexplored. This endeavor is particularly challenging on edge devices due to severe budget limitations and the dynamic variability of system resources.
Siqi Zhong, Mugeng Liu 0001, Haiyang Shen, Chongyang Pan, Yun Ma 0002
WWW3
2026 Beyond the Sum of Parts: Leveraging Entanglement for Bug Inducing Commit Localization
abstract
Modern software development often introduces bug inducing commits (BICs) that can degrade performance or cause crashes. Swift localization of BICs is crucial but challenging due to the entanglement among multiple kinds of information overlooked by existing methods that treat these elements independently. Understanding this entanglement is promising but faces two key challenges: (1) the entanglement representation problem, since simply concatenating diverse data types fails to capture their entanglement effectively; (2) the large input size problem, as identifying BICs requires analyzing a vast set of commits, making simultaneous processing infeasible. To address these challenges, we propose BICSleuth, a framework that encodes the entanglement for effective BIC localization through three stages. First, a small-model-based ranker efficiently narrows down commits despite large input sizes. Second, an LLM-based discriminator deepens the understanding of the entanglement through selective information integration. Third, a reranking strategy combines insights from both stages to enhance localization accuracy. Evaluated on a BIC dataset constructed from Defects4J v2.0.0, BICSleuth outperforms four state-of-the-art approaches, achieving 148.9% of the Mean Reciprocal Rank compared to the best spectrum-based baseline and 507.1% of the MRR of the best IR-based method. Additionally, BICSleuth ranks the BIC first in 70.0% of projects and within the top five in 84.6%. The results demonstrate that BICSleuth effectively leverages the entanglement for BIC localization, with all stages contributing to its success.
Guoqing Wang 0004, Zeyu Sun 0004, Haiyang Shen, Qingyuan Liang, Dan Hao 0001
IEEE Trans. Software Eng.5
2025 ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
abstract
Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, their ability to handle multi-dimensional difficulty levels, diverse task types, and real-world demands remains unknown. In this paper, we introduce \textsc{ShortcutsBench}, a large-scale benchmark for the comprehensive evaluation of API-based agents in solving real-world complex tasks. \textsc{ShortcutsBench} includes a wealth of real APIs from Apple Inc., refined user queries, human-annotated high-quality action sequences, detailed parameter filling values, and parameters requesting necessary input from the system or user. We revealed how existing benchmarks~/~datasets struggle to accommodate the advanced reasoning capabilities of existing more intelligent LLMs. Moreover, our extensive evaluation of agents built with $5$ leading open-source (size $\geq$ 57B) and $5$ closed-source LLMs (e.g. Gemini-1.5-Pro and GPT-4o-mini) with varying intelligence level reveals significant limitations of existing API-based agents in the whole process of handling complex queries related to API selection, parameter filling, and requesting necessary input from the system and the user. These findings highlight the great challenges that API-based agents face in effectively fulfilling real and complex user queries. All datasets, code, experimental logs, and results are available at https://github.com/EachSheep/ShortcutsBench
Haiyang Shen, Desong Meng, Dongqi Cai 0001, Li Zhang 0133, Mengwei Xu 0001, Yun Ma 0002
ICLR1
2025 WeInfer: Unleashing the Power of WebGPU on LLM Inference in Web Browsers
abstract
Web-based large language model (LLM) has garnered significant attention from both academia and industry as it combines the benefits of on-device computation with the accessibility and portability of Web applications. The advent of WebGPU, a modern browser API that enables Web applications to utilize a device's GPU, has opened up new possibilities for GPU-accelerated LLM inference within browsers. However, our experiment reveals that existing Web-based LLM inference frameworks exhibit inefficiencies in GPU utilization, limiting the inference speed. These inefficiencies primarily arise from underutilizing the full capabilities of WebGPU, particularly in resource management and execution synchronization. To address these limitations, we present WeInfer, an efficient Web-based LLM inference framework specifically designed to unleash the power of WebGPU. WeInfer incorporates two key innovations: 1) buffer reuse strategies that reduce the overhead associated with resource preparation, optimizing the lifecycle management of WebGPU buffers, and 2) an asynchronous pipeline that decouples resource preparation from GPU execution, enabling parallelized computation and deferred result fetching to improve overall efficiency. We conduct extensive evaluations across 9 different LLMs and 5 heterogeneous devices, covering a broad spectrum of model architectures and hardware configurations. The results demonstrate that WeInfer delivers substantial improvements in decoding speed, achieving up to a 3.76× performance boost compared with WebLLM, the state-of-the-art Web-based LLM inference framework.
Yun Ma 0002, Haiyang Shen, Mugeng Liu 0001
WWW3
2025 WebAssembly for Container Runtime: Are We There Yet?
abstract
To pursue more efficient software deployment with containers, WebAssembly (abbreviated as Wasm) has long been regarded as a promising alternative to native container runtime (such as Docker container) due to its features of secure memory sandbox, lightweight isolation, portability, and multi-language support. However, it remains unknown whether and how much Wasm indeed brings benefits for containerized software applications. To fill the knowledge gap, this paper presents the first measurement study on Wasm-based container runtime (i.e., Wasm container) by comparison with the Docker container and native standalone Wasm runtime for execution performance in terms of the startup, computation, system interface access, and resource consumption. Surprisingly, we find that the Wasm container does not achieve better performance versus the Docker container as expected and introduces significant overhead compared to the standalone Wasm runtime. Through comparison, we identify the main causes of performance degradation for Wasm containers. Some stem from the heavy containerization overhead similar to Docker containers, while others are inherently caused by Wasm VMs and the WASI interface. Our findings can help software developers, Wasm container developers and the Wasm community improve the efficiency of utilizing Wasm-based container runtime, ultimately optimizing software performance.
Mugeng Liu 0001, Haiyang Shen, Hong Mei 0001, Yun Ma 0002
ACM Trans. Softw. Eng. Methodol.2
2024 Characterizing the Developer Groups for Metaverse Services in Roblox
abstract
The Metaverse has experienced exponential growth in recent years. Most metaverse platforms enable users to create their own metaverse services to deliver immersive content via visualized no-code/low-code programming tools. Usually, users with similar interests form or join a developer group to create and maintain complex metaverse services. In this paper, we focus on one of the most successful metaverse platforms, Roblox, to reveal how developer groups perform to create metaverse services. We collected a snapshot of Roblox encompassing 960,000 developer groups and 18.73 million Roblox game services. This dataset allows us to analyze the development patterns of these groups and identify key factors influencing their creativity. Our observations reveal that developer groups have diverse roles and members, exhibiting remarkable creativity that offers valuable insights into effective creation modes within the Metaverse. To understand why certain groups create exceptional game services, we examined various features of users at different ranks within the groups and developed a model to assess group creativity. Our findings indicate that the organizational structure within groups significantly impacts group creativity. Notably, we discovered that middle-ranked users, rather than top-ranked ones, play a more critical role in fostering creativity, highlighting their pivotal position in the group dynamic.
Haiyang Shen, Yun Ma 0002
SSE1
2023 ADPal: Automatic Detection of Troubled Users in Online Service Systems via Page Access Logs
abstract
Online service providers rely on customer service to enhance the experience of troubled users who encounter problems when interacting with online services. Nowadays, the customer service usually follows a reactive style, i.e., users with problems actively resort to help, and the customer service passively solves problems. But a more ideal style pursued by online service providers is proactive customer service, where users with problems can be detected in advance and notified of possible solutions before they resort to help. However, is it possible to detect users with problems? To answer the question, in this paper, we collect user traces of page access logs and problem feedback through customer service from a commercial online service provider Fliggy. We first verify an intuition that the page access logs are a good indicator to detect users with problems. Based on this verification, we design ADPal, an approach to detecting users with problems. Given a user’s page access log, ADPal outputs whether he/she encounters problems. ADPal leverages the capability of Transformer to extract relationships between pages, achieving high effectiveness and high efficiency. Evaluations on real-world data sets show that ADPal can achieve P@1000 74.70%, outperforming state-of-the-art anomaly detection approaches.
Haiyang Shen, Yun Ma 0002, Deyu Tian, Tengfei He, Shenghua Luo
ICWS1
2022 Parallelizing DNN inference in mobile web browsers on heterogeneous hardware
abstract
Mobile Web apps are emerging to leverage DNN models to provide intelligent user experience. But the limited functionalities for heterogeneous hardware provided by mobile Web browsers challenge the Web apps to perform DNN inference efficiently. In this paper, we propose a novel DNN inference engine, named PipeEngine, to parallelize the DNN inference process on CPU and GPU in mobile Web browsers. The design of PipeEngine enables pipeline parallelism between two adjacent DNN inference tasks with heterogeneous hardware. Evaluation results show that PipeEngine can increase the inference throughput by up to 2.77×.
Deyu Tian, Haiyang Shen, Yun Ma 0002
MobiSys2
2021 Design of low-profile array antenna working at 110 GHz based on digital coding characterization
Weixiang Jiang, Haiyang Shen
Sci. China Inf. Sci.3