Bingheng Yan

dblp:49/6139 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
11since 2021 · last 2026
0009-0009-6255-1055ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Entangled Photon Source Pooling and Entanglement Distribution for Quantum Networks
Junyuan Shi, Yangming Zhao, Bingheng Yan, Chen Tian 0001, Chunming Qiao
IWQoS4
2026 Scaling NVMM-based file system on intensive shared file access
Qiqi Gu 0002, Chenpeng Wu, Bingheng Yan, Jianguo Yao 0002
J. Syst. Archit.4
2026 Cooling as You Wish: Component-Level Cooling for Heterogeneous Edge Datacenters
abstract
As computing shifts toward the edge, edge datacenters are becoming essential for supporting diverse real-time applications. Unlike traditional cloud datacenters, edge datacenters face unique cooling challenges due to their requirements forproximity to end users, high density, and hardware heterogeneity. While warm water cooling is a promising technique for this infrastructure, current one-size-fits-all cooling strategies significantly compromise efficiency due to severe inter- and intra-component hotspots. In this work, we present CoolEdge+, a cost-effective component–level water cooling system for enhancing the cooling efficiency of edge datacenters. Specifically, CoolEdge+dynamically adjusts the inlet water temperature for each component through a carefully designed water circulation architecture to mitigate inter-component hotspots. To address intra-component hotspots, it employs vapor chamber–based cold plates that rapidly dissipate heat without manual intervention or additional energy consumption. We further design a fine-grained cooling control framework that leverages a well-managed power capping approach to decide on customized inlet water temperatures and hardware power limits. Based on a hardware prototype and a real-world trace from Alibaba PAI, evaluation results show that CoolEdge+reduces cooling energy consumption by up to 27.19% compared to existing coarse-grained systems, while maintaining performance guarantees. Compared to the state-of-the-art CoolEdge, CoolEdge+saves 35.24% more cooling costs with comparable energy consumption and no latency violations.
Fangming Liu, Qiangyu Pei, Yongjie Yuan, Qixia Zhang, Ziyang Jia, Fei Xu 0009, Bingheng Yan
IEEE Trans. Computers10
2025 Metaverse Service Provisioning Empowered by Monitoring and Analytical Digital Twins in MEC
abstract
Metaverse, as the cyberspace against the real world, offers various immersive services enabling users to entertain, learn and work. The digital twin (DT) technology acts as a fundamental enabler of Metaverse services by mapping user devices to DTs and timely analyzing the status of user devices. Mobile edge computing (MEC) significantly improves the QoS of DT-empowered Metaverse services because it can deploy DTs in cloudlets closer to user devices. However, each user device often has a complex structure with multiple interdependent subsystems that frequently communicate. Further, the resources in MEC network are highly distributed and limited. Thus, provisioning DT-empowered Metaverse services in MEC networks faces the challenge of efficiently mapping intertwining subsystems in user devices to DTs, maximizing admitted service requests and resource utilization. In this paper, we first formulate throughput maximization problems for DT-empowered Metaverse services in an MEC network, with the aim to maximize the total data rate of service requests of Metaverse services while meeting resource capacity constraints of the MEC network. We then propose an approximation algorithm with a provable approximation ratio for the problem with a given set of service requests if both monitoring and analytical DTs are consolidated into a single edge server. Otherwise, we devise an efficient heuristic for the problem with monitoring and analytical DTs possibly being placed into different edge servers. We also consider a dynamic throughput maximization problem of Metaverse service provisioning for a given monitoring period, in which service requests arrive into the system dynamically without the knowledge of future arrivals, service resource demands, and service delay requirements, for which we devise an online learning algorithm. We finally evaluate the performance of the proposed algorithms by extensive simulations. Simulation results show that the throughputs of the proposed algorithms outperform their comparison counterparts by at least 50 %.
Guangyuan Xu, Zichuan Xu, Qiufen Xia, Bingheng Yan, Pengyuan Xu
HPCC4
2025 It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
abstract
Serverless platforms typically adopt an earlybinding approach for function sizing, requiring developers to specify an immutable size for each function within a workflow beforehand. Accounting for potential runtime variability, developers must size functions for worst-case scenarios to ensure service-level objectives (SLOs), resulting in significant resource inefficiency. To address this issue, we propose Janus, a novel resource adaptation framework for serverless platforms. Janus employs a late-binding approach, allowing function sizes to be dynamically adapted based on runtime conditions. The main challenge lies in the information barrier between the developer and the provider: developers lack access to runtime information, while providers lack domain knowledge about the workflow. To bridge this gap, Janus allows developers to provide hints containing rules and options for resource adaptation. Providers then follow these hints to dynamically adjust resource allocation at runtime based on real-time function execution information, ensuring compliance with SLOs. We implement Janus and conduct extensive experiments with real-world serverless workflows. Our results demonstrate that Janus enhances resource efficiency by up to 34.7% compared to the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Quanfeng Deng, Chen Yu 0003, Bingheng Yan, Fangming Liu
IPDPS6
2025 Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage Parallelization
abstract
This paper presents Nova, a real-time scheduling framework for serving agentic vision-language models (VLMs) on a single GPU with balanced per-request latency and overall request process throughput. Our design begins by enabling effective pipelining across vision encode, LLM prefill, and LLM decode stages of VLMs, by exploiting their heterogeneous resource demands during execution and incorporating elastic GPU spatial partitioning among stages to maximally utilize the compute and memory resources. Building on this, we introduce a realtime scheduling algorithm that adaptively calibrates resource allocation among stages based on a Pareto-optimal analysis of the latency-throughput trade-off, allowing the system to sustain responsiveness and resource efficiency under dynamic request loads. To further alleviate GPU memory pressure, we design a lightweight weight offloading strategy for vision encoders that preserves inference efficiency with minimized memory overhead. Extensive evaluations on both synthetic and real-world agent workloads demonstrate that Nova consistently outperforms the state-of-the-art baselines, improving the maximum latency by up to 23.3 %, while keeping competitive throughput.
Shengzhong Liu, Bingheng Yan, Fan Wu 0006, Guihai Chen
RTSS4
2025 Entangled qubit pricing for quantum networks
Yangming Zhao, Shouxi Luo, Haoze Chen, Chen Tian 0001, Bingheng Yan
Comput. Networks8
2025 Troubleshooting Programmable Data Planes via Real-Time Table Information Recording
abstract
While the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption.
Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.6
2025 Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge Datacenters
abstract
The proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications.
Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
IEEE Trans. Sustain. Comput.6
2024 InferCool: Enhancing AI Inference Cooling through Transparent, Non-Intrusive Task Reassignment
abstract
The increasing power consumption of AI inference in modern datacenters has escalated cooling demands significantly, necessitating the adoption of potent cooling approaches like water cooling. Unlike traditional cloud workloads, AI inference has unique characteristics that create substantial gaps in achieving optimal cooling efficiency. In this work, we present the first comprehensive measurement study of AI inference cooling across various models within an industrial-ready scheduling framework, highlighting significant inefficiencies and their causes. To fill the gap while following the fundamental requirements of cooling systems, we explore a new opportunity presented by modern Multi-Instance GPU-enabled inference serving, where the scheduling dimension is naturally orthogonal to the cooling dimension. Building on this insight, we develop InferCool, a cooling middleware designed to enhance cooling efficiency for inference serving through transparent, non-intrusive task reassignment. It includes a streamlined power and temperature prediction approach and a thermal-aware, adaptive application deployment and request scheduling mechanism. Real-world experiments on a water-cooled testbed and a three-node cluster demonstrate that InferCool can reduce the maximum GPU temperature by 5°C across eight A100 GPUs, equivalent to cooling energy savings of about 20%. Importantly, InferCool requires no modifications to existing cooling infrastructures and is compatible with existing scheduling systems.
Qiangyu Pei, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
SoCC4
2024 MpScope: Enabling multi-pipeline monitoring inside a switch
Chengyuan Huang, Tianfan Zhang, Li Wang 0110, Yibo Xiao, Chen Tian 0001, Xiaoliang Wang 0001, Bingheng Yan, Ahmed M. Abdelmoniem, Wan-Chun Dou, Guihai Chen
Comput. Networks9
2009 Reducing Communication Overhead in Threshold Monitoring with Arithmetic Aggregation
abstract
With increasing adoption of distributed systems, monitoring has become an important research topic in recent years. Monitoring itself introduces overhead to the system caused by communication for collecting measurement data between monitoring nodes and measuring nodes. Reducing communication frequency of monitoring is very significant, especially for threshold monitoring which only cares about whether some metric crosses certain threshold or not. Currently threshold monitoring only considers simple aggregation values such as sum or average of values of measurement data. However, in this paper, we take a further exploration to arithmetic aggregation which is more complicated arithmetic result of measurement data, not just the sum or average of them. We present an approach to solve the communication overhead problem of arithmetic aggregation in threshold monitoring, where corresponding algorithms are designed. In our approach, the global threshold can be split into many local thresholds which can be set remotely on distributed local monitoring nodes. Communication will only take place when the locally observed numerical value of measurement exceeds corresponding local thresholds. We argue that communications will be reduced significantly based on the assumption that numerical values of monitored measurements would not oscillate widely or rapidly all the time. By conducting experiments based on data from real scenes, we can demonstrate that our solutions outperform the traditional approaches.
Yuanqiang Huang, Yinan Ren, Bingheng Yan, Zhongzhi Luan, Depei Qian 0001
NAS4
2009 Data Currency in Replicated Distributed Storage System
abstract
Application-level storage aggregation provides a massive storage capacity with high scalability and low cost. However, these systems usually only support their special sites, which ignores the legacy storage systems. So we developed a distributed storage system to aggregate these popular storages and provide a uniform access and management interface for these sites. To ensure high data availability, our distributed storage system utilizes a sophisticated way known as data replication. In this paper, we present the data currency scenario employed in our distributed storage system, which can be also used in other replicated distributed storage systems. We propose a replica access service (RAS) to deal with data availability and efficient retrieval of current replicas based on version controlling, which can balance the data accessibility and replica consistency. We validate our solution's performance and scalability through simulation up to 10,000 distributed sites. The simulation results show that our algorithm used in RAS achieves major performance gains, in terms of response time, compared with a baseline algorithm.
Bingheng Yan, Depei Qian 0001, Yuanqiang Huang
NAS1
2008 EOMT: A Master-Slave Task Scheduling Strategy for Grid Environment
abstract
Task scheduling has been a key issue to improve parallel execution in distributed systems. Master-slave task scheduling, as a technique of mapping and scheduling loads to heterogeneous platforms, has aroused interests of many researchers. Although minimizing the master-slave application's makespan (the overall completion time) in general case is a NP-complete problem, it is still meaningful in some special fields. In this paper, we aim at improving the performance of the master-slave pattern applications in the case with a large number of equal-sized and independent tasks and propose a new strategy EOMT (equilibrium overhead with multi-cycle tasking) for task scheduling in the grid environment. The EOMT strategy is designed for the grid environment with heterogeneous resources. The main concept of EOMT is to make the workload assigned to each slave node as even as possible to reduce application's makespan. A detailed analysis for master-slave task scheduling is given in this paper. Experiment results show that our strategy outperforms other traditional task scheduling strategies in different computation and network resource combinations in the grid environment.
Yuanqiang Huang, Depei Qian 0001, Zhongzhi Luan, Zhongxin Wu, Bingheng Yan
HPCC6
2007 Context-Aware Web Service Selection Based on Multi-aspects Regulating
abstract
Context-aware Web service selection is an adaptive process of offering suitable Web service to Web service consumer. Service provider and consumer have their own context that affect their service access and supply, such as locations, input, output, devices and platforms. The paper develops Web service selection architecture for providing available context-aware Web service through clustering, context extracting, matching and scheduling methods. A prototype system integrates and validates all the methods.
Zhongxin Wu, Dongbo Yang, Bingheng Yan, Depei Qian 0001
APSCC4
2007 Experiences with the EUChinaGrid Project - Implementing Interoperation between gLite and GOS
abstract
Great changes have taken place in Grid Technology field and various grids are constructed for sharing data and collaborating in a large scale model of cross-organization and cross-region. However, technologies of existing grids are different from each other and many grid islands appear. The interoperation among different grids becomes more and more challenging for building a global grid and extending the scope of sharing and collaborating. In this paper, we present our experiences with implementing interoperation between two different grid middlewares: gLite and GOS, which is a summarization of the EUChinaGrid project and the contents focus on fourfold: job submission, data management, information service and schema, and security discuss.
Bingheng Yan, Zhongxin Wu, Dongbo Yang, Depei Qian 0001
APSCC1
2007 Agent-Based MADM Approach to the Dynamic Web Service Selection
abstract
The Business Process Execution Language (BPEL) has become the de-facto standard for the description of Web Service compositions. A variety of formal approaches to decide compatibility and consistency for BPEL processes has been presented. Nevertheless, these approaches suffer from high complexity and state explosion. Therefore we present a lean formalization of BPEL 2.0 based on the pi-calculus, that enables efficient reasoning. Due to our focus on behavioral compatibility and consistency checking (and not on comprehensive formalization), we are able to reduce effort needed for process verification. Besides the exemplary application of our approach, we also compare it to existing BPEL formalizations by means of complexity.
Dongbo Yang, Zhongxin Wu, Bingheng Yan, Depei Qian 0001, Zhongzhi Luan
APSCC3
2007 Interconnect EGEE and CNGRID e-Infrastructures through Interoperability between gLite and GOS Middlewares
abstract
Interoperability between the grid middlewares gLite, used in European grid infrastructure EGEE, and GOS, used in the Chinese grid infrastructure CNGRID infrastructure, is one of the main activities of the EUChinaGRID project. The EUChinaGRID project is an initiative, funded by the European Commission, to extend the European GRID infrastructure for e-Science to China. The first aim of EUChinaGRID is to facilitate scientific data transfer and processing in a first sample of scientific communities that have already strong collaborations between Europe and China. The interoperability between gLite and GOS middlewares and, then, the interconnection between EGEE and CNGRID infrastructures, is fundamental to achieve this goal. This paper presents the interoperability solution adopted by the EUChinaGRID project, including job submission, data transfer, and information services. The implementation of a gateway for achieving interoperability is discussed and preliminary results are presented.
Diego Scardaci, Bingheng Yan, Yuanqiang Huang
eScience3