Guihang Hong

dblp:409/8588 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
3 papers
Edge and fog computing · 70% Internet architecture and protocols · 23% Network optimization and economics · 7%
Artificial intelligence
1 paper
Language models and text generation · 56% Efficient and distributed learning · 44%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% Distributed systems · 50%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed inference
0.912025
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing · RTSS 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing · RTSS 2025
Information retrieval
retrieval-augmented generation
0.912025
AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge · INFOCOM 2025
Edge and fog computing
collaborative edge computing
0.912025
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing · RTSS 2025
Edge and fog computing
edge inference
0.912025
AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge · INFOCOM 2025
Edge and fog computing › resource management
edge resource provisioning
0.912025
Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches · IEEE Trans. Netw. 2025
Internet architecture and protocols › packet scheduling
hierarchical scheduling
0.912025
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing · RTSS 2025
Distributed systems › resource sharing
application co-location
0.912025
Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches · IEEE Trans. Netw. 2025
Cloud and datacenter computing
cluster resource management and scheduling
0.912025
Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches · IEEE Trans. Netw. 2025
Network optimization and economics
resource allocation
0.312025
Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches · IEEE Trans. Netw. 2025

Methods — techniques the papers use, named apart from their topics

stochastic subgradient · 1.7retrieval-augmented generation · 1.7reinforcement learning · 1.7proximal policy optimization · 1.7online convex optimization · 1.7machine learning prediction · 1.7lagrange relaxation · 1.7adaptive optimization · 1.7
YearPublicationVenuePosition
2025 AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge
Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 0006, Weigang Wu, Zhaobiao Lv, Xu Chen 0004
INFOCOM2
2025 CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
abstract
Motivated by the imperative for real-time responsiveness and data privacy preservation, large language models (LLMs) are increasingly deployed on resource-constrained edge devices to enable localized inference. To improve output quality, retrieval-augmented generation (RAG) is an efficient technique that seamlessly integrates local data into LLMs. However, existing edge computing paradigms primarily focus on single-node optimization, neglecting opportunities to holistically exploit distributed data and heterogeneous resources through cross-node collaboration. To bridge this gap, we propose CoEdge-RAG, a hierarchical scheduling framework for retrieval-augmented LLMs in collaborative edge computing. In general, privacy constraints preclude accurate a priori acquisition of heterogeneous data distributions across edge nodes, directly impeding RAG performance optimization. Thus, we first design an online query identification mechanism using proximal policy optimization (PPO), which autonomously infers query semantics and establishes cross-domain knowledge associations in an online manner. Second, we devise a dynamic inter-node scheduling strategy that balances workloads across heterogeneous edge nodes by synergizing historical performance analytics with real-time resource thresholds. Third, we develop an intra-node scheduler based on online convex optimization, adaptively allocating query processing ratios and memory resources to optimize the latency-quality trade-off under fluctuating assigned loads. Comprehensive evaluations across diverse QA benchmarks demonstrate that our proposed method significantly boosts the performance of collaborative retrieval-augmented LLMs, achieving performance gains of 4.23 % to 91.39% over baseline methods across all tasks.
Guihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004
RTSS1
2025 Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches
abstract
Due to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed loosely for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating powerful machine learning techniques, we then advocate a predictive provisioning approach, where future request arrivals can be estimated accurately. To align with practical deployments, we incorporate a tunable prediction window mechanism, which well balances the potential improvement and degradation of online performance in imperfect prediction scenarios. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge.
Tao Ouyang, Kongyange Zhao, Guihang Hong, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Netw.3