VLDB 2026 Research / reviewers in the wild / expert
Guanyu Gao
dblp:150/6757
· DBLP profile ↗
40ranked-venue papers
14as first author
26since 2021 · last 2026
0000-0001-8584-0532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 10 since 2021Computer networks · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning from Human Gaze: Human-like Robot Social Navigation in Dense CrowdsabstractRobot navigation in dense crowds requires understanding social cues that humans naturally use, yet existing methods struggle with real-world complexity. We investigate two questions: (1) Where do pedestrians look when navigating crowds? and (2) Can eye tracking improve robot navigation? To answer, we introduce GazeNav, an egocentric dataset collected via wearable eye trackers, featuring synchronized video, gaze, and trajectories in crowded environments. Analysis reveals that the gaze of pedestrians is closely related to the semantic presence and movement of other individuals, exhibiting distinct attention patterns across navigation behaviors. Building on this, we propose Gaze2Nav, a modular framework that first predicts human gaze to infer socially salient pedestrians, then incorporates the semantic attention into motion planning alongside visual inputs. Our method achieves 87.6% salient pedestrian prediction accuracy and reduces trajectory error by 15.4% over state-of-the-art baselines. By aligning with human gaze, our framework improves both performance and interpretability, advancing toward human-like, socially intelligent robot navigation. Zhecheng Yu, Yishuang Zhang, Bo Ling, Guanyu Gao, Weiwei Wu 0001, Brian Y. Lim |
AAAI | 8 |
| 2026 | Learn to Recover: Deep Reinforcement Learning for Failure Recovery in Large Networks
Zhiyu Fan, Guanyu Gao, Xueyong Xu, Vincent Chau, Weiwei Wu 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over QuantityabstractEdge computing-based video analytics faces data drift issues due to the occurrence of unseen objects or scenes in ever-changing environments. To maintain accuracy, continuous learning (CL) retrains stale models periodically with newly obtained data. However, it leads to unaffordable costs, as we must keep labeling drift data and retraining models. Regarding this concern, we first investigate video patterns across multiple cameras within an area and reveal significant data redundancies. We find that many of the same objects can be captured by multiple edge cameras or appear many times on the same edges. Our quantitative findings suggest that selecting a subset of high-quality data for CL is preferable over using a larger quantity. Yet, existing efforts for data acquisition have only focused on a single static dataset. These methods are not suitable for multi-edge video analytics scenarios, where videos are captured from multiple sources with non-iid data distribution. Hence, we propose a multi-edge collaborative active video acquisition (AVA) framework to collaboratively learn a reinforced video acquisition strategy to identify informative video frames from multiple edge nodes that best enhance model accuracy, avoiding redundancy across edges. Extensive experiments on three video datasets demonstrate that, our method achieves comparable performance to full-set video training while utilizing only 20% of the data in classification tasks. In object detection tasks, our methods can maintain productive accuracy with a reduction of nearly 70% in training video frames. Guanyu Gao, Haiyan Yin, Huaizheng Zhang |
AAAI | 2 |
| 2025 | SuperCache: Cost-Optimal Edge-Cloud Video Delivery through Integrated Caching, Super-Resolution, Transcoding, and Fetching
Luting Cao, Guanyu Gao |
APNet | 2 |
| 2025 | VCA: Tile Visual Contribution-Aware Streaming Adaptation for Volumetric VideoabstractVolumetric video enables immersive six Degrees of Freedom (6-DoF) viewing experiences but poses significant throughput challenges for streaming large-scale and high-quality volumetric video. This paper introduces VCA, a dynamic optimization method for volumetric video streaming based on a Quality of Experience (QoE) hierarchical optimization model utilizing 3D tile visual contribution. VCA optimally allocates bitrates to each 3D tile in real time through hierarchical optimization based on 3D tile visual contribution, network throughput, and buffer status. Our method features three key innovations: (1) a QoE prediction model based on 3D tile visual contribution that uses the Structural Similarity Index Measure (SSIM) as an objective quality metric and weights the rendered frame quality according to visual contribution; (2) a real-time algorithm that dynamically computes tile visual contribution by considering both intra-tile and inter-tile occlusion relationships; and (3) a method for real-time prediction of rendering quality for 3D tiles at arbitrary quality levels across arbitrary viewports. Experimental results demonstrate that VCA significantly outperforms baseline approaches, achieving QoE improvements of 9.79% and 33.27% in small-scale and large-scale scenarios, respectively, while maintaining throughput efficiency, making it a robust solution for high-quality volumetric video streaming. Zongzhou Xia, Guanyu Gao |
GLOBECOM | 2 |
| 2025 | ASDMR: Adaptive Scheduling of DNN Models and Resources for Collaborative Edge-Cloud Video AnalyticsabstractCollaborative edge-cloud video analytics is hampered by a fundamental trade-off among accuracy, latency, and resource cost. Cloud-centric processing suffers from prohibitive bandwidth consumption and delay, while edge-only solutions are constrained by limited computational power. To resolve this challenge, we propose an Adaptive Scheduling framework for DNN Models and Resources (ASDMR). ASDMR formulates the problem as a multiobjective optimization aimed at maximizing long-term system utility. The framework intelligently orchestrates key decisions in real-time, including adaptive model selection on heterogeneous edge nodes, dynamic task scheduling (local, edge-to-edge, or cloud), adaptive resolution scaling, content-aware result reuse, and load-based frame filtering. We employ Lyapunov optimization to decompose this long-term stochastic problem into a series of deterministic subproblems that are solved online in each time slot. This approach allows the system to dynamically balance analysis accuracy against resource cost while satisfying latency constraints, without needing to predict future workloads. Extensive experimental results demonstrate that ASDMR significantly outperforms baseline methods under diverse conditions, achieving a more effective trade-off across accuracy, cost, and latency. The framework thus offers a robust and practical solution for building efficient, large-scale collaborative video analytics systems. Yanfen Liang, Guanyu Gao |
ICPADS | 2 |
| 2025 | UniStream: Unifying In-Network Video Processing and Caching for Cost-Optimal Edge-Cloud Video StreamingabstractThe proliferation of multi-resolution video streaming, driven by Adaptive Bitrate (ABR) technology, imposes significant storage and computational burdens on content delivery networks. While edge computing alleviates latency, the cost of storing and processing numerous video versions remains a key challenge. To address this, we propose UniStream, an optimization framework that unifies in-network video processing—including Super-Resolution (SR) and transcoding—with intelligent caching and fetching strategies in a collaborative edge-cloud architecture. UniStream dynamically determines the most cost-effective action for each video request, deciding whether to serve content from cache, generate it on-the-fly via SR or transcoding, or fetch it from a peer node or the central cloud. We formulate this complex resource allocation challenge as a Mixed-Integer Linear Programming (MILP) problem, optimizing for minimal total cost under realistic storage, compute, and latency constraints. Extensive experiments using real-world datasets demonstrate that UniStream significantly reduces operational costs compared to baseline methods while maintaining competitive delivery latency, proving highly effective across diverse user request patterns and network conditions. Luting Cao, Guanyu Gao |
MMAsia | 2 |
| 2025 | FCLHet: Spatiotemporal Knowledge Continual Learning for Federated Heterogeneous ModelsabstractFederated Learning (FL) is a key distributed learning paradigm challenged by model heterogeneity and temporal shifts in data distribution. Existing methods to these problems are often impractical for heterogeneous environments, as they typically require the exchange of model parameters or the use of auxiliary public datasets. To address these limitations, we propose FCLHet, a novel Federated Continual Learning framework designed for resource-constrained, Heterogeneous environments. The core of FCLHet is a server-side spatiotemporal knowledge cache that collects data features from clients. This cache, combined with an active selection strategy guided by a client-side policy network, enables mining and sharing of personalized knowledge subsets. This mechanism enables collaborative learning without sharing model parameters or auxiliary data, allowing heterogeneous clients to continually learn and adapt while mitigating catastrophic forgetting. Extensive experiments on four benchmark datasets show that FCLHet improves average accuracy by 2–5% while keeping communication overhead comparable to state-of-the-art federated distillation techniques. Jinke Zhou, Guanyu Gao |
MMAsia | 2 |
| 2025 | QRALadder: QoE and Resource Consumption-Aware Encoding Ladder Optimization for Live Video Streaming
Yingqian Zhu 0002, Guanyu Gao |
MMM (3) | 2 |
| 2025 | Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret MinimizationabstractHuman-in-the-loop (HIL) imitation learning enables agents to learn complex behaviors safely through real-time human intervention. However, existing methods struggle to efficiently leverage agent-generated data due to dynamically evolving trajectory distributions and imperfections caused by human intervention delays, often failing to faithfully imitate the human expert policy. In this work, we propose Faithful Dynamic Imitation Learning (FaithDaIL) to address these challenges. We formulate HIL imitation learning as an online non-convex problem and employ dynamic regret minimization to adapt to the shifting data distribution and track high-quality policy trajectories.
To ensure faithful imitation of the human expert despite training on mixed agent and human data, we introduce an unbiased imitation objective and achieve it by weighting the behavior distribution relative to the human expert's as a proxy reward.
Extensive experiments on MetaDrive and CARLA driving benchmarks demonstrate that FaithDaIL achieves state-of-the-art performance in safety and task success with significantly reduced human intervention data compared to prior HIL baselines. Bo Ling, Zhengyu Gan, Wanyuan Wang, Guanyu Gao, Weiwei Wu 0001 |
NeurIPS | 4 |
| 2025 | CSVA: Complexity-Driven and Semantic-Aware Video Analytics via Edge-Cloud Collaboration
Guanyu Gao |
WASA (3) | 2 |
| 2025 | Spatial-Temporal Federated Learning for Lifelong Person Re-Identification on Distributed EdgesabstractData drift is a thorny challenge when deploying person re-identification (ReID) models into real-world devices, where the data distribution is significantly different from that of the training environment and keeps changing. To tackle this issue, we propose a federated spatial-temporal incremental learning approach, named FedSTIL, which leverages both lifelong learning and federated learning to continuously optimize models deployed on many distributed edge clients. Unlike previous efforts, FedSTIL aims to mine spatial-temporal correlations among the knowledge learnt from different edge clients. Specifically, the edge clients first periodically extract general representations of drifted data to optimize their local models. Then, the learnt knowledge from edge clients will be aggregated by centralized parameter server, where the knowledge will be selectively and attentively distilled from spatial- and temporal-dimension with carefully designed mechanisms. Finally, the distilled informative spatial-temporal knowledge will be sent back to correlated edge clients to further improve the recognition accuracy of each edge client with a lifelong learning method. Extensive experiments on a mixture of five real-world datasets demonstrate that our method outperforms others by nearly 4% in Rank-1 accuracy, while reducing communication cost by 62%. All implementation codes are publicly available on https://github.com/MSNLAB/Federated-Lifelong-Person-ReID. Guanyu Gao, Huaizheng Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Minimizing Age of Semantic Information for Analytics-Oriented Video Streaming SystemsabstractVideo streaming systems are critical for intelligent applications to transmit video data from end devices to servers for real-time analysis. In contrast to traditional human-centric streaming systems, which prioritize user-perceived metrics, machine-centric streaming systems are designed to continuously provide fresh and accurate information for analytics purposes. Although numerous studies have investigated policies to optimize streaming performance, most of them employ the segment-by-segment streaming framework from human-centric systems. Through comprehensive theoretical analysis and experimentation, we uncover that the segmented streaming approach is sub-optimal for machine-centric streaming systems compared to the straightforward frame-by-frame streaming approach. Furthermore, instead of relying on conventional frame-level metrics, we introduce a novel metric called the Age of Semantic Information (AoSI) to evaluate the performance of analytics-oriented streaming systems. This metric balances the quantity and timeliness of the semantic information. Consequently, we propose a compression ratio adaption method tailored to optimize AoSI performance for frame-by-frame streaming systems. This method leverages a deep learning (DL)-based predictor to discover the dynamic, latent relationships between compression and inference accuracy. Evaluated on actual streaming prototypes and real-world datasets, our method significantly surpasses both segmented and frame-by-frame baseline methods in terms of worst-case and average AoSI performance. Ziyao Huang 0001, Weiwei Wu 0001, Kui Wu 0001, Guanyu Gao, Jianping Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Towards Cost-Optimal Policies for DAGs to Utilize IaaS Clouds With Online LearningabstractPremier cloud service providers (CSPs) offer on-demand and spot instances (i.e., virtual machines) with time-varying features in availability and price. While interacting with a CSP, what concerns users is the process of cost-effectively utilizing these instances, possibly in addition to self-owned instances. A job in data-intensive applications is represented by a directed acyclic graph (DAG) whose nodes represent tasks and the DAG can further be transformed into a chain of tasks. The key to achieving cost efficiency is determining the allocation of a specific deadline to each task, as well as the allocation of different types of instances to the task. In this paper, we make some mild assumptions on the usage of various instances and give an analytical understanding of the expected behaviors while utilizing instances. Based on this, we propose informed heuristic policies to determine the allocation of deadlines and instances. The policies are parametric to support the usage of online learning to infer the optimal values against the dynamics of cloud markets. Finally, intuitive greedy policies are used as baselines to validate the effectiveness of the proposed analytical solutions. The cost improvement is up to 23.58% when spot and on-demand instances are considered and up to 50.08% when self-owned instances are also considered. Han Yu 0001, Giuliano Casale, Guanyu Gao |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | i-Rebalance: Personalized Vehicle Repositioning for Supply Demand BalanceabstractRide-hailing platforms have been facing the challenge of balancing demand and supply. Existing vehicle reposition techniques often treat drivers as homogeneous agents and relocate them deterministically, assuming compliance with the reposition. In this paper, we consider a more realistic and driver-centric scenario where drivers have unique cruising preferences and can decide whether to take the recommendation or not on their own. We propose i-Rebalance, a personalized vehicle reposition technique with deep reinforcement learning (DRL). i-Rebalance estimates drivers' decisions on accepting reposition recommendations through an on-field user study involving 99 real drivers. To optimize supply-demand balance and enhance preference satisfaction simultaneously, i-Rebalance has a sequential reposition strategy with dual DRL agents: Grid Agent to determine the reposition order of idle vehicles, and Vehicle Agent to provide personalized recommendations to each vehicle in the pre-defined order. This sequential learning strategy facilitates more effective policy training within a smaller action space compared to traditional joint-action methods. Evaluation of real-world trajectory data shows that i-Rebalance improves driver acceptance rate by 38.07% and total driver income by 9.97%. Peiyan Sun, Qiyuan Song, Wanyuan Wang, Weiwei Wu 0001, Wencan Zhang, Guanyu Gao |
AAAI | 7 |
| 2024 | SocialGAIL: Faithful Crowd Simulation for Social Robot NavigationabstractNavigation through crowded human environments is challenging for social robots. While reinforcement learning has been adopted for its capacity to capture complex interactions, the training process often relies on simulators to replicate realistic crowd behaviors, ensuring cost-efficiency. Existing crowd simulation methods typically rely on either handcrafted rules, which may lead to overly aggressive navigation, or learning from human trajectory demonstrations, which can be challenging to generalize effectively. In this paper, we introduce a data-driven crowd simulation method called SocialGAIL, which leverages Generative Adversarial Imitation Learning (GAIL) to emulate real pedestrian navigation in crowded environments. SocialGAIL utilizes an attention-based graph neural network to encode observations and employs a generator-discriminator architecture to closely mimic pedestrian behavior. We propose a set of metrics to evaluate the faithfulness of crowd simulation. Experimental results demonstrate that SocialGAIL outperforms baseline methods in terms of goal-reaching, intermediate state faithfulness, trajectory faithfulness, and adherence to global trajectory patterns. The code of our approach is available at https://github.com/William-island/SocialGAIL. Bo Ling, Guanyu Gao, Yi Shi 0011, Xueyong Xu, Weiwei Wu 0001 |
ICRA | 4 |
| 2024 | nHAS: Neural-Compensated Hybrid Adaptive Scheduling for Cloud Gaming
Qianyun Gong, Jiapei Xu, Jianxin Shi 0005, Xinjing Yuan, Jingdong Xu, Guanyu Gao, Lingjun Pu |
NPC (1) | 6 |
| 2024 | EdgeVision: Towards Collaborative Video Analytics on Distributed Edges for Performance MaximizationabstractDeep Neural Network (DNN)-based video analytics significantly improves recognition accuracy in computer vision applications. Deploying DNN models at edge nodes, closer to end users, reduces inference delay and minimizes bandwidth costs. However, these resource-constrained edge nodes may experience substantial delays under heavy workloads, leading to imbalanced workload distribution. While previous efforts focused on optimizing hierarchical device-edge-cloud architectures or centralized clusters for video analytics, we propose addressing these challenges through collaborative distributed and autonomous edge nodes. Despite the intricate control involved, we introduce EdgeVision, a Multiagent Reinforcement Learning (MARL)-based framework for collaborative video analytics on distributed edges. EdgeVision enables edge nodes to autonomously learn policies for video preprocessing, model selection, and request dispatching. Our approach utilizes an actor-critic-based MARL algorithm enhanced with an attention mechanism to learn optimal policies. To validate EdgeVision, we construct a multi-edge testbed and conduct experiments with real-world datasets. Results demonstrate a performance enhancement of 33.6% to 86.4% compared to baseline methods. Guanyu Gao, Yuqi Dong, Ran Wang 0004, Xin Zhou 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Edge-Assisted Joint Rate Adaptation and Quality Enhancement for 360-Degree Video Streamingabstract360-degree video offers users an immersive viewing experience by enabling them to look around during playback. However, this type of video consumes a significant amount of data, necessitating high bandwidth and minimal streaming delays. Super-Resolution (SR) is an emerging technique that can enhance video quality while reducing data transmission requirements. It achieves this by sending low-quality video and utilizing an SR model on the client side to enhance its quality. Nonetheless, SR is computationally demanding and typically requires high-performance hardware, making it impractical for many weak clients, particularly mobile devices. To overcome this limitation, we propose EdgeSR, an edge-assisted approach that combines rate adaptation and quality enhancement for 360-degree video streaming. In EdgeSR, the client can request low-quality video chunks from the cloud and perform SR at the edge to enhance video quality while minimizing transmission delays from the cloud to the edge. Simultaneously, the edge node can prefetch a video chunk from the cloud while the previous one is being transmitted from the edge to the client, thereby reducing delays. Experimental results under various settings demonstrate that EdgeSR outperforms the most competitive baseline methods, improving the overall Quality of Experience (QoE) by 27%. Kunkun Li, Guanyu Gao |
MMSP | 2 |
| 2023 | Deep Reinforcement Learning for UAV-Assisted Spectrum Sharing Under Partial ObservabilityabstractThis paper proposes a dynamic spectrum sharing scheme in an unmanned aerial vehicle (UAV) assisted cognitive radio network. The UAV serves as a secondary base station to provide communication services to multiple secondary users (SUs) by adaptively utilizing the spatio-temporal spectrum opportunities of multiple device-to-device primary users (PUs), where each PU’s spectrum occupancy follows a two-state Markov process. We jointly optimize the UAV’s trajectory and user association to maximize the expectation of its cumulative energy efficiency subject to the interference constraint of the PUs. We formulate this problem as a partially observable Markov decision process (POMDP), where the UAV can only observe the spectrum occupancy status of the adjacent PUs. Due to the lack of the PUs’ spectrum occupancy statistics, we propose a model-free reinforcement learning algorithm named partially observable double deep Q network (PO-DDQN) to obtain the near-optimal spectrum sharing policy. Simulation results show that our proposed algorithm outperforms the baseline policy gradient (PG) algorithm in terms of convergence speed and the UAV’s energy efficiency. Additionally, the spectrum utilization efficiency can be further enhanced when the UAV has wider observation radius, or if the PUs’ spectrum occupancy exhibits stronger temporal correlation. Sigen Zhang, Zhe Wang 0005, Guanyu Gao, Jun Li 0004, Jie Zhang 0006, Ziyan Yin |
VTC Fall | 3 |
| 2023 | Distributed Energy Trading and Scheduling Among Microgrids via Multiagent Reinforcement LearningabstractRenewable energy technologies empower microgrids to generate electricity to supply themselves and trade with others. Under this paradigm, microgrids have become autonomous entities that must intelligently determine their policies for energy trading and scheduling. Many factors influence a microgrid's decision-making, such as the complex microgrid infrastructure, the uncertain energy yield and demand, and the competition among the energy market players. These factors are usually hard to precisely model, and deriving the optimal policy for a microgrid is challenging. We propose a multiagent reinforcement learning (MARL) approach with an attention mechanism to learn the optimal policies for the microgrids without complex system modeling. We model each microgrid as an autonomous agent, which learns how to schedule energy resources and trade with others by collaborating with other agents. We adopt attention mechanism to enable intelligently selecting contextual information for the training of each agent. After training, an agent can make control decisions using only its local information, which can well preserve the microgrids' privacy and reduce the communication overhead among microgrids to facilitate distributed control. We implement a simulation environment and evaluate the performances of our proposed method using real-world datasets. The experimental results show that our method can significantly reduce the cost of the microgrids compared with the baseline methods. Guanyu Gao, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Dynamic DNN model selection and inference off loading for video analytics with edge-cloud collaborationabstractThe edge-cloud collaboration architecture can support Deep Neural Network-based (DNN) video analytics with low inference delays and high accuracy. However, the video analytics pipelines with edge-cloud collaboration are complex, involving the decision-making for many coupled control knobs. We propose a deep reinforcement learning-based approach, named ModelIO, for dynamic DNN Model selection and Inference Offloading for video analytics with edge-cloud collaboration. We jointly consider the decision-making for video pre-processing, DNN model selection, local inference, and offloading in a video analytics system to maximize performances. Our method can learn the optimal control policy for video analytics with the edge-cloud collaboration without complex system modeling. We implement a real-world testbed to conduct the experiments to evaluate the performances of our method. The results show that our method can significantly improve the system processing capacity, reduce average inference delays, and maximize overall rewards. Xuezhi Wang 0007, Guanyu Gao, Weiwei Wu 0001 |
NOSSDAV | 2 |
| 2022 | Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart CachingabstractContinuous learning (CL) has recently been adopted into edge video analytics, gaining huge success in maintaining high accuracy without constantly retraining DNN models by human intervention. Though existing solutions offer optimized processing pipelines, the cost brought by CL should not be neglected. This vision paper starts an investigation by exploring two kinds of cost, human labeling and edge storage. The former comes from the need for CL's automatically tuning, and the latter is due to an exemplar pool (including both drift and historical data) maintained to prevent catastrophic forgetting caused by naive retraining. To alleviate the costs, we propose a new CL-based edge video analytics system by incorporating an active learner mechanism. Specifically, we revisit the current CL video system design and develop an active CL pipeline atop them. The pipeline first accepts the drift data stored in drift pool and utilizes an active learner to sample a small partition of them for labeling. Then it mixes up both small labeled drifted data and some historical data to send them to an exemplar pool for CL. Our preliminary benchmark studies exhibit that the new system can achieve competitive accuracy by spending only 30% labeling and storage cost compared to other baselines, showing a promising research direction for future study. Guanyu Gao, Huaizheng Zhang |
SenSys | 2 |
| 2022 | FogChain: A Blockchain-Based Peer-to-Peer Solar Power Trading System Powered by Fog AIabstractMicrogrids, gaining traction from rising distributed generation for carbon reduction, demand novel solutions to regulate on- and off-grid operations, as well as both energy and monetary transfers between the microgrid and the central grid and among different microgrid participants. This research aims to develop and validate an intelligent microgrid management system to secure the competitiveness of Singapore’s energy market, by leveraging the inherent synergy between two emerging technologies, i.e., blockchain for Peer-to-Peer (P2P) solar power trading and fog computing for grid infrastructure management. For this vision, we have developed FogChain, an integrative, cost-effective, and scalable microgrid operating system (MGOS), consisting of three technical service layers: 1) a novel microgrid information infrastructure based on the fog-computing paradigm (i.e., intelligence on edge); 2) a blockchain-based microgrid service layer, providing smart contract and decentralized control capabilities for grid application development; and 3) a microgrid application layer (i.e., P2P energy trading) over the blockchain-based grid service. This MGOS would fundamentally transform how solar power is traded among participating electricity prosumers, leading to potentially new operational and business models. We have implemented the FogChain system and conducted extensive experiments to verify its performance advantages. Our results demonstrate that FogChain can efficiently process energy auction among 1000 participants with 1.1 s delay on average, reduce transmission cost up to 20% under the loss-aware trading mechanism, and reduce the solar yield prediction error to 0.11. Our system prototype suggests that FogChain provides a promising solution for efficient decentralized energy trading and intelligent distributed control for microgrids. Guanyu Gao, Chengru Song, T. G. Thusitha Asela Bandara, Meng Shen 0002, Fan Yang 0172, Wolf Posdorfer, Dacheng Tao, Yonggang Wen 0001 |
IEEE Internet Things J. | 1 |
| 2021 | SmartEye: An Open Source Framework for Real-Time Video Analytics with Edge-Cloud CollaborationabstractVideo analytics with Deep Neural Networks (DNNs) empowers many vision-based applications. However, deploying DNN models for video analytics services must address the challenges of computational capacity, service delay, and cost. Leveraging the edge-cloud collaboration to address these problems has become a growing trend. This paper provides the multimedia research community with an open source framework named SmartEye for real-time video analytics by leveraging the edge-cloud collaboration. The system consists of 1) an edge layer which enables video preprocessing, model selection, on-edge inference, and task offloading; 2) a request forwarding layer which serves as a gateway of the cloud and forwards the offloaded tasks to backend workers; and 3) a backend worker layer that processes the offloaded tasks with specified DNN models. One can easily customize the policies for preprocessing, offloading, model selection, and request forwarding. The framework can facilitate research and development in this field. The project is released as an open source project on GitHub at https://github.com/MSNLAB/SmartEye. Xuezhi Wang 0007, Guanyu Gao |
ACM Multimedia | 2 |
| 2021 | Toward Intelligent Multizone Thermal Control With Multiagent Deep Reinforcement LearningabstractEnergy usage and thermal comfort are the pillars of smart buildings. Many research works have been proposed to save energy while maintaining a comfortable thermal condition. However, most of them either make the oversimplified assumption on thermal comfort with unsatisfied comfort performance or deal with the single-zone thermal control only with limited practical impact. A few preliminary pieces of research on multizone control are available, but they fail to keep pace with the latest advancements in the deep-learning-based control techniques. In this article, we investigate the multizone thermal control with optimized energy usage and canonical thermal comfort modeling. We adopt the emerging multiagent deep reinforcement learning techniques and propose to model each zone as an agent. A multiagent framework is established to support the information exchange among the agents and enable intelligent thermal control in the heterogeneous zones. Accordingly, we mathematically formulate a problem to optimize both energy and comfort. A multizone thermal control algorithm (MOCA) is proposed to solve the problem by deriving optimal control policies. We validate the performance of MOCA through simulation in professional TRNSYS, configured based on our real-world laboratory. The results are promising with up to 15.4% energy saving as well as satisfied thermal comfort in different zones. Jie Li 0043, Wei Zhang 0082, Guanyu Gao, Yonggang Wen 0001, Guangyu Jin, George I. Christopoulos |
IEEE Internet Things J. | 3 |
| 2020 | Bridging the Web Data and Fine-Grained Visual Recognition via Alleviating Label Noise and Domain MismatchabstractTo distinguish the subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, we propose to directly leverage web images for fine-grained visual recognition. Our work mainly focuses on two critical issues including "label noise" and "domain mismatch" in the web images. Specifically, we propose an end-to-end deep denoising network (DDN) model to jointly solve these problems in the process of web images selection. To verify the effectiveness of our proposed approach, we first collect web images by using the labels in fine-grained datasets. Then we apply the proposed deep denoising network model for noise removal and domain mismatch alleviation. We leverage the selected web images as the training set for fine-grained categorization models learning. Extensive experiments and ablation studies demonstrate state-of-the-art performance gained by our proposed approach, which, at the same time, delivers a new pipeline for fine-grained visual categorization that is to be highly effective for real-world applications. Yazhou Yao, Xian-Sheng Hua 0001, Guanyu Gao, Zeren Sun, Zhibin Li 0002, Jian Zhang 0002 |
ACM Multimedia | 3 |
| 2020 | DeepComfort: Energy-Efficient Thermal Comfort Control in Buildings Via Reinforcement LearningabstractHeating, ventilation, and air conditioning (HVAC) are extremely energy consuming, accounting for 40% of total building energy consumption. It is crucial to design some energy-efficient building thermal comfort control strategy which can reduce the energy consumption of the HVAC while maintaining the comfort of the occupants. However, implementing such a strategy is challenging, because the changes of the thermal states in a building environment are influenced by various factors. The relationships among these influencing factors are hard to model and are always different in different building environments. To address this challenge, we propose a deep-reinforcement-learning-based framework, DeepComfort, for thermal comfort control in buildings. We formulate the thermal comfort control as a cost-minimization problem by jointly considering the energy consumption of the HVAC and the occupants' thermal comfort. We first design a deep feedforward neural network (FNN)-based approach for predicting the occupants' thermal comfort and then propose a deep deterministic policy gradients (DDPGs)-based approach for learning the optimal thermal comfort control policy. We implement a building thermal comfort control simulation environment and evaluate the performance under various settings. The experimental results show that our approaches can improve the performance of thermal comfort prediction by 14.5% and reduce the energy consumption of HVAC by 4.31% while improving the occupants' thermal comfort by 13.6%. Guanyu Gao, Jie Li 0043, Yonggang Wen 0001 |
IEEE Internet Things J. | 1 |
| 2020 | DeepQoE: A Multimodal Learning Framework for Video Quality of Experience (QoE) PredictionabstractRecently, many models have been developed to predict video Quality of Experience (QoE), yet the applicability of these models still faces significant challenges. Firstly, many models rely on features that are unique to a specific dataset and thus lack the capability to generalize. Due to the intricate interactions among these features, a unified representation that is independent of datasets with different modalities is needed. Secondly, existing models often lack the configurability to perform both classification and regression tasks. Thirdly, the sample size of the available datasets to develop these models is often very small, and the impact of limited data on the performance of QoE models has not been adequately addressed. To address these issues, in this work we develop a novel and end-to-end framework termed as DeepQoE. The proposed framework first uses a combination of deep learning techniques, such as word embedding and 3D convolutional neural network (C3D), to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. A learned representation will then serve as input for classification or regression tasks. We evaluate the performance of DeepQoE with three datasets. The results show that for small datasets (e.g., WHU-MVQoE2016 and Live-Netflix Video Database), the performance of state-of-the-art machine learning algorithms is greatly improved by using the QoE representation from DeepQoE (e.g., 35.71% to 44.82%); while for the large dataset (e.g., VideoSet), our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). In addition to the much improved performance, DeepQoE has the flexibility to fit different datasets, to learn QoE representation, and to perform both classification and regression problems. We also develop a DeepQoE based adaptive bitrate streaming (ABR) system to verify that our framework can be easily applied to multimedia communication service. The software package of the DeepQoE framework has been released to facilitate the current research on QoE. Huaizheng Zhang, Linsen Dong, Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Kyle Guan |
IEEE Trans. Multim. | 3 |
| 2019 | Content-Aware Personalised Rate Adaptation for Adaptive Streaming via Deep Video AnalysisabstractAdaptive bitrate (ABR) streaming is the de facto solution for achieving smooth viewing experiences under unstable network conditions. However, most of the existing rate adaptation approaches for ABR are content-agnostic, without considering the semantic information of the video content. Nevertheless, semantic information largely determines the informativeness and interestingness of the video content, and consequently affects the QoE for video streaming. One common case is that the user may expect higher quality for the parts of video content that are more interesting or informative so as to reduce overall subjective quality loss. This creates two main challenges for such a problem: First, how to determine which parts of the video content are more interesting? Second, how to allocate bitrate budgets for different parts of the video content with different significances? To address these challenges, we propose a Content-of-Interest (CoI) based rate adaptation scheme for ABR. We first design a deep learning approach for recognizing the interestingness of the video content, and then design a Deep Q-Network (DQN) approach for rate adaptation by incorporating video interestingness information. The experimental results show that our method can recognize video interestingness precisely, and the bitrate allocation for ABR can be aligned with the interestingness of video content while not compromising the performances on objective QoE metrics. Guanyu Gao, Linsen Dong, Huaizheng Zhang, Yonggang Wen 0001, Wenjun Zeng 0001 |
ICC | 1 |
| 2019 | Dynamic Priority-Based Resource Provisioning for Video Transcoding With Heterogeneous QoSabstractVideo transcoding is widely adopted in online video services to transcode videos into multiple representations for dynamic adaptive bitrate streaming. This solution may consume significant resources and incur intolerable processing delays. Meanwhile, different videos have different quality-of-service (QoS) requirements for transcoding. Delay-sensitive videos must be transcoded within a strict deadline, whereas delay-tolerant videos are not required to be transcoded immediately. Some intelligent policies are required for provisioning the right amount of resources in the transcoding system to meet the heterogeneous QoS requirements, especially under dynamic workloads. To this end, we develop a robust dynamic priority-based resource provisioning scheme for video transcoding. We adopt the preemptive resume priority discipline to design a multiple-priority transcoding mechanism. The system performs the transcoding for delay-tolerant videos by utilizing idle resources for improving resource utilization while not affecting the transcoding for delay-sensitive videos. We adopt the model predictive control framework to design an online algorithm for dynamic resource provisioning to accommodate time-varying workloads by predicting future workloads. To seek performance robustness against prediction noise, we improve the performance of our online algorithm via robust design. The experimental results demonstrate that our proposed method can satisfy the heterogeneous QoS requirements while significantly reducing computing resource consumption. Guanyu Gao, Yonggang Wen 0001, Cédric Westphal |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Deepqoe: A Unified Framework for Learning to Predict Video QoEabstractMotivated by the prowess of deep learning (DL) based techniques in prediction, generalization, and representation learning, we develop a novel framework called DeepQoE to predict video quality of experience (QoE). The end-to-end framework first uses a combination of DL techniques (e.g., word embeddings) to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. Such representations serve as inputs for classification or regression tasks. Evaluating the performance of DeepQoE with two datasets, we show that for the small dataset, the accuracy of all shallow learning algorithms is improved by using the representation derived from DeepQoE. For the large dataset, our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). Moreover, DeepQoE, also released as an open source tool, provides video QoE research much-needed flexibility in fitting different datasets, extracting generalized features, and learning representations. Huaizheng Zhang, Han Hu 0003, Guanyu Gao, Yonggang Wen 0001, Kyle Guan |
ICME | 3 |
| 2018 | Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest InferenceabstractRate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches. Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | QDLCoding: QoS-differentiated low-cost video encoding scheme for online video serviceabstractAdaptive bitrate (ABR) streaming is the de facto solution in online video services to cope with heterogeneous devices and varying network connections. However, this solution is computation intensive, demanding a large number of servers for encoding videos. Moreover, due to the time-varying nature of video generation, intelligent strategies are required in order to determine the right amount of resources for encoding. The situation is further complicated by the fact that, the two types of co-existing video content, live content and Video-on-Demand (VoD) content, have different QoS requirements for encoding. These observations posit daunting challenges for meeting the heterogeneous QoS requirements with a minimum computing capacity. This paper proposes the QoS-differentiated low-cost video encoding (QDLCoding) scheme to address these challenges. We develop a framework for scheduling the encoding workloads of the two types of videos with statistical QoS guarantees. Each type of videos is specified with a QoS criterion and a QoS loss bound. The objective is to provision the minimum amount of resources while keeping the QoS loss probabilities within the prescribed bounds. We design an online algorithm that can determine the minimum required capacity by learning content arrival distributions. The experiment results demonstrate that our method can greatly reduce the required capacity for encoding online videos while controlling the likelihood of QoS loss precisely. Guanyu Gao, Yonggang Wen 0001, Han Hu 0003 |
INFOCOM | 1 |
| 2017 | Resource Provisioning and Profit Maximization for Transcoding in Clouds: A Two-Timescale ApproachabstractTranscoding is widely adopted for content adaptation; however, it may incur excessive resource consumption and processing delays. Taking advantage of cloud infrastructure, cloud-based transcoding can elastically allocate resources under time-varying workloads and perform multiple transcodings in parallel to reduce delays. To provide transcoding as a cloud service, cloud transcoding systems require some intelligent mechanisms to provision resources and schedule tasks to satisfy user requirements while maximizing financial profit. To this end, we propose a two-timescale stochastic optimization framework for maximizing service profit while achieving performance requirements by jointly provisioning resources and scheduling tasks under a hierarchical control architecture. Our method analytically integrates service revenue, processing delay, and resource consumption in one optimization framework. We derive the offline exact solution and design some approximate online solutions for task scheduling and resource provisioning. We implement an open source cloud transcoding system, called Morph, and evaluate the performance of our method in a real environment. Empirical studies verify that our method can reduce resource consumption and achieve a higher profit compared with baseline schemes. Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Cédric Westphal |
IEEE Trans. Multim. | 1 |
| 2016 | Morph: A Fast and Scalable Cloud Transcoding SystemabstractMorph is an open source cloud transcoding system. It can leverage the scalability of the cloud infrastructure to encode and transcode video contents in fast speed, and dynamically provision the resources in cloud to accommodate the workload. The system is composed of a master node that performs the video file segmentation, concentration, and task scheduling operations; and multiple worker nodes that perform the transcoding for video blocks. Morph can transcode the video blocks of a video file on multiple workers in parallel to achieve fast speed, and automatically manage the data transfers and communications between the master node and the worker nodes. The worker nodes can join into or leave the transcoding cluster at any time for dynamic resource provisioning. The system is very modular, and all of the algorithms can be easily modified or replaced. We release the source code of Morph under MIT License, hoping that it can be shared among various research communities. Guanyu Gao, Yonggang Wen 0001 |
ACM Multimedia | 1 |
| 2016 | Dynamic Resource Provisioning with QoS Guarantee for Video Transcoding in Online Video Sharing ServiceabstractVideo transcoding is widely adopted in online video sharing services to encode video content into multiple representations. This solution, however, could consume huge amount of computing resource and incur excessive processing delays. Moreover, content has heterogeneous QoS requirements for transcoding. Some content must be transcoded in real time, while some are deferrable for transcoding. It needs to determine the strategy for intelligently provisioning the right amount of resource under dynamic workload to meet the heterogeneous QoS requirements. To this end, this paper develops a robust dynamic resource provisioning scheme for transcoding with heterogeneous QoS criteria. We adopt the Preemptive Resume Priority discipline for scheduling, so that the transcoding-deferrable content can utilize idle resources for transcoding to maximize resource utilization while remain transparent to delay-sensitive content. We leverage Model Predictive Control to design the online algorithm for dynamic resource provisioning using predictions to accommodate time-varying workload. To seek robustness of system performance against prediction noises, we improve our online algorithm through Robust Design. The experiment results in a real environment demonstrate that our proposed framework can achieve the QoS requirements while reducing 50% of resource consumption on average. Guanyu Gao, Yonggang Wen 0001, Cédric Westphal |
ACM Multimedia | 1 |
| 2015 | Cost-efficient and QoS-aware content management in media cloud: Implementation and evaluationabstractAdaptive bitrate streaming has been proposed to encode video contents into multiple versions for device heterogeneity and changing network conditions. This solution, however, could consume enormous computing and storage resource. In fact, only a small fraction of videos are frequently requested. Thus, caching multiple versions for unpopular contents is not cost efficient. In this paper, we design a cost-efficient and QoS-aware content management system for video streaming. The system consists of a set of streaming servers and a computing cluster, where streaming servers can cache video contents or transcode them in real time, and the computing cluster can perform transcoding tasks on behalf of streaming servers. Based on this architecture, to provide cost-efficient and QoS-aware video service, first, we design a cost-efficient content cache management module to minimize the operational cost, by dynamically determining whether a segment should be cached or transcoded on fly according to their popularity. Second, to reduce transcoding latency, we design a QoS-aware transcoding task delegation module to determine whether a transcoding task in streaming server should be delegated to the computing cluster according to the streaming server's workload. We implement the system and evaluate the performance in a real environment. The results demonstrate that our method can greatly reduce the operational cost and guarantee the QoS in providing video services. Guanyu Gao, Yonggang Wen 0001, Han Hu 0003 |
ICC | 1 |
| 2015 | Towards Cost-Efficient Video Transcoding in Media Cloud: Insights Learned From User Viewing PatternsabstractVideo transcoding in an adaptive bitrate streaming (ABR) system is demanded to support video streaming over heterogenous devices and varying networks. However, it could incur a tremendous cost. Meanwhile, most viewers terminate viewing sessions within 20% of their durations; only a small fraction of each video is consumed. Built upon this user viewing pattern, we propose a Partial Transcoding Scheme for content management in media clouds. Particularly, each content is encoded into different bitrates and split into segments. Some of the segments are stored in cache, resulting in storage cost; others are transcoded online in the case of cache miss, resulting in computing cost. We aim to minimize the long-term overall cost by determining whether a segment should be cached or transcoded online. We formulate it as a constrained stochastic optimization problem. Leveraging Lyapunov optimization framework and Lagrangian relaxation, we design an online algorithm which can achieve the optimal solution within provable upper bounds. Experiments demonstrate that our proposed method can reduce 30% of operational cost, compared with the scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Cost optimal video transcoding in media cloud: Insights from user viewing patternabstractVideo transcoding has been touted as an enabling technology to support growing media consumption over heterogenous devices. However, on-line transcoding could incur tremendous, if not prohibitive, cost in deploying or renting resources. In this research, we leverage an insight into the viewing pattern of video consumers to reduce the operating cost of video transcoding services. Specifically, it has been reported that viewers tend to terminate their session before the whole video is watched. As such, it is not cost-efficient for service providers to store or transcode all segments of the videos. Built upon this insight, we propose a partial transcoding scheme for content management in a media cloud to reduce the operating cost. Particularly, each content is split into multiple segments and stored in different files of varying playback rates. Some of the segments are stored in cache, resulting in storage cost; while some are transcoded in real-time in case of cache miss, resulting in computing cost. We aim to minimize the long-term operational cost by determining the number of segments for each playback rate to be cached or transcoded in real-time. We formulate this partial transcoding scheme as a constrained integer optimization problem. Leveraging Lagrangian relaxation and a subgradient method, we obtain the approximate solution to the integer program. Numerical results indicate that our proposed partial transcoding scheme can save more than 30% of operational cost, compared with a brute-force scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001, Yap-Peng Tan |
ICME | 1 |