Longhao Zou

dblp:137/6287 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-3477-7438ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network
Qi Zhao 0037, Shuchang Lyu, Longhao Zou
Pattern Recognit. Lett.5
2026 The Teacher-Student Interactive Cycle: Joint Optimization With Inner-Loop Self-Distillation in Prompted Foundation Models for Efficient Semantic Segmentation
abstract
In the field of semantic segmentation, the high computational cost of deep models poses a major barrier to deployment on edge devices. Among various efficiency-oriented methods, knowledge distillation has emerged as a promising technique for transferring knowledge from large models to lightweight networks. However, current knowledge distillation methods for efficient semantic segmentation still face two key challenges: (1) they often rely on large offline pre-trained teacher networks that remain fixed during training, and (2) they lack joint optimization mechanisms that enable effective teacher-student interaction in pixel-wise dense prediction. As a result, mutual learning strategies originally designed for image-level classification often fail to capture the fine-grained consistency required for semantic segmentation. To address these two challenges, we propose a novel training framework termed Teacher-Student Interactive Cycle (TSIC), which performs efficient semantic segmentation. Specifically, TSIC integrates a lightweight student network into a prompt-based foundation model as a prompted segmentor to assist an online-trained teacher. The student provides coarse mask prompts to guide the teacher, while the teacher offers fine-grained supervision through posterior probabilities and intermediate feature maps. This loop enables joint online optimization without relying on offline pre-trained teachers and fosters effective bidirectional communication. Extensive experiments conducted on several benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, demonstrate the effectiveness of TSIC. Compared to previous methods, TSIC achieves superior segmentation mIoU in most scenarios. Our code will be made publicly available at https://github.com/CV-ShuchangLyu/TSIC.
Qi Zhao 0037, Shuchang Lyu, Longhao Zou, Dingding Yao, Chenguang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Joint Bitrate and Resource Adaptation for Super-Resolution Video Streaming in Multi-Cluster Edge Networks: A New Online Learning Approach
abstract
Today's video streaming service providers have exploited cloud-edge collaborative networks for video delivery across geo-distributed edge clusters and end users. The existing content delivery network (CDN) scheduling and adaptive bitrate algorithms may not fully utilize edge resources or lack a global control to optimize resource sharing. The emerging super-resolution (SR) approach can unleash the potential of leveraging computation resources to compensate for bandwidth consumption, by producing high-quality videos from low-resolution contents. Yet the uncertain SR resource sensitivity and its interplay with bitrate adaptation are under-explored. In this work, we proposeRosevin, the first resource scheduler that jointly decides the bitrates and fine-grained resource allocation to perform SR at the edge, which can learn to optimize the long-term QoE for distributed end users. To handle the time-varying and complex space of decisions as well as a non-smooth objective function,Rosevinrealizes a novel online combinatorial learning algorithm, which nicely integrates convex optimization theories and online learning techniques, addressing the switching cost issues. In addition to theoretically analyzing its performance, we implement an SR-assisted video streaming prototype ofRosevinand demonstrate its advantages over several video delivery benchmarks.
Xiaoxi Zhang 0001, Longhao Zou, Jingpu Duan, Chuan Wu 0001, Yali Xue, Zuozhou Chen, Chaoqi Zhou, Xu Chen 0004
IEEE Trans. Mob. Comput.3
2025 Mixed Spiking NeRF: Towards a More Efficient Neural Radiance Fields
abstract
In recent years, 3D reconstruction has garnered significant attention. Traditional Neural Radiance Fields (NeRF) can generate high-quality, photorealistic 3D models, rendering object surfaces and texture details through computer vision tasks. However, traditional NeRF requires significant computational resources and extensive training time To further reduce power consumption and hardware requirements, this paper introduces a brain-inspired neural network mechanism aimed at optimizing low-power 3D reconstruction tasks and proposes a hybrid adaptive Spiking Neural Network (SNN) NeRF framework called Mixed Spiking NeRF. Spiking Neural Networks excel in handling sparse data, implementing information augmentation in the form of binary images, and utilizing Leaky Integrate-and-Fire (LIF) neurons to promote efficient communication across regions. In this work, we integrate the Group Neurons (GN) mechanism into the framework. We introduce comprehensive GN operations, including both horizontal and vertical GN, to enable efficient patch communication. Finally, through optimal group-based loss function optimization, across various datasets and scenarios, our framework achieved impressive results that indicate the energy consumption is reduced by 60%.
Kaian Wang, Longhao Zou
ICASSP2
2025 CWC-DNERF: Compact Dynamic Neural Radiance Field VIA Discrete Wavelet Transform And Learnable Codebooks
abstract
Neural radiance fields have significantly advanced dynamic scene reconstruction and novel view synthesis. However, relying on multiple implicit multi-layer perceptrons for reconstructing dynamic scenes is computationally expensive. Recent methods have alleviated this challenge by introducing explicit data structures, such as voxel grids and feature planes, but these significantly increase storage demands and complicate network transmission. We propose Cwc-DNeRF, a compact dynamic NeRF representation that leverages discrete wavelet transform (DWT) and learnable codebooks to achieve superior storage efficiency while maintaining competitive rendering quality compared to K-Planes. In Stage I, DWT and trainable masks are employed to optimize parameter efficiency, resulting in sparse space planes. In Stage II, learnable codebooks are introduced for the space-time planes to merge redundant spatio-temporal features further reducing the storage demand. Additionally, a data compression pipeline is applied to compress both sparse space plane parameters and codebooks. Experimental results on D-NeRF and DyNeRF datasets show that our method achieves state-of-the-art rendering quality within a 10MB storage budget while retaining the benefits of explicit feature planes.
Yaojian Xu, Qiudan Zhang, Longhao Zou, Qiong Liu 0001, Xu Wang 0006
ICIP4
2025 Calibration of Spatial Position Errors in Virtual and Real Spaces: Improving the Results of AR-Guided Wiring
abstract
Augmented Reality (AR) enables an immersive understanding of industrial operating behavior in a mixed environment of reality and virtual. Hence, an AR-guided navigation system for industrial wiring is proposed. The objective is to establish an interactive learning platform for the digital transmission of wiring skills to improve efficiency and quality. Firstly, the mapping relationship from virtual to real is constructed, and a reasonable spatial error threshold is set to reduce the error. Then, the visual positioning method is used to calibrate the spatial position errors of virtual parts to optimize the AR guidance effect. Finally, the virtual operation is accurately mapped to the real space for guidance. The results show that AR-guided wiring reduces operating time by 22% and errors by 74% compared to real wiring. Compared to general AR-guided wiring, the calibrated wiring exhibits an 80% reduction in operating errors. As a result, the virtual and real spatial calibration contributes to efficient AR-guided wiring for accurate digital training.
Haopeng He, Meipeng Huang, Jin Xie 0008, Linfeng Yang, Longhao Zou, Youbing Guo
Int. J. Hum. Comput. Interact.5
2024 Rosevin: Employing Resource- and Rate-Adaptive Edge Super-Resolution for Video Streaming
abstract
Today’s video streaming service providers have exploited cloud-edge collaborative networks for geo-distributed video delivery. The existing content delivery network (CDN) scheduling and adaptive bitrate algorithms may not fully utilize edge resources or lack a global control to optimize resource sharing. The emerging super-resolution (SR) approach can unleash the potential of leveraging computation resources to compensate for bandwidth consumption, by producing high-quality videos from low-resolution contents. Yet the uncertain SR resource sensitivity and its interplay with bitrate adaptation are underexplored. In this work, we propose Rosevin, the first resource scheduler that jointly decides the bitrates and fine-grained resource allocation to perform SR at the edge, which can learn to optimize the long-term QoE for distributed end users. To handle the time-varying and complex space of decisions as well as a non-smooth objective function, Rosevin realizes a novel online combinatorial learning algorithm, which nicely integrates convex optimization theories and online learning techniques. In addition to theoretically analyzing its performance, we implement an SR-assisted video streaming prototype of Rosevin and demonstrate its advantages over several video delivery benchmarks.
Xiaoxi Zhang 0001, Longhao Zou, Jingpu Duan, Chuan Wu 0001, Yali Xue, Zuozhou Chen, Xu Chen 0004
INFOCOM3
2024 CoSPAM: Multi-Robot Collaboration Simultaneous Path Planning and Semantic Mapping
abstract
In recent years, with the rapid development of mobile network communication and robotic vehicle technolo-gies, there has been a growing emphasis on the design and development of sustainable intelligent robotic vehicle systems. To address this trend, we propose CoSPAM, an innovative multi-robot collaboration for simultaneous path planning and semantic mapping in a large-scale environment. CoSPAM integrates local and global strategies to optimize planning and mapping processes. Each robot operates autonomously and executes local rapidly exploring random tree exploration within its designated spatial environment. The robots establish wireless communication to transmit local map information to the server. We employ Bayesian-based fusion and Bundle Adjustment-based optimization techniques in the server to construct a comprehensive global 3D semantic OctoMap. Based on spatial memory mechanisms and time constraints, we utilize knowledge from all robots to formulate a cohesive global path plan. The experimental results demonstrate that CoSPAM reduces exploration time by almost 56.3% compared to traditional methods. It optimally allocates each robot's exploration goals, minimizing the exploration path's length. Moreover, CoSPamdemonstrates its capability to gen-erate a high-quality 3D semantic OctoMap. CoSPAM offers a reliable, agile, and energy-efficient solution for large-scale environment planning and mapping among sustainable intelligent robotic vehicles.
Longhao Zou, Gabriel-Miro Muntean
VTC Spring2
2024 Enhancing Telecooperation Through Haptic Twin for Internet of Robotic Things: Implementation and Challenges
abstract
The Internet of Robotic Things (IoRT) serves as a bridge for the progress of upcoming immersive interaction technologies like XR, holographic communications, and the metaverse, enabling smooth interaction between the tangible and virtual domains. It enriches the capabilities of sensors and controllers in both domains, enabling precise mapping and control. Moreover, immersive interaction technologies enhance IoRT by offering superior visualization and user experiences. Here, we propose a haptic-twin-enhanced telecooperation system (HTTS) delving into tactile interaction technology for IoRT, unveiling a telecooperation-enhanced haptic twin technology. Our goal is to construct haptic gloves devices and its haptic twin entity that are centered on virtual collaborations (i.e., virtual handshakes and collaborative tasks), deploying an experimental system for the acquisition, processing, transmission, reconstruction and interaction of multisensory data over wireless networks. We simulate and demonstrate the perception of touch, friction, and virtual gravity using the proposed haptic twin-based experimental testbed. Compared to commercial product, namely, HTC Vive Controller, our approach achieves a more reliable and precise construction improvement of haptic twin entity that is based on our implemented haptic gloves in terms of simulated human hand posture, ensuring consistent haptic feedback in virtual control and collaboration. This offers a more user-friendly and immersive experience with adaptive haptic effects, paving the way for novel telecooperation applications in the future IoRT.
Meipeng Huang, Runhui Feng, Longhao Zou, Jin Xie 0008
IEEE Internet Things J.3
2024 Enabling Efficient Vehicle-Road Cooperation Through AIoT: A Deep Learning Approach to Computational Offloading
abstract
The integration of Artificial Intelligence with the Internet of Things significantly enhances the functionality of vehicle–road cooperation (VRC) systems by enabling smarter, real-time decision-making and resource optimization across interconnected vehicular networks. To tackle the challenges associated with resource constraints, this study introduces a method where vehicle users can offload tasks to nearby roadside units (RSUs) or service-oriented vehicles to ensure timely application execution. However, this task offloading introduces additional transmission delays and energy expenditures. Consequently, this article first conceptualizes the computation offloading problem, aiming to minimize the total task processing time and energy consumption under the constraints of resources provided by RSUs and service-oriented vehicles. We model the computation offloading issue within the VRC framework as a Markov decision process (MDP) and propose a multiagent reinforcement learning-based resource scheduling method. Each vehicle, acting as an intelligent agent, interacts with and influences decisions within this environment. The method integrates the twin delayed deep deterministic policy gradient algorithm to train deep neural networks for deciding on task offloading and computational resource allocation. Simulation results demonstrate that compared to existing algorithms, the proposed method more effectively utilizes the computational resources available through RSUs and service-oriented vehicles within the VRC system. It achieves joint optimization of latency and energy consumption, thus validating the efficacy of the proposed approach in enhancing the operational efficiency and sustainability of urban transportation systems.
Xin Wang 0134, Madini O. Alassafi, Fawaz E. Alsaadi, Xingsi Xue, Longhao Zou
IEEE Internet Things J.5
2023 EMS-SLAM: Edge-Assisted Multi-Agent System Simultaneous Localization and Mapping
abstract
In recent years, there has been a growing demand for robotic environment perception and autonomous driving due to the increasing popularity of visual and geometry-based localization and mapping techniques, such as simultaneous localization and mapping (SLAM). To address this trend, this paper proposes the EMS-SLAM framework, which utilizes cooperative adaptive wireless communication between servers and multi-robot agents to enhance environment perception and self-localization efficiency and accuracy. EMS-SLAM can reduce mapping time and CPU and memory utilization of individual robots while maintaining high accuracy OctoMap based on multi-map fusion and optimization. EMS-SLAM’s effectiveness and real-time performance have been validated and tested on publicly available datasets and real robots for real-world operations. The experimental results demonstrate that EMS-SLAM can reduce the CPU utilization of a single robot by approximately 10% and improve the efficiency of large-scale SLAM. The constructed OctoMap achieves centimeter-level accuracy. EMS-SLAM provides reliable, agile, and energy-efficient assistance for large-scale environment perception of robots.
Lei Zhan, Longhao Zou, Zuozhou Chen, Gabriel-Miro Muntean
VTC2023-Spring3
2023 DiffTREAT: Differentiated Traffic Scheduling Based on RNN in Data Centers
abstract
Transmission schemes in data centers are supposed to accurately distinguish flow types for different scheduling. However, prior efforts failed to meet the needs at all levels in a cost-effectively way. Nor the existing schemes proved applicable to all the diverse scenarios or dynamic traffic patterns. Therefore, we proposedDifferentiatedTraffic schEduling in dAta cenTers (DiffTREAT) based the Recurrent Neural Network (RNN), aiming to simplify the transmission in the dynamic and diverse network scenarios. First, DiffTREAT utilizes deep learning methods for traffic classification and flow size prediction. Second, according to the classified results of flows, DiffTREAT adopts multilevel priority queues to ensure the preferential transmission of latency-sensitive flows while optimizing the overall average flow completion time (FCT). Third, DiffTREAT employs the network cache to increase the capacity of data center networks (DCN), which effectively fights against the traffic burst and improves the throughput of latency-insensitive flows. DiffTREAT has been tested in different topologies in the contexts of diverse network loads and real-world workloads. Experiment results showed that compared with state-of-the-art schemes, DiffTREAT yielded both the lower average flow completion time for latency-sensitive flows and the higher throughput for latency-insensitive flow.
Ziqi Wei 0004, Qing Li 0006, Keke Zhu, Jianer Zhou, Longhao Zou, Yong Jiang 0001, Xi Xiao 0001
IEEE Trans. Cloud Comput.5
2023 A Super-Resolution Flexible Video Coding Solution for Improving Live Streaming Quality
abstract
In the context of the latest growing popularity of live video streaming, ensuring high video quality has become one of the most significant challenges faced by all live streaming platforms. Insufficient uplink bandwidth is an important factor that influences these live video transmissions, affecting their bitrate and latency and consequently the associated video streaming quality. This paper proposes a novel flexible super-resolution-based video coding and uploading framework (FlexSRVC) that improves the quality of live video streaming in limited uplink network bandwidth conditions. FlexSRVC includes a flexible video coding scheme, which compresses high-resolution key and non-key video frames to a lower bitrate in order to reduce the upload delay. A new flexible bitrate adaptation algorithm is also proposed to select dynamically the number of frames to be compressed and the compression ratio by jointly considering uplink network conditions and available cloud computing resources. Trace-driven emulations demonstrate that FlexSRVC provides the same quality while reducing up to 25% of the required bandwidth compared to the original encoding method (H.264). FlexSRVC improves users' QoE by at least 50% compared to a super resolution-based method which employs reconstruction of all video frames in uplink bandwidth constrained conditions.
Qing Li 0006, Aoyang Zhang, Yong Jiang 0001, Longhao Zou, Zhimin Xu 0001, Gabriel-Miro Muntean
IEEE Trans. Multim.5
2022 Learning-based Fuzzy Bitrate Matching at the Edge for Adaptive Video Streaming
abstract
The rapid growth of video traffic imposes significant challenges on content delivery over the Internet. Meanwhile, edge computing is developed to accelerate video transmission as well as release the traffic load of origin servers. Although some related techniques (e.g., transcoding and prefetching) are proposed to improve edge services, they cannot fully utilize cached videos. Therefore, we propose a Learning-based Fuzzy Bitrate Matching scheme (LFBM) at the edge for adaptive video streaming, which utilizes the capacity of network and edge servers. In accordance with user requests, cache states and network conditions, LFBM utilizes reinforcement learning to make a decision, either fetching the video of the exact bitrate from the origin server or responding with a different representation from the edge server. In the simulation, compared with the baseline, LFBM improves cache hit ratio by 128%. Besides, compared with the scheme without fuzzy bitrate matching, it improves Quality of Experience (QoE) by 45%. Moreover, the real-network experiments further demonstrate the effectiveness of LFBM. It increases the hit ratio by 84% compared with the baseline and improves the QoE by 51% compared with the scheme without fuzzy bitrate matching.
Wanxin Shi, Qing Li 0006, Longhao Zou, Gengbiao Shen, Pei Zhang 0003, Yong Jiang 0001
WWW4
2022 Learning-Based Joint QoE Optimization for Adaptive Video Streaming Based on Smart Edge
abstract
The latest increase in HTTP-based adaptive video streaming over the Internet enables a growing number of clients to compete for a shared bottleneck bandwidth. This competition may affect users’ Quality of Experience (QoE) negatively, especially in terms of fairness and stability. This paper presentsFlex-Steward, a solution that performs multi-client joint QoE optimization for adaptive video streaming during bottleneck bandwidth sharing. Joint QoE optimization refers to improving QoE fairness among clients with various video devices and availing from differentiated services with different priorities. Flex-Steward deploys an adaptive bitrate delivery algorithm based on Neural Networks (NN) and reinforcement learning at the network edge. It relies on a trained NN model to make appropriate bitrate recommendations in terms of video chunks to be requested by clients sharing the same bottleneck bandwidth. Flex-Steward is assessed in comparison with alternative state-of-the-art algorithms under different network conditions using a real-life prototype. Results show how Flex-Steward reduces the unfairness in terms of joint QoE optimization with between 10.9% and 41.7%.
Xiaoteng Ma, Qing Li 0006, Yong Jiang 0001, Gabriel-Miro Muntean, Longhao Zou
IEEE Trans. Netw. Serv. Manag.5
2021 Higher quality live streaming under lower uplink bandwidth: an approach of super-resolution based video coding
abstract
With the growing popularity of live streaming, high video quality and low latency with limited uplink bandwidth have become a significant challenge. In this study, we propose Live Super-Resolution Based Video Coding (LiveSRVC), a novel video uploading framework that improves the quality of live streaming with low latency under limited uplink bandwidth. We design a new super-resolution-based key frame coding module to improve the coding compression efficiency. LiveSRVC dynamically selects the bitrate and the compression ratio of key frames, mitigating the influence of uplink bandwidth capacity on live streaming quality. Trace-driven emulations verify that LiveSRVC can provide the same quality while reducing up to 50% of the required bandwidth compared to the original encoding method (H.264). LiveSRVC consumes at least 10X less GPU occupation time compared to the method of reconstructing all frames with super-resolution.
Qing Li 0006, Aoyang Zhang, Longhao Zou, Yong Jiang 0001, Zhimin Xu 0001, Zhenhui Yuan
NOSSDAV4
2020 Mulsemedia in Education: A Case Study on Learner Experience, Motivation and Knowledge Gain
Irina Tal, Longhao Zou, Margaret Farren, Gabriel-Miro Muntean
CSEDU (2)2
2020 Content-Aware Cubemap Projection for Panoramic Image via Deep Q-Learning
Xu Wang 0006, Yu Zhou 0027, Longhao Zou, Jianmin Jiang
MMM (2)4
2017 Can Multisensorial Media Improve Learner Experience?
abstract
In recent years, the emerging immersive technologies (e.g. Virtual/Augmented Reality, multisensorial media) bring brand-new multi-dimensional effects such as 3D vision, immersion, vibration, smell, airflow, etc. to gaming, video entertainment and other aspects of human life. This paper reports results from an European Horizon 2020 research project on the impact of multisensoral media (mulsemedia) on educational learner experience. A mulsemedia-enhanced test-bed was developed to perform delivery of video content enhanced with haptic, olfaction and airflow effects. The results of the quality rating and questionnaires show significant improvements in terms of mulsemedia-enhanced teaching.
Longhao Zou, Irina Tal, Alexandra Covaci, Eva Ibarrola, George Ghinea, Gabriel-Miro Muntean
MMSys1
2014 eDOAS: Energy-aware device-oriented adaptive multimedia scheme for Wi-Fi offload
abstract
Mobile devices became an essential part of every person daily routine enabling them to browse the Internet, watch videos, work and play online anytime and anywhere. However this led to a tremendous growth in user generated data traffic putting significant pressure on the underling network technology. Thus, in order to cope with this explosion of data traffic, Wi-Fi offload became a popular solution for network operators. The solution enables the network operators to accommodate more mobile users and keep up with their traffic demands. Moreover, with the energy conservation becoming a critical issue around the world, it provides motivation for this paper to propose an Energy-aware Device-Oriented Adaptive multimedia Scheme (eDOAS) for Wi-Fi Data Offload. eDOAS adapts the interactive multimedia application to the underlying Wi-Fi network conditions, device characteristics and device energy consumption, in order to prolong the battery lifetime of the mobile device and maintain an acceptable user perceived quality level. Real test-bed energy consumption measurements were conducted on five different devices and the performance of eDOAS was analyzed against other schemes from the literature, in terms of energy consumption, service outage, average throughput, packet loss and PSNR.
Longhao Zou, Ramona Trestian, Gabriel-Miro Muntean
WCNC1
2013 DOAS: Device-Oriented Adaptive Multimedia Scheme for 3GPP LTE systems
abstract
The growing popularity of the high-end mobile computing devices - smartphones, tablets, notebooks and more - equipped with high-speed network access, enables the mobile user to watch multimedia content from any source on any screen, at any time, while on the move or stationary. In this context, the network operators must ensure smooth video streaming with the lowest service delay, jitter, and packet loss. This paper proposes a resource efficient Device-Oriented Adaptive Multimedia Scheme (DOAS) built on top of the downlink scheduler in LTE-Advanced systems. DOAS bases its adaptation decision on the end-user device display resolution information and Quality of Service (QoS). DOAS is implemented on top of the Proportional Fair (PF) and the well-known Modified Largest Weighted Delay First (M-LWDF) scheduling algorithms within the 3GPP LTE/LTE-Advanced system. The performance of the proposed adaptive multimedia scheme was analyzed and compared against a non-adaptive solution in terms of throughput, packet loss and PSNR.
Longhao Zou, Ramona Trestian, Gabriel-Miro Muntean
PIMRC1