Dayou Zhang

dblp:323/1294 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping
abstract
Point-based geometric representations such as point clouds and Gaussian Splatting are fundamental for 3D understanding. However, the inherent irregularity and high-dimensional nature of point structures present significant challenges for direct 3D learning approaches, which often struggle with scalability and achieve suboptimal performance due to sparse data distributions. In contrast, 2D learning paradigms benefit from well-established architectures with superior optimization stability and efficiency. To bridge this gap, we propose Maniflat3D, a unified framework that systematically transforms volumetric point-based geometries into structured 2D representations through a two-stage process: a multilayer Ball-Pivoting reconstruction with adaptive density control, followed by Scalable Locally Injective Mapping (SLIM) to produce distortion-minimized, bijective UV parameterizations. Our approach explicitly encodes both geometric and attribute information into the flattened domain, enabling conventional 2D neural networks to effectively learn from complex 3D structures such as Gaussian Splatting. Experiments on the ShapeSplat dataset demonstrate that Maniflat3D achieves comparable performance while reducing parameter count by 90% compared to native 3D baselines, and simultaneously attains 21× compression ratio through neural encoding. These results establish a new paradigm for efficient geometric understanding, demonstrating successful transfer of planar learning advantages to challenging 3D manifold problems through dimensional reduction.
Zijian Cao 0007, Dayou Zhang, Zeyuan Liu, Zhicheng Liang, Fangxin Wang 0001
AAAI2
2026 Delta-STDP: Enabling Hardware-Friendly Supervised Learning in Spiking Neural Networks
abstract
The deployment of supervised learning in Spiking Neural Networks (SNNs) on energy-efficient hardware is hindered by the computational complexity of gradient-based algorithms like SpikeProp, involving intricate temporal derivative calculations and non-parallelizable operations. This paper introduces Delta-STDP, a novel learning rule through hardware-algorithm co-design. We simplify the postsynaptic potential (PSP) waveform from a complex alpha function to a ReLU shape, transforming its temporal derivative into a binary causal mask. Concurrently, we binarize the error signal to its sign, resulting in a weight update rule that factorizes into a global 1-bit error direction and a local spike-timing-dependent amplitude. Information-theoretic analysis shows ReLU optimally shifts information from error to PSP, minimizing loss from error sign binarization, while MNIST results confirm that under sign-based updates, ReLU outperforms alpha and exponential-rise PSPs, reversing their order with full-precision updates. Furthermore, we outline a direct mapping to hardware, where forward propagation simplifies to integrate-and-fire with square-pulse inputs, backpropagation becomes a causally gated memristor crossbar operation, and weight update translates to sign-modulated spike-timing-dependent plasticity (STDP). This work bridges the gap between algorithmic effectiveness and hardware efficiency, providing a pathway towards scalable neuromorphic learning accelerators.
Dayou Zhang, Xiangshui Miao, Yuhui He
ACM Great Lakes Symposium on VLSI1
2026 RAPIDS: Reliable Adaptive Priority-based Intelligent Delivery Streaming for Enhanced Error Resilience in Real-time Video
Dayou Zhang, Zijian Cao, Fangxin Wang 0001
IWQoS1
2026 Compute-in-memory compatible ANN-SNN conversion via ternary nonpolar neuron
Dayou Zhang, Xinyu Wen, Xiangshui Miao, Yuhui He
Neurocomputing2
2026 Live High-Fidelity Semantic Communication via Cross-Modal Fusion for Volumetric Video
abstract
Semantic communication (SC) emerges as a breakthrough paradigm for efficient data transmission in next-generation communication networks. However, SC is still in the infant stage with quite a few limitations, such as insufficient semantic representation capacity, high communication latency, and the susceptibility to channel noise. In this paper, we proposeLiFiSC, a cross-modal fusion based generative semantic communication framework with strong semantic compression capacity and high-fidelity semantic restoration. We then extend it toLiFiSCvv, specifically designed to achieve 3D volumetric video transmission and reconstruction with photorealistic visual quality and pixel-level visual consistency, providing end-to-end live watching experience with acceptable latency. We innovatively incorporate unified vision-language encoding into semantic communication, achieving superior semantic understanding and compression.LiFiSCvvcomprises three key components: (1) Information redundancy reduction through lightweight video analysis and structure-from-motion techniques, decreasing reconstruction cost; (2) Cross-modal fusion learning driven codec mechanisms that enable efficient semantic representation and robust transmission against channel impairments; and (3) Feed-forward vision transformer for rapid volumetric video reconstruction and rendering. Comprehensive evaluation results demonstrate thatLiFiSCvvachieves a live high-fidelity and visually consistent video watching experience, with only seconds of end-to-end latency and around 42× semantic compression, significantly outperforming SOTA methods in general image and volumetric video transmission. The project page is available here: https://inml-tygong.github.io/LiFiSCvv/.
Tianyi Gong, Zijian Cao 0007, Zhicheng Liang, Dayou Zhang, Fangxin Wang 0001, Shuguang Cui
IEEE J. Sel. Areas Commun.4
2026 4DGStream: Variable Bitrate Dynamic Gaussian Splatting Streaming
abstract
While 3D Gaussian Splatting (3DGS) has revolutionized static scene representation, the extension to dynamic scene, i.e., 3DGS video (GSV), faces challenges related to reconstruction quality, rendering speed, and storage requirements. The substantial data volume of current GSV poses significant hurdles for streaming applications, particularly in the realm of AR, VR and MR. To tackle these challenges, we introduce 4DGStream, a novel framework that integrates an efficient GSV compression method, Light4D, and a bitrate adaptation streaming strategy, QoSmooth, to ensure smooth playback while maintaining high visual quality. Light4D employs a binarizationassisted spatiotemporal deformation network to model the deformation of Gaussian primitive attributes over time, while a spatiotemporal-aware masking module prunes trivial Gaussians, further enhancing long-term reconstruction quality. To reduce storage, Light4D uses a binary hash grid to model the entropy of attributes for arithmetic coding, with its binary nature allowing efficient entropy modeling via a Bernoulli distribution. These components enable Light4D to improve the FPS/Storage metric by up to 12.4× over SpacetimeGS and 26.4× over 4DGS on the Neu3D dataset, with performance gains exceeding 3× orders of magnitude compared to other NeRF-based state-of-the-art (SOTA) methods. Here, FPS/Storage reflects the balance between rendering speed and data storage. Despite significant model size reductions, Light4D maintains or surpasses the reconstruction quality of 4DGS. Furthermore, QoSmooth provides effective rate control to enhance playback smoothness, reducing bitrate level switches by 61.6% and increasing time-average utility by 26.2%. All these improvements make 4DGStream highly suited for GSV streaming, improving QoE by 36.7% compared to SOTA methods.
Zhicheng Liang, Dayou Zhang, Linfeng Shen, Miao Zhang 0003, Jian Zhang 0054, Bin Ju, Mallesham Dasari, Fangxin Wang 0001, Jiangchuan Liu
IEEE Trans. Multim.2
2025 SRBF-Gaussian: Streaming-Optimized 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has emerged as a groundbreaking 3D scene representation technique, offering unprecedented visual quality and rendering efficiency. However, the substantial data volume of 3DGS scenes poses significant challenges for streaming applications. Existing research on 3DGS has primarily focused on compression and rendering efficiency, neglecting the specific requirements of streaming transmission. Moreover, the Spherical Harmonics color representation in 3DGS complicates viewport-based transmission partitioning. Achieving hierarchical Gaussian streaming without noticeable quality degradation also remains a significant challenge.To address these challenges, we propose SRBF-Gaussian, a new paradigm that revolutionizes the traditional 3DGS format. Our approach introduces viewport-dependent color encoding based on Spherical Radial Basis Functions (SRBFs) and HSL color space, enabling selective transmission of viewport-relevant color data. This reduces data transmission while maintaining visual quality. We implement adaptive Gaussian pruning and transmission, optimized for current viewports and network conditions. Additionally, we develop coherent multi-level Gaussian representations for smooth transitions between quality levels. Our system incorporates user-behavior-aware streaming strategies to anticipate and pre-fetch relevant data. In cloud VR scenarios, our approach demonstrates substantial improvements, achieving a 5.63% - 14.17% increase in PSNR, a 7.61% - 59.16% reduction in latency, and a 10.45% - 30.12% improvement in overall Quality of Experience (QoE).
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
VR1
2025 3DGStreaming: Spatial-Heterogeneity-Aware 3-D Gaussian Splatting Compression and Streaming
abstract
3-D Gaussian splatting (3DGS) has emerged as a promising technique for high-quality 3-D scene representation. However, streaming 3DGS scenes poses significant challenges due to large data volumes and complex spatial structures, resulting in nonsmooth scene loading, inferior visual quality, and ineffective streaming adaptation, ultimately impacting user experience adversely. To tackle these challenges and enhance user Quality of Experience (QoE), this article introduces a novel adaptive streaming framework called 3DGStreaming. Our framework comprises three key components: 1) smart Spatial Partitioning for efficient scene division, enabling selective streaming and seamless scene merging; 2) two-step Progressive Scene Generation, involving content-aware downsampling and attribute compression to create multibitrate 3DGS scenes; and 3) Field of View (FoV)-based Bitrate Adaptation using a decision transformer for viewport-based bitrate selection. Extensive experiments demonstrate the superiority of 3DGStreaming over existing state-of-the-art solutions. 3DGStreaming achieves a greater rendering quality with a 5.7%–25.5% increase in PSNR, a 27.8%–69.2% reduction in latency, a 54.7% reduction in training time, and a 17.2%–68.3% improvement in overall QoE.
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
IEEE Internet Things J.1
2024 VSAS: Decision Transformer-Based On-Demand Volumetric Video Streaming With Passive Frame Dropping
abstract
Volumetric video is becoming a popular application among various multimedia services, which is envisioned as a fundamental technology for VR, AR, and the emerging metaverse. The commodity RGB-D cameras provide an affordable solution for volumetric video capture, and the VR headset allows immersive and interactive display. From the networking perspective, the primary challenge lies in the smooth and high-quality transmission over the Internet, given the enormous data volume and limited bandwidth. MPEG V-PCC standard has stood out recently as a promising compression and streaming solution that can effectively reduce video size while maintaining high-visual quality. Since MPEG V-PCC is largely backward compatible with the 2-D video compression standard like H.264/AVC, it is natural to use DASH, the most widely used 2-D streaming framework, to stream it. We first propose an integrated framework based on DASH to support MPEG V-PCC Internet streaming. During this, we faced three challenges. First, the lack of a rate-distortion model for MPEG V-PCC encoding parameters. Second, the need for a new bitrate adaptation controller that not only considers the rate of the chunks but also chooses the chunks with proper frame rate. We align the decision transformer to this problem, which expands the success of transformer-based models in natural language processing to the decision problems. Finally, to solve the stalling time issue inherited from DASH, we use a frame-dropping mechanism to eliminate the stalling in DASH playback. Our evaluations show that VSAS achieves an average acrlong QoE improvement of$1.67\times $over a range of network conditions.
Fangxin Wang 0001, Dayou Zhang, Dan Wang 0002
IEEE Internet Things J.4
2024 ILCAS: Imitation Learning-Based Configuration- Adaptive Streaming for Live Video Analytics With Cross-Camera Collaboration
abstract
The high-accuracy and resource-intensive deep neural networks (DNNs) have been widely adopted by live video analytics (VA), where camera videos are streamed over the network to resource-rich edge/cloud servers for DNN inference. Common video encoding configurations (e.g., resolution and frame rate) have been identified with significant impacts on striking the balance between bandwidth consumption and inference accuracy and therefore their adaption scheme has been a focus of optimization. However, previous profiling-based solutions suffer from high profiling cost, while existing deep reinforcement learning (DRL) based solutions may achieve poor performance due to the usage of fixed reward function for training the agent, which fails to craft the application goals in various scenarios. In this paper, we proposeILCAS, the first imitation learning (IL) based configuration-adaptive VA streaming system. Unlike DRL-based solutions,ILCAStrains the agent with demonstrations collected from the expert which is designed as an offline optimal policy that solves the configuration adaption problem through dynamic programming. To tackle the challenge of video content dynamics,ILCASderives motion feature maps based on motion vectors which allowILCASto visually “perceive” video content changes. Moreover,ILCASincorporates a cross-camera collaboration scheme to exploit the spatio-temporal correlations of cameras for more proper configuration selection. Extensive experiments confirm the superiority ofILCAScompared with state-of-the-art solutions, with 2-20.9% improvement of mean accuracy and 19.9–85.3% reduction of chunk upload lag.
Duo Wu, Dayou Zhang, Miao Zhang 0003, Fangxin Wang 0001, Shuguang Cui
IEEE Trans. Mob. Comput.2
2024 TrimStream: Adaptive Realtime Video Streaming Through Intelligent Frame Retrospection in Adverse Network Conditions
abstract
Realtime video streaming (RVS) services are gaining popularity in various applications such as video conferencing, online education, and mixed reality. However, adverse network conditions can significantly damage video transmission, leading to a decline in users' Quality of Experience (QoE). Existing approaches have made considerable efforts to address these problems, including bitrate adaptation, FEC (forward error correction) encoding, and super-resolution techniques. Nevertheless, these methods either focus solely on adjusting transmission configurations (ABR) or consume additional network and computational resources to enhance QoE (FEC or super-resolution), making them suboptimal for adverse network conditions. In this paper, we analyze the limitations of conventional RVS systems when confronted with adverse network conditions and proposeTrimStream, a novel RVS solution based on intelligent frame retrospection, to effectively handle such scenarios. Our approach leverages the high similarity observed between frames in realtime video streaming. The core idea is to store a subset of correctly received frames and exploit frame similarity to minimize transmission while breaking down frame-level dependencies. We formulate the frame caching problem to maximize QoE in RVS and present an online frame cache algorithm. Furthermore, we design a vision-transformer-based, cost-effective frame matching framework that combines different levels of frame information. Our evaluation results demonstrate thatTrimStreamoutperforms state-of-the-art solutions by$14.8\% \sim 21.1\%$improvement in overall QoE.
Dayou Zhang, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.1
2024 DSJA: Distributed Server-Driven Joint Route Scheduling and Streaming Adaptation for Multi-Party Realtime Video Streaming
abstract
The widespread availability of convenient wireless network connection and video capture have fueled the development of multi-party realtime video streaming (MRVS) services, such as Zoom or Microsoft Teams. These services have transformed the generation and distribution of realtime streaming content and offer a new way of online communication, striving to provide high Quality-of-Experience (QoE) for individuals. However, delivering high QoE in MRVS is more challenging than in traditional video scenarios due to the stringent delay requirements and complex multi-party interactive architectures. In this paper, we propose DSJA, a distributed server-driven multi-party realtime video streaming framework that conquers the challenges. We first design an appropriate QoE model for MRVS services to capture the interplay among perceptual quality, variations, bitrate mismatch, loss damage, and streaming delay. We then model the QoE maximization problem in MRVS as a route scheduling and streaming adaptation problem. Afterward, we design DSJA which seamlessly integrates multiple selective forwarding units (SFU) architecture and server-driven approaches based on a two-step solution of route scheduling and streaming adaptation. DSJA first determines the most suitable SFU and streaming routes for each video session based on SFUs' job queuing delay and path latency. Then, the server conducts joint loss and bitrate adaptation decisions to optimize the streaming configuration of all clients, considering network conditions and QoE preferences. Our evaluations show that our framework outperforms state-of-the-art solutions by$23.1\% \sim 41.7\%$from the perspective of QoE, and reduces the backbone network transmission by$14.0\% \sim 36.6\%$.
Dayou Zhang, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.1
2023 SJA: Server-driven Joint Adaptation of Loss and Bitrate for Multi-Party Realtime Video Streaming
abstract
The outbreak of COVID-19 has dramatically promoted the explosive proliferation of multi-party realtime video streaming (MRVS) services, represented by Zoom and Microsoft Teams. Different from Video-on-Demand (VoD) or live streaming, MRVS enables all-to-all realtime video communication, bringing significant challenges to service providing. First, unreliable network transmission can cause network loss, resulting in delay increase and visual quality degradation. Second, the transformation from two-party to multi-party communication makes resource scheduling much more difficult. Moreover, optimizing the overall QoE requires a global coordination, which is quite challenging given the various impact factors such as bitrate and loss.In this paper, we propose the SJA framework, which is, to our best knowledge, the first server-driven joint loss and bitrate adaptation framework in multi-party realtime video streaming services towards maximized QoE. We comprehensively design an appropriate QoE model for MRVS services to capture the interplay among perceptual quality, variations, bitrate mismatch, loss damage, and streaming delay. We mathematically formulate the QoE maximization problem in MRVS services. A Lyapunov-based relaxation and the SJA algorithm are further designed to address the optimization problem with close-to-optimal performance. Evaluations show that our framework can outperform the SOTA solutions by 18.4% ∼ 46.5%.
Dayou Zhang, Zi Zhu, Lei Zhang 0066, Fangxin Wang 0001, Dan Wang 0002
INFOCOM2
2022 HARM: Hardware-Assisted Continuous Re-randomization for Microcontrollers
abstract
Microcontroller-based embedded systems have become ubiquitous with the emergence of IoT technology. Given its critical roles in many applications, its security is becoming increasingly important. Unfortunately, MCU devices are especially vulnerable. Code reuse attacks are particularly noteworthy since the memory address of firmware code is static. This work seeks to combat code reuse attacks, including ROP and more advanced JIT-ROP via continuous randomization. Previous proposals are geared towards full-fledged OSs with rich runtime environments, and therefore cannot be applied to MCUs. We propose the first solution for ARM-based MCUs. Our system, named HARM, comprises a secure runtime and a binary analysis tool with rewriting module. The secure runtime, protected inside the secure world, proactively triggers and performs non-bypassable randomization to the firmware running in a sandbox in the normal world. Our system does not rely on any firmware feature, and therefore is generally applicable to both bare-metal and RTOS-powered firmware. We have implemented a prototype on a development board. Our evaluation results indicate that HARM can effectively thaw code reuse attacks while keeping the performance and energy overhead low.
Jiameng Shi, Le Guan, Dayou Zhang, Ping Chen 0003, Ning Zhang 0017
EuroS&P4
2022 Towards Joint Loss and Bitrate Adaptation in Realtime Video Streaming
abstract
Recent years have seen booming development of realtime streaming services, highly improving user experience in remote work, online education, and entertainment. Unlike video-on-demand (VoD) or live services, realtime streaming service has extremely stringent delay requirements, rendering the TCP-based transmission no longer applicable. Existing works based on UDP (or its variants) either suffer from the packet loss problem or only focus on improving several QoS metrics, which cannot achieve satisfactory user QoE. Our insight is to slightly sacrifice the bitrate and video quality to trade for the most significant delay to maximize the overall QoE. We propose Oppugno‡‡Oppugno is a spell in Harry Potter that makes magical creatures attack the caster. It is a metaphor that we use an additional mechanism to mitigate the influence of packet loss., an integrated framework that achieves joint loss adaptation and bitrate adaption towards maximized QoE in realtime streaming services. Oppugno leverages existing UDP mechanisms and employs an advanced deep reinforcement learning algorithm Proximal Policy Optimization (PPO), to adaptively select optimal actions based on network conditions. Trace-driven experiments demonstrate the superiority of our framework, which outperforms the SOTA work by 3.9% ∼ 11.6%.
Dayou Zhang, Fangxin Wang 0001, Dan Wang 0002, Jiangchuan Liu
ICME1