Mufan Liu

dblp:268/8895 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Light4GS: Lightweight Compact 4D Gaussian Splatting Generation via Context Model
Mufan Liu, Qi Yang 0003, Zhenlong Yuan, Zhu Li 0001, Yiling Xu, Yunfeng Guan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 On the Efficient Adaptive Streaming of 3D Gaussian Splatting Over Dynamic Networks
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising representation for immersive media. Its explicit splat-based structure offers high visual quality and real-time rendering, making it particularly suitable for six degrees of freedom streaming applications. However, its deployment in practical streaming scenarios is still limited due to several key challenges such as the large data volume, and insufficient support for dynamic bitrate adaptation under fluctuating network conditions. This paper presents an efficient 3DGS streaming framework that operates directly on pre-generated 3DGS models without retraining or fine-tuning. First, a training-free perceptual pruning method, which removes visually redundant Gaussians according to the human visual system metrics, is introduced. The resulting 3DGS is then encoded into a compact representation using the extended 3D codecs, exploiting its point-based structure. Next, we build a scene-specific bitrate ladder through analyzing the trade-off between resolution, bitrate, and perceptual quality. This enables efficient and fine-grained representation selection. Finally, a progressive streaming mechanism is developed. It is driven by a reinforcement learning scheduler that adaptively decides whether to download new content or enhance previously buffered content based on real-time network feedback. Experiments on real-world 3DGS datasets and bandwidth traces show that the proposed method evidently improves the quality of experience and streaming efficiency in various network scenarios.
Mufan Liu, Qi Yang 0003, Le Yang 0001, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Anchor-Driven Compact Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods typically rely on per-Gaussian deformation from a canonical space to target frames, which overlooks the strong redundancy among spatially and temporally adjacent Gaussian primitives and leads to suboptimal efficiency. To address this limitation, we propose ADC-GS++, an anchor-driven compact Gaussian splatting framework for efficient and high-quality dynamic scene reconstruction. Specifically, ADC-GS++ organizes Gaussian primitives into an anchor-based canonical representation, enabling attribute sharing across local regions. To efficiently model dynamic scenes, we introduce a static-dynamic decomposition mechanism and further employ a coarse-to-fine deformation strategy driven by dynamic anchors at multiple granularities. In addition, a unified rate-distortion optimization is adopted to achieve a balanced trade-off between storage efficiency and reconstruction fidelity. Furthermore, a temporal significance-based anchor refinement strategy is employed to dynamically grow and prune anchors, allowing robust adaptation to complex and large-scale motions. Extensive experiments on multiple real-world dynamic scene datasets demonstrate that ADC-GS++ significantly improves rendering speed over deformation-based approaches by 300%-700%, while maintaining competitive rendering quality. Moreover, ADC-GS++ achieves a more favorable rate-distortion trade-off, resulting in substantially reduced storage consumption across different bitrate settings.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IEEE Trans. Vis. Comput. Graph.3
2026 Deformable 2D Gaussian Splatting for Efficient Wireless Radiance Field Rendering
abstract
Modeling the wireless radiance field (WRF) is fundamental to modern communication systems, enabling key tasks such as localization, sensing, and channel estimation. Traditional approaches, which rely on empirical formulas or physical simulations, often suffer from limited accuracy or require strong scene priors. Recent neural radiance field (NeRF)-based methods improve reconstruction fidelity through differentiable volumetric rendering, but their reliance on computationally expensive multilayer perceptron (MLP) queries hinders real-time deployment. To overcome these challenges, we introduce Gaussian splatting (GS) to the wireless domain, leveraging its efficiency in modeling optical radiance fields to enable compact and accurate WRF reconstruction. Specifically, we propose SwiftWRF, a deformable 2D Gaussian splatting framework that synthesizes WRF spectra at arbitrary positions under single-sided transceiver mobility. SwiftWRF employs CUDA-accelerated rasterization to render spectra at over 100 k FPS and uses the lightweight MLP to model the deformation of 2D Gaussians, effectively capturing mobility-induced WRF variations. In addition to novel spectrum synthesis, the efficacy of SwiftWRF is further underscored in its applications in angle-of-arrival (AoA) and received signal strength indicator (RSSI) prediction. Experiments conducted on both real-world and synthetic indoor scenes demonstrate that SwiftWRF can reconstruct WRF spectra up to 500x faster than existing state-of-the-art methods, while significantly enhancing its signal quality.
Mufan Liu, Cixiao Zhang, Qi Yang 0003, Yiling Xu, Yin Xu 0001, Shu Sun 0001, Mingzeng Dai, Yunfeng Guan 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Neural Adaptive Contextual Video Streaming
abstract
Video streaming services typically employ traditional codecs, such as H.264, to encode videos into multiple bitrate representations. These codecs are tightly limited by discrete quantization parameters (QPs), resulting in encoded rates that do not align with the target bitrate. Additionally, the subpar video quality produced by conventional codecs does not meet the demands of high-resolution communication. Considering the limitations of traditional codecs, we take a fresh new approach to video streaming by leveraging advanced deep learning-based video codecs. Specifically, we develop a neural adaptive contextual video streaming framework that incorporates: 1) an ensemble deep reinforcement learning based adaptive bitrate algorithm named TSAC that enables continuous bitrate adjustment to varying network conditions 2) a two-stage proportional-integral-derivative-based rate control module that dynamically fine-tunes QPs to ensure the encoded bitrate aligning with the target bitrate. Furthermore, we implement intra-GoP and inter-GoP techniques to accelerate the inference process of the contextual video codec for real-time processing needs. Our experiments demonstrate that the average relative error in bitrate remains below 2%, the quality of experience provided by our TSAC agents surpasses that of existing discrete algorithms by 13%-20%. Our optimization techniques enable real-time decoding at approximately 24 frames per second for quad high definition videos.
Jianchao Yang, Mufan Liu, Puyue Hou, Yiling Xu, Jun Sun 0005
ICASSP2
2025 Deep Joint Source-Channel Coding for Wireless Point Cloud Transmission
abstract
The growing demand for high-quality point cloud transmission over wireless networks presents significant challenges, primarily due to the large data sizes and the need for efficient encoding techniques. In response to these challenges, we introduce a novel system named Deep Point Cloud Semantic Transmission (PCST), designed for end-to-end wireless point cloud transmission. Our approach employs a progressive resampling framework using sparse convolution to project point cloud data into a semantic latent space. These semantic features are subsequently encoded through a deep joint source-channel (JSCC) encoder, generating the channel-input sequence. To enhance transmission efficiency, we use an adaptive entropy-based approach to assess the importance of each semantic feature, allowing transmission lengths to vary according to their predicted entropy. PCST is robust across diverse Signal-to-Noise Ratio (SNR) levels and supports an adjustable rate-distortion (RD) trade-off, ensuring flexible and efficient transmission. Experimental results indicate that PCST significantly outperforms traditional separate source-channel coding (SSCC) schemes, delivering superior reconstruction quality while achieving over a 50% reduction in bandwidth usage.
Cixiao Zhang, Mufan Liu, Yin Xu 0001, Yiling Xu, Dazhi He
ICASSP2
2025 ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods rely on per-Gaussian deformation from a canonical space to target frames, which overlooks redundancy among adjacent Gaussian primitives and result in suboptimal performance. To address this limitation, we propose Anchor-Driven Deformable and Compressed Gaussian Splatting (ADC-GS), a compact and efficient representation for dynamic scene reconstruction. Specifically, ADC-GS organizes Gaussian primitives into an anchor-based structure within the canonical space, enhanced by a temporal significance-based anchor refinement strategy. To reduce deformation redundancy, ADC-GS introduces a hierarchical coarse-to-fine pipeline that captures motions at varying granularities. Moreover, a rate-distortion optimization is adopted to achieve an optimal balance between bitrate consumption and representation fidelity. Experimental results demonstrate that ADC-GS outperforms the per-Gaussian deformation approaches in rendering speed by 300%-800% while achieving state-of-the-art storage efficiency without compromising rendering quality. The code is released at https://github.com/H-Huang774/ADC-GS.git.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IJCAI3
2025 Implicit Retinex Decomposition with Chromaticity Disentanglement for Low-Light Image Enhancement
Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma
ACM Multimedia1
2025 Video Streaming with Kairos: An MPC-Based ABR with Streaming-Aware Throughput Prediction
abstract
Throughput prediction in current adaptive bitrate (ABR) schemes often neglects streaming-aware characteristics, such as sequence irregularity and prediction smoothness, resulting in inaccurate predictions and suboptimal performance. To address these challenges, we propose Kairos, an MPC-based ABR scheme that integrates an attention-based throughput predictor with buffer-aware uncertainty control to enhancing both prediction accuracy and adaptability to dynamic network conditions. Specifically, Kairos employs a multi-time attention network (mTAN) to process irregularly sampled streaming data, producing uniformly spaced latent representations. Based on these, we introduce a percentile prediction network to estimate future throughput percentiles, along with a buffer-aware uncertainty control module that selects the optimal percentile based on the current buffer status. As smoothness is another key component of QoE, we incorporate a smoothness regularizer to ensure consistent throughput predictions, thereby facilitating smoother ABR decisions. Our Kairos design integrates sampling irregularity, prediction uncertainty, and smoothness into the throughput prediction, significantly enhancing bitrate decision making within the MPC framework. Extensive trace-driven and real-world experiments demonstrate that Kairos outperforms state-of-the-art ABR schemes, achieving a QoE improvement ranging from 6.42% to 29.45% across diverse network conditions.
Ziyu Zhong, Mufan Liu, Le Yang 0001, Yiling Xu, Jenq-Neng Hwang
NOSSDAV2
2025 SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo With Depth Restoration and Occlusion Constraint
abstract
Recently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas. However, existing approaches neglect to address the problem of deformation instability caused by easily overlooked edge-skipping, potentially leading to matching distortions, thus leaving room for further improvement. To fill this gap, we propose SED-MVS, which adopts panoptic segmentation and multi-trajectory diffusion strategy for segmentation-driven and edge-aligned patch deformation. Specifically, to prevent unanticipated edge-skipping, we first employ SAM2 for panoptic segmentation as depth-edge guidance to guide patch deformation, followed by multi-trajectory diffusion strategy to ensure patches are comprehensively aligned with depth edges. Moreover, to avoid potential inaccuracy of random initialization, we combine both sparse points from LoFTR and monocular depth map from DepthAnything V2 to restore reliable and realistic depth map for initialization and supervised guidance. Finally, we integrate the segmentation image with the monocular depth map to exploit inter-instance occlusion relationship, then further regard them as occlusion map to implement two distinct edge constraint, thereby facilitating occlusion-aware patch deformation. Extensive results on ETH3D, Tanks & Temples, BlendedMVS, Strecha and DL3DV-10K datasets validate the state-of-the-art performance and robust generalization capability of our proposed method.
Zhenlong Yuan, Zhidong Yang, Yujun Cai, Kuangxin Wu, Mufan Liu, Hao Jiang 0013, Zhaoxin Li
IEEE Trans. Circuits Syst. Video Technol.5
2024 Mesla: Neural Adaptive Layered Point Cloud Streaming with Enhanced Buffer Management
abstract
Point cloud video provides an immersive experience with six degrees of freedom (6DoF). However, streaming video imposes significant challenges due to the huge data volume involved, which increases the transmission burden and makes it difficult to maintain a high quality of experience (QoE) in bandwidth-constrained networks. Current methods either employ pre-specified control rules or make irreversible decisions to optimize the QoE, hindering adaptation to the varying network conditions and not employing buffer management. In this work, we present Mesla1, an adaptive layered point cloud streaming framework, which is empowered by deep reinforcement learning (DRL) and hierarchical buffer management. By maximizing the utility-driven QoE via proximal policy optimization (PPO)-based DRL, Mesla makes bitrate decisions based on the obtained network observations and viewing behaviors of users. The viewport-oriented utility is incorporated into the QoE objective for assigning point cloud tiles with different priorities. To achieve fine-grained rate adaptation, we partition point clouds into layers with different levels of details (LoD) and develop a buffering heuristic (i.e., the hierarchical buffer management) that enhances the quality of unconsumed content. Extensive simulations demonstrate the advantages of Mesla compared to the benchmarking schemes, corroborating the effectiveness of the proposed system.
Mufan Liu, Puyue Hou, Le Yang 0001, Yiling Xu
GLOBECOM2
2024 EVAN: Evolutional Video Streaming Adaptation via Neural Representation
abstract
Adaptive bitrate (ABR) using conventional codecs cannot further modify the bitrate once a decision has been made, exhibiting limited adaptation capability. This may result in either overly conservative or overly aggressive bitrate selection, which could cause either inefficient utilization of the network bandwidth or frequent re-buffering, respectively. Neural representation for video (NeRV), which embeds the video content into neural network weights, allows video reconstruction with incomplete models. Specifically, the recovery of one frame can be achieved without relying on the decoding of adjacent frames. NeRV has the potential to provide high video reconstruction quality and, more importantly, pave the way for developing more flexible ABR strategies for video transmission. In this work, a new framework, named Evolutional Video streaming Adaptation via Neural representation (EVAN), which can adaptively transmit NeRV models based on soft actor-critic (SAC) reinforcement learning, is proposed. EVAN is trained with a more exploitative strategy and utilizes progressive playback to avoid re-buffering. Experiments showed that EVAN can outperform existing ABRs with 50% reduction in re-buffering and achieve nearly 20% improvement in users’ quality of experience (QoE).
Mufan Liu, Le Yang 0001, Yiling Xu, Ye-Kui Wang, Jenq-Neng Hwang
ICME1
2023 Soft-Ack based Outer Loop Link Adaptation for Latency-constrained 5G Video Conferencing
abstract
The high reliability and low latency requirements of multimedia services necessitate the design of more efficient link adaptation methods. In this paper, we introduce instantaneous channel state information (CSI) reporting, specifically designed for 5G video conferencing, and enhance the outer loop link adaptation based on soft Acknowledgement (Soft-ACK). We also formulate a resource allocation problem in 5G physical downlink shared channel (PDSCH) to balance the uplink and downlink traffic in compliance with the specified latency constraints. Our proposed scheme operates in a relatively straightforward manner. It outperforms conventional link adaptation methods re-garding Block-Level Error Rate (BLER) and effectively adheres to stringent latency constraints in video transmission simulations.
Mufan Liu, Jie Chen 0088, Gang Wu 0001, Lei Ji 0003, Hao Wang 0179
GLOBECOM1