Zijian Cao 0007

dblp:401/6572 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Geometric modeling and processing · 63% Rendering · 28% Visual content generation and editing · 10%
Artificial intelligence
2 papers
3D vision · 69% Efficient and distributed learning · 10% Representation and self-supervised learning · 10%
Computer networks
2 papers
Content delivery and video streaming · 50% Physical-layer communications · 50%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d shape representation
1.012026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Computer vision › 3D vision
point cloud
1.012026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Geometric modeling and processing
surface parameterization
1.012026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Geometric modeling and processing › surface parameterization
UV mapping
1.012026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Physical-layer communications
semantic communication
1.012026
Live High-Fidelity Semantic Communication via Cross-Modal Fusion for Volumetric Video · IEEE J. Sel. Areas Commun. 2026
Content delivery and video streaming
volumetric video streaming
1.012026
Live High-Fidelity Semantic Communication via Cross-Modal Fusion for Volumetric Video · IEEE J. Sel. Areas Commun. 2026
Rendering › gaussian splatting
3d gaussian splatting
0.912025
SRBF-Gaussian: Streaming-Optimized 3D Gaussian Splatting · VR 2025
Machine learning › Representation and self-supervised learning › computational neuroscience
brain encoding
0.312026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
0.312026
Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping · AAAI 2026
Machine learning › Generative modeling
video generation
0.312026
Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model · NSDI 2026

Methods — techniques the papers use, named apart from their topics

vision transformer · 2.0vision foundation model · 2.0scalable locally injective mapping · 2.0generative model · 2.0dimensional reduction · 2.0ball-pivoting reconstruction · 2.0structure-from-motion · 1.0structure from motion · 1.0spherical radial basis functions · 0.9adaptive gaussian pruning · 0.9
YearPublicationVenuePosition
2026 Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping
abstract
Point-based geometric representations such as point clouds and Gaussian Splatting are fundamental for 3D understanding. However, the inherent irregularity and high-dimensional nature of point structures present significant challenges for direct 3D learning approaches, which often struggle with scalability and achieve suboptimal performance due to sparse data distributions. In contrast, 2D learning paradigms benefit from well-established architectures with superior optimization stability and efficiency. To bridge this gap, we propose Maniflat3D, a unified framework that systematically transforms volumetric point-based geometries into structured 2D representations through a two-stage process: a multilayer Ball-Pivoting reconstruction with adaptive density control, followed by Scalable Locally Injective Mapping (SLIM) to produce distortion-minimized, bijective UV parameterizations. Our approach explicitly encodes both geometric and attribute information into the flattened domain, enabling conventional 2D neural networks to effectively learn from complex 3D structures such as Gaussian Splatting. Experiments on the ShapeSplat dataset demonstrate that Maniflat3D achieves comparable performance while reducing parameter count by 90% compared to native 3D baselines, and simultaneously attains 21× compression ratio through neural encoding. These results establish a new paradigm for efficient geometric understanding, demonstrating successful transfer of planar learning advantages to challenging 3D manifold problems through dimensional reduction.
Zijian Cao 0007, Dayou Zhang, Zeyuan Liu, Zhicheng Liang, Fangxin Wang 0001
AAAI1
2026 Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
Tianyi Gong, Zijian Cao 0007, Zixing Zhang 0009, Jiangkai Wu, Xinggong Zhang, Shuguang Cui, Fangxin Wang 0001
NSDI2
2026 Live High-Fidelity Semantic Communication via Cross-Modal Fusion for Volumetric Video
abstract
Semantic communication (SC) emerges as a breakthrough paradigm for efficient data transmission in next-generation communication networks. However, SC is still in the infant stage with quite a few limitations, such as insufficient semantic representation capacity, high communication latency, and the susceptibility to channel noise. In this paper, we proposeLiFiSC, a cross-modal fusion based generative semantic communication framework with strong semantic compression capacity and high-fidelity semantic restoration. We then extend it toLiFiSCvv, specifically designed to achieve 3D volumetric video transmission and reconstruction with photorealistic visual quality and pixel-level visual consistency, providing end-to-end live watching experience with acceptable latency. We innovatively incorporate unified vision-language encoding into semantic communication, achieving superior semantic understanding and compression.LiFiSCvvcomprises three key components: (1) Information redundancy reduction through lightweight video analysis and structure-from-motion techniques, decreasing reconstruction cost; (2) Cross-modal fusion learning driven codec mechanisms that enable efficient semantic representation and robust transmission against channel impairments; and (3) Feed-forward vision transformer for rapid volumetric video reconstruction and rendering. Comprehensive evaluation results demonstrate thatLiFiSCvvachieves a live high-fidelity and visually consistent video watching experience, with only seconds of end-to-end latency and around 42× semantic compression, significantly outperforming SOTA methods in general image and volumetric video transmission. The project page is available here: https://inml-tygong.github.io/LiFiSCvv/.
Tianyi Gong, Zijian Cao 0007, Zhicheng Liang, Dayou Zhang, Fangxin Wang 0001, Shuguang Cui
IEEE J. Sel. Areas Commun.2
2025 SRBF-Gaussian: Streaming-Optimized 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has emerged as a groundbreaking 3D scene representation technique, offering unprecedented visual quality and rendering efficiency. However, the substantial data volume of 3DGS scenes poses significant challenges for streaming applications. Existing research on 3DGS has primarily focused on compression and rendering efficiency, neglecting the specific requirements of streaming transmission. Moreover, the Spherical Harmonics color representation in 3DGS complicates viewport-based transmission partitioning. Achieving hierarchical Gaussian streaming without noticeable quality degradation also remains a significant challenge.To address these challenges, we propose SRBF-Gaussian, a new paradigm that revolutionizes the traditional 3DGS format. Our approach introduces viewport-dependent color encoding based on Spherical Radial Basis Functions (SRBFs) and HSL color space, enabling selective transmission of viewport-relevant color data. This reduces data transmission while maintaining visual quality. We implement adaptive Gaussian pruning and transmission, optimized for current viewports and network conditions. Additionally, we develop coherent multi-level Gaussian representations for smooth transitions between quality levels. Our system incorporates user-behavior-aware streaming strategies to anticipate and pre-fetch relevant data. In cloud VR scenarios, our approach demonstrates substantial improvements, achieving a 5.63% - 14.17% increase in PSNR, a 7.61% - 59.16% reduction in latency, and a 10.45% - 30.12% improvement in overall Quality of Experience (QoE).
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
VR3
2025 3DGStreaming: Spatial-Heterogeneity-Aware 3-D Gaussian Splatting Compression and Streaming
abstract
3-D Gaussian splatting (3DGS) has emerged as a promising technique for high-quality 3-D scene representation. However, streaming 3DGS scenes poses significant challenges due to large data volumes and complex spatial structures, resulting in nonsmooth scene loading, inferior visual quality, and ineffective streaming adaptation, ultimately impacting user experience adversely. To tackle these challenges and enhance user Quality of Experience (QoE), this article introduces a novel adaptive streaming framework called 3DGStreaming. Our framework comprises three key components: 1) smart Spatial Partitioning for efficient scene division, enabling selective streaming and seamless scene merging; 2) two-step Progressive Scene Generation, involving content-aware downsampling and attribute compression to create multibitrate 3DGS scenes; and 3) Field of View (FoV)-based Bitrate Adaptation using a decision transformer for viewport-based bitrate selection. Extensive experiments demonstrate the superiority of 3DGStreaming over existing state-of-the-art solutions. 3DGStreaming achieves a greater rendering quality with a 5.7%–25.5% increase in PSNR, a 27.8%–69.2% reduction in latency, a 54.7% reduction in training time, and a 17.2%–68.3% improvement in overall QoE.
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
IEEE Internet Things J.3