Zihan Zhu

dblp:299/0630 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ProcTex: Consistent and Interactive Text-to-texture Synthesis for Part-based Procedural Models
abstract
Abstract Recent advances in generative modeling have driven significant progress in text‐guided texture synthesis. However, current methods focus on synthesizing texture for single static 3D object, and struggle to handle entire families of shapes, such as those produced by procedural programs. Applying existing methods naively to each procedural shape is too slow to support exploring different parameter configurations at interactive rates, and also results in inconsistent textures across the procedural shapes. To this end, we introduce ProcTex, the first text‐to‐texture system designed for part‐based procedural models. ProcTex enables consistent and real‐time text‐guided texture synthesis for families of shapes, which integrates seamlessly with the interactive design flow of procedural modeling. To ensure consistency, our core approach is to synthesize texture for a template shape from the procedural model, followed by a texture transfer stage to apply the texture to other procedural shapes via solving dense correspondence. To ensure interactiveness, we propose a novel correspondence network and show that dense correspondence can be effectively learned by a neural network for procedural models. We also develop several techniques, including a retexturing pipeline to support structural variation from procedural parameters, and part‐level UV texture map generation for local appearance editing. Extensive experiments on a diverse set of procedural models validate ProcTex's ability to produce high‐quality, visually consistent textures while supporting interactive applications. Code and data are available at: https://github.com/ruiqixu37/ProcTex.git
Zihan Zhu, Benjamin Ahlbrand, Srinath Sridhar 0002, Daniel Ritchie 0001
Comput. Graph. Forum2
2026 HACGNet: Hierarchical attention-driven cross-task alignment network for foggy vehicle re-identification
Zihan Zhu, Pengyang Wang, Haolin Xiang, Nuobing Li, Zhenfeng Zhao
Neurocomputing3
2026 Adaptive reinforcement learning based projected gradient descent attack
Zihan Zhu, Yuexin Zhang, Ayong Ye, Xiaoding Wang 0001, Chengling Wang, Tianqing Zhu
J. Supercomput.1
2026 A single image dehazing based on cross-scale feature aggregation and structural information guidance
Zhenfeng Zhao, Xinlong Yu, Zihan Zhu, Wenbang Fan, Quanli Zhao, Zihang Wu, Kun Meng
Vis. Comput.4
2025 WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
abstract
We present WildGS-SLAM, a robust and efficient monocular RGB SLAM system designed to handle dynamic environments by leveraging uncertainty-aware geometric mapping. Unlike traditional SLAM systems, which assume static scenes, our approach integrates depth and uncertainty information to enhance tracking, mapping, and rendering performance in the presence of moving objects. We introduce an uncertainty map, predicted by a shallow multi-layer perceptron and DI-NOv2 features, to guide dynamic object removal during both tracking and mapping. This uncertainty map enhances dense bundle adjustment and Gaussian map optimization, improving reconstruction accuracy. Our system is evaluated on multiple datasets and demonstrates artifact-free view synthesis. Results showcase WildGS-SLAM’s superior performance in dynamic environments compared to state-of-the-art methods.
Jianhao Zheng, Zihan Zhu, Valentin Bieri, Marc Pollefeys, Songyou Peng, Iro Armeni
CVPR2
2025 AlphaSparseTensor: Discovering Faster Sparse Matrix Multiplication Algorithms on GPUs for LLM Inference
Xuanzheng Wang, Shuo Miao, Zihan Zhu, Youhui Zhang
Euro-Par (2)3
2025 Decoding Rewards in Competitive Games: Inverse Game Theory with Entropy Regularization
abstract
Estimating the unknown reward functions driving agents' behavior is a central challenge in inverse games and reinforcement learning. This paper introduces a unified framework for reward function recovery in two-player zero-sum matrix games and Markov games with entropy regularization. Given observed player strategies and actions, we aim to reconstruct the underlying reward functions. This task is challenging due to the inherent ambiguity of inverse problems, the non-uniqueness of feasible rewards, and limited observational data coverage. To address these challenges, we establish reward function identifiability using the quantal response equilibrium (QRE) under linear assumptions. Building on this theoretical foundation, we propose an algorithm to learn reward from observed actions, designed to capture all plausible reward parameters by constructing confidence sets. Our algorithm works in both static and dynamic settings and is adaptable to incorporate other methods, such as Maximum Likelihood Estimation (MLE). We provide strong theoretical guarantees for the reliability and sample-efficiency of our algorithm. Empirical results demonstrate the framework’s effectiveness in accurately recovering reward functions across various scenarios, offering new insights into decision-making in competitive environments.
Junyi Liao, Zihan Zhu, Ethan X. Fang, Zhuoran Yang, Vahid Tarokh
ICML2
2025 Spatio-Temporal Correlated Network State Prediction and Dynamic Routing for Satellite Networks
abstract
Most existing routing algorithms for Low-Earth-Orbit (LEO) satellite networks neglect the spatio-temporal features inherent in the network states, leading to suboptimal performance in scenarios with dynamic network topologies. In this paper, we propose a Spatio-Temporal Graph Attention Network (STGAN) architecture for the extraction of spatio-temporal features from satellite network states. Building upon this, we propose a State Prediction based Dynamic Routing Algorithm (SP-DRA), which combines STGAN with the enhanced Shortest Path First algorithm (eSPF). SP-DRA utilizes STGAN to predict future network states by exploiting the spatio-temporal correlations of historic observations. Based on state predictions, each link is assigned with a delay-based prediction weight, which is used as the input for eSPF. The proposed eSPF dynamically selects the path with the smallest prediction weight to facilitate congestion avoidance and reduce transmission delay. Simulation results show that our STGAN architecture provides up to 18.6% prediction accuracy improvement than the existing deep learning-based methods, namely STGCN and ST-MGAT. Meanwhile, the SP-DRA algorithm outperforms existing routing strategies including SPF, Explicit load balancing algorithm (ELB), and Satellite networks Link State Routing algorithm (SLSR) in terms of packet loss rate and average end-to-end delay. Specifically, the performance gains increase with network traffic.
Zihan Zhu, Ke Wu 0012, Yunpeng Hou, Huasen He, Jian Yang 0014
WCNC2
2025 A Novel Traffic Prediction Method for Dynamic Satellite Networks Based on Graph Attention Networks
abstract
Satellite networks have been proposed as a vital component in 6G networks for providing global connections. In recent years, satellite network traffic has been increasing. However, the limited onboard resources and inter-satellite link bandwidth make network congestion a critical issue. In order to avoid network congestion and improve quality of service, satellite traffic prediction has received more attention. However, the traditional forecasting models do not fully consider the dynamic topology of satellite networks and the spatial-temporal features of traffic. Thus, we propose a novel traffic prediction method for dynamic satellite networks. Considering that there is traffic correlation between nodes without direct connection, our model introduces a graph generation module to generate adjacency matrices based on dynamic attributes. Moreover, to improve the accuracy of prediction, we further propose the spatial attention module and time sequence processing module to exploit the temporal and spatial correlations of satellite traffic respectively. The experimental results show that our method outperforms the compared algorithms and increases up to 6.84 % prediction accuracy.
Zihan Zhu, Ke Wu 0012, Yunpeng Hou, Huasen He, Jian Yang 0014
WCNC1
2025 Hybrid Mesh-Neural Representation for 3D Transparent Object Reconstruction
abstract
In this study, we propose a novel method to reconstruct the 3D shapes of transparent objects using images captured by handheld cameras under natural lighting conditions. It combines the advantages of an explicit mesh and multi-layer perceptron (MLP) network as a hybrid representation to simplify the capture settings used in recent studies. After obtaining an initial shape through multi-view silhouettes, we introduced surface-based local MLPs to encode the vertex displacement field (VDF) for reconstructing surface details. The design of local MLPs allowed representation of the VDF in a piecewise manner using two-layer MLP networks to support the optimization algorithm. Defining local MLPs on the surface instead of on the volume also reduced the search space. Such a hybrid representation enabled us to relax the ray-pixel correspondences that represent the light path constraint to our designed ray-cell correspondences, which significantly simplified the implementation of a single-image-based environment-matting algorithm. We evaluated our representation and reconstruction algorithm on several transparent objects based on ground truth models. The experimental results show that our method produces high-quality reconstructions that are superior to those of state-of-the-art methods using a simplified data-acquisition setup.
Jiamin Xu, Zihan Zhu, Hujun Bao, Weiwei Xu 0003
Comput. Vis. Media2
2024 NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM
abstract
Neural implicit representations have recently become popular in simultaneous localization and mapping (SLAM), especially in dense visual SLAM. However, existing works either rely on RGB-D sensors or require a separate monocular SLAM approach for camera tracking, and fail to produce high-fidelity 3D dense reconstructions. To address these shortcomings, we present NICER-SLAM, a dense RGB SLAM system that simultaneously optimizes for camera poses and a hierarchical neural implicit map representation, which also allows for high-quality novel view synthesis. To facilitate the optimization process for mapping, we integrate additional supervision signals including easy-to-obtain monocular geometric cues and optical flow, and also introduce a simple warping loss to further enforce geometric consistency. Moreover, to further boost performance in complex large-scale scenes, we also propose a local adaptive transformation from signed distance functions (SDFs) to density in the volume rendering equation. On multiple challenging indoor and outdoor datasets, NICER-SLAM demonstrates strong performance in dense mapping, novel view synthesis, and tracking, even competitive with recent RGB-D SLAM systems. Project page: https://nicer-slam.github.io/.
Zihan Zhu, Songyou Peng, Viktor Larsson, Zhaopeng Cui, Martin R. Oswald, Andreas Geiger 0001, Marc Pollefeys
3DV1
2024 NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild
abstract
Neural Radiance Fields (NeRFs) have shown remarkable success in synthesizing photorealistic views from multi-view images of static scenes, but face challenges in dynamic, real-world environments with distractors like moving objects, shadows, and lighting changes. Existing methods manage controlled environments and low occlusion ra-tios but fall short in render quality, especially under high occlusion scenarios. In this paper, we introduce NeRF On-the-go, a simple yet effective approach that enables the ro-bust synthesis of novel views in complex, in-the-wild scenes from only casually captured image sequences. Delving into uncertainty, our method not only efficiently eliminates dis-tractors, even when they are predominant in captures, but also achieves a notably faster convergence speed. Through comprehensive experiments on various scenes, our method demonstrates a significant improvement over state-of-the-art techniques. This advancement opens new avenues for NeRF in diverse and dynamic real-world applications.
Weining Ren, Zihan Zhu, Marc Pollefeys, Songyou Peng
CVPR2
2023 Online Performative Gradient Descent for Learning Nash Equilibria in Decision-Dependent Games
abstract
We study the multi-agent game within the innovative framework of decision-dependent games, which establishes a feedback mechanism that population data reacts to agents’ actions and further characterizes the strategic interactions between agents. We focus on finding the Nash equilibrium of decision-dependent games in the bandit feedback setting. However, since agents are strategically coupled, traditional gradient-based methods are infeasible without the gradient oracle. To overcome this challenge, we model the strategic interactions by a general parametric model and propose a novel online algorithm, Online Performative Gradient Descent (OPGD), which leverages the ideas of online stochastic approximation and projected gradient descent to learn the Nash equilibrium in the context of function approximation for the unknown gradient. In particular, under mild assumptions on the function classes defined in the parametric model, we prove that OPGD can find the Nash equilibrium efficiently for strongly monotone decision-dependent games. Synthetic numerical experiments validate our theory.
Zihan Zhu, Ethan X. Fang, Zhuoran Yang
NeurIPS1
2022 NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
abstract
Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over- smoothed scene reconstructions and have difficulty scaling up to large scenes. These limitations are mainly due to their simple fully-connected network architecture that does not incorporate local information in the observations. In this paper, we present NICE-SLAM, a dense SLAM system that incorporates multi-level local information by introducing a hierarchical scene representation. Optimizing this representation with pre-trained geometric priors enables detailed reconstruction on large indoor scenes. Compared to recent neural implicit SLAM systems, our approach is more scalable, efficient, and robust. Experiments on five challenging datasets demonstrate competitive results of NICE-SLAM in both mapping and tracking quality. Project page: https://pengsongyou.github.io/nice-slam.
Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu 0003, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, Marc Pollefeys
CVPR1
2022 Scalable neural indoor scene rendering
abstract
We propose a scalable neural scene reconstruction and rendering method to support distributed training and interactive rendering of large indoor scenes. Our representation is based on tiles. Tile appearances are trained in parallel through a background sampling strategy that augments each tile with distant scene information via a proxy global mesh. Each tile has two low-capacity MLPs: one for view-independent appearance (diffuse color and shading) and one for view-dependent appearance (specular highlights, reflections). We leverage the phenomena that complex view-dependent scene reflections can be attributed to virtual lights underneath surfaces at the total ray distance to the source. This lets us handle sparse samplings of the input scene where reflection highlights do not always appear consistently in input images. We show interactive free-viewpoint rendering results from five scenes, one of which covers an area of more than 100 m 2 . Experimental results show that our method produces higher-quality renderings than a single large-capacity MLP and five recent neural proxy-geometry and voxel-based baseline methods. Our code and data are available at project webpage https://xchaowu.github.io/papers/scalable-nisr.
Xiuchao Wu, Jiamin Xu, Zihan Zhu, Hujun Bao, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003
ACM Trans. Graph.3
2021 Scalable image-based indoor scene rendering with reflections
abstract
This paper proposes a novel scalable image-based rendering (IBR) pipeline for indoor scenes with reflections. We make substantial progress towards three sub-problems in IBR, namely, depth and reflection reconstruction, view selection for temporally coherent view-warping, and smooth rendering refinements. First, we introduce a global-mesh-guided alternating optimization algorithm that robustly extracts a two-layer geometric representation. The front and back layers encode the RGB-D reconstruction and the reflection reconstruction, respectively. This representation minimizes the image composition error under novel views, enabling accurate renderings of reflections. Second, we introduce a novel approach to select adjacent views and compute blending weights for smooth and temporal coherent renderings. The third contribution is a supersampling network with a motion vector rectification module that refines the rendering results to improve the final output's temporal coherence. These three contributions together lead to a novel system that produces highly realistic rendering results with various reflections. The rendering quality outperforms state-of-the-art IBR or neural rendering algorithms considerably.
Jiamin Xu, Xiuchao Wu, Zihan Zhu, Qixing Huang, Yin Yang 0002, Hujun Bao, Weiwei Xu 0003
ACM Trans. Graph.3