Hanxin Zhu

dblp:261/8127 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
abstract
Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the corresponding 4D scenes, posing a significant limitation to truly immersive audiovisual experiences. To mitigate this issue, we propose Sonic4D, a novel framework that enables spatial audio generation for immersive exploration of 4D scenes. Specifically, our method is composed of three stages: 1) To capture both the dynamic visual content and raw auditory information from a monocular video, we first employ pre-trained expert models to generate the 4D scene and its corresponding monaural audio. 2) Subsequently, to transform the monaural audio into spatial audio, we localize and track the sound sources within the 4D scene, where their 3D spatial coordinates at different timestamps are estimated via a pixel-level visual grounding strategy. 3) Based on the estimated sound source locations, we further synthesize plausible spatial audio that varies across different viewpoints and timestamps using physics-based simulation. Extensive experiments have demonstrated that our proposed method generates realistic spatial audio consistent with the synthesized 4D scene in a training-free manner, significantly enhancing the immersive experience for users.
Siyi Xie, Hanxin Zhu, Tianyu He, Xin Li 0082, Zhibo Chen 0001
AAAI2
2026 Res-P4DGS:Enhancing 4D Gaussian Splatting Compression with Scene-Depth Prior
Xinliang Gong, Hanxin Zhu, Henan Wang, Xin Li 0082, Zhibo Chen 0001
ISCAS2
2025 TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
abstract
Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and the textual description. (ii) how to improve the subjective quality of generated videos. To tackle the above challenges, we propose a new diffusion-based TI2V framework, termed TIV-Diffusion, via object-centric textual-visual alignment, intending to achieve precise control and high-quality video generation based on textual-described motion for different objects. Concretely, we enable our TIV-Diffuion model to perceive the textual-described objects and their motion trajectory by incorporating the fused textual and visual knowledge through scale-offset modulation. Moreover, to mitigate the problems of object disappearance and misaligned objects and motion, we introduce an object-centric textual-visual alignment module, which reduces the risk of misaligned objects/motion by decoupling the objects in the reference image and aligning textual features with each object individually. Based on the above innovations, our TIV-Diffusion achieves state-of-the-art high-quality video generation compared with existing TI2V methods.
Xingrui Wang, Xin Li 0082, Yaosi Hu, Hanxin Zhu, Chen Hou, Cuiling Lan, Zhibo Chen 0001
AAAI4
2025 MiNL: Micro-Images based Neural Representation for Light Fields
Hanxin Zhu, Henan Wang, Zhibo Chen 0001
ISCAS1
2024 SeD: Semantic-Aware Discriminator for Image Super-Resolution
abstract
Generative Adversarial Networks (GANs) have been widely used to recover vivid textures in image super-resolution (SR) tasks. In particular, one discriminator is utilized to enable the SR network to learn the distribution of real-world high-quality images in an adversarial training manner. However, the distribution learning is overly coarse-grained, which is susceptible to virtual textures and causes counter-intuitive generation results. To mitigate this, we propose the simple and effective Semantic-aware Discriminator (denoted as SeD), which encourages the SR network to learn the fine-grained distributions by introducing the semantics of images as a condition. Concretely, we aim to excavate the semantics of images from a well-trained semantic extractor. Under different semantics, the discriminator is able to distinguish the real-fake images individually and adaptively, which guides the SR network to learn the more fine-grained semantic-aware textures. To obtain accurate and abundant semantics, we take full advantage of recently popular pretrained vision models (PVMs) with extensive datasets, and then incorporate its semantic features into the discriminator through a well-designed spatial cross-attention module. In this way, our proposed semantic-aware discriminator empowered the SR network to produce more photo-realistic and pleasing images. Extensive experiments on two typical tasks, i.e., SR and Real SR have demonstrated the effectiveness of our proposed methods. The code will be available at https://github.com/1bc12345/SeD.
Bingchen Li 0001, Xin Li 0082, Hanxin Zhu, Yeying Jin, Ruoyu Feng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001
CVPR3
2024 Is Vanilla MLP in Neural Radiance Field Enough for Few-Shot View Synthesis?
abstract
Neural Radiance Field (NeRF) has achieved superior performance for novel view synthesis by modeling the scene with a Multi-Layer Perception (MLP) and a volume rendering procedure, however, when fewer known views are given (i.e., few-shot view synthesis), the model is prone to overfit the given views. To handle this issue, previous efforts have been made towards leveraging learned priors or introducing additional regularizations. In contrast, in this paper, we for the first time provide an orthogonal method from the perspective of network structure. Given the observation that trivially reducing the number of model parameters alleviates the overfitting issue, but at the cost of missing details, we propose the multi-input MLP (mi-MLP) that incorpo-rates the inputs (i.e., location and viewing direction) of the vanilla MLP into each layer to prevent the overfitting issue without harming detailed synthesis. To further reduce the artifacts, we propose to model colors and volume density separately and present two regularization terms. Ex-tensive experiments on multiple datasets demonstrate that: 1) although the proposed mi-MLP is easy to implement, it is surprisingly effective as it boosts the PSNR of the base-line from 14.73 to 24.23. 2) the overall framework achieves state-of-the-art results on a wide range of benchmarks.
Hanxin Zhu, Tianyu He, Xin Li 0082, Bingchen Li 0001, Zhibo Chen 0001
CVPR1
2024 UCIP: A Universal Framework for Compressed Image Super-Resolution Using Dynamic Prompt
Xin Li 0082, Bingchen Li 0001, Yeying Jin, Cuiling Lan, Hanxin Zhu, Yulin Ren, Zhibo Chen 0001
ECCV (47)5
2024 End-to-End Rate-Distortion Optimized 3D Gaussian Representation
Henan Wang, Hanxin Zhu, Tianyu He, Runsen Feng, Jiajun Deng, Jiang Bian 0002, Zhibo Chen 0001
ECCV (58)2
2024 Compositional 3D-aware Video Generation with LLM Director
abstract
Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual elements within the generated video, such as the movement and appearance of specific characters and the manipulation of viewpoints. In this work, we propose a novel paradigm that generates each element in 3D representation separately and then composites them with priors from Large Language Models (LLMs) and 2D diffusion models. Specifically, given an input textual query, our scheme consists of four stages: 1) we leverage the LLMs as the director to first decompose the complex query into several sub-queries, where each sub-query describes each element of the generated video; 2) to generate each element, pre-trained models are invoked by the LLMs to obtain the corresponding 3D representation; 3) to composite the generated 3D representations, we prompt multi-modal LLMs to produce coarse guidance on the scale, location, and trajectory of different objects; 4) to make the results adhere to natural distribution, we further leverage 2D diffusion priors and use score distillation sampling to refine the composition. Extensive experiments demonstrate that our method can generate high-fidelity videos from text with flexible control over each element.
Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen 0001, Jiang Bian 0002
NeurIPS1
2024 CMC: Few-shot Novel View Synthesis via Cross-view Multiplane Consistency
abstract
Neural Radiance Field (NeRF) has shown impressive results in novel view synthesis, particularly in Virtual Reality (VR) and Augmented Reality (AR), thanks to its ability to represent scenes continuously. However, when just a few input view images are available, NeRF tends to overfit the given views and thus make the estimated depths of pixels share almost the same value. Unlike previous methods that conduct regularization by introducing complex priors or additional supervisions, we propose a simple yet effective method that explicitly builds depth-aware consistency across input views to tackle this challenge. Our key insight is that by forcing the same spatial points to be sampled repeatedly in different input views, we are able to strengthen the interactions between views and therefore alleviate the overfitting problem. To achieve this, we build the neural networks on layered representations (i.e., multiplane images), and the sampling point can thus be resampled on multiple discrete planes. Furthermore, to regularize the unseen target views, we constrain the rendered colors and depths from different input views to be the same. Although simple, extensive experiments demonstrate that our proposed method can achieve better synthesis quality over state-of-the-art methods.
Hanxin Zhu, Zhibo Chen 0001
VR1
2023 Density-aware Swin Transformer for Compressed Point Cloud Geometry Artifacts Removal
abstract
Geometry-based point cloud compression (G-PCC), as a prevalent compression technique, has achieved remarkable compression efficiency, thereby significantly reducing the cost of transmission and storage. However, the compressed point clouds inevitably suffer from severe compression artifacts, i.e., geometry distortion, when the compression ratio increases. To address this, we propose the DensityFormer, the first transformer-based network to restore the geometry distortion in the compressed point cloud. Particularly, our approach focuses on two prominent challenges for this: i) the discrete points demand more stringent requirements for long-range contextual information modeling and ii) the point cloud exhibits a non-uniformed point distribution. For the first challenge, our DensityFormer introduce the Swin Transformer-based hierarchical encoder-decoder architecture, intending to model the multi-grained global contextual information for geometric restoration, based on the superior long-range dependency modeling capability of 3D Swin Transformer block. To solve the second challenge, we propose the density-aware Swin Transformer block on the basis of the intuition that the local density of the point cloud can identify the distribution of points, thereby enabling the adaptive non-uniformed restoration for compressed point clouds. By incorporating the above two advanced techniques, our DensityFormer has shown superior restoration capability on multiple typical benchmark datasets, which outperforms existing state-of-the-art (SOTA) methods by an average of 0.55 dB.
Xiqian Yu, Xin Li 0082, Hanxin Zhu, Zhibo Chen 0001
VCIP3
2022 Light Field Compression Based on Implicit Neural Representation
abstract
Light field, as a new data representation format in multimedia, has the ability to capture both intensity and direction of light rays. However, the additional angular information also brings a large volume of data. Classical coding methods are not effective to describe the relationship between different views, leading to redundancy left. To address this problem, we propose a novel light field compression scheme based on implicit neural representation to reduce redundancies between views. We store the information of a light field image implicitly in an neural network and adopt model compression methods to further compress the implicit representation. Extensive experiments have demonstrated the effectiveness of our proposed method, which achieves comparable rate-distortion performance as well as superior perceptual quality over traditional methods.
Henan Wang, Hanxin Zhu, Zhibo Chen 0001
PCS2