Zixin Zhu

dblp:298/2123 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
abstract
Textured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and generation. Existing metrics, such as Chamfer Distance, often fail to align with how humans evaluate the fidelity of 3D shapes. Recent learning-based metrics attempt to improve this by relying on rendered images and 2D image quality metrics. However, these approaches face limitations due to incomplete structural coverage and sensitivity to viewpoint choices. Moreover, most methods are trained on synthetic distortions, which differ significantly from real-world distortions, resulting in a domain gap. To address these challenges, we propose a new fidelity evaluation method that is based directly on 3D meshes with texture, without relying on rendering. Our method, named Textured Geometry Evaluation TGE, jointly uses the geometry and color information to calculate the fidelity of the input textured mesh with comparison to a reference colored shape. To train and evaluate our metric, we design a human-annotated dataset with real-world distortions. Experiments show that TGE outperforms rendering-based and geometry-only methods on real-world distortion dataset.
Tianyu Luan, Xuelu Feng, Zixin Zhu, Phani Nuney, Sheng Liu 0017, David S. Doermann, Chunming Qiao, Junsong Yuan 0001
AAAI3
2026 A Hierarchical Deep Reinforcement Learning Model for Joint Optimization of Urban Traffic Congestion and Cost in Intelligent Transportation Systems
abstract
Urban traffic congestion poses a persistent challenge to modern cities, resulting in increased travel delay, energy consumption, and economic loss. As a critical problem in intelligent transportation systems, urban congestion management requires coordinated and adaptive decision-making across both supply-side and demand-side control mechanisms. Existing traffic management approaches usually address congestion from either the supply side through traffic signal control or the demand side via congestion pricing, while their intrinsic interactions and multi-scale coordination remain insufficiently explored. To overcome these challenges, this paper proposes a hierarchical deep reinforcement learning framework, termed JCPO-DRL, for joint congestion and pricing optimization in urban traffic networks. The framework decomposes the joint control task into a two-layer decision process, where an upper-layer policy regulates network-level congestion pricing at a coarse temporal scale and a lower-layer policy performs adaptive traffic signal control conditioned on pricing decisions, enabling coordinated supply–demand regulation while avoiding the curse of dimensionality. Extensive experiments under diverse traffic scenarios show that JCPO-DRL consistently outperforms state-of-the-art signal-only and pricing-only baselines, reducing average travel delay by up to 19% and improving traffic throughput by up to 4% compared with representative baselines, and maintaining moderate congestion pricing cost. These results demonstrate the effectiveness of hierarchical joint optimization and highlight the potential of the proposed framework for intelligent urban traffic management in intelligent transportation systems.
Zixin Zhu
IEEE Internet Things J.1
2025 CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
abstract
In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes. Existing approaches typically train individual adapter for each attribute separately, overlooking the entanglement among multiple attributes. As a result, interference occurs among different attributes, preventing precise control of multiple attributes together. To address this challenge, we aim to disentangle multiple attributes in slider-based generation to enbale more reliable and independent attribute manipulation. Our approach, CompSlider, can generate a conditional prior for the T2I foundation model to control multiple attributes simultaneously. Furthermore, we introduce novel disentanglement and structure losses to compose multiple attribute changes while maintaining structural consistency within the image. Since CompSlider operates in the latent space of the conditional prior and does not require retraining the foundation model, it reduces the computational burden for both training and inference. We evaluate our approach on a variety of image attributes and highlight its generality by extending to video generation.
Zixin Zhu, Kevin Duarte, Mamshad Nayeem Rizve, Ratheesh Kalarot, Junsong Yuan 0001
ICCV1
2025 Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have emerged as a promising approach for energy-efficient and biologically plausible computation. However, due to limitations in existing training methods and inherent model constraints, SNNs often exhibit a performance gap when compared to Artificial Neural Networks (ANNs). Knowledge distillation (KD) has been explored as a technique to transfer knowledge from ANN teacher models to SNN student models to mitigate this gap. Traditional KD methods typically use Kullback-Leibler (KL) divergence to align output distributions. However, conventional KL-based approaches fail to fully exploit the unique characteristics of SNNs, as they tend to overemphasize high-probability predictions while neglecting low-probability ones, leading to suboptimal generalization. To address this, we propose Head-Tail Aware Kullback-Leibler (HTA-KL) divergence, a novel KD method for SNNs. HTA-KL introduces a cumulative probability-based mask to dynamically distinguish between high- and low-probability regions. It assigns adaptive weights to ensure balanced knowledge transfer, enhancing the overall performance. By integrating forward KL (FKL) and reverse KL (RKL) divergence, our method effectively align both head and tail regions of the distribution. We evaluate our methods on CIFAR-10, CIFAR-100 and Tiny ImageNet datasets. Our method outperforms existing methods on most datasets with fewer timesteps.
Tianqing Zhang, Zixin Zhu, Kairong Yu, Hongwei Wang 0001
IJCNN2
2025 Versatile Multimodal Controls for Expressive Talking Human Animation
Ruobing Zheng, Zixin Zhu, Sanping Zhou, Ming Yang 0007, Le Wang 0003
ACM Multimedia5
2025 GeoRemover: Removing Objects and Their Causal Visual Artifacts
abstract
Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image appearance-based methods either follow strictly mask-aligned training and fail to remove these casual effects which are not explicitly masked, or adopt loosely mask-aligned strategies that lack controllability and may unintentionally over-erase other objects. We identify that these limitations stem from ignoring the causal relationship between an object’s geometry presence and its visual effects. To address this limitation, we propose a geometry-aware two-stage framework that decouples object removal into (1) geometry removal and (2) appearance rendering. In the first stage, we remove the object directly from the geometry (e.g., depth) using strictly mask-aligned supervision, enabling structure-aware editing with strong geometric constraints. In the second stage, we render a photorealistic RGB image conditioned on the updated geometry, where causal visual effects are considered implicitly as a result of the modified 3D geometry. To guide learning in the geometry removal stage, we introduce a preference-driven objective based on positive and negative sample pairs, encouraging the model to remove objects as well as their causal visual artifacts while avoiding new structural insertions. Extensive experiments demonstrate that our method achieves state-of-the-art performance in removing both objects and their associated artifacts on two popular benchmarks. The project page is available at https://buxiangzhiren.github.io/GeoRemover.
Zixin Zhu, Xuelu Feng, He Wu, Chunming Qiao, Junsong Yuan 0001
NeurIPS1
2024 Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Zixin Zhu, Xuelu Feng, Dongdong Chen 0001, Junsong Yuan 0001, Chunming Qiao, Gang Hua 0001
ECCV (12)1
2023 ContextLoc++: A Unified Context Model for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching the local, global and multi-scale contexts in the popular two-stage temporal localization framework. Our proposed model, dubbed ContextLoc++, can be divided into three sub-networks: L-Net, G-Net, and M-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. Furthermore, the spatial and temporal snippet-level features, functioning as keys and values, are fused by temporal gating. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. M-Net further fuses the local and global contexts with multi-scale proposal features. Specially, proposal-level features from multi-scale video snippets can focus on different action characteristics. Short-term snippets with fewer frames pay attention to the action details while long-term snippets with more frames focus on the action variations. Experiments on the THUMOS14 and ActivityNet v1.3 datasets validate the efficacy of our method against existing state-of-the-art TAL algorithms.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Learning Disentangled Classification and Localization Representations for Temporal Action Localization
abstract
A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI1
2022 Local to Global Feature Learning for Salient Object Detection
Xuelu Feng, Sanping Zhou, Zixin Zhu, Le Wang 0003, Gang Hua 0001
Pattern Recognit. Lett.3
2021 Enriching Local and Global Contexts for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching both the local and global contexts in the popular two-stage temporal localization framework, where action proposals are first generated followed by action classification and temporal boundary regression. Our proposed model, dubbed ContextLoc, can be divided into three sub-networks: L-Net, G-Net and P-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. P-Net further models the context-aware inter-proposal relations. We explore two existing models to be the P-Net in our experiments. The efficacy of our proposed method is validated by experimental results on the THUMOS14 (54.3% at [email protected]) and ActivityNet v1.3 (56.01% at [email protected]) datasets, which outperforms recent states of the art. Code is available at https://github.com/buxiangzhiren/ContextLoc.
Zixin Zhu, Wei Tang 0016, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001
ICCV1