Ziyang Ren

dblp:353/7505 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0008-8159-4096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
abstract
3D occupancy and scene flow offer a detailed and dynamic representation of 3D scene. Recognizing the sparsity and complexity of 3D space, previous vision-centric methods have employed implicit learning-based approaches to model spatial and temporal information. However, these approaches struggle to capture local details and diminish the model’s spatial discriminative ability. To address these challenges, we propose a novel explicit state-based modeling method designed to leverage the occupied state to renovate the 3D features. Specifically, we propose a sparse occlusion-aware attention mechanism, integrated with a cascade refinement strategy, which accurately renovates 3D features with the guidance of occupied state information. Additionally, we introduce a novel method for modeling long-term dynamic interactions, which reduces computational costs and preserves spatial information. Compared to the previous state-of-the-art methods, our efficient explicit renovation strategy not only delivers superior performance in terms of RayloU and mAVE for occupancy and scene flow prediction but also markedly reduces GPU memory usage during training, bringing it down to 8.7GB. Our code is available on https://github.com/lzzzzzm/STCOcc
Zhimin Liao, Ping Wei 0001, Shuaijia Chen, Ziyang Ren
CVPR5
2025 Stochastic-Aware Mamba Diffusion for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction plays a crucial role in understanding human behavior and intentions. Due to the inherent randomness in human movement, current research constructs trajectories in stochastic space and uses diffusion models to reverse the denoising process. The commonly used denoising model, Transformer, is affected by uneven noise sampling, leading to confusion between temporal consistency and randomness across multiple time points. To address this issue, we propose a novel framework named Stochastic-Aware Mamba Diffusion (SAMD), which combines Stochastic-Aware Mamba with the Diffusion model to predict motion noise and motion states in the stochastic space. It utilizes a temporal aggregator to extract temporal consistency features. We construct a motion-selective state space model, which includes the adaptive transition between consistency motion states and stochastic states for balancing stability and diversity. The motion gating unit activates pertinent information within motion states to extract motion noise for trajectory prediction. Our approach achieves state-of-the-art results on the ETH-UCY and SDD datasets while significantly reducing the computational cost.
Ziyang Ren, Ping Wei 0001, Haowen Tang, Jialu Qin
ICASSP1
2025 $I^{\mathbf{2}}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
Zhimin Liao, Ping Wei 0001, Shuaijia Chen, Ziyang Ren
ICCV6
2025 TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent Diffusion
Ziyang Ren, Ping Wei 0001, Shangqi Deng, Haowen Tang, Jiapeng Li 0003
ICCV1
2024 Learning Scene-Goal-Aware Motion Representation for Trajectory Prediction
Ziyang Ren, Ping Wei 0001, Haowen Tang
BMVC1
2024 Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
abstract
Video moment retrieval and highlight detection are two highly valuable tasks in video understanding, but until recently they have been jointly studied. Although existing studies have made impressive advancement recently, they predominantly follow the data-driven bottom-up paradigm. Such paradigm overlooks task-specific and inter-task effects, resulting in poor model performance. In this paper, we propose a novel task-driven top-down framework TaskWeave for joint moment retrieval and highlight detection. The framework introduces a task-decoupled unit to capture task-specific and common representations. To investigate the interplay between the two tasks, we propose an inter-task feedback mechanism, which transforms the results of one task as guiding masks to assist the other task. Different from existing methods, we present a task-dependent joint loss function to optimize the model. Comprehensive experiments and in-depth ablation studies on QVHighlights, TVSum, and Charades-STA datasets corroborate the effectiveness and flexibility of the proposed framework. Codes are available at github.com/EdenGabriel/TaskWeave.
Ping Wei 0001, Ziyang Ren
CVPR4
2024 Gated Multi-Scale Transformer for Temporal Action Localization
abstract
Temporal action localization (TAL) is a critical task in video understanding. Effectively utilizing multi-scale information and handling interactions across various scales have consistently posed challenging issues within the realm of TAL. In this paper, we propose a novel gated multi-scale Transformer model (TransGMC) for temporal action localization. A gated control mechanism is designed to filter and aggregate the information at different scales, by which the contributions of contexts at different temporal scales are well characterized. To enhance the feature representation at each temporal scale, the rich global-local contexts are extracted at each temporal scale. A cascade attention module that contains two seamlessly integrated channel attention and moment attention is proposed for capturing global temporal contexts. We utilize a new regression loss function for locating the time boundaries. We conducted experiments on four challenging benchmark datasets, including two third-person view datasets and two first-person view datasets. Our method achieves an average mAP of 67.5% on THUMOS14, 36.1% on ActivityNet v1.3, 24.9% on EPIC-Kitchens 100, and 23.2% on Ego4D, which all outperform the previous state-of-the-arts methods. Extensive ablation studies also validate the effectiveness of the proposed method. Code is available at https://github.com/EdenGabriel/TransGMC.
Ping Wei 0001, Ziyang Ren, Nanning Zheng 0001
IEEE Trans. Multim.3