Ziliang Ren

dblp:188/4134 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0001-7940-294XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Human motion prediction using knowledge distillation with incremental learning
Ziliang Ren, Yuze Ma, Almazbek Arzybaev, Peiting Li, Fuyong Zhang
Pattern Anal. Appl.1
2026 Multimodal alignment of event and text streams in spiking neural networks for human action recognition
Ziliang Ren, Qieshi Zhang, Weiyu Yu, Fuyong Zhang
Pattern Recognit.1
2026 Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute Relationships
abstract
Expressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Language-Guided Multimodal Spiking Neural Networks for Event-Based Action Recognition
abstract
Event-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs.
Ziliang Ren, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Multim.1
2025 Incorporating Long and Short-Term Memory Networks and Quaternion Transformation for Human Motion Prediction
abstract
Improving the accuracy and stability of human-computer interaction requires machines to have the ability to predict human movements. However, existing methods perform poorly in capturing the temporal dependence between different motion sequences, making it difficult to generate natural motion sequences with good generalization capabilities. This paper proposes a novel approach that combines long short-term memory networks (LSTMs) and quaternion transform (QT)-based layer structures for human motion prediction. Our LSTM model adopts a five-layer architecture that can effectively capture feature information over a long time range. At the same time, the residual LSTM structure (LSTMR) and attention mechanism are introduced into the encoder-decoder structure, which further enhances the model’s time-dependent modeling capabilities and key motion feature extraction capabilities. In addition, layer normalization and hybrid loss functions are used in combination before and after the training process to optimize the performance of the quaternion-based temporal layer. Through experimental verification on the public Human3.6M and CMU-MoCap data sets, our method performs excellently within the 1000 millisecond prediction time range, significantly improving the accuracy of prediction compared to existing methods.
Miaomiao Jin, Ziliang Ren, Qieshi Zhang, Tiezhu Zhao
IJCNN2
2025 Two-Stage Modal Feature Enhancement for Multispectral Object Detection
Tichao Wang, Ziliang Ren, Qieshi Zhang, Yimin Zhou 0001, Jun Cheng 0002
PRCV (5)2
2025 Adaptive multi-frequency attention network for human motion prediction
abstract
Human motion prediction aims to predict future human motion sequences based on historical motion sequences. Existing deterministic methods typically predict only a single future sequence, ignoring the inherent stochasticity and diversity of human motion. To address this limitation, we propose a stochastic human action prediction network to achieve diverse motion prediction. Specifically, in the latent space, we introduce an adaptive multi-frequency attention module and a graph convolutional module. The graph convolutional network module encodes the sequence into a continuous latent representation, while the adaptive multi-frequency attention module captures multi-frequency information from historical action sequences and selects important frequency components for the encoder and decoder to improve prediction accuracy. Additionally, a set of motion queries and semantic latent directions are introduced to further enhance the diversity of prediction results. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both prediction accuracy and diversity.
Jianbo Shang, Ziliang Ren, Qieshi Zhang, Fuyong Zhang, Tiezhu Zhao
SMC2
2025 Re-parameterization Convolution Spiking Neural Network for Object Detection*
abstract
As the third-generation of neural networks, Spiking Neural Networks (SNN) have biological plausibility and low-power advantages over Artificial Neural Networks (ANNs). However, applying SNN to object detection tasks presents challenges in achieving both high detection accuracy and fast processing speed. To overcome the aforementioned problems, we propose a Re-parameterization SpikeYOLO (RepSpikeYOLO) for high-performance and energy-effcient object detection Our design revolves around network architecture and SNN residual block. Foremost, the SNN are difficult to train, mainly owing to their complex dynamics of neurons and non-differentiable spike operations. We design a YOLO architecture to solve this problem by training SNN with surrogate gradients. Second, object detection is more sensitive to gradient vanishing or exploding in training deep SNN. To address this challenge, we design a new SNN residual block, which can effectively extend the depth of the directly-trained with low power consumption. The proposed approach is validated on both COCO dataset and PASCAL VOC dataset. It is shown that our YOLO could achieve a comparable performance to the ANN with the same architecture. On the COCO dataset, we obtain 54% mAP@50 and 33.7% mAP@50:95, which is +3.9% and 3.7% higher than the prior state-of-the-art SNN, respectively. On the PASCAL VOC dataset, we achieve 75.1% mAP@50, which is +21.05% higher than the prior state-of-the-art SNN.
Ziliang Ren, Qieshi Zhang, Kadyrkulova Kyial Kudayberdievna, Taalaybekova Aizharkyn
SMC2
2025 Imitation learning and interactive game-based prediction-decision planning
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Weiyu Yu, Huiquan Zhang
Neurocomputing3
2025 Skeleton-guided and supervised learning of hybrid network for multi-modal action recognition
Ziliang Ren
J. Vis. Commun. Image Represent.1
2025 Heterogeneous information alignment and re-ranking for cross-modal pedestrian re-identification
Tiezhu Zhao, Xiaolun Liang, Kejing He 0001, Qiuhong Yang, Ziliang Ren
Multim. Tools Appl.5
2025 Masked cosine similarity prediction for self-supervised skeleton-based action representation learning
Ziliang Ren, Ronggui Liu, Xiangyang Gao, Qieshi Zhang
Pattern Anal. Appl.1
2024 MSGAT: Multi-Stage Graph Attention Network For Human Motion Prediction
abstract
Human motion prediction (HMP) refers to predicting the future body pose from the historical pose sequence. Many existing methods use Graph Convolutional Networks (GCN) to model the human body and convert the human pose from the pose space to the trajectory space or 3D coordinates. Furthermore, GCN treat human poses as a generic graph formed by links between each pair of body joints to encode the dependence of human spatial poses as well as temporal information by working in trajectory space. We design a multi-stage distributed processing network that includes Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). The multistage strategy enables us to gradually acquire smoother inputs. Additionally, we have incorporated an attention mechanism within the processing framework, which helps T-DGCN better capture temporal dependencies. As a result, the proposed network not only facilitates more effective feature extraction but also achieves state-of-the-art performance on the CMU-Mocap and 3DPW datasets. Our code is available at https://github.com/ihavenotgoodname/MSGAT.
Ziliang Ren, Gulin Wang, Qieshi Zhang
ICIP2
2024 AAGF: An Efficient Transformer With Mix-Features For Visual Place Recognition
abstract
Visual Place Recognition (VPR) is a task predicting the current location solely based on the visual features of images. It is susceptible to changes in perspective, lighting, and environmental conditions. Now the performance of the VPR method still relies on re-ranking, and the effectiveness of pure global retrieval is not ideal. To address this, we introduce a novel feature aggregation model based on the Transformer architecture, Agent-Attention with Gating Forward, which can aggregate the global relationships from feature maps obtained by a pre-trained backbone into a new global feature. Besides, a valid training strategy, Mix-Features Data Augment, is proposed to enhance the diversity of features and make the model more robust. Through experiments on multiple benchmarks, we demonstrate that our approach outperforms many existing techniques in terms of lightweight pre-trained backbone network aggregation.
Kuan Zhou, Zhenyu Xu 0014, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren, Xiangyang Gao
ICIP5
2024 Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to capture foundational semantic details that include the distinctive attributes for unique place identification. To address this problem, we propose a new visual token-guided VPR framework that contains a semantic-focused patch tokenizer and a multi-branch Mixer. To mitigate the inference from place-unrelated objects, the semantic-focused patch tokenizer exploits attention-based channel selection and spatial partition, which efficiently captures important semantic information within the channels and preserve spatial relationships among the backbone features. To extract abstract features with spatial structure information, the multi-branch Mixer utilizes a multi-branch structure to aggregate local and global position information, improving the robustness of global representations to environmental changes. Experimental results demonstrate that our method outperforms state-of-the-art methods, achieving 85.3% Recall@1 on the MSLS val dataset and 59.1% Recall@1 on the Nordland dataset when using ResNet18 as the backbone.
Zhenyu Xu 0014, Ziliang Ren, Qieshi Zhang, Jie Lou, Dacheng Tao, Jun Cheng 0002
ICRA2
2024 A Dense-Sparse Complementary Network for Human Action Recognition based on RGB and Skeleton Modalities
Qin Cheng, Jun Cheng 0002, Zhen Liu 0049, Ziliang Ren
Expert Syst. Appl.4
2023 Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
Qieshi Zhang, Zhenyu Xu 0014, Yuhang Kang, Fusheng Hao, Ziliang Ren, Jun Cheng 0002
Knowl. Based Syst.5
2023 Multi-scale spatial-temporal convolutional neural network for skeleton-based action recognition
Qin Cheng, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang
Pattern Anal. Appl.3
2022 Indoor Target-Driven Visual Navigation based on Spatial Semantic Information
abstract
Target-driven visual navigation is a widely focused learning-based approach in the field of computer vision. However, it faces two major challenges: poor generalization ability to unknown scenes, and poor navigation performance for increased number of scenes. In this paper, an end-to-end target-driven visual navigation method, which uses Spatial Semantic Information (SSI) to navigate the agent to the target, is presented. To fully integrate the spatial and semantic information in the scene, visual information is encoded into an 8-D spatial context vector. In addition, the size of detected bounding box is used to improve the reward function in end-to-end learning to solve the problem of sparse rewards. Experiments in interactive environment dataset AI2-THOR show that compared with state-of-the-art approaches, our approach has a higher success rate and a better route to target.
Jiaojie Yan, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren
ICIP4
2022 EEP-Net: Enhancing Local Neighborhood Features and Efficient Semantic Segmentation of Scale Point Clouds
Fuxiang Wu, Qieshi Zhang, Ziliang Ren, Jun Cheng 0002
PRCV (3)4
2022 Dual-stream cross-modality fusion transformer for RGB-D action recognition
Zhen Liu 0049, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Chengqun Song
Knowl. Based Syst.4
2022 Cross-Modality Compensation Convolutional Neural Networks for RGB-D Action Recognition
abstract
RGB-D-based human action recognition has attracted much attention recently because it can provide more complementary information than a single modality. However, it is difficult for two modalities to effectively learn spatial-temporal information from each other. To facilitate information interaction between different modalities, a cross-modality compensation convolutional neural network (ConvNet) is proposed for human action recognition, which enhances the discriminative ability by jointly learning compensation features from the RGB and depth modalities. Moreover, we design a cross-modality compensation block (CMCB) to extract compensation features from the RGB and depth modalities. Specifically, CMCB is incorporated into two typical network architectures, ResNet and VGG, to verify the ability to improve the performance of our model. The proposed architecture has been evaluated on three challenging datasets: NTU RGB+D 120, THU-READ and PKU-MMD. We experimentally verify that our proposed model with CMCB is effective for different input types, such as pairs of raw images and dynamic images constructed from the entire RGB-D sequence, and the experimental results show that the proposed framework achieves state-of-the-art performance on all three datasets.
Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Fusheng Hao
IEEE Trans. Circuits Syst. Video Technol.2
2021 VGG-CAE: Unsupervised Visual Place Recognition Using VGG16-Based Convolutional Autoencoder
Zhenyu Xu 0014, Qieshi Zhang, Fusheng Hao, Ziliang Ren, Yuhang Kang, Jun Cheng 0002
PRCV (2)4
2021 Segment spatial-temporal representation and cooperative learning of convolution neural networks for multimodal-based action recognition
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Fusheng Hao, Xiangyang Gao
Neurocomputing1
2021 Multi-modality learning for human action recognition
Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Pengyi Hao, Jun Cheng 0002
Multim. Tools Appl.1
2020 Multiple Time Scale Motion Images for Action Recognition
abstract
This paper proposes a simple and effective approach for RGB-based action recognition using multiple streams Convolutional Neural Networks (ConvNets). We utilize a new method to represent temporal structure of RGB videos, named Motion Image (MI), which is constructed from the difference between frames with a certain time scale. Considering the different duration of different actions, multiple time scale sampling MIs can obtain more temporal information. Furthermore, we adopt multiple streams ConvNets, including MIs and RGB streams, to learn spatial-temporal features for action recognition. Our approach has been evaluated on UCF-101 and HMDB-51, and the experimental results demonstrate the effectiveness and significantly improve action recognition rate at a small computational cost.
Qin Cheng, Ziliang Ren, Jun Cheng 0002
HealthCom2
2020 Phase-Sensitive Model for Temporal Action Proposal Generation
abstract
Temporal action proposal generation is an important and challenging task, aiming to localize the position where an action or event may occur in an untrimmed video. In this paper, we propose an efficient and end-to-end framework to generate temporal action proposals, named Phase-Sensitive Model (PSM), which fully understands all phases of temporal information. In particular, the PSM consists two modules: Boundary Phase Classification (BPC) and Action Phase Classification (APC). The BPC aims to provide two temporal boundary phase confidence maps by rich local information, while the APC is designed to generate an action phase confidence map by global features. Moreover, we introduce a new method boundary probability calculation to get the final score. Our experiments on ActivityNet-1.3 show a significant improvement with remarkable efficiency and generalizability.
Ziliang Ren, Lei Wang 0018, Jun Cheng 0002
HealthCom3
2020 ST-LSTM: Spatio-Temporal Graph Based Long Short-Term Memory Network For Vehicle Trajectory Prediction
abstract
Autonomous vehicles need the ability to predict the trajectory of surrounding vehicles, so as to make a rational decision planning, improve driving safety and ride comfort. In this paper, a new hierarchical Long Short-Term Memory (LSTM) based on Spatio-Temporal (ST) graph is proposed for vehicle trajectory prediction. Our ST-LSTM uses three layers of different LSTMs to capture the information of spatial, temporal and trajectory data, and LSTM-based encoder-decoder model as a whole, which is capable of accurately predicting future trajectories for vehicles on the highway. Our model trained and validated on the publicly available NGSIM US-101 and I-80 datasets. In comparison to state-of-art methods, our method could achieve a more accurate prediction trajectory over 5s time horizon.
Guangxi Chen, Qieshi Zhang, Ziliang Ren, Xiangyang Gao, Jun Cheng 0002
ICIP4