EDBT 2026 Demo / reviewers in the wild / expert
Ziliang Ren
dblp:188/4134
· DBLP profile ↗
28ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0001-7940-294XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human motion prediction using knowledge distillation with incremental learning
Ziliang Ren, Yuze Ma, Almazbek Arzybaev, Peiting Li, Fuyong Zhang |
Pattern Anal. Appl. | 1 |
| 2026 | Multimodal alignment of event and text streams in spiking neural networks for human action recognition
Ziliang Ren, Qieshi Zhang, Weiyu Yu, Fuyong Zhang |
Pattern Recognit. | 1 |
| 2026 | Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute RelationshipsabstractExpressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Language-Guided Multimodal Spiking Neural Networks for Event-Based Action RecognitionabstractEvent-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs. Ziliang Ren, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Trans. Multim. | 1 |
| 2025 | Incorporating Long and Short-Term Memory Networks and Quaternion Transformation for Human Motion PredictionabstractImproving the accuracy and stability of human-computer interaction requires machines to have the ability to predict human movements. However, existing methods perform poorly in capturing the temporal dependence between different motion sequences, making it difficult to generate natural motion sequences with good generalization capabilities. This paper proposes a novel approach that combines long short-term memory networks (LSTMs) and quaternion transform (QT)-based layer structures for human motion prediction. Our LSTM model adopts a five-layer architecture that can effectively capture feature information over a long time range. At the same time, the residual LSTM structure (LSTMR) and attention mechanism are introduced into the encoder-decoder structure, which further enhances the model’s time-dependent modeling capabilities and key motion feature extraction capabilities. In addition, layer normalization and hybrid loss functions are used in combination before and after the training process to optimize the performance of the quaternion-based temporal layer. Through experimental verification on the public Human3.6M and CMU-MoCap data sets, our method performs excellently within the 1000 millisecond prediction time range, significantly improving the accuracy of prediction compared to existing methods. Miaomiao Jin, Ziliang Ren, Qieshi Zhang, Tiezhu Zhao |
IJCNN | 2 |
| 2025 | Two-Stage Modal Feature Enhancement for Multispectral Object Detection
Tichao Wang, Ziliang Ren, Qieshi Zhang, Yimin Zhou 0001, Jun Cheng 0002 |
PRCV (5) | 2 |
| 2025 | Adaptive multi-frequency attention network for human motion predictionabstractHuman motion prediction aims to predict future human motion sequences based on historical motion sequences. Existing deterministic methods typically predict only a single future sequence, ignoring the inherent stochasticity and diversity of human motion. To address this limitation, we propose a stochastic human action prediction network to achieve diverse motion prediction. Specifically, in the latent space, we introduce an adaptive multi-frequency attention module and a graph convolutional module. The graph convolutional network module encodes the sequence into a continuous latent representation, while the adaptive multi-frequency attention module captures multi-frequency information from historical action sequences and selects important frequency components for the encoder and decoder to improve prediction accuracy. Additionally, a set of motion queries and semantic latent directions are introduced to further enhance the diversity of prediction results. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both prediction accuracy and diversity. Jianbo Shang, Ziliang Ren, Qieshi Zhang, Fuyong Zhang, Tiezhu Zhao |
SMC | 2 |
| 2025 | Re-parameterization Convolution Spiking Neural Network for Object Detection*abstractAs the third-generation of neural networks, Spiking Neural Networks (SNN) have biological plausibility and low-power advantages over Artificial Neural Networks (ANNs). However, applying SNN to object detection tasks presents challenges in achieving both high detection accuracy and fast processing speed. To overcome the aforementioned problems, we propose a Re-parameterization SpikeYOLO (RepSpikeYOLO) for high-performance and energy-effcient object detection Our design revolves around network architecture and SNN residual block. Foremost, the SNN are difficult to train, mainly owing to their complex dynamics of neurons and non-differentiable spike operations. We design a YOLO architecture to solve this problem by training SNN with surrogate gradients. Second, object detection is more sensitive to gradient vanishing or exploding in training deep SNN. To address this challenge, we design a new SNN residual block, which can effectively extend the depth of the directly-trained with low power consumption. The proposed approach is validated on both COCO dataset and PASCAL VOC dataset. It is shown that our YOLO could achieve a comparable performance to the ANN with the same architecture. On the COCO dataset, we obtain 54% mAP@50 and 33.7% mAP@50:95, which is +3.9% and 3.7% higher than the prior state-of-the-art SNN, respectively. On the PASCAL VOC dataset, we achieve 75.1% mAP@50, which is +21.05% higher than the prior state-of-the-art SNN. Ziliang Ren, Qieshi Zhang, Kadyrkulova Kyial Kudayberdievna, Taalaybekova Aizharkyn |
SMC | 2 |
| 2025 | Imitation learning and interactive game-based prediction-decision planning
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Weiyu Yu, Huiquan Zhang |
Neurocomputing | 3 |
| 2025 | Skeleton-guided and supervised learning of hybrid network for multi-modal action recognition
Ziliang Ren |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Heterogeneous information alignment and re-ranking for cross-modal pedestrian re-identification
Tiezhu Zhao, Xiaolun Liang, Kejing He 0001, Qiuhong Yang, Ziliang Ren |
Multim. Tools Appl. | 5 |
| 2025 | Masked cosine similarity prediction for self-supervised skeleton-based action representation learning
Ziliang Ren, Ronggui Liu, Xiangyang Gao, Qieshi Zhang |
Pattern Anal. Appl. | 1 |
| 2024 | MSGAT: Multi-Stage Graph Attention Network For Human Motion PredictionabstractHuman motion prediction (HMP) refers to predicting the future body pose from the historical pose sequence. Many existing methods use Graph Convolutional Networks (GCN) to model the human body and convert the human pose from the pose space to the trajectory space or 3D coordinates. Furthermore, GCN treat human poses as a generic graph formed by links between each pair of body joints to encode the dependence of human spatial poses as well as temporal information by working in trajectory space. We design a multi-stage distributed processing network that includes Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). The multistage strategy enables us to gradually acquire smoother inputs. Additionally, we have incorporated an attention mechanism within the processing framework, which helps T-DGCN better capture temporal dependencies. As a result, the proposed network not only facilitates more effective feature extraction but also achieves state-of-the-art performance on the CMU-Mocap and 3DPW datasets. Our code is available at https://github.com/ihavenotgoodname/MSGAT. Ziliang Ren, Gulin Wang, Qieshi Zhang |
ICIP | 2 |
| 2024 | AAGF: An Efficient Transformer With Mix-Features For Visual Place RecognitionabstractVisual Place Recognition (VPR) is a task predicting the current location solely based on the visual features of images. It is susceptible to changes in perspective, lighting, and environmental conditions. Now the performance of the VPR method still relies on re-ranking, and the effectiveness of pure global retrieval is not ideal. To address this, we introduce a novel feature aggregation model based on the Transformer architecture, Agent-Attention with Gating Forward, which can aggregate the global relationships from feature maps obtained by a pre-trained backbone into a new global feature. Besides, a valid training strategy, Mix-Features Data Augment, is proposed to enhance the diversity of features and make the model more robust. Through experiments on multiple benchmarks, we demonstrate that our approach outperforms many existing techniques in terms of lightweight pre-trained backbone network aggregation. Kuan Zhou, Zhenyu Xu 0014, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren, Xiangyang Gao |
ICIP | 5 |
| 2024 | Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place RecognitionabstractVisual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to capture foundational semantic details that include the distinctive attributes for unique place identification. To address this problem, we propose a new visual token-guided VPR framework that contains a semantic-focused patch tokenizer and a multi-branch Mixer. To mitigate the inference from place-unrelated objects, the semantic-focused patch tokenizer exploits attention-based channel selection and spatial partition, which efficiently captures important semantic information within the channels and preserve spatial relationships among the backbone features. To extract abstract features with spatial structure information, the multi-branch Mixer utilizes a multi-branch structure to aggregate local and global position information, improving the robustness of global representations to environmental changes. Experimental results demonstrate that our method outperforms state-of-the-art methods, achieving 85.3% Recall@1 on the MSLS val dataset and 59.1% Recall@1 on the Nordland dataset when using ResNet18 as the backbone. Zhenyu Xu 0014, Ziliang Ren, Qieshi Zhang, Jie Lou, Dacheng Tao, Jun Cheng 0002 |
ICRA | 2 |
| 2024 | A Dense-Sparse Complementary Network for Human Action Recognition based on RGB and Skeleton Modalities
Qin Cheng, Jun Cheng 0002, Zhen Liu 0049, Ziliang Ren |
Expert Syst. Appl. | 4 |
| 2023 | Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
Qieshi Zhang, Zhenyu Xu 0014, Yuhang Kang, Fusheng Hao, Ziliang Ren, Jun Cheng 0002 |
Knowl. Based Syst. | 5 |
| 2023 | Multi-scale spatial-temporal convolutional neural network for skeleton-based action recognition
Qin Cheng, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang |
Pattern Anal. Appl. | 3 |
| 2022 | Indoor Target-Driven Visual Navigation based on Spatial Semantic InformationabstractTarget-driven visual navigation is a widely focused learning-based approach in the field of computer vision. However, it faces two major challenges: poor generalization ability to unknown scenes, and poor navigation performance for increased number of scenes. In this paper, an end-to-end target-driven visual navigation method, which uses Spatial Semantic Information (SSI) to navigate the agent to the target, is presented. To fully integrate the spatial and semantic information in the scene, visual information is encoded into an 8-D spatial context vector. In addition, the size of detected bounding box is used to improve the reward function in end-to-end learning to solve the problem of sparse rewards. Experiments in interactive environment dataset AI2-THOR show that compared with state-of-the-art approaches, our approach has a higher success rate and a better route to target. Jiaojie Yan, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren |
ICIP | 4 |
| 2022 | EEP-Net: Enhancing Local Neighborhood Features and Efficient Semantic Segmentation of Scale Point Clouds
Fuxiang Wu, Qieshi Zhang, Ziliang Ren, Jun Cheng 0002 |
PRCV (3) | 4 |
| 2022 | Dual-stream cross-modality fusion transformer for RGB-D action recognition
Zhen Liu 0049, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Chengqun Song |
Knowl. Based Syst. | 4 |
| 2022 | Cross-Modality Compensation Convolutional Neural Networks for RGB-D Action RecognitionabstractRGB-D-based human action recognition has attracted much attention recently because it can provide more complementary information than a single modality. However, it is difficult for two modalities to effectively learn spatial-temporal information from each other. To facilitate information interaction between different modalities, a cross-modality compensation convolutional neural network (ConvNet) is proposed for human action recognition, which enhances the discriminative ability by jointly learning compensation features from the RGB and depth modalities. Moreover, we design a cross-modality compensation block (CMCB) to extract compensation features from the RGB and depth modalities. Specifically, CMCB is incorporated into two typical network architectures, ResNet and VGG, to verify the ability to improve the performance of our model. The proposed architecture has been evaluated on three challenging datasets: NTU RGB+D 120, THU-READ and PKU-MMD. We experimentally verify that our proposed model with CMCB is effective for different input types, such as pairs of raw images and dynamic images constructed from the entire RGB-D sequence, and the experimental results show that the proposed framework achieves state-of-the-art performance on all three datasets. Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Fusheng Hao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | VGG-CAE: Unsupervised Visual Place Recognition Using VGG16-Based Convolutional Autoencoder
Zhenyu Xu 0014, Qieshi Zhang, Fusheng Hao, Ziliang Ren, Yuhang Kang, Jun Cheng 0002 |
PRCV (2) | 4 |
| 2021 | Segment spatial-temporal representation and cooperative learning of convolution neural networks for multimodal-based action recognition
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Fusheng Hao, Xiangyang Gao |
Neurocomputing | 1 |
| 2021 | Multi-modality learning for human action recognition
Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Pengyi Hao, Jun Cheng 0002 |
Multim. Tools Appl. | 1 |
| 2020 | Multiple Time Scale Motion Images for Action RecognitionabstractThis paper proposes a simple and effective approach for RGB-based action recognition using multiple streams Convolutional Neural Networks (ConvNets). We utilize a new method to represent temporal structure of RGB videos, named Motion Image (MI), which is constructed from the difference between frames with a certain time scale. Considering the different duration of different actions, multiple time scale sampling MIs can obtain more temporal information. Furthermore, we adopt multiple streams ConvNets, including MIs and RGB streams, to learn spatial-temporal features for action recognition. Our approach has been evaluated on UCF-101 and HMDB-51, and the experimental results demonstrate the effectiveness and significantly improve action recognition rate at a small computational cost. Qin Cheng, Ziliang Ren, Jun Cheng 0002 |
HealthCom | 2 |
| 2020 | Phase-Sensitive Model for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an important and challenging task, aiming to localize the position where an action or event may occur in an untrimmed video. In this paper, we propose an efficient and end-to-end framework to generate temporal action proposals, named Phase-Sensitive Model (PSM), which fully understands all phases of temporal information. In particular, the PSM consists two modules: Boundary Phase Classification (BPC) and Action Phase Classification (APC). The BPC aims to provide two temporal boundary phase confidence maps by rich local information, while the APC is designed to generate an action phase confidence map by global features. Moreover, we introduce a new method boundary probability calculation to get the final score. Our experiments on ActivityNet-1.3 show a significant improvement with remarkable efficiency and generalizability. Ziliang Ren, Lei Wang 0018, Jun Cheng 0002 |
HealthCom | 3 |
| 2020 | ST-LSTM: Spatio-Temporal Graph Based Long Short-Term Memory Network For Vehicle Trajectory PredictionabstractAutonomous vehicles need the ability to predict the trajectory of surrounding vehicles, so as to make a rational decision planning, improve driving safety and ride comfort. In this paper, a new hierarchical Long Short-Term Memory (LSTM) based on Spatio-Temporal (ST) graph is proposed for vehicle trajectory prediction. Our ST-LSTM uses three layers of different LSTMs to capture the information of spatial, temporal and trajectory data, and LSTM-based encoder-decoder model as a whole, which is capable of accurately predicting future trajectories for vehicles on the highway. Our model trained and validated on the publicly available NGSIM US-101 and I-80 datasets. In comparison to state-of-art methods, our method could achieve a more accurate prediction trajectory over 5s time horizon. Guangxi Chen, Qieshi Zhang, Ziliang Ren, Xiangyang Gao, Jun Cheng 0002 |
ICIP | 4 |