Qieshi Zhang

dblp:24/70 · DBLP profile ↗
← Back
58ranked-venue papers
7as first author
38since 2021 · last 2026
0000-0001-6358-1840ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing robustness of neural ODEs via attractor dynamics
Qieshi Zhang, Jun Cheng 0002
Neurocomputing2
2026 Multimodal alignment of event and text streams in spiking neural networks for human action recognition
Ziliang Ren, Qieshi Zhang, Weiyu Yu, Fuyong Zhang
Pattern Recognit.3
2026 Language-Guided Multimodal Spiking Neural Networks for Event-Based Action Recognition
abstract
Event-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs.
Ziliang Ren, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Multim.4
2025 Incorporating Long and Short-Term Memory Networks and Quaternion Transformation for Human Motion Prediction
abstract
Improving the accuracy and stability of human-computer interaction requires machines to have the ability to predict human movements. However, existing methods perform poorly in capturing the temporal dependence between different motion sequences, making it difficult to generate natural motion sequences with good generalization capabilities. This paper proposes a novel approach that combines long short-term memory networks (LSTMs) and quaternion transform (QT)-based layer structures for human motion prediction. Our LSTM model adopts a five-layer architecture that can effectively capture feature information over a long time range. At the same time, the residual LSTM structure (LSTMR) and attention mechanism are introduced into the encoder-decoder structure, which further enhances the model’s time-dependent modeling capabilities and key motion feature extraction capabilities. In addition, layer normalization and hybrid loss functions are used in combination before and after the training process to optimize the performance of the quaternion-based temporal layer. Through experimental verification on the public Human3.6M and CMU-MoCap data sets, our method performs excellently within the 1000 millisecond prediction time range, significantly improving the accuracy of prediction compared to existing methods.
Miaomiao Jin, Ziliang Ren, Qieshi Zhang, Tiezhu Zhao
IJCNN3
2025 Two-Stage Modal Feature Enhancement for Multispectral Object Detection
Tichao Wang, Ziliang Ren, Qieshi Zhang, Yimin Zhou 0001, Jun Cheng 0002
PRCV (5)3
2025 Adaptive multi-frequency attention network for human motion prediction
abstract
Human motion prediction aims to predict future human motion sequences based on historical motion sequences. Existing deterministic methods typically predict only a single future sequence, ignoring the inherent stochasticity and diversity of human motion. To address this limitation, we propose a stochastic human action prediction network to achieve diverse motion prediction. Specifically, in the latent space, we introduce an adaptive multi-frequency attention module and a graph convolutional module. The graph convolutional network module encodes the sequence into a continuous latent representation, while the adaptive multi-frequency attention module captures multi-frequency information from historical action sequences and selects important frequency components for the encoder and decoder to improve prediction accuracy. Additionally, a set of motion queries and semantic latent directions are introduced to further enhance the diversity of prediction results. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both prediction accuracy and diversity.
Jianbo Shang, Ziliang Ren, Qieshi Zhang, Fuyong Zhang, Tiezhu Zhao
SMC3
2025 Re-parameterization Convolution Spiking Neural Network for Object Detection*
abstract
As the third-generation of neural networks, Spiking Neural Networks (SNN) have biological plausibility and low-power advantages over Artificial Neural Networks (ANNs). However, applying SNN to object detection tasks presents challenges in achieving both high detection accuracy and fast processing speed. To overcome the aforementioned problems, we propose a Re-parameterization SpikeYOLO (RepSpikeYOLO) for high-performance and energy-effcient object detection Our design revolves around network architecture and SNN residual block. Foremost, the SNN are difficult to train, mainly owing to their complex dynamics of neurons and non-differentiable spike operations. We design a YOLO architecture to solve this problem by training SNN with surrogate gradients. Second, object detection is more sensitive to gradient vanishing or exploding in training deep SNN. To address this challenge, we design a new SNN residual block, which can effectively extend the depth of the directly-trained with low power consumption. The proposed approach is validated on both COCO dataset and PASCAL VOC dataset. It is shown that our YOLO could achieve a comparable performance to the ANN with the same architecture. On the COCO dataset, we obtain 54% mAP@50 and 33.7% mAP@50:95, which is +3.9% and 3.7% higher than the prior state-of-the-art SNN, respectively. On the PASCAL VOC dataset, we achieve 75.1% mAP@50, which is +21.05% higher than the prior state-of-the-art SNN.
Ziliang Ren, Qieshi Zhang, Kadyrkulova Kyial Kudayberdievna, Taalaybekova Aizharkyn
SMC3
2025 Imitation learning and interactive game-based prediction-decision planning
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Weiyu Yu, Huiquan Zhang
Neurocomputing4
2025 Neural adaptive delay differential equations
Qieshi Zhang, Jun Cheng 0002
Neurocomputing2
2025 Progressive background-foreground difference enhancement for few-shot 3D point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Qieshi Zhang, Jun Cheng 0002
Image Vis. Comput.3
2025 Masked cosine similarity prediction for self-supervised skeleton-based action representation learning
Ziliang Ren, Ronggui Liu, Xiangyang Gao, Qieshi Zhang
Pattern Anal. Appl.5
2025 Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language Models
abstract
Due to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Signal Process. Lett.4
2025 Class-Irrelevant Feature Removal for Few-Shot Image Classification
abstract
Most existing few-shot image classification methods employ global pooling to aggregate class-relevant local features in a data-drive manner. Due to the difficulty and inaccuracy in locating class-relevant regions in complex scenarios, as well as the large semantic diversity of local features, the class-irrelevant information could reduce the robustness of the representations obtained by performing global pooling. Meanwhile, the scarcity of labeled images exacerbates the difficulties of data-hungry deep models in identifying class-relevant regions. These issues severely limit deep models' few-shot learning ability. In this work, we propose to remove the class-irrelevant information by making local features class relevant, thus bypassing the big challenge of identifying which local features are class irrelevant. The resulting class-irrelevant feature removal (CIFR) method consists of three phases. First, we employ the masked image modeling strategy to build an understanding of images' internal structures that generalizes well. Second, we design a semantic-complementary feature propagation module to make local features class relevant. Third, we introduce a weighted dense-connected similarity measure, based on which a loss function is raised to fine-tune the entire pipeline, with the aim of further enhancing the semantic consistency of the class-relevant local features. Visualization results show that CIFR achieves the removal of class-irrelevant information by making local features related to classes. Comparison results on four benchmark datasets indicate that CIFR yields very promising performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 MSGAT: Multi-Stage Graph Attention Network For Human Motion Prediction
abstract
Human motion prediction (HMP) refers to predicting the future body pose from the historical pose sequence. Many existing methods use Graph Convolutional Networks (GCN) to model the human body and convert the human pose from the pose space to the trajectory space or 3D coordinates. Furthermore, GCN treat human poses as a generic graph formed by links between each pair of body joints to encode the dependence of human spatial poses as well as temporal information by working in trajectory space. We design a multi-stage distributed processing network that includes Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). The multistage strategy enables us to gradually acquire smoother inputs. Additionally, we have incorporated an attention mechanism within the processing framework, which helps T-DGCN better capture temporal dependencies. As a result, the proposed network not only facilitates more effective feature extraction but also achieves state-of-the-art performance on the CMU-Mocap and 3DPW datasets. Our code is available at https://github.com/ihavenotgoodname/MSGAT.
Ziliang Ren, Gulin Wang, Qieshi Zhang
ICIP5
2024 AAGF: An Efficient Transformer With Mix-Features For Visual Place Recognition
abstract
Visual Place Recognition (VPR) is a task predicting the current location solely based on the visual features of images. It is susceptible to changes in perspective, lighting, and environmental conditions. Now the performance of the VPR method still relies on re-ranking, and the effectiveness of pure global retrieval is not ideal. To address this, we introduce a novel feature aggregation model based on the Transformer architecture, Agent-Attention with Gating Forward, which can aggregate the global relationships from feature maps obtained by a pre-trained backbone into a new global feature. Besides, a valid training strategy, Mix-Features Data Augment, is proposed to enhance the diversity of features and make the model more robust. Through experiments on multiple benchmarks, we demonstrate that our approach outperforms many existing techniques in terms of lightweight pre-trained backbone network aggregation.
Kuan Zhou, Zhenyu Xu 0014, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren, Xiangyang Gao
ICIP3
2024 Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to capture foundational semantic details that include the distinctive attributes for unique place identification. To address this problem, we propose a new visual token-guided VPR framework that contains a semantic-focused patch tokenizer and a multi-branch Mixer. To mitigate the inference from place-unrelated objects, the semantic-focused patch tokenizer exploits attention-based channel selection and spatial partition, which efficiently captures important semantic information within the channels and preserve spatial relationships among the backbone features. To extract abstract features with spatial structure information, the multi-branch Mixer utilizes a multi-branch structure to aggregate local and global position information, improving the robustness of global representations to environmental changes. Experimental results demonstrate that our method outperforms state-of-the-art methods, achieving 85.3% Recall@1 on the MSLS val dataset and 59.1% Recall@1 on the Nordland dataset when using ResNet18 as the backbone.
Zhenyu Xu 0014, Ziliang Ren, Qieshi Zhang, Jie Lou, Dacheng Tao, Jun Cheng 0002
ICRA3
2024 Two-stage feature distribution rectification for few-shot point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Guosheng Cui, Fuxiang Wu, Mengjie Yang, Qieshi Zhang, Jun Cheng 0002
Pattern Recognit. Lett.6
2024 BDR6D: Bidirectional Deep Residual Fusion Network for 6D Pose Estimation
abstract
Six-dimensional (6D) pose estimation is an important branch in the field of robotics focused on enhancing the ability of robots to manipulate and grasp objects. The latest research trend in 6D pose estimation is to directly predict the positions of two-dimensional (2D) keypoints from a single red, green, and blue (RGB) image through convolutional neural networks (CNNs) and establish a corresponding relationship with the three-dimensional (3D) keypoints of the model. Then, the perspective-n-point (PnP) algorithm is used to recover the 6D pose parameters. Currently, two challenges are encountered in pose estimation based on an RGB image. On the one hand, an RGB image lacks depth information, and it is thus difficult to directly obtain the corresponding geometric object information. On the other hand, when depth information is available, it is difficult to efficiently fuse the features of the RGB image with the features of the corresponding depth image. In this paper, we propose a bidirectional depth residual fusion network with a depth prediction (DP) network to estimate the 6D poses of objects (BDR6D). The BDR6D network predicts the depth information of objects using an RGB image, converts the depth information into point cloud information, and performs feature extraction and representation together with the RGB information during the feature extraction and representation stages. Specifically, the RGB image is fed into the BDR6D network, the DP network predicts the depth information of the objects in the image, and the depth map and RGB image are input into a point cloud network (PCN) and CNN, respectively, for feature extraction and representation. We build the bidirectional depth residual (BDR) structure so that the CNN and PCN can share information during feature extraction and representation. This approach allows the two networks to use each other’s local and global information to improve feature extraction and representation. For the keypoint selection stage, we propose an effective 2D keypoint selection method that considers the appearance and geometric information of the object of interest. We evaluate the proposed method with three benchmark datasets and compare it with other 6D pose estimation algorithms. The experimental results show that our method outperforms the state-of-the-art approach. Finally, we deploy our proposed method in conjunction with the Universal Robots 5 manipulator (UR5) robot to grasp and manipulate objects.Note to Practitioners—The purpose of this paper is to solve the problem of 6D pose estimation for robot grasping. The existing RGB image-based pose estimation approach faces two challenges. On the one hand, a single RGB image lacks depth information, so that it is difficult to directly obtain the corresponding geometric object information. On the other hand, when depth information is available, it is difficult to efficiently fuse the features of the RGB image with the features of the corresponding depth image. To solve the above problems, a novel network that can predict the depth information of objects from an RGB image and fuse the depth information with the RGB information to estimate the 6D pose of objects is proposed. Furthermore, we propose an effective 2D keypoint selection method that considers the appearance and geometric information of objects of interest. We evaluate the proposed approach based on three benchmark datasets and the UR5 robot platform and verify that our method is effective.
Penglei Liu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans Autom. Sci. Eng.2
2023 Efficiently Fusing Sparse Lidar for Enhanced Self-Supervised Monocular Depth Estimation
abstract
Monocular self-supervised depth estimation with a low-cost sensor is the mainstream solution to gathering dense depth maps for robots and autonomous driving. In this paper, based on the philosophy "less is more" (i.e., focusing only on valid pixels in sparse LiDAR), we propose a novel framework, Efficient Sparse Depth (EffisDepth), for predicting dense depth. The Sparse Feature Extractor (SFE) embedded in the proposed framework effectively handles sparse LiDAR by forming sparse tensors. The Slender Group Block (SGB) is the main building block in SFE, which extracts features from sparse tensors via a structure of two branches. Extensive experiments show that our method achieves state-of-the-art performance on the KITTI benchmark, demonstrating the effectiveness of each proposed component and the self-supervised learning framework.
Mingrong Gong, Qieshi Zhang, Jun Cheng 0002
ICASSP4
2023 GSNet: Model Reconstruction Network for Category-level 6D Object Pose and Size Estimation
abstract
Category-level 6D pose and size estimation is to estimate the rotation, translation and size of the observed instance objects from an arbitrary angle in a cluttered scene. Compared with instance-level 6D pose estimation, there are two main challenges for category-level 6D pose estimation. One is that the algorithm needs to estimate the 6D pose and size of unseen objects, and no 3D models are available. Another is that different instance objects of the same class of objects differ greatly in shape. This paper propose a novel method to estimate the 6D pose and size of unseen objects from an RGB-D image. To handle intra-class shape variation, we propose an autoencoder-decoder that is trained on a set of object models to learn structural feature-invariant and shape-variant features of intra-class objects, and constructs a category-level priori model containing the structure feature and shape feature. To solve the problem of 3D model, this paper proposes a model reconstruction network including 3D graph convolution and spherical convolution (GSNet), which can reconstruct the 3D model of the observed instance object from the input RGB-D image and the priori model, and establish a dense correspon-dence between the 3D model and the observed instance object. Finally, random sample consensus (RANSAC) algorithm and Umeyama algorithm are used to estimate the 6D pose and size of the object. Extensive experiments on benchmark datasets show that the proposed method achieves state-of-the-art performance in category-level 6D object pose estimation. In order to prove that our method can be applied to the grasping and operation tasks of robots in industry and life, we deploy our method to a physical UR5 robot to perform grasping tasks on unseen but category known instances, and the results validate the efficacy of our proposed method.
Penglei Liu, Qieshi Zhang, Jun Cheng 0002
ICRA2
2023 MixPose: 3D Human Pose Estimation with Mixed Encoder
Jisheng Cheng, Qin Cheng, Mengjie Yang, Zhen Liu 0049, Qieshi Zhang, Jun Cheng 0002
PRCV (8)5
2023 Prototype expansion and feature calibration for few-shot point cloud semantic segmentation
Qieshi Zhang, Tichao Wang, Fusheng Hao, Fuxiang Wu, Jun Cheng 0002
Neurocomputing1
2023 Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
Qieshi Zhang, Zhenyu Xu 0014, Yuhang Kang, Fusheng Hao, Ziliang Ren, Jun Cheng 0002
Knowl. Based Syst.1
2023 Semantic-Aware Feature Aggregation for Few-Shot Image Classification
Fusheng Hao, Fuxiang Wu, Fengxiang He, Qieshi Zhang, Chengqun Song, Jun Cheng 0002
Neural Process. Lett.4
2023 Multi-scale spatial-temporal convolutional neural network for skeleton-based action recognition
Qin Cheng, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang
Pattern Anal. Appl.4
2023 Mixer-Based Semantic Spread for Few-Shot Learning
abstract
Key semantics can come from everywhere on an image. Semantic alignment is a key part of few-shot learning but still remains challenging. In this paper, we design a Mixer-Based Semantic Spread (MBSS) algorithm that employs amixermodule to spread the key semantic on the whole image, so that one can directly compare the processed image pairs. We first adopt a convolutional neural network to extract features from both support and query images and separate each of them into multiple Local Descriptor-based Representations (LDRs). The LDRs are then fed into themixerfor semantic spread, where every LDR attracts complementary information from its peers. In this way, the objective semantic is made spread on the whole image in a data-driven manner. The overall pipeline is supervised by a voting-based loss, guaranteeing a goodmixer. Visualization results validate the feasibility of ourmixer. Comprehensive experiments on three benchmark datasets, miniImageNet, tieredImageNet, and CUB, show that our algorithm achieves the state-of-the-art performance in both 5-way 1-shot and 5-way 5-shot settings.
Jun Cheng 0002, Fusheng Hao, Fengxiang He, Liu Liu 0014, Qieshi Zhang
IEEE Trans. Multim.5
2023 InDecGAN: Learning to Generate Complex Images From Captions via Independent Object-Level Decomposition and Enhancement
abstract
Text-to-image synthesis is a challenging problem, in which a complex scene contains diverse objects of various sizes and sub-images of objects belonging to the same class have diverse forms from different perspectives. Thus, synthesis models have difficulty in capturing varied objects in the complex scene. To alleviate these problems, we devise an independent object-level decomposing and enhancing generative adversarial networks, denoted as InDecGAN, to synthesize complex images and capture varied objects in a complex scene. Specifically, InDecGAN fully utilizes the independent object-level information, bounding boxes and high-resolution images of objects in training, by employing independent object-level pathways to synthesize varied objects. The independent object-level pathway integrates an independent object-level adversarial loss and the bounding box information to learn the visual features of objects independently, then, the main pathway exploits the features provided by the object-level pathway to compose the full scene and synthesize images. In addition, we analyze the generalization properties of the proposed InDecGAN and demonstrate the improvement from the perspective of the model architecture. Moreover, extensive experiments conducted on a widely used dataset are presented to demonstrate that the proposed model with an independent object-level pathway produces synthesized images of significantly improved quality.
Jun Cheng 0002, Fuxiang Wu, Liu Liu 0014, Qieshi Zhang, Leszek Rutkowski, Dacheng Tao
IEEE Trans. Multim.4
2022 GAZEATTENTIONNET: Gaze Estimation with Attentions
abstract
Predicting gaze point on mobile devices without calibration in unconstrained environments has great significance on human computer interaction. Appearance-based gaze estimation methods have been improved due to the recent advance in convolutional neural network (CNN) models and the availability of large-scale datasets. CNN models have limitations on extracting the global information of features and ignore the important information of local features. In this paper, we propose a novel structure named GazeAttentionNet. To improve the accuracy of gaze estimation, we use the global and local attention modules to utilize both global and local features. Firstly, we use MobileNetV2 and the self-attention layers as the global attention module to extract global features. Secondly, we add the local attention module containing the spatial attention to extract local features. With GazeAttentionNet, we achieve an excellent result on the GazeCapture dataset. The average errors of mobile phones and tablets are 1.67 cm and 2.37 cm.
Haoxian Huang, Luqian Ren, Yinwei Zhan, Qieshi Zhang, Jujian Lv
ICASSP5
2022 Indoor Target-Driven Visual Navigation based on Spatial Semantic Information
abstract
Target-driven visual navigation is a widely focused learning-based approach in the field of computer vision. However, it faces two major challenges: poor generalization ability to unknown scenes, and poor navigation performance for increased number of scenes. In this paper, an end-to-end target-driven visual navigation method, which uses Spatial Semantic Information (SSI) to navigate the agent to the target, is presented. To fully integrate the spatial and semantic information in the scene, visual information is encoded into an 8-D spatial context vector. In addition, the size of detected bounding box is used to improve the reward function in end-to-end learning to solve the problem of sparse rewards. Experiments in interactive environment dataset AI2-THOR show that compared with state-of-the-art approaches, our approach has a higher success rate and a better route to target.
Jiaojie Yan, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren
ICIP2
2022 EEP-Net: Enhancing Local Neighborhood Features and Efficient Semantic Segmentation of Scale Point Clouds
Fuxiang Wu, Qieshi Zhang, Ziliang Ren, Jun Cheng 0002
PRCV (3)3
2022 Dual-stream cross-modality fusion transformer for RGB-D action recognition
Zhen Liu 0049, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Chengqun Song
Knowl. Based Syst.5
2022 Cross-Modality Compensation Convolutional Neural Networks for RGB-D Action Recognition
abstract
RGB-D-based human action recognition has attracted much attention recently because it can provide more complementary information than a single modality. However, it is difficult for two modalities to effectively learn spatial-temporal information from each other. To facilitate information interaction between different modalities, a cross-modality compensation convolutional neural network (ConvNet) is proposed for human action recognition, which enhances the discriminative ability by jointly learning compensation features from the RGB and depth modalities. Moreover, we design a cross-modality compensation block (CMCB) to extract compensation features from the RGB and depth modalities. Specifically, CMCB is incorporated into two typical network architectures, ResNet and VGG, to verify the ability to improve the performance of our model. The proposed architecture has been evaluated on three challenging datasets: NTU RGB+D 120, THU-READ and PKU-MMD. We experimentally verify that our proposed model with CMCB is effective for different input types, such as pairs of raw images and dynamic images constructed from the entire RGB-D sequence, and the experimental results show that the proposed framework achieves state-of-the-art performance on all three datasets.
Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Fusheng Hao
IEEE Trans. Circuits Syst. Video Technol.3
2022 Time-Varying Trajectory Tracking Formation H∞ Control for Multiagent Systems With Communication Delays and External Disturbances
abstract
Time-varying formation (TVF) and trajectory tracking$H_{\infty }$control problem of multiagent systems (MASs) subject to communication delays and external disturbances under the directed communication topology is studied. This article’s objective is for all agents to attain the desired TVF and track the pregiven formation center trajectory simultaneously. First, a distributed TVF and trajectory tracking control protocol employing neighborhood interaction information is developed in the presence of communication delays. Second, since the Laplacian matrix of a graph can be decomposed into the product of two specific matrices, the TVF and trajectory tracking$H_{\infty }$control problem is converted into the lower dimension asymptotic stability problem of a closed-loop system by applying an appropriate variable conversion. Third, a Lyapunov–Krasovskii functional is constructed to analyze the stability of MASs. Sufficient conditions are obtained in the form of linear matrix inequalities (LMIs) to ensure the completion of the TVF and formation center trajectory tracking of MASs. In the meantime, the maximum allowable communication delay can be calculated by the LMIs. Finally, the results of numerical simulations are presented to verify the validity of the approach this article proposes.
Jun Cheng 0002, Yuhang Kang, Bin Xin 0002, Qieshi Zhang, Shaolei Zhou
IEEE Trans. Syst. Man Cybern. Syst.4
2021 MFPN-6D : Real-time One-stage Pose Estimation of Objects on RGB Images
abstract
6D pose estimation of objects is an important part of robot grasping. The latest research trend on 6D pose estimation is to train a deep neural network to directly predict the 2D projection position of the 3D key points from the image, establish the corresponding relationship, and finally use Pespective-n-Point (PnP) algorithm performs pose estimation. The current challenge of pose estimation is that when the object texture-less, occluded and scene clutter, the detection accuracy will be reduced, and most of the existing algorithm models are large and cannot take the real-time requirements. In this paper, we introduce a Multi-directional Feature Pyramid Network, MFPN, which can efficiently integrate and utilize features. We combined the Cross Stage Partial Network (CSPNet) with MFPN to design a new network for 6D pose estimation, MFPN-6D. At the same time, we propose a new confidence calculation method for object pose estimation, which can fully consider spatial information and plane information. At last, we tested our method on the LINEMOD and Occluded-LINEMOD datasets. The experimental results demonstrate that our algorithm is robust to textureless materials and occlusion, while running more efficiently compared to other methods.
Penglei Liu, Qieshi Zhang, Jin Zhang 0013, Fei Wang 0066, Jun Cheng 0002
ICRA2
2021 VGG-CAE: Unsupervised Visual Place Recognition Using VGG16-Based Convolutional Autoencoder
Zhenyu Xu 0014, Qieshi Zhang, Fusheng Hao, Ziliang Ren, Yuhang Kang, Jun Cheng 0002
PRCV (2)2
2021 Segment spatial-temporal representation and cooperative learning of convolution neural networks for multimodal-based action recognition
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Fusheng Hao, Xiangyang Gao
Neurocomputing2
2021 Data Augmentation and Dense-LSTM for Human Activity Recognition Using WiFi Signal
abstract
Recent research has devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi coverage area could interfere with wireless signal propagation, that manifested as unique patterns for activity recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of two major challenges. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carries substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. Since only recording activities of limited subjects in a certain speed and scale, recent works commonly have a moderate amount of activity data for training the recognition model. The small-size data could often incur the overfitting issue that negative affect the traditional classification model. To address these challenges, we propose a WiFi-based human activity recognition system that synthesizes variant activities data through eight channel state information (CSI) transformation methods to mitigate the impact of activity inconsistency and subject-specific issues, and also design a novel deep-learning model that caters to the small-size WiFi activity data. We conduct extensive experiments and show synthetic data improve performance by up to 34.6% and our system achieves around 90% of accuracy with well robustness in adapting to small-size CSI data.
Jin Zhang 0013, Fuxiang Wu, Bo Wei 0003, Qieshi Zhang, Hui Huang 0014, Syed Wajid Ali Shah, Jun Cheng 0002
IEEE Internet Things J.4
2021 Multi-modality learning for human action recognition
Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Pengyi Hao, Jun Cheng 0002
Multim. Tools Appl.2
2020 Blockchain-Based Multi-Role Healthcare Data Sharing System
abstract
Blockchain has unique advantages in data privacy protection and data integrity. We can solve many security problems, such as the single point of failure and data sharing in the current centralized system through blockchain approach. Existing studies have demonstrated that the application of blockchain in the medical system could improve the patient's medical experience. However, we found that the blockchain-based medical systems didn't consider the problems of long insurance claim cycle and complicated procedures. There are also very few healthcare systems that provide targeted sharing protocols for medical data and personal health data. In this paper, we propose a multi-role healthcare data sharing system framework based on blockchain. In this system, we reduce the storage cost through the collaborative storage of blockchain and IPFS. Then we design a smart contract on insurance to help patients achieve automatic insurance claim. In addition, we design two different sharing protocols to realize the fine management of personal data. The system analysis shows that our proposed blockchain-based multi-role healthcare data sharing system can effectively address the actual needs of users and has perfect performance in data storage, privacy protection, insurance claims and personal data management.
Yao Yu 0002, Qieshi Zhang, Wenjian Hu, Shumei Liu
HealthCom3
2020 Link Fault Repair Algorithm of Wearable Wireless Sensor Networks based on Polygon Fermat Point
abstract
With the development of wearable wireless sensor network technology, the technology has been applied in more and more fields, such as military, medical rescue, transportation and so on. Especially in the field of medical rescue, because of the advantages of wearable wireless sensor network, such as portability, fast and flexible networking, it has a very high application prospect. However, due to the particularity of medical rescue scene, the communication quality of wearable wireless sensor network is facing a huge challenge. When there is a fault, it needs to repair the fault link timely and accurately. In this paper, a link fault repair algorithm based on polygon Fermat point (LFRA) is proposed. Firstly, the algorithm can repair the fault link by inserting the relay node after the fault occurs. Secondly, the algorithm considers the energy consumption of the nodes involved in the repair process, eliminates the nodes that do not meet the energy requirements, and avoids multiple repairs due to lack of energy. The simulation results show that the proposed algorithm has the advantages of fewer inserted relay nodes and longer maximum communication time between nodes.
Lincong Zhang, Ce Zhang 0003, Kefeng Wei, Qieshi Zhang
HealthCom4
2020 ST-LSTM: Spatio-Temporal Graph Based Long Short-Term Memory Network For Vehicle Trajectory Prediction
abstract
Autonomous vehicles need the ability to predict the trajectory of surrounding vehicles, so as to make a rational decision planning, improve driving safety and ride comfort. In this paper, a new hierarchical Long Short-Term Memory (LSTM) based on Spatio-Temporal (ST) graph is proposed for vehicle trajectory prediction. Our ST-LSTM uses three layers of different LSTMs to capture the information of spatial, temporal and trajectory data, and LSTM-based encoder-decoder model as a whole, which is capable of accurately predicting future trajectories for vehicles on the highway. Our model trained and validated on the publicly available NGSIM US-101 and I-80 datasets. In comparison to state-of-art methods, our method could achieve a more accurate prediction trajectory over 5s time horizon.
Guangxi Chen, Qieshi Zhang, Ziliang Ren, Xiangyang Gao, Jun Cheng 0002
ICIP3
2019 Accelerated Detail-Enhanced Ambient Occlusion
abstract
Ambient Occlusion (AO) is a technique to approximate the effect of environment lighting and add realism to a scene by accentuating surface details and adding soft shadows, which is widely used in multimedia applications. Neural Network Ambient Occlusion (NNAO) is a pioneer in introducing deep learning to accurate and real-time AO, but it has two limitations: 1) performance bottleneck under excessive amount of samples ; 2) low contrast and blurred edges leading to unreal effect. To overcome these two limitations, we propose Accelerated Detail-enhanced Ambient Occlusion (ADAO) method based on NNAO by adopting three image processing methods: 1) spiral sampling in screen space; 2) contrast enhancement of AO map; 3) normal-depth edge preserving bilateral filtering. Experimental results show that the proposed method is over 2 times faster than NNAO and produces shadows with more realistic details.
Yinwei Zhan, Jujian Lv, Qieshi Zhang, Wenxin Yu 0001
ICIP5
2019 End-to-End Panoptic Segmentation with Pixel-Level Non-Overlapping Embedding
abstract
Recent panoptic segmentation even instance segmentation methods usually rely on the region-based method or highly-specialized combination with heuristics module, followed by post-processing techniques. While most of the recent methods neglect low-fill rate linear objects and cannot recognize pixels located in bounding box margins. We propose a branched, end-to-end trainable multi-task architecture focusing on pixel-level grouping problems for panoptic segmentation. The embedding branch regress pixels into an embedding space, so that pixels from the same group are at close range while those from different groups have a specified margin. Every pixel can be considered in an image without overlapping. And semantic branch produces best seed scores with labels as clustering center. The further-embedding branch disentangles each pixel in pixel embedding space. Thus, we are able to segment both thing and stuff classes, and explain all the pixels in the image. We obtain state-of-the-art results on Pascal VOC2012 and Cityscapes.
Qieshi Zhang, Jun Cheng 0002, Cong Bai, Pengyi Hao
ICME2
2019 WiEnhance: Towards Data Augmentation in Human Activity Recognition Using WiFi Signal
abstract
Recent research have devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi spectrum could interfere wireless signal propagation which manifested as unique patterns for activities recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of a major challenge. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carry substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. To address this challenge, we propose WiEnhance, a WiFi based activity recognition system that synthesize variant activities data and mitigate the impact of activity inconsistency and subject-specific issues. We conduct extensive experiments and show an average 15.6% performance improvement on activity recognition.
Jin Zhang 0013, Fuxiang Wu, Wen Hu 0001, Qieshi Zhang, Weitao Xu, Jun Cheng 0002
MSN4
2018 UFSM VO: Stereo Odometry Based on Uniformly Feature Selection and Strictly Correspondence Matching
abstract
Robust visual feature plays a critical role in improving camera localization performance. However, it will cost much computation time for feature extracting and matching, such as SIFT or SURF. In this paper, we present a novel visual odometry (VO) algorithm based on stereo image sequences by performing uniformly feature selection and strict correspondence matching. Firstly, the stable and uniform feature selection is performed by setting adaptive feature thresholds and selecting limited number of features in each local region. Secondly, the precise correspondence matching is achieved by double verification based on the motion model. Finally, the translation vector and rotation matrix of camera are computed, the five-point method is combined with RANSAC-based outlier rejection scheme for initial rotation estimation. And then all inliers are used for minimizing reprojection error to get final camera pose. The experimental results show that the proposed method can achieve the average translational error lower than 1.16% with 12Hz on the public KITTI dataset [1].
Liangliang Pan, Jun Cheng 0002, Qieshi Zhang
ICIP3
2018 Co-consistent Regularization with Discriminative Feature for Zero-Shot Learning
Yanling Tian, Qieshi Zhang, Jun Cheng 0002, Pengyi Hao
ICONIP (1)3
2018 Recursive Inception Network for Super-Resolution
abstract
In this paper, we propose a novel network for super-resolution and achieve the state-of-the-art performance with limited parameters. Inspired by the previous methods, we use ResNet to learn the residual part of the input patches. In addition, we introduce an inception-like structure that helps to extract features and a weight sharing mechanism is utilized among these inception blocks. By cascading multi-scale filters with separate paths in a deep network, the proposed method can fully exploit the contextual information over large image regions. Besides, the residual learning module makes the training phase easy to converge. Extensive experiments demonstrate that the proposed method can achieve the same performance with fewer parameters compared with the previous state-of-the-art methods.
Xiaojun Wu 0002, Wuyang Shui, Shiqi Guo, Hao Fei 0002, Qieshi Zhang
ICPR8
2018 Selective Multi-Convolutional Region Feature Extraction based Iterative Discrimination CNN for Fine-Grained Vehicle Model Recognition
abstract
With the rapid rise of computer vision and driverless technology, vehicle model recognition plays a huge role in the common application and industry field. While fine-grained vehicle model recognition is often influenced by multi-level information, such as the image perspective, inter-feature similarity, vehicle details. Furthermore, pivotal regions extraction and fine-grained feature learning have become a vital obstacle to the fine-grained recognition of vehicle models. In this paper, we propose an iterative discrimination CNN (ID-CNN) based on selective multi-convolutional region (SMCR) feature extraction. The SMCR features, which consist of global and local SMCR features, are extracted from the original image with higher activation response value. As for ID-CNN, we use the global and local SMCR features iteratively to localize deep pivotal features and concatenate them together into a fully-connected fusion layer to predict the vehicle categories. We get better results and improve the accuracy to 91.8% on Stanford Cars-196 dataset and to 96.2% on CompCars dataset.
Yanling Tian, Qieshi Zhang, Xiaojun Wu 0002
ICPR3
2016 Learning discriminative and shareable patches for scene classification
abstract
This paper addresses the problem of scene classification and proposes learning discriminative and shareable patches (LDSP) method. The main idea of learning discriminative and shareable patches is to discover patches that exhibit both large between-class dissimilarity (discriminative) and large within-class similarity (shareable). A novel and efficient re-clustering, based on co-occurrence relationship of first-step clustering, is proposed and conducted to further enhance the visual similarity of patches within each cluster. In order to establish appropriate criteria for selecting desired patches, a condensed representation of image features called feature epitome is introduced. In the classification, a patch feature involving pre-trained convolutional neural network model is investigated. The experimental result outperforms existing single-feature methods on MIT 67 scene benchmark in term of mean Accuracy Precision.
Shoucheng Ni, Qieshi Zhang
ICASSP2
2016 A novel color space based on RGB color barycenter
abstract
Color space is one of the bases in the image processing area. Suitable color space can give the suitable description of colors for variant processing. However, in the image processing area, the existing color space cannot show the suitable distribution in color and lightness. In this paper, a novel color space based on RGB color barycenter (RGB-CB) is proposed to describe the color and lightness more intuitively. To prove the effectiveness of the proposed color space, YUV, HSV, L*a*b*, and IPT color spaces are discussed and compared. Experimental results show the proposed color space can perform better effect than other color space in image processing.
Qieshi Zhang
ICASSP1
2016 Adaptive sampling and wavelet tree based compressive sensing for MRI reconstruction
abstract
Magnetic Resonance Imaging (MRI) has been widely used in medical diagnose because of its non-invasive manner and excellent depiction of soft-tissue changes. Recently, the compressive sensing (CS) theory has been applied to reconstruct the MR image from highly down-sampled k-space data, which can reduce the scanning duration. To obtain useful information as much as possible with the same sampling rate, a weighted sampling strategy is studied. Moreover, based on the advantage of CS, a Wavelet tree based reconstruction approach is proposed. The experimental results demonstrate that the proposed method is preferable to other methods.
Qieshi Zhang
ICIP1
2015 Disparity refinement with stability-based tree for stereo matching
abstract
This paper proposes a disparity refinement method with stability-based tree. By developing stability-based tree to evaluate and reconstruct support regions for error parts, the proposed method achieves effective performance in removing outliers. This approach further improves the quality of raw disparity map in stereo matching, which makes the local methods results comparable to the global ones. Experiments exhibit that the proposed method reduces more than 70% aggregation time compared with traditional tree method without loss of accuracy. It also outperforms existing disparity refinement methods in removing large error parts.
Yuhang Ji, Qieshi Zhang, Kenjiro Sugimoto
Intelligent Vehicles Symposium2
2015 Fisheye Image Correction Based on Straight-Line Detection and Preservation
abstract
Fisheye lenses are widely used when the users want to capture the image/video with wide field of view (FoV) which is particularly suited to surveillance monitoring and vehicle camera. However, no projection from the actual scene in wide FoV image can avoid the distortion. If this problem cannot be solved, the fisheye image will difficult be used for object detection or analysis due to the distorted shapes of the scene objects. To correct this problem and obtain the natural-looking image, a two-step correction approach is proposed. Firstly, adaptive latitude and longitude correction are presented and the Hough transform is used to detect and estimate the straight-line. Secondly, the straight-line preserving and orientation, consistency based optimization is examined to obtain the final correction result. To compare the effectiveness of the proposed method, some fisheye correction methods are discussed. The experimental results demonstrate that the proposed method can obtain the coherent natural-looking.
Qieshi Zhang
SMC1
2014 Optimized curvelet-based empirical mode decomposition
abstract
The recent years has seen immense improvement in the development of signal processing based on Curvelet transform. The Curvelet transform provide a new multi-resolution representation. The frame elements of Curvelets exhibit higher direction sensitivity and anisotropic than the Wavelets, multi-Wavelets, steerable pyramids, and so on. These features are based on the anisotropic notion of scaling. In practical instances, time series signals processing problem is often encountered. To solve this problem, the time-frequency analysis based methods are studied. However, the time-frequency analysis cannot always be trusted. Many of the new methods were proposed. The Empirical Mode Decomposition (EMD) is one of them, and widely used. The EMD aims to decompose into their building blocks functions that are the superposition of a reasonably small number of components, well separated in the time-frequency plane. And each component can be viewed as locally approximately harmonic. However, it cannot solve the problem of directionality of high-dimensional. A reallocated method of Curvelet transform (optimized Curvelet-based EMD) is proposed in this paper. We introduce a definition for a class of functions that can be viewed as a superposition of a reasonably small number of approximately harmonic components by optimized Curvelet family. We analyze this algorithm and demonstrate its results on data. The experimental results prove the effectiveness of our method.
Renjie Wu 0002, Qieshi Zhang
ICMV2
2014 Disparity estimation from monocular image sequence
abstract
This paper proposes a novel method for estimating disparity accurately. To achieve the ideal result, an optimal adjusting framework is proposed to address the noise, occlusions, and outliners. Different from the typical multi-view stereo (MVS) methods, the proposed approach not only use the color constraint, but also use the geometric constraint associating multiple frame from the image sequence. The result shows the disparity with a good visual quality that most of the noise is eliminated, the errors in occlusion area are suppressed and the details of scene objects are preserved.
Qieshi Zhang
ICMV1
2014 Interactive object segmentation using color similarity based nearest neighbor regions mergence
abstract
An effective object segmentation is an important task in computer vision. Due to the automatic image segmentation is hard to segment the object from natural scenes, the interactive approach becomes a good solution. In this paper, a color similarity measure based region mergence approach is proposed with the interactive operation. Some local regions, which belong to the background and object, need to be interactively marked respectively. To judge whether two adjacent regions need to be merged or not, a color similarity measure is proposed with the help of mark. Execute merging operation based on the marks in background and the two regions with maximum similarity need to be merged until all candidate regions are examined. Consequently, the object is segmented by ignoring the merged background. The experiments prove that the proposed method can obtain more accurate result from the natural scenes.
Qieshi Zhang
ICMV2
2014 Sparse decomposition learning based dynamic MRI reconstruction
abstract
Dynamic MRI is widely used for many clinical exams but slow data acquisition becomes a serious problem. The application of Compressed Sensing (CS) demonstrated great potential to increase imaging speed. However, the performance of CS is largely depending on the sparsity of image sequence in the transform domain, where there are still a lot to be improved. In this work, the sparsity is exploited by proposed Sparse Decomposition Learning (SDL) algorithm, which is a combination of low-rank plus sparsity and Blind Compressed Sensing (BCS). With this decomposition, only sparsity component is modeled as a sparse linear combination of temporal basis functions. This enables coefficients to be sparser and remain more details of dynamic components comparing learning the whole images. A reconstruction is performed on the undersampled data where joint multicoil data consistency is enforced by combing Parallel Imaging (PI). The experimental results show the proposed methods decrease about 15~20% of Mean Square Error (MSE) compared to other existing methods.
Peifei Zhu, Qieshi Zhang
ICMV2
2008 Automatic road sign detection method based on Color Barycenters Hexagon model
abstract
Road sign detection is one of the major concerned topics in the field of driving safety and intelligent vehicle. In this paper, a novel model based on Color Barycenters Hexagon (CBH) is proposed and used to detect road sign usefully. In CBH model, full color images are calculated the color barycenters and get the barycenters region, then automatic select the idea threshold curves to separate the region of interest (ROI) of barycenters aiming to detect the road sign. Because of the practically images have many noise, and the existing color space cannot separate the ROI ideally. The proposed CBH model can thresholding the principal color of ROI and have high robust. With suitably thresholding and operations, road sign on various scene images can be detected.
Qieshi Zhang
ICPR1