Rui Liu 0015

dblp:42/469-15 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0002-5200-640XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoCoDiff: Modality-Aware Conditional Diffusion Model for 3D Brain Tumor Segmentation
Sijie Guo, Jing Dong 0009, Rui Liu 0015, Xiaopeng Wei
ICPR (5)5
2026 An adaptive multimodal semantic knowledge enhanced framework for sarcasm detection
Jing Dong 0009, Yu Sui, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang
Expert Syst. Appl.6
2026 Fine-grained face personalisation using a text-guided multi-attribute embedded diffusion model
Jing Dong 0009, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang
Expert Syst. Appl.6
2026 Sems-net: a semantic-enhanced and modality-shared collaborative network for multimodal sentiment analysis
Ruixia Duan, Zherui Li 0006, Rui Liu 0015, Jing Dong 0009
Multim. Syst.4
2026 Enhanced medical image segmentation via wavelet-deformable attention networks
Rui Liu 0015, Jing Dong 0009, Xiaopeng Wei
Vis. Comput.2
2025 Lightweight 2D Human Pose Estimation Based on Multi-scale Fusion and Attention Mechanism
Boyu Qi, Jing Dong 0009, Xiaoyong Fang, Rui Liu 0015
ICIC (11)6
2025 A sequential mixing fusion network for enhanced feature representations in multimodal sentiment analysis
Qiang Zhang 0008, Jing Dong 0009, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015
Knowl. Based Syst.6
2025 HSE-GNN: A hierarchical skeleton embedded graph neural network for 3D human pose estimation
Jing Dong 0009, Hui Fang 0003, Rui Liu 0015, Yu Sui
Pattern Recognit. Lett.4
2024 MCFNet: Multi-Attentional Class Feature Augmentation Network for Real-Time Scene Parsing
abstract
For real-time scene parsing tasks, capturing multi-scale semantic features and performing effective feature fusion is crucial. However, many existing solutions ignore stripe-shaped things like poles, traffic lights and are so computationally expensive that cannot meet the high real-time requirements. This article presents a novel model, the Multi-Attention Class Feature Augmentation Network (MCFNet) to address this challenge. MCFNet is designed to capture long-range dependencies across different scales with low computational cost and to perform a weighted fusion of feature maps. It features the BAM (Strip Matrix Based Attention Module) for extracting strip objects in images. The BAM module replaces the conventional self-attention method using square matrices with strip matrices, which allows it to focus more on strip objects while reducing computation. Additionally, MCFNet has a parallel branch that focuses on global information based on self-attention to avoid wasting computation. The two branches are merged to enhance the performance of traditional self-attention modules. Experimental results on two mainstream datasets demonstrate the effectiveness of MCFNet. On the Camvid and Cityscapes test sets, MCFNet achieved 207.5 FPS/73.5% mIoU and 136.1 FPS/71.63% mIoU, respectively. The experiments show that MCFNet outperforms other models on the Camvid dataset and can significantly improve the performance of real-time scene parsing tasks.
Xizhong Wang, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008
ACM Trans. Multim. Comput. Commun. Appl.2
2022 SCFNet: A Spatial-Channel Features Network Based on Heterocentric Sample Loss for Visible-Infrared Person Re-identification
Rui Liu 0015, Jing Dong 0009
ACCV (2)2
2022 Research on Depth-Adaptive Dual-Arm Collaborative Grasping Method
Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008
CollaborateCom (2)3
2022 SAMKR: Bottom-up Keypoint Regression Pose Estimation Method Based On Subspace Attention Module
abstract
As a hot research issue in computer vision, 2D human pose estimation plays an important role in human-computer interaction, intelligent monitoring, 3D human pose estimation and so on. Aiming at the problem of human scale inconsistency in the situation of multi-person, a keypoint regression method based on subspace attention module (SAMKR) is proposed in this paper for the 2D human pose estimation. Firstly, each keypoint is divided into independent regression branches, and then the feature mapping in each keypoint regression branch is evenly divided into a specified number of feature mapping subspaces, and different attention mappings are derived for each feature mapping subspace. By learning different attention maps in each feature subspace, multi-scale feature representation can be effectively improved. The experimental results show that SAMKR achieved 74.4 AP score on the CrowdPose test set, which may lead to an improvement of + 7.1AP, and reached 70.4 AP score on the COCO test-dev data set, which was 0.4AP higher than the baseline.
Rui Liu 0015, Qiang Zhang 0008
IJCNN3
2022 A Novel Movement-supported HRI Framework for Humanoid Robots
abstract
Current research related to human-robot interaction (HRI) of bipedal humanoid robots often assumes that the robot is in a standing stationary state, i.e., the relative position of the robot does not change, and rarely considers the effect of lower limb movement on interaction. However, HRI in the real world does not assume a moving or stationary state of the robot, and the equilibrium perturbations caused by movement can prevent HRI from functioning properly. In this paper, we propose a movement supported humanoid robot interaction method that empowers the robot to move stably while achieving HRI. First, a reinforcement learning-based neural network is run offline to generate interaction actions that satisfy the equilibrium constraint and support movement, and then an intention recognition network is introduced to run the movement-supported HRI framework online. It is demonstrated that the training method proposed in this paper can enable a bipedal robot to achieve a variety of interactive actions while moving stably.
Jing Dong 0009, Rui Liu 0015, Xiaopeng Wei, Qiang Zhang 0008
IJCNN4
2022 Semantic Image Synthesis via Location Aware Generative Adversarial Network
abstract
Semantic image synthesis aims to synthesize photo-realistic images through the given semantic segmentation masks. Most existing models use conditional batch normalization (CBN) to regulate normalization activation by spatially varying modulation parameters. It can prevent semantic information from being eliminated during normalization. But the modulation parameters in CBN lack location constraint, resulting in the lack of structural information in the synthetic image. And CBN is highly dependent on the batch size. To address these limitations, we propose location aware conditional group normalization (LACGN) and construct a location aware generative adversarial network (LAGAN) based on this method. LACGN can learn spatial location aware information in a weakly supervised manner that relies on the current image synthesis process to guide transformations spatially. It allows the synthetic image to have more structural information and detailed features. At the same time, group normalization(GN) replace the traditional BN to eliminate the dependence on batch size. Extensive experiments show that LAGAN is better than other methods.
Rui Liu 0015, Jing Dong 0009, Wanshu Fan
MSN2
2022 DEANet: A Real-Time Image Semantic Segmentation Method Based on Dual Efficient Attention Mechanism
Rui Liu 0015, Jing Dong 0009
WASA (2)2
2021 A Novel Gaze-Point-Driven HRI Framework for Single-Person
Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009
CollaborateCom (1)6
2021 A Novel and Efficient Distance Detection Based on Monocular Images for Grasp and Handover
Dianwen Liu, Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009
CollaborateCom (1)6
2021 Human-robot Interaction Method Combining Human Pose Estimation and Motion Intention Recognition
abstract
Although human pose estimation technology based on RGB images is becoming more and more mature, most of the current mainstream methods rely on depth camera to obtain human joints information. These interaction frameworks are affected by the infrared detection distance so that they cannot well adapt to the interaction scene of different distance. Therefore, the purpose of this paper is to build a modular interactive framework based on RGB images, which aims to alleviate the problem of high dependence on depth camera and low adaptability to distance in the current human-robot interaction (HRI) framework based on human body by using advanced human pose estimation technology. To enhance the adaptability of the HRI framework to different distances, we adopt optical cameras instead of depth cameras as acquisition equipment. Firstly, the human joints information is extracted by a human pose estimation network. Then, a joints sequence filter is designed in the intermediate stage to reduce the influence of unreasonable skeletons on the interaction results. Finally, a human intention recognition model is built to recognize the human intention from reasonable joints information, and drive the robot to respond according to the predicted intention. The experimental results show that our interactive framework is more robust in the distance than the framework based on depth camera and is able to achieve effective interaction under different distances, illuminations, costumes, customers, and scenes.
Yalin Cheng, Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008
CSCWD3
2021 Asymmetric Anomaly Detection for Human-Robot Interaction
abstract
Security in human-robot interaction is the focus of research in this field. Rapid detection of abnormal events that may cause danger in the interaction process can effectively reduce the probability of occurrence of danger. In general anomaly detection methods, 2D or 3D convolutional autoencoders are widely used for anomaly detection. Among them, 2D convolutional autoencoders are with good real-time performance and lower detection accuracy, while 3D convolutional autoencoders are with higher detection accuracy and insufficient real-time performance. In order to ensure realtime performance and obtain higher accuracy, an end-to-end asymmetric convolutional autoencoder network (ACANet) using both 2D and 3D convolutions is designed. Specifically, 3D convolution is used to build the encoder to learn comprehensive information in continuous input frames, and 2D convolution is used to build the decoder to model the information fast, a dimensional alignment module is constructed to connect the encoder and the decoder while avoiding a large number of calculations in the latent space of the 3D features output by the encoder, and the skip connections module is used to obtain accurate predictions. Anomaly detection can then be completed by evaluating the differences between results predicted by the ACANet and real frames. The experimental results show that our method achieves competitive accuracy on mainstream datasets and at the same time obtains the fastest speed. Compared with mainstream methods, this method is more suitable for anomaly detection tasks in human-robot interaction.
Rui Liu 0015, Yingkun Hou, Qiang Zhang 0008, Xiaopeng Wei
CSCWD3
2021 Deep Reinforcement Learning Visual Navigation Model Integrating Memory-prediction Mechanism
abstract
Deep reinforcement learning (DRL) has been widely used in the field of visual navigation. However, due to the lack of adaptability of DRL to the new tasks, the generalization ability of current visual navigation model using DRL is not desired. In order to improve this deficiency, we introduce the memory-prediction mechanism. By enhancing the memory of the scene, and combining the past experience of navigation to predict the next state, a more reasonable action can be obtained. First, we pass the image features extracted during the navigation process to an LSTM, and use LSTM to memorize the scene information in the image features. Then, we combine all the information (including state, target, and action) of each time step in the navigation process, and pass the historical information of multiple time steps to another LSTM to predict the next state. The action performed by the robot is determined by the predicted state. We use the AI2-THOR framework to carry out experiments. The results show that the proposed method can improve the navigation performance of the DRL visual navigation model and improve its adaptability to new tasks.
Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008
CSCWD3
2021 ASFNet: Adaptive multiscale segmentation fusion network for real-time semantic segmentation
abstract
Abstract Recently, the development of deep learning has facilitated continuous progress in the field of computer vision. Pixel‐level semantic segmentation serves as a fundamental task in computer vision. It achieves significant results by connecting wider and deeper backbone networks and building fine‐grained segmentation heads. However, applications such as self‐driving cars are more critical to the computational speed of the algorithms. The trade‐off between accuracy and real‐time performance of existing algorithms is still a challenging task. To address this challenge, this article proposes an adaptive multiscale segmentation fusion network to fuse multiscale contextual, which designs an adaptive multiscale segmentation fusion module based on an attention mechanism. Using segmentation fusion instead of feature fusion, the multiscale segmentation results are aggregated to obtain more precise segmentation results. The final results achieved 70.9% mIoU of accuracy in the Cityspace test set, processing images at 61 FPS when the input is 1024 × 2048. In addition, when adjusting the input size to 512 × 1024, the images are processed at 185 FPS.
Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Comput. Animat. Virtual Worlds2
2020 Efficient Attention Calibration Network for Real-Time Semantic Segmentation
abstract
In recent years, the attention mechanism has been widely used in computer vision. Semantic segmentation, as one of the fundamental tasks of computer vision, has been subject to tremendous development as a result. But because of its huge computing overhead, attention-based approaches are difficult to use for real-time applications such as self-driving. In this paper, we propose a self-calibration method baesd on self-attentiion that successfully applies the attention mechanism to real-time semantic segmentation. Specifically, a spatial attention module to adjust the edges of the coarse segmentation results which gained from the real-time semantic segmentation backbone network, and obtain more granular segmentation results. We refer to this method as the Efficient Attentional Calibration Network (EACNet). Experiments on the Cityscapes dataset validate the efficiency and performance of the method. With the high-resolution input and without any post-processing, EACNet achieved 72.4% mIoU of accuracy while running at 116.9 FPS. Compared to other state-of-the-art methods for real-time semantic segmentation, our network gained a better balance between performance and speed.
Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
ACML2