EDBT 2026 Demo / reviewers in the wild / expert
Ruyi Liu 0001
dblp:168/2163-1
· DBLP profile ↗
32ranked-venue papers
8as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly semantic-guided skeleton feature distillation for human action recognition
Ruyi Liu 0001, Qiguang Miao, Wentian Xin, Xiangzeng Liu, Long Li 0005 |
Expert Syst. Appl. | 1 |
| 2026 | A Systematic Review of Skeleton-Based Action Recognition: Methods, Challenges, and Future DirectionsabstractHuman action recognition (HAR), which aims to recognize and understand individual actions and intentions, has rapidly become a research hotspot in computer vision. Compared with other data modalities, skeleton data offers more efficient node semantics and more coherent spatio-temporal motion patterns, effectively reducing the impact of lighting and background changes. In recent years, many researchers have focused on skeleton-based action recognition methods and have made significant progress. However, we believe that the current skeleton-based action recognition methods still face three major challenges: 1) reducing reliance on expensive labeled data while maintaining model performance; 2) enabling the model to understand and recognize new behavior classes with a limited number of samples; and 3) addressing the challenges posed by the lack of skeleton information in single-modality spatio-temporal motion representation learning. Based on these challenges, we conduct a comprehensive review of the existing skeleton-based action recognition methods. Additionally, we provide an extensive review and analysis of publicly available action recognition datasets. This review aims to offer researchers a comprehensive perspective, stimulate more innovative ideas, and promote the application and breakthrough of skeleton action recognition in a wider range of computer vision tasks. Ruyi Liu 0001, Yuzhi Hu, Wentian Xin, Qiguang Miao, Shuai Wu 0001, Long Li 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | Boosting Semi-Supervised Medical Image Segmentation Through Inter-Instance Information ComplementarityabstractThe acquisition of expert-annotated data remains a critical bottleneck for medical image segmentation, thereby constraining the clinical applicability of highly accurate models. Crucially, despite this scarcity of labeled data, the intrinsic homogeneity in human anatomy across the cohort provides a fundamental basis (or: a promising leverage point) for enhancing model generalization and training efficiency by exploiting inter-instance anatomical complementarity. In this study, we propose a novel semi-supervised approach for medical image segmentation that fully exploits this inter-instance complementarity. The proposed model operates at two levels, integrating a sophisticated copy-paste augmentation module (CPAM) and a trainable region calibration mechanism (TRCM) within the simple mean teacher (MT) framework. Specifically, CPAM is a carefully designed copy-paste strategy that facilitates the exchange of informative regions between samples, thereby enhancing the diversity and robustness of the training data. TRCM leverages the predictions from labeled regions to guide and calibrate the trainable regions in unlabeled data. The calibrated regions typically yield high-quality pseudo-labels, which effectively improve model training. CPAM and TRCM work synergistically, complementing each other to enhance model performance. Experiments on diverse medical image datasets-including LA, ACDC, BraTS2019, and Pancreas-NIH-covering both MRI and CT modalities demonstrate the robust efficacy of our proposed model. In settings with limited annotated data, the model consistently outperforms current state-of-the-art methods across multiple evaluation metrics. The code is available at https://github.com/shuaiaihang/shuaiAIMedcalLab. Shuai Wu 0001, Ruyi Liu 0001, Hang Wei 0005, Linrunjia Liu, Jie Wen 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | An Effective Interval Normalization Weighting Method for Accurate Object DetectionabstractAn effective method for improving the object detection performance is to decrease the number of false positive (NFP) detection boxes and increase the number of true positive (NTP) detection boxes. In terms of the region-based object detection framework, an appropriate sample weighting strategy can help effectively achieve this goal without causing any inference efficiency loss. However, designing a suitable weighting method is not easy, and a reasonable guiding metric and comprehensive analysis are needed. This article directly sets the NFP and NTP as the evaluation metrics and examines how some preliminary weighting methods affect these two metrics. Based on the results of our analysis, we carefully design a simple yet effective sample weighting method, referred to as the interval normalization weighting strategy (INWS). Unlike some previous works, which only view sample losses as the weighting factor (e.g., focal losses), the INWS applies both the foreground score and the intersection over union (IoU) as the weighting factors. The INWS consists of two components: the IoU interval score normalization strategy (IISNS) for negative samples and the score interval IoU normalization strategy (SIINS) for positive samples. The IISNS can effectively decrease the NFP, and the SIINS is beneficial for increasing the NTP, especially under higher IoU thresholds. Furthermore, the INWS is convenient for application to most of the existing region-based object detection models. The experimental results on the mainstream benchmarks demonstrate that our INWS can achieve consistent improvements on various baselines. Shuai Wu 0001, Chunwei Tian, Ruyi Liu 0001, Hang Wei 0005, Yong Xu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Asymmetric Dual-Encoder and Width-Dependent Feature Interaction for Road Extraction in Remote Sensing ImagesabstractRoad extraction from remote sensing images provides basic data for urban planning and intelligent transportation, aiding efficient management and decision-making. Traditional methods struggle to maintain road continuity and integrity due to varying road widths, complex shapes, uneven feature distribution, and occlusions. Despite the powerful feature extraction capabilities of CNNs in this field, their single-branch structure has limited effectiveness in handling roads of varying widths and often results in boundary blurring and disconnections in complex scenes. To address these challenges, we propose an Asymmetric Dual-encoder Feature Interaction Network (ADFINet). This model integrates the local feature extraction capabilities of CNNs with the global context modeling of Transformers. It uses dual encoders to separately extract features for narrow and wide roads and integrates them via a Dual-branch Feature Interaction Fusion Module (DFIFM). Additionally, a Feature Refinement and Redundancy Removal Module (FR3M) enhances feature expression and detail retention. Experiments demonstrate that ADFINet achieves superior performance on the DeepGlobe dataset, significantly improving road extraction accuracy and robustness. Ruyi Liu 0001, Junhong Wu, Qiguang Miao, Panshi Guo |
CW | 1 |
| 2025 | Local and Global Spatial-Temporal Transformer for skeleton-based action recognition
Ruyi Liu 0001, Feiyu Gai, Qiguang Miao, Shuai Wu 0001 |
Neurocomputing | 1 |
| 2025 | One-shot handwriting imitation via self-supervised cross spatial transformer networks
Bocheng Zhao, Guanwen Feng, Wenxing Zhang, Yunan Li 0001, Qiguang Miao, Xiangzeng Liu, Ruyi Liu 0001 |
Neurocomputing | 7 |
| 2025 | A video course enhancement technique utilizing generated talking heads
Zixiang Lu, Bujia Tian, Qiguang Miao, Kun Xie 0011, Ruyi Liu 0001, Yi-Ning Quan |
Neural Comput. Appl. | 6 |
| 2025 | SG-CLR: Semantic representation-guided contrastive learning for self-supervised skeleton-based action recognition
Ruyi Liu 0001, Wentian Xin, Qiguang Miao, Xiangzeng Liu, Long Li 0005 |
Pattern Recognit. | 1 |
| 2025 | CAETFN: Context Adaptively Enhanced Text-Guided Fusion Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) is an active research area in recent years with the exponential development of the internet and social media, which aims to recognize the speaker’s sentiment in the video consisted of text, acoustic and visual cues, and has attracted attention from many applications such as smart education, intelligent medication and social security. The predominant approaches have devoted to developing more complicated fusion strategy to learn efficient multimodal representations. However, information from these modalities usually have different contributions to MSA task. More specifically, the text modality outperforms the non-verbal modalities since its highly condensed semantic information and the maturity of the pre-trained language models. Taking full advantage of the text modality while integrating the non-verbal sentiment-relevant contextual information becomes a substantial challenge. Thus, in this paper, we propose a Context Adaptively Enhanced Text-guided Fusion Network, which is embedded in the pre-trained language model and utilizes the text modality as the guide to reduce the redundancy and exploit the sentiment-relevant information and in turn uses these information to complement itself with the non-verbal sentiment contexts. Moreover, a novelly designed non-verbal feature enhancement module is introduced to capture long-range dependencies in two directions, with the substantial removal of the redundancy and the noise. Extensive experiments on two benchmark datasets CMU-MOSI and CMU-MOSEI demonstrate the competitive performance over the state-of-the-art methods. Ruyi Liu 0001, Qiguang Miao, Di Wang 0011, Xiangzeng Liu |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Adaptive Occlusion-Aware Network for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) is a challenging task due to some of the essential features are interfered by obstacles or other pedestrians. Multi-granularity local feature extraction and recognition can effectively improve the accuracy of ReID under occlusion. However, manual segmentation methods for local features can lead to feature misalignment. Feature alignment based on pose estimation often ignores non-body details (e.g., handbags, backpacks, etc.) while increasing the complexity of the model. To address the above challenges, we propose a novel Adaptive Occlusion-Aware Network (AOANet), which mainly consists of two modules, the Adaptive Position Extractor (APE) and the Occlusion Awareness Module (OAM). In order to adaptively extract distinguishing features of body parts, APE optimizes the representation of multi-granularity features by the guidance of attention mechanism and keypoint features. To further perceive the occluded region, the OAM is developed by adaptively calculating the occlusion weights for body parts. These weights can lead to highlighting the non-occluded parts and suppressing the occluded parts, which in turn improves the accuracy in the occluded situation. Extensive experimental results confirm the advantages of our method on the MSMT17, DukeMTMC-reID, Market-1501, Occluded-Duke and Occluded-ReID datasets. The comparative results demonstrate that our method outperforms comparable methods. Especially on the Occluded-Duke dataset, our method achieved 70.6% mAP and 81.2% Rank-1 accuracy. Xiangzeng Liu, Hao Chen 0187, Qiguang Miao, Ruyi Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Adaptive Pitfall: Exploring the Effectiveness of Adaptation in Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition by exploiting the adjacency topology of body representation. However, the adaptive strategy adopted by the previous methods to construct the adjacency matrix is not balanced between the performance and the computational cost. We assume this concept ofAdaptive Trap, which can be replaced by multiple autonomous submodules, thereby simultaneously enhancing the dynamic joint representation and effectively reducing network resources. To effectuate the substitution of the adaptive model, we unveil two distinct strategies, both yielding comparable effects. (1) Optimization.Individuality and Commonality GCNs (IC-GCNs)is proposed to specifically optimize the construction method of the associativity adjacency matrix for adaptive processing. The uniqueness and co-occurrence between different joint points and frames in the skeleton topology are effectively captured through methodologies like preferential fusion of physical information, extreme compression of multi-dimensional channels, and simplification of self-attention mechanism. (2) Replacement.Auto-Learning GCNs (AL-GCNs)is proposed to boldly remove popular adaptive modules and cleverly utilize human key points as motion compensation to provide dynamic correlation support. AL-GCNs construct a fully learnable group adjacency matrix in both spatial and temporal dimensions, resulting in an elegant and efficient GCN-based model. In addition, three effective tricks for skeleton-based action recognition (Skip-Block, Bayesian Weight Selection Algorithm, and Simplified Dimensional Attention) are exposed and analyzed in this paper. Finally, we employ the variable channel and grouping method to explore the hardware resource bound of the two proposed models. IC-GCN and AL-GCN exhibit impressive performance across NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UAV-Human datasets, with an exceptional parameter-cost ratio. Qiguang Miao, Wentian Xin, Ruyi Liu 0001, Cheng Shi 0002, Chi-Man Pun |
IEEE Trans. Multim. | 3 |
| 2024 | Language-Skeleton Pre-training to Collaborate with Self-Supervised Human Action Recognition
Ruyi Liu 0001, Wentian Xin, Qiguang Miao, Yuzhi Hu, Jiahao Qi |
PRCV (7) | 2 |
| 2024 | DiffTAD: Denoising diffusion probabilistic models for vehicle trajectory anomaly detection
Chaoneng Li, Guanwen Feng, Yunan Li 0001, Ruyi Liu 0001, Qiguang Miao, Liang Chang 0003 |
Knowl. Based Syst. | 4 |
| 2024 | Exploration of Class Center for Fine-Grained Visual ClassificationabstractDifferent from large-scale classification tasks, fine-grained visual classification is a challenging task due to two critical problems: 1) evident intra-class variances and subtle inter-class differences, and 2) overfitting owing to fewer training samples in datasets. Most existing methods extract key features to reduce intra-class variances, but pay no attention to subtle inter-class differences in fine-grained visual classification. To address this issue, we propose a loss function named exploration of class center, which consists of a multiple class-center constraint and a class-center label generation. This loss function fully utilizes the information of the class center from the perspective of features and labels. From the feature perspective, the multiple class-center constraint pulls samples closer to the target class center, and pushes samples away from the most similar nontarget class center. Thus, the constraint reduces intra-class variances and enlarges inter-class differences. From the label perspective, the class-center label generation utilizes class-center distributions to generate soft labels to alleviate overfitting. Our method can be easily integrated with existing fine-grained visual classification approaches as a loss function, to further boost excellent performance with only slight training costs. Extensive experiments are conducted to demonstrate consistent improvements achieved by our method on four widely-used fine-grained visual classification datasets. In particular, our method achieves state-of-the-art performance on the FGVC-Aircraft and CUB-200-2011 datasets. Hang Yao 0001, Qiguang Miao, Chaoneng Li, Guanwen Feng, Ruyi Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Is Really Correlation Information Represented Well in Self-Attention for Skeleton-based Action Recognition?abstractTransformer has shown significant advantages by various vision tasks. However, the lack of representation of correlation information about data properties makes it difficult to match the excellent results consistent with GCNs in skeleton-based action recognition. In this paper, we propose a Topology and Frames-guided Spatial-Temporal ConvFormer Network (TF-STCFormer), which is well suited for dynamically extracting topological and inter-frame uniqueness & co-occurrence information. Three essential components make up the proposed framework: (1) Grouped Physical-guided Spatial Transformer for focusing on learning essential spatial features and physical topology. (2) Global and Focal Temporal Transformer for promoting the relationship of different joints in consecutive frames and improving the representation of discriminative key-frames. (3) Grouped Dilation Temporal Convolution for connecting the intermediate output obtained by the previous transformers in the feature channels of different dilation. Experiments on four standard datasets (NTU RGB+D, NTU RGB+D 120, NW-UCLA, and UAV-Human) demonstrate that our approach prominently outperforms state-of-the-art methods on all benchmarks. Wentian Xin, Hongkai Lin, Ruyi Liu 0001, Qiguang Miao |
ICME | 3 |
| 2023 | Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionabstractVision Transformer, which performs well in various vision tasks, encounters a bottleneck in skeleton-based action recognition and falls short of advanced GCN-based methods. The root cause is that the current skeleton transformer depends on the self-attention mechanism of the complete channel of the global joint, ignoring the highly discriminative differential correlation within the channel, so it is challenging to learn the expression of the multivariate topology dynamically. To tackle this, we present Skeleton MixFormer, an innovative spatio-temporal architecture to effectively represent the physical correlations and temporal interactivity of the compact skeleton data. Two essential components make up the proposed framework: 1) Spatial MixFormer. The channel-grouping and mix-attention are utilized to calculate the dynamic multivariate topological relationships. Compared with the full-channel self-attention method, Spatial MixFormer better highlights the channel groups' discriminative differences and the joint adjacency's interpretable learning. 2) Temporal MixFormer, which consists of Multiscale Convolution, Temporal Transformer and Sequential Holding Module. The multivariate temporal models ensure the richness of global difference expression and realize the discrimination of crucial intervals in the sequence, thereby enabling more effective learning of long and short-term dependencies in actions. Our Skeleton MixFormer demonstrates state-of-the-art (SOTA) performance across seven different settings on four standard datasets, namely NTU-60, NTU-120, NW-UCLA, and UAV-Human. Related code will be available on https://github.com/ElricXin/Skeleton-MixFormer. Wentian Xin, Qiguang Miao, Ruyi Liu 0001, Chi-Man Pun, Cheng Shi 0002 |
ACM Multimedia | 4 |
| 2023 | Auto-Learning-GCN: An Ingenious Framework for Skeleton-Based Action Recognition
Wentian Xin, Ruyi Liu 0001, Qiguang Miao, Cheng Shi 0002, Chi-Man Pun |
PRCV (1) | 3 |
| 2023 | Deep mutual learning for brain tumor segmentation with the fusion network
Qiguang Miao, Daikai Ma, Ruyi Liu 0001 |
Neurocomputing | 4 |
| 2023 | Focus on hierarchical features: Soft-weighted hierarchical features network
Hongkai Lin, Wentian Xin, Shun Chang, Qianxue Yang, Qiguang Miao, Ruyi Liu 0001, Liang Chang 0003 |
Neurocomputing | 6 |
| 2023 | Transformer for Skeleton-based action recognition: A review of recent advances
Wentian Xin, Ruyi Liu 0001, Qiguang Miao |
Neurocomputing | 2 |
| 2023 | Refined probability distribution module for fine-grained visual categorization
Qiguang Miao, Hongsheng Li 0001, Ruyi Liu 0001, Yi-Ning Quan, Jianfeng Song |
Neurocomputing | 4 |
| 2022 | Multimodal medical image fusion using gradient domain guided filter random walk and side window filtering in framelet domain
Qiguang Miao, Ruyi Liu 0001, Yang Lei 0001 |
Inf. Sci. | 3 |
| 2021 | CA-PMG: Channel attention and progressive multi-granularity training network for fine-grained visual classificationabstractAbstract Fine‐grained visual classification is challenging due to the inherently subtle intra‐class object variations. To solve this issue, a novel framework named channel attention and progressive multi‐granularity training network, is proposed. It first exploits meaningful feature maps through the channel attention module and captures multi‐granularity features by the progressive multi‐granularity training module. For each feature map, the channel attention module is proposed to explore channel‐wise correlation. This allows the model to re‐weight the channels of the feature map according to the impact of their semantic information on performance. Furthermore, the progressive multi‐granularity training module is introduced to fuse features cross multi‐granularity. And the fused features pay more attention to the subtle differences between images. The model can be trained efficiently in an end‐to‐end manner without bounding box or part annotations. Finally, comprehensive experiments are conducted to show that the method achieves state‐of‐the‐art performances on the CUB‐200‐2011, Stanford Cars, and FGVC‐Aircraft datasets. Ablation studies demonstrate the effectiveness of each part in our module. Qiguang Miao, Hang Yao 0001, Xiangzeng Liu, Ruyi Liu 0001, Maoguo Gong |
IET Image Process. | 5 |
| 2021 | Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition
Qiguang Miao, Ruyi Liu 0001, Wentian Xin, Sheng Zhong 0006, Xuesong Gao |
Neurocomputing | 3 |
| 2020 | A Semi-Supervised High-Level Feature Selection Framework for Road Centerline ExtractionabstractAccurate road centerline extraction is very important for many vital applications. In the road extraction, the acquisition of labeled data is time-consuming; thus, there is only a small amount of labeled samples in reality. To solve the problem of limited labeled samples, a semi-supervised road centerline extraction is proposed, which incorporates high-level feature selection, Markov random field (MRF), and ridge transversal method. The proposed road extraction approach consists of three steps: multiple features extraction, semi-supervised road area extraction, and road centerlines extraction. To get more abstract and discriminative high-level features, we apply multiple-feature adaptive sparse representation in mid-level features in different views generated by different prototype sets. To obtain an accurate road area result, we combine the feature learning framework with MRF. Then, we integrate Gabor filters and nonmaxima suppression with the ridge transversal method to extract centerlines. It is verified the proposed method achieves comparable performance with the state-of-the-art methods in terms of visual and quantitative aspects. Ruyi Liu 0001, Qiguang Miao, Yi Zhang 0033, Maoguo Gong, Pengfei Xu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Architectural Style Classification Based on DNN Model
Qiguang Miao, Ruyi Liu 0001, Jianfeng Song |
PRCV (1) | 3 |
| 2019 | Multiscale road centerlines extraction from high-resolution aerial imagery
Ruyi Liu 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Yunan Li 0001, Pengfei Xu 0003 |
Neurocomputing | 1 |
| 2018 | A multi-scale fusion scheme based on haze-relevant features for single image dehazing
Yunan Li 0001, Qiguang Miao, Ruyi Liu 0001, Jianfeng Song, Yi-Ning Quan, Yuhui Huang |
Neurocomputing | 3 |
| 2016 | Road centerlines extraction from high resolution images based on an improved directional segmentation and road probability
Ruyi Liu 0001, Jianfeng Song, Qiguang Miao, Pengfei Xu 0003 |
Neurocomputing | 1 |
| 2016 | Dynamic character grouping based on four consistency constraints in topographic maps
Pengfei Xu 0003, Qiguang Miao, Ruyi Liu 0001, Xiaojiang Chen, Xunli Fan |
Neurocomputing | 3 |
| 2016 | Improved road centerlines extraction in high-resolution remote sensing images using shear transform, directional morphological filtering and enhanced broken lines connection
Ruyi Liu 0001, Qiguang Miao, Bormin Huang, Jianfeng Song, Johan Debayle |
J. Vis. Commun. Image Represent. | 1 |