Lei Lyu 0001

dblp:202/6940-1 · DBLP profile ↗
← Back
56ranked-venue papers
4as first author
55since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 4 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DHPL: Dynamic hyperedge refinement with spatio-temporal prototype learning for skeleton-based action recognition
Chen Pang 0001, Lei Lyu 0001
Expert Syst. Appl.4
2026 Multi-view transformer with hierarchical attention for action recognition
Yiliang Liu, Guangkuo Gao, Chen Pang 0001, Lei Lyu 0001
Neurocomputing4
2026 PHANet: Contrastive hypergraph structures and prototype memory for discriminative skeleton-based action recognition
Chen Pang 0001, Guangqi Wen, Chunmeng Kang, Xingyu Gao 0001, Lei Lyu 0001
Knowl. Based Syst.6
2026 Road Extraction via the Complementary Relationship Between Architectural and Nonarchitectural Areas
abstract
Road extraction from remote sensing images is crucial for applications such as autonomous driving and urban planning. However, existing methods often neglect the spatial correlation between roads and adjacent buildings, limiting extraction accuracy. This study proposes SAM BE Road, a road-building co-extraction framework based on the Segment Anything Model (SAM), which optimizes road extraction performance by integrating building spatial information. Extensive experiments are carried out on the Massachusetts dataset and the New York dataset, and the results demonstrate that our SAM BE Road outperforms other state-of-the-art methods in extraction accuracy and topological connectivity. The road labels extracted by our method exhibit preferable connectivity, especially in complex urban environments.
Mingyi Yu, Cun Ji, Lei Lyu 0001
IEEE Geosci. Remote. Sens. Lett.5
2026 KG-PTP: A Knowledge Graph-Driven Approach for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction in high-density and dynamic environments remains a significant challenge due to the limitations in modeling complex group interactions and evolving pedestrian-environment dependencies. This article presents KG-PTP, a knowledge graph-driven trajectory prediction framework that incorporates a dynamic trajectory knowledge graph (DTKG), a direction-aware clustering algorithm (D-DBSCAN), and an internal–external behavior modeling component (IN-EX). The DTKG enables semantic integration of heterogeneous spatial–temporal information and supports real-time updates. D-DBSCAN dynamically groups pedestrians based on motion direction and density, while IN-EX captures both individual motivations and external environmental influences. Experimental evaluations on ETH pedestrian dataset (ETH) and UCY crowd dataset (UCY) datasets demonstrate that KG-PTP achieves 12.3% and 15.7% reductions in average displacement error (ADE) and final displacement error (FDE), respectively, compared with state-of-the-art baselines. Furthermore, the proposed framework is validated through crowd evacuation simulations, confirming its applicability to public safety scenarios.
Xiling Cao, Chen Pang 0001, Hong Liu 0013, Lei Lyu 0001, Wenhao Li 0006, Jihao Duan
IEEE Trans. Comput. Soc. Syst.4
2025 Learning Adaptive Node Selection with External Attention for Human Interaction Recognition
abstract
Most GCN-based methods model interacting individuals as independent graphs, neglecting their inherent inter-dependencies. Although recent approaches utilize predefined interaction adjacency matrices to integrate participants, these matrices fail to adaptively capture the dynamic and context-specific joint interactions across different actions. In this paper, we propose the Active Node Selection with External Attention Network (ASEA), an innovative approach that dynamically captures interaction relationships without predefined assumptions. Our method models each participant individually using a GCN to capture intra-personal relationships, facilitating a detailed representation of their actions. To identify the most relevant nodes for interaction modeling, we introduce the Adaptive Temporal Node Amplitude Calculation (AT-NAC) module, which estimates global node activity by combining spatial motion magnitude with adaptive temporal weighting, thereby highlighting salient motion patterns while reducing irrelevant or redundant information. A learnable threshold, regularized to prevent extreme variations, is defined to selectively identify the most informative nodes for interaction modeling. To capture interactions, we design the External Attention (EA) module to operate on active nodes, effectively modeling the interaction dynamics and semantic relationships between individuals. Extensive evaluations show that our method captures interaction relationships more effectively and flexibly, achieving state-of-the-art performance.
Chen Pang 0001, Xuequan Lu, Qianyu Zhou 0001, Lei Lyu 0001
ACM Multimedia4
2025 DS-MAE: Dual-Siamese Masked Autoencoders for Point Cloud Analysis
abstract
Masked autoencoders (MAEs) have emerged as a powerful self-supervised approach for point cloud analysis. Nevertheless, existing methods often separately focus on global structures or multi-scale features, ignoring their complementary potential. In this paper, we propose a novel dual-Siamese masked autoencoder (DS-MAE) framework that explores integrating global and hierarchical feature learning in a unified architecture for point cloud analysis. In particular, we introduce a consistent dual-branch patch embedding strategy to partition the point cloud into patches using shared group centers, ensuring both global and hierarchical branches process point patches centered at the same spatial locations. Each branch employs dual-branch Siamese encoders to process original and augmented point patches, learning representations that capture both local details and global context. In addition, we have designed cross-attention Siamese decoders to reconstruct masked point patches and align features both within and between branches with cross-attention mechanisms. Comprehensive experiments demonstrate our method consistently achieves superior results to prior methods. Code is available at https://github.com/shaoandy1211/DS-MAE.git.
Di Shao, Yaping Jing, Xinkui Zhao, Shasha Mao, Lei Lyu 0001, Xiao Liu 0004, Xuequan Lu
Comput. Vis. Media5
2025 Correction of medical image segmentation errors through contrast learning with multi-branch
Tianlei Gao, Lei Lyu 0001, Nuo Wei, Tongze Liu, Minglei Shu
Eng. Appl. Artif. Intell.2
2025 Dual-branch feature Reinforcement Transformer for preoperative parathyroid gland segmentation
Lei Lyu 0001, Chen Pang 0001, Qinghan Yang, Kailin Liu, Chong Geng
Eng. Appl. Artif. Intell.1
2025 Adaptive Koopman contrastive learning for skeleton-based action recognition
Xiaohang Yu, Chen Pang 0001, Lei Lyu 0001
Neurocomputing4
2025 Sparse and Dense: Learning Confusion Representation Network for 3-D Action Recognition
abstract
In the field of the Internet of Medical Things (IoMT), the demand for Human action recognition (HAR) is growing. Due to the limitations of portability and privacy of traditional sensors, many endeavors have made significant progress in 3-D skeleton-based action recognition. However, existing methods ignore potential higher-order semantic information between joints and fail to perceive rapidly changing dynamic details, resulting in frequent confusion of actions with similar motion trajectories. To alleviate this issue, we propose a novel learning confusion representation network (LCR-Net). Specifically, a progressive feature enhancement module is first designed to utilize self-attention to gradually aggregate lower-order features to higher-order features, emphasizing the relative movement between body parts. Second, we design the enhanced spatio-temporal convolution to explore the potential spatio-temporal dependencies between joints by adding a mask matrix and an attention fusion mechanism. To further perceive the spatio-temporal relationships in subtle changes, we divide the interactive sparse-dense pathways at different spatio-temporal resolutions and enhance the complementary information between the two pathways through feature interaction. Finally, the frequency excitation learning module is proposed to efficiently learn the importance of different frequencies by cross-channel modeling, promoting the compactness of actions within classes and the separability of confusion actions. In addition, the lightweight LCR-Net${}^{\textbf {+}}$is achieved through model compression optimization to meet the deployment requirements of IoT systems. Comprehensive experiments conducted on three public datasets (NTU-RGB+D60&120, NW-UCLA) demonstrate the superior performance of our model.
Xinran Hou, Pei Geng, Tianchen Li, Yan Li 0046, Lei Lyu 0001
IEEE Internet Things J.6
2025 Multi-Scale Adaptive Large Kernel Graph Convolutional Network for Skeleton-Based Action Recognition
Yu-Qing Zhang, Chen Pang 0001, Pei Geng, Xuequan Lu, Lei Lyu 0001
J. Comput. Sci. Technol.5
2025 Skeleton-based action recognition through attention guided heterogeneous graph neural network
Tianchen Li, Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001
Knowl. Based Syst.5
2025 Efficiency-Driven Adaptive Task Planning for Household Robot Based on Hierarchical Item-Environment Cognition
abstract
Task planning focused on household robots represents a conventional yet complex research domain, necessitating the development of task plans that enable robots to execute unfamiliar household services. This area has garnered significant research interest due to its extensive applications in robotics, particularly concerning household robots. Nevertheless, the majority of task planning methodologies exhibit suboptimal performance regarding the success and efficiency of completing household tasks, primarily due to a lack of cognitive capacity of household items and home environments. To address these challenges, we propose an efficiency-driven adaptive task planning approach based on hierarchical item-environment cognition. Initially, we establish a multiple semantic attribute-based priori knowledge (MSAPK) framework to facilitate the attributive representation of household items. Utilizing MSAPK, we develop a long short-term memory (LSTM) based item cognition model that assigns relevant attributes and substitutes to specified household items, thereby enhancing the cognitive capabilities of household robots at the attribute level. Subsequently, we construct an environment cognition model that delineates the relationships between household items and room types, enabling household robots to locate target items more efficiently. Through hierarchical item-environment cognition, we introduce a strategy for adaptive task planning, empowering household robots to execute household tasks with both flexibility and efficiency. The generated plans are evaluated in both virtual and real-world experiments, with promising results affirming the effectiveness of our proposed methodology.
Mengyang Zhang, Guohui Tian, Yongcheng Cui, Hong Liu 0013, Lei Lyu 0001
IEEE Trans. Cybern.5
2024 EGLA-Net: Edge Guided with Lesion Aware Network for Medical image segmentation
abstract
Medical image segmentation plays a crucial role in diagnosis analysis and disease treatment. However, the boundaries of most lesion areas are blurred, and there are significant differences between the shape, size, and appearance of different lesion areas, which pose a challenge to many methods. To solve these problems, we propose an edge guided with lesion aware network (EGLA-Net). Specifically, we connect an edge attention (EA) module at each stage of the encoder to preserve more local edge features. And we design multiple global pyramidal guidance (GPG) modules to provide different levels of global information for the decoder. Further, we introduce a dynamic kernel generation (KG) and kernel update (KU) mechanism that utilizes continuously updated kernel parameters to learn and mine distinguishable regional features, resolving differences in shape, size, and appearance of different diseased regions. Extensive experiments show that the EGLA-Net can achieve superior segmentation performance.
Ruixue Qi, Chen Pang 0001, Mengyang Zhang, Lei Lyu 0001
ICME4
2024 HSNet: Crowd counting via hierarchical scale calibration and spatial attention
Ran Qi, Chunmeng Kang, Hong Liu 0013, Lei Lyu 0001
Eng. Appl. Artif. Intell.4
2024 Counting in congested crowd scenes with hierarchical scale-aware encoder-decoder network
Run Han, Ran Qi, Xuequan Lu, Lei Huang 0010, Lei Lyu 0001
Expert Syst. Appl.5
2024 CrowdUNet: Segmentation assisted U-shaped crowd counting network
Zhou Cao, Lei Lyu 0001, Ran Qi, Jihua Wang
Neurocomputing2
2024 Variation-aware directed graph convolutional networks for skeleton-based action recognition
Tianchen Li, Pei Geng, Guohui Cai, Xinran Hou, Xuequan Lu, Lei Lyu 0001
Knowl. Based Syst.6
2024 Understanding the role of pathways in a deep neural network
Lei Lyu 0001, Chen Pang 0001, Jihua Wang
Neural Networks1
2024 Dimensional Transformation Mixer for Ultra-High-Definition Industrial Camera Dehazing
abstract
Haze severely affects the reliability of vision-based industrial systems, and most of the current dehazing methods are not applicable to images captured by ultra-high-definition (UHD) industrial cameras. In this article, we propose a novel dimensional transformation mixer (DMixer) model for recovering haze-free images from UHD haze images. In DMixer, the dimensional transformation module encodes the complete image in multiple stages from different perspectives and associates features from different views by permuting the tensor for efficient long-range dependency modeling. In this way, the global perception capabilities of DMixer are complemented, allowing the quality of the reconstructed UHD images to be improved. Furthermore, DMixer employs a dual-stream network framework that combines local and multiscale features, allowing DMixer to better trade off the performance and efficiency for real industrial systems. Our proposed method enables the real-time processing of UHD images ($\sim$59 fps). Extensive results show that the proposed model outperforms current state-of-the-art methods, with PSNR improvements of 1.94 and 0.95 dB on the 4KID and O-HAZE datasets, respectively.
Yunliang Zhuang, Zhuoran Zheng, Lei Lyu 0001, Xiuyi Jia, Chen Lyu 0001
IEEE Trans. Ind. Informatics4
2024 IIAM: Intra and Inter Attention With Mutual Consistency Learning Network for Medical Image Segmentation
abstract
Medical image segmentation provides a reliable basis for diagnosis analysis and disease treatment by capturing the global and local features of the target region. To learn global features, convolutional neural networks are replaced with pure transformers, or transformer layers are stacked at the deepest layers of convolutional neural networks. Nevertheless, they are deficient in exploring local-global cues at each scale and the interaction among consensual regions in multiple scales, hindering the learning about the changes in size, shape, and position of target objects. To cope with these defects, we propose a novel Intra and Inter Attention with Mutual Consistency Learning Network (IIAM). Concretely, we design an intra attention module to aggregate the CNN-based local features and transformer-based global information on each scale. In addition, to capture the interaction among consensual regions in multiple scales, we devise an inter attention module to explore the cross-scale dependency of the object and its surroundings. Moreover, to reduce the impact of blurred regions in medical images on the final segmentation results, we introduce multiple decoders to estimate the model uncertainty, where we adopt a mutual consistency learning strategy to minimize the output discrepancy during the end-to-end training and weight the outputs of the three decoders as the final segmentation result. Extensive experiments on three benchmark datasets verify the efficacy of our method and demonstrate superior performance of our model to state-of-the-art techniques.
Chen Pang 0001, Xuequan Lu, Renfeng Zhang, Lei Lyu 0001
IEEE J. Biomed. Health Informatics5
2024 Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action Recognition
abstract
Supervised human action recognition methods based on skeleton data have achieved impressive performance recently. However, many current works emphasize the design of different contrastive strategies to gain stronger supervised signals, ignoring the crucial role of the model's encoder in encoding fine-grained action representations. Our key insight is that a superior skeleton encoder can effectively exploit the fine-grained dependencies between different skeleton information (e.g., joint, bone, angle) in mining more discriminative fine-grained features. In this paper, we devise an innovative hierarchical aggregated graph neural network (HA-GNN) that involves several core components. In particular, the proposed hierarchical graph convolution (HGC) module learns the complementary semantic information among joint, bone, and angle in a hierarchical manner. The designed pyramid attention fusion mechanism (PAFM) fuses the skeleton features successively to compensate for the action representations obtained by the HGC. We use the multi-scale temporal convolution (MSTC) module to enrich the expression capability of temporal features. In addition, to learn more comprehensive semantic representations of the skeleton, we construct a multi-task learning framework with simple contrastive learning and design the learnable data-enhanced strategy to acquire different data representations. Extensive experiments on NTU RGB+D 60/120, NW-UCLA, Kinetics-400, UAV-Human, and PKUMMD datasets prove that the proposed HA-GNN without contrastive learning achieves state-of-the-art performance in skeleton-based action recognition, and it achieves even better results with contrastive learning.
Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001
IEEE Trans. Multim.4
2024 Self-Adaptive Graph With Nonlocal Attention Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have achieved encouraging progress in modeling human body skeletons as spatial-temporal graphs. However, existing methods still suffer from two inherent drawbacks. Firstly, these models process the input data based on the physical structure of the human body, which leads to some latent correlations among joints being ignored. Furthermore, the key temporal relationships between nonadjacent frames are overlooked, preventing to fully learn the changes of the body joints along the temporal dimension. To address these issues, we propose an innovative spatial-temporal model by introducing a self-adaptive GCN (SAGCN) with global attention network, collectively termed SAGGAN. Specifically, the SAGCN module is proposed to construct two additional dynamic topological graphs to learn the common characteristics of all data and represent a unique pattern for each sample, respectively. Meanwhile, the global attention module (spatial attention (SA) and temporal attention (TA) modules) is designed to extract the global connections between different joints in a single frame and model temporal relationships between adjacent and nonadjacent frames in temporal sequences. In this manner, our network can capture richer features of actions for accurate action recognition and overcome the defect of the standard graph convolution. Extensive experiments on three benchmark datasets (NTU-60, NTU-120, and Kinetics) have demonstrated the superiority of our proposed method.
Chen Pang 0001, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Graph Convolutional Network with Long Time Memory for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition task has been widely studied in recent years. Currently, the most popular researches use graph convolutional network (GCN) to solve this task by modeling human joints data as spatio-temporal graph. However, a large number of long-term temporal motion relationships cannot be effectively captured by GCN. Thus, recurrent neural network (RNN) is introduced to solve this defect. In this work, we propose a model namely graph convolutional network with long time memory (GCN-LTM). Specifically, there are two task streams in our proposed model: GCN stream and RNN stream, respectively. The GCN stream aims to capture the spatial motion relationships as well as the RNN stream focuses on extracting the long-term temporal patterns. In addition, we introduce the contrastive learning strategy to better facilitate feature learning between these two streams. The multiple ablation experiments have verified the feasibility of our proposed model. Numerous experiments show that the proposed model is superior to the current state-of-the-art method under two large-scale datasets including NTU-RGBD and NTU-RGBD-120.
Yanpeng Qi, Chen Pang 0001, Yiliang Liu, Hong Liu 0013, Lei Lyu 0001
CSCWD5
2023 A Novel Method for Wearable Activity Recognition with Feature Evolvable Streams
Chunyu Hu 0001, Hong Liu 0013, Lei Lyu 0001, Lin Yuan 0001
MobiQuitous (1)4
2023 Semi-supervised Network for Thyroid Nodule Segmentation via Joint Consistency Learning and Co-Training
abstract
Thyroid nodule is a common clinical disease, although most nodules are benign, the incidence rate of thyroid cancer has risen rapidly in recent years. Even though many methods have achieved automated thyroid nodule segmentation based on deep learning, these methods are based on supervised learning and require a large amount of labeled data for training. However, the labeling work must be carried out by professional doctors, which results in a small number of datasets and difficulty in labeling. To address this problem, this paper proposes a semi-supervised thyroid nodule segmentation model via joint consistency learning and co-training. This model includes two branches: consistency learning and co-training. In the consistency learning branch, based on consistent regularization, the teacher model guides the student model to optimize. In order to make the teacher model more stable, we design a co-training framework to further optimize the teacher model. In co-training branch, the teacher model and TransUNet extract different representations of the same sample and teach each other to prevent consistent but incorrect predictions between the teacher model and the student model. This semi-supervised model can learns useful feature representations from unlabeled data, and effectively trains the model with a small amount of labeled data, reducing the dependence on labeled data during the model training process.
Guijuan Zhang, Lei Lyu 0001
SMC3
2023 Cascaded parallel crowd counting network with multi-resolution collaborative representation
Lei Lyu 0001, Run Han
Appl. Intell.1
2023 Environment-sensitive crowd behavior modeling method based on reinforcement learning
Chen Pang 0001, Lei Lyu 0001, Qinglin Zhou, Limei Zhou
Appl. Intell.2
2023 FedIERF: Federated Incremental Extremely Random Forest for Wearable Health Monitoring
Chunyu Hu 0001, Lisha Hu, Lin Yuan 0001, Dianjie Lu, Lei Lyu 0001, Yiqiang Chen 0001
J. Comput. Sci. Technol.5
2023 Focusing Fine-Grained Action by Self-Attention-Enhanced Graph Neural Networks With Contrastive Learning
abstract
With the aid of graph convolution neural network and transformer model, human action recognition has achieved significant performance based on skeleton data. However, the majority of existing works rarely focus on identifying fine-grained motion information (i.e., “read”, “write”, etc.). Furthermore, they tend to explore correlations between joints and bones ignoring the angular information. Consequently, the recognition accuracy for fine-grained actions with most models is still less desired. To address this issue, we first attempt to bring angular information as a complement to familiar joint and bone information, while learning the potential dependencies of the three kinds of information using graph neural networks. Based on this, we propose a self-attention-enhanced graph neural network (SAE-GNN), which consists of a kernel-unified graph convolution (KUGC) module and an enhanced attention graph convolution (EAGC) module. The KUGC module is devised to effectively extract rich features in the skeleton information. The EAGC consisting of a multi-scale enhanced graph convolution block and a multi-headed self-attention block is designed to learn the potential high-level semantic information in the features. Besides, we introduce contrastive learning in the two blocks to enhance feature representation by maximizing their mutual information. We conduct extensive experiments on four publicly available datasets, and results show that our model outperforms state-of-the-art methods in recognizing fine-grained actions.
Pei Geng, Xuequan Lu, Chunyu Hu 0001, Hong Liu 0013, Lei Lyu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Skeleton-Based Action Recognition Through Contrasting Two-Stream Spatial-Temporal Networks
abstract
For pursuing accurate skeleton-based action recognition, most prior methods use the strategy of combining Graph Convolution Networks (GCNs) with attention-based methods in a serial way. However, they regard the human skeleton as a complete graph, resulting in less variations between different actions (e.g., the connection between the elbow and head in action “clapping hands”). For this, we propose a novel Contrastive GCN-Transformer Network (ConGT) which fuses the spatial and temporal modules in a parallel way. The ConGT involves two parallel streams: Spatial-Temporal Graph Convolution stream (STG) and Spatial-Temporal Transformer stream (STT). The STG is designed to obtain action representations maintaining the natural topology structure of the human skeleton. The STT is devised to acquire action representations containing the global relationships among joints. Since the action representations produced from these two streams contain different characteristics, and each of them knows little information of the other, we introduce the contrastive learning paradigm to guide their output representations of the same sample to be as close as possible in a self-supervised manner. Through the contrastive learning, they can learn information from each other to enrich the action features by maximizing the mutual information between the two types of action representations. To further improve action recognition accuracy, we introduce the Cyclical Focal Loss (CFL) which can focus on confident training samples in early training epochs, with an increasing focus on hard samples during the middle epochs. We conduct experiments on three benchmark datasets, which demonstrate that our model achieves state-of-the-art performance in action recognition.
Chen Pang 0001, Xuequan Lu, Lei Lyu 0001
IEEE Trans. Multim.3
2023 Contrastive Multi-Level Graph Neural Networks for Session-Based Recommendation
abstract
Session-based recommendation (SBR) aims to predict the next item at a certain time point based on anonymous user behavior sequences. Existing methods typically model session representation based on simple item transition information. However, since session-based data consists of limited users' short-term interactions, modeling session representation by capturing fixed item transition information from a single dimension suffers from data sparsity. In this paper, we propose a novel contrastive multi-level graph neural networks (CM-GNN) to better exploit complex and high-order item transition information. Specifically, CM-GNN applies local-level graph convolutional network (L-GCN) and global-level graph convolutional network (G-GCN) on the current session and all the sessions respectively, to effectively capture pairwise relations over all the sessions by aggregation strategy. Meanwhile, CM-GNN applies hyper-level graph convolutional network (H-GCN) to capture high-order information among all the item transitions. CM-GNN further introduces an attention-based fusion module to learn pairwise relation-based session representation by fusing the item representations generated by L-GCN and G-GCN. CM-GNN averages the item representations obtained by H-GCN to obtain high-order relation-based session representation. Moreover, to convert the high-order item transition information into the pairwise relation-based session representation, CM-GNN maximizes the mutual information between the representations derived from the fusion module and the average pool layer by contrastive learning paradigm. We conduct extensive experiments on several widely used benchmark datasets to validate the efficacy of the proposed method. The encouraging results demonstrate that our proposed method outperforms the state-of-the-art SBR techniques.
Fuyun Wang, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Multim.4
2023 Dilated Convolution-based Feature Refinement Network for Crowd Localization
abstract
As an emerging computer vision task, crowd localization has received increasing attention due to its ability to produce more accurate spatially predictions. However, continuous scale variations in complex crowd scenes lead to tiny individuals at the edges, so that existing methods cannot achieve precise crowd localization. Aiming at alleviating the above problems, we propose a novel Dilated Convolution-based Feature Refinement Network (DFRNet) to enhance the representation learning capability. Specifically, the DFRNet is built with three branches that can capture the information of each individual in crowd scenes more precisely. More specifically, we introduce a Feature Perception Module to model long-range contextual information at different scales by adopting multiple dilated convolutions, thus providing sufficient feature information to perceive tiny individuals at the edge of images. Afterwards, a Feature Refinement Module is deployed at multiple stages of the three branches to facilitate the mutual refinement of feature information at different scales, thus further improving the expression capability of multi-scale contextual information. By incorporating the above modules, DFRNet can locate individuals in complex scenes more precisely. Extensive experiments on multiple datasets demonstrate that the proposed method has more advanced performance compared to existing methods and can be more accurately adapted to complex crowd scenes.
Xingyu Gao 0001, Jinyang Xie, Zhenyu Chen 0003, Anan Liu, Zhenan Sun, Lei Lyu 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2022 Deep Reinforcement Learning with Long-Time Memory Capability for Robot Mapless Navigation
abstract
Achieving autonomous navigation of indoor robots in a mapless environment is a long-standing research problem. Deep Reinforcement Learning (DRL) is widely used for robot navigation by virtue of learning through interaction with the environment. However, the large number of trials for training need lengthy computation times. To address this issue, we propose an innovative DRL model with long-time memory capability for mobile robot’s mapless navigation, which can achieve end-to-end navigation based only on laser-ranging data and target location. The long-time memory capability is realized by introducing a memory module based on the special structure of the Long Short-Term Memory (LSTM). The memory module ensures that the model derives information from previous navigation experiences and thus optimizes the model’s decision-making. In addition, to enable the model to explore the environment more effectively, we design a novel dual-noise mechanism consisting of Gaussian noise and Ornstein-Uhlenbeck noise. Extensive experiments are conducted on the Gazebo simulation platform and validate that the proposed approach can generate a smoother navigation path and exceed the state-of-the-art performance with less computation cost.
Qinglin Zhou, Lei Lyu 0001, Hong Liu 0013
CSCWD2
2022 MMF3: Neural Code Summarization Based on Multi-Modal Fine-Grained Feature Fusion
abstract
Background: Code summarization automatically generates the corresponding natural language descriptions according to the input code to characterize the function implemented by source code. Comprehensiveness of code representation is critical to code summarization task. However, most existing approaches typically use coarse-grained fusion methods to integrate multi-modal features. They generally represent different modalities of a piece of code, such as an Abstract Syntax Tree (AST) and a token sequence, as two embeddings and then fuse the two ones at the AST/code levels. Such a coarse integration makes it difficult to learn the correlations between fine-grained code elements across modalities effectively. Aims: This study intends to improve the model’s prediction performance for high-quality code summarization by accurately aligning and fully fusing semantic and syntactic structure information of source code at node/token levels. Method: This paper proposes a Multi-Modal Fine-grained Feature Fusion approach (MMF3) for neural code summarization. The method uses the Transformer architecture. In particular, we introduce a novel fine-grained fusion method, which allows fine-grained fusion of multiple code modalities at the token and node levels. Specifically, we use this method to fuse information from both token and AST modalities and apply the fused features to code summarization. Results: We conduct experiments on one Java and one Python datasets, and evaluate generated summaries using four metrics. The results show that: 1) the performance of our model outperforms the current state-of-the-art models, and 2) the ablation experiments show that our proposed fine-grained fusion method can effectively improve the accuracy of generated summaries. Conclusion: MMF3 can mine the relationships between cross-modal elements and perform accurate fine-grained element-level alignment fusion accordingly. As a result, more clues can be provided to improve the accuracy of the generated code summaries.
Yuexiu Gao, Lei Lyu 0001, Chen Lyu 0001
ESEM3
2022 HELoC: hierarchical contrastive learning of source code representation
abstract
Abstract syntax trees (ASTs) play a crucial role in source code representation. However, due to the large number of nodes in an AST and the typically deep AST hierarchy, it is challenging to learn the hierarchical structure of an AST effectively. In this paper, we propose HELoC, a hierarchical contrastive learning model for source code representation. To effectively learn the AST hierarchy, we use contrastive learning to allow the network to predict the AST node level and learn the hierarchical relationships between nodes in a self-supervised manner, which makes the representation vectors of nodes with greater differences in AST levels farther apart in the embedding space. By using such vectors, the structural similarities between code snippets can be measured more precisely. In the learning process, a novel GNN (called Residual Self-attention Graph Neural Network, RSGNN) is designed, which enables HELoC to focus on embedding the local structure of an AST while capturing its overall structure. HELoC is self-supervised and can be applied to many source code related downstream tasks such as code classification, code clone detection, and code clustering after pre-training. Our extensive experiments demonstrate that HELoC outperforms the state-of-the-art source code representation models.
Hongyu Zhang 0002, Chen Lyu 0001, Zhuoran Zheng, Lei Lyu 0001, Songlin Hu 0001
ICPC7
2022 Context-aware pyramid attention network for crowd counting
Lingyu Gu, Chen Pang 0001, Yanjun Zheng, Chen Lyu 0001, Lei Lyu 0001
Appl. Intell.5
2022 HRANet: Hierarchical region-aware network for crowd counting
Jinyang Xie, Lingyu Gu, Zhonghui Li, Lei Lyu 0001
Appl. Intell.4
2022 A CNN-based multi-task framework for weather recognition with multi-scale weather cues
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Lei Lyu 0001
Expert Syst. Appl.5
2022 Hierarchical feature aggregation network with semantic attention for counting large-scale crowd
abstract
The purpose of crowd counting is to estimate the number of people in an image. Due to the unconstrained imaging conditions, the scale variation and background occlusion in the images make it still challenging to achieve counting accurately. To tackle the two key challenges, we design a Hierarchical Feature Aggregation Network (HFANet) for accurate crowd counting in complex scenarios. The proposed method can extract multiple features at different levels and then aggregate them hierarchically to generate a high-quality density map. To highlight the crowd regions effectively from the cluttered background, we propose the Semantic Attention module to preserve the useful feature information through the attention mechanism for the low-level features extracted by VGG-16. Meanwhile, we employ convolutional kernels of different sizes to extract multi-scale features and global average pooling operations to preserve contextual information. Furthermore, we design the Feature Aggregation module to integrate the extracted multiple features through a progressive approach, which aims to take full advantage of the complementary properties between low-level and high-level features for efficient features aggregation. Finally, we evaluate the performance of the proposed HFANet on four challenging datasets. Extensive experimental results demonstrate that our proposed approach has better performance compared with most state-of-the-art approaches.
Chunmeng Kang, Lei Lyu 0001
Int. J. Intell. Syst.3
2022 GT-SimNet: Improving code automatic summarization via multi-modal similarity networks
Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001
J. Syst. Softw.6
2022 CGSNet: Contrastive Graph Self-Attention Network for Session-based Recommendation
Fuyun Wang, Xuequan Lu, Lei Lyu 0001
Knowl. Based Syst.3
2022 Adaptive multi-level graph convolution with contrastive learning for skeleton-based action recognition
Pei Geng, Fuyun Wang, Lei Lyu 0001
Signal Process.4
2022 Adaptive Multiview Graph Difference Analysis for Video Summarization
abstract
Adapting detection to different shot types is a significant challenge for video summarization methods based on shot boundary detection. In our recent work, a new graph model was introduced in the feature modelling of frames and analysed for changes in graph structure to improve the detection of shot boundaries. In this paper, we further explore the potential of graph models and propose a more general framework for online, real-time automatic video summarization. The framework develops a novel adaptive multiview graph difference analysis method to improve the algorithm’s robustness in detecting different shot transitions. Previous fusion methods typically used a priori knowledge to assign weights to the various feature differences from videos. In contrast, our framework can weigh and fuse the resulting differences by learning the importance of various video features from the structural changes of the corresponding multiview graphs. Additionally, we propose a new threshold-based adaptive decision method which can dynamically select the most accurate shot boundary decision threshold by analysing a small number of historical frames and learning the tolerance factor in the current shot. The experimental results show that the proposed method outperforms state-of-the-art methods in terms of precision and F-score on the VSUMM and YouTube datasets.
Caixia Ma, Lei Lyu 0001, Guoliang Lu, Chen Lyu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Multi-modal Code Summarization Fusing Local API Dependency Graph and AST
Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001
ICONIP (5)6
2021 MPANet: Multi-level Progressive Aggregation Network for Crowd Counting
Run Han, Chen Pang 0001, Chunmeng Kang, Chen Lyu 0001, Lei Lyu 0001
ICONIP (3)6
2021 Saliency Detection Framework Based on Deep Enhanced Attention Network
Xing Sheng 0001, Zhuoran Zheng, Chunmeng Kang, Yunliang Zhuang, Lei Lyu 0001, Chen Lyu 0001
ICONIP (4)6
2021 Code Representation Based on Hybrid Graph Modelling
Zhuoran Zheng, Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001
ICONIP (5)6
2021 HIANet: Hierarchical Interweaved Aggregation Network for Crowd Counting
Jinyang Xie, Jinfang Zheng, Lingyu Gu, Chen Lyu 0001, Lei Lyu 0001
ICONIP (6)5
2021 SS-CCN: Scale Self-guided Crowd Counting Network
Jinfang Zheng, Jinyang Xie, Chen Lyu 0001, Lei Lyu 0001
ICONIP (4)4
2021 A Dueling-DDPG Architecture for Mobile Robots Path Planning Based on Laser Range Findings
Panpan Zhao, Jinfang Zheng, Qinglin Zhou, Chen Lyu 0001, Lei Lyu 0001
PRICAI (1)5
2021 GCMNet: Gated Cascade Multi-scale Network for Crowd Counting
Jinfang Zheng, Panpan Zhao, Jinyang Xie, Chen Lyu 0001, Lei Lyu 0001
PRICAI (2)5
2021 TreeBERT: A tree-based pre-trained model for programming language
abstract
Source code can be parsed into the abstract syntax tree (AST) based on defined syntax rules. However, in pre-training, little work has considered the incorporation of tree structure into the learning process. In this paper, we present TreeBERT, a tree-based pre-trained model for improving programming language-oriented generation tasks. To utilize tree structure, TreeBERT represents the AST corresponding to the code as a set of composition paths and introduces node position embedding. The model is trained by tree masked language modeling (TMLM) and node order prediction (NOP) with a hybrid objective. TMLM uses a novel masking strategy designed according to the tree’s characteristics to help the model understand the AST and infer the missing semantics of the AST. With NOP, TreeBERT extracts the syntactical structure by learning the order constraints of nodes in AST. We pre-trained TreeBERT on datasets covering multiple programming languages. On code summarization and code documentation tasks, TreeBERT outperforms other pre-trained models and state-of-the-art models designed for these tasks. Furthermore, TreeBERT performs well when transferred to the pre-trained unseen programming language.
Zhuoran Zheng, Chen Lyu 0001, Lei Lyu 0001
UAI5
2021 Graph-based structural difference analysis for video summarization
Chunlei Chai, Guoliang Lu, Ruyun Wang, Chen Lyu 0001, Lei Lyu 0001, Peng Zhang 0009, Hong Liu 0013
Inf. Sci.5
2020 Stylized human motion warping method based on identity-independent coordinates
Lei Lyu 0001
Soft Comput.1