VLDB 2026 Research / reviewers in the wild / expert
Keyan Ren
dblp:85/7723
· DBLP profile ↗
15ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-7332-2882ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RCACF: 4D Radar-Camera Adaptive and Cooperative Fusion for 3D detection
Keyan Ren, Wenzhuo Zhu, Jianyu Gu, Qing Niu |
Pattern Recognit. | 2 |
| 2025 | DSPA: Dual-Spiral Pyramid Network with Multi-scale Attention for Temporal Action LocalizationabstractMost popular Feature Pyramid Networks (FPN) for temporal action localization (TAL) in videos encode multi-scale features during downsampling, which inevitably bring fine-grained feature loss. In addition, most popular TAL models that directly apply self-attention mechanisms, which impose equal importance on consecutive frames might lead to feature homogenization. In an attempt to address these problems, we propose a Dual-Spiral Pyramid Network with Multi-scale Attention (DSPA), which consists of three main modules: Feature Enhancement Module (FEM), Dual-Spiral Feature Pyramid Network (Ds-FPN), and Multi-Scale Dual-Spiral Attention Convolution Module (Ds-MAC). To be specific, we use the FEM to enhance features by exploring the relations on different temporal dimensions and channel dimensions. Moreover, the Ds-FPN integrates high-resolution temporal features from the base layer with fine-grained features processed by the FEM and sequentially propagates these fused features across adjacent layers to construct a hierarchical multi-scale video representation. Furthermore, the Ds-MAC adopts a hierarchical architecture with long-term and short-term temporal modeling and residual learning to capture global context and fine details, while enhancing feature diversity and reducing convergence risk by advanced nonlinear transformations. Extensive experimental results demonstrate that our approach achieves state-of-the-art performance on two public datasets, THUMOS14 and EPIC-Kitchens 100. Haiyu Liu, Hexiong Yang, Keyan Ren |
IJCNN | 5 |
| 2025 | PGLRNet: target pose-guided and feature loss-reduced network for oriented object detection in remote sensing images
Haiyu Liu, Keyan Ren, Zhiyu Nie |
Vis. Comput. | 3 |
| 2024 | Smart Feature is What You NeedabstractLack of shape guidance and label jitter caused by information deficiency of weak label are the main problems in 3D weakly-supervised object detection. Current weakly-supervised models often use heuristics or assumptions methods to infer information from weak labels without taking advantage of the inherent clues of weakly-supervised and fully-supervised methods, thus it is difficult to explore a method that combines data utilization efficiency and model accuracy. In an attempt to address these issues, we propose a novel plug-and-in point cloud feature representation network called Multi-scale Mixed Attention (MMA). MMA utilizes adjacency attention within neighborhoods and disparity attention at different density scales to build a feature representation network. The smart feature representation obtained from MMA has shape tendency and object existence area inference, which can constrain the region of the detection boxes, thereby alleviating the problems caused by the information default of weak labels. Extensive experiments show that in indoor weak label scenarios, the fully-supervised network can perform close to that of the weakly-supervised network merely through the improvement of point feature by MMA. At the same time, MMA can turn waste into treasure, reversing the label jitter problem that originally interfered with weakly-supervised detection into the source of data enhancement, strengthening the performance of existing weak supervision detection methods. Our code is available at https://github.com/hzx-9894/MMA. Zhaoxin Hu, Keyan Ren |
IJCNN | 2 |
| 2024 | Confidence sharing adaptation for out-of-domain human pose and shape estimation
Tianyi Yue, Keyan Ren, Qingyun Bian |
Comput. Vis. Image Underst. | 2 |
| 2024 | Ssman: self-supervised masked adaptive network for 3D human pose estimation
Tianyi Yue, Guoping He, Keyan Ren |
Mach. Vis. Appl. | 5 |
| 2024 | ASGSA: global semantic-aware network for action segmentation
Qingyun Bian, Keyan Ren, Tianyi Yue, Yunlu Zhang |
Neural Comput. Appl. | 3 |
| 2024 | TransFG: A Cross-View Geo-Localization of Satellite and UAVs Imagery Pipeline Using Transformer-Based Feature Aggregation and Gradient GuidanceabstractCross-view geo-localization of satellite and unmanned aerial vehicles (UAVs) imagery has attracted extensive attention due to its tremendous potential for global navigation satellite system (GNSS) denied navigation. However, inadequate feature representation across different views coupled with positional shifts and distance-scale uncertainty are key challenges. Most of the existing research mainly focused on extracting comprehensive and fine-grained information, yet effective feature representation and alignment should be imposed equal importance. In this article, we propose an innovative transformer-based pipeline TransFG for robust cross-view image matching, which incorporates feature aggregation (FA) and gradient guidance (GG) module. TransFG synergically takes advantage of FA and GG, achieving an effective balance in feature representation and alignment. Specifically, the proposed FA module implicitly learns salient features and dynamically aggregates contextual features from the vision transformer (ViT). The proposed GG module uses the gradient information of local features to further enhance the cross-view feature representation and aligns specific instances across different views. Extensive experiments demonstrate that our pipeline outperforms existing methods in cross-view geo-localization. It achieves an impressive improvement in R@1 and AP than the state-of-the-art (SOTA) methods. The code has been released athttps://github.com/happyboy1234/TransFG. Keyan Ren, Tianyi Yue |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Less is More: Decoupled High-Semantic Encoding for Action RecognitionabstractThis paper focuses on how to improve the efficiency of the action recognition framework by optimizing its complicated feature extraction pipelines and enhancing explainability, benefiting future adaptation to more complex visual understanding tasks (e.g. video captioning). To achieve this task, we propose a novel decoupled two-stream framework for action recognition - HSAR, which utilizes high-semantic features for increased efficiency and provides well-founded explanations in terms of spatial-temporal perceptions that will benefit further expansions on visual understanding tasks. The inputs are decoupled into spatial and temporal streams with designated encoders aiming to extract only the pinnacle of representations, gaining high-semantic features while reducing computation costs greatly. A lightweight Temporal Motion Transformer (TMT) module is proposed for globally modeling temporal features through self-attention, omitting redundant spatial features. Decoupled spatial-temporal embeddings are further merged dynamically by an attention fusion model to form a joint high-semantic representation. The visualization of the attention in each module offers intuitive interpretations of HSAR’s explainability. Extensive experiments on three widely-used benchmarks (Kinetics400, 600, and Sthv2) show that our framework achieves high prediction accuracy with significantly reduced computation (only 64.07 GFLOPs per clip), offering a great trade-off between accuracy and computational costs. Keyan Ren, Qingyun Bian |
ICMR | 2 |
| 2023 | CANET: A hierarchical CNN-Attention model for Network Intrusion Detection
Keyan Ren, Zhiqing Huang |
Comput. Commun. | 1 |
| 2023 | A detector for page-level handwritten music object recognition based on deep learning
Zhiqing Huang, Keyan Ren |
Neural Comput. Appl. | 4 |
| 2022 | SG-TCN: Semantic Guidance Temporal Convolutional Network for Action SegmentationabstractWhile recent approaches based on multi-stage temporal convolutional network (TCN) can achieve good accuracy in action segmentation, they cannot get an excellent F1-score, which makes them difficult to be applied in practice. The main issue we investigated is that the TCN lacks the max-pool and hence it is difficult to capture sufficient semantic information which leads to over-segmentation. To reduce the occurrence of over-segmentation, we propose the Semantic Guidance module (SG) to capture high-level semantic features and guide the TCN. In addition, we consider the role of each stage in a multi-stage architecture and deploy a lighter parameter-sharing TCN (PS-TCN) as the backbone, which achieves higher accuracy and reduces about 16% parameters than the most popular backbone. Simultaneously, our proposed Video Speed Prediction module (VSP) explores temporal information and improves temporal modeling ability. Combining PS-TCN with VSP and using SG for guidance yield an accurate and robust segmentation model. Extensive experiments demonstrate that our model is much better than the benchmark MS-TCN++ (e.g. from 45.9% to 56.4% F1@50 on Breakfast) and achieves state-of-the-art performance on two challenging datasets. The code is available at https://github.com/zhangylll/SG-TCN Yunlu Zhang, Keyan Ren, Tong Yan |
IJCNN | 2 |
| 2022 | FCNet: A feature context network based on ensemble framework for image retrievalabstractAbstract In order to train an end‐to‐end model for better performance in distinguishing superclasses and make the model more sensitive to inter‐class differences, the authors propose a Feature Context Network (FCNet), a novel ensemble framework for image retrieval based on clustering (EFC) and metric learning. In the authors' approach, the EFC framework consists of one common feature extractor and multiple learning branches. This tree‐like structure not only promotes the stability of the training process but also keeps the diversity by using multiple branches. Each branch is learnt by an independent random re‐labelling operation that distributes original classes into superclasses. In such a way, the framework thus can accelerate the training process. To make the EFC framework focus on the different attributes across the classes, the authors utilize k‐means clustering to divide the training set into subsets and train the model sequentially with those subsets. Meanwhile, to reduce the impact of over‐fitting problems, the authors design a branch weight structure during training. Further, FCNet fusions different scales of the neighbourhood information and consequently makes the feature map learn richer information with the different receptive fields. Simultaneously, it can reduce the layer of the network and the number of model parameters. Extensive experiments demonstrate that this approach outperforms the tested representative methods on CARS‐196, CUB‐200‐2011 datasets. Keyan Ren, Jianbo Shi |
IET Comput. Vis. | 5 |
| 2022 | LRATD: a lightweight real-time abnormal trajectory detection approach for road traffic surveillance
Keyan Ren |
Neural Comput. Appl. | 2 |
| 2018 | Prediction of Ejection State for a Pneumatic Valve-Controlled Micro-Droplet Generator by a BP Neural NetworkabstractPneumatic valve-controlled micro-droplet generation is a printing technique that has potential applications in many fields, especially in the field of biomedical printing. The droplet generation is controlled by a solenoid valve being briefly turned on, so that high pressure gas enters the liquid reservoir, forming a gas pressure pulse P(t), forcing the liquid out through a tiny nozzle to form a micro-droplet. Under the typical working conditions, P(t) is not consistent. Since P(t) is highly correlated with the micro-droplet ejection state, the inconsistency of P(t) results in fluctuation of ejection state. For each injection, the P(t) is acquired by a high speed pressure sensor, and the ejection state is obtained by machine vision processing. A machine learning method based on BP neural network is used to establish a prediction model with P(t) as the input and the droplet ejection state as the output. Experiments show that a BP neural network with only a single hidden layer and two neurons can accurately predict the number of droplets with an accuracy higher than 99%. Another experiment shows that a more complex double hidden layer BP neural network can improve the prediction accuracy for the position of droplets after a certain time delay. In summary, through pressure pulse P(t), the predictive model established by the machine learning method can effectively predict the micro-droplet ejection state. This technique may be used for real time monitoring and control of the pneumatic valve-controlled micro-droplet generator. Jiangeng Li, Weijie Bao, Zhixuan Er, Keyan Ren |
ICMLA | 7 |