EDBT 2026 Demo / reviewers in the wild / expert
Ke Li 0005
dblp:75/6627-5
· DBLP profile ↗
35ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0002-7873-1554ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-domain modulated spatio-temporal graph convolutional network for traffic flow prediction
Mengchao Liu, Bingchuan Jiang, Jingxu Liu, Ke Li 0005 |
Expert Syst. Appl. | 4 |
| 2025 | Dual-BEV Nav: Dual-Layer BEV-Based Heuristic Path Planning for Robotic Navigation in Unstructured Outdoor EnvironmentsabstractPath planning with strong environmental adaptability plays a crucial role in robotic navigation in unstructured outdoor environments, especially in the case of low-quality location and map information. The path planning ability of a robot depends on the identification of the traversability of global and local ground areas. In real-world scenarios, the complexity of outdoor open environments makes it difficult for robots to identify the traversability of ground areas that lack a clearly defined structure. Moreover, most existing methods have rarely analyzed the integration of local and global traversability identifications in unstructured outdoor scenarios. To address this problem, we propose a novel method, Dual-BEV Nav, first introducing Bird's Eye View (BEV) representations into local planning to generate high-quality traversable paths. Then, these paths are projected into the global traversability probability map generated by the global BEV planning model to obtain the optimal path. By integrating the traversability from both local and global BEV, we establish a dual-layer BEV heuristic planning paradigm, enabling long-distance navigation in unstructured outdoor environments. We test our approach through both public dataset evaluations and real-world robot deployments, yielding promising results. Compared to baselines, the Dual-BEV Nav improved temporal distance prediction accuracy by up to 18.26%. In the real-world deployment, under conditions significantly different from the training set and with notable occlusions in the global BEV, the Dual-BEV Nav successfully achieved a 65-meter-long outdoor navigation. Further analysis demonstrates that the local BEV representation significantly enhances the rationality of the planning, while the global BEV probability map ensures the robustness of the overall planning. Jian Yang 0034, Shibo Huang, Ke Li 0005, Xian Wei, Xiong You |
ICRA | 6 |
| 2025 | KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor EnvironmentsabstractAutonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large language models (LLMs) enhance semantic comprehension but lack spatial reasoning capabilities. Although diffusion models excel in local optimization, they fall short in large-scale long-distance navigation. To address these gaps, this paper proposes KiteRunner, a language-driven cooperative local-global navigation strategy that combines UAV orthophoto-based global planning with diffusion model-driven local path generation for long-distance navigation in open-world scenarios. Our method innovatively leverages real-time UAV orthophotography to construct a global probability map, providing traversability guidance for the local planner, while integrating large models like CLIP and GPT to interpret natural language instructions. Experiments demonstrate that KiteRunner achieves 5.6% and 12.8% improvements in path efficiency over state-of-the-art methods in structured and unstructured environments, respectively, with significant reductions in human interventions and execution time. Shibo Huang, Chenfan Shi, Jian Yang 0034, Jinpeng Mi, Ke Li 0005, Miao Ding, Peidong Liang, Xiong You, Xian Wei |
IROS | 6 |
| 2025 | NeuroLoc: Encoding Navigation Cells for 6-DOF Camera LocalizationabstractRecently, camera localization has been widely adopted in autonomous robotic navigation due to its efficiency and convenience. However, autonomous navigation in unknown environments often suffers from scene ambiguity, environmental disturbances, and dynamic object transformation in camera localization. To address this problem, inspired by the brain cognitive navigation mechanism (such as grid cells, place cells, and head direction cells), we propose a novel neurobiological camera location method, namely NeuroLoc. Firstly, we designed a Hebbian learning module driven by place cells to save and replay historical information, aiming to restore the details of historical representations and solve the issue of scene fuzziness. Secondly, we utilized the head direction cell-inspired internal direction learning as multi-head attention embedding to help restore the true orientation in similar scenes. Finally, we added a 3D grid center prediction in the pose regression module to reduce the final wrong prediction. We evaluate the proposed NeuroLoc on commonly used benchmark indoor and outdoor datasets. The experimental results show that our NeuroLoc can enhance the robustness in complex environments and improve the performance of pose regression by using only a single image. Jian Yang 0034, Fenli Jia, Muyu Wang, Jinpeng Mi, Jilin Hu, Peidong Liang, Ke Li 0005, Xiong You, Xian Wei |
IROS | 10 |
| 2025 | Multi-Scale Oriented Object Detection With Focus Error Ellipse LossabstractThe loss function and feature extraction framework are essential parts of the algorithm design and significantly affect the accuracy of oriented object detection in remote sensing images. Though considerable progress has been made, there are still challenges left to be explored, e.g., large variations in scales, arbitrary direction, and dense distribution of the objects, which may have some undesirable effects, such as inaccurate object position regression, high false alarm, and miss rate. To address the above problems, we propose a Focus Error Ellipse (FEE) loss function. This function bolsters the detection accuracy by narrowing the distance between the center points of the labeled and predicted bounding boxes based on the Error Ellipse. For the network part, we carefully crafted two unit modules: a Fine-grained and Context-augmented Module (FCM) and a Semantic Information Regrouping Module (SIRM). The FCM aligns fine-grained information with contextual information to establish dependencies between local and global features, which helps grasp the more holistic characteristics of objects. The SIRM reorganizes the acquired deep semantic features in the channel dimension, enhances the weight of task-beneficial semantic information, and further derives the optimal combination method of feature subsets for object detection. Based on the aforementioned work, we developed an oriented object detection framework, which further improves the detection accuracy of large aspect ratio objects and dense scenes. Experimental results show that the proposed method can produce competitive performance in oriented object detection compared to other state-of-the-art models. Xuanbei Lu, Ke Li 0005, Gong Cheng 0003, Xiong You |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Elevation Correction of a Large-Scale DEM Using ICESat-2 Laser Altimetry DataabstractThe digital elevation model (DEM) is a digital representation of the surface elevation, but it contains elevation errors arising from vegetation coverage and terrain undulation. Spaceborne LiDAR, with its large-scale and high-precision elevation measurement capabilities, can effectively correct these errors and has become an important means to improve the elevation accuracy of DEM. In this study, a hybrid incremental regression (HIR) model based on the Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) laser altimetry data is designed to correct large-scale DEM elevation errors. The model combines a neural network and a decision tree model to efficiently fit the elevation error through stepwise regression, which can be applied to different terrains and landforms over large areas. In the experiment, first, the proposed model was compared with four classic models [multiple linear regression (MLR), random forest (RF), backpropagation neural network (BPNN) and light gradient-boosting machine (LightGBM)] in seven different landform areas, and the results showed that our model was better than other models in terms of accuracy improvement and was applicable to various terrains. Then, the model was applied to China, and the higher precision, large-scale DEM datasets were produced. Compared with the original DEM, the root mean squared error (RMSE) and mean absolute error (MAE) of the optimized DEM were reduced by 1.603 and 1.453 m, respectively. The differences in accuracy under different slopes, land cover types, and vegetation heights were also analyzed, and it was found that the enhancement effect was the best in the high-vegetation cover area. Finally, it is validated using three regions of high-precision validation data, and the results showed that the RMSE had improved to varying degrees, ranging from 0.193 to 2.853 m. Weiqi Lian, Ziqi Nie, Guo Zhang 0001, Ke Li 0005, Xuefeng Cao, Anzhu Yu, Xin Li 0103 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | FuseFormer: A Manifold Metric Fusing Attention for Pedestrian Trajectory PredictionabstractAccurate pedestrian trajectory prediction is critical for ensuring the safety of autonomous vehicles and advancing higher levels of driving automation. However, the complex interpersonal interactions and highly dynamic trajectory patterns in real-world scenarios pose significant challenges to achieving precise predictions. Recently, Transformers have shown remarkable success in pedestrian trajectory prediction, primarily due to their effective modeling of temporal and spatial dependencies via Multi-Head Self-Attention (MHA) mechanisms. Despite these advancements, existing self-attention methods often rely on Euclidean distance-based metrics and dot-product operations, which are inadequate for capturing interaction-induced trajectory curvatures. To address this limitation, we propose a novel hybrid Transformer architecture, FuseFormer, that incorporates Geodesic Self-Attention (GSA) mechanisms. GSA utilizes geodesic distances to characterize interaction features effectively, complementing MHA, which excels in capturing local features and maintaining temporal correlations. FuseFormer employs a gating network to adaptively combine GSA and MHA embeddings, leveraging their complementary strengths. Additionally, FuseFormer integrates a Transformer-based Neural Ordinary Differential Equation (ODE) decoder to model trajectory temporal dynamics. This design enables the generation of future trajectories that align closely with motion trends while adapting the network depth to input sequence lengths. Experimental results demonstrate that FuseFormer achieves state-of-the-art performance across widely used pedestrian trajectory prediction datasets, including ETH/UCY, SDD, and NBA. These results underscore the model’s effectiveness and generalization capability in capturing complex interaction patterns and handling diverse scenarios. Kohsin Ko, Jian Yang 0034, Ke Li 0005, Xiong You, Jinpeng Mi, Mingsong Chen 0001, Xian Wei |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous DrivingabstractWith the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads, lane lines, vehicles and pedestrians) can be assembled into progressively more complex shapes to form a BEV representation of the whole road scene. Therefore, a BEV often has multiple levels of freedom on motion, i.e., the rotation and the moving shift of the whole BEV, and the random movements of objects (e.g., pedestrians and vehicles) inside the BEV. However, most of the current single-sensor or multi-sensor fusion-based BEV object detection methods have not yet taken into account capturing such multi-level motion in a BEV. To address this problem, we propose a product group equivariant object detection network framework that is equivariant with respect to multiple levels of symmetry groups based on multi-sensor fusion. The proposed framework extracts local equivariant features of objects in point clouds, while global equivariant features are extracted in both point clouds and images. Furthermore, the network learns diverse rotation-equivariant features and mitigates a significant amount of detection errors caused by rotations of BEV and objects inside a BEV, thereby further enhancing the performance of object detection. The experiment results show that the network architecture significantly improves object detection on mAP and NDS, respectively. In addition, in order to demonstrate the effectiveness of the proposed local-multi-global equivariant components, we conduct sufficient ablation experiments. The results show that the individual components are indispensable for the object detection performance improvement of the overall network architecture. Jian Yang 0034, Ke Li 0005, Jianzhang Zheng, Xihao Wang, Mingsong Chen 0001, Xiong You, Xian Wei |
ICRA | 4 |
| 2024 | Fewer is more: efficient object detection in large aerial images
Xingxing Xie, Gong Cheng 0003, Qingyang Li 0001, Shicheng Miao, Ke Li 0005, Junwei Han 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | Oriented R-CNN and Beyond
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xiwen Yao, Junwei Han 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Attention Prompt-Driven Source-Free Adaptation for Remote Sensing Images Semantic SegmentationabstractRecently, remote sensing images (RSIs) domain adaptation segmentation has been extensively studied. However, existing methods generally assume that source RSIs must be available, which is obviously an overly demanding condition and will increase unnecessary costs in practice. To this end, this letter takes the lead in exploring RSIs source-free adaptation segmentation, where only the offline model pretrained on the source domain and target RSIs are available. A novel method featuring prompt learning and vision foundation models is proposed, and the novelty design includes two aspects. First, to better adapt the general-purpose knowledge in the foundation model to different target RSIs, an attention-guided prompt tuning strategy is proposed, which can dynamically steer the knowledge at different layers and positions through prompts with different weights. Second, a feature alignment strategy with similarity distance is proposed for source-free domain adaptation by taking full advantage of the representation ability of the foundation model and the flexibility of prompt learning. Extensive experiments indicate that the performance of the proposed method is significantly superior to that of existing methods. Specifically, the mIoU of target RSIs has been improved by at least 3.14%~4.18%. Kuiliang Gao, Xiong You, Ke Li 0005, Juan Lei, Xibing Zuo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Integrating Multiple Sources Knowledge for Class Asymmetry Domain Adaptation Segmentation of Remote Sensing ImagesabstractIn the existing unsupervised domain adaptation (UDA) methods for remote sensing images (RSIs) semantic segmentation, class symmetry is a widely followed ideal assumption, where the source and target RSIs have exactly the same class space. In practice, however, it is often very difficult to find a source RSI with exactly the same classes as the target RSI. More commonly, there are multiple source RSIs available. And there is always an intersection or inclusion relationship between the class spaces of each source–target pair, which can be referred to as class asymmetry. Nevertheless, the class asymmetry domain adaptation segmentation of RSIs with multiple sources has not yet been explored. To this end, a novel class asymmetry RSIs domain adaptation method is proposed for the first time in this article, which consists of four key components. First, a multibranch segmentation network is built to learn an expert for each source RSI. Second, a novel collaborative learning method with the cross-domain mixing strategy is proposed, to supplement the class information for each source while achieving the domain adaptation of each source–target pair. Third, a pseudolabel generation strategy is proposed to effectively combine the strengths of different experts, which can be flexibly applied to two cases where the source class union is equal to or includes the target class set. Fourth, a multiview-enhanced knowledge integration module is developed for high-level knowledge routing and transfer from multiple domains to target predictions. The experimental results of six different class settings on airborne and spaceborne RSIs show that the proposed method can effectively perform the multisource domain adaptation in the case of class asymmetry, and the obtained segmentation performance of target RSIs is significantly better than the existing relevant methods. Kuiliang Gao, Anzhu Yu, Xiong You, Wenyue Guo, Ke Li 0005, Ningbo Huang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Practical Cross-System Shilling Attacks with Limited Access to DataabstractIn shilling attacks, an adversarial party injects a few fake user profiles into a Recommender System (RS) so that the target item can be promoted or demoted. Although much effort has been devoted to developing shilling attack methods, we find that existing approaches are still far from practical. In this paper, we analyze the properties a practical shilling attack method should have and propose a new concept of Cross-system Attack. With the idea of Cross-system Attack, we design a Practical Cross-system Shilling Attack (PC-Attack) framework that requires little information about the victim RS model and the target RS data for conducting attacks. PC-Attack is trained to capture graph topology knowledge from public RS data in a self-supervised manner. Then, it is fine-tuned on a small portion of target data that is easy to access to construct fake profiles. Extensive experiments have demonstrated the superiority of PC-Attack over state-of-the-art baselines. Our implementation of PC-Attack is available at https://github.com/KDEGroup/PC-Attack. Meifang Zeng, Ke Li 0005, Bingchuan Jiang, Liujuan Cao, Hui Li 0057 |
AAAI | 2 |
| 2023 | Mutual-Assistance Learning for Object DetectionabstractObject detection is a fundamental yet challenging task in computer vision. Despite the great strides made over recent years, modern detectors may still produce unsatisfactory performance due to certain factors, such as non-universal object features and single regression manner. In this paper, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a robust one-stage detector, referred as MADet, to address these weaknesses. First, the spirit of MA is manifested in the head design of the detector. Decoupled classification and regression features are reintegrated to provide shared offsets, avoiding inconsistency between feature-prediction pairs induced by zero or erroneous offsets. Second, the spirit of MA is captured in the optimization paradigm of the detector. Both anchor-based and anchor-free regression fashions are utilized jointly to boost the capability to retrieve objects with various characteristics, especially for large aspect ratios, occlusion from similar-sized objects, etc. Furthermore, we meticulously devise a quality assessment mechanism to facilitate adaptive sample selection and loss term reweighting. Extensive experiments on standard benchmarks verify the effectiveness of our approach. On MS-COCO, MADet achieves 42.5% AP with vanilla ResNet50 backbone, dramatically surpassing multiple strong baselines and setting a new state of the art. Xingxing Xie, Chunbo Lang, Shicheng Miao, Gong Cheng 0003, Ke Li 0005, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Perturbation-Seeking Generative Adversarial Networks: A Defense Framework for Remote Sensing Image Scene ClassificationabstractThe methods for remote sensing image (RSI) scene classification based on deep convolutional neural networks (DCNNs) have achieved prominent success. However, confronted with adversarial examples obtained by adding imperceptible perturbations to clean images, the great vulnerability of DCNNs makes it worth exploring effective defense methods. To date, numerous countermeasures for adversarial examples have been proposed, but how to improve the defensive ability for unknown attacks still to be answered. To address this issue, in this article, we propose an effective defense framework specified for RSI scene classification, named perturbation-seeking generative adversarial networks (PSGANs). In brief, a new training framework is designed to train the classifier by introducing the examples generated during the image reconstruction process, in addition to clean examples and adversarial ones. These generated examples can be random kinds of unknown attacks during training and thus are utilized to eliminate the blind spots of a classifier. To assist the proposed training framework, a reconstruction method is developed. First, instead of modeling the distribution of clean examples, we model the distributions of the perturbations added in adversarial examples. Second, to make a tradeoff between the diversity of the reconstructed examples and the optimization of PSGAN, a scale factor named seeking radius is introduced to scale the generated perturbations before they are subtracted by the given adversarial examples. Comprehensive and extensive experimental results on three widely used benchmarks for RSI scene classification demonstrate the great effectiveness of PSGAN when faced with both known and unknown attacks. Our source code is available athttps://github.com/xuxiangsun/PSGAN. Gong Cheng 0003, Xuxiang Sun 0001, Ke Li 0005, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Anchor-Free Oriented Proposal Generator for Object DetectionabstractOriented object detection is a practical and challenging task in remote sensing image interpretation. Nowadays, oriented detectors mostly use horizontal boxes as intermedium to derive oriented boxes from them. However, the horizontal boxes are inclined to get small Intersection-over-Unions (IoUs) with ground truths, which may have some undesirable effects, such as introducing redundant noise, mismatching with ground truths, detracting from the robustness of detectors, etc. In this paper, we propose a novel Anchor-free Oriented Proposal Generator (AOPG) that abandons horizontal box-related operations from the network architecture. AOPG first produces coarse oriented boxes by a Coarse Location Module (CLM) in an anchor-free manner and then refines them into high-quality oriented proposals. After AOPG, we apply a Fast R-CNN head to produce the final detection results. Furthermore, the shortage of large-scale datasets is also a hindrance to the development of oriented object detection. To alleviate the data insufficiency, we release a new dataset on the basis of our DIOR dataset and name it DIOR-R. Massive experiments demonstrate the effectiveness of AOPG. Particularly, without bells and whistles, we achieve the accuracy of 64.41%, 75.24% and 96.22% mAP on the DIOR-R, DOTA and HRSC2016 datasets respectively. Code and models are available at https://github.com/jbwang1997/AOPG. Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xingxing Xie, Chunbo Lang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Dual-Aligned Oriented DetectorabstractIn the past few years, object detection in remote sensing images has achieved remarkable progress. However, the detection of oriented and densely packed objects are still unsatisfactory due to the following spatial and feature misalignments. 1) Most two-stage oriented detectors only introduce an orientation regression branch in the detection head, while still leverage horizontal proposals for classification and regression. This inevitably results in the spatial misalignment problem between horizontal proposals and oriented objects. 2) The features used for classification are in fact extracted from the region proposals which have shifted to the final predictions via the regression branch. This leads to the feature misalignment problem between the classification and the localization tasks. In this article, we present a two-stage oriented object detection method, termed dual-aligned oriented detector (DODet), toward evading the aforementioned problems of spatial and feature misalignments. In DODet, the first stage is an oriented proposal network (OPN), which generates high-quality oriented proposals via a novel representation scheme of oriented objects. The second stage is a localization-guided detection head (LDH) that aims at alleviating the feature misalignment between classification and localization. Comprehensive and extensive evaluations on three benchmarks, including DIOR-R, DOTA, and HRSC2016, indicate that our method could obtain consistent and substantial gains compared with the baseline method. The source code is publicly available athttps://github.com/yanqingyao1994/DODet. Gong Cheng 0003, Shengyang Li, Ke Li 0005, Xingxing Xie, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Prototype-CNN for Few-Shot Object Detection in Remote Sensing ImagesabstractRecently, due to the excellent representation ability of convolutional neural networks (CNNs), object detection in remote sensing images has undergone remarkable development. However, when trained with a small number of samples, the performance of the object detectors drops sharply. In this article, we focus on the following three main challenges of few-shot object detection in remote sensing images: 1) since the sample number of novel classes is far less than base classes, object detectors would fail to quickly adapt to the features of novel classes, which would result in overfitting; 2) the scarcity of samples in novel classes leads to a sparse orientation space, while the objects in remote sensing images usually have arbitrary orientations; and 3) the distribution of object instances in remote sensing images is scattered and, therefore, it is hard to identify foreground objects from the complex background. To tackle these problems, we propose a simple yet effective method named prototype-CNN (P-CNN), which mainly consists of three parts: a prototype learning network (PLN) converting support images to class-aware prototypes, a prototype-guided region proposal network (P-G RPN) for better generation of region proposals, and a detector head extending the head of Faster region-based CNN (R-CNN) to further boost the performance. Comprehensive evaluations on the large-scale DIOR dataset demonstrate the effectiveness of our P-CNN. The source code is available athttps://github.com/Ybowei/P-CNN. Gong Cheng 0003, Bowei Yan, Peizhen Shi, Ke Li 0005, Xiwen Yao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | RTSfM: Real-Time Structure From Motion for Mosaicing and DSM Mapping of Sequential Aerial Images With Low OverlapabstractInspired by simultaneous localization and mapping (SLAM) style workflow, this article presented an online sequential structure from motion (SfM) solution for high-frequency video and large baseline high-resolution aerial images with high efficiency and novel precision. First, as traditional SLAM systems are not good in processing low overlap images, based on our novel hierarchical feature matching paradigm with multihomography and BoW, we proposed a robust tracking method where the relative pose and its scale are estimated separately followed by a joint optimization by considering both perspective-n-point (PnP) and epipolar constraints. Second, to further optimize the camera poses for the sparse map and dense pointcloud reconstruction, we provided a graph-based optimization with reprojection and GPS constraints, which make the camera trajectory and map georeferenced. We also incrementally generated the dense point cloud in real time from keyframes after local mapping optimization. Finally, we use a publicly available aerial image dataset with sequences of different environments, to evaluate the effectiveness of the proposed method, meanwhile, the robust performance of our solution is demonstrated with applications of high-quality aerial images mosaic and digital surface model (DSM) reconstruction in real time. Compared with the state-of-the-art SLAM and traditional SfM methods, the presented system can output large-scale high-quality ortho-mosaic and DSM in real time with the low computational cost. Lin Chen 0042, Xishan Zhang, Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han, Ke Li 0005 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | SPEX: A Generic Framework for Enhancing Neural Social RecommendationabstractSocial Recommender Systems (SRS) have attracted considerable attention since its accompanying service, social networks, helps increase user satisfaction and provides auxiliary information to improve recommendations. However, most existing SRS focus on social influence and ignore another essential social phenomenon, i.e., social homophily. Social homophily, which is the premise of social influence, indicates that people tend to build social relations with similar people and form influence propagation paths. In this article, we propose a generic framework Social PathExplorer (SPEX) to enhance neural SRS. SPEX treats the neural recommendation model as a black box and improves the quality of recommendations by modeling the social recommendation task, the formation of social homophily, and their mutual effect in the manner of multi-task learning. We design a Graph Neural Network based component for influence propagation path prediction to help SPEX capture the rich information conveyed by the formation of social homophily. We further propose an uncertainty based task balancing method to set appropriate task weights for the recommendation task and the path prediction task during the joint optimization. Extensive experiments have validated that SPEX can be easily plugged into various state-of-the-art neural recommendation models and help improve their performance. The source code of our work is available at: https://github.com/XMUDM/SPEX. Hui Li 0057, Lianyun Li, Guipeng Xv, Chen Lin 0001, Ke Li 0005, Bingchuan Jiang |
ACM Trans. Inf. Syst. | 5 |
| 2021 | HMMN: Online metric learning for human re-identification via hard sample mining memory network
Pengcheng Han, Qing Li 0018, Cunbao Ma, Shibiao Xu, Shuhui Bu, Ke Li 0005 |
Eng. Appl. Artif. Intell. | 7 |
| 2021 | Counting trees with point-wise supervised segmentation network
Pinmo Tong, Pengcheng Han, Suicheng Li, Shuhui Bu, Qing Li 0018, Ke Li 0005 |
Eng. Appl. Artif. Intell. | 7 |
| 2021 | PL-VSCN: Patch-level vision similarity compares network for image matchingabstractAbstract Image matching plays an important role in various computer vision tasks, such as image retrieval and loop closure detection in Simultaneous Localization and Mapping. The authors propose a discriminative patch‐based image matching method that converts the problem of whole image matching to that of local patch matching. To construct the patch representation, the Patch‐Level Vision Similarity Compare Network (PL‐VSCN) is proposed to produce the patch feature. In the image matching process, local patches that potentially contain objects within images are initially detected, and the discriminative feature of each patch is extracted based on the pre‐trained PL‐VSCN. Then, the similarities between the patch pairs are calculated to construct the similarity matrix, and the corresponding patch pairs are detected based on the mutual matching mechanism on the similarity matrix. Experimental results indicate that the proposed PL‐VSCN can generate the discriminative patch feature, which can accurately match the patch pairs with the corresponding content and distinguish those with non‐corresponding content. In addition, the comparison experiments demonstrate that the proposed image matching method outperforms existing approaches on most datasets and effectively completes the image matching task. Xiong You, Qin Li 0005, Ke Li 0005, Anzhu Yu, Shuhui Bu |
IET Comput. Vis. | 3 |
| 2020 | Change detection in images using shape-aware siamese convolutional network
Suicheng Li, Pengcheng Han, Shuhui Bu, Pinmo Tong, Qing Li 0018, Ke Li 0005 |
Eng. Appl. Artif. Intell. | 6 |
| 2020 | Mask-CDNet: A mask based pixel change detection network
Shuhui Bu, Qing Li 0018, Pengcheng Han, Pengyu Leng, Ke Li 0005 |
Neurocomputing | 5 |
| 2019 | Aerial image change detection using dual regions of interest networks
Pengcheng Han, Cunbao Ma, Qing Li 0018, Pengyu Leng, Shuhui Bu, Ke Li 0005 |
Neurocomputing | 6 |
| 2019 | Robust visual tracking via identifying multi-scale patches
Yun Liang 0003, Ke Li 0005, Jian Zhang 0026, Meihua Wang, Chen Lin 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong |
Pattern Recognit. | 1 |
| 2018 | Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing ImagesabstractMost of the existing deep-learning-based methods are difficult to effectively deal with the challenges faced for geospatial object detection such as rotation variations and appearance ambiguity. To address these problems, this paper proposes a novel deep-learning-based object detection framework including region proposal network (RPN) and local-contextual feature fusion network designed for remote sensing images. Specifically, the RPN includes additional multiangle anchors besides the conventional multiscale and multiaspect-ratio ones, and thus can deal with the multiangle and multiscale characteristics of geospatial objects. To address the appearance ambiguity problem, we propose a double-channel feature fusion network that can learn local and contextual properties along two independent pathways. The two kinds of features are later combined in the final layers of processing in order to form a powerful joint representation. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method. Ke Li 0005, Gong Cheng 0003, Shuhui Bu, Xiong You |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | 3D shape recognition and retrieval based on multi-modality deep learning
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Ke Li 0005 |
Neurocomputing | 5 |
| 2017 | Learning 3D faces from 2D images via Stacked Contractive Autoencoder
Jian Zhang 0026, Ke Li 0005, Yun Liang 0003 |
Neurocomputing | 2 |
| 2017 | Semi-direct tracking and mapping with RGB-D camera for MAV
Shuhui Bu, Ke Li 0005, Gong Cheng 0003, Zhenbao Liu |
Multim. Tools Appl. | 4 |
| 2016 | Place recognition based on deep feature and adaptive weighting of similarity matrix
Qin Li 0005, Ke Li 0005, Xiong You, Shuhui Bu, Zhenbao Liu |
Neurocomputing | 2 |
| 2014 | High-level semantic feature for 3D shape based on deep belief networksabstractDeep learning has emerged as a powerful technique to extract high-level features from low-level information, which shows that hierarchical representation can be easily achieved. However, applying deep learning into 3D shape is still a challenge. In this paper, we propose a novel high-level feature learning method for 3D shape retrieval based on deep learning. In this framework, the low-level 3D shape descriptors are first encoded into visual bag-of-words, and then highlevel shape features are generated via deep belief network, which facilitates a good semantic preserving ability for the tasks of shape classification and retrieval. Experiments on 3D shape recognition and retrieval demonstrate the superior performance of the proposed method in comparison to the state-of-the-art methods. Zhenbao Liu, Shaoguang Chen, Shuhui Bu, Ke Li 0005 |
ICME | 4 |
| 2014 | Shift-invariant ring feature for 3D shape
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Ke Li 0005, Junwei Han 0001 |
Vis. Comput. | 4 |