EDBT 2026 Demo / reviewers in the wild / expert
Liangjin Zhao
dblp:253/1981
· DBLP profile ↗
13ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-2590-591XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SemStereo: Semantic-Constrained Stereo Matching Network for Remote SensingabstractSemantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pioneering studies either utilize a loosely coupled parallel structure or engage in only implicit interactions, failing to capture the inherent connections. In this work, we explore the connections between the two tasks and propose a new network that imposes semantic constraints on the stereo matching task, both implicitly and explicitly. Implicitly, we transform the traditional parallel structure to a new cascade structure termed Semantic-Guided Cascade structure, where the deep features enriched with semantic information are utilized for the computation of initial disparity maps, enhancing semantic guidance. Explicitly, we propose a Semantic Selective Refinement (SSR) module and a Left-Right Semantic Consistency (LRSC) module. The SSR refines the initial disparity map under the guidance of the semantic map. The LRSC ensures semantic consistency between two views via reducing the semantic divergence after transforming the semantic map from one view to the other using the disparity map. Experiments on the US3D and WHU datasets demonstrate that our method achieves state-of-the-art performance for both semantic segmentation and stereo matching. Chen Chen 0036, Liangjin Zhao, Yuanchun He, Yingxuan Long, Kaiqiang Chen, Zhirui Wang 0003, Yanfeng Hu, Xian Sun 0001 |
AAAI | 2 |
| 2025 | Physics-Guided Deep Learning 3-D Inversion Based on Magnetic DataabstractThe 3-D inversion of magnetic data based on deep learning has achieved great success. This method relies on neural networks to extract features from a large amount of data and then generate structures, with high prediction accuracy and fast speed. However, this data-driven inversion method lacks a closed-form analytical expression and is likened to a “black box.” This feature makes it lack theoretical guidance and has poor interpretability. In addition, the training data often cannot fully cover all aspects of the target, so data-driven methods face the challenge of weak generalization ability in practical applications. Therefore, this letter proposes to integrate physical knowledge into the inversion method based on deep learning. The architecture and loss function of the deep learning model are designed based on the forward modeling of the magnetic field. This can not only improve the interpretability of the algorithm and provide a deeper understanding of the model but also help to make up for the lack of generalization capabilities of deep learning methods. Xiaoqing Shi, Zhirui Wang 0003, Xue Lu, Peirui Cheng, Luning Zhang, Liangjin Zhao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | SiamTHN: Siamese Target Highlight Network for Visual TrackingabstractSiamese network based trackers develop rapidly in the field of visual object tracking in recent years. The majority of Siamese network based trackers now in use treat each channel in the feature maps generated by the backbone network equally, making the similarity response map sensitive to background influence and hence challenging to focus on the target region. Additionally, there are no structural links between the classification and regression branches in these trackers, and the two branches are optimized separately during training. Therefore, there is a misalignment between the classification and regression branches, which results in less accurate tracking results. In this paper, a Target Highlight Module is proposed to help the generated similarity response maps to be more focused on the target region. To reduce the misalignment and produce more precise tracking results, we propose a corrective loss to train the model. The two branches of the model are jointly tuned with the use of corrective loss to produce more reliable prediction results. Experiments on 5 challenging benchmark datasets reveal that the method outperforms current models in terms of performance, and runs at 38 fps, proving its effectiveness and efficiency. Jiahao Bao, Kaiqiang Chen, Xian Sun 0001, Liangjin Zhao, Wenhui Diao, Menglong Yan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | UCDNet: Multi-UAV Collaborative 3-D Object Detection Network by Reliable Feature MappingabstractMulti-unmanned aerial vehicle (UAV) collaborative 3-D object detection can comprehend complex environments by integrating complementary information, with applications encompassing traffic monitoring, delivery services, and agricultural management. However, the extremely broad observations in aerial remote sensing and significant perspective differences across multiple UAVs make it challenging to achieve precise and consistent feature mapping from 2-D images to 3-D space in multi-UAV collaborative 3-D object detection paradigm. To address the problem, we propose an unparalleled camera-based multi-UAV collaborative 3-D object detection paradigm called UCDNet. Specifically, the depth information from the UAVs to the ground is explicitly utilized as a strong prior to provide a reference for more accurate and generalizable feature mapping. Additionally, we design a homologous point geometric consistency loss as an auxiliary self-supervision, which directly influences the feature mapping module, thereby strengthening the global consistency of multiview perception. Experiments on AeroCollab3D and CoPerception-UAVs datasets show that our method increases 4.7% and 10% mean Average Precision (mAP) respectively compared to the baseline, which demonstrates the superiority of UCDNet. Pengju Tian, Zhirui Wang 0003, Peirui Cheng, Zhechao Wang, Liangjin Zhao, Menglong Yan, Xue Yang 0005, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | AVCPNet: An AAV-Vehicle Collaborative Perception Network for 3-D Object DetectionabstractWith the advancement of collaborative perception, the role of autonomous aerial vehicle (AAV)–vehicle collaborative perception has become increasingly significant. The demand for collaborative perception from various perspectives to construct comprehensive perceptual information is rising. However, challenges emerge due to differences in the field of view (FOV) between cross-domain agents and their varying sensitivities to image information. Furthermore, accurate depth information is essential for collaboration to transform image features into bird’s eye view (BEV) features. To address these challenges, we propose a framework specifically designed for aerial-ground collaboration. First, to address the deficiency of datasets for aerial-ground collaboration, we have developed a virtual dataset named V2U-COO for our research. Second, we design a cross-domain cross-adaptation (CDCA) module to align the target information obtained from different domains, thereby achieving more accurate perception results. Finally, we introduce a collaborative depth optimization (CDO) module to obtain more precise depth estimation results, leading to more accurate perception results. We conduct extensive experiments on both our virtual dataset and a public dataset to validate the effectiveness of our framework. Our method resolves the feature fusion issue under significant height differences, a challenge that previous BEV generation methods struggled to address effectively. Our experiments on the V2U-COO and DAIR-V2X datasets demonstrate improvements in detection accuracy of 6.1% and 2.7%, respectively. Our code will be released athttps://github.com/wyccoo/uvcp. Zhirui Wang 0003, Peirui Cheng, Pengju Tian, Ziyang Yuan, Liangjin Zhao |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | RingMo-Galaxy: A Remote Sensing Distributed Foundation Model for Diverse Downstream TasksabstractRemote sensing lightweight foundation models have successfully achieved online perception, providing real-time intelligent interpretation. However, their capabilities are restricted to inferences solely based on their respective observations and models, thus lacking a comprehensive understanding of large-scale remote sensing scenarios. To address this limitation, we propose RingMo-Galaxy, a remote sensing distributed foundation model based on generalized information mapping and interaction. RingMo-Galaxy can realize online collaborative perception across multiple platforms and diverse downstream tasks by mapping observations into a unified space and implementing a task-agnostic information interaction strategy. Specifically, we leverage the ground-based geometric prior of remote sensing oblique observations to change feature mapping from absolute to relative depth estimation, thereby enhancing the model’s ability to extract generalized features across diverse heights and perspectives. In addition, we present a dual-branch information compression module to decouple high-frequency and low-frequency features, achieving feature-level compression while preserving critical task-agnostic details. To support our research, we collect a multitask simulation dataset named AirCo-MultiTasks, specifically designed for multi-unmanned aerial vehicle (UAV) collaborative observation. We also conduct extensive experiments, including 3-D object detection, instance segmentation, and trajectory prediction. The numerous results demonstrate that our proposed RingMo-Galaxy achieves state-of-the-art performance across various downstream tasks. Zhechao Wang, Zhirui Wang 0003, Peirui Cheng, Liangjin Zhao, Pengju Tian, Mingxin Chen, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | TeCCo: A Terminal-Cloud Cross-Domain Collaborative Framework for Remote Sensing Image ClassificationabstractThe terminal-cloud collaborative framework boosts precision and efficiency by integrating cloud computing power with low-latency terminal responsiveness, offering a suitable solution for the growing demands of multiplatform remote sensing (RS) image interpretation. However, the significant differences in data distribution across various RS platforms present a great challenge in balancing the cloud’s centralized processing capabilities with the local interpretation abilities of different terminals. To address this challenge, we propose a terminal-cloud cross-domain collaborative (TeCCo) framework that inherits the efficiency advantages of multiple platforms while ensuring high-accuracy interpretation of diverse data distributions from different terminals. First, the dual classifier co-learning (DCCL) module is designed to enhance cloud robustness. By combining a multilayer perceptron for instance-level classification and a graph convolutional network (GCN) for feature-level aggregation, it achieves mutual supervision and improves feature alignment across different data distributions. Second, the hypernetwork personalization (HNP) module is introduced to generate personalized classifier parameters for each terminal with little fine-tuning cost, allowing terminals to maintain their uniqueness while benefiting from the generalization advantages of collaborative training. Finally, a data-assisted progressive inference mechanism is proposed to enhance accuracy by jointly clustering the features transmitted from terminals and the features of supervised data in the cloud. Extensive experiments demonstrate that TeCCo effectively addresses data distribution challenges, enhancing both the generalization of the cloud model and the personalization of terminal models, achieving state-of-the-art (SOTA) performance in cross-domain and multiplatform RS image classification. Peirui Cheng, Liangjin Zhao, Zhirui Wang 0003, Lingyu Kong, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | FCIL-MSN: A Federated Class-Incremental Learning Method for Multisatellite NetworksabstractMulti-satellite networks have become the prevalent mode for remote sensing intelligent interpretation, with the onboard models requiring class-incremental updates to accommodate the new categories emerging in evolving data and tasks. Traditional model updating methods, which involve uploading models separately after ground-based updating, are inefficient due to limited uplink bandwidth and cumbersome ground update processes while underutilizing potential computing resources on satellites. To address the aforementioned problems, this paper innovatively proposes a collaborative in-orbit incremental update method termed FCIL-MSN, which leverages observational information and computing resources from multi-satellite networks. Firstly, FCIL-MSN achieves collaborative onboard model updates by introducing federated class-incremental learning into multi-satellite networks. Secondly, a bias calibration-guided relationship distillation module constructs a pseudo-feature set by collaborative multi-satellite networks, which alleviates the model bias caused by class imbalance from a global perspective, thereby enhancing model performance. Finally, a gradient information aggregation module is designed to facilitate the exclusion of unfavorable local updates by measuring the contribution of each terminal, thereby accelerating the convergence while obtaining the global model. We conduct extensive experiments on two datasets for scene classification tasks to verify the effectiveness of our proposed method. Experimental results demonstrate that FCIL-MSN outperforms existing general FCIL methods, improving average classification accuracy by 1.45% and decreasing the performance degradation rate by 6.40%. Ziqing Niu, Peirui Cheng, Zhirui Wang 0003, Liangjin Zhao, Xian Sun 0001, Zhi Guo |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid FrameworkabstractIn recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, a RS lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process on several tasks like classification, object detection, semantic segmentation, and change detection. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, a novelty-designed frequency-domain masked image modeling (FD-MIM) is employed during the pretraining stage for self-supervised learning, which combines the high-frequency and low-frequency characteristics of each image patch. This approach effectively captures the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future. Yuelei Wang, Liangjin Zhao, Zhechao Wang, Ziqing Niu, Peirui Cheng, Kaiqiang Chen, Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Elevation Estimation-Driven Building 3-D Reconstruction From Single-View Remote Sensing ImageryabstractBuilding 3D reconstruction from remote sensing images has a wide range of applications in smart cities, photogrammetry and other fields. Methods for automatic 3D urban building modeling typically employ multi-view images as input to algorithms to recover point clouds and 3D models of buildings. However, such models rely heavily on multi-view images of buildings, which are time-intensive and limit the applicability and practicality of the models. To solve these issues, we focus on designing an efficient DSM estimation-driven reconstruction framework (Building3D), which aims to reconstruct 3D building models from the input single-view remote sensing image. Existing DSM estimation networks suffer from the imbalance between local features and global features, which leads to over-smooth DSM estimates at instance boundaries. To address this issue, we propose a Semantic Flow Field-guided DSM Estimation (SFFDE) network, which utilizes the proposed concept of elevation semantic flow to achieve the registration of local and global features. First, in order to make the network semantics globally aware, we propose an Elevation Semantic Globalization (ESG) module to realize the semantic globalization of instances. Further, in order to alleviate the semantic span of global features and original local features, we propose a Local-to-Global Elevation Semantic Registration (L2G-ESR) module based on elevation semantic flow. Our Building3D is rooted in the SFFDE network for building elevation prediction, synchronized with a building extraction network for building masks, and then sequentially performs point cloud reconstruction and surface reconstruction (or CityGML model reconstruction). On this basis, our Building3D can optionally generate CityGML models or surface mesh models of the buildings. Extensive experiments on ISPRS Vaihingen and DFC2019 datasets on the DSM estimation task show that our SFFDE significantly improves upon state-of-the-art and δ1, δ2and δ3metrics of our SFFDE are improved to 0.595, 0.897 and 0.970. Furthermore, our Building3D achieves impressive results in the 3D point cloud and 3D model reconstruction process. Yongqiang Mao, Kaiqiang Chen, Liangjin Zhao, Deke Tang, Wenjie Liu 0016, Zhirui Wang 0003, Wenhui Diao, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | ASSD: Feature Aligned Single-Shot Detection for Multiscale Objects in Aerial ImageryabstractObject detection is a fundamental part of the interpretation of remote sensing imagery. The one-stage object detector has been adopted into this field because of its high computational efficiency. However, this detector suffers from the misalignment among predefined anchor, object, and feature extracted by standard convolution kernel both in spatial and scale. It limits the further improvement of performance, especially for the long-narrow and multiscale geospatial objects. In this article, the problem is defined asthe feature misalignmentproblem. To deal with this issue, an efficient feature aligned single-shot detector (ASSD) is proposed, which consists of two modules: a novel pseudo anchor proposal module (PAPM) and a flexible context-based feature alignment module (CFAM). The PAPM replaces the regular anchor group with the proposed core anchor and refines it to get aligned locations. It can tackle the spatial misalignment between anchors and their corresponding objects and alleviate the negative/positive imbalance problem. Then, the CFAM adaptively adjusts the sampling points of the convolution kernel and collects the context information according to the aligned core anchor. This plug-and-play module can effectively rectify the misalignment between kernel and objects and extract aligned and robust features. A series of comprehensive experiments are conducted on two large-scale public remote sensing object detection datasets. Experiment results suggest that the proposed method is effective to alleviate the misalignment problem. Compared with the baseline model, the detection accuracy is improved by 8.5% mAP and 11.0% mAP on the challenging benchmark for object detection in optical remote sensing image (DIOR) and a large-scale dataset for object detection in aerial image (DOTA) dataset, respectively. Our best-resulting model achieves the state-of-the-art performance, surpassing other one-stage detectors both on the two datasets at a high detection speed of 21 FPS. Tao Xu 0053, Xian Sun 0001, Wenhui Diao, Liangjin Zhao, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | FADA: Feature Aligned Domain Adaptive Object Detection in Remote Sensing ImageryabstractDeep learning-based object detectors have been widely adopted in the field of remote sensing imagery interpretation. These detectors heavily depend on the expensive large-scale labeled datasets, while the scarce remote sensing datasets limit the performance. The domain adaptive object detection can alleviate this problem. However, it struggles with the confusing feature’s alignment, damaging the domain generalization performance, especially for the remote sensing scene with sparse objects and diverse backgrounds. For that reason, a semisynthetic data generator (SDG) is proposed to automatically generate the remote sensing dataset with low cost and replace the real-world training dataset, afeature aligned domain adaptive object detector(FADA) is proposed to enhance the domain adaptation among the cross-domain remote sensing images. The FADA contains two proposed modules in addition to the base detector: an adversarial-based foreground alignment (AFA) and a prototype-based confusing feature alignment (PCFA). The AFA aligns the cross-domain foreground feature by adversarial training (AT), and it can filter the noisy background feature that is not suitable to transfer. Then, the PCFA adaptively aligns the confusing background and foreground feature, further promoting the domain adaptation performance. Comprehensive experiments validate the effectiveness of the proposed method. Compared with the baseline model trained on the semisynthetic source dataset, our FADA improves the generalized performance on the real-world target dataset a large-scale Dataset for Object deTection in Aerial images (DOTA) by 15.7% average precision (AP) and achieves state-of-the-art results. Tao Xu 0053, Xian Sun 0001, Wenhui Diao, Liangjin Zhao, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Geometrical Model for the Layover of Gable-Roofed Buildings and its Application in Building ReconstructionabstractBuilding reconstruction from SAR images is a hot topic in recent years. Currently, related methods mainly deal with on flat-roofed buildings. In this paper, we extend the research scope to gable-roofed buildings, and try to present a parameterized geometrical model for the layover of gable-roofed buildings. Based on this model, a top-down building reconstruction technique based on MCMC method is proposed. Through representing the layover with parameterized geometrical models, building reconstruction is converted into an optimization problem under the Bayesian scheme. In order to obtain global optima, simulated annealing algorithm with MCMC is used in the optimization stage. Two groups of transmission kernels which are responsible for model updates are designed according to the model. Experiments show that the layover model is accurate and the reconstruction method is effective. Yue Zhang 0016, Zhirui Wang 0003, Liangjin Zhao, Wenkai Zhang 0002, Menglong Yan, Xian Sun 0001 |
IGARSS | 3 |