EDBT 2026 Demo / reviewers in the wild / expert
Rongjun Qin
dblp:196/8678
· DBLP profile ↗
30ranked-venue papers
6as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Geometry-Aware Feature Matching for Large-Scale Structure from MotionabstractEstablishing consistent and dense correspondences across multiple images is crucial for Structure from Motion (SfM) systems. Significant view changes, such as air-to-ground with very sparse view overlap, pose an even greater challenge to the correspondence solvers. We present a novel optimization-based approach that significantly enhances existing feature matching methods by introducing geometry cues in addition to color cues. This helps fill gaps when there is less overlap in large-scale scenarios. Our method formulates geometric verification as an optimization problem, guiding feature matching within detector-free methods and using sparse correspondences from detector-based methods as anchor points. By enforcing geometric constraints via the Sampson Distance, our approach ensures that the denser correspondences from detector-free methods are geometrically consistent and more accurate. This hybrid strategy significantly improves correspondence density and accuracy, mitigates multi-view inconsistencies, and leads to notable advancements in camera pose accuracy and point cloud density. It outperforms state-of-the-art feature matching methods on benchmark datasets and enables feature matching in challenging extreme large-scale settings. Project page: https://xtcpete.github.io/geo-website/. Gonglin Chen, Jinsen Wu, Haiwei Chen, Wenbin Teng, Andrew Feng, Rongjun Qin |
3DV | 7 |
| 2025 | Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsabstractGenerating consistent ground-view images from satellite imagery is challenging, primarily due to the large discrepancies in viewing angles and resolution between satellite and ground-level domains. Previous efforts mainly concentrated on single-view generation, often resulting in inconsistencies across neighboring ground views. In this work, we propose a novel cross-view synthesis approach designed to overcome these challenges by ensuring consistency across ground-view images generated from satellite views. Our method, based on a fixed latent diffusion model, introduces two conditioning modules: satellite-guided de-noising, which extracts high-level scene layout to guide the denoising process, and satellite-temporal denoising, which captures camera motion to maintain consistency across multiple generated views. We further contribute a large-scale satellite-ground dataset containing over 100,000 perspective pairs to facilitate extensive ground scene or video generation. Experimental results demonstrate that our approach outperforms existing methods on perceptual and temporal metrics, achieving high photorealism and consistency in multi-view outputs. The project page is at https://gdaosu.github.io/sat2groundscape. Ningli Xu, Rongjun Qin |
CVPR | 2 |
| 2025 | OmniMesh: Addressing Findability Challenges in Distributed Nature Data Repositories
Arnab Nandi 0001, Wei-Lun Chao, Rongjun Qin, Carl Boettiger, Hilmar Lapp, Tanya Y. Berger-Wolf |
SSDBM | 3 |
| 2025 | Skyeyes: Ground Roaming using Aerial View ImagesabstractIntegrating aerial imagery-based scene generation into applications like autonomous driving and gaming enhances realism in 3D environments, but challenges remain in creating detailed content for occluded areas and ensuring real-time, consistent rendering. In this paper, we introduce Skyeyes, a novel framework that can generate photoreal-istic sequences of ground view images using only aerial view inputs, thereby creating a ground roaming experience. More specifically, we combine a 3D representation with a view consistent generation model, which ensures coher-ence between generated images. This method allows for the creation of geometrically consistent ground view images, even with large view gaps. The images maintain improved spatial-temporal coherence and realism, enhancing scene comprehension and visualization from aerial perspectives. To the best of our knowledge, there are no publicly avail-able datasets that contain pairwise geo-aligned aerial and ground view imagery. Therefore, we build a large, synthetic, and geo-aligned dataset using Unreal Engine. Both qualitative and quantitative analyses on this synthetic dataset display superior results compared to other leading syn-thesis approaches. See the project page for more results: chaoren2357.github.io/website-skyeyes/. Wenbin Teng, Gonglin Chen, Jinsen Wu, Ningli Xu, Rongjun Qin, Andrew Feng |
WACV | 6 |
| 2025 | Synthetic Data Matters: Retraining With Geo-Typical Synthetic Labels for Building DetectionabstractDeep learning has significantly advanced building segmentation in remote sensing, yet models struggle to generalize on data of diverse geographic regions due to variations in city layouts and the distribution of building types, sizes and locations. However, the amount of time-consuming annotated data for capturing worldwide diversity may never catch up with the demands of increasingly data-hungry models. Thus, we propose a novel approach: re-training models at test time using synthetic data tailored to the target region’s city layout. This method generates geo-typical synthetic data that closely replicates the urban structure of a target area by leveraging geospatial data such as street network from OpenStreetMap. Using procedural modeling and physics-based rendering, very high-resolution synthetic images are created, incorporating domain randomization in building shapes, materials, and environmental illumination. This enables the generation of virtually unlimited training samples that maintain the essential characteristics of the target environment. To overcome synthetic-to-real domain gaps, our approach integrates geo-typical data into an adversarial domain adaptation framework for building segmentation. Experiments demonstrate significant performance enhancements, with median improvements of up to 12%, depending on the domain gap. This scalable and cost-effective method blends partial geographic knowledge with synthetic imagery, providing a promising solution to the “model collapse” issue in purely synthetic datasets. It offers a practical pathway to improving generalization in remote sensing building segmentation without extensive real-world annotations. https://github.com/GDAOSU/geotypical_synthetic_label_building_detection. Shuang Song 0010, Yang Tang 0002, Rongjun Qin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Geospecific View Generation Geometry-Context Aware High-Resolution Ground View Inference from Satellite Views
Ningli Xu, Rongjun Qin |
ECCV (47) | 2 |
| 2024 | Automated Deep Learning-Based Point Cloud Classification on USGS 3DEP Lidar Data Using A TransformerabstractThe goal of the U.S. Geological Survey’s (USGS) 3D Elevation Program (3DEP) is to facilitate the acquisition of nationwide lidar data. Although data meet USGS lidar specifications, some point cloud tiles include noisy and incorrectly classified points. The enhanced accuracy of classified point clouds can improve support for many downstream applications such as hydrologic analysis, urban planning, and forest management. Despite noisy and incorrectly classified points, the current 3DEP classification specifications result in data that can be useful for Digital Terrain Model (DTM) extraction; however, the quality of the classification application can be improved to match state-of-the-art capabilities. Deep Learning (DL)-based approaches have been developed with outstanding performance for point cloud classification. This study will utilize the proven DL technologies to prepare for developing a user-friendly open-source toolkit that would automate classification to refine and enrich the results of existing and future 3DEP data. Jung-Kuan Liu, Rongjun Qin, Shuang Song 0010 |
IGARSS | 2 |
| 2024 | Multi-agent policy transfer via task relationship modeling
Rongjun Qin, Feng Chen 0042, Tonghan Wang 0001, Lei Yuan 0005, Xiaoran Wu, Yipeng Kang, Zongzhang Zhang, Chongjie Zhang, Yang Yu 0001 |
Sci. China Inf. Sci. | 1 |
| 2024 | Learning in games: a systematic review
Rongjun Qin |
Sci. China Inf. Sci. | 1 |
| 2024 | Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender SystemsabstractRecommender systems are expected to be assistants that help human users find relevant information automatically without explicit queries. As recommender systems evolve, increasingly sophisticated learning techniques are applied and have achieved better performance in terms of user engagement metrics such as clicks and browsing time. The increase in the measured performance, however, can have two possible attributions: a better understanding of user preferences, and a more proactive ability to utilize human bounded rationality to seduce user over-consumption. A natural following question is whether current recommendation algorithms are manipulating user preferences. If so, can we measure the manipulation level? In this article, we present a general framework for benchmarking the degree of manipulations of recommendation algorithms, in both slate recommendation and sequential recommendation scenarios. The framework consists of four stages, initial preference calculation, training data collection, algorithm training and interaction, and metrics calculation that involves two proposed metrics, Manipulation Score and Preference Shift. We benchmark some representative recommendation algorithms in both synthetic and real-world datasets under the proposed framework. We have observed that a high online click-through rate does not necessarily mean a better understanding of user initial preference, but ends in prompting users to choose more documents they initially did not favor. Moreover, we find that the training data have notable impacts on the manipulation degrees, and algorithms with more powerful modeling abilities are more sensitive to such impacts. The experiments also verified the usefulness of the proposed metrics for measuring the degree of manipulations. We advocate that future recommendation algorithm studies should be treated as an optimization problem with constrained user preference manipulations. Zhengbang Zhu, Rongjun Qin, Xinyi Dai, Yang Yu 0001, Yong Yu 0001, Weinan Zhang 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Mesh Conflation of Oblique Photogrammetric Models Using Virtual Cameras and Truncated Signed Distance FieldabstractConflating/stitching 2.5-D raster digital surface models (DSMs) into a large one has been a running practice in geoscience applications; however, conflating full-3-D mesh models, such as those from oblique photogrammetry, is extremely challenging. In this letter, we propose a novel approach to address this challenge by conflating multiple full-3-D oblique photogrammetric models into a single and seamless mesh for high-resolution site modeling. Given two or more individually collected and created photogrammetric meshes, we first propose to create a virtual camera field [with a panoramic field of view (FoV)] to incubate virtual spaces represented by truncated signed distance field (TSDF), an implicit volumetric field friendly for linear 3-D fusion; then, we adaptively leverage the truncated bound of meshes in TSDF to conflate them into a single and accurate full 3-D site model. With drone-based 3-D meshes, we show that our approach significantly improves upon traditional methods for model conflations, to drive new potentials to create excessively large and accurate full 3-D mesh models in support of geoscience and environmental applications. Shuang Song 0010, Rongjun Qin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Select-and-Combine (SAC): A Novel Multi-stereo Depth Fusion Algorithm for Point Cloud Generation via Efficient Local Markov NetletsabstractMany practical systems for image-based surface reconstruction employ a stereo/multi-stereo paradigm, due to its ability to scale for large scenes and its ease of implementation for out-of-core operations. In this process, multiple and abundant depth maps from stereo matching must be combined and fused into a single, consistent, and clean point cloud. However, the noises and outliers caused by stereo matching and the heterogenous geometric errors of the poses present a challenge for existing fusion algorithms, since they mostly assume Gaussian errors and predict fused results based on data from local spatial neighborhoods, which may inherit uncertainties from multiple depths resulting in lowered accuracy. In this paper, we propose a novel depth fusion paradigm, that instead of numerically fusing points from multiple depth maps, selects the best depth map per point, and combines them into a single and clean point cloud. This paradigm, called select-and-combine (SAC), is achieved through modeling the point level fusion using local Markov Netlets, a micro-network over point across neighboring views for depth/view selection, followed by a Netlets collapse process for point combination. The Markov Netlets are optimized such that they can inherently leverage spatial consistencies among depth maps of neighboring views, thus they can address errors beyond Gaussian ones. Our experiment results show that our approach outperforms existing depth fusion approaches by increasing the F1 score that considers both accuracy and completeness by 2.07% compared to the best existing method. Finally, our approach generates clearer point clouds that are 18% less redundant while with a higher accuracy before fusion. Mostafa M. El-Hashash, Rongjun Qin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Large-Scale and Efficient Texture Mapping Algorithm via Loopy Belief PropagationabstractTexture mapping as a fundamental task in 3D modeling has been well established for well-acquired aerial assets under consistent illumination, yet it remains a challenge when it is scaled to large datasets with images under varying views and illuminations. A well-performed texture mapping algorithm must be able to efficiently select views, fuse and map textures from these views to mesh models, at the same time, achieve consistent radiometry over the entire model. Existing approaches achieve efficiency either by limiting the number of images to one view per face, or simplifying global inferences to only achieve local color consistency. In this paper, we break this tie by proposing a novel and efficient texture mapping framework that allows the use of multiple views of texture per face, at the same time to achieve global color consistency. The proposed method leverages a loopy belief propagation algorithm to perform an efficient and global-level probabilistic inferences to rank candidate views per face, which enables face-level multi-view texture fusion and blending. The texture fusion algorithm, being non-parametric, brings another advantage over typical parametric post color correction methods, due to its improved robustness to non-linear illumination differences. The experiments on three different types of datasets (i.e. satellite dataset, unmanned-aerial vehicle dataset and close-range dataset) show that the proposed method has produced visually pleasant and texturally consistent results in all scenarios, with an added advantage of consuming less running time as compared to the state of the art methods, especially for large-scale dataset such as satellite-derived models. Rongjun Qin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | NeoRL: A Near Real-World Benchmark for Offline Reinforcement LearningabstractOffline reinforcement learning (RL) aims at learning effective policies from historical data without extra environment interactions. During our experience of applying offline RL, we noticed that previous offline RL benchmarks commonly involve significant reality gaps, which we have identified include rich and overly exploratory datasets, degraded baseline, and missing policy validation. In many real-world situations, to ensure system safety, running an overly exploratory policy to collect various data is prohibited, thus only a narrow data distribution is available. The resulting policy is regarded as effective if it is better than the working behavior policy; the policy model can be deployed only if it has been well validated, rather than accomplished the training. In this paper, we present a Near real-world offline RL benchmark, named NeoRL, to reflect these properties. NeoRL datasets are collected with a more conservative strategy. Moreover, NeoRL contains the offline training and offline validation pipeline before the online test, corresponding to real-world situations. We then evaluate recent state-of-the-art offline RL algorithms in NeoRL. The empirical results demonstrate that some offline RL algorithms are less competitive to the behavior cloning and the deterministic behavior policy, implying that they could be less effective in real-world tasks than in the previous benchmarks. We also disclose that current offline policy evaluation methods could hardly select the best policy. We hope this work will shed some light on future research and deploying RL in real-world systems. Rongjun Qin, Xingyuan Zhang, Songyi Gao, Xiong-Hui Chen, Weinan Zhang 0001, Yang Yu 0001 |
NeurIPS | 1 |
| 2022 | Concurrent events risk assessment generic models with enhanced reliability using Fault tree analysis and expanded rotational fuzzy sets
Nabeel Mahmood, Tarunjit Singh Butalia, Rongjun Qin, Maram Manasrah |
Expert Syst. Appl. | 3 |
| 2022 | Bispace Domain Adaptation Network for Remotely Sensed Semantic SegmentationabstractSupervised learning for semantic segmentation has achieved impressive success in remote sensing, while this normally has a high demand on pixel-level ground truth from the testing images (target domain). Labeling data for semantic segmentation is labor-intensive and time-consuming. To reduce the workload of manual labeling, domain adaptation (DA) utilizes preexisting labeled images from other sources (source domain) to classify the images in the target domain. In this article, we propose a bispace alignment network for DA named BSANet. BSANet is designed to have a dual-branch structure which is able to extract features in the image domain and the wavelet domain simultaneously. To minimize the discrepancy between the source and target domains, we propose a bispace adversarial learning strategy. Specifically, BSANet employs two discriminators in different spaces, one aligning the source and target feature distributions, and the other helping the classification outputs render reasonable spatial layouts. The proposed method shows the ability to train an end-to-end network for semantic segmentation without using any label in the target domain. Extensive experiments and ablation studies are conducted in cross-city scenarios. Comparative experiments with several state-of-the-art DA methods show that our method achieves the best performance. Wei Liu 0076, Fulin Su, Xinfei Jin, Rongjun Qin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Sat2Vid: Street-view Panoramic Video Synthesis from a Single Satellite ImageabstractWe present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attention. For geometrical and temporal consistency, our approach explicitly creates a 3D point cloud representation of the scene and maintains dense 3D-2D correspondences across frames that reflect the geometric scene configuration inferred from the satellite view. As for synthesis in the 3D space, we implement a cascaded network architecture with two hourglass modules to generate point-wise coarse and fine features from semantics and per-class latent vectors, followed by projection to frames and an up-sampling module to obtain the final realistic video. By leveraging computed correspondences, the produced street-view video frames adhere to the 3D geometric scene structure and maintain temporal consistency. Qualitative and quantitative experiments demonstrate superior results compared to other state-of-the-art synthesis approaches that either lack temporal consistency or realistic appearance. To the best of our knowledge, our work is the first one to synthesize cross-view images to videos.. Zuoyue Li, Zhenqiang Li 0002, Zhaopeng Cui, Rongjun Qin, Marc Pollefeys, Martin R. Oswald |
ICCV | 4 |
| 2021 | Vis2Mesh: Efficient Mesh Reconstruction from Unstructured Point Clouds of Large Scenes with Learned Virtual View VisibilityabstractWe present a novel framework for mesh reconstruction from unstructured point clouds by taking advantage of the learned visibility of the 3D points in the virtual views and traditional graph-cut based mesh generation. Specifically, we first propose a three-step network that explicitly employs depth completion for visibility prediction. Then the visibility information of multiple views is aggregated to generate a 3D mesh model by solving an optimization problem considering visibility in which a novel adaptive visibility weighting in surface determination is also introduced to suppress line of sight with a large incident angle. Compared to other learning-based approaches, our pipeline only exercises the learning on a 2D binary classification task, i.e., points visible or not in a view, which is much more generalizable and practically more efficient and capable to deal with a large number of points. Experiments demonstrate that our method with favorable transferability and robustness, and achieve competing performances w.r.t. state-of-the-art learning-based approaches on small complex objects and outperforms on large indoor and outdoor scenes. Code is available at https://github.com/GDAOSU/vis2mesh. Shuang Song 0010, Zhaopeng Cui, Rongjun Qin |
ICCV | 3 |
| 2020 | Geometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasabstractWe present a novel method for generating panoramic street-view images which are geometrically consistent with a given satellite image. Different from existing approaches that completely rely on a deep learning architecture to generalize cross-view image distributions, our approach explicitly loops in the geometric configuration of the ground objects based on the satellite views, such that the produced ground view synthesis preserves the geometric shape and the semantics of the scene. In particular, we propose a neural network with a geo-transformation layer that turns predicted ground-height values from the satellite view to a ground view while retaining the physical satellite-to-ground relation. Our results show that the synthesized image retains well-articulated and authentic geometric shapes, as well as texture richness of the street-view in various scenarios. Both qualitative and quantitative results demonstrate that our method compares favorably to other state-of-the-art approaches that lack geometric consistency. Xiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R. Oswald, Marc Pollefeys, Rongjun Qin |
CVPR | 6 |
| 2020 | Large-Scale Land Cover Mapping of Satellite Images Using Ensemble of Random Forests - IEEE Data Fusion Contest 2020 Track 1abstractThis paper describes our approach in the land cover classification with low-resolution labels challenge of 2020 IEEE Data Fusion Contest. The challenge features the large-scale land cover mapping based on weakly annotated samples. We firstly refine the samples based on analyzing the class confusions and the confidence of the sample. Subsequently, an ensemble of random forests is trained for classification and finally the results are generated via soft voting of the classifiers. This approach achieves an average accuracy of 0.5676 on the test dataset and achieves the 4thplace in Track 1 of the Data Fusion Contest 2020. Huijun Chen, Changlin Xiao, Rongjun Qin |
IGARSS | 4 |
| 2020 | Large-Scale Land Cover Mapping of Satellite Images Using Ensemble of Random Forests with Multi-Resolution Label - IEEE Data Fusion Contest 2020 Track 2abstractThis paper describes our approach in the land cover classification with low- and high-resolution labels challenge of 2020 IEEE Data Fusion Contest. The challenge features the large-scale land cover mapping based on weakly annotated samples. We firstly refine the samples based on prior knowledge on class confusion and the confidence of the samples. Subsequently, an ensemble of random forests is used for classification and a post-processing step is performed to improve the results. This approach achieves an average accuracy of 0.6142 on the test dataset and achieves the 1stplace in Track 2 of the Data Fusion Contest 2020. Huijun Chen, Changlin Xiao, Rongjun Qin |
IGARSS | 4 |
| 2020 | Effects of Unbalanced Data on Radiometric Transforming Model Fitting for Relative Radiometric NormalizationabstractMost of the existing (Relative Radiometric Normalization) RRN research focus on the automatic sample selection, while unbalanced data effect on the regression is not noted and the pre-selected sample set like the Pseudo-invariant Features (PIFs), or homogeneous pixels are directly used to solve the radiometric transforming models without considering the representativeness of the sample set. In this paper, we investigated the effects of unbalanced data, specifically on two aspects: 1) statistical properties of the estimated model parameters, 2) the normalizing accuracy of the fitted model. To make the work thoroughly, four regression methods are investigated, including Least Square Regression (LSQ), Theil-Sen estimator (TSR), Support vector machine regression (SVR), and Random forest regression (RFR). And Monte-Carlo Simulation is used to generated various sample sets with different distributions. It is demonstrated that the LSQ and the TSR are vulnerable to data unbalance, in terms of both the estimated model parameters and the normalizing accuracy of the fitted radiometric transforming model, whereas the SVR and RFR are not sensitive. Wenxia Gan, Jinying Xu, Weihang Yu, Huanning Yuan, Rongjun Qin |
IGARSS | 7 |
| 2020 | A MultiKernel Domain Adaptation Method for Unsupervised Transfer Learning on Cross-Source and Cross-Region Remote Sensing Data ClassificationabstractLabeling remote sensing data for classification is labor-intensive and time-consuming. Transfer learning (TL), under such context, is attracting increasing attention as it aims to harness information from data set of other regions where labels are readily available. The central topic of concern is to homogenize the large disparities of feature distribution of different data set through domain adaptation (DA). This article proposes a novel DA method for unsupervised TL, namely, multikernel jointly domain matching (MKJDM), which by definition considers multiple kernels as opposed to the currently popular single-kernel methods for measuring the distances between distributions. The single-kernel methods minimize the distances of feature distribution between the source domain (data set with training labels) and the target domain (data set to be classified) through, for example, maximum mean discrepancy (MMD) metric, formed under a kernel function mapping, while the multikernel version (MK-MMD) uses different kernel functions to encapsulate multiple aspects of distribution discrepancies, and is, therefore, more capable of distance minimization. Our MKJDM implementation also considers simultaneously aligning marginal and class conditional distributions and reweight for each instance, which further improves the performance. Two experiments performed on remote sensing images and multi-modal data sets (i.e., Orthophoto and Digital Surface Models), with regions of different countries with distinctly different land patterns serving as source and target domain data, show that the overall accuracies are improved by 37.28% and 46.62% after applications of our MKJDM method. An additional comparative experiment with five state-of-the-art DA methods also demonstrates that our method achieves the best performance. Rongjun Qin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Unsupervised Transfer Learning Using for Multi-Model Remote Sensing Data ClassificationabstractTraditional supervised classification has been successfully applied to remote sensing (RS) classification and thematic mapping. However, labeling data is usually time-consuming and expensive. Transfer learning (TL) is able to reduce the labeling labor by utilizing pre-existing labeled data (source domain). Most of the existing TL methods applied in remote sensing still need a few labeled samples in target domain (where to be classified) or interaction between the classifier and users. By transferring the data distributions between the source and target domain into a joint feature space (where the distribution discrepancy between the source and target domains is smaller than the original space), a novel unsupervised transfer learning framework is proposed in this paper. In our framework, robust remote sensing indices are extracted from Digital Surface Models (DSM) and the associated orthophoto to help us to estimate the conditional probability distribution of the target domain. After that, our proposed Dispersion-optimized Joint Distribution Alignment (DJDA) method is able to generate transformed feature representations of source and target feature sets through simultaneous reduction of the joint distribution discrepancy between two domains and optimizing the dispersion (both intra-class dispersion and inter-class dispersion). In our experiments, the proposed framework is able to improve the transferability between multi-model datasets (Orthophoto + DSM). The comparative experiments demonstrate that the proposed DJDA outperforms the compared state-of-the-art unsupervised domain adaptation methods on RS image classification. Wei Liu 0076, Rongjun Qin |
IGARSS | 2 |
| 2019 | Pairwise Stereo Image Disparity and Semantics Estimation with the Combination of U-Net and Pyramid Stereo Matching NetworkabstractStereo images are one of the most common resources for 3D reconstruction. In the pairwise semantic stereo challenge of the 2019 IEEE GRSS Data Fusion Contest, we generate the classification and the disparity map based on U-Net and Pyramid Stereo Matching Network (PSMNet), respectively. By using a dynamic and class-weighted loss function, the UNet is effectively trained with the imbalanced training samples. By voting classification results of the augmented prediction data with models trained under different epochs, we further refine the classification maps with the constraints of pseudo DSM, water index and mean-shift segmentation. Rongjun Qin, Xu Huang 0005, Changlin Xiao |
IGARSS | 1 |
| 2019 | Semantic 3D Reconstruction Using Multi-View High-Resolution Satellite Images Based on U-Net and Image-Guided Depth FusionabstractThis paper reports our workflow for multi-view semantic stereo reconstruction in the IEEE Data Fusion Contest (DFC) of 2019. The workflow mainly consists of two parts: semantic classification and multi-view stereo digital surface model (DSM) generation. Our strategy involves the interaction of these two components for better classification and DSM results. We firstly classify semantic objects using only original image information and then project them to the generated DSM for voting and post-refinement. An example-based pair selection strategy is used to pre-selected pairs for dense matching, and the fused DSM is performed through an image-guided median filter. Since the adjusted rational polynomial coefficient (RPC) parameters are not accurate in the dataset, these DSMs are co-registered after production. In our workflow, we only consider non-vegetation and non-water area for registration to obtain more robust co-registration result. These co-registered DSMs are fused and refined using the result from the semantic classification. Our method achieved 73.00% of mIoU-3 in the testing data in the contest. Rongjun Qin, Xu Huang 0005, Changlin Xiao |
IGARSS | 1 |
| 2019 | An Operational Pipeline for Generating Digital Surface Models from Multi-Stereo Satellite Images for Remote Sensing ApplicationsabstractMulti-view satellite images are particularly useful in providing large-scale geometric information and can benefit a wide range of remote sensing applications. This paper presents an operational and fully automated pipeline that performs stereo and multi-stereo reconstruction using RPC stereo processor (RSP) software. We demonstrate that the processed information such as per-pixel registered orthophotos and digital surface models may serve as attractive sources for remote sensing applications such as 3D building and canopy modeling, land-cover classification that might be traditionally limited in 2D remote sensing. Rongjun Qin |
IGARSS | 1 |
| 2019 | Urban Land-Cover Classification with Façade Feature From Oblique ImagesabstractIn the remote sensing community, land-cover classification is usually performed on the top-view images. However, besides the top-view features (including elevation), façade captured by the oblique images is useful but severely underutilized in the land-cover classification. The façade information of an object, is by nature more variable, thus can be extremely useful when the extracted features are used for land-cover classification. Hence, in this paper, we try to explore the use of façade from the oblique images to enhance the accuracy of land-cover classification. Firstly, we locate the façades by finding the elevation changes and the corresponding above-ground objects. Then, the façade images are cropped from oblique images and the color and Haar-like features are extracted as façade features. Finally, following the object-based land-cover classification, super-pixels are generated and used as the basic unit for the feature extraction and classification. Experiments are performed on five representative site using five-head oblique aerial images and their derived orthophoto and digital surface model (DSM). The results show that with the façade information, the classification performances have been steadily improved, especially for the buildings which has around 10% improvement. Changlin Xiao, Rongjun Qin, Hanning Yuan |
IGARSS | 2 |
| 2018 | Time-Series 3D Building Change Detection Based on Belief FunctionsabstractOne of the challenges of remote sensing image based building change detection is distinguishing building changes from other types of land cover alterations. Height information can be a great assistance for this task but its performance is limited to the quality of the height. Yet, the standard automatic methods for this task are still lacking. We propose a very high resolution stereo series data based building change detection approach that focuses on the use of time series information. In the first step, belief functions are explored to fuse the change features from the 2D and height maps to obtain an initial change detection result. In the second step, the building probability maps (BPMs) from the series data are adopted to refine the change detection results based on Dempster-Shafer theory. The final step is to fuse the series building change detection results in order to obtain a final change map. The advantages of the proposed approach are demonstrated by testing it on a set of time series data captured in North Korea. Jiaojiao Tian, Jean Dezert, Rongjun Qin |
FUSION | 3 |
| 2018 | Individual Tree Detection from Multi-View Satellite ImagesabstractIndividual tree detection is critical in forest monitoring and inventory. In this paper, we propose a novel method to use multi-view satellite images to detect individual trees and delineate their crowns. As compared to previous methods that only use image information, we generate the DSM from the multi-view high-resolution satellite images and combine it with the spectral information to detect the trees. Firstly, the vegetation areas are extracted to remove the non-vegetation objects while terrain areas are extracted to help estimate the tree height. Then, we utilize top-hat morphological operation to efficiently find the local maximal points as treetops and further refine them by checking their heights and doing non-maximum suppression. Finally, we use a revised superpixel segmentation algorithm to delineate the tree crowns which considered both 2D spectral and 3D structure similarities. To effectively assess the performance, we rigorously match and evaluate the detected and reference trees in a one-to-one relationship. A quantitative evaluation at three different sites shows that the proposed method is able to detect individual trees at different regions with high accuracy. Changlin Xiao, Rongjun Qin, Xu Huang 0005, Jiaqiang Li |
IGARSS | 2 |