José Marcato Junior

dblp:248/2262 · also José Marcato Jr. · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
23since 2021 · last 2025
0000-0002-9096-6866ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 MTCloud: Multi-type convolutional linkage network for point cloud instance segmentation
Jing Du 0007, Guo-Rong Cai, Zongyue Wang, Jinhe Su, Min Huang 0004, John S. Zelek, José Marcato Junior, Jonathan Li 0001
Expert Syst. Appl.7
2025 CityInsight: Incorporating Dual-Condition-Based Diffusion Model Into Building Footprint Segmentation From Remote Sensing Imagery
abstract
Accurately identifying urban building layouts plays a crucial role in understanding the complexity of urban construction and the level of economic development. Previous footprint segmentation methods have struggled to adapt to the diverse morphology of buildings, limiting the accurate extraction of building footprints from remote sensing imagery and impeding insights and understanding of urban areas. To this end, we propose a framework named CityInsight for analyzing urban building morphology from remote sensing imagery. First, we establish a semantic segmentation network, dual-condition diffusion network (DC-Net), based on a diffusion model to accurately identify building footprints from remote sensing images. Second, we use uncertainty attention and condition attention to generate spatial and semantic priors. Finally, we design a condition injection module to incorporate spatial and semantic information into the diffusion learning. Comprehensive experiments demonstrate the accuracy, robustness, and generalization of the proposed method. The$F_{1}$-scores of the DC-Net on the large-scale remote sensing datasets SpaceNet, WHU Building, Inria, and Massachusetts are 92.05%, 96.59%, 92.17%, and 92.86%, respectively. Furthermore, the footprint segmentation is utilized for subsequent urban function identification and urban analysis of the Zona Oeste of Rio de Janeiro, to emphasize the application value of CityInsight. Our code is available athttps://github.com/Ting-Devin-Han/CityInsight
Ting Han 0001, Chaolei Wang, Yang Luo 0002, Hongchao Fan, José Marcato Junior, Xinchang Zhang 0002, Yiping Chen 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 Prototypical Contrastive Network for Imbalanced Aerial Image Segmentation
abstract
Binary segmentation is the main task underpinning several remote sensing applications, which are particularly interested in identifying and monitoring a specific category/object. Although extremely important, such a task has several challenges, including huge intra-class variance for the background and data imbalance. Furthermore, most works tackling this task partially or completely ignore one or both of these challenges and their developments. In this paper, we propose a novel method to perform imbalanced binary segmentation of remote sensing images based on deep networks, prototypes, and contrastive loss. The proposed approach allows the model to focus on learning the foreground class while alleviating the class imbalance problem by allowing it to concentrate on the most difficult background examples. The results demonstrate that the proposed method outperforms state-of-the-art techniques for imbalanced binary segmentation of remote sensing images while taking much less training time.
Keiller Nogueira, Mayara Maezano Faita Pinheiro, Ana Paula Marques Ramos, Wesley Nunes Gonçalves, José Marcato Junior, Jefersson A. dos Santos
WACV5
2024 Crack-U2Net: Multiscale Feature Learning Network for Pavement Crack Detection From Large-Scale MLS Point Clouds
abstract
Deep learning-based algorithms detect pavement cracks in an end-to-end manner from Mobile Laser Scanning (MLS) point clouds, achieving impressive results. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding multiscale features and the limited training data. In this paper, we propose a novel pavement crack detection framework, Crack-U2Net, which innovatively incorporates a two-level nested U-Net architecture for feature learning. This design enables the learning of intra-stage multiscale features without introducing significant memory and computation costs, resulting in substantial improvements in accuracy. Moreover, to solve the challenge of insufficient training data, we propose a Geometry-based Data Augmentation (GDA) strategy, aiming to expand the pavement dataset while preserving the pavement geometry. Extensive experiments on the Qinghai-Tibet Highway point cloud dataset demonstrate the higher accuracy and efficiency of Crack-U2Net over the state-of-the-art methods, achieving an average precision, recall, F$1\text - $score, and accuracy of 83.8%, 77.6%, 80.1%, and 95.8%, respectively.
Huifang Feng 0002, Wen Li 0005, Lingfei Ma, Yiping Chen 0002, Haiyan Guan, Yongtao Yu, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.7
2023 Robust Image Captioning with Post-Generation Ensemble Method
abstract
Remote sensing image captioning is a research domain that aims to automatically generate natural language descriptions of the contents within remote sensed images. Providing accurate depictions of image contents holds great significance for downstream applications such as image retrieval and image understanding. While there is a pressing need for reliable results, current research predominantly focuses on single captioning algorithms, striving to enhance their performance on specific target-oriented datasets. Undoubtedly, this research trajectory is highly important. However, we believe that relying solely on the output of a single captioner may introduce a vulnerability from a robustness standpoint. This concern is particularly relevant in remote sensing, where the scarcity of large-scale datasets can limit the robustness and reliability of resulting algorithms. In this paper, we propose an approach that harnesses the advantages of ensembles to enhance accuracy and reliability in the context of image captioning. Our method introduces a novel technique for utilizing an ensemble of diverse captioning algorithms and automatically selecting the most suitable caption from the set of predictions. By decoupling the description generation and selection phases, this approach enables high flexibility of integration of architecturally different captioning algorithms in the pipeline.
Riccardo Ricci, Farid Melgani, José Marcato Junior, Wesley Nunes Gonçalves
IGARSS3
2023 A Click-Based Interactive Segmentation Network for Point Clouds
abstract
Interactive segmentation plays an essential role in several tasks involving point clouds. However, existing methods suffer from low segmentation accuracy and cannot adjust the segmentation results according to the user’s personal demands. This paper presents a novel deep learning-based interactive segmentation method, named Click Rough Segmentation Network (CRSNet), designed to handle point clouds. The method allows users to iteratively click to segment interesting objects. CRSNet consists of two key parts: a CRS module and a feature extraction module. First, the CRS module transforms click operations into an appropriate representation to input into the feature extraction module. The CRS module takes raw point clouds and click operations as input and outputs 3D Gaussian vectors and roughly segmented blocks, which adapt to different-sized and densely-distributed objects in complex environments. Second, the feature extraction module, which uses a novel mix loss-based analysis algorithm, extracts deep features and obtains instance segmentation results. The module is highly compatible because its backbones can be replaced by different deep learning architectures. Experimental results on the KITTI, Apolloscape, Roadmarking, Scannet, and SemanticKITTI datasets show that our method outperforms state-of-the-art semantic segmentation methods with one click. Moreover, our method can generalize well to unseen objects and datasets.
Wentao Sun, Yiping Chen 0002, Huxiong Li, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Weakly Supervised Few-Shot Segmentation via Meta-Learning
abstract
Semantic segmentation is a classic computer vision task with multiple applications, which includes medical and remote sensing image analysis. Despite recent advances with deep-based approaches, labeling samples (pixels) for training models is laborious and, in some cases, unfeasible. In this paper, we present two novel meta-learning methods, named WeaSeL and ProtoSeg, for the few-shot semantic segmentation task with sparse annotations. We conducted an extensive evaluation of the proposed methods in different applications (12 datasets) in medical imaging and agricultural remote sensing, which are very distinct fields of knowledge and usually subject to data scarcity. The results demonstrated the potential of our method, achieving suitable results for segmenting both coffee/orange crops and anatomical parts of the human body in comparison with full dense annotation.
Pedro H. T. Gama, Hugo N. Oliveira 0001, José Marcato Junior, Jefersson A. dos Santos
IEEE Trans. Multim.3
2022 Counting and locating high-density objects using convolutional neural network
Mauro dos Santos de Arruda, Lucas Prado Osco, Plabiany Rodrigo Acosta, Diogo Nunes Gonçalves, José Marcato Junior, Ana Paula Marques Ramos, Edson Takashi Matsubara, Jonathan Li 0001, Jonathan de Andrade Silva, Wesley Nunes Gonçalves
Expert Syst. Appl.5
2022 Building Instance Extraction Method Based on Improved Hybrid Task Cascade
abstract
Automatic building extraction from remote sensing imagery is crucial to urban construction and management. To address the main challenges of diverse building scale and appearance, this letter proposes an automatic building instance extraction method based on an improved hybrid task cascade (HTC). Our method consists of three components by obtaining high-resolution representation, defining guided anchor, and forming focal loss to boost the adaptability of automatic building instance extraction. Comprehensive experimental results on WHU aerial building data set demonstrated that compared with the mainstream Mask R-CNN method, our method increased AP and AR in bounding box branch and mask branch by 9.8%–6.5% and 10.7%–8.0% respectively, especially AP$_{S}$and AP$_{L}$in the two branches by 10.1%–6.9% and 3.4%–2.4%, respectively. We evaluated the effectiveness and complexity of these components separately and discussed the universality and practicability of deep learning method in automatic building extraction.
Yiping Chen 0002, Mingqiang Wei, Cheng Wang 0003, Wesley Nunes Gonçalves, José Marcato Junior, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 A Supervoxel Approach to Road Boundary Enhancement From 3-D LiDAR Point Clouds
abstract
Rapid and accurate enhancement of road boundaries from terrestrial laser scanning (TLS) 3-D point clouds has been a challenging task in road infrastructure inventory. To address the challenge with a lack of ability to enhance object boundaries when the supervoxel number is less, this letter proposes a novel supervoxel segmentation algorithm framework for enhancing road boundaries from 3-D point clouds. First, we utilize radius$k$nearest-neighbor search method to obtain the neighborhood information after partitioning points on octrees with seed points. Second, the iterative weighted least square algorithm and spatial structure judgment are used to segment point clouds based on seed points. Finally, an update method to adjust the supervoxel centroids is applied with surrounding information in the first part. To verify the excellent performance, we tested the proposed method on two publicly large-scale point clouds benchmarks—IQmulus and TerraMobilita (IQTM) and Semantic 3-D. The experimental results demonstrate that our approach achieved approximately 48.98% and 68.41% boundary recall higher than two existing classical methods in the street scene, and our running time is feasible and effective.
Zhengchuan Sha, Yiping Chen 0002, Yangbin Lin, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Intracity Temperature Estimation by Physics Informed Neural Network Using Modeled Forcing Meteorology and Multispectral Satellite Imagery
abstract
Estimating urban surface temperature at high resolution is crucial for effective urban planning for climate-driven risks. This high-resolution surface temperature over broader scales can usually be obtained via satellite remote sensing for historical period. However, it can be hard for future predictions. This paper presents a Physics Informed Hierarchical Perception (PIHP) network, a novel approach for accurate, high-resolution and generalizable urban surface temperature estimation. The key to our approach is leveraging the implied temperature-related physics information of the land surface structure from high-resolution multi-spectral satellite images, thus achieving precise estimation or prediction for high spatial resolution urban surface temperature. Specifically, a semantic category histogram is first designed to describe the land surface structures. Based on this, a hierarchical urban surface perception network is proposed to capture the complex relationship between the underlying land surface features, upper atmosphere conditions and the intracity temperature. The proposed PIHP-Net makes it possible to generate models that can generalize across different cities, thus to estimating or predicting high-resolution urban surface temperature when the satellite land surface temperature (LST) observation is not available. Experiments over various cities in different climate regions in China show, for the first time, errors less than 2 Kelvin (for most of the cases) at the high resolution (60-by-60 meters grids), thus making it possible to predict futureintracity temperaturefrom forcing meteorology and multi-spectral satellite imagery.
Donghang Wu, Weiquan Liu, Lei Zhao 0023, Shenlong Wang, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.9
2022 GCN-Based Pavement Crack Detection Using Mobile LiDAR Point Clouds
abstract
Mobile Laser Scanning (MLS) system can provide high-density and accurate 3D point clouds that enable rapid pavement crack detection for road maintenance tasks. Supervised learning-based algorithms have been proved pretty effective for handling such a large amount of inhomogeneous and unstructured point clouds. However, these algorithms often rely on a lot of annotated data, which is labor-intensive and time-consuming. This paper presents a semi-supervised point-level approach to overcome this challenge. We propose a graph-widen module to construct a reasonable graph structure for point clouds, increasing the detection performance of graph convolutional networks (GCN). The constructed graph characterizes the local features from a small amount of annotated data, avoiding information loss and dramatically reduces the dependence on annotated data. The MLS point clouds acquired by a commercial RIEGL VMX-450 system are used in this study. The experimental results demonstrate that our method outperforms the state-of-the-art point-level methods in terms of recall, F1 score, and efficiency while achieving comparable accuracy.
Huifang Feng 0002, Wen Li 0005, Yiping Chen 0002, Sarah Narges Fatholahi, Ming Cheng 0002, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.8
2022 Rapid Extraction of Urban Road Guardrails From Mobile LiDAR Point Clouds
abstract
Mobile Laser Scanning (MLS) systems provide highly dense 3D point clouds that enable the acquisition of accurate traffic facilities information for intelligent transportation system. Road guardrails with safety features that can separate traffic and define moving spaces for pedestrians and vehicles face challenges such as diverse guardrail types and continuous slopes in point clouds data. This paper proposes a novel approach for rapidly extracting urban road guardrails from MLS point clouds, combining a proposed multi-level filtering with a modified Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering, and adapting for most types of guardrails and rough slope roads. We develop a multi-level filter to detect the road surface and remove the undesirable points. Through a proposed modified DBSCAN clustering, the guardrails are extracted after a four-step screening, which includes the limits based on the number of points, the fitting error, the bounding box size and the average reflection intensity for each cluster. The proposed method achieves high precisions of 97.2% and 96.4% respectively for the lane-separating guardrails and the anti-fall guardrails on the dataset. Extensive experiments with test dataset captured by a RIEGL VMX-450 MLS, show that our method outperforms the state-of-the-art method to extract 3D guardrails from point clouds.
Jianlan Gao, Yiping Chen 0002, José Marcato Junior, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 BoundaryNet: Extraction and Completion of Road Boundaries With Deep Learning Using Mobile Laser Scanning Point Clouds and Satellite Imagery
abstract
Robust road boundary extraction and completion play an important role in providing guidance to all road users and supporting high-definition (HD) maps. The significant challenges remain in remarkable and accurate road boundary recovery from poor road boundary conditions. This paper presents a novel deep learning framework, named BoundaryNet, to extract and complete road boundaries by using both mobile laser scanning (MLS) point clouds and high-resolution satellite imagery. First, road boundaries are extracted by conducting a curb-based extraction method. Such extracted 3D road boundary lines are used as inputs to feed into a U-shaped network for erroneous boundary denoising. Then, a convolutional neural network (CNN) model is proposed to complete the road boundaries. Next, to achieve more complete and accurate road boundaries, a conditional deep convolutional generative adversarial network (c-DCGAN) with the assistance of road centerlines extracted from satellite images is developed. Finally, according to the completed road boundaries, the inherent road geometries are calculated. The proposed methods were evaluated using satellite imagery and four MLS point cloud datasets with varying densities and road conditions in urban environments. The quality evaluation metrics of 82.88%, 82.43%, 88.86%, and 84.89% were achieved for four data sets. The experimental results indicate that the BoundaryNet model can provide a promising solution for road boundary completion and road geometry estimation.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.4
2022 Robust Lane Extraction From MLS Point Clouds Towards HD Maps Especially in Curve Road
abstract
This article presents a semi-automated method to extract the lane features along the curved roads from mobile laser scanning (MLS) point clouds. The proposed method consists of four steps. After data pre-processing, a road edge detection algorithm is performed to distinguish road curbs and extract road surfaces. Then, textual and directional road markings such as arrows, symbols, and words, to inform drivers in necessary cases, are detected by intensity thresholding and conditional Euclidean clustering algorithms. Furthermore, lane markings are extracted by local intensity analysis and distance thresholding methods according to road design standards, because they are more regular along the road. Finally, centerline points on lanes are estimated based on the coordinates of extracted lane markings. Our method shows strong feasibility and robustness when creating high-definition (HD) maps from MLS data, by increasing the number of blocks in the curve and the distance threshold control in curved lane centerline extraction. Quantitative evaluations show that the average recall, precision, and F1-score obtained from four datasets for road marking extraction are 93.87%, 93.76%, and 93.73%, respectively. The generated lane centerlines are evaluated by overlaying them on manually labeled reference buffers from 4 cm resolution orthoimagery. The comparative study indicates that the proposed methods can achieve higher accuracy and robustness than most state-of-the-art methods.
Chengming Ye, He Zhao 0007, Lingfei Ma, Han Jiang 0005, Hongfu Li, Ruisheng Wang 0001, Michael A. Chapman, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.8
2022 3D Vehicle Detection Using Multi-Level Fusion From Point Clouds and Images
abstract
3D vehicle detectors based on point clouds generally have higher detection performance than detectors based on multi-sensors. However, with the lack of texture information, point-based methods get many missing detection of occluded and distant vehicles, and false detection with high-confidence of similarly shaped objects, which is a potential threat to traffic safety. Therefore, in the long run, fusion-based methods have more potential. This paper presents a multi-level fusion network for 3D vehicle detection from point clouds and images. The fusion network includes three stages: data-level fusion of point clouds and images, feature-level fusion of voxel and Bird’s Eye View (BEV) in the point cloud branch, and feature-level fusion of point clouds and images. Besides, a novel coarse-fine detection header is proposed, which simulates the two-stage detectors, generating coarse proposals on the encoder, and refining them on the decoder. Extensive experiments show that the proposed network has better detection performance on occluded and distant vehicles, and reduces the false detection of similarly shaped objects, proving its superiority over some state-of-the-art detectors on the challenging KITTI benchmark. Ablation studies have also demonstrated the effectiveness of each designed module.
Lingfei Ma, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.6
2021 Evaluating Different Deep Learning Models for Automatic Water Segmentation
abstract
Deep learning (DL) methods, integrated with imagery obtained by remote sensors, are considered a novel source of information in the field of hydrometry. Their results can be a support for standard gauging systems. In this research, three different model definitions based on the SegNet architecture were developed: training a new model, using a pre-trained model, and performing transfer learning. The main goal of this approach is the performance assessment of convolutional neural network (CNN) generalization to segment images containing different water bodies automatically. Diverse sensors were used to obtain RGB images from different areas of the world. The effectiveness of the CNN was estimated using pixel accuracy and IoU metrics. Training a new model and using transfer learning revealed similar high accuracies that were at least twice as accurate compared to the pre-trained model. However, the transfer learned model is preferred due to significantly lower training expenses.
Thales Akiyama, José Marcato Junior, Wesley Nunes Gonçalves, Mário de Araújo Carvalho, Anette Eltner
IGARSS2
2021 Deep Learning and Google Earth Engine Applied to Mapping Eucalyptus
abstract
The potential of integrating deep learning and Google Earth Engine (GEE) is a few explored in the literature. Here, we investigated their potential in the context of Eucalyptus mapping in Brazilian Savanah. Based on GEE API using python language, it is possible to integrate it with Google Colab. The experiments were conducted using the U-Net semantic segmentation method. A total of 704, 88, and 88 patches were used for training, validation, and test, respectively. The overall accuracy obtained in the test dataset was 96.88%, while the Jaccard index was 0.84. These results demonstrated the applicability of these platforms for using deep learning techniques for mapping based on satellite images.
João Otavio Nascimento Firigato, José Marcato Junior, Wesley Nunes Gonçalves, Vitor Matheus Bacani
IGARSS2
2021 Integration of Photogrammetry and Deep Learning in Earth Observation Applications
abstract
The integration of photogrammetry and deep learning methods can be powerful for Earth observation applications. Photogrammetry techniques allow the achievement of detailed geospatial products with em-level positional accuracy. Deep learning enables automatic image classification, segmentation, and object detection. For instance, when dealing with a large data set, photogrammetric processing steps, such as image orientation and dense point cloud generation, results in high computational costs. In contrast, deep learning methods are fast in the inference step. Here, we explore the complementarity of deep learning and photogrammetry, aiming to generate accurate and fast geospatial information. The main aim is to discuss the possibilities of using deep learning in the photogrammetric process. We conduct experiments to present the potential of the Mask R-CNN method trained on the COCO dataset to generate masks, essential to remove image observations from moving objects during the orientation (alignment) step.
José Marcato Junior, Pedro Zamboni, Mariana Batista Campos, Ana Paula Marques Ramos, Lucas Prado Osco, Jonathan de Andrade Silva, Wesley Nunes Gonçalves, Jonathan Li 0001
IGARSS1
2021 Segmentation of Tree Canopies in Urban Environments Using Dilated Convolutional Neural Network
abstract
Object detection and image segmentation are essential for environmental monitoring. This task can be performed and automated using machines with the processing capabilities of convolutional neural networks (CNN), achieving outstanding performance thanks to the current computing capacity and data available. Still, there are advances to perceive and questions on how to get the optimal performance and improve the results. With this aim, we assessed a state-of-the-art CNN, the Dynamic Dilated Convolution Neural Network (DDCN), to segment trees inside an urban environment. We chose DDCN because it exploits the paradigm of multi-context without increasing the number of trainable parameters of the network and defines, while training, the best patch size that should be used by the network in the test phase, helping to tune the CNN. With this technique, we achieved: pixel accuracy of 95.86%; average accuracy of 90.63%; F1-score of 90.93%; Kappa index of 81.87% and IoU of 72.78%. We are showing the capabilities of this CNN to segment complex images.
José Augusto Correa Martins, Keiller Nogueira, Pedro Zamboni, Paulo Tarso Sanches de Oliveira, Wesley Nunes Gonçalves, Jefersson A. dos Santos, José Marcato Junior
IGARSS7
2021 Retinanet Deep Learning-Based Approach to Detect Termite Mounds in Eucalyptus Forests
abstract
Brazilian eucalyptus plantations are largely used to produce wood products and play an important role in the country's economy. The presence of termites in these plantations can cause serious damage, requiring their identification and control. Here, we investigated the RetinaNet object detection deep learning method's potential to identify termite mounds in eucalyptus forests automatically. The experiments were conducted using terrestrial images acquired with a GoPro 360, composed of two fisheye cameras, providing a challenging task due to non-conventional geometry. In the tests, the model presented an Average Precision (AP) of 0.721, and after stipulating thresholds from which the objects would be used in training, the AP was 0.867.
Juan Sales, José Marcato Junior, Henrique Lopes Siqueira, Maurício de Souza, Edson Takashi Matsubara, Wesley Nunes Gonçalves
IGARSS2
2021 Assessment of CNN-Based Methods for Single Tree Detection on High-Resolution RGB Images in Urban Areas
abstract
Maintaining vegetation cover in cities is a key component to keep cities safe and resilient. The monitoring of trees is usually done with LiDAR data or multi and hyperspectral images. In this sense, remote sensing RGB images are presented as a cheaper and easier processing solution. Here, we proposed to evaluate deep learning-based methods combined with high-resolution RGB images to detect single-trees in the urban environment. Three state-of-the-art methods are tested: Faster-RCNN, RetinaNet, and ATSS. A total of 220 images were used, in which we manually labeled 3382 trees. For the proposal task, our findings show that ATSS is 3% more accurate than Faster-RCNN and 4% than RetinaNet. However, in a qualitative inspection, Faster-RCNN and RetinaNet seems to be better at this task. Our findings shows the need of further research for developing suitable tools for urban tree detection. This tools can help cities top achieve a more sustainable and resilient environment especially to face climate change.
Pedro Alberto Pereira ZamboniThgeThe, José Marcato Junior, Gabriela Takahashi Miyoshi, Jonathan de Andrade Silva, José Augusto Correa Martins, Wesley Nunes Gonçalves
IGARSS2
2021 Capsule-Based Networks for Road Marking Extraction and Classification From Mobile LiDAR Point Clouds
abstract
Accurate road marking extraction and classification play a significant role in the development of autonomous vehicles (AVs) and high-definition (HD) maps. Due to point density and intensity variations from mobile laser scanning (MLS) systems, most of the existing thresholding-based extraction methods and rule-based classification methods cannot deliver high efficiency and remarkable robustness. To address this, we propose a capsule-based deep learning framework for road marking extraction and classification from massive and unordered MLS point clouds. This framework mainly contains three modules. Module I is first implemented to segment road surfaces from 3D MLS point clouds, followed by an inverse distance weighting (IDW) interpolation method for 2D georeferenced image generation. Then, in Module II, a U-shaped capsule-based network is constructed to extract road markings based on the convolutional and deconvolutional capsule operations. Finally, a hybrid capsule-based network is developed to classify different types of road markings by using a revised dynamic routing algorithm and large-margin Softmax loss function. A road marking dataset containing both 3D point clouds and manually labeled reference data is built from three types of road scenes, including urban roads, highways, and underground garages. The proposed networks were accordingly evaluated by estimating robustness and efficiency using this dataset. Quantitative evaluations indicate the proposed extraction method can deliver 94.11% in precision, 90.52% in recall, and 92.43% in F1-score, respectively, while the classification network achieves an average of 3.42% misclassification rate in different road scenes.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, Yongtao Yu, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.5
2020 Characterization of MSS Channel Reflectance and Derived Spectral Indices for Building Consistent Landsat 1-5 Data Record
abstract
The Landsat 1-5 multispectral scanner system (MSS) collected records of land surface mainly during 1972-1992. Investigations on MSS have been relatively limited compared with the numerous investigations on its successors, such as Thematic Mapper (TM) and Enhanced TM Plus (ETM+). The benefits of the Landsat program are not fully accomplished without the inclusion of MSS archives. Investigations on the Landsat 1-5 MSS channel reflectance characteristics wereperformed followed by derived vegetation spectral indices and the Tasseled Cap (TC) transformed features mainly using a collection of synthesized records. On average, the Landsat 4 MSS is generally comparable to the Landsat 5 MSS. The Landsat 1-3 MSSs show disagreement in channel reflectance compared with the Landsat 5 MSS, especially for the red channel (600-700 nm) and the near-infrared channel (700-800 nm). Meanwhile, the relative differences for vegetation spectral indices of the Landsat 3 MSS are mainly from -16% to -5% with the median about -11.5%, while those of the Landsat 2 MSS are mainly from -15% to -7%. Cross-validation tests and two case applications suggested that between-sensor consistency was improved generally through the transformation models generated by ordinary least-squares regression. To improve the consistency of the vegetation indices and the TC greenness, direct strategy employing respective transformation models was more effective than calculations based on the transformed channel reflectance. Considering the shortages of the Landsat MSS archives, further efforts are needed to improve its comparability with observations by other successive Landsat sensors.
Feng Chen 0022, Qiancong Fan, Shenlong Lou, Martin Claverie, Cheng Wang 0003, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2019 Asphalt Pothole Detection in UAV Images Using Convolutional Neural Networks
abstract
Transportation infrastructure needs constant maintenance. Pavement management systems requires reliable and detailed data of the current state of the roads to make effective decisions. Currently, pavement condition evaluation methods are mostly performed manually with visual inspection and interpretations in situ, which is labor intensive, time consuming and expensive. In this paper an experiment was conducted where a set of different configurations and parameters for Convolutional Neural Networks (CNNs) were applied to automatically detect potholes from images taken by an Unmanned Aerial Vehicle (UAV). Results showed that the pre-trained Faster-RCNN Inception ResNet model with reduced anchor box stride and image augmentation applied provides better accuracy compared to several other models tested, obtaining accuracy for this experiment of 70.4% across five-fold cross validation.
Yuri V. Furusho Becker, Henrique Lopes Siqueira, Edson Takashi Matsubara, Wesley Nunes Gonçalves, José Marcato Junior
IGARSS5
2019 Machine Learning Applied to Uav Imagery in Precision Agriculture and Forest Monitoring in Brazililian Savanah
abstract
The Brazilian Savanah is one of the most important biomes of South America. It has an area of 2,036,448 km2(around 22% of the country). Several endangered tree species are protected by law and recently are threatened in the last years. In addition, the country is the second major producer of soybean in the world. However, some insect species have been causing great economic damage in the soybean fields, and the Integrated Pest Management is a key factor for the attack control of different species. We design and implement an end-to-end observing system based on UAV (Unmanned Aerial Vehicle) to support precision agriculture and the forest monitoring. The current research is one of the projects approved by GRSS Grand Challenge, and is under development. It was developed two UAVs, one for each problem. Machine learning techniques were used and the best accuracy obtained performance reaching a classification rate of 99.04%. Others preliminary results in both precision farming and forest monitoring applications are described in this paper.
David Robledo Di Martini, Veraldo Liesenberg, Everton Castelao Tetila, José Marcato Junior, Edson Takashi Matsubara, Henrique Lopes Siqueira, Amaury A. C. Junior, Márcio Santos Araújo, Carlos Henrique Monteiro, Hemerson Pistori
IGARSS4
2019 Image Segmentation and Classification with SLIC Superpixel and Convolutional Neural Network in Forest Context
abstract
Unmanned aerial vehicles (UAVs) are platforms suitable for obtaining information utilizing sensors in a great variety of areas and enviroments, thus in this context, this paper objective was to identify trees in images collected using UAV high-resolution imagery, with the digital approach of superpixel segmentation and convolutional neural networks. A forest environment was analyzed in the form of an orthomosaic that was produced using 423 images. Superpixels were generated using the Simple Linear Iterative Clustering (SLIC) method, considering only two classes for the classification: trees and background. The generation of superpixels (segmentation) and classification were performed considering the configurations: 2000, 3000 and 4000 segments, sigma 5 and compactness 10 for SLIC and 100% transfer of learning. For the purposes of classification, the deep convolutional networks were adopted through ResNet-50 architecture, the weights of this network were previously trained in an Image bank and later the bank of images of superpixels underwent a fine-tuning. The experiments obtained a maximum classification accuracy of 89% with 3000 superpixels distinction between canopies and Background.
José Augusto Correa Martins, José Marcato Junior, Geazy Vilharva Menezes, Hemerson Pistori, Diego Sant'Ana, Wesley Nunes Gonçalves
IGARSS2
2019 The Impact of Ground Control Point Quantity on Area and Volume Measurements with UAV SFM Photogrammetry Applied in Open Pit Mines
abstract
Even though the use of unmanned aerial vehicles (UAV) has been widely spread over the area of remote sensing, there are still uncertainties regarding the impact of parameters such as the number of ground control points (GCP), the image overlap or the ground sampling distance (GSD). In this study, we evaluate the accuracy of volumetric surveys performed with UAV imagery in open pit mines, assessing the influence of the number of GCPs on area and volume estimation. Our results suggest a density of 15 GCP/km2when performing precise volume measurements. In cases where only volume values are required and errors are acceptable below 10%, no GCPs are needed.
Henrique Lopes Siqueira, José Marcato Junior, Edson Takashi Matsubara, Anette Eltner, Reinaldo Almeida Colares, Fabio Martins Santos
IGARSS2