EDBT 2026 Demo / reviewers in the wild / expert
Naoto Yokoya
dblp:79/8993
· DBLP profile ↗
111ranked-venue papers
14as first author
59since 2021 · last 2026
0000-0002-7321-4590ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 66 · 13 first-author · 27 since 2021Artificial intelligence and machine learning · 28 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LandCraft: Designing the Structured 3D Landscapes via Text GuidanceabstractModeling large-scale landscapes is a foundational yet time-consuming task in many 3D applications, typically requiring substantial expertise. Recently, Text-to-3D techniques have emerged as a promising, beginner-friendly prototyping approach for generating 3D content from textual input. However, existing methods either produce unusable, problematic geometries, or fail to fully capture the user's complex intent from the input text—making it difficult to generate high-quality landscape assets with controllable spatial and geographic features. In this paper, we present LandCraft, a novel AI-assisted authoring tool that enables the rapid creation of high-quality landscape scenes based on user descriptions. Our system employs a coarse-to-fine generation process: Initially, large language and deep generative models concretize textual ideas into abstract representations that capture essential landscape features, such as spatial and geographic characteristics. Then, we leverage a comprehensive procedural generation module to synthesize the detailed, structurally consistent 3D landscapes based on these inferred representations. LandCraft can effectively generate production-ready 3D scene assets that can be seamlessly exported to external game engines or modeling software, enabling immediate practical use. Weihao Xuan, Naoto Yokoya |
AAAI | 4 |
| 2026 | The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsabstractAutonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge.A fundamental pillar of this trustworthiness is calibration, which refers to an agent's ability to express confidence that reliably reflects its actual performance.While calibration is well-established for static models, its dynamics in tool-integrated agentic workflows remain under-explored.In this work, we systematically investigate verbalized calibration in tooluse agents, revealing a fundamental confidence dichotomy driven by tool type.Specifically, our pilot study identifies that evidence tools (e.g., web search) systematically induce severe overconfidence due to inherent noise in retrieved information, while verification tools (e.g., code interpreters) can ground reasoning through deterministic feedback and mitigate miscalibration.To robustly improve calibration across tool types, we propose a reinforcement learning (RL) fine-tuning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs.We demonstrate that our trained agents not only achieve superior calibration but also exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning.Our results highlight the necessity of domain-specific calibration strategies for tooluse agents.More broadly, this work establishes a foundation for building self-aware agents that can reliably communicate uncertainty in highstakes, real-world deployments. Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Naoto Yokoya |
ACL (1) | 6 |
| 2026 | Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial ImageryabstractThree-dimensional geospatial analysis is critical to applications in urban planning, climate adaptation, and environmental assessment. Current methodologies depend on costly, specialized sensors (e.g., LiDAR and multispectral), which restrict global accessibility. Existing sensor-based and rule-driven methods further struggle with tasks requiring the integration of multiple 3D cues, handling diverse queries, and providing interpretable reasoning. We hereby present Geo3DVQA, a comprehensive benchmark for evaluating vision-language models (VLMs) in height-aware, 3D geospatial reasoning using RGB-only remote sensing imagery. Unlike conventional sensor-based frameworks, Geo3DVQA emphasizes realistic scenarios that integrate elevation, sky view factors, and land cover patterns. The benchmark encompasses 110k curated question–answer pairs spanning 16 task categories across three complexity levels: single-feature inference, multi-feature reasoning, and application-level spatial analysis. The evaluation of ten state-of-the-art VLMs highlights the difficulty of RGB-to-3D reasoning. GPT-4o and Gemini-2.5-Flash achieved only 28.6% and 33.0% accuracy respectively, while domain-specific fine-tuning of Qwen2.5-VL-7B achieved 49.6% (+24.8 points). These results reveal both the limitations of current VLMs and the effectiveness of domain adaptation. Geo3DVQA introduces a new frontier of challenges for scalable, accessible, and holistic 3D geospatial analysis. The dataset and code will be released upon publication at https://github.com/mm1129/Geo3DVQA. Mai Tsujimoto, Weihao Xuan, Naoto Yokoya |
WACV | 4 |
| 2026 | CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic SegmentationabstractDue to the substantial domain gaps in Remote Sensing (RS) images that are characterized by variabilities such as location, wavelength, and sensor type, Remote Sensing Domain Generalization (RSDG) has emerged as a critical and valuable research frontier, focusing on developing models that generalize effectively across diverse scenarios. However, research in this area remains underexplored: (1) Current cross-domain methods primarily focus on Domain Adaptation (DA), which adapts models to predefined domains rather than to unseen ones; (2) Few studies target the RSDG issue, especially for semantic segmentation tasks. Existing related models are developed for specific unknown domains, struggling with issues of underfitting on other unseen scenarios; (3) Existing RS foundation models tend to prioritize in-domain performance over cross-domain generalization. To this end, we introduce the first vision foundation model for RSDG semantic segmentation, CrossEarth. CrossEarth demonstrates strong cross-domain generalization through a specially designed data-level Earth-Style Injection pipeline and a model-level Multi-Task Training pipeline. In addition, for the semantic segmentation task, we have curated an RSDG benchmark comprising 32 semantic segmentation scenarios across various regions, spectral bands, platforms, and climates, providing comprehensive evaluations of the generalizability of future RSDG models. Extensive experiments on this collection demonstrate the superiority of CrossEarth over existing state-of-the-art methods. Ziyang Gong, Zhixiang Wei, Di Wang 0023, Xiaoxing Hu, Xianzheng Ma, Hongruixuan Chen, Yuru Jia, Yupeng Deng 0002, Zhenming Ji, Xiangwei Zhu, Xue Yang 0005, Naoto Yokoya, Jing Zhang 0037, Bo Du 0001, Junchi Yan, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2025 | Neural Hierarchical Decomposition for Single Image Plant ModelingabstractObtaining high-quality, practically usable 3D models of biological plants remains a significant challenge in computer vision and graphics. In this paper, we present a novel method for generating realistic 3D plant models from single-view photographs. Our approach employs a neural decomposition technique to learn a lightweight hierarchical box representation from the image, effectively capturing the structures and botanical features of plants. Then, this representation can be subsequently refined through a shape-guided parametric modeling module to produce complete 3D plant models. By combining hierarchical learning and parametric modeling, our method generates structured 3D plant assets with fine geometric details. Notably, through learning the decomposition in different levels of detail, our method can adapt to two distinct plant categories: outdoor trees and houseplants, each with unique appearance features. Within the scope of plant modeling, our method is the first comprehensive solution capable of reconstructing both plant categories from single-view images. Zhanglin Cheng, Naoto Yokoya |
CVPR | 3 |
| 2025 | Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language ModelsabstractUncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems.Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language models (LLMs).However, its effectiveness in vision-language models (VLMs) remains insufficiently studied.In this work, we conduct a comprehensive evaluation of verbalized confidence in VLMs, spanning three model categories, four task domains, and three evaluation scenarios.Our results show that current VLMs often display notable miscalibration across diverse tasks and settings.Notably, visual reasoning models (i.e., thinking with images) consistently exhibit better calibration, suggesting that modality-specific reasoning is critical for reliable uncertainty estimation.To further address calibration challenges, we introduce VI-SUAL CONFIDENCE-AWARE PROMPTING, a two-stage prompting strategy that improves confidence alignment in multimodal settings.Overall, our study highlights the inherent miscalibration in VLMs across modalities.More broadly, our findings underscore the fundamental importance of modality alignment and model faithfulness in advancing reliable multimodal systems.Mirko Borszukovszki, Ivo Pascal De Jong, and Matias Valdenegro-Toro.2025.Know what you do not know: Verbalized uncertainty estimation robustness on corrupted images in vision-language models.In Weihao Xuan, Qingcheng Zeng, Heli Qi, Naoto Yokoya |
EMNLP | 5 |
| 2025 | GaussianOcc: Fully Self-Supervised and Efficient 3D Occupancy Estimation with Gaussian SplattingabstractWe introduce GaussianOcc, a systematic method that investigates the two usages of Gaussian splatting for fully self-supervised and efficient 3D occupancy estimation in surround views. First, traditional methods for self-supervised 3D occupancy estimation still require ground truth 6D poses from sensors during training. To address this limitation, we propose Gaussian Splatting for Projection (GSP) module to provide accurate scale information for fully self-supervised training from adjacent view projection. Additionally, existing methods rely on volume rendering for final 3D voxel representation learning using 2D signals (depth maps, semantic maps), which is both time-consuming and less effective. We propose Gaussian Splatting from Voxel space (GSV) to leverage the fast rendering properties of Gaussian splatting. As a result, the proposed GaussianOcc method enables fully self-supervised (no ground truth pose) 3D occupancy estimation in competitive performance with low computational cost (2.7 times faster in training and 5 times faster in rendering). The relevant code is available in https://github.com/GANWANSHUI/GaussianOcc.git. Wanshui Gan, Ningkai Mo, Naoto Yokoya |
ICCV | 5 |
| 2025 | MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image RestorationabstractHyperspectral images (HSIs) often suffer from diverse and unknown degradations during imaging, leading to severe spectral and spatial distortions. Existing HSI restoration methods typically rely on specific degradation assumptions, limiting their effectiveness in complex scenarios. In this paper, we propose \textbf{MP-HSIR}, a novel multi-prompt framework that effectively integrates spectral, textual, and visual prompts to achieve universal HSI restoration across diverse degradation types and intensities. Specifically, we develop a prompt-guided spatial-spectral transformer, which incorporates spatial self-attention and a prompt-guided dual-branch spectral self-attention. Since degradations affect spectral features differently, we introduce spectral prompts in the local spectral branch to provide universal low-rank spectral patterns as prior knowledge for enhancing spectral reconstruction. Furthermore, the text-visual synergistic prompt fuses high-level semantic representations with fine-grained visual features to encode degradation information, thereby guiding the restoration process. Extensive experiments on 9 HSI restoration tasks, including all-in-one scenarios, generalization tests, and real-world cases, demonstrate that MP-HSIR not only consistently outperforms existing all-in-one methods but also surpasses state-of-the-art task-specific approaches across multiple tasks. The code and models are available at https://github.com/ZhehuiWu/MP-HSIR. Zhehui Wu, Yong Chen 0013, Naoto Yokoya, Wei He 0003 |
ICCV | 3 |
| 2025 | LR2Depth: Large-Region Aggregation at Low Resolution for Efficient Monocular Depth EstimationabstractMonocular depth estimation (MDE) is crucial for various computer vision applications, but existing methods often struggle to balance inference speed and accuracy when processing large-region visual information. This paper introduces LR2Depth, a novel MDE method that addresses this challenge by utilizing large-kernel convolution on low-resolution feature maps for efficient large-region feature aggregation. Our approach leverages the fact that each pixel on low-resolution feature maps corresponds to a larger region of the original image, allowing for fast and accurate depth predictions at a lower inference cost. Extensive experiments on NYU-Depth-V2, KITTI, and SUN RGB-D datasets demonstrate that LR2Depth not only achieves state-of-the-art performance but also operates approximately twice as fast as previous MDE methods. Notably, at the time of submission, LR2Depth secured the top-1 position on the KITTI depth prediction online benchmark. The code is available in the project page. Chao Ning 0011, Weihao Xuan, Wanshui Gan, Naoto Yokoya |
IROS | 4 |
| 2025 | DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and ResponseabstractLarge vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geographic regions, and satellite sensors have posed new challenges for VLM applications. To fill this gap, we curate the first remote sensing vision-language dataset (DisasterM3) for global-scale disaster assessment and response. DisasterM3 includes 26,988 bi-temporal satellite images and 123k instruction pairs across 5 continents, with three characteristics: **1) Multi-hazard**: DisasterM3 involves 36 historical disaster events with significant impacts, which are categorized into 10 common natural and man-made disasters. **2) Multi-sensor**: Extreme weather during disasters often hinders optical sensor imaging, making it necessary to combine Synthetic Aperture Radar (SAR) imagery for post-disaster scenes. **3) Multi-task**: Based on real-world scenarios, DisasterM3 includes 9 disaster-related visual perception and reasoning tasks, harnessing the full potential of VLM's reasoning ability with progressing from disaster-bearing body recognition to structural damage assessment and object relational reasoning, culminating in the generation of long-form disaster reports. We extensively evaluated 14 generic and remote sensing VLMs on our benchmark, revealing that state-of-the-art models struggle with the disaster tasks, largely due to the lack of a disaster-specific corpus, cross-sensor gap, and damage object counting insensitivity. Focusing on these issues, we fine-tune four VLMs using our dataset and achieve stable improvements (up to 10.4\%$\uparrow$QA, 2.1$\uparrow$Report, 40.8\%$\uparrow$Referring Seg.) with robust cross-sensor and cross-disaster generalization capabilities. Project: https://github.com/Junjue-Wang/DisasterM3. Weihao Xuan, Heli Qi, Kunyi Liu, Hongruixuan Chen, Jian Song 0010, Junshi Xia, Zhuo Zheng, Naoto Yokoya |
NeurIPS | 11 |
| 2025 | DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City UnderstandingabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite comprises 14,871 high-resolution (1.0m) multi-temporal images spanning 42 major cities in the U.S. from 2005 to 2023, organized into two components: DVL-Bench and DVL-Instruct. The DVL-Bench includes six urban understanding tasks, from fundamental change detection (pixel-level) to quantitative analyses (regional-level) and comprehensive urban narratives (scene-level), capturing diverse urban dynamics including expansion/transformation patterns, disaster assessment, and environmental challenges. We evaluate 18 state-of-the-art MLLMs and reveal their limitations in long-term temporal understanding and quantitative analysis. These challenges motivate the creation of DVL-Instruct, a specialized instruction-tuning dataset designed to enhance models' capabilities in multi-temporal Earth observation. Building upon this dataset, we develop DVLChat, a baseline model capable of both image-level question-answering and pixel-level segmentation, facilitating a comprehensive understanding of city dynamics through language interactions. Project: https://github.com/weihao1115/dynamicvl. Weihao Xuan, Heli Qi, Zihang Chen 0001, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya |
NeurIPS | 8 |
| 2025 | Wavelength- and Depth-Aware Deep Image Prior for Blind Hyperspectral Imagery Deblurring with Coarse Depth GuidanceabstractHyperspectral imagery (HSI) provides detailed spectral information, enabling precise analysis of materials. However, HSI imaging suffers from blurring degradation which results in the loss of fine details and hinders subsequent applications. The degree of blurriness is highly related to wavelength and depth, existing deblurring methods either lack the utilization of spectral correlation or ignore the depth variation since paired HSI and depth data are difficult to acquire and less discussed, leading to degraded performance when encountering wide-range HSIs of non-planar scenes. To address these challenges in both data acquisition and algorithm design, we propose a novel approach that simultaneously collects both modalities and integrates depth refinement into a blind HSI deblurring model with wavelength- and depth-aware deep image prior. Specifically, we capture blurred HSI and coarse depth map with separate devices, followed by registration. Our method performs depth-guided deblurring through depth-variant multi-channel kernel estimation and soft-weight map-based layer composition, while simultaneously refining the depth. The proposed approach effectively restores fine details with fewer artifacts, showing superior performance for both simulated blurred HSIs and real captured HSIs. Jiahuan Li, Wei He 0003, Naoto Yokoya |
WACV | 4 |
| 2025 | Multi-modal consistent loss diffusion model for Sentinel-3 single image super resolutionabstractAbstract In the context of Earth observation, the trade-off between spatial, spectral, and temporal resolution often limits the versatility of remote sensing images in many important applications. In response, this paper introduces a novel deep learning diffusion model, specifically tailored to improve the spatial resolution of the optical products acquired by the Sentinel-3 (S3) satellite. Our framework employs a diffusion probabilistic model, benefiting from the higher spatial resolution of the Sentinel-2 satellite during training via a new multi-modal loss formulation. This ensures consistency with the original S3 images while enhancing the spatial details. Two distinct conditional low-resolution encoders were experimented with, providing insights into their respective contributions to the diffusion process. The efficacy of the proposed model is demonstrated through extensive ablation studies and comparisons with state-of-the-art methods, using both synthetic and real S3 products. The findings indicate that our model successfully improves spatial resolution while maintaining the integrity of the spectral information, contributing to the field of remote sensing single-image super-resolution. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla, Naoto Yokoya, Junshi Xia |
Neural Comput. Appl. | 4 |
| 2025 | HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelabstractAccurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency. Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2025 | Inter-Sensor High-Resolution and Multi-Temporal Image Fusion for Unsupervised Domain Adaptation in Remote SensingabstractMotivated by the increasing demand for robust segmentation in unlabeled remote sensing data, we propose DAM-Former, a novel UDA model that fuses high-resolution multimodal imagery with multi-temporal multispectral data. Current UDA approaches in remote sensing rarely exploit the complementary strengths of spatial and temporal features. To address this gap, our framework integrates two interconnected branches: a transformer-based network for high-resolution multimodal data and a lightweight convolutional network with temporal attention for multi-temporal imagery. To improve segmentation accuracy and lower noise, the extracted features are robustly combined through a deep temporal fusion module and a new mixed loss with an ensemble pseudo-label strategy. Extensive experiments and an ablation study on the FLAIR-2 dataset demonstrate that DAM-Former outperforms state-of-the-art methods, marking the first in-depth study of temporal information fusion in UDA segmentation for remote sensing data. Damian Ibañez, Junshi Xia, Naoto Yokoya, Filiberto Pla, Rubén Fernández-Beltran |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | CoMiX: Cross-Modal Fusion With Deformable Convolutions for HSI-X Semantic SegmentationabstractImproving hyperspectral image (HSI) semantic segmentation by exploiting complementary information from supplementary modalities (termed X-modality) is promising but challenging due to significant differences in imaging sensors, image content, and resolution. Existing methods often underutilize the unique spatial–spectral features of HSIs by processing them uniformly with X-modality data. In addition, current cross-modality fusion strategies often suffer from limited intermodal interaction or significantly increased model complexity. To address these limitations, we propose CoMiX, an asymmetric encoder-decoder architecture with deformable convolutions (DCNs) for HSI-X semantic segmentation. CoMiX includes an encoder with two parallel, interacting backbones and a lightweight all-multilayer perceptron (ALL-MLP) decoder. The encoder consists of four stages, each incorporating 2D DCN blocks for the X-modality to accommodate geometric variations and 3D DCN blocks for HSIs to adaptively capture spatial-spectral features. Each stage also incorporates a Cross-Modality Feature enhancement and eXchange (CMFeX) module and a feature fusion module (FFM). CMFeX exploits spatial-spectral correlations across modalities to recalibrate and enhance modality-specific and modality-shared features, while adaptively exchanging complementary information. Its outputs are subsequently fused in the FFM and propagated to the next stage for further learning. Finally, the ALL-MLP decoder aggregates the fused features from all stages to produce the final predictions. Extensive experiments demonstrate that CoMiX achieves state-of-the-art performance and generalizes well to various multimodal datasets. The CoMiX code will be released soon. Xuming Zhang 0004, Naoto Yokoya, Xingfa Gu, Qingjiu Tian, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Cross-Scale Style-Guided Enhancement for Underwater Remote Sensing ImageryabstractUnderwater imaging is essential for marine remote sensing tasks, such as environmental monitoring, resource exploration, and autonomous navigation. However, images captured in underwater environments often suffer from complex degradations, including wavelength-dependent color distortion, contrast attenuation, and structural detail loss. To address these challenges, we propose a Cross-Scale Style-Guided Network (CSG-Net) for robust underwater image enhancement. CSG-Net employs a dual-stage collaborative framework that decouples global degradation modeling from local detail refinement. In the first stage, a Style Extraction Network (SE-Net) extracts multi-scale degradation-aware style priors that implicitly encode large-scale physical degradation patterns, such as red-channel attenuation and spectral imbalance. In the second stage, a Style-Guided Enhancement Network (SG-Net) leverages these style features to guide spatially adaptive enhancement, enabling consistent color correction and fine-grained detail recovery. To alleviate semantic degradation during scale transitions, CSG-Net introduces the proposed multi-resolution feature-preserving cross-scale interaction (MFPCSI) module, which enhances the preservation and integration of hierarchical features. Combined with the Multi-Stream Information Fusion (MSIF) module, this design enables the effective fusion of semantic and structural information across spatial scales. The proposed components enable the preservation of fine-grained details while adaptively integrating semantic and structural cues across multiple scales. Comprehensive experiments conducted on diverse and challenging underwater image datasets demonstrate that CSG-Net consistently surpasses state-of-the-art approaches in terms of PSNR, SSIM, and UIQM. Furthermore, the model exhibits strong cross-domain generalization and delivers high-fidelity visual results, underscoring its suitability for deployment in practical vision systems operating under complex, real-world environments. Jingchun Zhou, Dehuan Zhang, Xingcheng Han, Qiuping Jiang, Gemine Vivone, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | DeepTreeSketch: Neural Graph Prediction for Faithful 3D Tree Modeling from SketchesabstractWe present DeepTreeSketch, a novel AI-assisted sketching system that enables users to create realistic 3D tree models from 2D freehand sketches. Our system leverages a tree graph prediction network, TGP-Net, to learn the underlying structural patterns of trees from a large collection of 3D tree models. The TGP-Net simulates the iterative growth of botanical trees and progressively constructs the 3D tree structures in a bottom-up manner. Furthermore, our system supports a flexible sketching mode for both precise and coarse control of the tree shapes by drawing branch strokes and foliage strokes, respectively. Combined with a procedural generation strategy, users can freely control the foliage propagation with diverse and fine details. We demonstrate the expressiveness, efficiency, and usability of our system through various experiments and user studies. Our system offers a practical tool for 3D tree creation, especially for natural scenes in games, movies, and landscape applications. Fangyuan Tu, Ruiyuan Zhang, Zhanglin Cheng, Naoto Yokoya |
CHI | 6 |
| 2024 | Change Detection between Optical Remote Sensing Imagery and Map Data Via Segment Anything Model (SAM)abstractUnsupervised multimodal change detection is pivotal for time-sensitive tasks and comprehensive multi-temporal Earth monitoring. In this study, we explore unsupervised multimodal change detection between two key remote sensing data sources: optical high-resolution imagery and OpenStreetMap (OSM) data. Specifically, we propose to utilize the vision foundation model Segmentation Anything Model (SAM), for addressing our task. Leveraging SAM’s exceptional zeroshot transfer capability, high-quality segmentation maps of optical images can be obtained. Thus, we can directly compare these two heterogeneous data forms in the so-called segmentation domain. We then introduce two strategies for guiding SAM’s segmentation process: the ‘no-prompt’ and ‘box/mask prompt’ methods. The two strategies are designed to detect land-cover changes in general scenarios and to identify new land-cover objects within existing backgrounds, respectively. Experimental results on three datasets indicate that the proposed approach can achieve more competitive results compared to representative unsupervised multimodal change detection methods. Hongruixuan Chen, Jian Song 0010, Naoto Yokoya |
IGARSS | 3 |
| 2024 | When Daformer Meets Multi-Modality DatasetsabstractWe introduce innovative unsupervised domain adaptation (UDA) techniques that leverage the integration of DAFomer and cross-attention mechanisms, tailored to effectively handle multi-modal datasets. We investigate the methods on the FLAIR-1 dataset with different domains, including RGB, NIR bands, and height information. Experimental findings strongly support the idea that integrating inter-modal information significantly enhances segmentation accuracy, making the model more versatile and effective in handling multi-modal datasets across diverse conditions and domains. Damian Ibañez, Junshi Xia, Naoto Yokoya |
IGARSS | 3 |
| 2024 | OpenEarthMap Benchmark Suite and Its ApplicationsabstractWe present the OpenEarthMap benchmark suite, designed for global high-resolution land cover mapping and change analysis, and showcase its applications. This comprehensive expansion aims to strengthen OpenEarthMap’s versatility, covering various aspects such as increasing dataset diversity through synthetic data, developing lightweight models, improving land cover change detection with OpenStreetMap, and facilitating high-resolution mapping on a national scale. Naoto Yokoya, Junshi Xia, Clifford Broni-Bediako, Jian Song 0010, Hongruixuan Chen |
IGARSS | 1 |
| 2024 | SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing ImageryabstractGlobal semantic 3D understanding from single-view high-resolution remote sensing (RS) imagery is crucial for Earth observation (EO). However, this task faces significant challenges due to the high costs of annotations and data collection, as well as geographically restricted data availability. To address these challenges, synthetic data offer a promising solution by being unrestricted and automatically annotatable, thus enabling the provision of large and diverse datasets. We develop a specialized synthetic data generation pipeline for EO and introduce SynRS3D, the largest synthetic RS dataset. SynRS3D comprises 69,667 high-resolution optical images that cover six different city styles worldwide and feature eight land cover types, precise height information, and building change masks. To further enhance its utility, we develop a novel multi-task unsupervised domain adaptation (UDA) method, RS3DAda, coupled with our synthetic dataset, which facilitates the RS-specific transition from synthetic to real scenarios for land cover mapping and height estimation tasks, ultimately enabling global monocular 3D semantic understanding based on synthetic data. Extensive experiments on various real-world datasets demonstrate the adaptability and effectiveness of our synthetic dataset and the proposed RS3DAda method. SynRS3D and related codes are available at https://github.com/JTRNEO/SynRS3D. Jian Song 0010, Hongruixuan Chen, Weihao Xuan, Junshi Xia, Naoto Yokoya |
NeurIPS | 5 |
| 2024 | Understanding Dark Scenes by Contrasting Multi-Modal ObservationsabstractUnderstanding dark scenes based on multi-modal image data is challenging, as both the visible and auxiliary modalities provide limited semantic information for the task. Previous methods focus on fusing the two modalities but neglect the correlations among semantic classes when minimizing losses to align pixels with labels, resulting in inaccurate class predictions. To address these issues, we introduce a supervised multi-modal contrastive learning approach to increase the semantic discriminability of the learned multi-modal feature spaces by jointly performing cross-modal and intra-modal contrast under the supervision of the class correlations. The cross-modal contrast encourages same-class embeddings from across the two modalities to be closer and pushes different-class ones apart. The intra-modal contrast forces same-class or different-class embeddings within each modality to be together or apart. We validate our approach on a variety of tasks that cover diverse light conditions and image modalities. Experiments show that our approach can effectively enhance dark scene understanding based on multi-modal images with limited semantics by shaping semantic-discriminative feature spaces. Comparisons with previous methods demonstrate our state-of-the-art performance. Code and pretrained models are available at https://github.com/palmdong/SMMCL. Naoto Yokoya |
WACV | 2 |
| 2024 | SyntheWorld: A Large-Scale Synthetic Dataset for Land Cover Mapping and Building Change DetectionabstractSynthetic datasets, recognized for their cost effectiveness, play a pivotal role in advancing computer vision tasks and techniques. However, when it comes to remote sensing image processing, the creation of synthetic datasets becomes challenging due to the demand for larger-scale and more diverse 3D models. This complexity is compounded by the difficulties associated with real remote sensing datasets, including limited data acquisition and high annotation costs, which amplifies the need for high-quality synthetic alternatives. To address this, we present SyntheWorld, a synthetic dataset unparalleled in quality, diversity, and scale. It includes 40,000 images with submeter-level pixels and fine-grained land cover annotations of eight categories, and it also provides 40,000 pairs of bitemporal image pairs with building change annotations for building change detection. We conduct experiments on multiple benchmark remote sensing datasets to verify the effectiveness of SyntheWorld and to investigate the conditions under which our synthetic data yield advantages. The dataset is available at https://github.com/JTRNEO/SyntheWorld. Jian Song 0010, Hongruixuan Chen, Naoto Yokoya |
WACV | 3 |
| 2024 | Guest Editorial: Advanced image restoration and enhancement in the wildabstractImage restoration and enhancement has always been a fundamental task in computer vision and is widely used in numerous applications, such as surveillance imaging, remote sensing, and medical imaging. In recent years, remarkable progress has been witnessed with deep learning techniques. Despite the promising performance achieved on synthetic data, compelling research challenges remain to be addressed in the wild. These include: (i) degradation models for low-quality images in the real world are complicated and unknown, (ii) paired low-quality and high-quality data are difficult to acquire in the real world, and a large quantity of real data are provided in an unpaired form, (iii) it is challenging to incorporate cross-modal information provided by advanced imaging techniques (e.g. RGB-D camera) for image restoration, (iv) real-time inference on edge devices is important for image restoration and enhancement methods, and (v) it is difficult to provide the confidence or performance bounds of a learning-based method on different images/regions. This special issue invites original contributions in datasets, innovative architectures, and training methods for image restoration and enhancement to address these and other challenges. In this Special Issue, we have received 17 papers, of which 8 papers underwent the peer review process, while the rest were desk-rejected. Among these reviewed papers, 5 papers have been accepted and 3 papers have been rejected as they did not meet the criteria of IET Computer Vision. Thus, the overall submissions were of high quality, which marks the success of this Special Issue. The five eventually accepted papers can be clustered into two categories, namely video reconstruction and image super-resolution. The first category of papers aims at reconstructing high-quality videos. The papers in this category are of Zhang et al., Gu et al., and Xu et al. The second category of papers studies the task of image super-resolution. The papers in this category are of Dou et al. and Yang et al. A brief presentation of each of the paper in this special issue is as follows. Zhang et al. propose a point-image fusion network for event-based frame interpolation. Temporal information in event streams plays a critical role in this task as it provides temporal context cues complementary to images. Previous approaches commonly transform the unstructured event data to structured data formats through voxelisation and then employ advanced CNNs to extract temporal information. However, the voxelisation operation inevitably leads to information loss and introduces redundant computation. To address these limitations, the proposed method directly extracts temporal information from the events at the point level without relying on any voxelisation operation. Afterwards, a fusion module is adopted to aggregate complementary cues from both points and images for frame interpolation. Experiments on both synthetic and real-world datasets show that their method produces state-of-the-art accuracy with high efficiency. Gu et al. develop a temporal shift reconstruction network for compressive video sensing. To exploit the temporal cues between adjacent frames during the reconstruction of videos, most previous approaches commonly preform alignment between initial reconstructions. However, the estimated motions are usually too coarse to provide accurate temporal information. To remedy this, the proposed network employs stacked temporal shift reconstruction blocks to enhance the initial reconstruction progressively. Within each block, an efficient temporal shift operation is used to capture temporal structures in addition to computational overheads. Then, a bidirectional alignment module is adopted to capture the temporal dependencies in a video sequence. Different from previous methods that only extract supplementary information from the key frames, the proposed alignment module can receive temporal information from the whole video sequence via bidirectional propagations. Experiments demonstrate the superior performance of the proposed method. Qu et al. propose a lightweight video frame interpolation network with a three-scale encoding-decoding structure. Specifically, multi-scale motion information is first extracted from the input video. Then, recurrent convolutional layers are adopted to refine the resultant features. Afterwards, the resultant features are aggregated to generate high-quality interpolated frames. Experimental results on the CelebA and Helen datasets show that the proposed method outperforms state-of-the-art methods while using fewer parameters. Dou et al. introduce a decoder structure-guided CNN-Transformer network for face super-resolution. Most previous approaches follow a multi-task learning paradigm to perform landmark detection while super-resolving the low-resolution images. However, these methods require additional annotation cost, and the extracted facial prior structures are usually of low quality. To address these issues, the proposed network employs a global-local feature extraction unit to extract the global structure while capturing local texture details. In addition, a multi-state fusion module is incorporated to aggregate embeddings from different stages. Experiments show that the proposed method surpasses previous approaches by notable margins. Yang et al. study the problem of blind super-resolution and propose a method to exploit degradation information through degradation representation learning. Specifically, a generative adversarial network is employed to model the degradation process from HR images to LR images and constrain the data distribution of the synthetic LR images. Then, the learnt representation is adopted to super-resolve the input low-resolution images using a transformer-based SR network. Experiments on both synthetic and real-world datasets demonstrate the effectiveness and superiority of the proposed method. Longguang Wang received his BE and PhD degrees from Shandong University and National University of Defence Technology (NUDT) in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 40 peer-reviewed journals and conference publications (including TPAMI, TIP, CVPR, ICCV, and ECCV). He has organised three workshops at CVPR 2022 and 2023. His research interests include low-level vision and 3D vision, particularly on image restoration, image enhancement, image generation, depth estimation, point cloud understanding, and network acceleration. He received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM, and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, and the winner of 2019 ICCV-AIM. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an assistant professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany and at the Technical University of Munich, Munich, Germany. He is currently a lecturer with The University of Tokyo and a unit leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organised by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an associate editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at the University of Wurzburg, Germany. Also, he is a lecturer and a group leader at ETH Zurich, Switzerland. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI, and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of the Alexander von Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organiser of NTIRE, CLIC, AIM, Mobile AI, and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, and image/video compression, restoration, enhancement, and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from NUDT in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modelling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organised several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022, and ECCV 2022. He is a senior member of IEEE and ACM. Data sharing is not applicable to this article as no new data were created or analysed in this study. Longguang Wang received his B.E. and Ph.D. degrees from Shandong University and National University of Defense Technology in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 60 peer reviewed journal and conference publications (including TPAMI, TIP, CVPR, ICCV and ECCV). He served as a reviewer for more than 10 international journals (including TPAMI and TIP) and conferences (including CVPR, ICCV and ECCV). He has organized workshops at CVPR 2022/2023/2024. His research interests include low-level vision and 3D vision, particularly on image restoration, image generation, point cloud understanding, and network acceleration. His received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationalwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, the winner of 2019 ICCV-AIM, etc. Meanwhile, he served as a reviewer for more than 20 international journals and conferences. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an Assistant Professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany, and at the Technical University of Munich, Munich, Germany. He is currently a Lecturer with The University of Tokyo, and a Unit Leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images, with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organized by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an Associate Editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at theUniversity of Wurzburg, Germany. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of an Alexandervon Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organizer of NTIRE, CLIC, AIM, Mobile AI and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, image/video compression, restoration, enhancement and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from National University of Defense Technology (NUDT) in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modeling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organized several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022 and ECCV 2022. He is a Senior Member of IEEE and ACM. Longguang Wang, Juncheng Li 0003, Naoto Yokoya, Radu Timofte, Yulan Guo |
IET Comput. Vis. | 3 |
| 2024 | Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and BenchmarkabstractLearning with limited labeled data is a challenging problem in various applications, including remote sensing. Few-shot semantic segmentation is one approach that can encourage deep learning models to learn from few labeled examples for novel classes not seen during the training. The generalized few-shot segmentation setting has an additional challenge which encourages models not only to adapt to the novel classes but also to maintain strong performance on the training base classes. While previous datasets and benchmarks discussed the few-shot segmentation setting in remote sensing, we are the first to propose a generalized few-shot segmentation benchmark for remote sensing. The generalized setting is more realistic and challenging, which necessitates exploring it within the remote sensing context. We release the dataset augmenting OpenEarthMap (OEM) with additional classes labeled for the generalized few-shot evaluation setting. The dataset is released during the OEM land cover mapping generalized few-shot challenge in the learning with limited labeled data for image and video understanding (L3D-IVU) workshop in conjunction with computer vision and pattern recognition (CVPR) 2024. In this work, we summarize the dataset and challenge details in addition to providing the benchmark results on the two phases of the challenge for the validation and test sets. Clifford Broni-Bediako, Junshi Xia, Jian Song 0010, Hongruixuan Chen, Mennatullah Siam, Naoto Yokoya |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | SpectralGPT: Spectral Remote Sensing Foundation ModelabstractThe foundation model has recently garnered significant attention due to its potential to revolutionize the field of visual representation learning in a self-supervised manner. While most foundation models are tailored to effectively process RGB images for various visual tasks, there is a noticeable gap in research focused on spectral data, which offers valuable information for scene understanding, especially in remote sensing (RS) applications. To fill this gap, we created for the first time a universal RS foundation model, named SpectralGPT, which is purpose-built to handle spectral RS images using a novel 3D generative pretrained transformer (GPT). Compared to existing foundation models, SpectralGPT 1) accommodates input images with varying sizes, resolutions, time series, and regions in a progressive training fashion, enabling full utilization of extensive RS Big Data; 2) leverages 3D token generation for spatial-spectral coupling; 3) captures spectrally sequential patterns via multi-target reconstruction; and 4) trains on one million spectral RS images, yielding models with over 600 million parameters. Our evaluation highlights significant performance improvements with pretrained SpectralGPT models, signifying substantial potential in advancing spectral RS Big Data applications within the field of geoscience across four downstream tasks: single/multi-label scene classification, semantic segmentation, and change detection. Danfeng Hong, Bing Zhang 0001, Chenyu Li 0002, Jing Yao 0002, Naoto Yokoya, Hao Li 0019, Pedram Ghamisi, Xiuping Jia, Antonio Plaza, Paolo Gamba, Jón Atli Benediktsson, Jocelyn Chanussot |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided TransformerabstractOptical high-resolution imagery and OpenStreetMap (OSM) data are two important data sources of land-cover change detection (CD). Previous related studies focus on utilizing the information in OSM data to aid the CD on optical high-resolution images. This article pioneers the direct detection of land-cover changes utilizing paired OSM data and optical imagery, thereby expanding the scope of CD tasks. To this end, we propose an object-guided Transformer (ObjFormer) by naturally combining the object-based image analysis (OBIA) technique with the advanced vision Transformer architecture. This combination can significantly reduce the computational overhead in the self-attention module without adding extra parameters or layers. Specifically, ObjFormer has a hierarchical pseudo-Siamese encoder consisting of object-guided self-attention modules that extract multilevel heterogeneous features from OSM data and optical images; a decoder consisting of object-guided cross-attention modules can recover land-cover changes from the extracted heterogeneous features. Beyond basic binary CD (BCD), this article raises a new semi-supervised semantic CD (SCD) task that does not require any manually annotated land-cover labels to train semantic change detectors. Two lightweight semantic decoders are added to ObjFormer to accomplish this task efficiently. A converse cross-entropy (CCE) loss is designed to fully utilize negative samples, contributing to the great performance improvement in this task. A large-scale benchmark dataset called OpenMapCD containing 1287 map–image pairs covering 40 regions on six continents is constructed to conduct the detailed experiments. The results show the effectiveness of our methods in this new kind of CD task. In addition, case studies in two Japanese cities demonstrate the framework’s generalizability and practical potential. The code and dataset will be open-sourced inhttps://github.com/ChenHongruixuan/ObjFormer. Hongruixuan Chen, Cuiling Lan, Jian Song 0010, Clifford Broni-Bediako, Junshi Xia, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space ModelabstractConvolutional neural networks (CNNs) and Transformers have made impressive progress in the field of remote sensing change detection (CD). However, both architectures have inherent shortcomings: CNN is constrained by a limited receptive field that may hinder their ability to capture broader spatial contexts, while Transformers are computationally intensive, making them costly to train and deploy on large datasets. Recently, the Mamba architecture, based on state space models (SSMs), has shown remarkable performance in a series of natural language processing tasks, which can effectively compensate for the shortcomings of the above two architectures. In this article, we explore for the first time the potential of the Mamba architecture for remote sensing CD tasks. We tailor the corresponding frameworks, called MambaBCD, MambaSCD, and MambaBDA, for binary CD (BCD), semantic CD (SCD), and building damage assessment (BDA), respectively. All three frameworks adopt the cutting-edge Visual Mamba architecture as the encoder, which allows full learning of global spatial contextual information from the input images. For the change decoder, which is available in all three architectures, we propose three spatiotemporal relationship modeling mechanisms, which can be naturally combined with the Mamba architecture and fully utilize its attribute to achieve spatiotemporal interaction of multitemporal features, thereby obtaining accurate change information. On five benchmark datasets, our proposed frameworks outperform current CNN- and Transformer-based approaches without using any complex training strategies or tricks, fully demonstrating the potential of the Mamba architecture in CD tasks. Further experiments show that our architecture is quite robust to degraded data. The source code is available at:https://github.com/ChenHongruixuan/MambaCD. Hongruixuan Chen, Jian Song 0010, Chengxi Han, Junshi Xia, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Frequency-Based Optimal Style Mix for Domain Generalization in Semantic Segmentation of Remote Sensing ImagesabstractSupervised learning methods assume that training and test data are sampled from the same distribution. However, this assumption is not always satisfied in practical situations of land cover semantic segmentation when models trained in a particular source domain are applied to other regions. This is because domain shifts caused by variations in location, time, and sensor alter the distribution of images in the target domain from that of the source domain, resulting in significant degradation of model performance. To mitigate this limitation, domain generalization (DG) has gained attention as a way of generalizing from source domain features to unseen target domains. One approach is style randomization (SR), which enables models to learn domain-invariant features through randomizing styles of images in the source domain. Despite its potential, existing methods face several challenges, such as inflexible frequency decomposition, high computational and data preparation demands, slow speed of randomization, and lack of consistency in learning. To address these limitations, we propose a frequency-based optimal style mix (FOSMix), which consists of three components: 1) full mix (FM) enhances the data space by maximally mixing the style of reference images into the source domain; 2) optimal mix (OM) keeps the essential frequencies for segmentation and randomizes others to promote generalization; and 3) regularization of consistency ensures that the model can stably learn different images with the same semantics. Extensive experiments that require the model’s generalization ability, with domain shift caused by variations in regions and resolutions, demonstrate that the proposed method achieves superior segmentation in remote sensing. The source code is available athttps://github.com/Reo-I/FOSMix. Reo Iizuka, Junshi Xia, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Coupled Tensor Double-Factor Method for Hyperspectral and Multispectral Image FusionabstractHyperspectral and multispectral image fusion, denoted as HSI-MSI fusion, involves merging a pair of hyperspectral (HSI) and multispectral (MSI) images to generate a high spatial resolution hyperspectral image (HR-HSI). The primary challenge in HSI-MSI fusion is to find the best way to extract one-dimensional spectral features and two-dimensional (2-D) spatial features from HSI and MSI and harmoniously combine them. In recent times, coupled tensor decomposition (CTD)-based methods have shown promising performance in the fusion task. However, the tensor decompositions (TDs) used by these CTD-based methods face difficulties in extracting complex features and capturing 2-D spatial features, resulting in suboptimal fusion results. To address these issues, we introduce a novel method called Coupled Tensor Double-Factor Decomposition (CTDF). Specifically, we propose a Tensor Double-Factor (TDF) decomposition, representing a 3rd-order HR-HSI as a 4th-order spatial factor and a 3rd-order spectral factor, connected through tensor contraction. Compared to other TDs, the TDF has better feature extraction capability since it has a higher order factor than that of HR-HSI, whereas the other TDs only have the same order factor as the HR-HSI. Moreover, the TDF can extract 2-D spatial features using the 4th-order spatial factor. We apply the TDF to the HSI-MSI fusion problem and formulate the CTDF model. Furthermore, we design an algorithm based on proximal alternating minimization to solve this model and provide insights into its computational complexity and convergence analysis. The simulated and real experiments validate the effectiveness and efficiency of the proposed CTDF method. The code is available at https://github.com/tingxu113/CTDF. Ting-Zhu Huang, Liang-Jian Deng, Jin-Liang Xiao, Clifford Broni-Bediako, Junshi Xia, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Local-to-Global Cross-Modal Attention-Aware Fusion for HSI-X Semantic SegmentationabstractHyperspectral image (HSI) classification has recently reached its performance bottleneck. Multimodal data fusion is emerging as a promising approach to overcome this bottleneck by providing rich complementary information from the supplementary modality (X-modality). However, achieving comprehensive cross-modal interaction and fusion that can be generalized across different sensing modalities is challenging due to the disparity in imaging sensors, resolution, and content of different modalities. In this study, we propose a local-to-global cross-modal attention-aware fusion (LoGoCAF) framework for HSI-X segmentation. LoGoCAF adopts a two-branch semantic segmentation architecture to learn information from HSI and X modalities. The pipeline of LoGoCAF consists of a local-to-global encoder and a lightweight all multilayer perceptron (ALL-MLP) decoder. In the encoder, convolutions are used to encode local and high-resolution fine details in shallow layers, while transformers are used to integrate global and low-resolution coarse features in deeper layers. The ALL-MLP decoder aggregates information from the encoder for feature fusion and prediction. In particular, two cross-modality modules, the feature enhancement module (FEM) and the feature interaction and fusion module (FIFM), are introduced in each encoder stage. The FEM is used to enhance complementary information by combining information from the other modality across direction-aware, position-sensitive, and channel-wise dimensions. With the enhanced features, the FIFM is designed to promote cross-modality information interaction and fusion for the final semantic prediction. Extensive experiments demonstrate that our LoGoCAF achieves superior performance and generalizes well on various multimodal datasets. Code is available athttps://github.com/xumzhang. Xuming Zhang 0004, Naoto Yokoya, Xingfa Gu, Qingjiu Tian, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | V4D: Voxel for 4D Novel View SynthesisabstractNeural radiance fields have made a remarkable breakthrough in the novel view synthesis task at the 3D static scene. However, for the 4D circumstance (e.g., dynamic scene), the performance of the existing method is still limited by the capacity of the neural network, typically in a multilayer perceptron network (MLP). In this article, we utilize 3D Voxel to model the 4D neural radiance field, short as V4D, where the 3D voxel has two formats. The first one is to regularly model the 3D space and then use the sampled local 3D feature with the time index to model the density field and the texture field by a tiny MLP. The second one is in look-up tables (LUTs) format that is for the pixel-level refinement, where the pseudo-surface produced by the volume rendering is utilized as the guidance information to learn a 2D pixel-level refinement mapping. The proposed LUTs-based refinement module achieves the performance gain with little computational cost and could serve as the plug-and-play module in the novel view synthesis task. Moreover, we propose a more effective conditional positional encoding toward the 4D data that achieves performance gain with negligible computational burdens. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance at a low computational cost. Wanshui Gan, Yi Huang 0035, Shifeng Chen, Naoto Yokoya |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Combining Deep Learning and Numerical Simulation to Predict Flood Inundation DepthabstractCurrent flood mapping methods combine remote sensing and machine learning technologies to estimate the inundation area. Although these methods have shown great success, they mainly focus on the flood extent without additional information on the inundation depth. However, knowing the inundation level can significantly benefit first responders and rescue efforts. Recent advances in machine learning have boosted the development of advanced methods for disaster management. This paper integrates modern convolutional neural network (CNN) models and physics-based numerical simulation to develop a novel framework for automatic flood inundation depth mapping. Our framework builds a synthetic training dataset using numerical flood simulation in four geographical regions. Then, it trains CNN models to understand the nonlinear relationship between inundation depth and topographic features. Our experiments, designed to evaluate the strength of our methodology in a real-world application, demonstrate that it can estimate flood depth with acceptable accuracy (Root-Mean-Squared Error=0.2) in unseen areas during training. These results indicate that a worldwide flood inundation mapping could be achieved by including key areas with representative topographic features. Bruno Adriano, Naoto Yokoya, Kazuki Yamanoi, Satoru Oishi |
IGARSS | 2 |
| 2023 | OpenEarthMap: A Benchmark Dataset for Global High-Resolution Land Cover MappingabstractWe introduce OpenEarthMap, a benchmark dataset, for global high-resolution land cover mapping. OpenEarth-Map consists of 2.2 million segments of 5000 aerial and satellite images covering 97 regions from 44 countries across 6 continents, with manually annotated 8-class land cover labels at a 0.25–0.5m ground sampling distance. Se-mantic segmentation models trained on the OpenEarth-Map generalize worldwide and can be used as off-the-shelf models in a variety of applications. We evaluate the performance of state-of-the-art methods for unsupervised domain adaptation and present challenging problem settings suitable for further technical development. We also investigate lightweight models using automated neural architecture search for limited computational resources and fast mapping. The dataset is available at https: //open-earth-map.org. Junshi Xia, Naoto Yokoya, Bruno Adriano, Clifford Broni-Bediako |
WACV | 2 |
| 2023 | Guest Editorial: Spectral imaging powered computer visionabstractThe increasing accessibility and affordability of spectral imaging technology have revolutionised computer vision, allowing for data capture across various wavelengths beyond the visual spectrum.This advancement has greatly enhanced the capabilities of computers and AI systems in observing, understanding, and interacting with the world.Consequently, new datasets in various modalities, such as infrared, ultraviolet, fluorescent, multispectral, and hyperspectral, have been constructed, presenting fresh opportunities for computer vision research and applications.Although significant progress has been made in processing, learning, and utilising data obtained through spectral imaging technology, several challenges persist in the field of computer vision.These challenges include the presence of low-quality images, sparse input, high-dimensional data, expensive data labelling processes, and a lack of methods to effectively analyse and utilise data considering their unique properties.Many mid-level and high-level computer vision tasks, such as object segmentation, detection and recognition, image retrieval and classification, and video tracking and understanding, still have not leveraged the advantages offered by spectral information.Additionally, the problem of effectively and efficiently fusing data in different modalities to create robust vision systems remains unresolved.Therefore, there is a pressing need for novel computer vision methods and applications to advance this research area.This special issue aims to provide a venue for researchers to present innovative computer vision methods driven by the spectral imaging technology. Jun Zhou 0001, Fengchao Xiong, Naoto Yokoya, Pedram Ghamisi |
IET Comput. Vis. | 4 |
| 2023 | LaMIE: Large-Dimensional Multipass InSAR Phase Estimation for Distributed ScatterersabstractState-of-the-art phase linking (PL) methods for distributed scatterer interferometry (DSI) retrieve consistent phase histories from the sample coherence matrix or the one whose magnitudes are calibrated. To unify them, we first propose a framework consisting of sample coherence matrix estimation and Kullback–Leibler (KL) divergence minimization. Within such framework, we observe that the current state-of-the-art PL methods mainly focus on calibrating the magnitudes of sample coherence matrix while ignoring the errors caused by it exploited in the complex domain, especially when the PL problem is large-dimensional. In this paper, “large-dimensional” refers to the case where the temporal dimensionNof coherence matrices and the numberPof statistically homogeneous pixels (SHP) are at the same level. To solve this issue, we further propose a PL method, termed LaMIE, which is aimed at precise phase history retrieval from large-dimensional coherence matrices for DSI. It includes two steps: 1) sample coherence matrix shrinkage to calibrate the matrix in complex and real domains and 2) phase history retrieval via the flat coherence metric. Both simulated and real data experiments validate the effectiveness of the proposed method by comparing it with other PL methods. Through LaMIE, the densities of the selected points with stable phases can be significantly improved, and the displacement velocities for more regions can be obtained than with state-of-the-art methods. Yusong Bai, Jian Kang 0005, Anping Zhang, Zhe Zhang 0026, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Decoupled-and-Coupled Networks: Self-Supervised Hyperspectral Image Super-Resolution With Subpixel FusionabstractEnormous efforts have been recently made to super-resolve hyperspectral (HS) images with the aid of high spatial resolution multispectral (MS) images. Most prior works usually perform the fusion task by means of multifarious pixel-level priors. Yet the intrinsic effects of a large distribution gap between HS-MS data due to differences in the spatial and spectral resolution are less investigated. The gap might be caused by unknown sensor-specific properties or highly-mixed spectral information within one pixel (due to low spatial resolution). To this end, we propose a subpixel-level HS super-resolution framework by devising a novel decoupled-and-coupled network, called DC-Net, to progressively fuse HS-MS information from the pixel- to subpixel-level, from the image- to feature-level. As the name suggests, DC-Net first decouples the input into common (or cross-sensor) and sensor-specific components to eliminate the gap between HS-MS images before further fusion, and then thoroughly blends them by a model-guided coupled spectral unmixing (CSU) net. More significantly, we append a self-supervised learning module behind the CSU net by guaranteeing material consistency to enhance the detailed appearance of the restored HS product. Extensive experimental results show the superiority of our method both visually and quantitatively and achieve a significant improvement in comparison with the state-of-the-art. Danfeng Hong, Jing Yao 0002, Chenyu Li 0002, Deyu Meng, Naoto Yokoya, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | ES6D: A Computation Efficient and Symmetry-Aware 6D Pose Regression FrameworkabstractIn this paper, a computation efficient regression framework is presented for estimating the 6D pose of rigid objects from a single RGB-D image, which is applicable to handling symmetric objects. This framework is designed in a simple architecture that efficiently extracts point-wise features from RGB-D data using a fully convolutional network, called XYZNet, and directly regresses the 6D pose without any post refinement. In the case of symmetric object, one object has multiple ground-truth poses, and this one-to-many relationship may lead to estimation ambiguity. In order to solve this ambiguity problem, we design a symmetry-invariant pose distance metric, called average (maximum) grouped primitives distance or A(M)GPD. The proposed A(M)GPD loss can make the regression network converge to the correct state, i.e., all minima in the A(M)GPD loss surface are mapped to the correct poses. Extensive experiments on YCB-Video and TLESS datasets demonstrate the proposed framework's substantially superior performance in top accuracy and low computational cost. The relevant code is available in https://github.com/GANWANSHUI/ES6D.git. Ningkai Mo, Wanshui Gan, Naoto Yokoya, Shifeng Chen |
CVPR | 3 |
| 2022 | Learning Mutual Modulation for Self-supervised Cross-Modal Super-Resolution
Naoto Yokoya, Longguang Wang, Tatsumi Uezato |
ECCV (19) | 2 |
| 2022 | Spectrum-Aware and Transferable Architecture Search for Hyperspectral Image Restoration
Wei He 0003, Quanming Yao, Naoto Yokoya, Tatsumi Uezato, Hongyan Zhang 0001, Liangpei Zhang 0001 |
ECCV (19) | 3 |
| 2022 | EOD: The IEEE GRSS Earth Observation DatabaseabstractIn the era of deep learning, annotated datasets have become a crucial asset to the remote sensing community. In the last decade, a plethora of different datasets was published, each designed for a specific data type and with a specific task or application in mind. In the jungle of remote sensing datasets, it can be hard to keep track of what is available already. With this paper, we introduce EOD - the IEEE GRSS Earth Observation Database (EOD) - an interactive online platform for cataloguing different types of datasets leveraging remote sensing imagery. Michael Schmitt 0003, Pedram Ghamisi, Naoto Yokoya, Ronny Hänsch |
IGARSS | 3 |
| 2022 | Non-Local Meets Global: An Iterative Paradigm for Hyperspectral Image RestorationabstractNon-local low-rank tensor approximation has been developed as a state-of-the-art method for hyperspectral image (HSI) restoration, which includes the tasks of denoising, compressed HSI reconstruction and inpainting. Unfortunately, while its restoration performance benefits from more spectral bands, its runtime also substantially increases. In this paper, we claim that the HSI lies in a global spectral low-rank subspace, and the spectral subspaces of each full band patch group should lie in this global low-rank subspace. This motivates us to propose a unified paradigm combining the spatial and spectral properties for HSI restoration. The proposed paradigm enjoys performance superiority from the non-local spatial denoising and light computation complexity from the low-rank orthogonal basis exploration. An efficient alternating minimization algorithm with rank adaptation is developed. It is done by first solving a fidelity term-related problem for the update of a latent input image, and then learning a low-dimensional orthogonal basis and the related reduced image from the latent input image. Subsequently, non-local low-rank denoising is developed to refine the reduced image and orthogonal basis iteratively. Finally, the experiments on HSI denoising, compressed reconstruction, and inpainting tasks, with both simulated and real datasets, demonstrate its superiority with respect to state-of-the-art HSI restoration methods. Wei He 0003, Quanming Yao, Chao Li 0013, Naoto Yokoya, Qibin Zhao, Hongyan Zhang 0001, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Hyperspectral super-resolution via coupled tensor ring factorization
Wei He 0003, Yong Chen 0013, Naoto Yokoya, Chao Li 0013, Qibin Zhao |
Pattern Recognit. | 3 |
| 2022 | Synthesizing Optical and SAR Imagery From Land Cover Maps and Auxiliary Raster DataabstractWe synthesize both optical RGB and synthetic aperture radar (SAR) remote sensing images from land cover maps and auxiliary raster data using generative adversarial networks (GANs). In remote sensing, many types of data, such as digital elevation models (DEMs) or precipitation maps, are often not reflected in land cover maps but still influence image content or structure. Including such data in the synthesis process increases the quality of the generated images and exerts more control on their characteristics. Spatially adaptive normalization layers fuse both inputs and are applied to a full-blown generator architecture consisting of encoder and decoder to take full advantage of the information content in the auxiliary raster data. Our method successfully synthesizes medium (10 m) and high (1 m) resolution images when trained with the corresponding data set. We show the advantage of data fusion of land cover maps and auxiliary information using mean intersection over unions (mIoUs), pixel accuracy, and Fréchet inception distances (FIDs) using pretrained U-Net segmentation models. Handpicked images exemplify how fusing information avoids ambiguities in the synthesized images. By slightly editing the input, our method can be used to synthesize realistic changes, i.e., raising the water levels. The source code is available athttps://github.com/gbaier/rs_img_synth, and we published the newly created high-resolution data set athttps://ieee-dataport.org/open-access/geonrw. Gerald Baier, Antonin Deschemps, Michael Schmitt 0003, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised Multimodal Change Detection Based on Structural Relationship Graph Representation LearningabstractUnsupervised multimodal change detection is a practical and challenging topic that can play an important role in time-sensitive emergency applications. To address the challenge that multimodal remote sensing images cannot be directly compared due to their modal heterogeneity, we take advantage of two types of modality-independent structural relationships in multimodal images. In particular, we present a structural relationship graph representation learning framework for measuring the similarity of the two structural relationships. First, structural graphs are generated from preprocessed multimodal image pairs by means of an object-based image analysis approach. Then, a structural relationship graph convolutional autoencoder (SR-GCAE) is proposed to learn robust and representative features from graphs. Two loss functions aiming at reconstructing vertex information and edge information are presented to make the learned representations applicable for structural relationship similarity measurement. Subsequently, the similarity levels of two structural relationships are calculated from learned graph representations, and two difference images are generated based on the similarity levels. After obtaining the difference images, an adaptive fusion strategy is presented to fuse the two difference images. Finally, a morphological filtering-based postprocessing approach is employed to refine the detection results. Experimental results on six datasets with different modal combinations demonstrate the effectiveness of the proposed method. Hongruixuan Chen, Naoto Yokoya, Chen Wu 0003, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Coherence-Guided Complex Convolutional Sparse Coding for Interferometric Phase RestorationabstractInterferometric phase restoration is a crucial step in retrieving large-scale geophysical parameters from Synthetic Aperture Radar (SAR) images. Existing noise impacts the accuracy of parameter retrieval as a result of decorrelation effects. Most state-of-the-art filtering methods belong to the group of nonlocal filters. In this paper, we propose a novel convolutional sparse coding method in complex domain with the prior knowledge of coherence integrated into the optimization model, which is termed as CoComCSC. CoComCSC is not only capable of reducing noise in regions with continuous phase changes, but also of preserving the phase details prominently. The experiments results on simulated and real data demonstrate the effectiveness of CoComCSC by comparing with other state-of-the-art methods. Moreover, the obtained Digital Elevation Model (DEM) product by CoComCSC from RADARSAT-2 data indicates its superior filtering performance over regions with heterogeneous land-covers, which shows its great potential for generating high-resolution DEM products. Jian Kang 0005, Zhe Zhang 0026, Yan Huang 0018, Jialin Liu 0003, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Masked Auto-Encoding Spectral-Spatial Transformer for Hyperspectral Image ClassificationabstractDeep learning has certainly become the dominant trend in hyper-spectral (HS) remote sensing image classification owing to its excellent capabilities to extract highly discriminating spatial-spectral features. In this context, transformer networks have recently shown prominent results in distinguishing even the most subtle spectral differences because of their potential to characterize sequential spectral data. Nonetheless, many complexities affecting HS remote sensing data (e.g. atmospheric effects, thermal noise, quantization noise, etc.) may severely undermine such potential since no mode of relieving noisy feature patterns has still been developed within transformer networks. To address the problem, this paper presents a novel masked auto-encoding spectral-spatial transformer (MAEST), which gathers two different collaborative branches: (i) a reconstruction path, which dynamically uncovers the most robust encoding features based on a masking auto-encoding strategy; and (ii) a classification path, which embeds these features onto a transformer network to classify the data focusing on the features that better reconstruct the input. Unlike other existing models, this novel design pursues to learn refined transformer features considering the aforementioned complexities of the HS remote sensing image domain. The experimental comparison, including several state-of-the-art methods and benchmark datasets, shows the superior results obtained by MAEST. The codes of this paper will be available at https://github.com/ibanezfd/MAEST. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised and Unregistered Hyperspectral Image Super-Resolution With Mutual Dirichlet-NetabstractHyperspectral images (HSIs) provide rich spectral information that has contributed to the successful performance improvement of numerous computer vision and remote sensing tasks. However, it can only be achieved at the expense of images’ spatial resolution. HSI super-resolution (HSI-SR), thus, addresses this problem by fusing low-resolution (LR) HSI with the multispectral image (MSI) carrying much higher spatial resolution (HR). Existing HSI-SR approaches require the LR HSI and HR MSI to be well registered, and the reconstruction accuracy of the HR HSI relies heavily on the registration accuracy of different modalities. In this article, we propose an unregistered and unsupervised mutual Dirichlet-Net ($u^{2}$-MDN) to exploit the uncharted problem domain of HSI-SRwithout the requirement of multimodality registration. The success of this endeavor would largely facilitate the deployment of HSI-SR since registration requirement is difficult to satisfy in real-world sensing devices. The novelty of this work is threefold. First, to stabilize the fusion procedure of two unregistered modalities, the network is designed to extract spatial information and spectral information of two modalities with different dimensions through a shared encoder–decoder structure. Second, the mutual information (MI) is further adopted to capture the nonlinear statistical dependencies between the representations from two modalities (carrying spatial information) and their raw inputs. By maximizing the MI, spatial correlations between different modalities can be well characterized to further reduce the spectral distortion. We assume that the representations follow a similar Dirichlet distribution for their inherent sum-to-one and nonnegative properties. Third, a collaborative$l_{2,1}$-norm is employed as the reconstruction error instead of the more common$l_{2}$-norm to better preserve the spectral information. Extensive experimental results demonstrate the superior performance of$u^{2}$-MDN as compared to the state of the art. Ying Qu 0001, Hairong Qi 0001, Chiman Kwan, Naoto Yokoya, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DML: Differ-Modality Learning for Building Semantic SegmentationabstractThis work critically analyzes the problems arising from differ-modality building semantic segmentation in the remote sensing domain. With the growth of multimodality datasets, such as optical, synthetic aperture radar (SAR), light detection and ranging (LiDAR), and the scarcity of semantic knowledge, the task of learning multimodality information has increasingly become relevant over the last few years. However, multimodality datasets cannot be obtained simultaneously due to many factors. Assume that we have SAR images with reference information in one place and optical images without reference in another; how to learn relevant features of optical images from SAR images? We refer to it as differ-modality learning (DML). To solve the DML problem, we propose novel deep neural network architectures, which include image adaptation, feature adaptation, knowledge distillation, and self-training (SL) modules for different scenarios. We test the proposed methods on the differ-modality remote sensing datasets (very high-resolution SAR and RGB from SpaceNet 6) to build semantic segmentation and to achieve a superior efficiency. The presented approach achieves the best performance when compared with the state-of-the-art methods. Junshi Xia, Naoto Yokoya, Gerald Baier |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | An Iterative Regularization Method Based on Tensor Subspace Representation for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (HSI-SR) can be achieved by fusing a paired multispectral image (MSI) and hyperspectral image (HSI), which is a prevalent strategy. But, how to precisely reconstruct the high spatial resolution hyperspectral image (HR-HSI) by fusion technology is a challenging issue. In this paper, we propose an iterative regularization method based on tensor subspace representation (IR-TenSR) for MSI-HSI fusion, thus HSI-SR. First, we propose a tensor subspace representation (TenSR)-based regularization model that integrates the global spectral-spatial low-rank and the nonlocal self-similarity priors of HR-HSI. These two priors have been proven effective, but previous HSI-SR works cannot simultaneously exploit them. Subsequently, we design an iterative regularization procedure to utilize the residual information of acquired low-resolution images, which are ignored in other works that produce suboptimal results. Finally, we develop an effective algorithm based on the proximal alternating minimization method to solve the TenSR-regularization model. With that, we obtain the iterative regularization algorithm. Experiments implemented on the simulated and real datasets illustrate the advantages of the proposed IR-TenSR compared with state-of-the-art fusion approaches. The code is available at https://github.com/liangjiandeng/IR-TenSR. Ting-Zhu Huang, Liang-Jian Deng, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Breaking Limits of Remote Sensing by Deep Learning From Simulated Data for Flood and Debris-Flow MappingabstractWe propose a framework that estimates the inundation depth (maximum water level) and debris-flow-induced topographic deformation from remote sensing imagery by integrating deep learning and numerical simulation. A water and debris-flow simulator generates training data for various artificial disaster scenarios. We show that regression models based on Attention U-Net and LinkNet architectures trained on such synthetic data can predict the maximum water level and topographic deformation from a remote sensing-derived change detection map and a digital elevation model. The proposed framework has an inpainting capability, thus mitigating the false negatives that are inevitable in remote sensing image analysis. Our framework breaks limits of remote sensing and enables rapid estimation of inundation depth and topographic deformation, essential information for emergency response, including rescue and relief activities. We conduct experiments with both synthetic and real data for two disaster events that caused simultaneous flooding and debris flows and demonstrate the effectiveness of our approach quantitatively and qualitatively. Our code and data sets are available athttps://github.com/nyokoya/dlsim. Naoto Yokoya, Kazuki Yamanoi, Wei He 0003, Gerald Baier, Bruno Adriano, Hiroyuki Miura, Satoru Oishi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Endmember-Guided Unmixing Network (EGU-Net): A General Deep Learning Framework for Self-Supervised Hyperspectral UnmixingabstractOver the past decades, enormous efforts have been made to improve the performance of linear or nonlinear mixing models for hyperspectral unmixing (HU), yet their ability to simultaneously generalize various spectral variabilities (SVs) and extract physically meaningful endmembers still remains limited due to the poor ability in data fitting and reconstruction and the sensitivity to various SVs. Inspired by the powerful learning ability of deep learning (DL), we attempt to develop a general DL approach for HU, by fully considering the properties of endmembers extracted from the hyperspectral imagery, called endmember-guided unmixing network (EGU-Net). Beyond the alone autoencoder-like architecture, EGU-Net is a two-stream Siamese deep network, which learns an additional network from the pure or nearly pure endmembers to correct the weights of another unmixing network by sharing network parameters and adding spectrally meaningful constraints (e.g., nonnegativity and sum-to-one) toward a more accurate and interpretable unmixing solution. Furthermore, the resulting general framework is not only limited to pixelwise spectral unmixing but also applicable to spatial information modeling with convolutional operators for spatial-spectral unmixing. Experimental results conducted on three different datasets with the ground truth of abundance maps corresponding to each material demonstrate the effectiveness and superiority of the EGU-Net over state-of-the-art unmixing algorithms. The codes will be available from the website: https://github.com/danfenghong/IEEE_TNNLS_EGU-Net. Danfeng Hong, Lianru Gao, Jing Yao 0002, Naoto Yokoya, Jocelyn Chanussot, Uta Heiden, Bing Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | U-IMG2DSM: Unpaired Simulation of Digital Surface Models With Generative Adversarial NetworksabstractHigh-resolution digital surface models (DSMs) provide valuable height information about the Earth's surface, which can be successfully combined with other types of remotely sensed data in a wide range of applications. However, the acquisition of DSMs with high spatial resolution is extremely time-consuming and expensive with their estimation from a single optical image being an ill-possed problem. To overcome these limitations, this letter presents a new unpaired approach to obtain DSMs from optical images using deep learning techniques. Specifically, our new deep neural model is based on variational autoencoders (VAEs) and generative adversarial networks (GANs) to perform image-to-image translation, obtaining DSMs from optical images. Our newly proposed method has been tested in terms of photographic interpretation, reconstruction error, and classification accuracy using three well-known remotely sensed data sets with very high spatial resolution (obtained over Potsdam, Vaihingen, and Stockholm). Our experimental results demonstrate that the proposed approach obtains satisfactory reconstruction rates that allow enhancing the classification results for these images. The source code of our method is available from: https://github.com/mhaut/UIMG2DSM. Mercedes Eugenia Paoletti, Juan Mario Haut, Pedram Ghamisi, Naoto Yokoya, Javier Plaza, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Joint and Progressive Subspace Analysis (JPSA) With Spatial-Spectral Manifold Alignment for Semisupervised Hyperspectral Dimensionality ReductionabstractConventional nonlinear subspace learning techniques (e.g., manifold learning) usually introduce some drawbacks in explainability (explicit mapping) and cost effectiveness (linearization), generalization capability (out-of-sample), and representability (spatial-spectral discrimination). To overcome these shortcomings, a novel linearized subspace analysis technique with spatial-spectral manifold alignment is developed for a semisupervised hyperspectral dimensionality reduction (HDR), called joint and progressive subspace analysis (JPSA). The JPSA learns a high-level, semantically meaningful, joint spatial-spectral feature representation from hyperspectral (HS) data by: 1) jointly learning latent subspaces and a linear classifier to find an effective projection direction favorable for classification; 2) progressively searching several intermediate states of subspaces to approach an optimal mapping from the original space to a potential more discriminative subspace; and 3) spatially and spectrally aligning a manifold structure in each learned latent subspace in order to preserve the same or similar topological property between the compressed data and the original data. A simple but effective classifier, that is, nearest neighbor (NN), is explored as a potential application for validating the algorithm performance of different HDR approaches. Extensive experiments are conducted to demonstrate the superiority and effectiveness of the proposed JPSA on two widely used HS datasets: 1) Indian Pines (92.98%) and 2) the University of Houston (86.09%) in comparison with previous state-of-the-art HDR methods. The demo of this basic work (i.e., ECCV2018) is openly available at https://github.com/danfenghong/ECCV2018_J-Play. Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Jian Xu 0008, Xiao Xiang Zhu 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | More Diverse Means Better: Multimodal Deep Learning Meets Remote-Sensing Imagery ClassificationabstractClassification and identification of the materials lying over or beneath the earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS), and have garnered a growing concern owing to the recent advancements of deep learning techniques. Although deep networks have been successfully applied in single-modality-dominated classification tasks, yet their performance inevitably meets the bottleneck in complex scenes that need to be finely classified, due to the limitation of information diversity. In this work, we provide a baseline solution to the aforementioned difficulty by developing a general multimodal deep learning (MDL) framework. In particular, we also investigate a special case of multi-modality learning (MML)-cross-modality learning (CML) that exists widely in RS image classification applications. By focusing on “what,” “where,” and “how” to fuse, we show different fusion strategies as well as how to train deep networks and build the network architecture. Specifically, five fusion architectures are introduced and developed, further being unified in our MDL framework. More significantly, our framework is not only limited to pixel-wise classification tasks but also applicable to spatial information modeling with convolutional neural networks (CNNs). To validate the effectiveness and superiority of the MDL framework, extensive experiments related to the settings of MML and CML are conducted on two different multimodal RS data sets. Furthermore, the codes and data sets will be available at https://github.com/danfenghong/IEEE_TGRS_MDL-RS, contributing to the RS community. Danfeng Hong, Lianru Gao, Naoto Yokoya, Jing Yao 0002, Jocelyn Chanussot, Qian Du 0001, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Graph-Induced Aligned Learning on Subspaces for Hyperspectral and Multispectral DataabstractIn this article, we have great interest in investigating a common but practical issue in remote sensing (RS)-can a limited amount of one information-rich (or high-quality) data, e.g., hyperspectral (HS) image, improve the performance of a classification task using a large amount of another information-poor (low-quality) data, e.g., multispectral (MS) image? This question leads to a typical cross-modality feature learning. However, classic cross-modality representation learning approaches, e.g., manifold alignment, remain limited in effectively and efficiently handling such problems that the data from high-quality modality are largely absent. For this reason, we propose a novel graph-induced aligned learning (GiAL) framework by 1) adaptively learning a unified graph (further yielding a Laplacian matrix) from the data in order to align multimodality data (MS-HS data) into a latent shared subspace; 2) simultaneously modeling two regression behaviors with respect to labels and pseudo-labels under a multitask learning paradigm; and 3) dramatically updating the pseudo-labels according to the learned graph and refeeding the latest pseudo-labels into model learning of the next round. In addition, an optimization framework based on the alternating direction method of multipliers (ADMMs) is devised to solve the proposed GiAL model. Extensive experiments are conducted on two MS-HS RS data sets, demonstrating the superiority of the proposed GiAL compared with several state-of-the-art methods. Danfeng Hong, Jian Kang 0005, Naoto Yokoya, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Fast Hyperspectral Image Recovery of Dual-Camera Compressive Hyperspectral Imaging via Non-Iterative Subspace-Based FusionabstractCoded aperture snapshot spectral imaging (CASSI) is a promising technique for capturing three-dimensional hyperspectral images (HSIs), in which algorithms are used to perform the inverse problem of HSI reconstruction from a single coded two-dimensional (2D) measurement. Due to the ill-posed nature of this problem, various regularizers have been exploited to reconstruct 3D data from 2D measurements. Unfortunately, the accuracy and computational complexity are unsatisfactory. One feasible solution is to utilize additional information such as the RGB measurement in CASSI. Considering the combined CASSI and RGB measurements, in this paper, we propose a fusion model for HSI reconstruction. Specifically, we investigate the low-dimensional spectral subspace property of HSIs composed of a spectral basis and spatial coefficients. In particular, the RGB measurement is utilized to estimate the coefficients, while the CASSI measurement is adopted to provide the spectral basis. We further propose a patch processing strategy to enhance the spectral low-rank property of HSIs. The optimization of the proposed model requires neither iteration nor the spectral sensing matrix of the RGB detector. Extensive experiments on both simulated and real HSI datasets demonstrate that our proposed method not only outperforms previous state-of-the-art (iterative algorithms) methods in quality but also speeds up the reconstruction by more than 5000 times. Wei He 0003, Naoto Yokoya, Xin Yuan 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning Convolutional Sparse Coding on Complex Domain for Interferometric Phase RestorationabstractInterferometric phase restoration has been investigated for decades and most of the state-of-the-art methods have achieved promising performances for InSAR phase restoration. These methods generally follow the nonlocal filtering processing chain, aiming at circumventing the staircase effect and preserving the details of phase variations. In this article, we propose an alternative approach for InSAR phase restoration, that is, Complex Convolutional Sparse Coding (ComCSC) and its gradient regularized version. To the best of the authors' knowledge, this is the first time that we solve the InSAR phase restoration problem in a deconvolutional fashion. The proposed methods can not only suppress interferometric phase noise, but also avoid the staircase effect and preserve the details. Furthermore, they provide an insight into the elementary phase components for the interferometric phases. The experimental results on synthetic and realistic high- and medium-resolution data sets from TerraSAR-X StripMap and Sentinel-1 interferometric wide swath mode, respectively, show that our method outperforms those previous state-of-the-art methods based on nonlocal InSAR filters, particularly the state-of-the-art method: InSAR-BM3D. The source code of this article will be made publicly available for reproducible research inside the community. Jian Kang 0005, Danfeng Hong, Jialin Liu 0003, Gerald Baier, Naoto Yokoya, Begüm Demir |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Guided Deep Decoder: Unsupervised Image Pair Fusion
Tatsumi Uezato, Danfeng Hong, Naoto Yokoya, Wei He 0003 |
ECCV (6) | 3 |
| 2020 | Damage Characterization in Urban Environments from Multitemporal Remote Sensing Datasets Built from Previous EventsabstractDisasters such as earthquakes, hurricanes, and flooding are responsible for large-scale infrastructure damages and loss of human lives. Immediately after disaster strikes, one of the most critical and difficult tasks is accurately assessing the extent and severity of the disaster. This task is especially challenging in areas isolated by the disaster; in such cases, remote sensing information provides the best alternative to tackle this problem. This paper presents a damage mapping framework using remote sensing imagery acquired from previous disasters. The proposed deep learning-based framework is trained to learn features related to building damage using imagery from previous disasters that were collected from different regions around the world. Then, it is tested to recognize damage from a different urban environment. Bruno Adriano, Junshi Xia, Naoto Yokoya, Hiroyuki Miura, Masashi Matsuoka, Shunichi Koshimura |
IGARSS | 3 |
| 2020 | Learning-Shared Cross-Modality Representation Using Multispectral-LiDAR and Hyperspectral DataabstractDue to the ever-growing diversity of the data source, multimodality feature learning has attracted more and more attention. However, most of these methods are designed by jointly learning feature representation from multimodalities that exist in both training and test sets, yet they are less investigated in the absence of certain modality in the test phase. To this end, in this letter, we propose to learn a shared feature space across multimodalities in the training process. By this way, the out-of-sample from any of multimodalities can be directly projected onto the learned space for a more effective cross-modality representation. More significantly, the shared space is regarded as a latent subspace in our proposed method, which connects the original multimodal samples with label information to further improve the feature discrimination. Experiments are conducted on the multispectral-Light Detection and Ranging (LIDAR) and hyperspectral data set provided by the 2018 IEEE GRSS Data Fusion Contest to demonstrate the effectiveness and superiority of the proposed method in comparison with several popular baselines. Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Jian Kang 0005, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Hyperspectral Image Restoration Using Weighted Group Sparsity-Regularized Low-Rank Tensor DecompositionabstractMixed noise (such as Gaussian, impulse, stripe, and deadline noises) contamination is a common phenomenon in hyperspectral imagery (HSI), greatly degrading visual quality and affecting subsequent processing accuracy. By encoding sparse prior to the spatial or spectral difference images, total variation (TV) regularization is an efficient tool for removing the noises. However, the previous TV term cannot maintain the shared group sparsity pattern of the spatial difference images of different spectral bands. To address this issue, this article proposes a group sparsity regularization of the spatial difference images for HSI restoration. Instead of using ℓ1or ℓ2-norm (sparsity) on the difference image itself, we introduce a weighted ℓ2,1norm to constrain the spatial difference image cube, efficiently exploring the shared group sparse pattern. Moreover, we employ the well-known low-rank Tucker decomposition to capture the global spatial-spectral correlation from three HSI dimensions. To summarize, a weighted group sparsity-regularized low-rank tensor decomposition (LRTDGS) method is presented for HSI restoration. An efficient augmented Lagrange multiplier algorithm is employed to solve the LRTDGS model. The superiority of this method for HSI restoration is demonstrated by a series of experimental results from both simulated and real data, as compared with the other state-of-the-art TV-regularized low-rank matrix/tensor decomposition methods. Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang |
IEEE Trans. Cybern. | 3 |
| 2020 | Robust Nonlocal Low-Rank SAR Time Series Despeckling Considering Speckle Correlation by Total Variation RegularizationabstractOutliers and speckle both corrupt time series of synthetic aperture radar (SAR) acquisitions. Owing to the coherence between SAR acquisitions, their speckle can no longer be regarded as independent. In this study, we propose an algorithm for nonlocal low-rank time series despeckling, which is robust against outliers and also specifically addresses speckle correlation between acquisitions. By imposing total variation regularization on the signal's speckle component, the correlation between acquisitions can be identified, facilitating the extraction of outliers from unfiltered signals and the correlated speckle. This robustness against outliers also addresses matching errors and inaccuracies in the nonlocal similarity search. Such errors include mismatched data in the nonlocal estimation process, which degrade the denoising performance of conventional similarity-based filtering approaches. Multiple experiments on real and synthetic data assess the performance of the approach by comparing it with state-of-the-art methods. It provides filtering results of comparable quality but is not adversely affected by outliers. The source code is available at https://github.com/gbaier/nllrtv. Gerald Baier, Wei He 0003, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Nonlocal Tensor-Ring Decomposition for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is a fundamental problem in remote sensing and image processing. Recently, nonlocal low-rank tensor approximation-based denoising methods have attracted much attention due to their advantage of being capable of fully exploiting the nonlocal self-similarity and global spectral correlation. Existing nonlocal low-rank tensor approximation methods were mainly based on two common decomposition [Tucker or CANDECOMP/PARAFAC (CP)] methods and achieved the state-of-the-art results, but they are subject to certain issues and do not produce the best approximation for a tensor. For example, the number of parameters for Tucker decomposition increases exponentially according to its dimensions, and CP decomposition cannot better preserve the intrinsic correlation of the HSI. In this article, a novel nonlocal tensor-ring (TR) approximation is proposed for HSI denoising by using TR decomposition to explore the nonlocal self-similarity and global spectral correlation simultaneously. TR decomposition approximates a high-order tensor as a sequence of cyclically contracted third-order tensors, which has strong ability to explore these two intrinsic priors and to improve the HSI denoising results. Moreover, an efficient proximal alternating minimization algorithm is developed to optimize the proposed TR decomposition model efficiently. Extensive experiments on three simulated data sets under several noise levels and two real data sets verify that the proposed TR model provides better HSI denoising results than several state-of-the-art methods in terms of quantitative and visual performance evaluations. Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang, Xi-Le Zhao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Invariant Attribute Profiles: A Spatial-Frequency Joint Feature Extractor for Hyperspectral Image ClassificationabstractSo far, a large number of advanced techniques have been developed to enhance and extract the spatially semantic information in hyperspectral image processing and analysis. However, locally semantic change, such as scene composition, relative position between objects, spectral variability caused by illumination, atmospheric effects, and material mixture, has been less frequently investigated in modeling spatial information. Consequently, identifying the same materials from spatially different scenes or positions can be difficult. In this article, we propose a solution to address this issue by locally extracting invariant features from hyperspectral imagery (HSI) in both spatial and frequency domains, using a method called invariant attribute profiles (IAPs). IAPs extract the spatial invariant features by exploiting isotropic filter banks or convolutional kernels on HSI and spatial aggregation techniques (e.g., superpixel segmentation) in the Cartesian coordinate system. Furthermore, they model invariant behaviors (e.g., shift, rotation) by the means of a continuous histogram of oriented gradients constructed in a Fourier polar coordinate. This yields a combinatorial representation of spatial-frequency invariant features with application to HSI classification. Extensive experiments conducted on three promising hyperspectral data sets (Houston2013 and Houston2018) to demonstrate the superiority and effectiveness of the proposed IAP method in comparison with several state-of-the-art profile-related techniques. The codes will be available from the website: https://sites.google.com/view/danfeng-hong/data-code. Danfeng Hong, Xin Wu 0001, Pedram Ghamisi, Jocelyn Chanussot, Naoto Yokoya, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Hyperspectral Image Compressive Sensing Reconstruction Using Subspace-Based Nonlocal Tensor Ring DecompositionabstractHyperspectral image compressive sensing reconstruction (HSI-CSR) can largely reduce the high expense and low efficiency of transmitting HSI to ground stations by storing a few compressive measurements, but how to precisely reconstruct the HSI from a few compressive measurements is a challenging issue. It has been proven that considering the global spectral correlation, spatial structure, and nonlocal self-similarity priors of HSI can achieve satisfactory reconstruction performances. However, most of the existing methods cannot simultaneously capture the mentioned priors and directly design the regularization term to the HSI. In this article, we propose a novel subspace-based nonlocal tensor ring decomposition method (SNLTR) for HSI-CSR. Instead of designing the regularization of the low-rank approximation to the HSI, we assume that the HSI lies in a low-dimensional subspace. Moreover, to explore the nonlocal self-similarity and preserve the spatial structure of HSI, we introduce a nonlocal tensor ring decomposition strategy to constrain the related coefficient image, which can decrease the computational cost compared to the methods that directly employ the nonlocal regularization to HSI. Finally, a well-known alternating minimization method is designed to efficiently solve the proposed SNLTR. Extensive experimental results demonstrate that our SNLTR method can significantly outperform existing approaches for HSI-CSR. Yong Chen 0013, Ting-Zhu Huang, Wei He 0003, Naoto Yokoya, Xi-Le Zhao |
IEEE Trans. Image Process. | 4 |
| 2020 | Illumination Invariant Hyperspectral Image Unmixing Based on a Digital Surface ModelabstractAlthough many spectral unmixing models have been developed to address spectral variability caused by variable incident illuminations, the mechanism of the spectral variability is still unclear. This paper proposes an unmixing model, named illumination invariant spectral unmixing (IISU). IISU makes the first attempt to use the radiance hyperspectral data and a LiDAR-derived digital surface model (DSM) in order to physically explain variable illuminations and shadows in the unmixing framework. Incident angles, sky factors, visibility from the sun derived from the LiDAR-derived DSM support the explicit explanation of endmember variability in the unmixing process from radiance perspective. The proposed model was efficiently solved by a straightforward optimization procedure. The unmixing results showed that the other state-of-the-art unmixing models did not work well especially in the shaded pixels. On the other hand, the proposed model estimated more accurate abundances and shadow compensated reflectance than the existing models. Tatsumi Uezato, Naoto Yokoya, Wei He 0003 |
IEEE Trans. Image Process. | 2 |
| 2019 | Non-Local Meets Global: An Integrated Paradigm for Hyperspectral DenoisingabstractNon-local low-rank tensor approximation has been developed as a state-of-the-art method for hyperspectral image (HSI) denoising. Unfortunately, while their denoising performance benefits little from more spectral bands, the running time of these methods significantly increases. In this paper, we claim that the HSI lies in a global spectral low-rank subspace, and the spectral subspaces of each full band patch groups should lie in this global low-rank subspace. This motivates us to propose a unified spatial-spectral paradigm for HSI denoising. As the new model is hard to optimize, An efficient algorithm motivated by alternating minimization is developed. This is done by first learning a low-dimensional orthogonal basis and the related reduced image from the noisy HSI. Then, the non-local low-rank denoising and iterative regularization are developed to refine the reduced image and orthogonal basis, respectively. Finally, the experiments on synthetic and both real datasets demonstrate the superiority against the stateof-the-art HSI denoising methods. Wei He 0003, Quanming Yao, Chao Li 0013, Naoto Yokoya, Qibin Zhao |
CVPR | 4 |
| 2019 | Total-variation-regularized Tensor Ring Completion for Remote Sensing Image ReconstructionabstractIn recent studies, tensor ring (TR) decomposition has shown to be effective in data compression and representation. However, the existing TR-based completion methods only exploit the global low-rank property of the visual data. When applying them to remote sensing (RS) image processing, the spatial information in the RS image is ignored. In this paper, we introduce the TR decomposition to RS image processing and propose a tensor completion method for RS image reconstruction. We incorporate the total-variation regularization into the TR completion model to exploit the low-rank property and spatial continuity of the RS image simultaneously. The proposed algorithm is solved by the augmented Lagrange multiplier method and has shown the superior performance in hyperspectral image reconstruction and multi-temporal RS image cloud removal against the state-of-the-art algorithms. Wei He 0003, Longhao Yuan, Naoto Yokoya |
ICASSP | 3 |
| 2019 | Weighted Group Sparsity Regularized Low-Rank Tensor Decomposition for Hyperspectral Image RestorationabstractTotal variation (TV) regularization has been widely used in the hyperspectral image (HSI) mixed noise removal problem by utilizing ℓ1-norm to constrain the spatial difference image and promote the piecewise smooth structure. Unfortunately, it cannot depict the group sparse structure of spatial difference image along the spectral dimension. This paper proposes a new HSI restoration method using weighted group sparsity regularized low-rank tensor decomposition (LRTDGS). Specifically, we use a weighted group sparsity regularization which is denoted by ℓ2,1-norm to explore the group structure of spatial difference image along the spectral dimension. Moreover, the spatial-spectral correlation from three directions of HSI is depicted by low-rank Tucker decomposition. We use efficient augmented Lagrange multiplier method to optimize the proposed LRTDGS model, and a series of experimental results are presented to demonstrate the effectiveness of the proposed method. Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang |
IGARSS | 3 |
| 2019 | Total Variation Regularized Low-Rank Sparsity Decomposition for Blind Cloud and Cloud Shadow Removal from Multitemporal ImageryabstractThis paper proposes a spatial-spectral total variation (TV) regularized low-rank sparsity decomposition model for blind cloud and cloud shadow (cloud/shadow) detection and removal of multitemporal remote sensing imagery. Our concept is to decompose the contaminated image into the surface-reflected component and the cloud/shadow component. Low-rank regularization is utilized to model the spectral-temporal correlation of the surface-reflected component, meanwhile, the `1-norm and spatial-spectral total variation regularization is employed to describe the sparse prior and spatial-spectral continuity of the cloud/shadow component. To better preserve the information in cloud/shadow-free areas, the cloud/shadow detection results obtained as a by-product of our method are used to guide the information compensation from the original contaminated images. Several experiments are presented to demonstrate the effectiveness of the proposed method. Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang |
IGARSS | 3 |
| 2019 | Cross-Domain-Classification of Tsunami Damage Via Data Simulation and Residual-Network-Derived Features From Multi-Source ImagesabstractThis paper presents a novel application of remote sensing data and machine learning technologies for damage classification in a real-world cross-domain application. The proposed methodology trains models to learn the building damage characteristics recorded in the 2011 Tohoku Tsunami from multi-sensor and multi-temporal remote sensing images. Then, the trained models are tested in the recent 2018 Sulawesi Tsunami. Additionally, a simulation of high-resolution SAR image was carried to deal with missing data modality. Our initial results show that the ResNet-derived features from optical images acquired after the disaster together with moderate- and high-resolution synthetic aperture radar (SAR) post-event intensity data showed significant accuracy in classifying two levels of tsunami-induced damage, with an average f-score of approximately 0.72. Taking into account that no training data from the 2018 Sulawesi Tsunami was used, our methodology shows excellent potential for future implementation of a rapid response system based on a database of building damage constructed from previous majors disasters. Bruno Adriano, Naoto Yokoya, Junshi Xia, Gerald Baier, Shunichi Koshimura |
IGARSS | 2 |
| 2019 | Robust Nonlocal Low-Rank Sar Stack Despeckling With Application To Change DetectionabstractWe present a nonlocal low-rank denoising algorithm for synthetic aperture radar (SAR) image stacks. The method extends the widely known DespecKS algorithm by integrating low-rank approximation, outlier removal, and total variation (TV) regularization into the estimation process. Preliminary experiments shows increased robustness against outliers and comparable performance to state-of-the-art stack despeckling algorithms. Gerald Baier, Wei He 0003, Bruno Adriano, Junshi Xia, Naoto Yokoya |
IGARSS | 5 |
| 2019 | WU-Net: A Weakly-Supervised Unmixing Network for Remotely Sensed Hyperspectral ImageryabstractRecently, enormous efforts have been made to improve the performance of the linear or nonlinear mixing model for hyperspectral unmixing, yet their ability to handle spectral variability and extract physically meaningful endmembers remains limited. Based on the powerful learning ability of deep learning, we propose a weakly-supervised unmixing network, called WU-Net, to break the bottleneck. Beyond the autoencoder-like architecture, WU-Net learns an additional network from the pure or nearly-pure endmembers to correct the weights of another unmixing network towards a more accurate and interpretable unmixing solution, thus yielding a two-stream deep network. Experimental results conducted on two different datasets, one fully artificial simulation dataset and one simulated EnMap dataset generated from a real HyMap dataset, demonstrate the effectiveness and superiority of WU-Net over several state-of-the-art algorithms. Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Uta Heiden, Wieke Heldens, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2019 | Mangrove Species Mapping Using Sentinel-1 and Sentinel-2 Data in North VietnamabstractThis study employed Sentinel-1A C-band and Sentinel-2A multispectral data combined with the decision tree ensemble algorithms to map the spatial distribution of five mangrove communities in a coastal area in North Vietnam. The results show that the rotation forests (RoFs) model achieved better overall accuracy and kappa coefficient in mapping mangrove species than those of the canonical correlation forests (CCFs) and the random forests (RFs) models. This research demonstrates the potential of using optical and SAR data together with machine learning techniques to map mangrove species in tropical areas. Tien Dat Pham 0001, Junshi Xia, Gerald Baier, Nga Nhu Le, Naoto Yokoya |
IGARSS | 5 |
| 2019 | Building Damage Mapping Via Transfer LearningabstractThis paper presents building damage mapping based on transfer learning techniques. Due to the different spatial resolutions of optical (WorldView, 0.5m) and SAR (Sentinel-1, 10m), we adopt different methods: pixel-level for moderate-resolution SAR images, and patch-level for very high-resolution optical images. For SAR images, the performance of fast unsupervised transfer learning methods, such as overall centroid alignment (OCA) and CORrelation ALignment (CORAL), are investigated. For the optical images, two public databases are used to predict the building damage mapping of Palu with WorldView-3 images via ResNet50. Experimental results indicate the effectiveness of transfer learning on the building damage mapping using different data sources. Junshi Xia, Bruno Adriano, Gerald Baier, Naoto Yokoya |
IGARSS | 4 |
| 2019 | Land Cover Mapping without Human AnnotationabstractThe advent of satellite constellations has been rapidly increasing the acquisition frequency of Earth observation images, enabling continuous updates of land cover maps on a global scale with supervised learning. Since annotation of land cover semantic classes is expensive with a field survey, the number of labeled samples and the update frequency are limited. To compensate for the lack of ground-truth labels, we propose to substitute automatically labeled ground-shot images for human annotated data. In our methodology, unlabeled georeferenced ground-shot images are classified with a Convolutional Neural Network (CNN), which is trained by a set of ground-shot images annotated by humans in advance. We propose a new method that automatically grades human annotation to address noisy labels used for training the CNN. The outcomes of ground-shot image classification are used as labeled samples for pixel-wise land cover classification of multispectral satellite imagery. The resulted land cover classification map resembles the results supervised by human-annotated samples. The experimental results show that our methodology has the potential to associate satellite imagery with ground-shot imagery. Tatsuya Yamada, Naoto Yokoya, Takeo Tadono, Akira Iwasaki |
IGARSS | 2 |
| 2019 | Remote Sensing Image Reconstruction Using Tensor Ring Completion and Total VariationabstractTime-series remote sensing (RS) images are often corrupted by various types of missing information such as dead pixels, clouds, and cloud shadows that significantly influence the subsequent applications. In this paper, we introduce a new low-rank tensor decomposition model, termed tensor ring (TR) decomposition, to the analysis of RS data sets and propose a TR completion method for the missing information reconstruction. The proposed TR completion model has the ability to utilize the low-rank property of time-series RS images from different dimensions. To further explore the smoothness of the RS image spatial information, total-variation regularization is also incorporated into the TR completion model. The proposed model is efficiently solved using two algorithms, the augmented Lagrange multiplier (ALM) and the alternating least square (ALS) methods. The simulated and real-data experiments show superior performance compared to other state-of-the-art low-rank related algorithms. Wei He 0003, Naoto Yokoya, Longhao Yuan, Qibin Zhao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | CoSpace: Common Subspace Learning From Hyperspectral-Multispectral CorrespondencesabstractWith a large amount of open satellite multispectral (MS) imagery (e.g., Sentinel-2 and Landsat-8), considerable attention has been paid to global MS land cover classification. However, its limited spectral information hinders further improving the classification performance. Hyperspectral imaging enables discrimination between spectrally similar classes but its swath width from space is narrow compared to MS ones. To achieve accurate land cover classification over a large coverage, we propose a cross-modality feature learning framework, called common subspace learning (CoSpace), by jointly considering subspace learning and supervised classification. By locally aligning the manifold structure of the two modalities, CoSpace linearly learns a shared latent subspace from hyperspectral-MS (HS-MS) correspondences. The MS out-of-samples can be then projected into the subspace, which are expected to take advantages of rich spectral information of the corresponding hyperspectral data used for learning, and thus leads to a better classification. Extensive experiments on two simulated HS-MS data sets (University of Houston and Chikusei), where HS-MS data sets have tradeoffs between coverage and spectral resolution, are performed to demonstrate the superiority and effectiveness of the proposed method in comparison with previous state-of-the-art methods. Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | An Augmented Linear Mixing Model to Address Spectral Variability for Hyperspectral UnmixingabstractHyperspectral imagery collected from airborne or satellite sources inevitably suffers from spectral variability, making it difficult for spectral unmixing to accurately estimate abundance maps. The classical unmixing model, the linear mixing model (LMM), generally fails to handle this sticky issue effectively. To this end, we propose a novel spectral mixture model, called the augmented linear mixing model (ALMM), to address spectral variability by applying a data-driven learning strategy in inverse problems of hyperspectral unmixing. The proposed approach models the main spectral variability (i.e., scaling factors) generated by variations in illumination or typography separately by means of the endmember dictionary. It then models other spectral variabilities caused by environmental conditions (e.g., local temperature and humidity, atmospheric effects) and instrumental configurations (e.g., sensor noise), as well as material nonlinear mixing effects, by introducing a spectral variability dictionary. To effectively run the data-driven learning strategy, we also propose a reasonable prior knowledge for the spectral variability dictionary, whose atoms are assumed to be low-coherent with spectral signatures of endmembers, which leads to a well-known low-coherence dictionary learning problem. Thus, a dictionary learning technique is embedded in the framework of spectral unmixing so that the algorithm can learn the spectral variability dictionary and estimate the abundance maps simultaneously. Extensive experiments on synthetic and real datasets are performed to demonstrate the superiority and effectiveness of the proposed method in comparison with previous state-of-the-art methods. Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Joint and Progressive Learning from High-Dimensional Data for Multi-label Classification
Danfeng Hong, Naoto Yokoya, Jian Xu 0008, Xiao Xiang Zhu 0001 |
ECCV (8) | 2 |
| 2018 | A Comparative Study of Fusion-Based Change Detection Methods for Multi-Band Images with Different Spectral and Spatial ResolutionsabstractThis paper deals with a fusion-based change detection (CD) framework for multi-band images with different spatial and spectral resolutions. The first step of the considered CD framework consists in fusing the two observed images. The resulting fused image is subsequently spatially or spectrally degraded to produce two pseudo-observed images, with the same resolutions as the two observed images. Finally, CD can be performed through a pixel-wise comparison of the pseudo-observed and observed images since they share the same resolutions. Obviously, fusion is a key step in this framework. Thus, this paper proposes to quantitatively and qualitatively compare state-of-the-art fusion methods, gathered into four main families, namely component substitution, multi-resolution analysis, unmixing and Bayesian, with respect to the performance of the whole CD framework evaluated on simulated and real images. Vinicius Ferraris, Naoto Yokoya, Nicolas Dobigeon, Marie Chabert |
IGARSS | 2 |
| 2018 | Boosting for Domain Adaptation Extreme Learning Machines for Hyperspectral Image ClassificationabstractDomain adaptation and transfer learning adapt the priori information of source domain to train a classier used to predict the label in the target domain. The parameter and instance transfer methods have shown excellent performance. The former adjusts the parameters of transitional classifiers and the latter re-weights the training sample to the different training set, which is similar to the AdaBoost. To further improve the performance, we proposed to combine the two techniques mentioned above. More specifically, we select the Transfer Boosting and domain adaptation extreme learning machine (DAELM) as the instance and parameter transfer methods, respectively. We refer the proposed method to the boosting for DAELM (BDAELM). We compare the proposed method with DAELM and other methods on the real cross-domain hyperspectral remote sensing images acquired over a Japanese mixed forest, showing improved classification accuracies. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
IGARSS | 2 |
| 2018 | IMG2DSM: Height Simulation From Single Imagery Using Conditional Generative Adversarial NetabstractThis letter proposes a groundbreaking approach in the remote-sensing community to simulating the digital surface model (DSM) from a single optical image. This novel technique uses conditional generative adversarial networks whose architecture is based on an encoder-decoder network with skip connections (generator) and penalizing structures at the scale of image patches (discriminator). The network is trained on scenes where both the DSM and optical data are available to establish an image-to-DSM translation rule. The trained network is then utilized to simulate elevation information on target scenes where no corresponding elevation information exists. The capability of the approach is evaluated both visually (in terms of photographic interpretation) and quantitatively (in terms of reconstruction errors and classification accuracies) on subdecimeter spatial resolution data sets captured over Vaihingen, Potsdam, and Stockholm. The results confirm the promising performance of the proposed framework. Pedram Ghamisi, Naoto Yokoya |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Fusion of Hyperspectral and LiDAR Data With a Novel Ensemble ClassifierabstractDue to the development of sensors and data acquisition technology, the fusion of features from multiple sensors is a very hot topic. In this letter, the use of morphological features to fuse a hyperspectral (HS) image and a light detection and ranging (LiDAR)-derived digital surface model (DSM) is exploited via an ensemble classifier. In each iteration, we first apply morphological openings and closings with a partial reconstruction on the first few principal components (PCs) of the HS and LiDAR data sets to produce morphological features to model spatial and elevation information for HS and LiDAR data sets. Second, three groups of features (i.e., spectral and morphological features of HS and LiDAR data) are split into several disjoint subsets. Third, data transformation is applied to each subset and the features extracted in each subset are stacked as the input of a random forest classifier. Three data transformation methods, including PC analysis, linearity preserving projection, and unsupervised graph fusion, are introduced into the ensemble classification process. Finally, we integrate the classification results achieved at each step by a majority vote. Experimental results on coregistered HS and LiDAR-derived DSM demonstrate the effectiveness and potentialities of the proposed ensemble classifier. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Nonlocal Tensor Completion for Multitemporal Remotely Sensed Images' InpaintingabstractRemotely sensed images may contain some missing areas because of poor weather conditions and sensor failure. Information of those areas may play an important role in the interpretation of multitemporal remotely sensed data. This paper aims at reconstructing the missing information by a nonlocal low-rank tensor completion method. First, nonlocal correlations in the spatial domain are taken into account by searching and grouping similar image patches in a large search window. Then, low rankness of the identified fourth-order tensor groups is promoted to consider their correlations in spatial, spectral, and temporal domains, while reconstructing the underlying patterns. Experimental results on simulated and real data demonstrate that the proposed method is effective both qualitatively and quantitatively. In addition, the proposed method is computationally efficient compared with other patch-based methods such as the recently proposed patch matching-based multitemporal group sparse representation method. Teng-Yu Ji, Naoto Yokoya, Xiao Xiang Zhu 0001, Ting-Zhu Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Random Forest Ensembles and Extended Multiextinction Profiles for Hyperspectral Image ClassificationabstractClassification techniques for hyperspectral images based on random forest (RF) ensembles and extended multiextinction profiles (EMEPs) are proposed as a means of improving performance. To this end, five strategies - bagging, boosting, random subspace, rotation-based, and boosted rotation-based - are used to construct the RF ensembles. EPs, which are based on an extrema-oriented connected filtering technique, are applied to the images associated with the first informative components extracted by independent component analysis, leading to a set of EMEPs. The effectiveness of the proposed method is investigated on two benchmark hyperspectral images: the University of Pavia and Indian Pines. Comparative experimental evaluations reveal the superior performance of the proposed methods, especially those employing rotation-based and boosted rotation-based approaches. An additional advantage is that the CPU processing time is acceptable. Junshi Xia, Pedram Ghamisi, Naoto Yokoya, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | A novel ensemble classifier of hyperspectral and LiDAR data using morphological featuresabstractDue to the benefits and limitation of different remote sensing sensors, fusion of the features from multiple sensors, such as hyperspectral and light detection and ranging (LiDAR) is an effective method for land cover mapping. In this paper, we propose a novel ensemble classifier to fuse hyperspectral and LiDAR datasets for classification. First, morphological features are used to model spatial and elevation information from the first few principal components (PCs) of the original hyperspetcral (HS) image and LiDAR data. Second, we split different kinds of features (i.e., spectral bands, morphological features of hyperspectral and LiDAR), into several disjoint subsets and apply the data transformation method to each subset. In particular, three data transformation methods, including principal component analysis (PCA), linearity preserving projection (LPP) and unsupervised graph fusion (UGF) are considered. Third, the features extracted in each subset are concatenated to classify by a random forest (RF) classifier. Experimental results on a co-registered HS and LiDAR data provide the effectiveness and potentialities of the proposed ensemble classifier. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
ICASSP | 2 |
| 2017 | Learning a low-coherence dictionary to address spectral variability for hyperspectral unmixingabstractThis paper presents a novel spectral mixture model to address spectral variability in inverse problems of hyperspectral unmixing. Based on the linear mixture model (LMM), our model introduces a spectral variability dictionary to account for any residuals that cannot be explained by the LMM. Atoms in the dictionary are assumed to be low-coherent with spectral signatures of endmembers. A dictionary learning technique is proposed to learn the spectral variability dictionary while solving unmixing problems simultaneously. Experimental results on synthetic and real datasets demonstrate that the performance of the proposed method is superior to state-of-the-art methods. Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001 |
ICIP | 2 |
| 2017 | Hyperspectral image classification with partial least square forestabstractIn the hyperspectral remote sensing community, decision forests combine the predictions of multiple decision trees (DTs) to achieve better prediction performance. Two well-known and powerful decision forests are Random Forest (RF) and Rotation Forest (RoF). In this work, a novel decision forest, called Partial Least Square Forest (PLSF), is proposed. In the PLSF, we adapt PLS to obtain the components for the hyperplane splitting. Moreover, the projection bootstrap technique is used to retain the full spectral bands for the selection of split in the projected space. Experimental results on three hyperspectral datasets indicated the effectiveness of the proposed PLSF because it enhances the diversity and accuracy within the ensemble when compared to RF and RoF. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
IGARSS | 2 |
| 2017 | Ensemble of transfer component analysis for domain adaptation in hyperspectral remote sensing image classificationabstractIn this work, we address the problem of unsupervised domain transfer learning via an ensemble strategy in the context of classification between multiple hyperspectral images. The objective of domain adaption is to assign the label to an image of interest (the target image) using the labeled samples in the source image. The proposed method is based on the rotation-based ensemble and transfer component analysis (TCA). In this method, the feature space in both source and target image is divided into several disjoint feature subsets. Then, the features induced by the TCA technique in the source domain are used as the input space to a random forest (RF) classifier. Finally, the results achieved by each step are fused by a majority vote. We compare the proposed method, ensemble of TCA (E-TCA), with a regular RF and an RF with the reduced features by the TCA. Experiments on the real hyperspectral image acquired over a Japanese mixed forest show remarkable cross-image classification performances. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
IGARSS | 2 |
| 2017 | Multimodal, multitemporal, and multisource global data fusion for local climate zones classification based on ensemble learningabstractThis paper presents a new methodology for classification of local climate zones based on ensemble learning techniques. Landsat-8 data and open street map data are used to extract spectral-spatial features, including spectral reflectance, spectral indexes, and morphological profiles fed to subsequent classification methods as inputs. Canonical correlation forests and rotation forests are used for the classification step. The final classification map is generated by majority voting on different classification maps obtained by the two classifiers using multiple training subsets. The proposed method achieved an overall accuracy of 74.94% and a kappa coefficient of 0.71 in the 2017 IEEE GRSS Data Fusion Contest. Naoto Yokoya, Pedram Ghamisi, Junshi Xia |
IGARSS | 1 |
| 2017 | Hyperspectral Image Classification With Canonical Correlation ForestsabstractMultiple classifier systems or ensemble learning is an effective tool for providing accurate classification results of hyperspectral remote sensing images. Two well-known ensemble learning classifiers for hyperspectral data are random forest (RF) and rotation forest (RoF). In this paper, we proposed to use a novel decision tree (DT) ensemble method, namely, canonical correlation forest (CCF). More specifically, several individual canonical correlation trees (CCTs) that are binary DTs, which use canonical correlation components for the hyperplane splitting, are used to construct the CCF. Additionally, we adopt the projection bootstrap technique in CCF, in which the full spectral bands are retained for split selection in the projected space. The techniques aforementioned allow the CCF to improve the accuracy of member classifiers and diversity within the ensemble. Furthermore, the CCF is extended to the spectral-spatial frameworks that incorporate Markov random fields, extended multiattribute profiles (EMAPs), and the ensemble of independent component analysis and rolling guidance filter (E-ICA-RGF). Experimental results on six hyperspectral data sets are used to indicate the comparative effectiveness of the proposed method, in terms of accuracy and computational complexity, compared with RF and RoF, and it turns out that CCF is a promising approach for hyperspectral image classification not only with spectral information but also in the spectral-spatial frameworks. Junshi Xia, Naoto Yokoya, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Multisensor Coupled Spectral Unmixing for Time-Series AnalysisabstractWe present a new framework, called multisensor coupled spectral unmixing (MuCSUn), that solves unmixing problems involving a set of multisensor time-series spectral images in order to understand dynamic changes of the surface at a subpixel scale. The proposed methodology couples multiple unmixing problems based on regularization on graphs between the time-series data to obtain robust and stable unmixing solutions beyond data modalities due to different sensor characteristics and the effects of nonoptimal atmospheric correction. Atmospheric normalization and cross calibration of spectral response functions are integrated into the framework as a preprocessing step. The proposed methodology is quantitatively validated using a synthetic data set that includes seasonal and trend changes on the surface and the residuals of nonoptimal atmospheric correction. The experiments on the synthetic data set clearly demonstrate the efficacy of MuCSUn and the importance of the preprocessing step. We further apply our methodology to a real time-series data set composed of 11 Hyperion and 22 Landsat-8 images taken over Fukushima, Japan, from 2011 to 2015. The proposed methodology successfully obtains robust and stable unmixing results and clearly visualizes class-specific changes at a subpixel scale in the considered study area. Naoto Yokoya, Xiao Xiang Zhu 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Local manifold learning with robust neighbors selection for hyperspectral dimensionality reductionabstractManifold learning has been successfully applied to hyperspectral dimensionality reduction to embed nonlinear and nonconvex manifolds in the data. However, dimensionality reduction by manifold learning is sensitive to non-uniform data distribution and the selection of neighbors. To address the two issues to some extents, in this work a new manifold framework based on locality linear embedding (LLE), namely local normalization and local feature selection (LNLFS), is proposed. Classification is explored as a potential application to validate the proposed algorithm. Classification accuracy using data obtained using different dimensionality reduction methods is evaluated and compared, while applying two kinds of strategies for selecting the training and test samples: random sampling and region-based sampling. Experimental results show the classification accuracy obtained with LNLFS is superior to state-of-the-art dimensionality reduction methods. Danfeng Hong, Naoto Yokoya, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2016 | Hyperspectral Super-Resolution of Locally Low Rank Images From Complementary Multisource DataabstractRemote sensing hyperspectral images (HSIs) are quite often low rank, in the sense that the data belong to a low dimensional subspace/manifold. This has been recently exploited for the fusion of low spatial resolution HSI with high spatial resolution multispectral images in order to obtain super-resolution HSI. Most approaches adopt an unmixing or a matrix factorization perspective. The derived methods have led to state-of-the-art results when the spectral information lies in a low-dimensional subspace/manifold. However, if the subspace/manifold dimensionality spanned by the complete data set is large, i.e., larger than the number of multispectral bands, the performance of these methods mainly decreases because the underlying sparse regression problem is severely ill-posed. In this paper, we propose a local approach to cope with this difficulty. Fundamentally, we exploit the fact that real world HSIs are locally low rank, that is, pixels acquired from a given spatial neighborhood span a very low-dimensional subspace/manifold, i.e., lower or equal than the number of multispectral bands. Thus, we propose to partition the image into patches and solve the data fusion problem independently for each patch. This way, in each patch the subspace/manifold dimensionality is low enough, such that the problem is not ill-posed anymore. We propose two alternative approaches to define the hyperspectral super-resolution through local dictionary learning using endmember induction algorithms. We also explore two alternatives to define the local regions, using sliding windows and binary partition trees. The effectiveness of the proposed approaches is illustrated with synthetic and semi real data. Miguel Angel Veganzones, Miguel Simões, Giorgio Licciardi, Naoto Yokoya, José M. Bioucas-Dias, Jocelyn Chanussot |
IEEE Trans. Image Process. | 4 |
| 2015 | Detector anomaly detection and stripe correction of hyperspectral dataabstractSince SWIR hyperspectral data provides useful information on minerals and soils, the applications to geology, mining and agriculture are expected. In the SWIR region, hyperspectral sensors use mercury cadmium telluride detectors that have more defects and show non-uniform responses to input radiances. Therefore, the pushbroom scanning of hyperspectral sensors causes stripe noises that disturb the minute spectral analysis. This work correct the artifact due to stripe noise at first using the conventional correction method, such as momentum matching. The remaining stripe patterns are corrected using a sparse coding technique, which is examined using real data. Daisuke Niina, Naoto Yokoya, Akira Iwasaki |
IGARSS | 2 |
| 2015 | Generalized-hough-transform object detection using class-specific sparse representation for local-feature detectionabstractWe present a method for object detection based on sparse representations and Hough voting, which integrates sparse representations for local-feature detection into generalized-Hough-transform object detection. Object parts are detected via class-specific sparse image representations of patches using learned target and background dictionaries, and their cooccurrence is spatially integrated by Hough voting, which enables object detection. In this paper, a discriminative criterion is introduced into dictionary construction to improve the detection performance. Experiments performed on airplane detection and the identification of a specific ship show that the proposed method achieves state-of-the-art performance with the robustness against noise and occlusion using a small set of positive training samples. Naoto Yokoya, Akira Iwasaki |
IGARSS | 1 |
| 2014 | Facial alignment by using sparse initialization and random forestabstractOver the last decade, face alignment researches have been advancing rapidly and have vast applications related to face recognition, pose estimation, and human-robot interaction. Face alignment is typically performed in a two-stage fashion by alternatively using local landmark detector and global shape regularizer. While both landmark detector and shape regularizer have achieved impressive progress recently, shape initialization with mean shape remains a critical issue. In this paper, we present a unique sparse initialization method that is inspired by the sparse model. Experiment results with two datasets selected from the widely used Multi-PIE Database show that our method outperforms conventional initialization techniques. In addition to its capability to handle test data with high shape variability and potential occlusion, our method has merit of simplicity and can be easily integrated with other face alignment approaches. Chun Fui Liew, Naoto Yokoya, Takehisa Yairi |
ICIP | 2 |
| 2014 | Spectral unmixing of fluorescence fingerprint imagery for visualization of constituents in pie pastryabstractIn this work, we present a new method that combines fluorescence fingerprint (FF) imaging and spectral unmixing to visualize microstructures in food. The method is applied to visualization of three constituents, gluten, starch, and butter, in two types of pie pastry. It is challenging to discriminate between starch and butter because both of them can be represented by similar FFs of low intensities. Two optimization approaches of FF unmixing that consider qualitative knowledge are presented and validated by comparison to the conventional staining method. Although starch and butter were represented by very similar FFs, a constrained-least-squares method with abundance quantization successfully visualized the distributions of constituents in pie pastry. Naoto Yokoya, Mito Kokawa, Junichi Sugiyama |
ICIP | 1 |
| 2014 | Object localization based on sparse representation for remote sensing imageryabstractIn this paper, we propose a new object localization method named sparse representation based object localization (SROL), which is based on the generalized Hough-transform-based approach using sparse representations for parts detection. The proposed method was applied to car and ship detection in remote sensing images and its performance was compared to those of state-of-the-art methods. Experimental results showed that the SROL algorithm can accurately localize categorical objects or a specific object using a small size of training data. Naoto Yokoya, Akira Iwasaki |
IGARSS | 1 |
| 2014 | Airborne unmixing-based hyperspectral super-resolution using RGB imageryabstractThis paper presents an airborne experiment on unmixing-based hyperspectral super-resolution using RGB imagery. Preprocessing is described to ensure spatial and spectral consistency between hyperspectral and RGB images. An extended version of coupled nonnegative matrix factorization (CNMF) is introduced for multisensor hyperspectral super-resolution to deal with a challenging problem setting, i.e., only three spectral channels for higher spatial information and a 10-fold difference of ground sampling distance. The proposed method successfully estimated the high-spatial-resolution red-edge image. Numerical evaluation by comparing the high-spatial-resolution hyperspectral image to ground-measured spectra demonstrated recovery of pure-pixel spectra by the proposed method. Naoto Yokoya, Akira Iwasaki |
IGARSS | 1 |
| 2014 | Nonlinear Unmixing of Hyperspectral Data Using Semi-Nonnegative Matrix FactorizationabstractNonlinear spectral mixture models have recently received particular attention in hyperspectral image processing. In this paper, we present a novel optimization method of nonlinear unmixing based on a generalized bilinear model (GBM), which considers the second-order scattering of photons in a spectral mixture model. Semi-nonnegative matrix factorization (semi-NMF) is used for the optimization to process a whole image in matrix form. When endmember spectra are given, the optimization of abundance and interaction abundance fractions converge to a local optimum by alternating update rules with simple implementation. The proposed method is evaluated using synthetic datasets considering its robustness for the accuracy of endmember extraction and spectral complexity, and shows smaller errors in abundance fractions rather than conventional methods. GBM-based unmixing using semi-NMF is applied to the analysis of an airborne hyperspectral image taken over an agricultural field with many endmembers, and it visualizes the impact of a nonlinear interaction on abundance maps at reasonable computational cost. Naoto Yokoya, Jocelyn Chanussot, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Hyperspectral and multispectral data fusion mission on hyperspectral imager suite (HISUI)abstractHyperspectral imager suite (HISUI) is the Japanese next-generation earth-observing sensor composed of hyperspectral and multispectral imagers. Unmixing-based fusion of hyperspectral and multispectral data enables the production of high-spatial-resolution hyperspectral data. HISUI simulated imaging system combining two imagers was developed for verification experiments to investigate the feasibility and clarify the whole procedure of the hyperspectral and multispectral data fusion mission on HISUI. Airborne experiments are planned as simulation tests of HISUI higher-order products. The experimental results of the ground based observation showed the importance of the preprocessing and cross-calibration on the final quality of fused data, which contributes to the practical use of hyperspectral and multispectral data fusion. Naoto Yokoya, Akira Iwasaki |
IGARSS | 1 |
| 2012 | Similarity measure for spatial-spectral registration in hyperspectral eraabstractIn the hyperspectral era, the demand for data registration in spectral region is an important issue in addition to spatial region. Detection of smile and keystone phenomena that are caused by aberrations in spectrometer is related to registration activity, which is crucial for data fusion research. Hyperspectral Imager Suite (HISUI) is a next-generation Japanese optical sensor that is composed of a hyperspectral imager and a multispectral imager, which will be launched on Advanced Land Observation Satellite 3 (ALOS-3). Three similarity measures, normalized cross correlation (NCC), phase correlation (PC) and mutual information (MI), for spatial-spectral registration of hyperspectral data are discussed for Level-1 data processing of HISUI. Akira Iwasaki, Naoto Yokoya, Takeshi Arai, Norihide Miyamura |
IGARSS | 2 |
| 2012 | Generalized bilinear model based nonlinear unmixing using semi-nonnegative matrix factorizationabstractNonlinear spectral mixing models have recently been receiving attention in hyperspectral image processing. This work presents a novel optimization method for nonlinear unmixing based on a generalized bilinear model (GBM), which considers second-order scattering effects. Semi-nonnegative matrix factorization is used for optimization to process a whole image in a matrix form. The proposed method is applied to an airborne hyperspectral image with many endmembers and shows good performance both in unmixing quality and computational cost with simple implementation. The effect of endmember extraction on nonlinear unmixing is investigated and the impact of the nonlinearity on abundance maps is demonstrated. Naoto Yokoya, Jocelyn Chanussot, Akira Iwasaki |
IGARSS | 1 |
| 2012 | Coupled Nonnegative Matrix Factorization Unmixing for Hyperspectral and Multispectral Data FusionabstractCoupled nonnegative matrix factorization (CNMF) unmixing is proposed for the fusion of low-spatial-resolution hyperspectral and high-spatial-resolution multispectral data to produce fused data with high spatial and spectral resolutions. Both hyperspectral and multispectral data are alternately unmixed into end member and abundance matrices by the CNMF algorithm based on a linear spectral mixture model. Sensor observation models that relate the two data are built into the initialization matrix of each NMF unmixing procedure. This algorithm is physically straightforward and easy to implement owing to its simple update rules. Simulations with various image data sets demonstrate that the CNMF algorithm can produce high-quality fused data both in terms of spatial and spectral domains, which contributes to the accurate identification and classification of materials observed at a high spatial resolution. Naoto Yokoya, Takehisa Yairi, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | Coupled non-negative matrix factorization (CNMF) for hyperspectral and multispectral data fusion: Application to pasture classificationabstractCoupled non-negative matrix factorization (CNMF) is introduced for hyperspectral and multispectral data fusion. The CNMF fused data have little spectral distortion while enhancing spatial resolution of all hyperspectral band images owing to its unmixing based algorithm. CNMF is applied to the synthetic dataset generated from real airborne hyperspectral data taken over pasture area. The spectral quality of fused data is evaluated by the classification accuracy of pasture types. The experiment result shows that CNMF enables accurate identification and classification of observed materials at fine spatial resolution. Naoto Yokoya, Takehisa Yairi, Akira Iwasaki |
IGARSS | 1 |
| 2010 | Challenge of aster digital elevation modelabstractAccuracy of digital elevation model (DEM) obtained by the Advanced Spaceborne Thermal Emission and Reflection Radiometer (ASTER) that has along-track stereovision is investigated. The pointing offset and stability of the radiometer is one cause of the geometric deviation of the ASTER DEM attached with orthorectified image. The correction methodology to be implemented to the data processing is suggested. A fine-tuning of image matching procedure leads to better reproduction of the topography. The comparison with reference DEM is described. Akira Iwasaki, Masaru Koga, Hiroto Kanno, Naoto Yokoya, Tetsuya Okuda, Kojiro Saito |
IGARSS | 4 |
| 2010 | Detection and correction of spectral and spatial misregistrations for hyperspectral dataabstractHyperspectral imaging sensors suffer from spectral and spatial misregistrations. These artifacts prevent the accurate acquisition of the spectra and thus reduce classification accuracy. The main objective of this work is to detect and correct spectral and spatial misregistrations of hyperspectral images. The Hyperion visible near-infrared (VNIR) subsystem is used as an example. An image registration method using normalized cross-correlation for characteristic lines in spectrum image demonstrates its effectiveness for detection of the spectral and spatial misregistrations. Cubic spline interpolation using estimated properties makes it possible to modify the spectral signatures. The accuracy of the proposed postlaunch estimation of the Hyperion properties has been proven to be comparable to that of the prelaunch measurements, which enables the precise onboard calibration of hyperspectral sensors. Naoto Yokoya, Norihide Miyamura, Akira Iwasaki |
IGARSS | 1 |