Hongxiang Jiang

dblp:235/9020 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Orthogonal Subspace Representation for Generative Adversarial Networks
abstract
Disentanglement learning aims to separate explanatory factors of variation so that different attributes of the data can be well characterized and isolated, which promotes efficient inference for downstream tasks. Mainstream disentanglement approaches based on generative adversarial networks (GANs) learn interpretable data representation. However, most typical GAN-based works lack the discussion of the latent subspace, causing insufficient consideration of the variation of independent factors. Although some recent research analyzes the latent space on pretrained GANs for image editing, they do not emphasize learning representation directly from the subspace perspective. Appropriate subspace properties could facilitate corresponding feature representation learning to satisfy the independent variation requirements of the obtained explanatory factors, which is crucial for better disentanglement. In this work, we propose a unified framework for ensuring disentanglement, which fully investigates latent subspace learning (SL) in GAN. The novel GAN-based architecture explores orthogonal subspace representation (OSR) on vanilla GAN, named OSRGAN. To guide a subspace with strong correlation, less redundancy, and robust distinguishability, our OSR includes three stages, self-latent-aware, orthogonal subspace-aware, and structure representation-aware, respectively. First, the self-latent-aware stage promotes the latent subspace strongly correlated with the data space to discover interpretable factors, but with poor independence of variation. Second, the following orthogonal subspace-aware stage adaptively learns some 1-D linear subspace spanned by a set of orthogonal bases in the latent space. There is less redundancy between them, expressing the corresponding independence. Third, the structure representation-aware stage aligns the projection on the orthogonal subspace and the latent variables. Accordingly, feature representation in each linear subspace can be distinguishable, enhancing the independent expression of interpretable factors. In addition, we design an alternating optimization step, achieving a tradeoff training of OSRGAN on different properties. Despite it strictly constrains orthogonality, the loss weight coefficient of distinguishability induced by orthogonality could be adjusted and balanced with correlation constraint. To elucidate, this tradeoff training prevents our OSRGAN from overemphasizing any property and damaging the expressiveness of the feature representation. It takes into account both interpretable factors and their independent variation characteristics. Meanwhile, alternating optimization could keep the cost and efficiency of forward inference unchanged and will not burden the computational complexity. In theory, we clarify the significance of OSR, which brings better independence of factors, along with interpretability as correlation could converge to a high range faster. Moreover, through the convergence behavior analysis, including the objective functions under different constraints and the evaluation curve with iterations, our model demonstrates enhanced stability and definitely converges toward a higher peak for disentanglement. To depict the performance in downstream tasks, we compared the state-of-the-art GAN-based and even VAE-based approaches on different datasets. Our OSRGAN achieves higher disentanglement scores on FactorVAE, SAP, MIG, and VP metrics. All the experimental results illustrate that our novel GAN-based framework has considerable advantages on disentanglement.
Hongxiang Jiang, Xiaoyan Luo, Jihao Yin, Huazhu Fu, Fuxiang Wang
IEEE Trans. Neural Networks Learn. Syst.1
2024 Prototype-Guided Structural Learning from Visual Foundation Model for Few-Shot Aerial Image Semantic Segmentation
abstract
Few-shot aerial image semantic segmentation aims to segment query images with few annotated support samples. It is challenging due to intra-class variations and complex object details in remote aeiral images. However, these two issues are inadequately addressed in existing few-shot segmentation methods. In this paper, we propose a novel Prototype-Guided structural learning (PGSL) framework based on recently proposed segment anything model (SAM). Specifically, to accommodate intra-class variation in aerial image, a novel Prototype-Guided transformer is designed to interact the multiple prototypes from support images with query images, yielding initial segmentation map. Moreover, to improve the performance on object contours, we propose a refine branch based on the SAM, which adopts initial segmentation maps as prompt. This integrates the structural knowledge inherent in SAM into our model. Experiment on iSAID-5i dataset demonstrates the proposed PGSL framework outperforms other state-of-the-art methods.
Qixiong Wang, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IGARSS2
2024 Emergence of A Novel Domain Expert: A Generative AI-based Framework for Software Function Point Analysis
abstract
Estimating software functional size is a crucial initial step before development, impacting costs and timelines. This involves applying standard Function Point Analysis (FPA) to the Software Requirements Specification (SRS). However, manual analysis by Function Point (FP) analysts during the splitting of FP entries from SRS remains inefficient and costly. To address this issue, for the first time, we propose an AI-based domain expert for FPA, named FPA-EX. It employs a large language model (LLM), intelligently extracts software FP entries from SRS, providing automated support to enhance efficiency. Specifically, we construct a multi-domain FPA dataset through collecting and annotating 778 question-answer pairs related to various SRS. Based on this dataset, we present a novel densely supervised fine-tuning (DSFT) on LLM, which performs entries-level optimization over the human augmented text, ensuring precise FPs outputs. Finally, we design a ConceptAct Promting (CAP) process for correct logical reasoning. Experiments demonstrate the superior performance of FPA-EX, particularly higher than GPT3.5 by 0.491 on F1 scores. Furthermore, in practical application, FPA-EX significantly enhances the productivity of FP analysts, contributing to a shift towards more intelligent work patterns.
Hongxiang Jiang
ASE2
2024 S2JO: Spatial-Spectral Joint Optimization for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) annotation often suffers from noisy labels, which brings challenges in classification models training. Existing methods typically address this issue through two separate stages: noisy labels cleaning and robust model design. However, model performance is constrained by insufficient noisy label filtering. To explore an end-to-end joint optimization framework for HSI classification with noisy labels, we propose a spatial–spectral joint optimization (S2JO) network. The S2JO consists of a spatial nonuniform sampling (SNS) module and a spectral prototypes learning (SPL) module. The SNS module filters noisy labels dynamically by picking out the top-N confident points, purifying the input samples for the network. Meanwhile, the SPL module uses multicenter spectral prototypes to extract discriminate features accurately even if the input contains some noisy labels. Extensive experiments on the C2Seg-Beijing HSI subdataset demonstrate the superiority of the proposed S2JO over other state-of-the-art methods.
Jiaqi Feng 0001, Qixiong Wang, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.3
2024 Scene-Object Holistic Relation Network for Fine-Grained Airplane Detection
abstract
The airplane detection and fine-grained recognition in the remote sensing images are challenging due to high interclass indistinction. The subtle distinctions between classes make it difficult to accurately classify objects based purely on bounding box features without considering the broader context. However, recent studies on remote sensing object detection focuses on refining the representation of bounding boxes while ignoring holistic context knowledge in remote sensing scenarios. This letter addresses this gap by introducing the scene-object holistic relation (SOHR) network for fine-grained airplane detection. Specifically, the SOHR network distinctively exploits global scene-object context information through a novel lightweight scene context attention (SCA) module, which aggregates scene context feature and object position information. Furthermore, the object relation transformer (ORT) is designed to model interactions among all objects within the scene explicitly, thereby increasing the model performance for ambiguous hard samples. The experimental results obtained from the FAIR1M dataset demonstrate that the proposed SOHR-Net achieves a state-of-the-art detection accuracy of 56.110% mean average precision (mAP). Compared with the baseline, SOHR-Net exhibits an increase of 2.517%.
Weiyu Ning, Qixiong Wang, Jiaqi Feng 0001, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.4
2024 Balanced Orthogonal Subspace Separation Detector for Few-Shot Object Detection in Aerial Imagery
abstract
Few-shot object detection (FSOD) in remote sensing images (RSIs) aims to achieve object location and classification with only a few training samples. Currently, mainstream transfer-learning methods employ a two-stage approach: pretraining on data-abundant base classes and fine-tuning on few-shot novel classes. However, existing approaches suffer notable degradation in both base and novel classes during fine-tuning, because of gradient conflict and class imbalance. To address this, we construct the balanced orthogonal subspace separation (BOSS) detector, a novel two-stage framework for FSOD. Specifically, to avoid contradictory gradients, BOSS distinctly isolates the training of base and novel classes at both structural and feature levels. For structural separation, a low-rank subspace adapter (LoSA) is introduced to ensure network optimization for novel classes without hampering base classes’ pretraining performance, effectively addressing over-fitting in few-shot scenarios. For feature disentanglement, an orthogonal subspace extractor (OSE) is presented, enhancing class separability by learning class-specific, orthogonal basis-spanned subspace. Finally, a balanced classifier (BC) is proposed to equalize the imbalanced loss, with its dual-component design mitigating bias toward predicting background or base classes. Comparative evaluations on diverse remote sensing datasets demonstrate BOSS’s superiority, outperforming state-of-the-art methods in mean average precision (mAP). These results underscore BOSS’s effectiveness in FSOD, particularly in challenging remote sensing contexts.
Hongxiang Jiang, Qixiong Wang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.1
2024 Disentangled Foreground-Semantic Adapter Network for Generalized Aerial Image Few-Shot Semantic Segmentation
abstract
Semantic segmentation of remote sensing imagery requires extensive annotated samples for training, facing challenges in adapting to novel classes with few annotations. Few-shot semantic segmentation (FS-Seg) employs a support-to-query paradigm, which encounters many practical constraints. Recently, generalized few-shot semantic segmentation (GFS-Seg) has been proposed to align with general semantic segmentation paradigms, enabling the segmentation of all classes (both base and novel classes) in the image. However, existing GFS-Seg methods struggle with a large intra-class variance of background, degradation on base classes, and overfitting on novel classes during fine-tuning in aerial imagery. To address the above issues, we propose the disentangled foreground-semantic adapter network (DFSA-Net) for generalized aerial image FS-Seg. Specifically, to reduce the interference from background features, DFSA-Net employs a foreground-semantic decoder (FSD) to decompose semantic segmentation into foreground aggregation and multiclass refinement. To mitigate the base classes degradation and novel class overfitting during fine-tuning, we propose disentangled low-rank adapter (DLA) for fine-tuning phase, designed to preserve the base parameters while ensuring efficient adaptation to novel classes. Finally, we introduce an inference ensemble strategy that merges base and novel decoder prediction to achieve final output. Experimental results on NWPU and iSAID datasets demonstrate the superiority of our DFSA-Net over other compared methods.
Qixiong Wang, Jihao Yin, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang
IEEE Trans. Geosci. Remote. Sens.3
2021 Unsupervised Domain Adaptation for Semantic Segmentation via Self-Supervision
abstract
Recently, deep learning (DL) methods have been widely used for semantic segmentation of remote sensing and achieved significant progress. However, DL-based methods are time-consuming and labor intensive for the networks requiring abundant data with accurate labeling. To solve this issue, unsupervised domain adaption (UDA) has recently been used to transfer the information from labeled source domain to unlabeled target domain. In this paper, we propose a novel UDA approach based on the self-supervised theory for remote sensing image. Specifically, we firstly utilize the inter-domain adaptation to reduce the gap between the source and target domain. Secondly, based on our proposed spatial-frequency (SF) index, we detach the target domain into an easy and hard split. Ultimately, we adopt the intra-domain adaptation by self-supervised adaptation to improve the performance of hard split. Experimental results on ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness and rationality of our methods against the other state-of-the-art approaches.
Weifa Shen, Qixiong Wang, Hongxiang Jiang, Jihao Yin
IGARSS3