Qiwu Luo

dblp:179/3803 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-2822-5538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Prompt is All You Need: Prompting Foundation Models for Large-Scale Self-Supervised Semantic Segmentation
abstract
This paper addresses the important and challenging task of large-scale unsupervised semantic segmentation (LUSS). We present the first attempt to unleash the power of foundation models (FMs) for the challenging, dense prediction task LUSS, and our main objective is to present simple, effective yet efficient solutions for LUSS, namely Prompting foundation models for LUSS (PLUSS). Firstly, we proposed a cascade framework PLUSS$_\alpha$α by effectively marrying CLIPS, Grounding DINO, and SAM in a zero-shot manner. This cascade architecture automatically generates semantic and spatial prompts for SAM, establishing a strong baseline that significantly outperforms previous state-of-the-art methods. Building upon this foundation, we propose PLUSS$_\beta$β, which addresses the critical bottleneck of prompt quality through two novel tuner modules: a semantic tuner that enhances fine-grained category discrimination via visual prompt tuning, and a box tuner that improves object localization through cross-modal feature fusion. Both tuners are optimized by capitalizing on the knowledge already present within the foundation models themselves, deriving self-supervised signals from internal model consistency. This approach requires no external supervision or updates to the foundation models' parameters. Extensive experiments on ImageNet-S benchmarks demonstrate that PLUSS$_\beta$β achieves remarkable performance improvements, surpassing the previous best method by 39.6%, 27.3%, and 22.6% in mIoU for 50, 300, and 919 categories respectively. Our approach exhibits robust category-shape representation across varying object sizes and dataset scales, while maintaining strong generalization capabilities for open-vocabulary tasks. The proposed framework provides a solid baseline for adapting foundation models to downstream vision tasks.
Jiaojiao Su, Qiwu Luo, Shuzhou Sun, Yuenan Hou, Xinyu Zhang 0010, Janne Heikkilä, Chunhua Yang 0001, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Physics-Informed Fusion of Vision and Sensor Data for CO2 Prediction
abstract
Accurate real-time prediction of indoor Carbon dioxide concentration ($\text{CO}_{2}$) is crucial to optimizing ventilation, ensuring the well-being of the occupants, and improving energy efficiency. Traditional methods typically predict$\text{CO}_{2}$concentration with unimodal time series. The recent advance from multi-modal learning provides a new dimension of fusing visual and timeseries sensor data. Although multimodality offers a more holistic view of the environment, it introduces potential conflicts from physical dynamics that conventional fusion methods could simply overlook. This paper identifies and addresses two such challenges: (1) Multimodal Temporal Conflict, which arises when static visual cues correspond to dynamic$\text{CO}_{2}$changes; (2) Cross-Modal Causal Lag, which stems from the physical inertia of air volume that delays the sensor's response to occupancy changes. To address these challenges, we propose a novel physics-informed fusion framework to mitigate inconsistencies during transient physical states and model the time-varying delay between visual events and their corresponding$\text{CO}_{2}$responses. We validate our framework on a month-long, real-world multimodal dataset collected from an office environment. Extensive experiments demonstrate that our proposed methods significantly outperform both unimodal timeseries models and conventional fusion techniques, highlighting the critical importance of embedding physical priors into fusion models for robust and accurate environmental sensing.
Shouqi Wang, Lijuan Lan, Qiwu Luo
ICPADS3
2025 Few-Shot Parameter Efficient Finetuning for SAM in Salient Steel Surface Defect Detection
abstract
Current deep learning methods for strip steel defect detection prioritize accuracy and computational efficiency but struggle to generalize to new production environments with unseen defect types. While visual foundation models like segment anything model (SAM) offer broader generalization through pretrained visual knowledge, they face significant domain shifts when applied to industrial defect imagery. To address this, we propose AdaptedSAM, a parameter-efficient few-shot fine-tuning framework for SAM that optimizes adapter placement and feature interaction. We first demonstrate that adapter positioning within the encoder critically impacts few-shot adaptation performance. Building on this, we introduce an efficient global optimizer that dynamically allocates adapter weights via cross-block communication across transformer layers. Furthermore, we develop an adapted mask decoder to enhance detailed mask information. Our adaptedSAM achieves state-of-the-art defect detection with precise boundaries in five-shot learning, outperforming 12 state of the art (SOTA) models on the SD-saliency-900 steel dataset while ensuring computational efficiency. Cross-domain evaluations on six benchmarks (industrial defects, remote sensing, medical imaging, natural scenes) demonstrate robust generalization.
Jiaojiao Su, Qiwu Luo, Weihua Gui 0001, Chunhua Yang 0001
IEEE Trans. Ind. Informatics2
2024 PMSA-DyTr: Prior-Modulated and Semantic-Aligned Dynamic Transformer for Strip Steel Defect Detection
abstract
In-process hot-rolled strip steel is suffering from some complicated yet unavoidable surface defects due to its harsh production environment. The automated visual inspection on defects consistently faces challenges of interclass similarity, intraclass difference, low contrast, and overlapping issue, which tend to trigger false or missed detections. This article proposes a prior-modulated and semantic-aligned dynamic transformer, called PMSA-DyTr. In this framework, a long short-term self-attention embedded with local convolution is designed for assisting an encoder to eliminate noise ambiguity between defects and backgrounds. Then, a semantic aligner is cleverly bridged between the encoder and the decoder to align the sematic for speeding up the convergence, and prior-modulated cross attention is proposed to alleviate the deficiency of samples for a data-driven transformer. Furthermore, a gate controller is innovatively constructed to dynamically select the minimal number of encoder blocks while preserving detection accuracy. The proposed PMSA-DyTr outperforms 19 state-of-the-art models on mean average precision with an inference time of 54.67 ms and visually performs best in detecting low-contrast and multiple small defects.
Jiaojiao Su, Qiwu Luo, Chunhua Yang 0001, Weihua Gui 0001, Olli Silvén, Li Liu 0002
IEEE Trans. Ind. Informatics2
2022 Scale-selective and noise-robust extended local binary pattern for texture classification
Qiwu Luo, Jiaojiao Su, Chunhua Yang 0001, Olli Silvén, Li Liu 0002
Pattern Recognit.1
2022 SAR Target Classification Using the Multikernel-Size Feature Fusion-Based Convolutional Neural Network
abstract
It is well-known that the convolutional neural network (CNN) is an effective method for synthetic aperture radar (SAR) target classification. In the convolutional layer of CNN, convolutional kernels of different sizes can extract different feature information of the target. The small-size kernel can extract the local texture feature information, and the large-size kernel can extract the global contour feature information. Traditional CNN methods usually use fixed-size kernels for convolution, and they generally lose part of the target’s feature information, resulting in the inaccurate classification of the SAR targets. This article proposes a novel CNN model based on multikernel-size feature fusion (MKSFF-CNN) for SAR target classification. MKSFF-CNN designs a convolutional methodology with a multichannel parallel topology, it uses convolutional kernels of different sizes to extract the multikernel-size deep features of the SAR target, and then, these features are fused in an optimal way to acquire the lowest loss. Moreover, MKSFF-CNN concatenates the fused features extracted by the convolutional layers of different dimensions to achieve the finest classification. MKSFF-CNN greatly elevates the feature representation completeness of the SAR targets so that more useful feature information can be exploited for SAR target classification. Undoubtedly, MKSFF-CNN can achieve a better classification performance compared with traditional CNN models with a fixed kernel size. The superiority of MKSFF-CNN is validated on the moving and stationary target acquisition and recognition (MSTAR) dataset with the detailed objective and subjective evaluation.
Jiaqiu Ai, Yuxiang Mao, Qiwu Luo, Mengdao Xing
IEEE Trans. Geosci. Remote. Sens.3
2022 A Fine PolSAR Terrain Classification Algorithm Using the Texture Feature Fusion-Based Improved Convolutional Autoencoder
abstract
In order to more efficiently mine the features of polarimetric synthetic aperture radar (PolSAR) and establish a more appropriate classification model, this article proposes an improved convolutional autoencoder (ICAE) based on texture feature fusion (TFF-ICAE) for PolSAR terrain classification. First, TFF-ICAE specifically designs a multi-indicator squeeze-and-excitation (MI-SE) block and incorporates it into the CAE network. MI-SE can enhance the essential feature information while suppressing the interference information as much as possible, and it can effectively increase the between-class distance while reducing the within-class distance. Then, TFF-ICAE uses gray level co-occurrence matrix (GLCM) to capture the texture features, and it optimally fuses these texture features and the deep features extracted by ICAE to complete the multilevel feature fusion, elevating the feature representation completeness of the terrain. That is, TFF-ICAE effectively enhances the feature separation capability of different categories while greatly elevating the feature representation completeness. Experiments on the datasets of San Francisco, Oberpfaffenhofen, and Flevoland show that the proposed TFF-ICAE, respectively, achieves overall accuracies of 93.44%, 97.61%, and 97.78%, which are at least 0.92%, 1.52%, and 0.97% higher than other algorithms. Undoubtedly, the superiority of TFF-ICAE is verified on these datasets.
Jiaqiu Ai, Yuxiang Mao, Qiwu Luo, Baidong Yao, Mengdao Xing, Yanlan Wu
IEEE Trans. Geosci. Remote. Sens.4
2022 MPA-RNN: A Novel Attention-Based Recurrent Neural Networks for Total Nitrogen Prediction
abstract
Accurately predicting the short- and long-term variations of total nitrogen (TN) is vital for operating the wastewater treatment plants (WWTPs), considering the critical role TN plays in reflecting the eutrophication of wastewater. However, only a few relevant water quality parameters with limited samples can be obtained in WWTPs, which tremendously increases the difficulty in precisely predicting TN concentration. In this study, a multiphase attention-based recurrent neural network (MPA-RNN) is proposed. Benefited from its unique decomposition-summary attention structure, MPA-RNN first learns the temporal correlations and effectively excavates the useful information hidden in the historical data. Then, by designing a two-channel structure to transmit attention information, summary attention can integrate the decomposed information and learn the spatial relationships without information loss. Experimental results demonstrate that MPA-RNN achieves the best performance on both the SML2010 and practical TN datasets with the smallest root-mean-squared error, mean absolute error, and mean absolute percentage error when compared with the other state-of-the-art methods.
Jingxuan Geng, Chunhua Yang 0001, Yonggang Li 0002, Lijuan Lan, Qiwu Luo
IEEE Trans. Ind. Informatics5
2020 Jointly optimized echo state network for short-term channel state information prediction of fading channel
abstract
Accurately obtaining channel state information (CSI) in wireless systems is significant but challenging. This paper focuses the technique of machine-learning-based channel estimation. In particular, a jointly optimized echo state network (JOESN) is proposed to form a concept of the CSI prediction which is made up of two interacting aspects of output weight regularization and initial parameter optimization. First, in order to enhance noise robustness, a sparse regression based on L2 regularization is employed to finely learn the output weights of ESN. Second, vital reservoir parameters (i.e., global scaling factor, reservoir size, scaling coefficient and sparsity degree) are learned by a linear-weighted particle swarm optimization (LW-PSO) for further improve the prediction accuracy and reliability. The experiments about computational complexity and three evaluating metrics are carried out on two chaotic benchmarks and one real-world dataset. The analyzed results indicate that the JOESN performs promisingly on multivariate chaotic time series prediction.
Qiwu Luo, Yichuang Sun, Oluyomi Simpson
IWCMC1
2020 Outliers-Robust CFAR Detector of Gaussian Clutter Based on the Truncated-Maximum-Likelihood- Estimator in SAR Imagery
abstract
This paper proposes an outliers-robust constant false-alarm rate (OR-CFAR) detector of Gaussian clutter based on the truncated-maximum-likelihood estimator (TMLE) in SAR imagery. The proposed method aims at elevating the detection performance in multiple-target environment, where the sea clutter samples are often contaminated by the interfering target pixels, the azimuth ambiguities, and the breakwater. As a consequence, the parameters used for statistical modeling are over-estimated, resulting in a degradation of the CFAR detection rate. Inspired by the traditional two-parameter CFAR (TP-CFAR) detector of Gaussian clutter, OR-CFAR designs an adaptive threshold-based clutter truncation method to eliminate the high-intensity outliers from the clutter samples in the local reference window, and the probability density function (PDF) of the sea clutter can be accurately modeled through the newly raised TMLE. Furthermore, the optimal truncation depth used for clutter truncation and PDF modeling is evaluated and selected properly to get the best detection results. The OR-CFAR greatly enhances the CFAR detection rate in multiple-target environment, and it is computationally simple and efficient, which has a great application value. The Chinese Gaofen-3 SAR data are used for experiments to show the better detection performance of OR-CFAR.
Jiaqiu Ai, Qiwu Luo, Xuezhi Yang, Zhiping Yin
IEEE Trans. Intell. Transp. Syst.2
2019 Multi-Scale Rotation-Invariant Haar-Like Feature Integrated CNN-Based Ship Detection Algorithm of Multiple-Target Environment in SAR Imagery
abstract
This paper proposes a multi-scale rotation-invariant haar-like (MSRI-HL) feature integrated convolutional neural network (MSRIHL-CNN)-based ship detection algorithm of the multiple-target environment in synthetic aperture radar (SAR) imagery. Usually, ship detection includes preprocessing, prescreening, discrimination, and classification. Among them, prescreening and discrimination are the most two important stages so that they catch great intention. Based on our previous work, we propose a truncated-clutter-statistics-based joint, constant false alarm rate (CFAR) detector (TCS-JCFAR) for ship target prescreening in the multiple-target environment. TCS-JCFAR greatly enhances the prescreening rate in the multiple-target environment while achieving a low observed FAR. In the discrimination stage, conventional CNN extracts the deep features (high-level features); however, it will lose the local texture and edge information (low-level features) which are of great significance for target discrimination. Hence, the MSRI-HL features are used to represent the multi-scale, rotation-invariant texture, and edge information that conventional CNN fails to capture. The extracted low-level MSRI-HL features and the high-level deep features are optimally fused to a multi-layered feature vector. Finally, the multi-layered feature vector is fed into a typical support vector machine (SVM) classifier for ship target discrimination. The proposed MSRIHL-CNN combines the low-level texture and edge features and the high-level deep features; moreover, they are optimally fused to fully represent the ship targets. Undoubtedly, MSRIHL-CNN has better discrimination performance. The superiority of the proposed TCS-JCFAR-based prescreener and MSRIHL-CNN-based discriminator is validated on the Chinese Gaofen-3 SAR imagery.
Jiaqiu Ai, Ruitian Tian, Qiwu Luo, Bo Tang 0011
IEEE Trans. Geosci. Remote. Sens.3
2018 Time-effective Fault Diagnosis Algorithms for Analog and Mixed-signal Circuits Using Sparsity-aware Multi-class Relevance Vector Machine
abstract
Except for the advantages of supporting arbitrary kernels, probabilistic predictions and automatic estimation of hyper-parameters, relevance vector machine (RVM) also encounters some of training time increase and classification accuracy recession, compared with SVM. In order to suppress such `nuisance' imperfections, this paper proposed a sparsity-aware RVM model for multi-class classification (denoted as Sa-MRVM) by developing a configurable singular entropy decision mechanism. Multiple driven data sets captured from both emulational and actual circuits under test (CUTs) are involved to further improve the model's generalization ability and judging confidence. Experimental results carried out on two CUTs indicate that our proposed learning methodology is speedy and accurate enough for real world fault diagnosis tasks of analog and mixed-signal circuits.
Qiwu Luo, Yigang He 0001, Yichuang Sun, Lifen Yuan
ISCAS1