Xiaoning Song

dblp:82/6511 · DBLP profile ↗
← Back
108ranked-venue papers
15as first author
71since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 10 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 26 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HGPose: Geometry-Aware Part Discovery and Hybrid Graph Refinement for Animal Pose Estimation
Lianjun Geng, Xiaoning Song
ICIC (12)3
2026 Backtrack-on-Graph: Efficient Path Navigation via Candidate Stack and Role-Based Adaptive Update
Yufan Lu, Xiaoning Song
ICIC (24)4
2026 DFusion-DFI: A Multimodal Feature Fusion Model for Drug-Food Interaction Prediction
Xiaoning Song
ICIC (28)3
2026 FreqSeqNet: Dual-Domain Modeling for Sequential Face Manipulation Order Detection
Xiangxu Zhai, Lianjun Geng, Xiaoning Song
ICIC (19)3
2026 Coarse-to-fine dual flexible-competition hybrid collaborative-nonnegative representation method for image classification
Zi-Qi Li, Xiaoning Song, Tianyang Xu 0001
Expert Syst. Appl.5
2026 Positive and negative neighbor dual-flexible nonnegative representation method for image classification
Xiaoning Song, Tianyang Xu 0001
Inf. Process. Manag.5
2026 Vision Mamba-enhanced Multi-level Context Aggregation Network for polyp segmentation
Jiaye Chen, Tianyang Xu 0001, Rui Wang 0050, Xiaoning Song, Tao Zhou 0002
Pattern Recognit.5
2025 R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration
abstract
Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.
Yang Hua 0002, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Rui Wang 0050, Wenjie Zhang 0009, Xiaojun Wu 0001
AAAI3
2025 Catching Inter-Modal Artifacts: A Cross-Modal Framework for Temporal Forgery Localization
Yuhan Cai, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001
ICIC (6)4
2025 Document-Level Relation Extraction with Retrieval-Augmented
Qiyi Jiang, Wenjie Zhang 0009, Yueli Yang, Yang Hua 0002, Xiaoning Song
ICIC (24)5
2025 HAP-MT: Alternating Perturbation Strategies Across Data and Feature Levels in Semi-Supervised Medical Image Segmentation
Xiaoning Song
ICIC (28)4
2025 Decoupled Dual-Path Diffusion: Precise Spatial-Semantic Modeling for Human-Object Interaction Generation
Wenxiao Wan, Yang Hua 0002, Wenjie Zhang 0009, Yuhan Cai, Xiaoning Song
ICIC (21)5
2025 FreeForm-Prior: Parametric-Guided Model-Free 3D Human Mesh Reconstruction
Muyan Zhao, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song
ICIC (5)4
2025 A Correlation Manifold Self-Attention Network for EEG Decoding
abstract
Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geometrically capture the spatiotemporal dependencies inherent in time series data, e.g., electroencephalography (EEG). Recent studies have highlighted the full-rank correlation matrix as an advantageous alternative to the covariance matrix for data representation, owing to its invariance to the scale of variables. Motivated by these advancements, we propose the Correlation Attention Network (CorAtt) tailored for full-rank correlation matrices and implement it under the permutation-invariant and computationally efficient Off-Log and Log-Scaled geometries, respectively. Extensive evaluations on three benchmarking EEG datasets provide substantial evidence for the effectiveness of our introduced CorAtt. The code and supplementary material can be found at https://github.com/ChenHu-ML/CorAtt.
Rui Wang 0050, Xiaoning Song, Tao Zhou 0002, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001
IJCAI3
2025 3D Joint-Aware Features in GRU-based Kinematic Chain for Human Mesh Recovery
abstract
Model-based methods, which define pose parameters as 3D rotations of 24 joints relative to their parent joints, have proven highly effective in recovering 3D human meshes. Existing methods typically estimate these parameters by coupling features extracted from 2D images. However, the performance of these methods is always limited by spatial ambiguities due to dimensional inconsistencies and kinematic chain ambiguities caused by direct regression of all parameters. To overcome these limitations, we present a novel method that uses 3D joint-aware features within a GRU-based kinematic chain (AFGK) framework to obtain spatially accurate features. Specifically, we first use a self-attention mechanism to enable mutual interactions between features around keypoints, generating 3D joint-aware features, improving spatial accuracy of keypoints, and robustness to extreme poses. Then, a GRU-structured Kinematic chain update module (GRK) is introduced to decouple and update the 24 pose parameters. This module captures the kinematic relationships between keypoints and uses the features of parent joints to dynamically update the pose parameters of child joints, further reducing relative rotation errors. Extensive experimental results show that our method outperforms state-of-the-art competitors on standard benchmarks, significantly improving human reconstruction. The codes will be available on the project homepage: https://github.com/jingchenzhang/AFGK.
Jingchen Zhang, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song
IJCNN4
2025 Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
abstract
Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality.
Xiaoqing Luo, Zhancheng Zhang, Hui Li 0037, Rui Wang 0050, Zhenhua Feng 0001, Xiaoning Song
NeurIPS8
2025 Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds
abstract
Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored to specific geometries, limiting their applicability. On the other hand, recent studies show that several matrix manifolds, such as Symmetric Positive Definite (SPD), Symmetric Positive Semi-Definite (SPSD), and Grassmannian manifolds, admit gyrovector structures, which extend vector addition and scalar product into manifolds. Leveraging these properties, we propose a Gyro Attention (GyroAtt) framework over general gyrovector spaces, applicable to various matrix geometries. Empirically, we manifest GyroAtt on three gyro structures on the SPD manifold, three on the SPSD manifold, and one on the Grassmannian manifold. Extensive experiments on four electroencephalography (EEG) datasets demonstrate the effectiveness of our framework.
Rui Wang 0050, Xiaoning Song, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001
NeurIPS3
2025 Geometry-Aware Self-attention Network with Adaptive Log-Euclidean Metric for EEG Decoding
Zihao Bi, Rui Wang 0050, Tao Zhou 0002, Xiaoning Song, Xiaojun Wu 0001
PRCV (4)5
2025 Medical Vision Language Model With Multi-granularity Alignment and Learning Data Augmentation
Daoqiang Gao, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song
PRCV (13)4
2025 Fastere: a fast framework for entity relation extractions
Wenjie Zhang 0009, Tianyang Xu 0001, Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song
Data Min. Knowl. Discov.5
2025 Towards fine-grained adaptive video captioning via Quality-Aware Recurrent Feedback Network
Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
Expert Syst. Appl.3
2025 OCCO: LVM-Guided Infrared and Visible Image Fusion Framework Based on Object-Aware and Contextual Contrastive Learning
Hui Li 0037, Congcong Bian, Zeyang Zhang 0002, Xiaoning Song, Xi Li 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.4
2025 Learning Structure-Supporting Dependencies via Keypoint Interactive Transformer for General Mammal Pose Estimation
Tianyang Xu 0001, Jiyong Rao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.3
2025 Link prediction via adversarial knowledge distillation and feature aggregation
Xiaoning Song, Wenjie Zhang 0009, Yang Hua 0002, Xiaojun Wu 0001
Multim. Syst.2
2025 MMDG-DTI: Drug-target interaction prediction via multimodal feature fusion and domain generalization
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.3
2025 Large-Scale Retrieval and Quality Control of Leaf Area Index Based on ICESat-2 Spaceborne Photon-Counting Laser Altimeter
abstract
Spaceborne LiDAR provides a promising method for large-scale characterizing LAI. However, the quality of point cloud data from spaceborne LiDAR, especially ICESat-2, is susceptible to atmosphere and background noise, introducing considerable uncertainty in LAI retrieval. Thus, efficiently screening out the high-quality point cloud is a significant guarantee for high-quality LAI retrieval. In this study, we proposed a quality control (QC) method that employed the number of 10 m windows without ground points in the ICESat-2 100 m segment as the QC flag. This method divided segments into 11 QC flags from 0 to 10 and was applied to LAI retrieval across Chinese forests from 2019 to 2020. The field measurements at locations identical to ICESat-2 ground tracks were used to validate the ICESat-2 LAI at different QC flags. The results showed that the proposed method effectively improved point cloud quality recognition and LAI accuracy, with ICESat-2 LAI (QC < 3) reducing RMSE by 26.36% compared to all ICESat-2 LAIs. It also showed good agreement with MODIS and GLASS LAI and mitigated saturation issues in passive optical imagery. The ICESat-2 LAI with QC < 3 performed better in deciduous broadleaved, evergreen needle-leaved, deciduous needle-leaved, and mixed forests, but not in evergreen broadleaved forests. ICESat-2 LAI was particularly adept at capturing high LAI values, which had the highest proportion of LAI values over 6.0 compared to MODIS and GLASS LAI. The proposed method has the potential for large-scale and high-quality LAI retrieval using ICESat-2 data on a global scale.
Da Guo, Xiaoning Song, Ronghai Hu, Max Mallen-Cooper, Yuzhen Xing, Ruijin Li, Hong Zeng 0004, Guangjian Yan, Paul Kardol
IEEE Trans. Geosci. Remote. Sens.2
2025 S4Fusion: Saliency-Aware Selective State Space Model for Infrared and Visible Image Fusion
abstract
The preservation and the enhancement of complementary features between modalities are crucial for multi-modal image fusion and downstream vision tasks. However, existing methods are limited to local receptive fields (CNNs) or lack comprehensive utilization of spatial information from both modalities during interaction (transformers), which results in the inability to effectively retain useful information from both modalities in a comparative manner. Consequently, the fused images may exhibit a bias towards one modality, failing to adaptively preserve salient targets from all sources. Thus, a novel fusion framework (S4Fusion) based on the Saliency-aware Selective State Space is proposed. S4Fusion introduces the Cross-Modal Spatial Awareness Module (CMSA), which is designed to simultaneously capture global spatial information from all input modalities and promote effective cross-modal interaction. This enables a more comprehensive representation of complementary features. Furthermore, to guide the model in adaptively preserving salient objects, we propose a novel perception-enhanced loss function. This loss aims to enhance the retention of salient features by minimizing ambiguity or uncertainty, as measured at a pre-trained model's decision layer, within the fused images. The code is available at https://github.com/zipper112/S4Fusion.
Haolong Ma, Hui Li 0037, Chunyang Cheng, Gaoang Wang, Xiaoning Song, Xiaojun Wu 0001
IEEE Trans. Image Process.5
2025 ATMNet: Adaptive Two-Stage Modular Network for Accurate Video Captioning
abstract
In recent years, pretrained language-image models (PLIMs) have delivered advances in video captioning. However, existing PLIMs primarily focus on extracting global feature representations from still images and text sequences, while neglecting fine-grained semantic alignment and temporal variations between vision and text pairs. To this end, we propose a global-local alignment module and a temporal parsing module to reflect the detailed correspondence and temporal perception between the two modalities, respectively. In particular, the global-local alignment module enables cross-modal registration at two levels, i.e., the sentence-video level and the word-frame level, to obtain mixed-granularity semantic video features. The temporal parsing module is a dedicated self-attention structure that highlights temporal order cues along video frames, compensating for the limited temporal capacity of PLIMs. In addition, an adaptive two-stage gating structure is designed to leverage the linguistic predictions further. The linguistic information derived from the first stage prediction is dynamically routed through an adaptive decision gate, allowing for quality assessment of whether the information should proceed to the second stage. This structure can effectively reduce the computational burden for easy samples and further improve the accuracy of the prediction results. The experimental results obtained on several benchmark datasets demonstrate the effectiveness of the proposed solution, with improved performance compared to state-of-the-art methods.
Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 LabelPrompt: Effective prompt-based learning for relation classification
Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001
ACML2
2024 DRIVPocket: A Dual-stream Rotation Invariance in Feature Sampling and Voxel Fusion Approach for Protein Binding Site Prediction
Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song, Xiaojun Wu 0001
ICPR (12)4
2024 FewConv: Efficient Variant Convolution for Few-Shot Image Generation
Si-Hao Liu, Xiaoning Song, Jia-Sheng Chen, Xiaojun Wu 0001
ICPR (3)3
2024 SMFuse: Two-Stage Structural Map Aware Network for Multi-focus Image Fusion
Tianyu Shen, Hui Li 0037, Chunyang Cheng, Xiaoning Song
ICPR (22)5
2024 A Novel Loss for Contrastive Deep Supervision
Zhengming Ye, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
ICPR (25)4
2024 Attention-Based Patch Matching and Motion-Driven Point Association for Accurate Point Tracking
Han Zang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler
ICPR (16)4
2024 Cluster-Mined Negative Samples for Enhanced Unsupervised Sentence Representation Learning
Yuhang Zhang 0022, Wenjie Zhang 0009, Yang Hua 0002, Xiaoning Song, Xiaojun Wu 0001
ICPR (26)5
2024 A Grassmannian Manifold Self-Attention Network for Signal Classification
Rui Wang 0050, Ziheng Chen 0001, Xiaojun Wu 0001, Xiaoning Song
IJCAI5
2024 Updating Depth-aware Feature in the Feedback Loop for Human Mesh Recovery
abstract
Recently, depth ambiguity reduction has made significant progress in monocular human mesh reconstruction. However, regression-based methods struggle to perceive minor depth deviations due to insufficiently exploited monocular cues, such as linear perspective, shadows, and pose priors of the given image. To tackle this limitation, we propose a Depth-aware Updated Feature Feedback (DAUFF) loop to iteratively exploit spatial information from rendered images. Specifically, the normalized depth and dense correspondence images rendered from the ground truth human meshes are adopted as auxiliary supervision for more reliable and adequate cues. Moreover, we also use dense correspondences for feature update to improve the adaptability to new predictions in the loop. Remarkably, dense correspondences can facilitate the rectification of parameters by providing a fine-grained perception of the current prediction. Extensive quantitative comparison on standard benchmarks shows that DAUFF achieves better-aligned reconstruction results than existing approaches. The dataset and codes will be available on the project homepage: https://github.com/github1838/DAUFF.
Yang Hua 0002, Xiaoning Song, Wenjie Zhang 0009, Xiaojun Wu 0001
IJCNN3
2024 RRANet: A Reverse Region-Aware Network with Edge Difference for Accurate Breast Tumor Segmentation in Ultrasound Images
Xiaoning Song, Yang Hua 0002, Wenjie Zhang 0009
PRCV (14)2
2024 FedDCP: Personalized Federated Learning Based on Dual Classifiers and Prototypes
Yang Hua 0002, Xiaoning Song, Wenjie Zhang 0009, Xiaojun Wu 0001
PRCV (1)3
2024 CoMoFusion: Fast and High-Quality Fusion of Infrared and Visible Image with Consistency Model
Zhiming Meng, Hui Li 0037, Zeyang Zhang 0002, Yunlong Yu 0001, Xiaoning Song, Xiaojun Wu 0001
PRCV (8)6
2024 EDS: Exploring deeper into semantics for video captioning
Yibo Lou, Wenjie Zhang 0009, Xiaoning Song, Yang Hua 0002, Xiaojun Wu 0001
Pattern Recognit. Lett.3
2024 Towards accurate unsupervised video captioning with implicit visual feature injection and explicit
Tianyang Xu 0001, Xiaoning Song, Xuefeng Zhu 0003, Zhenhua Feng 0001, Xiaojun Wu 0001
Pattern Recognit. Lett.3
2024 Unified Referring Expression Generation for Bounding Boxes and Segmentations
abstract
Referring expression generation (REG) is a challenging task at the intersection of computer vision and natural language processing, which aims at generating natural language descriptions that uniquely refer to a specific object within an image. Existing REG approaches solely utilize bounding boxes in a rather primitive manner to specify target objects, and employ the classical Convolutional Neural Networks (CNNs) for image encoding, followed by recurrent layers for text generation. In this letter, we propose a novel end-to-end REG model. Our model highlights the target using bounding boxes and segmentations in a unified fashion. Specifically, we propose two settings for utilizing these signals: employing them as inputs to the model and as supervision signals for pre-training tasks. Additionally, we harness the power of the recently prevailed self-attention architecture to bridge targeted visual clues and text correspondence. During inference, our method achieves state-of-the-art performance in a one-stage manner, reflecting the potential of both bounding boxes and segmentation references in constructing REG solutions.
Zongtao Liu, Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
IEEE Signal Process. Lett.3
2024 APMG: 3D Molecule Generation Driven by Atomic Chemical Properties
abstract
Recently, mask-fill-based 3D Molecular Generation (MG) methods have become very popular in virtual drug design. However, the existing MG methods ignore the chemical properties of atoms and contain inappropriate atomic position training data, which limits their generation capability. To mitigate the above issues, this paper presents a novel mask-fill-based 3D molecule generation model driven by atomic chemical properties (APMG). Specifically, we construct a new attention-MPNN-based encoder and introduce the electronic information into atom representations to enrich chemical properties. Also, a multi-functional classifier is designed to predict the electronic information of each generated atom, guiding the type prediction of elements and bonds. By design, the proposed method uses the chemical properties of atoms and their correlations for high-quality molecule generation. Second, to optimize the atomic position training data, we propose a novel atomic training position generation approach using the Chi-Square distribution. We evaluate our APMG method on the CrossDocked dataset and visualize the docking states of the pockets and generated molecules. The obtained results demonstrate the superiority and merits of APMG over the state-of-the-art approaches.
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 Bottom-Up Estimation of Stand Leaf Area Index From Individual Tree Measurement Using Terrestrial Laser Scanning Data
abstract
Leaf area parameters are crucial in ecosystem studies. As ecophysiological models advance toward finer detail, accurately estimating LA at various scales becomes essential, particularly for diverse units like urban individual trees. Several algorithms based on terrestrial laser scanning (TLS) data have been developed to obtain the LA of individual trees. However, their use at the stand level needs further research. In this study, the comparative shortest-path algorithm (CSP) is introduced for the automatic individual tree segmentation, thereby facilitating the application of the path length distribution model (PATH) for leaf area estimation at the stand level. Using high-density TLS data, we presented a bottom-up estimation of stand leaf area index (LAI) from 50 individual tree measurements and validated the results at different scales. At the tree scale, the LA derived from TLS and allometric model were highly correlated, with an R-value of 0.83. At the stand scale, the proposed method provides consistent results with the allometric and TRAC instrument measurements, performing better than vertical upward photography. Generally, 23 shared stations under the forest are enough to accurately obtain the LA of 50 trees and the LAI in an urban forest stand. Sensitivity analysis shows that the method is not sensitive to TLS scan resolution and parameters used in tree crown envelope reconstruction. The proposed bottom-up approach provides a new way of estimating the LAI at stand level using TLS and has the advantage of providing multi-level leaf area information and avoiding the scale effect.
Yuzhen Xing, Ronghai Hu, Hengli Lin, Hong Zeng 0004, Da Guo, Guangjian Yan, Xiaoning Song, Pierre Kastendeuch, Marc Saudreau, Françoise Nerry, Kai Xue, Yanfen Wang
IEEE Trans. Geosci. Remote. Sens.7
2024 A Physical Process-Based Enhanced Adjacent Channel Retrieval Algorithm for Obtaining Cloudy-Sky Surface Temperature
abstract
Acquiring cloudy land surface temperature (LST) is crucial for terrestrial ecosystem monitoring and global climate observation. Although numerous methods have been proposed to retrieve cloudy-sky LST using microwave remote sensing, the physical significance of these methods is inadequate due to their oversimplification of the effects of atmospheric components and clouds. To obtain accurate cloudy-sky LST, by simultaneously considering the influences of water vapor and cloud properties on LST, a physics-based LST retrieval algorithm is developed for cloudy skies with adjacent three brightness temperatures (BTs) at 18.7, 23.8, and 36.5 GHz vertically polarization. We develop this algorithm by building a simulated database, which includes a large range of BTs, surface emissivities, air temperatures, and water vapor contents. Meanwhile, various cloudy atmospheric profiles are constructed to reveal multifarious weather conditions. The test results with the simulated database show that the algorithm has good accuracy with an RMSE of 1.75 K and MAE of 1.33 K under cloudy weather. Sensitivity analysis indicates that precipitable water vapor (PWV) and cloud liquid water (CLW) are indispensable for correcting cloudy LST. In particular, LST accuracy shows an evident sensitivity to PWV, and RMSEs are reduced with an increase in PWV. Meanwhile, the proposed algorithm was applied to AMSR-E BTs and ERA5 profile dataset over the China region in 2008 and validated with ground-based air temperatures. Results indicate that RMSEs between retrieved LSTs and true LSTs are 3.47 K for cloudy weather, and there is better performance at high water vapor status, with an RMSE of 2.05 K.
Xin-Ming Zhu, Xiaoning Song, Xiao-Tao Li, Fang-Cheng Zhou
IEEE Trans. Geosci. Remote. Sens.2
2023 Multi-modal Stream Fusion for Skeleton-Based Action Recognition
Ruixuan Pang, Rongchang Li 0001, Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
ICIG (3)4
2023 Disentangled Shape and Pose Based on Attention and Mesh Autoencoder
Xiaoning Song
ICIG (3)2
2023 LE2Fusion: A Novel Local Edge Enhancement Module for Infrared and Visible Image Fusion
Yongbiao Xiao, Hui Li 0037, Chunyang Cheng, Xiaoning Song
ICIG (1)4
2023 Semantic-Guided Multi-feature Fusion for Accurate Video Captioning
Tianyang Xu 0001, Xiaoning Song, Zhenghua Feng, Xiaojun Wu 0001
ICIG (4)3
2023 MFR-DTA: a multi-functional and robust model for predicting drug-target binding affinity and region
abstract
MOTIVATION: Recently, deep learning has become the mainstream methodology for drug-target binding affinity prediction. However, two deficiencies of the existing methods restrict their practical applications. On the one hand, most existing methods ignore the individual information of sequence elements, resulting in poor sequence feature representations. On the other hand, without prior biological knowledge, the prediction of drug-target binding regions based on attention weights of a deep neural network could be difficult to verify, which may bring adverse interference to biological researchers. RESULTS: We propose a novel Multi-Functional and Robust Drug-Target binding Affinity prediction (MFR-DTA) method to address the above issues. Specifically, we design a new biological sequence feature extraction block, namely BioMLP, that assists the model in extracting individual features of sequence elements. Then, we propose a new Elem-feature fusion block to refine the extracted features. After that, we construct a Mix-Decoder block that extracts drug-target interaction information and predicts their binding regions simultaneously. Last, we evaluate MFR-DTA on two benchmarks consistently with the existing methods and propose a new dataset, sc-PDB, to better measure the accuracy of binding region prediction. We also visualize some samples to demonstrate the locations of their binding sites and the predicted multi-scale interaction regions. The proposed method achieves excellent performance on these datasets, demonstrating its merits and superiority over the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/JU-HuaY/MFR.
Yang Hua 0002, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
Bioinform.2
2023 A Method for Estimating Daytime Average Evapotranspiration From Diurnal Land Surface Temperature Measurements
abstract
Evapotranspiration (ET) is an essential parameter in the water cycle and surface energy balance. Accurate estimation of daytime or daily ET is of great significance for many activities in human economic and social development. This study aims to propose a novel method for estimating daytime average ET by using temporal measurements of land surface temperature (LST) and net surface shortwave radiation (NSSR), following a previous developed elliptical relationship between diurnal cycles of LST and NSSR under cloud-free days. The method was primarily developed from the simulated data of a physics-based Atmosphere–Land Exchange (ALEX) model under different underlying surfaces and atmospheric conditions. Based on the simulated data, the proposed method showed considerable accuracy with the overall coefficient of determination (R2) of 0.958 and the root mean square error (RMSE) of 25.3 Wm-2. In addition, ground ET measurements at four Ameriflux sites (US-ARM, US-SRM, US-Whs, and US-Wkg) during the 2018 growing season were collected to assess the estimated daytime average ET. Results show an overall RMSE of 64.7 Wm-2for the estimated ET at the four sites, and the US-Whs site reveals a best accuracy (R2=0.825, RMSE=44.4 Wm-2). These results indicated a potential for generating daytime ET with geostationary satellite observations at regional scales in future development.
Yun-Jing Geng, Pei Leng, Xiaoning Song, Zhao-Liang Li
IEEE Geosci. Remote. Sens. Lett.3
2023 ETM-face: effective training sample selection and multi-scale feature learning for face detection
Junyuan He, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Multim. Tools Appl.2
2023 Global Context-Aware Feature Extraction and Visible Feature Enhancement for Occlusion-Invariant Pedestrian Detection in Crowded Scenes
Zhen Liu 0015, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Neural Process. Lett.2
2023 CPInformer for Efficient and Robust Compound-Protein Interaction Prediction
abstract
Recently, deep learning has become the mainstream methodology for Compound-Protein Interaction (CPI) prediction. However, the existing compound-protein feature extraction methods have some issues that limit their performance. First, graph networks are widely used for structural compound feature extraction, but the chemical properties of a compound depend on functional groups rather than graphic structure. Besides, the existing methods lack capabilities in extracting rich and discriminative protein features. Last, the compound-protein features are usually simply combined for CPI prediction, without considering information redundancy and effective feature mining. To address the above issues, we propose a novel CPInformer method. Specifically, we extract heterogeneous compound features, including structural graph features and functional class fingerprints, to reduce prediction errors caused by similar structural compounds. Then, we combine local and global features using dense connections to obtain multi-scale protein features. Last, we apply ProbSparse self-attention to protein features, under the guidance of compound features, to eliminate information redundancy, and to improve the accuracy of CPInformer. More importantly, the proposed method identifies the activated local regions that link a CPI, providing a good visualisation for the CPI state. The results obtained on five benchmarks demonstrate the merits and superiority of CPInformer over the state-of-the-art approaches.
Yang Hua 0002, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Exploring Photon-Counting Laser Altimeter ICESat-2 in Retrieving LAI and Correcting Clumping Effect
abstract
The Ice, Cloud, and Land Elevation Satellite-2 (ICESat-2) employs a unique multibeam photon counting approach to acquire a near-continuously sampled profile and provides more precise technology for mapping the leaf area index (LAI) at the global scale. The inversion accuracy of LAI is affected by the clumping effect, which has been an open question for spaceborne laser scanning (SLS). Here, we present a segmented method based on the path length distribution model to calculate the clumping-corrected LAI independently using ICESat-2 data. The results showed that the LAI derived by the proposed method with a 200 m segment was consistent with the airborne laser scanning (ALS)-derived LAI, with a root mean squared error (RMSE) of 0.37. A satisfactory agreement (RMSE$=1.03$) was also shown between moderate resolution imaging spectroradiometer (MODIS) LAI and ICESat-2 LAI. Moreover, the LAI derived by the proposed method was on average 31.72% higher than the LAIe derived by Beer’s law, which indicated that the proposed method achieved the purpose of correcting the clumping effect. The gap probability was calculated by the 200 m moving window and the path length distribution was obtained by the 1 m moving window as the model input had the highest accuracy. In addition, the limitation of the point cloud data and the time lag of ICESat-2 acquisitions and ALS observations may affect the inversion accuracy of LAI. This study proposed a feasible way to correct the clumping effect and invert LAI independently using ICESat-2 data, which has the potential to characterize vegetation structure precisely at regional and global scales.
Da Guo, Ronghai Hu, Xiaoning Song, Hengli Lin, Liang Gao 0010, Xinming Zhu
IEEE Trans. Geosci. Remote. Sens.3
2022 FoGMesh: 3D Human Mesh Recovery in Videos with Focal Transformer and GRU
Yihao He, Xiaoning Song, Tianyang Xu 0001, Yang Hua 0002, Xiaojun Wu 0001
BMVC2
2022 Memory-Token Transformer for Unsupervised Video Anomaly Detection
abstract
Video anomaly detection is crucial for behavior analysis, which has witnessed continuous progress in recent years with the auto-encoder based reconstruction framework. However, in some cases, abnormal frames may also be reconstructed well due to the strong representation ability of deep networks, increasing missed detection. To mitigate this issue, the existing methods usually the memory bank method. This method records normal patterns and assigns high errors for the reconstruction of abnormal frames into normal frames. In this paper, to better use the semantic information of normal videos recorded in the memory module, we introduce the Memory-Token Transformer (MTT) to boost the reconstruction performance on normal frames. We assume that the anomalies in a video mainly concentrate on the regions containing people and relevant objects. Therefore, during the decoding stage, we first extract the semantic concepts of a feature map and generate the corresponding semantic tokens. Then the tokens are combined with the proposed memory module. Last, we introduce a transformer to fuse the complex relationship among different tokens, and use 3D convolution with the pooling operator in our encoder to enhance spatio-temporal feature extraction as compared with 2D models. The experimental results obtained on various benchmarks demonstrate the effectiveness of the proposed method.
Youyu Li, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001
ICPR2
2022 KITPose: Keypoint-Interactive Transformer for Animal Pose Estimation
Jiyong Rao, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
PRCV (1)3
2022 A Simplified Approach to Retrieve the K-Band Microwave Surface Emissivity Under Clear Skies
abstract
Microwave land surface emissivity (MLSE) at the K band plays a key role in driving geophysical parameters, such as land surface temperature (LST). However, satellite-based MLSE currently is hard to be quickly retrieved in clear skies since the time cost is high in removing atmospheric contributions. In this letter, one clear-sky atmospheric profile dataset, including a wide range of precipitable water vapor (PWV) values, was constructed using the Thermodynamic Initial Guess Retrieval database for analyzing numerical relationships between PWV and atmospheric parameters. Then, a simplified algorithm was developed for accurately retrieving instantaneous K-bandMLSEs (18.7 and 23.8 GHz) under clear skies, which can significantly save the time of atmospheric correction. The sensitivity analysis shows that PWV is a key factor affecting MLSE estimation at 23.8 GHz, and the brightness temperature (BT) uncertainty has a greater impact on MLSE estimation than LST. Additionally, with LST derived from the Moderate-resolution Imaging Spectroradiometer, BT from the Advanced Microwave Scanning Radiometer Earth Observing System (AMSR-E), and the ERA5 reanalysis PWV in 2008, the proposed algorithm was respectively applied in Europe and the United States for presenting its applicability at a station scale and regional scale. The actual sounding profile and global AMSR-E MLSE product were used as validation datasets. Results indicate the simplified approach has a good performance with RMSEs less than 0.02 in the site and regional validations. Whereas, there are some apparent overestimations in estimating clear-skies MLSEs, especially for 23.8 GHz. We believe the proposed approach is promising for retrieving other parameters.
Xin-Ming Zhu, Xiaoning Song, Pei Leng, Xiao-Tao Li, Liang Gao 0010, Lirong Ding
IEEE Geosci. Remote. Sens. Lett.2
2022 Competitive Non-negative Representation with Image Gradient Orientations for Face Recognition
He-Feng Yin, Xiaojun Wu 0001, Xiaoning Song
Neural Process. Lett.3
2022 Using global information to refine local patterns for texture representation and classification
Xin Shu 0001, Xiaoning Song, Xiaojun Wu 0001
Pattern Recognit.4
2022 Target-Cognisant Siamese Network for Robust Visual Object Tracking
Yingjie Jiang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit. Lett.2
2022 Feature Alignment for Robust Acoustic Scene Classification Across Devices
abstract
This letter presents a feature alignment method for domain adaptive Acoustic Scene Classification (ASC) across recording devices. First, we design a two-stream network, in which each stream processes two features,i.e., Log-Mel spectrogram and delta-deltas, using two sub-networks. Second, we investigate different loss functions for feature alignment between the feature maps obtained by the source and target domains. Last, we present an alternate training strategy to deal with the data imbalance problem between paired and unpaired samples. The experimental results obtained on the DCASE benchmarks demonstrate the effectiveness and superiority of the proposed method. The source code of the proposed method is available athttps://github.com/Jingqiao-Zhao/FAASC.
Jingqiao Zhao, Qiuqiang Kong, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Signal Process. Lett.3
2022 Impact of Soil Salinity on Soil Dielectric Constant and Soil Moisture Retrieval From Active Microwave Remote Sensing
abstract
Soil salinity plays a key role in influencing the soil dielectric constant and soil backscatter coefficient. However, soil moisture (SM) retrieval models constructed based on active microwave data hardly consider soil salinity. Thus, obtaining the SM datasets with various salinity on regional and local scales is difficult. This study aimed to employ theoretical model simulation to investigate the errors of SM retrieval due to not considering the impact of soil salinity. Then, three typical saline soil dielectric constant models were validated and compared based on the experimental measurement datasets. Results show that the WYR saline soil dielectric constant model has excellent performance. The soil salinity mainly affects the imaginary part of the dielectric constant and the effect of salinity on the soil dielectric constant is more significant when the SM has larger values. In addition, in retrieving SM with soil salinity more than 10 g/kg, the retrieval result of SM has an absolute error of 0.04$\text{m}^{3}/\text{m}^{3}$and a relative error of 5% when not considering the soil salinity impact. In retrieving SM with soil salinity less than 10 g/kg, the retrieved SM error increased by 2%, and the absolute error increased by 0.01$\text{m}^{3}/\text{m}^{3}$as soil salinity increased by 3 g/kg. We believe that The study will give a theoretical reference for establishing the SM retrieval model in saline soil areas using microwave data.
Liang Gao 0010, Xiaoning Song, Pei Leng, Jian-Wei Ma, Xin-Ming Zhu, Ronghai Hu, Yanfen Wang, Dewei Yin
IEEE Trans. Geosci. Remote. Sens.2
2022 Estimate of Cloudy-Sky Surface Emissivity From Passive Microwave Satellite Data Using Machine Learning
abstract
The derivation of microwave land surface emissivity (MLSE) under various weather conditions from the microwave radiometer plays a crucial role in acquiring land surface and atmospheric parameters. Nevertheless, currently, most existing studies mainly focus on the clear-sky scenarios owing to a lack of cloudy-sky land surface temperature (LST) and uncertainties in simulating the scattering and emission properties of atmospheric hydrometeors. Under this background, with satellite observations and the random forest (RF) model, this study proposes a method to estimate the MLSE under cloudy skies. First, clear-sky MLSEs with satisfactory accuracy are retrieved by using the brightness temperatures (BTs) from the Advanced Microwave Scanning Radiometer-Earth sensor, LSTs from the Moderate Resolution Imaging Spectroradiometer, and atmospheric profiles from the ERA5 reanalysis. Then, the relation among the clear-sky MLSE and related impact factors is built with the RF and extended to the cloudy-sky environment for generating all-weather MLSEs with a 0.25°. The results show that the input datasets present a considerable impact on the calculation of instantaneous MLSE, and a 5.73 K bias of ERA5 LST may generate a 0.014-0.021 error in the MLSE from 6.9 to 89 GHz horizontal polarization, while the impacts of BT and profile uncertainties on the MLSE are smaller. The retrieved clear-sky MLSE is coincident with the existing MLSE for the spatiotemporal variations, and there is an average difference range from -0.035 to 0.035 in January 2008. Meanwhile, the constructed RF model can successfully apply to cloudy-sky status and recover the MLSE image gaps affected by cloud contamination.
Xin-Ming Zhu, Xiaoning Song, Pei Leng, Zhao-Liang Li, Xiao-Tao Li, Liang Gao 0010, Da Guo
IEEE Trans. Geosci. Remote. Sens.2
2021 MECT: Multi-Metadata Embedding based Cross-Transformer for Chinese Named Entity Recognition
abstract
Shuang Wu, Xiaoning Song, Zhenhua Feng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xiaoning Song
ACL/IJCNLP (1)2
2021 Why can deep convolutional neural networks improve protein fold recognition? A visual explanation by interpretation
abstract
As an essential task in protein structure and function prediction, protein fold recognition has attracted increasing attention. The majority of the existing machine learning-based protein fold recognition approaches strongly rely on handcrafted features, which depict the characteristics of different protein folds; however, effective feature extraction methods still represent the bottleneck for further performance improvement of protein fold recognition. As a powerful feature extractor, deep convolutional neural network (DCNN) can automatically extract discriminative features for fold recognition without human intervention, which has demonstrated an impressive performance on protein fold recognition. Despite the encouraging progress, DCNN often acts as a black box, and as such, it is challenging for users to understand what really happens in DCNN and why it works well for protein fold recognition. In this study, we explore the intrinsic mechanism of DCNN and explain why it works for protein fold recognition using a visual explanation technique. More specifically, we first trained a VGGNet-based DCNN model, termed VGGNet-FE, which can extract fold-specific features from the predicted protein residue-residue contact map for protein fold recognition. Subsequently, based on the trained VGGNet-FE, we implemented a new contact-assisted predictor, termed VGGfold, for protein fold recognition; we then visualized what features were extracted by each of the convolutional layers in VGGNet-FE using a deconvolution technique. Furthermore, we visualized the high-level semantic information, termed fold-discriminative region, of a predicted contact map from the localization map obtained from the last convolutional layer of VGGNet-FE. It is visually confirmed that VGGNet-FE could effectively extract distinct fold-discriminative regions for different types of protein folds, thereby accounting for the improved performance of VGGfold for protein fold recognition. In summary, this study is of great significance for both understanding the working principle of DCNNs in protein fold recognition and exploring the relationship between the predicted protein contact map and protein tertiary structure. This proposed visualization method is flexible and applicable to address other DCNN-based bioinformatics and computational biology questions. The online web server of VGGfold is freely available at http://csbio.njust.edu.cn/bioinf/vggfold/.
Yan Liu 0038, Yiheng Zhu 0001, Xiaoning Song, Jiangning Song, Dongjun Yu
Briefings Bioinform.3
2021 Impact of Atmospheric Correction on Spatial Heterogeneity Relations Between Land Surface Temperature and Biophysical Compositions
abstract
Investigating the relations between land surface temperature (LST) and biophysical compositions can help the understanding of the surface biophysical process. However, there are still uncertainties in determining the impacts of biophysical compositions on LST due to the atmospheric effects. In this article, four atmospheric correction algorithms were used to correct 12 Landsat 8 images in Xi'an, Beijing, Wuhan, and Guangzhou, China, including the Atmospheric Correction for Flat Terrain (ATCOR2), Quick Atmospheric Correction (QUAC), Fast Line-of-sight Atmospheric Analysis of Spectral Hypercube (FLAASH), and Second Simulation of Satellite Signal in the Solar Spectrum (6S). Then, geodetector was used to investigate the atmospheric correction differences in the spatial heterogeneity relationships between LST and normalized difference vegetation index (NDVI), normalized difference built-up index (NDBI), and bare soil index (BSI). Results indicate that the selected composition factors were greatly improved after atmospheric correction, and the relations between LST and three factors were characterized by obvious atmospheric correction differences in four study areas. On the whole, the 6S algorithm performed the best in improving the factor values and impacting the spatial heterogeneity relations between LST and biophysical compositions, followed by FLAASH, QUAC, and ATCOR2 algorithms. Except for Wuhan, 6S, FLAASH, and QUAC algorithms significantly enhanced the correlation between LST and NDVI. However, all algorithms weakened the correlations between LST, NDVI, and BSI, except Guangzhou. These findings have been verified using the regression analysis. In addition, with geodetector, combinations of any two composition factors all had strongly enhanced impacts on LST, and a combination between NDVI and NDBI performed the strongest in most cases.
Xin-Ming Zhu, Xiaoning Song, Pei Leng, Da Guo, Shuohao Cai
IEEE Trans. Geosci. Remote. Sens.2
2021 From RGB to Depth: Domain Transfer Network for Face Anti-Spoofing
Yahang Wang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.2
2021 SP-GAN: Self-Growing and Pruning Generative Adversarial Networks
abstract
This article presents a new Self-growing and Pruning Generative Adversarial Network (SP-GAN) for realistic image generation. In contrast to traditional GAN models, our SP-GAN is able to dynamically adjust the size and architecture of a network in the training stage by using the proposed self-growing and pruning mechanisms. To be more specific, we first train two seed networks as the generator and discriminator; each contains a small number of convolution kernels. Such small-scale networks are much easier and faster to train than large-capacity networks. Second, in the self-growing step, we replicate the convolution kernels of each seed network to augment the scale of the network, followed by fine-tuning the augmented/expanded network. More importantly, to prevent the excessive growth of each seed network in the self-growing stage, we propose a pruning strategy that reduces the redundancy of an augmented network, yielding the optimal scale of the network. Finally, we design a new adaptive loss function that is treated as a variable loss computational process for the training of the proposed SP-GAN model. By design, the hyperparameters of the loss function can dynamically adapt to different training stages. Experimental results obtained on a set of data sets demonstrate the merits of the proposed method, especially in terms of the stability and efficiency of network training. The source code of the proposed SP-GAN method is publicly available at https://github.com/Lambert-chen/SPGAN.git.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Dongjun Yu, Xiaojun Wu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2019 A Physical Method for Retrieving Microwave Land Surface Emissivity under all-Weather Conditions
abstract
Microwave land surface emissivity (LSE) is an important parameter for retrievals of land surface and atmospheric characteristics, and it is also crucial as an input parameter for numerical weather prediction model data assimilation. This study develops a method for retrieving all-weather LSE over China based on radiative transfer model through reconstructing spatial-temporal continuous land surface temperature (LST) data using China Land Data Assimilation System (CLDAS). Atmospheric effect is also removed with the relationships among atmospheric transmittances, atmospheric effective radiating temperature, precipitable water vapor (PWV), and cloud liquid water (CLW). The retrieved LSE are preliminarily validated by the simulations of Community Radiative Transfer Model (CRTM) and two LSEs show a determined parameter (R2) of 0.81 at 18.7 GHz.
Fang-Cheng Zhou, Shihao Tang, Hua Wu 0001, Zhao-Liang Li, Xiaoning Song, Xiuzhen Han, Shengli Wu 0002
IGARSS5
2019 Collaborative representation based face classification exploiting block weighted LBP and analysis dictionary learning
Xiaoning Song, Youming Chen, Zhenhua Feng 0001, Guosheng Hu, Tao Zhang 0010, Xiaojun Wu 0001
Pattern Recognit.1
2019 Fast SRC using quadratic optimisation in downsized coefficient solution subspace
Xiaoning Song, Guosheng Hu, Jian-Hao Luo, Zhenhua Feng 0001, Dongjun Yu, Xiaojun Wu 0001
Signal Process.1
2018 Application of the GF Satellite Data in Flood Disaster Monitoring
abstract
This article gives a detailed description of the High-resolution Earth Observation System, which is currently being carried out in China. The GF data have been successfully launched and their characteristics in flood disaster monitoring are analyzed. Followed by the introduction of the remote monitoring system of flood disaster has been completed. At last some examples of flood disaster monitoring in recent years based on the GF data are introduced.
Xiaotao Li, Jingxuan Lu, Xiaoning Song, Yayong Sun, Lin Li 0061, Tianjie Lei
IGARSS3
2018 Semi-supervised dictionary learning via local sparse constraints for violence detection
Tao Zhang 0010, Wenjing Jia, Chen Gong 0002, Jun Sun 0008, Xiaoning Song
Pattern Recognit. Lett.5
2018 Surface Soil Moisture Retrieval Using Optical/Thermal Infrared Remote Sensing Data
abstract
Surface soil moisture (SSM) plays significant roles in various scientific fields, including agriculture, hydrology, meteorology, and ecology. However, the spatial resolutions of microwave SSM products are too coarse for regional applications. Most current optical/thermal infrared SSM retrieval models cannot directly estimate the quantitative volumetric soil water content without establishing empirical relationships between ground-based SSM measurements and satellite-derived proxies of SSM. Therefore, in this paper, SSM is estimated directly from 5-km-resolution Chinese Geostationary Meteorological Satellite FY-2E data based on an elliptical-new SSM retrieval model developed from the synergistic use of diurnal cycles of land surface temperature (LST) and net surface shortwave radiation (NSSR). The elliptical-original model was constructed for bare soil and did not consider the impacts of different fractional vegetation cover (FVC) conditions. To optimize the elliptical-original model for regional-scale SSM estimates, it is improved in this paper by considering the influence of FVC, which is based on a dimidiate pixel model and a Moderate Resolution Imaging Spectroradiometer normalized difference vegetation index product. A preliminary validation of the model is conducted based on ground measurements from the counties of Maqu, Luqu, and Ruoergai in the source area of the Yellow River. A correlation coefficient (R) of 0.620, a root-mean-square error (RMSE) of 0.146 m3/m3, and a bias of 0.038 m3/m3were obtained when comparing the in situ measurements with the FY-2E-derived SSM using the elliptical-original model. In contrast, the FY-2E-derived SSM using the elliptical-new model exhibited greater consistency with the ground measurements, as evidenced by an R of 0.845, an RMSE of 0.064 m3/m3, and a bias of 0.017 m3/m3. To provide accurate SSM estimates, high-accuracy FVC, LST, and NSSR data are required. To complement the point-scale validation conducted here, cross-comparisons with other existing SSM products will be conducted in the future studies.
Yawei Wang 0001, Jian Peng 0006, Xiaoning Song, Pei Leng, Ralf Ludwig, Alexander Loew
IEEE Trans. Geosci. Remote. Sens.3
2018 Dictionary Integration Using 3D Morphable Face Models for Pose-Invariant Collaborative-Representation-Based Classification
abstract
The paper presents a dictionary integration algorithm using 3D morphable face models (3DMM) for pose-invariant collaborative-representation-based face classification. To this end, we first fit a 3DMM to the 2D face images of a dictionary to reconstruct the 3D shape and texture of each image. The 3D faces are used to render a number of virtual 2D face images with arbitrary pose variations to augment the training data, by merging the original and rendered virtual samples to create an extended dictionary. Second, to reduce the information redundancy of the extended dictionary and improve the sparsity of reconstruction coefficient vectors using collaborative-representation-based classification (CRC), we exploit an on-line class elimination scheme to optimise the extended dictionary by identifying the training samples of the most representative classes for a given query. The final goal is to perform pose-invariant face classification using the proposed dictionary integration method and the on-line pruning strategy under the CRC framework. Experimental results obtained for a set of well-known face data sets demonstrate the merits of the proposed method, especially its robustness to pose variations.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.1
2017 An algorithm for retrieving land surface temperature from AMSR-E data over the desert regions
abstract
Land surface temperature is an important driving force in the exchange of water, heat, and even CO2at the surface-atmosphere interface in the desert regions. The rapid and continuous measurements of land surface temperature are meaningful to the ecological and environmental researches. A physically based single-frequency and double-polarization algorithm for retrieving land surface temperature is developed in this study. The 18.7 GHz vertically polarized emissivities are firstly estimated from the Polarization Ratio (PR, defined as the ratio of the horizontal to vertical brightness temperature at the same frequency) at 18.7 GHz. And then the estimated emissivities can be directly used to retrieve land surface temperature without considering the atmospheric effect. A preliminary validation is done in the Taklimakan desert. The retrieved land surface temperatures are compared to the infrared land surface temperature products for all the year of 2007 with a Root Mean Square Error (RMSE) of 3.05 K.
Fang-Cheng Zhou, Zhao-Liang Li, Hua Wu 0001, Bo-Hui Tang, Ronglin Tang, Xiaoning Song, Guangjian Yan, Sibo Duan
IGARSS6
2017 Dynamic dictionary optimization for sparse-representation-based face classification using local difference images
Chang-Bin Shao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Yuhui Zheng
Inf. Sci.2
2017 Estimating Vegetation Water Content of Corn and Soybean Using Different Polarization Ratios Based on L- and S-Band Radar Data
abstract
Vegetation water content (VWC) is an important parameter of agriculture and forestry. In this letter, specific polarization ratios were evaluated for estimating VWC of corn and soybean. Backscattering coefficients (σhh, σvv, σvhand σhv), polarization ratios (σhh/σvv,σvv/σvh, and σhh/σhv), and the radar vegetation index derived from L-band (1.26 GHz) and S-band (3.15 GHz) radar data of the passive and active Land S-band sensor (PALS) in Soil Moisture Experiments 2002 were implemented to develop various linear relationship models with field VWC measurements for corn and soybean, respectively. L-band σhh/σvvwas found to be most correlated with corn VWC (R = 0.81), while for soybean, L-band σhh/σhvwas the best parameter to estimate VWC with an R of 0.90. Based upon these analyses, prediction equations for the estimation of corn and soybean VWC using the polarization ratios were developed. Results indicated that L-band σhh/σvvwas able to estimate corn VWC with a root mean square error (RMSE) of 0.53 kg/m2and a mean absolute relative error (MARE) of 11.48%. As for soybean, L-band σhh/σhvwas capable of estimating soybean VWC with an RMSE of 0.12 kg/m2and an MARE of 13.33%. The main reason for these differences is most likely due to the disparate structure features and VWC distribution of corn and soybean. This letter proposes an effective method for acquiring VWC in regional areas, and it is also considered to be a powerful supplement for the current methods based on optical remotely sensed data.
Jianwei Ma 0003, Shifeng Huang 0003, Jiren Li, Xiaotao Li, Xiaoning Song, Pei Leng, Yayong Sun
IEEE Geosci. Remote. Sens. Lett.5
2017 Converted-face identification: using synthesized images to replace original images for recognition
Chang-Bin Shao, Xiaoning Song, Xin Shu 0001, Xiaojun Wu 0001
Multim. Tools Appl.2
2017 Sparse representation-based classification using generalized weighted extended dictionary
Xiaoning Song, Chang-Bin Shao, Xibei Yang, Xiaojun Wu 0001
Soft Comput.1
2017 Half-Face Dictionary Integration for Representation-Based Classification
abstract
This paper presents a half-face dictionary integration (HFDI) algorithm for representation-based classification. The proposed HFDI algorithm measures residuals between an input signal and the reconstructed one, using both the original and the synthesized dual-column (row) half-face training samples. More specifically, we first generate a set of virtual half-face samples for the purpose of training data augmentation. The aim is to obtain high-fidelity collaborative representation of a test sample. In this half-face integrated dictionary, each original training vector is replaced by an integrated dual-column (row) half-face matrix. Second, to reduce the redundancy between the original dictionary and the extended half-face dictionary, we propose an elimination strategy to gain the most robust training atoms. The last contribution of the proposed HFDI method is the use of a competitive fusion method weighting the reconstruction residuals from different dictionaries for robust face classification. Experimental results obtained from the Facial Recognition Technology, Aleix and Robert, Georgia Tech, ORL, and Carnegie Mellon University-pose, illumination and expression data sets demonstrate the effectiveness of the proposed method, especially in the case of the small sample size problem.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Xiaojun Wu 0001
IEEE Trans. Cybern.1
2016 Estimating soil moisture in the agricultural areas using RADARSAT-2 Quad-olarization SAR data
abstract
The aim of this study was to estimate soil moisture from RADARSAT-2 Quad-polarization Synthetic Aperture Radar (SAR) data acquired in the agricultural areas. The adopted approach is based on the combination of semi-empirical Water-Cloud model and Dubois model. Firstly, VH backscattering coefficient was used to develop empirical relationship for crop water content estimation. Secondly, crop water content was then used to correct the semi-empirical Water-Cloud model for vegetation effects in order to get the VV and HH backscattering coefficient of soil surface in the absence of vegetation cover. Thirdly, the soil moisture was retrieved based on Dubois model using VV and HH backscattering coefficient of soil surface in the absence of vegetation cover. Finally, the soil moisture retrieved is evaluated over wheat crop fields using ground measurements. This paper proposes an effective method for acquiring soil moisture in the agricultural areas under any weather conditions.
Jianwei Ma 0003, Shifeng Huang 0003, Jiren Li, Xiaotao Li, Xiaoning Song, Pei Leng, Yayong Sun, Tianjie Lei
IGARSS5
2016 Estimation of surface soil moisture using FengYun-2E (FY-2E) data: A case study over the source area of the Yellow River
abstract
Surface soil moisture (SSM) is a significant variable in various fields of science. This paper aims to analyze and improve a SSM retrieval model to apply it to estimate SSM at the regional scale. Firstly, the model parameters were been analyzed. The rotation angle was transformed into exponential form and the ellipse center horizontal coordinate was decreased for the improved SSM retrieval model. After validation with the simulated data from Common Land Model (CoLM), the result indicated that the accuracy of improved model was not lower than the original one after one model coefficient removed. Subsequently, regional SSM was mapped with improved model from FengYun-2E (FY-2E) observation. In addition, the improved SSM retrieval model showed a good consistency with Climate Change Initiative soil moisture (CCI SM) product. Ultimately, a preliminary validation was conducted using the ground measurements in the source area of the Yellow River (SAYR). The result presented an R of 0.53, a RMSE of 0.06 m3/m3and a bias of 0.03 m3/m3.
Yawei Wang 0001, Xiaoning Song, Pei Leng
IGARSS2
2016 An algorithm for retrieving instantaneous microwave land surface emissivity from passive microwave brightness temperature and precipitable water vapor data
abstract
An algorithm has been developed for retrieving instantaneous microwave land surface emissivity using brightness temperature and precipitable water vapor data. Unlike previous algorithms, the new technique does not need infrared land surface temperature as the input data, and overcomes the limitation of previous algorithms under cloudy conditions. Compared with the values from physical retrieval algorithm, the result demonstrates that this new algorithm has a Root Mean Square Error of 0.038 and a bias of 0.012. Although the accuracy is worse than 1%, this new algorithm presents the potential to obtain the instantaneous microwave land surface emissivity under both cloud-free and cloudy conditions, which can be applied in some weather prediction models.
Fang-Cheng Zhou, Zhao-Liang Li, Hua Wu 0001, Bo-Hui Tang, Ronglin Tang, Xiaoning Song, Guangjian Yan
IGARSS6
2016 Towards multi-scale fuzzy sparse discriminant analysis using local third-order tensor model of face images
Xiaoning Song, Zhenhua Feng 0001, Xibei Yang, Xiaojun Wu 0001, Jing-Yu Yang 0001
Neurocomputing1
2016 Decision-theoretic rough set: A multicost strategy
Huili Dou, Xibei Yang, Xiaoning Song, Hualong Yu, Weizhi Wu 0001, Jing-Yu Yang 0001
Knowl. Based Syst.3
2016 Extended minimum-squared error algorithm for robust face recognition via auxiliary mirror samples
Chang-Bin Shao, Xiaoning Song, Xibei Yang, Xiaojun Wu 0001
Soft Comput.2
2015 A novel SRC fusion method using hierarchical multi-scale LBP and greedy search strategy
Zi Liu, Xiaoning Song, Zhenmin Tang
Neurocomputing2
2015 Fusing hierarchical multi-scale local binary patterns and virtual mirror samples to perform face recognition
Zi Liu, Xiaoning Song, Zhenmin Tang
Neural Comput. Appl.2
2015 Using idea of three-step sparse residuals measurement to perform discriminant analysis
Xiaoning Song, Zi Liu, Jing-Yu Yang 0001, Xiaojun Wu 0001
Soft Comput.1
2014 A parameterized fuzzy adaptive K-SVD approach for the multi-classes study of pursuit algorithms
Xiaoning Song, Zi Liu, Xibei Yang, Jing-Yu Yang 0001
Neurocomputing1
2014 Updating multigranulation rough approximations with increasing of granular structures
Xibei Yang, Yong Qi 0002, Hualong Yu, Xiaoning Song, Jing-Yu Yang 0001
Knowl. Based Syst.4
2014 A new sparse representation-based classification algorithm using iterative class elimination
Xiaoning Song, Zi Liu, Xibei Yang, Shang Gao 0001
Neural Comput. Appl.1
2014 Constructive and axiomatic approaches to hesitant fuzzy rough set
Xibei Yang, Xiaoning Song, Yunsong Qi, Jing-Yu Yang 0001
Soft Comput.2
2013 Test cost sensitive multigranulation rough set: Model and minimal cost selection
Xibei Yang, Yunsong Qi, Xiaoning Song, Jing-Yu Yang 0001
Inf. Sci.3
2011 An optimal symmetrical null space criterion of Fisher discriminant for feature extraction and recognition
Xiaoning Song, Jing-Yu Yang 0001, Xiaojun Wu 0001, Xibei Yang
Soft Comput.1
2010 Discriminant analysis approach using fuzzy fourfold subspaces model
Xiaoning Song, Xibei Yang, Jing-Yu Yang 0001, Xiaojun Wu 0001, Yu-Jie Zheng
Neurocomputing1
2009 Difference Relation-Based Rough Set and Negative Rules in Incomplete Information System
abstract
The purpose of this paper is to present a new rough set model for generating negative rules from the incomplete information system. A negative rule indicates that if an object does not satisfy the attribute-value pairs in the condition part, then we can exclude the decision part from such object. The proposed rough set model is constructed on the basis of a difference relation. Such difference relation is a binary relation without any constraints. Moreover, to simplify the negative rules generated from the difference relation-based rough approximations, the concepts of lower, upper approximate and rough reducts are also proposed. Some numerical examples are employed to substantiate the conceptual arguments.
Xibei Yang, Dongjun Yu, Jing-Yu Yang 0001, Xiaoning Song
Int. J. Uncertain. Fuzziness Knowl. Based Syst.4
2009 Credible rules in incomplete decision system based on descriptors
Xibei Yang, Xiaoning Song, Jing-Yu Yang 0001
Knowl. Based Syst.3
2007 The study of wetlands change in Yellow River delta Based on RS and GIS
abstract
In recent years, the water and sediment were dropped significantly in Yellow River, the average annual water and sediment discharge was 143.5 × 108 m3 and 3.75 × 108 t respectively, accounting for 29.2% and 31.2% of the 1950-1969 values. In addition, the day of flow depletion was also increased continuously and there was 21 years of flow depletion from 1972 to 1998. Especially in 1997 the day of flow depletion was up to 226 days in the section below Lijin Hydraulic Station. In the new hydrological environment the Yellow River delta eco- environment must be changed. Based on Remote Sensing and Geographic Information System (RS & GIS)(RS & GIS) technology, we study what and how the Yellow River delta wetlands changed in the new hydrologic conditions in recent years using the Landsat TM/ETM+ image of 1996, 2000 and 2004. The results show that the area of the wetlands in Yellow River Delta was increased and from 1996 to 2004 the area of the wetlands has been increased 22838.55 hm2and the increase rate was 2854.82hm2/y. The area of the nature wetlands and the artificial wetlands were increased as a whole, but their change character was different. The area of the nature wetlands was increased quickly from 1996 to 2000, which is related to that the national park of Yellow River delta was changed in 1996 and from 2000 to 2004 it was increased tardily. On the contrary, the area of the artificial wetlands was decreased tardily from 1996 to 2000 and from 2000 to 2004 it was increased quickly. In the mean time the area of river, bottomland and paddy field was decreased and the area of swampland, reed marsh, reservoir, saline and acequia was increased. To scrutinize the dynamic changes of wetlands, the simulation for transformation of the wetland types has been carried out by means of the ecological Markov Process. The results showed that the whole transfer velocity of the wetlands is 2.99% from 1996 to 2000 and the transfer among the wetland types is simple, but from 2000 to 2004 the transfer velocity of the wetlands is 6.95% and the transfer among the different wetland types is complex correspondingly. Moreover the reason of the wetlands changing has been analyzed in this paper. Firstly, the reciprocity between the decrease of the water and sediment and the regulating water and sediment is the main reason of the wetlands change. Aiming at the decrease of the water and sediment entering the Yellow River, the water and sediment regulation is an important measure which has positive impact on the Yellow River Delta Wetlands. Secondly, the process of the agriculture exploitation, city building and wetlands protection is also an important factor of the wetlands change. In order to resolve the problems of the excessive exploitation, wetlands pollution and so on, many protection measures, such as strengthening the management of the Yellow River, doing better program and dynamic monitoring, keeping harmony between the exploitation and recovery of the wetlands and paying attention to the wetland pollution, were carried out. Thirdly, the climate change and storm surge disaster are explicit factors of the wetlands change. This study has significance for building and protecting of eco-environment in Yellow River delta.
Xiaotao Li, Shifeng Huang 0003, Jiren Li, Mei Xu, Xiaoning Song
IGARSS5
2007 Vegetation water inversion using MODIS satellite data
abstract
The vegetation water condition is determined by soil water content. So soil water condition can be reflected by means of vegetation water condition indirectly. There have been many detailed studies on crop water stress index from the micrometeoological aspect before, for example, CWSI (crop water stress index) and crop water stress microclimatic model based on CWSI. The generation of these indices need many meteorological data and is not suitable to regional water stress study because of the difficulties in limitation of data acquisition of related data on the surface. Moreover, there are a few relevant studies using remote sensing techniques, such as WDI (water deficit index), VCI (vegetation condition index) and TCI (temperature condition index). Although these indices are suitable to regional study, they utilize the statistic value of remote sensing data for years and they are on the basis of pixel scale. Therefore, the precision of quantitatively assessing surface water is limited to some extent. In view of problems of these water stress indices, this article is going to discuss a not only simple but also reasonable method to derive vegetation water. Different vegetation water condition could cause the variation of vegetation spectrum, so vegetation water stress can be shown by remote sensing vegetation indices, such as vegetation condition index and anomaly vegetation index et al. In view of this, soil water content can be assessed indirectly. But vegetation indices neglect some environment factors, eg. temperature and precipitation. The less evaportranspiration, the higher vegetation and soil temperatures. Therefore the vegetation temperature is a direct indicator of vegetation water stressed and drought. In a word, the soil water is positive correlation with vegetation index and negative correlation with temperature. It takes the North-West semi-arid area - Xilingole district of Inner Mongolia as study area. Different degraded grasslands are selected as objects in this study. MODIS (moderate resolution imaging spectroradiometer) has thirty-six bands of visible/near-infrared and thermal infrared with abundant information. This study selects MODIS as data source. In consideration of the spectral character of MODIS data and different degraded grasslands, the reflective spectrum of vegetation is greatly affected by soil. MSAVI (modified soil- adjusted vegetation index) is selected to weaken the disturbanceof soil information. MSAVI and NDWI (Normalized Difference Water Index) are deduced using one visible band (0.66 mum) and two near-infrared bands (0.86 mum, 1.24 mum). In order to estimate the vegetation water more accurately. Vegetation component temperature is inversed using two thermal infrared bands (8.6 mum, 11 mum) according to the emissivity distribution of vegetation with the wavelength from eight to twelve micrometer and the correlation analysis of MODIS data. Sequentially, the vegetation water synthesis index is acquired by analyzing the coupling character of three indexes, which can reflect the vegetation water condition. Then the vegetation water content can be extracted using the synthesis index effectively. Lastly, the vegetation water content is validated using measured data. The matching results show that the synthesis index is directly proportional to the measured data. It proves that the synthesis index and the vegetation water content are credible and the method is reasonable. This study has discussed a new method to know the regional vegetation water condition from satellite remote sensing data directly and quickly.
Xiaoning Song, Qinhuo Liu, Xiaotao Li
IGARSS1
2005 An improved method of sensible heat flux
Xiaoning Song, Yingshi Zhao, Xiaoming Feng
IGARSS1
2004 Application of 3D remote sensing technique in tour resources investigation
abstract
Remotely sensed imagery can reflect characteristics of sceneries, such as the formation, structure and spatial correlation etc, in order to make people know the surface matters form a macroscopic and holistic view. However, two dimension plane showed in remote sensing image could not get the 3D space in the realistic world reappearance. As the 3D visualization technique is used wider and wider, the visualization processing on the basis of remote sensing technique not only ensures the realization of the real-time 3D visualization, but also is better for the interpretation of remote sensing image. This article discusses a simple and feasible way of recurring the 3D remote sensing image taking Mount Taishan as an example. At the same time, by means of intuitionistic character of the 3D image, the characteristic of geology structures, physiognomy and environment in this area could be evaluated combining with other correlated data. In the end, the types of physiognomy landscape are divided and the distribution character is depicted. This study enhances the practicability of remote sensing technique, as well as, it takes positive effect on the planning and development of tour resources of Mount Taishan
Xiaotao Li, Xiaoning Song, Fengjie Yang
IGARSS2
2004 A simplified surface albedo inverse model with MODIS data
abstract
Land surface albedo, which is a fundamental component needed for determining the radiation balance of the Earth-atmosphere system, is the surface hemispherical reflectivity integrated over the solar spectrum. Briefly said, it is the ratio of the total upwelling irradiance to the total downward irradiance. Given the limit of experiment field and instrument equipment, it is very difficult to obtain a great deal of surface measured data. Thereof, in our research, from the point of view that surface albedos are the ratio of flux, surface albedos can be reversed with less calibrated field synchronous observed data. During the conversion process of surface albedos, the first is to obtain surface reflectance from remotely sensed data, which can be completed by running 6S atmospheric radiative transfer model; the second is to reverse surface narrowband albedos, which is fulfilled by MODTRAN model. The outputs of MODTRAN simulations are upwelling and downwelling irradiance. The sensor spectral response functions are integrated with these outputs to generate narrowband albedos; the third is to calculate surface broadband albedos, which is focused on determining the converting coefficients from narrowbands to broadbands under the linear relationship between narrowband albedos and broadband albedos. The spectral distribution of solar irradiance at the surface is set as the weighting function for converting the narrowband albedos to broadband albedos in the research. The conversion of broadband albedos is fulfilled finally. The simplified surface broadband albedo reverse model is tested for the west of Inner Mongolia using MODIS image data and synchronous observed surface spectral data based on the above conversion algorithm. The results show that the method is convenient and feasible under a specific surface/atmospheric condition.
Yingshi Zhao, Xiaoning Song
IGARSS3
2004 Cloud detection and analysis of MODIS image
abstract
MODIS (Moderate Resolution Imaging Spectroradiometer) is a kind of new weather satellite data. Few weather satellite images obtained are all clear sky and they are always influenced by cloud more or less. Cloud is a large obstacle to remote sensing image processing and analysis all the while. In order to extract objective information more effective, cloud should be removed from the remote sensing images, which is an essential sector in the image preprocessing. Cloud detection is the most important processing before removing cloud. Taking it into account that MODIS data includes thirty-six bands, especially the infrared channels subdivided, it has realized cloud detection in MODIS images by multi-spectral synthesis method and cloud detection index in this paper. Owing to the limitation to a certainty of the above methods, an automatic cloud detection algorithm is applied based on the spatial texture analysis and neural network in this research. At last the cloud detection results gained by different ways are testified each other and analyzed by comparison. It found that the results are consistent, which shows that the cloud-contaminate pixels are detected successfully. It not only lays a good foundation for the cloud removing, but also can improve the precision of remote sensing image recognition, classification and inverse in this study.
Xiaoning Song, Yingshi Zhao
IGARSS1