Junqiang Song

dblp:72/4489 · DBLP profile ↗
← Back
45ranked-venue papers
0as first author
17since 2021 · last 2026
0009-0003-2686-566XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 11 since 2021Systems, architecture and hardware · 12Artificial intelligence and machine learning · 10 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KG-ART: Dual-Track Adversarial Reasoning for Knowledge Graph Question Answering
Yanan Guo 0006, Junqiang Song, Fukang Yin, Hongze Leng
KSEM (3)2
2026 Review on deep learning quantitative precipitation nowcasting: Advances and challenges
Jingnan Wang, Kefeng Deng, Di Zhang 0021, Chengwu Zhao, Hongze Leng, Yingfang Wen, Yudi Liu, Kaijun Ren, Junqiang Song
Expert Syst. Appl.10
2026 A Physics-Guided Hierarchical Transformer Framework for Sea Surface Temperature Forecasting and Marine Heatwave Detection
abstract
Accurate forecasting of sea surface temperature (SST) and early detection of marine heatwave (MHW) events are critical yet challenging tasks due to the complex, nonlinear, and multiscale dynamics of the ocean-atmosphere system. To address these challenges, we propose a novel physics-guided hierarchical Transformer framework that combines deep spatiotemporal learning with physical process constraints. The architecture integrates a U-Net-style encoder-decoder with a Temporal-Spatial Predictor (TSP) module. It introduces a physics-constrained branch based on the mixed-layer heat budget equation, enhancing physical consistency and interpretability. A data-driven anomaly compensation mechanism is further employed to adaptively fuse physically-derived predictions with complex dynamic corrections through a learnable weighting scheme. This dual-stream architecture enables robust multi-step rolling forecasting and accurate detection of both gradual SST trends and abrupt MHW events. Extensive experiments on high-resolution SST datasets show that our model significantly outperforms state-of-the-art deep learning baselines such as ConvLSTM, DeepONet, FNO, and hybrid CNN-Transformer models across various performance metrics. These results highlight the framework’s ability to bridge physical oceanography and modern AI, providing a powerful tool for operational ocean forecasting and climate risk assessment.
Yanan Guo 0006, Junqiang Song, Hongze Leng
IEEE Geosci. Remote. Sens. Lett.2
2026 A Novel Conditional Diffusion-Based Framework for Advanced Reconstruction of Cloud Vertical Structure
abstract
This study presents a framework based on conditional diffusion probabilistic models for reconstructing vertical cloud structures from passive satellite remote sensing observations. The retrieval is formulated as a conditional denoising diffusion process, in which randomly initialized latent fields are progressively refined through iterative sampling guided by Moderate Resolution Imaging Spectroradiometer (MODIS) measurements. Compared with generative adversarial network (GAN)-based approaches, the proposed model more effectively captures the intrinsic variability, multiscale organization, and stochastic nature of atmospheric cloud fields. It exhibits superior performance in reconstructing complex multilayer systems, intense convective structures, and mesoscale cloud features. Quantitative evaluation against CloudSat radar reflectivity data demonstrates that the method consistently attains high structural similarity and accurately reproduces the vertical distribution of cloud reflectivity. These findings indicate that conditional diffusion probabilistic models provide a novel generative modeling paradigm for atmospheric remote sensing, providing physically consistent reconstructions, quantitative uncertainty characterization, and robust generalization to diverse atmospheric regimes. Furthermore, the framework can be extended to the three-dimensional reconstruction of other meteorological variables, supporting broader application of generative AI in satellite-based atmospheric analyses.
Yanan Guo 0006, Junqiang Song, Chuanfeng Zhao, Hongze Leng
IEEE Geosci. Remote. Sens. Lett.2
2026 PGMNO: A physics-Guided mamba neural operator framework for partial differential equations
Yanan Guo 0006, Junqiang Song, Chuanfeng Zhao, Fukang Yin, Hongze Leng
Neural Networks2
2025 Data-Driven Super-Resolution Reconstruction of Quasi-Geostrophic Turbulence Via Enhanced Diffusion Model and Fourier Neural Operator
abstract
High-fidelity simulation and reconstruction of physical fields are essential in both scientific research and engineering, yet classical solvers can be prohibitively expensive at resolutions needed to capture fine-scale structures. We propose FNODiffSR, a data-driven super-resolution framework that couples a residual-guided diffusion model with an Adaptive Weighted Fourier Neural Operator (AWFNO). AWFNO models longrange spectral dependencies while selectively emphasizing highfrequency components, and the diffusion module employs a conditional probability-flow ODE instead of stochastic sampling to deterministically bridge low- and high-fidelity representations. Final reconstructions are obtained by integrating this ODE with an adaptive time-stepping solver. Experiments on quasigeostrophic turbulence across varied upsampling and sparsesampling regimes show that FNODiffSR consistently surpasses interpolation and learning-based baselines in reconstruction fidelity, structural similarity, and physical consistency (as assessed by a dimensionless equation-residual), while offering predictable runtime and scalability. These qualities make FNODiffSR a strong candidate for high-quality scientific data recovery and downstream analysis.
Yanan Guo 0006, Junqiang Song, Hongze Leng
ICDM2
2025 DRL-EnVar: an adaptive hybrid ensemble-variational data assimilation method based on deep reinforcement learning
abstract
Accurate estimation of the background error covariance matrix denoted as B remains a critical challenge in numerical weather prediction (NWP), directly influencing data assimilation (DA) performance and forecast accuracy. Although hybrid ensemble-variational (EnVar) methods combine static and flow-dependent matrices to improve assimilation, their effectiveness is constrained by empirically fixed weights. To address this limitation, we propose DRL-EnVar, an adaptive hybrid EnVar DA method enhanced with deep reinforcement learning. DRL-EnVar integrates deep learning (DL) components, including a novel cyclic convolution module to extract abstract features from data, and employs reinforcement learning (RL) to dynamically optimize hybrid weighting strategies. The system adaptively combines multiple ensemble-based flow-dependent matrices with one or more static matrices to construct a time-varying hybrid matrix B that better reflects real-time background errors. Experimental results demonstrate that DRL-EnVar performs better than the traditional ensemble Kalman filter (EnKF) and hybrid covariance DA (HCDA) methods, especially under sparse observations or transitional changes in state variables. It achieves competitive or superior assimilation accuracy with lower computational cost, and can be flexibly integrated into both three-dimensional variational assimilation (3DVar) and four-dimensional variational assimilation (4DVar) frameworks. Overall, DRL-EnVar offers a novel and efficient approach to adaptive DA, particularly valuable for improving forecast skill during transitional weather regimes.
Lilan Huang, Hongze Leng, Junqiang Song, Wuxin Wang, Ruisheng Hu
Frontiers Inf. Technol. Electron. Eng.3
2025 Skywave OTHR Full-Link Modeling and Simulation - Part I: Trans-Ionospheric Sea Clutter
abstract
Over-the-horizon radar (OTHR) utilizes the ionospheric refraction and reflection in high-frequency band for air-sea targets detection. However, non-stationary ionospheric dynamics induce inhomogeneous distortions in sea clutter and targets, significantly degrading detection and localization performance. To address this challenge, we present a comprehensive investigation on OTHR full-link modeling and simulation to systematically analyze the impact of various trans-ionospheric propagation effects on echo signals. In Part I, we develop a unified framework for full-link modeling of sea clutter that incorporate background ionospheric and oceanic conditions, enabling simulation and analysis of sea clutter characteristics. Firstly, we establish the models for radar cross section of sea clutter and ionospheric propagation effects. Key parameters, including the wave propagation range, actual range, group delay, phase disturbance, and propagation loss, are calculated based on the Appleton-Hartree formula and ray tracing technique. Secondly, we construct the echo signal models in the fast- and slow-time domain and derive the corresponding range-Doppler spectrum. The origin mechanism and intrinsic cause of Doppler shifting, broadening, and splitting, as well as range localization errors are theoretically analyzed. Finally, simulation experiments of three scenarios are designed to produce OTHR sea clutter data in sea and air modes, which are validated in comparison with the real data. The typical phenomena of Doppler shifting, broadening, and splitting observed in real data are reproduced, and the results indicate that the multi-mode propagation is the main obstacle to range localization and ionospheric decontamination. The full-link sea clutter model provides critical insights for the subsequent signal processing tasks including ionospheric decontamination, clutter suppression, target detection, and localization.
Yifei Ji, Zhen Dong 0001, Feixiang Tang, Weijian Liu 0001, Alei Chen, Ming Ou, Junqiang Song
IEEE Trans. Geosci. Remote. Sens.10
2025 Precipitation Nowcasting Diffusion Model Based on Fluid Dynamics and Multisource Data
abstract
Precipitation nowcasting is a long-standing challenge due to the inherent unpredictability, which often lead to significant risks and damage. Traditional approaches that model nonlinear relationships between initial and future precipitation states often fail to accurately capture precipitation dynamics, including distribution and intensity patterns. Current data-driven methods are limited in their ability to represent the chaotic nature of precipitation without guidance from physical theory. To address this, we present Rainfusion, a generative model that integrates Prandtl’s mixing length theory from fluid dynamics with computer vision diffusion models. This integration accounts for nonlinear interactions between large-scale evolution and turbulent fluctuations in precipitation, generating physically plausible predictions. Rainfusion significantly improves forecasting skill on two benchmark dataset over the next 3 hours. Furthermore, we enhance Rainfusion with a control network trained on multi-source data, particularly lightning observations, enabling more accurate and controllable predictions of precipitation’s spatial-temporal patterns. Weather forecasters can utilize Rainfusion to guide predictions toward either growth or decay based on their domain expertise. Our approach advances precipitation nowcasting, offering a robust framework that bridges physical theory with modern deep learning techniques.
Kefeng Deng, Di Zhang 0021, Hongze Leng, Yudi Liu, Kaijun Ren, Junqiang Song
IEEE Trans. Geosci. Remote. Sens.7
2023 DBSA-Net: Dual Branch Self-Attention Network for Underwater Acoustic Signal Denoising
abstract
Underwater acoustic signal denoising is a challenging task due to the complexity of the underwater environment. Most of the existing methods cannot effectively cope with the problem of underwater acoustic signal (UWAS) denoising at low signal-to-noise ratios (SNRs). According to the characteristics of UWAS, a novel idea is proposed to simultaneously model latent features from both the time and frequency dimensions of complex-valued spectrum in a dual-branch self-attention network, namely DBSA-Net. In this model, both magnitude and phase information in the complex spectrum are enhanced from different dimensions by two branches. Specifically, DBSA-Net is an encoder-decoder based network with several global-local-self-attention (GL-SA) blocks distributed on dual branches between encoder and decoder. Each GL-SA block incorporates global self-attention and local self-attention to capture distant context and fine-grained local dependencies along the temporal and frequency dimensions. Moreover, we also design an information interaction module between two branches to exchange complementary information. This interaction module together with a merge block fuse features extracted from different dimensions, thus enhancing the capability of our model to learn the target signal features. Extensive experiments are conducted to evaluate our model on a publicly available dataset. Results of the ablation experiments show that the different modules of DBSA-Net play their respective roles in improving denoising performance and are empirically valid. In both the seen ships and unseen ships scenarios, the proposed DBSA-Net outperforms existing approaches by a large margin on various evaluation metrics.
Aolong Zhou, Wen Zhang 0016, Guojun Xu, Xiaoyong Li 0002, Kefeng Deng, Junqiang Song
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 LPT-QPN: A Lightweight Physics-Informed Transformer for Quantitative Precipitation Nowcasting
abstract
Quantitative precipitation nowcasting (QPN) is a highly challenging task in weather forecasting. The ability to provide precise, immediate, and detailed QPN products is necessary for a variety of situations, including storm warnings, air travel, and large gatherings. To address this challenge, this article proposes a new transformer lightweight physics-informed transformer (LPT)-QPN for QPN tasks, utilizing vertical cumulative liquid water content (VIL) products. This model adopts novel transformer modules to model the long-term evolution of precipitation and incorporates multihead squared attention (MHSA) to model its highly nonlinear relationships while reducing computational complexity. The results of experimental evaluations demonstrate the superiority of LPT-QPN when compared to existing state-of-the-art QPN models. In particular, the LPT-QPN model demonstrates greater accuracy for long lead time and in high-intensity areas, confirmed in both quantitative and qualitative evaluations. In addition, through three customized fine-tuning schemes, we are able to further improve the predictability of the LPT-QPN model for specific precipitation events. By incorporating the physical constraints of the convection-diffusion equation, our approach offers novel perspectives for future explorations that combine physical prior knowledge and deep-learning (DL) techniques.
Kefeng Deng, Di Zhang 0021, Yudi Liu, Hongze Leng, Fukang Yin, Kaijun Ren, Junqiang Song
IEEE Trans. Geosci. Remote. Sens.8
2023 A Novel Cross-Attention Fusion-Based Joint Training Framework for Robust Underwater Acoustic Signal Recognition
abstract
Underwater acoustic signal recognition systems face challenges in achieving high accuracy when processing complex data with low signal-to-noise ratio (SNR) in underwater environments, leading to limited noise robustness. Conventional approaches typically employ pre-trained denoising models for preprocessing noisy signals. However, due to disparate optimization goals between denoising and recognition models, denoising methods might introduce signal distortion, hampering effective enhancement of system accuracy. To address this issue, this paper proposes a novel joint training framework with cross-attention fusion for robust underwater acoustic signal recognition (UASR), called CAF-JT. CAF-JT consists of a denoising module, a recognition module, and the CAF module. It addresses the mismatch problem arising from different optimization directions by jointly training the denoising frontend and the recognition backend. Additionally, inspired by the multi-condition training (MCT) method, the CAF module is designed to fuse characteristics from both denoised and noisy audio, thus incorporating noise information. This fusion mechanism enables the model to better adapt to the characteristics of the noisy environment and enhance its noise robustness. Furthermore, to improve the performance of UASR, TF-Transformer blocks are incorporated into both the denoising module and the recognition module to capture the spatio-temporal distribution of spectral features. The proposed approach is evaluated on two open-source underwater acoustic signal datasets, namely ShipsEar and DeepShip. Extensive experimental demonstrate the superiority of CAF-JT over conventional joint training approaches, showcasing its improved noise robustness. Particularly in low SNR conditions, CAF-JT achieves the best average recognition rates of 94.84% and 93.61% on the two datasets, respectively.
Aolong Zhou, Xiaoyong Li 0002, Wen Zhang 0016, Kefeng Deng, Kaijun Ren, Junqiang Song
IEEE Trans. Geosci. Remote. Sens.7
2023 A Novel Noise-Aware Deep Learning Model for Underwater Acoustic Denoising
abstract
Underwater acoustic signal denoising technology aims to overcome the challenge of recovering valuable ship target signals from noisy audios by suppressing underwater background noise. Traditional statistical-based denoising techniques are difficult to be applied effectively in complex underwater environments, especially in the case of extremely low signal-to-noise ratios (SNRs). To address these problems, we propose a noise-aware deep learning model with fullband-subband attention network (NAFSA-Net) for underwater acoustic signal denoising. NAFSA-Net adopts an encoder to extract the feature representation of the input audio. Subsequently, the noise subnet and the target subnet are designed to estimate the noise component and the target component simultaneously. Specifically, some stacked fullband-subband attention (FSA) blocks are deployed in each subnet to capture both global dependencies and fine-grained local dependencies of features. Furthermore, we introduce an interaction module to transmit auxiliary information from the noise subnet to the target subnet. Finally, we propose an improved weight SI-SNR loss function to optimize the training of our model. Experimental results show that our proposed NAFSA-Net substantially outperforms traditional methods and competitive DNN-based solutions in denoising underwater noisy signals with very low SNRs. More importantly, our proposals achieve equally excellent performance on both unseen datasets, which indicates that NAFSA-Net can be a more robust choice for real-world underwater acoustic denoising systems.
Aolong Zhou, Wen Zhang 0016, Xiaoyong Li 0002, Guojun Xu, Bingbing Zhang 0002, Yanxin Ma, Junqiang Song
IEEE Trans. Geosci. Remote. Sens.7
2022 FVec2vec: A Fast Nonlinear Dimensionality Reduction Approach for General Data
abstract
Dimensionality reduction is a fundamental technique to address the curse of dimensionality problem in real-world big datasets. However, most existing methods either only target raw datasets that contain explicit relationships between data points, or construct the complete neighborhood graph of the dataset by calculating pairwise similarities, and then generate contexts of data points by random walking to measure the structure of the dataset, which are computationally expensive. In this paper, we propose a fast nonlinear locality-preserving dimensionality reduction approach called FVec2vec, which extends the Skip-gram model to embedding representation of general numerical matrices. Specifically, instead of constructing neighborhood graph by calculating pairwise similarities between data points, we approximate the k-nearest neighbors (kNN) of each data point in matrices by exploring its neighbors’ neighbors first. Then, we design a novel sampling algorithm to randomly sample on the kNN to depict the structure of the dataset. Experimental results show that FVec2vec is faster than most existing methods while achieving acceptable accuracy, and the accuracy is even higher than the state-of-the-art method under certain similarity metrics.
Xiaoli Ren, Kefeng Deng, Kaijun Ren, Junqiang Song, Xiaoyong Li 0002
IEEE Big Data4
2022 A System For Hybrid 4DVar-EnKF Data Assimilation Based On Deep Learning
abstract
The accuracy of the initial field is crucial to the forecast results of numerical weather prediction (NWP). Data assimilation (DA) is a method to provide the initial field to the NWP. Currently, the hybrid 4DVar-EnKF DA method is the primary DA method used by operational NWP centres. The technique requires the derivation of the tangent linear and adjoint models for the nonlinear model, but it’s challenging to get the tangent linear and the adjoint models. Furthermore, this method usually adopts empirical coefficients to combine the four-dimensional variational assimilation (4DVar) and the ensemble Kalman filter (EnKF), which reduces the accuracy of assimilation results. This paper builds a hybrid DA system based on a deep learning model (DL-HDA) in response to the above problems. First, we establish a forecast model based on the bilinear neural network (BNN) and use the tangent linear and adjoint models of the BNN for the 4DVar. Then, we utilize the ResNet model to combine the analysis of the 4DVar and the EnKF. The experiments are carried out on the Lorenz-96 model, and then the DL-HDA is compared with the traditional method. The experimental results show that the DL-HDA can improve the precision of assimilation results and decrease the system’s running time.
Renze Dong, Hongze Leng, Junqiang Song, Chengwu Zhao, Jincai Li, Yunjie Lan
SMC3
2021 Improving Ocean Data Services with Semantics and Quick Index
Xiaoli Ren, Kaijun Ren, Zichen Xu 0001, Xiaoyong Li 0002, Aolong Zhou, Junqiang Song, Kefeng Deng
J. Comput. Sci. Technol.6
2021 Solving Boolean polynomial systems by parallelizing characteristic set method for cyber-physical systems
abstract
Summary Many cyber‐attach schemes and coding models established by algebra tools are build to address the problem of security of cyber‐pysical systems (CPS). As an important field of algebra computing, Boolean Polynomial System Solving (PoSSo) problem plays a very important role in many algebra applications. In this article, we propose an efficient Parallel Boolean Characteristic Set method (PBCS) under the high‐performance computing environment to improve the efficiency of solving Boolean polynomial systems. The PBCS is implemented based on the state‐of‐the‐art Boolean Characteristic Set method (BCS). It adopts a master‐slave parallel pattern, and distributes tasks based on the polynomial sets after initial zero decomposition. We design a strategy of dynamically reallocating tasks to ameliorate load imbalance, which is caused by dynamical zero decomposition of polynomials. Furthermore, we improve its performance by optimizing the parameter settings of PBCS, including the maximum number of polynomial branches that trigger the dynamic allocation policy and the scheduling time. Experimental results with solving several Boolean polynomial systems confirm that PBCS is efficient and scalable, especially for the equations generating from stream ciphers that have block triangular structure. Moreover, the method also has good scalability. It shows a stable speedup as well even extending to the size of thousands of CPU cores.
Juan Zhao 0006, Xiaoyong Li 0002, Zhenyu Huang 0004, Jincai Li, Junqiang Song
Softw. Pract. Exp.6
2020 A Hybrid 3DVar-EnKF Data Assimilation Approach Based on Multilayer Perceptron
abstract
The quality and accuracy of Numerical Weather Prediction (NWP) is based on its initial conditions (ICs), boundary conditions and forecast models. Data assimilation (DA) is a crucial procedure to optimally estimate the actual atmospheric state (known as the analysis field) as ICs for NWP by integrating available information, including the observation and the background field. Instead of only focusing on the speed-up for DA in virtue of the customized neural networks, this paper exploratively introduces the spatial-temporal peculiarities to construct a new hybrid data assimilation approach based on multilayer perceptron (MLP); and, its effectiveness and validity are verified in two classical nonlinear dynamic models. The results of experiments demonstrate that the Cache-MLP generally produces similar or smaller root mean square errors (RMSE) with much less time consuming, compared to the conventional 3D-Var and EnKF DA methods, and noticeably, the Cache-MLP has a more robust representation of turning points in the trajectories of the state variables. The final Backtracked-MLP learns appropriate weight matrix to couple previous two traditional DA methods and increases the accuracy by 10.32% in the Lorenz-63 system while 14.03% in the Lorenz-96 system, in comparison with the empirical hybrid DA method. To some extent, this method could be a reference to further researches to optimize the quality of the analysis field, in the meantime, saving significant computing time and resources by deep learning.
Lilan Huang, Hongze Leng, Junqiang Song, Juan Zhao 0006
IJCNN3
2020 pcIRM: Complex Ideal Ratio Masking for Speaker-Independent Monaural Source Separation with Utterance Permutation Invariant Training
abstract
Typical speech separation systems usually operate in the time-frequency (T-F) domain by enhancing the magnitude response and leaving the phase response unaltered. Recent studies, however, suggest that phase is important for perceptual quality, leading some researchers to consider magnitude and phase spectrum enhancements. The merging of the complex ideal ratio masking (cIRM) estimation and training with deep neural network (DNN) has been proved to be an effective way to improve speech separation. Furthermore, the label ambiguity (or permutation) problem has become a major barrier for speaker-independent multi-talker source separation, which prompts us to come up with new solutions. In this paper, to solve the problem of speaker-independent monaural source separation, we propose a novel method called pcIRM, which creatively achieves the cIRM estimation with the utterance-level permutation invariant training (uPIT). Specifically, pcIRM is implemented with the deep bidirectional LSTM (Bi-LSTM) RNN network, and evaluated with the WSJ0-2mix datasets. We report separation results for the proposed method and compare them to that of the existing state-of-the-art methods. Extensive experimental results demonstrate the advantages of our proposed pcIRM method in terms of the signal-to-distortion ratio (SDR) metric.
Wen Zhang 0016, Xiaoyong Li 0002, Aolong Zhou, Kaijun Ren, Junqiang Song
IJCNN5
2020 A Data and Task Co-Scheduling Algorithm for Scientific Cloud Workflows
abstract
Cloud computing has emerged as a promising computational infrastructure for cost-efficient workflow execution by provisioning on-demand resources in a pay-as-you-go manner. While scientific workflows require accessing community-wide resources, they usually need to be performed in collaborative cloud environments composed of multiple datacenters. Although such environments facilitate scientific collaboration, the movements of input and intermediate datasets across geographically distributed datacenters may cause intolerable latency that would hinder efficient execution of large-scale data-intensive scientific workflows. To address the problem, in this article we propose a novel multi-level K-cut graph partitioning algorithm to minimize the volume of data transfer across datacenters while satisfying load balancing and fixed data constraints. The algorithm first contracts the fixed input datasets in the same datacenter and their consuming tasks, and coarsens the contracted graph to a predefined scale in a level-by-level manner. Then, a K-cut algorithm is used to partition the resulted graph into K parts such that the cut size is minimized. After that, the partitioned graph is projected back to the original workflow graph, during which the load balancing constraint is maintained. We evaluate our algorithm using three real-world workflow applications and the results demonstrate that the proposed algorithm outperforms other state-of-the-art algorithms.
Kefeng Deng, Kaijun Ren, Junqiang Song
IEEE Trans. Cloud Comput.4
2019 Solving the Defect in Application of Compact Abating Probability to Convolutional Neural Network Based Open Set Recognition
abstract
Close set is a hypothesis utilized by the majority of machine-learning-based (ML-based) recognition algorithms, assuming all testing classes are known at training time. In real world, the more practical model is Open Set Recognition (OSR), which allows the presence of unknown classes at testing time, but requires the rejection ability of the model. The compact abating probability (CAP) model, which assumes the probability of class membership decreases in value (abates) as points move from known data toward open space, is first raised in traditional ML-based OSR method and soon become the basis of majority of later developed works. Most of convolutional-neural-network-based (CNN-based) OSR methods also adopted this model as their basis. During our exploration, however, we find that the application of CAP model to the CNN-based OSR method is restricted by the difference of its feature space from that of ML-based method. To the best of our knowledge, we are the first group who find this gap. To fill this gap, we propose a method called OpenSoftMax to transform the CNN-based methods' features by the process of SoftMax. In order to investigate performance, we further implement quantitative comparison between our OpenSoftMax method and the well-known CNN-based method OpenMax on caltech256 datasets. Extensive experiments have been conducted to verify the effectiveness and efficiency of our proposals.
Xiangyuan Sun, Xiaoyong Li 0002, Kaijun Ren, Junqiang Song
ICTAI4
2019 Parallelizing uncertain skyline computation against n-of-N data streaming model
abstract
Summary The skyline query over uncertain data streams, as an important aspect of big data analysis, plays a significant role in domains such as environment monitoring, decision‐making, and data mining. The skyline query over uncertain data streams with sliding window model always focuses on the most recent N streaming items, which cannot meet the query requirements of different window scales at the same time. To improve the query flexibility and efficiency, we propose an efficient parallel method for processing uncertain n‐of‐N skyline queries; that is, computing the skyline for the most recent n (∀n ≤ N) items in parallel. Specifically, we first propose a framework for parallelizing the query computation for uncertain n‐of‐N skylines. Furthermore, we put forward a sliding window partitioning strategy as well as a streaming items mapping strategy to realize the load balance for each node. In addition, we define a spatial index structure RST based on R‐tree to organize the elements within each individual sliding window and candidate set in each which can significantly improve the dominance tests. Most importantly, we provide an encoding interval scheme to transform the n‐of‐N query into stabbing query in each compute node, which can greatly minimize the query scope and improve the query efficiency. In addition, we use a red‐black tree named RBI to store all stabbing intervals. Extensive experimental results demonstrate that the proposals are efficient and can greatly meet the query requirement of users in real applications.
Jun Liu 0048, Xiaoyong Li 0002, Kaijun Ren, Junqiang Song
Concurr. Comput. Pract. Exp.4
2019 PAGCM: A scalable parallel spectral-based atmospheric general circulation model
abstract
Summary The Atmospheric General Circulation Model (AGCM) as one of the most important components of Climate System Model (CSM), has been proved to be an effective way for weather forecasting and climate prediction. Although lots of efforts have been conducted to improve the computing efficiency of AGCMs, such as exploit parallel algorithms, migrating codes, and even redesigning systems to adapt to the emerging computer architectures, it is not enough to match the real requirement, due to the limited scalability of the parallel algorithms themselves. Therefore, we design and implement a scalable parallel spectral‐based atmospheric circulation mode called PAGCM in this paper. Specifically, we first analyze the data dependencies of the dimensions in different spaces according to the calculation characteristics of spectral models, and based on which we propose a two‐dimensional decomposition algorithm in PAGCM to effectively increase the involving cores for the parallel computing, and thus reduce the overall computing time. Furthermore, to adapt to the novel data decomposition in each computing stage of dynamic framework, we propose three‐dimensional data transposition algorithms and data collection algorithms correspondingly, by considering of load balancing and communication optimization. Extensive experiments are conducted on Tianhe‐2 to validate the effectiveness and scalability of our proposals.
Xiaoli Ren, Juan Zhao 0006, Xiaoyong Li 0002, Kaijun Ren, Junqiang Song, Difu Sun
Concurr. Comput. Pract. Exp.5
2019 Rethinking compact abating probability modeling for open set recognition problem in Cyber-physical systems
Xiangyuan Sun, Xiaoyong Li 0002, Kaijun Ren, Junqiang Song, Zichen Xu 0001
J. Syst. Archit.4
2018 Parallel n-of-N Skyline Queries over Uncertain Data Streams
Jun Liu 0048, Xiaoyong Li 0002, Kaijun Ren, Junqiang Song, Zongshuo Zhang
DEXA (2)4
2018 PBCS: An Efficient Parallel Characteristic Set Method for Solving Boolean Polynomial Systems
abstract
Solving Boolean polynomial systems as an important aspect of symbolic computation, plays a fundamental role in various real applications. Although there exist many efficient sequential algorithms for solving Boolean polynomial systems, they are inefficient or even unavailable when the problem scale becomes large, due to the computational complexity of the problem and the limited processing capability of a single node. In this paper we propose an efficient parallel characteristic set method called PBCS for solving Boolean polynomial systems under the high-performance computing environment. Specifically, PBCS takes full advantage of the state-of-the-art characteristic set method and achieves load balancing by dynamically reallocating tasks. Moreover, the performance is further improved by optimizing the parameter setting. Extensive experiments are conducted to demonstrate that PBCS is efficient and scalable for solving Boolean equations, especially for the equations rasing from stream ciphers that have block triangular structure. In addition, the algorithm has good scalability and can be extended to the size of thousands CPU cores with a stable speedup.
Juan Zhao 0006, Junqiang Song, Jincai Li, Zhenyu Huang 0004, Xiaoyong Li 0002, Xiaoli Ren
ICPP2
2018 Comparison of Wind Speed from Quikscat, Ascat, Windsat, Era-Interim Reanalysis and Ship Measurements Over the China Sea
abstract
In this paper, we performed a comparison of wind speeds from the Quick Scatterometer (QuikSCAT), the MetOp-A Advanced Scatterometer (ASCAT), the WindSat Polarimetric Radiometer (WindSat) and ERA-Interim reanalysis using in situ ship measurements. The comparison was made over the China Sea during a 12-month period from January to December in 2008. The mean bias and Root Mean Square Error (RMSE) were calculated for the matchup dataset. The ASCAT wind speed product was observed more accurate and suitable for the China Sea during the research period with a relatively lower mean bias and RMSE. We also analyzed the accuracy of surface winds in different wind speed ranges. The statistical results show that the wind speeds of all products agree well in the ranges from 5m/s to 10m/s. However, underestimation at high wind speeds and overestimation at low wind speeds have been observed. Furthermore, the rain effects on the scattermeter wind measurements were considered, and ASCAT shows slightly better results and less affected by rain because of its C-band configuration compared with Ku-band QSCAT.
Dongxiang Zhang, Kaijun Ren, Jia Liu 0021, Junqiang Song
IGARSS5
2017 Parallel global atmospheric correction for FY3/MERSI data over land on multi-core and many-core architectures
abstract
For the accurate derivation of biophysical parameters based on surface reflectance, atmospheric correction is a necessary step to remove scattering and absorption effects by multiple atmospheric components. However, the huge amount of data and complex algorithms pose great computing challenges for massive operational tasks. Towards the global atmospheric correction for Medium Resolution Spectral Imager (MERSI) data onboard FY-3A and FY-3B, this paper describes an atmospheric correction algorithm considering the directional properties of the observed surface, and exploits its parallel implementations on multi-core and many-core architectures. The algorithm was developed with Open Multiprocessing (OpenMP) for multi-core processors and Compute Unified Device Architecture (CUDA) for Graphics Processing Units (GPU). Experimental results show the runtime was reduced from 187.19s to 42.63s and 10.11s when implemented on a multi-core processor and NVIDIA Tesla K80 respectively.
Jia Liu 0021, Jie Guang, Kaijun Ren, Junqiang Song, Yong Xue, Cheng Fan 0001, Shuchang Wang
IGARSS5
2015 DAG Scheduling for Heterogeneous Systems Using Biogeography-Based Optimization
abstract
Efficient scheduling algorithm is critical for DAG-based applications to obtain high-performance in heterogeneous computing systems. In comparison with heuristic-based algorithms, meta-heuristic based scheduling algorithms can produce better results by searching in a guided manner. Biogeography-based optimization (BBO) is a recently proposed optimization technique which has shown less parameters, faster convergency, and superior performance than existing meta-heuristics. In this article, we introduce this novel optimization technique into the field of DAG scheduling. To reduce scheduling overhead, the proposed algorithm only encodes task mapping while using a heuristic strategy to determine task ordering. Moreover, it uses heuristic-based algorithms as baseline algorithms to obtain better results. We evaluate the BBO-based scheduling algorithm using three real world DAG-based applications under various parameter settings. The results show that the BBO-based scheduling algorithm outperforms the state-of-the-art meta-heuristic based algorithms.
Kefeng Deng, Kaijun Ren, Junqiang Song
ICPADS4
2015 HLognGP: A parallel computation model for GPU clusters
abstract
Summary Parallel computation model is an abstraction for the performance characteristics of parallel computers, and should evolve with the development of computational infrastructure. The heterogeneous CPU/Graphics Processing Unit (GPU) systems have been and will be important platforms for scientific computing, which introduces an urgent demand for new parallel computation models targeting this kind of supercomputers. In this research, we propose a parallel computation model called HLognGP to abstract the computation and communication features of heterogeneous platforms like TH‐1A. All the substantial parameters of HLognGP are in vector form and deal with the new features in GPU clusters. A simplified version HLog3GP of the proposed model is mapped to a specific GPU cluster and verified with two typical benchmarks. Experimental results show that HLog3GP outperforms the other two evaluated models and can well model the new particularities of GPU clusters. Copyright © 2015 John Wiley & Sons, Ltd.
Fengshun Lu, Junqiang Song, Yufei Pang
Concurr. Comput. Pract. Exp.2
2013 Exploring portfolio scheduling for long-term execution of scientific workloads in IaaS clouds
abstract
Long-term execution of scientific applications often leads to dynamic workloads and varying application requirements. When the execution uses resources provisioned from IaaS clouds, and thus consumption-related payment, efficient and online scheduling algorithms must be found. Portfolio scheduling, which selects dynamically a suitable policy from a broad portfolio, may provide a solution to this problem. However, selecting online the right policy from possibly tens of alternatives remains challenging. In this work, we introduce an abstract model to explore this selection problem. Based on the model, we present a comprehensive portfolio scheduler that includes tens of provisioning and allocation policies. We propose an algorithm that can enlarge the chance of selecting the best policy in limited time, possibly online. Through trace-based simulation, we evaluate various aspects of our portfolio scheduler, and find performance improvements from 7% to 100% in comparison with the best constituent policies and high improvement for bursty workloads.
Kefeng Deng, Junqiang Song, Kaijun Ren, Alexandru Iosup
SC2
2013 A clustering based coscheduling strategy for efficient scientific workflow execution in cloud computing
abstract
SUMMARY Due to its advantages of cost‐effectiveness, on‐demand provisioning and easy for sharing, cloud computing has grown in popularity with the research community for deploying scientific applications such as workflows. Although such interests continue growing and scientific workflows are widely deployed in collaborative cloud environments that consist of a number of data centers, there is an urgent need for exploiting strategies which can place application datasets across globally distributed data centers and schedule tasks according to the data layout to reduce both latency and makespan for workflow execution. In this paper, by utilizing dependencies among datasets and tasks, we propose an efficient data and task coscheduling strategy that can place input datasets in a load balance way and meanwhile, group the mostly related datasets and tasks together. Moreover, data staging is used to overlap task execution with data transmission in order to shorten the start time of tasks. We build a simulation environment on Tianhe supercomputer for evaluating the proposed strategy and run simulations by random and realistic workflows. The results demonstrate that the proposed strategy can effectively improve scheduling performance while reducing the total volume of data transfer across data centers. Concurrency and Computation: Practice and Experience, 2013.© 2013 Wiley Periodicals, Inc.
Kefeng Deng, Kaijun Ren, Junqiang Song, Dong Yuan 0001, Yang Xiang 0001, Jinjun Chen
Concurr. Comput. Pract. Exp.3
2013 Notes and correspondence on ensemble-based three-dimensional variational filters
abstract
Several ensemble-based three-dimensional variational (3D-Var) filters are compared. These schemes replace the static background error covariance of the traditional 3D-Var with the ensemble forecast error covariance, but generate analysis ensemble anomalies (perturbations) in different ways. However, it is demonstrated in this paper that they are all theoretically equivalent to the ensemble transformation Kalman filter (ETKF). Furthermore, a new method named EnPSAS is presented. The analysis shows that EnPSAS has a small condition number and can apply covariance localization more easily than other ensemble-based 3D-Var methods.
Hongze Leng, Junqiang Song, Fukang Yin
J. Zhejiang Univ. Sci. C2
2013 A bargaining-driven global QoS adjustment approach for optimizing service composition execution path
Kaijun Ren, Junqiang Song, Nong Xiao 0001
J. Supercomput.2
2011 A Weighted K-Means Clustering Based Co-scheduling Strategy towards Efficient Execution of Scientific Workflows in Collaborative Cloud Environments
abstract
Due to the advantages of cost-effectiveness, on-demand resource provision and easy for sharing, cloud computing has grown in popularity with research community for deploying scientific applications such as workflows. When such interest continues growing and workflows are widely performed in collaborative cloud environments that consist of a number of data centers, there is an urgent need for exploiting strategies which can place the application data across globally distributed data centers and schedule tasks according to the data layout to reduce both the latency and make span for workflow execution. In this paper, by utilising dependencies among datasets and tasks, we propose an efficient data and task co scheduling strategy that can place input datasets in a load balance way and meanwhile group the mostly related datasets and tasks together. We build a simulation environment on Tianhe supercomputer to evaluate the proposed strategy and run simulations by random and realistic workflows. The results demonstrate that the proposed strategy can effectively improve workflows performance while reducing the total volume of data transfer across data centers.
Kefeng Deng, Lingmei Kong, Junqiang Song, Kaijun Ren, Dong Yuan 0001
DASC3
2011 The TianHe-1A Supercomputer: Its Hardware and Software
Xuejun Yang, Xiangke Liao, Kai Lu 0001, Qingfeng Hu, Junqiang Song, Jinshu Su
J. Comput. Sci. Technol.5
2010 MPIActor - A Multicore-Architecture Adaptive and Thread-Based MPI Program Accelerator
abstract
Improving MPI foundational software to suit multicore systems is a key issue for developing effective parallel software on high performance communication domain. Towards this issue, in this paper, we propose a novel technique, called MPI Accelerator or MPIActor in short, which is a transparent middleware to enhance conventional MPI libraries. The main idea is to optimize MPI routines for multicore systems by adopting threaded MPI mechanism and multicore architecture aware collectives in MPIActor. With the join of MPIActor, on one hand, all MPI processes in each node are mapped to several threads in one process. As a result, the overhead of intra-node point-to-point communications can greatly decrease. On the other hand, the collective routines are implemented by the cooperation of individual intra - and inter-node collective subroutines, and the intra-node collective subroutines can be further optimized by multicore architecture aware collective algorithms. Based on above idea, a framework involving an MPI_Reduce routine and a set of point-to-point communication routines has been implemented and evaluated on a 256 cores Nehalem platform. When compared to the performance of MVAPICH2, the final experimental results show that the performance by MPIActor can be significantly improved whatever by using OSU_LATENCY benchmark for point-to-point communications or IMB Reduce benchmark for reduction collectives. Especially, the performance results of using OSU_LATENCY benchmark even can be improved up to 321%.
Kaijun Ren, Junqiang Song
HPCC3
2010 MPIActor: A thread-based MPI program accelerator
abstract
Towards gaining the performance improvement benefited from threaded MPI while supporting MPI standard well, in this paper, we propose a thread-based MPI program accelerator (MPIActor). MPIActor is a transparent middleware to assist general MPI libraries. People can choose to adopt or abandon MPIActor freely in compiling time for any MPI program (Currently only support C code). With the join of MPIActor, in each node, the MPI processes will be mapped as several threads of one process, and the intra-node point-to-point communication and collective communication will have been enhanced by take advantage of thread based mechanism. We have implemented the point-to-point communication module of our design and evaluated it on a real platform. Comparing with MVAPICH2, the experimental results of OSU PINGPONG benchmark show a significant performance improvement from 114% to 321% for transferring messages which size is between 4KB and 2MB.
Junqiang Song, Shaoliang Peng
IWQoS2
2010 TH-1: China's first petaflop supercomputer
Xuejun Yang, Xiangke Liao, Weixia Xu 0001, Junqiang Song, Qingfeng Hu, Jinshu Su, Liquan Xiao, Kai Lu 0001, Qiang Dou, Juping Jiang, Canqun Yang
Frontiers Comput. Sci. China4
2009 Routing on Shortest Pair of Disjoint Paths with Bandwidth Guaranteed
abstract
QoS routing and multipath routing have been receiving much attention respectively in network communication. However, the research combining those two kinds of routing is rare. This paper integrated the ideas of QoS and multipath, and presented the problem of Shortest Pair of Disjoint Paths with Bandwidth Guaranteed. We proved it to be NP-Complete, and then proposed a heuristic algorithm. The analysis indicates that our algorithm shows good performance and it can produce optimal solutions in most cases.
Hongze Leng, Meilian Liang, Junqiang Song
DASC3
2009 Gradual Removal of QoS Constraint Violations by Employing Recursive Bargaining Strategy for Optimizing Service Composition Execution Path
abstract
A critical issue in service composition area is how to achieve an optimized overall end-to-end quality of service(QoS) requirements by effectively coordinating QoS constraints for individual service. However, this issue has not yet been well addressed. In this paper, we propose a novel method by employing a recursive bargaining Strategy to gradually remove QoS constraint violations for Optimizing service composition execution Path. Our method mainly exploits the hidden market competitive relationships which widely exist in real business world for developing a novel bargaining strategy. Based on this strategy, concessions can be made by service providers to offer better QoS values. By recursively using bargaining strategy, an initial execution path built by a local optimization policy for service composition, can be continually updated to be close to the optimal one by reselecting better service providers for meeting overall end-to-end QoS requirements. An experiment and evaluation have been made to demonstrate the feasibility and effectiveness of our proposed method.
Kaijun Ren, Nong Xiao 0001, Junqiang Song, Chi Yang, Jinjun Chen
ICWS3
2009 Building Quick Service Query list (QSQL) to support automated service discovery for scientific workflow
abstract
Abstract Scientific workflow is emerging as a promising scientific computing paradigm to offer the convenience for the scientists to resolve complex scientific problems. To successfully execute a scientific workflow, the workflow creation by depending on service discovery techniques should be made in the first place. Particularly, semantics have been proposed as a key to automatically solve service discovery issue for facilitating users to create a workflow. However, most of the semantic service discovery methods still remain at a low‐efficiency stage because they generally involve a large number of ontology reasoning that is often time consuming. To address this issue, we present an efficient service discovery method by building Quick Service Query list (QSQL) to support automated service discovery for creating a workflow. QSQL based on graph storage theory is an efficient service index list that is dynamically built by service publication algorithm. In QSQL, semantic relationships between the published services and all related ontology concepts can be processed in advance so that a large number of ontology reasoning can be avoided during service discovery. Further, our proposed discovery algorithm can efficiently select service models from QSQL to match a user query. The final experiments further demonstrate the feasibility and the efficiency of our proposed method. Copyright © 2009 John Wiley & Sons, Ltd.
Kaijun Ren, Jinjun Chen, Nong Xiao 0001, Junqiang Song
Concurr. Comput. Pract. Exp.4
2008 Building Quick Service Query List Using Wordnet for Automated Service Composition
abstract
Current existing semantic composition methods mainly rely on ontology reasoning to support automated service composition. However, in reality, ontologies are generally unavailable or ontology reasoning is time-consuming; thus existing semantic composition methods are becoming impractical in the general service integration field. To address this problem, in this paper, we present an innovative composition technique by combining Wordnet with ontologies together to build an extended quick service query list (EQSQL) for supporting automated service composition. In EQSQL, data structures are designed particularly to record service information and their associated semantic concepts by previously processing semantic-related computing during service publication stage. Based on EQSQL, not only a quick response can be achieved, but the semantically-similar composition quality as well even if there is a lack of concrete domain-dependent ontologies for a user query.
Kaijun Ren, Jinjun Chen, Nong Xiao 0001, Junqiang Song
APSCC4
2007 A Pre-reasoning Based Method for Service Discovery and Service Instance Selection in Service Grid Environments
abstract
Current service composition and coordination still remain at large amount of manual processing stage, which has brought about low efficiency. In this paper, we present an efficient algorithm for abstract service discovery and a service instance selection method. Our algorithm firstly builds up the special data structures of ontology concepts based on graph storage theories when publishing abstract services. Then, these data structures form a quick service query list. In our algorithm, the large number of ontology reasoning is processed at service publication stage, thus we can make sure the quick query response in service discovery without much reasoning. In addition, our service instance selection methods based on OWL QoS ontology can enable grid resource sharing and coordination more flexible.
Kaijun Ren, Junqiang Song, Jinjun Chen, Nong Xiao 0001, Cancan Liu
APSCC2
2005 High-performance navigation and rendering of very-large scale landscape and seascape
abstract
With fast development of graphics hardware, the current algorithms on rendering of landscape and seascape, which are focused on precise view-frustum culling and complicated decision of level of detail (LOD), lead the CPU to be the bottleneck of system. This paper presents a new framework based on some new data structures and new rendering algorithms for navigation and rendering of very large-scale landscape and seascape. Firstly, we construct a new terrain data model, which consists of terrain tile pyramid and terrain summary pyramid. Secondly, in order to stitch the terrain tile boundaries seamlessly, the index template is introduced, and an algorithm of LOD based on "remarkability" is described in detail. Finally, Perlin noise is used to simulate the ocean surface and a framework of navigation of very large-scale scene is completed. Experimental results show that this system can satisfy the requirements of real-time navigation of large-scale landscape and seascape.
Xuexian Pi, Sikun Li, Junqiang Song
CAD/Graphics4