Xiaoying Wang 0002

dblp:47/807-2 · DBLP profile ↗
← Back
36ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-1029-0358ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OSOW: A Resource-Efficient Parallelization of the Winograd Algorithm on Multicore CPUs
Haodong Bian, Xiaoying Wang 0002
KSEM (6)4
2026 PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUs
abstract
SpMV has been widely utilized and is regarded as a significant kernel in various scientific and engineering computing applications, where its parallel performance is heavily influenced by matrix sparsity and hardware architecture. Despite extensive prior research, static partitioning strategies that narrowly target computation or memory access remain a key performance bottleneck, severely stifling performance advancement of SpMV on modern multicore CPUs.
Haodong Bian, Youhui Zhang, Jianqiang Huang 0002, Xiaoying Wang 0002
PPoPP5
2026 Enhancing multivariate weather forecasting via temporal attention and spatiotemporal fusion
Zhibo Kong, Xiaoying Wang 0002, Guojing Zhang
Eng. Appl. Artif. Intell.3
2026 Dynamic pivotal node attention for spatio-temporal wind speed forecasting
Guojing Zhang, Jianqiang Huang 0002, Xiaoying Wang 0002
Neurocomputing5
2025 Research on Parallel Weighted Back-Projection Algorithm on Multi-CPU
Kaige Zheng, Kuangzheng Wu, Haodong Bian, Jianqiang Huang 0002, Xiaoying Wang 0002
ICIC (14)6
2025 Imputing Missing Temperature Data of Meteorological Stations Based on Global Spatiotemporal Attention Neural Network
abstract
Imputing missing meteorological site temperature data is necessary and valuable for researchers to analyze climate change and predict related natural disasters. Prior research often used interpolation-based methods, which basically ignored the temporal correlation existing in the site itself. Recently, researchers have attempted to leverage deep learning techniques. However, these models cannot fully utilize the spatiotemporal correlation in meteorological stations data. Therefore, this paper proposes a global spatiotemporal attention neural network (GSTA-Net), which consists of two sub networks, including the global spatial attention network and the global temporal attention network, respectively. The global spatial attention network primarily addresses the global spatial correlations among meteorological stations. The global temporal attention network predominantly captures the global temporal correlations inherent in meteorological stations. To further fully exploit and utilize spatiotemporal information from meteorological station data, adaptive weighting is applied to the outputs of the two sub-networks, thereby enhancing the imputation performance. Additionally, a progressive gated loss function has been designed to guide and accelerate GSTA-Net’s convergence. Finally, GSTA-Net has been validated through a large number of experiments on public dataset TND and QND with missing rates of 25%, 50%, and 75%, respectively. The experimental results indicate that GSTA-Net outperforms the latest models, including Linear, NLinear, DLinear, PatchTST, and STA-Net, across both the mean absolute error (MAE) and the root mean square error (RMSE) metrics.
Tianrui Hou, Xinshuai Guo, Xiaoying Wang 0002, Guojing Zhang, Jianqiang Huang 0002
SMC4
2025 Research on GPU transplantation optimization of PRM scalar advection scheme in GRAPES global forecast system
Zhangjie Tan, Jinfang Jia, Zhengsheng Ning, Jianqiang Huang 0002, Xiaoying Wang 0002
CCF Trans. High Perform. Comput.5
2025 Vector knowledge transfer-driven representation learning for heterogeneous hypernetworks
Yijian Chen, Xiaoying Wang 0002, Jianqiang Huang 0002
Knowl. Inf. Syst.3
2025 CPAT: cross-patch aggregated transformer for time series forecasting
Xiaoying Wang 0002, Jianqiang Huang 0002, Guojing Zhang
Mach. Learn.3
2025 GDSR: Global-Detail Integration Through Dual-Branch Network With Wavelet Losses for Remote Sensing Image Super-Resolution
abstract
In recent years, deep neural networks, including Convolutional Neural Networks, Transformers, and State Space Models, have achieved significant progress in Remote Sensing Image (RSI) Super-Resolution (SR). However, existing SR methods typically overlook the complementary relationship between global and local dependencies. These methods either focus on capturing local information or prioritize global information, which results in models that are unable to effectively capture both global and local features simultaneously. Moreover, their computational cost becomes prohibitive when applied to large-scale RSIs. To address these challenges, we introduce the novel application of Receptance Weighted Key Value (RWKV) to RSI-SR, which captures long-range dependencies with linear complexity. To simultaneously model global and local features, we propose the Global-Detail dual-branch structure, GDSR, which performs SR by paralleling RWKV and convolutional operations to handle large-scale RSIs. Furthermore, we introduce the Global-Detail Reconstruction Module (GDRM) as an intermediary between the two branches to bridge their complementary roles. In addition, we propose the Dual-Group Multi-Scale Wavelet Loss, a wavelet-domain constraint mechanism via dual-group subband strategy and cross-resolution frequency alignment for enhanced reconstruction fidelity in RSI-SR. Extensive experiments under two degradation methods on several benchmarks, including AID, UCMerced, and RSSRD-QH, demonstrate that GSDR outperforms the state-of-the-art Transformer-based method HAT by an average of 0.09 dB in PSNR, while using only 63% of its parameters and 51% of its FLOPs, achieving an inference speed 3.2 times faster.
Qiwei Zhu, Kai Li 0047, Guojing Zhang, Xiaoying Wang 0002, Jianqiang Huang 0002, Xilai Li
IEEE Trans. Geosci. Remote. Sens.4
2024 ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
abstract
Spatiotemporal prediction aims to generate future sequences by paradigms learned from historical contexts. It is essential in numerous domains, such as traffic flow prediction and weather forecasting. Recently, research in this field has been predominantly driven by deep neural networks based on autoencoder architectures. However, existing methods commonly adopt autoencoder architectures with identical receptive field sizes. To address this issue, we propose an Asymmetric Receptive Field Autoencoder (ARFA) model, which introduces corresponding sizes of receptive field modules tailored to the distinct functionalities of the encoder and decoder. In the encoder, we present a large kernel module for global spatiotemporal feature extraction. In the decoder, we develop a small kernel module for local spatiotemporal information reconstruction. Experimental results demonstrate that ARFA consistently achieves state-of-the-art performance on popular datasets. Additionally, we construct the RainBench, a large-scale radar echo dataset for precipitation prediction, to address the scarcity of meteorological data in the domain.
Xuechao Zou, Xiaoying Wang 0002, Jianqiang Huang 0002, Junliang Xing
ICASSP4
2024 DisRot: boosting the generalization capability of few-shot learning via knowledge distillation and self-supervised learning
Chenyu Ma, Jinfang Jia, Jianqiang Huang 0002, Xiaoying Wang 0002
Mach. Vis. Appl.5
2024 Interaction Trust-Driven Data Distribution for Vehicle Social Networks: A Matching Theory Approach
abstract
Due to the rapid expansion of the Internet of Vehicles (IoVs), service providers deploy roadside units (RSUs), and base stations (BSs) close to vehicles. They can provide vehicles with computational offloading services quickly. In the context of vehicle social networks, where vehicles can communicate and share data with each other, the security and efficiency of data distribution are crucial. Unfortunately, the open nature of RSU BSs makes them vulnerable to malicious attackers, hence affecting the quality of the user experience. This article proposes a security trust degree incentive-based evaluation mechanism that calculates the security trust degree of vehicle users to RSU BSs through the continuous interaction between them in order to effectively address the aforementioned issues. Additionally, taking into account the competitive nature of task computation offloading between vehicle users and BSs, a stable matching algorithm is used to match each vehicle user with the most appropriate BS so that they can work together to prevent competition in task offloading and improve task offloading efficiency. Due to the limited number of BS matches and the dynamic position changes of vehicle users, we further increase the data distribution efficiency by calculating the vehicle user degree of relationship and connection probability to match vehicle users with similar preferences. Finally, our proposed scheme is validated via numerous simulations with enhanced security service performance in terms of vehicle task offloading, while data distribution efficiency are effectively improved.
Jie Yi, Xiaoying Wang 0002, Changqiao Xu
IEEE Trans. Comput. Soc. Syst.3
2024 Heterogeneous Hypernetwork Representation Learning With Hyperedge Fusion
abstract
Most of the existing hypernetwork representation learning methods fail to fully consider the hyperedges, leading to the untapped potential of information contained within the hyperedges. To address this issue, this article proposes a heterogeneous hypernetwork representation learning method with hyperedge fusion abbreviated as HRHF. First, this method incorporates the hyperedges into random walk node sequences by means of incidence graph to enhance tuple relationships, i.e., the hyperedges among the nodes. Second, under the condition of the above random walk node sequences, the cognitive structure model, cognitive set model, and cognitive hyperedge model are jointly optimized to comprehensively consider pairwise relationships and tuple relationships among the nodes to learn high-quality node representation vectors. The experimental results demonstrate that, for the link prediction task, this method outperforms other optimal baseline methods, i.e., hyper-path-based random walks + hyper-gram (HPHG) by 0.99% points on the drug dataset, is comparable to the other optimal method, i.e., Event2vec on the global positioning system (GPS) dataset, is close to the performance of other optimal methods, i.e., hyper-path-based random walks + skip-gram (HPSG) on the MovieLens, and is close to the performance of other optimal methods, i.e., Hyper2vec on the WordNet dataset. For the hypernetwork reconstruction task, this method achieves superior average performance on the drug and GPS datasets compared with other baseline methods.
Xiaoying Wang 0002, Jianqiang Huang 0002
IEEE Trans. Comput. Soc. Syst.3
2024 PTCC: A Privacy-Preserving and Trajectory Clustering-Based Approach for Cooperative Caching Optimization in Vehicular Networks
abstract
5G vehicular networks provide abundant multimedia services among mobile vehicles. However, due to the mobility of vehicles, large-scale mobile traffic poses a challenge to the core network load and transmission latency. It is difficult for existing solutions to guarantee the quality of service (QoS) of vehicular networks. Besides, the sensitivity of vehicle trajectories also brings privacy concerns in vehicular networks. To address these problems, we propose a privacy-preserving and trajectory clustering-based framework for cooperative caching optimization (PTCC) in vehicular networks, which includes two tasks. Specifically, in the first task, we first apply differential privacy technologies to add noise to vehicle trajectories. In addition, a data aggregation model is provided to make the trade-off between aggregation accuracy and privacy protection. In order to analyze similar behavioral vehicles, trajectory clustering is then achieved by utilizing machine learning algorithms. In the second task, we construct a cooperative caching objective function with the transmission latency. Afterwards, the multi-agent deep Q network (MADQN) is leveraged to obtain the goal of caching optimization, which can achieve low delay. Finally, extensive simulation results verify that our framework respectively improves the QoS up to$9.8\%$and$12.8\%$with different file numbers and caching capacities, compared with other state-of-the-art solutions.
Zizhen Zhang, Xiaoying Wang 0002, Changqiao Xu
IEEE Trans. Sustain. Comput.3
2023 STA-Net: Reconstruct Missing Temperature Data of Meteorological Stations Using a Spatiotemporal Attention Neural Network
Tianrui Hou, Xinzhong Zhang, Xiaoying Wang 0002, Jianqiang Huang 0002
ICONIP (7)4
2023 PAS: A new powerful and simple quantum computing simulator
abstract
Abstract In recent years, many researchers have been using CPU for quantum computing simulation. However, in reality, the simulation efficiency of the large‐scale simulator is low on a single node. Therefore, striving to improve the simulator efficiency on a single node has become a serious challenge that many researchers need to solve. After many experiments, we found that much computational redundancy and frequent memory access are important factors that hinder the efficient operation of the CPU. This paper proposes a new powerful and simple quantum computing simulator: PAS (power and simple). Compared with existing simulators, PAS introduces four novel optimization methods: efficient hybrid vectorization, fast bitwise operation, memory access filtering, and quantum tracking. In the experiment, we tested the QFT (quantum Fourier transform) and RQC (random quantum circuits) of 21 to 30 qubits and selected the state‐of‐the‐art simulator QuEST (quantum exact simulation toolkit) as the benchmark. After experiments, we have concluded that PAS compared with QuEST can achieve a mean speedup of (QFT), (RQC) (up to , ) on the Intel Xeon E5‐2670 v3 CPU.
Haodong Bian, Jianqiang Huang 0002, Jiahao Tang, Runting Dong, Xiaoying Wang 0002
Softw. Pract. Exp.6
2022 A Two-stage Algorithm Based on Prediction and Search for Maxk-Truss Decomposition
abstract
The cohesive subgraph K-truss is often used to detect communities in large-scale social networks. The truss decomposition algorithm calculates the number of triangles formed by each edge, then peels off the edges with the number of triangles less than k-2 iteratively. The incremental maxk-truss algorithm must detect k-truss one by one making it inefficient. We propose a two-stage maxk-ktruss decomposition algorithm PSKT. The innovation points of PSKT are as follows: 1) The k-core structure is used to predict the maxk value, and the structure helps us avoid unnecessary mass calculations. 2) The characteristic that the k-truss of graph G must be included in the (k-l)-core of graph G is utilized to ensure that the proposed algorithm can safely prune in the graph. 3) A suitable triangle counting algorithm and the dynamic stream compression method are used to improve the algorithm efficiency. Comprehensive experiments demonstrate that PSKT is 15.26 times faster than the incremental algorithm FMT-dec and 3 times faster than the state-of-the-art maxk-truss algorithm FMT-max.
Jiahao Tang, Lingbin Liu, Jinfang Jia, Xiaoying Wang 0002, Jianqiang Huang 0002
ICPADS5
2022 Simulation of three-dimensional phase field model with LBM method using OpenCL
Chenyu Ma, Jinfang Jia, Jianqiang Huang 0002, Xiaoying Wang 0002
J. Supercomput.6
2022 $TC-Stream$TC-Stream: Large-Scale Graph Triangle Counting on a Single Machine Using GPUs
abstract
In this paper, we build a TC-Stream, a high-performance graph processing system specific for a triangle counting algorithm on graph data with up to tens of billions of edges, which significantly exceeds the device memory capacity of Graphics Processing Units (GPUs). The triangle counting problem is a broad research topic in data mining and social network analysis in the graph processing field. As the scale of the graph data grows, a portion of the graph data must be loaded iteratively. To solve the above problem, we propose TC-Stream. It focuses on three issues: 1) For power-law graphs, because the amount of tasks of each vertex or edge is inconsistent, it is bound to cause different demands of computing and memory resources for different task types. We propose a parallel vertex approach and the reordering of vertices for graph data that can be placed in the GPU device memory to ensure the maximum workload balancing; 2) A binary-search-based set intersection method is designed to achieve the maximum parallelism in GPU; 3) For the graph data that exceeds the GPU device memory capacity, we develop a novel vertical partition algorithm to guarantee the independent computing on each partition so that the three computation processes, i.e., the computation on GPU, the data transmission between main memory of CPU and SSD, and the communication between the CPU and the GPU can be perfectly overlapped. Extensive experiments conducted on large-scale datasets showed that the TC-stream running on a single Tesla V100 GPU performs 2.4 6 and 1.8 4.4 faster than the state-of-the-art single-machine in-memory triangle counting system and GPU-based triangle counting system, respectively, and achieves 2.4faster than the state-of-the-art out-of-core distributed system PDTL running on an 8-node cluster when processing the graph data with 42.5 billion edges.
Jianqiang Huang 0002, Haojie Wang 0004, Xiaoying Wang 0002
IEEE Trans. Parallel Distributed Syst.4
2021 ALBUS: A method for efficiently processing SpMV using SIMD and Load balancing
Haodong Bian, Jianqiang Huang 0002, Lingbin Liu, Dongqiang Huang, Xiaoying Wang 0002
Future Gener. Comput. Syst.5
2020 CSR2: A New Format for SIMD-accelerated SpMV
abstract
SpMV (Sparse matrix-vector multiplication) has attracted the attention of researchers in related fields at home and abroad. Of course, improving SpMV performance has also been a research hot spot for researchers in related fields. In this paper, we propose a new sparse matrix storage format CSR2 (Compressed Sparse Row 2) suitable for SIMD (Single Instruction Multiple Data)-accelerated SpMV. First, the format operation of CSR2 is easy to implement and has a low overhead of conversion. Second, CSR2 is a new single format and suitable for use on processor platforms with SIMD vectorization. We compare the SpMV algorithm based on CSR21with the one based on the current most advanced single format CSR5 (Compressed Sparse Row 5) on two mainstream high-performance processors: Intel Core i7-7700HQ CPU and Intel Xeon CPU E5-2670 v3. We choose 10 sets of regular matrices and 3 sets of irregular matrices to be used as benchmark suit. Experiments show that for the 13 sets of regular and irregular matrices in the benchmark suite, CSR2 has an average performance improvement of more than 50% compared to CSR5 (up to 125% on Intel Core i77700HQ CPU and 303% on Intel Xeon CPU E5-2670 v3). For applications with multiple iterations, in reality, using our CSR2 can bring low-overhead format conversion and high-throughput computing performance.
Haodong Bian, Jianqiang Huang 0002, Runting Dong, Lingbin Liu, Xiaoying Wang 0002
CCGRID5
2020 HpQC: A New Efficient Quantum Computing Simulator
Haodong Bian, Jianqiang Huang 0002, Runting Dong, Yuluo Guo, Xiaoying Wang 0002
ICA3PP (2)5
2020 Survey of single image super-resolution reconstruction
abstract
Image super‐resolution reconstruction refers to a technique of recovering a high‐resolution (HR) image (or multiple images) from a low‐resolution (LR) degraded image (or multiple images). Due to the breakthrough progress in deep learning in other computer vision tasks, people try to introduce deep neural network and solve the problem of image super‐resolution reconstruction by constructing a deep‐level network for end‐to‐end training. The currently used deep learning models can divide the SISR model into four types: interpolation‐based preprocessing‐based model, original image processing based model, hierarchical feature‐based model, and high‐frequency detail‐based model, or shared the network model. The current challenges for super‐resolution reconstruction are mainly reflected in the actual application process, such as encountering an unknown scaling factor, losing paired LR–HR images, and so on.
Kai Li 0047, Shenghao Yang 0003, Runting Dong, Xiaoying Wang 0002, Jianqiang Huang 0002
IET Image Process.4
2020 Survey of external memory large-scale graph processing on a multi-core system
Jianqiang Huang 0001, Xiaoying Wang 0002
J. Supercomput.3
2019 Pimiento: A Vertex-Centric Graph-Processing Framework on a Single Machine
Jianqiang Huang 0002, Xiaoying Wang 0002
ICA3PP (2)3
2018 Independent travel recommendation algorithm based on analytical hierarchy process and simulated annealing for professional tourist
Qingyi Pan, Xiaoying Wang 0002
Appl. Intell.2
2017 Airport Detection Based on a Multiscale Fusion Feature for Optical Remote Sensing Images
abstract
Automatically detecting airports from remote sensing images has attracted significant attention due to its importance in both military and civilian fields. However, the diversity of illumination intensities and contextual information makes this task difficult. Moreover, auxiliary features both within and surrounding the regions of interest are usually ignored. To address these problems, we propose a novel method that uses a multiscale fusion feature to represent the complementary information of each region proposal, which is extracted by constructing a GoogleNet with a light feature module model that has an additional light fully connected layer. Then, the fusion feature is input to a support vector machine whose performance is enhanced using a hard negative mining method. Finally, a simplified localization method is applied to tackle the problem of box redundancy and to optimize the locations of airports. An experiment demonstrates that the fusion feature outperforms other features on airport detection tasks from remote sensing images containing complicated contextual information.
Zhifeng Xiao, Yiping Gong, Yang Long 0002, DeRen Li, Xiaoying Wang 0002
IEEE Geosci. Remote. Sens. Lett.5
2012 An adaptive model-free resource and power management approach for multi-tier cloud environments
Xiaoying Wang 0002, Zhihui Du, Yinong Chen 0004
J. Syst. Softw.1
2011 Research on Adaptive QoS-Aware Resource Reservation Management in Cloud Service Environments
abstract
As cloud computing increasingly enters the commercial domain, the resource management issues are becoming a major concern. The dynamics and complexity of cloud environments pose some challenges in managing the resources to ensure the quality of service (QoS) continuously met under fluctuating workloads. In this paper, we focus on the advance resource reservation management issues for the cloud environments. The architecture of virtualization-based resource reservation is first described and then the formulation of an optimization problem concerning the reservation acceptance decision is presented. To solve this problem, an adaptive QoS-aware reservation management approach is proposed, which calculates the reservation acceptance gain to make the final decision. Performance evaluation results of simulation experiments demonstrate that by using the approach we designed, the QoS of both types of applications could be guaranteed and thus the total revenue of the resource provider could stably keep rising. Detailed analysis shows that proper choices can be made to deal with resource reservation requests in a QoS-aware manner.
Xiaoying Wang 0002, Yuanyuan Xue, Lihua Fan, Zhihui Du
APSCC1
2011 Design of a Robot Cloud Center
abstract
Service-oriented architecture and cloud computing have become the prevalent computing paradigm. In this paradigm, computing resources can be accessed like other utility services available in today's society. In the meantime, robotics applications are joining the trend. More and more robot applications are shifting from manufacture to non-manufacture and service industries. However, for the on-demand supply of the large-scale heterogeneous robots, It is still a problem have not yet been studied, including the fundamental management and efficiency issues in using of these resources. In this paper, we design a framework of "Robot Cloud Center" (RCC) following the general cloud computing paradigm to address the current limitations in capacity and versatility of robotic applications. In this framework, a robot can be provided as a service just like a public utility service so that everyone can access the powerful robotic services easily, efficiently, and cheaply. Based on a given scenario, a robot scheduling algorithm in RCC is proposed to take advantage of the heterogeneous robot resources to meet the end user's requirement with the minimum cost.
Zhihui Du, Weiqiang Yang, Yinong Chen 0004, Xin Sun 0003, Xiaoying Wang 0002
ISADS5
2011 Optimized QoS-aware replica placement heuristics and applications in astronomy data grid
Zhihui Du, Jingkun Hu, Yinong Chen 0004, Zhili Cheng, Xiaoying Wang 0002
J. Syst. Softw.5
2011 Typical Virtual Appliances: An optimized mechanism for virtual appliances provisioning and management
Zhihui Du, Yinong Chen 0004, Xiaoying Wang 0002
J. Syst. Softw.5
2008 A resource management framework for multi-tier service delivery in autonomic virtualized environments
abstract
Large data centers usually host many different services on a shared computing infrastructure, for which on-demand resource management is necessary to maximize providers' revenues by meeting service quality targets at least operational cost. This paper presents a novel architecture of autonomic resource management framework based on virtualized service-oriented computing (SOC) environment. A non-linear continuous optimization problem is defined for adaptive resource allocation and a model-based approach is adopted to solve this problem. Different from traditional approaches, the analytic model we established provides probabilistic performance guarantees and considers non-steady-state behavior assisted by admission control. Results of prototype experiments demonstrate that the performance of multiple services has been greatly improved by taking advantage of fine-grained resource sharing, while incurring much lower resource usage cost. Also, differentiated service qualities could be provided to different client classes through our dynamic resource allocation scheme.
Xiaoying Wang 0002, Dong-Jun Lan, Meng Ye 0001, Ying Chen 0004
NOMS1
2008 Virtualization-based autonomic resource management for multi-tier Web applications in shared data center
Xiaoying Wang 0002, Zhihui Du, Yinong Chen 0004, Sanli Li
J. Syst. Softw.1
2007 Multi-cluster Load Balancing Based on Process Migration
Xiaoying Wang 0002, Zhihui Du, Sanli Li
APPT1