Shixun Wu

dblp:57/10607 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-author · 10 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 cuHPX: GPU-Accelerated Differentiable Spherical Harmonic Transforms on HEALPix Grids
abstract
HEALPix (Hierarchical Equal Area isoLatitude Pixelization) is a widely adopted spherical grid system in astrophysics, cosmology, and Earth sciences. Its equal-area, iso-latitude structure makes it particularly well-suited for large-scale data analysis on the sphere. However, implementing high-performance spherical harmonic transforms (SHTs) on HEALPix grids remains challenging due to irregular pixel geometry, latitude-dependent alignments, and the demands for high-resolution transforms at scale. In this work, we present cuHPX, an optimized CUDA library that provides functionality for spherical harmonic analysis and related utilities on HEALPix grids. Beyond delivering substantial performance improvements, cuHPX ensures high numerical accuracy, analytic gradients for integration with deep learning frameworks, out-of-core memory-efficient optimization, and flexible regridding between HEALPix, equiangular, and other common spherical grid formats. Through evaluation, we show that cuHPX achieves rapid spectral convergence and delivers over 20 times speedup compared to existing libraries, while maintaining numerical consistency. By combining accuracy, scalability, and differentiability, cuHPX enables a broad range of applications in climate science, astrophysics, and machine learning, effectively bridging optimized GPU kernels with scientific workflows.
Xiaopo Cheng, Akshay Subramaniam, Shixun Wu, Noah D. Brenowitz
IPDPS3
2026 Hybrid CIR/ToA UWB Localization via Fingerprint Interval Prediction and GMM-PSO Optimization
abstract
To mitigate the significant degradation of Ultra-Wideband (UWB) positioning accuracy in non-line-of-sight (NLoS) scenarios, this paper proposes an integrated framework that synergizes Channel Impulse Response (CIR) fingerprinting with Time of Arrival (ToA) ranging. Unlike conventional methods that treat fingerprinting and ranging independently, we introduce a novel uncertainty-aware fusion strategy. Specifically, an attention-enhanced Convolutional Neural Network (CNN) is first employed to extract environmental features from CIR data. Crucially, instead of direct coordinate regression, we utilize a Copula-based conformal prediction(CP) method to quantify the positioning uncertainty and construct a reliable spatial search constraint. This constraint then guides a Particle Swarm Optimization (PSO) algorithm to robustly identify the optimal position within the predicted region when ToA ranging errors are modeled by a Gaussian Mixture Model (GMM). This mechanism effectively filters out large outliers by leveraging the complementary strengths of feature matching and geometric ranging. Experimental results on public datasets demonstrate that the proposed method achieves a mean error of 0.0911 meters, outperforming six state-of-the-art algorithms by at least 8.35% while satisfying high computational efficiency requirements.
Shixun Wu, Sen Teng, Zhongwei Hou, Zhangli Lan, Miao Zhang 0018, K. Cumanan
IEEE Internet Things J.1
2025 TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
abstract
GPU-based fast Fourier transform (FFT) is extremely important for scientific computing and signal processing. However, we find the inefficiency of existing FFT libraries and the absence of fault tolerance against soft error. To address these issues, we introduce TurboFFT, a new FFT prototype co-designed for high performance and online fault tolerance. For FFT, we propose an architecture-aware, padding-free, and template-based prototype to maximize hardware resource utilization, achieving a competitive or superior performance compared to the state-of-the-art closed-source library, cuFFT. For fault tolerance, we 1) explore algorithm-based fault tolerance (ABFT) at the thread and threadblock levels to reduce additional memory footprint, 2) address the error propagation by introducing a two-side ABFT with location encoding, and 3) further modify the threadblock-level FFT from 1-transaction to multi-transaction in order to bring more parallelism for ABFT. Our two-side strategy enables online correction without additional global memory while our multi-transaction design averages the expensive threadblock-level reduction in ABFT with zero additional operations. Experimental results on an NVIDIA A100 server GPU and a Tesla Turing T4 GPU demonstrate that TurboFFT without fault tolerance is comparable to or up to 300% faster than cuFFT and outperforms VkFFT. TurboFFT with fault tolerance maintains an overhead of 7% to 15%, even under tens of error injections per minute for both FP32 and FP64.
Shixun Wu, Jinyang Liu 0003, Jiajun Huang 0001, Zizhe Jian, Huangliang Dai, Sheng Di, Franck Cappello, Zizhong Chen
PPoPP1
2025 FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
abstract
Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in large-scale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for Transformers are designed at the operation level without architectural optimization, leading to significant computational and memory overhead, which in turn reduces protection efficiency and limits scalability to larger models. In this paper, we implement module-level protection for Transformers by treating the operations within the attention module as a single kernel and applying end-to-end fault tolerance. This method provides unified protection across multi-step computations, while achieving comprehensive coverage of potential errors in the nonlinear computations. For linear modules, we design a strided algorithm-based fault tolerance (ABFT) that avoids inter-thread communication. Experimental results show that our end-to-end fault tolerance achieves up to 7.56 × speedup over traditional methods with an average fault tolerance overhead of 13.9%.
Huangliang Dai, Shixun Wu, Jiajun Huang 0001, Zizhe Jian, Zizhong Chen
SC2
2025 Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
abstract
As high-performance computing architectures evolve, more scientific computing workflows are being deployed on advanced computing platforms such as GPUs. These workflows can produce raw data at extremely high throughputs, requiring urgent high-ratio and low-latency error-bounded data compression solutions. In this paper, we propose cuSZ-Hi, an optimized high-ratio GPU-based scientific error-bounded lossy compressor with a flexible, domain-irrelevant, and fully open-source framework design. Our novel contributions are: 1) We maximally optimize the parallelized interpolation-based data prediction scheme on GPUs, enabling the full functionalities of interpolation-based scientific data prediction that are adaptive to diverse data characteristics; 2) We thoroughly explore and investigate lossless data encoding techniques, then craft and incorporate the best-fit lossless encoding pipelines for maximizing the compression ratio of cuSZ-Hi; 3) We systematically evaluate cuSZ-Hi on benchmarking datasets together with representative baselines. Compared to existing state-of-the-art scientific lossy compressors, with comparative or better throughput than existing high-ratio scientific error-bounded lossy compressors on GPUs, cuSZ-Hi can achieve up to 249% compression ratio improvement under the same error bound, and up to 215% compression ratio improvement under the same decompression data PSNR.
Shixun Wu, Jinwen Pan, Jinyang Liu 0003, Jiannan Tian, Ziwei Qiu, Jiajun Huang 0001, Kai Zhao 0008, Xin Liang 0001, Sheng Di, Zizhong Chen, Franck Cappello
SC1
2025 TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
abstract
Fourier Neural Operators (FNO) are widely used for learning partial differential equation solution operators. However, FNO lacks architecture-aware optimizations,with its Fourier layers executing FFT, filtering, GEMM, zero padding, and iFFT as separate stages, incurring multiple kernel launches and significant global memory traffic. We propose TurboFNO, the first fully fused FFT-GEMM-iFFT GPU kernel with built-in FFT optimizations. We first develop FFT and GEMM kernels from scratch, achieving performance comparable to or faster than the closed-source SOTA cuBLAS and cuFFT. Additionally, our FFT kernel integrates a built-in high-frequency truncation, input zero-padding, and pruning feature to avoid additional memory copy kernels. To fuse the FFT and GEMM workloads, we propose an FFT variant in which a single thread block iterates over the hidden dimension, aligning with the k-loop in GEMM. Additionally, we design two shared memory swizzling patterns to achieve 100% memory bank utilization when forwarding FFT output to GEMM and enabling the iFFT to retrieve GEMM results directly from shared memory. Experimental result on an NVIDIA A100 GPU shows TurboFNO outperforms PyTorch, cuBLAS, and cuFFT by up to 150%.
Shixun Wu, Huangliang Dai, Zizhong Chen
SC1
2025 LCVAE-CNN: Indoor Wi-Fi Fingerprinting CNN Positioning Method Based on LCVAE
abstract
While Wi-Fi Received Signal Strength Indicator (RSSI) fingerprinting has emerged as a prominent solution for indoor positioning, its accuracy remains hindered by labor-intensive data collection and environmental variability. To overcome these challenges, we propose a novel LCVAE-CNN methodology that integrates a Location-Conditioned Variational Autoencoder (LCVAE) and a multi-task Convolutional Neural Network (CNN) to enhance data quality and positioning performance. The LCVAE employs a dual-encoder architecture to augment RSSI fingerprints by jointly modeling signal features and spatial dependencies, introducing three key innovations: (1) dual-stream encoding that decouples RSSI and location processing for more effective feature learning, (2) a geospatial loss function that enforces topological consistency in the generated data, and (3) conditional data augmentation that preserves physical constraints of indoor spaces. The multi-task CNN then leverages shared feature extraction to jointly optimize classification and regression tasks, enabling efficient and accurate positioning. Extensive evaluations on the UJIIndoorLoc and Tampere datasets demonstrate the superiority of the LCVAE-CNN that achieves 98.80% floor classification accuracy with a Mean Positioning Error (MPE) of 6.79 meters on UJIIndoorLoc, whereas 97.22% accuracy with a MPE of 5.44 meters on the Tampere dataset. Compared to five state-of-the-art methods, it improves floor accuracy by at least 1.9% and reduces MPE by over 19%, while maintaining comparable computational overhead, thereby achieving superior accuracy-efficiency tradeoffs.
Shixun Wu, Xinrui Zeng, Miao Zhang 0018, K. Cumanan, Abdulhamed Waraiet, Zheng Chu 0001
IEEE Internet Things J.1
2024 FT K-Means: A High-Performance K-Means on GPU with Fault Tolerance
abstract
K-means is a widely used algorithm in clustering, how-ever, its efficiency is primarily constrained by the computational cost of distance computing. Existing implementations suffer from suboptimal utilization of computational units and lack resilience against soft errors. To address these challenges, we introduce FT K-means, a high-performance GPU-accelerated implementation of K-means with online fault tolerance. We first present a step-wise optimization strategy that achieves competitive performance compared to NVIDIA's cuML library. We further improve FT K-means with a template-based code generation framework that supports different data types and adapts to different input shapes. A novel warp-level tensor-core error correction scheme is proposed to address the failure of existing fault tolerance methods due to mem-ory asynchronization during copy operations. Our experimental evaluations on NVIDIA T4 GPU and A100 GPU demonstrate that FT K-means without fault tolerance outperforms cuML's K-means implementation, showing a performance increase of 10%-300% in scenarios involving irregular data shapes. Moreover, the fault tolerance feature of FT K-means introduces only an overhead of 11%, maintaining robust performance even with tens of errors injected per second.
Shixun Wu, Yitong Ding, Jinyang Liu 0003, Jiajun Huang 0001, Zizhe Jian, Huangliang Dai, Sheng Di, Bryan M. Wong, Zizhong Chen, Franck Cappello
CLUSTER1
2024 CliZ: Optimizing Lossy Compression for Climate Datasets with Adaptive Fine-tuned Data Prediction
abstract
Benefiting from the cutting-edge supercomputers that support extremely large-scale scientific simulations, climate research has advanced significantly over the past decades. However, new critical challenges have arisen regarding efficiently storing and transferring large-scale climate data among distributed repositories and databases for post hoc analysis. In this paper, we develop CliZ, an efficient online error-controlled lossy compression method with optimized data prediction and encoding methods for climate datasets across various climate models. On the one hand, we explored how to take advantage of particular properties of the climate datasets (such as mask-map information, dimension permutation/fusion, and data periodicity pattern) to improve the data prediction accuracy. On the other hand, CliZ features a novel multi-Huffman encoding method, which can significantly improve the encoding efficiency. Therefore significantly improving compression ratios. We evaluated CliZ versus many other state-of-the-art error-controlled lossy compressors (including SZ3, ZFP, SPERR, and QoZ) based on multiple real-world climate datasets with different models. Experiments show that CliZ outperforms the second-best compressor (SZ3, SPERR, or QoZ1.1) on climate datasets by 20%-200% in compression ratio. CliZ can significantly reduce the data transfer cost between the two remote Globus endpoints by 32%-38%.
Zizhe Jian, Sheng Di, Jinyang Liu 0003, Kai Zhao 0008, Xin Liang 0001, Haiying Xu, Robert Underwood, Shixun Wu, Jiajun Huang 0001, Zizhong Chen, Franck Cappello
IPDPS8
2024 cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
abstract
Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. Compared to CPU-based compressors, GPU-based compressors exhibit substantially higher throughputs, fitting better for today’s HPC applications. However, the critical limitations of existing GPU-based compressors are their low compression ratios and qualities, severely restricting their applicability. To overcome these, we introduce a new GPU-based error-bounded scientific lossy compressor named CUSZ-i, with the following contributions: (1) A novel GPU-optimized interpolation-based prediction method significantly improves the compression ratio and decompression data quality. (2) The Huffman encoding module in CUSZ-i is optimized for better efficiency. (3) CUSZ-i is the first to integrate the NVIDIA Bitcomp-lossless as an additional compression-ratio-enhancing module. Evaluations show that CUSZ-i significantly outperforms other latest GPU-based lossy compressors in compression ratio under the same error bound (hence, the desired quality), showcasing a 476% advantage over the second-best. This leads to CUSZ-i’s optimized performance in several real-world use cases.
Jinyang Liu 0003, Jiannan Tian, Shixun Wu, Sheng Di, Boyuan Zhang 0002, Robert Underwood, Yafan Huang, Jiajun Huang 0001, Kai Zhao 0008, Guanpeng Li, Dingwen Tao, Zizhong Chen, Franck Cappello
SC3
2024 High-performance Effective Scientific Error-bounded Lossy Compression with Auto-tuned Multi-component Interpolation
abstract
Error-bounded lossy compression has been identified as a promising solution for significantly reducing scientific data volumes upon users' requirements on data distortion. For the existing scientific error-bounded lossy compressors, some of them (such as SPERR and FAZ) can reach fairly high compression ratios and some others (such as SZx, SZ, and ZFP) feature high compression speeds, but they rarely exhibit both high ratio and high speed meanwhile. In this paper, we propose HPEZ with newly-designed interpolations and quality-metric-driven auto-tuning, which features significantly improved compression quality upon the existing high-performance compressors, meanwhile being exceedingly faster than high-ratio compressors. The key contributions lie as follows: (1) We develop a series of advanced techniques such as interpolation re-ordering, multi-dimensional interpolation, and natural cubic splines to significantly improve compression qualities with interpolation-based data prediction. (2) The auto-tuning module in HPEZ has been carefully designed with novel strategies, including but not limited to block-wise interpolation tuning, dynamic dimension freezing, and Lorenzo tuning. (3) We thoroughly evaluate HPEZ compared with many other compressors on six real-world scientific datasets. Experiments show that HPEZ outperforms other high-performance error-bounded lossy compressors in compression ratio by up to 140% under the same error bound, and by up to 360% under the same PSNR. In parallel data transfer experiments on the distributed database, HPEZ achieves a significant performance gain with up to 40% time cost reduction over the second-best compressor.
Jinyang Liu 0003, Sheng Di, Kai Zhao 0008, Xin Liang 0001, Sian Jin, Zizhe Jian, Jiajun Huang 0001, Shixun Wu, Zizhong Chen, Franck Cappello
Proc. ACM Manag. Data8
2024 Hybrid Decision-Making for Intelligent High-Speed Train Operation: A Boundary Constraint and Pre-Evaluation Reinforcement Learning Approach
abstract
Deep Reinforcement Learning (DRL) is the most promising technology for improving high-speed train’s energy efficiency and operation quality. Existing solutions, however, suffer from three significant limitations: 1) They cannot effectively constrain the huge exploration space generated by high-speed trains under high temporal deformability and long-distance trips; 2) The reward function has no adaptability to the different energy-efficiency difficulties of different travel schedules, and the agent will receive incorrect reward signals, requiring manual adjustment; 3) They do not avoid the invalid action sequences of the agent well. To address this challenge, we propose a revolutionary Boundary Constrained and Pre-evaluated Reinforcement Learning (BCPRL) approach to alleviate these issues. This approach combines the Shrink Trajectory Exploration Space (STES) module, the Pre-evaluated Energy-efficiency Scenario Complexity (PESC) module, and the Twin Delayed Deep Deterministic Policy Gradient (TD3) module and uses a hybrid of STES and TD3 to make train operation decision-making to improve the operation quality and learning efficiency of the agent. Numerical experiments validate the effectiveness of the BCPRL approach, which, by drastically reducing the exploration space and getting the agent the correct reward signal, not only maintains excellence in efficiency and punctuality but also far surpasses the other baseline approaches in learning efficiency and robustness.
Haotong Zhang 0004, Deqing Huang, Deqiang He, Shixun Wu, Gang Xian
IEEE Trans. Intell. Transp. Syst.5
2023 Exploring Wavelet Transform Usages for Error-bounded Scientific Data Compression
abstract
To address the challenges raised by the data management of exascale scientific data, error-bounded lossy compression has been proposed and well-researched as a prominent solution. Among the existing works, a recent trend leverages wavelet transforms in the error-bounded lossy compression task to effectively capture long-term data correlations within the inputs. Applying those transforms as data preprocessors and decorrelators, wavelet-based lossy compressors have achieved optimized compression rate-distortion on several datasets. However, certain significant limitations of wavelet-based compressors have also been observed: On one hand, attributed to the high computational cost of wavelet transforms, wavelet-based compressors suffer from relatively low computational efficiencies compared to other state-of-the-art compressors. On the other hand, one certain type of wavelet transform cannot perform well on all variations of scientific data. Consequently, to further fine-tune the wavelet-based scientific data lossy compression, more in-depth and systematic research and analysis needs to be conducted. In this paper, based on the FAZ auto-tuning-based modular compression framework, we have integrated a great number of wavelet transforms into the framework and evaluated them with various real-world scientific datasets and fields. From the analysis of those evaluations and the comparison to existing state-of-the-art wavelet-based and non-wavelet-based error-bounded lossy compressors, we conclude and present several essential takeaways for designing and optimizing the wavelet-based scientific error-bounded lossy compressor.
Jiajun Huang 0001, Jinyang Liu 0003, Sheng Di, Zizhe Jian, Shixun Wu, Kai Zhao 0008, Zizhong Chen, Yanfei Guo, Franck Cappello
IEEE Big Data6
2023 FT-GEMM: A Fault Tolerant High Performance GEMM Implementation on x86 CPUs
abstract
General matrix/matrix multiplication (GEMM) is crucial for scientific computing and machine learning. However, the increased scale of the computing platforms raises concerns about hardware and software reliability. In this poster, we present FT-GEMM, a high-performance GEMM being capable of tolerating soft errors on-the-fly. We incorporate the fault tolerant functionality at algorithmic level by fusing the memory-intensive operations into the GEMM assembly kernels. We design a cache-friendly scheme for parallel FT-GEMM. Experimental results on Intel Cascade Lake demonstrate that FT-GEMM offers high reliability and performance -- faster than Intel MKL, OpenBLAS, and BLIS by 3.50%\sim 22.14% for both serial and parallel GEMM, even under hundreds of errors injected per minute.
Shixun Wu, Jiajun Huang 0001, Zizhe Jian, Zizhong Chen
HPDC1
2023 Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs
abstract
General Matrix Multiplication (GEMM) is a crucial algorithm for various applications such as machine learning and scientific computing since an efficient GEMM implementation is essential for the performance of these calculations. While researchers often strive for faster performance by using large computing platforms, the increased scale of these systems can raise concerns about hardware and software reliability. In this paper, we present a design of a high-performance GPU-based GEMM that integrates an algorithm-based fault tolerance scheme that detects and corrects silent data corruptions at computing units on-the-fly. We explore fault-tolerant designs for GEMM at the thread, warp, and threadblock levels, and also provide a baseline GEMM implementation that is competitive with or faster than the state-of-the-art, closed-source cuBLAS GEMM. We present a kernel fusion strategy to overlap and mitigate the memory latency due to fault tolerance with the original GEMM computation. To support a wide range of input matrix shapes and reduce development costs, we present a template-based approach for automatic code generation for both fault-tolerant and non-fault-tolerant GEMM implementations. We evaluate our work on NVIDIA Tesla T4 and A100 server GPUs. Our experimental results demonstrate that our baseline GEMM shows comparable or superior performance compared to the closed-source cuBLAS. Compared with the prior state-of-the-art non-fused fault-tolerant GEMM, our optimal fused strategy achieves a 39.04% speedup on average. In addition, our fault-tolerant GEMM incurs only a minimal overhead (8.89% on average) compared to cuBLAS even with hundreds of errors injected per minute. For irregularly shaped inputs, the code generator-generated kernels show remarkable speedups of 160% ~ 183.5% and 148.55% ~ 165.12% for fault-tolerant and non-fault-tolerant GEMMs, respectively, which outperforms cuBLAS by up to 41.40%.
Shixun Wu, Jinyang Liu 0003, Jiajun Huang 0001, Zizhe Jian, Bryan M. Wong, Zizhong Chen
ICS1
2021 Time Slot Detection-Based M -ary Tree Anticollision Identification Protocol for RFID Tags in the Internet of Things
abstract
Recently, a number of articles have proposed query tree algorithms based on bit tracking to solve the multitag collision problem in radio frequency identification systems. However, these algorithms still have problems such as idle slots and redundant prefixes. In this paper, a time slot detection‐based M‐ary tree (Time Slot Detection based M‐ary tree, TSDM) tag anticollision algorithm has been proposed. When a collision occurs, the reader sends a predetection command to detect the distribution of the m‐bit ID in the 2m subslots; then, the time slot after predetection is processed according to the format of the frame‐like. The idle time slots have been eliminate through the detection. Using a frame‐like mode, only the frame start command carries parameters, and the other time slot start commands do not carry any parameters, thereby reducing the communication of each interaction. Firstly, the research status of the anticollision algorithm is summarized, and then the TSDM algorithm is explained in detail. Finally, through theoretical analysis and simulation, it is proved that the time cost of the TSDM algorithm proposed in this paper is reduced by 12.57%, the energy cost is reduced by 12.65%, and the key performance outperforms the other anticollision algorithms.
Xiaojiao Yang, Bizao Wu, Shixun Wu, Xinxin Liu 0017, W. G. Will Zhao
Wirel. Commun. Mob. Comput.3
2019 Probability Weighting Localization Algorithm Based on NLOS Identification in Wireless Network
abstract
In this paper, a localization scenario that the home base station (BS) measures time of arrival (TOA) and angle of arrival (AOA) while the neighboring BSs only measure TOA is investigated. In order to reduce the effect of non-line of sight (NLOS) propagation, the probability weighting localization algorithm based on NLOS identification is proposed. The proposed algorithm divides these range and angle measurements into different combinations. For each combination, a statistic whose distribution is chi-square in LOS propagation is constructed, and the corresponding theoretic threshold is derived to identify each combination whether it is LOS or NLOS propagation. Further, if those combinations are decided as LOS propagation, the corresponding probabilities are derived to weigh the accepted combinations. Simulation results demonstrate that our proposed algorithm can provide better performance than conventional algorithms in different NLOS environments. In addition, computational complexity of our proposed algorithm is analyzed and compared.
Shixun Wu, Darong Huang 0002
Wirel. Commun. Mob. Comput.1
2011 An Improved Reference Selection Method in Linear Least Squares Localization for LOS and NLOS
abstract
Linear least squares (LLS) estimation is a low complexity but sub-optimum method for estimating the location of a mobile terminal (MT) from some measured distances. It requires selecting one of the known fixed terminals (FTs) as a reference FT for obtaining a linear set of expressions. In this paper, a new method for selecting the reference FT is proposed, which selects the reference FT based on the minimum residual rather than the smallest measured distance and improves the localization accuracy significantly in Line of sight (LOS) environment. In Non-line of sight (NLOS) environment, a new residual weighting algorithm is proposed, which is based on the proposed LOS algorithm. Moreover, by making use of the interior-point optimization method which can obtain the estimation of NLOS error, a new algorithm which is also based on the proposed LOS algorithm is proposed. Simulation results show that the proposed methods improve the positioning accuracy considerably comparing with other existing methods.
Shixun Wu, Jiping Li, Shouyin Liu
VTC Fall1