Le Yang 0001

dblp:79/2888-1 · DBLP profile ↗
← Back
54ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0001-7945-6323ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 13 · 5 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Computer networks · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-complexity BCJR symbol detection for sparse channels
Jun Tao 0004, Le Yang 0001
Signal Process.3
2026 Algebraic FDOA-only geolocation with three observers
Wenjun Zhang 0001, Xi Li 0020, Le Yang 0001, Fucheng Guo 0001
Signal Process.4
2026 From Images to Point Clouds: An Efficient Solution for Cross-Media Blind Quality Assessment Without Annotated Training
abstract
We present a novel quality assessment method which can predict the perceptual quality of point clouds from new scenes without available annotations by leveraging the rich prior knowledge in images, called the Distribution-Weighted Image-Transferred Point Cloud Quality Assessment (DWIT-PCQA). Recognizing the human visual system (HVS) as the decision-maker in quality assessment regardless of media types, we can emulate the evaluation criteria for human perception via neural networks and further transfer the capability of quality prediction from images to point clouds by leveraging the prior knowledge in the images. Specifically, domain adaptation (DA) can be leveraged to bridge the images and point clouds by aligning feature distributions of the two media in the same feature space. However, the different manifestations of distortions in images and point clouds make feature alignment a difficult task. To reduce the alignment difficulty and consider the different distortion distributions during alignment, we have derived formulas to decompose the optimization objective of the conventional DA into two suboptimization functions with distortion as a transition. Specifically, through network implementation, we propose the distortion-guided biased feature alignment which integrates existing/estimated distortion distribution into the adversarial DA framework, emphasizing common distortion patterns during feature alignment. Besides, we propose the quality-aware feature disentanglement to mitigate the destruction of the mapping from features to quality during alignment with biased distortions. Experimental results demonstrate that our proposed method exhibits reliable performance compared to general blind PCQA methods without needing point cloud annotations.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 On the Efficient Adaptive Streaming of 3D Gaussian Splatting Over Dynamic Networks
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising representation for immersive media. Its explicit splat-based structure offers high visual quality and real-time rendering, making it particularly suitable for six degrees of freedom streaming applications. However, its deployment in practical streaming scenarios is still limited due to several key challenges such as the large data volume, and insufficient support for dynamic bitrate adaptation under fluctuating network conditions. This paper presents an efficient 3DGS streaming framework that operates directly on pre-generated 3DGS models without retraining or fine-tuning. First, a training-free perceptual pruning method, which removes visually redundant Gaussians according to the human visual system metrics, is introduced. The resulting 3DGS is then encoded into a compact representation using the extended 3D codecs, exploiting its point-based structure. Next, we build a scene-specific bitrate ladder through analyzing the trade-off between resolution, bitrate, and perceptual quality. This enables efficient and fine-grained representation selection. Finally, a progressive streaming mechanism is developed. It is driven by a reinforcement learning scheduler that adaptively decides whether to download new content or enhance previously buffered content based on real-time network feedback. Experiments on real-world 3DGS datasets and bandwidth traces show that the proposed method evidently improves the quality of experience and streaming efficiency in various network scenarios.
Mufan Liu, Qi Yang 0003, Le Yang 0001, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.5
2026 Automatic Sleep Staging of Single-Channel Ear-EEG Signals With a Probabilistic Ensemble Learning Approach
abstract
Accurate sleep staging is crucial for the early diagnosis of neurodegenerative diseases and the management of sleep disorders. To provide a user-friendly, non-intrusive, and long-term monitoring solution, we explored the potential clinical applications of ear-electroencephalogram (ear-EEG). This study proposes a probabilistic ensemble learning approach for automatic sleep staging using single-channel ear-EEG data. The proposed method integrates Extreme Gradient Boosting (XGBoost) with Linear Discriminant Analysis (LDA), augmented by transition matrix correction and probability weighting strategies, to capture temporal sleep patterns without compromising data integrity or requiring intensive preprocessing. An ear-EEG with polysomnography (ear-PSG) dataset collected from twenty subjects using our custom-developed ear-EEG sensor, was compared with two public datasets, ear-Feature and Sleep-EDF, to validate both the reliability of the data and the effectiveness of the proposed approach. The results indicate that transition matrix correction is particularly effective when training and testing are conducted using single-epoch inputs, whereas model weighting demonstrates greater stability as the number of epochs increases. When using seven-epochs input sequences, leave-one-subject-out (LOSO) cross-validation achieved 0.814 accuracy with 0.749 kappa coefficient on ear-PSG (earL-R), and 0.841 accuracy with 0.779 kappa coefficient on the ear-Feature dataset. The design of a single-channel cross-ear intra-auricular ear-EEG configuration, combined with an ensemble learning framework, effectively balances device portability and classification performance, offering new insights for the clinical translation of wearable sleep monitoring technology and laying a foundation for the development of portable sleep monitoring devices.
Hongyu Liang, Yongxuan Wang, Le Yang 0001, Meimei Wu
IEEE J. Biomed. Health Informatics3
2025 On the Variational Gaussian Filtering with Natural Gradient Descent
abstract
Variational Gaussian filter (VGF) approximates the intractable posterior of the state of a non-linear non-Gaussian system using a single Gaussian density normally found through Kullback-Leibler divergence minimization. This paper focuses on the VGFs whose measurement update is realized by employing the natural gradient descent (NGD). Under the assumption that the state predictive distribution is also Gaussian, we re-examine the iterative NGD-based measurement update under two different parameterizations of the Gaussian posterior. The first one consists of the mean and covariance, while the other comprises the mean and precision matrix (i.e., the inverse of the covariance). Their NGD-based update rules are derived in an alternative but unified way using matrix calculus. They are compared against each other and with the one developed using the natural parameterization of the Gaussian density. Important new insights are obtained. Modifications to the established update rules, which guarantee the positive definiteness of the covariance/precision matrix of the Gaussian posterior, are re-visited as well. Simulations are used to corroborate the theoretical results and evaluate the performance of the developed algorithms in range-bearing tracking.
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova
FUSION3
2025 HybridGS: High-Efficiency Gaussian Splatting Data Compression using Dual-Channel Sparse Representation and Point Cloud Encoder
abstract
Most existing 3D Gaussian Splatting (3DGS) compression schemes focus on producing compact 3DGS representation via implicit data embedding. They have long encoding and decoding times and highly customized data format, making it difficult for widespread deployment. This paper presents a new 3DGS compression framework called HybridGS, which takes advantage of both compact generation and standardized point cloud data encoding. HybridGS first generates compact and explicit 3DGS data. A dual-channel sparse representation is introduced to supervise the primitive position and feature bit depth. It then utilizes a canonical point cloud encoder to carry out further data compression and form standard output bitstreams. A simple and effective rate control scheme is proposed to pivot the interpretable data compression scheme. HybridGS does not include any modules aimed at improving 3DGS quality during generation. But experiment results show that it still provides comparable reconstruction performance against state-of-the-art methods, with evidently faster encoding and decoding speed. The code is publicly available at https://github.com/Qi-Yangsjtu/HybridGS .
Qi Yang 0003, Le Yang 0001, Geert Van der Auwera, Zhu Li 0001
ICML2
2025 Video Streaming with Kairos: An MPC-Based ABR with Streaming-Aware Throughput Prediction
abstract
Throughput prediction in current adaptive bitrate (ABR) schemes often neglects streaming-aware characteristics, such as sequence irregularity and prediction smoothness, resulting in inaccurate predictions and suboptimal performance. To address these challenges, we propose Kairos, an MPC-based ABR scheme that integrates an attention-based throughput predictor with buffer-aware uncertainty control to enhancing both prediction accuracy and adaptability to dynamic network conditions. Specifically, Kairos employs a multi-time attention network (mTAN) to process irregularly sampled streaming data, producing uniformly spaced latent representations. Based on these, we introduce a percentile prediction network to estimate future throughput percentiles, along with a buffer-aware uncertainty control module that selects the optimal percentile based on the current buffer status. As smoothness is another key component of QoE, we incorporate a smoothness regularizer to ensure consistent throughput predictions, thereby facilitating smoother ABR decisions. Our Kairos design integrates sampling irregularity, prediction uncertainty, and smoothness into the throughput prediction, significantly enhancing bitrate decision making within the MPC framework. Extensive trace-driven and real-world experiments demonstrate that Kairos outperforms state-of-the-art ABR schemes, achieving a QoE improvement ranging from 6.42% to 29.45% across diverse network conditions.
Ziyu Zhong, Mufan Liu, Le Yang 0001, Yiling Xu, Jenq-Neng Hwang
NOSSDAV3
2025 Indoor 3-D Localization Based on Heterogeneous Sensing Data Fusion
abstract
To address the challenges of better balancing the cost and positioning accuracy in indoor localization and mitigating the performance degradation due to signal blocking and multipath propagation in complex three-dimensional (3D) space, this paper presents a 3D localization system that integrates heterogeneous sensing data. Specifically, by combining the cost-effective Bluetooth Low Energy (BLE) and high-ranging-precision Ultra-Wideband (UWB) technologies, the developed scheme realizes 3D indoor localization using no more than two UWB base stations (BSs). The system can dynamically switch between the single BS and dual BS modes according to the changes in the target mobility and environmental interference, enhancing its adaptability to various scenarios. The positioning results are found through solving a geometrically constrained optimization problem. A BS selection mechanism is further introduced to improve the performance. Real-world experimental results demonstrate that the proposed system has significantly enhanced indoor localization accuracy, with improvements ranging from 23.3% to 42.1% compared to conventional methods. Its deployment requires only at most two UWB BS and a limited number of BLE hotspots, which demonstrates its significant potentials in practical applications.
Ningning Qin, Le Yang 0001
IEEE Internet Things J.3
2024 On the Gaussian Filtering for Nonlinear Dynamic Systems Using Variational Inference
abstract
This paper introduces a new variational Gaussian filtering approach for estimating the state of a nonlinear dynamic system. We first assume that the predictive distribution of the state is Gaussian and derive an iterative method for updating the state posterior in the natural parameter space through KullbackLeibler divergence minimization. The obtained update rule is the same as that of the conjugate-computation variational inference technique in Bayesian learning. The derivation here is simpler and more insightful. We then impose a Wishart prior on the inverse of the state prediction covariance to take into account the impact of approximating the state predictive distribution using a Gaussian density on the state posterior estimation. The prediction covariance is identified jointly with the state using variational inference and the established state posterior update rule to achieve the desired Gaussian filtering. Simulation study examines the performance of the proposed filtering framework in target tracking based on bearing and range measurements.
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova
FUSION3
2024 Mesla: Neural Adaptive Layered Point Cloud Streaming with Enhanced Buffer Management
abstract
Point cloud video provides an immersive experience with six degrees of freedom (6DoF). However, streaming video imposes significant challenges due to the huge data volume involved, which increases the transmission burden and makes it difficult to maintain a high quality of experience (QoE) in bandwidth-constrained networks. Current methods either employ pre-specified control rules or make irreversible decisions to optimize the QoE, hindering adaptation to the varying network conditions and not employing buffer management. In this work, we present Mesla1, an adaptive layered point cloud streaming framework, which is empowered by deep reinforcement learning (DRL) and hierarchical buffer management. By maximizing the utility-driven QoE via proximal policy optimization (PPO)-based DRL, Mesla makes bitrate decisions based on the obtained network observations and viewing behaviors of users. The viewport-oriented utility is incorporated into the QoE objective for assigning point cloud tiles with different priorities. To achieve fine-grained rate adaptation, we partition point clouds into layers with different levels of details (LoD) and develop a buffering heuristic (i.e., the hierarchical buffer management) that enhances the quality of unconsumed content. Extensive simulations demonstrate the advantages of Mesla compared to the benchmarking schemes, corroborating the effectiveness of the proposed system.
Mufan Liu, Puyue Hou, Le Yang 0001, Yiling Xu
GLOBECOM4
2024 Enhanced Face Recognition using Intra-class Incoherence Constraint
abstract
The current face recognition (FR) algorithms has achieved a high level of accuracy, making further improvements increasingly challenging. While existing FR algorithms primarily focus on optimizing margins and loss functions, limited attention has been given to exploring the feature representation space. Therefore, this paper endeavors to improve FR performance in the view of feature representation space. Firstly, we consider two FR models that exhibit distinct performance discrepancies, where one model exhibits superior recognition accuracy compared to the other. We implement orthogonal decomposition on the features from the superior model along those from the inferior model and obtain two sub-features. Surprisingly, we find the sub-feature perpendicular to the inferior still possesses a certain level of face distinguishability. We adjust the modulus of the sub-features and recombine them through vector addition. Experiments demonstrate this recombination is likely to contribute to an improved facial feature representation, even better than features from the original superior model. Motivated by this discovery, we further consider how to improve FR accuracy when there is only one FR model available. Inspired by knowledge distillation, we incorporate the intra-class incoherence constraint (IIC) to solve the problem. Experiments on various FR benchmarks show the existing state-of-the-art method with IIC can be further improved, highlighting its potential to further enhance FR performance.
Yuanqing Huang 0002, Yinggui Wang, Le Yang 0001, Lei Wang 0251
ICLR3
2024 EVAN: Evolutional Video Streaming Adaptation via Neural Representation
abstract
Adaptive bitrate (ABR) using conventional codecs cannot further modify the bitrate once a decision has been made, exhibiting limited adaptation capability. This may result in either overly conservative or overly aggressive bitrate selection, which could cause either inefficient utilization of the network bandwidth or frequent re-buffering, respectively. Neural representation for video (NeRV), which embeds the video content into neural network weights, allows video reconstruction with incomplete models. Specifically, the recovery of one frame can be achieved without relying on the decoding of adjacent frames. NeRV has the potential to provide high video reconstruction quality and, more importantly, pave the way for developing more flexible ABR strategies for video transmission. In this work, a new framework, named Evolutional Video streaming Adaptation via Neural representation (EVAN), which can adaptively transmit NeRV models based on soft actor-critic (SAC) reinforcement learning, is proposed. EVAN is trained with a more exploitative strategy and utilizes progressive playback to avoid re-buffering. Experiments showed that EVAN can outperform existing ABRs with 50% reduction in re-buffering and achieve nearly 20% improvement in users’ quality of experience (QoE).
Mufan Liu, Le Yang 0001, Yiling Xu, Ye-Kui Wang, Jenq-Neng Hwang
ICME2
2024 Enhancing Real-Time Video Streaming with Joint Frame Size and Rate Adaptation
abstract
With advancing network technologies, real-time communication (RTC) scenarios like cloud gaming and video conferencing have gained more attention. However, when addressing the challenge of meeting users’ high-quality demands while dealing with network fluctuations, existing adaptive bitrate (ABR) or adaptive framerate (AFR) algorithms encounter limitations in enhancing Quality of Experience (QoE). This paper introduces Adaptive Frame Size and Rate (AFSR), a joint adaption algorithm based on DRL. AFSR dynamically adjusts the frame rate and frame size (magnitude of bits), and improves QoE in RTC by accurately assessing inter-frame quality and its impact on latency and stall. AFSR uses non-linear bitrate-quality relationships and precise end-to-end latency measurements. Comparative evaluations confirm AFSR’s superior performance in RTC video transmission.
Hengchao Wang, Ziyu Zhong, Jiaoyang Yin, Yiling Xu, Le Yang 0001
ISCAS5
2024 Annealed SOR-Based Gibbs Sampler for MIMO Detection in Frequency-Selective Channels
abstract
Asuccessive over-relaxation (SOR) based Gibbs sampler has been proposed for Markov chain Monte Carlo (MCMC) symbol detection in multiple-input multiple-output (MIMO) communication systems. It outperforms the conventional standard Gibbs sampler with even faster convergence. Such an advantage, however, is achieved only in the high signal-to-noise ratio (SNR) region. This paper proposes an annealed version of the SOR-based Gibbs sampler, not only expanding the working SNR range significantly but also improving the detection performance. Besides, we extend the MCMC symbol detection originally developed under frequency-flat channels to the more general case with frequency-selective channels. Numerical results corroborate the superiority of the proposed detection scheme.
Le Yang 0001, Jun Tao 0004, Yan Huang 0018
IEEE Signal Process. Lett.2
2024 TCDM: Transformational Complexity Based Distortion Metric for Perceptual Point Cloud Quality Assessment
abstract
The goal of objective point cloud quality assessment (PCQA) research is to develop quantitative metrics that measure point cloud quality in a perceptually consistent manner. Merging the research of cognitive science and intuition of the human visual system (HVS), in this article, we evaluate the point cloud quality by measuring the complexity of transforming the distorted point cloud back to its reference, which in practice can be approximated by the code length of one point cloud when the other is given. For this purpose, we first make space segmentation for the reference and distorted point clouds based on a 3D Voronoi diagram to obtain a series of local patch pairs. Next, inspired by the predictive coding theory, we utilize a space-aware vector autoregressive (SA-VAR) model to encode the geometry and color channels of each reference patch with and without the distorted patch, respectively. Assuming that the residual errors follow the multi-variate Gaussian distributions, the self-complexity of the reference and transformational complexity between the reference and distorted samples are computed using covariance matrices. Additionally, the prediction terms generated by SA-VAR are introduced as one auxiliary feature to promote the final quality prediction. The effectiveness of the proposed transformational complexity based distortion metric (TCDM) is evaluated through extensive experiments conducted on five public point cloud quality assessment databases. The results demonstrate that TCDM achieves state-of-the-art (SOTA) performance, and further analysis confirms its robustness in various scenarios.
Qi Yang 0003, Xiaozhong Xu, Le Yang 0001, Yiling Xu
IEEE Trans. Vis. Comput. Graph.5
2023 On the Approximation of the Quotient of Two Gaussian Densities for Multiple-Model Smoothing
abstract
The quotient of two multivariate Gaussian densities can be written as an unnormalized Gaussian density, which has been applied in some recently developed multiple-model fixed-interval smoothing algorithms. However, this expression is invalid if instead of being positive definite, the covariance of the unnormalized Gaussian density is indefinite (i.e., it has both positive and negative eigenvalues) or undefined (i.e., computing it requires inverting a singular matrix). This paper considers approximating the quotient of two Gaussian densities in this case using two different approaches to mitigate the caused numerical problems. The first approach directly replaces the indefinite covariance of the unnormalized Gaussian density with a positive definite matrix nearest to it. The second approach computes the approximation through solving, using the natural gradient, an optimization problem with a Kullback-Leibler divergence-based cost function. This paper illustrates the application of the theoretical results by incorporating them into an existing smoothing method for jump Markov systems and utilizing the obtained smoothers to track a maneuvering target.
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova, Yanbo Xue
FUSION3
2023 Optimizing Distributed Multi-Sensor Multi-Target Tracking Algorithm Based On Labeled Multi-Bernoulli Filter
abstract
In this paper, we propose an improved distributed fusion algorithm under the Labeled multi-Bernoulli (LMB) filter framework. Firstly, the LMB parameter set is augmented by a new group variable, which is able to record the matching information of the neighbour sensors. Then the matching LMB components of the survival targets between the sensors can be fused directly by checking the group variable, greatly reducing the time cost of the matching calculation and the interference from the newborn targets. While for the newborn targets, the Murty algorithm is employed and only performed once to find the best matching relation between the sensors. Finally, experimental results show that the proposed algorithm offers a better tracking performance than the state-of-the-art R-GCI-LMB algorithm with lower computational complexity and higher tracking accuracy.
Honggang Liu, Jinlong Yang 0002, Le Yang 0001
ICASSP4
2023 Privacy-Preserving End-to-End Spoken Language Understanding
abstract
Spoken language understanding (SLU), one of the key enabling technologies for human-computer interaction in IoT devices, provides an easy-to-use user interface. Human speech can contain a lot of user-sensitive information, such as gender, identity, and sensitive content. New types of security and privacy breaches have thus emerged. Users do not want to expose their personal sensitive information to malicious attacks by untrusted third parties. Thus, the SLU system needs to ensure that a potential malicious attacker cannot deduce the sensitive attributes of the users, while it should avoid greatly compromising the SLU accuracy. To address the above challenge, this paper proposes a novel SLU multi-task privacy-preserving model to prevent both the speech recognition (ASR) and identity recognition (IR) attacks. The model uses the hidden layer separation technique so that SLU information is distributed only in a specific portion of the hidden layer, and the other two types of information are removed to obtain a privacy-secure hidden layer. In order to achieve good balance between efficiency and privacy, we introduce a new mechanism of model pre-training, namely joint adversarial training, to further enhance the user privacy. Experiments over two SLU datasets show that the proposed method can reduce the accuracy of both the ASR and IR attacks close to that of a random guess, while leaving the SLU performance largely unaffected.
Yinggui Wang, Wei Huang 0039, Le Yang 0001
IJCAI3
2023 Point Cloud Quality Assessment: Dataset Construction and Learning-based No-reference Metric
abstract
Full-reference (FR) point cloud quality assessment (PCQA) has achieved impressive progress in recent years. However, in many cases, obtaining the reference point clouds is difficult, so no-reference (NR) metrics have become a research hotspot. Few researches about NR-PCQA are carried out due to the lack of a large-scale PCQA dataset. In this article, we first build a large-scale PCQA dataset named LS-PCQA, which includes 104 reference point clouds and more than 22,000 distorted samples. In the dataset, each reference point cloud is augmented with 31 types of impairments (e.g., Gaussian noise, contrast distortion, local missing, and compression loss) at 7 distortion levels. Besides, each distorted point cloud is assigned with a pseudo-quality score as its substitute of Mean Opinion Score. Inspired by the hierarchical perception system and considering the intrinsic attributes of point clouds, we propose a NR metric ResSCNN based on sparse convolutional neural network (CNN) to accurately estimate the subjective quality of point clouds. We conduct several experiments to evaluate the performance of the proposed NR metric. The results demonstrate that ResSCNN exhibits the state-of-the-art performance among all the existing NR-PCQA metrics and even outperforms some FR metrics. The dataset presented in this work will be made publicly accessible at https://smt.sjtu.edu.cn . The source code for the proposed ResSCNN can be found at https://github.com/lyp22/ResSCNN .
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Privacy-Preserving Face Recognition in the Frequency Domain
abstract
Some applications may require performing face recognition (FR) on third-party servers, which could be accessed by attackers with malicious intents to compromise the privacy of users’ face information. This paper advocates a practical privacy-preserving FR scheme without key management realized in the frequency domain. The new scheme first collects the components of the same frequency from different blocks of a face image to form component channels. Only part of the channels are retained and fed into the analysis network that performs an interpretable privacy-accuracy trade-off analysis to identify channels important for face image visualization but not crucial for maintaining high FR accuracy. For this purpose, the loss function of the analysis network consists of the empirical FR error loss and a face visualization penalty term, and the network is trained in an end-to-end manner. We find that with the developed analysis network, more than 94% of the image energy can be dropped while the face recognition accuracy stays almost undegraded. In order to further protect the remaining frequency components, we propose a fast masking method. Effectiveness of the new scheme in removing the visual information of face images while maintaining their distinguishability is validated over several large face datasets. Results show that the proposed scheme achieves a recognition performance and inference time comparable to ArcFace operating on original face images directly.
Yinggui Wang, Le Yang 0001
AAAI4
2022 On the Fixed-Interval Smoothing for Jump Markov Nonlinear Systems
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova, Yanbo Xue
FUSION3
2021 Enhanced Fixed-Interval Smoothing for Markovian Switching Systems
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova, Bing Deng
FUSION3
2021 Kld Minimization-Based Constrained Measurement Filtering For Two-Step TDOA Indoor Tracking
abstract
This paper presents an enhanced two-step method for tracking an indoor point target using the time difference of arrival (TDOA) measurements from an ultra wideband (UWB) positioning system. Again, the algorithm preprocesses the raw TDOAs and then feeds the results to a recursively bounded grid-based filter (RBGF) for position tracking. Different from the state-of-the-art, inequality constraints on the true TDOAs from the RBGF are exploited in the preprocessing step through constrained Kullback-Leibler divergence (KLD) minimization. In particular, a semidefinite programming (S-DP) problem is formulated and solved to find a Gaussian TDOA posterior closest in terms of KLD to the unconstrained one while satisfying all inequality constraints. Simulations show that the newly developed algorithm outperforms the one we recently proposed to impose the inequality constraints via probability density function (PDF) truncation.
Le Yang 0001, Jun Tao 0004, Yanbo Xue
ICASSP2
2021 Visual Quality Optimization for View-Dependent Point Cloud Compression
abstract
The video-based point cloud compression (V-PCC) is the state-of-the-art dynamic point cloud compression technique. V-PCC projects the 3D point cloud data patch by patch to its bounding box and organizes projected patches into a video frame, making full use of the well-developed video coding tools. Despite its high efficiency, cracks easily exist in the reconstructed point cloud in various viewing angles, which seriously degrades the visual quality. In this paper, we propose an efficient method to improve the visual quality of dynamic point cloud, especially for the main view from the content provider. The relationship between patches and views is exploited, and an algorithm intelligently reserving points that may be discarded in V-PCC is proposed. According to our subjective and perceptual objective evaluation experiments, compared with V-PCC, the overall visual quality of the reconstructed point could is evidently improved. In particular, cracks are mended with our proposed method. The Bjontegaard delta bit-rate reduction of up to 3.1% is achieved with respect to Point Cloud Quality Metric (PCQM), which partially verifies the improvement of subjective quality when adopting the proposed method.
Danying Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001
ISCAS5
2021 Point-Voting based Point Cloud Geometry Compression
abstract
The Geometry-based Point Cloud Compression (G-PCC) proposed by the Moving Picture Experts Group (MPEG) is the state-of-art point cloud compression algorithm. It provides an efficient lossy geometry compression technique called triangle soup (Trisoup) for static point clouds. Based on the pruned octree structure, Trisoup provides a local surface model consisting of multiple triangles and compresses vertices of the triangles instead of directly compressing the positions of the original points. Accordingly, we propose a point-voting based method to improve the triangle-construction within each leaf node. This new method leverages the node-based points distribution for more precise vertices determination, which better fits the local surface. Experimental results demonstrate the effectiveness of our point-voting based method for both objective and subjective quality evaluation.
Chaofei Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001
MMSP5
2021 A deep learning framework for autonomous flame detection
Lyudmila Mihaylova, Le Yang 0001
Neurocomputing3
2021 Energy-Efficient Resource Allocation for Blockchain-Enabled Industrial Internet of Things With Deep Reinforcement Learning
abstract
Industrial Internet of Things (IIoT) has emerged with the developments of various communication technologies. In order to guarantee the security and privacy of massive IIoT data, blockchain is widely considered as a promising technology and applied into IIoT. However, there are still several issues in the existing blockchain-enabled IIoT: 1) unbearable energy consumption for computation tasks; 2) poor efficiency of consensus mechanism in blockchain; and 3) serious computation overhead of network systems. To handle the above issues and challenges, in this article, we integrate mobile-edge computing (MEC) into blockchain-enabled IIoT systems to promote the computation capability of IIoT devices and improve the efficiency of the consensus process. Meanwhile, the weighted system cost, including the energy consumption and the computation overhead, are jointly considered. Moreover, we propose an optimization framework for blockchain-enabled IIoT systems to decrease consumption, and formulate the proposed problem as a Markov decision process (MDP). The master controller, offloading decision, block size, and computing server can be dynamically selected and adjusted to optimize the devices energy allocation and reduce the weighted system cost. Accordingly, due to the high-dynamic and large-dimensional characteristics, deep reinforcement learning (DRL) is introduced to solve the formulated problem. Simulation results demonstrate that our proposed scheme can improve system performance significantly compared to other existing schemes.
Le Yang 0001, Meng Li 0007, Pengbo Si, Ruizhe Yang, Enchang Sun, Yanhua Zhang
IEEE Internet Things J.1
2020 Outlier-Robust Schmidt-Kalman Filter Using Variational Inference
abstract
The Schmidt-Kalman filter (SKF) achieves filtering consistency in the presence of biases in system dynamic and measurement models through accounting for their impacts when updating the state estimate and covariance. However, the performance of the SKF may break down when the measurements are subject to non-Gaussian and heavy-tail noise. To address this, we impose the Wishart prior distribution on the precision matrix of measurement noise, such that the measurement likelihood now has heavier tails than the Gaussian distribution to deal with the potential occurrence of outliers. Variational inference is invoked to establish analytically tractable methods for computing the posterior of the system state, system biases, and the measurement noise precision matrix. The principle of the SKF considers the effect of system biases but does not actively estimate them when two variants of outlier-robust SKFs are incorporated. We evaluate their performance in terms of estimation accuracy and filtering consistency using simulations and real-world data. Promising results are obtained.
Xi Li 0020, Yanbo Xue, Stephen John Weddell, Le Yang 0001, Lyudmila Mihaylova
FUSION5
2020 Robust Tdoa Indoor Tracking Using Constrained Measurement Filtering and Grid-Based Filtering
abstract
This paper considers exploiting the time difference of arrival (TDOA) measurements from a ultra wideband (UWB) indoor positioning system to locate a moving point target. In indoor environments, measured TDOAs are subject to large errors due to multipath and/or non-line-of-sight (NLOS) propagation. Besides, they are nonlinearly related to the target position. This paper presents an enhanced two-step approach to achieve robust TDOA indoor tracking. Similar to the existing method, the first-step of the new algorithm preprocesses the raw TDOAs to mitigate the effect of large TDOA errors while its second step applies a recursively bounded grid-based filter (RBGF) to achieve target position tracking. To improve performance, in this work, the possible target position area, which is explicitly obtained by the RGBF, is fed back to the first-step such that a constrained TDOA measurement preprocessing is now performed. Extensive simulation results show that the newly proposed scheme offers better TDOA estimation and indoor target positioning accuracy over the original method and other benchmark algorithms.
Jun Tao 0004, Le Yang 0001, Yanbo Xue, Qisong Wu
ICASSP3
2020 Fast Video Saliency Detection based on Feature Competition
abstract
In this paper, we propose a light video saliency prediction model, named SalFCM, which achieves fixation detection rate of 110fps. It is known that the human attention is captured by objects that have always been present or newly appeared. To model this dynamic change, we propose an Inter-frame Feature Competition Module (IFCM) to make an adaptive choice between correlated and differential features of consecutive frames. Besides, it is noted that saliency is better explained by low-level rather than high-level features in some visual scenes. Hence, we design a Hierarchical Feature Competition Module (HFCM) to balance the influence of low-level and high-level features. Our model achieves a good trade-off between precision and processing speed. The developed SalFCM is evaluated on three video saliency datasets: DHF1K, Hollywood-2 and UCF-sports. We conduct ablation studies to verify the effectiveness of the proposed model.
Hang Yan 0006, Yiling Xu, Jun Sun 0005, Le Yang 0001, Wei Huang 0012
VCIP4
2019 Enhanced Multiple Model GPB2 Filtering Using Variational Inference
Xi Li 0020, Lyudmila Mihaylova, Le Yang 0001, Stephen John Weddell, Fucheng Guo 0001
FUSION4
2019 Multi-Band Image Fusion Using Gaussian Process Regression with Sparse Rational Quadratic Kernel
Fodio S. Longman, Lyudmila Mihaylova, Le Yang 0001, Konstantinos N. Topouzelis
FUSION3
2019 Joint Optimization of Networking and Computing Resources for Green M2M Communications Based on DRL
abstract
Recent advances in Internet of Things (IoT) provide plenty of opportunities for various areas. Nevertheless, the machine-to-machine (M2M) communications-based IoT develops rapidly but suffers from extra energy consumption, large data transmission latency as well as overmuch network cost, because various of machine-type communication devices (MTCDs) are deployed in the network. To meet the requirements of energy efficient M2M communications, in this paper, we introduce a promising technology named as mobile edge computing (MEC), and propose a performance optimization framework with MEC for M2M communications network based on deep reinforcement learning (DRL). According to dynamic decision process by DRL, the appropriate access networks and the computing servers can be determined and selected with the minimum system cost, which includes lower network cost, time cost and energy consumption for data transmission and computing tasks execution. Extensive simulation results with different system parameters show that our proposed framework can effectively improve the system performance for M2M communications compared to the existing schemes.
Meng Li 0007, Le Yang 0001, F. Richard Yu, Zhuwei Wang, Yanhua Zhang
GLOBECOM2
2019 Vehicle Positioning and Ranging with Static Traffic Camera based on 2D-3D Tracking and Re-Projection
abstract
Vehicle positioning and ranging are the current research hotspots. To attain the competition goal of the MMSP Witcomm Challenge 2019 and promote the development of the autonomous driving technology, a novel framework is proposed in this paper, which combines the 2D object tracking and 3D reprojection methodologies. Firstly, a correlation filter using the deep convolutional features is designed to detect the bounding box of the moving objects which is achieved via finding the maximum response of the initial object in the convolutional feature maps. Next the homography matrix is calculated based on the image coordinate points and the corresponding world coordinate points to project the vehicle 2D boundary into the real world. The validity of the proposed framework is verified using the dataset provided by the MMSP Witcomm Challenge 2019 competition, where our method was awarded the second-place prize.
Zhan Song, Yipeng Liu 0003, Yiling Xu, Le Yang 0001
MMSP4
2019 TDOA source positioning in the presence of outliers
abstract
Source localisation using time difference of arrival (TDOA) measurements has drawn considerable attention in the past few decades. The presence of outliers in TDOAs could deteriorate the localisation performance significantly. Under the reasonable assumption that outliers are sparse and they do not dominate in the raw measurement set, a computationally efficient outlier‐robust TDOA localisation method is proposed in this study. It integrates the half‐quadratic minimisation and reweighted least absolute shrinkage and selection operator to iteratively identify the outliers and find the source location estimate using the TDOA inliers only. Both additive and multiplicative forms of the proposed method are established. Simulation results demonstrate the computational efficiency and effectiveness of the developed algorithms.
Fuhe Ma, Le Yang 0001, Fucheng Guo 0001
IET Signal Process.2
2018 Ensemble Kalman Filtering for Online Gaussian Process Regression and Learning
abstract
Gaussian process regression is a machine learning approach which has been shown its power for estimation of unknown functions. However, Gaussian processes suffer from high computational complexity, as in a basic form they scale cubically with the number of observations. Several approaches based on inducing points were proposed to handle this problem in a static context. These methods though face challenges with real-time tasks and when the data is received sequentially over time. In this paper, a novel online algorithm for training sparse Gaussian process models is presented. It treats the mean and hyperparameters of the Gaussian process as the state and parameters of the ensemble Kalman filter, respectively. The online evaluation of the parameters and the state is performed on new upcoming samples of data. This procedure iteratively improves the accuracy of parameter estimates. The ensemble Kalman filter reduces the computational complexity required to obtain predictions with Gaussian processes preserving the accuracy level of these predictions. The performance of the proposed method is demonstrated on the synthetic dataset and real large dataset of UK house prices.
Danil Kuzin, Le Yang 0001, Olga Isupova, Lyudmila Mihaylova
FUSION2
2018 Enhanced GMM-Based Filtering with Measurement Update Ordering and Innovation-Based Pruning
abstract
The Gaussian mixture model (GMM) has been extensively investigated in nonlinear/non-Gaussian filtering problems. This paper presents two enhancements for GMM-based nonlinear filtering techniques, namely, the adaptive ordering of the measurement update and normalized innovation square (NIS)-based mixture component management. The first technique selects the order of measurement update by maximizing the marginal measurement likelihood to improve performance. The second approach takes the filtering history of a mixture component into account and prunes those components with NIS larger than a threshold to eliminate their impact on the filtering posterior. The advantage of the proposed enhancements is illustrated via simulations that consider source tracking using the time difference of arrival (TDOA) and frequency difference of arrival (FDOA) measurements received at two unmanned aerial vehicles (UAVs). A GMM-cubature quadrature Kalman filter (CQKF) is implemented and its performances with different measurement update and mixture component management strategies are compared. The superior performance obtained via the use of the two proposed techniques is demonstrated.
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova, Fucheng Guo 0001
FUSION2
2018 A Gaussian Process Regression Approach for Fusion of Remote Sensing Images for Oil Spill Segmentation
abstract
Synthetic Aperture Radar (SAR) satellite systems are very efficient in oil spill monitoring due to their capability to operate under all weather conditions. This paper presents a framework using Gaussian process (GP) to fuse SAR images of different modalities and to segment dark areas (assumed oil spill) for oil spill detection. A new covariance function; a product of an intrinsically sparse kernel and a Rational Quadratic Kernel (RQK) is used to model the prior of the estimated image allowing information to be transferred. The accuracy performance evaluation demonstrates that the proposed framework has 37% less RMSE per pixel and a compelling enhancement visually when compared with existing methods.
Fodio S. Longman, Lyudmila Mihaylova, Le Yang 0001
FUSION3
2017 Dual-satellite source geolocation with time and frequency offsets and satellite location errors
abstract
This paper considers locating a static source on Earth using the time difference of arrival (TDOA) and frequency difference of arrival (FDOA) measurements obtained by a dual-satellite geolocation system. The TDOA and FDOA from the source are subject to unknown time and frequency offsets because the two satellites are imperfectly time-synchronized or frequency-locked. The satellite locations are not known accurately as well. To make the source position identifiable and mitigate the effect of satellite location errors, calibration stations at known positions are used. Achieving the maximum likelihood (ML) geolocation performance usually requires jointly estimating the source position and extra variables (i.e., time and frequency offsets as well as satellite locations), which is computationally intensive. In this paper, a novel closed-form geolocation algorithm is proposed. It first fuses the TDOA and FDOA measurements from the source and calibration stations to produce a single pair of TDOA and FDOA for source geolocation. This measurement fusion step eliminates the time and frequency offsets while taking into account the presence of satellite location errors. The source position is then found via standard TDOA-FDOA geolocation. The developed algorithm has low complexity and performance analysis shows that it attains the Cramér-Rao lower bound (CRLB) under Gaussian noises and mild conditions. Simulations using a challenging scenario with a short-baseline dual-satellite system verify the theoretical developments and demonstrate the good performance of the proposed algorithm.
Chao Liu 0056, Le Yang 0001, Lyudmila Mihaylova
FUSION2
2017 Moving target localization in multistatic sonar using time delays, Doppler shifts and arrival angles
abstract
Identifying the location of a target is a fundamental application in multistatic sonar. Numerous attempts have been made to improve the accuracy, computational efficiency and robustness of target positioning. Previous studies mostly use time delay and angle measurements for localization, or time delays and Doppler shifts if relative motions exist among the transmitters, target and receivers. This paper considers the joint use of time delay, Doppler shift and angle measurements to locate a moving target. We develop an explicit algebraic solution to the problem, and illustrate the benefit of using all three kinds of measurements. The proposed solution is shown by theoretical performance analysis and confirmed by simulations to be able to reach the Cramer-Rao Bound (CRB) accuracy under Gaussian noise, when the noise level is not significant.
Liu Yang 0018, Le Yang 0001, K. C. Ho 0001
ICASSP2
2016 Source localization using a moving receiver and noisy TOA measurements
Fucheng Guo 0001, Le Yang 0001, Wenli Jiang
Signal Process.3
2016 Moving Target Localization in Multistatic Sonar by Differential Delays and Doppler Shifts
abstract
A moving target creates the Doppler effect on the transmitted signal, which can be exploited to improve the target localization accuracy in multistatic sonar that normally utilizes differential delay time measurements only. In this letter, we first examine the contribution of Doppler measurements via the Cr$\acute{\text{a}}$mer–Rao lower bound (CRLB) study, and then develop an algebraic closed-form solution for the moving target localization problem. The proposed algorithm is shown in both theory and simulation to be able to reach the CRLB performance under Gaussian noise, when the measurement error is small.
Liu Yang 0018, Le Yang 0001, K. C. Ho 0001
IEEE Signal Process. Lett.2
2015 TOA-based joint synchronization and source localization with random errors in sensor positions and sensor clock biases
Yinggui Wang, Le Yang 0001, Yanbo Xue
Ad Hoc Networks3
2014 Compressive detection of stochastic signals with the measurement matrix not necessarily orthonormal
Yinggui Wang, Le Yang 0001, Zheng Liu 0012, Fucheng Guo 0001, Wenli Jiang
FUSION2
2014 Block-sparse signal recovery with synthesized multitask compressive sensing
abstract
The paper considers the problem of reconstructing blocks-sparse signals. A new algorithm, called synthesized multitask compressive sensing (SMCS), is proposed. In contrast to existing methods that rely on the availability of the sparsity structure information, the SMCS algorithm resorts to the multitask compressive sensing (MCS) technique for signal recovery. The SMCS algorithm synthesizes new compressive sensing (CS) tasks via circular-shifting operations and utilizes the minimum description length (MDL) principle to determine the proper set of the synthesized CS tasks for signal reconstruction. An outstanding advantage of SMCS is that it can achieve good signal reconstruction performance without using prior information on the block-sparsity structure. Simulations corroborate the theoretical developments.
Yinggui Wang, Zheng Liu 0012, Wenli Jiang, Le Yang 0001
ICASSP4
2014 Improving noisy sensor positions using accurate inter-sensor range measurements
Ming Sun 0002, Le Yang 0001, Fucheng Guo 0001
Signal Process.2
2013 An efficient closed-form solution for joint synchronization and localization using TOA
Yanbo Xue, Le Yang 0001
Future Gener. Comput. Syst.3
2012 Circle fitting using semi-definite programming
abstract
The fitting of a collection of noisy data points to a circle is a nonlinear and challenging problem, and it plays an important role in many signal processing applications. This paper proposes a semi-definite programming solution for the circle fitting problem based on the semi-definite relaxation technique. The relaxation of the maximum likelihood estimation converts a nonconvex problem to an approximate but convex one that can be solved by using the semi-definite programming method. The performance of the proposed solution is examined via simulations and compared with the K?asa method.
Zhenhua Ma, Le Yang 0001, K. C. Ho 0001
ISCAS2
2012 Accurate sequential self-localization of sensor nodes in closed-form
Ming Sun 0002, Le Yang 0001, K. C. Ho 0001
Signal Process.2
2012 Efficient Joint Source and Sensor Localization in Closed-Form
abstract
This letter considers the problem of simultaneously locating multiple disjoint sources and refining erroneous sensor positions using TDOA measurements. The previous work by Yang and Hoto solve this problem cannot provide optimum accuracy for the sensor positions. The proposed estimator improves the previous method so that both the source and the sensor position estimates can achieve the Cramer–Rao lower bound (CRLB) accuracy. The theoretical derivation is corroborated by simulations.
Ming Sun 0002, Le Yang 0001, K. C. Ho 0001
IEEE Signal Process. Lett.2
2010 On using multiple calibration emitters and their geometric effects for removing sensor position errors in TDOA localization
abstract
The use of calibration emitters is known to be able to improve TDOA source localization accuracy when sensor positions are not accurate. This paper derives through CRLB analysis the conditions under which the sensor position errors can be completely eliminated in a source location estimate via deploying multiple calibration emitters whose positions can be erroneous. The implications on the geometric arrangements of calibration emitters to satisfy the conditions are elaborated. In particular, to fully remove the effect of sensor position errors, we need a sufficient number of calibration emitters and they together cannot lie in the same plane with any sensor. The theoretical developments are supported by simulations.
Le Yang 0001, K. C. Ho 0001
ICASSP1
2009 Solutions and comparison of Maximum Likelihood and Full-Least-Squares estimations for circle fitting
abstract
The fitting of a number of noisy data points with a circle has found numerous applications in image processing and pattern recognition. This paper examines two methods to estimate the circle parameters: the Maximum Likelihood (ML) method and the Full-Least-Squares (FLS) method. The ML method is based on the noisy model from the data while the FLS method minimizes the geometric distance square. We first provide the iterative solutions of them using Taylor-series linearization approach. We then show analytically that FLS does not yield the ML solution. This is in contrast to previous study that the FLS method gives the same solution as ML. FLS method approximates the ML estimation only if the noise power is much less than the circle radius square. Simulations are included to support the theoretical development.
Zhenhua Ma, K. C. Ho 0001, Le Yang 0001
ICASSP3
2007 Decoupled echo state networks with lateral inhibition
Yanbo Xue, Le Yang 0001, Simon Haykin 0001
Neural Networks2