Yanbing Yang 0001

dblp:183/6792-1 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-9266-8600ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
abstract
Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. However, dominated First-Order (FO) optimization is prone to trap models in suboptimal local minima, especially in limited exploration subspace within PEFT. To overcome this challenge, this paper pioneers a systematic exploration of adopting Zeroth-Order (ZO) optimization for PEFT-based VLCL. We first identify the incompatibility of naive full-ZO adoption in VLCL due to optimization process instability. We then investigate the application of ZO optimization from a modality branch-wise to a fine-grained layer-wise across various training units to identify an optimal strategy. Besides, a key theoretical insight reveals that vision modality exhibit higher variance than language counterparts in VLCL during the ZO optimization process, and we propose a modality-aware stabilized ZO strategy, which adopts gradient sign normalization in ZO and constrains vision modality perturbation to further improve performance. Benefiting from the adoption of ZO optimization, PEFT-based VLCL fulfills better ability to escape local minima during the optimization process, extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art results.
Ziwei Liu 0002, Borui Kang, Wei Li 0313, Hangjie Yuan, Yanbing Yang 0001, Yifan Zhu 0001, Tao Feng 0014, Jun Luo 0001
AAAI5
2026 Bi-VLP: Practical Bidirectional Visible Light Positioning via Uplink-Downlink RSS Fusion
Zhongren Jiang, Minggao Cao, Yimao Sun, Yanbing Yang 0001
INFOCOM6
2026 CS-VLP: Lightweight Parameter Update for Cross-Scene Passive Visible Light Positioning
Yihuai Xu, Yuxing Yang, Liangyi Zhang, Tianyi Pan, Yimao Sun, Yanbing Yang 0001
INFOCOM8
2026 FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition
abstract
Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution when low-light, obstructions, or privacy concerns render computer vision impractical. However, thescarcityof labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough offoundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi 2.0, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based, multi-person HAR systems. FM-Fi 2.0 first employs the intrinsic capabilities of FM and RF modality to associate both intra- and cross-modal features of each subject, while simultaneously filtering out irrelevant features to achieve better alignment between the two modalities. FM-Fi 2.0 also employs a cross-modalcontrastiveknowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi 2.0 rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi 2.0's generalizability across various environments.
Yuxuan Weng, Tianyue Zheng, Yanbing Yang 0001, Jun Luo 0001
IEEE Trans. Mob. Comput.3
2025 Extending MPR for Locating a Moving Object Based on TDOA and FDOA
abstract
Modified Polar Representation (MPR) has shown its superiority in integration of near and far field localization for a stationary source. This paper extends the researches and applications of the MPR to moving source, where time difference of arrival (TDOA) and frequency difference of arrival (FDOA) exist due to the relative motion between sensors and source. Theoretical analysis through hybrid Bhattacharyya-Barankin (HBB) bound is provided for evaluating the performance tighter than the Cramér-Rao lower bound (CRLB). We propose a Maximum Likelihood Estimator (MLE) implemented by Gauss-Newton (GN) iteration based on the extended MPR to estimate the position and velocity of source and test the performance under the new MPR. Simulation results confirm that MPR can eliminate the thresholding effect of source position as the source moves far away.
Beichuan Tang, Yimao Sun, Xiantao Heng, Yanbing Yang 0001, Liangyin Chen
ICASSP4
2025 FedUFD: Personalized Edge Computing Using Federated Uncertainty-Driven Feature Distillation
abstract
Recently, federated learning (FL) has been considered a promising and well-suited technique for edge computing applications, such as intelligent traffic control, autonomous driving, and mobile crowdsensing. However, since each edge device may perform individual-specific tasks, they often have heterogeneous data distributions that impact the performance of collaborative training models. Personalized FL (PFL) has then received considerable attention to tackle this problem. Many existing PFL works often employ knowledge distillation to mitigate the negative effects of data heterogeneity. Nevertheless, these works often neglect the fact that the knowledge transferred from the teacher models is not completely correct, which limits the personalization performance of edge devices. In this work, we leverage the knowledge contained in global features to explore the potential of global models and propose a novel uncertainty-driven feature distillation framework called FedUFD. Specifically, we design an uncertainty estimation module in local models, by estimating the uncertainty of the personalized feature distribution, FedUFD can measure the difficulty of learning different personalized features, and then combine the global features to distill the corresponding personalized features. Extensive experiments show that FedUFD outperforms fourteen state-of-the-art PFL frameworks in edge computing, beating the best-performing traditional and personalized baselines by up to 45.45% and 3.55%, respectively.
Zerui Shao, Beibei Li 0002, Zhibo Wang 0001, Yanbing Yang 0001, Peiran Wang, Jun Luo 0001
INFOCOM4
2025 Poster: Towards Backbone-Free VLC Networking via NLOS Optical Channels
abstract
Existing VLC backhaul solutions typically rely on rigid wired (e.g., Ethernet or power-line) or alignment-sensitive LOS links for inter-attocell connectivity, which renders the overall network vulnerable to backbone link failures. In this poster, we introduce a novel network architecture that exploits the inherent non-line-of-sight (NLOS) optical channels between adjacent attocells to enable inter-attocell communication. In particular, we design a chirp signal based on chirp spread spectrum (CSS) modulation to enhance the robustness of communication under low signal-to-noise ratio (SNR) conditions, along with a two-stage window alignment approach to achieve precise synchronization while reducing real-time decoding latency. The potential and feasibility of the newly suggested approach are demonstrated through implementations on both ESP32 and FPGA platforms.
Pinpin Zhang, Yanbing Yang 0001, Yimao Sun, Dié Wu
MobiCom2
2025 LR-ASD: Lightweight and Robust Network for Active Speaker Detection
Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang 0001, Liangyin Chen, Yanru Chen 0001
Int. J. Comput. Vis.5
2025 Exploring the Potential of Utilizing Nonline-of-Sight Channels for Networking in Visible Light Communication
abstract
Visible light communication (VLC), as one of the key technologies for new spectrum communication in 6G, has drawn much attention from both academia and industry. The femtocell-like deployment of VLC in indoor environments gives rise to the concept of optical attocells, where each light-emitting diode (LED) serves as an optical access point (AP), enabling illumination and communication simultaneously. However, the majority of existing optical attocell networks rely on wired backbone links (e.g., Ethernet or power-line connections) for inter-attocell connectivity, and this dependency renders the overall network vulnerable to backbone link failures. To this end, we introduce a novel network architecture that exploits the inherent non-line-of-sight (NLOS) optical channels between adjacent attocells to enable inter-attocell communication, enhancing network resilience and flexibility. In particular, we design a chirp signal based on chirp spread spectrum (CSS) modulation tailored for intensity-modulated VLC to improve the noise resilience in NLOS channels for reliable communication under low signalto-noise ratio (SNR) conditions. Moreover, we propose a twostage window alignment approach that integrates coarse-grained with fine-grained alignment to achieve precise synchronization while reducing real-time decoding latency. Finally, we build two prototypes of NLOS optical attocells based on different hardware platforms and conduct extensive field experiments using three modulation schemes, i.e., on-off keying (OOK), frequency-shift keying (FSK), and CSS, to evaluate their performance under variable parameters. Experimental results indicate that the CSS modulation scheme demonstrates superior robustness than the other two schemes, achieving bit error rates (BER) of 3.1×10-5 and 8.3×10-5 and packet reception rates (PRR) of 98.53% and 99.71% on two platforms, respectively, at a horizontal distance of 8 m between adjacent optical devices.
Pinpin Zhang, Chang Liu 0040, Yimao Sun, Chen Chen 0037, Yanbing Yang 0001, Jun Luo 0001
IEEE Internet Things J.7
2025 Algebraic Solution for Linear Array-Based 3D Localization Without Deployment Limitations
abstract
Localizing a three-dimensional (3D) source using linear arrays (LAs) is a promising new localization technology. Existing solutions are either designed for specific LA deployments, are computationally intensive, or rely on iterative methods that do not guarantee convergence. This paper presents a novel algebraic solution algorithm for 3D source localization using space angle (SA) measurements from LAs. We propose a new formulation of the SA measurement equation, which leads to a constrained weighted least squares (CWLS) problem. Solving it by Lagrangian multipliers, the optimal estimation is obtained with an error correction. The solution does not require specific arrangement and placement of LAs and effectively balances accuracy with computational efficiency. We analyze the performance and complexity of the proposed solution, demonstrating its ability to achieve the Cramér-Rao Lower Bound (CRLB) in the small error region under Gaussian noise with a low computational load. Simulations validate the analysis and confirm the superiority of the proposed solution compared to existing ones.
Beichuan Tang, Yanbing Yang 0001, Liangyin Chen, Yimao Sun
IEEE Signal Process. Lett.3
2025 A Novel Retrospective-Reading Model for Detecting Chinese Sarcasm Comments of Online Social Network
abstract
Through the use of sarcastic sentences on social media, people can express their strong emotions. Therefore, the detection of sarcasm in social media has received more and more attention over the past years. Classifying a sentence as sarcastic or nonsarcastic heavily relies on the contextual information of the sentence. However, only focusing on the features of target text is the main solution of most existing research. Moreover, the scale of publicly available Chinese sarcasm dataset is very small and does not contain the contextual information. To address the issues mentioned above, we build a Chinese sarcasm dataset from Bilibili, which is one of the most widely used social network platforms in China and has a significant number of sarcastic comments and contextual information. As far as we know, our dataset is the first publicly available large-scale Chinese sarcasm dataset including contextual information. Additionally, we have proposed a novel retrospective reading method for detecting sarcasm that leverages contextual information to improve model's performance. The experimental results show the effectiveness of the proposed model and the significance of contextual information for Chinese sarcasm detection: achieving the highest F-score of 0.6942, outperforming existing state-of-the-art (SOTA) approaches. The study presented in this article offers approaches and ideas for future Chinese sarcasm detection studies.
Lei Zhang 0103, Qinfeng Mao, Yanbing Yang 0001, Dong Li 0051, Haizhou Wang 0001
IEEE Trans. Comput. Soc. Syst.4
2025 ReflexGest: Recognizing Hand Gestures Under VLC-Capable Lamps
abstract
As a main approach towards touch-free human-computer interaction,hand gesture recognition(HGR) has long been a research focus for both academia and industry. Meanwhile,visible light communication(VLC) has become increasingly popular with VLC-ready commercial products (e.g., Philips lamps) available on the market. These facts provoke us to ask: can we leverage a VLC-ready lamp to realizeintegrated sensing and communication(ISAC) by conducting both HGR and VLC simultaneously? To this end, we propose ReflexGest as our answer to this question. ReflexGest is implemented upon a table lamp for the sake of practicality; this VLC-ready lamp is equipped with a ring-shaped light-emitting diode (LED) array and a photodiode (PD, for light intensity sensing) originally aiming for up/down-link VLCs. Demanding hand gestures to be performed between the lamp and a table surface, ReflexGest exploits the variation of the reflection and their unique correlation with the corresponding hand gestures to achieve HGR. In particular, ReflexGest first handles the limited sensing ability of the PD by enhancing the LED lamp and thus diversifying the light emission patterns. Moreover, ReflexGest combats the reflection interference from varying table surfaces via an adversarial learning technique to distill only the features relevant to hand gestures. Our extensive evaluations demonstrate that ReflexGest is able to deliver accurate HGR under realistic VLC traffic.
Ziwei Liu 0002, Jifei Zhu, Yimao Sun, Yanbing Yang 0001, Jun Luo 0001
IEEE Trans. Mob. Comput.5
2024 Multidimensional Scaling-Based TDOA Localization in Modified Polar Representation
abstract
Multidimensional scaling (MDS) is an attractive method for location-related applications due to its robustness against noise. This paper applies MDS to time difference of arrival (TDOA) localization in the modified polar representation (MPR) for integrating near-field and far-field localizations. The new MDS formulation yields a constrained optimization problem in terms of the source position. We then propose a computationally efficient and noise robust solution to solve this problem. The solution is closed-form and asymptotically unbiased. It can achieve better mean-square error (MSE) in the large noise region and possibly lower bias than the closed-form solutions from the literature, and also has attractive complexity.
Beichuan Tang, Yimao Sun, K. C. Ho 0001, Lei Zhang 0103, Yanbing Yang 0001
ICASSP5
2024 Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition
abstract
Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution for applications unamenable to techniques requiring computer visions. However, the scarcity of labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough of foundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based HAR systems. FM-Fi involves a novel cross-modal contrastive knowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. It also employs the intrinsic capabilities of FM and RF to remove extraneous features for better alignment between the two modalities. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi's generalizability across various environments.
Yuxuan Weng, Guoquan Wu, Tianyue Zheng, Yanbing Yang 0001, Jun Luo 0001
SenSys4
2024 MIMOCrypt: Multi-User Privacy-Preserving Wi-Fi Sensing via MIMO Encryption
abstract
Wi-Fi signals may help realize low-cost and noninvasive human sensing, yet it can also be exploited by eavesdroppers to capture private information. Very few studies rise to handle this privacy concern so far; they either jam all sensing attempts or rely on sophisticated technologies to support only a single sensing user, rendering them impractical for multi-user scenarios. Moreover, these proposals all fail to exploit Wi-Fi’s multiple-in multiple-out (MIMO) capability. To this end, we propose MIMOCrypt, a privacy-preserving Wi-Fi sensing framework to support realistic multi-user scenarios. To thwart unauthorized eavesdropping while retaining the sensing and communication capabilities for legitimate users, MIMOCrypt innovates in exploiting MIMO to physically encrypt Wi-Fi channels, treating the sensed human activities as physical plaintexts. The encryption scheme is further enhanced via an optimization framework, aiming to strike a balance among i) risk of eavesdropping, ii) sensing accuracy, and iii) communication quality, upon securely conveying decryption keys to legitimate users. We implement a prototype of MIMOCrypt on an SDR platform and perform extensive experiments to evaluate its effectiveness in common application scenarios, especially privacy-sensitive human gesture recognition.
Jun Luo 0001, Hangcheng Cao, Hongbo Jiang 0001, Yanbing Yang 0001, Zhe Chen 0015
SP4
2024 VehicleTalk: Lightweight V2V Network Enabled by Optical Wireless Communication and Sensing
abstract
Platooning has been proven to dramatically increase traffic flow and reduce fuel consumption, and vehicle-to-vehicle (V2V) communication and sensing are requisite for platooning stability. However, most of the existing works only address V2V communication or sensing functions respectively, which is far away from meeting the 6G requirements in availability and synchronization for platooning applications. Inspired by the recent advanced integrated sensing and communication (ISAC), in this paper, we propose VehicleTalk, a lightweight V2V communication and sensing framework. Essentially, VehicleTalk reuses the head/tail LED lights of the vehicles to construct communication/sensing channels for achieving concurrent message exchange and status awareness between adjacent vehicles. In particular, VehicleTalk innovates in both message transmission and vehicle sensing to improve communication robustness and lower system latency for platooning. It leverages Raptor Codes to combat serious packet loss in the V2V network. We also engineer a fast risk detection algorithm by simply monitoring the strength change of the received optical signals from the head/tail lights to improve safety in extreme cases such as emergency brake and cutting-in. Finally, we build a prototype of VehicleTalk with low-cost Commercial Off-The-Shelf (COTS) devices to quickly verify its effectiveness, and the extensive experimental results demonstrate the promising performance of VehicleTalk.
Ruoshen Mo, Pinpin Zhang, Zhengguo Sheng, Yimao Sun, Yanbing Yang 0001
VTC Spring7
2024 Joint Scheduling and Offloading Schemes for Multiple Interdependent Computation Tasks in Mobile Edge Computing
abstract
Mobile edge computing (MEC) can sufficiently meet the computing demands of complex application consists of multiple interdependent tasks which can be represented by a directed acyclic graph (DAG). For tasks in a DAG, different scheduling orders and offloading decisions will generate different completion time, which further affects the Quality of Experiences (QoEs). So it is important to study the scheduling and offloading schemes for tasks in MEC scenarios. To this end, we first designed a scheme that schedules tasks with the highest response ratio and offloads tasks to the optimal processor with the optimization method for a DAG, which is termed as HRRO algorithm. Then, considering the complexity of the reality, we extended the HRRO to the ultradense MEC system and achieved the optimal joint scheduling and offloading scheme for multi-DAG based on the genetic algorithm, which can be concluded as HRRO based on the genetic algorithm (HRRO-GA). Subsequently, to evaluate the performance of the algorithms, we conducted amounts of the simulation experiments and compared the results with several state-of-the-art algorithms, including distributed earliest finish-time offloading (DEFO), potential game-based offloading algorithm (PGOA), and GA-based multiuser earliest finish time (GA-MEFT). Meanwhile, we selected some random strategies to verify the schemes of HRRO-GA are the best. Finally, we concluded that HRRO-GA is more suitable for the ultradense MEC system.
Yanru Chen 0001, Yanbing Yang 0001, Lei Zhang 0103, Liangyin Chen
IEEE Internet Things J.4
2024 A Video Shot Occlusion Detection Algorithm Based on the Abnormal Fluctuation of Depth Information
abstract
To make the video more attractive, original video materials usually need postprocessing by video editors, especially to eliminate low-quality abnormal clips, which seriously affect the visual effect. One of the main reasons for the low-quality abnormal clips is that there are occluders that accidentally break into the shot to occlude the protagonist, resulting in the loss of the video protagonist’s information. However, it is time-consuming and laborious to manually find shot occlusion clips, so computer vision technology can be used to assist editors in completing this work. The previous solutions directly utilize neural networks to detect shot occlusion, so their performance is affected by the size and quality of the dataset. In contrast, inspired by the change of depth information in the frame caused by the occluder breaking into the shot, we propose an algorithm for video shot occlusion detection based on the fluctuation of depth information. This algorithm does not need occlusion data training and can detect shot occlusion well only by capturing the abnormal fluctuations of the frame depth information. Additionally, to overcome the defect in that the first video shot occlusion detection (VSOD) dataset released in our conference publication can only verify the sensitivity of detection methods, we expand the VSOD dataset to evaluate the comprehensive performance of detection algorithms. The plentiful experimental results show that, compared with state-of-the-art occlusion detection methods and self-designed baseline methods, our algorithm significantly improves the comprehensive performance of video shot occlusion detection. Furthermore, through verification on datasets with different data types and distributions, our shot occlusion detection algorithm can maintain an occlusion event recall of over 95%, while the false positive rate does not exceed 3%, demonstrating good generalization ability. To promote reproducible research, the code and dataset are available athttps://github.com/Junhua-Liao/VSOD.
Junhua Liao, Haihan Duan, Wanbing Zhao, Kanghui Feng, Yanbing Yang 0001, Liangyin Chen
IEEE Trans. Circuits Syst. Video Technol.5
2023 A Light Weight Model for Active Speaker Detection
abstract
Active speaker detection is a challenging task in audiovisual scenarios, with the aim to detect who is speaking in one or more speaker scenarios. This task has received considerable attention because it is crucial in many applications. Existing studies have attempted to improve the performance by inputting multiple candidate information and designing complex models. Although these methods have achieved excellent performance, their high memory and computational power consumption render their application to resource-limited scenarios difficult. Therefore, in this study, a lightweight active speaker detection architecture is constructed by reducing the number of input candidates, splitting 2D and 3D convolutions for audio-visual feature extraction, and applying gated recurrent units with low computational complexity for cross-modal modeling. Experimental results on the AVA-ActiveSpeaker dataset reveal that the proposed framework achieves competitive mAP performance (94.1% vs. 94.2%), while the resource costs are significantly lower than the state-of-the-art method, particularly in model parameters (1.0M vs. 22.5M, approximately 23×) and FLOPs (0.6G vs. 2.6G, approximately 4×). Additionally, the proposed framework also performs well on the Columbia dataset, thus demonstrating good robustness. The code and model weights are available at https://github.com/Junhua-Liao/Light-ASD.
Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang 0001, Liangyin Chen
CVPR5
2023 Robust Iterative Solution for Linear Array-Based 3-D Localization by Message Passing
abstract
Recent research has shown that using the 1-D signal arrival angles observed by linear arrays can locate a 3-D source in unique co-ordinates. Current methods to solve this localization problem are based on semidefinite programming (SDP) or gradient-based iteration, which are either computationally demanding or facing divergence or local convergence issues. This paper reformulates the maxi-mum likelihood (ML) estimation of the 3-D localization problem using the factor graph model, where an effective algorithm is designed through message passing. Although iterative, the proposed solution is more robust to measurement noise than the Gauss-Newton (GN) iterative solution, and the complexity is lower than the SDP solution without the need to introduce semidefinite relaxation error. Simulations validate the analytical performance and complexity, and con-firm the superiority on the convergence of the proposed solution.
Yimao Sun, K. C. Ho 0001, Yanbing Yang 0001, Lei Zhang 0103, Liangyin Chen
ICASSP3
2023 SemiGest: Recognizing Hand Gestures via Visible Light Sensing with Fewer Labels
abstract
Human-machine interaction (HMI) is much important in factories, and most HMI ways are contact which arise safety and health issues. To avoid such problems, contactless HMI ways such as in-air hand gesture recognition (HGR) via Wi-Fi or radar are widely studied by both academia and industry. However, these RF-based methods are not very appropriate for the industry because of the electromagnetic interference. As for visible light sensing, it is free of electromagnetic radiation and can reuse the existing devices, e.g., lamps on machines, hence utilizing visible light to realize HGR is a good solution for HMI in factories. The current visible-light-enabled HGR (VL-HGR) methods using deep learning algorithms are all supervised, which increases the cost of manual labeling and further hinders the industrial applications of VL-HGR. To this end, we propose SemiGest, a semi-supervised learning (SSL) method for VL-HGR, to facilitate the applications of VL-HGR in industry. The system prototype is built on a table lamp to mimic the lamp on a machine emitting lights at four distinct carrier frequencies, and the lights reflected by hands are collected by a receiver. SemiGest utilizes the variation and correlation of the lights to realize HGR with an SSL algorithm using only a small amount of labeled data and lots of unlabeled data. Furthermore, the SSL algorithm is designed not only for the visible light data but also can be generalized to other time-series data in the industry. To confirm the effectiveness and robustness of SemiGest, we perform various experiments to show the potential for practical implementation in the industry.
Jifei Zhu, Ziwei Liu 0002, Yimao Sun, Yanbing Yang 0001
MSN4
2023 An Asymptotically Optimal Estimator for Source Location and Propagation Speed by TDOA
abstract
The signal emitted by an acoustic source may be propagating in an environment in which the speed is not known, such as in solid or ocean. Localization of such a source through observing the signal by a number of sensors requires joint estimation with the propagation speed. This work applies the nullspace projection approach to the pseudo-linear formulation for the localization problem to obtain a closed-form solution, which is refined by error-compensation to reach the final estimation. In contrast to the methods from the literature that are either suboptimal or computationally demanding, the proposed method is both statistically and computationally efficient, and is shown analytically to achieve the Cramér-Rao Lower Bound accuracy.
Yimao Sun, K. C. Ho 0001, Yanbing Yang 0001, Liangyin Chen
IEEE Signal Process. Lett.3
2022 A Light Weight Model for Video Shot Occlusion Detection
abstract
The popularity of video social platforms (TikTok, etc.) shows that video is a popular information carrier at present. However, shot occlusion frequently occurs when people are shooting videos to record information. Since the shot occlusion seriously affects the viewers’ experience, the video editors need to find and delete such segments from the video material during post-processing. However, finding the shot occlusion from the video is a time-consuming and laborious task. To reduce the workload of editors, previous researchers proposed a shot occlusion detection algorithm using deep learning technology, which has promotion space in both recognition accuracy and computational efficiency. In this paper, we propose a neural network module, named SAT module, which can effectively extract spatio-temporal information with fewer parameters. We apply SAT module to construct a novel occlusion detection model, and improve the existing occlusion detection loss function for model training. The experimental results on the public dataset show that our method achieves the state-of-the-art performance of 88.25% accuracy and FPS of 130 with the least parameters. Code and models will be available at https://github.com/Junhua-Liao/ICASSP22-OcclusionDetection.
Junhua Liao, Haihan Duan, Wanbin Zhao, Yanbing Yang 0001, Liangyin Chen
ICASSP4
2022 CORE-lens: simultaneous communication and object recognition with disentangled-GAN cameras
abstract
Optical camera communication (OCC) enabled by LED and embedded cameras has attracted extensive attention, thanks to its rich spectrum availability and ready deployability. However, the close interactions between OCC and the indoor spaces have created two major challenges. On one hand, the stripe pattern incurred by OCC may greatly damage the accuracy of image-based object recognition. On the other hand, the patterns inherent to indoor spaces can significantly degrade the decoding performance of reflected OCC. To this end, we propose CORE-Lens as a pipeline to make the mutual interference transparent to existing OR and OCC algorithms. Essentially, CORE-Lens treats the two challenges as two sides of a signal mixture issue: the signals transmitted by OCC get mixed with background images so well that their features become entangled. Consequently, CORE-Lens exploits the idea of disentangled representation learning to separate the mixed signals in the feature space: while the GAN-reconstructed clean background images are used to perform object recognition, OCC decoding is conducted on the residual of the original image after subtracting the reconstructed background. Our extensive experiments on evaluating the real-life performance of CORE-Lens evidently demonstrate its superiority over conventional approaches.
Ziwei Liu 0002, Tianyue Zheng, Yanbing Yang 0001, Yimao Sun, Zhe Chen 0015, Liangyin Chen, Jun Luo 0001
MobiCom4
2022 Computationally Attractive and Location Robust Estimator for IoT Device Positioning
abstract
Locating a device is a basic element for many Internet of Things (IoT) applications. In particular, it often demands an algorithm having low complexity to limit the energy consumption and most important, sufficient robustness without knowing the device in the near-field for point localization or in the far-field for direction of arrival (DOA) estimation. This article proposes a new localization algorithm that can achieve the two purposes, with the theoretical analysis to validate the optimal accuracy and the real data experiment to support the promising performance. The first objective is achieved by a closed-form solution and the second is accomplished by using the modified polar representation (MPR) of the source position, based on a new formulation for the localization problem. While the MPR localization method has been introduced before, it is not sufficiently robust for IoT application to handle the large equal radius (LER) scenario or the presence of sensor position errors. The proposed algorithm uses a different MPR formulation, which is able to handle the LER scenario, sensor position errors, and has low computational complexity.
Yimao Sun, K. C. Ho 0001, Gang Wang 0007, Hongyang Chen 0001, Yanbing Yang 0001, Liangyin Chen, Qun Wan
IEEE Internet Things J.5
2022 Computationally attractive and statistically efficient estimator for noise resilient TOA localization
Yimao Sun, K. C. Ho 0001, Yanbing Yang 0001, Lei Zhang 0103, Liangyin Chen
Signal Process.3
2022 Minimizing the Longest Tour Time Among a Fleet of UAVs for Disaster Area Surveillance
abstract
In this paper, we study the employment of multiple Unmanned Aerial Vehicles (UAVs) to monitor Points of Interests (PoIs) in a disaster area, e.g., collapsed buildings after an earthquake, where the UAVs can take photos and videos for the people trapped at PoIs, because such valuable information is imperative to make rescue decisions. Unlike most existing studies that ignored the monitoring time of PoIs and simply minimized the longest flying distance among the UAVs, we observe that it takes time to monitor the PoIs. Then, it is possible that the flying distance of a UAV in its flying tour may not be too long, the tour however contains many densely-located PoIs. Therefore, it will take a very long time for the UAV to monitor the PoIs in its tour. In this paper, we first formulate a problem of finding flying tours for$K$given UAVs to collaboratively monitor PoIs in a disaster area, such that the maximum spent time of the$K$UAVs among their tours is minimized, where the spent time of a UAV in its tour consists of the flying time and the PoI monitoring time. We then propose a novel$5\frac{1}{3}$-approximation algorithm for the problem, improving the best approximation ratio 6 so far for the problem of minimizing the longest flying distance among the UAVs. In addition, we extend the proposed algorithm to the case that each UAV may not be able to monitor all PoIs assigned to it, due to its limited maximum flying time (e.g., 30 minutes), and the UAV must return to its depot to replace its battery. We finally evaluate the performance of the proposed algorithms via simulation environments, and experimental results show that the proposed algorithms are very promising. Especially, the maximum spent times of the$K$UAVs in their tours by the proposed algorithms are up to 30 percent shorter than those by existing algorithms. In addition, the empirical approximation ratios of the proposed algorithms are no more than 2.4, which are much smaller than their theoretical approximation ratios that are at least$5\frac{1}{3}$.
Qing Guo 0007, Jian Peng 0002, Wenzheng Xu, Weifa Liang, Xiaohua Jia, Zichuan Xu, Yanbing Yang 0001
IEEE Trans. Mob. Comput.7
2021 Pushing the Data Rate of Practical VLC via Combinatorial Light Emission
abstract
Visible light communication (VLC) systems relying on commercial-off-the-shelf (COTS) devices have gathered momentum recently, due to the pervasive adoption of LED lighting and mobile devices. However, the achievable throughput by such practical systems is still several orders below those claimed by controlled experiments with specialized devices. In this paper, we engineer CoLight aiming to boost the data rate of the VLC system purely built upon COTS devices. CoLight adopts COTS LEDs as its transmitter, but it innovates in its simple yet delicate driver circuit wiring an array of LED chips in a combinatorial manner. Consequently, modulated signals can directly drive the on-off procedures of individual chip groups, so that the spatially synthesized light emissions exhibit a varying luminance following exactly the modulation symbols. To obtain a readily usable receiver, CoLight interfaces a COTS PD with a smartphone through the audio jack, and it also has an alternative MCU-driven circuit to emulate a future integration into the phone. The evaluations on CoLight are both promising and informative: they demonstrate a throughput up to 80 kbps at a distance of 2 m, while suggesting various potentials to further enhance the performance.
Yanbing Yang 0001, Jun Luo 0001, Chen Chen 0037, Zequn Chen, Wen-De Zhong, Liangyin Chen
IEEE Trans. Mob. Comput.1
2020 Occlusion Detection for Automatic Video Editing
abstract
Videos have become the new preference comparing with images in recent years. However, during the recording of videos, the cameras are inevitably occluded by some objects or persons that pass through the cameras, which would highly increase the workload of video editors for searching out such occlusions. In this paper, for releasing the burden of video editors, a frame-level video occlusion detection method is proposed, which is a fundamental component of automatic video editing. The proposed method enhances the extraction of spatial-temporal information based on C3D yet only using around half amount of parameters, with an occlusion correction algorithm for correcting the prediction results. In addition, a novel loss function is proposed to better extract the characterization of occlusion and improve the detection performance. For performance evaluation, this paper builds a new large scale dataset, containing 1,000 video segments from seven different real-world scenarios, which could be available at: https://junhua-liao.github.io/Occlusion-Detection/. All occlusions in video segments are annotated frame by frame with bounding-boxes so that the dataset could be utilized in both frame-level occlusion detection and precise occlusion location. The experimental results illustrate that the proposed method could achieve good performance on video occlusion detection compared with the state-of-the-art approaches. To the best of our knowledge, this is the first study which focuses on occlusion detection for automatic video editing.
Junhua Liao, Haihan Duan, Yanbing Yang 0001, Wei Cai 0002, Yanru Chen 0001, Liangyin Chen
ACM Multimedia5
2020 Composite Amplitude-Shift Keying for Effective LED-Camera VLC
abstract
LED-Camera Visible Light Communication (VLC) is gaining increasing attention, thanks to its readiness to be implemented with Commercial Off-The-Shelf devices and its potential to deliver pervasive data services indoors. Nevertheless, existing LEDCamera VLC systems employ mainly low-order modulations such as On-Off Keying (OOK) given the simplicity of their implementation, yet such rudimentary modulations cannot yield a high throughput. In this paper, we investigate various opportunities of using a high-order modulation to boost the throughput of LED-Camera VLC systems, and we decide that Amplitude-Shift Keying (ASK) is the most suitable scheme given the limited operating frequency of such systems. However, directly driving an LED to emit different levels of luminance may suffer heavy distortions caused by the nonlinear behavior of LED. As a result, we innovatively propose to generate ASK using the composition of light emission. In other words, we digitally control the On-Off states of several groups of LED chips, so that their light emissions compose in the air to produce various ASK symbols. We build a prototype of this novel ASK-based VLC system and demonstrate its superior performance over existing systems: it achieves a rate of 2 kbps at a 1 m distance with only a single LED luminaire for static users and more than 1 kbps for mobile users.
Yanbing Yang 0001, Jun Luo 0001
IEEE Trans. Mob. Comput.1
2019 NOMA for MIMO Visible Light Communications: A Spatial Domain Perspective
abstract
In this paper, we propose a novel non-orthogonal multiple access (NOMA) technique from a spatial domain (SD) perspective for indoor multiple-input multiple-output visible light communication (MIMO- VLC) systems. By fully exploiting the spatial distributions of light-emitting diode (LED) transmitters in the ceiling and users over the receiving plane, SD-NOMA is achieved by assigning all the users to different LEDs in the MIMO-VLC system. Hence, each user only receives data from a specific LED and users assigned to the same LED can use the overall modulation bandwidth of the system. Moreover, a signal-to-noise ratio (SNR) based LED selection scheme is further proposed for each user to efficiently select its desired LED. The achievable rates of a general indoor MIMO-VLC system using conventional MIMO orthogonal frequency division multiple access (MIMO-OFDMA) and the proposed SD-NOMA are analytically derived. The superiority of SD-NOMA over conventional MIMO-OFDMA for multi-user MIMO-VLC systems is successfully verified by detailed analytical results.
Chen Chen 0037, Yanbing Yang 0001, Xiong Deng, Pengfei Du 0001, Helin Yang, Zhengchuan Chen, Wen-De Zhong
GLOBECOM2
2019 SynLight: Synthetic Light Emission for Fast Transmission in COTS Device-enabled VLC
abstract
Visible Light Communication (VLC) systems relying on commercial-off-the-shelf (COTS) devices have gathered momentum recently, due to the pervasive adoption of LED lighting and mobile devices. However, the achievable throughput by such practical systems is still several orders below those claimed by controlled experiments with specialized devices. In this paper, we engineer SynLight aiming to significantly improve the data rate of a practical VLC system. SynLight adopts COTS LEDs as its transmitter, but it innovates in its simple yet delicate driver circuit wiring an array of LED chips in a combinatorial manner. Consequently, modulated signals can directly drive the on-off procedures of individual chip groups, so that the spatially synthesized light emissions exhibit a varying luminance following exactly the modulation symbols. To obtain a readily usable receiver, SynLight interfaces a COTS Photo-Diode with a smartphone through the audio jack. The evaluations on SynLight are both promising and informative: they demonstrate a throughput up to 60 kbps, more than 50× of that achieved by state-of-the-art systems, while suggesting various potentials to further enhance the performance.
Yanbing Yang 0001, Jun Luo 0001, Chen Chen 0037, Wen-De Zhong, Liangyin Chen
INFOCOM1
2018 Boosting the Throughput of LED-Camera VLC via Composite Light Emission
abstract
LED-Camera Visible Light Communication (VLC) is gaining increasing attention, thanks to its readiness to be implemented with Commercial Off-The-Shelf devices and its potential to deliver pervasive data services indoors. Nevertheless, existing LED-Camera VLC systems employ mainly low-order modulations such as On-Off Keying (OOK) given the simplicity of their implementation, yet such rudimentary modulations cannot yield a high throughput. In this paper, we investigate various opportunities of using a high-order modulation to boost the throughput of LED-Camera VLC systems, and we decide that Amplitude-Shift Keying (ASK) is the most suitable scheme given the limited operating frequency of such systems. However, directly driving an LED to emit different levels of luminance may suffer heavy distortions caused by the nonlinear behavior of LED. As a result, we innovatively propose to generate ASK using the composition of light emission. In other words, we digitally control the On-Off states of several groups of LED chips, so that their light emissions compose in the air to produce various ASK symbols. We build a prototype of this novel ASK-based VLC system and demonstrate its superior performance over existing systems: it achieves a rate of 2 kbps at a 1 m distance with only a single LED luminaire.
Yanbing Yang 0001, Jun Luo 0001
INFOCOM1
2018 Counting via LED sensing: Inferring occupancy using lighting infrastructure
Yanbing Yang 0001, Jun Luo 0001, Jie Hao 0002, Sinno Jialin Pan
Pervasive Mob. Comput.1
2018 Trinity: Enabling Self-Sustaining WSNs Indoors with Energy-Free Sensing and Networking
abstract
Whereas a lot of efforts have been put on energy conservation in wireless sensor networks (WSNs), the limited lifetime of these systems still hampers their practical deployments. This situation is further exacerbated indoors, as conventional energy harvesting (e.g., solar) may not always work. To enable long-lived indoor sensing, we report in this article a self-sustaining sensing system that draws energy from indoor environments, adapts its duty-cycle to the harvested energy, and pays back the environment by enhancing the awareness of the indoor microclimate through an “energy-free” sensing. First of all, given the pervasive operation of heating, ventilation, and air conditioning (HVAC) systems indoors, our system harvests energy from airflow introduced by the HVAC systems to power each sensor node. Secondly, as the harvested power is tiny, an extremely low but synchronous duty-cycle has to be applied whereas the system gets no energy surplus to support existing synchronization schemes. So, we design two complementary synchronization schemes that cost virtually no energy. Finally, we exploit the feature of our harvester to sense the airflow speed in an energy-free manner. To our knowledge, this is the first indoor wireless sensing system that encapsulates energy harvesting, network operating, and sensing all together.
Feng Li 0002, Yanbing Yang 0001, Zicheng Chi, Yaowen Yang, Jun Luo 0001
ACM Trans. Embed. Comput. Syst.2
2017 ReflexCode: Coding with Superposed Reflection Light for LED-Camera Communication
abstract
As a popular approach to implementing Visible Light Communication (VLC) on commercial-off-the-shelf devices, LED-Camera VLC has attracted substantial attention recently. While such systems initially used reflected light as the communication media, direct light becomes the dominant media for the purpose of combating interference. Nonetheless, the data rate achievable by direct light LED-Camera VLC systems has hit its bottleneck: the dimension of the transmitters. In order to further improve the performance, we revisit the reflected light approach and we innovate in converting the potentially destructive interferences into collaborative transmissions. Essentially, our ReflexCode system codes information by superposing light emissions from multiple transmitters. It combines traditional amplitude demodulation with slope detection to "decode" the grayscale modulated signal, and it tunes decoding thresholds dynamically depending on the spatial symbol distribution. In addition, ReflexCode re-engineers the balanced codes to avoid flicker from individual transmitters. We implement ReflexCode as two prototypes and demonstrate that it can achieve a throughput up to 3.2kb/s at a distance of 3m.
Yanbing Yang 0001, Jiangtian Nie, Jun Luo 0001
MobiCom1
2017 Demo: Coding with Superposed Reflection Light for LED-Camera Communication
abstract
As a popular approach to implementing Visible Light Communication (VLC) on commercial-off-the-shelf devices, LED-Camera VLC has attracted substantial attention recently. While such systems initially used reflected light as the communication media, direct light becomes the dominant media for the purpose of combating interference. Nonetheless, the data rate achievable by direct light LED-Camera VLC systems has hit its bottleneck: the dimension of the transmitters. In this demo, we revisit the reflected light approach and propose a novel modulation mechanism, ReflexCode, which converts the potentially destructive interferences into collaborative transmissions. Essentially, our ReflexCode system codes information by superposing light emissions from multiple transmitters. It combines traditional amplitude demodulation with slope detection to "decode" the grayscale modulated signal, and it tunes decoding thresholds dynamically depending on the spatial symbol distribution. In addition, ReflexCode re-engineers the balanced codes to avoid flicker from individual transmitters. We implement ReflexCode as a prototype and demonstrate that it can achieve a promising throughput.
Yanbing Yang 0001, Jiangtian Nie, Jun Luo 0001
MobiCom1
2017 CeilingSee: Device-free occupancy inference through lighting infrastructure based LED sensing
abstract
As a key component of building management and security, occupancy inference through smart sensing has attracted a lot of research attentions for nearly two decades. Nevertheless, existing solutions mostly rely on either pre-deployed infrastructures or user device participation, thus hampering their wide adoption. This paper presents CeilingSee, a dedicated occupancy inference system free of heavy infrastructure deployments and user involvements. Building upon existing LED lighting systems, CeilingSee converts part of the ceiling-mounted LED luminaires to act as sensors, sensing the variances in diffuse reflection caused by occupants. In realizing CeilingSee, we first re-design the LED driver to leverage LED's photoelectric effect so as to transform a light emitter to a light sensor. In order to produce accurate occupancy inference, we then engineer efficient learning algorithms to fuse sensing information gathered by multiple LED luminaires. We build a testbed covering a 30m2office area; extensive experiments show that CeilingSee is able to achieve very high accuracy in occupancy inference.
Yanbing Yang 0001, Jie Hao 0002, Jun Luo 0001, Sinno Jialin Pan
PerCom1
2017 CeilingTalk: Lightweight Indoor Broadcast Through LED-Camera Communication
abstract
Although Visible Light Communication (VLC) is gaining increasing attention in research, developing a practical VLC system to harness its immediate benefits using Commercial Off-The-Shelf (COTS) devices is still an open issue. To this end, we develop and deploy CeilingTalk as a lightweight wireless broadcast system using COTS LED luminaries as transmitters and smartphone cameras as receivers so that it can be fully hosted in a smartphone and is feasible for all possible indoor environments. CeilingTalk innovates in both encoding and decoding to achieve an adequate throughput for realistic applications. On one hand, it employs Raptor coding to allow multiple LED luminaries to transmit collaboratively so as to benefit both throughput and reliability. On the other hand, it involves a lightweight decoding scheme to handle the asynchrony (both spatial and temporal) in transmissions. Moreover, we analyze the impact of various parameters on the performance of CeilingTalk, in order to derive a model for such VLC systems enabled by COTS devices and hence provide general guidance for future VLC deployments in larger scales. Finally, we conduct extensive field experiments to validate the effectiveness of our LED-camera VLC model, as well as to demonstrate the promising performance of CeilingTalk: up to 1.0 kb/s at a distance of 5 m.
Yanbing Yang 0001, Jie Hao 0002, Jun Luo 0001
IEEE Trans. Mob. Comput.1
2016 CeilingCast: Energy efficient and location-bound broadcast through LED-camera communication
abstract
Although Visible Light Communication (VLC) is gaining increasing attentions in research, developing a practical VLC system to harness its immediate benefits using Commercial Off-The-Shelf (COTS) devices is still an open issue. To this end, we develop and deploy CeilingCast as a location-bound wireless broadcast system using COTS LEDs as transmitters and smartphone cameras as receivers. CeilingCast innovates in its effective coding and efficient decoding schemes, so that it can be fully hosted in a smartphone and is feasible for all possible indoor environments. Moreover, we analyze the impact of various parameters on the performance of CeilingCast, in order to derive a model for such VLC systems enabled by COTS devices and hence provide general guidance for future VLC deployments in larger scales. Finally, we conduct extensive field experiments to validate the effectiveness of our LED-camera VLC model, as well as to demonstrate the promising performance of CeilingCast under various parameters.
Jie Hao 0002, Yanbing Yang 0001, Jun Luo 0001
INFOCOM2