EDBT 2026 Demo / reviewers in the wild / expert
Caili Guo
dblp:45/7794
· DBLP profile ↗
107ranked-venue papers
4as first author
56since 2021 · last 2026
0000-0001-8892-4520ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 52 · 3 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Aware Video Communication: Enhancing Traditional and Deep Video Encoders
Xiangben Zhu, Caili Guo, Yang Yang 0057, Chuanhong Liu, Kuiyuan Ding |
WCNC | 2 |
| 2026 | Adaptive U-Shaped Split Federated Learning for Image Coding at Resource-Constrained UAV NetworkabstractU-shaped Split Federated Learning (U-SFL) has been widely applied in image coding, as it can effectively balance parallel training, model privacy, and local computational cost. The performance of U-SFL largely depends on the selection of split points and the aggregation frequency, and thus, model splitting (MS) and model aggregation (MA) strategies are critical. In UAV scenarios, fluctuating communication links and computational resources can significantly impact U-SFL performance. To address this challenge, we propose a resource-adaptive U-SFL (AU-SFL) framework, which adaptively selects optimal MS and MA strategies based on the available computational and communication resources. Specifically, we first conduct a theoretical convergence analysis that systematically quantifies the individual and joint impacts of MS and MA strategies on convergence behavior. Then, we formulate an optimization problem aimed at minimizing training latency, grounded in the theoretical convergence analysis. Subsequently, an alternating MS-MA algorithm is proposed to solve the problem, which is decomposed into two subproblems and solved alternatively. Extensive experiments on multiple benchmark datasets demonstrate that AU-SFL achieves a 1.61 times reduction in convergence time while improving multi-scale structural similarity index (MS-SSIM) by 4.4%, conclusively validating our adaptive optimization strategy’s effectiveness. Caili Guo, Yang Yang 0057, Chuanhong Liu, Lin Hu 0008 |
IEEE Internet Things J. | 2 |
| 2025 | Joint Time-frequency-energy Resource Allocation for RIS-Assisted Underground Heterogeneous IoTabstractThe utilization of urban underground spaces poses new challenges for underground Internet of Things (IoT), such as energy constraint of sensor nodes (SNs) and channel blockage caused by complex spatial structures. Existing approaches rely on centralized access points for wireless energy transfer (WET), suffering from severe energy loss over long distances. Besides, time duplex energy-data transmission is inefficient. To address these issues, we propose a Reconfigurable Intelligent Surface (RIS)-assisted frequency division duplex (FDD) underground IoT architecture, which facilitates concurrent WET to SNs and uplink heterogeneous data transmission across different frequency bands. To minimize the weighted sum of energy consumption and trans-mission pressure—a metric designed to characterize transmission-task execution situations, we formulate an optimization problem that jointly coordinates the time-frequency-energy resource blocks and RIS phase shifts under heterogeneous Quality of Service (QoS) requirements. Then, we design a parameter-sharing Multi-Agent Proximal Policy Optimization (PS-MAPPO) algorithm which stably converging to the optimal solution of the proposed problem while reducing the complexity of the network. Simulation results demonstrate that the proposed method is superior in energy saving, latency guarantee and improving the task completion rate compared to existing methods. Biling Zhang, Caili Guo, Yang Yang 0057 |
GLOBECOM | 3 |
| 2025 | LLM4KG: Large Language Model Enabled Physical Layer Key Generation Scheme with Non-Ideal Channel Reciprocity
Zhenyang Mo, Caili Guo, Yang Yang 0057, Kuiyuan Ding |
GLOBECOM | 2 |
| 2025 | Task-Oriented Semantic Communication with Large Language Model Enabled Knowledge BaseabstractLarge Language Models (LLMs) have the potential to greatly enhance the performance of task-oriented semantic communication (TOSC) through their extensive knowledge. However, hallucinations from LLMs may cause semantic noise, thus degrading task performance. To address this issue, we propose a TOSC scheme with an LLM-enabled knowledge base (TOSCLKB). In the considered system, the transmitter transmits the features of the source over a wireless channel to accomplish downstream tasks at the receiver. Meanwhile, the transmitter utilizes the LLM-enabled Knowledge Base (KB) to provide additional data for downstream tasks. To effectively exploit the extensive knowledge of LLMs while mitigate its hallucination, a cross-domain fusion codec framework with a hallucination filtering phase and a cross-domain fusion phase is proposed. In particular, the first phase filters out data irrelevant to the source generated by the LLM-enabled KB based on semantic similarity. Then, a cross-domain fusion phase is proposed which fuses source data with LLM-generated data based on their semantic importance, thereby enhancing task performance. Experiment results on the text-based person retrieval task demonstrate that the proposed TOSC-LKB can achieve up to 26.7 % and 7.1 % performance gains without introducing additional time overhead compared to DeepSC and TOSC-LKB without LLMs. Wuxia Hu, Caili Guo, Yang Yang 0057, Chunyan Feng |
ICC | 2 |
| 2025 | Information Bottleneck Guided Joint Source-Channel Coding with HARQabstractDeep joint source-channel coding (JSCC) with hybrid automatic repeat request (HARQ) for image transmission has attracted increasing attention due to its flexibility and high efficiency. Existing researches mainly focus on minimizing the distortion of mutiple retransmissions while ignoring the redundancy in retransmitted signal and such redundancy may lead bandwidth waste and reconstruction quality degradation. In this paper, we propose an information bottleneck (IB) guided deep JSCC with HARQ system (HARQ-IBJSC), which aims at improving the reconstruction quality by compressing the redundancy in retransmitted signal. In particular, we first design a new IB objective for deep JSCC with HARQ system, which simultaneously reduces redundancy in the retransmitted signal and minimizes image transmission distortion. Since the mutual information terms in the designed IB objective is intractable, we then derive a differentiable lower bound on the IB objective and use the bound as the loss function of HARQ-IBJSC. Experimental results show that the proposed HARQ-IBJSC system can increase PSNR by up to 1 dB. Haoxuan Zhang, Lunan Sun, Caili Guo, Yang Yang 0057 |
WCNC | 3 |
| 2025 | Attribute-Aware Implicit Modality Alignment for text attribute person search
Fangfang Liu 0003, Xin Wang 0203, Zheng Li 0014, Caili Guo, Yang Yang 0057, Lin Hu 0008 |
Knowl. Based Syst. | 4 |
| 2025 | Multi-view visual semantic embedding for cross-modal image-text retrieval
Zheng Li 0014, Caili Guo, Xin Wang 0203, Hao Zhang 0161, Lin Hu 0008 |
Pattern Recognit. | 2 |
| 2025 | A unified framework of data augmentation using large language models for text-based cross-modal retrieval
Lijia Si, Caili Guo, Zheng Li 0014, Yang Yang 0057 |
Pattern Recognit. | 2 |
| 2025 | Spatial Geometry Theory-Based Algorithm for Accurate Positioning With Single Vision BeaconabstractQuick response (QR) code is one of the most popular beacons for indoor positioning. Existing QR code-based algorithms typically require several QR codes for precise positioning, which may need intense QR code deployment or have limited coverage. This paper proposes a visual positioning algorithm based on feature points and triangular structures on a QR code (V-FTQ), which can achieve precise position and orientation estimation using only a monocular camera and a single QR code beacon. In the considered system, a few QR codes are placed on the ceiling, and a user holds a camera in hand to capture the QR code for positioning. In particular, we first propose an efficient matching mechanism to match the feature points of the captured QR beacon to their projections in the image plane of the camera. Then, based on the geometry relations between the feature points and their projections, V-FTQ can calculate the coordinates of the feature points in the camera coordinate. Finally, the orientation and position of the user are estimated based on linear algebra and spatial geometry. Moreover, since V-FTQ may have a few ambiguous solutions, a singular solution elimination scheme is further proposed based on plane geometry and linear programming, which can effectively eliminate the incorrect solutions. Both simulation and experimental results demonstrate that V-FTQ can achieve a positioning accuracy within 10 cm with only one captured beacon, outperforming baseline algorithms in terms of precision and robustness to noise. Note to Practitioners—Existing positioning algorithms typically require multiple beacons for accurate positioning, which may limit its coverage or require dense beacon deployment. The paper proposes a novel positioning algorithm that achieves 3D orientation and position estimation with only a single beacon and a single camera. We applied the proposed method to the mobile phone and developed an APP for actual positioning tests. Experimental results verify that the positioning scheme reduces both the number of beacons required and the cost of the positioning system, making it suitable for the automatic navigation of robots and other applications in industrial automation. Yang Yang 0057, Caili Guo |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Selectively Hard Negative Mining for Alleviating Gradient Vanishing in Image-Text MatchingabstractMost Image-Text Matching (ITM) models adopt Triplet loss with Hard Negative mining (T-HN) as the optimization objective. T-HN mines the hardest negative samples in each batch for training and achieves impressive performance. However, we observe that these ITM models have bad training behaviors in the early phases of training. Model training is difficult to converge, and matching performance is slow to improve. In this paper, we find that the cause of bad training behavior is that the model suffers from gradient vanishing. Optimizing an ITM model using only the hardest negative samples can easily lead to gradient vanishing. Through gradient analysis, we first derive the condition under which the gradient vanishes during training. We explain why the gradient tends to zero under certain conditions. To alleviate gradient vanishing, we propose a Triplet loss with Selectively Hard Negative mining (T-SelHN), which decides whether to mine the hardest negative samples according to the gradient vanishing condition. T-SelHN can be applied to ITM models in a plug-and-play manner to improve their training behaviors. To further ensure the back-propagation of gradients, we construct a Residual Visual Semantic Embedding model with T-SelHN, denoted RVSE++, which has a simple network structure and efficient training and inference speeds. Extensive experiments on two ITM benchmarks demonstrate the strength of RVSE++, achieving state-of-the-art performance. The code is available athttps://github.com/AAA-Zheng/RVSEPP. Zheng Li 0014, Caili Guo, Xin Wang 0203, Zerun Feng, Zhongtian Du |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | On the Impact of Uncertainty and Calibration on Likelihood-Ratio Membership Inference AttacksabstractIn amembership inference attack(MIA), an attacker exploits the overconfidence exhibited by typical machine learning models to determine whether a specific data point was used to train a target model. In this paper, we analyze the performance of thelikelihood ratio attack(LiRA) within an information-theoretical framework that allows the investigation of the impact of thealeatoric uncertaintyin the true data generation process, of theepistemic uncertaintycaused by a limited training data set, and of thecalibration levelof the target model. We compare three different settings, in which the attacker receives decreasingly informative feedback from the target model:confidence vector(CV) disclosure, in which the output probability vector is released;true label confidence(TLC) disclosure, in which only the probability assigned to the true label is made available by the model; anddecision set(DS) disclosure, in which an adaptive prediction set is produced as in conformal prediction. We derive bounds on the advantage of an MIA adversary with the aim of offering insights into the impact of uncertainty and calibration on the effectiveness of MIAs. Simulation results demonstrate that the derived analytical bounds predict well the effectiveness of MIAs. Meiyi Zhu, Caili Guo, Chunyan Feng, Osvaldo Simeone |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Performance Analysis of Joint NOMA and JT-CoMP Based on Stienen ModelabstractFor fifth-generation wireless networks to transition to sixth-generation wireless networks, the integration of coordinated multipoint (CoMP) and non-orthogonal multiple access (NOMA) techniques is expected to overcome new challenges and enhance performance compared to the CoMP or NOMA scheme. The joint-transmission CoMP (JT-CoMP) technique is a typical technical implementation of the CoMP scheme. In this study, we investigate a downlink network with a joint JT-CoMP-NOMA scheme. Based on the generalized Stienen model from stochastic geometry, we divide far and near NOMA user equipment (UE) and develop a theoretical framework to analyze the system performance. Expressions for the coverage probabilities and average achievable rates of two types of UEs (named CoMP and non-CoMP UEs) are derived. By comparing analytical results with Monte Carlo simulations, we show that the approximations in the analytical derivations are tight. The impact of certain network parameters, such as the power allocation coefficient, on the system performance is also studied. Notably, the developed transmission scheme is shown to outperform the NOMA-only and the JT-CoMP-only schemes. Yunpei Chen, Martin Haenggi, Qi Zhu 0003, Caili Guo, Yifei Yuan 0003, Zhuhua Hu, Xiaohui Li 0008 |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Large Scale Model Enabled Semantic Communications Based on Robust Knowledge DistillationabstractLarge scale artificial intelligence (AI) models possess excellent capabilities in semantic representation and understanding, making them particularly well-suited for semantic encoding and decoding. However, the substantial scale of these AI models imposes unacceptable computational resources and communication delays. To address this issue, we propose a semantic communication scheme based on robust knowledge distillation (RKD-SC) for large scale model enabled semantic communications. In the considered system, a transmitter extracts the features of the source image for robust transmission and accurate image classification at the receiver. To effectively utilize the superior capability of large scale model while make the cost affordable, we first transfer knowledge from a large scale model to a smaller scale model to serve as the semantic encoder. Then, to enhance the robustness of the system against channel noise, we propose a channel-aware autoencoder (CAA) based on the Transformer architecture. Experimental results show that the encoder of proposed RKD-SC system can achieve over 93.3% of the performance of a large scale model while compressing 96.67% number of parameters. Code: https://github.com/echojayne/RKD-SC. Kuiyuan Ding, Fangfang Liu 0003, Yang Yang 0057, Mingzhe Chen, Caili Guo |
GLOBECOM | 5 |
| 2024 | Digital Task-Oriented Communication with Hardware-Limited Task-Based QuantizationabstractTask-oriented communication exploits the task to improve communication efficiency. Most existing works on task-oriented communication transmit analog signals without quantization, which limits its application in digital communication systems. This paper studies digital task-oriented communication systems with hardware-limited scalar quantization (TOC-SQ) for computation-limited scenarios such as the Internet of Things, where the source data is encoded by a source encoder and then quantized by hardware-limited quantizers for digital transmission. Accordingly, the receiver contains a source decoder to decode the received signal. Our goal is to minimize the mean squared error (MSE) between the transmitted and received task-relevant signals to achieve optimal task performance under a certain bit budget. In particular, we first establish a theoretical analysis framework for TOC-SQ. Then, the closed-form expressions of the source encoder and decoder are derived for linear tasks. Finally, the lower bound of the MSE between the transmitted and received task-relevant information is analyzed to evaluate the task performance. Simulation results verify that the proposed TOC-SQ achieves 6.9 dB MSE gains in frequency-selective channels compared with analog TOC systems. Wuxia Hu, Yang Yang 0057, Yonina C. Eldar, Chunyan Feng, Caili Guo |
ICASSP | 5 |
| 2024 | Privacy-Aware Joint Source-Channel Coding For Image Transmission Based On Disentangled Information BottleneckabstractCurrent privacy-aware joint source-channel coding (JSCC) works aim at avoiding private information transmission by adversarially training the JSCC encoder and decoder under specific signal-to-noise ratios (SNRs) of eavesdroppers. However, these approaches incur additional computational and storage requirements as multiple neural networks must be trained for various eavesdroppers’ SNRs to determine the transmitted information. To overcome this challenge, we propose a novel privacy-aware JSCC for image transmission based on disentangled information bottleneck (DIB-PAJSCC). In particular, we derive a novel disentangled information bottleneck objective to disentangle private and public information. Given the separate information, the transmitter can transmit only public information to the receiver while minimizing reconstruction distortion. Since DIB-PAJSCC transmits only public information regardless of the eavesdroppers’ SNRs, it can eliminate additional training adapted to eavesdroppers’ SNRs. Experimental results show that DIB-PAJSCC can reduce the eavesdropping accuracy on private information by up to 20% compared to existing methods. Lunan Sun, Caili Guo, Mingzhe Chen, Yang Yang 0057 |
ICASSP | 2 |
| 2024 | Performance Optimization for Task-Oriented CommunicationsabstractTask-oriented communication is a new paradigm that aims at providing efficient connectivity for accomplishing intelligent tasks rather than the reception of every transmitted bit. This paper proposes a deep learning-based task-oriented communication architecture for end-to-end (E2E) semantics transmission, where extracted semantics is compressed by the proposed adaptable semantic compression (ASC) method. However, accommodating multiple users in a delay-intolerant system poses a challenge. Higher compression ratios conserve channel re-sources but cause semantic distortion, while lower ratios demand more resources and may lead to transmission failure due to delay constraints. To address this, we optimize both compression ratio and resource allocation to maximize task success probability. Specifically, due to the nonconvexity of the problem, we propose a compression ratio and resource allocation (CRRA) algorithm that separates the problem into two subproblems and solving them iteratively. Simulation results show that the proposed algorithm can obtain at least 14.3% success gains over baseline algorithms. Chuanhong Liu, Caili Guo, Yang Yang 0057 |
ICC | 2 |
| 2024 | Semantic Redundancy-Aware Multi-View Edge Inference Based on Rate-Distortion in IoT SystemsabstractAs the Internet of Things (IoT) continues to evolve, it is feasible to deploy multiple edge devices to capture objects from various angles, yielding comprehensive multi-view data which provides sufficient information for edge inference. This escalation in data volume, however, presents formidable challenges to communication systems with constrained communication resources. To tackle this issue, existing work designed a Multi-View Convolutional Neural Network (MVCNN) to extract low-dimensional features with high information, thus improving the spectral efficiency. However, in spite of its endeavor to identify and extract feature information within views, it ignores the redundancy present in multi-view data, especially the possibility that the views themselves are redundancy. Taking advantage of advancements in semantic communication, this paper proposes a semantic redundancy-aware multi-view joint edge inference network based on rate-distortion designed for IoT systems, named Compression Multi-View Convolutional Neural Network (CMVCNN), addressing the challenge that MVCNN faces in identifying redundant views and reducing channel bandwidth. Specifically, a compression network is meticulously designed to identify and remove the semantic redundancy in multi-view data, while a optimization algorithm is designed based on rate-distortion theory to guide semantic compression. Simulation results validate the effectiveness of the proposed CMVCNN scheme. Caili Guo, Meiyi Zhu |
WCNC | 2 |
| 2024 | Explainable Semantic Communication for Text TasksabstractTask-oriented semantic communication has gained increasing attention due to its ability to reduce the amount of transmitted data without sacrificing task performance. Although some prior efforts have been dedicated to developing semantic communications, the semantics in these works remains to be unexplainable. Challenges related to explainable semantic representation and knowledge-based semantic compression have yet to be explored. In this article, we propose a triplet-based explainable semantic communication (TESC) scheme for representing text semantics efficiently. Specifically, we develop a semantic extraction method to convert text into triplets while using syntactic dependency analysis to enhance semantic completeness. Then, we design a semantic filtering method to further compress the duplicate and task-irrelevant triplets based on prior knowledge. The filtered triplets are encoded and transmitted to the receiver for completing intelligent tasks. Furthermore, we apply the proposed TESC scheme to two emblematic text tasks: 1) sentiment analysis and 2) question answering, in which the semantic codec is meticulously customized for each task. Experimental results demonstrate that 1) the TESC scheme outperforms benchmarks in terms of Top-1 accuracy and transmission efficiency and 2) the TESC scheme enjoys about 150% performance gain compared to the traditional communication method. Chuanhong Liu, Caili Guo, Yang Yang 0057, Wanli Ni, Yanquan Zhou, Lei Li 0009, Tony Q. S. Quek |
IEEE Internet Things J. | 2 |
| 2024 | Rate-Adaptable Multitask-Oriented Semantic Communication: An Extended Rate-Distortion Theory-Based SchemeabstractSemantic communication, as a new paradigm for next-generation communication, aims to transmit semantic symbols for artificial intelligence (AI) tasks. Existing research typically requires extracting and transmitting specialized semantics for each AI task when multiple target AI tasks exist. Considering that each AI task may share common semantics, this article proposes a joint source–channel coding scheme for a multitask semantic communication (MTSC) system, which can extract the common semantics required by multiple AI tasks, and thus reduce the overall semantic transmission. To this end, we first formulate the MTSC problem as a rate–distortion problem that simultaneously considers the rate of extracted semantics and the distortion of multiple AI tasks. Then, we derive a new form of rate–distortion, called extended rate–distortion, which can guide the compact semantics extraction of multiple AI tasks simultaneously. Additionally, we derive a self-consistent equation for this extended rate–distortion form, theoretically proving the effectiveness of this approach. To ensure proper tradeoff between the rates and distortions of multiple AI tasks, we further propose a rate adjustment module that can dynamically adjust the rate according to channel conditions. We validate our experimental results on multiple data sets, which show that the proposed method can reduce transmission overhead by 40%–50% and achieve a 7.6% improvement in multitask performance. Fangfang Liu 0003, Zhengfen Sun, Yang Yang 0057, Caili Guo |
IEEE Internet Things J. | 4 |
| 2024 | Integrating listwise ranking into pairwise-based image-text retrieval
Zheng Li 0014, Caili Guo, Xin Wang 0203, Hao Zhang 0161 |
Knowl. Based Syst. | 2 |
| 2024 | OFDM-Based Digital Semantic Communication With Importance AwarenessabstractSemantic communication (SemCom) has received considerable attention for its ability to reduce data transmission size while maintaining task performance. However, existing works mainly focus on analog SemCom with simple channel models, which may limit its practical application. To reduce this gap, we propose an orthogonal frequency division multiplexing (OFDM)-based SemCom system that is compatible with existing digital communication infrastructures. In the considered system, the extracted semantics is quantized by scalar quantizers, transformed into OFDM signal, and then transmitted over the frequency-selective channel. Moreover, we propose a semantic importance measurement method to build the relationship between target task and semantic features. Based on semantic importance, we formulate a sub-carrier and bit allocation problem to maximize communication performance. However, the optimization objective function cannot be accurately characterized using a mathematical expression due to the neural network-based semantic codec. Given the complex nature of the problem, we first propose a low-complexity sub-carrier allocation method that assigns sub-carriers with better channel conditions to more critical semantics. Then, we propose a deep reinforcement learning-based bit allocation algorithm with dynamic action space. Simulation results demonstrate that the proposed system achieves 9.7% and 28.7% performance gains compared to analog SemCom and conventional bit-based communication systems, respectively. Chuanhong Liu, Caili Guo, Yang Yang 0057, Wanli Ni, Tony Q. S. Quek |
IEEE Trans. Commun. | 2 |
| 2024 | Disentangled Information Bottleneck Guided Privacy-Protective Joint Source and Channel Coding for Image TransmissionabstractJoint source and channel coding (JSCC) has attracted increasing attention in semantic communications. However, JSCC is vulnerable to privacy issues due to the high relevance between the source image and channel input. In this paper, we propose a disentangled information bottleneck guided privacy-protective JSCC (DPJSCC) for image transmission, which aims at protecting private information and achieving superior image transmission performance. In particular, we propose a disentangled information bottleneck objective to compress the private information in public subcodewords and improve the reconstruction quality simultaneously. To optimize JSCC neural networks using the proposed objective, we derive a differentiable estimation based on variational approximation and the density-ratio trick. Additionally, we design a password-based privacy-protective algorithm that encrypts the private subcodewords, achieving joint optimization with JSCC neural networks. The proposed algorithm involves an encryptor for encrypting private information and a decryptor for recovering it at the legitimate receiver. A loss function is derived based on the maximum entropy principle for jointly training the encryptor, decryptor, and JSCC decoder to maximize eavesdropping uncertainty and improve reconstruction quality. Experimental results show that DPJSCC reduces eavesdropping accuracy on private information by up to 18% and decreases inference time by 10%. Lunan Sun, Yang Yang 0057, Mingzhe Chen, Caili Guo |
IEEE Trans. Commun. | 4 |
| 2024 | Visible Light Positioning With Visual Odometry: A Single Luminaire Based Positioning AlgorithmabstractVisible light positioning (VLP) is an accurate and low-cost positioning technique. However, existing VLP algorithms require multiple luminaires or multiple sensors to achieve the desired positioning accuracy, which may not be satisfied in practice. To circumvent this challenge, a novel visual odometry (VO) assisted VLP algorithm (VO-VLP) is proposed, which can achieve accurate positioning using only a single luminaire at the transmitter and a single camera at the receiver. In the considered model, the luminaires are equipped on the ceiling and consistently broadcast coordinate information of the luminaires by visible light communication (VLC). A user equipped with a camera captures photos of the ceiling so as to locate its position via VO-VLP. In particular, VO-VLP first uses the single luminaire’s circle feature and VLC information to obtain the pose and location of the user. However, there are dual solutions due to the limited received information in the single-luminaire scenario and the symmetry of the circular luminaire. Then, we propose a duality elimination method to eliminate the wrong one by introducing VO to exploit the visual features on the ceiling in two consecutive images, which are captured when the user moves. To verify the feasibility of our designed VO-VLP, a prototype is implemented. A cooperative multi-information image processing method is proposed for the prototype to ensure that the VLC information and the visual information of the luminaire and the ceiling can be simultaneously received for real-time positioning. Simulations and experiments are conducted to prove that VO-VLP can achieve accurate positioning with only a single luminaire and a camera without any extra sensors, such as an inertial measurement unit. In particular, simulation results show that the proposed indoor positioning algorithm can achieve a 97% positioning accuracy of around 10 cm, and experimental results show that the average positioning accuracy is less than 10 cm. Yang Yang 0057, Mingzhe Chen, Caili Guo, Jiangyi Hao, Shuguang Cui |
IEEE Trans. Commun. | 4 |
| 2024 | Quantify Wheat Canopy Leaf Angle Distribution Using Terrestrial Laser Scanning DataabstractLeaf angle distribution (LAD) is an important structural attribute of crop canopies as it influences photosynthesis and radiation transport. Terrestrial laser scanning (TLS) has shown promise as a tool for quantifying LAD. However, the current TLS-derived crop canopy LAD estimation lacks automatic segmentation for the special curved leaves of the crop. Furthermore, mutual shading between plants results in an uneven distribution of leaf point density in the crop canopy. We developed a novel voxel segmentation normal vector (VSNV) method for automatically segmenting and spatially normalizing curved leaves to address those concerns. In this methodology, the wheat canopy is divided into voxels, and LAD is derived by averaging the angles from the planes associated with each point within every voxel. The ray-tracing 3D radiative transfer model (LESS) was used to validate the effectiveness of the VSNV method, which produced better LAD results than the normal vector (NV) method. In addition, the mean leaf tilt angle (MTA) of wheat estimated by TLS using the VSNV approach correlated well with the measured value from LAI-2200C, especially at the booting stage (R2= 0.76 and RMSE = 1.40°). The result shows that the improved VSNV method can trace LAD characteristics among cultivars, nitrogen levels, growth stages, and canopy heights. Quantifying the variability of LAD could provide strong technical support for high-throughput phenotyping. Yangyang Gu, Jinxin Tang, Binbin Guo, Timothy A. Warner, Caili Guo, Hengbiao Zheng, Fumiki Hosoi, Tao Cheng 0003, Yan Zhu 0005, Weixing Cao, Xia Yao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Learning From Noisy Correspondence With Tri-Partition for Cross-Modal MatchingabstractDue to high labeling cost, it is inevitable to introduce a certain proportion of noisy correspondence into visual-text datasets, resulting in poor model robustness for cross-modal matching. Although recent methods divide the datasets into clean and noisy pair subsets to yield promising achievements, they still suffer from deep neural networks over-fitting on noisy correspondence. In particular, the similar positive pairs with partially relevant semantic correspondence are easily partitioned into noisy pair subset by mistake without carefully selection, which brings harmful impact on robust learning. Meanwhile, the similar negative pairs with partially relevant semantic correspondence lead to ambiguous distance relations in common space learning, which also damages the stability of performance. To solve the coarse-grained dataset division problem, we propose Correspondence Tri-Partition Rectifier (CTPR) to partition the training set into clean, hard, and noisy pair subsets based on the memorization effect of neural networks and prediction inconsistency. Then, we refine the correspondence labels for each subset to indicate the real semantic correspondence between visual-text pairs. The differences between rectified labels of anchors and hard negatives are recast as the adaptive margin in the improved triplet loss for robust training in a co-teaching manner. To verify the effectiveness and robustness of our method, we conduct experiments by implementing image-text and video-text matching as two showcases. Extensive experiments on Flickr30 K, MS-COCO, MSR-VTT, and LSMDC datasets verify that our method successfully partitions the visual-text pairs according to their semantic correspondence and improves performance under noisy data training. Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Lin Hu 0008 |
IEEE Trans. Multim. | 3 |
| 2024 | Integrating Language Guidance Into Image-Text Matching for Correcting False NegativesabstractImage-Text Matching (ITM) aims to establish the correspondence between images and sentences. ITM is fundamental to various vision and language understanding tasks. However, there are limitations in the way existing ITM benchmarks are constructed. The ITM benchmark collects pairs of images and sentences during construction. Therefore, only samples that are paired at collection are annotated as positive. All other samples are annotated as negative. Many correlations are missed in these samples that are annotated as negative. For example, a sentence matches only one image at the time of collection. Only this image is annotated as positive for the sentence. All other images are annotated as negative. However, these negative images may contain images that correspond to the sentences. These mislabeled samples are calledfalse negatives. Existing ITM models are optimized based on annotations containing mislabels, which can introduce noise during training. In this paper, we propose an ITM framework integrating Language Guidance (LG) for correcting false negatives. A language pre-training model is introduced into the ITM framework to identify false negatives. To correct false negatives, we propose language guidance loss, which adaptively corrects the locations of false negatives in the visual-semantic embedding space. Extensive experiments on two ITM benchmarks show that our method can improve the performance of existing ITM models. To verify the performance of correcting false negatives, we conduct further experiments on ECCV Caption. ECCV Caption is a verified dataset where false negatives in annotations have been corrected. The experimental results show that our method can recall more relevant false negatives. The code is available athttps://github.com/AAA-Zheng/LG_ITM. Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Zhongtian Du |
IEEE Trans. Multim. | 2 |
| 2023 | Visible Light Positioning Based on a Single Luminaire: A Novel Visual Odometry Assisted AlgorithmabstractVisible light positioning (VLP) is a promising positioning technique, which, however, typically requires multiple luminaires to achieve accurate positioning. This paper proposes a novel visual odometry (VO) assisted visible light positioning algorithm (VO-VLP) in achieving positioning with only a single luminaire. In the considered model, a user equipped with a camera jointly uses geometric features in the captured images and coordinates information obtained via visible light communication (VLC) for positioning. The proposed VLP algorithm does not rely on any extra inertial measurement unit and relaxes the tilted angle limitation at the user. In particular, VO-VLP first uses the circle feature of a luminaire to obtain dual normal vectors of the luminaire. Then, the basic principle of VO is used to eliminate the wrong normal vector by exploiting the geometric features in two consecutive images captured when the user moves. Finally, the pose and location of the user are obtained by using an artificially marked point on the luminaire's contour. VO-VLP can achieve accurate positioning with only a single luminaire and a camera. Simulation results show that the proposed indoor positioning algorithm can achieve a 97th-percentile positioning accuracy of around 10 cm. Yang Yang 0057, Mingzhe Chen, Caili Guo, Yipeng Bai |
ICC | 4 |
| 2023 | Deep Joint Source-Channel Coding Based on Semantics of Pixels for Wireless Image TransmissionabstractCurrent image coding methods for semantic communication typically concentrate on intelligent tasks or image reconstruction separately, and seldom consider both aspects simultaneously. To balance these two aspects during wireless image transmission, we propose a joint source-channel coding method based on the semantics of pixels (SP), which can retain both pixel information for reconstruction and semantic information for intelligent tasks. Specifically, we first design a gradient-based mechanism to quantify the semantic importance of downstream intelligent tasks on pixels. Then, we design the SP-based loss function to train the deep joint source-channel coding network. Experiment results demonstrate that the proposed method maintains reconstruction performance and improves the task performance by 1.61% and 4.06%, respectively, compared to the state-of-the-art deep joint source-channel coding method and traditional separate source-channel coding method at the same transmission rate and signal-to-noise ratio. Caili Guo, Yang Yang 0057, Chuanhong Liu |
PIMRC | 2 |
| 2023 | Feature Decomposition and Attribute Augmentation for Attribute-based Person SearchabstractThe attribute-based person search task aims to find matching pedestrian images by text attributes, which is relevant in scenarios where no query image is given. However, the existing methods exhibit inferior performance due to their inadequate representation of local fine-grained features, which hinders their ability to effectively model intra-class variations. In addition, there is a zero-shot retrieval problem due to the large number of unseen categories in the test set, resulting in suboptimal generalization performance. In this paper, we propose a novel Feature Decomposition and Attribute Augmentation (FDAA) framework to solve the above problems. Firstly, by decomposing the global features of the image from different directions, the local features are extracted at multiple granularities, thus effectively improving the discrimination ability of the model. Secondly, an attribute augmentation strategy is proposed that can effectively expand the combination of attributes during training and improve the generalization ability of the model. Extensive experiments on the PETA, Market-1501 Attribute, and PA100K datasets demonstrate the effectiveness of our proposed method, outperforming state-of-the-art methods. Xin Wang 0203, Fangfang Liu 0003, Caili Guo, Zheng Li 0014, Hao Zhang 0161 |
VCIP | 3 |
| 2023 | Feature Refinement with Masked Cascaded Network for Temporal Action LocalizationabstractDespite the great progress in temporal action localization (TAL), most existing methods directly use video encoders trained on trimmed Kinetics400 dataset to obtain clip-level visual features, ignoring the cross-dataset bias between Kinetics400 and TAL benchmarks. Such a dataset bias leads to poor visual representation, potentially hindering performance in both temporal detection and action recognition for TAL. In this paper, we propose a novel TAL method, termed feature refinement with masked cascaded network (FR-MCN), to tackle the above problem. Specifically, FR-MCN presents a new feature refinement strategy by developing clip-level feature classification task for both action and background clips to improve temporal sensitivity and enhance action semantics of visual features. Moreover, FR-MCN employs a masked cascaded paradigm for refinement to learn semantic disparities between action and background clips near boundary, enabling the starting and ending instants to be detected accurately for TAL. Extensive experimental results on THUMOS14 and ActivityNetv1.3 demonstrate that our FR-MCN, can significantly improve the action localization performance. Chunyang Feng, Hao Zhang 0161, Caili Guo, Zheng Li 0014 |
VCIP | 4 |
| 2023 | Boundary-Aware Proposal Generation Method for Temporal Action LocalizationabstractThe goal of Temporal Action Localization (TAL) is to find the categories and temporal boundaries of actions in an untrimmed video. Most TAL methods rely heavily on action recognition models that are sensitive to action labels rather than temporal boundaries. More importantly, few works consider the background frames that are similar to action frames in pixels but dissimilar in semantics, which also leads to inaccurate temporal boundaries. To address the challenge above, we propose a Boundary-Aware Proposal Generation (BAPG) method with contrastive learning. Specifically, we define the above background frames as hard negative samples. Contrastive learning with hard negative mining is introduced to improve the discrimination of BAPG. BAPG is independent of the existing TAL network architecture, so it can be applied plug-and-play to mainstream TAL models without training. Extensive experimental results on THUMOS14 and ActivityNet-1.3 demonstrate that BAPG can significantly improve the performance of TAL. Hao Zhang 0161, Chunyan Feng, Caili Guo, Zheng Li 0014, Xin Wang 0203 |
VCIP | 4 |
| 2023 | Dynamic Coded Caching in Cellular Networks with User Mobility: A Reinforcement Learning MethodabstractCoded caching manages to release cellular network traffic by increasing transmission rate via satisfying multiple user requests simultaneously. Specific contents stored in the private cache memory are used as side information to decode individual requests from the coded broadcasting messages. Considering local content popularity could improve caching performance dramatically. However, in the mobility scenario, local content popularity varies with user movements. Even worse, contents in the cache memory might become outdated when the user location changes. In this paper, we propose a dynamic coded caching scheme that reduces the loss of coded caching gain due to the user movement and local content popularity dynamic changing. We quantify the relationship between user preference, local popularity, and user mobility. We formulate a metric to measure the performance of the proposed coded caching scheme and propose a reinforcement learning problem to obtain the cache replacement strategy in the mobility scenario. Numerical results verify that our obtained replacement policy significantly outperforms the popularity-based, least-frequently-used, and multilayer replacement policy in terms of traffic offloading. Guangyu Zhu 0010, Caili Guo, Tiankui Zhang |
VTC Fall | 2 |
| 2023 | Task-Oriented Semantic Communication Based on Semantic TripletsabstractTask-oriented semantic communication has received growing interests, which can significantly reduce the amount of transmitted data without affecting task performance. In this paper, a novel semantic communication system based on semantic triplets (SCST) is proposed, in which the semantics is represented via the explainable semantic triplets. Specifically, we propose a semantic extraction method to convert the transmitted texts into semantic triplets, which can be further compressed via the designed semantic filtering method. The semantic triplets then will be encoded and transmitted via the wireless channel to complete intelligent tasks at the receiver. Moreover, we then apply the SCST to sentiment analysis task and question-answering task to verify the effectiveness, where the semantic encoder and decoder are designed respectively considering the final task. The experiment results show that the proposed SCST can obtain at least 43.5% and 52% accuracy gains, compared to the baselines using traditional communication method. Chuanhong Liu, Caili Guo, Dingxin Hu |
WCNC | 2 |
| 2023 | Joint 3-D Position Deployment and Traffic Offloading for Caching and Computing-Enabled UAV Under Asymmetric InformationabstractTo offload the rapidly increasing video traffic in the Internet of Things (IoT), the caching and computing-enabled unmanned aerial vehicles (UAVs) have become a paradigm in alleviating the pressure of wireless backhauls of the mobile network operator (MNO) and improving the Quality of Service (QoS) of subscribers of content providers (CPs), especially, in the situations where fixedly deployed small-cell base stations (SBSs) are not applicable. However, the resources in the UAV are costly and limited, while rational CPs are reluctant to reveal their willingness of leasing. Taking such asymmetric information into consideration, the problem that how the MNO leases the caching and computing resources to different CPs is formulated into a contract design problem, in which the UAV’s 3-D position deployment is also involved. To find the optimal solutions, the proposed problem is decoupled into two subproblems. In the position deployment subproblem, the 3-D position of the UAV is analyzed, and the hovering height and coverage radius are jointly optimized to maximize the offloading volume during the life of the UAV. With the derived optimal position of the UAV, the individual rationality (IR) and incentive compatibility (IC) constraints in the contract design subproblem are simplified, and a low complexity algorithm based on the alternating direction method of multipliers (ADMMs) is proposed to find the optimal contract. Finally, the effectiveness of the proposed scheme is verified by simulation, and a comparative analysis is carried out in terms of save latency, saved bandwidth and utility. Biling Zhang, Jung-Lang Yu, Caili Guo, Zhu Han 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Adaptive Information Bottleneck Guided Joint Source and Channel Coding for Image TransmissionabstractJoint source and channel coding (JSCC) for image transmission has attracted increasing attention due to its robustness and high efficiency. However, the existing deep JSCC research mainly focuses on minimizing the distortion between the transmitted and received information under a fixed number of available channels. Therefore, the transmitted rate may be far more than its required minimum value. In this paper, an adaptive information bottleneck (IB) guided joint source and channel coding (AIB-JSCC) method is proposed for image transmission. The goal of AIB-JSCC is to reduce the transmission rate while improving the image reconstruction quality. In particular, a new IB objective for image transmission is proposed so as to minimize the distortion and the transmission rate. A mathematically tractable lower bound on the proposed objective is derived, and then, adopted as the loss function of AIB-JSCC. To trade off compression and reconstruction quality, an adaptive algorithm is proposed to adjust the hyperparameter of the proposed loss function dynamically according to the distortion during the training. Experimental results show that AIB-JSCC can significantly reduce the required amount of transmitted data and improve the reconstruction quality and downstream task accuracy. Lunan Sun, Yang Yang 0057, Mingzhe Chen, Caili Guo, Walid Saad 0001, H. Vincent Poor |
IEEE J. Sel. Areas Commun. | 4 |
| 2023 | Information Bottleneck-Inspired Type Based Multiple Access for Remote Estimation in IoT SystemsabstractType-based multiple access (TBMA) is a semantics-aware multiple access protocol for remote inference. In TBMA, codewords are reused across transmitting sensors, with each codeword being assigned to a different observation value. Existing TBMA protocols are based on fixed shared codebooks and on conventional maximum-likelihood or Bayesian decoders, which require knowledge of the distributions of observations and channels. In this letter, we propose a novel design principle for TBMA based on the information bottleneck (IB). In the proposed IB-TBMA protocol, the shared codebook is jointly optimized with a decoder based on artificial neural networks (ANNs), so as to adapt to source, observations, and channel statistics based on data only. We also introduce the Compressed IB-TBMA (CIB-TBMA) protocol, which improves IB-TBMA by enabling a reduction in the number of codewords via an IB-inspired clustering phase. Numerical results demonstrate the importance of a joint design of codebook and neural decoder, and validate the benefits of codebook compression. Meiyi Zhu, Chunyan Feng, Caili Guo, Nan Jiang 0004, Osvaldo Simeone |
IEEE Signal Process. Lett. | 3 |
| 2023 | Joint SIC-Based Precoding and Sub-Connected Architecture Design for MIMO VLC SystemsabstractDensely deployed light emitting diodes (LEDs) typically lead to high spatial correlation in multiple-input multiple-output (MIMO) visible light communication (VLC) systems. Although precoding can effectively alleviate the spatial correlation issue, most existing precoding algorithms require a dedicated baseband chain for each LED, leading to high energy consumption and hardware complexity when a large number of LEDs are used. In this paper, a successive interference cancellation (SIC)-based precoding scheme with sub-connected architecture (SIC-SA) is proposed. In the considered model, each baseband chain is connected to an LED sub-array containing multiple LEDs to reduce the complexity. Since, in this case, SIC-based precoding can only determine the signal of each baseband chain for an LED sub-array, while its target is to mitigate the spatial correlation between individual LEDs, the electrical/optical power of each LED must also be jointly optimized to accomplish the target. This joint SIC-based precoding, power allocation, and direct current offset design problem is formulated as an achievable sum rate maximization problem under dimming control and electrical power constraints. A two-step iterative algorithm is proposed to solve this problem. In the first step, the SIC-based precoding is designed to alleviate the multi-user interference. In the second step, the power allocation of LEDs is optimized by matrix decomposition and convex optimization, and a closed-form solution of DC offset is derived. Furthermore, considering the dynamic scenarios, a SIC-based precoding scheme with dynamic sub-connected architecture (SIC-DSA) is proposed, in which a switch network is used to adaptively adjust the LED sub-array structure based on the channel state information. Simulation results show that the proposed SIC-SA and SIC-DSA respectively achieve 0.1340 bps/Hz/W and 0.1305 bps/Hz/W energy efficiency gains over the zero-forcing precoding scheme with SA, when the signal-to-noise ratio is 30 dB. Yang Yang 0057, Zhaohui Yang 0001, Chunyan Feng, Julian Cheng 0001, Caili Guo |
IEEE Trans. Commun. | 6 |
| 2023 | Temporal Multimodal Graph Transformer With Global-Local Alignment for Video-Text RetrievalabstractVideo-text retrieval is a crucial task that has been a powerful application for multi-media data analysis and attracted tremendous interest in the research area. The core steps are feature representations and alignment to overcome the heterogeneous gap between videos and texts. Existing methods not only take advantage of multi-modal information in videos but also explore local alignment to enhance retrieval accuracy. Although performing well, these methods seem deficient at three perspectives: a) The semantic correlations between different modal features are not considered, which introduces irrelevant noise in feature representations. b) The cross-modal relations and temporal associations are ambiguously learned by a single self-attention manipulation. c) The training signal to optimize the semantic topic assignment for local alignment is missing. In this paper, we proposed a novel Temporal Multi-modal Graph Transformer with Global-Local Alignment (TMMGT-GLA) for video-text retrieval. We model the input video as a sequence of semantic correlation graphs to exploit the structural information between multi-modal features. Graph and temporal self-attention layers are leveraged on the semantic correlation graphs to effectively learn cross-modal relations and temporal associations respectively. For local alignment, the encoded video and text features are assigned to a set of shared semantic topics, and the distances between residuals from the same ones are minimized. To optimize the assignments, a minimum entropy-based regularization term is proposed for training the overall framework. Experimental results are carried out on the MSR-VTT, LSMDC, and ActivityNet Captions datasets. Our method outperforms previous approaches by a large margin and achieves state-of-the-art performance. Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Physical Layer Authentication Based on Channel Polarization Response in Dual-Polarized Antenna Communication SystemsabstractThis study presents a novel approach for physical layer authentication based on channel polarization response (CPR). CPR is sensitive to variation in the physical properties of scatterers, and the CPR difference between various channels is higher than the channel frequency response (CFR) under rich scattering scenarios. Additionally, the estimation of CPR is continuous, the authentication interval can be adjusted according to the channel coherence time, then the proposed scheme can be applied to any rich scattering scenarios, including highly dynamic scenarios. Since the received polarization state is fixed during the channel coherence time, we can coherently stack the received polarization state to improve the signal to noise ratio (SNR) and the estimation accuracy of CPR, thereby achieving high authentication accuracy under ultra-low SNR. Moreover, since the transmitted polarization state of various transmitters is different, because of their unique hardware deficiencies, and since the CPR is dependent on the transmitted polarization state, the CPR of other transmitters is different, allowing the resolution of co-located attacks. We theoretically drive the false alarm probability, detection probability, optimal discriminant threshold, computational complexity, optimal stacking numbers, and optimal CPR points for authentication. Furthermore, extensive simulations and experiments are performed to verify the validity and effectiveness of the proposed scheme. Yuemei Wu, Dong Wei 0002, Caili Guo, Weiqing Huang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Positioning Using Visible Light Communications: A Perspective Arcs ApproachabstractVisible light positioning (VLP) is an accurate indoor positioning technology that uses luminaires as transmitters. In particular, circular luminaires are a common source type for VLP, which are typically treated only as point sources for positioning, while ignoring their geometry characteristics. In this paper, the arc feature of the circular luminaire and the coordinate information obtained via visible light communication (VLC) are jointly used for positioning, and a novel perspective arcs approach is proposed for VLC-enabled indoor positioning. The proposed approach does not rely on any inertial measurement unit and has no tilted angle limitation at the user. First, a VLC assisted perspective circle and arc algorithm (V-PCA) is proposed for a scenario in which a complete luminaire and an incomplete one can be captured by the user. Based on plane and solid geometry theory, the relationship between the luminaire and the user is exploited to estimate the orientation and the coordinate of the luminaire in the camera coordinate system. Then, the pose and location of the user in the world coordinate system are obtained by single-view geometry theory. Considering the cases in which parts of VLC links are blocked, an anti-occlusion VLC assisted perspective arcs algorithm (OA-V-PA) is proposed. In OA-V-PA, an approximation method is developed to estimate the projection of the luminaire’s center on the image and, then, to calculate the pose and location of the user. Simulation results show that the proposed indoor positioning algorithm can achieve a 90th percentile positioning accuracy of around 10 cm. Moreover, an experimental prototype is implemented to verify the feasibility. In the established prototype, a fused image processing method is proposed to simultaneously obtain the VLC information and the geometric information. Experimental results in the established prototype show that the average positioning accuracy is less than 5 cm for different tilted angles of the user. Caili Guo, Rongzhen Bao, Mingzhe Chen, Walid Saad 0001, Yang Yang 0057 |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | A Joint Transmit and Receive Design for Dimmable High Speed MC MU VLC SystemsabstractIn multi-cell multi-user multiple-input multiple-output (MC MU-MIMO) visible light communication (VLC) systems, each user is equipped with multiple closely-placed photodiodes (PDs) with similar channel gains, leading to severe inter-cell interference and intra-cell interference. To address this problem, this paper proposes a hybrid dimming (HD) scheme with MIMO VLC transceiver design, which jointly optimizes transmit and receive antenna selection (TRAS), cell clustering and precoding (TRASP-HD) for MC MU-MIMO VLC systems. In this scheme, a sum-rate maximization problem under the dimming level and illumination uniformity is formulated and solved by being divided into two sub-problems. In particular, The first sub-problem is on TRAS and cell formation based on the criterion of sum-rate maximization under the illumination uniformity constraint. With the same goal, the second sub-problem is on optimizing the precoding matrices of each cell. Finally, these two sub-problems are iteratively solved to obtain a convergent solution. Simulation results verify that in a typical indoor scenario, the mean bandwidth efficiency of TRASP-HD scheme is 2.36 bit/s/Hz higher than the conventional MC MU-MIMO system. Yang Yang 0057, Hailun Xia, Caili Guo, Mahdi Chehimi, Walid Saad 0001 |
GLOBECOM | 4 |
| 2022 | Visible Light Communication Assisted Perspective Circle and Arc Algorithm for Indoor PositioningabstractThis paper proposes a visible light communication (VLC) assisted perspective circle and arc algorithm (V-PCA) for indoor positioning. In the considered model, the coordinate information obtained via VLC and the geometric information of circle feature are jointly used for indoor positioning. In particular, V-PCA first designs a space-time coding model based on VLC, which can help determine the center and a mark point of the circular luminaire on the image plane. It also transmits the world coordinate information of the luminaire to the camera through the activated LEDs. Then, based on the plane and solid geometry, the geometric information of circular feature is exploited to estimate the orientation and the coordinate of the luminaire in the camera coordinate system. Finally, the pose and the location of the camera in the world coordinate system are obtained by the single-view geometry. V-PCA can achieve accurate positioning without tilting angle limitation and inertial measurement unit for the receiver. Simulation results show that the proposed indoor positioning algorithm can achieve a 95th percentile positioning accuracy of around 10 cm. Yang Yang 0057, Caili Guo, Chunyan Feng |
ICC | 3 |
| 2022 | Spatial-Temporal Alignment via Optimal Transport for Video-Text RetrievalabstractVideo-text retrieval is one of the most popular branches in the cross-modal research area facing the exponential growth of multimedia services. Present methods typically mea-sure cross-modal similarities only by the cosine metric in a common space. However, this function ignores the in-herent discrepancy between videos and texts for reflecting spatial-temporal contents and fails to align them under the weakly supervised setting. To address this issue, we propose a novel Spatial-Temporal Optimal Transport (STOT) frame-work building upon advances in optimal transport. The spatial and temporal characteristics of videos and texts are carefully considered in STOT and associated by calculating a spatial-temporal alignment distance, which can be implemented as a regularizer for existing models to improve retrieval accu-racy. Extensive experiments are conducted on two video-text datasets with two models as base architectures to demonstrate the outperformance of our framework due to the powerful spatial-temporal alignment capability. Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Yufeng Zhang 0008 |
ICME | 3 |
| 2022 | Jointly Learning Agent and Lane Information for Multimodal Trajectory PredictionabstractPredicting the plausible future trajectories of nearby agents is a core challenge for the safety of Autonomous Vehicles and it mainly depends on effectively unifying agent dynamics and scene context. Recent approaches have made great progress in characterizing the two part separately. However, the interdependence between them remains to further research. In this paper, we use lane as scene data and propose a novel network that Jointly learns Agent and Lane information for Multimodal Trajectory Prediction (JAL-MTP). JAL-MTP use a Social to Lane (S2L) module to jointly represent the social agents and lane proposals as instance-level lanes, a Recurrent Lane Attention (RLA) mechanism for utilizing the prior information of instance-level lanes to perform accurate prediction, two selectors to identify the typical and reasonable trajectories. The experiments conducted on the public Argoverse dataset demonstrate that JAL-MTP could predict accurate and reasonable multimodal trajectories from both quantitative and qualitative perspectives. Jie Wang 0124, Caili Guo, Minan Guo |
ICPR | 2 |
| 2022 | Semantic-driven Computation Offloading and Resource Allocation for UAV-assisted Monitoring System in Vehicular NetworksabstractIn the vehicular networks monitoring scenario where the unmanned aerial vehicle (UAV) assists the vehicle data collection and the edge-cloud server cooperates to complete the visual intelligent task (e.g. object detection), a large amount of video data needs to be transmitted. For task-oriented communication systems, traditional resource allocation schemes mainly focus on the network performance or user experience of video transmission, without considering the impact of semantic content on task performance. In this paper, we propose a semantic-driven computation offloading and resource allocation scheme, namely semantic-driven CO&RA. In the offloading decision, we propose a convolutional neural network (CNN) segmentation scheme, which makes full use of computing resources and better completes intelligent tasks. At the same time, facing the complex resource allocation problem, we design a multi-agent deep Q-network (DQN) algorithm. Last, the experimental results show that our proposed scheme has more advantages in the optimization of multiple objectives of latency, energy consumption and task performance. Caili Guo |
IECON | 3 |
| 2022 | Multi-View Visual Semantic EmbeddingabstractVisual Semantic Embedding (VSE) is a dominant method for cross-modal vision-language retrieval. Its purpose is to learn an embedding space so that visual data can be embedded in a position close to the corresponding text description. However, there are large intra-class variations in the vision-language data. For example, multiple texts describing the same image may be described from different views, and the descriptions of different views are often dissimilar. The mainstream VSE method embeds samples from the same class in similar positions, which will suppress intra-class variations and lead to inferior generalization performance. This paper proposes a Multi-View Visual Semantic Embedding (MV-VSE) framework, which learns multiple embeddings for one visual data and explicitly models intra-class variations. To optimize MV-VSE, a multi-view upper bound loss is proposed, and the multi-view embeddings are jointly optimized while retaining intra-class variations. MV-VSE is plug-and-play and can be applied to various VSE models and loss functions without excessively increasing model complexity. Experimental results on the Flickr30K and MS-COCO datasets demonstrate the superior performance of our framework. Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Xijun Xue |
IJCAI | 2 |
| 2022 | Deep Joint Source-Channel Coding for Wireless Image Transmission with Semantic ImportanceabstractThe sixth-generation mobile communication system proposes the vision of smart interconnection of everything, which requires accomplishing communication tasks while ensuring the performance of intelligent tasks. A joint source-channel coding method based on semantic importance is proposed, which aims at preserving semantic information during wireless image transmission and thereby boosting the performance of intelligent tasks for images at the receiver. Specifically, we first propose semantic importance weight calculation method, which is based on the gradient of intelligent task’s perception results with respect to the features. Then, we design the semantic loss function in the way of using semantic weights to weight the features. Finally, we train the deep joint source-channel coding network using the semantic loss function. Experiment results demonstrate that the proposed method achieves up to 57.7% and 9.1% improvement in terms of intelligent task’s performance compared with the source-channel separation coding method and the deep source-channel joint coding method without considering semantics at the same compression rate and signal-to-noise ratio, respectively. Caili Guo, Yang Yang 0057, Chuanhong Liu |
VTC Fall | 2 |
| 2022 | Blockchain-Inspired Secure Computation Offloading in a Vehicular Cloud NetworkabstractWith the emergence of computation-intensive vehicular applications, computation offloading based on mobile-edge computing (MEC) has become a promising paradigm in resource-constrained vehicular cloud networks (VCNs). However, when doing computation offloading in a VCN, malicious service providers can cause serious security concerns on the content offloading. To address that in this article, a blockchain-based secure computation offloading scheduling scheme is proposed. It embraces the blockchain-based trust management paradigm and smart contract-enabled deep reinforcement learning (DRL) algorithm. As for the trust management, the long-term reputation and short-term trust variability are jointly considered. Specifically, a novel three-valued subjective logic (3VSL) scheme is adopted to obtain a more comprehensive reputation, and the statistics of behavioral transitions can provide a short-term trust variability to timely capture the malicious behaviors. In addition, to securely update, validate, and store the trust information, we propose a hierarchical blockchain framework that comprises vehicular blockchain, roadside unit (RSU) blockchain, and cloud blockchain. Furthermore, a smart contract-enabled DRL algorithm is proposed to implement the secure and intelligent computation offloading scheduling in a VCN. Simulations are conducted to verify the effectiveness of the proposed scheme. Shilin Xu 0002, Caili Guo, Rose Qingyang Hu, Yi Qian 0001 |
IEEE Internet Things J. | 2 |
| 2022 | Computer Vision-Based Localization With Visible Light CommunicationsabstractVisible light positioning and computer vision-based localization have the potential to be cost-effective technologies for accurate indoor localization. However, the feasibility of existing methods in this domain is limited. In this paper, a novel visible light communication (VLC)-assisted perspective-four-line algorithm (V-P4L) is proposed for practical indoor localization. The basic idea of V-P4L is to jointly use VLC and computer vision techniques to achieve high localization accuracy regardless of LED height differences. In particular, the space-domain information is first exploited to estimate the orientation and coordinate information of a single rectangular LED luminaire in the camera coordinate system based on plane geometry theory and solid geometry theory. Then, by using time-domain information transmitted by VLC and the estimated luminaire information, the proposed V-P4L can estimate the position and pose of the camera using single-view geometry theory and the linear least square (LLS) method. To further mitigate the effect of height differences among LEDs on localization accuracy, a correction algorithm based on the LLS method and a simple optimization method is proposed. Due to the combination of time- and space-domain information, V-P4L can achieve accurate localization using a single luminaire without limitation on the correspondences between the features and their projections in conventional perspective-n-line (PnL) algorithms. Simulation results show that the position error caused by the proposed V-P4L algorithm is always less than 15 cm and the orientation error is always less than 4° using popular indoor luminaires. Experimental results with real hardware show that the average position error is less than 3 cm under both similar and different heights for the LEDs. Lin Bai 0004, Yang Yang 0057, Mingzhe Chen, Chunyan Feng, Caili Guo, Walid Saad 0001, Shuguang Cui |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Content Driven and Reinforcement Learning Based Resource Allocation Scheme in Vehicular NetworkabstractLots of videos are transmitted for content understanding tasks in vehicular networks, which puts pressure on limited communication resources. Traditional resource allocation schemes do not consider video content, so the performance of content understanding tasks performed on the transmitted videos is not optimal. To solve this problem, in this paper, we proposed a video frames priority driven and reinforcement Q-learning based resource allocation scheme. First, we propose a video frames content priority evaluation method from the perspective of video content, and the evaluation results are the basis of resource allocation. Then, we propose a Q-learning based resource allocation scheme in vehicular network scenario, which provides more reliable resources for video frames with higher priority. Finally, experiments on real datasets validate that the proposed scheme can help to improve the performance of video content understanding tasks, such as traffic accident detection tasks. Caili Guo, Chunyan Feng, Meiyi Zhu |
ICC | 2 |
| 2021 | Cooperative Multi-player Multi-Armed Bandit: Computation Offloading in a Vehicular Cloud NetworkabstractIn recent years, computation offloading has been considered as a promising technology to support computation- intensive vehicular applications. In this paper, we mainly focus on computation offloading in a vehicular cloud network (VCN), in which both vehicles and infrastructures with resource availability are defined as resource providers. However, due to the dynamically changing on-board resource distribution and the uncoordinated offloading strategies among vehicles, the computation offloading problem in a VCN is very challenging. In this paper we first model the problem as a multi-agent multi- armed bandit problem. We then propose a reshaped upper confidence bound (UCB) algorithm to estimate the on-board resource distribution with the reward estimation. We further utilize a novel multi-agent reinforcement learning algorithm to manage the computation offloading in a VCN. Simulation results demonstrate the performance gains by using the proposed algorithm. Shilin Xu 0002, Caili Guo, Rose Qingyang Hu, Yi Qian 0001 |
ICC | 2 |
| 2021 | Learning IoV in Edge: Deep Reinforcement Learning for Edge Computing Enabled Vehicular NetworksabstractThe development of artificial intelligence, wireless communication and smart sensor platform facilities the emergence of multitude of novel vehicular applications in recent years. These new vehicular applications usually are realized with BIG models, which are normally delay sensitive and computation intensive. To alleviate the heavy pressure on the resource-constrained vehicles, computation offloading has been regarded as a promising approach to circumvent this challenge. In this paper, we take the deep learning model as an BIG model example to investigate the computation offloading in a vehicular cloud network. Specifically, task division technology is utilized to decomposed the BIG model into several components, in which there are dependencies between multiple components. To satisfy the requirements of delay-sensitive vehicular applications, it is crucial to propose an efficient computation offloading scheme. To solve it, a novel deep reinforcement learning algorithm is proposed, wherein a common deep learning model is maintained by all agents. The reward mechanism is elaborately designed to combine the long-term reward and short-term reward. In the final, the proposed algorithm’s effectiveness is verified by the experimental simulations. Shilin Xu 0002, Caili Guo, Rose Qingyang Hu, Yi Qian 0001 |
ICC | 2 |
| 2021 | Optimization of User Selection and Bandwidth Allocation for Federated Learning in VLC/RF SystemsabstractLimited radio frequency (RF) resources restrict the number of users that can participate in federated learning (FL) thus affecting FL convergence speed and performance. In this paper, we first introduce visible light communication (VLC) as a supplement to RF in FL and build a hybrid VLC/RF communication system, in which each indoor user can use both VLC and RF to transmit its FL model parameters. Then, the problem of user selection and bandwidth allocation is studied for FL implemented over a hybrid VLC/RF system aiming to optimize the FL performance. The problem is first separated into two subproblems. The first subproblem is a user selection problem with a given bandwidth allocation, which is solved by a traversal algorithm. The second subproblem is a bandwidth allocation problem with a given user selection, which is solved by a numerical method. The final user selection and bandwidth allocation are obtained by iteratively solving these two subproblems. Simulation results show that the proposed FL algorithm that efficiently uses VLC and RF for FL model transmission can improve the prediction accuracy by up to 10% compared with a conventional FL system using only RF. Chuanhong Liu, Caili Guo, Yang Yang 0057, Mingzhe Chen, H. Vincent Poor, Shuguang Cui |
WCNC | 2 |
| 2021 | A Generalized Dimming Control Scheme for Visible Light CommunicationsabstractA novel dimming control scheme, termed as generalized dimming control (GDC), is proposed for visible light communication (VLC) systems. The basic idea of GDC is to exploit adaptively time, spatial and signal domain resources for enhancing communication performance under practical illumination constraints. In particular, to satisfy the illumination constraints, an incremental algorithm for index mapping is first proposed to achieve target optical power and uniform illumination. Next, GDC having the optimal activation pattern is investigated to improve the bit-error rate (BER) performance of VLC. In particular, the BER performance of GDC is analyzed using the union bound technique. Based on the analytical BER bound, the optimal activation pattern of GDC scheme having the minimum BER criterion (GDC-MBER) is obtained by exhaustively searching all conditional pairwise error probabilities. However, since GDC-MBER requires high search complexity, two low-complexity GDC schemes having the maximum free distance criterion (GDC-MFD) are proposed. The first low-complexity GDC-MFD scheme, termed as GDC-MFD1, is derived by a lower bound of the free distance using the Rayleigh-Ritz theorem. The second low-complexity GDC-MFD scheme, termed as GDC-MFD2, is proposed to reduce further the computation by exploiting the time-invariance characteristics of the VLC channel. Simulation and numerical results show that GDC-MBER, GDC-MFD1 and GDC-MFD2 have similar BER performance, and they can achieve more than 3 dB SNR gains compared with conventional dimming control schemes for a BER of 10-4and a dimming level of 50%. Yang Yang 0057, Chunyan Feng, Caili Guo, Julian Cheng 0001, Zhimin Zeng |
IEEE Trans. Commun. | 4 |
| 2021 | A High-Coverage Camera Assisted Received Signal Strength Ratio Algorithm for Indoor Visible Light PositioningabstractA high-coverage algorithm termed enhanced camera assisted received signal strength ratio (eCA-RSSR) positioning algorithm is proposed for visible light positioning (VLP) systems. The basic idea of eCA-RSSR is to utilize visual information captured by the camera to estimate first the incidence angles of visible lights. Based on the incidence angles, eCA-RSSR utilizes the received signal strength ratio (RSSR) calculated by the photodiode (PD) to estimate the ratios of the distances between the LEDs and the receiver. Based on an Euclidean plane geometry theorem, eCA-RSSR transforms the ratios of the distances into the absolute values. In this way, eCA-RSSR only requires three LEDs for both orientation-free 2D and 3D positioning, implying that eCA-RSSR can achieve high coverage. Based on the absolute values of the distances, the linear least square method is employed to estimate the position of the receiver. Therefore, for the receiver having a small distance between the PD and the camera, the accuracy of eCA-RSSR does not depend on the starting values of the non-linear least square method and the complexity of eCA-RSSR is low. Furthermore, since the distance between the PD and camera can significantly affect the performance of eCA-RSSR, we further propose a compensation algorithm for eCA-RSSR based on the single-view geometry. Experiment results show that positioning errors of less than five centimeters is achievable for eCA-RSSR. Simulation results show that eCA-RSSR can achieve 80th percentile accuracy of about four centimeters and can improve the coverage ratio at low cost. Lin Bai 0004, Yang Yang 0057, Chunyan Feng, Caili Guo, Julian Cheng 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2020 | Novel Visible Light Communication Assisted Perspective-Four-Line Algorithm for Indoor LocalizationabstractIn this paper, a novel visible light communication (VLC) assisted Perspective-four-Line algorithm (V-P4L) is proposed for indoor localization. The basic idea of V-P4L is to joint the coordinate information obtained by VLC and the geometric information in computer vision for practical indoor localization. In particular, V-P4L first exploits the geometric information to estimate the orientation and coordinate of a single rectangular LED luminaire in the camera coordinate system based on the plane and solid geometry. Then, VLC is used to transmit the world coordinate information of the luminaire to the camera. Next, the position and pose of the camera in the world coordinate system are obtained by the single-view geometry and the linear least square method. Due to the combination of VLC and computer vision, V-P4L only requires a single luminaire for localization and does not rely on the channel model. Therefore, V-P4L is more robust to dust and other factors that interfere visible light signals. Besides, unlike conventional Perspective-n-Line (PnL) algorithms, V-P4L does not require the 3 dimensional (3D)-2 dimensional (2D) correspondences. Simulation results show that for V-P4L the location error is always less than 18 cm and the orientation error is always less than 4.5°. Lin Bai 0004, Yang Yang 0057, Chunyan Feng, Caili Guo |
GLOBECOM | 4 |
| 2020 | Value Decomposition based Multi-Task Multi-Agent Deep Reinforcement Learning in Vehicular NetworksabstractWith the development of intelligent transportation system (ITS), a multitude of novel vehicular applications have been emerging. There is an urgent need for simultaneously supporting multi-tasks across a group of vehicles in a vehicular network, forming a typical multi-task multi-agent (MTMA) environment. Deep Reinforcement Learning (DRL) is deemed a promising approach to solving the highly complicated MTMA problem. However, owing to the extraordinarily growing computational complexity as well as the explosively increasing dimension of state and action spaces in the MTMA environment, the value functions in the DRL are usually bulky and could be difficult to be learned efficiently. In this way, by virtue of the correlations among multiple vehicular tasks, we adopt the value-decomposition mechanism (VDM) to decompose the complicated value function into several small pieces and then compute each sub-function separately. The proposed paradigm can yield great speed-up in learning and help substantially with a smaller state and action space but without degrading the performance. In this work, we consider an MTMA environment with three vehicular tasks to demonstrate the effectiveness of the proposed mechanism with simulation results. Shilin Xu 0002, Caili Guo, Rose Qingyang Hu, Yi Qian 0001 |
GLOBECOM | 2 |
| 2020 | Computational Resource Sharing in a Vehicular Cloud Network via Deep Reinforcement LearningabstractWith the explosive growth of the computation intensive vehicular applications, the demand for computational resource in vehicular networks has increased dramatically. However some vehicular networks may be deployed in an environment that lack resource-rich facilities to support computationally expensive vehicular applications. In this work we propose a new scheme that enables computational resource sharing among vehicles in vehicular cloud network (VCN), which can be formulated as a complex multi-knapsack problem. In order to solve it, a deep reinforcement learning (DRL) algorithm is developed. Considering the non-stationary behavior brought in by the parallel learning and exploring processes among vehicles, computational resource sharing in such a vehicular network is a typical multiagent problem, therefore we model the problem with a Markov game problem. In addition, to tackle the heterogeneity property of the computational resources, a multi-hot encoding scheme is designed to standardize the action space in DRL. Furthermore, we propose a centralized training and decentralized execution framework that can be solved by a multi-agent deep deterministic policy gradient (MADDPG) algorithm. The numerical simulation results demonstrate the effectiveness of the proposed scheme. Shilin Xu 0002, Caili Guo, Rose Qingyang Hu, Yi Qian 0001 |
GLOBECOM | 2 |
| 2020 | Hybrid Dimming Scheme based on Transmit Antenna Selection and Precoding for MU MC VLC SystemabstractThis paper proposes a hybrid dimming scheme based on joint design of transmit antenna selection and precoding (TASP-HD) for dimmable multi-cell multi-user multiple-input single-output visible light communications systems. In TASPHD, both the direct current bias and the number of activated light-emitting-diodes are dynamically adjusted to achieve efficient communications under the constraints of dimming level and illumination uniformity. With the goal of maximizing the sum-rate of users, a joint transmit antenna selection and precoding design problem is formulated. This problem is a non-convex, mixed integer problem, and thus it is non-deterministic polynomial time hard. To solve the problem, the original problem is separated into two subproblems. The first subproblem is a mixed integer problem in terms of transmit antenna selection, which is solved by the branch-and-bound algorithm. With the obtained transmit antenna selection matrix, the second subproblem in terms of the precoding matrix is solved by the Lagrangian dual method. Finally, these two subproblems are iteratively solved to obtain a convergent solution. The simulation results verify that the proposed TASP-HD can achieve 6 dB signal-to-noise ratio gains over the analog dimming scheme when the dimming level is 70% for a target bit error rate of 10-3. Yang Yang 0057, Caili Guo, Hailun Xia |
GLOBECOM | 4 |
| 2020 | A Novel Convolutional Architecture For Video-Text RetrievalabstractThe prevalent video-text retrieval methods usually use recurrent neural networks to encode sequences of frames in videos and sequences of words in text. In this paper, we introduce an encoding architecture based entirely on convolutional neural networks. Compared to recurrent models, the complexity is smaller, and computations over all elements can be fully parallelized during training to better exploit the GPU. We use the stacking of convolution kernels of different scales to realize the encoding of local and long-term features of video and text. Experiments validate that our method achieves a new state-of-the-art for the video-text retrieval on MSR-VTT and MSVD datasets with less training time. Zheng Li 0014, Caili Guo, Zerun Feng, Hao Zhang 0161 |
ICME | 2 |
| 2020 | Enhanced User Interest and Expertise Modeling for Expert RecommendationabstractThe rapid development of Community Question Answering (CQA) satisfies users' request for professional and personal knowledge. In CQA, one key issue is to recommend users with high expertise and willingness to answer the given questions, namely expert recommendation. However, most of existing methods for expert recommendation ignore some key information, such as time information and historical feedback information, degrading the performance. On the one hand, users' interest are changing over time. It is biased if we don't consider the dynamics. On the other hand, feedback information is critical to estimate users' expertise. To solve these problems, we propose a unified framework for expert recommendation to exploit user interest and expertise more precisely. Considering the inconsistency between them, we propose to learn their embeddings separately. We leverage Long Short-Term Memory (LSTM) to model user's short-term interest and combine it with long-term interest. The user expertise is learned by the designed user expertise network, which explicitly models feedback on users' historical behavior. The extensive experiments on a large-scale dataset from a realworld CQA site demonstrate the superior performance of our method than state-of-the-art solutions to the problem. Tongze He, Caili Guo, Yunfei Chu |
ICPR | 2 |
| 2020 | Exploiting Visual Semantic Reasoning for Video-Text RetrievalabstractVideo retrieval is a challenging research topic bridging the vision and language areas and has attracted broad attention in recent years. Previous works have been devoted to representing videos by directly encoding from frame-level features. In fact, videos consist of various and abundant semantic relations to which existing methods pay less attention. To address this issue, we propose a Visual Semantic Enhanced Reasoning Network (ViSERN) to exploit reasoning between frame regions. Specifically, we consider frame regions as vertices and construct a fully-connected semantic correlation graph. Then, we perform reasoning by novel random walk rule-based graph convolutional networks to generate region features involved with semantic relations. With the benefit of reasoning, semantic interactions between regions are considered, while the impact of redundancy is suppressed. Finally, the region features are aggregated to form frame-level features for further encoding to measure video-text similarity. Extensive experiments on two public benchmark datasets validate the effectiveness of our method by achieving state-of-the-art performance due to the powerful semantic reasoning. Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014 |
IJCAI | 3 |
| 2020 | Power Efficient Deployment of VLC-enabled UAVsabstractIn this paper, a power efficient deployment for visible light communication (VLC)-enabled unmanned aerial vehicles (UAVs) is studied. In the studied model, each UAV can provide communication service for ground users and illumination builds the VLC links between UAVs and users. Hence, each UAV’s signal transmission and illumination will affect other UAVs’ signal transmission and illumination. Therefore, to deploy VLC-enabled UAVs so as to efficiently service the ground users, the interference caused by the signal transmission and illumination of UAVs must be considered. This problem is formulated as an optimization problem whose goal is to optimize the deployment of UAVs so as to minimize the power consumption for signal transmission and illumination. An iterative algorithm is first proposed to transform the optimization problem into a series of interdependent subproblems, and the transformed problems are then solved by the Lagrangian dual method. In addition, convergence and complexity of the algorithm are also analyzed. Numerical results show that the proposed scheme can reduce at least 53.7% power consumption when compared to the baselines with UAVs at the center of each sub-area. Yang Yang 0057, Caili Guo, Mingzhe Chen, Shuguang Cui, H. Vincent Poor |
PIMRC | 3 |
| 2020 | Generalized Dimming Control Scheme with Optimal Dimming Control Pattern for VLCabstractThis paper proposes a simple dimming control scheme for multi-LED visible light communications (VLC), termed as generalized dimming control (GDC) scheme. The GDC performs dimming control by simultaneously adjusting the intensity of transmitted symbols and the number of active elements in a space-time matrix. The indices of active elements in each space-time matrix, as well as the modulated constellation symbols, are used to transmit information base on the space-time index modulation. Furthermore, the tradeoff between the intensity of transmitted symbols and the number of active elements in the space-time matrix is analyzed for enhancing the reliability of communication. To improve the communication reliability, a GDC with the minimum bit error rate (BER) criterion is first proposed to select the optimal dimming control pattern based on exhaustive search. Furthermore, to reduce the computational complexity, a low complexity GDC with the maximum free distance criterion is proposed. Simulation and numerical results show that the proposed scheme can outperform conventional dimming control schemes at both low and high dimming levels in terms of the BER performance. Yang Yang 0057, Caili Guo, Zhimin Zeng, Chunyan Feng |
WCNC | 3 |
| 2020 | A cross-domain hierarchical recurrent model for personalized session-based recommendations
Yaqing Wang 0004, Caili Guo, Yunfei Chu, Jenq-Neng Hwang, Chunyan Feng |
Neurocomputing | 2 |
| 2019 | A Multi-Agent Deep Reinforcement Learning Based Spectrum Allocation Framework for D2D CommunicationsabstractDevice-to-device (D2D) communication has been recognized as a promising technique to improve spectrum efficiency. However, D2D transmission as an underlay causes severe interference, which imposes a technical challenge to spectrum allocation. Existing centralized schemes require global information, which can cause serious signaling overhead. While existing distributed solution requires frequent information exchange between users and cannot achieve global optimization. In this paper, a distributed spectrum allocation framework based on multi-agent deep reinforcement learning is proposed, named Neighbor-Agent Actor Critic (NAAC). NAAC uses neighbor users' historical information for centralized training but is executed distributedly without that information, which not only has no signal interaction during execution, but also utilizes cooperation between users to further optimize system performance. The simulation results show that the proposed framework can effectively reduce the outage probability of cellular links, improve the sum rate of D2D links and have good convergence. Zheng Li 0014, Caili Guo, Yidi Xuan |
GLOBECOM | 2 |
| 2019 | A Novel Hybrid Dimming Scheme for MU-MIMO-OFDM VLC SystemabstractMultiuser visible light communication (MU-VLC) systems utilizing multiple-input multiple-output (MIMO) and orthogonal frequency-division multiplexing (OFDM) are gaining increased attentions. Proper dimming schemes are required to deploy MU-VLC in different illumination conditions. The current dimming schemes are typically based on the signal domain, such as analogue dimming (AD) scheme and digital dimming (DD) scheme, or on the spatial domain, like spatial dimming (SD) scheme. However, these single-domain based dimming methods have their own drawbacks, such as the performance loss in AD and DD and the dimming precision deficiency in SD. To tackle the dimming challenges in MU-MIMO-OFDM VLC system, this paper proposes a block diagonalization (BD) precoding based hybrid dimming (HD-BD) scheme. In addition, transmit antenna selection (TAS) algorithm is devised for the optimal LEDs subset to operate at different dimming levels. Compared with conventional schemes, the proposed HD-BD scheme can outperform its counterparts in terms of reliability and dimming precision, which is verified by simulation results. Caili Guo, Yang Yang 0057 |
ICC | 2 |
| 2019 | Inductive Embedding Learning on Attributed Heterogeneous Networks via Multi-task Sequence-to-Sequence LearningabstractIn the paper, we study the problem of inductive embedding learning on attributed heterogeneous networks, and propose a Multi-task sequence-to-sequence learning based Inductive Network Embedding framework (MINE) capturing the attribute similarity, network proximity, and partial label information simultaneously. In particular, MINE trains an encoder function that aggregates information from a node's long-range scope of contexts, with the node attribute sequences generated by the proposed type-guided heterogeneous random walk as inputs. We present an one-to-many multi-task sequence-to-sequence model where the encoder is shared between two related tasks: an unsupervised node identity sequence generation task to learn context-aware embeddings, and a semi-supervised label prediction task to learn semantics-rich embeddings. Extensive experiments on real-world datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods. Yunfei Chu, Caili Guo, Tongze He, Yaqing Wang 0004, Jenq-Neng Hwang, Chunyan Feng |
ICDM | 2 |
| 2019 | Solving the Sparsity Problem in Recommendations via Cross-Domain Item Embedding Based on Co-ClusteringabstractSession-based recommendations recently receive much attentions due to no available user data in many cases, e.g., users are not logged-in/tracked. Most session-based methods focus on exploring abundant historical records of anonymous users but ignoring the sparsity problem, where historical data are lacking or are insufficient for items in sessions. In fact, as users' behavior is relevant across domains, information from different domains is correlative, e.g., a user tends to watch related movies in a movie domain, after listening to some movie-themed songs in a music domain (i.e., cross-domain sessions). Therefore, we can learn a complete item description to solve the sparsity problem using complementary information from related domains. In this paper, we propose an innovative method, called Cross-Domain Item Embedding method based on Co-clustering (CDIE-C), to learn cross-domain comprehensive representations of items by collectively leveraging single-domain and cross-domain sessions within a unified framework. We first extract cluster-level correlations across domains using co-clustering and filter out noise. Then, cross-domain items and clusters are embedded into a unified space by jointly capturing item-level sequence information and cluster-level correlative information. Besides, CDIE-C enhances information exchange across domains utilizing three types of relations (i.e., item-to-context-item, item-to-context-co-cluster and co-cluster-to-context-item relations). Finally, we train CDIE-C with two efficient training strategies, i.e., joint training and two-stage training. Empirical results show CDIE-C outperforms the state-of-the-art recommendation methods on three cross-domain datasets and can effectively alleviate the sparsity problem. Yaqing Wang 0004, Chunyan Feng, Caili Guo, Yunfei Chu, Jenq-Neng Hwang |
WSDM | 3 |
| 2019 | A Relay-Assisted OFDM System for VLC Uplink TransmissionabstractUplink transmission is an issue for visible light communications due to the unpleasant irradiance from the source light when placed close to the users. To overcome this problem, this paper proposes the use of relays to lower the required source optical power. A popular multi-carrier modulation scheme, termed direct-current biased optical orthogonal frequency division multiplexing, is employed to achieve high spectral efficiency. In addition, both amplitude-and-forward (AF) and decode-and-forward (DF) protocols are used. The theoretical models of AF and DF protocols are also obtained and verified by simulations. To minimize the source optical power while satisfying reliable communications, this paper formulates the associated optimization problems for AF and DF protocols. Exhaustive search is first used to obtain the optimal configuration for the system. As exhaustive search requires high computation efforts and can be time-consuming, two low-complexity suboptimal designs for AF and DF protocols are then proposed, and the proposed suboptimal designs can approximate the performance of exhaustive search in high signal-to-noise ratio regions. Numerical and experimental results indicate that when compared with the counterpart without a relay, the proposed relay-assisted system requires much lower source optical power under the constraints of reliable transmissions. Yang Yang 0057, Zhimin Zeng, Julian Cheng 0001, Caili Guo, Chunyan Feng |
IEEE Trans. Commun. | 4 |
| 2019 | User identity linkage across social networks via linked heterogeneous network embedding
Yaqing Wang 0004, Chunyan Feng, Ling Chen 0006, Hongzhi Yin, Caili Guo, Yunfei Chu |
World Wide Web | 5 |
| 2018 | Social-Guided Representation Learning for Images via Deep Heterogeneous Hypergraph EmbeddingabstractRepresentation learning for images is widely recognized as critical to the performance of end tasks such as image classification and cross-modal retrieval. However, most existing methods extract features only from visual content, far from adequate in interpreting semantics latent in images. For social images, there also exists rich social context information, e.g. owners, tags and groups, which provides cues for interpreting semantics. In this paper, we propose a representation learning framework via deep heterogeneous hypergraph embedding (DHHE), considering both visual content and social contexts. In particular, images and their contexts are first represented as a heterogeneous hypergraph, which is then embedded into a low-dimensional space. To incorporate visual information and generalize for unseen images, we learn the mapping from visual content to the semantic space. We conduct experiments with the tasks of classification, cross-modal retrieval and recommendation, which demonstrates the effectiveness of our approach and the merits of social guidance. Yunfei Chu, Chunyan Feng, Caili Guo |
ICME | 3 |
| 2018 | Polarization Modulation based Phase Noise Cancellation for Massive MIMO-OFDM SystemsabstractIn massive multiple-input multiple-output (MIMO) uplink systems, phase noise introduced by oscillators can cause severe performance loss. It leads to common phase error and inter-carrier interference in massive MIMO-OFDM uplink. To solve the issue, a novel phase noise cancellation scheme based on polarization modulation is proposed. We first introduce the polarization modulation (PM) exploited in massive MIMO-OFDM uplink. Then, by exploiting zero-forcing detection, we analyze ICI and the distribution of the transformed noise. Furthermore, we demonstrate that phase noise can be asymptotically cancelled and only the transformed additive white Gaussian noise exists as the number of antennas at the base station is very lager. Moreover, we derive the instantaneous signal-to-noise ratio (SNR) on each subcarrier and analyze the ergodic capacity. The simulation results show that the proposed scheme can effectively mitigate phase noise and achieve a higher ergodic capacity. Yao Nie, Chunyan Feng, Fangfang Liu 0008, Caili Guo |
PIMRC | 4 |
| 2018 | Robust and Low-Complexity Cooperative Spectrum Sensing via Low-Rank Matrix Recovery in Cognitive Vehicular NetworksabstractIn cognitive vehicular networks (CVNs), many envisioned applications related to safety require highly reliable connectivity. This paper investigates the issue of robust and efficient cooperative spectrum sensing in CVNs. We propose robust cooperative spectrum sensing via low‐rank matrix recovery (LRMR‐RCSS) in cognitive vehicular networks to address the uncertainty of the quality of potentially corrupted sensing data by utilizing the real spectrum occupancy matrix and corrupted data matrix, which have a simultaneously low‐rank and joint‐sparse structure. Considering that the sensing data from crowd cognitive vehicles would be vast, we extend our robust cooperative spectrum sensing algorithm to dense cognitive vehicular networks via weighted low‐rank matrix recovery (WLRMR‐RCSS) to reduce the complexity of cooperative spectrum sensing. In the WLRMR‐RCSS algorithm, we propose a correlation‐aware selection and weight assignment scheme to take advantage of secondary user (SU) diversity and reduce the cooperation overhead. Extensive simulation results demonstrate that the proposed LRMR‐RCSS and WLRMR‐RCSS algorithms have good performance in resisting malicious SU behavior. Moreover, the simulations demonstrate that the proposed WLRMR‐RCSS algorithm could be successfully applied to a dense traffic environment. Xia Liu 0006, Zhimin Zeng, Caili Guo |
Wirel. Commun. Mob. Comput. | 3 |
| 2017 | An Amplify-and-Forward Based OFDM System for VLC Uplink TransmissionabstractUplink transmission design for visible light communication (VLC) is challenging due to the high optical power requirement for reliable links, which may result in unpleasant irradiance to humans. Current studies have proposed to use radio frequency (RF) communication for uplink while treating VLC as a downlink transmission technique. However, in typical scenarios such as hospitals or aircraft cabins, non-RF communications are preferred. Thus, it is beneficial to design a system which can utilize VLC for uplink transmission while minimizing its optical power. In this paper, we propose a amplify and forward based VLC uplink system which can significantly mitigate the optical power of the source light. The multi- carrier modulation scheme termed DC biased optical orthogonal frequency division multiplexing is employed, and the theoretical model of the proposed uplink system is studied. Based on this analytical model, the key system parameters are optimized. Numerical and simulation results verify that, when compared to an equal power allocation based relay system and a reference system without a relay, the proposed system can simultaneously achieve higher spectral efficiency and lower the source optical power. Yang Yang 0057, Zhimin Zeng, Julian Cheng 0001, Caili Guo |
GLOBECOM | 4 |
| 2017 | Blind polarization oblique projection based inter-user interference cancellation in full duplex multiuser MIMO systemabstractIn full duplex (FD) multiuser MIMO system, multiple mobile stations (MSs) with FD transmission simultaneously access a given base station (BS), significant interuser interferences exist, limiting FD system performance. Existing methods addressing in the inter-user interference cancellation mostly need to know the interference channel state information and to decode and subtract the interference from the receiver. In this paper, we propose a blind polarization oblique projection based inter-user interference cancellation method without any prior information of the interference channel. This method merely utilizes the polarization state (PS) information and the autocorrelation matrix information of received signals at the downlink MS receivers to construct an oblique projection operator, which can be used to project the inter-user interference to the null space of the operator. Numerical results demonstrate that the proposed method can cancel the inter-user interference effectively when compared to the system with inter-user interference existing, and it performs more robust when the signal to noise ratio (SNR) and the receive antennas number is small. Wen Zhao 0005, Chunyan Feng, Fangfang Liu 0008, Caili Guo, Yao Nie |
PIMRC | 4 |
| 2017 | Phase Noise Self-Cancellation Scheme with Orthogonal Polarization in the Polarization Dependent Loss Channel for OFDM SystemabstractTo cancel phase noise which causes the bit error rate of OFDM systems degradation, a novel orthogonal-polarization-based phase noise self- cancellation (OP-PNSC) scheme in the polarization dependent loss (PDL) channel is proposed. In the proposed scheme, the orthogonal polarizations are utilized to transmitted orthogonally polarized signals which are added together at the receiver to cancel phase noise. These orthogonally polarized signals transmitted in the same-frequency channel enable that the proposed scheme has the advantages of more efficient in cancelling phase noise and the spectral efficiency (SE) improvement. Considering the realistic wireless channel, the distortion of PDL is investigated and a PDL pre-compensated OP- PNSC (PPC-OP-PNSC) scheme is proposed to mitigate the power imbalance caused by PDL. Then, the signal- to-interference-plus-noise ratio (SINR) and SE performances are analyzed to evaluate the performance of the OP-PNSC scheme. Finally, the numerical results show that the OP-PNSC scheme achieves performance close to that of OFDM system without phase noise in the PDL channel. Yao Nie, Chunyan Feng, Fangfang Liu 0008, Caili Guo, Wen Zhao 0005, Tiankui Zhang |
VTC Spring | 4 |
| 2017 | Polarization and Power Optimization for Spectrum Sharing in Cognitive Heterogeneous Cellular NetworkabstractCognitive heterogeneous cellular network (CHCN), comprised of small cells overlaid with macrocells, is a promising topology to fulfill the explosive demand for high-data-rate transmission by employing cognitive radio technologies. In this paper, a joint polarization and power allocation (JPPA) scheme is proposed for spectrum sharing in CHCN. By exploiting spectrum opportunities in polarization and power domains simultaneously, the proposed algorithm jointly optimizes the polarization states and transmitting power of small cell to maximize its downlink capacity subject to its own transmit-power constraint and interference-power constraint of macrocell. Numerical results reveal that the proposed scheme provides capacity and spectrum efficiency improvements compared with polarization-based spectrum sharing and power allocation schemes. In addition, it approaches the performance of the optimal exhaustive search with a much lower computational complexity. Shuo Chen 0002, Zhimin Zeng, Caili Guo |
WCNC | 3 |
| 2016 | Exploiting Polarization for Underlay Spectrum Sharing in Cognitive Heterogeneous Cellular NetworkabstractThis paper proposes a novel polarization-based underlay spectrum sharing scheme in cognitive heterogeneous cellular network (CHCN). Distinguished from traditional spectrum sharing exploiting spectrum opportunities in time, frequency, space and power domains, we exploit spectrum opportunities in polarization domain to enable small cell to coexist with macrocell. The proposed scheme optimizes the transmitting and receiving polarization states (PSs) of small cell to maximize the downlink capacity of small cell under the interference constraint of macrocell. Virtual polarization adaptation is adopted in small cell to generate the desired PSs by digital signal processing. Simulation results reveal the superiority of the proposed scheme in interference avoidance and spectrum efficiency improvement by exploiting polarization in CHCN. Shuo Chen 0002, Zhimin Zeng, Caili Guo |
GLOBECOM | 3 |
| 2016 | Exploiting polarization to resist phase noise for digital self-interference cancellation in full-duplexabstractPhase noise caused by the unideal local oscillators limits the self-interference cancellation performance severely in full-duplex systems. In this paper, a novel digital polarization self-interference cancellation method is proposed to break the bottleneck of phase noise. The basic idea here is that the polarization is insensitive to the absolute phase and will not be affected by the random phase noise. Without any prior knowledge of the phase noise, the proposed method exploits the polarization signal processing to transform the multiplicative phase noise to a new additive white Gaussian noise, which benefits from the vectorial property of polarization. Based on the polarization system and signal models with the phase noises both in the upconversion and the downconversion, we demonstrate that the self-interference can be cancelled by regeneration where only a transformed additive white Gaussian noise exists. Moreover, the proposed method obtains an upper bound for digital self-interference cancellation when the phase noise exists if the dual-polarized channel estimation is perfect. The simulation is conducted in advanced design system, and results show that the proposed method cancels the self-interference to noise floor. Fangfang Liu 0008, Songlin Jia, Caili Guo, Chunyan Feng |
ICC | 3 |
| 2016 | Monitoring leaf area index after heading stage using hyperspectral remote sensing data in riceabstractLeaf area index (LAI), as an important characterization parameter, reflects the canopy structural characteristics of crops. It is commonly used to estimate foliage cover, as well as forecasting crop growth and yield [1,2,3]. Because LAI is functionally linked to the canopy spectral reflectance, its retrieval from remote sensing data has prompted many investigations and studies in recent years. The common and widely used approach has been to develop relationships between ground-measured LAI and vegetation indices [1,4,5]. These vegetation indices performed well at the early stage of crop growth, but the estimation accuracy are greatly decreased in the late growth stages, especially after heading stage. A major problem in the use of these indices arises from the fact that canopy reflectance, it is strongly dependent on both structural and biochemical properties of the canopy [6,7,8]. In the late period of crop growth, panicles changed the canopy structure of crops and affected the crop canopy spectral reflectance [9,10]. This study compared the accuracy of monitoring LAI by using the spectral reflectance that measured the entire canopy and those canopies with panicles removed, and proposed a convenient method to removal of the effect of panicles on canopy reflectance and to enhance the prediction accuracy of LAI after heading stage of rice. Jiaoyang He, Yehui Qin, Caili Guo, Liyun Zhao, Xia Yao, Tao Cheng 0003, Yongchao Tian |
IGARSS | 3 |
| 2016 | Antenna selection based dimming scheme for indoor MIMO visible light communication systems utilizing multiple lampsabstractVisible light communication (VLC), which can synergistically provide both illumination and data transmission, is garnering increasing attention. Dimming support is one of the main challenges for VLC systems. Under the requirements of the uniformity illuminance ratio (UIR) according to lighting engineers and the indoor illuminance range standardized by the international organization for standardization (ISO), an antenna selection based dimming (ASD) scheme is proposed for multiple-input multiple-output (MIMO) VLC systems equipped with multiple lamps. By fully taking advantage of the available channel state information (CSI) at both the transmitters and the receivers, the proposed scheme can select the best subset of light-emitting diodes (LEDs) from the lamps. To the best of the author's knowledge, this is the first design of a dimming scheme specially for MIMO-VLC systems with considering the UIR as well as the indoor illuminance range. Simulation results verify that the proposed ASD scheme has significant bit error ratio (BER) performance gains over the conventional dimming scheme for various dimming proportions and receiving positions, which makes it a more promising alternative dimming scheme for MIMO-VLC systems. Zhipei Wang, Caili Guo, Yang Yang 0057 |
PIMRC | 2 |
| 2016 | A cross-polarization discrimination compensation algorithm for polarization modulationabstractIn dual-polarized channel, the BER performance of polarization modulation (PM) can be decreased by cross-polarization discrimination (XPD) effect which can introduce cross power leakage. Therefore, a XPD compensation algorithm for PM is proposed. Based on the analyses about the effect of XPD on PM, the compensation factor is obtained by the channel state information. Then by compensating the state of polarization of the received signals, the cross power leakage between two orthogonal components of the states of polarization is mitigated, and the constellation distortion of PM is reduced. The analyses reveal that the compensation algorithm can effectively improve the BER performance of PM affected by XPD. Finally, simulation results show that when the XPD is a definite value, to achieve equivalent BER performance of 2PM, the SNR needed is 2.5dB reduction after compensation. Jinjin Yuan, Fangfang Liu 0008, Caili Guo, Chunyan Feng, Yao Nie |
PIMRC | 3 |
| 2016 | A Novel Dimming Scheme for Indoor MIMO Visible Light Communication Based on Antenna SelectionabstractIn the field of indoor wireless networks, visible light communication (VLC), which can synergistically provide both illumination and data transmission, is garnering increasing attention. Realizing dimming function without corrupting the communication performance is one of the main challenges for VLC systems. By fully taking advantage of the available channel state information (CSI) at both the transmitters and the receivers, a novel dimming scheme based on antenna selection is proposed for multiple- input multiple-output (MIMO) VLC systems. To the best of the author's knowledge, this is the first design of a dimming scheme specially for MIMO- VLC systems. Simulation results for various dimming proportions and receiving positions verify that the proposed scheme has significant bit error ratio (BER) performance gains over the conventional dimming scheme, which makes it a promising alternative dimming scheme for MIMO-VLC systems. Zhipei Wang, Caili Guo, Yang Yang 0057 |
VTC Spring | 2 |
| 2016 | Performance Analysis of Cooperative Spectrum Sensing in Cognitive Vehicular Networks with Dense TrafficabstractCognitive Radio (CR) is a promising technology to solve the spectrum scarcity. Spectrum sensing becomes more challenging in Cognitive Vehicular Networks (CVNs) due to Secondary Users (SUs) mobility. Current studies on cooperative spectrum sensing usually assume that sensors are static and independent, which is unreasonable in vehicular networks with dense traffic. In this paper, we investigate the cooperative spectrum sensing performance in dense traffic with the presence of SUs mobility and correlation. First of all, we establish a mobility model in dense traffic by analyzing the trajectory data provided by Next Generation SIMulation (NGSIM) program. Secondly, detection probability and false alarm probability are investigated with mobile SUs and spatial-temporal spectrum opportunities. Next, mobility-driven sensing capacity is proposed to evaluate the sensing capacity available for mobile SUs. Note that SU mobility increases the sensing performance by providing spatial diversity and also enables SUs to achieve higher sensing capacity because of the presence of spatial spectrum opportunities. We also indicate that in dense networks, correlation is a crucial factor that may affect the network performance. In addition, the effects of protection range and primary user activity on spectrum sensing are studied. The theoretical analysis is further validated through simulations. Caili Guo, Chunyan Feng |
VTC Spring | 2 |
| 2016 | An enhanced DCO-OFDM scheme for visible light communication systemsabstractIn visible light communication (VLC) systems, optical orthogonal frequency division multiplexing (O-OFDM) is an appealing modulation scheme. Recently, a number of O-OFDM schemes have been proposed. Direct current optical orthogonal frequency-division multiplexing (DCO-OFDM) is one of the known O-OFDM schemes for its high spectral efficiency and low-complexity. Since VLC involves a combination of illumination and communication, different optical power is often required to achieve a certain illumination level. However, the performance of DCO-OFDM will severely degrade when a relatively high or low optical power constraint is imposed. To solve this problem, an enhanced DCO-OFDM (eDCO-OFDM) scheme is proposed in this paper. By introducing a piecewise function with adaptive slopes according to the required optical power, eDCO-OFDM is able to demonstrate significant performance advantages over conventional DCO-OFDM in terms of spectral efficiency and biterror rate. Comprehensive simulations were conducted to verify the efficiency of the proposed scheme. Yang Yang 0057, Zhimin Zeng, Caili Guo |
WCNC | 3 |
| 2015 | Combination of spectrum allocation and multi-relay selection in overlay cognitive radio networkabstractIn overlay cognitive radio network, the available spectrum for secondary users are dynamically grabbed from a wideband of hundreds megahertz by spectrum sensing, which leads to remarkably differential path loss among different frequencies according to propagation theory. Adjusting the global path loss of cognitive relay system through spectrum allocation can serve nontrivial increment to the performance of multi-relay selection. In this paper, we propose a novel scheme combining multi-relay selection with spectrum allocation in the cognitive radio relay system to obtain maximum signal to noise ratio (SNR) at the receiver, called SAMS (spectrum allocation and multi-relay selection). Given the available frequencies, we first search all the possible resolutions for spectrum allocation, then make multi-relay selection under each possible spectrum allocation resolution to optimize the SNR value at the receiver. Finally, we choose the largest SNR, achieve the corresponding spectrum allocation resolution and also figure out the selected relays. Compared with conventional multi-relay selection scheme (CMS), simulation results demonstrate that our proposed scheme has a 4dB-increment of SNR value at the receiver. Caili Guo, Xuekang Sun, Chunyan Feng |
PIMRC | 2 |
| 2015 | Polarization mismatch based self-interference cancellation against power amplifier nonlinear distortion in full duplex systemsabstractFull duplex systems are more spectrally efficient than conventional half duplex systems if the self-interference (SI) can be significantly mitigated. Digital cancellation is one of the lowest complexity SI cancellation techniques in full duplex systems. However, its mitigation capability is mainly limited by the power amplifier (PA) nonlinear distortion. In this paper, to cancel the SI induced by the PA nonlinear distortion, a SI cancellation scheme based on polarization mismatch (PMC) is proposed. The proposed scheme takes advantage of the characteristic that the polarization state of the SI is immune to the PA nonlinearity. And it cancels the SI in the digital baseband by a polarization mismatch matrix which has an orthogonal polarization state to that of the SI. Analysis and numerical results demonstrate that the proposed scheme can cancel the nonlinear SI induced by the PA efficiently. Besides, the desired signal to interfere and noise (SINR) gain will not deteriorate even if the power of the SI is high compared with the existing cancellation scheme. Further, the achievable rate of it can reach more than two times compared with the half duplex when the similarity coefficient of the polarization states is greater than 0.5. Wen Zhao 0005, Chunyan Feng, Fangfang Liu 0008, Caili Guo, Yao Nie |
PIMRC | 4 |
| 2015 | Time-Efficient Wideband Spectrum Sensing Based on Compressive SamplingabstractCompressed spectrum sensing (CSS) is proposed to detect spectrum opportunities efficiently over a wideband. However, most of existing CSS approaches will cause high computation costs for signal recovery when spectrum bandwidth goes large. As a result, it prolongs time for spectrum detection, which however runs counter to the original purpose of finding out spectrum opportunities over a wideband as rapidly as possible. To reduce the time consumed in signal reconstruction and realize real- time detection, we propose a novel decomposition compressed spectrum sensing (D-CSS) scheme. In D- CSS, a sparse sampling matrix is constructed first, and then it equivalently means a decomposition of the reconstructing process into two recovery subtasks. In doing so, we can scale down the overall problem and reduce the entire time for wideband spectrum detection compared with current CSS methods for a given desired sensing accuracy. Furthermore, the sparse character of our designed sampling matrix not only facilitates the operations of signal sampling and signal recovery, but also relieves the burden on random seeds generator and memory storage, which alleviates the overall implementation cost in CR practice. Caili Guo, Xuekang Sun, Chunyan Feng |
VTC Spring | 2 |
| 2015 | Polarization Based Spectrum Sensing for Cognitive Radios in Presence of Arrival AngleabstractDue to the presence of arrival angle, which can reduce the received energy of primary signal and then degrade the performance of polarization detectors, we propose polarization information based spectrum sensing methods considering arrival angle in this paper. Using the generalized likelihood ratio test (GLRT) paradigm, we derive two algorithms with different statistical prior information and finally propose a blind Polarization and Arrival angle Eigenvalue Detection (PAED), which is the optimal polarization algorithm in the presence of arrival angle. Our results show that the proposed PAED method exhibits better performance than other existing techniques, particularly when the number of samples is small, which is critical in vehicular applications. Xiaoyu Yuan, Caili Guo, Shuo Chen 0002 |
VTC Spring | 2 |
| 2015 | Correlation-Statistics-Based Spectrum Sensing Exploiting Energy and Polarization for Dual-Polarized Cognitive RadiosabstractIn this paper, we consider the problem of spectrum sensing in cognitive radios by exploiting Stokes subvector, which can completely describe energy and polarization information of the received vector signal captured by dual-polarized antennas. We first find that both component correlation between Stokes variables (i.e., the elements of Stokes subvector) and vector correlation between Stokes subvectors containing signal and noise are different from that of noise only with high probability. Therefore, two new blind detectors, namely, component-correlation-based energy-polarization detection (CCB-EPD) and vector-correlation-based energy-polarization detection (VCB-EPD), are proposed, respectively. The analysis results reveal that CCB-EPD and VCB-EPD are all constant false alarm rate detectors, and the VCB-EPD method achieves better performance than CCB-EPD when channel is low depolarized and vice versa. Simulations show that the proposed two methods exhibit better performances than other multiantenna-based detectors whether priori polarization information of primary user is known or not. We also show that the two proposed methods have performance improvement with respect to existing polarization-based detectors due to the exploitation of both energy and polarization information and the unaffectedness by noise uncertainty. The experimental results verify that the proposed two methods can satisfy the performance requirement specified by the IEEE 802.22 standard. Caili Guo, Shuo Chen 0002, Chunyan Feng, Zhimin Zeng |
IEEE Trans. Wirel. Commun. | 1 |
| 2014 | Spectrum Sensing Based on EDCAF of Signal in Multipath-Doppler ChannelabstractThe communication channel is becoming more and more complicated, which increases difficulty of signal detection of the spectrum band. The performance of signal detection is degraded due to the frequency shift and fading of the doppler multipath channel. In this paper the weak signal detection based on cyclostationary property is considered and energy detection of cyclic-autocorrelation function (EDCAF) is proposed to detect the primary signal in doppler multipath channel. Simulation shows that the proposed method has better performance than that of other detection algorithms. Furthermore, the EDCAF is simple and direct, which is feasible for cognitive radio networks. Meimei Duan, Zhiming Zeng, Caili Guo |
VTC Fall | 3 |
| 2014 | Dynamic Spectrum Sharing for TD-LTE and FD-LTE Users Based on Joint Polarization Adaption and BeamformingabstractIn this work, a joint polarization adaption and beamforming technique is proposed for horizontal spectrum sharing for TD-LTE and FD-LTE users. The polarization and spatial information of multiple dual-polarized antennas on current LTE nodes is exploited to mitigate the mutual interference among the primary, TD-LTE and FD-LTE users. By this means, the opportunities in polarization and spatial domains are exploited, which means the licensed spectrum can be shared by TD-LTE and FD-LTE users with primary users simultaneously, thus the spectrum efficiency is improved. Simulation results show that the proposed technique can improve the spectrum efficiency in comparison traditional spectrum sharing approaches that utilize the spectrum opportunities in temporal domain, while acceptable interference is introduced to the primary user. Caili Guo, Zhimin Zeng, Xiaolin Lin |
VTC Spring | 2 |
| 2014 | Adaptive ABS Configuration Scheme with Joint Power Control for Macro-Pico Heterogeneous NetworksabstractIn Heterogeneous Network (HetNet), reasonable resource allocation has attracted extensive attention. To allocate the resource more effectively, we propose a novel algorithm from the perspective of both time domain and power domain. On the one hand, we propose an adaptive Almost Blank Subframe (ABS) configuration scheme to dynamically match the network resources with the real-time load. On the other hand, we propose a utility function of macrocells' power control and a corresponding scheduling scheme to make a tradeoff between the two-tier macro-pico networks and protect the victim users as well. The existence of the optimal solution to the problem is proved, and Differential Evolution (DE) is applied to find the optimal solution. System level simulation results show that the proposed algorithm can not only enhance the load balance between macrocell and picocell, but also provide a great improvement on the performance of edge users with little overall throughput cost. Hailun Xia, Caili Guo, Yaguang Wu |
VTC Fall | 3 |
| 2014 | Spectrum sensing algorithms based on correlation statistics of polarization vector
Caili Guo, Xiaobin Wu |
Signal Process. | 1 |
| 2013 | An optimal pre-compensation based joint polarization-amplitude-phase modulation scheme for the power amplifier energy efficiency improvementabstractA Joint Polarization-Amplitude-Phase Modulation (JPAPM) scheme in wireless communication is proposed to improve the Power Amplifier (PA) energy efficiency. The proposed scheme introduces the signal's Polarization State (PS), amplitude and phase as the information-bearing parameters. Thus, the data rate can be further enhanced on the basis of the traditional amplitude-phase modulation. Also, since the transmitted signal's PS completely manipulated by orthogonally dual-polarized antennas is unaffected by the PA, JPAPM can let PA work in its nonlinear region to acquire high PA conversion efficiency. Furthermore, to mitigate the polarization-based impairment to JPAPM caused by the wireless channel's polarization dependent loss effect, the optimal pre-compensation algorithm is also presented. Simulation under the same symbol error rate and channel state shows the JPAPM can improve the PA energy efficiency significantly compared with the traditional quadrature amplitude modulation. Dong Wei 0002, Chunyan Feng, Caili Guo |
ICC | 3 |
| 2013 | Energy-efficient component carrier configuration and power control for carrier aggregated systemsabstractThe energy efficiency (EE) optimization in downlink carrier aggregated networks is addressed in this paper. To maximize EE by joint component carrier (CC) configuration and power control, we first model the problem as a mixed integer non-linear programming problem, which is NP-hard. Then the existence of optimal solution is proved. Based on the properties of optimal solution obtained by analysis, an optimal algorithm is found. Finally, a low-complex suboptimal algorithm is proposed by exploiting inherent features of CC configuration and dividing power control into two subproblems, which can be solved optimally. Simulation results show that EE of proposed algorithms is significantly better than that of conventional algorithm, and the suboptimal algorithm can greatly reduce complexity with little loss of EE in comparison to the optimal algorithm. Shengsen Wang, Chunyan Feng, Caili Guo, Guoxiang Wang |
PIMRC | 3 |
| 2013 | A two-stage cooperative spectrum sensing method for energy efficiency improvement in cognitive radioabstractCooperative spectrum sensing (CSS) can improve the performance of spectrum sensing greatly in cognitive radio (CR), however, the energy consumption in CSS also increases because there are more second users taking part in spectrum sensing. To improve the energy efficiency while maintaining the sensing accuracy to a desired threshold, we propose a simple and practical CSS method called two-stage cooperative spectrum sensing. We seek to improve the energy efficiency by decreasing the average number of SUs carrying out spectrum sensing, so we divide the sensing procedure into two stages and let it stop at the first stage if the channel is sensed as occupied. Then particle swarm optimization (PSO) algorithm is introduced to maximize the energy efficiency. Simulation results show that the convergence performance of our method is quite good and the energy efficiency of our method is significantly improved compared with the single-stage CSS. Guoxiang Wang, Caili Guo, Shulan Feng, Chunyan Feng, Shengsen Wang |
PIMRC | 2 |
| 2013 | Downlink Joint Beamforming and Power Control for Energy Efficient Multiuser MISO SystemabstractWe address the energy-efficient optimization for MU-MISO (multi user multi-input single-output) downlink. An energy efficiency (EE) optimization model under the SINR constraints is proposed, and the feasibility conditions of the proposed EE optimization model is derived. Base on the model, a novel strategy that separates EE optimization to the sum of rates maximization and total transmit power approximation is proposed. Following this strategy, a convex approximation and geometric programming based algorithm is further developed to adjust beamformers and powers jointly. The algorithm is good at EE performance and practicality. The comprehensive numerical results are provide to demonstrate the significant EE performance gain achieved by our algorithm, and the convergence of proposed algorithm is proved by both theory analysis and simulation. Shengsen Wang, Chunyan Feng, Caili Guo |
VTC Spring | 3 |
| 2013 | Polarization Mode Dispersion Tolerant Subcarrier-Power Allocation for Improving the Power Amplifier Energy Efficiency of Joint Polarization-Amplitude-Phase ModulationabstractA Polarization Mode Dispersion Tolerant Subcarrier-power Allocation (PMDTSA) scheme is proposed to improve the Power Amplifier (PA) energy efficiency. The proposed scheme is applied in the multiuser downlink system based on Joint Polarization-Amplitude-Phase Modulation (JPAPM). Dut to the wireless channel's polarization mode dispersion effect, the polarization based impairment to JPAPM on each subcarrier will be diverse. Through the optimal subcarrier-power allocation, such diversity can be utilized to make PA energy efficiency optimal under the constraint of each user terminal's data rate demand. To reduce the allocation's computationally complexity, the PMDTSA scheme performs the subcarrier allocation and power allocation separately. Firstly, assuming the equal power distribution, the subcarrier is allocated through Particle Swarm Optimization (PSO); then based the obtained subcarrier allocation method, a two-step allocation algorithm is presented to distribute the PA input power on each subcarrier. Through numerical calculation, the optimal parameters setting of PSO is examined. Our results show that the proposed PMDTSA scheme is able to achieve significant improvement in PA energy efficiency. Dong Wei 0002, Chunyan Feng, Caili Guo |
VTC Spring | 3 |
| 2013 | A novel underlay TV spectrum sharing scheme based on polarization adaption for TD-LTE systemabstractSpectrum opportunities in time, space and power domain are exploited by secondary users in traditional cognitive radio studies. While spectrum opportunities in polarization domain, namely simultaneous utilization of spectrum between primary and secondary system with polarization signal characteristics, is exploited inadequately. With the merits of polarization mismatch, primary system lies in the orthogonal polarized domain of secondary system, and interference to primary system can be mitigated. This work studies the cognitive underlay TV spectrum sharing with the merits of polarization mismatch for TD-LTE system. In this direction, a novel Polarized Underlay Spectrum Sharing (PUSS) scheme is proposed for the coexistence of TD-LTE and DTV system. It aims to characterize the spectrum efficiency of the cognitive radio setup while transmission of Polarization States (PSs) is affected by channel depolarization effect and interference constraint of DTV receiver is taken into account. Transmitting and receiving PSs are optimized by Particle Swarm Optimization with the polarized channel state information and primary PSs. Theoretical analysis and numerical results indicate that the proposed PUSS scheme could enable TD-LTE system share TV spectrum without causing much interference to incumbent DTV communication, while improvement in spectrum efficiency is achieved. Caili Guo, Zhimin Zeng, Xiaolin Lin |
WCNC | 2 |
| 2013 | An Energy Efficient Subcarrier-power Allocation scheme for Polarization-Amplitude-Phase Modulation in channel with Polarization Mode DispersionabstractAn Energy Efficient Subcarrier-power Allocation (EESA) scheme is proposed to improve the Power Amplifier (PA) energy efficiency. The proposed scheme is applied for the Polarization-Amplitude-Phase Modulation (PAPM) in the channel with Polarization Mode Dispersion (PMD). Considering a multiuser downlink system, utilizing the diversity of the polarization characteristic on each subcarrier caused by PMD, EESA scheme can make PA energy efficiency optimal under the constraint of each user terminal's data demand. To reduce the allocation's computationally complexity, the EESA scheme performs the subcarrier allocation and power allocation separately. Firstly, by assuming the equal power distribution, the subcarrier is allocated through Differential Evolution (DE); then based the obtained subcarrier allocation method, a two-step allocation algorithm is presented to distribute the PA input power on each subcarrier. Through numerical calculation, the optimal parameters setting of DE is examined. Our results show that the EESA scheme is able to achieve significant improvement in PA energy efficiency. Dong Wei 0002, Chunyan Feng, Caili Guo |
WCNC | 3 |
| 2013 | Statistical characteristic of polarization dependent lossabstractIn wireless communications, wideband polarization dependent loss (PDL) has become significant with the development of signal processing in wideband polarization domain. Aiming at deriving PDL statistical characteristic, a polarized time-variant multipath channel transfer function matrix H which considers azimuth power spectra, power delay profile (PDP), polarization power imbalance and polarization correlations is proposed. Closed-form expressions have been derived for various autocorrelation functions and cross-correlation functions for the polarized channel to fully characterize its statistical property. PDL is defined as the ratio between maximum and minimum eigenvalues of HHHat frequency f. However, due to the difficulty to analytically derive PDL's statistical characteristic, i.e., the distribution of PDL, numerical sum-of-sinusoids simulator is realized to accurately emulate the channel correlations derived above. Inspired by Weibull distribution, the probability density function (PDF) of PDL is well curve fitted using raw data from numerical simulator. The function is parameterized by scale factor, shape factor and normalization factor which all relate to polarization power imbalance and polarization correlations of the polarized channel. Results show that the degree of attenuation unbalance towards eigen-polarizations is in the descending order as: NLOS macrocell > NLOS microcell > NLOS picocell. Xiaobin Wu, Caili Guo, Chunyan Feng |
WCNC | 2 |
| 2013 | Spectrum Sensing for Cognitive Radios Based on Directional Statistics of Polarization VectorsabstractIn this paper, we propose a new blind spectrum sensing method based on the polarization characteristic of the received signal, which is completely represented by the orientation of a polarization vector. We first discuss a spectrum sensing model based on polarization vectors' orientation. Then we develop the directional statistics of polarization vectors that contain both the signal and noise or noise only. The distinctive difference between the two statistics can be used to decide whether the primary signal exists or not. Based on this, by using the well-known generalized likelihood ratio test (GLRT) paradigm, a new polarization sensing algorithm GLRT-polarization vector (GLRT-PV) is proposed. By applying directional statistics, we derive closed-form expressions for the probability of false alarm and the probability of detection under both dual-polarized additive white Gaussian noise (AWGN) and Rayleigh-fading channels. Our numerical simulation and experimental results show that the proposed method exhibits better performance than other existing methods in the case of unknown primary transmitter polarization and/or presence of noise power uncertainty. Caili Guo, Xiaobin Wu, Chunyan Feng, Zhimin Zeng |
IEEE J. Sel. Areas Commun. | 1 |
| 2012 | Spectrum sensing algorithms for cognitive radio based on polarization vector's orientationabstractIn this paper, we propose a new blind spectrum sensing method based on the polarization vector's orientation characteristics of the received signal. We first discuss a spectrum sensing model based on polarization characteristic and develop the directional statistics of a polarization vector which can be used to discriminate the signal and noise. Then, a new polarization sensing algorithm (GLRT-PV) is introduced based on the generalized likelihood ratio test (GLRT) paradigm. Applying the recent advances in directional statistics, we derive closed-form expressions of both the probability of false alarm and the probability of detection. Our simulation results show that the proposed GLRT-PV method exhibits better performance than other existing methods in the case of unknown primary transmitter polarization and/or in the presence of noise power uncertainty. Caili Guo |
GLOBECOM | 1 |
| 2010 | Virtual Polarization Detection: A Vector Signal Sensing Method for Cognitive RadiosabstractA Virtual Polarization Detection (VPD) method based on the vector signal processing, is presented in this paper for the effective spectrum sensing of cognitive radios. The purpose of such VPD method is to exploit the orthogonal polarization component of primary signals besides the conjugate, in terms of vector information elements of signals. The spectrum sensing scenario of a single secondary user and multiple primary users is discussed, where the vector signals arrived at the secondary user are received by a pair of orthogonally polarized antennae. For the test of two-class problem (the primary user present class versus absent class), the secondary user optimizes the receiving polarization state to get the maximum primary signal to noise ratio by processing the received orthogonal polarization components. Specifically, it takes place in the processor of the secondary user virtually instead of antennae devices adaptation. The VPD method does not require any prior knowledge of the signal polarization states. The performance of our VPD method is compared with that of energy detection which uses scalar amplitude information only to sense the primary users. Simulation results show that the VPD method improves the spectrum sensing performance significantly. Fangfang Liu 0008, Chunyan Feng, Caili Guo, Yue Wang 0019, Dong Wei 0002 |
VTC Spring | 3 |