VLDB 2026 Research / reviewers in the wild / expert
Xiaoming Tao 0001
dblp:16/1093-1
· DBLP profile ↗
156ranked-venue papers
5as first author
87since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 71 · 3 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 19 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Frequency Pathways for Spatiotemporal ForecastingabstractSpatiotemporal forecasting is a fundamental task in areas such as traffic flow prediction, environmental sensing, and urban planning. Recent advances have shown that decomposing temporal signals into multiple frequencies and modeling them jointly with spatial structures can significantly enhance forecasting performance. However, existing multifrequency forecasting models still face two critical limitations. First, the importance of different temporal frequencies evolves over time, yet most models assume fixed or static frequency contributions. Second, spatial dependencies are inherently frequency-sensitive. For instance, low-frequency components often align with global spatial patterns, while highfrequency components tend to correspond to localized interactions. However, current approaches typically use a shared spatial information across all frequencies, introducing spatiotemporal inconsistency. To address these challenges, we propose a novel Adaptive Frequency Pathways (AdaFre) for spatiotemporal forecasting, which adaptively captures both dynamic frequency relevance and frequency-aligned spatial structures. AdaFre employs a multi-frequency routing mechanism to dynamically select and aggregate the most informative temporal frequency components, while associating each with its corresponding spatial representation derived from frequency-aware embeddings. Spatiotemporal backbones are then used to model each path independently before final aggregation. Extensive experiments on several real-world datasets demonstrate that AdaFre significantly outperforms state-of-the-art baselines. Yanjun Qin, Yuchen Fang 0001, Xinke Jiang, Hao Miao 0001, Xiaoming Tao 0001 |
AAAI | 5 |
| 2026 | Robust Multimodal Semantic Communications with Semantic Fusion and Compensation
Zhijin Qin, Xiaoming Tao 0001, Jianhua Lu |
ICC | 3 |
| 2026 | Perceptual Audio-Visual Quality Assessment for Compressed Professionally-Generated Content: A New Dataset and An Effective Method
Shuzhan Hu, Yiping Duan, Mingqiang Yuan, Xiaoming Tao 0001 |
ICC | 6 |
| 2026 | Gumbel Sparse Attention Spatio-Temporal Network: A framework for traffic risk prediction
Dongkun Wang, Jieyang Peng, Songsheng Wang, Xiaoming Tao 0001, Yiping Duan |
Adv. Eng. Informatics | 4 |
| 2026 | Assisted refinement network based on channel information interaction for camouflaged object detection
Kuan Wang 0007, Yanjun Qin, Mengge Lu, Xiaoming Tao 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Robust and Secure STAR-RIS-Assisted UAV Communications for Multi-User and Multi-EavesdropperabstractEnabled by 6G wireless technologies, Simultaneously Transmitting And Reflecting Reconfigurable Intelligent Surfaces (STAR-RISs) create a new dimension for optimizing performance in Uncrewed Aerial Vehicle (UAV) communications through fullspace signal coverage. However, existing research on STAR-RIS-assisted UAV secure communications still faces critical challenges, including the reliance on ideal Channel State Information (CSI) assumptions, amplitude optimization complexity under mode switching protocols, and limited scalability to meet multi-user communication demands. To address these challenges, we propose a robust and secure STAR-RIS-assisted UAV communication approach for a multi-user and multi-eavesdropper scenario. By jointly optimizing user scheduling, transmitting and reflecting coefficients of STAR-RIS, transmit power and flight trajectory of UAV, we aim to maximize the average worst-case achievable secrecy rate. To tackle the non-convexity and coupled decision variables of the formulated problem, we propose an alternating optimization framework, with a Lagrange multiplier method for power allocation, a deterministic model reformulated via S-procedure for CSI uncertainty quantification and robust handling, and a penalty-based double-loop iterative algorithm forcing the phase-shift matrix toward a rank-one solution. Finally, theoretical analysis and simulation results validate the superior secrecy performance of the proposed algorithm over other representative algorithms. Xiaojie Wang 0001, Zhaolong Ning, Xiaoming Tao 0001, Lei Guo 0005, Yan Zhang 0002 |
IEEE J. Sel. Areas Commun. | 5 |
| 2026 | Semantic encoding for image compression based on semantic segmentation maps
Zhipeng Xie, Yiping Duan, Wei Kun Kong, Qiyuan Du, Qinghua Liang, Xiaoming Tao 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | RIDE: Redensification-based intrinsic density estimation for knowledge graphs
Wei Kun Kong, Yiping Duan, Xiaoming Tao 0001, Yueran Zu |
Pattern Recognit. | 3 |
| 2026 | A Joint Dynamic Partial Offloading and Real-Time Scheduling Approach for LEO Satellite-Ground NetworksabstractLow Earth Orbit (LEO) satellite networks are expected to become a key component of Sixth Generation (6G) communication networks, to relieve the communication burden on ground networks. Driven by the rapid advancement of communication technologies and intelligent applications, dense traffic flow in the Internet of Vehicles (IoV) inevitably leads to a surge in task generation and an increased demand for network resources. The requirement for low latency further intensifies this challenge, making it difficult to rely solely on ground network resources to process tasks efficiently and promptly. Conversely, relying only on satellite networks for task processing results in high costs. Therefore, flexibly integrating LEO satellite links based on real-time traffic conditions, task demands, and the real-time state of ground network resources becomes an effective solution. However, achieving such a goal poses significant challenges in efficient allocation and balance between ground and LEO satellite network resources. Therefore, we propose a dynamic multi-task partial offloading algorithm based on LEO satellite-ground network collaboration to efficiently allocate resources between ground and satellite networks in real time. We first introduce the utility gain as a metric to evaluate task scheduling preference and design an improved iterative algorithm to jointly optimize the offloading ratio and channel allocation to maximize system utility. Finally, based on the real-world dataset of Shanghai (China), we demonstrate the significant advantages of the proposed strategy over representative methods in terms of delay, vehicle satisfaction, and system utility. Xiaojie Wang 0001, Zhaolong Ning, Xiaoming Tao 0001, Lei Guo 0005, Chunxiao Jiang, Song Guo 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Deep Video Coding With Bit-Depth ScalabilityabstractDeep video coding techniques have achieved significant advancements, leading to enhanced compression performance. However, existing approaches are primarily optimized for 8-bit content, thereby limiting their effectiveness in scenarios with different bit-depths. In this paper, we propose a deep bit-depth scalable video codec (DB-SVC) that supports two-layer scalability for different bit-depths. First, we design a base layer (BL) for low bit-depth (LBD) videos, incorporating a dual-stage multi-scale feature extraction module (DFEM) to enhance compression efficiency while providing reference features for subsequent coding. Second, we introduce an inter-layer bit-depth enhancement module (IBEM) that refines the bit-depth of BL reconstructed frames by leveraging interlayer information, thus enhancing the reference quality without increasing coding overhead. Third, we design an enhancement layer (EL) tailored for high bit-depth (HBD) videos, employing a bit-depth residual compression (BRC) method to achieve a more accurate reconstruction of HBD videos. DB-SVC supports progressive decoding of LBD and HBD videos, accommodating diverse display requirements. Experimental results demonstrate that DB-SVC outperforms state-of-the-art codecs in LBD and HBD scenarios. At the same PSNR/MS-SSIM levels, DB-SVC achieves average bit-rate savings of 11.94%/53.35% for 8-bit videos and 55.98%/73.16% for 10-bit videos while comparing with VTM13.2, showcasing its superior compression performance. Zhaoqing Pan, Tiesong Zhao, Xiaoming Tao 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | ISAC Enabled Anti-UAV: Joint Beamforming and Trajectory Design for Multi-UAVs
Xiaojie Wang 0001, Zhaolong Ning, Xiaoming Tao 0001, Tie Qiu 0001, Lei Guo 0005, Yan Zhang 0002 |
IEEE Trans. Wirel. Commun. | 4 |
| 2026 | Generative Semantic Communications for Robust Speech-to-Text TranslationabstractIn this article, we propose a robust semantic communication system for speech transmission, named Ross-S2T, to execute the speech-to-text translation (S2TT) transmission efficiently. First, a deep semantic encoder is developed to directly convert speech in the source language to textual features associated with the target language, facilitating the end-to-end (E2E) semantic exchange to perform the S2TT task and reducing the amount of transmission data without performance degradation. To mitigate semantic impairments inherent in the corrupted speech, a novel generative adversarial network (GAN)-enabled deep semantic compensator is established to estimate the hidden semantic information within the speech and extract deep semantic features simultaneously, which enables robust semantic transmission for corrupted speech. Furthermore, a semantic probe-aided compensator is devised to enhance the semantic fidelity of recovered semantic features and improve the understandability of the target text. According to simulation results, the proposed Ross-S2T exhibits superior S2TT performance compared to conventional approaches and high robustness against semantic impairments. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2026 | Joint Semantic-Channel Coding and Modulation for Token CommunicationsabstractIn recent years, the Transformer architecture has achieved outstanding performance across a wide range of tasks and modalities. Token is the unified input and output representation in Transformer-based models, which has become a fundamental information unit. In this work, we consider the problem of token communication, studying how to transmit tokens efficiently and reliably. Point cloud, a prevailing three-dimensional format which exhibits a more complex spatial structure compared to image or video, is chosen to be the information source. We utilize the set abstraction method to obtain point tokens. Subsequently, to get a more informative and transmission-friendly representation based on tokens, we propose a joint semantic-channel and modulation (JSCCM) scheme for the token encoder, mapping point tokens to standard digital constellation points (modulated tokens). Specifically, the JSCCM consists of two parallel Point Transformer-based encoders and a differential modulator which combines the Gumel-softmax and soft quantization methods. Besides, the rate allocator and channel adapter are developed, facilitating adaptive generation of high-quality modulated tokens conditioned on both semantic information and channel conditions. Extensive simulations demonstrate that the proposed method outperforms both joint semantic-channel coding and traditional separate coding, achieving over 1dB gain in reconstruction and more than 6× compression ratio in modulated symbols. Jingkai Ying, Zhijin Qin, Yulong Feng, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2025 | Lite3D: A Lightweight Hybrid CNN-Transformer Framework for 3D Video StabilizationabstractVideo stabilization plays a critical role in IoT, as it enhances the accuracy and reliability of visual data collected from mobile and unstable devices, enabling more effective monitoring and analysis. However, existing stabilization methods frequently lack sufficient accuracy in motion estimation or struggle with expensive computation. In this paper, we introduce Lite3D, a novel lightweight 3D deep learning-based video stabilization framework that combines a hybrid CNN and Transformer architecture for more efficient feature extraction. During training, our method utilizes estimated depth maps and relative camera poses to generate target views, jointly optimizing DepthNet and PoseNet through an unsupervised learning strategy. Compared to existing 3D video stabilization methods, our approach enables more accurate learning of motion information. In the inference phase, we smooth the estimated camera pose trajectory and synthesize stabilized frames using the optimal depth maps. Experimental results on the NUS dataset demonstrate that Lite3D achieves state-of-the-art stability, with additional tests on the DeepStab dataset confirming strong generalization. Moreover, our method contains fewer parameters compared to existing 3D deep learning-based video stabilization, offering both high stabilization quality and computational efficiency. Zejing Shan, Yiping Duan, Yue Wu 0004, Qiyuan Du, Xiaoming Tao 0001 |
ICC | 6 |
| 2025 | Joint Semantic-Channel Coding and Modulation for Point Cloud
Jingkai Ying, Zhijin Qin, Xiaoming Tao 0001 |
ICC | 4 |
| 2025 | Cloud-Edge-End Collaborative Surveillance Video Transmission with Object-Guided Video Super-ResolutionabstractCloud Video Surveillance (CVS) systems, as the backbone of distributed surveillance networks, face increasing challenges in transmitting large volumes of high-resolution video data. While cloud-end collaborative video transmission methods can reduce bitrates beyond conventional compression techniques, the periodic transmission of high-resolution keyframes consumes significant bandwidth due to semantically irrelevant background information. To address this, we propose a cloud-edge-end collaborative video transmission scheme based on object-guided video super-resolution. In this scheme, the edge extracts key objects from sparsely selected keyframes at the end and transmits them along with low-resolution video to the cloud, where our Keyframe-Guided Video Restoration Transformer (KG-VRT) is used to improve the video quality. Experimental results on public datasets show that our network outperforms state-of-the-art keyframe-based baselines with a 1.73 dB PSNR improvement and maintains robust performance even with keyframe intervals of up to 30 frames. A comparative analysis of two transmission strategies—transmitting full keyframes versus transmitting only key object regions—demonstrates a 60% – 80% reduction in keyframe bitrate while maintaining object detection accuracy at a significantly reduced overall system bitrate. This highlights the efficiency and scalability of our approach in bandwidth-constrained surveillance scenarios. Yiping Duan, Xiaoming Tao 0001, Wei Kun Kong, Qiyuan Du |
VTC2025-Fall | 3 |
| 2025 | A diffusion-based feature enhancement approach for driving behavior classification with EEG data
Yanjun Qin, Shanghang Zhang, Xiaoming Tao 0001 |
Adv. Eng. Informatics | 4 |
| 2025 | EEG-Driven Classification of Driver Mental Workload in Diverse Environments: A Dual-Branch Network for Efficient In-Vehicle ApplicationsabstractThe mental load of drivers can profoundly affect their driving performance, to the extent that it affects traffic safety. Therefore, monitoring mental workload has become a crucial aspect of sensor-based driver monitoring systems, especially in the context of the Industrial Internet of Things (IIoT), where driver status information can be exchanged between vehicles to enhance safety. However, the substantial energy consumption and transmission latency associated with traditional central-server-based IoT systems are prominent issues that necessitate the development of lighter algorithms for edge computing in individual vehicles. In this article, we focus on the impact of external traffic events and environmental changes on the mental load of drivers, as well as effective classification algorithms applied in monitoring systems. To analyze the physiological responses of drivers to road events and non driving related tasks under different weather conditions, we proposed a dual branch model, DMW-Net, based on attention mechanism branches and graph attention modules to discriminate the mental load level of drivers from physiological signals. The proposed method was validated on the manD dataset and achieved an accuracy of 90.07% in physiological signals of three different load levels, which is higher than the comparison models. This study provides innovative methods for driver monitoring systems, contributing to advanced driving assistance systems (ADAS) and traffic safety. Yanjun Qin, Shanghang Zhang, Yiping Duan, Xiaoming Tao 0001 |
IEEE Internet Things J. | 5 |
| 2025 | Object-Attribute-Relation Representation-Based Video Semantic CommunicationabstractWith the rapid growth of multimedia data volume, there is an increasing need for efficient video transmission in applications such as virtual reality and future video streaming services. Semantic communication is emerging as a vital technique for ensuring efficient and reliable transmission in low-bandwidth, high-noise settings. However, most current approaches focus on joint source-channel coding (JSCC) that depends on end-to-end training. These methods often lack an interpretable semantic representation and struggle with adaptability to various downstream tasks. In this paper, we introduce the use of object-attribute-relation (OAR) as a semantic framework for videos to facilitate low bit-rate coding and enhance the JSCC process for more effective video transmission. We utilize OAR sequences for both low bit-rate representation and generative video reconstruction. Additionally, we incorporate OAR into the image JSCC model to prioritize communication resources for areas more critical to downstream tasks. Our experiments on traffic surveillance video datasets assess the effectiveness of our approach in terms of video transmission performance. The empirical findings demonstrate that our OAR-based video coding method not only outperforms H.265 coding at lower bit-rates but also synergizes with JSCC to deliver robust and efficient video transmission. Qiyuan Du, Yiping Duan, Qianqian Yang 0002, Xiaoming Tao 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 4 |
| 2025 | Data-Free Cloud-Edge Distillation for Safe and Efficient Intelligent CommunicationsabstractEfficiency and security are the core challenges in the intelligent communication field. Lightweight neural networks have accelerated the information interpretation and communication efficiency, thereby fostering the rapid development of the Internet of Things. Enhancing the recognition capability of lightweight neural networks remains challenging. Knowledge distillation, a technique that transfers knowledge from a complex model to a smaller one, is often used to improve the recognition performance of lightweight networks. However, practical issues such as transmission constraints and user privacy make the original data required for knowledge distillation difficult to access directly. To tackle this issue, this paper proposes a Data-Free Cloud-Edge Knowledge Distillation (DF-CEKD) model, which uses a complex network in the cloud to provide training guidance for lightweight networks deployed on mobile devices. Specifically, DF-CEKD employs a novel Deep Inversion Diffusion Generation (DIDG) module to provide proxy data as input for the distillation process, thereby transferring the feature learning capability from the cloud network to the edge network. Meanwhile, a Multi-Layer Feature Joint Supervision Distillation (MLF-JSD) module is designed to further enhance the feature selection guidance provided by the teacher network in the cloud for training the lightweight student network. The simulation results demonstrate that the proposed DF-CEKD reduces the number of parameters to 1/20 and the floating-point operations to 1/12 when distilling from WRN40-2 to WRN16-1, resulting in only a 0.27% decrease in accuracy. Xiufang Li, Yiping Duan, Xiaoming Tao 0001, Qigong Sun, Qiyuan Du, Qianqian Yang 0002, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | A Robust Image Semantic Communication System With Multi-Scale Vision TransformerabstractSemantic communications have demonstrated exceptional performance across various tasks, yet they are susceptible to semantic impairments due to the inherent vulnerability of deep neural networks. This paper focuses on semantic impairments in images, particularly those stemming from adversarial perturbations. We introduce a novel metric for quantifying the level of semantic impairment and create a semantic impairment dataset. Furthermore, we propose a deep learning enabled semantic communication system for robust image transmission, termed as DeepSC-RI. The proposed system harnesses a multi-scale semantic extractor with a dual-branch design tailored for extracting semantics with varying granularity, thereby boosting the robustness of the system. The fine-grained branch incorporates a semantic importance evaluation module to identify and prioritize crucial semantics through self-attention score manipulations, while the coarse-grained branch adopts a hierarchical approach for progressively capturing the robust semantics. These two streams of semantics are seamlessly integrated via an advanced cross-attention-based semantic fusion module. Experimental results highlight the superior performance of DeepSC-RI under diverse channel conditions, across various levels of semantic impairment intensity, and in multiple tasks. Zhijin Qin, Xiaoming Tao 0001, Jianhua Lu, Khaled Ben Letaief |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Empowering Corner Case Detection in Autonomous Vehicles With Multimodal Large Language ModelsabstractObject detection powered by deep learning is an essential component in the realm of self-driving vehicles. However, the model may be affected by corner cases, which are rare or unusual objects and scenarios, and can significantly impact the reliability of object detection systems. In this paper, we applied a Multimodal Large Language Model (MLLM) to address the challenge of corner cases in autonomous driving systems. The MLLM consists of an image encoder, a text tokenizer, a modal alignment layer, and a pre-trained large language model, enabling the model to understand multimodal semantic information. We added text descriptions on the basis of corner case dataset CODA and constructed the CODA-REC dataset. This dataset is then used to perform instruction fine-tuning on the MLLM to adapt it to the object detection task. The proposed method leverages the extensive knowledge and zero-shot learning capabilities of LLMs to enhance the semantic understanding of text and images, enabling the detection and appropriate response to corner cases that were previously difficult to handle. The experimental results show that MLLM achieved better performance than baseline models, with an improvement of about 10% in mAR and mAP metrics compared to most closed-set models, and an improvement of 10% mAP compared to open set models. We hope that our work can inspire the application of MLLMs in the field of autonomous driving, contributing to more advanced intelligent transportation systems. Yanjun Qin, Shanghang Zhang, Xiaoming Tao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Brain-Inspired Video Quality Assessment via Visual-EEG Feature AlignmentabstractVideo quality assessment (VQA) is crucial in applications such as video calls, real-time meetings, and surveillance, where video quality directly impacts user experience greatly. Traditional objective methods like SSIM and PSNR fail to capture the subjective perception of video quality, while subjective Quality of Experience (QoE) assessment metrics like Mean Opinion Score (MOS) are not scalable for large-scale automated VQA tasks. To overcome these limitations, deep learning approaches have emerged, but mostly focusing only on a single video modality, extracting low-level visual features such as color and texture. Recently, electroencephalography (EEG) has been shown to align with users' subjective experiences, offering valuable insights into neural responses to visual content. Hence, in this letter, we propose a brain-inspired deep learning framework for VQA that aligns EEG and video features. We build a video distortion dataset annotated with both MOS and EEG signals to analyze the impact of video distortions on EEG responses and subjective ratings. We then employ an adaptive EEG feature learning network to extract EEG features linked to video distortions, and propose a video quality prediction network that aligns both video and EEG features using a three-stage training strategy. Our method outperforms existing techniques, showing strong alignment with human subjective ratings. Experimental results validate the effectiveness of EEG in enhancing VQA with a more human-centric approach. Shuzhan Hu, Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Group Image Compression for Dual Use of Machine and Human VisionabstractFaces in a scene of human group, if coded with sufficient precision, can be computer analyzed for machine vision tasks involving faces. But this requires storing and communicating them at a very high bit rate. Traditional ROI-based image compression methods are ill suited to code many faces at high precision against a complex background. In this work, we propose a novel group image compression neural network (GICNet) of two layers: 1) the face layer dedicated to machine analysis, in which face bounding boxes are first cropped out of the background and converted to a compression-friendly canonical sketch-guided representation of fixed resolution for compact coding and facilitating downstream tasks without additional preprocessing; 2) the background layer dedicated to overall human vision perceptual quality, in which face residuals and background elements are coded and appended to the code stream. Experimental results demonstrate the effectiveness of our proposed GICNet, conserving up to 13%-57% bitrate for machine vision applications while maintaining competitive perceptual quality. Xiaolin Wu 0001, Fan Li 0003, Yiping Duan, Xiaoming Tao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Diversifying Latent Flows for Safety-Critical Scenarios Generation With CARLA SimulatorabstractThe likelihood of encountering scenarios that lead to accidents, namely safety-critical scenarios, is minimal compared to long-term safe driving environments. The generation of repeatable and scalable safety-critical scenarios is essential for the advancement of human and autonomous driving capabilities. Compared with the high complexity and low practicality of existing scenario generation methods, in this paper we propose a real-time approach to automatically generate challenging scenarios and instantiate them in a CARLA-based simulator. First, the safety-critical scenario is decomposed into a perturbed and optimized vehicle trajectory and the remaining reusable Unreal Engine assets based on a hierarchical model. Second, a model that is based on a graph conditional variational autoencoder (VAE) is employed to predict future trajectories and head angles based on past information. Third, the safety-critical scene generation model is used to enhance the diversity of the scene by diversifying the latent variables over a pre-trained trajectory representation model. Finally, the trajectories of real-world vehicles are placed into the simulator by adapting them to enable the generation of safety-critical scenes in a three-dimensional environment. The results demonstrate that the proposed approach generates scenarios that are more plausible than those generated by the baselines, with a performance improvement of over 10% in collision metrics for scenario generation. The research facilitates the simplification of the long-tail scenario construction process for autonomous vehicles, which in turn facilitates the optimization of algorithms such as autonomous trajectory planning. Dingcheng Gao, Yanjun Qin, Xiaoming Tao 0001, Jianhua Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Cross-Scenario Vigilance Detection Based on EEG Analysis for Safety Driving in AutonomousabstractSafety driver vigilance is a prerequisite for the safe operation of autonomous vehicles. In contrast to vehicle behavioral trajectory detection, which suffers from high latency and low accuracy, vigilance detection based on physiological signals is currently the most reliable and accurate method. While vigilance monitoring methods using electroencephalograms (EEG) have made considerable progress in experimental scenarios, they remain a challenging problem in scenario-constrained conditions, such as high-speed moving autonomous vehicles. This is due to the low signal-to-noise ratio in EEG signal acquisition and the difficulty of real-time processing. Moreover, cumbersome data acquisition processes and the challenges of labeling have hindered progress in this area. Given the successful use of EEG for monitoring in experimental settings, we believe that the transfer of knowledge learned from these scenarios to new contexts is reasonably feasible. Thus, this work aims to bridge the domain gap between experimental and real-world scenarios while balancing the number of channels and accuracy. Specifically, we propose a framework for EEG vigilance detection capable ofCross-scenario,Cross-subject, andCross-device, calledCCC. The proposed framework leverages the standard montage structure of EEG channels, reducing the number of channels by considering the common regions of EEG channels across different scenarios. The results show that our proposed model achieves an average accuracy of 86.20% on the SEED-VIG dataset with 12 subjects, which is higher than the 82.21% achieved by state-of-the-art deep learning approaches. Finally, we investigate the role of the attention mechanism and transfer learning, and further attempt to explain the advantages of our proposed approach from a visualization perspective. Dingcheng Gao, Xiaoming Tao 0001, Xia Wu 0001, Yanjun Qin, Jianhua Lu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Synchronous Multi-Modal Semantic Communication System With Packet-Level CodingabstractAlthough the semantic communication with joint semantic-channel coding design has shown promising performance in transmitting data of different modalities over physical layer channels, the synchronization and packet-level forward error correction (FEC) of multimodal semantics have not been well studied. Synchronizing multimodal features in both the semantic and time domains is challenging due to the independent design of semantic encoders. In this paper, we take the facial video and speech transmission as an example and propose a Synchronous Multi-modal Semantic Communication System with Packet-Level Coding (SyncSC). To achieve semantic and time synchronization, 3D Morphable Mode (3DMM) coefficients and text are transmitted as semantics. We propose a semantic codec that achieves similar reconstruction quality with lower bandwidth. The visual-guided speech synthesis is designed to synchronize video, text and speech. We propose a packet-Level FEC method for video semantics, called PacSC, that maintains visual quality even at high packet loss rates. For text packets, a text packet loss concealment module, called TextPC, based on Bidirectional Encoder Representations from Transformers (BERT) is proposed, which improves the performance of traditional FEC methods. Simulation results show that SyncSC reduces transmission overhead while ensuring high-quality synchronous transmission of video and speech over the packet loss network. Jingkai Ying, Zhijin Qin, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | Semantic Encoding and Decoding for Talking Head TransmissionabstractSemantic communication focuses on transmitting only information of interest instead of raw data. It holds the potential to overcome the limitations of traditional communication methods. A critical aspect of semantic communication is semantic encoding and decoding, which deals with information compression and reconstruction. Existing key point-based talking head encoding and decoding methods encounter challenges when dealing with large changes in the pose of a human face. In this study, a semantic encoding and decoding method was proposed. Firstly, a sketch-based semantic decoding method based on generative adversarial networks was proposed. Secondly, a structure similarity loss was proposed to further enhance the performance of the sketch decoding performance. Finally, two pose adaptive semantic encoding and decoding methods were proposed to adaptively combine two modalities of semantic information. Experimental results showed that the proposed method achieved significantly better video reconstruction performance than baseline methods. The ablation study result showed that the proposed structure similarity loss can enhance the sketch decoding performance. The proposed method provides a novel approach to solving the facial distortion problem when the pose of the human face changes largely and improves the performance of semantic encoding and decoding systems. Baoping Cheng, Xiaoyan Xie, Xiaoming Tao 0001 |
GLOBECOM | 6 |
| 2024 | Brain-Inspired VR Video Quality Assessment Based on ElectroencephalographyabstractWith the rapid development of virtual reality (VR) technology, users are able to access a large number of new applications in their daily lives. VR expands users’ perceptual dimensions, bringing them entirely new experience. However, the user experience assessment for VR videos is still under exploration, which remains an unresolved issue. In such immersive scenarios, the methods based on user scoring require active feedback from users, which will interrupt the immersion experience. Besides, it is difficult to monitor the user experience status in real-time by user scoring. With the development of psychophysiological research, electroencephalographic (EEG) signal measurement is considered to have the potential to non-intrusively obtain the user experience. Hence, this paper employs EEG measurements to capture users’ EEG signals while watching VR videos with varying levels of stuttering, constructing a VR-EEG dataset. Subsequently, we analyze the dataset using time-frequency analysis methods to validate the feasibility of EEG signals reflecting user experience. Finally, we utilize machine learning methods to construct a QoE measurement network capable of analyzing users’ perceptual experience from single-trial EEG signals. Experimental results demonstrate that the proposed method establishes a relationship between brain activities and user experience and can effectively predict QoE scores from EEG signals. It provides a technical means for real-time, non-disturbing measurement of user experience in VR video playback. Shuzhan Hu, Jian Chu, Yiping Duan, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 4 |
| 2024 | A Robust Semantic Communication System for Image TransmissionabstractSemantic communications have gained significant attention as a promising approach to address the transmission bottleneck, especially with the continuous development of 6G techniques. Distinct from the well investigated physical channel impairments, this paper focuses on semantic impairments in images, particularly those arising from adversarial perturbations. Specifically, we propose a novel metric for quantifying the intensity of semantic impairment and develop a semantic impairment dataset. Furthermore, we introduce a deep learning enabled semantic communication system, termed as DeepSC-RI, to enhance the robustness of image transmission, which incorporates a multi-scale semantic extractor with a dual-branch architecture for extracting semantics with varying granularity, thereby improving the robustness of the system. The fine-grained branch incorporates a semantic importance evaluation module to identify and prioritize crucial semantics, while the coarse-grained branch adopts a hierarchical approach for capturing the robust semantics. These two streams of semantics are seamlessly integrated via an advanced cross-attention-based semantic fusion module. Experimental results demonstrate the superior performance of DeepSC-RI under various levels of semantic impairment intensity. Zhijin Qin, Xiaoming Tao 0001, Jianhua Lu, Khaled Ben Letaief |
GLOBECOM | 3 |
| 2024 | Semantic Security: A Digital Watermark Method for Image Semantic PreservationabstractDigital watermarking has long been used to protect digital images from abuse. However, applying digital watermarking to semantic communication remains a challenge. This work introduces a secure coding method that combines semantic coding and digital watermarking techniques. The proposed method selects points in the target image with high semantic importance and embeds watermark information into their position vectors, perpendicular to the embedding domains of previous works. The experiments conducted on the Cityscapes dataset demonstrate that our method integrates well with semantic communication systems. Compared to the previous approach, our proposed method can more completely preserve the structural features of the target image and is better suited for Machine Type Communication (MTC) tasks, such as target detection and semantic segmentation. Tianwei Zuo, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001 |
ICASSP | 4 |
| 2024 | Synchronous Semantic Communications for Video and SpeechabstractAlthough semantic communication has shown great performance in various types of data transmission, the problem of semantic synchronization between multimodal data has not been well studied. Semantic synchronization is a challenging issue that requires the transmitted information to be synchronized in both semantic and time domains. In this article, we propose a synchronous semantic communication system for video and speech transmission, which the real-time facial transmission is adopted as the use case. Particularly, to achieve time domain synchronization, we design an efficient semantic transmitter to send multimodal data packets. 3D Morphable Mode (3DMM) coefficients and text are employed as semantic information, achieving semantic interactivity and lower bandwidth. To address synchronization in semantic domain, we firstly employ the visual voice clone at the receiver. Visual-guided speech synthesis module is designed to align text and facial semantics. Thus, the generated speech is synchronized with video frames in both semantic and time domains. The simulation results show that our proposed system achieves high-quality synchronous transmission of video and speech with reducing transmission overhead. Jingkai Ying, Zhijin Qin, Xiaoming Tao 0001 |
ICC | 5 |
| 2024 | Computer Vision Based Link Scheduling in mmWave Multi-Hop V2X CommunicationsabstractIn this paper, we present a novel multi-hop link scheduling framework that utilizes the vision perception from cameras of the road-side unit (RSU) to support the large-capacity and reliable transmission of the high-speed dynamic vehicle network. Specifically, we propose a vision based link state identification method to determine whether the communication links between RSU and different vehicles are blocked or connected. The 3D detection technique is firstly used to obtain the vehicle spatial distribution in surrounding environment. Then, the geometric calculation is adopted to accurately analyze the link states between RSU and different vehicles. Moreover, we design an environmental statistical information based low-complexity link scheduling method. The joint statistical distribution of the residual transmission distance and the residual multi-hop latency is used to optimize the total multi-hop latency. Simulation results show that the proposed vision based link state identification method can significantly outperform the exiting methods, and the proposed link scheduling method can approximately achieve the optimal performance as that from the exhaustive search method but with much less computation overhead. Weihua Xu 0001, Feifei Gao 0001, Ling Xing 0001, Shaodan Ma, Xiaoming Tao 0001 |
WCNC | 5 |
| 2024 | Vision-Aided Reference Signal Receiving Power Prediction for Smart FactoryabstractSmart factory is a new intelligent platform requiring high throughput and millimeter wave (mmWave) technology has become an enabler for high speed communications in Industry 4.0. However, the sensitivity of mmWave signals to blockage poses serious challenges to the reliability of wireless networks in these frequency ranges. In this paper, we propose a vision-aided reference signal receiving power prediction (RSRP) framework for smart factory to avoid communications interruption caused by unexpected blockage. In particular, we design a feature extraction method to obtain communications-related features in environmental images. Then, we construct a joint image-channel dataset based on Blender and Wireless Insite software. Simulations show that the root mean square error (RMSE) of RSRP prediction 400 ms ahead reaches 2.88 dB. RSRP prediction can assist base station (BS) handover to avoid communications interruption. Hence, the proposed study provides a promising direction for enabling ultra-reliable communications under mmWave and even Terahertz bands in smart factory of Industry 4.0. Feifei Gao 0001, Xiaoming Tao 0001, Shaodan Ma, H. Vincent Poor |
WCNC | 3 |
| 2024 | DMGSTCN: Dynamic Multigraph Spatio-Temporal Convolution Network for Traffic ForecastingabstractTraffic forecasting belongs to intelligent transportation systems and is helpful for public property and life safety. Therefore, to forecast traffic accurately, researchers pay great attention to dealing with complex problems by mining intricate spatial and temporal dependencies of the traffic. However, some challenges still hold back traffic forecasting: 1) Most studies mainly focus on modeling correlations of traffic time series of close distances on the road network and ignore correlations of remote but similar traffic time series; 2) Previous static graph-based methods failed to reflect the dynamic changed spatial relations of multiple time series in the evolving traffic system. To tackle the above issues, we design a new dynamic multi-graph spatio-temporal convolution network (DMGSTCN) in this paper, which utilizes the gated causal convolution with the dynamic multi-graph convolution network (DMGCN) to simultaneously extract spatial and temporal information. Specifically, DMGCN uses not only distance-based graphs but also structure-based graphs to obtain spatial information from nearby and remote but similar traffic time series, respectively. Moreover, to dynamically model spatial correlations, DMGCN first splits neighbors of each traffic time series into different regions according to relative position relationships. Then DMGCN assigns different weights to different regions at different time slices. Empirical evaluations on four traffic forecasting benchmarks reveal that DMGSTCN outperforms existing methods. Yanjun Qin, Xiaoming Tao 0001, Yuchen Fang 0001, Haiyong Luo, Fang Zhao 0003, Chenxing Wang 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Decoding Brain-Controlled Intention for UAVs and IVs Based on Lightweight NetworkabstractBrain-Computer-Interface (BCI) plays an important role in the Internet of Things (IoT). With the development of electroencephalogram (EEG) signal processing and deep learning, researchers are beginning to decode intention for controlling smart movable devices by EEG signals, such as unmanned aerial vehicles (UAVs) and intelligent vehicles (IVs). However, the current related studies less consider the generic decoding of the two, and the performance of the adopted models needs to be improved in terms of both lightness and accuracy. In this paper, we are dedicated to the study of a generic brain-controlled intention decoding task serving UAVs and IVs. We adopt three different datasets, encompassing different BCI paradigms. A lightweight self-attention enhancement model is proposed, which incorporates self-attention into the model for feature enhancement. Experimental results show that our method outperforms the baselines on brain-controlled intention decoding for all three datasets. This research is useful for the work in the field of brain-controlled intention decoding, and also provides new ideas for lightweight network-based control of UAVs and IVs. Wenqi Zhang 0003, Yanjun Qin, Xiaoming Tao 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Federated Multi-View Synthesizing for MetaverseabstractThe metaverse is expected to provide immersive entertainment, education, and business applications. However, virtual reality (VR) transmission over wireless networks is data- and computation-intensive, making it critical to introduce novel solutions that meet stringent quality-of-service requirements. With recent advances in edge intelligence and deep learning, we have developed a novel multi-view synthesizing framework that can efficiently provide computation, storage, and communication resources for wireless content delivery in the metaverse. We propose a three-dimensional (3D)-aware generative model that uses collections of single-view images. These single-view images are transmitted to a group of users with overlapping fields of view, which avoids massive content transmission compared to transmitting tiles or whole 3D models. We then present a federated learning approach to guarantee an efficient learning process. The training performance can be improved by characterizing the vertical and horizontal data samples with a large latent feature space, while low-latency communication can be achieved with a reduced number of transmitted parameters during federated learning. We also propose a federated transfer learning framework to enable fast domain adaptation to different target domains. Simulation results have demonstrated the effectiveness of our proposed federated multi-view synthesizing framework for VR content delivery. Yiyu Guo, Zhijin Qin, Xiaoming Tao 0001, Geoffrey Ye Li |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | Brain-Inspired Image Perceptual Quality Assessment Based on EEG: A QoE PerspectiveabstractHuman-oriented image communication should take the quality of experience (QoE) as an optimization goal, which requires effective image perceptual quality metrics. However, traditional user-based assessment metrics are limited by the deviation caused by human high-level cognitive activities. To tackle this issue, in this paper, we construct a brain response-based image perceptual quality metric and develop a brain-inspired network to assess the image perceptual quality based on it. Our method aims to establish the relationship between image quality changes and underlying brain responses in image compression scenarios using the electroencephalography (EEG) approach. We first establish EEG datasets by collecting the corresponding EEG signals when subjects watch distorted images. Then, we design a measurement model to extract EEG features that reflect human perception to establish a new image perceptual quality metric: EEG perceptual score (EPS). To use this metric in practical scenarios, we embed the brain perception process into a prediction model to generate the EPS directly from the input images. Experimental results show that our proposed measurement model and prediction model can achieve better performance. The proposed brain response-based image perceptual quality metric can measure the human brain's perceptual state more accurately, thus performing a better assessment of image perceptual quality. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu, Guangyi Liu 0001, Zhimin Zheng, Chengkang Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | AI Empowered Wireless Communications: From Bits to SemanticsabstractArtificial intelligence (AI) and machine learning (ML) have shown tremendous potential in reshaping the landscape of wireless communications and are, therefore, widely expected to be an indispensable part of the next-generation wireless network. This article presents an overview of how AI/ML and wireless communications interact synergistically to improve system performance and provides useful tips and tricks on realizing such performance gains when training AI/ML models. In particular, we discuss in detail the use of AI/ML to revolutionize key physical layer and lower medium access control (MAC) layer functionalities in traditional wireless communication systems. In addition, we provide a comprehensive overview of the AI/ML-enabled semantic communication systems, including key techniques from data generation to transmission. We also investigate the role of AI/ML as an optimization tool to facilitate the design of efficient resource allocation algorithms in wireless communication networks at both bit and semantic levels. Finally, we analyze major challenges and roadblocks in applying AI/ML in practical wireless system design and share our thoughts and insights on potential solutions. Zhijin Qin, Le Liang, Shi Jin 0002, Xiaoming Tao 0001, Wen Tong, Geoffrey Ye Li |
Proc. IEEE | 5 |
| 2024 | MMPHGCN: A Hypergraph Convolutional Network for Detection of Driver Intention on Multimodal Physiological SignalsabstractDetecting driving behavior is crucial as it is closely related to driving safety. With the development of electroencephalogram (EEG) signal processing, researchers have started studying driving intentions through EEG signals. However, current research mainly focuses on a single intention, and EEG signals suffer from low spatial resolution, leading to less accurate detection. This paper aims to investigate a set of driving intentions, targeting a classification task based on multimodal physiological signals. We propose a model that incorporates hypergraph convolution for feature extraction. Experimental results demonstrate that our approach outperforms the baselines in detecting various types of driving intentions with 74.40% accuracy, 5.89% higher than the best baseline. This research contributes to both traffic safety and brain-computer interfaces. Wenqi Zhang 0003, Yanjun Qin, Xiaoming Tao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Vision-Aided Ultra-Reliable Low-Latency Communications for Smart FactoryabstractSmart factory is a new digital and intelligent platform requiring high throughput and ultra-reliable low-latency communications (URLLC). Industrial communications at sub-6 GHz faces spectrum congestion and bandwidth limitations, which seriously jeopardize the high data rate requirement of smart factory. Recently, millimeter wave (mmWave) and Terahertz technologies have become enablers for high speed communications and intelligent manufacturing in Industry 4.0 and beyond. However, the sensitivity of mmWave signals to blockage and the overhead of large-scale antenna beam sweeping pose serious challenges to the reliability and the latency of wireless networks in these frequency ranges. In this paper, we propose a vision-aided URLLC framework for smart factory that does not incur any overhead from channel training and beam sweeping. In particular, we design a feature extraction method to obtain communications-related features in environmental images for blockage prediction, reference signal receiving power (RSRP) prediction, and beam selection. Then, we construct a joint image-channel dataset covering images, annotations, blockage, and wireless channels based on Blender and Wireless Insite software. Simulations show that the accuracy of blockage prediction 400 ms ahead reaches 99.9%, the root mean square error (RMSE) of RSRP prediction 400 ms ahead reaches 2.78 dB, and the Top-5 accuracy of beam selection reaches 91.8%. Blockage and RSRP prediction can assist base station (BS) handover to avoid communications interruption, while beam selection can eliminate the overhead of channel training and beam sweeping. Hence, the proposed study provides a promising direction for enabling URLLC under mmWave and even Terahertz bands in smart factory of Industry 4.0. Feifei Gao 0001, Xiaoming Tao 0001, Shaodan Ma, H. Vincent Poor |
IEEE Trans. Commun. | 3 |
| 2024 | A Unified Multi-Task Semantic Communication System for Multimodal DataabstractTask-oriented semantic communications have achieved significant performance gains. However, the employed deep neural networks in semantic communications have to be updated when the task is changed or multiple models need to be stored for performing different tasks. To address this issue, we develop a unified deep learning-enabled semantic communication system (U-DeepSC), where a unified end-to-end framework can serve many different tasks with multiple modalities of data. As the number of required features varies from task to task, we propose a vector-wise dynamic scheme that can adjust the number of transmitted symbols for different tasks. Moreover, our dynamic scheme can also adaptively adjust the number of transmitted features under different channel conditions to optimize the transmission efficiency. Particularly, we devise a lightweight feature selection module (FSM) to evaluate the importance of feature vectors, which can hierarchically drop redundant feature vectors and significantly accelerate the inference. To reduce the transmission overhead, we then design a unified codebook for feature representation to serve multiple tasks, where only the indices of these task-specific features in the codebook are transmitted. According to the simulation results, the proposed U-DeepSC achieves comparable performance to the task-oriented semantic communication system designed for a specific task but with significant reduction in both transmission overhead and model size. Guangyi Zhang 0005, Qiyu Hu, Zhijin Qin, Yunlong Cai, Guanding Yu, Xiaoming Tao 0001 |
IEEE Trans. Commun. | 6 |
| 2024 | Optical Flow-Based Spatiotemporal Sketch for Video Representation: A Novel FrameworkabstractWith the rapid development of multimedia services and the dramatic growth of video data volume, efficient video representation and AI-generated content (AIGC) become critical parts of future multimedia communication systems. Sketch graph is a structured abstraction of key textures in an image, and video sketch graph further exploits the temporal continuity of videos to achieve a sparse representation. Sketch-based representation has potential applications in communication systems for both human subjective perception and machine vision tasks, and provides a new idea for AIGC. However, current video sketch extraction methods rely on human assistance and correction, and cannot be applied to end-to-end communication systems. We design a novel framework for spatiotemporal sketch extraction based on deep learning methods. In the proposed framework, sketch extraction and sparse coding are performed at the sender side using structural and temporal features of the video. The original videos are generatively reconstructed at the receiver side or applied to downstream machine vision tasks. We validate the performance of the proposed method on Cityscapes dataset with different metrics. Experiments show that our proposed framework can be end-to-end adapted to video communication tasks in different scenarios and can achieve efficient video characterization and transmission. Moreover, our proposed method enables sketch-based end-to-end AIGC for video generation. Qiyuan Du, Yiping Duan, Zhipeng Xie, Xiaoming Tao 0001, Linsu Shi, Zhijuan Jin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | OARNet: Object-Attribute-Relation Network for Predicting Soccer EventsabstractEvent prediction involves analyzing and forecasting events that occur at a specific time and location to inform decision-making and take the next actions. Current event prediction approaches primarily employ deep learning methods to analyze regular patterns from large amounts of historical data. However, predicting adversarial soccer events remains a significant challenge due to strong antagonism and complex relationships between players. With this consideration, we propose an objectattribute-relation (OAR) network for predicting soccer events using multimodal data, including spatiotemporal trajectory data and video data. The proposed scheme aims to enhance prediction performance by transforming multimodal data into an OAR space that integrates global and local relationships (adversarial information and multi-objective information). In particular, the scheme consists mainly of a relation module, an object attribute module, and a graph prediction module. We first use ConvLSTM to extract the visual features of players from video data and use LSTM to extract the movement features of players from spatiotemporal data. Additionally, we apply a multihead GRU attention mechanism to calculate the relation weights. These three components are then combined into an OAR graph of a clip in a soccer game. Finally, an OAR GNN is designed to determine the influence of different objects and predict events. The entire process constitutes an end-to-end event prediction learning framework. Extensive experimental results on the two challenging datasets, namely, soccER and SkillCorner, verify the effectiveness of the proposed framework. Yiping Duan, Xiaoming Tao 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2024 | Boosting Scene Graph Generation with Contextual InformationabstractScene graph generation (SGG) has been developed to detect objects and their relationships from the visual data and has attracted increasing attention in recent years. Existing works have focused on extracting object context for SGG. However, very few works have attempted to exploit implicit contextual correlations among relationships of the objects. Furthermore, most existing SGG schemes rely on high-level features to predict the predicates while overlooking the potential inherent association of low-level features with the object relationships. We present in this article a novel scheme to capture enhanced contextual information for both objects and relationships. We design a Dual-branch Context Analysis Transformer (DCAT) architecture to extract both object context and relationship context from the visual data with dual transformer branches and then effectively fuse both high-level and low-level features by an adaptive approach to facilitate relationship prediction. Specifically, we first conduct feature representation learning to enrich relation representations by the visual, spatial, and linguistic feature extractors. Next, two transformer branches are designed to leverage the modeling of global associative interaction and mine the hidden association among objects and relationships. Then, we devise a novel feature disentangling method to decouple contextualized high-level features with guidance from the visual semantics. Finally, we develop a refined attention module to perform low-level feature recalibration for the refinement of the final predicate prediction. Experiments on Visual Genome and Action Genome datasets demonstrate the effectiveness of DCAT for both image and video SGG settings. Moreover, we also test the quality of the generated image scene graphs to verify the generalizability on downstream tasks like sentence-to-graph retrieval and image retrieval. Shiqi Sun 0002, Danlan Huang, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Meta Federated Reinforcement Learning for Distributed Resource AllocationabstractIn cellular networks, resource allocation is usually performed in a centralized way, which brings huge computation complexity to the base station (BS) and high transmission overhead. This paper introduces a distributed resource allocation method that aims to maximize energy efficiency (EE) while ensuring quality of service (QoS) for users. Specifically, to address the challenge of fast-varying wireless channel conditions, we propose a robust meta federated reinforcement learning (MFRL) framework that enables local users to optimize transmit power and assign channels using locally trained neural network models. This approach offloads the computational burden from the cloud server to the local users, reducing transmission overhead associated with local channel state information. The BS performs the meta-learning procedure to initialize a general global model, enabling rapid adaptation to different environments and improved EE performance. The federated learning technique, based on decentralized reinforcement learning, promotes collaboration and mutual benefits among users. Analysis and numerical results demonstrate that the proposedMFRLframework accelerates the reinforcement learning process, decreases transmission overhead, and offloads computation, while outperforming the conventional decentralized reinforcement learning algorithm in terms of convergence speed and EE performance across various scenarios. Zelin Ji, Zhijin Qin, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Resource Optimization for Semantic-Aware Networks With Task OffloadingabstractThe limited capabilities of user equipment restrict the local implementation of computation-intensive applications. Edge computing, especially the edge intelligence system, enables local users to offload the computation tasks to the edge servers to reduce the computational energy consumption of user equipment and accelerate fast task execution. However, the limited bandwidth of upstream channels may increase the task transmission latency and affect the computation offloading performance. To overcome the challenge arising from scarce wireless communication resources, we propose a semantic-aware multi-modal task offloading system that facilitates the extraction and offloading of semantic task information to edge servers. To cope with the different tasks with multi-modal data, a unified quality of experience (QoE) criterion is designed. Furthermore, a proximal policy optimization-based multi-agent reinforcement learning algorithm (MAPPO) is proposed to coordinate the resource management for wireless communications and computation in a distributed and low computational complexity manner. Simulation results verify that the proposed MAPPO algorithm outperforms other reinforcement learning algorithms and fixed schemes in terms of task execution speed and the overall system QoE. Zelin Ji, Zhijin Qin, Xiaoming Tao 0001, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | A Robust Semantic Text Communication SystemabstractSemantic communication is increasingly viewed as a promising solution to improve the transmission efficiency. However, semantic communications are susceptible not only to physical channel impairments, but also to semantic impairments, which degrade semantic understanding at the receiver and disrupt the associated downstream tasks. Hence, we focus our attention on the robustness of semantic communications against semantic impairments. Specifically, we first categorize textual semantic impairments into three categories based on their sources. Then, we propose a robust deep learning enabled semantic communication system (R-DeepSC) by introducing a semantic corrector for robust semantic encoding so as to facilitate semantic transmission. Moreover, we develop a non-autoregressive version of R-DeepSC, namely NA-RDeepSC, which offers improved inference speed by relying on a non-autoregressive architecture and an adaptive generator embedded into the semantic decoder. NA-RDeepSC performs semantic decoding in parallel, hence reducing the decoding complexity fromO(n) toO(1) with a comparable performance to that of R-DeepSC. Our experimental results demonstrate the superior robustness of the proposed R-DeepSC and NA-RDeepSC architectures in eliminating semantic impairments, hence highlighting the significance of this work in advancing the development of robust semantic communications. Zhijin Qin, Xiaoming Tao 0001, Jianhua Lu, Lajos Hanzo |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Task-Oriented Scene Graph-Based Semantic Communications With Adaptive Channel CodingabstractSemantic communications have shown great potential in reducing the transmitted data amount through the powerful capability to extract and transmit essential semantic information. Although existing works have achieved certain transmission efficiency, challenges such as efficiently interpretable semantic extraction, and dynamically adaptive channel coding have not been fully explored. In this paper, we tackle these challenges by proposing a task-oriented scene graph-based semantic communication system with adaptive channel coding, named GRACE, to perform image retrieval task. To enhance the interpretability of semantic communications and reduce semantic redundancy, we introduce a scene graph semantic encoder. This encoder fully exploits informative scene graph semantics, effectively extracting scene graphs and performing further semantic coding. Additionally, to handle variable channel conditions in real-world scenarios, we develop a semantic-aware adaptive channel coding to adapt to channel conditions and reduce the communication resources. At the receiver, the image retrieval task is accomplished based on the recovered scene graph semantics. The experimental results demonstrate the superiority of the proposed system compared to other communication systems in terms of task-execution performance, robustness against channel variations, transmission efficiency, and computational complexity. Shiqi Sun 0002, Zhijin Qin, Huiqiang Xie, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Semantic MIMO Systems for Speech-to-Text TransmissionabstractSemantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the single-user multiple-input multiple-output (MIMO) and multi-user MIMO communication scenarios, named SAC-ST. Particularly, a semantic communication system to serve the speech-to-text task at the receiver is first designed, which compresses the semantic information and generates the low-dimensional semantic features by leveraging the transformer module. In addition, a novel semantic-aware network is proposed to facilitate transmission with high semantic fidelity by identifying the critical semantic information and guaranteeing its accurate recovery. Furthermore, we extend the SAC-ST with a neural network-enabled channel estimation network to mitigate the dependence on accurate channel state information and validate the feasibility of SAC-ST in practical communication environments. Simulation results will show that the proposed SAC-ST outperforms the communication framework without the semantic-aware network for speech-to-text transmission over the MIMO channels in terms of the speech-to-text metrics, especially in the low signal-to-noise regime. Moreover, the SAC-ST with the developed channel estimation network is comparable to the SAC-ST with perfect channel state information. Zhenzi Weng, Zhijin Qin, Huiqiang Xie, Xiaoming Tao 0001, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | QoE-Based Semantic-Aware Resource Allocation for Multi-Task NetworksabstractBy transmitting task-related information only, semantic communications yield significant performance gains over conventional communications. However, the lack of mature semantic theory about semantic information quantification and performance evaluation makes it challenging to perform resource allocation for semantic communications, especially when multiple tasks coexist in the network. To cope with this challenge, we propose a quality-of-experience (QoE) based semantic-aware resource allocation method for multi-task networks in this paper. First, semantic entropy is defined to quantify the semantic information for different tasks, and the relationship between semantic entropy and Shannon entropy is analyzed. Then, we develop a novel QoE model to formulate the semantic-aware resource allocation in terms of semantic compression, channel assignment, and transmit power. The compatibility of the formulated problem with conventional communications is further demonstrated. To solve this problem, we decouple it into two subproblems and solved them by a developed deep Q-network (DQN) based method and a proposed low-complexity matching algorithm, respectively. Finally, simulation results validate the effectiveness and superiority of the proposed method, as well as its compatibility with conventional communications. Lei Yan 0001, Zhijin Qin, Chunfeng Li, Rui Zhang 0026, Yongzhao Li, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2023 | Visual Attention Measurement Based on Electroencephalogram Feature LearningabstractWith the continuous development of multimedia technology, there is a growing demand for video applications. Among the huge number of videos, some video clips attract higher visual attention from viewers. Measuring this visual attention is not only crucial for evaluating the quality of experience (QoE) but also holds the potential for guiding video compression techniques. Therefore, there is an urgent need to propose an effective method to evaluate the changes of viewers' attention. To address this challenge, we employ electroencephalography (EEG) as a bridge to establish the relationship between video clips and visual attention. We introduce a visual attention measurement (VAM) method by combining EEG and machine learning approaches. Specifically, we design an EEG experiment to collect brain responses from viewers while watching various video clips, thus constructing a visual attention EEG dataset. Using these EEG signals, we propose a VAM network to identify whether viewers pay attention to the video clips. The experimental results show that our EEG experiment can capture the physiological signals related to visual attention. Moreover, the proposed VAM network can accurately determine viewers' attention levels to video clips. These results provide new insights into the evolution of human-oriented video communications. Jian Chu, Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001 |
GLOBECOM | 4 |
| 2023 | Task-Oriented Explainable Semantic Communications Based on Structured Scene GraphsabstractSemantic communications have been regarded as a promising solution for the next generation communication systems to alleviate the spectral resource shortage and the network congestion. Existing image semantic communication systems extract and transmit global semantics, which can cope with the downstream data reconstruction or intelligent tasks. However, the semantics involved in these methods are still severely redundant and uninterpretable. In this work, we propose a novel task-oriented semantic communication framework based on scene graph, named DeepSC-SG. Specifically, we first devise a scene graph based semantic encoder, which extracts the explainable scene graph semantics from the input images and encodes the semantics into informative graph embeddings. Then we design a joint source-channel (JSC) codec to combat physical channel impairment. After receiving the semantics, a semantic decoder is devised to achieve the downstream image retrieval task by computing scene graph similarities. Simulation results demonstrate that the proposed DeepSC-SG is fairly robust to the channel variations compared to the traditional communication systems, which has great potential in realizing downstream intelligent tasks like image retrieval with significantly reduced size of transmitted data. Shiqi Sun 0002, Zhijin Qin, Huiqiang Xie, Xiaoming Tao 0001 |
GLOBECOM | 4 |
| 2023 | Sketch Graph Representation for Multimedia Computational Communications: A Learning-Based MethodabstractMultimedia computational communications towards 6G can improve the transmission efficiency significantly by introducing intelligent computation in the communication process. This intelligent/smart communication architecture includes multimedia representation, coding, transmission and other parts from the perspective of semantics, where multimedia semantic representation is the core part and is mainly utilized to reduce the amount of multimedia data. In this paper, sketch graph is proposed as an effective representation of images to describe the pixel variations, geometric feature distribution and structural information and has potential applications in multimedia computational communications. Specifically, we developed a learning-based method to extract sketch graphs with edge detection, sketch point detection and sketch line detection by deep neural networks (DNNs). Moreover, we designed an end-to-end extraction method and achieved real-time processing. The experimental results on several datasets demonstrated the advanced performance in terms of classification and generation tasks. Image classification results on the HumanSketch, ImageNet, and Caltech datasets showed that sketch graphs extracted by our method had better describing ability than those extracted by traditional methods and other traditional image compression methods. On the other hand, image generation results on the Cityscapes dataset indicated the potential of the sketch-graph-based image compression codec. Qiyuan Du, Yiping Duan, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001 |
ICC | 3 |
| 2023 | Mem-DeepSC: A Semantic Communication System with MemoryabstractWhile semantic communications succeed in effectively transmitting due to the strong capability to extract the essential semantic information, it is still far from intelligent communications. In this paper, we introduce an essential component, memory, into semantic communications to mimic human communications. Particularly, we propose a deep learning (DL) based semantic communication system with memory, named Mem-DeepSC, by considering the scenario question answer as the task at the receiver. We exploit universal Transformer based transceiver to extract the semantic information and introduce the memory module to enhance the semantic decoding capability at the receiver. Moreover, we derive the semantic channel capacity and propose a consecutive dynamic transmission method to minimize the transmission latency. Numerical results show that Mem-DeepSC is superior to benchmarks in terms of answer accuracy and the number of transmitted symbols. Zhijin Qin, Huiqiang Xie, Xiaoming Tao 0001 |
ICC | 3 |
| 2023 | USGG: Union Message Based Scene Graph GenerationabstractScene graph generation (SGG) is designed to represent images by objects and their relationships. Existing works mainly attempt to strengthen object pair representations for SGG. However, most methods ignore the significant semantic information implied in union regions, which refers to the surrounding area of object pairs. In this paper, we propose a new union message based architecture, named as USGG, to profoundly exploit the relational semantics of unions to facilitate SGG. Concretely, we employ sufficient feature extraction to enhance the features of objects and unions. Next, we devise the Union Embedding Network to model the relational representations through two symmetric encoder-decoder branches. Moreover, the Union Fusion Network is designed to integrate the refined semantics by two-stage feature fusion. Extensive experiments are conducted on Visual Genome dataset, which demonstrates that the proposed approach achieves competitive performance against state-of-the-art methods on Recall, mean Recall and Zero Shot Recall metrics. Shiqi Sun 0002, Danlan Huang, Zhijin Qin, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001 |
ICIP | 4 |
| 2023 | A Hybrid Approach for Driving Behavior Recognition: Integration of CNN and Transformer-Encoder with EEG dataabstractHuman factors are considered as one of the main causes affecting road traffic safety. Therefore, it is highly necessary to establish a driving behavior model for predicting driver behaviors and states, which can be used for risk monitoring. As an objective indicator that can accurately measure human cognitive and emotional states, electroencephalography (EEG) has attracted widespread attention from researchers due to its suitability and high temporal resolution for driver state and behavior measurement. Our study proposes a novel EEG-based driving behavior recognition algorithm that combines CNN and Transformer-Encoder modules. We deploy CNN to extract spatial and temporal features from EEG signals, capturing local dependencies between different time points. Subsequently, the encoder module of Transformer is utilized to handle long-term dependencies, enhancing the correlation between different time steps and improving classification accuracy. Experimental results demonstrate that our proposed model outperforms the baseline model on a self-constructed driving behavior classification dataset, providing a foundation for further in-depth research into real-world road traffic safety. Yanjun Qin, Xiaoming Tao 0001 |
VTC Fall | 5 |
| 2023 | Dual-Transformer: A General Model for Traffic Accident PredictionabstractTraffic accident prediction aims to estimate the future traffic accident risk for improving urban public safety. Extensive efforts have been dedicated to leveraging spatial correlations and temporal dependencies in the evolving patterns of traffic accidents which are achieved by utilizing explicit spatial relations, such as geographical proximity, through predefined geographical structures. However, current methods for predicting traffic accident still neglect the implicit spatial correlations influenced by an area’s inherent geographic characteristics. To tackle this issue, we propose a Dual-Transformer model that automatically explores implicit spatial correlations, including a Spatial Transformer and a Temporal Transformer. The Spatial Transformer autonomously learns the implicit spatial correlations among road segments, extending beyond the boundaries of geographical structures. Meanwhile, the Temporal Transformer aims to capture the dynamic patterns of these implicit spatial correlations. Through extensive experiments conducted on two real-world traffic accident datasets, our proposed model demonstrates significant improvements compared to existing methods. Dongkun Wang, Jieyang Peng, Junkai Zhao, Yunfei Teng, Wenjing Xue, Xiaoming Tao 0001 |
VTC Fall | 6 |
| 2023 | Task-Oriented Semantic Communications for Speech TransmissionabstractSemantic communications execute intelligent tasks at the receiver by only transmitting necessary information. In this paper, we introduce TOS-ST, a task-oriented semantic communication system for speech transmission, which efficiently serves the semantic tasks at the receiver, including speech-to-text translation and speech-to-speech translation. Particularly, TOS-ST condenses the input speech in the source language and extracts the task-related semantics features prior to transmission. At the receiver, these features are recovered and utilized by the neural network-based semantic preserver and machine translation module to generate the uncorrupted text in the target language. To perform the speech-to-speech translation task, the translated text passes through a sophisticated neural network to obtain speech in the target language. According to the simulation results, the TOS-ST outperforms conventional speech transmission systems and exhibits higher robustness against channel impairment. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001 |
VTC Fall | 3 |
| 2023 | Video Reconstruction with Multimodal InformationabstractVideo reconstruction refers to generate videos through the high-level representations (edge map, labels and so on), while the reconstruction quality is always unsatisfactory due to sparse high-level representations, especially on video data. In order to improve the video reconstruction quality, we proposed a novel approach that generates realistic video from its multimodal information including structure features and color features. To extract color features, we mainly apply the k-means algorithm to segment labels and the structure features are extracted by an edge detection network. Video generation is regarded as learning the mapping from multimodal representations to the original videos. So, a conditional GAN is applied with a learning objective that models the temporal video dynamics. We use a spatio-temporal generator with attention to model the inter-frame dynamics and video consistency is improved in this way. Moreover, we use a multiscale discriminator to improve the improve the intra-frame quality of the video. Experimental results on Cityscapes, Apolloscape datasets demonstrate that our proposed approach performs better in both traditional and generative evaluating indicators. Zhipeng Xie, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Jiazhong Yu |
VTC Fall | 4 |
| 2023 | Adjustable Dielectric Resonator Antenna With Parasitic Elements for 5G SAGOI-IoT ApplicationsabstractWith the development of modern communication and the Internet of Things (IoT) in the need to provide seamless interconnection between heterogeneous devices, we design an adjustable-distance resonant antenna suitable for 5G communication. The antenna meets the multiangle radiation requirements of the antenna for flexible positioning of IoT devices and conforms to the concept of green Internet basic hardware, with low energy and small size. In our design, three parasitic elements will couple with the higher order modes through the slot-hole excitation of a higher order mode dielectric resonator antenna with a dielectric constant of 10. By controlling the distance between the three parasitic elements and changing the capacitor at their terminals, radiation variation in multiple directions can be achieved. The proposed model focuses on the relationships among the three element distances and the effect of the third parasitic element on the radiation angle. Good results were obtained for gain, bandwidth, and radiation angle. Through simulation, the dielectric resonator antenna works successfully in the 15-GHz frequency band. Thus, the antenna array can be from −34° to 34° in the horizontal direction and from 0° to −36° in the vertical direction. At the same time, the gains are kept at a good level and the bandwidths are greater than 2 GHz. Compared with other dielectric resonant antennas, the parameters of this antenna are not reduced, and the flexibility of the radiation direction angles is increased. These evaluation parameters are considered ideal conditions for device-to-device communication in 5G IoT applications. Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Spatio-temporal hierarchical MLP network for traffic forecasting
Yanjun Qin, Haiyong Luo, Fang Zhao 0003, Yuchen Fang 0001, Xiaoming Tao 0001, Chenxing Wang 0001 |
Inf. Sci. | 5 |
| 2023 | Toward Semantic Communications: Deep Learning-Based Image Semantic CodingabstractSemantic communications has received growing interest since it can remarkably reduce the amount of data to be transmitted without missing critical information. Most existing works explore the semantic encoding and transmission for text and apply techniques in Natural Language Processing (NLP) to interpret the meaning of the text. In this paper, we conceive the semantic communications for image data that is much more richer in semantics and bandwidth sensitive. We propose an reinforcement learning based adaptive semantic coding (RL-ASC) approach that encodes images beyond pixel level. Firstly, we define the semantic concept of image data that includes the category, spatial arrangement, and visual feature as the representation unit, and propose a convolutional semantic encoder to extract semantic concepts. Secondly, we propose the image reconstruction criterion that evolves from the traditional pixel similarity to semantic similarity and perceptual performance. Thirdly, we design a novel RL-based semantic bit allocation model, whose reward is the increase in rate-semantic-perceptual performance after encoding a certain semantic concept with adaptive quantization level. Thus, the task-related information is preserved and reconstructed properly while less important data is discarded. Finally, we propose the Generative Adversarial Nets (GANs) based semantic decoder that fuses both locally and globally features via an attention module. Experimental results demonstrate that the proposed RL-ASC is noise robust and could reconstruct visually pleasant and semantic consistent image in low bit rate condition. Danlan Huang, Feifei Gao 0001, Xiaoming Tao 0001, Qiyuan Du, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Environment Semantics Aided Wireless Communications: A Case Study of mmWave Beam Prediction and Blockage PredictionabstractIn this paper, we propose an environment semantics aided wireless communication framework to reduce the transmission latency and improve the transmission reliability, where semantic information is extracted from environment image data, selectively encoded based on its task-relevance, and then fused to make decisions for channel related tasks. As a case study, we develop an environment semantics aidednetwork architecturefor mmWave communication systems, which is composed of a semantic feature extraction network, a feature selection algorithm, a task-oriented encoder, and a decision network. With images taken from street cameras and user’s identification information as the inputs, the environment semantics aided network architecture is trained to predict the optimal beam index and the blockage state for the base station. It is seen that without pilot training or costly beam scans, the environment semantics aided network architecture can realize extremely efficient beam prediction and timely blockage prediction, thus meeting requirements for ultra-reliable and low-latency communications (URLLCs). Simulation results demonstrate that compared with existing works, the proposed environment semantics aided network architecture can reduce system overheads such as storage space and computational cost while achieving satisfactory prediction accuracy and protecting user privacy. Yuwen Yang, Feifei Gao 0001, Xiaoming Tao 0001, Guangyi Liu 0001, Chengkang Pan |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Sketch Assisted Face Image Coding for Human and Machine Vision: A Joint Training ApproachabstractImage coding is one of the most fundamental techniques and is widely used in image/video processing and multimedia communications. Current image coding methods are mainly human-oriented, and the visual quality is always unsatisfactory, especially at low bitrates. Moreover, the recent emergence of machine vision goes beyond the scope of current coding. With these considerations, we proposed a sketch assisted face image coding for human and machine vision by a joint training approach. In the proposed approach, we design a new feature representation: a color sketch, which aims to satisfy both low-frequency features of human vision and high-frequency features of machine analysis. Then, we present a novel end-to-end image codec framework with joint training that consists of three models: an image-to-image translation module, a coding module, and a two-stage reconstruction module. Specifically, the input image is first translated into the edge map with the Canny edge as the auxiliary label to merely preserve the structure information. Afterward, the backpropagation from reconstruction module guides the edge map to increase or decrease the information through joint training, which results in the generation of color sketch. Then, the generated sketch is compressed into the bitstream and decompressed back to a sketch in the coding module. Finally, the decompressed sketch is reconstructed to support the machine and human tasks, respectively. In this way, the color sketch is designed to bridge the gap between human and machine vision, and the joint training strategy helps to adjust the low-frequency information in the sketch. The experimental results on challenge datasets demonstrate that our proposed algorithm offers 40.9%-86.6% bitrate savings on machine vision and is comparable to state-of-the-art image coding methods on human vision. Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Audio Related Quality of Experience Evaluation in Urban Transportation Environments With Brain Inspired Graph LearningabstractThe fast advancement of urban transportation systems in the recent decades has on one hand improved efficiency in traffic control and management, yet on the other hand brought new obstacles and interferences in audio related services in transportation systems, which is one of the dominating components in urban transportation systems, such as end-to-end Voice over Internet Protocol (VoIP) communications, risk alerting, and personalised recommendation services. The movement of vehicles/trains and the growing complexity of transportation infrastructures has become a big threat to the audio related services. Hence it is crucial to evaluate the Quality of Experience (QoE) of audio related services. Different from traditional algorithms which use digital signal processing to evaluate the QoE of mobile users, in this paper, we propose a two-stage brain-alike neural network aided graph learning algorithm to evaluate the QoE of audio signals with the aid of EEG feature extraction. The results are evaluated by newly-collected on-site data in public transportation environments and are examined by a branch of human experts to show that our algorithm outperforms other benchmark algorithms in term of human perception and accuracy of classification. Wen-Long Shang, Xiaoming Tao 0001, Huibo Bi, Washington Yotto Ochieng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Computer Vision Aided Codebook Design for MIMO Communications SystemsabstractmmWave communications systems usually rely on analog or hybrid analog/digital architectures and thus need a predefined codebook to perform beamforming. Traditional codebooks are designed for universal environments, although in practice a particular BS will only serve a particular environment. In this paper, we propose novel site-specific codebook design methods by utilizing the visual information captured through cameras. Different from other site-specific codebook design methods that require a large amount of measured channel state information (CSI), the proposed ones need only a simple snapshot of the environment followed by efficient computer vision (CV) techniques. Thus the proposed CV-aided codebook design reduces the overhead of communications system, such as the cost of time, human resources, as well as the hardware installation and calibration. Specifically, we propose a CV-based approach that detects the LOS area around the BS and reconstructs the LOS channel vectors set (CVS). With this knowledge, we build a vision-based beam codebook using Lloyd algorithm. Further, we design a FusionNet to generate the codebook that can serve the non-line-of-sight (NLOS) users. The simulation results demonstrate the effectiveness of the proposed CV-aided codebook design methods and their superiority compared to the conventional methods. Feifei Gao 0001, Xiaoming Tao 0001, Guangyi Liu 0001, Chengkang Pan, Ahmed Alkhateeb |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Deep Learning Enabled Semantic Communications With Speech Recognition and SynthesisabstractIn this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication system, respectively. First, the speech recognition-related semantic features are extracted for transmission by a joint semantic-channel encoder and the text is recovered at the receiver based on the received semantic features, which significantly reduces the required amount of data transmission without performance degradation. Then, we perform speech synthesis at the receiver, which dedicates to re-generate the speech signals by feeding the recognized text and the speaker information into a neural network module. To enable the DeepSC-ST adaptive to dynamic channel environments, we identify a robust model to cope with different channel conditions. According to the simulation results, the proposed DeepSC-ST significantly outperforms conventional communication systems and existing DL-enabled communication systems, especially in the low signal-to-noise ratio (SNR) regime. A software demonstration is further developed as a proof-of-concept of the DeepSC-ST. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Computer Vision Aided mmWave Beam Alignment in V2X CommunicationsabstractVisual information, captured for example by cameras, can effectively reflect the sizes and locations of the environmental scattering objects, and thereby can be used to infer communications parameters like propagation directions, receiver powers, as well as the blockage status. In this paper, we propose a novel beam alignment framework that leverages images taken by cameras installed at the mobile user. Specifically, we utilize 3D object detection techniques to extract the size and location information of the dynamic vehicles around the mobile user, and design a deep neural network (DNN) to infer the optimal beam pair for transceivers without any pilot signal overhead. Moreover, to avoid performing beam alignment too frequently or too slowly, a beam coherence time (BCT) prediction method is developed based on the vision information. This can effectively improve the transmission rate compared with the beam alignment approach with the fixed BCT. Simulation results show that the proposed vision based beam alignment methods outperform the existing LIDAR and vision based solutions, and demand for much lower hardware cost and communication overhead. Weihua Xu 0001, Feifei Gao 0001, Xiaoming Tao 0001, Jianhua Zhang 0001, Ahmed Alkhateeb |
IEEE Trans. Wirel. Commun. | 3 |
| 2022 | Driver Vigilance Detection from EEG Signals using Transformer NetworksabstractSafety driver is one of the most important safety precautions to assure road safety in public road tests of autonomous vehicles (AVs), However, the decreasing vigilance of the safety driver has become a major cause of autonomous vehicle accidents in recent years. With the introduction of wireless, wearable Electroencephalography (EEG) devices and the enhanced computational capabilities of AVs, the ubiquitous monitoring of safety drivers as part of the emerging E-Health system has recently attracted great interest from researchers and industry. Unfortunately, we know little about the electrophysio-logical signals that assess safety driver vigilance. In this work, we propose a method for detecting driver vigilance based on EEG signals that combines frequency domain EEG features and a transformer model to maximize the prediction accuracy and recall rate. Experimental results show that trichotomous results based on the proposed model outperform previous dichotomous results, reaching 78% when tested with data from the Sustained-Attention Driving dataset published in Nature scientific data. Dingcheng Gao, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 3 |
| 2022 | A Robust Deep Learning Enabled Semantic Communication System for TextabstractWith the advent of the 6G era, the concept of semantic communication has attracted increasing attention. Compared with conventional communication systems, semantic communication systems are not only affected by physical noise existing in the wireless communication environment, e.g., additional white Gaussian noise, but also by semantic noise due to the source and the nature of deep learning-based systems. In this paper, we elaborate on the mechanism of semantic noise. In particular, we categorize semantic noise into two categories: literal semantic noise and adversarial semantic noise. The former is caused by written errors or expression ambiguity, while the latter is caused by perturbations or attacks added to the embedding layer via the semantic channel. To prevent semantic noise from influencing semantic communication systems, we present a robust deep learning enabled semantic communication system (R-DeepSC) that leverages a calibrated self-attention mechanism and adversarial training to tackle semantic noise. Compared with baseline models that only consider physical noise for text transmission, the proposed R-DeepSC achieves remarkable performance in dealing with semantic noise under different signal-to-noise ratios. Zhijin Qin, Danlan Huang, Xiaoming Tao 0001, Jianhua Lu, Guangyi Liu 0001, Chengkang Pan |
GLOBECOM | 4 |
| 2022 | Measuring Human Perception of Audiovisual Errors using EEGabstractAudiovisual synchronization is an essential indicator of video quality. The degradation in the quality of experience caused by such synchronization errors is often measured using the mean opinion score (MOS). However, this method is susceptible to emotion and bias. Electroencephalography (EEG), as an objective tool, overcomes the drawbacks of subjective testing for evaluating human perception. In this work, we measure the human perception of audio and video synchronization based on the EEG approach through a series of experiments. A scoring model was developed for audio and video under multilevel synchronization errors by extracting the power spectral density (PSD) of EEG signals as features. The use of EEG signals for perceptual ability assessment of audiovisual distortion provides a potential neurally informed approach that is more objective than the high-level cognitive activity in subjective data. Dingcheng Gao, Bingrui Geng, Yiping Duan, Xiaoming Tao 0001, Chengkang Pan |
VTC Fall | 4 |
| 2022 | Image Generation from Scene Graph with Object EdgesabstractSignificant progress has been made on methods for generating images from structured semantic descriptions, but the generated images only retain semantic information, and the appearance of objects cannot be constrained and effectively represented. Therefore, we propose a scene graph structure image generation method assisted by object edge information. Our model uses two graph convolution neural networks(GCN) to process scene graphs and obtains object features as well as relation features which aggregate related information. The object bounding boxes are predicted by a method a decoupling the size and position. Where auxiliary models are added to coordinate with segmentation mask network training. Our experiments show that the introduction of object edges provides clearer object appearance information for image generation, which can constrain object shapes and improve image quality greatly. Finally, the cascaded refinement network is used to generate images. Additionally, compared with other appearance features, such as object slices, edge information occupies a smaller quantity of data, which greatly improves the image quality with less increase in the input information. This feature also benefits semantic communication systems. A large number of experiments show that our method is significantly superior to the latest Sg2im method when evaluated on Visual Genome datasets. Chenxing Li, Yiping Duan, Qiyuan Du, Chengkang Pan, Guangyi Liu 0001, Xiaoming Tao 0001 |
VTC Fall | 6 |
| 2022 | Multi-Relational Pedestrian Trajectory Prediction in Complex ScenesabstractPedestrian trajectory prediction has an important impact on the construction of smart cities and the popularization of autonomous vehicles. Pedestrian trajectory prediction in complex scenes is challenging because the trajectories are largely disturbed by the surrounding social environment. To better model the relationship between pedestrians and the social environment, we propose a novel Multi-Relation Network based on Long Short-Term Memory(LSTM). We first use Convolution Neural Network(CNN) and LSTM for feature extraction of scenes and pedestrians respectively. We realize the interaction between pedestrians and social environment by introducing attention mechanism. The experiment results on two public datasets, i.e. ETH and UCY, demonstrate that the prediction accuracy can be effectively improved by considering the social interactions. Wenshuo Peng, Zhoujuan Cui, Yiping Duan, Xiaoming Tao 0001 |
VTC Fall | 4 |
| 2022 | Path-based Multimodal Trajectories PredictionabstractPredicting the future movement of other traffic participants in complex traffic is one of the key tasks to realizing automatic driving. However, it is a challenge to accurately predict the future movement of the target, because the driving behavior is inherently random and multimodal, and the behavior of the agent is also determined by the complex interaction of other agents. The target-based method has a good performance in generating multi-modal trajectories. In this paper, a path-based multimodal trajectories prediction method is proposed, which takes the path as the target and performs two tasks path classification and trajectory regression. In this method, different paths will produce different embedding, and the sparse attention method based on path is adopted to deal with the interaction between agents, which reduces the complexity of the model and the risk of mode collapse. Experiments on the Argoverse dataset show that the proposed method has achieved good performance in all metrics, and its outstanding performance in the MR proves that the method can ensure the multimodality of output. Yiping Duan, Xiaoming Tao 0001 |
VTC Fall | 3 |
| 2022 | Task-Oriented Multi-User Semantic CommunicationsabstractWhile semantic communications have shown the potential in the case of single-modal single-users, its applications to the multi-user scenario remain limited. In this paper, we investigate deep learning (DL) based multi-user semantic communication systems for transmitting single-modal data and multimodal data, respectively. We adopt three intelligent tasks, including, image retrieval, machine translation, and visual question answering (VQA) as the transmission goal of semantic communication systems. We propose a Transformer based framework to unify the structure of transmitters for different tasks. For the single-modal multi-user system, we propose two Transformer based models, named, DeepSC-IR and DeepSC-MT, to perform image retrieval and machine translation, respectively. In this case, DeepSC-IR is trained to optimize the distance in embedding space between images and DeepSC-MT is trained to minimize the semantic errors by recovering the semantic meaning of sentences. For the multimodal multi-user system, we develop a Transformer enabled model, named, DeepSC-VQA, for the VQA task by extracting text-image information at the transmitters and fusing it at the receiver. In particular, a novel layer-wise Transformer is designed to help fuse multimodal data by adding connection between each of the encoder and decoder layers. Numerical results show that the proposed models are superior to traditional communications in terms of the robustness to channels, computational complexity, transmission delay, and the task-execution performance at various task-specific metrics. Huiqiang Xie, Zhijin Qin, Xiaoming Tao 0001, Khaled Ben Letaief |
IEEE J. Sel. Areas Commun. | 3 |
| 2022 | Viewport-Based CNN: A Multi-Task Approach for Assessing 360° Video QualityabstractFor 360° video, the existing visual quality assessment (VQA) approaches are designed based on either the whole frames or the cropped patches, ignoring the fact that subjects can only access viewports. When watching 360° video, subjects select viewports through head movement (HM) and then fixate on attractive regions within the viewports through eye movement (EM). Therefore, this paper proposes a two-staged multi-task approach for viewport-based VQA on 360° video. Specifically, we first establish a large-scale VQA dataset of 360° video, called VQA-ODV, which collects the subjective quality scores and the HM and EM data on 600 video sequences. By mining our dataset, we find that the subjective quality of 360° video is related to camera motion, viewport positions and saliency within viewports. Accordingly, we propose a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, which has a novel multi-task architecture composed of a viewport proposal network (VP-net) and viewport quality network (VQ-net). The VP-net handles the auxiliary tasks of camera motion detection and viewport proposal, while the VQ-net accomplishes the auxiliary task of viewport saliency prediction and the main task of VQA. The experiments validate that our V-CNN approach significantly advances state-of-the-art VQA performance on 360° video and it is also effective in the three auxiliary tasks. Mai Xu, Lai Jiang 0004, Chen Li 0049, Zulin Wang, Xiaoming Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Human Perception Measurement by Electroencephalography for Facial Image CompressionabstractFacial images are the main focused contents in video conferences and many other applications. Therefore, it turns to a critical issue to measure and maintain the perceptual quality of facial images if transmitted over a bandwidth-limited communication system. In this letter, we propose a regional distortion perceptual threshold measurement model based on electroencephalography (EEG) to establish the relationship between image quality and human perception. Then, a facial image compression method is presented based on the model to improve the perceptual quality. Specifically, we construct a facial image dataset with regional distortion using the better portable graphics (BPG) compression. With the dataset, we design an EEG experiment and collect the brain responses to measure the human perception on regional distortions. By this method, a regional distortion perceptual threshold map (RDPTM) is constructed to guide the data rate allocation process for different regions of facial images. The experimental results show that our method can measure the human perception of regional distortion using EEG and improve the image perceptual quality by data rate allocation based on the RDPTM. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu |
IEEE Signal Process. Lett. | 3 |
| 2021 | Brain-Inspired Image Quality Assessment Method based on Electroencephalography Feature LearningabstractWith the explosion of multimedia data, quality of experience (QoE) has become a critical metric in multimedia transmission, and therefore, QoE-oriented image quality assess-ment (IQA) turns more important and urgent. However, the performance of the traditional user-based assessment methods is limited by the deviation caused by human cognitive activities. In this paper, we propose a brain-inspired IQA method based on electroencephalography (EEG) feature learning, which is a psychophysiological method for studying human perception for IQA. We first establish the EEG dataset by collecting the corresponding EEG signals when subjects watch distorted facial images and then design a siamese network to extract the EEG features that can distinguish image quality levels and measure user scores. The siamese network establishes the relationship between image quality and QoE that is reflected by the EEG scores. The relationship is then embedded into a prediction network that directly obtains the EEG scores from images with different qualities. In this way, EEG scores can be predicted through end-to-end learning. Experiment results show that our proposed method can not only better evaluate the perceptual quality of facial images and reflect real human perceptions but also achieve better score prediction performance on the facial image datasets. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu |
GLOBECOM | 3 |
| 2021 | Deep Learning-Based Image Semantic Coding for Semantic CommunicationsabstractThis paper presents the Generative Adversarial Networks (GANs)-based image semantic coding, the goal of which is semantic exchange rather than symbol transmission. State-of-the-art visually pleasing reconstruction and semantic preserving performance are obtained in extreme low bitrate via a rate-perception-distortion optimization framework. In particular, we investigate convolutional encoder, quantizer, conditional SPADE generator, residual coding as well as perceptual losses. In contrast to previous work, we designed a coarse-to-fine image semantic coding model for multimedia semantic communication system. The base layer of the image is fully generated and preserves semantic information while the enhancement layer restores the fine details. We explore the perception and distortion performance trade-off by tuning the rate of base layer and enhancement layer. Different from the existing methods that adopt pixel accuracy as distortion metric, we train and evaluate the proposed image semantic coding model with multiple perception metrics, in line with the purpose of semantic communications. Experimental results demonstrate that our model could achieve visually pleasant and semantic consistent reconstruction, as well as saving times of bitrate, compared to BPG, WebP, JPEG2000, JPEG, and other deep learning-based image codecs. Danlan Huang, Xiaoming Tao 0001, Feifei Gao 0001, Jianhua Lu |
GLOBECOM | 2 |
| 2021 | Video Quality Measurement For Buffering Time Based On EEG Frequency FeatureabstractCurrently, user-based quality of experience (QoE) measurement methods (e.g., mean opinion score, MOS) are often employed. However, their results might be affected by human subjective experience and thoughts. Physiological measurement methods can overcome these disadvantages. In the field of video quality models, the video buffering problem caused by poor network conditions is an important factor that affects QoE. In this paper, a reasonable psychophysiological measurement method, electroencephalography (EEG), is proposed to quantitatively analyze QoE changes when users face different levels of video buffering time. By extracting the band power of EEG signals as the feature and analyzing the correlation and variance, a more objective buffering time EEG (BT-EEG) score model of video quality under the single-factor video buffering time is established, which solves the problem of uneven subjective data quality. Bingrui Geng, Xiaoming Tao 0001, Yiping Duan, Dingcheng Gao, Shuzhan Hu |
ICIP | 3 |
| 2021 | Eeg Based Visual Classification With Multi-Feature Joint LearningabstractWith a significant boost in neuroscience and artificial intelligence, decoding the process of human vision has become a hot topic in the last few decades. Although many existing deep learning models are employed to explore and solve mysteries of human brain activity, the accuracy and reliability of the visual classification task based on electroencephalography (EEG) still have space for promotion. In our research, we design the experiments to collect the subjects’ EEG data when they are watching the different types of images. In this way, an image-EEG dataset corresponding to 80 ImageNet object classes was constructed. Afterward, we proposed a dual-EEGNet for joint feature learning for multi-category visual classification. Especially, one branch EEGNet is used to extract the spatio-temporal embeddings of EEG signals, and the other branch is used to extract the time-frequency embeddings of EEG signals. The experimental results demonstrate that EEG signals can reflect the human brain activity and distinguish the different types of images. Moreover, the proposed model with joint features has a better classification performance in terms of accuracy compared with other methods. Yiping Duan, Shuzhan Hu, Xiaoming Tao 0001, Ning Ge 0001 |
ICIP | 4 |
| 2021 | Asymmetric Adaptive Modulation for Uplink NOMA SystemsabstractNon-orthogonal multiple access (NOMA) as a promising technology, can achieve enhanced connectivity and spectral efficiency. However, the asymmetric channels in uplink NOMA systems can bring distinct uncertainty of the bit error rate (BER) and throughput performances to the users. To address this problem, an asymmetric adaptive modulation (AAM) framework for uplink NOMA systems is introduced in this paper. We first derive the closed-form BER expressions for uplink NOMA systems, then we propose a principle for reducing the performance uncertainty. Finally, we develop an AAM algorithm for uplink NOMA systems by studying the closed-form BER expressions. Numerical results demonstrate the correctness of the BER expressions we derived, and show that our AAM algorithm outperforms the benchmark in terms of the system sum throughput. Tianheng Xu, Honglin Hu, Xiaoming Tao 0001 |
IEEE Trans. Commun. | 5 |
| 2021 | Deep Learning Based Channel Covariance Matrix Estimation With User Location and Scene ImagesabstractChannel covariance matrix (CCM) is one critical parameter for designing the communications systems. In this paper, a novel framework of the deep learning (DL) based CCM estimation is proposed that exploits the perception of the transmission environment without any channel sample or the pilot signals. Specifically, as CCM is affected by the user’s movement, we design a deep neural network (DNN) to predict CCM from user location and user speed, and the corresponding estimation method is named as ULCCME. A location denoising method is further developed to reduce the positioning error and improve the robustness of ULCCME. For cases when user location information is not available, we propose an interesting way that uses the environmental 3D images to predict the CCM, and the corresponding estimation method is named as SICCME. Simulation results show that both the proposed methods are effective and will benefit the subsequent channel estimation. Weihua Xu 0001, Feifei Gao 0001, Jianhua Zhang 0001, Xiaoming Tao 0001, Ahmed Alkhateeb |
IEEE Trans. Commun. | 4 |
| 2021 | Semantic Perceptual Image Compression With a Laplacian Pyramid of Convolutional NetworksabstractThe existing image compression methods usually choose or optimize low-level representation manually. Actually, these methods struggle for the texture restoration at low bit rates. Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. To achieve better perceptual quality, generative models are widely used, especially generative adversarial networks (GAN). However, training GAN is intractable, especially for high-resolution images, with the challenges of unconvincing reconstructions and unstable training. To overcome these problems, we propose a novel DNN-based image compression framework in this paper. The key point is decomposing an image into multi-scale sub-images using the proposed Laplacian pyramid based multi-scale networks. For each pyramid scale, we train a specific DNN to exploit the compressive representation. Meanwhile, each scale is optimized with different aspects, including pixel, semantics, distribution and entropy, for a good "rate-distortion-perception" trade-off. By independently optimizing each pyramid scale, we make each stage manageable and make each sub-image plausible. Experimental results demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in providing improved visual quality. Additionally, a better performance in the down-stream visual analysis tasks which are conducted on the reconstructed images, validates the excellent semantics-preserving ability of the proposed method. Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
IEEE Trans. Image Process. | 3 |
| 2021 | Saliency Prediction on Omnidirectional Image With Generative Adversarial Imitation LearningabstractWhen watching omnidirectional images (ODIs), subjects can access different viewports by moving their heads. Therefore, it is necessary to predict subjects' head fixations on ODIs. Inspired by generative adversarial imitation learning (GAIL), this paper proposes a novel approach to predict saliency of head fixations on ODIs, named SalGAIL. First, we establish a dataset for attention on ODIs (AOI). In contrast to traditional datasets, our AOI dataset is large-scale, which contains the head fixations of 30 subjects viewing 600 ODIs. Next, we mine our AOI dataset and discover three findings: (1) the consistency of head fixations are consistent among subjects, and it grows alongside the increased subject number; (2) the head fixations exist with a front center bias (FCB); and (3) the magnitude of head movement is similar across the subjects. According to these findings, our SalGAIL approach applies deep reinforcement learning (DRL) to predict the head fixations of one subject, in which GAIL learns the reward of DRL, rather than the traditional human-designed reward. Then, multi-stream DRL is developed to yield the head fixations of different subjects, and the saliency map of an ODI is generated via convoluting predicted head fixations. Finally, experiments validate the effectiveness of our approach in predicting saliency maps of ODIs, significantly better than 11 state-of-the-art approaches. Our AOI dataset and code of SalGAIL are available online at https://github.com/yanglixiaoshen/SalGAIL. Mai Xu, Li Yang 0014, Xiaoming Tao 0001, Yiping Duan, Zulin Wang |
IEEE Trans. Image Process. | 3 |
| 2021 | Anonymization and De-Anonymization of Mobility Trajectories: Dissecting the Gaps Between Theory and PracticeabstractHuman mobility trajectories are increasingly collected by ISPs to assist academic research and commercial applications. Meanwhile, there is a growing concern that individual trajectories can be de-anonymized when the data is shared, using information from external sources (e.g., online social networks). To understand this risk, prior works either estimate the theoretical privacy bound or simulate de-anonymization attacks on synthetically created datasets. However, it is not clear how well the theoretical estimations are preserved in practice. In this article, we collected a large-scale ground-truth trajectory dataset from 2,161,500 users of a cellular network, and two matched external trajectory datasets from a large social network (56,683 users) and a check-in/review service (45,790 users) on the same user population. The two sets of large ground-truth data provide a rare opportunity to extensively evaluate a variety of de-anonymization algorithms (nine in total). We find that their performance in the real-world dataset is far from the theoretical bound. Further analysis shows that most algorithms have under-estimated the impact of spatio-temporal mismatches between the data from different sources, and the high sparsity of user generated data also contributes to the under-performance. Based on these insights, we propose four new algorithms that are specially designed to tolerate spatial or temporal mismatches (or both) and model location contexts and time contexts. Extensive evaluations show that our algorithms achieve more than 17 percent performance gain over the best existing algorithms, confirming our insights. Further, we propose two new location-privacy preserving mechanisms utilizing the spatio-temporal mismatches to better protect users' privacy against the de-anonymization attack. Evaluation results show that our proposed mechanisms can reduce the performance of de-anonymization attacks by over 8.0 percent, demonstrating the effectiveness of our insights. Huandong Wang, Yong Li 0008, Chen Gao 0001, Gang Wang 0011, Xiaoming Tao 0001, Depeng Jin |
IEEE Trans. Mob. Comput. | 5 |
| 2020 | Local-to-Global Semantic Supervised Learning for Image CaptioningabstractImage captioning is a challenging problem owing to the complexity of image content and the diverse ways of describing the content in natural language. Although current methods have made substantial progress in terms of objective metrics (such as BLEU, METEOR, ROUGE-L and CIDEr), there still exist some problems. Specifically, most of these methods are trained to maximize the log-likelihood or objective metrics. As a result, these methods often generate rigid and semantically incomplete captions. In this paper, we develop a new model that aims to generate captions conforming to human evaluation. The core idea is to use local-to-global semantic supervised learning by introducing the two-level optimization objective functions. At the word level, we match each word to the image regions using the local attention objective function; at the sentence level, we align the entire sentence and the image using the global semantic objective function. Experimentally, we compare the proposed model with current methods on MSCOCO dataset. We show that either local attention supervision or global semantic supervision is the necessary component for the success of our model through ablation studies. Furthermore, combining these two supervision objective functions achieves state-of-the-art performance in terms of both standard evaluation metrics and human judgment. Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Jianhua Lu |
ICC | 3 |
| 2020 | A Novel EEG Based Directed Transfer Function for Investigating Human Perception to Audio NoiseabstractAudio quality greatly affects users evaluation of multimedia communication, especially when the communication signal is disturbed, the noise in audio and video will decrease the quality of user experience. Psychophysiological indicators have high time resolution and precision, which can be used as important quality of experience characteristics. In this paper, electroencephalography is used as a psychophysiological method to assess brain connectivity in response to perceive the noise under different scenario. Specifically, we first record the response of the subjects' brainwaves to the audio quality using a high resolution electroencephalogram. Then, directed transfer function is used to analyze the directional information flow intensity between channels in the frequency domain, and 10% directed transfer function value are selected to construct the edge set of the directed graph to obtain the brain connectivity graph. Finally, the human perception to audio noise is obtained by using the weighted degree clustering method. In addition, the effectiveness of above results is verified by the small-world network coefficients experiment. Bingrui Geng, Yiping Duan, Qiwei Song, Xiaoming Tao 0001, Jianhua Lu, Jincheng Shi |
IWCMC | 5 |
| 2020 | EEG-Based Maritime Object Detection for IoT-Driven Surveillance Systems in Smart OceanabstractAutomated maritime object detection is a significant research challenge in intelligent marine surveillance systems for the Internet of Things (IoT) and smart ocean applications. In particular, ship detection is recognized as one of the core research issues of these IoT-driven intelligent marine surveillance systems. Traditional methods based on machine learning have made some achievements in detection tasks for specific objects. However, the ship objects are relatively small, and they are usually not accurately detected. In this article, we propose an electroencephalography (EEG)-based maritime object detection algorithm for IoT-driven surveillance systems in the smart ocean. For this purpose, we conduct experiments to record the EEG signals of subjects when they are watching the maritime image scenes. With the feature analysis of EEG signals, the event-related potential (ERP) components associated with detecting objects are induced, such as the$P3$and$N2$components. Employing classification based on linear discriminant analysis (LDA), the area under curve (AUC) of the receiver operating characteristic (ROC) is used to evaluate the detection accuracy. We use this novel method to determine and identify essential objects and areas from IoT devices, such as digital camera imaging sensors. Our proposed method can not only help to detect small objects accurately using fewer samples but can also be used to reduce the data volume needed to be stored and transmitted in IoT-driven marine surveillance systems. Yiping Duan, Xiaoming Tao 0001, Qiang Li 0035, Shuzhan Hu, Jianhua Lu |
IEEE Internet Things J. | 3 |
| 2020 | Toward Variable-Rate Generative Compression by Reducing the Channel RedundancyabstractCompressing large images with a generative model goes beyond typical image encoding standards under a notably low bitrate. In this paper, we step toward practical generative compression systems based on recent advances. Specifically, we show that the channel redundancy of the latent representation produced by an autoencoder network can be effectively compressed via mask compression. The mask compression performs quantization on the channel variance of latent representation instead of original values. Instead of training multiple models, changing the mask leads to a simple and efficient variable rate compression scheme. Then, we estimate the relative bitrate by measuring the L1 norm of the channel variance and hence obtain the rate-distortion formulation. The L1 regularizer assumes a Laplacian prior on the channel variance, through which model we develop corresponding methods to produce approximate images at a target bitrate. This eliminates the need for manually searching hyperparameters for our variable-rate compression. We conduct exhaustive experiments to demonstrate the advanced performance of the proposed method in preserving image quality and semantics. Chaoyi Han, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Hyperspectral Image Denoising via Matrix Factorization and Deep Prior RegularizationabstractDeep learning has been successfully introduced for 2D-image denoising, but it is still unsatisfactory for hyperspectral image (HSI) denosing due to the unacceptable computational complexity of the end-to-end training process and the difficulty of building a universal 3D-image training dataset. In this paper, instead of developing an end-to-end deep learning denoising network, we propose a hyperspectral image denoising framework for the removal of mixed Gaussian impulse noise, in which the denoising problem is modeled as a convolutional neural network (CNN) constrained non-negative matrix factorization problem. Using the proximal alternating linearized minimization, the optimization can be divided into three steps: the update of the spectral matrix, the update of the abundance matrix and the estimation of the sparse noise. Then, we design the CNN architecture and proposed two training schemes, which can allow the CNN to be trained with a 2D-image dataset. Compared with the state-of-the-art denoising methods, the proposed method has relatively good performance on the removal of the Gaussian and mixed Gaussian impulse noises. More importantly, the proposed model can be only trained once by a 2D-image dataset, but can be used to denoise HSIs with different numbers of channel bands. Baihong Lin, Xiaoming Tao 0001, Jianhua Lu |
IEEE Trans. Image Process. | 2 |
| 2020 | An EEG-Based Study on Perception of Video Distortion Under Various Content Motion ConditionsabstractHuman perception sensitivity to video distortion is vital for visual quality assessment (VQA). Different from the perception mechanism of image distortion that has been thoroughly studied, the perception of video distortion is inevitably influenced by motion of dynamic content due to the characteristics of the human visual system (HVS). In this paper, electroencephalography (EEG) is used as a novel psychophysiological method to study the human perception sensitivity to quantification-aroused video distortion under various content motion conditions. For this purpose, we conduct experiments to record the EEG signals of the subjects when they are watching distorted videos. According to the feature analysis of EEG data, the P300 component aroused by human perception of video quality change is selected as the indicator of human perception of distortion. By the means of classification based on linear discriminant analysis (LDA), it is found that the separability of the P300 component, which is measured by the area under curve (AUC) of the receiver operating characteristic (ROC), is positively correlated with the perceptibility of distortion. The correlation provides a valid psychophysiological method, which is exempt from being influenced by subjective bias due to human high-level cognitive activities, for evaluating distortion perceptibility. In addition, the regression analysis results demonstrate a sigmoid-typed quantitative relation between the perceptibility of distortion and separability of the P300 component. Based on such relation, the perceptibility thresholds of distortion corresponding to various content motion speeds are calibrated by EEG signals and it is found that the content motion speed has a significant impact on distortion perceptibility. Xiaoming Tao 0001, Mai Xu, Yafeng Zhan, Jianhua Lu |
IEEE Trans. Multim. | 2 |
| 2020 | Trace-Driven QoE-Aware Proactive Caching for Mobile Video Streaming in MetropolisabstractTo meet the ever-increasing demands for mobile video streaming, proactive caching over the network edge has been proposed as a promising solution for next generation wireless networks. In this paper, we consider the trace-driven cache-enabled video streaming design in the scenario of a metropolis to boost the spectral efficiency on the system side and the quality of experience (QoE) on the user side. A novel scheme to jointly provide proactive caching, power allocation, user association and adaptive video streaming is designed via the formation of a QoE-aware throughput maximization problem. Specifically, the caches are refreshed in the content placement phase according to the resource status and expected traffic, which is obtained by exploring the traces collected over a big city. In addition, users need to be associated with a proper small base station (SBS) in the content delivering phase to provide the highest attainable rate. We demonstrate the effectiveness of the proposed scheme via experiments conducted over real user trace datasets. Danlan Huang, Xiaoming Tao 0001, Chunxiao Jiang, Shuguang Cui, Jianhua Lu |
IEEE Trans. Wirel. Commun. | 2 |
| 2019 | Viewport Proposal CNN for 360deg Video Quality AssessmentabstractRecent years have witnessed the growing interest in visual quality assessment (VQA) for 360° video. Unfortunately, the existing VQA approaches do not consider the facts that: 1) Observers only see viewports of 360° video, rather than patches or whole 360° frames. 2) Within the viewport, only salient regions can be perceived by observers with high resolution. Thus, this paper proposes a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, considering both auxiliary tasks of viewport proposal and viewport saliency prediction. Our V-CNN approach is composed of two stages, i.e., viewport proposal and VQA. In the first stage, the viewport proposal network (VP-net) is developed to yield several potential viewports, seen as the first auxiliary task. In the second stage, a viewport quality network (VQ-net) is designed to rate the VQA score for each proposed viewport, in which the saliency map of the viewport is predicted and then utilized in VQA score rating. Consequently, another auxiliary task of viewport saliency prediction can be achieved. More importantly, the main task of VQA on 360° video can be accomplished via integrating the VQA scores of all viewports. The experiments validate the effectiveness of our V-CNN approach in significantly advancing the state-of-the-art performance of VQA on 360° video. In addition, our approach achieves comparable performance in two auxiliary tasks. The code of our V-CNN approach is available at https://github.com/Archer-Tatsu/V-CNN. Chen Li 0049, Mai Xu, Lai Jiang 0004, Shanyi Zhang, Xiaoming Tao 0001 |
CVPR | 5 |
| 2019 | A DenseNet Based Approach for Multi-frame In-loop Filter in HEVCabstractHigh efficiency video coding (HEVC) has brought outperforming efficiency for video compression. To reduce the compression artifacts of HEVC, we propose a DenseNet based approach as the in-loop filter of HEVC, which leverages multiple adjacent frames to enhance the quality of each encoded frame. Specifically, the higher-quality frames are found by a reference frame selector (RFS). Then, a deep neural network for multi-frame in-loop filter (named MIF-Net) is developed to enhance the quality of each encoded frame by utilizing the spatial information of this frame and the temporal information of its neighboring higher-quality frames. The MIF-Net is built on the recently developed DenseNet, benefiting from the improved generalization capacity and computational efficiency. Finally, experimental results verify the effectiveness of our multi-frame in-loop filter, outperforming the HM baseline and other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Xiaoming Tao 0001 |
DCC | 4 |
| 2019 | Mobility-Aware Centralized Reinforcement Learning for Dynamic Resource Allocation in HetNetsabstractHeterogeneous networks (HetNets) can improve resource efficiency and coverage range in cellular networks to meet the growing demand for wireless data rate. The main challenges faced by HetNets are load balancing and interference coordination, which needs to be addressed by effective user association and resource allocation (UARA) methods. In this paper, we propose a mobility- aware centralized reinforcement learning (MCRL) framework in order to achieve global optimality of dynamic resource allocation. A centralized agent is defined to select the values of the hyper parameters for UARA according to the real-time status of all users in HetNets. Besides, the state of the art Actor-Critic technique is employed in the training process to guarantee the convergence and performance of the agent's policy. Simulation results demonstrate the effectiveness of the proposed method and show the performance gain under different user distributions. Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 2 |
| 2019 | Optimizing QoE of Multiple Users over DASH: A Meta-learning ApproachabstractDynamic adaptive video streaming over HTTP (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches concentrate on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this paper, we formulate the QoE optimization of multi-user preferences as a multi-task deep reinforcement learning problem, in which QoE refers to the metrics of visual quality, fluctuation and rebuffing events. Then, we propose a meta-learning framework for multi-user preferences (MLMP) as a new DASH adaptation approach. Finally, the simulation results show that the proposed approach outperforms state-of-the-art DASH adaptation approaches in satisfying the different users' QoE preferences regarding the three metrics. Liangyu Huo, Zulin Wang, Mai Xu, Zhiguo Ding 0001, Xiaoming Tao 0001 |
ICASSP | 5 |
| 2019 | Geometry-Aware GAN for Face Attribute TransferabstractIn this paper, the geometry-aware GAN is proposed to address the issue of facial attribute transfer with unpaired data. To tackle the unpaired training sample problem, the CycleGAN architecture is applied, where the bilateral mappings between the source and target domains are learned. The deformation flow is learned to capture the geometric variation between two domains. We first warp the source face into desired pose and shape according to the flow. Then, the transfer sub-network is designed to refine the results by hallucinating new components on the warped image. The attribute is removed by the reconstruction sub-network, coupled with the warping process. Experiments on benchmark demonstrate the advantages of our method compared to baselines. Danlan Huang, Xiaoming Tao 0001, Jianhua Lu, Minh N. Do |
ICIP | 2 |
| 2019 | Semantic Perceptual Image Compression with a Laplacian Pyramid of Convolutional NetworksabstractRecently, deep neural network (DNN)-based image compression methods have achieved impressive results. These methods generally use thumbnail images or crop small patches from high-resolution images to train their networks. Instead of using patch-based training mode, we propose a novel DNN-based image compression framework in this paper. We apply the Laplacian pyramid to construct a multi-scale image representation. By learning the increasingly detailed representations, the proposed method is able to progressively restore an image. Furthermore, we use the adversarial networks for training to encourage the perceptual quality of the reconstructed image. Particularly at low bitrates, our model can only store the global semantics of an image and automatically synthesize the texture to achieve high subjective quality. Experimental results on demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in terms of visual quality. Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
ICIP | 2 |
| 2019 | Joint 3-D Shape Estimation and Landmark Localization From Monocular Cameras of Intelligent Vehiclesabstract3-D reconstruction is at the core for many driving applications of Internet of Intelligent Vehicles. Previous works on reconstruction of a 3-D point shape commonly use a two-step framework. Precisely localizing a series of feature points in an image is performed on the first step. Then the second procedure attempts to fit the 3-D data to the observations to get the real 3-D shape. Such an approach has high time consumption, and easily gets stuck into local minimum. To address this problem, we propose a method to jointly estimate the global 3-D geometric structure of car and localize 2-D landmarks from a single viewpoint image. First, we represent the 3-D shape with a set of predefined shape bases, while parametrizing it by the coefficients of the linear combination of them. Second, we adopt a cascaded regression framework to regress the global shape encoded by the prior bases, by jointly minimizing the appearance and shape fitting differences. The position fitting item can help cope with the description ambiguity of local appearance, and provide more information for 3-D reconstruction. We apply the proposed approach on a multiview car dataset. Experimental results demonstrate favorable improvements on pose estimation and shape prediction, compared with some previous methods. Yanan Miao, Xiaoming Tao 0001, Jianhua Lu |
IEEE Internet Things J. | 2 |
| 2019 | Rebuffering Optimization for DASH via Pricing and EEG-Based QoE ModelingabstractPricing is an effective mechanism for network resource allocation that can be used to achieve a desirable balance between efficiency and fairness. However, sophisticated utility models are needed to guarantee the performance of price-based resource allocation, especially with regard to video transmission. Among various performance indices of video transmission, rebuffering is an important one that influences user quality of experience (QoE). Therefore, a price-based bandwidth allocation scheme for a dynamic adaptive streaming over hypertext transfer protocol (DASH) system is proposed for mitigating the effect of rebuffering on QoE. The utility model of the proposed scheme considers the relationship between the allocated bandwidth and the rebuffering length, as well as the effect of rebuffering length on QoE. Specifically, electroencephalography (EEG) experiments are conducted, and the distribution of the subjects' time limits at which rebuffering arouse negative emotions is used for calibration. Based on this model, the DASH server collects the buffer state information of all DASH clients periodically to adjust its bandwidth allocation. Assuming the server protects itself from congestion by pricing the clients' requested bandwidth, a Stackelberg game is formulated to study the joint utility maximization on the revenue of the server and the utility of the clients. The Stackelberg equilibrium of the game is characterized, and an efficient searching algorithm is proposed for its solution. EEG experiments are repeated on another group of subjects to verify the generalization ability of the results, and simulation results are presented to validate the effectiveness of the proposed algorithm. The proposed algorithm is shown to have low complexity and outperforms both traditional price-based scheme and QoE maximized scheme in terms of rebuffering. Xiaoming Tao 0001, Zhao Chen 0002, Mai Xu, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | Learning QoE of Mobile Video Transmission With Deep Neural Network: A Data-Driven ApproachabstractQuality of experience (QoE) serves as a direct evaluation of users' experiences in mobile video transmission and thus essential for network management, such as network optimization. In this paper, we propose a deep learning-based QoE prediction approach with a large-scale QoE dataset for mobile video transmission. Specifically, we develop a mobile phone application for collecting user QoE data when viewing videos transmitted over the mobile internet in a practical environment. Then, we construct a large-scale dataset by collecting over 80000 piece of data with four kinds of subjective scores and 89 network parameters. Each QoE metric is related to only some of the 89 network parameters. Therefore, we apply the feature selection method to find the feature parameters related to user scores. Additionally, the boxplot method is used to clean the raw data by removing outliers. Finally, a deep neural network (DNN) is developed to learn the relationships between the network parameters and the subjective QoE scores. The proposed DNN can also be seen as a data-driven objective QoE prediction approach for mobile video transmission, which can be used to predict the user QoE scores. The experimental results show that the proposed approach can effectively remove most features irrelevant to QoE prediction. Moreover, the performance of QoE prediction by the proposed model outperforms other state-of-the-art approaches. Xiaoming Tao 0001, Yiping Duan, Mai Xu, Zhishen Meng, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | Joint Minimization of Wired and Wireless Traffic for Content Delivery by Multicast PushingabstractAs more mobile users become subscribers of content services, their subscribed content can be directly pushed from the content provider into the user equipment after the content is generated. In current and future network paradigms, a joint wired and wireless transmission design for this pushing is needed to guarantee the user experience without the extra deployment of communication infrastructures or consumption of resources. In this paper, we investigate a joint wired and wireless content delivery system that incorporates wired and wireless multicast. The users in the same group are served by wireless multicast from a base station (BS), while the BSs of the same content form a multicast tree in a backbone wired network. The sum of wired and wireless traffic is minimized by a joint design of user grouping, subchannel allocation, wired routing, and wired link usage. Exploiting the monotonicity of wired and wireless traffic with regard to the wired hop count, the original problem is converted for searching the optimal hop count vector that achieves the minimum sum of both types of traffic, which is solved by a monotonic optimization (MO)-based iterative algorithm. Compared with existing schemes and according to the numerical results, a reduction in total traffic of 43% can be achieved by our approach. Zhao Chen 0002, Xiaoming Tao 0001, Chunxiao Jiang, Victor C. M. Leung |
IEEE Trans. Wirel. Commun. | 2 |
| 2018 | Joint Wired and Wireless Traffic Minimization for Energy-Efficient Content Delivery NetworksabstractPushing contents from content providers (CPs) directly to user equipments (UEs) during off-peak hours can significantly reduce the incurring traffic during peak hours. Both wired and wireless transmission costs are considerable in this scenario, especially in current cellular networks where the backhaul is regarded as a bottleneck of transmission. In order to alleviate the traffic pressure over the network without loss of users' quality of experience, this paper presents a joint wired and wireless transmission scheme which considers grouping, subchannel allocation, wired routing, and wired bandwidth allocation. An iterative algorithm is proposed to achieve the tradeoff between those two kinds of traffic. The users in the same group are served by a single multicast transmission from a base station (BS), while the BSs which multicast the same content form a multicast tree in backbone wired network. Compared with traditional unicast scheme, our approach can reduce the wired, wireless, and total traffic by 23%, 46%, and 34% at most according to the simulation results. Zhao Chen 0002, Xiaoming Tao 0001, Chunxiao Jiang, Jianhua Lu |
GLOBECOM | 2 |
| 2018 | Multi-Scale Convolutional Neural Network for SAR Image Semantic SegmentationabstractAlthough the recent success of convolutional neural networks (CNNs) greatly advance the semantic segmentation of the natural images, few work has focused on the remote sensing images, especially the synthetic aperture radar (SAR) images. Specifically, the existing methods do not consider the speckle noise of the SAR images and the multi-scale characteristics contained in the SAR images. In this paper, we propose a multiscale convolutional neural network (CNN) model for SAR image semantic segmentation. The multi-scale CNN model includes noise removal stage, convolutional stage, feature concatenation stage and classification stage. In particular, we construct a sparse representation loss function to obtain a clear SAR image in noise removal stage. Then, the multi-scale convolutional stage is employed to learn the multi-scale deep features. The concatenation stage is used to connect the features with different scales and depths. Finally, softmax classifier is developed to obtain the labels of the SAR images with the multi-scale CNN model being trained in an end-to-end way. The experimental results on synthetic and real SAR images demonstrate the effectiveness of the proposed method. Yiping Duan, Xiaoming Tao 0001, Chaoyi Han, Xiaowci Qin, Jianhua Lu |
GLOBECOM | 2 |
| 2018 | Boundary Objectness Network for Object Detection and LocalizationabstractIn this paper, we present the boundary objectness network (BON), an effective convolutional neural network (CNN) for object detection. Its core contribution is to accurately localize the objects. Generally, the CNN-based localizers predict four bounding box coordinates by learning a regression function. This method shows a low Intersection-of-Union (IoU) with the ground truth box. In our work, the localization is formu-lated as a probabilistic problem. Specifically, the deep features inside the candidate proposal are mapped into a row and a column feature vector, which are called boundary object-ness. The boundary objectness indicates the existence of an object in the horizontal and vertical direction of the proposal, enabling us to elaborately localize the object. Moreover, the modules of object detection share the common convolution-al layers. Meanwhile, a multi-task loss function is designed for joint training strategy. Experimental results on the PAS-CAL VOC datasets demonstrate the competitive performance of our method. For the VGG16 model, we achieve 77.6 % mAP at a speed of 4 frame per second (FPS), thus having the potential for real-time processing. Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
ICASSP | 2 |
| 2018 | Semantic Conditional Random Field for Object Based SAR Image SegmentationabstractConditional random filed (CRF) model relaxes the conditional independence of the observed data and simultaneously captures the spatial contextual information. However, the single spatial contextual model is difficult to describe the heterogeneous structures of the synthetic aperture radar (SAR) images. This paper propose an semantic conditional random field (SCRF), which integrate the semantic space and pixel space for object-based SAR image segmentation. Specifically, the SAR image is divided into aggregated, structural and homogeneous subspaces by using the hierarchical semantic model. Then, we design gaussian kernel function, geometric kernel function and uniform kernel function to adaptively describe the spatial contextual constraints in the different subspaces. These kernel functions are incorporated into the pairwise potential of CRF model to improve the ability of the model. Afterwards, the piecewise training and Bayesian inference are proposed to achieve the object-based segmentation. Experiments on the synthetic and real SAR images demonstrate the effectiveness of the proposed method in the semantic consistency and detail preservations. Yiping Duan, Xiaoming Tao 0001, Chaoyi Han, Jianhua Lu |
ICIP | 2 |
| 2018 | Dense Convolution for Semantic SegmentationabstractState-of-the-art semantic segmentation methods adopt fully convolutional neural networks (FCNs) to solve this dense prediction problem. However, replacing fully connected layers with the standard 2D convolution layer is straightforward yet not optimal in generating segmentation results. In this paper we develop a dense convolution scheme that is more suitable for semantic segmentation. Instead of generating a single output, dense convolution produces the same number of output as its input and introduces spatial overlaps into current convolutions. Then each activation is obtained from multiple overlapped dense convolutions with learnable weights. Such dense convolution helps to reinforce local connections between activations and provide more flexible receptive fields for predictions. Experiments on benchmark dataset demonstrate the effectiveness of the proposed approach in semantic segmentation tasks. Chaoyi Han, Xiaoming Tao 0001, Yiping Duan, Jianhua Lu |
ICIP | 2 |
| 2018 | Calibrating Human Perception Threshold of Video Distortion Using EEGabstractHuman perception threshold of video distortion is vital for visual quality assessment. Traditionally, calibrating the perception threshold of distortion relies on subjective test, which may suffer from strategy and bias of the human. In this paper, electroencephalography (EEG) is used as a novel psychophysiological method to evaluate the human perception of quantification-aroused video distortion. By the means of classification based on linear discriminant analysis (LDA), the separability of event-related potentials (ERPs) aroused by human perception to video quality change is measured by the area under curve (AUC) of the receiver operating characteristic (ROC). Relating the EEG signals to behavior data, the sigmoid-typed relationship between the perceptibility of distortion and separability of the P300 component is discovered. Based on this relationship, the perceptibility of distortion can be evaluated by EEG signals alone, which provides a potential neurally informed method for the calibration of perception threshold of video distortion. Xiaoming Tao 0001, Yafeng Zhan |
ICIP | 2 |
| 2018 | Joint Monocular 3D Car Shape Estimation and Landmark Localization via Cascaded Regression
Yanan Miao, Jia Cui, Xiaoming Tao 0001 |
ICPRAM | 4 |
| 2018 | Hyperspectral Image Denoising via Nonnegative Matrix Factorization and Convolutional Neural NetworksabstractHyperspectral image (HSI) denoising plays an important role to enhance the image quality for subsequent applications. This paper proposes a novel denoising framework for the HSI, in which the denoising issue is modeled as a convolutional neural network (CNN) constrained non-negative matrix factorization problem. Then, by adopting the proximal alternating linearized minimization, the proposed approach can be decomposed into two iterative steps: In the first step, the spectral matrix is updated using a proximal operator to guarantee the non-negativity; in the second step, the abundance matrix is updated using a designed CNN. Finally, we propose a transfer learning scheme to train the design CNN with a natural gray image dataset. Exhaustive experiments show that the proposed approach outperforms the comparison state-of-the-art methods on two different test HSI datasets. Baihong Lin, Xiaoming Tao 0001, Xiaowei Qin, Yiping Duan, Jianhua Lu |
IGARSS | 2 |
| 2018 | THU Face Database for Real-Time Automatic Video Scoring ModelabstractSpeech consists of vocal sounds resulting from the synchronism of all parts of the human phonatory apparatus. The oral communication between individuals is supposed to sound pleasant, and this characteristic is strongly related to the subjective parameters studied by phonoaudiology, such as roughness, breathiness, and strain. Commonly, these characteristics are evaluated with a noninvasive vocal disorder test. This paper proposes the development of an intelligent computational tool to classify and point out the preponderance of such parameters from audio samples. Moreover, a comparative study was made to evaluate the efficiency of several Wavelet families as feature extraction methods. The features extracted were used with artificial neural networks and an automatic routine was set up to find the best topology of the Multi-Layer Perceptron (MLP) architecture used in the classification of such speech parameters. The results are promising and reliable, with accuracy rates higher than 98.9%. Yifeng Liu 0003, Xiaoming Tao 0001, Ailing Xiao |
IJCNN | 2 |
| 2018 | Visual information assisted UAV positioning using priori remote-sensing information
Xijia Liu, Xiaoming Tao 0001, Yiping Duan, Ning Ge 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Hierarchical objectness network for region proposal generation and object detection
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Yiping Duan, Jianhua Lu |
Pattern Recognit. | 2 |
| 2018 | Robust Monocular 3D Car Shape Estimation From 2D LandmarksabstractEstimating 3D shape of an object from 2D observations in a monocular image is fundamentally an inverse problem due to the ambiguity of the projection from 3D to 2D and becomes more challenging when there are undesirable outliers in the observations. In this paper, we develop a robust model to estimate 3D shape from 2D landmarks with an unknown camera pose. The 3D shape of the object is assumed as a linear combination of a group of prior shape bases. At the same time, we explicitly model the outliers as sparse noises to handle severely contaminated observations. The objective function is nonconvex and nonsmooth constrained on Stiefel manifold, where the coupling of underdetermined shape representation coefficients and camera pose makes it more difficult to solve. We first propose a numerical algorithm based on alternating direction method of multipliers for the no-outlier case. We set the orthogonality constraints into the smooth subproblem, which admits a closed-form solution, and the other subproblems are all well known and can be easily solved. We then extend this algorithm to the proposed robust model. The proposed algorithms can achieve convergence rapidly. The experimental results on both synthetic data and real data show that the proposed method outperforms the other methods. Yanan Miao, Xiaoming Tao 0001, Jianhua Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Adaptive Hierarchical Multinomial Latent Model With Hybrid Kernel Function for SAR Image Semantic SegmentationabstractSynthetic aperture radar (SAR) images have been one of the important tools to support earth observations and topographic measurements. It means that SAR images are essentially rich in structures. However, the single spatial relationship is difficult to deal with the heterogeneous structures of the SAR images. In this paper, we propose an adaptive hierarchical multinomial latent model with hybrid kernel function for SAR image semantic segmentation. In the proposed approach, we design a hybrid kernel function combing Gaussian radial basis function (GRBF) and ridgelet kernel function to adaptively describe the spatial relationships between the central pixel and the surrounding pixels. Then, based on the hybrid kernel function, adaptive methods are proposed for semantic segmentation. Specifically, an SAR image is divided into different characteristics subspaces, homogeneous, structural, and aggregated subspaces, by SAR hierarchical semantic model. For the homogeneous subspace, GRBF is used to describe the isotropic spatial relationships. Then, multilayer multinomial latent model with GRBF is used for segmentation to improve the labeling consistency and reduce the wrong segmentation. For the structural subspace, the ridgelet kernel function is used to describe the anisotropic spatial relationships. Then, we adopt the single-layer multinomial latent model with ridgelet kernel function for segmentation to preserve the details (such as edge, lines, and small objects). For aggregated subspace, bag-of-words model is used to extract the features of the aggregated portions, and then affinity propagation cluster is used for segmentation. Finally, the segmentation results of different subspaces are integrated together to obtain the final segmentation result. Comprehensive experiments on both synthetic and real SAR images demonstrate that the segmentation results by our proposed approach achieve the semantic consistency, labeling consistency, and detail preservation simultaneously. Yiping Duan, Fang Liu 0001, Licheng Jiao, Xiaoming Tao 0001, Jie Wu 0016, Cheng Shi 0002, Martin O. Wimmers |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Saliency Detection in Face Videos: A Data-Driven ApproachabstractRecently, videoconferencing has been popular in multimedia systems, such as FaceTime and Skype. In videoconferencing, almost every frame contains a human face. Therefore, it is important to predict human visual attention on face videos by saliency detection, as saliency may be used as a guide to the region of interest for the content-based applications of face videos. In this paper, we propose a data-driven approach for saliency detection in face videos. From the data-driven perspective, we first establish an eye-tracking database that contains fixations of 76 face videos viewed by 40 subjects. Upon the analysis of our database, we find that visual attention is significantly attracted by faces in videos. More important, the attention distribution within face regions varies with regard to mouth movement. Since previous works have investigated that it is efficient to model face saliency in still images using a Gaussian mixture model (GMM), the variation of visual attention in videos can be modeled by dynamic GMM (DGMM). Accordingly, we propose adopting the particle filter (PF) in modeling DGMM for saliency detection of face videos, which is called PF-DGMM. Finally, the experimental results show that our PF-DGMM approach significantly outperforms other state-of-the-art approaches in saliency detection of face videos. Mai Xu, Yun Ren, Zulin Wang, Jingxian Liu, Xiaoming Tao 0001 |
IEEE Trans. Multim. | 5 |
| 2017 | Latency-Efficient Video Streaming in Metropolis: A Caching FrameworkabstractThis paper presents a latency-efficient mobile video streaming design in the context of metropolis by incorporating caching. It shows that video traffic can be substantially offloaded from backhaul by caching predictable demands in the network edge. Notably, exploiting the spatial and temporal characteristics of video popularity, we focus on two sub problems: how to cache the content and how to associate users. Firstly, we investigate cache deployment strategy based on clients' viewing behavior in both downtown and suburb. The proposed hybrid collaborative filtering (CF)-based scheme guarantees high hit rate utilizing the available storage capacity in small base stations (SBS). Further, we formulate the dynamic user equipment and SBS (UE- SBS) optimal association problem into a convex optimization problem, so as to maximize the sum transmission rate of SBSs under the resource and quality-of-service constraint. Performance evaluation of real trace data demonstrates the significant advantage of our proposed framework. Danlan Huang, Xiaoming Tao 0001, Chunxiao Jiang, Yong Li 0008, Jianhua Lu |
GLOBECOM | 2 |
| 2017 | Variational inference for nonparametric subspace dictionary learning with hierarchical beta processabstractNonparametric Bayesian models have been implemented in dictionary learning. However, for signal samples from multiple subspaces, existing methods only learn one uniform dictionary and thus are not optimal for representing the subspace structures. To address this issue, we first utilize a combination of Dirichlet process and hierarchical Beta process as priors to infer the latent subspace number and dictionary dimension automatically; second, to derive tractable variational inference, we modify the priors with the Sethuraman's construction and further employ the multinomial approximation. Experimental results indicate that our approach can achieve a set of nonparametric subspace dictionaries, while showing performance enhancements in the tasks of image denoising. Shaoyang Li, Xiaoming Tao 0001, Jianhua Lu |
ICASSP | 2 |
| 2017 | Variational Bayesian inference for nonparametric signal compressive sensing on structured manifoldsabstractThe conventional sparsity-based compressive sensing (CS) has been extended to a more general framework based on the broad class of manifold models. Although some existing manifold-based CS methods use a mixture of factor analyzers to discover the low-dimensional geometric structures of the signals, they have two issues that may limit their practical use: First, the signal representation using manifolds assumes that the mixture components are independent and thus misses the potential overlapping structure of the factors; Second, the mixture model is not analytically tractable and requires a time-consuming stochastic technique. In this paper, we address these issues by: 1) capturing the correlation structure of mixture components via a hierarchical Beta process which is built with the Sethuraman's stick-breaking construction in the nonparametric Bayesian manner; 2) deriving an efficient variational inference for the modified model with the assistant of multinomial approximation. Experimental results on real dataset indicate that our proposed approach can outperform state-of-the-art manifold-based inversion algorithms in the application of CS reconstruction, while exhibiting satisfying time consumption compared to the stochastic sampling schemes. Shaoyang Li, Xiaoming Tao 0001, Jianhua Lu |
ICC | 2 |
| 2017 | Prior-Information-Based Remote Sensing Image Compression with Bayesian Dictionary LearningabstractRequirements for higher resolution remote sensing images lead to rapid increase of data amount in space communications. However, since satellite communications capacity is suffering from great pressure, seeking for more effective compression scheme is supposed to solve existing conflict between tremendous data and limited bandwidth. For this reason, this paper proposes a prior-information-based remote sensing image compression scheme. We firstly utilize prior information contained in historical remote sensing images for incremental image extraction, which is assumed to have removed redundant information possessed both on the satellite and ground. Moreover, Bayesian dictionary serves to sparsely represent the incremental image, generating finite number of representation coefficients in place of numerous pixels. Finally, quantization and encoding schemes are further designed for efficient data transmission. Experimental results show that the proposed scheme is competitive to existing general image compression schemes. Xiaoming Tao 0001, Shaoyang Li, Zizhuo Zhang, Xijia Liu, Juan Wang 0012, Jianhua Lu |
VTC Spring | 1 |
| 2017 | Online Bayesian Learning for Remote-Sensing Imagery CompressionabstractThis work investigates a statistical technique for high performance remote-sensing imagery compression. By exploiting existing remote-sensing data sets, useful structural and texture prior information can be learned. The main methodologies are Bayesian dictionary learning and stochastic approximation. A Bayesian network simulating the generation mechanism of remote- sensing images is modelled. The whole compression scheme is established. And the corresponding inference algorithm using Gibbs sampling is given, where the inference is realized in an online way. The performance of the proposed compressing scheme is evaluated over a high-resolution remote-sensing image data set captured by TH-1 series satellites. Experiment results have shown that our compression scheme outperforms JPEG-2000 by 3dB on average with same bits-per-pixel performance, and that Bayesian learning can provide a dictionary with high expressiveness for remote-sensing images. In addition, with online learning skills our proposed compression scheme can scale up to very large-scale training data. Zizhuo Zhang, Shaoyang Li, Xiaoming Tao 0001, Linhao Dong, Jianhua Lu |
VTC Spring | 3 |
| 2017 | Bayesian Hyperspectral and Multispectral Image Fusions via Double Matrix FactorizationabstractThis paper focuses on fusing hyperspectral and multispectral images with an unknown arbitrary point spread function (PSF). Instead of obtaining the fused image based on the estimation of the PSF, a novel model is proposed without intervention of the PSF under Bayesian framework, in which the fused image is decomposed into double subspace-constrained matrix-factorization-based components and residuals. On the basis of the model, the fusion problem is cast as a minimum mean square error estimator of three factor matrices. Then, to approximate the posterior distribution of the unknowns efficiently, an estimation approach is developed based on variational Bayesian inference. Different from most previous works, the PSF is not required in the proposed model and is not pre-assumed to be spatially invariant. Hence, the proposed approach is not related to the estimation errors of the PSF and has potential computational benefits when extended to spatially variant imaging system. Moreover, model parameters in our approach are less dependent on the input data sets and most of them can be learned automatically without manual intervention. Exhaustive experiments on three data sets verify that our approach shows excellent performance and more robustness to the noise with acceptable computational complexity, compared with other state-of-the-art methods. Baihong Lin, Xiaoming Tao 0001, Mai Xu, Linhao Dong, Jianhua Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Robust 3D Car Shape Estimation from Landmarks in Monocular Image
Yanan Miao, Xiaoming Tao 0001, Jianhua Lu |
BMVC | 2 |
| 2016 | Variational Bayesian image fusion based on combined sparse representationsabstractHyper-spectral image fusion has been a hot topic in medical imaging and remote sensing. This paper proposes a Bayesian fusion model which combines the panchromatic (PAN) image and the low spatial resolution hyper-spectral (HS) image under the same framework. Sparsity constraint is introduced as double "spike-and-slab" priors, and anisotropic Gaussian noise is adopted for accuracy. To achieve reduction in computational complexity, we turn the anisotropic Gaussian distribution into isotropic one with modified linear transformation and propose a variational Bayesian expectation maximization (EM) algorithm to calculate the result. Experiment results show that our solution can achieve comparable performance in pan-sharpening to other state-of-art algorithms while largely reducing the computational complexity. Baihong Lin, Xiaoming Tao 0001, Shaoyang Li, Linhao Dong, Jianhua Lu |
ICASSP | 2 |
| 2016 | Variational EM approach for high resolution hyper-spectral imaging based on probabilistic matrix factorizationabstractHigh resolution hyper-spectral imaging works as a scheme to obtain images with high spatial and spectral resolutions by merging a low spatial resolution hyper-spectral image (HSI) with a high spatial resolution multi-spectral image (MSI). In this paper, we propose a novel method based on probabilistic matrix factorization under Bayesian framework: First, Gaussian priors, as observations' distributions, are given upon two HSI-MSI-pair-based images, in which two variances share the same hyper-parameter to ensure fair and effective constraints on two observations. Second, to avoid the manual tuning process and learn a better setting automatically, hyper-priors are adopted for all hyper-parameters. To that end, a variational expectation-maximization (EM) approach is devised to figure out the result expectation for its simplicity and effectiveness. Exhaustive experiments of two different cases prove that our algorithm outperforms many state-of-the-art methods. Baihong Lin, Xiaoming Tao 0001, Linhao Dong, Jianhua Lu |
ICIP | 2 |
| 2016 | Mouse calibration aided real-time gaze estimation based on boost Gaussian Bayesian learningabstractIn this paper, we propose a novel gaze estimation method to evaluate the attention span of users upon on-screen content via a single webcam. Our method is based on supervised descent method for eye region of interest (ROI) extraction. Then, boost Gaussian Bayesian regressors are applied to learn a robust mapping from the input eye ROI to gaze coordinates. To get enough training samples, we implant our scheme as a plug-in into web browsers for data collection from users without bothering. To improve accuracy, we also introduce mouse click to help train the regressors. Experiment results show that our method outperforms the existing method and can provide gaze estimation data for user behaviour analysis in real-time implementation. Nanyang Ye 0001, Xiaoming Tao 0001, Linhao Dong, Ning Ge 0001 |
ICIP | 2 |
| 2016 | Quantization and Entropy Coding Scheme for Dictionary Learning Based Image CompressionabstractMost recently, there has been a growing interest in the study of dictionary learning based (DL-based) image compression, which has potential in relieving the bandwith-hungry bottleneck of visual communication. All existing DL-based image compression approaches mainly focus on the effective representation of images, thus losing sight of two basic elements of image compression, i.e., quantization and entropy coding. For this reason, this paper proposes a quantization and entropy coding scheme for DL-based image compression. In our scheme, the proposed Partition-Interval K-means (PIK) quantizer adaptively maps continuous coefficients to discrete values. The arithmetic coding, combined with differential coding technique, is applied to encode the indices of nonzero coefficients as well as the labels of quantization values. In our experiments, the proposed scheme is verified to be more effective than other quantization and entropy coding schemes for DL-based image compression. Juan Wang 0012, Xiaoming Tao 0001, Xijia Liu, Ning Ge 0001, Jianhua Lu |
VTC Fall | 2 |
| 2016 | The THU multi-view face database for videoconferences and baseline evaluations
Xiaoming Tao 0001, Linhao Dong, Yang Li 0005, Jianhua Lu |
Neurocomputing | 1 |
| 2016 | HEMS: Hierarchical Exemplar-Based Matching-Synthesis for Object-Aware Image ReconstructionabstractMotivated by the attention on salient objects, conventional region-of-interest (ROI)-based image coding approaches attempt to assign more bits to ROIs and fewer bits to other regions. Thus, the perceptual quality of salient object regions is improved by sacrificing the quality of non-ROI regions with unpleasant artifacts. To address this issue, we concentrate on the efficient compression of object-centered images by encoding salient objects and background features separately. To fully recover the object and background, we propose a hierarchical exemplar-based matching-synthesis (HEMS) approach to reconstruct the image from exemplars. In the proposed framework, once the salient object regions are encoded, only the quantized color features and local descriptors of the background are kept, achieving bit-rate reduction. To make it possible and practical to reconstruct background regions, the hierarchical framework is designed in three layers, including relevant image search, patch candidates matching, and distortion optimized image synthesis. In the hierarchical framework, firstly, image search from an external database returns relevant images, limiting the search space to a feasible number of patch candidates. Secondly, patches are matched by color features to select the appropriate candidates. Finally, the distortion optimized image synthesis further makes it possible to automatically choose the most suitable texture sample, and seamlessly reconstruct the image. Compared to the conventional ROI-based image coding schemes, the proposed approach can achieve better visual quality on both ROI and background regions. Yipeng Sun, Xiaoming Tao 0001, Yang Li 0005, Linhao Dong, Jianhua Lu |
IEEE Trans. Multim. | 2 |
| 2015 | A Nonparametric Bayesian Approach to Image Compressive Sensing on Manifolds with Correlation ConstraintsabstractThe broad class of manifold models are considered for extending the conventional compressive sensing (CS) to a more general framework. However, although the manifold-based CS approaches using a mixture of factor analyzers can learn latent geometric structures of high-dimensional signals, they have two issues that potentially limit their practical use: First, the manifold modeling does not take account of sharing factors among mixture components and thus misses the chance to enhance statistical strength; Second, introducing correlation constraints may cause intractable model inference. In this paper, we address these issues by: (1) depicting the mixture correlations with a hierarchical Beta process whose upper-layer process is an Indian buffet process in nonparametric Bayesian manner; (2) deriving a combination of collapsed Gibbs sampler and auxiliary-variable-based slice sampler to obtain the model accurate solutions. Experimental results on real dataset demonstrate that our proposed method provides significant performance improvements compared to the sparsity-based CS algorithms, while outperforming the state-of-the-art manifold-based inversion strategies for image reconstruction. Shaoyang Li, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 2 |
| 2015 | A New Radiation Correction Method for Remote Sensing Images Based on Change Detection
Juan Wang 0012, Xijia Liu, Xiaoming Tao 0001, Ning Ge 0001 |
ICIG (1) | 3 |
| 2015 | The THU multi-view face database for videoconferencesabstractIn this paper, we present a face video database that contains 31,500 videos of 100 individual volunteers. The primary purpose of building this database is to serve as a standardized test video sequences for any research related to video-conferences. Each of the volunteers was filmed by 9 groups of synchronized webcams under 7 illumination conditions, and was requested to complete a series designated actions. Thus, face variations on lip shape, occlusion, illumination, pose, and expression are presented in each video clip. Compared to the existing databases, THU face database provides multi-view video sequences with strict temporal synchronization, enabling evaluations on gaze-correction methods. Besides, based on our database, three well-known methods were tested, demonstrating the numerical performances under different circumstances. Free samples of this database can be downloaded at www.facedbv.com. Linhao Dong, Xiaoming Tao 0001, Yang Li 0005, Jichuan Lu, Zizhuo Zhang, Jingwen Cheng, Jianhua Lu |
ICIP | 2 |
| 2015 | Rate-distortion optimized inter-frame compression for parameter-driven animationabstractAs an important computer graphics technique, parameter-driven animation has seen increasing deployment on mobile platforms, where both communication bandwidth and device storage are subjected to stringent constraints. We propose an efficient rate-distortion optimized inter-frame compression scheme for parameter-driven animation capable of finding optimal bit-allocations for any given bit-rate. Experiments conducted on face animation based on active appearance models demonstrate that with the proposed method, the transmission and storage requirements of parameter-driven animation can be significantly reduced. Yang Li 0005, Xiaoming Tao 0001, Jianhua Lu |
PCS | 2 |
| 2015 | Large-scale structured sparse image reconstruction with correlated multiple-measurement vectors using Bayesian learningabstractThis paper proposes a Bayesian learning approach to structured sparse image reconstruction. In contrast to conventional paradigms which convert images into high-dimensional vectors and thus are impractical for recovering large-scale images, we formulate columns of image matrices into a multiple-measurement-vector (MMV) model to reduce the problem dimension. Besides, we simultaneously exploit the tree structure of image wavelet coefficients and the column correlations of image matrices in wavelet domain as two prior structured constraints to improve reconstruction accuracy. Experimental results reveal that our method significantly outperforms other MMV-based strategies in terms of reconstruction error and provides a practical and efficient alternative to large-scale structured sparse image reconstruction. Shaoyang Li, Xiaoming Tao 0001, Yang Li 0005, Jianhua Lu |
PCS | 2 |
| 2015 | Remote-Sensing Image Compression Using Priori-Information and Feature RegistrationabstractIn this paper, we focus on a high performance compression scheme for remote-sensing images, which is essential due to limited transmission bandwidth while explosively growing remote-sensing image data size. First, on the basis of intra-image spatial redundancy removal, which is used by JPEG 2000 and CCSDS, priori-information is introduced to eliminate temporal redundancy between historical and newly-captured images, at the same time. Second, feature registration technique is applied rather than motion estimation and compensation which is used in HEVC, to deal with the long-range non-linear correlation of remote-sensing image series. Numerical simulation results show that the proposed scheme outperforms JPEG 2000 and JPEG by over 1.37 times for lossless compression, and presents a 5 dB PSNR gain over JPEG 2000 and HEVC for lossy compression. Xijia Liu, Xiaoming Tao 0001, Ning Ge 0001 |
VTC Fall | 2 |
| 2015 | Efficient Multi-Cell Clustering for Coordinated Multi-Point Transmission with Blossom Tree AlgorithmabstractCoordinated multi-point(CoMP) transmission clustering schemes could provide significant gains of system performance, such as throughput and cell- edge user data rates. Due to limitations of the backhaul communication and signal processing capability of base stations(BSs), the intrinsic problem of CoMP is that the selection of which BSs shall cooperate as only a few of BSs can be grouped in a cluster. However, approximating the theoretical performance bound of this clustering problem in CoMP at present is seldom discussed due to its inherent combinatorial complexity. In this paper, a novel efficient multi-cell clustering scheme based on blossom tree algorithm is proposed for cellular networks, incorporating CoMP with two cells in each cluster. With blossom tree algorithm, the proposed scheme can find out the optimal clustering strategy and help the CoMP transmission reach its theoretical performance bound on data rate in real-time computing(milliseconds in MATLAB simulation for one clustering). The simulation results show that our proposed method outperforms the existing dynamic greedy method in terms of cell edge users' average achievable data rate. Besides, it can also maintain high performance when extended to larger clusters in that with 4-cell clustering, the proposed method can reach 23.8% higher data rates than dynamic greedy method. Nanyang Ye 0001, Linhao Dong, Xiaoming Tao 0001, Ning Ge 0001 |
VTC Fall | 3 |
| 2015 | Risk-based adaptive metric learning for nearest neighbour classification
Yanan Miao, Xiaoming Tao 0001, Yipeng Sun, Yang Li 0005, Jianhua Lu |
Neurocomputing | 2 |
| 2015 | Hybrid model-and-object-based real-time conversational video coding
Yang Li 0005, Xiaoming Tao 0001, Jianhua Lu |
Signal Process. Image Commun. | 2 |
| 2015 | Robust 2D Principal Component Analysis: A Structured Sparsity Regularized ApproachabstractPrincipal component analysis (PCA) is widely used to extract features and reduce dimensionality in various computer vision and image/video processing tasks. Conventional approaches either lack robustness to outliers and corrupted data or are designed for one-dimensional signals. To address this problem, we propose a robust PCA model for two-dimensional images incorporating structured sparse priors, referred to as structured sparse 2D-PCA. This robust model considers the prior of structured and grouped pixel values in two dimensions. As the proposed formulation is jointly nonconvex and nonsmooth, which is difficult to tackle by joint optimization, we develop a two-stage alternating minimization approach to solve the problem. This approach iteratively learns the projection matrices by bidirectional decomposition and utilizes the proximal method to obtain the structured sparse outliers. By considering the structured sparsity prior, the proposed model becomes less sensitive to noisy data and outliers in two dimensions. Moreover, the computational cost indicates that the robust two-dimensional model is capable of processing quarter common intermediate format video in real time, as well as handling large-size images and videos, which is often intractable with other robust PCA approaches that involve image-to-vector conversion. Experimental results on robust face reconstruction, video background subtraction data set, and real-world videos show the effectiveness of the proposed model compared with conventional 2D-PCA and other robust PCA algorithms. Yipeng Sun, Xiaoming Tao 0001, Yang Li 0005, Jianhua Lu |
IEEE Trans. Image Process. | 2 |
| 2014 | A low-complexity Bayesian approach to large-scale sparse image reconstruction with structured constraintsabstractThe known tree-structure of wavelet transform coefficients is considered for conventional sparse image reconstruction to enhance performance. However, although existing Bayesian approaches can learn the latent structures, they have two issues that potentially limit their applications to high resolution images: First, treating the wavelet coefficients of large-scale images as high-dimensional vectors leads to impractical problem dimension; Second, Bayesian learning methods based on Markov chain Monte Carlo can guarantee global optima, but they require infinite stochastic samplings and thus are of high time complexity. In this paper, we address these issues by: 1) representing the wavelet transform images as multiple-measurement vectors (MMV) to reduce the memory complexity; and 2) modifying the structured priors for the Bayesian model to derive an analytical variational Bayesian strategy which can converge with much less time complexity. Experimental results demonstrate that our proposed method provides a practical alternative to large-scale sparse image recovery with low memory and computational requirements, while exhibiting close reconstruction accuracy compared to the time-consuming exact solutions. Shaoyang Li, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 2 |
| 2014 | QoS-Guaranteed Energy-Efficient Power Allocation in downlink multi-user MIMO-OFDM systemsabstractEnergy efficiency has currently become one of the central topics in today's wireless communication industry. In this paper, we focus on the energy-efficient design in downlink multi-user MIMO-OFDM systems, so as to optimize the energy efficiency with users' quality of service (QoS) guarantees. Particularly, an optimization problem concerning power allocation is formulated, where the objective is the energy efficiency measured by “Bit per Joule” and the constraints are users' QoS demands. Based on a mathematical equivalence and Lagrange duality, we propose an effective algorithm, named QoS-Guaranteed Energy-Efficient Power Allocation, to address the problem and achieve the optimal energy efficiency. Simulation results reveal that our proposed algorithm brings about remarkable gains on energy efficiency and simultaneously satisfies users' QoS requirements. Xiao Xiao 0004, Xiaoming Tao 0001, Jianhua Lu |
ICC | 2 |
| 2014 | On optimal relay selection and subcarrier assignment in OFDMA relay networks with QoS guaranteesabstractThis paper investigates QoS-aware relay selection and subcarrier assignment in cooperative OFDMA networks. In contrast to some existing works, which solve the sum-rate maximization problem directly, we first simplify it by exploiting the structure of the optimal solution, then we solve the simplified problem. To characterize the optimal network sum-rate performance, we solve the simplified problem optimally by the branch-and-cut method. In addition to the sum-rate maximization problem, we demonstrate that its two variants can also be similarly simplified and then be solved optimally by the branch-and-cut method. The optimal method serving as the benchmark is particularly suitable for moderate scale problems. Simulations show that the branch-and-cut method can locate the optimal solution effectively, which may in turn provide some insights into the performance of the heuristic method. Xiaoming Tao 0001, Yang Li 0005, Ning Ge 0001, Jianhua Lu |
ICC | 2 |
| 2014 | A Variational Bayesian EM Approach to Structured Sparse Signal ReconstructionabstractThis paper investigates a variational Bayesian expectation maximization (VBEM) scheme to reconstruct structured sparse signals. To fully exploit available signal priors, the structured sparsity combines sparsity prior and structure prior together. In contrast to recent studies, inferring the unobserved variables via Markov chain Monte Carlo (MCMC) which demands infinite iterations to converge, this work employs probabilistic graphic models (PGMs) to factorize the complex joint distributions of signal models and utilizes VBEM algorithm to optimize the lower bound of the model marginal likelihood. In this way, analytical posterior distributions of signal coefficients and model parameters can be obtained by a two-step iterative algorithm. The proposed method reduces computational time consumption with little reconstruction performance loss compared to long-time MCMC, and outperforms state-of-art recovery algorithms. Thus, it offers a practical and stable approach to large-scale Bayesian recovery applications. Shaoyang Li, Xiaoming Tao 0001, Jianhua Lu |
VTC Fall | 2 |
| 2014 | Dictionary Learning for Image Coding Based on Multisample Sparse RepresentationabstractIn this brief we propose a multisample sparse representation (MSR)-based online dictionary-learning approach to encode images more efficiently. To minimize the reconstructed error while handling a variety of image samples, we develop a multisample sparse representation method capable of obtaining sparser coefficients combined with learning dictionaries on-the-fly. With a well-learned dictionary, we further derive an MSR-based image coding approach to encode the quantized sparse coefficients with reduced reconstructed errors. Experimental results demonstrate rapid convergence of the proposed dictionary-learning algorithm and improved rate-distortion performance over other competitive image compression schemes both subjectively and quantitatively, validating the effectiveness of the proposed approach. Yipeng Sun, Xiaoming Tao 0001, Yang Li 0005, Jianhua Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | QoS provisioning scheduling with joint optimization of base station and relay power allocation in cooperative OFDMA systemsabstractMost existing works on scheduling for cooperative OFDMA systems have focused on homogeneous users with the same service and demand. Although scheduling for service differentiated cooperative OFDMA systems have been discussed in some previous works, their schedule problems do not take the joint optimization of base-station and relay power into account. In this paper, we investigate the optimization problem for joint base-station and relay power allocation, relay selection and subcarrier assignment with QoS guarantees and service support. We formulate the QoS provisioning joint schedule problem as a Mixed Binary Integer Program (MBIP). By introducing power and QoS prices, this MBIP is transformed into its corresponding dual problem, and a two-level dual decomposition method is proposed to solve the problem. Simulation results demonstrate that our proposed method outperforms previous works in terms of spectrum efficiency and QoS satisfaction. Xiaoming Tao 0001, Yang Li 0005, Jianhua Lu |
ICC | 2 |
| 2013 | Robust two-dimensional principal component analysis via alternating optimizationabstractTo extract two-dimensional principal components from image samples while being insensitive to outliers, we propose a robust model for two-dimensional principal component analysis (robust 2D-PCA) by regularizing sparse penalty term. Moveover, we develop a novel iterative algorithm for robust 2D-PCA via alternating optimization, learning the projection matrices by bi-directional decomposition. To further speed up the iteration, we develop an alternating greedy approach, minimizing over the low-dimensional feature matrix and the sparse error matrix. Experimental results on dynamic background subtraction are evaluated to show the effectiveness of the proposed model, compared with conventional 2D-PCA and robust PCA algorithms. Yipeng Sun, Xiaoming Tao 0001, Yang Li 0005, Jianhua Lu |
ICIP | 2 |
| 2013 | Information Theory Analysis of Blind Detection for PCMA Satellite Communication SystemsabstractPaired Carrier Multiple Access (PCMA) is widely used in bandwidth limited satellite network systems for its high frequency efficiency and compatibility with existing communication methods. With blind detection for PCMA signal being an important topic, various blind detection methods for PCMA signals with specific properties have been proposed. However, information theory analysis is still required for general blind detection method design. This paper introduces information theoretical bound for blind detection using a simulation based computation method, and applies a Viterbi detection method to verify the bound. Given a PCMA signal, the information theoretical bound helps to evaluate whether blind detection is possible, and guides how to design specific blind detection methods efficiently. Simulation result shows how mutual information carried by PCMA signal of the communicating peers is influenced by signal fading, propagation delay and other parameters numerically. Xijia Liu, Xiaoming Tao 0001, Xiang Chen 0007, Ning Ge 0001 |
VTC Fall | 2 |
| 2012 | Energy efficiency optimization in multi-user cellular systems with radio resource constraintsabstractEnergy efficiency has currently become one of the central issues in today's wireless communication industry. In this paper, we focus on the optimal energy efficiency in a multiuser cellular system and investigate its achievable upper bound from the perspective of radio resource available. Particularly, an optimization problem is formulated concerning radio resource scheduling, and an analytical optimal solution is obtained via mathematical analysis. Furthermore, an energy-efficient radio resource scheduling (EE-RRS) algorithm is proposed for numerical applications. Simulation results show that our proposed EE-RRS algorithm can realize significant improvements on energy efficiency with the optimal solution, and there is always a tradeoff between energy efficiency and transmission capacity. Xiao Xiao 0004, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 2 |
| 2012 | QoS Aware Scheduling with Optimization of Base Station Power Allocation in Downlink Cooperative OFDMA SystemsabstractQuality-of-Service (QoS) aware scheduling in cooperative orthogonal frequency division multiple access (OFDMA) systems has long been a critical yet challenging research topic for achieving full system utilization. The optimal performance of such systems can only be achieved by joint optimization of various resources (power, bandwidth and etc). In this paper, we investigate the optimization of base station power allocation with QoS requirements in cooperative OFDMA systems. We formulate the QoS aware resource allocation problem and exploit its structure to derive an efficient solution. By employing Lagrangian relaxation, the problem decouples on each subcarrier into a hierarchy of subproblems. Combined with an iterative procedure, some separable and parallel structures could be exploited to obtain an efficient solution. Simulation results demonstrate that our proposed method outperforms previous works in terms of spectrum efficiency and QoS satisfaction. Xiaoming Tao 0001, Jianhua Lu |
VTC Fall | 2 |
| 2011 | Resource Allocation for Layered Multicast Streaming in Wireless OFDMA NetworksabstractIn this paper, we focus on subcarrier and power allocation for layered multicast streams in OFDMA cellular networks, where the multicast stream is composed of a basic layer and an enhancement layer. Our goal is to maximize the system total throughput with a total power constraint and a minimum rate requirement. A low-complexity allocation algorithm is proposed, which combined the suitability-based subcarrier allocation (SSA) with the traditional water filling (TWF) for the base layer and an advanced water filling (AWF) for the enhancement layer. Besides, we present a throughput-based user selection (TUS) algorithm which selects a proper set of users to serve when the minimum rate requirement can not be satisfied. Simulation results show that our proposed algorithm can improve the system throughput and outage probability. Xiaoming Tao 0001, Tengfei Xing, Jianhua Lu |
ICC | 2 |
| 2011 | QoS-Aware Resource Allocation for Mixed Multicast and Unicast Traffic in OFDMA NetworksabstractThis paper focuses on the subchannel and power allocation for mixed multicast and unicast traffic in wireless OFDMA networks, where the multicast data is divided into basic layer and enhancement layer data. Our goal is to maximize the network total throughput with a total power constraint while guaranteeing the minimum rate requirement of both the unicast and muticast traffic. A heuristic allocation algorithm is proposed, which combines a cost-based subchannel allocation (CSA) with the traditional water filling (TWF) and an advanced water filling (AWF). The TWF is used for the subchannels allocated to satisfy the rate requirements of each unicast traffic and multicast traffic while the AWF is used for the remaining subchannels. Simulation results show that our proposed algorithm can improve the network throughput and outage probability compared with other algorithms. Xiaoming Tao 0001, Jianhua Lu |
VTC Fall | 2 |
| 2011 | A QoS-Aware Power Optimization Scheme in OFDMA Systems with Integrated Device-to-Device (D2D) CommunicationsabstractThis paper proposes a power optimization scheme with joint resource allocation (i.e. subcarrier and bit allocation) and mode selection in an OFDMA system with integrated D2D communications. Through the proper control of the base station (BS), users can communicate with each other either directly or via the BSs as in traditional cellular networks. Particularly, an optimization problem is formulated to minimize total downlink transmission power constrained by users' QoS demands; while a heuristic scheme exploiting joint subcarrier allocation, adaptive modulation and mode selection is contrived to solve the problem. Simulation results show that our proposed scheme may not only conserve total downlink transmission power effectively, but also save overall power consumption of BSs significantly, compared with existing algorithms used in traditional OFDMA systems. Xiao Xiao 0004, Xiaoming Tao 0001, Jianhua Lu |
VTC Fall | 2 |
| 2011 | An energy-efficient hybrid structure with resource allocation in OFDMA networksabstractEnergy consumption of information and communication industry has generated increasing concerns recently, aiming at reducing power consumption or increasing energy efficiency. This paper proposes an energy efficient hybrid structure which introduces micro sites into traditional OFDMA cellular systems with only central macro base stations (BSs). The deployment of small and low power BSs alongside the central macro BSs is capable of advancing radio coverage, enhancing system capacity, and simultaneously increasing energy efficiency. Upon the proposed structure, two efficient solutions respectively based on Lagrangian dual decomposition (LDD) and serial carrier and power allocation (SCPA), are proposed to optimize energy efficiency. Simulation results show that our proposed hybrid structure can dramatically increase the energy efficiency and the downlink system capacity in OFDMA cellular systems. Xiao Xiao 0004, Xiaoming Tao 0001, Jianhua Lu |
WCNC | 2 |
| 2008 | Adaptive SD-OFDM in Time-Frequency Selective Fading ChannelabstractSD-OFDM has been proposed for information transmission with high spectrum efficiency. In this paper, an adaptive SD-OFDM (ASD-OFDM) is proposed for time- frequency selective fading channel. Compared with traditional OFDM, the ASD-OFDM improves the performance, i. e., spectrum efficiency and interference cancelation capability, in the high speed mobile scenario as well as short-delay multipath environment. Xiaoming Tao 0001, Jianhua Lu |
ICC | 1 |