Shuoyao Wang

dblp:200/8365 · DBLP profile ↗
← Back
44ranked-venue papers
14as first author
41since 2021 · last 2026
0000-0003-1395-4383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 26 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive-Smooth LiDAR-Camera Knowledge Distillation with Heterogeneous Fusion for Multi-View 3D Object Detection
abstract
Multi-view 3D object detection has garnered increasing attention, particularly due to its success in autonomous driving systems. Although multi-view systems possess rich semantic information, their spatial-geometric reasoning capabilities remain limited. Recent studies employ simulated point cloud generation mechanisms to facilitate LiDAR-camera multi-modal knowledge distillation, achieving formal structural consistency. Despite advancements, these methods still face two main issues: i) alignment challenges caused by discrepancies between LiDAR and camera data, and ii) prediction errors from simulated point clouds that compromise the semantic information extracted from images during fusion. To address these problems, we propose adaptive-smooth distillation to optimize alignment granularity based on feature discrepancies for improved LiDAR-camera knowledge distillation. Specifically, this work considers both LIDAR-to-camera cross-modal distillation and LiDAR-camera fusion to simulated point cloud-camera fusion multi-modal distillation. Then, we introduce a heterogeneous fusion module to strategically bias the fusion process toward the extracted camera features, thereby enhancing the robustness of the fusion feature. Additionally, soft-weighted response distillation is proposed to facilitate the student model to selectively mimic the high-quality output of the teacher model. Extensive experiments have demonstrated the superiority of our method, achieving statistically significant improvements of 4.9% in mean Average Precision (mAP) and 4.5% in NuScenes Detection Score (NDS) over the benchmark.
Rui Zhao 0029, Shuoyao Wang, Xinhu Zheng, Shijian Gao
AAAI2
2026 Adaptive Video Streaming in Heterogeneous Wireless Networks with Mixture of Experts
Shuoyao Wang, Xiaowen Cao 0001
ICC1
2026 An Immersive Image Transmission System with Mutual Information based Feature Separation
Chonghua Yang, Shuoyao Wang, Suzhi Bi
ICC2
2026 Generalizing Adaptive Video Streaming With Mixture of Experts in Heterogeneous Wireless Networks
abstract
Adaptive video streaming has become a core technology for modern video delivery, particularly in cellular networks. However, the growing dynamics of mobile environments and the diversity of user preferences present major challenges for adaptive bitrate (ABR) algorithms. Existing approaches often struggle to maintain a balance between high in-distribution performance and strong generalization across heterogeneous network conditions and personalized Quality of Experience (QoE) demands. To address these challenges, we propose NMoEABR, a unified ABR decision-making framework that integrates a nonlinear Mixtureof- Experts (NMoE) architecture with preference-aware metareinforcement learning. Specifically, we design an NMoE-based actor network that adaptively aggregates expert policies through dynamic convolution conditioned on real-time network states, thereby enhancing robustness and cross-network generalization in a zero-hot manner. Furthermore, to mitigate convergence difficulties arising from the joint optimization of expert policies and expert-weight prediction, we introduce a preference-aware meta-RL strategy that incorporates user preference embeddings and virtual preference synthesis to stabilize meta-policy updates. Comprehensive evaluations on real-world traces and wireless testbed demonstrate that NMoEABR consistently outperforms mainstream ABR benchmarks in terms of average QoE, stability, and adaptability, particularly under unseen network conditions and diverse user preference distributions.
Shuoyao Wang, Xiaowen Cao 0001, Lifeng Xie
IEEE Trans. Mob. Comput.1
2026 Digital Semantic Communications: An Alternating Multi-Phase Training Strategy With Mask Attack
abstract
Semantic communication (SemComm) has emerged as new paradigm shifts. Most existing SemComm systems transmit continuously distributed signals in analog fashion. However, the analog paradigm is not compatible with current digital communication frameworks. In this paper, we propose an alternating multi-phase training strategy (AMP) to enable the joint training of the networks in the encoder and decoder through non-differentiable digital processes. AMP contains three training phases, aiming at feature extraction (FE), robustness enhancement (RE), and training-testing alignment (TTA), respectively. In particular, in the FE stage, we learn the representation ability of semantic information by jointly training the encoder and decoder in an analog manner. When we take digital communication into consideration, the domain shift between digital and analog demands the fine-tuning for encoder and decoder. To cope with joint training process within the non-differentiable digital processes, we propose the alternation between updating the decoder individually and jointly training the codec in RE phase. To boost robustness further, we investigate a mask-attack (MATK) in RE to simulate an evident and severe bit-flipping effect in a differentiable manner. To address the training-testing inconsistency introduced by MATK, we employ an additional TTA phase, fine-tuning the decoder without MATK. Combining with AMP and an information restoration network, we propose a digital joint source-channel coding system for image transmission, named AMP-SC1. Comparing with the representative benchmark, AMP-SC achieves 0.82 ~ 1.65dB higher average reconstruction performance among several representative datasets at different scales and a wide range of signal-to-noise ratios.
Mingze Gong, Shuoyao Wang, Suzhi Bi, Yuan Wu 0001, Li Ping Qian 0001
IEEE Trans. Wirel. Commun.2
2026 Near Field Sparse Representation and Dictionary Design Using Discrete Fresnel Transform
abstract
Extremely large-scale antenna arrays (ELAAs) and millimeter wave (mmWave) communication are prominent enablers in advanced 6G networks. To accommodate the unique characteristics introduced by these trends in communication system design, near-field spherical-wave propagation modeling is imperative, departing from the classical far-field planar-wave model. In the millimeter wave band, where wavelengths are small compared to the array aperture, many propagation phenomena resemble those in optics. Inspired by Fresnel diffraction analysis in classical optics, we explore the utilization of the Fresnel Transform for channel modeling, especially in the near-field region. In this paper, we propose a Discrete Fresnel Transform (DFnT)-based dictionary towards effectively characterizing and sparsely representing the channel vector in both near and far-field scenarios. Specifically, we bridge the gap between the continuous Fresnel Transform and the near-field DFnT dictionary by discretizing and parameterizing the transform, adapting it to ELAA codewords. We assess the near-field representation capability of the dictionary by employing it for channel estimation using Orthogonal Matching Pursuit, a compressive channel estimation algorithm whose performance relies heavily on the sparsity of the representation. Simulation results underscore the superiority of the DFnT dictionary over the existing far-field and the Polar dictionaries.
Shuoyao Wang, Ying-Jun Angela Zhang
IEEE Trans. Wirel. Commun.2
2025 JS3C:Joint Source-Channel-Check Coding for Reliable Semantic Communication
abstract
Semantic communication, as a novel paradigm for enhancing transmission quality and data efficiency, has attracted significant attention in recent years. However, most existing studies focus on improving average performance, neglecting to ensure reliability. Taking reliability into consideration, recent researches have introduced retransmission strategies based on semantic error detection (SED). However, the retransmission signals typically focus on adding redundant data to improve reliability, sacrificing limited retransmission efficiency. Moreover, existing SED typical introduce additional overhead, where efficient SED remains an open challenge. To address these issues, we propose JS3C, a Joint Source-Channel-Check Coding framework for efficient and reliable semantic communications. Specifically, we propose a fidelity-aware check coding scheme where the semantic information embedded in error detection codes is also reused to enhance semantic reconstruction, thereby improving coding efficiency. Furthermore, we incorporate quality estimation feedback into the retransmission encoder to balance redundancy and refinement information, improving retransmission efficiency. Experimental results demonstrate that, compared with existing harq based semantic communication system, the proposed JS3C scheme a 2.75 dB improvement in 95th percentile PSNR.
Mingze Gong, Shuoyao Wang
GLOBECOM3
2025 Recovery Condition-Oriented Sensing Matrix Optimization for Block-Sparse Near-Field Channel Estimation
abstract
Millimeter wave channels exhibit sparsity under certain basis in the far-field, such as Discrete Fourier Transform (DFT) basis. This property enables the use of compressed sensing (CS) techniques to reconstruct channel vectors with reduced pilot overhead. However, with increasing antenna array sizes and higher frequencies, users more frequently operate in the radiating near-field, where spherical wavefronts dominate. In this regime, the DFT basis no longer yields exact sparsity. Instead, the channel exhibits a block-sparse structure, where non-zero elements tend to cluster in the angular domain. To exploit this structure, we propose a block sparsity-aware CS approach for near-field channel estimation. Specifically, we derive the block sparse signal recovery condition based on block coherence measurements. Based on the condition, we adopt a mixed l2/l1norm-based optimization that extends the conventional l1-based methods by explicitly modeling block sparsity of the near-field channel. Additionally, we design an optimized sensing matrix that outperforms conventional random sensing matrices. This method effectively exploits the block structure of non-zero elements, leading to more accurate channel estimation in both near-field and far-field regions.
Shuoyao Wang, Ying-Jun Angela Zhang
GLOBECOM2
2025 Low-Rate Semantic Communication with Codebook-Based Conditional Generative Models
abstract
Generative semantic communication models are reshaping semantic communication frameworks by moving beyond pixel-wise optimization to align with human perception. However, many existing approaches prioritize image-level perceptual quality, often neglecting alignment with downstream tasks, which can lead to suboptimal semantic representation. This paper introduces an Ultra-Low Bitrate Semantic Communication (ULBSC) system that employs a conditional generative model and a learnable condition codebook. By integrating saliency conditions and image-level semantic information, the proposed method enables high-perceptual-quality and controllable task-oriented image transmission. Recognizing shared patterns among objects, we propose a codebook-assisted condition transmission method, integrated with joint source-channel coding (JSCC)-based text transmission to establish ULBSC. The codebook serves as a knowledge base, reducing communication costs to achieve ultra-low bitrate while enhancing robustness against noise and inaccuracies in saliency detection. Simulation results indicate that, under ultra-low bitrate conditions with an average compression ratio of 0.57 %, the proposed system delivers superior visual quality compared to traditional JSCC techniques and achieves higher saliency similarity between the generated and source images compared to state-of-the-art generative semantic communication methods.
Kailang Ye, Mingze Gong, Shuoyao Wang, Daquan Feng
VTC2025-Spring3
2025 Predictive Target-to-User Association in Complex Scenarios via Hybrid-Field ISAC Signaling
abstract
This paper presents a novel and robust target-to-user (T2U) association framework to support reliable vehicle-to-infrastructure (V2I) networks that potentially operate within the hybrid field (near-field and far-field). To address the challenges posed by complex vehicle maneuvers and user association ambiguity, an interacting multiple-model filtering scheme is developed, which combines coordinated turn and constant velocity models for predictive beamforming. Building upon this foundation, a lightweight association scheme leverages user-specific integrated sensing and communication (ISAC) signaling while employing probabilistic data association to manage clutter measurements in dense traffic. Numerical results validate that the proposed framework significantly outperforms conventional methods in terms of both tracking accuracy and association reliability.
Yifeng Yuan, Miaowen Wen, Xinhu Zheng, Shuoyao Wang, Shijian Gao
VTC2025-Spring4
2025 Adaptive 360-Degree Streaming: Optimizing With Multi-Window and Stochastic Viewport Prediction
abstract
The tile-based approach is widely adopted in adaptive 360-degree video streaming systems, due to its efficiency in managing limited bandwidth resources. Recently, significant research efforts have been devoted to viewport-prediction-enabled bitrate adaptation for tile-based 360-degree Adaptive Bit-Rate (ABR) streaming, towards improving the average video quality while reducing rebuffering. However, the inherent uncertainty of users’ viewports has posed limitations on users’ Quality of Experience (QoE) for tile-based 360-degree ABR streaming. In this paper, we introduce a multi-window and stochastic viewport prediction approach to address the viewport uncertainty. In particular, considering our goal of maximizing the expectation of future QoE, we investigate a viewport distribution prediction model, to cope with the inherent randomness. Additionally, to accommodate the varying gap between the playback and the download process, we explore the multiple-window viewport prediction models to capture different prediction gaps. Even with the utilization of distributional prediction and multi-window models, predicting viewports far into the future is still inherently challenging. Accordingly, we propose a patience pattern temporarily suspending the download process, allowing for the accumulation of additional head movement trajectory data. Finally, we employ a model predictive control (MPC) approach for sequential decision-making, formulating the MPC problem as a mixed-integer non-linear programming (MINLP) task. To mitigate the computational burden associated with solving MINLP, we introduce a mixed-integer linear programming transformation to achieve efficient decision-making. Extensive experiments, utilizing real-world traces and user head movement trajectories, demonstrate that the proposed method outperforms state-of-the-art methods, improving overall QoE performance by 16.75% –18.91% .
Weichao Feng, Shuoyao Wang
IEEE Trans. Mob. Comput.2
2025 Geometric Continuity and Consistency Learning for Self-Supervised Point Cloud Completion
abstract
Point cloud completion aims to infer the complete point clouds from incomplete ones. In real-world scenarios, where the paired data is absent, self-supervised methods have emerged as a promising solution. Although existing self-supervised methods perform well at relatively low resolutions, they suffer significant performance degradation at higher resolution primarily because they focus on point cloud reconstruction at patch-level or point-level. In this paper, we propose a self-supervised method based on Geometric Continuity and Consistency Learning (GCCL) at multi-scale level to improve the accuracy of predicting local details and global shapes of point clouds. Specifically, to capture local details, we employ a patch-topoint strategy and a coarse-fine manner for geometric continuity learning. To constrain the global shapes, we construct multiple branches for mutual supervision and utilize class priors to build a memory queue for contrasting current features, enhancing the network focus on geometric consistency learning. We evaluate GCCL on multiple datasets, and the results show that our method outperforms existing self-supervised methods by a 4.4 improvement in CD-$\ell_2$on the synthetic PCN dataset and can generate more uniformly distributed completion results on realworld datasets
Junkang Ma, Shuoyao Wang, Xiaochun Mai
IEEE Trans. Multim.2
2025 Scalable Multi-Task Edge Sensing via Task-Oriented Joint Information Gathering and Broadcast
abstract
The recent advance of edge computing technology enables significant sensing performance improvement of Internet of Things (IoT) networks. In particular, an edge server (ES) is responsible for gathering sensing data from distributed sensing devices, and immediately executing different sensing tasks to accommodate the heterogeneous service demands of mobile users. However, as the number of users surges and the sensing tasks become increasingly compute-intensive, the huge amount of computation workloads and data transmissions may overwhelm the edge system of limited resources. Accordingly, we propose in this paper a scalable edge sensing framework for multi-task execution, in the sense that the computation workload and communication overhead of the ES do not increase with the number of downstream users or tasks. By exploiting the task-relevant correlations, the proposed scheme implements a unified encoder at the ES, which produces a common low-dimensional message from the sensing data and broadcasts it to all users to execute their individual tasks. To achieve high sensing accuracy, we extend the well-known information bottleneck theory to a multi-task scenario to jointly optimize the information gathering and broadcast processes. We also develop an efficient two-step training procedure to optimize the parameters of the neural network-based codecs deployed in the edge sensing system. Experiment results show that the proposed scheme significantly outperforms the considered representative benchmark methods in multi-task inference accuracy. Besides, the proposed scheme is scalable to the network size, which maintains almost constant computation delay with less than 1% degradation of inference performance when the user number increases by four times.
Huawei Hou, Suzhi Bi, Xian Li 0005, Shuoyao Wang, Li Ping Qian 0001, Zhi Quan
IEEE Trans. Wirel. Commun.4
2025 Scenario-Adaptive Meta-Learning for mmWave Beam Alignment
abstract
In millimeter wave communication systems, achieving high-quality data transmission demands efficient and rapid beam alignment. Conventional deep learning-based methods, although promising, often rely on the assumption that training and testing channels share identical distribution. This assumption may not hold in practical settings, potentially leading to significant performance degradation when the deployment environment changes. To address this issue, we introduce SAMBA, a novel meta-learning-based approach for adaptive beam alignment without requiring Channel State Information (CSI). SAMBA enables swift adaptation to unknown scenarios using a minimal set of newly labeled data. Specifically, we adopt a probing beam-search strategy to obviate the need for CSI. Furthermore, we employ Model-Agnostic Meta-Learning (MAML) for parameter pre-training and fine-tuning to enhance our model’s adaptability. Confronted with the challenge of numerous beam candidates in the narrow beam selection problem, which complicates the straightforward replication of MAML, we develop a novel training task generation strategy. In our experimental assessments, we subjected SAMBA to a wide range of challenging scenarios using ray-tracing simulations. These scenarios encompassed various frequency bands, distinct base station layouts, and outdoor-to-indoor transitions. Our results demonstrate that SAMBA consistently outperforms learning-based baseline models, showcasing its superior domain adaptation capabilities in dynamic and diverse channel settings.
Shuoyao Wang, Ying-Jun Angela Zhang
IEEE Trans. Wirel. Commun.2
2024 Revisiting Domain-Adaptive Object Detection in Adverse Weather by the Generation and Composition of High-Quality Pseudo-labels
Rui Zhao 0029, Huibin Yan, Shuoyao Wang
ECCV (36)3
2024 Adaptive 360-Degree Video Streaming with Multi-window and Stochastic Viewport Prediction
abstract
The tile-based approach is widely adopted in adaptive 360-degree video streaming systems, due to its efficiency in managing limited bandwidth resources. However, the inherent uncertainty of users' viewports, making viewport prediction difficult, has posed limitations on the performance of tile-based bitrate adaptive streaming. In this paper, we introduce a Multi-window and Stochastic Viewport Prediction approach to address the uncertainty in viewport changes. In particular, considering our goal of maximizing the expectation of future Quality of Experience (QoE), we investigate a viewport distribution prediction model, to cope with the inherent randomness. Additionally, to accommodate the varying gap between the playback and the downloading process, we explore the multiple-windows viewport prediction model to capture different prediction windows. Even with the utilization of distributional prediction and multi-window models, predicting viewports far into the future is still inherently challenging. Accordingly, we propose a patience pattern temporarily suspending the download process, allowing for the accumulation of additional head movement trajectory data. Finally, within the framework of Model Prediction Control (MPC)-based rate adaptation, we introduce a multi-window Probabilistic view-port prediction-based Tile-level Bitrate Adaptation Streaming (PTBAS) algorithm. Extensive experiments, utilizing real-world traces and user head movement trajectories, demonstrate that PTBAS outperforms state-of-the-art methods, improving overall QoE performance by 15.96%.
Weichao Feng, Shuoyao Wang
ICC3
2024 DFMDA-Net: Dense Fusion and Multi-dimension Aggregation Network for Image Restoration
Huibin Yan, Shuoyao Wang
IJCAI2
2024 Beyond Profit: A Multi-Objective Framework for Electric Vehicle Charging Station Operations
abstract
This paper explores the pricing and scheduling strategies of the electric vehicle charging stations in response to the rising demand for cleaner transportation. Most of the existing methods focus on maximizing the energy efficiency or the charging station profit, however, the reputation of EVs is also a key factor for the long-term charging station operations. To address these gaps, we propose a novel framework for jointly optimizing pricing and continuous-multiple charging rates. Our approach aims to maximize both charging station profit and reputation, considering multi-objective optimization and continuous rate control within physical constraints. Introducing a pricing fluctuating penalty for reputation modeling and a linear programming-based safe layer for constraints, we confront the complexity of continuous charging rates' action space. To enhance convergence, we explore a soft action critic framework with novel entropy temperature tunning technique. The experiments conducted with real data demonstrate that the proposed method can provide extra 25.45%-52.20% average JPR than the representative baselines.
Shuoyao Wang
VTC Spring1
2024 A Scalable Re-Ranking Optimization Approach for Real-Time Network-Friendly Recommendations
abstract
Network-friendly recommendation (NFR) has emerged as a promising method to enhance network performance. However, current NFR methods incur significant computational overhead and predominantly focus on content recommendations, often neglecting the importance of content ranking, which can limit recommendation quality. This article addresses these challenges by proposing a recommendation re-ranking algorithm designed to maximize ranking-aware recommendation quality while accommodating diverse network scenarios and conditions. We formulate the problem as an integer programming (IP) problem that maximizes the mean reciprocal rank of the baseline recommendation under network constraints. To tackle the computational challenges posed by large-scale IP problems, we propose a low-complexity algorithm that enables real-time NFR by re-ranking only the candidate list instead of the entire content list. Results on the real-world data sets manifest that compared with the state-of-the-art (SOTA) NFR baseline, our method achieved a 40.3% increase in cache hit rate (CHR) and a 25.32% reduction in transmission latency. Additionally, the recommendation quality is further enhanced, with improvements of more than 3.5% in CHR and 2.62% in latency reduction scenarios. Moreover, computational time is reduced by over 97%.
Jiayin Hou, Shuoyao Wang
IEEE Internet Things J.3
2024 Cooperative Charging Stations Management Under Irrational Hierarchy EV Behaviors
abstract
The Internet of Things (IoT) technology connects various aspects of society and enhances human life. Electric vehicles (EVs) and charging stations (CSs) are essential components of the IoT system, offering business opportunities for companies like the aggregator. The aggregator can maximize profit by implementing various CSs management strategies, such as pricing and charging scheduling. However, managing CSs presents challenges due to irrational human behavior, particularly uncertain CS selections by EV users. To address this issue, we propose a cooperative model for CSs that incorporates the cognitive hierarchy quantum response CS selection model. In this model, EV users are considered to possess${k}$-level rationality, and the charging environment is formulated as a Markov decision process (MDP) based on this user model. To make optimal decisions for CSs from the MDP, we propose a denoising autoencoder-based deep reinforcement learning (DEEDRL) method. This method learns from the uncertain environment and effectively denoises the state space using a pretrained autoencoder. In addition, to reduce the computational burden caused by time-varying strategies, we design a discretization strategy for action space based on the current market rule of tiered pricing and CS types. Our experiments with real data demonstrate that our proposed method accurately portrays the CS selection behavior of irrational users in realistic scenarios. Furthermore, our method outperforms noncooperative modes and benchmark cooperative algorithms regarding profitability, such as DDPG and DQN.
Jie Liu 0061, Shuoyao Wang, Xiaoying Tang 0002
IEEE Internet Things J.2
2024 Image Intrinsic Components Guided Conditional Diffusion Model for Low-Light Image Enhancement
abstract
Through formulating the image restoration as a generation problem, the conditional diffusion model has been applied to low-light image enhancement (LIE) to restore the details in dark regions. However, in the previous diffusion model based LIE methods, the conditions used for guiding generation are degraded images, such as low-light image, signal-to-noise ratio map and color map, which suffer from severe degradation and are simply fed into diffusion model by rigidly concatenating with the noise. To avoid using degraded conditions resulting in sub-optimal performance in recovering details and enhancing brightness, we use the image intrinsic components originating from the Retinex model as guidance, whose multi-scale features are flexibly integrated into the diffusion model, and propose a novel conditional diffusion model for LIE. Specifically, the input low-light image is decomposed into reflectance and illumination by a Retinex decomposition module, where two components contain abundant physical property and lighting conditions of the scene. Then, we extract the latent features from two conditions through a component-dependent feature extraction module, which is designed according to the physical property of components. Finally, instead of previous rigid concatenation manner, a well-designed feature fusion mechanism is equipped to adaptively embed generative conditions into diffusion model. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods, and is capable of effectively restoring the local details while brightening the dark regions. Our codes are available athttps://github.com/Knossosc/ICCDiff.
Sicong Kang, Shuaibo Gao, Wenhui Wu 0001, Xu Wang 0006, Shuoyao Wang, Guoping Qiu
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Two-Stage Deep Reinforcement Learning Framework for MEC-Enabled Adaptive 360-Degree Video Streaming
abstract
The emerging multi-access edge computing (MEC) technology effectively enhances the wireless streaming performance of 360-degree videos. By connecting a user's head-mounted device (HMD) to a smart MEC platform, the edge server (ES) can efficiently perform adaptive tile-based video streaming to improve the user's viewing experience. Under constrained wireless channel capacity, the ES can predict the user's field of view (FoV) and transmit to the HMD high-resolution video tiles only within the predicted FoV. In practice, the video streaming performance is challenged by the random FoV prediction error and wireless channel fading effects. For this, we propose in this paper a novel two-stage adaptive 360-degree video streaming scheme that maximizes the user's quality of experience (QoE) to attain stable and high-resolution video playback. Specifically, we divide the video file into groups of pictures (GOPs) of fixed playback interval, where each GOP consists of a number of video frames. At the beginning of each GOP (i.e., the inter-GOP stage), the ES predicts the FoV of the next GOP and allocates an encoding bitrate for transmitting (precaching) the video tiles within the predicted FoV. Then, during the real-time video playback of the current GOP (i.e., the intra-GOP stage), the ES observes the user's true FoV of each frame and transmits the missing tiles to compensate for the FoV prediction errors. To maximize the user's QoE under random variations of FoV and wireless channel, we propose a double-agent deep reinforcement learning framework, where the two agents operate in different time scales to decide the bitrates of inter- and intra-GOP stages, respectively. Experiments based on real-world measurements show that the proposed scheme can effectively mitigate FoV prediction errors and maintain stable QoE performance under different scenarios, achieving over 22.1% higher QoE than some representative benchmark methods.
Suzhi Bi, Haoguo Chen, Xian Li 0005, Shuoyao Wang, Yuan Wu 0001, Li Ping Qian 0001
IEEE Trans. Mob. Comput.4
2024 Imitation Learning for Adaptive Video Streaming With Future Adversarial Information Bottleneck Principle
abstract
Adaptive video streaming plays a crucial role in ensuring high-quality video streaming services. Despite extensive research efforts devoted to Adaptive BitRate (ABR) techniques, the current reinforcement learning (RL)-based ABR algorithms may benefit the average Quality of Experience (QoE) but suffers from fluctuating performance in individual video sessions. In this paper, we present a novel approach that combines imitation learning with the information bottleneck technique, to learn from the complex offline optimal scenario rather than inefficient exploration. In particular, we leverage the deterministic offline bitrate optimization problem with the future throughput realization as the expert and formulate it as a mixed-integer non-linear programming (MINLP) problem. To enable large-scale training for improved performance, we propose an alternative optimization algorithm that efficiently solves the formulated MINLP problem. To address the overfitting issues due to the future information leakage in MINLP, we incorporate an adversarial information bottleneck framework. By compressing the video streaming state into a latent space, we retain only action-relevant information. Additionally, we introduce a future adversarial term to mitigate the influence of future information leakage, where Model Prediction Control (MPC) policy without any future information is employed as the adverse expert. Experimental results demonstrate the effectiveness of our proposed approach in significantly enhancing the quality of adaptive video streaming, providing a 7.30% average QoE improvement and a 30.01% average ranking reduction.
Shuoyao Wang, Fangwei Ye
IEEE Trans. Mob. Comput.1
2024 Adaptive Video Streaming in Multi-Tier Computing Networks: Joint Edge Transcoding and Client Enhancement
abstract
With the advancement of multimedia technology and wireless networks, there is a growing demand for high-quality video streaming. Delivering stable video streaming in extremely dynamic wireless networks, nevertheless, is still an open problem. Recent developments in client computing and mobile edge computing (MEC) technologies have both shown promise in enhancing the adaptive bitrate (ABR) streaming services. In this paper, we consider a video streaming system in multi-tier computing networks, enabled by joint edge-side video transcoding and client-side video enhancement. By “enhancement,” we mean that the client improves the video chunk quality via client-side image processing modules. In particular, we aim to design a joint bitrate adaptation, edge transcoding, and client image-processing algorithm, maximizing the quality of experience (QoE) of streaming services. The majority of the prior art has concentrated on super-resolution-enabled video streaming. Contrarily, we show that the video enhancement method outperforms the super-resolution approach in terms of signal-to-noise ratio and frames per second, implying a superior alternative for client-side processing in ABR streaming. We formulate the problem as an event-triggered Markov decision process (E-MDP), and propose a deep reinforcement learning (DRL)-based framework, named EDTEA. To deal with the delayed feedback induced by multi-tier computing, the entropy and the expected re-buffering terms are introduced to the objective and the reward, respectively. Extensive simulations based on real-world videos and bandwidth traces manifest that compared with state-of-the-art approaches, EDTEA provides$10.4\%\sim 78.4\%$extra QoE while reducing re-buffering time by$85.5\%\sim 91.7\%$.
Shuoyao Wang, Suzhi Bi
IEEE Trans. Mob. Comput.1
2024 MMVS: Enabling Robust Adaptive Video Streaming for Wildly Fluctuating and Heterogeneous Networks
abstract
With the advancement of wireless technology, the fifth-generation mobile communication network (5G) has the capability to provide exceptionally high bandwidth for supporting high-quality video streaming services. Nevertheless, this network exhibits substantial fluctuations, posing a significant challenge in ensuring the reliability of video streaming services. This research introduces a novel algorithm, the Multi-type data perception-based Meta-learning-enabled adaptive Video Streaming algorithm (MMVS), designed to adapt to diverse network conditions, encompassing 3G and mmWave 5G networks. The proposed algorithm integrates the proximal policy optimization technique with the meta-learning framework to cope with the gradient estimation noise in network fluctuation. To further improve the robustness of the algorithm, MMVS introduces meta advantage normalization. Additionally, MMVS treats network information as multiple types of input data, thus enabling the precise definition of distinct network structures for perceiving them accurately. The experimental results on network trace datasets in real-world scenarios illustrate that MMVS is capable of delivering an additional 6% average QoE in mmWave 5G network, and outperform the representative benchmarks in six pairs of heterogeneous networks and user preferences.
Shuoyao Wang
IEEE Trans. Multim.1
2024 Compression Before Fusion: Broadcast Semantic Communication System for Heterogeneous Tasks
abstract
Semantic communication has emerged as new paradigm shifts in 6G from the conventional syntax-oriented communications. Recently, the wireless broadcast technology has been introduced to support semantic communication system toward higher communication efficiency. Nevertheless, existing broadcast semantic communication systems target on general representation within one stage and fail to balance the inference accuracy among users. In this paper, the broadcast encoding process is decomposed into compression and fusion to improve communication efficiency with adaptation to tasks and channels. Particularly, we propose multiple task-channel-aware sub-encoders (TCEs) and a channel-aware feature fusion sub-encoder (CFE) towards compression and fusion, respectively. In TCEs, multiple local-channel-aware attention blocks are employed to extract and compress task-relevant information for each user. In GFE, we introduce a global-channel-aware fine-tuning block to merge these compressed task-relevant signals into a compact broadcast signal. Notably, we retrieve the bottleneck in DeepBroadcast and leverage information bottleneck theory to further optimize the parameter tuning of TCEs and CFE. We substantiate our approach through experiments on a range of heterogeneous tasks across various channels with additive white Gaussian noise (AWGN) channel, Rayleigh fading channel, and Rician fading channel. Simulation results evidence that the proposed DeepBroadcast outperforms the state-of-the-art methods.
Mingze Gong, Shuoyao Wang, Fangwei Ye, Suzhi Bi
IEEE Trans. Wirel. Commun.2
2023 A Scalable Multi-Device Semantic Communication System for Multi - Task Execution
abstract
Motivated by the success of deep learning, semantic communication has emerged as new paradigm shifts in 6G from the conventional data-oriented communications. However, the semantic communication systems suffer performance degradation at the receiver side or computation latency accumulation at the transmitter side, when serves multiple task execution. To address the issue, we develop a VisionTransformer based multi-device semantic communication system called MDSC to effectively perform multiple tasks. In particular, we introduce the shared semantic encoder to the transmitter to extract global semantic information, preventing for computation accumulation at the transmitter side. To cope with the fact that the semantic information for different task may differ from each other, we propose multiple-encoder-multiple-decoder channel codec architecture, improving the compression efficiency with task-specific codec. The compression efficiency leads to higher noise robustness and downstream task execution accuracy at the receiver side. In the experiments, we validate the proposed system with two tasks in NYUD-v2 and four tasks in PASCAL-Context, respectively. Compared with the state-of-the-art multi-task semantic communication system, MDSC achieves higher performance simultaneously for all tasks in both datasets.
Mingze Gong, Shuoyao Wang, Suzhi Bi
GLOBECOM2
2023 Network-Friendly Sequential Recommendation with Quality Constraints: A Safe Deep Reinforcement Learning Approach
abstract
Network-friendly recommendation have emerged as a promising approach to relieve data traffic congestion without sacrificing user preference. Most of existing works focus on one-stage recommendations, maximizing recommendation quality and reducing network latency for the next request. In this paper, we focus on policy optimization for Network-friendly Sequential Recommendation (NSR), towards maximizing the recommendation quality as well as network performance for the whole session with hard quality constraints. To achieve this goal, we first formulate the NSR problem as a Markov Decision Process (MDP) problem. To characterize the fundamental performance limit, we consider the offline solution by assuming the distributional knowledge of user behavior is known as a prior. In this case, we solve the offline problem through policy iteration. However, user behavior in real-world scenarios is unpredictable, which makes it difficult to know the distributional knowledge of users. To handle this issue, we propose a proximal policy optimization-based algorithm with a safe layer, NSR-PPOSL, to seek NSR online solution. Through extensive simulations, we show that the proposed online method achieves over 80.0% performance of the offline method, under the condition of unknown user behavior. Moreover, our proposed online method outperforms representative benchmark by 13.5% under various network conditions and user behaviors.
Junxian Lu, Shuoyao Wang
GLOBECOM2
2023 Improving robustness of learning-based adaptive video streaming in wildly fluctuating networks
abstract
With the development of wireless technology, the fifth-generation mobile communication network (5G) can offer ultra-high bandwidth to support high-quality video streaming services. However, such a network shows much wild fluctuation especially in the case of mmWave 5G, making it challenging to ensure a robust video streaming service. In this paper, we propose a robust adaptive bitrate algorithm based on deep reinforcement learning, named MAP-DRL. In "robust", we mean that MAP-DRL provides high quality of experience (QoE) streaming services under various network conditions and different QoE preferences. In particular, to avoid the gradient explosion caused by network fluctuation, we introduce several reinforcement learning designs into MAP-DRL, such as advantage function normalization and reward scaling. To fully utilize the limited environment feedback, we process environmental information via a multi-type data manner and design the corresponding network structures for each type of them. The experimental results on real-world network trace datasets demonstrate that MAP-DRL can succeed in improving the average QoE by 29% versus heterogeneous networks and user preferences as compared with the representative benchmarks.
Shuoyao Wang
ICME2
2023 Improving Federated Person Re-Identification through Feature-Aware Proximity and Aggregation
abstract
Person re-identification (ReID) is a challenging task that aims to identify individuals across multiple non-overlapping camera views. To enhance the performance and robustness of ReID models, it is crucial to train them over multiple data sources. However, the traditional centralized approach poses a significant challenge to privacy as it requires collecting data from distributed data owners. To overcome this challenge, we employ the federated learning approach, which enables distributed model training without compromising data privacy. In this paper, we propose a novel feature-aware local proximity and global aggregation method for federated ReID to extract robust feature representations. Specifically, we introduce a proximal term and a feature regularization term for local model training to improve local training accuracy while ensuring global aggregation convergence. Furthermore, we use the cosine distance of backbone features to determine the global aggregation weight of each local model. Our proposed method significantly improves the performance and generalization of the global model. Extensive experiments demonstrate the effectiveness of our proposal. Specifically, our method achieves an additional 27.3% Rank-1 average accuracy in federated full supervision and an extra 20.3% mean Average Precision (mAP) on DukeMTMC in federated domain generalization.
Pengling Zhang, Huibin Yan, Wenhui Wu 0001, Shuoyao Wang
ACM Multimedia4
2023 Maximizing Ranking-Aware Recommendation Quality for Low-Complexity Network-Friendly Recommendation
abstract
Network-friendly recommendation (NFR) has been identified as a promising method to facilitate the network performance. However, the current NFR approaches introduce non-negligible computation overhead and primarily focus on the content to recommend, while neglecting the impact of ranking among contents, sharing limited recommendation quality. In this paper, we aim to design a recommendation re-ranking algorithm to maximize the ranking-aware recommendation quality while satisfying the network constraint. We formulate the problem as an integer programming (IP) problem that maximizes the mean reciprocal rank of the baseline recommendation under network constraint. To address the computational challenge posed by the large-scale IP problem, we propose a low-complexity algorithm that enables real-time NFR. Specifically, our approach reduces the computational complexity by re-ranking only the measurement list, instead of the entire content list. Empirical results on real-world datasets demonstrate that compared with state-of-the-art NFR baselines, the proposed scheme delivers an improvement of more than 4.8% in both precision and coverage, while reducing 97.5% computation time.
Jiayin Hou, Shuoyao Wang
VTC Fall3
2023 ResMon: Domain-Adaptive Wireless Respiration State Monitoring via Few-Shot Bayesian Deep Learning
abstract
Under the outbreak of the COVID-19 pandemic, respiration state monitoring plays an important role in assisting respiratory disease diagnosis and treatment. Thanks to the nonintrusive nature and low deployment cost, Wi-Fi-based wireless respiration state monitoring methods have gained increasing popularity. By analyzing the variation of channel state information (CSI) of Wi-Fi signals, the respiration states of a target person under the wireless coverage, such as cough, sneeze, and yawn, can be accurately detected. A major problem of the current wireless respiration state monitoring methods is being overly domain-dependent. That is, a sensing algorithm fine-tuned to a specific device placement and background setting (i.e., a domain) can result in drastic drop in detection accuracy when applied to a dissimilar new domain. To enhance the robustness of wireless sensing and reduce the sensing cost across different domains, we propose in this article a domain-adaptive respiration state monitoring system (ResMon) that achieves highly accurate cross-domain detection performance while requiring very limited labeled samples in the new domain. In a nutshell, the proposed ResMon consists of a source domain meta-training stage and a target domain meta-testing stage. In the meta-training stage, we leverage the rich source domain labeled data set to train an embedding model as a feature extractor of high-dimensional CSI data measurements. In particular, we apply the statistical Bayesian deep learning technique to improve the generalization performance of the embedding model in cross-domain applications. In the meta-testing stage, we combine the embedding model with a few-shot learning technique to train a domain-specific classifier using very limited labeled samples in the target domain. Experiment results show that the proposed ResMon can achieve on average 87.26% cross-domain detection accuracy in a 4-class respiration state classification task using only five labeled samples per class, which significantly outperforms the considered benchmark methods.
Suzhi Bi, Shuoyao Wang, Zhi Quan, Xian Li 0005, Xiaohui Lin 0001, Hui Wang 0022
IEEE Internet Things J.3
2023 Improving Beam Alignment Accuracy in mmWave Communication Systems With Auxiliary Tasks
abstract
Beam alignment is essential for high-quality data transmission in millimeter wave (mmWave) communication systems. Recent studies have revealed that the beam alignment method can train a small-scale probing codebook that is customized to specific sites, and leverage the codebook measurements to determine the optimal transmit beam. However, existing approaches still necessitate a certain number of probing beams in order to achieve high-accuracy alignment. In this work, we propose a multi-task learning-based beam alignment method, leveraging channel reconstruction and contrastive representation as auxiliary tasks, to improve the primary task of joint probing codebook design and optimal beam selection. Specifically, we offer a channel reconstruction module that estimates the wireless channel with the probing measurements, to improve the channel sensing efficiency of the probing codebook. Likewise, we also expose a contrastive representation module to improve the beam selector's representation robustness against the noisy channel. Results obtained from simulations using realistic public datasets indicate that the proposed method surpasses current state-of-the-art beam alignment techniques. The proposed method demonstrates superior performance in terms of alignment accuracy, achieved throughput, and beam sweeping complexity.
Shuoyao Wang, Suzhi Bi
IEEE Signal Process. Lett.1
2023 Edge Video Analytics With Adaptive Information Gathering: A Deep Reinforcement Learning Approach
abstract
With growing popularity of enormous public safety and transportation infrastructure cameras, there are increasing demands for automatic mobile video analytics. The emerging multi-access edge computing (MEC) technology has been recently applied to improve the accuracy-latency tradeoff of mobile video analytics. In this paper, we study an MEC-enabled multi-device video analytics system and formulate the problem as a Markov decision process (MDP) to meet two practical challenges: i) the absence of ground truth in real-time and ii) the content-varying degradation-accuracy relation. In particular, we aim to design an online joint frame degradation and bandwidth allocation algorithm with the time-varying function and limited feedback from each device. Thanks to the MDP formulation and$n$-step return technique, the long-term goal offers adaptive information gathering and thus improves the average accuracy and latency. For sample efficiency, we decompose the MDP problem into discrete degradation adaptation subproblems and continuous bandwidth allocation subproblems. Based on the decomposition, we propose a deep reinforcement learning (DRL) based framework, referred to as DBAG, to solve the decomposed subproblems. DBAG integrates model-based optimization and model-free DRL to solve the MDP problem with a discrete-continuous hybrid action space. Under various network setups and public datasets, DBAG greatly improves the accuracy-latency tradeoff.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
IEEE Trans. Wirel. Commun.1
2022 Enhancement or Super-Resolution: Learning-based Adaptive Video Streaming with Client-Side Video Processing
abstract
The rapid development of multimedia and communication technology has resulted in an urgent need for high-quality video streaming. However, robust video streaming under fluctuating network conditions and heterogeneous client computing capabilities remains a challenge. In this paper, we consider an enhancement-enabled video streaming network under a time-varying wireless network and limited computation capacity. "Enhancement" means that the client can improve the quality of the downloaded video segments via image processing modules. We aim to design a joint bitrate adaptation and client-side enhancement algorithm toward maximizing the quality of experience (QoE). We formulate the problem as a Markov decision process (MDP) and propose a deep reinforcement learning (DRL)-based framework, named ENAVS. As video streaming quality is mainly affected by video compression, we demonstrate that the video enhancement algorithm outperforms the super-resolution algorithm in terms of signal-to-noise ratio and frames per second, suggesting a better solution for client processing in video streaming. Ultimately, we implement ENAVS and demonstrate extensive testbed results under real-world bandwidth traces and videos. The simulation shows that ENAVS is capable of delivering 5%−14% more QoE under the same bandwidth and computing power conditions as conventional ABR streaming.
Shuoyao Wang
ICC3
2022 Deep Reinforcement Learning With Communication Transformer for Adaptive Live Streaming in Wireless Edge Networks
abstract
The emerging mobile edge computing (MEC) technology has been recently applied to improve the Quality of Experience (QoE) of network services, such as live video streaming. In this paper, we study an energy-aware adaptive live streaming scheme in wireless edge networks. In particular, we aim to design a joint uplink transmission and edge transcoding algorithm maximizing the video followers’ QoE, while minimizing the energy consumption of the video streamer. We formulate the problem as a Markov decision process (MDP), and propose a deep reinforcement learning (DRL) based framework, named SACCT, to determine the streamer’s encoding bitrate, the uploading power as well as the edge transcoding bitrates and frequency. We decompose the MDP problem into inter-frame and intra-frame problems to address the key design challenges that arise from continuous-discrete hybrid action space, time-varying state and action spaces, and unknown network variation. By doing so, SACCT integrates model-based optimization and model-free DRL to determine the intra-frame continuous resource allocation decisions and the inter-frame discrete bitrate adaptation decisions, respectively. To integrate both the numerical features (e.g., channel gain) and the categorical features (e.g., bitrate), we propose a communication Transformer (CT) as a backbone of SACCT by representing network states as communication tokens and running Transformers to model multi-scale dependencies. Extensive simulations manifest that compared with state-of-the-art approaches, SACCT can provide 128.23% (on average) extra reward. As such, by leveraging joint uplink adaption and edge transcoding, the proposed scheme enables an intelligent wireless network edge with QoE-assured and energy-aware live streaming services.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
IEEE J. Sel. Areas Commun.1
2022 Structure and Texture Preserving Network for Real-World Image Super-Resolution
abstract
Real-world image super-resolution (Real-SR) is a challenging task due to the unknown complex image degradation. Recent research on Real-SR has achieved remarkable progress by degradation process modeling; however, there are still undesired structural distortions and over-smoothed textures in the recovered images. In this letter, we propose a structure and texture preserving network, towards reducing the structural distortions while refining the perceptual-pleasant textures. Specifically, we propose a structure tensor (ST) branch to guide the restoration of high-resolution images by extracting channel-aggregated structural information. To further adaptively optimize different local texture, we replace the global discriminator with a global-local discriminator. By “local,” we mean that the discriminator loss, imposed on local areas randomly selected from the generated SR image, is minimized to generate textures with great visual perception in the selected local areas. Experimental results on five real-world datasets demonstrate the superiority of our methods in restoring structures, generating visually realistic SR images, as well as handling images of different degradation levels.
Bijun Zhou, Huibin Yan, Shuoyao Wang
IEEE Signal Process. Lett.3
2022 Adaptive Wireless Video Streaming: Joint Transcoding and Transmission Resource Allocation
abstract
The emerging mobile edge computing (MEC) technology has been recently applied to improve adaptive bitrate (ABR) streaming service quality under time-varying wireless channels. In this paper, we consider a heterogeneous multi-user MEC-enabled video streaming network with time-varying wireless channels in sequential time frames. In particular, we aim to design an online joint transcoding and transmission resource allocation algorithm to maximize the ABR streaming user’s quality of experience (QoE) subject to the bandwidth and CPU constraints. The algorithm is “online” in the sense that the bitrate and resource allocation decisions made at each frame depend only on the observation of past events. We formulate the problem as a mixed integer non-linear programming (MINLP) that jointly determines bitrate adaptation, bandwidth allocation, and CPU cycle assignment. To cope with the challenge arising from the coupling decisions of adjacent frames, we propose a low-complexity online algorithm, named OCCA. Specifically, by introducing queueing model constraints, we transform the offline non-conex MINLP problem into a multi-frame problem. Then, we analytically decouple the multi-stage(frame) problem to multiple per-frame convex subproblems that can be solved with high robustness and low computational complexity. We perform simulations with realistic scenarios to evaluate the performance of the proposed algorithm. Results manifest that compared with state-of-the-art approaches, our proposed algorithm can provide 97.84% (on average) extra QoE.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
IEEE Trans. Wirel. Commun.1
2021 ℓ2 Norm is all Your Need: Infrared-Visible Image Fusion VIA Guided Transformation Minimization
abstract
Among the key sub-topics of image fusion, infrared-visible image fusion technology is widely used in modern military and civilian domain. Driven by the data-free and end-to-end advancedness, growing research efforts have been devoted to convert the image fusion task into a norm minimization problem. To simultaneously keep the highlighted thermal information in the infrared image and the clear appearance information in the visible image, we propose a low-complexity fusion algorithm via guided transformation minimization, namely L2GTM. In particular, we formulate the fusion task as a solely `2norm minimization problem, which enjoys the uniqueness of the optimal solution. To adaptively integrate information from both infrared and visible images, we introduce guided weights to determine the degree of information retention. Experiments on public datasets validate the competitiveness of L2GTM from both the visual quality and objective evaluation perspective.
Huibin Yan, Shuoyao Wang
ICME2
2021 FCGP: Infrared and Visible Image Fusion via Joint Contrast and Gradient Preservation
abstract
With the fast development of multi-sensor technologies, image fusion has played an essential role in modern military and civilian. To better integrate thermal radiation information in infrared images and detailed appearance information in visible images, we investigate a novel norm formulation via joint contrast and gradient preservation, for infrared-visible image fusion. Specifically, we employ a structure tensor measurement to characterize the similarity between the fused image and the infrared image in terms of thermal radiation information, to better integrate visible appearance details. Since natural image gradients follow the hyper-Laplacian distribution, we employ$ \ell_{ p\in\left[0.5,0.8\right]} $norm instead of$ \ell_0 $or$ \ell_1 $norm to measure the gradient term to further extract texture information from visible images. To cope with the computational problem introduced by the coupled structure tensor and non-convex$ \ell_{p\in\left[0.5,0.8\right]}$norm, we propose a computational efficient solver based on half-quadratic splitting scheme. Experiments on public datasets validate the competitiveness of FCGP from both the subjective and objective evaluation perspective.
Huibin Yan, Shuoyao Wang
IEEE Signal Process. Lett.2
2021 Reinforcement Learning for Real-Time Pricing and Scheduling Control in EV Charging Stations
abstract
This article proposes a reinforcement-learning (RL) approach for optimizing charging scheduling and pricing strategies that maximize the system objective of a public electric vehicle (EV) charging station. The proposed algorithm is “online” in the sense that the charging and pricing decisions made at each time depend only on the observation of past events, and is “model-free” in the sense that the algorithm does not rely on any assumed stochastic models of uncertain events. To cope with the challenge arising from the time-varying continuous state and action spaces in the RL problem, we first show that it suffices to optimize the total charging rates to fulfill the charging requests before departure times. Then, we propose a feature-based linear function approximator for the state-value function to further enhance the efficiency and generalization ability of the proposed algorithm. Through numerical simulations with real-world data, we show that the proposed RL algorithm achieves on average 138.5% higher charging-station profit than representative benchmark algorithms.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
IEEE Trans. Ind. Informatics1
2020 Interpretable Multimodal Learning for Intelligent Regulation in Online Payment Systems
abstract
With the explosive growth of transaction activities in online payment systems, effective and real-time regulation becomes a critical problem for payment service providers. Thanks to the rapid development of artificial intelligence (AI), AI-enable regulation emerges as a promising solution. One main challenge of the AI-enabled regulation is how to utilize multimedia information, i.e., multimodal signals, in Financial Technology (FinTech). Inspired by the attention mechanism in nature language processing, we propose a novel cross-modal and intra-modal attention network (CIAN) to investigate the relation between the text and transaction. More specifically, we integrate the text and transaction information to enhance the text-trade joint-embedding learning, which clusters positive pairs and push negative pairs away from each other. Another challenge of intelligent regulation is the interpretability of complicated machine learning models. To sustain the requirements of financial regulation, we design a CIAN-Explainer to interpret how the attention mechanism interacts the original features, which is formulated as a low-rank matrix approximation problem. With the real datasets from the largest online payment system, WeChat Pay of Tencent, we conduct experiments to validate the practical application value of CIAN, where our method outperforms the state-of-the-art methods.
Shuoyao Wang, Diwei Zhu
IJCAI1
2020 Locational Detection of the False Data Injection Attack in a Smart Grid: A Multilabel Classification Approach
abstract
State estimation is critical to the monitoring and control of smart grids. Recently, the false data injection attack (FDIA) is emerging as a severe threat to state estimation. Conventional FDIA detection approaches are limited by their strong statistical knowledge assumptions, complexity, and hardware cost. Moreover, most of the current FDIA detection approaches focus on detecting the presence of FDIA, while the important information of the exact injection locations is not attainable. Inspired by the recent advances in deep learning, we propose a deep-learning-based locational detection architecture (DLLD) to detect the exact locations of FDIA in real time. The DLLD architecture concatenates a convolutional neural network (CNN) with a standard bad data detector (BDD). The BDD is used to remove the low-quality data. The followed CNN, as a multilabel classifier, is employed to capture the inconsistency and co-occurrence dependency in the power flow measurements due to the potential attacks. The proposed DLLD is “model-free” in the sense that it does not leverage any prior statistical assumptions. It is also “cost-friendly” in the sense that it does not alter the current BDD system and the runtime of the detection process is only hundreds of microseconds on a household computer. Through extensive experiments in the IEEE bus systems, we show that DLLD can perform locational detection precisely under various noise and attack conditions. In addition, we also demonstrate that the employed multilabel classification approach effectively enhances the presence-detection accuracy.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
IEEE Internet Things J.1
2018 The Impacts of Energy Customers Demand Response on Real-Time Electricity Market Participants
abstract
In this paper, we consider the profit-maximizing demand response of an energy customer in the real-time electricity market. In a real-time electricity market, the market clearing price is determined by the random deviation of actual power supply and demand from the predicted values in the day-ahead market. An energy customer, which requires a total amount of energy over a certain period of time, has the flexibility of shifting its energy usage in time, and therefore is in perfect position to exploit the volatile real-time market price through demand response. We show that the profit-maximizing demand response strategy can be obtained by solving a finite-horizon continuous-state Markov decision process (MDP) problem. Through rigorous analysis, we show that the optimal actual demand policy exhibits a threshold structure, which can solve the MDP without the need of discretizing the state and action spaces. We demonstrate through extensive simulations that the proposed demand response strategy not only maximizes the profit of the energy customer, but also alleviates the supply-demand imbalance in the power grid, and even reduces the bills of other market participants. On average, the proposed demand response strategy increases the energy customer's profit by 53.8% and saves the bills of other utilities by 80.4% comparing with the benchmark algorithms.
Shuoyao Wang, Suzhi Bi, Ying-Jun Angela Zhang
ICC1