Lijun He 0001

dblp:95/7206-1 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 11 since 2021Computer networks · 12 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative by leveraging their rich world knowledge. However, these methods typically rely on textual outputs, which introduces information loss, exhibits normalcy bias, and suffers from prompt sensitivity, making them insufficient for capturing subtle anomalous cues. To address these constraints, we propose HeadHunt-VAD, a novel tuning-free VAD paradigm that bypasses textual generation by directly hunting robust anomaly-sensitive internal attention heads within the frozen MLLM. Central to our method is a Robust Head Identification module that systematically evaluates all attention heads using a multi-criteria analysis of saliency and stability, identifying a sparse subset of heads that are consistently discriminative across diverse prompts. Features from these expert heads are then fed into a lightweight anomaly scorer and a temporal locator, enabling efficient and accurate anomaly detection with interpretable outputs. Extensive experiments show that HeadHunt-VAD achieves state-of-the-art performance among tuning-free methods on two major VAD benchmarks while maintaining high efficiency, validating head-level probing in MLLMs as a powerful and practical solution for real-world anomaly detection.
Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Haixia Bi, Lijun He 0001
AAAI5
2026 Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving
abstract
Modern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detection accuracy, current deep neural network-based models remain highly susceptible to adversarial examples. The underlying safety concerns motivate us to investigate realistic adversarial attacks in AD scenarios. Previous work has demonstrated the feasibility of placing adversarial posters on the road surface to induce hallucinations in the detector. However, the unnatural appearance of the posters makes them easily noticeable by humans, and their fixed content can be readily targeted and defended. To address these limitations, we propose the AdvRoad to generate diverse road-style adversarial posters. The adversaries have naturalistic appearances resembling the road surface while compromising the detector to perceive non-existent objects at the attack locations. We employ a two-stage approach, termed Road-Style Adversary Generation and Scenario-Associated Adaptation, to maximize the attack effectiveness on the input scene while ensuring the natural appearance of the poster, allowing the attack to be carried out stealthily without drawing human attention. Extensive experiments show that AdvRoad generalizes well to different detectors, scenes, and spoofing locations. Moreover, physical attacks further demonstrate the practical threats in real-world environments.
Jian Wang 0113, Lijun He 0001, Yixing Yong, Haixia Bi, Fan Li 0003
AAAI2
2026 Polarimetric diffusion model with hybrid convolutional Transformer for PolSAR image classification
Zuzheng Kuang, Haixia Bi, Lijun He 0001, Fan Li 0003
Pattern Recognit.4
2025 A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision
abstract
Extracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper layers generally contain more contextual information than shallower layers. To help the shallower layers enhance the ability of task-relevant feature extraction, we propose a self-distillation and multi-level supervision (SD-MS) method, which leverages the feature distillation from the deepest layer to the shallower one and multi-level supervision to guide network training. Visualization analysis shows that by introducing SD-MS, the same-level layer in the network can obtain a clearer representation of fundamental frequency components, while the shallower layers can even learn more task-relevant semantic information. Ablation study results indicate that SD-MS applies to existing melody extraction models and can consistently improve performance. Experimental results show that our proposed method, MF-TFA with SD-MS, outperforms six compared state-of-the-art methods, achieving overall accuracy (OA) scores of 87.1%, 89.9%, and 76.6% on the ADC 2004, MIREX 05, and MEDLEY DB datasets, respectively. The main code will be available at https://github.com/SmoothJing/MF-TFA_SD-MS.
Ying Hu 0005, Jiabo Jing, Fan Li 0003, Lijun He 0001, Wenzhong Yang
ICASSP4
2025 TDE-VC: Timbre Disentanglement and Extraction Via Consistency for Zero-Shot Voice Conversion
abstract
Voice conversion (VC) transforms certain characteristics of speech from a source to a target while preserving the original linguistic content. This paper focuses on timbre conversion, a key type of VC. Current VC methods face two challenges: retaining source speaker information in the extracted content and inadequately capturing timbre features, often leading to suboptimal speaker similarity in the converted speech. To address these issues, we propose the TDE-VC model, a zero-shot voice conversion framework that incorporates a phased-trained content extractor, combining the strengths of adversarial speaker classifier and data perturbation to extract cleaner content. Critically, we introduce a timbre disentanglement and extraction strategy, based on a multi-level consistency constraint, which effectively disentangles timbre from content and guides the timbre encoder to focus solely on timbre extraction. Additionally, we present an effective multi-scale timbre encoder. Experimental results demonstrate that TDE-VC significantly improves speaker similarity, especially for unseen target speakers, while maintaining competitive naturalness compared to existing methods. The demo page is publicly available.1.
Ying Hu 0005, Shangkun Tu, Fan Li 0003, Lijun He 0001, Hai Yan
ICME4
2025 MADRL-Based Collaborative Computation Offloading and Resource Orchestration for Multitask Data Sharing in Smart Agriculture
abstract
Multiple different computation tasks may be simultaneously offloaded to mobile edge computing (MEC) servers in smart agriculture scenarios, where the redundant transmission of shared data among different tasks leads to insufficient utilization of system resources (i.e., computing resources, communication resources, and caching resources) and lagging processing efficiency. Existing schemes optimizing multitask computation offloading with shared data almost focus on identical tasks, which are difficult to apply in real-world scenarios with different tasks to meet various service demands of fairness, low latency, and low energy consumption. In this article, we propose a fair, real-time, and green collaborative optimization scheme of computation offloading and resource orchestration for multitask data sharing in smart agriculture based on multiagent deep reinforcement learning (MADRL), aiming to improve offloading efficiency and system resource utilization to meet diverse tasks’ service demands. First, a collaborative optimization problem of computation offloading and resource orchestration is formulated to minimize the system latency, energy consumption, and caching space occupancy under constraints of redundant data transmission and limited system resources. It is difficult for traditional optimization methods to solve the formulated optimization problem characterized by dynamics, high-dimensionality, nonlinearity, and mixed-integer. Then, we propose an MADRL algorithm named MATD3-CO-RO-MDS based on a hierarchical reward mechanism to solve it and approximate the optimal offloading and orchestration strategy. Finally, experimental results prove that our proposed algorithm achieves smaller latency, energy consumption, and caching space occupancy compared with existing algorithms. It even has a 76.7% advantage in reducing latency when more tasks participate in offloading.
Tantan Zhao, Miao Zhang 0041, Lijun He 0001, Fan Li 0003
IEEE Internet Things J.3
2025 A Unified Framework for Adversarial Patch Attacks Against Visual 3D Object Detection in Autonomous Driving
abstract
The rapid development of vision-based 3D perceptions, in conjunction with the inherent vulnerability of deep neural networks to adversarial examples, motivates us to investigate realistic adversarial attacks for the 3D detection models in autonomous driving scenarios. Due to the perspective transformation from 3D space to the image and object occlusion, current 2D image attacks are difficult to generalize to 3D detectors and are limited by physical feasibility. In this work, we propose a unified framework to generate physically printable adversarial patches with different attack goals: 1)instance-level hiding—pasting the learned patches to any target vehicle allows it to evade the detection process; 2)scene-level creating—placing the adversarial patch in the scene induces the detector to perceive plenty of fake objects. Both crafted patches areuniversal, which can take effect across a wide range of objects and scenes. To achieve above attacks, we first introduce the differentiable image-3D rendering algorithm that makes it possible to learn a patch located in 3D space. Then, two novel designs are devised to promote effective learning of patch content: 1) a Sparse Object Sampling Strategy is proposed to ensure that the rendered patches follow the perspective criterion and avoid being occluded during training, and 2) a Patch-Oriented Adversarial Optimization is used to facilitate the learning process focused on the patch areas. Both digital and physical-world experiments are conducted and demonstrate the effectiveness of our approaches, revealing potential threats when confronted with malicious attacks. We also investigate the defense strategy using adversarial augmentation to further improve the model’s robustness.
Jian Wang 0113, Fan Li 0003, Lijun He 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 A Unified Framework for Generating Diverse and Stealthy Adversarial Patches Against Aerial Object Detection
abstract
Deep neural network (DNN) has become important in aerial object detection field. However, due to the vulnerabilities of DNN, adversarial patch is proved to be an efficient method to attack the DNN models, which can maliciously manipulate model predictions, potentially causing critical perception and decision-making errors in safety-sensitive systems. Current adversarial patch methods, which directly adopt pixel-level optimization, can obtain a good attack performance, but the generated patches inevitably introduce abstract content that significantly different from surrounding environment, creating easily identified patterns that undermine stealthiness. Moreover, optimized adversarial pattern is single and fixed after training, making them easy to defend against. In this work, we propose a unified adversarial attack framework, named Diverse and Stealthy Adversarial Patches (DSAP). The DSAP can generate arbitrary number of patches with different patterns, which are visually integrated with the surrounding environment and difficult to recognize. By placing the patches on the ground, the detector will identify the patches as non-existing target objects. We train a generator to implicitly encode adversarial content, which can directly map distinct latent vectors to patches with different adversarial patterns, making them difficult to defend against. In addition, we construct a source domain image set based on the scene background and use a discriminator to constrain the similarity between generated patches and background environment, ensuring stealthiness while keeping attack performance. Extensive experimental results demonstrate the superiority of our patch in terms of stealthiness and defense difficulty compared to ordinary adversarial patches, while the physical domain experiments indicate that our attack can transfer from digital to physical domain, posing a threat in the real world.
Yixing Yong, Jian Wang 0113, Lijun He 0001, Fan Li 0003
IEEE Trans. Geosci. Remote. Sens.3
2025 Physically Realizable Adversarial Creating Attack Against Vision-Based BEV Space 3D Object Detection
abstract
Vision-based 3D object detection, a cost-effective alternative to LiDAR-based solutions, plays a crucial role in modern autonomous driving systems. Meanwhile, deep models have been proven susceptible to adversarial examples, and attacking detection models can lead to serious driving consequences. Most previous adversarial attacks targeted 2D detectors by placing the patch in a specific region within the object's bounding box in the image, allowing it to evade detection. However, attacking 3D detector is more difficult because the adversary may be observed from different viewpoints and distances, and there is a lack of effective methods to differentiably render the 3D space poster onto the image. In this paper, we propose a novel attack setting where a carefully crafted adversarial poster (looks like meaningless graffiti) is learned and pasted on the road surface, inducing the vision-based 3D detectors to perceive a non-existent object. We show that even a single 2D poster is sufficient to deceive the 3D detector with the desired attack effect, and the poster is universal, which is effective across various scenes, viewpoints, and distances. To generate the poster, an image-3D applying algorithm is devised to establish the pixel-wise mapping relationship between the image area and the 3D space poster so that the poster can be optimized through standard backpropagation. Moreover, a ground-truth masked optimization strategy is presented to effectively learn the poster without interference from scene objects. Extensive results including real-world experiments validate the effectiveness of our adversarial attack. The transferability and defense strategy are also investigated to comprehensively understand the proposed attack.
Jian Wang 0113, Fan Li 0003, Song Lv, Lijun He 0001, Chao Shen 0001
IEEE Trans. Image Process.4
2024 Fine-Grained Dynamic Network for Generic Event Boundary Detection
Ziwei Zheng, Lijun He 0001, Le Yang 0007, Fan Li 0003
ECCV (43)2
2024 Multi-class Token-Guided End-to-End Weakly Supervised Image Semantic Segmentation Method
Lijun He 0001, Fan Li 0003
PRCV (13)2
2024 Secure Video Offloading in Multi-UAV-Enabled MEC Networks: A Deep Reinforcement Learning Approach
Tantan Zhao, Fan Li 0003, Lijun He 0001
IEEE Internet Things J.3
2024 VADiffusion: Compressed Domain Information Guided Conditional Diffusion for Video Anomaly Detection
abstract
The demand for security surveillance has grown exponentially, making video anomaly detection particularly crucial. Existing image-domain based anomaly detection algorithms face implementation challenges due to several drawbacks, including latency during long-distance transmission, the need for complete decoding, and the complexity of network inference structures. Moreover, current frame prediction methods using generative models suffer from low prediction quality and mode collapse. To tackle these challenges, we propose VADiffusion, a compressed domain information guided conditional diffusion framework. VADiffusion adopts a dual-branch structure that combines motion vector reconstruction and I-frame prediction, effectively addressing the limitations of the reconstruction method in identifying sudden anomalies and the struggles of the frame prediction method in detecting persistent anomalies. Furthermore, our proposed framework incorporates the diffusion model into the realm of video anomaly detection, thereby improving the stability and accuracy of the model. Specifically, we employ sparse sampling of the compressed video, utilizing I-frames to capture appearance information and motion vectors to represent motion-related details. Different from the existing independent two-branch mechanism, we adopt a reconstruction-assisted prediction strategy, leveraging I-frames and the reconstructed motion vectors from the reconstruction branch as conditions for the diffusion model utilized in frame prediction. Ultimately, we perform decision fusion of reconstruction and prediction branches to determine anomalies. Through extensive experiments, we demonstrate that our algorithm achieves an effective trade-off between detection accuracy and model complexity.
Hao Liu 0059, Lijun He 0001, Miao Zhang 0041, Fan Li 0003
IEEE Trans. Circuits Syst. Video Technol.2
2024 Dynamic Spatial Focus for Efficient Compressed Video Action Recognition
abstract
Recent years have witnessed a growing interest in compressed video action recognition due to the rapid growth of online videos. It remarkably reduces the storage by replacing raw videos with sparsely sampled RGB frames and other compressed motion cues (motion vectors and residuals). However, existing compressed video action recognition methods face two main issues: First, the inefficiency caused by the usage of coarse-level information under full resolution, and second, the disturbing due to the noisy dynamics in motion vectors. To address the two issues, this paper proposes a dynamic spatial focus method for efficient compressed video action recognition (CoViFocus). Specifically, we first use a light-weighted two-stream architecture to localize the task-relevant patches for both the RGB frames and motion vectors. Then the selected patch pair will be processed by a high-capacity two-stream deep model for the final prediction. Such a patch selection strategy crops out the irrelevant motion noise in motion vectors, as well as reduces the spatial redundancy of the inputs, leading to the high efficiency of our method in the compressed domain. Moreover, we found that the motion vectors can help our method to address the possibly happened static-issue, which means that the focus patches get stuck at some regions related to static objects rather than target actions, which further improves our method. Extensive results on both the HMDB-51 and UCF-101 datasets demonstrate the effectiveness and efficiency of our method in compressed video action recognition tasks.
Ziwei Zheng, Le Yang 0007, Yulin Wang 0002, Miao Zhang 0041, Lijun He 0001, Gao Huang 0001, Fan Li 0003
IEEE Trans. Circuits Syst. Video Technol.5
2024 Unsupervised Pansharpening Based on Double-Cycle Consistency
abstract
Multispectral (MS) pansharpening can improve the spatial resolution of MS images by fusing panchromatic (PAN) images, which have important applications in the fields of smart agriculture and environmental monitoring. However, existing supervised algorithms treat the original MS images as ground truth and generate training data under Wald’s protocol, resulting in a gap between the learned degradation process of the model and reality. This leads to the model having poor generalization and impractical. Unsupervised pansharpening methods often struggle to fully explore the rich information contained in images, leading to suboptimal pansharpening outcomes. In this work, we propose an unsupervised pansharpening algorithm based on double-cycle consistency that can learn directly from the original MS images without relying on artificially simulated degradation processes. Specifically, the network with cross-domain correlation information interaction is developed to achieve a deep fusion of spatial and spectral features. To address the inaccurate degradation mechanism representation of MS images, a spatial information extraction module based on scale invariance is developed to achieve an accurate representation. Meanwhile, double-cycle consistency loss is proposed to reduce the information loss caused by simulated degradation during the cycle process. Experimental results show that this method outperforms existing unsupervised pansharpening methods in both quantitative and qualitative evaluation of full-resolution images.
Lijun He 0001, Zhihan Ren 0002, Wanyue Zhang, Fan Li 0003, Shaohui Mei
IEEE Trans. Geosci. Remote. Sens.1
2024 Secure Video Offloading in MEC-Enabled IIoT Networks: A Multicell Federated Deep Reinforcement Learning Approach
abstract
Wireless video offloading in mobile-edge-computing (MEC)-enabled Industrial Internet of Things imposes a risk of exposing users' private data to eavesdroppers. It is difficult for existing secure video offloading schemes to simultaneously guarantee security, reduce latency and energy consumption in privacy-sensitive multicell scenarios where users are unwilling to offload data to other cells. In this article, a secure video offloading scheme based on multicell federated (MCF) deep reinforcement learning (DRL) is proposed to facilitate a secure, real-time, and efficient MEC network by efficient orchestration of limited resources. We formulate a collaborative optimization problem of video frame resolution and resources to minimize latency and energy consumption while maximizing the security rate subject to analytic accuracy and limited resources. To solve the formulated NP-hard problem, a MCF DRL algorithm based on the frameworks of multicell horizontal federated learning (FL) and hierarchical reward function-based twin delayed deep deterministic policy gradient (TD3) is proposed. First of all, hierarchical reward function-based TD3 is employed to solve the collaborative optimization NP-hard problem formulated for each single cell, where the optimal solution can be efficiently approached by the agent under the guidance of the innovatively designed hierarchical reward function. Then, multicell horizontal FL is applied on TD3 to obtain a model with higher model quality by averagely aggregating multiple individual TD3 models. Simulation results reveal that the proposed algorithm outperforms comparison algorithms in terms of utility, cost, latency, energy consumption, and security rate.
Tantan Zhao, Fan Li 0003, Lijun He 0001
IEEE Trans. Ind. Informatics3
2023 Frame importance and temporal memory effect-based fast video quality assessment for user-generated content
Yuan Zhang 0023, Lijun He 0001
Appl. Intell.4
2023 DRL-Based Secure Aggregation and Resource Orchestration in MEC-Enabled Hierarchical Federated Learning
abstract
Federated learning (FL) provides a new paradigm for protecting data privacy by enabling model training at devices and model aggregation at servers. However, data information may be leaked to honest-but-curious aggregation servers by updated model parameters. The existing secure methods do not fully exploit the potentiality of data characteristics in enhancing security, which makes it impossible to optimize limited system resources overall to achieve secure, fair, and efficient FL systems. In this article, a DRL-based joint secure aggregation and resource orchestration scheme is proposed to guarantee security and fairness, and improve efficiency for hierarchical FL (HFL) assisted by untrusted mobile-edge computing (MEC) servers. We formulate a joint optimization problem of data size, payment, and resource orchestration, to maximize the long-term social welfare subject to secure aggregation and limited resources. Since the formulated problem is a complex mixed integer dynamic optimization problem with NP-hardness, where multiple mixed integer optimization variables are highly coupled in time-varying constraints and objective function, it is difficult to obtain its optimal solution via traditional optimization methods. Thus, we propose a hierarchical reward function-based DRL algorithm (MATD3) to guide the agents to approach the optimal policy of secure aggregation and resource orchestration. Simulation results show that the proposed algorithm MATD3 can achieve superior performance over comparison algorithms and the MEC-enabled HFL framework outperforms two-layer FL frameworks.
Tantan Zhao, Fan Li 0003, Lijun He 0001
IEEE Internet Things J.3
2023 Multimodal Mutual Attention-Based Sentiment Analysis Framework Adapted to Complicated Contexts
abstract
Sentiment analysis has broad application prospects in the field of social opinion mining. The openness and invisibility of the internet makes users’ expression styles more diverse and thus results in the blooming of complicated contexts in which different unimodal data have inconsistent sentiment tendencies. However, most sentiment analysis algorithms only focus on designing multimodal fusion methods without preserving the individual semantics of each unimodal data. To avoid misunderstandings caused by ambiguity and sarcasm in complicated contexts, we propose a multimodal mutual attention-based sentiment analysis (MMSA) framework adapted to complicated contexts, which consists of three levels of subtasks to preserve the unimodal unique semantics and enhance the common semantics, to mine the association between unique semantics and common semantics and to balance decisions from unique and common semantics. In the framework, a multiperspective and hierarchical fusion (MHF) module is developed to fully fuse multimodal data, in which different modalities are mutually constrained and the fusion order is adjusted in the next step to enhance cross-modal complementarity. To balance the data, we calculate the loss by applying different weights to positive and negative samples. The experimental results on the CH-SIMS multimodal dataset show that our method outperforms existing multimodal sentiment analysis algorithms.The code of this work is available athttps://gitee.com/viviziqing/mmsacode.
Lijun He 0001, Fan Li 0003
IEEE Trans. Circuits Syst. Video Technol.1
2023 DRL-Based Joint Resource Allocation and Device Orchestration for Hierarchical Federated Learning in NOMA-Enabled Industrial IoT
abstract
Federated learning (FL) provides a new paradigm for protecting data privacy in Industrial Internet of Things (IIoT). To reduce network burden and latency brought by FL with a parameter server at the cloud, hierarchical federated learning (HFL) with mobile edge computing (MEC) servers is proposed. However, HFL suffers from a bottleneck of communication and energy overhead before reaching satisfying model accuracy as IIoT devices dramatically increase. In this article, a deep reinforcement learning (DRL)-based joint resource allocation and IIoT device orchestration policy using nonorthogonal multiple access is proposed to achieve a more accurate model and reduce overhead for MEC-assisted HFL in IIoT. We formulate a multiobjective optimization problem to simultaneously minimize latency, energy consumption, and model accuracy under the constraints of computing capacity and transmission power of IIoT devices. To solve it, we propose a DRL algorithm based on deep deterministic policy gradient. Simulation results show proposed algorithm outperforms others.
Tantan Zhao, Fan Li 0003, Lijun He 0001
IEEE Trans. Ind. Informatics3
2022 Unsupervised defect inspection algorithm based on cascaded GAN with edge repair feature fusion
Lijun He 0001, Nan Shi, Kainnat Malik, Fan Li 0003
Appl. Intell.1
2022 MTRFN: Multiscale Temporal Receptive Field Network for Compressed Video Action Recognition at Edge Servers
abstract
With the wide deployment of Internet of Things monitoring terminals, a tremendous number of videos are accumulated continuously. Big data processing and analysis-based action recognition has an increasingly important role in making cities simpler, better, and smarter. The traditional cloud server-centered analysis mode has to spend extra time transmitting vast video data terminals to remote cloud servers, which is always violated in real implementation. Edge servers with limited caching and computation capacities near the monitoring terminals enable implementation. However, due to the dependency on training data and the high complexity of extracting information and network architecture, existing image domain-based methods cannot be implemented at edge servers. Moreover, recognizing actions with different durations is still challenging. Due to these issues, we extend the traditional image domain to the compressed domain to efficiently extract the information of$I$frames and physical knowledge motion vectors (MVs), which can reflect the multiscale temporal feature just by partial decoding. To recognize the actions with different durations, a multiscale temporal receptive field network (MTRFN), including short-term and long-term branches, is proposed to simultaneously capture the action’s instant change based on the extracted MVs, the long temporal feature between adjacent$I$frames, and the interaction between them. The results show that our algorithm can achieve a better balance between accuracy and computational complexity.
Lijun He 0001, Miao Zhang 0041, Sijin Zhang, Fan Li 0003
IEEE Internet Things J.1
2022 Revenue and Energy Efficiency-Driven Delay-Constrained Computing Task Offloading and Resource Allocation in a Vehicular Edge Computing Network: A Deep Reinforcement Learning Approach
abstract
For in-vehicle application, task type and vehicle state information, i.e., vehicle speed, bear a significant impact on the task delay requirement. However, the joint impact of task type and vehicle speed on the task delay constraint has not been studied, and this lack of study may cause a mismatch between the requirement of the task delay and allocated computation and wireless resources. In this article, we propose a joint task type and vehicle speed-aware task offloading and resource allocation strategy to decrease the vehicle’s energy cost for executing tasks and increase the revenue of the vehicle for processing tasks within the delay constraint. First, we establish the joint task type and vehicle speed-aware delay constraint model. Then, the delay, energy cost, and revenue for task execution in the vehicular edge computing (VEC) server, local terminal, and terminals of other vehicles are calculated. Based on the energy cost and revenue from task execution, the utility function of the vehicle is acquired. Next, we formulate a joint optimization of task offloading and resource allocation to maximize the utility level of the vehicles subject to the constraints of task delay, computation resources, and wireless resources. To obtain a near-optimal solution of the formulated problem, a joint offloading and resource allocation based on the multiagent deep deterministic policy gradient (JORA-MADDPG) algorithm is proposed to maximize the utility level of vehicles. Simulation results show that our algorithm can achieve superior performance in task completion delay, vehicles’ energy cost, and processing revenue.
Lijun He 0001, Xing Chen 0007, Fan Li 0003
IEEE Internet Things J.2
2022 DRL-Based Secure Video Offloading in MEC-Enabled IoT Networks
abstract
Wireless offloading in mobile-edge-computing (MEC)-enabled Internet of Things (IoT) networks inevitably suffers the risk of eavesdropping. Physical-layer security (PLS) approaches can be applied to prevent eavesdropping. However, the existing PLS techniques are not well targeted for videos due to the fact that video’s distortion characteristics, which allow encoding parameters to be flexibly adjusted to enhance security in offloading, are ignored. A deep reinforcement learning (DRL)-based real-time, secure, and efficient video offloading scheme is proposed in this article, where video frame resolution, one key parameter of video’s distortion characteristics, is introduced and jointly optimized with PLS scheme to guarantee video’s security, improve users’ Quality of Experience (QoE) and save energy consumption. We formulate a joint optimization problem of video frame resolution selection, computation offloading, and resource allocation strategy, to minimize energy consumption and maximize QoE in terms of delay and analytic accuracy, while subject to security rate, computing capability, and transmission power. To solve the formulated NP-hard problem with the form of high-dimensional nonlinear mixed-integer programming, the hierarchical reward-function-based DRL (JVFRS-CO-RA-MADDPG) algorithm is proposed to guide the agents to obtain the optimal policy efficiently. Finally, the simulation results show that the proposed algorithm outperforms the existing algorithms in terms of delay, energy consumption, and security level.
Tantan Zhao, Lijun He 0001, Fan Li 0003
IEEE Internet Things J.2
2021 Vehicle theft recognition from surveillance video based on spatiotemporal attention
Lijun He 0001, Shuai Wen, Fan Li 0003
Appl. Intell.1
2021 Efficient attention based deep fusion CNN for smoke detection in fog environment
Lijun He 0001, Xiaoli Gong, Sirou Zhang, Fan Li 0003
Neurocomputing1
2020 A More Refined Mobile Edge Cache Replacement Scheme For Adaptive Video Streaming With Mutual Cooperation In Multi-Mec Servers
abstract
Instead of only focusing on the hit ratio of the videos cached in Mobile Edge Computing (MEC) server, we propose a more refined video segment content and client statusbased MEC cache update strategy, to improve clients' Quality of Experience (QoE). First, based on both the segment popularity and importance, we divide MEC cache into three parts which can be flexibly transformed into each other by combing the requested times of segments, transmission capability and clients' playback status together. Furthermore, we present the client's cache priority utility function and formulate a problem to maximize the utility function subject to the constraints of MEC cache size and transmission capacity. The brand and branch method is employed to obtain the optimal solution. Simulation results show that our algorithm can improve system throughput, hit ratio of video segment, playback frozen time as well as backhaul traffic.
Lijun He 0001, Xing Chen 0007, Guizhong Liu, Fan Li 0003
ICME2
2020 Playback experience driven cross layer optimisation of APP, transport and MAC layer for video clients over long-term evolution system
abstract
In the traditional communication system, information of application (APP) layer, transport layer, and media access control (MAC) layer has not been fully interacted. To solve the problem, the authors propose a joint optimisation framework, which consists of the APP layer, transport layer, and MAC layer, to improve the video clients' playback experience and system throughput. First, a client requirement aware autonomous packet drop strategy, based on packet importance, channel condition, and playback status, is developed to decrease the network load and the probability of rebuffering events. Furthermore, TCP state aware downlink and uplink resource allocation schemes are proposed to achieve smooth video transmission and steady acknowledgment (ACK) feedback, respectively. For the downlink scheme, the maximum transmission capacity required for each client is calculated based on feedback ACK information from the transport layer to avoid allocating excessive resource to the client, whose ACK feedback is blocked due to bad uplink channel condition. For the uplink scheme, information of retransmission timeout and TCP congestion window are utilised to indicate ACK scheduling priority. The simulation results show that their algorithm can significantly improve the system throughput and the clients' playback continuity with the acceptable video quality.
Lijun He 0001
IET Commun.2
2019 Hit Ratio Driven Mobile Edge Caching Scheme for Video on Demand Services
abstract
More and more scholars focus on mobile edge computing (MEC) technology, because the strong storage and computing capabilities of MEC servers can reduce the long transmission delay, bandwidth waste, energy consumption, and privacy leaks in the data transmission process. In this paper, we study the cache placement problem to determine how to cache videos and which videos to be cached in a mobile edge computing system. First, we derive the video request probability by taking into account video popularity, user preference and the characteristic of video representations. Second, based on the acquired request probability, we formulate a cache placement problem with the objective to maximize the cache hit ratio subject to the storage capacity constraints. Finally, in order to solve the formulated problem, we transform it into a grouping knapsack problem and develop a dynamic programming algorithm to obtain the optimal caching strategy. Simulation results show that the proposed algorithm can greatly improve the cache hit ratio.
Xing Chen 0007, Lijun He 0001, Shang Xu, Shibo Hu, Qingzhou Li, Guizhong Liu
ICME2
2018 Playback continuity and video quality driven optimisation for dynamic adaptive streaming over HTTP clients over wireless networks
abstract
In this study, the authors focus on segment characteristic analysis and joint optimisation of modulation and coding scheme (MCS) selection, resource block (RB) assignment and the block error ratio (BLER) determination and segment adaptation scheme to satisfy dynamic adaptive streaming over HTTP (DASH) clients over wireless networks. First, the authors define a utility function as the continuous playback time that the packets scheduled with the allocated resource can support. Then, the authors formulate the MCS selection, RB assignment and BLER determination into a mathematical model. The relationship among the above three factors can be explored instead of performing MCS selection and RB assignment with the fixed BLER. By decomposing the original problem into some sub‐problems, the authors can get a solution to the original problem with low complexity. At the client level, the authors develop an adaptive segment request strategy based on the playback information, the segments' characteristics and the estimated transmission rate. To decrease the influence of the inaccurate estimations, an adaptive guard time interval based on the real playback information and transmission information of the previous segments is introduced. Simulation results show that the proposed algorithm can efficiently improve playback continuity and video quality over the existing algorithms.
Lijun He 0001, Guizhong Liu
IET Commun.1
2017 Low-complexity multi-service power optimisation algorithm based on MOS models
abstract
In this study, the authors consider the power optimisation based on the MOS (mean of score) models of different services over multi‐user femtocell systems. They first formulate a mathematic model to minimise the total power consumption, subject to the MOS requirements and the network resource constraint. To solve the formulated problem, they first propose an optimal power allocation (OPA) scheme, in which the optimal power allocation and RB (resource block) assignment can be obtained by employing the Lagrange dual decomposition method. However, iterative update of the Lagrange variables leads to high complexity. To deal with this, they further propose a utility‐based suboptimal power optimisation algorithm (SPA). The utility of assigning each RB to each user is defined as the power consumption difference caused by adding the RB to the set of the RBs already assigned to the user. To calculate the utility, they develop a low‐complexity optimal power allocation scheme. Based on the calculated utility, each RB is assigned to the user with the maximum utility. Experimental results demonstrate that OPA can indeed acquire an upper bound of the system performance. SPA can improve the MOS of the users with acceptable complexity compared with other existing algorithms.
Lijun He 0001, Guizhong Liu
IET Commun.1
2014 Playback continuity driven cross-layer design for HTTP streaming in LTE systems
abstract
In this paper, a playback continuity driven cross-layer design is proposed for HTTP streaming in LTE systems in which the stringent delay deadline of the application layer is translated into the transmission capacity requirement at the MAC layer of eNodeB. First the client buffer information such as the fullness of the buffer and the playback information is easily obtained at the client. Based on the buffer information, the continuous playback time interval that the completely received segments in the client buffer can support is calculated. To keep continuous playback, the remaining packets of the un-completely received segment should arrive at the client during the time interval. Considering the video information in MPD from the HTTP server and the buffer information about the un-completely received segment, the sum size of the remaining packets in MAC queue can be obtained. From the time interval and the sum size, the transmission rate requirement that the MAC layer should satisfy can be acquired. For a scheduling period with fixed duration called TTI, the transmission capacity at the MAC layer can be easily acquired. Then a new mathematical model is proposed to maximize the system throughput subject to the transmission capacity requirements of the clients and the total RB constraint of the physical layer. Then a resource allocation scheme with low complexity is developed to solve the proposed problem. Simulation results show that the proposed algorithm can efficiently improve the playback continuity compared with other existing algorithms.
Lijun He 0001, Guizhong Liu
WoWMoM1
2014 A multiuser simulation system for video transmission over HSDPA
Qinli Wang, Guizhong Liu, Lijun He 0001
Multim. Tools Appl.6
2014 Quality-Driven Cross-Layer Design for H.264/AVC Video Transmission over OFDMA System
abstract
Video applications over wireless network have attracted more and more attention in recent years. However, due to the unique characteristics of video, limited available resources and the time-varying wireless channel state, video transmission over wireless network is still a challenging task. In this paper, a novel cross-layer design is proposed to maximize the overall received video quality with the limited network resource and the stringent deadline constraints over OFDMA network. Considering the information extracted from the Application layer, Media Access Control layer and Physical layer, we formulate the packet scheduling, subcarrier assignment and power allocation into a mathematical model. By employing Lagrange dual decomposition, asymptotically optimal packet scheduling and resource allocation strategy can be obtained. To reduce the computation complexity, we further develop a suboptimal solution which performs packet scheduling and subcarrier assignment jointly but power optimization independently. Finally, we validate our proposed algorithms in a multi-user scenario where all the video sequences are pre-encoded by H.264/AVC encoder. Experimental results demonstrate that our proposed algorithms can indeed improve the received video quality significantly compared with other existing algorithms.
Lijun He 0001, Guizhong Liu
IEEE Trans. Wirel. Commun.1
2012 Optimal cross layer design for video transmission over OFDMA system
abstract
Video transmission over OFDMA (orthogonal frequency-division multiple-access) system is a challenging problem which involves better received video quality and stringent transmission delays. In this paper, we present a cross-layer design based on the joint optimization of resource allocation and packet scheduling in which the importance of each video packet is taken into consideration. The objective is to maximize the received video quality of all the users subject to the network resource constraint. By employing the Lagrange dual decomposition method, we can obtain the global optimal solution to the optimization problem. We also propose a method of suboptimal solution to reduce the high complexity. Simulation results show that the proposed algorithms have superior performance.
Lijun He 0001, Guizhong Liu
ICC1
2009 Application-driven cross-layer design of multiuser H.264 video transmission over wireless networks
abstract
An application-driven cross-layer design of multiuser H.264/AVC video transmission over wireless networks is proposed in this paper. An objective function for cross-layer optimization is developed based on the parameters abstracted from the application layer, the media access control layer and the physical layer. Our objective is to maximize the video perceptual quality after delivery with constraint of the limited wireless resources. Simulation results show that the proposed scheme performs significantly than the conventional scheduling schemes for video transmission.
Fan Li 0003, Guizhong Liu, Lijun He 0001
IWCMC3