EDBT 2026 Demo / reviewers in the wild / expert
Fan Li 0003
dblp:73/237-3
· DBLP profile ↗
84ranked-venue papers
12as first author
64since 2021 · last 2026
0000-0002-7566-1634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 8 first-author · 34 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Computer networks · 13 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 13 since 2021Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative by leveraging their rich world knowledge. However, these methods typically rely on textual outputs, which introduces information loss, exhibits normalcy bias, and suffers from prompt sensitivity, making them insufficient for capturing subtle anomalous cues. To address these constraints, we propose HeadHunt-VAD, a novel tuning-free VAD paradigm that bypasses textual generation by directly hunting robust anomaly-sensitive internal attention heads within the frozen MLLM. Central to our method is a Robust Head Identification module that systematically evaluates all attention heads using a multi-criteria analysis of saliency and stability, identifying a sparse subset of heads that are consistently discriminative across diverse prompts. Features from these expert heads are then fed into a lightweight anomaly scorer and a temporal locator, enabling efficient and accurate anomaly detection with interpretable outputs. Extensive experiments show that HeadHunt-VAD achieves state-of-the-art performance among tuning-free methods on two major VAD benchmarks while maintaining high efficiency, validating head-level probing in MLLMs as a powerful and practical solution for real-world anomaly detection. Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Haixia Bi, Lijun He 0001 |
AAAI | 2 |
| 2026 | Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous DrivingabstractModern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detection accuracy, current deep neural network-based models remain highly susceptible to adversarial examples. The underlying safety concerns motivate us to investigate realistic adversarial attacks in AD scenarios. Previous work has demonstrated the feasibility of placing adversarial posters on the road surface to induce hallucinations in the detector. However, the unnatural appearance of the posters makes them easily noticeable by humans, and their fixed content can be readily targeted and defended. To address these limitations, we propose the AdvRoad to generate diverse road-style adversarial posters. The adversaries have naturalistic appearances resembling the road surface while compromising the detector to perceive non-existent objects at the attack locations. We employ a two-stage approach, termed Road-Style Adversary Generation and Scenario-Associated Adaptation, to maximize the attack effectiveness on the input scene while ensuring the natural appearance of the poster, allowing the attack to be carried out stealthily without drawing human attention. Extensive experiments show that AdvRoad generalizes well to different detectors, scenes, and spoofing locations. Moreover, physical attacks further demonstrate the practical threats in real-world environments. Jian Wang 0113, Lijun He 0001, Yixing Yong, Haixia Bi, Fan Li 0003 |
AAAI | 5 |
| 2026 | Polarimetric diffusion model with hybrid convolutional Transformer for PolSAR image classification
Zuzheng Kuang, Haixia Bi, Lijun He 0001, Fan Li 0003 |
Pattern Recognit. | 5 |
| 2026 | IDEAL: Independent domain embedding augmentation learning
Zelin Yang, Lin Xu 0001, Shiyang Yan, Haixia Bi, Fan Li 0003 |
Pattern Recognit. | 5 |
| 2026 | LiDAR point clouds segmentation in adverse weather conditions
Yi An, Fan Li 0003 |
Signal Process. | 2 |
| 2026 | Composable Multimodal Semantic Communication: A Lightweight Large AI Model Approach
Tantan Zhao, Fan Li 0003, Arumugam Nallanathan |
IEEE Trans. Commun. | 2 |
| 2026 | LSFMamba: Local-Enhanced Spiral Fusion Mamba for Multi-Modal Land Cover ClassificationabstractMulti-modal learning, which fuses complementary information from different modalities, has significantly improved the accuracy of land cover classification, especially under adverse conditions like cloudy or rainy weather. Recent advancements in multi-modal remote sensing land cover classification (MMRLC) have witnessed the efficacy of approaches based on CNN and Transformer. However, CNN exhibits limitations in capturing long-range dependencies, whereas Transformer suffers from high computational complexity. Recently, Mamba has garnered widespread attention due to its superior long-range modeling capabilities with linear complexity. Nevertheless, Mamba demonstrates notable limitations when directly applied to MMRLC, including limited local contextual modeling capacity, suboptimal multi-modal feature fusion and lack of a task-specific spatial continuity scanning strategy. Hence, to fully explore the potential of Mamba in multi-modal land cover classification, we propose LSFMamba, which comprises multiple hierarchically connected local-enhanced fusion Mamba (LFM) modules. Within each LFM module, a local-enhanced visual state space (LVSS) block is designed to extract features from different modalities, while a cross-modal interaction state space (CISS) block is created to fuse these multi-modal features. In the LVSS block, we integrate a multi-kernel CNN block into the gating branch in Mamba to enhance its local modeling capabilities. In the CISS block, features from different modalities are interleaved, facilitating cross-modal feature interaction through the state space model. Furthermore, we introduce a novel spiral scanning strategy to reassess the significance of central pixels, a design driven by the unique characteristics of pixel-wise classification task. Extensive experimental results on three multi-modal remote sensing datasets demonstrate that the proposed LSFMamba achieves state-of-the-art performance with lower complexity. The code will be released at https://github.com/hhchhang78/LSFMamba. Honghao Chang, Haixia Bi, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Singing Melody Extraction Network Via Self-Distillation and Multi-Level SupervisionabstractExtracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper layers generally contain more contextual information than shallower layers. To help the shallower layers enhance the ability of task-relevant feature extraction, we propose a self-distillation and multi-level supervision (SD-MS) method, which leverages the feature distillation from the deepest layer to the shallower one and multi-level supervision to guide network training. Visualization analysis shows that by introducing SD-MS, the same-level layer in the network can obtain a clearer representation of fundamental frequency components, while the shallower layers can even learn more task-relevant semantic information. Ablation study results indicate that SD-MS applies to existing melody extraction models and can consistently improve performance. Experimental results show that our proposed method, MF-TFA with SD-MS, outperforms six compared state-of-the-art methods, achieving overall accuracy (OA) scores of 87.1%, 89.9%, and 76.6% on the ADC 2004, MIREX 05, and MEDLEY DB datasets, respectively. The main code will be available at https://github.com/SmoothJing/MF-TFA_SD-MS. Ying Hu 0005, Jiabo Jing, Fan Li 0003, Lijun He 0001, Wenzhong Yang |
ICASSP | 3 |
| 2025 | A Task-Oriented Real-Time and Robust Feature Compression and Selection Method in Collaborative Intelligence SystemabstractThe emerging autonomous driving has stringent requirements for latency and reliability. In this paper, we propose a task-oriented real-time and robust feature compression and selection method in collaborative intelligence system. Our design, consisting of a three-dimensional channel compression (TDCC) module and a one-dimensional feature selection (ODFS) module, can efficiently reduce feature size in different dimensional spaces. The TDCC initially reduces the number of feature channels in three-dimensional space. Subsequently, the ODFS flattens the three-dimensional feature maps into a one-dimensional feature vector and selects the most effective feature dimensions of the vector for transmission. Specifically, in ODFS, the most relevant information is first aggregated into specific dimensions through an information-theory-based inductive loss function. Then, the important features are selected using the mask generated by the mask generator. Extensive experiments demonstrate that the proposed method outperforms the baseline in both communication latency and task accuracy, while exhibiting robustness against poor channel conditions. Kaile Wang, Tantan Zhao, Fan Li 0003 |
ICASSP | 3 |
| 2025 | TDE-VC: Timbre Disentanglement and Extraction Via Consistency for Zero-Shot Voice ConversionabstractVoice conversion (VC) transforms certain characteristics of speech from a source to a target while preserving the original linguistic content. This paper focuses on timbre conversion, a key type of VC. Current VC methods face two challenges: retaining source speaker information in the extracted content and inadequately capturing timbre features, often leading to suboptimal speaker similarity in the converted speech. To address these issues, we propose the TDE-VC model, a zero-shot voice conversion framework that incorporates a phased-trained content extractor, combining the strengths of adversarial speaker classifier and data perturbation to extract cleaner content. Critically, we introduce a timbre disentanglement and extraction strategy, based on a multi-level consistency constraint, which effectively disentangles timbre from content and guides the timbre encoder to focus solely on timbre extraction. Additionally, we present an effective multi-scale timbre encoder. Experimental results demonstrate that TDE-VC significantly improves speaker similarity, especially for unseen target speakers, while maintaining competitive naturalness compared to existing methods. The demo page is publicly available.1. Ying Hu 0005, Shangkun Tu, Fan Li 0003, Lijun He 0001, Hai Yan |
ICME | 3 |
| 2025 | HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMsabstractVideo Anomaly Detection (VAD) aims to identify and locate deviations from normal patterns in video sequences. Traditional methods often struggle with substantial computational demands and a reliance on extensive labeled datasets, thereby restricting their practical applicability. To address these constraints, we propose HiProbe-VAD, a novel framework that leverages pre-trained Multimodal Large Language Models (MLLMs) for VAD without requiring fine-tuning. In this paper, we discover that the intermediate hidden states of MLLMs contain information-rich representations, exhibiting higher sensitivity and linear separability for anomalies compared to the output layer. To capitalize on this, we propose a Dynamic Layer Saliency Probing (DLSP) mechanism that intelligently identifies and extracts the most informative hidden states from the optimal intermediate layer during the MLLMs reasoning. Then a lightweight anomaly scorer and temporal localization module efficiently detects anomalies using these extracted hidden states and finally generate explanations. Experiments on the UCF-Crime and XD-Violence datasets demonstrate that HiProbe-VAD outperforms existing training-free and most traditional approaches. Furthermore, our framework exhibits remarkable cross-model generalization capabilities in different MLLMs without any tuning, unlocking the potential of pre-trained MLLMs for video anomaly detection and paving the way for more practical and scalable solutions. Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Yanjun Qin |
ACM Multimedia | 2 |
| 2025 | Single-layer feedforward neural networks with dynamic width for domain adaptation
Le Yang 0007, Zelin Yang, Fan Li 0003, C. L. Philip Chen |
Sci. China Inf. Sci. | 3 |
| 2025 | MADRL-Based Collaborative Computation Offloading and Resource Orchestration for Multitask Data Sharing in Smart AgricultureabstractMultiple different computation tasks may be simultaneously offloaded to mobile edge computing (MEC) servers in smart agriculture scenarios, where the redundant transmission of shared data among different tasks leads to insufficient utilization of system resources (i.e., computing resources, communication resources, and caching resources) and lagging processing efficiency. Existing schemes optimizing multitask computation offloading with shared data almost focus on identical tasks, which are difficult to apply in real-world scenarios with different tasks to meet various service demands of fairness, low latency, and low energy consumption. In this article, we propose a fair, real-time, and green collaborative optimization scheme of computation offloading and resource orchestration for multitask data sharing in smart agriculture based on multiagent deep reinforcement learning (MADRL), aiming to improve offloading efficiency and system resource utilization to meet diverse tasks’ service demands. First, a collaborative optimization problem of computation offloading and resource orchestration is formulated to minimize the system latency, energy consumption, and caching space occupancy under constraints of redundant data transmission and limited system resources. It is difficult for traditional optimization methods to solve the formulated optimization problem characterized by dynamics, high-dimensionality, nonlinearity, and mixed-integer. Then, we propose an MADRL algorithm named MATD3-CO-RO-MDS based on a hierarchical reward mechanism to solve it and approximate the optimal offloading and orchestration strategy. Finally, experimental results prove that our proposed algorithm achieves smaller latency, energy consumption, and caching space occupancy compared with existing algorithms. It even has a 76.7% advantage in reducing latency when more tasks participate in offloading. Tantan Zhao, Miao Zhang 0041, Lijun He 0001, Fan Li 0003 |
IEEE Internet Things J. | 4 |
| 2025 | Deep Reinforcement Learning- and Information Bottleneck-Enabled Task-Oriented Semantic CommunicationabstractTask-oriented semantic communication offers a promising solution for providing real-time computer vision services. However, existing research on semantic communication ignores the connection between the key performance indicators (KPIs) used to measure semantic encoding-decoding networks (i.e., inference accuracy) and wireless semantic communication networks (i.e., transmission latency). Therefore, the designed semantic communication schemes are difficult to simultaneously meet the requirements of low latency and high accuracy of emerging intelligent applications. In this paper, by deeply exploring the relationship and interdependencies between the two kinds of KPIs, we propose a real-time and efficient task-oriented end-to-end semantic communication scheme enabled by deep reinforcement learning (DRL) and information bottleneck to improve both inference and communication efficiency. Specifically, we initially use information bottleneck theory to model the optimal tradeoff between inference accuracy and communication latency, which is subsequently reformulated by variational inference to be differentiable and tractable. Then, we introduce DRL to address the non-differentiability of dynamic stochastic fading channels and channel mismatch between the training phase and deployment phase, enabling accurate selection of the most task-relevant semantic feature dimensions for transmission under dynamic fading channels. Finally, extensive experiments show that our proposed scheme achieves better performance in latency and accuracy than comparison methods. Tantan Zhao, Fan Li 0003, Hongyang Du 0001, Li Sun 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Group Image Compression for Dual Use of Machine and Human VisionabstractFaces in a scene of human group, if coded with sufficient precision, can be computer analyzed for machine vision tasks involving faces. But this requires storing and communicating them at a very high bit rate. Traditional ROI-based image compression methods are ill suited to code many faces at high precision against a complex background. In this work, we propose a novel group image compression neural network (GICNet) of two layers: 1) the face layer dedicated to machine analysis, in which face bounding boxes are first cropped out of the background and converted to a compression-friendly canonical sketch-guided representation of fixed resolution for compact coding and facilitating downstream tasks without additional preprocessing; 2) the background layer dedicated to overall human vision perceptual quality, in which face residuals and background elements are coded and appended to the code stream. Experimental results demonstrate the effectiveness of our proposed GICNet, conserving up to 13%-57% bitrate for machine vision applications while maintaining competitive perceptual quality. Xiaolin Wu 0001, Fan Li 0003, Yiping Duan, Xiaoming Tao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Unified Framework for Adversarial Patch Attacks Against Visual 3D Object Detection in Autonomous DrivingabstractThe rapid development of vision-based 3D perceptions, in conjunction with the inherent vulnerability of deep neural networks to adversarial examples, motivates us to investigate realistic adversarial attacks for the 3D detection models in autonomous driving scenarios. Due to the perspective transformation from 3D space to the image and object occlusion, current 2D image attacks are difficult to generalize to 3D detectors and are limited by physical feasibility. In this work, we propose a unified framework to generate physically printable adversarial patches with different attack goals: 1)instance-level hiding—pasting the learned patches to any target vehicle allows it to evade the detection process; 2)scene-level creating—placing the adversarial patch in the scene induces the detector to perceive plenty of fake objects. Both crafted patches areuniversal, which can take effect across a wide range of objects and scenes. To achieve above attacks, we first introduce the differentiable image-3D rendering algorithm that makes it possible to learn a patch located in 3D space. Then, two novel designs are devised to promote effective learning of patch content: 1) a Sparse Object Sampling Strategy is proposed to ensure that the rendered patches follow the perspective criterion and avoid being occluded during training, and 2) a Patch-Oriented Adversarial Optimization is used to facilitate the learning process focused on the patch areas. Both digital and physical-world experiments are conducted and demonstrate the effectiveness of our approaches, revealing potential threats when confronted with malicious attacks. We also investigate the defense strategy using adversarial augmentation to further improve the model’s robustness. Jian Wang 0113, Fan Li 0003, Lijun He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | ECP-Mamba: An Efficient Multiscale Self-Supervised Contrastive Learning Method With State Space Model for PolSAR Image ClassificationabstractRecently, polarimetric synthetic aperture radar (PolSAR) image classification has been greatly promoted by deep neural networks. However, current deep learning-based PolSAR image classification methods are caught in the dilemma of obtaining high accuracy with sparse labels while maintaining high computational efficiency. To solve this issue, we present ECP-Mamba, an efficient framework integrating multi-scale self-supervised contrastive learning with a state space model backbone. Specifically, we design a cross-scale predictive pretext task, which learns representations via aligning local and global polarimetric features, effectively mitigating the annotation scarcity issue. To enhance computational efficiency, we introduce Mamba architecture to PolSAR image classification for the first time. A spiral scanning strategy tailored for pixel-wise classification task is proposed within this framework, prioritizing causally relevant features near the central pixel. Additionally, a lightweight cross Mamba module is proposed to facilitate complementary multi-scale feature interaction. Extensive experiments on four benchmark datasets demonstrate the effectiveness of ECP-Mamba in balancing high accuracy with computational efficiency. On the Flevoland 1989 dataset, ECP-Mamba achieves state-of-the-art performance with an overall accuracy of 99.70%, an average accuracy of 99.64% and a Kappa coefficient of 0.9962. Our code will be available at https://github.com/HaixiaBi1982/ECP_Mamba. Zuzheng Kuang, Haixia Bi, Fan Li 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Unified Framework for Generating Diverse and Stealthy Adversarial Patches Against Aerial Object DetectionabstractDeep neural network (DNN) has become important in aerial object detection field. However, due to the vulnerabilities of DNN, adversarial patch is proved to be an efficient method to attack the DNN models, which can maliciously manipulate model predictions, potentially causing critical perception and decision-making errors in safety-sensitive systems. Current adversarial patch methods, which directly adopt pixel-level optimization, can obtain a good attack performance, but the generated patches inevitably introduce abstract content that significantly different from surrounding environment, creating easily identified patterns that undermine stealthiness. Moreover, optimized adversarial pattern is single and fixed after training, making them easy to defend against. In this work, we propose a unified adversarial attack framework, named Diverse and Stealthy Adversarial Patches (DSAP). The DSAP can generate arbitrary number of patches with different patterns, which are visually integrated with the surrounding environment and difficult to recognize. By placing the patches on the ground, the detector will identify the patches as non-existing target objects. We train a generator to implicitly encode adversarial content, which can directly map distinct latent vectors to patches with different adversarial patterns, making them difficult to defend against. In addition, we construct a source domain image set based on the scene background and use a discriminator to constrain the similarity between generated patches and background environment, ensuring stealthiness while keeping attack performance. Extensive experimental results demonstrate the superiority of our patch in terms of stealthiness and defense difficulty compared to ordinary adversarial patches, while the physical domain experiments indicate that our attack can transfer from digital to physical domain, posing a threat in the real world. Yixing Yong, Jian Wang 0113, Lijun He 0001, Fan Li 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Physically Realizable Adversarial Creating Attack Against Vision-Based BEV Space 3D Object DetectionabstractVision-based 3D object detection, a cost-effective alternative to LiDAR-based solutions, plays a crucial role in modern autonomous driving systems. Meanwhile, deep models have been proven susceptible to adversarial examples, and attacking detection models can lead to serious driving consequences. Most previous adversarial attacks targeted 2D detectors by placing the patch in a specific region within the object's bounding box in the image, allowing it to evade detection. However, attacking 3D detector is more difficult because the adversary may be observed from different viewpoints and distances, and there is a lack of effective methods to differentiably render the 3D space poster onto the image. In this paper, we propose a novel attack setting where a carefully crafted adversarial poster (looks like meaningless graffiti) is learned and pasted on the road surface, inducing the vision-based 3D detectors to perceive a non-existent object. We show that even a single 2D poster is sufficient to deceive the 3D detector with the desired attack effect, and the poster is universal, which is effective across various scenes, viewpoints, and distances. To generate the poster, an image-3D applying algorithm is devised to establish the pixel-wise mapping relationship between the image area and the 3D space poster so that the poster can be optimized through standard backpropagation. Moreover, a ground-truth masked optimization strategy is presented to effectively learn the poster without interference from scene objects. Extensive results including real-world experiments validate the effectiveness of our adversarial attack. The transferability and defense strategy are also investigated to comprehensively understand the proposed attack. Jian Wang 0113, Fan Li 0003, Song Lv, Lijun He 0001, Chao Shen 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
Le Yang 0007, Ziwei Zheng, Yizeng Han, Hao Cheng 0015, Shiji Song, Gao Huang 0001, Fan Li 0003 |
ECCV (46) | 7 |
| 2024 | Fine-Grained Dynamic Network for Generic Event Boundary Detection
Ziwei Zheng, Lijun He 0001, Le Yang 0007, Fan Li 0003 |
ECCV (43) | 4 |
| 2024 | Energy-based Active Learning for Bringing Beam-induced Domain Gap for 3D Object DetectionabstractIn many real-world applications, 16-beam LiDAR-based 3D object detection (3DOD) is indispensable in scene understanding. However, the absence of well-labeled large-scale 16-beam LiDAR datasets impedes the development of these 3DOD methods. To avoid annotation costs in developing datasets, we proposed an energy-based active learning method for cross-beam domain adaptation, which effectively transfers the knowledge from the existing well-labeled 64-beam counterpart. Specifically, the cross-beam domain gap between the source (64-beam) and the target (16-beam) domain is reduced by aligning the deep features based on an energy-based feature-matching loss term during training. Moreover, the proposed energy-based active learning method enables the sampling strategy to shed light on selecting the most valuable 16-beam target samples to be manually labeled, which are then added to the training set. Experimental results show that our method can effectively transfer the knowledge from the 64-beam domain to the 16-beam one, and successfully learns a high-performance 16-beam 3DOD model with only a small portion of unlabeled data to annotate. Le Yang 0007, Yixuan Yan, Hao Cheng 0015, Fan Li 0003 |
MobiCom | 5 |
| 2024 | Multi-class Token-Guided End-to-End Weakly Supervised Image Semantic Segmentation Method
Lijun He 0001, Fan Li 0003 |
PRCV (13) | 4 |
| 2024 | Secure Video Offloading in Multi-UAV-Enabled MEC Networks: A Deep Reinforcement Learning Approach
Tantan Zhao, Fan Li 0003, Lijun He 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Class-Aware Prediction: A Solution to Center Point Collision for Anchor-Free Object Detection in Aerial ImagesabstractObject detection in aerial images has emerged as a critically important field in recent years, with anchor-free detectors garnering considerable attention from researchers. However, these detectors often encounter challenges when objects of different classes overlap, leading to false detection due to key point collisions. To solve this problem, we propose a novel method called class aware prediction, which performs bounding box predictions across all potential categories. However, class aware prediction requires detector to perform bounding box regression and prediction at the appropriate class simultaneously, which greatly aggravates the regressing difficulty of branches. Therefore, a gaussian prediction method is introduced to simplify the regression difficulty by predicting the distribution of parameters with a re-designed Negative Log-likelihood Loss. Extensive experiments conducted on the DOTA dataset have demonstrated the efficacy of our proposed methods. Our method equipping ResNet-101 obtains 79.49% mAP on the challenging DOTA dataset, achieving the top ranked accuracy among the mainstream single-stage object detector in the field of aerial image object detection. Addition results on DIOR-R dataset also show the effectiveness of our method. The results indicate a marked improvement in the detection of certain object classes that are prone to overlap in aerial images, underscoring the significance of our contributions to the field. Yixing Yong, Jian Wang 0113, Fan Li 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Toward Robust LiDAR-Camera Fusion in BEV Space via Mutual Deformable Attention and Temporal AggregationabstractLiDAR and camera are two critical sensors that can provide complementary information for accurate 3D object detection. Most works are devoted to improving the detection performance of fusion models on the clean and well-collected datasets. However, the collected point clouds and images in real scenarios may be corrupted to various degrees due to potential sensor malfunctions, which greatly affects the robustness of the fusion model and poses a threat to safe deployment. In this paper, we first analyze the shortcomings of most fusion detectors, which rely mainly on the LiDAR branch, and the potential of the bird’s eye-view (BEV) paradigm in dealing with partial sensor failures. Based on that, we present a robust LiDAR-camera fusion pipeline in unified BEV space with two novel designs under four typical LiDAR-camera malfunction cases. Specifically, a mutual deformable attention is proposed to dynamically model the spatial feature relationship and reduce the interference caused by the corrupted modality, and a temporal aggregation module is devised to fully utilize the rich information in the temporal domain. Together with the decoupled feature extraction for each modality and holistic BEV space fusion, the proposed detector, termed RobBEV, can work stably regardless of single-modality data corruption. Extensive experiments on the large-scale nuScenes dataset under robust settings demonstrate the effectiveness of our approach. Jian Wang 0113, Fan Li 0003, Yi An, Xuchong Zhang, Hongbin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | VADiffusion: Compressed Domain Information Guided Conditional Diffusion for Video Anomaly DetectionabstractThe demand for security surveillance has grown exponentially, making video anomaly detection particularly crucial. Existing image-domain based anomaly detection algorithms face implementation challenges due to several drawbacks, including latency during long-distance transmission, the need for complete decoding, and the complexity of network inference structures. Moreover, current frame prediction methods using generative models suffer from low prediction quality and mode collapse. To tackle these challenges, we propose VADiffusion, a compressed domain information guided conditional diffusion framework. VADiffusion adopts a dual-branch structure that combines motion vector reconstruction and I-frame prediction, effectively addressing the limitations of the reconstruction method in identifying sudden anomalies and the struggles of the frame prediction method in detecting persistent anomalies. Furthermore, our proposed framework incorporates the diffusion model into the realm of video anomaly detection, thereby improving the stability and accuracy of the model. Specifically, we employ sparse sampling of the compressed video, utilizing I-frames to capture appearance information and motion vectors to represent motion-related details. Different from the existing independent two-branch mechanism, we adopt a reconstruction-assisted prediction strategy, leveraging I-frames and the reconstructed motion vectors from the reconstruction branch as conditions for the diffusion model utilized in frame prediction. Ultimately, we perform decision fusion of reconstruction and prediction branches to determine anomalies. Through extensive experiments, we demonstrate that our algorithm achieves an effective trade-off between detection accuracy and model complexity. Hao Liu 0059, Lijun He 0001, Miao Zhang 0041, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Dynamic Spatial Focus for Efficient Compressed Video Action RecognitionabstractRecent years have witnessed a growing interest in compressed video action recognition due to the rapid growth of online videos. It remarkably reduces the storage by replacing raw videos with sparsely sampled RGB frames and other compressed motion cues (motion vectors and residuals). However, existing compressed video action recognition methods face two main issues: First, the inefficiency caused by the usage of coarse-level information under full resolution, and second, the disturbing due to the noisy dynamics in motion vectors. To address the two issues, this paper proposes a dynamic spatial focus method for efficient compressed video action recognition (CoViFocus). Specifically, we first use a light-weighted two-stream architecture to localize the task-relevant patches for both the RGB frames and motion vectors. Then the selected patch pair will be processed by a high-capacity two-stream deep model for the final prediction. Such a patch selection strategy crops out the irrelevant motion noise in motion vectors, as well as reduces the spatial redundancy of the inputs, leading to the high efficiency of our method in the compressed domain. Moreover, we found that the motion vectors can help our method to address the possibly happened static-issue, which means that the focus patches get stuck at some regions related to static objects rather than target actions, which further improves our method. Extensive results on both the HMDB-51 and UCF-101 datasets demonstrate the effectiveness and efficiency of our method in compressed video action recognition tasks. Ziwei Zheng, Le Yang 0007, Yulin Wang 0002, Miao Zhang 0041, Lijun He 0001, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | OStr-DARTS: Differentiable Neural Architecture Search Based on Operation StrengthabstractDifferentiable architecture search (DARTS) has emerged as a promising technique for effective neural architecture search, and it mainly contains two steps to find the high-performance architecture. First, the DARTS supernet that consists of mixed operations will be optimized via gradient descent. Second, the final architecture will be built by the selected operations that contribute the most to the supernet. Although DARTS improves the efficiency of neural architecture search (NAS), it suffers from the well-known degeneration issue which can lead to deteriorating architectures. Existing works mainly attribute the degeneration issue to the failure of its supernet optimization, while little attention has been paid to the selection method. In this article, we cease to apply the widely-used magnitude-based selection method and propose a novel criterion based on operation strength that estimates the importance of an operation by its effect on the final loss. We show that the degeneration issue can be effectively addressed by using the proposed criterion without any modification of supernet optimization, indicating that the magnitude-based selection method can be a critical reason for the instability of DARTS. The experiments on NAS-Bench-201 and DARTS search spaces show the effectiveness of our method. Le Yang 0007, Ziwei Zheng, Yizeng Han, Shiji Song, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Cybern. | 6 |
| 2024 | Deep Symmetric Fusion Transformer for Multimodal Remote Sensing Data ClassificationabstractIn recent years, multimodal remote sensing data classification (MMRSC) has evoked growing attention due to its more comprehensive and accurate delineation of Earth’s surface compared to its single-modal counterpart. However, it remains challenging to capture and integrate local and global features from single-modal data. Moreover, how to fully excavate and exploit the interactions between different modalities is still an intricate issue. To this end, we propose a novel dual-branch transformer-based framework named deep symmetric fusion transformer (DSymFuser). Within the framework, each branch contains a stack of local-global mixture (LGM) blocks, to extract hierarchical and discriminative single-modal features. In each LGM block, a local-global feature mixer with learnable weights is specifically devised to adaptively aggregate the local and global features extracted with a convolutional neural network (CNN)–transformer network. Furthermore, we innovatively design a symmetric fusion transformer (SFT) that trails behind each LGM block. The elaborately designed SFT symmetrically facilitates cross-modal correlation excavation, comprehensively exploiting the complementary cues underlying heterogeneous modalities. The hierarchical construction of the LGM and SFT blocks enables feature extraction and fusion in a multilevel manner, further promoting the completeness and descriptiveness of the learned features. We conducted extensive ablation studies and comparative experiments on three benchmark datasets, and the experimental results validated the effectiveness and superiority of the proposed method. The source code of the proposed method will be available publicly athttps://github.com/HaixiaBi1982/DSymFuser. Honghao Chang, Haixia Bi, Fan Li 0003, Jocelyn Chanussot, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Unsupervised Pansharpening Based on Double-Cycle ConsistencyabstractMultispectral (MS) pansharpening can improve the spatial resolution of MS images by fusing panchromatic (PAN) images, which have important applications in the fields of smart agriculture and environmental monitoring. However, existing supervised algorithms treat the original MS images as ground truth and generate training data under Wald’s protocol, resulting in a gap between the learned degradation process of the model and reality. This leads to the model having poor generalization and impractical. Unsupervised pansharpening methods often struggle to fully explore the rich information contained in images, leading to suboptimal pansharpening outcomes. In this work, we propose an unsupervised pansharpening algorithm based on double-cycle consistency that can learn directly from the original MS images without relying on artificially simulated degradation processes. Specifically, the network with cross-domain correlation information interaction is developed to achieve a deep fusion of spatial and spectral features. To address the inaccurate degradation mechanism representation of MS images, a spatial information extraction module based on scale invariance is developed to achieve an accurate representation. Meanwhile, double-cycle consistency loss is proposed to reduce the information loss caused by simulated degradation during the cycle process. Experimental results show that this method outperforms existing unsupervised pansharpening methods in both quantitative and qualitative evaluation of full-resolution images. Lijun He 0001, Zhihan Ren 0002, Wanyue Zhang, Fan Li 0003, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Polarimetry-Inspired Contrastive Learning for Class-Imbalanced PolSAR Image ClassificationabstractIn recent years, deep neural networks have significantly boosted the performance of polarimetric synthetic aperture radar (PolSAR) image classification. However, existing deep learning-based approaches still suffer from the following limitations. First, the performance of them is subject to the availability of massive annotations which are difficult to acquire for PolSAR images. Secondly, the class imbalance in PolSAR data greatly hinders the correct classification of minority yet equally pivotal classes. To overcome the above shortcomings, we propose a polarimetry-inspired contrastive learning PolSAR image classification approach, in the hope of elevating the classification accuracy by taking advantage of the polarimetric domain knowledge. Firstly, a complex-valued contrastive learning framework is designed, via which powerful polarimetric representations are learnt without any manual annotations. Specifically, we innovatively design two distribution-inspired positive sample generation strategies, i.e., WishartPSG and NoisePSG, to enable discriminative and domain-specific representation learning. A novel hybrid anti-imbalance scheme is further devised to tackle the class imbalance issue, which combines a contextual consistency-based pseudo-label generation and a weighted feature-level synthetic data over-sampling technique. It should be highlighted that the domain knowledge of PolSAR, including the data and noise distributions, complex-valued characteristics and the spatial consistency prior, is fully exploited throughout our model design. Extensive experiments on four benchmark datasets demonstrated the effectiveness of the proposed model. For the Flevoland 1989 dataset, our method improves the overall accuracy, average accuracy and Kappa metrics by 3.54%, 6.81% and 7.29% respectively, compared to existing state-of-the-art method. Our code will be available at https://github.com/HaixiaBi1982/PiCL. Zuzheng Kuang, Haixia Bi, Fan Li 0003, Jian Sun 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Joint Spatial and Spectral Graph-Based Consistent Self-Representation for Unsupervised Hyperspectral Band SelectionabstractBand selection (BS), which effectively reduces spectral dimensionality, stands out as a leading focus within hyperspectral image (HSI) analysis. Self-representation (SR) has surfaced as a favored technique in this domain due to its applicability to BS and unsupervised nature. However, the existing SR-based BS approaches only leverage either spatial or spectral relationships, with few integrating both while concentrating on the representation level rather than the selection level. In addition, employing all spatial pixels for spatial relationship utilization leads to considerable computational complexity. Therefore, this article proposes joint spatial and spectral graph-based consistent SR (JSSGCSR) to more effectively exploit spatial and spectral relationships for BS, which separately conducts SR to handle each view of spatial and spectral graphs to better consider two different structure characteristics, and ultimately integrates two SR results to achieve a unified and robust representative band set by imposing consistent sparsity pattern on their joint representation coefficients. In addition, the spatial and spectral relationships are integrated into different data spaces, that is, spectral graph SR and spatial graph SR are, respectively, conducted in the original HSI and the segmented and pooled HSI, which not only reduces the influence of superpixel segmentation on spectral relationships, but also improves the efficiency of spatial relationship utilization. Experimental results on three benchmark datasets have demonstrated the effectiveness of the proposed JSSGCSR in HSI classification tasks. Mingyang Ma 0004, Fan Li 0003, Zhiyong Wang 0001, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Secure Video Offloading in MEC-Enabled IIoT Networks: A Multicell Federated Deep Reinforcement Learning ApproachabstractWireless video offloading in mobile-edge-computing (MEC)-enabled Industrial Internet of Things imposes a risk of exposing users' private data to eavesdroppers. It is difficult for existing secure video offloading schemes to simultaneously guarantee security, reduce latency and energy consumption in privacy-sensitive multicell scenarios where users are unwilling to offload data to other cells. In this article, a secure video offloading scheme based on multicell federated (MCF) deep reinforcement learning (DRL) is proposed to facilitate a secure, real-time, and efficient MEC network by efficient orchestration of limited resources. We formulate a collaborative optimization problem of video frame resolution and resources to minimize latency and energy consumption while maximizing the security rate subject to analytic accuracy and limited resources. To solve the formulated NP-hard problem, a MCF DRL algorithm based on the frameworks of multicell horizontal federated learning (FL) and hierarchical reward function-based twin delayed deep deterministic policy gradient (TD3) is proposed. First of all, hierarchical reward function-based TD3 is employed to solve the collaborative optimization NP-hard problem formulated for each single cell, where the optimal solution can be efficiently approached by the agent under the guidance of the innovatively designed hierarchical reward function. Then, multicell horizontal FL is applied on TD3 to obtain a model with higher model quality by averagely aggregating multiple individual TD3 models. Simulation results reveal that the proposed algorithm outperforms comparison algorithms in terms of utility, cost, latency, energy consumption, and security rate. Tantan Zhao, Fan Li 0003, Lijun He 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Adversarial Obstacle Generation Against LiDAR-Based 3D Object DetectionabstractLiDAR sensors are widely used in many safety-critical applications such as autonomous driving and drone control, and the collected data called point clouds are subsequently processed by 3D object detectors for visual perception. Recent works have shown that attackers can inject virtual points into LiDAR sensors by strategically transmitting laser pulses to them; additionally, deep visual models have been found to be vulnerable to carefully crafted adversarial examples. Therefore, a LiDAR-based perception may be maliciously attacked with serious safety consequences. In this article, we present a highly-deceptive adversarial obstacle generation algorithm against deep 3D detection models, to mimic fake obstacles within the effective detection range of LiDAR using a limited number of points. To achieve this goal, we first perform a physical LiDAR simulation to construct sparse obstacle point clouds. Then, we devise a strong attack strategy to adversarially perturb prototype points along each direction of the ray. Our method achieves a high attack success rate while complying with physical laws at the hardware level. We perform comprehensive experiments on different types of 3D detectors and determine that the voxel-based detectors are more vulnerable to adversarial attacks than the point-based methods. For example, our approach achieves an 89% mean attack success rate against PV-RCNN by using only 20 points to spoof a fake car. Jian Wang 0113, Fan Li 0003, Xuchong Zhang, Hongbin Sun 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Low-Rate Feature Compression for Collaborative Intelligence: Reducing Redundancy in Spatial and Statistical LevelsabstractTo distribute the storage and computation load caused by growing capacity of deep neural network (DNN), collaborative intelligence (CI) framework has been proposed, where a deep model is split and executed in two distributed devices respectively. Intermediate feature must be transferred from the front end to the back in order to perform distributed inference, thus transmission process is the bottleneck that influences the inference efficiency in terms of accuracy and delay. Specifically for a bandwidth-limited human-in-loop visual analysis task, feature compression approach needs exploration to reduce the data volume to be transmitted, in order to achieve low transmission delay as well as maintain analysis performance and human perception ability. In this article, the redundancy of intermediate feature both in spatial and statistical levels are firstly analyzed. A mathematical expression for the goal of feature compression is formulated, based on which a two-level redundancy removal based low-rate feature compression approach is proposed. For the front-end device, an information squeezing (IS) module is developed to squeeze the key information of input image and inject them into a low-resolution image. Then a backbone network is split into two parts with respects to the application demands of CI, and can be deployed at the front and back ends correspondingly. With a specifically designed objective function, IS module and the partitioned backbone network are optimized collaboratively to reduce the two-level redundancy, thus compressing the intermediate feature. A generative adversarial network (GAN)-based restoration module is proposed to recover an image with original resolution from the compressed feature, for satisfying human perception. Comprehensive experiments are conduct to validate the efficiency of the proposed method. Fan Li 0003, Yuan Zhang 0023 |
IEEE Trans. Multim. | 2 |
| 2023 | Complex-Valued Self-Supervised PolSAR Image Classification Integrating Attention MechanismabstractPolarimetric synthetic aperture radar (PolSAR) image classification is acknowledged as a critical task in remote sensing image processing, and the performance of this task has witnessed a substantial improvement owing to developing deep neural networks. However, existing approaches still suffer from at least one of the following limitations. First, their performance mainly depends on huge amounts of annotations. Secondly, the physical mechanism and characteristics of PolSAR data are not fully exploited in these methods. To tackle the label scarcity issue, we establish an attention mechanism-incorporated self-supervision framework for PolSAR image classification via designing a predictive auxiliary learning task. According to the properties of PolSAR data, we adapt the framework to complex-valued. Additionally, noise injection augmentation scheme considering the speckle noise distribution is designed to enhance the robustness of our model. Involving the complex-valued characteristics of the PolSAR data and noise into the architecture and loss function design makes it essentially distinctive from existing self-supervised methods. Zuzheng Kuang, Haixia Bi, Fan Li 0003 |
IGARSS | 3 |
| 2023 | Domain-specific knowledge-driven pan-sharpening algorithm
Nan Shi, Ping Wang 0009, Fan Li 0003 |
Neurocomputing | 3 |
| 2023 | DRL-Based Secure Aggregation and Resource Orchestration in MEC-Enabled Hierarchical Federated LearningabstractFederated learning (FL) provides a new paradigm for protecting data privacy by enabling model training at devices and model aggregation at servers. However, data information may be leaked to honest-but-curious aggregation servers by updated model parameters. The existing secure methods do not fully exploit the potentiality of data characteristics in enhancing security, which makes it impossible to optimize limited system resources overall to achieve secure, fair, and efficient FL systems. In this article, a DRL-based joint secure aggregation and resource orchestration scheme is proposed to guarantee security and fairness, and improve efficiency for hierarchical FL (HFL) assisted by untrusted mobile-edge computing (MEC) servers. We formulate a joint optimization problem of data size, payment, and resource orchestration, to maximize the long-term social welfare subject to secure aggregation and limited resources. Since the formulated problem is a complex mixed integer dynamic optimization problem with NP-hardness, where multiple mixed integer optimization variables are highly coupled in time-varying constraints and objective function, it is difficult to obtain its optimal solution via traditional optimization methods. Thus, we propose a hierarchical reward function-based DRL algorithm (MATD3) to guide the agents to approach the optimal policy of secure aggregation and resource orchestration. Simulation results show that the proposed algorithm MATD3 can achieve superior performance over comparison algorithms and the MEC-enabled HFL framework outperforms two-layer FL frameworks. Tantan Zhao, Fan Li 0003, Lijun He 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Sketch Assisted Face Image Coding for Human and Machine Vision: A Joint Training ApproachabstractImage coding is one of the most fundamental techniques and is widely used in image/video processing and multimedia communications. Current image coding methods are mainly human-oriented, and the visual quality is always unsatisfactory, especially at low bitrates. Moreover, the recent emergence of machine vision goes beyond the scope of current coding. With these considerations, we proposed a sketch assisted face image coding for human and machine vision by a joint training approach. In the proposed approach, we design a new feature representation: a color sketch, which aims to satisfy both low-frequency features of human vision and high-frequency features of machine analysis. Then, we present a novel end-to-end image codec framework with joint training that consists of three models: an image-to-image translation module, a coding module, and a two-stage reconstruction module. Specifically, the input image is first translated into the edge map with the Canny edge as the auxiliary label to merely preserve the structure information. Afterward, the backpropagation from reconstruction module guides the edge map to increase or decrease the information through joint training, which results in the generation of color sketch. Then, the generated sketch is compressed into the bitstream and decompressed back to a sketch in the coding module. Finally, the decompressed sketch is reconstructed to support the machine and human tasks, respectively. In this way, the color sketch is designed to bridge the gap between human and machine vision, and the joint training strategy helps to adjust the low-frequency information in the sketch. The experimental results on challenge datasets demonstrate that our proposed algorithm offers 40.9%-86.6% bitrate savings on machine vision and is comparable to state-of-the-art image coding methods on human vision. Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Multimodal Mutual Attention-Based Sentiment Analysis Framework Adapted to Complicated ContextsabstractSentiment analysis has broad application prospects in the field of social opinion mining. The openness and invisibility of the internet makes users’ expression styles more diverse and thus results in the blooming of complicated contexts in which different unimodal data have inconsistent sentiment tendencies. However, most sentiment analysis algorithms only focus on designing multimodal fusion methods without preserving the individual semantics of each unimodal data. To avoid misunderstandings caused by ambiguity and sarcasm in complicated contexts, we propose a multimodal mutual attention-based sentiment analysis (MMSA) framework adapted to complicated contexts, which consists of three levels of subtasks to preserve the unimodal unique semantics and enhance the common semantics, to mine the association between unique semantics and common semantics and to balance decisions from unique and common semantics. In the framework, a multiperspective and hierarchical fusion (MHF) module is developed to fully fuse multimodal data, in which different modalities are mutually constrained and the fusion order is adjusted in the next step to enhance cross-modal complementarity. To balance the data, we calculate the loss by applying different weights to positive and negative samples. The experimental results on the CH-SIMS multimodal dataset show that our method outperforms existing multimodal sentiment analysis algorithms.The code of this work is available athttps://gitee.com/viviziqing/mmsacode. Lijun He 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Spectral Correlation-Based Diverse Band Selection for Hyperspectral Image ClassificationabstractBand selection which can reduce the spectral dimensionality effectively, has become one of the most popular topics in hyperspectral image (HSI) analysis. Recently, sparse representation based band selection (BS) has emerged as a popular tool. The existing sparse models mainly focus on minimizing reconstruction error and sparsity, while do not fully exploit the unique correlations among hundreds of continuous bands, which may cause representative bands missed and highly-correlated bands selected. Therefore, this paper proposes the spectral correlation based diverse band selection (SCDBS) for HSIs to improve representativeness and diversity of the selected bands. Specifically, a correlation derived weight is used to perform weighted sparse reconstruction to select the bands that are more correlated to the whole HSI, and a correlation minimization term is designed to remove the highly-correlated bands simultaneously. In addition, the proposed method imposes an adjustable sparse constraint by using an ℓ2,0 Mingyang Ma 0004, Shaohui Mei, Fan Li 0003, Yaoyang Ge, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | DRL-Based Joint Resource Allocation and Device Orchestration for Hierarchical Federated Learning in NOMA-Enabled Industrial IoTabstractFederated learning (FL) provides a new paradigm for protecting data privacy in Industrial Internet of Things (IIoT). To reduce network burden and latency brought by FL with a parameter server at the cloud, hierarchical federated learning (HFL) with mobile edge computing (MEC) servers is proposed. However, HFL suffers from a bottleneck of communication and energy overhead before reaching satisfying model accuracy as IIoT devices dramatically increase. In this article, a deep reinforcement learning (DRL)-based joint resource allocation and IIoT device orchestration policy using nonorthogonal multiple access is proposed to achieve a more accurate model and reduce overhead for MEC-assisted HFL in IIoT. We formulate a multiobjective optimization problem to simultaneously minimize latency, energy consumption, and model accuracy under the constraints of computing capacity and transmission power of IIoT devices. To solve it, we propose a DRL algorithm based on deep deterministic policy gradient. Simulation results show proposed algorithm outperforms others. Tantan Zhao, Fan Li 0003, Lijun He 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | End-to-End Blind Video Quality Assessment Based on Visual and Memory Attention ModelingabstractDeveloping an objective quality assessment model for user-generated content (UGC) videos is significant for multimedia applications, and also a challenge due to the diversity of video content and unpredictability of distortions. To predict the perceived quality, it is necessary to consider the human visual system, in which attention in visual and memory domains is an essential component. With the idea that the stimulus-driven bottom-up mechanism and cognition-driven top-down mechanism work in synergy to generate quality-aware attention, we propose an end-to-end blind video quality assessment (VQA) algorithm based on visual and memory attention modeling. First, a quality-aware visual attention module is established to obtain spatial-temporal attention-guided representations for frame-level quality perception. Specifically, an attention selection and confluence method is developed by circularly integrating the quality-aware attention information to spatial-temporal content features. Then, with the aid of a quality-aware memory attention module, the video-level attention-guided features are inferred through the dimension and attention reshaping of frame-level representations. The video quality is predicted with the guidance of frame-level visual attention and video-level memory attention in an end-to-end structure. Experimental results on five UGC-VQA databases (CVD2014, LIVE-Qualcomm, KoNViD-1 k, LIVE-VQC and Youtube-UGC) demonstrate the effectiveness of our modules. Xiaodi Guan, Fan Li 0003, Yangfan Zhang, Pamela C. Cosman |
IEEE Trans. Multim. | 2 |
| 2022 | Mutual Learning Inspired Prediction Network for Video Anomaly Detection
Yuan Zhang 0023, Fan Li 0003, Lu Yu 0003 |
PRCV (3) | 3 |
| 2022 | Unsupervised defect inspection algorithm based on cascaded GAN with edge repair feature fusion
Lijun He 0001, Nan Shi, Kainnat Malik, Fan Li 0003 |
Appl. Intell. | 4 |
| 2022 | MTRFN: Multiscale Temporal Receptive Field Network for Compressed Video Action Recognition at Edge ServersabstractWith the wide deployment of Internet of Things monitoring terminals, a tremendous number of videos are accumulated continuously. Big data processing and analysis-based action recognition has an increasingly important role in making cities simpler, better, and smarter. The traditional cloud server-centered analysis mode has to spend extra time transmitting vast video data terminals to remote cloud servers, which is always violated in real implementation. Edge servers with limited caching and computation capacities near the monitoring terminals enable implementation. However, due to the dependency on training data and the high complexity of extracting information and network architecture, existing image domain-based methods cannot be implemented at edge servers. Moreover, recognizing actions with different durations is still challenging. Due to these issues, we extend the traditional image domain to the compressed domain to efficiently extract the information of$I$frames and physical knowledge motion vectors (MVs), which can reflect the multiscale temporal feature just by partial decoding. To recognize the actions with different durations, a multiscale temporal receptive field network (MTRFN), including short-term and long-term branches, is proposed to simultaneously capture the action’s instant change based on the extracted MVs, the long temporal feature between adjacent$I$frames, and the interaction between them. The results show that our algorithm can achieve a better balance between accuracy and computational complexity. Lijun He 0001, Miao Zhang 0041, Sijin Zhang, Fan Li 0003 |
IEEE Internet Things J. | 5 |
| 2022 | Revenue and Energy Efficiency-Driven Delay-Constrained Computing Task Offloading and Resource Allocation in a Vehicular Edge Computing Network: A Deep Reinforcement Learning ApproachabstractFor in-vehicle application, task type and vehicle state information, i.e., vehicle speed, bear a significant impact on the task delay requirement. However, the joint impact of task type and vehicle speed on the task delay constraint has not been studied, and this lack of study may cause a mismatch between the requirement of the task delay and allocated computation and wireless resources. In this article, we propose a joint task type and vehicle speed-aware task offloading and resource allocation strategy to decrease the vehicle’s energy cost for executing tasks and increase the revenue of the vehicle for processing tasks within the delay constraint. First, we establish the joint task type and vehicle speed-aware delay constraint model. Then, the delay, energy cost, and revenue for task execution in the vehicular edge computing (VEC) server, local terminal, and terminals of other vehicles are calculated. Based on the energy cost and revenue from task execution, the utility function of the vehicle is acquired. Next, we formulate a joint optimization of task offloading and resource allocation to maximize the utility level of the vehicles subject to the constraints of task delay, computation resources, and wireless resources. To obtain a near-optimal solution of the formulated problem, a joint offloading and resource allocation based on the multiagent deep deterministic policy gradient (JORA-MADDPG) algorithm is proposed to maximize the utility level of vehicles. Simulation results show that our algorithm can achieve superior performance in task completion delay, vehicles’ energy cost, and processing revenue. Lijun He 0001, Xing Chen 0007, Fan Li 0003 |
IEEE Internet Things J. | 5 |
| 2022 | Human-Machine Interaction-Oriented Image Coding for Resource-Constrained Visual Monitoring in IoTabstractVisual monitoring supported by the Internet of Things (IoT) increasingly relies on analyzing a mass of image data with human–machine interactive mechanisms. However, maintaining the efficiency of such a monitoring system in complex environments with energy or bandwidth constraints poses challenges. While both human perception and machine analysis performance should be satisfied, transmitting extra information to satisfy both should be avoided for efficient resource utilization. To this end, we propose a human–machine interaction-oriented image coding (HMI-IC) framework based on deep learning. In this framework, machines should provide early monitoring messages consisting of analysis results and preview images, and then humans can additionally request high-quality images of objects of interest. Each collected image is compressed into a layered data stream by HMI-IC to fulfill the demands of analysis, preview visualization, and high-quality reconstruction. Adaptive coding transmission can fit different demands in two stages, according to resource constraints. Experimental results show that both accuracy and inference speed on compressed images are improved by our method, with entire coding efficiency comparable to JPEG2000. To validate HMI-IC’s efficiency in practical terms, we provide two use cases (energy constrained and bandwidth constrained) for visual monitoring. Fan Li 0003, Jing Xu 0003, Pamela C. Cosman |
IEEE Internet Things J. | 2 |
| 2022 | DRL-Based Secure Video Offloading in MEC-Enabled IoT NetworksabstractWireless offloading in mobile-edge-computing (MEC)-enabled Internet of Things (IoT) networks inevitably suffers the risk of eavesdropping. Physical-layer security (PLS) approaches can be applied to prevent eavesdropping. However, the existing PLS techniques are not well targeted for videos due to the fact that video’s distortion characteristics, which allow encoding parameters to be flexibly adjusted to enhance security in offloading, are ignored. A deep reinforcement learning (DRL)-based real-time, secure, and efficient video offloading scheme is proposed in this article, where video frame resolution, one key parameter of video’s distortion characteristics, is introduced and jointly optimized with PLS scheme to guarantee video’s security, improve users’ Quality of Experience (QoE) and save energy consumption. We formulate a joint optimization problem of video frame resolution selection, computation offloading, and resource allocation strategy, to minimize energy consumption and maximize QoE in terms of delay and analytic accuracy, while subject to security rate, computing capability, and transmission power. To solve the formulated NP-hard problem with the form of high-dimensional nonlinear mixed-integer programming, the hierarchical reward-function-based DRL (JVFRS-CO-RA-MADDPG) algorithm is proposed to guide the agents to obtain the optimal policy efficiently. Finally, the simulation results show that the proposed algorithm outperforms the existing algorithms in terms of delay, energy consumption, and security level. Tantan Zhao, Lijun He 0001, Fan Li 0003 |
IEEE Internet Things J. | 4 |
| 2022 | A novel transmission approach based on video content for 360-degree streaming
Fan Li 0003, Zhisheng Yan |
Multim. Tools Appl. | 2 |
| 2022 | Study of Subjective and Objective Quality Assessment of Night-Time VideosabstractWith the widespread usage of video capture devices and social media videos, videos are dominating the multimedia landscape. There is an emerging need for video quality assessment (VQA) that forms the backbone of advanced video systems. Night-time videos play an important role in user capturing, hence being able to accurately assess their quality is critical. However, the characteristics of night-time videos differ from those of general in-capture videos; and VQA algorithms that have been developed for general-purpose videos cannot accurately assess the quality of night-time videos. Research is needed to gain a better understanding of how humans perceive the quality of night-time videos, and use this new understanding to develop reliable VQA algorithms. To this end, we construct a large-scale night-time VQA database, namely Mobile In-capture Night-time Database for Video Quality (MIND-VQ), containing 1181 night-time videos, 435 subjects, and over 130000 opinion scores. We perform thorough analyses to reveal subjective quality assessment behaviors of night-time videos. Furthermore, we propose a new VQA model, namely Visibility-based Night-time Video Quality Assessment Network, VINIA. Spatial and temporal visibility-aware components are characterized to reflect properties of human perception of night-time VQA task. A series of experiments are conducted to compare our VINIA with other existing VQA algorithms using our new MIND-VQ database and other public VQA databases. Experimental results show that our subjective VQA database provides new insights and our new VINIA model achieves superior performance in accessing night-time video quality. Xiaodi Guan, Fan Li 0003, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Gaussian Focal Loss: Learning Distribution Polarized Angle Prediction for Rotated Object Detection in Aerial ImagesabstractWith the increasing availability of aerial data, object detection in aerial images has aroused more and more attention in remote sensing community. The difficulty lies in accurately predicting the angular information for each target when using the oriented bounding boxes to represent the arbitrary oriented objects, as the periodicity of the angle could cause inconsistency between target angle values. To resolve the problem, recent works propose to perform angular prediction from a regression problem to a classification task with circular smooth label. However, we find that current loss functions applying to binary soft labels need to approximate the soft label values at each position. When summed over all the negative angle categories, these relatively insignificant loss values can overwhelm the target angle category, thus preventing the network from predicting precise angle information. In this paper, we propose a novel loss function that acts as a more effective alternative to the classification-based rotated detectors. By constructing the classification loss with adaptive Gaussian attenuation on the negative locations, our training objective can not only avoid discontinuous angle boundaries but also enable the network to obtain more accurate angle predictions with higher response at peaks. Moreover, an aspect ratio-aware factor was proposed based on our loss function to enhance the robustness of the model for determining the orientation for square-like objects. Extensive experiments on aerial image datasets DOTA, HRSC2016, and UCAS-AOD demonstrated the effectiveness and superior performances of our approaches. Jian Wang 0113, Fan Li 0003, Haixia Bi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Learning-Based Rate Control for Video-Based Point Cloud CompressionabstractDue to limited transmission resources and storage capacity, efficient rate control is important in Video-based Point Cloud Compression (V-PCC). In this paper, we propose a learning-based rate control method to improve the rate-distortion (RD) performance of V-PCC. A low-latency synchronous rate control structure is designed to reduce the overhead of pre-coding. The basic unit (BU) parameters are predicted accurately based on our proposed CNN-LSTM neural network, instead of the online updating approach, which can be inaccurate due to low consistency between adjacent 2D frames in V-PCC. When determining the quantization parameters for the BU, a patch-based clipping method is proposed to avoid unnecessary clipping. This approach is able to improve the RD performance and subjective dynamic point cloud quality. Experiments show that our proposed rate control method outperforms present approaches. Taiyu Wang, Fan Li 0003, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2022 | Learning-Based Scalable Image Compression With Latent-Feature Reuse and PredictionabstractRecently, learning-based image compression model has attracted much attention due to its impressive performance and ease of optimization, compared with traditional DCT and wavelet-based image compression standards. Most learning-based image compression models are trained to minimize joint rate-distortion (RD) loss on one single RD trade-off point. However, in many multimedia applications, due to communication constraints, or display adaptation needs for different spatial formats, bit rates or power, it is necessary to provide a variety of image versions for different client devices. To fulfill this requirement, typical end-to-end image compression methods have to compress an image into several bit streams independently by a number of pre-trained networks, which are resource-consuming because of redundancy among these streams. To address this problem, inspired by traditional scalable video coding framework, we propose a learning-based end-to-end quality and spatial scalable image compression (QSSIC) model in multi-layer structure, in which each layer could generate one bitstream corresponding to a specified resolution and image fidelity. This scalability is achieved by exploring the potential of feature-domain representation prediction and reuse. To be specific, firstly, bitstreams of previous layers are used to predict the current layer representations which contains the enhancement information, and then only prediction residuals need to be coded in enhancement layers. Secondly, previous bitstreams are reused in image reconstruction in higher layers to provide basic information. The proposed model could be optimized in an end-to-end manner. Extensive experiments show that our method outperforms state-of-art deep neural networks (DNN)-based auto-encoders in simulcast scenarios. In addition, our method has a better performance than the traditional scalable image compression method scalable extension of H.264/AVC (SVC) and is comparable to scalable extension of H.265/HEVC (SHVC). Yixin Mei, Li Li 0040, Zhu Li 0001, Fan Li 0003 |
IEEE Trans. Multim. | 4 |
| 2021 | Learn A Compression for Objection Detection - VAE with a BridgeabstractRecent advances in sensor technology and wide deployment of visual sensors lead to a new application whereas compression of images are not mainly for pixel recovery for human consumption, instead it is for communication to cloud side machine vision tasks like classification, identification, detection and tracking. This opens up new research dimensions for a learning based compression that directly optimizes loss function in vision tasks, and therefore achieves better compression performance vis-a-vis the pixel recovery and then performing vision tasks computing. In this work, we developed a learning based compression scheme that learns a compact feature representation and appropriate bitstreams for the task of visual object detection. Variational Auto-Encoder (VAE) framework is adopted for learning a compact representation, while a bridge network is trained to drive the detection loss function. Simulation results demonstrate that this approach is achieving a new state-of-the-art in task driven compression efficiency, compared with pixel recovery approaches, including both learning based and handcrafted solutions. Yixin Mei, Fan Li 0003, Li Li 0040, Zhu Li 0001 |
VCIP | 2 |
| 2021 | Vehicle theft recognition from surveillance video based on spatiotemporal attention
Lijun He 0001, Shuai Wen, Fan Li 0003 |
Appl. Intell. | 4 |
| 2021 | Feature fusion quality assessment model for DASH video streamingabstractAbstract Dynamic Adaptive Streaming over HTTP (DASH) employs the flexible rate adaptation scheme to combat with time‐varying channel conditions. In addition to compression impairment, DASH video streaming suffers from transmission impairment, such as rate switches and stalling events. Both of them severely degrade users' Quality of Experience (QoE). Herein, an assessment model is established for DASH video streaming by directly jointing multiple QoE influential factors, which quantify impairments resulting from the compression and transmission. To demonstrate the influence of video content characteristics on users' QoE, spatio‐temporal content perceptual features are employed to represent the compression impairment. When reflecting temporal characteristics, a novel motion vector padding method is proposed to quantify the influence of intra macroblock on the human visual system. The proposed model is evaluated on a newly public Waterloo SQoE‐III database, which is available for DASH video streaming. Experimental results demonstrate that the authors' model outperforms the comparative models and owns the strong generalization ability to different video contents. Moreover, the proposed model is statistically superior to the existing models. Fan Li 0003, Zhisheng Yan |
IET Image Process. | 2 |
| 2021 | Efficient attention based deep fusion CNN for smoke detection in fog environment
Lijun He 0001, Xiaoli Gong, Sirou Zhang, Fan Li 0003 |
Neurocomputing | 5 |
| 2021 | Deep image compression with multi-stage representation
Guiguang Ding, Jungong Han, Fan Li 0003 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Convolutional neural network based low complexity HEVC intra encoder
Fan Li 0003 |
Multim. Tools Appl. | 2 |
| 2021 | MMMNet: An End-to-End Multi-Task Deep Convolution Neural Network With Multi-Scale and Multi-Hierarchy Fusion for Blind Image Quality AssessmentabstractAs the evaluation of image quality depends on the human visual system (HVS), many existing image quality assessment (IQA) methods focus on modeling the HVS to account for subjective perception. The visual attention of the HVS makes humans more sensitive to distortion on the attended regions than on regions which are not the focus of attention. Therefore, we propose an end-to-end multi-task deep convolution neural network with multi-scale and multi-hierarchy fusion (MMMNet), in which the IQA and saliency subtasks are jointly optimized to improve saliency-guided IQA performance. Particularly, the incorporation of saliency information is achieved by fusing saliency features with IQA features hierarchically to progressively improve the IQA features over network depth. A multi-scale feature extraction module (MSFE) is proposed to provide effective saliency features for the IQA network. Based on the saliency fusion, MMMNet introduces an auxiliary saliency task, achieving the multi-task learning to improve the generalization of the IQA task. Experimental results show that MMMNet achieves state-of-the-art performance and strong generalization ability on IQA databases. Fan Li 0003, Yangfan Zhang, Pamela C. Cosman |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Low-Complexity Error Resilient HEVC Video Coding: A Deep Learning ApproachabstractIntra/inter switching-based error resilient video coding effectively enhances the robustness of video streaming when transmitting over error-prone networks. But it has a high computation complexity, due to the detailed end-to-end distortion prediction and brute-force search for rate-distortion optimization. In this article, a Low Complexity Mode Switching based Error Resilient Encoding (LC-MSERE) method is proposed to reduce the complexity of the encoder through a deep learning approach. By designing and training multi-scale information fusion-based convolutional neural networks (CNN), intra and inter mode coding unit (CU) partitions can be predicted by the networks rapidly and accurately, instead of using brute-force search and a large number of end-to-end distortion estimations. In the intra CU partition prediction, we propose a spatial multi-scale information fusion based CNN (SMIF-Intra). In this network a shortcut convolution architecture is designed to learn the multi-scale and multi-grained image information, which is correlated with the CU partition. In the inter CU partition, we propose a spatial-temporal multi-scale information fusion-based CNN (STMIF-Inter), in which a two-stream convolution architecture is designed to learn the spatial-temporal image texture and the distortion propagation among frames. With information from the image, and coding and transmission parameters, the networks are able to accurately predict CU partitions for both intra and inter coding tree units (CTUs). Experiments show that our approach significantly reduces computation time for error resilient video encoding with acceptable quality decrement. Taiyu Wang, Fan Li 0003, Xiaoya Qiao, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2021 | TTL-IQA: Transitive Transfer Learning Based No-Reference Image Quality AssessmentabstractImage quality assessment (IQA) based on deep learning faces the overfitting problem due to limited training samples available in existing IQA databases. Transfer learning is a plausible solution to the problem, in which the shared features derived from the large-scale Imagenet source domain could be transferred from the original recognition task to the intended IQA task. However, the Imagenet source domain and the IQA target domain as well as their corresponding tasks are not directly related. In this paper, we propose a new transitive transfer learning method for no-reference image quality assessment (TTL-IQA). First, the architecture of the multi-domain transitive transfer learning for IQA is developed to transfer the Imagenet source domain to the auxiliary domain, and then to the IQA target domain. Second, the auxiliary domain and the auxiliary task are constructed by a new generative adversarial network based on distortion translation (DT-GAN). Furthermore, a TTL network of the semantic features transfer (SFTnet) is proposed to optimize the shared features for the TTL-IQA. Experiments are conducted to evaluate the performance of the proposed method on various IQA databases, including the LIVE, TID2013, CSIQ, LIVE multiply distorted and LIVE challenge. The results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, our proposed method demonstrates a strong generalization ability. Fan Li 0003, Hantao Liu |
IEEE Trans. Multim. | 2 |
| 2020 | A More Refined Mobile Edge Cache Replacement Scheme For Adaptive Video Streaming With Mutual Cooperation In Multi-Mec ServersabstractInstead of only focusing on the hit ratio of the videos cached in Mobile Edge Computing (MEC) server, we propose a more refined video segment content and client statusbased MEC cache update strategy, to improve clients' Quality of Experience (QoE). First, based on both the segment popularity and importance, we divide MEC cache into three parts which can be flexibly transformed into each other by combing the requested times of segments, transmission capability and clients' playback status together. Furthermore, we present the client's cache priority utility function and formulate a problem to maximize the utility function subject to the constraints of MEC cache size and transmission capacity. The brand and branch method is employed to obtain the optimal solution. Simulation results show that our algorithm can improve system throughput, hit ratio of video segment, playback frozen time as well as backhaul traffic. Lijun He 0001, Xing Chen 0007, Guizhong Liu, Fan Li 0003 |
ICME | 5 |
| 2020 | Deep feature importance awareness based no-reference image quality prediction
Fan Li 0003, Hantao Liu |
Neurocomputing | 2 |
| 2019 | A Comparative Study of DNN-Based Models for Blind Image Quality PredictionabstractRecently, deep learning methods have gained substantial attention in the research community and have proven useful for blind image quality assessment (BIQA). Although previous study of deep neural networks (DNN) methods is presented, some novelty methods, which are recently proposed, are not summarized. In this paper, we provide a comparative study on the application of DNN methods for BIQA. First, we systematically analyze the existing DNN-based quality assessment methods. Then, we compare the predictive performance of various methods in synthetic and authentic databases, providing important information that can help understand the underlying properties between different methods. Finally, we describe some emerging challenges in designing and training DNN-based BIQA, along with few directions that are worth further investigations in the future. Fan Li 0003, Hantao Liu |
ICIP | 2 |
| 2019 | Joint rate adaptation and resource allocation for real-time H.265/HEVC video transmission over uplink OFDMA systems
Fan Li 0003, Taiyu Wang, Pamela C. Cosman |
Multim. Tools Appl. | 1 |
| 2018 | Message-Prioritization Based Unequal Secrecy Protection for Untrusted Two-Way Relaying NetworksabstractThis paper proposes a message-prioritization based unequal secrecy protection framework for untrusted two- way relaying networks, where two terminal users communicate bidirectionally with the assistance of an untrusted relay. Each user is assumed to have two messages with distinct priorities: high priority and low priority. The high priority messages (HPM) are first transmitted, for which we devise a constellation overlapping method such that the received signals at the relay overlap with each other and a high error floor is created to prevent the untrusted relay from deciphering the information. Upon the completion of HPM exchange, users transmit their low priority messages (LPM) using a noise aggregation approach. To be specific, each user superposes its LPM onto the previously decoded HPM from the other user. By exploiting the difference between the error patterns for HPM at the terminal users and the relay, channel noises in various time slots can be aggregated at the relay to secure the LPM transmission. It is shown from simulation results that HPM is guaranteed to have greater reliability and higher secrecy level than that of LPM. However, the transmission of LPM enjoys lower implementation complexity and reduced system overhead, which fully demonstrates that the proposed unequal secrecy protection scheme can realize a good performance-complexity tradeoff. Li Sun 0001, Fan Li 0003 |
ICC | 3 |
| 2018 | Towards Enhanced Security for Two-Way Untrusted Relaying Systems: A Constellation Overlapping SchemeabstractThis paper proposes a constellation overlapping scheme to secure two-way untrusted relaying systems, where the relay acts as both a helper facilitating data transmission and an eavesdropper intercepting users' messages. A truncated-channel-inversion based approach is developed to make the signals transmitted from two users experience the same equivalent channel, thereby realizing full constellation overlapping at the relay. Consequently, it is extremely difficult for the relay to recover the users' individual signals, and data confidentiality is thus protected. The achieved error floor level at the untrusted relay is analyzed, and the truncation thresholds are optimized to maximize the sum rate for end-to-end information exchange. Simulation results demonstrate the superiority of our scheme in terms of security and transmission efficiency compared with the existing alternatives. Li Sun 0001, Fan Li 0003 |
ICC | 3 |
| 2018 | Scene-Aware Soccer Video QoE Assessment - A Compressed-Domain ApproachabstractThe small screen of mobile devices and bandwidth limitations of communication networks greatly affect users' quality of experience (QoE), especially for soccer video, which is characterized by rapid movement and small objects. In this paper, a Compressed-domain Soccer Video Quality assessment Model (CSVQM) is proposed based on the fact that soccer video includes three distinct scene types, which cause different concerns for viewers. To reduce complexity and operate in real-time, all model parameters are derived from the compressed video stream without resorting to complete video decoding. The validation shows that CSVQM significantly outperforms conventional models in terms of accuracy, consistency, and complexity. Fan Li 0003, Yixin Mei, Pamela C. Cosman |
ICME | 1 |
| 2018 | A Cost-Constrained Video Quality Satisfaction Study on Mobile DevicesabstractMobile videos on smartphones have been widely used and enjoyed by increasing numbers of people in recent years; however, due to the high bit rates of video streams, viewers must pay high data communication costs when watching videos over cellular networks. Thus, the viewers' quality of experience (QoE) for streaming videos is influenced by their psychological states under this cost pressure. This paper evaluates the cost-constrained video quality. First, a cost-constrained QoE framework that considers technical factors, content types and cost pressure is developed. Then, a cost-constrained video quality satisfaction (CVQS) model is proposed to assess users' video quality satisfaction (QS) considering the three aspects. Finally, the proposed model is verified by empirical analysis, mathematical fitting and experimental simulations. The proposed CVQS model maximizes the QS of videos by considering the cost to users, the users' personality profiles and objective video quality. An example application that employs the CVQS model is provided to demonstrate the applicability and effectiveness of the model. Fan Li 0003, Fu Shuang, Xueming Qian |
IEEE Trans. Multim. | 1 |
| 2017 | Object tracking via a cooperative appearance model
Fan Li 0003 |
Knowl. Based Syst. | 1 |
| 2017 | Compressed-domain-based no-reference video quality assessment model considering fast motion and scene change
Fan Li 0003 |
Multim. Tools Appl. | 2 |
| 2016 | Region-of-interest based rate control algorithm for H.264/AVC video coding
Fan Li 0003 |
Multim. Tools Appl. | 1 |
| 2015 | Packet importance based scheduling strategy for H.264 video transmission in wireless networks
Fan Li 0003 |
Multim. Tools Appl. | 1 |
| 2013 | Multiuser multimedia communication over orthogonal frequency-division multiple access downlink systemsabstractSUMMARY This paper investigates the problem of multiuser multimedia communication over orthogonal frequency‐division multiple‐access downlink systems. A cross‐layer design is proposed to maximize the received video quality of all the users subject to the network resource constraint. With the optimal joint subcarrier assignment and power allocation, the proposed scheme can maximally satisfy the requirements of the packet scheduling from the higher layer. Employing the Lagrange dual decomposition method, we can obtain the global optimal solution to the optimization problem. Simulation results show that the proposed algorithm has superior performance compared with the existing alternatives. Copyright © 2012 John Wiley & Sons, Ltd. Fan Li 0003 |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Joint Sensing and Transmission for AF Relay Assisted PU Transmission in Cognitive Radio NetworksabstractRecent measurements show that there are abundant spectrum opportunities across the licensed cellular bands, over which the relay stations (RSs) are widely employed for both coverage extension and throughput improvement. However, current efforts in cognitive radios focus on exploiting spectrum holes over the licensed direct transmission links. To exploit new spectrum opportunities over the licensed relay assisted transmission links, we propose a novel two-phase joint sensing and transmission scheme (TP-JSTS) for the dual-hop Amplify-and-Forward (AF) relay assisted primary user (PU) transmission in cognitive radio networks. The TP-JSTS takes into account the relay behaviors of the PU system. In the first phase, the secondary user (SU) senses the signal transmitted by the PU base station (BS). In the second phase, the SU detects the signals sent by the PU-Relays. To protect the PU from harmful interference, the SU transmits within the current band in the rest of each phase only when the PU is detected to be absent. We investigate the sensing performance, throughput performance and delay performance of our proposed TP-JSTS when all the PU-Relays adopt the AF relay protocol. We obtain the optimal sensing-time allocation strategy that maximizes the achievable throughput of the SU and the optimal sensing-time allocation strategy that minimizes the average transmission delay of the SU. Simulation results show that there exists an optimal pair of sensing-time durations that maximizes the achievable throughput for the SU, while there exists another optimal pair of sensing-time durations that minimizes the average transmission delay for the SU. Wenshan Yin, Pinyi Ren, Fan Li 0003, Qinghe Du |
IEEE J. Sel. Areas Commun. | 3 |
| 2012 | Cross-layer based power allocation over cognitive wireless relay link with statistical delay QoS guaranteesabstractSUMMARY In this paper, we propose a cross‐layer based power allocation scheme with statistical delay QoS guarantees for the cognitive (secondary) amplify‐and‐forward relay link, which coexists with one primary link by sharing particular portion of the spectrum. Specifically, our derived power allocation scheme aims at maximizing the effective capacity of the cognitive relay link, which can be seen as the maximum arrival rate supported by the system under given QoS constraints. In our work, not only the average total transmit power and average interference power constraints are considered, but also the impact of the interference from the primary link to the cognitive relay link is taken into consideration. Simulation results show that the effective capacity of the cognitive relay link varies with the statistical QoS constraints. In particular, the stringent QoS constraint will cause low effective capacity. Moreover, we observe that the average total transmit power and average interference power are two important parameters, which will obviously impact the performance of the cognitive relay link. In addition, we find that the transmission of the primary link will significantly affect the performance of the cognitive relay link, such that a larger transmit power of the primary link will cause the performance degradation of the cognitive relay link. Copyright © 2011 John Wiley & Sons, Ltd. Yichen Wang 0002, Pinyi Ren, Fan Li 0003, Zhou Su 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2011 | ROI-based error resilient coding of H.264 for conversational video communicationabstractTransmission of compressed video over packet lossy networks suffers from the problem of error propagation which deteriorate the reconstruction quality of a number of successive frames. For conversational video communication, this problem becomes more severe due to the localization and predominance of the attention-attracted areas, where the motion compensated prediction errors have more energy compared to that of the relatively invariable background. In this paper, we propose an error resilient coding scheme for conversational video communication. The proposed scheme combines the region-of-interest (ROI) segmentation and the long-term reference (LTR) frame coding. The scheme can be easily integrated into the H.264/AVC encoder, and efficiently improve the video quality of the ROI areas. Simulation results verify the robustness and coding efficiency as well as improved subjective visual quality. Fan Li 0003 |
IWCMC | 1 |
| 2011 | Utility Max-Min Fair Rate Allocation for Multiuser Multimedia Communications
Guizhong Liu, Fan Li 0003 |
MMM (1) | 3 |
| 2009 | Application-driven cross-layer design of multiuser H.264 video transmission over wireless networksabstractAn application-driven cross-layer design of multiuser H.264/AVC video transmission over wireless networks is proposed in this paper. An objective function for cross-layer optimization is developed based on the parameters abstracted from the application layer, the media access control layer and the physical layer. Our objective is to maximize the video perceptual quality after delivery with constraint of the limited wireless resources. Simulation results show that the proposed scheme performs significantly than the conventional scheduling schemes for video transmission. Fan Li 0003, Guizhong Liu, Lijun He 0001 |
IWCMC | 1 |
| 2009 | Compressed-Domain-Based Transmission Distortion Modeling for Precoded H.264/AVC VideoabstractTransmission distortion analysis for video streams is a considerably challenging task. In this letter, a compressed-domain-based (CDB) transmission distortion model for precoded H.264/advanced video coding video streams is developed. Unlike the earlier schemes, which were based on pixel domain and required a complete decoding of the compressed video streams, the CDB model only requires some information on the video features, which can be directly extracted from the compressed video streams. Therefore, the complexity of the calculations is substantially reduced, which is well suited for real-time applications. More specifically, the model is applicable to the real-time transmission for precoded video streams, such as video on demand and mobile video. The experimental results demonstrate high accuracy of the model. Furthermore, an application example using the CDB model in resource allocation in real-time multiuser video communication reveals the applicability and effectiveness of the model. Fan Li 0003, Guizhong Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | A Novel Marker System for Real-Time H.264 Video Delivery Over Diffserv NetworksabstractPacket marking plays a key role in video transmission over DiffServ Networks. Legacy marker systems mark video packets mainly based on the coding types of corresponding frames in order to achieve priority transmission. In this paper, we propose a novel packet marker system for realtime H.264 video delivery. Our system marks packets depending on the corresponding frames' contribution to the video quality. The frame relative importance index (FRII) is defined to estimate each frame's contribution, using the degradation of video quality caused by the loss of information in that frame. Consequently, packets of each frame are differently protected according to the FRII in delivery and the video perceptual quality is improved. Simulation results show that our system outperforms the legacy marker systems in terms of the quality of the delivered H.264 video streams. Fan Li 0003, Guizhong Liu |
ICME | 1 |