Li Zhuo 0001

dblp:12/5504 · DBLP profile ↗
← Back
122ranked-venue papers
10as first author
65since 2021 · last 2026
0000-0002-9937-2669ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 38 · 7 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 17 since 2021Computer networks · 6 · 3 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal driver behavior recognition based on frame-adaptive convolution and feature fusion
Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001
Comput. Vis. Image Underst.5
2026 CC-mamba: Mamba-based color constancy with illumination prior-guided dynamic feature modulation and wavelet-domain attention mechanism
Li Zhuo 0001, Hui Zhang 0049, Haokui Xu
Neurocomputing2
2026 TEFormer: Texture-Aware and Edge-Guided Transformer for Semantic Segmentation of Urban Remote Sensing Images
abstract
Accurate semantic segmentation of urban remote sensing images (URSIs) is essential for urban planning and environmental monitoring. However, it remains challenging due to the subtle texture differences and similar spatial structures among geospatial objects, which cause semantic ambiguity and misclassification. Additional complexities arise from irregular object shapes, blurred boundaries, and overlapping spatial distributions of objects, resulting in diverse and intricate edge morphologies. To address these issues, we propose TEFormer, a texture-aware and edge-guided Transformer. Our model features a texture-aware module (TaM) in the encoder to capture fine-grained texture distinctions between visually similar categories, thereby enhancing semantic discrimination. The decoder incorporates an edge-guided tri-branch decoder (Eg3Head) to preserve local edges and details while maintaining multiscale context-awareness. Finally, an edge-guided feature fusion module (EgFFM) effectively integrates contextual, detail, and edge information to achieve refined semantic segmentation. Extensive evaluation demonstrates that TEFormer yields mIoU scores of 88.57% on Potsdam and 81.46% on Vaihingen, exceeding the next best methods by 0.73% and 0.22%. On the LoveDA dataset, it secures the second position with an overall mIoU of 53.55%, trailing the optimal performance by a narrow margin of 0.19%.
Guoyu Zhou, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
IEEE Geosci. Remote. Sens. Lett.5
2026 DualBranchEdgeNet: a dual branch polyp segmentation algorithm based on LoG edge enhancement and quadtree attention transformer
Zhenyu Ni, Suyu Wang, Li Zhuo 0001
Pattern Anal. Appl.3
2026 Asymmetric simulation-enhanced flow reconstruction for incomplete multimodal learning
Jiacheng Yao, Jing Zhang 0023, Li Zhuo 0001
Pattern Recognit.4
2026 CWEFS: Brain Volume Conduction Effects Inspired Channel-Wise EEG Feature Selection for Multi-Dimensional Emotion Recognition
abstract
Due to the intracranial volume conduction effects, high-dimensional multi-channel electroencephalography (EEG) features often contain substantial redundant and irrelevant information. This issue not only hinders the extraction of discriminative emotional representations but also compromises real-time performance. Feature selection has been established as an effective approach to address the challenges while enhancing the transparency and interpretability of emotion recognition models. However, existing EEG feature selection research overlooks the influence of latent EEG feature structures on emotional label correlations and assumes uniform importance across various channels, directly limiting the precise construction of EEG feature selection models for multi-dimensional affective computing. To overcome these issues, this paper proposes a channel-wise EEG feature selection (CWEFS) method for multi-dimensional emotion recognition. Inspired by the volume conduction effects, CWEFS models feature selection within a shared latent structure that captures a consensus representation across EEG channels. This consensus space is jointly learned with a latent semantic analysis of emotional labels to preserve local geometric structure. Furthermore, CWEFS incorporates adaptive channel-weight learning to automatically assess the contribution of each channel. Comprehensive experimental results, compared against nineteen popular feature selection methods, demonstrate that the EEG feature subsets chosen by CWEFS achieve optimal emotion recognition performance across six evaluation metrics.
Xueyuan Xu, Wenjia Dong, Zhijian Gong, Fulin Wei, Li Zhuo 0001, Xia Wu 0001
IEEE Trans. Affect. Comput.6
2026 TSHR-Net: Text Semantic Homogenization Recognition Network for Short Video Title Overlays
abstract
Short videos follow the trend of creation, leading to a proliferation of homogenized video content. Textual overlays such as titles in short videos often reflect semantic homogeneity. This phenomenon manifests not only in the syntactic structure and overall thematic expression, but also in local semantic elements such as words and phrases. Based on information retrieval techniques, we propose a text semantic homogenization recognition network (TSHR-Net) for short video title overlays. The framework comprises key components: (1) a dynamic semantic representation that incorporates contextual information of title overlays using RoBERTa pre-trained word embeddings; (2) a dual-path semantic parser that integrates global semantics via BiLSTM-Attention and local semantics via multi-scale TextCNN; (3) a ranking loss optimization is designed to measure cosine similarity between semantic features, thereby improving homogenization recognition accuracy. Experimental results show that our TSHR-Net achieves the competitive performance in Chinese text semantic homogenization recognition, with ρ and ρ X,Y reaching 80.16% and 78.23% on LCQMC, 81.05% and 80.19% on STS-B(ZH), and 92.47% and 92.19% on our self-built BJUT-HCD. The model also exhibits generalization ability in English, attaining 79.68% and 79.89% on STS-B(EN), and 73.56% and 72.21% on SICK dataset, respectively.
Jing Zhang 0023, Shuying Zhang, Li Zhuo 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2026 LRGFormer: A Multiscale Feature Fusion Transformer for Image Restoration
abstract
Adverse weather conditions can significantly degrade image quality and impair the capture of critical information. Existing restoration networks struggle to effectively combine local, regional, and global features, thereby limiting their ability to handle diverse impacts of such weather. This study proposes the local-region-global transformer (LRGFormer), a transformer-based image restoration model for multiscale feature perception. The model comprises a basic module composed of multi-scale fusion attention (MSFSA) and a channel-spatial dual-attention feed-forward network (CSDF). Specifically, this study designs an MSFSA module. For the first time, it combines rotation-equivariant convolution with local attention for local information extraction and introduces a frequency-domain adaptive attention mechanism. By incorporating a query-aware global adaptive sparse attention mechanism for global information extraction, the network gradually fuses along the channel dimension, enabling progressive capture of spatial and frequency-domain information from the local and regional to global scale. Secondly, a CSDF network structure was designed to enhance channel-spatial interaction and improve the representational capacity of the model. By constructing a basic U-Net framework, the excellent basic modules for image restoration proposed in recent years are compared on a unified framework. Experimental results demonstrated that the proposed basic module can not only better extracts multi-scale features of images and restores image distortion caused by various degradation factors, and also exhibits good universality and generalization.
Jiafeng Li 0001, Wanying Hu, Tongyao Jia, Jiaqi Jin, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 BEV-CMHF: A Cross-Modality Hybrid Fusion Framework for BEV 3D Object Detection With Feature Interaction and Temporal Fusion
abstract
Autonomous driving technology has garnered significant attention for its potential to reduce driver burden and enhance road safety. Modern autonomous driving systems rely on a variety of sensors to perceive complex driving environments. Many existing methods map heterogeneous data into the bird’s eye view (BEV) space for feature fusion. However, they often fail to fully exploit the cross-modal interactions between cameras and LiDAR, or incorporate temporal information, resulting in suboptimal performance. Furthermore, commonly used fusion strategies are often overly simplistic. This study proposes BEV-CMHF, a cross-modality hybrid fusion framework for BEV 3D object detection with feature interaction and temporal fusion. By introducing an interactive cross-attention module and a long-short-term temporal module, the proposed framework enhances the representational power of fused BEV features. Specifically, a feature-interaction attention module that facilitates effective interaction between the camera and LiDAR BEV features using deformable attention is designed, providing guidance and supervision for the camera BEV features. Subsequently, a historical feature temporal fusion module that integrates the long-short-term temporal module is introduced to incorporate additional critical temporal information into the BEV features. Moreover, a dynamic hybrid feature-fusion module is designed to fuse the BEV features of the camera and LiDAR effectively through a hybrid attention mechanism that combines coarse and fine attention. Extensive experiments conducted on the nuScenes benchmark validate the effectiveness of the proposed method, achieving 70.87% mAP and 74.00% NDS on the test set. Using a single NVIDIA GeForce RTX 4090, the method attained an inference speed of 5.79 images per second (5.79 img/s), corresponding to an inference time of 172.64 ms on the nuScenes dataset. The source code will be released athttps://github.com/BJUTsipl/BEV-CMHF
Jiafeng Li 0001, Jinquan Xu, Mengxun Zhi, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Intell. Transp. Syst.5
2026 HDMDN: Hierarchical-Decoupling Based Meta-Knowledge Single Image Dehazing Network
abstract
Images suffer from color shift and detail distortion owing to limitations of contrast and visibility in hazy scenes, affecting their subjective perception. However, the performance of existing algorithms on real-world hazy images remains limited as scenes can be complex and haze degradation varies in outdoor visual systems. This study proposes a meta-knowledge single image dehazing algorithm based on hierarchical decoupling, combining the advantages of convolutional neural networks (CNNs) and transformers. We propose a novel dual-branch decoupling network that decouples low-level features from high level semantic information in images, leveraging the hierarchical properties of the network. It combines a CNN and cross dual branch transformer network (CDual transformer) in the encoder network to fully extract local and global features of images. To disentangle high-level semantic features, a style transfer module is designed to transform the style of hazy images while retaining the remaining semantic information. Afterward, the low-level features of images and the transformed high-level semantic features are used to reconstruct the dehazed images, fully capitalizing on the multilevel features. Furthermore, we built a meta-semi-supervised training strategy to improve the decoupling performance of the model and accumulated style knowledge of clear images from both synthetic and real-world hazy data, improving the generalizability of the model. Extensive experiments on both synthetic and real datasets show that the proposed algorithm effectively removes haze and offers better generalization abilities than similar methods. The project is publicly available at https://github.com/BJUTsipl/HDMDN.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu
IEEE Trans. Multim.3
2026 RcFormer: Reconfigurable Self-Attention Transformer for Image Restoration
abstract
Adverse weather and imaging environments may degrade image quality and pose a significant challenge to the visual perception systems of multimedia. Various image restoration tasks necessitate the modeling of multiscale features, which is highly demanding on networks. To date, Vision Transformer has exhibited impressive image restoration performance. However, in this model, global self-attention is computationally expensive and local self-attention typically limits the interaction domain of each token. To solve this problem, we propose a novel reconfigurable self-attention transformer called RcFormer, which is designed to adequately model multiscale image features. This is achieved through a cross-grouped transformer (CGTransformer) block that uses convolution, area self-attention, and row-column self-attention for different head groups. CGTransformer is combined with an intragroup operation interaction structure. Moreover, an intergroup reconfigurable mechanism is implemented based on CGTransformer and channel circulation. The combination of multiple operations effectively enhances the modeling capability in the spatial and channel dimensions for various image recovery tasks. The performance of the proposed RcFormer is compared with low-level vision modules in a unified framework. Extensive experiments demonstrated that RcFormer exhibited a superior performance for the following image restoration tasks: image dehazing, rain streak removal, raindrop removal, snow removal, and single image deblurring. The source code is publicly available athttps://github.com/dehazing/RcFormer.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu
IEEE Trans. Multim.3
2026 HMS2Net: Heterogeneous Multimodal State Space Network via CLIP for Dynamic Scene Classification in Livestreaming
abstract
Livestreaming platforms attract countless daily active users, making online content regulation imperative. The complex and diverse multimodal content elements in dynamic livestreaming scene pose a great challenge to video content understanding. Thanks to the success of contrastive language-image pre-training (CLIP) for dynamic scene classification, which is one of the basic tasks of video content understanding. We propose a heterogeneous multimodal state space network (HMS2Net) for dynamic scene classification in livestreaming via CLIP. (1) To fully and efficiently mine the dynamic scene elements in livestreaming, we design a heterogeneous teacher-student Transformer (HT-SFormer) with CLIP to extract multimodal features in an energy-efficient unified pipeline; (2) To cope with the possible information conflicts in heterogeneous feature fusion, we introduce a cross-modal adaptive feature filter and fusion (CMAF) module to generate more complete information complementarity by adjusting multimodal feature composition; (3) For temporal context-awareness of dynamic scene, we establish a dynamic state space memory (DSSM) structure for capturing the correlation of multimodal data between neighboring video frames. A series of comparative experiments are conducted on the publicly available datasets DAVIS, Mini-kinetics, HMDB51, and the self-built BJUT-LCD. Our HMS2Net produce competitive results of 71.09%, 95.40%, 53.64%, and 82.36%, respectively, demonstrating the effectiveness and superiority of dynamic scene classification in livestreaming.
Jing Zhang 0023, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Multim.3
2026 On the Adversarial Robustness of Learning-Based Image Compression Against Rate-Distortion Attacks
abstract
Despite demonstrating superior Rate-Distortion (RD) performance, Learning-based Image Compression (LIC) algorithms have been found to be vulnerable to malicious perturbations in recent studies. However, the adversarial attacks considered in existing literature remain divergent from real-world scenarios, both in terms of the attack direction and bitrate. Additionally, existing methods focus solely on empirical observations of the model vulnerability, neglecting to identify the origin of it. These limitations hinder the comprehensive investigation and in-depth understanding of the adversarial robustness of LIC algorithms. To address the aforementioned issues, this paper considers the arbitrary nature of the attack direction and the uncontrollable compression ratio faced by adversaries, and presents two practical rate-distortion attack paradigms,i.e., Specific-ratio Rate-Distortion Attack (SRDA) and Agnostic-ratio Rate-Distortion Attack (ARDA). To the best of our knowledge, we are the first to conduct joint rate-distortion attacks on LIC algorithms. Using the performance variations as indicators, we evaluate the adversarial robustness of eight predominant LIC algorithms against diverse attacks. Furthermore, we propose two novel analytical tools for in-depth analysis,i.e., Entropy Causal Intervention and Layer-wise Distance Magnify Ratio, and reveal thathyperpriorsignificantly increases the bitrate andInverse Generalized Divisive Normalization (IGDN)significantly amplifies input perturbations when under attack. Lastly, we examine the efficacy of adversarial training and introduce the use of online updating for defense. By comparing their advantages and disadvantages, we provide a reference for constructing more robust LIC algorithms against the rate-distortion attacks.
Qingbo Wu 0001, Lei Wang 0186, Fanman Meng, King Ngi Ngan, Li Zhuo 0001, Hongliang Li 0001
IEEE Trans. Multim.7
2026 ToGCN-LLM: Tri-optimization Graph Convolutional Network with LLM for Group Activity Recognition
abstract
Graph networks face substantial challenges in handing large-scale graph-structured data. Although graph convolutional networks (GCNs) have been widely applied to group activity recognition, they still struggle with high graph structure complexity and semantic gap, especially in complex scenarios. To address these limitations, we harness large language models (LLMs) and multimodal learning to enhance semantic understanding as well as bridge the gap between graph structures and contextual meanings. For the simultaneous optimization of model efficiency and recognition accuracy in group activities, we propose a tri-optimization graph convolutional network with LLM (ToGCN-LLM) from the perspective of graph structure knowledge distillation. (1) To tackle high complexity of teacher network in GCN, we adopt a sparsification strategy to prune irrelevant edges and nodes, reducing computational overhead and enhancing training efficiency. (2) To mitigate information redundancy during knowledge distillation, we design a hierarchical optimization module combined with a hierarchical sampling mechanism, exploiting graph hierarchical structure and adjacency relationships to improve knowledge transfer efficiency. (3) Considering the student network’s varying learning performance across training stages and the limitations of fixed learning strategies, we introduce a dynamic adaptive weight decay mechanism to achieve fine-grained convergence under different gradient updates, thereby boosting overall recognition accuracy. (4) We use the Qwen LLM to extract text description tokens, which are fused with the student GCN’s last-layer features to enable multi-model, multi-optimization learning for group activity recognition. Six experiments on CAD, CAED, and BJUT-GAD dataset demonstrate that our ToGCN-LLM achieves competitive MPCA scores of 94.89%, 93.61%, and 95.77%, respectively.
Junpeng Kang, Jing Zhang 0023, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2026 KdM-Net: Knowledge-Driven Memory Network for Long-Tailed Human-Object Interaction in Livestreaming
abstract
With the continuous advancement of the livestreaming industry, streamers as content producers pose significant challenges to the timeliness of regulatory response mechanisms, emerging as a critical weak link in cyberspace governance. Human-object interaction (HOI) detection plays a pivotal role in understanding multimodal livestreaming videos. In mainstream healthy online ecosystems, normal HOI categories dominate, while rare ones are extremely scarce, i.e., long-tail distribution that hinders HOI models from effectively detecting streamer violations. Driven by the transformative potential of foundation models (FMs) in multimodal video understanding, we propose a knowledge-driven memory network (KdM-Net) for long-tailed HOI in livestreaming, leveraging the extensive capabilities of the contrastive language-image pretraining (CLIP) model. First, human-object (HO) pairs are generated and modeled using general object detector/tracker. After converting each HOI label into a short sentence description, text embeddings are extracted via the CLIP text encoder to initialize classifier weights. Notably, we introduce a visual-textual knowledge transfer strategy to align visual and text features, complementing for the multimodal knowledge deficit of rare categories that plague long-tailed HOI distributions. Finally, a knowledge-driven memory module is designed to dynamically assign adaptive weights and attention to interaction categories based on their long-tailed distribution characteristics, mitigating model forgetting tail data and enhancing HOI detection performance. Experimental results demonstrate that KdM-Net achieves HOI detection accuracies of 37.33@full, 50.63%@non-rare, and 27.14%@rare on the publicly available VidHOI dataset, and 45.41%@full, 61.77%@non-rare, and 31.95%@rare on the self-built BJUT-HOI dataset. These fundings validate the generalization power of our KdM-Net for HOI detection and its competitiveness in livestreaming scenarios.
Menghui Zhang, Jing Zhang 0023, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 From Echo Distillation to Chat-Driven: A Multi-Task VLM is OK for Object Detection and Fine-Grained Classification in Remote Sensing Images
abstract
Vision-language models (VLMs) have demonstrated attracting performance in natural image domains, and their application to multitask remote sensing analysis and interpretation holds great potential. However, realizing this potential in remote sensing images (RSIs) is hindered by limited feature extraction in lightweight detectors, loss of fine-grained details in classification, and the lack of unified multi-task frameworks. In response, we propose a multi-task vision-language model (MT-VLM) to address these limitations, offering an effective solution for both object detection and fine-grained classification in RSIs. Specifically, we introduce EchoDistill, an echo distillation framework that combines teacher knowledge transfer and self-knowledge reinforcement to iteratively enhance feature learning, significantly boosting the detection performance of the lightweight models. Additionally, we develop a chat-driven VLM with LoRA fine-tuning, enabling the model to capture fine-grained details and improve classification accuracy using multimodal embeddings. To support our framework, we collected BJUT-FGC, a comprehensive image-text labeled dataset for fine-grained classification in RSIs, and built a dual-mode multi-task benchmark to facilitate both task-specific adaptation and integrated multi-task refinement. Experimental results on DOTA 2.0, VEDAI, and self-built BJUT-FGC datasets demonstrate the effectiveness of our approach, achieving 58.71% and 82.06% mAP on the object detection task and 88.08% top-1 accuracy on the fine-grained classification task.
Liuqian Wang, Jing Zhang 0023, Guangming Mi, Li Zhuo 0001
IJCNN4
2025 Prototype Embedding Optimization for Human-Object Interaction Detection in Livestreaming : PeO-HOI
abstract
Livestreaming frequently features interactions between streamers and objects, which is critical for understanding and regulating online content. While human-object interaction (HOI) detection has advanced significantly for general video tasks, its application to recognizing streamer-object interactions in livestreaming often exhibits object bias: an excessive focuses on the objects at the expense of interactions with the streamer. To address this challenge, we propose a prototype embedding optimization for human-object interaction detection (PeO-HOI). Our approach first preprocesses the livestreaming using object detection and tracking to extract features of human-object (HO) pairs. Subsequently, prototype embedding optimization is applied to mitigate object bias effects. Finally, after modeling the spatio-temporal context among HO pairs, the HOI detection results are generated by the prediction head. Experimental results demonstrate that the PeO-HOI achieves detection accuracies of 37.19%@full, 51.42%@non-rare, and 26.20%@rare on the publicly available VidHOI dataset, and 45.13%@full, 62.78%@non-rare, and 30.37%@rare on our self-built BJUT-HOI dataset. These results confirm that PeO-HOI effectively enhances HOI detection performance in livestreaming scenarios.
Menghui Zhang, Jing Zhang 0023, Li Zhuo 0001
MMSP4
2025 RWGCN: Random walk graph convolutional network for group activity recognition
Junpeng Kang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Appl. Intell.5
2025 Explainable graph convolutional network based on catastrophe theory and its application to group activity recognition
Junpeng Kang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Eng. Appl. Artif. Intell.5
2025 DR-YOLO: dual reconstructed YOLO for logo detection in livestreaming
Chenyu Yuan, Jing Zhang 0023, Li Zhuo 0001
Multim. Syst.4
2025 Embedded multi-label feature selection via orthogonal regression
Xueyuan Xu, Fulin Wei, Tianze Yu, Jinxin Lu, Aomei Liu, Li Zhuo 0001, Feiping Nie 0001, Xia Wu 0001
Pattern Recognit.6
2025 Cellular spatial-semantic embedding for multi-label classification of cell clusters in thyroid fine needle aspiration biopsy whole slide images
Juntao Gao, Jing Zhang 0023, Li Zhuo 0001
Pattern Recognit. Lett.4
2025 Position Guided Dynamic Receptive Field Network: A Small Object Detection Friendly to Optical and SAR Images
abstract
Object detection in remote sensing images (RSIs), including optical and SAR images, has emerged as a rapidly advancing field. However, the abundance of small objects in RSIs poses a significant challenge in designing a network structure with effective receptive fields to support accurate localization and classification. In this paper, we propose a position guided dynamic receptive field network (PG-DRFNet) for small object detection friendly to optical and SAR images. Specifically, PG-DRFNet overcomes the problem of small objects vanishing or being submerged in features by establishing a positional guidance relationship of small objects between different feature layers. Then, we design a combination head structure that utilizes additional supervised information extracted from small objects to make the model more effective and flexible. Moreover, a dynamic perception algorithm based on feature construction is developed to dynamically optimize the perception regions and feature hierarchies of the model, while seeking the optimal tradeoff between model accuracy and inference speed. Without bells and whistles, our model is robust to two modalities of remote sensing data, and our experiments are conducted on four benchmark RSI datasets, including DOTA-v2.0, VEDAI, SSDD, and HRSID. The experimental results achieve competitive performance with 59.01%, 84.06%, 90.06%, and 80.59% mAP, respectively. Code and models are released athttps://github.com/BJUT-AIVBD/PG-DRFNet.
Liuqian Wang, Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Hybrid-MambaCD: Hybrid Mamba-CNN Network for Remote Sensing Image Change Detection With Region-Channel Attention Mechanism and Iterative Global-Local Feature Fusion
abstract
Mamba has gained significant attention for its outstanding long-range context modeling capability while maintaining linear complexity, compared with Transformer. In this article, a hybrid Mamba and convolutional neural network (CNN) architecture is proposed for remote sensing image change detection (RSICD), named Hybrid-MambaCD, which leverages the advantages of CNN for local detail information extraction and Mamba for global context information extraction, providing an efficient solution for RSICD tasks. First, the region-channel attention mechanism (RCAM) is designed to enhance the CNN features from both channel and region dimensions, enabling the network to focus more on change regions while suppressing interference from background areas. Second, an iterative global-local feature fusion (IGLFF) strategy is proposed, which performs an adaptive weighted fusion of global and local features across multiple scales in a progressive manner, enhancing the representation ability of the features. Experimental results on three public datasets of LEVIR-CD, WHU-CD, and DSIFN-CD show that compared to the existing RSICD methods, the proposed Hybrid-MambaCD achieves the state-of-the-art (SOTA) detection performance.
Li Zhuo 0001, Hui Zhang 0049, Jiafeng Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 TSTrack: A Lightweight Transformer-Based Spatiotemporal Feature Refinement Tracking Algorithm
abstract
Single-object tracking is a fundamental enabling technology in the field of remote sensing observation. It plays a crucial role in tasks such as unmanned aerial vehicle route surveillance and maritime vessel trajectory prediction. However, because of challenges such as the weak discriminative power of target features, interference from complex environments, and frequent viewpoint changes, existing trackers often suffer from insufficient temporal modeling capabilities and low computational efficiency, which limit their practical deployment. To address these challenges, we propose TSTrack, a novel lightweight single-object tracking framework that integrates Transformer and Mamba-based spatiotemporal modeling. First, we propose the target-aware feature purification preprocessor (TAFPP) , designed to dynamically enhance target representation through a synergistic combination of the dynamic position acuity module (DPAM) and spectral channel recalibrator (SCR). Second, we introduce the recurrent Mamba interaction pyramid (RM-IP) to replace traditional recurrent neural network-based structures, leveraging a state-space model for efficient and expressive temporal modeling with significantly reduced parameter overhead. Finally, we propose the elastic reconstructive multi-scale fusion (ERMSF) module, which adopts a four-branch parallel architecture to achieve effective multiscale feature fusion and dynamic shape adaptation, thereby enhancing robustness against target deformations and scale variations. Extensive experiments conducted on benchmark datasets, including LaSOT, TrackingNet, and GOT-10k, demonstrate the effectiveness of TSTrack. The results show that TSTrack achieves a superior tracking accuracy while maintaining a lightweight design, significantly outperforming existing state-of-the-art methods. The source code is publicly available at https://github.com/BJUTsipl/TSTrack.
Jiafeng Li 0001, Shengyao Sun, Yang Wang 0023, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Accurate Multi-Landmark Localization in 3D Ultra-High Resolution CT Images of the Ears Via Deep Reinforcement Learning and Transformer
abstract
Automated landmark localization can help radiologists quickly determine the locations of key structures or lesion areas from medical images. However, when facing large-volume 3D medical images, existing methods have very high computational complexity due to the need to encode the global image. That is to say, it is difficult for existing methods to achieve accurate landmark localization in 3D medical images at a faster localization speed. In this paper, an accurate multi-landmark localization method for ear 3D Ultra-High Resolution CT (U-HRCT) images is proposed. This method adopts a novel localization pipeline that combines Deep Reinforcement Learning (DRL) and Transformer. Firstly, the DRL algorithm is used to quickly collect landmark-related local features. Secondly, Transformer is used to extract the spatial position relationship between anatomical structures from these discrete local features to infer the coordinate position of the landmark. Because the complex process of encoding the global image is avoided, the proposed method can achieve fast localization of ear multi-landmark in 3D U-HRCT images. Finally, we proposed a refinement module based on dual-branch hybrid Multi-Layer Perceptron, which can use the fast localization results of multi-landmark to learn the spatial position relationship between landmarks, thereby further improving the accuracy and stability of landmark localization. Experimental results on the self-built ear 3D U-HRCT dataset and the publicly available 2D cephalometric dataset demonstrate that, the proposed method can achieve Successful Detection Rate of 96.71% and 89.97% respectively within the precision range of 2.0 mm, surpassing the state-of-the-art multi-landmark localization methods and has a faster localization speed.
Zhiwei Qu, Li Zhuo 0001, Hongxia Yin, Zhenchang Wang
IEEE J. Biomed. Health Informatics2
2025 CycFormer: Unsupervised Rain Removal Network Based on CycleGAN and Transformer
abstract
Rainy weather presents significant challenges for applications relying on visual perception in intelligent transportation systems. The scarcity of real paired training data complicates single-image rain removal tasks, prompting an increasing interest in unsupervised methods capable of handling real-world rainy images without paired data. At present, most unsupervised rain removal methods are based on the CycleGAN framework; however, the combination of this framework and transformer is not satisfactory owing to most Transformers’ insufficient ability to model real rain features with global inhomogeneous distributions, which prevents them from being fully applicable to unsupervised tasks. This study devised an unsupervised rain removal network based on CycleGAN and the DerainFormer transformer. First, a deformable sparse attention mechanism was developed to improve the Transformer’s suitability for unsupervised tasks in CycleGAN architectures. Subsequently, a two-stage alternating transformer structure was designed to enhance its global non-uniform modeling capabilities for real rain images, In addition, a dual-channel parallel feed-forward network was used to establish the correlation between multiscale rain stripes. Finally, since rain removal is considered a decomposition task, a rain layer unsupervised training method for joint positional contrastive learning was proposed to separate the rain streaks effectively. We conducted several experiments on different real and synthetic rain datasets and the results confirmed that our unsupervised rain removal method performed well. The source code will be released athttps://github.com/derainsipl/CycFormer.
Jiafeng Li 0001, Shuhao Yan, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Intell. Transp. Syst.5
2025 V2V Cooperative Perception With Adaptive Communication Loss for Autonomous Driving
Jingyue Shi, Junhui Zhao 0001, Li Zhuo 0001, Xiaoming Wang 0011, Xiaohuang Zhan
IEEE Trans. Intell. Transp. Syst.3
2025 Cross-Modal Tri-Semantic Correlation-CLIP for Short Video Homogenization Recognition
abstract
Short videos are one of the most popular social media in the world, triggering a proliferation of copycat creations leading to homogenized video content, with visual and textual homogenization being the most prevalent. Unlike near-duplicate video retrieval, which relies on visual appearance similarity, homogenization recognition emphasizes identifying videos with similar semantic units. Short videos exhibit multimodal features, in which there is a many-to-many mapping relationship between visual and text elements, and the two modalities are relatively independent and semantically correlated. Therefore, cross-modal semantic correlation needs to be explored and established to achieve homogenization recognition of short videos. Based on the idea of divide-and-conquer and joint processing, we propose a cross-modal tri-semantic correlation-CLIP (CS 3 C-CLIP) for short video homogenization recognition. First, visual and text features in the shared subspace are extracted using the contrastive language-image pre-training visual-text dual encoder. Then, features at the patch, frame, and video levels are generated using the patch selection module and the temporal encoder, while the word-level and sentence-level features are respectively derived from text features and [EOS] token. After establishing cross-modal tri-semantic correlations by constructing a triple semantic (i.e., video-sentence, frame-sentence, and patch-word) correlation, homogenized short videos are recognized by measuring the aggregated cross-modal similarity between pairs of short videos. Experimental results on three publicly available datasets demonstrate that our CS 3 C-CLIP outperforms state-of-the-art methods, achieving 85.7% R@1 and 94.4% R@5 on self-built BJUT-HCD, 49.4% R@1 and 74.6% R@5 on MSR-VTT, and 49.8% R@1 and 78.1% R@5 on MSVD, respectively.
Jiacheng Yao, Jing Zhang 0023, Shuying Zhang, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Unpaved road segmentation of UAV imagery via a global vision transformer with dilated cross window self-attention for dynamic map
Jing Zhang 0023, Jiafeng Li 0001, Li Zhuo 0001
Vis. Comput.4
2024 A Coarse to Fine Detection Method for Prohibited Object in X-ray Images Based on Progressive Transformer Decoder
Chunjie Ma, Lina Du, Zan Gao 0001, Li Zhuo 0001, Meng Wang 0001
ACM Multimedia4
2024 WSEL: EEG Feature Selection with Weighted Self-expression Learning for Incomplete Multi-dimensional Emotion Recognition
abstract
Due to the small size of valid samples, multi-source EEG features with high dimensionality can easily cause problems such as overfitting and poor real-time performance of the emotion recognition classifier. Feature selection has been demonstrated as an effective means to solve these problems. Current EEG feature selection research assumes that all dimensions of emotional labels are complete. However, owing to the open acquisition environment, subjective variability, and border ambiguity of individual perceptions of emotion, the training data in the practical application often includes missing information, i.e., multi-dimensional emotional labels of several instances are incomplete. The aforementioned incomplete information directly restricts the accurate construction of the EEG feature selection model for multi-dimensional emotion recognition. To wrestle with the aforementioned problem, we propose a novel EEG feature selection model with weighted self-expression learning (WSEL). The model utilizes self-representation learning and least squares regression to reconstruct the label space through the second-order correlation and higher-order correlation within the multi-dimensional emotional labels and simultaneously realize the EEG feature subset selection under the incomplete information. We have utilized two multimedia-induced emotion datasets with EEG recordings, DREAMER and DEAP, to confirm the effectiveness of WSEL in the missing multi-dimensional emotional feature selection challenge. Compared to nine state-of-the-art feature selection approaches, the experimental results demonstrate that the EEG feature subsets chosen by WSEL can achieve optimal performance in terms of six performance metrics.
Xueyuan Xu, Li Zhuo 0001, Jinxin Lu, Xia Wu 0001
ACM Multimedia2
2024 MKP-Net: Memory knowledge propagation network for point-supervised temporal action localization in livestreaming
Jing Zhang 0023, Yian Zhang, Junpeng Kang, Li Zhuo 0001
Comput. Vis. Image Underst.5
2024 LCMA-Net: A light cross-modal attention network for streamer re-identification in live video
Jiacheng Yao, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Comput. Vis. Image Underst.4
2024 BEV perception for autonomous driving: State of the art and future perspectives
Junhui Zhao 0001, Jingyue Shi, Li Zhuo 0001
Expert Syst. Appl.3
2024 Content-Adaptive Residual Learning and Context-Aware Entropy Model for SAR Image Compression
abstract
Synthetic aperture radar (SAR) images are pivotal in remote sensing applications. However, due to the physical characteristics of coherent imaging, current SAR image compression methods are susceptible to speckle noise, leading to distortion and higher compression rates. To address these issues, we propose a content-adaptive transformation network that dynamically adjusts the receptive field size based on image content, thereby mitigating noise impact and capturing detailed features more effectively. In addition, we developed a context-aware entropy model (CAEM) to better explore channel correlations within the latent feature space, which helps to reduce redundancy in latent features. Experimental results demonstrate that our method achieves state-of-the-art performance compared to traditional image compression standards and deep learning-based models, significantly enhancing the compression ratio and image reconstruction quality for the Sandia and ICEYE datasets.
Shaoman Fu, Hui Zhang 0049, Haoxuan Feng, Li Zhuo 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 HDUD-Net: heterogeneous decoupling unsupervised dehaze network
Jiafeng Li 0001, Lingyan Kuang, Jiaqi Jin, Li Zhuo 0001, Jing Zhang 0023
Neural Comput. Appl.4
2024 RaSTFormer: region-aware spatiotemporal transformer for visual homogenization recognition in short videos
Shuying Zhang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Neural Comput. Appl.4
2024 Self-guided disentangled representation learning for single image dehazing
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023
Neural Networks3
2024 MPLA-Net: Multiple Pseudo Label Aggregation Network for Weakly Supervised Video Salient Object Detection
abstract
Weakly Supervised Video Salient Object Detection (WSVSOD) only requires coarse-grained manual annotations, which can achieve a good trade-off between labeling efficiency and detection performance. In this paper, a Multiple Pseudo Label Aggregation Network (MPLA-Net) is proposed for WSVSOD. Firstly, the video frames that can obtain high-quality pseudo labels are selected to generate multiple pseudo labels, so as to avoid the prejudice of the single label. Moreover, the pseudo label with fine edge information is used to generate the Edge Information Map (EIM). Secondly, MPLA-Net is designed to adequately excavate and utilize the comprehensive saliency cues in multiple pseudo labels to improve the detection accuracy, in which ResNet-50 is adopted as the backbone network. Edge loss, pseudo label loss, self-supervised loss and fusion loss are exploited to jointly supervise and optimize the network training to obtain a robust detection model. Experimental results on five benchmark datasets demonstrate that, compared with existing weakly supervised methods, the proposed method can achieve state-of-the-art detection accuracy with less model parameters and higher detection speed. And the detected salient objects have fine boundaries.
Chunjie Ma, Lina Du, Li Zhuo 0001, Jiafeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 GCFormer: Global Context-Aware Transformer for Remote Sensing Image Change Detection
abstract
In recent years, Transformer-based Change Detection (CD) in Remote Sensing Images has achieved significant advances, making it an emerging hot research topic. However, the current CD methods suffer from some problems, such as incomplete detection of change regions and missed detection of small change regions. In this paper, a Global Context-aware Transformer is proposed for CD tasks, named GCFormer, to address above issues by efficiently enhancing the global context information. It is fulfilled from two aspects based on the hybrid Convolutional Neural Network (CNN)+Transformer framework. Firstly, a Multi-Receptive Field Conv-Attention (MRFCA) mechanism is designed, which combines dilated convolutions with multiple rates and Conv-Attention, fully leveraging the advantages of convolution operation and self-attention mechanism. It is embedded at the highest layer of CNN to extract multi-receptive-field global context information. Secondly, a Context-aware Relative Position Encoding (CRPE) mode is proposed to replace the Absolute Position Encoding (APE) mode of Transformer. As a result, it can capture long-range dependency more efficiently and further enhance the global context information extraction and representation ability of the network. Experimental results on three public benchmark datasets of LEVIR-CD, WHU-CD and DSIFN-CD show that, the proposed GCFormer achieves superior detection performance with lower model complexity than the state-of-the-art Transformer-based CD methods. The source code is available at: https://github.com/yuwanting828/yuwanting828.github.io.
Wanting Yu, Li Zhuo 0001, Jiafeng Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 WDFF-Net: Weighted Dual-Branch Feature Fusion Network for Polyp Segmentation With Object-Aware Attention Mechanism
abstract
Colon polyps in colonoscopy images exhibit significant differences in color, size, shape, appearance, and location, posing significant challenges to accurate polyp segmentation. In this paper, a Weighted Dual-branch Feature Fusion Network is proposed for Polyp Segmentation, named WDFF-Net, which adopts HarDNet68 as the backbone network. First, a dual-branch feature fusion network architecture is constructed, which includes a shared feature extractor and two feature fusion branches, i.e. Progressive Feature Fusion (PFF) branch and Scale-aware Feature Fusion (SFF) branch. The branches fuse the deep features of multiple layers for different purposes and with different fusion ways. The PFF branch is to address the under-segmentation or over-segmentation problems of flat polyps with low-edge contrast by iteratively fusing the features from low, medium, and high layers. The SFF branch is to tackle the the problem of drastic variations in polyp size and shape, especially the missed segmentation problem for small polyps. These two branches are complementary and play different roles, in improving segmentation accuracy. Second, an Object-aware Attention Mechanism (OAM) is proposed to enhance the features of the target regions and suppress those of the background regions, to interfere with the segmentation performance. Third, a weighted dual-branch the segmentation loss function is specifically designed, which dynamically assigns the weight factors of the loss functions for two branches to optimize their collaborative training. Experimental results on five public colon polyp datasets demonstrate that, the proposed WDFF-Net can achieve a superior segmentation performance with lower model complexity and faster inference speed, while maintaining good generalization ability.
Zhiwei Qu, Li Zhuo 0001, Hui Zhang 0049
IEEE J. Biomed. Health Informatics4
2024 Semi-Supervised Single-Image Dehazing Network via Disentangled Meta-Knowledge
abstract
Captured outdoor scene images are easily affected by haze. Most image dehazing methods have limited generalization capabilities for real-world hazy images owing to the complexities of real-world environments and domain gaps in the training datasets. This article proposes a semi-supervised single-image dehazing network based on disentangled meta-knowledge. The symmetric and heterogeneous design of the disentangled network is conducive to the separation of the content and mask features of hazy images and these features are used as meta-knowledge to guide feature fusion in the dehazing network. Moreover, functions describing constant-color and disentangled-reconstruction-checking losses are designed to ensure the subjective qualities of the generated dehazed images. The results of extensive experiments conducted on synthetic datasets and real-world images indicate that the proposed algorithm outperforms state-of-the-art single-image dehazing algorithms. In addition, the algorithm effectively improves the performance of object-detection tasks.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Tianjian Yu
IEEE Trans. Multim.3
2023 Graph Disentangled Representation Based Semi-supervised Single Image Dehazing Network
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001
ICIC (2)3
2023 Deeply supervised vestibule segmentation network for CT images with global context-aware pyramid feature extraction
abstract
Abstract Accurate vestibule segmentation for CT images is of great significance for the clinical diagnosis of congenital ear malformations and cochlear implant. However, it is still a challenging task due to extremely small size and irregular shape of vestibule. Here, a vestibule segmentation network for CT images is proposed under the basic encoder‐decoder framework. Firstly, a residual block based on channel attention mechanism, named Res‐CA block, is designed to guide the network to enhance the important features for the segmentation tasks while suppressing the irrelevant ones. And then, a global context‐aware pyramid feature extraction (GCPFE) module is proposed to capture multi‐receptive‐field global context information. Finally, active contour with elastic (ACE) loss function is adopted to guide network learning more detailed information of the boundary. Furthermore, deep supervision (DS) mechanism is employed to locate the boundaries finely, improving the robustness of the network. The experiments are conducted on the self‐established VestibuleDataset and UHRCT‐Dataset, as well as publicly available retinal dataset, namely DRIVE, to comprehensively verify the robustness and generalization capability of the proposed segmentation network. The experimental results show that the proposed network can achieve a superior performance.
Meijuan Chen, Li Zhuo 0001, Ziyao Zhu, Hongxia Yin, Zhenchang Wang
IET Image Process.2
2023 Occluded prohibited object detection in X-ray images with global Context-aware Multi-Scale feature Aggregation
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
Neurocomputing2
2023 Short video fingerprint extraction: from audio-visual fingerprint fusion to multi-index hashing
Shuying Zhang, Jing Zhang 0023, Li Zhuo 0001
Multim. Syst.4
2023 Few-Shot Remote Sensing Scene Classification With Spatial Affinity Attention and Class Surrogate-Based Supervised Contrastive Learning
abstract
Learning powerful and discriminative representation is critical for boosting the performance of Few-Shot Remote Sensing Scene Classification (FSRSSC). The remote sensing images have unique characteristics, such as, complex background and co-occurrence of multiple objects, making FSRSSC challenging. To address this problem, in this paper, a novel FSRSSC method is proposed. Firstly, a Spatial Affinity Attention (SAA) mechanism is designed to encourage the network model to focus on critical regions. The SAA infers attention maps from both channel and spatial dimensions, and encodes the mean values and affinities of feature nodes in each channel along the vertical and horizontal directions. Secondly, a Class Surrogate-based Supervised Contrastive Learning (CSSCL) strategy is proposed to promote intra-class compactness and inter-class dispersion. Different from Supervised Contrastive Learning (SCL), the CSSCL learns a surrogate for each class in the contrastive space and select the corresponding class surrogate rather than the samples from the same class for the anchor to form positive pairs. It can alleviate the impact of too hard or too simple positive sample pairs on model generalization that exists when SCL is introduced into FSRSSC. The model is trained under the joint supervision of Cross-Entropy (CE) loss and CSSCL loss on the merged base class data instead of using a meta-learning strategy to learn a robust base feature extractor. Extensive experimentations on three public remote sensing benchmark datasets show that our proposed method can achieve a competitive or state-of-the-art classification performance.
Li Zhuo 0001, Jiafeng Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Efficient Fine-Grained Object Recognition in High-Resolution Remote Sensing Images From Knowledge Distillation to Filter Grafting
abstract
With the development of high-resolution remote sensing images (HR-RSIs) and the escalating demand for intelligent analysis, fine-grained recognition of geospatial objects has become a more practical and challenging task. Although deep learning-based object recognition has achieved superior performance, it is inflexible to be directly utilized to the fine-grained object recognition tasks of HR-RSIs under the limitation of the size of geospatial objects. An efficient fine-grained object recognition method in HR-RSIs from knowledge distillation to filter grafting is proposed. Specifically, fine-grained object recognition consists of two stages: Stage 1 utilizes oriented region convolutional neural network (oriented R-CNN) to accurately locate and preliminarily classify geospatial objects. At the same time, it serves as a teacher network to guide students’ effective learning of fine-grained object recognition; in Stage 2, we design a coarse-to-fine object recognition network (CF-ORNet), as the second teacher network, which realizes fine-grained recognition through feature learning and category correction. After that, we propose a lightweight model from knowledge distillation to filter grafting on two teacher networks to achieve efficient fine-grained object recognition. The experimental results on VEDAI and HRSC2016 datasets achieve competitive performance.
Liuqian Wang, Jing Zhang 0023, Jimiao Tian, Jiafeng Li 0001, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 BARRN: A Blind Image Compression Artifact Reduction Network for Industrial IoT Systems
abstract
Most industrial Internet of Things (IoT) devices reduce the capture image size using high-ratio joint photographic experts group (JPEG) compression, saving storage space, and transmission bandwidth consumption. However, the resulting compression artifacts considerably affect the accuracy of subsequent tasks. Most artifact reduction algorithms do not consider the limitations of storage space and computing power of edge devices. In this study, a blind artifact reduction recurrent network (BARRN), which can reduce compression artifacts when the quality factors are unknown, is proposed. First, a structure based on recurrent convolution is designed for the specific requirements of industrial IoT image acquisition devices; the network can be scaled according to system resource constraints. Second, a more efficient convolution group, capable of adaptively processing different degradation levels, is proposed for optimal use of the limited computational resources. The experimental results demonstrate that the proposed BARRN can meet the needs of industrial systems with high computational efficiency.
Jiafeng Li 0001, Yuqi Gao, Li Zhuo 0001, Jing Zhang 0023
IEEE Trans. Ind. Informatics4
2023 TP-Net: Two-Path Network for Retinal Vessel Segmentation
abstract
Refined and automatic retinal vessel segmentation is crucial for computer-aided early diagnosis of retinopathy. However, existing methods often suffer from mis-segmentation when dealing with thin and low-contrast vessels. In this paper, a two-path retinal vessel segmentation network is proposed, namely TP-Net, which consists of three core parts, i.e., main-path, sub-path, and multi-scale feature aggregation module (MFAM). Main-path is to detect the trunk area of the retinal vessels, and the sub-path to effectively capture edge information of the retinal vessels. The prediction results of the two paths are combined by MFAM, obtaining refined segmentation of retinal vessels. In the main-path, a three-layer lightweight backbone network is elaborately designed according to the characteristics of retinal vessels, and then a global feature selection mechanism (GFSM) is proposed, which can autonomously select features that are more important for the segmentation task from the features at different layers of the network, thereby, enhancing the segmentation capability for low-contrast vessels. In the sub-path, an edge feature extraction method and an edge loss function are proposed, which can enhance the ability of the network to capture edge information and reduce the mis-segmentation of thin vessels. Finally, MFAM is proposed to fuse the prediction results of main-path and sub-path, which can remove background noises while preserving edge details, and thus, obtaining refined segmentation of retinal vessels. The proposed TP-Net has been evaluated on three public retinal vessel datasets, namely DRIVE, STARE, and CHASE DB1. The experimental results show that the TP-Net achieved a superior performance and generalization ability with fewer model parameters compared with the state-of-the-art methods.
Zhiwei Qu, Li Zhuo 0001, Hongxia Yin, Zhenchang Wang
IEEE J. Biomed. Health Informatics2
2023 USID-Net: Unsupervised Single Image Dehazing Network via Disentangled Representations
abstract
Captured images of outdoor scenes usually exhibit low visibility in cases of severe haze, which interferes with optical imaging and degrades image quality. Most of the existing methods solve the single-image dehazing problem by applying supervised training on paired images; however, in practice, the pairing of real-world images is not viable. Additionally, the processing speed of individual dehazing models is important in practical applications. In this study, a novel unsupervised single image dehazing network (USID-Net) based on disentangled representations without paired training images is explored. Furthermore, considering the trade-off between performance and memory storage, a compact multi-scale feature attention (MFA) module is developed, integrating multi-scale feature representation and attention mechanism to facilitate feature representation. To effectively extract haze information, a mechanism referred to as OctEncoder is designed to include multi-frequency representations that can capture more global information. Extensive experiments show that USID-Net achieves competitive dehazing results and a relatively high processing speed compared to state-of-the-art methods. The source code is available athttps://github.com/dehazing/USID-Net.
Jiafeng Li 0001, Yaopeng Li, Li Zhuo 0001, Lingyan Kuang, Tianjian Yu
IEEE Trans. Multim.3
2023 Cascade Transformer Decoder Based Occluded Pedestrian Detection With Dynamic Deformable Convolution and Gaussian Projection Channel Attention Mechanism
abstract
Occluded pedestrian detection is very challenging in computer vision, because the pedestrians are frequently occluded by various obstacles or persons, especially in crowded scenarios. In this article, an occluded pedestrian detection method is proposed under a basic DEtection TRansformer (DETR) framework. Firstly, Dynamic Deformable Convolution (DyDC) and Gaussian Projection Channel Attention (GPCA) mechanism are proposed and embedded into the low layer and high layer of ResNet50 respectively, to improve the representation capability of features. Secondly, Cascade Transformer Decoder (CTD) is proposed, which aims to generate high-score queries, avoiding the influence of low-score queries in the decoder stage, further improving the detection accuracy. The proposed method is verified on three challenging datasets, namely CrowdHuman, WiderPerson, and TJU-DHD-pedestrian. The experimental results show that, compared with the state-of-the-art methods, it can obtain a superior detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IEEE Trans. Multim.2
2022 Prohibited Object Detection in X-ray Images with Dynamic Deformable Convolution and Adaptive IoU
abstract
Due to the variety and complexity of objects in X-ray images, how to detect the prohibited items automatically and accurately is a challenging problem. In this paper, an X-ray image prohibited object detection method based on Dynamic Deformable Convolution (DyDC) and adaptive Intersection over Union (IoU) is proposed based on Cascade R-CNN framework. The main contributions are as follows. First, DyDC is proposed to cope with the diversity of the prohibited objects in X-ray images and to improve the feature representation capability. Then, adaptive IoU mechanism is proposed, which can dynamically adjust the IoU threshold during the training process to generate high quality proposals. The proposed method is extensively evaluated on two publicly available benchmark datasets, namely SIXray and OPIXray, and the experimental results show that it can achieve the state-of-the-art detection accuracy, compared with other existing methods.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
ICIP2
2022 EAOD-Net: Effective anomaly object detection networks for X-ray images
abstract
Abstract Anomaly object detection is the core technology in the application for X‐ray images. However, the accuracy of current X‐ray anomaly object detection method still needs to be improved. In this paper, an effective anomaly object detection network is proposed to improve the detection accuracy of anomaly object for X‐ray images. Firstly, learnable Gabor convolution layer, deformable convolution, and spatial attention mechanism are introduced to enhance the representative ability of features in ResNeXt. Then, dense local regression is applied to predict the offset of multiple dense boxes in region proposal to locate the object accurately. At last, bigger discriminative RoI pooling is proposed to classify the candidate boxes to improve the accuracy of object classification. Experimental results on the SIXray and OPIXray datasets show that compared with the state‐of‐the‐art methods, the proposed EAOD‐Net can achieve the competitive detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IET Image Process.2
2022 VCFNet: video clarity-fluency network for quality of experience evaluation model of HTTP adaptive video streaming services
Lina Du, Jiafeng Li 0001, Li Zhuo 0001
Multim. Tools Appl.3
2022 Meta-Learning Paradigm and CosAttn for Streamer Action Recognition in Live Video
abstract
As an emerging field of network content production, live video has been in the vacuum zone of cyberspace governance for a long time. Streamer action recognition is conducive to the supervision of live video content. In view of the diversity and imbalance of streamer actions, it is attractive to introduce few-shot learning to realize streamer action recognition. Therefore, a meta-learning paradigm and CosAttn for streamer action recognition method in live video is proposed, including: (1) the training set samples similar to the streamer action to be recognized are pretrained to improve the backbone network; (2) video-level features are extracted by R(2+1)D-18 backbone and global average pooling in the meta-learning paradigm; (3) the streamer action is recognized by calculating cosine similarity after sending the video-level features to CosAttn to generate a streamer action category prototype. Experimental results on several real-world action recognition datasets demonstrate the effectiveness of our method.
Jing Zhang 0023, Jiacheng Yao, Li Zhuo 0001, Qi Tian 0001
IEEE Signal Process. Lett.4
2022 Effective Meta-Attention Dehazing Networks for Vision-Based Outdoor Industrial Systems
abstract
Haze seriously affects the reliability of industrial systems, especially vision-based outdoor industrial systems such as autopilot systems. A majority of existing dehazing methods are not specifically designed for industrial systems and do not consider the reliability and resource cost of industrial system implementation. In this article, a novel meta-attention dehazing network (MADN) is proposed for direct restoration of clear images from hazy images without using the physical scattering model. Combined with parallel operation and enhancement modules, the meta-network automatically selects the most suitable dehazing network structure based on the current input hazy image by a meta-attention module. In addition, a novel feature loss calculated by the meta-network is proposed, which can accelerate the convergence of the dehazing network to meet the application requirements of practical industrial systems. A large number of experimental results on synthetic and real-world datasets show that the proposed MADN satisfies the needs of industrial systems.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Guoqiang Li 0001
IEEE Trans. Ind. Informatics3
2022 Detecting Absence of Bone Wall in Jugular Bulb by Image Transformation Surrogate Tasks
Yichao Zhou 0002, Hongxia Yin, Zhenchang Wang, Li Zhuo 0001, Hui Zhang 0049
IEEE Trans. Medical Imaging5
2022 Computation Offloading With Instantaneous Load Billing for Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a promising approach that can reduce the latency of task processing by offloading tasks from user equipments (UEs) to MEC servers. Existing works always assume that the MEC server is capable of executing the offloaded tasks, without considering the impact of improper load on task processing efficiency. In this article, we present a two-stage computing offloading scheme to minimize the task processing delay while managing the server load properly. To minimize the task processing delay, each UE optimizes how much workload to be offloaded to the MEC server. To improve the task processing efficiency of the server, we arrange the processing order of offloading tasks by introducing an aggregative game with an instantaneous load billing mechanism. The proposed game can obtain the optimal task offloading and processing strategy with limited information and a small number of iterations. Simulation results show that our scheme approaches the optimal offloading strategy in terms of minimizing task processing delay for each UE and improving processing efficiency for the server.
Mingjin Gao, Rujing Shen, Jun Li 0004, Shihao Yan, Yonghui Li 0001, Jinglin Shi, Zhu Han 0001, Li Zhuo 0001
IEEE Trans. Serv. Comput.8
2021 Crowd activity recognition in live video streaming via 3D-ResNet and region graph convolution network
abstract
Abstract Since the era of we‐media, live video industry has shown an explosive growth trend. For large‐scale live video streaming, especially those containing crowd events that may cause great social impact, how to identify and supervise the crowd activity in live video streaming effectively is of great value to push the healthy development of live video industry. The existing crowd activity recognition mainly uses visual information, rarely fully exploiting and utilizing the correlation or external knowledge between crowd content. Therefore, a crowd activity recognition method in live video streaming is proposed by 3D‐ResNet and regional graph convolution network (ReGCN). (1) After extracting deep spatiotemporal features from live video streaming with 3D‐ResNet, the region proposals are generated by region proposal network. (2) A weakly supervised ReGCN is constructed by making region proposals as graph nodes and their correlations as edges. (3) Crowd activity in live video streaming is recognised by combining the output of ReGCN, the deep spatiotemporal features and the crowd motion intensity as external knowledge. Four experiments are conducted on the public collective activity extended dataset and a real‐world dataset BJUT‐CAD. The competitive results demonstrate that our method can effectively recognise crowd activity in live video streaming.
Junpeng Kang, Jing Zhang 0023, Li Zhuo 0001
IET Image Process.4
2021 A discriminative self-attention cycle GAN for face super-resolution and recognition
abstract
Abstract Face image captured via surveillance videos in an open environment is usually of low quality, which seriously affects the visual quality and recognition accuracy. Most image super‐resolution methods adopt paired high‐quality and its interpolated low‐resolution version to train the super‐resolution network. It is difficult to achieve contented visual quality and restoring discriminative features in real scenarios. A discriminative self‐attention cycle generative adversarial network is proposed for real‐world face image super‐resolution. Based on the cycle GAN framework, unpaired samples are adopted to train a degradation network and a reconstruction network simultaneously. A self‐attention mechanism is employed to capture the contextual information for details restoring. A Siamese face recognition network is introduced to provide a constraint on identify consistency. In addition, an asymmetric perceptual loss is introduced to handle the imbalance between the degradation model and the reconstruction model. Experimental results show that the observation model achieved more realistic low‐quality face images, and the super‐resolved face images have shown better subjective quality and higher face recognition performance.
Jianglu Huang, Li Zhuo 0001, Jiafeng Li 0001
IET Image Process.4
2021 Blind image quality assessment with channel attention based deep residual network and extended LargeVis dimensionality reduction
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Meng Wang 0001
J. Vis. Commun. Image Represent.2
2021 Kernelized Multiview Subspace Analysis By Self-Weighted Learning
abstract
With the popularity of multimedia technology, information is always represented from multiple views. Even though multiview data can reflect the same sample from different perspectives, multiple views are consistent to some extent because they are representations of the same sample. Most of the existing algorithms are graph-based ones to learn the complex structures within multiview data but overlook the information within data representations. Furthermore, many existing works treat multiple views discriminatively by introducing some hyperparameters, which is undesirable in practice. To this end, abundant multiview-based methods have been proposed for dimension reduction. However, there is still no research that leverages the existing work into a unified framework. In this paper, we propose a general framework for multiview data dimension reduction, named kernelized multiview subspace analysis (KMSA) to handle multiview feature representation in the kernel space, providing a feasible channel for multiview data with different dimensions. Compared with the graph-based methods, KMSA can fully exploit information from multiview data with nothing to lose. Since different views have different influences on KMSA, we propose a self-weighted strategy to treat different views discriminatively. A co-regularized term is proposed to promote the mutual learning from multiviews. KMSA combines self-weighted learning with the co-regularized term to learn the appropriate weights for all views. We evaluate our proposed framework on 6 multiview datasets for classification and image retrieval. The experimental results validate the advantages of our proposed method.
Huibing Wang, Yang Wang 0023, Zhao Zhang 0001, Xianping Fu, Li Zhuo 0001, Mingliang Xu 0001, Meng Wang 0001
IEEE Trans. Multim.5
2021 RMoR-Aion: Robust Multioutput Regression by Simultaneously Alleviating Input and Output Noises
abstract
Multioutput regression, referring to simultaneously predicting multiple continuous output variables with a single model, has drawn increasing attention in the machine learning community due to its strong ability to capture the correlations among multioutput variables. The methodology of output space embedding, built upon the low-rank assumption, is now the mainstream for multioutput regression since it can effectively reduce the parameter numbers while achieving effective performance. The existing low-rank methods, however, are sensitive to the noises of both inputs and outputs, referring to the noise problem. In this article, we develop a novel multioutput regression method by simultaneously alleviating input and output noises, namely, robust multioutput regression by alleviating input and output noises (RMoR-Aion), where both the noises of the input and output are exploited by leveraging auxiliary matrices. Furthermore, we propose a prediction output manifold constraint with the correlation information regarding the output variables to further reduce the adversarial effects of the noise. Our empirical studies demonstrate the effectiveness of RMoR-Aion compared with the state-of-the-art baseline methods, and RMoR-Aion is more stable in the settings with artificial noise.
Ximing Li 0002, Yang Wang 0023, Zhao Zhang 0001, Richang Hong, Li Zhuo 0001, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2020 3D Hand Pose Estimation with Disentangled Cross-Modal Latent Space
abstract
Estimating 3D hand pose from a single RGB image is a challenging task because of its ill-posed nature (i.e., depth ambiguity). Recently, various generative approaches have been proposed to predict the 3D joints of an RGB hand image by learning a unified latent space between two modalities (i.e., RGB image and 3D joints). However, projecting multi-modal data (i.e., RGB images and 3D joints) into a unified latent space is difficult as the modality-specific features usually interfere the learning of the optimal latent space. Hence in this paper, we propose to disentangle the latent space into two sub-latent spaces: modality- specific latent space and pose-specific latent space for 3D hand pose estimation. Our proposed method, namely Disentangled Cross-Modal Latent Space (DCMLS), consists of two variational autoencoder networks and auxiliary components which connect the two VAEs to align underlying hand poses and transfer modality-specific context from RGB to 3D. For the hand pose latent space, we align it with the two modalities by using a cross-modal discriminator with an adversarial learning strategy. For the context latent space, we learn a context translator to gain access to the cross-modal context. Experimental results on two widely used public benchmark datasets RHD and STB demonstrate that our proposed DCMLS method is able to clearly outperform the state-of-the-art ones on single image based 3D hand pose estimation.
Jiajun Gu, Zhiyong Wang 0001, Wanli Ouyang, Jiafeng Li 0001, Li Zhuo 0001
WACV6
2020 Coarse-to-fine object detection in unmanned aerial vehicle imagery using lightweight convolutional neural network and deep motion saliency
Jing Zhang 0023, Xi Liang 0003, Meng Wang 0001, Liheng Yang, Li Zhuo 0001
Neurocomputing5
2020 Multi-level prediction Siamese network for real-time UAV visual tracking
Hui Zhang 0049, Jing Zhang 0023, Li Zhuo 0001
Image Vis. Comput.4
2020 A 3D deep supervised densely network for small organs of human temporal bone segmentation in CT images
Zhaopeng Gong, Hongxia Yin, Hui Zhang 0049, Zhenchang Wang, Li Zhuo 0001
Neural Networks6
2020 Multilevel fusion of multimodal deep features for porn streamer recognition in live video
Jing Zhang 0023, Meng Wang 0001, Jimiao Tian, Li Zhuo 0001
Pattern Recognit. Lett.5
2020 Small Object Detection in Unmanned Aerial Vehicle Images Using Feature Fusion and Scaling-Based Single Shot Detector With Spatial Context Analysis
abstract
Objects in unmanned aerial vehicle (UAV) images are generally small due to the high-photography altitude. Although many efforts have been made in object detection, how to accurately and quickly detect small objects is still one of the remaining open challenges. In this paper, we propose a feature fusion and scaling-based single shot detector (FS-SSD) for small object detection in the UAV images. The FS-SSD is an enhancement based on FSSD, a variety of the original single shot multibox detector (SSD). We add an extra scaling branch of the deconvolution module with an average pooling operation to form a feature pyramid. The original feature fusion branch is adjusted to be better suited to the small object detection task. The two feature pyramids generated by the deconvolution module and feature fusion module are utilized to make predictions together. In addition to the deep features learned by the FS-SSD, to further improve the detection accuracy, spatial context analysis is proposed to incorporate the object spatial relationships into object redetection. The interclass and intraclass distances between different object instances are computed as a spatial context, which proves effective for multiclass small object detection. Six experiments are conducted on the PASCAL VOC dataset and the two UAV image datasets. The experimental results demonstrate that the proposed method can achieve a comparable detection speed but an accuracy superior to those of the six state-of-the-art methods.
Xi Liang 0003, Jing Zhang 0023, Li Zhuo 0001, Yuzhao Li, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Porn Streamer Recognition in Live Video Streaming via Attention-Gated Multimodal Deep Features
abstract
Live video streaming platforms have attracted millions of streamers and daily active users. For profit and popularity accumulation, some streamers mix pornography content into live content to avoid online supervision. Therefore, accurate recognition of porn streamers in live video streaming has become a challenging task. Porn streamers in live video present multimodal characteristics including visual and acoustic content. Therefore, a porn streamer recognition method in live video streaming is proposed that uses attention-gated multimodal deep features. Our contribution includes the following: (1) multimodal deep features, i.e., spatial, motion and audio, are extracted from live video streaming using convolutional neural networks (CNNs), in which the temporal context of multimodal features is obtained with a bi-directional gated recurrent unit (Bi-GRU); (2) the tri-attention gated mechanism is applied to map the associations between different modalities by assigning higher weights to important features for further reduction in the redundancy of multimodal features; (3) porn streamers in live video streaming are recognized via the attention-gated multimodal deep features. Six experiments are conducted on a real-world dataset, and the competitive results demonstrate that our method can effectively recognize porn streamers in live video streaming.
Jing Zhang 0023, Qi Tian 0001, Li Zhuo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 Effective Data-Driven Technology for Efficient Vision-Based Outdoor Industrial Systems
abstract
Vision systems are the core information collection module in outdoor industrial systems such as factory inspection robots. However, haze greatly reduces working efficiency. Existing dehazing methods have two problems-first, they are not specifically designed for the industrial systems; second, these methods include several assumptions in their design processes and imaging models, leading to unsatisfactory results. In this article, an approach for single image dehazing is proposed to improve the efficiency of outdoor vision-based systems. First, a novel haze imaging model is proposed based on the dichromatic atmospheric scattering model. It considers the effects of multiple scattering and involves fewer assumptions. Then a data-driven technique called sparse representation is used to solve this model. Considering a haze image, a distorted and blurred version of a fine image, every patch is presented using dedicatedly prepared over-complete dictionaries and is traced back to a haze-free image. Quantitative and qualitative comparisons on a number of real-world haze images demonstrate that the proposed approach not only is more stable but also leads to better dehazing results.
Jiafeng Li 0001, Li Zhuo 0001, Hong Zhang 0018, Guoqiang Li 0001, Naixue Xiong
IEEE Trans. Ind. Informatics2
2020 Prediction and Communication Co-Design for Ultra-Reliable and Low-Latency Communications
abstract
Ultra-reliable and low-latency communications (URLLC) are considered as one of three new application scenarios in the fifth generation cellular networks. In this work, we aim to reduce the user experienced delay through prediction and communication co-design, where each mobile device predicts its future states and sends them to a data center in advance. Since predictions are not error-free, we consider prediction errors and packet losses in communications when evaluating the reliability of the system. Then, we formulate an optimization problem that maximizes the number of URLLC services supported by the system by optimizing time and frequency resources and the prediction horizon. Simulation results verify the effectiveness of the proposed method, and show that the tradeoff between user experienced delay and reliability can be improved significantly via prediction and communication co-design. Furthermore, we carried out an experiment on the remote control in a virtual factory, and validated our concept on prediction and communication co-design with the practical mobility data generated by a real tactile device.
Zhanwei Hou, Changyang She, Yonghui Li 0001, Li Zhuo 0001, Branka Vucetic
IEEE Trans. Wirel. Commun.4
2020 Hybrid-Precoding for mmWave Multi-User Communications in the Presence of Beam-Misalignment
abstract
In this paper, we propose the hybrid-precoding design that alleviates the performance loss caused by beam-misalignment in the mmWave multi-user communication systems. To this end, we firstly design the beam-misalignment aware fully-digital precoders for two distinct scenarios. First, for a base-station (BS) with full estimated channel-state-information (CSI), the minimum-mean-squared-error metric incorporating the `error-statistics' of the beam-misalignment error is used to analytically derive a closed-form expression for the fully-digital precoder, which maximizes the array gain while suppressing the inter-user interference for each user-equipment (UE). Second, for a BS which can only acquire partial estimated CSI, a min-max non-convex optimization is considered to obtain the fully-digital precoder, which minimizes the maximum loss in the array gains of the expected beam-misalignment `error-range' over the UEs while cancelling the inter-user interference. Subsequently, we propose the hybrid-precoding design that approximates the fully-digital designs based on the gradient-projection method, which is mathematically proven to converge to an approximate local solution with further reduced complexity compared to the state-of-the-art algorithms. Finally, the proposed hybrid-precoding design is further extended to the wideband mmWave communication systems. Numerical results show that the proposed hybrid-precoding design can effectively alleviate the performance degradation incurred by the beam-misalignment.
Chandan Pradhan, Ang Li 0003, Li Zhuo 0001, Yonghui Li 0001, Branka Vucetic
IEEE Trans. Wirel. Commun.3
2019 Face Super-Resolution via Discriminative-Attributes
Jiafeng Li 0001, Li Zhuo 0001
PRCV (2)4
2019 Deep-network based method for joint image deblocking and super-resolution
abstract
Many pieces of research have been conducted on image‐restoration techniques to recover high‐quality images from their low‐quality versions, but they usually aim to handle a single degraded factor. However, captured images usually suffer from various degradation factors, such as low resolution and compression distortion, in the procedures of image acquisition, compression, and transmission simultaneously. Ignoring the correlation of different degraded factors may result in the limited efficiency of the existing image‐restoration methods for captured images. A joint deep‐network‐based image‐restoration algorithm is proposed to establish a restoration framework for image deblocking and super‐resolution. The proposed convolutional neural network is made up of two stages. A deblocking network is constructed with two cascade deblocking subnets first, then, super‐resolution is performed by a very deep network with skipping links. Cascading these two stages forms a novel deep network. An end‐to‐end training scheme is developed, which makes the two stages be trained jointly so as to achieve better performance. Intensive evaluations have been conducted to measure the performance of the authors’ method both in general images and face images. Experimental results on several datasets demonstrate that the proposed method outperforms other state‐of‐the‐art methods, in terms of both subjective and objective performances.
Kin-Man Lam 0001, Li Zhuo 0001, Jiafeng Li 0001
IET Image Process.4
2019 Personalized Recommendation of Social Images by Constructing a User Interest Tree With Deep Features and Tag Trees
abstract
In view of the great diversity and complexity of social images, it is of great significance to improve the performance of personalized recommendation by learning a user interest from large-scale social images. Deep learning, as the latest research in the field of artificial intelligence, provides a new personalized recommendation solution of social images for learning a user's interest. Moreover, social image sharing websites (such as Flickr) allow users to tag uploaded images with tags. As an important image semantic cue, effective tags not only represent the latent image information but also show personalized user interest. Therefore, a personalized recommendation method of social image is proposed by constructing a user-interest tree with deep features and tag trees in this paper. The main contributions of our paper are as follows: first, to efficiently make use of tags, a tag tree of social images is created by the re-ranked tags; second, for compactly representing the image content, deep features are learned by training the AlexNet network; third, a user-interest tree is constructed with deep features and tag trees that include the user-interest tree of social images and the user-interest tree of tags, respectively, and finally, a personalized recommendation system of social images is built based on a user-interest tree. Experiments on the NUS-WIDE dataset have shown that our method outperforms state-of-the-art methods in terms of both precision and recall of personalized recommendations.
Jing Zhang 0023, Ying Yang 0018, Li Zhuo 0001, Qi Tian 0001, Xi Liang 0003
IEEE Trans. Multim.3
2019 Cross-Modality Retrieval by Joint Correlation Learning
abstract
As an indispensable process of cross-media analyzing, comprehending heterogeneous data faces challenges in the fields of visual question answering (VQA), visual captioning, and cross-modality retrieval. Bridging the semantic gap between the two modalities is still difficult. In this article, to address the problem in cross-modality retrieval, we propose a cross-modal learning model with joint correlative calculation learning. First, an auto-encoder is used to embed the visual features by minimizing the error of feature reconstruction and a multi-layer perceptron (MLP) is utilized to model the textual features embedding. Then we design a joint loss function to optimize both the intra- and the inter-correlations among the image-sentence pairs, i.e., the reconstruction loss of visual features, the relevant similarity loss of paired samples, and the triplet relation loss between positive and negative examples. In the proposed method, we optimize the joint loss based on a batch score matrix and utilize all mutual mismatched paired samples to enhance its performance. Our experiments in the retrieval tasks demonstrate the effectiveness of the proposed method. It achieves comparable performance to the state-of-the-art on three benchmarks, i.e., Flickr8k, Flickr30k, and MS-COCO.
Shuo Wang 0008, Dan Guo 0001, Xin Xu 0007, Li Zhuo 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2018 An efficient method of content-targeted online video advertising
Guanyao Wang, Li Zhuo 0001, Jiafeng Li 0001, Dongyue Ren, Jing Zhang 0023
J. Vis. Commun. Image Represent.2
2018 Vehicle color recognition using Multiple-Layer Feature Representations of lightweight convolutional neural network
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Hui Zhang 0049
Signal Process.2
2017 Automatic Tongue Image Segmentation for Traditional Chinese Medicine Using Deep Neural Network
Panling Qu, Hui Zhang 0049, Li Zhuo 0001, Jing Zhang 0023, Guoying Chen
ICIC (1)3
2017 Tag tree creation of social image for personalized recommendation
abstract
The tags are usually tagged by different users in social image sharing websites, which can indicate image semantic information and imply user's preference. Therefore, the tags can contribute to personalized recommendation of social image. However, the present social image tags models only consider single tag, resulting in the relationships among tags are ignored. In this paper, we propose a novel method to create tag tree of social image for personalized recommendation. Firstly, the tag ranking is realized to remove noisy tags. Then, the first layer tags are selected from re-ranked tags lists. To sufficiently express tag's significances, the tag subtrees can be created based on different image categories and combined with first layer tags to create tag tree. Finally, the personalized recommendation of social image is achieved by using tag tree. Experimental results show that our tag tree can effectively express the relationships among tags as well as obtain satisfactory results in personalized recommendation of social image.
Ying Yang 0018, Jing Zhang 0023, Jihong Liu, Jiafeng Li 0001, Li Zhuo 0001
ICIP5
2017 A joint deep-network-based image restoration algorithm for multi-degradations
abstract
In the procedures of image acquisition, compression, and transmission, captured images usually suffer from various degradations, such as low-resolution and compression distortion. Although there have been a lot of research done on image restoration, they usually aim to deal with a single degraded factor, ignoring the correlation of different degradations. To establish a restoration framework for multiple degradations, a joint deep-network-based image restoration algorithm is proposed in this paper. The proposed convolutional neural network is composed of two stages. Firstly, a de-blocking subnet is constructed, using two cascaded neural network. Then, super-resolution is carried out by a 20-layer very deep network with skipping links. Cascading these two stages forms a novel deep network. Experimental results on the Set5, Setl4 and BSD100 benchmarks demonstrate that the proposed method can achieve better results, in terms of both the subjective and objective performances.
Li Zhuo 0001, Kin-Man Lam 0001, Jiafeng Li 0001
ICME3
2017 A CBIR System for Hyperspectral Remote Sensing Images Using Endmember Extraction
abstract
With the rapid development of remote sensing technology, searching the similar image is a challenge for hyperspectral remote sensing image processing. Meanwhile, the dramatic growth in the amount of hyperspectral remote sensing data has stimulated considerable research on content-based image retrieval (CBIR) in the field of remote sensing technology. Although many CBIR systems have been developed, few studies focused on the hyperspectral remote sensing images. A CBIR system for hyperspectral remote sensing image using endmember extraction is proposed in this paper. The main contributions of our method are that: (1) the endmembers as the spectral features are extracted from hyperspectral remote sensing image by improved automatic pixel purity index (APPI) algorithm; (2) the spectral information divergence and spectral angle match (SID–SAM) mixed measure method is utilized as a similarity measurement between hyperspectral remote sensing images. At last, the images are ranked with descending and the top-[Formula: see text] retrieved images are returned. The experimental results on NASA datasets show that our system can yield a superior performance.
Jing Zhang 0023, Qianlan Zhou, Li Zhuo 0001, Wenhao Geng, Suyu Wang
Int. J. Pattern Recognit. Artif. Intell.3
2017 Vehicle classification for large-scale traffic surveillance videos using Convolutional Neural Networks
Li Zhuo 0001, Liying Jiang, Jiafeng Li 0001, Jing Zhang 0023
Mach. Vis. Appl.1
2017 Road Recognition From Remote Sensing Imagery Using Incremental Learning
abstract
Roads, as important artificial objects, are the main body of modern traffic system, providing many conveniences for human civilization. With the development of Intelligent Transportation Systems (ITS), the road structure is changing frequently. Road recognition is to identify the road type from remote sensing imagery, and road types depend largely on the characteristics of roads. Thus, how to extract road features and further making road classification efficient have become a popular and challenging research topic. In this paper, we propose a road recognition method for remote sensing imagery using incremental learning. In principle, our method includes the following steps: 1) the non-road remote sensing imagery is first filtered by using support vector machine; 2) the road network is obtained from the road remote sensing imagery by computing multiple saliency features; 3) the road features are extracted from road network and background environment; and 4) the roads are recognized as three road types according to the classification results of incremental learning algorithm. The experimental results show that our method has higher road recognition rate as well as less recognition time than the other popular algorithms.
Jing Zhang 0023, Chao Wang 0017, Li Zhuo 0001, Qi Tian 0001, Xi Liang 0003
IEEE Trans. Intell. Transp. Syst.4
2017 Personalized Social Image Recommendation Method Based on User-Image-Tag Model
abstract
In the social image sharing websites (such as Flickr), users are allowed to upload images and tag them with tags. Due to the diversities of users' interests, different users may tag the same image with different tags. Therefore, tags not only reveal some important image semantic clues, but also show user's preference, which can provide a new effective solution for overcoming the semantic gap as well as realizing a personalized recommendation. In this paper, a personalized social image recommendation method based on user-image-tag model is proposed. The main contributions of our work are 1) to efficiently make use of tags, social image tags are re-ranked according to the image content; 2) to obtain user preference, a user-image-tag model is constructed with tripartite graph according to the correlation among users, images and top-ranking tags; and 3) a personalized social recommendation system is implemented based on user-image-tag model. Experimental results proved that our method can significantly improve the accuracy of personalized image recommendation.
Jing Zhang 0023, Ying Yang 0018, Qi Tian 0001, Li Zhuo 0001, Xin Liu 0016
IEEE Trans. Multim.4
2016 Extreme Weather Recognition Using Convolutional Neural Networks
abstract
Extreme weather always brings potential risk to driving, which leads to people's life and property being put into great dangers. Therefore, the automatic recognition of extreme weather plays an important role in the application of the highway traffic condition warning, automobile auxiliary driving, climate analysis and so on. Generally, multiple sensors are adopted in traditional methods of automatic extreme weather recognition with artificial participation and low accuracy. A new extreme weather recognition method based on images by using computer vision manners has been proposed in this paper. Since the weather is affected by many factors, features that can accurately represent various weather characteristics are difficult to be extracted. Therefore, in this paper, convolutional neural networks (CNNs) are applied to settle this problem. Features of extreme weather and recognition models are generated from big data. Moreover, a large-scale extreme weather dataset, "WeatherDataset", has been collected, in which 16635 extreme weather images are divided into four classes (sunny, rainstorm, blizzard, and fog), and complex scenes are coverd. A recognition model for extreme weather is obtained through two steps: Pre-training and Fine Tuning. In Pre-training step, ILSVRC-2012 Dataset is trained to obtain the model of ILSVRC using GoogLeNet. A more accurate model for extreme weather recognition is obatined by further fine-tuning GoogLeNet on WeatherDataset. The experimental results show that the proposed method is able to achieve a high performance with the recognition accuracy rate of 94.5% and can meet the requirements of some real applications.
Li Zhuo 0001, Panling Qu, Kailong Zhou, Jing Zhang 0023
ISM2
2016 Automatic Endmember Extraction Using Pixel Purity Index for Hyperspectral Imagery
Qianlan Zhou, Jing Zhang 0023, Qi Tian 0001, Li Zhuo 0001, Wenhao Geng
MMM (2)4
2016 ORB feature based web pornographic image recognition
Li Zhuo 0001, Zhen Geng, Jing Zhang 0023
Neurocomputing1
2016 A K-PLSR-based color correction method for TCM tongue images under different illumination conditions
Li Zhuo 0001, Pei Zhang 0012, Panling Qu, Yuanfan Peng, Jing Zhang 0023
Neurocomputing1
2015 Creating descriptive visual words for tag ranking of compressed social image
abstract
Visual words description method has been widely applied in the fields of social image's tag ranking, tag recommendation and annotation. At present, visual words are usually obtained by unsupervised clustering methods which lead to generate many unnecessary and non-descriptive words. Therefore, how to make visual words be descriptive has become a very meaningful task for tag ranking of social image. However, for compressed social image on the network, visual words are created after fully decompressing a compressed image into pixel domain. In this paper, creating descriptive visual words in compressed domain is proposed for tag ranking of compressed social image. Firstly, the traditional visual words are created by using the partly decoded data; then the descriptive visual words are selected from traditional visual words by the VisualWordRank ranking algorithm; finally the descriptive visual words are applied to rank the tag of social image. Experimental results show the descriptive visual words can improve the accuracy of tag ranking, which further prove our method has more descriptive ability. Besides that, our method also reduces the processing time for compressed social image greatly.
Xin Liu 0016, Jing Zhang 0023, Li Zhuo 0001, Ying Yang 0018
ICIP3
2015 An image deblocking method based on pre-classified sparse representation
abstract
Block-based Discrete Cosine Transform (BDCT) image compression method inevitably produces annoying blocking artifacts in the case of low-bit-rate compression, as each block is transformed and quantized independently. Blocking artifacts not only seriously affect the subjective image quality, but also affect the performance of automatic analysis. In this paper, we propose a pre-classified sparse representation based deblocking method. We combine the human visual sensitivity based classification and the sparse representation method together. For different contents, the reconstruction threshold can been adaptively adjusted. Experimental results show that the proposed method can improve the deblocking quality of the compressed images effectively.
Li Zhuo 0001
MMSP4
2015 Social images tag ranking based on visual words in compressed domain
Jing Zhang 0023, Xin Liu 0016, Li Zhuo 0001, Chao Wang 0017
Neurocomputing3
2015 A Compressed-Domain Image Filtering and Re-Ranking Approach for Multi-Agent Image Retrieval
abstract
For the limited transmission capacity and compressed images in the network environment, a compressed-domain image filtering and re-ranking approach for multi-agent image retrieval is proposed in this paper. Firstly, the distributed image retrieval platform with multi-agent is constructed by using Aglet development system, the lifecycle and the migration mechanism of agent is designed and planned for multi-agent image retrieval by using the characteristics of mobile agent. Then, considering the redundant image brought by distributed multi-agent retrieval, the duplicate images in distributed retrieval results are filtered based on the perceptual hashing feature extracted in the compressed-domain. Finally, weight-based hamming distance is utilized to re-rank the retrieval results. The experimental results show that the proposed approach can effectively filter the duplicate images in distributed image retrieval results as well as improve the accuracy and speed of compressed-domain image retrieval.
Jing Zhang 0023, Li Zhuo 0001, Xin Liu 0016, Ying Yang 0018
Int. J. Pattern Recognit. Artif. Intell.3
2015 An iteratively reweighting algorithm for dynamic video summarization
Pei Dong, Yong Xia 0001, Shanshan Wang 0002, Li Zhuo 0001, David Dagan Feng
Multim. Tools Appl.4
2014 Uniform color space based facial complexion recognition for Traditional Chinese Medicine
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the facial complexion to obtain the disease diagnostic results. Color space based on human visual system will be more conducive to facial complexion recognition, which is more suitable to measure and distinguish facial complexion. Uniform color space based facial complexion recognition for TCM is proposed in this paper, which include: (1) the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately; (2) the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature; (3) the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the proposed complexion feature can achieve good performance, with the facial complexion recognition rate up to 81%.
Jing Zhang 0023, Chao Wang 0017, Li Zhuo 0001, Yuncong Yang
ICARCV3
2014 Automatic tongue color analysis of traditional Chinese medicine based on image retrieval
abstract
Content-Based Image Retrieval (CBIR) characterizes the image content by extracting visual features, and measures the similarity according to the distance between the two features. This paper adopts CBIR to perform automatic tongue color analysis of Traditional Chinese Medicine (TCM). Firstly, we extract the visual features of tongue images to be analyzed, especially the color features; and then retrieve the similar tongue images from the database, which have been labeled by TCM doctors in advance. Finally, statistical decision method is exploited based on the retrieval results to classify the tongue color. Experimental results show that the proposed method can achieve the classification accuracy of 87.85% and 88.54% respectively for the colors of tongue substance and tongue coating. The proposed method in this paper can provide a new means for the tongue color automatic analysis of TCM, and it is also a new application of CBIR.
Li Zhuo 0001, Pei Zhang 0012, Bo Cheng 0007, Jing Zhang 0023
ICARCV1
2014 Surf feature extraction in encrypted domain
abstract
Signal processing in the encrypted domain has become a hot research topic, which enable signal processing tasks in a secure and privacy-preserving manner. Taken the fact that SURF (Speeded Up Robust Feature) has been widely utilized in various applications into account, SURF feature extraction method in the encrypted domain has been proposed in this paper. Because all steps must be implemented in the encrypted domain, Paillier homomorphic encryption method is adopted. Experimental results demonstrate that the number and location of SURF features extracted from the encrypted data are the same as those from the plaintext data. And the error between the descriptors obtained from the plaintext data and the encrypted data is only 0.0002932%. We also provide security analysis and complexity analysis. The proposed method can be used in the encrypted domain based applications, such as secure image processing and image retrieval.
Yu Bai 0008, Li Zhuo 0001, Bo Cheng 0007, Yuanfan Peng
ICME2
2014 An efficient motion reference structure based selective encryption algorithm for H.264 videos
abstract
In this study, based on both the prediction mechanism of H.264 encoder and the syntax of H.264 bitstream, an efficient selective video encryption algorithm is proposed. The contributions of the study include two aspects. First, motion reference ratio (MRR) of macroblock (MB) is proposed to describe the inter‐frame dependency among the adjacent frames. At the MB layer, MRRs of MBs are statistically analysed, and MBs to be encrypted are selected based on the statistical results. Second, at the bitstream layer of MBs, bit‐sensitivity is proposed to represent the degree of importance of each bit in the compressed bitstream for reconstructed video quality. The most significant bits for reconstructed video quality are selected to be encrypted based on the bit‐sensitivity of H.264 bitstream. The intra‐prediction mode codewords, the sign bits of the non‐zero coefficients and the info_suffix of motion vector difference codewords are extracted to be encrypted. The proposed two‐layer selection scheme improves the encryption efficiency significantly. Experimental results demonstrate that both perceptual security and cryptographic security are achieved, and compared with the existing SEH264 algorithm, the proposed selective encryption algorithm can reduce the computational complexity by 50% on average.
Haojie Shen, Li Zhuo 0001, Yingdi Zhao
IET Inf. Secur.2
2014 A comparative study of dimensionality reduction methods for large-scale image retrieval
Li Zhuo 0001, Bo Cheng 0007, Jing Zhang 0023
Neurocomputing1
2014 An SA-GA-BP neural network-based color correction algorithm for TCM tongue images
Li Zhuo 0001, Jing Zhang 0023, Pei Dong, Yingdi Zhao
Neurocomputing1
2014 Human Facial Complexion Recognition of Traditional Chinese Medicine Based on Uniform Color Space
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the human facial complexion to obtain the disease diagnostic results. The morbidity of the organs can be revealed by the human facial complexion, so the color space based on human visual system will be more conducive to facial complexion recognition. It is much suitable to measure and distinguish facial complexion by uniform Lab color space, as it has the characteristic of isometry and high resolving power. First, the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately. Second, the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature. At last, the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the complexion feature proposed in this paper can achieve the better performance, with the facial complexion recognition rate up to 81%.
Li Zhuo 0001, Yuncong Yang, Jing Zhang 0023
Int. J. Pattern Recognit. Artif. Intell.1
2013 Comparative Study on Dimensionality Reduction in Large-Scale Image Retrieval
abstract
Dimensionality reduction plays a significant role for the performance of large-scale image retrieval. In this paper, various dimensionality reduction methods are compared to validate their own performance in image retrieval. For this purpose, first, the Scale Invariant Feature Transform (SIFT) features and HSV (Hue, Saturation, Value) histogram are extracted as image features. Second, the Principal Component Analysis (PCA), Fisher Linear Discriminant Analysis (FLDA), Local Fisher Discriminant Analysis (LFDA), Isometric Mapping (ISOMAP), Locally Linear Embedding (LLE), and Locality Preserving Projections (LPP) are respectively applied to reduce the dimensions of SIFT feature descriptors and color information, which can be used to generate vocabulary trees. Finally, through setting the match weights of vocabulary trees, large-scale image retrieval scheme is implemented. By comparing multiple sets of experimental data from several platforms, it can be concluded that dimensionality reduction method of LLE and LPP can effectively reduce the computational cost of image features, and maintain the high retrieval performance as well.
Bo Cheng 0007, Li Zhuo 0001, Jing Zhang 0023
ISM2
2013 Layered-based exposure fusion algorithm
abstract
Owing to the limitation of dynamic range, a single still image is usually insufficient to describe a high contrast scene. Fusing multi‐exposure images of the same scene can produce a resulting image with details both in the bright and the dark regions. However, they may be sensitive to the exposure parameters of the input images. In this study, a global layer is introduced to improve the robustness of the fusion method. The global layer is employed to preserve the overall luminance of a real scene and avoid possible luminance reversion artefacts. Then, details are recovered in the gradient domain by a Poisson solver. Experimental results show the superior performance of our approach in terms of robustness and details preservation.
Fenghui Li, Li Zhuo 0001, David Dagan Feng
IET Image Process.3
2013 Pornographic image region detection based on visual attention model in compressed domain
abstract
According to biological attention mechanism, a region of interest (ROI) detection based on visual attention model is closer to human visual system. Taken into account the characteristics of pornographic image during regions detection, a pornographic image region detection method based on visual attention model in compressed domain is proposed in this study, which includes the following four steps: (i) the skin colour regions of pornographic images are detected in compressed domain; (ii) visual saliency map in compressed domain is computed to construct visual attention model; (iii) threshold segmentation method is used for visual saliency map, and then the torso information is retained as pornographic regions; and (iv) four features of colour, texture, intensity and skin are extracted to represent pornographic region. The experimental results show that the proposed method can perform well on the speed/accuracy of pornographic regions detection and representation.
Jing Zhang 0023, Lei Sui, Li Zhuo 0001
IET Image Process.3
2013 An approach of bag-of-words based on visual attention model for pornographic images recognition in compressed domain
Jing Zhang 0023, Lei Sui, Li Zhuo 0001, Yuncong Yang
Neurocomputing3
2013 Semantic context based refinement for news video annotation
Zhiyong Wang 0001, Genliang Guan, Li Zhuo 0001, David Dagan Feng
Multim. Tools Appl.4
2013 Learning realistic facial expressions from web images
Kaimin Yu, Zhiyong Wang 0001, Li Zhuo 0001, Zheru Chi, David Dagan Feng
Pattern Recognit.3
2013 Compressed domain based pornographic image recognition using multi-cost sensitive decision trees
Li Zhuo 0001, Jing Zhang 0023, Yingdi Zhao
Signal Process.1
2012 Region of Interest Detection Based on Visual Perception Model
abstract
According to human vision theory, the image is conveyed from human visual system to brain when people have a look at. Different from previous work, the study reported in this paper attempts to simulate a more real and complex method for region of interest (ROI) detection and quantitatively analyze the correlation between users' visual perception and ROI. In this paper, a visual perception model-based ROI detection is proposed, which can be realized with an ordinary web camera. Visual perception model employs a combination of visual attention model and gaze tracking data to objectively detect ROIs. The work includes pre-ROI estimation using visual attention model, gaze data collection and ROI detection. Pre-ROIs are segmented by the visual attention model. Since eye feature extraction is critical to the accuracy and performance of gaze tracking, adaptive eye template and neural network are employed to predict gaze points. By computing the density of the gaze points, ROIs are ranked. Experimental results show that the accuracy of our ROI detection method can be raised as high as 97% and it is also demonstrated that our model can efficiently adapt to users' interests and match the objective ROI.
Jing Zhang 0023, Li Zhuo 0001, Yingdi Zhao
Int. J. Pattern Recognit. Artif. Intell.2
2011 Real-time moving object segmentation and tracking for H.264/AVC surveillance videos
abstract
With increased use of H.264/AVC in various applications including video surveillance systems, feature extraction and knowledge representation in compressed domain are becoming attractive. A real-time H.264/AVC compressed domain moving object segmentation and tracking algorithm for surveillance videos is proposed in this paper. This algorithm consists of moving object detection, bounding box matching, spatiotemporal merge and split reasoning and trajectory smoothing, with major innovation in incorporating the information provided by the prediction modes into the framework of motion detection and trajectory construction. The experimental results on both indoor and outdoor surveillance videos demonstrate that the adaptive use of the information from motion vectors, DCT coefficients and prediction modes can substantially improve the performance of moving object segmentation and tracking.
Pei Dong, Yong Xia 0001, Li Zhuo 0001, David Dagan Feng
ICIP3
2011 Regions of Interest Extraction Based on Visual Saliency in Compressed Domain
abstract
Recently bag-of-words (BoW) model having been widely used in textual information processing has been extended into many tasks in visual domain such as image classification, scene analysis, image annotation and image retrieval, namely bag-of-visual-words (BoVW) model. Therefore, it is essential to create an effective visual vocabulary. Most of existing approaches create visual vocabularies from image in pixel domain, which requires extra processing time in decompressed images, since most images are stored in compressed format. In this paper we propose to create a visual vocabulary based on Scale Invariant Feature Transform(SIFT) descriptor in compressed domain with the following three steps, (1) constructing low-resolution images in compressed domain, (2) extracting SIFT descriptor from low-resolution images, and (3) creating a visual vocabulary based on extracted SIFT descriptors. In order to evaluate the performance of the visual words, experiments have been conducted on identifying pornographic images. Experimental results indicate that the proposed method can recognize pornographic images accurately with much reduced computational time.
Lei Sui, Jing Zhang 0023, Li Zhuo 0001, Yuncong Yang
ISM3
2011 Cooperative multi-object tracking method for Wireless Video Sensor Networks
abstract
For the enormous number and the limited energy of network nodes in the wireless video sensor networks (WVSN) environment, to fulfil the complicated tasks, multiple sensor nodes should collaborate with each other. A cooperative multi-object tracking method for Wireless Video Sensor Networks is proposed in this paper. The proposed method is focused on the solution of cooperative multi-object tracking among multiple sensor nodes when an object leaves the view field of the tracking node. The main contributions of our proposed method are that: (1) the sensing model of a video sensor and Kalman filter is utilized to achieve optimal sensor selection. (2) Projective Invariants are employed to integrate information from the related nodes. The experimental results show that the proposed method is effective for resolving the problem of tracking relay.
Zheng Chu 0001, Li Zhuo 0001, Yingdi Zhao
MMSP2
2010 A Personalized Image Retrieval Based on User Interest Model
abstract
In order to narrow the semantic gap, user interest model plays an important role in personalized image retrieval. A novel personalized image retrieval approach based on user interest model is proposed in this study. User interest model is developed on the basis of short-tem and long-term interests. (1) Short-term interests are represented by collecting visual and semantic features. Visual features are collected by MARS relevance feedback. Semantic features are constructed by building a mapping from image low-level visual features to high-level semantic features on the basis of SVM. (2) Long-term interests are inferred by inference engine from the collected short-term interests. Long-term visual features are collected by the nonlinear gradual forgetting interest inference algorithm and semantic features are obtained by clustering algorithm. After applying to image retrieval, experimental results show that the average recall/precision is significantly improved and a better user satisfaction rate is achieved as well. Furthermore, it demonstrates our model can be efficiently adapted to user interests and matches personalized image retrieval.
Jing Zhang 0023, Li Zhuo 0001, Lansun Shen
Int. J. Pattern Recognit. Artif. Intell.2
2010 Complexity scalable control for H.264 motion estimation and mode decision under energy constraints
Xuejuan Gao, Kin-Man Lam 0001, Li Zhuo 0001, Lansun Shen
Signal Process.3
2008 A weight matrix based blind Super Resolution restoration algorithm
abstract
Approximate description of the observation model has severely prevented performance improvement of the existing Super Resolution(SR) algorithms. To solve this problem, a new weight-matrix based blind super resolution algorithm is proposed in this paper. A new observation model based on a motion compensate matrix and a weight-matrix is defined first. Then it is introduced into the traditional framework of Maximum a Posterior (MAP), where the problem of super resolution is converted to joint estimation of the weight matrix and the High Resolution (HR) image. Evaluated by both still and active image sequences, the proposed algorithm can model the degrading process of the acquired low resolution images accurately. It shows obvious performance improvement compared with the traditional MAP based super resolution algorithm. For several cases, it has even exceeded the results when the observation model is known.
Suyu Wang, Li Zhuo 0001, Lansun Shen
MMSP2
2008 A low complexity method for real-time gaze tracking
abstract
The paper addresses a novel, fast and low complexity gaze tracking method, which remedies computation complexity and real-time processing problems of most existing gaze tracking methods. Face images are captured and pre-processed under infrared lamp-houses with dual thresholds. Pupil area is verified by the run-length in eight directions, and the gaze point is located by mapping the vectors of pupil and reflection point or using neural network. The experimental results show that the real-time processing of gaze tracking, with high resolution and high accuracy simultaneously, is achieved in this new system.
Jing Zhang 0023, Mengkai Zhao, Li Zhuo 0001, Lansun Shen
MMSP3
2007 Optimization and Implementation of H.264 Encoder on DSP Platform
abstract
Compared with MPEG-4 and other previous standards, H.264 standard has achieved great breakthrough in coding performance. In this paper, the optimization and implementation of H.264 baseline profile encoder on the TMS320DM642 has been presented. Based on the architectural features of TMS320DM642 and the computational complexity analysis of H.264 encoder, the H.264 encoder has been optimized from three aspects: algorithms, data transfer and memory/Cache use. The experimental results demonstrate that, for the video sequences with CIF format, the optimized H.264 encoder can achieve the encoding speed of more than 24 frames per second, which can meet the real-time requirements ofthe applications.
Li Zhuo 0001, David Dagan Feng, Lansun Shen
ICME1
2007 Concept Constrained Image Region Annotation
abstract
Annotating image regions has been a challenging open issue in many areas such as image content understanding and image retrieval. In this paper, rather than solely rely on visual features of image regions, a novel approach is proposed to improve region annotation by taking concept constraints into account, since high level conceptual information such as image categories can increase the confidence of possible region labels as well as decrease the confidence of impossible region labels. We employ statistical models to learn the relationships among visual features, image concepts, and region labels. As a result, a set of possible region labels can be derived from a set of visual feature vectors of a given image so as to refine the annotation output obtained by using visual feature only. Promising experimental results have been demonstrated on 8462 regions of the University of Washington image dataset with diverse concepts for the proposed approach.
Zhiyong Wang 0001, Kelly Lam, Li Zhuo 0001, David Dagan Feng
MMSP3
2006 Channel-adaptive error protection for streaming stored MPEG-4 FGS over error-prone environments
abstract
In this paper, we investigate an adaptive channel error-protection scheme for streaming stored MPEG-4 fine granular scalability (FGS) bitstreams over error-prone environments. The human eye is sensitive to the variations in reconstructed image quality. Therefore, instead of using only the minimal average frame distortion as the optimization criterion, we propose a rate-distortion (R-D)-based bit-allocation method to determine the source coding rate and the channel coding rate under a given channel condition so as to minimize both the average image distortion and quality variation. Based on our proposed piecewise model of the MPEG-4 FGS enhancement layer, a balance between the average frame quality and quality variation will be obtained. Experimental results show that, compared with other existing methods, our proposed method can achieve less quality variation and comparable average frame distortion under various channel conditions.
Li Zhuo 0001, Kin-Man Lam 0001, Lansun Shen
IEEE Trans. Circuits Syst. Video Technol.1