Jing Zhang 0023

dblp:05/3499-23 · DBLP profile ↗
← Back
81ranked-venue papers
15as first author
44since 2021 · last 2026
0000-0003-1290-0738ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 31 · 8 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 11 since 2021Computer networks · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Multimodal driver behavior recognition based on frame-adaptive convolution and feature fusion
Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001
Comput. Vis. Image Underst.4
2026 TEFormer: Texture-Aware and Edge-Guided Transformer for Semantic Segmentation of Urban Remote Sensing Images
abstract
Accurate semantic segmentation of urban remote sensing images (URSIs) is essential for urban planning and environmental monitoring. However, it remains challenging due to the subtle texture differences and similar spatial structures among geospatial objects, which cause semantic ambiguity and misclassification. Additional complexities arise from irregular object shapes, blurred boundaries, and overlapping spatial distributions of objects, resulting in diverse and intricate edge morphologies. To address these issues, we propose TEFormer, a texture-aware and edge-guided Transformer. Our model features a texture-aware module (TaM) in the encoder to capture fine-grained texture distinctions between visually similar categories, thereby enhancing semantic discrimination. The decoder incorporates an edge-guided tri-branch decoder (Eg3Head) to preserve local edges and details while maintaining multiscale context-awareness. Finally, an edge-guided feature fusion module (EgFFM) effectively integrates contextual, detail, and edge information to achieve refined semantic segmentation. Extensive evaluation demonstrates that TEFormer yields mIoU scores of 88.57% on Potsdam and 81.46% on Vaihingen, exceeding the next best methods by 0.73% and 0.22%. On the LoveDA dataset, it secures the second position with an overall mIoU of 53.55%, trailing the optimal performance by a narrow margin of 0.19%.
Guoyu Zhou, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
IEEE Geosci. Remote. Sens. Lett.2
2026 Asymmetric simulation-enhanced flow reconstruction for incomplete multimodal learning
Jiacheng Yao, Jing Zhang 0023, Li Zhuo 0001
Pattern Recognit.2
2026 TSHR-Net: Text Semantic Homogenization Recognition Network for Short Video Title Overlays
abstract
Short videos follow the trend of creation, leading to a proliferation of homogenized video content. Textual overlays such as titles in short videos often reflect semantic homogeneity. This phenomenon manifests not only in the syntactic structure and overall thematic expression, but also in local semantic elements such as words and phrases. Based on information retrieval techniques, we propose a text semantic homogenization recognition network (TSHR-Net) for short video title overlays. The framework comprises key components: (1) a dynamic semantic representation that incorporates contextual information of title overlays using RoBERTa pre-trained word embeddings; (2) a dual-path semantic parser that integrates global semantics via BiLSTM-Attention and local semantics via multi-scale TextCNN; (3) a ranking loss optimization is designed to measure cosine similarity between semantic features, thereby improving homogenization recognition accuracy. Experimental results show that our TSHR-Net achieves the competitive performance in Chinese text semantic homogenization recognition, with ρ and ρ X,Y reaching 80.16% and 78.23% on LCQMC, 81.05% and 80.19% on STS-B(ZH), and 92.47% and 92.19% on our self-built BJUT-HCD. The model also exhibits generalization ability in English, attaining 79.68% and 79.89% on STS-B(EN), and 73.56% and 72.21% on SICK dataset, respectively.
Jing Zhang 0023, Shuying Zhang, Li Zhuo 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2026 LRGFormer: A Multiscale Feature Fusion Transformer for Image Restoration
abstract
Adverse weather conditions can significantly degrade image quality and impair the capture of critical information. Existing restoration networks struggle to effectively combine local, regional, and global features, thereby limiting their ability to handle diverse impacts of such weather. This study proposes the local-region-global transformer (LRGFormer), a transformer-based image restoration model for multiscale feature perception. The model comprises a basic module composed of multi-scale fusion attention (MSFSA) and a channel-spatial dual-attention feed-forward network (CSDF). Specifically, this study designs an MSFSA module. For the first time, it combines rotation-equivariant convolution with local attention for local information extraction and introduces a frequency-domain adaptive attention mechanism. By incorporating a query-aware global adaptive sparse attention mechanism for global information extraction, the network gradually fuses along the channel dimension, enabling progressive capture of spatial and frequency-domain information from the local and regional to global scale. Secondly, a CSDF network structure was designed to enhance channel-spatial interaction and improve the representational capacity of the model. By constructing a basic U-Net framework, the excellent basic modules for image restoration proposed in recent years are compared on a unified framework. Experimental results demonstrated that the proposed basic module can not only better extracts multi-scale features of images and restores image distortion caused by various degradation factors, and also exhibits good universality and generalization.
Jiafeng Li 0001, Wanying Hu, Tongyao Jia, Jiaqi Jin, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 BEV-CMHF: A Cross-Modality Hybrid Fusion Framework for BEV 3D Object Detection With Feature Interaction and Temporal Fusion
abstract
Autonomous driving technology has garnered significant attention for its potential to reduce driver burden and enhance road safety. Modern autonomous driving systems rely on a variety of sensors to perceive complex driving environments. Many existing methods map heterogeneous data into the bird’s eye view (BEV) space for feature fusion. However, they often fail to fully exploit the cross-modal interactions between cameras and LiDAR, or incorporate temporal information, resulting in suboptimal performance. Furthermore, commonly used fusion strategies are often overly simplistic. This study proposes BEV-CMHF, a cross-modality hybrid fusion framework for BEV 3D object detection with feature interaction and temporal fusion. By introducing an interactive cross-attention module and a long-short-term temporal module, the proposed framework enhances the representational power of fused BEV features. Specifically, a feature-interaction attention module that facilitates effective interaction between the camera and LiDAR BEV features using deformable attention is designed, providing guidance and supervision for the camera BEV features. Subsequently, a historical feature temporal fusion module that integrates the long-short-term temporal module is introduced to incorporate additional critical temporal information into the BEV features. Moreover, a dynamic hybrid feature-fusion module is designed to fuse the BEV features of the camera and LiDAR effectively through a hybrid attention mechanism that combines coarse and fine attention. Extensive experiments conducted on the nuScenes benchmark validate the effectiveness of the proposed method, achieving 70.87% mAP and 74.00% NDS on the test set. Using a single NVIDIA GeForce RTX 4090, the method attained an inference speed of 5.79 images per second (5.79 img/s), corresponding to an inference time of 172.64 ms on the nuScenes dataset. The source code will be released athttps://github.com/BJUTsipl/BEV-CMHF
Jiafeng Li 0001, Jinquan Xu, Mengxun Zhi, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Intell. Transp. Syst.4
2026 HDMDN: Hierarchical-Decoupling Based Meta-Knowledge Single Image Dehazing Network
abstract
Images suffer from color shift and detail distortion owing to limitations of contrast and visibility in hazy scenes, affecting their subjective perception. However, the performance of existing algorithms on real-world hazy images remains limited as scenes can be complex and haze degradation varies in outdoor visual systems. This study proposes a meta-knowledge single image dehazing algorithm based on hierarchical decoupling, combining the advantages of convolutional neural networks (CNNs) and transformers. We propose a novel dual-branch decoupling network that decouples low-level features from high level semantic information in images, leveraging the hierarchical properties of the network. It combines a CNN and cross dual branch transformer network (CDual transformer) in the encoder network to fully extract local and global features of images. To disentangle high-level semantic features, a style transfer module is designed to transform the style of hazy images while retaining the remaining semantic information. Afterward, the low-level features of images and the transformed high-level semantic features are used to reconstruct the dehazed images, fully capitalizing on the multilevel features. Furthermore, we built a meta-semi-supervised training strategy to improve the decoupling performance of the model and accumulated style knowledge of clear images from both synthetic and real-world hazy data, improving the generalizability of the model. Extensive experiments on both synthetic and real datasets show that the proposed algorithm effectively removes haze and offers better generalization abilities than similar methods. The project is publicly available at https://github.com/BJUTsipl/HDMDN.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu
IEEE Trans. Multim.4
2026 RcFormer: Reconfigurable Self-Attention Transformer for Image Restoration
abstract
Adverse weather and imaging environments may degrade image quality and pose a significant challenge to the visual perception systems of multimedia. Various image restoration tasks necessitate the modeling of multiscale features, which is highly demanding on networks. To date, Vision Transformer has exhibited impressive image restoration performance. However, in this model, global self-attention is computationally expensive and local self-attention typically limits the interaction domain of each token. To solve this problem, we propose a novel reconfigurable self-attention transformer called RcFormer, which is designed to adequately model multiscale image features. This is achieved through a cross-grouped transformer (CGTransformer) block that uses convolution, area self-attention, and row-column self-attention for different head groups. CGTransformer is combined with an intragroup operation interaction structure. Moreover, an intergroup reconfigurable mechanism is implemented based on CGTransformer and channel circulation. The combination of multiple operations effectively enhances the modeling capability in the spatial and channel dimensions for various image recovery tasks. The performance of the proposed RcFormer is compared with low-level vision modules in a unified framework. Extensive experiments demonstrated that RcFormer exhibited a superior performance for the following image restoration tasks: image dehazing, rain streak removal, raindrop removal, snow removal, and single image deblurring. The source code is publicly available athttps://github.com/dehazing/RcFormer.
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu
IEEE Trans. Multim.4
2026 HMS2Net: Heterogeneous Multimodal State Space Network via CLIP for Dynamic Scene Classification in Livestreaming
abstract
Livestreaming platforms attract countless daily active users, making online content regulation imperative. The complex and diverse multimodal content elements in dynamic livestreaming scene pose a great challenge to video content understanding. Thanks to the success of contrastive language-image pre-training (CLIP) for dynamic scene classification, which is one of the basic tasks of video content understanding. We propose a heterogeneous multimodal state space network (HMS2Net) for dynamic scene classification in livestreaming via CLIP. (1) To fully and efficiently mine the dynamic scene elements in livestreaming, we design a heterogeneous teacher-student Transformer (HT-SFormer) with CLIP to extract multimodal features in an energy-efficient unified pipeline; (2) To cope with the possible information conflicts in heterogeneous feature fusion, we introduce a cross-modal adaptive feature filter and fusion (CMAF) module to generate more complete information complementarity by adjusting multimodal feature composition; (3) For temporal context-awareness of dynamic scene, we establish a dynamic state space memory (DSSM) structure for capturing the correlation of multimodal data between neighboring video frames. A series of comparative experiments are conducted on the publicly available datasets DAVIS, Mini-kinetics, HMDB51, and the self-built BJUT-LCD. Our HMS2Net produce competitive results of 71.09%, 95.40%, 53.64%, and 82.36%, respectively, demonstrating the effectiveness and superiority of dynamic scene classification in livestreaming.
Jing Zhang 0023, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Multim.2
2026 ToGCN-LLM: Tri-optimization Graph Convolutional Network with LLM for Group Activity Recognition
abstract
Graph networks face substantial challenges in handing large-scale graph-structured data. Although graph convolutional networks (GCNs) have been widely applied to group activity recognition, they still struggle with high graph structure complexity and semantic gap, especially in complex scenarios. To address these limitations, we harness large language models (LLMs) and multimodal learning to enhance semantic understanding as well as bridge the gap between graph structures and contextual meanings. For the simultaneous optimization of model efficiency and recognition accuracy in group activities, we propose a tri-optimization graph convolutional network with LLM (ToGCN-LLM) from the perspective of graph structure knowledge distillation. (1) To tackle high complexity of teacher network in GCN, we adopt a sparsification strategy to prune irrelevant edges and nodes, reducing computational overhead and enhancing training efficiency. (2) To mitigate information redundancy during knowledge distillation, we design a hierarchical optimization module combined with a hierarchical sampling mechanism, exploiting graph hierarchical structure and adjacency relationships to improve knowledge transfer efficiency. (3) Considering the student network’s varying learning performance across training stages and the limitations of fixed learning strategies, we introduce a dynamic adaptive weight decay mechanism to achieve fine-grained convergence under different gradient updates, thereby boosting overall recognition accuracy. (4) We use the Qwen LLM to extract text description tokens, which are fused with the student GCN’s last-layer features to enable multi-model, multi-optimization learning for group activity recognition. Six experiments on CAD, CAED, and BJUT-GAD dataset demonstrate that our ToGCN-LLM achieves competitive MPCA scores of 94.89%, 93.61%, and 95.77%, respectively.
Junpeng Kang, Jing Zhang 0023, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2026 KdM-Net: Knowledge-Driven Memory Network for Long-Tailed Human-Object Interaction in Livestreaming
abstract
With the continuous advancement of the livestreaming industry, streamers as content producers pose significant challenges to the timeliness of regulatory response mechanisms, emerging as a critical weak link in cyberspace governance. Human-object interaction (HOI) detection plays a pivotal role in understanding multimodal livestreaming videos. In mainstream healthy online ecosystems, normal HOI categories dominate, while rare ones are extremely scarce, i.e., long-tail distribution that hinders HOI models from effectively detecting streamer violations. Driven by the transformative potential of foundation models (FMs) in multimodal video understanding, we propose a knowledge-driven memory network (KdM-Net) for long-tailed HOI in livestreaming, leveraging the extensive capabilities of the contrastive language-image pretraining (CLIP) model. First, human-object (HO) pairs are generated and modeled using general object detector/tracker. After converting each HOI label into a short sentence description, text embeddings are extracted via the CLIP text encoder to initialize classifier weights. Notably, we introduce a visual-textual knowledge transfer strategy to align visual and text features, complementing for the multimodal knowledge deficit of rare categories that plague long-tailed HOI distributions. Finally, a knowledge-driven memory module is designed to dynamically assign adaptive weights and attention to interaction categories based on their long-tailed distribution characteristics, mitigating model forgetting tail data and enhancing HOI detection performance. Experimental results demonstrate that KdM-Net achieves HOI detection accuracies of 37.33@full, 50.63%@non-rare, and 27.14%@rare on the publicly available VidHOI dataset, and 45.41%@full, 61.77%@non-rare, and 31.95%@rare on the self-built BJUT-HOI dataset. These fundings validate the generalization power of our KdM-Net for HOI detection and its competitiveness in livestreaming scenarios.
Menghui Zhang, Jing Zhang 0023, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 From Echo Distillation to Chat-Driven: A Multi-Task VLM is OK for Object Detection and Fine-Grained Classification in Remote Sensing Images
abstract
Vision-language models (VLMs) have demonstrated attracting performance in natural image domains, and their application to multitask remote sensing analysis and interpretation holds great potential. However, realizing this potential in remote sensing images (RSIs) is hindered by limited feature extraction in lightweight detectors, loss of fine-grained details in classification, and the lack of unified multi-task frameworks. In response, we propose a multi-task vision-language model (MT-VLM) to address these limitations, offering an effective solution for both object detection and fine-grained classification in RSIs. Specifically, we introduce EchoDistill, an echo distillation framework that combines teacher knowledge transfer and self-knowledge reinforcement to iteratively enhance feature learning, significantly boosting the detection performance of the lightweight models. Additionally, we develop a chat-driven VLM with LoRA fine-tuning, enabling the model to capture fine-grained details and improve classification accuracy using multimodal embeddings. To support our framework, we collected BJUT-FGC, a comprehensive image-text labeled dataset for fine-grained classification in RSIs, and built a dual-mode multi-task benchmark to facilitate both task-specific adaptation and integrated multi-task refinement. Experimental results on DOTA 2.0, VEDAI, and self-built BJUT-FGC datasets demonstrate the effectiveness of our approach, achieving 58.71% and 82.06% mAP on the object detection task and 88.08% top-1 accuracy on the fine-grained classification task.
Liuqian Wang, Jing Zhang 0023, Guangming Mi, Li Zhuo 0001
IJCNN2
2025 RealText: Realistic Text Image Generation based on Glyph and Scene Aware Inpainting
abstract
Text-to-image generation models can create diverse, high-quality images, but they frequently encounter challenges in accurately rendering text within those images due to the insufficient representation of desired text. In this study, we introduce RealText, a method for generating scene text images that excels in producing precise and realistic scene text images in any language. We disentangle scene text images generation into three stages: background and glyph image generation, text deformation, and whole image generation. Initially, we utilize prompts to guide the creation of well-organized background images. By identifying optimal text placements on these backgrounds, we render the glyph images of target text using user-specified font, effectively eliminating incorrect characters. In the next stage, we propose scene sensing to perceive text carrier surfaces and viewpoints through 3D scene reconstruction using depth and normal map to apply text deformation, thereby enhancing the realism of generated images. The final stage involves generating complete image with the aid of background and glyph guidance. Thanks to glyph disentangling, scene sensing, and text inpainting, we can exert more precise control over scene text image generation process. We have developed a unified framework which supports major generation models. Extensive experiments illustrate the exceptional performance of our method in generating images with multilingual text. The codes will soon be available at https://github.com/cccvl/RealText.
Zihou Liu, Dongming Zhang 0004, Jing Zhang 0023, Yongdong Zhang 0001
ACM Multimedia3
2025 Prototype Embedding Optimization for Human-Object Interaction Detection in Livestreaming : PeO-HOI
abstract
Livestreaming frequently features interactions between streamers and objects, which is critical for understanding and regulating online content. While human-object interaction (HOI) detection has advanced significantly for general video tasks, its application to recognizing streamer-object interactions in livestreaming often exhibits object bias: an excessive focuses on the objects at the expense of interactions with the streamer. To address this challenge, we propose a prototype embedding optimization for human-object interaction detection (PeO-HOI). Our approach first preprocesses the livestreaming using object detection and tracking to extract features of human-object (HO) pairs. Subsequently, prototype embedding optimization is applied to mitigate object bias effects. Finally, after modeling the spatio-temporal context among HO pairs, the HOI detection results are generated by the prediction head. Experimental results demonstrate that the PeO-HOI achieves detection accuracies of 37.19%@full, 51.42%@non-rare, and 26.20%@rare on the publicly available VidHOI dataset, and 45.13%@full, 62.78%@non-rare, and 30.37%@rare on our self-built BJUT-HOI dataset. These results confirm that PeO-HOI effectively enhances HOI detection performance in livestreaming scenarios.
Menghui Zhang, Jing Zhang 0023, Li Zhuo 0001
MMSP2
2025 RWGCN: Random walk graph convolutional network for group activity recognition
Junpeng Kang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Appl. Intell.2
2025 Explainable graph convolutional network based on catastrophe theory and its application to group activity recognition
Junpeng Kang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Eng. Appl. Artif. Intell.2
2025 DR-YOLO: dual reconstructed YOLO for logo detection in livestreaming
Chenyu Yuan, Jing Zhang 0023, Li Zhuo 0001
Multim. Syst.2
2025 Cellular spatial-semantic embedding for multi-label classification of cell clusters in thyroid fine needle aspiration biopsy whole slide images
Juntao Gao, Jing Zhang 0023, Li Zhuo 0001
Pattern Recognit. Lett.2
2025 Position Guided Dynamic Receptive Field Network: A Small Object Detection Friendly to Optical and SAR Images
abstract
Object detection in remote sensing images (RSIs), including optical and SAR images, has emerged as a rapidly advancing field. However, the abundance of small objects in RSIs poses a significant challenge in designing a network structure with effective receptive fields to support accurate localization and classification. In this paper, we propose a position guided dynamic receptive field network (PG-DRFNet) for small object detection friendly to optical and SAR images. Specifically, PG-DRFNet overcomes the problem of small objects vanishing or being submerged in features by establishing a positional guidance relationship of small objects between different feature layers. Then, we design a combination head structure that utilizes additional supervised information extracted from small objects to make the model more effective and flexible. Moreover, a dynamic perception algorithm based on feature construction is developed to dynamically optimize the perception regions and feature hierarchies of the model, while seeking the optimal tradeoff between model accuracy and inference speed. Without bells and whistles, our model is robust to two modalities of remote sensing data, and our experiments are conducted on four benchmark RSI datasets, including DOTA-v2.0, VEDAI, SSDD, and HRSID. The experimental results achieve competitive performance with 59.01%, 84.06%, 90.06%, and 80.59% mAP, respectively. Code and models are released athttps://github.com/BJUT-AIVBD/PG-DRFNet.
Liuqian Wang, Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 TSTrack: A Lightweight Transformer-Based Spatiotemporal Feature Refinement Tracking Algorithm
abstract
Single-object tracking is a fundamental enabling technology in the field of remote sensing observation. It plays a crucial role in tasks such as unmanned aerial vehicle route surveillance and maritime vessel trajectory prediction. However, because of challenges such as the weak discriminative power of target features, interference from complex environments, and frequent viewpoint changes, existing trackers often suffer from insufficient temporal modeling capabilities and low computational efficiency, which limit their practical deployment. To address these challenges, we propose TSTrack, a novel lightweight single-object tracking framework that integrates Transformer and Mamba-based spatiotemporal modeling. First, we propose the target-aware feature purification preprocessor (TAFPP) , designed to dynamically enhance target representation through a synergistic combination of the dynamic position acuity module (DPAM) and spectral channel recalibrator (SCR). Second, we introduce the recurrent Mamba interaction pyramid (RM-IP) to replace traditional recurrent neural network-based structures, leveraging a state-space model for efficient and expressive temporal modeling with significantly reduced parameter overhead. Finally, we propose the elastic reconstructive multi-scale fusion (ERMSF) module, which adopts a four-branch parallel architecture to achieve effective multiscale feature fusion and dynamic shape adaptation, thereby enhancing robustness against target deformations and scale variations. Extensive experiments conducted on benchmark datasets, including LaSOT, TrackingNet, and GOT-10k, demonstrate the effectiveness of TSTrack. The results show that TSTrack achieves a superior tracking accuracy while maintaining a lightweight design, significantly outperforming existing state-of-the-art methods. The source code is publicly available at https://github.com/BJUTsipl/TSTrack.
Jiafeng Li 0001, Shengyao Sun, Yang Wang 0023, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Recurrent Semantic Change Detection in VHR Remote Sensing Images Using Visual Foundation Models
abstract
Semantic change detection (SCD) involves the simultaneous extraction of changed regions and their corresponding semantic classifications (pre- and post-change) in remote sensing images (RSIs). Despite recent advancements in vision foundation models (VFMs), the fast-segment anything model has demonstrated insufficient performance in SCD. In this article, we propose a novel VFMs architecture for SCD, designated as VFM-ReSCD. This architecture integrates a side adapter (SA) into the VFM-ReSCD to fine-tune the fast segment anything model (FastSAM) network, enabling zero-shot transfer to novel image distributions and tasks. This enhancement facilitates the extraction of spatial features from very high-resolution (VHR) RSIs. Moreover, we introduce a recurrent neural network (RNN) to model semantic correlation and capture feature changes. We evaluated the proposed methodology on two benchmark datasets. Extensive experiments show that our method achieves state-of-the-art (SOTA) performances over existing approaches and outperforms other CNN-based methods on two RSI datasets.
Jing Zhang 0023, Lei Ding 0008, Tingyuan Zhou, Jian Wang 0138, Peter M. Atkinson, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2025 CycFormer: Unsupervised Rain Removal Network Based on CycleGAN and Transformer
abstract
Rainy weather presents significant challenges for applications relying on visual perception in intelligent transportation systems. The scarcity of real paired training data complicates single-image rain removal tasks, prompting an increasing interest in unsupervised methods capable of handling real-world rainy images without paired data. At present, most unsupervised rain removal methods are based on the CycleGAN framework; however, the combination of this framework and transformer is not satisfactory owing to most Transformers’ insufficient ability to model real rain features with global inhomogeneous distributions, which prevents them from being fully applicable to unsupervised tasks. This study devised an unsupervised rain removal network based on CycleGAN and the DerainFormer transformer. First, a deformable sparse attention mechanism was developed to improve the Transformer’s suitability for unsupervised tasks in CycleGAN architectures. Subsequently, a two-stage alternating transformer structure was designed to enhance its global non-uniform modeling capabilities for real rain images, In addition, a dual-channel parallel feed-forward network was used to establish the correlation between multiscale rain stripes. Finally, since rain removal is considered a decomposition task, a rain layer unsupervised training method for joint positional contrastive learning was proposed to separate the rain streaks effectively. We conducted several experiments on different real and synthetic rain datasets and the results confirmed that our unsupervised rain removal method performed well. The source code will be released athttps://github.com/derainsipl/CycFormer.
Jiafeng Li 0001, Shuhao Yan, Jing Zhang 0023, Li Zhuo 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Cross-Modal Tri-Semantic Correlation-CLIP for Short Video Homogenization Recognition
abstract
Short videos are one of the most popular social media in the world, triggering a proliferation of copycat creations leading to homogenized video content, with visual and textual homogenization being the most prevalent. Unlike near-duplicate video retrieval, which relies on visual appearance similarity, homogenization recognition emphasizes identifying videos with similar semantic units. Short videos exhibit multimodal features, in which there is a many-to-many mapping relationship between visual and text elements, and the two modalities are relatively independent and semantically correlated. Therefore, cross-modal semantic correlation needs to be explored and established to achieve homogenization recognition of short videos. Based on the idea of divide-and-conquer and joint processing, we propose a cross-modal tri-semantic correlation-CLIP (CS 3 C-CLIP) for short video homogenization recognition. First, visual and text features in the shared subspace are extracted using the contrastive language-image pre-training visual-text dual encoder. Then, features at the patch, frame, and video levels are generated using the patch selection module and the temporal encoder, while the word-level and sentence-level features are respectively derived from text features and [EOS] token. After establishing cross-modal tri-semantic correlations by constructing a triple semantic (i.e., video-sentence, frame-sentence, and patch-word) correlation, homogenized short videos are recognized by measuring the aggregated cross-modal similarity between pairs of short videos. Experimental results on three publicly available datasets demonstrate that our CS 3 C-CLIP outperforms state-of-the-art methods, achieving 85.7% R@1 and 94.4% R@5 on self-built BJUT-HCD, 49.4% R@1 and 74.6% R@5 on MSR-VTT, and 49.8% R@1 and 78.1% R@5 on MSVD, respectively.
Jiacheng Yao, Jing Zhang 0023, Shuying Zhang, Li Zhuo 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Unpaved road segmentation of UAV imagery via a global vision transformer with dilated cross window self-attention for dynamic map
Jing Zhang 0023, Jiafeng Li 0001, Li Zhuo 0001
Vis. Comput.2
2024 MKP-Net: Memory knowledge propagation network for point-supervised temporal action localization in livestreaming
Jing Zhang 0023, Yian Zhang, Junpeng Kang, Li Zhuo 0001
Comput. Vis. Image Underst.2
2024 LCMA-Net: A light cross-modal attention network for streamer re-identification in live video
Jiacheng Yao, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Comput. Vis. Image Underst.2
2024 HDUD-Net: heterogeneous decoupling unsupervised dehaze network
Jiafeng Li 0001, Lingyan Kuang, Jiaqi Jin, Li Zhuo 0001, Jing Zhang 0023
Neural Comput. Appl.5
2024 RaSTFormer: region-aware spatiotemporal transformer for visual homogenization recognition in short videos
Shuying Zhang, Jing Zhang 0023, Hui Zhang 0049, Li Zhuo 0001
Neural Comput. Appl.2
2024 Self-guided disentangled representation learning for single image dehazing
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023
Neural Networks4
2024 Joint Spatio-Temporal Modeling for Semantic Change Detection in Remote Sensing Images
abstract
Semantic Change Detection (SCD) refers to the task of simultaneously extracting the changed areas and the semantic categories (before and after the changes) in Remote Sensing Images (RSIs). This is more meaningful than Binary Change Detection (BCD) since it enables detailed change analysis in the observed areas. Previous works established triple-branch Convolutional Neural Network (CNN) architectures as the paradigm for SCD. However, it remains challenging to exploit semantic information with a limited amount of change samples. In this work, we investigate to jointly consider the spatio-temporal dependencies to improve the accuracy of SCD. First, we propose a Semantic Change Transformer (SCanFormer) to explicitly model the ’from-to’ semantic transitions between the bi-temporal RSIs. Then, we introduce a semantic learning scheme to leverage the spatio-temporal constraints, which are coherent to the SCD task, to guide the learning of semantic changes. The resulting network (SCanNet) significantly outperforms the baseline method in terms of both detection of critical semantic changes and semantic consistency in the obtained bi-temporal results. It achieves the SOTA accuracy on two benchmark datasets for the SCD.
Lei Ding 0008, Jing Zhang 0023, Haitao Guo, Kai Zhang 0010, Bing Liu 0018, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.2
2023 Bi-Directional Temporal Modelling for Semantic Change Detection in Remote Sensing Images
abstract
Semantic change detection (SCD) is a branch of change detection (CD) that provides detailed land-cover/land-use (LCLU) change information. It presents not only the changed information but also the bi-temporal LCLU semantic maps (in the changed areas). Studies have recently highlighted [1] that SCD can be addressed through a triple-branch Convolutional Neural Network (CNN), which contains two multi-temporal branches and a change detection branch. However, in this architecture, the two temporal branches learn insufficient LCLU transition information. In this paper, we present a novel architecture that combines CNN and RNN for the SCD of remote sensing images. It employs a Siamese CNN to learn semantic information from two temporal images, followed by a Bidirectional Recurrent Neural Network (Bi-directional RNN) to learn temporal dependencies of the LCLU classes. The resulting CNN-RNN architecture can model better the LCLU transitions, thus enhancing the semantic representation of the bi-temporal features. Experimental results on a benchmark dataset show that the proposed method obtains significant accuracy improvements over the existing approaches. It also shows advantages in recognizing the minority LCLU changes.
Jing Zhang 0023, Lei Ding 0008, Lorenzo Bruzzone
IGARSS1
2023 Occluded prohibited object detection in X-ray images with global Context-aware Multi-Scale feature Aggregation
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
Neurocomputing5
2023 Short video fingerprint extraction: from audio-visual fingerprint fusion to multi-index hashing
Shuying Zhang, Jing Zhang 0023, Li Zhuo 0001
Multim. Syst.2
2023 Efficient Fine-Grained Object Recognition in High-Resolution Remote Sensing Images From Knowledge Distillation to Filter Grafting
abstract
With the development of high-resolution remote sensing images (HR-RSIs) and the escalating demand for intelligent analysis, fine-grained recognition of geospatial objects has become a more practical and challenging task. Although deep learning-based object recognition has achieved superior performance, it is inflexible to be directly utilized to the fine-grained object recognition tasks of HR-RSIs under the limitation of the size of geospatial objects. An efficient fine-grained object recognition method in HR-RSIs from knowledge distillation to filter grafting is proposed. Specifically, fine-grained object recognition consists of two stages: Stage 1 utilizes oriented region convolutional neural network (oriented R-CNN) to accurately locate and preliminarily classify geospatial objects. At the same time, it serves as a teacher network to guide students’ effective learning of fine-grained object recognition; in Stage 2, we design a coarse-to-fine object recognition network (CF-ORNet), as the second teacher network, which realizes fine-grained recognition through feature learning and category correction. After that, we propose a lightweight model from knowledge distillation to filter grafting on two teacher networks to achieve efficient fine-grained object recognition. The experimental results on VEDAI and HRSC2016 datasets achieve competitive performance.
Liuqian Wang, Jing Zhang 0023, Jimiao Tian, Jiafeng Li 0001, Li Zhuo 0001, Qi Tian 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 BARRN: A Blind Image Compression Artifact Reduction Network for Industrial IoT Systems
abstract
Most industrial Internet of Things (IoT) devices reduce the capture image size using high-ratio joint photographic experts group (JPEG) compression, saving storage space, and transmission bandwidth consumption. However, the resulting compression artifacts considerably affect the accuracy of subsequent tasks. Most artifact reduction algorithms do not consider the limitations of storage space and computing power of edge devices. In this study, a blind artifact reduction recurrent network (BARRN), which can reduce compression artifacts when the quality factors are unknown, is proposed. First, a structure based on recurrent convolution is designed for the specific requirements of industrial IoT image acquisition devices; the network can be scaled according to system resource constraints. Second, a more efficient convolution group, capable of adaptively processing different degradation levels, is proposed for optimal use of the limited computational resources. The experimental results demonstrate that the proposed BARRN can meet the needs of industrial systems with high computational efficiency.
Jiafeng Li 0001, Yuqi Gao, Li Zhuo 0001, Jing Zhang 0023
IEEE Trans. Ind. Informatics5
2023 Cascade Transformer Decoder Based Occluded Pedestrian Detection With Dynamic Deformable Convolution and Gaussian Projection Channel Attention Mechanism
abstract
Occluded pedestrian detection is very challenging in computer vision, because the pedestrians are frequently occluded by various obstacles or persons, especially in crowded scenarios. In this article, an occluded pedestrian detection method is proposed under a basic DEtection TRansformer (DETR) framework. Firstly, Dynamic Deformable Convolution (DyDC) and Gaussian Projection Channel Attention (GPCA) mechanism are proposed and embedded into the low layer and high layer of ResNet50 respectively, to improve the representation capability of features. Secondly, Cascade Transformer Decoder (CTD) is proposed, which aims to generate high-score queries, avoiding the influence of low-score queries in the decoder stage, further improving the detection accuracy. The proposed method is verified on three challenging datasets, namely CrowdHuman, WiderPerson, and TJU-DHD-pedestrian. The experimental results show that, compared with the state-of-the-art methods, it can obtain a superior detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IEEE Trans. Multim.5
2022 Prohibited Object Detection in X-ray Images with Dynamic Deformable Convolution and Adaptive IoU
abstract
Due to the variety and complexity of objects in X-ray images, how to detect the prohibited items automatically and accurately is a challenging problem. In this paper, an X-ray image prohibited object detection method based on Dynamic Deformable Convolution (DyDC) and adaptive Intersection over Union (IoU) is proposed based on Cascade R-CNN framework. The main contributions are as follows. First, DyDC is proposed to cope with the diversity of the prohibited objects in X-ray images and to improve the feature representation capability. Then, adaptive IoU mechanism is proposed, which can dynamically adjust the IoU threshold during the training process to generate high quality proposals. The proposed method is extensively evaluated on two publicly available benchmark datasets, namely SIXray and OPIXray, and the experimental results show that it can achieve the state-of-the-art detection accuracy, compared with other existing methods.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
ICIP5
2022 EAOD-Net: Effective anomaly object detection networks for X-ray images
abstract
Abstract Anomaly object detection is the core technology in the application for X‐ray images. However, the accuracy of current X‐ray anomaly object detection method still needs to be improved. In this paper, an effective anomaly object detection network is proposed to improve the detection accuracy of anomaly object for X‐ray images. Firstly, learnable Gabor convolution layer, deformable convolution, and spatial attention mechanism are introduced to enhance the representative ability of features in ResNeXt. Then, dense local regression is applied to predict the offset of multiple dense boxes in region proposal to locate the object accurately. At last, bigger discriminative RoI pooling is proposed to classify the candidate boxes to improve the accuracy of object classification. Experimental results on the SIXray and OPIXray datasets show that compared with the state‐of‐the‐art methods, the proposed EAOD‐Net can achieve the competitive detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IET Image Process.5
2022 Meta-Learning Paradigm and CosAttn for Streamer Action Recognition in Live Video
abstract
As an emerging field of network content production, live video has been in the vacuum zone of cyberspace governance for a long time. Streamer action recognition is conducive to the supervision of live video content. In view of the diversity and imbalance of streamer actions, it is attractive to introduce few-shot learning to realize streamer action recognition. Therefore, a meta-learning paradigm and CosAttn for streamer action recognition method in live video is proposed, including: (1) the training set samples similar to the streamer action to be recognized are pretrained to improve the backbone network; (2) video-level features are extracted by R(2+1)D-18 backbone and global average pooling in the meta-learning paradigm; (3) the streamer action is recognized by calculating cosine similarity after sending the video-level features to CosAttn to generate a streamer action category prototype. Experimental results on several real-world action recognition datasets demonstrate the effectiveness of our method.
Jing Zhang 0023, Jiacheng Yao, Li Zhuo 0001, Qi Tian 0001
IEEE Signal Process. Lett.2
2022 Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images
abstract
Semantic change detection (SCD) extends the multiclass change detection (MCD) task to provide not only the change locations but also the detailed land-cover/land-use (LCLU) categories before and after the observation intervals. This fine-grained semantic change information is very useful in many applications. Recent studies indicate that the SCD can be modeled through a triple-branch convolutional neural network (CNN), which contains two temporal branches and a change branch. However, in this architecture, the communications between the temporal branches and the change branch are insufficient. To overcome the limitations in existing methods, we propose a novel CNN architecture for the SCD, where the semantic temporal features are merged in a deep CD unit. Furthermore, we elaborate on this architecture to reason the bi-temporal semantic correlations. The resulting bi-temporal semantic reasoning network (Bi-SRNet) contains two types of semantic reasoning blocks to reason both single-temporal and cross-temporal semantic correlations, as well as a novel loss function to improve the semantic consistency of change detection results. Experimental results on a benchmark dataset show that the proposed architecture obtains significant accuracy improvements over the existing approaches, while the added designs in the Bi-SRNet further improve the segmentation of both semantic categories and the changed areas. The codes in this article are accessible athttps://github.com/ggsDing/Bi-SRNet.
Lei Ding 0008, Haitao Guo, Sicong Liu 0001, Lichao Mou, Jing Zhang 0023, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.5
2022 Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Long-range contextual information is crucial for the semantic segmentation of high-resolution (HR) remote sensing images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the perception of long-range contexts in large RSIs. To overcome this limitation, we propose a wide-context network (WiCoNet) for the semantic segmentation of HR RSIs. Apart from extracting local features with a conventional convolutional neural network (CNN), the WiCoNet has an extra context branch to aggregate information from a larger image area. Moreover, we introduce a context transformer to embed contextual information from the context branch and selectively project it onto the local features. The context transformer extends the vision transformer, an emerging kind of neural networks, to model the dual-branch semantic correlations. It overcomes the locality limitation of CNNs and enables the WiCoNet to see the bigger picture before segmenting the land-cover/land-use (LCLU) classes. Ablation studies and comparative experiments conducted on several benchmark datasets demonstrate the effectiveness of the proposed method. In addition, we present a new Beijing Land-Use (BLU) dataset. This is a large-scale HR satellite dataset with high-quality and fine-grained reference labels, which can facilitate future studies in this field.
Lei Ding 0008, Dong Lin, Shaofu Lin, Jing Zhang 0023, Xiaojie Cui, Yuebin Wang, Hao Tang 0005, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2021 Crowd activity recognition in live video streaming via 3D-ResNet and region graph convolution network
abstract
Abstract Since the era of we‐media, live video industry has shown an explosive growth trend. For large‐scale live video streaming, especially those containing crowd events that may cause great social impact, how to identify and supervise the crowd activity in live video streaming effectively is of great value to push the healthy development of live video industry. The existing crowd activity recognition mainly uses visual information, rarely fully exploiting and utilizing the correlation or external knowledge between crowd content. Therefore, a crowd activity recognition method in live video streaming is proposed by 3D‐ResNet and regional graph convolution network (ReGCN). (1) After extracting deep spatiotemporal features from live video streaming with 3D‐ResNet, the region proposals are generated by region proposal network. (2) A weakly supervised ReGCN is constructed by making region proposals as graph nodes and their correlations as edges. (3) Crowd activity in live video streaming is recognised by combining the output of ReGCN, the deep spatiotemporal features and the crowd motion intensity as external knowledge. Four experiments are conducted on the public collective activity extended dataset and a real‐world dataset BJUT‐CAD. The competitive results demonstrate that our method can effectively recognise crowd activity in live video streaming.
Junpeng Kang, Jing Zhang 0023, Li Zhuo 0001
IET Image Process.2
2021 Streamer action recognition in live video with spatial-temporal attention and deep dictionary learning
Jing Zhang 0023, Jiacheng Yao
Neurocomputing2
2021 Blind image quality assessment with channel attention based deep residual network and extended LargeVis dimensionality reduction
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Meng Wang 0001
J. Vis. Commun. Image Represent.4
2020 Coarse-to-fine object detection in unmanned aerial vehicle imagery using lightweight convolutional neural network and deep motion saliency
Jing Zhang 0023, Xi Liang 0003, Meng Wang 0001, Liheng Yang, Li Zhuo 0001
Neurocomputing1
2020 Multi-level prediction Siamese network for real-time UAV visual tracking
Hui Zhang 0049, Jing Zhang 0023, Li Zhuo 0001
Image Vis. Comput.3
2020 Multilevel fusion of multimodal deep features for porn streamer recognition in live video
Jing Zhang 0023, Meng Wang 0001, Jimiao Tian, Li Zhuo 0001
Pattern Recognit. Lett.2
2020 Small Object Detection in Unmanned Aerial Vehicle Images Using Feature Fusion and Scaling-Based Single Shot Detector With Spatial Context Analysis
abstract
Objects in unmanned aerial vehicle (UAV) images are generally small due to the high-photography altitude. Although many efforts have been made in object detection, how to accurately and quickly detect small objects is still one of the remaining open challenges. In this paper, we propose a feature fusion and scaling-based single shot detector (FS-SSD) for small object detection in the UAV images. The FS-SSD is an enhancement based on FSSD, a variety of the original single shot multibox detector (SSD). We add an extra scaling branch of the deconvolution module with an average pooling operation to form a feature pyramid. The original feature fusion branch is adjusted to be better suited to the small object detection task. The two feature pyramids generated by the deconvolution module and feature fusion module are utilized to make predictions together. In addition to the deep features learned by the FS-SSD, to further improve the detection accuracy, spatial context analysis is proposed to incorporate the object spatial relationships into object redetection. The interclass and intraclass distances between different object instances are computed as a spatial context, which proves effective for multiclass small object detection. Six experiments are conducted on the PASCAL VOC dataset and the two UAV image datasets. The experimental results demonstrate that the proposed method can achieve a comparable detection speed but an accuracy superior to those of the six state-of-the-art methods.
Xi Liang 0003, Jing Zhang 0023, Li Zhuo 0001, Yuzhao Li, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Porn Streamer Recognition in Live Video Streaming via Attention-Gated Multimodal Deep Features
abstract
Live video streaming platforms have attracted millions of streamers and daily active users. For profit and popularity accumulation, some streamers mix pornography content into live content to avoid online supervision. Therefore, accurate recognition of porn streamers in live video streaming has become a challenging task. Porn streamers in live video present multimodal characteristics including visual and acoustic content. Therefore, a porn streamer recognition method in live video streaming is proposed that uses attention-gated multimodal deep features. Our contribution includes the following: (1) multimodal deep features, i.e., spatial, motion and audio, are extracted from live video streaming using convolutional neural networks (CNNs), in which the temporal context of multimodal features is obtained with a bi-directional gated recurrent unit (Bi-GRU); (2) the tri-attention gated mechanism is applied to map the associations between different modalities by assigning higher weights to important features for further reduction in the redundancy of multimodal features; (3) porn streamers in live video streaming are recognized via the attention-gated multimodal deep features. Six experiments are conducted on a real-world dataset, and the competitive results demonstrate that our method can effectively recognize porn streamers in live video streaming.
Jing Zhang 0023, Qi Tian 0001, Li Zhuo 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Semantic Segmentation of Large-Size VHR Remote Sensing Images Using a Two-Stage Multiscale Training Architecture
abstract
Very-high resolution (VHR) remote sensing images (RSIs) have significantly larger spatial size compared to typical natural images used in computer vision applications. Therefore, it is computationally unaffordable to train and test classifiers on these images at a full-size scale. Commonly used methodologies for semantic segmentation of RSIs perform training and prediction on cropped image patches. Thus, they have the limitation of failing to incorporate enough context information. In order to better exploit the correlations between ground objects, we propose a deep architecture with a two-stage multiscale training strategy that is tailored to the semantic segmentation of large-size VHR RSIs. In the first stage of the training strategy, a semantic embedding network is designed to learn high-level features from downscaled images covering a large area. In the second training stage, a local feature extraction network is designed to introduce low-level information from cropped image patches. The resulting training strategy is able to fuse complementary information learned from multiple levels to make predictions. Experimental results on two data sets show that it outperforms local-patch-based training models in terms of both accuracy and stability.
Lei Ding 0008, Jing Zhang 0023, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.2
2019 Personalized Recommendation of Social Images by Constructing a User Interest Tree With Deep Features and Tag Trees
abstract
In view of the great diversity and complexity of social images, it is of great significance to improve the performance of personalized recommendation by learning a user interest from large-scale social images. Deep learning, as the latest research in the field of artificial intelligence, provides a new personalized recommendation solution of social images for learning a user's interest. Moreover, social image sharing websites (such as Flickr) allow users to tag uploaded images with tags. As an important image semantic cue, effective tags not only represent the latent image information but also show personalized user interest. Therefore, a personalized recommendation method of social image is proposed by constructing a user-interest tree with deep features and tag trees in this paper. The main contributions of our paper are as follows: first, to efficiently make use of tags, a tag tree of social images is created by the re-ranked tags; second, for compactly representing the image content, deep features are learned by training the AlexNet network; third, a user-interest tree is constructed with deep features and tag trees that include the user-interest tree of social images and the user-interest tree of tags, respectively, and finally, a personalized recommendation system of social images is built based on a user-interest tree. Experiments on the NUS-WIDE dataset have shown that our method outperforms state-of-the-art methods in terms of both precision and recall of personalized recommendations.
Jing Zhang 0023, Ying Yang 0018, Li Zhuo 0001, Qi Tian 0001, Xi Liang 0003
IEEE Trans. Multim.1
2018 An efficient method of content-targeted online video advertising
Guanyao Wang, Li Zhuo 0001, Jiafeng Li 0001, Dongyue Ren, Jing Zhang 0023
J. Vis. Commun. Image Represent.5
2018 Vehicle color recognition using Multiple-Layer Feature Representations of lightweight convolutional neural network
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Hui Zhang 0049
Signal Process.4
2017 Automatic Tongue Image Segmentation for Traditional Chinese Medicine Using Deep Neural Network
Panling Qu, Hui Zhang 0049, Li Zhuo 0001, Jing Zhang 0023, Guoying Chen
ICIC (1)4
2017 Tag tree creation of social image for personalized recommendation
abstract
The tags are usually tagged by different users in social image sharing websites, which can indicate image semantic information and imply user's preference. Therefore, the tags can contribute to personalized recommendation of social image. However, the present social image tags models only consider single tag, resulting in the relationships among tags are ignored. In this paper, we propose a novel method to create tag tree of social image for personalized recommendation. Firstly, the tag ranking is realized to remove noisy tags. Then, the first layer tags are selected from re-ranked tags lists. To sufficiently express tag's significances, the tag subtrees can be created based on different image categories and combined with first layer tags to create tag tree. Finally, the personalized recommendation of social image is achieved by using tag tree. Experimental results show that our tag tree can effectively express the relationships among tags as well as obtain satisfactory results in personalized recommendation of social image.
Ying Yang 0018, Jing Zhang 0023, Jihong Liu, Jiafeng Li 0001, Li Zhuo 0001
ICIP2
2017 A CBIR System for Hyperspectral Remote Sensing Images Using Endmember Extraction
abstract
With the rapid development of remote sensing technology, searching the similar image is a challenge for hyperspectral remote sensing image processing. Meanwhile, the dramatic growth in the amount of hyperspectral remote sensing data has stimulated considerable research on content-based image retrieval (CBIR) in the field of remote sensing technology. Although many CBIR systems have been developed, few studies focused on the hyperspectral remote sensing images. A CBIR system for hyperspectral remote sensing image using endmember extraction is proposed in this paper. The main contributions of our method are that: (1) the endmembers as the spectral features are extracted from hyperspectral remote sensing image by improved automatic pixel purity index (APPI) algorithm; (2) the spectral information divergence and spectral angle match (SID–SAM) mixed measure method is utilized as a similarity measurement between hyperspectral remote sensing images. At last, the images are ranked with descending and the top-[Formula: see text] retrieved images are returned. The experimental results on NASA datasets show that our system can yield a superior performance.
Jing Zhang 0023, Qianlan Zhou, Li Zhuo 0001, Wenhao Geng, Suyu Wang
Int. J. Pattern Recognit. Artif. Intell.1
2017 Vehicle classification for large-scale traffic surveillance videos using Convolutional Neural Networks
Li Zhuo 0001, Liying Jiang, Jiafeng Li 0001, Jing Zhang 0023
Mach. Vis. Appl.5
2017 Road Recognition From Remote Sensing Imagery Using Incremental Learning
abstract
Roads, as important artificial objects, are the main body of modern traffic system, providing many conveniences for human civilization. With the development of Intelligent Transportation Systems (ITS), the road structure is changing frequently. Road recognition is to identify the road type from remote sensing imagery, and road types depend largely on the characteristics of roads. Thus, how to extract road features and further making road classification efficient have become a popular and challenging research topic. In this paper, we propose a road recognition method for remote sensing imagery using incremental learning. In principle, our method includes the following steps: 1) the non-road remote sensing imagery is first filtered by using support vector machine; 2) the road network is obtained from the road remote sensing imagery by computing multiple saliency features; 3) the road features are extracted from road network and background environment; and 4) the roads are recognized as three road types according to the classification results of incremental learning algorithm. The experimental results show that our method has higher road recognition rate as well as less recognition time than the other popular algorithms.
Jing Zhang 0023, Chao Wang 0017, Li Zhuo 0001, Qi Tian 0001, Xi Liang 0003
IEEE Trans. Intell. Transp. Syst.1
2017 Personalized Social Image Recommendation Method Based on User-Image-Tag Model
abstract
In the social image sharing websites (such as Flickr), users are allowed to upload images and tag them with tags. Due to the diversities of users' interests, different users may tag the same image with different tags. Therefore, tags not only reveal some important image semantic clues, but also show user's preference, which can provide a new effective solution for overcoming the semantic gap as well as realizing a personalized recommendation. In this paper, a personalized social image recommendation method based on user-image-tag model is proposed. The main contributions of our work are 1) to efficiently make use of tags, social image tags are re-ranked according to the image content; 2) to obtain user preference, a user-image-tag model is constructed with tripartite graph according to the correlation among users, images and top-ranking tags; and 3) a personalized social recommendation system is implemented based on user-image-tag model. Experimental results proved that our method can significantly improve the accuracy of personalized image recommendation.
Jing Zhang 0023, Ying Yang 0018, Qi Tian 0001, Li Zhuo 0001, Xin Liu 0016
IEEE Trans. Multim.1
2016 Extreme Weather Recognition Using Convolutional Neural Networks
abstract
Extreme weather always brings potential risk to driving, which leads to people's life and property being put into great dangers. Therefore, the automatic recognition of extreme weather plays an important role in the application of the highway traffic condition warning, automobile auxiliary driving, climate analysis and so on. Generally, multiple sensors are adopted in traditional methods of automatic extreme weather recognition with artificial participation and low accuracy. A new extreme weather recognition method based on images by using computer vision manners has been proposed in this paper. Since the weather is affected by many factors, features that can accurately represent various weather characteristics are difficult to be extracted. Therefore, in this paper, convolutional neural networks (CNNs) are applied to settle this problem. Features of extreme weather and recognition models are generated from big data. Moreover, a large-scale extreme weather dataset, "WeatherDataset", has been collected, in which 16635 extreme weather images are divided into four classes (sunny, rainstorm, blizzard, and fog), and complex scenes are coverd. A recognition model for extreme weather is obtained through two steps: Pre-training and Fine Tuning. In Pre-training step, ILSVRC-2012 Dataset is trained to obtain the model of ILSVRC using GoogLeNet. A more accurate model for extreme weather recognition is obatined by further fine-tuning GoogLeNet on WeatherDataset. The experimental results show that the proposed method is able to achieve a high performance with the recognition accuracy rate of 94.5% and can meet the requirements of some real applications.
Li Zhuo 0001, Panling Qu, Kailong Zhou, Jing Zhang 0023
ISM5
2016 Automatic Endmember Extraction Using Pixel Purity Index for Hyperspectral Imagery
Qianlan Zhou, Jing Zhang 0023, Qi Tian 0001, Li Zhuo 0001, Wenhao Geng
MMM (2)2
2016 ORB feature based web pornographic image recognition
Li Zhuo 0001, Zhen Geng, Jing Zhang 0023
Neurocomputing3
2016 A K-PLSR-based color correction method for TCM tongue images under different illumination conditions
Li Zhuo 0001, Pei Zhang 0012, Panling Qu, Yuanfan Peng, Jing Zhang 0023
Neurocomputing5
2015 An Assessment Method of Tongue Image Quality Based on Random Forest in Traditional Chinese Medicine
Xinfeng Zhang 0002, Yazhen Wang, Guangqin Hu, Jing Zhang 0023
ICIC (3)4
2015 Preliminary Study of Tongue Image Classification Based on Multi-label Learning
Xinfeng Zhang 0002, Jing Zhang 0023, Guangqin Hu, Yazhen Wang
ICIC (3)2
2015 Creating descriptive visual words for tag ranking of compressed social image
abstract
Visual words description method has been widely applied in the fields of social image's tag ranking, tag recommendation and annotation. At present, visual words are usually obtained by unsupervised clustering methods which lead to generate many unnecessary and non-descriptive words. Therefore, how to make visual words be descriptive has become a very meaningful task for tag ranking of social image. However, for compressed social image on the network, visual words are created after fully decompressing a compressed image into pixel domain. In this paper, creating descriptive visual words in compressed domain is proposed for tag ranking of compressed social image. Firstly, the traditional visual words are created by using the partly decoded data; then the descriptive visual words are selected from traditional visual words by the VisualWordRank ranking algorithm; finally the descriptive visual words are applied to rank the tag of social image. Experimental results show the descriptive visual words can improve the accuracy of tag ranking, which further prove our method has more descriptive ability. Besides that, our method also reduces the processing time for compressed social image greatly.
Xin Liu 0016, Jing Zhang 0023, Li Zhuo 0001, Ying Yang 0018
ICIP2
2015 Social images tag ranking based on visual words in compressed domain
Jing Zhang 0023, Xin Liu 0016, Li Zhuo 0001, Chao Wang 0017
Neurocomputing1
2015 A Compressed-Domain Image Filtering and Re-Ranking Approach for Multi-Agent Image Retrieval
abstract
For the limited transmission capacity and compressed images in the network environment, a compressed-domain image filtering and re-ranking approach for multi-agent image retrieval is proposed in this paper. Firstly, the distributed image retrieval platform with multi-agent is constructed by using Aglet development system, the lifecycle and the migration mechanism of agent is designed and planned for multi-agent image retrieval by using the characteristics of mobile agent. Then, considering the redundant image brought by distributed multi-agent retrieval, the duplicate images in distributed retrieval results are filtered based on the perceptual hashing feature extracted in the compressed-domain. Finally, weight-based hamming distance is utilized to re-rank the retrieval results. The experimental results show that the proposed approach can effectively filter the duplicate images in distributed image retrieval results as well as improve the accuracy and speed of compressed-domain image retrieval.
Jing Zhang 0023, Li Zhuo 0001, Xin Liu 0016, Ying Yang 0018
Int. J. Pattern Recognit. Artif. Intell.1
2014 Uniform color space based facial complexion recognition for Traditional Chinese Medicine
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the facial complexion to obtain the disease diagnostic results. Color space based on human visual system will be more conducive to facial complexion recognition, which is more suitable to measure and distinguish facial complexion. Uniform color space based facial complexion recognition for TCM is proposed in this paper, which include: (1) the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately; (2) the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature; (3) the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the proposed complexion feature can achieve good performance, with the facial complexion recognition rate up to 81%.
Jing Zhang 0023, Chao Wang 0017, Li Zhuo 0001, Yuncong Yang
ICARCV1
2014 Automatic tongue color analysis of traditional Chinese medicine based on image retrieval
abstract
Content-Based Image Retrieval (CBIR) characterizes the image content by extracting visual features, and measures the similarity according to the distance between the two features. This paper adopts CBIR to perform automatic tongue color analysis of Traditional Chinese Medicine (TCM). Firstly, we extract the visual features of tongue images to be analyzed, especially the color features; and then retrieve the similar tongue images from the database, which have been labeled by TCM doctors in advance. Finally, statistical decision method is exploited based on the retrieval results to classify the tongue color. Experimental results show that the proposed method can achieve the classification accuracy of 87.85% and 88.54% respectively for the colors of tongue substance and tongue coating. The proposed method in this paper can provide a new means for the tongue color automatic analysis of TCM, and it is also a new application of CBIR.
Li Zhuo 0001, Pei Zhang 0012, Bo Cheng 0007, Jing Zhang 0023
ICARCV5
2014 A comparative study of dimensionality reduction methods for large-scale image retrieval
Li Zhuo 0001, Bo Cheng 0007, Jing Zhang 0023
Neurocomputing3
2014 An SA-GA-BP neural network-based color correction algorithm for TCM tongue images
Li Zhuo 0001, Jing Zhang 0023, Pei Dong, Yingdi Zhao
Neurocomputing2
2014 Human Facial Complexion Recognition of Traditional Chinese Medicine Based on Uniform Color Space
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the human facial complexion to obtain the disease diagnostic results. The morbidity of the organs can be revealed by the human facial complexion, so the color space based on human visual system will be more conducive to facial complexion recognition. It is much suitable to measure and distinguish facial complexion by uniform Lab color space, as it has the characteristic of isometry and high resolving power. First, the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately. Second, the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature. At last, the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the complexion feature proposed in this paper can achieve the better performance, with the facial complexion recognition rate up to 81%.
Li Zhuo 0001, Yuncong Yang, Jing Zhang 0023
Int. J. Pattern Recognit. Artif. Intell.3
2013 Comparative Study on Dimensionality Reduction in Large-Scale Image Retrieval
abstract
Dimensionality reduction plays a significant role for the performance of large-scale image retrieval. In this paper, various dimensionality reduction methods are compared to validate their own performance in image retrieval. For this purpose, first, the Scale Invariant Feature Transform (SIFT) features and HSV (Hue, Saturation, Value) histogram are extracted as image features. Second, the Principal Component Analysis (PCA), Fisher Linear Discriminant Analysis (FLDA), Local Fisher Discriminant Analysis (LFDA), Isometric Mapping (ISOMAP), Locally Linear Embedding (LLE), and Locality Preserving Projections (LPP) are respectively applied to reduce the dimensions of SIFT feature descriptors and color information, which can be used to generate vocabulary trees. Finally, through setting the match weights of vocabulary trees, large-scale image retrieval scheme is implemented. By comparing multiple sets of experimental data from several platforms, it can be concluded that dimensionality reduction method of LLE and LPP can effectively reduce the computational cost of image features, and maintain the high retrieval performance as well.
Bo Cheng 0007, Li Zhuo 0001, Jing Zhang 0023
ISM3
2013 Pornographic image region detection based on visual attention model in compressed domain
abstract
According to biological attention mechanism, a region of interest (ROI) detection based on visual attention model is closer to human visual system. Taken into account the characteristics of pornographic image during regions detection, a pornographic image region detection method based on visual attention model in compressed domain is proposed in this study, which includes the following four steps: (i) the skin colour regions of pornographic images are detected in compressed domain; (ii) visual saliency map in compressed domain is computed to construct visual attention model; (iii) threshold segmentation method is used for visual saliency map, and then the torso information is retained as pornographic regions; and (iv) four features of colour, texture, intensity and skin are extracted to represent pornographic region. The experimental results show that the proposed method can perform well on the speed/accuracy of pornographic regions detection and representation.
Jing Zhang 0023, Lei Sui, Li Zhuo 0001
IET Image Process.1
2013 An approach of bag-of-words based on visual attention model for pornographic images recognition in compressed domain
Jing Zhang 0023, Lei Sui, Li Zhuo 0001, Yuncong Yang
Neurocomputing1
2013 Compressed domain based pornographic image recognition using multi-cost sensitive decision trees
Li Zhuo 0001, Jing Zhang 0023, Yingdi Zhao
Signal Process.2
2012 Region of Interest Detection Based on Visual Perception Model
abstract
According to human vision theory, the image is conveyed from human visual system to brain when people have a look at. Different from previous work, the study reported in this paper attempts to simulate a more real and complex method for region of interest (ROI) detection and quantitatively analyze the correlation between users' visual perception and ROI. In this paper, a visual perception model-based ROI detection is proposed, which can be realized with an ordinary web camera. Visual perception model employs a combination of visual attention model and gaze tracking data to objectively detect ROIs. The work includes pre-ROI estimation using visual attention model, gaze data collection and ROI detection. Pre-ROIs are segmented by the visual attention model. Since eye feature extraction is critical to the accuracy and performance of gaze tracking, adaptive eye template and neural network are employed to predict gaze points. By computing the density of the gaze points, ROIs are ranked. Experimental results show that the accuracy of our ROI detection method can be raised as high as 97% and it is also demonstrated that our model can efficiently adapt to users' interests and match the objective ROI.
Jing Zhang 0023, Li Zhuo 0001, Yingdi Zhao
Int. J. Pattern Recognit. Artif. Intell.1
2011 Regions of Interest Extraction Based on Visual Saliency in Compressed Domain
abstract
Recently bag-of-words (BoW) model having been widely used in textual information processing has been extended into many tasks in visual domain such as image classification, scene analysis, image annotation and image retrieval, namely bag-of-visual-words (BoVW) model. Therefore, it is essential to create an effective visual vocabulary. Most of existing approaches create visual vocabularies from image in pixel domain, which requires extra processing time in decompressed images, since most images are stored in compressed format. In this paper we propose to create a visual vocabulary based on Scale Invariant Feature Transform(SIFT) descriptor in compressed domain with the following three steps, (1) constructing low-resolution images in compressed domain, (2) extracting SIFT descriptor from low-resolution images, and (3) creating a visual vocabulary based on extracted SIFT descriptors. In order to evaluate the performance of the visual words, experiments have been conducted on identifying pornographic images. Experimental results indicate that the proposed method can recognize pornographic images accurately with much reduced computational time.
Lei Sui, Jing Zhang 0023, Li Zhuo 0001, Yuncong Yang
ISM2
2010 A Personalized Image Retrieval Based on User Interest Model
abstract
In order to narrow the semantic gap, user interest model plays an important role in personalized image retrieval. A novel personalized image retrieval approach based on user interest model is proposed in this study. User interest model is developed on the basis of short-tem and long-term interests. (1) Short-term interests are represented by collecting visual and semantic features. Visual features are collected by MARS relevance feedback. Semantic features are constructed by building a mapping from image low-level visual features to high-level semantic features on the basis of SVM. (2) Long-term interests are inferred by inference engine from the collected short-term interests. Long-term visual features are collected by the nonlinear gradual forgetting interest inference algorithm and semantic features are obtained by clustering algorithm. After applying to image retrieval, experimental results show that the average recall/precision is significantly improved and a better user satisfaction rate is achieved as well. Furthermore, it demonstrates our model can be efficiently adapted to user interests and matches personalized image retrieval.
Jing Zhang 0023, Li Zhuo 0001, Lansun Shen
Int. J. Pattern Recognit. Artif. Intell.1
2008 A low complexity method for real-time gaze tracking
abstract
The paper addresses a novel, fast and low complexity gaze tracking method, which remedies computation complexity and real-time processing problems of most existing gaze tracking methods. Face images are captured and pre-processed under infrared lamp-houses with dual thresholds. Pupil area is verified by the run-length in eight directions, and the gaze point is located by mapping the vectors of pupil and reflection point or using neural network. The experimental results show that the real-time processing of gaze tracking, with high resolution and high accuracy simultaneously, is achieved in this new system.
Jing Zhang 0023, Mengkai Zhao, Li Zhuo 0001, Lansun Shen
MMSP1