Wenjing Jia

dblp:24/669 · DBLP profile ↗
← Back
97ranked-venue papers
5as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 38 · 1 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Wrinkles in time: Multi-scale patching and super-resolution for efficient time series forecasting
Yuwei Chen 0007, Wenjing Jia, Qiang Wu 0001
Neurocomputing2
2026 ECMVC: Entropy-aware Curriculum-guided Multi-view Contrastive Clustering
abstract
Multi-view clustering aims to integrate complementary information from multiple views to achieve better performance than single-view clustering. However, in practical scenarios, the quality of each view is often inconsistent, with some views containing substantial noise or redundant information, which may adversely affect the overall clustering performance. Moreover, existing contrastive learning techniques typically employ overly simplistic strategies for negative sample selection, making them prone to local optima during training and compromising model effectiveness. To address these challenges, this paper proposes a novel information fusion-based deep multi-view contrastive clustering algorithm, termed ECMVC. The proposed method explicitly models both consistency and complementarity among views and leverages a feature fusion network to enhance the stability and accuracy of clustering in noisy and redundant environments. In addition, we further propose a curriculum-guided contrastive learning approach, where an entropy-driven dynamic scheduler adaptively selects informative negative samples and progressively increases the training difficulty. This curriculum-guided mechanism enables faster convergence and more stable optimization. Experiments on multiple benchmark datasets demonstrate the effectiveness of the proposed method.
Yuquan Shao, Yangyang Zhou, Wenjing Jia
Neural Process. Lett.4
2026 Reliable-Teacher: Uncertainty-Guided Collaborative Learning for Nighttime Object Detection
abstract
Nighttime object detection presents significant challenges due to the scarcity of large-scale, high-quality annotations across diverse nighttime scenarios. To circumvent the need for manual nighttime image annotation, researchers have explored Unsupervised Domain Adaptive Object Detection (UDA-OD), which transfers knowledge from labeled daytime datasets to unlabeled nighttime data through pseudo-labeling. While existing approaches have shown promising results, their effectiveness remains limited by the low quality of pseudo labels, restricting model adaptation to nighttime conditions. To address these limitations, we propose Reliable-Teacher, a novel mutual-learning framework that comprehensively leverages target domain knowledge through Uncertainty-Guided Collaborative Learning. Specifically, our approach consists of three key components: 1) A Collaborative Pseudo-Label Construction module that intelligently integrates reliable Teacher-generated pseudo-labels into Student proposals, significantly enhancing pseudo-label quality; 2) An Uncertainty-Guided Consistency Reasoning module that enforces inter-category consistency between Teacher and Student predictions at both anchor and bounding box levels; 3) A Reliability-Weighted Classification Loss that minimizes the influence of unreliable predictions to further enhance uncertainty-guided learning. Extensive experiments demonstrate that Reliable-Teacher significantly outperforms state-of-the-art methods, achieving performance gain of up to 3.1%, 2.2% and 1.7% mAP on BDD100K [1], SHIFT [2], and VisDrone [3] benchmarks, respectively. Upon acceptance, our code will be released to facilitate further research in this domain.
Wenjing Jia, Jiaqi Xiao, Jinchang Ren, Di Yuan 0002, Qiguang Miao, Xiangjian He
IEEE Trans. Circuits Syst. Video Technol.3
2026 SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation
abstract
The Vision Transformer (ViT) has achieved notable success in computer vision, with its variants widely validated across various downstream tasks, including semantic segmentation. However, as general-purpose visual encoders, ViT backbones often do not fully address the specific requirements of task decoders, highlighting opportunities for designing decoders optimized for efficient semantic segmentation. This paper proposes Strip Cross-Attention (SCASeg), an innovative decoder head specifically designed for semantic segmentation. Instead of relying on the conventional skip connections, we utilize lateral connections between encoder and decoder stages, leveraging encoder features as Queries in cross-attention modules. Additionally, we introduce a Cross-Layer Block (CLB) that integrates hierarchical feature maps from various encoder and decoder stages to form a unified representation for Keys and Values. The CLB also incorporates the local perceptual strengths of convolution, enabling SCASeg to capture both global and local context dependencies across multiple layers, thus enhancing feature interaction at different scales and improving overall efficiency. To further optimize computational efficiency, SCASeg compresses the channels of queries and keys into one dimension, creating strip-like patterns that reduce memory usage and increase inference speed compared to traditional vanilla cross-attention. Experiments show that SCASeg's adaptable decoder delivers competitive performance across various setups, outperforming leading segmentation architectures on benchmark datasets, including ADE20K, Cityscapes, COCO-Stuff 164k, and Pascal VOC2012, even under diverse computational constraints.
Guoan Xu, Jiaming Chen 0001, Wenfeng Huang, Wenjing Jia, Guangwei Gao, Guo-Jun Qi
IEEE Trans. Image Process.4
2025 Versatile and Efficient Medical Image Super-Resolution Via Frequency-Gated Mamba
abstract
Medical image super-resolution (SR) is essential for enhancing diagnostic accuracy while reducing acquisition cost and scanning time. However, modeling both long-range anatomical structures and fine-grained frequency details with low computational overhead remains challenging. We propose FGMamba, a novel frequency-aware gated state-space model that unifies global dependency modeling and fine-detail enhancement into a lightweight architecture. Our method introduces two key innovations: a Gated Attention-enhanced State-Space Module (GASM) that integrates efficient state-space modeling with dualbranch spatial and channel attention, and a Pyramid Frequency Fusion Module (PFFM) that captures high-frequency details across multiple resolutions via FFT-guided fusion. Extensive evaluations across five medical imaging modalities (Ultrasound, OCT, MRI, CT, and Endoscopic) demonstrate that FGMamba achieves superior PSNR/SSIM while maintaining a compact parameter footprint (<0.75M), outperforming CNN-based and Transformerbased SOTAs. Our results validate the effectiveness of frequencyaware state-space modeling for scalable and accurate medical image enhancement. Source code and dataset will be made publicly available.
Wenfeng Huang, Xiangyun Liao, Wei Cao 0008, Wenjing Jia, Weixin Si
BIBM4
2025 PointSR: Self-Regularized Point Supervision for Drone-View Object Detection
abstract
Point-Supervised Object Detection (PSOD) in a discriminative style has recently gained significant attention for its impressive detection performance and cost-effectiveness. However, accurately predicting high-quality pseudo-box labels for drone-view images, which often feature densely packed small objects, remains a challenge. This difficulty arises primarily from the limitation of rigid sampling strategies, which hinder the pseudo-box optimization process. To address this, we propose PointSR, an effective and robust point-supervised object detection framework with self-regularized sampling that integrates temporal and informative constraints throughout the pseudo-box generation process. Specifically, the framework comprises three key components: Temporal-Ensembling Encoder (TE Encoder), Coarse Pseudo-box Prediction, and Pseudo-box Refinement. The TE Encoder builds an anchor prototype library by aggregating temporal information for dynamic anchor adjustment. In Coarse Pseudo-box Prediction, anchors are refined using the prototype library, and a set of informative samples is collected for subsequent refinement. During Pseudo-box Refinement, these informative negative samples are used to suppress low-confidence candidate positive samples, thereby improving the quality of the pseudo-boxes. Experimental results on benchmark datasets demonstrate that PointSR significantly outperforms state-of-the-art methods, achieving up to 2.6% ∼ 7.2% higher AP50using only point supervision. Additionally, it exhibits strong robustness to perturbation in human-labeled points.
Weizhuo Li, Wenjing Jia, Zehao Zhang, Xiangzeng Liu, Qiguang Miao
CVPR3
2025 CD-Net: Context-Driven Ultrasound Image Enhancement, a 2.5D Approach for Scoliosis Assessment
abstract
Recent developments have established ultrasound imaging as a promising new standard for scoliosis assessment. However, its limited penetration in soft tissue inevitably produces acoustic shadowing and loss of deeper structural information in individual B-mode slices. Critical spine-related features could therefore be absent or inconsistently represented across adjacent slices, and effective feature extraction is further undermined by speckle noise, a major additive object, which further decreases local contrast. Together, these limitations impede the detection of key anatomical landmarks such as thoracic bony features (TBF) and lamellar bone features (LBF), blocking accurate downstream analysis. To overcome these challenges, we present the Contextual-Driven Ultrasound Enhancement Network (CD-Net), a self-supervised framework that fuses information across slices and refines local detail. CD-Net comprises two core modules: (1) the Contextual Cross-Attention Transfer (CCAT) module, which captures and transfers interslice spatial relationships, and (2) the Localized Attention Contrast Enhancement (LoCE) module, which selectively sharpens and enhances feature regions. On a dataset of 309 patients, CD-Net boosts the detection rate from 78.25% to 93.18% and achieves a Structural Similarity Index (SSIM) of 89.2% (σ = 0.045). Enhanced visibility of TBF and LBF across varying image qualities demonstrates CD-Net’s potential to significantly improve the reliability and efficiency of ultrasound-based scoliosis diagnosis in clinical practice.
Sumartini Dana, Wenjing Jia, Sai-Ho Ling
SMC3
2025 ISC-Swin: Inter Sample Contrastive Enhancement for Swin-Transformer in Ultrasound Spine Feature Segmentation
abstract
Scoliosis, a three-dimensional spinal deformity, requires early and accurate detection for effective treatment. While Cobb’s angle measurement from radiographs remains the clinical gold standard, the associated radiation exposure underscores the need for safer alternatives. Ultrasound imaging offers a non-invasive solution; however, it presents significant challenges, including low contrast, high noise, and irregular anatomical structures, which complicate the accurate estimation of the Ultrasound Curve Angle (UCA).Previous studies have attempted to improve segmentation performance on ultrasound images, often relying on limited paired datasets. While these methods can enhance results, they risk overfitting due to the small sample size and lack a broader understanding of inter-sample relationships.To address these limitations, we propose ISC-Swin, a Swin Transformer-based framework integrated with inter-sample contrastive learning for more robust spine feature segmentation in ultrasound images. Our architecture leverages both local and global contextual information through a novel Inter-Sample Contrastive Bank (ISCB), which dynamically extracts multilevel features across diverse samples. By explicitly modeling inter-class and intra-class differences, ISC-Swin improves the detection of subtle spinal features in challenging, noisy environments.Experimental results show that ISC-Swin achieves a 1–5% improvement in both Dice Similarity Coefficient and Intersection over Union (IoU) metrics, surpassing current state-of-the-art models in bone feature detection and enhancing diagnostic precision in ultrasound-based scoliosis assessment.
Wenjing Jia, Sai-Ho Ling
SMC2
2025 Empowering large language models for automated clinical assessment with generation-augmented retrieval and hierarchical chain-of-thought
abstract
BACKGROUND: Understanding and extracting valuable information from electronic health records (EHRs) is important for improving healthcare delivery and health outcomes. Large language models (LLMs) have demonstrated significant proficiency in natural language understanding and processing, offering promises for automating the typically labor-intensive and time-consuming analytical tasks with EHRs. Despite the active application of LLMs in the healthcare setting, many foundation models lack real-world healthcare relevance. Applying LLMs to EHRs is still in its early stage. To advance this field, in this study, we pioneer a generation-augmented prompting paradigm "GAPrompt" to empower generic LLMs for automated clinical assessment, in particular, quantitative stroke severity assessment, using data extracted from EHRs. METHODS: The GAPrompt paradigm comprises five components: (i) prompt-driven selection of LLMs, (ii) generation-augmented construction of a knowledge base, (iii) summary-based generation-augmented retrieval (SGAR); (iv) inferencing with a hierarchical chain-of-thought (HCoT), and (v) ensembling of multiple generations. RESULTS: GAPrompt addresses the limitations of generic LLMs in clinical applications in a progressive manner. It efficiently evaluates the applicability of LLMs in specific tasks through LLM selection prompting, enhances their understanding of task-specific knowledge from the constructed knowledge base, improves the accuracy of knowledge and demonstration retrieval via SGAR, elevates LLM inference precision through HCoT, enhances generation robustness, and reduces hallucinations of LLM via ensembling. Experiment results demonstrate the capability of our method to empower LLMs to automatically assess EHRs and generate quantitative clinical assessment results. CONCLUSION: Our study highlights the applicability of enhancing the capabilities of foundation LLMs in medical domain-specific tasks, i.e., automated quantitative analysis of EHRs, addressing the challenges of labor-intensive and often manually conducted quantitative assessment of stroke in clinical practice and research. This approach offers a practical and accessible GAPrompt paradigm for researchers and industry practitioners seeking to leverage the power of LLMs in domain-specific applications. Its utility extends beyond the medical domain, applicable to a wide range of fields.
Zhanzhong Gu, Wenjing Jia, Massimo Piccardi, Ping Yu 0004
Artif. Intell. Medicine2
2025 ReviveDiff: A Universal Diffusion Model for Restoring Images in Adverse Weather Conditions
abstract
Images captured in challenging environments-such as nighttime, smoke, rainy weather, and underwater-often suffer from significant degradation, resulting in a substantial loss of visual quality. The effective restoration of these degraded images is critical for the subsequent vision tasks. While many existing approaches have successfully incorporated specific priors for individual tasks, these tailored solutions limit their applicability to other degradations. In this work, we propose a universal network architecture, dubbed "ReviveDiff", which can address various degradations and restore images to their original quality by enhancing and restoring their details. Our approach is inspired by the observation that, unlike degradation caused by movement or electronic issues, quality degradation under adverse conditions primarily stems from natural media (such as fog, water, and low luminance), which generally preserves the original structures of objects. To restore the quality of such images, we leveraged the latest advancements in diffusion models and developed ReviveDiff to restore image quality from both macro and micro levels across some key factors determining image quality, such as sharpness, distortion, noise level, dynamic range, and color accuracy. We rigorously evaluated ReviveDiff on seven benchmark datasets covering five types of degrading conditions: Rainy, Underwater, Low-light, Smoke, and Nighttime Hazy. Our experimental results demonstrate that ReviveDiff outperforms the state-of-the-art methods both quantitatively and visually.
Wenfeng Huang, Guoan Xu, Wenjing Jia, Stuart W. Perry, Guangwei Gao
IEEE Trans. Image Process.3
2025 S2AFormer: Strip Self-Attention for Efficient Vision Transformer
abstract
The Vision Transformer (ViT) has achieved remarkable success in computer vision due to its powerful token mixer, which effectively captures global dependencies among all tokens. However, the quadratic complexity of standard self-attention with respect to the number of tokens severely hampers its computational efficiency in practical deployment. Although recent hybrid approaches have sought to combine the strengths of convolutions and self-attention to improve the performance-efficiency trade-off, the costly pairwise token interactions and heavy matrix operations in conventional self-attention remain a critical bottleneck. To overcome this limitation, we introduce S2AFormer, an efficient Vision Transformer architecture built around a novel Strip Self-Attention (SSA) mechanism. Our design incorporates lightweight yet effective Hybrid Perception Blocks (HPBs) that seamlessly fuse the local inductive biases of CNNs with the global modeling capability of Transformer-style attention. The core innovation of SSA lies in simultaneously reducing the spatial resolution of the key ( $K$ ) and value ( $V$ ) tensors while compressing the channel dimension of the query ( $Q$ ) and key ( $K$ ) tensors. This joint spatial-and-channel compression dramatically lowers computational cost without sacrificing representational power, achieving an excellent balance between accuracy and efficiency. We extensively evaluate S2AFormer on a wide range of vision tasks, including image classification (ImageNet-1K), semantic segmentation (ADE20K), and object detection/instance segmentation (COCO). Experimental results consistently show that S2AFormer delivers substantial accuracy improvements together with superior inference speed and throughput across both GPU and non-GPU platforms, establishing it as a highly competitive solution in the landscape of efficient Vision Transformers.
Guoan Xu, Wenfeng Huang, Wenjing Jia, Jiamao Li, Guangwei Gao, Guo-Jun Qi
IEEE Trans. Image Process.3
2024 Zero Trust for Intrusion Detection System: A Systematic Literature Review
Abeer Z. Alalmaie, Nazar Waheed, Mohrah Alalyan, Priyadarsi Nanda, Wenjing Jia, Xiangjian He
ICAART (3)5
2024 MFPNet: A Multi-scale Feature Propagation Network for Lightweight Semantic Segmentation
Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao
ICANN (3)2
2024 Seeing Text in the Dark: Algorithm and Benchmark
abstract
Localizing text in low-light environments is challenging due to visual degradations. Although a straightforward solution involves a two-stage pipeline with low-light image enhancement (LLE) as the initial step followed by detection, LLE is primarily designed for human vision rather than machine vision and can accumulate errors. In this work, we propose an efficient and effective single-stage approach for localizing text in the dark that circumvents the need for LLE. We introduce a constrained learning module as an auxiliary mechanism during the training stage of the text detector. This module is designed to guide the text detector in preserving textual spatial features amidst feature map resizing, thus minimizing the loss of spatial information in texts under low-light visual degradations. Specifically, we incorporate spatial reconstruction and spatial semantic constraints within this module to ensure the text detector acquires essential positional and contextual range knowledge. Our approach enhances the original text detector's ability to identify text's local topological features using a dynamic snake feature pyramid network and adopts a bottom-up contour shaping strategy with a novel rectangular accumulation technique for accurate delineation of streamlined text features. In addition, we present a comprehensive low-light dataset for arbitrary-shaped text, encompassing diverse scenes and languages. Notably, our method achieves state-of-the-art results on this low-light dataset and exhibits comparable performance on standard normal light datasets. The code and dataset will be released.
Chengpei Xu, Hao Fu 0004, Long Ma 0002, Wenjing Jia, Chengqi Zhang, Feng Xia 0001, Xiaoyu Ai, Binghao Li, Wenjie Zhang 0001
ACM Multimedia4
2024 Fine-scale deep learning model for time series forecasting
abstract
Abstract Time series data, characterized by large volumes and wide-ranging applications, requires accurate predictions of future values based on historical data. Recent advancements in deep learning models, particularly in the field of time series forecasting, have shown promising results by leveraging neural networks to capture complex patterns and dependencies. However, existing models often overlook the influence of short-term cyclical patterns in the time series, leading to a lag in capturing changes and accurately tracking fluctuations in forecast data. To overcome this limitation, this paper introduces a new method that utilizes an interpolation technique to create a fine-scaled representation of the cyclical pattern, thereby alleviating the impact of the irregularity in the cyclical component and hence enhancing prediction accuracy. The proposed method is presented along with evaluation metrics and loss functions suitable for time series forecasting. Experiment results on benchmark datasets demonstrate the effectiveness of the proposed approach in effectively capturing cyclical patterns and improving prediction accuracy.
Yuwei Chen 0007, Wenjing Jia, Qiang Wu 0001
Appl. Intell.2
2024 Automatic quantitative stroke severity assessment based on Chinese clinical named entity recognition with domain-adaptive pre-trained large language model
abstract
BACKGROUND: Stroke is a prevalent disease with a significant global impact. Effective assessment of stroke severity is vital for an accurate diagnosis, appropriate treatment, and optimal clinical outcomes. The National Institutes of Health Stroke Scale (NIHSS) is a widely used scale for quantitatively assessing stroke severity. However, the current manual scoring of NIHSS is labor-intensive, time-consuming, and sometimes unreliable. Applying artificial intelligence (AI) techniques to automate the quantitative assessment of stroke on vast amounts of electronic health records (EHRs) has attracted much interest. OBJECTIVE: This study aims to develop an automatic, quantitative stroke severity assessment framework through automating the entire NIHSS scoring process on Chinese clinical EHRs. METHODS: Our approach consists of two major parts: Chinese clinical named entity recognition (CNER) with a domain-adaptive pre-trained large language model (LLM) and automated NIHSS scoring. To build a high-performing CNER model, we first construct a stroke-specific, densely annotated dataset "Chinese Stroke Clinical Records" (CSCR) from EHRs provided by our partner hospital, based on a stroke ontology that defines semantically related entities for stroke assessment. We then pre-train a Chinese clinical LLM coined "CliRoberta" through domain-adaptive transfer learning and construct a deep learning-based CNER model that can accurately extract entities directly from Chinese EHRs. Finally, an automated, end-to-end NIHSS scoring pipeline is proposed by mapping the extracted entities to relevant NIHSS items and values, to quantitatively assess the stroke severity. RESULTS: Results obtained on a benchmark dataset CCKS2019 and our newly created CSCR dataset demonstrate the superior performance of our domain-adaptive pre-trained LLM and the CNER model, compared with the existing benchmark LLMs and CNER models. The high F1 score of 0.990 ensures the reliability of our model in accurately extracting the entities for the subsequent automatic NIHSS scoring. Subsequently, our automated, end-to-end NIHSS scoring approach achieved excellent inter-rater agreement (0.823) and intraclass consistency (0.986) with the ground truth and significantly reduced the processing time from minutes to a few seconds. CONCLUSION: Our proposed automatic and quantitative framework for assessing stroke severity demonstrates exceptional performance and reliability through directly scoring the NIHSS from diagnostic notes in Chinese clinical EHRs. Moreover, this study also contributes a new clinical dataset, a pre-trained clinical LLM, and an effective deep learning-based CNER model. The deployment of these advanced algorithms can improve the accuracy and efficiency of clinical assessment, and help improve the quality, affordability and productivity of healthcare services.
Zhanzhong Gu, Xiangjian He, Ping Yu 0004, Wenjing Jia, Xiguang Yang, Penghui Hu, Shiyan Chen, Yiguang Lin
Artif. Intell. Medicine4
2024 CARD: Semantic Segmentation With Efficient Class-Aware Regularized Decoder
abstract
Semantic segmentation has recently achieved notable advances by exploiting “class-level” contextual information during learning, e.g., the Object Contextual Representation (OCR) and Context Prior (CPNet) approaches. However, these approaches simply concatenate class-level information to pixel features to boost pixel representation learning, which cannot fully utilize intra-class and inter-class contextual information. Moreover, these approaches learn soft class centers based on coarse mask prediction, which is prone to error accumulation. To better exploit class-level information, we propose a universal Class-Aware Regularization (CAR) approach to optimize the intra-class variance and inter-class distance during feature learning, motivated by the fact that humans can recognize an object by itself no matter which other objects it appears with. Moreover, we design a dedicated decoder for CAR (named CARD), which consists of a novel spatial token mixer and an upsampling module, to maximize its gain for existing baselines while being highly efficient in terms of computational cost. Specifically, CAR consists of three novel loss functions. The first loss function encourages more compact class representations within each class, the second directly maximizes the distance between different class centers, and the third further pushes the distance between inter-class centers and pixels. Furthermore, the class center in our approach is directly generated from ground truth instead of from the error-prone coarse prediction. CAR can be directly applied to most existing segmentation models during training, including OCR and CPNet, and can largely improve their accuracy at no additional inference overhead. Extensive experiments and ablation studies conducted on multiple benchmark datasets demonstrate that the proposed CAR can boost the accuracy of all baseline models by up to 2.23% mIOU with superior generalization ability. CARD outperforms state-of-the-art approaches on multiple benchmarks with a highly efficient architecture. The code will be available at https://github.com/edwardyehuang/CAR.
Liang Chen 0026, Wenjing Jia, Xiangjian He, Lixin Duan, Xuefei Zhe, Linchao Bao
IEEE Trans. Circuits Syst. Video Technol.4
2024 Detection-Driven Exposure-Correction Network for Nighttime Drone-View Object Detection
abstract
Drone-view object detection (DroneDet) models typically suffer a significant performance drop when applied to nighttime scenes. Existing solutions attempt to employ an exposure-adjustment module to reveal objects hidden in dark regions before detection. However, most exposure-adjustment models are only optimized for human perception, where the exposure-adjusted images may not necessarily enhance recognition. To tackle this issue, we propose a novel Detection-driven Exposure-correction network for nighttime DroneDet, called DEDet. The DEDet conducts adaptive, nonlinear adjustment of pixel values in a spatially fine-grained manner to generate DroneDet-friendly images. Specifically, we develop a fine-grained parameter predictor (FPP) to estimate pixelwise parameter maps of the image filters. These filters, along with the estimated parameters, are used to adjust pixel values of the low-light image based on nonuniform illuminations in drone-captured images. In order to learn the nonlinear transformation from the original nighttime images to their DroneDet-friendly counterparts, we propose a progressive filtering module that applies recursive filters to iteratively refine the exposed image. Furthermore, to evaluate the performance of the proposed DEDet, we have built a dataset NightDrone to address the scarcity of the datasets specifically tailored for this purpose. Extensive experiments conducted on four nighttime datasets show that DEDet achieves a superior accuracy compared with the state-of-the-art (SOTA) methods. Furthermore, ablation studies and visualizations demonstrate the validity and interpretability of our approach. Our NightDrone dataset can be downloaded fromhttps://github.com/yuexiemail/NightDrone-Dataset.
Wenjing Jia, Qiguang Miao, Junmei Feng, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.2
2024 HAFormer: Unleashing the Power of Hierarchy-Aware Features for Lightweight Semantic Segmentation
abstract
Both Convolutional Neural Networks (CNNs) and Transformers have shown great success in semantic segmentation tasks. Efforts have been made to integrate CNNs with Transformer models to capture both local and global context interactions. However, there is still room for enhancement, particularly when considering constraints on computational resources. In this paper, we introduce HAFormer, a model that combines the hierarchical features extraction ability of CNNs with the global dependency modeling capability of Transformers to tackle lightweight semantic segmentation challenges. Specifically, we design a Hierarchy-Aware Pixel-Excitation (HAPE) module for adaptive multi-scale local feature extraction. During the global perception modeling, we devise an Efficient Transformer (ET) module streamlining the quadratic calculations associated with traditional Transformers. Moreover, a correlation-weighted Fusion (cwF) module selectively merges diverse feature representations, significantly enhancing predictive accuracy. HAFormer achieves high performance with minimal computational overhead and compact model size, achieving 74.2% mIoU on Cityscapes and 71.1% mIoU on CamVid test datasets, with frame rates of 105FPS and 118FPS on a single 2080Ti GPU. The source codes are available at https://github.com/XU-GITHUB-curry/HAFormer.
Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao
IEEE Trans. Image Process.2
2024 Point Clouds are Specialized Images: A Knowledge Transfer Approach for 3D Understanding
abstract
Self-supervised representation learning (SSRL) has gained increasing attention in point cloud understanding, in addressing the challenges posed by 3D data scarcity and high annotation costs. This paper presents PCExpert, a novel SSRL approach that reinterprets point clouds as “specialized images”. This conceptual shift allows PCExpert to leverage knowledge derived from large-scale image modality in a more direct and deeper manner, via extensively sharing the parameters with a pre-trained image encoder in a multi-way Transformer architecture. The parameter sharing strategy, combined with an additional pretext task for pre-training, i.e., transformation estimation, empowers PCExpert to outperform the state of the arts in a variety of tasks, with a remarkable reduction in the number of trainable parameters. Notably, PCExpert's performance underLINEARfine-tuning (e.g., yielding a 90.02% overall accuracy on ScanObjectNN) has already closely approximated the results obtained withFULLmodel fine-tuning (92.66%), demonstrating its effective representation capability.
Jiachen Kang, Wenjing Jia, Xiangjian He, Kin-Man Lam 0001
IEEE Trans. Multim.2
2023 Learning Hierarchical Semantic Information for Efficient Low-Light Image Enhancement
abstract
Low-light environments can cause a variety of complex degradation problems, which result in poor visibility in images. As a classical vision task, low-light image enhancement has attracted an increasing interest in the research community. However, the existing methods tend to require a large number of parameters, making them difficult to implement and optimize, especially on resource-constrained devices. In this paper, we mainly focus on the lightweight of the method and propose a novel end-to-end two-stage CNN-ViT architecture (HSINet) to learn hierarchical semantic information (HSI) from low-light images efficiently. The HSINet consists of two stages: the first stage is a CNN-based low-level semantic (LS) Stage, and the second stage is ViT-based high-level semantic (HS) Stage. The LS Stage contains an efficient multi-scale convolution block, MLS Block, for low-level semantic information extraction. The HS stage, on the other hand, aims to learn the high-level semantic features via ViT's excellent global-learning capability. We propose a hierarchical Swin Transformer-based block, HS Block, to gradually enlarge Swin Transformer's window size as the network becomes deeper, to learn hierarchical high-level semantic information. Benefiting from the efficient architecture, our model only contains 0.6M parameters, far fewer than the existing SOTAs. We evaluated the method on three challenging benchmark datasets: LOL, VE-LOL, and MIT-Adobe FiveK, using three popular evaluation metrics. The quantitative and qualitative results both show that the proposed method not only outperforms the state of the arts in terms of PSNR, SSIM, LPIPS, and visual effects, but also with better efficiency.
Wenfeng Huang, Xiangyun Liao, Yinling Qian, Wenjing Jia
IJCNN4
2023 ABUSDet: A Novel 2.5D deep learning model for automated breast ultrasound tumor detection
Xudong Song, Xiaoyang Lu, Gengfa Fang, Xiangjian He, Xiaochen Fan, Le Cai, Wenjing Jia
Appl. Intell.7
2023 Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision
Jiachen Kang, Wenjing Jia, Xiangjian He
Neurocomputing2
2023 Cross-domain learning for underwater image enhancement
Fei Li 0030, Jiangbin Zheng 0001, Yuan-fang Zhang, Wenjing Jia, Qianru Wei, Xiangjian He
Signal Process. Image Commun.4
2023 Arbitrary-Shape Scene Text Detection via Visual-Relational Rectification and Contour Approximation
abstract
One trend in the latest bottom-up approaches for arbitrary-shape scene text detection is to determine the links between text segments using Graph Convolutional Networks (GCNs). However, the performance of these bottom-up methods is still inferior to that of state-of-the-art top-down methods even with the help of GCNs. We argue that a cause of this is that bottom-up methods fail to make proper use of visual-relational features, which results in accumulated false detection, as well as the error-prone route-finding used for grouping text segments. In this paper, we improve classic bottom-up text detection frameworks by fusing the visual-relational features of text with two effective false positive/negative suppression (FPNS) mechanisms and developing a new shape-approximation strategy. First, dense overlapping text segments depicting the “characterness” and “streamline” properties of text are constructed and used in weakly supervised node classification to filter the falsely detected text segments. Then, relational features and visual features of text segments are fused with a novel Location-Aware Transfer (LAT) module and Fuse Decoding (FD) module to jointly rectify the detected text segments. Finally, a novel multiple-text-map-aware contour-approximation strategy is developed based on the rectified text segments, instead of the error-prone route-finding process, to generate the final contour of the detected text. Experiments conducted on five benchmark datasets demonstrate that our method outperforms the state-of-the-art performance when embedded in a classic text detection framework, which revitalizes the strengths of bottom-up methods.
Chengpei Xu, Wenjing Jia, Tingcheng Cui, Ruomei Wang 0001, Yuan-fang Zhang, Xiangjian He
IEEE Trans. Multim.2
2023 MorphText: Deep Morphology Regularized Accurate Arbitrary-Shape Scene Text Detection
abstract
Bottom-up text detection methods play an important role in arbitrary-shape scene text detection but there are two restrictions preventing them from achieving their great potential, i.e., 1) the accumulation of false text segment detections, which affects subsequent processing, and 2) the difficulty of building reliable connections between text segments. Targeting these two problems, we propose a novel approach, named ``MorphText", to capture the regularity of texts by embedding deep morphology for arbitrary-shape text detection. Towards this end, two deep morphological modules are designed to regularize text segments and determine the linkage between them. First, a Deep Morphological Opening (DMOP) module is constructed to remove false text segment detections generated in the feature extraction process. Then, a Deep Morphological Closing (DMCL) module is proposed to allow text instances of various shapes to stretch their morphology along their most significant orientation while deriving their connections.Extensive experiments conducted on four challenging benchmark datasets (CTW1500, Total-Text, MSRA-TD500 and ICDAR2017) demonstrate that our proposed MorphText outperforms both top-down and bottom-up state-of-the-art arbitrary-shape scene text detection approaches.
Chengpei Xu, Wenjing Jia, Ruomei Wang 0001, Xiangjian He
IEEE Trans. Multim.2
2022 Channelized Axial Attention - considering Channel Relation within Spatial Attention for Semantic Segmentation
abstract
Spatial and channel attentions, modelling the semantic interdependencies in spatial and channel dimensions respectively, have recently been widely used for semantic segmentation. However, computing spatial and channel attentions separately sometimes causes errors, especially for those difficult cases. In this paper, we propose Channelized Axial Attention (CAA) to seamlessly integrate channel attention and spatial attention into a single operation with negligible computation overhead. Specifically, we break down the dot-product operation of the spatial attention into two parts and insert channel relation in between, allowing for independently optimized channel attention on each spatial location. We further develop grouped vectorization, which allows our model to run with very little memory consumption without slowing down the running speed. Comparative experiments conducted on multiple benchmark datasets, including Cityscapes, PASCAL Context, and COCO-Stuff, demonstrate that our CAA outperforms many state-of-the-art segmentation models (including dual attention) on all tested datasets.
Wenjing Jia, Liu Liu 0014, Xiangjian He
AAAI3
2022 CAR: Class-Aware Regularizations for Semantic Segmentation
Liang Chen 0026, Xuefei Zhe, Wenjing Jia, Linchao Bao, Xiangjian He
ECCV (28)5
2022 An anchor-free object detector based on soften optimized bi-directional FPN
Tao Zhang 0010, Bo Jin 0013, Wenjing Jia
Comput. Vis. Image Underst.3
2022 Deep RGB-D Saliency Detection Without Depth
abstract
The existing saliency detection models based on RGB colors only leverage appearance cues to detect salient objects. Depth information also plays a very important role in visual saliency detection and can supply complementary cues for saliency detection. Although many RGB-D saliency models have been proposed, they require to acquire depth data, which is expensive and not easy to get. In this paper, we propose to estimate depth information from monocular RGB images and leverage the intermediate depth features to enhance the saliency detection performance in a deep neural network framework. Specifically, we first use an encoder network to extract common features from each RGB image and then build two decoder networks for depth estimation and saliency detection, respectively. The depth decoder features can be fused with the RGB saliency features to enhance their capability. Furthermore, we also propose a novel dense multiscale fusion model to densely fuse multiscale depth and RGB features based on the dense ASPP model. A new global context branch is also added to boost the multiscale features. Experimental results demonstrate that the added depth cues and the proposed fusion model can both improve the saliency detection performance. Finally, our model not only outperforms state-of-the-art RGB saliency models, but also achieves comparable results compared with state-of-the-art RGB-D saliency models.
Yuan-fang Zhang, Jiangbin Zheng 0001, Wenjing Jia, Wenfeng Huang, Long Li 0008, Nian Liu 0002, Fei Li 0030, Xiangjian He
IEEE Trans. Multim.3
2021 Synthetic CT images for semi-sequential detection and segmentation of lung nodules
Mohammad Hesam Hesamian, Wenjing Jia, Xiangjian He, Paul J. Kennedy
Appl. Intell.2
2021 Nighttime image dehazing based on Retinex and dark channel prior using Taylor series expansion
Qunfang Tang, Jie Yang 0022, Xiangjian He, Wenjing Jia, Qingnian Zhang
Comput. Vis. Image Underst.4
2021 Collaborative algorithms that combine AI with IoT towards monitoring and control system
Tao Zhang 0010, Wenjing Jia, Mu-Yen Chen
Future Gener. Comput. Syst.3
2021 PDANet: Pyramid density-aware attention based network for accurate crowd counting
Saeed Amirgholipour Kasmani, Wenjing Jia, Lei Liu 0036, Xiaochen Fan, Dadong Wang, Xiangjian He
Neurocomputing2
2021 See more than once: Kernel-sharing atrous convolution for semantic segmentation
Wenjing Jia, Yue Lu 0001, Xiangjian He
Neurocomputing3
2021 AMDFNet: Adaptive multi-level deformable fusion network for RGB-D saliency detection
Fei Li 0030, Jiangbin Zheng 0001, Yuanfang Zhang, Nian Liu 0002, Wenjing Jia
Neurocomputing5
2021 Accurate and automatic tooth image segmentation model with deep convolutional neural networks and level set method
Yunyun Yang, Ruicheng Xie, Wenjing Jia, Yunna Yang, Lipeng Xie, Benxiang Jiang
Neurocomputing3
2021 Rethinking feature aggregation for deep RGB-D salient object detection
Yuanfang Zhang, Jiangbin Zheng 0001, Long Li 0008, Nian Liu 0002, Wenjing Jia, Xiaochen Fan, Chengpei Xu, Xiangjian He
Neurocomputing5
2021 Single image deraining using Context Aggregation Recurrent Network
Qunfang Tang, Jie Yang 0022, Zhiqiang Guo, Wenjing Jia
J. Vis. Commun. Image Represent.5
2021 Double level set segmentation model based on mutual exclusion of adjacent regions with application to brain MR images
Yunyun Yang, Ruicheng Xie, Wenjing Jia
Knowl. Based Syst.3
2021 Level set framework with transcendental constraint for robust and fast image segmentation
Yunyun Yang, Xiu Shu, Chong Feng 0002, Ruicheng Xie, Wenjing Jia, Chunming Li
Pattern Recognit.6
2021 UDR: An Approximate Unbiased Difference-Ratio Edge Detector for SAR Images
abstract
Edge detection is a critical component of synthetic aperture radar (SAR) image interpretation. Due to serious speckle noise, the core problems for SAR edge detection are how to keep a constant false alarm rate (CFAR) and how to achieve unbiased localization of the edges. Aiming at these problems, this article proposes a novel edge detector with a unique structure for noise-contaminated SAR images, which creatively integrates the difference operation with ratio operation (hence named as “UDR: unbiased difference-ratio” edge detector). Theoretical analysis proves that the difference operation effectively affords the UDR unbiased localization ability for both ideal and nonideal edges, and the ratio operation provides the UDR the property of CFAR under the influence of speckle noise. Experimental results on both simulated and real-world SAR images demonstrate the unbiased localization ability of the proposed UDR edge detection, insensitive to the changes of edge contrast, the width of the transition zone and the noise level. Benefited from the superior localization precision and insensitivity to noise, the average true positive detection rate of the proposed detector is improved to 95%, outperforming the compared state-of-the-art methods.
Qian-Ru Wei, Da-Zheng Feng, Wenjing Jia
IEEE Trans. Geosci. Remote. Sens.3
2021 DENet: A Universal Network for Counting Crowd With Varying Densities and Scales
abstract
Counting people or objects with significantly varying scales and densities has attracted much interest from the research community and yet it remains an open problem. In this paper, we propose a simple but efficient and effective network, named DENet, which is composed of two components,i.e., a detection network (DNet) and an encoder-decoder estimation network (ENet). We first run the DNet on the input image to detect and count individuals who can be segmented clearly. Then, the ENet is utilized to estimate the density maps of the remaining areas, typically with low resolution and high densities where individuals cannot be detected. For this purpose, we propose a modified Xception network as the encoder for feature extraction and a combination of dilated convolution and transposed convolution as the decoder. When evaluated on the ShanghaiTech Part A, UCF and WorldExpo’10 datasets, our DENet has achieved lower Mean Absolute Error (MAE) than those of the state-of-the-art methods.
Lei Liu 0036, Jie Jiang 0005, Wenjing Jia, Saeed Amirgholipour Kasmani, Yi Wang 0037, Michelle Zeibots, Xiangjian He
IEEE Trans. Multim.3
2021 A New Algorithm for Sketch-Based Fashion Image Retrieval Based on Cross-Domain Transformation
abstract
Due to the rise of e‐commerce platforms, online shopping has become a trend. However, the current mainstream retrieval methods are still limited to using text or exemplar images as input. For huge commodity databases, it remains a long‐standing unsolved problem for users to find the interested products quickly. Different from the traditional text‐based and exemplar‐based image retrieval techniques, sketch‐based image retrieval (SBIR) provides a more intuitive and natural way for users to specify their search need. Due to the large cross‐domain discrepancy between the free‐hand sketch and fashion images, retrieving fashion images by sketches is a significantly challenging task. In this work, we propose a new algorithm for sketch‐based fashion image retrieval based on cross‐domain transformation. In our approach, the sketch and photo are first transformed into the same domain. Then, the sketch domain similarity and the photo domain similarity are calculated, respectively, and fused to improve the retrieval accuracy of fashion images. Moreover, the existing fashion image datasets mostly contain photos only and rarely contain the sketch‐photo pairs. Thus, we contribute a fine‐grained sketch‐based fashion image retrieval dataset, which includes 36,074 sketch‐photo pairs. Specifically, when retrieving on our Fashion Image dataset, the accuracy of our model ranks the correct match at the top‐1 which is 96.6%, 92.1%, 91.0%, and 90.5% for clothes, pants, skirts, and shoes, respectively. Extensive experiments conducted on our dataset and two fine‐grained instance‐level datasets, i.e., QMUL‐shoes and QMUL‐chairs, show that our model has achieved a better performance than other existing methods.
Hao-Peng Lei, Mingwen Wang 0001, Xiangjian He, Wenjing Jia
Wirel. Commun. Mob. Comput.5
2020 Shared subspace least squares multi-label linear discriminant analysis
Wenjing Jia
Appl. Intell.3
2020 FACLSTM: ConvLSTM with focused attention for scene text recognition
Wenjing Jia, Xiangjian He, Michael Blumenstein, Shujing Lyu, Yue Lu 0001
Sci. China Inf. Sci.3
2020 Beyond context: Exploring semantic similarity for small object detection in crowded scenes
Jiangbin Zheng 0001, Xiangjian He, Wenjing Jia, Yefan Xie, Mingchen Feng, Xiuxiu Li
Pattern Recognit. Lett.4
2020 Simultaneous segmentation and correction model for color medical and natural images with intensity inhomogeneity
Yunyun Yang, Wenjing Jia, Boying Wu
Vis. Comput.2
2019 Atrous Convolution for Binary Semantic Segmentation of Lung Nodule
abstract
Accurately estimating the size of tumours and reproducing their boundaries from lung CT images provides crucial information for early diagnosis, staging and evaluating patients response to cancer therapy. This paper presents an advanced solution to segment lung nodules from CT images by employing a deep residual network structure with Atrous convolution. The Atrous convolution increases the field of view of the filters and helps to improve classification accuracy. Moreover, in order to address the significant class imbalance issue between the nodule pixels and background non-nodule pixels, a weighted loss function is proposed. We evaluate our proposed solution on the widely adopted benchmark dataset LIDC. A promising result of an average DCS of 81.24% is achieved, outperforming the state of the arts. This demonstrates the effectiveness and importance of applying the Atrous convolution and weighted loss for such problems.
Mohammad Hesam Hesamian, Wenjing Jia, Xiangjian He, Paul J. Kennedy
ICASSP2
2019 DeepText: Detecting Text from the Wild with Multi-ASPP-Assembled DeepLab
abstract
In this paper, we address the issue of scene text detection in the way of direct regression and successfully adapt an effective semantic segmentation model, DeepLab v3+ [1], for this application. In order to handle texts with arbitrary orientations and sizes and improve the recall of small texts, we propose to extract features of multiple scales by inserting multiple Atrous Spatial Pyramid Pooling (ASPP) layers to the DeepLab after the feature maps with different resolutions. Then, we set multiple auxiliary IoU losses at the decoding stage and make auxiliary connections from the intermediate encoding layers to the decoder to assist network training and enhance the discrimination ability of lower encoding layers. Experiments conducted on the benchmark scene text dataset ICDAR2015 demonstrate the superior performance of our proposed network, named as DeepText, over the state-of-the-art approaches.
Wenjing Jia, Xiangjian He, Yue Lu 0001, Michael Blumenstein, Shujing Lyu
ICDAR2
2019 Targeting malware discrimination based on reversed association task
abstract
Summary Regarding the current situation that the recognition rate of malware is decreasing, the article points out that the reason for this dilemma is that more and more targeting malware have emerged, which share little or no common feature with traditional malware. The premise of malware recognition judging whether a software is malicious or benign is actually a decision problem. We propose that malware discrimination should resort to the corresponding task or purpose. We first present a formal definition of a task and then provide further classifications of malicious tasks. Based on the decidable theory, we prove that task performed by any software is recursive and determinable. By establishing a mapping from software to task, we prove that software is many‐to‐one reducible to corresponding tasks. Thus, we demonstrate that software, including malware, is also recursive and can be determined by the corresponding tasks. Finally, we present the discrimination process of our method. Nine real malwares are presented, which were firstly discriminated by our method but at that time could not be identified by Kaspersky, McAfee, Symantec Norton, or Kingsoft Antivirus.
Lansheng Han, Shuxia Han, Wenjing Jia, Changhua Sun, Cai Fu
Concurr. Comput. Pract. Exp.4
2019 Efficient and robust segmentation and correction model for medical images
abstract
Accurate segmentation of medical images plays a very important role in clinical diagnosis so that the segmentation technology for medical images attracts more and more attention. However, most medical images usually suffer from severe intensity inhomogeneity and make accurate segmentation difficult. In this study, the authors propose an efficient and robust active contour model for simultaneous image segmentation and correction. The proposed model not only can accurately segment images with severe intensity inhomogeneity and serious noise but also can eliminate the intensity varying information to get the homogeneous correction images. They first present the level set formulation of the two‐phase model, which is then extended to the multi‐phase formulation. The split Bregman method is applied to efficiently minimise the proposed energy functionals. The proposed model is tested with lots of synthetic images and medical images with promising results. Experimental results demonstrate that the proposed model can accurately segment and correct the inhomogeneous images with serious noise. Quantitative comparison results of the proposed model and other models illustrate the proposed model is more accurate and more efficient. What's more, the proposed model not only is insensitive to the initial contour, but also is robust to the noise.
Yunyun Yang, Wenjing Jia
IET Image Process.2
2019 Performance-enhancing network pruning for crowd counting
Lei Liu 0036, Saeed Amirgholipour Kasmani, Jie Jiang 0005, Wenjing Jia, Michelle Zeibots, Xiangjian He
Neurocomputing4
2019 Intrusion detection model of wireless sensor networks based on game theory and an autoregressive model
Lansheng Han, Wenjing Jia, Zakaria Dalil, Xingbo Xu
Inf. Sci.3
2019 Multi-atlas segmentation and correction model with level set formulation for 3D brain MR images
abstract
We present an efficient multi-atlas segmentation and correction model with level set formulation for 3D brain MR images in this paper. We define a new energy functional by combining a weighted label fusion term, a bias field based image information fitting term and a regularization term together. More image information is taken into consideration in the new image data term to substantially improve the segmentation accuracy, especially when serious inhomogeneity and bias field exist in regions of interest in MR images. We introduce a spatially weight function and incorporate it into the label fusion term to increase the robustness of our segmentation algorithm to atlases with different registration accuracy. The new energy functional is in the form of L1 regularization problems, and we minimize it with the split Bregman method to ensure the segmentation efficiency. We apply the proposed model to segment six tissues in 3D brain MR images, including the amygdala, caudate, hippocampus, pallidum, putamen and thalamus. Experimental results have shown that our model can segment regions of interest accurately and eliminate bias field simultaneously. Quantitative comparisons with related methods have demonstrated the superiority of our model in terms of accuracy, efficiency and robustness.
Yunyun Yang, Wenjing Jia, Yunna Yang
Pattern Recognit.2
2018 A-CCNN: Adaptive CCNN for Density Estimation and Crowd Counting
abstract
Crowd counting, for estimating the number of people in a crowd using vision-based computer techniques, has attracted much interest in the research community. Although many attempts have been reported, real-world problems, such as huge variation in subjects' sizes in images and serious occlusion among people, make it still a challenging problem. In this paper, we propose an Adaptive Counting Convolutional Neural Network (A-CCNN) and consider the scale variation of objects in a frame adaptively so as to improve the accuracy of counting. Our method takes advantages of contextual information to provide more accurate and adaptive density maps and crowd counting in a scene. Extensively experimental evaluation is conducted using different benchmark datasets for object-counting and shows that the proposed approach is effective and outperforms state-of-the-art approaches.
Saeed Amirgholipour Kasmani, Xiangjian He, Wenjing Jia, Dadong Wang, Michelle Zeibots
ICIP3
2018 Beyond Context: Exploring Semantic Similarity for Tiny Face Detection
abstract
Tiny face detection aims to find faces with high degrees of variability in scale, resolution and occlusion in cluttered scenes. Due to the very little information available on tiny faces, it is not sufficient to detect them merely based on the information presented inside the tiny bounding boxes or their context. In this paper, we propose to exploit the semantic similarity among all predicted targets in each image to boost current face detectors. To this end, we present a novel framework to model semantic similarity as pairwise constraints within the metric learning scheme, and then refine our predictions with the semantic similarity by utilizing the graph cut techniques. Experiments conducted on three widely-used benchmark datasets have demonstrated the improvement over the-state-of-the-arts gained by applying this idea.
Jiangbin Zheng 0001, Xiangjian He, Wenjing Jia
ICIP4
2018 Fast and robust road sign detection in driver assistance systems
Tao Zhang 0010, Wenjing Jia
Appl. Intell.3
2018 Owner based malware discrimination
Lansheng Han, Shuxia Han, Wenjing Jia, Jingwei Lei
Future Gener. Comput. Syst.4
2018 SUDMAD: Sequential and unsupervised decomposition of a multi-author document based on a hidden markov model
abstract
Decomposing a document written by more than one author into sentences based on authorship is of great significance due to the increasing demand for plagiarism detection, forensic analysis, civil law (i.e., disputed copyright issues), and intelligence issues that involve disputed anonymous documents. Among existing studies for document decomposition, some were limited by specific languages, according to topics or restricted to a document of two authors, and their accuracies have big room for improvement. In this paper, we consider the contextual correlation hidden among sentences and propose an algorithm for Sequential and Unsupervised Decomposition of a Multi‐Author Document (SUDMAD) written in any language, disregarding topics, through the construction of a Hidden Markov Model (HMM) reflecting the authors' writing styles. To build and learn such a model, an unsupervised, statistical approach is first proposed to estimate the initial values of HMM parameters of a preliminary model, which does not require the availability of any information of author's or document's context other than how many authors contributed to writing the document. To further boost the performance of this approach, a boosted HMM learning procedure is proposed next, where the initial classification results are used to create labeled training data to learn a more accurate HMM. Moreover, the contextual relationship among sentences is further utilized to refine the classification results. Our proposed approach is empirically evaluated on three benchmark datasets that are widely used for authorship analysis of documents. Comparisons with recent state‐of‐the‐art approaches are also presented to demonstrate the significance of our new ideas and the superior performance of our approach.
Khaled Aldebei, Xiangjian He, Wenjing Jia, Wei-Chang Yeh 0001
J. Assoc. Inf. Sci. Technol.3
2018 Semi-supervised dictionary learning via local sparse constraints for violence detection
Tao Zhang 0010, Wenjing Jia, Chen Gong 0002, Jun Sun 0008, Xiaoning Song
Pattern Recognit. Lett.2
2018 Fast and robust occluded face detection in ATM surveillance
Tao Zhang 0010, Wenjing Jia, Jun Sun 0008
Pattern Recognit. Lett.3
2017 MoWLD: a robust motion image descriptor for violence detection
Tao Zhang 0010, Wenjing Jia, Baoqing Yang, Jie Yang 0002, Xiangjian He, Zhonglong Zheng
Multim. Tools Appl.2
2017 Discriminative Dictionary Learning With Motion Weber Local Descriptor for Violence Detection
abstract
Automatic violence detection from video is a hot topic for many video surveillance applications. However, there has been little success in developing an algorithm that can detect violence in surveillance videos with high performance. In this paper, following our recently proposed idea of motion Weber local descriptor (WLD), we make two major improvements and propose a more effective and efficient algorithm for detecting violence from motion images. First, we propose an improved WLD (IWLD) to better depict low-level image appearance information, and then extend the spatial descriptor IWLD by adding a temporal component to capture local motion information and hence form the motion IWLD. Second, we propose a modified sparse-representation-based classification model to both control the reconstruction error of coding coefficients and minimize the classification error. Based on the proposed sparse model, a class-specific dictionary containing dictionary atoms corresponding to the class labels is learned using class labels of training samples. With this learned dictionary, not only the representation residual but also the representation coefficients become discriminative. A classification scheme integrating the modified sparse model is developed to exploit such discriminative information. The experimental results on three benchmark data sets have demonstrated the superior performance of the proposed approach over the state of the arts.
Tao Zhang 0010, Wenjing Jia, Xiangjian He, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2016 Unsupervised Multi-Author Document Decomposition Based on Hidden Markov Model
abstract
© 2016 Association tor Computational Linguistics. This paper proposes an unsupervised approach for segmenting a multiauthor document into authorial components. The key novelty is that we utilize the sequential patterns hidden among document elements when determining their authorships. For this purpose, we adopt Hidden Markov Model (HMM) and construct a sequential probabilistic model to capture the dependencies of sequential sentences and their authorships. An unsupervised learning method is developed to initialize the HMM parameters. Experimental results on benchmark datasets have demonstrated the significant benefit of our idea and our approach has outperformed the state-of-the-arts on all tests. As an example of its applications, the proposed approach is applied for attributing authorship of a document and has also shown promising results.
Khaled Aldebei, Xiangjian He, Wenjing Jia, Jie Yang 0002
ACL (1)3
2016 A new method for violence detection in surveillance scenes
Tao Zhang 0010, Wenjing Jia, Baoqing Yang, Jie Yang 0002, Xiangjian He
Multim. Tools Appl.3
2015 Sparse coding-based spatiotemporal saliency for action recognition
abstract
In this paper, we address the problem of human action recognition by representing image sequences as a sparse collection of patch-level spatiotemporal events that are salient in both space and time domain. Our method uses a multi-scale volumetric representation of video and adaptively selects an optimal space-time scale under which the saliency of a patch is most significant. The input image sequences are first partitioned into non-overlapping patches. Then, each patch is represented by a vector of coefficients that can linearly reconstruct the patch from a learned dictionary of basis patches. We propose to measure the spatiotemporal saliency of patches using Shannon's self-information entropy, where a patch's saliency is determined by information variation in the contents of the patch's spatiotemporal neighborhood. Experimental results on two benchmark datasets demonstrate the effectiveness of our proposed method.
Tao Zhang 0010, Jie Yang 0002, Wenjing Jia
ICIP5
2015 A New Image Decomposition and Reconstruction Approach - Adaptive Fourier Decomposition
Can He, Liming Zhang 0002, Xiangjian He, Wenjing Jia
MMM (2)4
2015 Fast and robust head detection with arbitrary pose and occlusion
Tao Zhang 0010, Wenjing Jia, Qiang Wu 0001, Jie Yang 0002, Xiangjian He
Multim. Tools Appl.3
2014 Unsupervised Segmentation Using Cluster Ensembles
Wei Zhang 0362, Jie Yang 0002, Wenjing Jia, Nikola K. Kasabov, Zhenhong Jia, Lei Zhou 0003
ICONIP (3)3
2014 Agent-Based Modeling of Oxygen-Responsive Transcription Factors in Escherichia coli
abstract
In the presence of oxygen (O2) the model bacterium Escherichia coli is able to conserve energy by aerobic respiration. Two major terminal oxidases are involved in this process - Cyo has a relatively low affinity for O2 but is able to pump protons and hence is energetically efficient; Cyd has a high affinity for O2 but does not pump protons. When E. coli encounters environments with different O2 availabilities, the expression of the genes encoding the alternative terminal oxidases, the cydAB and cyoABCDE operons, are regulated by two O2-responsive transcription factors, ArcA (an indirect O2 sensor) and FNR (a direct O2 sensor). It has been suggested that O2-consumption by the terminal oxidases located at the cytoplasmic membrane significantly affects the activities of ArcA and FNR in the bacterial nucleoid. In this study, an agent-based modeling approach has been taken to spatially simulate the uptake and consumption of O2 by E. coli and the consequent modulation of ArcA and FNR activities based on experimental data obtained from highly controlled chemostat cultures. The molecules of O2, transcription factors and terminal oxidases are treated as individual agents and their behaviors and interactions are imitated in a simulated 3-D E. coli cell. The model implies that there are two barriers that dampen the response of FNR to O2, i.e. consumption of O2 at the membrane by the terminal oxidases and reaction of O2 with cytoplasmic FNR. Analysis of FNR variants suggested that the monomer-dimer transition is the key step in FNR-mediated repression of gene expression.
Matthew D. Rolfe, Wenjing Jia, Simon Coakley, Robert K. Poole, Jeffrey Green, Mike Holcombe
PLoS Comput. Biol.3
2014 Characterness: An Indicator of Text in the Wild
abstract
Text in an image provides vital information for interpreting its contents, and text in a scene can aid a variety of tasks from navigation to obstacle avoidance and odometry. Despite its value, however, detecting general text in images remains a challenging research problem. Motivated by the need to consider the widely varying forms of natural text, we propose a bottom-up approach to the problem, which reflects the characterness of an image region. In this sense, our approach mirrors the move from saliency detection methods to measures of objectness. In order to measure the characterness, we develop three novel cues that are tailored for character detection and a Bayesian method for their integration. Because text is made up of sets of characters, we then design a Markov random field model so as to exploit the inherent dependencies between characters. We experimentally demonstrate the effectiveness of our characterness cues as well as the advantage of Bayesian multicue integration. The proposed text detector outperforms state-of-the-art methods on a few benchmark scene text detection data sets. We also show that our measurement of characterness is superior than state-of-the-art saliency detection models when applied to the same task.
Yao Li 0003, Wenjing Jia, Chunhua Shen, Anton van den Hengel
IEEE Trans. Image Process.2
2013 Text detection in born-digital images using multiple layer images
abstract
In this paper, a new framework for detecting text from webpage and email images is presented. The original image is split into multiple layer images based on the maximum gradient difference (MGD) values to detect text with both strong and weak contrasts. Connected component processing and text detection are performed in each layer image. A novel texture descriptor named T-LBP, is proposed to further filter out non-text candidates with a trained SVM classifier. The ICDAR 2011 born-digital image dataset is used to evaluate and demonstrate the performance of the proposed method. Following the same performance evaluation criteria, the proposed method outperforms the winner algorithm of the ICDAR 2011 Robust Reading Competition Challenge 1.
Wenjing Jia, Xiangjian He
ICASSP2
2013 Leveraging surrounding context for scene text detection
abstract
Finding text in natural images has been a challenging task in vision. At the core of state-of-the-art scene text detection algorithms are a set of text-specific features within extracted regions. In this paper, we attempt to solve this problem from a different prospective. We show that characters and non-character interferences are separable by leveraging the surrounding context. Surrounding context, in our work, is composed of two components which are computed in an information-theoretic fashion. Minimization of an energy cost function yields a binary label for each region, which indicates the category it belongs to. The proposed algorithm is fast, discriminative and tolerant to character variations and involves minimal parameter tuning.
Yao Li 0003, Chunhua Shen, Wenjing Jia, Anton van den Hengel
ICIP3
2012 Efficient Super-Resolution by Finer Sub-Pixel Motion Prediction and Bilateral Filtering
abstract
Super-resolution reconstruction produces high-resolution images from a set of low-resolution images of the same scene. In the last two and a half decades, many super-resolution algorithms have been proposed. These algorithms are very sensitive to their assumed models of motion and noise, and computationally expensive for many practical applications. In this paper we adopt earlier reported fast prediction based sub-pixel motion estimation and a novel interpolation scheme based on the bilateral filter to produce a fast color super-resolution reconstruction that can accommodate arbitrary local motion patterns. The proposed algorithm exploits photometric proximity and available finer fractional motion information in the high resolution grid, to reconstruct enhanced super-resolved image frames. Experiments show a PSNR performance comparable to the state-of-the-art but at a fraction of their computational cost.
Damith J. Mudugamuwa, Xiangjian He, Wenjing Jia
ICME3
2012 Battle-Lemarie wavelet pyramid for improved GSM image denoising
Damith J. Mudugamuwa, Xiangjian He, Wenjing Jia
ICPR3
2011 An overcomplete pyramid representation for improved gsm image denoising
abstract
Removing noise from a digital image is a challenging problem. Application of Gaussian Scale Mixtures (GSM) in the wavelet domain has been reported to be one of the most effective denoising algorithms, published to date. In this paper we investigate the impact of overcomplete wavelet image representations on the GSM image denoising algorithm. We explore the desirable local characteristics of wavelet coefficients that can enhance the efficiency of GSM denoising and based on the findings, we devise an improved over-complete pyramid representation to enhance the GSM denoising performance. We present the experimental denoising results using the proposed pyramid representation, and they outperform state-of-the-art GSM denoising results reported in the literature.
Damith J. Mudugamuwa, Wenjing Jia, Xiangjian He, Jie Yang 0002
ICME2
2011 Learning Global and Local Features for License Plate Detection
Sheng Wang 0003, Wenjing Jia, Qiang Wu 0001, Xiangjian He, Jie Yang 0002
ICONIP (3)2
2011 Facial Expression Recognition on Hexagonal Structure Using LBP-Based Histogram Variances
Xiangjian He, Ruo Du, Wenjing Jia, Qiang Wu 0001, Wei-Chang Yeh 0001
MMM (2)4
2011 More on Weak Feature: Self-correlate Histogram Distances
Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Wenjing Jia
PSIVT (1)4
2010 Canny Edge Detection Using Bilateral Filter on Real Hexagonal Structure
Xiangjian He, Daming Wei, Kin-Man Lam 0001, Wenjing Jia, Qiang Wu 0001
ACIVS (1)6
2010 ECCH: A novel color coocurrence histogram
abstract
In this paper, a novel color cooccurrence histogram method, named eCCH which stands for color cooccurrence histogram at edge points, is proposed to describe the spatial-color joint distribution of images. Unlike all existing ideas, we only investigate the color distribution of pixels located at the two sides of edge points on gradient direction lines. When measuring the similarity of two eCCHs, the Gaussian weighted histogram intersection method is adopted, where both identical and similar color pairs are considered to compensate color variations. Comparative experimental results demonstrate the performance of the proposed eCCH in terms of robustness to color variance and small computational complexity.
Wenjing Jia, Xiangjian He, Qiang Wu 0001
ICASSP1
2010 A Two-Tier System for Web Attack Detection Using Linear Discriminant Method
Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001, Wenjing Jia, Wei-Chang Yeh 0001
ICICS6
2009 Facial expression recognition using histogram variances faces
abstract
In human's expression recognition, the representation of expression features is essential for the recognition accuracy. In this work we propose a novel approach for extracting expression dynamic features from facial expression videos. Rather than utilising statistical models e.g. Hidden Markov Model (HMM), our approach integrates expression dynamic features into a static image, the Histogram Variances Face (HVF), by fusing histogram variances among the frames in a video. The HVFs can be automatically obtained from videos with different frame rates and immune to illumination interference. In our experiments, for the videos picturing the same facial expression, e.g., surprise, happy and sadness etc., their corresponding HVFs are similar, even though the performers and frame rates are different. Therefore the static facial recognition approaches can be utilised for the dynamic expression recognition. We have applied this approach on the well-known Cohn-Kanade AU-Coded Facial Expression database then classified HVFs using PCA and Support Vector Machine (SVMs), and found the accuracy of HVFs classification is very encouraging.
Ruo Du, Qiang Wu 0001, Xiangjian He, Wenjing Jia, Daming Wei
WACV4
2008 An approach of canny edge detection with virtual hexagonal image structure
abstract
Edge detection plays an important role in the areas of image processing, multimedia and computer vision. Gradient-based edge detection is a straightforward method to identify the edge points in the original grey-level image. It is intuitive that, in the human vision system, the edge points always appear where the gradient magnitude assumes a maximum. Hexagonal structure is an image structure alternative to traditional square image structure. The geometrical arrangement of pixels on a hexagonal structure can be described as a collection of hexagonal pixels. Because all the existing hardware for capturing image and for displaying image are produced based on square structure, an approach that uses bilinear interpolation and tri-linear interpolation is applied for conversion between square and hexagonal structures. Based on this approach, an edge detection method is proposed. This method performs Gaussian filtering to suppress image noise and computes gradients on the hexagonal structure. The pixel edge strengths on the square structure are then estimated before Canny' edge detector is applied to determine the final edge map. The experimental results show that the proposed method improves the edge detection accuracy and efficiency.
Xiangjian He, Wenjing Jia, Qiang Wu 0001
ICARCV2
2008 Segmentation of characters on car license plates
abstract
License plate recognition usually contains three steps, namely license plate detection/localization, character segmentation and character recognition. When reading characters on a license plate one by one after license plate detection step, it is crucial to accurately segment the characters. The segmentation step may be affected by many factors such as license plate boundaries (frames). The recognition accuracy will be significantly reduced if the characters are not properly segmented. This paper presents an efficient algorithm for character segmentation on a license plate. The algorithm follows the step that detects the license plates using an AdaBoost algorithm. It is based on an efficient and accurate skew and slant correction of license plates, and works together with boundary (frame) removal of license plates. The algorithm is efficient and can be applied in real-time applications. The experiments are performed to show the accuracy of segmentation.
Xiangjian He, Lihong Zheng, Qiang Wu 0001, Wenjing Jia, Bijan Samali, Marimuthu Palaniswami
MMSP4
2007 Parallel Edge Detection on a Virtual Hexagonal Structure
Xiangjian He, Wenjing Jia, Qiang Wu 0001, Tom Hintz
GPC2
2007 Local Binary Patterns for Human Detection on Hexagonal Structure
abstract
Local binary pattern (LBP) was designed and has been widely used for efficient texture classification. LBP provides a simple and effective way to represent texture patterns. Uniform LBPs play an important role for LBP-based pattern/object recognition as they include majority of LBPs. On the other hand, Human detection based on Mahalanobis distance map (MDM) recognizes appearance of human based on geometrical structure. Each MDM shows a clear texture pattern that can be classified using LBPs. In this paper, we compute LBPs of MDMs on a hexagonal structure. The circular pixel arrangement in hexagonal structure results in higher accuracy for LBP representation than on square structure. Chi-square as a measure is used for human detection based on uniform LBPs obtained. We show that our method using LBPs built on MDMs has a higher human detection rate and a lower false positive rate compared to the method merely based on MDMs. We will also show using experimental results that LBPs on hexagonal structure lead to more robust human classification.
Xiangjian He, Yan Chen 0020, Qiang Wu 0001, Wenjing Jia
ISM5
2007 Region-based license plate detection
Wenjing Jia, Huaifeng Zhang, Xiangjian He
J. Netw. Comput. Appl.1
2006 Symmetric Color Ratio in Spiral Architecture
Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001
ACCV (2)1
2006 A Comparison on Histogram Based Image Matching Methods
abstract
Using colour histogram as a stable representation over change in view has been widely used for object recognition. In this paper, three newly proposed histogram-based methods are compared with other three popular methods, including conventional histogram intersection (HI) method, Wong and Cheung's merged palette histogram matching (MPHM) method, and Gevers' colour ratio gradient (CRG) method. These methods are tested on vehicle number plate images for number plate classification. Experimental results disclose that, the CRG method is the best choice in terms of speed, and the GWHI method can give the best classification results. Overall, the CECH method produces the best performance when both speed and classification performance are concerned.
Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001
AVSS1
2006 Car Plate Detection Using Cascaded Tree-Style Learner Based on Hybrid Object Features
abstract
Car plate detection is a key component in automatic license plate recognition system. This paper adopts an enhanced cascaded tree style learner framework for car plate detection using the hybrid object features including the simple statistical features and Harr-like features. The statistical features are useful for simplifying the process on cascade classifier. The cascaded tree-style detector design will further reduce the false alarm and the false dismissal while retaining a high detection ratio. The experimental results obtained by the proposed algorithm exhibit the encouraging performance.
Qiang Wu 0001, Huaifeng Zhang, Wenjing Jia, Xiangjian He, Jie Yang 0002, Tom Hintz
AVSS3
2006 Uniformly Partitioning Images on Virtual Hexagonal Structure
abstract
Hexagonal structure is different from the traditional square structure for image representation. The geometrical arrangement of pixels on hexagonal structure can be described in terms of a hexagonal grid. Uniformly separating image into seven similar copies with a smaller scale has commonly been used for parallel and accurate image processing on hexagonal structure. However, all the existing hardware for capturing image and for displaying image are produced based on square architecture. It has become a serious problem affecting the advanced research based on hexagonal structure. Furthermore, the current techniques used for uniform separation of images on hexagonal structure do not coincide with the rectangular shape of images. This has been an obstacle in the use of hexagonal structure for image processing. In this paper, we briefly review a newly developed virtual hexagonal structure that is scalable. Based on this virtual structure, algorithms for uniform image separation are presented. The virtual hexagonal structure retains image resolution during the process of image separation, and does not introduce distortion. Furthermore, images can be smoothly and easily transferred between the traditional square structure and the hexagonal structure while the image shape is kept in rectangle
Xiangjian He, Huaqing Wang, Namho Hur, Wenjing Jia, Qiang Wu 0001, Jinwoong Kim, Tom Hintz
ICARCV4
2006 Image Matching Using Colour Edge Cooccurrence Histograms
abstract
In this paper, a novel colour edge cooccurrence histogram (CECH) method is proposed to match images by measuring similarities between their CECH histograms. Unlike the previous colour edge cooccurrence histogram proposed by Crandall and Luo (2004 ) we only investigate those pixels which are located at the two sides of edge points in their gradient direction lines and at a distance away from the edge points. When measuring similarities between two CECH histograms, a newly proposed Gaussian weighted histogram intersection (GWHI) method is extended for this purpose. Both identical colour pairs and similar colour pairs are taken into account in our algorithm, and the weights are decided by the larger distance between two colour pairs involved in matching. The proposed algorithm is tested for matching vehicle number plate images captured under various illumination conditions. Experimental results demonstrate that the proposed algorithm can be used to compare images in real-time, and is robust to illumination variations and insensitive to the model images selected.
Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001
SMC1
2006 A Fast Algorithm for License Plate Detection in Various Conditions
abstract
This paper proposes a fast algorithm detecting license plates in various conditions. There are three main contributions in this paper. The first contribution is that we define a new vertical edge map, with which the license plate detection algorithm is extremely fast. The second contribution is that we construct a cascade classifier which is composed of two kinds of classifiers. The classifiers based on statistical features decrease the complexity of the system. They are followed by the classifiers based on Haar-features, which make it possible to detect license plate in various conditions. Our algorithm is robust to the variance of the illumination, view angle, the position, size and color of the license plates when working in complex environment. The third contribution is that we experimentally analyze the relations of the scaling factor with detection rate and processing time. On the basis of the analysis, we select the optimal scaling factor in our algorithm. In the experiments, both high detection rate (with low false positive rate) and high speed are achieved when the algorithm is used to detect license plates in various complex conditions.
Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001
SMC2
2006 Real-Time License Plate Detection Under Various Conditions
Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001
UIC2
2005 Modified Color Ratio Gradient
abstract
Color ratio gradient is an efficient method used for color image retrieval and object recognition, which is shown to be illumination-independent and geometry-insensitive when tested on scenery images. However, color ratio gradient produces unsatisfied matching result while dealing with relatively uniform objects without rich color texture. In addition, performance of color ratio gradient degenerates while processing unsaturated color image objects. In this paper, a scheme with modified color ratio gradient is presented, which addresses the two problems above. Experimental results using the proposed method in this paper exhibit more robust performance
Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001
MMSP2