Jiaxu Leng

dblp:219/1664 · DBLP profile ↗
← Back
72ranked-venue papers
25as first author
65since 2021 · last 2026
0000-0003-2802-8139ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 14 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic-Static Collaboration for Unsupervised Domain Adaptive Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) aims to match pedestrian sequences across modalities for all-day surveillance. While supervised methods have shown progress, their dependence on large-scale cross-modal annotations limits scalability. We investigate the task of unsupervised domain adaptation for VVI-ReID (UDA-VVI-ReID), where a model trained on a labeled source domain is adapted to an unlabeled target domain. Directly extending existing image-based unsupervised VI-ReID methods to video scenarios by simply averaging frame-level features is suboptimal, as this naive strategy neglects the rich temporal dynamics in video data and leads to unreliable pseudo-labels due to occlusion-induced noise. To overcome these limitations, we propose a Dynamic-Static Collaboration (DSC) framework that explicitly leverages the complementary strengths of motion and appearance cues. The Dynamic-Static Label Unification (DSLU) module refines pseudo-labels by validating the consistency between static and dynamic predictions. Based on these labels, the Dynamic-Static Joint Learning (DSJL) module performs neighbor-aware contrastive learning in both feature spaces, promoting robust representation learning under cross-modal and temporal variations. Experiments on HITSZ-VCM and BUPTCampus show that DSC sets a strong baseline for this new task, enabling robust cross-modal video ReID without target labels.
Jiaxu Leng, Xinbo Gao 0001
AAAI1
2026 PortraitSR: Artist-Inspired Prior Learning for Progressive Face Super-Resolution
abstract
Face super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from low-resolution (LR) inputs. While recent methods have advanced this task through architectural innovations and generative modeling, but they often leads to semantically inconsistent structures and unrealistic textures, particularly under high magnification. To mitigate these limitations, we draw inspiration from the human artistic process of “structuring before detailing” and propose a progressive prior-guided restoration strategy. Specifically, we first introduce a Sketching Structure Prior (SSP) module that embeds global semantics and refines local geometry through implicit parsing guidance and explicit spatial modulation. Then, an Associative Texture Prior (ATP) module leverages a High-Quality Dictionary (HD) learned from high-quality reconstruction to guide fine-grained detail recovery. Finally, to unify structure and detail features, we design a Holistic Prior Fusion (HPF) module that adaptively integrates them within semantically consistent facial regions. Our method surpasses state-of-the-art on CelebA and Helen in both structural fidelity and texture realism.
Miaoqing Wang, Jiaxu Leng, Changjiang Kuang
AAAI2
2026 Retrieval-Guided Contextual Inference for Training-Free Video Anomaly Detection in Low-Light Scenarios
abstract
Real-world surveillance often operates in low-light environments, where degraded visual evidence can make training-free anomaly reasoning unreliable. However, current training-free methods typically assume sufficiently clean inputs, which can lead to hallucinated semantics and unstable anomaly scores under visual degradation. To address this issue, we propose Retrieval-augmented Contextual Inference (ReCI), a training-free framework that leverages retrieval-augmented context for robust low-light video anomaly detection. ReCI constructs Semantic Context (SC) through hierarchical captioning by aggregating clip-level local captions into a video-level global description. It then performs contextualized anomaly inference using Reference Context (RC) retrieved from a reference pool built during inference. The VLM outputs anomaly scores and associated confidence values, which we use as a heuristic reliability signal for temporal refinement. Experiments on XD-Violence and UCF-Crime show that ReCI consistently improves over prior training-free baselines, with particularly clear gains on the low-light subset.
Mengjingcheng Mo, Jiankang Zheng, Jiaxu Leng, Xinbo Gao 0001
ICMR3
2026 Future tells the goal: Future semantic learning for unsupervised video anomaly prediction
Mingpi Tan, Jiaxu Leng, Zhanjie Wu, Mengjingcheng Mo, Xinbo Gao 0001
Neurocomputing2
2026 Revisiting multi-scale feature representation and fusion for UAV-based road distress detection
Peng Wang 0151, Jiamei Liu, Haofeng Chen, Jiaxu Leng, Gang Ma 0008, Wanjing Ma
Neurocomputing4
2026 TumorAL: Evidence-aware active learning for 3D tumor segmentation
Hongyi Wang 0006, Jiaxu Leng, Yue Zhao 0012, Weikai Li 0003, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing2
2026 Transformer-Based 3-D Hand Pose Estimation via Bidirectional Multiscale Fusion and Learnable Anchor Guidance
abstract
Accurate 3D hand pose estimation faces inherent challenges, including self-occlusions, joint similarities, and high degrees of freedom. Although most existing CNN-based or Transformer-based methods leverage global contexts, they often fail to capture fine-grained local details and robustly handle occluded joints. To address these limitations, we propose a novel Transformer framework for depth-based 3D hand pose estimation, which incorporates two key designs: the Bidirectional Multiscale Fusion and Learnable Anchor Guidance. Firstly, we propose bidirectional multiscale fusion that sequentially propa-gates features from different encoder levels in both top-down and bottom-up directions, followed by a final aggregation of all scale features. Such an intricate feature interaction eventually enables effective joint modeling of fine-grained local details (e.g., fingertip positions) and high-level semantic context information (e.g., palm orientation). Secondly, we introduce the learnable anchor query as prior guidance to dynamically guide the decoder to localize ambiguous joints better. The learnable anchors are derived from joint-specific attention maps under 3D ground-truth supervision and then are concatenated with static grid anchors to form hybrid anchors, which effectively enable more precise 3D hand pose estimation, especially for occlusions. To further demonstrate the practicality of our framework for IoT-oriented deployment, we conduct edge-device experiments to validate its deployment feasibility. Experiments on benchmark datasets (including NYU, ICVL, MSRA, and DexYCB) demonstrate the superiority of our method over previous state-of-the-art approaches.
Ji Gan, Weiqiang Wang 0001, Feng Gao 0005, Jiaxu Leng, Haosheng Chen 0001, Xinbo Gao 0001
IEEE Internet Things J.5
2026 A spatial-frequency hybrid restoration network for JPEG compressed image deblurring
Shu Tang, Xinbo Gao 0001, Shuli Yang, Jiaxu Leng, Zengdan Pan
Neural Networks5
2026 PiercingEye: Dual-Space Video Violence Detection With Hyperbolic Vision-Language Guidance
abstract
Existing weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet semantically distinct events due to limited hierarchical modeling and insufficient ambiguous training samples. To address this challenge, we propose PiercingEye, a novel dual-space learning framework that synergizes Euclidean and hyperbolic geometries to enhance discriminative feature representation. Specifically, PiercingEye introduces a layer-sensitive hyperbolic aggregation strategy with hyperbolic Dirichlet energy constraints to progressively model event hierarchies, and a cross-space attention mechanism to facilitate complementary feature interactions between Euclidean and hyperbolic spaces. Furthermore, to mitigate the scarcity of ambiguous samples, we leverage large language models to generate logic-guided ambiguous event descriptions, enabling explicit supervision through a hyperbolic vision-language contrastive loss that prioritizes high-confusion samples via dynamic similarity-aware weighting. Extensive experiments on XD-Violence and UCF-Crime benchmarks demonstrate that PiercingEye achieves state-of-the-art performance, with particularly strong results on a newly curated ambiguous event subset, validating its superior capability in fine-grained violence detection.
Jiaxu Leng, Zhanjie Wu, Mingpi Tan, Mengjingcheng Mo, Jiankang Zheng, Ji Gan, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 A multi-scale gate network for high-quality image deblurring
Shu Tang, Yufu Lin, Xinbo Gao 0001, Shuli Yang, Jiaxu Leng, Zengdan Pan
Pattern Recognit.5
2026 Efficient bidirectional fusion multi-gate network for lightweight single image super-resolution
Shuli Yang, Shu Tang, Xinbo Gao 0001, Jiaxu Leng, Xianzhong Xie
Pattern Recognit.4
2026 HandJoKe: Joint-Guided Keypoint Denoising Transformer for Depth-Based 3D Hand Pose Estimation
abstract
Existing depth-based 3D hand pose estimation methods typically estimate hand joints from either 2D depth images or 3D point clouds, whereas the approaches that fuse multimodal data remain underexplored. Furthermore, previous methods often struggle to learn geometric-facilitated features and precise joint correlations, especially for occluded hands, due to the lack of explicit prior guidance and insufficient cross-dimensional interaction. By taking advantage of multi-modal fusion, cross-dimensional interaction, and prior guidance, we propose a novel joint-guided keypoint denoising Transformer (named HandJoKe) to achieve more precise hand pose estimation, which can iteratively estimate hand poses based on keypoint features from both 2D depth images and 3D point clouds under explicit joint guidance within only several denoising steps. Rather than directly applying existing multi-modal fusion to perform redundant interactions among many background pixels and irrelevant points, HandJoKe focuses on modeling correlations and capturing dependencies among local informative hand regions (i.e., keypoints), thus attaining higher learning capability with lower computation redundancy. Moreover, a novel joint-guided denoising estimation strategy is introduced to adequately fuse cross-modal keypoint features under explicit joint guidance, achieving geometric-facilitated cross-modal keypoint interaction in both 2D and 3D spaces. The effectiveness of joint guidance can be further strengthened through iterative denoising, since it can subsequently update cross-modal keypoint features based on previous denoised hand poses and thus can help better locate confused joints, especially for occluded hands. Extensive experiments show that HandJoKe has achieved state-of-the-art performance on four public challenging benchmarks, including single-hand datasets NYU and ICVL, and hand-object datasets DexYCB and HO3D.
Ji Gan, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 A Lightweight Frequency-Selection-Based Progressive Patch Transformer Network for Single Image Super-Resolution
abstract
Recently, lightweight networks for single image super-resolution (SISR) have surged due to the need of resource-constrained devices, where divide-and-conquer multi-route model exhibits impressive trade-off between performance and computational cost. However, most existing divide-and-conquer multi-route models face two key limitations: (1) possible suboptimal decoupling of image components (e.g. smooth regions, edges and texture details) due to spatial-domain-only processing, and (2) inability to model global dependencies explicitly and capture structural information, hindering further performance gains. To address these drawbacks, we propose a lightweight frequency-selection-based progressive patch Transformer network (FSPPTN) for higher-quality SISR reconstruction. Specifically, we first propose a frequency selection module, in which we develop a frequency enhancement branch (FEB) to dynamically decouple different image components by introducing the window-based Fast Fourier transform (WFFT) and a learnable weight matrix, and a spatial restoration branch (SRB) to recalibrate and fuse cross-granularity features by designing a multi-gate mechanism for reconstructing the component information screened out by the FEB at current level. Secondly, we propose a lightweight multi-branch gradient-guided inter-patch self-attention to explicitly capture global structural similarities by summarizing structural information of each patch into a lower-dimensional space using the statistical properties of first-order gradients, thereby achieving explicit global dependencies modeling and lightweight. Extensive experimental results demonstrate that, in the vast majority of cases, FSPPTN outperforms state-of-the-art lightweight SISR methods in terms of both performance and computational overhead, especially for ×3 and ×4 SR, e.g. FSPPTN outperforms MaIR-Small by even 0.14dB PSNR on Manga109 dataset for ×4 SR even with 48.3% fewer parameters and 63.5% lower FLOPs. The code is available at: https://github.com/yslyangshuli/FSPPTN-main.
Shuli Yang, Shu Tang, Xinbo Gao 0001, Xianzhong Xie, Jiaxu Leng
IEEE Trans. Circuits Syst. Video Technol.5
2026 Causal Bootstrapped Alignment for Unsupervised Video-Based Visible-Infrared Person Re-Identification
Jiaxu Leng, Changjiang Kuang, Mingpi Tan, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2026 Hierarchical Causal Learning for Face Age Synthesis
abstract
Face age synthesis (FAS) predicts a person's future or past facial appearance. In FAS, modifying one facial attribute usually affects the generation of other attributes during face image generation. Current models directly learn entangled representations of age-related features, resulting in insufficient feature disentanglement, which consequently impairs their causal reasoning capability for FAS tasks. To this end, we propose a hierarchical causal learning model for face age synthesis (HCFace), which integrates hierarchical structures and causal relationships into the facial generative model. Specifically, we propose to leverage hierarchical causal relationships to align with facial features for feature disentanglement. Furthermore, we design a novel nonlinear mapping function that captures the true patterns of facial attribute changes with age, enhancing the disentanglement of these attributes. We conduct extensive experiments to validate the superiority of our proposed model. Compared to other advanced baseline methods, HCFace improves overall accuracy by 2.47%, with improvements of 9.75% and 9.69% in certain age-related attributes, such as skin and hair. Our source code is available at https://github.com/SE-hash/HCFace.
Ye Wang 0006, Pan Sun, Lifeng Shen, Jiaxu Leng, Guoyin Wang 0001, Hong Yu 0007
IEEE Trans. Image Process.5
2026 Transferring and Refining Visual-Semantic Priors via Graph-Enhanced CLIP for 3D Hand Pose Estimation
abstract
3D hand pose estimation is crucial for many human-computer interaction applications. However, existing deep neural networks (DNNs) for 3D hand pose estimation suffer from poor generalizability due to data scarcity and a lack of domain-specific knowledge. In contrast, humans remain far better than DNNs at learning; Humans require fewer samples for learning new concepts under the guidance of their prior knowledge. Inspired by this, we propose a graph-enhanced CLIP to deliver visual-semantic priors to DNNs, and provide refined domain-specific knowledge for better 3D hand pose estimation. Specifically, we first introduce a pre-trained CLIP to guide the hand estimation model in learning the semantic-aware visual features, and text-free contrastive learning is proposed to effectively transfer high-level visual-semantic priors from the pre-trained large multimodal models. Notably, our strategy is data-agnostic and avoids designing hand-crafted text prompts for various visual inputs. Second, we introduce novel graph Transformers to refine the domain-specific knowledge by fully exploiting the local adjacent relations of hand joints and capturing the global structure representations of hand poses. The introduced graph Transformers are supposed to further refine the generalized CLIP feature for the downstream task (i.e., hand pose estimation) with better performance. Experiments show that our proposed graph-enhanced CLIP achieves state-of-the-art performances on benchmark datasets, demonstrating its effectiveness for 3D hand pose estimation. The source code is available athttps://github.com/TLiu2832/TRVSP-GE-CLIP.
Ji Gan, Jiaxu Leng, Xinbo Gao 0001
IEEE Trans. Multim.3
2026 Hybrid supervised learning for enhanced open-world video anomaly detection
Weijie Gao, Jiaxu Leng, Xiangqi Meng, Changhe Tu
Vis. Comput.2
2025 Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling
abstract
The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images. To tackle these issues, a novel Bidirectional Reference Image Quality Assessment (BRIQA) method is proposed, centering on leveraging bidirectional reference images and content-quality correlation modeling. First, triplets of content-aligned low-quality and content-non-aligned high-quality reference images are generated using two easily accessible approaches. To prevent the extraction of redundant information, two feature extractors pretrained through unsupervised contrastive learning are utilized to independently extract content and quality features for the triplet images. Then, an attention-mixer is introduced to further mine quality difference information and enhance content feature. Finally, a content-quality correlation modeler is proposed to model the relationship between quality differences and visual contents. Experimental results on benchmark datasets demonstrate that the BRIQA outperforms existing state-of-the-art methods.
Bo Hu 0008, Wenzhi Chen, Chunyi Li 0001, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001
ICASSP4
2025 Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual Learning
abstract
Existing handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to fully process handwritten text images alongside their corresponding geometric graphs, which consists of one shared cross-modal encoder and two parallel inverse decoders. Specifically, the encoder simultaneously extracts visual and geometric information from the cross-modal inputs, and the decoders fuse the multi-modal features for prediction under the guidance of cross-modal fusion. Moreover, two parallel decoders sequentially aggregate cross-modal features in inverse orders (V→G and G→V) but are enhanced through mutual distillation at each time-step, which involves one-to-one knowledge transfer and fully leverages complementary cross-modal information from both directions. Notably, only one branch of GCM is activated in inference, thus avoiding the increase of the model parameters and computation costs for testing. Experiments show that our method outperforms previous state-of-the-art methods on public benchmarks such as IAM, RIMES, and ICDAR-2013 when no extra training data is utilized.
Ji Gan, Yupeng Zhou, Jiaxu Leng, Xinbo Gao 0001
IJCAI4
2025 DichotomyIR: Universal Image Reconstruction via Dichotomy Classification and Uncertainty Elimination
Yan Zhang 0108, Shiwen He, Lin Yuan 0002, Jiaxu Leng, Xinbo Gao 0001
ACM Multimedia4
2025 A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
abstract
While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view scenarios.To bridge this gap, we introduce A2Seek (Aerial Anomaly Seek), a large-scale, reasoning-centric benchmark dataset for aerial anomaly understanding. This dataset covers various scenarios and environmental conditions, providing high-resolution real-world aerial videos with detailed annotations, including anomaly categories, frame-level timestamps, region-level bounding boxes, and natural language explanations for causal reasoning. Building on this dataset, we propose A2Seek-R1, a novel reasoning framework that generalizes R1-style strategies to aerial anomaly understanding, enabling a deeper understanding of “Where” anomalies occur and “Why” they happen in aerial frames.To this end, A2Seek-R1 first employs a graph-of-thought (GoT)-guided supervised fine-tuning approach to activate the model's latent reasoning capabilities on A2Seek. Then, we introduce Aerial Group Relative Policy Optimization (A-GRPO) to design rule-based reward functions tailored to aerial scenarios. Furthermore, we propose a novel “seeking” mechanism that simulates UAV flight behavior by directing the model's attention to informative regions.Extensive experiments demonstrate that A2Seek-R1 achieves up to a 22.04\% improvement in AP for prediction accuracy and a 13.9\% gain in mIoU for anomaly localization, exhibiting strong generalization across complex environments and out-of-distribution scenarios. Our dataset and code are released at https://2-mo.github.io/A2Seek/.
Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng, Jiankang Zheng, Haosheng Chen 0001, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001
NeurIPS4
2025 Dual-Space Video Person Re-identification
Jiaxu Leng, Changjiang Kuang, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001
Int. J. Comput. Vis.1
2025 Ranking-based adaptive query generation for DETRs in crowded pedestrian detection
Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001
Neurocomputing2
2025 RC-DETR: Improving DETRs in crowded pedestrian detection via rank-based contrastive learning
Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001
Neural Networks2
2025 GCapNet-FSD: A heterogeneous Graph Capsule Network for Few-Shot object Detection
Jiaxu Leng, Qianru Chen, Taiyue Chen, Feng Gao 0005, Ji Gan, Changjun Gu, Xinbo Gao 0001
Neural Networks1
2025 Shape-centered representation learning for visible-infrared person re-identification
Jiaxu Leng, Ji Gan, Mengjingcheng Mo, Xinbo Gao 0001
Pattern Recognit.2
2025 Peer Is Your Pillar: A Data-Unbalanced Conditional GANs for Few-Shot Image Generation
abstract
Few-shot image generation aims to train generative models using a small number of training images. When there are few images available for training (e.g. 10 images), Learning From Scratch (LFS) methods often generate images that closely resemble the training data while Transfer Learning (TL) methods try to improve performance by leveraging prior knowledge from GANs pre-trained on large-scale datasets. However, current TL methods may not allow for sufficient control over the degree of knowledge preservation from the source model, making them unsuitable for setups where the source and target domains are not closely related. To address this, we propose a novel pipeline called Peer is your Pillar (PIP), which combines a target few-shot dataset with a peer dataset to create a data-unbalanced conditional generation. Our approach includes a class embedding method that separates the class space from the latent space, and we use a direction loss based on pre-trained CLIP to improve image diversity. Experiments on various few-shot datasets demonstrate the advancement of the proposed PIP, especially reduces the training requirements of few-shot image generation.
Ziqiang Li 0001, Xue Rui, Jiaxu Leng, Zhangjie Fu 0001, Bin Li 0025
IEEE Trans. Circuits Syst. Video Technol.5
2025 Video-Level Language-Driven Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based Visible-Infrared Person Re-Identification (VVI-ReID) aims to match pedestrian sequences across modalities by extracting modality-invariant sequence-level features. As a high-level semantic representation, language provides a consistent description of pedestrian characteristics in both infrared and visible modalities. Leveraging the Contrastive Language-Image Pre-training (CLIP) model to generate video-level language prompts and guide the learning of modality-invariant sequence-level features is theoretically feasible. However, the challenge of generating and utilizing modality-shared video-level language prompts to address modality gaps remains a critical problem. To address this problem, we propose a simple yet powerful framework, video-level language-driven VVI-ReID (VLD), which consists of two core modules: invariant-modality language prompting (IMLP) and spatial-temporal prompting (STP). IMLP employs a joint fine-tuning strategy for the visual encoder and the prompt learner to effectively generate modality-shared text prompts and align them with visual features from different modalities in CLIP’s multimodal space, thereby mitigating modality differences. Additionally, STP models spatiotemporal information through two submodules, the spatial-temporal hub (STH) and spatial-temporal aggregation (STA), which further enhance IMLP by incorporating spatiotemporal information into text prompts. The STH aggregates and diffuses spatiotemporal information into the [CLS] token of each frame across the vision transformer (ViT) layers, whereas STA introduces dedicated identity-level loss and specialized multihead attention to ensure that the STH focuses on identity-relevant spatiotemporal feature aggregation. The VLD framework achieves state-of-the-art results on two VVI-ReID benchmarks. On the HITSZ-VCM dataset, it improves the Rank-1 accuracy by 7.3% and mAP by 7.6% (infrared-to-visible) and the Rank-1 accuracy by 10.4% and the mAP accuracy by 9.3% (visible to infrared) and requires only 2 hours of training, 2.39M additional parameters, and 0.12G FLOPs. The code will be released at https://github.com/Visuang/VLD.
Jiaxu Leng, Changjiang Kuang, Mingpi Tan, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Dual-Space Normalizing Flow for Unsupervised Video Anomaly Detection
abstract
Conventional reconstruction-based video anomaly detection (VAD) methods implicitly model normality in latent spaces, which is limited by the generalization ability of latent features. Normalizing Flow (NF)-based methods have been introduced to address this issue, as they explicitly model the distribution of input data and achieve significant performance in VAD. However, existing NF-based methods are confined to Euclidean space, limiting their ability to model action hierarchies. While effective at capturing local joint dynamics and short-term temporal variations, they fail to encode kinematic dependencies and long-term pose evolution, ultimately struggling to discern ambiguous anomalies that deviate minimally from normal motion. In contrast, hyperbolic representation learning, with its ability to model hierarchical and complex relationships among actions, offers a promising solution to enhance the discriminative power between similar skeletal actions. Motivated by this, we propose a novel Dual-Space Normalizing Flow (DSNF) method. Specifically, we design a Dual-Space Parallel Graph Convolutional Network (DSPGCN) that synergistically integrates the strengths of both Euclidean and hyperbolic geometries to simultaneously capture local detail features of poses and intrinsic hierarchical relationships of actions. To enhance the model's focus on discriminative features, we design an Adaptive Weighted Approximation Mass (AWAM) loss that dynamically adjusts weights to impose stronger constraints on regions with low discriminability in the dual space, encouraging the model to focus more on key discriminative features in hyperbolic space that reflect complex relationships between actions. Extensive experiments on public datasets demonstrate the effectiveness and robustness of our method in various VAD scenarios.
Jiaxu Leng, Mingpi Tan, Changjiang Kuang, Zhanjie Wu, Ji Gan, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 Difficulty-Guided Variant Degradation Learning for Blind Image Super-Resolution
abstract
Recent blind super-resolution (BSR) methods are explored to handle unknown degradations and achieve impressive performance. However, the prevailing assumption in most BSR methods is the spatial invariance of degradation kernels across the entire image, which leads to significant performance declines when faced with spatially variant degradations caused by object motion or defocusing. Additionally, these methods do not account for the human visual system's tendency to focus differently on areas of varying perceptual difficulty, as they uniformly process each pixel during reconstruction. To cope with these issues, we propose a difficulty-guided variant degradation learning network for BSR, named difficulty-guided degradation learning (DDL)-BSR, which explores the relationship between reconstruction difficulty and degradation estimation. Accordingly, the proposed DDL-BSR consists of three customized networks: reconstruction difficulty prediction (RDP), space-variant degradation estimation (SDE), and degradation and difficulty-informed reconstruction (DDR). Specifically, RDP learns the reconstruction difficulty with the proposed reconstruction-distance supervision. Then, SDE is designed to estimate space-variant degradation kernels according to the difficulty map. Finally, both degradation kernels and reconstruction difficulty are fed into DDR, which takes into account such two prior knowledge information to guide super-resolution (SR). Experimental analysis on various synthetic datasets demonstrates that DDL-BSR invariably surpasses state-of-the-art (SOTA) methods, producing SR images with enhanced realism and texture quality. Code is available at https://github.com/JiaWang0704/DDL-BSR.
Jiaxu Leng, Jia Wang 0036, Mengjingcheng Mo, Ji Gan, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator
abstract
In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all simply learn discriminative patterns by modelling low-level relationships between adjacent points of trajectories, while completely ignoring the inherent geometric structures of characters. Instead, we propose a novel Graph-guided Cross-modality Translator for IAHTR, which further explicitly exploits the geometric structures of characters for guiding the decoding of trajectories via graph-guided cross-modality attention mechanism without introducing extra annotation costs. Experiments on benchmarks IAHEW-UCAS2016 & IAM-OnDB show that our method has achieved state-of-the-art performance for handwritten text recognition.
Yuyan Chen, Ji Gan, Jiaxu Leng, Yan Zhang 0108, Xinbo Gao 0001
ICASSP4
2024 MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval
abstract
Text-to-image person retrieval aims to recognize target pedestrians based on specified text. Existing methods mainly obtain image and text features separately through distinct feature extractors, subsequently embedding them into a unified feature space and calculating their similarity. Despite great success, current methods still suffer from the lack of information interaction between images and text. To address this issue, we propose Mutual-guidance Representation Learning (MGRL) for text-to-image person retrieval, which captures the key features for matching via text-image information interaction. Accordingly, our MGRL consists of two customized modules: iterative text-guided feature extraction (ITFE) and vision-assisted specific mask complement (VSMC). Specifically, ITFE is first designed to extract the matching information between the text and the image concerning the local feature attention of the target pedestrians by iterative text guidance. Then, to further ensure the image features extracted by ITFE contain the text description, VSMC is designed to utilize the extracted image features to help complete masked text where the mask is difficult to complete with only unmasked text information. Experiments are conducted on CUHK-PEDES and ICFG-PEDES datasets, and experimental results demonstrate the superiority of the proposed MGRL.
Tianle Lv, Jiaxu Leng, Xinbo Gao 0001
ICASSP3
2024 Modality-Free Violence Detection via Cross-Modal Causal Attention and Feature Distillation
abstract
In this paper, we propose a novel framework, Modality-Free Violence Detection (MFVD), which captures the causal relationships among multimodal cues and ensures stable performance even in the absence of audio information. Specifically, we design a novel Cross-Modal Causal Attention mechanism (CCA) to deal with modality asynchrony by utilizing relative temporal distance and semantic correlation to obtain causal attention between audio and visual information instead of merely calculating correlation scores between audio and visual features. Moreover, to ensure our framework can work well when the audio modality is missing, we design a Cross-Modal Feature Distillation module (CFD), leveraging the common parts of the fused features obtained from CCA to guide the enhancement of visual features. Experimental results on the XD-Violence dataset demonstrate the superior performance of the proposed method in both vision-only and audio-visual modalities, surpassing state-of-the-art methods for both tasks.
Jiaxu Leng, Zhanjie Wu, Mengjingcheng Mo, Mingpi Tan, Xinbo Gao 0001
ICME1
2024 Dual Space Embedding Learning For Weakly Supervised Audio-Visual Violence Detection
abstract
In this paper, we propose Dual Space Embedding Learning (DSEL) for weakly supervised audio-visual violence detection, which excavates violence information deeply in both Euclidean and Hyperbolic spaces to distinguish violence from non-violence semantically and alleviate the asynchronous issue of violent cues in audio-visual patterns. Specifically, we first design a dual space visual feature interaction module (DSVFI) to deeply investigate the violence information in visual modality, which contains richer information compared to audio counterpart. Then, considering the modality asynchrony between the two modalities, we employ a late modality fusion method and design an asynchrony-aware audio-visual fusion module (AAF), in which visual features receive the violent prompt from the audio features after interacting among snippets and learning the violence information from each other. Experimental results show that our method achieves state-of-the-art performance on XD-Violence.
Zhanjie Wu, Mengjingcheng Mo, Ji Gan, Jiaxu Leng, Xinbo Gao 0001
ICME5
2024 Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection
abstract
While numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In contrast, hyperbolic representation learning, renowned for its ability to model hierarchical and complex relationships between events, has the potential to amplify the discrimination between visually similar events. Inspired by these, we develop a novel Dual-Space Representation Learning (DSRL) method for weakly supervised VVD to utilize the strength of both Euclidean and hyperbolic geometries, capturing the visual features of events while also exploring the intrinsic relations between events, thereby enhancing the discriminative capacity of the features. DSRL employs a novel information aggregation strategy to progressively learn event context in hyperbolic spaces, which selects aggregation nodes through layer-sensitive hyperbolic association degrees constrained by hyperbolic Dirichlet energy. Furthermore, DSRL attempts to break the cyber-balkanization of different spaces, utilizing cross-space attention to facilitate information interactions between Euclidean and hyperbolic space to capture better discriminative features for final violence detection. Comprehensive experiments demonstrate the effectiveness of our proposed DSRL.
Jiaxu Leng, Zhanjie Wu, Mingpi Tan, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001
NeurIPS1
2024 SCGTracker: Spatio-temporal correlation and graph neural networks for multiple object tracking
Yongquan Liang 0001, Jiaxu Leng, Zhihui Wang 0003
Pattern Recognit.3
2024 Invertible Image Obfuscation for Facial Privacy Protection via Secure Flow
abstract
This paper presents a fresh paradigm for protecting facial privacy via an invertible image obfuscation framework that incorporates multiple characteristics including anonymity, diversity, reversibility, security, and lightweight all at once. We name the framework PRO-Face S, an acronym for Privacy-preserving Reversible Obfuscation of Face images via Secure flow. The core of the proposed framework is a flow-based generative model (or invertible neural network), which takes as input a face image along with its pre-obfuscated form, and outputs the privacy-protected image that visually mirrors the pre-obfuscated one. The pre-obfuscation applied can be in various forms with different types and strengths. The invertibility of the flow-based model ensures that the original image can be easily recovered from the protected image in high fidelity. An elaborate secret key mechanism is devised to securely guide the mutual transformations of privacy protection and image recovery, such that the correct recovery is only possible upon the availability of the correct secret, pre-specified by the user in the protection stage. Two modes of wrong recovery are investigated to deal with malicious recovery attempts in different scenarios. Finally, extensive experiments conducted on multiple image datasets demonstrate the superiority of the proposed framework over state-of-the-art methods.
Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Jiaxu Leng, Tao Wu 0003, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 CRNet: Context-guided Reasoning Network for Detecting Hard Objects
abstract
Recent studies have shown impressive performance in object detection. However, most current detectors only explore the appearance feature to locate and classify objects but disregard or underestimate the valuable contextual information in the image, which limits the detection performance for those hard objects, such as small objects, occluded objects, blurred objects, etc. In this article, we instead seek to build a novel context modeling framework and conduct more effective context reasoning for object detection. Specifically, we design a Context-guided Reasoning Network (CRNet) to explore the relationships between objects and use easy detected objects to help understand hard ones. In our CRNet, an image is modeled as a graph and local features of objects are viewed as nodes of the graph to learn the relationships between objects. By passing contextual information in the built graph, the features of hard objects can be updated to discriminative features. To this end, we first develop a cascaded center prediction module built upon CenterNet to produce a set of high-quality proposals viewed as nodes of the graph. In addition, to maximize the value of global context information, we present a multi-granularity feature fusion network to encode the whole scene information which is also viewed as nodes of the graph. Then, the spatial and semantic relationships between objects are learned to initialize edges of the graph. Finally, context reasoning is conducted to update the node states iteratively. Extensive experiments are conducted on MS COCO and Pascal VOC to demonstrate the effectiveness of the proposed CRNet. Experimental results show that the proposed CRNet greatly improves the detection performance over existing context-based detectors, and it is comparable with state-of-the-art detectors.
Jiaxu Leng, Xinbo Gao 0001, Zhihui Wang 0003
IEEE Trans. Multim.1
2024 Contextual Learning in Fourier Complex Field for VHR Remote Sensing Images
abstract
Very high-resolution (VHR) remote sensing (RS) image classification is the fundamental task for RS image analysis and understanding. Recently, Transformer-based models demonstrated outstanding potential for learning high-order contextual relationships from natural images with general resolution ( pixels) and achieved remarkable results on general image classification tasks. However, the complexity of the naive Transformer grows quadratically with the increase in image size, which prevents Transformer-based models from VHR RS image ( pixels) classification and other computationally expensive downstream tasks. To this end, we propose to decompose the expensive self-attention (SA) into real and imaginary parts via discrete Fourier transform (DFT) and, therefore, propose an efficient complex SA (CSA) mechanism. Benefiting from the conjugated symmetric property of DFT, CSA is capable to model the high-order contextual information with less than half computations of naive SA. To overcome the gradient explosion in Fourier complex field, we replace the Softmax function with the carefully designed Logmax function to normalize the attention map of CSA and stabilize the gradient propagation. By stacking various layers of CSA blocks, we propose the Fourier complex Transformer (FCT) model to learn global contextual information from VHR aerial images following the hierarchical manners. Universal experiments conducted on commonly used RS classification datasets demonstrate the effectiveness and efficiency of FCT, especially on VHR RS images. The source code of FCT will be available at https://github.com/Gao-xiyuan/FCT.
Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Jiaxu Leng, Xiao Pu 0002, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Selecting Learnable Training Samples is All DETRs Need in Crowded Pedestrian Detection
abstract
DEtection TRansformer (DETR) and its variants (DETRs) achieved impressive performance in general object detection. However, in crowded pedestrian detection, the performance of DETRs is still unsatisfactory due to the inappropriate sample selection method which results in more false positives. To settle the issue, we propose a simple but effective sample selection method for DETRs, Sample Selection for Crowded Pedestrians (SSCP), which consists of the constraint-guided label assignment scheme (CGLA) and the utilizability-aware focal loss (UAFL). Our core idea is to select learnable samples for DETRs and adaptively regulate the loss weights of samples based on their utilizability. Specifically, in CGLA, we proposed a new cost function to ensure that only learnable positive training samples are retained and the rest are negative training samples. Further, considering the utilizability of samples, we designed UAFL to adaptively assign different loss weights to learnable positive samples depending on their gradient ratio and IoU. Experimental results show that the proposed SSCP effectively improves the baselines without introducing any overhead in inference. Especially, Iter Deformable DETR is improved to 39.7(-2.0)% MR on Crowdhuman and 31.8(-0.4)% MR on Citypersons.
Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001
ACM Multimedia2
2023 A location-aware siamese network for high-speed visual tracking
Lifang Zhou, Weisheng Li 0001, Jiaxu Leng, Bang Jun Lei, Weibin Yang
Appl. Intell.4
2023 Where to look: Multi-granularity occlusion aware for video person re-identification
Jiaxu Leng, Xinbo Gao 0001, Yan Zhang 0108, Ye Wang 0006, Mengjingcheng Mo
Neurocomputing1
2023 Lexical knowledge enhanced text matching via distilled word sense disambiguation
Xiao Pu 0002, Lin Yuan 0002, Jiaxu Leng, Tao Wu 0003, Xinbo Gao 0001
Knowl. Based Syst.3
2023 Characters as graphs: Interpretable handwritten Chinese character recognition via Pyramid Graph Transformer
Ji Gan, Yuyan Chen, Bo Hu 0008, Jiaxu Leng, Weiqiang Wang 0001, Xinbo Gao 0001
Pattern Recognit.4
2023 Pareto Refocusing for Drone-View Object Detection
abstract
Drone-view Object Detection (DOD) is a meaningful but challenging task. It hits a bottleneck due to two main reasons: (1) The high proportion of difficult objects (e.g., small objects, occluded objects, etc.) makes the detection performance unsatisfactory. (2) The unevenly distributed objects make detection inefficient. These two factors also lead to a phenomenon, obeying the Pareto principle, that some challenging regions occupying a low area proportion of the image have a significant impact on the final detection while the vanilla regions occupying the major area have a negligible impact due to the limited room for performance improvement. Motivated by the human visual system that naturally attempts to invest unequal energies in things of hierarchical difficulty for recognizing objects effectively, this paper presents a novel Pareto Refocusing Detection (PRDet) network that distinguishes the challenging regions from the vanilla regions under reverse-attention guidance and refocuses the challenging regions with the assistance of the region-specific context. Specifically, we first propose a Reverse-attention Exploration Module (REM) that excavates the potential position of difficult objects by suppressing the features which are salient to the commonly used detector. Then, we propose a Region-specific Context Learning Module (RCLM) that learns to generate specific contexts for strengthening the understanding of challenging regions. It is noteworthy that the specific context is not shared globally but unique for each challenging region with the exploration of spatial and appearance cues. Extensive experiments and comprehensive evaluations on the VisDrone2021-DET and UAVDT datasets demonstrate that the proposed PRDet can effectively improve the detection performance, especially for those difficult objects, outperforming state-of-the-art detectors. Furthermore, our method also achieves significant performance improvements on the DTU-Drone dataset for power inspection.
Jiaxu Leng, Mengjingcheng Mo, Yinghua Zhou, Chenqiang Gao, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Correlation Filter Tracker With Sample-Reliability Awareness and Self-Guided Update
abstract
In visual tracking, unreliable samples always exist because of occlusion, illumination variation, motion blur, etc. Existing studies have effectively improved the performance of trackers by enhancing the quality of online samples. However, an underappreciated view is that not all samples are equally essential to model training. In this paper, we propose a Sample-Aware Adaptive Updating (SAAU) strategy which can actively adjust the update formula by sensing the reliability of samples. Specifically, the Sample-Reliability Awareness (SRA) module can quantify sample reliability by calculating three specific indicators, where the Residual Peak-to-Correlation Energy (RPCE) is designed to cooperate with the other two introduced indicators to obtain credit scores on each sample. Besides, the Self-Guided Update (SGU) module provides a tracker with an unfixed learning rate that matches with the reliability label during updating, where our label annotator generates the label. Extensive experiments on several public benchmarks demonstrate the outstanding compatibility of SAAU and the superiority of our tracker (SAAU-CF) over state-of-the-art approaches.
Lifang Zhou, Bang Jun Lei, Weisheng Li 0001, Jiaxu Leng
IEEE Trans. Circuits Syst. Video Technol.5
2023 Multiple Pedestrian Tracking With Graph Attention Map on Urban Road Scene
abstract
Pedestrians are often vulnerable users of urban roads and ensuring their safety is a pressing challenge in the filed of intelligent transportation. Multiple pedestrian tracking is one of the key technologies for traffic statistics and abnormal behavior analysis, etc. Detection-based tracking methods have achieved remarkable results and have become mainstream processing schemes. However, target association is still immature and less effective in complex scenarios. In the proposed tracking system, several candidates surrounding each detected pedestrian are selected sparsely, and the associating relationship between the target and these candidates is determined based on a graph attention map. These graph attention maps contain positional correlations of the matching pairs and are applicable with pedestrians’ posture variations. The weighted correlating value is estimated with the positional weighted matrix and merged attention map. The correlating relationship is confirmed with the weighted correlating value and distance matching loss. To enhance the computation efficiency of the graph attention maps for these tracked targets and candidates, feature extraction is processed separately. Convolutional features extracted from one specific middle layer of the backbone network are used to represent each target or candidate. The experiment results show that the proposed tracker achieves better performance than the other five state-of-the-art trackers on three publicly available databases.
Zhihui Wang 0003, Zhiyuan Li 0013, Jiaxu Leng, Ming Li 0065, Lu Bai 0001
IEEE Trans. Intell. Transp. Syst.3
2023 HiGAN+: Handwriting Imitation GAN with Disentangled Representations
abstract
Humans remain far better than machines at learning, where humans require fewer examples to learn new concepts and can use those concepts in richer ways. Take handwriting as an example, after learning from very limited handwriting scripts, a person can easily imagine what the handwritten texts would like with other arbitrary textual contents (even for unseen words or texts). Moreover, humans can also hallucinate to imitate calligraphic styles from just a single reference handwriting sample (that even have never seen before). Humans can do such hallucinations, perhaps because they can learn to disentangle the textual contents and calligraphic styles from handwriting images. Inspired by this, we propose a novel handwriting imitation generative adversarial network (HiGAN+) for realistic handwritten text synthesis based on disentangled representations. The proposed HiGAN+ can achieve a precise one-shot handwriting style transfer by introducing the writer-specific auxiliary loss and contextual loss, and it also attains a good global & local consistency by refining local details of synthetic handwriting images. Extensive experiments, including human evaluations, on the benchmark dataset validate our superiority in terms of visual quality, scalability, compactness, and style transferability compared with the state-of-the-art GANs for handwritten text synthesis.
Ji Gan, Weiqiang Wang 0001, Jiaxu Leng, Xinbo Gao 0001
ACM Trans. Graph.3
2022 RCNet: Recurrent Collaboration Network Guided by Facial Priors for Face Super-Resolution
abstract
In this paper, we present a novel FSR method, called RCNet, which progressively improves the performance of FSR and Landmark Estimation (LE) via recurrent collaboration. In our approach, FSR and LE complement each other. Different from previous FSR methods that directly estimate the facial landmarks on the low-resolution face images, the proposed RCNet conducts LE on the super-resolution face image obtained through multiple iterations. Benefiting from the super-resolution face images, facial landmarks are precisely estimated, which boosts the performance of FSR in turn. Furthermore, we design a Component-Aware Fusion Module (CAFM) for better recovering facial details, which adaptively fine-tunes the estimated landmarks and groups them into face components to maximize the guiding role of the facial land-marks. In addition, the iterative feature aggregation is developed to preferably capture the information from the LR/SR face images. Experimental results show that the proposed RCNet outperforms the state-of-the-art methods in both quantitative and qualitative aspects for super-resolving very low-resolution faces.
Jiaxu Leng, Ye Wang 0006
ICME1
2022 Anomaly Warning: Learning and Memorizing Future Semantic Patterns for Unsupervised Ex-ante Potential Anomaly Prediction
abstract
Existing video anomaly detection methods typically utilize reconstruction or prediction error to detect anomalies in the current frame. However, these methods cannot predict ex-ante potential anomalies in future frames, which is imperative in real scenes. Inspired by the ex-ante prediction ability of humans, we propose an unsupervised Ex-ante Potential Anomaly Prediction Network (EPAP-Net), which learns to build a semantic pool to memorize the normal semantic patterns of future frames for indirect anomaly prediction. At the training time, the memorized patterns are encouraged to be discriminated through our Semantic Pool Building Module (SPBM) with the novel padding and updating strategies. Moreover, we present a novel Semantic Similarity Loss (SSLoss) at the feature level to maximize the semantic consistency of memorized items and corresponding future frames. Specially, to enhance the value of our work, we design a Multiple Frames Prediction module (MFP) to achieve anomaly prediction in future multiple frames. At the test time, we utilize the trained semantic pool instead of ground truth to evaluate the anomalies of future frames. Besides, to obtain better feature representations for our task, we introduce a novel Channel-selected Shift Encoder (CSE), which shifts channels along the temporal dimension between the input frames to capture motion information without generating redundant features. Experimental results demonstrate that the proposed EPAP-Net can effectively predict the potential anomalies in future frames and exhibit superior or competitive performance on video anomaly detection.
Jiaxu Leng, Mingpi Tan, Xinbo Gao 0001, Wen Lu 0004, Zongyi Xu
ACM Multimedia1
2022 ICNet: Joint Alignment and Reconstruction via Iterative Collaboration for Video Super-Resolution
abstract
Most previous frameworks either cost too much time or adopt some fixed modules resulting in alignment error in video super-resolution (VSR). In this paper, we propose a novel many-to-many VSR framework with Iterative Collaboration (ICNet), which employs the concurrent operation by iterative collaboration between alignment and reconstruction proving to be more efficient and effective than existing recurrent and sliding-window frameworks. With the proposed iterative collaboration, alignment can be conducted on super-resolved features from reconstruction while accurate alignment boosts reconstruction in return. In each iteration, the features of low-resolution video frames are first fed into the alignment and reconstruction subnetworks, which can generate temporal aligned features and spatial super-resolved features. Then, both outputs are fed into the proposed Tidy Two-stream Fusion (TTF) subnetwork that shares inter-frame temporal information and intra-frame spatial information without redundancy. Moreover, we design the Frequency Separation Reconstruction (FSR) subnetwork to not only model high-frequency and low-frequency information separately but also take benefit of each other for better reconstruction. Extensive experiments on benchmark datasets demonstrate that the proposed ICNet outperforms state-of-the-art VSR methods in terms of PSNR/SSIM values and visual quality, respectively.
Jiaxu Leng, Jia Wang 0036, Xinbo Gao 0001, Bo Hu 0008, Ji Gan, Chenqiang Gao
ACM Multimedia1
2022 Context augmentation for object detection
Jiaxu Leng, Ying Liu 0039
Appl. Intell.1
2022 Sampling-invariant fully metric learning for few-shot object detection
Jiaxu Leng, Taiyue Chen, Xinbo Gao 0001, Mengjingcheng Mo, Yongtao Yu, Yan Zhang 0108
Neurocomputing1
2022 Semantic-aware conditional variational autoencoder for one-to-many dialogue generation
Ye Wang 0006, Jingbo Liao, Hong Yu 0007, Jiaxu Leng
Neural Comput. Appl.4
2022 MTCNet: Multi-task collaboration network for rotation-invariance face detection
Lifang Zhou, Jiaxu Leng
Pattern Recognit.3
2022 Semantic-refined spatial pyramid network for crowd counting
Lifang Zhou, Peiwen Wang, Weisheng Li 0001, Jiaxu Leng, Bang Jun Lei
Pattern Recognit. Lett.4
2022 Hierarchical discrepancy learning for image restoration quality assessment
Bo Hu 0008, Shuaijian Wang, Leida Li, Jiaxu Leng, Yuzhe Yang 0001, Xinbo Gao 0001
Signal Process.4
2022 CrossNet: Detecting Objects as Crosses
abstract
With the use of deep learning, object detection has achieved great breakthroughs. However, existing object detection methods still can not cope with challenging environments, such as dense objects, small objects, and object scale variations. To address these issues, this paper proposes a novel keypoint-based detection framework, called CrossNet, which significantly improves detection performance with minimal costs. In our approach, an object is modeled as a cross that consists of a center keypoint and a specific size, which eliminates the need of hand-craft anchor design. The proposed CrossNet outputs three types of maps: the center map, size map, and offset map, where both center map and offset map are to predict the center keypoints of objects and the size map is to estimate the sizes (width and height) of objects. Specifically, we first design a cascaded center prediction method that introduces a coarse-to-fine idea to improve center prediction. Furthermore, since center prediction considered as a classification task is easier than size regression relatively, we design a center-attention size regression module that uses the detection results of centers to assist the size prediction. In addition, a slightly modified hourglass network is designed to enhance the quality of feature maps for center and size prediction. Extensive experiments are conducted to demonstrate the effectiveness of CrossNet on the challenging PASCAL VOC, COCO, KITTI, and WiderFace datasets. Empirical studies show that CrossNet achieves competitive results with top-ranked one-stage and two-stage detectors while being time-efficient.
Jiaxu Leng, Ying Liu 0039, Zhihui Wang 0003, Haibo Hu 0002, Xinbo Gao 0001
IEEE Trans. Multim.1
2021 CSFQGD: Chinese Sentence Fill-in-the-blank Question Generation Dataset for Examination
abstract
Fill-in-the-blank question generation has become enormously popular and attracted lots of attention recently. However, most of the existing question generation datasets are developed for machine reading comprehension, which are not specifically designed for examination. To fill in the gap, in this paper, we propose a Chinese sentence fill-in-the-blank question generation dataset for examination (named CSFQGD), which will be released to the public11Resources are available at The dataset is composed of 20.5K questions from many real examinations in Chinese that cover a wide spectrum of learning subjects. Based on the proposed dataset, we test several well-known methods for fill-in-the-blank question generation and compare their performance. Our baseline study on this dataset shows that CSFQGD is a challenging test bed for further research.
Zhenyu Cui, Jiaxu Leng, Ying Liu 0039
CSCWD3
2021 Selective region enlargement network for fast object detection in high resolution images
Jiaxu Leng, Ying Liu 0039, Xinbo Gao 0001
Neurocomputing1
2021 Realize your surroundings: Exploiting context information for small object detection
Jiaxu Leng, Yihui Ren 0002, Xiaoding Sun, Ye Wang 0006
Neurocomputing1
2021 A method of radar target detection based on convolutional neural network
Yihui Ren 0002, Ying Liu 0039, Jiaxu Leng
Neural Comput. Appl.4
2021 Single-shot augmentation detector for object detection
Jiaxu Leng, Ying Liu 0039
Neural Comput. Appl.1
2021 SKNet: Detecting Rotated Ships as Keypoints in Optical Remote Sensing Images
abstract
Detecting rotated ships is difficult in optical remote sensing images due to the challenges of complex scenes. Existing advanced rotated ship detectors are typically anchor-based algorithms that require plenty of predefined anchors. However, the use of anchors brings three critical problems: 1) a large number of anchors bring a huge amount of calculation; 2) the attributes (e.g., size and aspect ratios) of anchors are designed viaad hocheuristics; and 3) only a tiny fraction of anchors that overlap with ground-truth bounding boxes of ships tightly can be considered as positive samples, which causes an extreme imbalance between positive and negative samples. As a result, the detection accuracy will be influenced seriously when the design of anchors is not suitable. To address the above problems, this article proposes a novel anchor-free rotated ship detection framework, called SKNet, which detects rotated ships as keypoints in optical remote sensing images. In SKNet, a ship target is modeled as its center keypoint and morphological sizes, including the width, height, and rotation angle. Accordingly, we design two customized modules: orthogonal pooling and soft-rotate-nonmaximum suppression (NMS), where the former is to improve the prediction accuracy of the center keypoint and the morphological size, and the latter is to effectively remove redundant rotated ship detection results. Extensive experiments are conducted to demonstrate the effectiveness of SKNet on three optical remote sensing image data sets: HRSC2016, DOTA-ship, and HPDM-OSOD, which is collected by ourselves and published in this article. Empirical studies show that SKNet achieves state-of-the-art detection performance while being time-efficient. Overall, SKNet achieves the best speed–accuracy tradeoff.
Zhenyu Cui, Jiaxu Leng, Ying Liu 0039, Pei Quan
IEEE Trans. Geosci. Remote. Sens.2
2021 Selective Domain-Invariant Feature Alignment Network for Face Anti-Spoofing
abstract
One primary challenge in face anti-spoofing refers to suffering a sharp performance drop in cross-domain scenes, where training and testing images are collected from different datasets. Recent methods have achieved promising results by aligning the features of all images among the available source domains. However, due to significant distribution discrepancies among non-face regions of all images, it is challenging to capture domain-invariant features for these regions. In this paper, we propose a novel Selective Domain-invariant Feature Alignment Network (SDFANet) for cross-domain face anti-spoofing, which aims to seek common feature representations by fully exploring the generalization of different regions of images. Different from previous works that align the whole features directly, the proposed SDFANet leverages multiple domain discriminators with the same architecture to balance the generalization of different regions of the all images. Specifically, we firstly design a multi-grained feature alignment network composed of a local-region and global-image alignment subnetworks to learn more generalized feature space for real faces. Besides, the domain adapter module, which aims to alleviate the large domain discrepancy with the help of the domain attention strategy, is adopted to facilitate the learning of our multi-grained feature alignment network. In addition, a multi-scale attention fusion module is designed in our feature generator to refine the different levels of features effectively. Experimental results show that the proposed SDFANet can greatly improve the generalization ability of face anti-spoofing, and that is superior to the existing methods.
Lifang Zhou, Xinbo Gao 0001, Weisheng Li 0001, Bang Jun Lei, Jiaxu Leng
IEEE Trans. Inf. Forensics Secur.6
2020 Deep learning for drug-drug interaction extraction from the literature: a review
abstract
Drug-drug interactions (DDIs) are crucial for drug research and pharmacovigilance. These interactions may cause adverse drug effects that threaten public health and patient safety. Therefore, the DDIs extraction from biomedical literature has been widely studied and emphasized in modern biomedical research. The previous rules-based and machine learning approaches rely on tedious feature engineering, which is labourious, time-consuming and unsatisfactory. With the development of deep learning technologies, this problem is alleviated by learning feature representations automatically. Here, we review the recent deep learning methods that have been applied to the extraction of DDIs from biomedical literature. We describe each method briefly and compare its performance in the DDI corpus systematically. Next, we summarize the advantages and disadvantages of these deep learning models for this task. Furthermore, we discuss some challenges and future perspectives of DDI extraction via deep learning methods. This review aims to serve as a useful guide for interested researchers to further advance bioinformatics algorithms for DDIs extraction from the literature.
Jiaxu Leng, Ying Liu 0039
Briefings Bioinform.2
2020 Robust Obstacle Detection and Recognition for Driver Assistance Systems
abstract
This paper proposes a robust obstacle detection and recognition method for driver assistance systems. Unlike existing methods, our method aims to detect and recognize obstacles on the road rather than all the obstacles in the view. The proposed method involves two stages aiming at an increased quality of the results. The first stage is to locate the positions of obstacles on the road. In order to accurately locate the on-road obstacles, we propose an obstacle detection method based on the U-V disparity map generated from a stereo vision system. The proposed U-V disparity algorithm makes use of the V-disparity map that provides a good representation of the geometric content of the road region to extract the road features, and then detects the on-road obstacles using our proposed realistic U-disparity map that eliminates the foreshortening effects caused by the perspective projection of pinhole imaging. The proposed realistic U-disparity map greatly improves the detection accuracy of the distant obstacles compared with the conventional U-disparity map. Second, the detection results of our proposed U-V disparity algorithm are put into a context-aware Faster-RCNN that combines the interior and contextual features to improve the recognition accuracy of small and occluded obstacles. Specifically, we propose a context-aware module and apply it into the architecture of Faster-RCNN. The experimental results on two public datasets show that our proposed method achieves state-of-the-art performance under various driving conditions.
Jiaxu Leng, Ying Liu 0039, Dawei Du, Pei Quan
IEEE Trans. Intell. Transp. Syst.1
2019 Alzheimer's Disease Diagnosis Using Enhanced Inception Network Based on Brain Magnetic Resonance Image
abstract
An estimated 24 million people worldwide have dementia, the majority of whom are thought to have Alzheimer's disease(AD). Nowadays, Alzheimer's disease represents a significant public health concern and has been identified as a research priority. Most unfortunately, there is little chance of a cure for Alzheimer's disease, and the disease is difficult to detect before the dominant characteristics such as memory loss are manifested. Therefore, the diagnosis of Alzheimer's disease has become an urgent problem today. Studies have shown that mild cognitive impairment(MCI) is a state between Alzheimer's disease and normal, and the chance of it turning into Alzheimer's disease is high. Therefore, if machines can automatically learn the characteristics of three kinds of human brain magnetic resonance(MR) images through deep learning, and help doctors to diagnose patients with mild cognitive impairment or Alzheimer's disease accurately, it will be beneficial for the early diagnosis of Alzheimer's disease. In this paper, we improve the Inception(V3) neural network and further test the effectiveness of the enhanced network based on the international Alzheimer's disease data set, which consists of brain magnetic resonance images. The results show that the average accuracy of our approach can reach 85.7%.
Zhenyu Cui, Zhiao Gao, Jiaxu Leng, Pei Quan
BIBM3
2019 A Novel Neuron Connection Model Mimicking Human Beings
abstract
Neural Networks have achieved great success in many computer vision tasks, especially in image recognition. However, as neural networks grow deeper and deeper, to some extend, we've found them becoming difficult to train, and requiring samples in large scale dramatically, even with the help of Dropout and Dropconnect methods, which do improve the accuracy a bit but burdens the training process as a sacrifice. To overcome this, we propose a novel neuron connection model to generate dynamic graphs of computation. As synapses have two kinds: excitatory and inhibitory ones, our model also has two kinds of connections for neurons. In addition, we propose a training algorithm that deals with non-differentiable because the equations of the connections and activation function of neurons in our model are not really differentiable. To evaluate the effectiveness the proposed method, we apply it to the image recognition task, and the results show that our proposed model achieves state-of-the-art performance on three public datasets: MNIST, CIFAR-10, and CIFAR-100.
Jiaxu Leng, Zhenyu Cui, Chao Xiang, Pei Quan
BIBM1
2019 An enhanced SSD with feature fusion and visual reasoning for object detection
Jiaxu Leng, Ying Liu 0039
Neural Comput. Appl.1
2019 Context-aware attention network for image recognition
Jiaxu Leng, Ying Liu 0039, Shang Chen
Neural Comput. Appl.1
2018 Context-Aware U-Net for Biomedical Image Segmentation
Jiaxu Leng, Ying Liu 0039, Pei Quan, Zhenyu Cui
BIBM1