VLDB 2026 Research / reviewers in the wild / expert
Jie Li 0001
dblp:17/2703-1
· DBLP profile ↗
228ranked-venue papers
3as first author
118since 2021 · last 2026
0000-0001-7950-4233ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 109 · 1 first-author · 59 since 2021Artificial intelligence and machine learning · 105 · 2 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 16 since 2021Security and privacy · 6 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional AnnotationabstractThe development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline. Longhao Li, Zhao Guo, Hongjie Chen 0001, Yuhang Dai, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Hui Bu, Jie Li 0001, Jian Kang 0006, Ruibin Yuan, Ziya Zhou, Wei Xue 0002, Lei Xie 0001 |
AAAI | 12 |
| 2026 | Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision AlignmentabstractDecoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally asymmetric, characterized by two critical gaps: a Fidelity Gap (stemming from EEG's inherent noise and signal degradation, vs. vision's high-fidelity features) and a Semantic Gap (arising from EEG's shallow conceptual representation, vs. vision's rich semantic depth). Previous methods often overlook this asymmetry, forcing alignment between the two modalities as if they were equal partners and thereby leading to poor generalization. To address this, we propose the adaptive teaching paradigm. This paradigm empowers the ``teacher" modality (vision) to dynamically shrink and adjust its knowledge structure under task guidance, tailoring its semantically dense features to match the ``student" modality (EEG)'s capacity. We implement this paradigm with the ShrinkAdapter, a simple yet effective module featuring a residual-free design and a bottleneck structure. Through extensive experiments, we validate the underlying rationale and effectiveness of our paradigm. Our method achieves a top-1 accuracy of 60.2% on the zero-shot brain-to-image retrieval task, surpassing previous state-of-the-art methods by a margin of 9.8%. Our work introduces a new perspective for asymmetric alignment: the teacher must shrink and adapt to bridge the vision-brain gap. Lukun Wu, Jie Li 0001, Ziqi Ren, Kaifan Zhang, Xinbo Gao 0001 |
AAAI | 2 |
| 2026 | Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud RemovalabstractCloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct structured information beneath cloud-covered areas. In this work, we propose a Visibility-guided Semantic Estimation and Reconstruction network for cloud removal (VISER-CR), which reformulates CR as a structure-guided completion problem. Specifically, VISER-CR explicitly models cloud interference via spatial masking, encouraging the model to reason beyond pixel-level appearance and enhance scene-level structural understanding. Moreover, to further improve the representation of structural information, we introduce Patch Saliency Encoding, a self-guided mechanism that implicitly models structural alignment among patches, significantly enhancing clustering consistency and semantic separability in the latent space. This adaptive mechanism guides the network to focus on learning and reconstructing structurally important regions, thereby reducing redundancy and improving overall cloud removal performance. Extensive experiments on multiple benchmark datasets demonstrate the superior effectiveness of our method. Jingwei Xin, Jie Li 0001, Nannan Wang 0001 |
AAAI | 3 |
| 2026 | DIFFA: Large Language Diffusion Models Can Listen and UnderstandabstractRecent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Jiaming Zhou 0001, Hongjie Chen 0001, Shiwan Zhao, Jian Kang 0006, Jie Li 0001, Enzhi Wang, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xuelong Li 0001 |
AAAI | 5 |
| 2026 | Mitigating bias in chest X-ray disease diagnosis via de-biased disentangled representation learning
Xinwei Lai, Jie Li 0001, Xinbo Gao 0001, Zhicheng Jiao, Zhusi Zhong |
Artif. Intell. Medicine | 2 |
| 2026 | A Multi-Granularity Scene-Aware Graph Convolution Method for Weakly Supervised Person Search
De Cheng, Haichun Tai, Nannan Wang 0001, Xiangqian Zhao, Jie Li 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | EKPC: Elastic Knowledge Preservation and Compensation for Class-Incremental Learning
Huaijie Wang, De Cheng, Yan Li 0125, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | One-step diffusion-based real-world image super-resolution with visual perception distillation
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Neurocomputing | 5 |
| 2026 | SSD: Making Face Forgery Clues Evident Again With Self-Steganographic DetectionabstractThe rapid development of generative AI techniques enables the synthesis of highly realistic facial images, posing significant challenges for the accurate detection of face forgeries. In contrast to solely elevating detector awareness, proactively reducing the intrinsic difficulty of forgery detection can streamline detector complexity while improving both generalization and robustness. This insight motivates our defense strategy to make face forgery clues more evident. Specifically, a novel proactive approach dubbed Self-Steganographic Detection (SSD) is proposed to imperceptibly embed facial images into themselves as a form of detection evidence. The recovery process is designed to remain robust under normal manipulations while exhibiting deliberate degradation under malicious manipulations, thereby clearly revealing potential forgeries. Unlike embedding bit-level vectors, pixel-level images are informative to ensure the generalization of our approach. Due to the similarity between the protected and embedded images, SSD performs detection without storing any embedded information in advance. To support practical deployment, our approach incorporates a dual detection scheme that aims to identify unprotected images and determine the authenticity of protected images. Extensive experiments using 8 face forgery techniques demonstrate the effectiveness of our approach compared to state-of-the-art methods. Ruiyang Xia, Dawei Zhou 0004, Lin Yuan 0002, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DINO-PCB: Two-stage vision foundation model pretraining and distillation for real-time circuit-board defect detection
Junjie Ke, Lihuo He, Jing Zhang 0037, Yuqi Ji, Hui Chen 0013, Jie Li 0001, Sicheng Zhao, Guiguang Ding, Xinbo Gao 0001 |
Pattern Recognit. | 7 |
| 2026 | PSC-UDA: Point-cloud Structure Constrained Unsupervised Domain Adaptation for contour-based kidney segmentationabstractCross-domain medical image segmentation has gained increasing interest for its potential to reduce annotation efforts and improve clinical generalization capabilities. Domain adaptation aims to tackle the domain shift that appears in different image modalities. In cross-domain segmentation, generative models often suffer from limited accuracy due to their lack of domain-specific representations. Besides, many transfer learning approaches rely on additional manual annotations for supervision, emerging paradigms such as Unsupervised Domain Adaptation (UDA) facilitate effective knowledge transfer even when labels in the target domain are entirely absent. In this study, we propose a novel Point-cloud Structure Constrained Unsupervised Domain Adaptation (PSC-UDA) framework based on a Contour-Aware Segmentation (CAS) model with a 3D contour point cloud to bridge the domain gaps appearing in cross-site and cross-domain medical images. The CAS model distills the domain-invariant kidney structure from image texture to distinguish the point cloud and characterize the kidney contour in a coarse-to-fine way. With point-to-voxel self-learning on 3D structure constraints, the proposed PSC-UDA framework addresses visual domain shift, adapting discriminative information of the kidney from the labeled source domain (CT) to the unlabeled target domain (CT/MRI), so that it realizes precise cross-domain kidney segmentation with limited labels. Experimental results prove that the proposed method outperforms the generative UDA methods and the source-free methods on three cross-domain kidney segmentation datasets, outperforming even without a target domain adaptation strategy. The source code is available at https://github.com/zzs95/PSC-UDA . Yang Li 0111, Zhusi Zhong, Jie Li 0001, Helen Zhang, Mihir Khunte, Lulu Bi, Scott Collins, Harrison X. Bai, Michael Atalay, Ihab Kamel, Xinbo Gao 0001, Zhicheng Jiao |
Pattern Recognit. | 3 |
| 2026 | TransFA: Transformer-based representation for face attribute evaluation
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 5 |
| 2026 | IMEVSI: Online Adaptive Video Stream Interpolation via Inertia-Aware Motion EstimationabstractRecent video frame interpolation (VFI) methods rely on computationally heavy modules (e.g., global attention module) to handle large motions, incurring prohibitive costs which hinders their practical real-time deployment. In this work, we revisit the core objective of VFI: enhancing the temporal resolution of videos. We identify that previous VFI’s frame-isolated processing ignores continuous temporal modeling, introduces computational redundancy in video streaming scenarios. To address this, we propose an online learning recurrent net with inertia-aware motion estimation(IMEVSI). It consists of implicit motion propagation( IMP), explicit motion propagation(EMP) and adaptive online learning strategy(AOL). IMP and EMP are used to high order inter-frame motion modeling considering motion inertia, AOL are proposed to bridge the motion domain gap between training and deployment. For IMP, we initiate from explicit physical motion modeling, progressively integrating learnable parameters into inertia-ware motion extraction and finally unify motion propagation and extraction within our recurrent motion propagation Transformer(RMPT). For EMP, we directly inject adjacent motion into current flow estimation recognizing its inertia contribution. For AOL,we leverage cycle consistency to dynamically adjust intermediate flow estimator and maintains an adaptive threshold to control parameter update. Extensive experiments demonstrate that our method outperforms state-of-the-art (SOTA) approaches on regular, large-motion, and high-resolution benchmarks while achieving excellent inference speed and FLOPs. Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality AssessmentabstractModeling visual perception in a manner consistent with human subjective evaluation has become a central direction in both video quality assessment (VQA) and broader visual understanding tasks. While free-energy-guided self-repair mechanisms—reflecting human observational experience—have proven effective in image quality assessment, extending them to VQA remains non-trivial. In addition, biologically inspired paradigms such as holistic perception, local analysis, and gaze-driven scanning have achieved notable success in high-level vision tasks, yet their potential within the VQA context remains largely underexplored. To address these issues, we propose EyeSimVQA, a novel VQA framework that incorporates free-energy-based self-repair. It adopts a dual-branch architecture, with an aesthetic branch for global perceptual evaluation and a technical branch for fine-grained structural and semantic analysis. Each branch integrates specialized enhancement modules tailored to distinct visual inputs—resized full-frame images and patch-based fragments—to simulate adaptive repair behaviors. We also explore a principled strategy for incorporating high-level visual features without disrupting the original backbone. In addition, we design a biologically inspired prediction head that models sweeping gaze dynamics to better fuse global and local representations for quality prediction. Experiments on five public VQA benchmarks demonstrate that EyeSimVQA achieves competitive or superior performance compared to state-of-the-art methods, while offering improved interpretability through its biologically grounded design. Our code will be publicly available at https://github.com/handsomewzy/EyeSim-VQA. Zhaoyang Wang 0003, Wen Lu 0004, Jie Li 0001, Lihuo He, Maoguo Gong, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Video Frame Interpolation via Appearance-Based Intermediate Flow EstimationabstractIntermediate flow estimation is an important part of video frame interpolation (VFI). Most previous works use interpolation to derive the intermediate flow assuming localized linear motion. However, this method is not effective when dealing with extreme motions. In this work, we assume that the motion trajectory of an object is determined by the appearance characteristics of this object. Based on this assumption, we propose a new intermediate flow estimation method, which obtains the motion features of intermediate frames from image appearance and inter-frame motion features. In addition, in order to fully extract the inter-frame features, we rethink the difference of VFI and previous works on using Swin-Transformer and compute the appearance features and motion features within the adaptive neighborhood by cyclically shifting the window. Experimental results show that our method achieves state-of-the-art performance on different datasets for both fixed-time and arbitrary-time interpolation. Moreover, our proposed method outperforms models that require inputting a sequence of four frames when handling videos with extremely large motion. The source code is available from https://github.com/chen12304/IFE-VFI. Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | DPENet: A Dual Prototype-Enhanced Network for Few-Shot Object DetectionabstractExisting meta-learning based few-shot object detection methods suffer from limitations in learning representative prototypes. Specifically, directly aggregating bounding box contents from support images into prototypes renders these methods vulnerable to background noise and the morphological intricacies of objects. Furthermore, these methods neglect the varied contributions of intra-class image-specific prototypes and fail to leverage semantic information effectively during prototype generation, resulting in suboptimal class representations due to naive average aggregation. To address these issues, we propose a Dual Prototype-Enhancement Network (DPENet), designed to optimize prototypes by improving support feature representation and enhancing prototype discriminability. Specifically, we introduce an Object Enhancement Module (OEM) based on dynamic hypergraph construction. This module employs hypergraph convolution to adaptively capture complex high-order semantic interactions among highly similar regions within support features, thereby highlighting salient features of target regions, suppressing background noise, and enhancing support feature representation. Moreover, we propose a Semantic Fusion Perception Module (SFPM) that generates more discriminative class-specific prototypes by integrating weighted intra-class prototype representations with text-based semantic embeddings. Experimental results demonstrate that DPENet significantly outperforms existing methods on the PASCAL VOC and MS COCO datasets. Jingling Huang, Hanzi Wang, Qiangqiang Wu, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | WSformer: Wavelet-Based Sparse Transformer for Blind Image RestorationabstractAs a fundamental task in image processing, blind image restoration (BIR) faces significant challenges due to the unknown nature of the degradation process. While transformer-based methods have shown promise in various applications, they encounter difficulties in BIR. One key challenge is that the complexity of degradation easily leads to incorporate irrelevant information into their attention mechanisms, thereby hindering restoration performance. To address this challenge, sparsification strategies have been commonly adopted. However, existing sparse transformer-based methods typically determine sparse members through fixed patterns such as constant thresholds or predefined sources, making their sparsification strategies too rigid. To tackle this issue, we propose WSformer, a Wavelet-based Sparse transformer tailored for BIR, which offers three key advantages. First, we design a Sparse Reciprocal Multi-head Self-Attention (SR-MSA) mechanism in the attention layer. This mechanism employs sparse and reciprocal strategies to adaptively select reliable information, while operating across channels to reduce computational complexity. Second, recognizing that feed-forward networks in existing transformer blocks fail to effectively leverage global information, we develop a Recalibrated Feed-Forward Network (RFFN). It fully exploits the fusion of local and global information, enhancing the robustness of feature learning. Finally, to mitigate the increased computational burden introduced by these innovations, we equip WSformer with wavelet transform. Combined with a U-shaped architecture, it enables WSformer to achieve an optimal balance between performance and inference time. Extensive experiments on multiple BIR tasks validate WSformer's effectiveness in both quantitative metrics and visual quality. The code is available at https://github.com/CanZhang01/WSformer. Zhonggui Sun, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Interpretable General Image Fusion via Scalable Autoregressive ModelingabstractExisting image fusion methods have developed increasingly sophisticated network architectures for exploiting modality-shared and modality-specific features. However, despite these advancements in feature extraction, most methods ultimately rely on relatively simple implicit or explicit fusion strategies, which can compromise interpretability and limit fusion accuracy. In this paper, we incorporate visual autoregressive modeling to bridge the gap between implicit feature extraction and explicit modality fusion. First, the proposed approach conducts a low-to-high resolution autoregressive objective with modality-specific features, introducing a scalable feature autoregressive mechanism. It aggregates local and global contextual dependencies while enhancing implicit cross-scale interaction. Furthermore, to promote the consistency and complementarity across modalities, we embed an explicit high-order fusion strategy within the progressive modality-specific feature extraction process. This integration facilitates a next-scale synergistic relationship between implicit learning and explicit fusion. Our High-order Feature AutoRegressive Fusion framework (HFARFusion) provides a robust and interpretable solution for general image fusion tasks, effectively balancing fusion performance and transparency through the strengths of autoregressive learning. Extensive experiments demonstrate the outstanding performance of the proposed method in several classical fusion tasks, including infrared-visible, medical, multi-focus, and multi-exposure image fusion. Our code is available at https://github.com/happysbn/HFARFusion. Jingwei Xin, Boneng Shi, Zhen Li 0026, Xuehao Song, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | PropMambaSR: Lightweight Image Super-Resolution With Propagation State Space ModelabstractState Space Models (SSMs), particularly Mamba, have emerged as promising alternatives to Transformers for lightweight Single Image Super-Resolution (SISR) due to their linear complexity in modeling long-range dependencies. Advanced Mamba-based Super-Resolution models, such as MambaIR and VMambaIR, have demonstrated compelling performance by integrating bidirectional and multi-directional scanning mechanisms. However, these methods are still constrained by several critical challenges that limit their ultimate reconstruction capabilities. Firstly, the inherent requirement for 2D-to-1D serialization disrupts local spatial coherence, which is vital for preserving fine-grained texture details during reconstruction. Second, stacking multiple SSM blocks in deep networks leads to progressive decay of fine-grained information, as hidden states are typically confined within individual layers without cross-layer propagation. To address these challenges, we propose PropMambaSR, a novel hybrid architecture that synergistically combines global context modeling with local feature extraction. Our core contribution is the Propagation State Space Model (PropSSM), which establishes explicit cross-layer hidden state routing to enable direct propagation of fine-grained information from shallow to deep layers, effectively mitigating information decay while preserving reconstruction details. Additionally, we introduce an Enhanced Spatial Feature Block (ESFB) that employs multi-scale residual distillation to explicitly capture local texture information compromised by the serialization process. Finally, a dynamic Fusion Model (FM) adaptively integrates global contextual features from PropSSM with local spatial features from ESFB through learnable gating mechanisms. Extensive experiments demonstrate that PropMambaSR achieves state-of-the-art performance across multiple benchmarks while maintaining computational efficiency. Notably, on the Urban100 dataset for$\times 2$SR, our model surpasses the next-best method by 0.33 dB in PSNR, demonstrating superior capability in reconstructing high-frequency textures and regular structural patterns. Wushuai Jin, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Infrared and Visible Image Fusion with Hierarchical Human PerceptionabstractImage fusion combines images from multiple domains into one image, containing complementary information from source domains. Existing methods take pixel intensity, texture and high-level vision task information as the standards to determine preservation of information, lacking enhancement for human perception. We introduce an image fusion method, Hierarchical Perception Fusion (HPFusion), which leverages Large Vision-Language Model to incorporate hierarchical human semantic priors, preserving complementary information that satisfies human visual system. We propose multiple questions that humans focus on when viewing an image pair, and answers are generated via the Large Vision-Language Model according to images. The texts of answers are encoded into the fusion network, and the optimization also aims to guide the human semantic distribution of the fused image more similarly to source images, exploring complementary information within the human perception domain. Extensive experiments demonstrate our HPFusoin can achieve high-quality fusion results both for information preservation and human visual enhancement Our code is available at https://github.com/SSyangguangZHPFusion. Jie Li 0001, Zhusi Zhong, Xinbo Gao 0001 |
ICASSP | 2 |
| 2025 | Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
Hongjie Chen 0001, Qing Wang 0039, Hang Lv 0006, Jian Kang 0006, Jie Li 0001, Zhennan Lin, Lei Xie 0001 |
INTERSPEECH | 6 |
| 2025 | WM-SORT: Modeling multi-object tracking motion patterns in real world
Yizhuo Jiang, Weiyu Zhao, Yan Gao 0025, Xinhang Niu, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 5 |
| 2025 | DETrack: Depth information is predictable for tracking
Weiyu Zhao, Yizhuo Jiang, Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2025 | Inferring normality from noised samples: Enhanced deep autoencoder with image denoising for anomaly detection
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001 |
Inf. Sci. | 4 |
| 2025 | Unsupervised Face Super-Resolution via Integrating Faithful 3D Facial PriorsabstractRecently, unsupervised face super-resolution (FSR) has attracted significant attention due to its remarkable generalization performance. However, existing methods neglect the incorporation of facial priors, which can effectively guide the restoration of face images. The root cause of this issue lies in the significant challenges associated with incorporating facial priors into unsupervised frameworks. First, unsupervised methods often face the challenge of real-world low-quality (LQ) images that are severely corrupted, making it unrealistic to extract reliable prior information from them. Second, the estimation of facial priors exponentially increases the model’s parameters and computational complexity, contradicting the purpose of unsupervised methods for practical deployment. In this work, we fundamentally address the aforementioned challenges and proposeFaith3D-FSR, a novel approach that incorporates faithful 3D facial priors into unsupervised FSR. Specifically, we introduceFaith3Dmechanism for faithful prior integration, which deconstructs super-resolution images into 3D elements and uses the 3D priors from real high-quality (HQ) images as reference for calibration solely during the training phase. This strategy enables more precise guidance on the super-resolution in a high-dimensional space, without requiring additional prior estimation during inference. It successfully overcomes the aforementioned challenges, making it more suitable for real-world applications, and offers a plug-and-play solution for incorporating 3D priors into unsupervised FSR. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on multiple benchmark datasets and across a range of evaluation metrics. The code is available here. Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Vision-Language Models Empowered Nighttime Object Detection With Consistency Sampler and Hallucination Feature GeneratorabstractCurrent object detectors often suffer performance degradation when applied to cross-domain scenarios, particularly under challenging visual conditions such as nighttime scenes. This is primarily due to the I3 problems: Inadequate sampling of instance-level features, Indistinguishable feature representation across domains and Inaccurate generation for identical category participation. To address these challenges, we propose a domain-adaptive detection framework that enables robust generalization across different visual domains without introducing any additional inference overhead. The framework comprises three key components. Specifically, the centerness-category consistency sampler alleviates inadequate sampling by selecting representative instance-level features, while the paired centerness consistency loss enforces alignment between classification and localization. Second, VLM-based orthogonality enhancement leverages frozen vision-language encoders with an orthogonal projection loss to improve cross-domain feature distinguishability. Third, hallucination feature generator synthesizes robust instance-level features for missing categories, ensuring balanced category participation across domains. Extensive experiments on multiple datasets covering various domain adaptation and generalization settings demonstrate that our method consistently outperforms state-of-the-art detectors, achieving up to 5.5 mAP improvement, with particularly strong gains in nighttime adaptation. Lihuo He, Junjie Ke, Jie Li 0001, Qi Wang 0009, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Diffusion Model-Based Visual Compensation Guidance and Visual Difference Analysis for No-Reference Image Quality AssessmentabstractExisting free-energy guided No-Reference Image Quality Assessment (NR-IQA) methods continue to face challenges in effectively restoring complexly distorted images. The features guiding the main network for quality assessment lack interpretability, and efficiently leveraging high-level feature information remains a significant challenge. As a novel class of state-of-the-art (SOTA) generative model, the diffusion model exhibits the capability to model intricate relationships, enhancing image restoration effectiveness. Moreover, the intermediate variables in the denoising iteration process exhibit clearer and more interpretable meanings for high-level visual information guidance. In view of these, we pioneer the exploration of the diffusion model into the domain of NR-IQA. We design a novel diffusion model for enhancing images with various types of distortions, resulting in higher quality and more interpretable high-level visual information. Our experiments demonstrate that the diffusion model establishes a clear mapping relationship between image reconstruction and image quality scores, which the network learns to guide quality assessment. Finally, to fully leverage high-level visual information, we design two complementary visual branches to collaboratively perform quality evaluation. Extensive experiments are conducted on seven public NR-IQA datasets, and the results demonstrate that the proposed model outperforms SOTA methods for NR-IQA. The codes will be available at https://github.com/handsomewzy/DiffV2IQA. Zhaoyang Wang 0003, Bo Hu 0008, Mingyang Zhang 0002, Jie Li 0001, Leida Li, Maoguo Gong, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | MVFusion: Generative Representation Learning With Masked Variational Autoencoders for Multi-Modality Image FusionabstractCreating a comprehensively representative image while maintaining the merits of various modalities is a key focus of current Multi-Modality Image Fusion research. Existing unified methods often struggle to handle varying types of degradation while extracting modality-shared and modality-specific information from source images, leading to limitations in their generative or representation capabilities under different conditions. To address the challenge, we propose MVFusion, a novel self-supervised masked variational autoencoder framework that simultaneously enhances generative training and representation learning. It is designed to cope with varying image quality and dataset composition with a unified framework while ensuring effective fusion of modality information. Specifically, MVFusion employs a self-supervised masked autoencoder to reduce the impact of redundancy and degradation in the source images, and thus learns the latent distribution of degraded input images in the generative training stage. In addition, we incorporate variational feature learning to further preserve the distinctive modality features in the representation learning stage. Extensive experiments demonstrate that our model achieves promising results in several classical fusion tasks, including infrared-visible, multi-focus, multi-exposure, and medical image fusion. The code is available at https://github.com/shiboneng/MVFusion. Jingwei Xin, Boneng Shi, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Semantic-Driven Global-Local Fusion Transformer for Image Super-ResolutionabstractImage Super-Resolution (SR) has seen remarkable progress with the emergence of transformer-based architectures. However, due to the high computational cost, many existing transformer-based SR methods limit their attention to local windows, which hinders their ability to model long-range dependencies and global structures. To address these challenges, we propose a novel SR framework named Semantic-Driven Global-Local Fusion Transformer (SGLFT). The proposed model enhances the receptive field by combining a Hybrid Window Transformer (HWT) and a Scalable Transformer Module (STM) to jointly capture local textures and global context. To further strengthen the semantic consistency of reconstruction, we introduce a Semantic Extraction Module (SEM) that distills high-level semantic priors from the input. These semantic cues are adaptively integrated with visual features through an Adaptive Feature Fusion Semantic Integration Module (AFFSIM). Extensive experiments on standard benchmarks demonstrate the effectiveness of SGLFT in producing visually faithful and structurally consistent SR results. The code will be available at https://github.com/kbzhang0505/SGLFT. Kaibing Zhang, Zhouwei Cheng, Xin He 0029, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Multi-Modality Regional Alignment Network for Covid X-Ray Survival Prediction and Report GenerationabstractIn response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic analysis. This study proposes a Multi-modality Regional Alignment Network (MRANet), an explainable model for radiology report generation and survival prediction that focuses on high-risk regions. By learning spatial correlation in the detector, MRANet visually grounds region-specific descriptions, providing robust anatomical regions with a completion strategy. The visual features of each region are embedded using a novel survival attention mechanism, offering spatially and risk-aware features for sentence encoding while maintaining global coherence across tasks. A cross-domain LLMs-Alignment is employed to enhance the image-to-text transfer process, resulting in sentences rich with clinical detail and improved explainability for radiologists. Multi-center experiments validate the overall performance and each module's composition within the model, encouraging further advancements in radiology report generation research emphasizing clinical interpretation and trustworthiness in AI models applied to medical studies. Zhusi Zhong, Jie Li 0001, John Sollee, Scott Collins, Harrison X. Bai, Terrance Healey, Michael Atalay, Xinbo Gao 0001, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Boosting Modal-Specific Representations for Sentiment Analysis With Incomplete ModalitiesabstractMultimodal sentiment analysis aims at exploiting complementary information from multiple modalities or data sources to enhance the understanding and interpretation of sentiment. While existing multi-modal fusion techniques offer significant improvements in sentiment analysis, real-world scenarios often involve missing modalities, introducing complexity due to uncertainty of which modalities may be absent. To tackle the challenge of incomplete modality-specific feature extraction caused by missing modalities, this paper proposes a Cosine Margin-Aware Network (CMANet) which centers on the Cosine Margin-Aware Distillation (CMAD) module. The core module measures distance between samples and the classification boundary, enabling CMANet to focus on samples near the boundary. So, it effectively captures the unique features of different modal combinations. To address the issue of modality imbalance during modality-specific feature extraction, this paper proposes a Weak Modality Regularization (WMR) strategy, which aligns the feature distributions between strong and weak modalities at the dataset-level, while also enhancing the prediction loss of samples at the sample-level. This dual mechanism improves the recognition robustness of weak modality combination. Extensive experiments demonstrate that the proposed method outperforms the previous best model, MMIN, with a 3.82% improvement in unweighted accuracy. These results underscore the robustness of the approach under conditions of uncertain and missing modalities. Lihuo He, Fei Gao 0006, Kaifan Zhang, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | ETC: Temporal Boundary Expand Then Clarify for Weakly Supervised Video Grounding With Multimodal Large Language ModelabstractEarly weakly supervised video grounding (WSVG) methods often struggle with incomplete boundary detection due to the absence of temporal boundary annotations. To bridge the gap between video-level and boundary-level annotations, explicit supervision methods (i.e., generating pseudo-temporal boundaries for training) have achieved great success. However, data augmentation in these methods might disrupt critical temporal information, yielding poor pseudo-temporal boundaries. In this paper, we propose a new perspective that maintains the integrity of the original temporal content while introducing more valuable information for expanding the incomplete boundaries. To this end, we proposeETC(ExpandthenClarify), first using the additional information to expand the initial incomplete pseudo-temporal boundaries, and subsequently refining these expanded ones to achieve precise boundaries. Motivated by video continuity, i.e., visual similarity across adjacent frames, we use powerful multi-modal large language models (MLLMs) to annotate each frame within the initial pseudo-temporal boundaries, yielding more comprehensive descriptions for expanded boundaries. To further clarify the noise in expanded boundaries, we combine mutual learning with a tailored proposal-level contrastive objective to use a learnable approach to harmonize a balance between incomplete yet clean (initial) and comprehensive yet noisy (expanded) boundaries for more precise ones. Experiments demonstrate the superiority of our method on two challenging WSVG datasets. Guozhang Li, Xinpeng Ding, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Language Knowledge-Assisted Representation Learning for Skeleton-Based Action RecognitionabstractHow humans understand and recognize the actions of others is a complex neuroscientific problem that involves a combination of cognitive mechanisms and neural networks. Research has shown that humans have brain areas that recognize actions that process top-down attentional information, such as the temporoparietal association area. Also, humans have brain regions dedicated to understanding the minds of others and analyzing their intentions, such as the medial prefrontal cortex of the temporal lobe. Skeleton-based action recognition creates mappings for the complex connections between the human skeleton movement patterns and behaviors. Although existing studies encoded meaningful node relationships and synthesized action representations for classification with good results, few of them considered incorporating a priori knowledge to aid potential representation learning for better performance. LA-GCN proposes a graph convolution network using large-scale language models (LLM) knowledge assistance. First, the LLM knowledge is mapped into a priori global relationship (GPR) topology and a priori category relationship (CPR) topology between nodes. The GPR guides the generation of new “bone” representations, aiming to emphasize essential node information from the data level. The CPR mapping simulates category prior knowledge in human brain regions, encoded by the PC-AC module and used to add additional supervision—forcing the model to learn class-distinguishable features. In addition, to improve information transfer efficiency in topology modeling, we propose multi-hop attention graph convolution. It aggregates each node's k-order neighbor simultaneously to speed up model convergence. LA-GCN reaches state-of-the-art on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets. Yan Gao 0025, Zheng Hui, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | An Information Compensation Framework for Zero-Shot Skeleton-Based Action RecognitionabstractZero-shot human skeleton-based action recognition aims to construct a model that can recognize actions outside the categories seen during training. Previous research has focused on aligning sequences' visual and semantic spatial distributions. However, these methods extract semantic features simply. They ignore that proper prompt design for rich and fine-grained action cues can provide robust representation space clustering. In order to alleviate the problem of insufficient information available for skeleton sequences, we design an information compensation learning framework from an information-theoretic perspective to improve zero-shot action recognition accuracy with a multi-granularity semantic interaction mechanism. Inspired by ensemble learning, we propose a multi-level alignment (MLA) approach to compensate information for action classes. MLA aligns multi-granularity embeddings with visual embedding through a multi-head scoring mechanism to distinguish semantically similar action names and visually similar actions. Furthermore, we introduce a new loss function sampling method to obtain a tight and robust representation. Finally, these multi-granularity semantic embeddings are synthesized to form a proper decision surface for classification. Significant action recognition performance is achieved when evaluated on the challenging NTU RGB+D, NTU RGB+D 120, and PKU-MMD benchmarks and validate that multi-granularity semantic features facilitate the differentiation of action clusters with similar visual features. Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | I²NQ: Inter and Intra Nonuniform Quantization for Single Image Super-ResolutionabstractQuantizing neural network is an efficient model compression technique that converts weights and activations from floating-point to integer. However, existing model quantization methods are primarily designed for high-level visual tasks. They do not sufficiently consider the unique characteristics of feature distribution in image super-resolution (SR) reconstruction models. On the one hand, the objective of SR is to restore high-frequency and fine-detail information while preserving the overall feature distribution. Therefore, the regularization techniques are removed to maintain the original distribution. However, vanilla quantization methods often employ regularization techniques to normalize the features for stable network training, which destroys the inherent information of the feature distribution. On the other hand, the feature distribution in SR models exhibits a nonuniform bell-shaped form. Common quantization methods adopt a uniform quantization strategy with equal quantization intervals. This fails to effectively capture the nonuniform feature distribution in SR. To address the above issue, we propose a novel method named Inter and Intra Nonuniform Quantization, which takes into account the specific characteristics of the feature distribution in the context of SR reconstruction models. Additionally, we propose a weight adjustment method called flex-scale-weight-adjust (FSWA). It can maintain the diversity of weight information and reduce quantization errors. Extensive experiments demonstrate that our proposed method surpasses other quantization methods in both the evaluation of reconstruction metrics and visual reconstruction performance. Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Rectified Binary Network for Single-Image Super-ResolutionabstractBinary neural network (BNN) is an effective approach to reduce the memory usage and the computational complexity of full-precision convolutional neural networks (CNNs), which has been widely used in the field of deep learning. However, there are different properties between BNNs and real-valued models, making it difficult to draw on the experience of CNN composition to develop BNN. In this article, we study the application of binary network to the single-image super-resolution (SISR) task in which the network is trained for restoring original high-resolution (HR) images. Generally, the distribution of features in the network for SISR is more complex than those in recognition models for preserving the abundant image information, e.g., texture, color, and details. To enhance the representation ability of BNN, we explore a novel activation-rectified inference (ARI) module that achieves a more complete representation of features by combining observations from different quantitative perspectives. The activations are divided into several parts with different quantification intervals and are inferred independently. This allows the binary activations to retain more image detail and yield finer inference. In addition, we further propose an adaptive approximation estimator (AAE) for gradually learning the accurate gradient estimation interval in each layer to alleviate the optimization difficulty. Experiments conducted on several benchmarks show that our approach is able to learn a binary SISR model with superior performance over the state-of-the-art methods. The code will be released at https://github.com/jwxintt/Rectified-BSR. Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object TrackingabstractThe global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature poses storage space requirements that challenge algorithm handling long videos. Currently, commonly used methods are still generated trajectories by building one-forward associations across frames. Such matches produced under the guidance of first-order similarity information may not be optimal from a longer-time perspective. Moreover, they often lack an end-to-end scheme for correcting mismatches. This paper proposes the Composite Node Message Passing Network (CoNo-Link), a multi-scene generalized framework for modeling ultra-long frames information for association. CoNo-Link's solution is a low-storage overhead method for building constrained connected graphs. In addition to the previous method of treating objects as nodes, the network innovatively treats object trajectories as nodes for information interaction, improving the graph neural network's feature representation capability. Specifically, we formulate the graph-building problem as a top-k selection task for some reliable objects or trajectories. Our model can learn better predictions on longer-time scales by adding composite nodes. As a result, our method outperforms the state-of-the-art in several commonly used datasets. Yan Gao 0025, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
AAAI | 3 |
| 2024 | Trades++: Enhancing Multi-Object Tracking of Real Low Confidence Targets Using a Pyramid-Like Self-Attention ModelabstractIn reality, multi-object tracking (MOT) is used in a wide range of scenarios. Maintaining the motion trajectory of the target, especially in high-density pedestrian scenarios, is often difficult. The tracking quality of most multi-object trackers correlates strongly with the detector quality and they often ignore the low-scoring detection boxes obtained by the detector. In this paper, we propose a TraDeS-based method called TraDeS++ that enhances the detection features using a pyramid-like self-attention model, significantly reducing the model training time and achieving a reduction of half the training epochs. The second motivation is to focus on the association method. We use a two-stage matching strategy with GIoU constraints, effectively improving HOTA. Experimental results show that our component effectively improves the metrics of MOT, especially MOTA, HOTA, and IDF1. Competitive results are achieved on the popular MOT16 and MOT17 datasets. Chenxin Wen, Yan Gao 0025, Jie Li 0001 |
ICASSP | 3 |
| 2024 | Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search
Haichun Tai, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IJCAI | 3 |
| 2024 | Advancing Generalized Deepfake Detector with Forgery Perception GuidanceabstractOne of the serious impacts brought by artificial intelligence is the abuse of deepfake techniques. Despite the proliferation of deepfake detection methods aimed at safeguarding the authenticity of media across the Internet, they mainly consider the improvement of detector architecture or the synthesis of forgery samples. The forgery perceptions, including the feature responses and prediction scores for forgery samples, have not been well considered. As a result, the generalization across multiple deepfake techniques always comes with complicated detector structures and expensive training costs. In this paper, we shift the focus to real-time perception analysis in the training process and generalize deepfake detectors through an efficient method dubbed Forgery Perception Guidance (FPG). In particular, after investigating the deficiencies of forgery perceptions, FPG adopts a sample refinement strategy to pertinently train the detector, thereby elevating the generalization efficiently. Moreover, FPG introduces more sample information as explicit optimizations, which makes the detector further adapt the sample diversities. Experiments demonstrate that FPG improves the generality of deepfake detectors with small training costs, minor detector modifications, and the acquirement of real data only. In particular, our approach not only outperforms the state-of-the-art on both the cross-dataset and cross-manipulation evaluation but also surpasses the baseline that needs more than 3× training time. Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Lin Yuan 0002, Shuodi Wang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
ACM Multimedia | 6 |
| 2024 | Coordinate Attention Guided Dual-Teacher Adaptive Knowledge Distillation for image classification
Dongtong Ma, Kaibing Zhang, Qizhi Cao, Jie Li 0001, Xinbo Gao 0001 |
Expert Syst. Appl. | 4 |
| 2024 | A multi-scale information integration framework for infrared and visible image fusion
Jie Li 0001, Hanxiao Lei, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2024 | A locally weighted, correlated subdomain adaptive network employed to facilitate transfer learning
Tuo Xu, Bing Han 0003, Jie Li 0001, Yuefan Du |
Image Vis. Comput. | 3 |
| 2024 | ProFPN: Progressive feature pyramid network with soft proposal assignment for object detection
Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Xinbo Gao 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Brain-driven facial image reconstruction via StyleGAN inversion with improved identity consistency
Ziqi Ren, Jie Li 0001, Lukun Wu, Xuetong Xue, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2024 | Coupled discriminative manifold alignment for low-resolution face recognition
Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2024 | A Non-Local Block With Adaptive Regularization StrategyabstractNon-local block (NLB) is a breakthrough technology in computer vision. It greatly boosts the capability of deep convolutional neural networks (CNNs) to capture long-range dependencies. As the critical component of NLB, non-local operation can be considered a network-based implementation of the well-known non-local means filter (NLM). Drawing on the solid theoretical foundation of NLM, we provide an innovative interpretation of the non-local operation. Specifically, it is formulated as an optimization problem regularized by Shannon entropy with a fixed parameter. Building on this insight, we further introduce an adaptive regularization strategy to enhance NLB and get a novel non-local block named ARNLB. Preliminary experiments on semantic segmentation demonstrate its effectiveness. Zhonggui Sun, Huichao Sun, Jie Li 0001, Xinbo Gao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Deep Convolution Modulation for Image Super-ResolutionabstractRecently, deep-learning-based super-resolution methods have achieved excellent performances, but mainly focus on training a single generalized deep network by feeding numerous samples. Yet intuitively, each image has its specific representation, and is expected to acquire an adaptive model. For this issue, we propose a novel convolution modulation (CoMo) mechanism to build image-specific deep networks, by exploiting the principal information of the feature to generate a modulation weight, and thereby adaptively modulating the kernel weights of convolution without any additional parameters, which outperforms the vanilla convolution and several existing attention mechanisms when embedding into the state-of-the-art architectures. To optimize the modulated convolutions in mini-batch training, we introduce an image-specific optimization (IsO) algorithm, which tackles the infeasibility of the conventional optimization algorithms on this issue. Furthermore, we investigate the effect of CoMo on state-of-the-art architectures and design a new CoMoNet architecture by employing the U-style residual learning and hourglass dense block learning, which is an appropriate architecture to utmost improve the effectiveness of CoMo theoretically. Extensive experiments on benchmarks show that the proposed methods achieve superior performances and higher flexibility against the state-of-the-art SISR and blind SR methods. The code is available at github.com/YuanfeiHuang/CoMoNet. Yuanfei Huang, Jie Li 0001, Yanting Hu, Hua Huang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Universal Heterogeneous Face Analysis via Multi-Domain Feature DisentanglementabstractHeterogeneous face analysis is an important and challenge problem in face recognition community, because of the large modality discrepancy between heterogeneous face images. Existing methods either focus on transforming heterogeneous faces into the same style via face synthesis process, or intend to directly recognize heterogeneous face via modality invariant descriptors. However, the tasks of cross modality face synthesis and face recognition share a common purpose, which is to disentangle an inherent explainable representation. To this end, we propose a novel universal heterogenous face analysis method via multi-domain feature disentanglement, which does not need any face domain label. The proposed method explores to disentangle factors of variations of cross modality faces in an unsupervised manner. Then we could translate cross modality faces through modifying semantic factors, and the extracted inherent explainable representation still maintains being discriminative for heterogeneous face recognition. Experimental results on multiple cross modality face databases demonstrate the effectiveness of the proposed method. These experimental results also inspire us that the unsupervised disentangled module could help to analyze the interpretability of heterogenous face representation. Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | MMNet: Multi-Collaboration and Multi-Supervision Network for Sequential Deepfake DetectionabstractAdvanced manipulation techniques have provided criminals with opportunities to make social panic or gain illicit profits through the generation of deceptive media, such as forgery face images. In response, various deepfake detection methods have been proposed to assess image authenticity. Sequential deepfake detection, which is an extension of deepfake detection, aims to identify forged facial regions with the correct sequence for recovery. Nonetheless, due to the different combinations of spatial and sequential manipulations, forgery face images exhibit substantial discrepancies that severely impact detection performance. Additionally, the recovery of forged images requires knowledge of the manipulation model to implement inverse transformations, which is difficult to ascertain as relevant techniques are often concealed by attackers. To address these issues, we propose Multi-Collaboration and Multi-Supervision Network (MMNet) that handles various spatial scales and sequential permutations in forgery face images and achieve recovery without requiring knowledge of the corresponding manipulation method. Furthermore, existing evaluation metrics only consider detection accuracy at a single inferring step, without accounting for the matching degree with ground-truth under continuous multiple steps. To overcome this limitation, we propose a novel evaluation metric called Complete Sequence Matching (CSM), which considers the detection accuracy at multiple inferring steps, reflecting the ability to detect integrally forged sequences. Extensive experiments on several typical datasets demonstrate that MMNet achieves state-of-the-art detection performance and independent recovery performance. Code will be available at https://github.com/xarryon/MMNet. Ruiyang Xia, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Large Pose Face Recognition via Facial Representation LearningabstractOvercoming image acquisition perspectives and face pose variations is a key problem in unconstrained face recognition tasks. One of the practical approaches is by reconstructing the face with extreme pose into a version that is more easily recognized by the discriminator, such as a frontal face. Often, existing methods attempt to balance the accuracy of downstream tasks with human visual perception, but ignore the differences in propensity between the two. Besides, large-scale datasets of profile-frontal paired face images are absent, which further hinders the training of models. In this work, we investigate a variety of face reconstruction approaches and propose a very simple, but very effective method to match face images across different scenes, named facial representation learning (FRL). The core idea of FRL is to introduce a representation generator in front of a pre-trained face recognition model, which can extract face representations from arbitrary faces that are more suitable for recognition model discrimination. In particular, the representation generator reconstructs the facial representation by minimising identity differences from the frontal face and adds pixel-level and adversarial constraints to cater for discriminator preferences. Extensive benchmark experiments show that the proposed method not only achieves better performance than state-of-the-art methods, but also can further squeeze the inference potential of existing face recognition models. Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | BPMTrack: Multi-Object Tracking With Detection Box Application Pattern MiningabstractThe key to multi-object tracking is its stability and the retention of identity information. A common problem with most detection-based approaches is trusting and using all the detector outputs for the association. However, some settings of detectors can affect stable long-range tracking. Based on the principle of reducing the association noise in the detection processing step, we propose a new framework, the Box application Pattern Mining Tracker (BPMTrack), to address this issue. Specifically, we worked on three main aspects: output threshold, association strategy, and motion model. Due to the problem of inconsistency between classification scores and localization accuracy, we propose the Box Quality Estimation Network (BQENet) to predict the localization quality scores of all detections in the current frame, reserving high-quality boxes for the tracker. In addition, based on observations of intensive scenarios, we propose a simple and effective data association method, the Non-Maximum Suppression Integration (NMSI) matching strategy. It recovers the Non-Maximum Suppression (NMS) detection, inputs them into BQENet, and then performs hierarchical matching with reasonable control of box priority to alleviate the problem of absent objects caused by occlusion. Finally, we propose an improved Measurement Correct and Noise Scale (MCNS) Kalman algorithm to improve the prediction accuracy of object positions and, thus, the association quality. We performed an extensive ablation evaluation of the proposed framework to prove its effectiveness. Moreover, the three tracking benchmarks show our method's accuracy and long-distance performance. Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Neighbor-Guided Pseudo-Label Generation and Refinement for Single-Frame Supervised Temporal Action LocalizationabstractDue to the sparse single-frame annotations, current Single-Frame Temporal Action Localization (SF-TAL) methods generally employ threshold-based pseudo-label generation strategies. However, these approaches suffer from inefficient data utilization, as only parts of unlabeled frames with confidence scores surpassing a predefined threshold are selected for training. Moreover, the variability of single-frame annotations and unreliable model predictions introduce pseudo-label noise. To address these challenges, we propose two strategies by using the relationship of the video segments with their neighbors': 1) temporal neighbor-guided soft pseudo-label generation (TNPG); and 2) semantic neighbor-guided pseudo-label refinement (SNPR). TNPG utilizes a local-global self-attention mechanism in a transformer encoder to capture temporal neighbor information while focusing on the whole video. Then the generated self-attention map is multiplied by the network predictions to propagate information between labeled and unlabeled frames, and produce soft pseudo-label for all segments. Despite this, label noise persists due to unreliable model predictions. To mitigate this, SNPR refines pseudo-labels based on the assumption that predictions should resemble their semantic nearest neighbors'. Specifically, we search for semantic nearest neighbors of each video segment by cosine similarity in the feature space. Then the refined soft pseudo-labels can be obtained by a weight combination of the original pseudo-label and the semantic nearest neighbors'. Finally, the model can be trained with the refined pseudo-labels, and the performance has been greatly improved. Comprehensive experimental results on different benchmarks show that we achieve state-of-the-art performances on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets. Guozhang Li, De Cheng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Inspector for Face Forgery Detection: Defending Against Adversarial Attacks From Coarse to FineabstractThe emergence of face forgery has raised global concerns on social security, thereby facilitating the research on automatic forgery detection. Although current forgery detectors have demonstrated promising performance in determining authenticity, their susceptibility to adversarial perturbations remains insufficiently addressed. Given the nuanced discrepancies between real and fake instances are essential in forgery detection, previous defensive paradigms based on input processing and adversarial training tend to disrupt these discrepancies. For the detectors, the learning difficulty is thus increased, and the natural accuracy is dramatically decreased. To achieve adversarial defense without changing the instances as well as the detectors, a novel defensive paradigm called Inspector is designed specifically for face forgery detectors. Specifically, Inspector defends against adversarial attacks in a coarse-to-fine manner. In the coarse defense stage, adversarial instances with evident perturbations are directly identified and filtered out. Subsequently, in the fine defense stage, the threats from adversarial instances with imperceptible perturbations are further detected and eliminated. Experimental results across different types of face forgery datasets and detectors demonstrate that our method achieves state-of-the-art performances against various types of adversarial perturbations while better preserving natural accuracy. Code is available on https://github.com/xarryon/Inspector. Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | De-Biased Disentanglement Learning for Pulmonary Embolism Survival Prediction on Multimodal DataabstractHealth disparities among marginalized populations with lower socioeconomic status significantly impact the fairness and effectiveness of healthcare delivery. The increasing integration of artificial intelligence (AI) into healthcare presents an opportunity to address these inequalities, provided that AI models are free from bias. This paper aims to address the bias challenges by population disparities within healthcare systems, existing in the presentation of and development of algorithms, leading to inequitable medical implementation for conditions such as pulmonary embolism (PE) prognosis. In this study, we explore the diverse bias in healthcare systems, which highlights the demand for a holistic framework to reducing bias by complementary aggregation. By leveraging de-biasing deep survival prediction models, we propose a framework that disentangles identifiable information from images, text reports, and clinical variables to mitigate potential biases within multimodal datasets. Our study offers several advantages over traditional clinical-based survival prediction methods, including richer survival-related characteristics and bias-complementary predicted results. By improving the robustness of survival analysis through this framework, we aim to benefit patients, clinicians, and researchers by enhancing fairness and accuracy in healthcare AI systems. Zhusi Zhong, Jie Li 0001, Helen Zhang, Fayez H. Fayad, Yang Li 0111, Scott Collins, Harrison X. Bai, Sun Ho Ahn, Michael Atalay, Xinbo Gao 0001, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | VLDadaptor: Domain Adaptive Object Detection With Vision-Language Model DistillationabstractDomain adaptive object detection (DAOD) aims to develop a detector trained on labeled source domains to identify objects in unlabeled target domains. A primary challenge in DAOD is the domain shift problem. Most existing methods learn domain-invariant features within single domain embedding space, often resulting in heavy model biases due to the intrinsic data properties of source domains. To mitigate the model biases, this paper proposes VLDadaptor, a domain adaptive object detector based on vision-language models (VLMs) distillation. Firstly, the proposed method integrates domain-mixed contrastive knowledge distillation between the visual encoder of CLIP and the detector by transferring category-level instance features, which guarantees the detector can extract domain-invariant visual instance features across domains. Then, VLDadaptor employs domain-mixed consistency distillation between the text encoder of CLIP and detector by aligning text prompt embeddings with visual instance features, which helps to maintain the category-level feature consistency among the detector, text encoder and the visual encoder of VLMs. Finally, the proposed method further promotes the adaptation ability by adopting a prompt-based memory bank to generate semantic-complete features for graph matching. These contributions enable VLDadaptor to extract visual features into the visual-language embedding space without any evident model bias towards specific domains. Extensive experimental results demonstrate that the proposed method achieves state-of-the-art performance on Pascal VOC to Clipart adaptation tasks and exhibits high accuracy on driving scenario tasks with significantly less training time. Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Di Wang 0011, Xinbo Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Toward Pixel-Level Precision for Binary Super-Resolution With Mixed Binary RepresentationabstractBinary neural network (BNN) is an effective method for reducing model computational and memory cost, which has achieved much progress in the super-resolution (SR) field. However, there is still a noticeable performance gap between a binary SR network and its full-precision counterpart. Considering that the information density in quantization features is far lower than full-precision features, we aim to improve the precision of quantization features to produce rich-enough output activations for SR task. First, we make several observations that a multibit value could be approximated by multiple 1-bit values, and the computation power of binary convolution could be improved by approximating the multibit convolution process. Then, we propose a mixed binary representation set to approximate multibit activations, which is effective in compensating the quantization precision loss. Finally, we present a new precision-driven binary convolution (PDBC) module, which increases the convolution precision and protects image detail information without extra computation. Compared with normal binary convolution, our method could largely reduce the information loss caused by binarization. In experiments, our methods consistently show superior performance over the baseline models and can surpass state-of-the-art methods in terms of peak signal to noise ratio (PSNR) and visual quality. Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Weakly Supervised Temporal Action Localization With Bidirectional Semantic Consistency ConstraintabstractWeakly supervised temporal action localization (WTAL) aims to classify and localize temporal boundaries of actions for the video, given only video-level category labels in the training datasets. Due to the lack of boundary information during training, existing approaches formulate WTAL as a classification problem, i.e., generating the temporal class activation map (T-CAM) for localization. However, with only classification loss, the model would be suboptimized, i.e., the action-related scenes are enough to distinguish different class labels. Regarding other actions in the action-related scene (i.e., the scene same as positive actions) as co-scene actions, this suboptimized model would misclassify the co-scene actions as positive actions. To address this misclassification, we propose a simple yet efficient method, named bidirectional semantic consistency constraint (Bi-SCC), to discriminate the positive actions from co-scene actions. The proposed Bi-SCC first adopts a temporal context augmentation to generate an augmented video that breaks the correlation between positive actions and their co-scene actions in the inter-video. Then, a semantic consistency constraint (SCC) is used to enforce the predictions of the original video and augmented video to be consistent, hence suppressing the co-scene actions. However, we find that this augmented video would destroy the original temporal context. Simply applying the consistency constraint would affect the completeness of localized positive actions. Hence, we boost the SCC in a bidirectional way to suppress co-scene actions while ensuring the integrity of positive actions, by cross-supervising the original and augmented videos. Finally, our proposed Bi-SCC can be applied to current WTAL approaches and improve their performance. Experimental results show that our approach outperforms the state-of-the-art methods on THUMOS14 and ActivityNet. The code is available at https://github.com/lgzlIlIlI/BiSCC. Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Local Means Binary Networks for Image Super-ResolutionabstractThe success of modern single image super-resolution (SISR) algorithms is inspired by the development of deep convolutional neural networks (CNNs). However, these CNN-based methods require considerable computation and complexity, making it impossible for these methods to perform real-time calculations in edge devices. Thus, lightweight model design has become a development trend in the super-resolution field, including pruning, quantization, and other methods. The 1-bit quantization is an extreme lightweight method which can reduce the calculation amount of the model in an extreme manner and is friendly to hardware such as edge devices. Most existing binary quantization approaches lead to a large information loss during forward propagation, especially in detailed color information (e.g., edge, texture, and contrast). The loss of color information makes modern binary methods unsuitable for SISR tasks. We think the loss occurs because these methods typically utilize a uniform threshold to quantize the weights and activations. Thus, in this article, we thoroughly analyze the difference between normal classification tasks and SISR tasks, and present a binarization scheme based on local means. The proposed method can maintain more detailed information in feature maps using dynamic thresholds during quantization. Specifically, each value in the full precision activations has a corresponding threshold during the quantization process, and those thresholds are determined by the full precision values of the surroundings. In addition, a gradient approximator is introduced to adaptively optimize the gradient for updating binary weights. We then verify the effectiveness of our method for training binary networks on several SISR benchmarks including VDSR and SRResNet. Experimental results show that the proposed method can outperform the state-of-the-art algorithms to obtain binary networks for image super-resolution with better peak signal-to-noise ratio (PSNR) values and visual quality. Nannan Wang 0001, Jingwei Xin, Jie Li 0001, Xinbo Gao 0001, Kai Han 0002, Yunhe Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identificationabstractVisible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to reduce the discrepancy is to filter out identity-irrelevant interference and effectively learn modality-invariant person representations. In this paper, we propose a novel Modality Restitution and Compensation Network (MRCN) to narrow the gap between the two modalities. Specifically, we first reduce the modality discrepancy by using two Instance Normalization (IN) layers. Next, to reduce the influence of IN layers on removing discriminative information and to reduce modality differences, we propose a Modality Restitution Module (MRM) and a Modality Compensation Module (MCM) to respectively distill modality-irrelevant and modality-relevant features from the removed information. Then, the modality-irrelevant features are used to restitute to the normalized visible and infrared features, while the modality-relevant features are used to compensate for the features of the other modality. Furthermore, to better disentangle the modality-relevant features and the modality-irrelevant features, we propose a novel Center-Quadruplet Causal (CQC) loss to encourage the network to effectively learn the modality-relevant features and the modality-irrelevant features. Extensive experiments are conducted to validate the superiority of our method on the challenging SYSU-MM01 and RegDB datasets. More remarkably, our method achieves 95.1% in terms of Rank-1 and 89.2% in terms of mAP on the RegDB dataset. Yan Yan 0001, Jie Li 0001, Hanzi Wang |
AAAI | 3 |
| 2023 | Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object TrackingabstractRecently, some weakly supervised multi-object tracking (MOT) methods learn identity embedding features with pseudo identity labels rather than the high-cost manual ones. However, these pseudo identity labels may contain many false or missing identities, which adversely affect the optimization of tracking networks, resulting in interrupted trajectories of occluded targets. To effectively reconnect the interrupted trajectories caused by noisy pseudo labels, we propose a novel weakly supervised MOT method based on a Trajectory-Reconnecting Transformer (TRTMOT). TRT-MOT performs feature decoupling to extract discriminative embedding features for reconnecting trajectories of occluded targets. Experimental results show that TRTMOT outperforms previous weakly supervised MOT methods by at least +3.6 and +5.6 on MOTA for the MOT17 and MOT20 datasets, respectively. Yu-Lei Li, Yang Lu 0009, Jie Li 0001, Hanzi Wang |
ICASSP | 3 |
| 2023 | A Dual-Stream Convolutional Feature Fusion Network for Hyperspectral UnmixingabstractAs an important research element in the field of hyperspectral remote sensing, the purpose of hyperspectral unmixing is to decompose the mixed pixels in a hyperspectral image into endmembers and abundance. With the development of deep learning, it also has good prospects for application in the field of hyperspectral unmixing. Existing hyperspectral unmixing networks tend to focus only on the spectral information in the image which neglecting the spatial information. Therefore, a network based on dual-stream convolutional feature fusion for hyperspectral unmixing is proposed to make full use of the correlation between neighboring pixels in the paper, which includes three parts: dual-stream feature extraction, feature fusion and abundance estimation. The unmixing performance of the network is verified on two real datasets and exhibits better performance compared with other state-of-the-art methods. Haoyue Hua, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001 |
IGARSS | 2 |
| 2023 | 3D-Mglnet: Moving Vehicle Detection in Satellite Videos with 3D Motion-Guided Lightweight NetworkabstractObject detectors based on convolutional neural networks have been widely-applied to detect moving vehicles in satellite videos. However, many detectors render superior detection accuracy at the expense of increased computational complexity and decreased inference speed. This prevents these detectors from being deployed into mobile devices. In this paper, an efficient 3D motion-guided lightweight network (3D-MGLNet) is proposed. Specifically, 3D-MGLNet constructs a motion-guided module based on 3D convolution to extract motion cues from spatial-temporal information. This module uses model compression strategies to detect moving vehicles in real-time while following the principle of "fewer channels, smaller convolution kernels," significantly reducing the number of parameters and computational complexity. Extensive experiments are conducted on the Jilin-1 and SkySat satellite video datasets. The results demonstrate that 3D-MGLNet gains strong performance by striking an excellent tradeoff between resource and accuracy, resulting in the fewest parameters (0.35M) and fastest speed (66.84 fps) compared to other popular models. Jie Li 0001, Jie Feng 0003, Quanpeng Jiang, Xiangrong Zhang, Licheng Jiao |
IGARSS | 2 |
| 2023 | Improving Outcome Prediction of Pulmonary Embolism by De-biased Multi-modality Model
Zhusi Zhong, Jie Li 0001, Yang Li 0111, Fayez H. Fayad, Helen Zhang, Sun Ho Ahn, Harrison X. Bai, Xinbo Gao 0001, Michael Atalay, Zhicheng Jiao |
MICCAI (5) | 2 |
| 2023 | Spatio-Temporal Self-supervision for Few-Shot Action Recognition
Wanchuan Yu, Hanyu Guo, Yan Yan 0001, Jie Li 0001, Hanzi Wang |
PRCV (1) | 4 |
| 2023 | Advanced Binary Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Multi-modal deep convolutional dictionary learning for image denoising
Zhonggui Sun, Huichao Sun, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2023 | Reconstructing controllable faces from brain activity with hierarchical multiview representations
Ziqi Ren, Jie Li 0001, Xuetong Xue, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Xinbo Gao 0001 |
Neural Networks | 2 |
| 2023 | Transitional Learning: Exploring the Transition States of Degradation for Blind Super-resolutionabstractBeing extremely dependent on iterative estimation of the degradation prior or optimization of the model from scratch, the existing blind super-resolution (SR) methods are generally time-consuming and less effective, as the estimation of degradation proceeds from a blind initialization and lacks interpretable representation of degradations. To address it, this article proposes a transitional learning method for blind SR using an end-to-end network without any additional iterations in inference, and explores an effective representation for unknown degradation. To begin with, we analyze and demonstrate the transitionality of degradations as interpretable prior information to indirectly infer the unknown degradation model, including the widely used additive and convolutive degradations. We then propose a novel Transitional Learning method for blind Super-Resolution (TLSR), by adaptively inferring a transitional transformation function to solve the unknown degradations without any iterative operations in inference. Specifically, the end-to-end TLSR network consists of a degree of transitionality (DoT) estimation network, a homogeneous feature extraction network, and a transitional learning module. Quantitative and qualitative evaluations on blind SR tasks demonstrate that the proposed TLSR achieves superior performances and costs fewer complexities against the state-of-the-art blind SR methods. The code is available at github.com/YuanfeiHuang/TLSR. Yuanfei Huang, Jie Li 0001, Yanting Hu, Xinbo Gao 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning a High Fidelity Identity Representation for Face FrontalizationabstractThis paper considers the problem of face frontalization in the wild, which transforms a face image with profile views into a frontal face. Face frontalization provides an effective solution to the face recognition problem in uncontrolled scenes. However, the existing methods either focus on deep learning techniques as an end-to-end framework or combine other explicit facial prior estimation tasks, such as 3D representation, optical flow estimation and so on, where computation is highly redundant and facial identity cannot be well represented. In this paper, we focus on how to maximise the potential of the model for identity learning and representation, and propose an accurate and lightweight face frontalization approach, named identity-preserving model (IPM). IPM has a well-designed encoder-decoder architecture which restores input face to a frontal counterpart. The encoder is constructed to extract representation from the input face, where a contrastive loss function is applied that encourages representations to form compact clusters, while preserving their relationships across the corpora. Then a cross-domain rectification module is proposed to eliminate the representation differences between the recognition and reconstruction domains, thus improving the accuracy of the reconstructed face. Extensive experiments on benchmark datasets show that the proposed IPM approach not only outperforms the state-of-the-art on public datasets but also can cope with images in the uncontrolled scenes. Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Unsupervised Across Domain Consistency- Difference Network for Hyperspectral Image Super-ResolutionabstractWithout reducing the spectral resolution, hyperspectral image super-resolution has achieved remarkable progress thanks to the success of deep neural networks. However, existing methods can not fully excavate the latent high-frequency details only in the single spatial domain. Different from existing methods that only achieves the super-resolution task in spatial domain, we optimize the amplitude spectrum and phase spectrum in frequency domain to obtain high resolution hyperspectral image (HR-HSI). We propose a new unsupervised framework to reconstruct HR-HSI using only the observed low resolution HSI and HR multispectral image. Based on triple-level modeling, the encoder-decoder learns abundant features including contextual information from multiple scales. In addition, we propose iterative across domain consistency-difference (ADCD) module, which is embedded between encoder and decoder. In ADCD module, three parallel convolution streams, (amplitude spectrum adjustment branch, phase spectrum adjustment branch and spatial domain branch) are used to explore the consistency-difference between each other, which is preserved by memory units within the module. Particularly, we embed the dilated causal convolution in the frequency domain processing branch, which is convenient to flexibly adjust the receptive field and adapt to different domains. Extensive experiments are conducted on widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the proposed method. Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Task-Specific Heterogeneous Network for Object Detection in Aerial ImagesabstractObject detection in aerial images has attracted increasing attention in recent years. Due to the complex background and arbitrary-oriented objects, it is challenging to accurately locate the objects of interest in the images. Many methods have been developed for improving localization accuracy of oriented objects. However, classification and localization tasks require different features due to the unique characteristics of aerial images, which are still not fully considered in previous methods. Therefore, we propose a Task-Specific Heterogeneous (TSH) network for aerial object detection. Specifically, we design an Interference-Suppression Module (ISM) to reduce both the background and inter-class interference, which can provide discriminative features for classification. To produce more reliable localization confidence, we propose a Joint-learning Quality Estimation (JQE) module to adaptively combine the classification and regression features, thereby achieving accurate classification and localization quality estimation simultaneously. Moreover, we propose a Point-Based Localization (PBL) branch. In the PBL, the learnable points can effectively adapt to objects with diverse shapes and orientations, and the dynamic information aggregation module can enhance the relationships between the dispersible points to promote localization accuracy. The proposed TSH is evaluated extensively on four widely used aerial datasets, demonstrating its state-of-the-art performance. Ablation study and visualizations further verify the effectiveness of our method. Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Refined Hybrid Network for Object Detection in Aerial ImagesabstractAerial object detection is a challenging task that needs to detect objects with large variations in scale and orientation. Previous dense object detectors rely on heuristic Non-Maximum Suppression (NMS) to filter out redundant detections. This may reduce the recall rate for objects with arbitraty orientations and large aspect ratios. Recently proposed sparse object detectors treat object detection as a set prediction task, effectively eliminating the need for hand-crafted components. However, applying this paradigm directly to aerial images achieves inferior performance. In this paper, we develop an effective refined hybrid network for object detection in aerial images. Our method combines the advantages of both dense and sparse detectors, achieving outstanding performance for aerial objects with large variations. Specifically, considering the highly diverse orientations of objects, we first apply a dynamic query generation module to produce high-quality oriented queries. These queries can effectively locate the foreground objects in an image, ensuring a high recall rate. Then, the object queries are sent to a query decoder for further refinement. This refinement stage adopts one-to-one matching to eliminate the negative impact caused by NMS. Moreover, an adaptive feature fusion module is designed to learn stronger modeling capabilities for rotated objects at different scales. In addition, we propose a practical mixed query sampling strategy that utilizes many-to-one assignment as an auxiliary scheme to help the detector training. Extensive experiments conducted on several aerial datasets demonstrate the superior performance of the proposed method in comparison with other state-of-the-art approaches. Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | FABNet: Frequency-Aware Binarized Network for Single Image Super-ResolutionabstractRemarkable achievements have been obtained with binary neural networks (BNN) in real-time and energy-efficient single-image super-resolution (SISR) methods. However, existing approaches often adopt the Sign function to quantize image features while ignoring the influence of image spatial frequency. We argue that we can minimize the quantization error by considering different spatial frequency components. To achieve this, we propose a frequency-aware binarized network (FABNet) for single image super-resolution. First, we leverage the wavelet transformation to decompose the features into low-frequency and high-frequency components and then employ a "divide-and-conquer" strategy to separately process them with well-designed binary network structures. Additionally, we introduce a dynamic binarization process that incorporates learned-threshold binarization during forward propagation and dynamic approximation during backward propagation, effectively addressing the diverse spatial frequency information. Compared to existing methods, our approach is effective in reducing quantization error and recovering image textures. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed methods could surpass state-of-the-art approaches in terms of PSNR and visual quality with significantly reduced computational costs. Our codes are available at https://github.com/xrjiang527/FABNet-PyTorch. Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Multi-Branch and Progressive Network for Low-Light Image EnhancementabstractLow-light images incur several complicated degradation factors such as poor brightness, low contrast, color degradation, and noise. Most previous deep learning-based approaches, however, only learn the mapping relationship of single channel between the input low-light images and the expected normal-light images, which is insufficient enough to deal with low-light images captured under uncertain imaging environment. Moreover, too deeper network architecture is not conducive to recover low-light images due to extremely low values in pixels. To surmount aforementioned issues, in this paper we propose a novel multi-branch and progressive network (MBPNet) for low-light image enhancement. To be more specific, the proposed MBPNet is comprised of four different branches which build the mapping relationship at different scales. The followed fusion is performed on the outputs obtained from four different branches for the final enhanced image. Furthermore, to better handle the difficulty of delivering structural information of low-light images with low values in pixels, a progressive enhancement strategy is applied in the proposed method, where four convolutional long short-term memory networks (LSTM) are embedded in four branches and an recurrent network architecture is developed to iteratively perform the enhancement process. In addition, a joint loss function consisting of the pixel loss, the multi-scale perceptual loss, the adversarial loss, the gradient loss, and the color loss is framed to optimize the model parameters. To evaluate the effectiveness of proposed MBPNet, three popularly used benchmark databases are used for both quantitative and qualitative assessments. The experimental results confirm that the proposed MBPNet obviously outperforms other state-of-the-art approaches in terms of quantitative and qualitative results. The code will be available at https://github.com/kbzhang0505/MBPNet. Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001, Minqi Li |
IEEE Trans. Image Process. | 3 |
| 2022 | CSTNET: Enhancing Global-To-Local Interactions for Image CaptioningabstractImage captioning aims to generate descriptions of images, which requires capturing complex interactions between local regions and global context within image. However, effective global context modeling from image remains a challenging research topic. Existing approaches incorporate global-level information into the initialized input mainly based on transformer architecture. Unlike previous methods that may not be able to capture rich global contextual information, we propose a novel method named Context-Sensitive Transformer (CSTNet), which can discover the inherent global context and further empower the global-to-local interactions. Experimental results on the MSCOCO dataset show that the proposed model can significantly improve the performance of image captioning. Ying Wang 0007, HaiShun Chen, Jie Li 0001 |
ICIP | 4 |
| 2022 | Hyperspectral and Multispectral Data Fusion with 1D-Convolution on SpectrumabstractFusion of hyperspectral and multispectral data to obtain hyperspectral data with high-spatial-resolution has been an important topic in recent years. The fusion methods that from model-driven to data-driven have been proposed constantly. Data driven models, especially the deep neural networks, are widely used owing to their excellent performance. And the design of its structure is highly related to the task. It is noteworthy that the hyperspectral and multispectral spectrum contains abundant information. However, most neural net-work based approaches mainly focus on spatial features and ignore the spectral information, which is also very important for fusion task. In this paper, we propose a ID-convolutional neural network, it can extract hyperspectral and multispectral features, simultaneously, and also can capture the spectral correlation and cross-correlation between hyperspectral and multispectral spectrum. Groups of simulation experiments demonstrate that compared with the state-of-the-art methods, paying more attention to the spectrum can obtain better performance. Jinchi Xie, Ying Wang 0007, Jie Li 0001 |
IGARSS | 3 |
| 2022 | A local-nonlocal mathematical morphology
Zhonggui Sun, Meiqi Lyu, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2022 | Face photo-sketch synthesis via full-scale identity supervision
Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Qinghua Hu, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2022 | Spatiotemporal consistency-enhanced network for video anomaly detection
Jie Li 0001, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2022 | Collaborative boundary-aware context encoding networks for error map prediction
Chunna Tian, Xinbo Gao 0001, Jie Li 0001, Zhicheng Jiao, Zhusi Zhong |
Pattern Recognit. | 4 |
| 2022 | Systemic distortion analysis with deep distortion directed image quality assessment models
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001 |
Signal Process. Image Commun. | 4 |
| 2022 | Image Super-Resolution With Self-Similarity Prior Guided Network and Sample-Discriminating LearningabstractThe nonlocal self-similarity in natural image provides an effective prior for single image super-resolution (SISR), which is beneficial to contextual information capture and performance improvement, as demonstrated by conventional SISR methods. However, it is little explored to utilize this property in deep neural networks. In this paper, we propose a self-similarity prior guided (SSPG) network to incorporate self-similarity-based nonlocal operation into deep neural network for SISR. Specifically, we design a cross-scale nearest-neighbor residual (CSNNR) block via introducing cross-scale$k$-nearest neighbors (KNN) matching into a residual block, which can be flexibly integrated into deep networks to capture long-range correlations among multi-scale and multi-level features. Meanwhile, by stacking a CSNNR block and a sequence of wide-activated residual blocks with a local skip-connection, a multi-level residual self-similarity (MRSS) module is developed to effectively employ local and nonlocal information for detail recovery. Thus, through cascading multiple MRSS modules, the proposed SSPG network performs both self-similarity-based nonlocal operation and convolution-based local operation on multi-level features to reconstruct informative features for accurate SISR. In addition, for pursuing visually pleasing results, we apply our SSPG network to the perception-oriented SISR field by following the framework of generative adversarial networks. In particular, we explore a sample-discriminating learning mechanism based on the statistical descriptions of training samples, and include it in optimization procedure to automatically tune the contributions of different samples according to their characteristics and then focus the network on creating realistic results. Extensive quantitative and qualitative evaluations on benchmark datasets illustrate the superiority of our proposed models over the state-of-the-art methods for both distortion-oriented and perception-oriented image super-resolution tasks. Yanting Hu, Jie Li 0001, Yuanfei Huang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Locality-Adaptive Structured Dictionary Learning for Cross-Domain RecognitionabstractDictionary learning has achieved remarkable success on a wide range of machine learning-based applications. In this paper, a locality-adaptive structured dictionary learning (LASDL) algorithm for cross-domain recognition is proposed. In the LASDL, a projective structured double reconstruction strategy is developed to train class-oriented sub-dictionaries from the specific classes of cross-domain samples. The strategy benefits to make full advantage of the discriminative information of cross-domain data and bridge the distribution divergence between two different domains. Meanwhile, an adaptive geometrical structure preserving function is designed to not only preserve the local manifold structures spanned by the representation coefficient spaces of the source and target domains, but also impose a constraint that the coefficients should keep closer to their class centers, which is propitious to reduce the distribution divergence and make the representation more accurate. With the structured linear coding technique, the final cross-domain recognition can be efficiently performed by determining class-specific reconstruction error. The optimization of the proposed LASDL model can be efficiently solved by simple least square method and alternating direction method of multipliers (ADMM) algorithm. Extensive experimental results validate the superiority of the proposed algorithm in contrast to other state-of-the-art predecessors. Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | External-Internal Attention for Hyperspectral Image Super-ResolutionabstractIn recent years, hyperspectral image (HSI) super-resolution has made significant progress by leveraging convolution neural network. Existing methods with spectral or spatial attention, which only consider the spectral similarity or pixel-pixel similarity, ignore sample-sample correlations and sparsity. Therefore, based on the fusion of HSI and multispectral image, we propose a new HSI super-resolution model with external-internal attention. Instead of considering a single sample, external attention module is employed to exploit the incorporating correlations between different samples to get a better feature representation. In addition, an internal attention module based on non-local operation is designed to explore the long-range dependencies information. Particularly, oriented to high mapping precision and low computational cost inference, spherical locality sensitive hashing is used to divide features into different hash buckets so that every query point is calculated in the hash bucket assigned to it, rather than based a weight sum of features across all positions. The sequential external-internal attention greatly improves the generalization ability and robustness of the model by modeling at the dataset level and at the sample level. Extensive experiments are conducted on five widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the method we proposed. Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Cascade Rotated Anchor-Aided Detector for Ship Detection in Remote Sensing ImagesabstractAutomatic ship detection in high-resolution remote sensing images has attracted increasing research interest due to its numerous practical applications. However, there still exist challenges when directly applying state-of-the-art object detection methods to real ship detection, which greatly limits the detection accuracy. In this article, we propose a novel cascade rotated anchor-aided detection network to achieve high-precision performance for detecting arbitrary-oriented ships. First, we develop a data preprocessing embedded cascade structure to reduce large amounts of false positives generated on blank areas in large-size remote sensing images. Second, to achieve accurate arbitrary-oriented ship detection, we design a rotated anchor-aided detection module. This detection module adopts a coarse-to-fine architecture with a cascade refinement module (CRM) to refine the rotated boxes progressively. Meanwhile, it utilizes an anchor-aided strategy similar to anchor-free, thus breaking through the bottlenecks of anchor-based methods and leading to a more flexible detection manner. Besides, a rotated align convolution layer is introduced in CRM to extract features from rotated regions accurately. Experimental results on the challenging DOTA and HRSC2016 data sets show that the proposed method achieves 84.12% and 90.79% AP, respectively, outperforming other state-of-the-art methods. Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Object Detection for Aerial Images With Feature Enhancement and Soft Label AssignmentabstractObject detection in aerial images, different from general object detection, faces with several challenges such as arbitrary-oriented objects and extremely imbalanced foreground-background distribution. Although some recent proposed aerial object detection methods achieve promising results, they are mainly anchor-based detectors which rely heavily on pre-defined anchor boxes and the final detection performance is sensitive to anchor-related hyper-parameters. In contrast, in this paper, we present an anchor-free detector with feature enhancement and soft label assignment (FSDet) which adopts a simpler design and achieves competitive performance. Specifically, to address the feature misalignment for detecting oriented objects, we propose an oriented feature refinement module to align the features with oriented objects. To alleviate the background issue, we design a class-aware context aggregation module to integrate the intra-class context information and suppress background context. Moreover, we propose a soft label assignment mechanism to measure the weight of training samples within the arbitrary-oriented objects, which can concentrate more on representative items with regard to their potential to detect oriented objects, achieving a more stable optimization during training. Extensive experiments on several datasets suggest that the proposed method is superior to the state-of-the-art methods and achieves a better trade-off between speed and accuracy. Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | An Object Point Set Inductive Tracker for Multi-Object Tracking and SegmentationabstractMulti-object tracking and segmentation (MOTS) is a derivative task of multi-object tracking (MOT). The new setting encourages the learning of more discriminative high-quality embeddings. In this paper, we focus on the problem of exploring the relationship between the segmenter and the tracker, and propose an efficient Object Point set Inductive Tracker (OPITrack) based on it. First, we discover that after a single attention layer, the high-dimensional, key point embedding will show feature averaging. To alleviate this phenomenon, we propose an embedding generalization training strategy for sparse training and dense testing. This strategy allows the network to increase randomness in training and encourages the tracker to learn more discriminative features. In addition, to learn the desired embedding space, we propose a general Trip-hard sample augmentation loss. The loss uses patches that are not distinguishable by the segmenter to join the feature learning and force the embedding network to learn the difference between false positives and true positives. Our method was validated on two MOTS benchmark datasets and achieved promising results. In addition, our OPITrack can achieve better performance for the raw model while costing less video memory (VRAM) at training time. Yan Gao 0025, Yu Zheng 0006, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Seeking Subjectivity in Visual Emotion Distribution LearningabstractVisual Emotion Analysis (VEA), which aims to predict people's emotions towards different visual stimuli, has become an attractive research topic recently. Rather than a single label classification task, it is more rational to regard VEA as a Label Distribution Learning (LDL) problem by voting from different individuals. Existing methods often predict visual emotion distribution in a unified network, neglecting the inherent subjectivity in its crowd voting process. In psychology, the Object-Appraisal-Emotion model has demonstrated that each individual's emotion is affected by his/her subjective appraisal, which is further formed by the affective memory. Inspired by this, we propose a novel Subjectivity Appraise-and-Match Network (SAMNet) to investigate the subjectivity in visual emotion distribution. To depict the diversity in crowd voting process, we first propose the Subjectivity Appraising with multiple branches, where each branch simulates the emotion evocation process of a specific individual. Specifically, we construct the affective memory with an attention-based mechanism to preserve each individual's unique emotional experience. A subjectivity loss is further proposed to guarantee the divergence between different individuals. Moreover, we propose the Subjectivity Matching with a matching loss, aiming at assigning unordered emotion labels to ordered individual predictions in a one-to-one correspondence with the Hungarian algorithm. Extensive experiments and comparisons are conducted on public visual emotion distribution datasets, and the results demonstrate that the proposed SAMNet consistently outperforms the state-of-the-art methods. Ablation study verifies the effectiveness of our method and visualization proves its interpretability. Jingyuan Yang 0002, Jie Li 0001, Leida Li, Xiumei Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Heterogeneous Face Interpretable Disentangled Representation for Joint Face Recognition and SynthesisabstractHeterogeneous faces are acquired with different sensors, which are closer to real-world scenarios and play an important role in the biometric security field. However, heterogeneous face analysis is still a challenging problem due to the large discrepancy between different modalities. Recent works either focus on designing a novel loss function or network architecture to directly extract modality-invariant features or synthesizing the same modality faces initially to decrease the modality gap. Yet, the former always lacks explicit interpretability, and the latter strategy inherently brings in synthesis bias. In this article, we explore to learn the plain interpretable representation for complex heterogeneous faces and simultaneously perform face recognition and synthesis tasks. We propose the heterogeneous face interpretable disentangled representation (HFIDR) that could explicitly interpret dimensions of face representation rather than simple mapping. Benefited from the interpretable structure, we further could extract latent identity information for cross-modality recognition and convert the modality factor to synthesize cross-modality faces. Moreover, we propose a multimodality heterogeneous face interpretable disentangled representation (M-HFIDR) to extend the basic approach suitable for the multimodality face recognition and synthesis. To evaluate the ability of generalization, we construct a novel large-scale face sketch data set. Experimental results on multiple heterogeneous face databases demonstrate the effectiveness of the proposed method. Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Wavelet-Based Dual Recursive Network for Image Super-ResolutionabstractAlthough remarkable progress has been made on single-image super-resolution (SISR), deep learning methods cannot be easily applied to real-world applications due to the requirement of its heavy computation, especially for mobile devices. Focusing on the fewer parameters and faster inference SISR approach, we propose an efficient and time-saving wavelet transform-based network architecture, where the image super-resolution (SR) processing is carried out in the wavelet domain. Different from the existing methods that directly infer high-resolution (HR) image with the input low-resolution (LR) image, our approach first decomposes the LR image into a series of wavelet coefficients (WCs) and the network learns to predict the corresponding series of HR WCs and then reconstructs the HR image. Particularly, in order to further enhance the relationship between WCs and image deep characteristics, we propose two novel modules [wavelet feature mapping block (WFMB) and wavelet coefficients reconstruction block (WCRB)] and a dual recursive framework for joint learning strategy, thus forming a WCs prediction model to realize the efficient and accurate reconstruction of HR WCs. Experimental results show that the proposed method can outperform state-of-the-art methods with more than a 2× reduction in model parameters and computational complexity. Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Heng Huang 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Knowledge Distillation for Face Photo-Sketch SynthesisabstractSignificant progress has been made with face photo-sketch synthesis in recent years due to the development of deep convolutional neural networks, particularly generative adversarial networks (GANs). However, the performance of existing methods is still limited because of the lack of training data (photo-sketch pairs). To address this challenge, we investigate the effect of knowledge distillation (KD) on training neural networks for the face photo-sketch synthesis task and propose an effective KD model to improve the performance of synthetic images. In particular, we utilize a teacher network trained on a large amount of data in a related task to separately learn knowledge of the face photo and knowledge of the face sketch and simultaneously transfer this knowledge to two student networks designed for the face photo-sketch synthesis task. In addition to assimilating the knowledge from the teacher network, the two student networks can mutually transfer their own knowledge to further enhance their learning. To further enhance the perception quality of the synthetic image, we propose a KD+ model that combines GANs with KD. The generator can produce images with more realistic textures and less noise under the guide of knowledge. Extensive experiments and a user study demonstrate the superiority of our models over the state-of-the-art methods. Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning the Non-Differentiable Optimization for Blind Super-ResolutionabstractPrevious convolutional neural network (CNN) based blind super-resolution (SR) methods usually adopt an iterative optimization way to approximate the ground-truth (GT) step-by-step. This solution always involves more computational costs to bring about time-consuming inference. At present, most blind SR algorithms are dedicated to obtaining high-fidelity results; their loss function generally employs L1 loss. To further improve the visual quality of SR results, perceptual metric, such as NIQE, is necessary to guide the network optimization. However, due to the non-differentiable property of NIQE, it cannot be as the loss function. Towards these issues, we propose an adaptive modulation network (AMNet) for multiple degradations SR, which is composed of the pivotal adaptive modulation layer (AMLayer). It is an efficient yet lightweight fusion layer between blur kernel and image features. Equipped with the blur kernel predictor, we naturally upgrade the AMNet to the blind SR model. Instead of considering iterative strategy, we make the blur kernel predictor trainable in the whole blind SR model, in which AMNet is well-trained. Also, we fit deep reinforcement learning into the blind SR model (AMNet-RL) to tackle the non-differentiable optimization problem. Specifically, the blur kernel predictor will be the actor to estimate the blur kernel from the input low-resolution (LR) image. The reward is designed by the pre-defined differentiable or non-differentiable metric. Extensive experiments show that our model can outperform state-of-the-art methods in both fidelity and perceptual metrics. Zheng Hui, Jie Li 0001, Xiumei Wang 0002, Xinbo Gao 0001 |
CVPR | 2 |
| 2021 | Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style TransferabstractArtistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize complex style, especially when holistic global and local patterns exist. Inspired by the common painting process of drawing a draft and revising the details, we introduce a novel feed-forward method named Laplacian Pyramid Network (LapStyle). LapStyle first transfers global style patterns in low-resolution via a Drafting Network. It then revises the local details in high-resolution via a Revision Network, which hallucinates a residual image according to the draft and the image textures extracted by Laplacian filtering. Higher resolution details can be easily generated by stacking Revision Networks with multiple Laplacian pyramid levels. The final stylized image is obtained by aggregating outputs of all pyramid levels. Experiments demonstrate that our method can synthesize high quality stylized images in real time, where holistic style patterns are properly transferred. Zhuoqi Ma, Fu Li 0003, Dongliang He, Xin Li 0106, Errui Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
CVPR | 8 |
| 2021 | A Circular-Structured Representation for Visual Emotion Distribution LearningabstractVisual Emotion Analysis (VEA) has attracted increasing attention recently with the prevalence of sharing images on social networks. Since human emotions are ambiguous and subjective, it is more reasonable to address VEA in a label distribution learning (LDL) paradigm rather than a single-label classification task. Different from other LDL tasks, there exist intrinsic relationships between emotions and unique characteristics within them, as demonstrated in psychological theories. Inspired by this, we propose a well-grounded circular-structured representation to utilize the prior knowledge for visual emotion distribution learning. To be specific, we first construct an Emotion Circle to unify any emotional state within it. On the proposed Emotion Circle, each emotion distribution is represented with an emotion vector, which is defined with three attributes (i.e., emotion polarity, emotion type, emotion intensity) as well as two properties (i.e., similarity, additivity). Besides, we design a novel Progressive Circular (PC) loss to penalize the dissimilarities between predicted emotion vector and labeled one in a coarse-to-fine manner, which further boosts the learning process in an emotion-specific way. Extensive experiments and comparisons are conducted on public visual emotion distribution datasets, and the results demonstrate that the proposed method outperforms the state-of-the-art methods. Jingyuan Yang 0002, Jie Li 0001, Leida Li, Xiumei Wang 0002, Xinbo Gao 0001 |
CVPR | 2 |
| 2021 | Captioning Transformer With Scene Graph GuidingabstractImage captioning is a challenging task which aims to generate descriptions of images. Most existing approaches adopt the encoder-decoder architecture, where encoder takes the image as input and decoder predicts corresponding word sequence. However, a common defect of these methods is that the abundant semantic relationships between relevant regions are ignored, leading the decoder to give a misled caption. To alleviate this issue, we propose a novel model, which utilizes sufficient semantic relationships provided by scene graph to guide the word generation process. To some extent, the scene graph narrows the semantic gap between images and descriptions, and hence improves the quality of generated sentences. Extensive experimental results demonstrate that our model achieves superior performance on various quantitative metrics. HaiShun Chen, Ying Wang 0007, Jie Li 0001 |
ICIP | 4 |
| 2021 | Remote Sensing Imagery Scene Classification Based on Spiking Neural NetworkabstractIn order to overcome the large computational cost of deep neural networks (DNNs), spiking neural networks (SNNs) have been proposed, which are more biologically reasonable. It has the potential to achieve energy efficiency while maintaining performance comparable to DNNs. Although SNNs have achieved good results on the MNIST and CIFAR10 data sets, its potential in remote sensing has not been studied and explored. This paper adopts the idea of converting a trained DNN into an SNN, and proposes a multi-bit-based SNN, and introduces channel normalization (channel-norm) to replace the previous layer normalization to achieve remote sensing imagery scene classification tasks. By achieving multi-bit spiking and channel-norm in the way of DNN conversion to SNN, the SNN proposed in this paper achieves lossless conversion on the UC Merced data set and WHU-RS data set. Saifei Wu, Jie Li 0001, Lin Qi 0004, Xinbo Gao 0001 |
IGARSS | 2 |
| 2021 | CRNet: Centroid Radiation Network for Temporal Action Localization
Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
PRCV (1) | 3 |
| 2021 | Weakly Supervised Temporal Action Localization with Segment-Level Labels
Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
PRCV (1) | 3 |
| 2021 | Learning Deep Patch representation for Probabilistic Graphical Model-Based Face Sketch Synthesis
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Deep blind image quality assessment based on multiple instance regression
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001 |
Neurocomputing | 4 |
| 2021 | Video quality assessment with dense features and ranking pooling
Yu Zhang 0062, Lihuo He, Wen Lu 0004, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2021 | Quality-driven deep active learning method for 3D brain MRI segmentation
Jie Li 0001, Chunna Tian, Zhusi Zhong, Zhicheng Jiao, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2021 | Progressive perception-oriented network for single image super-resolution
Zheng Hui, Jie Li 0001, Xinbo Gao 0001, Xiumei Wang 0002 |
Inf. Sci. | 2 |
| 2021 | Visual relationship detection with region topology structure
Le Zhang 0001, Ying Wang 0007, HaiShun Chen, Jie Li 0001 |
Inf. Sci. | 4 |
| 2021 | Evaluation and comparison of accurate automated spinal curvature estimation algorithms with spinal anterior-posterior X-Ray images: The AASCE2019 challenge
Liansheng Wang 0002, Kailin Chen, Dalong Cheng, Florian Dubost, Benjamin Collery, Bidur Khanal, Bishesh Khanal, Rong Tao, Shangliang Xu, Upasana Upadhyay Bharadwaj, Zhusi Zhong, Jie Li 0001, Shuo Li 0001 |
Medical Image Anal. | 15 |
| 2021 | Iterative local re-ranking with attribute guided synthesis for face sketch recognition
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Chunlei Peng, Jie Li 0001 |
Pattern Recognit. | 5 |
| 2021 | Single image super-resolution with multi-scale information cross-fusion network
Yanting Hu, Xinbo Gao 0001, Jie Li 0001, Yuanfei Huang, Hanzi Wang |
Signal Process. | 3 |
| 2021 | Patch-based co-occurrence filter with fast adaptive kernel
Zhonggui Sun, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001 |
Signal Process. | 3 |
| 2021 | Learning stacking regression for no-reference super-resolution image quality assessment
Kaibing Zhang, Danni Zhu, Jie Li 0001, Xinbo Gao 0001, Fei Gao 0006 |
Signal Process. | 3 |
| 2021 | Attention Multibranch Convolutional Neural Network for Hyperspectral Image Classification Based on Adaptive Region SearchabstractConvolutional neural networks (CNNs) have demonstrated outstanding performance on image classification. To classify the hyperspectral images (HSIs), existing CNN-based approaches commonly adopt the architecture using single or several fixed spatial windows as inputs. This kind of architecture may lose contextual information or incorporate heterogeneous information due to the neglect of various land-cover distributions in HSIs. To deal with this problem, a novel attention multibranch CNN method based on adaptive region search (RS-AMCNN) is proposed for HSI classification. In RS-AMCNN, sizes and locations of spatial windows are searched in the nonlocal candidate region adaptively according to sample-specific distribution. These flexible spatial windows are input into several branches of RS-AMCNN. In each branch, convolutional long short-term memories (ConvLSTMs) are merged into CNN from shallow to deep layers, which not only extracts joint spatial-spectral features, but also exploits complementary information among different layers. Then, a branch attention mechanism is devised to emphasize more discriminative branches and suppress less useful ones. It forces RS-AMCNN to extract multiscale and multicontextual attention features for classification. Finally, RS-AMCNN is optimized end-to-end by combining the losses from the ramose classifiers of different branches and the main classifier. Experiments carried on several benchmark HSI data sets demonstrate that RS-AMCNN provides promising classification performance, especially in edge preservation and region uniformity. Jie Feng 0003, Xiande Wu, Ronghua Shang, Chenhong Sui, Jie Li 0001, Licheng Jiao, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Adaptive Edge Preserving Maps in Markov Random Fields for Hyperspectral Image ClassificationabstractThis article presents a novel adaptive edge preserving (aEP) scheme in Markov random fields (MRFs) for hyperspectral image (HSI) classification. MRF regularization usually suffered from over-smoothing at boundaries and insufficient refinement within class objects. This work divides and conquers this problem class-by-class, and integrates${K}$(${K} -1$)/2 (${K}$is the class number) aEP maps (aEPMs) in MRF model. Spatial label dependence measure (SLDM) is designed to estimate the interpixel label dependence for given spectral similarity measure. For each class pair, aEPM is optimized by maximizing the difference between intraclass and interclass SLDM. Then, aEPMs are integrated with multilevel logistic (MLL) model to regularize the raw pixelwise labeling obtained by spectral and spectral–spatial methods, respectively. The graph-cuts-based$\alpha ~\beta $-swap algorithm is modified to optimize the designed energy function. Moreover, to evaluate the final refined results at edges and small details thoroughly, segmentation evaluation metrics are introduced. Experiments conducted on real HSI data denote the superiority of aEPMs in evaluation metrics and region consistency, especially in detail preservation. Chao Pan 0006, Xiuping Jia, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Soft Semantic Representation for Cross-Domain Face RecognitionabstractThe problem of cross-domain face recognition aims to identify facial images obtained across different domains, which attracts increasing attentions because of its wide applications on law-enforcement identification and camera surveillance. The problem is challenging due to the huge domain discrepancy. Despite great progress achieved in recent years, existing algorithms usually fail to fully exploit the semantic information for identifying cross-domain faces, which could be a strong clue for recognition. In this article, we propose an effective algorithm for cross-domain face recognition by exploiting semantic information integrated with deep convolutional neural networks (CNN). We first introduce a soft face parsing algorithm where the boundaries of facial components are measured as probabilistic values. By taking the original face image as the guidance to improve face parsing result, each pixel may belong partially to the facial component to avoid inaccurate segmentation around component boundaries. We then propose a hierarchical soft semantic representation framework for cross-domain face recognition. Both the soft semantic level and contour level deep features obtained via CNN are computed and combined together, which could fully exploit the identical semantic clue among cross-domain faces. We provide extensive experiments to demonstrate that the proposed soft semantic representation algorithm performs superior against state-of-the-art methods. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | KFC: An Efficient Framework for Semi-Supervised Temporal Action LocalizationabstractIn temporal action localization (TAL), semi-supervised learning is a promising technique to mitigate the cost of precise boundary annotations. Semi-supervised approaches employing consistency regularization (CR), encouraging models to be robust to the perturbed inputs, have achieved great success in image classification problems. The success of CR is largely depended on the perturbations, where instances are perturbed to train a robust model without altering their semantic information. However, the perturbations for image or video classification tasks are not fit to apply to TAL. Since videos in TAL are too long to train the model with raw videos in an end-to-end manner. In this paper, we devise a method named K-farthest crossover to construct perturbations based on video features and apply it to TAL. Motivated by the observation that features in the same action instance become more and more similar during the training process while those in different action instances or backgrounds become more and more divergent, we add perturbations to each feature along temporal axis and adopt CR to encourage the model to retain this observation. Specifically, for a feature, we first find the top-k dissimilar features and average them to form a perturbation. Then, similar to chromosomal crossover, we select a large part of the feature and a small part of the perturbation to recombine a perturbed feature, which preserves the feature semantics yet enough discrepancy. Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu |
IEEE Trans. Image Process. | 4 |
| 2021 | Interpretable Detail-Fidelity Attention Network for Single Image Super-ResolutionabstractBenefiting from the strong capabilities of deep CNNs for feature representation and nonlinear mapping, deep-learning-based methods have achieved excellent performance in single image super-resolution. However, most existing SR methods depend on the high capacity of networks that are initially designed for visual recognition, and rarely consider the initial intention of super-resolution for detail fidelity. To pursue this intention, there are two challenging issues that must be solved: (1) learning appropriate operators which is adaptive to the diverse characteristics of smoothes and details; (2) improving the ability of the model to preserve low-frequency smoothes and reconstruct high-frequency details. To solve these problems, we propose a purposeful and interpretable detail-fidelity attention network to progressively process these smoothes and details in a divide-and-conquer manner, which is a novel and specific prospect of image super-resolution for the purpose of improving detail fidelity. This proposed method updates the concept of blindly designing or using deep CNNs architectures for only feature representation in local receptive fields. In particular, we propose a Hessian filtering for interpretable high-profile feature representation for detail inference, along with a dilated encoder-decoder and a distribution alignment cell to improve the inferred Hessian features in a morphological manner and statistical manner respectively. Extensive experiments demonstrate that the proposed method achieves superior performance compared to the state-of-the-art methods both quantitatively and qualitatively. The code is available at github.com/YuanfeiHuang/DeFiAN. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Yanting Hu, Wen Lu 0004 |
IEEE Trans. Image Process. | 2 |
| 2021 | Multi-Hierarchical Category Supervision for Weakly-Supervised Temporal Action LocalizationabstractWeakly Supervised Temporal Action Localization (WTAL) aims to localize action segments in untrimmed videos with only video-level category labels in the training phase. In WTAL, an action generally consists of a series of sub-actions, and different categories of actions may share the common sub-actions. However, to distinguish different categories of actions with only video-level class labels, current WTAL models tend to focus on discriminative sub-actions of the action, while ignoring those common sub-actions shared with different categories of actions. This negligence of common sub-actions would lead to the located action segments incomplete, i.e., only containing discriminative sub-actions. Different from current approaches of designing complex network architectures to explore more complete actions, in this paper, we introduce a novel supervision method named multi-hierarchical category supervision (MHCS) to find more sub-actions rather than only the discriminative ones. Specifically, action categories sharing similar sub-actions will be constructed as super-classes through hierarchical clustering. Hence, training with the new generated super-classes would encourage the model to pay more attention to the common sub-actions, which are ignored training with the original classes. Furthermore, our proposed MHCS is model-agnostic and non-intrusive, which can be directly applied to existing methods without changing their structures. Through extensive experiments, we verify that our supervision method can improve the performance of four state-of-the-art WTAL methods on three public datasets: THUMOS14, ActivityNet1.2, and ActivityNet1.3. Guozhang Li, Jie Li 0001, Nannan Wang 0001, Xinpeng Ding, Zhifeng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Stimuli-Aware Visual Emotion AnalysisabstractVisual emotion analysis (VEA) has attracted great attention recently, due to the increasing tendency of expressing and understanding emotions through images on social networks. Different from traditional vision tasks, VEA is inherently more challenging since it involves a much higher level of complexity and ambiguity in human cognitive process. Most of the existing methods adopt deep learning techniques to extract general features from the whole image, disregarding the specific features evoked by various emotional stimuli. Inspired by the Stimuli-Organism-Response (S-O-R) emotion model in psychological theory, we proposed a stimuli-aware VEA method consisting of three stages, namely stimuli selection (S), feature extraction (O) and emotion prediction (R). First, specific emotional stimuli (i. e., color, object, face) are selected from images by employing the off-the-shelf tools. To the best of our knowledge, it is the first time to introduce stimuli selection process into VEA in an end-to-end network. Then, we design three specific networks, i. e., Global-Net, Semantic-Net and Expression-Net, to extract distinct emotional features from different stimuli simultaneously. Finally, benefiting from the inherent structure of Mikel's wheel, we design a novel hierarchical cross-entropy loss to distinguish hard false examples from easy ones in an emotion-specific manner. Experiments demonstrate that the proposed method consistently outperforms the state-of-the-art approaches on four public visual emotion datasets. Ablation study and visualizations further prove the validity and interpretability of our method. Jingyuan Yang 0002, Jie Li 0001, Xiumei Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Beyond Vision: A Multimodal Recurrent Attention Convolutional Neural Network for Unified Image Aesthetic Prediction TasksabstractOver the past few years, image aesthetic prediction has attracted increasing attention because of its wide applications, such as image retrieval, photo album management and aesthetic-driven image enhancement. However, previous studies in this area only achieve limited success because 1) they primarily depend on visual features and ignore textual information. 2) they tend to focus equally on to each part of images and ignore the selective attention mechanism. This paper overcomes these limitations by proposing a novel multimodal recurrent attention convolutional neural network (MRACNN). More specifically, the MRACNN consists of two streams: the vision stream and the language stream. The former employs the recurrent attention network to tune out irrelevant information and focuses on some key regions to extract visual features. The latter utilizes the Text-CNN to capture the high-level semantics of user comments. Finally, a multimodal factorized bilinear (MFB) pooling approach is used to achieve effective fusion of textual and visual features. Extensive experiments demonstrate that the proposed MRACNN significantly outperforms state-of-the-art methods for unified aesthetic prediction tasks: (i) aesthetic quality classification; (ii) aesthetic score regression; and (iii) aesthetic score distribution prediction. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Jie Li 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Facial Attribute Capsules for Noise Face Super ResolutionabstractExisting face super-resolution (SR) methods mainly assume the input image to be noise-free. Their performance degrades drastically when applied to real-world scenarios where the input image is always contaminated by noise. In this paper, we propose a Facial Attribute Capsules Network (FACN) to deal with the problem of high-scale super-resolution of noisy face image. Capsule is a group of neurons whose activity vector models different properties of the same entity. Inspired by the concept of capsule, we propose an integrated representation model of facial information, which named Facial Attribute Capsule (FAC). In the SR processing, we first generated a group of FACs from the input LR face, and then reconstructed the HR face from this group of FACs. Aiming to effectively improve the robustness of FAC to noise, we generate FAC in semantic, probabilistic and facial attributes manners by means of integrated learning strategy. Each FAC can be divided into two sub-capsules: Semantic Capsule (SC) and Probabilistic Capsule (PC). Them describe an explicit facial attribute in detail from two aspects of semantic representation and probability distribution. The group of FACs model an image as a combination of facial attribute information in the semantic space and probabilistic space by an attribute-disentangling way. The diverse FACs could better combine the face prior information to generate the face images with fine-grained semantic attributes. Extensive benchmark experiments show that our method achieves superior hallucination results and outperforms state-of-the-art for very low resolution (LR) noise face image super resolution. Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001 |
AAAI | 4 |
| 2020 | Video Face Super-Resolution with Motion-Adaptive Feedback CellabstractVideo super-resolution (VSR) methods have recently achieved a remarkable success due to the development of deep convolutional neural networks (CNN). Current state-of-the-art CNN methods usually treat the VSR problem as a large number of separate multi-frame super-resolution tasks, at which a batch of low resolution (LR) frames is utilized to generate a single high resolution (HR) frame, and running a slide window to select LR frames over the entire video would obtain a series of HR frames. However, duo to the complex temporal dependency between frames, with the number of LR input frames increase, the performance of the reconstructed HR frames become worse. The reason is in that these methods lack the ability to model complex temporal dependencies and hard to give an accurate motion estimation and compensation for VSR process. Which makes the performance degrade drastically when the motion in frames is complex. In this paper, we propose a Motion-Adaptive Feedback Cell (MAFC), a simple but effective block, which can efficiently capture the motion compensation and feed it back to the network in an adaptive way. Our approach efficiently utilizes the information of the inter-frame motion, the dependence of the network on motion estimation and compensation method can be avoid. In addition, benefiting from the excellent nature of MAFC, the network can achieve better performance in the case of extremely complex motion scenarios. Extensive evaluations and comparisons validate the strengths of our approach, and the experimental results demonstrated that the proposed framework is outperform the state-of-the-art methods. Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001 |
AAAI | 3 |
| 2020 | Binarized Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Heng Huang 0001, Xinbo Gao 0001 |
ECCV (4) | 4 |
| 2020 | Spatial-Spectral Autoencoder Networks for Hyperspectral UnmixingabstractWe present a spatial-spectral autoencoder (SSAE) for hyperspectral unmixing, including a net for endmember extraction (EENet) and a net for abundance estimation (AENet). The EENet exploits the spatial information in hyperspectral image by a “many to one” strategy, i.e., the abundance of a pixel is combined by the abundances of its adjacent pixels. The idea is based on the assumption: once an endmember is mixed in a pixel, it is mixed in the surrounding pixels with high probability. The strategy promotes a continuous and smooth spatial distribution of abundances, and it is more effective than the other methods for endmember extraction. Besides, to make full use of the rich spectral information and obtain more accurate abundances, we design an AENet, which applies the deep convolutional neural network to estimate the abundances with the endmembers acquired from the EENet. The experiments are conducted on two real datasets, which show the SSAE outperforms the state-of-the-art methods. Yongfa Huang, Jie Li 0001, Lin Qi 0004, Ying Wang 0007, Xinbo Gao 0001 |
IGARSS | 2 |
| 2020 | Hyperspectral Unmixing via Recurrent Neural Network With Chain ClassifierabstractRecently, the development of deep learning brings new opportunities for hyperspectral unmixing. In this paper, we propose a new architecture based on Recurrent Neural Networks and Bidirectional-LSTM (BiLSTM) that consists of two blocks: the feature extraction stage based on BiLSTM and the abundance estimation stage via Chain Classifier. BiLSTM can capture the long-distance relation among spectral bands better and Chain Classifier is more suitable to deal with multi-label task that exists in abundance estimation process. We evaluate the proposed method on two real HSI data sets including Jasper Ridge and Urban compared with three state-of-the-art approaches, and the results show significant improvements in accuracy of unmixing. Mingyu Lei, Jie Li 0001, Lin Qi 0004, Ying Wang 0007, Xinbo Gao 0001 |
IGARSS | 2 |
| 2020 | Image Super-Resolution via Deep Feature Recalibration Network
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
PRCV (1) | 4 |
| 2020 | Semantic-related image style transfer with dual-consistency loss
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2020 | CBFNet: Constraint balance factor for semantic segmentation
Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2020 | Image style transfer with collection representation space and semantic-guided reconstruction
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Neural Networks | 2 |
| 2020 | Modality adversarial neural network for visible-thermal person re-identification
Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2020 | Deep spectral convolution network for hyperspectral image unmixing with spectral library
Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Mingyu Lei, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2020 | Channel-Wise and Spatial Feature Modulation Network for Single Image Super-ResolutionabstractThe performance of single image super-resolution has achieved significant improvement by utilizing deep convolutional neural networks (CNNs). The features in deep CNN contain different types of information which make different contributions to image reconstruction. However, the most CNN-based models lack discriminative ability for different types of information and deal with them equally, which results in the representational capacity of the models being limited. On the other hand, as the depth of neural network grows, the long-term information coming from preceding layers is easy to be weaken or lost at later layers, which is adverse to super-resolving image. To capture more informative features and maintain long-term information for image super-resolution, we propose a channel-wise and spatial feature modulation (CSFM) network in which a series of feature modulation memory (FMM) modules are cascaded with a densely connected structure to transform shallow features to high informative features. In each FMM module, we construct a set of channel-wise and spatial attention residual (CSAR) blocks and stack them in a chain structure to dynamically modulate the multi-level features in global and local manners. This feature modulation strategy enables the valuable information to be enhanced and the redundant information to be suppressed. Meanwhile, for long-term information persistence, a gated fusion (GF) node is attached at the end of the FMM module to adaptively fuse hierarchical features and distill more effective information via the dense skip connections and the gating mechanism. The extensive quantitative and qualitative evaluations on benchmark datasets illustrate the superiority of our proposed method over the state-of-the-art methods. Yanting Hu, Jie Li 0001, Yuanfei Huang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Spectral-Spatial-Weighted Multiview Collaborative Sparse Unmixing for Hyperspectral ImagesabstractSpectral unmixing is an important task in hyperspectral image (HSI) analysis and processing. Sparse representation has become a promising semisupervised method for remotely sensed hyperspectral unmixing and incorporating the spectral or spatial information to improve the spectral unmixing results under a weighted sparse unmixing framework is a recent trend. While most methods focus on analyzing HSI by exploring the spatial information, it is known that hyperspectral data are characterized by its large contiguous set of wavelengths. This information can be naturally used to improve the representation of pixels in HSI. In order to take the advantage of the hyper spectral information as well as the spatial information for hyperspectral unmixing, in this article, we explore and introduce a multiview data processing approach through spectral partitioning to benefit from the abundant spectral information in HSI. Some important findings on the application of multiview data set in sparse unmixing are discussed. Meanwhile, we develop a new spectral–spatial-weighted multiview collaborative sparse unmixing (MCSU) model to tackle such a multiview data set. The MCSU uses a weighted sparse regularizer, which includes both multiview spectral and spatial weighting factors to further impose sparsity on the fractional abundances. The weights are adaptively updated associated with the abundances, and the proposed MCSU can be solved by the alternating direction method of multipliers efficiently. The experimental results on both the simulated and real hyperspectral data sets demonstrate the effectiveness of the proposed MCSU, which can significantly improve the abundance estimation results. Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Yongfa Huang, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Face Sketch Synthesis in the Wild via Deep Patch Representation-Based Probabilistic Graphical ModelabstractThis paper considers the problem of face sketch synthesis in the wild, which transforms a face photo into a face sketch. Face sketch synthesis is widely applied in law enforcement as well as digital entertainment fields. However, the existing methods either focus on hand-crafted techniques where prior human experience is relied on or adopt deep learning techniques as an end-to-end framework, where facial details cannot be well represented. In this paper, we propose a novel approach for face sketch synthesis in the wild via a deep patch representation-based probabilistic graphical model (DeepPGM). A Siamese network is constructed to extract deep patch representation from a raw facial patch, where the representative detail information for robust face sketch synthesis can be exploited. The generated deep patch representation and facial image patches are then optimally combined through a probabilistic graphical model. The proposed DeepPGM approach not only outperforms the state-of-the-art on public face sketch datasets but also can cope with forensic photos in the wild conditions, including varying lightings, poses, occlusions, skin colors, and ethnic origins. The superiority of the proposed method is demonstrated by extensive experiments on two public face sketch datasets and real-world forensic photos in the wild. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Group Feedback Capsule NetworkabstractIn capsule networks (CapsNets), the capsule is made up of collections of neurons. Their adjacent capsule layers are connected using routing-by-agreement mechanisms in an unsupervised way. The routing-by-agreement mechanisms have two main drawbacks: a) too many parameters and high computation complexity; b) the cluster distribution assumptions of these routing mechanisms may not hold in some complex real-world data. In this paper, we propose a novel Group Feedback Capsule Network (GF-CapsNet) which adopts a supervised routing strategy called group-routing. Compared with the previous routing strategies which globally transform each capsule, Group-routing equally splits capsules into groups where capsules locally share the same transformation weights, reducing routing parameters. To address the second drawback, we devise a distance network to directly predict capsules in a supervised way without making distribution assumptions. Our proposed group-routing captures local information of low-level capsules by group-wise transformation and supervisedly predicts high-level ones in a feedback way to address two drawbacks respectively. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts. Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu |
IEEE Trans. Image Process. | 4 |
| 2020 | Universal Face Photo-Sketch Style Transfer via Multiview Domain TranslationabstractFace photo-sketch style transfer aims to convert a representation of a face from the photo (or sketch) domain to the sketch (respectively, photo) domain while preserving the character of the subject. It has wide-ranging applications in law enforcement, forensic investigation and digital entertainment. However, conventional face photo-sketch synthesis methods usually require training images from both the source domain and the target domain, and are limited in that they cannot be applied to universal conditions where collecting training images in the source domain that match the style of the test image is unpractical. This problem entails two major challenges: 1) designing an effective and robust domain translation model for the universal situation in which images of the source domain needed for training are unavailable, and 2) preserving the facial character while performing a transfer to the style of an entire image collection in the target domain. To this end, we present a novel universal face photo-sketch style transfer method that does not need any image from the source domain for training. The regression relationship between an input test image and the entire training image collection in the target domain is inferred via a deep domain translation framework, in which a domain-wise adaption term and a local consistency adaption term are developed. To improve the robustness of the style transfer process, we propose a multiview domain translation method that flexibly leverages a convolutional neural network representation with hand-crafted features in an optimal way. Qualitative and quantitative comparisons are provided for universal unconstrained conditions of unavailable training images from the source domain, demonstrating the effectiveness and superiority of our method for universal face photo-sketch style transfer. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Weighted Guided Image Filtering With Steering KernelabstractDue to its local property, guided image filter (GIF) generally suffers from halo artifacts near edges. To make up for the deficiency, a weighted guided image filter (WGIF) was proposed recently by incorporating an edge-aware weighting into the filtering process. It takes the advantages of local and global operations, and achieves better performance in edge-preserving. However, edge direction, a vital property of the guidance image, is not considered fully in these guided filters. In order to overcome the drawback, we propose a novel version of GIF, which can leverage the edge direction more sufficiently. In particular, we utilize the steering kernel to adaptively learn the direction and incorporate the learning results into the filtering process to improve the filter's behavior. Theoretical analysis shows that the proposed method can get more powerful performance with preserving edges and reducing halo artifacts effectively. Similar conclusions are also reached through the thorough experiments including edge-aware smoothing, detail enhancement, denoising and dehazing. Zhonggui Sun, Bo Han 0004, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Progressive Sub-Band Residual-Learning Network for MR Image Super ResolutionabstractHigh-resolution (HR) magnetic resonance images (MRI) provide more detailed information for clinical application. However, HR MRI is less available because of the longer scan time and lower signal-to-noise ratio. Spatial resolution is one of the key parameters of MRI. The image post-processing technique super-resolution (SR) is an alternative approach to improve the spatial resolution of MR images. Inspired by advanced deep learning based SR methods, we propose an MRI SR model named progressive sub-band residual learning SR network (PSR-SRN). The proposed model contains two parallel progressive learning streams, where one stream learns on missed high-frequency residuals by sub-band residual learning unit (ISRL) and the other focuses on reconstructing refined MR image. These two streams complement each other and enable to learn complex mappings between "Low-" and "High-" resolution MR images. Besides, we introduce brain-like mechanisms (in-depth supervision and local feedback mechanism) and progressive sub-band learning strategy to emphasize variant textures of MRI. Compared with traditional and deep learning MRI SR methods, our PSR-SRN model shows superior performance. Xuetong Xue, Ying Wang 0007, Jie Li 0001, Zhicheng Jiao, Ziqi Ren, Xinbo Gao 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Coupled Attribute Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is a challenging problem in face recognition and subject to large textural and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. In addition, limited training samples of cross-modality pairs make HFR more challenging due to the complex generation procedure of these images. Despite the great progress that has been achieved in recent years, existing works mainly focus on HFR from only cross-modality image matching. However, it is more practical to obtain both facial images and semantic descriptions about facial attributes in real-world situations, in which the semantic description clues are nearly always obtained during the process of image generation. Motivated by human cognitive mechanisms, we naturally utilize the explicit invariant semantic description, i.e., face attributes, to help address the gap among face images of different modalities. Existing facial attributes-related face recognition methods primarily regard attributes as the high-level features used to enhance recognition performance, ignoring the inherent relationship between face attributes and identities. In this article, we propose novel coupled attribute learning for the HFR (CAL-HFR) method without labeling the attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances are taken as dissimilarities of pairs. Coupled attribute guided triplet loss (CAGTL) is designed to train an end-to-end HFR network that can effectively eliminate defects of incorrectly estimated attributes. Extensive experiments on multiple heterogeneous scenarios demonstrate that the proposed method achieves superior performance compared with that of state-of-the-art methods. Furthermore, we make publicly available our generated pairwise annotated heterogeneous facial attribute database for evaluation and promoting related research. Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Chunlei Peng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | HSME: Hypersphere Manifold Embedding for Visible Thermal Person Re-IdentificationabstractPerson Re-identification(re-ID) has great potential to contribute to video surveillance that automatically searches and identifies people across different cameras. Heterogeneous person re-identification between thermal(infrared) and visible images is essentially a cross-modality problem and important for night-time surveillance application. Current methods usually train a model by combining classification and metric learning algorithms to obtain discriminative and robust feature representations. However, the combined loss function ignored the correlation between classification subspace and feature embedding subspace. In this paper, we use Sphere Softmax to learn a hypersphere manifold embedding and constrain the intra-modality variations and cross-modality variations on this hypersphere. We propose an end-to-end dualstream hypersphere manifold embedding network(HSMEnet) with both classification and identification constraint. Meanwhile, we design a two-stage training scheme to acquire decorrelated features, we refer the HSME with decorrelation as D-HSME. We conduct experiments on two crossmodality person re-identification datasets. Experimental results demonstrate that our method outperforms the state-of-the-art methods on two datasets. On RegDB dataset, rank-1 accuracy is improved from 33.47% to 50.85%, and mAP is improved from 31.83% to 47.00%. Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
AAAI | 3 |
| 2019 | Residual Attribute Attention Network for Face Image Super-ResolutionabstractFacial prior knowledge based methods recently achieved great success on the task of face image super-resolution (SR). The combination of different type of facial knowledge could be leveraged for better super-resolving face images, e.g., facial attribute information with texture and shape information. In this paper, we present a novel deep end-to-end network for face super resolution, named Residual Attribute Attention Network (RAAN), which realizes the efficient feature fusion of various types of facial information. Specifically, we construct a multi-block cascaded structure network with dense connection. Each block has three branches: Texture Prediction Network (TPN), Shape Generation Network (SGN) and Attribute Analysis Network (AAN). We divide the task of face image reconstruction into three steps: extracting the pixel level representation information from the input very low resolution (LR) image via TPN and SGN, extracting the semantic level representation information by AAN from the input, and finally combining the pixel level and semantic level information to recover the high resolution (HR) image. Experiments on benchmark database illustrate that RAAN significantly outperforms state-of-the-arts for very low-resolution face SR problem, both quantitatively and qualitatively. Jingwei Xin, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
AAAI | 4 |
| 2019 | Improving Image Super-Resolution via Feature Re-Balancing FusionabstractRecently, benefiting from the strong ability of feature representation, deep-learning-based methods have achieved excellent performance in single image super-resolution (SR). Furthermore, skip connection and feature fusion have been demonstrated to be a commendable strategy to deal with various features in different depth for informative reconstruction. Nevertheless, cross-layer features show diverse characteristic in detail representation, blindly fusion then introduces unavoidable interference of features. In this paper, we propose a novel feature fusion unit by utilizing alternative dilated convolutions for re-balancing diverse cross-layer features, named Feature Re-Balancing Fusion Network (RBFNet), which is theoretically and experimentally demonstrated to be robust to the interference in feature fusion for SR. Extensive experiments show that the proposed method achieves excellent performances quantitatively and qualitatively against the state-of-the-art methods. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Wen Lu 0004, Yanting Hu |
ICME | 2 |
| 2019 | Non-Local Compressive Network for Hyperspectral and Multispectral Data FusionabstractHyperspectral imagery enhancing is very important to remote sensing interpretation. Recently, fusing hyperspectral data with its corresponding high spatial resolution multispectral data has been an important technology to obtain high-resolution imagery. Deep neural networks are widely used because of their significant performance, while they are still suffered from the problem of been insensitive to non-local information and the problem of overfitting. In this paper, a novel fusion method based on the non-local compressive network is proposed. This network can extract non-local information and also reduce overfitting when dealing with fusion tasks. Because of the introduction of non-local structure, the global relation information is also effectively enhanced according to the self-attention mechanism of features. Experiments on two datasets have shown that the proposed method can obtain better performance than the state-of-art methods. Junbo Hao, Ying Wang 0007, Jie Li 0001, Xinbo Gao 0001 |
IGARSS | 3 |
| 2019 | Multi-Margin based Decorrelation Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching face images acquired from different domains with wide applications in security scenarios. However, HFR is still a challenging problem due to the significant cross-domain discrepancy and the lacking of sufficient training data in different domains. This paper presents a deep neural network approach namely Multi-Margin based Decorrelation Learning (MMDL) to extract decorrelation representations in a hyperspherical space for cross-domain face images. The proposed framework can be divided into two components: heterogeneous representation network and decorrelation representation learning. First, we employ a large scale of accessible visual face images to train heterogeneous representation network. The decorrelation layer projects the output of the first component into decorrelation latent subspace and obtain decorrelation representation. In addition, we design a multi-margin loss (MML), which consists of tetradmargin loss (TML) and heterogeneous angular margin loss (HAML), to constrain the proposed framework. Experimental results on two challenging heterogeneous face databases show that our approach achieves superior performance on both verification and recognition tasks, comparing with state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001 |
IJCAI | 4 |
| 2019 | Group Reconstruction and Max-Pooling Residual Capsule NetworkabstractIn capsule networks, the mapping of low-level capsules to high-level capsules is achieved by a routing-by-agreement algorithm. Since the capsule is made up of collections of neurons and the routing mechanism involves all the capsules instead of simply discarding some of the neurons like Max-Pooling, the capsule network has stronger representation ability than the traditional neural network. However, considering too much low-level capsules' information will cause its corresponding upper layer capsules to be interfered by other irrelevant information or noise capsules. Therefore, the original capsule network does not perform well on complex data structure. What's worse, computational complexity becomes a bottleneck in dealing with large data networks. In order to solve these shortcomings, this paper proposes a group reconstruction and max-pooling residual capsule network (GRMR-CapsNet). We build a block in which all capsules are divided into different groups and perform group reconstruction routing algorithm to obtain the corresponding high-level capsules. Between the lower and higher layers, Capsule Max-Pooling is adopted to prevent overfitting. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts. Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002 |
IJCAI | 4 |
| 2019 | Face Photo-Sketch Synthesis via Knowledge TransferabstractDespite deep neural networks have demonstrated strong power in face photo-sketch synthesis task, their performance, however, are still limited by the lack of training data (photo-sketch pairs). Knowledge Transfer (KT), which aims at training a smaller and fast student network with the information learned from a larger and accurate teacher network, has attracted much attention recently due to its superior performance in the acceleration and compression of deep neural networks. This work has brought us great inspiration that we can train a relatively small student network on very few training data by transferring knowledge from a larger teacher model trained on enough training data for other tasks. Therefore, we propose a novel knowledge transfer framework to synthesize face photos from face sketches or synthesize face sketches from face photos. Particularly, we utilize two teacher networks trained on large amount of data in related task to learn the knowledge of face photos and face sketches separately and transfer them to two student networks simultaneously. In addition, the two student networks, one for photo ? sketch task and the other for sketch ? photo task, can transfer their knowledge mutually. With the proposed method, we can train our model which has superior performance using a small set of photo-sketch pairs. We validate the effectiveness of our method across several datasets. Quantitative and qualitative evaluations illustrate that our model outperforms other state-of-the-art methods in generating face sketches (or photos) with high visual quality and recognition ability. Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001 |
IJCAI | 4 |
| 2019 | Refined Segmentation R-CNN: A Two-Stage Convolutional Neural Network for Punctate White Matter Lesion Segmentation in Preterm Infants
Yalong Liu, Jie Li 0001, Ying Wang 0007, Xianjun Li, Zhicheng Jiao, Jian Yang 0003, Xinbo Gao 0001 |
MICCAI (3) | 2 |
| 2019 | An Attention-Guided Deep Regression Model for Landmark Detection in Cephalograms
Zhusi Zhong, Jie Li 0001, Zhicheng Jiao, Xinbo Gao 0001 |
MICCAI (6) | 2 |
| 2019 | Dual-alignment Feature Embedding for Cross-modality Person Re-identificationabstractPerson re-identification aims at searching pedestrians across different cameras, which is a key problem in video surveillance. With requirements in night environment, RGB-infrared person re-identification which could be regarded as a cross-modality matching problem, has gained increasing attention in recent years. Aside from cross-modality discrepancy, RGB-infrared person re-identification also suffers from human pose and view point differences. We design a dual-alignment feature embedding method to extract discriminative modality-invariant features. The concept of dual-alignment is two folds: spatial and modality alignments. We adopt the part-level features to extract fine-grained camera-invariant information. We introduce distribution loss function and correlation loss function to align the embedding features across visible and infrared modalities. Finally, we can extract modality-invariant features with robust and rich identity embeddings for cross-modality person re-identification. Experiment confirms that the proposed baseline and improvement achieves competitive results with the state-of-the-art methods on two datasets. For instance, We achieve (57.5+12.6)% rank-1 accuracy and (57.3+11.8)% mAP on the RegDB dataset. Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002 |
ACM Multimedia | 4 |
| 2019 | A novel joint dictionary framework for sparse hyperspectral unmixing incorporating spectral library
Lin Qi 0004, Jie Li 0001, Xinbo Gao 0001, Ying Wang 0007, Chongyue Zhao, Yu Zheng 0006 |
Neurocomputing | 2 |
| 2019 | Region-Based Multiview Sparse Hyperspectral Unmixing Incorporating Spectral LibraryabstractHyperspectral image (HSI) is characterized by its huge contiguous set of wavelengths. It is possible and needed to benefit from the “hyper” spectral information as well as the spatial information. For this purpose, we propose a new multiview data generation approach that takes full advantage of the rich spectral and spatial information in HSI, by dividing the original HSI into several spatially homogeneous regions with different band margins. Then, a new sparse unmixing algorithm, called region-based multiview sparse unmixing (RMSU), is presented to tackle such a multiview data model in this letter. The RMSU algorithm combines the multiview learning anda prioriinformation to improve the performance of sparse unmixing by incorporating the multiview information and spectral library into the dictionary learning framework. We also show that RMSU can serve as a dictionary pruning algorithm, which provides a possibility that unmixing algorithms could have higher accuracy and efficiency. Experimental results on both simulated and real hyperspectral data demonstrate the effectiveness of the proposed RMSU algorithm both visually and quantitatively. Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Structural Reweight Sparse Subspace Clustering
Bing Han 0003, Jie Li 0001, Xinbo Gao 0001 |
Neural Process. Lett. | 3 |
| 2019 | DLFace: Deep local descriptor for cross-modality face recognition
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2019 | Sparse graphical representation based discriminant analysis for heterogeneous face recognition
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
Signal Process. | 4 |
| 2019 | Learning recurrent residual regressors for single image super-resolution
Kaibing Zhang, Zhen Wang 0037, Jie Li 0001, Xinbo Gao 0001, Zenggang Xiong |
Signal Process. | 3 |
| 2019 | Markov Random Fields Integrating Adaptive Interclass-Pair Penalty and Spectral Similarity for Hyperspectral Image ClassificationabstractThis paper presents a novel Markov random field (MRF) method integrating adaptive interclass-pair penalty (aICP2) and spectral similarity information (SSI) for hyper-spectral image (HSI) classification. aICP2structurally combines$K(K - 1)/2$(K is the number of classes) classical “Potts model” with$K(K - 1)/2$interaction coefficients. aICP2tries a new way to solve the key problems, insufficient correction within homogeneous regions, and over-smoothness at class boundaries, in MRF-based HSI classification. It is assumed that different class pairs should be assigned with various degrees of penalties in MRF smoothness process, according to pairwise class separability and spatial class confusion in raw classification map. The Fisher ratio is modified to measure pairwise class separability with a training set. And, gray level co-occurrence matrix is used to measure spatial class confusion degree. Then, aICP2is constructed by combining Fisher ratio and GCLM. aICP2applies larger penalty on class pairs that confuse with each other seriously to provide sufficient smoothness, and vice versa. In addition, to protect class edges and details, SSI is introduced to make the penalty of related neighboring pixels small. aICP2ssi denotes the integration of aICP2and SSI. The further improved method is both interclass-pair and interpixel adaptive in spatial term. A graph-cut-based$\alpha - \beta $-swap method is introduced to optimize the proposed energy function. The experimental results on real HSI data indicate that the proposed method outperforms compared MRF-based and other spectral–spatial approaches in terms of classification accuracies and region uniformity. Chao Pan 0006, Xinbo Gao 0001, Ying Wang 0007, Jie Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Unsupervised Semantic-Preserving Adversarial Hashing for Image SearchabstractHashing plays a pivotal role in nearest-neighbor searching for large-scale image retrieval. Recently, deep learning-based hashing methods have achieved promising performance. However, most of these deep methods involve discriminative models, which require large-scale, labeled training datasets, thus hindering their real-world applications. In this paper, we propose a novel strategy to exploit the semantic similarity of the training data and design an efficient generative adversarial framework to learn binary hash codes in an unsupervised manner. Specifically, our model consists of three different neural networks: an encoder network to learn hash codes from images, a generative network to generate images from hash codes, and a discriminative network to distinguish between pairs of hash codes and images. By adversarially training these networks, we successfully learn mutually coherent encoder and generative networks, and can output efficient hash codes from the encoder network. We also propose a novel strategy, which utilizes both feature and neighbor similarities, to construct a semantic similarity matrix, then use this matrix to guide the hash code learning process. Integrating the supervision of this semantic similarity matrix into the adversarial learning framework can efficiently preserve the semantic information of training data in Hamming space. The experimental results on three widely used benchmarks show that our method not only significantly outperforms several state-of-the-art unsupervised hashing methods, but also achieves comparable performance with popular supervised hashing methods. Cheng Deng 0002, Erkun Yang, Tongliang Liu, Jie Li 0001, Wei Liu 0005, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2019 | Re-Ranking High-Dimensional Deep Local Representation for NIR-VIS Face RecognitionabstractHeterogeneous face recognition refers to matching facial images captured from different sensors or sources, which has wide applications in public security and law enforcement. Because of the great differences in sensing and creating procedure, there are huge feature gap between heterogeneous facial images. Existing methods merely focus on comparing the probe image with the gallery in feature space, while the true target may not appear at the first rank due to the appearance variations caused by different sensing patterns. In order to exploit valuable information from initial ranking result, this paper proposes to re-rank high-dimensional deep local representation for matching near-infrared (NIR) and visual (VIS) facial images, i.e. NIR-VIS face recognition. A high-dimensional deep local representation is firstly constructed by extracting and concatenating deep features on local facial patches via a convolutional neural network (CNN). The initial NIR-VIS recognition ranking results can be obtained by comparing the compressed deep features. We then propose a novel and efficient locally linear re-ranking (LLRe-Rank) technique to refine the initial ranking results, which can explore valuable information from initial ranking result. The proposed re-ranking method does not require any human interaction or data annotation, and can be served as an unsupervised post processing technique. Experimental results on the most challenging Oulu-CASIA NIR-VIS database and CASIA NIR-VIS 2.0 database demonstrate the effectiveness of our method. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Dual-Transfer Face Sketch-Photo SynthesisabstractRecognizing the identity of a sketched face from a face photograph dataset is a critical yet challenging task in many applications, not least law enforcement and criminal investigations. An intelligent sketched face identification system would rely on automatic face sketch synthesis from photographs, thereby avoiding the cost of artists manually drawing sketches. However, conventional face sketch-photo synthesis methods tend to generate sketches that are consistent with the artists'drawing styles. Identity-specific information is often overlooked, leading to unsatisfactory identity verification and recognition performance. In this paper, we discuss the reasons why conventional methods fail to recover identity-specific information. Then, we propose a novel dual-transfer face sketch-photo synthesis framework composed of an inter-domain transfer process and an intra-domain transfer process. In the inter-domain transfer, a regressor of the test photograph with respect to the training photographs is learned and transferred to the sketch domain, ensuring the recovery of common facial structures during synthesis. In the intra-domain transfer, a mapping characterizing the relationship between photographs and sketches is learned and transferred across different identities, such that the loss of identity-specific information is suppressed during synthesis. The fusion of information recovered by the two processes is straightforward by virtue of an ad hoc information splitting strategy. We employ both linear and nonlinear formulations to instantiate the proposed framework. Experiments on The Chinese University of Hong Kong face sketch database demonstrate that compared to the current state-of-the-art the proposed framework produces more identifiable facial structures and yields higher face recognition performance in both the photo and sketch domains. Mingjin Zhang, Ruxin Wang 0002, Xinbo Gao 0001, Jie Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2019 | Data Augmentation-Based Joint Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is the process of matching face images captured from different sources. HFR plays an important role in security scenarios. However, HFR remains a challenging problem due to the considerable discrepancies (i.e., shape, style, and color) between cross-modality images. Conventional HFR methods utilize only the information involved in heterogeneous face images, which is not effective because of the substantial differences between heterogeneous face images. To better address this issue, this paper proposes a data augmentation-based joint learning (DA-JL) approach. The proposed method mutually transforms the cross-modality differences by incorporating synthesized images into the learning process. The aggregated data augments the intraclass scale, which provides more discriminative information. However, this method also reduces the interclass diversity (i.e., discriminative information). We develop the DA-JL model to balance this dilemma. Finally, we obtain the similarity score between heterogeneous face image pairs through the log-likelihood ratio. Extensive experiments on a viewed sketch database, forensic sketch database, near-infrared image database, thermal-infrared image database, low-resolution photo database, and image with occlusion database illustrate that the proposed method achieves superior performance in comparison with the state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | A Deep Collaborative Framework for Face Photo-Sketch SynthesisabstractGreat breakthroughs have been made in the accuracy and speed of face photo-sketch synthesis in recent years. Regression-based methods have gained increasing attention, which benefit from deeper and faster end-to-end convolutional neural networks. However, most of these models typically formulate the mapping from photo domain X to sketch domain Y as a unidirectional feedforward mapping, G: X → Y , and vice versa, F: Y → X ; thus, the utilization of mutual interaction between two opposite mappings is lacking. Therefore, we proposed a collaborative framework for face photo-sketch synthesis. The concept behind our model was that a middle latent domain ~Z between the photo domain X and the sketch domain Y can be learned during the learning procedure of G: X → Y and F: Y → X by introducing a collaborative loss that makes full use of two opposite mappings. This strategy can constrain the two opposite mappings and make them more symmetrical, thus making the network more suitable for the photo-sketch synthesis task and obtaining higher quality generated images. Qualitative and quantitative experiments demonstrated the superior performance of our model in comparison with the existing state-of-the-art solutions. Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Asymmetric Joint Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching a probe face image taken from one modality to face images acquired from another modality. It plays an important role in security scenarios. However, HFR is still a challenging problem due to great discrepancies between cross-modality images. This paper proposes an asymmetric joint learning (AJL) approach to handle this issue. The proposed method transforms the cross-modality differences mutually by incorporating the synthesized images into the learning process which provides more discriminative information. Although the aggregated data would augment the scale of intra-classes, it also reduces the diversity (i.e. discriminative information) for inter-classes. Then, we develop the AJL model to balance this dilemma. Finally, we could obtain the similarity score between two heterogeneous face images through the log-likelihood ratio. Extensive experiments on viewed sketch database, forensic sketch database and near infrared image database illustrate that the proposed AJL-HFR method achieve superior performance in comparison to state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
AAAI | 4 |
| 2018 | Deep Attribute Guided Representation for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is a challenging problem in face recognition, subject to large texture and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. Motivated by human cognitive mechanism, we naturally utilize the explicit invariant semantic information (face attributes) to help address the gap of different modalities. Existing related face recognition methods mostly regard attributes as the high level feature integrated with other engineering features enhancing recognition performance, ignoring the inherent relationship between face attributes and identities. In this paper, we propose a novel deep attribute guided representation based heterogeneous face recognition method (DAG-HFR) without labeling attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances mean similarities of pairs. An attribute guided triplet loss (AGTL) is designed to train an end-to-end HFR network which could effectively eliminate defects of incorrectly detected attributes. Extensive experiments on multiple heterogeneous scenarios (composite sketches, resident ID cards) demonstrate that the proposed method achieves superior performances compared with state-of-the-art methods. Decheng Liu, Nannan Wang 0001, Chunlei Peng, Jie Li 0001, Xinbo Gao 0001 |
IJCAI | 4 |
| 2018 | From Reality to Perception: Genre-Based Neural Image Style TransferabstractWe introduce a novel thought for integrating artists’ perceptions on the real world into neural image style transfer process. Conventional approaches commonly migrate color or texture patterns from style image to content image, but the underlying design aspect of the artist always get overlooked. We want to address the in-depth genre style, that how artists perceive the real world and express their perceptions in the artwork. We collect a set of Van Gogh’s paintings and cubist artworks, and their semantically corresponding real world photos. We present a novel genre style transfer framework modeled after the mechanism of actual artwork production. The target style representation is reconstructed based on the semantic correspondence between real world photo and painting, which enable the perception guidance in style transfer. The experimental results demonstrate that our method can capture the overall style of a genre or an artist. We hope that this work provides new insight for including artists’ perceptions into neural style transfer process, and helps people to understand the underlying characters of the artist or the genre. Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
IJCAI | 4 |
| 2018 | Composite components-based face sketch recognition
Decheng Liu, Jie Li 0001, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2018 | Collaborative learning for hyperspectral image classification
Chao Pan 0006, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2018 | A hybrid spatio-temporal model for detection and severity rating of Parkinson's disease from gait data
Aite Zhao, Lin Qi 0004, Jie Li 0001, Junyu Dong, Hui Yu 0001 |
Neurocomputing | 3 |
| 2018 | A parasitic metric learning net for breast mass classification based on mammography
Zhicheng Jiao, Xinbo Gao 0001, Ying Wang 0007, Jie Li 0001 |
Pattern Recognit. | 4 |
| 2018 | Deep Convolutional Neural Networks for mental load classification based on EEG data
Zhicheng Jiao, Xinbo Gao 0001, Ying Wang 0007, Jie Li 0001 |
Pattern Recognit. | 4 |
| 2018 | Face recognition from multiple stylistic sketches: Scenarios, datasets, and evaluation
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
Pattern Recognit. | 4 |
| 2018 | Random sampling for fast face sketch synthesis
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
Pattern Recognit. | 3 |
| 2018 | Back projection: An effective postprocessing method for GAN-based face sketch synthesis
Nannan Wang 0001, Wenjin Zha, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Learning local dictionaries and similarity structures for single image super-resolution
Kaibing Zhang, Jie Li 0001, Xiuping Liu, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2018 | Anchored Neighborhood Index for Face Sketch SynthesisabstractExemplar-based face sketch synthesis has long been impeded by the difficulty of accurate neighbor selection. Given a test patch extracted from the test photograph, the K-nearest neighbor (K-NN) matching algorithm is generally performed by existing methods to find K-nearest photograph patches in the training data set, which contains some pairs of face sketches and photographs. Then, the training sketch patches corresponding to the selected nearest photograph patches are taken as the candidate to synthesize the target sketch patch. In the aforementioned neighbor selection process, training sketch patches is not taken into consideration in the process of K-nearest neighbor selection. In this paper, we proposed a simple yet effective neighbor selection algorithm, namely, anchored neighborhood index (ANI), to boost the synthesis performance by taking training sketch patches into the consideration. In addition, the proposed ANI can be conducted offline and, thus, it does not increase the computational complexity. Extensive experiments on public available database demonstrate that the proposed algorithm achieves superior performance compared with the state-of-the-art methods in terms of both objective image quality scores and face recognition accuracy. Nannan Wang 0001, Xinbo Gao 0001, Leiyu Sun, Jie Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Compositional Model-Based Sketch Generator in Facial EntertainmentabstractFace sketch synthesis (FSS) plays an important role in facial entertainment, which includes face sketch morphing among two styles, multiview FSS and face sketch expression manipulation. For facial entertainment, most existing FSS methods generate sketches with over-smoothing effects, i.e., fine details are suppressed more or less. In this paper, we propose a face sketch generator based on the compositional model to handle this issue. It decomposes a face into different components instead of patches as before, and each component has several candidate templates. Multilevel B-spline approximation is utilized to delicately polish the chosen templates of all components. To fuse these components, Poisson blending is employed instead of the weighted average operator. The proposed compositional method crucially reduces the high frequency loss and improves the synthesis performance in comparison to the state-of-the-art methods. Experiments on face sketch morphing, expression manipulation, and multiview FSS, make further efforts to demonstrate the effectiveness of the proposed method. Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | An Integrated Spatio-Spectral-Temporal Sparse Representation Method for Fusing Remote-Sensing Images With Different ResolutionsabstractDifferent spectral, spatial, and temporal features have been widely used in the remote-sensing image analysis. The further development of multiple sensor remote-sensing technologies has made it necessary to explore new methods of remote-sensing image fusion using different optical image data sets which provide complementary image properties and a tradeoff among spatial, spectral, and temporal resolutions. However, due to problems in assessing correlations between different types of satellite data with different resolutions, a few efforts have been made to explore spatio-spectral–temporal features. For this purpose, we propose a novel sparse representation model to generate synthesized frequent high-spectral and high-spatial resolution data by blending multiple types: spatio-temporal data fusion, spectral–temporal data fusion, spatio-spectral data fusion, and spatio-spectral–temporal data fusion. The proposed method exploits high-spectral correlation across spectral domains and high self-similarity across spatial domains to learn the spatio-spectral fusion basis. Then, it associates temporal changes using a local constraint sparse representation. The integrated spatio-spectral–temporal sparse representation model based on the learned spectral–spatial and temporal change features strengthens the model’s ability to provide high-resolution data needed to address demanding work in real-world applications. Finally, the proposed method is not restricted to a certain type of data, but it can associate any type of remote-sensing data and be applied to dynamic changes in heterogeneous landscapes. The experimental results illustrate the effectiveness and efficiency of the proposed method. Chongyue Zhao, Xinbo Gao 0001, William J. Emery, Ying Wang 0007, Jie Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Single Image Super-Resolution via Multiple Mixture Prior ModelsabstractExample learning-based single image super-resolution (SR) is a promising method for reconstructing a high-resolution (HR) image from a single-input low-resolution (LR) image. Lots of popular SR approaches are more likely either time-or space-intensive, which limit their practical applications. Hence, some research has focused on a subspace view and delivered state-of-the-art results. In this paper, we utilize an effective way with mixture prior models to transform the large nonlinear feature space of LR images into a group of linear subspaces in the training phase. In particular, we first partition image patches into several groups by a novel selective patch processing method based on difference curvature of LR patches, and then learning the mixture prior models in each group. Moreover, different prior distributions have various effectiveness in SR, and in this case, we find that student-t prior shows stronger performance than the well-known Gaussian prior. In the testing phase, we adopt the learned multiple mixture prior models to map the input LR features into the appropriate subspace, and finally reconstruct the corresponding HR image in a novel mixed matching way. Experimental results indicate that the proposed approach is both quantitatively and qualitatively superior to some state-of-the-art SR methods. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
IEEE Trans. Image Process. | 2 |
| 2018 | Shared Predictive Cross-Modal Deep QuantizationabstractWith explosive growth of data volume and ever-increasing diversity of data modalities, cross-modal similarity search, which conducts nearest neighbor search across different modalities, has been attracting increasing interest. This paper presents a deep compact code learning solution for efficient cross-modal similarity search. Many recent studies have proven that quantization-based approaches perform generally better than hashing-based approaches on single-modal similarity search. In this paper, we propose a deep quantization approach, which is among the early attempts of leveraging deep neural networks into quantization-based cross-modal similarity search. Our approach, dubbed shared predictive deep quantization (SPDQ), explicitly formulates a shared subspace across different modalities and two private subspaces for individual modalities, and representations in the shared subspace and the private subspaces are learned simultaneously by embedding them to a reproducing kernel Hilbert space, where the mean embedding of different modality distributions can be explicitly compared. In addition, in the shared subspace, a quantizer is learned to produce the semantics preserving compact codes with the help of label alignment. Thanks to this novel network architecture in cooperation with supervised quantization training, SPDQ can preserve intramodal and intermodal similarities as much as possible and greatly reduce quantization error. Experiments on two popular benchmarks corroborate that our approach outperforms state-of-the-art methods. Erkun Yang, Cheng Deng 0002, Chao Li 0033, Wei Liu 0005, Jie Li 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Semantic Segmentation Based Automatic Two-Tone Portrait Synthesis
Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
ICIG (3) | 4 |
| 2017 | Deep Graphical Feature Learning for Face Sketch SynthesisabstractThe exemplar-based face sketch synthesis method generally contains two steps: neighbor selection and reconstruction weight representation. Pixel intensities are widely used as features by most of the existing exemplar-based methods, which lacks of representation ability and robustness to light variations and clutter backgrounds. We present a novel face sketch synthesis method combining generative exemplar-based method and discriminatively trained deep convolutional neural networks (dCNNs) via a deep graphical feature learning framework. Our method works in both two steps by using deep discriminative representations derived from dCNNs. Instead of using it directly, we boost its representation capability by a deep graphical feature learning framework. Finally, the optimal weights of deep representations and optimal reconstruction weights for face sketch synthesis can be obtained simultaneously. With the optimal reconstruction weights, we can synthesize high quality sketches which is robust against light variations and clutter backgrounds. Extensive experiments on public face sketch databases show that our method outperforms state-of-the-art methods, in terms of both synthesis quality and recognition ability. Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
IJCAI | 4 |
| 2017 | Adaptive representation-based face sketch-photo synthesis
Jie Li 0001, Xinye Yu, Chunlei Peng, Nannan Wang 0001 |
Neurocomputing | 1 |
| 2017 | Data-driven vs. model-driven: Fast face sketch synthesis
Nannan Wang 0001, Mingrui Zhu, Jie Li 0001, Bin Song 0001, Zan Li 0001 |
Neurocomputing | 3 |
| 2017 | Graphical Representation for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching face images acquired from different sources (i.e., different sensors or different wavelengths) for identification. HFR plays an important role in both biometrics research and industry. In spite of promising progresses achieved in recent years, HFR is still a challenging problem due to the difficulty to represent two heterogeneous images in a homogeneous manner. Existing HFR methods either represent an image ignoring the spatial information, or rely on a transformation procedure which complicates the recognition task. Considering these problems, we propose a novel graphical representation based HFR method (G-HFR) in this paper. Markov networks are employed to represent heterogeneous image patches separately, which takes the spatial compatibility between neighboring image patches into consideration. A coupled representation similarity metric (CRSM) is designed to measure the similarity between obtained graphical representations. Extensive experiments conducted on multiple HFR scenarios (viewed sketch, forensic sketch, near infrared image, and thermal infrared image) show that the proposed method outperforms state-of-the-art methods. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Fast single image super-resolution using sparse Gaussian process regression
Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
Signal Process. | 4 |
| 2017 | Unified framework for face sketch synthesis
Nannan Wang 0001, Shengchuan Zhang, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001 |
Signal Process. | 4 |
| 2017 | Superpixel-Based Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis technique has attracted growing attention in many computer vision applications, such as law enforcement and digital entertainment. Existing methods either simply perform the face sketch-photo synthesis on the holistic image or divide the face image into regular rectangular patches ignoring the inherent structure of the face image. In view of such situations, this paper presents a novel superpixel-based face sketch-photo synthesis method by estimating the face structures through image segmentation. In our proposed method, face images are first segmented into superpixels, which are then dilated to enhance the compatibility of neighboring superpixels. Each input face image induces a specific graphical structure modeled by Markov networks. We employ a two-stage synthesis process to learn the face structures through Markov networks constructed from two scales of dilation, respectively. Experiments on several public databases demonstrate that our proposed face sketch-photo synthesis method achieves superior performance compared with the state-of-the-art methods. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Face Sketch Synthesis From a Single Photo-Sketch PairabstractFace sketch synthesis is crucial in many practical applications, such as digital entertainment and law enforcement. Previous methods relying on many photo-sketch pairs have made great progress. State-of-the-art face sketch synthesis algorithms adopt Bayesian inference (BI) (e.g., Markov random fields) to select local sketch patches around corresponding position from a set of training data. However, these methods have two limitations: 1) they depend on many training photo-sketch pairs and 2) they cannot tackle nonfacial factors (e.g., hairpins, glasses, backgrounds, and image size) if these factors are excluded in training data. In this paper, we propose a novel face sketch synthesis method that is capable of handling nonfacial factors only using a single photo-sketch pair from coarse to fine. Our method proposes a cascaded image synthesis (CIS) strategy and integrates sparse representation-based greedy search (SRGS) and BI for face sketch synthesis. We first apply SRGS to select candidate sketch patches from the whole training photo-sketch pairs sampled from the only photo-sketch pair. We then employ BI to estimate an initial sketch. Afterward, the input photo and the estimated initial sketch are taken as an additional photo-sketch pair for training. Finally, we adopt CIS with the given two photo-sketch pairs to further improve the quality of the initial sketch. The experimental results on several databases demonstrate that our algorithm outperforms state-of-the-art methods. Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Bayesian Face Sketch SynthesisabstractExemplar-based face sketch synthesis has been widely applied to both digital entertainment and law enforcement. In this paper, we propose a Bayesian framework for face sketch synthesis, which provides a systematic interpretation for understanding the common properties and intrinsic difference in different methods from the perspective of probabilistic graphical models. The proposed Bayesian framework consists of two parts: the neighbor selection model and the weight computation model. Within the proposed framework, we further propose a Bayesian face sketch synthesis method. The essential rationale behind the proposed Bayesian method is that we take the spatial neighboring constraint between adjacent image patches into consideration for both aforementioned models, while the state-of-the-art methods neglect the constraint either in the neighbor selection model or in the weight computation model. Extensive experiments on the Chinese University of Hong Kong face sketch database demonstrate that the proposed Bayesian method could achieve superior performance compared with the state-of-the-art methods in terms of both subjective perceptions and objective evaluations. Nannan Wang 0001, Xinbo Gao 0001, Leiyu Sun, Jie Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Single Image Super-Resolution Using Gaussian Process Regression With Dictionary-Based Sampling and Student-t LikelihoodabstractGaussian process regression (GPR) is an effective statistical learning method for modeling non-linear mapping from an observed space to an expected latent space. When applying it to example learning-based super-resolution (SR), two outstanding issues remain. One is its high computational complexity restricts SR application when a large data set is available for learning task. The other is that the commonly used Gaussian likelihood in GPR is incompatible with the true observation model for SR reconstruction. To alleviate the above two issues, we propose a GPR-based SR method by using dictionary-based sampling (DbS) and student-t likelihood. Considering that dictionary atoms effectively span the original training sample space, we adopt a DbS strategy by combining all the neighborhood samples of each atom into a compact representative training subset so as to reduce the computational complexity. Based on statistical tests, we statistically validate that student-t likelihood is more suitable to build the observation model for the SR problem. Extensive experimental results show that the proposed method outperforms other competitors and produces more pleasing details in texture regions. Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Coarse-to-Fine Learning for Single-Image Super-ResolutionabstractThis paper develops a coarse-to-fine framework for single-image super-resolution (SR) reconstruction. The coarse-to-fine approach achieves high-quality SR recovery based on the complementary properties of both example learning-and reconstruction-based algorithms: example learning-based SR approaches are useful for generating plausible details from external exemplars but poor at suppressing aliasing artifacts, while reconstruction-based SR methods are propitious for preserving sharp edges yet fail to generate fine details. In the coarse stage of the method, we use a set of simple yet effective mapping functions, learned via correlative neighbor regression of grouped low-resolution (LR) to high-resolution (HR) dictionary atoms, to synthesize an initial SR estimate with particularly low computational cost. In the fine stage, we devise an effective regularization term that seamlessly integrates the properties of local structural regularity, nonlocal self-similarity, and collaborative representation over relevant atoms in a learned HR dictionary, to further improve the visual quality of the initial SR estimation obtained in the coarse stage. The experimental results indicate that our method outperforms other state-of-the-art methods for producing high-quality images despite that both the initial SR estimation and the followed enhancement are cheap to implement. Kaibing Zhang, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2016 | An Integrated Method for Road Network Centerline Detection from Multispectral Imagery
Yuefan Du, Jie Li 0001, Ying Wang 0007 |
IDEAL | 2 |
| 2016 | Hyperspectral image super-resolution using sparse spectral unmixing and low-rank constraintsabstractHyperspectral images play an important role in real-world applications, such as recognition and remote sensing, etc. How to enhance the spatial resolution of hyperspectral image is still a challenging problem in this field. In this paper, we propose a novel hyperspectral image super-resolution approach by jointly incorporating the sparse, low-rank constraints and spectral mixture priori into a linear unmixing framework, which will make the unmixing framework more consistent with the real-world scenarios of the spectral mixture. Experiments on two public databases show that our proposed approach achieves much lower average reconstruction errors than other state-of-the-art methods. Chao Li 0033, Cheng Deng 0002, Jie Li 0001 |
IGARSS | 4 |
| 2016 | Multispectral image classification based on improved weighted MRF Bayesian
Zhaobin Cui, Ying Wang 0007, Xinbo Gao 0001, Jie Li 0001, Yu Zheng 0006 |
Neurocomputing | 4 |
| 2016 | A deep feature based framework for breast masses classification
Zhicheng Jiao, Xinbo Gao 0001, Ying Wang 0007, Jie Li 0001 |
Neurocomputing | 4 |
| 2016 | Evaluation on synthesized face sketches
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001 |
Neurocomputing | 3 |
| 2016 | Image super-resolution using non-local Gaussian process regression
Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
Neurocomputing | 4 |
| 2016 | A novel dimensionality reduction method with discriminative generalized eigen-decomposition
Xiumei Wang 0002, Weifang Liu, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2016 | Single image super-resolution using regularization of non-local steering kernel regression
Kaibing Zhang, Xinbo Gao 0001, Jie Li 0001, Hongxing Xia |
Signal Process. | 3 |
| 2016 | Efficient Multiple-Feature Learning-Based Hyperspectral Image Classification With Limited Training SamplesabstractLinearly derived features have been widely used in hyperspectral image classification to find linear separability of certain classes in recent years. Moreover, nonlinearly transformed features are more effective for class discrimination in real analysis scenarios. However, few efforts have attempted to combine both linear and nonlinear features in the same framework even if they can demonstrate some complementary properties. Moreover, conventional multiple-feature learning-based approaches deal with different features equally, which is not reasonable. This paper proposes an efficient multiple-feature learning-based model with adaptive weights for effectively classifying complex hyperspectral images with limited training samples. A new diversity kernel function is proposed first to simulate the vision perception and analysis procedure of human beings. It could simultaneously evaluate the contrast differences of global features and spatial coherence. Since existing multiple-kernel feature models are always time-consuming, we then design a new adaptive weighted multiple kernel learning method. It employs kernel projection, which could lower the dimensionalities and also learn kernel weights to further discriminate the classification boundaries. For combining both linear and nonlinear features, this paper also proposes a novel decision fusion strategy. The method combines linear and multiple kernel features to balance the classification results of different classifiers. The proposed scheme is tested on several hyperspectral data sets and extended to multisource feature classification environment. The experimental results show that the proposed classification method outperforms most of the existing ones and significantly reduces the computational complexity. Chongyue Zhao, Xinbo Gao 0001, Ying Wang 0007, Jie Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | Single-Image Super-Resolution Using Active-Sampling Gaussian Process RegressionabstractAs well known, Gaussian process regression (GPR) has been successfully applied to example learning-based image super-resolution (SR). Despite its effectiveness, the applicability of a GPR model is limited by its remarkably computational cost when a large number of examples are available to a learning task. For this purpose, we alleviate this problem of the GPR-based SR and propose a novel example learning-based SR method, called active-sampling GPR (AGPR). The newly proposed approach employs an active learning strategy to heuristically select more informative samples for training the regression parameters of the GPR model, which shows significant improvement on computational efficiency while keeping higher quality of reconstructed image. Finally, we suggest an accelerating scheme to further reduce the time complexity of the proposed AGPR-based SR by using a pre-learned projection matrix. We objectively and subjectively demonstrate that the proposed method is superior to other competitors for producing much sharper edges and finer details. Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Robust Face Sketch Style SynthesisabstractHeterogeneous image conversion is a critical issue in many computer vision tasks, among which example-based face sketch style synthesis provides a convenient way to make artistic effects for photos. However, existing face sketch style synthesis methods generate stylistic sketches depending on many photo-sketch pairs. This requirement limits the generalization ability of these methods to produce arbitrarily stylistic sketches. To handle such a drawback, we propose a robust face sketch style synthesis method, which can convert photos to arbitrarily stylistic sketches based on only one corresponding template sketch. In the proposed method, a sparse representation-based greedy search strategy is first applied to estimate an initial sketch. Then, multi-scale features and Euclidean distance are employed to select candidate image patches from the initial estimated sketch and the template sketch. In order to further refine the obtained candidate image patches, a multi-feature-based optimization model is introduced. Finally, by assembling the refined candidate image patches, the completed face sketch is obtained. To further enhance the quality of synthesized sketches, a cascaded regression strategy is adopted. Compared with the state-of-the-art face sketch synthesis methods, experimental results on several commonly used face sketch databases and celebrity photos demonstrate the effectiveness of the proposed method. Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Multiple Representations-Based Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis plays an important role in law enforcement and digital entertainment. Most of the existing methods only use pixel intensities as the feature. Since face images can be described using features from multiple aspects, this paper presents a novel multiple representations-based face sketch-photo-synthesis method that adaptively combines multiple representations to represent an image patch. In particular, it combines multiple features from face images processed using multiple filters and deploys Markov networks to exploit the interacting relationships between the neighboring image patches. The proposed framework could be solved using an alternating optimization strategy and it normally converges in only five outer iterations in the experiments. Our experimental results on the Chinese University of Hong Kong (CUHK) face sketch database, celebrity photos, CUHK Face Sketch FERET Database, IIIT-D Viewed Sketch Database, and forensic sketches demonstrate the effectiveness of our method for face sketch-photo synthesis. In addition, cross-database and database-dependent style-synthesis evaluations demonstrate the generalizability of this novel method and suggest promising solutions for face identification in forensic science. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2016 | Shape-Constrained Sparse and Low-Rank Decomposition for Auroral Substorm DetectionabstractAn auroral substorm is an important geophysical phenomenon that reflects the interaction between the solar wind and the Earth's magnetosphere. Detecting substorms is of practical significance in order to prevent disruption to communication and global positioning systems. However, existing detection methods can be inaccurate or require time-consuming manual analysis and are therefore impractical for large-scale data sets. In this paper, we propose an automatic auroral substorm detection method based on a shape-constrained sparse and low-rank decomposition (SCSLD) framework. Our method automatically detects real substorm onsets in large-scale aurora sequences, which overcomes the limitations of manual detection. To reduce noise interference inherent in current SLD methods, we introduce a shape constraint to force the noise to be assigned to the low-rank part (stationary background), thus ensuring the accuracy of the sparse part (moving object) and improving the performance. Experiments conducted on aurora sequences in solar cycle 23 (1996-2008) show that the proposed SCSLD method achieves good performance for motion analysis of aurora sequences. Moreover, the obtained results are highly consistent with manual analysis, suggesting that the proposed automatic method is useful and effective in practice. Xi Yang 0011, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Bing Han 0003, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2015 | Dirichlet-Based Concentric Circle Feature Transform for Breast Mass ClassificationabstractBreast cancer has caused more and more attention in recent years since the mortality rate is increasing and age of onset is trend to be younger than before. Using computer vision technology for automatic classifying benign and masses malignant ones could assist doctors in diagnosing condition. However, the margins and shapes of masses are various and which are very similar with surrounding tissues, there are still a lot of problems in classification for which the extracted features couldn't express the original image very well. Hence, in this paper, a new mass feature extraction method is proposed for enhancing the performance of classification. First, the concentric circle bag-of-words (BOW) of the mass is captured spatially from the mammographic images. Then, the Dirichlet fisher kernel is conducted to enhance discrimination ability of the feature. Finally, the SVM classifier is employed for mass classification. Experiments are conducted on DDSM dataset and achieve good classification accuracy. Min Pang, Ying Wang 0007, Jie Li 0001 |
ICTAI | 3 |
| 2015 | A level set method with shape priors by using locality preserving projections
Bin Wang 0027, Xinbo Gao 0001, Jie Li 0001, Xuelong Li 0001, Dacheng Tao |
Neurocomputing | 3 |
| 2015 | Recognition of facial sketch styles
Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2015 | Large-scale multi-task image labeling with adaptive relevance discovery and feature hashing
Cheng Deng 0002, Xianglong Liu 0001, Yadong Mu, Jie Li 0001 |
Signal Process. | 4 |
| 2015 | An Efficient MRF Embedded Level Set Method for Image SegmentationabstractThis paper presents a fast and robust level set method for image segmentation. To enhance the robustness against noise, we embed a Markov random field (MRF) energy function to the conventional level set energy function. This MRF energy function builds the correlation of a pixel with its neighbors and encourages them to fall into the same region. To obtain a fast implementation of the MRF embedded level set model, we explore algebraic multigrid (AMG) and sparse field method (SFM) to increase the time step and decrease the computation domain, respectively. Both AMG and SFM can be conducted in a parallel fashion, which facilitates the processing of our method for big image databases. By comparing the proposed fast and robust level set method with the standard level set method and its popular variants on noisy synthetic images, synthetic aperture radar (SAR) images, medical images, and natural images, we comprehensively demonstrate the new method is robust against various kinds of noises. In particular, the new level set method can segment an image of size 500 × 500 within 3 s on MATLAB R2010b installed in a computer with 3.30-GHz CPU and 4-GB memory. Xi Yang 0011, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | Face Sketch Synthesis via Sparse Representation-Based Greedy SearchabstractFace sketch synthesis has wide applications in digital entertainment and law enforcement. Although there is much research on face sketch synthesis, most existing algorithms cannot handle some nonfacial factors, such as hair style, hairpins, and glasses if these factors are excluded in the training set. In addition, previous methods only work on well controlled conditions and fail on images with different backgrounds and sizes as the training set. To this end, this paper presents a novel method that combines both the similarity between different image patches and prior knowledge to synthesize face sketches. Given training photo-sketch pairs, the proposed method learns a photo patch feature dictionary from the training photo patches and replaces the photo patches with their sparse coefficients during the searching process. For a test photo patch, we first obtain its sparse coefficient via the learnt dictionary and then search its nearest neighbors (candidate patches) in the whole training photo patches with sparse coefficients. After purifying the nearest neighbors with prior knowledge, the final sketch corresponding to the test photo can be obtained by Bayesian inference. The contributions of this paper are as follows: 1) we relax the nearest neighbor search area from local region to the whole image without too much time consuming and 2) our method can produce nonfacial factors that are not contained in the training set and is robust against image backgrounds and can even ignore the alignment and image size aspects of test photos. Our experimental results show that the proposed method outperforms several state-of-the-arts in terms of perceptual and objective metrics. Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Mingjin Zhang |
IEEE Trans. Image Process. | 4 |
| 2014 | A Comprehensive Survey to Face Hallucination
Nannan Wang 0001, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2014 | Latent feature mining of spatial and marginal characteristics for mammographic mass classification
Ying Wang 0007, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2014 | A shape-initialized and intensity-adaptive level set method for auroral oval segmentation
Xi Yang 0011, Xinbo Gao 0001, Jie Li 0001, Bing Han 0003 |
Inf. Sci. | 3 |
| 2013 | Heterogeneous image transformation
Nannan Wang 0001, Jie Li 0001, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001 |
Pattern Recognit. Lett. | 2 |
| 2013 | Transductive Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis plays a critical role in many applications, such as law enforcement and digital entertainment. Recently, many face sketch-photo synthesis methods have been proposed under the framework of inductive learning, and these have obtained promising performance. However, these inductive learning-based face sketch-photo synthesis methods may result in high losses for test samples, because inductive learning minimizes the empirical loss for training samples. This paper presents a novel transductive face sketch-photo synthesis method that incorporates the given test samples into the learning process and optimizes the performance on these test samples. In particular, it defines a probabilistic model to optimize both the reconstruction fidelity of the input photo (sketch) and the synthesis fidelity of the target output sketch (photo), and efficiently optimizes this probabilistic model by alternating optimization. The proposed transductive method significantly reduces the expected high loss and improves the synthesis performance for test samples. Experimental results on the Chinese University of Hong Kong face sketch data set demonstrate the effectiveness of the proposed method by comparing it with representative inductive learning-based face sketch-photo synthesis methods. Nannan Wang 0001, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2012 | Robust Reversible Watermarking via Clustering and Enhanced Pixel-Wise MaskingabstractRobust reversible watermarking (RRW) methods are popular in multimedia for protecting copyright, while preserving intactness of host images and providing robustness against unintentional attacks. However, conventional RRW methods are not readily applicable in practice. That is mainly because 1) they fail to offer satisfactory reversibility on large-scale image datasets; 2) they have limited robustness in extracting watermarks from the watermarked images destroyed by different unintentional attacks; and 3) some of them suffer from extremely poor invisibility for watermarked images. Therefore, it is necessary to have a framework to address these three problems, and further improve its performance. This paper presents a novel pragmatic framework, wavelet-domain statistical quantity histogram shifting and clustering (WSQH-SC). Compared with conventional methods, WSQH-SC ingeniously constructs new watermark embedding and extraction procedures by histogram shifting and clustering, which are important for improving robustness and reducing run-time complexity. Additionally, WSQH-SC includes the property inspired pixel adjustment (PIPA) to effectively handle overflow and underflow of pixels. This results in satisfactory reversibility and invisibility. Furthermore, to increase its practical applicability, WSQH-SC designs an enhanced pixel-wise masking (EPWM) to balance robustness and invisibility. We perform extensive experiments over natural, medical, and synthetic aperture radar (SAR) images to show the effectiveness of WSQH-SC by comparing with the histogram rotation (HR)-based and histogram distribution constrained (HDC) methods. Lingling An, Xinbo Gao 0001, Xuelong Li 0001, Dacheng Tao, Cheng Deng 0002, Jie Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2011 | Single-step-lag OOSM algorithm based on unscented transformation
Jinguang Chen, Jie Li 0001, Xinbo Gao 0001 |
Sci. China Inf. Sci. | 2 |
| 2011 | Chinese text location under complex background using Gabor filter and SVM
Jianqiang Yan, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2010 | Semi-supervised Gaussian process latent variable model with pairwise constraints
Xiumei Wang 0002, Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao, Jie Li 0001 |
Neurocomputing | 5 |
| 2010 | Incremental pairwise discriminant analysis based visual tracking
Xinbo Gao 0001, Xuelong Li 0001, Dacheng Tao, Jie Li 0001 |
Neurocomputing | 5 |
| 2010 | Incremental tensor biased discriminant analysis: A new color-based visual tracking method
Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao, Jie Li 0001 |
Neurocomputing | 5 |
| 2010 | Photo-sketch synthesis and recognition based on subspace learning
Bing Xiao 0003, Xinbo Gao 0001, Dacheng Tao, Yuan Yuan 0001, Jie Li 0001 |
Neurocomputing | 5 |
| 2010 | An integrated aurora image retrieval system: AuroraEye
Xinbo Gao 0001, Xuelong Li 0001, Dacheng Tao, Yongjun Jian, Jie Li 0001, Hongqiao Hu, Huigen Yang |
J. Vis. Commun. Image Represent. | 6 |
| 2009 | Sinusoidal Signals Pattern Based Robust Video Watermarking in the 3D-CWT DomainabstractThis paper presents a robust video watermarking scheme. It mainly contains three characteristics: 1) the sinusoidal signals pattern is embedded as watermark in the selected sub-bands of 3D dual-tree complex wavelet transform (3D-CWT), which is employed to preserve the image quality and improve the robustness of watermark; 2) at the detection end, the detected peaks response are used to achieve attacks estimation and rectification; and 3) the watermark is confirmed simply and objectively by detecting the prominent peaks in frequency domain. Experimental results conducted on several popular true color video sequences demonstrate that the proposed watermarking scheme has good performance in terms of common video operations as well as geometric distortions. Cheng Deng 0002, Jie Li 0001, Xinbo Gao 0001 |
IAS | 2 |
| 2009 | The Gabor-Based Tensor Level Set Method for Multiregional Image Segmentation
Bin Wang 0027, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001 |
CAIP | 5 |
| 2008 | Face Sketch Synthesis Algorithm Based on E-HMM and Selective EnsembleabstractSketch synthesis plays an important role in face sketch-photo recognition system. In this manuscript, an automatic sketch synthesis algorithm is proposed based on embedded hidden Markov model (E-HMM) and selective ensemble strategy. First, the E-HMM is adopted to model the nonlinear relationship between a sketch and its corresponding photo. Then based on several learned models, a series of pseudo-sketches are generated for a given photo. Finally, these pseudo-sketches are fused together with selective ensemble strategy to synthesize a finer face pseudo-sketch. Experimental results illustrate that the proposed algorithm achieves satisfactory effect of sketch synthesis with a small set of face training samples. Xinbo Gao 0001, Juanjuan Zhong, Jie Li 0001, Chunna Tian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | A Feature Analysis Approach to Mass Detection in Mammography Based on RF-SVMabstractA new approach to mass detection in mammography is presented. The main obstacle of building a mass detection system is the similar appearance between masses and density tissues in breast. Hence, the various features of the extracted regions of interest (ROIs) are analyzed by synthesis. Then the support vector machine (SVM), which is designed later to distinguish masses from normal areas, is employed to classify these ROIs exactly. To further improve the performance of SVM, the relevance feedback (RF) is introduced to filter out the false positives. The experimental results illustrate that SVM classifier can effectively detect the mass areas, and the RF-SVM scheme can be efficiently incorporated into this learning framework to further improve detection performance. Ying Wang 0007, Xinbo Gao 0001, Jie Li 0001 |
ICIP (5) | 3 |
| 2006 | A Valid Multi-View Face Detection Tree Based on Floatboost LearningabstractA novel face detection tree based on floatboost learning is proposed to accommodate the in-class variability of multi-view faces. The tree splitting procedure is realized through dividing face training examples into the optimal sub-clusters using the fuzzy c-means (FCM) algorithm together with a new cluster validity function based on the modified partition fuzzy degree. Then each sub-cluster of face examples is conquered with the floatboost learning to construct branches in the node of the detection tree. During training, the proposed algorithm is much faster than the original detection tree. The experimental results on the CMU and our home-brew test database illustrate that the proposed detection tree is more efficient than the original one while keeping its detection speed. Chunna Tian, Xinbo Gao 0001, Jie Li 0001 |
ICIP | 3 |
| 2006 | A Cartoon Video Detection Method Based on Active Relevance Feedback and SVM
Xinbo Gao 0001, Jie Li 0001 |
ISNN (2) | 2 |
| 2005 | A Novel Clustering Method Based on SVM
Jie Li 0001, Xinbo Gao 0001, Licheng Jiao |
ISNN (2) | 1 |
| 2004 | A novel clustering method with network structure based on clonal algorithmabstractIn the field of cluster analysis, the objective function based clustering algorithm is one of the most widely applied methods. However, this type of algorithm, which needs the priori knowledge about the cluster number and the form of clustering prototypes, can only process data sets with the same type of prototypes. Moreover, these algorithms are very sensitive to the initialization and easy to get trapped into local optima. This paper presents a novel clustering method, with network structure based on a clonal algorithm, to realize the automatization of cluster analysis. By analyzing the neurons of the obtained network with a minimal spanning tree, one can easily get the cluster number and the related classification information. The test results with various data sets illustrate that the novel algorithm achieves more effective performance on cluster analyzing data sets with mixed numeric values and categorical values. Jie Li 0001, Xinbo Gao 0001, Licheng Jiao |
ICASSP (5) | 1 |