VLDB 2026 Research / reviewers in the wild / expert
Jianfeng Lu 0003
dblp:82/6187-3
· DBLP profile ↗
130ranked-venue papers
2as first author
72since 2021 · last 2026
0000-0002-9190-507XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 1 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 1 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 19 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Systems, architecture and hardware · 4 · 2 since 2021Security and privacy · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM AgentsabstractXiangyu Wu, Yuwei Hu, Tianyu Cui, Yueying Tian, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yang Yang, Jianfeng Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyu Cui, Yueying Tian, Weihua Luo, Kaifu Zhang, Yang Yang 0074, Jianfeng Lu 0003 |
ACL (1) | 10 |
| 2026 | Feature interaction modelling using a similarity-based adaptive graph attention network for click-through rate predictionabstractIn the field of recommendation systems, click-through rate (CTR) prediction is essential for measuring user engagement and product interest. However, input features are typically high-dimensional and sparse, requiring effective modeling of high-order feature interactions. Existing methods learn low-dimensional representations and identify useful feature combinations, but they often struggle to capture implicit interactions in non-Euclidean spaces, leading to noisy feature relationships, reduced interpretability, and suboptimal prediction performance. To address these challenges, we introduce a similarity-based adaptive graph attention network (SGAT) for CTR prediction. The SGAT employs a dual-attention mechanism that integrates similarity-based adaptive attention with softmax attention to selectively emphasize beneficial feature interactions while suppressing irrelevant ones. This strategy reduces noise, enhances interaction modelling, and mitigates overfitting. Experimental results on two public benchmark datasets show that SGAT achieves superior performance compared with several state-of-the-art baseline methods. Aniqa Nawaz, Zhengwang Xia, Jianfeng Lu 0003 |
Connect. Sci. | 3 |
| 2026 | Bridging the resolution gap: Semantic-aware alignment for cross-resolution change detection
Wang Hao, Fengchao Xiong, Jianfeng Lu 0003, Jingzhou Chen, Yuntao Qian |
Pattern Recognit. | 4 |
| 2026 | A unified spatial-spectral-temporal network for hyperspectral object tracking
Zhuanfeng Li, Jing Wang 0062, Jue Zhang 0001, Dong Zhao 0005, Guanyiman Fu, Jianfeng Lu 0003 |
Pattern Recognit. | 7 |
| 2026 | Joint Dynamic Brain Network Estimation and Graph Representation Learning for the Recognition of Neurological DisordersabstractRecently, Graph Neural Networks (GNNs) have shown significant improvements in the recognition of neurological disorders by incorporating brain networks/graphs. However, most existing approaches have three main limitations. First, these methodologies rely on precomputed brain networks as input, typically derived from statistical metrics (e.g., Pearson correlation), which are inherently not learnable. Second, methods often assume that the magnitude of the brain interactions remains constant across the whole scan duration. Third, representations produced by models often lack interpretability and robustness when applied across brain disorders. To address these limitations, we propose a novel model called the Effective Brain Inference Graph Neural Network (EBIGNN), which infers dynamic Effective Connectivity (dEC) to characterize brain networks trained with direct feedback from downstream tasks within a unified end-to-end framework. EBIGNN is highly flexible in learning the most relevant graph structures customized to the specific underlying brain condition. The proposed model offers strong interpretability, providing valuable insights into the temporal evolution and altered connectivity patterns essential for understanding brain disorders. The model is validated on three publicly available datasets, demonstrating superior performance compared to other state-of-the-art methods. Moreover, the findings are consistent with previous neuroimaging-derived evidence of biomarkers, underscoring the model's robustness in clinical settings. Saqib Mamoon, Zhengwang Xia, Wang Jin, Amani Alfakih, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multi-Label Test-Time Adaptation with Bound Entropy MinimizationabstractMainstream test-time adaptation (TTA) techniques endeavor to mitigate distribution shifts via entropy minimization for multi-class classification, inherently increasing the probability of the most confident class. However, when encountering multi-label instances, the primary challenge stems from the varying number of labels per image, and prioritizing only the highest probability class inevitably undermines the adaptation of other positive labels. To address this issue, we investigate TTA within multi-label scenario (ML--TTA), developing Bound Entropy Minimization (BEM) objective to simultaneously increase the confidence of multiple top predicted labels. Specifically, to determine the number of labels for each augmented view, we retrieve a paired caption with yielded textual labels for that view. These labels are allocated to both the view and caption, called weak label set and strong label set with the same size k. Following this, the proposed BEM considers the highest top-k predicted labels from view and caption as a single entity, respectively, learning both view and caption prompts concurrently. By binding top-k predicted labels, BEM overcomes the limitation of vanilla entropy minimization, which exclusively optimizes the most confident class. Across the MSCOCO, VOC, and NUSWIDE multi-label datasets, our ML--TTA framework equipped with BEM exhibits superior performance compared to the latest SOTA methods, across various model architectures, prompt initialization, and varying label scenarios. The code is available at https://github.com/Jinx630/ML-TTA. Feng Yu 0030, Yang Yang 0074, Jianfeng Lu 0003 |
ICLR | 5 |
| 2025 | Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental ControlabstractGenerating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions. While diffusion models have shown promise in generative tasks, their ability to maintain strict pose alignment throughout the diffusion process is limited. In this paper, we propose a novel Iterative Homography Adjustment (IHA) scheme applied during the denoising process, which effectively addresses pose misalignment and ensures spatial consistency in the generated street-view images. Additionally, currently, available datasets for satellite-to-street-view generation are limited in their diversity of illumination and weather conditions, thereby restricting the generalizability of the generated outputs. To mitigate this, we introduce a text-guided illumination and weather-controlled sampling strategy that enables fine-grained control over the environmental factors. Extensive quantitative and qualitative evaluations demonstrate that our approach significantly improves pose accuracy and enhances the diversity and realism of generated street-view images, setting a new benchmark for satellite-to-street-view generation tasks. Xianghui Ze, Zhenbo Song, Jianfeng Lu 0003, Yujiao Shi 0002 |
ICLR | 4 |
| 2025 | Safety-constrained Reinforcement Learning with Interaction-aware for Decision-making of Autonomous DrivingabstractReinforcement learning(RL) has made significant advancements in autonomous driving(AD). However, the stochastic nature of dynamic traffic scenario and the diversity of road type make it challenging for autonomous vehicles to make safe and efficient decisions. To tackle these problems, this paper proposes a novel RL framework that incorporates the motion prediction model to enhance the agent’s decision-making capability. We first utilize Transformer to model driving scenarios and capture interaction-aware relationships between the ego vehicle and scenarios, then design a safety-constraint and integrate it into the Proximal Policy Optimization (PPO) algorithm so as to guarantee the safety and feasibility of the policy. To improve data efficiency and filter noisy samples, we construct a dual network to communicate and guide each other. Experimental results show that compared with popular RL algorithms, our method demonstrates superior performance in success rate, completion time, safety, and data efficiency. Haonan Luo 0002, Honglin Dong, Jianfeng Lu 0003 |
ICME | 4 |
| 2025 | Gradient-Based Adversarial Attacks on Deep LiDAR OdometryabstractAdversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attacks on consecutive-frame tasks, i.e. the LiDAR odometry. In this paper, we propose a gradient optimization-based adversarial attack towards deep LiDAR odometry networks. To generate point clouds consistent with real-world scenarios, we constrain adversarial points within the range of a small object, e.g. a traffic cone, and render new points to simulate real LiDAR measurements. By incorporating such adversarial points in consecutive frames, we demonstrate a significant decrease in pose estimation accuracy of current popular LiDAR odometry networks. In addition, we also evaluate traditional geometric odometry approaches and report their robustness against adversarial points. Extensive experiments on the KITTI and Waymo datasets illustrate the effectiveness of the proposed attack method and the vulnerability of deep LiDAR odometry networks against adversarial points. Zhenbo Song, Xuanzhu Chen, Zhenyuan Zhang 0001, Kaihao Zhang, Jianfeng Lu 0003 |
ICRA | 5 |
| 2025 | Text as Any-Modality for Zero-Shot Classification by Consistent Prompt TuningabstractThe integration of prompt tuning with multimodal learning has shown significant generalization abilities for various downstream tasks. Despite advancements, existing methods heavily depend on massive modality-specific labeled data (e.g., video, audio, and image), or are customized for a single modality. In this study, we present Text as Any-Modality by Consistent Prompt Tuning (TaAM-CPT), a scalable approach for constructing a general representation model toward unlimited modalities using solely text data. TaAM-CPT comprises modality prompt pools, text construction, and modality-aligned text encoders from pre-trained models, which allows for extending new modalities by simply adding prompt pools and modality-aligned text encoders. To harmonize the learning across different modalities, TaAM-CPT designs intra- and inter-modal learning objectives, which can capture category details within modalities while maintaining semantic consistency across different modalities. Benefiting from its scalable architecture and pre-trained models, TaAM-CPT can be seamlessly extended to accommodate unlimited modalities. Remarkably, without any modality-specific labeled data, TaAM-CPT achieves leading results on diverse datasets spanning various modalities, including video classification, image classification, and audio classification. The code is available at https://github.com/Jinx630/TaAM-CPT. Feng Yu 0030, Yang Yang 0074, Jianfeng Lu 0003 |
ACM Multimedia | 4 |
| 2025 | UV-GA: UV-Guided Gaussian Avatar Reconstruction from Single Image
Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003 |
PRCV (10) | 4 |
| 2025 | V2DGS:Visual Voxel Map-Based 2D Gaussian Splatting for Accurate Outdoor Reconstruction
Zhenbo Song, Qigeng Duan, Benyun Zhao, Jianfeng Lu 0003 |
PRCV (10) | 5 |
| 2025 | Dynamic brain effective connectivity network for identifying neurological disorders
Saqib Mamoon, Zhengwang Xia, Amani Alfakih, Jianfeng Lu 0003 |
Appl. Intell. | 4 |
| 2025 | Multi-domain universal representation learning for hyperspectral object tracking
Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jing Wang 0062, Diqi Chen, Jun Zhou 0001, Yuntao Qian |
Pattern Recognit. | 3 |
| 2025 | Spatial-Spectral-Temporal Correlation Filter for Hyperspectral Object TrackingabstractObject tracking with hyperspectral videos (HSVs) offers significant advantages due to the captured spectral fingerprint information, which provides detailed physical material characteristics. While correlation filter (CF)-based tracking methods align well with the high-dimensional nature of HSVs, they often fall short of fully utilizing the spatial–spectral–temporal structure inherent in these data. In this article, we introduce a spatial–spectral–temporal CF (SSTCF) framework to address these limitations. SSTCF employs the spatial-spectral histogram of gradients and fractional abundances as features to characterize the spatial-spectral structure of the object. A low-rank constraint is integrated into the CF framework to enhance the global spectral semantic dependencies among learned filters. In addition, a temporal constraint is incorporated to ensure filter consistency across consecutive frames, further improving tracking continuity between nearby frames. Extensive experiments demonstrate that our SSTCF tracker achieves more accurate and stable performance. The source code will be publicly available athttps://github.com/bearshng/SSTCF Fengchao Xiong, Yongle Sun, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial AttacksabstractImage processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios. Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Enhanced Fine-Grained Motion Diffusion for Text-Driven Human Motion SynthesisabstractThe emergence of text-driven motion synthesis technique provides animators with great potential to create efficiently. However, in most cases, textual expressions only contain general and qualitative motion descriptions, while lack fine depiction and sufficient intensity, leading to the synthesized motions that either (a) semantically compliant but uncontrollable over specific pose details, or (b) even deviates from the provided descriptions, bringing animators with undesired cases. In this paper, we propose DiffKFC, a conditional diffusion model for text-driven motion synthesis with KeyFrames Collaborated, enabling realistic generation with collaborative and efficient dual-level control: coarse guidance at semantic level, with only few keyframes for direct and fine-grained depiction down to body posture level. Unlike existing inference-editing diffusion models that incorporate conditions without training, our conditional diffusion model is explicitly trained and can fully exploit correlations among texts, keyframes and the diffused target frames. To preserve the control capability of discrete and sparse keyframes, we customize dilated mask attention modules where only partial valid tokens participate in local-to-global attention, indicated by the dilated keyframe mask. Additionally, we develop a simple yet effective smoothness prior, which steers the generated frames towards seamless keyframe transitions at inference. Extensive experiments show that our model not only achieves state-of-the-art performance in terms of semantic fidelity, but more importantly, is able to satisfy animator requirements through fine-grained guidance without tedious labor. Dong Wei 0007, Xiaoning Sun, Huaijiang Sun, Shengxiang Hu 0001, Bin Li 0084, Jianfeng Lu 0003 |
AAAI | 7 |
| 2024 | Universal Video Face Restoration Method Based on Vision-Language Model
Yipiao Xu, Zhenbo Song, Jianfeng Lu 0003 |
ACML | 3 |
| 2024 | Fast Adaptation for Human Pose Estimation via Meta-OptimizationabstractDomain shift is a challenge for supervised human pose estimation, where the source data and target data come from different distributions. This is why pose estimation methods generally perform worse on the test set than on the training set. Recently, test-time adaptation has proven to be an effective way to deal with domain shift in human pose estimation. Although the performance on the target domain has been improved, existing methods require a large number of weight updates for convergence, which is time-consuming and brings catastrophic forgetting. To solve these issues, we propose a meta-auxiliary learning method to achieve fast adaptation for domain shift during inference. Specifically, we take human pose estimation as the supervised primary task, and propose body-specific image inpainting as a self-supervised auxiliary task. First, we Jointly train the primary and auxiliary tasks to get a pre-trained model on the source domain. Then, meta-training correlates the performance of the two tasks to learn a good weight initialization. Finally, meta-testing adapts the meta-learned model to the target data through self-supervised learning. Benefiting from the meta-learning paradigm, the proposed method enables fast adaptation to the target domain while preserving the source domain knowledge. The carefully designed auxiliary task better pays attention to human-related semantics in a single image. Extensive experiments demonstrate the effectiveness of our test-time fast adaptation. Shengxiang Hu 0001, Huaijiang Sun, Bin Li 0084, Dong Wei 0007, Jianfeng Lu 0003 |
CVPR | 6 |
| 2024 | MoML: Online Meta Adaptation for 3D Human Motion PredictionabstractIn the academic field, the research on human motion pre-diction tasks mainly focuses on exploiting the observed in-formation to forecast human movements accurately in the near future horizon. However, a significant gap appears when it comes to the application field, as current models are all trained offline, with fixed parameters that are inher-ently suboptimal to handle the complex yet ever-changing nature of human behaviors. To bridge this gap, in this pa-per, we introduce the task of online meta adaptation for hu-man motion prediction, based on the insight that finding “smart weights” capable of swift adjustments to suit dif-ferent motion contexts along the time is a key to improving predictive accuracy. We propose MoML, which ingeniously borrows the bilevel optimization spirit of model-agnostic meta-learning, to transform previous predictive mistakes into strong inductive biases to guide online adaptation. This is achieved by our MoAdapter blocks that can learn er-ror information by facilitating efficient adaptation via a few gradient steps, which fine-tunes our meta-learned “smart” initialization produced by the generic predictor. Considering real-time requirements in practice, we further propose Fast-MoML, a more efficient variant of MoML that features a closed-form solution instead of conventional gradient up-date. Experimental results show that our approach can ef-fectively bring many existing offline motion prediction mod-els online, and improves their predictive accuracy. Xiaoning Sun, Huaijiang Sun, Bin Li 0084, Dong Wei 0007, Jianfeng Lu 0003 |
CVPR | 6 |
| 2024 | Human Motion Forecasting in Dynamic Domain Shifts: A Homeostatic Continual Test-Time Adaptation Framework
Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084 |
ECCV (31) | 4 |
| 2024 | NeRM: Learning Neural Representations for High-Framerate Human Motion SynthesisabstractGenerating realistic human motions with high framerate is an underexplored task, due to the varied framerates of training data, huge memory burden brought by high framerates and slow sampling speed of generative models. Recent advances make a compromise for training by downsampling high-framerate details away and discarding low-framerate samples, which suffer from severe information loss and restricted-framerate generation. In this paper, we found that the recent emerging paradigm of Implicit Neural Representations (INRs) that encode a signal into a continuous function can effectively tackle this challenging problem. To this end, we introduce NeRM, a generative model capable of taking advantage of varied-size data and capturing variational distribution of motions for high-framerate motion synthesis. By optimizing latent representation and a auto-decoder conditioned on temporal coordinates, NeRM learns continuous motion fields of sampled motion clips that ingeniously avoid explicit modeling of raw varied-size motions. This expressive latent representation is then used to learn a diffusion model that enables both unconditional and conditional generation of human motions. We demonstrate that our approach achieves competitive results with state-of-the-art methods, and can generate arbitrary framerate motions. Additionally, we show that NeRM is not only memory-friendly, but also highly efficient even when generating high-framerate motions. Dong Wei 0007, Huaijiang Sun, Bin Li 0084, Xiaoning Sun, Shengxiang Hu 0001, Jianfeng Lu 0003 |
ICLR | 7 |
| 2024 | Semantic-Aware Alignment Network for Cross-Resolution Change DetectionabstractCross-resolution change detection (CRCD) is of significant practical importance in disaster assessment, rapid urban transitions, and various applications. Conventional change detection methods are primarily tailored for bitemporal images with consistent spatial resolution, rendering them unsuitable for direct application to CRCD tasks. This limitation stems from the substantial scale differences and pixel-wise misalignment prevalent in cross-resolution remote sensing images. In response to these challenges, we introduce a semantic-aware alignment network (SA-Net). SA-Net utilizes cross-attention to map bitemporal images into a shared semantic space, effectively alleviating the difficulties of the subsequent alignment arising from semantic mismatches. Furthermore, a joint transformer featuring an encoder-decoder architecture is employed to extract global information and learn the geometric parameters for spatial alignment between bitemporal images. Experimental evaluations on two real-collected datasets, HTCD and MRCDD, showcase the superior performance of our proposed SA-Net in CRCD tasks. Fengchao Xiong, Jianfeng Lu 0003, Minchao Ye, Jun Zhou 0001, Yuntao Qian |
IGARSS | 3 |
| 2024 | TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
Qing-Yuan Jiang, Yang Yang 0074, Yi-Feng Wu, Jianfeng Lu 0003 |
IJCAI | 6 |
| 2024 | Customized Relationship Graph Neural Network for Brain Disorder Identification
Zhengwang Xia, Tao Zhou 0002, Jianfeng Lu 0003 |
MICCAI (2) | 5 |
| 2024 | On the Robustness of Deep Face Inpainting: An Adversarial Perspective
Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003 |
MMAsia | 4 |
| 2024 | Vehicle Re-identification with a Pose-Aware Discriminative Part Learning Model
Jianfeng Lu 0003, Minxian Li, Gang Ren 0005, Jingfeng Ma |
PRCV (13) | 2 |
| 2024 | AS-FIBA: Adaptive Selective Frequency-Injection for Backdoor Attack on Deep Face RestorationabstractDeep learning-based face restoration models, increasingly prevalent in smart devices, have become targets for sophisticated backdoor attacks. Through subtle trigger injection into input face images, these attacks can lead to unexpected restoration outcomes. Unlike conventional methods focused on classification tasks, our approach introduces a unique degradation objective tailored for attacking restoration models. Moreover, we propose the Adaptive Selective Frequency Injection Backdoor Attack (AS-FIBA) framework, employing a neural network for input-specific trigger generation in the frequency domain, seamlessly blending triggers with benign images. This results in imperceptible yet effective attacks, guiding restoration predictions towards subtly degraded outputs rather than conspicuous targets. Extensive experiments demonstrate the efficacy of the degradation objective on state-of-the-art face restoration models. Additionally, it is notable that AS-FIBA can insert effective backdoors that are more imperceptible than existing backdoor attack methods, including WaNet, ISSBA, and FIBA. Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003 |
TrustCom | 4 |
| 2024 | Inferring brain causal and temporal-lag networks for recognizing abnormal patterns of dementia
Zhengwang Xia, Tao Zhou 0002, Saqib Mamoon, Jianfeng Lu 0003 |
Medical Image Anal. | 4 |
| 2024 | Unified Privileged Knowledge Distillation Framework for Human Motion PredictionabstractPrevious works on human motion prediction follow the pattern of building an extrapolation mapping between the sequence observed and the one to be predicted. However, the inherent difficulty of time-series extrapolation and complexity of human motion data still result in many failure cases. In this paper, we explore a longer horizon of sequence with more poses following behind, which breaks the limit in extrapolation problems that data/information on the other side of the predictive target is completely unknown. As these poses are unavailable for testing, we regard them as a privileged sequence, and propose a Two-stage Privileged Knowledge Distillation framework that incorporates privileged information in the forecasting process while avoiding direct use of it. Specifically, in the first stage, both the observed and privileged sequence are encoded for interpolation, with Privileged-sequence-Encoder (Priv-Encoder) learning privileged knowledge (PK) simultaneously. Then, in the second stage where privileged sequence is not observable, a novel PK-Simulator distills PK by approximating the behavior of Priv-Encoder, but only taking as input the observed sequence, to enable a PK-aware prediction pattern. Moreover, we present a One-stage version of this framework, using Shared Encoder that integrates the observation encoding in both interpolation and prediction branches to realize parallel training, which helps produce the most conducive PK to prediction pipeline. Experimental results show that our frameworks are model-agnostic, and can be applied to existing motion prediction models with encoder-decoder architecture to achieve improved performance. Xiaoning Sun, Huaijiang Sun, Dong Wei 0007, Jin Wang 0005, Bin Li 0084, Jianfeng Lu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | SSUMamba: Spatial-Spectral Selective State Space Model for Hyperspectral Image DenoisingabstractDenoising is a crucial preprocessing step for hyperspectral images (HSIs) due to noise arising from intraimaging mechanisms and environmental factors. Long-range spatial-spectral correlation modeling is beneficial for HSI denoising but often comes with high complexity. Based on the state space model (SSM), Mamba is known for its remarkable long-range dependency modeling capabilities and computational efficiency. Building on this, we introduce a memory-efficient spatial-spectral UMamba (SSUMamba) for HSI denoising, with the spatial-spectral continuous scan (SSCS) Mamba being the core component. SSCS Mamba alternates the row, column, and band in six different orders to generate the sequence and uses the bidirectional SSM to exploit long-range spatial-spectral dependencies. In each order, the images are rearranged between adjacent scans to ensure spatial-spectral continuity. In addition, 3-D convolutions are embedded into the SSCS Mamba to enhance local spatial-spectral modeling. Experiments demonstrate that SSUMamba achieves superior denoising results with lower memory consumption per batch compared with transformer-based methods. The source code is available at:https://github.com/lronkitty/SSUMamba. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hyperspectral Image Denoising via Spatial-Spectral Recurrent TransformerabstractHyperspectral images (HSIs) often suffer from noise arising from both intra-imaging mechanisms and environmental factors. Leveraging domain knowledge specific to HSIs, such as global spectral correlation (GSC) and non-local spatial self-similarity (NSS), is crucial for effective denoising. Existing methods tend to independently utilize each of these knowledge components with multiple blocks, overlooking the inherent 3D nature of HSIs where domain knowledge is strongly interlinked, resulting in suboptimal performance. To address this challenge, this paper introduces a spatial-spectral recurrent transformer U-Net (SSRT-UNet) for HSI denoising. The proposed SSRT-UNet integrates NSS and GSC properties within a single SSRT block. This block consists of a spatial branch and a spectral branch. The spectral branch employs a combination of transformer and recurrent neural network to perform recurrent computations across bands, allowing for GSC exploitation beyond a fixed number of bands. Concurrently, the spatial branch encodes NSS for each band by sharingkeysandvalueswith the spectral branch under the guidance of GSC. The interaction between the two branches enables the joint utilization of NSS and GSC, avoiding their independent treatment. Experimental results demonstrate that our method outperforms several alternative approaches. The source code will be available at https://github.com/lronkitty/SSRT. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Jiantao Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Material-Guided Multiview Fusion Network for Hyperspectral Object TrackingabstractHyperspectral videos (HSVs) have more potential in object tracking than color videos thanks to their material identification ability. Nevertheless, previous works have not fully explored the benefits of the material information, resulting in limited representation ability and tracking accuracy. To address this issue, this paper introduces a material-guided multi-view fusion network for improved tracking. Specifically, we combine false-color information, hyperspectral information, and material information obtained by hyperspectral unmixing to provide a rich multi-view representation of the object. Cross-material attention is employed to capture the interaction among materials, enabling the network to focus on the most relevant materials for the target. Furthermore, leveraging the discriminative ability of material view, a novel material-guided multi-view fusion module is proposed to capture both intra-view and cross-view long-range spatial dependencies for effective feature aggregation. Thanks to the enhanced representation ability of each view and the integration of the complementary advantages of all views, our network is more capable of suppressing the tracking drift in various challenging scenes and achieving accurate object localization. Extensive experiments show that our tracker achieves state-of-the-art tracking performance. The source code will be available at https://github.com/hscv/MMF-Net. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Wavelet Siamese Network With Semi-Supervised Domain Adaptation for Remote Sensing Image Change DetectionabstractChange detection is a crucial technique in remote sensing image analysis and faces challenges, such as background complexity and appearance shift, resulting in incomplete change boundaries and pseudochanges. This article introduces a novel wavelet Siamese network with semi-supervised domain adaptation (DA) to address these issues, named WS-Net++. WS-Net++ establishes spatial–frequency interactions between bitemporal images to enhance the completeness of the change boundaries. The spatial-domain interaction highlights the pixelwise differences. The frequency-domain interaction first adaptively adjusts the contributions from different frequency components based on image context. Within-frequency and between-frequency interactions are further constructed to capture the frequency-domain differences, enabling the adaptive and effective handling of both overall and subtle changes. In addition, WS-Net++ employs a semi-supervised DA strategy to mitigate the appearance shifts between bitemporal images. By categorizing regions into changed, unchanged, and regions of no interest in a semi-supervised manner, the network minimizes intraclass discrepancies within unchanged regions and maximizes interclass discrepancies between changed regions, reducing the domain gap. Experimental results on the LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that our WS-Net++ outperforms alternative methods, achieving the$F1$scores of 91.31%, 94.52%, and 79.77%, respectively. The code and models will be publicly available athttps://github.com/JiTaiTai/WS-Net_Plusfor reproducible research. Fengchao Xiong, Tianhan Li, Yi Yang 0071, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Meta-Auxiliary Learning for Adaptive Human Pose PredictionabstractPredicting high-fidelity future human poses, from a historically observed sequence, is crucial for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on external datasets and then directly apply it to all test samples, emerge as the dominant solution to solve this issue. Despite encouraging progress, they remain non-optimal, as the unique properties (e.g., motion style, rhythm) of a specific sequence cannot be adapted. More generally, once encountering out-of-distributions, the predicted poses tend to be unreliable. Motivated by this observation, we propose a novel test-time adaptation framework that leverages two self-supervised auxiliary tasks to help the primary forecasting network adapt to the test sequence. In the testing phase, our model can adjust the model parameters by several gradient updates to improve the generation quality. However, due to catastrophic forgetting, both auxiliary tasks typically have a low ability to automatically present the desired positive incentives for the final prediction performance. For this reason, we also propose a meta-auxiliary learning scheme for better adaptation. Extensive experiments show that the proposed approach achieves higher accuracy and more realistic visualization. Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084 |
AAAI | 3 |
| 2023 | Human Joint Kinematics Diffusion-Refinement for Stochastic Motion PredictionabstractStochastic human motion prediction aims to forecast multiple plausible future motions given a single pose sequence from the past. Most previous works focus on designing elaborate losses to improve the accuracy, while the diversity is typically characterized by randomly sampling a set of latent variables from the latent prior, which is then decoded into possible motions. This joint training of sampling and decoding, however, suffers from posterior collapse as the learned latent variables tend to be ignored by a strong decoder, leading to limited diversity. Alternatively, inspired by the diffusion process in nonequilibrium thermodynamics, we propose MotionDiff, a diffusion probabilistic model to treat the kinematics of human joints as heated particles, which will diffuse from original states to a noise distribution. This process not only offers a natural way to obtain the "whitened'' latents without any trainable parameters, but also introduces a new noise in each diffusion step, both of which facilitate more diverse motions. Human motion prediction is then regarded as the reverse diffusion process that converts the noise distribution into realistic future motions conditioned on the observed sequence. Specifically, MotionDiff consists of two parts: a spatial-temporal transformer-based diffusion network to generate diverse yet plausible motions, and a flexible refinement network to further enable geometric losses and align with the ground truth. Experimental results on two datasets demonstrate that our model yields the competitive performance in terms of both diversity and accuracy. Dong Wei 0007, Huaijiang Sun, Bin Li 0084, Jianfeng Lu 0003, Xiaoning Sun, Shengxiang Hu 0001 |
AAAI | 4 |
| 2023 | Robust Single Image Reflection Removal Against Adversarial AttacksabstractThis paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustness study, we first conduct diverse adversarial attacks specifically for the SIRR problem, i.e. towards different attacking targets and regions. Then we propose a robust SIRR model, which integrates the cross-scale attention module, the multi-scale fusion module, and the adversarial image discriminator. By exploiting the multi-scale mechanism, the model narrows the gap between features from clean and adversarial images. The image discriminator adaptively distinguishes clean or noisy inputs, and thus further gains reliable robustness. Extensive experiments on Nature, SIR2, and Real datasets demonstrate that our model remarkably improves the robustness of SIRR across disparate scenes. Zhenbo Song, Zhenyuan Zhang 0001, Kaihao Zhang, Wenhan Luo, Zhaoxin Fan, Wenqi Ren, Jianfeng Lu 0003 |
CVPR | 7 |
| 2023 | DeFeeNet: Consecutive 3D Human Motion Prediction with Deviation FeedbackabstractLet us rethink the real-world scenarios that require human motion prediction techniques, such as human-robot collaboration. Current works simplify the task of predicting human motions into a one-off process of forecasting a short future sequence (usually no longer than 1 second) based on a historical observed one. However, such simplification may fail to meet practical needs due to the neglect of the fact that motion prediction in real applications is not an isolated “observe then predict” unit, but a consecutive process composed of many rounds of such unit, semi-overlapped along the entire sequence. As time goes on, the predicted part of previous round has its corresponding ground truth observable in the new round, but their deviation in-between is neither exploited nor able to be captured by existing isolated learning fashion. In this paper, we propose DeFeeNet, a simple yet effective network that can be added on existing one-off prediction models to realize deviation perception and feedback when applied to consecutive motion prediction task. At each prediction round, the deviation generated by previous unit is first encoded by our DeFeeNet, and then incorporated into the existing predictor to enable a deviation-aware prediction manner, which, for the first time, allows for information transmit across adjacent prediction units. We design two versions of DeFeeNet as MLP-based and GRU-based, respectively. On Human3.6M and more complicated BABEL, experimental results indicate that our proposed network improves consecutive human motion prediction performance regardless of the basic model. Xiaoning Sun, Huaijiang Sun, Bin Li 0084, Dong Wei 0007, Jianfeng Lu 0003 |
CVPR | 6 |
| 2023 | Test-time Personalizable Forecasting of 3D Human PosesabstractCurrent motion forecasting approaches typically train a deep end-to-end model from the source domain data, and then apply it directly to target subjects. Despite promising results, they remain non-optimal, due to privacy considerations, the test person and his/her natural properties (e.g., behavioral trait) are typically unseen in training. In this case, the source pre-trained model has a low ability to adapt to these out-of-source characteristics, resulting in an unreliable prediction. To tackle this issue, we propose a novel helper-predictor test-time personalization approach (H/P-TTP), which allows for a generalizable representation of out-of-source subjects to gain more realistic predictions. Concretely, the helper is preceded by explicit and implicit augmenters, where the former yields noisy sequences to improve robustness, while the latter is to generate novel-domain data with an adversarial learning paradigm. Then, the domain-generalizable learning is achieved where the helper can extract cross-subject invariant-knowledge to update the predictor. At test time, given a new person, the predictor is able to be further optimized to empower personalized capabilities to the specific properties. Extensive experiments show that with H/P-TTP, the existing models are significantly improved for various unseen subjects. The project page is available at https://sites.google.com/view/hp-ttp. Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084, Hongwei Yi |
ICCV | 3 |
| 2023 | Incorporating Global Correlation and Local Aggregation for Efficient Visual Localization
Jianfeng Lu 0003, Zhenbo Song, Xuanzhu Chen |
ICIG (2) | 2 |
| 2023 | Learning Dense Flow Field for Highly-accurate Cross-view Camera LocalizationabstractThis paper addresses the problem of estimating the 3-DoF camera pose for a ground-level image with respect to a satellite image that encompasses the local surroundings. We propose a novel end-to-end approach that leverages the learning of dense pixel-wise flow fields in pairs of ground and satellite images to calculate the camera pose. Our approach differs from existing methods by constructing the feature metric at the pixel level, enabling full-image supervision for learning distinctive geometric configurations and visual appearances across views. Specifically, our method employs two distinct convolution networks for ground and satellite feature extraction. Then, we project the ground feature map to the bird's eye view (BEV) using a fixed camera height assumption to achieve preliminary geometric alignment. To further establish the content association between the BEV and satellite features, we introduce a residual convolution block to refine the projected BEV feature. Optical flow estimation is performed on the refined BEV feature map and the satellite feature map using flow decoder networks based on RAFT. After obtaining dense flow correspondences, we apply the least square method to filter matching inliers and regress the ground camera pose. Extensive experiments demonstrate significant improvements compared to state-of-the-art methods. Notably, our approach reduces the median localization error by 89\%, 19\%, 80\%, and 35\% on the KITTI, Ford multi-AV, VIGOR, and Oxford RobotCar datasets, respectively. Zhenbo Song, Xianghui Ze, Jianfeng Lu 0003, Yujiao Shi 0002 |
NeurIPS | 3 |
| 2023 | FPGA-oriented lightweight multi-modal free-space detection networkabstractFor autonomous vehicles, free-space detection is an essential part of visual perception. With the development of multi-modal convolutional neural networks (CNNs) in recent years, the performance of driving scene semantic segmentation algorithms has been dramatically improved. Therefore most free-space detection algorithms are developed based on multiple sensors. However, multi-modal CNNs have high data throughput and contain a large number of computationally intensive convolution calculations, limiting their feasibility for real-time applications. Field Programmable Gate Arrays (FPGAs) provide a unique combination of flexibility, performance, and low power for these problems to accommodate multi-modal data and the computational acceleration of different compression algorithms. Network lightweight methods offer great assurance for facilitating the deployment of CNNs on such resource-constrained devices. In this paper, we propose a network lightweight method for a multi-modal free-space detection algorithm. We first propose an FPGA-friendly multi-modal free-space detection lightweight network. It comprises operators that FPGA prefers and achieves a 95.54% MaxF score on the test set of KITTI-Road free-space detection tasks and 81 ms runtime when running on 700 W GPU devices. Then we present a pruning approach for this network to reduce the number of parameters in case the complete model exceeds the FPGA chip memory. The pruning is in two parts. For the feature extractors, we propose a data-dependent filter pruner according to the principle that the low-rank feature map contains less information. To not compromise the integrity of the multi-modal information, the pruner is independent for each modality. For the segmentation decoder, we apply a channel pruning approach to remove redundant parameters. Finally, we implement our designs on an FPGA board using 8-bit quantisation, and the accelerator achieves outstanding performance. A real-time application of scene segmentation on KITTI-Road is used to evaluate our algorithm, and the model achieves a 94.39% MaxF score and minimum 14 ms runtime on 20W FPGA devices. Feiyi Fang, Junzhu Mao, Jianfeng Lu 0003 |
Connect. Sci. | 4 |
| 2023 | Shared and individual representation learning with Feature Diversity for Deep MultiView Clustering
Sheng Wang 0015, Liyong Chen, Ning Zheng 0003, Furong Peng, Jianfeng Lu 0003 |
Inf. Sci. | 6 |
| 2023 | Deep semantic-aware remote sensing image deblurring
Zhenbo Song, Zhenyuan Zhang 0001, Feiyi Fang, Zhaoxin Fan, Jianfeng Lu 0003 |
Signal Process. | 5 |
| 2023 | Multitask Sparse Representation Model-Inspired Network for Hyperspectral Image DenoisingabstractHyperspectral images (HSIs) are prone to noise because of the imaging mechanism and environment. This paper proposes a multitask sparse representation (SR) model inspired neural network for HSI denoising. Unlike other deep learning-based methods, our network is interpretable, whose network architecture is induced by unfolding the iterative optimization of a multitask sparse representation model. On the one hand, the model globally represents the common structure among bands, such as image edges, with the shared sparse coefficients. On the other hand, it separately encodes the unique structure of individual bands with unshared ones to capture image details. Accordingly, our network has three modules: the shared SR module, the unshared SR module, and the image reconstruction (IR) module. All the modules are connected with a specific operation of the iterative optimization algorithm, equipping the network with clear physical interpretation. Experimental results on both synthetic and real-world datasets demonstrate the superior performance of our method, visually and quantitatively. The codes will be publicly available at https://github.com/bearshng/mtsrnn for reproducible research. Fengchao Xiong, Jiantao Zhou 0001, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Deep Parameterized Neural Networks for Hyperspectral Image DenoisingabstractSparse representation (SR)-based hyperspectral image (HSI) denoising methods normally average the local denoising results of multiple overlapped cubes to recover the whole HSI. Though interpretable, they rely on cumbersome hyperparameter settings and ignore the relationship between overlapped cubes, leading to poor denoising performance. This article combines SR and convolutional neural networks and introduces a deep parameterized sparse neural network (DPNet-S) to address the above issues. DPNet-S parameterizes the SR-based HSI denoising model with two modules: 1) sparse optimizer to extract sparse feature maps from noisy HSIs via recurrent usage of convolution, deconvolution, and soft shrinkage operations; and 2) image reconstructor to recover the denoised HSI from its sparse feature maps via deconvolution operations. We further replace the soft shrinkage operator with U-Net architecture to account for general HSI priors and more effectively capture the complex structures of HSIs, resulting in DPNet-U. Both networks directly learn the parameters from data and perform denoising on the whole HSI, which overcomes the limitations of SR-based methods. Moreover, our networks are generated from the denoising model and optimization procedures, thus leveraging the knowledge embedded and relying less on the number of training samples. Extensive experiments on both synthetic and real-world HSIs show that our DPNet-S and DPNet-U achieve remarkable results when compared with state-of-the-art methods. The codes will be publicly available athttps://github.com/bearshng/dpnetsfor reproducible research. Fengchao Xiong, Jun Zhou 0001, Jiantao Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Learning a Deep Ensemble Network With Band Importance for Hyperspectral Object TrackingabstractAttributing to material identification ability powered by a large number of spectral bands, hyperspectral videos (HSVs) have great potential for object tracking. Most hyperspectral trackers employ manually designed features rather than deeply learned features to describe objects due to limited available HSVs for training, leaving a huge gap to improve the tracking performance. In this paper, we propose an end-to-end deep ensemble network (SEE-Net) to address this challenge. Specifically, we first establish a spectral self-expressive model to learn the band correlation, indicating the importance of a single band in forming hyperspectral data. We parameterize the optimization of the model with a spectral self-expressive module to learn the nonlinear mapping from input hyperspectral frames to band importance. In this way, the prior knowledge of bands is transformed into a learnable network architecture, which has high computational efficiency and can fast adapt to the changes of target appearance because of no iterative optimization. The band importance is further exploited from two aspects. On the one hand, according to the band importance, each frame of HSVs is divided into several three-channel false-color images which are then used for deep feature extraction and location. On the other hand, based on the band importance, the importance of each false-color image is computed, which is then used to assemble the tracking results from individual false-color images. In this way, the unreliable tracking caused by false-color images of low importance can be suppressed to a large extent. Extensive experimental results show that SEE-Net performs favorably against the state-of-the-art approaches. The source code will be available at https://github.com/hscv/SEE-Net. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Image Process. | 4 |
| 2023 | A Structure-Guided Effective and Temporal-Lag Connectivity Network for Revealing Brain Disorder MechanismsabstractBrain network provides important insights for the diagnosis of many brain disorders, and how to effectively model the brain structure has become one of the core issues in the domain of brain imaging analysis. Recently, various computational methods have been proposed to estimate the causal relationship (i.e., effective connectivity) between brain regions. Compared with traditional correlation-based methods, effective connectivity can provide the direction of information flow, which may provide additional information for the diagnosis of brain diseases. However, existing methods either ignore the fact that there is a temporal-lag in the information transmission across brain regions, or simply set the temporal-lag value between all brain regions to a fixed value. To overcome these issues, we design an effective temporal-lag neural network (termed ETLN) to simultaneously infer the causal relationships and the temporal-lag values between brain regions, which can be trained in an end-to-end manner. In addition, we also introduce three mechanisms to better guide the modeling of brain networks. The evaluation results on the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate the effectiveness of the proposed method. Zhengwang Xia, Tao Zhou 0002, Saqib Mamoon, Amani Alfakih, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Overlooked Poses Actually Make Sense: Distilling Privileged Knowledge for Human Motion Prediction
Xiaoning Sun, Qiongjie Cui, Huaijiang Sun, Bin Li 0084, Jianfeng Lu 0003 |
ECCV (5) | 6 |
| 2022 | Material-Guided Siamese Fusion Network for Hyperspectral Object TrackingabstractHyperspectral videos (HSVs) have more potential in target tracking than color videos thanks to the material identification capability provided by abundant spectral bands. Due to limited HSVs for training, most current hyperspectral trackers are based on hand-crafted features rather than deeply learned ones, resulting in poor tracking performance. This paper introduces a material-guided Siamese fusion network (SiamF) for hyperspectral object tracking to make up this gap. Belonging to the Siamese tracker family and SiamF aims to model the appearance of hyperspectral objects using backbone networks trained on color images. Specifically, SiamF splits each hyperspectral frame into multiple groups of false-color images according to their band importance. Then SiamF employs a hyperspectral feature fusion (HFF) module with a dense connection architecture to integrate the extracted features from different layers and band groups, producing a multi-scale multilevel spatial-spectral representation of the targets. Instead of direct addition or concatenation, HFF employs global-local channel attention for feature fusion, so that yielded features capture the global and local structure of a specific object. Moreover, online spatial and material classifiers are developed to inject spatial and material appearance changes information into SiamF for adaptively online tracking. Experimental results demonstrate our tracker outperforms alternative methods. Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
ICASSP | 3 |
| 2022 | Multitask Sparse Neural Network for Hyperspectral Image DenoisingabstractData-driven deep learning (DL)-based methods directly learn the nonlinear mapping between noisy hyperspectral images (HSIs) and corresponding clean ones. However, DLbased methods neglect the prior knowledge of HSIs embodied by physical models. Consequently, they require complex network architectures and a large number of training samples. To address the above issues, this paper introduces a multitask sparse neural network (MTSNN) which bridges the sparsity prior of HSIs with data-driven deep learning for HSI denoising. Specifically, we first build a multitask sparse (MTS) denoising model which shares sparse coefficients among bands to exploit the spectral-spatial correlation and learns a dictionary for each band to depict the distinct spatial structure among bands. The iterative optimization of the MTS model is then unfolded to yield our MTSNN by introducing some learnable parameters. MTSNN is a multi-branch network. Each branch performs a single denoising task for an individual band. All branches are connected by shared coefficients, forming multitask denoising for all bands. The hybrid advantages of the MTS model and data-driven learning equip MTSNN with strong denoising ability, preferable learning capability, superior interpretability, and higher generalization capacity. Experimental results demonstrate that our method achieves state-of-the-art denoising performance compared with several alternative approaches. Fengchao Xiong, Minchao Ye, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
ICASSP | 4 |
| 2022 | Ques-to-Visual Guided Visual Question AnsweringabstractVisual question answering (VQA) answers text-based questions about images. The difficulty of VQA lies in the accurate localization of the region related to the question. In this paper, we introduce the ques-to-visual (q2v) feature as the additional input of VQA to tackle this problem. The q2v feature is generated according to the semantics of the question, containing visual semantics that is helpful to locate the region related to the question. We then use self-attention to model the intra-relationship in each modality to enhance different features, i.e., q2v, image, and text features. The enhanced features are then fused by spatial guided-attention and multi-scale channel attention modules for the answer prediction. Experimental results on the VQA2.0 benchmark dataset show that our method achieves higher performance when compared with other methods. Jianfeng Lu 0003, Zhuanfeng Li, Fengchao Xiong |
ICIP | 2 |
| 2022 | PilotAttnNet: Multi-modal Attention Network for End-to-End Steering Control
Jincan Zhang, Zhenbo Song, Jianfeng Lu 0003, Xingwei Qu, Zhaoxin Fan |
PRCV (3) | 3 |
| 2022 | Consensus graph learning for auto-weighted multi-view projection clustering
Xiaoshuang Sang, Jianfeng Lu 0003 |
Inf. Sci. | 2 |
| 2022 | Nonlocal Spatial-Spectral Neural Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an essential preprocessing step to improve the quality of HSIs. The difficulty of HSI denoising lies in effectively modeling the intrinsic characteristics of HSIs, such as spatial-spectral correlation, global spectral correlation, and nonlocal spatial correlation. This paper introduces a nonlocal spatial-spectral neural network (NSSNN) for HSI denoising by considering the above three factors in a unified network. More specifically, NSSNN is based on the residual U-Net and embedded with the introduced spatial-spectral recurrent (SSR) blocks and nonlocal self-similarity (NSS) blocks. The SSR block comprises 3D convolutions, one light recurrence, and one highway network. 3D convolution helps exploit the spatial-spectral correlation. The light recurrence and highway network make up the recurrent computation component and refined component, respectively, to model the global spectral correlation. NSS block is based on crisscross attention and can exploit the long-range spatial contexts effectively and efficiently. Attributing to effective modeling of the spatial-spectral correlation, the global spectral correlation, and the nonlocal spatial correlation, our NSSNN has a strong denoising ability. Extensive experiments show the superior denoising effectiveness of our method on synthetic and real-world datasets when compared to alternative methods. The source code will be available at https://github.com/lronkitty/NSSNN. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | MAC-Net: Model-Aided Nonlocal Neural Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an ill-posed inverse problem. The underlying physical model is always important to tackle this problem, which is unfortunately ignored by most of the current deep learning (DL)-based methods, producing poor denoising performance. To address this issue, this article introduces an end-to-end model-aided nonlocal neural network (MAC-Net) which simultaneously takes the spectral low-rank model and spatial deep prior into account for HSI noise reduction. Specifically, motivated by the success of the spectral low-rank model in depicting the strong spectral correlations and the nonlocal similarity prior in capturing spatial long-range dependencies, we first build a spectral low-rank model and then integrate a nonlocal U-Net into the model. In this way, we obtain a hybrid model-based and DL-based HSI denoising method where the spatial local and nonlocal multi-scale and spectral low-rank structures are effectively exploited. After that, we cast the optimization and denoising procedure of the hybrid method as a forward process of a neural network and introduce a set of learnable modules to yield our MAC-Net. Compared with traditional model-based methods, our MAC-Net overcomes the difficulties of accurate modeling, thanks to the strong learning and representation ability of DL. Unlike most “black-box” DL-based methods, the spectral low-rank model is beneficial to increase the generalization ability of the network and decrease the requirement of training samples. Experimental results on the natural and remote-sensing HSIs show that MAC-Net achieves state-of-the-art performance over both model-based and DL-based methods. The source code and data of this article will be made publicly available athttps://github.com/bearshng/mac-netfor reproducible research. Fengchao Xiong, Jun Zhou 0001, Qinling Zhao, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SNMF-Net: Learning a Deep Alternating Neural Network for Hyperspectral UnmixingabstractHyperspectral unmixing is recognized as an important tool to learn the constituent materials and corresponding distribution in a scene. The physical spectral mixture model is always important to tackle this problem because of its highly ill-posed nature. In this article, we introduce a linear spectral mixture model (LMM)-based end-to-end deep neural network named SNMF-Net for hyperspectral unmixing. SNMF-Net shares an alternating architecture and benefits from both model-based methods and learning-based methods. On the one hand, SNMF-Net is of high physical interpretability as it is built by unrolling$L_{p}$sparsity constrained nonnegative matrix factorization ($L_{p}$-NMF) model belonging to LMM families. On the other hand, all the parameters and submodules of SNMF-Net can be seamlessly linked with the alternating optimization algorithm of$L_{p}$-NMF and unmixing problem. This enables us to reasonably integrate the prior knowledge on unmixing, the optimization algorithm, and the sparse representation theory into the network for robust learning, so as to improve unmixing. Experimental results on the synthetic and real-world data show the advantages of the proposed SNMF-Net over many state-of-the-art methods. Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SMDS-Net: Model Guided Spectral-Spatial Network for Hyperspectral Image DenoisingabstractDeep learning (DL) based hyperspectral images (HSIs) denoising approaches directly learn the nonlinear mapping between noisy and clean HSI pairs. They usually do not consider the physical characteristics of HSIs. This drawback makes the models lack interpretability that is key to understanding their denoising mechanism and limits their denoising ability. In this paper, we introduce a novel model-guided interpretable network for HSI denoising to tackle this problem. Fully considering the spatial redundancy, spectral low-rankness, and spectral-spatial correlations of HSIs, we first establish a subspace-based multidimensional sparse (SMDS) model under the umbrella of tensor notation. After that, the model is unfolded into an end-to-end network named SMDS-Net, whose fundamental modules are seamlessly connected with the denoising procedure and optimization of the SMDS model. This makes SMDS-Net convey clear physical meanings, i.e., learning the low-rankness and sparsity of HSIs. Finally, all key variables are obtained by discriminative training. Extensive experiments and comprehensive analysis on synthetic and real-world HSIs confirm the strong denoising ability, strong learning capability, promising generalization ability, and high interpretability of SMDS-Net against the state-of-the-art HSI denoising methods. The source code and data of this article will be made publicly available at https://github.com/bearshng/smds-net for reproducible research. Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Jiantao Zhou 0001, Yuntao Qian |
IEEE Trans. Image Process. | 4 |
| 2022 | Self-Supervised Multi-Modal Hybrid Fusion Network for Brain Tumor SegmentationabstractAccurate medical image segmentation of brain tumors is necessary for the diagnosing, monitoring, and treating disease. In recent years, with the gradual emergence of multi-sequence magnetic resonance imaging (MRI), multi-modal MRI diagnosis has played an increasingly important role in the early diagnosis of brain tumors by providing complementary information for a given lesion. Different MRI modalities vary significantly in context, as well as in coarse and fine information. As the manual identification of brain tumors is very complicated, it usually requires the lengthy consultation of multiple experts. The automatic segmentation of brain tumors from MRI images can thus greatly reduce the workload of doctors and buy more time for treating patients. In this paper, we propose a multi-modal brain tumor segmentation framework that adopts the hybrid fusion of modality-specific features using a self-supervised learning strategy. The algorithm is based on a fully convolutional neural network. Firstly, we propose a multi-input architecture that learns independent features from multi-modal data, and can be adapted to different numbers of multi-modal inputs. Compared with single-modal multi-channel networks, our model provides a better feature extractor for segmentation tasks, which learns cross-modal information from multi-modal data. Secondly, we propose a new feature fusion scheme, named hybrid attentional fusion. This scheme enables the network to learn the hybrid representation of multiple features and capture the correlation information between them through an attention mechanism. Unlike popular methods, such as feature map concatenation, this scheme focuses on the complementarity between multi-modal data, which can significantly improve the segmentation results of specific regions. Thirdly, we propose a self-supervised learning strategy for brain tumor segmentation tasks. Our experimental results demonstrate the effectiveness of the proposed model against other state-of-the-art multi-modal medical segmentation methods. Feiyi Fang, Yazhou Yao, Tao Zhou 0002, Guosen Xie, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Self-Supervised Depth Completion From Direct Visual-LiDAR Odometry in Autonomous DrivingabstractIn this work, a simple yet effective deep neural network is proposed to generate the dense depth map of the scene by exploiting both LiDAR sparse point cloud and the monocular camera image. Specifically, a feature pyramid network is firstly employed to extract feature maps from images across time. Then the relative pose is calculated by minimizing the feature distance between aligned pixels from inter-frame feature maps. Finally, the feature maps and the relative pose are further applied to compute the feature-metric loss for training the depth completion network. The key novelty of this work lies in that a self-supervised mechanism is presented to train the depth completion network by directly using visual-LiDAR odometry between consecutive frames. Comprehensive experiments and ablation studies on benchmark dataset KITTI demonstrate the superior performance over other state-of-the-art methods in terms of pose estimation and depth completion. The detailed performance of the proposed approach (referred to asSelfCompDVLO) can be found on the KITTI depth completion benchmark. The source code, models, and data have been made available at GitHub. Zhenbo Song, Jianfeng Lu 0003, Yazhou Yao, Jian Zhang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Exploiting Web Images for Fine-Grained Visual Recognition via Dynamic Loss Correction and Global Sample SelectionabstractTo distinguish subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, acquiring manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, directly leveraging web images for learning fine-grained models becomes a natural choice. Nevertheless, due to the existence of label noise, this learning paradigm tends to have a poor performance. In this work, we propose an end-to-end approach by combining dynamic loss correction and global sample selection to alleviate the problem of label noise. Specifically, we leverage the network to predict all samples, record the predictions of recent several epochs, and calculate the uncertainly-based dynamic loss for global sample selection. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed approach. The source code of our approach has been released on the website:https://github.com/NUST-Machine-Intelligence-Laboratory/dlc. Huafeng Liu 0004, Haofeng Zhang 0001, Jianfeng Lu 0003, Zhenmin Tang |
IEEE Trans. Multim. | 3 |
| 2021 | NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral UnmixingabstractHyperspectral unmixing is an important tool to learn the material constitution and distribution of a scene. Model-based unmixing methods depend on well-designed iterative optimization algorithms, which is usually time consuming. Learning-based methods perform unmixing in a data-driven manner but heavily rely on the quality and quantity of the training samples due to the lack of physical interpretability. In this paper, we combine the advantages of both model-based and learning-based methods and propose a nonnegative matrix factorization (NMF) inspired sparse autoencoder (NMF-SAE) for hyperspectral unmixing. NMF-SAE consists of an encoder and a decoder, both of which are constructed by unrolling the iterative optimization rules of L1sparsity-constrained NMF for the linear spectral mixture model. All parameters in our method are obtained by end-to-end training in a data-driven manner. Our network is not only physically interpretable and flexible but also has higher learning capacity with fewer parameters. Experimental results on both synthetic and real-world data demonstrate that our method is capable of producing desirable unmixing results when compared against several alternative approaches. Fengchao Xiong, Jun Zhou 0001, Minchao Ye, Jianfeng Lu 0003, Yuntao Qian |
ICASSP | 4 |
| 2021 | Target-targeted Domain Adaptation for Unsupervised Semantic SegmentationabstractSemantic segmentation has attracted increasing attention due to its important role in self-driving, and it is often realized by supervised learning with large number of well labeled maps. However, the labeled images are hard to be obtained in most circumstances, and the common way for unsupervised semantic segmentation is usually implemented by transferring the knowledge from source supervised domain to target unsupervised domain. Most researches focus on encouraging target predictions to be closer to the source ones through a weight-sharing network, and achieve certain performance. However, these methods often suffer from the domain shift problem that the networks are often trained towards the source domain and lead to performance degradation. In this paper, we propose a target-targeted domain adaptation approach by focusing the training on target domain. Our model consists of two components: the Image-to-image Translation (IIT) module to translate the source image to target domain and the Target-targeted Segmentation Adaptation (TSA) module to focus the semantic segmentation on target domain. The IIT module deals with image space alignment while the TSA module bridges the domain gap at the segmentation map level. In addition, we design a closed-loop learning to promote each other by employing feedback from TSA to IIT. Extensive experiments on GTA5 and SYNTHIA to Cityscapes demonstrate the effectiveness of our method in domain adaptation of unsupervised semantic segmentation. Xiaohong Zhang 0009, Haofeng Zhang 0001, Jianfeng Lu 0003, Ling Shao 0001, Jing-Yu Yang 0001 |
ICRA | 3 |
| 2021 | Learning a Model-Based Deep Hyperspectral Denoiser from a Single Noisy Hyperspectral ImageabstractHyperspectral image (HSI) denoising is a crucial preprocessing procedure to improve the quality of HSI. Model-based methods take the degradation model and the structure of underlying clean HSI into account for denoising but require a large number of numerical iterations and exhausting parameter tuning. Deep-learning-based (DL-based) methods directly learn the nonlinear transformation of clean and noisy image HSI pairs, but rely on large-scale high-quality training samples because of its “black box” denoising mechanism. In this paper, we propose a model-based DL method for HSI denoising to combine the advantages of model-based methods and DL-based methods. Specifically, we first build a HSI denoising model based on sparse representation. Then, we unfold the iterative optimization under the framework of gradient descent with momentum to yield a Gradient Momentum Sparse Coding Network (GMSC-Net) for denoising. In order to overcome the unavailability of noisy-clean HSI pairs for training, we directly learn GMSC-Net from a single HSI. The observed noisy HSI is grouped into a number of clusters containing local cubes. The cluster centers are treated as “clean” cubes and are polluted by noises, yielding a set of “noisy-clean” pairs for training. Extensive experiments show the effectiveness of our method on both synthetic and real-world datasets. Guanyiman Fu, Fengchao Xiong, Shuyin Tao, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IGARSS | 4 |
| 2021 | Multi-Scale Spatial Transformer Network for LiDAR-Camera 3D Object DetectionabstractAccurate 3D object detection has recently aroused interest in the context of emerging autonomous driving technologies. Existing approaches predominantly use LiDAR-Camera fusion method to fulfill this challenging task, while neglecting the fact that LiDAR and camera data are spatially correlated, and cannot well retain the edge information. To solve these problems, in this paper, we propose a novel LiDAR-Camera 3D object detection method, namely the Multi-scale Spatial Transformer Network (MST-Net). The proposed method exploits an innovative spatial alignment scheme based on the projection transformer network (PTN) to mitigate the effects of the perspective view caused by sensors. In the process of generating 3D bounding boxes, the Atrous Spatial Pyramid Pool (ASPP) is applied to spatially aligned fusion features in order to preserve edge information to the greatest extent. Extensive experiments are conducted on the popular dataset KITTI, and the results can demonstrate the superiority of the proposed method. In addition, the effectiveness of these two strategies has been illustrated in ablation studies. Zhifan Wang, Xiaohong Zhang 0009, Tong Xin 0002, Haofeng Zhang 0001, Jianfeng Lu 0003 |
IJCNN | 6 |
| 2021 | MPN: Multi-scale Progressive Restoration Network for Unsupervised Defect Detection
Xuefei Liu, Kaitao Song, Jianfeng Lu 0003 |
PRCV (2) | 3 |
| 2021 | Real-Time Gait-Based Age Estimation and Gender Classification from a Single ImageabstractIn this paper, we propose a unified real-time framework for gait-based age estimation and gender classification that uses just a single image, which reduces the latency in video capturing compared with the existing methods based on a gait cycle. To cope with the problem of lacking motion information in the input single image, we first reconstruct a gait cycle of a silhouette sequence from the input image via a gait cycle reconstruction network. The reconstructed gait cycle is then fed into a state-of-the-art gait recognition network for feature representation learning, which is further used to obtain the class of the gender and the estimated probability distribution of integer age labels. Unlike the existing methods focusing on the gait sequences captured from the side view, the proposed method is applicable to the gait images from an arbitrary view with a single trained model, which is more suitable for real-world application scenarios (e.g., automatic access control). Stand-alone and client-server online systems were implemented based on the proposed method, which validates the real-time/online property in actual scenes. The experiments on the world's largest multi-view gait dataset demonstrate the effectiveness of the proposed method, which achieves performance improvement compared with the benchmark algorithms. Chi Xu 0003, Yasushi Makihara, Ruochen Liao, Hirotaka Niitsuma, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
WACV | 7 |
| 2021 | Nonconvex regularizer and latent pattern based robust regression for face recognition
Xiaoshuang Sang, Faen Zhang, Jianfeng Lu 0003 |
Inf. Sci. | 5 |
| 2021 | An adaptive two phase blind image deconvolution algorithm for an iterative regularization model
Shuyin Tao, Wende Dong, Jianfeng Lu 0003, Guili Xu |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Coarse-to-fine: A dual-view attention network for click-through rate prediction
Kaitao Song, Qingkang Huang, Faen Zhang, Jianfeng Lu 0003 |
Knowl. Based Syst. | 4 |
| 2021 | Cross-View Gait Recognition Using Pairwise Spatial Transformer NetworksabstractIn this paper, we propose a pairwise spatial transformer network (PSTN) for cross-view gait recognition, which reduces unwanted feature mis-alignment due to view differences before a recognition step for better performance. The proposed PSTN is a unified CNN architecture that consists of a pairwise spatial transformer (PST) and subsequent recognition network (RN). More specifically, given a matching pair of gait features from different source and target views, the PST estimates a non-rigid deformation field to register the features in the matching pair into their intermediate view, which mitigates distortion by registration compared with the case of direct deformation from the source view to target view. The registered matching pair is then fed into the RN to output a dissimilarity score. Although registration may reduce not only intra-subject variations but also inter-subject variations, we can still achieve a good trade-off between them using a loss function designed to optimize recognition accuracy. Experiments on three publicly available gait datasets demonstrate that the proposed method yields superior performance for both verification and identification scenarios by combining any gait recognition network benchmarks with the PST. Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | When Visual Disparity Generation Meets Semantic Segmentation: A Mutual Encouragement ApproachabstractSemantic segmentation and depth estimation play important roles in the field of autonomous driving. In recent years, the advantages of Convolutional Neural Networks (CNNs) have allowed these two topics to flourish. However, people always solve these two tasks separately and rarely solve them in a united model. In this paper, we propose a Mutual Encouragement Network (MENet), which includes a semantic segmentation branch and a disparity regression branch, and simultaneously generates semantic map and visual disparity. In the cost volume construction phase, the depth information is embedded in the semantic segmentation branch to increase contextual understanding. Similarly, the semantic information is also included in the disparity regression branch to generate more accurate disparity. Two branches mutually promote each other during training phase and inference phase. We conducted our method on the popular dataset KITTI, and the experimental results show that our method can outperform the state-of-the-art methods on both visual disparity generation and semantic segmentation. In addition, extensive ablation studies also demonstrate that the two tasks in our method can facilitate each other significantly with the proposed approach. Xiaohong Zhang 0009, Yi Chen 0023, Haofeng Zhang 0001, Shuihua Wang, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Gait Recognition from a Single Image Using a Phase-Aware Gait Cycle Reconstruction Network
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
ECCV (19) | 5 |
| 2020 | DeformGait: Gait Recognition under Posture Changes using Deformation Patterns between Gait Feature PairsabstractIn this paper, we propose a unified convolutional neural network (CNN) framework for robust gait recognition against posture changes (e.g., those induced by walking speed changes). In order to mitigate the posture changes, we first register an input matching pair of gait features with different postures by a deformable registration network, which estimates a deformation field to transform the input pair both into their intermediate posture. The pair of the registered features is then fed into a recognition network. Furthermore, ways of the deformation (i.e., deformation patterns) can differ between the same subject pairs (e.g., only posture deformation) and different subject pairs (e.g., not only posture deformation but also body shape deformation), which implies the deformation pattern can be another cue to distinguish the same subject pairs from the different subject pairs. We therefore introduce another recognition network whose input is the deformation pattern. Finally, the deformable registration network, and the two recognition networks for the registered features and the deformation patterns, constitute the whole framework, named DeformGait, and they are trained in an end-to-end manner by minimizing a loss function which is appropriately designed for each of verification and identification scenario. Experiments on the publicly available dataset containing the largest speed variations demonstrate that the proposed method achieves the state-of-the-art performance in both identification and verification scenarios. Chi Xu 0003, Daisuke Adachi, Yasushi Makihara, Yasushi Yagi, Jianfeng Lu 0003 |
IJCB | 5 |
| 2020 | BAE-Net: A Band Attention Aware Ensemble Network for Hyperspectral Object TrackingabstractHyperspectral videos contain images with a large number of light wavelength indexed bands that can facilitate material identification for object tracking. Most hyperspectral trackers use hand-crafted features rather than deep learning generated features for image representation due to limited training samples. To fill this gap, this paper introduces a band attention aware ensemble network (BAE-Net) for deep hyperspectral object tracking, which takes advantages of deep models trained on color videos for feature representation. Specifically, an autoencoder-like band attention block is introduced to learn the dependencies among bands and generate band-wise weights. Guided by these weights, hyperspectral images are then divided into a number of three-channel images. These three-channel images are fed into a deep color tracking network, producing several weak trackers. Finally, weak trackers are fused using ensemble learning for target location. Experimental results on hyperspectral datasets show the effectiveness and advantages of the proposed deep hyperspectral tracker. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jing Wang 0062, Jianfeng Lu 0003, Yuntao Qian |
ICIP | 5 |
| 2020 | Hsi Road: A Hyper Spectral Image Dataset For Road SegmentationabstractRoad segmentation is a challenging task in the field of self-driving research. This paper present a road dataset built by hyper spectral imaging (HSI) cameras instead of the widely-used RGB cameras. HSI image is informative in spectrums and full of potential for natural environment perception. In this article, a first-of-its-kind HSI road segmentation dataset is built with careful annotation in both urban and rural scenes. It contains 3799 scenes with RGB and NIR bands as well as their respective masks. Unlike many existing datasets that provide urban scenes in RGB images only, our dataset expands the sensing spectrum to 28 bands and includes various kinds of road surfaces, such as asphalt, cement, dirt and sand, under rural and natural scenes. We also provide benchmark performances based on the recently popular segmentation algorithms on this dataset. The dataset is released at github‡.‡https://github.com/NUST-Machine-Intelligence-Laboratory/hsi_road Jiarou Lu, Huafeng Liu 0004, Yazhou Yao, Shuyin Tao, Zhenmin Tang, Jianfeng Lu 0003 |
ICME | 6 |
| 2020 | End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular CameraabstractInter-vehicle distance and relative velocity estimations are two basic functions for any ADAS (Advanced driver-assistance systems). In this paper, we propose a monocular camera based inter-vehicle distance and relative velocity estimation method based on end-to-end training of a deep neural network. The key novelty of our method is the integration of multiple visual clues provided by any two time-consecutive monocular frames, which include deep feature clue, scene geometry clue, as well as temporal optical flow clue. We also propose a vehicle-centric sampling mechanism to alleviate the effect of perspective distortion in the motion field (i.e. optical flow). We implement the method by a light-weight deep neural network. Extensive experiments are conducted which confirm the superior performance of our method over other state-of-the-art methods, in terms of estimation accuracy, computational speed, and memory footprint. Zhenbo Song, Jianfeng Lu 0003, Tong Zhang 0023, Hongdong Li |
ICRA | 2 |
| 2020 | Nonlocal Low-Rank Nonnegative Tensor Factorization for Hyperspectral UnmixingabstractHyperspectral unmixing decomposes hyperspectral images (HSI) into a collection of constituent materials or end-members and their fractions, i.e., abundances. Nonnegative tensor factorization (NTF) has been utilized thanks to its ability of preserving all the information in HSI. However, NTF based unmixing only makes use of global spatial-spectral information without considering detailed local/non-local spatial information, making it vulnerable to real-world disturbance such as noises. To this end, in this paper, we extend NTF by introducing non-local low-rank constraint to abundance maps. The additional regularization on abundances facilities tensor factorization avoid being trapped into a large number of suspicious solutions, so as to preserve the non-local spatial structure on abundance maps. Experimental results on synthetic data and real-world data show that the proposed method outperforms the state-of-the-art methods. Fengchao Xiong, Kun Qian 0015, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IGARSS | 3 |
| 2020 | Neural Machine Translation with Error CorrectionabstractNeural machine translation (NMT) generates the next target token given as input the previous ground truth target tokens during training while the previous generated target tokens during inference, which causes discrepancy between training and inference as well as error propagation, and affects the translation accuracy. In this paper, we introduce an error correction mechanism into NMT, which corrects the error information in the previous generated tokens to better predict the next token. Specifically, we introduce two-stream self-attention from XLNet into NMT decoder, where the query stream is used to predict the next token, and meanwhile the content stream is used to correct the error information from the previous predicted tokens. We leverage scheduled sampling to simulate the prediction errors during training. Experiments on three IWSLT translation datasets and two WMT translation datasets demonstrate that our method achieves improvements over Transformer baseline and scheduled sampling. Further experimental analyses also verify the effectiveness of our proposed error correction mechanism to improve the translation quality. Kaitao Song, Xu Tan 0003, Jianfeng Lu 0003 |
IJCAI | 3 |
| 2020 | MPNet: Masked and Permuted Pre-training for Language UnderstandingabstractBERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this problem. However, XLNet does not leverage the full position information of a sentence and thus suffers from position discrepancy between pre-training and fine-tuning. In this paper, we propose MPNet, a novel pre-training method that inherits the advantages of BERT and XLNet and avoids their limitations. MPNet leverages the dependency among predicted tokens through permuted language modeling (vs. MLM in BERT), and takes auxiliary position information as input to make the model see a full sentence and thus reducing the position discrepancy (vs. PLM in XLNet). We pre-train MPNet on a large-scale dataset (over 160GB text corpora) and fine-tune on a variety of down-streaming tasks (GLUE, SQuAD, etc). Experimental results show that MPNet outperforms MLM and PLM by a large margin, and achieves better results on these tasks compared with previous state-of-the-art pre-trained methods (e.g., BERT, XLNet, RoBERTa) under the same model setting. We attach the code in the supplemental materials. Kaitao Song, Xu Tan 0003, Tao Qin 0001, Jianfeng Lu 0003, Tie-Yan Liu |
NeurIPS | 4 |
| 2020 | Community detection method using improved density peak clustering and nonnegative matrix factorization
Zhu Shen, Xiaoshuang Sang, Jianfeng Lu 0003 |
Neurocomputing | 5 |
| 2020 | SPSSNet: a real-time network for image semantic segmentationabstractAlthough deep neural networks (DNNs) have achieved great success in semantic segmentation tasks, it is still challenging for real-time applications. A large number of feature channels, parameters, and floating-point operations make the network sluggish and computationally heavy, which is not desirable for real-time tasks such as robotics and autonomous driving. Most approaches, however, usually sacrifice spatial resolution to achieve inference speed in real time, resulting in poor performance. In this paper, we propose a light-weight stage-pooling semantic segmentation network (SPSSN), which can efficiently reuse the paramount features from early layers at multiple stages, at different spatial resolutions. SPSSN takes input of full resolution 2048×1024 pixels, uses only 1.42 × 10 6 parameters, yields 69.4% mIoU accuracy without pre-training, and obtains an inference speed of 59 frames/s on the Cityscapes dataset. SPSSN can run directly on mobile devices in real time, due to its light-weight architecture. To demonstrate the effectiveness of the proposed network, we compare our results with those of state-of-the-art networks. Saqib Mamoon, Muhammad Arslan Manzoor, Faen Zhang, Zakir Ali, Jianfeng Lu 0003 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2020 | Robust mixed-norm constrained regression with application to face recognitions
Xiaoshuang Sang, Yesong Xu, Zakir Ali, Jianfeng Lu 0003 |
Neural Comput. Appl. | 6 |
| 2020 | Community Detection in Complex Networks Using Nonnegative Matrix Factorization and Density-Based Clustering Algorithm
Xiaoshuang Sang, Jianfeng Lu 0003 |
Neural Process. Lett. | 4 |
| 2020 | Bi-Modal Progressive Mask Attention for Fine-Grained RecognitionabstractTraditional fine-grained image recognition is required to distinguish different subordinate categories (e.g., birds species) based on the visual cues beneath raw images. Due to both small inter-class variations and large intra-class variations, it is desirable to capture the subtle differences between these sub-categories, which is crucial but challenging for fine-grained recognition. Recently, language modality aggregation has been proved as a successful technique to improve visual recognition in the experience. In this paper, we introduce an end-to-end trainable Progressive Mask Attention (PMA) model for fine-grained recognition by leveraging both visual and language modalities. Our Bi-Modal PMA model can not only stage-by-stage capture the most discriminative part in the visual modality by our mask-based fashion, but also explore the out-of-visual-domain knowledge from the language modality in an interactional alignment paradigm. Specifically, at each stage, a self-attention module is proposed to attend to the key patch from images or text descriptions. Besides, a query-relational module is designed to seize the key words/phrases of texts and further bridge the connection between two modalities. Later, the learned representations of bi-modality from multiple stages are aggregated as the final features for recognition. Our Bi-Modal PMA model only needs raw images and raw text descriptions, without requiring bounding boxes/part annotations in images or key word annotations in texts. By conducting comprehensive experiments on fine-grained benchmark datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines, on either vision and language bi-modality or single visual modality. Kaitao Song, Xiu-Shen Wei, Xiangbo Shu, Renjie Song, Jianfeng Lu 0003 |
IEEE Trans. Image Process. | 5 |
| 2019 | Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention NetworksabstractAs the visual reflections of our daily lives, images are frequently shared on the social network, which generates the abundant 'metadata' that records user interactions with images. Due to the diverse contents and complex styles, some images can be challenging to recognise when neglecting the context. Images with the similar metadata, such as 'relevant topics and textual descriptions', 'common friends of users' and 'nearby locations', form a neighbourhood for each image, which can be used to assist the annotation. In this paper, we propose a Metadata Neighbourhood Graph Co-Attention Network (MangoNet) to model the correlations between each target image and its neighbours. To accurately capture the visual clues from the neighbourhood, a co-attention mechanism is introduced to embed the target image and its neighbours as graph nodes, while the graph edges capture the node pair correlations. By reasoning on the neighbourhood graph, we obtain the graph representation to help annotate the target image. Experimental results on three benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods. Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003 |
CVPR | 5 |
| 2019 | MASS: Masked Sequence to Sequence Pre-training for Language GenerationabstractPre-training and fine-tuning, e.g., BERT \citep{devlin2018bert}, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to Sequence pre-training (MASS) for the encoder-decoder based language generation tasks. MASS adopts the encoder-decoder framework to reconstruct a sentence fragment given the remaining part of the sentence: its encoder takes a sentence with randomly masked fragment (several consecutive tokens) as input, and its decoder tries to predict this masked fragment. In this way, MASS can jointly train the encoder and decoder to develop the capability of representation extraction and language modeling. By further fine-tuning on a variety of zero/low-resource language generation tasks, including neural machine translation, text summarization and conversational response generation (3 tasks and totally 8 datasets), MASS achieves significant improvements over the baselines without pre-training or with other pre-training methods. Especially, we achieve the state-of-the-art accuracy (30.02 in terms of BLEU score) on the unsupervised English-French translation, even beating the early attention-based supervised model \citep{bahdanau2015neural}. Kaitao Song, Xu Tan 0003, Tao Qin 0001, Jianfeng Lu 0003, Tie-Yan Liu |
ICML | 4 |
| 2019 | Speed-Invariant Gait Recognition Using Single-Support Gait Energy ImageabstractGait is one of the most popular behavioral biometrics because it can be authenticated at a distance from a camera without subject cooperation. Speed differences between matching pairs, however, cause significant performance drops in gait recognition, and gait mode difference (i.e., walking versus running) makes gait recognition further challenging. We therefore propose a speed-invariant gait representation called single-support GEI (SSGEI), which realizes a good trade-off between speed invariance and stability by aggregating multiple frames around single-support phases. In addition, to mitigate the pose differences between walking and running modes at single-support phases, we morph walking and running SSGEIs into intermediate SSGEIs between walking and running mode, where we exploit a free-form deformation field from the walking or running modes to the intermediate mode obtained by training data. We finally apply Gabor filtering and spatial metric learning as postprocessing for further accuracy improvement. Experiments on two publicly available datasets, the OU-ISIR Treadmill Dataset A and the CASIA-C Dataset demonstrate that the proposed method yields the state-of-the-art accuracies in both identification and verification scenarios with a low computational cost. Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
Multim. Tools Appl. | 5 |
| 2019 | Gait-based age progression/regression: a baseline and performance evaluation by age group classification and cross-age gait identificationabstractGait is believed to be an advanced behavioral biometric that can be perceived at a large distance from a camera without subject cooperation and hence is favorable for many applications in surveillance and forensics. However, appearance differences caused by human aging may significantly reduce the performance of gait recognition. Modeling the aging process on gait features is one of the possible solutions to this problem, and it may inspire more potential applications, such as finding lost children and examining health status. To the best of our knowledge, this topic has not been studied in the literature. Motivated by the fact that aging effects are mainly reflected in the shape and appearance deformations of the gait feature, we propose a baseline algorithm for gait-based age progression and regression using a generic geometric transformation between different age groups, in conjunction with the gait energy image, which is an appearance-based gait feature frequently used in the gait analysis community, to render gait aging and reverse aging effects simultaneously. Various evaluations were conducted through gait-based age group classification and cross-age gait identification to validate the performance of the proposed method, in addition to providing several insights for future research on the subject. Chi Xu 0003, Yasushi Makihara, Yasushi Yagi, Jianfeng Lu 0003 |
Mach. Vis. Appl. | 4 |
| 2019 | Statistical performance of convex low-rank and sparse tensor recovery
Xiangrui Li, Andong Wang, Jianfeng Lu 0003, Zhenmin Tang |
Pattern Recognit. | 3 |
| 2019 | Heritage image annotation via collective knowledge
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003, Qiang Wu 0001 |
Pattern Recognit. | 5 |
| 2018 | Kill Two Birds With One Stone: Weakly-Supervised Neural Network for Image Annotation and Tag RefinementabstractThe number of social images has exploded by the wide adoption of social networks, and people like to share their comments about them. These comments can be a description of the image, or some objects, attributes, scenes in it, which are normally used as the user-provided tags. However, it is well-known that user-provided tags are incomplete and imprecise to some extent. Directly using them can damage the performance of related applications, such as the image annotation and retrieval. In this paper, we propose to learn an image annotation model and refine the user-provided tags simultaneously in a weakly-supervised manner. The deep neural network is utilized as the image feature learning and backbone annotation model, while visual consistency, semantic dependency, and user-error sparsity are introduced as the constraints at the batch level to alleviate the tag noise. Therefore, our model is highly flexible and stable to handle large-scale image sets. Experimental results on two benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods. Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003 |
AAAI | 5 |
| 2018 | Double Path Networks for Sequence to Sequence LearningabstractEncoder-decoder based Sequence to Sequence learning (S2S) has made remarkable progress in recent years. Different network architectures have been used in the encoder/decoder. Among them, Convolutional Neural Networks (CNN) and Self Attention Networks (SAN) are the prominent ones. The two architectures achieve similar performances but use very different ways to encode and decode context: CNN use convolutional layers to focus on the local connectivity of the sequence, while SAN uses self-attention layers to focus on global semantics. In this work we propose Double Path Networks for Sequence to Sequence learning (DPN-S2S), which leverage the advantages of both models by using double path information fusion. During the encoding step, we develop a double path architecture to maintain the information coming from different paths with convolutional layers and self-attention layers separately. To effectively use the encoded context, we develop a gated attention fusion module and use it to automatically pick up the information needed during the decoding step, which is also a double path network. By deeply integrating the two paths, both types of information are combined and well exploited. Experiments show that our proposed method can significantly improve the performance of sequence to sequence learning over state-of-the-art systems. Kaitao Song, Xu Tan 0003, Di He 0001, Jianfeng Lu 0003, Tao Qin 0001, Tie-Yan Liu |
COLING | 4 |
| 2018 | Single Image Water Hazard Detection Using FCN with Reflection Attention Units
Chuong V. Nguyen, Shaodi You, Jianfeng Lu 0003 |
ECCV (6) | 4 |
| 2018 | Goal-Oriented Visual Question Generation via Intermediate Rewards
Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003, Anton van den Hengel |
ECCV (5) | 5 |
| 2018 | A Least Squares Approach to Region Selection
Liantao Wang, Jianfeng Lu 0003 |
ICONIP (5) | 3 |
| 2018 | Part-based Multi-stream Model for Vehicle SearchingabstractDue to the enormous requirement in public security and intelligent transportation system, searching an identical vehicle has become more and more important. Current studies usually treat vehicle as an integral object and then train a distance metric to measure the similarity among vehicles. However, these raw images may be exactly similar to ones with different identification and include some pixels in background that may disturb the distance metric learning. In this paper, we propose a novel and useful method to segment an original vehicle image into several discriminative foreground parts, and these parts consist of some fine grained regions that are named discriminative patches. After that, these parts combined with the raw image are fed into the proposed deep learning network. We can easily measure the similarity of two vehicle images by computing the Euclidean distance of the features from FC layer. Two main contributions of this paper are as follows. Firstly, a method is proposed to estimate if a patch in a raw vehicle image is discriminative or not. Secondly, a new Part-based Multi-Stream Model (PMSM) is designed and optimized for vehicle retrieval and re-identification tasks. We evaluate the proposed method on the VehicleID dataset, and the experimental results show that our method can outperform the baseline. Minxian Li, Jianfeng Lu 0003 |
ICPR | 3 |
| 2018 | Generating Adversarial Examples With Conditional Generative Adversarial NetabstractRecently, deep neural networks have significant progress and successful application in various fields, but they are found vulnerable to attack instances, e.g., adversarial examples. State-of-art attack methods can generate attack images by adding small perturbation to the source image. These attack images can fool the classifier but have little impact to human. Therefore, such attack instances are difficult to generate by searching the feature space. How to design an effective and robust generating method has become a spotlight. Inspired by adversarial examples, we propose two novel generative models to produce adaptive attack instances directly, in which conditional generative adversarial network is adopted and distinctive strategy is designed for training. Compared with the common method, such as Fast Gradient Sign Method, our models can reduce the generating cost and improve robustness and has about one fifth running time for producing attack instance. Kaitao Song, Jianfeng Lu 0003 |
ICPR | 3 |
| 2018 | Attention-based Neural Network for Traffic Sign DetectionabstractExisting object detection pipelines can show superior performance for large objects with high resolution but fail to detect very small objects such as traffic signs. So, detecting traffic signs is a proverbially challenging problem. In this paper, we propose a novel end-to-end architecture that improves small object detection by combining Faster R-CNN with the attention mechanism. Specifically, we focus on channel-wise features and utilize the attention mechanism to enhance the feature responses by explicitly modeling the interdependencies between channel-wise features. Finally, the regression of bounding boxes and the classification of traffic signs are generated after selecting the discriminative features by the attention mechanism. Extensive evaluations of the largest traffic sign dataset demonstrate that the attention mechanism improves the performance of detecting objects, especially the small targets. For traffic sign detection task, our method achieves better performance compared with many state-of-the-art approaches on the largest traffic sign detection dataset, Tsinghua-Tencent 100K. Le Hui, Jianfeng Lu 0003, Yuhua Zhu |
ICPR | 3 |
| 2018 | Fully Convolutional Neural Networks for Road Detection with Multiple Cues IntegrationabstractRoad detection from images is a key task in autonomous driving. The recent advent of deep learning (and in particular, CNN or convolutional neural networks) has greatly improved the performance of road detection algorithms. In this paper, we show how to fuse multiple different cues under the same convolutional network framework. Specifically, we adopt a pre-trained Resnet-lOl to extract feature maps from RGB images; we then connect it with three extra deconvolution layers. These deconvolution layers is trained conditioning on appropriate image cues, and in our case they are a height image (i.e. elevation map obtained by e.g. Lidar scanner), image gradient, and position map. We also design two skip layers to speed up the convergence. Experiments on KITTI benchmark show competitive performance of our new networks. Jianfeng Lu 0003, Chunxia Zhao, Hongdong Li |
ICRA | 2 |
| 2018 | Non-concept density estimation via kernel regression for concept ranking in weakly labelled dataabstractAutomatic object annotation for weakly labelled images/videos has attracted great research interests. In the literature, the idea of negative mining has been proposed for the task. Following existing works, the authors start with image/video over‐segmentation. With the assumption that the noisy segments in the concept images and the strongly labelled non‐concept segments are drawn from the same distribution, the authors plan to estimate the non‐concept distribution and apply it to the ambiguous segments to generate a concept ranking. Although this idea was proposed in existing work and was shown ineffective when combined with a naive kernel density estimation strategy, in this study, the authors explore improved density estimation techniques for the ranking and propose a kernel regression model whose parameters are estimated by a maximum likelihood estimation. Experimental results validate the effectiveness of their method. Liantao Wang, Qingwu Li, Jianfeng Lu 0003 |
IET Comput. Vis. | 3 |
| 2018 | Semisupervised and Weakly Supervised Road Detection Based on Generative Adversarial NetworksabstractRoad detection is a key component of autonomous driving; however, most fully supervised learning road detection methods suffer from either insufficient training data or high costs of manual annotation. To overcome these problems, we propose a semisupervised learning (SSL) road detection method based on generative adversarial networks (GANs) and a weakly supervised learning (WSL) method based on conditional GANs. Specifically, in our SSL method, the generator generates the road detection results of labeled and unlabeled images, and then they are fed into the discriminator, which assigns a label on each input to judge whether it is labeled. Additionally, in WSL method we add another network to predict road shapes of input images and use them in both generator and discriminator to constrain the learning progress. By training under these frameworks, the discriminators can guide a latent annotation process on the unlabeled data; therefore, the networks can learn better representations of road areas and leverage the feature distributions on both labeled and unlabeled data. The experiments are carried out on KITTI ROAD benchmark, and the results show our methods achieve the state-of-the-art performances. Jianfeng Lu 0003, Chunxia Zhao, Shaodi You, Hongdong Li |
IEEE Signal Process. Lett. | 2 |
| 2018 | Multilabel Image Classification With Regional Latent Semantic DependenciesabstractDeep convolution neural networks (CNNs) have demonstrated advanced performance on single-label image classification, and various progress also has been made to apply CNN methods on multilabel image classification, which requires annotating objects, attributes, scene categories, etc., in a single shot. Recent state-of-the-art approaches to the multilabel image classification exploit the label dependencies in an image, at the global level, largely improving the labeling capacity. However, predicting small objects and visual concepts is still challenging due to the limited discrimination of the global visual features. In this paper, we propose a regional latent semantic dependencies model (RLSD) to address this problem. The utilized model includes a fully convolutional localization architecture to localize the regions that may contain multiple highly dependent labels. The localized regions are further sent to the recurrent neural networks to characterize the latent semantic dependencies at the regional level. Experimental results on several benchmark datasets show that our proposed model achieves the best performance compared to the state-of-the-art models, especially for predicting small objects occurring in the images. Also, we set up an upper bound model (RLSD+ft-RPN) using bounding-box coordinates during training, and the experimental results also show that our RLSD can approach the upper bound without using the bounding-box annotations, which is more realistic in the real world. Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003 |
IEEE Trans. Multim. | 5 |
| 2017 | Davies Bouldin Index based hierarchical initialization K-meansabstractK-means algorithm is an effective clustering algorithm based on partition, which has been widely used for clustering analysis. However, there are two main problems for K-means algorithm: how to provide appropriate number of clusters and how to determine initial cluster centers automatically. Plenty of methods have been proposed to address the above problems. In our previous work, we proposed the hierarchical initialization approach to determine initial cluster centers, but we cannot provide the number of clusters automatically. In this paper, in order to determine the number of clusters automatically, we propose the Davies Bouldin Index (DBI) based hierarchical K-means (DHIKM) algorithm on the basis of our previous work. The proposed algorithm can integrate DBI metric into our hierarchical K-means algorithm and can determine the number of clusters with low time cost. Experiments on UCI datasets and synthetic data demonstrate the effectiveness and feasibility of the proposed algorithm. Junwei Xiao, Jianfeng Lu 0003, Xiangyu Li 0005 |
Intell. Data Anal. | 2 |
| 2017 | Learning arbitrary-shape object detector from bounding-box annotation by searching region-graph
Liantao Wang, Jianfeng Lu 0003, Xiangyu Li 0005, Zhan Huan, Jiuzhen Liang, Shuyue Chen |
Pattern Recognit. Lett. | 2 |
| 2017 | Instance Annotation via Optimal BoW for Weakly Supervised Object LocalizationabstractIn this paper, we aim at irregular-shape object localization under weak supervision. With over-segmentation, this task can be transformed into multiple-instance context. However, most multiple-instance learning methods only emphasize single most positive instance in a positive bag to optimize bag-level classification, and leads to imprecise or incomplete localization. To address this issue, we propose a scheme for instance annotation, where all of the positive instances are detected by labeling each instance in each positive bag. Inspired by the successful application of bag-of-words (BoW) to feature representation, we leverage it at instance-level to model the distributions of the positive class and negative class, and then incorporate the BoW learning and instance labeling in a single optimization formulation. We also demonstrate that the scheme is well suited to weakly supervised object localization of irregular-shape. Experimental results validate the effectiveness both for the problem of generic instance annotation and for the application of weakly supervised object localization compared to some existing methods. Liantao Wang, Deyu Meng, Xuelei Hu, Jianfeng Lu 0003, Ji Zhao 0001 |
IEEE Trans. Cybern. | 4 |
| 2016 | Speed Invariance vs. Stability: Cross-Speed Gait Recognition Using Single-Support Gait Energy Image
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
ACCV (2) | 5 |
| 2016 | MetricRec: Metric Learning for Cold-Start Recommendations
Furong Peng, Xuan Lu 0001, Jianfeng Lu 0003, Chao Ma 0002, Jing-Yu Yang 0001 |
ADMA | 3 |
| 2016 | Exploring Brain Networks via Structured Sparse Representation of fMRI Data
Jianfeng Lu 0003, Jinglei Lv, Xi Jiang 0001, Shijie Zhao 0001, Tianming Liu 0001 |
MICCAI (1) | 2 |
| 2016 | N-dimensional Markov random field prior for cold-start recommendation
Furong Peng, Jianfeng Lu 0003, Chao Ma 0002, Jing-Yu Yang 0001 |
Neurocomputing | 2 |
| 2016 | Unsupervised discriminant canonical correlation analysis based on spectral clustering
Sheng Wang 0015, Jianfeng Lu 0003, Xingjian Gu, Benjamin Asubam Weyori, Jing-Yu Yang 0001 |
Neurocomputing | 2 |
| 2016 | Canonical principal angles correlation analysis for two-view data
Sheng Wang 0015, Jianfeng Lu 0003, Xingjian Gu, Chunhua Shen, Jing-Yu Yang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Semi-supervised linear discriminant analysis for dimension reduction and classification
Sheng Wang 0015, Jianfeng Lu 0003, Xingjian Gu, Haishun Du, Jing-Yu Yang 0001 |
Pattern Recognit. | 2 |
| 2015 | Fiber Connection Pattern-Guided Structured Sparse Representation of Whole-Brain fMRI Signals for Functional Network Inference
Xi Jiang 0001, Jianfeng Lu 0003, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 4 |
| 2015 | Distance Networks for Morphological Profiling and Characterization of DICCCOL Landmarks
Hanbo Chen, Jianfeng Lu 0003, Tianming Liu 0001 |
MICCAI (2) | 3 |
| 2015 | Multi-scale and Multimodal Fusion of Tract-Tracing, Myelin Stain and DTI-derived Fibers in Macaque Brains
Ke Jing, Hanbo Chen, Xi Jiang 0001, Longchuan Li, Lei Guo 0002, Jianfeng Lu 0003, Xiaoping Hu 0001, Tianming Liu 0001 |
MICCAI (2) | 8 |
| 2015 | Active learning via query synthesis and nearest neighbour search
Liantao Wang, Xuelei Hu, Bo Yuan 0003, Jianfeng Lu 0003 |
Neurocomputing | 4 |
| 2014 | Street view cross-sourced point cloud matching and registrationabstractObject registration has been widely discussed with the development of various range sensing technologies. In most work, however, the point clouds of reference and target are generated by the same technology, such as a Kinect range camera, LiDAR sensor, or Structure from Motion technique. Cases in which reference and target point clouds are generated by different technologies are rarely discussed. Due to the significant differences across various point cloud data in terms of point cloud density, sensing noise, scale, occlusion etc., object registration between such different point clouds becomes extremely difficult. In this study, we address for the first time an even more challenging case in which the differently-sourced point clouds are acquired from a real street view. One is generated on the basis of an image sequence through the SfM process, and the other is produced directly by the LiDAR system. We propose a two-stage matching and registration algorithm to achieve object registration between these two different point clouds. The experiments are based on real building object point cloud data and demonstrate the effectiveness and efficiency of the proposed solution. The newly proposed solution can be further developed to contribute to several related applications, such as Location Based Service. Furong Peng, Qiang Wu 0001, Lixin Fan, Jian Zhang 0002, Yu You, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
ICIP | 6 |
| 2014 | Weakly supervised object localization via maximal entropy random walkabstractIn this paper, we investigate the problem of weakly supervised object localization in images. For such a problem, the goal is to predict the locations of objects in test images while the labels of the training images are given at image-level. That means a label only indicates whether an image contains objects or not, but does not provide the exact locations of the objects. We propose to address this problem using Maximal Entropy Random Walk (MERW). Specifically, we first train a linear SVM classifier with the weakly labeled data. Based on bag-of-words feature representation, the response of a region to the linear SVM classifier can be formulated as the sum of the feature-weights within the region. For a test image, by properly constructing a graph on the feature-points, the stationary distribution of a MERW can indicate the region with the densest positive feature-weights, and thus provides a probabilistic object localization. Experiments compared with state-of-the-art methods on two datasets validate the performance of our method. Liantao Wang, Ji Zhao 0001, Xuelei Hu, Jianfeng Lu 0003 |
ICIP | 4 |
| 2014 | Unsupervised Discriminant Canonical Correlation Analysis for Feature FusionabstractCanonical correlation analysis (CCA) has been widely applied to information fusion. It only considers the correlated information of the paired data, but ignores the correlated information between the samples in the same class. Furthermore, class information is useful for CCA, but there is little class information in the scenarios of real applications. Thus, it is difficult to utilize the correlated information between the samples in the same class. To utilize the correlated information between the samples, we propose a method named Unsupervised Discriminant Canonical Correlation Analysis (UDCCA). In UDCCA, the class membership and mapping are iteratively computed by using the normalized spectral clustering and generalized Eigen value methods alternatively. The experimental results on the MFD dataset and ORL dataset show that UDCCA outperforms traditional CCA and its variants in most situations. Sheng Wang 0015, Xingjian Gu, Jianfeng Lu 0003, Jing-Yu Yang 0001, Ruili Wang 0001, Jian Yang 0003 |
ICPR | 3 |
| 2014 | Multiple kernel clustering based on centered kernel alignment
Yanting Lu, Liantao Wang, Jianfeng Lu 0003, Jing-Yu Yang 0001, Chunhua Shen |
Pattern Recognit. | 3 |
| 2013 | Instance Selection and Instance Weighting for Cross-Domain Sentiment Classification via PU Learning
Xuelei Hu, Jianfeng Lu 0003, Jian Yang 0003, Chengqing Zong |
IJCAI | 3 |
| 2012 | Adaptive kernel learning based on centered alignment for hierarchical classification
Yanting Lu, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
ICPR | 2 |
| 2012 | Density-based hierarchical clustering for streaming data
Q. Tu, Jianfeng Lu 0003, Bo Yuan 0003, J. B. Tang, Jing-Yu Yang 0001 |
Pattern Recognit. Lett. | 2 |
| 2011 | An augmented Lagrangian method for fast gradient vector flow computationabstractGradient vector flow (GVF) and its generalization have been widely applied in many image processing applications. The high cost of GVF computation, however, has restricted their potential applications to images with large size. In this paper, motivated by progress in fast image restoration algorithms, we reformulate the GVF computation problem as a convex optimization model with an equality constraint, and solve it using a fast algorithm, inexact augmented Lagrangian method (ALM). With fast Fourier transform (FFT), we provide a novel simple and efficient algorithm for GVF computation. Experimental results show that the proposed method can improve the computational speed by an order of magnitude, and is even more efficient for images with large sizes. Jianfeng Lu 0003, Wangmeng Zuo, David Zhang 0001 |
ICIP | 1 |
| 2011 | Graph attribute embedding via Riemannian submersion learning
Haifeng Zhao 0002, Antonio Robles-Kelly, Jun Zhou 0001, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2010 | Automated Cell Phase Classification for Zebrafish Fluorescence Microscope ImagesabstractAutomated cell phenotype image classification is an interesting bioinformatics problem. In this paper, an automated cell phase classification framework is investigated for zebra fish presomitic mesoderm (PSM) images. Low image resolution, gradual transitions between adjacent categories and irregularity of real cell images make this classification task tough but intriguing. The proposed framework first segments zebra fish image into cell patches by a two-stage segmentation procedure, then extracts feature set NF9, which designed especially for this low resolution image set, on each cell patch, and finally employs support vector machine (SVM) as cell classifier. At present, the total accuracy by NF9 is 75%. Yanting Lu, Jianfeng Lu 0003, Tianming Liu 0001, Jing-Yu Yang 0001 |
ICPR | 2 |
| 2008 | Hierarchical initialization approach for K-Means clustering
Jianfeng Lu 0003, J. B. Tang, Zhenmin Tang, Jing-Yu Yang 0001 |
Pattern Recognit. Lett. | 1 |
| 2004 | An efficient renovation on kernel Fisher discriminant analysis and face recognition experiments
Yong Xu 0001, Jing-Yu Yang 0001, Jianfeng Lu 0003, Dongjun Yu |
Pattern Recognit. | 3 |
| 2003 | Feature fusion: parallel strategy vs. serial strategy
Jian Yang 0003, Jing-Yu Yang 0001, David Zhang 0001, Jianfeng Lu 0003 |
Pattern Recognit. | 4 |