EDBT 2026 Demo / reviewers in the wild / expert
Jianjun Qian
dblp:03/3289
· DBLP profile ↗
113ranked-venue papers
9as first author
71since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 6 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 2 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shaping Without Tearing: Controllable Diffeomorphic Deformations for Topology-Preserving 3D Point Cloud AugmentationabstractPoint cloud data augmentation is critical to improving the generalization of 3D deep learning models. However, existing methods often fail to preserve the underlying manifold structure, leading to semantic distortion or topology violation. This causes models to learn untrustworthy features, thereby limiting the representational ability of the model. To overcome these limitations, we propose ManiPoint, a novel point cloud augmentation framework based on diffeomorphism that explicitly preserves manifold structure during deformation. ManiPoint constructs diffeomorphic transformations via continuous differentiable mappings, ensuring topological consistency and geometric continuity between original and augmented data. To prevent excessive distortion and ensure semantic consistency, we introduce a controllable deformation mechanism that quantitatively constrains the augmentation magnitude and enables fine-grained control over the deformation space. We further provide theoretical analysis, indicating that, compared with topologically inconsistent methods, ManiPoint reduces empirical and vicinal risks by generating diverse and structurally reliable samples. Extensive experiments and visualizations on object-level datasets demonstrate that ManiPoint produces high-quality augmentations and consistently improves model robustness over existing baselines. Meanwhile, the scalability of our method was further verified on the scene-level datasets. Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
AAAI | 3 |
| 2026 | Discriminative region learning for point cloud-based place recognition
Le Hui, Yun Zhu 0011, Jianjun Qian, Yigong Zhang, Jin Xie 0001 |
Neural Networks | 4 |
| 2026 | Structure-aware spherical density steered cross-domain learning for effective point cloud understanding
Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2026 | Leaning geometrical diffusion network via power spherical distribution for point clouds generation
Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0024 |
Pattern Recognit. | 3 |
| 2026 | TranSpike: Pixel-wise frequency reconstruction and spike interaction for remote photoplethysmography
Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Chuanfei Hu, Shuo Chen 0003, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2026 | Learning From Past and Future: A Unified Instantaneous Pedestrian Intent Prediction Framework Based on Privileged Knowledge Distillation for Autonomous Driving
Xiaobo Chen 0001, Wei Xu 0052, Jianjun Qian |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Physics-Guided Posterior Sampling for Diffusion-Based Real-World Dehazing and Image EnhancementabstractReal-world image dehazing is a challenging task due to the collection of aligned hazy/clear image pairs under unpredictable and complex environments. To address this limitation, we propose a Physical-Guided Posterior Sampling (PGPS) method that designs a dehazing reconstruction posterior to sample an RGB and depth from pre-trained unconditional diffusion generation process. First, we introduce a Hybrid Degradation Atmospheric Scattering Model (HD-ASM) to adapt the diffusion model, enabling the generation of high-fidelity dehazed images from posterior samples without relying on the aligned hazy/clear image pairs. Second, we propose a two-stage sampling strategy with piecewise loss to improve sampling quality and stability, along with a post-processing technique to remove JPEG compression artifacts amplified by dehazing. Extensive experiments show that our method outperforms state-of-the-art techniques in image dehazing, and in the RTTS dataset’s complex human-vehicle environment. Additionally, our approach also surpasses other benchmarks in object detection, exhibiting superior generalization performance. Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Jianjun Qian, Heyou Chang, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | WeatherCycle: Unpaired Multi-Weather Restoration via Color Space Decoupled Cycle LearningabstractUnsupervised image restoration under multi-weather conditions remains a fundamental yet underexplored challenge. While existing methods often rely on task-specific physical priors, their narrow focus limits scalability and generalization to diverse real-world weather scenarios. In this work, we propose WeatherCycle, a unified unpaired framework that reformulates weather restoration as a bidirectional degradation-content translation cycle, guided by degradation-aware curriculum regularization. At its core, WeatherCycle employs alumina-chroma decompositionstrategy to decouple degradation from content without modeling complex weather, enabling domain conversion between degraded and clean images. To model diverse and complex degradations, we propose aLumina Degradation Guidance Module(LDGM), which learns luminance degradation priors from a degraded image pool and injects them into clean images via frequency-domain amplitude modulation, enabling controllable and realistic degradation modeling. Additionally, we incorporate aDifficulty-Aware Contrastive Regularization(DACR) module that identifies hard samples via a CLIP-based classifier and enforces contrastive alignment between hard samples and restored features to enhance semantic consistency and robustness. Extensive experiments across serve multi-weather datasets, demonstrate that our method achieves state-of-the-art performance among unsupervised approaches, with strong generalization to complex weather degradations. Wenxuan Fang 0001, Jiangwei Weng, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | TDHF: Task-Driven Hierarchical Image Fusion Under Low-Light ConditionsabstractVision analysis tasks often experience substantial performance drops when processing images captured in lowlight environments. Existing approaches introduce infrared images to provide complementary information, and fusion strategies are consequently developed to combine the advantages of visible light images and infrared images. While fusion methods are typically designed to enhance visual quality with the expectation of improving task performance, many existing approaches tend to overemphasize perceptual fidelity, which can inadvertently compromise task-specific feature. To address this issue, we propose a task-driven hierarchical fusion (TDHF) framework designed to retain task-relevant information throughout the fusion process. Specifically, TDHF adopts a multi-scale hierarchical architecture to capture rich feature representations and incorporates a multi-head attention mechanism to model cross-modal interactions. In addition, we introduce a single-step denoising generation module that guides the fusion of infrared edge features and texture details from low-light images progressively. This process ultimately reconstructs the task-critical Y channel in the YCrCb color space. Extensive experiments on object detection under low-light conditions across four benchmark datasets demonstrate that TDHF effectively enhances task performance, achieving up to 1.1% higher mAP on LLVIP and 1.3% higher mAP on FLIR compared with state-of-the-art methods, without relying on excessive optimization of image quality (PSNR). Daoheng Li, Mengkai Yan, Jun Li 0027, Jian Yang 0003, Jianjun Qian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | ASGNet: Adaptive Spectrum Guidance Network for Automatic Polyp SegmentationabstractEarly identification and removal of polyps can reduce the risk of developing colorectal cancer. However, the diverse morphologies, complex backgrounds and often concealed nature of polyps make polyp segmentation in colonoscopy images highly challenging. Despite the promising performance of existing deep learning-based polyp segmentation methods, their perceptual capabilities remain biased toward local regions, mainly because of the strong spatial correlations between neighboring pixels in the spatial domain. This limitation makes it difficult to capture the complete polyp structures, ultimately leading to sub-optimal segmentation results. In this paper, we propose a novel adaptive spectrum guidance network, called ASGNet, which addresses the limitations of spatial perception by integrating spectral features with global attributes. Specifically, we first design a spectrum-guided non-local perception module that jointly aggregates local and global information, therefore enhancing the discriminability of polyp structures, and refining their boundaries. Moreover, we introduce a multi-source semantic extractor that integrates rich high-level semantic information to assist in the preliminary localization of polyps. Furthermore, we construct a dense cross-layer interaction decoder that effectively integrates diverse information from different layers and strengthens it to generate high-quality representations for accurate polyp segmentation. Extensive quantitative and qualitative results demonstrate the superiority of our ASGNet approach over 21 state-of-the-art methods across five widely-used polyp segmentation benchmarks. The code will be publicly available at: https://github.com/CSYSI/ASGNet. Hengmin Zhang, Jianjun Qian, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | FeverMamba: Highlight Fever in Crowd via Mamba for Remote Fever ScreeningabstractRemote fever screening based on thermal infrared images can screen feverish faces in real time and play an important role in vital sign monitoring. Screening methods that rely on short-term crowd facial temperature differences can overcome environmental effects without the need for additional sensors, but effectively encoding these differences to highlight fever face remains a challenge. To this end, we develop a fever screening framework based on Mamba, which exploits the context capturing capability of Mamba to model the temperature differences of thermal infrared face images. Furthermore, considering that local contextual associations in the crowd will limit the construction of global differences, we propose a shuffle scanning method to break the local contextual associations and construct multiple shuffle scanning to achieve global difference representation. In addition, we design a temperature-aware self-supervised loss function to cope with the situation where data of fever faces is difficult to collect. Finally, we achieve state-of-the-art performance in both supervised and self-supervised cases on the thermal infrared face dataset. Mengkai Yan, Jianjun Qian, Jindi Bao, Jian Yang 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Dual Manifold Regularization Steered Robust Representation Learning for Point Cloud AnalysisabstractWith the rapid advancement of 3D scanning technology, point clouds have become a crucial data type in computer vision and machine learning. However, learning robust representations for point clouds remains a significant challenge due to their irregularity and sparsity. In this paper, we propose a novel Dual Manifold Regularization (DMR) framework that makes full use of the properties of positive and negative curvature in manifolds to improve the representation of point clouds. Specifically, we leverage DMR based on hyperbolic and hyperspherical manifolds to address the limitations of traditional single-manifold regularization techniques, including inadequate generalization ability and adaptability to data diversity, as well as the difficulty of capturing complex relationships between data. To begin, we utilize the tree-like structure of the hyperbolic manifold to model the part-whole hierarchical relationships within point clouds. This allows for a more comprehensive representation of the data, improving the model's capability to understand complex shapes. Additionally, we construct positive samples through topological consistency augmentation and employ contrastive learning techniques in the hyperspherical manifold to capture more discriminative features within the data. Our experimental results show that our method outperforms traditional supervised learning and single-manifold regularization techniques in point cloud analysis. Specifically, for shape classification, DMR achieves a new State-Of-The-Art (SOTA) performance with 94.8% Overall Accuracy (OA) on ModelNet40 and 90.7% OA on ScanObjectNN, surpassing the recent SOTA model without increasing the baseline parameters. Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
AAAI | 3 |
| 2025 | Learning Generalized Residual Exchange-Correlation-Uncertain Functional for Density Functional TheoryabstractDensity Functional Theory (DFT) stands as a widely used and efficient approach for addressing the many-electron Schrödinger equation across various domains such as physics, chemistry, and biology. However, a core challenge that persists over the long term pertains to refining the exchange-correlation (XC) approximation. This approximation significantly influences the triumphs and shortcomings observed in DFT applications. Nonetheless, a prevalent issue among XC approximations is the presence of systematic errors, stemming from deviations from the mathematical properties of the exact XC functional. For example, although both B3LYP and DM21 (DeepMind 21) exhibit improvements over previous benchmarks, there is still potential for further refinement. In this paper, we propose a strategy for enhancing XC approximations by estimating the neural uncertainty of the XC functional, named Residual XC-Uncertain Functional. Specifically, our approach involves training a neural network to predict both the mean and variance of the XC functional, treating it as a Gaussian distribution. To ensure stability in each sampling point, we construct the mean by combining traditional XC approximations with our neural predictions, mitigating the risk of divergence or vanishing values. It is crucial to highlight that our methodology excels particularly in cases where systematic errors are pronounced. Empirical outcomes from three benchmark tests substantiate the superiority of our approach over existing state-of-the-art methods. Our approach not only surpasses related techniques but also significantly outperforms both the popular B3LYP and the recent DM21 methods, achieving average RMSE improvements of 62% and 37%, respectively, across the three benchmarks: W4-17, G21EA, and G21IP. Sizhuo Jin, Jianjun Qian, Ying Tai |
AAAI | 3 |
| 2025 | NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationabstractLearning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long-range dependencies across spatio-temporal cues, which leads to suboptimal state perception and potential decision failures. In this paper, we introduce NaviFormer, an effective encoder-decoder navigation transformer, to aggregate discriminative spatio-temporal context information for object navigation. Our navigation encoder not only encodes spatial layouts and temporal agent poses but also innovatively constructs and encodes a passable frontier map, enriching the original state encoding with cues of potential exploration regions. Furthermore, our navigation decoder employs spatio-temporal self-attention and cross-attention mechanisms to model the dependencies among spatial layout encoding, temporal pose encoding, and passable frontier encoding, thereby facilitating comprehensive contextual state feature aggregation. Finally, we leverage these learned spatio-temporal contextual state representations for PPO-based navigation decisions. Extensive experiments on the Gibson, Habitat-Matterport3D (HM3D) and Matterport3D (MP3D) datasets demonstrate the superiority of our approach. Wei Xie 0019, Haobo Jiang, Yun Zhu 0011, Jianjun Qian, Jin Xie 0001 |
AAAI | 4 |
| 2025 | Remote Photoplethysmography in Real-World and Extreme Lighting ScenariosabstractPhysiological activities can be manifested by the sensitive changes in facial imaging. While they are barely observable to our eyes, computer vision manners can, and the derived remote photoplethysmography (rPPG) has shown considerable promise. However, existing studies mainly rely on spatial skin recognition and temporal rhythmic interactions, so they focus on identifying explicit features under ideal light conditions, but perform poorly in-the-wild with intricate obstacles and extreme illumination exposure. In this paper, we propose an end-to-end video transformer model for rPPG. It strives to eliminate complex and unknown external time-varying interferences, whether they are sufficient to occupy subtle biosignal amplitudes or exist as periodic perturbations that hinder network training. In the specific implementation, we utilize global interference sharing, subject background reference, and self-supervised disentanglement to eliminate interference, and further guide learning based on spatiotemporal filtering, reconstruction guidance, and frequency domain and biological prior constraints to achieve effective rPPG. To the best of our knowledge, this is the first robust rPPG model for real outdoor scenarios based on natural face videos, and is lightweight to deploy. Extensive experiments show the competitiveness and performance of our model in rPPG prediction across datasets and scenes. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shuo Chen 0003, Jian Yang 0003 |
CVPR | 3 |
| 2025 | WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusionabstract3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model, and the fidelity of the generated data is quite limited. This paper presents WeatherGen, the first unified diverse-weather LiDAR data diffusion generation framework, significantly improving fidelity. Specifically, we first design a map-based data producer, which can provide a vast amount of high-quality diverse-weather data for training purposes. Then, we utilize the diffusion-denoising paradigm to construct a diffusion model. Among them, we propose a spider mamba generator to restore the disturbed diverse weather data gradually. The spider mamba models the feature interactions by scanning the Li-Dar beam circle or central ray, excellently maintaining the physical structure of the LiDAR data. Subsequently, following the generator to transfer real-world knowledge, we design a latent feature aligner. Afterward, we devise a contrastive learning-based controller, which equips weather control signals with compact semantic knowledge through language supervision, guiding the diffusion model to generate more discriminative data. Extensive evaluations demonstrate the high generation quality of WeatherGen. Through WeatherGen, we construct the mini-weather dataset, promoting the performance of the downstream task under adverse weather conditions. Code is available: https://github.com/wuyang98/weathergen Yun Zhu 0011, Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
CVPR | 4 |
| 2025 | Learning Class Prototypes for Unified Sparse-Supervised 3D Object DetectionabstractBoth indoor and outdoor scene perceptions are essential for embodied intelligence. However, current sparse supervised 3D object detection methods focus solely on outdoor scenes without considering indoor settings. To this end, we propose a unified sparse supervised 3D object detection method for both indoor and outdoor scenes through learning class prototypes to effectively utilize unlabeled objects. Specifically, we first propose a prototype-based object mining module that converts the unlabeled object mining into a matching problem between class prototypes and unlabeled features. By using optimal transport matching results, we assign prototype labels to high-confidence features, thereby achieving the mining of unlabeled objects. We then present a multi-label cooperative refinement module to effectively recover missed detections through pseudo label quality control and prototype label cooperation. Experiments show that our method achieves state-of-the-art performance under the one object per scene sparse supervised setting across indoor and outdoor datasets. With only one labeled object per scene, our method achieves about 78%, 90%, and 96% performance compared to the fully supervised detector on ScanNet V2, SUN RGB-D, and KITTI, respectively, highlighting the scalability of our method. Code is available at https://github.com/zyrant/CPDet3D. Yun Zhu 0011, Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
CVPR | 4 |
| 2025 | GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse Views
Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
ICCV | 3 |
| 2025 | Rethinking Point Cloud Data Augmentation: Topologically Consistent DeformationabstractData augmentation has been widely used in machine learning. Its main goal is to transform and expand the original data using various techniques, creating a more diverse and enriched training dataset. However, due to the disorder and irregularity of point clouds, existing methods struggle to enrich geometric diversity and maintain topological consistency, leading to imprecise point cloud understanding. In this paper, we propose SinPoint, a novel method designed to preserve the topological structure of the original point cloud through a homeomorphism. It utilizes the Sine function to generate smooth displacements. This simulates object deformations, thereby producing a rich diversity of samples. In addition, we propose a Markov chain Augmentation Process to further expand the data distribution by combining different basic transformations through a random process. Our extensive experiments demonstrate that our method consistently outperforms existing Mixup and Deformation methods on various benchmark point cloud datasets, improving performance for shape classification and part segmentation tasks. Specifically, when used with PointNet++ and DGCNN, our method achieves a state-of-the-art accuracy of 90.2 in shape classification with the real-world ScanObjectNN dataset. We release the code at https://github.com/CSBJian/SinPoint. Jian Bi, Qianliang Wu, Xiang Li 0041, Shuo Chen 0003, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
ICML | 5 |
| 2025 | Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing ImagesabstractAutomatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former. Jiexi Yan, Jianjun Qian, Chunyan Xu, Jian Yang 0003, Lei Luo 0001 |
IJCAI | 3 |
| 2025 | Self-Supervised Vision Graph Neural Networks Based on Contrastive LearningabstractIn the field of computer vision, Vision Graph Neural Networks (ViG) have demonstrated significant potential in image understanding. By treating the divided image patches as nodes and constructing connection relationships based on neighbor attributes, ViG can efficiently model global dependencies within images with the help of graph attributes. However, most existing ViG methods have the problem of high computational complexity in graph construction and may not be able to effectively and fully explore the graph structure information. Besides, the heavy reliance on manual annotation labels limits the application potential of ViG in practical scenarios. To this end, in this paper, we propose a novel self-supervised vision graph contrastive learning method (S2ViG) based on image mixing strategy for efficient vision graph representation learning. It aims to use self-supervised method to alleviate the dependence on manual annotation and enhance the understanding of the global structure of the graph using two different vision graph construction methods. Specifically, we first employ image mixing strategy to uncover latent semantic relationships among multiple images. Then, we construct dynamic graph structures for image patches from local and global perspectives to obtain augmented contrastive samples. Finally, the multilevel contrastive loss is constructed to optimize the network. Experimental results show that our method achieves excellent performance on multiple datasets such as ImageNet-1K and CIFAR. Yuehui Han, Jianjun Qian, Jian Yang 0003 |
ACM Multimedia | 3 |
| 2025 | Cross-View Geometric Collaboration for Generalizable Sparse View Neural Surface ReconstructionabstractGeneralizable neural implicit surface reconstruction aims to recover accurate surfaces with sparse views from unseen scenes. Most existing methods suffer from severe incompleteness and inaccuracies in the case of reconstruction with large viewpoint variations, as significant perspective distortions across views lead to unreliable feature correspondence and geometry representations. In this paper, we propose a cross-view geometric collaboration framework for generalizable neural surface reconstruction, which exploits cross-view complementary geometric information to improve the accuracy and robustness of reconstruction from sparse views. Specifically, we propose a cross-view geometry complement module that utilizes the reliable geometric information of different views to refine geometric representations. In addition, we construct a distortion-robust patch-based consistency volume to provide supplementary geometric cues for uncertain regions. For the rendering process, we develop a cross-view geometry transformer to adaptively aggregate reliable cross-view point features by considering geometric context along the ray. Finally, we render per-view depth maps and fuse them to reconstruct the final surface. Extensive experimental results on the DTU, BlendedMVS, and Tanks and Temples datasets demonstrate the superior reconstruction quality and view-combination generalizability of our solution. Le Hui, Jianjun Qian, Jian Yang 0003, Yigong Zhang, Jin Xie 0001 |
ACM Multimedia | 3 |
| 2025 | AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group SwappingabstractFusing cross-category objects to a single coherent object has gained increasing attention in text-to-image (T2I) generation due to its broad applications in virtual reality, digital media, film, and gaming. However, existing methods often produce biased, visually chaotic, or semantically inconsistent results due to overlapping artifacts and poor integration. Moreover, progress in this field has been limited by the absence of a comprehensive benchmark dataset. To address these problems, we propose Adaptive Group Swapping (AGSwap), a simple yet highly effective approach comprising two key components: (1) Group-wise Embedding Swapping, which fuses semantic attributes from different concepts through feature manipulation, and (2) Adaptive Group Updating, a dynamic optimization mechanism guided by a balance evaluation score to ensure coherent synthesis. Additionally, we introduce Cross-category Object Fusion (COF), a large-scale, hierarchically structured dataset built upon ImageNet-1K and WordNet. COF includes 95 superclasses, each with 10 subclasses, enabling 451,250 unique fusion pairs. Extensive experiments demonstrate that AGSwap outperforms state-of-the-art compositional T2I methods, including GPT-Image-1 using simple and complex prompts. Project Page Zedong Zhang, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
SIGGRAPH Asia | 3 |
| 2025 | RagNet3D: Learning distinguishable representation for pooled grids in 3D object detection
Jiaxin Chen 0001, Yuehui Han, Zhiqiang Yan 0001, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
Neurocomputing | 4 |
| 2025 | Cascading enhancement representation for face anti-spoofing
Yimei Ma, Yangwei Dong, Jianjun Qian, Jian Yang 0003 |
Pattern Recognit. Lett. | 3 |
| 2025 | SVD-Guided Multimodal Feature Fusion for Emotion Recognition From Facial VideosabstractMultimodal emotion recognition based on facial videos aims to extract features from different modalities to identify human emotions. The previous work focus on designing various fusion schemes to combine heterogeneous modal data. However, most studies have overlooked the role of different modalities in emotion recognition and have not fully utilized the intrinsic connections between modalities. Furthermore, the multimodal data from facial videos also contain various distractions bad for emotion analysis. How to reduce the impact of distractions and enable a model to mine effective information for emotion recognition from different modalities is still a challenge problem. To address above issue, we propose a SVD-guided multimodal feature fusion method based on facial video for emotion recognition, which uses a hierarchical fusion mechanism and adopts different loss strategies at each level to learn multimodal feature representation. Specifically, we fuse the facial expression and rPPG signal (or Point-of-Gaze) by using the weak supervision strategy and contrastive learning. Subsequently, the fused feature of facial expression and rPPG signal and the fused feature of facial expression and Point-of-Gaze are combined together to construct the unified multimodal feature matrix. Based on this, Singular Value Decomposition (SVD) is used to refine the redundancy information caused by the multimodal fusion and guide the neural network to learn discriminative emotion feature. At the same time, a consistent loss is developed to enhance the multimodal representation. Experiments on three public datasets show that the proposed method achieves better results over the compared methods. Jindi Bao, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Non-Aligned Supervision for Real Image DehazingabstractRemoving haze from real-world images is challenging due to unpredictable weather conditions, resulting in the misalignment of hazy and clear image pairs. In this paper, we propose an innovative dehazing framework that operates under non-aligned supervision. This framework is grounded in the atmospheric scattering model, and consists of three interconnected networks: dehazing, airlight, and transmission networks. In particular, we explore a non-alignment scenario that a clear reference image, unaligned with the input hazy image, is utilized to supervise the dehazing network. To implement this, we present a multi-scale reference loss that compares the feature representations between the referred image and the dehazed output. Our scenario makes it easier to collect hazy/clear image pairs in real-world environments, even under conditions of misalignment and shift views. To showcase the effectiveness of our scenario, we have collected a new hazy dataset including 415 image pairs captured by mobile Phone in both rural and urban areas, called "Phone-Hazy". Furthermore, we introduce a self-attention network based on mean and variance for modeling real infinite airlight, using the dark channel prior as positional guidance. Experimental results demonstrate the superior performance of our framework over existing state-of-the-art techniques in the real-world image dehazing task. Phone-Hazy and code will be available at https://fanjunkai1.github.io/projectpage/NSDNet/index.html. Junkai Fan, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Self-Supervised Temperature Representation Learning for Fever ScreeningabstractUtilizing thermal infrared facial imaging for fever screening in public spaces has become a common strategy to curb the spread of influenza viruses. However, it is difficult to capture larger number of faces with fever labels, which makes learning facial temperature representation extremely difficult. To overcome this limitation, we propose a self-supervised fever screening framework (SelfFS) to learn temperature representation from infrared face images. Specifically, SelfFS employs rate reduction theory to guide the network to focus on temperature features by expanding the coding rate of faces with different temperatures and compressing the coding rate of faces with the same temperature but different appearances. Furthermore, we impose sparsity constraints on the network parameters, which facilitates the extraction of simple temperature features with a limited number of neurons while filtering complex appearance features. Experiments demonstrate that our SelfFS framework outperforms existing fever screening techniques and achieves the comparable results with the supervised methods. Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 2 |
| 2025 | Daytime-Mixed Non-Aligned Learning for Real Nighttime Image EnhancementabstractEnhancing real nighttime images is a significant challenge due to the deterioration of visual quality caused by limited perceptibility under adverse illumination conditions, leading to loss of details and color deviation. In this paper, we propose a novel nighttime image enhancement framework using daytime-mixed non-aligned supervision. It aims to couple the information between non-aligned daytime and nighttime image pairs. Specifically, our framework consists of a simple yet effective daytime-mixed supervised learning phase and a Retinex-based reconstruction phase. In the first phase, we employ a multi-instance with adaptive information fusion (AIF) module integrated within a UNet enhancement network called MIFUNet, which is trained via a daytime-mixed supervised loss. In the second phase, the Retinex-based reconstruction employs both a light-effect estimation network and an illumination adjustment network to restore the nighttime image, guided by physical principles. To evaluate the effectiveness of our approach, we collect a real non-aligned day-night dataset named the NANE dataset, which contains 748 non-aligned image pairs and 100 nighttime images solely for testing. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art image enhancement methods. Jiangwei Weng, Junkai Fan, Jianjun Qian, Haiyang Zou, Ying Tai, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Video-Based Multiphysiological Disentanglement and Remote Robust Estimation for RespirationabstractRemote noncontact respiratory rate estimation by facial visual information has great research significance, providing valuable priors for health monitoring, clinical diagnosis, and anti-fraud. However, existing studies suffer from disturbances in epidermal specular reflections induced by head movements and facial expressions. Furthermore, diffuse reflections of light in the skin-colored subcutaneous tissue caused by multiple time-varying physiological signals independent of breathing are entangled with the intention of the respiratory process, leading to confusion in current research. To address these issues, this article proposes a novel network for natural light video-based remote respiration estimation. Specifically, our model consists of a two-stage architecture that progressively implements vital measurements. The first stage adopts an encoder-decoder structure to recharacterize the facial motion frame differences of the input video based on the gradient binary state of the respiratory signal during inspiration and expiration. Then, the obtained generative mapping, which is disentangled from various time-varying interferences and is only linearly related to the respiratory state, is combined with the facial appearance in the second stage. To further improve the robustness of our algorithm, we design a targeted long-term temporal attention module and embed it between the two stages to enhance the network's ability to model the breathing cycle that occupies ultra many frames and to mine hidden timing change clues. We train and validate the proposed network on a series of publicly available respiration estimation datasets, and the experimental results demonstrate its competitiveness against the state-of-the-art breathing and physiological prediction frameworks. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shangbing Gao, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Driving-Video Dehazing with Non-Aligned Regularization for Safety AssistanceabstractReal driving-video dehazing poses a significant challenge due to the inherent difficulty in acquiring precisely aligned hazy/clear video pairs for effective model training, especially in dynamic driving scenarios with unpredictable weather conditions. In this paper, we propose a pioneering approach that addresses this challenge through a nonaligned regularization strategy. Our core concept involves identifying clear frames that closely match hazy frames, serving as references to supervise a video dehazing network. Our approach comprises two key components: reference matching and video dehazing. Firstly, we introduce a non-aligned reference frame matching module, leveraging an adaptive sliding window to match high-quality reference frames from clear videos. Video dehazing incorporates flow-guided cosine attention sampler and deformable cosine attention fusion modules to enhance spatial multi-frame alignment and fuse their improved information. To validate our approach, we collect a GoProHazy dataset captured effortlessly with GoPro cameras in diverse rural and urban road environments. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art methods in the challenging task of real driving-video dehazing. Project page. Junkai Fan, Jiangwei Weng, Kun Wang 0042, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
CVPR | 5 |
| 2024 | Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning
Yuehui Han, Can Xu 0006, Rui Xu 0021, Jianjun Qian, Jin Xie 0001 |
ECCV (76) | 4 |
| 2024 | Text2LiDAR: Text-Guided LiDAR Point Cloud Generation via Equirectangular Transformer
Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
ECCV (56) | 3 |
| 2024 | Dense Voxel Representation Network for Implicit Scene CompletionabstractImplicit scene completion aims to learn an implicit representation of dense point clouds from incomplete ones. Since point clouds are disordered and irregular, some implicit scene completion methods learn representations from voxelized point clouds with sparse convolution. Despite achieving promising results, they lack deep exploration of feature learning on empty voxels, which is beneficial for implicit scene completion task. To address this, we propose a dense voxel representation network for implicit scene completion. First, we design a Bird’s-Eye View (BEV) assisted enhancement module to enhance non-empty voxel features by incorporating the information contained in the learned dense BEV features into them through deformable cross-attention. Second, we construct a feature adaptive completion module to adaptively complete voxel features using deformable self-attention, realizing the transfer of the information from non-empty voxels to empty voxels. Extensive experiments on SemanticKITTI and SemanticPOSS datasets demonstrate our method achieves state-of-the-art performance. Fan Dai, Yun Zhu 0011, Yaqi Shen, Jin Xie 0001, Jianjun Qian |
ICME | 5 |
| 2024 | ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceabstractDiffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framework for fully consistent and high-fidelity talking avatar generation. Instead of directly employing multi-modal conditions to the diffusion process, our method learns to first model the temporal representation for stability between adjacent frames. Specifically, we propose a Temporally-Sensitive Detail (TSD) map containing high-frequency feature and contours that vary significantly along the time axis. Using a temporal consistent diffusion module, we learn to align TSD of the initial result to that of the video frame ground truth. The final avatar is generated by a fully consistent diffusion module, conditioned on the aligned TSD, rough head normal, and emotion prompt embedding. We find that the aligned TSD, which represents the temporal patterns, constrains the diffusion process to generate temporally stable talking head. Further, its reliable guidance complements the inaccuracy of other conditions, suppressing the accumulated error while improving the consistency on various aspects. Extensive experiments demonstrate that ConsistentAvatar outperforms the state-of-the-art methods on the generated appearance, 3D, expression and temporal consistency. Haijie Yang, Zhenyu Zhang 0005, Hao Tang 0005, Jianjun Qian, Jian Yang 0003 |
ACM Multimedia | 4 |
| 2024 | MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State SpaceabstractRecent advances in low light image enhancement have been dominated by Retinex-based learning framework, leveraging convolutional neural networks (CNNs) and Transformers. However, the vanilla Retinex theory primarily addresses global illumination degradation and neglects local issues such as noise and blur in dark conditions. Moreover, CNNs and Transformers struggle to capture global degradation due to their limited receptive fields. While state space models (SSMs) have shown promise in the long-sequence modeling, they face challenges in combining local invariants and global context in visual data. In this paper, we introduce MambaLLIE, an implicit Retinex-aware low light enhancer featuring a global-then-local state space design. We first propose a Local-Enhanced State Space Module (LESSM) that incorporates an augmented local bias within a 2D selective scan mechanism, enhancing the original SSMs by preserving local 2D dependency. Additionally, an Implicit Retinex-aware Selective Kernel module (IRSK) dynamically selects features using spatially-varying operations, adapting to varying inputs through an adaptive kernel selection process. Our Global-then-Local State Space Block (GLSSB) integrates LESSM and IRSK with layer normalization (LN) as its core. This design enables MambaLLIE to achieve comprehensive global long-range modeling and flexible local feature aggregation. Extensive experiments demonstrate that MambaLLIE significantly outperforms state-of-the-art CNN and Transformer-based methods. Our code is available at https://github.com/wengjiangwei/MambaLLIE. Jiangwei Weng, Zhiqiang Yan 0001, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
NeurIPS | 4 |
| 2024 | Learning Fully Parametric Subspace Clustering
Xuanrong Chen, Jianjun Qian, Shuo Chen 0003, Jian Yang 0003, Jun Li 0027 |
PRCV (1) | 2 |
| 2024 | Learning Robust Facial Representation From the View of Diversity and Closeness
Chaoyu Zhao, Jianjun Qian, Shumin Zhu, Jin Xie 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 2 |
| 2024 | Contrastive subspace distribution learning for novel category discovery in high-dimensional visual data
Shuai Wei, Shuo Chen 0003, Jian Yang 0003, Jianjun Qian, Jun Li 0027 |
Knowl. Based Syst. | 5 |
| 2024 | Dual feature disentanglement for face anti-spoofing
Yimei Ma, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2024 | FeverNet: Enabling accurate and robust remote fever screening
Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2024 | TranPhys: Spatiotemporal Masked Transformer Steered Remote Photoplethysmography EstimationabstractSubtle variations are invisible to the naked eyes in human physiological signals can reflect important biological and health indicators. Although numerous computer vision methods have been proposed to recover and magnify these changes, most of them either only focus on identifying and recognizing explicit features such as shapes and textures, or are weak in long-term temporal modeling and spatiotemporal interactive perception of implicit biometrics. Therefore, it is difficult for them to robustly overcome various disturbances that affect detection performance. To address these issues, this paper presents TranPhys, a novel remote photoplethysmography (rPPG) network for facial video-based heart rate estimation. Specifically, first, we argue that facial subregions vary over time due to their biological personalities. So we split the input face video into multiple spatiotemporal tubes, build the 3D vision transformer with encoders and decoders to adequately model the high-dimensional representations of the respective regulars in each subregion, and globally coordinate their feedback on the cardiac pulsing waveform. Second, we design the temporal pooling attention to more finely mine the subtle changes hidden in the skin color over time and their long-term contextual rhythm cues. Third, we leverage the self-supervised masked autoencoding paradigm to overcome redundancy to enhance the robustness of our model, and construct the targeted spatiotemporal sampling maps instead of raw input sequences as the pretrained constraint labels to fully inspire self-supervision. We train, validate, and practice our TranPhys on multiple public datasets to demonstrate that our method achieves the competitive performance in remote heart rate estimation. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Linear Regression Problem Relaxations Solved by Nonconvex ADMM With Convergence AnalysisabstractIn this work, we focus on studying the differentiable relaxations of several linear regression problems, where the original formulations are usually both nonsmooth with one nonconvex term. Unfortunately, in most cases, the standard alternating direction method of multipliers (ADMM) cannot guarantee global convergence when addressing these kinds of problems. To address this issue, by smoothing the convex term and applying a linearization technique before designing the iteration procedures, we employ nonconvex ADMM to optimize challenging nonconvex-convex composite problems. In our theoretical analysis, we prove the boundedness of the generated variable sequence and then guarantee that it converges to a stationary point. Meanwhile, a potential function is derived from the augmented Lagrange function, and we further verify that the objective function is monotonically nonincreasing. Under the Kurdyka-Łojasiewicz (KŁ) property, the global convergence is analyzed step by step. Finally, experiments on face reconstruction, image classification, and subspace clustering tasks are conducted to show the superiority of our algorithms over several state-of-the-art ones. Hengmin Zhang, Junbin Gao, Jianjun Qian, Jian Yang 0003, Chunyan Xu, Bob Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Efficient Image Classification via Structured Low-Rank Matrix Factorization RegressionabstractIn real-world applications involving sparse coding and low-rank matrix recovery problems, linear regression methods usually struggle to effectively capture the structured correlations present in data matrices. This limitation arises from representation approaches that treat images as vectors and handle testing samples individually, overlooking these correlations. To address these challenges, we propose a novel approach that leverages the low-rank property to capture the global and intrinsic structure of residual and coefficient matrices, departing from the assumption of independent and identically distributed (I.I.D) data. Our method introduces nonconvex and nonsmooth low-rank matrix regression models guided by the extended matrix variate power exponential distribution (M.P.E.D). By incorporating factorization strategies into the regression coefficient matrix and utilizing the Schatten-$p$norm with three distinct values of$p$, we enhance computational efficiency. Our formulation enables efficient subproblem solving through the introduction of auxiliary variables and the use of singular value threshold operators. We achieve closed-form solutions using the proposed multi-variable alternating direction method of multipliers (ADMM). Theoretical analysis establishes the local convergence properties and computational complexity of our optimization algorithm. Furthermore, we conduct numerical experiments on various image datasets, including face, object, and digital, to demonstrate the superior performance and computational efficiency of our methods compared to several related regression approaches. The source codes for our method are available athttps://github.com/ZhangHengMin/TIFS_SLRMFR. Hengmin Zhang, Jian Yang 0003, Jianjun Qian, Guangwei Gao, Xiangyuan Lan, Zhiyuan Zha, Bihan Wen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Learning Structured Relation Embeddings for Fine-Grained Fashion Attribute RecognitionabstractFashion attribute recognition is a not-new topic, but rather a core task in understanding fashion from the perspective of computer vision. This article proposes a structured relation-aware network (sRA-Net), which exploits multiple hidden relations in fashion images to enrich and achieve accurate attribute representations to boost the performance of fashion attribute recognition. Specifically, it deconstructs the features of a clothing fashion item into three levels, including low-level attribute-related image region information, mid-level attribute dependency information, and high-level clothing look information. To learn these multi-relational embeddings, we present three relation-aware attention mechanisms. The attribute attention mechanism describes the relationship among different attribute vectors through self-attention and uses the attention map to update the attribute embedding. Then, the spatial attention mechanism associates the attribute with the image features and enhances the attribute embedding by leveraging the attribute-related image region. Finally, the channel attention mechanism selects attribute-related image feature channels to obtain a more fine-grained attribute embedding. Furthermore, we introduce structure-aware embedding to constrain attribute recognition in images from a global perspective by identifying the inner structure of the clothing. Without bells and whistles, sRA-Net outperforms all state-of-the-art attribute recognition methods on two mainstream fashion attribute datasets, namely the DeepFashion-C dataset and iFashion-Attribute dataset, with over 1%-3% improvement. Shumin Zhu, Xingxing Zou, Jianjun Qian, Wai Keung Wong |
IEEE Trans. Multim. | 3 |
| 2024 | Unified Framework for Faster Clustering via Joint Schatten p-Norm Factorization With Optimal MeanabstractTo enhance the effectiveness and efficiency of subspace clustering in visual tasks, this work introduces a novel approach that automatically eliminates the optimal mean, which is embedded in the subspace clustering framework of low-rank representation (LRR) methods, along with the computationally factored formulation of Schatten p -norm. By addressing the issues related to meaningful computations involved in some LRR methods and overcoming biased estimation of the low-rank solver, we propose faster nonconvex subspace clustering methods through joint Schatten p -norm factorization with optimal mean (JS p NFOM), forming a unified framework for enhancing performance while reducing time consumption. The proposed approach employs tractable and scalable factor techniques, which effectively address the disadvantages of higher computational complexity, particularly when dealing with large-scale coefficient matrices. The resulting nonconvex minimization problems are reformulated and further iteratively optimized by multivariate weighting algorithms, eliminating the need for singular value decomposition (SVD) computations in the developed iteration procedures. Moreover, each subproblem can be guaranteed to obtain the closed-form solver, respectively. The theoretical analyses of convergence properties and computational complexity further support the applicability of the proposed methods in real-world scenarios. Finally, comprehensive experimental results demonstrate the effectiveness and efficiency of the proposed nonconvex clustering approaches compared to existing state-of-the-art methods on several publicly available databases. The demonstrated improvements highlight the practical significance of our work in subspace clustering tasks for visual data analysis. The source code for the proposed algorithms is publicly accessible at https://github.com/ZhangHengMin/TRANSUFFC. Hengmin Zhang, Jiaoyan Zhao, Bob Zhang 0001, Chen Gong 0002, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Recurrent Structure Attention Guidance for Depth Super-resolutionabstractImage guidance is an effective strategy for depth super-resolution. Generally, most existing methods employ hand-crafted operators to decompose the high-frequency (HF) and low-frequency (LF) ingredients from low-resolution depth maps and guide the HF ingredients by directly concatenating them with image features. However, the hand-designed operators usually cause inferior HF maps (e.g., distorted or structurally missing) due to the diverse appearance of complex depth maps. Moreover, the direct concatenation often results in weak guidance because not all image features have a positive effect on the HF maps. In this paper, we develop a recurrent structure attention guided (RSAG) framework, consisting of two important parts. First, we introduce a deep contrastive network with multi-scale filters for adaptive frequency-domain separation, which adopts contrastive networks from large filters to small ones to calculate the pixel contrasts for adaptive high-quality HF predictions. Second, instead of the coarse concatenation guidance, we propose a recurrent structure attention block, which iteratively utilizes the latest depth estimation and the image features to jointly select clear patterns and boundaries, aiming at providing refined guidance for accurate depth recovery. In addition, we fuse the features of HF maps to enhance the edge structures in the decomposed LF maps. Extensive experiments show that our approach obtains superior performance compared with state-of-the-art depth super-resolution methods. Our code is available at: https://github.com/Yuanjiayii/DSR-RSAG. Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
AAAI | 4 |
| 2023 | Structure Flow-Guided Network for Real Depth Super-resolutionabstractReal depth super-resolution (DSR), unlike synthetic settings, is a challenging task due to the structural distortion and the edge noise caused by the natural degradation in real-world low-resolution (LR) depth maps. These defeats result in significant structure inconsistency between the depth map and the RGB guidance, which potentially confuses the RGB-structure guidance and thereby degrades the DSR quality. In this paper, we propose a novel structure flow-guided DSR framework, where a cross-modality flow map is learned to guide the RGB-structure information transferring for precise depth upsampling. Specifically, our framework consists of a cross-modality flow-guided upsampling network (CFUNet) and a flow-enhanced pyramid edge attention network (PEANet). CFUNet contains a trilateral self-attention module combining both the geometric and semantic correlations for reliable cross-modality flow learning. Then, the learned flow maps are combined with the grid-sampling mechanism for coarse high-resolution (HR) depth prediction. PEANet targets at integrating the learned flow map as the edge attention into a pyramid network to hierarchically learn the edge-focused guidance feature for depth edge refinement. Extensive experiments on real and synthetic DSR datasets verify that our approach achieves excellent performance compared to state-of-the-art methods. Our code is available at: https://github.com/Yuanjiayii/DSR-SFG. Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
AAAI | 4 |
| 2023 | Graph Spectral Perturbation for 3D Point Cloud Contrastive Learningabstract3D point cloud contrastive learning has attracted increasing attention due to its efficient learning ability. By distinguishing the similarity relationship between positive and negative samples in the feature space, it can learn effective point cloud feature representations without manual annotation. However, most point cloud contrastive learning methods construct contrastive samples by perturbing point clouds in data space or introducing multi-modality/format data, which may be difficult to control the intensity of the perturbation or introduce interference from different modalities/formats. To this end, in this paper, we propose a novel graph spectral perturbation based contrastive learning framework (GSPCon) for efficient and robust self-supervised 3D point cloud representation learning. It aims to perform perturbations in the graph spectral domain to construct contrastive samples of the point cloud. Specifically, we first naturally represent the point cloud as a k-nearest neighbors (KNN) graph, and adaptively transform the coordinates of the points into the graph spectral domain based on the graph Fourier transform (GFT). Then we implement data augmentation in the graph spectral domain by perturbing the spectral representations. Finally, the contrastive samples are generated by employing the inverse graph Fourier transform (IGFT) to transform the augmented spectral representations back to the point clouds. Experimental results show that our method achieves the state-of-the-art performance on various downstream tasks. Source code is available at https://github.com/yh-han/GSPCon.git. Yuehui Han, Jiaxin Chen 0001, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 3 |
| 2023 | Transformer-based Point Cloud Generation NetworkabstractPoint cloud generation is an important research topic in 3D computer vision, which can provide high-quality datasets for various downstream tasks. However, efficiently capturing the geometry of point clouds remains a challenging problem due to their irregularities. In this paper, we propose a novel transformer-based 3D point cloud generation network to generate realistic point clouds. Specifically, we first develop a transformer-based interpolation module that utilizes k-nearest neighbors at different scales to learn global and local information about point clouds in the feature space. Based on geometric information, we interpolate new point features to upsample the point cloud features. Then, the upsampled features are used to generate a coarse point cloud with spatial coordinate information. We construct a transformer-based refinement module to enhance the upsampled features in feature space with geometric information in coordinate space. Finally, we use a multi-layer perceptron on the upsampled features to generate the final point cloud. Extensive experiments on ShapeNet and ModelNet demonstrate the effectiveness of our proposed method. Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 4 |
| 2023 | Scene Graph Masked Variational Autoencoders for 3D Scene GenerationabstractGenerating realistic 3D indoor scenes requires a deep understanding of objects and their spatial relationships. However, existing methods often fail to generate realistic 3D scenes due to the limited understanding of object relationships. To tackle this problem, we propose a Scene Graph Masked Variational Auto-Encoder (SG-MVAE) framework that fully captures the relationships between objects to generate more realistic 3D scenes. Specifically, we first introduce a relationship completion module that adaptively learns the missing relationships between objects in the scene graph. To accurately predict the missing relationships, we employ multi-group attention to capture the correlations between the objects with missing relationships and other objects in the scene. After obtaining the complete scene relationships, we mask the relationships between objects and use a decoder to reconstruct the scene. The reconstruction process enhances the model's understanding of relationships, generating more realistic scenes. Extensive experiments on benchmark datasets show that our model outperforms state-of-the-art methods. Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 4 |
| 2023 | OTFace: Hard Samples Guided Optimal Transport Loss for Deep Face RepresentationabstractFace representation in the wild is extremely hard due to the large scale face variations. Some deep convolutional neural networks (CNNs) have been developed to learn discriminative feature by designing properly margin-based losses, which perform well on easy samples but fail on hard samples. Although some methods mainly adjust the weights of hard samples in training stage to improve the feature discrimination, they overlook the distribution property of feature. It is worth noting that the miss-classified hard samples may be corrected from the feature distribution view. To overcome this problem, this paper proposes the hard samples guided optimal transport (OT) loss for deep face representation, OTFace in short. OTFace aims to enhance the performance of hard samples by introducing the feature distribution discrepancy while maintaining the performance on easy samples. Specifically, we embrace triplet scheme to indicate hard sample groups in one mini-batch during training. OT is then used to characterize the distribution differences of features from the high level convolutional layer. Finally, we integrate the margin-based-softmax (e.g. ArcFace or AM-Softmax) and OT together to guide deep CNN learning. Extensive experiments were conducted on several benchmark databases. The quantitative results demonstrate the advantages of the proposed OTFace over state-of-the-art methods. Jianjun Qian, Shumin Zhu, Chaoyu Zhao, Jian Yang 0003, Wai Keung Wong |
IEEE Trans. Multim. | 1 |
| 2023 | Cross-View Panorama Image SynthesisabstractIn this paper, we tackle the problem of ground-view panorama image generation conditioning on top-view aerial image, which is a challenging problem due to large gap between image domains associated with different view-points. Instead of learning the underlying cross-view mapping by a feedforward network as previous methods, we propose a novel adversarial feedback GAN framework named PanoGAN consisting of two key components: an adversarial feedback module and a dual branch discrimination strategy. First, aerial image is fed into the generator to produce panorama image and segmentation map, which facilitates the use of semantic layout information for model training. Second, the discriminator's feature responses of model outputs are encoded by the adversarial feedback module and then fed back to the generator for the next round of generation. Continual improvement of generated image quality is achieved through an iterative generation process. Third, to pursue high-fidelity and semantic consistency of the generated panorama image, we propose a pixel-segmentation alignment mechanism under the dual branch discrimiantion strategy that promotes the cooperation between generator and discriminator. Extensive experimental results on two challenging cross-view image datasets show that the proposed PanoGAN generates high-quality panorama images with more convincing details than state-of-the-art methods. The source code and trained models are available at https://github.com/sswuai/PanoGAN. Songsong Wu, Hao Tang 0005, Xiaoyuan Jing, Haifeng Zhao 0002, Jianjun Qian, Nicu Sebe, Yan Yan 0002 |
IEEE Trans. Multim. | 5 |
| 2023 | Incorporating Linear Regression Problems Into an Adaptive Framework With Feasible OptimizationsabstractAccompanied with the increasing popularity of linear regression approaches, most of the existing minimization problems are related with several convex measurements, e.g.,$\ell_1$/$\ell_2$/$\ell_{2,1}$-norm of a vector and$L_1$/$L_{2,1}$/Frobenius/nuclear norm of a matrix, where the regularized function and the loss function are usually studied for two objective terms case by case, respectively. To address this issue, this work combines these linear regression problems into a unified expression framework by employing an adaptive and flexible function, in which we need to choose different variable elements and adjust an inner parameter, properly. Besides this, they are equipped with some corresponding relationships and their interesting properties. Intuitively speaking, the proposed framework can generalize several traditional linear regression formulations and even more complex ones into an extended representation. For further optimizations, an iteratively re-weighted penalty solution (IRwPS) is devised without any inner loops, making the iteration programming easy to perform. Meanwhile, the theoretical results are provided for guaranteeing that the mathematical convergence analysis is solid and meaningful. Finally, by performing real-world applications in supervised, unsupervised, and semi-supervised tasks, numerical experiments are conducted to validate the theoretical properties and the superiority over some of the state-of-the-art. Hengmin Zhang, Feng Qian 0004, Bob Zhang 0001, Wenli Du, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | Generalized Nonconvex Nonsmooth Low-Rank Matrix Recovery Framework With Feasible Algorithm Designs and Convergence AnalysisabstractDecomposing data matrix into low-rank plus additive matrices is a commonly used strategy in pattern recognition and machine learning. This article mainly studies the alternating direction method of multiplier (ADMM) with two dual variables, which is used to optimize the generalized nonconvex nonsmooth low-rank matrix recovery problems. Furthermore, the minimization framework with a feasible optimization procedure is designed along with the theoretical analysis, where the variable sequences generated by the proposed ADMM can be proved to be bounded. Most importantly, it can be concluded from the Bolzano-Weierstrass theorem that there must exist a subsequence converging to a critical point, which satisfies the Karush-Kuhn-Tucher (KKT) conditions. Meanwhile, we further ensure the local and global convergence properties of the generated sequence relying on constructing the potential objective function. Particularly, the detailed convergence analysis would be regarded as one of the core contributions besides the algorithm designs and the model generality. Finally, the numerical simulations and the real-world applications are both provided to verify the consistence of the theoretical results, and we also validate the superiority in performance over several mostly related solvers to the tasks of image inpainting and subspace clustering. Hengmin Zhang, Feng Qian 0004, Peng Shi 0001, Wenli Du, Yang Tang 0001, Jianjun Qian, Chen Gong 0002, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Linearity-Aware Subspace ClusteringabstractObtaining a good similarity matrix is extremely important in subspace clustering. Current state-of-the-art methods learn the similarity matrix through self-expressive strategy. However, these methods directly adopt original samples as a set of basis to represent itself linearly. It is difficult to accurately describe the linear relation between samples in the real-world applications, and thus is hard to find an ideal similarity matrix. To better represent the linear relation of samples, we present a subspace clustering model, Linearity-Aware Subspace Clustering (LASC), which can consciously learn the similarity matrix by employing a linearity-aware metric. This is a new subspace clustering method that combines metric learning and subspace clustering into a joint learning framework. In our model, we first utilize the self-expressive strategy to obtain an initial subspace structure and discover a low-dimensional representation of the original data. Subsequently, we use the proposed metric to learn an intrinsic similarity matrix with linearity-aware on the obtained subspace. Based on such a learned similarity matrix, the inter-cluster distance becomes larger than the intra-cluster distances, and thus successfully obtaining a good subspace cluster result. In addition, to enrich the similarity matrix with more consistent knowledge, we adopt a collaborative learning strategy for self-expressive subspace learning and linearity-aware subspace learning. Moreover, we provide detailed mathematical analysis to show that the metric can properly characterize the linear correlation between samples. Yesong Xu, Shuo Chen 0003, Jun Li 0027, Jianjun Qian |
AAAI | 4 |
| 2022 | Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper, we propose a novel domain disentangled generative adversarial network (DD-GAN) for zero-shot sketch-based 3D retrieval, which can retrieve the unseen categories that are not accessed during training. Specifically, we first generate domain-invariant features and domain-specific features by disentangling the learned features of sketches and 3D shapes, where the domain-invariant features are used to align with the corresponding word embeddings. Then, we develop a generative adversarial network that combines the domain-specific features of the seen categories with the aligned domain-invariant features to synthesize samples, where the synthesized samples of the unseen categories are generated by using the corresponding word embeddings. Finally, we use the synthesized samples of the unseen categories combined with the real samples of the seen categories to train the network for retrieval, so that the unseen categories can be recognized. In order to reduce the domain shift problem, we utilize unlabeled unseen samples to enhance the discrimination ability of the discriminator. With the discriminator distinguishing the generated samples from the unlabeled unseen samples, the generator can generate more realistic unseen samples. Extensive experiments on the SHREC'13 and SHREC'14 datasets show that our method significantly improves the retrieval performance of the unseen categories. Rui Xu 0021, Zongyan Han, Le Hui, Jianjun Qian, Jin Xie 0001 |
AAAI | 4 |
| 2022 | Emphasizing Closeness and Diversity Simultaneously for Deep Face Representation
Chaoyu Zhao, Jianjun Qian, Shumin Zhu, Jin Xie 0001, Jian Yang 0003 |
ACCV (4) | 2 |
| 2022 | Generative Subgraph Contrast for Self-Supervised Graph Representation Learning
Yuehui Han, Le Hui, Haobo Jiang, Jianjun Qian, Jin Xie 0001 |
ECCV (30) | 4 |
| 2022 | Unsupervised Domain Adaptation for Point Cloud Semantic Segmentation via Graph MatchingabstractUnsupervised domain adaptation for point cloud semantic segmentation has attracted great attention due to its effectiveness in learning with unlabeled data. Most of existing methods use global-level feature alignment to transfer the knowledge from the source domain to the target domain, which may cause the semantic ambiguity of the feature space. In this paper, we propose a graph-based framework to explore the local-level feature alignment between the two domains, which can reserve semantic discrimination during adaptation. Specifically, in order to extract local-level features, we first dynamically construct local feature graphs on both domains and build a memory bank with the graphs from the source domain. In particular, we use optimal transport to generate the graph matching pairs. Then, based on the assignment matrix, we can align the feature distributions between the two domains with the graph-based local feature loss. Furthermore, we consider the correlation between the features of different categories and formulate a category-guided contrastive loss to guide the segmentation model to learn discriminative features on the target domain. Extensive experiments on different synthetic-to-real and real-to-real domain adaptation scenarios demonstrate that our method can achieve state-of-the-art performance. Our code is available at https://github.com/BianYikai/PointUDA. Yikai Bian, Le Hui, Jianjun Qian, Jin Xie 0001 |
IROS | 3 |
| 2022 | Spatial-Channel Mixed Attention Based Network for Remote Heart Rate Estimation
Bixiao Ling, Jianjun Qian, Jian Yang 0003 |
PRCV (2) | 3 |
| 2022 | Cross-view panorama image synthesis with progressive attention GANs
Songsong Wu, Hao Tang 0005, Xiaoyuan Jing, Jianjun Qian, Nicu Sebe, Yan Yan 0002 |
Pattern Recognit. | 4 |
| 2022 | Joint Optimal Transport With Convex Regularization for Robust Image ClassificationabstractThe critical step of learning the robust regression model from high-dimensional visual data is how to characterize the error term. The existing methods mainly employ the nuclear norm to describe the error term, which are robust against structure noises (e.g., illumination changes and occlusions). Although the nuclear norm can describe the structure property of the error term, global distribution information is ignored in most of these methods. It is known that optimal transport (OT) is a robust distribution metric scheme due to that it can handle correspondences between different elements in the two distributions. Leveraging this property, this article presents a novel robust regression scheme by integrating OT with convex regularization. The OT-based regression with$L_{2} $norm regularization (OTR) is first proposed to perform image classification. The alternating direction method of multipliers is developed to handle the model. To further address the occlusion problem in image classification, the extended OTR (EOTR) model is then presented by integrating the nuclear norm error term with an OTR model. In addition, we apply the alternating direction method of multipliers with Gaussian back substitution to solve EOTR and also provide the complexity and convergence analysis of our algorithms. Experiments were conducted on five benchmark datasets, including illumination changes and various occlusions. The experimental results demonstrate the performance of our robust regression model on biometric image classification against several state-of-the-art regression-based classification methods. Jianjun Qian, Wai Keung Wong, Hengmin Zhang, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 1 |
| 2022 | Global Convergence Guarantees of (A)GIST for a Family of Nonconvex Sparse Learning ProblemsabstractIn recent years, most of the studies have shown that the generalized iterated shrinkage thresholdings (GISTs) have become the commonly used first-order optimization algorithms in sparse learning problems. The nonconvex relaxations of the$\ell _{0}$-norm usually achieve better performance than the convex case (e.g.,$\ell _{1}$-norm) since the former can achieve a nearly unbiased solver. To increase the calculation efficiency, this work further provides an accelerated GIST version, that is, AGIST, through the extrapolation-based acceleration technique, which can contribute to reduce the number of iterations when solving a family of nonconvex sparse learning problems. Besides, we present the algorithmic analysis, including both local and global convergence guarantees, as well as other intermediate results for the GIST and AGIST, denoted as (A)GIST, by virtue of the Kurdyka-Łojasiewica (KŁ) property and some milder assumptions. Numerical experiments on both synthetic data and real-world databases can demonstrate that the convergence results of objective function accord to the theoretical properties and nonconvex sparse learning methods can achieve superior performance over some convex ones. Hengmin Zhang, Feng Qian 0004, Fanhua Shang, Wenli Du, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Cybern. | 5 |
| 2022 | CBi-GNN: Cross-Scale Bilateral Graph Neural Network for 3D Object Detectionabstract3D object detection from LiDAR point clouds is a challenging task, since the point clouds are irregular and sparse. Existing one-stage methods mainly predict the 3D bounding box of 3D objects by extracting deep down-scaled features of point clouds from low-level (high-resolution, HR) feature maps to high-level (low-resolution, LR). Nonetheless, most of these methods ignore geometric context information of the down-scaled feature maps across scales, especially only using the LR feature will result in incomplete structure and less location accuracy of 3D objects. In this paper, we propose a novel cross-scale graph network-based one-stage 3D object detector to fully exploit the geometric contexts of the voxels between the down-scaled feature maps. Specifically, we first employ a 3D sparse convolution neural network to form different resolutions of feature maps of voxels. We then dynamically construct a cross-scale bilateral graph to search the neighbor non-empty voxels in the HR feature map with a fixed radius for each non-empty voxel in the LR feature map. In the constructed graph, we present a bilateral attention mechanism (i.e., self-attention and spatial attention) in the HR feature map and encode each non-empty voxel in the LR feature map by aggregating the HR features to obtain the attention features. In addition, we design a non-local part pooling operation to improve the score of the detected bounding box of 3D objects. Finally, we formulate a multi-task loss to train our network for regression of the 3D bounding box of the 3D objects. Experiments on the challenging KITTI’s 3D/BEV benchmark show that our proposed detector outperforms all one-stage 3D object detectors and is comparable to two-stage 3D object detectors. Our code is available athttps://github.com/csjxchen/CBi-GNN. Jiaxin Chen 0001, Xiang Li 0041, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationabstractIn this paper, by modeling the point cloud registration task as a Markov decision process, we propose an end-to-end deep model embedded with the cross-entropy method (CEM) for unsupervised 3D registration. Our model consists of a sampling network module and a differentiable CEM module. In our sampling network module, given a pair of point clouds, the sampling network learns a prior sampling distribution over the transformation space. The learned sampling distribution can be used as a "good" initialization of the differentiable CEM module. In our differentiable CEM module, we first propose a maximum consensus criterion based alignment metric as the reward function for the point cloud registration task. Based on the reward function, for each state, we then construct a fused score function to evaluate the sampled transformations, where we weight the current and future rewards of the transformations. Particularly, the future rewards of the sampled transforms are obtained by performing the iterative closest point (ICP) algorithm on the transformed state. By selecting the top-k transformations with the highest scores, we iteratively update the sampling distribution. Furthermore, in order to make the CEM differentiable, we use the sparse-max function to replace the hard top-k selection. Finally, we formulate a Geman-McClure estimator based loss to train our end-to-end registration model. Extensive experimental results demonstrate the good registration performance of our method on benchmark datasets. Code is available at https://github.com/Jiang-HB/CEMNet. Haobo Jiang, Yaqi Shen, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
ICCV | 5 |
| 2021 | Robust Recovery of Low Rank Matrix by Nonconvex Rank Regularization
Hengmin Zhang, Wei Luo 0006, Wenli Du, Jianjun Qian, Jian Yang 0003, Bob Zhang 0001 |
ICIG (2) | 4 |
| 2021 | Planning with Learned Dynamic Model for Unsupervised Point Cloud RegistrationabstractPoint cloud registration is a fundamental problem in 3D computer vision. In this paper, we cast point cloud registration into a planning problem in reinforcement learning, which can seek the transformation between the source and target point clouds through trial and error. By modeling the point cloud registration process as a Markov decision process (MDP), we develop a latent dynamic model of point clouds, consisting of a transformation network and evaluation network. The transformation network aims to predict the new transformed feature of the point cloud after performing a rigid transformation (i.e., action) on it while the evaluation network aims to predict the alignment precision between the transformed source point cloud and target point cloud as the reward signal. Once the dynamic model of the point cloud is trained, we employ the cross-entropy method (CEM) to iteratively update the planning policy by maximizing the rewards in the point cloud registration process. Thus, the optimal policy, i.e., the transformation between the source and target point clouds, can be obtained via gradually narrowing the search space of the transformation. Experimental results on ModelNet40 and 7Scene benchmark datasets demonstrate that our method can yield good registration performance in an unsupervised manner. Haobo Jiang, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
IJCAI | 2 |
| 2021 | Non-significant Information Enhancement Based Attention Network for Face Anti-spoofing
Yangwei Dong, Jianjun Qian, Jian Yang 0003 |
PRCV (3) | 2 |
| 2021 | Structured discriminative tensor dictionary learning for unsupervised domain adaptation
Songsong Wu, Yan Yan 0002, Hao Tang 0005, Jianjun Qian, Jian Zhang 0002, Xiaoyuan Jing |
Neurocomputing | 4 |
| 2021 | Dual robust regression for pattern classification
Jianjun Qian, Shumin Zhu, Wai Keung Wong, Hengmin Zhang, Zhihui Lai 0001, Jian Yang 0003 |
Inf. Sci. | 1 |
| 2020 | Progressive Point Cloud Deconvolution Generation Network
Le Hui, Rui Xu 0021, Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
ECCV (15) | 4 |
| 2020 | PUI-Net: A Point Cloud Upsampling and Inpainting Network
Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
PRCV (1) | 3 |
| 2020 | The Devil is in the Detail: Deep Feature Based Disguised Face Recognition Method
Shumin Zhu, Jianjun Qian, Yangwei Dong, Wai Keung Wong |
PRCV (2) | 2 |
| 2020 | Image decomposition based matrix regression with applications to robust face recognition
Jianjun Qian, Jian Yang 0003, Yong Xu 0001, Jin Xie 0001, Zhihui Lai 0001, Bob Zhang 0001 |
Pattern Recognit. | 1 |
| 2020 | Low-Rank Matrix Recovery via Modified Schatten-p Norm Minimization With Convergence GuaranteesabstractIn recent years, low-rank matrix recovery problems have attracted much attention in computer vision and machine learning. The corresponding rank minimization problems are both combinational and NP-hard in general, which are mainly solved by both nuclear norm and Schatten-p (0<p<1) norm based optimization algorithms. However, inspired by weighted nuclear norm and Schatten-p norm as the relaxations of rank function, the main merits of this work firstly provide a modified Schatten-p norm in the affine matrix rank minimization problem, denoted as the modified Schatten-p norm minimization (MSpNM). Secondly, its surrogate function is constructed and the equivalence relationship with the MSpNM is further achieved. Thirdly, the iterative singular value thresholding algorithm (ISVTA) is devised to optimize it, and its accelerated version, i.e., AISVTA, is also obtained to reduce the number of iterations through the well-known Nesterov's acceleration strategy. Most importantly, the convergence guarantees and their relationship with objective function, stationary point and variable sequence generated by the proposed algorithms are established under some specific assumptions, e.g., Kurdyka-Łojasiewicz (KŁ) property. Finally, numerical experiments demonstrate the effectiveness of the proposed algorithms in the matrix completion problem for image inpainting and recommender systems. It should be noted that the accelerated algorithm has a much faster convergence speed and a very close recovery precision when comparing with the proposed non-accelerated one. Hengmin Zhang, Jianjun Qian, Bob Zhang 0001, Jian Yang 0003, Chen Gong 0002, Yang Wei 0003 |
IEEE Trans. Image Process. | 2 |
| 2019 | DSFD: Dual Shot Face DetectorabstractRecently, Convolutional Neural Network (CNN) has achieved great success in face detection. However, it remains a challenging problem for the current face detection methods owing to high degree of variability in scale, pose, occlusion, expression, appearance and illumination. In this Paper, we propose a novel detection network named Dual Shot face Detector(DSFD). which inherits the architecture of SSD and introduces a Feature Enhance Module (FEM) for transferring the original feature maps to extend the single shot detector to dual shot detector. Specially, progressive anchor loss (PAL) computed by using two set of anchors is adopted to effectively facilitate the features. Additionally, we propose an improved anchor matching (IAM) method by integrating novel data augmentation techniques and anchor design strategy in our DSFD to provide better initialization for the regressor. Extensive experiments on popular benchmarks: WIDER FACE (easy: 0.966, medium: 0.957, hard: 0.904) and FDDB ( discontinuous: 0.991, continuous: 0.862 ) demonstrate the superiority of DSFD over the state-of-the-art face detection methods (e.g., PyramidBox and SRN). Code will be made available upon publication. Jian Li 0062, Yabiao Wang, Changan Wang, Ying Tai, Jianjun Qian, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 5 |
| 2019 | Assignment Problem Based Deep Embedding
Ruishen Zheng, Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
PRCV (2) | 3 |
| 2019 | Obstacle Detection by Fusing Point Clouds and Monocular Image
Yang Wei 0003, Jian Yang 0003, Chen Gong 0002, Shuo Chen 0003, Jianjun Qian |
Neural Process. Lett. | 5 |
| 2019 | Efficient Recovery of Low-Rank Matrix via Double Nonconvex Nonsmooth Rank MinimizationabstractRecently, there is a rapidly increasing attraction for the efficient recovery of low-rank matrix in computer vision and machine learning. The popular convex solution of rank minimization is nuclear norm-based minimization (NNM), which usually leads to a biased solution since NNM tends to overshrink the rank components and treats each rank component equally. To address this issue, some nonconvex nonsmooth rank (NNR) relaxations have been exploited widely. Different from these convex and nonconvex rank substitutes, this paper first introduces a general and flexible rank relaxation function named weighted NNR relaxation function, which is actually derived from the initial double NNR (DNNR) relaxations, i.e., DNNR relaxation function acts on the nonconvex singular values function (SVF). An iteratively reweighted SVF optimization algorithm with continuation technology through computing the supergradient values to define the weighting vector is devised to solve the DNNR minimization problem, and the closed-form solution of the subproblem can be efficiently obtained by a general proximal operator, in which each element of the desired weighting vector usually satisfies the nondecreasing order. We next prove that the objective function values decrease monotonically, and any limit point of the generated subsequence is a critical point. Combining the Kurdyka-Łojasiewicz property with some milder assumptions, we further give its global convergence guarantee. As an application in the matrix completion problem, experimental results on both synthetic data and real-world data can show that our methods are competitive with several state-of-the-art convex and nonconvex matrix completion methods. Hengmin Zhang, Chen Gong 0002, Jianjun Qian, Bob Zhang 0001, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Scalable Proximal Jacobian Iteration Method With Global Convergence Analysis for Nonconvex Unconstrained Composite Optimizationsabstract-norm and rank function minimization problems. However, due to the absence of convexity in these nonconvex problems, developing efficient algorithms with convergence guarantee becomes very challenging. Inspired by the basic ideas of both the Jacobian alternating direction method of multipliers (JADMMs) for solving linearly constrained problems with separable objectives and the proximal gradient methods (PGMs) for optimizing the unconstrained problems with one variable, this paper focuses on extending the PGMs to the proximal Jacobian iteration methods (PJIMs) for handling with a family of nonconvex composite optimization problems with two splitting variables. To reduce the total computational complexity by decreasing the number of iterations, we devise the accelerated version of PJIMs through the well-known Nesterov's acceleration strategy and further extend both to solve the multivariable cases. Most importantly, we provide a rigorous convergence analysis, in theory, to show that the generated variable sequence globally converges to a critical point by exploiting the Kurdyka-Łojasiewica (KŁ) property for a broad class of functions. Furthermore, we also establish the linear and sublinear convergence rates of the obtained variable sequence in the objective function. As the specific application to the nonconvex sparse and low-rank recovery problems, several numerical experiments can verify that the newly proposed algorithms not only keep fast convergence speed but also have high precision. Hengmin Zhang, Jianjun Qian, Junbin Gao, Jian Yang 0003, Chunyan Xu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Episode-Experience Replay Based Tree-Backup Method for Off-Policy Actor-Critic Algorithm
Haobo Jiang, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
PRCV (1) | 2 |
| 2018 | Joint Bayesian guided metric learning for end-to-end face verification
Chunyan Xu, Jian Yang 0003, Jianjun Qian, Yuhui Zheng, LinLin Shen |
Neurocomputing | 4 |
| 2017 | Object detection via feature fusion based single networkabstractThis paper presents a novel network, coined single unified fully convolutional network (SingleNet), for object detection. The proposed method mainly combines two ideas: (1) Our approach aggregates hierarchical features and map them into a uniform space. So, it can further enhance the feature representation ability to degrade the recognition error; (2) To approximate the ground-truth box, we design a set of dense boxes over different aspect ratios and scales per feature map pixel to regress bounding box. It's easy to improve the location performance using dense boxes scheme. Additionally, SingleNet is convenient to train and can be integrated into detection system. Experimental results (mAP: 0.776) on VOC 2007 test demonstrate the advantages of the proposed method over state-of-the-art methods. Jian Li 0062, Jianjun Qian, Jian Yang 0003 |
ICIP | 2 |
| 2017 | Recurrent neural network for facial landmark detection
Yu Chen 0037, Jian Yang 0003, Jianjun Qian |
Neurocomputing | 3 |
| 2017 | Kernel orthogonal Procrustes regression for face recognition across pose
Ying Tai, Jian Yang 0003, Lei Luo 0001, Jianjun Qian |
Neurocomputing | 4 |
| 2017 | Bi-weighted robust matrix regression for face recognition
Jianchun Xie, Jian Yang 0003, Jianjun Qian, Lei Luo 0001 |
Neurocomputing | 3 |
| 2017 | Nonconvex relaxation based matrix regression for face recognition with structural noise and mixed noise
Hengmin Zhang, Jian Yang 0003, Jianjun Qian, Wei Luo 0006 |
Neurocomputing | 3 |
| 2017 | Weighted sparse coding regularized nonconvex matrix regression for robust face recognition
Hengmin Zhang, Jian Yang 0003, Jianchun Xie, Jianjun Qian, Bob Zhang 0001 |
Inf. Sci. | 4 |
| 2017 | Nuclear Norm Based Matrix Regression with Applications to Face Recognition with Occlusion and Illumination ChangesabstractRecently, regression analysis has become a popular tool for face recognition. Most existing regression methods use the one-dimensional, pixel-based error model, which characterizes the representation error individually, pixel by pixel, and thus neglects the two-dimensional structure of the error image. We observe that occlusion and illumination changes generally lead, approximately, to a low-rank error image. In order to make use of this low-rank structural information, this paper presents a two-dimensional image-matrix-based error model, namely, nuclear norm based matrix regression (NMR), for face representation and classification. NMR uses the minimal nuclear norm of representation error image as a criterion, and the alternating direction method of multipliers (ADMM) to calculate the regression coefficients. We further develop a fast ADMM algorithm to solve the approximate NMR model and show it has a quadratic rate of convergence. We experiment using five popular face image databases: the Extended Yale B, AR, EURECOM, Multi-PIE and FRGC. Experimental results demonstrate the performance advantage of NMR over the state-of-the-art regression-based methods for face recognition in the presence of occlusion and illumination variations. Jian Yang 0003, Lei Luo 0001, Jianjun Qian, Ying Tai, Fanlong Zhang, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Robust Nuclear Norm-Based Matrix Regression With Applications to Robust Face RecognitionabstractFace recognition (FR) via regression analysis-based classification has been widely studied in the past several years. Most existing regression analysis methods characterize the pixelwise representation error via l1-norm or l2-norm, which overlook the 2D structure of the error image. Recently, the nuclear norm-based matrix regression model is proposed to characterize low-rank structure of the error image. However, the nuclear norm cannot accurately describe the low-rank structural noise when the incoherence assumptions on the singular values does not hold, since it overpenalizes several much larger singular values. To address this problem, this paper presents the robust nuclear norm to characterize the structural error image and then extends it to deal with the mixed noise. The majorization-minimization (MM) method is applied to derive a iterative scheme for minimization of the robust nuclear norm optimization problem. Then, an efficiently alternating direction method of multipliers (ADMM) method is used to solve the proposed models. We use weighted nuclear norm as classification criterion to obtain the final recognition results. Experiments on several public face databases demonstrate the effectiveness of our models in handling with variations of structural noise (occlusion, illumination, and so on) and mixed noise. Jianchun Xie, Jian Yang 0003, Jianjun Qian, Ying Tai, Hengmin Zhang |
IEEE Trans. Image Process. | 3 |
| 2017 | Robust Image Regression Based on the Extended Matrix Variate Power Exponential Distribution of Dependent NoiseabstractDealing with partial occlusion or illumination is one of the most challenging problems in image representation and classification. In this problem, the characterization of the representation error plays a crucial role. In most current approaches, the error matrix needs to be stretched into a vector and each element is assumed to be independently corrupted. This ignores the dependence between the elements of error. In this paper, it is assumed that the error image caused by partial occlusion or illumination changes is a random matrix variate and follows the extended matrix variate power exponential distribution. This has the heavy tailed regions and can be used to describe a matrix pattern of l × m dimensional observations that are not independent. This paper reveals the essence of the proposed distribution: it actually alleviates the correlations between pixels in an error matrix E and makes E approximately Gaussian. On the basis of this distribution, we derive a Schatten p-norm-based matrix regression model with Lqregularization. Alternating direction method of multipliers is applied to solve this model. To get a closed-form solution in each step of the algorithm, two singular value function thresholding operators are introduced. In addition, the extended Schatten p-norm is utilized to characterize the distance between the test samples and classes in the design of the classifier. Extensive experimental results for image reconstruction and classification with structural noise demonstrate that the proposed algorithm works much more robustly than some existing regression-based methods. Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai, Gui-Fu Lu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Volume measurement based tensor completionabstractThis paper presents a new tensor completion method named minimum volume constraint tensor completion. Unlike the nuclear norm penalization based methods, our method extends the conception of the matrix volume to the tensor volume, and uses the volume measurement as the penalization to address the tensor completion problem. The alternating direction method of multipliers (ADMM) algorithm is then employed to solve the optimization problem of the proposed model. Experimental results on several popular databases show superior performance of our method compared to the nuclear norm penalization based methods in terms of the accuracy and robustness. Jianchun Xie, Jian Yang 0003, Ying Tai, Jianjun Qian |
ICIP | 4 |
| 2016 | Face alignment with Cascaded Bidirectional LSTM Neural NetworksabstractFace alignment is an important issue in many computer vision problems. The key problem is to find the nonlinear mapping from face image or feature to landmark locations. In this paper, we propose a novel cascaded approach with bidirectional Long Short Term Memory (LSTM) neural networks to approximate this nonlinear mapping. The cascaded structure is used to reduce the complexity of this problem and accelerate the algorithm by conducting the coarse-to-fine search. In each cascaded module, features of landmarks are delivered as inputs into the bidirectional LSTM network. The depth of the network guarantees the ability to learn highly complex mapping. The recurrent connections in LSTM explore the relationships of different landmarks and ensure that the shape of the face is maintained. On several challenging public databases, our approach achieves state-of-the-art performances. Yu Chen 0037, Jianjun Qian, Jian Yang 0003, Zhong Jin |
ICPR | 2 |
| 2016 | Structural Orthogonal Procrustes Regression for Face Recognition with Pose Variations and MisalignmentabstractRegression based method is a hot topic in the face recognition community and has achieved interesting results when dealing with well-aligned frontal face images. However, most of the existing regression analysis based methods are sensitive to pose variations. In this paper, we firstly introduce the orthogonal Procrustes problem (OPP), which is simple but effective, as a model to handle pose variations in two-dimensional face images. OPP seeks an optimal transformation between two images to correct the pose from one to the other. We integrate OPP into the regression model and propose the structural orthogonal Procrustes regression (SOPR) using the nuclear norm constraint on the error term to keep image's structural information. Moreover, a subject-wise strategy is adopted to address the problem that the gallery images may span over different poses. The proposed model is solved by an efficient iteratively reweighted algorithm and experimental results on popular face databases demonstrate the effectiveness of our method. Ying Tai, Jian Yang 0003, Fanlong Zhang, Yigong Zhang, Lei Luo 0001, Jianjun Qian |
SDM | 6 |
| 2016 | Exploring deep gradient information for biometric image feature representation
Jianjun Qian, Jian Yang 0003, Ying Tai |
Neurocomputing | 1 |
| 2016 | Adaptive noise dictionary construction via IRRPCA for face recognition
Yu Chen 0037, Jian Yang 0003, Lei Luo 0001, Hengmin Zhang, Jianjun Qian, Ying Tai, Jian Zhang 0025 |
Pattern Recognit. | 5 |
| 2016 | Learning discriminative singular value decomposition representation for face recognition
Ying Tai, Jian Yang 0003, Lei Luo 0001, Fanlong Zhang, Jianjun Qian |
Pattern Recognit. | 5 |
| 2016 | Tree-Structured Nuclear Norm Approximation With Applications to Robust Face RecognitionabstractStructured sparsity, as an extension of standard sparsity, has shown the outstanding performance when dealing with some highly correlated variables in computer vision and pattern recognition. However, the traditional mixed (L1, L2) or (L1, L∞) group norm becomes weak in characterizing the internal structure of each group since they cannot alleviate the correla-tions between variables. Recently, nuclear norm has been vali-dated to be useful for depicting a spatially structured matrix variable. It considers the global structure of the matrix variable but overlooks the local structure. To combine the advantages of structured sparsity and nuclear norm, this paper presents a tree-structured nuclear norm approximation (TSNA) model as-suming that the representation residual with tree-structured prior is a random matrix variable and follows a dependent matrix dis-tribution. The Extended Alternating Direction Method of Multi-pliers (EADMM) is utilized to solve the proposed model. An effi-cient bound condition based on the extended restricted isometry constants is provided to show the exact recovery of the proposed model under the given noisy case. In addition, TSNA is connected with some newest methods such as sparse representation based classifier (SRC), nuclear-L1 norm joint regression (NL1R) and nuclear norm based matrix regression (NMR), which can be re-garded as the special cases of TSNA. Experiments with face re-construction and recognition demonstrate the benefits of TSNA over other approaches. Lei Luo 0001, Liang Chen 0003, Jian Yang 0003, Jianjun Qian, Bob Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Face Recognition With Pose Variations and Misalignment via Orthogonal Procrustes RegressionabstractA linear regression-based method is a hot topic in face recognition community. Recently, sparse representation and collaborative representation-based classifiers for face recognition have been proposed and attracted great attention. However, most of the existing regression analysis-based methods are sensitive to pose variations. In this paper, we introduce the orthogonal Procrustes problem (OPP) as a model to handle pose variations existed in 2D face images. OPP seeks an optimal linear transformation between two images with different poses so as to make the transformed image best fits the other one. We integrate OPP into the regression model and propose the orthogonal Procrustes regression (OPR) model. To address the problem that the linear transformation is not suitable for handling highly non-linear pose variation, we further adopt a progressive strategy and propose the stacked OPR. As a practical framework, OPR can handle face alignment, pose correction, and face representation simultaneously. We optimize the proposed model via an efficient alternating iterative algorithm, and experimental results on three popular face databases, such as CMU PIE database, CMU Multi-PIE database, and LFW database, demonstrate the effectiveness of our proposed method. Ying Tai, Jian Yang 0003, Yigong Zhang, Lei Luo 0001, Jianjun Qian, Yu Chen 0037 |
IEEE Trans. Image Process. | 5 |
| 2015 | Nearest orthogonal matrix representation for face recognition
Jian Zhang 0025, Jian Yang 0003, Jianjun Qian |
Neurocomputing | 3 |
| 2015 | Nuclear-L1 norm joint regression for face reconstruction and recognition with mixed noise
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
Pattern Recognit. | 3 |
| 2015 | Robust nuclear norm regularized regression for face recognition with occlusion
Jianjun Qian, Lei Luo 0001, Jian Yang 0003, Fanlong Zhang, Zhouchen Lin |
Pattern Recognit. | 1 |
| 2015 | Matrix Variate Distribution-Induced Sparse Representation for Robust Image ClassificationabstractSparse representation learning has been successfully applied into image classification, which represents a given image as a linear combination of an over-complete dictionary. The classification result depends on the reconstruction residuals. Normally, the images are stretched into vectors for convenience, and the representation residuals are characterized by l2 -norm or l1 -norm, which actually assumes that the elements in the residuals are independent and identically distributed variables. However, it is hard to satisfy the hypothesis when it comes to some structural errors, such as illuminations, occlusions, and so on. In this paper, we represent the image data in their intrinsic matrix form rather than concatenated vectors. The representation residual is considered as a matrix variate following the matrix elliptically contoured distribution, which is robust to dependent errors and has long tail regions to fit outliers. Then, we seek the maximum a posteriori probability estimation solution of the matrix-based optimization problem under sparse regularization. An alternating direction method of multipliers (ADMMs) is derived to solve the resulted optimization problem. The convergence of the ADMM is proven theoretically. Experimental results demonstrate that the proposed method is more effective than the state-of-the-art methods when dealing with the structural errors. Jian Yang 0003, Lei Luo 0001, Jianjun Qian, Wei Xu 0052 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Nuclear Norm-Based 2-DPCA for Extracting Features From ImagesabstractThe 2-D principal component analysis (2-DPCA) is a widely used method for image feature extraction. However, it can be equivalently implemented via image-row-based principal component analysis. This paper presents a structured 2-D method called nuclear norm-based 2-DPCA (N-2-DPCA), which uses a nuclear norm-based reconstruction error criterion. The nuclear norm is a matrix norm, which can provide a structured 2-D characterization for the reconstruction error image. The reconstruction error criterion is minimized by converting the nuclear norm-based optimization problem into a series of F-norm-based optimization problems. In addition, N-2-DPCA is extended to a bilateral projection-based N-2-DPCA (N-B2-DPCA). The virtue of N-B2-DPCA over N-2-DPCA is that an image can be represented with fewer coefficients. N-2-DPCA and N-B2-DPCA are applied to face recognition and reconstruction and evaluated using the Extended Yale B, CMU PIE, FRGC, and AR databases. Experimental results demonstrate the effectiveness of the proposed methods. Fanlong Zhang, Jian Yang 0003, Jianjun Qian, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Nuclear-L1 Norm Joint Regression for Face Reconstruction and Recognition
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
ACCV (2) | 3 |
| 2014 | Nuclear Norm Regularized Sparse CodingabstractPartially occluded or illuminated faces pose a significant obstacle for robust, real-world face recognition. The problem of how to characterize the error caused by occlusion or illumination is still a challenging task. There must exist some close relationship between the error metric and error distribution. However, some metric (e.g. Z2-norm) can't characterize this error distribution completely. By some experiments, we found that nuclear norm is more suitable for characterizing the occluded or illuminated error distribution. Thus, a nuclear norm regularized sparse coding model is presented. Such a problem is solved by using ALM (or ADMM). In addition, we use nuclear norm as a metric to characterize the distance between reconstruction samples and classes. The experiments for image classification and face reconstruction demonstrate that our algorithm is robust to some face variations such as occlusion and illumination, and thus can act as a fast solver for matrix regression problem. Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Jing-Yu Yang 0001 |
ICPR | 3 |
| 2014 | Integration of multiple orientation and texture information for finger-knuckle-print verification
Guangwei Gao, Jian Yang 0003, Jianjun Qian, Lin Zhang 0014 |
Neurocomputing | 3 |
| 2014 | Histogram of visual words based on locally adaptive regression kernels descriptors for image feature extraction
Jianjun Qian, Jian Yang 0003, Nan Zhang 0037, Zhangjing Yang |
Neurocomputing | 1 |
| 2013 | Discriminative histograms of local dominant orientation (D-HLDO) for biometric image feature extraction
Jianjun Qian, Jian Yang 0003, Guangwei Gao |
Pattern Recognit. | 1 |
| 2013 | Local Structure-Based Image Decomposition for Feature Extraction With Applications to Face RecognitionabstractThis paper presents a robust but simple image feature extraction method, called image decomposition based on local structure (IDLS). It is assumed that in the local window of an image, the macro-pixel (patch) of the central pixel, and those of its neighbors, are locally linear. IDLS captures the local structural information by describing the relationship between the central macro-pixel and its neighbors. This relationship is represented with the linear representation coefficients determined using ridge regression. One image is actually decomposed into a series of sub-images (also called structure images) according to a local structure feature vector. All the structure images, after being down-sampled for dimensionality reduction, are concatenated into one super-vector. Fisher linear discriminant analysis is then used to provide a low-dimensional, compact, and discriminative representation for each super-vector. The proposed method is applied to face recognition and examined using our real-world face image database, NUST-RWFR, and five popular, publicly available, benchmark face image databases (AR, Extended Yale B, PIE, FERET, and LFW). Experimental results show the performance advantages of IDLS over state-of-the-art algorithms. Jianjun Qian, Jian Yang 0003, Yong Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Component-based global k-NN classifier for small sample size problems
Nan Zhang 0037, Jian Yang 0003, Jianjun Qian |
Pattern Recognit. Lett. | 3 |
| 2012 | Fuzzy local maximal marginal embedding for feature extraction
Cairong Zhao, Chuancai Liu, Xingjian Gu, Jianjun Qian |
Soft Comput. | 5 |