VLDB 2026 Research / reviewers in the wild / expert
Lei Luo 0001
dblp:82/3419-1
· DBLP profile ↗
80ranked-venue papers
14as first author
49since 2021 · last 2026
0000-0002-9976-0442ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 13 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 8 first-author · 28 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shaping Without Tearing: Controllable Diffeomorphic Deformations for Topology-Preserving 3D Point Cloud AugmentationabstractPoint cloud data augmentation is critical to improving the generalization of 3D deep learning models. However, existing methods often fail to preserve the underlying manifold structure, leading to semantic distortion or topology violation. This causes models to learn untrustworthy features, thereby limiting the representational ability of the model. To overcome these limitations, we propose ManiPoint, a novel point cloud augmentation framework based on diffeomorphism that explicitly preserves manifold structure during deformation. ManiPoint constructs diffeomorphic transformations via continuous differentiable mappings, ensuring topological consistency and geometric continuity between original and augmented data. To prevent excessive distortion and ensure semantic consistency, we introduce a controllable deformation mechanism that quantitatively constrains the augmentation magnitude and enables fine-grained control over the deformation space. We further provide theoretical analysis, indicating that, compared with topologically inconsistent methods, ManiPoint reduces empirical and vicinal risks by generating diverse and structurally reliable samples. Extensive experiments and visualizations on object-level datasets demonstrate that ManiPoint produces high-quality augmentations and consistently improves model robustness over existing baselines. Meanwhile, the scalability of our method was further verified on the scene-level datasets. Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
AAAI | 4 |
| 2026 | Small but Mighty: Dynamic Wavelet Expert-Guided Fine-Tuning of Large-Scale Models for Optical Remote Sensing Object SegmentationabstractAccurately localizing and segmenting relevant objects from optical remote sensing images (ORSIs) is critical for advancing remote sensing applications. Existing methods are typically built upon moderate-scale pre-trained models and employ diverse optimization strategies to achieve promising performance under full-parameter fine-tuning. In fact, deeper and larger-scale foundation models can provide stronger support for performance improvement. However, due to their massive number of parameters, directly adopting full-parameter fine-tuning leads to pronounced training difficulties, such as excessive GPU memory consumption and high computational costs, which result in extremely limited exploration of large-scale models in existing works. In this paper, we propose a novel dynamic wavelet expert-guided fine-tuning paradigm with fewer trainable parameters, dubbed WEFT, which efficiently adapts large-scale foundation models to ORSIs segmentation tasks by leveraging the guidance of wavelet experts. Specifically, we introduce a task-specific wavelet expert extractor to model wavelet experts from different perspectives and dynamically regulate their outputs, thereby generating trainable features enriched with task-specific information for subsequent fine-tuning. Furthermore, we construct an expert-guided conditional adapter that first enhances the fine-grained perception of frozen features for specific tasks by injecting trainable features, and then iteratively updates the information of both types of feature, allowing for efficient fine-tuning. Extensive experiments show that our WEFT not only outperforms 21 state-of-the-art (SOTA) methods on three ORSIs datasets, but also achieves optimal results in camouflage, natural, and medical scenarios. Jian Yang 0003, Lei Luo 0001 |
AAAI | 4 |
| 2026 | Structure-aware spherical density steered cross-domain learning for effective point cloud understanding
Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2026 | Leaning geometrical diffusion network via power spherical distribution for point clouds generation
Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0024 |
Pattern Recognit. | 4 |
| 2026 | TranSpike: Pixel-wise frequency reconstruction and spike interaction for remote photoplethysmography
Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Chuanfei Hu, Shuo Chen 0003, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2026 | CoMPR: Efficient point cloud dataset condensation via bidirectional matching and point recycling
Hongliang Zhang 0002, Xiaoqi An, Jiawei Lian, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2026 | ERDDCI: Exact Reversible Diffusion via Dual-Chain Inversion for High-Quality Image EditingabstractDiffusion models (DMs) have been successfully applied to real image editing. These models typically invert images into latent noise vectors during the inversion process, and then edit them during the inference process. However, DMs often rely on the local linearization assumption, which assumes that the noise injected during the inversion process approximates the noise removed during the inference process. While DMs efficiently generate images under this assumption, it also accumulates errors during the diffusion process due to the assumption, ultimately negatively impacting the quality of real image reconstruction and editing. To address this issue, we propose a novel ERDDCI (Exact Reversible Diffusion via Dual-Chain Inversion). ERDDCI uses the new Dual-Chain Inversion (DCI) for joint inference to derive an exact reversible diffusion process. Using DCI, our method avoids the cumbersome optimization process in existing inversion approaches and achieves high-quality image editing. Additionally, to accommodate image operations under high guidance scales, we introduce a dynamic control strategy that enables more refined image reconstruction and editing. Our experiments demonstrate that ERDDCI significantly outperforms state-of-the-art methods in a 50-step diffusion process. It achieves rapid and precise image reconstruction with SSIM of 0.999 and LPIPS of 0.001, and delivers competitive results in image editing. The source code is available at: https://github.com/daii-y/ERDDCI. Jimin Dai, Yingzhen Zhang, Shuo Chen 0003, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | ASGNet: Adaptive Spectrum Guidance Network for Automatic Polyp SegmentationabstractEarly identification and removal of polyps can reduce the risk of developing colorectal cancer. However, the diverse morphologies, complex backgrounds and often concealed nature of polyps make polyp segmentation in colonoscopy images highly challenging. Despite the promising performance of existing deep learning-based polyp segmentation methods, their perceptual capabilities remain biased toward local regions, mainly because of the strong spatial correlations between neighboring pixels in the spatial domain. This limitation makes it difficult to capture the complete polyp structures, ultimately leading to sub-optimal segmentation results. In this paper, we propose a novel adaptive spectrum guidance network, called ASGNet, which addresses the limitations of spatial perception by integrating spectral features with global attributes. Specifically, we first design a spectrum-guided non-local perception module that jointly aggregates local and global information, therefore enhancing the discriminability of polyp structures, and refining their boundaries. Moreover, we introduce a multi-source semantic extractor that integrates rich high-level semantic information to assist in the preliminary localization of polyps. Furthermore, we construct a dense cross-layer interaction decoder that effectively integrates diverse information from different layers and strengthens it to generate high-quality representations for accurate polyp segmentation. Extensive quantitative and qualitative results demonstrate the superiority of our ASGNet approach over 21 state-of-the-art methods across five widely-used polyp segmentation benchmarks. The code will be publicly available at: https://github.com/CSYSI/ASGNet. Hengmin Zhang, Jianjun Qian, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Localized Background-Aware Generative Distillation for Enhanced Remote Sensing Object Detection
Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Dual Manifold Regularization Steered Robust Representation Learning for Point Cloud AnalysisabstractWith the rapid advancement of 3D scanning technology, point clouds have become a crucial data type in computer vision and machine learning. However, learning robust representations for point clouds remains a significant challenge due to their irregularity and sparsity. In this paper, we propose a novel Dual Manifold Regularization (DMR) framework that makes full use of the properties of positive and negative curvature in manifolds to improve the representation of point clouds. Specifically, we leverage DMR based on hyperbolic and hyperspherical manifolds to address the limitations of traditional single-manifold regularization techniques, including inadequate generalization ability and adaptability to data diversity, as well as the difficulty of capturing complex relationships between data. To begin, we utilize the tree-like structure of the hyperbolic manifold to model the part-whole hierarchical relationships within point clouds. This allows for a more comprehensive representation of the data, improving the model's capability to understand complex shapes. Additionally, we construct positive samples through topological consistency augmentation and employ contrastive learning techniques in the hyperspherical manifold to capture more discriminative features within the data. Our experimental results show that our method outperforms traditional supervised learning and single-manifold regularization techniques in point cloud analysis. Specifically, for shape classification, DMR achieves a new State-Of-The-Art (SOTA) performance with 94.8% Overall Accuracy (OA) on ModelNet40 and 90.7% OA on ScanObjectNN, surpassing the recent SOTA model without increasing the baseline parameters. Jian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
AAAI | 4 |
| 2025 | Towards Better Spherical Sliced-Wasserstein Distance Learning with Data-Adaptive Discriminative Projection DirectionabstractSpherical Sliced-Wasserstein (SSW) has recently been proposed to measure the discrepancy between spherical data distributions in various fields, such as geology, medical domains, computer vision, and deep representation learning. However, in the original SSW, all projection directions are treated equally, which is too idealistic and cannot accurately reflect the importance of different projection directions for various data distributions. To address this issue, we propose a novel data-adaptive Discriminative Spherical Sliced-Wasserstein (DSSW) distance, which utilizes a projected energy function to determine the discriminative projection direction for SSW. In our new DSSW, we introduce two types of projected energy functions to generate the weights for projection directions with complete theoretical guarantees. The first type employs a non-parametric deterministic function that transforms the projected Wasserstein distance into its corresponding weight in each projection direction. This improves the performance of the original SSW distance with negligible additional computational overhead. The second type utilizes a neural network-induced function that learns the projection direction weight through a parameterized neural network based on data projections. This further enhances the performance of the original SSW distance with less extra computational overhead. Finally, we evaluate the performance of our proposed DSSW by comparing it with several state-of-the-art methods across a variety of machine learning tasks, including gradient flows, density estimation on real earth data, and self-supervised learning. Hongliang Zhang 0002, Shuo Chen 0003, Lei Luo 0001, Jian Yang 0003 |
AAAI | 3 |
| 2025 | Remote Photoplethysmography in Real-World and Extreme Lighting ScenariosabstractPhysiological activities can be manifested by the sensitive changes in facial imaging. While they are barely observable to our eyes, computer vision manners can, and the derived remote photoplethysmography (rPPG) has shown considerable promise. However, existing studies mainly rely on spatial skin recognition and temporal rhythmic interactions, so they focus on identifying explicit features under ideal light conditions, but perform poorly in-the-wild with intricate obstacles and extreme illumination exposure. In this paper, we propose an end-to-end video transformer model for rPPG. It strives to eliminate complex and unknown external time-varying interferences, whether they are sufficient to occupy subtle biosignal amplitudes or exist as periodic perturbations that hinder network training. In the specific implementation, we utilize global interference sharing, subject background reference, and self-supervised disentanglement to eliminate interference, and further guide learning based on spatiotemporal filtering, reconstruction guidance, and frequency domain and biological prior constraints to achieve effective rPPG. To the best of our knowledge, this is the first robust rPPG model for real outdoor scenarios based on natural face videos, and is lightweight to deploy. Extensive experiments show the competitiveness and performance of our model in rPPG prediction across datasets and scenes. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shuo Chen 0003, Jian Yang 0003 |
CVPR | 2 |
| 2025 | Cross-modal Gaussian Localization Distillation for Optical Information guided SAR Object DetectionabstractSynthetic Aperture Radar (SAR) images contain a dense clutter of objects that can be better characterized using bounding boxes with angles. However, accurately detecting the angles of objects remains challenging due to the imaging mechanism of SAR. To address this issue, we propose a novel knowledge distillation method called cross-modal Gaussian Localization Distillation (GaLD). It aims to improve SAR object detection performance by utilizing the angle information from optical images. Specifically, we convert the oriented bounding box into a Gaussian distribution and design a Gaussian Angle Distillation (GAD) loss function to align the angle information between optical and SAR images. In addition, to mitigate the negative impact of low-quality angle information on the network, we design an Adaptive Weighting Strategy (AWS) to guide the student network to prioritize high-quality angle information. The lack of high-quality oriented labels and objects in the OGSOD-1.0 dataset has hindered the progress in related fields. Therefore, we have added high-quality oriented labels and images to the OGSOD-1.0 dataset to construct a new dataset. Extensive experiments demonstrate the effectiveness and superiority of our proposed GaLD over existing methods. The dataset and code are available at: https://github.com/wchao0601/GaLD. Lei Luo 0001, Wenxuan Fang 0001, Jian Yang 0003 |
ICASSP | 2 |
| 2025 | Straighten Viscous Rectified Flow via Noise OptimizationabstractThe Reflow operation aims to straighten the inference trajectories of the rectified flow during training by constructing deterministic couplings between noises and images, thereby improving the quality of generated images in single-step or few-step generation. However, we identify critical limitations in Reflow, particularly its inability to rapidly generate high-quality images due to a distribution gap between images in its constructed deterministic couplings and real images. To address these shortcomings, we propose a novel alternative called Straighten Viscous Rectified Flow via Noise Optimization (VRFNO), which is a joint training framework integrating an encoder and a neural velocity field. VRFNO introduces two key innovations: (1) a historical velocity term that enhances trajectory distinction, enabling the model to more accurately predict the velocity of the current trajectory, and (2) the noise optimization through reparameterization to form optimized couplings with real images which are then utilized for training, effectively mitigating errors caused by Reflow's limitations. Comprehensive experiments on synthetic data and real datasets with varying resolutions show that VRFNO significantly mitigates the limitations of Reflow, achieving state-of-the-art performance in both one-step and few-step generation tasks. Jimin Dai, Jiexi Yan, Jian Yang 0003, Lei Luo 0001 |
ICCV | 4 |
| 2025 | Controllable-Lpmoe: Adapting to Challenging Object Segmentation Via Dynamic Local Priors From Mixture-Of-Experts
Jiawei Lian, Lei Luo 0001 |
ICCV | 4 |
| 2025 | Rethinking Point Cloud Data Augmentation: Topologically Consistent DeformationabstractData augmentation has been widely used in machine learning. Its main goal is to transform and expand the original data using various techniques, creating a more diverse and enriched training dataset. However, due to the disorder and irregularity of point clouds, existing methods struggle to enrich geometric diversity and maintain topological consistency, leading to imprecise point cloud understanding. In this paper, we propose SinPoint, a novel method designed to preserve the topological structure of the original point cloud through a homeomorphism. It utilizes the Sine function to generate smooth displacements. This simulates object deformations, thereby producing a rich diversity of samples. In addition, we propose a Markov chain Augmentation Process to further expand the data distribution by combining different basic transformations through a random process. Our extensive experiments demonstrate that our method consistently outperforms existing Mixup and Deformation methods on various benchmark point cloud datasets, improving performance for shape classification and part segmentation tasks. Specifically, when used with PointNet++ and DGCNN, our method achieves a state-of-the-art accuracy of 90.2 in shape classification with the real-world ScanObjectNN dataset. We release the code at https://github.com/CSBJian/SinPoint. Jian Bi, Qianliang Wu, Xiang Li 0041, Shuo Chen 0003, Jianjun Qian, Lei Luo 0001, Jian Yang 0003 |
ICML | 6 |
| 2025 | Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing ImagesabstractAutomatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former. Jiexi Yan, Jianjun Qian, Chunyan Xu, Jian Yang 0003, Lei Luo 0001 |
IJCAI | 6 |
| 2025 | SSAIM: Not All Self-Attentions Contain Effective Spatial Structure in Diffusion Models for Text-to-Image EditingabstractWith the rapid progress of diffusion-based Text-to-Image Generation (TIG), Text-to-Image Editing (TIE) has become increasingly important for enabling controllable visual content creation. A core challenge in TIE is generating text-guided edits while preserving the spatial structure of the original image. Recent methods attempt to address this by leveraging self-attention maps from diffusion models, as these encode rich spatial information. However, we identify two key limitations: (1) not all self-attention maps contribute meaningfully to spatial structure, and (2) over-reliance on them can suppress desired editing effects. To address this, we propose the Spatial Information Score (SIS), a novel metric that quantifies the spatial structure encoded in each self-attention map. Leveraging SIS, we develop Selective Self-Attention-based Image Manipulation (SSAIM), which selectively utilizes self-attention maps with effective spatial structure (high SIS) to preserve the structural of the original image and reduce excessive reliance on self-attention maps with ineffective spatial structure (low SIS) to enhance editing performance in TIE tasks. Extensive experiments across diverse TIE tasks demonstrate that SSAIM significantly improves both structural fidelity and editing quality. Zhenbo Yu, Jimin Dai, Yingzhen Zhang, Jian Yang 0003, Lei Luo 0001 |
ACM Multimedia | 5 |
| 2025 | DCNOT: Diffusion-Cascaded Neural Optimal Transport for Scalable Multi-Domain Image-to-Image Translation
Yingzhen Zhang, Jimin Dai, Qianliang Wu, Jian Yang 0003, Lei Luo 0001 |
ACM Multimedia | 5 |
| 2025 | Self-Supervised Temperature Representation Learning for Fever ScreeningabstractUtilizing thermal infrared facial imaging for fever screening in public spaces has become a common strategy to curb the spread of influenza viruses. However, it is difficult to capture larger number of faces with fever labels, which makes learning facial temperature representation extremely difficult. To overcome this limitation, we propose a self-supervised fever screening framework (SelfFS) to learn temperature representation from infrared face images. Specifically, SelfFS employs rate reduction theory to guide the network to focus on temperature features by expanding the coding rate of faces with different temperatures and compressing the coding rate of faces with the same temperature but different appearances. Furthermore, we impose sparsity constraints on the network parameters, which facilitates the extraction of simple temperature features with a limited number of neurons while filtering complex appearance features. Experiments demonstrate that our SelfFS framework outperforms existing fever screening techniques and achieves the comparable results with the supervised methods. Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 4 |
| 2025 | MSOD: A Large-Scale Multiscene Dataset and a Novel Diagonal-Geometry Loss for SAR Object DetectionabstractSynthetic Aperture Radar (SAR) has attracted significant attention due to its excellent all-weather imaging capabilities. However, SAR image object detection methods face two major challenges: 1) Most existing datasets are small in volume and single in category and scene. 2) Existing IoU-based loss functions cannot fully capture the relationship between prediction and target bounding boxes. To further advance the development of the SAR object detection method, we construct a large-scale multi-scene SAR object detection dataset called MSOD. It comprises three distinct scenarios, containing 40K images and about 1M instances of interest classified into six categories. In addition, we propose a novel diagonal-based similarity loss, Diagonal-Geometry IoU (DGIoU), to optimize the performance of SAR object detection by measuring the similarity between the diagonal of the prediction and target boxes. Specifically, we equivalently represent a rectangular box as a diagonal, and then define DGIoU based on the similarity of a set of sampling points between the diagonals of the predicted box and the target box. DGIoU effectively characterizes the difference between the predicted box and the target box, particularly in box inclusion and separation cases, resulting in improved localization accuracy. Numerous experimental results demonstrate that MSOD is closer to practical application and more challenging than existing SAR image datasets, and serves as a strong benchmark for evaluating the effectiveness of various IoU loss functions. The dataset and code are available at: https://github.com/wchao0601/MSOD-DGIoU. Wenxuan Fang 0001, Xiang Li 0041, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Video-Based Multiphysiological Disentanglement and Remote Robust Estimation for RespirationabstractRemote noncontact respiratory rate estimation by facial visual information has great research significance, providing valuable priors for health monitoring, clinical diagnosis, and anti-fraud. However, existing studies suffer from disturbances in epidermal specular reflections induced by head movements and facial expressions. Furthermore, diffuse reflections of light in the skin-colored subcutaneous tissue caused by multiple time-varying physiological signals independent of breathing are entangled with the intention of the respiratory process, leading to confusion in current research. To address these issues, this article proposes a novel network for natural light video-based remote respiration estimation. Specifically, our model consists of a two-stage architecture that progressively implements vital measurements. The first stage adopts an encoder-decoder structure to recharacterize the facial motion frame differences of the input video based on the gradient binary state of the respiratory signal during inspiration and expiration. Then, the obtained generative mapping, which is disentangled from various time-varying interferences and is only linearly related to the respiratory state, is combined with the facial appearance in the second stage. To further improve the robustness of our algorithm, we design a targeted long-term temporal attention module and embed it between the two stages to enhance the network's ability to model the breathing cycle that occupies ultra many frames and to mine hidden timing change clues. We train and validate the proposed network on a series of publicly available respiration estimation datasets, and the experimental results demonstrate its competitiveness against the state-of-the-art breathing and physiological prediction frameworks. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shangbing Gao, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | GLCONet: Learning Multisource Perception Representation for Camouflaged Object DetectionabstractRecently, the biological perception has been a powerful tool for handling the camouflaged object detection (COD) task. However, most existing methods are heavily dependent on the local spatial information of diverse scales from convolutional operations to optimize initial features. A commonly neglected point in these methods is the long-range dependencies between feature pixels from different scale spaces that can help the model build a global structure of the object, inducing a more precise image representation. In this article, we propose a novel global-local collaborative optimization network called GLCONet. Technically, we first design a collaborative optimization strategy (COS) from the perspective of multisource perception to simultaneously model the local details and global long-range relationships, which can provide features with abundant discriminative information to boost the accuracy in detecting camouflaged objects. Furthermore, we introduce an adjacent reverse decoder (ARD) that contains cross-layer aggregation and reverse optimization to integrate complementary information from different levels for generating high-quality representations. Extensive experiments demonstrate that the proposed GLCONet method with different backbones can effectively activate potentially significant pixels in an image, outperforming 20 state-of-the-art (SOTA) methods on three public COD datasets. The source code is available at: https://github.com/CSYSI/GLCONet. Hanyu Xuan, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Frequency-Spatial Entanglement Learning for Camouflaged Object Detection
Chunyan Xu, Jian Yang 0003, Hanyu Xuan, Lei Luo 0001 |
ECCV (6) | 5 |
| 2024 | Diff-Reg: Diffusion Model in Doubly Stochastic Matrix Space for Registration Problem
Qianliang Wu, Haobo Jiang, Lei Luo 0001, Jun Li 0027, Yaqing Ding 0001, Jin Xie 0001, Jian Yang 0003 |
ECCV (65) | 3 |
| 2024 | Hybrid Graph Representation Learning: Integrating Euclidean and Hyperbolic Space
Lening Li, Lei Luo 0001 |
ICPR (9) | 2 |
| 2024 | Efficiency Calibration of Implicit Regularization in Deep Networks via Self-paced Curriculum-Driven Singular Value Selection
Shuo Chen 0003, Jian Yang 0003, Lei Luo 0001 |
IJCAI | 4 |
| 2024 | SGNet: Salient Geometric Network for Point Cloud RegistrationabstractPoint Cloud Registration (PCR) is a critical and challenging task in computer vision and robotics. One of the primary difficulties in PCR is identifying salient and meaningful points that exhibit consistent semantic and geometric properties across different scans. Previous methods have encountered challenges with ambiguous matching due to the similarity among patch blocks throughout the entire point cloud and the lack of consideration for efficient global geometric consistency. To address these issues, we propose a new framework that includes several novel techniques. Firstly, we introduce a semantic-aware geometric encoder that combines object-level and patch-level semantic information. This encoder significantly improves registration recall by reducing ambiguity in patch-level superpoint matching. Additionally, we incorporate a prior knowledge approach that utilizes an intrinsic shape signature to identify salient points. This enables us to extract the most salient super points and meaningful dense points in the scene. Secondly, we introduce an innovative transformer that encodes High-Order (HO) geometric features. These features are crucial for identifying salient points within initial overlap regions while considering global high-order geometric consistency. We introduce an anchor node selection strategy to optimize this high-order transformer further. By encoding inter-frame triangle or polyhedron consistency features based on these anchor nodes, we can effectively learn high-order geometric features of salient super points. These high-order features are then propagated to dense points and utilized by a Sinkhorn matching module to identify critical correspondences for successful registration. The experiments conducted on the 3DMatch/3DLoMatch and KITTI datasets demonstrate the effectiveness of our method. Qianliang Wu, Yaqing Ding 0001, Lei Luo 0001, Haobo Jiang, Shuo Gu, Chuanwei Zhou, Jin Xie 0001, Jian Yang 0003 |
IROS | 3 |
| 2024 | Robust Audio-Visual Contrastive Learning for Proposal-Based Self-Supervised Sound Source Localization in VideosabstractBy observing a scene and listening to corresponding audio cues, humans can easily recognize where the sound is. To achieve such cross-modal perception on machines, existing methods take advantage of the maps obtained by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for prospective practical applications, we argue that these map-based methods only offer a coarse-grained and indirect description of the sound source. Additionally, these methods utilize a single audio-visual tuple at a time during self-supervised learning, causing the model to lose the crucial chance to reason about the data distribution of large-scale audio-visual samples. Although the introduction of Audio-Visual Contrastive Learning (AVCL) can effectively alleviate this issue, the contrastive set constructed by randomly sampling is based on the assumption that the audio and visual segments from all other videos are not semantically related. Since the resulting contrastive set contains a large number of faulty negatives, we believe that this assumption is rough. In this paper, we advocate a novel proposal-based solution that directly localizes the semantic object-level sound source, without any manual annotations. The Global Response Map (GRM) is incorporated as an unsupervised spatial constraint to filter those instances corresponding to a large number of sound-unrelated regions. As a result, our proposal-based Sound Source Localization (SSL) can be cast into a simpler Multiple Instance Learning (MIL) problem. To overcome the limitation of random sampling in AVCL, we propose a novel Active Contrastive Set Mining (ACSM) to mine the contrastive sets with informative and diverse negatives for robust AVCL. Our approaches achieve state-of-the-art (SOTA) performance when compared to several baselines on multiple SSL datasets with diverse scenarios. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Bo Jiang 0002, Lei Luo 0001, Xavier Alameda-Pineda, Yan Yan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | FeverNet: Enabling accurate and robust remote fever screening
Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2024 | Few-shot learning with long-tailed labels
Hongliang Zhang 0002, Shuo Chen 0003, Lei Luo 0001 |
Pattern Recognit. | 3 |
| 2024 | TranPhys: Spatiotemporal Masked Transformer Steered Remote Photoplethysmography EstimationabstractSubtle variations are invisible to the naked eyes in human physiological signals can reflect important biological and health indicators. Although numerous computer vision methods have been proposed to recover and magnify these changes, most of them either only focus on identifying and recognizing explicit features such as shapes and textures, or are weak in long-term temporal modeling and spatiotemporal interactive perception of implicit biometrics. Therefore, it is difficult for them to robustly overcome various disturbances that affect detection performance. To address these issues, this paper presents TranPhys, a novel remote photoplethysmography (rPPG) network for facial video-based heart rate estimation. Specifically, first, we argue that facial subregions vary over time due to their biological personalities. So we split the input face video into multiple spatiotemporal tubes, build the 3D vision transformer with encoders and decoders to adequately model the high-dimensional representations of the respective regulars in each subregion, and globally coordinate their feedback on the cardiac pulsing waveform. Second, we design the temporal pooling attention to more finely mine the subtle changes hidden in the skin color over time and their long-term contextual rhythm cues. Third, we leverage the self-supervised masked autoencoding paradigm to overcome redundancy to enhance the robustness of our model, and construct the targeted spatiotemporal sampling maps instead of raw input sequences as the pretrained constraint labels to fully inspire self-supervision. We train, validate, and practice our TranPhys on multiple public datasets to demonstrate that our method achieves the competitive performance in remote heart rate estimation. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | United Domain Cognition Network for Salient Object Detection in Optical Remote Sensing ImagesabstractRecently, deep learning-based salient object detection (SOD) in optical remote sensing images (ORSIs) have achieved significant breakthroughs. We observe that existing ORSIs-SOD methods consistently center around optimizing pixel features in the spatial domain, progressively distinguishing between backgrounds and objects. However, pixel information represents local attributes, which are often correlated with their surrounding context. Even with strategies expanding the local region, spatial features remain biased toward local characteristics, lacking the ability of global perception. To address this problem, we introduce the Fourier transform that generate global frequency features and achieve an image-size receptive field. To be specific, we propose a novel united domain cognition network (UDCNet) to jointly explore the global-local information in the frequency and spatial domains. Technically, we first design a frequency-spatial domain transformer (FSDT) block that mutually amalgamates the complementary local spatial and global frequency features to strength the capability of initial input features. Furthermore, a dense semantic excavation (DSE) module is constructed to capture higher level semantic for guiding the positioning of remote sensing objects. Finally, we devise a dual-branch joint optimization (DJO) decoder that applies the saliency and edge branches to generate high-quality representations for predicting salient objects. Experimental results demonstrate the superiority of the proposed UDCNet method over 24 state-of-the-art models, through extensive quantitative and qualitative comparisons in three widely used ORSIs-SOD datasets. The source code is available at:https://github.com/CSYSI/UDCNet. Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Curriculum Temperature for Knowledge DistillationabstractMost existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search. In general, the temperature controls the discrepancy between two distributions and can faithfully determine the difficulty level of the distillation task. Keeping a constant temperature, i.e., a fixed level of task difficulty, is usually sub-optimal for a growing student during its progressive learning stages. In this paper, we propose a simple curriculum-based technique, termed Curriculum Temperature for Knowledge Distillation (CTKD), which controls the task difficulty level during the student's learning career through a dynamic and learnable temperature. Specifically, following an easy-to-hard curriculum, we gradually increase the distillation loss w.r.t. the temperature, leading to increased distillation difficulty in an adversarial manner. As an easy-to-use plug-in technique, CTKD can be seamlessly integrated into existing knowledge distillation frameworks and brings general improvements at a negligible additional computation cost. Extensive experiments on CIFAR-100, ImageNet-2012, and MS-COCO demonstrate the effectiveness of our method. Zheng Li 0028, Xiang Li 0041, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo 0001, Jun Li 0027, Jian Yang 0003 |
AAAI | 6 |
| 2023 | Faster Fair Machine via Transferring Fairness Constraints to Virtual SamplesabstractFair classification is an emerging and important research topic in machine learning community. Existing methods usually formulate the fairness metrics as additional inequality constraints, and then embed them into the original objective. This makes fair classification problems unable to be effectively tackled by some solvers specific to unconstrained optimization. Although many new tailored algorithms have been designed to attempt to overcome this limitation, they often increase additional computation burden and cannot cope with all types of fairness metrics. To address these challenging issues, in this paper, we propose a novel method for fair classification. Specifically, we theoretically demonstrate that all types of fairness with linear and non-linear covariance functions can be transferred to two virtual samples, which makes the existing state-of-the-art classification solvers be applicable to these cases. Meanwhile, we generalize the proposed method to multiple fairness constraints. We take SVM as an example to show the effectiveness of our new idea. Empirically, we test the proposed method on real-world datasets and all results confirm its excellent performance. Zhou Zhai, Lei Luo 0001, Heng Huang 0001, Bin Gu 0001 |
AAAI | 2 |
| 2023 | Denoising Multi-Similarity Formulation: A Self-Paced Curriculum-Driven Approach for Robust Metric LearningabstractDeep Metric Learning (DML) is a group of techniques that aim to measure the similarity between objects through the neural network. Although the number of DML methods has rapidly increased in recent years, most previous studies cannot effectively handle noisy data, which commonly exists in practical applications and often leads to serious performance deterioration. To overcome this limitation, in this paper, we build a connection between noisy samples and hard samples in the framework of self-paced learning, and propose a Balanced Self-Paced Metric Learning (BSPML) algorithm with a denoising multi-similarity formulation, where noisy samples are treated as extremely hard samples and adaptively excluded from the model training by sample weighting. Especially, due to the pairwise relationship and a new balance regularization term, the sub-problem w.r.t. sample weights is a nonconvex quadratic function. To efficiently solve this nonconvex quadratic problem, we propose a doubly stochastic projection coordinate gradient algorithm. Importantly, we theoretically prove the convergence not only for the doubly stochastic projection coordinate gradient algorithm, but also for our BSPML algorithm. Experimental results on several standard data sets demonstrate that our BSPML algorithm has better generalization ability and robustness than the state-of-the-art robust DML approaches. Chenkang Zhang, Lei Luo 0001, Bin Gu 0001 |
AAAI | 2 |
| 2023 | Local-Fusion Diffusion Model for Enhancing Few-Shot Image Generation
Jishuai Hou, Lei Luo 0001, Jian Yang 0003 |
ICIG (1) | 2 |
| 2023 | Graph Matching Optimization Network for Point Cloud RegistrationabstractPoint Cloud Registration is a fundamental and challenging problem in 3D computer vision. Recent works often utilize geometric structure features in downsampled points (patches) to seek correspondences, then propagate these sparse patch correspondences to the dense level in the corresponding patches' neighborhood. However, they neglect the explicit global scale rigid constraint at the dense level point matching. We claim that the explicit isometry-preserving constraint in the dense level on a global scale is also important for improving feature representation in the training stage. To this end, we propose a Graph Matching Optimization based Network (GMONet for short), which utilizes the graph-matching optimizer to explicitly exert the isometry preserving constraints in the point feature training to improve the point feature representation. Specifically, we exploit a partial graph-matching optimizer to enhance the super point (i.e., down-sampled key points) features and a full graph-matching optimizer to improve the dense level point features in the overlap region. Meanwhile, we leverage the inexact proximal point method and the mini-batch sampling technique to accelerate these two graph-matching optimizers. Given high discriminative point features in the evaluation stage, we utilize the RANSAC approach to estimate the transformation between the scanned pairs. The proposed method has been evaluated on the 3DMatch/3DLoMatch and the KITTI datasets. The experimental results show that our method performs competitively compared to state-of-the-art baselines. Qianliang Wu, Yaqi Shen, Haobo Jiang, Guofeng Mei, Yaqing Ding 0001, Lei Luo 0001, Jin Xie 0001, Jian Yang 0003 |
IROS | 6 |
| 2023 | Doubly Robust AUC Optimization against Noisy and Adversarial SamplesabstractArea under the ROC curve (AUC) is an important and widely used metric in machine learning especially for imbalanced datasets. In current practical learning problems, not only adversarial samples but also noisy samples seriously threaten the performance of learning models. Nowadays, there have been a lot of research works proposed to defend the adversarial samples and noisy samples separately. Unfortunately, to the best of our knowledge, none of them with AUC optimization can secure against the two kinds of harmful samples simultaneously. To fill this gap and also address the challenge, in this paper, we propose a novel doubly robust dAUC optimization (DRAUC) algorithm. Specifically, we first exploit the deep integration of self-paced learning and adversarial training under the framework of AUC optimization, and provide a statistical upper bound to the AUC adversarial risk. Inspired by the statistical upper bound, we propose our optimization objective followed by an efficient alternatively stochastic descent algorithm, which can effectively improve the performance of learning models by guarding against adversarial samples and noisy samples. Experimental results on several standard datasets demonstrate that our DRAUC algorithm has better noise robustness and adversarial robustness than the state-of-the-art algorithms. Chenkang Zhang, Wanli Shi, Lei Luo 0001, Bin Gu 0001 |
KDD | 3 |
| 2023 | Hyperbolic embedding steered spatiotemporal graph convolutional network for video-based remote heart rate estimation
Hang Shao 0001, Lei Luo 0001, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Adaptive Hierarchical Similarity Metric Learning With Noisy LabelsabstractDeep Metric Learning (DML) plays a critical role in various machine learning tasks. However, most existing deep metric learning methods with binary similarity are sensitive to noisy labels, which are widely present in real-world data. Since these noisy labels often cause a severe performance degradation, it is crucial to enhance the robustness and generalization ability of DML. In this paper, we propose an Adaptive Hierarchical Similarity Metric Learning method. It considers two noise-insensitive information, i.e., class-wise divergence and sample-wise consistency. Specifically, class-wise divergence can effectively excavate richer similarity information beyond binary in modeling by taking advantage of Hyperbolic metric learning, while sample-wise consistency can further improve the generalization ability of the model using contrastive augmentation. More importantly, we design an adaptive strategy to integrate this information in a unified view. It is noteworthy that the new method can be extended to any pair-based metric loss. Extensive experimental results on benchmark datasets demonstrate that our method achieves state-of-the-art performance compared with current deep metric learning approaches. Jiexi Yan, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Noise Is Also Useful: Negative Correlation-Steered Latent Contrastive LearningabstractHow to effectively handle label noise has been one of the most practical but challenging tasks in Deep Neural Networks (DNNs). Recent popular methods for training DNNs with noisy labels mainly focus on directly filtering out samples with low confidence or repeatedly mining valuable information from low-confident samples. However, they cannot guarantee the robust generalization of models due to the ignorance of useful information hidden in noisy data. To address this issue, we propose a new effective method named as LaCoL (Latent Contrastive Learning) to leverage the negative correlations from the noisy data. Specifically, in label space, we exploit the weakly-augmented data to filter samples and adopt classification loss on strong augmentations of the selected sample set, which can preserve the training diversity. While in metric space, we utilize weakly-supervised contrastive learning to excavate these negative correlations hidden in noisy data. Moreover, a cross-space similarity consistency regularization is provided to constrain the gap between label space and metric space. Extensive experiments have validated the superiority of our approach over existing state-of-the-art methods. Jiexi Yan, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
CVPR | 2 |
| 2021 | Unsupervised Hyperbolic Metric LearningabstractLearning feature embedding directly from images without any human supervision is a very challenging and essential task in the field of computer vision and machine learning. Following the paradigm in supervised manner, most existing unsupervised metric learning approaches mainly focus on binary similarity in Euclidean space. However, these methods cannot achieve promising performance in many practical applications, where the manual information is lacking and data exhibits non-Euclidean latent anatomy. To address this limitation, we propose an Unsupervised Hyperbolic Metric Learning method with Hierarchical Similarity. It considers the natural hierarchies of data by taking advantage of Hyperbolic metric learning and hierarchical clustering, which can effectively excavate richer similarity information beyond binary in modeling. More importantly, we design a new loss function to capture the hierarchical similarity among samples to enhance the stability of the proposed method. Extensive experimental results on benchmark datasets demonstrate that our method achieves state-of-the-art performance compared with current unsupervised deep metric learning approaches. Jiexi Yan, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
CVPR | 2 |
| 2021 | Learning Better Visual Data Similarities via New Grouplet Non-Euclidean EmbeddingabstractIn many computer vision problems, it is desired to learn the effective visual data similarity such that the prediction accuracy can be enhanced. Deep Metric Learning (DML) methods have been actively studied to measure the data similarity. Pair-based and proxy-based losses are the two major paradigms in DML. However, pair-wise methods involve expensive training costs, while proxy-based methods are less accurate in characterizing the relationships between data points. In this paper, we provide a hybrid grouplet paradigm, which inherits the accurate pair-wise relationship in pair-based methods and the efficient training in proxy-based methods. Our method also equips a non-Euclidean space to DML, which employs a hierarchical representation manifold. More specifically, we propose a unified graph perspective — different DML methods learn different local connecting patterns between data points. Based on the graph interpretation, we construct a flexible subset of data points, dubbed grouplet. Our grouplet doesn’t require explicit pair-wise relationships, instead, we encode the data relationships in an optimal transport problem regarding the proxies, and solve this problem via a differentiable implicit layer to automatically determine the relationships. Extensive experimental results show that our method significantly outperforms state-of-the-art baselines on several benchmarks. The ablation studies also verify the effectiveness of our method. Yanfu Zhang, Lei Luo 0001, Wenhan Xian, Heng Huang 0001 |
ICCV | 2 |
| 2021 | Unified Fairness from Data to Learning AlgorithmabstractIn classification problems, individual fairness prevents discrimination against individuals based on protected attributes. Fairness-aware methods usually consist of two stages, first determining a fair metric concerning the similarity between different instances and then learning the fairness-aware model. However, existing works usually consider these two stages separately and only focus on improving the individual stage. Moreover, the choice of fair metric is heavily dependent on the task or dataset of interest, which requires ad-hoc domain knowledge and introduces extra difficulty into algorithm designing. As such, this discrepancy presumably leads to sub-optimal fairness-aware pipelines for different applications. In this paper, we propose to fill in the fairness learning gap between these two stages by automatically learning an effective metric integrated into the fairness of both data and classifiers. Specifically, we formulate the fairness-aware classification as a distributional robustness optimization problem based on deep metric learning and propose an effective optimization algorithm to solve it. Meanwhile, we establish the asymptotically unbiased generalization bounds for the proposed algorithm using the techniques of U-statistics. The experimental results on popular benchmark datasets demonstrate that the proposed approach achieves consistent improvement concerning several fairness assessments. Yanfu Zhang, Lei Luo 0001, Heng Huang 0001 |
ICDM | 2 |
| 2021 | A new nonlocal means based framework for mixed noise removal
Jielin Jiang, Jian Yang 0003, Zhi-Xin Yang 0001, Yadang Chen, Lei Luo 0001 |
Neurocomputing | 6 |
| 2021 | Learnable low-rank latent dictionary for subspace clustering
Yesong Xu, Shuo Chen 0003, Jun Li 0027, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2021 | δ-Norm-Based Robust Regression With Applications to Image Analysisabstract-norm, etc.) have been widely leveraged to form the loss function of different regression models, and have played an important role in image analysis. However, the previous regression models adopting the existing norms are sensitive to outliers and, thus, often bring about unsatisfactory results on the heavily corrupted images. This is because their adopted norms for measuring the data residual can hardly suppress the negative influence of noisy data, which will probably mislead the regression process. To address this issue, this paper proposes a novel δ (delta)-norm to count the nonzero blocks around an element in a vector or matrix, which weakens the impacts of outliers and also takes the structure property of examples into account. After that, we present the δ -norm-based robust regression (DRR) in which the data examples are mapped to the kernel space and measured by the proposed δ -norm. By exploring an explicit kernel function, we show that DRR has a closed-form solution, which suggests that DRR can be efficiently solved. To further handle the influences from mixed noise, DRR is extended to a multiscale version. The experimental results on image classification and background modeling datasets validate the superiority of the proposed approach to the existing state-of-the-art robust regression models. Shuo Chen 0003, Jian Yang 0003, Yang Wei 0003, Lei Luo 0001, Gui-Fu Lu, Chen Gong 0002 |
IEEE Trans. Cybern. | 4 |
| 2021 | Discriminative Cross-Modality Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationabstractIt is theoretically insufficient to construct a complete set of semantics in the real world using single-modality data. As a typical application of multi-modality perception, the audio-visual event localization task aims to match audio and visual components to identify the simultaneous events of interest. Although some recent methods have been proposed to deal with this task, they cannot handle the practical situation of temporal inconsistency that is widespread in the audio-visual scene. Inspired by the human system which automatically filters out event-unrelated information when performing multi-modality perception, we propose a discriminative cross-modality attention network to simulate such a process. Similar to human mechanism, our network can adaptively select "where" to attend, "when" to attend and "which" to attend for audio-visual event localization. In addition, to prevent our network from getting trivial solutions, a novel eigenvalue-based objective function is proposed to train the whole network to better fuse audio and visual signals, which can obtain discriminative and nonlinear multi-modality representation. In this way, even with large temporal inconsistency between audio and visual sequence, our network is able to adaptively select event-valuable information for audio-visual event localization. Furthermore, we systemically investigate three subtasks of audio-visual event localization, i.e., temporal localization, weakly-supervised spatial localization and cross-modality localization. The visualization results also help us better understand how our network works. Hanyu Xuan, Lei Luo 0001, Zhenyu Zhang 0005, Jian Yang 0003, Yan Yan 0002 |
IEEE Trans. Image Process. | 2 |
| 2020 | Adversarial Nonnegative Matrix FactorizationabstractNonnegative Matrix Factorization (NMF) has become an increasingly important research topic in machine learning. Despite all the practical success, most of existing NMF models are still vulnerable to adversarial attacks. To overcome this limitation, we propose a novel Adversarial NMF (ANMF) approach in which an adversary can exercise some control over the perturbed data generation process. Different from the traditional NMF models which focus on either the regular input or certain types of noise, our model considers potential test adversaries that are beyond the pre-defined constraints, which can cope with various noises (or perturbations). We formulate the proposed model as a bilevel optimization problem and use Alternating Direction Method of Multipliers (ADMM) to solve it with convergence analysis. Theoretically, the robustness analysis of ANMF is established under mild conditions dedicating asymptotically unbiased prediction. Extensive experiments verify that ANMF is robust to a broad categories of perturbations, and achieves state-of-the-art performances on distinct real-world benchmark datasets. Lei Luo 0001, Yanfu Zhang, Heng Huang 0001 |
ICML | 1 |
| 2020 | Sinkhorn RegressionabstractThis paper introduces a novel Robust Regression (RR) model, named Sinkhorn regression, which imposes Sinkhorn distances on both loss function and regularization. Traditional RR methods target at searching for an element-wise loss function (e.g., Lp-norm) to characterize the errors such that outlying data have a relatively smaller influence on the regression estimator. Due to the neglect of the geometric information, they often lead to the suboptimal results in the practical applications. To address this problem, we use a cross-bin distance function, i.e., Sinkhorn distances, to capture the geometric knowledge of real data. Sinkhorn distances is invariant in movement, rotation and zoom. Thus, our method is more robust to variations of data than traditional regression models. Meanwhile, we leverage Kullback-Leibler divergence to relax the proposed model with marginal constraints into its unbalanced formulation to adapt more types of features. In addition, we propose an efficient algorithm to solve the relaxed model and establish its complete statistical guarantees under mild conditions. Experiments on the five publicly available microarray data sets and one mass spectrometry data set demonstrate the effectiveness and robustness of our method. Lei Luo 0001, Jian Pei 0001, Heng Huang 0001 |
IJCAI | 1 |
| 2019 | Orthogonality-Promoting Dictionary Learning via Bayesian InferenceabstractDictionary Learning (DL) plays a crucial role in numerous machine learning tasks. It targets at finding the dictionary over which the training set admits a maximally sparse representation. Most existing DL algorithms are based on solving an optimization problem, where the noise variance and sparsity level should be known as the prior knowledge. However, in practice applications, it is difficult to obtain these knowledge. Thus, non-parametric Bayesian DL has recently received much attention of researchers due to its adaptability and effectiveness. Although many hierarchical priors have been used to promote the sparsity of the representation in non-parametric Bayesian DL, the problem of redundancy for the dictionary is still overlooked, which greatly decreases the performance of sparse coding. To address this problem, this paper presents a novel robust dictionary learning framework via Bayesian inference. In particular, we employ the orthogonality-promoting regularization to mitigate correlations among dictionary atoms. Such a regularization, encouraging the dictionary atoms to be close to being orthogonal, can alleviate overfitting to training data and improve the discrimination of the model. Moreover, we impose Scale mixture of the Vector variate Gaussian (SMVG) distribution on the noise to capture its structure. A Regularized Expectation Maximization Algorithm is developed to estimate the posterior distribution of the representation and dictionary with orthogonality-promoting regularization. Numerical results show that our method can learn the dictionary with an accuracy better than existing methods, especially when the number of training signals is limited. Lei Luo 0001, Jie Xu 0012, Cheng Deng 0002, Heng Huang 0001 |
AAAI | 1 |
| 2019 | Robust Metric Learning on Grassmann Manifolds with Generalization GuaranteesabstractIn recent research, metric learning methods have attracted increasing interests in machine learning community and have been applied to many applications. However, the existing metric learning methods usually use a fixed L2-norm to measure the distance between pairwise data samples in the projection space, which cannot provide an effective mechanism to automatically remove the noise that exist in data samples. To address this issue, we propose a new robust formulation of metric learning. Our new model constructs a projection from higher dimensional Grassmann manifold into the one in a relative low-dimensional with more discriminative capability, where the errors between sample points are considered as an MLE (maximum likelihood estimation)-like estimator. An efficient iteratively reweighted algorithm is derived to solve the proposed metric learning model. More importantly, we establish the generalization bounds for the proposed algorithm by utilizing the techniques of U-statistics. Experiments on six benchmark datasets clearly show that the proposed method achieves consistent improvements in discrimination accuracy, in comparison to state-of-the-art methods. Lei Luo 0001, Jie Xu 0012, Cheng Deng 0002, Heng Huang 0001 |
AAAI | 1 |
| 2019 | Curvilinear Distance Metric LearningabstractDistance Metric Learning aims to learn an appropriate metric that faithfully measures the distance between two data points. Traditional metric learning methods usually calculate the pairwise distance with fixed distance functions (\emph{e.g.,}\ Euclidean distance) in the projected feature spaces. However, they fail to learn the underlying geometries of the sample space, and thus cannot exactly predict the intrinsic distances between data points. To address this issue, we first reveal that the traditional linear distance metric is equivalent to the cumulative arc length between the data pair's nearest points on the learned straight measurer lines. After that, by extending such straight lines to general curved forms, we propose a Curvilinear Distance Metric Learning (CDML) method, which adaptively learns the nonlinear geometries of the training data. By virtue of Weierstrass theorem, the proposed CDML is equivalently parameterized with a 3-order tensor, and the optimization algorithm is designed to learn the tensor parameter. Theoretical analysis is derived to guarantee the effectiveness and soundness of CDML. Extensive experiments on the synthetic and real-world datasets validate the superiority of our method over the state-of-the-art metric learning models. Shuo Chen 0003, Lei Luo 0001, Jian Yang 0003, Chen Gong 0002, Jun Li 0027, Heng Huang 0001 |
NeurIPS | 2 |
| 2019 | Nesting-structured nuclear norm minimization for spatially correlated matrix variate
Lei Luo 0001, Jian Yang 0003, Yigong Zhang, Yong Xu 0001, Heng Huang 0001 |
Pattern Recognit. | 1 |
| 2018 | Matrix Variate Gaussian Mixture Distribution Steered Robust Metric LearningabstractMahalanobis Metric Learning (MML) has been actively studied recently in machine learning community. Most of existing MML methods aim to learn a powerful Mahalanobis distance for computing similarity of two objects. More recently, multiple methods use matrix norm regularizers to constrain the learned distance matrixMto improve the performance. However, in real applications, the structure of the distance matrix M is complicated and cannot be characterized well by the simple matrix norm. In this paper, we propose a novel robust metric learning method with learning the structure of the distance matrix in a new and natural way. We partition M into blocks and consider each block as a random matrix variate, which is fitted by matrix variate Gaussian mixture distribution. Different from existing methods, our model has no any assumption on M and automatically learns the structure of M from the real data, where the distance matrix M often is neither sparse nor low-rank. We design an effective algorithm to optimize the proposed model and establish the corresponding theoretical guarantee. We conduct extensive evaluations on the real-world data. Experimental results show our method consistently outperforms the related state-of-the-art methods. Lei Luo 0001, Heng Huang 0001 |
AAAI | 1 |
| 2018 | Multi-Level Metric Learning via Smoothed Wasserstein DistanceabstractTraditional metric learning methods aim to learn a single Mahalanobis distance metric M, which, however, is not discriminative enough to characterize the complex and heterogeneous data. Besides, if the descriptors of the data are not strictly aligned, Mahalanobis distance would fail to exploit the relations among them. To tackle these problems, in this paper, we propose a multi-level metric learning method using a smoothed Wasserstein distance to characterize the errors between any two samples, where the ground distance is considered as a Mahalanobis distance. Since smoothed Wasserstein distance provides not only a distance value but also a flow-network indicating how the probability mass is optimally transported between the bins, it is very effective in comparing two samples whether they are aligned or not. In addition, to make full use of the global and local structures that exist in data features, we further model the commonalities between various classification through a shared distance matrix and the classification-specific idiosyncrasies with additional auxiliary distance matrices. An efficient algorithm is developed to solve the proposed new model. Experimental evaluations on four standard databases show that our method obviously outperforms other state-of-the-art methods. Jie Xu 0012, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
IJCAI | 2 |
| 2018 | New Robust Metric Learning Model Using Maximum Correntropy Criterionabstracttopic with many real-world applications. Most existing metric learning methods aim to learn an optimal Mahalanobis distance matrix M, under which data samples from the same class are forced to be close to each other and those from different classes are pushed far away. The Mahalanobis distance matrix M can be factorized as M = L'L, and the Mahalanobis distance induced by L is equivalent to the Euclidean distance after linear projection of the feature vectors on the rows of L. However, the Euclidean distance is only suitable for characterizing Gaussian noise, thus the traditional metric learning algorithms are not robust to achieve good performance when they are applied to the occlusion data, which often appear in image and video data mining applications. To overcome this limitation, we propose a new robust metric learning approach by introducing the maximum correntropy criterion to deal with real-world malicious occlusions or corruptions. In our new model, we enforce the intra-class reconstruction residual of each sample to be smaller than the inter-class reconstruction residual by a large margin. Meanwhile, we employ correntropy induced metric to fit the reconstruction residual, which has been proved to be useful in non-Gaussian data processing. Leveraging the half-quadratic optimization technique, we derive an efficient algorithm to solve the proposed new model and provide its convergence guarantee as well. Extensive experiments on various occluded data sets indicate that our proposed model can achieve more promising performance than other related methods. Jie Xu 0012, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
KDD | 2 |
| 2018 | Bilevel Distance Metric Learning for Robust Image RecognitionabstractMetric learning, aiming to learn a discriminative Mahalanobis distance matrix M that can effectively reflect the similarity between data samples, has been widely studied in various image recognition problems. Most of the existing metric learning methods input the features extracted directly from the original data in the preprocess phase. What's worse, these features usually take no consideration of the local geometrical structure of the data and the noise existed in the data, thus they may not be optimal for the subsequent metric learning task. In this paper, we integrate both feature extraction and metric learning into one joint optimization framework and propose a new bilevel distance metric learning model. Specifically, the lower level characterizes the intrinsic data structure using graph regularized sparse coefficients, while the upper level forces the data samples from the same class to be close to each other and pushes those from different classes far away. In addition, leveraging the KKT conditions and the alternating direction method (ADM), we derive an efficient algorithm to solve the proposed new model. Extensive experiments on various occluded datasets demonstrate the effectiveness and robustness of our method. Jie Xu 0012, Lei Luo 0001, Cheng Deng 0002, Heng Huang 0001 |
NeurIPS | 2 |
| 2018 | An adaptive line search scheme for approximated nuclear norm based matrix regression
Lei Luo 0001, Qinghua Tu, Jian Yang 0003, Jing-Yu Yang 0001 |
Neurocomputing | 1 |
| 2018 | Nonparametric Bayesian Correlated Group Regression With Applications to Image ClassificationabstractSparse Bayesian learning has emerged as a powerful tool to tackle various image classification tasks. The existing sparse Bayesian models usually use independent Gaussian distribution as the prior knowledge for the noise. However, this assumption often contradicts to the practical observations in which the noise is long tail and pixels containing noise are spatially correlated. To handle the practical noise, this paper proposes to partition the noise image into several 2-D groups and adopt the long-tail distribution, i.e., the scale mixture of the matrix Gaussian distribution, to model each group to capture the intragroup correlation of the noise. Under the nonparametric Bayesian estimation, the low-rank-induced prior and the matrix Gamma distribution prior are imposed on the covariance matrix of each group, respectively, to induce two Bayesian correlated group regression (BCGR) methods. Moreover, the proposed methods are extended to the case with unknown group structure. Our BCGR method provides an effective way to automatically fit the noise distribution and integrates the long-tail attribute and structure information of the practical noise into model. Therefore, the estimated coefficients are better for reconstructing the desired data. We apply BCGR to address image classification task and utilize the learned covariance matrices to construct a grouped Mahalanobis distance to measure the reconstruction residual of each class in the design of a classifier. Experimental results demonstrate the effectiveness of our new BCGR model. Lei Luo 0001, Jian Yang 0003, Bob Zhang 0001, Jielin Jiang, Heng Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Kernel orthogonal Procrustes regression for face recognition across pose
Ying Tai, Jian Yang 0003, Lei Luo 0001, Jianjun Qian |
Neurocomputing | 3 |
| 2017 | Bi-weighted robust matrix regression for face recognition
Jianchun Xie, Jian Yang 0003, Jianjun Qian, Lei Luo 0001 |
Neurocomputing | 4 |
| 2017 | Nuclear Norm Based Matrix Regression with Applications to Face Recognition with Occlusion and Illumination ChangesabstractRecently, regression analysis has become a popular tool for face recognition. Most existing regression methods use the one-dimensional, pixel-based error model, which characterizes the representation error individually, pixel by pixel, and thus neglects the two-dimensional structure of the error image. We observe that occlusion and illumination changes generally lead, approximately, to a low-rank error image. In order to make use of this low-rank structural information, this paper presents a two-dimensional image-matrix-based error model, namely, nuclear norm based matrix regression (NMR), for face representation and classification. NMR uses the minimal nuclear norm of representation error image as a criterion, and the alternating direction method of multipliers (ADMM) to calculate the regression coefficients. We further develop a fast ADMM algorithm to solve the approximate NMR model and show it has a quadratic rate of convergence. We experiment using five popular face image databases: the Extended Yale B, AR, EURECOM, Multi-PIE and FRGC. Experimental results demonstrate the performance advantage of NMR over the state-of-the-art regression-based methods for face recognition in the presence of occlusion and illumination variations. Jian Yang 0003, Lei Luo 0001, Jianjun Qian, Ying Tai, Fanlong Zhang, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Low-Rank Latent Pattern Approximation With Applications to Robust Image ClassificationabstractThis paper develops a novel method to address the structural noise in samples for image classification. Recently, regression-related classification methods have shown promising results when facing the pixelwise noise. However, they become weak in coping with the structural noise due to ignoring of relationships between pixels of noise image. Meanwhile, most of them need to implement the iterative process for computing representation coefficients, which leads to the high time consumption. To overcome these problems, we exploit a latent pattern model called low-rank latent pattern approximation (LLPA) to reconstruct the test image having structural noise. The rank function is applied to characterize the structure of the reconstruction residual between test image and the corresponding latent pattern. Simultaneously, the error between the latent pattern and the reference image is constrained by Frobenius norm to prevent overfitting. LLPA involves a closed-form solution by the virtue of a singular value thresholding operator. The provided theoretic analysis demonstrates that LLPA indeed removes the structural noise during classification task. Additionally, LLPA is further extended to the form of matrix regression by connecting multiple training samples, and alternating direction of multipliers method with Gaussian back substitution algorithm is used to solve the extended LLPA. Experimental results on several popular data sets validate that the proposed methods are more robust to image classification with occlusion and illumination changes, as compared to some existing state-of-the-art reconstruction-based methods and one deep neural network-based method. Shuo Chen 0003, Jian Yang 0003, Lei Luo 0001, Yang Wei 0003, Kaihua Zhang 0001, Ying Tai |
IEEE Trans. Image Process. | 3 |
| 2017 | Robust Image Regression Based on the Extended Matrix Variate Power Exponential Distribution of Dependent NoiseabstractDealing with partial occlusion or illumination is one of the most challenging problems in image representation and classification. In this problem, the characterization of the representation error plays a crucial role. In most current approaches, the error matrix needs to be stretched into a vector and each element is assumed to be independently corrupted. This ignores the dependence between the elements of error. In this paper, it is assumed that the error image caused by partial occlusion or illumination changes is a random matrix variate and follows the extended matrix variate power exponential distribution. This has the heavy tailed regions and can be used to describe a matrix pattern of l × m dimensional observations that are not independent. This paper reveals the essence of the proposed distribution: it actually alleviates the correlations between pixels in an error matrix E and makes E approximately Gaussian. On the basis of this distribution, we derive a Schatten p-norm-based matrix regression model with Lqregularization. Alternating direction method of multipliers is applied to solve this model. To get a closed-form solution in each step of the algorithm, two singular value function thresholding operators are introduced. In addition, the extended Schatten p-norm is utilized to characterize the distance between the test samples and classes in the design of the classifier. Extensive experimental results for image reconstruction and classification with structural noise demonstrate that the proposed algorithm works much more robustly than some existing regression-based methods. Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai, Gui-Fu Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Dual approximated nuclear norm based matrix regression via adaptive line search schemeabstractFace recognition with partial occlusion is one of the urgent and challenging problems in the pattern recognition research. Using the Alternating Direction Method of Multipliers (ADMM), the recently proposed nuclear norm based matrix regression model (NMR) has been shown a great potential in dealing with the structural noise. And yet, ADMM needs to bring into an auxiliary variable and only exploits the convexity of NMR. Compared with ADMM, the gradient based methods are simpler. To make use of these methods, this paper considers the Approximated NMR (ANMR) model. Utilizing the singular value shrinkage operator and strong convexity of ANMR, the dual problem of ANMR (DANMR) is derived and a crucial result is obtained: the primal optimal solution of ANMR can be converted as the matrix function associated with the dual optimal solution. Due to the differentiability of DANMR, an adaptive line search scheme is developed to solve it. This approach combines the advantages of the accelerated gradient technique and adaptive parameters updating strategy. Therefore, a convergence rate of O(1/N2) can be guaranteed. Experimental results show the superiority of the proposed algorithm over some existing methods. Lei Luo 0001, Qinghua Tu, Jian Yang 0003, Yigong Zhang |
ICPR | 1 |
| 2016 | Structural Orthogonal Procrustes Regression for Face Recognition with Pose Variations and MisalignmentabstractRegression based method is a hot topic in the face recognition community and has achieved interesting results when dealing with well-aligned frontal face images. However, most of the existing regression analysis based methods are sensitive to pose variations. In this paper, we firstly introduce the orthogonal Procrustes problem (OPP), which is simple but effective, as a model to handle pose variations in two-dimensional face images. OPP seeks an optimal transformation between two images to correct the pose from one to the other. We integrate OPP into the regression model and propose the structural orthogonal Procrustes regression (SOPR) using the nuclear norm constraint on the error term to keep image's structural information. Moreover, a subject-wise strategy is adopted to address the problem that the gallery images may span over different poses. The proposed model is solved by an efficient iteratively reweighted algorithm and experimental results on popular face databases demonstrate the effectiveness of our method. Ying Tai, Jian Yang 0003, Fanlong Zhang, Yigong Zhang, Lei Luo 0001, Jianjun Qian |
SDM | 5 |
| 2016 | Schatten p-norm based principal component analysis
Heyou Chang, Lei Luo 0001, Jian Yang 0003, Meng Yang 0001 |
Neurocomputing | 2 |
| 2016 | Adaptive noise dictionary construction via IRRPCA for face recognition
Yu Chen 0037, Jian Yang 0003, Lei Luo 0001, Hengmin Zhang, Jianjun Qian, Ying Tai, Jian Zhang 0025 |
Pattern Recognit. | 3 |
| 2016 | Learning discriminative singular value decomposition representation for face recognition
Ying Tai, Jian Yang 0003, Lei Luo 0001, Fanlong Zhang, Jianjun Qian |
Pattern Recognit. | 3 |
| 2016 | Double Low Rank Matrix Recovery for Saliency FusionabstractIn this paper, we address the problem of fusing various saliency detection methods such that the fusion result outperforms each of the individual methods. We observe that the saliency regions shown in different saliency maps are with high probability covering parts of the salient object. With image regions being represented by the saliency values of multiple saliency maps, the object regions have strong correlation and thus lie in a low-dimensional subspace. Meanwhile, most of background regions tend to have lower saliency values in various saliency maps. They are also strongly correlated and lie in a lowdimensional subspace that is independent of the object subspace. Therefore, an image can be represented as the combination of two low rank matrices. To obtain a unified low rank matrix that represents the salient object, this paper presents a double low rank matrix recovery model for saliency fusion. The inference process is formulated as a constrained nuclear norm minimization problem, which is convex and can be solved efficiently with the alternating direction method of multipliers (ADMM). Furthermore, to reduce the computational complexity of the proposed saliency fusion method, a saliency model selection strategy based on the sparse representation is proposed. Experiments on five datasets show that our method consistently outperforms each individual saliency detection approach and other state-of-the-art saliency fusion methods. Junxia Li, Lei Luo 0001, Fanlong Zhang, Jian Yang 0003, Deepu Rajan |
IEEE Trans. Image Process. | 2 |
| 2016 | Tree-Structured Nuclear Norm Approximation With Applications to Robust Face RecognitionabstractStructured sparsity, as an extension of standard sparsity, has shown the outstanding performance when dealing with some highly correlated variables in computer vision and pattern recognition. However, the traditional mixed (L1, L2) or (L1, L∞) group norm becomes weak in characterizing the internal structure of each group since they cannot alleviate the correla-tions between variables. Recently, nuclear norm has been vali-dated to be useful for depicting a spatially structured matrix variable. It considers the global structure of the matrix variable but overlooks the local structure. To combine the advantages of structured sparsity and nuclear norm, this paper presents a tree-structured nuclear norm approximation (TSNA) model as-suming that the representation residual with tree-structured prior is a random matrix variable and follows a dependent matrix dis-tribution. The Extended Alternating Direction Method of Multi-pliers (EADMM) is utilized to solve the proposed model. An effi-cient bound condition based on the extended restricted isometry constants is provided to show the exact recovery of the proposed model under the given noisy case. In addition, TSNA is connected with some newest methods such as sparse representation based classifier (SRC), nuclear-L1 norm joint regression (NL1R) and nuclear norm based matrix regression (NMR), which can be re-garded as the special cases of TSNA. Experiments with face re-construction and recognition demonstrate the benefits of TSNA over other approaches. Lei Luo 0001, Liang Chen 0003, Jian Yang 0003, Jianjun Qian, Bob Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Face Recognition With Pose Variations and Misalignment via Orthogonal Procrustes RegressionabstractA linear regression-based method is a hot topic in face recognition community. Recently, sparse representation and collaborative representation-based classifiers for face recognition have been proposed and attracted great attention. However, most of the existing regression analysis-based methods are sensitive to pose variations. In this paper, we introduce the orthogonal Procrustes problem (OPP) as a model to handle pose variations existed in 2D face images. OPP seeks an optimal linear transformation between two images with different poses so as to make the transformed image best fits the other one. We integrate OPP into the regression model and propose the orthogonal Procrustes regression (OPR) model. To address the problem that the linear transformation is not suitable for handling highly non-linear pose variation, we further adopt a progressive strategy and propose the stacked OPR. As a practical framework, OPR can handle face alignment, pose correction, and face representation simultaneously. We optimize the proposed model via an efficient alternating iterative algorithm, and experimental results on three popular face databases, such as CMU PIE database, CMU Multi-PIE database, and LFW database, demonstrate the effectiveness of our proposed method. Ying Tai, Jian Yang 0003, Yigong Zhang, Lei Luo 0001, Jianjun Qian, Yu Chen 0037 |
IEEE Trans. Image Process. | 4 |
| 2015 | Mixed noise removal by weighted low rank model
Jielin Jiang, Jian Yang 0003, Yan Cui 0007, Lei Luo 0001 |
Neurocomputing | 4 |
| 2015 | Nuclear-L1 norm joint regression for face reconstruction and recognition with mixed noise
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
Pattern Recognit. | 1 |
| 2015 | Robust nuclear norm regularized regression for face recognition with occlusion
Jianjun Qian, Lei Luo 0001, Jian Yang 0003, Fanlong Zhang, Zhouchen Lin |
Pattern Recognit. | 2 |
| 2015 | Matrix Variate Distribution-Induced Sparse Representation for Robust Image ClassificationabstractSparse representation learning has been successfully applied into image classification, which represents a given image as a linear combination of an over-complete dictionary. The classification result depends on the reconstruction residuals. Normally, the images are stretched into vectors for convenience, and the representation residuals are characterized by l2 -norm or l1 -norm, which actually assumes that the elements in the residuals are independent and identically distributed variables. However, it is hard to satisfy the hypothesis when it comes to some structural errors, such as illuminations, occlusions, and so on. In this paper, we represent the image data in their intrinsic matrix form rather than concatenated vectors. The representation residual is considered as a matrix variate following the matrix elliptically contoured distribution, which is robust to dependent errors and has long tail regions to fit outliers. Then, we seek the maximum a posteriori probability estimation solution of the matrix-based optimization problem under sparse regularization. An alternating direction method of multipliers (ADMMs) is derived to solve the resulted optimization problem. The convergence of the ADMM is proven theoretically. Experimental results demonstrate that the proposed method is more effective than the state-of-the-art methods when dealing with the structural errors. Jian Yang 0003, Lei Luo 0001, Jianjun Qian, Wei Xu 0052 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Nuclear-L1 Norm Joint Regression for Face Reconstruction and Recognition
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
ACCV (2) | 1 |
| 2014 | Nuclear Norm Regularized Sparse CodingabstractPartially occluded or illuminated faces pose a significant obstacle for robust, real-world face recognition. The problem of how to characterize the error caused by occlusion or illumination is still a challenging task. There must exist some close relationship between the error metric and error distribution. However, some metric (e.g. Z2-norm) can't characterize this error distribution completely. By some experiments, we found that nuclear norm is more suitable for characterizing the occluded or illuminated error distribution. Thus, a nuclear norm regularized sparse coding model is presented. Such a problem is solved by using ALM (or ADMM). In addition, we use nuclear norm as a metric to characterize the distance between reconstruction samples and classes. The experiments for image classification and face reconstruction demonstrate that our algorithm is robust to some face variations such as occlusion and illumination, and thus can act as a fast solver for matrix regression problem. Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Jing-Yu Yang 0001 |
ICPR | 1 |