VLDB 2026 Research / reviewers in the wild / expert
Yi-Gang Cen
dblp:22/7330 · also Yigang Cen
· DBLP profile ↗
74ranked-venue papers
2as first author
51since 2021 · last 2026
0000-0001-6255-9422ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 26 since 2021Artificial intelligence and machine learning · 33 · 1 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemDNet: Semantic-guided despeckling network for SAR images
Fuyu Bo, Yi Jin 0001, Xiaole Ma, Yi-Gang Cen, Shaohai Hu, Yidong Li |
Expert Syst. Appl. | 4 |
| 2026 | Query-guided predicate decoupling and prototype approximation learning for scene graph generation
Shichao Kan, Yue Zhang 0065, Yi-Gang Cen, Wanru Xu, Yi Jin 0001, Yidong Li |
Expert Syst. Appl. | 4 |
| 2026 | Causal learning with uncertainty-aware transformer for vision-and-language navigation
Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Wangsheng He |
Neurocomputing | 5 |
| 2026 | Integrating spatial features and dynamically learned temporal features via contrastive learning for video temporal grounding in LLM
Peifu Wang, Yixiong Liang, Yi-Gang Cen, Jin Liu 0012, Shichao Kan |
Image Vis. Comput. | 3 |
| 2026 | Progressively multi-scale feature fusion for semantic segmentation
Shichao Kan, Yi-Gang Cen, Qi Cao 0002, Yansen Huang, Ming Zeng 0012 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | DePoint: Improving rotation robustness of 3D point cloud analysis via decreasing entropy
Lu Shi 0004, Gaoyun An, Yi-Gang Cen, Yansen Huang, Fei Gan |
Neural Networks | 3 |
| 2026 | Visual perception-inspired 3D point cloud samplingabstractTask-oriented sampling aims to predict the importance of points of a point cloud to better serve downstream tasks, which has attracted increasing attention in the fields of computer vision and visualization in recent years. However, existing methods cannot sufficiently leverage both global saliency and local saliency cues, resulting in suboptimal performance that requires further improvement. To tackle this challenge, we propose a novel 3D point cloud sampling method inspired by the human visual perception mechanism in this study, which can effectively extract important point cloud subsets from critical regions to better adapt to downstream tasks, thereby maintaining superior sampling performance. The proposed Visual Perception-inspired 3D Point Cloud Sampling (VPI-3DPS) method simulates the human visual system’s dynamic attention-shifting strategy by combining coarse-grained attention-driven sampling with fine-grained detail preservation. This allows our approach to adaptively capture both global context and local details within point cloud data, safeguarding downstream task performance. By leveraging Gated Recurrent Units (GRUs) for long-term dependency modeling and integrating Graph Convolutional Networks (GCNs) to capture local structures, VPI-3DPS obtains an integrated representation of regional correlation and detail awareness. Extensive experiments show that VPI-3DPS outperforms existing methods. Compared to the best-performing approaches, it achieves an average increase of 1.29% in classification accuracy, an average reduction of 13.20% in registration MRE, and an average decrease of 4.29% in Chamfer Distance for reconstruction. Xu Wang 0053, Yi Jin 0001, Hui Yu 0001, Yi-Gang Cen, Yidong Li |
Pattern Recognit. | 4 |
| 2026 | Towards efficient and robust correntropy-based anchor tensor learning for multi-view subspace clustering
Shuqin Wang 0001, Yongli Wang 0004, Fang Qiu, Yongyong Chen, Yi-Gang Cen, Fanghui Zhang |
Signal Process. | 5 |
| 2026 | Vision-Semantics-Label: A New Two-Step Paradigm for Action Recognition With Large Language ModelabstractIn recent years, the rapid advancement of multi-modal large language models has propelled the development of video-based conversation models. Due to their exceptional video understanding capabilities, there is often an expectation that these models can handle all video-related tasks, including action recognition. However, because action recognition datasets typically lack semantic information, limiting the performance of dialogue models. Additionally, as these dialogue models are designed for video understanding, they frequently overlook critical information required for action recognition—continuous motion—in their model architecture and training dataset configurations. To address these challenges, we first propose a novel two-step mapping framework based on large language models, termed “Vision-Semantics-Label” mapping, to better adapt video-based large language models for action recognition. In the first step, we proposed a visual-skeletal collaborative learning large language model (VS-LLM), which utilizes human keypoints to compensate for the missing motion details without increasing the input token length of the large language model. In the second step, we designed two mapping methods: verb noun match (VN-Match) and all text match (ALL-Match), which can effectively extract relevant action descriptions from the text. Finally, we construct semantic action recognition datasets to ensure that the training data inherently contains action details, enabling the model to better achieve action recognition. We evaluate our approach on five benchmark datasets, demonstrating the state-of-the-art performance of large language models in action recognition. The source code and dataset are publicly available at https://github.com/xiaoyu92568/VS-LLM. Wanru Xu, Shichao Kan, Linna Zhang, Yi Jin 0001, Yi-Gang Cen, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | UAGM: Uncertainty-Aware Geometric Modeling for Multi-Scenario 3-D Object Detection in Autonomous VehiclesabstractVision-based 3-D object detection is a core task for autonomous driving and intelligent transport systems. However, the idealized geometric assumptions and single-view modeling methods relied upon by existing methods are susceptible to geometric uncertainties, caused by external parameter perturbations and depth estimation noise in practical multi-scenario deployments. Building a unified and robust perception model suitable for multi-scenarios of both ego-vehicle and roadside remains challenging. To address these challenges, we propose an Uncertainty-Aware Geometric Modeling (UAGM) method that explicitly handles geometric uncertainty to achieve robust perception across multiple scenarios. We use a dual-branch architecture to establish a robust geometric foundation: the height branch introduces dynamic virtual coordinate calibration to compensate for camera extrinsic parameter perturbations in real time, while explicitly modeling height prediction uncertainty through Multi-Hypothesis Projection (MHP), thereby constructing a robust global geometric representation. Meanwhile, the depth branch integrates a Probabilistic Depth Smoothing (PDS) module that employs Conditional Random Fields (CRF) to model spatial consistency constraints, effectively mitigating geometric discontinuities arising from pixel-level predictions. To facilitate better information fusion and interaction, we first propose a Temporal Pyramid Fusion (TPF) module to effectively capture multi-scale spatio-temporal dynamics to reduce the uncertainty in single-frame estimation, instead of error-prone dynamic ego-motion compensation. Subsequently, our Hierarchical Refinement Decoder (HRD) refines BEV proposal localization by fusing image features with depth embeddings to compensate for spatial distortions caused by forward projection. Experimental results demonstrate that UAGM not only achieves state-of-the-art detection performance on both the nuScenes and DAIR-V2X benchmarks, but more importantly, it successfully demonstrates the strong generalization capability of a single model across different viewpoints and deployment conditions. Zhaojie Sun, Wanru Xu, Lu Shi 0004, Yi-Gang Cen, Yi Jin 0001, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | ConfMan Web 3.0: Decentralized Academic Conference Management System with Rust and Web 3.0abstractAcademic conferences serve as a useful platform for researchers and educators to share their work and ideas. Upon the paper acceptance, authors need to register their papers, pay registration fees, and present their work at conferences. International conferences often involve multiple parties of users and cross-border payments with multiple channels that incur additional service charges. Conference organisers and reviewers typically contribute voluntarily without any monetary rewards. The emergence of Web 3.0 technology, leveraging the decentralized and secure nature of blockchain, presents an opportunity for this domain. This research creates a decentralized academic conference management system using Web 3.0 (ConfMan Web 3.0) and Rust programming language that simplifies cross-border payments on different channels by using cryptocurrencies instead of fiat currency, thereby reducing service charges. It aims to provide rewards, recognition of contributions, and consolidated historical records to all conference contributors, including authors, reviewers, programme chairs, and organising committees. The Solana blockchain is used to store conference- related data, and a web application is developed for ConfMan Web 3.0. Various testings are conducted to evaluate its performance. The findings highlight the potential of Web 3.0 technology in transforming the academic conference management landscape. Chian Min Gan, Chee Kiat Seow, Sye Loong Keoh, Dezhong Yao 0002, Yi-Gang Cen, Yiyu Cai, Nisha Jain, Qi Cao 0002 |
COMPSAC | 5 |
| 2025 | Injecting Cross-modal Fine-Grained Perception into LLMs for 3D Object-of-Interest UnderstandingabstractRecent advancements in 3D Large Language Models (LLMs) have revealed significant potential in enhancing the understanding of 3D scenes. However, previous methods have struggled with extracting and utilizing fine-grained information of 3D objects for the coarsness of point clouds, resulting in limitations in understanding object-of-interested (OoI) within the scene. To address this issue, we introduce the object-centric 2D-3D interaction module for enhancing the ability of LLMs for 3D understanding tasks, which consists of the fine-grained 2D representation perception and the object-centric 3D scene representation perception. Specifically, the 2D representation associated with 3D objects is captured based on cross-modal semantic consistency without any spatial projector. Experimental results show that our model significantly outperforms existing methods on benchmarks including ScanRefer and ScanQA. Qianqian Sun, Lu Shi 0004, Linna Zhang, Gaoyun An, Yi Jin 0001, Yidong Li, Yi-Gang Cen |
ICME | 7 |
| 2025 | Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization for Scene Graph GenerationabstractScene Graph Generation (SGG) is a fundamental task in visual understanding, aimed at providing more precise local detail comprehension for downstream applications. Existing SGG methods often overlook the diversity of predicate representations and the consistency among similar predicates when dealing with long-tail distributions. As a result, the model's decision layer fails to effectively capture details from the tail end, leading to biased predictions. To address this, we propose a Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization (NoDIS) method. On the one hand, expanding the predicate representation space enhances the model's ability to learn both common and rare predicates, thus reducing prediction bias caused by data scarcity. We propose a conditional diffusion model to reconstructs features and increase the diversity of representations for same category predicates. On the other hand, independent predicate representations in the decision phase increase the learning complexity of the decision layer, making accurate predictions more challenging. To address this issue, we introduce a discretization mapper that learns consistent representations among similar predicates, reducing the learning difficulty and decision ambiguity in the decision layer. To validate the effectiveness of our method, we integrate NoDIS with various SGG baseline models and conduct experiments on multiple datasets. The results consistently demonstrate superior performance. Shichao Kan, Fanghui Zhang, Wanru Xu, Yue Zhang 0065, Yi-Gang Cen |
ICML | 6 |
| 2025 | Hierarchical Meta-prototypes Network for Few-shot Action RecognitionabstractExisting few-shot action recognition (FSAR) studies predominantly follow a metric learning framework, where prototypes are generated directly from features extracted by an encoder, and classification is performed via distance-based matching. However, due to the limited number of available samples, significant variations exist between different video features of the same class. As a result, the same query video may yield different classification results when matched against different sets of support videos. To address this issue, we propose a novel Hierarchical Meta-Prototypes Network (HMP-Net). The key innovation of our approach lies in the introduction of a category-agnostic and feature-agnostic meta-prototype module, which guides video feature mapping into a more suitable feature space. To optimize this meta-prototype, we design an alternating meta-prototype training strategy, where the model first learns to transform features under a fixed meta-prototype, and then the meta-prototype is refined to better guide feature mapping. Additionally, to adapt image-based metric learning models to video-based FSAR tasks, we introduce a series of lightweight adaptation modules. Specifically, we integrate an adapter into the encoder to improve video frame feature extraction, design a hierarchical prototype generation mechanism to enhance overall video understanding, and incorporate a task-specific perception module to extract unique features for each task. These adaptations make our model better suited for FSAR, significantly improving performance. We evaluate HMP-Net on five challenging benchmarks, and experimental results demonstrate that our model achieves new state-of-the-art performance on HMDB51, UCF101, Kinetics, and SthSthV2-Small. Extensive empirical evaluations further highlight the effectiveness and robustness of HMP-Net. Yi-Gang Cen, Wanru Xu, Yue Zhang 0065, Yi Jin 0001, Yidong Li, Linna Zhang |
ACM Multimedia | 2 |
| 2025 | Tree of Prompts: Aligning Hierarchical Visual Prior for Continual Generalized Category DiscoveryabstractContinual Generalized Category Discovery (C-GCD) aims to incrementally identify both known and novel classes from unlabeled data streams while preserving previously acquired knowledge. However, current approaches face a critical limitation we term unstructured knowledge interference, a critical issue that arises when unconstrained parameter updates entangle discriminative representations across classes, severely contaminating the feature space and introducing significant transfer and bias risks. To address these challenges, we propose the Tree of Prompts (ToP), a novel hierarchical prompting framework that facilitates structured knowledge adaptation through multi-granular parameter regulation. ToP hierarchically integrates three synergistic components: (1) Stage-level prompts preserve historical knowledge by isolating task-specific parameters, thereby mitigating conflicts between incremental tasks; (2) Centroid-level prompts disentangle category semantics through learnable prototype calibration, sharpening decision boundaries in the feature space; and (3) Context-level prompts dynamically capture discriminative local features to suppress contamination from superficial similarities. Experimental results demonstrate that ToP markedly outperforms existing methods and provides a comprehensive and efficient solution for C-GCD. Yiqing Hao, Yangru Huang, Yi Jin 0001, Tao Wang 0011, Yidong Li, Yi-Gang Cen |
ACM Multimedia | 6 |
| 2025 | Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental LearningabstractMultimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modality or task significantly increases inference costs. Existing incremental learning methods focus on task expansion within a single modality, whereas MBIIL seeks to train a unified model incrementally across modalities. The MBIIL faces two challenges: I) How to preserve previously learned knowledge during incremental updates? II) How to effectively leverage knowledge acquired from existing modalities to support new modalities? To address these challenges, we propose MSLoRA-CR, a method that fine-tunes Modality-Specific LoRA modules while incorporating Contrastive Regularization to enhance intra-modality knowledge sharing and promote inter-modality knowledge differentiation. Our approach builds upon a large vision-language model (LVLM), keeping the pretrained model frozen while incrementally adapting new LoRA modules for each modality or task. Experiments on the incremental learning of biomedical images demonstrate that MSLoRA-CR outperforms both the state-of-the-art (SOTA) approach of training separate models for each modality and the general incremental learning method (incrementally fine-tuning LoRA). Specifically, MSLoRA-CR achieves a 1.88% improvement in overall performance compared to unconstrained incremental learning methods while maintaining computational efficiency. Our code is publicly available at https://github.com/VentusAislant/MSLoRA_CR. Yixiong Liang, Hulin Kuang, Yi-Gang Cen, Min Zeng 0004, Shichao Kan |
ACM Multimedia | 6 |
| 2025 | Pedestrian Open-Attribute Recognition via Dynamic Semantic Masking
Yue Zhang 0065, Sen Feng, Fanghui Zhang, Guoqi Liu, Yi-Gang Cen |
PRCV (7) | 6 |
| 2025 | CroCaps: A CLIP-assisted cross-domain video captioner
Wanru Xu, Yenan Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
Expert Syst. Appl. | 4 |
| 2025 | Feature Transformation Reconstruction (FTR) Network for Unsupervised Anomaly DetectionabstractThe goal of the feature reconstruction network based on an autoencoder in the training phase is to force the network to reconstruct the input features well. The network tends to learn shortcuts of “identity mapping,” which leads to the network outputting abnormal features as they are in the inference phase. As such, the abnormal features based on reconstruction error cannot be distinguished from normal features, significantly limiting the detection performance of such methods. To address this issue, we propose a feature transformation reconstruction (FTR) network, which can avoid the identity mapping problem. Specifically, we use a normalizing flow model as a feature transformation (FT) network to transform input features into other forms. The training goal of the feature reconstruction (FR) network is no longer to reconstruct the input features but to reconstruct the transformed features, effectively avoiding the shortcut of learning the “identity map.” Furthermore, this paper proposes a masked convolutional attention (MCA) module, which randomly masks the input features in the training phase and reconstructs the input features in a self‐supervised manner. In the testing phase, the MCA can effectively suppress the excessive reconstruction of abnormal features and further improve anomaly detection performance. FTR achieves the scores of the area under the receiver operating characteristic curve (AUROC) at 99.5% and 97.8% on the MVTec AD and BTAD datasets, respectively, outperforming other state‐of‐the‐art methods. Moreover, FTR is faster than the existing methods, with a high speed of 137 frames per second (FPS) on a 3080ti GPU. Linna Zhang, Lanyao Zhang, Qi Cao 0002, Shichao Kan, Yi-Gang Cen, Fugui Zhang, Yansen Huang |
Int. J. Intell. Syst. | 5 |
| 2025 | Attention redirection transformer with semantic oriented learning for unbiased scene graph generation
Gaoyun An, Yi-Gang Cen, Qiuqi Ruan |
Pattern Recognit. | 3 |
| 2025 | Cross-scene visual context parsing with large vision-language modelabstractRelation analysis is crucial for image-based applications such as visual reasoning and visual question answering . Current relation analysis such as scene graph generation (SGG) only focuses on building relationships among objects within a single image. However, in real-world applications, relationships among objects across multiple images, as seen in video understanding , may hold greater significance as they can capture global information. This is still a challenging and unexplored task. In this paper, we aim to explore the technique of Cross-Scene Visual Context Parsing (CS-VCP) using a large vision-language model. To achieve this, we first introduce a cross-scene dataset comprising 10,000 pairs of cross-scene visual instruction data, with each instruction describing the common knowledge of a pair of cross-scene images. We then propose a Cross-Scene Visual Symbiotic Linkage (CS-VSL) model to understand both cross-scene relationships and objects by analyzing the rationales in each scene. The model is pre-trained on 100,000 cross-scene image pairs and validated on 10,000 image pairs. Both quantitative and qualitative experiments demonstrate the effectiveness of the proposed method. Our method has been released on GitHub: https://github.com/gavin-gqzhang/CS-VSL . Shichao Kan, Lu Shi 0004, Wanru Xu, Gaoyun An, Yi-Gang Cen |
Pattern Recognit. | 6 |
| 2025 | V2PNet: A Voxel-to-Point Network Framework for Task-Oriented Point Cloud SamplingabstractTask-oriented point cloud sampling is a fundamental technique in 3D computer vision and has become a crucial step in numerous 3D applications. However, most state-of-the-art task-oriented sampling methods adopt a point-wise analysis strategy, making them susceptible to data redundancy. Taking inspiration from the abstract-to-detailed recognition process of the human visual system, we propose a novel voxel-to-point network framework called V2PNet for task-oriented point cloud sampling. Specifically, we first design a lightweight coarse-grained sampling module named Important Voxel Prediction (IMVP). This module adaptively outputs points from significant regions of the point cloud by explicitly modeling inter-region relationships, thereby reducing interference from redundant points. Then, the V2PNet framework seamlessly integrates the IMVP module with existing point-wise and task-oriented sampling networks, enabling joint training with downstream tasks. This creates a task-oriented coarse-to-fine-grained sampling pipeline that effectively samples representative and informative points from significant regions to represent the original point cloud. Moreover, to mitigate disturbances across similar regions, we introduce a voxel simplification loss function to enhance the discriminative voxel prediction. Extensive experiments demonstrate that V2PNet improves the performance of existing state-of-the-art task-oriented sampling models. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Multi-Modal Self-Perception Enhanced Large Language Model for 3D Region-of-Interest Captioning With Limited Dataabstract3D Region-of-Interest (RoI) Captioning involves translating a model's understanding of specific objects within a complex 3D scene into descriptive captions. Recent advancements in Large Language Models (LLMs) have shown great potential in this area. Existing methods capture the visual information from RoIs as input tokens for LLMs. However, this approach may not provide enough detailed information for LLMs to generate accurate region-specific captions. In this paper, we introduce Self-RoI, a Large Language Model with multi-modal self-perception capabilities for 3D RoI captioning. To ensure LLMs receive more precise and sufficient information, Self-RoI incorporates Implicit Textual Info. Perception to construct a multi-modal vision-language information. This module utilizes a simple mapping network to generate textual information about basic properties of RoI from vision-following response of LLMs. This textual information is then integrated with the RoI's visual representation to form a comprehensive multi-modal instruction for LLMs. Given the limited availability of 3D RoI-captioning data, we propose a two-stage training strategy to optimize Self-RoI efficiently. In the first stage, we align 3D RoI vision and caption representations. In the second stage, we focus on 3D RoI vision-caption interaction, using a disparate contrastive embedding module to improve the reliability of the implicit textual information and employing language modeling loss to ensure accurate caption generation. Our experiments demonstrate that Self-RoI significantly outperforms previous 3D RoI captioning models. Moreover, the Implicit Textual Info. Perception can be integrated into other multi-modal LLMs for performance enhancement. We will make our code available for further research. Lu Shi 0004, Shichao Kan, Yi Jin 0001, Linna Zhang, Yi-Gang Cen |
IEEE Trans. Multim. | 5 |
| 2025 | LighTN: Light-Weight Transformer Network for Performance-Overhead Tradeoff in Point Cloud DownsamplingabstractDownsampling is a crucial task for processing large scale and/or dense point clouds with limited resources. Owing to the development of deep learning, approaches of task-oriented point cloud downsampling have significant performance gains in preserving geometric information. However, most downsamling methods are limited by the disordered and unstructured point cloud data, making it difficult to continually improve the performance. To address this issue, we propose a light-weight Transformer network (LighTN) for the task-oriented point cloud downsampling as an end-to-end solution. In LighTN, we design an energy-efficient and permutation invariant single-head self-correlation module to extract refined global geometric features. Moreover, we present a novel sampling loss function to guide LighTN to focus on critical point cloud regions with more uniform distributions and prominent point coverage. Extensive experiments on classification, registration, and reconstruction tasks demonstrate that LighTN can achieve the state-of-the-art performance-overhead tradeoff and high-quality qualitative results. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Tao Wang 0011, Bowen Tang 0001, Yidong Li |
IEEE Trans. Multim. | 3 |
| 2025 | Low-Shot Unsupervised Visual Anomaly Detection via Sparse Feature RepresentationabstractVisual anomaly detection is an essential component in modern industrial manufacturing. Existing studies using notions of pairwise similarity distance between a test feature and nominal features have achieved great breakthroughs. However, the absolute similarity distance lacks certain generalizations, making it challenging to extend the comparison beyond the available samples. This limitation could potentially hamper anomaly detection performance in scenarios with limited samples. This article presents a novel sparse feature representation anomaly detection (SFRAD) framework, which formulates the anomaly detection as a sparse feature representation problem; and notably proposes an anomaly score by orthogonal matching pursuit (ASOMP) as a novel detection metric. Specifically, SFRAD calculates the Gaussian kernel distance between the test feature and its sparse representation in the nominal feature space for anomaly detection. Here, the orthogonal matching pursuit (OMP) algorithm is adopted to achieve the sparse feature representation. Moreover, to construct a low-redundancy memory bank storing the basis features for sparse representation, a novel basis feature sampling (BFS) algorithm is proposed by considering both the maximum coverage and the optimum feature representation simultaneously. As a result, SFRAD incorporates both the advantages of absolute similarity and linear representation; and this enhances the generalization in low-shot scenarios. Extensive experiments on the MVTec anomaly detection (MVTec AD), Kolektor surface-defect dataset (KolektorSDD), Kolektor surface-defect dataset 2 (KolektorSDD2), MVTec logical constraints anomaly detection (MVTec LOCO AD), Visual anomaly (VISA), Modified national institute of standards and technology (MNIST), and CIFAR-10 datasets demonstrate that our proposed SFRAD outperforms the previous methods and achieves state-of-the-art unsupervised anomaly detection performance. Notably, significantly improved outcomes and results have also been achieved on low-shot anomaly detection. Code is available at https://github.com/fanghuisky/SFRAD. Fanghui Zhang, Haiyue Zhu, Yi-Gang Cen, Shichao Kan, Linna Zhang, Prahlad Vadakkepat, Tong Heng Lee |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningabstractVisual abduction reasoning aims to find the most plausible explanation for incomplete observations, and suffers from inherent uncertainties and ambiguities, which mainly stem from the latent causal relations, incomplete observations, and the reasoning itself. To address this, we propose a probabilistic model named Uncertainty-Guided Probabilistic Distillation Transformer (UPD-Trans) to model uncertainties for Visual Abductive Reasoning. In order to better discover the correct cause-effect chain, we model all the potential causal relations into a unified reasoning framework, thus both the direct relations and latent relations are considered. In order to reduce the effect of the stochasticity and uncertainty for reasoning: 1) we extend the deterministic Transformer to a probabilistic Transformer by considering those uncertain factors as Gaussian random variables and explicitly modeling their distribution; 2) we introduce a distillation mechanism between the posterior branch with complete observations and the prior branch with incomplete observations to transfer posterior knowledge. Evaluation results on the benchmark datasets, consistently demonstrate the commendable performance of our UPD-Trans, with significant improvements after latent relation modeling and uncertainty modeling. Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
ACM Multimedia | 4 |
| 2024 | Synergetic Prototype Learning Network for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) is an important cross-modal task in scene understanding, aiming to detect visual relations in an image. However, due to the various appearance features, the feature distributions of different categories have suffered from a severe overlap, which makes the decision boundaries ambiguous. The current SGG methods mainly attempt to re-balance the data distribution, which is dataset-dependent and limits the generalization. To solve this problem, a Synergetic Prototype Learning Network (SPLN) is proposed here, where the generalized semantic space is modeled and the synergetic effect among different semantic subspaces is delved into. In SPLN, a Collaboration-induced Prototype Learning method is proposed to model the interaction of visual semantics and structural semantics. The conventional visual semantics is focused on with a residual-driven representation enhancement module to capture details. And the intersection of structural semantics and visual semantics is explicitly modeled as conceptual semantics, which has been ignored by existing methods. Meanwhile, to alleviate the noise of unrelated and meaningless words, an Intersection-induced Prototype Learning method is also proposed specially for conceptual semantics with an essence-driven prototype enhancement module. Moreover, a Selective Fusion Module is proposed to synergetically integrate the results of visual, structural, conceptual branches and the generalized semantics projection. Experiments on VG and GQA datasets show that our method achieves state-of-the-art performance on the unbiased metrics. Ziwei Shang, Zhaoqilin Yang, Shan Cao 0002, Yi-Gang Cen, Gaoyun An |
ACM Multimedia | 6 |
| 2024 | CAST: Cross-Modal Retrieval and Visual Conditioning for image captioning
Shan Cao 0002, Gaoyun An, Yi-Gang Cen, Zhaoqilin Yang, Weisi Lin |
Pattern Recognit. | 3 |
| 2024 | SAR Image Speckle Reduction Based on Nuclear Norm Minus Frobenius Norm RegularizationabstractSynthetic aperture radar (SAR) is a powerful imaging system with all-day and all-weather capabilities, making it suitable for a wide range of applications. However, SAR images often suffer from coherent speckle noise, which degrades image quality and hampers subsequent analysis and interpretation. Recently, methods based on the Fisher-Tippett (FT) distribution and nonlocal low-rank (NLR) techniques have shown great potential in SAR despeckling. Building upon these methods, this article proposes a novel SAR image despeckling method named SAR nuclear norm minus Frobenius norm (SAR-NNFN). This method effectively restores clean images using singular value shrinkage and allows for adaptive shrinkage without the need for additional weighting parameters. SAR-NNFN utilizes NNFN to achieve rank relaxation, resulting in a more robust low-rank solution for speckle reduction. The proposed model comprises two components: a data fidelity term that captures the statistical characteristics of SAR images using the FT distribution in the logarithmic domain, and an NNFN regularization term that enhances low-rank approximations. The optimization problem associated with SAR-NNFN is solved using the alternating direction method of multipliers (ADMM) algorithm. Extensive experiments conducted on both simulated and real SAR images demonstrate that SAR-NNFN can not only adequately suppress speckle noise but also preserve fine textures. Fuyu Bo, Xiaole Ma, Yi-Gang Cen, Shaohai Hu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR. Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan |
ACM Multimedia | 5 |
| 2023 | End-to-end feature diversity person search with rank constraint of cross-class matrix
Yue Zhang 0065, Shuqin Wang 0001, Shichao Kan, Yi-Gang Cen, Linna Zhang |
Neurocomputing | 4 |
| 2023 | Multiscale spatial temporal attention graph convolution network for skeleton-based anomaly behavior detection
Shichao Kan, Fanghui Zhang, Yi-Gang Cen, Linna Zhang, Damin Zhang |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Contrastive Bayesian Analysis for Deep Metric LearningabstractRecent methods for deep metric learning have been focusing on designing different contrastive loss functions between positive and negative pairs of samples so that the learned feature embedding is able to pull positive samples of the same class closer and push negative samples from different classes away from each other. In this work, we recognize that there is a significant semantic gap between features at the intermediate feature layer and class labels at the final output layer. To bridge this gap, we develop a contrastive Bayesian analysis to characterize and model the posterior probabilities of image labels conditioned by their features similarity in a contrastive learning setting. This contrastive Bayesian analysis leads to a new loss function for deep metric learning. To improve the generalization capability of the proposed method onto new classes, we further extend the contrastive Bayesian loss with a metric variance constraint. Our experimental results and ablation studies demonstrate that the proposed contrastive Bayesian metric learning method significantly improves the performance of deep metric learning in both supervised and pseudo-supervised scenarios, outperforming existing methods by a large margin. Shichao Kan, Zhiquan He, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | A graph model-based multiscale feature fitting method for unsupervised anomaly detection
Fanghui Zhang, Shichao Kan, Damin Zhang, Yi-Gang Cen, Linna Zhang, Vladimir Mladenovic |
Pattern Recognit. | 4 |
| 2023 | RS-TNet: point cloud transformer with relation-shape awareness for fine-grained 3D visual processing
Xu Wang 0053, Yuqiao Zeng, Yi Jin 0001, Yi-Gang Cen, Baifu Liu, Shaohua Wan 0001 |
Soft Comput. | 4 |
| 2023 | Robustness Meets Low-Rankness: Unified Entropy and Tensor Learning for Multi-View Subspace ClusteringabstractIn this paper, we develop the weighted error entropy-regularized tensor learning method for multi-view subspace clustering (WETMSC), which integrates the noise disturbance removal and subspace structure discovery into one unified framework. Unlike most existing methods which focus only on the affinity matrix learning for the subspace discovery by different optimization models and simply assume that the noise is independent and identically distributed (i.i.d.), our WETMSC method adopts the weighted error entropy to characterize the underlying noise by assuming that noise is independent and piecewise identically distributed (i.p.i.d.). Meanwhile, WETMSC constructs the self-representation tensor by storing all self-representation matrices from the view dimension, preserving high-order correlation of views based on the tensor nuclear norm. To solve the proposed nonconvex optimization method, we design a half-quadratic (HQ) additive optimization technology and iteratively solve all subproblems under the alternating direction method of multipliers framework. Extensive comparison studies with state-of-the-art clustering methods on real-world datasets and synthetic noisy datasets demonstrate the ascendancy of the proposed WETMSC method. Shuqin Wang 0001, Yongyong Chen, Zhiping Lin 0001, Yi-Gang Cen, Qi Cao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Bi-Nuclear Tensor Schatten-p Norm Minimization for Multi-View Subspace ClusteringabstractMulti-view subspace clustering aims to integrate the complementary information contained in different views to facilitate data representation. Currently, low-rank representation (LRR) serves as a benchmark method. However, we observe that these LRR-based methods would suffer from two issues: limited clustering performance and high computational cost since (1) they usually adopt the nuclear norm with biased estimation to explore the low-rank structures; (2) the singular value decomposition of large-scale matrices is inevitably involved. Moreover, LRR may not achieve low-rank properties in both intra-views and inter-views simultaneously. To address the above issues, this paper proposes the Bi-nuclear tensor Schatten- p norm minimization for multi-view subspace clustering (BTMSC). Specifically, BTMSC constructs a third-order tensor from the view dimension to explore the high-order correlation and the subspace structures of multi-view features. The Bi-Nuclear Quasi-Norm (BiN) factorization form of the Schatten- p norm is utilized to factorize the third-order tensor as the product of two small-scale third-order tensors, which not only captures the low-rank property of the third-order tensor but also improves the computational efficiency. Finally, an efficient alternating optimization algorithm is designed to solve the BTMSC model. Extensive experiments with ten datasets of texts and images illustrate the performance superiority of the proposed BTMSC method over state-of-the-art methods. Shuqin Wang 0001, Zhiping Lin 0001, Qi Cao 0002, Yi-Gang Cen, Yongyong Chen |
IEEE Trans. Image Process. | 4 |
| 2023 | Local Correlation Ensemble with GCN Based on Attention Features for Cross-domain Person Re-IDabstractPerson re-identification (Re-ID) has achieved great success in single-domain. However, it remains a challenging task to adapt a Re-ID model trained on one dataset to another one. Unsupervised domain adaption (UDA) was proposed to migrate a model from a labeled source domain to an unlabeled target domain. The main difference in the cross-domain is different background styles. Although the style transfer approach effectively reduces inter-domain gaps, it ignores the reduction of intra-class differences. Clustering-based pipelines maintain state-of-the-art performance for UDA by learning domain-independent features; however, most existing models do not sufficiently exploit the rich unlabeled samples in target domains due to unsatisfactory clustering. Thus, we propose a novel local correlation ensemble model that focuses on the diversity of intra-class information and the reliability of class centers. Specifically, a pedestrian attention module is proposed to enable the encoder to pay more attention to the person’s features to relieve interference caused by the shared background style. Furthermore, we propose a priority-distance graph convolutional network (PDGCN) module that employs a graph convolutional network network to predict the priority of a node as a class center and then calculates the distance between nodes with high priority values to screen out the class center nodes. Finally, the encoder features (local) and PDGCN features (context-aware) are combined to perform person Re-ID. The results of experiments on the large-scale public Re-ID datasets verified the effectiveness of the proposed method. Yue Zhang 0065, Fanghui Zhang, Yi Jin 0001, Yi-Gang Cen, Viacheslav V. Voronin, Shaohua Wan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Coded Residual Transform for Generalizable Deep Metric LearningabstractA fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded residual transform (CRT) for deep metric learning to significantly improve its generalization capability. Specifically, we learn a set of diversified prototype features, project the feature map onto each prototype, and then encode its features using their projection residuals weighted by their correlation coefficients with each prototype. The proposed CRT method has the following two unique characteristics. First, it represents and encodes the feature map from a set of complimentary perspectives based on projections onto diversified prototypes. Second, unlike existing transformer-based feature representation approaches which encode the original values of features based on global correlation analysis, the proposed coded residual transform encodes the relative differences between the original features and their projected prototypes. Embedding space density and spectral decay analysis show that this multi perspective projection onto diversified prototypes and coded residual representation are able to achieve significantly improved generalization capability in metric learning. Finally, to further enhance the generalization performance, we propose to enforce the consistency on their feature similarity matrices between coded residual transforms with different sizes of projection prototypes and embedding dimensions. Our extensive experimental results and ablation studies demonstrate that the proposed CRT method outperform the state-of-the-art deep metric learning methods by large margins and improving upon the current best method by up to 4.28% on the CUB dataset. Shichao Kan, Yixiong Liang, Min Li 0007, Yi-Gang Cen, Jianxin Wang 0001, Zhihai He |
NeurIPS | 4 |
| 2022 | Nonconvex low-rank and sparse tensor representation for multi-view subspace clustering
Shuqin Wang 0001, Yongyong Chen, Yi-Gang Cen, Linna Zhang, Hengyou Wang, Viacheslav V. Voronin |
Appl. Intell. | 3 |
| 2022 | VSLN: View-aware sphere learning network for cross-view vehicle re-identificationabstractCross-view vehicle Reidentification (ReID) has attracted widespread attention as an increasingly important vision task in intelligent transportation and urban surveillance. Benefiting from Convolutional Neural Network (CNN), recent studies have promoted the development of vehicle ReID by extracting discriminative local features. However, two fundamental challenges of small interclass discrepancy caused by different views and large intraclass distance caused by similar appearance still hinder the performance of cross-view vehicle ReID. In this paper, a novel View-aware Sphere Learning Network (VSLN) is proposed to alleviate the above issues while maintaining the merits of CNN-based approaches to generate view-aware sphere-based features. First, a Sphere Feature Embedding Network (SFEN) is proposed to constrain the images into hypersphere for extracting sphere features. On the other hand, this study presents a sphere similarity triple loss to help SFEN concentrate more on robust and discriminative vehicle parts. Second, since the vehicle images are usually captured from different viewpoints, this study further extends SFEN by introducing a Vehicle Viewpoint Predictor (VVP) combined with global attention mechanism to enlarge the discrepancy of interclass and shorten the distance of intraclass. Moreover, a city-scale data set, named Vehicle from Different Viewpoints, containing image-level viewpoint labels, is collected for training VVP. As a result, the proposed VLSN can achieve 96.31% Top-1 accuracy and 79.46% Top-1 accuracy on VeRi-776 and VRIC data sets, respectively. Overall, extensive experimental results on two benchmark data sets show that the proposed VSLN outperforms state-of-the-art methods. Xu Wang 0053, Yi Jin 0001, Chenning Li, Yi-Gang Cen, Yidong Li |
Int. J. Intell. Syst. | 4 |
| 2022 | A GAN-based input-size flexibility model for single image dehazing
Shichao Kan, Yue Zhang 0065, Fanghui Zhang, Yi-Gang Cen |
Signal Process. Image Commun. | 4 |
| 2022 | Local Semantic Correlation Modeling Over Graph Neural Networks for Deep Feature Embedding and Image RetrievalabstractDeep feature embedding aims to learn discriminative features or feature embeddings for image samples which can minimize their intra-class distance while maximizing their inter-class distance. Recent state-of-the-art methods have been focusing on learning deep neural networks with carefully designed loss functions. In this work, we propose to explore a new approach to deep feature embedding. We learn a graph neural network to characterize and predict the local correlation structure of images in the feature space. Based on this correlation structure, neighboring images collaborate with each other to generate and refine their embedded features based on local linear combination. Graph edges learn a correlation prediction network to predict the correlation scores between neighboring images. Graph nodes learn a feature embedding network to generate the embedded feature for a given image based on a weighted summation of neighboring image features with the correlation scores as weights. Our extensive experimental results under the image retrieval settings demonstrate that our proposed method outperforms the state-of-the-art methods by a large margin, especially for top-1 recalls. Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
IEEE Trans. Image Process. | 2 |
| 2021 | Relative Order Analysis and Optimization for Unsupervised Deep Metric LearningabstractIn unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to image A than image B. In this work, we propose to explore how this relative order can be used to learn discriminative features with an unsupervised metric learning method. Instead of resorting to clustering or self-supervision to create pseudo labels for an absolute decision, which often suffers from high label error rates, we construct reliable relative orders for groups of image samples and learn a deep neural network to predict these relative orders. During training, this relative order prediction network and the feature embedding network are tightly coupled, providing mutual constraints to each other to improve overall metric learning performance in a cooperative manner. During testing, the predicted relative orders are used as constraints to optimize the generated features and refine their feature distance-based image retrieval results using a constrained optimization procedure. Our experimental results demonstrate that the proposed relative orders for unsupervised learning (ROUL) method is able to significantly improve the performance ofunsupervised deep metric learning. Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
CVPR | 2 |
| 2021 | PST-NET: Point Cloud Sampling via Point-Based Transformer
Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Congyan Lang, Yidong Li |
ICIG (3) | 3 |
| 2021 | Cross-domain Person Re-identification Based on the Sample Relation Guidance
Yue Zhang 0065, Fanghui Zhang, Shichao Kan, Linna Zhang, Jiaping Zong, Yi-Gang Cen |
ICIG (2) | 6 |
| 2021 | Block-based image matching for image retrieval
Ruizhen Zhao, Liequan Liang, Xinwei Zheng, Yi-Gang Cen, Shichao Kan |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Error-robust low-rank tensor approximation for multi-view clustering
Shuqin Wang 0001, Yongyong Chen, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Linna Zhang |
Knowl. Based Syst. | 4 |
| 2021 | Pedestrian detection with super-resolution reconstruction for low-quality image
Yi Jin 0001, Yue Zhang 0065, Yi-Gang Cen, Yidong Li, Vladimir Mladenovic, Viacheslav V. Voronin |
Pattern Recognit. | 3 |
| 2021 | Zero-Shot Learning to Index on Semantic Trees for Scalable Image RetrievalabstractIn this study, we develop a new approach, called zero-shot learning to index on semantic trees (LTI-ST), for efficient image indexing and scalable image retrieval. Our method learns to model the inherent correlation structure between visual representations using a binary semantic tree from training images which can be effectively transferred to new test images from unknown classes. Based on predicted correlation structure, we construct an efficient indexing scheme for the whole test image set. Unlike existing image index methods, our proposed LTI-ST method has the following two unique characteristics. First, it does not need to analyze the test images in the query database to construct the index structure. Instead, it is directly predicted by a network learnt from the training set. This zero-shot capability is critical for flexible, distributed, and scalable implementation and deployment of the image indexing and retrieval services at large scales. Second, unlike the existing distance-based index methods, our index structure is learnt using the LTI-ST deep neural network with binary encoding and decoding on a hierarchical semantic tree. Our extensive experimental results on benchmark datasets and ablation studies demonstrate that the proposed LTI-ST method outperforms existing index methods by a large margin while providing the above new capabilities which are highly desirable in practice. Shichao Kan, Yi-Gang Cen, Vladimir Mladenovic, Yang Li 0091, Zhihai He |
IEEE Trans. Image Process. | 3 |
| 2021 | Snowball: Iterative Model Evolution and Confident Sample Discovery for Semi-Supervised Learning on Very Small Labeled DatasetsabstractIn this work, we develop a joint sample discovery and iterative model evolution method for semi-supervised learning on very small labeled training sets. We propose a master-teacher-student model framework to provide multi-layer guidance during the model evolution process with multiple iterations and generations. The teacher model is constructed by performing an exponential moving average of the student models obtained from past training steps. The master network combines the knowledge of the student and teacher models with additional access to newly discovered samples. The master and teacher models are then used to guide the training of the student network by enforcing the consistency between their predictions of unlabeled samples and evolve all models when more and more samples are discovered. Our extensive experiments demonstrate that the process of discovering confident samples from the unlabeled dataset, once coupled with the master-teacher-student network evolution, can significantly improve the overall semi-supervised learning performance. For example, on the CIFAR-10 dataset, with a small set of 250 labeled samples, our method achieves an error rate of 11.58%, more than 38% lower than Mean-Teacher (49.91%). When coupled with the MixMatch augmentation and loss function, the improvements are also significant. Yang Li 0091, Zhiqun Zhao, Hao Sun 0024, Yi-Gang Cen, Zhihai He |
IEEE Trans. Multim. | 4 |
| 2020 | Graph-regularized least squares regression for multi-view subspace clustering
Yongyong Chen, Shuqin Wang 0001, Fangying Zheng, Yi-Gang Cen |
Knowl. Based Syst. | 4 |
| 2020 | Metric learning-based kernel transformer with triplets and label constraints for feature fusion
Shichao Kan, Linna Zhang, Zhihai He, Yi-Gang Cen, Shiming Chen 0001, Jikun Zhou |
Pattern Recognit. | 4 |
| 2020 | Abnormal event detection in surveillance videos based on low-rank and compact coefficient dictionary learning
Zhenjiang Miao, Yi-Gang Cen, Xiao-Ping Zhang 0002, Linna Zhang, Shiming Chen 0001 |
Pattern Recognit. | 3 |
| 2020 | Multi-Matrices Low-Rank Decomposition With Structural Smoothness for Image DenoisingabstractIn this paper, we propose a multi-matrices lowrank decomposition method for image denoising. In this new method, the total variation (TV) norm is incorporated into lowrank approximation analysis to achieve structural smoothness and to improve quality of the recovered images. Our proposed mathematical framework for multi-matrices low-rank decomposition combines the nuclear norm, TV norm, and L1norm, which allows us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Based on the iterative alternating direction method, we develop an algorithm to solve the proposed challenging optimization problem. We conduct extensive experiments and perform evaluations on multi-images denoising and multi-frames video prediction. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for images with large sparse noise. Hengyou Wang, Yang Li 0091, Yi-Gang Cen, Zhihai He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Dual-scale weighted structural local sparse appearance model for object trackingabstractIt is a great challenge to develop an effective appearance model for robust visual tracking due to various interfering factors, such as pose change, occlusion, background clutter etc. More and more visual tracking methods tend to exploit the local appearance model to deal with the above challenges. In this study, the authors present a simple yet effective weighted structural local sparse appearance model, which can better describe the target appearance information through patch‐based generative weight. To further improve the robustness of tracking, they implement this appearance model on two‐scale patches. The two derived appearance models are then combined to form a collaborative model to play their advantages. Extensive experiments on the tracking benchmark dataset show that the proposed method performs favourably against several state‐of‐the‐art methods. Xianyou Zeng, Long Xu 0001, Yi-Gang Cen, Ruizhen Zhao, Wanli Feng |
IET Comput. Vis. | 3 |
| 2019 | Robust discriminant low-rank representation for subspace clustering
Gaoyun An, Yi-Gang Cen, Hengyou Wang, Ruizhen Zhao |
Soft Comput. | 3 |
| 2019 | A supervised learning to index model for approximate nearest neighbor image retrieval
Shichao Kan, Xinwei Zheng, Yi-Gang Cen, Zhenmin Zhu, Hengyou Wang |
Signal Process. Image Commun. | 4 |
| 2019 | Supervised Deep Feature Embedding With Handcrafted FeatureabstractImage representation methods based on deep convolutional neural networks (CNNs) have achieved the state-of-the-art performance in various computer vision tasks, such as image retrieval and person re-identification. We recognize that more discriminative feature embeddings can be learned with supervised deep metric learning and handcrafted features for image retrieval and similar applications. In this paper, we propose a new supervised deep feature embedding with a handcrafted feature model. To fuse handcrafted feature information into CNNs and realize feature embeddings, a general fusion unit is proposed (called Fusion-Net). We also define a network loss function with image label information to realize supervised deep metric learning. Our extensive experimental results on the Stanford online products' data set and the in-shop clothes retrieval data set demonstrate that our proposed methods outperform the existing state-of-the-art methods of image retrieval by a large margin. Moreover, we also explore the applications of the proposed methods in person re-identification and vehicle re-identification; the experimental results demonstrate both the effectiveness and efficiency of the proposed methods. Shichao Kan, Yi-Gang Cen, Zhihai He, Zhi Zhang 0005, Linna Zhang |
IEEE Trans. Image Process. | 2 |
| 2018 | Multi-separable dictionary learning
Fengzhen Zhang, Yi-Gang Cen, Ruizhen Zhao, Shaohai Hu, Vladimir Mladenovic |
Signal Process. | 2 |
| 2018 | Reweighted Low-Rank Matrix Analysis With Structural Smoothness for Image DenoisingabstractIn this paper, we develop a new low-rank matrix recovery algorithm for image denoising. We incorporate the total variation (TV) norm and the pixel range constraint into the existing reweighted low-rank matrix analysis to achieve structural smoothness and to significantly improve quality in the recovered image. Our proposed mathematical formulation of the low-rank matrix recovery problem combines the nuclear norm, TV norm, and norm, thereby allowing us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Using the iterative alternating direction and fast gradient projection methods, we develop an algorithm to solve the proposed challenging non-convex optimization problem. We conduct extensive performance evaluations on single-image denoising, hyper-spectral image denoising, and video background modeling from corrupted images. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for large random noise. For example, when the density of random sparse noise is 30%, for single-image denoising, our proposed method is able to improve the quality of the restored image by up to 4.21 dB over existing methods. Hengyou Wang, Yi-Gang Cen, Zhiquan He, Zhihai He, Ruizhen Zhao, Fengzhen Zhang |
IEEE Trans. Image Process. | 2 |
| 2017 | Separable vocabulary and feature fusion for image retrieval based on sparse representation
Yi-Gang Cen, Ruizhen Zhao, Shaohai Hu, Viacheslav V. Voronin, Hengyou Wang |
Neurocomputing | 2 |
| 2017 | Fast smooth rank function approximation based on matrix tri-factorization
Hengyou Wang, Yi-Gang Cen, Ruizhen Zhao, Viacheslav V. Voronin, Fengzhen Zhang |
Neurocomputing | 2 |
| 2017 | Analytic separable dictionary learning based on oblique manifold
Fengzhen Zhang, Yi-Gang Cen, Ruizhen Zhao, Hengyou Wang, Lihong Cui, Shaohai Hu |
Neurocomputing | 2 |
| 2017 | SURF binarization and fast codebook construction for image retrieval
Shichao Kan, Yi-Gang Cen, Viacheslav V. Voronin, Vladimir Mladenovic, Ming Zeng 0012 |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Anomaly detection using sparse reconstruction in crowded scenes
Zhenjiang Miao, Yi-Gang Cen |
Multim. Tools Appl. | 3 |
| 2017 | Robust Generalized Low-Rank Decomposition of Multimatrices for Image RecoveryabstractLow-rank approximation has been successfully used for dimensionality reduction, image noise removal, and image restoration. In existing work, input images are often reshaped to a matrix of vectors before low-rank decomposition. It has been observed that this procedure will destroy the inherent two-dimensional correlation within images. To address this issue, the generalized low-rank approximation of matrices (GLRAM) method has been recently developed, which is able to perform low-rank decomposition of multiple matrices directly without the need for vector reshaping. In this paper, we propose a new robust generalized low-rank matrices decomposition method, which further extends the existing GLRAM method by incorporating rank minimization into the decomposition process. Specifically, our method aims to minimize the sum of nuclear norms and l1-norms. We develop a new optimization method, called alternating direction matrices tri-factorization method, to solve the minimization problem. We mathematically prove the convergence of the proposed algorithm. Our extensive experimental results demonstrate that our method significantly outperforms existing GLRAM methods. Hengyou Wang, Yi-Gang Cen, Zhihai He, Ruizhen Zhao, Fengzhen Zhang |
IEEE Trans. Multim. | 2 |
| 2016 | Abnormal event detection based on sparse reconstruction in crowded scenesabstractIn this paper, we propose an algorithm of abnormal event detection in crowded scenes using sparse representation over the bases of normal motion feature descriptors. To construct an over-complete dictionary, we extract the histogram of maximal optical flow projection (HMOFP) feature from a set of normal training frames. Then the K-SVD dictionary training method is used to get a redundant dictionary after a process of selecting the training samples, which is better than the dictionary simply composed by the HMOFP feature of the whole training frames. In order to detect whether a frame is normal or not, we use the U-norm of the sparse reconstruction coefficients (i.e., the sparse reconstruction cost, SRC) to show the anomaly of the testing frame, which is simple but very effective. The experiment results on UMN dataset and the comparison to the state-of-the-art methods show that our algorithm is promising. Zhenjiang Miao, Yi-Gang Cen, Qinghua Liang |
ICASSP | 3 |
| 2016 | Global anomaly detection in crowded scenes based on optical flow saliencyabstractIn this paper, an algorithm of global anomaly detection in crowded scenes using the saliency in optical flow field is proposed. Before the process of extracting the histogram of maximal optical flow projection (HMOFP), the scale invariant feature transforms (SIFT) method is utilized to get the saliency map of optical flow field. On the basis of the HMOFP feature of normal frames, the online dictionary learning algorithm is used to train an optimal dictionary with proper redundancy after a process of selecting the training samples, which is better than the dictionary simply composed by the HMOFP feature of the whole training frames. In order to detect whether a frame is normal or not, we use the ℓ1-norm of the sparse reconstruction coefficients (i.e., the sparse reconstruction cost, SRC) to show the anomaly of the testing frame, which is simple but very effective. The experiment results on UMN dataset and the comparison to the state-of-the-art methods show that our algorithm is promising. Zhenjiang Miao, Yi-Gang Cen |
MMSP | 3 |
| 2016 | Video restoration based on PatchMatch and reweighted low-rank matrix recovery
Bo-Hua Xu, Yi-Gang Cen, Ruizhen Zhao, Zhenjiang Miao |
Multim. Tools Appl. | 2 |
| 2015 | Pedestrian detection based on Region Proposal FusionabstractAlmost all existing state-of-the-art pedestrian detection methods use combination of hand-crafted features, which cannot well handle the particular challenges in real-world situation. In this paper, we take advantage of Regions with Convolution Neural Networks features (R-CNN) to extract more robust pedestrian features for effective pedestrian detection in complicated environments. To further improve the performance: 1) we propose a Region Proposal Fusion algorithm to get effective region proposals since after careful observation, we found that the quality of region proposals is crucially important for detection performance. 2) we exploit a pedestrian detection expansion method based on image retrieval with color moment features due to R-CNN's requirements of large number of training samples to avoid overfitting. Consequently, the final average miss rate is greatly reduced to 23% in the INRIA pedestrian detection dataset, which is much (23%) lower than that of original HOG (46%). Sheng Tang, Ruizhen Zhao, Yi-Gang Cen |
MMSP | 5 |
| 2015 | Defect inspection for TFT-LCD images based on the low-rank matrix reconstruction
Yi-Gang Cen, Ruizhen Zhao, Lihong Cui, Zhenjiang Miao |
Neurocomputing | 1 |
| 2014 | A new approach of conditions on δ 2s (Φ) for s-sparse recovery
Yi-Gang Cen, Ruizhen Zhao, Zhenjiang Miao, Lihong Cui |
Sci. China Inf. Sci. | 1 |
| 2014 | Rank adaptive atomic decomposition for low-rank matrix completion and its application on image recovery
Hengyou Wang, Ruizhen Zhao, Yi-Gang Cen |
Neurocomputing | 3 |