VLDB 2026 Research / reviewers in the wild / expert
Yi Jin 0001
dblp:38/4674-1
· DBLP profile ↗
91ranked-venue papers
7as first author
64since 2021 · last 2026
0000-0001-8408-3816ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-author · 35 since 2021Artificial intelligence and machine learning · 37 · 1 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Single-Speed Reasoning: Coordinating Fast and Slow Dynamics for Efficient World ModelingabstractModel-based reinforcement learning (MBRL) enables efficient decision-making by learning predictive world modelsof environment dynamics. Despite recent advances, existingmodels often struggle to reconcile accurate short-term transitions with coherent long-term planning, especially in partially observable or long-horizon settings. We argue that thislimitation often stems from modeling all transitions at a single temporal resolution, which makes it challenging to simultaneously capture fine-grained local dynamics and abstractglobal structures. To this end, we propose SF-RSSM (Slow-Fast Recurrent State-Space Model), a novel method that decouples short-term and long-term dynamics via a dualbranchdesign. The fast branch captures short-horizon transitions using residual prediction, while the slow branch models long-range dependencies with a GRU-based recurrent pathway.A distillation mechanism is developed to enable cooperationacross timescales, with the slow model providing soft targetsto guide the fast model. Additionally, a curiosity module encourages exploration by promoting learning in regions wherethe fast and slow branches exhibit divergent dynamics. Experiments on CARLA, DMControl and Atari benchmarks showthat SF-RSSM outperforms strong baselines in policy performance. Yangru Huang, Xu Wang 0053, Yi Jin 0001 |
AAAI | 5 |
| 2026 | SemDNet: Semantic-guided despeckling network for SAR images
Fuyu Bo, Yi Jin 0001, Xiaole Ma, Yi-Gang Cen, Shaohai Hu, Yidong Li |
Expert Syst. Appl. | 2 |
| 2026 | Query-guided predicate decoupling and prototype approximation learning for scene graph generation
Shichao Kan, Yue Zhang 0065, Yi-Gang Cen, Wanru Xu, Yi Jin 0001, Yidong Li |
Expert Syst. Appl. | 6 |
| 2026 | Adaptive spatial-temporal graph ODE networks for traffic flow forecasting
Shixiang Han, Xu Wang 0053, Yi Jin 0001, Songhe Feng, Congyan Lang, Yidong Li |
Multim. Syst. | 3 |
| 2026 | Visual perception-inspired 3D point cloud samplingabstractTask-oriented sampling aims to predict the importance of points of a point cloud to better serve downstream tasks, which has attracted increasing attention in the fields of computer vision and visualization in recent years. However, existing methods cannot sufficiently leverage both global saliency and local saliency cues, resulting in suboptimal performance that requires further improvement. To tackle this challenge, we propose a novel 3D point cloud sampling method inspired by the human visual perception mechanism in this study, which can effectively extract important point cloud subsets from critical regions to better adapt to downstream tasks, thereby maintaining superior sampling performance. The proposed Visual Perception-inspired 3D Point Cloud Sampling (VPI-3DPS) method simulates the human visual system’s dynamic attention-shifting strategy by combining coarse-grained attention-driven sampling with fine-grained detail preservation. This allows our approach to adaptively capture both global context and local details within point cloud data, safeguarding downstream task performance. By leveraging Gated Recurrent Units (GRUs) for long-term dependency modeling and integrating Graph Convolutional Networks (GCNs) to capture local structures, VPI-3DPS obtains an integrated representation of regional correlation and detail awareness. Extensive experiments show that VPI-3DPS outperforms existing methods. Compared to the best-performing approaches, it achieves an average increase of 1.29% in classification accuracy, an average reduction of 13.20% in registration MRE, and an average decrease of 4.29% in Chamfer Distance for reconstruction. Xu Wang 0053, Yi Jin 0001, Hui Yu 0001, Yi-Gang Cen, Yidong Li |
Pattern Recognit. | 2 |
| 2026 | GrassNet: State space model meets graph neural network
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling |
Pattern Recognit. | 3 |
| 2026 | Vision-Semantics-Label: A New Two-Step Paradigm for Action Recognition With Large Language ModelabstractIn recent years, the rapid advancement of multi-modal large language models has propelled the development of video-based conversation models. Due to their exceptional video understanding capabilities, there is often an expectation that these models can handle all video-related tasks, including action recognition. However, because action recognition datasets typically lack semantic information, limiting the performance of dialogue models. Additionally, as these dialogue models are designed for video understanding, they frequently overlook critical information required for action recognition—continuous motion—in their model architecture and training dataset configurations. To address these challenges, we first propose a novel two-step mapping framework based on large language models, termed “Vision-Semantics-Label” mapping, to better adapt video-based large language models for action recognition. In the first step, we proposed a visual-skeletal collaborative learning large language model (VS-LLM), which utilizes human keypoints to compensate for the missing motion details without increasing the input token length of the large language model. In the second step, we designed two mapping methods: verb noun match (VN-Match) and all text match (ALL-Match), which can effectively extract relevant action descriptions from the text. Finally, we construct semantic action recognition datasets to ensure that the training data inherently contains action details, enabling the model to better achieve action recognition. We evaluate our approach on five benchmark datasets, demonstrating the state-of-the-art performance of large language models in action recognition. The source code and dataset are publicly available at https://github.com/xiaoyu92568/VS-LLM. Wanru Xu, Shichao Kan, Linna Zhang, Yi Jin 0001, Yi-Gang Cen, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | UAGM: Uncertainty-Aware Geometric Modeling for Multi-Scenario 3-D Object Detection in Autonomous VehiclesabstractVision-based 3-D object detection is a core task for autonomous driving and intelligent transport systems. However, the idealized geometric assumptions and single-view modeling methods relied upon by existing methods are susceptible to geometric uncertainties, caused by external parameter perturbations and depth estimation noise in practical multi-scenario deployments. Building a unified and robust perception model suitable for multi-scenarios of both ego-vehicle and roadside remains challenging. To address these challenges, we propose an Uncertainty-Aware Geometric Modeling (UAGM) method that explicitly handles geometric uncertainty to achieve robust perception across multiple scenarios. We use a dual-branch architecture to establish a robust geometric foundation: the height branch introduces dynamic virtual coordinate calibration to compensate for camera extrinsic parameter perturbations in real time, while explicitly modeling height prediction uncertainty through Multi-Hypothesis Projection (MHP), thereby constructing a robust global geometric representation. Meanwhile, the depth branch integrates a Probabilistic Depth Smoothing (PDS) module that employs Conditional Random Fields (CRF) to model spatial consistency constraints, effectively mitigating geometric discontinuities arising from pixel-level predictions. To facilitate better information fusion and interaction, we first propose a Temporal Pyramid Fusion (TPF) module to effectively capture multi-scale spatio-temporal dynamics to reduce the uncertainty in single-frame estimation, instead of error-prone dynamic ego-motion compensation. Subsequently, our Hierarchical Refinement Decoder (HRD) refines BEV proposal localization by fusing image features with depth embeddings to compensate for spatial distortions caused by forward projection. Experimental results demonstrate that UAGM not only achieves state-of-the-art detection performance on both the nuScenes and DAIR-V2X benchmarks, but more importantly, it successfully demonstrates the strong generalization capability of a single model across different viewpoints and deployment conditions. Zhaojie Sun, Wanru Xu, Lu Shi 0004, Yi-Gang Cen, Yi Jin 0001, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Federated Privacy Re-identification via Frequency Domain Splitting
Xuanwen Su, Xu Wang 0053, Tengfei Liang, Yi Jin 0001, Yidong Li |
ICIG (3) | 4 |
| 2025 | CFF: Coarse-to-Fine-to-Fusion Semantic Prototype Generation for Zero-Shot ClassificationabstractZero-Shot Learning focuses on recognizing images from unseen classes with the model trained only on seen classes and auxiliary information. Auxiliary information represents the semantic concepts of classes and is crucial for unseen class generalization. Existing works have tried human-annotated attributes, word embedding, or texts as auxiliary information. However, these auxiliary information is semantically insufficient and visually misaligned, constraining model performance. In this work, we propose the Coarse-to-Fine-to-Fusion prototype generation network (CFF). To obtain vision-oriented text corpora, we design the Coarse-to-Fine Text Generation (CFTG) paradigm, utilizing large language models to generate coarse- and fine-grained texts. For text embedding and fusion, we propose the Semantic Prototype Generation (SPG) module, fusing learnable prompts with coarse- and fine-grained embedding, enabling fine-grained and fusion prototype generation. Moreover, we propose the Visual-Semantic Alignment (VSA) loss to narrow the domain gap. Extensive experiments on AWA2, SUN, and CUB datasets demonstrate the effectiveness of our method. Xuanwen Su, Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li |
ICME | 4 |
| 2025 | Injecting Cross-modal Fine-Grained Perception into LLMs for 3D Object-of-Interest UnderstandingabstractRecent advancements in 3D Large Language Models (LLMs) have revealed significant potential in enhancing the understanding of 3D scenes. However, previous methods have struggled with extracting and utilizing fine-grained information of 3D objects for the coarsness of point clouds, resulting in limitations in understanding object-of-interested (OoI) within the scene. To address this issue, we introduce the object-centric 2D-3D interaction module for enhancing the ability of LLMs for 3D understanding tasks, which consists of the fine-grained 2D representation perception and the object-centric 3D scene representation perception. Specifically, the 2D representation associated with 3D objects is captured based on cross-modal semantic consistency without any spatial projector. Experimental results show that our model significantly outperforms existing methods on benchmarks including ScanRefer and ScanQA. Qianqian Sun, Lu Shi 0004, Linna Zhang, Gaoyun An, Yi Jin 0001, Yidong Li, Yi-Gang Cen |
ICME | 5 |
| 2025 | Generative Adversarial Network-based Image and Tabular Data Generation with Differential PrivacyabstractMachine learning and artificial intelligence technologies have become integral to various industries, driven by the availability of large-scale data. However, the use of sensitive industrial and personal data introduces significant privacy risks. Generative Adversarial Networks (GANs) are employed to generate synthetic data, thereby rendering them a feasible privacy-preserving technique. Despite their potential, existing private GANs face two major limitations: they are restricted to single-modal data generation or fail to ensure strict privacy guarantees. To tackle these issues, we propose Differential Privacy GAN of Image and Table—DPGAN-IT, which is capable of generating both image and tabular data simultaneously while enforcing robust privacy protection. Extensive experiments validate high utility of multi-modal synthetic data and demonstrate an effective balance between privacy and data utility. Jiming Yang, Xu Wang 0053, Yi Jin 0001, Yidong Li, Hui Yu 0001 |
ICME | 3 |
| 2025 | Hierarchical Meta-prototypes Network for Few-shot Action RecognitionabstractExisting few-shot action recognition (FSAR) studies predominantly follow a metric learning framework, where prototypes are generated directly from features extracted by an encoder, and classification is performed via distance-based matching. However, due to the limited number of available samples, significant variations exist between different video features of the same class. As a result, the same query video may yield different classification results when matched against different sets of support videos. To address this issue, we propose a novel Hierarchical Meta-Prototypes Network (HMP-Net). The key innovation of our approach lies in the introduction of a category-agnostic and feature-agnostic meta-prototype module, which guides video feature mapping into a more suitable feature space. To optimize this meta-prototype, we design an alternating meta-prototype training strategy, where the model first learns to transform features under a fixed meta-prototype, and then the meta-prototype is refined to better guide feature mapping. Additionally, to adapt image-based metric learning models to video-based FSAR tasks, we introduce a series of lightweight adaptation modules. Specifically, we integrate an adapter into the encoder to improve video frame feature extraction, design a hierarchical prototype generation mechanism to enhance overall video understanding, and incorporate a task-specific perception module to extract unique features for each task. These adaptations make our model better suited for FSAR, significantly improving performance. We evaluate HMP-Net on five challenging benchmarks, and experimental results demonstrate that our model achieves new state-of-the-art performance on HMDB51, UCF101, Kinetics, and SthSthV2-Small. Extensive empirical evaluations further highlight the effectiveness and robustness of HMP-Net. Yi-Gang Cen, Wanru Xu, Yue Zhang 0065, Yi Jin 0001, Yidong Li, Linna Zhang |
ACM Multimedia | 5 |
| 2025 | Tree of Prompts: Aligning Hierarchical Visual Prior for Continual Generalized Category DiscoveryabstractContinual Generalized Category Discovery (C-GCD) aims to incrementally identify both known and novel classes from unlabeled data streams while preserving previously acquired knowledge. However, current approaches face a critical limitation we term unstructured knowledge interference, a critical issue that arises when unconstrained parameter updates entangle discriminative representations across classes, severely contaminating the feature space and introducing significant transfer and bias risks. To address these challenges, we propose the Tree of Prompts (ToP), a novel hierarchical prompting framework that facilitates structured knowledge adaptation through multi-granular parameter regulation. ToP hierarchically integrates three synergistic components: (1) Stage-level prompts preserve historical knowledge by isolating task-specific parameters, thereby mitigating conflicts between incremental tasks; (2) Centroid-level prompts disentangle category semantics through learnable prototype calibration, sharpening decision boundaries in the feature space; and (3) Context-level prompts dynamically capture discriminative local features to suppress contamination from superficial similarities. Experimental results demonstrate that ToP markedly outperforms existing methods and provides a comprehensive and efficient solution for C-GCD. Yiqing Hao, Yangru Huang, Yi Jin 0001, Tao Wang 0011, Yidong Li, Yi-Gang Cen |
ACM Multimedia | 3 |
| 2025 | Differential Contrastive Training for Gaze EstimationabstractThe complex application scenarios have raised critical requirements for precise and generalizable gaze estimation methods. Recently, the pre-trained CLIP has achieved remarkable performance on various vision tasks, but its potentials have not been fully exploited in gaze estimation. In this paper, we propose a novel Differential Contrastive Training strategy, which boosts gaze estimation performance with the help of the CLIP. Accordingly, a Differential Contrastive Gaze Estimation network (DCGaze) composed of a Visual Appearance-aware branch and a Semantic Differential-aware branch is introduced. The Visual Appearance-aware branch is essentially a primary gaze estimation network and it incorporates an Adaptive Feature-refinement Unit (AFU) and a Double-head Gaze Regressor (DGR), which both help the primary network to extract informative and gaze-related appearance features. Moreover, the Semantic Difference-aware branch is designed on the basis of the CLIP's text encoder to reveal the semantic difference of gazes. This branch could further empower the Visual Appearance-aware branch with the capability of characterizing the gaze-related semantic information. Extensive experimental results on four challenging datasets over within and cross-domain tasks demonstrate the effectiveness of our DCGaze. The code is available at https://github.com/LinZhang-bjtu/DCGaze. Xiyun Wang, Wanru Xu, Yi Jin 0001 |
ACM Multimedia | 5 |
| 2025 | Open World Adaptive Pseudo Contrastive Learning for Generalized Category DiscoveryabstractIn this work, we investigate the challenging task of Generalized Category Discovery (GCD). Given datasets collected from open-world scenarios comprising both labeled and unlabeled images, GCD aims to classify all unlabeled images while simultaneously identifying unlabeled novel categories. The fundamental challenge in GCD tasks stems from inherent annotation discrepancies between seen and novel classes within the dataset. The lack of reliable label supervision for novel classes in unlabeled data leads to significant disparities in the model’s learning between old and novel classes, which is termed the bias risk. Recent advancements in GCD have employed the entropy maximization algorithm to alleviate the bias risk. However, they fail to provide debiased optimization for unlabeled data, leading to models that struggle with extracting discriminative features from such data. To address these challenges, we have created an Open-world pseudo-contrastive learning framework named OpcGCD. Our OpcGCD framework implements a dynamic category-wise threshold mechanism, which employs a parametric prototype classifiers to generate debiased pseudo-labels for unlabeled samples. To facilitate the learning of discriminative feature representations, our proposed OpcGCD employs debiased pseudo-labels in the formulation of a contrastive learning loss. Extensive evaluations conducted on multiple GCD benchmark datasets demonstrate the robustness and effectiveness of the approach. Yiqing Hao, Xu Wang 0053, Yi Jin 0001, Tao Wang 0011, Yidong Li, Shuoyan Liu, Chao Li 0026, Hui Yu 0001 |
SMC | 3 |
| 2025 | The Cascaded Forward algorithm for neural network training
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling |
Pattern Recognit. | 3 |
| 2025 | V2PNet: A Voxel-to-Point Network Framework for Task-Oriented Point Cloud SamplingabstractTask-oriented point cloud sampling is a fundamental technique in 3D computer vision and has become a crucial step in numerous 3D applications. However, most state-of-the-art task-oriented sampling methods adopt a point-wise analysis strategy, making them susceptible to data redundancy. Taking inspiration from the abstract-to-detailed recognition process of the human visual system, we propose a novel voxel-to-point network framework called V2PNet for task-oriented point cloud sampling. Specifically, we first design a lightweight coarse-grained sampling module named Important Voxel Prediction (IMVP). This module adaptively outputs points from significant regions of the point cloud by explicitly modeling inter-region relationships, thereby reducing interference from redundant points. Then, the V2PNet framework seamlessly integrates the IMVP module with existing point-wise and task-oriented sampling networks, enabling joint training with downstream tasks. This creates a task-oriented coarse-to-fine-grained sampling pipeline that effectively samples representative and informative points from significant regions to represent the original point cloud. Moreover, to mitigate disturbances across similar regions, we introduce a voxel simplification loss function to enhance the discriminative voxel prediction. Extensive experiments demonstrate that V2PNet improves the performance of existing state-of-the-art task-oriented sampling models. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | M3-ReID: Unifying Multi-View, Granularity, and Modality for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based visible-infrared person re-identification (VVI-ReID) task focuses on cross-modality retrieval of pedestrian videos, which are captured in visible and infrared modalities by non-overlapping cameras across diverse scenes, and holds significant value for security surveillance scenarios. The challenges of this task mainly stem from three issues: the difficulty of capturing comprehensive spatio-temporal cues, intra-class variations within video sequences, and inter-modality discrepancies between visible and infrared data. Existing methods mainly try to address the modality gap or focus on one of the other aspects, but rarely do they jointly consider these key factors. Motivated by these core challenges, we propose the M3-ReID (Multi-View & Granularity & Modality) method, a unified framework that simultaneously enhances spatio-temporal feature extraction, intra-class discrimination, and cross-modality consistency. Specifically, to capture diverse spatio-temporal patterns, we design a Multi-View Learning module that leverages different spatial and temporal-spatial perspectives to adaptively emphasize diverse key regions and motion cues. To enhance intra-class modeling of each identity, we introduce a Multi-Granularity Representation strategy that optimizes features across both fine-grained frame level and coarse-grained video level by minimizing mutual information among redundant frames while enhancing identity representations. Furthermore, to bridge the visible-infrared gap, we propose a Multi-Modality Alignment mechanism that explicitly aligns metric learning and cross-modality matching goals, transforming features into a unified embedding space with modality consistency and class discrimination. Extensive experiments on benchmark VVI-ReID datasets demonstrate the superiority of our proposed M3-ReID framework against existing methods. Tengfei Liang, Yi Jin 0001, Zhun Zhong, Xin Chen 0003, Xianjia Meng, Tao Wang 0011, Yidong Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | 'Disengage AND Integrate': Personalized Causal Network for Gaze EstimationabstractGaze estimation task aims to predict a 3D gaze direction or a 2D gaze point given a face or eye image. To improve generalization of gaze estimation models to unseen new users, existing methods either disentangle personalized information of all subjects from their gaze features, or integrate unrefined personalized information into blended embeddings. Their methodologies are not rigorous whose performance is still unsatisfactory. In this paper, we put forward a comprehensive perspective named 'Disengage AND Integrate' to deal with personalized information, which elaborates that for specified users, their irrelevant personalized information should be discarded while relevant one should be considered. Accordingly, a novel Personalized Causal Network (PCNet) for generalizable gaze estimation has been proposed. The PCNet adopts a two-branch framework, which consists of a subject-deconfounded appearance sub-network (SdeANet) and a prototypical personalization sub-network (ProPNet). The SdeANet aims to explore causalities among facial images, gazes, and personalized information and extract a subject-invariant appearance-aware feature of each image by means of causal intervention. The ProPNet aims to characterize customized personalization-aware features of arbitrary users with the help of a prototype-based subject identification task. Furthermore, our whole PCNet is optimized in a hybrid episodic training paradigm, which further improve its adaptability to new users. Experiments on three challenging datasets over within-domain and cross-domain gaze estimation tasks demonstrate the effectiveness of our method. Xiyun Wang, Sihui Zhang, Wanru Xu, Yi Jin 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Multi-Modal Self-Perception Enhanced Large Language Model for 3D Region-of-Interest Captioning With Limited Dataabstract3D Region-of-Interest (RoI) Captioning involves translating a model's understanding of specific objects within a complex 3D scene into descriptive captions. Recent advancements in Large Language Models (LLMs) have shown great potential in this area. Existing methods capture the visual information from RoIs as input tokens for LLMs. However, this approach may not provide enough detailed information for LLMs to generate accurate region-specific captions. In this paper, we introduce Self-RoI, a Large Language Model with multi-modal self-perception capabilities for 3D RoI captioning. To ensure LLMs receive more precise and sufficient information, Self-RoI incorporates Implicit Textual Info. Perception to construct a multi-modal vision-language information. This module utilizes a simple mapping network to generate textual information about basic properties of RoI from vision-following response of LLMs. This textual information is then integrated with the RoI's visual representation to form a comprehensive multi-modal instruction for LLMs. Given the limited availability of 3D RoI-captioning data, we propose a two-stage training strategy to optimize Self-RoI efficiently. In the first stage, we align 3D RoI vision and caption representations. In the second stage, we focus on 3D RoI vision-caption interaction, using a disparate contrastive embedding module to improve the reliability of the implicit textual information and employing language modeling loss to ensure accurate caption generation. Our experiments demonstrate that Self-RoI significantly outperforms previous 3D RoI captioning models. Moreover, the Implicit Textual Info. Perception can be integrated into other multi-modal LLMs for performance enhancement. We will make our code available for further research. Lu Shi 0004, Shichao Kan, Yi Jin 0001, Linna Zhang, Yi-Gang Cen |
IEEE Trans. Multim. | 3 |
| 2025 | LighTN: Light-Weight Transformer Network for Performance-Overhead Tradeoff in Point Cloud DownsamplingabstractDownsampling is a crucial task for processing large scale and/or dense point clouds with limited resources. Owing to the development of deep learning, approaches of task-oriented point cloud downsampling have significant performance gains in preserving geometric information. However, most downsamling methods are limited by the disordered and unstructured point cloud data, making it difficult to continually improve the performance. To address this issue, we propose a light-weight Transformer network (LighTN) for the task-oriented point cloud downsampling as an end-to-end solution. In LighTN, we design an energy-efficient and permutation invariant single-head self-correlation module to extract refined global geometric features. Moreover, we present a novel sampling loss function to guide LighTN to focus on critical point cloud regions with more uniform distributions and prominent point coverage. Extensive experiments on classification, registration, and reconstruction tasks demonstrate that LighTN can achieve the state-of-the-art performance-overhead tradeoff and high-quality qualitative results. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Tao Wang 0011, Bowen Tang 0001, Yidong Li |
IEEE Trans. Multim. | 2 |
| 2024 | Enhancing Multimedia Applications by Removing Dynamic Objects in Neural Radiance Fields
XianBen Yang, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li |
ACCV (10) | 4 |
| 2024 | DFA-GNN: Forward Learning of Graph Neural Networks by Direct Feedback AlignmentabstractGraph neural networks (GNNs) are recognized for their strong performance across various applications, with the backpropagation (BP) algorithm playing a central role in the development of most GNN models. However, despite its effectiveness, BP has limitations that challenge its biological plausibility and affect the efficiency, scalability and parallelism of training neural networks for graph-based tasks. While several non-backpropagation (non-BP) training algorithms, such as the direct feedback alignment (DFA), have been successfully applied to fully-connected and convolutional network components for handling Euclidean data, directly adapting these non-BP frameworks to manage non-Euclidean graph data in GNN models presents significant challenges. These challenges primarily arise from the violation of the independent and identically distributed (i.i.d.) assumption in graph data and the difficulty in accessing prediction errors for all samples (nodes) within the graph. To overcome these obstacles, in this paper we propose DFA-GNN, a novel forward learning framework tailored for GNNs with a case study of semi-supervised learning. The proposed method breaks the limitations of BP by using a dedicated forward training mechanism. Specifically, DFA-GNN extends the principles of DFA to adapt to graph data and unique architecture of GNNs, which incorporates the information of graph topology into the feedback links to accommodate the non-Euclidean characteristics of graph data. Additionally, for semi-supervised graph learning tasks, we developed a pseudo error generator that spreads residual errors from training data to create a pseudo error for each unlabeled node. These pseudo errors are then utilized to train GNNs using DFA. Extensive experiments on 10 public benchmarks reveal that our learning framework outperforms not only previous non-BP methods but also the standard BP methods, and it exhibits excellent robustness against various types of noise and attacks. Gongpei Zhao, Tao Wang 0011, Congyan Lang, Yi Jin 0001, Yidong Li, Haibin Ling |
NeurIPS | 4 |
| 2024 | Enhancing Point Cloud Sampling Quality with Dual-Branch Fusion NetworksabstractTask-oriented point cloud sampling methods have attracted considerable attention for their ability to adaptively select important point sets based on downstream tasks, achieving an excellent balance between data simplification and task performance. However, existing task-oriented sampling models, primarily based on single-branch designs, struggle to fully extract features from input point clouds that comprehensively reflect multi-dimensional key information, thus limiting their sampling performance. In this paper, we introduce a dual-branch sampling network, named DBS-NET, which conducts crucial point sampling from both the global and local importance perspectives separately before merging them, thereby preserving multi-dimensional key information of the input data during the sampling process. Qualitative and quantitative experimental results demonstrate the competitive performance of DBS-NET on the classification benchmark task. Yi Jin 0001, Xu Wang 0053, Mengxia Hu, Hui Yu 0001, Yidong Li, Tao Wang 0011, Songhe Feng, Congyan Lang |
SMC | 1 |
| 2024 | GLAN: A graph-based linear assignment network
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li |
Pattern Recognit. | 5 |
| 2024 | Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person Re-IDentification (VI-ReID) is a challenging cross-modality image retrieval task that aims to match pedestrians’ images across visible and infrared cameras. To solve the modality gap, existing mainstream methods adopt a learning paradigm converting the image retrieval task into an image classification task with cross-entropy loss and auxiliary metric learning losses. These losses follow the strategy of adjusting the distribution of extracted embeddings to reduce the intra-class distance and increase the inter-class distance. However, such objectives do not precisely correspond to the final test setting of the retrieval task, resulting in a new gap at the optimization level. By rethinking these keys of VI-ReID, we propose a simple and effective method, the Multi-level Cross-modality Joint Alignment (MCJA), bridging both the modality and objective-level gap. For the former, we design the Visible-Infrared Modality Coordinator in the image space and propose the Modality Distribution Adapter in the feature space, effectively reducing modality discrepancy of the feature extraction process. For the latter, we introduce a new Cross-Modality Retrieval loss. It is the first work to constrain from the perspective of the ranking list in the VI-ReID, aligning with the goal of the testing stage. Moreover, to strengthen the robustness and cross-modality retrieval ability, we further introduce a Multi-Spectral Enhanced Ranking strategy for the testing phase. Based on the global feature only, our method outperforms existing methods by a large margin, achieving the remarkable rank-1 of 89.51% and mAP of 87.58% on the most challenging single-shot setting and all-search mode of the SYSU-MM01 dataset. Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Tao Wang 0011, Songhe Feng, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A Generative-Based Image Fusion Strategy for Visible-Infrared Person Re-IdentificationabstractCross-modality person re-identification task is a challenging task aiming to recognize images of the same identity between different modalities. To alleviate the cross-modality discrepancies between images, existing approaches mainly guide models to mine modality invariant features. Although those approaches are effective, they lose the modality-specific features that include important information beneficial to VI-ReID. Therefore, some approaches are using generative adversarial networks to compensate for modality information. However, the quality of images generated by these methods is usually poor, and most of them focus only on the learning of modality-sharable features. To solve these problems, this paper proposes a generative-based cross-modality image fusion strategy (GC-IFS), which can generate high-quality cross-modality paired images and fuse the information of the two modalities. Firstly, considering the importance of the identity discriminative information of the generated image, we propose a contrastive-learning image generation (CLIG) network to generate cross-modality paired images. Meanwhile, to fully integrate and utilize the information of the two modalities and eliminate the influence of cross-modality discrepancies, we design a part-based dual multi-modality feature fusion (P-DMFF) module to extract the unified feature representation. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that our strategy outperforms the state-of-the-art methods for the VI-ReID task. Jia Qi, Tengfei Liang, Wu Liu 0005, Yidong Li, Yi Jin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Transformer Driven Matching Selection Mechanism for Multi-Label Image ClassificationabstractGraph Matching has recently emerged as an attractive technique applied to various computer vision tasks. Graph Matching based multi-label image classification, in particular, treats each image as a bag of instances and reformulates the classification task as an instance-label matching selection problem, achieving state-of-the-art results on diverse benchmarks. However, the generalization and scalability of such learned model cannot be well guaranteed due to its manually predetermined graph structure and high-dimension embedding of dense connections between instances and labels. To address these limitations, in this work, we propose a novel${T}$ransformer Driven${M}$atching${S}$election framework for Multi-Label Image${C}$lassification (C-TMS), where instance structural relationships, class-wise global dependencies, and the co-occurrence possibility of varying instance-label assignments are simultaneously taken into consideration in a unified and adaptive manner. Moreover, the parallelization capability of the Transformer enables efficient computation, making our model scalable to large-scale datasets. Specifically, we first represent instances and labels as nodes in the visual space and label space respectively, and then compute the hidden representation of each node in its individual space, by attending a self-attention strategy over its entire neighborhood. Subsequently, the cross-attention is adopted to excavate the correct assignments between instances and labels, and further interprets how classifying each label depends on the instances within an image and its interaction with other labels. Finally, an asymmetric focal loss is designed to optimize the instance-label correspondence, and read out image-level category confidences. Extensive experiments conducted on various multi-label image datasets demonstrate the superiority of our proposed method. Songhe Feng, Gongpei Zhao, Yi Jin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MMI-Det: Exploring Multi-Modal Integration for Visible and Infrared Object DetectionabstractThe Visible-Infrared (VIS-IR) object detection is a challenging detection task, which combines visible and infrared data to give information on the category and location of objects in the scene. Therefore, the core of this task is to combine complementary information in the visible and infrared modalities to provide more object detection results for detection. The existing methods mainly face the problem of insufficient ability to perceive and combine visible-infrared modal information and have difficulty in balancing the optimization directions of the fusion and detection tasks. To solve these problem, we propose the MMI-Det which is a multi-modal fusion method for visible and infrared object detection. The method can provide a good combination of complementary information in the visible-infrared modalities and output accurate and robust object information. Specifically, to improve the ability of the model to perceive environment at the visible-infrared image level, we designed the Contour Enhancement Module. Furthermore, to extract complementary information from VIS and IR modalities, we design the Fusion Focus Module. It can extract different frequency spectral features of the visible and infrared modalities and focus on the key information of the object at different spatial locations. Moreover, we design the Contrast Bridge Module to improve the ability to extract modal invariant features in the visible-infrared scene. Finally, to ensure that our model can balance the optimization directions of image fusion and object detection, we design the Info Guided Module as a way to improve the effectiveness of the model’s training optimization. We implement extensive experiments on the public FLIR, M3FD, LLVIP, TNO and MSRS datasets, and compared with previous methods, our method achieves better performance with powerful multi-modal information perception capabilities. Yuqiao Zeng, Tengfei Liang, Yi Jin 0001, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Neighborhood Pattern Is Crucial for Graph Convolutional Networks Performing Node ClassificationabstractGraph convolutional networks (GCNs) are widely believed to perform well in the graph node classification task, and homophily assumption plays a core rule in the design of previous GCNs. However, some recent advances on this area have pointed out that homophily may not be a necessity for GCNs. For deeper analysis of the critical factor affecting the performance of GCNs, we first propose a metric, namely, neighborhood class consistency (NCC), to quantitatively characterize the neighborhood patterns of graph datasets. Experiments surprisingly illustrate that our NCC is a better indicator, in comparison to the widely used homophily metrics, to estimate GCN performance for node classification. Furthermore, we propose a topology augmentation graph convolutional network (TA-GCN) framework under the guidance of the NCC metric, which simultaneously learns an augmented graph topology with higher NCC score and a node classifier based on the augmented graph topology. Extensive experiments on six public benchmarks clearly show that the proposed TA-GCN derives ideal topology with higher NCC score given the original graph topology and raw features, and it achieves excellent performance for semi-supervised node classification in comparison to several state-of-the-art (SOTA) baseline algorithms. Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Editorial to Special Issue on Multimedia Cognitive Computing for Intelligent Transportation SystemabstractNo abstract available. Shaohua Wan 0001, Yi Jin 0001, Guandong Xu, Michele Nappi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | MetaZSCIL: A Meta-Learning Approach for Generalized Zero-Shot Class Incremental LearningabstractGeneralized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Standard GZSL cannot handle dynamic addition of new seen and unseen classes. In order to address this limitation, some recent attempts have been made to develop continual GZSL methods. However, these methods require end-users to continuously collect and annotate numerous seen class samples, which is unrealistic and hampers the applicability in the real-world. Accordingly, in this paper, we propose a more practical and challenging setting named Generalized Zero-Shot Class Incremental Learning (CI-GZSL). Our setting aims to incrementally learn unseen classes without any training samples, while recognizing all classes previously encountered. We further propose a bi-level meta-learning based method called MetaZSCIL to directly optimize the network to learn how to incrementally learn. Specifically, we sample sequential tasks from seen classes during the offline training to simulate the incremental learning process. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Note that our optimization can be flexibly equipped with most existing generative methods to tackle CI-GZSL. This work introduces a feature generative framework that leverages visual feature distribution alignment to produce replayed samples of previously seen classes to reduce catastrophic forgetting. Extensive experiments conducted on five widely used benchmarks demonstrate the superiority of our proposed method. Tengfei Liang, Songhe Feng, Yi Jin 0001, Gengyu Lyu, Haojun Fei, Yang Wang 0003 |
AAAI | 4 |
| 2023 | Localized Knowledge Distillation Helps IoT Devices Provide High-performance Visual ServicesabstractDeploying high-performance convolutional neural networks (CNNs) on ubiquitous Internet of Things (IoT) devices to provide convenient services has attracted increasing attention. However, most existing studies have two limitations, (i) low performance due to lack of effective learning strategies; (ii) difficulty to measure resource consumption due to lack of real deployment. To this end, this paper proposes a novel localized knowledge distillation (LKD) method to train the resource-efficient CNN and implements a cloud-assisted system to evaluate the on-device performance. The proposed LKD follows the layer-wise heterogeneous information distribution in the CNN and distills the knowledge from features that contains the most crucial knowledge to the resource-efficient CNN. Thus, the knowledge guided to learn by limited parameters is reduced to the crucial part, which is more reliable for the resource-efficient CNN. The cloud-assisted system consists cloud server and IoT devices, which allows the training and deployment of the resource-efficient CNN, and the measurement of the corresponding resource consumption. Experimental results show that the proposed LKD could improve the performance on standard benchmarks close to the counterpart CNN, and the robustness on corrupted data by approximately 11.3% for the resource-efficient CNN. The measurements on the cloud-assisted system also demonstrate the resource efficiency for transmission and on-device running. Chuntao Ding, Yi Jin 0001, Yidong Li |
ICWS | 4 |
| 2023 | Context-guided coarse-to-fine detection model for bird nest detection on high-speed railway catenaryabstractAbstract As a critical component of ensuring the safe and stable operation of trains, the detection of bird’s nests on the rail catenary has always been essential. Low-resolution images and the lack of labelled data, however, make it difficult to detect smaller bird’s nests (those occupying small pixels in the input image). Previous solution relies on manual online patrol or offline video playback, which severely limits the detection efficiency. Previously, this challenge was addressed by manual online patrol or offline video playback, which severely limits detection efficiency. We propose in this work a context-guided coarse-to-fine detection model (CG-CFDM) for solving the bird’s nest detection problem. This solution consists of a context reasoning module and a coarse-to-fine detection network. By detecting domains and matching templates, the context reasoning module generates new labelled context bounding boxes, thereby reducing the burden of annotation. As a result of its delicately designed architecture and powerful representation learning ability, this trained coarse-to-fine detection network further facilitates the detection of bird’s nests in an efficient and accurate manner. Extensive experiments demonstrate that the proposed approach is superior to existing methods in terms of performance and has a great deal of potential for detecting bird’s nests. Siquan Wu, Yidong Li, Yi Jin 0001 |
Multim. Syst. | 5 |
| 2023 | Joint Graph Learning and Matching for Semantic Feature Correspondence
Tao Wang 0011, Yidong Li, Congyan Lang, Yi Jin 0001, Haibin Ling |
Pattern Recognit. | 5 |
| 2023 | RS-TNet: point cloud transformer with relation-shape awareness for fine-grained 3D visual processing
Xu Wang 0053, Yuqiao Zeng, Yi Jin 0001, Yi-Gang Cen, Baifu Liu, Shaohua Wan 0001 |
Soft Comput. | 3 |
| 2023 | Prior Knowledge Regularized Self-Representation Model for Partial Multilabel LearningabstractPartial multilabel learning (PML) aims to learn from training data, where each instance is associated with a set of candidate labels, among which only a part is correct. The common strategy to deal with such a problem is disambiguation, that is, identifying the ground-truth labels from the given candidate labels. However, the existing PML approaches always focus on leveraging the instance relationship to disambiguate the given noisy label space, while the potentially useful information in label space is not effectively explored. Meanwhile, the existence of noise and outliers in training data also makes the disambiguation operation less reliable, which inevitably decreases the robustness of the learned model. In this article, we propose a prior label knowledge regularized self-representation PML approach, called PAKS, where the self-representation scheme and prior label knowledge are jointly incorporated into a unified framework. Specifically, we introduce a self-representation model with a low-rank constraint, which aims to learn the subspace representations of distinct instances and explore the high-order underlying correlation among different instances. Meanwhile, we incorporate prior label knowledge into the above self-representation model, where the prior label knowledge is regarded as the complement of features to obtain an accurate self-representation matrix. The core of PAKS is to take advantage of the data membership preference, which is derived from the prior label knowledge, to purify the discovered membership of the data and accordingly obtain more representative feature subspace for model induction. Enormous experiments on both synthetic and real-world datasets show that our proposed approach can achieve superior or comparable performance to state-of-the-art approaches. Gengyu Lyu, Songhe Feng, Yi Jin 0001, Tao Wang 0011, Congyan Lang, Yidong Li |
IEEE Trans. Cybern. | 3 |
| 2023 | Distance-Preserving Embedding Adaptive Bipartite Graph Multi-View Learning with Application to Multi-Label ClassificationabstractGraph-based multi-view learning has attracted much attention due to the efficacy of fusing the information from different views. However, most of them exhibit high computational complexity. We propose an anchor-based bipartite graph embedding approach to accelerate the learning process. Specifically, different from existing anchor-based methods where anchors are obtained from key samples by clustering or weighted averaging strategies, in this article, the anchors are learned in a principled fashion which aims at constructing a distance-preserving embedding for each view from samples to their representations, whose elements are the weights of the edges linking corresponding samples and anchors. In addition, the consistency among different views can be explored by imposing a low-rank constraint on the concatenated embedding representations. We further design a concise yet effective feature collinearity guided feature selection scheme to learn tight multi-label classifiers. The objective function is optimized in an alternating optimization fashion. Both theoretical analysis and experimental results on different multi-label image datasets verify the effectiveness and efficiency of the proposed method. Songhe Feng, Gengyu Lyu, Yi Jin 0001, Congyan Lang |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Weakly Supervised Object Detection With Class Prototypical NetworkabstractIn this paper, we aim to devise a new framework to compel the network to be equipped with the capability of detecting objects using image-level class labels as supervision. The challenge of such a weakly supervised setting mainly lies in how to make the network accurately understand both semantics and objectness of a given proposal without bounding box annotations. To this end, we contribute a concise framework, named Class Prototypical Network (CPNet). Concretely, our CPNet defines a set of learnable class prototypes to help classify object proposals. To endow the prototypes be not only discriminative for classes but also sensitive for proposals' objectness, we conduct both class-aware cross-attention and location-aware cross-attention between the feature embeddings of the learnable prototypes and the proposals. The learned attention scores are then used to form the proposal-level category information into the image-level one, making the entire framework be trained without any bounding box annotations. Besides, by applying these two kinds of attention mechanisms, the knowledge from both proposals' location and its class information can be successfully transferred into the corresponding prototypes. With the help of prototypes, our CPNet detects true positive object proposals. In addition, the CPNet further introduces a multi-head detection head to perform complementary training, preventing the model from falling into local discriminative parts and improving the model's performance on challenging non-rigid categories. We examine our CPNet on popular benchmarks,i.e., PASCAL VOC 2007, 2012 and MS COCO 2014. Extensive experiments show our CPNet is a simple and effective framework. Yidong Li, Yuanzhouhan Cao, Yushan Han, Yi Jin 0001, Yunchao Wei |
IEEE Trans. Multim. | 5 |
| 2023 | Cross-Modality Transformer With Modality Mining for Visible-Infrared Person Re-IdentificationabstractThe visible-infrared person re-identification (VI-ReID) is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to bridge the huge gap between these two modalities. The existing methods mainly face the problem of insufficient perception of modality information, and can not learn good discriminative modality-invariant embeddings for identities, which limits their performance. To solve these problems, we propose a new cross-modality transformer-based method (CMTR) for this visible-infrared person re-identification task, which can explicitly mine the information of each modality and generate better discriminative features based on it. Specifically, to capture inherent characteristics of modalities, we design the novel modality embeddings, which are fused with token embeddings to encode modality information directly. Moreover, to enhance representation of modality embeddings and adjust the distribution of embeddings, we further propose a modality-aware enhancement loss based on the learned modality information, which contains two components to reduce intra-class distance and enlarging inter-class distance simultaneously. To our knowledge, this is the first exploration of applying pure transformer network to the cross-modality re-identification task. We implement extensive experiments on the public SYSU-MM01 and RegDB datasets, and compared with previous methods, our method achieves good performance with more compact and powerful embeddings for the cross-modality retrieval. Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Yidong Li |
IEEE Trans. Multim. | 2 |
| 2023 | Local Correlation Ensemble with GCN Based on Attention Features for Cross-domain Person Re-IDabstractPerson re-identification (Re-ID) has achieved great success in single-domain. However, it remains a challenging task to adapt a Re-ID model trained on one dataset to another one. Unsupervised domain adaption (UDA) was proposed to migrate a model from a labeled source domain to an unlabeled target domain. The main difference in the cross-domain is different background styles. Although the style transfer approach effectively reduces inter-domain gaps, it ignores the reduction of intra-class differences. Clustering-based pipelines maintain state-of-the-art performance for UDA by learning domain-independent features; however, most existing models do not sufficiently exploit the rich unlabeled samples in target domains due to unsatisfactory clustering. Thus, we propose a novel local correlation ensemble model that focuses on the diversity of intra-class information and the reliability of class centers. Specifically, a pedestrian attention module is proposed to enable the encoder to pay more attention to the person’s features to relieve interference caused by the shared background style. Furthermore, we propose a priority-distance graph convolutional network (PDGCN) module that employs a graph convolutional network network to predict the priority of a node as a class center and then calculates the distance between nodes with high priority values to screen out the class center nodes. Finally, the encoder features (local) and PDGCN features (context-aware) are combined to perform person Re-ID. The results of experiments on the large-scale public Re-ID datasets verified the effectiveness of the proposed method. Yue Zhang 0065, Fanghui Zhang, Yi Jin 0001, Yi-Gang Cen, Viacheslav V. Voronin, Shaohua Wan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Boundary Corrected Multi-Scale Fusion Network for Real-Time Semantic SegmentationabstractImage semantic segmentation aims at the pixel-level classification of images, which has requirements for both accuracy and speed in practical application. Existing semantic segmentation methods mainly rely on the high-resolution input to achieve high accuracy and do not meet the requirements of inference time. Although some methods focus on high-speed scene parsing with lightweight architectures, they can not fully mine semantic features under low computation with relatively low performance. To realize the real-time and high-precision segmentation, we propose a new method named Boundary Corrected Multi-scale Fusion Network, which uses the designed Low-resolution Multi-scale Fusion Module to extract semantic information. Moreover, to deal with boundary errors caused by low-resolution feature map fusion, we further design an additional Boundary Corrected Loss to constrain overly smooth features. Extensive experiments show that our method achieves a state-of-the-art balance of accuracy and speed for the real-time semantic segmentation. Tianjiao Jiang, Yi Jin 0001, Tengfei Liang, Xu Wang 0053, Yidong Li |
ICIP | 2 |
| 2022 | Camera-Aware Style Separation and Contrastive Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID) is a challenging task without data annotation to guide discriminative learning. Existing methods attempt to solve this problem by clustering extracted embeddings to generate pseudo labels. However, most methods ignore the intra-class gap caused by camera style variance, and some methods are relatively complex and indirect although they try to solve the negative impact of the camera style on feature distribution. To solve this problem, we propose a camera-aware style separation and contrastive learning method (CA-UReID), which directly separates camera styles in the feature space with the designed camera-aware attention module. It can explicitly divide the learnable feature into camera-specific and camera-agnostic parts, reducing the influence of different cameras. Moreover, to further narrow the gap across cameras, we design a camera-aware contrastive center loss to learn more discriminative embeddings for each identity. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods on the unsupervised person ReID task. Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li |
ICME | 3 |
| 2022 | Keypoint-Guided Modality-Invariant Discriminative Learning for Visible-Infrared Person Re-identificationabstractThe visible-infrared person re-identification (VI-ReID) task aims to retrieve images of pedestrians across cameras with different modalities. In this task, the major challenges arise from two aspects: intra-class variations among images of the same identity, and cross-modality discrepancies between visible and infrared images. Existing methods mainly focus on the latter, attempting to alleviate the impact of modality discrepancy, which ignore the former issue of identity variations and achieve limited discrimination. To address both aspects, we propose a Keypoint-guided Modality-invariant Discriminative Learning (KMDL) method, which can simultaneously adapt to intra-ID variations and bridge the cross-modality gap. By introducing human keypoints, our method makes further exploration in the image space, feature space and loss constraints to solve the above issues. Specifically, considering the modality discrepancy in original images, we first design a Hue Jitter Augmentation (HJA) strategy, introducing the hue disturbance to alleviate color dependence in the input stage. To obtain discriminative fine-grained representation for retrieval, we design the Global-Keypoint Graph Module (GKGM) in feature space, which can directly extract keypoint-aligned features and mine relationships within global and keypoint embeddings. Based on these semantic local embeddings, we further propose the Keypoint-Aware Center (KAC) loss that can effectively adjust the feature distribution under the supervision of ID and keypoint to learn discriminative representation for the matching. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our KMDL method. Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Songhe Feng, Tao Wang 0011, Yidong Li |
ACM Multimedia | 2 |
| 2022 | Task Adaptive Modeling for Few-shot Action RecognitionabstractCollecting action recognition datasets is time-consuming and labor-intensive. To solve this problem, a few-shot action recognition task that uses episode training to learn the model appears. However, due to the randomness of few-shot learning task sampling, there are great differences between each task, and the characteristics of classes are also diverse. Most of the current methods simply use the same processing flow for action recognition, ignoring the correlation between tasks. To solve this problem, we propose a task adaptive network for few-shot action recognition, which utilizes the dependency of support set and query set categories. Our method mainly includes two key points: Firstly, we add an attention module after the feature extraction module, which can use the attention mechanism to focus the obtained feature representation on more important local information. Secondly, we design a task adaptive module, which uses the support set samples to strengthen all samples of the current task. The module strengthens the common features within each class of the support set and expands the query set to highlight the differences between classes. We have conducted a large number of experiments on two commonly used action recognition data sets: HMDB51 and UCF101. The results of experiments show that our method has strong competitiveness and performs well in the field of few-shot action recognition. Yi Jin 0001, Songhe Feng, Yidong Li |
MMSP | 2 |
| 2022 | Object representation enhancement for self-supervised colocalizationabstractSelf-supervised colocalization is to localize common objects in the data set containing only one superclass without using human-annotated labels. Existing methods achieve impressive results by employing self-supervised pretext learning. However, a common limitation still exists. They either tend to overextend activations to the background, or they tend to activate the most discriminative object part. To alleviate this problem, we propose an object representation enhancement model to weaken background distraction and to mine complementary object regions during the object representation learning. Specifically, we first propose an Object-aware Representation Enhancement (ORE) module to estimate an object mask for each input image, guiding the model to disregard the background content and focus on the foreground object. The ORE module and the subsequent self-supervised learning can mutually reinforce each other. Then we propose a Masked Self-supervised Learning branch and design a masked attention consistency objective to induce the model to activate complementary parts of the object effectively. Extensive experiments on four fine-grained data sets demonstrate the superiority of the proposed model. Yidong Li, Yi Jin 0001, Tao Wang 0011 |
Int. J. Intell. Syst. | 3 |
| 2022 | VSLN: View-aware sphere learning network for cross-view vehicle re-identificationabstractCross-view vehicle Reidentification (ReID) has attracted widespread attention as an increasingly important vision task in intelligent transportation and urban surveillance. Benefiting from Convolutional Neural Network (CNN), recent studies have promoted the development of vehicle ReID by extracting discriminative local features. However, two fundamental challenges of small interclass discrepancy caused by different views and large intraclass distance caused by similar appearance still hinder the performance of cross-view vehicle ReID. In this paper, a novel View-aware Sphere Learning Network (VSLN) is proposed to alleviate the above issues while maintaining the merits of CNN-based approaches to generate view-aware sphere-based features. First, a Sphere Feature Embedding Network (SFEN) is proposed to constrain the images into hypersphere for extracting sphere features. On the other hand, this study presents a sphere similarity triple loss to help SFEN concentrate more on robust and discriminative vehicle parts. Second, since the vehicle images are usually captured from different viewpoints, this study further extends SFEN by introducing a Vehicle Viewpoint Predictor (VVP) combined with global attention mechanism to enlarge the discrepancy of interclass and shorten the distance of intraclass. Moreover, a city-scale data set, named Vehicle from Different Viewpoints, containing image-level viewpoint labels, is collected for training VVP. As a result, the proposed VLSN can achieve 96.31% Top-1 accuracy and 79.46% Top-1 accuracy on VeRi-776 and VRIC data sets, respectively. Overall, extensive experimental results on two benchmark data sets show that the proposed VSLN outperforms state-of-the-art methods. Xu Wang 0053, Yi Jin 0001, Chenning Li, Yi-Gang Cen, Yidong Li |
Int. J. Intell. Syst. | 2 |
| 2022 | VARID: Viewpoint-Aware Re-IDentification of Vehicle Based on Triplet LossabstractWith the increasing prevalence of intelligent traffic control and monitoring, research on vehicle re-identification (Re-ID) draws substantial attention in recent years. Different from other cross-view searching tasks such as person Re-ID, the vehicle Re-ID problem is more challenging and unpredictable as viewpoint variations can greatly affect the appearance of vehicles. Existing studies mainly focus on extracting global features based on visual appearance to represent the identity of the target vehicle, while the impact of viewpoint variation is rarely considered. In this paper, we take the view information into account to boost vehicle Re-ID, and introduce latent view labels by clustering and incorporates view information into deep metric learning to tackle the challenge. We also develop a stricter center constraint to further improve the intra-class compactness of feature space. Moreover, we adopt an orthogonal regularization to increase the separability between different vehicles. VARID achieves 79.3% mAP on VeRi-776 and 88.5% mAP on VehicleID which surpasses state-of-the-arts a lot. More comprehensive experimental analyses and evaluations on four benchmarks demonstrate that the proposed method outperforms significantly state-of-the-arts methods. Yidong Li, Yi Jin 0001, Tao Wang 0011, Weipeng Lin |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Unsupervised Domain Adaptation for Person Re-identification via Heterogeneous Graph AlignmentabstractUnsupervised person re-identification (re-ID) is becoming increasingly popular due to its power in real-world systems such as public security and intelligent transportation systems. However, the person re-ID task is challenged by the problems of data distribution discrepancy across cameras and lack of label information. In this paper, we propose a coarse-to-fine heterogeneous graph alignment (HGA) method to find cross-camera person matches by characterizing the unlabeled data as a heterogeneous graph for each camera. In the coarse-alignment stage, we assign a projection for each camera and utilize an adversarial learning based method to align coarse-grained node groups from different cameras into a shared space, which consequently alleviates the distribution discrepancy between cameras. In the fine-alignment stage, we exploit potential fine-grained node groups in the shared space and introduce conservative alignment loss functions to constrain the graph aligning process, resulting in reliable pseudo labels as learning guidance. The proposed domain adaptation framework not only improves model generalization on target domain, but also facilitates mining and integrating the potential discriminative information across different cameras. Extensive experiments on benchmark datasets demonstrate that the proposed approach outperforms the state-of-the-arts. Minying Zhang, Yidong Li, Shihui Guo, Hongtao Duan 0003, Yimin Long, Yi Jin 0001 |
AAAI | 7 |
| 2021 | PST-NET: Point Cloud Sampling via Point-Based Transformer
Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Congyan Lang, Yidong Li |
ICIG (3) | 2 |
| 2021 | GM-MLIC: Graph Matching based Multi-Label Image ClassificationabstractMulti-Label Image Classification (MLIC) aims to predict a set of labels that present in an image. The key to deal with such problem is to mine the associations between image contents and labels, and further obtain the correct assignments between images and their labels. In this paper, we treat each image as a bag of instances, and reformulate the task of MLIC as a instance-label matching selection problem. To model such problem, we propose a novel deep learning framework named Graph Matching based Multi-Label Image Classification (GM-MLIC), where Graph Matching (GM) scheme is introduced owing to its excellent capability of excavating the instance and label relationship. Specifically, we first construct an instance spatial graph and a label semantic graph respectively, and then incorporate them into a constructed assignment graph by connecting each instance to all labels. Subsequently, the graph network block is adopted to aggregate and update all nodes and edges state on the assignment graph to form structured representations for each instance and label. Our network finally derives a prediction score for each instance-label correspondence and optimizes such correspondence with a weighted cross-entropy loss. Extensive experiments conducted on various datasets demonstrate the superiority of our proposed method. Songhe Feng, Yi Jin 0001, Gengyu Lyu, Zizhang Wu |
IJCAI | 4 |
| 2021 | MSO: Multi-Feature Space Joint Optimization Network for RGB-Infrared Person Re-IdentificationabstractThe RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to eliminate the discrepancy between the two modalities in the final common feature space, which ignore the single space of each modality in the shallow layers. To solve it, in this paper, we present a novel multi-feature space joint optimization (MSO) network, which can learn modality-sharable features in both the single-modality space and the common space. Firstly, based on the observation that edge information is modality-invariant, we propose an edge features enhancement module to enhance the modality-sharable features in each single-modality space. Specifically, we design a perceptual edge features (PEF) loss after the edge fusion strategy analysis. According to our knowledge, this is the first work that proposes explicit optimization in the single-modality feature space on cross-modality ReID task. Moreover, to increase the difference between cross-modality distance and class distance, we introduce a novel cross-modality contrastive-center (CMCC) loss into the modality-joint constraints in the common feature space. The PEF loss and CMCC loss jointly optimize the model in an end-to-end manner, which markedly improves the network's performance. Extensive experiments demonstrate that the proposed model significantly outperforms state-of-the-art methods on both the SYSU-MM01 and RegDB datasets. Yajun Gao, Tengfei Liang, Yi Jin 0001, Xiaoyan Gu 0001, Wu Liu 0005, Yidong Li, Congyan Lang |
ACM Multimedia | 3 |
| 2021 | Video saliency prediction via spatio-temporal reasoning
Jiazhong Chen, Zongyi Li, Yi Jin 0001, Dakai Ren |
Neurocomputing | 3 |
| 2021 | Entropy-aware self-training for graph convolutional networks
Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang |
Neurocomputing | 4 |
| 2021 | Action unit analysis enhanced facial expression recognition by deep neural network evolution
Ruicong Zhi, Caixia Zhou, Shuai Liu 0003, Yi Jin 0001 |
Neurocomputing | 5 |
| 2021 | Text to photo-realistic image synthesis via chained deep recurrent generative adversarial network
Congyan Lang, Songhe Feng, Tao Wang 0011, Yi Jin 0001, Yidong Li |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Error-robust low-rank tensor approximation for multi-view clustering
Shuqin Wang 0001, Yongyong Chen, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Linna Zhang |
Knowl. Based Syst. | 3 |
| 2021 | Pedestrian detection with super-resolution reconstruction for low-quality image
Yi Jin 0001, Yue Zhang 0065, Yi-Gang Cen, Yidong Li, Vladimir Mladenovic, Viacheslav V. Voronin |
Pattern Recognit. | 1 |
| 2021 | Model Latent Views With Multi-Center Metric Learning for Vehicle Re-IdentificationabstractMulti-view vehicle re-identification (Re-ID) aims to retrieve all images of a target vehicle from a large gallery where the vehicles are captured from non-overlapping cameras. However, the drastic variation in vehicle appearance under different viewpoints greatly affects the performance of the multi-view vehicle Re-ID model, so the key issue in multi-view vehicle Re-ID is learning an effective feature representation that is robust to both dramatic intra-class variability and small inter-class variability. To achieve this goal, we have proposed a multi-center metric learning framework for multi-view vehicle Re-ID. In our approach, we model latent views from vehicle visual appearance directly without any extra labels except ID. Firstly, we introduce several latent view clusters for a vehicle to model latent multi-view information and each view cluster has a learnable center. Then multi-view vehicle matching task can be transformed into two subproblems, cross-view matching and cross-target matching. Finally, an intra-class ranking loss with cross-view center constraint and a cross-class ranking loss with cross-vehicle center constraint are proposed to address the two subproblems, respectively. Extensive experimental evaluations on three widely used benchmarks show the superiority of the proposed framework in contrast to a series of existing state-of-the-arts. Yi Jin 0001, Chenning Li, Yidong Li, Peixi Peng, George A. Giannopoulos |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Account Guarantee Scheme: Making Anonymous Accounts Supervised in BlockchainabstractIn blockchain networks, reaching effective supervision while maintaining anonymity to the public has been an ongoing challenge. In existing solutions, certification authorities need to record all pairs of identities and pseudonyms, which is demanding and costly. This article proposed an account guarantee scheme to realize feasible supervision for existing anonymous blockchain networks with lower storage costs. Users are able to guarantee anonymous accounts with account guarantee key pairs generated from certificated polynomial functions, which inherently maintains one-to-n mapping certifications. Single or limited account guarantee key pairs do not leak privacy. Victims are able to request TCs to screen a cheater or disclose a cheater with enough fraud transactions by themselves. Detailed security and privacy analysis showed that the account guarantee scheme preserves user privacy and realizes feasible supervision. Experimental results demonstrated that the account guarantee scheme is efficient and practical. Lichen Cheng, Jiqiang Liu, Yi Jin 0001, Yidong Li, Wei Wang 0012 |
ACM Trans. Internet Techn. | 3 |
| 2021 | SPGAN: Face Forgery Using Spoofing Generative Adversarial NetworksabstractCurrent face spoof detection schemes mainly rely on physiological cues such as eye blinking, mouth movements, and micro-expression changes, or textural attributes of the face images [9]. But none of these methods represent a viable mechanism for makeup-induced spoofing, especially since makeup has been widely used. Compared with face alteration techniques such as plastic surgery, makeup is non-permanent and cost efficient, which makes makeup-induced spoofing become a realistic threat to the integrity of a face recognition system. To solve this problem, we propose a generative model to construct spoofing face images (confusing face images) for improving the accuracy and robustness of automatic face recognition. Our network structure is composed of two separate parts, with one using inter-attention mechanism to obtain interested face region, and another using intra-attention to translate imitation style with preserving imitation style-excluding details. These two attention mechanisms can precisely learn imitation style, where inter-attention pays more attention to imitation regions of image and intra-attention learns face attributes with long distance in image. To effectively discriminate generated images, we introduce an imitation style discriminator. Our model (SPGAN) generates face images that transfer the imitation style from target to subject image and preserve the imitation-excluding features. Experimental results demonstrate the performance of our model in improving quality of imitated face images. Yidong Li, Wenhua Liu, Yi Jin 0001, Yuanzhouhan Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Trust-based federated learning for network anomaly detectionabstractWith the rapid development of social networks and the massive popularity of intelligent mobile terminals, network anomaly detection is becoming increasingly important. In daily work and life, edge nodes store a large number of network local connection data and audit data, which can be used to analyze network abnormal behavior. With the increasingly close network communication, the amount of network connection and other related data collected by each network terminal is increasing. Machine learning has become a classification method to analyze the features of big data in the network. Face to the problems of excessive data and long response time for network anomaly detection, we propose a trust-based Federated learning anomaly detection algorithm. We use the edge nodes to train the local data model, and upload the machine learning parameters to the central node. Meanwhile, according to the performance of edge nodes training, we set different weights to match the processing capacity of each terminal which will obtain faster convergence speed and better attack classification accuracy. The user’s private information will only be processed locally and will not be uploaded to the central server, which can reduce the risk of information disclosure. Finally, we compare the basic federated learning model and TFCNN algorithm on KDD Cup 99 dataset and MNIST dataset. The experimental results show that the TFCNN algorithm can improve accuracy and communication efficiency. Naiyue Chen, Yi Jin 0001, Yinglong Li, Luxin Cai |
Web Intell. | 2 |
| 2021 | Image-Based Indoor Localization Using Smartphone CameraabstractWith the increasing demand for location‐based services such as railway stations, airports, and shopping malls, indoor positioning technology has become one of the most attractive research areas. Due to the effects of multipath propagation, wireless‐based indoor localization methods such as WiFi, bluetooth, and pseudolite have difficulty achieving high precision position. In this work, we present an image‐based localization approach which can get the position just by taking a picture of the surrounding environment. This paper proposes a novel approach which classifies different scenes based on deep belief networks and solves the camera position with several spatial reference points extracted from depth images by the perspective‐n‐point algorithm. To evaluate the performance, experiments are conducted on public data and real scenes; the result demonstrates that our approach can achieve submeter positioning accuracy. Compared with other methods, image‐based indoor localization methods do not require infrastructure and have a wide range of applications that include self‐driving, robot navigation, and augmented reality. Baoguo Yu, Yi Jin 0001, Lu Huang 0001, Heng Zhang 0041, Xiaohu Liang |
Wirel. Commun. Mob. Comput. | 3 |
| 2020 | Domain Adaptive Attention Learning for Unsupervised Person Re-IdentificationabstractPerson re-identification (Re-ID) across multiple datasets is a challenging task due to two main reasons: the presence of large cross-dataset distinctions and the absence of annotated target instances. To address these two issues, this paper proposes a domain adaptive attention learning approach to reliably transfer discriminative representation from the labeled source domain to the unlabeled target domain. In this approach, a domain adaptive attention model is learned to separate the feature map into domain-shared part and domain-specific part. In this manner, the domain-shared part is used to capture transferable cues that can compensate cross-dataset distinctions and give positive contributions to the target task, while the domain-specific part aims to model the noisy information to avoid the negative transfer caused by domain diversity. A soft label loss is further employed to take full use of unlabeled target data by estimating pseudo labels. Extensive experiments on the Market-1501, DukeMTMC-reID and MSMT17 benchmarks demonstrate the proposed approach outperforms the state-of-the-arts. Yangru Huang, Peixi Peng, Yi Jin 0001, Yidong Li, Junliang Xing |
AAAI | 3 |
| 2020 | Learning Combinatorial Solver for Graph MatchingabstractLearning-based approaches to graph matching have been developed and explored for more than a decade, have grown rapidly in scope and popularity in recent years. However, previous learning-based algorithms, with or without deep learning strategy, mainly focus on the learning of node and/or edge affinities generation, and pay less attention on the learning of the combinatorial solver. In this paper we propose a fully trainable framework for graph matching, in which learning of affinities and solving for combinatorial optimization are not explicitly separated as in many previous arts. We firstly convert the problem of building node correspondences between two input graphs to the problem of selecting reliable nodes from a constructed assignment graph. Subsequently, the graph network block module is adopted to perform computation on the graph to form structured representations for each node. It finally predicts a label for each node that is used for node classification, and the training is performed under the supervision of both permutation differences and the one-to-one matching constraints. The proposed method is evaluated on four public benchmarks in comparison with several state-of-the-art algorithms, and the experimental results illustrate its excellent performance. Tao Wang 0011, Yidong Li, Yi Jin 0001, Xiaohui Hou, Haibin Ling |
CVPR | 4 |
| 2020 | 3DPC-Net: 3D Point Cloud Network for Face Anti-spoofingabstractFace anti-spoofing plays a vital role in face recognition systems. Most deep learning-based methods directly use 2D images assisted with temporal information (i.e., motion, rPPG) or pseudo-3D information (i.e., Depth). The main drawback of the mentioned methods is that another extra network is needed to generate the depth/rPPG information to assist the backbone network for face anti-spoofing. Different from these methods, we propose a novel method named 3D Point Cloud Network (3DPC-Net). It is an encoder-decoder network that can predict the 3DPC maps to discriminate live faces from spoofing ones. The main traits of the proposed method are that: 1) It is the first time that 3DPC is used for face anti-spoofing; 2) 3DPC-Net is simple and effective and it only relies on 3DPC supervision. Extensive experiments on four databases (i.e., Oulu-NPU, SiW, CASIA-FASD, Replay Attack) have demonstrated that the 3DPC-Net is comparative to the state-of-the-art methods. Jun Wan 0001, Yi Jin 0001, Ajian Liu 0001, Guodong Guo, Stan Z. Li |
IJCB | 3 |
| 2020 | Partial Label Learning via Self-Paced Curriculum Strategy
Gengyu Lyu, Songhe Feng, Yi Jin 0001, Yidong Li |
ECML/PKDD (2) | 3 |
| 2020 | Hierarchical discriminant feature learning for cross-modal face recognition
Yidong Li, Yi Jin 0001 |
Multim. Tools Appl. | 3 |
| 2020 | View-specific subspace learning and re-ranking for semi-supervised person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Shiming Ge |
Pattern Recognit. | 3 |
| 2020 | HERA: Partial Label Learning by Combining Heterogeneous Loss with Sparse and Low-Rank RegularizationabstractPartial label learning (PLL) aims to learn from the data where each training instance is associated with a set of candidate labels, among which only one is correct. Most existing methods deal with this type of problem by either treating each candidate label equally or identifying the ground-truth label iteratively. In this article, we propose a novel PLL approach named HERA, which simultaneously incorporates the HeterogEneous Loss and the SpaRse and Low-rAnk procedure to estimate the labeling confidence for each instance while training the desired model. Specifically, the heterogeneous loss integrates the strengths of both the pairwise ranking loss and the pointwise reconstruction loss to provide informative label ranking and reconstruction information for label identification, whereas the embedded sparse and low-rank scheme constrains the sparsity of ground-truth label matrix and the low rank of noise label matrix to explore the global label relevance among the whole training data, for improving the learning model. Comprehensive ablation study demonstrates the effectiveness of our employed heterogeneous loss, and extensive experiments on both artificial and real-world datasets demonstrate that our method achieves superior or comparable performance against state-of-the-art methods. Gengyu Lyu, Songhe Feng, Yidong Li, Yi Jin 0001, Guojun Dai, Congyan Lang |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2020 | Learning part-alignment feature for person re-identification with spatial-temporal-based re-ranking method
Yi Jin 0001, Yidong Li, Congyan Lang, Songhe Feng, Tao Wang 0011 |
World Wide Web | 2 |
| 2019 | Partial Multi-Label Learning by Low-Rank and Sparse DecompositionabstractMulti-Label Learning (MLL) aims to learn from the training data where each example is represented by a single instance while associated with a set of candidate labels. Most existing MLL methods are typically designed to handle the problem of missing labels. However, in many real-world scenarios, the labeling information for multi-label data is always redundant , which can not be solved by classical MLL methods, thus a novel Partial Multi-label Learning (PML) framework is proposed to cope with such problem, i.e. removing the the noisy labels from the multi-label sets. In this paper, in order to further improve the denoising capability of PML framework, we utilize the low-rank and sparse decomposition scheme and propose a novel Partial Multi-label Learning by Low-Rank and Sparse decomposition (PML-LRS) approach. Specifically, we first reformulate the observed label set into a label matrix, and then decompose it into a groundtruth label matrix and an irrelevant label matrix, where the former is constrained to be low rank and the latter is assumed to be sparse. Next, we utilize the feature mapping matrix to explore the label correlations and meanwhile constrain the feature mapping matrix to be low rank to prevent the proposed method from being overfitting. Finally, we obtain the ground-truth labels via minimizing the label loss, where the Augmented Lagrange Multiplier (ALM) algorithm is incorporated to solve the optimization problem. Enormous experimental results demonstrate that PML-LRS can achieve superior or competitive performance against other state-of-the-art methods. Songhe Feng, Tao Wang 0011, Congyan Lang, Yi Jin 0001 |
AAAI | 5 |
| 2019 | Unsupervised Person Re-identification Based on Clustering and Domain-Invariant Network
Yangru Huang, Yi Jin 0001, Peixi Peng, Congyan Lang, Yidong Li |
ICIG (3) | 2 |
| 2019 | FERLrTc: 2D+3D facial expression recognition via low-rank tensor completion
Yunfang Fu, Qiuqi Ruan, Ziyan Luo, Yi Jin 0001, Gaoyun An, Jun Wan 0001 |
Signal Process. | 4 |
| 2018 | Constrained Confidence Matching for Planar Object TrackingabstractTracking planar objects has a wide range of applications in robotics. Conventional template tracking algorithms, however, often fail to observe fast object motion or drift significantly after a period of time, due to drastic object appearance change. To address such challenges, we propose a novel constrained confidence matching algorithm for motion estimation and a robust Kalman filter for template updating. Integrated with an accurate occlusion detector, our approach achieves accurate motion estimation in presence of partial occlusion, by excluding occluded pixels from computation of motion parameters. Furthermore, the proposed Kalman filter employs a novel control-input model to handle the object appearance change, which brings our tracker high robustness against sudden illumination change and heavy motion blur. For evaluation, we compare the proposed tracker with several state-of-the-art planar object trackers on two public benchmark datasets. Experimental results show that our algorithm achieves robust tracking results against various environmental variations, and outperforms baseline algorithms remarkably on both datasets. Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li |
ICRA | 5 |
| 2018 | Spectrum-Centric Differential Privacy for Hypergraph Spectral Clustering
Xiaochun Wang, Yidong Li, Yi Jin 0001, Wei Wang 0012 |
PDCAT | 3 |
| 2018 | Hierarchical Discriminant Feature Learning for Heterogeneous Face RecognitionabstractHeterogeneous Face Recognition (HFR) refers to the problem of recognizing faces across different visual domains and has attached great attention owing to its tremendous potential benefits in practical applications. In this paper, a novel feature learning approach named hierarchical discriminant feature learning (HDFL) has been proposed for HFR. Different from traditional feature learning based HFR approaches, the proposed HDFL aims to learn the most discriminative information via a two-layer hierarchical boosting network (HBN), where the hierarchical discriminative information can be exploited in the learned features and the appearance difference can be effectively reduced, simultaneously. Extensive experiments on three different heterogeneous face databases demonstrate that our approach consistently outperforms the state-of-the-art methods. Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng, Tao Wang 0011 |
VCIP | 3 |
| 2018 | A novel hypergraph matching algorithm based on tensor refining
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Robust Object Tracking Based on Temporal and Spatial Deep NetworksabstractRecently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net, a Temporal Net, and a Spatial Net. The Feature Net extracts general feature representations of the target. With these feature representations, the Temporal Net encodes the trajectory of the target and directly learns temporal correspondences to estimate the object state from a global perspective. Based on the learning results of the Temporal Net, the Spatial Net further refines the object tracking state using local spatial object information. Extensive experiments on four of the largest tracking benchmarks, including VOT2014, VOT2016, OTB50, and OTB100, demonstrate competing performance of the proposed tracker over a number of state-of-the-art algorithms. Zhu Teng, Junliang Xing, Qiang Wang 0051, Congyan Lang, Songhe Feng, Yi Jin 0001 |
ICCV | 6 |
| 2017 | Multiple metric learning with query adaptive weights and multi-task re-weighting for person re-identification
Jieru Jia, Qiuqi Ruan, Gaoyun An, Yi Jin 0001 |
Comput. Vis. Image Underst. | 4 |
| 2016 | Geometric Preserving Local Fisher Discriminant Analysis for person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001 |
Neurocomputing | 3 |
| 2016 | Ensemble based extreme learning machine for cross-modality face matching
Yi Jin 0001, Jiuwen Cao, Ruicong Zhi |
Multim. Tools Appl. | 1 |
| 2015 | 3D human behavior recognition based on spatiotemporal texture featuresabstractNowadays, more and more activity recognition algorithms begin to improve recognition performance by combining the RGB and depth information. Although, the space-time volumes (STV) algorithm and the space-time local features algorithm can combine the RGB and depth information effectively, they also have their own defects. Such as they need expensive computational cost and they are not suitable for modeling nonperiodic activity. In this paper, we propose a novel algorithm for three dimensional human activity recognition that combines spatial-domain local texture features and spatio-temporal local texture features. On the one hand, in order to extract spatial local texture features, we mix the RGB and depth image sequence which have been applied with ViBe (Visual Background extractor) and binarization operator. Then we obtain the RGB-MOHBBI and depth-MOBHBI respectively and perform intersect operation on them. Afterwards, we extract LBP feature from the mixed MOHBBI to describe spatial domain feature. On the other hand, we follow the same background subtraction and binarization method to process the RGB and depth image sequences and get the spatial-temporal local texture features. And then, we project the three dimensional image volume on plane X-T and plane Y-T to get the spatio-temporal behavior volume change image to which we apply LBP operator to extract features that can represent human activity feature in spatio-temporal domain. At last, we combine the two local features that are extracted by LBP algorithm as one integrated feature of our model final output. Extensive experiments are conducted on the BUPT Arm Activity Dataset and the BUPT Arm And Finger Activity Dataset. The experimental results demonstrate the algorithm we proposed in this paper can make up for the deficiency of traditional activity recognition algorithms effectively and provide excellent experiment results on different databases of various complexities. Chunxiao Fan 0001, Lei Tian 0002, Guangchao Wang, Yue Ming 0001, Jiakun Shi, Yi Jin 0001 |
HSI | 6 |
| 2015 | Multiple strategies to enhance automatic 3D facial expression recognition
Xiaoli Li 0009, Qiuqi Ruan, Gaoyun An, Yi Jin 0001, Ruizhen Zhao |
Neurocomputing | 4 |
| 2015 | Fully automatic 3D facial expression recognition using polytypic multi-block local binary patterns
Xiaoli Li 0009, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Ruizhen Zhao |
Signal Process. | 3 |
| 2015 | Coupled Discriminative Feature Learning for Heterogeneous Face RecognitionabstractThis paper presents a coupled discriminative feature learning (CDFL) method for heterogeneous face recognition (HFR). Different from most existing HFR approaches which use hand-crafted feature descriptors for face representation, our CDFL directly learns discriminative features from raw pixels for face representation. In particular, a couple of image filters is learned in CDFL to simultaneously exploit discriminative information and to reduce the appearance difference of face images captured across different modalities. With the help of the learned filters, CDFL can maximize the interclass variations and minimize the intraclass variations of the learned feature vectors, and meanwhile maximize the correlation of face images of the same person from different modalities by solving a generalized eigenvalue problem. Experimental results on three different heterogeneous face recognition applications show the effectiveness of our proposed approach. Yi Jin 0001, Jiwen Lu, Qiuqi Ruan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Complete discriminative feature learning: A new approach for heterogeneous face recognitionabstractIn this paper, we propose a new feature learning approach called complete discriminative feature learning (CDFL) for heterogeneous face recognition. Unlike most existing heterogeneous face recognition methods where hand-crafted feature descriptors are used for face representation, the proposed CD-FL aims to learn an optimal weighted discriminative image filter to improve learning discriminative filters, so that complete discriminative information is exploited and the feature difference between different modalities is effectively reduced, simultaneously. Experimental results shows that our approach consistently outperforms the state-of-the-art methods. Yi Jin 0001, Jiwen Lu, Qiuqi Ruan, Yap-Peng Tan |
ICME | 1 |
| 2012 | Orthogonal tensor rank one differential graph preserving projections with its application to facial expression recognition
Shuai Liu 0003, Qiuqi Ruan, Yi Jin 0001 |
Neurocomputing | 3 |
| 2008 | Fusing Global and Local Complete Linear Discriminant Features by Fuzzy Integral for Face RecognitionabstractFace recognition becomes very difficult in a complex environment, and the combination of multiple classifiers is a good solution to this problem. A novel face recognition algorithm GLCFDA-FI is proposed in this paper, which fuses the complementary information extracted by complete linear discriminant analysis from the global and local features of a face to improve the performance. The Choquet fuzzy integral is used as the fusing tool due to its suitable properties for information aggregation. Experiments are carried out on the CAS-PEAL-R1 database, the Harvard database and the FERET database to demonstrate the effectiveness of the proposed method. Results also indicate that the proposed method GLCFDA-FI outperforms five other commonly used algorithms — namely, Fisherfaces, null space-based linear discriminant analysis (NLDA), cascaded-LDA, kernel-Fisher discriminant analysis (KFDA), and null-space based KFDA (NKFDA). Qiuqi Ruan, Yi Jin 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Gabor-Based Improved Locality Preserving Projections for Face RecognitionabstractA novel Gabor-based improved locality preserving projections for face recognition is presented in this paper. This new algorithm is based on a combination of Gabor wavelets representation of face images and improved locality preserving projections for face recognition and it is robust to changes in illumination and facial expressions and poses. In this paper, Gabor filter is first designed to extract the features from the whole face images, and then a locality preserving projections, which is improved by two-directional 2DPCA to eliminate redundancy among Gabor features, is used to subject these feature vectors onto locality subspace projection. Experiments based on the ORL face database demonstrate the effectiveness and efficiency of the new method. Results show that our new algorithm outperforms the other popular approaches reported in the literature and achieves a much higher accurate recognition rate. Yi Jin 0001, Qiuqi Ruan |
ICIP (1) | 1 |