VLDB 2026 Research / reviewers in the wild / expert
Yilong Yin
dblp:94/458
· DBLP profile ↗
307ranked-venue papers
1as first author
180since 2021 · last 2027
0000-0002-8465-1294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 142 · 87 since 2021Graphics, computer vision, multimedia, augmented reality and games · 140 · 1 first-author · 86 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 17 since 2021Databases, data management, data science and information retrieval · 27 · 12 since 2021Security and privacy · 11 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Computer networks · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Principled efficient triple-weighting for AUC-oriented imbalanced covariate shift
Yan Zhang 0145, Guoqiang Wu, Yilong Yin |
Expert Syst. Appl. | 6 |
| 2026 | PEOCH: Online Cross-Modal Hashing with Semi-Supervised Streaming Data Driving Prototype EvolutionabstractThe exponential growth of streaming multi-modal data presents critical challenges for cross-modal retrieval: distribution shifts, modality gap, and scarce labels. Semi-supervised online cross-modal hashing has gained increasing interest due to its ability to encode complex streaming data and update hash functions simultaneously. Nevertheless, existing methods can hardly generate high-quality unsupervised hash codes, which fundamentally limits diversity and flexibility during the retrieval process. To this end, we propose a novel method named Prototype Evolution Online Cross-modal Hashing (PEOCH). By driving prototype evolution with semi-supervised streaming data, precise and stable hash codes are generated for both labeled and unlabeled data. Specifically, two prototype updates with stability guarantee are conducted: labeled samples push semantic knowledge into the supervised prototypes, while unlabeled samples perform clustering to generate unsupervised prototypes. Simultaneously, a co-optimization mechanism is designed to ensure the prototypes continuously evolve and preserve the consistency of the entire streaming data. Besides, an elasticity regularizer integrates discriminability and smoothness constraints, improving the reliability of prototypes. Extensive experiments on three benchmark datasets demonstrate that PEOCH outperforms state-of-the-art methods, achieving an average improvement of 6.7% in mAP@all across various retrieval tasks. Xiao Kang, Xingbo Liu, Shuo Pan, Xuening Zhang, Xiushan Nie, Yilong Yin |
AAAI | 6 |
| 2026 | MTRL-CG: Multi-Task Reinforcement Learning Method with Spectral Clustering-Based Task GroupingabstractMulti-task reinforcement learning (RL) aims to enhance agent performance across multiple tasks by enabling effective knowledge transfer. However, these methods adopt a fully shared policy across all tasks without explicitly distinguishing between related and conflicting ones, making them suffer from negative interference issue, where updates beneficial to one task adversely affect others and lead to degraded overall performance. In this paper, we propose a multi-task reinforcement learning method with spectral clustering-based task grouping (MTRL-CG), which leverages spectral clustering to group related tasks and separate conflicting ones, enabling group-wise policy learning to mitigate negative interference. We first quantify inter-task affinity by measuring the influence of task-specific updates on others within a shared model, and construct an affinity matrix to capture these relationships. Spectral clustering is then applied to partition tasks via spectral embedding and k-means clustering. Each task group is trained with a dedicated policy network to promote focused learning. Built upon the Soft Actor-Critic (SAC) algorithm, MTRL-CG can be readily integrated into existing SAC-based multi-task RL methods. Extensive experiments on the Meta-World benchmark demonstrate the effectiveness of the proposed MTRL-CG method. Wenjia Meng, Haoliang Sun, Yilong Yin |
AAAI | 4 |
| 2026 | Retriever Encoder Selection Matters for In-Context Learning-based Medical SegmentationabstractIn-context learning-based medical segmentation (ICLM) enables foundation models to generalize to unseen cases without retraining. To enhance performance on test queries, existing methods typically follow a two-stage process: (1) using a retrieval encoder (RE) to map both queries and training samples into a shared feature space, and (2) retrieving and utilizing the top-k most similar training samples. While current methods fix the RE and focus on optimizing stage (2), we show that the choice of RE in stage (1) alone can account for over 70% of the performance variation, highlighting RE selection as a critical yet often overlooked factor in ICLM. In this paper, we conduct an analysis of the RE selection and make two main findings: (1) dynamically selecting the RE for each query outperforms selecting a fixed RE for the entire task; and (2) feature-space heuristics (e.g., intra-class compactness and inter-class separability) fail to predict RE quality. To this end, we propose the instance-adaptive retrieval encoder selection (IRES) method that can select the optimal RE for each query based on output predictions. IRES is based on the intuition that a good RE retrieves relevant demonstrations, helping the ICL model generate more accurate and stable segmentation masks. Thus, we introduce the shape stability score (S³), which evaluates the morphological stability of predicted masks under iterative erosion. Experiments show S³ correlates strongly with true RE quality (Pearson > 0.8), serving as a reliable selection proxy. To reduce S³’s per-query cost, we propose parallel prediction with reciprocal neighbor reuse (P2R), which accelerates inference by parallelizing encoding and reusing encoder selections across reciprocal neighbors, avoiding redundant computation. Built on S³ and P2R, IRES improves ICLM performance across FUNDUS, Brain MRI, and Chest X-ray datasets, with up to 10.6% gain on fundus segmentation. Zhongyi Han, Yongshun Gong, Yilong Yin |
AAAI | 4 |
| 2026 | Recent advances of local mechanisms in vision foundation models: A survey and outlook
Qiangchang Wang, Jing Li 0175, Yilong Yin, Huimin Lu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2026 | STADNN: Spatio-temporal adaptive decomposition neural network for traffic prediction
Yongshun Gong, Xiushan Nie, Yilong Yin |
Neurocomputing | 6 |
| 2026 | SA-Diff: Semantic-Aware graph outlier generation via diffusion models for graph out-of-Distribution detection
Yicong Dong, Rundong He, Zhongyi Han, Jieming Shi 0001, Yilong Yin |
Knowl. Based Syst. | 5 |
| 2026 | Typical-smoothed truncation for out-of-distribution detection
Feichao Wang, Rundong He, Yilong Yin |
Knowl. Based Syst. | 4 |
| 2026 | Dual-level self-adaptive threshold learning for semi-supervised CNV classification
Jie Guo 0012, Lingzhao Meng, Ying Guo 0030, Fengxiang Li, Yipeng Ning, Lishan Qiao, Nianying Sun, Xiaoming Xi, Yilong Yin |
Pattern Recognit. | 11 |
| 2026 | Class-mismatched semi-supervised learning from a new perspective
Rundong He, Zhongyi Han, Xiushan Nie, Qi Wei 0004, Yilong Yin |
Pattern Recognit. | 6 |
| 2026 | Multi-feature embedding, fusion and enhancement for partial finger vein recognition
Enyan Li, Lu Yang 0005, Qiangchang Wang, Yilong Yin |
Pattern Recognit. | 5 |
| 2026 | A unified framework to learn invariant representations of graph neural networks for ECG biometrics
Tianbang Ma, Chunying Liu, Yilong Yin, Gongping Yang 0001, Jinshan Pan |
Pattern Recognit. | 4 |
| 2026 | Diffusion classifier-driven reward for offline preference-based reinforcement learning
Teng Pang, Bingzheng Wang, Guoqiang Wu, Yilong Yin |
Pattern Recognit. | 4 |
| 2026 | Internal-External Context Interaction Network for Person Re-IdentificationabstractCapturing discriminative cues with attention mechanisms is crucial for solving the high inter-class similarity problem of person re-identification (Re-ID). Self-attention (SA) learns its own contextual information within a single sample using self-affinity between elements, and some works have demonstrated its superiority in person Re-ID. However, SA weakens some subtle semantic cues and additional visual cues such as backpacks, which makes it difficult to distinguish similar-looking persons. In this paper, we propose an internal-external context interaction (IEI) attention mechanism, which aims to exploit the interaction of inter-sample latent context information and intra-sample local context information to enhance the feature representation of each element. The mechanism is able to capture subtle differences between persons and additional visual cues using inter-sample difference information and rich detail information within the element neighborhood, improving the ability to distinguish similar persons. Based on this mechanism, we propose an internal-external context interaction network (IEINet) for extracting discriminative features from multiple dimensions. In addition, to capture more discriminative information, we propose a region-diverse loss to constrain the network. Many experiments validate the effectiveness of our IEINet and demonstrate that our approach attains state-of-the-art performance on several large-scale person Re-ID datasets. Tongxin Liu, Xiyu Pang, Gangwu Jiang, Xiushan Nie, Meifeng Zheng, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Prior Distribution Guided Gaussian Mixture Variational Autoencoder (PDGM-VAE) for Image GenerationabstractVariational Autoencoder(VAE) combines the ideas of autoencoders and variational inference, introducing the concept of latent space and variational inference to endow autoencoders to generate new images. VAE typically assumes that data follows a Gaussian distribution, but real data may follow other distributions. This inconsistency between the assumption and the true distribution can affect the modeling and reconstruction capabilities of VAE, which makes it difficult for traditional models to accurately capture the true distribution. To address the aforementioned issues, we propose a Prior Distribution Guided Gaussian Mixture Variational Autoencoder(PDGM-VAE). Specifically, we construct a Gaussian Mixture Prior Learner (GMPL) to capture complex features of the data distribution, enabling the model to learn and obtain a Gaussian mixture distribution that is reasonable and close to the real data distributions, which is then used as the prior distribution in the network. Furthermore, we build a Semantic-Aware Module with Embedded Prior Distribution (SAMEPD), integrating data and label information to learn the distribution parameters, enabling the network to learn and utilize the semantic knowledge contained in the labels. During training, by approximating the posterior distribution to the prior distribution, we enhance the model’s modeling and reconstruction capabilities, improving the quality of generated images. We evaluated the image generation task on five public datasets, and based on the FID metric, our proposed method outperformed other VAE methods. Jingqi Song, Yipeng Ning, Xiaoming Xi, Jie Guo 0012, Xiushan Nie, Lishan Qiao, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Local Refinement and Global Strengthening Network for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) aims to retrieve vehicle images with the same identity as the query from an image library. Currently, the vehicle Re-ID task mainly faces two challenges: large intra-class variance and small inter-class variance. Learning discriminative local features and global features of vehicles is crucial to address both challenges, and the attention mechanism adequately learns local features and global features in vehicle images without the aid of an auxiliary model. Self-attention mechanism as a special kind of attention mechanism, it mainly contains two forms of local self-attention for extracting local features and Full self-attention for extracting global features. However, these two approaches have their own limitations: the window mode of local self-attention hinders adequate learning of the local detailed information of vehicles; the remote connections in the global context modeled by full self-attention are usually weak, which limits the full learning of the overall information about vehicles. To address the above problem, we propose two complementary modules: local refinement module (LRM) and global strengthening module (GSM). The LRM aims to learn the refined local representation, which captures the rich correlation information between adjacent pixels through the interactions of the target pixel with its nearest pixels. The GSM aims to learn the strengthened global representation, which first disperses attention at the target pixel into various windows to emphasize important remote dependence within each region and then aggregates globally meaningful remote connections by cross-window interaction. In addition, we construct a multi-branch network, local refinement and global strengthening network (LRGS-Net), which uses LRM and GSM to learn discriminative local features and global features to address vehicle Re-ID challenges. We validate the effectiveness of our method on three datasets, VeRi-776, VehicleID, and VERI-Wild. Meifeng Zheng, Xiyu Pang, Xiushan Nie, Houren Zhou, Yilong Yin |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2026 | Difficulty-Aware Pseudo-Label Correction Network for Fine-Grained Classification of Choroidal Neovascularization in OCT ImagesabstractChoroidal neovascularization (CNV) classification is a fine-grained classification task. Accurate classification of CNV in optical coherence tomography (OCT) images is crucial for clinical treatment. However, image acquisition noise degrades image quality and exacerbates confirmation bias from class imbalance in medical datasets. Moreover, significant inter-class ambiguity in fine-grained categories can misclassify informative samples (e.g., hard samples or minority class samples) when generating pseudo-labels, leading to sub-optimal classifiers. To address these challenges, we propose a difficulty-aware pseudo-label correction network (DPLC-Net). Specifically, we designed a robust feature mining module using feature similarity loss to maintain consistency between generated adversarial and original samples, enabling noise-resistant feature learning. A difficulty-aware pseudo-label correction module mines and corrects potential noisy pseudo-labels to improve classification performance. Finally, to alleviate data bias and leverage all unlabeled samples, we integrated a hybrid consistency and pseudo-labeling module comprising adaptive weighted consistency loss (AWCL) and class-aware dynamic threshold strategy (CDTS). AWCL adaptively learns weights for unlabeled samples, effectively utilizing all unlabeled data through weighted consistency loss. CDTS dynamically adjusts confidence thresholds based on class distribution and model learning status, improving pseudo-label quantity and quality. Experiments on private and public OCT datasets demonstrate that our method outperforms state-of-the-art methods. Lingzhao Meng, Xiaoming Xi, Lishan Qiao, Yilong Yin, Xinjian Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Semi-Supervised Online Cross-Modal HashingabstractOnline cross-modal hashing has gained increasing interest due to its ability to encode streaming data and update hash functions simultaneously. Existing online methods often assume either fully supervised or completely unsupervised settings. However, they overlook the prevalent and challenging scenario of semi-supervised cross-modal streaming data, where diverse data types, including labeled/unlabeled, paired/unpaired, and multi-modal, are intertwined. To address this issue, we propose Semi-Supervised Online Cross-modal Hashing (SSOCH). It presents an alignment-free pseudo-labeling strategy that extracts semantic information from unlabeled streaming data without relying on pairing relations. Furthermore, we design an online tri-consistent preserving scheme, integrating pseudo-labeled data regularization, discriminative label embedding, and fine-grained similarity preservation. This scheme fully explores consistency across data annotation, modalities, and streaming chunks, improving the model's adaptiveness in these challenging scenarios. Extensive experiments on benchmark datasets demonstrate the superiority of SSOCH under various scenarios, highlighting the importance of semi-supervised learning for online cross-modal hashing. Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin |
AAAI | 6 |
| 2025 | Generalized Debiased Semi-Supervised Hashing for Large-Scale Image RetrievalabstractSemi-supervised hashing has shown promising efficacy in large-scale image retrieval, which learns similarity-preserving codes from both labeled and unlabeled data. To enable the use of advanced supervised hashing techniques, pseudo labels are widely applied. However, existing methods typically suffer from a biased learning issue due to pseudo label noise, which can be further aggravated during optimization. Although such a bias can adversely affect hashing accuracy, it has not been investigated sufficiently. In view of this, we present a comprehensive discussion on potential causes of biases, involving processes of pseudo-labeling, hash learning and optimization. Accordingly, a novel Generalized Debiased Semi-supervised Hashing (GDSH) method is proposed as a unified solution to mitigate the biases. Specifically, reliable pseudo labels are first predicted via a robust label completion strategy. Secondly, a debiased hash learning module is designed by combining label denoising and similarity updating. This can not only refine the supervision, but also obtain hash codes that are semantically debiased in both category and sample levels. Finally, a discrete semi-supervised hashing algorithm is proposed to alleviate the bias arising from optimization. Experimental results on three single-label and three multi-label image benchmarks demonstrate that GDSH remarkably outperforms the state-of-the-arts in different semi-supervised settings. Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin |
AAAI | 5 |
| 2025 | Towards Macro-AUC Oriented Imbalanced Multi-Label Continual LearningabstractIn Continual Learning (CL), while existing work primarily focuses on the multi-class classification task, there has been limited research on Multi-Label Learning (MLL). In practice, MLL datasets are often class-imbalanced, making it inherently challenging, a problem that is even more acute in CL. Due to its sensitivity to imbalance, Macro-AUC is an appropriate and widely used measure in MLL. However, there is no research to optimize Macro-AUC in MLCL specifically. To fill this gap, in this paper, we propose a new memory replay-based method to tackle the imbalance issue for Macro-AUC-oriented MLCL. Specifically, inspired by recent theory work, we propose a new Reweighted Label-Distribution-Aware Margin (RLDAM) loss. Furthermore, to be compatible with the RLDAM loss, a new memory-updating strategy named Weight Retain Updating (WRU) is proposed to maintain the numbers of positive and negative instances of the original dataset in memory. Theoretically, we provide superior generalization analyses of the RLDAM-based algorithm in terms of Macro-AUC, separately in batch MLL and MLCL settings. This is the first work to offer theoretical generalization analyses in MLCL to our knowledge. Finally, a series of experimental results illustrate the effectiveness of our method over several baselines. Yan Zhang 0145, Guoqiang Wu, Bingzheng Wang, Teng Pang, Haoliang Sun, Yilong Yin |
AAAI | 6 |
| 2025 | Empowering Multimodal Models via Active in-Context Learning for Test-Time Medical ImagingabstractIn-context learning (ICL) has enabled large multimodal models (LMMs) to achieve effective medical image classification through the strategic utilization of relevant examples from pre-existing training sets. However, building high-quality training sets for ICL in medical imaging remains challenging due to costly manual annotations and strict privacy constraints. In this paper, we propose Active In-Context Learning (AICL), a novel paradigm that eliminates the need for pre-existing training sets. AICL dynamically selects and annotates a small, informative set of medical samples at dynamic test time during the query phase, by continuously retrieving relevant ICL examples to optimize LMM performance without relying on traditional datasets. To construct an optimal active set, we introduce Neighbor-relaxed Representative Sampling, which applies spectral clustering within each batch to select class-balanced and representative samples. By incorporating neighbor relaxation across batches, this module ensures sample diversity and better captures the overall data distribution. To fully utilize the active set, we propose Similarity-enhanced TopK Prompt Construction, which retrieves the most relevant multimodal examples using a TopK similarity strategy and embeds their visual similarities with the query samples into the text prompts. This enhances LMMs' understanding of relationships, enabling more accurate and context-aware predictions. Experiments on nine specialized medical datasets across four LMMs show the effectiveness of our method. Zhongyi Han, Yilong Yin |
BIBM | 3 |
| 2025 | SeqMvRL: A Sequential Fusion Framework for Multi-view Representation LearningabstractMulti-view representation learning integrates multiple observable views of an entity into a unified representation to facilitate downstream tasks. Current methods predominantly focus on distinguishing compatible components across views, followed by a single-step parallel fusion process. However, this parallel fusion is static in essence, overlooking potential conflicts among views and compromising representation ability. To address this issue, this paper proposes a novel Sequential fusion framework for Multi-view Representation Learning, termed SeqMvRL. Specifically, we model multi-view fusion as a sequential decision-making problem and construct a pairwise integrator (PI) and a next-view selector (NVS), which represent the environment and agent in reinforcement learning, respectively. PI merges the current fused feature with the selected view, while NVS is introduced to determine which view to fuse subsequently. By adaptively selecting the next optimal view for fusion based on the current fusion state, SeqMvRL thereby effectively reduces conflicts and enhances unified representation quality. Additionally, an elaborate novel reward function encourages the model to prioritize views that enhance the discriminability of the fused features. Experimental results demonstrate that SeqMvRL outperforms parallel fusion schemes in classification and clustering tasks. Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Chuanhui Zuo, Yongshun Gong, Yilong Yin, Wenjia Meng |
CVPR | 6 |
| 2025 | Disparity-Guided Cross-View Transformer For Stereo Image Super-ResolutionabstractAlthough transformer-based methods excel in stereo image super-resolution, the full potential of the distinctive, complementary information inherent in stereo images has not been fully utilized. We propose a Disparity-Guided Cross-View Transformer (DCT) to extract features across dimensions and views, achieving a more comprehensive feature representation. The proposed method introduces mutual attention within the transformer architecture, establishing the difference between left and right views through cross-view interaction. The proposed algorithm effectively harnesses the complementary information present in stereo image pairs, enhancing the restoration performance. Furthermore, we propose a disparity-guided cross-modal residual fusion module that leverages disparity information as prior knowledge to substantially improve image reconstruction. This module significantly complements the missing information in stereo images, enabling the network to comprehend more effectively and reconstruct the image content with greater accuracy. Extensive experimental results and ablation studies demonstrate the effectiveness of our method. Bingting Li, Wenjing Shang, Yongshun Gong, Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
ICASSP | 6 |
| 2025 | LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy LabelsabstractFacial expression recognition (FER) faces unique challenges from expression ambiguity and noisy labels, degrading performance in real-world applications. While leveraging attention, existing methods frequently neglect attention dynamic mechanism of dispersion followed by focus and the spatially structural knowledge essential for effectively guiding this dynamic dispersion of attention. To address this, we propose the Last fOcus First dIsperse (LOFI), which harnesses attention dynamics and spatial structure information dynamically refining focus during classification to mitigate the impact of noise labels. LOFI comprises two modules: Spatial Keypoint-enhanced Fused Attention (SKFA), which disperses focus on subtle, critical features near facial landmarks, and Hybrid Consistency-Calibrated Loss (HCCL), which employs consistency and re-weighting strategies focusing attention to boost performance. The synergy between these modules enables LOFI to adapt to various noise levels and challenging classes. Extensive experiments demonstrate that LOFI outperforms existing state-of-the-art (SOTA) methods in noisy FER, offering a robust solution for real-world applications. Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
ICASSP | 4 |
| 2025 | N3C: Towards Replay-based Novelty Continual Clustering with Class-OverlappingabstractDeep clustering has excelled in batch settings, but little work has addressed the more practical and challenging continual clustering (CC) with shifting data distributions. Additionally, class-overlapping, also a challenging issue, where classes recur across tasks, is common in real-world scenarios. In this paper, we introduce a new framework for CC with class-overlapping, integrating OOD detection to distinguish between old and new classes and a two-step deep clustering process: contrastive learning for feature representation and rehearsal-based learning to retain previous knowledge. We also propose a memory-updating strategy for handling unsupervised data. Experiments validate our approach, examining factors like OOD detection, class-overlapping levels, etc. This work advances continual clustering toward real-world applications. Yan Zhang 0145, Guoqiang Wu, Bingzheng Wang, Teng Pang, Yilong Yin |
ICASSP | 5 |
| 2025 | A Conditional Probability Framework for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of known objects and attributes by leveraging knowledge from previously seen compositions. Traditional approaches primarily focus on disentangling attributes and objects, treating them as independent entities during learning. However, this assumption overlooks the semantic constraints and contextual dependencies inside a composition. For example, certain attributes naturally pair with specific objects (e.g., "striped" applies to "zebra" or "shirts" but not "sky" or "water"), while the same attribute can manifest differently depending on context (e.g., "young" in "young tree" vs. "young dog"). Thus, capturing attribute-object interdependence remains a fundamental yet long-ignored challenge in CZSL. In this paper, we adopt a Conditional Probability Framework (CPF) to explicitly model attribute-object dependencies. We decompose the probability of a composition into two components: the likelihood of an object and the conditional likelihood of its attribute. To enhance object feature learning, we incorporate textual descriptors to highlight semantically relevant image regions. These enhanced object features then guide attribute learning through a cross-attention mechanism, ensuring better contextual alignment. By jointly optimizing object likelihood and conditional attribute likelihood, our method effectively captures compositional dependencies and generalizes well to unseen compositions. Extensive experiments on multiple CZSL benchmarks demonstrate the superiority of our approach. Code is available at here. Peng Wu 0014, Qiuxia Lai, Hao Fang 0010, Guosen Xie, Yilong Yin, Xiankai Lu, Wenguan Wang |
ICCV | 5 |
| 2025 | Re-Evaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning ModelabstractSemi-supervised learning (SSL) effectively leverages unlabeled data and has been proven successful across various fields. Current safe SSL methods believe that unseen classes in unlabeled data harm the performance of SSL models. However, previous methods for assessing the impact of unseen classes on SSL model performance are flawed. They fix the size of the unlabeled dataset and adjust the proportion of unseen classes within the unlabeled data to assess the impact. This process contravenes the principle of controlling variables. Adjusting the proportion of unseen classes in unlabeled data alters the proportion of seen classes, meaning the decreased classification performance of seen classes may not be due to an increase in unseen class samples in the unlabeled data, but rather a decrease in seen class samples. Thus, the prior flawed assessment standard that "unseen classes in unlabeled data can damage SSL model performance" may not always hold true. This paper strictly adheres to the principle of controlling variables, maintaining the proportion of seen classes in unlabeled data while only changing the unseen classes across five critical dimensions, to investigate their impact on SSL models from global robustness and local robustness. Experiments demonstrate that unseen classes in unlabeled data do not necessarily impair the performance of SSL models; in fact, under certain conditions, unseen classes may even enhance them. Rundong He, Yicong Dong, Lanzhe Guo, Yilong Yin, Tailin Wu |
ICLR | 4 |
| 2025 | Semantic-Aware Adaptation with Hierarchical Multimodal Prompts for Few-Shot LearningabstractFew-shot learning aims to recognize novel classes with limited labeled samples. Existing methods often utilize semantic information from natural language but integrate it after visual feature extraction, overlooking fine-grained cross-modal interactions. Moreover, they struggle with spatial variations, as target objects often appear in varying regions. To address these limitations, we propose Semantic-Aware Adaptation (SAA) which consists of Hierarchical Multimodal Prompts (HMP) and Global-Local Adaptation (GLA). SAA leverages textual prompts encoded by CLIP and adaptively modulated by learnable visual prompts to better align text and vision feature distributions. During visual extraction, these fused prompts are integrated with visual patches in channel and spatial dimensions, dynamically enhancing visual features, while a consistency loss is used to regularize them to prevent bias and overfitting. By the deep cross-modal interactions at different scales, HMP is thus constructed for improving robustness to spatial variations. In GLA, patch-level soft labels based on rich semantics further emphasize class-specific visual patches, improving token dependency learning. Experiments on five benchmarks demonstrate the effectiveness of our SAA, with an average accuracy improvement of 1.6% on challenging 1-shot tasks. Wenhao Li 0011, Qiangchang Wang, Jing Li 0175, Mindi Ruan, Yilong Yin |
ICME | 6 |
| 2025 | Pixel-wise Single Image Reflection Removal Method Based on Reinforcement LearningabstractSingle image reflection removal is particularly important in improving image quality. However, existing single image reflection removal methods cannot remove reflection in a pixel-by-pixel manner, significantly reducing their effectiveness. To address this issue, we propose a pixel-wise single image reflection removal method based on reinforcement learning. Specifically, we formalize the image reflection removal process as a sequential decision-making process and remove the single image reflection pixel by pixel. We first introduce the single image reflection removal method at each time step. We then describe the reinforcement learning setup for image reflection, including state, action, reward, and agent network design. Finally, we conducted experiments on benchmark datasets, and the results show that our method outperforms other single image reflection removal methods. Xueshi Yu, Zhengzhe Zhang, Xiankai Lu, Yilong Yin, Wenjia Meng |
ICME | 5 |
| 2025 | Improving Generalization in Meta-Learning via Meta-Gradient AugmentationabstractMeta-learning methods typically follow a two-loop framework, where each loop potentially suffers from notorious overfitting, hindering rapid adaptation and generalization to new tasks. Existing methods address this by enhancing the mutual-exclusivity or diversity of training samples, but these data manipulation strategies are data-dependent and insufficiently flexible. This work proposes a data-independent Meta-Gradient Augmentation (MGAug) method from the perspective of gradient regularization. The key idea is first to break the rote memories by network pruning to address memorization overfitting in the inner loop, then use the gradients of pruned sub-networks to augment meta-gradients, alleviating overfitting in the outer loop. Specifically, we explore three pruning strategies, including random width pruning, random parameter pruning, and a newly proposed catfish pruning that measures a Meta-Memorization Carrying Amount (MMCA) score for each parameter and prunes high-score ones to break rote memories. The proposed MGAug is theoretically guaranteed by the generalization bound from the PAC-Bayes framework. Extensive experiments on multiple few-shot learning benchmarks validate MGAug's effectiveness and significant improvement over various meta-baselines. Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Xinxin Zhang 0004, Yilong Yin |
IJCAI | 5 |
| 2025 | VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningabstractFew-shot learning (FSL) aims to recognize novel concepts from only a few labeled support samples. Recent studies enhance support features by incorporating additional semantic information (e.g., class descriptions) or designing complex semantic fusion modules. However, these methods still suffer from hallucinating semantics that contradict the visual evidence due to the lack of grounding in actual instances, resulting in noisy guidance and costly corrections. To address these issues, we propose a novel framework, bridging Vision and Text with LLMs for Few-Shot Learning (VT-FSL), which constructs precise cross-modal prompts conditioned on Large Language Models (LLMs) and support images, seamlessly integrating them through a geometry-aware alignment mechanism. It mainly consists of Cross-modal Iterative Prompting (CIP) and Cross-modal Geometric Alignment (CGA). Specifically, the CIP conditions an LLM on both class names and support images to generate precise class descriptions iteratively in a single structured reasoning pass. These descriptions not only enrich the semantic understanding of novel classes but also enable the zero-shot synthesis of semantically consistent images. The descriptions and synthetic images act respectively as complementary textual and visual prompts, providing high-level class semantics and low-level intra-class diversity to compensate for limited support data. Furthermore, the CGA jointly aligns the fused textual, support, and synthetic visual representations by minimizing the kernelized volume of the 3-dimensional parallelotope they span. It captures global and nonlinear relationships among all representations, enabling structured and consistent multimodal integration. The proposed VT-FSL method establishes new state-of-the-art performance across ten diverse benchmarks, including standard, cross-domain, and fine-grained few-shot learning scenarios. Code is available at https://github.com/peacelwh/VT-FSL. Wenhao Li 0011, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin |
NeurIPS | 5 |
| 2025 | From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical ModelsabstractFoundation models pretrained on web-scale data drive contemporary transfer learning in vision, language, and multimodal tasks. Recent work shows that mild label noise in these corpora may lift in-distribution accuracy yet sharply reduce out-of-distribution generalization, an effect known as catastrophic inheritance. Medical data is especially sensitive because annotations are scarce, domain shifts are large, and pretraining sources are noisy.
We present the first systematic analysis of catastrophic inheritance in medical models. Controlled label-corruption experiments expose a clear structural collapse: as noise rises, the skewness and kurtosis of feature and logit distributions decline, signaling a flattened representation space and diminished discriminative detail. These higher-order statistics form a compact, interpretable marker of degradation in fine-grained tasks such as histopathology.
Guided by this finding, we introduce a fine-tuning objective that restores skewness and kurtosis through two scalar regularizers added to the task loss. The method leaves the backbone unchanged and incurs negligible overhead. Tests on PLIP models trained with Twitter pathology images, as well as other large-scale vision and language backbones, show consistent gains in robustness and cross-domain accuracy under varied noise levels. Hao Sun 0002, Zhongyi Han, Hao Chen 0102, Jindong Wang 0001, Xin Gao 0001, Yilong Yin |
NeurIPS | 6 |
| 2025 | Dynamic prompt allocation and tuning for continual test-time adaptation
Chaoran Cui, Yongrui Zhen, Shuai Gong, Chunyun Zhang, Hui Liu 0016, Yilong Yin |
Sci. China Inf. Sci. | 6 |
| 2025 | Data transmission optimization based on multi-objective deep reinforcement learningabstractAbstract Simultaneously reducing network energy consumption and delay is a hot topic today. This paper addresses this issue by designing a novel multi-objective data transmission optimization algorithm based on deep reinforcement learning. A three-layer back propagation (BP) neural network is designed to improve the accuracy of environmental state prediction, by learning from historical state and action sequence data, which can help the agent make better decision for routing selection in complex network environment. Based on this, we use Q-Learning to find routing for transmission demands, aggregating more traffic through less links and routers, to reduce energy consumption and delay. To enhance the efficiency and robustness of the algorithm, a new reward mechanism is designed based on the traffic demand and the link state. The algorithm divides candidate links into three levels for path selection so that a better solution can be obtained on the basis of ensuring feasible solutions are obtained. Continuous updating of the Pareto set through multiple state steps approximates the optimal solution. We leverage the Euclidean distance to the reference point to measure the optimization effect of the two objectives. The simulation results show that this algorithm outperforms existing algorithms in reducing energy consumption and network delay. Cuiping Wang, Xiaole Li, Jinwei Tian, Yilong Yin |
Comput. J. | 4 |
| 2025 | Cross-graph meta matching correction for noisy graph matching
Fangkai Li, Feiyu Pan, Wenjia Meng, Haoliang Sun, Xiushan Nie, Yilong Yin, Xiankai Lu |
Comput. Vis. Image Underst. | 6 |
| 2025 | Leveraging spatio-temporal multi-task learning for potential urban flow prediction in newly developed regions
Wenqian Mu, Yongshun Gong, Xiushan Nie, Yilong Yin |
Expert Syst. Appl. | 5 |
| 2025 | Enhancing origin-destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution
Piao Yu, Xu Zhang 0039, Yongshun Gong, Jian Zhang 0002, Haoliang Sun, Junjie Zhang 0002, Xinxin Zhang 0004, Yilong Yin |
Expert Syst. Appl. | 8 |
| 2025 | Variational Rectification Inference for Learning with Noisy Labels
Haoliang Sun, Qi Wei 0004, Lei Feng 0006, Yupeng Hu 0003, Fan Liu 0008, Hehe Fan, Yilong Yin |
Int. J. Comput. Vis. | 7 |
| 2025 | Correction: Variational Rectification Inference for Learning with Noisy Labels
Haoliang Sun, Qi Wei 0004, Lei Feng 0006, Yupeng Hu 0003, Fan Liu 0008, Hehe Fan, Yilong Yin |
Int. J. Comput. Vis. | 7 |
| 2025 | RehearMixup: Improving rehearsal-based continual learning
Yan Zhang 0145, Kaiyuan Qi, Guoqiang Wu, Yilong Yin |
Neurocomputing | 5 |
| 2025 | A noise-robust and generalizable framework for facial expression recognition
Qiangchang Wang, Jing Li 0175, Yilong Yin |
Inf. Sci. | 4 |
| 2025 | Diverse Information Aggregation with Adaptive Graph Construction and prompts for deepfake detection
Zhenhua Bai, Qiangchang Wang, Lu Yang 0005, Xinxin Zhang 0004, Yanbo Gao, Yilong Yin |
Image Vis. Comput. | 6 |
| 2025 | Active source-free open-set domain adaptation
Zhongyi Han, Hao Sun 0002, Yilong Yin |
Knowl. Based Syst. | 4 |
| 2025 | Private-library-oriented code generation with large language models
Daoguang Zan, Bei Chen 0008, Yongshun Gong, Junzhi Cao, Fengji Zhang, Bingchao Wu, Bei Guan, Yilong Yin, Yongji Wang 0002 |
Knowl. Based Syst. | 8 |
| 2025 | GeM: Gaussian embeddings with Multi-hop graph transfer for next POI recommendation
Wenqian Mu, Jiyuan Liu 0013, Yongshun Gong, Ji Zhong, Wei Liu 0007, Haoliang Sun, Xiushan Nie, Yilong Yin, Yu Zheng 0004 |
Neural Networks | 8 |
| 2025 | Diverse Teacher-Students for deep safe semi-supervised learning under class mismatch
Qikai Wang, Rundong He, Yongshun Gong, Chunxiao Ren, Haoliang Sun, Xiaoshui Huang, Yilong Yin |
Neural Networks | 7 |
| 2025 | QFAE: Q-Function guided Action Exploration for offline deep reinforcement learning
Teng Pang, Guoqiang Wu, Yan Zhang 0145, Bingzheng Wang, Yilong Yin |
Pattern Recognit. | 5 |
| 2025 | CTPT: Continual Test-time Prompt Tuning for vision-language models
Zhongyi Han, Xingbo Liu, Yilong Yin, Xin Gao 0001 |
Pattern Recognit. | 4 |
| 2025 | Consistency and label constrained transfer low-rank representation for cross-light finger vein recognition
Lu Yang 0005, Kuikui Wang, Xiaoming Xi, Xiushan Nie, Gongping Yang 0001, Yilong Yin |
Pattern Recognit. | 7 |
| 2025 | Adaptive division and priori reinforcement part learning network for vehicle re-identification
Xiaoying Zhou, Houren Zhou, Xiyu Pang, Jiachen Tian 0003, Xiushan Nie, Yilong Yin |
Pattern Recognit. | 8 |
| 2025 | Dual Difficulty-Aware Adaptive Pseudo Labeling for Semi-Supervised CNV SegmentationabstractIn clinical practice, obtaining a large amount of labeled CNV data is very difficult. Semi-supervised learning can effectively utilize a large amount of unlabeled CNV data. Since CNV has complex features such as blurred and unevenly distributed pixels on the edges, there are differences in the segmentation difficulty between pixels in the same image. Existing semi-supervised segmentation methods do not consider the segmentation difficulty of pixels, which will reduce the segmentation accuracy. To address this problem, we propose a dual difficulty-aware adaptive pseudo-label learning (D2APL) method for semi-supervised CNV segmentation. The proposed dual difficulty awareness includes segmentation difficulty perception of pixels in labeled and unlabeled data. For labeled data, we propose a classification confidence-guided difficulty perception method. For unlabeled data, we propose a model stability-guided difficulty perception method. Finally, we propose a difficulty-aware self-training method to dynamically adjust the threshold of pseudolabels according to the difficulty, thereby improving the utilization of difficult-to-segment pixels in unlabeled data. Experimental results show that our method outperforms the state-of-the-art method in CNV segmentation. Jie Guo 0012, Liangyun Sun, Lishan Qiao, Xiushan Nie, Jixin Yang, Weicui Li, Ying Guo 0030, Xiaoming Xi, Xinjian Chen 0001, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | LAC-PS: A Light Direction Selection Policy Under the Accuracy Constraint for Photometric StereoabstractPhotometric stereo (PS) methods recover surface normals from appearance changes under varying light directions, excelling in tasks like 3D surface reconstruction and defect inspection. However, collecting the illumination images is expensive, and current PS methods cannot obtain the light direction set that satisfies the pre-defined accuracy constraint, limiting their adaptability to various applications with varying accuracy requirements. To address this issue, we propose the LAC-PS, a light direction selection policy under the accuracy constraint for photometric stereo, which optimizes the light direction set to meet target reconstruction accuracy. In our method, we develop an accuracy assessment network that estimates reconstruction accuracy without ground truth. With this estimated accuracy, we put forward a reinforcement learning-based method that can utilize policy to sequentially select light directions and obtain the light directions satisfying the desired PS recovery accuracy constraint. Experimental results on real and synthetic datasets demonstrate that our method effectively selects light directions that satisfy accuracy constraints. Wenjia Meng, Huimin Han, Xiankai Lu, Yilong Yin, Gang Pan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MGRL4RE: A Multi-Graph Representation Learning Approach for Urban Region EmbeddingabstractUsing multi-modal data to learn region representations has gained popularity for its ability to reveal diverse socioeconomic features in cities. However, many studies focus solely on semantic features from points-of-interest (POIs), neglecting the issue of spatial imbalance. This article introduces a Multi-Graph Representation Learning framework for Region Embedding (MGRL4RE), which leverages both inter-region and intra-region correlations through two main components: multi-graph construction based on various region correlations and multi-graph representation learning. The construction module creates a multi-graph reflecting various correlations among regions, utilizing geo-tagged POIs, region data, and human mobility data. Specifically, we assess a region’s importance relative to its spatial context (neighborhood) and develop spatially invariant semantic features to address spatial imbalance. Furthermore, the representation learning module generates comprehensive and effective region representations via multi-view embedding fusion. Our extensive experiments across various downstream tasks, including land use clustering, region popularity prediction, and crime prediction, confirm that our model significantly outperforms existing state-of-the-art region embedding methods. Meng Chen 0003, Zechen Li 0003, Hongwei Jia, Min Yang 0006, Yilong Yin |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2025 | Off-OAB: Off-Policy Policy Gradient Method With Optimal Action-Dependent BaselineabstractThe policy-based methods have achieved remarkable success in solving challenging reinforcement learning (RL) problems. Among these methods, the off-policy policy gradient (OPPG) methods are particularly important because they can benefit from off-policy data. However, these methods suffer from the high variance of the OPPG estimator, which results in poor sample efficiency during training. In this article, we propose an off-policy policy gradient method with the optimal action-dependent baseline (Off-OAB) to mitigate this variance issue. Specifically, this baseline maintains the OPPG estimator's unbiasedness while theoretically minimizing its variance. To enhance practical computational efficiency, we design an approximated version of this optimal baseline. Utilizing this approximation, our method (Off-OAB) aims to decrease the OPPG estimator's variance during policy optimization. We evaluate the proposed Off-OAB method on six representative tasks from OpenAI Gym and MuJoCo, where it demonstrably surpasses the state-of-the-art methods on the majority of these tasks. Wenjia Meng, Long Yang 0004, Yilong Yin, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Exploring Channel-Aware Typical Features for Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) data is essential to ensure the reliability of machine learning models when deployed in real-world scenarios. Different from most previous test-time OOD detection methods that focus on designing OOD scores, we delve into the challenges in OOD detection from the perspective of typicality and regard the feature’s high-probability region as the feature’s typical set. However, the existing typical-feature-based OOD detection method implies an assumption: the proportion of typical feature sets for each channel is fixed. According to our experimental analysis, each channel contributes differently to OOD detection. Adopting a fixed proportion for all channels results in several channels losing too many typical features or incorporating too many abnormal features, resulting in low performance. Therefore, exploring the channel-aware typical features is crucial to better-separating ID and OOD data. Driven by this insight, we propose expLoring channel-Aware tyPical featureS (LAPS). Firstly, LAPS obtains the channel-aware typical set by calibrating the channel-level typical set with the global typical set from the mean and standard deviation. Then, LAPS rectifies the features into channel-aware typical sets to obtain channel-aware typical features. Finally, LAPS leverages the channel-aware typical features to calculate the energy score for OOD detection. Theoretical and visual analyses verify that LAPS achieves a better bias-variance trade-off. Experiments verify the effectiveness and generalization of LAPS under different architectures and OOD scores. Rundong He, Zhongyi Han, Wan Su, Yilong Yin, Tongliang Liu, Yongshun Gong |
AAAI | 6 |
| 2024 | DiffAIL: Diffusion Adversarial Imitation LearningabstractImitation learning aims to solve the problem of defining reward functions in real-world decision-making tasks. The current popular approach is the Adversarial Imitation Learning (AIL) framework, which matches expert state-action occupancy measures to obtain a surrogate reward for forward reinforcement learning. However, the traditional discriminator is a simple binary classifier and doesn't learn an accurate distribution, which may result in failing to identify expert-level state-action pairs induced by the policy interacting with the environment. To address this issue, we propose a method named diffusion adversarial imitation learning (DiffAIL), which introduces the diffusion model into the AIL framework. Specifically, DiffAIL models the state-action pairs as unconditional diffusion models and uses diffusion loss as part of the discriminator's learning objective, which enables the discriminator to capture better expert demonstrations and improve generalization. Experimentally, the results show that our method achieves state-of-the-art performance and significantly surpasses expert demonstration on two benchmark tasks, including the standard state-action setting and state-only settings. Bingzheng Wang, Guoqiang Wu, Teng Pang, Yan Zhang 0145, Yilong Yin |
AAAI | 5 |
| 2024 | Navigating the Unknown: A Novel MGUAN Framework for Medical Image Recognition Across Dynamic DomainsabstractMachine learning has significantly advanced medical image recognition, enhancing diagnostic accuracy in various applications. However, these advancements primarily apply to scenarios with consistent data distributions, a condition rarely met in real-world clinical settings. In real-world clinical environments, variations in device specifications and patient demographics introduce distribution shifts, class imbalance and unknown class challenges, undermining model robustness. Addressing this, we present the Medical image recognition under Generalized Universal Domain Adaptation (MGUDA) concept, targeting distribution shifts, class imbalance and unknown class detection. Our innovative Medical Dual-Prototype Adaptation Network (MDPAN) framework, integrating dual prototype learning, dual prototype employment, and weighted multi-class adversarial alignment, adeptly confronts these issues. Extensive evaluations on diverse medical image datasets validate MDPAN’s superiority in managing class imbalances and enhancing target domain classification, marking a pivotal step in robust medical image recognition across variable domains. Wan Su, Rundong He, Zhongyi Han, Yilong Yin |
BIBM | 5 |
| 2024 | Discriminability-Driven Channel Selection for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is essential for deploying machine learning models in open-world environments. Activation-based methods are a key approach in OOD detection, working to mitigate overconfident predictions of OOD data. These techniques rectifying anomalous activations, enhancing the distinguishability between in-distribution (ID) data and OOD data. However, they assume by default that every channel is necessary for OOD detection, and rectify anomalous activations in each channel. Empirical evidence has shown that there is a significant difference among various channels in OOD detection, and discarding some channels can greatly enhance the performance of OOD detection. Based on this insight, we propose Discriminability-Driven Channel Selection (DDCS), which leverages an adaptive channel selection by estimating the discriminative score of each channel to boost OOD detection. The discriminative score takes inter-class similarity and inter-class variance of training data into account. However, the estimation of discriminative score itself is susceptible to anomalous activations. To better estimate score, we pre-rectify anomalous activations for each channel mildly. The experimental results show that DDCS achieves state-of-the-art performance on CIFAR and ImageNet-1K benchmarks. Moreover, DDCS can generalize to different backbones and OOD scores. Rundong He, Yicong Dong, Zhongyi Han, Yilong Yin |
CVPR | 5 |
| 2024 | Distribution-Aware Contrastive Learning for Robust Medical Image SegmentationabstractMedical image segmentation is pivotal in quantifying tissue volumes, facilitating diagnoses, and enabling other critical medical applications. However, accurately segmenting medical images can be challenging because the complex intensity distribution inherent in the data arises from the highly complex interaction of many latent factors (data heterogeneity). In this context, we propose a novel method called Distribution-aware Contrastive Learning for Robust Segmentation (DCL-Seg) to address the inconsistency in medical image segmentation. Based on the assumption of content separability, we use learnable parameters to construct positive samples with a potential structure invariance via contrastive learning. In this way, our method can mitigate the negative effects of data heterogeneity to separate overlapped class distribution and structural solid boundary. We are in one public dataset and two clinical datasets for Breast tumor and Retinal vessel segmentation, which have achieved excellent results and widely proved the superiority of our method. Zheyun Qin, Xiaoming Xi, Yilong Yin |
ICASSP | 3 |
| 2024 | Two-phase Parametric Registration for Retinal ImagesabstractWe propose a two-phase parametric registration algorithm for retinal images. Our algorithm focuses on dealing with the geometric transformation and the intensity transformation in the retinal image registration problem. In the first phase, we efficiently detect only one pair of feature points in the source and the target retinal images to estimate a translation transformation and get a warped source image. In the second phase, we estimate both the intensity and the geometric transformations between the target image and the warped source image by fitting parametric expressions. The displacement field is generated by a super fast and accurate coarse-to-fine elastic registration algorithm—local all-pass filters algorithm (LAP). At each iteration of the LAP, the elastic displacement field and the intensity difference take turns being fitted by two different low-order polynomial functions. The fitting steps are performed by solving linear systems of equations efficiently. Experiments on real retinal image datasets demonstrated the high accuracy and computational efficiency of the proposed retinal image registration method. Xinxin Zhang 0004, Xiankai Lu, Jizhou Li, Yongshun Gong, Qiangchang Wang, Yilong Yin |
ICME | 6 |
| 2024 | Unsupervised Online Cross-modal Hashing With Multiple Association ExploitationabstractUnsupervised online cross-modal hashing has gained increasing attention for its effectiveness in streaming data retrieval. However, existing methods primarily focus on exploiting shared properties, overlooking semantic shifts among chunks and specific properties of each modality. To address these challenges, we propose a novel method called Unsupervised Online Cross-Modal Hashing with multiple association exploitation, UOCMH in short. Specifically, we design a hierarchical matrix factorization framework. It skillfully constructs robust orthogonal bases, multi-modality specific representations, and unified common representations, thereby capturing semantic associations among multi-modality streaming data more sufficiently. Additionally, we present a semantic auto-encoder scheme as hash functions. It builds the association between features and hash codes, facilitating the stability of the hashing process. Extensive experiments on the widely-used benchmark datasets demonstrate the superiority of the proposed UOCMH. Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin |
ICME | 7 |
| 2024 | Fast Multi-view Clustering With Binary Anchor GraphabstractMulti-view clustering has achieved remarkable efficacy in integrating multi-view information, and received much research interest. Although anchor-based clustering algorithms have been well-investigated in past years, the separation of graph construction and category partitioning, can lead to suboptimal clustering performance and learning efficiency. To address these challenges, we propose a novel fast clustering algorithm named FAST-BAG. The proposed method can integrate the anchor graph construction and clustering partitioning seamlessly, breaking the separation between data fusion and task processes. Specifically, the multi-view data is unified into a consistent binary anchor graph with linear time complexity. Additionally, we leverage the high efficiency of binary distance computation to expedite the category partitioning process. Experiments conducted on five benchmark datasets validate the effectiveness and efficiency of the proposed method. Xingbo Liu, Xiao Kang, Xuening Zhang, Xiushan Nie, Yilong Yin |
ICME | 7 |
| 2024 | Completely Unpaired Cross-Modal Hashing Based on Coupled SubspaceabstractUnpaired cross-modal hashing which requires no supervision is a promising candidate to support large-scale retrieval across heterogeneous data. However, existing works focus on recovering pairwise relationships, which are usually time-consuming and sensitive to outliers. To tackle this issue, we propose a novel method termed Completely Unpaired Crossmodal Hashing (CUCH), which is applicable to scenarios where neither pairwise correspondence nor label information is available. The proposed CUCH creatively combines the merits of subspace recovery and cross-modal hashing, producing an effective subspace with both robustness and high efficiency. It first discovers robust subspace from each modality by excluding outliers. Then latent space translation is elaborated to obtain coupled subspace, based on which intermodal similarities can be captured. Moreover, the similarity-preservation property for CUCH is guaranteed. By manipulating subspaces rather than pairwise relations, CUCH reduces computational cost significantly. Experimental results demonstrate its advantages in various settings. Xuening Zhang, Xingbo Liu, Xiao Kang, Xiushan Nie, Yilong Yin |
ICME | 7 |
| 2024 | Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial TopologyabstractWith the expansion and growth of cities, profiling urban areas with the advent of multi-modal urban datasets (e.g., points-of-interest and street view imagery) has become increasingly important in urban planing and management. Particularly, street view images have gained popularity for understanding the characteristics of urban areas due to its abundant visual information and inherent correlations with human activities. In this study, we define a street segment represented by multiple street view images as the minimum spatial unit for analysis and predict its functional and socioeconomic indicators, which presents several challenges in modeling spatial distributions of images on a street and the spatial topology (adjacency) of streets. Meanwhile, Large Language Models are capable of understanding imagery data based on its extraordinary knowledge base and unveil a remarkable opportunity for profiling streets with images. In view of the challenges and opportunity, we present a semi-supervised Urban Street Profiling Model (USPM) based on street view imagery and spatial adjacency of urban streets. Specifically, given a street with multiple images, we first employ a newly designed spatial context-based contrastive learning method to generate feature vectors of images and then apply the LSTM-based fusion method to encode multiple images on a street to yield the street visual representation; we then create the descriptions of street scenes for street view images based on the SPHINX (a large language model) and produce the street textual representation; finally, we build an urban street graph based on spatial topology (adjacency) and employ a semi-supervised graph learning algorithm to further encode the street representations for prediction. We conduct thorough experiments with real-world datasets to assess the proposed USPM. The experimental results demonstrate that USPM considerably outperforms baseline methods in two urban prediction tasks. Meng Chen 0003, Zechen Li 0003, Weiming Huang 0001, Yongshun Gong, Yilong Yin |
KDD | 5 |
| 2024 | Hypergraph-guided Intra- and Inter-category Relation Modeling for Fine-grained Visual RecognitionabstractFine-grained Visual Recognition (FGVR) aims to distinguish objects within similar subcategories. Humans adeptly perform this challenging task by leveraging both intra-category distinctiveness and inter-category similarity. However, previous methods fail to combine these two complementary dimensions and mine the intrinsic relations among various semantic features. To address these limitations, we propose HI2R, a Hypergraph-guided Intra- and Inter-category Relation Modeling approach, which simultaneously extracts the intra-category structural information and inter-category relation for more precise reasoning. Specifically, we exploit a Hypergraph-guided Structure Learning (HSL) module, which employs hypergraphs to capture high-order structural relations, transcending traditional graph-based methods that are limited to pairwise linkages. This advancement allows the model to adapt to significant intra-category variations. Additionally, we propose an Inter-category Relation Perception (IRP) module to improve feature discrimination across categories by extracting and analyzing semantic relations among them. Our objective is to alleviate the robustness issue associated with exclusive reliance on intra-category discriminative features. Furthermore, a random semantic consistency (RSC) loss is introduced to direct the model's attention to commonly overlooked yet distinctive regions, indirectly enhancing the representation ability of both HSL and IRP modules. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of HI2R. Qiangchang Wang, Yilong Yin |
ACM Multimedia | 4 |
| 2024 | KNN Transformer with Pyramid Prompts for Few-Shot LearningabstractFew-Shot Learning (FSL) aims to recognize new classes with limited labeled data. Recent studies have attempted to address the challenge of rare samples with textual prompts to modulate visual features. However, they usually struggle to capture complex semantic relationships between textual and visual features. Moreover, vanilla self-attention is heavily affected by useless information in images, severely constraining the potential of semantic priors in FSL due to the confusion of numerous irrelevant tokens during interaction. To address these aforementioned issues, a K-NN Transformer with Pyramid Prompts (KTPP) is proposed to select discriminative information with K-NN Context Attention (KCA) and adaptively modulate visual features with Pyramid Cross-modal Prompts (PCP). First, for each token, the KCA only selects the K most relevant tokens to compute the self-attention matrix and incorporates the mean of all tokens as the context prompt to provide the global context in three cascaded stages. As a result, irrelevant tokens can be progressively suppressed. Secondly, pyramid prompts are introduced in the PCP to emphasize visual features via interactions between text-based class-aware prompts and multi-scale visual features. This allows the ViT to dynamically adjust the importance weights of visual features based on rich semantic information at different scales, making models robust to spatial variations. Finally, augmented visual features and class-aware prompts are interacted via the KCA to extract class-specific features. Consequently, our model further enhances noise-free visual representations via deep cross-modal interactions, extracting generalized visual representation in scenarios with few labeled samples. Extensive experiments on four benchmark datasets demonstrate significant gains over the state-of-the-art methods, especially for the 1-shot task with 2.28% improvement on average due to semantically enhanced visual representations. Wenhao Li 0011, Qiangchang Wang, Peng Zhao 0016, Yilong Yin |
ACM Multimedia | 4 |
| 2024 | Dual-track spatio-temporal learning for urban flow prediction with adaptive normalization
Yongshun Gong, Wei Liu 0007, Yilong Yin, Yu Zheng 0004, Liqiang Nie |
Artif. Intell. | 4 |
| 2024 | Visual Out-of-Distribution Detection in Open-Set Noisy Environments
Rundong He, Zhongyi Han, Xiushan Nie, Yilong Yin, Xiaojun Chang |
Int. J. Comput. Vis. | 4 |
| 2024 | Multi-axis interactive multidimensional attention network for vehicle re-identification
Xiyu Pang, Yanli Zheng, Xiushan Nie, Yilong Yin |
Image Vis. Comput. | 4 |
| 2024 | Cascaded Cross-modal Alignment for Visible-Infrared Person Re-Identification
Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
Knowl. Based Syst. | 5 |
| 2024 | Generalized Universal Domain Adaptation
Wan Su, Zhongyi Han, Xingbo Liu, Yilong Yin |
Knowl. Based Syst. | 4 |
| 2024 | BIAS: Bridging Inactive and Active Samples for active source free domain adaptation
Zhongyi Han, Yilong Yin |
Knowl. Based Syst. | 3 |
| 2024 | Learning sample-aware threshold for semi-supervised learning
Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Rundong He, Yilong Yin |
Mach. Learn. | 6 |
| 2024 | Correction: Learning sample-aware threshold for semi-supervised learning
Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Rundong He, Yilong Yin |
Mach. Learn. | 6 |
| 2024 | A survey of micro-video analysis
Jie Guo 0012, Yuling Ma, Meng Liu 0006, Xiaoming Xi, Xiushan Nie, Yilong Yin |
Multim. Tools Appl. | 7 |
| 2024 | Heterogeneous context interaction network for vehicle re-identification
Xiyu Pang, Meifeng Zheng, Xiushan Nie, Houren Zhou, Yilong Yin |
Neural Networks | 7 |
| 2024 | MetaKernel: Learning Variational Random Features With Limited LabelsabstractFew-shot learning deals with the fundamental and challenging problem of learning from a few annotated samples, while being able to generalize well on new tasks. The crux of few-shot learning is to extract prior knowledge from related tasks to enable fast adaptation to a new task with a limited amount of data. In this paper, we propose meta-learning kernels with random Fourier features for few-shot learning, we call MetaKernel. Specifically, we propose learning variational random features in a data-driven manner to obtain task-specific kernels by leveraging the shared knowledge provided by related tasks in a meta-learning setting. We treat the random feature basis as the latent variable, which is estimated by variational inference. The shared knowledge from related tasks is incorporated into a context inference of the posterior, which we achieve via a long-short term memory module. To establish more expressive kernels, we deploy conditional normalizing flows based on coupling layers to achieve a richer posterior distribution over random Fourier bases. The resultant kernels are more informative and discriminative, which further improves the few-shot learning. To evaluate our method, we conduct extensive experiments on both few-shot image classification and regression tasks. A thorough ablation study demonstrates that the effectiveness of each introduced component in our method. The benchmark results on fourteen datasets demonstrate MetaKernel consistently delivers at least comparable and often better performance than state-of-the-art alternatives. Yingjun Du, Haoliang Sun, Xiantong Zhen, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Discrete online cross-modal hashing with consistency preservation
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin |
Pattern Recognit. | 6 |
| 2024 | Discriminative atoms embedding relation dual network for classification of choroidal neovascularization in OCT images
Xiaoming Xi, Longsheng Xu, Xiushan Nie, Jianhua Nie, Xianjing Meng, Xinjian Chen 0001, Yilong Yin |
Pattern Recognit. | 10 |
| 2024 | Scalable Unsupervised Hashing via Exploiting Robust Cross-Modal ConsistencyabstractUnsupervised cross-modal hashing has received increasing attention because of its efficiency and scalability for large-scale data retrieval and analysis. However, existing unsupervised cross-modal hashing methods primarily focus on learning shared feature embedding, ignoring robustness and consistency across different modalities. To this end, this study proposes a novel method called scalable unsupervised hashing (SUH) for large-scale cross-modal retrieval. In the proposed method, latent semantic information and common semantic embedding within heterogeneous data are simultaneously exploited using multimodal clustering and collective matrix factorization, respectively. Furthermore, the robust norm is seamlessly integrated into the two processes, making SUH insensitive to outliers. Based on the robust consistency exploited from the latent semantic information and feature embedding, hash codes can be learned discretely to avoid cumulative quantitation loss. The experimental results on five benchmark datasets demonstrate the effectiveness of the proposed method under various scenarios. Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin |
IEEE Trans. Big Data | 6 |
| 2024 | Online Discriminative Cross-Modal HashingabstractOnline cross-modal hashing has received increasing research attention due to its capability of encoding streaming data and updating hash functions simultaneously. Despite significant progress, there is still room for further improving accuracy from two aspects,i.e., 1) enhancing discrimination of hash codes with an efficient training process; 2) elevating generalization performance by harmonizing the training and retrieval process. Inspired by this, we propose an Online Discriminative Cross-modal Hashing method, called ODCH. To enlarge the inter-class margin and magnify the intra-class similarity, ODCH skillfully constructs a discriminative semantic space and seamlessly integrates bit balance and uncorrelation constraints, discrete optimization, and asymmetric strategy for embedding the discriminative semantic information into hamming space. Furthermore, ODCH attempts to boost the generalization process by bridging the gap between learning and generalization. It develops adaptive bit-wise weights to reflect different learning conditions among bits and transmits them into the generalization process. Besides, the proposed discriminative embedding and adaptive weighting can be adopted by existing supervised cross-modal hashing methods, achieving more precise performance than the original versions. Extensive experiments on three benchmarked datasets show that ODCH achieves up to an average of 4.17% mAP score gains compared to state-of-the-art online cross-modal hashing methods, indicating its superiority. Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Semi-Supervised Semi-Paired Cross-Modal HashingabstractLarge-scale cross-modal hashing has drawn extensive attention due to its attractive efficiency in both storage and retrieval. Existing methods exhibit poor performance when exploiting the semantic correlations implied in unsupervised and unpaired data during training process. To deal with this issue, we propose a novel hashing method, named Semi-supervised Semi-paired Cross-modal Hashing (SSCH). By leveraging a general and flexible two-step scheme, the proposed method can handle the complex training data effectively and efficiently, where both the common semantics and the modality-specific optimal pseudo semantics are well captured. Specifically, the proposed SSCH performs an alignment-free pseudo-labeling process to get strengthened semantic information. Furthermore, hash representations for various data are learned via a label-enhanced strategy, through which the cross-modal correlations are strengthened and preserved with considering efficiency. The semantic-preserving proof of SSCH is given based on statistical analysis. Also, we prove the stability of the proposed time-saving algorithm using properties of Bregman divergence. Experimental results on three benchmark datasets show that SSCH can obtain satisfactory precision and scalability in various scenarios. Xuening Zhang, Xingbo Liu, Xiushan Nie, Xiao Kang, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Video Corpus Moment Retrieval via Deformable Multigranularity Feature Fusion and Adversarial TrainingabstractAs a new emerging task, video corpus moment retrieval (VCMR) aims to find the video segments relevant to a given natural language query from a large number of untrimmed videos. It mainly includes two subtasks, finding the most relevant video based on the query text (video retrieval), and locating the segment most relevant to a given query in a video (moment localization). At the same time, since videos often contain rich multi-modal information such as audio, text, and images, how to align and interact with the multi-modal information of videos and the text information of natural language queries across modalities is the core issue of this task. This article proposes a Deformable Multigranularity Feature Fusion with Adversarial Training Network (DMFAT), first inputs the subtitle and frame multi-modal information of the video into our Multi-Scale Deformable Attention module and performs multi-granularity feature fusion through Deformable Attention respectively. Then, guided by the query, adaptive weights are generated to fuse the two multi-granularity modality features of the video. Finally, the cross-modal representation of the query and video features is obtained through a bidirectional attention module, and an adversarial contrastive learning objective is introduced to enhance more precise moment localization. Our model is evaluated on two representative video corpus moment retrieval benchmarks: TVR and DiDeMo. Extensive experiments have been conducted to demonstrate that our method outperforms existing work. Peng Zhao 0016, Jinsheng Ji, Xiankai Lu, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Characterizing Hierarchical Semantic-Aware Parts With Transformers for Generalized Zero-Shot LearningabstractThis paper presents a novel Transformer architecture for zero-shot learning (ZSL), termed TransZSL, which can characterize hierarchical semantic-aware parts. It consists of an adaptive token refinement (ATR), a hierarchical token aggregation (HTA), and semantic-aware prototypes (SAP). Firstly, the ViT is used as the backbone that provides comprehensive local information without missing details. To address the different degrees of noise caused by large appearance variations, the ATR is proposed to highlight important tokens and suppress useless ones adaptively. However, due to the complex image structure, some important tokens may be incorrectly discarded. Therefore, a random perturbation is proposed to reactivate discarded tokens randomly, reducing the risk of missing discriminative information. Secondly, dataset descriptions contain both low- and high-level attributes. To this end, the HTA aggregates complementary hierarchical tokens from multiple ViT layers. Thirdly, semantically similar content may be distributed in different tokens. To overcome this issue, the SAP is proposed to group semantically identical tokens into one prototype, focusing on semantic-aware parts. Besides, diversity loss is used to encourage networks to learn diverse prototypes that discover diverse parts. Both qualitative and quantitative results on several challenging tasks demonstrate the usefulness and effectiveness of our proposed methods. Peng Zhao 0016, Xiaoming Xi, Qiangchang Wang, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Biomarkers-Aware Asymmetric Bibranch GAN With Adaptive Memory Batch Normalization for Prediction of Anti-VEGF Treatment Response in Neovascular Age-Related Macular DegenerationabstractThe emergence of anti-vascular endothelial growth factor (anti-VEGF) therapy has revolutionized neovascular age-related macular degeneration (nAMD). Post-therapeutic optical coherence tomography (OCT) imaging facilitates the prediction of therapeutic response to anti-VEGF therapy for nAMD. Although the generative adversarial network (GAN) is a popular generative model for post-therapeutic OCT image generation, it is realistically challenging to gather sufficient pre- and post-therapeutic OCT image pairs, resulting in overfitting. Moreover, the available GAN-based methods ignore local details, such as the biomarkers that are essential for nAMD treatment. To address these issues, a Biomarkers-aware Asymmetric Bibranch GAN (BAABGAN) is proposed to efficiently generate post-therapeutic OCT images. Specifically, one branch is developed to learn prior knowledge with a high degree of transferability from large-scale data, termed the source branch. Then, the source branch transfer knowledge to another branch, which is trained on small-scale paired data, termed the target branch. To boost the transferability, a novel Adaptive Memory Batch Normalization (AMBN) is introduced in the source branch, which learns more effective global knowledge that is impervious to noise via memory mechanism. Also, a novel Adaptive Biomarkers-aware Attention (ABA) module is proposed to encode biomarkers information into latent features of target branches to learn finer local details of biomarkers. The proposed method outperforms traditional GAN models and can produce high-quality post-treatment OCT pictures with limited data sets, as shown by the results of experiments. Peng Zhao 0016, Xian Song, Xiaoming Xi, Xiushan Nie, Xianjing Meng, Yilong Yin |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Spatio-Temporal Enhanced Contrastive and Contextual Learning for Weather ForecastingabstractWeather forecasting is of great importance for human life and various real-world fields, e.g., traffic prediction, agricultural production, and tourist industry. Existing methods can be roughly divided into two categories: theory-driven (e.g., numerical weather prediction (NWP)) and data-driven methods. Theory-driven methods require a complex simulation of the physical evolution process in the atmosphere model using supercomputers, while most data-driven methods learn the underlying laws from the historical weather records via deep learning models. However, some data-driven methods simply regard all weather variables of monitoring stations as a whole and fail to more granularly exploit complex correlations across different stations, while others prefer to construct large neural networks with massive learnable parameters. To alleviate these defects, we propose a spatio-temporal contrastive self-supervision method and a generative contextual self-supervised technique to capture spatial and temporal dependencies from the station-level and variable-level, respectively. Through these well-designed self-supervised tasks, uncomplicated networks obtain strong capability to capture latent representations for weather changes with time-varying. Thereafter, an effective encoder-decoder based fine-tuning framework is proposed, consisting of three self-supervised encoders. Extensive experiments conducted on four real-world weather condition datasets demonstrate that our method outperforms the state-of-the-art models and also empirically validates the feasibility of each self-supervised task. Yongshun Gong, Tiantian He 0004, Meng Chen 0003, Bin Wang 0045, Liqiang Nie, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | SAFER-STUDENT for Safe Deep Semi-Supervised Learning With Unseen-Class Unlabeled DataabstractDeep semi-supervised learning (SSL) methods aim to utilize abundant unlabeled data to improve the seen-class classification. However, in the open-world scenario, collected unlabeled data tend to contain unseen-class data, which would degrade the generalization to seen-class classification. Formally, we define the problem as safe deep semi-supervised learning with unseen-class unlabeled data. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the lack of suitable score function, the uncalibrated model, and the small number of labeled data. To this end, we propose a safe SSL method called SAFER-STUDENT from the teacher-student view. First, to enhance the ability of teacher model to identify seen and unseen classes, we propose a general scoring framework calledDiscrepancy withRaw (DR). Second, based on unseen-class data mined by teacher model from unlabeled data, we calibrate student model by newly proposedUnseen-classEnergy-boundedCalibration (UEC) loss. Third, based on seen-class data mined by teacher model from unlabeled data, we proposeWeightedConfirmationBiasElimination (WCBE) loss to boost seen-class classification of student model. Extensive studies show that SAFER-STUDENT remarkably outperforms the state-of-the-art, verifying the effectiveness of our method in the under-explored problem. Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo SelectionabstractNowadays, capturing cherished moments results in an abundance of photos, which necessitates the selection of the finest one from a pool of akin images—a process both intricate and time-intensive. Thus, series photo selection (SPS) techniques have been developed to recommend the optimal moment from nearly identical photos through the use of aesthetic quality assessment. However, addressing SPS proves demanding due to the subtle nuances within such imagery. Existing approaches predominantly rely on diverse feature types (e.g., color, layout, generic features) extracted from original images to discern the qualified shot, yet they disregard disentangling generality and specificity at the feature level. This study aims to detect subtle aesthetic distinctions among akin photos. We propose a feature separation model that captures all label-relevant information through an encoder. We introduce Information Bottleneck (IB) learning to obtain non-redundant representations of image pairs and filter out noise information from the representations. Our model segregates image features into shared and specific attributes by employing feature constraints to boost mutual information across images and guide meaningful information within individual images. This process filters out extraneous data within individual images, thus significantly enhancing the representation of similar image pairs. Extensive experiments on the Phototriage dataset show that our model can accentuate subtle disparities and achieve superior results when compared to alternative methods. Yongshun Gong, Xinxin Zhang 0004, Jian Zhang 0002, Yilong Yin |
IEEE Trans. Multim. | 6 |
| 2024 | Modeling Multiple Aesthetic Views for Series Photo SelectionabstractNumerous photos are taken in daily life, and sorting them is laborious and time consuming. The large number of similar images exacerbates the difficulty of album management, under this scenario, serial photo selection (SPS) emerges. As an important branch of image aesthetic quality assessment, it focuses on identifying the best image among a series of almost identical photos. Currently, most existing SPS methods focus only on extracting features from the original image, while neglecting the fact that multiple views of the image can provide much more detailed aesthetic information. In this article, we propose a Siamese network structure called SPSNet to enhance the representation learning of multi-view features by acquiring the depth, generic, and handcrafted features of images. In specific, we implement a parallel structure to extract deep and shallow features, fusing local and global representations at different resolutions interactively. The aggregation of multiple views of image via a self-attentive module with adaptive weights enables the model to discriminate the importance of each view. Moreover, we employ a graph neural network to construct the relationships among the multi-view features. Our proposed method, which is trained by a Siamese network, can effectively distinguish the nuances of similar images, and thus, select the best one from a series of almost identical photos. Extensive experiments conducted on the aesthetic dataset demonstrate that our method outperforms other state-of-the-art SPS methods, which achieves the 75.36% accuracy on the Phototriage dataset. Besides, our model is up to 3.04% better than the baseline methods in terms of the average accuracy. Yongshun Gong, Lu Zhang 0062, Jian Zhang 0002, Liqiang Nie, Yilong Yin |
IEEE Trans. Multim. | 6 |
| 2024 | Learning Feature Semantic Matching for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks. Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 7 |
| 2024 | Relational Network via Cascade CRF for Video Language GroundingabstractVideo Language Grounding is one of the most challenging cross-modal video understanding tasks. This task aims to localize a target moment semantically corresponding to a given language query in an untrimmed video. Many existing VLG methods rely on the proposal-based framework, despite the dominant performance achieved, they usually focus on interacting a few internal frames with the query to score segment proposals, trapping in the long-range dependencies when the proposal feature is limited. Meanwhile, adjacent proposals share similar visual semantics, making VLG models hard to align the accurate semantics of video-query contents and degenerating the ranking performance. To remedy the above limitations, we propose VLG-CRF by introducing the conditional random fields (CRFs) to handle the discrete yet indistinguishable proposals. Specifically, VLG-CRF consists of two cascade CRF-based modules. The AttentiveCRFs is developed for multi-modal feature fusion to better integrate temporal and semantic relation between modalities. We also devise a new variant of ConvCRFs to capture the relation of discrete segments and rectify the predicting scores to make relatively high prediction scores clustered in a range. Experiments on three benchmark datasets,i.e., Charades-STA, ActivityNet-Caption, and TACoS, show the superiority of our method and the state-of-the-art performance is achieved. Xiankai Lu, Hao Zhang 0048, Xiushan Nie, Yilong Yin, Jianbing Shen |
IEEE Trans. Multim. | 5 |
| 2024 | Tri-Branch Convolutional Neural Networks for Top-k Focused Academic Performance PredictionabstractAcademic performance prediction aims to leverage student-related information to predict their future academic outcomes, which is beneficial to numerous educational applications, such as personalized teaching and academic early warning. In this article, we reveal the students' behavior trajectories by mining campus smartcard records, and capture the characteristics inherent in trajectories for academic performance prediction. Particularly, we carefully design a tri-branch convolutional neural network (CNN) architecture, which is equipped with rowwise, columnwise, and depthwise convolutions and attention operations, to effectively capture the persistence, regularity, and temporal distribution of student behavior in an end-to-end manner, respectively. However, different from existing works mainly targeting at improving the prediction performance for the whole students, we propose to cast academic performance prediction as a top-k ranking problem, and introduce a top-k focused loss to ensure the accuracy of identifying academically at-risk students. Extensive experiments were carried out on a large-scale real-world dataset, and we show that our approach substantially outperforms recently proposed methods for academic performance prediction. For the sake of reproducibility, our codes have been released at https://github.com/ZongJ1111/Academic-Performance-Prediction. Chaoran Cui, Jian Zong, Yuling Ma, Xinhua Wang 0003, Lei Guo 0008, Meng Chen 0003, Yilong Yin |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Online Cross-modal Hashing With Dynamic PrototypeabstractOnline cross-modal hashing has received increasing attention due to its efficiency and effectiveness in handling cross-modal streaming data retrieval. Despite the promising performance, these methods mainly focus on the supervised learning paradigm, demanding expensive and laborious work to obtain clean annotated data. Existing unsupervised online hashing methods mostly struggle to construct instructive semantic correlations among data chunks, resulting in the forgetting of accumulated data distribution. To this end, we propose a Dynamic Prototype-based Online Cross-modal Hashing method, called DPOCH. Based on the pre-learned reliable common representations, DPOCH generates prototypes incrementally as sketches of accumulated data and updates them dynamically for adapting streaming data. Thereafter, the prototype-based semantic embedding and similarity graphs are designed to promote stability and generalization of the hashing process, thereby obtaining globally adaptive hash codes and hash functions. Experimental results on benchmarked datasets demonstrate that the proposed DPOCH outperforms state-of-the-art unsupervised online cross-modal hashing methods. Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Fast Unsupervised Cross-Modal Hashing with Robust Factorization and Dual ProjectionabstractUnsupervised hashing has attracted extensive attention in effectively and efficiently tackling large-scale cross-modal retrieval task. Existing methods typically try to mine the latent common subspace across multimodal data without any category annotation. Despite the exciting progress, there are still three challenges that need to be further addressed: (1) efficiently improving the robustness during latent common subspace learning; (2) harmoniously embedding the intra-modal inherence and inter-modal relevance of multimodal data into Hamming space; and (3) effectively reducing the training time complexity and making the model scalable for large-scale datasets. To well address the above challenges, this study proposes a method named Fast Unsupervised Cross-Modal Hashing (FUCH). Specifically, FUCH proposes a semantic-aware collective matrix factorization to learn robust representation via exploiting latent category-specific attributes, and introduces Cauchy loss to measure the factorization process. Accordingly, the above process can effectively embed potential discriminative information into common space, while making the model insensitive for outliers. Moreover, FUCH designs a dual projection learning scheme, which not only learns modality-unique hash functions to excavate individual properties but also learns modality-mutual hash functions to multimodal correlational properties. Experimental results on three benchmark datasets verify the effectiveness of FUCH under various scenarios. Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Complex Scenario Image Retrieval via Deep Similarity-aware HashingabstractWhen performing hashing-based image retrieval, it is difficult to learn discriminative hash codes especially for the multi-label, zero-shot and fine-grained settings. This is due to the fact that the similarities vary, even within the same category, under the conditions of complex scenario settings. To address this problem, this study develops a deep similarity-aware hashing method for complex scenario image retrieval (DEPISH). DEPISH more focuses on the samples that are difficult to distinguish from other images (i.e., “difficult samples”), such as images that contain multiple semantics. It dynamically divides attention among samples according to their difficulty levels with a margin weighting strategy. Furthermore, by adding special terms in the model, DEPISH is capable of avoiding the inconsistency between the hash code representation and true similarity among negative samples. In addition, unlike the existing methods that use a pre-defined similarity matrix with fixed values, the DEPISH adopts an adaptive similarity matrix, which accurately captures the various similarities among all samples. The results of our experiment on multiple benchmark datasets containing complex scenarios (i.e., multi-label, zero-shot, and fine-grained datasets) verify the effectiveness of this method. Xiushan Nie, Weili Guan, Yilong Yin |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Discriminability and Transferability Estimation: A Bayesian Source Importance Estimation Approach for Multi-Source-Free Domain AdaptationabstractSource free domain adaptation (SFDA) transfers a single-source model to the unlabeled target domain without accessing the source data. With the intelligence development of various fields, a zoo of source models is more commonly available, arising in a new setting called multi-source-free domain adaptation (MSFDA). We find that the critical inborn challenge of MSFDA is how to estimate the importance (contribution) of each source model. In this paper, we shed new Bayesian light on the fact that the posterior probability of source importance connects to discriminability and transferability. We propose Discriminability And Transferability Estimation (DATE), a universal solution for source importance estimation. Specifically, a proxy discriminability perception module equips with habitat uncertainty and density to evaluate each sample's surrounding environment. A source-similarity transferability perception module quantifies the data distribution similarity and encourages the transferability to be reasonably distributed with a domain diversity loss. Extensive experiments show that DATE can precisely and objectively estimate the source importance and outperform prior arts by non-trivial margins. Moreover, experiments demonstrate that DATE can take the most popular SFDA networks as backbones and make them become advanced MSFDA solutions. Zhongyi Han, Zhiyan Zhang, Rundong He, Wan Su, Xiaoming Xi, Yilong Yin |
AAAI | 7 |
| 2023 | Off-Policy Proximal Policy OptimizationabstractProximal Policy Optimization (PPO) is an important reinforcement learning method, which has achieved great success in sequential decision-making problems. However, PPO faces the issue of sample inefficiency, which is due to the PPO cannot make use of off-policy data. In this paper, we propose an Off-Policy Proximal Policy Optimization method (Off-Policy PPO) that improves the sample efficiency of PPO by utilizing off-policy data. Specifically, we first propose a clipped surrogate objective function that can utilize off-policy data and avoid excessively large policy updates. Next, we theoretically clarify the stability of the optimization process of the proposed surrogate objective by demonstrating the degree of policy update distance is consistent with that in the PPO. We then describe the implementation details of the proposed Off-Policy PPO which iteratively updates policies by optimizing the proposed clipped surrogate objective. Finally, the experimental results on representative continuous control tasks validate that our method outperforms the state-of-the-art methods on most tasks. Wenjia Meng, Gang Pan 0001, Yilong Yin |
AAAI | 4 |
| 2023 | Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsabstractSelf-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence predictive models require mining more inherent properties in videos, such as structural information. In this work, we propose the VideoHiGraph, a space-time correspondence framework based on a learnable graph kernel. Concerning the video as the spatial-temporal graph, the learning objectives of VideoHiGraph are emanated in a self-supervised manner for predicting unobserved hidden graphs via graph kernel manner. We learn a representation of the temporal coherence across frames in which pairwise similarity defines the structured hidden graph, such that a biased random walk graph kernel along the sub-graph can predict long-range correspondence. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The self-supervision of the model training is formed by the structural and temporal consistency of the graph. VideoHiGraph achieves superior performance and demonstrates its robustness across the benchmark of label propagation tasks involving objects, semantic parts, keypoints, and instances. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/VideoHiGraph. Zheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin, Jianbing Shen |
AAAI | 4 |
| 2023 | MetaViewer: Towards A Unified Multi-View RepresentationabstractExisting multi-view representation learning methods typically follow a specific-to-uniform pipeline, extracting latent features from each view and then fusing or aligning them to obtain the unified object representation. However, the manually pre-specified fusion functions and aligning criteria could potentially degrade the quality of the derived representation. To overcome them, we propose a novel uniform-to-specific multi-view learning framework from a meta-learning perspective, where the unified representation no longer involves manual manipulation but is automatically derived from a meta-learner named MetaViewer. Specifically, we formulated the extraction and fusion of view-specific latent features as a nested optimization problem and solved it by using a bi-level optimization scheme. In this way, MetaViewer automatically fuses view-specific features into a unified one and learns the optimal fusion scheme by observing reconstruction processes from the unified to the specific over all views. Extensive experimental results in downstream classification and clustering tasks demonstrate the efficiency and effectiveness of the proposed method. Ren Wang 0011, Haoliang Sun, Yuling Ma, Xiaoming Xi, Yilong Yin |
CVPR | 5 |
| 2023 | MHPL: Minimum Happy Points Learning for Active Source Free Domain AdaptationabstractSource free domain adaptation (SFDA) aims to transfer a trained source model to the unlabeled target domain without accessing the source data. However, the SFDA setting faces a performance bottleneck due to the absence of source data and target supervised information, as evidenced by the limited performance gains of the newest SFDA methods. Active source free domain adaptation (ASFDA) can break through the problem by exploring and exploiting a small set of informative samples via active learning. In this paper, we first find that those satisfying the proper-ties of neighbor-chaotic, individual-different, and source-dissimilar are the best points to select. We define them as the minimum happy (MH) points challenging to explore with existing methods. We propose minimum happy points learning (MHPL) to explore and exploit MH points actively. We design three unique strategies: neighbor environment uncertainty, neighbor diversity relaxation, and one-shot querying, to explore the MH points. Further, to fully exploit MH points in the learning process, we design a neighbor focal loss that assigns the weighted neighbor purity to the cross entropy loss of MH points to make the model focus more on them. Extensive experiments verify that MHPL remarkably exceeds the various types of baselines and achieves significant performance gains at a small cost of labeling. Zhongyi Han, Zhiyan Zhang, Rundong He, Yilong Yin |
CVPR | 5 |
| 2023 | Fine-Grained Classification with Noisy LabelsabstractLearning with noisy labels (LNL) aims to ensure model generalization given a label-corrupted training set. In this work, we investigate a rarely studied scenario of LNL on fine-grained datasets (LNL-FG), which is more practical and challenging as large inter-class ambiguities among fine-grained classes cause more noisy labels. We empirically show that existing methods that work well for LNL fail to achieve satisfying performance for LNL-FG, arising the practical need of effective solutions for LNL-FG. To this end, we propose a novel framework called stochastic noise-tolerated supervised contrastive learning (SNSCL) that confronts label noise by encouraging distinguishable representation. Specifically, we design a noise-tolerated supervised contrastive learning loss that incorporates a weight-aware mechanism for noisy label correction and selectively updating momentum queue lists. By this mechanism, we mitigate the effects of noisy anchors and avoid inserting noisy labels into the momentum-updated queue. Besides, to avoid manually-defined augmentation strategies in contrastive learning, we propose an efficient stochastic module that samples feature embeddings from a generated distribution, which can also enhance the representation ability of deep models. SNSCL is general and compatible with prevailing robust LNL strategies to improve their performance for LNL-FG. Extensive experiments demonstrate the effectiveness of SNSCL. Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Chenhui Guo, Yilong Yin |
CVPR | 6 |
| 2023 | Fine-grained Urban Flow Inference with Unobservable Data via Space-Time Attraction LearningabstractFine-grained urban flow inference focuses on inferring fine-grained urban flows based solely on coarse-grained observations, which is essential for the city management and transportation services. However, most of the existing methods assume that partial urban flows in coarse-grained regions cannot be observable. In this study, we propose a multi-task framework known as UrbanSTA with space-time attraction learning to estimate missing values in coarse-grained urban flow map and forecast fine-grained urban flows simultaneously. Specifically, UrbanSTA comprises two parts: the flow completion network STA and the fine-grained flow inference network FIN. STA captures space-time features with a separable space-time attention encoder and recovers the missing flow features with a decoder. FIN directly uses complete coarse-grained flow features for further decoding, and reconstructs fine-grained flow features based on the complex associations between coarse- and fine-grained urban flows, relying on upsampling constraints. Extensive experiments conducted on two real-world datasets demonstrate that our proposed model yields the best results compared to other state-of-the-art methods. The source code has been provided at https://github.com/Wangzheaos/UrbanSTA. Yuansheng Liu, Yongshun Gong, Wei Liu 0007, Meng Chen 0003, Yilong Yin, Yu Zheng 0004 |
ICDM | 6 |
| 2023 | Towards Understanding Generalization of Macro-AUC in Multi-label LearningabstractMacro-AUC is the arithmetic mean of the class-wise AUCs in multi-label learning and is commonly used in practice. However, its theoretical understanding is far lacking. Toward solving it, we characterize the generalization properties of various learning algorithms based on the corresponding surrogate losses w.r.t. Macro-AUC. We theoretically identify a critical factor of the dataset affecting the generalization bounds: the label-wise class imbalance. Our results on the imbalance-aware error bounds show that the widely-used univariate loss-based algorithm is more sensitive to the label-wise class imbalance than the proposed pairwise and reweighted loss-based ones, which probably implies its worse performance. Moreover, empirical results on various datasets corroborate our theory findings. To establish it, technically, we propose a new (and more general) McDiarmid-type concentration inequality, which may be of independent interest. Guoqiang Wu, Chongxuan Li, Yilong Yin |
ICML | 3 |
| 2023 | Multi-View Representation Learning via View-Aware ModulationabstractMulti-view (representation) learning derives an entity's representation from its multiple observable views to facilitate various downstream tasks. The most challenging topic is how to model unobserved entities and their relationships to specific views. To this end, this work proposes a novel multi-view learning method using a View-Aware parameter Modulation mechanism, termed VAM. The key idea is to use trainable parameters as proxies for unobserved entities and views, such that modeling entity-view relationships is converted into modeling the relationship between proxy parameters. Specifically, we first build a set of trainable parameters to learn a mapping from multi-view data to the unified representation as the entity proxy. Then we learn a prototype for each view and design a Modulation Parameter Generator (MPG) that learns a set of view-aware scale and shift parameters from prototypes to modulate the entity proxy and obtain view proxies. By constraining the representativeness, uniqueness, and simplicity of the proxies and proposing an entity-view contrastive loss, parameters are alternatively updated. We end up with a set of discriminative prototypes, view proxies, and an entity proxy that are flexible enough to yield robust representations for out-of-sample entities. Extensive experiments on five datasets show that the results of our VAM outperform existing methods in both classification and clustering tasks. Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yuxiu Lin, Xiaoming Xi, Yilong Yin |
ACM Multimedia | 6 |
| 2023 | Topological Structure Learning for Weakly-Supervised Out-of-Distribution DetectionabstractOut-of-distribution~(OOD) detection is the key to deploying models safely in the open world. For OOD detection, collecting sufficient in-distribution~(ID) labeled data is usually more time-consuming and costly than unlabeled data. When ID labeled data is limited, the previous OOD detection methods are no longer superior due to their high dependence on the amount of ID labeled data. Based on limited ID labeled data and sufficient unlabeled data, we define a new setting called Weakly-Supervised Out-of-Distribution Detection (WSOOD). To solve the new problem, we propose an effective method called Topological Structure Learning (TSL). Firstly, TSL uses a contrastive learning method to build the initial topological structure space for ID and OOD data. Secondly, TSL mines effective topological connections in the initial topological space. Finally, based on limited ID labeled data and mined topological connections, TSL reconstructs the topological structure in a new topological space to increase the separability of ID and OOD instances. Extensive studies on several representative datasets show that TSL remarkably outperforms the state-of-the-art, verifying the validity and robustness of our method in the new setting of WSOOD. Rundong He, Rongxue Li, Zhongyi Han, Xihong Yang, Yilong Yin |
ACM Multimedia | 5 |
| 2023 | Personalized Single Image Reflection Removal Network through Adaptive Cascade RefinementabstractIn this paper, we aim to restore a reflection-free image from a single reflection-contaminated image captured through the glass. Many deep-learning-based methods attempt to solve the challenging problem by utilizing a uniform model obtained from training data for all test images. Hence, the distinctive characteristics of the test images are not considered. Besides, several methods use a cascade structure in image restoration to refine the results. But they blindly cascade modules with the same weights, improving the model's performance only to a certain extent. To address these problems, we propose a personalized single-image reflection removal network through adaptive cascade refinement (PNACR) based on meta-learning and self-supervised learning. While meta-learning can rapidly adapt to a new task with a few samples, PNACR can remove reflections of a new image with its distinctive characteristics learned by self-supervised learning. Furthermore, the proposed adaptive cascade model can adjust the weights of the model at the next iteration according to the output of the model at the current iteration, significantly improving the model's performance. Hence, the proposed model can learn information from both external training data and the new input image to provide a personalized reflection removal model for each new input image. Extensive comparison and ablation experiments on publicly available datasets demonstrate the validity of the proposed method in quantitative evaluation metrics and qualitative visualization. Mengyi Wang 0001, Xinxin Zhang 0004, Yongshun Gong, Yilong Yin |
ACM Multimedia | 4 |
| 2023 | Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosabstractHuman mesh reconstruction (HMR) from monocular video is the key step to many mixed reality and robotic applications. Although existing methods show promising results by capturing frames' temporal information, these methods predict human mesh with the design of implicit temporal learning modules in a sequence to frame manner. To mine more temporal information from the video, we present a bi-level clip inference network for HMR, which leverages both local motion and global context explicitly for dense 3D reconstruction. Specifically, we propose a novel bi-level temporal fusion strategy that takes both neighboring and long-range relations into consideration. In addition, different from traditional frame-wise operation, we investigate an alternative perspective by treating video-based HMR as clip-wise inference. We evaluate the proposed method on multiple datasets (3DPW, Human3.6M, and MPI-INF-3DHP) quantitatively and qualitatively, demonstrating a significant improvement over existing methods (in terms of PA-MPJPE, ACC-Error etc). Furthermore, we extend the proposed method on more challenging Multiple Shots HMR task to demonstrate its generalizability. Some visual demos can be seen https://github.com/bicf0/bicf_demo. Peng Wu 0014, Xiankai Lu, Jianbing Shen, Yilong Yin |
ACM Multimedia | 4 |
| 2023 | LHAct: Rectifying Extremely Low and High Activations for Out-of-Distribution DetectionabstractIn recent years, out-of-distribution (OOD) detection has emerged as a crucial research area, especially when deploying AI products in real-world scenarios. OOD detection researchers have made significant efforts to mitigate the adverse effects of abnormal activation values (abbr. activations) that refer to the outputs of the activation function acted on feature maps. Since abnormal activations would cause difficulty in separating ID and OOD data, the previous unified solution is to rectify the extremely high abnormal activations by clipping them with a pre-defined threshold or filtering them with a low-pass filter. However, it ignores the extremely low abnormal activations, and the proposed rectification strategy is always suboptimal because the used rectification function is non-convergence or high-intensity convergence, leading to under-rectification or over-rectification. In this paper, we propose an approach called Rectifying Extremely Low and High Activations (LHAct). LHAct includes a newly-designed function to rectify the extremely low and high activations at the same time. Specifically, LHAct increases the difference of means between ID and OOD activation distributions while decreasing their variances after processing the original activations. Our theoretical analyses demonstrate that LHAct significantly enhances the separability of ID and OOD data. By conducting extensive experiments, we demonstrate that LHAct surpasses previous activation-based methods significantly and generalizes well to other architectures and OOD scores. Code is available at: https://github.com/ystyuan/LHAct.git. Rundong He, Zhongyi Han, Yilong Yin |
ACM Multimedia | 4 |
| 2023 | M3R: Masked Token Mixup and Cross-Modal Reconstruction for Zero-Shot LearningabstractIn the zero-shot learning (ZSL), learned representation spaces are often biased toward seen classes, thus limiting the ability to predict previously unseen classes. In this paper, we propose Masked token Mixup and cross-Modal Reconstruction for zero-shot learning, termed as M3R, which can significantly alleviate the bias toward seen classes. The M3R mainly consists of Random Token Mixup (RTM), Unseen Class Detection (UCD), and Hard Cross-modal Reconstruction (HCR). Firstly, mappings without proper adaptations to unseen classes would cause the bias toward seen classes. To address this issue, the RTM is introduced to generate diverse unseen class agents, thereby broadening the representation space to cover unknown classes. It is applied at a randomly selected layer in the Vision Transformer, producing smooth low- and high-level representation space boundaries to cover rich attributes. Secondly, it should be noted that unseen class agents generated by the RTM may be mixed with seen class samples. To overcome this challenge, the UCD is designed to generate greater entropy values for unseen classes, thereby distinguishing seen classes from unseen classes. Thirdly, to further mitigate the bias toward seen classes and explore associations between semantics and visual images, the HCR is proposed, which can reconstruct masked pixels based on few discriminative tokens and attribute embeddings. This approach can enable models to have a deep understanding of image contents and build powerful connections between semantic attributes and visual information. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of our proposed M3R model. Peng Zhao 0016, Qiangchang Wang, Yilong Yin |
ACM Multimedia | 3 |
| 2023 | Unified 3D Segmenter As Prototypical ClassifiersabstractThe task of point cloud segmentation, comprising semantic, instance, and panoptic segmentation, has been mainly tackled by designing task-specific network architectures, which often lack the flexibility to generalize across tasks, thus resulting in a fragmented research landscape. In this paper, we introduce ProtoSEG, a prototype-based model that unifies semantic, instance, and panoptic segmentation tasks. Our approach treats these three homogeneous tasks as a classification problem with different levels of granularity. By leveraging a Transformer architecture, we extract point embeddings to optimize prototype-class distances and dynamically learn class prototypes to accommodate the end tasks. Our prototypical design enjoys simplicity and transparency, powerful representational learning, and ad-hoc explainability. Empirical results demonstrate that ProtoSEG outperforms concurrent well-known specialized architectures on 3D point cloud benchmarks, achieving 72.3%, 76.4% and 74.2% mIoU for semantic segmentation on S3DIS, ScanNet V2 and SemanticKITTI, 66.8% mCov and 51.2% mAP for instance segmentation on S3DIS and ScanNet V2, 62.4% PQ for panoptic segmentation on SemanticKITTI, validating the strength of our concept and the effectiveness of our algorithm. The code and models are available at https://github.com/zyqin19/PROTOSEG. Zheyun Qin, Cheng Han 0001, Qifan Wang 0001, Xiushan Nie, Yilong Yin, Xiankai Lu |
NeurIPS | 5 |
| 2023 | Abductive subconcept learning
Zhongyi Han, Le-Wen Cai, Wang-Zhou Dai, Yu-Xuan Huang, Benzheng Wei, Yilong Yin |
Sci. China Inf. Sci. | 7 |
| 2023 | Difficulty-aware prior-guided hierarchical network for adaptive segmentation of breast tumors
Sumaira Hussain, Xiaoming Xi, Inam Ullah 0002, Syeda Wajiha Naim, Kashif Shaheed, Cuihuan Tian, Yilong Yin |
Sci. China Inf. Sci. | 7 |
| 2023 | Deep visual-linguistic fusion network considering cross-modal inconsistency for rumor detection
Yang Yang 0074, Ran Bao, Weili Guo, De-Chuan Zhan, Yilong Yin, Jian Yang 0003 |
Sci. China Inf. Sci. | 5 |
| 2023 | PESTA: An Elastic Motion Capture Data Retrieval Method
Zifei Jiang, Wei Li 0143, Yan Huang 0003, Yilong Yin, C.-C. Jay Kuo, Jingliang Peng |
J. Comput. Sci. Technol. | 4 |
| 2023 | Vehicle re-identification based on grouping aggregation attention and cross-part interaction
Xiyu Pang, Xiushan Nie, Yilong Yin, Gangwu Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Deep regional detail-aware hashing
Yuling Ma, Jie Guo 0012, Xiushan Nie, Yilong Yin |
Multim. Syst. | 6 |
| 2023 | Robust seed selection of foreground and background priors based on directional blocks for saliency-detection system
Muwei Jian, Ruihong Wang, Hui Yu 0001, Junyu Dong, Gongfa Li, Yilong Yin, Kin-Man Lam 0001 |
Multim. Tools Appl. | 7 |
| 2023 | Missingness-Pattern-Adaptive Learning With Incomplete DataabstractMany real-world problems deal with collections of data with missing values, e.g., RNA sequential analytics, image completion, video processing, etc. Usually, such missing data is a serious impediment to a good learning achievement. Existing methods tend to use a universal model for all incomplete data, resulting in a suboptimal model for each missingness pattern. In this paper, we present a general model for learning with incomplete data. The proposed model can be appropriately adjusted with different missingness patterns, alleviating competitions between data. Our model is based on observable features only, so it does not incur errors from data imputation. We further introduce a low-rank constraint to promote the generalization ability of our model. Analysis of the generalization error justifies our idea theoretically. In additional, a subgradient method is proposed to optimize our model with a proven convergence rate. Experiments on different types of data show that our method compares favorably with typical imputation strategies and other state-of-the-art models for incomplete data. More importantly, our method can be seamlessly incorporated into the neural networks with the best results achieved. The source code is released at https://github.com/YS-GONG/missingness-patterns. Yongshun Gong, Zhibin Li 0002, Wei Liu 0007, Xiankai Lu, Xinwang Liu 0002, Ivor W. Tsang, Yilong Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Towards Accurate and Robust Domain Adaptation Under Multiple Noisy EnvironmentsabstractIn many non-stationary environments, machine learning algorithms usually confront the distribution shift scenarios. Previous domain adaptation methods have achieved great success. However, they would lose algorithm robustness in multiple noisy environments where the examples of source domain become corrupted by label noise, feature noise, or open-set noise. In this paper, we report our attempt toward achieving noise-robust domain adaptation. We first give a theoretical analysis and find that different noises have disparate impacts on the expected target risk. To eliminate the effect of source noises, we propose offline curriculum learning minimizing a newly-defined empirical source risk. We suggest a proxy distribution-based margin discrepancy to gradually decrease the noisy distribution distance to reduce the impact of source noises. We propose an energy estimator for assessing the outlier degree of open-set-noise examples to defeat the harmful influence. We also suggest robust parameter learning to mitigate the negative effect further and learn domain-invariant feature representations. Finally, we seamlessly transform these components into an adversarial network that performs efficient joint optimization for them. A series of empirical studies on the benchmark datasets and the COVID-19 screening task show that our algorithm remarkably outperforms the state-of-the-art, with over 10% accuracy improvements in some transfer tasks. Zhongyi Han, Xian-Jin Gui, Haoliang Sun, Yilong Yin, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Neighborhood-based credibility anchor learning for universal domain adaptation
Wan Su, Zhongyi Han, Rundong He, Benzheng Wei, Xueying He, Yilong Yin |
Pattern Recognit. | 6 |
| 2023 | Attentional prototype inference for few-shot segmentation
Haoliang Sun, Xiankai Lu, Yilong Yin, Xiantong Zhen, Cees Snoek, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2023 | Detail enhancement-based vehicle re-identification with orientation-guided re-ranking
Ziruo Sun, Xiushan Nie, Xiaopeng Bi, Yilong Yin |
Pattern Recognit. | 5 |
| 2023 | Triple-attention interaction network for breast tumor classification based on multi-modality images
Xiaoming Xi, Kesong Wang, Liangyun Sun, Lingzhao Meng, Xiushan Nie, Lishan Qiao, Yilong Yin |
Pattern Recognit. | 8 |
| 2023 | Supervised Discrete Multiple-Length Hashing for Image RetrievalabstractHashing can facilitate efficient retrieval and storage for large-scale images due to the binary representation. In the real applications, the trade-off between retrieval accuracy and speed is essential for designing a hashing framework, which is reflected by variable hash code lengths. In light of this, the existing hashing methods need to train different models for different lengths of hash codes, leading to considerable training time cost and hashing flexibility reduction. Given that a sample can be represented by various hash codes with different lengths, there are some helpful relationships that can boost the performance of hashing methods. However, the existing hashing methods do not fully utilize these relationships. To address the aforementioned issues, we propose a new model, known as supervised discrete multiple-length hashing (SDMLH), to simultaneously learn hash codes with multiple lengths. In this proposed SDMLH method, three types of information are respectively derived, from the hash codes with different lengths. The original features of the samples, and the label, are applied for hash learning. Unlike the existing hashing methods, SDMLH can fully employ the assistance among hash codes with different lengths and learn them in one step. Furthermore, given a hash length meeting the demand of users, we propose a hash fusion strategy to obtain the hash code with this desirable length by fusing the multiple-length hash codes. This obtained hash code outperforms the one learned directly. In addition, SDMLH can generate the hash code of any length that is shorter than the sum length of given multiple hash codes with the fusion strategy. To the best of our knowledge, SDMLH is one of the first attempts for learning multiple-length hash codes simultaneously. We conduct extensive experiments based on three benchmark datasets, demonstrating the superiority of this proposed method. Xiushan Nie, Xingbo Liu, Jie Guo 0012, Yilong Yin |
IEEE Trans. Big Data | 5 |
| 2023 | Single Image Reflection Removal Based on Dark Channel Sparsity PriorabstractThe major task of reflection removal methods is to restore a reflection-free image from a reflection-contaminated image taken through glass. We propose an algorithm to remove reflections from a single image by means of the$l_{0}$-regularized dark channel sparsity prior and an$l_{0}$gradient sparsity prior. In addition, we analyze the difference between the dark channel map in the reflection-contaminated image and the reflection-free image empirically and mathematically. Moreover, a new data fidelity term is introduced to handle strong reflections and preserve high-frequency details in the recovered transmission image. Different from the model used in most state-of-the-art methods, our reflection removal model does not rely on the assumption of out-of-focus objects in the reflection layer. Quantitative evaluation on several publicly available real-world image datasets including ground-truth demonstrates the high accuracy of our algorithm. Qualitative evaluation of extensive experimental results on real-world images shows the competitive performance of the proposed method compared with the state-of-the-art reflection removal methods. Xinxin Zhang 0004, Kaixin Xing, Da Chen 0002, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Small-Area Finger Vein RecognitionabstractRecently, finger vein sensors have been embedded in all kinds of electronic devices for personal identification, such as intelligent door locks and attendance machines. The embedded sensors are generally small, thus capturing only part of the finger vein. However, prior studies have focused on near-full finger vein recognition, without considering the partial finger vein image caused by the small imaging window of the finger vein sensor. This paper aims to study personal identification based on partial finger vein images, known as small-area finger vein recognition. The effect of the small-area finger vein on recognition performance is first analyzed by cutting out the local part from the near-full finger vein image to model a small-area finger vein image. Second, a small-area finger vein database is built using a commercial finger vein imaging device, in which the vein pattern from approximately one-third of one adult finger is captured. To explore more discriminative information from small-area finger vein images, we propose a locality-constrained consistent dictionary learning (LCDL) method to fuse multiple features for small-area finger vein recognition. Finally, the proposed method is evaluated on the self-built small-area finger vein database and four synthetic small-area finger vein databases. Experimental results show the promising recognition performance of the proposed method. Lu Yang 0005, Gongping Yang 0001, Jun Wang 0071, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence LearningabstractSelf-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph. Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui |
IEEE Trans. Image Process. | 5 |
| 2023 | Missing Value Imputation for Multi-View Urban Statistical Data via Spatial Correlation LearningabstractAs a developing trend of urbanization, massive amounts of urban statistical data with multiple views (e.g., views of Population and Economy) are increasingly collected and benefited to diverse domains, including transportation service, regional analysis, etc. Unfortunately, these statistical data that are divided into fine-grained regions usually suffer from missing value problem during the acquisition and storage processes. It is mianly caused by some inevitable circumstances, e.g., the document defacement, statistical difficulty in remote districts, and inaccurate information cleaning, etc. Those missing entries which make valuable information invisible may distort the further urban analysis. To improve the quality of missing data imputation, we propose an improved spatial multi-kernel learning method to guide the imputation process incorporating with the adaptive-weight non-negative matrix factorization strategy. Our model takes into account the regional latent similarities and the real geographical positions as well as the correlations among various views that are able to complete missing values precisely. We conduct intensive experiments to evaluate our method and compare with other state-of-the-art approaches on real-world datasets. All the empirical results show that the proposed model outperforms all the other state-of-the-art methods. Additionally, our model represents a strong generalization ability across multiple cities. Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Yilong Yin, Yu Zheng 0004 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Forecasting Fine-Grained Urban Flows Via Spatio-Temporal Contrastive Self-SupervisionabstractAs a critical task of the urban traffic services, fine-grained urban flow inference (FUFI) benefits in many fields including intelligent transportation management, urban planning, public safety. FUFI is a technique that focuses on inferring fine-grained urban flows depending solely on observed coarse-grained data. However, existing methods always require massive learnable parameters and the complex network structures. To reduce these defects, we formulate a contrastive self-supervision method to predict fine-grained urban flows taking into account all correlated spatial and temporal contrastive patterns. Through several well-designed self-supervised tasks, uncomplicated networks have a strong ability to capture high-level representations from flow data. Then, a fine-tuning network combining with three pre-training encoder networks is proposed. We conduct experiments to evaluate our model and compare with other state-of-the-art methods by using two real-world datasets. All the empirical results not only show the superiority of our model against other comparative models, but also demonstrate its effectiveness in the resource-limited environment. Yongshun Gong, Meng Chen 0003, Junbo Zhang 0004, Yu Zheng 0004, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Zero-Shot Hashing via Asymmetric Ratio Similarity MatrixabstractZero-shot hashing targets to learn the hash codes of images in unseen classes based on the limited training data provided by seen classes. In zero-shot hashing, transferring the supervised knowledge, such as attributes and semantic relations, from seen classes to unseen ones is a widely employed method, where the performance is always subject to the ability to capture these supervised knowledge (which is always difficult to obtain). Therefore, in this study, we propose a new methodology for zero-shot hashing via an asymmetric ratio similarity matrix (ASZH), which only needs to calculate the semantic similarity among seen classes for hash learning. Specifically, we use an asymmetric ratio matrix in the similarity calculation to further explore the influence of similarity, where the values of positive weights for similar samples are not equivalent to those of negative ones for dissimilar samples. Additionally, a theoretical analysis regarding the utilization of an asymmetric ratio matrix is provided in this study. The experiments on three large benchmark datasets indicate that the proposed method achieves excellent performance than several state-of-the-art hashing methods. Xiushan Nie, Xingbo Liu, Lu Yang 0005, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Not All Parameters Should Be Treated Equally: Deep Safe Semi-supervised Learning under Class Distribution MismatchabstractDeep semi-supervised learning (SSL) aims to utilize a sizeable unlabeled set to train deep networks, thereby reducing the dependence on labeled instances. However, the unlabeled set often carries unseen classes that cause the deep SSL algorithm to lose generalization. Previous works focus on the data level that they attempt to remove unseen class data or assign lower weight to them but could not eliminate their adverse effects on the SSL algorithm. Rather than focusing on the data level, this paper turns attention to the model parameter level. We find that only partial parameters are essential for seen-class classification, termed safe parameters. In contrast, the other parameters tend to fit irrelevant data, termed harmful parameters. Driven by this insight, we propose Safe Parameter Learning (SPL) to discover safe parameters and make the harmful parameters inactive, such that we can mitigate the adverse effects caused by unseen-class data. Specifically, we firstly design an effective strategy to divide all parameters in the pre-trained SSL model into safe and harmful ones. Then, we introduce a bi-level optimization strategy to update the safe parameters and kill the harmful parameters. Extensive experiments show that SPL outperforms the state-of-the-art SSL methods on all the benchmarks by a large margin. Moreover, experiments demonstrate that SPL can be integrated into the most popular deep SSL networks and be easily extended to handle other cases of class distribution mismatch. Rundong He, Zhongyi Han, Yang Yang 0074, Yilong Yin |
AAAI | 4 |
| 2022 | SNAIL: Semi-Separated Uncertainty Adversarial Learning for Universal Domain Adaptation
Zhongyi Han, Wan Su, Rundong He, Yilong Yin |
ACML | 4 |
| 2022 | Transferable Discriminative Learning for Medical Open-Set Domain Adaptation: Application to Pneumonia ClassificationabstractPrevious pneumonia classification algorithms have succeeded in the clinic under closed and static environments. However, in the real world, the emergence of new categories (e.g., COVID-19) and changes in data distribution will cause the existing methods to lose their robustness. In this paper, we formalize this problem as medical open-set domain adaptation under open and dynamic environments. The critical challenge of this problem is to accurately detect the open class samples with subtle differences from the common class. To achieve that, we propose transferable discriminative learning that remarkably achieves robust pneumonia classification with distribution shift and open class emerging. First, we propose the transferable high-density clustering module to detect open class samples and obtain reliable common class samples by considering the density degree. Secondly, we present the transferable triplet loss to enlarge the semantic feature difference between common class and open class samples. Finally, we design the transferable scoring function to detect open class samples effectively. A series of empirical studies show that our algorithm remarkably outperforms state-of-the-art methods. This result demonstrates its potential as a clinical tool for medical open-set domain adaptation. Wan Su, Zhongyi Han, Yilong Yin |
BIBM | 4 |
| 2022 | Safe-Student for Safe Deep Semi-Supervised Learning with Unseen-Class Unlabeled DataabstractDeep semi-supervised learning (SSL) methods aim to take advantage of abundant unlabeled data to improve the algorithm performance. In this paper, we consider the problem of safe SSL scenario where unseen-class instances appear in the unlabeled data. This setting is essential and commonly appears in a variety of real applications. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the small number of labeled data and ignoring the availability of unlabeled data. To take advantage of these unseen-class data and ensure performance, we propose a safe SSL method called SAFE-STUDENT from the teacher-student view. Firstly, a new scoring function called energy-discrepancy (ED) is proposed to help the teacher model improve the security of instances selection. Then, a novel unseen-class label distribution learning mechanism mitigates the unseen-class perturbation by calibrating the unseen-class label distribution. Finally, we propose an iterative optimization strategy to facilitate teacher-student network learning. Extensive studies on several representative datasets show that SAFE-STUDENT remarkably outperforms the state-of-the-art, verifying the feasibility and robustness of our method in the under-explored problem. Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
CVPR | 4 |
| 2022 | Exploring Domain-Invariant Parameters for Source Free Domain AdaptationabstractSource-free domain adaptation (SFDA) newly emerges to transfer the relevant knowledge of a well-trained source model to an unlabeled target domain, which is critical in various privacy-preserving scenarios. Most existing methods focus on learning the domain-invariant representations depending solely on the target data, leading to the obtained representations are target-specific. In this way, they cannot fully address the distribution shift problem across domains. In contrast, we provide a fascinating insight: rather than attempting to learn domain-invariant representations, it is better to explore the domain-invariant parameters of the source model. The motivation behind this insight is clear: the domain-invariant representations are dominated by only partial parameters of an available deep source model. We devise the Domain-Invariant Parameter Exploring (DIPE) approach to capture such domain-invariant parameters in the source model to generate domain-invariant representations. A distinguishing method is developed correspondingly for two types of parameters, i.e., domain-invariant and domain-specific parameters, as well as an effective update strategy based on the clustering correction technique and a target hypothesis is proposed. Extensive experiments verify that DIPE successfully exceeds the current state-of-the-art models on many domain adaptation datasets. Zhongyi Han, Yongshun Gong, Yilong Yin |
CVPR | 4 |
| 2022 | Self-Filtering: A Noise-Aware Sample Selection for Label Noise with Confidence Penalization
Qi Wei 0004, Haoliang Sun, Xiankai Lu, Yilong Yin |
ECCV (30) | 4 |
| 2022 | Online Ecg Biometrics Via Hadamard CodeabstractIn recent years, Electrocardiogram (ECG) biometrics has gained extensive attention. However, most existing methods adopted offline batch learning, which means that they need to accumulate all data and retrain the model when new data comes. Therefore, it is inefficient and unpractical for them to handle the online scenario where new data may continually come. To overcome the above limitation, we propose a novel ECG biometrics framework, termed Online ECG Biomet-rics based on Hadamard Codes. Firstly, we leverage matrix factorization to learn discriminative representations for ECG signals from their base feature space. Considering to leverage the orthogonal property of the Hadamard matrix, we use it to construct Hadamard codes to represent individuals and further guide the learning of representations. Furthermore, we develop an online optimization algorithm, which is efficient and effective to investigate the incremental problem in the context of ECG biometrics. The experimental results on two benchmark datasets indicate the merits of the proposed framework over the state-of-the-art. Kuikui Wang, Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
ICASSP | 5 |
| 2022 | Joint Dual-Domain Matrix Factorization for ECG Biometric RecognitionabstractElectrocardiogram (ECG) biometrics has aroused extensive attention in the research field of biometric recognition. How-ever, most existing methods either only consider a single do-main (time domain or frequency domain) to extract features or extract multi-features while ignoring the specific proper-ties of each domain. In this paper, we propose a novel ECG biometrics framework termed Joint Dual-domain Matrix Factorization (JDMF). JDMF learns latent spaces for each do-main by exploring the cross-correlations between them and preserving domain-specific properties. To endow the latent spaces with more powerful representation capabilities, JDMF further makes full use of the supervised information and could automatically learn the weights of domains. The experimental results on two widely-used datasets indicate that the proposed framework can outperform state-of-the-arts. Kuikui Wang, Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
ICASSP | 5 |
| 2022 | Exploring Linear Feature Disentanglement for Neural NetworksabstractNon-linear activation functions, e.g., Sigmoid, ReLU, and Tanh, have achieved great success in neural networks (NNs). Due to the complex non-linear characteristic of samples, the objective of those activation functions is to project samples from their original feature space to a linear separable feature space. This phenomenon ignites our interest in exploring whether all features need to be transformed by all nonlinear functions in current typical NNs, i.e., whether there exists a part of features arriving at the linear separable feature space in the intermediate layers, that does not require further non-linear variation but an affine transformation instead. To validate the above hypothesis, we explore the problem of linear feature disentanglement for neural networks in this paper. Specifically, we devise a learnable mask module to distinguish between linear and non-linear features. Through our designed experiments we found that some features reach the linearly separable space earlier than the others and can be detached partly from the NNs. The explored method also provides a readily feasible pruning strategy which barely affects the performance of the original model. We conduct our experiments on four datasets and present promising results. Tiantian He 0004, Zhibin Li 0002, Yongshun Gong, Yazhou Yao, Xiushan Nie, Yilong Yin |
ICME | 6 |
| 2022 | Series Photo Selection via Multi-View Graph LearningabstractSeries photo selection (SPS) is an important branch of the image aesthetics quality assessment, which focuses on finding the best one from a series of nearly identical photos. While a great progress has been observed, most of the existing SPS approaches concentrate solely on extracting features from the original image, neglecting that multiple views, e.g, saturation level, color histogram and depth of field of the image, will be of benefit to successfully reflecting the subtle aesthetic changes. Taken multi-view into consideration, we leverage a graph neural network to construct the relationships between multi-view features. Besides, multiple views are aggregated with an adaptive-weight self-attention module to verify the significance of each view. Finally, a siamese network is proposed to select the best one from a series of nearly identical photos. Experimental results demonstrate that our model accomplish the highest success rates compared with competitive methods. Lu Zhang 0062, Yongshun Gong, Jian Zhang 0002, Xiushan Nie, Yilong Yin |
ICME | 6 |
| 2022 | RONF: Reliable Outlier Synthesis under Noisy Feature Space for Out-of-Distribution DetectionabstractOut-of-distribution~(OOD) detection is fundamental to guaranteeing the reliability of multimedia applications during deployment in the open world. However, due to the lack of supervision signals from OOD data, the current model easily outputs overconfident predictions to OOD data during the inference phase. Several previous methods rely on large-scale auxiliary OOD datasets for model regularization. However, obtaining suitable and clean large-scale auxiliary OOD datasets is usually challenging. In this paper, we present Reliable Outlier synthesis under Noisy Feature space (RONF), which synthesizes reliable virtual outliers in noisy feature space to provide supervision signals for model regularization. Specifically, RONF first introduces a novel virtual outlier synthesis strategy Boundary Feature Mixup (BFM), which mixes up samples from the low-likelihood region of the class-conditional distribution in the feature space. However, the feature space is noisy due to the spurious features, which cause unreliable outlier synthesizing. To mitigate this problem, RONF then introduces Optimal Parameter Learning (OPL) to obtain desirable features and remove spurious features. Alongside, RONF proposes a provable and effective scoring function called Energy with Energy Discrepancy (EED) for the uncertainty measurement of OOD data. Extensive studies on several representative datasets of multimedia applications show that RONF outperforms the state-of-the-arts remarkably Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin |
ACM Multimedia | 4 |
| 2022 | Difficulty-aware bi-network with spatial attention constrained graph for axillary lymph node segmentation
Xiaoming Xi, Xianjing Meng, Zheyun Qin, Xiushan Nie, Yongjian Wu 0001, Chenglong Li 0004, Yilong Yin |
Sci. China Inf. Sci. | 10 |
| 2022 | Learning disentangled representation for self-supervised video object segmentation
Wenjie Hou, Zheyun Qin, Xiaoming Xi, Xiankai Lu, Yilong Yin |
Neurocomputing | 5 |
| 2022 | Enabling the interpretability of pretrained venue representations using semantic categories
Ning An 0004, Meng Chen 0003, Li Lian, Kai Zhang 0039, Xiaohui Yu 0001, Yilong Yin |
Knowl. Based Syst. | 7 |
| 2022 | Towards safe and robust weakly-supervised anomaly detection under subpopulation shift
Rundong He, Zhongyi Han, Yilong Yin |
Knowl. Based Syst. | 3 |
| 2022 | SNIP-FSL: Finding task-specific lottery jackpots for few-shot learning
Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yilong Yin |
Knowl. Based Syst. | 4 |
| 2022 | Robust multi-feature collective non-negative matrix factorization for ECG biometrics
Gongping Yang 0001, Kuikui Wang, Yilong Yin |
Pattern Recognit. | 5 |
| 2022 | Learning to rectify for robust learning with noisy labels
Haoliang Sun, Chenhui Guo, Qi Wei 0004, Zhongyi Han, Yilong Yin |
Pattern Recognit. | 5 |
| 2022 | Context-related video anomaly detection via generative adversarial network
Daoheng Li, Xiushan Nie, Yilong Yin |
Pattern Recognit. Lett. | 5 |
| 2022 | Adaptive Feature Aggregation in Deep Multi-Task Convolutional Neural NetworksabstractMulti-task learning in Convolutional Neural Networks (CNNs) has led to remarkable success in a variety of applications of computer vision. Towards effective multi-task CNN architectures, recent studies automatically learn the optimal combinations of task-specific features at single network layers. However, they generally learn an unchanged operation of feature combination after training, regardless of the characteristic changes of task-specific features across different inputs. In this paper, we propose a novel Adaptive Feature Aggregation (AFA) layer for multi-task CNNs, in which a dynamic aggregation mechanism is designed to allow each task adaptively determines the degree to which the knowledge sharing or preserving between tasks is needed based on the characteristics of inputs. We introduce two types of aggregation modules to the AFA layer, which realize the adaptive feature aggregation by capturing the feature dependencies of different tasks along the channel and spatial axes, respectively. The AFA layer is a plug-and-play component with low parameter and computation overheads, and can be trained end-to-end along with backbone networks. For both pixel-level and image-level tasks, we empirically show that our approach substantially outperforms the previous state-of-the-art methods of multi-task CNNs. The code and models are available athttps://github.com/zhenshen-mla/AFANet. Chaoran Cui, Zhen Shen 0001, Meng Chen 0003, Mingliang Xu 0001, Meng Wang 0001, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Learning Transferable Parameters for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) enables a learning machine to adapt from a labeled source domain to an unlabeled target domain under the distribution shift. Thanks to the strong representation ability of deep neural networks, recent remarkable achievements in UDA resort to learning domain-invariant features. Intuitively, the goal is that a good feature representation and the hypothesis learned from the source domain can generalize well to the target domain. However, the learning processes of domain-invariant features and source hypotheses inevitably involve domain-specific information that would degrade the generalizability of UDA models on the target domain. The lottery ticket hypothesis proves that only partial parameters are essential for generalization. Motivated by it, we find in this paper that only partial parameters are essential for learning domain-invariant information. Such parameters are termed transferable parameters that can generalize well in UDA. In contrast, the rest parameters tend to fit domain-specific details and often cause the failure of generalization, which are termed untransferable parameters. Driven by this insight, we propose Transferable Parameter Learning (TransPar) to reduce the side effect of domain-specific information in the learning process and thus enhance the memorization of domain-invariant information. Specifically, according to the distribution discrepancy degree, we divide all parameters into transferable and untransferable ones in each training iteration. We then perform separate update rules for the two types of parameters. Extensive experiments on image classification and regression tasks (keypoint detection) show that TransPar outperforms prior arts by non-trivial margins. Moreover, experiments demonstrate that TransPar can be integrated into the most popular deep UDA networks and be easily extended to handle any data distribution shift scenarios. Zhongyi Han, Haoliang Sun, Yilong Yin |
IEEE Trans. Image Process. | 3 |
| 2022 | Supervised Adaptive Similarity Matrix HashingabstractCompact hash codes can facilitate large-scale multimedia retrieval, significantly reducing storage and computation. Most hashing methods learn hash functions based on the data similarity matrix, which is predefined by supervised labels or a distance metric type. However, this predefined similarity matrix cannot accurately reflect the real similarity relationship among images, which results in poor retrieval performance of hashing methods, especially in multi-label datasets and zero-shot datasets that are highly dependent on similarity relationships. Toward this end, this study proposes a new supervised hashing method called supervised adaptive similarity matrix hashing (SASH) via feature-label space consistency. SASH not only learns the similarity matrix adaptively, but also extracts the label correlations by maintaining consistency between the feature and the label space. This correlation information is then used to optimize the similarity matrix. The experiments on three large normal benchmark datasets (including two multi-label datasets) and three large zero-shot benchmark datasets show that SASH has an excellent performance compared with several state-of-the-art techniques. Xiushan Nie, Xingbo Liu, Yilong Yin |
IEEE Trans. Image Process. | 5 |
| 2022 | Learning Binary Semantic Embedding for Large-Scale Breast Histology Image AnalysisabstractWith the progress of clinical imaging innovation and machine learning, the computer-assisted diagnosis of breast histology images has attracted broad attention. Nonetheless, the use of computer-assisted diagnoses has been blocked due to the incomprehensibility of customary classification models. In view of this question, we propose a novel method for Learning Binary Semantic Embedding (LBSE). In this study, bit balance and uncorrela-tion constraints, double supervision, discrete optimization and asymmetric pairwise similarity are seamlessly integrated for learning binary semantic-preserving embedding. Moreover, a fusion-based strategy is carefully designed to handle the intractable problem of parameter setting, saving huge amounts of time for boundary tuning. Based on the above-mentioned proficient and effective embedding, classification and retrieval are simultaneously performed to give interpretable image-based deduction and model helped conclusions for breast histology images. Extensive experiments are conducted on three benchmark datasets to approve the predominance of LBSE in different situations. Xingbo Liu, Xiao Kang, Xiushan Nie, Jie Guo 0012, Yilong Yin |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Regularized Two Granularity Loss Function for Weakly Supervised Video Moment RetrievalabstractWeakly supervised video moment retrieval or weakly supervised language moment retrieval aims to search the most relevant moment given a language query. In order to guide the model to capture the most matching video segments with the text description, we design a two-granularity loss function that simultaneously considers both video-level and instance-level relationships. Specifically, we first generate coarse video segments and regard each video segment as an instance. For video-level regularized multiple instance loss (MIL), we leverage the latent alignment between all intra-video segments (ie., positive bag) and text descriptions. Then, we classify these segments by regarding this procedure as a supervised learning task under noisy labels. With the instance-level regularized loss function, our model can learn to correct noisy instance-level labels so as to locate the more accurate frame boundary from all the positive instances. Comprehensive experimental results onActivityNetandDiDeModemonstrate that the proposed loss function sets a new state-of-the-art. Junya Teng, Xiankai Lu, Yongshun Gong, Xinfang Liu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 6 |
| 2022 | Embedding Hierarchical Structures for Venue Category RepresentationabstractVenue categories used in location-based social networks often exhibit a hierarchical structure, together with the category sequences derived from users’ check-ins. The two data modalities provide a wealth of information for us to capture the semantic relationships between those categories. To understand the venue semantics, existing methods usually embed venue categories into low-dimensional spaces by modeling the linear context (i.e., the positional neighbors of the given category) in check-in sequences. However, the hierarchical structure of venue categories, which inherently encodes the relationships between categories, is largely untapped. In this article, we propose a venue C ategory E mbedding M odel named Hier-CEM , which generates a latent representation for each venue category by embedding the Hier archical structure of categories and utilizing multiple types of context. Specifically, we investigate two kinds of hierarchical context based on any given venue category hierarchy and show how to model them together with the linear context collaboratively. We apply Hier-CEM to three tasks on two real check-in datasets collected from Foursquare. Experimental results show that Hier-CEM is better at capturing both semantic and sequential information inherent in venues than state-of-the-art embedding methods. Meng Chen 0003, Lei Zhu 0002, Ronghui Xu 0001, Yang Liu 0008, Xiaohui Yu 0001, Yilong Yin |
ACM Trans. Inf. Syst. | 6 |
| 2022 | $\hbox {PISEP}{^2}$: pseudo-image sequence evolution-based 3D pose prediction
Jianqin Yin, Huaping Liu 0001, Yilong Yin |
Vis. Comput. | 4 |
| 2021 | Learning Binary Semantic Embedding for Breast Histology Image Classification and RetrievalabstractWith the development of medical imaging technology and machine learning, the computer-assisted diagnosis has attracted extensive research attention, which can provide beneficial reference to pathologists. However, the exponential growth of medical images and uninterpretability of traditional classification models have hindered the applications of the computer-assisted diagnosis. To address this issues, we propose a novel method for Learning Binary Semantic Embedding (LBSE). Based on this efficient and effective embedding, classification and retrieval are performed to provide interpretable computer-assisted diagnosis for histology images. Furthermore, double supervision, bit uncorrelation and balance constraint, asymmetric strategy and discrete optimization are seamlessly integrated in the proposed method for learning binary embedding. Experiments conducted on three benchmark datasets validate the superiority of LBSE under various scenarios. Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin |
ICASSP | 4 |
| 2021 | Joint Learning of Image Aesthetic Quality Assessment and Semantic Recognition Based on Feature EnhancementabstractAesthetic quality assessment and semantic recognition are the two fundamental aspects of image perception and understanding tasks. Though these two tasks are related, most of the current research generally treats them as independent problems without any interaction. In this paper, we explore the relationships between aesthetic quality assessment and semantic recognition task, and employ a multi-task convolutional neural network with feature enhancement mechanism to effectively integrate these two tasks. A novel Enhanced Aggregation of Features Network (EAFNet) for joint learning of the two tasks is proposed to enhance the valid features and suppress the invalid features of each task in both channel and spatial dimensions. Experiments conducted on two benchmark datasets well verify the superior performance of EAFNet in handling aesthetic quality assessment and semantic recognition tasks. Xiangfei Liu, Xiushan Nie, Zhen Shen 0001, Yilong Yin |
ICASSP | 4 |
| 2021 | ECCL: Explicit Correlation-Based Convolution Boundary Locator for Moment LocalizationabstractMoment localization in videos using natural language refers to finding the most relevant segment from the video with given a query in natural language form. In this paper, we present a new boundary-determining strategy called explicit correlation-based convolution boundary locator (ECCL), which can handle any lengths of videos and moments while leveraging fine-grained matching relationships. In this method, we first train a deep network to obtain the correlation scores between video clips and query statements. Subsequently, with the correlation scores, we utilize a convolution kernel to generate the boundary probability distribution. Finally, the start and end time indexes of the video moment are calculated with an optimization problem. Experiments on two publicly available datasets demonstrate the feasibility of ECCL. Xinfang Liu, Xiushan Nie, Junya Teng, Fanchang Hao, Yilong Yin |
ICASSP | 5 |
| 2021 | Label-Guided Dictionary Pair Learning for ECG Biometric RecognitionabstractECG biometric recognition has received plenty of attention in biometrics area. In recent years, various classical sparse representation and dictionary learning methods have been utilized in ECG biometric recognition. However, to produce better classification results, lP-norm is used to regularize the representation coefficients, which undoubtedly brings time cost problem. To overcome this limitation, our method, namely label-guided dictionary pair learning, aims to learn a projective dictionary and reconstructed dictionary jointly, which achieves signal representation and reconstruction simultaneously. Introduction of label information with each dictionary item and Fisher-like regularization on projective dictionary enforce discriminability during the dictionary learning process. Alternating direction method of multipliers is then exploited to optimize the corresponding objective function. Extensive experiments on two databases demonstrate that our method can achieve better performance compared with state-of-the-art ECG biometric recognition methods. Mingzhu Ma, Gongping Yang 0001, Kuikui Wang, Yilong Yin |
ICASSP | 5 |
| 2021 | STERLING: Towards Effective ECG Biometric RecognitionabstractElectrocardiogram (ECG) biometric recognition has recently attracted considerable attention and various promising approaches have been proposed. However, due to the real nonstationary ECG noise environment, it is still challenging to perform this technique robustly and precisely. In this paper, we propose a novel ECG biometrics framework named robuSt semanTic spacE leaRning with Local sImilarity preserviNG (STERLING) to learn a latent space where ECG signals can be robustly and discriminatively represented with semantic information and local structure being preserved. Specifically, in the proposed framework, a novel loss function is proposed to learn robust semantic representation by introducing l2,1-norm loss and making full use of the supervised information. In addition, a graph regularization is imposed to preserve the local structure information in each subject. Finally, in the learnt latent space, matching can be effectively done. The experimental results on three widely-used datasets indicate that the proposed framework can outperform the state-of-the-arts. Kuikui Wang, Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
IJCB | 5 |
| 2021 | Learning Hierarchical Embedding for Video Instance SegmentationabstractIn this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability. Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin |
ACM Multimedia | 5 |
| 2021 | Deep Adaptive Attention Triple HashingabstractRecent studies have verified that learning compact hash codes can facilitate big data retrieval processing. In particular, learning the deep hash function can greatly improve the retrieval performance. However, the existing deep supervised hashing algorithm treats all the samples in the same way, which leads to insufficient learning of difficult samples. Therefore, we cannot obtain the accurate learning of the similarity relation, making it difficult to achieve satisfactory performance. In light of this, this work proposes a deep supervised hashing model, called deep adaptive attention triple hashing (DAATH), which weights the similarity prediction scores of positive and negative samples in the form of triples, thus giving different degrees of attention to different samples. Compared with the traditional triple loss, it places a greater emphasis on the difficult triple, dramatically reducing the redundant calculation. Extensive experiments have been conducted to show that DAAH consistently outperforms the state-of-the-arts, confirmed its the effectiveness. Xiushan Nie, Yilong Yin |
MMAsia | 5 |
| 2021 | Visual saliency detection by integrating spatial position prior of object with background cues
Muwei Jian, Hui Yu 0001, Guodong Wang 0001, Xianjing Meng, Lu Yang 0005, Junyu Dong, Yilong Yin |
Expert Syst. Appl. | 8 |
| 2021 | Finger vein recognition based on zone-based minutia matching
Xianjing Meng, Jinwen Zheng, Xiaoming Xi, Yilong Yin |
Neurocomputing | 5 |
| 2021 | Global context-aware multi-scale features aggregative network for salient object detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Li Lian, Zafar Ali, Imran Qureshi, Jie Guo 0012, Yilong Yin |
Neurocomputing | 8 |
| 2021 | Attention based consistent semantic learning for micro-video scene recognition
Jie Guo 0012, Xiushan Nie, Yuling Ma, Kashif Shaheed, Inam Ullah 0002, Yilong Yin |
Inf. Sci. | 6 |
| 2021 | Multi-Scale Deep Cascade Bi-Forest for Electrocardiogram Biometric Recognition
Gongping Yang 0001, Kuikui Wang, Yilong Yin |
J. Comput. Sci. Technol. | 5 |
| 2021 | Unifying neural learning and symbolic reasoning for spinal medical report generation
Zhongyi Han, Benzheng Wei, Xiaoming Xi, Bo Chen 0013, Yilong Yin, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2021 | DSFMA: deeply supervised fully convolutional neural networks based on multi-level aggregation for saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Li Lian, Hui Yu 0001, Kashif Shaheed, Yilong Yin |
Multim. Tools Appl. | 8 |
| 2021 | Multi-view discriminant analysis with sample diversity for ECG biometric recognition
Gongping Yang 0001, Kuikui Wang, Yilong Yin |
Pattern Recognit. Lett. | 4 |
| 2021 | Reinforced Short-Length HashingabstractGiven that retrieval and storage have compelling efficiency, similarity-preserving hashing has been extensively employed to approximate nearest neighbor search in large-scale image retrieval. Hash codes that are extremely compact not only can further lower the storage cost, but also accelerate the retrieval speed. However, existing methods perform poorly in retrieval based on an extremely short-length hash code, which attributes to the weak ability of classification and poor distribution of hash bit. To tackle this issue, in this study, we propose a novel reinforced short-length hashing (RSLH). In particular, this proposed method applies the mutual reconstruction between the hash representation and semantic label to retain the semantic information. Furthermore, to enhance the accuracy of hash representation, a pairwise similarity matrix is designed to make a balance between accuracy and training expenditure on memory. Besides, we integrate a parameter boosting strategy to strengthen the precision with the consideration of bit balance and uncorrelation constraints. Extensive experiments on three large-scale image benchmarks demonstrate the superior performance of RSLH under various short-length hashing scenarios. Xingbo Liu, Xiushan Nie, Qi Dai 0001, Yupan Huang, Li Lian, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Fast Unmediated Hashing for Cross-Modal RetrievalabstractCross-modal hashing is for the purpose of compressing heterogeneous multi-modal data into compact binary codes for the cross-modal retrieval, where accuracy and efficiency are two primary issues. To achieve high accuracy and efficiency, we put forward a novel method named Fast Unmediated Hashing (FUH) for cross-modal retrieval. For this method, motivated by the fact that label vector is a natural binary representation of samples for retrieval, we directly learn the cross-modal hash codes from semantic labels without any intermediate representation. This will capture more relations among different modalities, and reduce the number of variables. However, directly learning hash codes from labels would weaken the discrimination of hash codes. To address this issue, double supervision involving label information and pairwise similarity is proposed to enhance the discrimination. In addition, to decrease the training time, we present a strategy to bypass the similarity matrix-related operation in each iteration of optimization, thus some other related terms can also be computed offline to lower training complexity. Compared to several state-of-the-art techniques on three public datasets, the experimental results have manifested the superiority of FUH concerning efficiency and accuracy. Xiushan Nie, Xingbo Liu, Xiaoming Xi, Chenglong Li 0004, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Deep Multiscale Fusion Hashing for Cross-Modal RetrievalabstractOwing to the rapid development of deep learning and the high efficiency of hashing, hashing methods based on deep learning models have been extensively adopted in the area of cross-modal retrieval. In general, in existing deep model-based methods, modality-specific features play an important role during the hash learning. However, most existing methods only use the modality-specific features from the final fully connected layer, ignoring the semantic relevance among modality-specific features with different scales in multiple layers. To address this issue, in this study, we put forward an end-to-end deep hashing method called deep multiscale fusion hashing (DMFH) for cross-modal retrieval. For the proposed DMFH, we first design different network branches for two modalities and then adopt multiscale fusion models for each branch network to fuse the multiscale semantics, which can be used to explore the semantic relevance. Furthermore, the multi-fusion models also embed the multiscale semantics into the final hash codes, making the final hash codes more representative. In addition, the proposed DMFH can learn common hash codes directly without a relaxation, thereby avoiding a loss in accuracy during hash learning. Experimental results on three benchmark datasets prove the relative superiority of the proposed method. Xiushan Nie, Bowei Wang, Fanchang Hao, Muwei Jian, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Normality Learning in Multispace for Video Anomaly DetectionabstractVideo anomaly detection is a challenging task owing to the rare and diverse nature of abnormal events. However, most of the existing methods only learn the normality in a single space, focusing on low-level detailed features, which is easily affected by unimportant pixels. To address this issue, in this study, we propose a semi-supervised method based on the generative adversarial network and frame prediction, wherein the normality is learned in both the original image space and latent space, and the events deviating from the normality are detected as anomalies. In particular, given a video clip, we first predict a future frame and minimize the prediction errors between the generated frame and its ground truth. Thereafter, we encode the predicted frames and their ground truths in the latent space and minimize their differences. In the testing phase, we calculate the normal scores of each frame in both the image and latent spaces to obtain a comprehensive evaluation. Utilizing the multispace can capture more normality distribution information of the data, which can benefit anomaly detection. The results of experiments on three benchmark datasets demonstrate the effectiveness of the proposed method. Xiushan Nie, Rundong He, Meng Chen 0003, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Learning Joint and Specific Patterns: A Unified Sparse Representation for Off-the-Person ECG Biometric RecognitionabstractDevices such as smartphones and tablets have spurred interest in off-the-person electrocardiogram (ECG) biometric recognition. While the advantage of using multi-feature information for establishing identities has been widely recognized, computational sparse representation models for multi-feature biometric recognition have only recently received more attention. We propose a unified sparse representation framework which collaboratively exploits joint and specific patterns for ECG biometric recognition. In particular, unlike joint sparse representation, which only considers the consistency among sparsity patterns of multiple features, we combine the consistent and pairwise constraints, which not only learn latent discriminant representations for all features but capture the interactions between them. In addition, our framework is universal and easily adapts to other multi-feature sparse representation models by just tuning the regularization parameters. The optimization problem is solved by an efficient alternating direction method of multipliers (ADMM). Extensive experiments on two publicly available off-the-person datasets demonstrate that our method can achieve competitive or even superior performance compared to state-of-the-art ECG biometric recognition methods. Gongping Yang 0001, Kuikui Wang, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Finger Vein Recognition via Sparse Reconstruction Error Constrained Low-Rank RepresentationabstractVein pattern-based methods have powerfully promoted the performance of finger vein recognition. However, it is not easy to precisely extract vein patterns from images, especially from low-quality images, and the non-vein area have been proved to be helpful for recognition. This paper proposes to use low-rank representation to extract as much noiseless discriminative information as possible from finger vein images. However, image deformation and image quality variations weaken the correlation of genuine images, and therefore damage the low-rank linear representation. To further deal with this problem, the class labels of training images and the local geometric structure between testing images and training images, reflected by sparse reconstruction errors of testing images, are used as constraints of low-rank coefficients. In particular, vein backbone decomposition based sparse representation is proposed to fast compute the deformation-robust reconstruction errors of each testing image. The reconstruction errors on sub-backbones of one training image are summed and modified as the constraint of the low-rank coefficient on this training image. We evaluate the proposed method on three widely used finger vein databases, and experimental results show that the proposed method performs well in finger vein recognition. Lu Yang 0005, Gongping Yang 0001, Kuikui Wang, Fanchang Hao, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Correction to "Finger Vein Code: From Indexing to Matching"abstractIn second paragraph of the footnote on the first page of[1], the institution information of Lu Yang and Xiaoming Xi is inaccurate. The correct institution name is “School of Computer Science and Technology, Shandong University of Finance and Economics.” So this paragraph should be corrected as: Lu Yang 0005, Gongping Yang 0001, Xiaoming Xi, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Deep Hashing With Weighted Spatial ImportanceabstractHashing method has been widely used in big data retrieval because of its low computational complexity. Most of existing hashing methods learn the final hash code from the semantic information of the whole image. However, different spatial regions of an image have different influences during the hash learning. To tackle this issue, we propose a new deep hashing with weighted spatial importance (DWSH) in this paper. Specifically, the proposed DWSH first utilizes a spatial attention model to learn the importance of different spatial regions in the original image, and then assigns different weights to these spatial regions according to their importance. The final hash codes are learned based on the weighted spatial information. In addition, two strategies are designed to utilize the spatial importance, including discrete weight strategy and continuous weight strategy, which weight the spatial information with discrete and continuous values, respectively. The results of extensive experiments conducted on three benchmark datasets show that the proposed DWSH method is superior to the state-of-the-art hashing method based on different evaluation protocols. Xiushan Nie, Meng Chen 0003, Li Lian, Yilong Yin |
IEEE Trans. Multim. | 5 |
| 2021 | Single-shot Semantic Matching Network for Moment Localization in VideosabstractMoment localization in videos using natural language refers to finding the most relevant segment from videos given a natural language query. Most of the existing methods require video segment candidates for further matching with the query, which leads to extra computational costs, and they may also not locate the relevant moments under any length evaluated. To address these issues, we present a lightweight single-shot semantic matching network (SSMN) to avoid the complex computations required to match the query and the segment candidates, and the proposed SSMN can locate moments of any length theoretically. Using the proposed SSMN, video features are first uniformly sampled to a fixed number, while the query sentence features are generated and enhanced by GloVe, long-term short memory (LSTM), and soft-attention modules. Subsequently, the video features and sentence features are fed to an enhanced cross-modal attention model to mine the semantic relationships between vision and language. Finally, a score predictor and a location predictor are designed to locate the start and stop indexes of the query moment. We evaluate the proposed method on two benchmark datasets and the experimental results demonstrate that SSMN outperforms state-of-the-art methods in both precision and efficiency. Xinfang Liu, Xiushan Nie, Junya Teng, Li Lian, Yilong Yin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Focusing on Detail: Deep Hashing Based on Multiple Region Details (Student Abstract)abstractFast retrieval efficiency and high performance hashing, which aims to convert multimedia data into a set of short binary codes while preserving the similarity of the original data, has been widely studied in recent years. Majority of the existing deep supervised hashing methods only utilize the semantics of a whole image in learning hash codes, but ignore the local image details, which are important in hash learning. To fully utilize the detailed information, we propose a novel deep multi-region hashing (DMRH), which learns hash codes from local regions, and in which the final hash codes of the image are obtained by fusing the local hash codes corresponding to local regions. In addition, we propose a self-similarity loss term to address the imbalance problem (i.e., the number of dissimilar pairs is significantly more than that of the similar ones) of methods based on pairwise similarity. Xiushan Nie, Xingbo Liu, Yilong Yin |
AAAI | 5 |
| 2020 | Discrete Spatial Importance-Based Deep Weighted Hashing
Xiushan Nie, Xiaoming Xi, Yilong Yin |
ACCV (3) | 5 |
| 2020 | Deep Adaptive Feature Aggregation in Multi-task Convolutional Neural NetworksabstractConvolutional Neural Network (CNN) based multi-task learning methods have been widely used in a variety of applications of computer vision. Towards effective multi-task CNN architectures, recent studies automatically learn the optimal combinations of task-specific features at single network layers. However, they generally construct an unchanged operation of feature aggregation after training, regardless of the characteristics of input features. In this paper, we propose a novel Adaptive Feature Aggregation (AFA) layer for multi-task CNNs, in which a dynamic aggregation mechanism is designed to allow each task to adaptively determine the degree to which the feature aggregation of different tasks is needed according to the feature dependencies. On both pixel-level and image-level tasks, we demonstrate that our approach significantly outperforms the previous state-of-the-art methods of multi-task CNNs. Zhen Shen 0001, Chaoran Cui, Jian Zong, Meng Chen 0003, Yilong Yin |
CIKM | 6 |
| 2020 | Behavior-driven Student Performance Prediction with Tri-branch Convolutional Neural NetworkabstractStudent performance prediction aims to leverage student-related information to predict their future academic outcomes, which may be beneficial to numerous educational applications, such as personalized teaching and academic early warning. In this paper, we seek to address the problem by analyzing students' daily studying and living behavior, which is comprehensively recorded via campus smart cards. Different from previous studies, we propose an end-to-end student performance prediction model, namely Tri-branch CNN, which is equipped with three types of convolutional filters, i.e., the row-wise convolution, column-wise convolution, and group-wise convolution, to effectively capture the duration, periodicity, and location-aware characteristic of student behavior, respectively. We also introduce the attention mechanism and cost-sensitive learning strategy to further improve the accuracy of our approach. Extensive experiments on a large-scale real-world dataset demonstrate the potential of our approach for student performance prediction. Jian Zong, Chaoran Cui, Yuling Ma, Meng Chen 0003, Yilong Yin |
CIKM | 6 |
| 2020 | Learning Multi-Scale Attentive Features for Series Photo SelectionabstractPeople used to take a series of nearly identical photos about the same subject, but it is usually a tedious chore to select the reversed ones from them. Despite the remarkable progress, most existing studies on image aesthetics assessment fail to fulfill the task of series photo selection. In this paper, we develop a novel deep CNN architecture that aggregates multi-scale features from different network layers, in order to capture the subtle differences between series photos. To reduce the risk of redundant or even interfering features, we introduce the spatial-channel self-attention mechanism to adaptively recalibrate the features at each layer, so that informative features can be selectively emphasized and less useful ones suppressed. Extensive experiments on a benchmark dataset well demonstrate the potential of our approach for series photo selection. Chaoran Cui, Chunyun Zhang, Zhen Shen 0001, Yilong Yin |
ICASSP | 6 |
| 2020 | Deep Multi-Region HashingabstractHashing has been widely used for large-scale approximate nearest neighbors retrieval own to its high efficiency. In the existing hashing methods, deep supervised hashing methods have achieved the best performance by utilizing the semantic labels on data with deep learning. However, most of these methods only consider the semantics of whole image but ignore the local information which contains much more semantic details. Evidently, the semantic details are beneficial for hash learning. To address this issue, in this paper, we proposed a novel Deep Multi-Region Hashing (DMRH) method to fully utilize the semantic details, which uses overlapping N × N regions of an image to learn N2hash codes for getting a final hash code. Extensive experimental results with three datasets show that DMRH can achieve state-of-the-art performance. Xiushan Nie, Xingbo Liu, Yilong Yin |
ICASSP | 5 |
| 2020 | Learning to Learn Kernels with Variational Random FeaturesabstractWe introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the latent variable. We formulate the optimization of MetaVRF as a variational inference problem by deriving an evidence lower bound under the meta-learning framework. To incorporate shared knowledge from related tasks, we propose a context inference of the posterior, which is established by an LSTM architecture. The LSTM-based inference network can effectively integrate the context information of previous tasks with task-specific information, generating informative and adaptive features. The learned MetaVRF can produce kernels of high representational power with a relatively low spectral sampling rate and also enables fast adaptation to new tasks. Experimental results on a variety of few-shot regression and classification tasks demonstrate that MetaVRF delivers much better, or at least competitive, performance compared to existing meta-learning alternatives. Xiantong Zhen, Haoliang Sun, Yingjun Du, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek |
ICML | 5 |
| 2020 | Multi-Scale and Attention based ResNet for Heartbeat ClassificationabstractThis paper presents a novel deep learning framework for the electrocardiogram (ECG) heartbeat classification. Although there have been some studies with excellent overall accuracy, these studies have not been very accurate in the diagnosis of arrhythmia classes especially such as supraventricular ectopic beat (SVEB) and ventricular ectopic beat (VEB). In our work, we propose a Multi-Scale and Attention based Res Net for heartbeat classification in intra-patient and inter-patient paradigms respectively. Firstly, we extract shallow features from a convolutional layer. Secondly, the shallow features are sent into three branches with different convolution kernels in order to combine receptive fields of different sizes. Finally, fully connected layers are used to classify the heartbeat. Besides, we design a new attention mechanism based on the characteristics of heartbeat data. At last, extensive experiments on benchmark dataset demonstrate the effectiveness of our proposed model. Gongping Yang 0001, Yilong Yin |
ICPR | 5 |
| 2020 | Towards Accurate and Robust Domain Adaptation under Noisy EnvironmentsabstractIn non-stationary environments, learning machines usually confront the domain adaptation scenario where the data distribution does change over time. Previous domain adaptation works have achieved great success in theory and practice. However, they always lose robustness in noisy environments where the labels and features of examples from the source domain become corrupted. In this paper, we report our attempt towards achieving accurate noise-robust domain adaptation. We first give a theoretical analysis that reveals how harmful noises influence unsupervised domain adaptation. To eliminate the effect of label noise, we propose an offline curriculum learning for minimizing a newly-defined empirical source risk. To reduce the impact of feature noise, we propose a proxy distribution based margin discrepancy. We seamlessly transform our methods into an adversarial network that performs efficient joint optimization for them, successfully mitigating the negative influence from both data corruption and distribution shift. A series of empirical studies show that our algorithm remarkably outperforms state of the art, over 10% accuracy improvements in some domain adaptation tasks under noisy environments. Zhongyi Han, Xian-Jin Gui, Chaoran Cui, Yilong Yin |
IJCAI | 4 |
| 2020 | CFVMNet: A Multi-branch Network for Vehicle Re-identification Based on Common Field of ViewabstractVehicle re-identification (re-ID) aims to retrieve the image of the same vehicles across multiple cameras. It has attracted wide attention in the field of computer vision owing to the deployment of surveillance system. However, some unfavorable factors restrict the retrieval accuracy of re-ID; minor inter-class difference and orientation variation are two main issues. In this study, we proposed a multi-branch network based on common field of view (CFVMNet) to address these issues. In the proposed method, we extracted and fused the global and local detail features using four branches and the Batch DropBlock (BDB) strategy to accentuate inter-class difference. We also considered some other attributes (i.e., color, type, and model) in the feature extraction process to make the final features more recognizable. For the issue of orientation variation that could lead to large intra-class difference, we learned two different metrics according to whether there is common field of view of two vehicle images, respectively, which can enable the proposed CFVMNet to focus on different regions. Extensive experiments on two public datasets, VeRi-776 and VehicleID, show that the proposed method outperformed the state-of-the-art approaches to vehicle re-ID. Ziruo Sun, Xiushan Nie, Xiaoming Xi, Yilong Yin |
ACM Multimedia | 4 |
| 2020 | Recursive narrative alignment for movie narrating
Zhongyi Han, Hongbo Wu, Benzheng Wei, Yilong Yin, Shuo Li 0001 |
Sci. China Inf. Sci. | 4 |
| 2020 | Multi-task MIML learning for pre-course student performance prediction
Yuling Ma, Chaoran Cui, Jie Guo 0012, Gongping Yang 0001, Yilong Yin |
Frontiers Comput. Sci. | 6 |
| 2020 | Erratum to: Cuckoo search with varied scaling factor
Lijin Wang, Yilong Yin, Yiwen Zhong |
Frontiers Comput. Sci. | 2 |
| 2020 | Personalized image quality assessment with Social-Sensed aesthetic preference
Chaoran Cui, Wenya Yang, Meng Wang 0001, Xiushan Nie, Yilong Yin |
Inf. Sci. | 6 |
| 2020 | Saliency detection using multiple low-level priors and a propagation mechanism
Muwei Jian, Junyu Dong, Chaoran Cui, Xiushan Nie, Yilong Yin |
Multim. Tools Appl. | 6 |
| 2020 | Short Term ECG Classification with Residual-Concatenate Network and Metric Learning
Xinjing Song, Gongping Yang 0001, Kuikui Wang, Yilong Yin |
Multim. Tools Appl. | 6 |
| 2020 | A brief survey of visual saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Hui Yu 0001, Xing Wang 0002, Yilong Yin |
Multim. Tools Appl. | 7 |
| 2020 | Modality correlation-based video summarization
Xingrun Wang, Xiushan Nie, Xingbo Liu, Binze Wang, Yilong Yin |
Multim. Tools Appl. | 5 |
| 2020 | Multi-scale differential feature for ECG biometrics with collective matrix factorization
Kuikui Wang, Gongping Yang 0001, Yilong Yin |
Pattern Recognit. | 4 |
| 2020 | Robust ECG biometrics using GNMF and sparse representation
Gongping Yang 0001, Kuikui Wang, Yilong Yin |
Pattern Recognit. Lett. | 6 |
| 2020 | Structural sparse representation with class-specific dictionary for ECG biometric recognition
Jingxiao Xu, Gongping Yang 0001, Kuikui Wang, Yilong Yin |
Pattern Recognit. Lett. | 6 |
| 2020 | Model Optimization Boosting Framework for Linear Model Hash LearningabstractEfficient hashing techniques have attracted extensive research interests in both storage and retrieval of highdimensional data, such as images and videos. In existing hashing methods, a linear model is commonly utilized owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider the inherent characteristics and neighborhood information of samples. Differing from existing hashing methods, in this study, we propose a self-improvement framework called Model Boost (MoBoost) to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, for a linear-based hashing method, we first repeatedly execute the hashing method to obtain several hash codes to training samples. Then, utilizing two novel fusion strategies, these codes are fused into a single set. We also propose two new criteria to evaluate the goodness of hash bits during the fusion process. Based on the fused set of hash codes, we learn new parameters for the linear hash function that can significantly improve the accuracy. In general, the proposed MoBoost can be adopted by existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods, and adopting the proposed MoBoost will incur negligible time and space costs. To evaluate the proposed MoBoost, we performed extensive experiments on four benchmark datasets, and the results demonstrate superior performance. Xingbo Liu, Xiushan Nie, Liqiang Nie, Yilong Yin |
IEEE Trans. Image Process. | 5 |
| 2020 | Joint Multi-View Hashing for Large-Scale Near-Duplicate Video RetrievalabstractMulti-view hashing can well support large-scale near-duplicate video retrieval, due to its desirable advantages of mutual reinforcement of multiple features, low storage cost, and fast retrieval speed. However, there are still two limitations that impede its performance. First, existing methods only consider local structures in multiple features. They ignore the global structure that is important for near-duplicate video retrieval, and cannot fully exploit the dependence and complementarity of multiple features. Second, existing works always learn hashing functions bit by bit, which unfortunately increases the time complexity of hash function learning. In this paper, we propose a supervised hashing scheme, termed as joint multi-view hashing (JMVH), to address the aforementioned problems. It jointly preserves the global and local structures of multiple features while learning hashing functions efficiently. Specially, JMVH considers features of video as items, based on which an underlying Hamming space is learned by simultaneously preserving their local and global structures. In addition, a simple but efficient multi-bit hash function learning based on generalized eigenvalue decomposition is devised to learn multiple hash functions within a single step. It can significantly reduce the time complexity of conventional hash function learning processes that sequentially learn multiple hash functions bit by bit. The proposed JMVH is evaluated on two public databases: CC_WEB_VIDEO and UQ_VIDEO. Experimental results demonstrate that the proposed JMVH achieves more than a 5 percent improvement compared to several state-of-the-art methods which indicates the superior performance of JMVH. Xiushan Nie, Weizhen Jing, Chaoran Cui, Chen Zhang 0013, Lei Zhu 0002, Yilong Yin |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Learning the Traditional Art of Chinese Calligraphy via Three-Dimensional Reconstruction and AssessmentabstractThe traditional art of Chinese calligraphy, reflecting the wisdom of the grass-roots community, is the soul of Chinese culture. Just like many other types of craftsmanship, it is part of the historical heritage and is worth conserving, from generation to generation. Since the movements of an ink brush are in a 3D style when Chinese calligraphy is written, they embody “The Power of Beauty,” comprising various reflectance properties and rough-surface geometry. To truly understand the powerful significance and beauty of the art of Chinese calligraphy, in this paper, a 3D calligraphy reconstruction method, based on Photometric Stereo, is designed to capture the detailed appearance of the calligraphy's 3D surface geometry. For assessment, an Iterative Closest Point (ICP) algorithm is applied for registration of 3D intrinsic shapes between the Chinese calligraphy and the calligraphy fans' handwriting. Through matching these two sets of calligraphy characters, the designed system can give a score to the handwriting of a user. Experiments have been performed on Chinese calligraphy from different historical dynasties to evaluate the effectiveness of the proposed scheme, and experimental results show that the developed system is useful and provides a convenient method of calligraphy appreciation and assessment. Muwei Jian, Junyu Dong, Maoguo Gong, Hui Yu 0001, Liqiang Nie, Yilong Yin, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Social-sensed Image Aesthetics AssessmentabstractImage aesthetics assessment aims to endow computers with the ability to judge the aesthetic values of images, and its potential has been recognized in a variety of applications. Most previous studies perform aesthetics assessment purely based on image content. However, given the fact that aesthetic perceiving is a human cognitive activity, it is necessary to consider users’ perception of an image when judging its aesthetic quality. In this article, we regard users’ social behavior as the reflection of their perception of images and harness these additional clues to improve image aesthetics assessment. Specifically, we first merge the raw social interactions between users and images into clusters as the social labels of images, so the collective social behavioral information associated with an image can be well represented over a structured and compact space. Then, we develop a novel deep multi-task network to jointly learn social labels in different modalities from social images and apply it to common web images. In this manner, our approach is readily generalized to web images without social behavioral information. Finally, we introduce a high-level fusion sub-network to the aesthetics model, in which the social and visual representations of images are well balanced for aesthetics assessment. Experimental results on two benchmark datasets well verify the effectiveness of our approach and highlight the benefits of different types of social behavioral information for image aesthetics assessment. Chaoran Cui, Peiguang Lin, Xiushan Nie, Muwei Jian, Yilong Yin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Jointly Multiple Hash LearningabstractHashing can compress heterogeneous high-dimensional data into compact binary codes while preserving the similarity to facilitate efficient retrieval and storage, and thus hashing has recently received much attention from information retrieval researchers. Most of the existing hashing methods first predefine a fixed length (e.g., 32, 64, or 128 bit) for the hash codes before learning them with this fixed length. However, one sample can be represented by various hash codes with different lengths, and thus there must be some associations and relationships among these different hash codes because they represent the same sample. Therefore, harnessing these relationships will boost the performance of hashing methods. Inspired by this possibility, in this study, we propose a new model jointly multiple hash learning (JMH), which can learn hash codes with multiple lengths simultaneously. In the proposed JMH method, three types of information are used for hash learning, which come from hash codes with different lengths, the original features of the samples and label. In contrast to the existing hashing methods, JMH can learn hash codes with different lengths in one step. Users can select appropriate hash codes for their retrieval tasks according to the requirements in terms of accuracy and complexity. To the best of our knowledge, JMH is one of the first attempts to learn multi-length hash codes simultaneously. In addition, in the proposed model, discrete and closed-form solutions for variables can be obtained by cyclic coordinate descent, thereby making the proposed model much faster during training. Extensive experiments were performed based on three benchmark datasets and the results demonstrated the superior performance of the proposed method. Xingbo Liu, Xiushan Nie, Yingxin Wang, Yilong Yin |
AAAI | 4 |
| 2019 | MoBoost: A Self-improvement Framework for Linear-based HashingabstractThe linear model is commonly utilized in hashing methods owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider neighborhood information. In this study, we propose a novel generalized framework called Model Boost (MoBoost), which can achieve the self-improvement of the linear-based hashing. The proposed MoBoost is used to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, given a linear-based hashing method, we first execute the method several times to get several different hash codes for training samples, and then combine these different hash codes into one set utilizing one novel fusion strategy. Based on this set of hash codes, we learn some new parameters for the linear hash function that can significantly improve accuracy. The proposed MoBoost can be generally adopted in existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods while imposing negligible added expenditure in terms of time and space. Extensive experiments are performed based on three benchmark datasets, and the results demonstrate the superior performance of the proposed framework. Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin |
CIKM | 5 |
| 2019 | Variable-Length Quantization Strategy for HashingabstractHashing is widely used to solve fast Approximate Nearest Neighbor (ANN) search problems, involves converting the original real-valued samples to binary-valued representations. The conventional quantization strategies, such as Single-Bit Quantization and Multi-Bit quantization, are considered ineffective, because of their serious information loss. To address this issue, we propose a novel variable-length quantization (VLQ) strategy for hashing. In the proposed VLQ technique, we divide all samples into different regions in each dimension firstly given the real-valued features of samples. Then we compute the dispersion degrees of these regions. Subsequently, we attempt to optimally assign different number of bits to each dimensions to obtain the minimum dispersion degree. Our experiments show that the VLQ strategy achieves not only superior performance over the state-of-the-art methods, but also has a faster retrieval speed on public datasets. Xiushan Nie, Xiaoming Xi, Yilong Yin |
ICIP | 5 |
| 2019 | Learning the Set Graphs: Image-Set Classification Using Sparse Graph Convolutional NetworksabstractImage-set classification has recently made great progress in computer vision. Compared with traditional image classification tasks, set-based classification exhibits great challenges due to huge intra-class variability and high inter-class ambiguity. In this paper, we propose to model the image set as a graph and formulate image set classification as the graph matching task. Without relying on the strong structure assumption, we build the first end-to-end graph convolutional network, the Deep SetNet, to learn the graph structure of an image set. Specifically, the SetNet consists of one convolutional network (CNN) to sufficiently extract the discriminative vertex, one graph convolutional Network (GCN) to faithfully learn the substructure in set graphs and graph pooling layers to aggregate the vertex features from the GCN. Moreover, we propose imposing the ℓ1,2-norm based sparsity constraint to select vertex features, which largely improves the model generalization capability. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art methods, showing its great effectiveness in set-based image classification. Haoliang Sun, Xiantong Zhen, Yilong Yin |
ICIP | 3 |
| 2019 | Towards Unified Aesthetics and Emotion Prediction in ImagesabstractAesthetics assessment and emotion recognition are two fundamental problems in user perception understanding. While the two tasks are correlated and mutually beneficial, they are usually solved separately in existing studies. In this paper, we resort to multi-task learning to deal with aesthetics assessment and emotion recognition for images in a unified framework. Towards this goal, we extend a large scale emotion dataset by further manually rating the aesthetic qualities of images. To our best knowledge, the new dataset is the first collection of images that are associated with both aesthetic and emotional labels. Besides, we present a novel Aesthetics-Emotion hybrid Network (AENet) for multi-task learning on aesthetics assessment and emotion recognition. Task-specific and shared features have been explicitly separated by different network streams, and effectively fused at multiple network levels. Experiments on our new and benchmark datasets verify the effectiveness of our approach for unified aesthetics and emotion prediction. Chaoran Cui, Leilei Geng, Yuling Ma, Yilong Yin |
ICIP | 5 |
| 2019 | Supervised Short-Length HashingabstractHashing can compress high-dimensional data into compact binary codes, while preserving the similarity, to facilitate efficient retrieval and storage. However, when retrieving using an extremely short length hash code learned by the existing methods, the performance cannot be guaranteed because of severe information loss. To address this issue, in this study, we propose a novel supervised short-length hashing (SSLH). In this proposed SSLH, mutual reconstruction between the short-length hash codes and original features are performed to reduce semantic loss. Furthermore, to enhance the robustness and accuracy of the hash representation, a robust estimator term is added to fully utilize the label information. Extensive experiments conducted on four image benchmarks demonstrate the superior performance of the proposed SSLH with short-length hash codes. In addition, the proposed SSLH outperforms the existing methods, with long-length hash codes. To the best of our knowledge, this is the first linear-based hashing method that focuses on both short and long-length hash codes for maintaining high precision. Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin |
IJCAI | 6 |
| 2019 | Supervised Discrete Hashing With Mutual Linear RegressionabstractSupervised linear hashing can compress high-dimensional data into compact binary codes owing to its efficiency. Generally, the relation between label and hash codes is widely used in the existing hashing methods because of its effectiveness of improving the accuracy. The existing hashing methods always use two different projections to represent the mutual regression between hash codes and class labels. In contrast to the existing methods, we propose a novel learning-based hashing method termed supervised discrete hashing with mutual linear regression (SDHMLR) in this study, where only one stable projection is used to describe the linear correlation between hash codes and corresponding labels. To the best of our knowledge, this strategy has not been used for hashing previously. In addition, we further use a boosting strategy to improve the final performance of the proposed method without adding extra constraints and with little extra expenditure in terms of time and space. Extensive experiments conducted on three image benchmarks demonstrate the superior performance of the proposed method. Xingbo Liu, Xiushan Nie, Yilong Yin |
ACM Multimedia | 4 |
| 2019 | Visual Urban Perception with Deep Semantic-Aware Network
Yongchao Xu, Qizheng Yang, Chaoran Cui, Guangle Song, Xiaohui Han, Yilong Yin |
MMM (2) | 7 |
| 2019 | Dual Path Convolutional Neural Network for Student Performance Prediction
Yuling Ma, Jian Zong, Chaoran Cui, Chunyun Zhang, Qizheng Yang, Yilong Yin |
WISE | 6 |
| 2019 | Anchor-based manifold binary pattern for finger vein recognition
Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
Sci. China Inf. Sci. | 5 |
| 2019 | Pre-course student performance prediction with multi-instance multi-label learning
Yuling Ma, Chaoran Cui, Xiushan Nie, Gongping Yang 0001, Kashif Shaheed, Yilong Yin |
Sci. China Inf. Sci. | 6 |
| 2019 | Non-negative locality-constrained vocabulary tree for finger vein image retrieval
Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
Frontiers Comput. Sci. | 5 |
| 2019 | Learning personalized binary codes for finger vein recognition
Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
Neurocomputing | 4 |
| 2019 | Human identification using finger vein and ECG signals
Gongping Yang 0001, Lu Yang 0005, Dunfeng Li, Yilong Yin |
Neurocomputing | 7 |
| 2019 | Multi-view face hallucination using SVD and a mapping model
Muwei Jian, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Liqiang Nie, Yilong Yin |
Inf. Sci. | 6 |
| 2019 | Joint multi-label classification and label correlations with missing labels and feature selection
Zhifen He, Ming Yang 0014, Yang Gao 0001, Hui-Dong Liu, Yilong Yin |
Knowl. Based Syst. | 5 |
| 2019 | Automated segmentation of choroidal neovascularization in optical coherence tomography images using multi-scale convolutional neural networks with structure prior
Xiaoming Xi, Xianjing Meng, Lu Yang 0005, Xiushan Nie, Gongping Yang 0001, Haoyu Chen 0002, Yilong Yin, Xinjian Chen 0001 |
Multim. Syst. | 8 |
| 2019 | Binary feature representation learning for scene retrieval in micro-video
Jie Guo 0012, Xiushan Nie, Muwei Jian, Yilong Yin |
Multim. Tools Appl. | 4 |
| 2019 | Assessment of feature fusion strategies in visual attention mechanism for saliency detection
Muwei Jian, Chaoran Cui, Xiushan Nie, Hanjiang Luo, Yilong Yin |
Pattern Recognit. Lett. | 7 |
| 2019 | Learning binary hash codes for finger vein image retrieval
Gongping Yang 0001, Lu Yang 0005, Dunfeng Li, Yilong Yin |
Pattern Recognit. Lett. | 6 |
| 2019 | One-Shot SADI-EPE: A Visual Framework of Event Progress EstimationabstractIn many practical engineering applications, the number of actions that have been finished should be known, particularly for an untrimmed video sequence that includes an event with a series of actions, it is important to know the number of actions that have been finished. In this paper, we termed this process as visual event progress estimation (EPE). However, the research related to this problem is few in the research community. To solve this problem, a visual human action analysis-based framework, namely one-shot simultaneously action detection and identification (SADI)-EPE, is presented in this paper. The visual EPE is modeled as an online one-shot learning-based problem; sliding window and attention-based bag of key poses formulate our framework. Unlike most of the action analysis methods relying on a number of training data of some predefined classes, our method can realize SADI for any event if one sample of the event is given, which makes it feasible for practical applications. At the same time, not only SADI but also the progress estimation of the event can be realized by our algorithm. In terms of methodology, the key pose is defined by an invariant pose descriptor from skeletal data and silhouette data. Moreover, in order to extract representative and discriminative poses from one training sample, we present a new bidirectional kNN-based attention weighted key pose selection method, which can filter the unrelated actions and model different importance of various key poses. In addition, an attention-based multi-modal fusion scheme, which addresses the difficulty of high-dimensional features and few training samples, is proposed to augment the performance of our algorithm. Finally, we propose an evaluation criterion for the estimation problem. Extensive results demonstrated the efficacy of our proposed framework. Jianqin Yin, Fuchun Sun 0001, Huaping Liu 0001, Bin Wang 0045, Jun Liu 0007, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2019 | Finger Vein Code: From Indexing to MatchingabstractVein pattern-based methods powerfully boost the recognition accuracy of finger veins, but real-time recognition cannot be guaranteed, especially in large-scale applications. Moreover, previous studies focused on either the matching task to enhance the accuracy or the indexing task to improve the efficiency. This paper proposes a finger vein code indexing method and combines it with a finger vein pattern matching method into an integration framework for improving both accuracy and efficiency. With the extracted vein patterns, the direction of each vein segment is detected and represented by the elliptical direction map as a feature for indexing, which will be encoded into a binary code by the angle K-means. The similarity between vein direction codes is measured by the grouped hamming distance in indexing, and further weighted by the overlap degree of the corresponding vein patterns to return the candidates for the probe. In addition, based on the above distance measurement, only vein segments with the same direction code are considered in following probe-to-candidate matching. Experimental results indicate that our indexing method outperforms the state-of-the-art methods and has competitive potential in performing the matching task. The results also indicate that the integration framework highly improves the identification efficiency with a slight improvement on the accuracy. Lu Yang 0005, Gongping Yang 0001, Xiaoming Xi, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2019 | Distribution-Oriented Aesthetics Assessment With Semantic-Aware Hybrid NetworkabstractImage aesthetics assessment has emerged as a hot topic in recent years due to its potential in numerous high-level vision applications. In this paper, distinguished from existing studies relying on a single label, we propose quantifying image aesthetics by a distribution over multiple quality levels. The distribution-based representation characterizes the disagreement among users' aesthetic preferences regarding the same image, and is also compatible with the traditional task of aesthetic label prediction. Our framework is developed based on fully convolutional networks and enables inputs of varying sizes. In this way, we circumvent the fixed-size constraint of prevalent convolutional neural networks, and avoid the risk of impairing the intrinsic aesthetic appeal of images. Moreover, given the fact that aesthetic perceiving is coupled with semantic understanding, we present a novel semantic-aware hybrid NEtwork (SANE), which harvests the information from object categorization and scene recognition to enhance image aesthetics assessment. Experiments on two benchmark datasets have well verified the effectiveness of our approach in both scenarios of aesthetic distribution prediction and aesthetic label prediction, and highlighted the benefits of input preserving as well as semantic understanding for images. Chaoran Cui, Tao Lian, Liqiang Nie, Lei Zhu 0002, Yilong Yin |
IEEE Trans. Multim. | 6 |
| 2019 | Global-view hashing: harnessing global relations in near-duplicate video retrieval
Weizhen Jing, Xiushan Nie, Chaoran Cui, Xiaoming Xi, Gongping Yang 0001, Yilong Yin |
World Wide Web | 6 |
| 2018 | Deep Correlation Structure Preserved Label Space Embedding for Multi-label ClassificationabstractLabel embedding is an effective and efficient method which can jointly extract the information of all labels for better performance of multi-label classification. However, most existing embedding methods ignore information of feature space or intrinsic structure of previous label space, such that their learned latent space will not have strong predictability and discriminant ability. We propose a novel deep neural network (DNN) based model, namely Deep Correlation Structure Preserved Label Space Embedding (DCSPE). Specifically, DCSPE derives a deep latent space by performing feature-aware label space embedding with deep canonical correlation analysis (DCCA) and preserving the intrinsic structure of the previous label space with proposed deep multidimensional scaling (DMDS). Our DCSPE is achieved by integrating the DNN architectures of the two DNN based models and can learn a feature-aware structure preserved deep latent space. Furthermore, extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms. Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin |
ACML | 4 |
| 2018 | Modality-Specific Structure Preserving Hashing for Cross-Modal RetrievalabstractHashing-based methods have made great advancements in cross-modal retrieval in both computational efficiency and storage. Learning a common space from different modalities is the common strategy of hashing-based methods, however, relational and structural information between samples in each modality, namely, a modality-specific structure, is always discarded during learning. In addition, cross-modality samples sometimes suffer from inter-class ambiguity and intra-class variability because of the uncertainty of manual labeling. To address these issues, we propose a novel method named Modality-specific structure Preserving Hashing (MsPH), which learns hashes by preserving the local structure and relations between samples in each modality. Moreover, label enhancement is utilized in MsPH to address label ambiguity and variability. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of MsPH under various cross-modal scenarios. Xingbo Liu, Haoliang Sun, Xiushan Nie, Chaoran Cui, Yilong Yin |
ICASSP | 5 |
| 2018 | Structural Compact Core Tensor Dictionary Learning for Multispec-Tral Remote Sensing Image DeblurringabstractThe multispectral remote sensing image (MS-RSI) is blurred existing multispectral camera due to various hardware limitations. In this paper, we propose a novel structural compact core tensor dictionary learning (SCCTDL) model for MS-RSI deblurring. First, the multispectral patch is modeled by three-order tensor and high-order singular value decomposition is applied to the tensor. Then the task of MS-RSI deblurring is formulated as a minimum sparse core tensor estimation problem. To improve the accuracy of core tensor coding, the core tensor estimation based on the structural compact principle is introduced into the SCCTDL model to exploit abundant structural similarity in image. Experimental results suggest that our method outperforms several existing MS-RSI deblurring methods in both subjective image quality and visual perception. Leilei Geng, Xiushan Nie, Sijie Niu, Yilong Yin |
ICIP | 4 |
| 2018 | Fully convolutional network and graph-based method for co-segmentation of retinal layer on macular OCT imagesabstractRetinal layer segmentation in optical coherence tomography (OCT) images is crucial for the diagnosis and study of retinal diseases. Graph-based methods are commonly used in layer segmentation. However, most of these methods require a lot of human efforts for determining an appropriate model to compute good edge weights. In this paper, we propose a novel automatic method for segmenting retinal layers in macular OCT images. Specially, we propose a new fully convolutional deep learning architecture with a side output layer to directly learn optimal graph-edge weights from raw pixels. The architecture can automatically learn multi-scale and multi-level features to generate accurate boundary probabilities as good edge weights without hand-crafted appropriate models. The boundaries are finalized by using graph segmentation method. The proposed method is evaluated on a dataset with 130 OCT B-scans. The experimental results show the mean absolute boundary positioning differences are 1.48±0.34 pixel. Yun Liu 0039, Gongping Yang 0001, Xiaoming Xi, Xinjian Chen 0001, Yilong Yin |
ICPR | 6 |
| 2018 | Robust ECG Biometrics Using Two-Stage ModelabstractECG biometrics has achieved great success on high quality ECG signals. However, it is still a challenging problem to apply ECG biometrics on mobile devices due to the low quality signals. In this paper, we propose a robust two-stage model. In first stage, we utilize 1D CNN model to remove the invalid heartbeats from ECG recording. And then, we combine the raw signal with the hidden feature of 1D CNN as the feature representation of heartbeat. In second stage, we group a certain number of heartbeat representations as input sequence. Attention-based bidirectional LSTM is used to aggregate input sequence and generate discriminative identity features for recognition. We evaluate our method on two public datasets, and the results show that our two-stage model can achieve the state-of-the-art performance compared with other existing methods. Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
ICPR | 4 |
| 2018 | Deep Cross-View Label Embedding with Correlation and Structure Preserved for Multi-Label ClassificationabstractLabel embedding is an important family of multi-label classification algorithms which can jointly extract the information of all labels for better performance. However, few works have been done on label embedding methods which consider the structure information of original feature and label space simultaneously. We propose a novel deep neural network (DNN) based model for learning an effective deep latent space, namely Deep Cross-view label space Embedding with Correlation and Structure preserved (DCECS). In DCECS, the latent space correlates with feature and label spaces closely by virtue of the deep cross-view embedding. Meanwhile, the latent space is also learned under the guidance of label correlation and local structure of feature space which are exploited by hypergraph and graph regularizations. The overall framework achieves the complementarity and correspondence between information of feature and label space, therefore the feature-aware deep latent space we learned has strong predictability and discriminant ability. Extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms. Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin |
ICTAI | 4 |
| 2018 | Deep Random Walk for Drusen Segmentation from Fundus Images
Fang Yan 0003, Jia Cui, Yu Wang 0228, Hong Liu 0013, Hui Liu 0007, Benzheng Wei, Yilong Yin, Yuanjie Zheng |
MICCAI (2) | 7 |
| 2018 | Fast Discrete Cross-modal Hashing With Regressing From Semantic LabelsabstractHashing has recently received great attention in cross-modal retrieval. Cross-modal retrieval aims at retrieving information across heterogeneous modalities (e.g., texts vs. images). Cross-modal hashing compresses heterogeneous high-dimensional data into compact binary codes with similarity preserving, which provides efficiency and facility in both retrieval and storage. In this study, we propose a novel fast discrete cross-modal hashing (FDCH) method with regressing from semantic labels to take advantage of supervised labels to improve retrieval performance. In contrast to existing methods that learn the projection from hash codes to semantic labels, the proposed FDCH regresses the semantic labels of training examples to the corresponding hash codes with a drift. It not only accelerates the hash learning process, but also helps generate stable hash codes. Furthermore, the drift can adjust the regression and enhance the discriminative capability of hash codes. Especially in the case of training efficiency, FDCH is much faster than existing methods. Comparisons with several state-of-the-art techniques on three benchmark datasets have demonstrated the superiority of FDCH under various cross-modal retrieval scenarios. Xingbo Liu, Xiushan Nie, Wenjun Zeng 0001, Chaoran Cui, Lei Zhu 0002, Yilong Yin |
ACM Multimedia | 6 |
| 2018 | Image Aesthetic Distribution Prediction with Fully Convolutional Network
Huidi Fang, Chaoran Cui, Xiang Deng 0002, Xiushan Nie, Muwei Jian, Yilong Yin |
MMM (1) | 6 |
| 2018 | Finger vein recognition based on deformation information
Xianjing Meng, Xiaoming Xi, Gongping Yang 0001, Yilong Yin |
Sci. China Inf. Sci. | 4 |
| 2018 | Learned local similarity prior embedding active contour model for choroidal neovascularization segmentation in optical coherence tomography images
Xiaoming Xi, Xianjing Meng, Lu Yang 0005, Xiushan Nie, Zhilou Yu, Chunyun Zhang, Haoyu Chen 0002, Yilong Yin, Xinjian Chen 0001 |
Sci. China Inf. Sci. | 8 |
| 2018 | Polygene-based evolutionary algorithms with frequent pattern mining
Shuaiqiang Wang, Yilong Yin |
Frontiers Comput. Sci. | 2 |
| 2018 | Sequential quadratic programming enhanced backtracking search algorithm
Wenting Zhao 0004, Lijin Wang, Yilong Yin, Yuchun Tang |
Frontiers Comput. Sci. | 3 |
| 2018 | Geometric shape analysis based finger vein deformation detection and correction
Lu Yang 0005, Gongping Yang 0001, Yilong Yin |
Neurocomputing | 4 |
| 2018 | Graph cut based automatic aorta segmentation with an adaptive smoothness constraint in 3D abdominal CT images
Xiang Deng 0002, Yuanjie Zheng, Xiaoming Xi, Yilong Yin |
Neurocomputing | 6 |
| 2018 | Fast and effective optic disk localization based on convolutional neural network
Xianjing Meng, Xiaoming Xi, Lu Yang 0005, Yilong Yin, Xinjian Chen 0001 |
Neurocomputing | 5 |
| 2018 | Identifying advisor-advisee relationships from co-author networks via a novel deep model
Zhongying Zhao 0001, Liqiang Nie, Yilong Yin, Yong Zhang 0001 |
Inf. Sci. | 5 |
| 2018 | Integrating QDWD with pattern distinctness and local contrast for underwater saliency detection
Muwei Jian, Qiang Qi, Junyu Dong, Yilong Yin, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Saliency detection based on background seeds by object proposals and extended random walk
Muwei Jian, Runxia Zhao, Xin Sun 0003, Hanjiang Luo, Wenyin Zhang, Huaxiang Zhang 0001, Junyu Dong, Yilong Yin, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 8 |
| 2018 | Saliency detection based on directional patches extraction and principal local color contrast
Muwei Jian, Wenyin Zhang, Hui Yu 0001, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Yilong Yin |
J. Vis. Commun. Image Represent. | 7 |
| 2018 | Content-based image retrieval via a hierarchical-local-feature extraction scheme
Muwei Jian, Yilong Yin, Junyu Dong, Kin-Man Lam 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Finger Vein Recognition With Anatomy Structure AnalysisabstractFinger vein recognition has received a lot of attention recently and is viewed as a promising biometric trait. In related methods, vein pattern-based methods explore intrinsic finger vein recognition, but their performance remains unsatisfactory owing to defective vein networks and weak matching. One important reason may be the neglect of deep analysis of the vein anatomy structure. By comprehensively exploring the anatomy structure and imaging characteristic of vein patterns, this paper proposes a novel finger vein recognition framework, including an anatomy structure analysis-based vein extraction algorithm and an integration matching strategy. Specifically, the vein pattern is extracted from the orientation map-guided curvature based on the valley- or half valley-shaped cross-sectional profile. In addition, the extracted vein pattern is further thinned and refined to obtain a reliable vein network. In addition to the vein network, the relatively clear vein branches in the image are mined from the vein pattern, referred to as the vein backbone. In matching, the vein backbone is used in vein network calibration to overcome finger displacements. The similarity of two calibrated vein networks is measured by the proposed elastic matching and further recomputed by integrating the overlap degree of corresponding vein backbones. Extensive experiments on two public finger vein databases verify the effectiveness of the proposed framework. Lu Yang 0005, Gongping Yang 0001, Yilong Yin, Xiaoming Xi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Robust Image Fingerprinting Based on Feature Point Relationship MiningabstractLocal feature points have been widely employed in robust image fingerprinting. One of their intrinsic advantages is their invariance under geometric transforms. However, their robustness against certain attacks that modify the positions of points, such as additive noising and blurring, is limited. In addition, local-feature-point-based approaches ignore the distribution of the feature points. In this paper, we harness feature point relationships, including local structures and global relevance, to overcome these limitations. In the relationship mining strategy, Delaunay triangulation is first applied to the feature points to capture their geometric structures. Subsequently, local structures are represented by searching for an independent set in the mapping graph constructed via Delaunay triangulation, whereas the global relevance is represented by the Laplacian of the graph. Finally, the local structures and global relevance are used as input to the quantization process of the image fingerprinting system. In the process of quantization, we propose an unsupervised quantization strategy called between-cluster distance-based quantization to preserve the neighborhood structure between the binary fingerprint space and the original feature space. Experimental results show that the proposed method achieves effective performance under common modifications. Xiushan Nie, Yane Chai, Chaoran Cui, Xiaoming Xi, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2018 | Multiscale Rotation-Invariant Convolutional Neural Networks for Lung Texture ClassificationabstractWe propose a new multiscale rotation-invariant convolutional neural network (MRCNN) model for classifying various lung tissue types on high-resolution computed tomography. MRCNN employs Gabor-local binary pattern that introduces a good property in image analysis-invariance to image scales and rotations. In addition, we offer an approach to deal with the problems caused by imbalanced number of samples between different classes in most of the existing works, accomplished by changing the overlapping size between the adjacent patches. Experimental results on a public interstitial lung disease database show a superior performance of the proposed method to state of the art. Qiangchang Wang, Yuanjie Zheng, Gongping Yang 0001, Weidong Jin, Xinjian Chen 0001, Yilong Yin |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | Patch-Based Image Inpainting via Two-Stage Low Rank ApproximationabstractTo recover the corrupted pixels, traditional inpainting methods based on low-rank priors generally need to solve a convex optimization problem by an iterative singular value shrinkage algorithm. In this paper, we propose a simple method for image inpainting using low rank approximation, which avoids the time-consuming iterative shrinkage. Specifically, if similar patches of a corrupted image are identified and reshaped as vectors, then a patch matrix can be constructed by collecting these similar patch-vectors. Due to its columns being highly linearly correlated, this patch matrix is low-rank. Instead of using an iterative singular value shrinkage scheme, the proposed method utilizes low rank approximation with truncated singular values to derive a closed-form estimate for each patch matrix. Depending upon an observation that there exists a distinct gap in the singular spectrum of patch matrix, the rank of each patch matrix is empirically determined by a heuristic procedure. Inspired by the inpainting algorithms with component decomposition, a two-stage low rank approximation (TSLRA) scheme is designed to recover image structures and refine texture details of corrupted images. Experimental results on various inpainting tasks demonstrate that the proposed method is comparable and even superior to some state-of-the-art inpainting algorithms. Qiang Guo 0003, Shanshan Gao 0003, Xiaofeng Zhang 0003, Yilong Yin, Caiming Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Personalized Image Aesthetics AssessmentabstractAutomatically assessing image quality from an aesthetic perspective is of great interest to the high-level vision research community. Existing methods are typically non-personalized and quantify image aesthetics with a universal label. However, given the fact that aesthetics is a subjective perception, how to understand user aesthetic perceptions poses a formidable challenge to image aesthetics assessment. In this paper, we propose to model user aesthetic perceptions using a set of exemplar images from social media platforms, and realize personalized aesthetics assessment by transferring this knowledge to adapt the results of the trained generic model. In this way, image aesthetics is measured from both aspects of visual quality and user tastes. Extensive experiments on two benchmark datasets well verified the potential of our approach for personalized image aesthetics assessment. Xiang Deng 0002, Chaoran Cui, Huidi Fang, Xiushan Nie, Yilong Yin |
CIKM | 5 |
| 2017 | Learning Deep Match Kernels for Image-Set ClassificationabstractImage-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this paper, we propose deep match kernels (DMK) to directly measure the similarity between image sets in the match kernel framework. Specifically, we build deep local match kernels between images upon arc-cosine kernels, which can faithfully characterize the similarity between images by mimicking deep neural networks, we introduce anchors to aggregate those deep local match kernels into a global match kernel between image sets, which is learned in a supervised way by kernel alignment and therefore more discriminative. The DMK provides the first match kernel framework for image-set classification, which removes specific assumptions usually required in previous approaches and is computationally more efficient. We conduct extensive experiments on four datasets for three diverse image-set classification tasks. The DMK achieves high performance and consistently surpasses state-of-the-art methods, showing its great effectiveness for image-set classification. Haoliang Sun, Xiantong Zhen, Yuanjie Zheng, Gongping Yang 0001, Yilong Yin, Shuo Li 0001 |
CVPR | 5 |
| 2017 | DFVR: Deformable finger vein recognitionabstractAlthough some developments have been achieved in finger vein recognition recently, the image deformation problem has received relatively less attention and still intractable. In this paper, the reason and the harmfulness of this problem are analyzed firstly. And then, a deformable finger vein recognition framework is proposed to deal with this problem, consisting of the improved vein PCA-SIFT feature and bidirectional deformable spatial pyramid matching (BDSPM). Furthermore, we build a finger vein deformation database to imitate image deformation in real application. The experimental results, on the self-built deformation database and one public database, prove the effectiveness of the proposed framework for dealing with the image deformation problem. Lu Yang 0005, Gongping Yang 0001, Yilong Yin, Xianjing Meng |
ICASSP | 4 |
| 2017 | Finger vein image retrieval via affinity-preserving K-means hashingabstractEfficient identification of finger veins is still a challenging problem due to the increasing size of the finger vein database. Most leading finger vein image identification methods have high-dimensional real-valued features, which result in extremely high computation complexity. Hashing algorithms are extraordinary effective ways to facilitate finger vein image retrieval. Therefore, in this paper, we proposed a finger vein image retrieval scheme based on Affinity-Preserving K-means Hashing (APKMH) algorithm and bag of subspaces based image feature. At first, we represent finger vein image by Nonlinearly Sub-space Coding (NSC) method which can obtain the discriminative finger vein image features. Then the features space is partitioned into multiple subsegments. In each subsegment, we employ the APKMH algorithm, which can simultaneously construct the visual codebook by directly k-means clustering and encode the feature vector as the binary index of the codeword. Experimental results on a large fused finger vein dataset demonstrate that our hashing method outperforms the state-of-the-art finger vein retrieval methods. Gongping Yang 0001, Lu Yang 0005, Yilong Yin |
IJCB | 4 |
| 2017 | Integration of discriminative features and similarity-preserving encoding for finger vein image retrievalabstractAlthough some image retrieval methods were proposed to accelerate finger vein recognition, the insufficient feature (e.g., the number of vein point) and unfavorable encoding (e.g., predefined threshold based binarization) limited retrieval performance largely. In view of this problem, we develop a new retrieval framework, based on the integration of discriminative texture features and similarity-preserving binary codes. In detail, the vector and scalar features, measuring the gray level, gray difference, and gray gathering of image patch, are both used to represent finger vein image. And to improve the retrieval efficiency, the high-dimensional decimal features are further encoded into the compact binary patterns by principal component analysis (PCA) and similarity-preserving iterative quantization (ITQ). Experimental results on one large finger vein database prove that the proposed method can powerfully improve the retrieval accuracy and efficiency. Kuikui Wang, Lu Yang 0005, Gongping Yang 0001, Yilong Yin |
ICIP | 4 |
| 2017 | An Adaptive Sentence Representation Learning Model Based on Multi-gram CNNabstractNature Language Processing has been paid more attention recently. Traditional approaches for language model primarily rely on elaborately designed features and complicated natural language processing tools, which take a large amount of human effort and are prone to error propagation and data sparse problem. Deep neural network method has been shown to be able to learn implicit semantics of text without extra knowledge. To better learn deep underlying semantics of sentences, most deepneuralnetworklanguagemodelsutilizemulti-gramstrategy. However, the current multi-gram strategies in CNN framework are mostly realized by concatenating trained multi-gram vectors to form the sentence vector, which can increase the number of parameters to be learned and is prone to over fitting. To alleviate the problem mentioned above, we propose a novel adaptive sentence representation learning model based on multigram CNN framework. It learns adaptive importance weights of different n-gram features and forms sentence representation by using weighted sum operation on extracted n-gram features, which can largely reduce parameters to be learned and alleviate the threat of over fitting. Experimental results show that the proposed method can improve performances when be used in sentiment and relation classification tasks. Chunyun Zhang, Baolin Zhao, Lu Yang 0005, Xiaoming Xi, Chaoran Cui, Yilong Yin |
Intelligent Environments | 7 |
| 2017 | Finger Vein Image Retrieval via Coding Scale-varied Superpixel FeatureabstractFinger vein image retrieval is one significant technique for performing fast identification especially in large-scale applications. However, most existing retrieval methods were based on fixed-scale feature of non-overlapped rectangular image block, in which the representation ability of feature and the local consistency of vein pattern were both overlooked. And the weak encoding (e.g., predefined threshold based binarization) was also limited the retrieval performance. Focusing on these problems, this paper proposes a novel finger vein image retrieval framework based on similarity-preserving encoding of scale-varied superpixel feature. In the framework, locally consistent pixels in one superpixel are used as a unit of feature representation, and the feature length is varied with the category of the superpixel classified by the variance of lowest dimensional feature. Additionally, the feature compaction and feature rotation based encoding can minimize the quantization loss and preserve the similarity between the scale-varied feature and the encoded binary codes. Experimental results on six public finger vein databases demonstrate that the superiority of the proposed coding scale-varied superpixel feature based retrieval approach over the state-of-the-arts. Kuikui Wang, Lu Yang 0005, Gongping Yang 0001, Xin Luo 0006, Yilong Yin |
ICMR | 6 |
| 2017 | Distribution-oriented Aesthetics Assessment for Image SearchabstractAesthetics has become increasingly prominent for image search to enhance user satisfaction. Therefore, image aesthetics assessment is emerging as a promising research topic in recent years. In this paper, distinguished from existing studies relying on a single label, we propose to quantify the image aesthetics by a distribution over quality levels. The distribution representation can effectively characterize the disagreement among the aesthetic perceptions of users regarding the same image. Our framework is developed on the foundation of label distribution learning, in which the reliability of training examples and the correlations between quality levels are fully taken into account. Extensive experiments on two benchmark datasets well verified the potential of our approach for aesthetics assessment. The role of aesthetics in image search was also rigorously investigated. Chaoran Cui, Huidi Fang, Xiang Deng 0002, Xiushan Nie, Hongshuai Dai, Yilong Yin |
SIGIR | 6 |
| 2017 | Choroid segmentation from Optical Coherence Tomography with graph-edge weights learned from deep convolutional neural networks
Xiaodan Sui, Yuanjie Zheng, Benzheng Wei, Hongsheng Bi, Xuemei Pan, Yilong Yin, Shaoting Zhang 0001 |
Neurocomputing | 7 |
| 2017 | Corrigendum to "Hierarchical retinal blood vessel segmentation based on feature and ensemble learning" [Neurocomputing 149 (2015) 708-717]
Shuangling Wang, Yilong Yin, Guibao Cao, Benzheng Wei, Yuanjie Zheng, Gongping Yang 0001 |
Neurocomputing | 2 |
| 2017 | Breast tumor segmentation with prior knowledge learning
Xiaoming Xi, Lingyan Han, Tingwen Wang, Hong Yu Ding, Yuchun Tang, Yilong Yin |
Neurocomputing | 8 |
| 2017 | Robust texture analysis of multi-modal images using Local Structure Preserving Ranklet and multi-task learning for breast tumor diagnosis
Xiaoming Xi, Chunyun Zhang, Hong Yu Ding, Yuchun Tang, Yilong Yin |
Neurocomputing | 8 |
| 2017 | Coronal Mass Ejections detection using multiple features based ensemble learning
Jianqin Yin, Hai Yao, Jiaben Lin, Yilong Yin, Zhiquan Feng |
Neurocomputing | 4 |
| 2017 | Hybrid textual-visual relevance learning for content-based image retrieval
Chaoran Cui, Peiguang Lin, Xiushan Nie, Yilong Yin, Qingfeng Zhu |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Learning discriminative binary codes for finger vein recognition
Xiaoming Xi, Lu Yang 0005, Yilong Yin |
Pattern Recognit. | 3 |
| 2017 | Superpixels by Bilateral Geodesic DistanceabstractWe present a novel superpixel generation algorithm based on a new definition of geodesic distance, called bilateral geodesic distance. In contrast to the traditional geodesic distance, the new bilateral geodesic distance of two pixels considers the distance between their positions as well as their color difference. Superpixel generation is essentially a problem of clustering image pixels with respect to a set of properly selected seeds. We first use an adaptive hexagonal subdivision method to determine the initial seed-based image gradient. Then, we use the bilateral geodesic distance to measure the similarity between the pixels and the seeds. We apply an improved fast marching method to generate superpixels’ contour regions with the expansion velocities dependent on a new gradient formulation that depends on the seeds’ properties. The experimental results indicate that our algorithm is not only much faster than the structure-based method, which uses conventional geodesic distance, but also outperforms the existing methods in terms of region compactness and region boundary regularity. Yuanfeng Zhou, Wenping Wang 0001, Yilong Yin, Caiming Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | A Maximal Clique Based Multiobjective Evolutionary Algorithm for Overlapping Community DetectionabstractDetecting community structure has become one important technique for studying complex networks. Although many community detection algorithms have been proposed, most of them focus on separated communities, where each node can belong to only one community. However, in many real-world networks, communities are often overlapped with each other. Developing overlapping community detection algorithms thus becomes necessary. Along this avenue, this paper proposes a maximal clique based multiobjective evolutionary algorithm (MOEA) for overlapping community detection. In this algorithm, a new representation scheme based on the introduced maximal-clique graph is presented. Since the maximal-clique graph is defined by using a set of maximal cliques of original graph as nodes and two maximal cliques are allowed to share the same nodes of the original graph, overlap is an intrinsic property of the maximal-clique graph. Attributing to this property, the new representation scheme allows MOEAs to handle the overlapping community detection problem in a way similar to that of the separated community detection, such that the optimization problems are simplified. As a result, the proposed algorithm could detect overlapping community structure with higher partition accuracy and lower computational cost when compared with the existing ones. The experiments on both synthetic and real-world networks validate the effectiveness and efficiency of the proposed algorithm. Xuyun Wen, Weineng Chen, Ying Lin 0001, Tianlong Gu, Huaxiang Zhang 0001, Yun Li 0002, Yilong Yin, Jun Zhang 0003 |
IEEE Trans. Evol. Comput. | 7 |
| 2017 | Comprehensive Feature-Based Robust Video Fingerprinting Using Tensor ModelabstractContent-based near-duplicate video detection (NDVD) is essential for effective search and retrieval, and robust video fingerprinting is a good solution for NDVD. Most existing video fingerprinting methods use a single feature or concatenate different features to generate video fingerprints, and show good performance under single-mode modifications such as noise addition and blurring. However, when they suffer combined modifications, the performance is degraded to a certain extent because such features cannot characterize the video content completely. By contrast, the assistance and consensus among different features can improve the performance of video fingerprinting. Therefore, in the present study, we mine the assistance and consensus among different features based on a tensor model, and we present a new comprehensive feature to fully use them in the proposed video fingerprinting framework. We also analyze what the comprehensive feature really is for representing the original video. In this framework, the video is initially set as a high-order tensor that consists of different features, and the video tensor is decomposed via the Tucker model with a solution that determines the number of components. Subsequently, the comprehensive feature is generated by the low-order tensor obtained from tensor decomposition. Finally, the video fingerprint is computed using this feature. A matching strategy used for narrowing the search is also proposed based on the core tensor. The robust video fingerprinting framework is resistant not only to single-mode modifications but also to their combination. Xiushan Nie, Yilong Yin, Jiande Sun 0001, Chaoran Cui |
IEEE Trans. Multim. | 2 |
| 2016 | Best Guided Backtracking Search Algorithm for Numerical Optimization Problems
Wenting Zhao 0004, Lijin Wang, Yilong Yin |
KSEM | 4 |
| 2016 | Automated Diagnosis of Neural Foraminal Stenosis Using Synchronized Superpixels Representation
Xiaoxu He, Yilong Yin, Manas Sharma, Gary Brahm, Ashley Mercado, Shuo Li 0001 |
MICCAI (2) | 2 |
| 2016 | Multi-task Shape Regression for Medical Image SegmentationabstractIn this paper, we propose a general segmentation framework of Multi-Task Shape Regression (MTSR) which formulates segmentation as multi-task learning to leverage its strength of jointly solving multiple tasks enhanced by capturing task correlations. The MTSR entirely estimates coordinates of all points on shape contours by multi-task regression, where estimation of each coordinate corresponds to a regression task; the MTSR can jointly handle nonlinear relationships between image appearance and shapes while capturing holistic shape information by encoding coordinate correlations, which enables estimation of highly variable shapes, even with vague edge or region inhomogeneity. The MTSR achieves a long-desired general framework without relying on any specific assumptions or initialization, which enables flexible and fully automatic segmentation of multiple objects simultaneously, for different applications irrespective of modalities. The MTSR is validated on six representative applications of diverse images, achieves consistently high performance with dice similarity coefficient (DSC) up to 0.93 and largely outperforms state of the arts in each application, which demonstrates its effectiveness and generality for medical image segmentation. Xiantong Zhen, Yilong Yin, Mousumi Bhaduri, Ilanit Ben Nachum, David T. Laidley, Shuo Li 0001 |
MICCAI (3) | 2 |
| 2016 | Spherical torus-based video hashing for near-duplicate video detection
Xiushan Nie, Yane Chai, Jiande Sun 0001, Yilong Yin |
Sci. China Inf. Sci. | 5 |
| 2016 | Unmatched minutiae: Useful information to boost fingerprint recognition
Yilong Yin, Gongping Yang 0001 |
Neurocomputing | 2 |
| 2016 | Finger Vein Recognition Based on Stable and Discriminative SuperpixelsabstractFinger vein pattern, as a promising hand-based biometric technology, has been well studied in recent years. In this paper, a new superpixel-based finger vein recognition method is presented. In the proposed method, we develop two types of effective superpixels, i.e. stable superpixel and discriminative superpixel to represent finger vein image and these superpixels are expected to play different roles in matching stage. In detail, the stable and discriminative superpixels are firstly learned from the training images for each enrolled class. When verifying a testing image, we just compare the superpixels at the same location as the two types of superpixels in template. Then, the two types of superpixels are combined utilizing a reversible weight-based fusion method in score level. Additionally, to further improve the recognition performance, we explore the superpixel context feature (SPCF). For each superpixel the SPCF is obtained by comparing the current superpixel with its surrounding neighbors. In the final matching stage, we integrate the matching score of two types of superpixels and it of the SPCF using the weighted SUM fusion method. The experimental results on two open finger vein databases, i.e. PolyU and SDUMLA-FV, show that our method not only performs better than the existing superpixel-based method, but also has advantages in comparison with some traditional ones. Lizhen Zhou, Gongping Yang 0001, Yilong Yin, Lu Yang 0005, Kuikui Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2016 | An Improved Neural Network with Random Weights Using Backtracking Search Algorithm
Lijin Wang, Yilong Yin, Wenting Zhao 0004, Yuchun Tang |
Neural Process. Lett. | 3 |
| 2016 | k-Multi-preference query over road networks
Peiguang Lin, Yilong Yin, Peiyao Nie |
Pers. Ubiquitous Comput. | 2 |
| 2016 | A Hybrid Evolutionary Immune Algorithm for Multiobjective Optimization ProblemsabstractIn recent years, multiobjective immune algorithms (MOIAs) have shown promising performance in solving multiobjective optimization problems (MOPs). However, basic MOIAs only use a single hypermutation operation to evolve individuals, which may induce some difficulties in tackling complicated MOPs. In this paper, we propose a novel hybrid evolutionary framework for MOIAs, in which the cloned individuals are divided into several subpopulations and then evolved using different evolutionary strategies. An example of this hybrid framework is implemented, in which simulated binary crossover and differential evolution with polynomial mutation are adopted. A fine-grained selection mechanism and a novel elitism sharing strategy are also adopted for performance enhancement. Various comparative experiments are conducted on 28 test MOPs and our empirical results validate the effectiveness and competitiveness of our proposed algorithm in solving MOPs of different types. Qiuzhen Lin, Jianyong Chen, Zhi-hui Zhan, Weineng Chen, Carlos A. Coello Coello, Yilong Yin, Chih-Min Lin, Jun Zhang 0003 |
IEEE Trans. Evol. Comput. | 6 |
| 2015 | A hybrid biometric identification framework for high security applications
Xuzhou Li, Yilong Yin, Yanbin Ning, Gongping Yang 0001 |
Frontiers Comput. Sci. | 2 |
| 2015 | Cuckoo search with varied scaling factor
Lijin Wang, Yilong Yin, Yiwen Zhong |
Frontiers Comput. Sci. | 2 |
| 2015 | Hierarchical retinal blood vessel segmentation based on feature and ensemble learning
Shuangling Wang, Yilong Yin, Guibao Cao, Benzheng Wei, Yuanjie Zheng, Gongping Yang 0001 |
Neurocomputing | 2 |
| 2015 | Finger Vein Verification with Vein TextonsabstractFinger vein pattern has become one of the most promising biometric identifiers. In this paper, a robust method based on Bag-of-Words (BoW) is developed for finger vein verification. Firstly, some robust and discriminative visual words are learned from local base features such as Local Binary Pattern (LBP), Mean Curvature and Webber Local Descriptor (WLD). We name these visual words as Finger Vein Textons (FVTs). Secondly, each image is mapped into a FVTs matrix. Finally, spatial pyramid matching (SPM) method is applied to maintain spatial layout information by representing each image as pyramid histogram which is performed for matching by histogram intersection function. Experimental results show that the proposed method achieves satisfactory performance both on our database and the open PolyU database. In addition, our method also has strong robustness and high accuracy on the self-built rotation and illumination databases. Lumei Dong, Gongping Yang 0001, Yilong Yin, Xiaoming Xi, Lu Yang 0005, Fei Liu 0010 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | Finger vein verification based on a personalized best patches mapabstractFinger vein pattern has become one of the most promising biometric identifiers. In this paper, we propose a robust finger vein verification method based on a personalized best patches map (PBPM). Firstly, some robust and discriminative visual words of finger vein are learned from traditional base feature such as local binary pattern (LBP). These visual words are named as finger vein textons (FVTs), which can well represent the visual primitives of finger vein. Secondly, we represent the finger vein image as a finger vein textons map (FVTM) by mapping each patch of the image into the closest FVT. Thirdly, by rejecting inconsistent patches, the PBPM of a certain individual is learned from these FVTMs which are extracted from the training samples of the same finger. Finally, the matched best patch ratio is used to measure similarity between the extracted FVTM of the input finger and the PBPM of a certain individual. Experimental results show that our method achieves satisfactory performance on the open PolyU database. In addition, it also has strong robustness and high accuracy on the self-built rotation and translation databases. Lumei Dong, Gongping Yang 0001, Yilong Yin, Fei Liu 0010, Xiaoming Xi |
IJCB | 3 |
| 2014 | Finger vein recognition with superpixel-based featuresabstractFinger veins based biometrics, as a new approach to personal identification, has received much attention in recent years. The methods based on low level feature, for instance the gray, texture of finger vein, are the mainstream, but they are usually faced with many challenges, such as sensitivity to noise and low local consistency. In fact, finger vein recognition based on high level feature representation has been proved to be a promising way to effectively overcome the above limitations and improve the system performance. Thus, in this paper, we present a novel identification framework, which utilizes superpixel-based features (SPFs) of finger vein for high level feature representation. When comparing two finger veins, the features of each pixel are firstly extracted as base attributes by traditional way. Then, after superpixel over-segmentation, the SPF of each finger vein can be obtained based on its base attributes by some statistical techniques. Lastly, a weighted spatial pyramid matching (WSPM) scheme is utilized to implement matching. Our experiments have yielded some very good results evidenced by an EER of 0.0147 on the benchmark database PolyU. Fei Liu 0010, Yilong Yin, Gongping Yang 0001, Lumei Dong, Xiaoming Xi |
IJCB | 2 |
| 2014 | An Improved Backtracking Search Algorithm for Constrained Optimization Problems
Wenting Zhao 0004, Lijin Wang, Yilong Yin, Yushan Yin |
KSEM | 3 |
| 2014 | Singular value decomposition based minutiae matching method for finger vein recognition
Fei Liu 0010, Gongping Yang 0001, Yilong Yin, Shuaiqiang Wang |
Neurocomputing | 3 |
| 2014 | Exploring soft biometric trait with finger vein recognition
Lu Yang 0005, Gongping Yang 0001, Yilong Yin, Xiaoming Xi |
Neurocomputing | 3 |
| 2014 | Bilinear discriminative dictionary learning for face recognition
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Yilong Yin |
Pattern Recognit. | 4 |
| 2014 | A Novel Serial Multimodal Biometrics Framework Based on Semisupervised Learning TechniquesabstractWe propose in this paper a novel framework for serial multimodal biometric systems based on semisupervised learning techniques. The proposed framework addresses the inherent issues of user inconvenience and system inefficiency in parallel multimodal biometric systems. Further, it advances the serial multimodal biometric systems by promoting the discriminating power of the weaker but more user convenient trait(s) and saving the use of the stronger but less user convenient trait(s) whenever possible. This is in contrast to other existing serial multimodal biometric systems that suggest optimized orderings of the traits deployed and parameterizations of the corresponding matchers but ignore the most important requirements of common applications. In terms of methodology, we propose to use semisupervised learning techniques to strengthen the matcher(s) on the weaker trait(s), utilizing the coupling relationship between the weaker and the stronger traits. A dimensionality reduction method for the weaker trait(s) based on dependence maximization is proposed to achieve this purpose. Experiments on two prototype systems clearly demonstrate the advantages of the proposed framework and methodology. Yilong Yin, De-Chuan Zhan, Jingliang Peng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Automatic measurement on CT images for patella dislocation diagnosisabstractTo diagnose the patella dislocation, various angles and distances need to be measured on knee CT images, which was traditionally done by doctors manually. In this work, we propose a novel scheme for automatic measurement on knee CT images to assist doctors diagnosis of patella dislocation. Specifically, we first segment the femur and the patella regions on the CT images, then adopt optimal fitting to obtain the central planes of the femur and the patella bones and, based on which, make the measurement. As experimentally demonstrated, the measured results obtained with our system are highly consistent with those manually made by experienced doctors. Qi Kong, Shaoshan Wang, Jiushan Yang, Ruiqi Zou, Yan Huang 0003, Yilong Yin, Jingliang Peng |
ICIP | 6 |
| 2012 | Semi-supervised Gait Recognition Based on Self-TrainingabstractTraditional gait recognition researches focus on supervised learning methods that use only a limited number of labeled sequences to train, which will definitely restrict the recognition ability of the gait recognition system. Meanwhile, training with more typical gait sequences can improve the generalization ability of gait recognition system and eventually achieve better recognition accuracy. However, it is difficult, expensive, time consuming and boring to capture enough gait sequences comparing with capturing other biometric traits such as fingerprint, face and iris during the enrolment stage. To address the problem, a semi-supervised gait recognition algorithm based on self-training is proposed to optimize the performance of gait recognition system with both a few labeled sequences and a large amount of unlabeled sequences. Nearest Neighbor (NN) classifier and K-Nearest Neighbor (KNN) classifier are carried out to recognize the different subjects. Experimental results show that the proposed algorithm has an encouraging recognition performance even with only one labeled sequence each class. Yilong Yin, Shaohua Pang, Qiuhong Yu |
AVSS | 2 |
| 2012 | Polygene-based evolution: a novel framework for evolutionary algorithmsabstractIn this paper, we introduce polygene-based evolution, a novel framework for evolutionary algorithms (EAs) that features distinctive operations in the evolution process. In traditional EAs, the primitive evolution unit is gene, where genes are independent components during evolution. In polygene-based evolutionary algorithms (PGEAs), the evolution unit is polygene, i.e., a set of co-regulated genes. Discovering and maintaining quality polygenes can play an effective role in evolving quality individuals. Polygenes generalize genes, and PGEAs generalize EAs. Implementing the PGEA framework involves three phases: polygene discovery, polygene planting, and polygene-compatible evolution. Extensive experiments on function optimization benchmarks in comparison with the conventional and state-of-the-art EAs demonstrate the potential of the approach in accuracy and efficiency improvement. Shuaiqiang Wang, Byron J. Gao, Shuangling Wang, Guibao Cao, Yilong Yin |
CIKM | 5 |
| 2012 | Importance weighted passive learningabstractImportance weighted active learning (IWAL) introduces a weighting scheme to measure the importance of each instance for correcting the sampling bias of the probability distributions between training and test datasets. However, the weighting scheme of IWAL involves the distribution of the test data, which can be straightforwardly estimated in active learning by interactively querying users for labels of selected test instances, but difficult for conventional learning where there are no interactions with users, referred as passive learning. In this paper, we investigate the insufficient sampling bias problem, i.e., bias occurs only because of insufficient samples, but the sampling process is unbiased. In doing this, we present two assumptions on the sampling bias, based on which we propose a practical weighting scheme for the empirical loss function in conventional passive learning, and present IWPL, an importance weighted passive learning framework. Furthermore, we provide IWSVM, an importance weighted SVM for validation. Extensive experiments demonstrate significant advantages of IWSVM on benchmarks and synthetic datasets. Shuaiqiang Wang, Xiaoming Xi, Yilong Yin |
CIKM | 3 |
| 2012 | A Novel Method Using Videos for Fingerprint VerificationabstractTraditional fingerprint verifications use single image for matching. However, the verification accuracy cannot meet the need of some application domains. In this paper, we propose to use videos for fingerprint verification. To take full use of the information contained in fingerprint videos, we present a novel method to use the dynamic as well as the static information in fingerprint videos. After preprocessing and aligning processes, the Inclusion Ratio of two matching fingerprint videos is calculated and used to represent the similarity between these two videos. Experimental results show that video-based method can access better accuracy than the method based on single fingerprint. Yilong Yin |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Relevance feature mapping for content-based multimedia information retrieval
Guang-Tong Zhou, Kai Ming Ting, Fei Tony Liu, Yilong Yin |
Pattern Recognit. | 4 |
| 2011 | Singular points detection based on multi-resolution in fingerprint images
Dawei Weng, Yilong Yin |
Neurocomputing | 2 |
| 2010 | Video-based fingerprint verificationabstractIn this paper, fingerprint videos are used to improve the accuracy of a fingerprint verification system. We define the “inside-similarity” and “outside-similarity” to represent the similarity within a video and between two videos, respectively. A new method is proposed to define and calculate the matching score of two videos according to the similarity and the effect on the error probability of this method is analyzed theoretically. Experimental results confirm our arguments in the analysis and indicate that the proposed method can lead a much better performance than the method using a single impression. Therefore, we believe that video-based method is an effective approach to improve the accuracy of fingerprint system. Yilong Yin, Chunxiao Ren |
ICASSP | 2 |
| 2010 | A hybrid fusion method of fingerprint identification for high security applicationsabstractThough fingerprint identification is widely used now, its imperfect performance for some high security applications, such as ATM, the access control of nuclear power stations and exchequers, etc, is still a challenge. In high security applications, an extremely low false accept rate and as low as possible false reject rate are desired at the same time, which is called Double Low problem in this paper. It is to be noted that even a fingerprint system with very low equal error rate can not achieve such a Double Low goal. It is difficult to solve Double Low problem only by improving the performance of a certain individual fingerprint identification algorithm, and the fusion of various fingerprint identification algorithms becomes a promising way. In this paper, a hybrid fusion method of fingerprint identification is proposed to solve Double Low problem. Firstly, minutiae-based and ridge-based matching algorithms are used orderly, which is a kind of serial fusion strategy. Secondly, a rank-level fusion is used, which is a kind of parallel fusion strategy. Experiment results on FVC2002DB1 and FVC2002DB2 indicate that only 6.6% fingerprints are falsely rejected on the average under zero false accept rate with our method, while 14.8%, 9.4% fingerprints are falsely rejected under zero false accept rate with the serial fusion strategy and the parallel fusion strategy, respectively. Yilong Yin, Yanbin Ning |
ICIP | 1 |
| 2010 | Detecting malignant patients via modified boosted tree
Cailing Dong 0003, Yilong Yin, Xiukun Yang |
Sci. China Inf. Sci. | 2 |
| 2010 | Finger vein recognition with manifold learning
Zhi Liu 0004, Yilong Yin, Hongjun Wang 0004, Shangling Song, Qingli Li |
J. Netw. Comput. Appl. | 2 |
| 2009 | Personalized Fingerprint Segmentation
Xinjian Guo, Yilong Yin, Zhichen Shi |
ICONIP (1) | 2 |
| 2009 | A Novel Method of Score Level Fusion Using Multiple Impressions for Fingerprint VerificationabstractHow to improve the performance of an existing biometric system is always interesting and meaningful. In this paper, we present a novel method of score level fusion using multiple enrolled impressions to achieve higher verification accuracy of existing fingerprint systems. The main idea of the method is to build a representation of the biometric reference as a polyhedron by taking into account the matching results of multiple enrolled impressions. The verification step consists in measuring a distance between the centroid of the polyhedron and the acquired image. This novel method outperforms the traditional uni-matcher based scheme over a wide range of FAR and FRR values. The equal error rate of our method is observed to be 2.25%, while that of the uni-matcher is 5.75%. Chunxiao Ren, Yilong Yin, Jun Ma 0001, Gongping Yang 0001 |
SMC | 2 |
| 2009 | Feature Selection for Sensor Interoperability: A Case Study in Fingerprint SegmentationabstractThe need for sensor interoperability has increased tremendously in many fingerprint large-scale application areas such as e-commerce, welfare-disbursement and e-education. However, the problem of feature selection for sensor interoperability has received limited attention in the literature. In this paper, the relationships among person, sensor and feature are discussed. Especially, a feature selection method for sensor interoperability is proposed. Some experimental results of feature selection for sensor interoperability in fingerprint segmentation are presented as a case study. Experiments show that the various features exhibit different sensor interoperability on different sensors. Chunxiao Ren, Yilong Yin, Jun Ma 0001, Gongping Yang 0001 |
SMC | 2 |
| 2008 | Fingerprint Scaling
Chunxiao Ren, Yilong Yin, Jun Ma 0001 |
ICIC (1) | 2 |