Shuo Yang 0006

dblp:78/1102-6 · DBLP profile ↗
← Back
32ranked-venue papers
14as first author
32since 2021 · last 2026
0000-0001-6145-0150ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 14 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
abstract
Vision-Language Models (VLMs), exemplified by CLIP, have emerged as foundational for multimodal intelligence. However, their capacity for logical understanding remains significantly underexplored, resulting in critical **''logical blindspots''** that limit their reliability in practical applications. To systematically diagnose this, we introduce **LogicBench**, a comprehensive benchmark with over 50,000 vision-language pairs across 9 logical categories and 4 diverse scenarios: images, videos, anomaly detection, and medical diagnostics. Our evaluation reveals that existing VLMs, even the state-of-the-art ones, fall at over 40 accuracy points below human performance, particularly in challenging tasks like Causality and Conditionality, highlighting their reliance on surface semantics over critical logical structures. To bridge this gap, we propose **LogicCLIP**, a novel training framework designed to boost VLMs' logical sensitivity through advancements in both data generation and optimization objectives. LogicCLIP utilizes logic-aware data generation and a contrastive learning strategy that combines coarse-grained alignment, a fine-grained multiple-choice objective, and a novel logical structure-aware objective. Extensive experiments demonstrate LogicCLIP's substantial improvements in logical comprehension across all LogicBench domains, significantly outperforming baselines. Moreover, LogicCLIP retains, and often surpasses, competitive performance on general vision-language benchmarks, demonstrating that the enhanced logical understanding does not come at the expense of general alignment. We believe LogicBench and LogicCLIP will be important resources for advancing VLM logical capabilities.
Yuchen Zhou 0002, Jiayu Tang, Shuo Yang 0006, Xiaoyan Xiao, Yuqin Dai, Chao Gou, Xiaobo Xia, Tat-Seng Chua
AAAI3
2026 MEGG: replay via maximally extreme GGscore in incremental learning for neural recommendation models
abstract
Abstract Neural collaborative filtering (NCF)-based recommendation models have been widely adopted in practical recommender systems due to their effectiveness. However, these models are typically developed under the static deep learning paradigm, where training is conducted on fixed datasets with the implicit assumption of a static data distribution. This approach is ill-suited for dynamic environments, such as those encountered in real-world platforms, where user preferences and collaborative filtering patterns evolve continuously. To address this limitation, incremental learning-a paradigm designed to integrate new knowledge while preserving previously learned information-emerges as a promising alternative. Despite its potential, the direct application of conventional incremental learning methods, which are prevalent in domains like computer vision and natural language processing, is hindered by unique challenges in recommender systems. These include the distinct task paradigm, data complexity, and sparsity issues. Moreover, existing incremental learning approaches tailored for neural recommendation models remain scarce and often suffer from limited generalizability. To bridge this gap, we propose an innovative experience replay-based incremental learning framework specifically designed for neural recommendation models, termed Replay Samples with Maximally Extreme GGscore (MEGG). At the core of MEGG is a novel metric, the GGscore, which quantifies the influence of individual samples on model training. By selectively replaying samples with the most extreme GGscores, our method effectively mitigates catastrophic forgetting, thereby maintaining high predictive performance over time. A key advantage of MEGG lies in its data-centric nature, which renders it agnostic to the underlying model architecture. This ensures broad applicability across various neural recommendation models and seamless integration with existing incremental learning frameworks to further enhance performance. Extensive experiments conducted on three neural recommendation models across four benchmark datasets demonstrate the superior effectiveness of MEGG compared to state-of-the-art methods. Furthermore, additional evaluations highlight its scalability, efficiency, and robustness. The implementation of MEGG will be made publicly available upon acceptance.
Yunxiao Shi, Shuo Yang 0006, Haimin Zhang 0001, Li Wang 0064, Yongze Wang, Qiang Wu 0001, Min Xu 0001
Data Min. Knowl. Discov.2
2026 Noisy Correspondence Rectification in Multimodal Clustering Space for Cross-Modal Matching
abstract
As one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose BiCro++ (Improved Bidirectional Cross-modal Similarity Consistency). This module can be integrated into existing cross-modal matching models, enhancing their robustness against noisy data through self-adaptive soft labels that dynamically reflect the true correspondence of data pairs. The basic idea of BiCro++ is motivated by that - taking image-text matching as an example - similar images should have similar textual descriptions and vice versa. This bidirectional similarity consistency can be directly translated into soft labels as a self-supervision signal to train the matching model. To further refine soft label quality, BiCro++ first introduces a Diagonal-Dominance Purification process to identify reliable anchor points from noisy dataset as the reference for soft label estimation. Then it employs a Hybrid-level Codebook Alignment mechanism that establishes enhanced consistency in bidirectional cross-modal similarity. The experiments on three popular cross-modal matching datasets show that our method significantly improves the noise-robustness of various matching models, and surpasses the state-of-the-art method by an average of 5.3%, 3.1% and 6.4% in terms of recall, respectively.
Shuo Yang 0006, Yancheng Long, Zeke Xie, Hongxun Yao, Min Xu 0001, Liqiang Nie
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Learning from Ambiguous Data with Hard Labels
abstract
Real-world data often contains intrinsic ambiguity that the common single-hard-label annotation paradigm ignores. Standard training using ambiguous data with these hard labels may produce overly confident models and thus leading to poor generalization. In this paper, we propose a novel framework called Quantized Label Learning (QLL) to alleviate this issue. First, we formulate QLL as learning from (very) ambiguous data with hard labels: ideally, each ambiguous instance should be associated with a ground-truth soft-label distribution describing its corresponding probabilistic weight in each class, however, this is usually not accessible; in practice, we can only observe a quantized label, i.e., a hard label sampled (quantized) from the corresponding ground-truth soft-label distribution, of each instance, which can be seen as a biased approximation of the ground-truth soft-label. Second, we propose a Class-wise Positive-Unlabeled (CPU) risk estimator that allows us to train accurate classifiers from only ambiguous data with quantized labels. Third, to simulate ambiguous datasets with quantized labels in the real world, we design a mixing-based ambiguous data generation procedure for empirical evaluation. Experiments demonstrate that our CPU method can significantly improve model generalization performance and outperform the baselines.
Zeke Xie, Nan Lu 0001, Lichen Bai, Shuo Yang 0006, Mingming Sun 0001, Ping Li 0001
ICASSP6
2025 Stable Fair Graph Representation Learning with Lipschitz Constraint
abstract
Group fairness based on adversarial training has gained significant attention on graph data, which was implemented by masking sensitive attributes to generate fair feature views. However, existing models suffer from training instability due to uncertainty of the generated masks and the trade-off between fairness and utility. In this work, we propose a stable fair Graph Neural Network (SFG) to maintain training stability while preserving accuracy and fairness performance. Specifically, we first theoretically derive a tight upper Lipschitz bound to control the stability of existing adversarial-based models and employ a stochastic projected subgradient algorithm to constrain the bound, which operates in a block-coordinate manner. Additionally, we construct the uncertainty set to train the model, which can prevent unstable training by dropping some overfitting nodes caused by chasing fairness. Extensive experiments conducted on three real-world datasets demonstrate that SFG is stable and outperforms other state-of-the-art adversarial-based methods in terms of both fairness and utility performance. Codes are available at https://github.com/sh-qiangchen/SFG.
Qiang Chen 0016, Zhongze Wu, Xiu Su, Xi Lin 0003, Shan You, Shuo Yang 0006, Chang Xu 0002
ICML7
2025 BOOD: Boundary-based Out-Of-Distribution Data Generation
abstract
Harnessing the power of diffusion models to synthesize auxiliary training data based on latent space features has proven effective in enhancing out-of-distribution (OOD) detection performance. However, extracting effective features outside the in-distribution (ID) boundary in latent space remains challenging due to the difficulty of identifying decision boundaries between classes. This paper proposes a novel framework called Boundary-based Out-Of-Distribution data generation (BOOD), which synthesizes high-quality OOD features and generates human-compatible outlier images using diffusion models. BOOD first learns a text-conditioned latent feature space from the ID dataset, selects ID features closest to the decision boundary, and perturbs them to cross the decision boundary to form OOD features. These synthetic OOD features are then decoded into images in pixel space by a diffusion model. Compared to previous works, BOOD provides a more training efficient strategy for synthesizing informative OOD features, facilitating clearer distinctions between ID and OOD data. Extensive experimental results on common benchmarks demonstrate that BOOD surpasses the state-of-the-art method significantly, achieving a 29.64\% decrease in average FPR95 (40.31\% vs. 10.67\%) and a 7.27\% improvement in average AUROC (90.15\% vs. 97.42\%) on the Cifar-100 dataset.
Qilin Liao, Shuo Yang 0006, Bo Zhao 0038, Ping Luo 0002, Hengshuang Zhao
ICML2
2025 UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
abstract
Data augmentation using generative models has emerged as a powerful paradigm for enhancing performance in computer vision tasks. However, most existing augmentation approaches primarily focus on optimizing intrinsic data attributes -- such as fidelity and diversity -- to generate visually high-quality synthetic data, while often neglecting task-specific requirements. Yet, it is essential for data generators to account for the needs of downstream tasks, as training data requirements can vary significantly across different tasks and network architectures. To address these limitations, we propose UtilGen, a novel utility-centric data augmentation framework that adaptively optimizes the data generation process to produce task-specific, high-utility training data via downstream task feedback. Specifically, we first introduce a weight allocation network to evaluate the task-specific utility of each synthetic sample. Guided by these evaluations, UtilGen iteratively refines the data generation process using a dual-level optimization strategy to maximize the synthetic data utility: (1) model-level optimization tailors the generative model to the downstream task, and (2) instance-level optimization adjusts generation policies -- such as prompt embeddings and initial noise -- at each generation round. Extensive experiments on eight benchmark datasets of varying complexity and granularity demonstrate that UtilGen consistently achieves superior performance, with an average accuracy improvement of 3.87\% over previous SOTA. Further analysis of data influence and distribution reveals that UtilGen produces more impactful and task-relevant synthetic data, validating the effectiveness of the paradigm shift from visual characteristics-centric to task utility-centric data augmentation.
Jiyu Guo, Shuo Yang 0006, Yiming Huang 0001, Yancheng Long, Xiaobo Xia, Xiu Su, Bo Zhao 0038, Zeke Xie, Liqiang Nie
NeurIPS2
2025 L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
abstract
Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained in both contextual coverage and inference efficiency due to its inherently sequential process. To overcome these challenges, we propose leap multi-token prediction~(L-MTP), an innovative token prediction method that extends the capabilities of multi-token prediction (MTP) by introducing a leap-based mechanism. Unlike conventional MTP, which generates multiple tokens at adjacent positions, L-MTP strategically skips over intermediate tokens, predicting non-sequential ones in a single forward pass. This structured leap not only enhances the model's ability to capture long-range dependencies but also enables a decoding strategy specially optimized for non-sequential leap token generation, effectively accelerating inference. We theoretically demonstrate the benefit of L-MTP in improving inference efficiency. Experiments across diverse benchmarks validate its merit in boosting both LLM performance and inference speed. The source code is available at https://github.com/Xiaohao-Liu/L-MTP.
Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang, Xianzhi Yu, Xiu Su, Shuo Yang 0006, See-Kiong Ng, Tat-Seng Chua
NeurIPS7
2025 Visible-Infrared Person Re-Identification With Real-World Label Noise
abstract
In recent years, growing needs for advanced security and traffic management have significantly heightened the prominence of the visible-infrared person re-identification community (VI-ReID), garnering considerable attention. A critical challenge in VI-ReID is the performance degradation attributable to label noise, an issue that becomes even more pronounced in cross-modal scenarios due to an increased likelihood of data confusion. While previous methods have achieved notable successes, they often overlook the complexities of instance-dependent and real-world noise, creating a disconnect from the practical applications of person re-identification. To bridge this gap, our research analyzes the primary sources of label noise in real-world settings, which include a) instantiated identities, b) blurry infrared images, and c) annotators’ errors. In response to these challenges, we develop a Robust Hybrid Loss function (RHL) that enables targeted recognition and retrieval optimization through a more fine-grained division of the noisy dataset. The proposed method categorises data into three sets: clean, obviously noisy, and indistinguishably noisy, with bespoke loss calculations for each category. The identification loss is structured to address the varied nature of these sets specifically. For the retrieval sub-task, we utilize an enhanced triplet loss, adept at handling noisy correspondences. Furthermore, to empirically validate our method, we have re-annotated a real-world dataset, SYSU-Real. Our experiments on SYSU-MM01 and RegDB, conducted under various noise ratios of random and instance-dependent label noise, demonstrate the generalized robustness and effectiveness of our proposed approach.
Ruiheng Zhang 0001, Zhe Cao 0001, Yan Huang 0023, Shuo Yang 0006, Lixin Xu 0001, Min Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Cognition-Driven Structural Prior for Instance-Dependent Label Transition Matrix Estimation
abstract
The label transition matrix has emerged as a widely accepted method for mitigating label noise in machine learning. In recent years, numerous studies have centered on leveraging deep neural networks to estimate the label transition matrix for individual instances within the context of instance-dependent noise. However, these methods suffer from low search efficiency due to the large space of feasible solutions. Behind this drawback, we have explored that the real murderer lies in the invalid class transitions, that is, the actual transition probability between certain classes is zero but is estimated to have a certain value. To mask the invalid class transitions, we introduced a human-cognition-assisted method with structural information from human cognition. Specifically, we introduce a structured transition matrix network (STMN) designed with an adversarial learning process to balance instance features and prior information from human cognition. The proposed method offers two advantages: 1) better estimation effectiveness is obtained by sparing the transition matrix and 2) better estimation accuracy is obtained with the assistance of human cognition. By exploiting these two advantages, our method parametrically estimates a sparse label transition matrix, effectively converting noisy labels into true labels. The efficiency and superiority of our proposed method are substantiated through comprehensive comparisons with state-of-the-art methods on three synthetic datasets and a real-world dataset. Our code will be available at https://github.com/WheatCao/STMN-Pytorch.
Ruiheng Zhang 0001, Zhe Cao 0001, Shuo Yang 0006, Lingyu Si, Lixin Xu 0001, Fuchun Sun 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Revisiting Context Aggregation for Image Matting
abstract
Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle the context scale shift caused by the difference in image size during training and inference, resulting in matting performance degradation. In this paper, we revisit the context aggregation mechanisms of matting networks and find that a basic encoder-decoder network without any context aggregation modules can actually learn more universal context aggregation, thereby achieving higher matting performance compared to existing methods. Building on this insight, we present AEMatter, a matting network that is straightforward yet very effective. AEMatter adopts a Hybrid-Transformer backbone with appearance-enhanced axis-wise learning (AEAL) blocks to build a basic network with strong context aggregation learning capability. Furthermore, AEMatter leverages a large image training strategy to assist the network in learning context aggregation from data. Extensive experiments on five popular matting datasets demonstrate that the proposed AEMatter outperforms state-of-the-art matting methods by a large margin. The source code is available at https://github.com/aipixel/AEMatter.
Qinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li 0004, Xiangyuan Lan, Shuo Yang 0006, Shengping Zhang, Liqiang Nie
ICML6
2024 Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary
abstract
Existing paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily link the geometry of data distribution with models’ generalization capability in theoretics. Leveraging these theoretical insights, we propose a novel coreset construction method by selecting training samples to reconstruct the decision boundary of a deep neural network learned on the full dataset. Extensive experiments across various popular benchmarks demonstrate the superiority of our method over multiple competitors. For the first time, our method achieves a 50% data pruning rate on the ImageNet-1K dataset while sacrificing less than 1% in accuracy. Additionally, we showcase and analyze the remarkable cross-architecture transferability of the coresets derived from our approach.
Shuo Yang 0006, Zhe Cao 0001, Ruiheng Zhang 0001, Ping Luo 0002, Shengping Zhang, Liqiang Nie
ICML1
2024 Concentrating Estimation Attention: Human Prior Constrained Methods for Robust Classification
Zhe Cao 0001, Shuo Yang 0006, Hongbin Pei, Yan Huang 0023, Yushu Yu, Ruiheng Zhang 0001
PRCV (15)2
2024 Data-efficient Fine-tuning for LLM-based Recommendation
abstract
Leveraging Large Language Models (LLMs) for recommendation has recently garnered considerable attention, where fine-tuning plays a key role in LLMs' adaptation. However, the cost of fine-tuning LLMs on rapidly expanding recommendation data limits their practical application. To address this challenge, few-shot fine-tuning offers a promising approach to quickly adapt LLMs to new recommendation data. We propose the task of data pruning for efficient LLM-based recommendation, aimed at identifying representative samples tailored for LLMs' few-shot fine-tuning. While coreset selection is closely related to the proposed task, existing coreset selection methods often rely on suboptimal heuristic metrics or entail costly optimization on large-scale recommendation data.
Xinyu Lin 0001, Wenjie Wang 0007, Yongqi Li 0001, Shuo Yang 0006, Fuli Feng, Yinwei Wei, Tat-Seng Chua
SIGIR4
2024 U2D2Net: Unsupervised Unified Image Dehazing and Denoising Network for Single Hazy Image Enhancement
abstract
Hazy images captured under ill-posed scenarios with scattering medium (i.e. haze, fog, or smoke) are contaminated in visibility. Inevitably, these images are further degraded by noises owing to real-world imaging. Most existing hazy image enhancement methods perform image dehazing and denoising stage by stage, with the undesirable result that the estimation error of the former stage has to be propagated and amplified in the latter stage, e.g., noise amplification after dehazing. To address this inconsistent degradation, we present an Unsupervised Unified Image Dehazing and Denoising Network, U2D2Net, to remove the haze and suppress the noise simultaneously for a single hazy image. U2D2Net is mainly comprised of an unsupervised dehazing module, an unsupervised denoising module, and a region-similarity fusion strategy. Specifically, we propose an unsupervised transmission-aware dehazing module to restore visibility and suppress depth-dependent noise propagation in the dehazing module. Besides, we design an unsupervised network with a Mean/Max Sub-Sampler in the denoising module. To exploit the correlation and complementary between the previous outputs, a region-similarity fusion strategy is developed to compute the final qualified result. Extensive experiments on both synthetic and real-world datasets illustrate that U2D2Net outperforms other state-of-the-art dehazing and denoising methods in terms of PSNR, SSIM, and subjective visual effects.
Bosheng Ding, Ruiheng Zhang 0001, Lixin Xu 0001, Guanyu Liu, Shuo Yang 0006, Qi Zhang 0004
IEEE Trans. Multim.5
2023 BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency
abstract
As one of the most fundamental techniques in multi-modal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose a general framework called BiCro (Bidirectional Cross-modal similarity consistency), which can be easily integrated into existing cross-modal matching models and improve their robustness against noisy data. Specifically, BiCro aims to estimate soft labels for noisy data pairs to reflect their true correspondence degree. The basic idea of BiCro is motivated by that – taking image-text matching as an example – similar images should have similar textual descriptions and vice versa. Then the consistency of these two similarities can be recast as the estimated soft labels to train the matching model. The experiments on three popular cross-modal matching datasets demonstrate that our method significantly improves the noise-robustness of various matching models, and surpass the state-of-the-art by a clear margin. The code is available at https://github.com/xu5zhao/BiCro.
Shuo Yang 0006, Zhaopan Xu, Kai Wang 0036, Yang You 0001, Hongxun Yao, Tongliang Liu, Min Xu 0001
CVPR1
2023 Speech4Mesh: Speech-Assisted Monocular 3D Facial Reconstruction for Speech-Driven 3D Facial Animation
abstract
Recent audio2mesh-based methods have shown promising prospects for speech-driven 3D facial animation tasks. However, some intractable challenges are urgent to be settled. For example, the data-scarcity problem is intrinsically inevitable due to the difficulty of 4D data collection. Besides, current methods generally lack controllability on the animated face. To this end, we propose a novel framework named Speech4Mesh to consecutively generate 4D talking head data and train the audio2mesh network with the reconstructed meshes. In our framework, we first reconstruct the 4D talking head sequence based on the monocular videos. For precise capture of the talking-related variation on the face, we exploit the audio-visual alignment information from the video by employing a contrastive learning scheme. We next can train the audio2mesh network (e.g., FaceFormer) based on the generated 4D data. To get control of the animated talking face, we encode the speaking-unrelated factors (e.g., emotion, etc.) into an emotion embedding for manipulation. Finally, a differentiable renderer guarantees more accurate photometric details of the reconstruction and animation results. Empirical experiments demonstrate that the Speech4Mesh framework can not only outperform state-of-the-art reconstruction methods, especially on the lower-face part but also achieve better animation performance both perceptually and objectively after pre-trained on the synthesized data. Besides, we also verify that the proposed framework is able to explicitly control the emotion of the animated talking face.
Shuo Yang 0006, Pengcheng Xia 0002, Cong Liu 0006, Li-Rong Dai 0001, Chang Xu 0002
ICCV3
2023 Dataset Pruning: Reducing Training Data by Examining Generalization Influence
Shuo Yang 0006, Zeke Xie, Hanyu Peng, Min Xu 0001, Mingming Sun 0001, Ping Li 0001
ICLR1
2023 A Parametrical Model for Instance-Dependent Label Noise
abstract
In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited on class-dependent label-noise (wherein all samples in a clean class share the same label transition matrix). However, the CLTM cannot handle the more common instance-dependent label-noise well (wherein the clean-to-noisy label transition matrix needs to be estimated at the instance level by considering the input quality). Motivated by the fact that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-Label Transition Matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have no uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, this work proposes a parametrical model for estimating the instance-dependent label-noise transition matrix by employing a deep neural network, leading to better generalization and superior classification performance.
Shuo Yang 0006, Songhua Wu, Erkun Yang, Bo Han 0003, Yang Liu 0018, Min Xu 0001, Gang Niu 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Adversarial Recurrent Time Series Imputation
abstract
For the real-world time series analysis, data missing is a ubiquitously existing problem due to anomalies during data collecting and storage. If not treated properly, this problem will seriously hinder the classification, regression, or related tasks. Existing methods for time series imputation either impose too strong assumptions on the distribution of missing data or cannot fully exploit, even simply ignore, the informative temporal dependencies and feature correlations across different time steps. In this article, inspired by the idea of conditional generative adversarial networks, we propose a generative adversarial learning framework for time series imputation under the condition of observed data (as well as the labels, if possible). In our model, we employ a modified bidirectional RNN structure as the generator G, which is aimed at generating the missing values by taking advantage of the temporal and nontemporal information extracted from the observed time series. The discriminator D is designed to distinguish whether each value in a time series is generated or not so that it can help the generator to make an adjustment toward a more authentic imputation result. For an empirical verification of our model, we conduct imputation and classification experiments on several real-world time series data sets. The experimental results show an eminent improvement compared with state-of-the-art baseline models.
Shuo Yang 0006, Minjing Dong, Yunhe Wang 0001, Chang Xu 0002
IEEE Trans. Neural Networks Learn. Syst.1
2022 CAFE: Learning to Condense Dataset by Aligning Features
abstract
Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intuitive motivation and promising results, such gradient-based methods, by nature, easily overfit to a biased set of samples that produce dominant gradients, and thus lack a global supervision of data distribution. In this paper, we propose a novel scheme to Condense dataset by Aligning FEatures (CAFE), which explicitly attempts to preserve the real-feature distribution as well as the discriminant power of the resulting synthetic set, lending itself to strong generalization capability to various architectures. At the heart of our approach is an effective strategy to align features from the real and synthetic data across various scales, while accounting for the classification of real samples. Our scheme is further backed up by a novel dynamic bi-level optimization, which adaptively adjusts parameter updates to prevent over-/under-fitting. We validate the proposed CAFE across various datasets, and demonstrate that it generally outperforms the state of the art: on the SVHN dataset, for example, the performance gain is up to 11%. Extensive experiments and analysis verify the effectiveness and necessity of proposed designs.
Kai Wang 0036, Bo Zhao 0038, Shuo Yang 0006, Shuo Wang 0001, Guan Huang 0003, Hakan Bilen, Xinchao Wang, Yang You 0001
CVPR5
2022 One Size Does NOT Fit All: Data-Adaptive Adversarial Training
Shuo Yang 0006, Chang Xu 0002
ECCV (5)1
2022 Self-Attention Gated Cognitive Diagnosis For Faster Adaptive Educational Assessments
abstract
Cognitive diagnosis models map observations onto psychological internal traits, which have been widely used to construct personalized educational assessments. However, the challenges of the latent ability representation, accurate performance with explanatory and cold-start problems caused by limited handcraft skills still exist. In this paper, we introduce a selfadaptive Attention Gate Cognitive Diagnosis Model (AGCDM) based on a multi-layer hierarchical structure. Specifically, in the first two layers, our model embeds the item response logs into a higher expressive latent space, and the specific self-attention mechanism captures the rich information from the inner-dependencies correlation among the examinee, quiz and knowledge. In the next layer, the gate mechanism addresses the challenge of noisy information (slip and guess) over psychological reasons. We verify the empirical performance of our model on different educational tasks and all comparisons study conducted on the same training recipe. On the item response prediction task, the experiment results demonstrate our hierarchical model achieves state-of-the-art performance. Through the estimation of adaptive learning, we also validate the model’s effectiveness in addressing the limited skill problem in real scenarios. The learning tracks show that our model empowered by the meta-learning algorithm can be adaptive faster than the base model to the new different skill distribution set. Furthermore, we present the ablation study on the hierarchical architecture and detailed hyperparameters selections. The comparison results demonstrate the attention blocks yield the scores improvement, and the estimation of the noisy errors shows the gate mechanism boosts the detection accuracy on the slip and guess. Code is available at https://github.com/TerryPei/AGCDM.
Xiaohuan Pei, Shuo Yang 0006, Chang Xu 0002
ICDM2
2022 Objects in Semantic Topology
Shuo Yang 0006, Peize Sun, Yi Jiang 0009, Xiaobo Xia, Ruiheng Zhang 0001, Zehuan Yuan, Changhu Wang, Ping Luo 0002, Min Xu 0001
ICLR1
2022 Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network
abstract
In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to learn a clean label classifier by employing the noisy data. Motivated by that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-label transition matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have less uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, we estimate the BLTM parametrically by employing a deep neural network, leading to better generalization and superior classification performance.
Shuo Yang 0006, Erkun Yang, Bo Han 0003, Yang Liu 0018, Min Xu 0001, Gang Niu 0001, Tongliang Liu
ICML1
2022 An efficient multitask neural network for face alignment, head pose estimation and face tracking
abstract
While Convolutional Neural Networks (CNNs) have significantly boosted the performance of face related algorithms, maintaining accuracy and efficiency simultaneously in practical use remains challenging. The state-of-the-art methods employ deeper networks for better performance, which makes it less practical for mobile applications because of more parameters and higher computational complexity. Therefore, we propose an efficient multitask neural network, Alignment & Tracking & Pose Network (ATPN) for face alignment, face tracking and head pose estimation . Specifically, to achieve better performance with fewer layers for face alignment, we introduce a shortcut connection between shallow-layer and deep-layer features. We find the shallow-layer features are highly correspond to facial boundaries that can provide the structural information of face and it is crucial for face alignment. Moreover, we generate a cheap heatmap based on the face alignment result and fuse it with features to improve the performance of the other two tasks. Based on the heatmap, the network can utilize both geometric information of landmarks and appearance information for head pose estimation. The heatmap also provides attention clues for face tracking. The face tracking task also saves us the face detection procedure for each frame, which also significantly boost the real-time capability for video-based tasks. We experimentally validate ATPN on four benchmark datasets, WFLW, 300VW, WIDER Face and 300W-LP. The experimental results demonstrate that it achieves better performance with much less parameters and lower computational complexity compared to other light models.
Jiahao Xia 0001, Haimin Zhang 0001, Shiping Wen 0001, Shuo Yang 0006, Min Xu 0001
Expert Syst. Appl.4
2022 Graph-based few-shot learning with transformed feature propagation and optimal class allocation
Ruiheng Zhang 0001, Shuo Yang 0006, Qi Zhang 0004, Lixin Xu 0001, Yang He 0002, Fan Zhang 0007
Neurocomputing2
2022 Bridging the Gap Between Few-Shot and Many-Shot Learning via Distribution Calibration
abstract
A major gap between few-shot and many-shot learning is the data distribution empirically oserved by the model during training. In few-shot learning, the learned model can easily become over-fitted based on the biased distribution formed by only a few training examples, while the ground-truth data distribution is more accurately uncovered in many-shot learning to learn a well-generalized model. In this paper, we propose to calibrate the distribution of these few-sample classes to be more unbiased to alleviate such an over-fitting problem. The distribution calibration is achieved by transferring statistics from the classes with sufficient examples to those few-sample classes. After calibration, an adequate number of examples can be sampled from the calibrated distribution to expand the inputs to the classifier. Specifically, we assume every dimension in the feature representation from the same class follows a Gaussian distribution so that the mean and the variance of the distribution can borrow from that of similar classes whose statistics are better estimated with an adequate number of samples. Extensive experiments on three datasets,miniImageNet,tieredImageNet, and CUB, show that a simple linear classifier trained using the features sampled from our calibrated distribution can outperform the state-of-the-art accuracy by a large margin. Besides the favorable performance, the proposed method also exhibits high flexibility by showing consistent accuracy improvement when it is built on top of any off-the-shelf pretrained feature extractors and classification models without extra learnable parameters. The visualization of these generated features demonstrates that our calibrated distribution is an accurate estimation thus the generalization ability gain is convincing. We also establish a generalization error bound for the proposed distribution-calibration-based few-shot learning, which consists of thedistribution assumption error, thedistribution approximation error, and theestimation error. This generalization error bound theoretically justifies the effectiveness of the proposed method.
Shuo Yang 0006, Songhua Wu, Tongliang Liu, Min Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Adversarial Robustness through Disentangled Representations
abstract
Despite the remarkable empirical performance of deep learning models, their vulnerability to adversarial examples has been revealed in many studies. They are prone to make a susceptible prediction to the input with imperceptible adversarial perturbation. Although recent works have remarkably improved the model's robustness under the adversarial training strategy, an evident gap between the natural accuracy and adversarial robustness inevitably exists. In order to mitigate this problem, in this paper, we assume that the robust and non-robust representations are two basic ingredients entangled in the integral representation. For achieving adversarial robustness, the robust representations of natural and adversarial examples should be disentangled from the non-robust part and the alignment of the robust representations can bridge the gap between accuracy and robustness. Inspired by this motivation, we propose a novel defense method called Deep Robust Representation Disentanglement Network (DRRDN). Specifically, DRRDN employs a disentangler to extract and align the robust representations from both adversarial and natural examples. Theoretical analysis guarantees the mitigation of the trade-off between robustness and accuracy with good disentanglement and alignment performance. Experimental results on benchmark datasets finally demonstrate the empirical superiority of our method.
Shuo Yang 0006, Tianyu Guo 0001, Yunhe Wang 0001, Chang Xu 0002
AAAI1
2021 Single-View 3D Object Reconstruction From Shape Priors in Memory
abstract
Existing methods for single-view 3D object reconstruction directly learn to transform image features into 3D representations. However, these methods are vulnerable to images containing noisy backgrounds and heavy occlusions because the extracted image features do not contain enough information to reconstruct high-quality 3D shapes. Humans routinely use incomplete or noisy visual cues from an image to retrieve similar 3D shapes from their memory and reconstruct the 3D shape of an object. Inspired by this, we propose a novel method, named Mem3D, that explicitly constructs shape priors to supplement the missing information in the image. Specifically, the shape priors are in the forms of "image-voxel" pairs in the memory network, which is stored by a well-designed writing strategy during training. We also propose a voxel triplet loss function that helps to retrieve the precise 3D shapes that are highly related to the input image from shape priors. The LSTM-based shape encoder is introduced to extract information from the retrieved 3D shapes, which are useful in recovering the 3D shape of an object that is heavily occluded or in complex environments. Experimental results demonstrate that Mem3D significantly improves reconstruction quality and performs favorably against state-of-the-art methods on the ShapeNet and Pix3D datasets.
Shuo Yang 0006, Min Xu 0001, Haozhe Xie, Stuart W. Perry, Jiahao Xia 0001
CVPR1
2021 Structure-Aware Stabilization of Adversarial Robustness with Massive Contrastive Adversaries
abstract
Recent researches indicate that the impact of adversarial perturbations on deep learning models is reflected not only on the alteration of predicted labels but also on the distortion of data structure in the representation space. Significant improvement of the model’s adversarial robustness can be achieved by reforming the structure-aware representation distortion. Current methods generally utilize the one-to-one representation alignment or the triplet information between the positive and negative pairs. However, in this paper, we show that the representation structure of the natural and adversarial examples cannot be well and stably captured if we only focus on a localized range of contrastive examples. To achieve better and more stable adversarial robustness, we propose to adjust the adversarial distortion of representation structure by using Massive Contrastive Adversaries (MCA). Inspired by the Noise-Contrastive Estimation (NCE), MCA exploits the contrastive information by employing m negative instances. Compared with existing methods, our method recruits a much wider range of negative examples per update, so a better and more stable representation relationship between the natural and adversarial examples can be captured. Theoretical analysis shows that the proposed MCA inherently maximizes a lower bound of the mutual information (MI) between the representations of the natural and adversarial examples. Empirical experiments on benchmark datasets demonstrate that MCA can achieve better and more stable intra-class compactness and inter-class divergence, which further induces better adversarial robustness.
Shuo Yang 0006, Zeyu Feng, Bo Du 0001, Chang Xu 0002
ICDM1
2021 Free Lunch for Few-shot Learning: Distribution Calibration
Shuo Yang 0006, Min Xu 0001
ICLR1