VLDB 2026 Research / reviewers in the wild / expert
Yangyang Shu
dblp:201/7247
· DBLP profile ↗
16ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPASRNN: boosting spiking recurrent neural networks using sparse gradient decent for semantic representation learning
Yansong Chua, Qian Zhang 0035, Yangyang Shu |
Neural Comput. Appl. | 6 |
| 2026 | LearnMat: Semantic-Aware Self-Supervision Fine-Grained Visual RecognitionabstractSelf-supervised learning has shown potential in fine-grained visual recognition (FGVR). However, existing self-supervised learning methods are often susceptible to irrelevant patterns during training and lack the ability to capture the critical subtle differences in FGVR, leading to suboptimal performance. Moreover, existing approaches focus primarily on uni-modal visual concepts. Despite the emergence of powerful vision-language models (VLMs) in various high-level vision tasks, their potential in self-supervised FGVR remains largely unexplored. To this end, we propose a novel self-supervised learning (LearnMat) framework, that effectively filters out irrelevant feature interference and extracts more important and subtle discriminative features during training. Specifically, LearnMat consists of two key modules: the semantic awareness module (SAM) and the insight extraction module (IEM). In the SAM, we introduce a novel vision-language-grounded semantic distillation strategy using a corpus of generic, category-agnostic textual attributes, that injects explicit semantic constraints into self-supervised training and improves robustness to background interference. Complementarily, the IEM exploits gradient-based signals from the input image to highlight subtle differences and localize key discriminative regions, mitigating inter-class similarity and intra-class variation, and enhancing fine-grained discrimination. Extensive experiments across multiple popular FGVR datasets show that LearnMat significantly outperforms recent state-of-the-art methods, highlighting its marked effectiveness. Our code is avaliable at https://github.com/Heng-CHY/LearnMat. ShuaiHeng Li, Fan Zhang 0070, Yangyang Shu, Guanbin Li, Junyu Dong, Lingqiao Liu, David Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
Yangyang Shu, Hye-Young Paik, Yulei Sui |
ICFEM | 2 |
| 2025 | MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention FusionabstractThe combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has attracted significant attention due to the great potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT, a novel spike-driven Transformer architecture, which firstly uses multi-scale spiking attention (MSSA) to enrich the capability of spiking attention blocks. We validate our approach across various main data sets. The experimental results indicate that our MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among NN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT. Chenlin Zhou, Jibin Wu, Yansong Chua, Yangyang Shu |
IJCAI | 5 |
| 2025 | CIT: Rethinking class-incremental semantic segmentation with a Class Independent TransformationabstractClass-incremental semantic segmentation (CSS) requires that a model learn to segment new classes without forgetting how to segment previous ones: this is typically achieved by distilling the current knowledge and incorporating the latest data. However, bypassing iterative distillation by directly transferring outputs of initial classes to the current learning task is not supported in existing class-specific CSS methods. Via Softmax, they enforce dependency between classes and adjust the output distribution at each learning step, resulting in a large probability distribution gap between initial and current tasks. We introduce a simple, yet effective Class Independent Transformation (CIT) that converts the outputs of existing semantic segmentation models into class-independent forms with negligible cost or performance loss. By utilizing class-independent predictions facilitated by CIT, we establish an accumulative distillation framework, ensuring equitable incorporation of all class information. We conduct extensive experiments on various segmentation architectures, including DeepLabV3, Mask2Former, and SegViTv2. Results from these experiments show minimal task forgetting across different datasets, with less than 5% for ADE20K in the most challenging 11 task configurations and less than 1% across all configurations for the PASCAL VOC 2012 dataset. • Softmax interdependency causes incremental forgetting in continual learning. • We introduce a class-independent transformation (CIT) to reduce forgetting. • CIT reformulates segmentation as class-agnostic, enhancing CSS training pipelines. • Our method significantly reduces forgetting on ADE20K compared to CSS baselines. • CIT achieves near-zero forgetting ( ≤ 1%) in Pascal-VOC 2012 settings. Jinchao Ge, Bowen Zhang 0009, Akide Liu, Vu Minh Hieu Phan, Qi Chen 0014, Yangyang Shu, Yang Zhao 0019 |
Pattern Recognit. | 6 |
| 2024 | Unlocking the Potential of Pre-Trained Vision Transformers for Few-Shot Semantic Segmentation through Relationship DescriptorsabstractThe recent advent of pre-trained vision transformers has unveiled a promising property: their inherent capability to group semantically related visual concepts. In this paper, we explore to harnesses this emergent feature to tackle few-shot semantic segmentation, a task focused on classifying pixels in a test image with a few example data. A critical hurdle in this endeavor is preventing overfitting to the limited classes seen during training the few-shot segmentation model. As our main discovery, we find that the concept of “relationship descriptors”, initially conceived for enhancing the CLIP model for zero-shot semantic segmentation, offers a potential solution. We adapt and refine this concept to craft a relationship descriptor construction tailored for few-shot semantic segmentation, extending its application across multiple layers to enhance performance. Building upon this adaptation, we proposed a few-shot semantic segmentation framework that is not only easy to implement and train but also effectively scales with the number of support examples and categories. Through rigorous experimentation across various datasets, including PASCAL-5iand COCO-20i, we demonstrate a clear advantage of our method in diverse few-shot semantic segmentation scenarios, and a range of pre-trained vision transformer models. The findings clearly show that our method significantly outperforms current state-of-the-art techniques, highlighting the effectiveness of harnessing the emerging capabilities of vision transformers for few-shot semantic segmentation. We release the code at https://github.com/ZiqinZhou66/FewSegwithRD.git. Ziqin Zhou, Yangyang Shu, Lingqiao Liu |
CVPR | 3 |
| 2024 | Semi-Supervised Adversarial Learning for Attribute-Aware Photo Aesthetic AssessmentabstractAesthetic attributes are crucial for aesthetics because they explicitly present some photo quality cues that a human expert might use to evaluate a photo’s aesthetic quality. However, annotating aesthetic attributes is a time-consuming, costly, and error-prone task, which leads to the issue that photos available are partially annotated with attributes. To alleviate this issue, we propose a novel semi-supervised adversarial learning method for photo aesthetic assessment from partially attribute-annotated photos, which can greatly reduce the reliance on manual attribute annotation. Specifically, the proposed method consists of a score-attributes generator$R$, a photo generator$G$, and a discriminator$D$. The score-attributes generator learns the aesthetic score and attributes simultaneously to capture their dependencies and construct better feature representations. The photo generator reconstructs the photo by feeding aesthetic attributes, score, and informative feature representation. A discriminator is used to force the convergence of the features-attributes-score tuples generated from the score-attributes generator, the photo generator, and the ground-truth distribution in labeled data for training data. The proposed method significantly outperforms the state of the art, increasing the Spearman rank-order correlation coefficient (SRCC) from the existing best reported of 0.726 to 0.761 onAesthetic and attributes databaseand 0.756 to 0.774 onAesthetic visual analysis database, respectively. Yangyang Shu, Qian Li 0003, Lingqiao Liu, Guandong Xu |
IEEE Trans. Multim. | 1 |
| 2023 | Learning Common Rationale to Improve Self-Supervised Representation for Fine-Grained Visual Recognition ProblemsabstractSelf-supervised learning (SSL) strategies have demonstrated remarkable performance in various recognition tasks. However, both our preliminary investigation and recent studies suggest that they may be less effective in learning representations for fine-grained visual recognition (FGVR) since many features helpful for optimizing SSL objectives are not suitable for characterizing the subtle differences in FGVR. To overcome this issue, we propose learning an additional screening mechanism to identify discriminative clues commonly seen across instances and classes, dubbed as common rationales in this paper. Intuitively, common rationales tend to correspond to the discriminative patterns from the key parts of foreground objects. We show that a common rationale detector can be learned by simply exploiting the GradCAM induced from the SSL objective without using any pre-trained object parts or saliency detectors, making it seamlessly to be integrated with the existing SSL process. Specifically, we fit the GradCAM with a branch with limited fitting capacity, which allows the branch to capture the common rationales and discard the less common discriminative patterns. At the test stage, the branch generates a set of spatial weights to selectively aggregate features representing an instance. Extensive experimental results on four visual tasks demonstrate that the proposed method can lead to a significant improvement in different evaluation settings.11The source code will be publicly available at:https://github.com/GANPerf/LCR Yangyang Shu, Anton van den Hengel, Lingqiao Liu |
CVPR | 1 |
| 2022 | Improving Fine-Grained Visual Recognition in Low Data Regimes via Self-boosting Attention Mechanism
Yangyang Shu, Baosheng Yu, Lingqiao Liu |
ECCV (25) | 1 |
| 2022 | Privileged multi-task learning for attribute-aware aesthetic assessment
Yangyang Shu, Qian Li 0003, Lingqiao Liu, Guandong Xu |
Pattern Recognit. | 1 |
| 2022 | V-SVR+: Support Vector Regression With Variational Privileged InformationabstractMany regression tasks encounter an asymmetric distribution of information between training and testing phases where the additional information available in training, the so-called privileged information (PI), is often inaccessible in testing. In practice, the privileged information in training data might be expressed in different formats, such as continuous, ordinal, or binary values. However, most the existing learning using privileged information (LUPI) paradigms primarily deal with the continuous form of PI, preventing them from managing variational PI, which motivates this research. Therefore, in this paper, we propose a unified framework to systematically address the aforementioned three forms of privileged information. The proposed V-SVR+ method integrates continuous, ordinal, and binary PI into the learning process of support vector regression (SVR) via three losses. For continuous privileged information, we define a linear correcting (slack) function in the privileged information space to estimate slack variables in the standard SVR method using privileged information. For the ordinal relations of privileged information, we first rank the privileged information and then, regard this ordinal privileged information as auxiliary information used in the learning process of the SVR model. For the binary or Boolean privileged information, we infer a probabilistic dependency between the privileged information and labels from the summarized privileged information knowledge. Then, we transfer the privileged information knowledge to constraints and form a constrained optimization problem. We evaluate the proposed method in three applications: music emotion recognition from songs with the help of implicit information about music elements judged by composers; multiple object recognition from images with the help of implicit information about the object’s importance conveyed by the list of manually annotated image tags; and photo aesthetic assessment enhanced by high-level aesthetic attributes hidden in photos. Experiment results demonstrate that the proposed methods are superior to the classic learning paradigm when solving practical problems. Yangyang Shu, Qian Li 0003, Chang Xu 0002, Shaowu Liu, Guandong Xu |
IEEE Trans. Multim. | 1 |
| 2021 | Video Affective Content Analysis by Exploring Domain KnowledgeabstractFilm grammar is often used to invoke certain emotional experiences from audiences through changing visual, speech, and musical elements of videos. Such film grammar, referred to as domain knowledge, is of great importance for video affective content analysis but has not been thoroughly examined in research. In this paper, we propose an improved method for emotion recognition and regression from videos through exploring domain knowledge. We first investigate the domain knowledge of visual, speech, and musical elements, and infer probabilistic dependencies between elements and emotions from the summarized film grammar. Then, we transfer the summarized dependencies between elements and emotions as constraints, and formulate video affective content analysis, including both emotion recognition and emotion regression from video content, as a constrained optimization problem. Experiments on the LIRIS-ACCEDE database, the FilmStim database, and the DEAP database demonstrate that the proposed video affective content analysis method can successfully leverage well-established film grammar to improve emotion recognition and regression from video content. Shangfei Wang, Can Wang 0007, Tanfang Chen, Yangyang Shu |
IEEE Trans. Affect. Comput. | 5 |
| 2020 | Perf-AL: Performance Prediction for Configurable Software through Adversarial LearningabstractContext: Many software systems are highly configurable. Different configuration options could lead to varying performances of the system. It is difficult to measure system performance in the presence of an exponential number of possible combinations of these options. Yangyang Shu, Yulei Sui, Hongyu Zhang 0002, Guandong Xu |
ESEM | 1 |
| 2020 | Learning with privileged information for photo aesthetic assessment
Yangyang Shu, Qian Li 0003, Shaowu Liu, Guandong Xu |
Neurocomputing | 1 |
| 2019 | Emotion Recognition from Music Enhanced by Domain Knowledge
Yangyang Shu, Guandong Xu |
PRICAI (1) | 1 |
| 2017 | Emotion recognition through integrating EEG and peripheral signalsabstractThe inherent dependencies among multiple physiological signals are crucial for multimodal emotion recognition, but have not been thoroughly exploited yet. This paper propose to use restricted Boltzmann machine (RBM) to model such dependencies.Specifically, the visible nodes of RBM represent EEG and peripheral physiological signals, and thus the connections between visible nodes and hidden nodes capture the intrinsic relations among multiple physiological signals. The RBM generates new representation from multiple physiological signals. Then, a support vector machine is adopted to recognize users' emotion states from the generated features. Furthermore, we extend the proposed fusion method for incomplete datas, since physiological signals are often corrupted due to artifacts. Specifically, we pre-train the RBM using all the complete data, then we update missing values and RBM parameters to minimize free energy of visible vectors using both complete and incomplete data. Experiments on two benchmark databases demonstrate the effectiveness of the proposed methods. Yangyang Shu, Shangfei Wang |
ICASSP | 1 |