VLDB 2026 Research / reviewers in the wild / expert
Sai Yang
dblp:28/5126
· DBLP profile ↗
20ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Vision Mamba With Contrastive Regularization Network for Single Image DehazingabstractABSTRACT Benefiting from the powerful nonlinear fitting ability of neural networks, deep learning‐based methods have gradually emerged as the dominate solutions for single image dehazing. However, supervised learning–based methods require paired image samples for training. To address this, an unsupervised Vision Mamba with Contrastive Regularization network (VMCR) is proposed. The network is designed based on the DisentGAN framework, and its main module is the Vision Mamba. This module performs very competitively compared to transformers, while maintaining linear time complexity and constant memory complexity with respect to the input size. Furthermore, a contrastive regularization method based on contrastive learning is proposed to enhance the reconstruction capabilities of the network and achieve superior dehazing results. Our VMCR‐Net outperforms state‐of‐the‐art unsupervised image dehazing methods, as evidenced by experimental results on several benchmarks. This research successfully proposes an enhanced unsupervised image dehazing approach, overcoming the limitations of existing methods and achieving superior dehazing performance. Bin Hu 0023, Wanzhi Wen, Sai Yang |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | Image restoration model compression via mamba-oriented heterogeneous knowledge distillation
Sai Yang, Bin Hu 0023, Xiaoxin Wu 0004, Fan Liu 0003, Wanzhi Wen |
Neural Networks | 1 |
| 2025 | Collaborative Semantic Contrastive for All-in-one Image Restoration
Bin Hu 0023, Sai Yang, Fan Liu 0003, Weiping Ding 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | BTCDNet: Bayesian Tile Attention Network for Hyperspectral Image Change DetectionabstractHyperspectral images (HSI) provide detailed spectral information, which are effective for change detection (CD). Prior knowledge has been proven to improve the robustness of models in HSI processing. However, current CD methods do not fully utilize prior knowledge and research on hyperspectral mangroves CD is limited. In this letter, we propose a general hyperspectral CD model with Bayesian prior guided module (BPGM) and tile attention block (TAB) called BTCDNet. BPGM leverages prior information to steer the model training process under limited labeled samples condition, while TAB can reduce complexity and improve performance by tile attention. Moreover, a novel and restricted hyperspectral CD dataset Shenzhen has been annotated for hyperspectral mangroves CD reference. Experiments demonstrate that our proposal achieves state-of-the-art (SOTA) performances on this dataset and two other public benchmark datasets. Our code and datasets are available at https://github.com/JeasunLok/BTCDNet. Junshen Luo, Jiahe Li 0017, Xinlin Chu, Sai Yang, Lingjun Tao, Qian Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | IPT-ILR: Image Pyramid Transformer Coupled With Information Loss Regularization for All-in-One Image RestorationabstractAll-in-one image restoration has recently developed to be a new research trend in the low-level computer vision field, aiming to tackle multiple image degradation types simultaneously in a unified model. As a typical multi-task learning, existing approaches focus on modeling either the specificity or commonality among different image restoration tasks. To exploit the unique strengths of both worlds, we propose a method of Image Pyramid Transformer coupled with Information Loss Regularization (IPT-ILR), in which the multi-scale architecture structure can excavate more information for multiple restoration tasks concurrently, while the learning strategy can identify the difference among multiple restoration tasks depending on the degree of information loss in each restoration task. Specifically, it first establishes a new Image Pyramid Transformer Network (IPT-Network) to accommodate multiple image restoration tasks. Given original degraded images, the IPT-Network exploits the image pyramid technique to establish a series of images with different scales, which are then restored by transformer-like auto-encoders. Moreover, the restored image on a low-level scale is referenced to assist restoring the degraded image on a high-level scale. Next, Information Loss Regularization (ILR) is presented to optimize the IPT-Network. ILR calculates the average distance between degraded images and their clean counterparts as the weights, which automatically implement different penalties for different image restoration tasks, thus avoiding the short-cut phenomenon for the easy task while encouraging the hard task. Extensive experiments have been conducted with 6 image restoration tasks in the all-in-one setting. The results show our method performs favorably against numerous state-of-the-art methods across most tasks, including image denoising, image deblurring, image dehazing, image deraining, image desnowing, as well as low-light enhancement. Sai Yang, Bin Hu 0023, Fan Liu 0003, Xiaoxin Wu 0004, Weiping Ding 0001, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Few-shot classification guided by generalization error bound
Fan Liu 0003, Sai Yang, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
Pattern Recognit. | 2 |
| 2024 | Few-Shot Classification Model Compression via School LearningabstractFew-shot classification (FSC) is a challenging task due to limitation in accessing training data. Recent methods often employ highly complex networks to obtain high-quality features, but this may not be suitable for resource-limited applications. To tackle this challenge, we introduce Few-Shot Classification Model Compression (FSC-MC), a new task aimed at enhancing the FSC performance of lightweight and low-capacity models by learning from more complex models. We also propose a novel two-level learning strategy called School Learning to accomplish the FSC-MC task by mimicking the real learning process in the social school life. In this new learning paradigm, the first level performs preview learning, in which each student is equipped with a preparer to perform self-learning on the base set. The second level is the team learning, consisting of a complex teacher network and several lightweight student networks organized into a team. One student network is randomly chosen as the leader network, while the remaining student networks serve as member networks. The leader network simultaneously learns knowledge from the teacher network and all member networks. Conversely, each member network receives knowledge from both the teacher network and the leader network. Ultimately, the leader network is deployed for FSC evaluation, resulting in effective model compression. Extensive experiments in the FSC-MC setting demonstrate that School Learning outperforms 17 state-of-the-art knowledge distillation methods including both offline methods and online methods, enabling lightweight models to achieve outstanding FSC performance. Sai Yang, Fan Liu 0003, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Few-shot Classification via Ensemble Learning with Multi-Order StatisticsabstractTransfer learning has been widely adopted for few-shot classification. Recent studies reveal that obtaining good generalization representation of images on novel classes is the key to improving the few-shot classification accuracy. To address this need, we prove theoretically that leveraging ensemble learning on the base classes can correspondingly reduce the true error in the novel classes. Following this principle, a novel method named Ensemble Learning with Multi-Order Statistics (ELMOS) is proposed in this paper. In this method, after the backbone network, we use multiple branches to create the individual learners in the ensemble learning, with the goal to reduce the storage cost. We then introduce different order statistics pooling in each branch to increase the diversity of the individual learners. The learners are optimized with supervised losses during the pre-training phase. After pre-training, features from different branches are concatenated for classifier evaluation. Extensive experiments demonstrate that each branch can complement the others and our method can produce a state-of-the-art performance on multiple few-shot classification benchmark datasets. Sai Yang, Fan Liu 0003, Delong Chen, Jun Zhou 0001 |
IJCAI | 1 |
| 2023 | Safety evaluation of buildings adjacent to shield construction in karst areas: An improved extension cloud approach
Hongyu Chen 0007, Sai Yang, Zongbao Feng, Yang Liu 0261, Yawei Qin |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Intelligent multiobjective optimization for high-performance concrete mix proportion design: A hybrid machine learning approach
Sai Yang, Hongyu Chen 0007, Zongbao Feng, Yawei Qin, Yang Liu 0261 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Few-shot classification using Gaussianisation prototypical classifierabstractAbstract Few‐shot classification (FSC) aims at classifying query samples into correct classes given only a few labelled samples. Prototypical Classifier (PC) can be chosen to be an ideal classifier for settling this problem, as it has good properties of low‐capacity and parameter‐free. However, the mean‐based prototypes suffer from the issue of deviating from its ground‐truth centre. In order to solve such problem of prototype bias, Gaussianisation Prototypical Classifier (GPC) is proposed, which is a kind of one‐step prototype rectification method. Specifically, the authors first perform Gaussianisation operation over the feature extracted from the backbone network so that the features fit the particular Gaussian distribution. Second, the authors use prototype feature of the base class as prior information and employs Maximum a Posteriori estimation method to obtain the reliable prototype for each novel class. Finally, the query sample of novel class is classified to be its nearest prototype with non‐parametric classifiers. Extensive experiments have been conducted on multiple FSC benchmarks. Comparative results also demonstrate that the authors’ method is superior to existing state‐of‐the‐art FSC methods. Fan Liu 0003, Sai Yang |
IET Comput. Vis. | 3 |
| 2023 | Lite general network and MagFace CNN for micro-expression spotting in long videos
Quan-Lin Gu, Sai Yang, Tianxing Yu |
Multim. Syst. | 2 |
| 2022 | Feature hallucination in hypersphere space for few-shot classificationabstractAbstract Few‐shot classification (FSC) targeting at classifying unseen classes with few labelled samples is still a challenging task. Recent works show that transfer‐learning based approaches are competitive with meta‐learning ones, which usually pre‐train a convolutional neural networks (CNN)‐based network using cross‐entropy (CE) loss and throw away the last layer to post‐process the novel classes. Hereby, they still suffer the issue of getting a more transferable extractor and lacking enough labelled novel samples. Thus, the authors propose the algorithm of feature hallucination in hypersphere space (FHHS) for FSC. On the first stage, the authors pre‐train a more transferable feature extractor using a hypersphere loss (HL), which supplies CE with supervised contrastive (SC) loss and self‐supervised loss (SSL), in which SC can map the base and novel images onto the hypersphere space densely. On the second stage, the authors generate new samples for unseen classes using their novel algorithm of synthetic novel sampling with the base (SNSB), which linearly interpolate between each novel class prototype and its K nearest neighbour base class prototypes. Comprehensive experiments on multiple popular FSC demonstrate that HL loss can enhance the performance of backbone network and the authors’ feature hallucination method is superior to the existing hallucination‐based methods. Sai Yang, Fan Liu 0003 |
IET Image Process. | 1 |
| 2022 | Self-Supervised Music Motion Synchronization Learning for Music-Driven Conducting Motion Generation
Fan Liu 0003, Delong Chen, Ruizhi Zhou, Sai Yang, Feng Xu 0008 |
J. Comput. Sci. Technol. | 4 |
| 2021 | Feature hallucination via Maximum A Posteriori for few-shot learning
Ning Dong 0001, Fan Liu 0003, Sai Yang, Jinglu Hu |
Knowl. Based Syst. | 4 |
| 2019 | Single sample face recognition via BoF using multistage KNN collaborative codingabstractIn this paper, we propose a multistage KNN collaborative coding based Bag-of-Feature (MKCC-BoF) method to address SSPP problem, which tries to weaken the semantic gap between facial features and facial identification. First, local descriptors are extracted from the single training face images and a visual dictionary is obtained offline by clustering a large set of descriptors with K-means. Then, we design a multistage KNN collaborative coding scheme to project local features into the semantic space, which is much more efficient than the most commonly used non-negative sparse coding algorithm in face recognition. To describe the spatial information as well as reduce the feature dimension, the encoded features are then pooled on spatial pyramid cells by max-pooling, which generates a histogram of visual words to represent a face image. Finally, a SVM classifier based on linear kernel is trained with the concatenated features from pooling results. Experimental results on three public face databases show that the proposed MKCC-BoF is much superior to those specially designed methods for SSPP problem. Moreover, it also has great robustness to expression, illumination, occlusion and, time variation. Fan Liu 0003, Sai Yang, Yuhua Ding, Feng Xu 0008 |
Multim. Tools Appl. | 2 |
| 2017 | Robust scene matching method based on sparse representation and iterative correction
Sai Yang, Bo Xiao 0006, Yuanqing Xia, Mengyin Fu, Yang Liu 0038 |
Image Vis. Comput. | 1 |
| 2016 | Local structure based multi-phase collaborative representation for face recognition with single sample per person
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Ye Bi, Sai Yang |
Inf. Sci. | 5 |
| 1988 | Two-dimensional spline interpolation for image reconstruction
Kazuo Toraichi, Sai Yang, Masaru Kamada, Ryoichi Mori |
Pattern Recognit. | 2 |
| 1986 | An automatic analyzing system of left ventricular cineangiogramsabstractA new system for automatically analyzing left ventricular cineangiograms (LVC) is presented. The system processes images of LVC to obtain parameters available for assessing left ventricular functions. Special attention is paid on the left ventricular boundary extraction, since it is of crucial importance to all other processes. The method for reconstructing three-dimensional view of left ventricle is also described. Kazuki Katagishi, Kazuo Toraichi, Ryoichi Mori, Sai Yang |
ICASSP | 4 |