VLDB 2026 Research / reviewers in the wild / expert
Shuyi Ouyang
dblp:353/9998
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-4507-4153ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language ModelsabstractHallucination in Large Vision-Language Models (LVLMs) remains a critical challenge, undermining their reliability in real-world applications. Existing studies have investigated the causes of hallucination at the modality level and proposed effective strategies. However, interaction patterns beyond the modality level remain insufficiently explored. In this paper, we conduct a token-level analysis and identify two key phenomena: (1) a small subset of textual tokens in LVLMs exert disproportionate influence in the visual-active layers, surpassing that of the visual modality and potentially misleading visual understanding; (2) while LVLMs can correctly identify key visual information, insufficient focus on these cues can sometimes lead to hallucinations. Based on such observation, we attribute hallucinations in LVLMs to two token-level causes: the disproportionate influence of certain textual tokens (phantom tokens) and the underutilization of critical visual cues (anchor tokens). To mitigate these issues, we introduce Token-Asymmetric Filtering (TAF)—a training-free, plug-and-play method that modulates intermediate attention maps in LVLMs. TAF isolates the influence of phantom tokens and emphasizes the influence of anchor tokens in the visual-active layers. Experimental results across multiple benchmarks demonstrate that TAF significantly mitigates hallucinations across a range of state-of-the-art LVLMs. Shuyi Ouyang, Hongyi Wang 0002, Gongfan Fang, Xinyin Ma, Lanfen Lin, Xinchao Wang |
AAAI | 1 |
| 2026 | Lightweight Skin Lesion Images Segmentation Based on CNN and TransformerabstractABSTRACT Accurate segmentation of skin lesion images is essential for diagnosing and treating skin disorders. While current research primarily aims to enhance segmentation accuracy through the use of complex network models, the large size of these models restricts their practical application in clinical settings. To address this challenge, we propose SMedt (skin‐medical transformer), a high‐precision, parameter‐efficient model for skin lesion image segmentation. SMedt combines CNN and transformer architectures within a dual‐branch structure to extract both global and local features effectively. The model's global branch employs a dual‐attention mechanism that integrates channel and spatial attention, along with skip connection cross attention (SCCA) between the encoder and decoder layers, to enhance global feature decoding. The local branch incorporates an All‐aggregation decoder (all decoder) method, enabling the capture of multi‐scale features, while pyramid stacking improves the extraction of local features across different channel dimensions. We evaluated SMedt on the ISIC2016 and ISIC2018 datasets. On the ISIC2016 dataset, the model achieved a 0.17% improvement in the Dice coefficient over the second‐best model, FAT‐Net, while reducing model parameters by 91%. On the ISIC2018 dataset, SMedt improved the Dice coefficient by 1.1% over the state‐of‐the‐art Efficient UNet, while reducing model parameters by 73.26%. Compared with the latest lightweight model UCM‐Net, SMedt improves segmentation accuracy (Dice) by 1.92% on ISIC2018, achieving a better balance between accuracy and model size. By leveraging the strengths of both CNN and Transformer models, SMedt maintains high segmentation accuracy while keeping a low parameter count, thereby enhancing diagnostic accuracy and clinical efficiency. Tuoyu Ouyang, Huimin Quan, Guocai Liu, Shuyi Ouyang |
Concurr. Comput. Pract. Exp. | 4 |
| 2026 | S2Match: Revisiting Weak-to-Strong Consistency From a Semantic Similarity Perspective for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) for medical image segmentation is a challenging yet highly practical task, which reduces reliance on large-scale labeled datasets by leveraging unlabeled samples. Among SSL techniques, the weak-to-strong consistency framework, popularized by FixMatch, has emerged as a state-of-the-art method in classification tasks. Notably, such a simple pipeline has also shown competitive performance in medical image segmentation. However, two key limitations still persist, impeding its efficient adaptation: (1) the neglect of contextual dependencies results in inconsistent predictions for similar semantic features, leading to incomplete object segmentation; (2) the lack of exploitation on semantic similarity between labeled and unlabeled data induces considerable class-distribution discrepancy. To address these limitations, we propose a novel SSL framework for medical image segmentation, named S2Match, powered by two appealing designs from a semantic similarity perspective: (1) rectifying pixel-wise prediction by reasoning about the intra-image pair-wise affinity map, thus integrating contextual dependencies explicitly into the final prediction; (2) bridging labeled and unlabeled data via a feature querying mechanism for compact class representation learning, which fully considers cross-image anatomical similarities. As the reliable semantic similarity extraction depends on robust features, we further introduce an effective Spatial-aware Fusion Module (SFM) to explore distinctive information from multiple scales. Experiments show that S2Match yields consistent improvements over the state-of-the-art methods across five public medical image segmentation benchmarks, exhibiting competitive performance on both 2D and 3D tasks. Shiao Xie, Hongyi Wang 0002, Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology ImagesabstractThe advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the ST expressions from digital pathology images is desired. Current methods usually adopt existing regression backbones along with patch-sampling for this task, which ignores the inherent multi-scale information embedded in the pyramidal data structure of digital pathology images, and wastes the inter-spot visual information crucial for accurate gene expression prediction. To address these limitations, we propose M2OST, a many-to-one regression Transformer that can accommodate the hierarchical structure of the pathology images via a decoupled multi-scale feature extractor. Unlike traditional models that are trained with one-to-one image-label pairs, M2OST uses multiple images from different levels of the digital pathology image to jointly predict the gene expressions in their common corresponding spot. Built upon our many-to-one scheme, M2OST can be easily scaled to fit different numbers of inputs, and its network structure inherently incorporates nearby inter-spot features, enhancing regression performance. We have tested M2OST on three public ST datasets and the experimental results show that M2OST can achieve state-of-the-art performance with fewer parameters and floating-point operations (FLOPs). Hongyi Wang 0002, Xiuju Du, Jing Liu 0041, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
AAAI | 4 |
| 2025 | Region-Aware Anchoring Mechanism for Efficient Referring Visual Grounding
Shuyi Ouyang, Ziwei Niu, Hongyi Wang 0002, Yen-Wei Chen 0001, Lanfen Lin |
ICCV | 1 |
| 2024 | IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image SegmentationabstractSingle-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performance and complex inference. In response to these challenges, we present a novel single domaininvariant representation learning approach for medical image segmentation, called IRLSG, with two appealing designs: (1) A Classscale Photo-metric Augmentation is first proposed to simulate unseen target domain that is sufficient in diversity and informativeness. After that, a Dual-Consistency Framework is further designed to constrain the consistency of intermediate features and segmentation results between the original and the augmented images, which helps to explore the domain-invariant representation. (2) A simple and effective Style Feature Whitening is designed to decouple and remove the domain-specific style from higher-order covariance statistics, which can further improve the modeling and generalization capability of the network. Experimental results on different benchmarks demonstrate that our IRLSG outperforms the current state-of-the-art methods in tackling single-domain generalization. Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Shiao Xie, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin |
ICASSP | 3 |
| 2023 | MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And GeneralizationabstractConventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications. A paradigm for tackling this problem is to train multiple local models and aggregate a generalized central model without data sharing. Recent methods have made remarkable advancements in this paradigm by exploiting parameter alignment and aggregation. But when sources domain variety increases, directly aligning and aggregating local parameters becomes more challenging. Adapting a different approach in this work, we devised a data-free semantic collaborative distillation strategy to learn domain-invariant representation for both federated UDA and DG. Each local model transmits its predictions to the central server and derives its target distribution from the average of other local models' distributions to facilitate the mutual transfer of domain-specific knowledge. When unlabeled target data is available, we introduce a novel UDA strategy termed knowledge filter to adapt the central model to the target data. Extensive experiments on four UDA and DG datasets demonstrate that our method has a competitive performance compared with the state-of-the-art methods. Ziwei Niu, Hongyi Wang 0002, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
ICASSP | 4 |
| 2023 | SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image SegmentationabstractReferring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques have been proposed to efficiently leverage long-range cross-modal dependencies, enhancing performance for referring segmentation. However, existing methods consider visual feature extraction and cross-modal fusion separately, resulting in insufficient visual-linguistic alignment in semantic space. In addition, they employ sequential structures and hence lack multi-scale information interaction. To address these limitations, we propose a Scale-Wise Language-Guided Vision Transformer (SLViT) with two appealing designs: (1) Language-Guided Multi-Scale Fusion Attention, a novel attention mechanism module for extracting rich local visual information and modeling global visual-linguistic relationships in an integrated manner. (2) An Uncertain Region Cross-Scale Enhancement module that can identify regions of high uncertainty using linguistic features and refine them via aggregated multi-scale features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that SLViT surpasses state-of-the-art methods with lower computational cost. The code is publicly available at: https://github.com/NaturalKnight/SLViT. Shuyi Ouyang, Hongyi Wang 0002, Shiao Xie, Ziwei Niu, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin |
IJCAI | 1 |
| 2023 | HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationabstractThe task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and essential visual features presented in the image, tight visual-linguistic interactions play a vital role in improving classification performance. Moreover, given the potential variance in object size and appearance within a single image, attention to features of different scales can help to discover possible objects in the image. Recently, Transformer-based methods have achieved great success in multi-label image classification by leveraging the advantage of modeling long-range dependencies, but they have several limitations. Firstly, existing methods treat visual feature extraction and cross-modal fusion as separate steps, resulting in insufficient visual-linguistic alignment in the joint semantic space. Additionally, they only extract visual features and perform cross-modal fusion at a single scale, neglecting objects with different characteristics. To address these issues, we propose a Hierarchical Scale-Aware Vision-Language Transformer (HSVLT) with two appealing designs: (1)A hierarchical multi-scale architecture that involves a Cross-Scale Aggregation module, which leverages joint multi-modal features extracted from multiple scales to recognize objects of varying sizes and appearances in images. (2)Interactive Visual-Linguistic Attention, a novel attention mechanism module that tightly integrates cross-modal interaction, enabling the joint updating of visual, linguistic and multi-modal features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that HSVLT surpasses state-of-the-art methods with lower computational cost. Shuyi Ouyang, Hongyi Wang 0002, Ziwei Niu, Zhenjia Bai, Shiao Xie, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 1 |