VLDB 2026 Research / reviewers in the wild / expert
Yiran Song
dblp:193/2610
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the gap between LLMs and structured program vulnerability analysis: An agent reasoning approach with first-order logic modeling
Zhongxu Yin, Yiran Song, Liya Kong |
Expert Syst. Appl. | 3 |
| 2026 | Retrieval-augmented in-context learning for multimodal large language models in disease classification
Zaifu Zhan, Shuang Zhou 0012, Xiaoshan Zhou, Yongkang Xiao, Yiran Song, Mingquan Lin, Rui Zhang 0028 |
J. Biomed. Informatics | 9 |
| 2025 | SU-SAM: A Simple Unified Framework for Adapting SAM in Underperformed SceneabstractSegment Anything Model (SAM) excels in common vision tasks but struggles with specialized data. Recent methods fine-tune SAM using parameter-efficient techniques and task-specific designs, but they rely heavily on handcrafting and pre/post-processing, limiting the generalizability. In this paper, we propose SU-SAM, a simple and unified framework that adapts SAM efficiently without task-specific designs, improving its adaptability to underperforming scenes. SU-SAM abstracts parameter-efficient modules into basic design elements, offering four variants: series, parallel, mixed, and LoRA structures. Experiments across nine datasets and six tasks, including medical and defect segmentation, demonstrate SU-SAM’s superior performance. We analyze the effectiveness of different parameter-efficient designs and present a generalized model and benchmark, highlighting SU-SAM’s adaptability across diverse datasets. Yiran Song, Qianyu Zhou 0001, Xuequan Lu, Zhiwen Shao, Lizhuang Ma |
ICME | 1 |
| 2024 | BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything ModelabstractIn this paper, we address the challenge of image resolution variation for the Segment Anything Model (SAM). SAM, known for its zero-shot generalizability, exhibits a performance degradation when faced with datasets with varying image sizes. Previous approaches tend to resize the image to a fixed size or adopt structure modifications, hindering the preservation of SAM's rich prior knowledge. Besides, such task-specific tuning necessitates a complete retraining of the model, which is cost-expensive and unacceptable for deployment in the downstream tasks. In this paper, we reformulate this challenge as a length extrapolation problem, where token sequence length varies while maintaining a consistent patch size for images with different sizes. To this end, we propose a Scalable Bias-Mode Attention Mask (BA-SAM) to enhance SAM's adaptability to varying image resolutions while eliminating the need for structure modifications. Firstly, we introduce a new scaling factor to ensure consistent magnitude in the attention layer's dot product values when the token sequence length changes. Secondly, we present a bias-mode attention mask that allows each token to prioritize neighboring information, mitigating the impact of untrained distant information. Our BA-SAM demonstrates efficacy in two scenarios: zero-shot and finetuning. Extensive evaluation of diverse datasets, including DIS5K, DUTS, ISIC, COD10K, and COCO, reveals its ability to significantly mitigate performance degradation in the zero-shot setting and achieve state-of-the-art performance with minimal fine-tuning. Furthermore, we propose a generalized model and benchmark, showcasing BA-SAM's generalizability across all four datasets simultaneously. Yiran Song, Qianyu Zhou 0001, Xiangtai Li, Deng-Ping Fan, Xuequan Lu, Lizhuang Ma |
CVPR | 1 |
| 2024 | Source-Free Test-Time Adaptation For Online Surface-Defect Detection
Yiran Song, Qianyu Zhou 0001, Lizhuang Ma |
ICPR (9) | 1 |
| 2024 | CFRL: Coarse-Fine Decoupled Representation Learning For Long-Tailed RecognitionabstractData often faces a severe class imbalance issue in the real world, meaning that the number of instances within classes varies greatly, following a long-tailed distribution.In this case, the direct application of supervised learning yields poor performance.Existing long-tailed recognition (LTR) methods often heavily rely on the label information to enhance tail classes' accuracy at the expense of head class by an image-level end-to-end resampling strategy to address data distribution imbalance.Nevertheless, they neglect label bias, which can severely affect the LTR model's accuracy.In this paper, we propose a novel approach, namely Coarse-Fine Decoupled Representation Learning (CFRL) for LTR.Our core idea is to decouple data representations from the classifier and decompose representation learning into two stages: image-level and patch-level.Specifically, in the image-level stage, we leverage unsupervised learning on image-level information to reduce the impact of label bias caused by imbalanced datasets.In the patch-level stage, we introduce patch-level rotation augmentation as negative samples, forcing the model to acquire more comprehensive information.Our theoretical and empirical analyses demonstrate that the approach does not sacrifice the accuracy of head classes while significantly reducing the overfitting of tail classes, improving both of them.We showcase state-of-the-art results on CIFAR, ImageNet, and iNaturalist datasets.Furthermore, we illustrate that this training methodology can be combined with various existing Long-Tailed Recognition (LTR) methods, further enhancing their performance. Yiran Song, Qianyu Zhou 0001, Kun Hu 0008, Lizhuang Ma, Xuequan Lu |
MMAsia | 1 |
| 2024 | Discovering API usage specifications for security detection using two-stage code miningabstractAbstract An application programming interface (API) usage specification, which includes the conditions, calling sequences, and semantic relationships of the API, is important for verifying its correct usage, which is in turn critical for ensuring the security and availability of the target program. However, existing techniques either mine the co-occurring relationships of multiple APIs without considering their semantic relationships, or they use data flow and control flow information to extract semantic beliefs on API pairs but difficult to incorporate when mining specifications for multiple APIs. Hence, we propose an API specification mining approach that efficiently extracts a relatively complete list of the API combinations and semantic relationships between APIs. This approach analyzes a target program in two stages. The first stage uses frequent API set mining based on frequent common API identification and filtration to extract the maximal set of frequent context-sensitive API sequences. In the second stage, the API relationship graph is constructed using three semantic relationships extracted from the symbolic path information, and the specifications containing semantic relationships for multiple APIs are mined. The experimental results on six popular open-source code bases of different scales show that the proposed two-stage approach not only yields better results than existing typical approaches, but also can effectively discover the specifications along with the semantic relationships for multiple APIs. Instance analysis shows that the analysis of security-related API call violations can assist in the cause analysis and patch of software vulnerabilities. Zhongxu Yin, Yiran Song, Guoxiao Zong |
Cybersecur. | 2 |
| 2023 | CLMAE: A Liter and Faster Masked AutoencodersabstractSelf-supervised pre-training has been widely utilized on various vision tasks and gains a great success. However, pre-training on big datasets suffers a lengthy training schedule and large memory consumption. To alleviate these problems, we propose a light-weighted model called Convolutional Lite Masked AutoEncoder (CLMAE). To improve the convergence speed of the transformer during pre-training. We introduce two-stage convolutional progressive patch embedding and an additional convolution in the feed-forward layer, which promote better correlation among patches in the spatial dimensions. The most important design is called cross-layer parameter sharing mechanism, which reduces model parameters with little impact on the performance. We find that sharing parameters among layers not only improves the parameter efficiency, but also acts as a form of regularization that stabilizes the training. Experimental results on downstream tasks show the effectiveness and generalization ability of CLMAE, which accelerates the training process significantly (by 5× for ViT-B and MAE) and reduces a quarter of parameters (by 25M fewer for ViT-B), with a competitive accuracy (82.8% on ImageNet-1K). Yiran Song, Lizhuang Ma |
ICASSP | 1 |
| 2023 | Rethinking Implicit Neural Representations For Vision LearnersabstractImplicit Neural Representations (INRs) are powerful to parameterize continous signals in computer vision. However, almost all INRs methods are limited to low-level tasks, e.g., image/video compression, super-resolution, and image generation. The questions on how to explore INRs to high-level tasks and deep networks are still under-explored. Existing INRs methods suffer from two problems: 1) narrow theoretical definitions of INRs are inapplicable to high-level tasks; 2) lack of representation capabilities to deep networks. Motivated by above facts, we reformulate the definitions of INRs from a novel perspective, and propose an innovative Implicit Neural Representation Network (INRN), which is the first study of INRs to tackle both low-level and high-level tasks. Specifically, we present three key designs for basic blocks in INRN along with two different stacking ways and corresponding loss functions. Extensive experiments with analysis on both low-level task (image fitting) and high-level vision tasks (image classification, object detection, instance segmentation) demonstrate the effectiveness of the proposed method. Yiran Song, Qianyu Zhou 0001, Lizhuang Ma |
ICASSP | 1 |