Yin Lin

dblp:19/3936 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 4Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking the I/O Bottleneck: I/O Coordination Optimization for Efficient Large-Scale LLM Fine-Tuning
abstract
Large Language Models (LLMs) with tens or even hundreds of billions of parameters have become the foundation of modern AI applications. However, fine-tuning such massive models is severely constrained by the limited GPU memory. Existing memory-saving systems, such as ZeRO-based offloading in DeepSpeed, reduce GPU memory usage but inevitably incur substantial I/O overhead, especially when model states reside on slow storage devices, such as NVMe SSDs. As a result, the memory bottleneck in large-scale fine-tuning is transformed into an I/O bottleneck. Although prior systems have employed strategies like parameter prefetching and partial asynchronous execution, they remain limited by synchronous I/O-communication dependencies and the lack of fine-grained read/write I/O scheduling. To address these limitations, we propose IOC, an I/O Coordination Optimization framework that maximizes pipeline parallelism across different phases of LLM fine-tuning. IOC introduces three key mechanisms: (1) An All-Gather prefetching technique based on an I/O state hash table, which completely decouples All-Gather prefetching from parameter I/O, achieving continuous overlap among I/O, communication, and computation; (2) The parameter update phase is refactored into an asynchronous pipeline with explicit I/O isolation, where the optimizer state write-back is executed in a semi-asynchronous manner, thereby mitigating read/write contention and reducing synchronization stalls; (3) Multi-disk parallelism is leveraged by introducing an additional disk to further relieve I/O contention and defer synchronization waits to the latest possible time point. Experimental results demonstrate that IOC significantly accelerates LLM finetuning while preserving low memory consumption. The end-toend fine-tuning time on the Llama-70B model is reduced by $\mathbf{2 1. 5 \%}$ and 34.3% in single-disk and multi-disk configurations compared to the baseline.
Ziyang Shen, Hongchao Du, Kaihuan Lin, Yin Lin, Qiao Li 0001, Chun Jason Xue
ISPASS5
2026 Lightweight ensemble vision transformer framework for non-invasive survival prediction in glioblastoma
abstract
Glioblastoma is the most aggressive primary brain tumor, characterized by rapid progression and poor prognosis. Predicting overall survival (OS) at diagnosis can support clinicians in tailoring treatments to patient-specific risk levels. Traditional approaches often rely on tumor segmentation or pre-trained models based on natural images, both of which present limitations in scalability and clinical practicality. In this work, we introduce GLiT-Net, a lightweight ensemble framework based on Vision Transformers (ViTs), designed for OS classification using only MRI scans and patient age, without the need for segmentation or external pre-training. To address data scarcity, GLiT-Net operates on reduced-dimensional MRI volumes and leverages targeted data augmentation strategies to enhance generalizability. We evaluate the model on the BraTS 2020 dataset, selecting 118 glioblastoma patients with gross total resection. Survival times are discretized into three balanced classes using percentile thresholds. Our ensemble framework integrates predictions from ten compact ViT models, trained under a repeated nested cross-validation scheme. GLiT-Net achieves an average test accuracy of 70.3% and an F1-score of 68.7%, outperforming existing approaches in the literature which report accuracies up to 65%. The model maintains predictive strength despite reduced image resolution and limited data availability. These findings underscore the feasibility of segmentation-free, non-invasive survival prediction using simplified ViT-based architectures. While external validation and integration of additional clinical or molecular features remain future directions, GLiT-Net offers a practical and effective solution for automated risk stratification in glioblastoma care. • Accurate prediction of overall survival in glioblastoma is essential for treatment planning, yet current approaches rely heavily on tumor segmentation to achieve reliable performance. • A compact ensemble of vision transformers predicts survival directly from reduced volumetric MRI without requiring tumor segmentation and remains stable across varying contrast conditions, matching the reliability of more complex state-of-the-art pipelines. • This framework delivers immediate prognostic insight following MRI acquisition, enabling faster, contrast-invariant, segmentation-free support for personalized treatment planning and more efficient patient management.
Yin Lin, Domenico Aquino, Giuseppe Lauria, Marina Grisoli, Alberto Redaelli, Riccardo Barbieri, Simona Ferrante
Neurocomputing1
2026 Mixgaze: a dually supervised mixed attention network for gaze estimation
Ziyang Wu, Yin Lin, Hu Cheng, Caihua Kong, Wengang Zhou 0001, Houqiang Li
Multim. Syst.2
2026 SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001
Pattern Recognit.3
2026 Exploring a novel data- and parameter-efficient fine-tuning method for robust face recognition
Yin Lin, Qidong Huang, Yang Cao 0010, Zengfu Wang
Pattern Recognit.1
2025 Cross-modulated Attention Transformer for RGBT Tracking
abstract
Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and search-template correlation. Nevertheless, the independent search-template correlation calculations are prone to be affected by low-quality data, which might result in contradictory and ambiguous correlation weights. It not only limits the intra-modal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Cross-modulated Attention Transformer (CAFormer), which innovatively integrates inter-modality interaction into the search-template correlation computation within typical attention mechanism, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed correlation modulated enhancement module, which can modify inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we design a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against state-of-the-art methods.
Yun Xiao 0003, Jiacong Zhao, Andong Lu, Chenglong Li 0002, Yin Lin, Cong Liu 0006
AAAI6
2025 Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging Scenarios
abstract
Face recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale training data, effectively transferring pre-trained face recognition models to these scenarios has become a challenge. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have shown great potential in natural language processing tasks, but their effectiveness in computer vision tasks, especially in face recognition tasks, remains under-explored. In this paper, we propose a Data-Parameter-Efficient Fine-Tuning (DPEFT) approach for the face recognition tasks, encompassing two kinds of fine-tuning strategies. With these strategies, the DPEFT method requires only an additional 2.7% learnable parameters and 20% of the training data during the training phase to achieve competitive results. Moreover, by further integrating the concept of structural re-parameterization, our approach maintains the same model architecture and parameters as the pre-trained model during inference. Extensive experimental results on both holistic and occluded face datasets demonstrate that our approach achieves performance comparable to or better than the fully fine-tuning methods, and significantly lower training costs. Our DPEFT enables the pre-trained face recognition model to adapt efficiently and effectively to a variety of scenarios, indicating its potential in practical applications.
Yin Lin, Ziyang Wu, Qidong Huang, Jinshui Hu, Zengfu Wang
ICASSP1
2025 Exploring Part-Informed Visual-Language Learning for Person Re-Identification
abstract
Recently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while neglecting supervision on fine-grained part features, thus lacking constraints for local feature semantic consistency. To this end, we propose Part-Informed Visual-language Learning (π-VL) to enhance fine-grained visual features with part-informed language supervisions for ReID tasks. Specifically, π-VL introduces a human parsing-guided prompt tuning strategy and a hierarchical visual-language alignment paradigm to ensure within-part feature semantic consistency. The former combines both identity labels and human parsing maps to constitute pixel-level text prompts, and the latter fuses multi-scale visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our π-VL achieves performance comparable to or better than state-of-the-art methods on four commonly used ReID benchmarks. Notably, it reports 91.0% Rank-1 and 76.9% mAP on the challenging MSMT17 database, without bells and whistles.
Yin Lin, Yehansen Chen, Jinshui Hu, Cong Liu 0006, Zengfu Wang
ICME1
2025 Multi-scale count-task guided feature enhancement face detection
Ziyang Wu, Yin Lin, Qidong Huang, Wengang Zhou 0001, Houqiang Li
Multim. Syst.2
2024 SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions
Yin Lin, Bolin Ding, H. V. Jagadish, Jingren Zhou 0001
CIDR1
2024 Mitigating Subgroup Unfairness in Machine Learning Classifiers: A Data-Driven Approach
abstract
Fairness in machine learning, particularly in classifiers, is receiving increasing attention. However, most studies on this topic focus on fairness metrics for a limited number of predefined groups and do not address fairness across intersectional subgroups. In this paper, we investigate ways to improve subgroup fairness where subgroups are defined by the intersection of protected attributes. Specifically, our paper reveals the correlation between the representation bias of training data and model fairness. We demonstrate that biased sample collection due to historical biases and a lack of control over data collection can lead to unfairness in learned models. We introduce the concept of an “Implicit Biased Set (IBS)”, which refers to regions in the intersectional attribute space where positive and negative examples are not proportionately represented. For example, if our training data set has a disproportionate representation of black male recidivists, then criminal risk assessment tools are more likely to discriminate against black males, even if they are innocent. We propose an efficient pre-processing approach that initially identifies IBS and then employs techniques to remedy the data collection within IBS. Our evaluation shows that our method effectively mitigates various subgroup biases regardless of the downstream machine learning models used.
Yin Lin, Samika Gupta, H. V. Jagadish
ICDE1
2024 Attention-Guided Contrastive Masked Autoencoders for Self-supervised Cross-Modal Biometric Matching
Jiaxiang Wang 0001, Hanqin Shi, Zhenda Yu, Yin Lin
ICDF2C (2)5
2023 Predicate Pushdown for Data Science Pipelines
abstract
Predicate pushdown is a widely adopted query optimization. Existing systems and prior work mostly use pattern-matching rules to decide when a predicate can be pushed through certain operators like join or groupby. However, challenges arise in optimizing for data science pipelines due to the widely used non-relational operators and user-defined functions (UDF) that existing rules would fail to cover. In this paper, we present MagicPush, which decides predicate pushdown using a search-verification approach.MagicPush searches for candidate predicates on pipeline input, which is often not the same as the predicate to be pushed down, and verifies that the pushdown does not change pipeline output with full correctness guarantees. Our evaluation on TPC-H queries and 200 real-world pipelines sampled from GitHub Notebooks shows that MagicPush substantially outperforms a strong baseline that uses a union of rules from prior work - it is able to discover new pushdown opportunities and better optimize 42 real-world pipelines with up to 99% reduction in running time, while discovering all pushdown opportunities found by the existing baseline on remaining cases.
Cong Yan, Yin Lin, Yeye He
Proc. ACM Manag. Data2
2022 OREO: Detection of Cherry-picked Generalizations
abstract
Data analytics often make sense of large data sets by generalization: aggregating from the detailed data to a more general context. Given a dataset, misleading generalizations can sometimes be drawn from a cherry-picked level of aggregation to obscure substantial subgroups that oppose the generalization. Our goal is to detect and explain cherry-picked generalizations by refining the corresponding aggregate queries. We demonstrate OREO, a system to compute a support score of the given statement to quantify the quality of the generalization; that is, whether the aggregated result is an accurate reflection of the data. To better understand the resulting score, our system also identifies significant counterexamples and alternative statements that better represent the data at hand. We will demonstrate the utility of OREO for investigating generalizations, by interacting with the VLDB'22 participants who will use the OREO interface for statement validation and explanation.
Yin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish, Tova Milo
Proc. VLDB Endow.1
2021 On Detecting Cherry-picked Generalizations
abstract
Generalizing from detailed data to statements in a broader context is often critical for users to make sense of large data sets. Correspondingly, poorly constructed generalizations might convey misleading information even if the statements are technically supported by the data. For example, a cherry-picked level of aggregation could obscure substantial sub-groups that oppose the generalization. We present a framework for detecting and explaining cherry-picked generalizations by refining aggregate queries. We present a scoring method to indicate the appropriateness of the generalizations. We design efficient algorithms for score computation. For providing a better understanding of the resulting score, we also formulate practical explanation tasks to disclose significant counterexamples and provide better alternatives to the statement. We conduct experiments using real-world data sets and examples to show the effectiveness of our proposed evaluation metric and the efficiency of our algorithmic framework.
Yin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish, Tova Milo
Proc. VLDB Endow.1
2020 Identifying Insufficient Data Coverage in Databases with Multiple Relations
Yin Lin, Abolfazl Asudeh, H. V. Jagadish
Proc. VLDB Endow.1
2018 R^2 -Tree: An Efficient Indexing Scheme for Server-Centric Data Center Networks
Yin Lin, Xinyi Chen 0004, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen
DEXA (1)1
2013 Peer-assisted content distribution in Akamai netsession
abstract
Content distribution systems have traditionally adopted one of two architectures: infrastructure-based content delivery networks (CDNs), in which clients download content from dedicated, centrally managed servers, and peer-to-peer CDNs, in which clients download content from each other. The advantages and disadvantages of each architecture have been studied in great detail. Recently, hybrid, or 'peer-assisted', CDNs have emerged, which combine elements from both architectures. The properties of such systems, however, are not as well understood.
Mingchen Zhao, Paarijaat Aditya, Ang Chen 0001, Yin Lin, Andreas Haeberlen, Peter Druschel, Bruce M. Maggs, Bill Wishon, Miroslav Ponec
Internet Measurement Conference4
2013 DataSpotting: Exploiting naturally clustered mobile devices to offload cellular traffic
abstract
The proliferation of pictures and videos in the Internet is imposing heavy demands on mobile data networks. Though emerging wireless technologies will provide more bandwidth, the increase in demand will easily consume the additional capacity. To alleviate this problem, we explore the possibility of serving user requests from other mobile devices located geographically close to the user. For instance, when Alice reaches areas with high device density - Data Spots - the cellular operator learns Alice's content request, and guides her device to nearby devices that have the requested content. Importantly, communication between the nearby devices can be mediated by servers, avoiding many of the known problems of pure ad hoc communication. This paper argues this viability through systematic prototyping, measurements, and measurement-driven analysis.
Xuan Bao, Yin Lin, Uichin Lee, Ivica Rimac, Romit Roy Choudhury
INFOCOM2
2013 Less pain, most of the gain: incrementally deployable ICN
abstract
Information-Centric Networking (ICN) has seen a significant resurgence in recent years. ICN promises benefits to users and service providers along several dimensions (e.g., performance, security, and mobility). These benefits, however, come at a non-trivial cost as many ICN proposals envision adding significant complexity to the network by having routers serve as content caches and support nearest-replica routing. This paper is driven by the simple question of whether this additional complexity is justified and if we can achieve these benefits in an incrementally deployable fashion. To this end, we use trace-driven simulations to analyze the quantitative benefits attributed to ICN (e.g., lower latency and congestion). Somewhat surprisingly, we find that pervasive caching and nearest-replica routing are not fundamentally necessary---most of the performance benefits can be achieved with simpler caching architectures. We also discuss how the qualitative benefits of ICN (e.g., security, mobility) can be achieved without any changes to the network. Building on these insights, we present a proof-of-concept design of an incrementally deployable ICN architecture.
Seyed Kaveh Fayaz, Yin Lin, Amin Tootoonchian, Ali Ghodsi 0002, Teemu Koponen, Bruce M. Maggs, K. C. Ng, Vyas Sekar, Scott Shenker
SIGCOMM2
2012 Reliable Client Accounting for P2P-Infrastructure Hybrids
Paarijaat Aditya, Mingchen Zhao, Yin Lin, Andreas Haeberlen, Peter Druschel, Bruce M. Maggs, Bill Wishon
NSDI3