VLDB 2026 Research / reviewers in the wild / expert
Xiaotian Lin
dblp:225/7491
· DBLP profile ↗
18ranked-venue papers
2as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
Xiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas, Chengliang Chai, Nan Tang 0001, Yuyu Luo |
Proc. VLDB Endow. | 1 |
| 2025 | A Robotic Micromanipulation System for Homogeneous Organoid CultureabstractOrganoids are cell clusters cultured in vitro that maintain the structure and function of the donor organs. They have found important applications in biomedicine, such as drug screening and personalized therapy. However, conventional organoid culture methods lack control of physical properties like size and distribution, leading to increased heterogeneity and very low batch-to-batch reproducibility, which significantly limits their widespread use. Controlling these properties at the microscale is challenging, particularly for fragile fragments, which are the main source for culturing organoids. To address this issue, we present a robotic micromanipulation system that allows operators to select fragments of particular sizes and automatically transfer them into a customized in-situ organoid chip (IOC) for culture. The chip was designed with microwell arrays to uniform the culture environment and facilitate imaging analysis. The transfer of fragments is modeled based on computational fluid dynamics (CFD) and is enabled by designing a robust model predictive control (RMPC) framework. Simulation and experiment results demonstrated the effectiveness of the model and controller. In colorectal cancer organoid culture experiments, our system significantly improved the morphological homogeneity of organoids. Note to Practitioners—Organoids have been demonstrated to be one of the most promising in vitro models. Lacking control of its size and distribution results in significant heterogeneity and low batch-to-batch reproducibility, which limits its wide uses. Here, we report a robotic micromanipulation system that allows operators to select fragments of particular sizes and morphologies and automatically transfer them into a customized organoid chip for culture. The results of colorectal cancer organoids culture experiments verified the effectiveness of our system in reducing the morphological heterogeneity among organoids. Xiaotian Lin, Xinghu Yu, Qiong Mo, Mingsi Tong, Songlin Zhuang, Huijun Gao |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | A Model Predictive Control Approach of Optimal Autonomous Laboratory ManagementabstractThe development of autonomous laboratories has significantly advanced with the integration of computer vision, simultaneous localization and mapping, cloud computation technologies, etc. These advancements have enhanced the automation and efficiency of experimental processes. However, the optimal management of complex task scheduling within such environments remains underexplored, especially in the face of challenges such as managing numerous tasks, adhering to strict time and state constraints, and ensuring the sustainable stability and performance of the entire laboratory operation. This article introduces a novel model predictive control (MPC)-based strategy for the optimal management of task scheduling in autonomous laboratories. Our approach begins with the abstraction of the scheduling problem as a finite state machine, which lays the foundation for a systematic analysis. We then employ concepts of invariant sets and stability to ensure that the pro posed scheduling strategy is not only efficient but also resilient to operational uncertainties. The proposed approach ensures recursive feasibility, which guarantees the adaptability of the scheduling strategy over time. Through a series of simulations, we demonstrate the efficacy of our MPC-based management strategy in optimizing task scheduling, thereby significantly enhancing the laboratory's operational efficiency, stability, and sustainability. Our findings offer promising insights into the future of autonomous laboratory management, providing a robust framework for tackling the complexities of task scheduling in such environments. Xiaotian Lin, Juan J. Rodríguez-Andina, Zhengkai Li |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A Chinese Grammatical Error Correction Model Based On Grammatical Generalization And Parameter SharingabstractAbstract Chinese grammatical error correction (CGEC) is a significant challenge in Chinese natural language processing. Deep-learning-based models tend to have tens of millions or even hundreds of millions of parameters since they model the target task as a sequence-to-sequence problem. This may require a vast quantity of annotated corpora for training and parameter tuning. However, there are currently few open-source annotated corpora for the CGEC task; the existing researches mainly concentrate on using data augmentation technology to alleviate the data-hungry problem. In this paper, rather than expanding training data, we propose a competitive CGEC model from a new insight for reducing model parameters. The model contains three main components: a sequence learning module, a grammatical generalization module and a parameter sharing module. Experimental results on two Chinese benchmarks demonstrate that the proposed model could achieve competitive performance over several baselines. Even if the parameter number of our model is reduced by 1/3, it could reach a comparable $F_{0.5}$ value of 30.75%. Furthermore, we utilize English datasets to evaluate the generalization and scalability of the proposed model. This could provide a new feasible research direction for CGEC research. Nankai Lin, Xiaotian Lin, Yingwen Fu, Shengyi Jiang, Lianxi Wang 0001 |
Comput. J. | 2 |
| 2023 | An efficient framework for few-shot skeleton-based temporal action segmentation
Leiyang Xu, Qiang Wang 0001, Xiaotian Lin |
Comput. Vis. Image Underst. | 3 |
| 2023 | Skeleton-based Tai Chi action segmentation using trajectory primitives and content
Leiyang Xu, Qiang Wang 0001, Xiaotian Lin |
Neural Comput. Appl. | 3 |
| 2023 | CL-XABSA: Contrastive Learning for Cross-Lingual Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA), an extensively researched area in the field of natural language processing (NLP), predicts the sentiment expressed in a text relative to the corresponding aspect. Unfortunately, most languages lack sufficient annotation resources; thus, an increasing number of recent researchers have focused on cross-lingual aspect-based sentiment analysis (XABSA). However, most recent studies focus only on cross-lingual data alignment instead of model alignment. Therefore, we propose a novel framework, CL-XABSA: contrastive learning for cross-lingual aspect-based sentiment analysis. Based on contrastive learning, we close the distance between samples with the same label in different semantic spaces, achieving convergence of semantic spaces of different languages. Specifically, we design two contrastive objectives, token-level contrastive learning of token embeddings (TL-CTE) and sentiment-level contrastive learning of token embeddings (SL-CTE), to unify the semantic space of source and target languages. Since CL-XABSA can receive datasets in multiple languages during training, it can be further extended to multilingual aspect-based sentiment analysis (MABSA). To further improve the model performance, we perform knowledge distillation with target-language unlabeled data. In the distillation XABSA task, we further explore the effectiveness of different data (source dataset, translated dataset, and code-switched dataset). The results demonstrate that the proposed method has a certain improvement in the three XABSA tasks, distillation XABSA and MABSA. The source code of this paper is publicly available athttps://github.com/GKLMIP/CL-XABSA. Nankai Lin, Yingwen Fu, Xiaotian Lin, Dong Zhou 0001, Aimin Yang 0002, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Automatic Dataset Generation for Specific Object DetectionabstractIn the past decade, object detection tasks are defined mostly by large public datasets. However, building object detection datasets is not scalable due to inefficient image collecting and labeling. Furthermore, most labels are still in the form of bounding boxes, which provide much less information than the real human visual system. In this paper, we present a method to synthesize object-in-scene images, which can preserve the objects' detailed features without bringing irrelevant information. In brief, given a set of images containing a target object, our algorithm first trains a model to find an approximate center of the object as an anchor, then makes an outline regression to estimate its boundary, and finally blends the object into a new scene. Our result shows that in the synthesized image, the boundaries of objects blend very well with the background. Experiments also show that SOTA segmentation models work well with our synthesized data. Xiaotian Lin, Leiyang Xu, Qiang Wang 0001 |
ICIP | 1 |
| 2022 | Temporal-spatial Feature Fusion for Few-shot Skeleton-based Action RecognitionabstractRecognizing new action categories from a few reference samples is an encouraging research field because the cost of labeling data is expensive. This work presents a method for few-shot (or one-shot) skeleton-based action recognition by fusing temporal and spatial features of actions. Trajectory primitives are proposed to characterize the temporal features, which can be obtained by segmenting and clustering the trajectories of joints. After that, we modify the original dynamic time warping (DTW) algorithm and use it to measure the similarity between trajectory primitive sequences. Besides, we compute the joint angles as spatial feature vectors. Support vector machines (SVM) are used to classify the joint angle vectors. In this way, the temporal distance matrix can be calculated by modified DTW, and the spatial distance matrix can be obtained by trained SVM. Finally, we fuse temporal and spatial distance matrices by adjusting a parameter to improve recognition accuracy. Furthermore, extensive experiments are conducted on three small-scale datasets to verify the effectiveness of our proposed method. Leiyang Xu, Qiang Wang 0001, Xiaotian Lin |
IECON | 3 |
| 2022 | A Fine-Grained Social Bias Measurement Framework for Open-Domain Dialogue Systems
Aimin Yang 0002, Qifeng Bai, Jigang Wang, Nankai Lin, Xiaotian Lin, Guanqiu Qin, Junheng He |
NLPCC (2) | 5 |
| 2022 | Multi-label emotion classification based on adversarial multi-task learning
Nankai Lin, Sihui Fu, Xiaotian Lin, Lianxi Wang 0001 |
Inf. Process. Manag. | 3 |
| 2022 | Batch covariance neural network for image recognition
Tianyou Zheng, Qiang Wang 0001, Xiaotian Lin |
Image Vis. Comput. | 5 |
| 2022 | Gradient rectified parameter unit of the fully connected layer in convolutional neural networks
Tianyou Zheng, Qiang Wang 0001, Xiaotian Lin |
Knowl. Based Syst. | 4 |
| 2022 | High-resolution rectified gradient-based visual explanations for weakly supervised segmentation
Tianyou Zheng, Qiang Wang 0001, Xiaotian Lin |
Pattern Recognit. | 5 |
| 2021 | Research on Pseudo-label Technology for Multi-label News Classification
Lianxi Wang 0001, Xiaotian Lin, Nankai Lin |
ICDAR (2) | 2 |
| 2021 | Pre-trained Language Models for Tagalog with Multi-source Data
Shengyi Jiang, Yingwen Fu, Xiaotian Lin, Nankai Lin |
NLPCC (1) | 3 |
| 2021 | A Framework for Indonesian Grammar Error CorrectionabstractGrammatical Error Correction (GEC) is a challenge in Natural Language Processing research. Although many researchers have been focusing on GEC in universal languages such as English or Chinese, few studies focus on Indonesian, which is a low-resource language. In this article, we proposed a GEC framework that has the potential to be a baseline method for Indonesian GEC tasks. This framework treats GEC as a multi-classification task. It integrates different language embedding models and deep learning models to correct 10 types of Part of Speech (POS) error in Indonesian text. In addition, we constructed an Indonesian corpus that can be utilized as an evaluation dataset for Indonesian GEC research. Our framework was evaluated on this dataset. Results showed that the Long Short-Term Memory model based on word-embedding achieved the best performance. Its overall macro-average F 0.5 in correcting 10 POS error types reached 0.551. Results also showed that the framework can be trained on a low-resource dataset. Nankai Lin, Xiaotian Lin, Kanoksak Wattanachote, Shengyi Jiang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Multi-domain Sentiment Classification on Self-constructed Indonesian Dataset
Nankai Lin, Sihui Fu, Xiaotian Lin, Shengyi Jiang |
NLPCC (1) | 4 |