VLDB 2026 Research / reviewers in the wild / expert
Anqi Lin
dblp:237/0948
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A bi-directional flow weighted regression for interpreting social media sentiment identified by large language modelsabstractSocial networks, combined with location-based services, offer valuable opportunities to examine social media sentiment and interactions across regions. Information flow within social networks is often highly directional and intense, transcending geographic distances. As a result, Geographically Weighted Regression (GWR), a traditional model that uses geographic distance to measure spatial proximity, falls short in explaining the factors influencing social media sentiment. To address this limitation, this study proposed a bi-directional flow weighted regression (BDFWR) model, supported by large language models, to interpret influencing factors of social media sentiment. The results demonstrated that the BDFWR model outperformed the GWR model by effectively capturing the relationship between social media sentiment and socioeconomic factors. This approach revealed deeper insights into the spatial heterogeneity of social media sentiment across diverse regions, enhancing the accuracy of modelling social media sentiment distribution. Incorporating bi-directional flow distance significantly improved the model’s performance, particularly in cases involving ‘closely low interflow’ and ‘remotely high interflow’ phenomena—critical aspects often neglected in conventional geographic models. Moreover, large language models excelled in detecting implicit positive and negative trends within textual data, offering a promising avenue for advancing sentiment analysis research. Anqi Lin, Hengyuan Liu, Hao Wu 0004 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | Bridging artificial intelligence and biological sciences: a comprehensive review of large language models in bioinformaticsabstractLarge language models (LLMs), representing a breakthrough advancement in artificial intelligence, have demonstrated substantial application value and development potential in bioinformatics research, particularly showing significant progress in the processing and analysis of complex biological data. This comprehensive review systematically examines the development and applications of LLMs in bioinformatics, with particular emphasis on their advancements in protein and nucleic acid structure prediction, omics analysis, drug design and screening, and biomedical literature mining. This work highlights the distinctive capabilities of LLMs in end-to-end learning and knowledge transfer paradigms. Additionally, this paper thoroughly discusses the major challenges confronting LLMs in current applications, including key issues such as model interpretability and data bias. Furthermore, this review comprehensively explores the potential of LLMs in cross-modal learning and interdisciplinary development. In conclusion, this paper aims to systematically summarize the current research status of LLMs in bioinformatics, objectively evaluate their advantages and limitations, and provide insights and recommendations for future research directions, thereby positioning LLMs as essential tools in bioinformatics research and fostering innovative developments in the biomedical field. Anqi Lin, Junpu Ye, Chang Qi, Lingxuan Zhu, Weiming Mou, Wenyi Gan, Dongqiang Zeng, Bufu Tang, Mingjia Xiao, Guangdi Chu, Shengkun Peng, Hank Z. H. Wong, Lin Zhang 0058, Hengguo Zhang, Xinpei Deng, Kailai Li 0003, Jian Zhang 0104, Aimin Jiang, Zhengrui Li, Peng Luo 0005 |
Briefings Bioinform. | 1 |
| 2024 | CPADS: a web tool for comprehensive pancancer analysis of drug sensitivityabstractDrug therapy is vital in cancer treatment. Accurate analysis of drug sensitivity for specific cancers can guide healthcare professionals in prescribing drugs, leading to improved patient survival and quality of life. However, there is a lack of web-based tools that offer comprehensive visualization and analysis of pancancer drug sensitivity. We gathered cancer drug sensitivity data from publicly available databases (GEO, TCGA and GDSC) and developed a web tool called Comprehensive Pancancer Analysis of Drug Sensitivity (CPADS) using Shiny. CPADS currently includes transcriptomic data from over 29 000 samples, encompassing 44 types of cancer, 288 drugs and more than 9000 gene perturbations. It allows easy execution of various analyses related to cancer drug sensitivity. With its large sample size and diverse drug range, CPADS offers a range of analysis methods, such as differential gene expression, gene correlation, pathway analysis, drug analysis and gene perturbation analysis. Additionally, it provides several visualization approaches. CPADS significantly aids physicians and researchers in exploring primary and secondary drug resistance at both gene and pathway levels. The integration of drug resistance and gene perturbation data also presents novel perspectives for identifying pivotal genes influencing drug resistance. Access CPADS at https://smuonco.shinyapps.io/CPADS/ or https://robinl-lab.com/CPADS. Anqi Lin, Jiayi Xie, Jianguo Zhou, Shamus R. Carr, Zaoqu Liu, Jian Zhang 0104, David S. Schrump, Peng Luo 0005 |
Briefings Bioinform. | 3 |
| 2024 | A review of crowdsourced geographic information for land-use and land-cover mapping: current progress and challengesabstractThe emergence of crowdsourced geographic information (CGI) has markedly accelerated the evolution of land-use and land-cover (LULC) mapping. This approach taps into the collective power of the public to share spatial information, providing a relevant data source for producing LULC maps. Through the analysis of 262 papers published from 2012 to 2023, this work provides a comprehensive overview of the field, including prominent researchers, key areas of study, major CGI data sources, mapping methods, and the scope of LULC research. Additionally, it evaluates the pros and cons of various data sources and mapping methods. The findings reveal that while applying CGI with LULC labels is a common way by using spatial analysis, it is limited by incomplete CGI coverage and other data quality issues. In contrast, extracting semantic features from CGI for LULC interpretation often requires integrating multiple CGI datasets and remote sensing imagery, alongside advanced methods such as ensemble and deep learning. The paper also delves into the challenges posed by the quality of CGI data in LULC mapping and explores the promising potential of introducing large language models to overcome these hurdles. Hao Wu 0004, Yan Li 0114, Anqi Lin, Hongchao Fan, Kaixuan Fan, Junyang Xie, Wenting Luo |
Int. J. Geogr. Inf. Sci. | 3 |
| 2024 | PESSA: A web tool for pathway enrichment score-based survival analysis in cancerabstractThe activation levels of biologically significant gene sets are emerging tumor molecular markers and play an irreplaceable role in the tumor research field; however, web-based tools for prognostic analyses using it as a tumor molecular marker remain scarce. We developed a web-based tool PESSA for survival analysis using gene set activation levels. All data analyses were implemented via R. Activation levels of The Molecular Signatures Database (MSigDB) gene sets were assessed using the single sample gene set enrichment analysis (ssGSEA) method based on data from the Gene Expression Omnibus (GEO), The Cancer Genome Atlas (TCGA), The European Genome-phenome Archive (EGA) and supplementary tables of articles. PESSA was used to perform median and optimal cut-off dichotomous grouping of ssGSEA scores for each dataset, relying on the survival and survminer packages for survival analysis and visualisation. PESSA is an open-access web tool for visualizing the results of tumor prognostic analyses using gene set activation levels. A total of 238 datasets from the GEO, TCGA, EGA, and supplementary tables of articles; covering 51 cancer types and 13 survival outcome types; and 13,434 tumor-related gene sets are obtained from MSigDB for pre-grouping. Users can obtain the results, including Kaplan-Meier analyses based on the median and optimal cut-off values and accompanying visualization plots and the Cox regression analyses of dichotomous and continuous variables, by selecting the gene set markers of interest. PESSA (https://smuonco.shinyapps.io/PESSA/ OR http://robinl-lab.com/PESSA) is a large-scale web-based tumor survival analysis tool covering a large amount of data that creatively uses predefined gene set activation levels as molecular markers of tumors. Anqi Lin, Chang Qi, Zaoqu Liu, Kai Miao, Jian Zhang 0104, Peng Luo 0005 |
PLoS Comput. Biol. | 3 |
| 2022 | CAMOIP: a web server for comprehensive analysis on multi-omics of immunotherapy in pan-cancerabstractImmune checkpoint inhibitors (ICIs) have completely changed the approach pertaining to tumor diagnostics and treatment. Similarly, immunotherapy has also provided much needed data about mutation, expression and prognosis, affording an unprecedented opportunity for discovering candidate drug targets and screening for immunotherapy-relevant biomarkers. Although existing web tools enable biologists to analyze the expression, mutation and prognostic data of tumors, they are currently unable to facilitate data mining and mechanism analyses specifically related to immunotherapy. Thus, we effectively developed our own web-based tool, called Comprehensive Analysis on Multi-Omics of Immunotherapy in Pan-cancer (CAMOIP), in which we are able to successfully screen various prognostic markers and analyze the mechanisms involved in biomarker expression and function, as well as immunotherapy. The analyses include information relevant to survival analysis, expression analysis, mutational landscape analysis, immune infiltration analysis, immunogenicity analysis and pathway enrichment analysis. This comprehensive analysis of biomarkers for immunotherapy can be carried out by a click of CAMOIP, and the software should greatly encourage the further development of immunotherapy. CAMOIP provides invaluable evidence that bridges the information between the data of cancer genomics based on immunotherapy, providing comprehensive information to users and assisting in making the value of current ICI-treated data available to all users. CAMOIP is available at https://www.camoip.net. Anqi Lin, Chang Qi, Zaoqu Liu, Peng Luo 0005, Jian Zhang 0104 |
Briefings Bioinform. | 1 |
| 2022 | A Double Dictionary-Based Nonlinear Representation Model for Hyperspectral Subpixel Target DetectionabstractDue to the limitations of hardware technology and budget constraints, there always exists a tradeoff between spatial and spectral resolutions in a hyperspectral image (HSI). Because of the limited spatial resolution, mixed pixels are a common issue in HSIs, and consequently, some targets appear as subpixels. The effectiveness of hyperspectral target detection is affected greatly by the subpixel targets, especially when the size of the targets is small. In this article, we proposed a double dictionary-based nonlinear representation model for hyperspectral subpixel target detection (DDNRTD). DDNRTD represents HSIs with a nonlinear model based on background and target dictionaries, which fully considers the spatial property of background and targets and can separate background and targets reliably, especially for small-sized subpixel targets. In addition, we designed an over-completed background dictionary construction strategy to represent the background part more effectively, which integrates spectral angle distance (SAD) with sparse representation. Experiments on two simulated and five real HSI datasets showed that the proposed DDNRTD method produced more accurate detection results than six state-of-the-art methods. Xiaoyi Wang 0004, Liguo Wang 0001, Hao Wu 0004, Kaipeng Sun, Anqi Lin, Qunming Wang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | A comprehensive quality assessment framework for linear features from Volunteered Geographic InformationabstractThe majority of spatial data provided as Volunteered Geographic Information (VGI) are roads and other linear map features. Such data have been widely used in routing and navigation, road network update, emergency response, urban planning and more. Due to the lack of cartographic standards and issues with volunteer credibility, the quality of VGI linear features remains a concern and could seriously hinder the broad application of VGI data. This research proposes a comprehensive quality assessment framework for VGI linear features which adopts factor analysis to integrate two novel quality metrics with six other commonly used metrics, and further examines the spatial autocorrelation and semantic correlation of VGI linear feature quality. The OpenStreetMap road network of Allegheny County, Pennsylvania (USA) was selected as an example to test the proposed framework. Our results suggest that the proposed metrics, Box-counting dimension difference and Link accuracy are feasible for detecting quality issues and are important supplements to the common quality metrics. The findings also show that significant spatial autocorrelation exists in spatial completeness, positional accuracy, and logical consistency. Road type such as Tertiary, Residential, Service and Link has been proven to be a typical indicator of the different quality elements for VGI linear features. Hao Wu 0004, Anqi Lin, Keith C. Clarke, Wenzhong Shi, Abraham Cardenas-Tristan, Zhenfa Tu |
Int. J. Geogr. Inf. Sci. | 2 |
| 2019 | Examining the sensitivity of spatial scale in cellular automata Markov chain simulation of land use changeabstractUnderstanding the spatial scale sensitivity of cellular automata is crucial for improving the accuracy of land use change simulation. We propose a framework based on a response surface method to comprehensively explore spatial scale sensitivity of the cellular automata Markov chain (CA-Markov) model, and present a hybrid evaluation model for expressing simulation accuracy that merges the strengths of the Kappa coefficient and of Contagion index. Three Landsat-Thematic Mapper remote sensing images of Wuhan in 1987, 1996, and 2005 were used to extract land use information. The results demonstrate that the spatial scale sensitivity of the CA-Markov model resulting from individual components and their combinations are both worthy of attention. The utility of our proposed hybrid evaluation model and response surface method to investigate the sensitivity has proven to be more accurate than the single Kappa coefficient method and more efficient than traditional methods. The findings also show that the CA-Markov model is more sensitive to neighborhood size than to cell size or neighborhood type considering individual component effects. Particularly, the bilateral and trilateral interactions between neighborhood and cell size result in a more remarkable scale effect than that of a single cell size. Hao Wu 0004, Keith C. Clarke, Wenzhong Shi, Linchuan Fang, Anqi Lin |
Int. J. Geogr. Inf. Sci. | 6 |