VLDB 2026 Research / reviewers in the wild / expert
Yanguang Liu
dblp:78/9719
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-3086-9586ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 44% Vision and language · 44% Trustworthy machine learning · 13% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding |
1.0 | 1 | 2026 | FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › document analysis
financial text analysis |
1.0 | 1 | 2026 | FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction · ACL (1) 2026 |
Bioinformatics and computational biology › drug discovery › drug design
de novo drug design |
0.8 | 1 | 2024 | TransGEM: a molecule generation model based on Transformer with gene expression data · Bioinform. 2024 |
Bioinformatics and computational biology
drug discovery |
0.8 | 1 | 2024 | TransGEM: a molecule generation model based on Transformer with gene expression data · Bioinform. 2024 |
Bioinformatics and computational biology
gene expression analysis |
0.8 | 1 | 2024 | TransGEM: a molecule generation model based on Transformer with gene expression data · Bioinform. 2024 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation |
0.8 | 1 | 2024 | TransGEM: a molecule generation model based on Transformer with gene expression data · Bioinform. 2024 |
Methods — techniques the papers use, named apart from their topics
large vision-language model · 1.0large language model · 1.0transformer · 0.8gene expression encoder · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise PredictionabstractPredicting corporate earnings surprises is a profitable yet challenging task, as accurate forecasts can inform significant investment decisions.However, progress in this domain has been constrained by a reliance on expensive, proprietary, and text-only data, limiting the development of advanced models.To address this gap, we introduce FinCall-Surprise (Financial Conference Call for Earning Surprise Prediction), the first large-scale, open-source, and multi-modal dataset for earnings surprise prediction.Comprising 2,688 unique corporate conference calls from 2019 to 2021, our dataset features word-to-word conference call textual transcripts, full audio recordings, and corresponding presentation slides.We establish a comprehensive benchmark by evaluating 26 state-of-the-art unimodal and multimodal LLMs.Our findings reveal that (1) while many models achieve high accuracy, this performance is often an illusion caused by significant class imbalance in the realworld data.(2) Some specialized financial models demonstrate unexpected weaknesses in instruction-following and language generation.(3) Although incorporating audio and visual modalities provides some performance gains, current models still struggle to leverage these signals effectively.These results highlight critical limitations in the financial reasoning capabilities of existing LLMs and establish a challenging new baseline for future research.The FinCall-Surprise dataset is available at https://github.com/Tizzzzy/ FinCall-Surprise. Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du |
ACL (1) | 2 |
| 2026 | FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language ModelsabstractLarge vision-language models (LVLMs) have made significant progress in chart understanding.However, financial charts, characterized by complex temporal structures and domainspecific terminology, remain notably underexplored.We introduce FinChart-Bench, the first benchmark specifically focused on realworld financial charts.FinChart-Bench comprises 1,200 financial chart images collected from 2015 to 2024, each annotated with True/-False (TF), Multiple Choice (MC), and Question Answering (QA) questions, totaling 7,016 questions.We conducted a comprehensive evaluation of 26 state-of-the-art LVLMs on FinChart-Bench.Our evaluation reveals critical insights: (1) the performance gap between open-source and closed-source models is narrowing, (2) performance degradation occurs in upgraded models within families, (3) many models struggle with instruction following, (4) both advanced models show significant limitations in spatial reasoning abilities, and (5) current LVLMs are not reliable enough to serve as automated evaluators.These findings highlight important limitations in current LVLM capabilities for financial chart understanding.The FinChart-Bench dataset is available at https: //github.com/Tizzzzy/FinChart-Bench. Dong Shu, Haoyang Yuan, Yanguang Liu, Huopu Zhang, Mengnan Du |
ACL (1) | 4 |
| 2025 | MCPybarra: A Multi-agent Framework for Low-Cost, High-Quality MCP Service Generation
Bocheng Peng, Yanguang Liu, Congcong Tian, Zhongjie Wang 0003 |
ICSOC (1) | 3 |
| 2024 | TransGEM: a molecule generation model based on Transformer with gene expression dataabstractMOTIVATION: It is difficult to generate new molecules with desirable bioactivity through ligand-based de novo drug design, and receptor-based de novo drug design is constrained by disease target information availability. The combination of artificial intelligence and phenotype-based de novo drug design can generate new bioactive molecules, independent from disease target information. Gene expression profiles can be used to characterize biological phenotypes. The Transformer model can be utilized to capture the associations between gene expression profiles and molecular structures due to its remarkable ability in processing contextual information. RESULTS: We propose TransGEM (Transformer-based model from gene expression to molecules), which is a phenotype-based de novo drug design model. A specialized gene expression encoder is used to embed gene expression difference values between diseased cell lines and their corresponding normal tissue cells into TransGEM model. The results demonstrate that the TransGEM model can generate molecules with desirable evaluation metrics and property distributions. Case studies illustrate that TransGEM model can generate structurally novel molecules with good binding affinity to disease target proteins. The majority of genes with high attention scores obtained from TransGEM model are associated with the onset of the disease, indicating the potential of these genes as disease targets. Therefore, this study provides a new paradigm for de novo drug design, and it will promote phenotype-based drug discovery. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/hzauzqy/TransGEM. Yanguang Liu, Xinya Duan, Yao Ruan, Qingye Zhang |
Bioinform. | 1 |
| 2010 | Graph based automatic centralized PCI assignment in LTEabstractPhysical Cell Identity (PCI) is cell identifier on the physical layer which can be used to create synchronization signals. There is a mapping between synchronization signals and physical cell identity, and then User Equipment (UE) can search cells through this mapping. Traditionally, PCI was configured during network planning process. While with 3G Long Term Evolution (LTE) has been prevalent in the future network, auto configuration of radio parameters has become a necessity. This paper proposed an automatic centralized PCI assignment mechanism using OAM as central server to collect cell information of the network, and create an abstract graph using information given which reflects the relationship in real-world network. After that, an enhanced graph-coloring algorithm which can greatly reduce time complexity while keeping a high PCI utility ratio was provided to create a collision and confusion free PCI for the new cell. Yanguang Liu, Weihao Lu 0001 |
ISCC | 1 |