VLDB 2026 Research / reviewers in the wild / expert
Peijia Qin
dblp:372/8155
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 43% Generative modeling · 19% Representation and self-supervised learning · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
1.0 | 1 | 2026 | BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models · ACL (1) 2026 |
Machine learning › Representation and self-supervised learning › vector quantization
codebook learning |
0.9 | 1 | 2025 | Learning to Quantize for Training Vector-Quantized Networks · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model |
0.9 | 1 | 2025 | Learning to Quantize for Training Vector-Quantized Networks · ICML 2025 |
Machine learning › Generative modeling › vector-quantized generative models
vector-quantized autoencoder |
0.9 | 1 | 2025 | Learning to Quantize for Training Vector-Quantized Networks · ICML 2025 |
Bioinformatics and computational biology
biomedical text mining |
0.3 | 1 | 2026 | BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
in-context learning · 2.0fine-tuning · 2.0vector quantization · 0.9meta-learning · 0.9hypernetwork · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language ModelsabstractDespite the success of large language models (LLMs) on general-purpose tasks, their performance in highly specialized domains such as biomedicine remains unsatisfactory.A key limitation is the inability of LLMs to effectively leverage biomedical tools, which clinical experts and biomedical researchers rely on extensively in daily workflows.While recent general-domain tool-calling datasets have substantially improved the capabilities of LLM agents, existing efforts in the biomedical domain largely rely on in-context learning and restrict models to a small set of tools.To address this gap, we introduce BIOTOOL, a comprehensive biomedical tool-calling dataset designed for fine-tuning LLMs.BIOTOOL comprises 34 frequently used tools collected from the NCBI, Ensembl, and UniProt databases, along with 7,040 high-quality, human-verified query-API call pairs spanning variation, genomics, proteomics, evolution, and general biology.Fine-tuning a 4-billion-parameter LLM on BIOTOOL yields substantial improvements in biomedical tool-calling performance, outperforming cutting-edge commercial LLMs such as GPT-5.1.Furthermore, human expert evaluations demonstrate that integrating a BIOTOOLfine-tuned tool caller significantly improves downstream answer quality compared to the same LLM without tool usage, highlighting the effectiveness of BIOTOOL in enhancing the biomedical capabilities of LLMs.The full dataset and evaluation code are available at https://github.com/gxx27/BioTool. Meixi Du, Peijia Qin, Pengtao Xie |
ACL (1) | 4 |
| 2026 | Online Learning in Open Data Space
Zhi Cao 0001, Peijia Qin, Chin-Teng Lin, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Learning to Quantize for Training Vector-Quantized NetworksabstractDeep neural networks incorporating discrete latent variables have shown significant potential in sequence modeling.
A notable approach is to leverage vector quantization (VQ) to generate discrete representations within a codebook.
However, its discrete nature prevents the use of standard backpropagation, which has led to challenges in efficient codebook training.
In this work, we introduce **Meta-Quantization (MQ)**, a novel vector quantization training framework inspired by meta-learning.
Our method separates the optimization of the codebook and the auto-encoder into two levels.
Furthermore, we introduce a hyper-net to replace the embedding-parameterized codebook, enabling the codebook to be dynamically generated based on the feedback from the auto-encoder.
Different from previous VQ objectives, our innovation results in a meta-objective that makes the codebook training task-aware.
We validate the effectiveness of MQ with VQVAE and VQGAN architecture on image reconstruction and generation tasks.
Experimental results showcase the superior generative performance of MQ, underscoring its potential as a robust alternative to existing VQ methods. Peijia Qin |
ICML | 1 |