VLDB 2026 Research / reviewers in the wild / expert
Guoxin Kang
dblp:286/9664
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5180-7053ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 46% Database system architecture and tuning · 31% Data integration and cleaning · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 100% |
Topics — the 3 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
1.4 | 2 | 2025 | BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases · Proc. VLDB Endow. 2025 OLxPBench: Real-time, Semantically Consistent, and Domain-specific are Essential in Benchmarking, Designing, and Implementing HTAP Systems · ICDE 2022 |
Information retrieval › similarity search
nearest neighbor search |
0.9 | 1 | 2025 | BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases · Proc. VLDB Endow. 2025 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.6 | 1 | 2022 | OLxPBench: Real-time, Semantically Consistent, and Domain-specific are Essential in Benchmarking, Designing, and Implementing HTAP Systems · ICDE 2022 |
Methods — techniques the papers use, named apart from their topics
benchmark suite design · 2.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OpenClinicalAI: An open and dynamic model for Alzheimer's Disease diagnosis
Yunyou Huang, Xiaoshuang Liang, Jiyue Xie, Xiangjiang Lu, Xiuxia Miao, Fan Zhang 0047, Guoxin Kang, Suqin Tang, Jianfeng Zhan |
Expert Syst. Appl. | 8 |
| 2025 | BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector DatabasesabstractVector databases are designed to effectively store, organize, and retrieve high-dimensional vectors, enabling faster and more accurate querying and analysis. This study highlights that the performance of cutting-edge vector databases hinges on their proficiency in managing heterogeneous data embedding and handling compound queries. The former task revolves around converting varied data types into a cohesive vector format, while the latter involves processing multimodal or single-modal queries with precise constraints. The paper advocates for evaluating these dual tasks within an integrated benchmark framework. However, state-of-the-art vector database benchmarks overlook heterogeneous data embedding and compound queries, creating a gap in evaluating vector database performance. To address this gap, we introduce BigVectorBench, a benchmark suite designed to evaluate vector database performance. BigVectorBench contributes by defining and evaluating the embedding performance of heterogeneous data. Additionally, it abstracts compound queries, which are increasingly used in real-world applications, replacing unimodal vector searches. Our rigorous evaluations validate the two design decisions of BigVectorBench and identify performance bottlenecks of mainstream vector databases. Its source code and user manual are available from https://github.com/BenchCouncil/BigVectorBench. Guoxin Kang, Zhongxin Ge, Jingpei Hu, Xueya Zhang, Lei Wang 0004, Jianfeng Zhan |
Proc. VLDB Endow. | 1 |
| 2022 | OLxPBench: Real-time, Semantically Consistent, and Domain-specific are Essential in Benchmarking, Designing, and Implementing HTAP SystemsabstractAs real-time analysis on the fresh data become in-creasingly compelling, more organizations deploy Hybrid Trans-actional/Analytical Processing (HTAP) systems to support real-time queries on data recently generated by online transaction processing. This paper argues that real-time queries, semantically consistent schema, and domain-specific workloads are essential in benchmarking, designing, and implementing HTAP systems. However, most state-of-the-art and state-of-the-practice bench-marks ignore those critical factors. Hence, at best, they are incommensurable and, at worst, misleading in benchmarking, designing, and implementing HTAP systems. This paper presents OLxPBench, a composite HTAP benchmark suite. OLxPBench proposes: (1) the abstraction of a hybrid transaction, performing a real-time query in-between an online transaction, to model widely-observed behavior pattern - making a quick decision while consulting real-time analysis; (2) a semantically consistent schema to express the relationships between OLTP and OLAP schema; (3) the combination of domain-specific and general benchmarks to characterize diverse application scenarios with varying resource demands. Our evaluations justify the three design decisions of OLxPBench and pinpoint the bottlenecks of two mainstream distributed HTAP DBMSs. International Open Benchmark Council (Bench Council) sets up the OLxP-Bench homepage at https://www.benchcouncil.org/olxpbench/. Its source code is available from https://github.com/BenchCouncil/olxpbench.git. Guoxin Kang, Lei Wang 0004, Wanling Gao, Fei Tang 0003, Jianfeng Zhan |
ICDE | 1 |