Boyan Niu

dblp:381/6129 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0000-3242-6864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data models and query languages · 55% Data mining · 37% Distributed and cloud data management · 8%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
exploratory data analysis
1.022025
Chat2Query: A Zero-Shot Automatic Exploratory Data Analysis System with Large Language Models · ICDE 2024
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models · Proc. VLDB Endow. 2025
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.912025
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models · Proc. VLDB Endow. 2025
Visualization and visual analytics
data visualization
0.912025
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models · Proc. VLDB Endow. 2025
Visualization and visual analytics › visualization generation › automated visualization generation
natural language to visualization
0.912025
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models · Proc. VLDB Endow. 2025
Data models and query languages › natural language interface
natural language interface to database
0.812024
Chat2Query: A Zero-Shot Automatic Exploratory Data Analysis System with Large Language Models · ICDE 2024
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.812024
Chat2Query: A Zero-Shot Automatic Exploratory Data Analysis System with Large Language Models · ICDE 2024

Methods — techniques the papers use, named apart from their topics

large language model · 3.4hierarchical data context · 2.6data-to-chart generation · 0.8
YearPublicationVenuePosition
2025 Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models
abstract
Exploratory data analysis (EDA), coupled with SQL, is essential for data analysts involved in data exploration and analysis. However, data analysts often encounter two primary challenges: (1) the need to craft SQL queries skillfully and (2) the requirement to generate suitable visualization types that enhance the interpretation of query results. Due to its significance, substantial research efforts have been made to explore different approaches to address these challenges, including leveraging large language models (LLMs). However, existing methods fail to meet real-world data exploration requirements primarily due to (1) complex database schema, (2) unclear user intent, (3) limited cross-domain generalization capability, and (4) insufficient end-to-end text-to-visualization capability. This paper presents TiInsight, an automated SQL-based cross-domain exploratory data analysis system. First, we propose a hierarchical data context (i.e., HDC), which leverages LLMs to summarize the contexts related to the database schema, which is crucial for open-world EDA systems to generalize across data domains. Second, the EDA system is divided into four components (i.e., stages): HDC generation, question clarification and decomposition, text-to-SQL generation (i.e., TiSQL), and data visualization (i.e., TiChart). Finally, we implemented an end-to-end EDA system with a user-friendly GUI in the production environment at PingCAP. We have also open-sourced all APIs of TiInsight to facilitate research within the EDA community. Through extensive evaluations by a real-world user study, we demonstrate that TiInsight offers remarkable performance compared to human experts. Additionally, TiSQL achieves an execution accuracy of 86.3% on the Spider dataset when using GPT-4. It also attains an execution accuracy of 60.98% on the Bird test dataset.
Jun-Peng Zhu, Boyan Niu, Peng Cai 0001, Zheming Ni, Jianwei Wan, Kai Xu 0003, Xuan Zhou 0001, Guanglei Bao
Proc. VLDB Endow.2
2024 Chat2Query: A Zero-Shot Automatic Exploratory Data Analysis System with Large Language Models
abstract
Data analysts often encounter two primary challenges while conducting exploratory data analysis by SQL: (1) the need to skillfully craft SQL queries, and (2) the requirement to generate suitable visualizations that enhance the interpretation of query results. The emergence of large language models (LLMs) has inaugurated a paradigm shift in text-to-SQL and data-to-chart. This paper presents Chat2Query, an LLM -empowered zero-shot automatic exploration data analysis system. Firstly, Chat2Query provides a user-friendly interface that allows users to employ natural languages to interact with the database directly. Secondly, Chat2Query offers an LLM -empowered text-to-SQL generator, SQL rewriter, SQL formatter, and data-to-chart generator. Thirdly, Chat2Query is uniquely distinguished by its underlying incorporation of the TiDB Serverless, fostering superior elasticity and scalability. This strategic integration empowers Chat2Query with the capability to seamlessly adapt to change workloads, aligning with the evolving demands of the user. We have implemented and deployed Chat2Query in the production environment, and demonstrate its usability and efficiency in three representative real-world scenarios.
Jun-Peng Zhu, Boyan Niu, Zheming Ni, Jianwei Wan
ICDE3