VLDB 2026 Research / reviewers in the wild / expert
Jinguo You
dblp:39/395
· DBLP profile ↗
13ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-9118-3775ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | G²SQL: guided & guarded Text-to-SQL generation with two-stage verification
Jinguo You, Heng Li 0005, Jun Peng 0001, Ziheng Guo |
Expert Syst. Appl. | 2 |
| 2025 | Query Weak Equivalence and its Verification in Analytical DatabasesabstractModern database applications operate on massive data and support a range of complex queries, especially OLAP queries which are time-consuming. To accelerate query processing, a variety of methods for automatically verifying query equivalence have been proposed to avoid redundant executions of equivalent queries, mainly in a semantic sense. However, we have observed some queries that are not semantically equivalent also return the same tuples under the specific data distribution, which cannot be detected by most current automated verification of query equivalence. To deal with this issue, this paper proposes weak equivalence for identifying queries that are not semantically equivalent but produce the same results under the read-mostly scenarios such as OLAP. Specifically, for posed queries, we extract their filter condition expressions, which are then transformed into symbolic representations, namely first-order logic formulae. In terms of their partial order, i.e. containment relationship, we introduce Query Lattice, a novel structure that is constructed as a lattice which is partitioned into equivalence classes that are convex to answer queries if we determine they belong to the classes. The equivalence class enables stored queries to respond to future unseen queries so that redundant generation of query plan and execution can be bypassed. Experimental evaluation of Query Lattice built on top of a prevailing open-source DBMS, PostgreSQL shows that the maximum improvement that Query Lattice can achieve is 44.95 % over the original PostgreSQL, when running on the datasets of both TPC-H and TPC-H Skew benchmarks. Jinguo You, Wanting Fu, Peilei He, Quanqing Xu |
ICDE | 1 |
| 2025 | A Unified Computation Framework of Lattices in Hierarchical Data Analysis
Wen Shang, Jinguo You, Xingrui Huang, Jialin Xu |
KSEM (4) | 3 |
| 2025 | MatSciES: Automated Knowledge Extraction and Summarization from Materials Science Literature with Large Language Models
Jialin Xu, Jinguo You, Huaze Huang, Jingmei Tao, Jianhong Yi |
KSEM (4) | 2 |
| 2025 | SOC: A Succinct Adaptive Semantic OLAP CachingabstractAbstract In big data analysis, a large quantity of OLAP and aggregate queries exists, which have much stronger semantic context relationships (e.g. drill down and roll up) than generic SQL queries. Caching query results in memory playing an important role in accelerating data queries. Nevertheless, traditional query caching schemata neither fully utilize the features of OLAP, such as drill down and roll up semantics, nor compress the cached results, as the memory space is limited. In this paper, we propose a succinct, adaptive semantic OLAP caching, where the cache items are the cube lattice equivalence classes with only the bounds in a class stored. With further queries, the bound ranges are extended or expanded, indicating more query-answering ability which is assessed by the proposed covering capacity. The bounds of equivalence classes that more covering capacity are preferentially preserved in caching. We further empower our cache with some inference ability to derive more new data cells without posing extra queries and develop efficient query and update algorithms. The extensive experimental evaluation is conducted on synthetic and real data sets with various parameter settings. Our cache outperforms the common caching like LRU and LFU. Furthermore, it is robust to the non-repeated-pattern queries, still with a 30% hit ratio. Jinguo You, Xingrui Huang, Zhenrui Yi, Wanting Fu, Pengchen Zhang |
Data Sci. Eng. | 1 |
| 2025 | CM-SQL: A cross-model consistency framework for text-to-SQL
Jinguo You, Ziheng Guo |
Neurocomputing | 2 |
| 2025 | A multidimensional feature grouping sampling algorithm based on dynamic feedback of prior bias
Zongkai Shen, Jinguo You, Xiaoxia Zhao |
Inf. Sci. | 4 |
| 2025 | Efficient group based Hilbert encoding and decoding algorithms
Lianyin Jia, Songyu Wang, Shaowen Sun, Jiaman Ding, Mengjuan Li, Jinguo You, Shaojie Qiao |
Pattern Recognit. | 6 |
| 2022 | Civil airline fare prediction with a multi-attribute dual-stage attention mechanismabstractAirfare price prediction is one of the core facilities of the decision support system in civil aviation, which includes departure time, days of purchase in advance and flight airline. The traditional airfare price prediction system is limited by the nonlinear interrelationship of multiple factors and fails to deal with the impact of different time steps, resulting in low prediction accuracy. To address these challenges, this paper proposes a novel civil airline fare prediction system with a Multi-Attribute Dual-stage Attention (MADA) mechanism integrating different types of data extracted from the same dimension. In this method, the Seq2Seq model is used to add attention mechanisms to both the encoder and the decoder. The encoder attention mechanism extracts multi-attribute data from time series, which are optimized and filtered by the temporal attention mechanism in the decoder to capture the complex time dependence of the ticket price sequence. Extensive experiments with actual civil aviation data sets were performed, and the results suggested that MADA outperforms airfare prediction models based on the Auto-Regressive Integrated Moving Average (ARIMA), random forest, or deep learning models in MSE, RMSE, and MAE indicators. And from the results of a large amount of experimental data, it is proven that the prediction results of the MADA model proposed in this paper on different routes are at least 2.3% better than the other compared models. Jinguo You, Guoyu Gan, Xiaowu Li, Jiaman Ding |
Appl. Intell. | 2 |
| 2019 | A Parallel Uncertain Frequent Itemset Mining Algorithm with SparkabstractFrequent Itemset Mining (FIM) from large-scale databases has emerged as an important problem in the data mining and knowledge discovery research community. However, FIM suffers from three important limitations with the rapidly expanding of big data in all domains. First, it assumes that all items have the same importance. Second, it ignores the fact that data collected in a real-life environment is often inaccurate. Third, it is also a data-intensive and computation-intensive process which makes the FIM algorithm very time-consuming over large datasets. To address these issues, we propose a Parallel uncertain frequent itemset mining algorithm with spark (Pufim). Pufim firstly expresses item uncertainty by considering both the probability and weight, and calculates the maximum probability weight value of 1-items. Next, a distributed Pufim-tree structure is designed inspiring by FP-Tree for reducing the times of scanning the databases. Each node of Pufim-tree stores an item and its maximum probability weight value. Finally, experiments on publicly available UCI datasets demonstrate that Pufim achieves more prominent results than other related approaches across various metrics. In addition, the empirical study also shows Pufim has a good scalability. Jiaman Ding, Lianyin Jia, Jinguo You |
PDCAT | 5 |
| 2015 | Fast T-overlap query algorithms using graphics processor units and its applications in web data query
Mengjuan Li, Lianyin Jia, Jinguo You, Jianqing Xi, HaiFei Qin |
World Wide Web | 3 |
| 2012 | Bizard: An Online Multi-dimensional Data Analysis Visualization Tool
Zhuoluo Yang, Jinguo You |
APWeb | 2 |
| 2010 | Double Table Switch: An Efficient Partitioning Algorithm for Bottom-Up Computation of Data Cubes
Jinguo You, Lianyin Jia, Qingsong Huang, Jianqing Xi |
ADMA (2) | 1 |