VLDB 2026 Research / reviewers in the wild / expert
Jingyi Qiu
dblp:257/6872
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-7014-8375ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task LearningabstractRelation Extraction (RE) is a key task in table understanding, aiming to extract semantic relations between columns.However, complex tables with hierarchical headers are hard to obtain high-quality textual formats (e.g., Markdown) for input under practical scenarios like webpage screenshots and scanned documents, while table images are more accessible and intuitive.Besides, existing works overlook the need of mining relations among multiple columns rather than just the semantic relation between two specific columns in real-world practice.In this work, we explore utilizing Multimodal Large Language Models (MLLMs) to address RE in tables with complex structures.We creatively extend the concept of RE to include calculational relations, enabling multi-task learning of both semantic and calculational RE for mutual reinforcement.Specifically, we reconstruct table images into graph structure based on neighboring nodes to extract graph-level visual features.Such feature enhancement alleviates the insensitivity of MLLMs to the positional information within table images.We then propose a Chain-of-Thought distillation framework with self-correction mechanism to enhance MLLMs' reasoning capabilities without increasing parameter scale.Our method significantly outperforms most baselines on wide datasets.Additionally, we release a benchmark dataset for calculational RE in complex tables. Aibo Song, Jingyi Qiu, Jiahui Jin 0001, Tianbo Zhang, Xiaolin Fang 0001 |
ACL (1) | 3 |
| 2025 | Graph Representation-Aware Online Aggregations over Knowledge Graph
Jingyi Qiu, Aibo Song, Tongwei Liu, Tianbo Zhang |
KSEM (2) | 1 |
| 2024 | Relation-oriented few-shot knowledge graph prototype networks
Yingying Xue, Aibo Song, Jiahui Jin 0001, Jingyi Qiu, Xiaolin Fang 0001, Xiaorui Zhai |
Neurocomputing | 5 |
| 2024 | Matching Tabular Data to Knowledge Graph with Effective Core Column Set DiscoveryabstractMatching tabular data to a knowledge graph (KG) is critical for understanding the semantic column types, column relationships, and entities of a table. Existing matching approaches rely heavily on core columns that represent primary subject entities on which other columns in the table depend. However, discovering these core columns before understanding the table’s semantics is challenging. Most prior works use heuristic rules, such as the leftmost column, to discover a single core column, while an insightful discovery of the core column set that accurately captures the dependencies between columns is often overlooked. To address these challenges, we introduce Dependency-aware Core Column Set Discovery ( DaCo ), an iterative method that uses a novel rough matching strategy to identify both inter-column dependencies and the core column set. Additionally, DaCo can be seamlessly integrated with pre-trained language models, as proposed in the optimization module. Unlike other methods, DaCo does not require labeled data or contextual information, making it suitable for real-world scenarios. In addition, it can identify multiple core columns within a table, which is common in real-world tables. We conduct experiments on six datasets, including five datasets with single core columns and one dataset with multiple core columns. Our experimental results show that DaCo outperforms existing core column set detection methods, further improving the effectiveness of table understanding tasks. Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Jiaoyan Chen 0001, Xiaolin Fang 0001, Tianbo Zhang |
ACM Trans. Web | 1 |
| 2023 | Dependency-Aware Core Column Discovery for Table Understanding
Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Tianbo Zhang, Jingyi Ding, Xiaolin Fang 0001, Jianguo Qian |
ISWC | 1 |
| 2021 | BSDP: A Novel Balanced Spark Data PartitionerabstractAs a memory-based distributed big data computing framework, Spark has been widely used in big data processing systems. However, during the execution of Spark, due to the imbalance of input data distribution and the shortage of existing data partitioners in Spark, it is easy to cause partition skew problem and reduce the execution efficiency of Spark. Aiming at this problem, this paper proposes a balanced Spark data partitioner called BSDP (Balanced Spark Data Partitioner). By deeply analyzing the partitioning characteristics of Shuffle intermediate data, the Spark Shuffle intermediate data equalization partitioning model is established. The model aims to minimize the partition skew and find a Shuffle intermediate data equalization partitioning strategy. Based on the model, this paper designs and implements a data equalization partitioning algorithm of BSDP. This algorithm transforms the Shuffle intermediate data equalization partitioning problem into a classic List-Scheduling task scheduling problem, effectively realizes the balanced partitioning of Shuffle intermediate data. The experiment verifies that the BSDP can effectively realize the balanced partitioning of the Shuffle intermediate data and improve the execution efficiency of Spark. Aibo Song, Bowen Peng, Jingyi Qiu, Yingying Xue, Mingyang Du |
ICPADS | 3 |
| 2019 | Query optimization Approach with Shuffle Intermediate Cache Layer for Spark SQLabstractSpark SQL is a big data processing tool for structured data query and analysis. However, due to the execution of Spark SQL, there are multiple times to write intermediate data to the disk, which reduces the execution efficiency of Spark SQL. Targeting on the existing issues, we design and implement an intermediate data cache layer between the underlying file system and the upper Spark core to reduce the cost of random disk I/O. By using the query pre-analysis module, we can dynamically adjust the capacity of cache layer for different queries. And the allocation module can allocate proper memory for each node in cluster. This paper develops the SSO (Spark SQL optimizer) module and integrates it into the original Spark system to achieve the above functions. This paper compares the query performance with the existing Spark SQL by experiment data generated by TPC-H tool. The experimental results show that the SSO module can effectively improve the query efficiency, reduce the disk I/O cost and make full use of the cluster memory resources. Mingyu Zhai, Aibo Song, Jingyi Qiu, Xuechun Ji, Qingxi Wu |
IPCCC | 3 |