EDBT 2026 Demo / reviewers in the wild / expert
Jiaming Tian
dblp:232/4862
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0002-3224-4803ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement Learning with Verbalized Probabilities for LLM ClassificationabstractWhile Large Language Models (LLMs) excel at many reasoning tasks, their native inability to produce calibrated, multi-class probability distributions limits their use in high-stakes Web applications like content moderation and fraud detection. Existing methods to elicit probabilities from LLMs either sacrifice their crucial Chain-of-Thought (CoT) reasoning capabilities or suffer from poor calibration. To address this, we introduce a new paradigm, Verbalized Probability Distribution, and a novel training framework, RLVP (Reinforcement Learning with Verbalized Probabilities). RLVP fine-tunes an LLM to generate both an interpretable CoT and a complete, verbalized probability distribution. We overcome the ''insufficient reward granularity'' problem in standard Reinforcement Learning (RL) for classification by using soft probabilities from expert tabular models as a dense reward curriculum. Through large-scale joint training on 169 tabular tasks, we demonstrate that a single RLVP-trained model can surpass a strong, task-specific XGBoost baseline on up to 55% of tasks. More importantly, the trained model achieves state-of-the-art few-shot performance on unseen, heterogeneous Web benchmarks that mix structured data with free text, achieving performance comparable to or superior than expert models trained on the same limited data. This showcases a strong capability for generalization and knowledge transfer to complex Web data. Our work presents a viable path toward building general-purpose, probabilistically-sound, and interpretable foundation models for the Web. Liyao Li, Hao Chen 0081, Jiaming Tian, Wentao Ye, Lirong Gao, Chao Ye 0002, Ningtao Wang, Yu Cheng 0005, Haobo Wang 0001, Gang Chen 0001, Junbo Zhao 0002 |
WWW | 3 |
| 2026 | Toward real-world Table Agents: capabilities, workflows, and design principles for LLM-based table intelligence
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang 0001, Lingxin Wang, Lihua Yu, Zujie Ren, Gang Chen 0001, Junbo Zhao 0002 |
World Wide Web (WWW) | 1 |
| 2025 | MARAG: Multi‑agent Retrieval‑Augmented Generation for Mitigating Knowledge Conflicts in Large Language Models
Jiaming Tian, Weixin Zeng, Jibing Wu, Lihua Liu 0002, Xiang Zhao 0002 |
WISA | 1 |
| 2025 | Dual RAG: An Effective Graph-Based RAG Framework with Adaptively Integrating Knowledge Graphs and Chunks
Jiaming Tian, Zhenbo Fu, Qiange Wang, Chaoyi Chen, Minghe Yu 0001, Yanfeng Zhang 0001, Ge Yu 0001 |
IEEE Big Data | 1 |
| 2025 | Table as a Modality for Large Language ModelsabstractTo migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed tabular data. Despite that, in this work, by showing a probing experiment on our proposed StructQA benchmark, we postulate that even the most advanced LLMs (such as GPTs) may still fall short of coping with tabular data. More specifically, the current scheme often simply relies on serializing the tabular data, together with the meta information, then inputting them through the LLMs. We argue that the loss of structural information is the root of this shortcoming. In this work, we further propose TAMO, which bears an ideology to treat the tables as an independent modality integrated with the text tokens. The resulting model in TAMO is a multimodal framework consisting of a hypergraph neural network as the global table encoder seamlessly integrated with the mainstream LLM. Empirical results on various benchmarking datasets, including HiTab, WikiTQ, WikiSQL, FeTaQA, and StructQA, have demonstrated significant improvements on generalization with an average relative gain of **42.65%**. Liyao Li, Chao Ye 0002, Wentao Ye, Haobo Wang 0001, Jiaming Tian, Yiming Zhang 0023, Ningtao Wang, Gang Chen 0001, Junbo Zhao 0002 |
NeurIPS | 7 |
| 2025 | Understanding LLMs: A comprehensive overview from training to inference
Tianle Han, Jiaming Tian, Yutong Zhang 0019, Jiaqi Wang 0010, Xiaohui Gao, Tianyang Zhong, Yi Pan 0001, Shaochen Xu, Zihao Wu 0001, Zhengliang Liu, Xin Zhang 0151, Shu Zhang 0001, Xintao Hu, Ning Qiang, Tianming Liu 0001, Bao Ge |
Neurocomputing | 6 |
| 2025 | Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph ProcessingabstractLeveraging GPUs' high parallelism can significantly improve the real-time computation efficiency of streaming graph processing. However, when a large-scale graph exceeds GPU memory capacity, CPU-GPU cooperative processing often results in substantial and irregular CPU-to-GPU data transfer overhead. This stems from the extensive redundant graph accesses during continuous computation, which can hardly be addressed by existing solutions. In this work, we present Grapin, an out-of-memory GPU streaming graph processing system designed to minimize graph data transfer via two effective techniques for eliminating redundant accesses: (1) Extending advanced incremental processing algorithms to GPUs by converting their heavyweight data dependency processing into GPU-friendly forms, eliminating redundant graph accesses from the computation side; and (2) providing a lightweight yet efficient GPU hot subgraph management framework that finely caches the frequently accessed dynamic subgraphs in a vertex-centric manner. Experimental results demonstrate that Grapin can efficiently process large-scale streaming graphs with billions of edges on a single NVIDIA A5000 GPU. Enabling incremental computation reduces data transfer by 61%, and the integration of GPU hot subgraph reuse further reduces the remaining transfer by 72%, resulting in a total reduction of 89%. Compared with CPU-based solutions, Grapin achieves speedups ranging from 1.8x to 96.9x (17.9x on average). Qiange Wang, Yongze Yan, Hongshi Tan, Cheng Chen 0008, Cheng Zhao 0001, Jiaming Tian, Xiaoliang Cong, Yanfeng Zhang 0001, Ge Yu 0001, Weng-Fai Wong, Bingsheng He |
Proc. VLDB Endow. | 6 |