VLDB 2026 Research / reviewers in the wild / expert
Tianshi Zheng
dblp:341/1619
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0008-7053-6253ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 40% Knowledge representation and reasoning · 31% Efficient and distributed learning · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Knowledge graphs · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph construction |
1.0 | 1 | 2026 | AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM routing |
1.0 | 1 | 2026 | InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling · ACL (1) 2026 |
Machine learning › Reinforcement learning
reinforcement learning for structured prediction |
1.0 | 1 | 2026 | AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction · ACL (1) 2026 |
Knowledge graphs
knowledge graph construction |
1.0 | 1 | 2026 | AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora · ACL (1) 2026 |
Knowledge graphs › relation extraction
knowledge triple extraction |
1.0 | 1 | 2026 | AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora · ACL (1) 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery · EMNLP 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
logical entailment |
0.9 | 1 | 2025 | Enhancing Transformers for Generalizable First-Order Logical Entailment · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
0.9 | 1 | 2025 | LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning · EMNLP 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
scientific discovery |
0.9 | 1 | 2025 | From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery · EMNLP 2025 |
Natural language and speech › Language models and text generation › text generation
large language model generation |
0.8 | 1 | 2024 | Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › structure prediction
text-to-table generation |
0.8 | 1 | 2024 | Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction · EMNLP 2024 |
Knowledge graphs
knowledge graph reasoning |
0.8 | 1 | 2024 | Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation · ACL (1) 2024 |
Natural language and speech › Language models and text generation
large language model |
0.6 | 2 | 2026 | AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction · ACL (1) 2026 AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2026 | InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling · ACL (1) 2026 |
Natural language and speech › Language models and text generation › knowledge manipulation
knowledge augmentation |
0.3 | 1 | 2026 | AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora · ACL (1) 2026 |
Machine learning › Learning theory
model selection |
0.3 | 1 | 2026 | InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.3 | 1 | 2025 | LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning · EMNLP 2025 |
Computational science and engineering › AI for science
AI for scientific discovery |
0.3 | 1 | 2025 | From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.0conceptualization · 2.0reinforcement learning · 1.8survey · 1.7knowledge profiling · 1.0graph neural network · 1.0capability profiling · 1.0transformer architecture · 0.9large language model prompting · 0.9generative model · 0.8fine-tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaabstractWe present AutoSchemaKG, a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. Our system leverages large language models to simultaneously extract knowledge triples and induce comprehensive schemas directly from text, modeling both entities and events while employing conceptualization to organize instances into semantic categories. Processing over 50 million documents, we construct ATLAS (Automated Triple Linking And Schema induction), a family of knowledge graphs with 900+ million nodes and 5.9 billion edges. This approach outperforms state-of-the-art baselines on multi-hop QA tasks and enhances LLM factuality. Notably, our schema induction achieves 92\% semantic alignment with human-crafted schemas with zero manual intervention, demonstrating that billion-scale knowledge graphs with dynamically induced schemas can effectively complement parametric knowledge in large language models. Jiaxin Bai, Wei Fan 0001, Qing Zong, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Tianshi Zheng, Xi Peng 0006, Xin Yao 0008, Huiwen Yang, Leijie Wu, J. I Yi, Gong Zhang 0001, Renhai Chen, Yangqiu Song |
ACL (1) | 12 |
| 2026 | InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge ProfilingabstractHaochen Shi, Tianshi Zheng, Weiqi Wang, Baixuan Xu, Chunyang Li, Chunkit Chan, Tao Fan, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianshi Zheng, Weiqi Wang 0001, Baixuan Xu, Chunkit Chan, Tao Fan 0002, Yangqiu Song |
ACL (1) | 2 |
| 2026 | AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph ConstructionabstractHong Ting Tsang, Jiaxin Bai, Haoyu Huang, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hong Ting Tsang, Jiaxin Bai, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu 0001, Yangqiu Song |
ACL (1) | 5 |
| 2025 | Enhancing Transformers for Generalizable First-Order Logical EntailmentabstractTianshi Zheng, Jiazheng Wang, Zihao Wang, Jiaxin Bai, Hang Yin, Zheye Deng, Yangqiu Song, Jianxin Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianshi Zheng, Zihao Wang 0001, Jiaxin Bai, Hang Yin 0008, Zheye Deng, Yangqiu Song, Jianxin Li 0002 |
ACL (1) | 1 |
| 2025 | From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryabstractLarge Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and human-AI collaboration.This survey systematically charts this burgeoning field, placing a central focus on the changing roles and escalating capabilities of LLMs in science.Through the lens of the scientific method, we introduce a foundational three-level taxonomy-Tool, Analyst, and Scientist-to delineate their escalating autonomy and evolving responsibilities within the research lifecycle.We further identify pivotal challenges and future research trajectories such as robotic automation, self-improvement, and ethical governance.Overall, this survey provides a conceptual architecture and strategic foresight to navigate and shape the future of AI-driven scientific discovery, fostering both rapid innovation and responsible advancement. Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang 0001, Jiaxin Bai, Zihao Wang 0001, Yangqiu Song |
EMNLP | 1 |
| 2025 | LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM ReasoningabstractTianshi Zheng, Cheng Jiayang, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Tianshi Zheng, Cheng Jiayang, Zihao Wang 0001, Jiaxin Bai, Yangqiu Song, Ginny Y. Wong, Simon See |
EMNLP | 1 |
| 2024 | Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis GenerationabstractAbductive reasoning is the process of making educated guesses to provide explanations for observations.Although many applications require the use of knowledge for explanations, the utilization of abductive reasoning in conjunction with structured knowledge, such as a knowledge graph, remains largely unexplored.To fill this gap, this paper introduces the task of complex logical hypothesis generation, as an initial step towards abductive logical reasoning with KG.In this task, we aim to generate a complex logical hypothesis so that it can explain a set of observations.We find that the supervised trained generative model can generate logical hypotheses that are structurally closer to the reference hypothesis.However, when generalized to unseen observations, this training objective does not guarantee better hypothesis generation.To address this, we introduce the Reinforcement Learning from Knowledge Graph (RLF-KG) method, which minimizes differences between observations and conclusions drawn from generated hypotheses according to the KG.Experiments show that, with RLF-KG's assistance, the generated hypotheses provide better explanations, and achieve stateof-the-art results on three widely used KGs. 1 Jiaxin Bai, Tianshi Zheng, Xin Liu 0039, Yangqiu Song |
ACL (1) | 3 |
| 2024 | Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple ExtractionabstractThe task of condensing large chunks of textual information into concise and structured tables has gained attention recently due to the emergence of Large Language Models (LLMs) and their potential benefit for downstream tasks, such as text summarization and text mining.Previous approaches often generate tables that directly replicate information from the text, limiting their applicability in broader contexts, as text-to-table generation in real-life scenarios necessitates information extraction, reasoning, and integration.However, there is a lack of both datasets and methodologies towards this task.In this paper, we introduce LIVESUM, a new benchmark dataset created for generating summary tables of competitions based on real-time commentary texts.We evaluate the performances of state-of-the-art LLMs on this task in both fine-tuning and zero-shot settings, and additionally propose a novel pipeline called T3 (Text-Tuple-Table ) to improve their performances.Extensive experimental results demonstrate that LLMs still struggle with this task even after fine-tuning, while our approach can offer substantial performance gains without explicit training.Further analyses demonstrate that our method exhibits strong generalization abilities, surpassing previous approaches on several other text-to-table datasets. Zheye Deng, Chunkit Chan, Weiqi Wang 0001, Yuxi Sun 0010, Wei Fan 0001, Tianshi Zheng, Yauwai Yim, Yangqiu Song |
EMNLP | 6 |