VLDB 2026 Research / reviewers in the wild / expert
Tianyang Liu 0003
dblp:89/1676-3
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-7754-7029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsabstractYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang, Yifei Shao, Shibo Hao, Yi Gu, Jieyuan Liu, Somanshu Singla, Tianyang Liu, Eric P. Xing, Zhengzhong Liu, Haojian Jin, Zhiting Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yanbin Yin, Kun Zhou 0002, Zhen Wang 0041, Yifei Shao, Shibo Hao, Yi Gu 0002, Jieyuan Liu, Somanshu Singla, Tianyang Liu 0003, Eric P. Xing, Zhengzhong Liu 0001, Haojian Jin, Zhiting Hu |
ACL (1) | 10 |
| 2025 | Symbolic Representation for Any-to-Any Generative TasksabstractWe propose a symbolic generative task description language and a corresponding inference engine capable of representing arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models that rely on large-scale training and implicit neural representations to learn cross-modal mappings—often at high computational cost and with limited flexibility—our framework introduces an explicit symbolic representation comprising three core primitives: $\color{blue}{\text{functions}}$, $\color{green}{\text{parameters}}$, and $\color {Purple}{\text{topological}}\,{\text{logic}}$. Leveraging a pre-trained language model, our inference engine maps natural language instructions directly to symbolic workflows in a training-free manner. Our framework successfully performs over 12 diverse multimodal generative tasks, demonstrating strong performance and flexibility without the need for task-specific tuning. Experiments show that our method not only matches or outperforms existing state-of-the-art unified models in content quality, but also offers greater efficiency, editability, and interruptibility. We believe that symbolic task representations provide a cost-effective and extensible foundation for advancing the capabilities of generative AI. Xiaoye Zhu, Tianyang Liu 0003, Chak Tou Leong, Yifei Ke, Yiwen Yuan, Julian J. McAuley, Li-jia Li |
CVPR | 4 |
| 2025 | Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsabstractDayu Yang, Tianyang Liu, Daoan Zhang, Antoine Simoulin, Xiaoyi Liu, Yuwei Cao, Zhaopu Teng, Xin Qian, Grey Yang, Jiebo Luo, Julian McAuley. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dayu Yang, Tianyang Liu 0003, Daoan Zhang, Antoine Simoulin, Yuwei Cao, Zhaopu Teng, Grey Yang, Jiebo Luo 0001, Julian J. McAuley |
EMNLP | 2 |
| 2025 | Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain PerspectiveabstractReinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, covering six reasoning domains: Math, Code, Science, Logic, Simulation, and Tabular, each with corresponding verifiers. We build \ours via a careful data-curation pipeline, including sourcing, deduplication, reward design, and domain-specific and difficulty-based filtering, to facilitate the systematic investigation of cross-domain RL generalization. Our study using \ours suggests the efficacy of a simple mixed-domain RL training approach and reveals several key aspects affecting cross-domain transferability. We further train two models {\ours}-7B and {\ours}-32B purely with RL on our curated data and observe largely improved performance over leading open RL reasoning model baselines, with gains of 7.3\% and 7.8\% respectively on an extensive 17-task, six-domain evaluation suite. We are releasing our dataset, code, and evaluation suite to the community, aiming to support further research and development of more general RL-enhanced reasoning models. Jorge (Zhoujun) Cheng, Shibo Hao, Tianyang Liu 0003, Yuexin Bian, Nilabjo Dey, Yonghao Zhuang 0001, Yuheng Zha, Yi Gu 0002, Kun Zhou 0002, Yuan Li 0032, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Taylor W. Killian, Haonan Li 0002, Mikhail Yurochkin, Eric P. Xing, Zhengzhong Liu 0001, Zhiting Hu |
NeurIPS | 3 |
| 2024 | Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language ModelsabstractAligning Large Language Models (LLMs) traditionally relies on costly training and human preference annotations.Self-alignment aims to reduce these expenses by aligning models by themselves.To further minimize the cost and enable LLM alignment without any expensive tuning and annotations, we introduce a new tuning-free approach for self-alignment, called Dynamic Rewarding with Prompt Optimization (DRPO).Our approach leverages a search-based optimization framework that allows LLMs to iteratively self-improve and design the best alignment instructions without the need for additional training or human intervention.The core of DRPO is a dynamic rewarding mechanism, which identifies and rectifies model-specific alignment weaknesses, allowing LLMs to adapt efficiently to diverse alignment challenges.Empirical evaluations on eight recent LLMs, both open-and closed-source, reveal that DRPO significantly enhances alignment performance, with base models outperforming their SFT/RLHF-tuned counterparts.Moreover, DRPO's automatically optimized prompts surpass those curated by human experts, further validating the effectiveness of our approach.Our findings highlight the great potential of current LLMs to be adaptively selfaligned through inference-time optimization, complementing existing tuning-based alignment research. Somanshu Singla, Zhen Wang 0041, Tianyang Liu 0003, Abdullah Ashfaq, Zhiting Hu, Eric P. Xing |
EMNLP | 3 |
| 2024 | RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsabstractLarge Language Models (LLMs) have greatly advanced code auto-completion systems, with a potential for substantial productivity enhancements for developers. However, current benchmarks mainly focus on single-file tasks, leaving an assessment gap for more complex, real-world, multi-file programming scenarios. To fill this gap, we introduce RepoBench, a new benchmark specifically designed for evaluating repository-level code auto-completion systems. RepoBench consists of three interconnected evaluation tasks: RepoBench-R (Retrieval), RepoBench-C (Code Completion), and RepoBench-P (Pipeline). Each task respectively measures the system's ability to retrieve the most relevant code snippets from other files as cross-file context, predict the next line of code with cross-file and in-file context, and handle complex tasks that require a combination of both retrieval and next-line prediction. RepoBench aims to facilitate a more complete comparison of performance and encouraging continuous improvement in auto-completion systems. RepoBench is actively maintained with the latest code, serving as a live benchmark publicly available at https://github.com/Leolty/repobench. Tianyang Liu 0003, Canwen Xu, Julian J. McAuley |
ICLR | 1 |
| 2024 | Rethinking Tabular Data Understanding with Large Language ModelsabstractTianyang Liu, Fei Wang, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tianyang Liu 0003, Fei Wang 0060, Muhao Chen 0001 |
NAACL-HLT | 1 |
| 2023 | ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool EmbeddingsabstractIntegrating large language models (LLMs) with various tools has led to increased attention in the field. Existing approaches either involve fine-tuning the LLM, which is both computationally costly and limited to a fixed set of tools, or prompting LLMs by in-context tool demonstrations. Although the latter method offers adaptability to new tools, it struggles with the inherent context length constraint of LLMs when many new tools are presented, and mastering a new set of tools with few-shot examples remains challenging, resulting in suboptimal performance. To address these limitations, we propose a novel solution, named **ToolkenGPT**, wherein LLMs effectively learn to master tools as predicting tokens through **tool embeddings** for solving complex tasks. In this framework, each tool is transformed into vector embeddings and plugged into the language model head. Once the function is triggered during text generation, the LLM enters a special function mode to execute the tool calls. Our experiments show that function embeddings effectively help LLMs understand tool use and improve on several tasks, including numerical reasoning, knowledge-based question answering and embodied decision-making. Shibo Hao, Tianyang Liu 0003, Zhen Wang 0041, Zhiting Hu |
NeurIPS | 2 |
| 2023 | Architecture Decisions in AI-based Systems Development: An Empirical StudyabstractArtificial Intelligence (AI) technologies have been developed rapidly, and AI-based systems have been widely used in various application domains with opportunities and challenges. However, little is known about the architecture decisions made in AI-based systems development, which has a substantial impact on the success and sustainability of these systems. To this end, we conducted an empirical study by collecting and analyzing the data from Stack Overflow (SO) and GitHub. More specifically, we searched on SO with six sets of keywords and explored 32 AI-based projects on GitHub, and finally we collected 174 posts and 128 GitHub issues related to architecture decisions. The results show that in AI-based systems development (1) architecture decisions are expressed in six linguistic patterns, among which Solution Proposal and Information Giving are most frequently used, (2) Technology Decision, Component Decision, and Data Decision are the main types of architecture decisions made, (3) Game is the most common application domain among the eighteen application domains identified, (4) the dominant quality attribute considered in architecture decision-making is Performance, and (5) the main limitations and challenges encountered by practitioners in making architecture decisions are Design Issues and Data Issues. Our results suggest that the limitations and challenges when making architecture decisions in AI-based systems development are highly specific to the characteristics of AI-based systems and are mainly of technical nature, which need to be properly confronted. Beiqi Zhang, Tianyang Liu 0003, Peng Liang 0001, Chong Wang 0004, Mojtaba Shahin |
SANER | 2 |
| 2023 | RoseMatcher: Identifying the impact of user reviews on app updates
Tianyang Liu 0003, Chong Wang 0004, Peng Liang 0001, Beiqi Zhang, Maya Daneva, Marten van Sinderen |
Inf. Softw. Technol. | 1 |
| 2021 | The Role of User Reviews in App Updates: A Preliminary Investigation on App Release NotesabstractRelease planning for mobile apps has recently become an area of active research. Prior research in this area concentrated on the analysis of release notes and on tracking user reviews to support app evolution with issue trackers. However, little is known about the impact of user reviews on the evolution of mobile apps. Our work explores the role of user reviews in app updates based on release notes. For this purpose, we collected user reviews and release notes of Spotify, the ߢnumber one’ app in the ‘Music’ category in Apple App Store, as the research data. Then, we manually removed non-informative parts of each release note, and manually determined the relevance of the app reviews with respect to the release notes. We did this by using Word2Vec calculation techniques based on the top 80 app release notes with the highest similarities. Our empirical results show that more than 60 % of the matched reviews are actually irrelevant to the corresponding release notes. When zooming in at these relevant user reviews, we found that around half of them were posted before the new release and referred to requests, suggestions, and complaints. Whereas, the other half of the relevant user reviews were posted after updating the apps and concentrated more on bug reports and praise. Chong Wang 0004, Tianyang Liu 0003, Peng Liang 0001, Maya Daneva, Marten van Sinderen |
APSEC | 2 |