VLDB 2026 Research / reviewers in the wild / expert
Tiancheng Hu
dblp:195/8059
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0006-7354-1088ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 25% Language models and text generation · 22% Planning, search and constraint satisfaction · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 50% Program analysis · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information |
1.0 | 1 | 2026 | Value of Information: A Framework for Human-Agent Communication · ACL (1) 2026 |
Collaborative and social computing
creative collaboration |
1.0 | 1 | 2026 | TRACE: A Corpus of Team Creative Discussions · ACL (1) 2026 |
Collaborative and social computing › social interaction
group dynamics |
1.0 | 1 | 2026 | TRACE: A Corpus of Team Creative Discussions · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › abusive language detection
hate speech detection |
0.9 | 1 | 2025 | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them · EMNLP 2025 |
Machine learning › Generative modeling › synthetic data generation
LLM-based data generation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Natural language and speech › Machine translation
low-resource machine translation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMs · EMNLP 2025 |
Web and social media mining
content moderation |
0.9 | 1 | 2025 | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them · EMNLP 2025 |
Multimedia analysis and retrieval › multimedia analysis
multimodal news analysis |
0.9 | 1 | 2025 | iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News · ACL (1) 2025 |
Cloud and datacenter computing
context caching |
0.9 | 1 | 2025 | EPIC: Efficient Position-Independent Caching for Serving Large Language Models · ICML 2025 |
Cloud and datacenter computing
KV cache reuse |
0.9 | 1 | 2025 | EPIC: Efficient Position-Independent Caching for Serving Large Language Models · ICML 2025 |
Cloud and datacenter computing › inference serving
LLM serving |
0.9 | 1 | 2025 | EPIC: Efficient Position-Independent Caching for Serving Large Language Models · ICML 2025 |
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation |
0.8 | 1 | 2024 | Quantifying the Persona Effect in LLM Simulations · ACL (1) 2024 |
Natural language and speech › Language models and text generation › prompting
persona prompting |
0.8 | 1 | 2024 | Quantifying the Persona Effect in LLM Simulations · ACL (1) 2024 |
Program analysis › program representation
abstract syntax tree analysis |
0.7 | 1 | 2023 | Fine-Grained Code Clone Detection with Block-Based Splitting of Abstract Syntax Tree · ISSTA 2023 |
Software maintenance and evolution
code clone detection |
0.7 | 1 | 2023 | Fine-Grained Code Clone Detection with Block-Based Splitting of Abstract Syntax Tree · ISSTA 2023 |
Methods — techniques the papers use, named apart from their topics
personalized classification · 2.6boundary classifier · 2.6zero-shot prediction · 1.7multimodal annotation · 1.7in-context learning · 1.7persona prompting · 1.5linear regression · 1.5sentence embeddings · 1.0information theory · 1.0factor analysis · 1.0legolink algorithm · 0.9back-translation · 0.9attention sink mitigation · 0.9LLM-based data generation · 0.9tree-based analysis · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Value of Information: A Framework for Human-Agent CommunicationabstractYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulić, Andreea Bobu, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulic, Andreea Bobu, Nigel Collier |
ACL (1) | 2 |
| 2026 | TRACE: A Corpus of Team Creative DiscussionsabstractUnderstanding how discussion dynamics shape team creativity has been limited by the difficulty of measuring process at scale.We introduce TRACE, a corpus of 309 group discussions from 103 teams (421 participants) across six creative problem-solving tasks.The dataset follows an input-process-output framework, integrating team composition (demographics, personalities), full discussion transcripts, and creativity outcomes.Using sentence embeddings and factor analysis, we identify four interpretable discussion dimensions:Coherence, Exploration, Convergence, and Participation.Analysis reveals a depth-breadth trade-off: coherent idea development inversely relates to semantic exploration.Larger teams explore more broadly but converge less effectively while team diversity shapes participation patterns more than discussion content.Novelty and usefulness in the creativity outcomes follow distinct pathways: Exploration and Convergence predict novelty, whereas Coherence predicts usefulness.These findings ground our understanding of how teams talk their way to creative solutions and provide guidance for designing multiagent systems. Yixuan Jiang, Tiancheng Hu, José Hernández-Orallo, David Stillwell, Luning Sun 0001 |
ACL (1) | 2 |
| 2025 | iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to NewsabstractUnderstanding how individuals perceive and react to information is fundamental for advancing social and behavioral sciences and developing human-centered AI systems.Current approaches often lack the granular data needed to model these personalized responses, relying instead on aggregated labels that obscure the rich variability driven by individual differences.We introduce iNews, a novel large-scale dataset specifically designed to facilitate the modeling of personalized affective responses to news content.Our dataset comprises annotations from 291 demographically diverse UK participants across 2,899 multimodal Facebook news posts from major UK outlets, with an average of 5.18 annotators per sample.For each post, annotators provide multifaceted labels including valence, arousal, dominance, discrete emotions, content relevance judgments, sharing likelihood, and modality importance ratings.Crucially, we collect comprehensive annotator persona information covering demographics, personality, media trust, and consumption patterns, which explain 15.2% of annotation variance -substantially higher than existing NLP datasets.Incorporating this information yields a 7% accuracy gain in zero-shot prediction and remains beneficial even with 32-shot in-context learning. Tiancheng Hu, Nigel Collier |
ACL (1) | 1 |
| 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMsabstractOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ona de Gibert Bonet, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann |
EMNLP | 7 |
| 2025 | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce ThemabstractPersonalized content moderation can protect users from harm while facilitating free expression by tailoring moderation decisions to individual preferences rather than enforcing universal rules.However, content moderation that is fully personalized to individual preferences, no matter what these preferences are, may lead to even the most hazardous types of content being propagated on social media.In this paper, we explore this risk using hate speech as a case study.Certain types of hate speech are illegal in many countries.We show that, while fully personalized hate speech detection models increase overall user welfare (as measured by user-level classification performance), they also make predictions that violate such legal hate speech boundaries, especially when tailored to users who tolerate highly hateful content.To address this problem, we enforce legal boundaries in personalized hate speech detection by overriding predictions from personalized models with those from a boundary classifier.This approach significantly reduces legal violations while minimally affecting overall user welfare.Our findings highlight both the promise and the risks of personalized moderation, and offer a practical solution to balance user preferences with legal and ethical obligations. Emanuele Moscato, Tiancheng Hu, Matthias Orlikowski, Paul Röttger, Debora Nozza |
EMNLP | 2 |
| 2025 | EPIC: Efficient Position-Independent Caching for Serving Large Language ModelsabstractLarge Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become more complex. Context caching improves serving performance by reusing Key-Value (KV) vectors, the intermediate representations of tokens that are repeated across requests. However, existing context caching requires exact prefix matches across requests, limiting reuse cases in settings such as few-shot learning and retrieval-augmented generation, where immutable content (e.g., documents) remains unchanged across requests but is preceded by varying prefixes. Position-Independent Caching (PIC) addresses this issue by enabling modular reuse of the KV vectors regardless of prefixes. We formalize PIC and advance prior work by introducing EPIC, a serving system incorporating our new LegoLink algorithm, which mitigates the inappropriate “attention sink” effect at every document beginning, to maintain accuracy with minimal computation. Experiments show that EPIC achieves up to 8$\times$ improvements in Time-To-First-Token (TTFT) and 7$\times$ throughput gains over existing systems, with negligible or no accuracy loss. Wenrui Huang, Haoyi Wang, Tiancheng Hu, Xusheng Chen, Yizhou Shan, Tao Xie 0001 |
ICML | 5 |
| 2024 | Quantifying the Persona Effect in LLM SimulationsabstractLarge language models (LLMs) have shown remarkable promise in simulating human language and behavior.This study investigates how integrating persona variables-demographic, social, and behavioral factors-impacts LLMs' ability to simulate diverse perspectives.We find that persona variables account for <10% variance in annotations in existing subjective NLP datasets.Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements.Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor.Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting.In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations.However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.1 Tiancheng Hu, Nigel Collier |
ACL (1) | 1 |
| 2023 | Quotatives Indicate Decline in Objectivity in U.S. Political NewsabstractAccording to journalistic standards, direct quotes should be attributed to sources with objective quotatives such as ``said'' and ``told,'' since nonobjective quotatives, e.g., ``argued'' and ``insisted,'' would influence the readers' perception of the quote and the quoted person. In this paper, we analyze the adherence to this journalistic norm to study trends in objectivity in political news across U.S. outlets of different ideological leanings. We ask: 1) How has the usage of nonobjective quotatives evolved? 2) How do news outlets use nonobjective quotatives when covering politicians of different parties? To answer these questions, we developed a dependency-parsing-based method to extract quotatives and applied it to Quotebank, a web-scale corpus of attributed quotes, obtaining nearly 7 million quotes, each enriched with the quoted speaker's political party and the ideological leaning of the outlet that published the quote. We find that, while partisan outlets are the ones that most often use nonobjective quotatives, between 2013 and 2020, the outlets that increased their usage of nonobjective quotatives the most were ``moderate'' centrist news outlets (around 0.6 percentage points, or 20% in relative percentage over seven years). Further, we find that outlets use nonobjective quotatives more often when quoting politicians of the opposing ideology (e.g., left-leaning outlets quoting Republicans) and that this ``quotative bias'' is rising at a swift pace, increasing up to 0.5 percentage points, or 25% in relative percentage, per year. These findings suggest an overall decline in journalistic objectivity in U.S. political news. Tiancheng Hu, Manoel Horta Ribeiro, Robert West 0001, Andreas Spitz |
ICWSM | 1 |
| 2023 | Fine-Grained Code Clone Detection with Block-Based Splitting of Abstract Syntax TreeabstractCode clone detection aims to find similar code fragments and gains increasing importance in the field of software engineering. There are several types of techniques for detecting code clones. Text-based or token-based code clone detectors are scalable and efficient but lack consideration of syntax, thus resulting in poor performance in detecting syntactic code clones. Although some tree-based methods have been proposed to detect syntactic or semantic code clones with decent performance, they are mostly time-consuming and lack scalability. In addition, these detection methods can not realize fine-grained code clone detection. They are unable to distinguish the concrete code blocks that are cloned. In this paper, we design Tamer, a scalable and fine-grained tree-based syntactic code clone detector. Specifically, we propose a novel method to transform the complex abstract syntax tree into simple subtrees. It can accelerate the process of detection and implement the fine-grained analysis of clone pairs to locate the concrete clone parts of the code. To examine the detection performance and scalability of Tamer, we evaluate it on a widely used dataset BigCloneBench. Experimental results show that Tamer outperforms ten state-of-the-art code clone detection tools (i.e., CCAligner, SourcererCC, Siamese, NIL, NiCad, LVMapper, Deckard, Yang2018, CCFinder, and CloneWorks). Tiancheng Hu, Zijing Xu, Yilin Fang, Yueming Wu 0001, Bin Yuan 0002, Deqing Zou, Hai Jin 0001 |
ISSTA | 1 |
| 2022 | The Causal News Corpus: Annotating Causal Relations in Event Sentences from NewsabstractDespite the importance of understanding causality, corpora addressing causal relations are limited. There is a discrepancy between existing annotation guidelines of event causality and conventional causality corpora that focus more on linguistics. Many guidelines restrict themselves to include only explicit relations or clause-based arguments. Therefore, we propose an annotation schema for event causality that addresses these concerns. We annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not. Our corpus is known as the Causal News Corpus (CNC). A neural network built upon a state-of-the-art pre-trained language model performed well with 81.20% F1 score on test set, and 83.46% in 5-folds cross-validation. CNC is transferable across two external corpora: CausalTimeBank (CTB) and Penn Discourse Treebank (PDTB). Leveraging each of these external datasets for training, we achieved up to approximately 64% F1 on the CNC test set without additional fine-tuning. CNC also served as an effective training and pre-training dataset for the two external corpora. Lastly, we demonstrate the difficulty of our task to the layman in a crowd-sourced annotation exercise. Our annotated corpus is publicly available, providing a valuable resource for causal text mining researchers. Fiona Anting Tan, Ali Hurriyetoglu, Tommaso Caselli, Nelleke Oostdijk, Tadashi Nomoto, Hansi Hettiarachchi, Iqra Ameer, Onur Uca, Farhana Ferdousi Liza, Tiancheng Hu |
LREC | 10 |
| 2022 | Temporal Head Pose Estimation From Point Cloud in Naturalistic Driving ConditionsabstractHead pose estimation is an important problem as it facilitates tasks such as gaze estimation and attention modeling. In the automotive context, head pose provides crucial information about the driver’s mental state, including drowsiness, distraction and attention. It can also be used for interaction with in-vehicle infotainment systems. While computer vision algorithms using RGB cameras are reliable in controlled environments, head pose estimation is a challenging problem in the car due to sudden illumination changes, occlusions and large head rotations that are common in a vehicle. These issues can be partially alleviated by using depth cameras. Head rotation trajectories are continuous with important temporal dependencies. Our study leverages this observation, proposing a novel temporal deep learning model for head pose estimation from point cloud. The approach extracts discriminative feature representation directly from point cloud data, leveraging the 3D spatial structure of the face. The frame-based representations are then combined withbidirectional long short term memory(BLSTM) layers. We train this model on the newly collectedmultimodal driver monitoring(MDM) dataset, achieving better results compared to non-temporal algorithms using point cloud data, and state-of-the-art models using RGB images. We further show quantitatively and qualitatively that incorporating temporal information provides large improvements not only in accuracy, but also in the smoothness of the predictions. Tiancheng Hu, Carlos Busso |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | The Multimodal Driver Monitoring Database: A Naturalistic Corpus to Study Driver AttentionabstractA smart vehicle should be able to monitor the actions and behaviors of the human driver to provide critical warnings or intervene when necessary. Recent advancements in deep learning and computer vision have shown great promise in monitoring human behavior and activities. While these algorithms work well in a controlled environment, naturalistic driving conditions add new challenges such as illumination variations, occlusions, and extreme head poses. A vast amount of in-domain data is required to train models that provide high performance in predicting driving related tasks to effectively monitor driver actions and behaviors. Toward building the required infrastructure, this paper presents themultimodal driver monitoring(MDM) dataset, which was collected with 59 subjects that were recorded performing various tasks. We use the Fi-Cap device that continuously tracks the head movement of the driver using fiducial markers, providing frame-based annotations to train head pose algorithms in naturalistic driving conditions. We ask the driver to look at predetermined gaze locations to obtain accurate correlation between the driver’s facial image and visual attention. We also collect data when the driver performs common secondary activities such as navigation using a smart phone and operating the in-car infotainment system. All of the driver’s activities are recorded with high definition RGB cameras and a time-of-flight depth camera. We also record thecontroller area network-bus(CAN-Bus), extracting important information. These high quality recordings serve as the ideal resource to train various efficient algorithms for monitoring the driver, providing further advancements in the field of in-vehicle safety systems. Mohamed F. Marzban, Tiancheng Hu, Mohamed Hany Mahmoud, Naofal Al-Dhahir, Carlos Busso |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Robust Driver Head Pose Estimation in Naturalistic Conditions from Point-Cloud DataabstractHead pose estimation has been a key task in computer vision since a broad range of applications often requires accurate information about the orientation of the head. Achieving this goal with regular RGB cameras faces challenges in automotive applications due to occlusions, extreme head poses and sudden changes in illumination. Most of these challenges can be attenuated with algorithms relying on depth cameras. This paper proposes a novel point-cloud based deep learning approach to estimate the driver's head pose from depth camera data, addressing these challenges. The proposed algorithm is inspired by the PointNet++ framework, where points are sampled and grouped before extracting discriminative features. We demonstrate the effectiveness of our algorithm by evaluating our approach on a naturalistic driving database consisting of 22 drivers, where the benchmark for the orientation of the driver's head is obtained with the Fi-Cap device. The experimental evaluation demonstrates that our proposed approach relying on point-cloud data achieves predictions that are almost always more reliable than state-of-the-art head pose estimation methods based on regular cameras. Furthermore, our approach provides predictions even for extreme rotations, which is not the case for the baseline methods. To the best of our knowledge, this is the first study to propose head pose estimation using deep learning on point-cloud data. Tiancheng Hu, Carlos Busso |
IV | 1 |