EDBT 2026 Demo / reviewers in the wild / expert
Tao Chen 0008
dblp:69/510-8
· DBLP profile ↗
12ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-5334-219XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question AnsweringabstractPersonalization is well studied in search and recommendation, but personalized question answering remains underexplored due to challenges in inferring preferences from long, noisy, implicit contexts and generating responses that are both accurate and aligned with user expectations. To address this, we propose Pathways of Thoughts (PoT), an inference-stage method that applies to any large language model (LLM) without task-specific fine-tuning. PoT models the thinking as an iterative decision process, where the model dynamically selects among cognitive operations such as reasoning, revision, personalization, and clarification. This enables exploration of multiple reasoning trajectories, producing diverse candidate responses that capture different perspectives. PoT then aggregates and reweights these candidates according to inferred user preferences, yielding a final personalized response that benefits from the complementary strengths of diverse reasoning paths. Experiments on the LaMP-QA benchmark show that PoT consistently outperforms competitive baselines, achieving up to a 10.8% relative improvement. Human evaluation further validates these improvements, with annotators preferring PoT in 66% of cases compared to the best-performing baseline and reporting ties in 15% of cases. Alireza Salemi, Cheng Li 0012, Mingyang Zhang 0001, Qiaozhu Mei, Zhuowan Li, Spurthi Amba Hombaiah, Weize Kong, Tao Chen 0008, Hamed Zamani, Michael Bendersky |
WWW | 8 |
| 2023 | End-to-End Query Term WeightingabstractBag-of-words based lexical retrieval systems are still the most commonly used methods for real-world search applications. Recently deep learning methods have shown promising results to improve this retrieval performance but are expensive to run in an online fashion, non-trivial to integrate into existing production systems, and might not generalize well in out-of-domain retrieval scenarios. Instead, we build on top of lexical retrievers by proposing a Term Weighting BERT (TW-BERT) model. TW-BERT learns to predict the weight for individual n-gram (e.g., uni-grams and bi-grams) query input terms. These inferred weights and terms can be used directly by a retrieval system to perform a query search. To optimize these term weights, TW-BERT incorporates the scoring function used by the search engine, such as BM25, to score query-document pairs. Given sample query-document pairs we can compute a ranking loss over these matching scores, optimizing the learned query term weights in an end-to-end fashion. Aligning TW-BERT with search engine scorers minimizes the changes needed to integrate it into existing production applications, whereas existing deep learning based search methods would require further infrastructure optimization and hardware requirements. The learned weights can be easily utilized by standard lexical retrievers and by other retrieval techniques such as query expansion. We show that TW-BERT improves retrieval performance over strong term weighting baselines within MSMARCO and in out-of-domain retrieval on TREC datasets. Karan Samel, Cheng Li 0012, Weize Kong, Tao Chen 0008, Mingyang Zhang 0001, Shaleen Kumar Gupta, Swaraj Khadanga, Wensong Xu, Kashyap Kolipaka, Michael Bendersky, Marc Najork |
KDD | 4 |
| 2023 | APPCorp: a corpus for Android privacy policy document structure analysis
Shuang Liu 0007, Baiyang Zhao, Renjie Guo, Tao Chen 0008, Meishan Zhang |
Frontiers Comput. Sci. | 5 |
| 2022 | Out-of-Domain Semantics to the Rescue! Zero-Shot Hybrid Retrieval Models
Tao Chen 0008, Mingyang Zhang 0001, Jing Lu 0014, Michael Bendersky, Marc Najork |
ECIR (1) | 1 |
| 2021 | Dynamic Language Models for Continuously Evolving ContentabstractThe content on the web is in a constant state of flux. New entities,issues, and ideas continuously emerge, while the semantics of the existing conversation topics gradually shift. In recent years, pre-trained language models like BERT greatly improved the state-of-the-art for a large spectrum of content understanding tasks.Therefore, in this paper, we aim to study how these language models can be adapted to better handle continuously evolving web content.In our study, we first analyze the evolution of 2013 - 2019 Twitter data, and unequivocally confirm that a BERT model trained on past tweets would heavily deteriorate when directly applied to data from later years. Then, we investigate two possible sources of the deterioration: the semantic shift of existing tokens and the sub-optimal or failed understanding of new tokens. To this end, we both explore two different vocabulary composition methods, as well as propose three sampling methods which help in efficient incremental training for BERT-like models. Compared to a new model trained from scratch offline, our incremental training (a) reduces the training costs, (b) achieves better performance on evolving content, and (c)is suitable for online deployment. The superiority of our methods is validated using two downstream tasks. We demonstrate significant improvements when incrementally evolving the model from a particular base year, on the task of Country Hashtag Prediction, as well as on the OffensEval 2019 task. Spurthi Amba Hombaiah, Tao Chen 0008, Mingyang Zhang 0001, Michael Bendersky, Marc Najork |
KDD | 2 |
| 2019 | Identifying vulnerable older adult populations by contextualizing geriatric syndrome information in clinical notes of electronic health recordsabstractOBJECTIVE: Geriatric syndromes such as functional disability and lack of social support are often not encoded in electronic health records (EHRs), thus obscuring the identification of vulnerable older adults in need of additional medical and social services. In this study, we automatically identify vulnerable older adult patients with geriatric syndrome based on clinical notes extracted from an EHR system, and demonstrate how contextual information can improve the process. MATERIALS AND METHODS: We propose a novel end-to-end neural architecture to identify sentences that contain geriatric syndromes. Our model learns a representation of the sentence and augments it with contextual information: surrounding sentences, the entire clinical document, and the diagnosis codes associated with the document. We trained our system on annotated notes from 85 patients, tuned the model on another 50 patients, and evaluated its performance on the rest, 50 patients. RESULTS: Contextual information improved classification, with the most effective context coming from the surrounding sentences. At sentence level, our best performing model achieved a micro-F1 of 0.605, significantly outperforming context-free baselines. At patient level, our best model achieved a micro-F1 of 0.843. DISCUSSION: Our solution can be used to expand the identification of vulnerable older adults with geriatric syndromes. Since functional and social factors are often not captured by diagnosis codes in EHRs, the automatic identification of the geriatric syndrome can reduce disparities by ensuring consistent care across the older adult population. CONCLUSION: EHR free-text can be used to identify vulnerable older adults with a range of geriatric syndromes. Tao Chen 0008, Mark Dredze, Jonathan P. Weiner, Hadi Kharrazi |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | How Personality Affects our Likes: Towards a Better Understanding of Actionable ImagesabstractMessages like "If You Drink Don't Drive", "Each water drop count" or "Smoking causes cancer" are often paired with visual content in order to persuade an audience to perform specific actions, such as clicking a link, retweeting a post or purchasing a product. Despite its usefulness, the current way of discovering actionable images is entirely manual and typically requires marketing experts to filter over thousands of candidate images. To help understand the audience, marketers and social scientists have been investigating for years the role of personality in personalized services by leveraging AI technologies and social network data. In this work, we analyze how personality affects user actions on images in a social network website, and which visual stimuli contained in image content influence actions from users with certain Big Five traits. In order to achieve this goal, we ground this research on psychological studies which investigate the interplay between personality and emotions. Given a public Twitter dataset containing 1.6 million user-image timeline retweet actions, we carried out two extensive statistical analysis, which show significant correlation between personality traits and affective visual concepts in image content. We then proposed a novel model that combines user personality traits and image visual concepts for the task of predicting user actions in advance. This work is the first attempt to integrate personality traits and multimedia features, and moves an important step towards building personalized systems for automatically discovering actionable multimedia content. Francesco Gelli, Xiangnan He 0001, Tao Chen 0008, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2016 | Context-aware Image Tweet Modelling and RecommendationabstractWhile efforts have been made on bridging the semantic gap in image understanding, the in situ understanding of social media images is arguably more important but has had less progress. In this work, we enrich the representation of images in image tweets by considering their social context. We argue that in the microblog context, traditional image features, e.g., low-level SIFT or high-level detected objects, are far from adequate in interpreting the necessary semantics latent in image tweets. To bridge this gap, we move from the images' pixels to their context and propose a context-aware image bf tweet modelling (CITING) framework to mine and fuse contextual text to model such social media images' semantics. We start with tweet's intrinsic contexts, namely, 1) text within the image itself and 2) its accompanying text; and then we turn to the extrinsic contexts: 3) the external web page linked to by the tweet's embedded URL, and 4) the Web as a whole. These contexts can be leveraged to benefit many fundamental applications. To demonstrate the effectiveness our framework, we focus on the task of personalized image tweet recommendation, developing a feature-aware matrix factorization framework that encodes the contexts as a part of user interest modelling. Extensive experiments on a large Twitter dataset show that our proposed method significantly improves performance. Finally, to spur future studies, we have released both the code of our recommendation model and our image tweet dataset. Tao Chen 0008, Xiangnan He 0001, Min-Yen Kan |
ACM Multimedia | 1 |
| 2015 | VELDA: Relating an Image Tweet's Text and ImagesabstractImage tweets are becoming a prevalent form of socialmedia, but little is known about their content — textualand visual — and the relationship between the two mediums.Our analysis of image tweets shows that while visualelements certainly play a large role in image-text relationships, other factors such as emotional elements, also factor into the relationship. We develop Visual-Emotional LDA (VELDA), a novel topic model to capturethe image-text correlation from multiple perspectives (namely, visual and emotional). Experiments on real-world image tweets in both Englishand Chinese and other user generated content, show that VELDA significantly outperforms existingmethods on cross-modality image retrieval. Even in other domains where emotion does not factor in imagechoice directly, our VELDA model demonstrates good generalization ability, achieving higher fidelity modeling of such multimedia documents. Tao Chen 0008, Hany SalahEldeen, Xiangnan He 0001, Min-Yen Kan, Dongyuan Lu |
AAAI | 1 |
| 2015 | #mytweet via Instagram: Exploring User Behaviour across Multiple Social NetworksabstractWe study how users of multiple online social networks (OSNs) employ and share information by studying a common user pool that use six OSNs -- Flickr, Google+, Instagram, Tumblr, Twitter, and YouTube. We analyze the temporal and topical signature of users' sharing behaviour, showing how they exhibit distinct behaviorial patterns on different networks. We also examine cross-sharing (i.e., the act of user broadcasting their activity to multiple OSNs near-simultaneously), a previously-unstudied behaviour and demonstrate how certain OSNs play the roles of originating source and destination sinks. Bang Hui Lim, Dongyuan Lu, Tao Chen 0008, Min-Yen Kan |
ASONAM | 3 |
| 2015 | TriRank: Review-aware Explainable Recommendation by Modeling AspectsabstractMost existing collaborative filtering techniques have focused on modeling the binary relation of users to items by extracting from user ratings. Aside from users' ratings, their affiliated reviews often provide the rationale for their ratings and identify what aspects of the item they cared most about. We explore the rich evidence source of aspects in user reviews to improve top-N recommendation. By extracting aspects (i.e., the specific properties of items) from textual reviews, we enrich the user--item binary relation to a user--item--aspect ternary relation. We model the ternary relation as a heterogeneous tripartite graph, casting the recommendation task as one of vertex ranking. We devise a generic algorithm for ranking on tripartite graphs -- TriRank -- and specialize it for personalized recommendation. Experiments on two public review datasets show that it consistently outperforms state-of-the-art methods. Most importantly, TriRank endows the recommender system with a higher degree of explainability and transparency by modeling aspects in reviews. It allows users to interact with the system through their aspect preferences, assisting users in making informed decisions. Xiangnan He 0001, Tao Chen 0008, Min-Yen Kan, Xiao Chen 0004 |
CIKM | 2 |
| 2013 | Understanding and classifying image tweetsabstractSocial media platforms now allow users to share images alongside their textual posts. These image tweets make up a fast-growing percentage of tweets, but have not been studied in depth unlike their text-only counterparts. We study a large corpus of image tweets in order to uncover what people post about and the correlation between the tweet's image and its text. We show that an important functional distinction is between visually-relevant and visually-irrelevant tweets, and that we can successfully build an automated classifier utilizing text, image and social context features to distinguish these two classes, obtaining a macro F1 of 70.5%. Tao Chen 0008, Dongyuan Lu, Min-Yen Kan, Peng Cui 0001 |
ACM Multimedia | 1 |