VLDB 2026 Research / reviewers in the wild / expert
I-Ta Lee
dblp:49/9859
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 82% Knowledge representation and reasoning · 16% Representation and self-supervised learning · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 67% Information retrieval · 33% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › narrative understanding
narrative extraction |
1.0 | 1 | 2026 | A Structured Clustering Approach for Inducing Media Narratives · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › text mining
text clustering |
1.0 | 1 | 2026 | A Structured Clustering Approach for Inducing Media Narratives · ACL (1) 2026 |
Information retrieval › evaluation
benchmark |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Recommender systems
recommender system evaluation |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Recommender systems › content recommendation
web page recommendation |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis
event embedding |
0.7 | 2 | 2019 | Multi-Relational Script Learning for Discourse Relations · ACL (1) 2019 FEEL: Featured Event Embedding Learning · AAAI 2018 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse relation recognition |
0.4 | 1 | 2019 | Multi-Relational Script Learning for Discourse Relations · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.3 | 1 | 2018 | FEEL: Featured Event Embedding Learning · AAAI 2018 |
Natural language and speech › Information extraction and text analysis › event analysis › event understanding
event sequence modeling |
0.3 | 1 | 2018 | FEEL: Featured Event Embedding Learning · AAAI 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
script learning |
0.3 | 1 | 2018 | FEEL: Featured Event Embedding Learning · AAAI 2018 |
Machine learning › Representation and self-supervised learning › text embedding
text representation learning |
0.1 | 1 | 2018 | FEEL: Featured Event Embedding Learning · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
clustering · 1.0prompted LLM baseline · 0.9event embedding · 0.7multi-relational learning · 0.4narrative cloze · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Structured Clustering Approach for Inducing Media NarrativesabstractRohan Das, Advait Deshmukh, Alexandria Leto, Zohar Naaman, I-Ta Lee, Maria Leonor Pacheco. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rohan Das 0001, Advait Deshmukh, Alexandria Leto, Zohar Naaman, I-Ta Lee, Maria Leonor Pacheco |
ACL (1) | 5 |
| 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden TestsabstractRecommender systems are among the most impactful AI applications, interacting with billions of users every day, guiding them to relevant products, services, or information tailored to their preferences.However, the research and development of recommender systems are hindered by existing datasets that fail to capture realistic user behaviors and inconsistent evaluation settings that lead to ambiguous conclusions.This paper introduces the \textbf{O}pen \textbf{R}ecommendation \textbf{B}enchmark for Reproducible Research with H\textbf{I}dden \textbf{T}ests (\textbf{ORBIT}), a unified benchmark for consistent and realistic evaluation of recommendation models. ORBIT offers a standardized evaluation framework of public datasets with reproducible splits and transparent settings for its public leaderboard. Additionally, ORBIT introduces a new webpage recommendation task, ClueWeb-Reco, featuring web browsing sequences from 87 million public, high-quality webpages. ClueWeb-Reco is a synthetic dataset derived from real, user-consented, and privacy-guaranteed browsing data. It aligns with modern recommendation scenarios and is reserved as the hidden test part of our leaderboard to challenge recommendation models' generalization ability. ORBIT measures 12 representative recommendation models on its public benchmark and introduces a prompted LLM baseline on the ClueWeb-Reco hidden test.Our benchmark results reflect general improvements of recommender systems on the public datasets, with variable individual performances.The results on the hidden test reveal the limitations of existing approaches in large-scale webpage recommendation and highlight the potential for improvements with LLM integrations.ORBIT benchmark, leaderboard, and codebase are available at \url{https://www.open-reco-bench.ai}. Jingyuan He, Jiongnan Liu 0001, Vishan Vishesh Oberoi, Bolin Wu, Mahima Jagadeesh Patel, Kangrui Mao, Chuning Shi, I-Ta Lee, Arnold Overwijk, Chenyan Xiong |
NeurIPS | 8 |
| 2021 | Modeling Human Mental States with an Entity-based Narrative GraphabstractUnderstanding narrative text requires capturing characters' motivations, goals, and mental states.This paper proposes an Entity-based Narrative Graph (ENG) to model the internalstates of characters in a story.We explicitly model entities, their interactions and the context in which they appear, and learn rich representations for them.We experiment with different task-adaptive pre-training objectives, in-domain training, and symbolic inference to capture dependencies between different decisions in the output space.We evaluate our model on two narrative understanding tasks: predicting character mental states, and desire fulfillment, and conduct a qualitative analysis. I-Ta Lee, Maria Leonor Pacheco, Dan Goldwasser |
NAACL-HLT | 1 |
| 2019 | Multi-Relational Script Learning for Discourse RelationsabstractModeling script knowledge can be useful for a wide range of NLP tasks.Current statistical script learning approaches embed the events, such that their relationships are indicated by their similarity in the embedding.While intuitive, these approaches fall short of representing nuanced relations, needed for downstream tasks.In this paper, we suggest to view learning event embedding as a multi-relational problem, which allows us to capture different aspects of event pairs.We model a rich set of event relations, such as Cause and Contrast, derived from the Penn Discourse Tree Bank.We evaluate our model on three types of tasks, the popular Mutli-Choice Narrative Cloze and its variants, several multi-relational prediction tasks, and a related downstream task-implicit discourse sense classification. I-Ta Lee, Dan Goldwasser |
ACL (1) | 1 |
| 2019 | ACE - An Anomaly Contribution Explainer for Cyber-Security ApplicationsabstractIn this paper we introduce Anomaly Contribution Explainer or ACE, a tool to explain security anomaly detection models in terms of the model features through a regression framework, and its variant, ACE-KL, which highlights the important anomaly contributors. ACE and ACE-KL provide insights in diagnosing which attributes significantly contribute to an anomaly by building a specialized linear model to locally approximate the anomaly score that a black-box model generates. We conducted experiments with these anomaly detection models to detect security anomalies on both synthetic data and real data. In particular, we evaluate performance on three public data sets: CERT insider threat, netflow logs, and Android malware. The experimental results are encouraging: our methods consistently identify the correct contributing feature in the synthetic data where ground truth is available; similarly, for real data sets, our methods point a security analyst in the direction of the underlying causes of an anomaly, including in one case leading to the discovery of previously overlooked network scanning activity. We have made our source code publicly available. Xiao Zhang 0017, Manish Marwah, I-Ta Lee, Martin F. Arlitt, Dan Goldwasser |
IEEE BigData | 3 |
| 2018 | FEEL: Featured Event Embedding LearningabstractStatistical script learning is an effective way to acquire world knowledge which can be used for commonsense reasoning. Statistical script learning induces this knowledge by observing event sequences generated from texts. The learned model thus can predict subsequent events, given earlier events. Recent approaches rely on learning event embeddings which capture script knowledge. In this work, we suggest a general learning model–Featured Event Embedding Learning (FEEL)–for injecting event embeddings with fine grained information. In addition to capturing the dependencies between subsequent events, our model can take into account higher level abstractions of the input event which help the model generalize better and account for the global context in which the event appears. We evaluated our model over three narrative cloze tasks, and showed that our model is competitive with the most recent state-of-the-art. We also show that our resulting embedding can be used as a strong representation for advanced semantic tasks such as discourse parsing and sentence semantic relatedness. I-Ta Lee, Dan Goldwasser |
AAAI | 1 |
| 2011 | A cooperative multicast routing protocol for mobile ad hoc networks
I-Ta Lee, Guann-Long Chiou, Shun-Ren Yang |
Comput. Networks | 1 |