VLDB 2026 Research / reviewers in the wild / expert
Aria Haghighi
dblp:13/1739
· DBLP profile ↗
28ranked-venue papers
11as first author
5since 2021 · last 2023
0000-0002-4997-0353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 11 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | bfkNN-Embed: Locally Smoothed Embedding Mixtures for Multi-interest Candidate Retrieval
Ahmed El-Kishky, Thomas Markovich, Kenny Leung, Frank Portman, Aria Haghighi |
PAKDD (3) | 5 |
| 2023 | Learning Stance Embeddings from Signed Social GraphsabstractA challenge in social network analysis, is understanding the position, or stance, of people on a large set of topics. While past work has modeled (dis)agreement in social networks using signed graphs, these approaches have not modeled agreement patterns across a range of correlated topics. For instance, disagreement on one topic may make disagreement (or agreement) more likely for related topics. Recognizing topics influence agreement and disagreement, we propose the Stance Embeddings Model (SEM), which jointly learns embeddings for each user and topic in signed social graphs with distinct edge types for each topic. By jointly learning user and topic embeddings, SEM can perform cold-start topic stance detection, predicting the stance of a user on topics for which we have not observed their engagement. We demonstrate the effectiveness of SEM using two large-scale Twitter signed graph datasets that we open-source. One dataset, TwitterSG, labels (dis)agreements using engagements between users via tweets to derive topic-informed, signed edges. The other, BirdwatchSG, leverages community reports on misinformation and misleading content. On TwitterSG and BirdwatchSG, SEM shows a 39% and 26% error reduction respectively against strong topic-agnostic baselines. John Pougue Biyong, Aria Haghighi, Ahmed El-Kishky |
WSDM | 3 |
| 2022 | Graph-based Representation Learning for Web-scale Recommender SystemsabstractRecommender systems are fundamental building blocks of modern consumer web applications that seek to predict user preferences to better serve relevant items. As such, high-quality user and item representations as inputs to recommender systems are crucial for personalized recommendation. To construct these user and item representations, self-supervised graph embedding has emerged as a principled approach to embed relational data such as user social graphs, user membership graphs, user-item engagements, and other heterogeneous graphs. In this tutorial we discuss different families of approaches to self-supervised graph embedding. Within each family, we outline a variety of techniques, their merits and disadvantages, and expound on latest works. Finally, we demonstrate how to effectively utilize the resultant large embedding tables to improve candidate retrieval and ranking in modern industry-scale deep-learning recommender systems. Ahmed El-Kishky, Michael M. Bronstein, Aria Haghighi |
KDD | 4 |
| 2022 | TwHIN: Embedding the Twitter Heterogeneous Information Network for Personalized RecommendationabstractSocial networks, such as Twitter, form a heterogeneous information network (HIN) where nodes represent domain entities (e.g., user, content, advertiser, etc.) and edges represent one of many entity interactions (e.g, a user re-sharing content or "following" another). Interactions from multiple relation types can encode valuable information about social network entities not fully captured by a single relation; for instance, a user's preference for accounts to follow may depend on both user-content engagement interactions and the other users they follow. In this work, we investigate knowledge-graph embeddings for entities in the Twitter HIN (TwHIN); we show that these pretrained representations yield significant offline and online improvement for a diverse range of downstream recommendation and classification tasks: personalized ads rankings, account follow-recommendation, offensive content detection, and search ranking. We discuss design choices and practical challenges of deploying industry-scale HIN embeddings, including compressing them to reduce end-to-end model latency and handling parameter drift across versions. Ahmed El-Kishky, Thomas Markovich, Se Rim Park, Chetan Verma, Baekjin Kim, Ramy Eskander, Yury Malkov, Frank Portman, Sofía Samaniego, Aria Haghighi |
KDD | 11 |
| 2022 | TweetNERD - End to End Entity Linking Benchmark for TweetsabstractNamed Entity Recognition and Disambiguation (NERD) systems are foundational for information retrieval, question answering, event detection, and other natural language processing (NLP) applications. We introduce TweetNERD, a dataset of 340K+ Tweets across 2010-2021, for benchmarking NERD systems on Tweets. This is the largest and most temporally diverse open sourced dataset benchmark for NERD on Tweets and can be used to facilitate research in this area. We describe evaluation setup with TweetNERD for three NERD tasks: Named Entity Recognition (NER), Entity Linking with True Spans (EL), and End to End Entity Linking (End2End); and provide performance of existing publicly available methods on specific TweetNERD splits. TweetNERD is available at: https://doi.org/10.5281/zenodo.6617192 under Creative Commons Attribution 4.0 International (CC BY 4.0) license. Check out more details at https://github.com/twitter-research/TweetNERD. Shubhanshu Mishra, Aman Saini, Raheleh Makki, Sneha Mehta, Aria Haghighi, Ali Mollahosseini |
NeurIPS | 5 |
| 2020 | Entity Matching in the Wild: A Consistent and Versatile Framework to Unify Data in Industrial ApplicationsabstractEntity matching -- the task of clustering duplicated database records to underlying entities -- has become an increasingly critical component in modern data integration management. Amperity provides a platform for businesses to manage customer data that utilizes a machine-learning approach to entity matching, resolving billions of customer records on a daily basis. We face several challenges in deploying entity matching to industrial applications at scale, and they are less prominent in the literature. These challenges include: (1) Providing not just a single entity clustering, but supporting clusterings at multiple confidence levels to enable downstream applications with varying precision/recall trade-off needs. (2) Many customer record attributes may be systematically missing from different sources of data, creating many pairs of records in a cluster that appear to not match due to incomplete, rather than conflicting information. Allowing these records to connect transitively without introducing conflicts is invaluable to businesses because they can acquire a more comprehensive profile of their customers without incorrect entity merges. (3) How to cluster records over time and assign persistent cluster IDs that can be used for downstream use cases such as A/B tests or predictive model training; this is made more challenging by the fact that we receive new customer data every day and clusters naturally evolving over time still require persistent IDs that refer to the same entity. In this work, we describe Amperity's entity matching framework, Fusion, and how its design provides solutions to these challenges. In particular, we describe our pairwise matching model based on ordinal regression that permits a well-defined way to produce entity clusterings at different confidence levels, a novel clustering algorithm that separates conflicting record pairs in clusters while allowing for pairs that may appear dissimilar due to missing data, and a persistent ID generation algorithm which balances stability of the identifier with ever-evolving entities. Stephen Meyles, Aria Haghighi, Dan Suciu |
SIGMOD Conference | 3 |
| 2011 | Event Discovery in Social Media Feeds
Edward Benson, Aria Haghighi, Regina Barzilay |
ACL | 2 |
| 2011 | Ordering Prenominal Modifiers with a Reranking Approach
Jenny Liu, Aria Haghighi |
ACL | 2 |
| 2011 | Content Models with Attitude
Christina Sauper, Aria Haghighi, Regina Barzilay |
ACL | 2 |
| 2011 | Modeling Syntactic Context Improves Morphological Segmentation
Yoong Keok Lee, Aria Haghighi, Regina Barzilay |
CoNLL | 2 |
| 2011 | Structured Relation Discovery using Generative Models
Limin Yao, Aria Haghighi, Sebastian Riedel 0001, Andrew McCallum |
EMNLP | 2 |
| 2010 | Simple Type-Level Unsupervised POS Tagging
Yoong Keok Lee, Aria Haghighi, Regina Barzilay |
EMNLP | 2 |
| 2010 | Incorporating Content Structure into Text Analysis Applications
Christina Sauper, Aria Haghighi, Regina Barzilay |
EMNLP | 2 |
| 2010 | Coreference Resolution in a Modular, Entity-Centered Model
Aria Haghighi, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2009 | Better Word Alignments with Supervised ITG Models
Aria Haghighi, John Blitzer, John DeNero, Daniel Klein 0001 |
ACL/IJCNLP | 1 |
| 2009 | Simple Coreference Resolution with Rich Syntactic and Semantic Features
Aria Haghighi, Daniel Klein 0001 |
EMNLP | 1 |
| 2009 | Exploring Content Models for Multi-Document Summarization
Aria Haghighi, Lucy Vanderwende |
HLT-NAACL | 1 |
| 2008 | Learning Bilingual Lexicons from Monolingual Corpora
Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
ACL | 1 |
| 2008 | Coarse-to-Fine Syntactic Machine Translation using Language Projections
Slav Petrov, Aria Haghighi, Daniel Klein 0001 |
EMNLP | 2 |
| 2008 | Fully distributed EM for very large datasetsabstractIn EM and related algorithms, E-step computations distribute easily, because data items are independent given parameters. For very large data sets, however, even storing all of the parameters in a single node for the M-step can be impractical. We present a framework that fully distributes the entire EM procedure. Each node interacts only with parameters relevant to its data, sending messages to other nodes along a junction-tree topology. We demonstrate improvements over a MapReduce topology, on two tasks: word alignment and topic modeling. Jason Andrew Wolfe, Aria Haghighi, Daniel Klein 0001 |
ICML | 2 |
| 2008 | A Global Joint Model for Semantic Role LabelingabstractWe present a model for semantic role labeling that effectively captures the linguistic intuition that a semantic argument frame is a joint structure, with strong dependencies among the arguments. We show how to incorporate these strong dependencies in a statistical joint model with a rich set of features over multiple argument phrases. The proposed model substantially outperforms a similar state-of-the-art local model that does not include dependencies among different arguments. We evaluate the gains from incorporating this joint information on the Propbank corpus, when using correct syntactic parse trees as input, and when using automatically derived parse trees. The gains amount to 24.1% error reduction on all arguments and 36.8% on core arguments for gold-standard parse trees on Propbank. For automatic parse trees, the error reductions are 8.3% and 10.3% on all and core arguments, respectively. We also present results on the CoNLL 2005 shared task data set. Additionally, we explore considering multiple syntactic analyses to cope with parser noise and uncertainty. Kristina Toutanova, Aria Haghighi, Christopher D. Manning |
Comput. Linguistics | 2 |
| 2007 | A* Search via Approximate Factoring
Aria Haghighi, John DeNero, Daniel Klein 0001 |
AAAI | 1 |
| 2007 | Unsupervised Coreference Resolution in a Nonparametric Bayesian Model
Aria Haghighi, Daniel Klein 0001 |
ACL | 1 |
| 2007 | Approximate Factoring for A* Search
Aria Haghighi, John DeNero, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2006 | Prototype-Driven Grammar InductionabstractWe investigate prototype-driven learning for primarily unsupervised grammar induction. Prior knowledge is specified declaratively, by providing a few canonical examples of each target phrase type. This sparse prototype information is then propagated across a corpus using distributional similarity features, which augment an otherwise standard PCFG model. We show that distributional features are effective at distinguishing bracket labels, but not determining bracket locations. To improve the quality of the induced trees, we combine our PCFG induction with the CCM model of Klein and Manning (2002), which has complementary stengths: it identifies brackets but does not label them. Using only a handful of prototypes, we show substantial improvements over naive PCFG induction for English and Chinese grammar induction. Aria Haghighi, Daniel Klein 0001 |
ACL | 1 |
| 2006 | Prototype-Driven Learning for Sequence Models
Aria Haghighi, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2005 | Joint Learning Improves Semantic Role LabelingabstractDespite much recent progress on accurate semantic role labeling, previous work has largely used independent classifiers, possibly combined with separate label sequence models via Viterbi decoding. This stands in stark contrast to the linguistic observation that a core argument frame is a joint structure, with strong dependencies between arguments. We show how to build a joint model of argument frames, incorporating novel features that model these interactions into discriminative log-linear models. This system achieves an error reduction of 22% on all arguments and 32% on core arguments over a state-of-the art independent classifier for gold-standard parse trees on PropBank. Kristina Toutanova, Aria Haghighi, Christopher D. Manning |
ACL | 2 |
| 2005 | A Joint Model for Semantic Role Labeling
Aria Haghighi, Kristina Toutanova, Christopher D. Manning |
CoNLL | 1 |