Megan Leszczynski

dblp:217/1811 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-8065-7763ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2023 Beyond Single Items: Exploring User Preferences in Item Sets with the Conversational Playlist Curation Dataset
abstract
Users in consumption domains, like music, are often able to more efficiently provide preferences over a set of items (e.g. a playlist or radio) than over single items (e.g. songs). Unfortunately, this is an underexplored area of research, with most existing recommendation systems limited to understanding preferences over single items. Curating an item set exponentiates the search space that recommender systems must consider (all subsets of items!): this motivates conversational approaches-where users explicitly state or refine their preferences and systems elicit preferences in natural language-as an efficient way to understand user needs. We call this task conversational item set curation and present a novel data collection methodology that efficiently collects realistic preferences about item sets in a conversational setting by observing both item-level and set-level feedback. We apply this methodology to music recommendation to build the Conversational Playlist Curation Dataset (CPCD), where we show that it leads raters to express preferences that would not be otherwise expressed. Finally, we propose a wide range of conversational retrieval models as baselines for this task and evaluate them on the dataset.
Arun Tejasvi Chaganty, Megan Leszczynski, Ravi Ganti, Krisztian Balog, Filip Radlinski
SIGIR2
2021 Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Laurel J. Orr, Megan Leszczynski, Neel Guha, Sen Wu 0002, Simran Arora, Christopher Ré
CIDR2
2021 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems
abstract
The industrial machine learning pipeline requires iterating on model features, training and deploying models, and monitoring deployed models at scale. Feature stores were developed to manage and standardize the engineer's workflow in this end-to-end pipeline, focusing on traditional tabular feature data. In recent years, however, model development has shifted towards using self-supervised pretrained embeddings as model features. Managing these embeddings and the downstream systems that use them introduces new challenges with respect to managing embedding training data, measuring embedding quality, and monitoring downstream models that use embeddings. These challenges are largely unaddressed in standard feature stores. Our goal in this tutorial is to introduce the feature store system and discuss the challenges and current solutions to managing these new embedding-centric pipelines.
Laurel J. Orr, Atindriyo Sanyal, Karan Goel, Megan Leszczynski
Proc. VLDB Endow.5
2020 Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
Tri Dao, Nimit Sharad Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra, Christopher Ré
ICLR6