Katherine Lin

dblp:187/9727 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ENCO: Deploying Production-Scale Engineering Copilots
abstract
Software engineers frequently grapple with the challenge of accessing fragmented documentation and telemetry data, such as Troubleshooting Guides (TSGs), incident reports, code repositories, and internal tools maintained by different teams. In this work, we introduced ENCO, a comprehensive framework for developing, deploying, and managing copilots tailored to improve productivity in large scale production scenarios. The framework combines an innovative NL2SearchQuery module with a lightweight hierarchical agentic planner to enable accurate and efficient retrieval-augmented generation (RAG) for code, semi-structured data and documents. These components allow the copilot to retrieve relevant information from diverse sources and invoke the right skills with low latency to answer highly complex technical questions. Since its launch in September 2023, ENCO has demonstrated its effectiveness through widespread adoption, enabling tens of thousands of interactions and engaging over 1,000 monthly active users (MAUs). The system has been continuously optimized based on usage patterns and user feedback, resulting in measurable improvements in response relevance, latency, and user satisfaction.
Mathieu B. Demarne, Wenjing Wang 0005, Nutan Sahoo, Hannah Lerner, Anjali Bhavan, Divya Vermareddy, Yunlei Lu, Swati Bararia, William Zhang 0001, Katherine Lin, Miso Cilimdzic, Subru Krishnan
KDD (1)13
2025 Crafting a Personal Journaling Practice: Negotiating Ecosystems of Materials, Personal Context, and Community in Analog Journaling
abstract
Analog journaling has grown in popularity, with journaling on paper encompassing a range of motivations, styles, and practices including planning, habit-tracking, and reflecting.Journalers develop strong personal preferences around the tools they use, the ideas they capture, and the layout in which they represent their ideas and memories.Understanding how analog journaling practices are individually shaped and crafted over time is critical to supporting the varied benefits associated with journaling, including improved mental health and positive support for identity development.To understand this development, we qualitatively analyzed publicly-shared journaling content from YouTube and Instagram and interviewed 11 journalers.We report on our identification of the journaling ecosystem in which journaling practices are shaped by materials, personal context, and communities, sharing how this ecosystem plays a role in the practices and identities of journalers as they customize their journaling routine to best suit their personal goals.Using these insights, we discuss design opportunities for how future tools can better align with and reflect the rich affordances and practices of journaling on paper.
Katherine Lin, Juna Kawai-Yue, Adira Sklar, Lucy Hecht, Sarah Sterman, Tiffany Tseng
Creativity & Cognition1
2025 FLAIR: Feedback Learning for Adaptive Information Retrieval
abstract
Recent advances in Large Language Models (LLMs) have driven the adoption of copilots in complex technical scenarios, underscoring the growing need for specialized information retrieval solutions. In this paper, we introduce FLAIR, a lightweight, feedback learning framework that adapts copilot systems' retrieval strategies by integrating domain-specific expert feedback. FLAIR operates in two stages: an offline phase obtains indicators from (1) user feedback and (2) questions synthesized from documentation, storing these indicators in a decentralized manner. An online phase then employs a two-track ranking mechanism to combine raw similarity scores with the collected indicators. This iterative setup refines retrieval performance for any query. Extensive real-world evaluations of FLAIR demonstrate significant performance gains on both previously seen and unseen queries, surpassing state-of-the-art approaches. The system has been successfully integrated into Copilot DECO, serving thousands of users at Microsoft, demonstrating its scalability and effectiveness in operational environments.
William Zhang 0001, Yunlei Lu, Mathieu B. Demarne, Wenjing Wang 0005, Nutan Sahoo, Katherine Lin, Miso Cilimdzic, Subru Krishnan
CIKM8
2025 Horizon: Robust Checks for SQL Migration Using LLMs
abstract
Large language models (LLMs) have recently demonstrated strong capabilities in code migration across languages, making them promising for SQL schema migration. However, achieving reliable and accurate SQL migration with LLMs remains a challenge. This paper presents the first comprehensive approach for practical and effective SQL schema migration using LLMs. We highlight the necessity of robust evaluation and iterative query refinement to achieve highly accurate migrations. Building on traditional database tools along with LLMs, we introduce novel checks to guide LLMs towards syntactically complete and functionally equivalent translations. Our approach supports all schema object types, including complex procedural constructs. Our demonstrations offer audience opportunities to explore our system using a variety of configurations, datasets and custom inputs, providing useful insights into the underlying techniques, their strengths, and limitations.
K. Venkatesh Emani, Wenjing Wang 0005, Neel Ball, Kumaraswamy Boora, Carlo Curino, Avrilia Floratou, Manan Goenka, Paridhi Gupta, Katherine Lin, Nick Litombe, Jared Meade, Suryakant Mutnal, Raghu Ramakrishnan 0001, Sudhir Raparla, Dhruv Relwani, Shyam Sai, Vaibhave Sekar, Roneet Shaw, Harmeet Singh, Prasanna Sridharan, Sunidhi Tiwari
Proc. VLDB Endow.12
2024 NL2SQL is a solved problem... Not!
Avrilia Floratou, Fotis Psallidas, Fuheng Zhao, Shaleen Deep, Gunther Hagleither, Wangda Tan, Joyce Cahoon, Rana Alotaibi, Jordan Henkel, Abhik Singla, Alex Van Grootel, Brandon Chow, Katherine Lin, Marcos Campos, K. Venkatesh Emani, Vivek Pandit, Victor Shnayder, Wenjing Wang 0005, Carlo Curino
CIDR14
2024 NL2Code-Reasoning and Planning with LLMs for Code Development
abstract
There is huge value in making software development more productive with AI. An important component of this vision is the capability to translate natural language to a programming language ("NL2Code") and thus to significantly accelerate the speed at which code is written.
Ye Xing, Jun Huan, Wee Hyong Tok, Cong Shen 0001, Johannes Gehrke, Katherine Lin, Arjun Guha, Omer Tripp, Murali Krishna Ramanathan
KDD6
2023 Stitcher: Learned Workload Synthesis from Historical Performance Footprints
Chengcheng Wan 0001, Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Alexandra M. Ciortea, Konstantinos Karanasos, Subru Krishnan
EDBT5
2022 Doppler: Automated SKU Recommendation in Migrating SQL Workloads to the Cloud
abstract
Selecting the optimal cloud target to migrate SQL estates from on-premises to the cloud remains a challenge. Current solutions are not only time-consuming and error-prone, requiring significant user input, but also fail to provide appropriate recommendations. We present Doppler, a scalable recommendation engine that provides right-sized Azure SQL Platform-as-a-Service (PaaS) recommendations without requiring access to sensitive customer data and queries. Doppler introduces a novel price-performance methodology that allows customers to get a personalized rank of relevant cloud targets solely based on low-level resource statistics, such as latency and memory usage. Doppler supplements this rank with internal knowledge of Azure customer behavior to help guide new migration customers towards one optimal target. Experimental results over a 9-month period from prospective and existing customers indicate that Doppler can identify optimal targets and adapt to changes in customer workloads. It has also found cost-saving opportunities among over-provisioned cloud customers, without compromising on capacity or other requirements. Doppler has been integrated and released in the Azure Data Migration Assistant v5.5, which receives hundreds of assessment requests daily.
Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Chengcheng Wan 0001, Alexandra M. Ciortea, Sreraman Narasimhan, Subru Krishnan
Proc. VLDB Endow.4
2021 Toto - Benchmarking the Efficiency of a Cloud Service
abstract
Microsoft aims to increase the efficiency of Azure SQL DB by maximizing the number of databases that can be hosted in a cluster. However, resource contention among customers increases when changing the configurations, policies, and features that control database co-location on cluster nodes. Tuning and evaluating the efficiency and customer impact of these variables in a scientific manner in production, with a dynamic system and customer workloads, is difficult or infeasible. Here, we present Toto, a benchmark framework for evaluating the efficiency of any cloud service that leverages orchestrators like Service Fabric or Kubernetes. Toto allows for reliable and repeatable specification of a benchmarking scenario of arbitrary scale, complexity, and time-length. An implementation of Toto is deployed in all SQL DB staging clusters and is used to evaluate system efficiency and behaviors. As an example of Toto's capabilities, we present a study to explore the balance between cluster database density and quality of service.
Justin Moeller, Katherine Lin, Willis Lang
SIGMOD Conference3
2016 Habitsourcing: Sensing the Environment through Immersive, Habit-Building Experiences
abstract
Citizen science and communitysensing applications allow everyday citizens to collect data about the physical world to benefit science and society. Yet despite successes, current approaches are still limited by the number of domain-interested volunteers who are willing and able to contribute useful data. In this paper we introduce habitsourcing, an alternative approach that harnesses the habit-building practices of millions of people to collect environmental data. To support the design and development of habitsourcing apps, we present (1) interaction techniques and design principles for sensing through actuation, a method for acquiring sensing data from cued interactions; and (2) ExperienceKit, an iOS library that makes it easy for developers to build and test habitsourcing applications. In two experiments, we show that our two proof-of-concept apps, ZenWalk and Zombies Interactive, compare favorably to their non-data collecting counterparts, and that we can effectively extract environmental data using simple detection techniques.
Katherine Lin, Henry Spindell, Scott Allen Cambo, Yongsung Kim
UIST1