EDBT 2026 Demo / reviewers in the wild / expert
Katherine Lin
dblp:187/9727
· DBLP profile ↗
8ranked-venue papers in the field
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 5Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ENCO: Deploying Production-Scale Engineering CopilotsabstractSoftware engineers frequently grapple with the challenge of accessing fragmented documentation and telemetry data, such as Troubleshooting Guides (TSGs), incident reports, code repositories, and internal tools maintained by different teams. In this work, we introduced ENCO, a comprehensive framework for developing, deploying, and managing copilots tailored to improve productivity in large scale production scenarios. The framework combines an innovative NL2SearchQuery module with a lightweight hierarchical agentic planner to enable accurate and efficient retrieval-augmented generation (RAG) for code, semi-structured data and documents. These components allow the copilot to retrieve relevant information from diverse sources and invoke the right skills with low latency to answer highly complex technical questions. Since its launch in September 2023, ENCO has demonstrated its effectiveness through widespread adoption, enabling tens of thousands of interactions and engaging over 1,000 monthly active users (MAUs). The system has been continuously optimized based on usage patterns and user feedback, resulting in measurable improvements in response relevance, latency, and user satisfaction. Mathieu B. Demarne, Wenjing Wang 0005, Nutan Sahoo, Hannah Lerner, Anjali Bhavan, Divya Vermareddy, Yunlei Lu, Swati Bararia, William Zhang 0001, Katherine Lin, Miso Cilimdzic, Subru Krishnan |
KDD (1) | 13 |
| 2025 | FLAIR: Feedback Learning for Adaptive Information RetrievalabstractRecent advances in Large Language Models (LLMs) have driven the adoption of copilots in complex technical scenarios, underscoring the growing need for specialized information retrieval solutions. In this paper, we introduce FLAIR, a lightweight, feedback learning framework that adapts copilot systems' retrieval strategies by integrating domain-specific expert feedback. FLAIR operates in two stages: an offline phase obtains indicators from (1) user feedback and (2) questions synthesized from documentation, storing these indicators in a decentralized manner. An online phase then employs a two-track ranking mechanism to combine raw similarity scores with the collected indicators. This iterative setup refines retrieval performance for any query. Extensive real-world evaluations of FLAIR demonstrate significant performance gains on both previously seen and unseen queries, surpassing state-of-the-art approaches. The system has been successfully integrated into Copilot DECO, serving thousands of users at Microsoft, demonstrating its scalability and effectiveness in operational environments. William Zhang 0001, Yunlei Lu, Mathieu B. Demarne, Wenjing Wang 0005, Nutan Sahoo, Katherine Lin, Miso Cilimdzic, Subru Krishnan |
CIKM | 8 |
| 2025 | Horizon: Robust Checks for SQL Migration Using LLMsabstractLarge language models (LLMs) have recently demonstrated strong capabilities in code migration across languages, making them promising for SQL schema migration. However, achieving reliable and accurate SQL migration with LLMs remains a challenge. This paper presents the first comprehensive approach for practical and effective SQL schema migration using LLMs. We highlight the necessity of robust evaluation and iterative query refinement to achieve highly accurate migrations. Building on traditional database tools along with LLMs, we introduce novel checks to guide LLMs towards syntactically complete and functionally equivalent translations. Our approach supports all schema object types, including complex procedural constructs. Our demonstrations offer audience opportunities to explore our system using a variety of configurations, datasets and custom inputs, providing useful insights into the underlying techniques, their strengths, and limitations. K. Venkatesh Emani, Wenjing Wang 0005, Neel Ball, Kumaraswamy Boora, Carlo Curino, Avrilia Floratou, Manan Goenka, Paridhi Gupta, Katherine Lin, Nick Litombe, Jared Meade, Suryakant Mutnal, Raghu Ramakrishnan 0001, Sudhir Raparla, Dhruv Relwani, Shyam Sai, Vaibhave Sekar, Roneet Shaw, Harmeet Singh, Prasanna Sridharan, Sunidhi Tiwari |
Proc. VLDB Endow. | 12 |
| 2024 | NL2SQL is a solved problem... Not!
Avrilia Floratou, Fotis Psallidas, Fuheng Zhao, Shaleen Deep, Gunther Hagleither, Wangda Tan, Joyce Cahoon, Rana Alotaibi, Jordan Henkel, Abhik Singla, Alex Van Grootel, Brandon Chow, Katherine Lin, Marcos Campos, K. Venkatesh Emani, Vivek Pandit, Victor Shnayder, Wenjing Wang 0005, Carlo Curino |
CIDR | 14 |
| 2024 | NL2Code-Reasoning and Planning with LLMs for Code DevelopmentabstractThere is huge value in making software development more productive with AI. An important component of this vision is the capability to translate natural language to a programming language ("NL2Code") and thus to significantly accelerate the speed at which code is written. Ye Xing, Jun Huan, Wee Hyong Tok, Cong Shen 0001, Johannes Gehrke, Katherine Lin, Arjun Guha, Omer Tripp, Murali Krishna Ramanathan |
KDD | 6 |
| 2023 | Stitcher: Learned Workload Synthesis from Historical Performance Footprints
Chengcheng Wan 0001, Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Alexandra M. Ciortea, Konstantinos Karanasos, Subru Krishnan |
EDBT | 5 |
| 2022 | Doppler: Automated SKU Recommendation in Migrating SQL Workloads to the CloudabstractSelecting the optimal cloud target to migrate SQL estates from on-premises to the cloud remains a challenge. Current solutions are not only time-consuming and error-prone, requiring significant user input, but also fail to provide appropriate recommendations. We present Doppler, a scalable recommendation engine that provides right-sized Azure SQL Platform-as-a-Service (PaaS) recommendations without requiring access to sensitive customer data and queries. Doppler introduces a novel price-performance methodology that allows customers to get a personalized rank of relevant cloud targets solely based on low-level resource statistics, such as latency and memory usage. Doppler supplements this rank with internal knowledge of Azure customer behavior to help guide new migration customers towards one optimal target. Experimental results over a 9-month period from prospective and existing customers indicate that Doppler can identify optimal targets and adapt to changes in customer workloads. It has also found cost-saving opportunities among over-provisioned cloud customers, without compromising on capacity or other requirements. Doppler has been integrated and released in the Azure Data Migration Assistant v5.5, which receives hundreds of assessment requests daily. Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Chengcheng Wan 0001, Alexandra M. Ciortea, Sreraman Narasimhan, Subru Krishnan |
Proc. VLDB Endow. | 4 |
| 2021 | Toto - Benchmarking the Efficiency of a Cloud ServiceabstractMicrosoft aims to increase the efficiency of Azure SQL DB by maximizing the number of databases that can be hosted in a cluster. However, resource contention among customers increases when changing the configurations, policies, and features that control database co-location on cluster nodes. Tuning and evaluating the efficiency and customer impact of these variables in a scientific manner in production, with a dynamic system and customer workloads, is difficult or infeasible. Here, we present Toto, a benchmark framework for evaluating the efficiency of any cloud service that leverages orchestrators like Service Fabric or Kubernetes. Toto allows for reliable and repeatable specification of a benchmarking scenario of arbitrary scale, complexity, and time-length. An implementation of Toto is deployed in all SQL DB staging clusters and is used to evaluate system efficiency and behaviors. As an example of Toto's capabilities, we present a study to explore the balance between cluster database density and quality of service. Justin Moeller, Katherine Lin, Willis Lang |
SIGMOD Conference | 3 |