Kamile Stankeviciute

dblp:319/4868 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-2489-9615ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 48% Language models and text generation · 33% Time series and sequential data · 19%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 87% Machine learning and data management · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Medical and health informatics
electronic health records
0.912025
MEDS: Building Models and Tools in a Reproducible Health AI Ecosystem · KDD (2) 2025
Information retrieval › evaluation › benchmark
benchmark construction
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Information retrieval
retrieval evaluation
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.512021
Conformal Time-series Forecasting · NeurIPS 2021
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.512021
Conformal Time-series Forecasting · NeurIPS 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Conformal Time-series Forecasting · NeurIPS 2021
Machine learning › Trustworthy machine learning
data leakage
0.312025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025

Methods — techniques the papers use, named apart from their topics

question-answer pair generation · 1.7document corpus generation · 1.7data standardization · 1.7benchmarking · 1.7recurrent neural network · 0.5inductive conformal prediction · 0.5
YearPublicationVenuePosition
2025 PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
abstract
High-quality benchmarks are essential for evaluating reasoning and retrieval capabilities of large language models (LLMs). However, curating datasets for this purpose is not a permanent solution as they are prone to data leakage and inflated performance results. To address these challenges, we propose PhantomWiki: a pipeline to generate unique, factually consistent document corpora with diverse question-answer pairs. Unlike prior work, PhantomWiki is neither a fixed dataset, nor is it based on any existing data. Instead, a new PhantomWiki instance is generated on demand for each evaluation. We vary the question difficulty and corpus size to disentangle reasoning and retrieval capabilities, respectively, and find that PhantomWiki datasets are surprisingly challenging for frontier LLMs. Thus, we contribute a scalable and data leakage-resistant framework for disentangled evaluation of reasoning, retrieval, and tool-use abilities.
Albert Gong, Kamile Stankeviciute, Chao Wan, Anmol Kabra, Raphael Thesmar, Johann Lee, Julius Klenke, Carla P. Gomes, Kilian Q. Weinberger
ICML2
2025 MEDS: Building Models and Tools in a Reproducible Health AI Ecosystem
abstract
Health AI suffers from a systemic reproducibility crisis that irreparably hinders research across both academia and industry [4,5].One key tool poised to solve this crisis is the Medical Event Data Standard (MEDS), a comprehensive data format and open-source ecosystem designed to enhance reproducibility and interoperability of AI research using longitudinal Electronic Health Records (EHR) [6].Currently adopted by over 15 institutions globally, MEDS encompasses various open-source tools, published models, and data processing pipelines, enabling streamlined model development and robust benchmarking.In this tutorial, participants will gain key hands-on experience in working with the MEDS format to perform efficient, reproducible, state-of-the-art AI research over real health data.Participants will transform data into the MEDS format, preprocess data, build predictive models, and contribute to the decentralized MEDS-DEV benchmarking platform.Interactive exercises using Jupyter notebooks will provide hands-on experience and practical skills for reproducible health AI research.Attendees will leave equipped
Matthew B. A. McDermott, Justin Xu, Teya S. Bergamaschi, Hyewon Jeong, Simon A. Lee, Nassim Oufattole, Patrick Rockenschaub, Kamile Stankeviciute, Ethan Steinberg, Jimeng Sun 0001, Robin Van De Water, Michael Wornow, John Wu, Zhenbang Wu
KDD (2)8
2021 Conformal Time-series Forecasting
abstract
Current approaches for multi-horizon time series forecasting using recurrent neural networks (RNNs) focus on issuing point estimates, which is insufficient for decision-making in critical application domains where an uncertainty estimate is also required. Existing approaches for uncertainty quantification in RNN-based time-series forecasts are limited as they may require significant alterations to the underlying model architecture, may be computationally complex, may be difficult to calibrate, may incur high sample complexity, and may not provide theoretical guarantees on frequentist coverage. In this paper, we extend the inductive conformal prediction framework to the time-series forecasting setup, and propose a lightweight algorithm to address all of the above limitations, providing uncertainty estimates with theoretical guarantees for any multi-horizon forecast predictor and any dataset with minimal exchangeability assumptions. We demonstrate the effectiveness of our approach by comparing it with existing benchmarks on a variety of synthetic and real-world datasets.
Kamile Stankeviciute, Ahmed Alaa 0001, Mihaela van der Schaar
NeurIPS1