EDBT 2026 Demo / reviewers in the wild / expert
Christian Pölitz
dblp:66/4776
· DBLP profile ↗
10ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-1690-0540ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 30% Multi-agent systems · 30% Language models and text generation · 30% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 56% Data mining · 44% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions · AAAI 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.9 | 1 | 2025 | MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQL · AAAI 2025 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.9 | 1 | 2025 | MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQL · AAAI 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions · AAAI 2025 |
Visualization and visual analytics
sensitivity analysis |
0.1 | 1 | 2012 | Identifying Place Histories from Activity Traces with an Eye to Parameter Impact · IEEE Trans. Vis. Comput. Graph. 2012 |
Visualization and visual analytics › visual analytics
spatiotemporal visual analytics |
0.1 | 1 | 2012 | Identifying Place Histories from Activity Traces with an Eye to Parameter Impact · IEEE Trans. Vis. Comput. Graph. 2012 |
Visualization and visual analytics
visual analytics |
0.1 | 1 | 2012 | Identifying Place Histories from Activity Traces with an Eye to Parameter Impact · IEEE Trans. Vis. Comput. Graph. 2012 |
Data mining › dimensionality reduction
feature selection |
0.1 | 1 | 2011 | Learning to rank under tight budget constraints · SIGIR 2011 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2011 | Learning to rank under tight budget constraints · SIGIR 2011 |
Smart cities and intelligent transportation › urban informatics
human mobility analysis |
0.0 | 1 | 2012 | Identifying Place Histories from Activity Traces with an Eye to Parameter Impact · IEEE Trans. Vis. Comput. Graph. 2012 |
Information retrieval
retrieval evaluation |
0.0 | 1 | 2011 | Learning to rank under tight budget constraints · SIGIR 2011 |
Methods — techniques the papers use, named apart from their topics
perplexity · 0.9large language model prompting · 0.9in-context learning · 0.9RLHF · 0.9statistical summarization · 0.3geocomputation · 0.3prefix selection · 0.1feature pruning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synthetic Function Demonstrations Improve Generation in Low-Resource Programming LanguagesabstractA key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python compared to Excel. Typical training data for programming languages consists of real program demonstrations coupled with explanatory human-written comments. In this work we present a novel approach to the creation of such data for low resource programming languages, which lack naturally occurring data. Our process generates synthetic, textbook-quality demonstrations of how to use library functions, which we show makes for good model finetuning data. We demonstrate in an example domain of Excel Formulas. First, we collate language documentation, then we use this to augment a powerful teacher model which generates synthetic training data, and finally finetune student models on the demonstrations. Our technique improves student performance on 2 question-answering datasets: WikiTQ and TAT-QA. We also show advantages of finetuning over standard RAG approaches, which can offer only modest improvement due to the unfamiliarity of the target domain to student models. Nick McKenna, Xinnuo Xu, Jack Williams 0001, Nicholas C. Wilson, Benjamin Van Durme, Christian Pölitz |
LREC | 6 |
| 2026 | A Teacher-Student Approach to Creating Verified Synthetic Clarification and Correction Dialogues for TableQA TasksabstractReal dialogues with AI assistants for solving data-centric tasks often follow dynamic, unpredictable paths due to imperfect information provided by the user or in the data, which must be caught and handled. Developing datasets which capture such user-AI interactions is difficult and time-consuming. In this work, we develop a novel framework for synthetically generating controlled, multi-turn conversations between a user and AI assistant for the task of table-based question answering, which can be generated from an existing dataset with fully specified table QA examples for any target domain. Each conversation aims to solve a table-based reasoning question through collaborative effort, modeling one of two real-world scenarios: (1) an AI-initiated clarification, or (2) a user-initiated correction. Critically, we employ a strong teacher LLM to verify the correctness of our synthetic conversations, ensuring high quality. We demonstrate synthetic datasets generated from TAT-QA and WikiTableQuestions as benchmarks of frontier LLMs. We find that even larger models struggle to effectively issuing clarification questions and accurately integrate user feedback for corrections. Christian Pölitz, Nick McKenna |
LREC | 1 |
| 2025 | MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQLabstractSelf-correction in text-to-SQL is the process of prompting large language model (LLM) to revise its previously incorrectly generated SQL, and commonly relies on manually crafted self-correction guidelines by human experts that are not only labor-intensive to produce but also limited by the human ability in identifying all potential error patterns in LLM responses. We introduce MAGIC, a novel multi-agent method that automates the creation of the self-correction guideline. MAGIC uses three specialized agents: a manager, a correction, and a feedback agent. These agents collaborate on the failures of an LLM-based method on the training set to iteratively generate and refine a self-correction guideline tailored to LLM mistakes, mirroring human processes but without human involvement. Our extensive experiments show that MAGIC's guideline outperforms expert human's created ones. We empirically find out that the guideline produced by MAGIC enhances the interpretability of the corrections made, providing insights in analyzing the reason behind the failures and successes of LLMs in self-correction. Arian Askari, Christian Pölitz, Xinye Tang |
AAAI | 2 |
| 2025 | Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation InstructionsabstractLLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Human Feedback), state-of-the-art LLMs like GPT4 and Llama3 are expected to have strong alignment with human preferences when prompted for a quality judgement, such as the coherence of a text. While this seems beneficial, it is not clear whether the assessments by an LLM-as-a-judge constitute only an evaluation based on the instructions in the prompts, or reflect its preference for high-quality data similar to its fine-tune data. To investigate how much influence prompting the LLMs-as-a-judge has on the alignment of AI judgements to human judgements, we analyze prompts with increasing levels of instructions about the target quality of an evaluation, for several LLMs-as-a-judge. Further, we compare to a prompt-free method using model perplexity as a quality measure instead. We aggregate a taxonomy of quality criteria commonly used across state-of-the-art evaluations with LLMs and provide this as a rigorous benchmark of models as judges. Overall, we show that the LLMs-as-a-judge benefit only little from highly detailed instructions in prompts and that perplexity can sometimes align better with human judgements than prompting, especially on textual quality. Bhuvanashree Murugadoss, Christian Pölitz, Ian Drosos, Vu Le 0002, Nick McKenna, Carina Negreanu, Chris Parnin, Advait Sarkar |
AAAI | 2 |
| 2016 | Interpretable domain adaptation via optimization over the Stiefel manifold
Christian Pölitz, Wouter Duivesteijn, Katharina Morik |
Mach. Learn. | 1 |
| 2015 | Distance Based Active Learning for Domain Adaptation
Christian Pölitz |
ICPRAM (1) | 1 |
| 2014 | Kernel Completion for Learning Consensus Support Vector Machines in Bandwidth-limited Sensor NetworksabstractAbstract: Recent developments in sensor technology allows for capturing dynamic patterns in vehicle movements, tem-perature changes, and sea-level fluctuations, just to name a few. A usual way for decision making on sensor networks, such as detecting exceptional surface level changes across the Pacific ocean, involves collecting measurement data from all sensors to build a predictor in a central processing station. However, data col-lection becomes challenging when communication bandwidth is limited, due to communication distance or low-energy requirements. Also, such settings will introduce unfavorable latency for making predictions on unseen events. In this paper, we propose an alternative strategy for such scenarios, aiming to build a consensus support vector machine (SVM) in each sensor station by exchanging a small amount of sampled information from local kernel matrices amongst peers. Our method is based on decomposing a “global ” kernel defined with all features into “local ” kernels defined only with attributes stored in each sensor station, sampling few entries of the decomposed kernel matrices that belong to other stations, and filling in unsampled entries in kernel matrices by matrix completion. Experiments on benchmark data sets illustrate that a consensus SVM can be built in each station using limited communication, which is competent in prediction performance to an SVM built with accessing all features. 1 Sangkyun Lee 0002, Christian Pölitz |
ICPRAM | 2 |
| 2012 | Identifying Place Histories from Activity Traces with an Eye to Parameter ImpactabstractEvents that happened in the past are important for understanding the ongoing processes, predicting future developments, and making informed decisions. Important and/or interesting events tend to attract many people. Some people leave traces of their attendance in the form of computer-processable data, such as records in the databases of mobile phone operators or photos on photo sharing web sites. We developed a suite of visual analytics methods for reconstructing past events from these activity traces. Our tools combine geocomputations, interactive geovisualizations, and statistical methods to enable integrated analysis of the spatial, temporal, and thematic components of the data, including numeric attributes and texts.We also support interactive investigation of the sensitivity of the analysis results to the parameters used in the computations. For this purpose, statistical summaries of computation results obtained with different combinations of parameter values are visualized in a way facilitating comparisons. We demonstrate the utility of our approach on two large real data sets, mobile phone calls in Milano during 9 days and flickr photos made on British Isles during 5 years. Gennady L. Andrienko, Natalia V. Andrienko, Martin Mladenov, Michael Mock, Christian Pölitz |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2011 | Learning to rank under tight budget constraintsabstractThis paper investigates the influence of pruning feature lists to keep a given budget for the evaluation of ranking methods. We learn from a given training set how important the individual prefixes are for the ranking quality. Based on there importance we choose the best prefixes to calculate the ranking while keeping the budget. Christian Pölitz, Ralf Schenkel |
SIGIR | 1 |
| 2010 | Extracting Events from Spatial Time SeriesabstractAn important task in exploration of data about phenomena and processes that develop over time is detection of significant changes that happened to the studied phenomenon. Our research is focused on supporting detection of significant changes, called events, in multiple time series of numeric values. We developed a suite of visual analytics techniques that combines interactive visualizations on time-aware displays and maps with statistical event detection methods implemented in R. We demonstrate the utility of our approach using two large data sets. Gennady L. Andrienko, Natalia V. Andrienko, Martin Mladenov, Michael Mock, Christian Pölitz |
IV | 5 |