EDBT 2026 Demo / reviewers in the wild / expert
Chrysoula Zerva
dblp:190/2097
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0002-4031-9492ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 73% Representation and self-supervised learning · 12% Machine translation · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.3 | 2 | 2024 | Non-Exchangeable Conformal Risk Control · ICLR 2024 Disentangling Uncertainty in Machine Translation Evaluation · EMNLP 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.8 | 1 | 2024 | Non-Exchangeable Conformal Risk Control · ICLR 2024 |
Machine learning › Trustworthy machine learning › risk control
conformal risk control |
0.8 | 1 | 2024 | Non-Exchangeable Conformal Risk Control · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | Non-Exchangeable Conformal Risk Control · ICLR 2024 |
Bioinformatics and computational biology
biomedical text mining |
0.6 | 2 | 2018 | LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway models · Bioinform. 2018 Using uncertainty to link and rank evidence from biomedical literature for model curation · Bioinform. 2017 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.6 | 1 | 2022 | Learning Disentangled Representations of Negation and Uncertainty · ACL (1) 2022 |
Natural language and speech › Machine translation
machine translation evaluation |
0.6 | 1 | 2022 | Disentangling Uncertainty in Machine Translation Evaluation · EMNLP 2022 |
Bioinformatics and computational biology › biomedical text mining
biomedical event extraction |
0.3 | 1 | 2018 | LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway models · Bioinform. 2018 |
Bioinformatics and computational biology › systems bioinformatics
pathway model curation |
0.3 | 1 | 2018 | LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway models · Bioinform. 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.2 | 1 | 2022 | Learning Disentangled Representations of Negation and Uncertainty · ACL (1) 2022 |
Visualization and visual analytics › visual analytics
visual text analytics |
0.1 | 1 | 2018 | LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway models · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
risk control · 0.8conformal prediction · 0.8text mining · 0.7semi-supervised learning · 0.7interactive visualization · 0.7variational autoencoder · 0.6mutual information minimization · 0.6monte carlo dropout · 0.6heteroscedastic regression · 0.6divergence minimization · 0.6direct uncertainty prediction · 0.6deep ensembles · 0.6adversarial learning · 0.6subjective logic theory · 0.3rule induction · 0.3machine learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Non-Exchangeable Conformal Risk ControlabstractSplit conformal prediction has recently sparked great interest due to its ability to provide formally guaranteed uncertainty sets or intervals for predictions made by black-box neural models, ensuring a predefined probability of containing the actual ground truth. While the original formulation assumes data exchangeability, some extensions handle non-exchangeable data, which is often the case in many real-world scenarios. In parallel, some progress has been made in conformal methods that provide statistical guarantees for a broader range of objectives, such as bounding the best $F_1$-score or minimizing the false negative rate in expectation. In this paper, we leverage and extend these two lines of work by proposing non-exchangeable conformal risk control, which allows controlling the expected value of any monotone loss function when the data is not exchangeable. Our framework is flexible, makes very few assumptions, and allows weighting the data based on its relevance for a given test example; a careful choice of weights may result in tighter bounds, making our framework useful in the presence of change points, time series, or other forms of distribution drift. Experiments with both synthetic and real world data show the usefulness of our method. António Farinhas, Chrysoula Zerva, Dennis Ulmer, André F. T. Martins |
ICLR | 2 |
| 2024 | Conformal Prediction for Natural Language Processing: A SurveyabstractAbstract The rapid proliferation of large language models and natural language processing (NLP) applications creates a crucial need for uncertainty quantification to mitigate risks such as Hallucinations and to enhance decision-making reliability in critical applications. Conformal prediction is emerging as a theoretically sound and practically useful framework, combining flexibility with strong statistical guarantees. Its model-agnostic and distribution-free nature makes it particularly promising to address the current shortcomings of NLP systems that stem from the absence of uncertainty quantification. This paper provides a comprehensive survey of conformal prediction techniques, their guarantees, and existing applications in NLP, pointing to directions for future research and open challenges. Margarida M. Campos, António Farinhas, Chrysoula Zerva, Mário A. T. Figueiredo, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Conformalizing Machine Translation EvaluationabstractAbstract Several uncertainty estimation methods have been recently proposed for machine translation evaluation. While these methods can provide a useful indication of when not to trust model predictions, we show in this paper that the majority of them tend to underestimate model uncertainty, and as a result, they often produce misleading confidence intervals that do not cover the ground truth. We propose as an alternative the use of conformal prediction, a distribution-free method to obtain confidence intervals with a theoretically established guarantee on coverage. First, we demonstrate that split conformal prediction can “correct” the confidence intervals of previous methods to yield a desired coverage level, and we demonstrate these findings across multiple machine translation evaluation metrics and uncertainty quantification methods. Further, we highlight biases in estimated confidence intervals, reflected in imbalanced coverage for different attributes, such as the language and the quality of translations. We address this by applying conditional conformal prediction techniques to obtain calibration subsets for each data subgroup, leading to equalized coverage. Overall, we show that, provided access to a calibration set, conformal prediction can help identify the most suitable uncertainty quantification methods and adapt the predicted confidence intervals to ensure fairness with respect to different attributes.1 Chrysoula Zerva, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | BLEU Meets COMET: Combining Lexical and Neural Metrics Towards Robust Machine Translation EvaluationabstractAlthough neural-based machine translation evaluation metrics, such as COMET or BLEURT, have achieved strong correlations with human judgements, they are sometimes unreliable in detecting certain phenomena that can be considered as critical errors, such as deviations in entities and numbers. In contrast, traditional evaluation metrics such as BLEU or chrF, which measure lexical or character overlap between translation hypotheses and human references, have lower correlations with human judgements but are sensitive to such deviations. In this paper, we investigate several ways of combining the two approaches in order to increase robustness of state-of-the-art evaluation methods to translations with critical errors. We show that by using additional information during training, such as sentence-level features and word-level tags, the trained metrics improve their capability to penalize translations with specific troublesome phenomena, which leads to gains in correlations with humans and on the recent DEMETR benchmark on several language pairs. Taisiya Glushkova, Chrysoula Zerva, André F. T. Martins |
EAMT | 2 |
| 2023 | Context-aware Neural Machine Translation for English-Japanese Business Scene DialoguesabstractDespite the remarkable advancements in machine translation, the current sentence-level paradigm faces challenges when dealing with highly-contextual languages like Japanese. In this paper, we explore how context-awareness can improve the performance of the current Neural Machine Translation (NMT) models for English-Japanese business dialogues translation, and what kind of context provides meaningful information to improve translation. As business dialogue involves complex discourse phenomena but offers scarce training resources, we adapted a pretrained mBART model, finetuning on multi-sentence dialogue data, which allows us to experiment with different contexts. We investigate the impact of larger context sizes and propose novel context tokens encoding extra-sentential information, such as speaker turn and scene type. We make use of Conditional Cross-Mutual Information (CXMI) to explore how much of the context the model uses and generalise CXMI to study the impact of the extra sentential context. Overall, we find that models leverage both preceding sentences and extra-sentential context (with CXMI increasing with context size) and we provide a more focused analysis on honorifics translation. Regarding translation quality, increased source-side context paired with scene and speaker information improves the model performance compared to previous work and our context-agnostic baselines, measured in BLEU and COMET metrics. Sumire Honda, Patrick Fernandes, Chrysoula Zerva |
MTSummit (1) | 3 |
| 2022 | Learning Disentangled Representations of Negation and UncertaintyabstractNegation and uncertainty modeling are long-standing tasks in natural language processing. Linguistic theory postulates that expressions of negation and uncertainty are semantically independent from each other and the content they modify. However, previous works on representation learning do not explicitly model this independence. We therefore attempt to disentangle the representations of negation, uncertainty, and content using a Variational Autoencoder. We find that simply supervising the latent representations results in good disentanglement, but auxiliary objectives based on adversarial learning and mutual information minimization can provide additional disentanglement gains. Jake Vasilakes, Chrysoula Zerva, Makoto Miwa, Sophia Ananiadou |
ACL (1) | 2 |
| 2022 | DeepSPIN: Deep Structured Prediction for Natural Language ProcessingabstractDeepSPIN is a research project funded by the European Research Council (ERC) whose goal is to develop new neural structured prediction methods, models, and algorithms for improving the quality, interpretability, and data-efficiency of natural language processing (NLP) systems, with special emphasis on machine translation and quality estimation. We describe in this paper the latest findings from this project. André F. T. Martins, Ben Peters, Chrysoula Zerva, Chunchuan Lyu, Gonçalo M. Correia, Marcos V. Treviso, Pedro Henrique Martins, Tsvetomila Mihaylova |
EAMT | 3 |
| 2022 | Disentangling Uncertainty in Machine Translation EvaluationabstractTrainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data.Recent work has attempted to mitigate this with simple uncertainty quantification techniques (Monte Carlo dropout and deep ensembles), however these techniques (as we show) are limited in several ways -for example, they are unable to distinguish between different kinds of uncertainty, and they are time and memory consuming.In this paper, we propose more powerful and efficient uncertainty predictors for MT evaluation, and we assess their ability to target different sources of aleatoric and epistemic uncertainty.To this end, we develop and compare training objectives for the COMET metric to enhance it with an uncertainty prediction output, including heteroscedastic regression, divergence minimization, and direct uncertainty prediction.Our experiments show improved results on uncertainty prediction for the WMT metrics task datasets, with a substantial reduction in computational costs.Moreover, they demonstrate the ability of these predictors to address specific uncertainty causes in MT evaluation, such as low quality references and outof-domain data. 1 Chrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. Martins |
EMNLP | 1 |
| 2022 | MLQE-PE: A Multilingual Quality Estimation and Post-Editing DatasetabstractWe present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains annotations for eleven language pairs, including both high- and low-resource languages. Specifically, it is annotated for translation quality with human labels for up to 10,000 translations per language pair in the following formats: sentence-level direct assessments and post-editing effort, and word-level binary good/bad labels. Apart from the quality-related scores, each source-translation sentence pair is accompanied by the corresponding post-edited sentence, as well as titles of the articles where the sentences were extracted from, and information on the neural MT models used to translate the text. We provide a thorough description of the data collection and annotation process as well as an analysis of the annotation distribution for each language pair. We also report the performance of baseline systems trained on the MLQE-PE dataset. The dataset is freely available and has already been used for several WMT shared tasks. Marina Fomicheva, Erick Rocha Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, André F. T. Martins |
LREC | 4 |
| 2018 | LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway modelsabstractMotivation: Pathway models are valuable resources that help us understand the various mechanisms underpinning complex biological processes. Their curation is typically carried out through manual inspection of published scientific literature to find information relevant to a model, which is a laborious and knowledge-intensive task. Furthermore, models curated manually cannot be easily updated and maintained with new evidence extracted from the literature without automated support. Results: We have developed LitPathExplorer, a visual text analytics tool that integrates advanced text mining, semi-supervised learning and interactive visualization, to facilitate the exploration and analysis of pathway models using statements (i.e. events) extracted automatically from the literature and organized according to levels of confidence. LitPathExplorer supports pathway modellers and curators alike by: (i) extracting events from the literature that corroborate existing models with evidence; (ii) discovering new events which can update models; and (iii) providing a confidence value for each event that is automatically computed based on linguistic features and article metadata. Our evaluation of event extraction showed a precision of 89% and a recall of 71%. Evaluation of our confidence measure, when used for ranking sampled events, showed an average precision ranging between 61 and 73%, which can be improved to 95% when the user is involved in the semi-supervised learning process. Qualitative evaluation using pair analytics based on the feedback of three domain experts confirmed the utility of our tool within the context of pathway model exploration. Availability and implementation: LitPathExplorer is available at http://nactem.ac.uk/LitPathExplorer_BI/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Axel J. Soto, Chrysoula Zerva, Riza Theresa Batista-Navarro, Sophia Ananiadou |
Bioinform. | 2 |
| 2017 | Using uncertainty to link and rank evidence from biomedical literature for model curationabstractMOTIVATION: In recent years, there has been great progress in the field of automated curation of biomedical networks and models, aided by text mining methods that provide evidence from literature. Such methods must not only extract snippets of text that relate to model interactions, but also be able to contextualize the evidence and provide additional confidence scores for the interaction in question. Although various approaches calculating confidence scores have focused primarily on the quality of the extracted information, there has been little work on exploring the textual uncertainty conveyed by the author. Despite textual uncertainty being acknowledged in biomedical text mining as an attribute of text mined interactions (events), it is significantly understudied as a means of providing a confidence measure for interactions in pathways or other biomedical models. In this work, we focus on improving identification of textual uncertainty for events and explore how it can be used as an additional measure of confidence for biomedical models. RESULTS: We present a novel method for extracting uncertainty from the literature using a hybrid approach that combines rule induction and machine learning. Variations of this hybrid approach are then discussed, alongside their advantages and disadvantages. We use subjective logic theory to combine multiple uncertainty values extracted from different sources for the same interaction. Our approach achieves F-scores of 0.76 and 0.88 based on the BioNLP-ST and Genia-MK corpora, respectively, making considerable improvements over previously published work. Moreover, we evaluate our proposed system on pathways related to two different areas, namely leukemia and melanoma cancer research. AVAILABILITY AND IMPLEMENTATION: The leukemia pathway model used is available in Pathway Studio while the Ras model is available via PathwayCommons. Online demonstration of the uncertainty extraction system is available for research purposes at http://argo.nactem.ac.uk/test. The related code is available on https://github.com/c-zrv/uncertainty_components.git. Details on the above are available in the Supplementary Material. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chrysoula Zerva, Riza Theresa Batista-Navarro, Philip Day, Sophia Ananiadou |
Bioinform. | 1 |