VLDB 2026 Research / reviewers in the wild / expert
Lina Maria Rojas-Barahona
dblp:80/7169
· DBLP profile ↗
33ranked-venue papers
7as first author
10since 2021 · last 2026
0009-0009-8439-4695ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Detection of Synthetic Tabular Data Under Schema VariabilityabstractThe rise of powerful generative models has sparked concerns over data authenticity. While detection methods have been extensively developed for images and text, the case of tabular data, despite its ubiquity, has been largely overlooked. Yet, detecting synthetic tabular data is especially challenging due to its heterogeneous structure and unseen formats at test time. We address the underexplored task of detecting synthetic tabular data "in the wild", i.e. when the detector is deployed on tables with variable and previously unseen schemas. We introduce a novel datum-wise transformer architecture that significantly outperforms the only previously published baseline, improving both AUC and accuracy by 7 points. By incorporating a table-adaptation component, our model gains an additional 7 accuracy points, demonstrating enhanced robustness. This work provides the first strong evidence that detecting synthetic tabular data in real-world conditions is feasible, and demonstrates substantial improvements over previous approaches. The code will be made available in the extended version. G. Charbel N. Kindji, Élisa Fromont, Lina Maria Rojas-Barahona, Tanguy Urvoy |
AAAI | 3 |
| 2026 | Entrainment detection using DNN
Jay Kejriwal, Stefan Benus, Lina Maria Rojas-Barahona |
Comput. Speech Lang. | 3 |
| 2025 | Synthetic Tabular Data Detection in the Wild
G. Charbel N. Kindji, Élisa Fromont, Lina Maria Rojas-Barahona, Tanguy Urvoy |
IDA | 3 |
| 2025 | Tabular data generation models: An in-depth survey and performance benchmarks with extensive tuning
G. Charbel N. Kindji, Lina Maria Rojas-Barahona, Élisa Fromont, Tanguy Urvoy |
Neurocomputing | 2 |
| 2024 | KGConv, a Conversational Corpus Grounded in WikidataabstractWe present KGConv, a large corpus of 71k English conversations where each question-answer pair is grounded in a Wikidata fact. Conversations contain on average 8.6 questions and for each Wikidata fact, we provide multiple variants (12 on average) of the corresponding question using templates, human annotations, hand-crafted rules and a question rewriting neural model. We provide baselines for the task of Knowledge-Based, Conversational Question Generation. KGConv can further be used for other generation and analysis tasks such as single-turn question generation from Wikidata triples, question rewriting, question answering from conversation or from knowledge graphs and quiz generation. Quentin Brabant, Lina Maria Rojas-Barahona, Gwénolé Lecorvé, Claire Gardent |
LREC/COLING | 2 |
| 2023 | Investigating the Effect of Relative Positional Embeddings on AMR-to-Text Generation with Structural AdaptersabstractSebastien Montella, Alexis Nasr, Johannes Heinecke, Frederic Bechet, Lina M. Rojas Barahona. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Sébastien Montella, Alexis Nasr, Johannes Heinecke, Frédéric Béchet, Lina Maria Rojas-Barahona |
EACL | 5 |
| 2023 | Unsupervised Auditory and Semantic Entrainment Models with Deep Neural NetworksabstractSpeakers tend to engage in adaptive behavior, known as entrainment, when they become similar to their interlocutor in various aspects of speaking. We present an unsupervised deep learning framework that derives meaningful representation from textual features for developing semantic entrainment. We investigate the model's performance by extracting features using different variations of the BERT model (DistilBERT and XLM-RoBERTa) and Google's universal sentence encoder (USE) embeddings on two human-human (HH) corpora (The Fisher Corpus English Part 1, Columbia games corpus) and one human-machine (HM) corpus (Voice Assistant Conversation Corpus (VACC)). In addition to semantic features we also trained DNN-based models utilizing two auditory embeddings (TRIpLet Loss network (TRILL) vectors, Low-level descriptors (LLD) features) and two units of analysis (Inter pausal unit and Turn). The results show that semantic entrainment can be assessed with our model, that models can distinguish between HH and HM interactions and that the two units of analysis for extracting acoustic features provide comparable findings. Jay Kejriwal, Stefan Benus, Lina Maria Rojas-Barahona |
INTERSPEECH | 3 |
| 2022 | CoQAR: Question Rewriting on CoQAabstractQuestions asked by humans during a conversation often contain contextual dependencies, i.e., explicit or implicit references to previous dialogue turns. These dependencies take the form of coreferences (e.g., via pronoun use) or ellipses, and can make the understanding difficult for automated systems. One way to facilitate the understanding and subsequent treatments of a question is to rewrite it into an out-of-context form, i.e., a form that can be understood without the conversational context. We propose CoQAR, a corpus containing 4.5K conversations from the Conversational Question-Answering dataset CoQA, for a total of 53K follow-up question-answer pairs. Each original question was manually annotated with at least 2 at most 3 out-of-context rewritings. CoQA originally contains 8k conversations, which sum up to 127k question-answer pairs. CoQAR can be used in the supervised learning of three tasks: question paraphrasing, question rewriting and conversational question answering. In order to assess the quality of CoQAR’s rewritings, we conduct several experiments consisting in training and evaluating models for these three tasks. Our results support the idea that question rewriting can be used as a preprocessing step for (conversational and non-conversational) question answering models, thereby increasing their performances. Quentin Brabant, Gwénolé Lecorvé, Lina Maria Rojas-Barahona |
LREC | 3 |
| 2022 | Graph Neural Network Policies and Imitation Learning for Multi-Domain Task-Oriented DialoguesabstractTask-oriented dialogue systems are designed to achieve specific goals while conversing with humans.In practice, they may have to handle simultaneously several domains and tasks.The dialogue manager must therefore be able to take into account domain changes and plan over different domains/tasks in order to deal with multidomain dialogues.However, learning with reinforcement in such context becomes difficult because the state-action dimension is larger while the reward signal remains scarce.Our experimental results suggest that structured policies based on graph neural networks combined with different degrees of imitation learning can effectively handle multi-domain dialogues.The reported experiments underline the benefit of structured policies over standard policies. Thibault Cordier, Tanguy Urvoy, Fabrice Lefèvre, Lina Maria Rojas-Barahona |
SIGDIAL | 4 |
| 2022 | "Do you follow me?": A Survey of Recent Approaches in Dialogue State TrackingabstractWhile communicating with a user, a taskoriented dialogue system has to track the user's needs at each turn according to the conversation history.This process called dialogue state tracking (DST) is crucial because it directly informs the downstream dialogue policy.DST has received a lot of interest in recent years with the text-to-text paradigm emerging as the favored approach.In this review paper, we first present the task and its associated datasets.Then, considering a large number of recent publications, we identify highlights and advances of research in 2021-2022.Although neural approaches have enabled significant progress, we argue that some critical aspects of dialogue systems such as generalizability are still underexplored.To motivate future studies, we propose several research avenues. Léo Jacqmin, Lina Maria Rojas-Barahona, Benoît Favre |
SIGDIAL | 2 |
| 2019 | Graph2Bots, Unsupervised Assistance for Designing ChatbotsabstractWe present Graph2Bots, a tool for assisting conversational agent designers.It extracts a graph representation from human-human conversations by using unsupervised learning.The generated graph contains the main stages of the dialogue and their inner transitions.The graphical user interface (GUI) then allows graph editing. Jean Léon Bouraoui, Sonia Le Meitour, Romain Carbou, Lina Maria Rojas-Barahona, Vincent Lemaire 0001 |
SIGdial | 4 |
| 2019 | Spoken Conversational Search for General KnowledgeabstractLina M. Rojas Barahona, Pascal Bellec, Benoit Besset, Martinho Dossantos, Johannes Heinecke, Munshi Asadullah, Olivier Leblouch, Jeanyves. Lancien, Geraldine Damnati, Emmanuel Mory, Frederic Herledan. Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue. 2019. Lina Maria Rojas-Barahona, Pascal Bellec, Benoit Besset, Martinho Dos-Santos, Johannes Heinecke, Munshi Asadullah, Olivier Le Blouch, Jean Y. Lancien, Géraldine Damnati, Emmanuel Mory, Frédéric Herledan |
SIGdial | 1 |
| 2018 | Addressing Objects and Their Relations: The Conversational Entity Dialogue ModelabstractStefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Lina M. Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve Young, Milica Gašić. Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue. 2018. Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Lina Maria Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve J. Young, Milica Gasic |
SIGDIAL Conference | 4 |
| 2017 | A Network-based End-to-End Trainable Task-oriented Dialogue SystemabstractTsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, Steve Young. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Tsung-Hsien Wen, David Vandyke, Nikola Mrksic, Milica Gasic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, Steve J. Young |
EACL (1) | 5 |
| 2017 | Domain-Independent User Satisfaction Reward Estimation for Dialogue Policy LearningabstractLearning suitable and well-performing dialogue behaviour in statistical spoken dialogue systems has been in the focus of research for many years. While most work which is based on reinforcement learning employs an objective measure like task success for modelling the reward signal, we propose to use a reward based on user satisfaction. We will show in simulated experiments that a live user satisfaction estimation model may be applied resulting in higher estimated satisfaction whilst achieving similar success rates. Moreover, we will show that one satisfaction estimation model which has been trained on one domain may be applied in many other domains which cover a similar task. We will verify our findings by employing the model to one of the domains for learning a policy from real users and compare its performance to policies using the user satisfaction and task success acquired directly from the users as reward. Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Tsung-Hsien Wen, Milica Gasic, Steve J. Young |
INTERSPEECH | 5 |
| 2017 | Sub-domain Modelling for Dialogue Management with Hierarchical Reinforcement LearningabstractPaweł Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrkšić, Tsung-Hsien Wen, Iñigo Casanueva, Lina M. Rojas-Barahona, Milica Gašić. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017. Pawel Budzianowski, Stefan Ultes, Pei-hao Su, Nikola Mrksic, Tsung-Hsien Wen, Iñigo Casanueva, Lina Maria Rojas-Barahona, Milica Gasic |
SIGDIAL Conference | 7 |
| 2017 | DialPort, Gone Live: An Update After A Year of DevelopmentabstractKyusong Lee, Tiancheng Zhao, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David Traum, Stefan Ultes, Lina M. Rojas-Barahona, Milica Gasic, Steve Young, Maxine Eskenazi. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017. Kyusong Lee, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David R. Traum, Stefan Ultes, Lina Maria Rojas-Barahona, Milica Gasic, Steve J. Young, Maxine Eskénazi |
SIGDIAL Conference | 9 |
| 2017 | Reward-Balancing for Statistical Spoken Dialogue Systems using Multi-objective Reinforcement LearningabstractStefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, Tsung-Hsien Wen, Milica Gašić, Steve Young. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017. Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Tsung-Hsien Wen, Milica Gasic, Steve J. Young |
SIGDIAL Conference | 5 |
| 2017 | Dialogue manager domain adaptation using Gaussian process reinforcement learningabstractSpoken dialogue systems allow humans to interact with machines using natural speech. As such, they have many benefits. By using speech as the primary communication medium, a computer interface can facilitate swift, human-like acquisition of information. In recent years, speech interfaces have become ever more popular, as is evident from the rise of personal assistants such as Siri, Google Now, Cortana and Amazon Alexa. Recently, data-driven machine learning methods have been applied to dialogue modelling and the results achieved for limited-domain applications are comparable to or out-perform traditional approaches. Methods based on Gaussian processes are particularly effective as they enable good models to be estimated from limited training data. Furthermore, they provide an explicit estimate of the uncertainty which is particularly useful for reinforcement learning. This article explores the additional steps that are necessary to extend these methods to model multiple dialogue domains. We show that Gaussian process reinforcement learning is an elegant framework that naturally supports a range of methods, including prior knowledge, Bayesian committee machines and multi-agent learning, for facilitating extensible and adaptable dialogue systems. Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young |
Comput. Speech Lang. | 3 |
| 2016 | On-line Active Reward Learning for Policy Optimisation in Spoken Dialogue SystemsabstractThe ability to compute an accurate reward function is essential for optimising a dialogue policy via reinforcement learning. In real-world applications, using explicit user feedback as the reward signal is often unreliable and costly to collect. This problem can be mitigated if the user's intent is known in advance or data is available to pre-train a task success predictor off-line. In practice neither of these apply for most real world applications. Here we propose an on-line learning framework whereby the dialogue policy is jointly trained alongside the reward model via active learning with a Gaussian process model. This Gaussian process operates on a continuous space dialogue representation generated in an unsupervised fashion using a recurrent neural network encoder-decoder. The experimental results demonstrate that the proposed framework is able to significantly reduce data annotation costs and mitigate noisy user feedback in dialogue policy learning. Pei-hao Su, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young |
ACL (1) | 4 |
| 2016 | Exploiting Sentence and Context Representations in Deep Neural Models for Spoken Language UnderstandingabstractThis paper presents a deep learning architecture for the semantic decoder component of a Statistical Spoken Dialogue System. In a slot-filling dialogue, the semantic decoder predicts the dialogue act and a set of slot-value pairs from a set of n-best hypotheses returned by the Automatic Speech Recognition. Most current models for spoken language understanding assume (i) word-aligned semantic annotations as in sequence taggers and (ii) delexicalisation, or a mapping of input words to domain-specific concepts using heuristics that try to capture morphological variation but that do not scale to other domains nor to language variation (e.g., morphology, synonyms, paraphrasing ). In this work the semantic decoder is trained using unaligned semantic annotations and it uses distributed semantic representation learning to overcome the limitations of explicit delexicalisation. The proposed architecture uses a convolutional neural network for the sentence representation and a long-short term memory network for the context representation. Results are presented for the publicly available DSTC2 corpus and an In-car corpus which is similar to DSTC2 but has a significantly higher word error rate (WER). Lina Maria Rojas-Barahona, Milica Gasic, Nikola Mrksic, Pei-hao Su, Stefan Ultes, Tsung-Hsien Wen, Steve J. Young |
COLING | 1 |
| 2016 | Conditional Generation and Snapshot Learning in Neural Dialogue SystemsabstractTsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, David Vandyke, Steve Young. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, David Vandyke, Steve J. Young |
EMNLP | 4 |
| 2016 | Counter-fitting Word Vectors to Linguistic ConstraintsabstractNikola Mrkšić, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, Tsung-Hsien Wen, Steve Young. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Nikola Mrksic, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gasic, Lina Maria Rojas-Barahona, Pei-hao Su, David Vandyke, Tsung-Hsien Wen, Steve J. Young |
HLT-NAACL | 5 |
| 2016 | Multi-domain Neural Network Language Generation for Spoken Dialogue SystemsabstractTsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, Steve Young. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, David Vandyke, Steve J. Young |
HLT-NAACL | 4 |
| 2014 | Bayesian Inverse Reinforcement Learning for Modeling Conversational Agents in a Virtual Environment
Lina Maria Rojas-Barahona, Christophe Cerisara |
CICLing (1) | 1 |
| 2013 | Using Paraphrases and Lexical Semantics to Improve the Accuracy and the Robustness of Supervised Models in Situated Dialogue SystemsabstractThis paper explores to what extent lemmatisation, lexical resources, distributional semantics and paraphrases can increase the accuracy of supervised models for dialogue management.The results suggest that each of these factors can help improve performance but that the impact will vary depending on their combination and on the evaluation mode. Claire Gardent, Lina Maria Rojas-Barahona |
EMNLP | 2 |
| 2013 | Weakly and Strongly Constrained Dialogues for Language Learning
Claire Gardent, Alejandra Lorenzo, Laura Perez-Beltrachini, Lina Maria Rojas-Barahona |
SIGDIAL Conference | 4 |
| 2013 | Unsupervised structured semantic inference for spoken dialog reservation tasks
Alejandra Lorenzo, Lina Maria Rojas-Barahona, Christophe Cerisara |
SIGDIAL Conference | 2 |
| 2012 | Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora
Fabrice Lefèvre, Djamel Mostefa, Laurent Besacier, Yannick Estève, Matthieu Quignard, Nathalie Camelin, Benoît Favre, Bassam Jabaian, Lina Maria Rojas-Barahona |
LREC | 9 |
| 2012 | Building and Exploiting a Corpus of Dialog Interactions between French Speaking Virtual and Human Agents
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent |
LREC | 1 |
| 2012 | An End-to-End Evaluation of Two Situated Dialog Systems
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent |
SIGDIAL Conference | 1 |
| 2011 | An Incremental Architecture for the Semantic Annotation of Dialogue Corpora with High-Level Structures. A case of study for the MEDIA corpus
Lina Maria Rojas-Barahona, Matthieu Quignard |
SIGDIAL Conference | 1 |
| 2009 | HomeNL: Homecare Assistance in Natural Language. An Intelligent Conversational Agent for Hypertensive Patients Management
Lina Maria Rojas-Barahona, Silvana Quaglini, Mario Stefanelli |
AIME | 1 |