VLDB 2026 Research / reviewers in the wild / expert
Philippe Muller
dblp:77/5733
· DBLP profile ↗
31ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-6765-4020ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 36% Knowledge representation and reasoning · 31% Trustworthy machine learning · 30% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 77% User interface design and tools · 23% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Human-AI interaction › large language models
large language model evaluation |
1.0 | 1 | 2026 | iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision · CHI 2026 |
Machine learning › Trustworthy machine learning › interpretability › logic-based explanation
abductive explanation |
0.7 | 1 | 2023 | Leveraging Argumentation for Generating Robust Sample-based Explanations · IJCAI 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
argumentation |
0.7 | 1 | 2023 | Leveraging Argumentation for Generating Robust Sample-based Explanations · IJCAI 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › argumentation
argumentation-based explanation |
0.7 | 1 | 2023 | Leveraging Argumentation for Generating Robust Sample-based Explanations · IJCAI 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Leveraging Argumentation for Generating Robust Sample-based Explanations · IJCAI 2023 |
Natural language and speech › Information extraction and text analysis › text segmentation
discourse segmentation |
0.5 | 1 | 2021 | Weakly supervised discourse segmentation for multiparty oral conversations · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse parsing |
0.3 | 1 | 2017 | How much progress have we made on RST discourse parsing? A replication study of recent results on the RST-DT · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse parsing
rhetorical structure theory |
0.3 | 1 | 2017 | How much progress have we made on RST discourse parsing? A replication study of recent results on the RST-DT · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › text similarity
semantic similarity |
0.2 | 1 | 2014 | Predicting the relevance of distributional semantic similarity with contextual information · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis
temporal information extraction |
0.1 | 1 | 2011 | Predicting Globally-Coherent Temporal Structures from Texts via Endpoint Inference and Graph Decomposition · IJCAI 2011 |
Natural language and speech › Information extraction and text analysis › temporal information extraction
temporal ordering |
0.1 | 1 | 2011 | Predicting Globally-Coherent Temporal Structures from Texts via Endpoint Inference and Graph Decomposition · IJCAI 2011 |
Empirical software engineering › reproducibility
replication study |
0.1 | 1 | 2017 | How much progress have we made on RST discourse parsing? A replication study of recent results on the RST-DT · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse coherence |
0.1 | 1 | 2014 | Predicting the relevance of distributional semantic similarity with contextual information · ACL (1) 2014 |
Mathematical optimization
integer programming |
0.0 | 1 | 2011 | Predicting Globally-Coherent Temporal Structures from Texts via Endpoint Inference and Graph Decomposition · IJCAI 2011 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
qualitative reasoning |
0.0 | 1 | 1998 | A Qualitative Theory of Motion Based on Spatio-Temporal Primitives · KR 1998 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › qualitative reasoning
qualitative spatial reasoning |
0.0 | 1 | 1998 | A Qualitative Theory of Motion Based on Spatio-Temporal Primitives · KR 1998 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.0argumentation · 0.7parseval evaluation · 0.6distributed representations · 0.6weak supervision · 0.5latent model · 0.5heuristic rules · 0.5integer linear programming · 0.2constraint optimization · 0.2distributional analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision
Jingwen Bai 0006, Wei Soon Cheong, Philippe Muller, Brian Y. Lim |
CHI | 3 |
| 2026 | SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV SeriesabstractWe introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adjacent frames or clips. SagaQA addresses this gap by requiring high-level comprehension of extended multimodal narratives in entire TV shows. A distinguishing feature of SagaQA is the granularity of its reasoning steps. Our dataset necessitates long-range reasoning hops to connect information across completely different episodes. This requires models to reason over entire events and actions, demanding a deep understanding of the show’s narration and progression at a multimodal level. Motivated by recent progress in agentic methods, we further study how different planning strategies handle such complex reasoning. We categorize these approaches into three classes—parallel, sequential, and hybrid planners—and evaluate their ability to generate coherent and complete reasoning plans. Our results on SagaQA suggest that hybrid planners consistently produce higher-quality plans and exhibit stronger capabilities for complex, high-level narrative understanding in TV shows. Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen |
SIGDIAL | 4 |
| 2024 | DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse ProcessingabstractThis paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks: RST, SDRT, PDTB, and Discourse Dependencies. We present an overview of the data, its development across three NLP shared tasks on discourse processing carried out in the past five years, and the latest modifications and added extensions. We also carry out an evaluation of state-of-the-art multilingual systems trained on the data for each task, showing plateau performance on segmentation, but important room for improvement for connective identification and relation classification. The DISRPT benchmark employs a unified format that we make available on GitHub and HuggingFace in order to encourage future work on discourse processing across languages, domains, and frameworks. Chloé Braud, Amir Zeldes, Laura Rivière, Yang Janet Liu, Philippe Muller, Damien Sileo, Tatsuya Aoyama |
LREC/COLING | 5 |
| 2024 | Zero-shot Learning for Multilingual Discourse Relation ClassificationabstractClassifying discourse relations is known as a hard task, relying on complex indices. On the other hand, discourse-annotated data is scarce, especially for languages other than English: many corpora, of limited size, exist for several languages but the domain is split between different theoretical frameworks that have a huge impact on the nature of the textual spans to be linked, and the label set used. Moreover, each annotation project implements modifications compared to the theoretical background and other projects. These discrepancies hinder the development of systems taking advantage of all the available data to tackle data sparsity and work on transfer between languages is very limited, almost nonexistent between frameworks, while it could improve our understanding of some theoretical aspects and enhance many applications. In this paper, we propose the first experiments on zero-shot learning for discourse relation classification and investigate several paths in the way source data can be combined, either based on languages, frameworks, or similarity measures. We demonstrate how difficult transfer is for the task at hand, and that the most impactful factor is label set divergence, where the notion of underlying framework possibly conceals crucial disagreements. Eleni Metheniti, Philippe Muller, Chloé Braud, Margarita Hernández Casas |
LREC/COLING | 2 |
| 2023 | Leveraging Argumentation for Generating Robust Sample-based ExplanationsabstractExplaining predictions made by inductive classifiers has become crucial with the rise of complex models acting more and more as black-boxes. Abductive explanations are one of the most popular types of explanations that are provided for the purpose. They highlight feature-values that are sufficient for making predictions. In the literature, they are generated by exploring the whole feature space, which is unreasonable in practice. This paper solves the problem by introducing explanation functions that generate abductive explanations from a sample of instances. It shows that such functions should be defined with great care since they cannot satisfy two desirable properties at the same time, namely existence of explanations for every individual decision (success) and correctness of explanations (coherence). The paper provides a parameterized family of argumentation-based explanation functions, each of which satisfies one of the two properties. It studies their formal properties and their experimental behaviour on different datasets. Leila Amgoud, Philippe Muller, Henri Trenquier |
IJCAI | 2 |
| 2022 | A Pragmatics-Centered Evaluation Framework for Natural Language UnderstandingabstractNew models for natural language understanding have recently made an unparalleled amount of progress, which has led some researchers to suggest that the models induce universal text representations. However, current benchmarks are predominantly targeting semantic phenomena; we make the case that pragmatics needs to take center stage in the evaluation of natural language understanding. We introduce PragmEval, a new benchmark for the evaluation of natural language understanding, that unites 11 pragmatics-focused evaluation datasets for English. PragmEval can be used as supplementary training data in a multi-task learning setup, and is publicly available, alongside the code for gathering and preprocessing the datasets. Using our evaluation suite, we show that natural language inference, a widely used pretraining task, does not result in genuinely universal representations, which presents a new challenge for multi-task learning. Damien Sileo, Philippe Muller, Tim Van de Cruys, Camille Pradel |
LREC | 2 |
| 2021 | Weakly supervised discourse segmentation for multiparty oral conversationsabstractDiscourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks.While segmentation models for written text tend to perform well, they are not directly applicable to spontaneous, oral conversation, which has linguistic features foreign to written text.Segmentation is less studied for this type of language, where annotated data is scarce, and existing corpora more heterogeneous.We develop a weak supervision approach to adapt, using minimal annotation, a state of the art discourse segmenter trained on written text to French conversation transcripts.Supervision is given by a latent model bootstrapped by manually defined heuristic rules that use linguistic and acoustic information.The resulting model improves the original segmenter, especially in contexts where information on speaker turns is lacking or noisy, gaining up to 13% in F-score.Evaluation is performed on data like those used to define our heuristic rules, but also on transcripts from two other corpora. Lila Gravellier, Julie Hunter 0001, Philippe Muller, Thomas Pellegrini, Isabelle Ferrané |
EMNLP (1) | 3 |
| 2020 | DiscSense: Automated Semantic Analysis of Discourse MarkersabstractUsing a model trained to predict discourse markers between sentence pairs, we predict plausible markers between sentence pairs with a known semantic relation (provided by existing classification datasets). These predictions allow us to study the link between discourse markers and the semantic relations annotated in classification datasets. Handcrafted mappings have been proposed between markers and discourse relations on a limited set of markers and a limited set of categories, but there exists hundreds of discourse markers expressing a wide variety of relations, and there is no consensus on the taxonomy of relations between competing discourse theories (which are largely built in a top-down fashion). By using an automatic prediction method over existing semantically annotated datasets, we provide a bottom-up characterization of discourse markers in English. The resulting dataset, named DiscSense, is publicly available. Damien Sileo, Tim Van de Cruys, Camille Pradel, Philippe Muller |
LREC | 4 |
| 2019 | Which aspects of discourse relations are hard to learn? Primitive decomposition for discourse relation classificationabstractDiscourse relation classification has proven to be a hard task, with rather low performance on several corpora that notably differ on the relation set they use.We propose to decompose the task into smaller, mostly binary tasks corresponding to various primitive concepts encoded into the discourse relation definitions.More precisely, we translate the discourse relations into a set of values for attributes based on distinctions used in the mappings between discourse frameworks proposed by Sanders et al. (2018).This arguably allows for a more robust representation of discourse relations, and enables us to address usually ignored aspects of discourse relation prediction, namely multiple labels and underspecified annotations.We study experimentally which of the conceptual primitives are harder to learn from the Penn Discourse Treebank English corpus, and propose a correspondence to predict the original labels, with preliminary empirical comparisons with a direct model. Charlotte Roze, Chloé Braud, Philippe Muller |
SIGdial | 3 |
| 2018 | A Dependency Perspective on RST Discourse Parsing and EvaluationabstractComputational text-level discourse analysis mostly happens within Rhetorical Structure Theory (RST), whose structures have classically been presented as constituency trees, and relies on data from the RST Discourse Treebank (RST-DT); as a result, the RST discourse parsing community has largely borrowed from the syntactic constituency parsing community. The standard evaluation procedure for RST discourse parsers is thus a simplified variant of PARSEVAL, and most RST discourse parsers use techniques that originated in syntactic constituency parsing. In this article, we isolate a number of conceptual and computational problems with the constituency hypothesis. We then examine the consequences, for the implementation and evaluation of RST discourse parsers, of adopting a dependency perspective on RST structures, a view advocated so far only by a few approaches to discourse parsing. While doing that, we show the importance of the notion of headedness of RST structures. We analyze RST discourse parsing as dependency parsing by adapting to RST a recent proposal in syntactic parsing that relies on head-ordered dependency trees, a representation isomorphic to headed constituency trees. We show how to convert the original trees from the RST corpus, RST-DT, and their binarized versions used by all existing RST parsers to head-ordered dependency trees. We also propose a way to convert existing simple dependency parser output to constituent trees. This allows us to evaluate and to compare approaches from both constituent-based and dependency-based perspectives in a unified framework, using constituency and dependency metrics. We thus propose an evaluation framework to compare extant approaches easily and uniformly, something the RST parsing community has lacked up to now. We can also compare parsers’ predictions to each other across frameworks. This allows us to characterize families of parsing strategies across the different frameworks, in particular with respect to the notion of headedness. Our experiments provide evidence for the conceptual similarities between dependency parsers and shift-reduce constituency parsers, and confirm that dependency parsing constitutes a viable approach to RST discourse parsing. Mathieu Morey, Philippe Muller, Nicholas Asher |
Comput. Linguistics | 2 |
| 2017 | How much progress have we made on RST discourse parsing? A replication study of recent results on the RST-DTabstractThis article evaluates purported progress over the past years in RST discourse parsing.Several studies report a relative error reduction of 24 to 51% on all metrics that authors attribute to the introduction of distributed representations of discourse units.We replicate the standard evaluation of 9 parsers, 5 of which use distributed representations, from 8 studies published between 2013 and 2017, using their predictions on the test set of the RST-DT.Our main finding is that most recently reported increases in RST discourse parser performance are an artefact of differences in implementations of the evaluation procedure.We evaluate all these parsers with the standard Parseval procedure to provide a more accurate picture of the actual RST discourse parsers performance in standard evaluation settings.Under this more stringent procedure, the gains attributable to distributed representations represent at most a 16% relative error reduction on fully-labelled structures. Mathieu Morey, Philippe Muller, Nicholas Asher |
EMNLP | 2 |
| 2016 | A Supervised Approach for Enriching the Relational Structure of Frame Semantics in FrameNetabstractFrame semantics is a theory of linguistic meanings, and is considered to be a useful framework for shallow semantic analysis of natural language. FrameNet, which is based on frame semantics, is a popular lexical semantic resource. In addition to providing a set of core semantic frames and their frame elements, FrameNet also provides relations between those frames (hence providing a network of frames i.e. FrameNet). We address here the limited coverage of the network of conceptual relations between frames in FrameNet, which has previously been pointed out by others. We present a supervised model using rich features from three different sources: structural features from the existing FrameNet network, information from the WordNet relations between synsets projected into semantic frames, and corpus-collected lexical associations. We show large improvements over baselines consisting of each of the three groups of features in isolation. We then use this model to select frame pairs as candidate relations, and perform evaluation on a sample with good precision. Shafqat Mumtaz Virk, Philippe Muller, Juliette Conrath |
COLING | 2 |
| 2016 | Corpus Annotation within the French FrameNet: a Domain-by-domain Methodology
Marianne Djemaa, Marie Candito, Philippe Muller, Laure Vieu |
LREC | 3 |
| 2016 | A General Framework for the Annotation of Causality Based on FrameNet
Laure Vieu, Philippe Muller, Marie Candito, Marianne Djemaa |
LREC | 2 |
| 2014 | Predicting the relevance of distributional semantic similarity with contextual informationabstractUsing distributional analysis methods to compute semantic proximity links between words has become commonplace in NLP.The resulting relations are often noisy or difficult to interpret in general.This paper focuses on the issues of evaluating a distributional resource and filtering the relations it contains, but instead of considering it in abstracto, we focus on pairs of words in context.In a discourse, we are interested in knowing if the semantic link between two items is a byproduct of textual coherence or is irrelevant.We first set up a human annotation of semantic links with or without contextual information to show the importance of the textual context in evaluating the relevance of semantic similarity, and to assess the prevalence of actual semantic relations between word tokens.We then built an experiment to automatically predict this relevance, evaluated on the reliable reference data set which was the outcome of the first annotation.We show that in-document information greatly improve the prediction made by the similarity level alone. Philippe Muller, Cécile Fabre, Clémentine Adam |
ACL (1) | 1 |
| 2014 | Unsupervised extraction of semantic relations using discourse cues
Juliette Conrath, Stergos D. Afantenos, Nicholas Asher, Philippe Muller |
COLING | 4 |
| 2014 | Developing a French FrameNet: Methodology and First results
Marie Candito, Pascal Amsili, Lucie Barque, Farah Benamara, Gaël de Chalendar, Marianne Djemaa, Pauline Haas, Richard Huyghe, Yvette Yannick Mathieu, Philippe Muller, Benoît Sagot, Laure Vieu |
LREC | 10 |
| 2013 | Expressivity and comparison of models of discourse structure
Antoine Venant, Nicholas Asher, Philippe Muller, Pascal Denis, Stergos D. Afantenos |
SIGDIAL Conference | 3 |
| 2012 | Constrained Decoding for Text-Level Discourse Parsing
Philippe Muller, Stergos D. Afantenos, Pascal Denis, Nicholas Asher |
COLING | 1 |
| 2012 | An empirical resource for discovering cognitive principles of discourse organisation: the ANNODIS corpus
Stergos D. Afantenos, Nicholas Asher, Farah Benamara, Myriam Bras, Cécile Fabre, Lydia-Mai Ho-Dac, Anne Le Draoulec, Philippe Muller, Marie-Paule Péry-Woodley, Laurent Prévot 0001, Josette Rebeyrolle, Ludovic Tanguy, Marianne Vergez-Couret, Laure Vieu |
LREC | 8 |
| 2011 | Predicting Globally-Coherent Temporal Structures from Texts via Endpoint Inference and Graph DecompositionabstractAn elegant approach to learning temporal orderings from texts is to formulate this problem as a constraint optimization problem, which can be then given an exact solution using Integer Linear Programming. This works well for cases where the number of possible relations between temporal entities is restricted to the mere precedence relation [Bramsen et al., 2006; Chambers and Jurafsky, 2008], but becomes impractical when considering all possible interval relations. This paper proposes two innovations, inspired from work on temporal reasoning, that control this combinatorial blow-up, therefore rendering an exact ILP inference viable in the general case. First, we translate our network of constraints from temporal intervals to their endpoints, to handle a drastically smaller set of constraints, while preserving the same temporal information. Second, we show that additional efficiency is gained by enforcing coherence on particular subsets of the entire temporal graphs. We evaluate these innovations through various experiments on TimeBank 1.2, and compare our ILP formulations with various baselines and oracle systems. 1 Pascal Denis, Philippe Muller |
IJCAI | 2 |
| 2011 | Evaluating Temporal Graphs Built from Texts via Transitive ReductionabstractTemporal information has been the focus of recent attention in information extraction, leading to some standardization effort, in particular for the task of relating events in a text. This task raises the problem of comparing two annotations of a given text, because relations between events in a story are intrinsically interdependent and cannot be evaluated separately. A proper evaluation measure is also crucial in the context of a machine learning approach to the problem. Finding a common comparison referent at the text level is not obvious, and we argue here in favor of a shift from event-based measures to measures on a unique textual object, a minimal underlying temporal graph, or more formally the transitive reduction of the graph of relations between event boundaries. We support it by an investigation of its properties on synthetic data and on a well-know temporal corpus. Xavier Tannier, Philippe Muller |
J. Artif. Intell. Res. | 2 |
| 2010 | Comparison of different algebras for inducing the temporal structure of texts
Pascal Denis, Philippe Muller |
COLING | 2 |
| 2010 | Learning Recursive Segments for Discourse Parsing
Stergos D. Afantenos, Pascal Denis, Philippe Muller, Laurence Danlos |
LREC | 3 |
| 2008 | Evaluation Metrics for Automatic Temporal Annotation of Texts
Xavier Tannier, Philippe Muller |
LREC | 2 |
| 2005 | Using Inference for Evaluating Models of Temporal DiscourseabstractThis paper addresses the problem of building and evaluating models of the temporal interpretation of a discourse in natural language. The extraction of temporal information is a complicated task as it is not limited to finding pieces of information at specific places in a text. A lot of temporal data is made of relations between events, or relations between events and dates. Building such information is highly context-dependent, taking into account information more than a sentence at a time. Moreover it is not clear what the target representation should be: the way it is done by human beings is still a subject of study in itself. It seems to require some sort of reasoning, either purely temporal or involving complex world knowledge. This is the reason why evaluating this task is also problematic when trying to design a system for it. We present a method for enriching the detection of event-to-event relations with a basic reasoning model that can be also used for helping to compare the extraction of temporal information by a system and by a human being. We have experimented with this method on a set of texts, comparing a very basic model of tense interpretation with a more complex model inspired by the Reichenbach 's well-known theory of narrative discourse. Philippe Muller, Axel Reymonet |
TIME | 1 |
| 2004 | Word Sense Disambiguation using a dictionary for sense similarity measure
Bruno Gaume, Nabil Hathout, Philippe Muller |
COLING | 3 |
| 2004 | Annotating and measuring temporal relations in texts
Philippe Muller, Xavier Tannier |
COLING | 1 |
| 2002 | Topological Spatio-Temporal Reasoning and RepresentationabstractWe present here a theory of motion from a topological point of view, in a symbolic perspective. Taking space–time histories of objects as primitive entities, we introduce temporal and topological relations on the thus defined space–time to characterize classes of spatial changes. The theory thus accounts for qualitative spatial information, dealing with underspecified, symbolic information when accurate data are not available or unnecessary. We show that these structures give a basis for commonsense spatio–temporal reasoning by presenting a number of significant deductions in the theory. This can serve as a formal basis for languages describing motion events in a qualitative way. Philippe Muller |
Comput. Intell. | 1 |
| 2001 | Plausible reasoning from spatial observations
Jérôme Lang, Philippe Muller |
UAI | 2 |
| 1998 | A Qualitative Theory of Motion Based on Spatio-Temporal Primitives
Philippe Muller |
KR | 1 |