Nicholas Asher

dblp:33/3380 · DBLP profile ↗
← Back
48ranked-venue papers
14as first author
14since 2021 · last 2026
0000-0002-7689-8246ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 10 first-author · 13 since 2021Theory of computation · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SSA: Improving Performance With a Better Scoring Function
abstract
While transformer models exhibit strong incontext learning (ICL) abilities, they often fail to generalize under simple distribution shifts.We analyze these failures and identify Softmax, the scoring function in the attention mechanism, as a contributing factor.We propose Scaled Signed Averaging (SSA), a novel attention scoring function that mitigates these failures.SSA significantly improves performance on our ICL tasks and outperforms transformer models with Softmax on several NLP benchmarks and linguistic probing tasks, in both decoder-only and encoder-only architectures.
Omar Naim, Swarnadeep Bhar, Jérôme Bolte, Nicholas Asher
ACL (1)4
2026 EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions
abstract
Mainstream multilingual LLMs are generally trained on a much higher proportion of English than multilingual data, raising questions about their ability to capture linguistic features particular to non-English languages or to capture information important to non-anglophone cultures.We add to a growing effort to increase multilingual sensitivity in LLMs by developing a benchmark, EIFFEL, testing mastery of French idiomatic expressions in context.We fully explain the methodology, which exploits input from native French speakers, to make it reproducible for other languages.We compare mainstream multilingual LLMs with Frenchfocused LLMs both on standard LLM benchmarks and EIFFEL; EIFFEL brings out the benefits of higher proportions of French data and shows limitations of standard benchmarks for measuring multilingual competence.We also train from scratch a series of 1B SLMs with different proportions of French and English pretraining data that confirm EIFFEL's lessons.
Charlotte Noel, Nicholas Asher, Olivier Gouvert, Farah Benamara, Julie Hunter 0001
ACL (1)2
2026 COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following
abstract
Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actions. We identify this failure as arising from a lack of coupling between detection and execution, and propose that reliable behavior requires enforcing missing information as a precondition for action. We instantiate this principle in Cocoreli, a modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification. In Cocoreli, detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution. We evaluate Cocoreli in a controlled construction environment isolating underspecification and sequential execution. Cocoreli makes execution reliable by imposing explicit task structure: instructions are represented as parameterized executable objects and execution under unresolved specifications is blocked by construction, eliminating hallucinated actions. In contrast, chain-of-thought, prompt-chaining, and ReAct-style reasoning may still execute under incomplete specifications despite high detection rates. The same representation supports abstraction and reuse, and generalizes to API workflow tasks on ToolBench. These results show that reliable collaborative execution under underspecified instructions requires architectural enforcement, not just model capability.
Swarnadeep Bhar, Omar Naim, Eleni Metheniti, Loïc Cabannes, Bastien Navarri, Morteza Kamaladdini Ezzabady, Nicholas Asher
SIGDIAL7
2026 SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
abstract
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adjacent frames or clips. SagaQA addresses this gap by requiring high-level comprehension of extended multimodal narratives in entire TV shows. A distinguishing feature of SagaQA is the granularity of its reasoning steps. Our dataset necessitates long-range reasoning hops to connect information across completely different episodes. This requires models to reason over entire events and actions, demanding a deep understanding of the show’s narration and progression at a multimodal level. Motivated by recent progress in agentic methods, we further study how different planning strategies handle such complex reasoning. We categorize these approaches into three classes—parallel, sequential, and hybrid planners—and evaluate their ability to generate coherent and complete reasoning plans. Our results on SagaQA suggest that hybrid planners consistently produce higher-quality plans and exhibit stronger capabilities for complex, high-level narrative understanding in TV shows.
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen
SIGDIAL3
2026 KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
abstract
Abstract Chain-of-thought (CoT) traces have been shown to improve performance of large language models on a plethora of reasoning tasks, yet there is no consensus on the mechanism by which this boost is achieved. To shed more light on this, we introduce Causal CoT Graphs (CCGraphs), which are directed acyclic graphs automatically extracted from reasoning traces that model finegrained causal dependencies in language-model outputs. A collection of 1671 mathematical reasoning problems from MATH500, GSM8K, and AIME, together with their associated CCGraphs, has been compiled into our dataset—KisMATH. Our detailed empirical analysis with 15 open-weight LLMs shows that (i) reasoning nodes in the CCGraphs are causal contributors to the final answer, which we argue is constitutive of reasoning; and (ii) LLMs emphasize the reasoning paths captured by the CCGraphs, indicating that the models internally realize structures similar to our graphs. KisMATH enables controlled, graph-aligned interventions and opens avenues for further investigation into the role of CoT in LLM reasoning.
Soumadeep Saha, Akshay Chaturvedi, Saptarshi Saha, Utpal Garain, Nicholas Asher
Trans. Assoc. Comput. Linguistics5
2025 DIMSUM: Discourse in Mathematical Reasoning as a Supervision Module
abstract
We look at reasoning on GSM8k, a dataset of short texts presenting primary school, math problems. We find, with Mirzadeh et al (2024), that current LLM progress on the data set may not be explained by better reasoning but by exposure to a broader pretraining data distribution. We then introduce a novel information source for helping models with less data or inferior training reason better: discourse structure. We show that discourse structure improves performance for models like Llama2 13b by up to 160%. Even for models that have most likely memorized the data set, adding discourse structural information to the model still improves predictions and dramatically improves large model performance on out of distribution examples.
Krish Sharma, Niyar R. Barman, Akshay Chaturvedi, Nicholas Asher
SIGDIAL4
2024 Discourse Structure for the Minecraft Corpus
abstract
We provide a new linguistic resource: The Minecraft Structured Dialogue Corpus (MSDC), a discourse annotated version of the Minecraft Dialogue Corpus (MDC; Narayan-Chen et al., 2019), with complete, situated discourse structures in the style of SDRT (Asher and Lascarides, 2003). Our structures feature both linguistic discourse moves and nonlinguistic actions. To show computational tractability, we train a discourse parser with a novel “2 pass architecture” on MSDC that gives excellent results on attachment prediction and relation labeling tasks especially long distance attachments.
Kate Thompson, Julie Hunter 0001, Nicholas Asher
LREC/COLING3
2024 On Explaining with Attention Matrices
abstract
This paper explores the much discussed, possible explanatory link between attention weights (AW) in transformer models and predicted output. Contrary to intuition and early research on attention, more recent prior research has provided formal arguments and empirical evidence that AW are not explanatorily relevant. We show that the formal arguments are incorrect. We introduce and effectively compute efficient attention, which isolates the effective components of attention matrices in tasks and models in which AW play an explanatory role. We show that efficient attention has a causal role (provides minimally necessary and sufficient conditions) for predicting model output in NLP tasks requiring contextual information, and we show, contrary to [7], that efficient attention matrices are probability distributions and are effectively calculable. Thus, they should play an important part in the explanation of attention based model behavior. We offer empirical experiments in support of our method illustrating various properties of efficient attention with various metrics on four datasets.
Omar Naim, Nicholas Asher
ECAI2
2024 Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question Answering
abstract
Abstract Transformer-based language models have been shown to be highly effective for several NLP tasks. In this article, we consider three transformer models, BERT, RoBERTa, and XLNet, in both small and large versions, and investigate how faithful their representations are with respect to the semantic content of texts. We formalize a notion of semantic faithfulness, in which the semantic content of a text should causally figure in a model’s inferences in question answering. We then test this notion by observing a model’s behavior on answering questions about a story after performing two novel semantic interventions—deletion intervention and negation intervention. While transformer models achieve high performance on standard question answering tasks, we show that they fail to be semantically faithful once we perform these interventions for a significant number of cases (∼ 50% for deletion intervention, and ∼ 20% drop in accuracy for negation intervention). We then propose an intervention-based training regime that can mitigate the undesirable effects for deletion intervention by a significant margin (from ∼ 50% to ∼ 6%). We analyze the inner-workings of the models to better understand the effectiveness of intervention-based training for deletion intervention. But we show that this training does not attenuate other aspects of semantic unfaithfulness such as the models’ inability to deal with negation intervention or to capture the predicate–argument structure of texts. We also test InstructGPT, via prompting, for its ability to handle the two interventions and to capture predicate–argument structure. While InstructGPT models do achieve very high performance on predicate–argument structure task, they fail to respond adequately to our deletion and negation interventions.
Akshay Chaturvedi, Swarnadeep Bhar, Soumadeep Saha, Utpal Garain, Nicholas Asher
Comput. Linguistics5
2024 Transport-based Counterfactual Models
abstract
Counterfactual frameworks have grown popular in machine learning for both explaining algorithmic decisions but also defining individual notions of fairness, more intuitive than typical group fairness conditions. However, state-of-the-art models to compute counterfactuals are either unrealistic or unfeasible. In particular, while Pearl's causal inference provides appealing rules to calculate counterfactuals, it relies on a model that is unknown and hard to discover in practice. We address the problem of designing realistic and feasible counterfactuals in the absence of a causal model. We define transport-based counterfactual models as collections of joint probability distributions between observable distributions, and show their connection to causal counterfactuals. More specifically, we argue that optimal-transport theory defines relevant transport-based counterfactual models, as they are numerically feasible, statistically-faithful, and can coincide under some assumptions with causal counterfactual models. Finally, these models make counterfactual approaches to fairness feasible, and we illustrate their practicality and efficiency on fair learning. With this paper, we aim at laying out the theoretical foundations for a new, implementable approach to counterfactual thinking.
Lucas de Lara, Alberto González-Sanz, Nicholas Asher, Laurent Risser, Jean-Michel Loubes
J. Mach. Learn. Res.3
2023 A simple but effective model for attachment in discourse parsing with multi-task learning for relation labeling
abstract
We present a discourse parsing model for conversation trained on the STAC corpus (Asher et al., 2016).We fine-tune a BERT-based model to encode pairs of discourse units and use a simple linear layer to predict discourse attachments.We then exploit a multi-task setting to predict relation labels, which effectively aids in the difficult task of relation type prediction; our F1-score equals or surpasses the state of the art in the approaches we have reimplemented using code from the authors with no loss in performance for attachment, confirming the intuitive interdependence of these two tasks.Our method also improves over other discourse parsing models in the literature in permitting attachments in which one node has multiple parents, an important feature of multiparty conversation.
Zineb Bennis, Julie Hunter 0001, Nicholas Asher
EACL3
2022 Tractable Explanations for d-DNNF Classifiers
abstract
Compilation into propositional languages finds a growing number of practical uses, including in constraint programming, diagnosis and machine learning (ML), among others. One concrete example is the use of propositional languages as classifiers, and one natural question is how to explain the predictions made. This paper shows that for classifiers represented with some of the best-known propositional languages, different kinds of explanations can be computed in polynomial time. These languages include deterministic decomposable negation normal form (d-DNNF), and so any propositional language that is strictly less succinct than d-DNNF. Furthermore, the paper describes optimizations, specific to Sentential Decision Diagrams (SDDs), which are shown to yield more efficient algorithms in practice.
Xuanxiang Huang, Yacine Izza, Alexey Ignatiev, Martin C. Cooper, Nicholas Asher, João Marques-Silva 0001
AAAI5
2022 BelElect: A New Dataset for Bias Research from a "Dark" Platform
Sviatlana Höhn, Sjouke Mauw, Nicholas Asher
ICWSM3
2021 Fair and Adequate Explanations
Nicholas Asher, Soumya Paul, Chris Russell 0001
CD-MAKE1
2019 Data Programming for Learning Discourse Structure
abstract
This paper investigates the advantages and limits of data programming for the task of learning discourse structure.The data programming paradigm implemented in the Snorkel framework allows a user to label training data using expert-composed heuristics, which are then transformed via the "generative step" into probability distributions of the class labels given the training candidates.These results are later generalized using a discriminative model.Snorkel's attractive promise to create a large amount of annotated data from a smaller set of training data by unifying the output of a set of heuristics has yet to be used for computationally difficult tasks, such as that of discourse attachment, in which one must decide where a given discourse unit attaches to other units in a text in order to form a coherent discourse structure.Although approaching this problem using Snorkel requires significant modifications to the structure of the heuristics, we show that weak supervision methods can be more than competitive with classical supervised learning approaches to the attachment problem.
Sonia Badene, Kate Thompson, Jean-Pierre Lorré, Nicholas Asher
ACL (1)4
2019 Weak Supervision for Learning Discourse Structure
abstract
Sonia Badene, Kate Thompson, Jean-Pierre Lorré, Nicholas Asher. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sonia Badene, Kate Thompson, Jean-Pierre Lorré, Nicholas Asher
EMNLP/IJCNLP (1)4
2018 A Dependency Perspective on RST Discourse Parsing and Evaluation
abstract
Computational text-level discourse analysis mostly happens within Rhetorical Structure Theory (RST), whose structures have classically been presented as constituency trees, and relies on data from the RST Discourse Treebank (RST-DT); as a result, the RST discourse parsing community has largely borrowed from the syntactic constituency parsing community. The standard evaluation procedure for RST discourse parsers is thus a simplified variant of PARSEVAL, and most RST discourse parsers use techniques that originated in syntactic constituency parsing. In this article, we isolate a number of conceptual and computational problems with the constituency hypothesis. We then examine the consequences, for the implementation and evaluation of RST discourse parsers, of adopting a dependency perspective on RST structures, a view advocated so far only by a few approaches to discourse parsing. While doing that, we show the importance of the notion of headedness of RST structures. We analyze RST discourse parsing as dependency parsing by adapting to RST a recent proposal in syntactic parsing that relies on head-ordered dependency trees, a representation isomorphic to headed constituency trees. We show how to convert the original trees from the RST corpus, RST-DT, and their binarized versions used by all existing RST parsers to head-ordered dependency trees. We also propose a way to convert existing simple dependency parser output to constituent trees. This allows us to evaluate and to compare approaches from both constituent-based and dependency-based perspectives in a unified framework, using constituency and dependency metrics. We thus propose an evaluation framework to compare extant approaches easily and uniformly, something the RST parsing community has lacked up to now. We can also compare parsers’ predictions to each other across frameworks. This allows us to characterize families of parsing strategies across the different frameworks, in particular with respect to the notion of headedness. Our experiments provide evidence for the conceptual similarities between dependency parsers and shift-reduce constituency parsers, and confirm that dependency parsing constitutes a viable approach to RST discourse parsing.
Mathieu Morey, Philippe Muller, Nicholas Asher
Comput. Linguistics3
2017 How much progress have we made on RST discourse parsing? A replication study of recent results on the RST-DT
abstract
This article evaluates purported progress over the past years in RST discourse parsing.Several studies report a relative error reduction of 24 to 51% on all metrics that authors attribute to the introduction of distributed representations of discourse units.We replicate the standard evaluation of 9 parsers, 5 of which use distributed representations, from 8 studies published between 2013 and 2017, using their predictions on the test set of the RST-DT.Our main finding is that most recently reported increases in RST discourse parser performance are an artefact of differences in implementations of the evaluation procedure.We evaluate all these parsers with the standard Parseval procedure to provide a more accurate picture of the actual RST discourse parsers performance in standard evaluation settings.Under this more stringent procedure, the gains attributable to distributed representations represent at most a 16% relative error reduction on fully-labelled structures.
Mathieu Morey, Philippe Muller, Nicholas Asher
EMNLP3
2017 A logic of sights
abstract
International audience
Cédric Dégremont, Soumya Paul, Nicholas Asher
J. Log. Comput.3
2016 Discourse Structure and Dialogue Acts in Multiparty Dialogue: the STAC Corpus
Nicholas Asher, Julie Hunter 0001, Mathieu Morey, Farah Benamara, Stergos D. Afantenos
LREC1
2016 Parallel Discourse Annotations on a Corpus of Short Texts
Manfred Stede, Stergos D. Afantenos, Andreas Peldszus, Nicholas Asher, Jérémy Perret
LREC4
2016 Integer Linear Programming for Discourse Parsing
abstract
Jérémy Perret, Stergos Afantenos, Nicholas Asher, Mathieu Morey. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Jérémy Perret, Stergos D. Afantenos, Nicholas Asher, Mathieu Morey
HLT-NAACL3
2016 Integrating Type Theory and Distributional Semantics: A Case Study on Adjective-Noun Compositions
abstract
In this article, we explore an integration of a formal semantic approach to lexical meaning and an approach based on distributional methods. First, we outline a formal semantic theory that aims to combine the virtues of both formal and distributional frameworks. We then proceed to develop an algebraic interpretation of that formal semantic theory and show how at least two kinds of distributional models make this interpretation concrete. Focusing on the case of adjective–noun composition, we compare several distributional models with respect to the semantic information that a formal semantic theory would need, and we show how to integrate the information provided by distributional models back into the formal semantic framework.
Nicholas Asher, Tim Van de Cruys, Antoine Bride, Márta Abrusán
Comput. Linguistics1
2015 A Generalisation of Lexical Functions for Composition in Distributional Semantics
abstract
Antoine Bride, Tim Van de Cruys, Nicholas Asher. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Antoine Bride, Tim Van de Cruys, Nicholas Asher
ACL (1)3
2015 Discourse parsing for multi-party chat dialogues
abstract
In this paper we present the first ever, to the best of our knowledge, discourse parser for multi-party chat dialogues.Discourse in multi-party dialogues dramatically differs from monologues since threaded conversations are commonplace rendering prediction of the discourse structure compelling.Moreover, the fact that our data come from chats renders the use of syntactic and lexical information useless since people take great liberties in expressing themselves lexically and syntactically.We use the dependency parsing paradigm as has been done in the past (Muller et al., 2012;Li et al., 2014).We learn local probability distributions and then use MST for decoding.We achieve 0.680 F 1 on unlabelled structures and 0.516 F 1 on fully labeled structures which is better than many state of the art systems for monologues, despite the inherent difficulties that multi-party chat dialogues have.
Stergos D. Afantenos, Eric Kow, Nicholas Asher, Jérémy Perret
EMNLP3
2014 Unsupervised extraction of semantic relations using discourse cues
Juliette Conrath, Stergos D. Afantenos, Nicholas Asher, Philippe Muller
COLING3
2014 What have we learned in formal semantics about ontology?
Nicholas Asher
FOIS1
2013 Measuring the Effect of Discourse Structure on Sentiment Analysis
Baptiste Chardon, Farah Benamara, Yvette Yannick Mathieu, Vladimir Popescu, Nicholas Asher
CICLing (2)5
2013 Grounding Strategic Conversation: Using Negotiation Dialogues to Predict Trades in a Win-Lose Game
abstract
This paper describes a method that predicts which trades players execute during a winlose game.Our method uses data collected from chat negotiations of the game The Settlers of Catan and exploits the conversation to construct dynamically a partial model of each player's preferences.This in turn yields equilibrium trading moves via principles from game theory.We compare our method against four baselines and show that tracking how preferences evolve through the dialogue and reasoning about equilibrium moves are both crucial to success.
Anaïs Cadilhac, Nicholas Asher, Farah Benamara, Alex Lascarides
EMNLP2
2013 Expressivity and comparison of models of discourse structure
Antoine Venant, Nicholas Asher, Philippe Muller, Pascal Denis, Stergos D. Afantenos
SIGDIAL Conference2
2012 Constrained Decoding for Text-Level Discourse Parsing
Philippe Muller, Stergos D. Afantenos, Pascal Denis, Nicholas Asher
COLING4
2012 An empirical resource for discovering cognitive principles of discourse organisation: the ANNODIS corpus
Stergos D. Afantenos, Nicholas Asher, Farah Benamara, Myriam Bras, Cécile Fabre, Lydia-Mai Ho-Dac, Anne Le Draoulec, Philippe Muller, Marie-Paule Péry-Woodley, Laurent Prévot 0001, Josette Rebeyrolle, Ludovic Tanguy, Marianne Vergez-Couret, Laure Vieu
LREC2
2011 Commitments to Preferences in Dialogue
Anaïs Cadilhac, Nicholas Asher, Farah Benamara, Alex Lascarides
SIGDIAL Conference2
2010 Testing SDRT's Right Frontier
Stergos D. Afantenos, Nicholas Asher
COLING2
2010 Extracting and Modelling Preferences from Dialogue
Nicholas Asher, Elise Bonzon, Alex Lascarides
IPMU1
2008 Categorizing Opinion in Discourse
abstract
International audience
Nicholas Asher, Farah Benamara, Yvette Yannick Mathieu
ECAI1
2008 A Type Driven Theory of Predication with Complex Types
Nicholas Asher
Fundam. Informaticae1
1995 Toward a Geometry of Common Sense: A Semantics and a Complete Axiomatization of Mereotopology
Nicholas Asher, Laure Vieu
IJCAI (1)1
1994 Intentions and Indormation in Discourse
abstract
This paper is about the flow of inference between communicative intentions, discourse structure and the domain during discourse processing. We augment a theory of discourse interpretation with a theory of distinct mental attitudes and reasoning about them, in order to provide an account of how the attitudes interact with reasoning about discourse structure.
Nicholas Asher, Alex Lascarides
ACL1
1994 Reasoning About Action and Time with Epistemic Conditionals
Nicholas Asher
ISMIS1
1993 A Semantics and Pragmatics for the Pluperfect
Alex Lascarides, Nicholas Asher
EACL2
1992 Interferring Discourse Relations in Context
abstract
We investigate various contextual effects on text interpretation, and account for them by providing contextual constraints in a logical theory of text interpretation. On the basis of the way these constraints interact with the other knowledge sources, we draw some general conclusions about the role of domain-specific information, top-down and bottom-up discourse information flow, and the usefulness of formalisation in discourse theory.
Alex Lascarides, Nicholas Asher, Jon Oberlander
ACL2
1991 Discourse Relations and Defeasible Knowledge
abstract
This paper presents a formal account of the temporal interpretation of text. The distinct natural interpretations of texts with similar syntax are explained in terms of defeasible rules characterising causal laws and Gricean-style pragmatic maxims. Intuitively compelling patterns of defeasible entailment that are supported by the logic in which the theory is expressed are shown to underly temporal interpretation.
Alex Lascarides, Nicholas Asher
ACL2
1991 Commonsense Entailment: A Modal Theory of Non-monotonic Reasoning
Nicholas Asher, Michael Morreau
IJCAI1
1990 Intentional Paradoxes and an Inductive Theory of Propositional Quantification
Nicholas Asher
TARK1
1988 Reasoning about Belief and Knowledge with Self-Reference and Time
Nicholas Asher
TARK1
1986 BUILDRS: An Implementation of DR Theory and LFG
Hajime Wada, Nicholas Asher
COLING2
1986 The Knower's Paradox and Representational Theories of Attitudes
Nicholas Asher, Johan A. W. Kamp
TARK1