Jan Buys

dblp:157/2104 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1994-5832ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2026 MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages
Anri M. Lombard, Simbarashe Mawere, Temi Aina, Ethan Wolff, Sbonelo Gumede, Elan Novick, Francois Meyer, Jan Buys
LREC8
2025 Cross-Lingual Knowledge Projection and Knowledge Enhancement for Zero-Shot Question Answering in Low-Resource Languages
abstract
Knowledge bases (KBs) in low-resource languages (LRLs) are often incomplete, posing a challenge for developing effective question answering systems over KBs in those languages. On the other hand, the size of training corpora for LRL language models is also limited, restricting the ability to do zero-shot question answering using multilingual language models. To address these issues, we propose a two-fold approach. First, we introduce LeNS-Align, a novel cross-lingual mapping technique which improves the quality of word alignments extracted from parallel English-LRL text by combining lexical alignment, named entity recognition, and semantic alignment. LeNS-Align is applied to perform cross-lingual projection of KB triples. Second, we leverage the projected KBs to enhance multilingual language models’ question answering capabilities by augmenting the models with Graph Neural Networks embedding the projected knowledge. We apply our approach to map triples from two existing English KBs, ConceptNet and DBpedia, to create comprehensive LRL knowledge bases for four low-resource South African languages. Evaluation on three translated test sets show that our approach improves zero-shot question answering accuracy by up to 17% compared to baselines without KB access. The results highlight how our approach contributes to bridging the knowledge gap for low-resource languages by expanding knowledge coverage and question answering capabilities.
Sello Ralethe, Jan Buys
COLING2
2024 Multipath parsing in the brain
abstract
Humans understand sentences word-by-word, in the order that they hear them.This incrementality entails resolving temporary ambiguities about syntactic relationships.We investigate how humans process these syntactic ambiguities by correlating predictions from incremental generative dependency parsers with timecourse data from people undergoing functional neuroimaging while listening to an audiobook.In particular, we compare competing hypotheses regarding the number of developing syntactic analyses in play during word-by-word comprehension: one vs more than one.This comparison involves evaluating syntactic surprisal from a state-of-the-art dependency parser with LLM-adapted encodings against an existing fMRI dataset.In both English and Chinese data, we find evidence for multipath parsing.Brain regions associated with this multipath effect include bilateral superior temporal gyrus.
Berta Franzluebbers, Donald Dunagan, Milos Stanojevic, Jan Buys, John T. Hale
ACL (1)4
2024 Neural Machine Translation between Low-Resource Languages with Synthetic Pivoting
abstract
Training neural models for translating between low-resource languages is challenging due to the scarcity of direct parallel data between such languages. Pivot-based neural machine translation (NMT) systems overcome data scarcity by including a high-resource pivot language in the process of translating between low-resource languages. We propose synthetic pivoting, a novel approach to pivot-based translation in which the pivot sentences are generated synthetically from both the source and target languages. Synthetic pivot sentences are generated through sequence-level knowledge distillation, with the aim of changing the structure of pivot sentences to be closer to that of the source or target languages, thereby reducing pivot translation complexity. We incorporate synthetic pivoting into two paradigms for pivoting: cascading and direct translation using synthetic source and target sentences. We find that the performance of pivot-based systems highly depends on the quality of the NMT model used for sentence regeneration. Furthermore, training back-translation models on these sentences can make the models more robust to input-side noise. The results show that synthetic data generation improves pivot-based systems translating between low-resource Southern African languages by up to 5.6 BLEU points after fine-tuning.
Khalid Ahmed, Jan Buys
LREC/COLING2
2024 Triples-to-isiXhosa (T2X): Addressing the Challenges of Low-Resource Agglutinative Data-to-Text Generation
abstract
Most data-to-text datasets are for English, so the difficulties of modelling data-to-text for low-resource languages are largely unexplored. In this paper we tackle data-to-text for isiXhosa, which is low-resource and agglutinative. We introduce Triples-to-isiXhosa (T2X), a new dataset based on a subset of WebNLG, which presents a new linguistic context that shifts modelling demands to subword-driven techniques. We also develop an evaluation framework for T2X that measures how accurately generated text describes the data. This enables future users of T2X to go beyond surface-level metrics in evaluation. On the modelling side we explore two classes of methods - dedicated data-to-text models trained from scratch and pretrained language models (PLMs). We propose a new dedicated architecture aimed at agglutinative data-to-text, the Subword Segmental Pointer Generator (SSPG). It jointly learns to segment words and copy entities, and outperforms existing dedicated models for 2 agglutinative languages (isiXhosa and Finnish). We investigate pretrained solutions for T2X, which reveals that standard PLMs come up short. Fine-tuning machine translation models emerges as the best method overall. These findings underscore the distinct challenge presented by T2X: neither well-established data-to-text architectures nor customary pretrained methodologies prove optimal. We conclude with a qualitative analysis of generation errors and an ablation study.
Francois Meyer, Jan Buys
LREC/COLING2
2024 NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages
abstract
The Nguni languages have over 20 million home language speakers in South Africa. There has been considerable growth in the datasets for Nguni languages, but so far no analysis of the performance of NLP models for these languages has been reported across languages and tasks. In this paper we study pretrained language models for the 4 Nguni languages - isiXhosa, isiZulu, isiNdebele, and Siswati. We compile publicly available datasets for natural language understanding and generation, spanning 6 tasks and 11 datasets. This benchmark, which we call NGLUEni, is the first centralised evaluation suite for the Nguni languages, allowing us to systematically evaluate the Nguni-language capabilities of pretrained language models (PLMs). Besides evaluating existing PLMs, we develop new PLMs for the Nguni languages through multilingual adaptive finetuning. Our models, Nguni-XLMR and Nguni-ByT5, outperform their base models and large-scale adapted models, showing that performance gains are obtainable through limited language group-based adaptation. We also perform experiments on cross-lingual transfer and machine translation. Our models achieve notable cross-lingual transfer improvements in the lower resourced Nguni languages (isiNdebele and Siswati). To facilitate future use of NGLUEni as a standardised evaluation suite for the Nguni languages, we create a web portal to access the collection of datasets and publicly release our models.
Francois Meyer, Haiyue Song, Abhisek Chakrabarty, Jan Buys, Raj Dabre, Hideki Tanaka
LREC/COLING4
2023 Policy-based Reinforcement Learning for Generalisation in Interactive Text-based Environments
abstract
Text-based environments enable RL agents to learn to converse and perform interactive tasks through natural language.However, previous RL approaches applied to text-based environments show poor performance when evaluated on unseen games.This paper investigates the improvement of generalisation performance through the simple switch from a value-based update method to a policy-based one, within text-based environments.We show that by replacing commonly used value-based methods with REINFORCE with baseline, a far more general agent is produced.The policy-based agent is evaluated on Coin Collector and Question Answering with interactive text (QAit), two text-based environments designed to test zero-shot performance.We see substantial improvements on a variety of zero-shot evaluation experiments, including tripling accuracy on various QAit benchmark configurations.The results indicate that policy-based RL has significantly better generalisation capabilities than value-based methods within such text-based environments, suggesting that RL agents could be applied to more complex natural language environments.
Edan Toledo, Jan Buys, Jonathan Shock
EACL2
2022 Generic Overgeneralization in Pre-trained Language Models
abstract
Generic statements such as “ducks lay eggs” make claims about kinds, e.g., ducks as a category. The generic overgeneralization effect refers to the inclination to accept false universal generalizations such as “all ducks lay eggs” or “all lions have manes” as true. In this paper, we investigate the generic overgeneralization effect in pre-trained language models experimentally. We show that pre-trained language models suffer from overgeneralization and tend to treat quantified generic statements such as “all ducks lay eggs” as if they were true generics. Furthermore, we demonstrate how knowledge embedding methods can lessen this effect by injecting factual knowledge about kinds into pre-trained language models. To this end, we source factual knowledge about two types of generics, minority characteristic generics and majority characteristic generics, and inject this knowledge using a knowledge embedding model. Our results show that knowledge injection reduces, but does not eliminate, generic overgeneralization, and that majority characteristic generics of kinds are more susceptible to overgeneralization bias.
Sello Ralethe, Jan Buys
COLING2
2021 Discourse Understanding and Factual Consistency in Abstractive Summarization
abstract
Saadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Saadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi 0001
EACL5
2020 The Curious Case of Neural Text Degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, Yejin Choi 0001
ICLR2
2019 BottleSum: Unsupervised and Self-supervised Sentence Summarization using the Information Bottleneck Principle
abstract
Peter West, Ari Holtzman, Jan Buys, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Peter West, Ari Holtzman, Jan Buys, Yejin Choi 0001
EMNLP/IJCNLP (1)3
2018 Learning to Write with Cooperative Discriminators
abstract
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, Yejin Choi. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, Yejin Choi 0001
ACL (1)2
2018 Neural Syntactic Generative Models with Exact Marginalization
abstract
We present neural syntactic generative models with exact marginalization that support both dependency parsing and language modeling.Exact marginalization is made tractable through dynamic programming over shiftreduce parsing and minimal RNN-based feature sets.Our algorithms complement previous approaches by supporting batched training and enabling online computation of next word probabilities.For supervised dependency parsing, our model achieves a stateof-the-art result among generative approaches.We also report empirical results on unsupervised syntactic models and their role in language modeling.We find that our model formulation of latent dependencies with exact marginalization do not lead to better intrinsic language modeling performance than vanilla RNNs, and that parsing accuracy is not correlated with language modeling perplexity in stack-based models.
Jan Buys, Phil Blunsom
NAACL-HLT1
2017 Robust Incremental Neural Semantic Graph Parsing
abstract
Parsing sentences to linguisticallyexpressive semantic representations is a key goal of Natural Language Processing.Yet statistical parsing has focussed almost exclusively on bilexical dependencies or domain-specific logical forms.We propose a neural encoder-decoder transition-based parser which is the first full-coverage semantic graph parser for Minimal Recursion Semantics (MRS).The model architecture uses stack-based embedding features, predicting graphs jointly with unlexicalized predicates and their token alignments.Our parser is more accurate than attention-based baselines on MRS, and on an additional Abstract Meaning Representation (AMR) benchmark, and GPU batch processing makes it an order of magnitude faster than a high-precision grammar-based parser.Further, the 86.69%Smatch score of our MRS parser is higher than the upper-bound on AMR parsing, making MRS an attractive choice as a semantic representation.
Jan Buys, Phil Blunsom
ACL (1)1
2016 Cross-Lingual Morphological Tagging for Low-Resource Languages
abstract
Morphologically rich languages often lack the annotated linguistic resources required to develop accurate natural language processing tools.We propose models suitable for training morphological taggers with rich tagsets for low-resource languages without using direct supervision.Our approach extends existing approaches of projecting part-of-speech tags across languages, using bitext to infer constraints on the possible tags for a given word type or token.We propose a tagging model using Wsabie, a discriminative embeddingbased model with rank-based learning.In our evaluation on 11 languages, on average this model performs on par with a baseline weakly-supervised HMM, while being more scalable.Multilingual experiments show that the method performs best when projecting between related language pairs.Despite the inherently lossy projection, we show that the morphological tags predicted by our models improve the downstream performance of a parser by +0.6 LAS on average. 1 This extends, but is not fully consistent with, the set of 12 tags proposed by Petrov et al. (2012).
Jan Buys, Jan A. Botha
ACL (1)1
2016 Online Segment to Segment Neural Transduction
abstract
We introduce an online neural sequence to sequence model that learns to alternate between encoding and decoding segments of the input as it is read.By independently tracking the encoding and decoding representations our algorithm permits exact polynomial marginalization of the latent segmentation during training, and during decoding beam search is employed to find the best alignment path together with the predicted output sequence.Our model tackles the bottleneck of vanilla encoder-decoders that have to read and memorize the entire input sequence in their fixedlength hidden states before producing any output.It is different from previous attentive models in that, instead of treating the attention weights as output of a deterministic function, our model assigns attention weights to a sequential latent variable which can be marginalized out and permits online generation.Experiments on abstractive sentence summarization and morphological inflection show significant performance gains over the baseline encoder-decoders.
Lei Yu 0008, Jan Buys, Phil Blunsom
EMNLP2