VLDB 2026 Research / reviewers in the wild / expert
Martijn Wieling 0001
dblp:35/2985 · also Martijn B. Wieling
· DBLP profile ↗
19ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-0434-1526ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR
Reihaneh Amooie, Wietse de Vries, Rik van Noord, Martijn Wieling 0001 |
LREC | 5 |
| 2025 | Enhancing Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource PerformanceabstractAutomatic Speech Recognition (ASR) performance for low-resource languages is still far behind that of higher-resource languages such as English, due to a lack of sufficient labeled data. State-of-the-art methods deploy self-supervised transfer learning where a model pre-trained on large amounts of data is fine-tuned using little labeled data in a target low-resource language. In this paper, we present and examine a method for fine-tuning an SSL-based model in order to improve the performance for Frisian and its regional dialects (Clay Frisian, Wood Frisian, and South Frisian). We show that Frisian ASR performance can be improved by using multilingual (Frisian, Dutch, English and German) fine-tuning data and an auxiliary language identification task. In addition, our findings show that performance on dialectal speech suffers substantially, and, importantly, that this effect is moderated by the elicitation approach used to collect the dialectal data. Our findings also particularly suggest that relying solely on standard language data for ASR evaluation may underestimate real-world performance, particularly in languages with substantial dialectal variation. Reihaneh Amooie, Wietse de Vries, Jelske Dijkstra, Matt Coler, Martijn Wieling 0001 |
ICASSP | 6 |
| 2025 | Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancerabstractMeaningful speech assessment is vital in clinical phonetics and therapy monitoring. This study examined the link between perceptual speech assessments and objective acoustic measures in a large head and neck cancer (HNC) dataset. Trained listeners provided ratings of intelligibility, articulation, voice quality, phonation, speech rate, nasality, and background noise on speech. Strong correlations were found between subjective intelligibility, articulation, and voice quality, likely due to a shared underlying cause of speech symptoms in our speaker population. Objective measures of intelligibility and speech rate aligned with their subjective counterpart. Our results suggest that a single intelligibility measure may be sufficient for the clinical monitoring of speakers treated for HNC using concomitant chemoradiation. Bence Mark Halpern, Thomas Tienkamp, Teja Rebernik, R. J. J. H. van Son, Martijn Wieling 0001, Defne Abur, Tomoki Toda |
INTERSPEECH | 5 |
| 2025 | Articulatory clarity and variability before and after surgery for tongue cancerabstractSurgical treatment for tongue cancer can negatively affect the mobility and musculature of the tongue, which can influence articulatory clarity and variability. In this study, we investigated articulatory clarity through the vowel articulation index (VAI) and variability through vowel formant dispersion (VFD). Using a sentence reading task, we assessed 11 individuals pre and six months post tongue cancer surgery, alongside 11 sex- and age-matched typical speakers. Our results show that while the VAI was significantly smaller post-surgery compared to pre-surgery, there was no significant difference between patients and typical speakers at either time point. Post-surgery, speakers had higher VFD values for /i/ compared to pre-surgery and typical speakers, signalling higher variability. Taken together, our results suggest that while articulatory clarity remained within typical ranges following surgery for tongue cancer for the speakers in our study, articulatory variability increased. Thomas Tienkamp, Fleur van Ast, Roos van der Veen, Teja Rebernik, Raoul Buurke, Nikki Hoekzema, Katharina Polsterer, Hedwig Sekeres, R. J. J. H. van Son, Martijn Wieling 0001, Max J. H. Witjes, Sebastiaan A. H. J. de Visscher, Defne Abur |
INTERSPEECH | 10 |
| 2024 | Quantifying the effect of speech pathology on automatic and human speaker verificationabstractThis study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verification (ASV) system. Using two recently collected Dutch datasets with parallel pre and post-surgery audio from the same speaker, NKI-OC-VC and SPOKE, we assess the extent to which speech pathology influences ASV performance, and whether objective/subjective measures of speech severity are correlated with the performance. Finally, we carry out a perceptual study to compare judgements of ASV and human listeners. Our findings reveal that pathological speech negatively affects ASV performance, and the severity of the speech is negatively correlated with the performance. There is a moderate agreement in perceptual and objective scores of speaker similarity and severity, however, we could not clearly establish in the perceptual study, whether the same phenomenon also exists in human perception. Bence Mark Halpern, Thomas Tienkamp, Wen-Chin Huang, Lester Phillip Violeta, Teja Rebernik, Sebastiaan A. H. J. de Visscher, Max J. H. Witjes, Martijn Wieling 0001, Defne Abur, Tomoki Toda |
INTERSPEECH | 8 |
| 2024 | Exploring Self-Supervised Speech Representations for Cross-lingual Acoustic-to-Articulatory InversionabstractAcoustic-to-articulatory inversion (AAI) is the process of inferring vocal tract movements from acoustic speech signals. Despite its diverse potential applications, AAI research in languages other than English is scarce due to the challenges of collecting articulatory data. In recent years, self-supervised learning (SSL) based representations have shown great potential for addressing low-resource tasks. We utilize wav2vec 2.0 representations and English articulatory data for training AAI systems and investigates their effectiveness for a different language: Dutch. Results show that using mms-1b features can reduce the cross-lingual performance drop to less than 30%. We found that increasing model size, selecting intermediate rather than final layers, and including more pre-training data improved AAI performance. By contrast, fine-tuning on an ASR task did not. Our results therefore highlight promising prospects for implementing SSL in AAI for languages with limited articulatory data. Reihaneh Amooie, Wietse de Vries, Thomas Tienkamp, Rik van Noord, Martijn Wieling 0001 |
INTERSPEECH | 6 |
| 2023 | Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data AugmentationabstractThe performance of automatic speech recognition (ASR) systems has advanced substantially in recent years, particularly for languages for which a large amount of transcribed speech is available.Unfortunately, for low-resource languages, such as minority languages, regional languages or dialects, ASR performance generally remains much lower.In this study, we investigate whether data augmentation techniques could help improve low-resource ASR performance, focusing on four typologically diverse minority languages or language variants (West Germanic: Gronings, West-Frisian; Malayo-Polynesian: Besemah, Nasal).For all four languages, we examine the use of selftraining, where an ASR system trained with the available human-transcribed data is used to generate transcriptions, which are then combined with the original data to train a new ASR system.For Gronings, for which there was a preexisting text-to-speech (TTS) system available, we also examined the use of TTS to generate ASR training data from text-only sources.We find that using a self-training approach consistently yields improved performance (a relative WER reduction up to 20.5% compared to using an ASR system trained on 24 minutes of manually transcribed speech).The performance gain from TTS augmentation for Gronings was even stronger (up to 25.5% relative reduction in WER compared to a system based on 24 minutes of manually transcribed speech).In sum, our results show the benefit of using selftraining or (if possible) TTS-generated data as an efficient solution to overcome the limitations of data availability for resource-scarce languages in order to improve ASR performance. Martijn Bartelds, Nay San, Bradley McDonnell, Daniel Jurafsky, Martijn Wieling 0001 |
ACL (1) | 5 |
| 2023 | DUMB: A Dutch Model BenchmarkabstractWe introduce the Dutch Model Benchmark: DUMB.The benchmark includes a diverse set of datasets for low-, medium-and highresource tasks.The total set of nine tasks includes four tasks that were previously not available in Dutch.Instead of relying on a mean score across tasks, we propose Relative Error Reduction (RER), which compares the DUMB performance of language models to a strong baseline which can be referred to in the future even when assessing different sets of language models.Through a comparison of 14 pre-trained language models (monoand multi-lingual, of varying sizes), we assess the internal consistency of the benchmark tasks, as well as the factors that likely enable high performance.Our results indicate that current Dutch monolingual models under-perform and suggest training larger Dutch models with other architectures and pretraining objectives.At present, the highest performance is achieved by DeBERTaV3 large , XLM-R large and mDeBERTaV3 base .In addition to highlighting best strategies for training larger Dutch models, DUMB will foster further research on Dutch.A public leaderboard is available at dumbench.nl. WSD BERTje{1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 67.9 65.9 WSD RobBERT V1 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 61.3 60.6 WSD RobBERT V2 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.3 64.1 WSD RobBERT 2022 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 67.0 63.7 WSD mBERT cased {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 68.2 68.5 WSD XLM-R base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.7 66.5 WSD mDeBERTaV3 base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 69.8 69.6 WSD XLM-R large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 73.0 73.1 WSD BERT base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 60.0 58.2 WSD RoBERTa base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 62.6 61.1 WSD DeBERTaV3 base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.4 64.4 WSD BERT large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 60.2 57.2 WSD RoBERTa large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 64.3 59.1 WSD DeBERTaV3 large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 71.3 70.2 Wietse de Vries, Martijn Wieling 0001, Malvina Nissim |
EMNLP | 2 |
| 2022 | Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 LanguagesabstractCross-lingual transfer learning with large multilingual pre-trained models can be an effective approach for low-resource languages with no labeled training data.Existing evaluations of zero-shot cross-lingual generalisability of large pre-trained models use datasets with English training data, and test data in a selection of target languages.We explore a more extensive transfer learning setup with 65 different source languages and 105 target languages for part-of-speech tagging.Through our analysis, we show that pre-training of both source and target language, as well as matching language families, writing systems, word order systems, and lexical-phonetic distance significantly impact cross-lingual performance.The findings described in this paper can be used as indicators of which factors are important for effective zero-shot cross-lingual transfer to zero-and low-resource languages. Wietse de Vries, Martijn Wieling 0001, Malvina Nissim |
ACL (1) | 2 |
| 2022 | Quantifying Language Variation Acoustically with Few Resourcesabstractof authors shown on this cover page is limited to 10 maximum. Martijn Bartelds, Martijn Wieling 0001 |
NAACL-HLT | 2 |
| 2021 | Automatically Identifying Eviction Cases and Outcomes Within Case Law of Dutch Courts of First InstanceabstractIn this paper we attempt to identify eviction judgements within all case law published by Dutch courts in order to automate data collection, previously conducted manually. To do so we performed two experiments. The first focused on identifying judgements related to eviction, while the second focused on identifying the outcome of the cases in the judgements (eviction vs. dismissal of the landlord’s claim). In the process of conducting the experiments for this study, we have created a manually annotated dataset with eviction-related judgements and their outcomes. Masha Medvedeva, Thijmen Dam, Martijn Wieling 0001, Michel Vols |
JURIX | 3 |
| 2020 | JURI SAYS: An Automatic Judgement Prediction System for the European Court of Human RightsabstractIn this paper we present the web platform JURI SAYS that automatically predicts decisions of the European Court of Human Rights based on communicated cases, which are published by the court early in the proceedings and are often available many years before the final decision is made. Our system therefore predicts future judgements of the court. The platform is available at jurisays.com and shows the predictions compared to the actual decisions of the court. It is automatically updated every month by including the prediction for the new cases. Additionally, the system highlights the sentences and paragraphs that are most important for the prediction (i.e. violation vs. no violation of human rights). Masha Medvedeva, Martijn Wieling 0001, Michel Vols |
JURIX | 3 |
| 2018 | Project PiPeNovel: Pilot on Post-editing NovelsabstractGiven (i) the rise of a new paradigm to machine translation based on neural networks that results in more fluent and less literal output than previous models and (ii) the maturity of machine-assisted translation via post-editing in industry, project PiPeNovel studies the feasibility of the post-editing workflow for literary text conducting experiments with professional literary translators. Antonio Toral, Martijn Wieling 0001, Sheila Castilho, Joss Moorkens, Andy Way |
EAMT | 2 |
| 2018 | Reproducibility in Computational Linguistics: Are We Willing to Share?abstractThis study focuses on an essential precondition for reproducibility in computational linguistics: the willingness of authors to share relevant source code and data. Ten years after Ted Pedersen’s influential “Last Words” contribution in Computational Linguistics, we investigate to what extent researchers in computational linguistics are willing and able to share their data and code. We surveyed all 395 full papers presented at the 2011 and 2016 ACL Annual Meetings, and identified whether links to data and code were provided. If working links were not provided, authors were requested to provide this information. Although data were often available, code was shared less often. When working links to code or data were not provided in the paper, authors provided the code in about one third of cases. For a selection of ten papers, we attempted to reproduce the results using the provided data and code. We were able to reproduce the results approximately for six papers. For only a single paper did we obtain the exact same results. Our findings show that even though the situation appears to have improved comparing 2016 to 2011, empiricism in computational linguistics still largely remains a matter of faith. Nevertheless, we are somewhat optimistic about the future. Ensuring reproducibility is not only important for the field as a whole, but also seems worthwhile for individual researchers: The median citation count for studies with working links to the source code is higher. Martijn Wieling 0001, Josine Rawee, Gertjan van Noord |
Comput. Linguistics | 1 |
| 2017 | Analysis of Acoustic-to-Articulatory Speech Inversion Across Different Accents and Languages
Ganesh Sivaraman, Carol Y. Espy-Wilson, Martijn Wieling 0001 |
INTERSPEECH | 3 |
| 2017 | Sharing Is Caring: The Future of Shared TasksabstractShared tasks are indisputably drivers of progress and interest for problems in NLP. This is reflected by their increasing popularity, as well as by the fact that new shared tasks regularly emerge for under-researched and under-resourced topics, especially at workshops and smaller conferences.The general procedures and conventions for organizing a shared task have arisen organically over time (Paroubek, Chaudiron, and Hirschman, 2007, Section 7). There is no consistent framework that describes how shared tasks should be organized. This is not a harmful thing per se, but we believe that shared tasks, and by extension the field in general, would benefit from some reflection on the existing conventions. This, in turn, could lead to the future harmonization of shared task procedures.Shared tasks revolve around two aspects: research advancement and competition. We see research advancement as the driving force and main goal behind organizing them. Competition is an instrument to encourage and promote participation. However, just because these two forces are intrinsic to shared tasks does not mean that they always act in the same direction: Ensuring that the competition is fair is not a necessary requirement for advancing the field, and might even slow down progress.Our position in this respect is clear: We do believe that (i) advancing the field should be given priority over ensuring fair competition, also because (ii) inequality is partly unsolvable and intrinsic to life. In other words: Equality between competitors is desirable if it does not hinder research advancement.In the recently established workshop on ethics in NLP,1Parra Escartín et al. (2017) raise a set of considerations involving shared tasks, mainly focusing on areas where general ethical concerns regarding good scientific practice intersect with certain aspects of shared tasks. We find that they raise valid concerns, and in this contribution, we address some of them. However, we take a different perspective. Instead of focusing on ethical issues and potential negative effects of the competition aspect, we rather concentrate on how to bolster scientific progress.We make a simple proposal for the improvement of shared tasks, and discuss how it can help to mitigate the problems raised by Parra Escartín et al. (2017), while not necessarily tackling them directly. We start with assessing the concrete impact and significance of such concerns first.Recently, Parra Escartín et al. (2017) drew attention to a list of potential negative effects and ethical issues concerning shared tasks in NLP. In this section, we take this list as a starting point and examine each problem with respect to the main goal of shared tasks—to advance research in the field. Some issues were regarded as potential concerns rather than definite problems, because it is unclear how large their actual impact is. We believe that some of these concerns needed to be quantified in order to be properly assessed.To this end, we reviewed about 100 recent shared tasks from various campaigns (SemEval, EVALITA, CoNLL, WMT, and CLEF) between 2014 and 2016. We focused on several aspects, such as participation of companies, participation of organizers, closed versus open tracks, and the submission of papers by participants. Note that this is not an exhaustive overview of all shared tasks in NLP, but rather an arbitrary sample to investigate general trends in recent times. We use information drawn from this annotation exercise for assessing some of the problems we report in the following sections. The figures that are relevant for the discussion are reported in Table 1. The annotated spreadsheets used to collect this information are publicly available, together with some basic statistics and additional explanations.2Some potential concerns, although being ethically relevant, are not necessarily a problem in terms of research advancement, and fixing them directly should not be a priority. Here, we assess issues raised by Parra Escartín et al. (2017) that we believe fall into this category.Potential Conflicts of Interest. Parra Escartín et al. (2017) state that participation of organizers or annotators in their own shared task raises questions about inequality among participants, as organizers have earlier access to the data than the regular participants. In our survey, we found that in 5.8% of shared tasks, organizers did indeed participate. However, we also observe that this happens in connection with few participants (average 3.5 compared with 12.1, see Table 1), thus typically smaller shared tasks. This indicates that organizers' participation is more common in small, specialized tasks. The low number of participants can also explain why organizers perform better on average compared with non-organizers (see average normalized rank in Table 1).Unequal Playing Field. An unequal playing field mainly reflects the starting point that the different teams have. Parra Escartín et al. (2017) report on the issue of differences in processing power. An extreme example of this issue is the submission of Durrani et al. (2013) at WMT13, in which they reached the highest scores because they were able to boost the BLEU score by approximately 0.8% by making use of 1TB RAM, which was probably unavailable to the other teams at that time. Computational resources are not the only reason for an unequal playing field, though. There are many other causes that could lead to an unequal playing field—for example, some teams might have access to more proprietary data, proprietary software, or research equipment.The competitive nature of shared tasks can be fun, and stimulating for a variety of reasons (visibility, grant applications, beating state of the art, etc.). Such reasons might not necessarily be positively correlated with advancing the field, though. Here we discuss issues also raised by Parra Escartín et al. (2017) that we believe fall into this category.Secretiveness. As a result of the competitive nature of shared tasks, it can be desirable for participating teams to keep their “secret sauce” private, as this could mean an advantage for a re-run of the same task, or a shared task on a related problem. As a possible effect of secretiveness, Parra Escartín et al. (2017) also mention “Unconscious overlooking of ethical concerns,” actually referring to an unacceptable level of vagueness in papers. In other words, participants may unintentionally describe their systems in an abstract and vague way due to a previously established practice in systems' descriptions.Lack of Description of Negative Results. Given that negative results are informative, their under-representation in shared tasks is a concern. A lack of knowledge about negative results might lead to a research redundancy, which is clearly undesirable. The issue of under-represented negative results is a global concern for the entire field, and for science in general. However, shared tasks provide an excellent opportunity for publishing negative results, as the acceptance of papers for publication does not particularly favor positive results.Redundancy and Replicability in the Field. Parra Escartín et al. (2017) raise issues concerning two types of redundancy, (a) when optimal parameter settings of a previous shared task do not carry over to the new version of the task, therefore it is not clear what is learned; and (b) when algorithms are reimplemented for replicability purposes.Regarding (a), we think that differences in used parameter settings are not actually a bad thing; we learn from this that we overfit on the previous task, or that we need to adapt our systems to another data set or domain.Regarding (b), this is a real problem because starting from scratch to reimplement existing systems is unnecessarily time-consuming. In addition, it would always be desirable to be able to directly reproduce the same results of the same model for the same task (Pedersen, 2008; Fokkens et al., 2013).Withdrawal from Competition. Participants may withdraw from a shared task if their ranking in the competition can negatively affect their reputation and/or future funding. For example, Parra Escartín et al. (2017) suggest that companies might prefer to withdraw from the competition if they are not highly ranked, to avoid blemishing their reputation. This is something that we could not quantify in our survey, as in case of withdrawal there would be no evidence of participation in reports. There are two aspects, though, that we can quantify. The first aspect is the number of teams that do not publish their system's description, which amounts to approximately 9%, and could indeed be related to withdrawals. However, exactly because the paper is missing, information on why a team withdrew is not available. The second aspect is the total number of industry participants, which in our sample amounts to 20% (“Some company” and “Only company” in Table 1). Thus, although there is not much that can be done about withdrawal—and this might not be a problem anyway—we believe that, considering the substantial presence and interest of industries so far, their participation should be accommodated.Potentially Gaming the System. Shared tasks are usually bound to data sets and evaluation metrics. This could lead to competition-oriented participants focusing more on tuning their systems on a given data set and metrics rather than finding a scientifically sound and scalable method for solving the problem. This can, in turn, result in an “unfair” ranking or a misleading relation between a methodology and its value with respect to the research problem. A potential negative outcome of the latter is a scenario where “optimal” methods of a shared task do not carry over to related shared tasks. These problems become more severe when system gaming is combined with a secretive attitude. While tackling this issue, we should take into account that participants might be less eager to write about ad hoc solutions, for example tuning pre-processing components or tailoring a system too closely to specifics of the annotation.Our proposal for future shared tasks is not revolutionary. It simply revolves around the key aspects of sharing, not only resources but also experiences, including negative ones. Specifically, with research progress in mind, we believe sharing should be encouraged and even partially enforced. We therefore suggest an explicit setting for shared tasks in NLP, and reflect on the issue of what organizers could do in order to maximize sharing of information regarding participating systems. We also show how such a simple strategy can help to overcome the problems raised that can hinder research advancement.One of the challenges faced by shared tasks is to ensure a level playing field permitting a transparent comparison of the merits of different methods. The problem is that system A might come out on top of system B not because its method is superior, but because, for example, it was trained on more data. This would favor teams with access to more resources, like companies with large quantities of proprietary in-house data.Traditionally, this problem has been mitigated by establishing “closed tracks.” In closed tracks, participating teams are not allowed to use any training data other than that provided by the shared task organizers. The rationale behind this is that if all systems use exactly the same data, the playing field is equal, and the competition results will show the strengths of the different methods. In order to study the effect of additional training data, many shared tasks have a separate competition, the so-called “open track.”However, this open–closed division is increasingly impractical and ineffective. The main problem is that it is only concerned with training data, whereas the performance of systems can crucially depend on other resources. Examples are external components with pre-trained models, such as part-of-speech taggers and dependency parsers, auxiliary data-derived resources like word embeddings, and other influential factors like the availability of computational resources. Because such models are almost always derived from external language data, it is unclear where to draw the line between closed and open. Should such data be disallowed or not? If not, teams still do not really participate on an equal footing.From a research perspective, banning external models is completely impractical and nonsensical, as most state-of-the-art systems now depend on them. Likewise, trying to force all teams to use the same set of external models, and no other, would place a heavy burden on both organizers and participants. Moreover, restricting the resources participants can use is questionable, because, for research to progress quickly, teams should use the best resources available, or the resources best fitting their system.It is therefore unsurprising that the use of closed tracks has declined in shared tasks in general in the last few years, as we have observed during our review of shared tasks. However, the original problem of unequal playing field, and thus a bias in favor of teams with ample resources, remains.We propose, then, to rethink the problem, not in terms of equal training data, but in terms of equal opportunities. This is closely connected to the wider issue of reproducibility and replicability: Like all published research, shared task results should ideally be fully reproducible by anyone (Pedersen, 2008; Fokkens et al., 2013).3 Moreover, it should be easy to build on others' work to try out new variations of a method, without having to reimplement things from scratch. To ensure this, it is desirable that everything needed to reproduce experimental results is publicly and freely available, including code, data, pre-trained models, and so on. Interestingly, at the CoNLL-2013 shared task a similar step was taken, but only in terms of pre-condition: “While all teams in the shared task use the NUCLE corpus, they are also allowed to use additional external resources (both corpora and tools) so long as they are publicly available and not proprietary” (Ng et al., 2013). We would like to take this a step further, by enforcing the sharing of whatever resource teams might choose to use, so as to favor the injection of new resources in the field.Applying this principle to shared tasks in practice, we propose making the primary competition a “public track,” where participants can use any code, data, and pre-trained models they want, as long as others can then freely obtain them. In other words: All resources used to participate in the shared task should be subsequently shared with the community. Although this does not ensure equal access to resources for the current edition, it will still ensure a progressively more equal footing for the future. We believe this is the crucial step to move the field forward, as everyone will have access to the resources used in state-of-the-art systems. To keep participation possible for teams who cannot or will not make all resources available, a secondary, “proprietary track” can be established.Ranking of systems forms a large part of the appeal of shared tasks. However, rankings should not be overemphasized and are far from being the final goal of shared tasks. Research is supposed to teach us about the merits and characteristics of methods, including insights of what does not work, rather than about which team built the system that performed best on the test data.Negative results are very informative for future developments. Although publishing negative results is difficult, shared tasks do provide the ideal context for disclosing and explaining low performance methods and choices. We believe that shared task organizers should explicitly and strongly solicit the inclusion of what did not work in the reports written by participating teams. This could be even solicited via an online form that participants submit after the evaluation phase, where they comment on what worked well (as commonly done), but also provides a separate section to explain what did not work. This information could in turn be valuable data for organizers when compiling the overview report. Moreover, a clear explanation of what did not work, in connection with availability of code, would help to better understand whether something does not work as an idea or because of a specific implementation.More generally, organizers should encourage—and to some point ensure through the reviewing process—that all participants provide exhaustive reports, potentially including ablation/addition tests, so as to have a picture as comprehensive as possible. Because participating in shared tasks directly implies getting a paper accepted for publication, not everyone describes their system to the satisfaction of external reviewers. This should change, and acceptance should be conditional on clarity and exhaustiveness.We stated that the goal of advancing research should be prioritized over competition. This is especially the case when focusing on the competition aspect would encourage undesired practices like secretiveness and gaming the system. We suggest a simple solution based on the principle of maximizing resource- and information-sharing. As a byproduct, some problematic competition-related issues will be overcome, too. Some outstanding ethical issues cannot be solved, as they are intrinsic in human nature and cannot be controlled for by means of specific guidelines.Introducing proprietary and public tracks will stimulate participants to release their systems and resources. This will directly reduce Secretiveness and the issue of Redundancy and Replicability in the Field. It will also partially address the Unequal Playing Field problem, at least in the long run: Even if at the same competition different teams will have access to different resources, all resources will be available to everyone for the next round. Moreover, being able to access and run systems on different data sets will uncover limitations that might have been due to tailoring systems to the specifics of a given shared task (Potential Gaming the System). The presence of a proprietary track still allows for industrial participation (see Withdrawal from Competition in Section 2.2), where distribution of resources might not be as easy as for other teams.Encouraging participants to write comprehensive reports that include negative results will be a valid instrument towards advancing research, at the same time solving some outstanding problems. The reviewers should probably spend extra time in assessing the single reports and accept them conditionally on clarity requirements, but we believe this is worth the effort. Indeed, enforcing that systems are described properly will ensure and the of there are some issues We believe these are issues that cannot or need not be Withdrawal of teams cannot be controlled if of negative results is encouraged and common practice, it is possible that teams will choose to the competition. on progress and sharing rather than will also The of interest is not relevant in our We do believe that organizers should be allowed to and this is especially for shared tasks that might a number of participating teams due to the nature of the As long as these are explicitly reported in both the overview paper and the system the of results and ranking is to the closed tracks also an equal playing will it make it to different methods over the same We do not think Equality will be increasingly by resource sharing, to the that it is as inequality is part of the As we in Section comparison of methods has not been transparent in closed tracks the use of resources is not clarity in reports and release of systems will make it possible for the to assess which methods work and which do not, and to progressively on the state of the that our and will discussion on shared tasks, to make them more for driving progress in and each task will to have their own settings that the organizers will most However, we do believe that participants to release their in terms of resources, and of and should be a common to from long we together and on various with many It is to everyone for their However, we to mention a few who have to into better The discussion we at of the of the Computational at the of was the actual for this are to everyone and in to and for their and valuable has and to the on what does not work with the open versus closed track setting as it We and Parra Escartín for on earlier of this We are also to for Malvina Nissim, Lasha Abzianidze, Kilian Evang, Rob van der Goot, Hessel Haagsma, Barbara Plank, Martijn Wieling 0001 |
Comput. Linguistics | 7 |
| 2016 | ALT Explored: Integrating an Online Dialectometric Tool and an Online Dialect Atlas
Martijn Wieling 0001, Eva Sassolini, Sebastiana Cucurullo, Simonetta Montemagni |
LREC | 1 |
| 2013 | Word frequency, vowel length and vowel quality in speech production: an EMA study of the importance of experienceabstractA frequently replicated finding is that higher frequency words tend to be shorter and contain more strongly reduced vowels. However, little is known about potential differences in the articulatory gestures for high vs. low frequency words. The present study made use of electromagnetic articulography to investigate the production of two German vowels, [i] and [a], embedded in high and low frequency words. We found that word frequency differently affected the production of [i] and [a] at the temporal as well as the gestural level. Higher frequency of use predicted greater acoustic durations for long vowels; reduced durations for short vowels; articulatory trajectories with greater tongue height for [i] and more pronounced downward articulatory trajectories for [a]. These results show that the phonological contrast between short and long vowels is learned better with experience, and challenge both the Smooth Signal Redundancy Hypothesis and current theories of German phonology. Fabian Tomaschek, Martijn Wieling 0001, Denis Arnold, R. Harald Baayen |
INTERSPEECH | 2 |
| 2011 | Bipartite spectral graph partitioning for clustering dialect varieties and detecting their linguistic features
Martijn Wieling 0001, John Nerbonne |
Comput. Speech Lang. | 1 |