VLDB 2026 Research / reviewers in the wild / expert
Stephan Oepen
dblp:36/5477
· DBLP profile ↗
41ranked-venue papers
8as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 8 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Information extraction and text analysis · 69% Language models and text generation · 23% Learning paradigms · 8% | |
| Theoretical computer science
1 paper |
Logic in computer science · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
multilingual dataset |
0.9 | 1 | 2025 | An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
graph-based dependency parsing |
0.5 | 1 | 2021 | Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.5 | 1 | 2021 | Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
structured sentiment analysis |
0.5 | 1 | 2021 | Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021 |
Machine learning › Learning paradigms
multi-task learning |
0.3 | 1 | 2018 | Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › lexical semantics › multiword expression
noun compound interpretation |
0.3 | 1 | 2018 | Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › text classification
semantic classification |
0.3 | 1 | 2018 | Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › semantic analysis
negation scope detection |
0.2 | 1 | 2014 | Simple Negation Scope Resolution through Deep Parsing: A Semantic Solution to a Semantic Problem · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.1 | 1 | 2011 | Parser Evaluation over Local and Non-Local Deep Dependencies in a Large Corpus · EMNLP 2011 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 1 | 2011 | Parser Evaluation over Local and Non-Local Deep Dependencies in a Large Corpus · EMNLP 2011 |
Logic in computer science › semantics
underspecified semantics |
0.1 | 1 | 2014 | Simple Negation Scope Resolution through Deep Parsing: A Semantic Solution to a Semantic Problem · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis › data annotation
annotation efficiency |
0.1 | 1 | 2005 | High Precision Treebanking-Blazing Useful Trees Using POS Information · ACL 2005 |
Natural language and speech › Information extraction and text analysis › data annotation › corpus annotation
treebank |
0.1 | 1 | 2005 | High Precision Treebanking-Blazing Useful Trees Using POS Information · ACL 2005 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser evaluation |
0.0 | 1 | 2011 | Parser Evaluation over Local and Non-Local Deep Dependencies in a Large Corpus · EMNLP 2011 |
Natural language and speech › Language models and text generation
text generation |
0.0 | 1 | 2006 | Re-Usable Tools for Precision Machine Translation · ACL 2006 |
Methods — techniques the papers use, named apart from their topics
graph-based parsing · 0.5dependency parsing · 0.5semantic representation · 0.4deep parsing · 0.4transfer learning · 0.3neural classification · 0.3multi-task learning · 0.3symbolic resources · 0.1stochastic process · 0.1statistical ranking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context Is (Almost) Everything: Llama-3 on Structured Output and AMR Parsing
Maja Buljan, Stephan Oepen, Lilja Øvrelid |
LREC | 2 |
| 2026 | HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained ModelsabstractWe present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation. Stephan Oepen, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Georges Gabriel Charpentier, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Gema Ramírez-Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dusan Varis, Fedor Vitiugin, Jaume Zaragoza |
LREC | 1 |
| 2025 | An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)abstractLaurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Laurie Burchell, Ona de Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Tereza Vojtechová, Jaume Zaragoza-Bernabeu |
ACL (1) | 25 |
| 2025 | HPLT's Second Data ReleaseabstractWe describe the progress of the High Performance Language Technologies (HPLT) project, a 3-year EU-funded project that started in September 2022. We focus on the up-to-date results on the release of free text datasets derived from web crawls, one of the central objectives of the project. The second release used a revised processing pipeline, and an enlarged set of input crawls. From 4.5 petabytes of web crawls we extracted 7.6T tokens of monolingual text in 193 languages, plus 380 million parallel sentences in 51 language pairs. We also release MultiHPLT, a cross-combination of the parallel data, which produces 1,275 pairs, as well as releasing the containing documents for all parallel sentences in order to enable research in document-level MT. We report changes in the pipeline, analysis and evaluation results for the second parallel data release based on machine translation systems. All datasets are released under a permissive CC0 licence. Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Laurie Burchell, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu |
MTSummit (2) | 20 |
| 2024 | A New Massive Multilingual Dataset for High-Performance Language TechnologiesabstractWe present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the Internet Archive. We describe our methods for data acquisition, management and processing of large corpora, which rely on open-source software tools and high-performance computing. Our monolingual collection focuses on low- to medium-resourced languages and covers 75 languages and a total of ≈ 5.6 trillion word tokens de-duplicated on the document level. Our English-centric parallel corpus is derived from its monolingual counterpart and covers 18 language pairs and more than 96 million aligned sentence pairs with roughly 1.4 billion English tokens. The HPLT language resources are one of the largest open text corpora ever released, providing a great resource for language modeling and machine translation training. We publicly release the corpora, the software, and the tools used in this work. Ona de Gibert Bonet, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, Jörg Tiedemann |
LREC/COLING | 12 |
| 2021 | Structured Sentiment Analysis as Dependency Graph ParsingabstractJeremy Barnes, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jeremy Barnes 0001, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal |
ACL/IJCNLP (1) | 3 |
| 2020 | A Tale of Three Parsers: Towards Diagnostic Evaluation for Meaning Representation ParsingabstractWe discuss methodological choices in contrastive and diagnostic evaluation in meaning representation parsing, i.e. mapping from natural language utterances to graph-based encodings of its semantic structure. Drawing inspiration from earlier work in syntactic dependency parsing, we transfer and refine several quantitative diagnosis techniques for use in the context of the 2019 shared task on Meaning Representation Parsing (MRP). As in parsing proper, moving evaluation from simple rooted trees to general graphs brings along its own range of challenges. Specifically, we seek to begin to shed light on relative strenghts and weaknesses in different broad families of parsing techniques. In addition to these theoretical reflections, we conduct a pilot experiment on a selection of top-performing MRP systems and one of the five meaning representation frameworks in the shared task. Empirical results suggest that the proposed methodology can be meaningfully applied to parsing into graph-structured target representations, uncovering hitherto unknown properties of the different systems that can inform future development and cross-fertilization across approaches. Maja Buljan, Joakim Nivre, Stephan Oepen, Lilja Øvrelid |
LREC | 3 |
| 2018 | Transfer and Multi-Task Learning for Noun-Noun Compound InterpretationabstractIn this paper, we empirically evaluate the utility of transfer and multi-task learning on a challenging semantic classification task: semantic interpretation of noun-noun compounds.Through a comprehensive series of experiments and in-depth error analysis, we show that transfer learning via parameter initialization and multi-task learning via parameter sharing can help a neural classification model generalize over a highly skewed distribution of relations.Further, we demonstrate how dual annotation with two distinct sets of relations over the same set of compounds can be exploited to improve the overall accuracy of a neural classifier and its F 1 scores on the less frequent, but more difficult relations. Murhaf Fares, Stephan Oepen, Erik Velldal |
EMNLP | 2 |
| 2016 | Towards Comparability of Linguistic Graph Banks for Semantic Parsing
Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Silvie Cinková, Dan Flickinger, Jan Hajic 0001, Angelina Ivanova, Zdenka Uresová |
LREC | 1 |
| 2016 | Towards a Catalogue of Linguistic Graph BanksabstractGraphs exceeding the formal complexity of rooted trees are of growing relevance to much NLP research. Although formally well understood in graph theory, there is substantial variation in the types of linguistic graphs, as well as in the interpretation of various structural properties. To provide a common terminology and transparent statistics across different collections of graphs in NLP, we propose to establish a shared community resource with an open-source reference implementation for common statistics. Marco Kuhlmann, Stephan Oepen |
Comput. Linguistics | 2 |
| 2014 | Simple Negation Scope Resolution through Deep Parsing: A Semantic Solution to a Semantic ProblemabstractIn this work, we revisit Shared Task 1 from the 2012 *SEM Conference: the au-tomated analysis of negation. Unlike the vast majority of participating systems in 2012, our approach works over explicit and formal representations of proposi-tional semantics, i.e. derives the notion of negation scope assumed in this task from the structure of logical-form meaning rep-resentations. We relate the task-specific interpretation of (negation) scope to the concept of (quantifier and operator) scope in mainstream underspecified semantics. With reference to an explicit encoding of semantic predicate-argument structure, we can operationalize the annotation deci-sions made for the 2012 *SEM task, and demonstrate how a comparatively simple system for negation scope resolution can be built from an off-the-shelf deep parsing system. In a system combination setting, our approach improves over the best pub-lished results on this task to date. 1 Woodley Packard, Emily M. Bender, Jonathon Read, Stephan Oepen, Rebecca Dridan |
ACL (1) | 4 |
| 2014 | Towards an Encyclopedia of Compositional Semantics: Documenting the Interface of the English Resource Grammar
Dan Flickinger, Emily M. Bender, Stephan Oepen |
LREC | 3 |
| 2014 | Semantic Technologies for Querying Linguistic Annotations: An Experiment Focusing on Graph-Structured Data
Milen Kouylekov, Stephan Oepen |
LREC | 2 |
| 2014 | Off-Road LAF: Encoding and Processing Annotations in NLP Workflows
Emanuele Lapponi, Erik Velldal, Stephan Oepen, Rune Lain Knudsen |
LREC | 3 |
| 2013 | Machine Learning for High-Quality Tokenization Replicating Variable Tokenization Schemes
Murhaf Fares, Stephan Oepen, Yi Zhang 0003 |
CICLing (1) | 2 |
| 2012 | The WeSearch Corpus, Treebank, and Treecache - A Comprehensive Sample of User-Generated Content
Jonathon Read, Dan Flickinger, Rebecca Dridan, Stephan Oepen, Lilja Øvrelid |
LREC | 4 |
| 2012 | Speculation and Negation: Rules, Rankers, and the Role of SyntaxabstractThis article explores a combination of deep and shallow approaches to the problem of resolving the scope of speculation and negation within a sentence, specifically in the domain of biomedical research literature. The first part of the article focuses on speculation. After first showing how speculation cues can be accurately identified using a very simple classifier informed only by local lexical context, we go on to explore two different syntactic approaches to resolving the in-sentence scopes of these cues. Whereas one uses manually crafted rules operating over dependency structures, the other automatically learns a discriminative ranking function over nodes in constituent trees. We provide an in-depth error analysis and discussion of various linguistic properties characterizing the problem, and show that although both approaches perform well in isolation, even better results can be obtained by combining them, yielding the best published results to date on the CoNLL-2010 Shared Task data. The last part of the article describes how our speculation system is ported to also resolve the scope of negation. With only modest modifications to the initial design, the system obtains state-of-the-art results on this task also. Erik Velldal, Lilja Øvrelid, Jonathon Read, Stephan Oepen |
Comput. Linguistics | 4 |
| 2011 | Parser Evaluation over Local and Non-Local Deep Dependencies in a Large Corpus
Emily M. Bender, Dan Flickinger, Stephan Oepen, Yi Zhang 0003 |
EMNLP | 3 |
| 2011 | Treeblazing: Using External Treebanks to Filter Parse Forests for Parse Selection and Treebanking
Andrew MacKinlay, Rebecca Dridan, Dan Flickinger, Stephan Oepen, Timothy Baldwin |
IJCNLP | 4 |
| 2011 | Deep open-source machine translation
Francis Bond, Stephan Oepen, Eric Nichols, Dan Flickinger, Erik Velldal, Petter Haugereid |
Mach. Transl. | 2 |
| 2010 | Syntactic Scope Resolution in Uncertainty Analysis
Lilja Øvrelid, Erik Velldal, Stephan Oepen |
COLING | 3 |
| 2010 | WikiWoods: Syntacto-Semantic Annotation for English Wikipedia
Dan Flickinger, Stephan Oepen, Gisle Ytrestøl |
LREC | 2 |
| 2009 | Automatic Translation of Norwegian Noun Compounds
Lars Bungum, Stephan Oepen |
EAMT | 2 |
| 2008 | Some Fine Points of Hybrid Natural Language Parsing
Peter Adolphs, Stephan Oepen, Ulrich Callmeier, Berthold Crysmann, Dan Flickinger, Bernd Kiefer |
LREC | 2 |
| 2006 | Re-Usable Tools for Precision Machine TranslationabstractThe LOGON MT demonstrator assembles independently valuable general-purpose NLP components into a machine translation pipeline that capitalizes on output quality. The demonstrator embodies an interesting combination of hand-built, symbolic resources and stochastic processes. Jan Tore Lønning, Stephan Oepen |
ACL | 2 |
| 2006 | Using a Bi-Lingual Dictionary in Lexical Transfer
Lars Nygaard, Jan Tore Lønning, Torbjørn Nordgård, Stephan Oepen |
EAMT | 4 |
| 2006 | Statistical Ranking in Tactical Generation
Erik Velldal, Stephan Oepen |
EMNLP | 2 |
| 2006 | Discriminant-Based MRS Banking
Stephan Oepen, Jan Tore Lønning |
LREC | 1 |
| 2005 | High Precision Treebanking-Blazing Useful Trees Using POS InformationabstractIn this paper we present a quantitative and qualitative analysis of annotation in the Hinoki treebank of Japanese, and investigate a method of speeding annotation by using part-of-speech tags. The Hinoki treebank is a Redwoods-style treebank of Japanese dictionary definition sentences. 5,000 sentences are annotated by three different annotators and the agreement evaluated. An average agreement of 65.4% was found using strict agreement, and 83.5% using labeled precision. Exploiting POS tags allowed the annotators to choose the best parse with 19.5% fewer decisions. Takaaki Tanaka, Francis Bond, Stephan Oepen, Sanae Fujita |
ACL | 3 |
| 2005 | Holistic regression testing for high-quality MT: some methodological and technological reflections
Stephan Oepen, Helge Dyvik, Dan Flickinger, Jan Tore Lønning, Paul Meurer, Victoria Rosén |
EAMT | 1 |
| 2005 | High Efficiency Realization for a Wide-Coverage Unification Grammar
John Carroll 0001, Stephan Oepen |
IJCNLP | 2 |
| 2005 | SEM-I Rational MT: Enriching Deep Grammars with a Semantic Interface for Scalable Machine TranslationabstractIn the LOGON machine translation system where semantic transfer using Minimal Recursion Semantics is being developed in conjunction with two existing broad-coverage grammars of Norwegian and English, we motivate the use of a grammar-specific semantic interface (SEM-I) to facilitate the construction and maintenance of a scalable translation engine. The SEM-I is a theoretically grounded component of each grammar, capturing several classes of lexical regularities while also serving the crucial engineering function of supplying a reliable and complete specification of the elementary predications the grammar can realize. We make extensive use of underspecification and type hierarchies to maximize generality and precision. Dan Flickinger, Jan Tore Lønning, Helge Dyvik, Stephan Oepen, Francis Bond |
MTSummit | 4 |
| 2005 | Maximum Entropy Models for Realization RankingabstractIn this paper we describe and evaluate different statistical models for the task of realization ranking, i.e. the problem of discriminating between competing surface realizations generated for a given input semantics. Three models are trained and tested; an n-gram language model, a discriminative maximum entropy model using structural features, and a combination of these two. Our realization component forms part of a larger, hybrid MT system. Erik Velldal, Stephan Oepen |
MTSummit | 2 |
| 2004 | Road-testing the English Resource Grammar Over the British National Corpus
Timothy Baldwin, Emily M. Bender, Dan Flickinger, Ara Kim, Stephan Oepen |
LREC | 5 |
| 2004 | A Lexicon Module for a Grammar Development Environment
Ann A. Copestake, Fabre Lambeau, Benjamin Waldron, Francis Bond, Dan Flickinger, Stephan Oepen |
LREC | 6 |
| 2002 | The LinGO Redwoods Treebank: Motivation and Preliminary Applications
Stephan Oepen, Kristina Toutanova, Stuart M. Shieber, Christopher D. Manning, Dan Flickinger, Thorsten Brants |
COLING | 1 |
| 2000 | Parser engineering and performance profiling
Stephan Oepen, John Carroll 0001 |
Nat. Lang. Eng. | 1 |
| 2000 | Introduction to this Special Issue
Stephan Oepen, Dan Flickinger, Hans Uszkoreit, Jun'ichi Tsujii |
Nat. Lang. Eng. | 1 |
| 1998 | Towards systematic grammar profiling.Test suite technology 10 years after
Stephan Oepen, Dan Flickinger |
Comput. Speech Lang. | 1 |
| 1996 | TSNLP - Test Suites for Natural Language Processing
Sabine Lehmann, Stephan Oepen, Sylvie Regnier-Prost, Klaus Netter, Veronika Lux, Judith Klein, Kirsten Falkedal, Frederik Fouvry, Dominique Estival, Eva Dauphin, Herve Compagnion, Judith Baur, Lorna Balkan, Doug Arnold |
COLING | 2 |
| 1994 | DISCO-An HPSG-based NLP System and its Application for Appointment Scheduling Project Note
Hans Uszkoreit, Rolf Backofen, Stephan Busemann, Abdel Kader Diagne, Elizabeth A. Hinkelman, Walter Kasper, Bernd Kiefer, Hans-Ulrich Krieger, Klaus Netter, Günter Neumann, Stephan Oepen, Stephen P. Spackman |
COLING | 11 |