VLDB 2026 Research / reviewers in the wild / expert
Benjamin Swanson
dblp:117/3984 · also Ben Swanson
· DBLP profile ↗
6ranked-venue papers
4as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 50% Machine translation · 25% Transfer learning and domain adaptation · 25% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › sequence labeling
binary sequence labeling |
0.8 | 1 | 2024 | BinaryAlign: Word Alignment as Binary Sequence Labeling · ACL (1) 2024 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.8 | 1 | 2024 | Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection |
0.8 | 1 | 2024 | Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection · EMNLP 2024 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.8 | 1 | 2024 | BinaryAlign: Word Alignment as Binary Sequence Labeling · ACL (1) 2024 |
Automata and formal languages
tree adjoining grammar |
0.2 | 1 | 2013 | A Context Free TAG Variant · ACL (1) 2013 |
Programming languages and type systems
grammar formalisms |
0.0 | 1 | 2013 | A Context Free TAG Variant · ACL (1) 2013 |
Methods — techniques the papers use, named apart from their topics
zero-shot cross-lingual transfer · 0.8synthetic data generation · 0.8multilingual foundation models · 0.8fine-tuning · 0.8formal language theory · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BinaryAlign: Word Alignment as Binary Sequence LabelingabstractReal world deployments of word alignment are almost certain to cover both high and low resource languages.However, the state-ofthe-art for this task recommends a different model class depending on the availability of gold alignment training data for a particular language pair.We propose BinaryAlign, a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios, offering a unifying approach to the task.Additionally, we vary the specific choice of multilingual foundation model, perform stratified error analysis over alignment error type, and explore the performance of BinaryAlign on non-English language pairs.We make our source code publicly available.1 Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson |
ACL (1) | 3 |
| 2024 | Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error DetectionabstractGrammatical Error Detection (GED) methods rely heavily on human annotated error corpora.However, these annotations are unavailable in many low-resource languages.In this paper, we investigate GED in this context.Leveraging the zero-shot cross-lingual transfer capabilities of multilingual pre-trained language models, we train a model using data from a diverse set of languages to generate synthetic errors in other languages.These synthetic error corpora are then used to train a GED model.Specifically we propose a two-stage fine-tuning pipeline where the GED model is first fine-tuned on multilingual synthetic data from target languages followed by fine-tuning on human-annotated GED corpora from source languages.This approach outperforms current state-of-the-art annotation-free GED methods.We also analyse the errors produced by our method and other strong baselines, finding that our approach produces errors that are more diverse and more similar to human errors. Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson |
EMNLP | 3 |
| 2014 | Data Driven Language Transfer HypothesesabstractLanguage transfer, the preferential second language behavior caused by similarities to the speaker's native language, requires considerable expertise to be detected by humans alone.Our goal in this work is to replace expert intervention by data-driven methods wherever possible.We define a computational methodology that produces a concise list of lexicalized syntactic patterns that are controlled for redundancy and ranked by relevancy to language transfer.We demonstrate the ability of our methodology to detect hundreds of such candidate patterns from currently available data sources, and validate the quality of the proposed patterns through classification experiments. Benjamin Swanson, Eugene Charniak |
EACL | 1 |
| 2013 | A Context Free TAG Variant
Benjamin Swanson, Elif Yamangil, Eugene Charniak, Stuart M. Shieber |
ACL (1) | 1 |
| 2013 | Extracting the Native Language Signal for Second Language Acquisition
Benjamin Swanson, Eugene Charniak |
HLT-NAACL | 1 |
| 2012 | Correction Detection and Error Type Selection as an ESL Educational Aid
Benjamin Swanson, Elif Yamangil |
HLT-NAACL | 1 |