Benjamin Swanson

dblp:117/3984 · also Ben Swanson · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 50% Machine translation · 25% Transfer learning and domain adaptation · 25%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › sequence labeling
binary sequence labeling
0.812024
BinaryAlign: Word Alignment as Binary Sequence Labeling · ACL (1) 2024
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.812024
Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection · EMNLP 2024
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection
0.812024
Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection · EMNLP 2024
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.812024
BinaryAlign: Word Alignment as Binary Sequence Labeling · ACL (1) 2024
Automata and formal languages
tree adjoining grammar
0.212013
A Context Free TAG Variant · ACL (1) 2013
Programming languages and type systems
grammar formalisms
0.012013
A Context Free TAG Variant · ACL (1) 2013

Methods — techniques the papers use, named apart from their topics

zero-shot cross-lingual transfer · 0.8synthetic data generation · 0.8multilingual foundation models · 0.8fine-tuning · 0.8formal language theory · 0.3
YearPublicationVenuePosition
2024 BinaryAlign: Word Alignment as Binary Sequence Labeling
abstract
Real world deployments of word alignment are almost certain to cover both high and low resource languages.However, the state-ofthe-art for this task recommends a different model class depending on the availability of gold alignment training data for a particular language pair.We propose BinaryAlign, a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios, offering a unifying approach to the task.Additionally, we vary the specific choice of multilingual foundation model, perform stratified error analysis over alignment error type, and explore the performance of BinaryAlign on non-English language pairs.We make our source code publicly available.1
Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson
ACL (1)3
2024 Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection
abstract
Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora.However, these annotations are unavailable in many low-resource languages.In this paper, we investigate GED in this context.Leveraging the zero-shot cross-lingual transfer capabilities of multilingual pre-trained language models, we train a model using data from a diverse set of languages to generate synthetic errors in other languages.These synthetic error corpora are then used to train a GED model.Specifically we propose a two-stage fine-tuning pipeline where the GED model is first fine-tuned on multilingual synthetic data from target languages followed by fine-tuning on human-annotated GED corpora from source languages.This approach outperforms current state-of-the-art annotation-free GED methods.We also analyse the errors produced by our method and other strong baselines, finding that our approach produces errors that are more diverse and more similar to human errors.
Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson
EMNLP3
2014 Data Driven Language Transfer Hypotheses
abstract
Language transfer, the preferential second language behavior caused by similarities to the speaker's native language, requires considerable expertise to be detected by humans alone.Our goal in this work is to replace expert intervention by data-driven methods wherever possible.We define a computational methodology that produces a concise list of lexicalized syntactic patterns that are controlled for redundancy and ranked by relevancy to language transfer.We demonstrate the ability of our methodology to detect hundreds of such candidate patterns from currently available data sources, and validate the quality of the proposed patterns through classification experiments.
Benjamin Swanson, Eugene Charniak
EACL1
2013 A Context Free TAG Variant
Benjamin Swanson, Elif Yamangil, Eugene Charniak, Stuart M. Shieber
ACL (1)1
2013 Extracting the Native Language Signal for Second Language Acquisition
Benjamin Swanson, Eugene Charniak
HLT-NAACL1
2012 Correction Detection and Error Type Selection as an ESL Educational Aid
Benjamin Swanson, Elif Yamangil
HLT-NAACL1