Yingqiang Gao

dblp:264/0145 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 36% Machine translation · 31% Information extraction and text analysis · 31%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
SwiLTra-Bench: The Swiss Legal Translation Benchmark · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
lexical semantics
0.912025
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
SwiLTra-Bench: The Swiss Legal Translation Benchmark · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
text segmentation
0.712023
GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information · EMNLP 2023
Natural language and speech › Machine translation › neural machine translation
character-level translation
0.412020
Character-Level Translation with Self-attention · ACL 2020
Natural language and speech › Machine translation
neural machine translation
0.412020
Character-Level Translation with Self-attention · ACL 2020
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.312025
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › text similarity
semantic similarity
0.212023
GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information · EMNLP 2023
Machine learning › Deep learning architectures and training
transformer
0.112020
Character-Level Translation with Self-attention · ACL 2020

Methods — techniques the papers use, named apart from their topics

zero-shot prompting · 0.9fine-tuning · 0.9contrastive learning · 0.9normalized mutual information · 0.7greedy optimization · 0.7self-attention · 0.4convolution · 0.4
YearPublicationVenuePosition
2025 ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords
abstract
Sina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Sina Ahmadi, Micha David Hess, Elena Álvarez Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich
ACL (1)7
2025 SwiLTra-Bench: The Swiss Legal Translation Benchmark
abstract
In Switzerland legal translation is uniquely important due to the country’s four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.
Joel Niklaus, Jakob Merane, Luka Nenadic, Sina Ahmadi, Yingqiang Gao, Cyrill A. H. Chevalley, Claude Humbel, Christophe Gösken, Lorenzo Tanzi, Thomas Lüthi, Stefan Palombo, Spencer Poff, Boling Yang, Matthew Guillod, Robin Mamié, Daniel Brunner, Julio Pereyra, Niko Grupen
ACL (1)5
2023 GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information
abstract
The abstracts of scientific papers typically contain both premises (e.g., background and observations) and conclusions.Although conclusion sentences are highlighted in structured abstracts, in non-structured abstracts the concluding information is not explicitly marked, which makes the automatic segmentation of conclusions from scientific abstracts a challenging task.In this work, we explore Normalized Mutual Information (NMI) as a means for abstract segmentation.We consider each abstract as a recurrent cycle of sentences and place two segmentation boundaries by greedily optimizing the NMI score between the two segments, assuming that conclusions are strongly semantically linked with preceding premises.On nonstructured abstracts, our proposed unsupervised approach GreedyCAS achieves the best performance across all evaluation metrics; on structured abstracts, GreedyCAS outperforms all baseline methods measured by P k .The strong correlation of NMI to our evaluation metrics reveals the effectiveness of NMI for abstract segmentation.1
Yingqiang Gao, Jessica Lam, Nianlong Gu, Richard H. R. Hahnloser
EMNLP1
2022 Local Citation Recommendation with Hierarchical-Attention Text Encoder and SciBERT-Based Reranking
Nianlong Gu, Yingqiang Gao, Richard H. R. Hahnloser
ECIR (1)2
2020 Character-Level Translation with Self-attention
abstract
We explore the suitability of self-attention models for character-level neural machine translation.We test the standard transformer model, as well as a novel variant in which the encoder block combines information from nearby characters using convolutions.We perform extensive experiments on WMT and UN datasets, testing both bilingual and multilingual translation to English using up to three input languages (French, Spanish, and Chinese).Our transformer variant consistently outperforms the standard transformer at the character-level and converges faster while learning more robust character-level alignments. 1
Yingqiang Gao, Nikola I. Nikolov, Yuhuang Hu, Richard H. R. Hahnloser
ACL1