Siyu Lai

dblp:75/10303 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Machine translation · 50% Language models and text generation · 50%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model
0.612022
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment · EMNLP 2022
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.612022
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

two-stage training · 0.6translation language modeling · 0.6self-attention · 0.6cross-attention · 0.6
YearPublicationVenuePosition
2026 Research on Core Technology Identification Methods Based on High-Order Dependency Metrics
Siyu Lai, Huijun Zheng, Longyun Wang, Wenchuan Yang, Xin Lu 0002
DATA (1)2
2022 Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment
abstract
Word alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing.Recent studies in this area have yielded substantial improvements by generating alignments from contextualized embeddings of the pre-trained multilingual language models.However, we find that the existing approaches capture few interactions between the input sentence pairs, which degrades the word alignment quality severely, especially for the ambiguous words in the monolingual context.To remedy this problem, we propose Cross-Align to model deep interactions between the input sentence pairs, in which the source and target sentences are encoded separately with the shared self-attention modules in the shallow layers, while cross-lingual interactions are explicitly constructed by the crossattention modules in the upper layers.Besides, to train our model effectively, we propose a two-stage training framework, where the model is trained with a simple Translation Language Modeling (TLM) objective in the first stage and then finetuned with a self-supervised alignment objective in the second stage.Experiments show that the proposed Cross-Align achieves the state-of-the-art (SOTA) performance on four out of five language pairs. 1
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP1
2022 Generating Authentic Adversarial Examples beyond Meaning-preserving with Doubly Round-trip Translation
abstract
Siyu Lai, Zhen Yang, Fandong Meng, Xue Zhang, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
NAACL-HLT1