Yuwei Yin

dblp:297/8984 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2025 McEval: Massively Multilingual Code Evaluation
abstract
Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval.
Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001
ICLR4
2025 SWI: Speaking with Intent in Large Language Models
abstract
Intent, typically clearly formulated and planned, functions as a cognitive framework for communication and problem-solving. This paper introduces the concept of Speaking with Intent (SWI) in large language models (LLMs), where the explicitly generated intent encapsulates the model’s underlying intention and provides high-level planning to guide subsequent analysis and action. By emulating deliberate and purposeful thoughts in the human mind, SWI is hypothesized to enhance the reasoning capabilities and generation quality of LLMs. Extensive experiments on text summarization, multi-task question answering, and mathematical reasoning benchmarks consistently demonstrate the effectiveness and generalizability of Speaking with Intent over direct generation without explicit intent. Further analysis corroborates the generalizability of SWI under different experimental settings. Moreover, human evaluations verify the coherence, effectiveness, and interpretability of the intent produced by SWI. The promising results in enhancing LLMs with explicit intents pave a new avenue for boosting LLMs’ generation and reasoning abilities with cognitive notions.
Yuwei Yin, Eunjeong Hwang, Giuseppe Carenini
INLG1
2024 UniCoder: Scaling Code Large Language Model via Universal Code
abstract
Tao Sun, Linzheng Chai, Jian Yang, Yuwei Yin, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, Zhoujun Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tao Sun 0016, Linzheng Chai, Jian Yang 0030, Yuwei Yin, Hongcheng Guo, Liqun Yang, Zhoujun Li 0001
ACL (1)4
2024 m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt
abstract
Multilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario.
Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001
LREC/COLING3
2024 A Word and Local Feature Aware Network for Chinese Spoken Language Understanding
abstract
Chinese SLU suffers from unclear word boundaries. Previous works were either missing the word information or afflicted by word segmentation errors. In this paper, we propose a novel Word and Local feature Aware Network (WLAN), which incorporates both word information and local features for Chinese SLU while avoiding word segmentation. In addition, in our knowledge we first time use the potential slots likely to co-occur with the predicted intent as specific guidance. Experimental results on two Chinese SLU datasets show that our model gets very competitive performance.
Yuwei Yin, Shuang-Hua Yang
IJCNN2
2024 Dialogue Discourse Parsing as Generation: A Sequence-to-Sequence LLM-based Approach
abstract
Discourse analysis studies the sentence organization within a document, aiming to reveal its underlying structural information.Existing works on dialogue discourse parsing mostly use encoder-only models and sophisticated decoding strategies to extract structures.Despite recent advances in Large Language Models (LLMs), applying directly these models on discourse parsing is challenging.To fully leverage the rich semantic and discourse knowledge in LLMs, we propose to transform discourse parsing into a generation task using a text-to-text paradigm.Our approach is intuitive and requires no modification of the LLM architecture.Experimental results on STAC and Molweni datasets show that a sequence-tosequence model such as T0 can perform reasonably well.Notably, our improved transitionbased sequence-to-sequence system achieves new state-of-the-art performance on Molweni.Furthermore, our systems can generate richer discourse structures such as graphs, whereas previous methods are mostly limited to trees. 1
Chuyuan Li, Yuwei Yin, Giuseppe Carenini
SIGDIAL2
2024 mt4CrossOIE: Multi-stage tuning for cross-lingual open information extraction
Tongliang Li, Linzheng Chai, Jian Yang 0030, Jiaqi Bai 0001, Yuwei Yin, Hongcheng Guo, Liqun Yang, Hebboul Zine El Abidine, Zhoujun Li 0001
Expert Syst. Appl.6
2024 Dark-DSAR: Lightweight one-step pipeline for action recognition in dark videos
Yuwei Yin, Renjie Yang, Yuanzhong Liu, Zhigang Tu 0001
Neural Networks1
2023 GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator
abstract
Jian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001
ACL (1)6
2023 HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei
DASFAA (3)2
2023 GTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation
abstract
Transformer structure, stacked by a sequence of encoder and decoder network layers, achieves significant development in neural machine translation. However, vanilla Transformer mainly exploits the top-layer representation, assuming the lower layers provide trivial or redundant information and thus ignoring the bottom-layer feature that is potentially valuable. In this work, we propose theGroup-Transformer model (GTrans) that flexibly divides multi-layer representations of both encoder and decoder into different groups and then fuses these group features to generate target words. To corroborate the effectiveness of the proposed method, extensive experiments and analytic experiments are conducted on three bilingual translation benchmarks and three multilingual translation tasks, including the IWLST-14, IWLST-17, LDC, WMT-14, WMT-21 and OPUS-100 benchmark. Experimental and analytical results demonstrate that our model outperforms its Transformer counterparts by a consistent gain. Furthermore, it can be successfully scaled up to 60 encoder layers and 36 decoder layers.
Jian Yang 0030, Yuwei Yin, Liqun Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Furu Wei, Zhoujun Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Exploring Entity Interactions for Few-Shot Relation Learning (Student Abstract)
abstract
Few-shot relation learning refers to infer facts for relations with a few observed triples. Existing metric-learning methods mostly neglect entity interactions within and between triples. In this paper, we explore this kind of fine-grained semantic meaning and propose our model TransAM. Specifically, we serialize reference entities and query entities into sequence and apply transformer structure with local-global attention to capture intra- and inter-triple entity interactions. Experiments on two public datasets with 1-shot setting prove the effectiveness of TransAM.
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
AAAI4
2022 High-resource Language-specific Training for Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation (MNMT) trained in multiple language pairs has attracted considerable attention due to fewer model parameters and lower training costs by sharing knowledge among multiple languages. Nonetheless, multilingual training is plagued by language interference degeneration in shared parameters because of the negative interference among different translation directions, especially on high-resource languages. In this paper, we propose the multilingual translation model with the high-resource language-specific training (HLT-MT) to alleviate the negative interference, which adopts the two-stage training with the language-specific selection mechanism. Specifically, we first train the multilingual model only with the high-resource pairs and select the language-specific modules at the top of the decoder to enhance the translation quality of high-resource directions. Next, the model is further trained on all available corpora to transfer knowledge from high-resource languages (HRLs) to low-resource languages (LRLs). Experimental results show that HLT-MT outperforms various strong baselines on WMT-10 and OPUS-100 benchmarks. Furthermore, the analytic experiments validate the effectiveness of our method in mitigating the negative interference in multilingual training.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Furu Wei
IJCAI2
2022 UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine Translation
abstract
Most translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enables one-pass translation using shared semantic space for all languages compared to the two-pass pivot translation but often underperforms the pivot-based method. In this paper, we propose a novel method, named as Unified Multilingual Multiple teacher-student Model for NMT (UM4). Our method unifies source-teacher, target-teacher, and pivot-teacher models to guide the student model for the zero-resource translation. The source teacher and target teacher force the student to learn the direct source-target translation by the distilled knowledge on both source and target sides. The monolingual corpus is further leveraged by the pivot-teacher model to enhance the student model. Experimental results demonstrate that our model of 72 directions significantly outperforms previous methods on the WMT benchmark.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Hongcheng Guo, Zhoujun Li 0001, Furu Wei
IJCAI2
2022 Tackling Solitary Entities for Few-Shot Knowledge Graph Completion
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
KSEM (1)4
2022 Explore Modeling Relation Information and Direction Information in KBQA
Shuai Zhao 0001, Bo Cheng 0001, Yuwei Yin, Hao Yang 0006
Neurocomputing4
2022 Toward Tweet Entity Linking With Heterogeneous Information Networks
abstract
Twitter, a microblogging platform, has developed into an increasingly invaluable information source, where millions of users post a great quantity of tweets with various topics per day. Heterogeneous information networks consisting of multi-type objects and relations are becoming more and more prevalent as an organization form of knowledge and information. The task of linking an entity mention in a tweet with its corresponding entity in a heterogeneous information network is of great importance, for the purpose of enriching heterogeneous information networks with the abundant and fresh knowledge embedded in tweets. However, the entity mention is ambiguous. Additionally, tweets are short and informal, making it difficult to mine enough information from a single tweet for entity linking. In this paper, we propose an unsupervised iterative clustering framework TELHIN to link multiple similar tweets with a heterogeneous information network jointly. Our framework takes three dimensions of tweet similarity into consideration: (1) content similarity, (2) temporal similarity, and (3) user similarity. The appropriate weights of different similarity dimensions for each entity mention are learned iteratively based on the metric learning algorithm by leveraging the pairwise constraints generated automatically. Experiments on real data demonstrate the effectiveness of our framework in comparison with the baselines.
Wei Shen 0004, Yuwei Yin, Yang Yang 0008, Jiawei Han 0001, Jianyong Wang 0001, Xiaojie Yuan
IEEE Trans. Knowl. Data Eng.2