Yichen Dong

dblp:378/9181 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 57% Language models and text generation · 33% Efficient and distributed learning · 10%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
preference optimization
1.012026
M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation · ACL (1) 2026
Cloud and datacenter computing
cluster resource management and scheduling
1.012026
Resource-Aware Distributed Training Job Placement for GPU Cluster Defragmentation · IEEE Trans. Netw. 2026
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
GPU cluster scheduling
1.012026
Resource-Aware Distributed Training Job Placement for GPU Cluster Defragmentation · IEEE Trans. Netw. 2026
Cloud and datacenter computing › job scheduling
job placement
1.012026
Resource-Aware Distributed Training Job Placement for GPU Cluster Defragmentation · IEEE Trans. Netw. 2026
Natural language and speech › Machine translation
document-level machine translation
0.912025
Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement · ACL (1) 2025
Natural language and speech › Machine translation
translation refinement
0.912025
Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement · ACL (1) 2025
Machine learning › Efficient and distributed learning
distributed training
0.312026
Resource-Aware Distributed Training Job Placement for GPU Cluster Defragmentation · IEEE Trans. Netw. 2026

Methods — techniques the papers use, named apart from their topics

submodular optimization · 2.0NP-hardness proof · 2.0multi-perspective preference optimization · 1.0multi-pair preference optimization · 1.0greedy approximation algorithms · 1.0greedy approximation algorithm · 1.0self-refinement · 0.9fine-tuning · 0.9
YearPublicationVenuePosition
2026 M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
abstract
Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linlong Xu, Liangying Shao, Yichen Dong, Xinwei Wu 0001, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo
ACL (1)8
2026 Resource-Aware Distributed Training Job Placement for GPU Cluster Defragmentation
abstract
Distributed training (DT) has emerged as a solution to address the growing computational resource demands of training large-scale machine learning models. To meet this need, cloud providers typically build GPU clusters to accommodate DT jobs. For DT job requests, cloud providers need to determine in which GPUs place workers (i.e., job placement). Existing approaches usually place workers on as few idle machines as possible to minimize communication time. However, this scheme will lead to aresource fragmentation problem, which degrades the resource utilization rate of the GPU cluster and increases training costs for cloud providers. In this paper, we propose$\textsf {Titan}$, a novel job placement scheme that mitigates the influence of resource fragmentation by enhancing the utilization of non-idle machines. To further optimize resource allocation, we introduce a dynamic defragmentation algorithm that migrates fragmented jobs to consolidate GPU resources, enabling efficient placement of large-scale training jobs.$\textsf {Titan}$formulates a multi-objective non-linear optimization problem and proves its NP-hardness. To solve this problem,$\textsf {Titan}$presents an effective submodular-based greedy algorithm with a tight approximation ratio ($1-\frac {1}{e}$). We evaluate$\textsf {Titan}$with a large-scale simulation employing real-world job traces and a small-scale testbed consisting of 8 servers with 32 logical GPUs. Experimental results show that$\textsf {Titan}$can achieve near-optimal training throughput while improving the efficiency of the cluster by 74.9% compared to the state-of-the-art solutions.
Gongming Zhao, Yichen Dong, Hongli Xu 0001, Baoyi An 0002, Gangyi Luo
IEEE Trans. Netw.2
2025 Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
abstract
Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub.
Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ACL (1)1