Yanlin Tang

dblp:134/9034 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 64% Program synthesis and code generation · 36%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › code generation
assembly code generation
0.912025
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code · NeurIPS 2025
Program synthesis and code generation
code generation with language models
0.912025
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code · NeurIPS 2025
Compilers and program optimization › parallelization
automatic parallelization
0.612022
BabelTower: Learning to Auto-parallelized Program Translation · ICML 2022
Program synthesis and code generation
code translation
0.612022
BabelTower: Learning to Auto-parallelized Program Translation · ICML 2022
Compilers and program optimization
compiler construction
0.312025
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

self-evolving prompt optimization · 0.9self-debugging · 0.9large language model · 0.9neural machine translation · 0.6discriminative reranker · 0.6back-translation · 0.6
YearPublicationVenuePosition
2025 QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
abstract
Compilers, while essential, are notoriously complex systems that demand prohibitively expensive human expertise to develop and maintain. The recent advancements in Large Language Models (LLMs) offer a compelling new paradigm: Neural Compilation, which could potentially simplify compiler development for new architectures and facilitate the discovery of innovative optimization techniques. However, several critical obstacles impede its practical adoption. Firstly, a significant lack of dedicated benchmarks and robust evaluation methodologies hinders objective assessment and tracking of progress in the field. Secondly, systematically enhancing the reliability and performance of LLM-generated assembly remains a critical challenge. Addressing these challenges, this paper introduces NeuComBack, a novel benchmark dataset specifically designed for IR-to-assembly compilation. Leveraging this dataset, we first define a foundational Neural Compilation workflow and conduct a comprehensive evaluation of the capabilities of recent frontier LLMs on Neural Compilation, establishing new performance baselines. We further propose a self-evolving prompt optimization method that enables LLMs to iteratively evolve their internal prompt strategies by extracting insights from prior self-debugging traces, thereby enhancing their neural compilation capabilities. Experiments demonstrate that our method significantly improves both the functional correctness and the performance of LLM-generated assembly code. Compared to baseline prompts, the functional correctness rates improved from 44% to 64% on x86_64 and from 36% to 58% on aarch64, respectively. More significantly, among the 16 correctly generated x86_64 programs using our method, 14 (87.5%) surpassed clang-O3 performance. These consistent improvements across diverse architectures (x86_64 and aarch64) and program distributions (NeuComBack L1 and L2) validate our method's superiority over conventional approaches and its potential for broader adoption in low-level neural compilation.
Hainan Fang, Yuanbo Wen 0001, Jun Bi, Tonghui He, Yanlin Tang, Jiaming Guo, Rui Zhang 0040, Qi Guo 0001, Yunji Chen
NeurIPS6
2022 BabelTower: Learning to Auto-parallelized Program Translation
abstract
GPUs have become the dominant computing platforms for many applications, while programming GPUs with the widely-used CUDA parallel programming model is difficult. As sequential C code is relatively easy to obtain either from legacy repositories or by manual implementation, automatically translating C to its parallel CUDA counterpart is promising to relieve the burden of GPU programming. However, because of huge differences between the sequential C and the parallel CUDA programming model, existing approaches fail to conduct the challenging auto-parallelized program translation. In this paper, we propose a learning-based framework, i.e., BabelTower, to address this problem. We first create a large-scale dataset consisting of compute-intensive function-level monolingual corpora. We further propose using back-translation with a discriminative reranker to cope with unpaired corpora and parallel semantic conversion. Experimental results show that BabelTower outperforms state-of-the-art by 1.79, 6.09, and 9.39 in terms of BLEU, CodeBLEU, and specifically designed ParaBLEU, respectively. The CUDA code generated by BabelTower attains a speedup of up to 347x over the sequential C code, and the developer productivity is improved by at most 3.8x.
Yuanbo Wen 0001, Qi Guo 0001, Xiaqing Li, Jianxing Xu, Yanlin Tang, Yongwei Zhao 0001, Xing Hu 0001, Zidong Du, Ling Li 0001, Chao Wang 0003, Xuehai Zhou, Yunji Chen
ICML6
2016 Robust Quantile Analysis for Accelerated Life Test Data
abstract
We propose a quantile regression framework to model accelerated life tests (ALT) data. The quantile of the failure time distribution at the usage level can be easily estimated using quantile regression. Compared with traditional parametric regression methods, quantile regression is distribution-free, efficient in the presence of censoring, and more flexible in modeling ALT relations. More importantly, we show that it is able to handle ALT data with a failure-free life, which is a great challenge in the ALT literature. We use extensive simulation studies and two real ALT case studies to demonstrate the effectiveness of the proposed method.
Nan Chen 0002, Yanlin Tang, Zhisheng Ye 0001
IEEE Trans. Reliab.2
2015 Partition-based range query for uncertain trajectories in road networks
Ling Chen 0001, Yanlin Tang, Mingqi Lv, Gencai Chen
GeoInformatica2