Yue Fang 0001

dblp:92/4710-1 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Logic in computer science · 70% Automated reasoning and model checking · 30%
Artificial intelligence
1 paper
Reinforcement learning · 62% Language models and text generation · 38%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation · AAAI 2026
Logic in computer science
formal specification
1.012026
RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation · AAAI 2026
Logic in computer science › temporal logic
signal temporal logic
1.012026
RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation · AAAI 2026
Program synthesis and code generation
code generation from natural language
0.912025
NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learning · EMNLP 2025
Automated reasoning and model checking › theorem proving
interactive theorem proving
0.912025
NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learning · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312026
RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation · AAAI 2026
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.312026
RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation · AAAI 2026

Methods — techniques the papers use, named apart from their topics

reward modeling · 2.0reinforcement learning · 2.0proximal policy optimization · 2.0curriculum learning · 2.0multi-aspect reinforcement learning · 1.7
YearPublicationVenuePosition
2026 RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation
abstract
Signal Temporal Logic (STL) is a powerful formal language for specifying real-time specifications of Cyber-Physical Systems (CPS). Transforming specifications written in natural language into STL formulas automatically has attracted increasing attention. Existing rule-based methods depend heavily on rigid pattern matching and domain-specific knowledge, limiting their generalizability and scalability. Recently, Supervised Fine-Tuning (SFT) of large language models (LLMs) has been successfully applied to transform natural language into STL. However, the lack of fine-grained supervision on atomic proposition correctness, semantic fidelity, and formula readability often leads SFT-based methods to produce formulas misaligned with the intended meaning. To address these issues, we propose RESTL, a reinforcement learning (RL)-based framework for the transformation from natural language to STL. RESTL introduces multiple independently trained reward models that provide fine-grained, multi-faceted feedback from four perspectives, i.e., atomic proposition consistency, semantic alignment, formula succinctness, and symbol matching. These reward models are trained with a curriculum learning strategy to improve their feedback accuracy, and their outputs are aggregated into a unified signal that guides the optimization of the STL generator via Proximal Policy Optimization (PPO). Experimental results demonstrate that RESTL significantly outperforms state-of-the-art methods in both automatic metrics and human evaluations.
Yue Fang 0001, Zhi Jin 0001, Jie An 0001, Hongshen Chen, Xiaohong Chen 0001, Naijun Zhan
AAAI1
2025 NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learning
abstract
Yue Fang, Shaohan Huang, Xin Yu, Haizhen Huang, Zihan Zhang, Weiwei Deng, Furu Wei, Feng Sun, Qi Zhang, Zhi Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yue Fang 0001, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066, Zhi Jin 0001
EMNLP1
2022 From spoken dialogue to formal summary: An utterance rewriting for dialogue summarization
abstract
Yue Fang, Hainan Zhang, Hongshen Chen, Zhuoye Ding, Bo Long, Yanyan Lan, Yanquan Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yue Fang 0001, Hainan Zhang 0001, Hongshen Chen, Zhuoye Ding, Bo Long, Yanyan Lan, Yanquan Zhou
NAACL-HLT1