Yutao Sun

dblp:01/9758 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Block Length Identification of Block Ciphers Based on Bit-Level Markov Mapping
Rongna Xie, Chunlin Zou, Yutao Sun, Zezheng Gao
ICIC (9)4
2025 FocusLLM: Precise Understanding of Long Context by Dynamic Condensing
abstract
Zhenyu Li, Yike Zhang, Tengyu Pan, Yutao Sun, Zhichao Duan, Junjie Fang, Rong Han, Zixuan Wang, Jianyong Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhenyu Li 0008, Tengyu Pan, Yutao Sun, Zhichao Duan 0001, Junjie Fang, Jianyong Wang 0001
ACL (1)4
2025 StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He, Yujie Hou, Xiaoyang Xing, Huiyan Li, Bing Han 0008, Yanmin Qian
ASRU3
2025 ISO 26262-Aligned Functional Safety Verification Framework with Explainable Graph Neural Network
abstract
The growing complexity and integration of automotive electronic systems, driven by advancements in intelligent vehicles and autonomous driving, make functional safety (FuSa) verification critical for ensuring system reliability. Traditional fault injection (FI) methods face inefficiencies, scalability limitations, and interpretability gaps, hard to meet the stringent requirements of ISO 26262 standards for safety-critical automotive systems. This paper proposes an explainable Graph Neural Network(GNN)-based framework for FuSa verification in automotive electronics through three core contributions: graph neural networks for modeling circuit structures to identify critical fault nodes, gradient-driven feature importance analysis to optimize selective hardening with minimal resource overhead and GNNExplainer to visualize critical nodes and connections driving fault-criticality predictions through subgraph analysis. Validated across diverse circuits the framework achieves up to 99.6% precision and 99.8% F1-score in fault detection while significantly reducing the simulation time. Notably, the framework improves accuracy by approximately 5% compared to state-of-the-art (SOTA) methods while requiring only half the fault injection data. Integrated explainable artificial intelligence techniques provide transparent decision traces, ensuring compliance with ISO 26262 traceability requirements. Through feature selection, the framework achieves comparable accuracy with a minimal feature set, significantly reducing computational overhead while maintaining performance. By bridging AI-driven automation with rigorous safety certification, this work establishes a scalable, efficient, and interpretable solution for FuSa verification in automotive SoCs.
Yutao Sun, Jiehua Huang, Xiangping Liao, Liping Liang 0001
ICCAD1
2025 Differential Transformer
abstract
Transformer tends to overallocate attention to irrelevant context. In this work, we introduce Diff Transformer, which amplifies attention to the relevant context while canceling noise. Specifically, the differential attention mechanism calculates attention scores as the difference between two separate softmax attention maps. The subtraction cancels noise, promoting the emergence of sparse attention patterns. Experimental results on language modeling show that Diff Transformer outperforms Transformer in various settings of scaling up model size and training tokens. More intriguingly, it offers notable advantages in practical applications, such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers. By being less distracted by irrelevant context, Diff Transformer can mitigate hallucination in question answering and text summarization. For in-context learning, Diff Transformer not only enhances accuracy but is also more robust to order permutation, which was considered as a chronic robustness issue. The results position Diff Transformer as a highly effective and promising architecture for large language models.
Tianzhu Ye, Li Dong 0004, Yuqing Xia, Yutao Sun, Gao Huang 0001, Furu Wei
ICLR4
2025 Horae: A Domain-Agnostic Language for Automated Service Regulation
abstract
Artificial intelligence is rapidly encroaching on the field of service regulation. However, existing AI-based regulation techniques are often tailored to specific application domains and thus are difficult to generalize in an automated manner. This paper presents Horae, a unified specification language for modeling (multimodal) regulation rules across a diverse set of domains. We showcase how Horae facilitates an intelligent service regulation pipeline by further exploiting a fine-tuned large language model named RuleGPT that automates the Horae modeling process, thereby yielding an end-to-end framework for fully automated intelligent service regulation. The feasibility and effectiveness of our framework are demonstrated over a benchmark of various real-world regulation domains. In particular, we show that our open-sourced, fine-tuned RuleGPT with 7B parameters suffices to outperform GPT-3.5 and perform on par with GPT-4o.
Yutao Sun, Mingshuai Chen, Kangjia Zhao, Jintao Chen 0001, Zhongyi Wang 0004, Liqiang Lu, Xinkui Zhao, Shuiguang Deng, Jianwei Yin
IJCAI1
2025 Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
Chenyang Le, Yinfeng Xia, Huiyan Li, Manhong Wang, Yutao Sun, Xingyang Ma, Yanmin Qian
INTERSPEECH5
2025 MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
Yinfeng Xia, Huiyan Li, Chenyang Le, Manhong Wang, Yutao Sun, Xingyang Ma, Yanmin Qian
INTERSPEECH5
2025 Graph Attention Networks Based Fault Prediction Framework for Functional Safety Verification
abstract
As automotive chips grow in complexity, the cost of functional safety(FuSa) verification rises sharply. This paper proposes a Graph Attention Network (GAT) based framework that extracts fault propagation features and iteratively identifies critical nodes prone to single-point failures under ISO 26262 guidance. Tested on five open-source circuits, the framework achieves 97.89% accuracy and 98.42% F1score. Leave-one-out cross-validation yields 97.12% accuracy and 96.86% F1-score, demonstrating strong generalization from small to large-scale circuits. Compared to traditional RTL fault injection and neural network methods, it reduces fault simulation time and training data usage by about half. The framework outperforms existing techniques in both accuracy and efficiency, offering a practical solution for scalable automotive circuit safety verification.
Yutao Sun, Jiehua Huang, Xiangping Liao, Liping Liang 0001
ITC1
2024 An FPGA-Based Emulation Platform for Functional Safety Verification in Automotive SoC Systems
abstract
The increasing complexity of automotive SoC systems, particularly in the context of autonomous vehicles, demands rigorous functional safety verification methods. This paper proposes a novel Functional Safety FPGA Fault Injection Tool (FSF-FIT) which builds an FPGA-based emulation platform by using FPGA fault injection techniques. The core innovation of FSFFIT lies in its comprehensive FPGA functional safety evaluation process, targeting specific components within the Design Under Test (DUT), down to individual LUTs or registers. This detailed approach allows for accurate reliability level (RL) and diagnostic coverage (DC) assessments of safety mechanisms. The speed and accuracy of FSFFIT were validated through experiments conducted on the XuanTie C906 RISC-V processor as well as on a multiplier with different security mechanisms. The results demonstrated that FSFFIT’s performance is consistent with that of SSIM, a certified ISO26262-compliant functional safety tool, while also achieving faster execution times. Additionally, by comparing the results with some state-of-the-art (SOTA) FPGA fault injection tools, our method is superior in terms of injection speed.
Yutao Sun, Zean Huang, Liping Liang 0001
ATS1
2024 Fine-Grained Legal Argument-Pair Extraction via Coarse-Grained Pre-training
abstract
Legal Argument-Pair Extraction (LAE) is dedicated to the identification of interactive arguments targeting the same subject matter within legal complaints and corresponding defenses. This process serves as a foundation for automatically recognizing the focal points of disputes. Current methodologies predominantly conceptualize LAE as a supervised sentence-pair classification problem and usually necessitate extensive manual annotations, thereby constraining their scalability and general applicability. To this end, we present an innovative approach to LAE that focuses on fine-grained alignment of argument pairs, building upon coarse-grained complaint-defense pairs. This strategy stems from two key observations: 1) In general, every argument presented in a legal complaint is likely to be addressed by at least one corresponding argument in the defense. 2) It’s rare for multiple complaint arguments to be addressed by a single defense argument; rather, each complaint argument usually corresponds to a unique defense argument. Motivated by these insights, we develop a specialized pre-training framework. Our model employs pre-training objectives designed to exploit the coarse-grained supervision signals. This enables expressive representations of legal arguments for LAE, even when working with a limited amount of labeled data. To verify the effectiveness of our model, we construct the largest LAE datasets from two representative causes, private lending, and contract dispute. The experimental results demonstrate that our model can effectively capture informative argument knowledge from unlabeled complaint-defense pairs and outperform the unsupervised and supervised baselines by 3.7 and 2.4 points on average respectively. Besides, our model can reach superior accuracy with only half manually annotated data. The datasets and code can be found in https://github.com/thunlp/LAE.
Chaojun Xiao, Yutao Sun, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001
LREC/COLING2
2024 Horae: A Domain-Agnostic Modeling Language for Automating Multimodal Service Regulation⋆
abstract
Artificial intelligence is rapidly encroaching on the field of service regulation. This work-in-progress article presents the design principles behind Horae, a unified specification language to model multimodal regulation rules across a diverse set of domains. We show how Horae facilitates an intelligent service regulation pipeline by further exploiting a fine-tuned large language model named RuleGPT that automates the Horae modeling process, thereby yielding an end-to-end framework for fully automated intelligent service regulation.
Yutao Sun, Mingshuai Chen, Kangjia Zhao, Jintao Chen 0001
ICWS1
2024 You Only Cache Once: Decoder-Decoder Architectures for Language Models
abstract
We introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes.
Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei
NeurIPS1
2023 A Length-Extrapolatable Transformer
abstract
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yutao Sun, Li Dong 0004, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Furu Wei
ACL (1)1
2023 Prototypical Calibration for Few-shot Learning of Language Models
Zhixiong Han, Yaru Hao, Li Dong 0004, Yutao Sun, Furu Wei
ICLR4