Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yihan He

dblp:280/1719 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0001-1640-4355ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 24% Learning theory · 21% Optimization for machine learning · 16%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
convergence analysis
0.922024
Global Convergence in Training Large-Scale Transformers · NeurIPS 2024
Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript · ICML 2020
Natural language and speech › Language models and text generation
in-context learning
0.812024
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context · NeurIPS 2024
Machine learning › Learning theory › statistical learning theory › statistical physics of learning
mean-field analysis
0.812024
Global Convergence in Training Large-Scale Transformers · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer
transformer training
0.812024
Global Convergence in Training Large-Scale Transformers · NeurIPS 2024
Robotics › Autonomous driving
safety-critical scenario generation
0.612022
SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles · NeurIPS 2022
Machine learning › Trustworthy machine learning
safety evaluation
0.612022
SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles · NeurIPS 2022
Privacy and data protection › differential privacy
differentially private graph algorithms
0.512021
DPGraph: A Benchmark Platform for Differentially Private Graph Analysis · SIGMOD Conference 2021
Privacy and data protection
differential privacy
0.512021
DPGraph: A Benchmark Platform for Differentially Private Graph Analysis · SIGMOD Conference 2021
Privacy and data protection
privacy configuration
0.512021
DPGraph: A Benchmark Platform for Differentially Private Graph Analysis · SIGMOD Conference 2021
Machine learning › Efficient and distributed learning
distributed training
0.412020
Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript · ICML 2020
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.412020
Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript · ICML 2020
Machine learning › Efficient and distributed learning
model compression
0.412020
Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript · ICML 2020
Machine learning › Efficient and distributed learning › model compression
quantization
0.412020
Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript · ICML 2020
Machine learning › Optimization for machine learning
gradient flow
0.212024
Global Convergence in Training Large-Scale Transformers · NeurIPS 2024
Algorithms and data structures › similarity search
nearest neighbor search
0.212024
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context · NeurIPS 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.212022
SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

softmax attention · 1.5non-convex optimization · 1.5gradient descent · 1.5wasserstein gradient flow · 0.8partial differential equations · 0.8mean-field theory · 0.8deep reinforcement learning · 0.6differential privacy · 0.5weibull distribution · 0.4non-uniform quantization · 0.4
YearPublicationVenuePosition
2026 KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Gang Liao, Hongsen Qin, Alicia Golden, Michael Kuchnik, Yavuz Yetim, Ruichao Xiao, Jia Jiunn Ang, Chunli Fu, Yihan He, Samuel Hsia, Zewei Jiang, Roman Levenstein, Dianshi Li, Liyuan Li, Ajit Mathews, Varna Puvvada, Feng Shi 0001, Nathan Yan, Xiayu Yu, Uladzimir Pashkevich, Matt Steiner, Carole-Jean Wu, Gaoxiang Liu
ISCA10
2025 Few-shot radar emitter signal recognition based on prototype network with filter system
Yanping Liao, Shengwen Lin, Yihan He
J. Supercomput.3
2024 Global Convergence in Training Large-Scale Transformers
abstract
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in training Transformers with weight decay regularization. First, we construct the mean-field limit of large-scale Transformers, showing that as the model width and depth go to infinity, gradient flow converges to the Wasserstein gradient flow, which is represented by a partial differential equation. Then, we demonstrate that the gradient flow reaches a global minimum consistent with the PDE solution when the weight decay regularization parameter is sufficiently small. Our analysis is based on a series of novel mean-field techniques that adapt to Transformers. Compared with existing tools for deep networks (Lu et al., 2020) that demand homogeneity and global Lipschitz smoothness, we utilize a refined analysis assuming only $\textit{partial homogeneity}$ and $\textit{local Lipschitz smoothness}$. These new techniques may be of independent interest.
Yuan Cao 0006, Yihan He, Mengdi Wang 0001, Han Liu 0001, Jason M. Klusowski, Jianqing Fan
NeurIPS4
2024 One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
abstract
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, they are still able to solve unseen tasks well purely based on task-specific prompts. In this paper, we study the capability of one-layer transformers in learning the one-nearest neighbor prediction rule. Under a theoretical framework where the prompt contains a sequence of labeled training data and unlabeled test data, we show that, although the loss function is nonconvex, when trained with gradient descent, a single softmax attention layer can successfully learn to behave like a one-nearest neighbor classifier. Our result gives a concrete example on how transformers can be trained to implement nonparametric machine learning algorithms, and sheds light on the role of softmax attention in transformer models.
Yuan Cao 0006, Yihan He, Han Liu 0001, Jason M. Klusowski, Jianqing Fan, Mengdi Wang 0001
NeurIPS4
2024 A Universal Representation Mechanism for Multisource Remote Sensing Image
abstract
Deep learning-based processing of hyperspectral remote sensing images (HSIs) has emerged as a research hotspot. In recent years, the delivery of numerous novel HSI acquisition platforms has resulted in an exponential growth in the number of available HSI datasets from various sources. However, due to variations among hyperspectral sensors, the acquired data frequently consists of various spectral dimensions. This results in a challenge since standard deep learning approaches often require a different model for each HSI source, which impedes the construction of fundamental models for HSIs. To address this issue, we propose a unified representation mechanism for multisource HSIs that can transform spectra from numerous dimensions to a shared representation space, yielding a scalable pretraining model. Compared to existing methods, our strategy has the following advantages: 1) compatibility with HSIs of arbitrary spectral dimensions, ranges, and resolutions; 2) fully leveraging existing multisource HSIs; 3) spontaneously capturing spectral features via self-supervised learning; and 4) pretrained on large-scale multisource HSI datasets and a considerable enhancement in classification accuracy. In three reconstruction test sets, the PSNRs are 27.07, 22.27, and 29.35 dB, and the SSIMs are 0.93, 0.82, and 0.88, respectively. Compared with the randomly initialized model, there are 8.28% and 1.4% improvements on the Indian Pines dataset and Pavia University dataset respectively.
Baisen Liu, Xiaojun Bi 0002, Yihan He
IEEE Geosci. Remote. Sens. Lett.4
2022 SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles
abstract
As shown by recent studies, machine intelligence-enabled systems are vulnerable to test cases resulting from either adversarial manipulation or natural distribution shifts. This has raised great concerns about deploying machine learning algorithms for real-world applications, especially in safety-critical domains such as autonomous driving (AD). On the other hand, traditional AD testing on naturalistic scenarios requires hundreds of millions of driving miles due to the high dimensionality and rareness of the safety-critical scenarios in the real world. As a result, several approaches for autonomous driving evaluation have been explored, which are usually, however, based on different simulation platforms, types of safety-critical scenarios, scenario generation algorithms, and driving route variations. Thus, despite a large amount of effort in autonomous driving testing, it is still challenging to compare and understand the effectiveness and efficiency of different testing scenario generation algorithms and testing mechanisms under similar conditions. In this paper, we aim to provide the first unified platform SafeBench to integrate different types of safety-critical testing scenarios, scenario generation algorithms, and other variations such as driving routes and environments. In particular, we consider 8 safety-critical testing scenarios following National Highway Traffic Safety Administration (NHTSA) and develop 4 scenario generation algorithms considering 10 variations for each scenario. Meanwhile, we implement 4 deep reinforcement learning-based AD algorithms with 4 types of input (e.g., bird’s-eye view, camera) to perform fair comparisons on SafeBench. We find our generated testing scenarios are indeed more challenging and observe the trade-off between the performance of AD agents under benign and safety-critical testing scenarios. We believe our unified platform SafeBench for large-scale and effective autonomous driving testing will motivate the development of new testing scenario generation and safe AD algorithms. SafeBench is available at https://safebench.github.io.
Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Yihan He, Hanjiang Hu, Ding Zhao, Bo Li 0026
NeurIPS6
2021 DPGraph: A Benchmark Platform for Differentially Private Graph Analysis
abstract
Differential privacy has become an appealing choice for analyzing sensitive data while offering strong privacy protection, even for complex data types like graphs. Despite a decade of academic efforts in designing differentially private algorithms for graph analysis, few works have been used in practice. This is due to their complexity in the choice of privacy guarantees and parameter/environmental configurations, or due to their scalability issues for large datasets.
Siyuan Xia, Beizhen Chang, Karl Knopf, Yihan He, Yuchao Tao, Xi He 0001
SIGMOD Conference4
2020 Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScript
abstract
Recent years have witnessed intensive research interests on training deep neural networks (DNNs) more efficiently by quantization-based compression methods, which facilitate DNNs training in two ways: (1) activations are quantized to shrink the memory consumption, and (2) gradients are quantized to decrease the communication cost. However, existing methods mostly use a uniform mechanism that quantizes the values evenly. Such a scheme may cause a large quantization variance and slow down the convergence in practice. In this work, we introduce TinyScript, which applies a non-uniform quantization algorithm to both activations and gradients. TinyScript models the original values by a family of Weibull distributions and searches for ”quantization knobs” that minimize quantization variance. We also discuss the convergence of the non-uniform quantization algorithm on DNNs with varying depths, shedding light on the number of bits required for convergence. Experiments show that TinyScript always obtains lower quantization variance, and achieves comparable model qualities against full precision training using 1-2 bits less than the uniform-based counterpart.
Fangcheng Fu, Yuzheng Hu, Yihan He, Jiawei Jiang 0001, Yingxia Shao, Ce Zhang 0001, Bin Cui 0001
ICML3