Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Andy Yang

dblp:92/3951 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Automata and formal languages · 62% Logic in computer science · 38%
Artificial intelligence
1 paper
Language models and text generation · 50% Deep learning architectures and training · 50%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › compositional generalization
length generalization
0.912025
A Formal Framework for Understanding Length Generalization in Transformers · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
A Formal Framework for Understanding Length Generalization in Transformers · ICLR 2025
Automata and formal languages › formal language classes
formal language expressiveness
0.812024
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages · NeurIPS 2024
Logic in computer science › temporal logic
linear temporal logic
0.812024
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages · NeurIPS 2024
Automata and formal languages › regular languages
star-free languages
0.212024
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

norm-based regularizer · 1.7causal transformers · 0.9causal transformer · 0.9absolute positional encodings · 0.9absolute positional encoding · 0.9hard attention · 0.8attention masking · 0.8Boolean RASP · 0.8
YearPublicationVenuePosition
2026 Simulating Hard Attention Using Soft Attention
abstract
Abstract We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores.
Andy Yang, Lena Strobl, David Chiang 0001, Dana Angluin
Trans. Assoc. Comput. Linguistics1
2025 A Formal Framework for Understanding Length Generalization in Transformers
abstract
A major challenge for transformers is generalizing to sequences longer than those observed during training. While previous works have empirically shown that transformers can either succeed or fail at length generalization depending on the task, theoretical understanding of this phenomenon remains limited. In this work, we introduce a rigorous theoretical framework to analyze length generalization in causal transformers with learnable absolute positional encodings. In particular, we characterize those functions that are identifiable in the limit from sufficiently long inputs with absolute positional encodings under an idealized inference scheme using a norm-based regularizer. This enables us to prove the possibility of length generalization for a rich family of problems. We experimentally validate the theory as a predictor of success and failure of length generalization across a range of algorithmic and formal language tasks. Our theory not only explains a broad set of empirical observations but also opens the way to provably predicting length generalization capabilities in transformers.
Xinting Huang, Andy Yang, Satwik Bhattamishra, Yash Raj Sarrof, Andreas Krebs, Hattie Zhou, Preetum Nakkiran, Michael Hahn 0001
ICLR2
2025 Knee-Deep in C-RASP: A Transformer Depth Hierarchy
abstract
It has been observed that transformers with greater depth (that is, more layers) have more capabilities, but can we establish formally which capabilities are gained? We answer this question with a theoretical proof followed by an empirical study. First, we consider transformers that round to fixed precision except inside attention. We show that this subclass of transformers is expressively equivalent to the programming language $\textsf{C}$-$\textsf{RASP}$ and this equivalence preserves depth. Second, we prove that deeper $\textsf{C}$-$\textsf{RASP}$ programs are more expressive than shallower $\textsf{C}$-$\textsf{RASP}$ programs, implying that deeper transformers are more expressive than shallower transformers (within the subclass mentioned above). The same is also proven for transformers with positional encodings (like RoPE and ALiBi). These results are established by studying a temporal logic with counting operators equivalent to $\textsf{C}$-$\textsf{RASP}$. Finally, we provide empirical evidence that our theory predicts the depth required for transformers without positional encodings to length-generalize on a family of sequential dependency tasks.
Andy Yang, Michaël Cadilhac, David Chiang 0001
NeurIPS1
2024 Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
abstract
The expressive power of transformers over inputs of unbounded size can be studied through their ability to recognize classes of formal languages. In this paper, we establish exact characterizations of transformers with hard attention (in which all attention is focused on exactly one position) and attention masking (in which each position only attends to positions on one side). With strict masking (each position cannot attend to itself) and without position embeddings, these transformers are expressively equivalent to linear temporal logic (LTL), which defines exactly the star-free languages. A key technique is the use of Boolean RASP as a convenient intermediate language between transformers and LTL. We then take numerous results known for LTL and apply them to transformers, showing how position embeddings, strict masking, and depth all increase expressive power.
Andy Yang, David Chiang 0001, Dana Angluin
NeurIPS1
2024 Variant graph craft (VGC): a comprehensive tool for analyzing genetic variation and identifying disease-causing variants
abstract
BACKGROUND: The variant call format (VCF) file is a structured and comprehensive text file crucial for researchers and clinicians in interpreting and understanding genomic variation data. It contains essential information about variant positions in the genome, along with alleles, genotype calls, and quality scores. Analyzing and visualizing these files, however, poses significant challenges due to the need for diverse resources and robust features for in-depth exploration. RESULTS: To address these challenges, we introduce variant graph craft (VGC), a VCF file visualization and analysis tool. VGC offers a wide range of features for exploring genetic variations, including extraction of variant data, intuitive visualization, and graphical representation of samples with genotype information. VGC is designed primarily for the analysis of patient cohorts, but it can also be adapted for use with individual probands or families. It integrates seamlessly with external resources, providing insights into gene function and variant frequencies in sample data. VGC includes gene function and pathway information from Molecular Signatures Database (MSigDB) for GO terms, KEGG, Biocarta, Pathway Interaction Database, and Reactome. Additionally, it dynamically links to gnomAD for variant information and incorporates ClinVar data for pathogenic variant information. VGC supports the Human Genome Assembly Hg37 and Hg38, ensuring compatibility with a wide range of data sets, and accommodates various approaches to exploring genetic variation data. It can be tailored to specific user needs with optional phenotype input data. CONCLUSIONS: In summary, VGC provides a comprehensive set of features tailored to researchers working with genomic variation data. Its intuitive interface, rapid filtering capabilities, and the flexibility to perform queries using custom groups make it an effective tool in identifying variants potentially associated with diseases. VGC operates locally, ensuring data security and privacy by eliminating the need for cloud-based VCF uploads, making it a secure and user-friendly tool. It is freely available at https://github.com/alperuzun/VGC .
Jennifer Li, Andy Yang, Benedito A. Carneiro, Ece D. Gamsiz Uzun, Lauren Massingham, Alper Uzun
BMC Bioinform.2
2021 e-Health for Older Adults: Navigating Misinformation
Amira Ghenai, Xueguang Ma, Robin Cohen, Karyn Moffatt, Andy Yang, Yipeng Ji
ICT4AWE5