Yi-Lun Liao

dblp:225/6644 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-5299-6749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Graph learning · 39% Efficient and distributed learning · 26% Speech recognition and synthesis · 19%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant graph neural network
1.422024
EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations · ICLR 2024
Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic Graphs · ICLR 2023
Machine learning › Graph learning
graph neural network
1.422024
EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations · ICLR 2024
Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic Graphs · ICLR 2023
Machine learning › Generative modeling
diffusion model
0.912025
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials · ICML 2025
Computational science and engineering
computational chemistry
0.912025
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials · ICML 2025
Machine learning › Efficient and distributed learning › automated machine learning
architecture optimization
0.512021
NetAdaptV2: Efficient Neural Architecture Search With Fast Super-Network Training and Architecture Optimization · CVPR 2021
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.512021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
low-resource speech recognition
0.512021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.512021
NetAdaptV2: Efficient Neural Architecture Search With Fast Super-Network Training and Architecture Optimization · CVPR 2021
Machine learning › Efficient and distributed learning › model compression
pruning
0.512021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021
Natural language and speech › Speech recognition and synthesis › speech representation learning
self-supervised speech representation
0.512021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular property prediction
0.212024
EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations · ICLR 2024
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.112021
PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

transformer · 2.4latent diffusion · 1.7autoencoder · 1.7separable layer normalization · 0.8separable activation · 0.8eSCN convolutions · 0.8attention re-normalization · 0.8equivariant attention · 0.7multi-layer coordinate descent · 0.5channel-level bypass connections · 0.5
YearPublicationVenuePosition
2025 All-atom Diffusion Transformers: Unified generative modelling of molecules and materials
abstract
Diffusion models are the standard toolkit for generative modelling of 3D atomic systems. However, for different types of atomic systems -- such as molecules and materials -- the generative processes are usually highly specific to the target system despite the underlying physics being the same. We introduce the All-atom Diffusion Transformer (ADiT), a unified latent diffusion framework for jointly generating both periodic materials and non-periodic molecular systems using the same model: (1) An autoencoder maps a unified, all-atom representations of molecules and materials to a shared latent embedding space; and (2) A diffusion model is trained to generate new latent embeddings that the autoencoder can decode to sample new molecules or materials. Experiments on MP20, QM9 and GEOM-DRUGS datasets demonstrate that jointly trained ADiT generates realistic and valid molecules as well as materials, obtaining state-of-the-art results on par with molecule and crystal-specific models. ADiT uses standard Transformers with minimal inductive biases for both the autoencoder and diffusion model, resulting in significant speedups during training and inference compared to equivariant diffusion models. Scaling ADiT up to half a billion parameters predictably improves performance, representing a step towards broadly generalizable foundation models for generative chemistry. Open source code: https://github.com/facebookresearch/all-atom-diffusion-transformer
Chaitanya K. Joshi, Xiang Fu 0005, Yi-Lun Liao, Vahe Gharakhanyan, Benjamin Kurt Miller, Anuroop Sriram, Zachary W. Ulissi
ICML3
2024 EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations
abstract
Equivariant Transformers such as Equiformer have demonstrated the efficacy of applying Transformers to the domain of 3D atomistic systems. However, they are limited to small degrees of equivariant representations due to their computational complexity. In this paper, we investigate whether these architectures can scale well to higher degrees. Starting from Equiformer, we first replace $SO(3)$ convolutions with eSCN convolutions to efficiently incorporate higher-degree tensors. Then, to better leverage the power of higher degrees, we propose three architectural improvements – attention re-normalization, separable $S^2$ activation and separable layer normalization. Putting these all together, we propose EquiformerV2, which outperforms previous state-of-the-art methods on large-scale OC20 dataset by up to 9% on forces, 4% on energies, offers better speed-accuracy trade-offs, and 2$\times$ reduction in DFT calculations needed for computing adsorption energies. Additionally, EquiformerV2 trained on only OC22 dataset outperforms GemNet-OC trained on both OC20 and OC22 datasets, achieving much better data efficiency. Finally, we compare EquiformerV2 with Equiformer on QM9 and OC20 S2EF-2M datasets to better understand the performance gain brought by higher degrees.
Yi-Lun Liao, Brandon M. Wood, Tess E. Smidt
ICLR1
2023 Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic Graphs
Yi-Lun Liao, Tess E. Smidt
ICLR1
2022 On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis
abstract
Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoffs between sparsity and its subsequent effects on synthetic speech. Additionally, we explore several aspects of TTS pruning: amount of finetuning data versus sparsity, TTS-Augmentation to utilize unspoken text, and combining knowledge distillation and pruning. Our findings suggest that not only are end-to-end TTS models highly prunable, but also, perhaps surprisingly, pruned TTS models can produce synthetic speech with equal or higher naturalness and intelligibility, with similar prosody. All of our experiments are conducted on publicly available models, and findings in this work are backed by large-scale subjective tests and objective measures. Code and 200 pruned models are made available to facilitate future research on efficiency in TTS1.
Cheng-I Lai, Erica Cooper, Yang Zhang 0001, Shiyu Chang, Kaizhi Qian, Yi-Lun Liao, Yung-Sung Chuang, Alexander H. Liu, Junichi Yamagishi, David D. Cox, James R. Glass
ICASSP6
2021 NetAdaptV2: Efficient Neural Architecture Search With Fast Super-Network Training and Architecture Optimization
abstract
Neural architecture search (NAS) typically consists of three main steps: training a super-network, training and evaluating sampled deep neural networks (DNNs), and training the discovered DNN. Most of the existing efforts speed up some steps at the cost of a significant slowdown of other steps or sacrificing the support of non-differentiable search metrics. The unbalanced reduction in the time spent per step limits the total search time reduction, and the inability to support non-differentiable search metrics limits the performance of discovered DNNs.In this paper, we present NetAdaptV2 with three innovations to better balance the time spent for each step while supporting non-differentiable search metrics. First, we propose channel-level bypass connections that merge network depth and layer width into a single search dimension to reduce the time for training and evaluating sampled DNNs. Second, ordered dropout is proposed to train multiple DNNs in a single forward-backward pass to decrease the time for training a super-network. Third, we propose the multi-layer coordinate descent optimizer that considers the interplay of multiple layers in each iteration of optimization to improve the performance of discovered DNNs while supporting non-differentiable search metrics. With these innovations, NetAdaptV2 reduces the total search time by up to 5.8× on ImageNet and 2.4× on NYU Depth V2, respectively, and discovers DNNs with better accuracy-latency/accuracy-MAC trade-offs than state-of-the-art NAS works. Moreover, the discovered DNN outperforms NAS-discovered MobileNetV3 by 1.8% higher top-1 accuracy with the same latency.1
Tien-Ju Yang, Yi-Lun Liao, Vivienne Sze
CVPR2
2021 PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
abstract
Self-supervised speech representation learning (speech SSL) has demonstrated the benefit of scale in learning rich representations for Automatic Speech Recognition (ASR) with limited paired data, such as wav2vec 2.0. We investigate the existence of sparse subnetworks in pre-trained speech SSL models that achieve even better low-resource ASR results. However, directly applying widely adopted pruning methods such as the Lottery Ticket Hypothesis (LTH) is suboptimal in the computational cost needed. Moreover, we show that the discovered subnetworks yield minimal performance gain compared to the original dense network.We present Prune-Adjust-Re-Prune (PARP), which discovers and finetunes subnetworks for much better performance, while only requiring a single downstream ASR finetuning run. PARP is inspired by our surprising observation that subnetworks pruned for pre-training tasks need merely a slight adjustment to achieve a sizeable performance boost in downstream ASR tasks. Extensive experiments on low-resource ASR verify (1) sparse subnetworks exist in mono-lingual/multi-lingual pre-trained speech SSL, and (2) the computational advantage and performance gain of PARP over baseline pruning methods.In particular, on the 10min Librispeech split without LM decoding, PARP discovers subnetworks from wav2vec 2.0 with an absolute 10.9%/12.6% WER decrease compared to the full model. We further demonstrate the effectiveness of PARP via: cross-lingual pruning without any phone recognition degradation, the discovery of a multi-lingual subnetwork for 10 spoken languages in 1 finetuning run, and its applicability to pre-trained BERT/XLNet for natural language tasks1.
Cheng-I Lai, Yang Zhang 0001, Alexander H. Liu, Shiyu Chang, Yi-Lun Liao, Yung-Sung Chuang, Kaizhi Qian, Sameer Khurana, David D. Cox, James R. Glass
NeurIPS5
2019 Learning Pose-aware 3D Reconstruction via 2D-3D Self-consistency
abstract
3D reconstruction, inferring 3D shape information from a single 2D image, has drawn attention from learning and vision communities. In this paper, we propose a framework for learning pose-aware 3D shape reconstruction. Our proposed model learns deep representation for recovering the 3D object, with the ability to extract camera pose information but without any direct supervision of ground truth camera pose. This is realized by exploitation of 2D-3D self-consistency between 2D masks and 3D voxels. Experiments qualitatively and quantitatively demonstrate the effectiveness and robustness of our model, which performs favorably against state-of-the-art methods.
Yi-Lun Liao, Yao-Cheng Yang, Yuan-Fang Lin, Pin-Jung Chen, Chia-Wen Kuo, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
ICASSP1
2018 Adaptively Banded Smith-Waterman Algorithm for Long Reads and Its Hardware Accelerator
abstract
In this paper, we propose hardware-compatible Adaptively Banded Smith-Waterman algorithm (ABSW) to align long genomic sequences. By utilizing banded Smith-Waterman algorithm to align subsequences of fixed lengths, ABSW finds alignment of a pair of arbitrarily long sequences with constant memory. In addition, a heuristic algorithm, dynamic overlapping, is proposed to make overlaps of bands of subsequences to improve accuracy. To enable hardware acceleration of ABSW, we further propose the hardware architecture of banded Smith-Waterman with traceback. Experiments show that ABSW produces near optimal alignment scores for sequences with up to 40% error rates. Our hardware implementation of ABSW demonstrates more than 200× speedup over software implementation.
Yi-Lun Liao, Nae-Chyun Chen, Yi-Chang Lu
ASAP1