Tien-Ju Yang

dblp:08/10809 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
6since 2021 · last 2024
0000-0003-4728-0321ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2024 Contemplative Mechanism for Speech Recognition: Speech Encoders can Think
Tien-Ju Yang, Andrew Rosenberg, Bhuvana Ramabhadran
INTERSPEECH1
2023 Online Model Compression for Federated Learning with Large Models
abstract
This paper addresses the challenges of training large neural networks under federated learning settings: high on-device memory usage and communication cost. The proposed Online Model Compression (OMC) provides a framework that stores model parameters in a compressed format and decompresses them only when needed. We use quantization as the compression method in this paper and propose three methods, (1) per-variable transformation, (2) weight-matrix-only quantization, and (3) partial variable quantization, to minimize its impact on model accuracy. Our experiments on two recent neural networks for speech recognition and two different datasets show that OMC can reduce memory usage and communication cost of model parameters by up to 59% while attaining comparable accuracy and training speed when compared with full-precision federated learning.
Tien-Ju Yang, Yonghui Xiao, Giovanni Motta, Françoise Beaufays, Rajiv Mathews, Mingqing Chen
ICASSP1
2022 Enabling On-Device Training of Speech Recognition Models With Federated Dropout
abstract
Federated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with the size of the model being trained, and are significant for state-of-the-art automatic speech recognition models.We propose using federated dropout to reduce the size of client models while training a full-size model server-side. We provide empirical evidence of the effectiveness of federated dropout, and propose a novel approach to vary the dropout rate applied at each layer. Furthermore, we find that federated dropout enables a set of smaller sub-models within the larger model to independently have low word error rates, making it easier to dynamically adjust the size of the model deployed for inference.
Dhruv Guliani, Lillian Zhou, Changwan Ryu, Tien-Ju Yang, Harry Zhang, Yonghui Xiao, Françoise Beaufays, Giovanni Motta
ICASSP4
2022 Partial Variable Training for Efficient on-Device Federated Learning
abstract
This paper aims to address the major challenges of Federated Learning (FL) on edge devices: limited memory and expensive communication. We propose a novel method, called Partial Variable Training (PVT), that only trains a small subset of variables on edge devices to reduce memory usage and communication cost. With PVT, we show that network accuracy can be maintained by utilizing more local training steps and devices, which is favorable for FL involving a large population of devices. According to our experiments on two state-of-the-art neural networks for speech recognition and two different datasets, PVT can reduce memory usage by up to 1.9× and communication cost by up to 593× while attaining comparable accuracy when compared with full network training.
Tien-Ju Yang, Dhruv Guliani, Françoise Beaufays, Giovanni Motta
ICASSP1
2022 Federated Pruning: Improving Neural Network Efficiency with Federated Learning
abstract
Automatic Speech Recognition models require large amount of speech data for training, and the collection of such data often leads to privacy concerns.Federated learning has been widely used and is considered to be an effective decentralized technique by collaboratively learning a shared prediction model while keeping the data local on different clients devices.However, the limited computation and communication resources on clients devices present practical difficulties for large models.To overcome such challenges, we propose Federated Pruning to train a reduced model under the federated setting, while maintaining similar performance compared to the full model.Moreover, the vast amount of clients data can also be leveraged to improve the pruning results compared to centralized training.We explore different pruning schemes and provide empirical evidence of the effectiveness of our methods.
Rongmei Lin, Yonghui Xiao, Tien-Ju Yang, Ding Zhao, Li Xiong 0001, Giovanni Motta, Françoise Beaufays
INTERSPEECH3
2021 NetAdaptV2: Efficient Neural Architecture Search With Fast Super-Network Training and Architecture Optimization
abstract
Neural architecture search (NAS) typically consists of three main steps: training a super-network, training and evaluating sampled deep neural networks (DNNs), and training the discovered DNN. Most of the existing efforts speed up some steps at the cost of a significant slowdown of other steps or sacrificing the support of non-differentiable search metrics. The unbalanced reduction in the time spent per step limits the total search time reduction, and the inability to support non-differentiable search metrics limits the performance of discovered DNNs.In this paper, we present NetAdaptV2 with three innovations to better balance the time spent for each step while supporting non-differentiable search metrics. First, we propose channel-level bypass connections that merge network depth and layer width into a single search dimension to reduce the time for training and evaluating sampled DNNs. Second, ordered dropout is proposed to train multiple DNNs in a single forward-backward pass to decrease the time for training a super-network. Third, we propose the multi-layer coordinate descent optimizer that considers the interplay of multiple layers in each iteration of optimization to improve the performance of discovered DNNs while supporting non-differentiable search metrics. With these innovations, NetAdaptV2 reduces the total search time by up to 5.8× on ImageNet and 2.4× on NYU Depth V2, respectively, and discovers DNNs with better accuracy-latency/accuracy-MAC trade-offs than state-of-the-art NAS works. Moreover, the discovered DNN outperforms NAS-discovered MobileNetV3 by 1.8% higher top-1 accuracy with the same latency.1
Tien-Ju Yang, Yi-Lun Liao, Vivienne Sze
CVPR1
2019 SegSort: Segmentation by Discriminative Sorting of Segments
abstract
Almost all existing deep learning approaches for semantic segmentation tackle this task as a pixel-wise classification problem. Yet humans understand a scene not in terms of pixels, but by decomposing it into perceptual groups and structures that are the basic building blocks of recognition. This motivates us to propose an end-to-end pixel-wise metric learning approach that mimics this process. In our approach, the optimal visual representation determines the right segmentation within individual images and associates segments with the same semantic classes across images. The core visual learning problem is therefore to maximize the similarity within segments and minimize the similarity between segments. Given a model trained this way, inference is performed consistently by extracting pixel-wise embeddings and clustering, with the semantic label determined by the majority vote of its nearest neighbors from an annotated set. As a result, we present the SegSort, as a first attempt using deep learning for unsupervised semantic segmentation, achieving 76% performance of its supervised counterpart. When supervision is available, SegSort shows consistent improvements over conventional approaches based on pixel-wise softmax training. Additionally, our approach produces more precise boundaries and consistent region predictions. The proposed SegSort further produces an interpretable result, as each choice of label can be easily understood from the retrieved nearest segments.
Jyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins, Tien-Ju Yang, Liang-Chieh Chen
ICCV5
2019 FastDepth: Fast Monocular Depth Estimation on Embedded Systems
abstract
Depth sensing is a critical function for robotic tasks such as localization, mapping and obstacle detection. There has been a significant and growing interest in depth estimation from a single RGB image, due to the relatively low cost and size of monocular cameras. However, state-of-the-art single-view depth estimation algorithms are based on fairly complex deep neural networks that are too slow for real-time inference on an embedded platform, for instance, mounted on a micro aerial vehicle. In this paper, we address the problem of fast depth estimation on embedded systems. We propose an efficient and lightweight encoder-decoder network architecture and apply network pruning to further reduce computational complexity and latency. In particular, we focus on the design of a low-latency decoder. Our methodology demonstrates that it is possible to achieve similar accuracy as prior work on depth estimation, but at inference speeds that are an order of magnitude faster. Our proposed network, FastDepth, runs at 178 fps on an NVIDIA Jetson TX2 GPU and at 27 fps when using only the TX2 CPU, with active power consumption under 10 W. FastDepth achieves close to state-of-the-art accuracy on the NYU Depth v2 dataset. To the best of the authors' knowledge, this paper demonstrates real-time monocular depth estimation using a deep neural network with the lowest latency and highest throughput on an embedded platform that can be carried by a micro aerial vehicle.
Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, Vivienne Sze
ICRA3
2018 MorphNet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks
abstract
We present MorphNet, an approach to automate the design of neural network structures. MorphNet iteratively shrinks and expands a network, shrinking via a resource-weighted sparsifying regularizer on activations and expanding via a uniform multiplicative factor on all layers. In contrast to previous approaches, our method is scalable to large networks, adaptable to specific resource constraints (e.g. the number of floating-point operations per inference), and capable of increasing the network's performance. When applied to standard network architectures on a wide variety of datasets, our approach discovers novel structures in each domain, obtaining higher performance while respecting the resource constraint.
Ariel Gordon, Elad Eban, Ofir Nachum, Bo Chen 0019, Tien-Ju Yang, Edward Choi 0003
CVPR6
2018 NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
Tien-Ju Yang, Andrew G. Howard, Bo Chen 0019, Alec Go, Mark Sandler 0002, Vivienne Sze, Hartwig Adam
ECCV (10)1
2017 Designing Energy-Efficient Convolutional Neural Networks Using Energy-Aware Pruning
abstract
Deep convolutional neural networks (CNNs) are indispensable to state-of-the-art computer vision algorithms. However, they are still rarely deployed on battery-powered mobile devices, such as smartphones and wearable gadgets, where vision algorithms can enable many revolutionary real-world applications. The key limiting factor is the high energy consumption of CNN processing due to its high computational complexity. While there are many previous efforts that try to reduce the CNN model size or the amount of computation, we find that they do not necessarily result in lower energy consumption. Therefore, these targets do not serve as a good metric for energy cost estimation. To close the gap between CNN design and energy consumption optimization, we propose an energy-aware pruning algorithm for CNNs that directly uses the energy consumption of a CNN to guide the pruning process. The energy estimation methodology uses parameters extrapolated from actual hardware measurements. The proposed layer-by-layer pruning algorithm also prunes more aggressively than previously proposed pruning methods by minimizing the error in the output feature maps instead of the filter weights. For each layer, the weights are first pruned and then locally fine-tuned with a closed-form least-square solution to quickly restore the accuracy. After all layers are pruned, the entire network is globally fine-tuned using back-propagation. With the proposed pruning method, the energy consumption of AlexNet and GoogLeNet is reduced by 3.7x and 1.6x, respectively, with less than 1% top-5 accuracy loss. We also show that reducing the number of target classes in AlexNet greatly decreases the number of weights, but has a limited impact on energy consumption.
Tien-Ju Yang, Vivienne Sze
CVPR1
2017 Efficient Processing of Deep Neural Networks: A Tutorial and Survey
abstract
Deep neural networks (DNNs) are currently widely used for many artificial intelligence (AI) applications including computer vision, speech recognition, and robotics. While DNNs deliver state-of-the-art accuracy on many AI tasks, it comes at the cost of high computational complexity. Accordingly, techniques that enable efficient processing of DNNs to improve energy efficiency and throughput without sacrificing application accuracy or increasing hardware cost are critical to the wide deployment of DNNs in AI systems. This article aims to provide a comprehensive tutorial and survey about the recent advances toward the goal of enabling efficient processing of DNNs. Specifically, it will provide an overview of DNNs, discuss various hardware platforms and architectures that support DNNs, and highlight key trends in reducing the computation cost of DNNs either solely via hardware design changes or via joint hardware design and DNN algorithm changes. It will also summarize various development resources that enable researchers and practitioners to quickly get started in this field, and highlight important benchmarking metrics and design considerations that should be used for evaluating the rapidly growing number of DNN hardware designs, optionally including algorithmic codesigns, being proposed in academia and industry. The reader will take away the following concepts from this article: understand the key design considerations for DNNs; be able to evaluate different DNN hardware implementations with benchmarks and comparison metrics; understand the tradeoffs between various hardware architectures and platforms; be able to evaluate the utility of various DNN design techniques for efficient processing; and understand recent implementation trends and opportunities.
Vivienne Sze, Tien-Ju Yang, Joel S. Emer
Proc. IEEE3
2012 A high speed feature matching architecture for real-time video stabilization
abstract
An efficient feature matching architecture targets at real-time video stabilization is revealed in this paper. For some applications, such as vehicular application, real-time video stabilization is needed to provide instant stable video input. However, feature matching is usually the bottleneck to achieve high performance. High speed feature matching architecture is proposed to accelerate the performance of video stabilization. Locality sensitive hashing (LSH) helps us realize the feature matching procedure in hardware implementation. By applying the proposed dynamic table allocation and on-chip cache mechanism, this work achieves 422K queries/s and real-time feature matching with 90% in memory reduction and more than 50% in relieving the feature bus burden of the system.
Keng-Yen Huang, Yi-Min Tsai, Tien-Ju Yang, Liang-Gee Chen
ISCAS3
2012 WarmL1: A warm-start homotopy-based reconstruction algorithm for sparse signals
abstract
A sparse signal can be reconstructed from a small amount of random and linear measurements by solving a system of underdetermined equations. In this paper, we study the reconstruction problem while the system undergoes dynamic modifications. Resolving this problem from scratch requires high computational efforts. Therefore, we propose an efficient homotopy-based reconstruction algorithm with warmstart, named WarmL1. WarmL1 quickly updates the previous solution to the desired one. Based on the concept of homotopy, WarmL1 breaks the reconstruction procedure into simple steps, and solves the problem iteratively. Four possible applications are presented and discussed to demonstrate the usage of WarmL1 for different warm-start situations. Experiments on these applications are performed. The results show that WarmL1 achieves 3.2× to 37.5× speeding up or up to 1/5100 l2-error at the same computational cost compared to related works.
Tien-Ju Yang, Yi-Min Tsai, Chung-Te Li, Liang-Gee Chen
ISIT1
2011 Smart display: A mobile self-adaptive projector-camera system
abstract
Owing to the diversity of projection surfaces, an effective mobile display system must be adaptive to the surface to avoid introducing a clipped scene. In this paper, we propose a smart mobile display system which automatically adapts to the location and motion of a surface. Firstly, an imperceptible structured light technique is adopted and continuous adaptation is accomplished. Secondly, a specifically designed code image with high distortion tolerance is proposed. Thirdly, we present a priority-based correction method to revise previous decoding results. Finally, a matching procedure resisting noise interruption is introduced. The system achieves 95% correct rate under the common indoor illuminance. In addition, the system performance is independent of the projected content and the surface shapes.
Tien-Ju Yang, Yi-Min Tsai, Liang-Gee Chen
ICME1