Edward C. Lin

dblp:51/3478 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 50% Reconfigurable computing and FPGAs · 50%
Artificial intelligence
2 papers
Speech recognition and synthesis · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
0.222009
A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizer · FPGA 2009
A 1000-word vocabulary, speaker-independent, continuous live-mode speech recognizer implemented in a single FPGA · FPGA 2007
Hardware accelerators and domain-specific architectures › signal processing accelerator
speech recognition accelerator
0.222009
A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizer · FPGA 2009
A 1000-word vocabulary, speaker-independent, continuous live-mode speech recognizer implemented in a single FPGA · FPGA 2007
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition
large vocabulary continuous speech recognition
0.012009
A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizer · FPGA 2009

Methods — techniques the papers use, named apart from their topics

viterbi search · 0.3beam search · 0.2acoustic modeling · 0.1
YearPublicationVenuePosition
2009 A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizer
abstract
Today's best quality speech recognition systems are implemented in software. These systems fully occupy the resources of a high-end server to deliver results at real-time speed: each hour of audio requires a significant fraction of an hour of computation for recognition. This is profoundly limiting for applications that require extreme recognition speed, for example, high-volume tasks such as video indexing (e.g., YouTube), or high-speed tasks such as triage of homeland security intelligence. We describe the architecture and implementation of one critical component -- the backend search stage -- of a high-speed, large-vocabulary recognizer. Implemented on a multi-FPGA Berkeley Emulation Engine 2 (BEE2) platform, we handle a standard 5000-word Wall Street Journal speech benchmark. Our backend search engine can decode on average 10 times faster than real-time running at 100 MHz, i.e, 10x faster than real-time, with negligible degradation in accuracy, running at a clock rate ~ 30x slower than a conventional server. To the best of our knowledge, this is both the most complex, and the fastest recognizer ever to be realized in a hardware form.
Edward C. Lin, Rob A. Rutenbar
FPGA1
2007 A 1000-word vocabulary, speaker-independent, continuous live-mode speech recognizer implemented in a single FPGA
abstract
The Carnegie Mellon In Silico Vox project seeks to move best-quality speech recognition technology from its current software-only form into a range of efficient all-hardware implementations. The central thesis is that, like graphics chips, the application is simply too performance hungry, and too power sensitive, to stay as a large software application. As a first step in this direction, we describe the design and implementation of a fully functional speech-to-text recognizer on a single Xilinx XUP platform. The design recognizes a 1000 word vocabulary, is speaker-independent, recognizes continuous (connected) speech, and is a "live mode" engine, wherein recognition can start as soon as speech input appears. To the best of our knowledge, this is the most complex recognizer architecture ever fully committed to a hardware-only form. The implementation is extraordinarily small, and achieves the same accuracy as state-of-the-art software recognizers, while running at a fraction of the clock speed.
Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen
FPGA1
2006 In silico vox: Towards speech recognition in silicon
Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen
Hot Chips Symposium1
2006 Moving speech recognition from software to silicon: the in silico vox project
abstract
To achieve much faster decoding, or much lower power consumption, we need to liberate speech recognition from the artificial constraints of its current software-only form, and move the essential computations directly into silicon. There are vast efficiencies waiting to be unlocked in this application – we need the proper architecture to do so. We report results from a firstgeneration hardware architecture simulated at bit-level, and a complete, working FPGA-based prototype. Simulation results show that rather modest hardware designs, running 10-20X slower than conventional processors, can already decode at 0.6 xRT, running the standard 5K Wall Street Journal benchmark.
Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen
INTERSPEECH1