Aaron Smith

dblp:19/3885 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-authorArtificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021Theory of computation · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 The Needle is a Thread: Finding Planted Paths in Noisy Process Trees
Maya Le, Pawel Pralat, Aaron Smith, François Théberge
WAW3
2026 Live and let die: Lysis time variability and resource limitation shape lytic bacteriophage fitness
abstract
Bacteriophages (phages) play a critical role in controlling bacterial populations, both in nature and as potential therapeutic agents. Their ability to replicate, compete against each other, and eradicate target cell populations is usually understood through a number of 'life history parameters', traditionally measured by population-level assays, which implicitly average the parameter's value across a large number of infection events. Recent experiments suggest that bacteriophage life history parameters are subject to considerable stochasticity, raising the question of whether experimental and modelling efforts that do not account for this variability may overlook important factors in phage's behaviour, competitive fitness or therapeutic viability. Here, using agent-based simulations, we investigate the importance of stochasticity in lysis time and burst size of lytic bacteriophages in two common laboratory competition experiments: serial passage of well-mixed populations and plaque expansion across a bacterial lawn. We find that a phage's analytic growth rate in isolation can be a poor predictor of its fitness advantage in simulated competition experiments. Specifically, when lysis times are tightly distributed, we identify a novel effect we name "population resonance", through which a bacteriophage can display a significant fitness advantage over a competitor with a much greater growth rate in isolation. Our simulations also show that both serial passage and plaque expansion reward variability in lysis time more than expected, by increasing the phage resilience when resources are scarce.
Aaron Smith, Michael Hunter, Somenath Bakshi, Diana Fusco
PLoS Comput. Biol.1
2025 Improving Community Detection via Community Association Strength Scores
Jordan Barrett, Ryan DeWolfe, Bogumil Kaminski, Pawel Pralat, Aaron Smith, François Théberge
WAW5
2025 The Artificial Benchmark for Community Detection with Outliers and Overlapping Communities ($\mathbf {ABCD{+}o}^2$)
Jordan Barrett, Ryan DeWolfe, Bogumil Kaminski, Pawel Pralat, Aaron Smith, François Théberge
WAW5
2024 On cyclical MCMC sampling
Aaron Smith, Aguemon Y. Atchadé
AISTATS3
2022 Rapid Convergence of Informed Importance Tempering
abstract
Informed Markov chain Monte Carlo (MCMC) methods have been proposed as scalable solutions to Bayesian posterior computation on high-dimensional discrete state spaces, but theoretical results about their convergence behavior in general settings are lacking. In this article, we propose a class of MCMC schemes called informed importance tempering (IIT), which combine importance sampling and informed local proposals, and derive generally applicable spectral gap bounds for IIT estimators. Our theory shows that IIT samplers have remarkable scalability when the target posterior distribution concentrates on a small set. Further, both our theory and numerical experiments demonstrate that the informed proposal should be chosen with caution: the performance may be very sensitive to the shape of the target distribution. We find that the “square-root proposal weighting” scheme tends to perform well in most settings.
Aaron Smith
AISTATS2
2021 Integrating a functional pattern-based IR into MLIR
abstract
The continued specialization in hardware and software due to the end of Moore's law forces us to question fundamental design choices in compilers, and in particular for domain specific languages. The days where a single universal compiler intermediate representation (IR) was sufficient to perform all important optimizations are over. We need novel IRs and ways for them to interact with one another while leveraging established compiler infrastructures.
Martin Paul Lücke, Michel Steuwer, Aaron Smith
CC3
2019 Mixing of Hamiltonian Monte Carlo on strongly log-concave distributions 2: Numerical integrators
abstract
We obtain quantitative bounds on the mixing properties of the Hamiltonian Monte Carlo (HMC) algorithm with target distribution in d-dimensional Euclidean space, showing that HMC mixes quickly whenever the target log-distribution is strongly concave and has Lipschitz gradients. We use a coupling argument to show that the popular leapfrog implementation of HMC can sample approximately from the target distribution in a number of gradient evaluations which grows like d^1/2 with the dimension and grows at most polynomially in the strong convexity and Lipschitz-gradient constants. Our results significantly extend and improve on the dimension dependence of previous quantitative bounds on the mixing of HMC and of the unadjusted Langevin algorithm in this setting.
Oren Mangoubi, Aaron Smith
AISTATS2
2019 Raising binaries to LLVM IR with MCTOLL (WIP paper)
abstract
The need to analyze and execute binaries from legacy ISAs on new or different ISAs has been addressed in a variety of ways over the past few decades. Solutions using complementary static and dynamic binary translation techniques have been deployed in most real-world situations. As new ISAs are designed and legacy ISAs re-examined, the need for binary translation infrastructure re-emerges, and needs to be re- engineered all over again.
S. Bharadwaj Yadavalli, Aaron Smith
LCTES2
2018 An Investigation of the Interactions Between Pre-Trained Word Embeddings, Character Models and POS Tags in Dependency Parsing
abstract
We provide a comprehensive analysis of the interactions between pre-trained word embeddings, character models and POS tags in a transition-based dependency parser.While previous studies have shown POS information to be less important in the presence of character models, we show that in fact there are complex interactions between all three techniques.In isolation each produces large improvements over a baseline system using randomly initialised word embeddings only, but combining them quickly leads to diminishing returns.We categorise words by frequency, POS tag and language in order to systematically investigate how each of the techniques affects parsing quality.For many word categories, applying any two of the three techniques is almost as good as the full combined system.Character models tend to be more important for low-frequency open-class words, especially in morphologically rich languages, while POS tags can help disambiguate highfrequency function words.We also show that large character embedding sizes help even for languages with small character sets, especially in morphologically rich languages.
Aaron Smith, Miryam de Lhoneux, Sara Stymne, Joakim Nivre
EMNLP1
2017 A Study of Dynamic Phase Adaptation Using a Dynamic Multicore Processor
abstract
Heterogeneous processors such as ARM’s big.LITTLE have become popular for embedded systems. They offer a choice between running workloads on a high performance core or a low-energy core leading to increased energy efficiency. However, the core configurations are fixed at design time which offers a limited amount of adaptation. Dynamic Multicore Processors (DMPs) bridge the gap between homogeneous and fully reconfigurable systems. Cores can fuse dynamically to adapt the computational resources to the needs of different workloads. There exists multiple examples of DMPs in the literature, yet the focus has mainly been on static partitioning. This paper conducts the first thorough study of the potential for dynamic reconfiguration of DMPs at runtime. We study how performance varies with static partitioning and what software optimizations are required to achieve high performance. We show that energy consumption is reduced considerably when adapting the number of cores to program phases, and introduce a simple online model which predicts the optimal number of cores to use to minimize energy consumption while maintaining high performance. Using the San Diego Vision Benchmark Suite as a use case, the dynamic scheme leads to ∼40% energy savings on average without decreasing performance.
Paul-Jules Micolet, Aaron Smith, Christophe Dubach
ACM Trans. Embed. Comput. Syst.2
2016 Parallel Markov Chain Monte Carlo via Spectral Clustering
abstract
As it has become common to use many computer cores in routine applications, finding good ways to parallelize popular algorithms has become increasingly important. In this paper, we present a parallelization scheme for Markov chain Monte Carlo (MCMC) methods based on spectral clustering of the underlying state space, generalizing earlier work on parallelization of MCMC methods by state space partitioning. We show empirically that this approach speeds up MCMC sampling for multimodal distributions and that it can be usefully applied in greater generality than several related algorithms. Our algorithm converges under reasonable conditions to an ‘optimal’ MCMC algorithm. We also show that our approach can be asymptotically far more efficient than naive parallelization, even in situations such as completely flat target distributions where no unique optimal algorithm exists. Finally, we combine theoretical and empirical bounds to provide practical guidance on the choice of tuning parameters.
Guillaume W. Basse, Aaron Smith, Natesh S. Pillai
AISTATS2
2016 Climbing Mont BLEU: The Strange World of Reachable High-BLEU Translations
Aaron Smith, Christian Hardmeier, Jörg Tiedemann
EAMT1
2016 A machine learning approach to mapping streaming workloads to dynamic multicore processors
abstract
Dataflow programming languages facilitate the design of data intensive programs such as streaming applications commonly found in embedded systems. They also expose parallelism that can be exploited using multicore processors which are now part of the mobile landscape. In recent years a shift has occurred towards heterogeneity ( ARM big.LITTLE) and reconfigurability. Dynamic Multicore Processors (DMPs) bridge the gap between fully reconfigurable processors and homogeneous multicore systems. They can re-allocate their resources at runtime to create larger more powerful logical processors fine-tuned to the workload. Unfortunately, there exists no accurate method to determine how to partition the cores in a DMP among application threads. Often programmers rely on analyzing the application manually and using a set of hand picked heuristics. This leads to sub-optimal performance, reducing the potential of DMPs. What is needed is a way to determine the optimal partitioning and grouping of resources to maximize performance. As a first step, this paper studies the effect of thread partitioning and hardware resource allocation on a set of StreamIt applications. We show that the resulting space is not trivial and exhibits a large performance variation depending on the combination of parameters. We introduce a machine-learning based methodology to tackle the space complexity. Our machine-learning model is able to directly predict the best combination of parameters using static code features. The predicted set of parameters leads to performance on-par with the best performance found in a space of more than 32,000 configurations per application.
Paul-Jules Micolet, Aaron Smith, Christophe Dubach
LCTES2
2016 Evaluating and optimizing OpenCL kernels for high performance computing with FPGAs
abstract
We evaluate the power and performance of the Rodinia benchmark suite using the Altera SDK for OpenCL targeting a Stratix V FPGA against a modern CPU and GPU. We study multiple OpenCL kernels per benchmark, ranging from direct ports of the original GPU implementations to loop-pipelined kernels specifically optimized for FPGAs. Based on our results, we find that even though OpenCL is functionally portable across devices, direct ports of GPU-optimized code do not perform well compared to kernels optimized with FPGA-specific techniques such as sliding windows. However, by exploiting FPGA-specific optimizations, it is possible to achieve up to 3.4x better power efficiency using an Altera Stratix V FPGA in comparison to an NVIDIA K20c GPU, and better run time and power efficiency in comparison to CPU. We also present preliminary results for Arria 10, which, due to hardened FPUs, exhibits noticeably better performance compared to Stratix V in floating-point-intensive benchmarks.
Hamid Reza Zohouri, Naoya Maruyama, Aaron Smith, Motohiko Matsuda, Satoshi Matsuoka
SC3
2014 EVX: Vector execution on low power EDGE cores
abstract
In this paper, we present a vector execution model that provides the advantages of vector processors on low power, general purpose cores, with limited additional hardware. While accelerating data-level parallel (DLP) workloads, the vector model increases the efficiency and hardware resources utilization. We use a modest dual issue core based on an Explicit Data Graph Execution (EDGE) architecture to implement our approach, called EVX. Unlike most DLP accelerators which utilize additional hardware and increase the complexity of low power processors, EVX leverages the available resources of EDGE cores, and with minimal costs allows for specialization of the resources. EVX adds a control logic that increases the core area by 2.1%. We show that EVX yields an average speedup of 3x compared to a scalar baseline and outperforms multimedia SIMD extensions.
Milovan Duric, Oscar Palomar, Aaron Smith, Osman S. Unsal, Adrián Cristal, Mateo Valero, Doug Burger
DATE3
2014 A reconfigurable fabric for accelerating large-scale datacenter services
abstract
Datacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution—or, while maintaining equivalent throughput, reduces the tail latency by 29%.
Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fowers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim 0001, Sitaram Lanka, James R. Larus, Eric Peterson, Simon Pope, Aaron Smith, Jason Thong, Phillip Yi Xiao, Doug Burger
ISCA20
2014 ParCor 1.0: A Parallel Pronoun-Coreference Corpus to Support Statistical MT
Liane Guillou, Christian Hardmeier, Aaron Smith, Jörg Tiedemann, Bonnie L. Webber
LREC3
2014 Scaling Power and Performance viaProcessor Composability
abstract
Power dissipation trends are leading high-performance processors to a regime in which all chip elements cannot be operated simultaneously at maximum frequency. Consequently, energy-efficiency will increase even more in importance, and performance must be achieved within strict power budgets. Current designs employ techniques such as dynamic voltage and frequency scaling (DVFS) to provide power-performance tradeoffs for both single and multi-threaded workloads. In power-dominated regimes, processors will be run at or near the minimum voltage. Frequency can be reduced to save power, but there is no scaling strategy for increasing performance with high energy-efficiency if the processor is operating at its maximum frequency (and minimum voltage). In this paper, we evaluate the energy-efficiency of processor composability—dynamically aggregating small energy-efficient physical cores into larger logical processors—as a method of scaling single-threaded performance up and down, comparing composability to the energy-efficiency of voltage and frequency scaling. We measure the power breakdowns of the baseline composable microarchitecture (the TFlex microarchitecture, based on an EDGE ISA) and compare the energy-efficiency and performance to one processor designed for power-efficiency (XScale) and another designed for high-performance (a variant of the Power-4) using normalized power models for as fair a comparison as possible. The study shows that composing multiple dual-issue cores (up to eight) provides performance scaling that is as energy-efficient as frequency scaling in a balanced microarchitecture, and is considerably more efficient than scaling the voltage to achieve additional performance once the maximum frequency at the minimum voltage is attained.
Madhu Saravana Sibi Govindan, Behnam Robatmili, Bertrand A. Maher, Aaron Smith, Stephen W. Keckler, Doug Burger
IEEE Trans. Computers5
2013 How to implement effective prediction and forwarding for fusable dynamic multicore architectures
abstract
Dynamic multicore architectures, that fuse and split cores at run time, potentially offer a level of performance/energy agility that static multicore designs cannot achieve. Conventional ISAs, however, have scalability limits to fusion. EDGE-based designs offer greater scalability but to date have been performance limited by significant microarchitectural bottlenecks. This paper addresses these issues and makes three major contributions. First, it proposes Iterative Path Prediction to address low next block prediction accuracy and low speculation rates. It achieves close to taken/not-taken prediction accuracy for multi-exit instruction blocks while also speculating the predicated execution path within the block. Second, the paper proposes Exposed Operand Broadcasts to address the overhead of operand delivery for high fanout instructions by exposing a small number of broadcast operands in the ISA. Third, we present a scalable composable architecture called T3 that uses these mechanisms and show it can operate across a wide range of power and performance spectrum by increasing energy efficiency and performance significantly. Compared to previous EDGE designs, T3 improves energy efficiency by about 2x and performance by up to 50%.
Behnam Robatmili, Hadi Esmaeilzadeh, Madhu Saravana Sibi Govindan, Aaron Smith, Andrew Putnam, Doug Burger, Stephen W. Keckler
HPCA5
2013 Provenance Representation for the National Climate Assessment in the Global Change Information System
abstract
The important topic of global climate change builds on a huge collection of scientific research. It is common for agencies releasing climate change information to be served with requests for all supporting materials resulting in a particular conclusion. Capturing and presenting global change provenance, linking to the research papers, data sets, models, analyses, observations, satellites, etc., that support the key research findings in this domain can increase understanding and aid in reproducibility of results and conclusions. The U.S. Global Change Research Program is now coordinating the production of a national climate assessment (NCA) that presents our best understanding of global change. We are now developing a global change information system that will present the content of that report and its provenance, including the scientific support for the findings of the assessment. We are using an approach that will present this information both through a human accessible Web site as well as a machine-readable interface for automated mining of the provenance graph. We plan to use the developing World Wide Web Consortium (W3C) PROV data model and ontology for this system. This paper will describe an overview of the process of developing the NCA and how the provenance trail of the report and each of the technical inputs can be captured and represented using the W3C PROV ontology. This will improve the visibility into the assessment process, increase understanding and possibility of reproducibility, and ultimately increase the credibility and trust of the resulting report.
Curt Tilmes, Peter Fox 0001, Xiaogang Ma 0001, Deborah L. McGuinness, Ana Pinheiro Privette, Aaron Smith, Anne Waple, Stephan Zednik, Jinguang Zheng
IEEE Trans. Geosci. Remote. Sens.6
2009 An evaluation of the TRIPS computer system
abstract
The TRIPS system employs a new instruction set architecture (ISA) called Explicit Data Graph Execution (EDGE) that renegotiates the boundary between hardware and software to expose and exploit concurrency. EDGE ISAs use a block-atomic execution model in which blocks are composed of dataflow instructions. The goal of the TRIPS design is to mine concurrency for high performance while tolerating emerging technology scaling challenges, such as increasing wire delays and power consumption. This paper evaluates how well TRIPS meets this goal through a detailed ISA and performance analysis. We compare performance, using cycles counts, to commercial processors. On SPEC CPU2000, the Intel Core 2 outperforms compiled TRIPS code in most cases, although TRIPS matches a Pentium 4. On simple benchmarks, compiled TRIPS code outperforms the Core 2 by 10% and hand-optimized TRIPS code outperforms it by factor of 3. Compared to conventional ISAs, the block-atomic model provides a larger instruction window, increases concurrency at a cost of more instructions executed, and replaces register and memory accesses with more efficient direct instruction-to-instruction communication. Our analysis suggests ISA, microarchitecture, and compiler enhancements for addressing weaknesses in TRIPS and indicates that EDGE architectures have the potential to exploit greater concurrency in future technologies.
Mark Gebhart, Bertrand A. Maher, Katherine E. Coons, Jeffrey R. Diamond, Paul Gratz, Mario Marino, Nitya Ranganathan, Behnam Robatmili, Aaron Smith, James H. Burrill, Stephen W. Keckler, Doug Burger, Kathryn S. McKinley
ASPLOS9
2006 Compiling for EDGE Architectures
abstract
Explicit data graph execution (EDGE) architectures offer the possibility of high instruction-level parallelism with energy efficiency. In EDGE architectures, the compiler breaks a program into a sequence of structured blocks that the hardware executes atomically. The instructions within each block communicate directly, instead of communicating through shared registers. The TRIPS EDGE architecture imposes restrictions on its blocks to simplify the microarchitecture: each TRIPS block has at most 128 instructions, issues at most 32 loads and/or stores, and executes at most 32 register bank reads and 32 writes. To detect block completion, each TRIPS block must produce a constant number of outputs (stores and register writes) and a branch decision. The goal of the TRIPS compiler is to produce TRIPS blocks full of useful instructions while enforcing these constraints. This paper describes a set of compiler algorithms that meet these sometimes conflicting goals, including an algorithm that assigns load and store identifiers to maximize the number of loads and stores within a block. We demonstrate the correctness of these algorithms in simulation on SPEC2000, EEMBC, and microbenchmarks extracted from SPEC2000 and others. We measure speedup in cycles over an Alpha 21264 on microbenchmarks.
Aaron Smith, Jon Gibson, Bertrand A. Maher, Nicholas Nethercote, Bill Yoder, Doug Burger, Kathryn S. McKinley, James H. Burrill
CGO1
2006 Merging Head and Tail Duplication for Convergent Hyperblock Formation
abstract
VLIW and EDGE (explicit data graph execution) architectures rely on compilers to form high-quality hyper-blocks for good performance. These compilers typically perform hyperblock formation, loop unrolling, and scalar optimizations in a fixed order. This approach limits the compiler's ability to exploit or correct interactions among these phases. EDGE architectures exacerbate this problem by imposing structural constraints on hyperblocks, such as instruction count and instruction composition. This paper presents convergent hyperblock formation, which iteratively applies if-conversion, peeling, unrolling, and scalar optimizations until converging on hyperblocks that are as close as possible to the structural constraints. To perform peeling and unrolling, convergent hyperblock formation generalizes tail duplication, which removes side entrances to acyclic traces, to remove back edges into cyclic traces using head duplication. Simulation results for an EDGE architecture show that convergent hyperblock formation improves code quality over discrete-phase approaches with heuristics for VLIW and EDGE. This algorithm offers a solution to hyperblock phase ordering problems and can be configured to implement a wide range of policies
Bertrand A. Maher, Aaron Smith, Doug Burger, Kathryn S. McKinley
MICRO2
2006 Dataflow Predication
abstract
Predication facilitates high-bandwidth fetch and large static scheduling regions, but has typically been too complex to implement comprehensively in out-of-order micro architectures. This paper describes dataflow predication, which provides per-instruction predication in a dataflow ISA, low predication computation overheads similar to VLIW ISAs, and low complexity out-of-order issue. A two-bitfield in each instruction specifies whether an instruction is predicated, in which case, an arriving predicate token determines whether an instruction should execute. Dataflow predication incorporates three features that reduce predication overheads. First, dataflow predicate computation permits computation of compound predicates with virtually no overhead instructions. Second, early mispredication termination squashes in-flight instructions with false predicates at any time, eliminating the overhead of falsely predicated paths. Finally, implicit predication mitigates the fanout overhead of dataflow predicates by reducing the number of explicitly predicated instructions, by predicating only the heads of dependence chains. Dataflow predication also exposes new compiler optimizations - such as disjoint instruction merging and path-sensitive predicate removal - for increased performance of predicated code in an out-of-order design
Aaron Smith, Ramadass Nagarajan, Karthikeyan Sankaralingam, Robert G. McDonald, Doug Burger, Stephen W. Keckler, Kathryn S. McKinley
MICRO1
1997 Speech recognition HMM training on reconfigurable parallel processor
abstract
Armstrong III is a 20 node multi-computer that is currently operational. In addition to a RISC processor, each node contains reconfigurable resources implemented with FPGAs. The in-circuit reprogramability of static RAM based FPGAs allows the computational capabilities of a node to be dynamically matched to the computational requirements of an application. Most reconfigurable computers in existence today rely solely on a large number of FPGAs to perform computations. In contrast, the paper demonstrates the utility of a small number of FPGAs coupled to a RISC processor with a simple interconnect. The article describes a substantive example application that performs HMM training for speech recognition with the reconfigurable platform.
Hyun-Kyu Yun, Aaron Smith, Harvey F. Silverman
FCCM2
1995 Implementing a genetic algorithm on a parallel custom computing machine
abstract
Genetic algorithms (GAs) are a currently popular method for nonlinear optimization that can be used to provide a solution for the chip partitioning problem. Unfortunately, GAs usually require prohibitively large computation times on current workstations. This paper demonstrates the utility of the Armstrong III architecture by addressing the computational problems associated with partitioning large designs using GAs. An example GA is presented for chip partitioning that runs on Armstrong III. GA computation bottlenecks are identified and hardware implementation strategies are discussed. Results are presented that show the Armstrong III architecture can be adapted to execute a GA in significantly less time than current workstations.
Nathan Sitkoff, Michael E. Wazlowski, Aaron Smith, Harvey F. Silverman
FCCM3
1995 Performing Log-Scale Addition on a Distributed Memory MIMD Multicomputer with Reconfigurable Computing Capabilities
Michael E. Wazlowski, Aaron Smith, Ricardo Citro, Harvey F. Silverman
ICPP (3)2
1992 Distributed hidden Markov model training on loosely-coupled multiprocessor networks
abstract
An explicit-duration hidden Markov model (HMM) algorithm for speech recognition has been proposed that potentially provides a more precise and versatile duration model than the implicit models ordinarily used, but at the cost of increased computation. The authors address the computational issues involved in conventional and explicit-duration HMM training by providing an analysis of the algorithm, suggesting serial enhancements and two efficient parallel implementations, and presenting experimental results on both common network workstations as well as a parallel system.>
Jonathan Foote, Mike Hochberg, Peter M. Athanas, Aaron Smith, Michael E. Wazlowski, Harvey F. Silverman
ICASSP4