Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Michael A. Kozuch

dblp:59/5000 · also Michael Kozuch · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
3since 2021 · last 2024
0009-0009-0939-3297ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 45 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 1 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
26 papers
Memory systems · 39% Cloud and datacenter computing · 30% Hardware accelerators and domain-specific architectures · 10%
Software engineering, system software, and programming languages
6 papers
Operating systems · 51% Program analysis · 32% Concurrent programming · 17%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 30 heaviest of 62, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Textual-Visual Logic Challenge: Understanding and Reasoning in Text-to-Image Generation · ECCV (5) 2024
Cloud and datacenter computing
serverless computing
0.712023
Memento: Architectural Support for Ephemeral Memory Management in Serverless Environments · MICRO 2023
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.622018
3Sigma: distribution-based cluster scheduling for runtime uncertainty · EuroSys 2018
TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters · EuroSys 2016
Memory systems
non-volatile memory
0.512021
NVOverlay: Enabling Efficient and Scalable High-Frequency Snapshotting to NVM · ISCA 2021
Memory systems
processing-in-memory
0.522017
Ambit: in-memory accelerator for bulk bitwise operations using commodity DRAM technology · MICRO 2017
RowClone: fast and energy-efficient in-DRAM bulk data copy and initialization · MICRO 2013
Storage systems
distributed storage
0.432014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Integrating Portable and Distributed Storage · FAST 2004
Program analysis
dynamic analysis
0.432014
Guardrail: a high fidelity approach to protecting hardware devices from buggy drivers · ASPLOS 2014
ParaLog: enabling and accelerating online parallel monitoring of multithreaded applications · ASPLOS 2010
Butterfly analysis: adapting dataflow analysis to dynamic parallel monitoring · ASPLOS 2010
Cloud and datacenter computing
cluster resource management and scheduling
0.422016
TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters · EuroSys 2016
AutoScale: Dynamic, Robust Capacity Management for Multi-Tier Data Centers · ACM Trans. Comput. Syst. 2012
Memory systems
DRAM
0.422015
Gather-scatter DRAM: in-DRAM address translation to improve the spatial locality of non-unit strided accesses · MICRO 2015
RowClone: fast and energy-efficient in-DRAM bulk data copy and initialization · MICRO 2013
Memory systems
cache
0.422014
Mitigating Prefetcher-Caused Pollution Using Informed Caching Policies for Prefetched Blocks · ACM Trans. Archit. Code Optim. 2014
The Dirty-Block Index · ISCA 2014
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312018
3Sigma: distribution-based cluster scheduling for runtime uncertainty · EuroSys 2018
Hardware accelerators and domain-specific architectures
video processing accelerator
0.312018
Mainstream: Dynamic Stem-Sharing for Multi-Tenant Video Processing · USENIX ATC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
in-memory computing accelerator
0.312017
Ambit: in-memory accelerator for bulk bitwise operations using commodity DRAM technology · MICRO 2017
Operating systems › resource management
memory management
0.322023
Memento: Architectural Support for Ephemeral Memory Management in Serverless Environments · MICRO 2023
Page overlays: an enhanced virtual memory framework to enable fine-grained memory management · ISCA 2015
Memory systems › cache management
cache replacement
0.322015
Mitigating Prefetcher-Caused Pollution Using Informed Caching Policies for Prefetched Blocks · ACM Trans. Archit. Code Optim. 2014
Exploiting compressed block size as an indicator of future reuse · HPCA 2015
Embedded and real-time systems › real-time scheduling
reservation-based scheduling
0.212016
TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters · EuroSys 2016
Memory systems
cache management
0.212015
Exploiting compressed block size as an indicator of future reuse · HPCA 2015
Memory systems
memory access patterns
0.212015
Gather-scatter DRAM: in-DRAM address translation to improve the spatial locality of non-unit strided accesses · MICRO 2015
Memory systems › memory access patterns
strided access
0.212015
Gather-scatter DRAM: in-DRAM address translation to improve the spatial locality of non-unit strided accesses · MICRO 2015
Memory systems › memory management
virtual memory
0.212015
Page overlays: an enhanced virtual memory framework to enable fine-grained memory management · ISCA 2015
Operating systems › i/o › i/o subsystem › device drivers
device driver reliability
0.212014
Guardrail: a high fidelity approach to protecting hardware devices from buggy drivers · ASPLOS 2014
Operating systems
virtualization
0.212014
Gleaner: Mitigating the Blocked-Waiter Wakeup Problem for Virtualized Multicore Applications · USENIX ATC 2014
Memory systems › cache management › cache interference
cache pollution
0.212014
Mitigating Prefetcher-Caused Pollution Using Informed Caching Policies for Prefetched Blocks · ACM Trans. Archit. Code Optim. 2014
Cloud and datacenter computing
cloud storage
0.212014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Storage systems
data migration
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Cloud and datacenter computing › datacenter storage
elastic storage
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Memory systems › memory compression
cache compression
0.212013
Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013
Memory systems › memory compression
main memory compression
0.212013
Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013
Memory systems
memory controller
0.212013
Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013
Processor architecture and microarchitecture › multiprocessor architecture
multi-socket systems
0.112021
NVOverlay: Enabling Efficient and Scalable High-Frequency Snapshotting to NVM · ISCA 2021

Methods — techniques the papers use, named apart from their topics

architectural support · 1.3multimodal reasoning · 0.8invariant checking · 0.4hardware logging · 0.4dynamic stem-sharing · 0.3distribution-based runtime estimation · 0.3bulk-bitwise operation · 0.3DRAM technology · 0.3plan-ahead scheduling · 0.2calendaring · 0.2reaching expressions · 0.2reaching definitions · 0.2hardware acceleration · 0.2data flow analysis · 0.2address translation · 0.2prototype implementation · 0.1emulation · 0.1metadata-TLBs · 0.1
YearPublicationVenuePosition
2024 Textual-Visual Logic Challenge: Understanding and Reasoning in Text-to-Image Generation
Peixi Xiong, Michael A. Kozuch, Nilesh Jain
ECCV (5)2
2023 Memento: Architectural Support for Ephemeral Memory Management in Serverless Environments
abstract
Serverless computing is an increasingly attractive paradigm in the cloud due to its ease of use and fine-grained pay-for-what-you-use billing. However, serverless computing poses new challenges to system design due to its short-lived function execution model. Our detailed analysis reveals that memory management is responsible for a major amount of function execution cycles. This is because functions pay the full critical-path costs of memory management in both userspace and the operating system without the opportunity to amortize these costs over their short lifetimes.
Ziqi Wang 0007, Kaiyang Zhao 0002, Andrew Jacob, Michael A. Kozuch, Todd C. Mowry, Dimitrios Skarlatos 0002
MICRO5
2021 NVOverlay: Enabling Efficient and Scalable High-Frequency Snapshotting to NVM
abstract
The ability to capture frequent (per millisecond) persistent snapshots to NVM would enable a number of compelling use cases. Unfortunately, existing NVM snapshotting techniques suffer from a combination of persistence barrier stalls, write amplification to NVM, and/or lack of scalability beyond a single socket. In this paper, we present NVOverlay, which is a scalable and efficient technique for capturing frequent persistent snapshots to NVM such that they can be randomly accessed later. NVOverlay uses Coherent Snapshot Tracking to efficiently track changes to memory (since the previous snapshot) across multi-socket parallel systems, and it uses Multi-snapshot NVM Mapping to store these snapshots to NVM while avoiding excessive write amplification. Our experiments demonstrate that NVOverlay successfully hides the overhead of capturing these snapshots while reducing write amplification by 29%–47% compared with state-of-the-art logging-based snapshotting techniques.
Ziqi Wang 0007, Chul-Hwan Choo, Michael A. Kozuch, Todd C. Mowry, Gennady Pekhimenko, Vivek Seshadri, Dimitrios Skarlatos 0002
ISCA3
2019 Multiversioned Page Overlays: Enabling Faster Serializable Hardware Transactional Memory
abstract
Practical and efficient support for multiversioning memory systems would offer a number of potential advantages, including improving the performance and functionality of hardware transactional memory (HTM). This paper presents a new approach to multiversioning support (Multiversioned Page Overlays) along with a new HTM design that it enables: OverlayTM. Compared with existing HTM designs, OverlayTM takes advantage of multiversioning to reduce unnecessary transaction aborts while providing full serializable semantics (in contrast with multiversioning HTMs that improve performance at the expense of being vulnerable to write skew anomalies). Our performance results demonstrate that OverlayTM is especially advantageous in read-heavy workloads.
Ziqi Wang 0007, Michael A. Kozuch, Todd C. Mowry, Vivek Seshadri
PACT2
2018 3Sigma: distribution-based cluster scheduling for runtime uncertainty
abstract
The 3Sigma cluster scheduling system uses job runtime histories in a new way. Knowing how long each job will execute enables a scheduler to more effectively pack jobs with diverse time concerns (e.g., deadline vs. the-sooner-the-better) and placement preferences on heterogeneous cluster resources. But, existing schedulers use single-point estimates (e.g., mean or median of a relevant subset of historical runtimes), and we show that they are fragile in the face of real-world estimate error profiles. In particular, analysis of job traces from three different large-scale cluster environments shows that, while the runtimes of many jobs can be predicted well, even state-of-the-art predictors have wide error profiles with 8--23% of predictions off by a factor of two or more. Instead of reducing relevant history to a single point, 3Sigma schedules jobs based on full distributions of relevant runtime histories and explicitly creates plans that mitigate the effects of anticipated runtime uncertainty. Experiments with workloads derived from the same traces show that 3Sigma greatly outperforms a state-of-the-art scheduler that uses point estimates from a state-of-the-art predictor; in fact, the performance of 3Sigma approaches the end-to-end performance of a scheduler based on a hypothetical, perfect runtime predictor. 3Sigma reduces SLO miss rate, increases cluster goodput, and improves or matches latency for best effort jobs.
Jun Woo Park, Alexey Tumanov, Angela H. Jiang, Michael A. Kozuch, Gregory R. Ganger
EuroSys4
2018 Mainstream: Dynamic Stem-Sharing for Multi-Tenant Video Processing
Angela H. Jiang, Daniel Lin-Kit Wong, Christopher Canel, Lilia Tang, Ishan Misra, Michael Kaminsky, Michael A. Kozuch, Padmanabhan Pillai, David G. Andersen, Gregory R. Ganger
USENIX ATC7
2017 WorkloadCompactor: reducing datacenter cost while providing tail latency SLO guarantees
abstract
Service providers want to reduce datacenter costs by consolidating workloads onto fewer servers. At the same time, customers have performance goals, such as meeting tail latency Service Level Objectives (SLOs). Consolidating workloads while meeting tail latency goals is challenging, especially since workloads in production environments are often bursty. To limit the congestion when consolidating workloads, customers and service providers often agree upon rate limits. Ideally, rate limits are chosen to maximize the number of workloads that can be co-located while meeting each workload's SLO. In reality, neither the service provider nor customer knows how to choose rate limits. Customers end up selecting rate limits on their own in some ad hoc fashion, and service providers are left to optimize given the chosen rate limits.
Timothy Zhu, Michael A. Kozuch, Mor Harchol-Balter
SoCC2
2017 RBAY: A Scalable and Extensible Information Plane for Federating Distributed Datacenter Resources
abstract
While many institutions, whether industrial, academic, or governmental, satisfy their computing needs through public cloud providers, many others still manage their own resources, often as geographically distributed datacenters. Spare capacity from these geographically distributed datacenters could be offered to others, provided there were a mechanism to discover, and then request these resources. Unfortunately, single datacenter administrators tend not to cooperate due to issues of scalability, diverse administrative policies, and site-specific monitoring infrastructure. This paper describes RBAY, an integrated information plane that enables secure and scalable sharing between geographically distributed datacenters. RBAY's key design features are twofold. First, RBAY employs a decentralized `hierarchical aggregation tree' structure to seamlessly aggregate spare resources from geographically distributed datacenters to a global information plane. Second, RBAY attaches to each participating server a `admin-customized' handler, which follows site-specific policy to expose, hide, add, remove resources to RBAY, and thus fulfill the task of `which resource to expose to whom, when, and how'. An experimental evaluation on eight real-world geo-distributed sites demonstrates RBAY's rapid response to composite queries, as well as its extensible, scalable, and lightweight nature.
Liting Hu, Douglas M. Blough, Michael A. Kozuch, Matthew Wolf
ICDCS4
2017 Ambit: in-memory accelerator for bulk bitwise operations using commodity DRAM technology
abstract
Many important applications trigger bulk bitwise operations, i.e., bitwise operations on large bit vectors. In fact, recent works design techniques that exploit fast bulk bitwise operations to accelerate databases (bitmap indices, BitWeaving) and web search (BitFunnel). Unfortunately, in existing architectures, the throughput of bulk bitwise operations is limited by the memory bandwidth available to the processing unit (e.g., CPU, GPU, FPGA, processing-in-memory).
Vivek Seshadri, Donghyuk Lee, Thomas Mullins, Hasan Hassan, Amirali Boroumand, Jeremie S. Kim, Michael A. Kozuch, Onur Mutlu, Phillip B. Gibbons, Todd C. Mowry
MICRO7
2016 TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters
abstract
TetriSched is a scheduler that works in tandem with a calendaring reservation system to continuously re-evaluate the immediate-term scheduling plan for all pending jobs (including those with reservations and best-effort jobs) on each scheduling cycle. TetriSched leverages information supplied by the reservation system about jobs' deadlines and estimated runtimes to plan ahead in deciding whether to wait for a busy preferred resource type (e.g., machine with a GPU) or fall back to less preferred placement options. Plan-ahead affords significant flexibility in handling mis-estimates in job runtimes specified at reservation time. Integrated with the main reservation system in Hadoop YARN, TetriSched is experimentally shown to achieve significantly higher SLO attainment and cluster utilization than the best-configured YARN reservation and CapacityScheduler stack deployed on a real 256 node cluster.
Alexey Tumanov, Timothy Zhu, Jun Woo Park, Michael A. Kozuch, Mor Harchol-Balter, Gregory R. Ganger
EuroSys4
2015 Tracking and Reducing Uncertainty in Dataflow Analysis-Based Dynamic Parallel Monitoring
abstract
Dataflow analysis-based dynamic parallel monitoring (DADPM) is a recent approach for identifying bugs in parallel software as it executes, based on the key insight of explicitly modeling a sliding window of uncertainty across parallel threads. While this makes the approach practical and scalable, it also introduces the possibility of false positives in the analysis. In this paper, we improve upon the DADPM framework through two observations. First, by explicitly tracking new “uncertain” states in the metadata lattice, we can distinguish potential false positives from true positives. Second, as the analysis tool runs dynamically, it can use the existence (or absence) of observed uncertain states to adjust the tradeoff between precision and performance on-the-fly. For example, we demonstrate how the epoch size parameter can be adjusted dynamically in response to uncertainty in order to achieve better performance and precision than when the tool is statically configured. This paper shows how to adapt a canonical dataflow analysis problem (reaching definitions) and a popular security monitoring tool (TAINTCHECK) to our new uncertainty-tracking framework, and provides new provable guarantees that reported true errors are now precise.
Michelle L. Goodstein, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
PACT3
2015 Exploiting compressed block size as an indicator of future reuse
abstract
We introduce a set of new Compression-Aware Management Policies (CAMP) for on-chip caches that employ data compression. Our management policies are based on two key ideas. First, we show that it is possible to build a more efficient management policy for compressed caches if the compressed block size is directly used in calculating the value (importance) of a block to the cache. This leads to Minimal-Value Eviction (MVE), a policy that evicts the cache blocks with the least value, based on both the size and the expected future reuse. Second, we show that, in some cases, compressed block size can be used as an efficient indicator of the future reuse of a cache block. We use this idea to build a new insertion policy called Size-based Insertion Policy (SIP) that dynamically prioritizes cache blocks using their compressed size as an indicator. We compare CAMP (and its global variant G-CAMP) to prior on-chip cache management policies (both size-oblivious and size-aware) and find that our mechanisms are more effective in using compressed block size as an extra dimension in cache management decisions. Our results show that the proposed management policies (i) decrease off-chip bandwidth consumption (by 8.7% in single-core), (ii) decrease memory subsystem energy consumption (by 7.2% in single-core) for memory intensive workloads compared to the best prior mechanism, and (iii) improve performance (by 4.9%/9.0%/10.2% on average in single-/two-/four-cor e workload evaluations and up to 20.1%) CAMP is effective for a variety of compression algorithms and different cache designs with local and global replacement strategies.
Gennady Pekhimenko, Tyler Huberty, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
HPCA6
2015 Page overlays: an enhanced virtual memory framework to enable fine-grained memory management
abstract
Many recent works propose mechanisms demonstrating the potential advantages of managing memory at a fine (e.g., cache line) granularity---e.g., fine-grained deduplication and fine-grained memory protection. Unfortunately, existing virtual memory systems track memory at a larger granularity (e.g., 4 KB pages), inhibiting efficient implementation of such techniques. Simply reducing the page size results in an unacceptable increase in page table overhead and TLB pressure.
Vivek Seshadri, Gennady Pekhimenko, Olatunji Ruwase, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry, Trishul M. Chilimbi
ISCA6
2015 Gather-scatter DRAM: in-DRAM address translation to improve the spatial locality of non-unit strided accesses
abstract
Many data structures (e.g., matrices) are typically accessed with multiple access patterns. Depending on the layout of the data structure in physical address space, some access patterns result in non-unit strides. In existing systems, which are optimized to store and access cache lines, non-unit strided accesses exhibit low spatial locality. Therefore, they incur high latency, and waste memory bandwidth and cache space.
Vivek Seshadri, Thomas Mullins, Amirali Boroumand, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
MICRO6
2014 Guardrail: a high fidelity approach to protecting hardware devices from buggy drivers
abstract
Device drivers are an Achilles' heel of modern commodity operating systems, accounting for far too many system failures. Previous work on driver reliability has focused on protecting the kernel from unsafe driver side-effects by interposing an invariant-checking layer at the driver interface, but otherwise treating the driver as a black box. In this paper, we propose and evaluate Guardrail, which is a more powerful framework for run-time driver analysis that performs decoupled instruction-grain dynamic correctness checking on arbitrary kernel-mode drivers as they execute, thereby enabling the system to detect and mitigate more challenging correctness bugs (e.g., data races, uninitialized memory accesses) that cannot be detected by today's fault isolation techniques. Our evaluation of Guardrail shows that it can find serious data races, memory faults, and DMA faults in native Linux drivers that required fixes, including previously unknown bugs. Also, with hardware logging support, Guardrail can be used for online protection of persistent device state from driver bugs with at most 10% overhead on the end-to-end performance of most standard I/O workloads.
Olatunji Ruwase, Michael A. Kozuch, Phillip B. Gibbons, Todd C. Mowry
ASPLOS2
2014 PriorityMeister: Tail Latency QoS for Shared Networked Storage
abstract
Meeting service level objectives (SLOs) for tail latency is an important and challenging open problem in cloud computing infrastructures. The challenges are exacerbated by burstiness in the workloads. This paper describes PriorityMeister -- a system that employs a combination of per-workload priorities and rate limits to provide tail latency QoS for shared networked storage, even with bursty workloads. PriorityMeister automatically and proactively configures workload priorities and rate limits across multiple stages (e.g., a shared storage stage followed by a shared network stage) to meet end-to-end tail latency SLOs. In real system experiments and under production trace workloads, PriorityMeister outperforms most recent reactive request scheduling approaches, with more workloads satisfying latency SLOs at higher latency percentiles. PriorityMeister is also robust to mis-estimation of underlying storage device performance and contains the effect of misbehaving workloads.
Timothy Zhu, Alexey Tumanov, Michael A. Kozuch, Mor Harchol-Balter, Gregory R. Ganger
SoCC3
2014 SpringFS: bridging agility and performance in elastic distributed storage
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
FAST6
2014 The Dirty-Block Index
abstract
On-chip caches maintain multiple pieces of metadata about each cached block—e.g., dirty bit, coherence information, ECC. Traditionally, such metadata for each block is stored in the corresponding tag entry in the tag store. While this approach is simple to implement and scalable, it necessitates a full tag store lookup for any metadata query—resulting in high latency and energy consumption. We Vnd that this approach is inefficient and inhibits several cache optimizations.
Vivek Seshadri, Abhishek Bhowmick 0002, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
ISCA5
2014 Gleaner: Mitigating the Blocked-Waiter Wakeup Problem for Virtualized Multicore Applications
Xiaoning Ding, Phillip B. Gibbons, Michael A. Kozuch, Jianchen Shan
USENIX ATC3
2014 Mitigating Prefetcher-Caused Pollution Using Informed Caching Policies for Prefetched Blocks
abstract
Many modern high-performance processors prefetch blocks into the on-chip cache. Prefetched blocks can potentially pollute the cache by evicting more useful blocks. In this work, we observe that both accurate and inaccurate prefetches lead to cache pollution, and propose a comprehensive mechanism to mitigate prefetcher-caused cache pollution. First, we observe that over 95% of useful prefetches in a wide variety of applications are not reused after the first demand hit (in secondary caches). Based on this observation, our first mechanism simply demotes a prefetched block to the lowest priority on a demand hit. Second, to address pollution caused by inaccurate prefetches, we propose a self-tuning prefetch accuracy predictor to predict if a prefetch is accurate or inaccurate. Only predicted-accurate prefetches are inserted into the cache with a high priority. Evaluations show that our final mechanism, which combines these two ideas, significantly improves performance compared to both the baseline LRU policy and two state-of-the-art approaches to mitigating prefetcher-caused cache pollution (up to 49%, and 6% on average for 157 two-core multiprogrammed workloads). The performance improvement is consistent across a wide variety of system configurations.
Vivek Seshadri, Samihan Yedkar, Hongyi Xin, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
ACM Trans. Archit. Code Optim.6
2014 Agility and Performance in Elastic Distributed Storage
abstract
Elastic storage systems can be expanded or contracted to meet current demand, allowing servers to be turned off or used for other tasks. However, the usefulness of an elastic distributed storage system is limited by its agility: how quickly it can increase or decrease its number of servers. Due to the large amount of data they must migrate during elastic resizing, state of the art designs usually have to make painful trade-offs among performance, elasticity, and agility. This article describes the state of the art in elastic storage and a new system, called SpringFS, that can quickly change its number of active servers, while retaining elasticity and performance goals. SpringFS uses a novel technique, termed bounded write offloading , that restricts the set of servers where writes to overloaded servers are redirected. This technique, combined with the read offloading and passive migration policies used in SpringFS, minimizes the work needed before deactivation or activation of servers. Analysis of real-world traces from Hadoop deployments at Facebook and various Cloudera customers and experiments with the SpringFS prototype confirm SpringFS’s agility, show that it reduces the amount of data migrated for elastic resizing by up to two orders of magnitude, and show that it cuts the percentage of active servers required by 67--82%, outdoing state-of-the-art designs by 6--120%.
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
ACM Trans. Storage6
2013 Linearly compressed pages: a low-complexity, low-latency main memory compression framework
abstract
Data compression is a promising approach for meeting the increasing memory capacity demands expected in future systems. Unfortunately, existing compression algorithms do not translate well when directly applied to main memory because they require the memory controller to perform non-trivial computation to locate a cache line within a compressed memory page, thereby increasing access latency and degrading system performance. Prior proposals for addressing this performance degradation problem are either costly or energy inefficient.
Gennady Pekhimenko, Vivek Seshadri, Yoongu Kim, Hongyi Xin, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
MICRO7
2013 RowClone: fast and energy-efficient in-DRAM bulk data copy and initialization
abstract
Several system-level operations trigger bulk data copy or initialization. Even though these bulk data operations do not require any computation, current systems transfer a large quantity of data back and forth on the memory channel to perform such operations. As a result, bulk data operations consume high latency, bandwidth, and energy--degrading both system performance and energy efficiency.
Vivek Seshadri, Yoongu Kim, Chris Fallin, Donghyuk Lee, Rachata Ausavarungnirun, Gennady Pekhimenko, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
MICRO10
2012 Chrysalis analysis: incorporating synchronization arcs in dataflow-analysis-based parallel monitoring
abstract
Software lifeguards, or tools that monitor applications at runtime, are an effective way of identifying program errors and security exploits. Parallel programs are susceptible to a wider range of possible errors than sequential programs, making them even more in need of online monitoring. Unfortunately, monitoring parallel applications is difficult due to inter-thread data dependences. In prior work, we introduced a new software framework for online parallel program monitoring inspired by dataflow analysis, called Butterfly Analysis. Butterfly Analysis uses bounded windows of uncertainty to model the finite upper bound on delay between when an instruction is issued and when all its effects are visible throughout the system. While Butterfly Analysis offers many advantages, it ignored one key source of ordering information which affected its false positive rate: explicit software synchronization, and the corresponding high-level happens-before arcs.
Michelle L. Goodstein, Shimin Chen, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
PACT4
2012 Base-delta-immediate compression: practical data compression for on-chip caches
abstract
Cache compression is a promising technique to increase on-chip cache capacity and to decrease on-chip and off-chip bandwidth usage. Unfortunately, directly applying well-known compression algorithms (usually implemented in software) leads to high hardware complexity and unacceptable decompression/compression latencies, which in turn can negatively affect performance. Hence, there is a need for a simple yet efficient compression technique that can effectively compress common in-cache data patterns, and has minimal effect on cache access latency.
Gennady Pekhimenko, Vivek Seshadri, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
PACT5
2012 The evicted-address filter: a unified mechanism to address both cache pollution and thrashing
abstract
Off-chip main memory has long been a bottleneck for system performance. With increasing memory pressure due to multiple on-chip cores, effective cache utilization is important. In a system with limited cache space, we would ideally like to prevent 1) cache pollution, i.e., blocks with low reuse evicting blocks with high reuse from the cache, and 2) cache thrashing, i.e., blocks with high reuse evicting each other from the cache.
Vivek Seshadri, Onur Mutlu, Michael A. Kozuch, Todd C. Mowry
PACT3
2012 Heterogeneity and dynamicity of clouds at scale: Google trace analysis
abstract
To better understand the challenges in developing effective cloud-based resource schedulers, we analyze the first publicly available trace data from a sizable multi-purpose cluster. The most notable workload characteristic is heterogeneity: in resource types (e.g., cores:RAM per machine) and their usage (e.g., duration and resources needed). Such heterogeneity reduces the effectiveness of traditional slot- and core-based scheduling. Furthermore, some tasks are constrained as to the kind of machine types they can use, increasing the complexity of resource assignment and complicating task migration. The workload is also highly dynamic, varying over time and most workload features, and is driven by many short jobs that demand quick scheduling decisions. While few simplifying assumptions apply, we find that many longer-running jobs have relatively stable resource utilizations, which can help adaptive resource schedulers.
Charles Reiss, Alexey Tumanov, Gregory R. Ganger, Randy H. Katz, Michael A. Kozuch
SoCC5
2012 alsched: algebraic scheduling of mixed workloads in heterogeneous clouds
abstract
As cloud resources and applications grow more heterogeneous, allocating the right resources to different tenants' activities increasingly depends upon understanding tradeoffs regarding their individual behaviors. One may require a specific amount of RAM, another may benefit from a GPU, and a third may benefit from executing on the same rack as a fourth. This paper promotes the need for and an approach for accommodating diverse tenant needs, based on having resource requests indicate any soft (i.e., when certain resource types would be better, but are not mandatory) and hard constraints in the form of composable utility functions. A scheduler that accepts such requests can then maximize overall utility, perhaps weighted by priorities, taking into account application specifics. Experiments with a prototype scheduler, called alsched, demonstrate that support for soft constraints is important for efficiency in multi-purpose clouds and that composable utility functions can provide it.
Alexey Tumanov, James Cipar, Gregory R. Ganger, Michael A. Kozuch
SoCC4
2012 SOFTScale: Stealing Opportunistically for Transient Scaling
Anshul Gandhi, Timothy Zhu, Mor Harchol-Balter, Michael A. Kozuch
Middleware4
2012 AutoScale: Dynamic, Robust Capacity Management for Multi-Tier Data Centers
abstract
Energy costs for data centers continue to rise, already exceeding $15 billion yearly. Sadly much of this power is wasted. Servers are only busy 10--30% of the time on average, but they are often left on, while idle, utilizing 60% or more of peak power when in the idle state. We introduce a dynamic capacity management policy, AutoScale , that greatly reduces the number of servers needed in data centers driven by unpredictable, time-varying load, while meeting response time SLAs. AutoScale scales the data center capacity, adding or removing servers as needed. AutoScale has two key features: (i) it autonomically maintains just the right amount of spare capacity to handle bursts in the request rate; and (ii) it is robust not just to changes in the request rate of real-world traces, but also request size and server efficiency. We evaluate our dynamic capacity management approach via implementation on a 38-server multi-tier data center, serving a web site of the type seen in Facebook or Amazon, with a key-value store workload. We demonstrate that AutoScale vastly improves upon existing dynamic capacity management policies with respect to meeting SLAs and robustness.
Anshul Gandhi, Mor Harchol-Balter, Ram Raghunathan, Michael A. Kozuch
ACM Trans. Comput. Syst.4
2011 Switching the optical divide: fundamental challenges for hybrid electrical/optical datacenter networks
abstract
Recent proposals to build hybrid electrical (packet-switched) and optical (circuit switched) data center interconnects promise to reduce the cost, complexity, and energy requirements of very large data center networks. Supporting realistic traffic patterns, however, exposes a number of unexpected and difficult challenges to actually deploying these systems "in the wild." In this paper, we explore several of these challenges, uncovered during a year of experience using hybrid interconnects. We discuss both the problems that must be addressed to make these interconnects truly useful, and the implications of these challenges on what solutions are likely to be ultimately feasible.
Hamid Hajabdolali Bazzaz, Malveeka Tewari, George Porter, T. S. Eugene Ng, David G. Andersen, Michael Kaminsky, Michael A. Kozuch, Amin Vahdat
SoCC8
2010 Butterfly analysis: adapting dataflow analysis to dynamic parallel monitoring
abstract
Online program monitoring is an effective technique for detecting bugs and security attacks in running applications. Extending these tools to monitor parallel programs is challenging because the tools must account for inter-thread dependences and relaxed memory consistency models. Existing tools assume sequential consistency and often slow down the monitored program by orders of magnitude. In this paper, we present a novel approach that avoids these pitfalls by not relying on strong consistency models or detailed inter-thread dependence tracking. Instead, we only assume that events in the distant past on all threads have become visible; we make no assumptions on (and avoid the overheads of tracking) the relative ordering of more recent events on other threads. To overcome the potential state explosion of considering all the possible orderings among recent events, we adapt two techniques from static dataflow analysis, reaching definitions and reaching expressions, to this new domain of dynamic parallel monitoring. Significant modifications to these techniques are proposed to ensure the correctness and efficiency of our approach. We show how our adapted analysis can be used in two popular memory and security tools. We prove that our approach does not miss errors, and sacrifices precision only due to the lack of a relative ordering among recent events. Moreover, our simulation study on a collection of Splash-2 and Parsec 2.0 benchmarks running a memory-checking tool on a hardware-assisted logging platform demonstrates the potential benefits in trading off a very low false positive rate for (i) reduced overhead and (ii) the ability to run on relaxed consistency models.
Michelle L. Goodstein, Evangelos Vlachos, Shimin Chen, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry
ASPLOS5
2010 ParaLog: enabling and accelerating online parallel monitoring of multithreaded applications
abstract
Instruction-grain lifeguards monitor the events of a running application at the level of individual instructions in order to identify and help mitigate application bugs and security exploits. Because such lifeguards impose a 10-100X slowdown on existing platforms, previous studies have proposed hardware designs to accelerate lifeguard processing. However, these accelerators are either tailored to a specific class of lifeguards or suitable only for monitoring singlethreaded programs.
Evangelos Vlachos, Michelle L. Goodstein, Michael A. Kozuch, Shimin Chen, Babak Falsafi, Phillip B. Gibbons, Todd C. Mowry
ASPLOS3
2010 Robust and flexible power-proportional storage
abstract
Power-proportional cluster-based storage is an important component of an overall cloud computing infrastructure. With it, substantial subsets of nodes in the storage cluster can be turned off to save power during periods of low utilization. Rabbit is a distributed file system that arranges its data-layout to provide ideal power-proportionality down to very low minimum number of powered-up nodes (enough to store a primary replica of available datasets). Rabbit addresses the node failure rates of large-scale clusters with data layouts that minimize the number of nodes that must be powered-up if a primary fails. Rabbit also allows different datasets to use different subsets of nodes as a building block for interference avoidance when the infrastructure is shared by multiple tenants. Experiments with a Rabbit prototype demonstrate its power-proportionality, and simulation experiments demonstrate its properties at scale.
Hrishikesh Amur, James Cipar, Varun Gupta 0004, Gregory R. Ganger, Michael A. Kozuch, Karsten Schwan
SoCC5
2010 c-Through: part-time optics in data centers
abstract
Data-intensive applications that operate on large volumes of data have motivated a fresh look at the design of data center networks. The first wave of proposals focused on designing pure packet-switched networks that provide full bisection bandwidth. However, these proposals significantly increase network complexity in terms of the number of links and switches required and the restricted rules to wire them up. On the other hand, optical circuit switching technology holds a very large bandwidth advantage over packet switching technology. This fact motivates us to explore how optical circuit switching technology could benefit a data center network. In particular, we propose a hybrid packet and circuit switched data center network architecture (or HyPaC for short) which augments the traditional hierarchy of packet switches with a high speed, low complexity, rack-to-rack optical circuit-switched network to supply high bandwidth to applications. We discuss the fundamental requirements of this hybrid architecture and their design options. To demonstrate the potential benefits of the hybrid architecture, we have built a prototype system called c-Through. c-Through represents a design point where the responsibility for traffic demand estimation and traffic demultiplexing resides in end hosts, making it compatible with existing packet switches. Our emulation experiments show that the hybrid architecture can provide large benefits to unmodified popular data center applications at a modest scale. Furthermore, our experimental experience provides useful insights on the applicability of the hybrid architecture across a range of deployment scenarios.
David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, T. S. Eugene Ng, Michael A. Kozuch, Michael P. Ryan
SIGCOMM6
2010 Optimality analysis of energy-performance trade-off for server farm management
Anshul Gandhi, Varun Gupta 0004, Mor Harchol-Balter, Michael A. Kozuch
Perform. Evaluation4
2009 Cluster fault-tolerance: An experimental evaluation of checkpointing and MapReduce through simulation
abstract
Traditionally, cluster computing has employed checkpointing to address fault tolerance. Recently, new models for parallel applications have grown in popularity namely MapReduce and Dryad, with runtime systems providing their own re-execute based fault tolerance mechanisms, but with no analysis of their failure characteristics. Another development is the availability of failure data spanning years for systems of significant size at Los Alamos National Labs (LANL), but the time between failure (TBF) for these systems is a poor fit to the exponential distribution assumed by optimization work in checkpointing, bringing these results into question. The work in this paper describes a discrete event simulation driven by the LANL data and by models of parallel checkpointing and MapReduce tasks. The simulation allows us to then evaluate and assess the fault tolerance characteristics of these tasks with the goal of minimizing the expected running time of a parallel program in a cluster in the presence of faults for both fault tolerance models.
Thomas C. Bressoud, Michael A. Kozuch
CLUSTER2
2009 Your Data Center Is a Router: The Case for Reconfigurable Optical Circuit Switched Paths
David G. Andersen, Michael Kaminsky, Michael A. Kozuch, T. S. Eugene Ng, Konstantina Papagiannaki, Madeleine Glick, Lily B. Mummert
HotNets4
2009 Migration without Virtualization
Michael A. Kozuch, Michael Kaminsky, Michael P. Ryan
HotOS1
2008 Flexible Hardware Acceleration for Instruction-Grain Program Monitoring
abstract
Instruction-grain program monitoring tools, which check and analyze executing programs at the granularity of individual instructions, are invaluable for quickly detecting bugs and security attacks and then limiting their damage (via containment and/or recovery). Unfortunately, their fine-grain nature implies very high monitoring overheads for software-only tools, which are typically based on dynamic binary instrumentation. Previous hardware proposals either focus on mechanisms that target specific bugs or address only the cost of binary instrumentation. In this paper, we propose a flexible hardware solution for accelerating a wide range of instruction-grain monitoring tools. By examining a number of diverse tools (for memory checking, security tracking, and data race detection), we identify three significant common sources of overheads and then propose three novel hardware techniques for addressing these overheads: Inheritance Tracking, Idempotent Filters, and Metadata-TLBs. Together, these constitute a general-purpose hardware acceleration framework. Experimental results show our framework reduces overheads by 2-3X over the previous state-of-the-art, while supporting the needed flexibility.
Shimin Chen, Michael A. Kozuch, Theodoros Strigkos, Babak Falsafi, Phillip B. Gibbons, Todd C. Mowry, Vijaya Ramachandran, Olatunji Ruwase, Michael P. Ryan, Evangelos Vlachos
ISCA2
2008 Provably good multicore cache performance for divide-and-conquer algorithms
Guy E. Blelloch, Rezaul Alam Chowdhury, Phillip B. Gibbons, Vijaya Ramachandran, Shimin Chen, Michael A. Kozuch
SODA6
2008 Parallelizing dynamic information flow tracking
abstract
Dynamic information flow tracking (DIFT) is an important tool for detecting common security attacks and memory bugs. A DIFT tool tracks the flow of information through a monitored program's registers and memory locations as the program executes, detecting and containing/fixing problems on-the-fly. Unfortunately, sequential DIFT tools are quite slow, and DIFT is quite challenging to parallelize. In this paper, we present a new approach to parallelizing DIFT-like functionality. Extending our recent work on accelerating sequential DIFT, we consider a variant of DIFT that tracks the information flow only through unary operations relaxed DIFT, and yet makes sense for detecting security attacks and memory bugs. We present a parallel algorithm for relaxed DIFT, based on symbolic inheritance tracking, which achieves linear speed-up asymptotically. Moreover, we describe techniques for reducing the constant factors, so that speed-ups can be obtained even with just a few processors. We implemented the algorithm in the context of a Log-Based Architectures (LBA) system, which provides hardware support for logging a program trace and delivering it to other (monitoring) processors. Our simulation results on SPEC benchmarks and a video player show that our parallel relaxed DIFT reduces the overhead to as low as 1.2X using 9 monitoring cores on a 16-core chip multiprocessor.
Olatunji Ruwase, Phillip B. Gibbons, Todd C. Mowry, Vijaya Ramachandran, Shimin Chen, Michael A. Kozuch, Michael P. Ryan
SPAA6
2008 Adaptive File Transfers for Diverse Environments
Himabindu Pucha, Michael Kaminsky, David G. Andersen, Michael A. Kozuch
USENIX ATC4
2007 Scheduling threads for constructive cache sharing on CMPs
abstract
In chip multiprocessors (CMPs), limiting the number of offchip cache misses is crucial for good performance. Many multithreaded programs provide opportunities for constructive cache sharing, in which concurrently scheduled threads share a largely overlapping working set. In this paper, we compare the performance of two state-of-the-art schedulers proposed for fine-grained multithreaded programs: Parallel Depth First (PDF), which is specifically designed for constructive cache sharing, and Work Stealing (WS), which is a more traditional design. Our experimental results indicate that PDF scheduling yields a 1.3--1.6X performance improvement relative to WS for several fine-grain parallel benchmarks on projected future CMP configurations; we also report several issues that may limit the advantage of PDF in certain applications. These results also indicate that PDF more effectively utilizes off-chip bandwidth, making it possible to trade-off on-chip cache for a larger number of cores. Moreover, we find that task granularity plays a key role in cache performance. Therefore, we present an automatic approach for selecting effective grain sizes, based on a new working set profiling algorithm that is an order of magnitude faster than previous approaches. This is the first paper demonstrating the effectiveness of PDF on real benchmarks, providing a direct comparison between PDF and WS, revealing the limiting factors for PDF in practice, and presenting an approach for overcoming these factors.
Shimin Chen, Phillip B. Gibbons, Michael A. Kozuch, Vasileios Liaskovitis, Anastasia Ailamaki, Guy E. Blelloch, Babak Falsafi, Limor Fix, Nikos Hardavellas, Todd C. Mowry, Chris Wilkerson
SPAA3
2006 Parallel depth first vs. work stealing schedulers on CMP architectures
abstract
In chip multiprocessors (CMPs), limiting the number of off-chip cache misses is crucial for good performance. Many multithreaded programs provide opportunities for constructive cache sharing, in which concurrently scheduled threads share a largely overlapping working set. In this brief announcement, we highlight our ongoing study [4] comparing the performance of two schedulers designed for fine-grained multithreaded programs: Parallel Depth First (PDF) [2], which is designed for constructive sharing, and Work Stealing (WS) [3], which takes a more traditional approach.Overview of schedulers. In PDF, processing cores are allocated ready-to-execute program tasks such that higher scheduling priority is given to those tasks the sequential program would have executed earlier. As a result, PDF tends to co-schedule threads in a way that tracks the sequential execution. Hence, the aggregate working set is (provably) not much larger than the single thread working set [1]. In WS, each processing core maintains a local work queue of readyto-execute threads. Whenever its local queue is empty, the core steals a thread from the bottom of the first non-empty queue it finds. WS is an attractive scheduling policy because when there is plenty of parallelism, stealing is quite rare. However, WS is not designed for constructive cache sharing, because the cores tend to have disjoint working sets.CMP configurations studied. We evaluated the performance of PDF and WS across a range of simulated CMP configurations. We focused on designs that have fixed-size private L1 caches and a shared L2 cache on chip. For a fixed die size (240 mm2), we varied the number of cores from 1 to 32. For a given number of cores, we used a (default) configuration based on current CMPs and realistic projections of future CMPs, as process technologies decrease from 90nm to 32nm.Summary of findings. We studied a variety of benchmark programs to show the following findings.For several application classes, PDF enables significant constructive sharing between threads, leading to better utilization of the on-chip caches and reducing off-chip traffic compared to WS. In particular, bandwidth-limited irregular programs and parallel divide-and-conquer programs present a relative speedup of 1.3-1.6X over WS, observing a 13- 41% reduction in off-chip traffic. An example is shown in Figure 1, for parallel merge sort. For each schedule, the number of L2 misses (i.e., the off-chip traffic) is shown on the left and the speed-up over running on one core is shown on the right, for 1 to 32 cores. Note that reducing the offchip traffic has the additional benefit of reducing the power consumption. Moreover, PDF's smaller working sets provide opportunities to power down segments of the cache without increasing the running time. Furthermore, when multiple programs are active concurrently, the PDF version is also less of a cache hog and its smaller working set is more likely to remain in the cache across context switches.For several other applications classes, PDF and WS have roughly the same execution times, either because there is only limited data reuse that can be exploited or because the programs are not limited by off-chip bandwidth. In the latter case, the constructive sharing PDF enables does provide the power and multiprogramming benefits discussed above.Finally, most parallel benchmarks to date, written for SMPs, use such a coarse-grained threading that they cannot exploit the constructive cache behavior inherent in PDF.We find that mechanisms to finely grain multithreaded applications are crucial to achieving good performance on CMPs.
Vasileios Liaskovitis, Shimin Chen, Phillip B. Gibbons, Anastasia Ailamaki, Guy E. Blelloch, Babak Falsafi, Limor Fix, Nikos Hardavellas, Michael A. Kozuch, Todd C. Mowry, Chris Wilkerson
SPAA9
2006 Design Tradeoffs in Applying Content Addressable Storage to Enterprise-scale Systems Based on Virtual Machines
Partho Nath, Michael A. Kozuch, David R. O'Hallaron, Jan Harkes, Mahadev Satyanarayanan, Niraj Tolia, Matt Toups
USENIX ATC, General Track2
2005 Towards seamless mobility on pervasive hardware
Mahadev Satyanarayanan, Michael A. Kozuch, Casey Helfrich, David R. O'Hallaron
Pervasive Mob. Comput.2
2004 Integrating Portable and Distributed Storage
Niraj Tolia, Jan Harkes, Michael A. Kozuch, Mahadev Satyanarayanan
FAST3
2003 Opportunistic Use of Content Addressable Storage for Distributed File Systems
Niraj Tolia, Michael A. Kozuch, Mahadev Satyanarayanan, Brad Karp, Thomas C. Bressoud, Adrian Perrig
USENIX ATC, General Track2
2000 An Experimental Analysis of Digital Video Library Servers
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
Multim. Syst.1
1997 An Approach to Network Caching for Multimedia Objects
abstract
Caching is an important mechanism for improving both the performance and operational cost of multimedia networks. This paper presents a new approach to network caching for large multimedia objects, the Caching Tree Algorithm, which is based on previous research regarding the File Allocation Problem. This novel algorithm is both distributed and computationally tractable. We also describe some practical considerations in the application of the algorithm to modern networks.
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
ICCD1
1996 New Challenges for Video Servers: Performance of Non-Linear Applications under User Choice
abstract
This paper outlines a framework for classifying video application types according to linearity and user-choice response constraint. It then presents the first analyses of video server performance under various non-linear video application loads.
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
ICCD1
1994 Compression of Embedded System Programs
abstract
Embedded systems are often sensitive to space, weight, and cost considerations. Reducing the size of stored programs can significantly improve these factors. This paper discusses a program compression methodology based on existing processor architectures. The authors examine practical and theoretical measures for the maximum compression rate of a suite of programs across six modern architectures. The theoretical compression rate is reported in terms of the zeroth and first-order entropies, while the practical compression rate is reported in terms of the Huffman-encoded format of the proposed compression methodology and the GNU file compression utility, gzip. These experiments indicate that a practical increase of 15%-30% and a theoretical increase of over 100% in code density can be expected using the techniques examined. In addition, a novel, greedy, variable-length-to variable-length encoding algorithm is presented with preliminary results.>
Michael A. Kozuch, Andrew Wolfe
ICCD1