Cristian Ungureanu

dblp:51/6366 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-authorDatabases, data management, data science and information retrieval · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Storage systems · 34% Distributed systems · 32% Parallel and multicore computing · 11%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
data consistency
0.422015
Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015
Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015
Parallel and multicore computing › parallel computing › parallel machine learning
data-parallel training
0.212015
MALT: distributed data-parallelism for existing ML applications · EuroSys 2015
Distributed systems
data synchronization
0.212015
Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015
Parallel and multicore computing › data-parallel programming
distributed data-parallel execution
0.212015
MALT: distributed data-parallelism for existing ML applications · EuroSys 2015
Distributed systems
distributed machine learning
0.212015
MALT: distributed data-parallelism for existing ML applications · EuroSys 2015
Storage systems
mobile applications
0.212015
Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015
Cloud and datacenter computing
mobile cloud computing
0.212015
Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015
Distributed systems › consistency models
tunable consistency
0.212015
Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015
Embedded and real-time systems
mobile computing
0.222012
Revisiting storage for smartphones · ACM Trans. Storage 2012
Revisiting storage for smartphones · FAST 2012
Memory systems › cache management
cache replacement
0.212013
TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013
Memory systems › cache management › storage caching
flash cache
0.212013
TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013
Storage systems › flash and SSD › flash memory
flash storage
0.112012
Revisiting storage for smartphones · ACM Trans. Storage 2012
Storage systems › storage devices › storage media
mobile storage
0.112012
Revisiting storage for smartphones · FAST 2012
Embedded and real-time systems › mobile computing
smartphone storage
0.112012
Revisiting storage for smartphones · FAST 2012
Distributed systems
fault tolerance
0.122015
Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015
Failure detection and localization in component based systems by online tracking · KDD 2005
Storage systems
content-addressable storage
0.112010
HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010
Storage systems › data reduction › data deduplication
content-defined chunking
0.112010
Bimodal Content Defined Chunking for Backup Streams · FAST 2010
Storage systems › data reduction
data deduplication
0.112010
Bimodal Content Defined Chunking for Backup Streams · FAST 2010
Storage systems › file systems
distributed file system
0.112010
HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010
Storage systems
file systems
0.112010
HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010
Storage systems
distributed storage
0.112009
HYDRAstor: A Scalable Secondary Storage · FAST 2009
Storage systems
scalable storage
0.112009
HYDRAstor: A Scalable Secondary Storage · FAST 2009
Storage systems › storage hierarchy
secondary storage
0.112009
HYDRAstor: A Scalable Secondary Storage · FAST 2009
Machine learning and data management
scalable machine learning
0.112015
MALT: distributed data-parallelism for existing ML applications · EuroSys 2015
Memory systems
cache management
0.012013
TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013
Distributed systems › peer-to-peer systems
distributed hash table
0.012004
FPN: A Distributed Hash Table for Commercial Applications · HPDC 2004
Storage systems › storage management › backup and recovery
data backup
0.012010
Bimodal Content Defined Chunking for Backup Streams · FAST 2010

Methods — techniques the papers use, named apart from their topics

data parallelism · 0.4bloom filter · 0.2LRU approximation · 0.2storage characterization · 0.1pilot solution design · 0.1measurement study · 0.1subspace decomposition · 0.1squared prediction error · 0.1sequentially discounting expectation maximization · 0.1hotelling t2 · 0.1
YearPublicationVenuePosition
2015 MALT: distributed data-parallelism for existing ML applications
abstract
Machine learning methods, such as SVM and neural networks, often improve their accuracy by using models with more parameters trained on large numbers of examples. Building such models on a single machine is often impractical because of the large amount of computation required.
Hao Li 0022, Asim Kadav, Erik Kruus, Cristian Ungureanu
EuroSys4
2015 Simba: tunable end-to-end data consistency for mobile apps
abstract
Developers of cloud-connected mobile apps need to ensure the consistency of application and user data across multiple devices. Mobile apps demand different choices of distributed data consistency under a variety of usage scenarios. The apps also need to gracefully handle intermittent connectivity and disconnections, limited bandwidth, and client and server failures. The data model of the apps can also be complex, spanning inter-dependent structured and unstructured data, and needs to be atomically stored and updated locally, on the cloud, and on other mobile devices.
Dorian Jean Perkins, Nitin Agrawal 0001, Akshat Aranya, Curtis Yu, Younghwan Go, Harsha V. Madhyastha, Cristian Ungureanu
EuroSys7
2015 Reliable, Consistent, and Efficient Data Sync for Mobile Apps
Younghwan Go, Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu
FAST4
2013 Mobile Data Sync in a Blink
Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu
HotStorage3
2013 TBF: A memory-efficient replacement policy for flash-based caches
abstract
The performance and capacity characteristics of flash storage make it attractive to use as a cache. Recency-based cache replacement policies rely on an in-memory full index, typically a B-tree or a hash table, that maps each object to its recency information. Even though the recency information itself may take very little space, the full index for a cache holding N keys requires at least log N bits per key. This metadata overhead is undesirably high when used for very large flash-based caches, such as key-value stores with billions of objects. To solve this problem, we propose a new RAM-frugal cache replacement policy that approximates the least-recently-used (LRU) policy. It uses two in-memory Bloom sub-filters (TBF) for maintaining the recency information and leverages an on-flash key-value store to cache objects. TBF requires only one byte of RAM per cached object, making it suitable for implementing very large flash-based caches. We evaluate TBF through simulation on traces from several block stores and key-value stores, as well as evaluate it using the Yahoo! Cloud Serving Benchmark in a real system implementation. Evaluation results show that TBF achieves cache hit rate and operations per second comparable to those of LRU in spite of its much smaller memory requirements.
Cristian Ungureanu, Biplob Debnath, Stephen Rago, Akshat Aranya
ICDE1
2013 Building a Delay-Tolerant Cloud for Mobile Data
abstract
Mobile data usage is on a tremendous rise, due not only to increasing number of users but also to an increase in the number of applications that transfer data over the network. Moreover, applications for sharing, sensing, and collaboration have become more popular, causing significant amounts of data to be generated on devices. Managing this data -syncing it to the cloud, or with other users or devices- is a crucial and often challenging part of writing mobile apps and services. In spite of plenty of good advice and best practices from OS vendors and network operators, storing and transferring mobile data is fraught with issues. On the one hand, an app developer needs to worry about the semantics of data storage and synchronization, while on the other, about the end-user experience, which maybe impacted by poor and intermittent network connectivity. To address the needs of the app developers and the end-users, we have built Izzy: a platform to rapidly develop and deploy data-centric mobile apps. Izzy provides well-defined and easy to use semantics for accessing local storage and for synchronizing data with a remote, scalable, global store. Izzy also provides global store access to the cloud-resident part of the applications (if any) through a similar server API. Last but not least, Izzy is designed to be frugal: it conserves mobile device resources by applying delay-tolerance and data reduction techniques (message coalescing and compression) across applications on a mobile device. In this paper we present the design of Izzy and our early experiences with using it.
Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu
MDM (1)4
2012 Revisiting storage for smartphones
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu
FAST3
2012 Revisiting storage for smartphones
abstract
Conventional wisdom holds that storage is not a big contributor to application performance on mobile devices. Flash storage (the type most commonly used today) draws little power, and its performance is thought to exceed that of the network subsystem. In this article, we present evidence that storage performance does indeed affect the performance of several common applications such as Web browsing, maps, application install, email, and Facebook. For several Android smartphones, we find that just by varying the underlying flash storage, performance over WiFi can typically vary between 100% and 300% across applications; in one extreme scenario, the variation jumped to over 2000%. With a faster network (set up over USB), the performance variation rose even further. We identify the reasons for the strong correlation between storage and application performance to be a combination of poor flash device performance, random I/O from application databases, and heavy-handed use of synchronous writes. Based on our findings, we implement and evaluate a set of pilot solutions to address the storage performance deficiencies in smartphones.
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu
ACM Trans. Storage3
2011 Using Eager Strategies to Improve NFS I/O Performance
abstract
Typical NFS clients write in a lazy fashion: they leave dirty pages in the page cache and defer writing to the server until later. This reduces network traffic when applications repeatedly modify the same set of pages. However, this approach can lead to memory pressure, when the number of available pages on the client system is so low that the system must work harder to reclaim dirty pages. System performance is poor under memory pressure. We show examples of this problem and present two mechanisms to solve it: eager write back and eager page laundering. These mechanisms change the client's data management policy from lazy to eager, resulting in higher throughput for sequential writes. In addition, we show that NFS servers suffer from out-of-order file operations, which further reduce performance. We introduce request ordering, a server mechanism to process operations (as much as possible) in the order they were sent by the client, which improves read performance substantially. We have implemented these techniques in the Linux operating system. I/O performance is improved, with the most pronounced improvement visible for sequential access to large files. We see about 33% improvement in the performance of streaming write workloads and more than triple the performance of streaming read workloads. We evaluate several nonsequential workloads and show that these techniques do not degrade performance, and can sometimes improve performance. We also design and evaluate an adversarial workload to show that the eager policies can perform worse in some pathological cases.
Stephen Rago, Aniruddha Bohra, Cristian Ungureanu
NAS3
2010 Bimodal Content Defined Chunking for Backup Streams
Erik Kruus, Cristian Ungureanu, Cezary Dubnicki
FAST2
2010 HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System
Cristian Ungureanu, Benjamin Atkin, Akshat Aranya, Salil Gokhale, Stephen Rago, Grzegorz Calkowski, Cezary Dubnicki, Aniruddha Bohra
FAST1
2010 KVZone and the Search for a Write-Optimized Key-Value Store
Salil Gokhale, Nitin Agrawal 0001, Sean Noonan, Cristian Ungureanu
HotStorage4
2009 HYDRAstor: A Scalable Secondary Storage
Cezary Dubnicki, Leszek Gryz, Lukasz Heldt, Michal Kaczmarczyk, Wojciech Kilian, Przemyslaw Strzelczak, Jerzy Szczepkowski, Cristian Ungureanu, Michal Welnicki
FAST8
2007 Online Tracking of Component Interactions for Failure Detection and Localization in Distributed Systems
abstract
This paper proposes a novel failure-detection approach that can handle high-dimensional observation and frequent system changes. We extract two statistics from the subspace decomposition of observations, and use the mixture of Gaussians to model their probability density. Instead of monitoring the original data, the density model of extracted statistics is adaptively updated and examined regularly to detect failures. We also present a localization method to identify the faulty components once the failure happens. Applying our technique to monitor the component interactions in an e-commerce application shows satisfactory results in detecting a variety of injected failures.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C3
2007 Multiresolution Abnormal Trace Detection Using Varied-Length n-Grams and Automata
abstract
Detection and diagnosis of faults in a large-scale distributed system is a formidable task. Interest in monitoring and using traces of user requests for fault detection has been on the rise recently. In this paper we propose novel fault detection methods based on abnormal trace detection. One essential problem is how to represent the large amount of training trace data compactly as an oracle. Our key contribution is the novel use of varied-length n-grams and automata to characterize normal traces. A new trace is compared against the learned automata to determine whether it is abnormal. We develop algorithms to automatically extract n-grams and construct multiresolution automata from training data. Further, both deterministic and multihypothesis algorithms are proposed for detection. We inspect the trace constraints of real application software and verify the existence of long n-grams. Our approach is tested in a real system with injected faults and achieves good results in experiments
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C3
2005 Failure detection and localization in component based systems by online tracking
abstract
The increasing complexity of today's systems makes fast and accurate failure detection essential for their use in mission-critical applications. Various monitoring methods provide a large amount of data about system's behavior. Analyzing this data with advanced statistical methods holds the promise of not only detecting the errors faster, but also detecting errors which are difficult to catch with current monitoring tools. Two challenges to building such detection tools are: the high dimensionality of observation data, which makes the models expensive to apply, and frequent system changes, which make the models expensive to update. In this paper, we present algorithms to reduce the dimensionality of data in a way that makes it easy to adapt to system changes. We decompose the observation data into signal and noise subspaces. Two statistics, the Hotelling T2 score and squared prediction error (SPE) are calculated to represent the data characteristics in signal and noise subspaces respectively. Instead of tracking the original data, we use a sequentially discounting expectation maximization (SDEM) algorithm to learn the distribution of the two extracted statistics. A failure event can then be detected based on the abnormal change of the distribution. Applying our technique to component interaction data in a simple e-commerce application shows better accuracy than building independent profiles for each component. Additionally, experiments on synthetic data show that the detection accuracy is high even for changing systems.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
KDD3
2004 FPN: A Distributed Hash Table for Commercial Applications
Cezary Dubnicki, Cristian Ungureanu, Wojciech Kilian
HPDC2
2000 Concurrency Analysis for Java
Cristian Ungureanu, Suresh Jagannathan
SAS1
1997 Formal Models of Distributed Memory Management
abstract
We develop am abstract model of memory management in distributed systems. The model is low-level enough so that we can express communication, allocation and garbage collection, but otherwise hides many of the lower-level details of an actual implementation.Recently, such formal models have been developed for memory management in a functional, sequential setting [10]. The models are rewriting systems whose terms are programs. Programs have both the "code" (control string) and the "store" syntactically apparent. Evaluation is expressed as conditional rewriting and includes store operations. Garbage collection becomes a rewriting relation that removes part of the store without affecting the behavior of the program.Distribution adds another dimension to an already complex problem. By using techniques developed for communicating and concurrent systems [9], we extend their work to a distributed environment. Sending and receiving messages is also made apparent at the syntactic level. A very general garbage collection rule based on reachability is introduced and proved correct. Now proving correct a specific collection strategy is reduced to showing that the relation between programs defined by the strategy is a sub-relation of the general relation. Any actual implementation which is capable of providing the transitions (including their atomicity constraints) specified by the strategy is therefore correct.This model allows us to specify and prove correct in a compact manner two garbage collectors; the first one does a simple garbage collection local to a node. The second garbage collector uses migration of data in order to be able to reclaim inter-node cyclic garbage.
Cristian Ungureanu, Benjamin Goldberg 0001
ICFP1
1997 Run-Time versus Compile-Time Instruction Scheduling in Superscalar (RISC) Processors: Performance and Trade-Off
abstract
The RISC revolution has spurred the development of processors with increasing degrees ofinstruction level parallelism(ILP). In order to realize the full potential of these processors, multiple instructions must continuously be issued and executed in a single cycle. Consequently,instruction schedulingplays a crucial role as an optimization in this context. While early attempts at instruction scheduling were limited to compile-time approaches, the current trends are aimed at providingdynamicsupport in hardware. In this paper, we present the results of a detailed comparative study of the performance advantages to be derived by the spectrum of instruction scheduling approaches: from limited basic-block schedulers in the compiler, to novel and aggressive schedulers in hardware. A significant portion of our experimental study via simulations, is devoted to understanding the performance advantages of run-time scheduling. Our results indicate it to be effective in extracting the ILP inherent to the program trace being scheduled, over a wide range of machine and program parameters. Furthermore, we also show that this effectiveness can be further enhanced by a simple basic-block scheduler in the compiler, which optimizes for the presence of the run-time scheduler in the target; current basic-block schedulers are not designed to take advantage of this feature. We demonstrate this fact by presenting a novel basic-block scheduling algorithm that is sensitive to the lookahead hardware in the target processor. Finally, we outline a simple analytical characterization of the performance advantage that run-time schedulers have to offer.
Allen Leung, Krishna V. Palem, Cristian Ungureanu
J. Parallel Distributed Comput.3
1996 Run-time versus compile-time instruction scheduling in superscalar (RISC) processors: performance and tradeoffs
abstract
The RISC revolution has spurred the development of processors with increasing degrees of instruction level parallelism (ILP). In order to realize the full potential of these processors, multiple instructions must continuously be issued and executed in a single cycle. Consequently, instruction scheduling plays a crucial role as an optimization in this context. While early attempts at instruction scheduling were limited to compile-time approaches, the current trends are aimed at providing dynamic support in hardware. In this paper, we present the results of a detailed comparative study of the performance advantages to be derived by the spectrum of instruction scheduling approaches: from limited basic-block schedulers in the compiler, to novel and aggressive schedulers in hardware. A significant portion of our experimental study via simulations, is devoted to understanding the performance advantages of run-time scheduling. Our results indicate it to be effective in extracting the ILP inherent to the program trace being scheduled, over a wide range of machine and program parameters. Furthermore, we also show that this effectiveness can be further enhanced by a simple basic-block scheduler in the compiler, which optimizes for the presence of the run-time scheduler in the target; current basic-block schedulers are not designed to take advantage of this feature. We demonstrate this fact by presenting a novel basic-block scheduling algorithm that is sensitive to the lookahead hardware in the target processor.
Allen Leung, Krishna V. Palem, Cristian Ungureanu
HiPC3