VLDB 2026 Research / reviewers in the wild / expert
Cristian Ungureanu
dblp:51/6366
· DBLP profile ↗
21ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 1 first-authorDatabases, data management, data science and information retrieval · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Storage systems · 34% Distributed systems · 32% Parallel and multicore computing · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 27 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
data consistency |
0.4 | 2 | 2015 | Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015 Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015 |
Parallel and multicore computing › parallel computing › parallel machine learning
data-parallel training |
0.2 | 1 | 2015 | MALT: distributed data-parallelism for existing ML applications · EuroSys 2015 |
Distributed systems
data synchronization |
0.2 | 1 | 2015 | Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015 |
Parallel and multicore computing › data-parallel programming
distributed data-parallel execution |
0.2 | 1 | 2015 | MALT: distributed data-parallelism for existing ML applications · EuroSys 2015 |
Distributed systems
distributed machine learning |
0.2 | 1 | 2015 | MALT: distributed data-parallelism for existing ML applications · EuroSys 2015 |
Storage systems
mobile applications |
0.2 | 1 | 2015 | Reliable, Consistent, and Efficient Data Sync for Mobile Apps · FAST 2015 |
Cloud and datacenter computing
mobile cloud computing |
0.2 | 1 | 2015 | Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015 |
Distributed systems › consistency models
tunable consistency |
0.2 | 1 | 2015 | Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015 |
Embedded and real-time systems
mobile computing |
0.2 | 2 | 2012 | Revisiting storage for smartphones · ACM Trans. Storage 2012 Revisiting storage for smartphones · FAST 2012 |
Memory systems › cache management
cache replacement |
0.2 | 1 | 2013 | TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013 |
Memory systems › cache management › storage caching
flash cache |
0.2 | 1 | 2013 | TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013 |
Storage systems › flash and SSD › flash memory
flash storage |
0.1 | 1 | 2012 | Revisiting storage for smartphones · ACM Trans. Storage 2012 |
Storage systems › storage devices › storage media
mobile storage |
0.1 | 1 | 2012 | Revisiting storage for smartphones · FAST 2012 |
Embedded and real-time systems › mobile computing
smartphone storage |
0.1 | 1 | 2012 | Revisiting storage for smartphones · FAST 2012 |
Distributed systems
fault tolerance |
0.1 | 2 | 2015 | Simba: tunable end-to-end data consistency for mobile apps · EuroSys 2015 Failure detection and localization in component based systems by online tracking · KDD 2005 |
Storage systems
content-addressable storage |
0.1 | 1 | 2010 | HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010 |
Storage systems › data reduction › data deduplication
content-defined chunking |
0.1 | 1 | 2010 | Bimodal Content Defined Chunking for Backup Streams · FAST 2010 |
Storage systems › data reduction
data deduplication |
0.1 | 1 | 2010 | Bimodal Content Defined Chunking for Backup Streams · FAST 2010 |
Storage systems › file systems
distributed file system |
0.1 | 1 | 2010 | HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010 |
Storage systems
file systems |
0.1 | 1 | 2010 | HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System · FAST 2010 |
Storage systems
distributed storage |
0.1 | 1 | 2009 | HYDRAstor: A Scalable Secondary Storage · FAST 2009 |
Storage systems
scalable storage |
0.1 | 1 | 2009 | HYDRAstor: A Scalable Secondary Storage · FAST 2009 |
Storage systems › storage hierarchy
secondary storage |
0.1 | 1 | 2009 | HYDRAstor: A Scalable Secondary Storage · FAST 2009 |
Machine learning and data management
scalable machine learning |
0.1 | 1 | 2015 | MALT: distributed data-parallelism for existing ML applications · EuroSys 2015 |
Memory systems
cache management |
0.0 | 1 | 2013 | TBF: A memory-efficient replacement policy for flash-based caches · ICDE 2013 |
Distributed systems › peer-to-peer systems
distributed hash table |
0.0 | 1 | 2004 | FPN: A Distributed Hash Table for Commercial Applications · HPDC 2004 |
Storage systems › storage management › backup and recovery
data backup |
0.0 | 1 | 2010 | Bimodal Content Defined Chunking for Backup Streams · FAST 2010 |
Methods — techniques the papers use, named apart from their topics
data parallelism · 0.4bloom filter · 0.2LRU approximation · 0.2storage characterization · 0.1pilot solution design · 0.1measurement study · 0.1subspace decomposition · 0.1squared prediction error · 0.1sequentially discounting expectation maximization · 0.1hotelling t2 · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | MALT: distributed data-parallelism for existing ML applicationsabstractMachine learning methods, such as SVM and neural networks, often improve their accuracy by using models with more parameters trained on large numbers of examples. Building such models on a single machine is often impractical because of the large amount of computation required. Hao Li 0022, Asim Kadav, Erik Kruus, Cristian Ungureanu |
EuroSys | 4 |
| 2015 | Simba: tunable end-to-end data consistency for mobile appsabstractDevelopers of cloud-connected mobile apps need to ensure the consistency of application and user data across multiple devices. Mobile apps demand different choices of distributed data consistency under a variety of usage scenarios. The apps also need to gracefully handle intermittent connectivity and disconnections, limited bandwidth, and client and server failures. The data model of the apps can also be complex, spanning inter-dependent structured and unstructured data, and needs to be atomically stored and updated locally, on the cloud, and on other mobile devices. Dorian Jean Perkins, Nitin Agrawal 0001, Akshat Aranya, Curtis Yu, Younghwan Go, Harsha V. Madhyastha, Cristian Ungureanu |
EuroSys | 7 |
| 2015 | Reliable, Consistent, and Efficient Data Sync for Mobile Apps
Younghwan Go, Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu |
FAST | 4 |
| 2013 | Mobile Data Sync in a Blink
Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu |
HotStorage | 3 |
| 2013 | TBF: A memory-efficient replacement policy for flash-based cachesabstractThe performance and capacity characteristics of flash storage make it attractive to use as a cache. Recency-based cache replacement policies rely on an in-memory full index, typically a B-tree or a hash table, that maps each object to its recency information. Even though the recency information itself may take very little space, the full index for a cache holding N keys requires at least log N bits per key. This metadata overhead is undesirably high when used for very large flash-based caches, such as key-value stores with billions of objects. To solve this problem, we propose a new RAM-frugal cache replacement policy that approximates the least-recently-used (LRU) policy. It uses two in-memory Bloom sub-filters (TBF) for maintaining the recency information and leverages an on-flash key-value store to cache objects. TBF requires only one byte of RAM per cached object, making it suitable for implementing very large flash-based caches. We evaluate TBF through simulation on traces from several block stores and key-value stores, as well as evaluate it using the Yahoo! Cloud Serving Benchmark in a real system implementation. Evaluation results show that TBF achieves cache hit rate and operations per second comparable to those of LRU in spite of its much smaller memory requirements. Cristian Ungureanu, Biplob Debnath, Stephen Rago, Akshat Aranya |
ICDE | 1 |
| 2013 | Building a Delay-Tolerant Cloud for Mobile DataabstractMobile data usage is on a tremendous rise, due not only to increasing number of users but also to an increase in the number of applications that transfer data over the network. Moreover, applications for sharing, sensing, and collaboration have become more popular, causing significant amounts of data to be generated on devices. Managing this data -syncing it to the cloud, or with other users or devices- is a crucial and often challenging part of writing mobile apps and services. In spite of plenty of good advice and best practices from OS vendors and network operators, storing and transferring mobile data is fraught with issues. On the one hand, an app developer needs to worry about the semantics of data storage and synchronization, while on the other, about the end-user experience, which maybe impacted by poor and intermittent network connectivity. To address the needs of the app developers and the end-users, we have built Izzy: a platform to rapidly develop and deploy data-centric mobile apps. Izzy provides well-defined and easy to use semantics for accessing local storage and for synchronizing data with a remote, scalable, global store. Izzy also provides global store access to the cloud-resident part of the applications (if any) through a similar server API. Last but not least, Izzy is designed to be frugal: it conserves mobile device resources by applying delay-tolerance and data reduction techniques (message coalescing and compression) across applications on a mobile device. In this paper we present the design of Izzy and our early experiences with using it. Nitin Agrawal 0001, Akshat Aranya, Cristian Ungureanu |
MDM (1) | 4 |
| 2012 | Revisiting storage for smartphones
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu |
FAST | 3 |
| 2012 | Revisiting storage for smartphonesabstractConventional wisdom holds that storage is not a big contributor to application performance on mobile devices. Flash storage (the type most commonly used today) draws little power, and its performance is thought to exceed that of the network subsystem. In this article, we present evidence that storage performance does indeed affect the performance of several common applications such as Web browsing, maps, application install, email, and Facebook. For several Android smartphones, we find that just by varying the underlying flash storage, performance over WiFi can typically vary between 100% and 300% across applications; in one extreme scenario, the variation jumped to over 2000%. With a faster network (set up over USB), the performance variation rose even further. We identify the reasons for the strong correlation between storage and application performance to be a combination of poor flash device performance, random I/O from application databases, and heavy-handed use of synchronous writes. Based on our findings, we implement and evaluate a set of pilot solutions to address the storage performance deficiencies in smartphones. Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu |
ACM Trans. Storage | 3 |
| 2011 | Using Eager Strategies to Improve NFS I/O PerformanceabstractTypical NFS clients write in a lazy fashion: they leave dirty pages in the page cache and defer writing to the server until later. This reduces network traffic when applications repeatedly modify the same set of pages. However, this approach can lead to memory pressure, when the number of available pages on the client system is so low that the system must work harder to reclaim dirty pages. System performance is poor under memory pressure. We show examples of this problem and present two mechanisms to solve it: eager write back and eager page laundering. These mechanisms change the client's data management policy from lazy to eager, resulting in higher throughput for sequential writes. In addition, we show that NFS servers suffer from out-of-order file operations, which further reduce performance. We introduce request ordering, a server mechanism to process operations (as much as possible) in the order they were sent by the client, which improves read performance substantially. We have implemented these techniques in the Linux operating system. I/O performance is improved, with the most pronounced improvement visible for sequential access to large files. We see about 33% improvement in the performance of streaming write workloads and more than triple the performance of streaming read workloads. We evaluate several nonsequential workloads and show that these techniques do not degrade performance, and can sometimes improve performance. We also design and evaluate an adversarial workload to show that the eager policies can perform worse in some pathological cases. Stephen Rago, Aniruddha Bohra, Cristian Ungureanu |
NAS | 3 |
| 2010 | Bimodal Content Defined Chunking for Backup Streams
Erik Kruus, Cristian Ungureanu, Cezary Dubnicki |
FAST | 2 |
| 2010 | HydraFS: A High-Throughput File System for the HYDRAstor Content-Addressable Storage System
Cristian Ungureanu, Benjamin Atkin, Akshat Aranya, Salil Gokhale, Stephen Rago, Grzegorz Calkowski, Cezary Dubnicki, Aniruddha Bohra |
FAST | 1 |
| 2010 | KVZone and the Search for a Write-Optimized Key-Value Store
Salil Gokhale, Nitin Agrawal 0001, Sean Noonan, Cristian Ungureanu |
HotStorage | 4 |
| 2009 | HYDRAstor: A Scalable Secondary Storage
Cezary Dubnicki, Leszek Gryz, Lukasz Heldt, Michal Kaczmarczyk, Wojciech Kilian, Przemyslaw Strzelczak, Jerzy Szczepkowski, Cristian Ungureanu, Michal Welnicki |
FAST | 8 |
| 2007 | Online Tracking of Component Interactions for Failure Detection and Localization in Distributed SystemsabstractThis paper proposes a novel failure-detection approach that can handle high-dimensional observation and frequent system changes. We extract two statistics from the subspace decomposition of observations, and use the mixture of Gaussians to model their probability density. Instead of monitoring the original data, the density model of extracted statistics is adaptively updated and examined regularly to detect failures. We also present a localization method to identify the faulty components once the failure happens. Applying our technique to monitor the component interactions in an e-commerce application shows satisfactory results in detecting a variety of injected failures. Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2007 | Multiresolution Abnormal Trace Detection Using Varied-Length n-Grams and AutomataabstractDetection and diagnosis of faults in a large-scale distributed system is a formidable task. Interest in monitoring and using traces of user requests for fault detection has been on the rise recently. In this paper we propose novel fault detection methods based on abnormal trace detection. One essential problem is how to represent the large amount of training trace data compactly as an oracle. Our key contribution is the novel use of varied-length n-grams and automata to characterize normal traces. A new trace is compared against the learned automata to determine whether it is abnormal. We develop algorithms to automatically extract n-grams and construct multiresolution automata from training data. Further, both deterministic and multihypothesis algorithms are proposed for detection. We inspect the trace constraints of real application software and verify the existence of long n-grams. Our approach is tested in a real system with injected faults and achieves good results in experiments Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2005 | Failure detection and localization in component based systems by online trackingabstractThe increasing complexity of today's systems makes fast and accurate failure detection essential for their use in mission-critical applications. Various monitoring methods provide a large amount of data about system's behavior. Analyzing this data with advanced statistical methods holds the promise of not only detecting the errors faster, but also detecting errors which are difficult to catch with current monitoring tools. Two challenges to building such detection tools are: the high dimensionality of observation data, which makes the models expensive to apply, and frequent system changes, which make the models expensive to update. In this paper, we present algorithms to reduce the dimensionality of data in a way that makes it easy to adapt to system changes. We decompose the observation data into signal and noise subspaces. Two statistics, the Hotelling T2 score and squared prediction error (SPE) are calculated to represent the data characteristics in signal and noise subspaces respectively. Instead of tracking the original data, we use a sequentially discounting expectation maximization (SDEM) algorithm to learn the distribution of the two extracted statistics. A failure event can then be detected based on the abnormal change of the distribution. Applying our technique to component interaction data in a simple e-commerce application shows better accuracy than building independent profiles for each component. Additionally, experiments on synthetic data show that the detection accuracy is high even for changing systems. Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira |
KDD | 3 |
| 2004 | FPN: A Distributed Hash Table for Commercial Applications
Cezary Dubnicki, Cristian Ungureanu, Wojciech Kilian |
HPDC | 2 |
| 2000 | Concurrency Analysis for Java
Cristian Ungureanu, Suresh Jagannathan |
SAS | 1 |
| 1997 | Formal Models of Distributed Memory ManagementabstractWe develop am abstract model of memory management in distributed systems. The model is low-level enough so that we can express communication, allocation and garbage collection, but otherwise hides many of the lower-level details of an actual implementation.Recently, such formal models have been developed for memory management in a functional, sequential setting [10]. The models are rewriting systems whose terms are programs. Programs have both the "code" (control string) and the "store" syntactically apparent. Evaluation is expressed as conditional rewriting and includes store operations. Garbage collection becomes a rewriting relation that removes part of the store without affecting the behavior of the program.Distribution adds another dimension to an already complex problem. By using techniques developed for communicating and concurrent systems [9], we extend their work to a distributed environment. Sending and receiving messages is also made apparent at the syntactic level. A very general garbage collection rule based on reachability is introduced and proved correct. Now proving correct a specific collection strategy is reduced to showing that the relation between programs defined by the strategy is a sub-relation of the general relation. Any actual implementation which is capable of providing the transitions (including their atomicity constraints) specified by the strategy is therefore correct.This model allows us to specify and prove correct in a compact manner two garbage collectors; the first one does a simple garbage collection local to a node. The second garbage collector uses migration of data in order to be able to reclaim inter-node cyclic garbage. Cristian Ungureanu, Benjamin Goldberg 0001 |
ICFP | 1 |
| 1997 | Run-Time versus Compile-Time Instruction Scheduling in Superscalar (RISC) Processors: Performance and Trade-OffabstractThe RISC revolution has spurred the development of processors with increasing degrees ofinstruction level parallelism(ILP). In order to realize the full potential of these processors, multiple instructions must continuously be issued and executed in a single cycle. Consequently,instruction schedulingplays a crucial role as an optimization in this context. While early attempts at instruction scheduling were limited to compile-time approaches, the current trends are aimed at providingdynamicsupport in hardware. In this paper, we present the results of a detailed comparative study of the performance advantages to be derived by the spectrum of instruction scheduling approaches: from limited basic-block schedulers in the compiler, to novel and aggressive schedulers in hardware. A significant portion of our experimental study via simulations, is devoted to understanding the performance advantages of run-time scheduling. Our results indicate it to be effective in extracting the ILP inherent to the program trace being scheduled, over a wide range of machine and program parameters. Furthermore, we also show that this effectiveness can be further enhanced by a simple basic-block scheduler in the compiler, which optimizes for the presence of the run-time scheduler in the target; current basic-block schedulers are not designed to take advantage of this feature. We demonstrate this fact by presenting a novel basic-block scheduling algorithm that is sensitive to the lookahead hardware in the target processor. Finally, we outline a simple analytical characterization of the performance advantage that run-time schedulers have to offer. Allen Leung, Krishna V. Palem, Cristian Ungureanu |
J. Parallel Distributed Comput. | 3 |
| 1996 | Run-time versus compile-time instruction scheduling in superscalar (RISC) processors: performance and tradeoffsabstractThe RISC revolution has spurred the development of processors with increasing degrees of instruction level parallelism (ILP). In order to realize the full potential of these processors, multiple instructions must continuously be issued and executed in a single cycle. Consequently, instruction scheduling plays a crucial role as an optimization in this context. While early attempts at instruction scheduling were limited to compile-time approaches, the current trends are aimed at providing dynamic support in hardware. In this paper, we present the results of a detailed comparative study of the performance advantages to be derived by the spectrum of instruction scheduling approaches: from limited basic-block schedulers in the compiler, to novel and aggressive schedulers in hardware. A significant portion of our experimental study via simulations, is devoted to understanding the performance advantages of run-time scheduling. Our results indicate it to be effective in extracting the ILP inherent to the program trace being scheduled, over a wide range of machine and program parameters. Furthermore, we also show that this effectiveness can be further enhanced by a simple basic-block scheduler in the compiler, which optimizes for the presence of the run-time scheduler in the target; current basic-block schedulers are not designed to take advantage of this feature. We demonstrate this fact by presenting a novel basic-block scheduling algorithm that is sensitive to the lookahead hardware in the target processor. Allen Leung, Krishna V. Palem, Cristian Ungureanu |
HiPC | 3 |