Haralampos Pozidis

dblp:75/6947 · also Haralambos Pozidis, Haris Pozidis · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
8since 2021 · last 2024
0000-0001-5084-6651ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 5 since 2021Computer networks · 11 · 4 first-authorArtificial intelligence and machine learning · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2024 WannaLaugh: A Configurable Ransomware Emulator - Learning to Mimic Malicious Storage Traces
abstract
Ransomware, a fearsome and an evolving cybersecurity threat, continues to inflict severe consequences on individuals and organizations worldwide. Traditional detection methods, reliant on static signatures and application behavioral patterns, are challenged by the dynamic nature of these threats. This paper introduces two primary contributions to address this challenge. First, we introduce the WannaLaugh ransomware emulator. This tool is designed to safely mimic ransomware attacks without causing actual harm or spreading malware, making it a unique solution for studying ransomware behavior. Second, we show how this emulator can be used to mimic the I/O behavior of existing ransomware. Experimental results show that WannaLaugh can mimic six real ransomware with high accuracy. Both the emulator and its mimicking application aim to represent significant steps forward in ransomware detection in the era of machine-learning-driven cybersecurity.
Dionysios Diamantopoulos, Roman A. Pletka, Slavisa Sarafijanovic, A. L. Narasimha Reddy, Haralampos Pozidis
SYSTOR5
2023 xCloudServing: Automated ML Serving Across Clouds
abstract
As machine learning (ML) models have grown in complexity, so too have the expenses they incur when deployed in the cloud. In order to reduce the costs associated with ML serving, it is necessary to optimize the choice of cloud infrastructure used. Additionally, the chosen infrastructure must be able to deliver on the latency constraints that are typically defined for cloud services. This problem is made more challenging since today's organizations often need to work with more than one cloud provider, and each provider offers its own unique set of interfaces and infrastructure choices. In this work we present xCloudServing - a novel system for consistent and automated deployment of ML inference services across multi-ple cloud providers and regions. We describe the architecture and implementation of xCloudServing, as well as the different optimization algorithms implemented internally. These include established methods from the literature, as well as Niebo - our novel algorithm for minimizing cost whilst satisfying the tail latency constraint. We present simulation results for 5 different ML models over 3 cloud providers and multiple tail latency constraints that indicate that on average, Niebo outperforms state-of-the-art algorithms by 37%. Additionally, we evaluate xCloudServing with live runs and demonstrate that it is robust to nondeterministic effects and exhibits reproducible behavior.
Malgorzata Lazuka, Andreea Anghel, Parikshit Ram, Haralampos Pozidis, Thomas P. Parnell
CLOUD4
2023 Multivariate Anomaly Detection with Domain Clustering
abstract
Existing time-series anomaly detection (AD) pipelines for cloud monitoring at scale commonly rely on isolated training per cloud service or cloud infrastructure component. However, with the increasing volume of data generated from thousands of services and components, there is an untapped opportunity for a more effective approach to detect key performance indicator (KPI) anomalies by capitalizing on the abundance of data available. In this paper, we propose MADDoC, an unsupervised transfer learning framework for reconstruction based anomaly detection on multivariate time-series data. We show how to efficiently leverage available KPIs in the realm of cloud infrastructure monitoring to generalize unsupervised time-series AD across infrastructure components. Compared to state-of-the-art approaches relying on isolated component-wise training, the MADDoC framework achieves superior Precision and F1 scores on public and internal time-series AD datasets, by learning a strong reconstruction backbone on the time-series data across many components, before fine-tuning to a specific component. Moreover, MADDoC achieves substantial cost savings in model training, with reductions of 60% to 75% when monitoring thousands of storage infrastructure components. Further, the framework overcomes the trade-off between training efficiency and AD performance of previous AD transfer learning approaches.
Frederic Boesel, Livio Schläpfer, Haralampos Pozidis, Mitchell Gusat
SoCC3
2023 Acceleration of Decision-Tree Ensemble Models on the IBM Telum Processor
abstract
This paper presents a tensor-based algorithm that leverages a hardware accelerator for inferencing decision-tree-based machine learning models. The algorithm has been integrated in a public software library and is demonstrated on an IBM z16 server, using the Telum processor with the Integrated Accelerator for AI. We describe the architecture and implementation of the algorithm and present experimental results that demonstrate its superior runtime performance compared with popular CPU-based machine learning inference implementations.
Nikolaos Papandreou, Jan van Lunteren, Andreea Anghel, Thomas P. Parnell, Martin Petermann, Milos Stanisavljevic, Cédric Lichtenau, Andrew Sica, Dominic Röhm, Elpida Tzortzatos, Haralampos Pozidis
ISCAS11
2022 Search-based Methods for Multi-Cloud Configuration
abstract
Multi-cloud computing has become increasingly popular with enterprises looking to avoid vendor lock-in. While most cloud providers offer similar functionality, they may differ significantly in terms of performance and/or cost. A customer looking to benefit from such differences will naturally want to solve the multi-cloud configuration problem: given a workload, which cloud provider should be chosen and how should its nodes be configured in order to minimize runtime or cost? In this work, we consider possible solutions to this multi-cloud optimization problem. We develop and evaluate possible adaptations of state-of-the-art cloud configuration solutions to the multi-cloud domain. Furthermore, we identify an analogy between multi-cloud configuration and the selection-configuration problems that are commonly studied in the automated machine learning (AutoML) field. Inspired by this connection, we utilize popular optimizers from AutoML to solve multi-cloud configuration. Finally, we propose a new algorithm for solving multi-cloud configuration, CloudBandit. It treats the outer problem of cloud provider selection as a best-arm identification problem, in which each arm pull corresponds to running an arbitrary black-box optimizer on the inner problem of node configuration. Our extensive experiments indicate that (a) many state-of-the-art cloud configuration solutions can be adapted to multi-cloud, with best results obtained for adaptations which utilize the hierarchical structure of the multi-cloud configuration domain, (b) hierarchical methods from AutoML can be used for the multi-cloud configuration task and can outperform state-of-the-art cloud configuration solutions and (c) CloudBandit achieves competitive or lower regret relative to other tested algorithms, whilst also identifying configurations that have 65% lower median cost and 20% lower median runtime in production, compared to choosing a random provider and configuration.
Malgorzata Lazuka, Thomas P. Parnell, Andreea Anghel, Haralampos Pozidis
CLOUD4
2022 CCA: An ML Pipeline for Cloud Anomaly Troubleshooting
abstract
Cloud Causality Analyzer (CCA) is an ML-based analytical pipeline to automate the tedious process of Root Cause Analysis (RCA) of Cloud IT events. The 3-stage pipeline is composed of 9 functional modules, including dimensionality reduction (feature engineering, selection and compression), embedded anomaly detection, and an ensemble of 3 custom explainability and causality models for Cloud Key Performance Indicators (KPI). Our challenge is: How to apply a reduced (sub)set of judiciously selected KPIs to detect Cloud performance anomalies, and their respective root causal culprits, all without compromising accuracy?
Lili Georgieva, Ioana Giurgiu, Serge Monney, Haralampos Pozidis, Viviane Potocnik, Mitchell Gusat
AAAI4
2022 AI accelerator on IBM telum processor: industrial product
abstract
IBM Telum is the next generation processor chip for IBM Z and LinuxONE systems. The Telum design is focused on enterprise class workloads and it achieves over 40% per socket performance growth compared to IBM z15. The IBM Telum is the first server-class chip with a dedicated on-chip AI accelerator that enables clients to gain real time insights from their data as it is getting processed.
Cédric Lichtenau, Alper Buyuktosunoglu, Ramon Bertran Monfort, Peter Figuli, Christian Jacobi 0002, Nikolaos Papandreou, Haralampos Pozidis, Anthony Saporito, Andrew Sica, Elpida Tzortzatos
ISCA7
2021 High-Throughput ECC with Integrated Chipkill Protection for Nonvolatile Memory Arrays
abstract
New coding schemes based on generalized concatenated codes are proposed for emerging nonvolative memory technologies. Key requirements are high code rate, low latency, high throughput and the ability to correct chipkill failures while sustaining high data reliability despite raw bit error rates up to 10-3. New concatenated codes based on Reed-Solomon codes have been designed for payload sizes of 512B, 1kB, and 2kB; they have high rates above 0.8 and a high data reliability with a decoder target BER of 10-15. An FPGA-based implementation of the decoder validates the low latency and high throughput: for an operating clock frequency of 250MHz, the decoding latency is 236ns and a 9.3GB/s throughput is achieved.
Thomas Mittelholzer, Milos Stanisavljevic, Nikolaos Papandreou, Haralampos Pozidis
ISCAS4
2020 Weighted Sampling for Combined Model Selection and Hyperparameter Tuning
Dimitrios Sarigiannis, Thomas P. Parnell, Haralampos Pozidis
AAAI3
2020 An Anomaly Detection and Explainability Framework using Convolutional Autoencoders for Data Storage Systems
abstract
Anomaly detection in data storage systems is a challenging problem due to the high dimensional sequential data involved, and lack of labels. The state of the art for automating anomaly detection in these systems typically relies on hand crafted rules and thresholds which mainly allow to distinguish between normal and abnormal behavior of each indicator in isolation. In this work we present an end-to-end framework based on convolutional autoencoders which not only allows for anomaly detection on multivariate time series data, but also provides explainability. This is done by identifying similar historic anomalies and extracting the most influential indicators. These are then presented to relevant personnel such as system designers and architects, or to support engineers for further analysis. We demonstrate the application of this framework along with an intuitive interactive web interface which was developed for data storage system anomaly detection. We discuss how this framework along with its explainability aspects enables support engineers to effectively tackle abnormal behaviors, all while allowing for crucial feedback.
Roy Assaf, Ioana Giurgiu, Jonas Pfefferle, Serge Monney, Haralampos Pozidis, Anika Schumann
IJCAI5
2020 Improving NAND flash performance with read heat separation
abstract
The continuous growth in 3D-NAND flash storage density has primarily been enabled by 3D stacking and by increasing the number of bits stored per memory cell. Unfortunately, these desirable flash device design choices are adversely affecting reliability and latency characteristics. In particular, increasing the number of bits stored per cell results in having to apply additional voltage thresholds during each read operation, therefore increasing the read latency characteristics. While most NAND flash challenges can be mitigated through appropriate background processing, the flash read latency characteristics cannot be hidden and remains the biggest challenge, especially for the newest flash generations that store four bits per cell. In this paper, we introduce read heat separation (RHS), a new heat-aware data-placement technique that exploits the skew present in real-world workloads to place frequently read user data on low-latency flash pages. Although conceptually simple, such a technique is difficult to integrate in a flash controller, as it introduces a significant amount of complexity, requires more metadata, and is further constrained by other flash-specific peculiarities. To overcome these challenges, we propose a novel flash controller architecture supporting read heat-aware data placement. We first discuss the trade-offs that such a new design entails and analyze the key aspects that influence the efficiency of RHS. Through both, extensive simulations and an implementation we realized in a commercial enterprise-grade solid-state drive controller, we show that our architecture can indeed significantly reduce the average read latency. For certain workloads, it can reverse the system-level read latency trends when using recent multi-bit flash generations and hence outperform SSDs using previous faster flash generations.
Roman A. Pletka, Nikolaos Papandreou, Radu Stoica, Haralampos Pozidis, Nikolas Ioannou, Timothy Fisher, Aaron Fry, Kip Ingram, Andrew Walls
MASCOTS4
2020 SnapBoost: A Heterogeneous Boosting Machine
abstract
Modern gradient boosting software frameworks, such as XGBoost and LightGBM, implement Newton descent in a functional space. At each boosting iteration, their goal is to find the base hypothesis, selected from some base hypothesis class, that is closest to the Newton descent direction in a Euclidean sense. Typically, the base hypothesis class is fixed to be all binary decision trees up to a given depth. In this work, we study a Heterogeneous Newton Boosting Machine (HNBM) in which the base hypothesis class may vary across boosting iterations. Specifically, at each boosting iteration, the base hypothesis class is chosen, from a fixed set of subclasses, by sampling from a probability distribution. We derive a global linear convergence rate for the HNBM under certain assumptions, and show that it agrees with existing rates for Newton's method when the Newton direction can be perfectly fitted by the base hypothesis at each boosting iteration. We then describe a particular realization of a HNBM, SnapBoost, that, at each boosting iteration, randomly selects between either a decision tree of variable depth or a linear regressor with random Fourier features. We describe how SnapBoost is implemented, with a focus on the training complexity. Finally, we present experimental results, using OpenML and Kaggle datasets, that show that SnapBoost is able to achieve better generalization loss than competing boosting frameworks, without taking significantly longer to tune.
Thomas P. Parnell, Andreea Anghel, Malgorzata Lazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, Haralampos Pozidis
NeurIPS8
2020 Tera-scale coordinate descent on GPUs
Thomas P. Parnell, Celestine Dünner, Kubilay Atasu, Manolis Sifalakis, Haralampos Pozidis
Future Gener. Comput. Syst.5
2019 Understanding the Design Trade-Offs of Hybrid Flash Controllers
abstract
Over the last few years, NAND flash manufacturers have steadily increased the number of bits stored per cell to achieve significant cost reductions. However, the increased density does not come without drawbacks. All key flash performance metrics, including latency and endurance, significantly degrade as bit density increases. Particularly, sustained write throughput is the worst affected as writes are roughly one order of magnitude slower than reads and further require precursory block erases in the background. As a result, many recent flash controllers operate flash blocks both in single-bit (high endurance and performance) and in multi-bit (high density) mode. In theory, such hybrid controllers are a great way of hiding flash technology limitations. A controller can use a small percentage of the flash blocks in single-bit mode as a cache which allows orders of magnitude higher write bandwidth and endurance in environments where the access patterns of the workload are skewed and bursty. In practice, however, many devices fall short of expectations when write performance varies significantly and utilization increases. We argue that a principled approach is required to understand the design trade-offs of hybrid NAND flash controllers. To this end, we develop a modeling framework for estimating the performance and endurance of hybrid controllers. The modeling framework computes the internal data movement generated by a hybrid controller by relying on advanced analytical models that offer both accurate and fast predictions. The data flow is then translated into higher-level metrics that quantify upper bounds for the overall performance of an SSD such as write throughput, latency, and device endurance. Using our modeling framework, we compare different controller architectures, identify their strong and weak points, and show that there is room to improve the efficiency of the hybrid controllers used today.
Radu Stoica, Roman A. Pletka, Nikolas Ioannou, Nikolaos Papandreou, Sasa Tomic, Haralampos Pozidis
MASCOTS6
2019 Accelerated ML-Assisted Tumor Detection in High-Resolution Histopathology Images
Nikolas Ioannou, Milos Stanisavljevic, Andreea Anghel, Nikolaos Papandreou, Sonali Andani, Jan Hendrik Rüschoff, Peter Wild, Maria Gabrani, Haralampos Pozidis
MICCAI (1)9
2019 Snap ML - Accelerated Machine Learning for Big Data (Keynote Abstract)
abstract
Snap Machine Learning (Snap ML) is a new software library for training popular machine learning models, characterized by very high performance, scalability to TB-scale datasets and high resource efficiency. It continuously evolves and currently supports generalized linear models, decision trees, random forests and gradient boosting machines. Snap ML has been built to address the needs of business applications, which often have to deal with high-volume data, react fast to changing environments, and use resources efficiently to drive down cost. The high efficiency of Snap ML, in particular in dealing with big data, comes from innovations in distributed optimization, among other things. This talk will review the principles of the Snap ML library, explain how it achieves high speed and scalability, and present several cases of business workloads that demonstrate the benefits offered by Snap ML. Haris Pozidis manages the Cloud Storage and Analytics group at IBM Research in Zurich, Switzerland. He was with Philips Research, Eindhoven, The Netherlands, before joining IBM. He has worked on read channel design for DVD and Blu-ray Disc at Philips, and played a key role in developing the first scanning probe-based data storage system at IBM, the “Millipede”. His current focus is on the development of Flash memory controllers for all-flash arrays, on phase change memory technology and system solutions, and on accelerated software libraries for machine learning. He holds over 120 US patents, has co-authored more than 120 publications, is an IBM Principal Research Scientist and Master Inventor, and a Senior Member of the IEEE.
Haralampos Pozidis
OPODIS1
2018 Exploiting the non-linear current-voltage characteristics for resistive memory readout
abstract
Various resistive memory technologies are finding application in the space of storage-class memory and emerging non-von Neumann computing systems. For both applications, a key enabling technology is the ability to store multiple resistance levels in a single memory cell. The resistance states of these devices are typically measured in the low-field regime, where the electrical transport can be assumed to be Ohmic. However, when biased at slightly higher voltages, they exhibit significantly nonlinear I-V characteristics. In this paper, we demonstrate how this field dependence of the resistance values can be exploited in various applications. We present simulation and experimental results where readout schemes based on the non-linear I-V behavior are used to enhance the readout margin and also to compensate for resistance drift.
Nikolaos Papandreou, Abu Sebastian, Haralampos Pozidis
ISCAS3
2018 Drift-Invariant Detection for Multilevel Phase-Change Memory
abstract
Next-generation memory (NGM) technologies present a major opportunity but also a significant challenge, due to their intricate reliability issues. In particular, multilevel-cell (MLC) storage is highly desirable for increasing storage capacity and lowering total cost-per-bit. In phase-change memory (PCM), MLC storage is hampered by sensitivity to temperature variations and resistance drift. A novel drift-invariant detection (DID) scheme that estimates variable read thresholds based on ordered statistics and clustering of the soft read-back signals from a small block of 32 cells has been developed and implemented in hardware to improve reliability and prolong data retention. A low-complexity implementation of the DID on a FPGA platform comprises 20'000 LUTs and 6'000 flip-flops and has a latency of 90ns. We present results from an extensive performance verification that ascertains highly reliable data retrieval up to 13 orders of magnitude in time after programming. Such elevated reliability is necessary for the most anticipated application of NGM, namely persistent far-memory, where the NGM is used as a large memory pool, possibly together with DRAM.
Milos Stanisavljevic, Thomas Mittelholzer, Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis
ISCAS5
2018 Snap ML: A Hierarchical Framework for Machine Learning
abstract
We describe a new software framework for fast training of generalized linear models. The framework, named Snap Machine Learning (Snap ML), combines recent advances in machine learning systems and algorithms in a nested manner to reflect the hierarchical architecture of modern computing systems. We prove theoretically that such a hierarchical system can accelerate training in distributed environments where intra-node communication is cheaper than inter-node communication. Additionally, we provide a review of the implementation of Snap ML in terms of GPU acceleration, pipelining, communication patterns and software architecture, highlighting aspects that were critical for achieving high performance. We evaluate the performance of Snap ML in both single-node and multi-node environments, quantifying the benefit of the hierarchical scheme and the data streaming functionality, and comparing with other widely-used machine learning software frameworks. Finally, we present a logistic regression benchmark on the Criteo Terabyte Click Logs dataset and show that Snap ML achieves the same test loss an order of magnitude faster than any of the previously reported results, including those obtained using TensorFlow and scikit-learn.
Celestine Dünner, Thomas P. Parnell, Dimitrios Sarigiannis, Nikolas Ioannou, Andreea Anghel, Gummadi Ravi, Madhusudanan Kandasamy, Haralampos Pozidis
NeurIPS8
2018 Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets
abstract
Despite its widespread use in consumer devices and enterprise storage systems, NAND flash faces a growing number of challenges. While technology advances have helped to increase the storage density and reduce costs, they have also led to reduced endurance and larger block variations, which cannot be compensated solely by stronger ECC or read-retry schemes but have to be addressed holistically. Our goal is to enable low-cost NAND flash in enterprise storage for cost efficiency. We present novel flash-management approaches that reduce write amplification, achieve better wear leveling, and enhance endurance without sacrificing performance. We introduce block calibration, a technique to determine optimal read-threshold voltage levels that minimize error rates, and novel garbage-collection as well as data-placement schemes that alleviate the effects of block health variability and show how these techniques complement one another and thereby achieve enterprise storage requirements. By combining the proposed schemes, we improve endurance by up to 15× compared to the baseline endurance of NAND flash without using a stronger ECC scheme. The flash-management algorithms presented herein were designed and implemented in simulators, hardware test platforms, and eventually in the flash controllers of production enterprise all-flash arrays. Their effectiveness has been validated across thousands of customer deployments since 2015.
Roman A. Pletka, Ioannis Koltsidas, Nikolas Ioannou, Sasa Tomic, Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Aaron Fry, Timothy Fisher
ACM Trans. Storage7
2017 Linear-complexity relaxed word Mover's distance with GPU acceleration
abstract
The amount of unstructured text-based data is growing every day. Querying, clustering, and classifying this big data requires similarity computations across large sets of documents. Whereas low-complexity similarity metrics are available, attention has been shifting towards more complex methods that achieve a higher accuracy. In particular, the Word Mover's Distance (WMD) method proposed by Kusner et al. is a promising new approach, but its time complexity grows cubically with the number of unique words in the documents. The Relaxed Word Mover's Distance (RWMD) method, again proposed by Kusner et al., reduces the time complexity from qubic to quadratic and results in a limited loss in accuracy compared with WMD. Our work contributes a low-complexity implementation of the RWMD that reduces the average time complexity to linear when operating on large sets of documents. Our linear-complexity RWMD implementation, henceforth referred to as LC-RWMD, maps well onto GPUs and can be efficiently distributed across a cluster of GPUs. Our experiments on real-life datasets demonstrate 1) a performance improvement of two orders of magnitude with respect to our GPU-based distributed implementation of the quadratic RWMD, and 2) a performance improvement of three to four orders of magnitude with respect to our distributed WMD implementation that uses GPU-based RWMD for pruning.
Kubilay Atasu, Thomas P. Parnell, Celestine Dünner, Manolis Sifalakis, Haralampos Pozidis, Vasileios Vasileiadis, Michail Vlachos, Cesar Berrospi, Abdel Labbi
IEEE BigData5
2017 Understanding and optimizing the performance of distributed machine learning applications on apache spark
abstract
In this paper we explore the performance limits of Apache Spark for machine learning applications. We begin by analyzing the characteristics of a state-of-the-art distributed machine learning algorithm implemented in Spark and compare it to an equivalent reference implementation using the high performance computing framework MPI. We identify critical bottlenecks of the Spark framework and carefully study their implications on the performance of the algorithm. In order to improve Spark performance we then propose a number of practical techniques to alleviate some of its overheads. However, optimizing computational efficiency and framework related overheads is not the only key to performance — we demonstrate that in order to get the best performance out of any implementation it is necessary to carefully tune the algorithm to the respective trade-off between computation time and communication latency. The optimal trade-off depends on both the properties of the distributed algorithm as well as infrastructure and framework-related characteristics. Finally, we apply these technical and algorithmic optimizations to three different distributed linear machine learning algorithms that have been implemented in Spark. We present results using five large datasets and demonstrate that by using the proposed optimizations, we can achieve a reduction in the performance difference between Spark and MPI from 20x to 2x.
Celestine Dünner, Thomas P. Parnell, Kubilay Atasu, Manolis Sifalakis, Haralampos Pozidis
IEEE BigData5
2017 High-Performance Recommender System Training Using Co-Clustering on CPU/GPU Clusters
abstract
Recommender systems are becoming the crystal ball of the Internet because they can anticipate what the users may want, even before the users know they want it. However, the machine-learning algorithms typically involved in the training of such systems can be computationally expensive, and often may require several days for retraining. Here, we present a distributed approach for load-balancing the training of a recommender system based on state-of-art non-negative matrix factorization principles. The approach can exploit the presence of a cluster of mixed CPUs and GPUs, and results in a 466-fold performance improvement compared with the serial CPU implementation, and a 15-fold performance improvement compared with the best previously reported results for the popular Netflix data set.
Kubilay Atasu, Thomas P. Parnell, Celestine Dünner, Michail Vlachos, Haralampos Pozidis
ICPP5
2016 Controller architecture for low-latency access to phase-change memory in OpenPOWER systems
abstract
Novel forms of nonvolatile memory, such as phase-change memory (PCM), promise low latency and small granularity of read and write access at high storage density. They also feature very high endurance. These characteristics make them highly desirable for emerging high-capacity (hybrid) memory applications such as in-memory databases and in-memory processing. In this work, we present the architecture, implementation and experimental performance results of an FPGA-based PCM memory controller for OpenPOWER servers. The memory controller leverages the Coherent Accelerator Processor Interface (CAPI) of the POWER processor in order to offer low-latency access to the CPU memory space. In addition, the memory controller implements an efficient management protocol that supports a dynamic size of pending read and write requests in order to offer high bandwidth under mixed-type workloads. We describe the architecture and implementation details of the memory controller and we demonstrate its performance using a prototype platform based on different types of OpenPOWER servers equipped with CAPI-enabled FPGA cards. The developed PCM controller is evaluated in terms of sustained data rates (MBps) and access latency (us). Experimental results are based on legacy commercial 90nm PCM chips as well as on accurate HW emulation of next generation PCM chips.
Antonios Prodromakis, Nikolaos Papandreou, Eleni Bougioukou, Urs Egger, Nikos Toulgaridis, Theodore Antonakopoulos 0001, Haralampos Pozidis, Evangelos Eleftheriou
FPL7
2016 Improving the error-floor performance of binary half-product codes
Thomas Mittelholzer, Thomas P. Parnell, Nikolaos Papandreou, Haralampos Pozidis
ISITA4
2016 Guest Editorial Channel Modeling, Coding and Signal Processing for Novel Physical Memory Devices and Systems
abstract
The digital universe is doubling every two years and expected to reach an unwieldy 44 zettabytes into the next decade. To cope with the ever increasing need for storing, transmitting and retrieving huge amounts of data, cloud storage, data centers and other massively distributed storage networks have emerged. These rely on efficient memory technologies at the physical level for speed, reliability and energy efficiency.
Shayan Garani Srinivasa, Tong Zhang 0002, Ravi Motwani, Haralampos Pozidis, Bane Vasic
IEEE J. Sel. Areas Commun.4
2015 Endurance limits of MLC NAND flash
abstract
An extensive effort is being undertaken by the flash community to develop signal processing and error-correction coding schemes that make use of soft information. Using experimental data from a state-of-the-art MLC flash device we demonstrate that the theoretical endurance improvement that such schemes can bring is limited. To investigate further, we develop a parametric channel model that takes into account the effects of cell-to-cell interference and demonstrate that it is the presence of programming errors in the channel that restricts the potential endurance enhancement that soft information can offer.
Thomas P. Parnell, Celestine Dünner, Thomas Mittelholzer, Nikolaos Papandreou, Haralampos Pozidis
ICC5
2015 Symmetry-based subproduct codes
abstract
Recently, a new type of product-like codes, known as half-product codes, have been studied for OTN applications. Motivated by these codes, new classes of symmetry-invariant subproduct codes are proposed and investigated under iterative hard-decision decoding. A subset of the new class of quarter product codes has lower error floors than comparable half-product codes in terms of length, rate and performance.
Thomas Mittelholzer, Thomas P. Parnell, Nikolaos Papandreou, Haralampos Pozidis
ISIT4
2015 Enhancing the Reliability of MLC NAND Flash Memory Systems by Read Channel Optimization
abstract
NAND flash memory is not only the ubiquitous storage medium in consumer applications but has also started to appear in enterprise storage systems as well. MLC and TLC flash technology made it possible to store multiple bits in the same silicon area as SLC, thus reducing the cost per amount of data stored. However, at current sub-20nm technology nodes, MLC flash devices fail to provide the levels of raw reliability, mainly cycling endurance, that are required by typical enterprise applications. Advanced signal processing and coding schemes are needed to improve the flash bit error rate and thus elevate the device reliability to the desired level. In this article, we report on the use of adaptive voltage thresholds and cell-to-cell interference cancellation in the read operation of NAND flash devices. We discuss how the optimal read voltage thresholds can be determined and assess the benefit of cancelling cell-to-cell interference in terms of cycling endurance, data retention, and resilience to read disturb.
Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Thomas Mittelholzer, Evangelos Eleftheriou, Charles Camp, Thomas Griffin, Gary A. Tressler, Andrew Walls
ACM Trans. Design Autom. Electr. Syst.3
2014 Modelling of the threshold voltage distributions of sub-20nm NAND flash memory
abstract
The proliferation of NAND flash memory in consumer devices has driven their aggressive cost reduction by continuous scaling to smaller technology nodes. However, this relentless cost per capacity improvement has diminished the reliability of flash memory to a degree that advanced signal processing and error correction are needed to enhance signal integrity in current flash-based systems. Accurate models of flash readback signals are necessary to properly design such advanced signal enhancement schemes. We propose a new parametric model of the flash readback signal based on fitting threshold voltage distributions from NAND flash devices. We show accurate fitting results for flash devices cycled up to 10 times longer than their nominal endurance specification, and provide simple expressions of the model parameters as a function of program/erase cycles. Finally, we also demonstrate that the proposed model can be used to capture effects such as programming errors, that occur in over-stressed flash devices.
Thomas P. Parnell, Nikolaos Papandreou, Thomas Mittelholzer, Haralampos Pozidis
GLOBECOM4
2014 Using adaptive read voltage thresholds to enhance the reliability of MLC NAND flash memory systems
abstract
NAND Flash memory is not only the ubiquitous storage medium in consumer applications, but has also started to appear in enterprise storage systems as well. MLC and TLC Flash technology made it possible to store multiple bits in the same silicon area as SLC, thus reducing the cost per amount of data stored. However, at current sub-20nm technology nodes, MLC Flash devices fail to provide the levels of raw reliability, mainly cycling endurance, that are required by typical enterprise applications. Advanced signal-processing and coding schemes are needed to improve the Flash bit error rate and thus elevate the device reliability to the desired level. In this paper, we report on the use of adaptive voltage thresholds in the read operation of NAND Flash devices. We discuss how the optimal read voltage thresholds can be determined, and assess the benefit of adapting the read voltage thresholds in terms of cycling endurance, data retention and resilience to read disturb.
Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Thomas Mittelholzer, Evangelos Eleftheriou, Charles Camp, Thomas Griffin, Gary A. Tressler, Andrew Walls
ACM Great Lakes Symposium on VLSI3
2011 Programming algorithms for multilevel phase-change memory
abstract
Phase-change memory (PCM) has emerged as one among the most promising technologies for next-generation non-volatile solid-state memory. Multilevel storage, namely storage of non-binary information in a memory cell, is a key factor for reducing the total cost-per-bit and thus increasing the competiveness of PCM technology in the nonvolatile memory market. In this paper, we present a family of advanced programming schemes for multilevel storage in PCM. The proposed schemes are based on iterative write-and-verify algorithms that exploit the unique programming characteristics of PCM in order to achieve significant improvements in resistance-level packing density, robustness to cell variability, programming latency, energy- per-bit and cell storage capacity. Experimental results from PCM test-arrays are presented to validate the proposed programming schemes. In addition, the reliability issues of multilevel PCM in terms of resistance drift and read noise are discussed.
Nikolaos Papandreou, Haralampos Pozidis, Angeliki Pantazi, Abu Sebastian, Matthew J. Breitwisch, Chung Hon Lam, Evangelos Eleftheriou
ISCAS2
2010 Channel Modeling and Signal Processing for Probe Storage Channels
abstract
Probe-storage devices employ large arrays of probes to write/read data in parallel in some storage medium, and combine ultra-high density, low access times, and low power consumption. A particular probe-storage technique utilizes thermomechanical means to store and retrieve information in thin polymer films. In this paper, a system-level channel model for the thermomechanical probe-storage channel is presented. Each of the components of the proposed model is derived by extensive characterization of experimentally obtained readback signals from probe recording tests. Moreover, detection techniques that are actually utilized in a probe-storage prototype implementation are described, followed by coding techniques for added reliability in the presence of particles or other impurities of the storage medium. In addition to low-complexity coding constructs, a concatenated coding scheme with an outer LDPC and inner modulation code is considered, in order to establish a benchmark for overall system performance. A novel methodology for joint decoding of outer LDPC and inner (d,k) modulation codes is developed. Furthermore, an optimal soft decoder for the modulation code is proposed, based on a modification of the decoder metrics to accurately account for the probe storage channel output statistics. Experimental results are used throughout the paper to validate the channel model and identify its relevant parameters, as well as to verify the system performance obtained by simulations.
Haralampos Pozidis, Giovanni Cherubini, Angeliki Pantazi, Abu Sebastian, Evangelos Eleftheriou
IEEE J. Sel. Areas Commun.1
2009 Performance Evaluation of the Probe Storage Channel
abstract
Scanning probes can be used for storage of data at ultra-high areal densities, beyond those achieved by other techniques. Thermo-mechanical probe storage is one variant of scanning-probe technology in which data is stored in the form of indentations in thin polymer films. A simplified channel model of thermo-mechanical probe storage is first introduced, derived from a comprehensive characterization study of experimentally extracted read-back signals. This model is then used as basis for deriving analytical expressions for the probability density function of the storage channel output signal in the presence of stochastic impairments such as noise and jitter. These expressions are used to construct an optimal bit-by-bit channel detector and to derive a theoretical formula for its bit error rate. Finally, the results of the analytical study are validated by comparison with simulation results. Analytical expressions for the channel output statistics are valuable as they enable quick performance comparisons between candidate detection schemes without resorting to lengthy simulations. Furthermore, they can be used to design probabilistic/soft detectors for soft-input error-correction schemes.
Thomas P. Parnell, Haralampos Pozidis, Oleg V. Zaboronski
GLOBECOM2
2009 Probabilistic Data Detection for Probe-Based Storage Channels in the Presence of Jitter
abstract
Probe-storage devices employ large arrays of probes to write/read data in parallel in some storage medium. Because of their inherent parallelism and the independence of data retrieved from different probes, these devices lend themselves naturally to low-density parity-check (LDPC) codes and associated soft/iterative decoding techniques. In this paper, a concatenated coding scheme for a particular probe-storage channel is presented that comprises an inner (d, k)-constrained code and an outer LDPC code, and the problem of probabilistic data retrieval for this channel is addressed. In particular, soft information is generated by explicitly accounting for the channel output statistics in the presence of jitter and additive noise, based on derived analytical expressions.
Haralampos Pozidis, Giovanni Cherubini
ICC1
2008 Forward Message Passing Detector for Probe Storage
abstract
We propose a simple soft-output detection scheme for communication channels characterized by data-dependent noise and/or inter-symbol interference (ISI). The proposed detection algorithm is based on the idea of forward message passing. We discuss a particular embodiment of the corresponding detector for a thermomechanical probe storage read channel. We perform an extensive performance investigation of the forward message passing detector using a simplified probe storage channel model characterized by non-linear inter-symbol interference and non-linear jitter distortion.
Thomas P. Parnell, Haralampos Pozidis, Oleg V. Zaboronski
ICC2
2007 Jitter Investigation and Performance Evaluation of a Small-Scale Probe Storage Device Prototype
abstract
MEMS-based scanning-probe data storage devices are emerging as potential ultra-high-density, low-access-time, and low-power alternatives to conventional data storage. Thermomechanical probe-based storage on thin polymer films is arguably the most advanced scanning-probe data storage scheme. The performance evaluation of a small-scale storage device prototype based on this concept is presented. The emphasis is on understanding the timing jitter in the read-back signals. Experiments are performed that confirm that the primary source of timing-jitter is the nanometer-scale perturbations of the micro-scanner while positioning the recording medium relative to the read/write transducers. Analytical estimates of these micro-scanner perturbations are obtained. An extensive performance evaluation, using the experimentally identified channel and medium-noise spectral characteristics, is conducted to study the impact of the microscanner perturbations on the performance of the storage device.
Abu Sebastian, Angeliki Pantazi, Haralampos Pozidis
GLOBECOM3
2005 Signal processing for probe storage
abstract
Scanning-probe data storage is emerging as a viable alternative to conventional data storage, offering ultra-high density, low access times, and low power consumption. One probe-storage technique utilizes a thermomechanical means to store and retrieve information in thin polymer films. We describe the readback signal path and characterize the thermomechanical-based probe-storage recording channel. It is shown that this channel exhibits a particular nonlinear behavior at high storage densities or high recording power, that is, the energy per unit time used to write a bit of information. A simple model is proposed that accurately captures the characteristics of this nonlinearity. Experimental results from single-probe recording setups are used to verify the validity of this model and identify its relevant parameters.
Haralampos Pozidis, Peter Bächtold, Giovanni Cherubini, Evangelos Eleftheriou, Christoph Hagleitner, Angeliki Pantazi, Abu Sebastian
ICASSP (5)1
2003 A Nanotechnology-based Approach to Data Storage
Evangelos Eleftheriou, Peter Bächtold, Giovanni Cherubini, Ajay Dholakia, Christoph Hagleitner, Teddy Loeliger, Angeliki Pantazi, Haralampos Pozidis, T. R. Albrecht, Gerd Karl Binnig, Michel Despont, Ute Drechsler, Urs Dürig, Bernd Gotsmann, Daniel Jubin, Walter Häberle, Mark A. Lantz, Hugo E. Rothuizen, Richard Stutz, Peter Vettiger, Dorothea Wiesmann
VLDB8
2002 Modeling and compensation of asymmetry in optical recording
abstract
Asymmetry, also known as domain bloom, is a systematic imperfection caused by the writing process in optical or magneto-optical recording. At the reading end of the system, asymmetry causes shifts of adjacent signal transitions in opposite directions. We present a simple nonlinear model for the replay signal in the presence of asymmetry. The model is specified by a single parameter, which is proportional to asymmetry, and its accuracy is demonstrated by application to replay signals from digital video disk drives. Based on the proposed model, a maximum-likelihood sequence detector is designed for replay signals with asymmetry. Simple modifications of the proposed detector lead to significant reductions in complexity, while the attainable performance, evaluated both analytically and through error-rate simulations, is superior to that of conventional techniques.
Haralampos Pozidis, Jan W. M. Bergmans, Wim M. J. Coene
IEEE Trans. Commun.1
2001 Run-length limited parity-check coding for transition-shift errors in optical recording
abstract
Random channel errors that may occur in the run-length limited (RLL) channel bit-stream after bit-detection, are commonly repaired by error correction (de)coding (ECC) after demodulation of the RLL bitstream. RLL parity-check coding makes it possible to correct random errors already at the level of the channel bit-stream, at a much lower overhead than needed for standard ECC. With parity-check coding, a certain parity-check constraint is realized on segments of the RLL channel bit-stream. Violation of the parity-check constraint in the as-detected RLL bit-stream enables error detection; for error correction, some side information from the channel waveform is needed. One scheme from the literature considers parsing of parity-check blocks into the original RLL bit-stream. Another scheme is concatenated parity-check coding. While the parsing scheme has the drawback of low coding efficiency, its advantage is its simplicity and the lack of error propagation. The concatenated scheme has a high efficiency but it suffers from error propagation, and local demodulation is needed to derive the parity checks. As an alternative solution, we propose to use a parity-check coding scheme based on a combination of codes. Such a scheme is called 'combi-code', a term that was proposed within the framework of DC-free RLL codes. Apart from a main RLL code, a second RLL code, the parity-check enabling code, is required. The latter code is used to set the parity-check constraint of a segment to a predetermined value. Parity-check coding via combi-codes combines the advantages of the two other schemes: simplicity, a high coding efficiency, and no error propagation.
Wim M. J. Coene, Haralampos Pozidis, Jan W. M. Bergmans
GLOBECOM2
2000 A Simple Nonlinear Model for the Optical Recording Channel
abstract
This study is concerned with the development of a simple yet sufficiently accurate model for the replay signal in optical discs. A method is proposed which models data storage (the write channel) as a nonlinear process and data retrieval (the read channel) as a linear one. The result is a low-complexity nonlinear model, of which the linear binary-PAM model is a special case. The performance of the proposed model is demonstrated by application to experimentally measured DVD data, and is found to be both accurate as well as robust to different recording media and conditions.
Haralampos Pozidis, Wim M. J. Coene, Jan W. M. Bergmans
ICC (1)1
1998 Use of selected HOS information for low-variance estimation of bandlimited systems with short data records
abstract
Although the reconstruction of a nonminimum-phase system excited by a stationary non-Gaussian white input is only possible using higher-order statistics (HOS) of the system output, there has been a lot of criticism in the literature against the amount of data required for keeping estimation errors low, and the complexity involved. Several attempts for reducing the variance of the HOS estimates have appeared. In the case of bandlimited signals, we have demonstrated via simulations that the estimation variance can be reduced if "good" slices, instead of the whole bispectrum, are used. This suggests a potential reduction of the variance in the system estimates, without having to resort to long observations. We justify theoretically the dependence of the system estimate variance on the bispectrum slice, and the criterion of slice selection. We also present simulation results, where the selected-slices approach appears to result in much lower estimation variance, as compared to other entire-bispectrum based approaches, for data lengths as low as 64 samples.
Haralampos Pozidis, Athina P. Petropulu
ICASSP1
1997 Signal reconstruction from phase only information and application to blind system estimation
abstract
We propose a method for the reconstruction of a complex signal from its Fourier phase only, where the phase is known within a linear phase term, and the sequence's length is unknown. The case of the phase known exactly has received a lot of attention in the past, however, in most cases the phase can be estimated up to a linear phase term whose slope is unknown. Moreover, in most cases of interest, the exact length of the sequence which is to be recovered is unknown. As an application of the reconstruction from phase technique, we propose a method for blind channel identification.
Haralampos Pozidis, Athina P. Petropulu
ICASSP1