EDBT 2026 Demo / reviewers in the wild / expert
Nikolaos Papandreou
dblp:65/6743
· DBLP profile ↗
24ranked-venue papers
7as first author
4since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 3 since 2021Computer networks · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Acceleration of Decision-Tree Ensemble Models on the IBM Telum ProcessorabstractThis paper presents a tensor-based algorithm that leverages a hardware accelerator for inferencing decision-tree-based machine learning models. The algorithm has been integrated in a public software library and is demonstrated on an IBM z16 server, using the Telum processor with the Integrated Accelerator for AI. We describe the architecture and implementation of the algorithm and present experimental results that demonstrate its superior runtime performance compared with popular CPU-based machine learning inference implementations. Nikolaos Papandreou, Jan van Lunteren, Andreea Anghel, Thomas P. Parnell, Martin Petermann, Milos Stanisavljevic, Cédric Lichtenau, Andrew Sica, Dominic Röhm, Elpida Tzortzatos, Haralampos Pozidis |
ISCAS | 1 |
| 2022 | AI accelerator on IBM telum processor: industrial productabstractIBM Telum is the next generation processor chip for IBM Z and LinuxONE systems. The Telum design is focused on enterprise class workloads and it achieves over 40% per socket performance growth compared to IBM z15. The IBM Telum is the first server-class chip with a dedicated on-chip AI accelerator that enables clients to gain real time insights from their data as it is getting processed. Cédric Lichtenau, Alper Buyuktosunoglu, Ramon Bertran Monfort, Peter Figuli, Christian Jacobi 0002, Nikolaos Papandreou, Haralampos Pozidis, Anthony Saporito, Andrew Sica, Elpida Tzortzatos |
ISCA | 6 |
| 2021 | Differentially Private Stochastic Coordinate DescentabstractIn this paper we tackle the challenge of making the stochastic coordinate descent algorithm differentially private. Compared to the classical gradient descent algorithm where updates operate on a single model vector and controlled noise addition to this vector suffices to hide critical information about individuals, stochastic coordinate descent crucially relies on keeping auxiliary information in memory during training. This auxiliary information provides an additional privacy leak and poses the major challenge addressed in this work. Driven by the insight that under independent noise addition, the consistency of the auxiliary information holds in expectation, we present DP-SCD, the first differentially private stochastic coordinate descent algorithm. We analyze our new method theoretically and argue that decoupling and parallelizing coordinate updates is essential for its utility. On the empirical side we demonstrate competitive performance against the popular stochastic gradient descent alternative (DP-SGD) while requiring significantly less tuning. Georgios Damaskinos, Celestine Dünner, Rachid Guerraoui, Nikolaos Papandreou, Thomas P. Parnell |
AAAI | 4 |
| 2021 | High-Throughput ECC with Integrated Chipkill Protection for Nonvolatile Memory ArraysabstractNew coding schemes based on generalized concatenated codes are proposed for emerging nonvolative memory technologies. Key requirements are high code rate, low latency, high throughput and the ability to correct chipkill failures while sustaining high data reliability despite raw bit error rates up to 10-3. New concatenated codes based on Reed-Solomon codes have been designed for payload sizes of 512B, 1kB, and 2kB; they have high rates above 0.8 and a high data reliability with a decoder target BER of 10-15. An FPGA-based implementation of the decoder validates the low latency and high throughput: for an operating clock frequency of 250MHz, the decoding latency is 236ns and a 9.3GB/s throughput is achieved. Thomas Mittelholzer, Milos Stanisavljevic, Nikolaos Papandreou, Haralampos Pozidis |
ISCAS | 3 |
| 2020 | Improving NAND flash performance with read heat separationabstractThe continuous growth in 3D-NAND flash storage density has primarily been enabled by 3D stacking and by increasing the number of bits stored per memory cell. Unfortunately, these desirable flash device design choices are adversely affecting reliability and latency characteristics. In particular, increasing the number of bits stored per cell results in having to apply additional voltage thresholds during each read operation, therefore increasing the read latency characteristics. While most NAND flash challenges can be mitigated through appropriate background processing, the flash read latency characteristics cannot be hidden and remains the biggest challenge, especially for the newest flash generations that store four bits per cell. In this paper, we introduce read heat separation (RHS), a new heat-aware data-placement technique that exploits the skew present in real-world workloads to place frequently read user data on low-latency flash pages. Although conceptually simple, such a technique is difficult to integrate in a flash controller, as it introduces a significant amount of complexity, requires more metadata, and is further constrained by other flash-specific peculiarities. To overcome these challenges, we propose a novel flash controller architecture supporting read heat-aware data placement. We first discuss the trade-offs that such a new design entails and analyze the key aspects that influence the efficiency of RHS. Through both, extensive simulations and an implementation we realized in a commercial enterprise-grade solid-state drive controller, we show that our architecture can indeed significantly reduce the average read latency. For certain workloads, it can reverse the system-level read latency trends when using recent multi-bit flash generations and hence outperform SSDs using previous faster flash generations. Roman A. Pletka, Nikolaos Papandreou, Radu Stoica, Haralampos Pozidis, Nikolas Ioannou, Timothy Fisher, Aaron Fry, Kip Ingram, Andrew Walls |
MASCOTS | 2 |
| 2020 | SnapBoost: A Heterogeneous Boosting MachineabstractModern gradient boosting software frameworks, such as XGBoost and LightGBM, implement Newton descent in a functional space. At each boosting iteration, their goal is to find the base hypothesis, selected from some base hypothesis class, that is closest to the Newton descent direction in a Euclidean sense. Typically, the base hypothesis class is fixed to be all binary decision trees up to a given depth. In this work, we study a Heterogeneous Newton Boosting Machine (HNBM) in which the base hypothesis class may vary across boosting iterations. Specifically, at each boosting iteration, the base hypothesis class is chosen, from a fixed set of subclasses, by sampling from a probability distribution. We derive a global linear convergence rate for the HNBM under certain assumptions, and show that it agrees with existing rates for Newton's method when the Newton direction can be perfectly fitted by the base hypothesis at each boosting iteration. We then describe a particular realization of a HNBM, SnapBoost, that, at each boosting iteration, randomly selects between either a decision tree of variable depth or a linear regressor with random Fourier features. We describe how SnapBoost is implemented, with a focus on the training complexity. Finally, we present experimental results, using OpenML and Kaggle datasets, that show that SnapBoost is able to achieve better generalization loss than competing boosting frameworks, without taking significantly longer to tune. Thomas P. Parnell, Andreea Anghel, Malgorzata Lazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, Haralampos Pozidis |
NeurIPS | 7 |
| 2019 | Understanding the Design Trade-Offs of Hybrid Flash ControllersabstractOver the last few years, NAND flash manufacturers have steadily increased the number of bits stored per cell to achieve significant cost reductions. However, the increased density does not come without drawbacks. All key flash performance metrics, including latency and endurance, significantly degrade as bit density increases. Particularly, sustained write throughput is the worst affected as writes are roughly one order of magnitude slower than reads and further require precursory block erases in the background. As a result, many recent flash controllers operate flash blocks both in single-bit (high endurance and performance) and in multi-bit (high density) mode. In theory, such hybrid controllers are a great way of hiding flash technology limitations. A controller can use a small percentage of the flash blocks in single-bit mode as a cache which allows orders of magnitude higher write bandwidth and endurance in environments where the access patterns of the workload are skewed and bursty. In practice, however, many devices fall short of expectations when write performance varies significantly and utilization increases. We argue that a principled approach is required to understand the design trade-offs of hybrid NAND flash controllers. To this end, we develop a modeling framework for estimating the performance and endurance of hybrid controllers. The modeling framework computes the internal data movement generated by a hybrid controller by relying on advanced analytical models that offer both accurate and fast predictions. The data flow is then translated into higher-level metrics that quantify upper bounds for the overall performance of an SSD such as write throughput, latency, and device endurance. Using our modeling framework, we compare different controller architectures, identify their strong and weak points, and show that there is room to improve the efficiency of the hybrid controllers used today. Radu Stoica, Roman A. Pletka, Nikolas Ioannou, Nikolaos Papandreou, Sasa Tomic, Haralampos Pozidis |
MASCOTS | 4 |
| 2019 | Accelerated ML-Assisted Tumor Detection in High-Resolution Histopathology Images
Nikolas Ioannou, Milos Stanisavljevic, Andreea Anghel, Nikolaos Papandreou, Sonali Andani, Jan Hendrik Rüschoff, Peter Wild, Maria Gabrani, Haralampos Pozidis |
MICCAI (1) | 4 |
| 2018 | Exploiting the non-linear current-voltage characteristics for resistive memory readoutabstractVarious resistive memory technologies are finding application in the space of storage-class memory and emerging non-von Neumann computing systems. For both applications, a key enabling technology is the ability to store multiple resistance levels in a single memory cell. The resistance states of these devices are typically measured in the low-field regime, where the electrical transport can be assumed to be Ohmic. However, when biased at slightly higher voltages, they exhibit significantly nonlinear I-V characteristics. In this paper, we demonstrate how this field dependence of the resistance values can be exploited in various applications. We present simulation and experimental results where readout schemes based on the non-linear I-V behavior are used to enhance the readout margin and also to compensate for resistance drift. Nikolaos Papandreou, Abu Sebastian, Haralampos Pozidis |
ISCAS | 1 |
| 2018 | Drift-Invariant Detection for Multilevel Phase-Change MemoryabstractNext-generation memory (NGM) technologies present a major opportunity but also a significant challenge, due to their intricate reliability issues. In particular, multilevel-cell (MLC) storage is highly desirable for increasing storage capacity and lowering total cost-per-bit. In phase-change memory (PCM), MLC storage is hampered by sensitivity to temperature variations and resistance drift. A novel drift-invariant detection (DID) scheme that estimates variable read thresholds based on ordered statistics and clustering of the soft read-back signals from a small block of 32 cells has been developed and implemented in hardware to improve reliability and prolong data retention. A low-complexity implementation of the DID on a FPGA platform comprises 20'000 LUTs and 6'000 flip-flops and has a latency of 90ns. We present results from an extensive performance verification that ascertains highly reliable data retrieval up to 13 orders of magnitude in time after programming. Such elevated reliability is necessary for the most anticipated application of NGM, namely persistent far-memory, where the NGM is used as a large memory pool, possibly together with DRAM. Milos Stanisavljevic, Thomas Mittelholzer, Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis |
ISCAS | 3 |
| 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency TargetsabstractDespite its widespread use in consumer devices and enterprise storage systems, NAND flash faces a growing number of challenges. While technology advances have helped to increase the storage density and reduce costs, they have also led to reduced endurance and larger block variations, which cannot be compensated solely by stronger ECC or read-retry schemes but have to be addressed holistically. Our goal is to enable low-cost NAND flash in enterprise storage for cost efficiency. We present novel flash-management approaches that reduce write amplification, achieve better wear leveling, and enhance endurance without sacrificing performance. We introduce block calibration, a technique to determine optimal read-threshold voltage levels that minimize error rates, and novel garbage-collection as well as data-placement schemes that alleviate the effects of block health variability and show how these techniques complement one another and thereby achieve enterprise storage requirements. By combining the proposed schemes, we improve endurance by up to 15× compared to the baseline endurance of NAND flash without using a stronger ECC scheme. The flash-management algorithms presented herein were designed and implemented in simulators, hardware test platforms, and eventually in the flash controllers of production enterprise all-flash arrays. Their effectiveness has been validated across thousands of customer deployments since 2015. Roman A. Pletka, Ioannis Koltsidas, Nikolas Ioannou, Sasa Tomic, Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Aaron Fry, Timothy Fisher |
ACM Trans. Storage | 5 |
| 2016 | Controller architecture for low-latency access to phase-change memory in OpenPOWER systemsabstractNovel forms of nonvolatile memory, such as phase-change memory (PCM), promise low latency and small granularity of read and write access at high storage density. They also feature very high endurance. These characteristics make them highly desirable for emerging high-capacity (hybrid) memory applications such as in-memory databases and in-memory processing. In this work, we present the architecture, implementation and experimental performance results of an FPGA-based PCM memory controller for OpenPOWER servers. The memory controller leverages the Coherent Accelerator Processor Interface (CAPI) of the POWER processor in order to offer low-latency access to the CPU memory space. In addition, the memory controller implements an efficient management protocol that supports a dynamic size of pending read and write requests in order to offer high bandwidth under mixed-type workloads. We describe the architecture and implementation details of the memory controller and we demonstrate its performance using a prototype platform based on different types of OpenPOWER servers equipped with CAPI-enabled FPGA cards. The developed PCM controller is evaluated in terms of sustained data rates (MBps) and access latency (us). Experimental results are based on legacy commercial 90nm PCM chips as well as on accurate HW emulation of next generation PCM chips. Antonios Prodromakis, Nikolaos Papandreou, Eleni Bougioukou, Urs Egger, Nikos Toulgaridis, Theodore Antonakopoulos 0001, Haralampos Pozidis, Evangelos Eleftheriou |
FPL | 2 |
| 2016 | Tactile Identification of Embossed Raised Lines and Raised Squares with Variable Dot Elevation by Persons Who Are Blind
Georgios Kouroupetroglou, Aineias Martos, Nikolaos Papandreou, Konstantinos Papadopoulos 0001, Vassilis Argyropoulos, Georgios D. Sideridis |
ICCHP (2) | 3 |
| 2016 | Improving the error-floor performance of binary half-product codes
Thomas Mittelholzer, Thomas P. Parnell, Nikolaos Papandreou, Haralampos Pozidis |
ISITA | 3 |
| 2016 | Capacity of the MLC NAND Flash ChannelabstractIn this paper, we develop a framework for evaluating the symmetric capacity of multilevel-cell (MLC) NAND flash devices while making very few assumptions regarding the underlying device physics. A set of recursive equations are derived that allow one to measure the symmetric capacity for any given page in a flash device using simple conditional statistics that can be extracted experimentally. Using data captured from two different 1y nm MLC devices, we demonstrate that the symmetric capacity of a flash page not only depends on the amount of program/erase cycling and data retention stress that has accumulated, but also on the position of the page within the flash block. We then study the effect on symmetric capacity of using optimized read-back schemes (both hard and soft) and show that while there is significant benefit, not all pages in the block are improved by the same amount. Finally, we show that it is possible to design error correction architectures that harness the inherent variation of symmetric capacity within a flash block to dramatically extend the program/erase cycling endurance of flash-based storage systems. Thomas P. Parnell, Celestine Dünner, Thomas Mittelholzer, Nikolaos Papandreou |
IEEE J. Sel. Areas Commun. | 4 |
| 2015 | Endurance limits of MLC NAND flashabstractAn extensive effort is being undertaken by the flash community to develop signal processing and error-correction coding schemes that make use of soft information. Using experimental data from a state-of-the-art MLC flash device we demonstrate that the theoretical endurance improvement that such schemes can bring is limited. To investigate further, we develop a parametric channel model that takes into account the effects of cell-to-cell interference and demonstrate that it is the presence of programming errors in the channel that restricts the potential endurance enhancement that soft information can offer. Thomas P. Parnell, Celestine Dünner, Thomas Mittelholzer, Nikolaos Papandreou, Haralampos Pozidis |
ICC | 4 |
| 2015 | Symmetry-based subproduct codesabstractRecently, a new type of product-like codes, known as half-product codes, have been studied for OTN applications. Motivated by these codes, new classes of symmetry-invariant subproduct codes are proposed and investigated under iterative hard-decision decoding. A subset of the new class of quarter product codes has lower error floors than comparable half-product codes in terms of length, rate and performance. Thomas Mittelholzer, Thomas P. Parnell, Nikolaos Papandreou, Haralampos Pozidis |
ISIT | 3 |
| 2015 | Enhancing the Reliability of MLC NAND Flash Memory Systems by Read Channel OptimizationabstractNAND flash memory is not only the ubiquitous storage medium in consumer applications but has also started to appear in enterprise storage systems as well. MLC and TLC flash technology made it possible to store multiple bits in the same silicon area as SLC, thus reducing the cost per amount of data stored. However, at current sub-20nm technology nodes, MLC flash devices fail to provide the levels of raw reliability, mainly cycling endurance, that are required by typical enterprise applications. Advanced signal processing and coding schemes are needed to improve the flash bit error rate and thus elevate the device reliability to the desired level. In this article, we report on the use of adaptive voltage thresholds and cell-to-cell interference cancellation in the read operation of NAND flash devices. We discuss how the optimal read voltage thresholds can be determined and assess the benefit of cancelling cell-to-cell interference in terms of cycling endurance, data retention, and resilience to read disturb. Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Thomas Mittelholzer, Evangelos Eleftheriou, Charles Camp, Thomas Griffin, Gary A. Tressler, Andrew Walls |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2014 | Modelling of the threshold voltage distributions of sub-20nm NAND flash memoryabstractThe proliferation of NAND flash memory in consumer devices has driven their aggressive cost reduction by continuous scaling to smaller technology nodes. However, this relentless cost per capacity improvement has diminished the reliability of flash memory to a degree that advanced signal processing and error correction are needed to enhance signal integrity in current flash-based systems. Accurate models of flash readback signals are necessary to properly design such advanced signal enhancement schemes. We propose a new parametric model of the flash readback signal based on fitting threshold voltage distributions from NAND flash devices. We show accurate fitting results for flash devices cycled up to 10 times longer than their nominal endurance specification, and provide simple expressions of the model parameters as a function of program/erase cycles. Finally, we also demonstrate that the proposed model can be used to capture effects such as programming errors, that occur in over-stressed flash devices. Thomas P. Parnell, Nikolaos Papandreou, Thomas Mittelholzer, Haralampos Pozidis |
GLOBECOM | 2 |
| 2014 | Using adaptive read voltage thresholds to enhance the reliability of MLC NAND flash memory systemsabstractNAND Flash memory is not only the ubiquitous storage medium in consumer applications, but has also started to appear in enterprise storage systems as well. MLC and TLC Flash technology made it possible to store multiple bits in the same silicon area as SLC, thus reducing the cost per amount of data stored. However, at current sub-20nm technology nodes, MLC Flash devices fail to provide the levels of raw reliability, mainly cycling endurance, that are required by typical enterprise applications. Advanced signal-processing and coding schemes are needed to improve the Flash bit error rate and thus elevate the device reliability to the desired level. In this paper, we report on the use of adaptive voltage thresholds in the read operation of NAND Flash devices. We discuss how the optimal read voltage thresholds can be determined, and assess the benefit of adapting the read voltage thresholds in terms of cycling endurance, data retention and resilience to read disturb. Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Thomas Mittelholzer, Evangelos Eleftheriou, Charles Camp, Thomas Griffin, Gary A. Tressler, Andrew Walls |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Programming algorithms for multilevel phase-change memoryabstractPhase-change memory (PCM) has emerged as one among the most promising technologies for next-generation non-volatile solid-state memory. Multilevel storage, namely storage of non-binary information in a memory cell, is a key factor for reducing the total cost-per-bit and thus increasing the competiveness of PCM technology in the nonvolatile memory market. In this paper, we present a family of advanced programming schemes for multilevel storage in PCM. The proposed schemes are based on iterative write-and-verify algorithms that exploit the unique programming characteristics of PCM in order to achieve significant improvements in resistance-level packing density, robustness to cell variability, programming latency, energy- per-bit and cell storage capacity. Experimental results from PCM test-arrays are presented to validate the proposed programming schemes. In addition, the reliability issues of multilevel PCM in terms of resistance drift and read noise are discussed. Nikolaos Papandreou, Haralampos Pozidis, Angeliki Pantazi, Abu Sebastian, Matthew J. Breitwisch, Chung Hon Lam, Evangelos Eleftheriou |
ISCAS | 1 |
| 2009 | Architecture and DSP Implementation of a DVB-S2 Baseband DemodulatorabstractThis paper presents the design and implementation of a baseband demodulator for DVB-S2 satellite receivers. In order to meet the requirements of different complex and multidomain signal processing stages of the DVB-S2 baseband signal-flow, the presented architecture is based on efficient fixed-point implementation of the various demodulation algorithms and on the use of a dynamic time-sharing scheduler for the various DSP software tasks. The prototyping of the demodulator and its verification in the design of a complete digital DVB-S2 satellite receiver using a versatile testbed is also presented. Panayiotis Savvopoulos, Nikolaos Papandreou, Theodore Antonakopoulos 0001 |
DSD | 2 |
| 2008 | A low-complexity bandwidth allocation algorithm for frequency-selective multiuser OFDM systems
Nikolaos Papandreou, Theodore Antonakopoulos 0001 |
Comput. Commun. | 1 |
| 2005 | A new computationally efficient discrete bit-loading algorithm for DMT applicationsabstractThis letter presents a new bit-loading algorithm for discrete multitone systems that converges faster to the same bit allocation as the optimal discrete bit-filling and bit-removal methods. The algorithm exploits the differences between the subchannel gain-to-noise ratios in order to determine an initial bit allocation and then performs a multiple-bits loading procedure for achieving the requested target rate. Numerical results using asymmetric digital subscriber test loops demonstrate the computational efficiency of the proposed algorithm. Nikolaos Papandreou, Theodore Antonakopoulos 0001 |
IEEE Trans. Commun. | 1 |