EDBT 2026 Demo / reviewers in the wild / expert
Aleksandar Milenkovic
dblp:54/1358
· DBLP profile ↗
44ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-9359-4594ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 4Computer networks · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SRAM DNA: Spatial Signatures in Power-Up States for Memory Family IdentificationabstractEnsuring the authenticity of integrated circuits in complex semiconductor supply chains requires noninvasive methods for identifying device provenance. This work investigates whether SRAM power-up states contain structural signatures, referred to in this work as SRAM DNA, that distinguish memory families beyond traditional device-unique Physical Unclonable Functions (PUFs). We analyze eight commercial 4-Mbit SRAM families (15 chips per family) and collect multiple power-up measurements from each device. Each power-up state is reshaped into a 2048×2048 binary grid, enabling spatial analysis of the full memory array. From these grids we extract interpretable multi-scale descriptors capturing global bias, spatial periodicity, texture statistics, and directional anisotropy. A Random Forest classifier evaluated under chip-level holdout validation (Leave-M-Chips-Per-Family) achieves more than 99.5% accuracy and macro-F1 score, demonstrating strong generalization to unseen devices with only a few power-up measurements per chip. These results show that SRAM power-up states encode reproducible family-level spatial organization reflecting memory-array structure and manufacturing characteristics. Sayan Samanta, Biswajit Ray, Aleksandar Milenkovic |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | Exploring Machine Learning Algorithms for Analysing Students' Attitudes Towards Distance Mathematics LearningabstractABSTRACT This study investigates the application of machine learning algorithms to analyse students' attitudes towards distance mathematics education, focusing on perceived effectiveness and students' ability to successfully learn and adopt mathematical content in an online setting. Data were collected from 1154 students at various educational levels using a 28‐item Likert scale questionnaire on distance mathematics learning. An ML pipeline incorporating multiple data preparation techniques and machine learning algorithms was applied to two key prediction questions. Unlike predominantly descriptive or single model studies in this area, this approach evaluates both predictive performance and the stability of selected survey items across many model configurations, providing more robust and interpretable pedagogical insights. The results show that Recursive Feature Elimination was the most effective feature selection method, while Random Forest, Ridge Classifier and Categorical Naive Bayes achieved the strongest overall predictive performance across the two questions. These findings confirm the value of combining feature selection techniques and machine learning algorithms to derive robust and interpretable insights from educational survey data, while also highlighting the importance of well‐structured distance learning strategies in mathematics education. The methodology is readily adaptable to other academic disciplines, providing educators with a data‐driven framework for improving the design and effectiveness of online learning. Marina R. Svicevic, Aleksandar Milenkovic, Lazar Krstic, Milos Pavkovic |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | Page-Overwrite Data Sanitization in 3D NAND Flash: Challenges, Feasibility, and the PULSE SolutionabstractInstant data deletion (or sanitization) in NAND flash devices is essential for achieving data privacy, but it remains challenging due to the mismatch between erase and write granularities, which leads to high overhead and accelerated wear. While page-overwrite-based instant data sanitization has proven effective for 2D NAND, its applicability to 3D NAND is limited due to the unique sub-block architecture. In this study, we experimentally evaluate page-overwrite-based sanitization on commercial 3D NAND flash memory chips and uncover significant threshold voltage disturbances in erased cells on adjacent pages within the same layer but across different sub-blocks. Our key findings reveal that page-overwrite sanitization increases the median raw bit error rate (RBER) beyond correction limits (exceeding 0.93%) in Floating-Gate (FG) Single-Level Cell (SLC) technology, whereas Charge-Trap (CT) SLC 3D NAND flash memories exhibit higher robustness. In Triple-Level Cell (TLC) 3D NAND, page-overwrite sanitization proves impractical, with the median RBER of ∼13% for FG and ∼5% for CT devices. To overcome these challenges, we propose PULSE , a low-disturbance sanitization technique that balances sanitization efficiency ( \({{\eta }_{san}}\) ) and data integrity (RBER). Experimental results show that PULSE eliminates RBER increases in SLC devices and reduces the median RBER to below 0.57% for FG and 0.79% for CT in fresh TLC blocks, demonstrating its practical viability for 3D NAND flash sanitization. Matchima Buddhanoy, Aleksandar Milenkovic, Sudeep Pasricha, Biswajit Ray |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Improving Block Management in 3D NAND Flash SSDs with Sub-Block First Write SequencingabstractContinual vertical scaling in 3D NAND flash solid-state drives (SSDs) results in larger memory blocks, causing performance degradation due to big-block management issues. Pages within a 3D NAND flash block are traditionally written using layer first write sequencing. This paper introduces and explores the benefits of an alternative sub-block first write sequence. This method when coupled with sub-block erase operations promises to alleviate the big-block problem. Our evaluation on a commercial 32-layer 3D NAND flash SSD chip shows that though the proposed method increases the raw bit error rate (RBER), it remains below the threshold that can be corrected by error correction codes (ECCs). Simulation analysis further shows that our proposed method reduces garbage collection overhead, resulting in 36.0% lower response time and 9.6% reduction in additional writes due to garbage collection compared to traditional 3D NAND flash SSDs. Matchima Buddhanoy, Kamil Khan, Aleksandar Milenkovic, Sudeep Pasricha, Biswajit Ray |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Hide-and-Seek: Hiding Secrets in Threshold Voltage Distributions of NAND Flash Memory CellsabstractIn this paper, we propose a new page-writing technique to hide secret information using the threshold voltage variation of programmed memory cells. We demonstrate the proposed technique on the state-of-the-art commercial 3D NAND flash memory chips by utilizing common user mode commands. We explore the design space metrics of interest for data hiding: bit accuracy of public and secret data and detectability of holding secret data. The proposed method ensures more than 97% accuracy of recovered secret data, with negligible accuracy loss in the public data. Our analysis shows that the proposed technique introduces negligible distortions in the threshold voltage distributions. These distortions are lower than the inherent threshold voltage variations of program states. As a result, the proposed method provides a hiding technique that is undetectable, even by a powerful adversary with low-level access to the memory chips. Md Raquibuzzaman, Aleksandar Milenkovic, Biswajit Ray |
HotStorage | 2 |
| 2022 | Instant data sanitization on multi-level-cell NAND flash memoryabstractDeleting data instantly from NAND flash memories incurs hefty overheads, and increases wear level. Existing solutions involve unlinking the physical page addresses making data inaccessible through standard interfaces, but they carry the risk of data leakage. An all-zero-in-place data overwrite has been proposed as a countermeasure, but it applies only to SLC flash memories. This paper introduces an instant page data sanitization method for MLC flash memories that prevents leakage of deleted information without any negative effects on valid data in shared pages. We implement and evaluate the proposed method on commercial 2D and 3D NAND flash memory chips. Md Raquibuzzaman, Matchima Buddhanoy, Aleksandar Milenkovic, Biswajit Ray |
SYSTOR | 3 |
| 2021 | Microcontroller Fingerprinting Using Partially Erased NOR Flash Memory CellsabstractElectronic device fingerprints, unique bit vectors extracted from device's physical properties, are used to differentiate between instances of functionally identical devices. This article introduces a new technique that extracts fingerprints from unique properties of partially erased NOR flash memory cells in modern microcontrollers. NOR flash memories integrated in modern systems-on-a-chip typically hold firmware and read-only data, but they are increasingly in-system-programmable, allowing designers to erase and program them during normal operation. The proposed technique leverages partial erase operations of flash memory segments that bring them into the state that exposes physical properties of the flash memory cells through a digital interface. These properties reflect semiconductor process variations and defects that are unique to each microcontroller or a flash memory segment within a microcontroller. The article explores threshold voltage variation in NOR flash memory cells for generating fingerprints and describes an algorithm for extracting fingerprints. The experimental evaluation utilizing a family of commercial microcontrollers demonstrates that the proposed technique is cost-effective, robust, and resilient to changes in voltage and temperature as well as to aging effects. Prawar Poudel, Biswajit Ray, Aleksandar Milenkovic |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Flashmark: Watermarking of NOR Flash Memories for Counterfeit DetectionabstractCounterfeit electronics represents a significant concern because recycled, over-produced, out-of-spec, or cloned chips can enter globalized supply chains. This paper introduces Flashmark, a technique for watermarking of NOR flash memories for counterfeit detection. Flashmark relies on novel approaches for (a) imprinting watermarks into dedicated blocks of flash memory chips by repeated stressing, thus irreversibly changing physical properties of flash memory cells, and (b) reading out watermarks by sensing the changes in physical properties of flash cells through standard digital interfaces. The paper demonstrates Flashmark on embedded NOR flash memories in a family of low-cost microcontrollers. Prawar Poudel, Biswajit Ray, Aleksandar Milenkovic |
DAC | 3 |
| 2019 | SPEC CPU2017: Performance, Event, and Energy Characterization on the Core i7-8700KabstractComputer engineers in academia and industry rely on a standardized set of benchmarks to quantitatively evaluate the performance of computer systems and research prototypes. SPEC CPU2017 is the most recent incarnation of standard benchmarks designed to stress a system's processor, memory subsystem, and compiler. This paper describes the results of measurement-based studies focusing on characterization, performance, and energy-efficiency analyses of SPEC CPU2017 on the Intel's Core i7-8700K. Intel and GNU compilers are used to create executable files utilized in performance studies. The results show that executables produced by the Intel compilers are superior to those produced by GNU compilers. We characterize all the benchmarks, perform a top-down microarchitectural analysis to identify performance bottlenecks, and test benchmark scalability with respect to performance and energy. Findings from these studies can be used to guide future performance evaluations and computer architecture research Ranjan Hebbar, Aleksandar Milenkovic |
ICPE | 2 |
| 2019 | Microcontroller TRNGs Using Perturbed States of NOR Flash Memory CellsabstractThis paper introduces a new technique that perturbs split-gate NOR Flash memory cells and extracts randomness of read noise to generate true random numbers. Flash memory cells exhibit threshold voltage fluctuations during read operations caused by thermal noise and random telegraph noise effects. Recent proposals demonstrate how these inherent properties of Flash memory cells can be used to create true random numbers in modern NAND Flash memories. However, they cannot be directly applied to NOR Flash memories in microcontrollers that have different architecture, improved data retention, high endurance, and are not as susceptible to noise as high-density NAND Flash memories. The proposed technique is experimentally demonstrated and evaluated using a family of commercial microcontrollers. The evaluation shows that it enables extraction of high-throughput random sequences that pass the NIST statistical tests. Advantages of the proposed technique are as follows: (a) it does not require any special hardware and/or interface modifications, (b) it is robust, cost-effective, and high-throughput, (c) it is entirely implemented in software, and (d) it is flexible and can be tailored to work in low-end microcontrollers that are often resource- or cost-constrained. Prawar Poudel, Biswajit Ray, Aleksandar Milenkovic |
IEEE Trans. Computers | 3 |
| 2019 | Enabling On-the-Fly Hardware Tracing of Data Reads in MulticoresabstractSoftware debugging is one of the most challenging aspects of embedded system development due to growing hardware and software complexity, limited visibility of system components, and tightening time-to-market. To find software bugs faster, developers often rely on on-chip trace modules with large buffers to capture program execution traces with minimum interference with program execution. However, the high volumes of trace data and the high cost of trace modules limit the visibility into the system operation to short program segments. This article introduces a new hardware/software technique for capturing and filtering read data value traces in multicores that enables a complete reconstruction of parallel program execution. The proposed technique exploits tracking of data reads in data caches and cache coherence protocol states to minimize the number of trace messages streamed out of the target platform to the software debugger. The effectiveness of the proposed technique is determined by analyzing the required trace port bandwidth and trace buffer sizes as a function of the data cache size and the number of processor cores. The results show that the proposed technique significantly reduces the required trace port bandwidth, from 12.2 to 73.9 times, when compared to the Nexus-like read data value tracing, thus enabling continuous on-the-fly data tracing at modest hardware cost. Mounika Ponugoti, Aleksandar Milenkovic |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | A framework for optimizing file transfers between mobile devices and the cloudabstractAn exponential growth of data traffic that originates on mobile devices and a shift toward cloud computing necessitate finding new approaches to optimize file transfers. Whereas compression utilities can improve effective throughput of file transfers between mobile devices and the cloud, finding the best-performing utility for a given file transfer is a challenging task. In this paper we introduce a framework for optimizing file transfers that relies on agents running on both mobile devices and the cloud. The agents are responsible to select an effective transfer mode by considering characteristics of files to be transferred, network conditions, and mobile device performance. The framework is implemented and experimentally evaluated for file uploads and downloads initiated from a smartphone, while varying WLAN network conditions. The results of the evaluation show that the framework effectively increases upload throughputs in range from 1.46 to 2.38 times relative to uncompressed uploads and from 1.02 to 2.97 times relative to default compressed uploads. Similarly, it improves download throughputs in range of 1.47 to 2.5 times relative to uncompressed downloads and up to 1.2 times relative to default compressed downloads. Armen Dzhagaryan, Aleksandar Milenkovic |
PIMRC | 2 |
| 2017 | Data summarization method for chronic disease tracking
Dejan Aleksic, Petar Rajkovic, Dusan Vuckovic, Dragan Jankovic, Aleksandar Milenkovic |
J. Biomed. Informatics | 5 |
| 2016 | On-the-fly load data value tracing in multicoresabstractSoftware testing and debugging of modern multicore-based embedded systems is a challenging proposition because of growing hardware and software complexity, increased integration, and tightening time-to-market. To find more bugs faster, software developers of real-time embedded systems increasingly rely on on-chip trace and debug resources, including hefty on-chip buffers and wide trace ports. However, these resources often offer limited visibility of the system, increase the system cost, and do not scale well with a growing number of cores. This paper introduces mlvCFiat, a hardware/software mechanism for capturing and filtering load data value traces in multicores. It relies on first-access tracking in data caches and equivalent modules in the software debugger to significantly reduce the number of trace events streamed out of the target platform. Our experimental evaluation explores the effectiveness of the proposed technique as a function of cache sizes, encoding mechanism, and the number of cores. The results show that mlvCFiat significantly reduces the total trace port bandwidth. The improvements relative to the existing Nexus-like load data value tracing range from 15 to 33 times for a single core and from 14 to 20 times for an octa core. Mounika Ponugoti, Amrish K. Tewar, Aleksandar Milenkovic |
CASES | 3 |
| 2016 | Adaption of medical information system's e-learning extension to a simple suggestion toolabstractDeveloping suggestion tools in the scope of health information systems can be a complex task, followed by a risk of not being accepted by the end users. Thus, we decide to start the implementation around the existing functionality. In this paper we present a case study showing the adaptation of e-learning medical information system extension to a set of simple suggestion tools. While some features of initial system had to be modified, the domain specific knowledge collected for the e-learning extension is used to suppress potential errors. Presented suggestion tool is based on highly configurable lists of pre-defined entities that can be easily selected, and after the verification from the medical practitioner, copied into an active visit. After four years of active use, and several iteration of update, described suggestion tools are mostly accepted among the general practitioners, especially within certain scenarios where faster medication prescription is a must. Petar Rajkovic, Dragan Jankovic, Aleksandar Milenkovic |
HealthCom | 3 |
| 2016 | Models for Evaluating Effective Throughputs for File Transfers in Mobile ComputingabstractThe importance of optimizing data transfers between mobile computing devices and the cloud is increasing with an exponential growth of mobile data traffic. Lossless data compression can be essential in increasing communication throughput, reducing communication latency, achieving energy-efficient communication, and making effective use of available storage. In this paper we introduce analytical models for estimating effective through-put of uncompressed data transfers and compressed data transfers that utilize common compression utilities. The proposed analytical models are experimentally verified using state-of-the-art smartphones as mobile devices. These models are instrumental in developing a framework for seamless optimization of file transfers in mobile computing. Armen Dzhagaryan, Aleksandar Milenkovic |
ICCCN | 2 |
| 2016 | Exploiting cache coherence for effective on-the-fly data tracing in multicoresabstractSoftware testing and debugging of modern embedded computer systems become increasingly a challenging task due to growing hardware and software complexity, increased integration and miniaturization, and ever tightening time-to-market. To find software bugs faster, developers often rely on on-chip trace and debug resources. However, these resources offer limited visibility of the system, increase the system cost, and do not scale well with a growing number of processor cores. This paper introduces a new hardware/software mechanism for capturing and filtering load data value traces in multicores that enables a complete reconstruction of a parallel program execution. The proposed mechanism exploits data caches and cache coherence protocol states to minimize the number of trace events that are necessary to stream out of the target platform to the software debugger. The mechanism relies on a single trace bit per data cache block, thus minimizing the cost of hardware implementation. Our experimental evaluation explores the effectiveness of the proposed technique by measuring the trace port bandwidth as a function of the cache size and the number of processor cores. The results show that the proposed mechanism significantly reduces the required trace port bandwidth when compared to the Nexus-like load data value tracing. Depending on data cache size, the improvements range from 9.9 to 23.5 times for single cores and from 18.6 to 37.3 times for octa cores. Mounika Ponugoti, Aleksandar Milenkovic |
ICCD | 2 |
| 2015 | On effectiveness of lossless compression in transferring mHealth data filesabstractThe health and fitness data traffic originating on mobile devices has been continually increasing, with an exponential increase in the number of personal wearable devices and mobile health monitoring applications. Lossless data compression can increase throughput, reduce latency, and achieve energy-efficient communication between personal devices and the cloud. This paper experimentally explores the effectiveness of common compression utilities on mobile devices when uploading and downloading a representative mHealth data set. Based on the results of our study, we develop recommendations for effective data transfers that can assist mHealth application developers. Armen Dzhagaryan, Aleksandar Milenkovic |
HealthCom | 2 |
| 2015 | Smart Button: A wearable system for assessing mobility in elderlyabstractContinuous advances in sensors, semiconductors, wireless networks, mobile and cloud computing enable the development of integrated wearable computing systems for continuous health monitoring. These systems can be used as a part of diagnostic procedures, in the optimal maintenance of chronic conditions, in the monitoring of adherence to treatment guidelines, and for supervised recovery. In this paper, we describe a wearable system called Smart Button designed to assess mobility of elderly. The Smart Button is easily mounted on the chest of an individual and currently quantifies the Timed-Upand- Go and 30-Second Chair Stand tests. These two tests are routinely used to assess mobility, balance, strength of the lower extremities, and fall risk of elderly and people with Parkinson's disease. The paper describes the design of the Smart Button, parameters used to quantify the tests, signal processing used to extract the parameters, and integration of the Smart Button into a broader mHealth system. Armen Dzhagaryan, Aleksandar Milenkovic, Emil Jovanov, Mladen Milosevic |
HealthCom | 2 |
| 2015 | Quantifying Benefits of Lossless Compression Utilities on Modern SmartphonesabstractThe data traffic originating on mobile computing devices has been growing exponentially over the last several years. Lossless data compression and decompression can be essential in increasing communication throughput, reducing communication latency, achieving energy-efficient communication, and making effective use of available storage. This paper experimentally evaluates several compression utilities and configurations on a modern smartphone. We characterize each utility in terms of its compression ratio, compression and decompression throughput, and energy efficiency for representative use cases. We find a wide variety of energy costs associated with data compression and decompression and provide practical guidelines for selecting the most energy efficient configurations for each use case. For data transfers over WLAN, the best configurations provide a 2.1-fold and 2.7-fold improvement in energy efficiency for compressed uploads and downloads, respectively, when compared to uncompressed data transfers. For data transfers over a mobile broadband network, the best configurations provide a 2.7-fold and 3-fold improvement in energy efficiency for compressed uploads and downloads, respectively. Armen Dzhagaryan, Aleksandar Milenkovic, Martin Burtscher |
ICCCN | 2 |
| 2015 | mcfTRaptor: Toward unobtrusive on-the-fly control-flow tracing in multicores
Amrish K. Tewar, Albert R. Myers, Aleksandar Milenkovic |
J. Syst. Archit. | 3 |
| 2014 | Using Branch Predictors and Variable Encoding for On-the-Fly Program TracingabstractUnobtrusive capturing of program execution traces in real-time is crucial for debugging many embedded systems. However, tracing even limited program segments is often cost-prohibitive, requiring wide trace ports and large on-chip trace buffers. This paper introduces a new cost-effective technique for capturing and compressing program execution traces on-the-fly. It relies on branch predictor-like structures in the trace module and corresponding software modules in the debugger to significantly reduce the number of events that need to be streamed out of the target system. Coupled with an effective variable encoding scheme that adapts to changing program patterns, our technique requires merely 0.029 bits per instruction of trace port bandwidth, providing a 34-fold improvement over the commercial state-of-the-art and a five-fold improvement over academic proposals, at the low cost of under 5,000 logic gates. Vladimir Uzelac, Aleksandar Milenkovic, Milena Milenkovic, Martin Burtscher |
IEEE Trans. Computers | 2 |
| 2013 | Smartphones for smart wheelchairsabstractIndividuals with limited ambulatory skills are at high risk for all physical inactivity-related diseases, such as coronary disease and diabetes. Increased physical activity can significantly lower risks of these diseases. However, quantifying recommendations for increased physical activity remain challenging for individuals who use wheelchairs for mobility. In this paper we introduce a smart wheelchair that utilizes a smartphone with its built-in sensors to capture and record physical activity of manual wheelchair users in both unstructured and structured environments. We develop algorithms for data acquisition and processing on the smartphone and implement them in an Android application called mWheelness. The application is successfully tested in laboratory and free-living experiments using several modern smartphones. Aleksandar Milenkovic, Mladen Milosevic, Emil Jovanov |
BSN | 1 |
| 2013 | Quantifying Timed-Up-and-Go test: A smartphone implementationabstractTimed-Up-and-Go (TUG) is a simple, easy to administer, and frequently used test for assessing balance and mobility in elderly and people with Parkinson's disease. An instrumented version of the test (iTUG) has been recently introduced to better quantify subject's movements during the test. The subject is typically instrumented by a dedicated device designed to capture signals from inertial sensors that are later analyzed by healthcare professionals. In this paper we introduce a smartphone application called sTUG that completely automates the iTUG test so it can be performed at home. sTUG captures the subject's movements utilizing smartphone's built-in accelerometer and gyroscope sensors, determines the beginning and the end of the test and quantifies its individual phases, and optionally uploads test descriptors into a medical database. We describe the parameters used to quantify the iTUG test and algorithms to extract the parameters from signals captured by the smartphone sensors. Mladen Milosevic, Emil Jovanov, Aleksandar Milenkovic |
BSN | 3 |
| 2013 | A software model of mobile notification system for medication misuse preventionabstractThe one of the most common problems affecting older patients is the misuse of prescribed therapies. The patients, in many cases, forget to take their therapy, use it twice or even use the wrong medication. Since all of the mentioned cases can trigger further serious health problems, and since the majority of the population in Serbia already uses mobile phones, we decided to extend our existing medical information system with a prescription service that will allow the development of mobile-device based notification system. The notification system will then help patients with medication timing and dosing issues to take their therapies more accurate and to prevent possible health complications. In this paper we present the existing mobile services offered by the medical information system Medis.NET, as well as an architectural solution for the extension of the existing software with a prescription and notification service. The architectural solution is followed by the example of an initial implementation, description of the most important use-cases and related test results. In a near future, we plan to improve the proposed solution by offering new options to the patients, easing the connection with third party services and by developing a smartphone application that can improve the complete interaction with patients. Petar Rajkovic, Dragan Jankovic, Aleksandar Milenkovic |
Healthcom | 3 |
| 2013 | Energy efficiency of lossless data compression on a mobile device: An experimental evaluationabstractLossless compression and decompression are routinely used in mobile computing devices to reduce the costs of communicating and storing data. This paper presents the results of an experimental evaluation of common compression utilities on Pandaboard, a development platform similar to current commercial mobile devices. We study the compression ratio, compression and decompression throughput, and energy efficiency of different usage scenarios typical for mobile computing. We observe a wide variety of energy costs associated with data compression and provide practical guidelines for selecting the most energy-efficient configurations. Armen Dzhagaryan, Aleksandar Milenkovic, Martin Burtscher |
ISPASS | 2 |
| 2013 | Performance and Energy Consumption of Lossless Compression/Decompression Utilities on Mobile Computing PlatformsabstractData compression and decompression utilities can be critical in increasing communication throughput, reducing communication latencies, achieving energy-efficient communication, and making effective use of available storage. This paper experimentally evaluates several such utilities for multiple compression levels on systems that represent current mobile platforms. We characterize each utility in terms of its compression ratio, compression and decompression through-put, and energy efficiency. We consider different use cases that are typical for modern mobile environments. We find a wide variety of energy costs associated with data compression and decompression and provide practical guidelines for selecting the most energy efficient configurations for each use case. The best performing configurations provide 6-fold and 4-fold improvements in energy efficiency for compressed uploads and downloads over WLAN, respectively, when compared to uncompressed data transfers. Aleksandar Milenkovic, Armen Dzhagaryan, Martin Burtscher |
MASCOTS | 1 |
| 2013 | Hardware-Based Load Value Trace Filtering for On-the-Fly DebuggingabstractCapturing program and data traces during program execution unobtrusively on-the-fly is crucial in debugging and testing of cyber-physical systems. However, tracing a complete program unobtrusively is often cost-prohibitive, requiring large on-chip trace buffers and wide trace ports. This article describes a new hardware-based load data value filtering technique called Cache First-access Tracking. Coupled with an effective variable encoding scheme, this technique achieves a significant reduction of load data value traces, from 5.86 to 56.39 times depending on the data cache size, thus enabling cost-effective, unobtrusive on-the-fly tracing and debugging. Vladimir Uzelac, Aleksandar Milenkovic |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2011 | Caches and Predictors for Real-Time, Unobtrusive, and Cost-Effective Program Tracing in Embedded SystemsabstractThe increasing complexity of modern embedded computer systems makes software development and system verification the most critical steps in system development. To expedite verification and program debugging, chip manufacturers increasingly consider hardware infrastructure for program debugging and tracing, including logic to capture and filter traces, buffers to store traces, and a trace port through which the trace is read by the debug tools. In this paper, we introduce a new approach to capture and compress program execution traces in hardware. The proposed trace compressor encompasses two cost-effective structures, a stream descriptor cache, and a last stream predictor. Information about the program flow is translated into a sequence of hit and miss events in these structures, thus dramatically reducing the number of bits that need to be sent out of the chip. We evaluate the efficiency of the proposed mechanism by measuring the trace port bandwidth on a set of benchmark programs. Our mechanism requires only 0.15 bits/instruction/CPU on average on the trace port, which is a sixfold improvement over state-of-the-art commercial solutions. The trace compressor requires an on-chip area that is equivalent to one third of a 1 kilobyte cache and it allows for continual and unobtrusive program tracing in real time. Aleksandar Milenkovic, Vladimir Uzelac, Milena Milenkovic, Martin Burtscher |
IEEE Trans. Computers | 1 |
| 2010 | Hardware-based data value and address trace filtering techniquesabstractCapturing program and data traces during program execution unobtrusively in real-time is crucial in debugging and testing of cyber-physical systems. However, tracing a complete program unobtrusively is often cost-prohibitive, requiring large on-chip trace buffers and wide trace ports. Whereas program execution traces can be efficiently compressed in hardware, compression of data address and data value traces is much more challenging due to limited redundancy. In this paper we describe two hardware-based filtering techniques for data traces: cache first-access tracking for load data values and data address filtering using partial register-file replay. The results of our experimental analysis indicate that the proposed filtering techniques can significantly reduce the size of the data traces (~5 20 times for the load data value trace, depending on the data cache size; and ~5 times for the data address trace) at the cost of rather small hardware structures in the trace module. Vladimir Uzelac, Aleksandar Milenkovic |
CASES | 2 |
| 2010 | Real-time unobtrusive program execution trace compression using branch predictor eventsabstractUnobtrusive capturing of program execution traces in real-time is crucial in debugging cyber-physical systems. However, tracing even limited program segments is often cost-prohibitive, requiring wide trace ports and large on-chip trace buffers. This paper introduces a new cost-effective technique for capturing and compressing program execution traces in real time. It uses branch predictor-like structures in the trace module to losslessly compress the traces. This approach results in high compression ratios because it only has to transmit misprediction events to the software debugger. Coupled with an effective variable encoding scheme, our technique requires merely 0.036 bits/instruction of trace port bandwidth (a 28-fold improvement over the commercial state-of-the-art) at a cost of roughly 5,200 logic gates. Vladimir Uzelac, Aleksandar Milenkovic, Martin Burtscher, Milena Milenkovic |
CASES | 2 |
| 2009 | A real-time program trace compressor utilizing double move-to-front methodabstractThis paper introduces a new unobtrusive and cost-effective method for the capture and compression of program execution traces in real-time, which is based on a double move-to-front transformation. We explore its effectiveness and describe a cost-effective hardware implementation. The proposed trace compressor requires only 0.12 bits per instruction of trace port bandwidth, at the cost of 25K gates. Vladimir Uzelac, Aleksandar Milenkovic |
DAC | 2 |
| 2009 | Real-time, unobtrusive, and efficient program execution tracing with stream caches and last stream predictorsabstractThis paper introduces a new hardware mechanism for capturing and compressing program execution traces unobtrusively in real-time. The proposed mechanism is based on two structures called stream cache and last stream predictor. We explore the effectiveness of a trace module based on these structures and analyze the design space. We show that our trace module, with less than 600 bytes of state, achieves a trace-port bandwidth of 0.15 bits/instruction/processor, which is over six times better than state-of-the-art commercial designs. Vladimir Uzelac, Aleksandar Milenkovic, Milena Milenkovic, Martin Burtscher |
ICCD | 2 |
| 2009 | Experiment flows and microbenchmarks for reverse engineering of branch predictor structuresabstractInsights into branch predictor organization and operation can be used in architecture-aware compiler optimizations to improve program performance. Unfortunately, such details are rarely publicly disclosed. In this paper we introduce a set of experiment flows and corresponding microbenchmarks for reverse engineering cache-like branch target and outcome predictor structures, indexed by branch address or program path information. The experiment flows are demonstrated on the Intel Pentium M branch predictor. We have been able to determine the size, organization, internal operation, and interactions between various hardware structures used in the Pentium M branch predictor, namely the branch target buffer, indirect branch target buffer, loop branch predictor buffer, global predictor, and bimodal predictor. These findings have been validated using a functional PIN model. Vladimir Uzelac, Aleksandar Milenkovic |
ISPASS | 2 |
| 2007 | Algorithms and Hardware Structures for Unobtrusive Real-Time Compression of Instruction and Data Address TracesabstractInstruction and data address traces are widely used by computer designers for quantitative evaluations of new architectures and workload characterization, as well as by software developers for program optimization, performance tuning, and debugging. Such traces are typically very large and need to be compressed to reduce the storage, processing, and communication bandwidth requirements. However, preexisting general-purpose and trace-specific compression algorithms are designed for software implementation and are not suitable for runtime compression. Compressing program execution traces at runtime in hardware can deliver insights into the behavior of the system under test without any negative interference with normal program execution. Traditional debugging tools, on the other hand, have to stop the program frequently to examine the state of the processor. Moreover, software developers often do not have access to the entire history of computation that led to an erroneous state. In addition, stepping through a program is a tedious task and may interact with other system components in such a way that the original errors disappear, thus preventing any useful insight. The need for unobtrusive tracing is further underscored by the development of computer systems that feature multiple processing cores on a single chip. In this paper, we introduce a set of algorithms for compressing instruction and data address traces that can easily be implemented in an on-chip trace compression module and describe the corresponding hardware structures. The proposed algorithms are analytically and experimentally evaluated. Our results show that very small hardware structures suffice to achieve a compression ratio similar to that of a software implementation of gzip while being orders of magnitude faster. A hardware structure with slightly over 2 KB of state achieves a compression ratio of 125.9 for instruction address traces, whereas gzip achieves a compression ratio of 87.4. For data address traces, a hardware structure with 5 KB of state achieves a compression ratio of 6.1, compared to 6.8 achieved by gzip Milena Milenkovic, Aleksandar Milenkovic, Martin Burtscher |
DCC | 2 |
| 2007 | A low overhead hardware technique for software integrity and confidentialityabstractSoftware integrity and confidentiality play a central role in making embedded computer systems resilient to various malicious actions, such as software attacks; probing and tampering with buses, memory, and I/O devices; and reverse engineering. In this paper we describe an efficient hardware mechanism that protects software integrity and guarantees software confidentiality. To provide software integrity, each instruction block is signed during program installation with a cryptographically secure signature. The signatures embedded in the code are verified during program execution. Software confidentiality is provided by encrypting instruction blocks. To achieve low performance overhead, the proposed mechanism combines several architectural enhancements: a variation of one-time-pad encryption, parallelizable signatures, and conditional execution of unverified instructions. A relatively high memory overhead due to embedded signatures can be reduced by protecting multiple instruction blocks with one signature, with minimal effects on complexity and performance overhead. Austin Rogers, Milena Milenkovic, Aleksandar Milenkovic |
ICCD | 3 |
| 2007 | A 1GHz Direct Digital Frequency Synthesizer Based on the Quasi-Linear Interpolation MethodabstractThe paper presents a novel architecture for a direct digital frequency synthesizer (DDFS) based on the Quasi-Linear interpolation (QLIP) method. The four-segment QLIP is utilized to realize a DDFS with a spurious free dynamic range (SFDR) of 63.2dBc. The DDFS chip featuring a 5-stage pipeline is implemented in TSMC 0.13μm technology. The chip occupies 9874μm2, consumes 8.2μW/MHz, and runs at 1GHz clock rate. Ashkan Ashrafi, Aleksandar Milenkovic, Reza R. Adhami |
ISCAS | 2 |
| 2006 | Wireless sensor networks for personal health monitoring: Issues and an implementation
Aleksandar Milenkovic, Chris Otto, Emil Jovanov |
Comput. Commun. | 1 |
| 2005 | Hardware support for code integrity in embedded processorsabstractComputer security becomes increasingly important with continual growth of the number of interconnected computing platforms. Moreover, as capabilities of embedded processors increase, the applications running on these systems also grow in size and complexity, and so does the number of security vulnerabilities. Attacks that impair code integrity by injecting and executing malicious code are one of the major security issues. This problem can be addressed at different levels, from more secure software and operating systems, down to solutions that require hardware support. Most of the existing techniques tackle the problem of security flaws at the software level, but this approach lacks generality and often induces prohibitive overhead in performance and cost, or generates a significant number of false alarms. On the other hand, a further increase in the number of transistors on a single chip enables integrated hardware support for functions that formerly were restricted to the software domain. Hardware-supported defense techniques have the potential to be more general and more efficient than solely software solutions. This paper proposes four new architectural extensions to ensure complete run-time code integrity using instruction block signature verification. The experimental analysis shows that the proposed techniques have low performance and energy overhead. In addition, the proposed mechanism has low hardware complexity, and does not impose either changes to the compiler or changes to the existing instruction set architecture. Milena Milenkovic, Aleksandar Milenkovic, Emil Jovanov |
CASES | 2 |
| 2004 | Microbenchmarks for determining branch predictor organizationabstractAbstract In order to achieve an optimum performance of a given application on a given computer platform, a program developer or compiler must be aware of computer architecture parameters, including those related to branch predictors. Although dynamic branch predictors are designed with the aim of automatically adapting to changes in branch behavior during program execution, code optimizations based on the information about predictor structure can greatly increase overall program performance. Yet, exact predictor implementations are seldom made public, even though processor manuals provide valuable optimization tips. This paper presents an experimental flow with a series of microbenchmarks that determine the organization and size of a branch predictor using on‐chip performance monitoring registers. Such knowledge can be used either for manual code optimization or for design of new, more architecture‐aware compilers. Three examples illustrate how insight into exact branch predictor organization can be directly applied to code optimization. The proposed experimental flow is illustrated with microbenchmarks tuned for Intel Pentium III and Pentium 4 processors, although they can easily be adapted for other architectures. The described approach can also be used during processor design for performance evaluation of various branch predictor organizations and for testing and validation during implementation. Copyright © 2004 John Wiley & Sons, Ltd. Milena Milenkovic, Aleksandar Milenkovic, Jeffrey H. Kulick |
Softw. Pract. Exp. | 2 |
| 2000 | Cache Injection: A Novel Technique for Tolerating Memory Latency in Bus-Based SMPs
Aleksandar Milenkovic, Veljko M. Milutinovic |
Euro-Par | 1 |
| 2000 | Scowl: A Tool for Characterization of Parallel Workload and its Use on Splash-2 Application SuiteabstractConcentrates on the problem of defining and measuring parameters that characterize typical behavior of parallel applications targeted to distributed shared memory (DSM) systems and shared-memory multiprocessors (SMPs). These parameters can be used as input to various models for performance evaluation in this research area. Furthermore, typical application behaviors can be recognized, which can help to generate new ideas for improvements to memory consistency protocols, adapting them to specific application characteristics. Our study encompasses a variety of parameters, such as frequencies of operations of various access types (private read/writes, shared read/writes, lock operations, barrier operations), the average number of accessed blocks per interval, the average number of modified words, etc. The results presented in this paper are based on the SPLASH-2 (Stanford Parallel Applications for SHared Memory) application suite. The developed instrumentation tool Scowl, along with the applied simulation environment Limes, are publicly available and applicable for performing measurements on other parallel applications as well. Darko Marinov, Davor Magdic, Aleksandar Milenkovic, Jelica Protic, Igor Tartalja, Veljko M. Milutinovic |
MASCOTS | 3 |
| 1998 | Cache Injection on Bus Based MultiprocessorsabstractSoftware-controlled cache prefetching and data forwarding are widely used techniques for tolerating memory latency in shared memory multiprocessors. However, some previous studies show that cache prefetching is not so effective on bus-based multiprocessors, while the effectiveness of data forwarding has not been explored in this environment, yet. In this paper, a novel technique called cache injection is proposed. Cache injection, tuned to the properties of bus-based architectures, combines advantages of both cache prefetching and data forwarding. Some preliminary experiments show that the proposed solution can significantly help in reducing the overall miss ratio and bus traffic in applications where write-shared data prevails. Aleksandar Milenkovic, Veljko M. Milutinovic |
SRDS | 1 |
| 1997 | The Cache InjectionKofetch Architecture: Initial Performance EvaluationabstractOne of the major problems in a number of SM (shared memory) and DSM (distributed shared memory) applications is the overall cost of read misses in conditions when: (a) system latencies are relatively large, and (b) a shared data item is read relatively few times by each of the processors in the system; modern SM and DSM systems are typically based on off-the-shelf microprocessors which do not include any support for the described problem. Consequently, the major goal of our research is to come up with a new concept to be incorporated into the next generation microprocessors, so they can became more efficient in the sense described above. Existing 64-bit processors support only data prefetching (PF) as a method to fight against negative effects of the described problem. Our research introduces a new concept referred to as cache injection (CI), as well as the related cache injection/cofetch architecture (CICA). Initial performance evaluation is performed using a simulation methodology based on the set of synthetic benchmarks. Veljko M. Milutinovic, Aleksandar Milenkovic, Gad Sheaffer |
MASCOTS | 2 |