VLDB 2026 Research / reviewers in the wild / expert
Mantas Mikaitis
dblp:202/6177
· DBLP profile ↗
8ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0001-8706-1436ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MATLAB Simulator of Level-Index ArithmeticabstractLevel-index arithmetic appeared in the 1980s. One of its principal purposes is to abolish the issues caused by underflows and overflows in floating point. However, level-index arithmetic does not expand the set of numbers but spaces out the numbers of large magnitude even more than floating-point representations to move the infinities further away from zero: gaps between numbers on both ends of the range become very large. We revisit level index by presenting a custom precision simulator in MATLAB. This toolbox is useful for exploring performance of level-index arithmetic in research projects, such as using 8-bit and 16-bit representations in machine learning algorithms where narrow bit-width is desired but overflow/underflow of floating-point representations causes difficulties. Mantas Mikaitis |
ARITH | 1 |
| 2024 | Monotonicity of Multi-term Floating-Point AddersabstractIn the literature on algorithms for computing multi-term addition$s_n=\sum_{i=1}^n x_i$in floating-point arithmetic it is often shown that a hardware unit that has single normalization and rounding improves precision, area, latency, and power consumption, compared with the use of standard add or fused multiply–add units. However, non-monotonicity can appear when computing sums with a subclass of multi-term addition units, which is currently not explored in the literature. We prove that computing multi-term floating-point addition withn≥ 4, without normalization of intermediate quantities, can result in non-monotonicity—increasing one of the addendsxidecreases the sumsn. Summation is required in dot product and matrix multiplication operations, operations that are increasingly appearing in the hardware of high-performance computers, and knowing where monotonicity is preserved can be of interest to the developers and users. Non-monotonicity of summation in existent hardware devices that implement a specific class of multi-term adders may have appeared unintentionally as a consequence of design choices that reduce circuit area and other metrics. To demonstrate our findings we simulate non-monotonic multi-term adders in MATLAB using theCPFloatcustom-precision floating-point simulator. Mantas Mikaitis |
IEEE Trans. Computers | 1 |
| 2023 | CPFloat: A C Library for Simulating Low-precision ArithmeticabstractOne can simulate low-precision floating-point arithmetic via software by executing each arithmetic operation in hardware and then rounding the result to the desired number of significant bits. For IEEE-compliant formats, rounding requires only standard mathematical library functions, but handling subnormals, underflow, and overflow demands special attention, and numerical errors can cause mathematically correct formulae to behave incorrectly in finite arithmetic. Moreover, the ensuing implementations are not necessarily efficient, as the library functions these techniques build upon are typically designed to handle a broad range of cases and may not be optimized for the specific needs of rounding algorithms. CPFloat is a C library for simulating low-precision arithmetics. It offers efficient routines for rounding, performing mathematical computations, and querying properties of the simulated low-precision format. The software exploits the bit-level floating-point representation of the format in which the numbers are stored and replaces costly library calls with low-level bit manipulations and integer arithmetic. In numerical experiments, the new techniques bring a considerable speedup (typically one order of magnitude or more) over existing alternatives in C, C++, and MATLAB. To our knowledge, CPFloat is currently the most efficient and complete library for experimenting with custom low-precision floating-point arithmetic. Massimiliano Fasi, Mantas Mikaitis |
ACM Trans. Math. Softw. | 2 |
| 2021 | Algorithms for Stochastically Rounded Elementary Arithmetic Operations in IEEE 754 Floating-Point ArithmeticabstractPublished in "IEEE Transactions on Emerging Topics in Computing, Volume: 9, Issue: 3, JulySeptember 2021" and orally presented at ARITH 2021. Massimiliano Fasi, Mantas Mikaitis |
ARITH | 2 |
| 2021 | Stochastic Rounding: Algorithms and Hardware AcceleratorabstractWe present algorithms and a hardware accelerator for performing stochastic rounding (SR). Our main goal is to augment the ARM M4F-based multi-core processor SpiNNaker2 with a more flexible rounding functionality than is available in the ARM processor itself. The motivation of adding such functionality in hardware is based on our previous results showing improvements in numerical accuracy of ODE solvers in fixed-point arithmetic with SR, compared to a standard round to nearest mode (RN) or bit truncation. Performing SR purely in software can be expensive due to requirement of multiple masking and shifting instructions, and an addition operation per each rounding. Also, saturation of values is included since it is required on overflows, which is common in fixed-point arithmetic due to a narrow dynamic range. The main intended use of the accelerator is to round fixed-point multiplier outputs, which are returned unrounded by the ARM processor in a wider fixed-point format than the arguments. The proposed accelerator is not specific to SpiNNaker, and is a generally applicable rounding unit provided a pseudorandom number generator is available that can supply random bits to it. Additionally, to the best of our knowledge, this is a first exploration of a stochastic rounding accelerator with a programmable bit position and multiple data type support. Mantas Mikaitis |
IJCNN | 1 |
| 2020 | Issues with rounding in the GCC implementation of the ISO 18037: 2008 standard fixed-point arithmeticabstractWe describe various issues caused by the lack of round-to-nearest mode in the gcc compiler implementation of the fixed-point arithmetic data types and operations. We demonstrate that round-to-nearest is not performed in the conversion of constants, conversion from one numerical type to a less precise type and results of multiplications. Furthermore, we show that mixed-precision operations in fixed-point arithmetic lose precision on arguments, even before carrying out arithmetic operations. The ISO 18037:2008 standard was created to standardize C language extensions, including fixed-point arithmetic, for embedded systems. Embedded systems are usually based on ARM processors, of which approximately 100 billion have been manufactured by now. Therefore, the observations about numerical issues that we discuss in this paper can be rather dangerous and are important to address, given the wide ranging type of applications that these embedded systems are running. Mantas Mikaitis |
ARITH | 1 |
| 2018 | Approximate Fixed-Point Elementary Function Accelerator for the SpiNNaker-2 Neuromorphic ChipabstractNeuromorphic chips are used to model biologically inspired Spiking-Neural-Networks (SNNs) where most models are based on differential equations. Equations for most SNN algorithms usually contain variables with one or more excomponents. SpiNNaker is a digital neuromorphic chip that has so far been using pre-calculated look-up tables for exponential function. However this approach is limited because the memory requirements grow as more complex neural models are developed. To save already limited memory resources in the next generation SpiNNaker chip, we are including a fast exponential function in the silicon. In this paper we analyse iterative algorithms for elementary functions and show how to build a single hardware accelerator for exp and natural log, for a neuromorphic chip prototype, to be manufactured in a 22 nm FDSOI process. We present the accelerator that has algorithmic level approximation control, allowing it to trade precision for latency and energy efficiency. As an addition to neuromorphic chip application, we provide analysis of a parameterized elementary function unit that can be tailored for other systems with different power, area, accuracy and latency constraints. Mantas Mikaitis, David R. Lester, Delong Shang, Steve Furber, Gengting Liu, Jim D. Garside, Stefan Scholze, Sebastian Höppner, Andreas Dixius |
ARITH | 1 |
| 2017 | Brewing the first ever automatic memory management utility for SpiNNaker: Real-time garbage collection for STDP simulationsabstractFirst generation SpiNNaker chip uses ARM968, with highly limited internal memory space, as its core element. In simulations of learning algorithms, many biologically plausible learning rules require history traces of each neuron's activity to be stored. As a result, the history traces of neurons rapidly fill the internal memory space eventually reaching the limits of ARM968. To lower the possibility of memory overflow, we propose to introduce a memory management routine working in the background, which must respect the biological timing constraints of the SpiNNaker simulations. Real-time garbage collection is an automatic memory management technique that can satisfy these requirements. This study presents the first ever implementation of real-time garbage collector for SpiNNaker architecture and evaluates the performance, carefully considering the biological real-time constraints of the system. Mantas Mikaitis, David R. Lester |
IJCNN | 1 |