VLDB 2026 Research / reviewers in the wild / expert
Paolo Montuschi
dblp:68/2893
· DBLP profile ↗
63ranked-venue papers
17as first author
10since 2021 · last 2026
0000-0003-2563-2250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 12 first-author · 3 since 2021Theory of computation · 11 · 4 first-authorComputer networks · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Fusion of Face and Gait for Person Identification in Automotive ApplicationsabstractSmart and secure access to vehicles is a crucial aspect of the evolving automotive industry. This paper focuses on the development of an end-to-end multimodal biometric recognition framework that identifies people walking towards a vehicle from an RGB video feed. The framework is based on a deep-learning pipeline for person detection and tracking, face and gait feature extraction, and fusion of the two modalities at the score and feature level. Traditional face recognition systems can suffer from variations in lighting and occlusions. In order to deal with these issues, the proposed framework integrates face and gait features with the aim to enhance accuracy. The pipeline is modular, enabling seamless integration of new models for each step of person identification without the need for additional training. Baseline face and gait recognition models, as well as score- and feature-level fusion techniques are evaluated on subsets of the CASIA-A and CASIA-B datasets. Experimental results show that weighted mean score-level fusion significantly improves both Rank-1 accuracy and verification accuracy (TAR@FAR=10−5) over unimodal baselines. Overall, reported work provides insights into current limitations and suggests directions for future research about secure identity verification in vehicles. Federico Boscolo, Fabrizio Lamberti, Paolo Montuschi, Mario Testa |
IEEE Internet Things J. | 3 |
| 2026 | Runtime Feature Compression for Adaptive Keyword Spotting on Embedded SystemsabstractVoice user interfaces rely on keyword spotting (KWS) to detect wake-word commands, enabling low-power devices to switch from drowsy to active states and initiate more complex tasks. In embedded systems, KWS combines handcrafted acoustic features extraction with lightweight neural network classifiers to achieve accurate detection within strict resource constraints. Adapting KWS to time-varying energy budgets requires optimization strategies that operate at runtime. Most existing approaches adjust the complexity of the neural model but overlook that a substantial amount of latency, and thus energy consumption, is due to feature extraction, which remains unaffected by model scaling. This work introducesRuntime Feature Compression(RFC), a dynamic rescaling strategy that modulates the workload of the entire KWS pipeline. RFC promotes thehop-lengthparameter of the Short-Time Fourier Transform as a runtime control knob to adjust the number of time frames in speech features, allowing a single model to operate across multiple latency modes. To support this flexibility, we introduce two training-time techniques:HopAugment, a data augmentation scheme that exposes the model to variable hop lengths during training, andMasked Layers, which preserve consistent activation statistics during training and inference under compressed feature settings. Evaluations on four KWS datasets using the TC-ResNet model family show that RFC outperforms model scaling techniques, offering a wider range of latency-accuracy trade-offs. RFC achieves up to 31.8% lower latency without accuracy degradation, or up to 0.30% higher accuracy within equivalent latency bounds. That proves RFC improves adaptability in energy-constrained IoT speech interfaces. A set of ablation studies further demonstrates the robustness of RFC by evaluating the role of its training components, batching strategies, ability to preserve accuracy with a shared weight set, scalability across operating modes, and applicability to different model architectures. Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi |
IEEE Internet Things J. | 4 |
| 2026 | An OpenSSL Engine for Secure and Transparent Offloading of Cryptographic Operations to OP-TEE in Embedded IoT PlatformsabstractEmbedded IoT platforms are frequently deployed in hostile or physically exposed environments, where compromise of the operating system is a realistic threat. In conventional deployments, OpenSSL executes entirely in user space, leaving cryptographic keys and intermediate material potentially exposed in the presence of a compromised OS or privileged malware. This work presents a portable OpenSSL engine that enables secure and transparent offloading of cryptographic operations to OP-TEE, a software stack leveraging a Trusted Execution Environment (TEE) to confine cryptographic key material and security-critical computations within secure-world memory without modifying existing applications or altering the OpenSSL EVP interface. The proposed engine uses a GlobalPlatform Client API implementation as an underlying communication layer and supports both Copy-Based (CB) data transfer and a Shared-Memory (SM) mode, enabling performance tuning under the resource constraints typical of embedded IoT platforms. An experimental evaluation on an NXP i.MX7 industrial gateway shows that secure-world execution introduces bounded overhead dominated by REE–TEE transitions and memory-management operations. For bandwidth-intensive primitives such as SHA-256, SM mode improves throughput by approximately 10–15% over CB transfers. For symmetric encryption algorithms such as AES-256-CBC, SM communication provides a substantial throughput improvement of approximately 40%, indicating that data-movement overhead plays a significant role on embedded platforms. In contrast, asymmetric primitives (e.g. RSA-1024) remain largely insensitive to transfer optimisations because they operate on small, fixed-size operands, making the overall execution time dominated by invocation overheads and internal big-number computations rather than data movement. Tina Sayarmoafi, Francesco Barchi, Lorenzo Bottaccioli, Andrea Acquaviva, Edoardo Patti, Paolo Montuschi, Luca Barbierato |
IEEE Internet Things J. | 6 |
| 2026 | The inNuCE Research Infrastructure and the Neuromorphic MLOps for AIoT PrototypingabstractNeuromorphic computing promises significant improvements in latency and energy efficiency for machine intelligence at the edge. However, its adoption in the IoT domain is still limited by the heterogeneity of the HW, the immaturity of the toolchains, and the poor reproducibility of experiments. The present paper sets out the inNuCE RI, a two-pillar facility composed of a physical inNuCE Lab and a cloud-based inNuCE HPP. The purpose of the inNuCE RI is to enable developers to prototype, evaluate and compare neuromorphic and conventional end-to-end digital solutions. From a methodological perspective, we formalize the adaptation of MLOps to event-driven sensing and brain-inspired computation as NMLOps. We illustrate how inNuCE RI instantiates NMLOps through containerized toolchains orchestrated with Kubernetes and Slurm-managed heterogeneous resources (neuromorphic chips, FPGAs, GPUs, MCUs). The approach is analyzed on representative AIoT use cases, including HAR, Braille reading, event-based gesture recognition, Hi-Co semantization of memories, navigation tracking, and constraint satisfaction problems. The development of inNuCE RI has been driven by the need to facilitate the transition from prototype (in nuce) to engineered AIoT systems for lower entry barriers and enforce reproducibility. This paves the way for future system-of-systems engineering. Gianvito Urgese, Vittorio Fra, Andrea Pignata, Giuseppe Fanuli, Walter Gallego Gomez, Riccardo Pignari, Michelangelo Barocci, Benedetto Leto, Salvatore Tilocca, Nicola Cassetta, Paolo Montuschi, Enrico Macii |
IEEE Internet Things J. | 11 |
| 2025 | Dealing With Challenged IoT Networks in Hierarchical Federated LearningabstractFederated Learning has revolutionized the way in which mobile devices and IoT can share common knowledge in data analytics. However, some challenges arise when dealing with heterogeneous and challenged networks, especially in gradient synchronization. For example, some clients (referred to as stragglers) may take much longer to report their output than other nodes. Current solutions addressing the straggling problems either propose a distributed coordination (but introduce new synchronization issues) or deadline-based approaches to discard clients after a fixed deadline (but introduce the problem of determining a suitable deadline). To this end, we propose to set a dynamic deadline in which the central server selects the best IoT nodes via an online learning approach based on predicting the response time of each client. Moreover, to further mitigate synchronization and scalability issues, we also consider a hierarchical approach in which clients send model parameters to intermediate aggregation edge servers. Our results demonstrate that this approach can lower network overhead by 78% compared to the widely adopted FedAvg and 49% to the best alternative. At the same time, the model accuracy is preserved, and the training time in challenged networks is reduced by 52% w.r.t. FedAvg and 32% w.r.t. recent solutions. Alessio Sacco, Doriana Monaco, Guido Marchetto, Paolo Montuschi |
IEEE Internet Things J. | 4 |
| 2024 | Dynamic Decision Tree Ensembles for Energy-Efficient Inference on IoT Edge NodesabstractWith the increasing popularity of Internet of Things (IoT) devices, there is a growing need for energy-efficient machine learning (ML) models that can run on constrained edge nodes. Decision tree ensembles, such as random forests (RFs) and gradient boosting (GBTs), are particularly suited for this task, given their relatively low complexity compared to other alternatives. However, their inference time and energy costs are still significant for edge hardware. Given that said costs grow linearly with the ensemble size, this article proposes the use of dynamic ensembles, that adjust the number of executed trees based both on a latency/energy target and on the complexity of the processed input, to tradeoff computational cost and accuracy. We focus on deploying these algorithms on multicore low-power IoT devices, designing a tool that automatically converts a Python ensemble into optimized C code, and exploring several optimizations that account for the available parallelism and memory hierarchy. We extensively benchmark both static and dynamic RFs and GBTs on three state-of-the-art IoT-relevant data sets, using an 8-core ultralow-power System-on-Chip (SoC), GAP8, as the target platform. Thanks to the proposed early stopping mechanisms, we achieve an energy reduction of up to 37.9% with respect to static GBTs (8.82 uJ versus 14.20 uJ per inference) and 41.7% with respect to static RFs (2.86 uJ versus 4.90 uJ per inference), without losing accuracy compared to the static model. Francesco Daghero, Alessio Burrello, Enrico Macii, Paolo Montuschi, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 4 |
| 2024 | Automatic Layer Freezing for Communication Efficiency in Cross-Device Federated LearningabstractFederated learning (FL) is a collaborative machine learning paradigm where network-edge clients train a global model under the orchestration of a central server. Unlike traditional distributed learning, each participating client keeps its data locally, ensuring privacy protection by default. However, state-of-the-art FL implementations suffer from massive information exchange between clients and the server. This issue prevents the adoption in constrained environments, typical of the Internet of Things domain, where the communication bandwidth and the energy budget are severely limited. To achieve higher efficiency at scale, the future of FL calls for additional optimizations to reach high-quality learning capability with lower communication pressure. To address this challenge, we propose automatic layer freezing (ALF), an embedded mechanism that gradually drops a growing portion of the model out of the training and synchronization phases of the learning loop, reducing the volume of exchanged data with the central server. ALF monitors the evolution of model updates and identifies layers that have reached a stable representation, where further weight updates would have minimal impact on accuracy. By freezing these layers, ALF achieves substantial savings in communication bandwidth and energy consumption. The proposed implementation of the ALF mechanism is compatible with any FL strategy, requiring minimal effort and without interfering with existing optimizations. The extensive experiments conducted using a representative set of FL strategies applied to two image classification tasks show that ALF improves the communication efficiency of the baseline FL implementations, ensuring up to 83.91% of data volume savings with no or marginal losses of accuracy. Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi |
IEEE Internet Things J. | 5 |
| 2023 | Exact and Approximate Squarers for Error-Tolerant ApplicationsabstractApproximate computing is considered an innovative paradigm with wide applications to high performance and low power systems. These applications have relaxed requirements for accuracy, so they can tolerate errors in results and achieve high performance. In approximate computing, multipliers have been widely studied, but squarers (as similar schemes) have not received much attention. In this paper, an accurate squarer is designed based on a Radix-8 Booth-folding square algorithm to reduce the number of partial products and the depth of the partial product array. Several approximate squarers (R8AS1, R8AS2 and R8AS3) are proposed based on the exact squarer to reduce power and delay. Two approximate partial product generators are also designed to simplify the Radix-8 Booth square encoder in R8AS1 and R8AS2. In addition, approximate compressors with compensation are used in the partial product compression stage to reduce additional area and power consumption in R8AS3. Synthesis results for power, area, and delay at 28 nm CMOS technology are presented. Compared with designs in the technical literature with the same accuracy, the proposed 16-bit designs reduce the PDP by 37%; in general, the PDP is decreased by up to 51%. Finally, the proposed approximate squarers are implemented in a square-law detector as a communication application and achieve an SNR close to 30 dB. Also, the three proposed approximate squarers are applied to the k-means clustering algorithm for machine learning to accomplish high performance in classification. Ke Chen 0018, Chenyu Xu, Haroon Waris, Weiqiang Liu 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 5 |
| 2023 | Tolerance of Siamese Networks (SNs) to Memory Errors: Analysis and DesignabstractThis article considers memory errors in a Siamese Network (SN) through an extensive analysis and proposes two schemes (using a weight filter and a code) to provide efficient hardware solutions for error tolerance. Initially the impact of memory errors on the weights of the SN (stored as floating-point (FP) numbers) is analyzed; this shows that the degradation is mostly caused by outliers in weights. Two schemes are subsequently proposed. An analysis is pursued to establish the filter's bounds selection by the maximum/minimum values of the weight distributions, by which outliers can be removed from the operation of the SN. A code scheme for protecting the sign and exponent bits of each weight in an FP number, is also proposed; this code incurs in no memory overhead by utilizing the 4 least significant bits (LSB) to store parity bits. Simulation shows that the filter has a better performance for multi-bit errors correction (a reduction of 95.288% in changed predictions), while the code achieves superior results in single-bit errors correction (a reduction of 99.775% in changed predictions). The combined method that uses the two proposed schemes, retains their advantages, so adaptive to all scenarios; The ASIC-based FP designs of the SN using serial and hybrid implementations are also presented; these pipelined designs utilize a novel multi-layer perceptron (MLP) (as branch networks of the SN) that operates at a frequency of 681.2 MHz (at a 32nm technology node), so significantly higher than existing designs found in the technical literature. The proposed error-tolerant approaches also show advantages in overheads comparing with for example traditional error correction code (ECC). These error-tolerant MLP-based designs are well suited to hardware/power-constrained platforms. Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 5 |
| 2021 | Less-is-Better Protection (LBP) for memory errors in kNNs classifiers
Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi |
Future Gener. Comput. Syst. | 3 |
| 2020 | Security in Approximate Computing and Approximate Computing for Security: Challenges and OpportunitiesabstractApproximate computing is an advanced computational technique that trades the accuracy of computation results for better utilization of system resources. It has emerged as a new preferable paradigm over traditional computing architectures for many applications where inaccurate results are acceptable. However, approximate computing also introduces security vulnerabilities mainly due to the fact that the uncertain and unpredictable intrinsic errors during approximate execution may be indistinguishable from malicious modification of the input data, the execution process, and the results. On the other hand, interestingly, approximate computing presents new opportunities to secure the system and the computation. Existing work on the security of approximate computing covers threat models, countermeasures, and evaluations but lacks a framework for analysis and comparison. In this article, we provide a classification of the state-of-the-art works in this research field, including threat models in approximate computing and promising security approaches using approximate computing. Open questions and potential future research directions are also discussed. Weiqiang Liu 0001, Chongyan Gu, Máire O'Neill, Gang Qu 0001, Paolo Montuschi, Fabrizio Lombardi |
Proc. IEEE | 5 |
| 2020 | Modeling and Simulation of Cyber-Physical Electrical Energy Systems With SystemC-AMSabstractModern cyber-physical electrical energy systems (CPEES) are characterized by wider adoption of sustainable energy sources and by an increased attention to optimization, with the goal of reducing pollution and wastes. This imposes a need for instruments supporting the design flow, to simulate and validate the behavior of system components and to apply additional optimization and exploration steps. Additionally, each system might be tested with a number of management policies, to evaluate their economic impact. It is thus evident that simulation is a key ingredient in the design flow of CPEES. This paper proposes a framework for CPEES modeling and simulation, that relies on the open-source standard SystemC-AMS. The paper formalizes the information and energy flow in a generic CPEES, by focusing on both AC and DC components, and by including support for mechanical and physical models that represent multiple energy sources and loads. Experimental results, applied to a complex CPEES case study, will prove the effectiveness of the proposed solution, in terms of accuracy, speed up w.r.t. the current state-of-the-art Matlab/Simulink, and support for the design flow. Yukai Chen, Sara Vinco, Daniele Jahier Pagliari, Paolo Montuschi, Enrico Macii, Massimo Poncino |
IEEE Trans. Sustain. Comput. | 4 |
| 2019 | Thank-You State of the Journal Editorial by the Outgoing Editor-in-ChiefabstractPresents comments from the output Editor-in-Chief for this issue of the publication. Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2018 | Combining Restoring Array and Logarithmic Dividers into an Approximate Hybrid DesignabstractThis paper proposes a new design of an approximate hybrid divider (AXHD), which combines the restoring array and the logarithmic dividers to achieve an excellent tradeoff between accuracy and hardware performance. Exact restoring divider cells (EXDCrs) are used to generate the MSBs of the quotient for attaining a high accuracy; the other quotient digits are processed by a logarithmic divider as inexact scheme to improve figures of merit such as power consumption, area and delay. The proposed AXHD is evaluated and analyzed using error and hardware metrics. The proposed design is also compared with the exact restoring divider (EXDr) and previous approximate restoring dividers (AXDrs). The results show that the proposed design achieves very good performance in terms of accuracy and hardware; case studies for image processing also show the validity of the proposed designs. Weiqiang Liu 0001, Jing Li 0117, Chenghua Wang, Paolo Montuschi, Fabrizio Lombardi |
ARITH | 5 |
| 2018 | Design and Application of an Approximate 2-D Convolver with Error CompensationabstractThis paper proposes an error compensation scheme of two-dimensional (2D) convolver in which both approximate circuit- and algorithm-level techniques are utilized in the design. Truncation and voltage scaling are used as circuit techniques, while bit-width reduction is utilized at the algorithm level. These different techniques are related to the configuration of the convolver by which its operation can be configured to meet different and often contrasting figures of merit. An extensive evaluation of different error metrics is performed. An error analysis is also presented to substantiate the simulation results; an error compensation scheme is introduced to remedy a loss of accuracy in computation. Convolution for image processing is treated in detail to show the effectiveness of the proposed approach. The design, the analysis and the simulation results show that the approximate techniques utilized in the inexact convolver can operate in synergy. Ke Chen 0018, Jie Han 0001, Paolo Montuschi, Weiqiang Liu 0001, Fabrizio Lombardi |
ISCAS | 3 |
| 2018 | State of the JournalabstractOverviews current society and chapter news and events. Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2018 | Virtual Character Animation Based on Affordable Motion Capture and Reconfigurable Tangible InterfacesabstractSoftware for computer animation is generally characterized by a steep learning curve, due to the entanglement of both sophisticated techniques and interaction methods required to control 3D geometries. This paper proposes a tool designed to support computer animation production processes by leveraging the affordances offered by articulated tangible user interfaces and motion capture retargeting solutions. To this aim, orientations of an instrumented prop are recorded together with animator's motion in the 3D space and used to quickly pose characters in the virtual environment. High-level functionalities of the animation software are made accessible via a speech interface, thus letting the user control the animation pipeline via voice commands while focusing on his or her hands and body motion. The proposed solution exploits both off-the-shelf hardware components (like the Lego Mindstorms EV3 bricks and the Microsoft Kinect, used for building the tangible device and tracking animator's skeleton) and free open-source software (like the Blender animation tool), thus representing an interesting solution also for beginners approaching the world of digital animation for the first time. Experimental results in different usage scenarios show the benefits offered by the designed interaction strategy with respect to a mouse & keyboard-based interface both for expert and non-expert users. Fabrizio Lamberti, Gianluca Paravati, Valentina Gatteschi, Alberto Cannavò, Paolo Montuschi |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | Design of Approximate High-Radix Dividers by Inexact Binary Signed-Digit AdditionabstractApproximate high radix dividers (HR-AXDs) are proposed and investigated in this paper. High-radix division is reviewed and inexact computing is introduced at different levels. Design parameters such as number of bits (N) and radix (r) are considered in the analysis; the replacement schemes with inexact cells and truncation schemes of exact cells in the binary signed-digit adder array is introduced. Circuit-level performance and the error characteristics of the inexact high radix dividers are analyzed for the proposed designs. The combined assessment of the normal error distance, power dissipation and delay is investigated and applications of approximate high-radix dividers are treated in detail. The simulation results show that the proposed approximate dividers offer extensive saving in terms of power dissipation, circuit complexity and delay, while only incurring in a small degradation in accuracy thus making them possibly suitable and interesting to some applications and domains such as low power/mobile computing. Linbin Chen, Fabrizio Lombardi, Paolo Montuschi, Jie Han 0001, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | State of the JournalabstractPresents an editorial on the curent status of this publication journal. Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2016 | State of the JournalabstractThe first year of the Editor-in-Chief's (EiC's) term has passed and the IEEE Transactions on Computers has strengthened its reputation and consolidated its role as the flagship Transactions of the Computer Society. The EiC offers thanks to all readers and supporters of the Journal, the members of the Editorial Board, our contributors and our reviewers. An additional acknowledgment for her day-by-day support in all operations, always with a smile, is owed to the TC Assistant, Ms. Natalie Cicero, and to all Computer Society Staff in support to our Journal. As a part of the natural turnaround, some Colleagues have left the editorial board and others have joined. On behalf of the Communities served by TC the EiC personally thanks the following Colleagues for their thoughtful service: Habib Ammari, Abhishek Chandra, Jinjun Chen, Eui-Young Chung, Amr El Abbadi, Vincenzo Eramo, Tian He, Michael Hsiao, Samee Khan, Hai Jin, Dan Li, Keqiu Li, Yingshu Li, Cristina Nita-Rotaru, Yi Pan, Manish Parashar, Meikang Qiu, Sanjay Ranka, Berk Sunar, Zahir Tari, Mateo Valero, Bharadwaj Veeravalli, Julio Villalba, Laurence Yang. At the same time, the EiC welcomes the new members of the editorial board. Their short biographies are found herein. Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2016 | State of the JournalabstractDiscusses the current state of the journal, reports on current and future areas of exploration and research, and presents new editors. Paolo Montuschi, Edward J. McCluskey, Samarjit Chakraborty, Jason Cong, Ramón M. Rodríguez-Dagnino, Fred Douglis, Lieven Eeckhout, Gernot Heiser, Sushil Jajodia, Ruby B. Lee, Dinesh Manocha, Tomás F. Pena, Isabelle Puaut, Hanan Samet, Donatella Sciuto |
IEEE Trans. Computers | 1 |
| 2015 | Unequal Error Protection of Memories in LDPC DecodersabstractMemories are one of the most critical components of many systems: due to exposure to energetic particles, fabrication defects and aging they are subject to various kinds of permanent and transient errors. In this scenario, Unequal error protection (UEP) techniques have been proposed in the past to encode stored information, allowing to detect and possibly recover from errors during load operations, while offering different levels of protection to partitions of codewords according to their importance. Low-density parity-check (LDPC) codes are used in many communication standards to encode the transmitted information: at reception, LDPC decoders heavily rely on memories to store and correct the received information. To ensure efficient and reliable decoding of information, the need to protect the memories used in LDPC decoders is of primary importance. In this paper we present a study on how to efficiently design UEP techniques for LDPC decoder memories. The devised UEP method is divided in four adjustable levels, each one offering a different degree of protection. The full UEP, along with simplified versions, has been implemented within an existing decoder and its area occupation and power consumption evaluated. Comparison with the literature on the subject shows an unmatched level of protection from errors at a small complexity and energy cost. Carlo Condo, Guido Masera, Paolo Montuschi |
IEEE Trans. Computers | 3 |
| 2015 | Design and Analysis of Approximate Compressors for MultiplicationabstractInexact (or approximate) computing is an attractive paradigm for digital processing at nanometric scales. Inexact computing is particularly interesting for computer arithmetic designs. This paper deals with the analysis and design of two new approximate 4-2 compressors for utilization in a multiplier. These designs rely on different features of compression, such that imprecision in computation (as measured by the error rate and the so-called normalized error distance) can meet with respect to circuit-based figures of merit of a design (number of transistors, delay and power consumption). Four different schemes for utilizing the proposed approximate compressors are proposed and analyzed for a Dadda multiplier. Extensive simulation results are provided and an application of the approximate multipliers to image processing is presented. The results show that the proposed designs accomplish significant reductions in power dissipation, delay and transistor count compared to an exact design; moreover, two of the proposed multiplier designs provide excellent capabilities for image multiplication with respect to average normalized error distance and peak signal-to-noise ratio (more than 50 dB for the considered image examples). Amir Momeni, Jie Han 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2015 | Editorial from the New Editor in ChiefabstractPresents the introductory editorial for this issue of the publication Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2012 | An Algorithmic and Architectural Study on Montgomery Exponentiation in RNSabstractThe modular exponentiation on large numbers is computationally intensive. An effective way for performing this operation consists in using Montgomery exponentiation in the Residue Number System (RNS). This paper presents an algorithmic and architectural study of such exponentiation approach. From the algorithmic point of view, new and state-of-the-art opportunities that come from the reorganization of operations and precomputations are considered. From the architectural perspective, the design opportunities offered by well-known computer arithmetic techniques are studied, with the aim of developing an efficient arithmetic cell architecture. Furthermore, since the use of efficient RNS bases with a low Hamming weight are being considered with ever more interest, four additional cell architectures specifically tailored to these bases are developed and the tradeoff between benefits and drawbacks is carefully explored. An overall comparison among all the considered algorithmic approaches and cell architectures is presented, with the aim of providing the reader with an extensive overview of the Montgomery exponentiation opportunities in RNS. Filippo Gandino, Fabrizio Lamberti, Gianluca Paravati, Jean-Claude Bajard, Paolo Montuschi |
IEEE Trans. Computers | 5 |
| 2011 | A General Approach for Improving RNS Montgomery Exponentiation Using Pre-processingabstractThe hardware implementation of modular exponentiation for very large integers is a well-known topic in digital arithmetic. An effective approach for obtaining parallel and carry-free implementations consists in using the Montgomery exponentiation algorithm and executing the necessary operations in RNS. Two efficient methods for performing the RNS Montgomery exponentiation have been proposed by Kawamura et al. and by Bajard and Imbert. The above approaches mainly differ in the algorithm used for implementing the base extension. This paper presents a modified RNS Montgomery exponentiation algorithm, where several multiplications are moved outside the main execution loop and replaced by an effective pre-processing stage producing a significant saving on the overall delay with respect to state-of-the-art approaches. Since the proposed modification should be applied to both of the above algorithms, two versions are specifically discussed. Filippo Gandino, Fabrizio Lamberti, Paolo Montuschi, Jean-Claude Bajard |
IEEE Symposium on Computer Arithmetic | 3 |
| 2011 | Reducing the Computation Time in (Short Bit-Width) Two's Complement MultipliersabstractTwo's complement multipliers are important for a wide range of applications. In this paper, we present a technique to reduce by one row the maximum height of the partial product array generated by a radix-4 Modified Booth Encoded multiplier, without any increase in the delay of the partial product generation stage. This reduction may allow for a faster compression of the partial product array and regular layouts. This technique is of particular interest in all multiplier designs, but especially in short bit-width two's complement multipliers for high-performance embedded cores. The proposed method is general and can be extended to higher radix encodings, as well as to any size square and m \times n rectangular multipliers. We evaluated the proposed approach by comparison with some other possible solutions; the results based on a rough theoretical analysis and on logic synthesis showed its efficiency in terms of both area and delay. Fabrizio Lamberti, Nikolaos Andrikos, Elisardo Antelo, Paolo Montuschi |
IEEE Trans. Computers | 4 |
| 2010 | Improved Design of High-Performance Parallel Decimal MultipliersabstractThe new generation of high-performance decimal floating-point units (DFUs) is demanding efficient implementations of parallel decimal multipliers. In this paper, we describe the architectures of two parallel decimal multipliers. The parallel generation of partial products is performed using signed-digit radix-10 or radix-5 recodings of the multiplier and a simplified set of multiplicand multiples. The reduction of partial products is implemented in a tree structure based on a decimal multioperand carry-save addition algorithm that uses unconventional (non BCD) decimal-coded number systems. We further detail these techniques and present the new improvements to reduce the latency of the previous designs, which include: optimized digit recoders for the generation of 2n-tuples (and 5-tuples), decimal carry-save adders (CSAs) combining different decimal-coded operands, and carry-free adders implemented by special designed bit counters. Moreover, we detail a design methodology that combines all these techniques to obtain efficient reduction trees with different area and delay trade-offs for any number of partial products generated. Evaluation results for 16-digit operands show that the proposed architectures have interesting area-delay figures compared to conventional Booth radix-4 and radix--8 parallel binary multipliers and outperform the figures of previous alternatives for decimal multiplication. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
IEEE Trans. Computers | 3 |
| 2009 | Guest Editors' Introduction: Special Section on Computer Arithmetic
Peter Kornerup, Paolo Montuschi, Jean-Michel Muller, Eric Schwarz |
IEEE Trans. Computers | 2 |
| 2008 | A Radix-2 Digit-by-Digit Architecture for Cube RootabstractA radix-2 digit-recurrence algorithm and architecture for the computation of the cube root are presented in this paper. The original recurrence based on the concept of completing the cube is modified to allow an efficient implementation of the algorithm, and the cycle time and area cost of the resulting architecture are estimated as 7.5 times the delay of a full adder and around 9000 $nand2$ cells, respectively, for double-precision computations. Alex Piñeiro, Javier D. Bruguera, Fabrizio Lamberti, Paolo Montuschi |
IEEE Trans. Computers | 4 |
| 2007 | A New Family of High.Performance Parallel Decimal MultipliersabstractThis paper introduces two novel architectures for parallel decimal multipliers. Our multipliers are based on a new algorithm for decimal carry-save multioperand addition that uses a novel BCD-4221 recoding for decimal digits. It significantly improves the area and latency of the partial product reduction tree with respect to previous proposals. We also present three schemes for fast and efficient generation of partial products in parallel. The recoding of the BCD-8421 multiplier operand into minimally redundant signed-digit radix-10, radix-4 and radix-5 representations using new recoders reduces the complexity of partial product generation. In addition, SD radix-4 and radix-5 recodings allow the reuse of a conventional parallel binary radix-4 multiplier to perform combined binary/decimal multiplications. Evaluation results show that the proposed architectures have interesting area-delay figures compared to conventional Booth radix-4 and radix-8 parallel binary multipliers and other representative alternatives for decimal multiplication. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 3 |
| 2007 | A radix-10 SRT divider based on alternative BCD codingsabstractIn this paper we present the algorithm and architecture a radix-10 floating-point divider based on an SRT non-restoring digit-by-digit algorithm. The algorithm uses conventional techniques developed to speed-up radix-2kdivision such as signed-digit (SD) redundant quotient and digit selection by constant comparison using a carry-save estimate of the partial remainder. To optimize area and latency for decimal, we include novel features such as the use of alternative BCD codings to represent decimal operands, estimates by truncation at any binary position inside a decimal digit, a single customized fast carry propagate decimal adder for partial remainder computation, initial odd multiple generation and final normalization with rounding, and register placement to exploit advanced high fanin mux-latch circuits. The rough area-delay estimations performed show that the proposed divider has a similar latency but less hardware complexity (1.3 area ratio) than a recently published high performance digit-by-digit implementation. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
ICCD | 3 |
| 2007 | A Digit-by-Digit Algorithm for mth Root ExtractionabstractA general digit-recurrence algorithm for the computation of the mth root (with an m integer) is presented in this paper. Based on the concept of completing the mth root, a detailed analysis of the convergence conditions is performed and iteration- independent digit-selection rules are obtained for any radix and redundant digit set. A radix-2 version for mth rooting is also studied, together with closed formulas for both the digit selection rules and the number of bits required to perform correct selections. Paolo Montuschi, Javier D. Bruguera, Luigi Ciminiera, José-Alejandro Piñeiro |
IEEE Trans. Computers | 1 |
| 2005 | Low Latency Digit-Recurrence Reciprocal and Square-Root Reciprocal Algorithm and ArchitectureabstractThe reciprocal and square-root reciprocal operations are important in several applications. For these operations, we present algorithms that combine a digit-by-digit module and one iteration of a quadratic-convergence approximation. The latter is implemented by a digit-recurrence, which uses the digits produced by the digit-by-digit part. In this way, both parts execute in an overlapped manner, so that the total number of cycles is about half of the number that would be required by the digit-by-digit part alone. Because of the approximation, correct rounding of the result cannot be obtained directly in all cases; we propose a variable-time implementation that produces the correctly rounded result with a small average overhead. Radix-4 implementations are described and have been synthesized. They achieve the same cycle time as the standard digit-by-digit implementation, resulting in a speed-up of about 2 and, because of the approximation part, the area factor is also about 2. We also show a combined implementation for both operations that has essentially the same complexity as that for square-root reciprocal alone. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Symposium on Computer Arithmetic | 3 |
| 2005 | Digit-Recurrence Dividers with Reduced Logical DepthabstractIn this paper, we propose a class of division algorithms with the aim of reducing the delay of the selection of the quotient digit by introducing more concurrency and flexibility in its computation. From the proposed class of algorithms, we select one that moves part of the selection function out of the critical path, with a corresponding reduction in the critical path compared with existing alternatives: we present the algorithm and describe the architectures for radix 4 and for radix 16. For radix 16, we use the scheme of overlapping two radix-4 stages. In both cases, radix 4 and radix 16, we show that our algorithms allow the design of units with well-balanced critical paths with consequent decreases of the cycle times. Moreover, in the radix-16 case, we include some additional speculation techniques. To estimate the speedup, we used a rough timing model based on logical effort. For both radices, we estimate a speedup of about 25 percent with respect to previous implementations. In the radix-4 case, this is achieved by using roughly the same area, while, in the radix-16 case, the area is increased by about 30 percent. We verified our estimations by performing a synthesis of the radix-4 units. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Trans. Computers | 3 |
| 2002 | Fast Radix-4 Retimed Division with Selection by ComparisonsabstractSince a large portion of the critical path in an implementation of radix-4 division corresponds to the delay of the quotient-digit selection module, it is of interest to reduce this delay. The proposal of this paper extends the approach presented recently of prestoring the selection constants corresponding to the actual value of the divisor and to perform the determination of the quotient digit by carry-free subtraction and sign detection. This extension consists in advancing the subtraction so that it is outside of the critical path. This advancement also provides the possibility of placing the registers so as to minimize the cycle time. We present the method and report results of synthesis using a family of standard cells. We conclude that the extension results in a speedup of 1.35 with respect to the basic implementation and of 1.3 with respect to the previously mentioned approach. We estimate that the areas of all three units are about the same. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
ASAP | 3 |
| 2001 | Visualizing vector fields: the thick oriented stream-line algorithm (TOSL)
Andrea Sanna, Bartolomeo Montrucchio, Paolo Montuschi, Amelia Carolina Sparavigna |
Comput. Graph. | 3 |
| 2001 | Boosting Very-High Radix Division with Prescaling and Selection by RoundingabstractAn extension of the very-high radix division with prescaling and selection by rounding is presented. This extension consists of increasing the effective radix of the implementation by obtaining a few additional bits of the quotient per iteration, without increasing the complexity of the unit to obtain the prescaling factor or the delay of an iteration. As a consequence, for some values of the effective radix, it permits an implementation with a smaller area and the same execution time of the original scheme. Details of the algorithm and the implementation are presented. Estimations of the execution time and area are given for 54 bit and 114 bit quotients and compared with those of other division units. Paolo Montuschi, Tomás Lang |
IEEE Trans. Computers | 1 |
| 1999 | Boosting Very-High Radix Division with Prescaling and Selection by RoundingabstractAn extension of the very-high radix division with prescaling and selection by rounding is presented. This extension consists in increasing the effective radix of the implementation by obtaining a few additional bits of the quotient per iteration, without increasing the complexity of the unit to obtain the prescaling factor nor the delay of an iteration. As a consequence, for some values of the effective radix, it permits an implementation with a smaller area and the same execution time than the original scheme. Estimations are given for 54-bit and 114-bit quotients. Paolo Montuschi, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 1 |
| 1999 | Very High Radix Square Root with Prescaling and Rounding and a Combined Division/Square Root UnitabstractAn algorithm for square root with prescaling and selection by rounding is developed and combined with a similar scheme for division. Since division is usually more frequent than square root, the main concern of the combined implementation is to maintain the low execution time of division, while accepting a somewhat larger execution time for square root. The algorithm is presented in detail, including the mathematical development of bounds for the first square-root digit and for the scaling factor. The proposed implementation is described, evaluated and compared with other combined div/sqrt units. The comparisons show that the proposed scheme potentially produces a significant speed-up for division, whereas, for square root, the speed-up is small. Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 1998 | A Flexible Algorithm for Multiprocessor Ray TracingabstractRay tracing programs are widely used to generate photo-realistic images, but the high computation time may discourage their implementation on single-processor machines; moreover, cost reduction of multi-processor general purpose architectures makes parallel rendering an attractive field of research. We propose a new algorithm which addresses the main issues of a parallel implementation of ray tracing on a message-passing-based machine. We adopt efficient strategies for dynamic workload distribution among processors, task synchronization and communication delay reduction. The resulting implementation is highly flexible, since any number of processors can be employed without introducing synchronization problems. We show an implementation of our algorithm on an nCUBE 2 supercomputer which is a general purpose parallel architecture with distributed memory. A theoretical evaluation of our algorithm allows us to identify a decreasing function for rendering times; the considered examples confirm the theoretical expectations showing that the efficiency of our system may reach up to the 91% of the value achievable by dividing the sequential rendering time by the number of processors employed. Andrea Sanna, Paolo Montuschi, Massimo Rossi |
Comput. J. | 2 |
| 1998 | An efficient algorithm for ray casting of CSG animation framesabstractThis paper presents a new algorithm to generate ray-cast CSG animation frames. We consider sequences of frames where only the objects can move; in this way, we take advantage of the high screen area coherence of this kind of animation. A new definition of bounding box allows us to reduce the number of pixels to be computed for the frames after the first. We associate a CSG subtree and two new flags, denoting if the box has changed in the current frame and if it will change in the next frame, with each box. We show with three examples the advantages of our technique when compared with an algorithm which entirely renders each frame of an animation. Intersections with CSG objects may be reduced to about one-fifth, while the rendering may be computed up to four times faster for the test sequences. © 1998 John Wiley & Sons, Ltd. Andrea Sanna, Paolo Montuschi |
Comput. Animat. Virtual Worlds | 2 |
| 1997 | A Q-Coder Algorithm with Carry Free AdditionabstractThe Q-Coder algorithm is a very efficient compression technique for bi-level images based on the arithmetic coding. The paper presents a new and fast version of the Q-Coder algorithm in which the carry-propagated adders have been replaced by carry-save adders. In this way, all the additions can be performed with a delay time of a single fill adder, independently of the length of the operands. The compression method is faster than the traditional Q-Coder algorithm with an almost unnoticeable increasing of the hardware requirements. Gianluca Cena, Paolo Montuschi, Luigi Ciminiera, Andrea Sanna |
IEEE Symposium on Computer Arithmetic | 2 |
| 1997 | A New Algorithm for the Rendering of CSG ScenesabstractThe generation of 3-D solid objects, and more generally solid geometric modelling, is very important in Computer Aided Design (CAD). An important role is played by the Constructive Solid Geometry (CSG) representation scheme. IN CSG, objects are described by trees of Boolean operations on half-spaces or boundaries of primitive solids. The study of techniques to speed up the rendering of scenes modelled with the CSG scheme is an attractive field of research; in this paper we propose a new algorithm which reduces the computational complexity for ray casting approaches. Our strategy identifies a set of areas on the plane of view where the rays starting from the observer have to be traced; for each zone, only a portion of the entire CSG tree has to be considered for intersection tests, instead of the whole database of the primitive objects. A comparison of our algorithm with a ray caster that adopts bounding volume hierarchies and with a freeware ray tracer called POV-Ray shows that, for the examples considered, we may reduce the intersection tests to one third of those performed when standard optimizations are adopted. Andrea Sanna, Paolo Montuschi, Antonio Fisone, Bartolomeo Montrucchio |
Comput. J. | 2 |
| 1996 | Carry-Save Multiplication Schemes without Final AdditionabstractCarry-save multipliers require an adder at the last step to convert the carry-sum representation of the most significant half of the result into a non-redundant form. This paper presents n/spl times/n multiplication schemes where this conversion is performed with a circuit operating in parallel with the carry-save array. The most relevant feature of the proposed multipliers is that the full 2n-bit result is produced, unlike similar multiplication schemes presented in the literature. Luigi Ciminiera, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 1995 | Very-high radix combined division and square root with prescaling and selection by roundingabstractAn algorithm for square root with prescaling is developed and combined with a similar scheme for division. An implementation is described, evaluated and compared with other combined div/sqrt implementations.> Tomás Lang, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 2 |
| 1995 | A Remark on "Reducing Iteration Time when Result Digit is Zero for Radix-2 SRT Division and Square Root with Redundant Remainders"abstractIn a previous paper by P. Montuschi and L. Ciminiera (ibid., vol. 42, no.2 p239-246, Feb 1993), an architecture for shared radix 2 division and square root has been presented whose main characteristic is the ability to avoid any addition/subtraction, when the digit 0 has been selected. Here, we emphasize the characteristics of the digit selection mechanism used by Montuschi and Ciminiera by presenting a small modification of the digit selection hardware, which has the benefit to further reduce the computation delay with respect to the time estimated in that work.> Paolo Montuschi, Luigi Ciminiera |
IEEE Trans. Computers | 1 |
| 1994 | Very-High Radix Division with Prescaling and Selection by RoundingabstractA division algorithm in which the quotient-digit selection is performed by rounding the shifted residual in carry-save form is presented. To allow the use of this simple function, the divisor (and dividend) is prescaled to a range close to one. The implementation presented results in a fast iteration because of the use of carry-save forms and suitable recodings. The execution time is calculated and several convenient values of the radix are selected. Comparison with other dividers for radices 2/sup 9/ to 2/sup 18/ is performed using the same assumptions.> Milos D. Ercegovac, Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 3 |
| 1994 | Over Redundant Digit Sets and the Design of Digit-by-Digit Division UnitsabstractOver-redundant digit sets are defined as those ranging from /spl minus/s to +s, with s/spl ges/B, B being the radix. This paper presents new techniques for the direct computation of division, that use an over-redundant digit set for representing the quotient, instead of simply redundant ones used previously. In particular, general criteria for synthesizing the digit selection rules and remainder updating are given for any radix and index of redundancy. A methodology combining the use of over-redundant digit sets with the prescaling of the divisor is also studied in order to achieve radix-B division units with trivial digit selection functions. It is also shown, for the specific case of radix-4 that using a prescaling slightly wider than in a radix-4 unit by M.D. Ercegovac and T. Lang (1990) possible to avoid the digit selection table. The paper also presents a modified algorithm for on-the-fly conversion of the result into the irredundant form. The proposed methodology can be considered as an alternative to existing division techniques.> Paolo Montuschi, Luigi Ciminiera |
IEEE Trans. Computers | 1 |
| 1993 | Very high radix division with selection by rounding and prescalingabstractA division algorithm in which the quotient-digit selection is performed by rounding the shifted residual in carry-save form is presented. To allow the use of this simple function, the divisor (and dividend) is prescaled to a range close to one. The implementation presented results in a fast iteration because of the use of carry-save forms and suitable recodings. The execution time is calculated, and several convenient values of the radix are selected. Comparison with other high-radix dividers is performed using the same assumptions.> Milos D. Ercegovac, Tomás Lang, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 3 |
| 1993 | n × n carry-save multipliers without final additionabstractCarry-save multipliers require an adder at the last step to convert the carry-sum representation of the most significant half of the result into an irredundant form. A multiplication scheme where by this conversion is performed with a circuit operating in parallel with the carry-save array is presented. The resulting implementation, when a radix-2 adder array is used, produces a result on 2n bits with a delay comparable to that of the multiplier proposed by M.D. Ercegovac and T. Lang (1990). When a radix-4 array is used, the proposed unit is almost twice as fast as units proposed previously.> Paolo Montuschi, Luigi Ciminiera |
IEEE Symposium on Computer Arithmetic | 1 |
| 1993 | Reducing Iteration Time When Result Digit is Zero for Radix 2 SRT Division and Square Root with Redundant RemaindersabstractA new architecture is presented for shared radix 2 division and square root whose main characteristic is the ability to avoid any addition/subtraction, when the digit 0 has been selected. The solution presented uses a redundant representation of the partial remainder, while keeping the advantages of classical solutions. It is shown how the next digit of the result can be selected even when the remainder is not updated, and the subsequent tradeoff is presented. The proposed architecture is also extended in order to consider other implementations.> Paolo Montuschi, Luigi Ciminiera |
IEEE Trans. Computers | 1 |
| 1992 | Throughput analysis of timed token protocols in double ring networksabstractA study of the throughput and time characteristics of double-ring networks is presented, assuming that the two channels are used for balancing the traffic (when both are working) and for synchronous traffic generated according to a generic but periodic pattern. In addition, reconfiguration using portions of both rings to circumvent faulty elements in a double ring network is shown to be equivalent to the single ring configuration already studied. Examples of applications of the results are also illustrated.> Claudio Giovanni Demartini, Paolo Montuschi, Adriano Valenzano, Luigi Ciminiera, Riccardo Sisto |
LCN | 2 |
| 1992 | Higher Radix Square Root with PrescalingabstractA scheme for performing higher radix square root based on prescaling of the radicand is presented to reduce the complexity of the result-digit selection. The scheme requires several steps, namely multiplication for prescaling the radicand, square root, multiplication for prescaling for the division, and division. Online algorithms are used to reduce the overall time and pipelining to reuse the different modules. An estimate of the execution time for a radix-256 unit for double-precision square root and a comparison with other implementations indicate that the proposed approach is an alternative to consider when designing a square-root unit.> Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 1992 | Design of a Radix 4 Division Unit with Simple Selection TableabstractA radix 4 division architecture is presented which partially overlaps the updating of the remainder with the digit selection procedure. It is obtained by separating the radix 4 digit selection process into two concurrent substeps. The proposed unit requires a simple selection table and involves a small extra expense for the additional hardware compared to the usual radix 4 division units. Four possible implementations are derived from the general model, with different types of substeps. The high level evaluation shows that the proposed architectures offer an efficient alternative.> Paolo Montuschi, Luigi Ciminiera |
IEEE Trans. Computers | 1 |
| 1991 | Simple radix 2 division and square root with skipping of some addition stepsabstractThe authors present a novel algorithm for shared radix 2 division and square root whose main characteristic is the ability to avoid any addition when the digit 0 has been selected. The solution presented uses a redundant representation of the partial remainder, while keeping the advantages of classical solutions. It is shown how the next digit of the result can be selected even when the remainder is not updated; the tradeoff arising is also indicated. The average occurrences of 0 digit selections are also estimated in order to assess the benefits of the algorithm presented.> Paolo Montuschi, Luigi Ciminiera |
IEEE Symposium on Computer Arithmetic | 1 |
| 1991 | On the Equivalence of IEEE 802.4 and FDDI Timed Token ProtocolsabstractTwo timed token protocols, the IEEE 802.4 and FDDI (fiber distributed data interface), are considered. Although IEEE 802.4 and FDDI have similar protocol rules, the different method of measuring the duration of the generic token rotation has imposed separate analyses for the two protocols. Presented is a general proof of equivalence between IEEE 802.4 and FDDI networks, where both consist of queues belonging to two priority classes. Having proved this equivalence, it is possible to extend results formally demonstrated for one protocol to the other. In particular, the results regarding the bounds on mean and maximum token rotation time derived for FDDI by Sevcik and Johnson (1987) can also be extended to IEEE 802.4 networks.> Paolo Montuschi, Adriano Valenzano, Luigi Ciminiera |
INFOCOM | 1 |
| 1990 | Higher Radix Square RootingabstractA general discussion on nonrestoring square root algorithms is presented, showing bounds and constraints delimiting the space of feasible algorithms, for all the choices of radix, digit set and representation of the partial remainder. Two classes of algorithms are then derived from the general discussion, and it is shown how it is possible to determine two parameters with a relevant impact on the implementation: the number of radicand bits to be inspected in order to obtain a starting value, and the number of partial remainder bits to be examined for digit selection. The algorithms for the specific case of radix 4 digit set (-2, -1, 0, +1, +2), and partial remainder represented in carry-save form are derived in order to show that the algorithms introduced can lead to better results than those obtained with algorithms previously presented.> Luigi Ciminiera, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 1990 | Some Properties of Timed Token Medium Access ProtocolsabstractTimed-token protocols are used to handle, on the same local area network, both real-time and non-real-time traffic. The authors analyze this type of protocol, giving worst-case values for the throughput of non-real-time traffic and the average token rotation time. Results are obtained for synchronous traffic generated according to a generic periodic pattern under heavy conditions for non-real-time traffic and express not only theoretical lower bounds but values deriving from the analysis of some real networks. A model which addresses the asynchronous overrun problem is presented. The influence of introducing multiple priority classes for non-real-time traffic on the total throughput of this type of message is shown. It is also shown that the differences between the values obtained under worst-case assumptions are close to those obtained under best-case assumptions; the method may therefore be used to provide important guidelines in properly tuning timed-token protocol parameters for each specific network installation.> Adriano Valenzano, Paolo Montuschi, Luigi Ciminiera |
IEEE Trans. Software Eng. | 2 |
| 1989 | On the efficient implementation of higher radix square root algorithmsabstractSquare root nonrestoring algorithms operating with a radix higher than two (but power of 2) are discussed. Formulas are derived delimiting the feasibility space of the class of algorithms considered as a function of the different parameters. This definition leads to the determination of some of these parameters; in particular, it is possible to compute the number of partial reminder bits to be inspected for digit selection and the number of operand bits to be inspected to generate the first radicand value, as both parameters have a relevant impact on the implementation. The specific case of radix 4, digit set (-2, -1, 0, +1, +2) and partial remainder represented by the sum of two numbers is considered.> Paolo Montuschi, Luigi Ciminiera |
IEEE Symposium on Computer Arithmetic | 1 |
| 1989 | On the Behavior of Control Token Protocols with Asynchronous and Synchronous TrafficabstractTimed token protocols are both used to handle, on the same local area network, both real-time and non-real-time traffic. The authors analyze this type of protocols, giving worst-case values for the throughput of non-real-time traffic and the average token rotation time. The results were obtained for synchronous traffic generated according to a generic periodic pattern, under heavy traffic conditions. Finally, it is shown that the difference between the values obtained under worst-case assumptions are close to those obtained under best-case assumptions. Therefore, the method presented here may be used to provide important guidelines so as to properly tune timed token protocol parameters for each specific network installation.> Adriano Valenzano, Paolo Montuschi, Luigi Ciminiera |
INFOCOM | 2 |
| 1989 | Some Properties of Double-Ring Networks with Real-Time ConstraintsabstractTimed-token protocols are used in local area networks to achieve bounded access times for a class of messages referred to as synchronous. An analysis is made of the behavior of this type of access protocol in a network with one redundant channel. Stations connected to both rings and stations connected to either ring are considered. A method of computing the minimum guaranteed throughput for asynchronous messages is shown, assuming that the two channels are used for balancing the traffic (when both are working) and for synchronous traffic generated according to a generic but periodic pattern.> Luigi Ciminiera, Paolo Montuschi, Adriano Valenzano |
RTSS | 2 |
| 1989 | Implementation of algorithms for graphic surface modeling using transputers
Adriano Valenzano, Paolo Montuschi, Luigi Ciminiera |
Microprocess. Microprogramming | 2 |