Giuseppe Coviello

dblp:49/8419 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bifröst: Peer-to-Peer Load-Balancing for Function Execution in Agentic AI Systems
Giuseppe Coviello, Kunal Rao, Mohammad Ali Amir Khojastepour, Srimat T. Chakradhar
Euro-Par (1)1
2025 XPF: Agentic AI System for Business Workflow Automation
abstract
In this paper, we propose a novel agentic AI system called XPF, which enables users to create "agents" using just natural language, where each agent is capable of executing complex, real-world business workflows in an accurate and reliable manner. XPF provides an interface to develop and iterate over the agent creation process and then deploy the agent in production when satisfactory results are produced consistently. The key components of XPF include: (a) planner, which leverages LLM to generate a step-by-step plan, which can further be edited by a human (b) compiler, which leverages LLM to compile the plan into a flow graph (c) executor, which handles distributed execution of the flow graph (using LLM, tools, RAG, etc.) on an underlying cluster and (d) verifier, which helps in verification of the output (through human generated tests or auto-generated tests using LLM). We develop five different agents using XPF and conduct experiments to evaluate one particular aspect i.e. difference in accuracy and reliability of the five agents with "human-generated" vs "auto-generated" plans. Our experiments show that we can get much more accurate and reliable response for a business workflow when step-by-step instructions (in natural language) are given by a human familiar with the workflow, rather than letting the LLM figure out the execution plan steps. In particular, we observe that "human-generated" plan almost always gives 100% accuracy whereas "auto-generated" plan almost never gives 100% accuracy. In terms of reliability, we observe through Rouge-L, Blue and Meteor scores, that the output from "human-generated" plan is much more reliable than "auto-generated" plan.
Kunal Rao, Giuseppe Coviello, Gennaro Mellone, Ciro Giuseppe De Vita, Srimat T. Chakradhar
HPDC2
2025 An Open Testbed for Mixed Reality Precise Rotation Guidance: Comparative Case Study of Arrow, Gestalt and Magnifier Cues
abstract
A wide range of tools, machines from diverse industries are operated by the precise rotation of physical devices (e.g., knobs, encoders). Literature proved that Mixed Reality can provide rotation guidance by superimposing digital cues onto the real environment. However, existing studies primarily focus on speed rather than precision and lack proven solutions with standard comparison methods. We presented a novel open MR testbed to evaluate MR rotation guidance cue designs. It comprises (i) a 3D-printed structure with a controller mount for MR cue tracking, (ii) an industrial encoder (resolution 1.8º), (iii) an Esp32 microcontroller, (iv) a firmware for USB connection, and (v) a USB foot pedal to confirm rotation, (vi) test software and documentation. The testbed uses available off-the-shelf components, and it is distributed with an open source license, Unity package with cues examples and scripts for running tests with randomized targets and collecting performance data (error count and speed). We evaluated the MR testbed by comparing the Arrow cue (baseline) and two novel designs: Magnifier with a zoomed region and Gestalt leveraging related perception principles. A within-subjects study ($\mathrm{N}=36,180^{\prime \prime}, 18$ọ,$\mathrm{M}=27.58, \pm 3.55$) showed that the Magnifier is the most precise ($p<.001$), while the Gestalt is the fastest ($p<.001$). Moreover, males (higher motor and gaming skills) ($\mathrm{p}<.002, \mathrm{t}=3.89$) resulted in faster ($\mathrm{r}=-0.458, \mathrm{p}<$. 005). Participants who declared with higher gaming skills demonstrated less task load during the experiment ($\mathrm{r}=+0.548, \mathrm{p}<.005$). Overall, the participants slightly preferred the Gestalt (18) over the Magnifier (15), while the Arrow was selected by a few (3). Ultimately, the testbed has proven to be an effective research and evaluation platform for future mixed reality interfaces.
Mine Dastan, Fabio Vangi, Francesco Musolino, Giuseppe Coviello, Michele Fiorentino 0001
ISMAR4
2025 G-Litter Marine Litter Dataset Augmentation with Diffusion Models and Large Language Models on GPU Acceleration
abstract
Marine litter detection is crucial for environmental monitoring, yet the imbalance in existing datasets limits model performance in identifying various types of waste accurately. This paper presents an efficient data augmentation pipeline that combines generative diffusion models (e.g., Stable Diffusion) and Large Language Models (LLMs) to expand the G-Litter dataset, a marine litter dataset designed for autonomous detection in heterogeneous environments. Leveraging scalable diffusion models for image generation and Alpaca LLMs for diverse prompt generation, our approach augments underrepresented classes by generating over 200 additional images per class, significantly improving the dataset’s balance. Training G-Litter augmented dataset using YOLOv8 for object detection demonstrated an increase in detection performance, improving recall by 7.82% and mAP50 by 3.87% (compared with baseline results). This study emphasizes the potential for combining generative AI with HPC resources to automate data augmentation on large-scale, unstructured datasets, particularly in edge computing contexts for real-time marine monitoring. The models were tested on real videos captured during simulated missions, demonstrating a superior ability to detect submerged objects in dynamic scenarios. These results highlight the potential of generative AI techniques to improve dataset quality and detection model performance, laying the foundation for further expansion in real-time marine monitoring.
Gennaro Mellone, Ciro Giuseppe De Vita, Emanuel Di Nardo, Giuseppe Coviello, Diana Di Luccio, Pietro Aucelli, Angelo Ciaramella, Raffaele Montella
PDP4
2025 CamTuner: Adaptive Video Analytics Pipelines via Real-Time Automated Camera Parameter Tuning
abstract
In Video Analytics Pipelines (VAP), Analytics Units (AUs) such as object detection and face recognition operating on remote servers rely heavily on surveillance cameras to capture high-quality video streams to achieve high accuracy. Modern network cameras offer an array of parameters that directly influence video quality. While a few of such parameters, e.g., exposure, focus and white balance, are automatically adjusted by the camera internally, the others are not. We denote such camera parameters as non-automated (NAUTO) parameters. In this work, we first show that in a typical surveillance camera deployment, environmental condition changes can have significant adverse effect on the accuracy of insights from the AUs, but such adverse impact can potentially be mitigated by dynamically adjusting NAUTO camera parameters in response to changes in environmental conditions. Second, since most end-users lack the skill or understanding to appropriately configure these parameters and typically use a fixed parameter setting, we presentCamTuner, to our knowledge, the first framework that dynamically adapts NAUTO camera parameters to optimize the accuracy of AUs in a VAP in response to adverse changes in environmental conditions.CamTuneris based on SARSA reinforcement learning and it incorporates two novel components: a light-weight analytics quality estimator and a virtual camera that drastically speed up offline RL training. Our controlled experiments and real-world VAP deployment show that compared to a VAP using the default camera setting,CamTunerenhances VAP accuracy by detecting 15.9% additional persons and 2.6% –4.2% additional cars (without any false positives) in a large enterprise parking lot.CamTuneropens up new avenues for elevating video analytics accuracy, transcending mere incremental enhancements achieved through refining deep-learning models.
Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Y. Charlie Hu, Srimat T. Chakradhar
IEEE Trans. Mob. Comput.3
2024 DiCE-M: Distributed Code Generation and Execution for Marine Applications - An Edge-Cloud Approach
abstract
Edge computing has emerged as a transformative technology that reduces application latency, improves cost efficiency, enhances security, and enables large-scale deployment of applications across various domains. In environmental monitoring, systems such as MegaSense[49], use low-cost sensors to gather and process real-time air quality data through edge-cloud collaboration, highlighting the critical role of edge computing in enabling scalable, efficient solutions. Similarly, marine science increasingly requires real-time processing and analysis of marine data from remote, resource-constrained environments. In this paper, we extend the power of edge computing by integrating it with Generative Artificial Intelligence(GenAI),specifically large language models (LLMs), to address challenges in marine science applications. We propose DiCE-M (Distributed Code generation and Execution for Marine applications), a robust system that uses LLM to generate distributed code for marine applications and then utilizes a runtime to efficiently execute it on an edge+cloud computing infrastructure. Specifically, DiCE-M leverages edge computing to execute lightweight AI models locally on unmanned surface vehicles(USVs)while offloading complex tasks to the cloud, thus balancing computational load and enabling realtime monitoring in marine environments. We use marine litter identification as an example application to demonstrate the utility of DiCE-M. Our results show that DiCE-M reduces latency by more than 2X when marine litter is not detected and cuts cloud computing costs by more than half compared to traditional cloud-based approaches. By selectively cropping and transmitting relevant image portions, DiCE-M further improves bandwidth efficiency, making it a reliable and cost-effective solution for deploying AI-drivenapplications on resource-constrained USVs in dynamic marine environments.
Giuseppe Coviello, Kunal Rao, Gennaro Mellone, Ciro Giuseppe De Vita, Srimat T. Chakradhar
SEC1
2024 LARA: Latency-Aware Resource Allocator for Stream Processing Applications
abstract
One of the key metrics of interest for stream processing applications is “latency”, which indicates the total time it takes for the application to process and generate insights from streaming input data. For mission-critical video analytics applications like surveillance and monitoring, it is of paramount importance to report an incident as soon as it occurs so that necessary actions can be taken right away. Stream processing applications are typically developed as a chain of microser-vices and are deployed on container orchestration platforms like Kubernetes. Allocation of system resources like “cpu” and “memory” to individual application microservices has direct impact on “latency”. Kubernetes does provide ways to allocate these resources e.g. through fixed resource allocation or through vertical pod autoscaler (VPA), however there is no straight-forward way in Kubernetes to prioritize “latency” for an end-to-end application pipeline. In this paper, we present LARA, which is specifically designed to improve “latency” of stream processing application pipelines. LARA uses a regression-based technique for resource allocation to individual microservices. We implement four real-world video analytics application pipelines i.e. license plate recognition, face recognition, human attributes detection and pose detection, and show that compared to fixed allocation, LARA is able to reduce latency by up to 2.8X and is consistently better than VPA. While reducing latency, LARA is also able to deliver over 2X throughput compared to fixed allocation and is almost always better than VPA.
Priscilla Benedetti, Giuseppe Coviello, Kunal Rao, Srimat T. Chakradhar
PDP2
2024 Improving Real-Time Data Streams Performance on Autonomous Surface Vehicles using DataX
abstract
In the evolving Artificial Intelligence (AI) era, the need for real-time algorithm processing in marine edge en-vironments has become a crucial challenge. Data acquisition, analysis, and processing in complex marine situations require sophisticated and highly efficient platforms. This study optimizes real-time operations on a containerized distributed processing platform designed for Autonomous Surface Vehicles (ASV) to help safeguard the marine environment. The primary objective is to improve the efficiency and speed of data processing by adopting a microservice management system called DataX. DataX leverages containerization to break down operations into modular units, and resource coordination is based on Kubernetes. This combination of technologies enables more efficient resource management and real-time operations optimization, contributing significantly to the success of marine missions. The platform was developed to address the unique challenges of managing data and running advanced algorithms in a marine context, which often involves limited connectivity, high latencies, and energy restrictions. Finally, as a proof of concept to justify this platform's evolution, experiments were carried out using a cluster of single-board computers equipped with GPUs, running an AI-based marine litter detection application and demonstrating the tangible benefits of this solution and its suitability for the needs of maritime missions.
Gennaro Mellone, Ciro Giuseppe De Vita, Giuseppe Coviello, Pietro Aucelli, Angelo Ciaramella, Raffaele Montella
PDP3
2024 A Reconfigurable Full-Digital Architecture for Angle of Arrival Estimation
abstract
In recent years, Real-Time Localization Systems (RTLS) are gaining significant attention from the scientific community. Usually, RLTS rely on positioning techniques based on the electromagnetic (e.m.) propagation properties of the received signals. Among the many possible approaches, RTLS based on the Angle-of-Arrival (AoA) estimation allows for accurate results with a comparatively lower number of receivers. In this work, we propose a novel, fully digital synchronous architecture for AoA estimation based on phase interferometry. After reviewing the state-of-the-art on full-hardware techniques for localization, the main blocks composing the proposed architecture, are discussed in detail, along with their dimensioning equations. The granularity with which the system estimates the phase shifts used to compute the AoA is reconfigurable according to the desired accuracy. The architecture is lightweight and computes the AoA in real-time. To validate the proposed approach, the RTLS architecture has been implemented on an Intel Cyclone IV E EP4CE115F29C7 FPGA board interfaced with a custom-designed receiver front-end, and experimental results on benchmark applications have been collected and analyzed.
Antonello Florio, Gianfranco Avitabile, Claudio Talarico, Giuseppe Coviello
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Citizen Science for the Sea with Information Technologies: An Open Platform for Gathering Marine Data and Marine Litter Detection from Leisure Boat Instruments
abstract
Data crowdsourcing is an increasingly pervasive and lifestyle-changing technology due to the flywheel effect that results from the interaction between the Internet of Things and Cloud Computing. This paper presents the Citizen Science for the Sea with Information Technologies (C4Sea-IT) framework. It is an open platform for gathering marine data from leisure boat instruments. C4Sea-IT aims to provide a coastal marine data gathering, moving, processing, exchange, and sharing platform using the existing navigation instruments and sensors for today's leisure and professional vessels. In this work, a use case for the detection and tracking of marine litter is shown. The final goal is weather/ocean forecasts argumentation with Artificial Intelligence prediction models trained with crowdsourced data.
Ciro Giuseppe De Vita, Gennaro Mellone, Dante D. Sánchez-Gallegos, Giuseppe Coviello, Diego Romano, Marco Lapegna, Angelo Ciaramella
e-Science4
2023 Content-aware auto-scaling of stream processing applications on container orchestration platforms
abstract
Modern applications are designed as an interacting set of microservices, and these applications are typically deployed on container orchestration platforms like Kubernetes. Several attractive features in Kubernetes make it a popular choice for deploying applications, and automatic scaling is one such feature. The default horizontal scaling technique in Kubernetes is the Horizontal Pod Autoscaler (HPA). It scales each microservice independently while ignoring the interactions among the microservices in an application. In this paper, we show that ignoring such interactions by HPA leads to inefficient scaling, and the optimal scaling of different microservices in the application varies as the stream content changes. To automatically adapt to variations in stream content, we present a novel system called DataX AutoScaler that leverages knowledge of the entire stream processing application pipeline to efficiently auto-scale different microservices by taking into account their complex interactions. Through experiments on real-world video analytics applications, such as face recognition and pose classification, we show that DataX AutoScaler adapts to variations in stream content and achieves up to 43% improvement in overall application performance compared to a baseline system that uses HPA.
Giuseppe Coviello, Kunal Rao, Ciro Giuseppe De Vita, Gennaro Mellone, Priscilla Benedetti, Srimat T. Chakradhar
PDP1
2023 Elixir: A System to Enhance Data Quality for Multiple Analytics on a Video Stream
abstract
IoT sensors, especially video cameras, are ubiquitously deployed around the world to perform a variety of computer vision tasks in several verticals including retail, health-care, safety and security, transportation, manufacturing, etc. To amortize their high deployment effort and cost, it is desirable to perform multiple video analytics tasks, which we refer to as Analytical Units (AUs), off the video feed coming out of every camera. As AUs typically use deep-learning based AI/ML models, their performances depend on the quality of the input video. The most recent work has shown that dynamically adjusting the camera setting exposed by popular network cameras can help improve the quality of the video feed and hence the AU accuracy, in a single AU setting. In this paper, we first show that in a multi-AU setting, changing the camera setting has disproportionate impact on different AUs performance. In particular, the optimal setting for one AU may severely degrade the performance for another AU, and further, the impact on different AUs varies as the environmental condition changes. We then present Elixir, a system to enhance the video stream quality for multiple analytics on a video stream. Elixir leverages Multi-Objective Reinforcement Learning (MORL), where the RL agent caters to the objectives from different AUs and adjusts the camera setting to simultaneously enhance the performance of all AUs. To define the multiple objectives in MORL, we develop new AU-specific quality estimator values for each individual AU. We evaluate Elixir through real-world experiments on a testbed with three cameras deployed next to each other (overlooking a large enterprise parking lot) running Elixir and two baseline approaches, respectively. Elixir correctly detects 7.1% (22,068) and 5.0% (15,731) more cars, 94% (551) and 72% (478) more faces, and 670.4% (4975) and 158.6% (3507) more persons than the default-setting and time-sharing approaches, respectively. It also detects 115 license plates, far more than the time-sharing approach (7) and the default setting (0).
Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Y. Charlie Hu, Srimat T. Chakradhar
SMARTCOMP3
2023 AnB: Application-in-a-Box to Rapidly Deploy and Self-optimize 5G Apps
abstract
We present "Application in a Box" (AnB) product concept aimed at simplifying the deployment and operation of remote 5G applications. AnB comes pre-configured with all necessary hardware and software components, including sensors like cameras, hardware and software components for a local 5G wireless network, and 5G-ready apps. Enterprises can easily download additional apps from an App Store. Setting up a 5G infrastructure and running applications on it is a significant challenge, but AnB is designed to make it fast, convenient, and easy, even for those without extensive knowledge of software, computers, wireless networks, or AI-based analytics. With AnB, customers only need to open the box, set up the sensors, turn on the 5G networking and edge computing devices, and start running their applications. Our system software automatically deploys and optimizes the pipeline of microservices in the application on a tiered computing infrastructure that includes device, edge, and cloud computing. Application scalability, dynamic resource management, placement of critical tasks for low-latency response, and dynamic network bandwidth allocation for efficient 5G network usage are all automatically orchestrated.AnB offers cost savings, simplified setup and management, and increased reliability and security. We’ve implemented several real-world applications, such as collision prediction at busy traffic light intersections and remote construction site monitoring using video analytics. With AnB, deployment and optimization effort can be reduced from several months to just a few minutes. This is the first-of-its-kind approach to easing deployment effort and automating self-optimization of the application during system operation.
Kunal Rao, Murugan Sankaradass, Giuseppe Coviello, Ciro Giuseppe De Vita, Gennaro Mellone, Wang-Pin Hsiung, Srimat T. Chakradhar
SMARTCOMP3
2022 Enhancing Video Analytics Accuracy via Real-time Automated Camera Parameter Tuning
abstract
In Video Analytics Pipelines (VAP), Analytics Units (AUs) such as object detection and face recognition running on remote servers critically rely on surveillance cameras to capture high-quality video streams in order to achieve high accuracy. Modern IP cameras come with a large number of camera parameters that directly affect the quality of the video stream capture. While a few of such parameters, e.g., exposure, focus, white balance are automatically adjusted by the camera internally, the remaining ones are not. We denote such camera parameters as non-automated (NAUTO) parameters. In this paper, we first show that environmental condition changes can have significant adverse effect on the accuracy of insights from the AUs, but such adverse impact can potentially be mitigated by dynamically adjusting NAUTO camera parameters in response to changes in environmental conditions. We then present CamTuner, to our knowledge, the first framework that dynamically adapts NAUTO camera parameters to optimize the accuracy of AUs in a VAP in response to adverse changes in environmental conditions. CamTuner is based on SARSA reinforcement learning and it incorporates two novel components: a light-weight analytics quality estimator and a virtual camera that drastically speed up offline RL training. Our controlled experiments and real-world VAP deployment show that compared to a VAP using the default camera setting, CamTuner enhances VAP accuracy by detecting 15.9% additional persons and 2.6%--4.2% additional cars (without any false positives) in a large enterprise parking lot and 9.7% additional cars in a 5G smart traffic intersection scenario, which enables a new usecase of accurate and reliable automatic vehicle collision prediction (AVCP). CamTuner opens doors for new ways to significantly enhance video analytics accuracy beyond incremental improvements from refining deep-learning models.
Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Oliver Po, Y. Charlie Hu, Srimat T. Chakradhar
SenSys3
2021 ECO: Edge-Cloud Optimization of 5G applications
abstract
Centralized cloud computing with 100+ milliseconds network latencies cannot meet the tens of milliseconds to sub-millisecond response times required for emerging 5G applications like autonomous driving, smart manufacturing, tactile internet, and augmented or virtual reality. We describe a new, dynamic runtime that enables such applications to make effective use of a 5G network, computing at the edge of this network, and resources in the centralized cloud, at all times. Our runtime continuously monitors the interaction among the microservices, estimates the data produced and exchanged among the microservices, and uses a novel graph min-cut algorithm to dynamically map the microservices to the edge or the cloud to satisfy application-specific response times. Our runtime also handles temporary network partitions, and maintains data consistency across the distributed fabric by using microservice proxies to reduce WAN bandwidth by an order of magnitude, all in an application-specific manner by leveraging knowledge about the application's functions, latency-critical pipelines and intermediate data. We illustrate the use of our runtime by successfully mapping two complex, representative real-world video analytics applications to the AWS/Verizon Wavelength edge-cloud architecture, and improving application response times by 2x when compared with a static edge-cloud implementation.
Kunal Rao, Giuseppe Coviello, Wang-Pin Hsiung, Srimat T. Chakradhar
CCGRID2
2021 Magic-Pipe: self-optimizing video analytics pipelines
abstract
Microservices-based video analytics pipelines routinely use multiple deep convolutional neural networks. We observe that the best allocation of resources to deep learning engines (or microservices) in a pipeline, and the best configuration of parameters for each engine vary over time, often at a timescale of minutes or even seconds based on the dynamic content in the video. We leverage these observations to develop Magic-Pipe, a self-optimizing video analytic pipeline that leverages AI techniques to periodically self-optimize. First, we propose a new, adaptive resource allocation technique to dynamically balance the resource usage of different microservices, based on dynamic video content. Then, we propose an adaptive microservice parameter tuning technique to balance the accuracy and performance of a microservice, also based on video content. Finally, we propose two different approaches to reduce unnecessary computations due to unavoidable mismatch of independently designed, re-usable deep-learning engines: a deep learning approach to improve the feature extractor performance by filtering inputs for which no features can be extracted, and a low-overhead graph-theoretic approach to minimize redundant computations across frames. Our evaluation of Magic-Pipe shows that pipelines augmented with self-optimizing capability exhibit application response times that are an order of magnitude better than the original pipelines, while using the same hardware resources, and achieving similar high accuracy.
Giuseppe Coviello, Yi Yang 0018, Kunal Rao, Srimat T. Chakradhar
Middleware1
2021 F3S: Free Flow Fever Screening
abstract
Identification of people with elevated body temperature can reduce or dramatically slow down the spread of infectious diseases like COVID-19. We present a novel fever-screening system, F3S, that uses edge machine learning techniques to accurately measure core body temperatures of multiple individuals in a free-flow setting. F3S performs real-time sensor fusion of visual camera with thermal camera data streams to detect elevated body temperature, and it has several unique features: (a) visual and thermal streams represent very different modalities, and we dynamically associate semantically-equivalent regions across visual and thermal frames by using a new, dynamic alignment technique that analyzes content and context in real-time, (b) we track people through occlusions, identify the eye (inner canthus), forehead, face and head regions where possible, and provide an accurate temperature reading by using a prioritized refinement algorithm, and (c) we robustly detect elevated body temperature even in the presence of personal protective equipment like masks, or sunglasses or hats, all of which can be affected by hot weather and lead to spurious temperature readings. F3S has been deployed at over a dozen large commercial establishments, providing contact-less, free-flow, real-time fever screening for thousands of employees and customers in indoors and outdoor settings.
Kunal Rao, Giuseppe Coviello, Min Feng 0001, Biplob Debnath, Wang-Pin Hsiung, Murugan Sankaradass, Yi Yang 0018, Oliver Po, Utsav Drolia, Srimat T. Chakradhar
SMARTCOMP2
2018 An Integrated Phase Shifting Frequency Synthesizer for Active Electronically Scanned Arrays
abstract
This paper presents the design of an integrated phase shifting frequency synthesizer for active electronically scanned arrays. The proposed frequency synthesizer, implements a 3-channel beam steering unit based on a local oscillator (LO) phase shifting approach that relies on the joint use of external loop filters and voltage-controlled oscillators. The design has been fabricated using AMS 0.35-μm SiGe BiCMOS process and it takes a die area of 9.38-mm2. Based on a frequency independent architecture and targeting space and weight constrained applications, the integrated circuit (IC) allows to program, its output channels, with a 10-bit precision 0° to 360° mutual phase shifts.
Giulio D'Amato, Giovanni Piccinni, Gianfranco Avitabile, Giuseppe Coviello, Claudio Talarico
DDECS4
2017 Narrowband distance evaluation technique for indoor positioning systems based on Zadoff-Chu sequences
abstract
A rugged method for distance estimation based on an OFDM signal combined with the Zadoff-Chu sequences is introduced. A software algorithm extracts a coarse distance in time domain exploiting the Zadoff-Chu properties. The distance, then, is fine adjusted evaluating the time offset from the subcarriers phase. Several tests were performed in laboratory considering a very harsh indoor environment and using a pair of Right-Hand circularly polarized helical antennas in order to demonstrate the technique effectiveness. The system is able to estimate the distance with a very good level of precision and accuracy with only 50 MHz of signal bandwidth. In fact, the results show a maximum error of 1.23 cm, in distance estimation, and a maximum standard deviation of 0.8 cm. The proposed technique can be used in the implementation of indoor positioning systems.
Giovanni Piccinni, Gianfranco Avitabile, Giuseppe Coviello
WiMob3
2015 High Precision Digital Based 3.8GHz Phase Shifter
abstract
The paper reports a highly repeatable 10-bits 3.8 GHz phase shifter implemented in a standard 0.35 μm SiGe technology for high precision applications. The prototype principle of operations and design are thoroughly described and demonstrated. The generated 0.35° phase step exhibits a rms error lower than 0.32°, mainly determined by the PLL inserted in the architecture.
Francesco Cannone, Gianfranco Avitabile, Giuseppe Coviello, Giovanni Piccinni
DDECS3
2014 Snapify: capturing snapshots of offload applications on xeon phi manycore processors
abstract
Intel Xeon Phi coprocessors provide excellent performance acceleration for highly parallel applications and have been deployed in several top-ranking supercomputers. One popular approach of programming the Xeon Phi is the offload model, where parallel code is executed on the Xeon Phi, while the host system executes the sequential code. However, Xeon Phi's Many Integrated Core Platform Software Stack (MPSS) lacks fault-tolerance support for offload applications. This paper introduces Snapify, a set of extensions to MPSS that provides three novel features for Xeon Phi offload applications: checkpoint and restart, process swapping, and process migration. The core technique of Snapify is to take consistent process snapshots of the communicating offload processes and their host processes. To reduce the PCI latency of storing and retrieving process snapshots, Snapify uses a novel data transfer mechanism based on remote direct memory access (RDMA). Snapify can be used transparently by single-node and MPI applications, or be triggered directly by job schedulers through Snapify's API. Experimental results on OpenMP and MPI offload applications show that Snapify adds a runtime overhead of at most 5%, and this overhead is low enough for most use cases in practice.
Arash Rezaei, Giuseppe Coviello, Cheng-Hong Li, Srimat T. Chakradhar, Frank Mueller 0001
HPDC2
2014 A Coprocessor Sharing-Aware Scheduler for Xeon Phi-Based Compute Clusters
abstract
We propose a cluster scheduling technique for compute clusters with Xeon Phi coprocessors. Even though the Xeon Phi runs Linux which allows multiprocessing, cluster schedulers generally do not allow jobs to share coprocessors because sharing can cause oversubscription of coprocessor memory and thread resources. It has been shown that memory or thread oversubscription on a many core like the Phi results in job crashes or drastic performance loss. We first show that such an exclusive device allocation policy causes severe coprocessor underutilization: for typical workloads, on average only 38% of the Xeon Phi cores are busy across the cluster. Then, to improve coprocessor utilization, we propose a scheduling technique that enables safe coprocessor sharing without resource oversubscription. Jobs specify their maximum memory and thread requirements, and our scheduler packs as many jobs as possible on each coprocessor in the cluster, subject to resource limits. We solve this problem using a greedy approach at the cluster level combined with a knapsack-based algorithm for each node. Every coprocessor is modeled as a knapsack and jobs are packed into each knapsack with the goal of maximizing job concurrency, i.e., as many jobs as possible executing on each coprocessor. Given a set of jobs, we show that this strategy of packing for high concurrency is a good proxy for (i) reducing make span, without the need for users to specify job execution times and (ii) reducing coprocessor footprint, or the number of coprocessors required to finish the jobs without increasing make span. We implement the entire system as a seamless add on to Condor, a popular distributed job scheduler, and show make span and footprint reductions of more than 50% across a wide range of workloads.
Giuseppe Coviello, Srihari Cadambi, Srimat T. Chakradhar
IPDPS1
2013 COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors
Srihari Cadambi, Giuseppe Coviello, Cheng-Hong Li, Rajat Phull, Kunal Rao, Murugan Sankaradass, Srimat T. Chakradhar
HPDC2
2010 A GPGPU Transparent Virtualization Component for High Performance Computing Clouds
Giulio Giunta, Raffaele Montella, Giuseppe Agrillo, Giuseppe Coviello
Euro-Par (1)4