Lucas Francisco Wanner

dblp:w/LucasFranciscoWanner · also Lucas Wanner 0001 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-5564-698XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 A heuristic approach for near Pareto-optimal design space exploration in Approximate High-Level Synthesis
Tiago Almeida 0002, Isaías B. Felzmann, Lucas Francisco Wanner
Integr.3
2025 Best practices for responsible machine learning in credit scoring
Giovani Valdrighi, Athyrson M. Ribeiro, Jansen Silva de Brito Pereira, Vitória Guardieiro, Arthur Hendricks, Décio Miranda Filho, Juan David Nieto Garcia, Felipe F. Bocca, Thalita B. Veronese, Lucas Francisco Wanner, Marcos M. Raimundo
Neural Comput. Appl.10
2023 Enhancing Disaster Management of Guyed Towers through Machine Learning-Based Data Fusion
abstract
Power grid networks eventually require guyed towers as support structures for transmission lines. In these cases, the cables that support the structure can experience long-term degradation as a result of environmental forces. Long-term degradation can result in tower collapse leading to power outages in essential public services. We propose a Structure Health Monitor (SHM) system to improve transmission line reliability. It is based on data fusion using machine learning algorithms that monitor acceleration signals collected from multiple locations of the guyed tower that can allow the identification of loose cables and the amount of tightness. We select the most relevant individual measurable properties as a set of uncorrelated input sources, which can be in the time or frequency domain. The overall results for loose cable estimation of all investigated methods show a balanced accuracy in the range of 85% up to 96% and identification of tightness shows values between 91% and 96% for respectively inference using selected features set by investigated methods and all features available.
Juliane Regina de Oliveira, German Efrain Casteñeda Jimenez, Claudio Ferreira Dias, Eduardo Rodrigues de Lima, Janito Vaqueiro Ferreira, Larissa Medeiros de Almeida, Lucas Francisco Wanner
FUSION7
2022 Data fusion strategies for improving resilience to sensor noise in cable-stayed tower monitoring
Juliane Regina de Oliveira, Claudio Ferreira Dias, Eduardo Rodrigues de Lima, Larissa Medeiros de Almeida, Lucas Francisco Wanner
FUSION5
2022 Approximate Memory with Protected Static Allocation
abstract
Approximate memories provide energy savings or performance improvements at the cost of occasional errors in stored data. Applications that tolerate errors on their data profit from this trade-off by controlling these errors to not affect critical data. This control usually involves programmer intervention with annotations in the source code. To avoid annotations, some techniques protect critical data that are common on many applications, isolating specific memory regions from errors. In this work, we propose and explore alternatives for the protection of application critical data by managing a supervisor execution environment with an approximate memory system. We expose only dynamically allocated data to errors with secure data manipulation through an approximate allocation scheme that divide stored data based on the approximation of the heap area. We evaluate 6 applications with different data access profiles and obtain up to 20% of energy savings.
João Fabrício Filho, Isaías B. Felzmann, Lucas Francisco Wanner
SBAC-PAD3
2022 Prof5: A RISC-V profiler tool
abstract
RISC-V is supported by a series of design and simulation tools that enable simple instruction set customization and rapid exploration of application-specific accelerators. Evaluating the performance and energy impact of specific design choices and optimizations on applications remains, however, challenging. Traditional RT- or Gate-level simulation, while fairly precise, is complex and slow, and is, therefore, typically limited to small fractions of code. Functional simulation, while faster, is typically imprecise and lacks the detailed information presented by traditional profilers. We introduce Prof5, a profiler for RISC-V designs that combines functional simulation with precise energy and timing models calibrated from RTL simulation and power analysis. Prof5 is based on the Spike simulator and provides detailed, function-level timing and energy statistics that can be used to guide design and optimization choices, and enable rapid design-space exploration. Prof5 can furthermore aid the user in creating new timing and energy models for custom designs and architecture variations. Energy and timing estimation with Prof5 is 8000x faster than traditional synthesis-based analysis with an average of 95% accuracy for an embedded RISC-V processor.
Jonathas Silveira, Lucas Castro, Rodrigo Zeli, Daniel Lazari, Marcelo Guedes, Rodolfo Azevedo, Lucas Francisco Wanner
SBAC-PAD8
2021 AxPIKE: Instruction-level Injection and Evaluation of Approximate Computing
abstract
Representing the interaction between accurate and approximate hardware modules at the architecture level is essential to understand the impact of Approximate Computing in a general-purpose computing scenario. However, extensive effort is required to model approximations into a baseline instruction-level simulator and collect its execution metrics. In this work, we present the AxPIKE ISA simulation environment, a tool that allows designers to inject models of hardware approximation at the instruction level and evaluate their impact on the quality of results. AxPIKE embeds a high-level representation of a RISC-V system and produces a dedicated control mechanism, that allows the simulated software to manage the approximate behavior of compatible execution scenarios. The environment also provides detailed execution statistics that are forwarded to dedicated tools for energy accounting. We apply the AxPIKE environment to inject integer multiplication and memory access approximations into different applications and demonstrate how the generated statistics are translated into energy-quality trade-offs.
Isaías B. Felzmann, João Fabrício Filho, Lucas Francisco Wanner
DATE3
2021 Special Session: How much quality is enough quality? A case for acceptability in approximate designs
abstract
Approximate systems are designed to offer improved efficiency with potentially reduced quality of results. Quality of output in these systems is typically quantified in comparison to a precise result using metrics such as RMSE, MAE, PSNR, or application-specific metrics such as structural similarity of images (SSIM). Furthermore, systems are typically designed to maximize efficiency for a given minimum quality requirement. It is often difficult to determine what this quality requirement should be for an application, let alone a system. Thus, a fixed quality requirement may be overly conservative, and leave optimization opportunities on the table. In this work, we present a different approach to evaluate approximate systems based on the usefulness of results instead of quality. Our method qualitatively determines the acceptability of approximate results within different processing pipelines. To demonstrate the method, we implement three image and signal processing applications featuring scenarios of image classification, image recognition, and frequency estimation. Our results show that designing approximate systems to guarantee acceptability can produce up to 20% more valid results than the conservative quality thresholds commonly adopted in the literature, allowing for higher error rates and, consequently, lower energy cost.
Isaías B. Felzmann, João Fabrício Filho, Juliane Regina de Oliveira, Lucas Francisco Wanner
ICCD4
2021 Functional Approximation and Approximate Parallelization with the ACCEPT compiler
abstract
Approximate computing can aid in the use of energy on cloud and mobile systems by exchanging result accuracy for faster processing times and energy efficiency. One fundamental challenge in applying approximate techniques arises when trying to identify which parts of the application are resilient to approximations, since resilience is not fixed and is dependent on the sensibility of the code. ACCEPT [1] is an approximate computing framework that applies multiple approximation techniques at the compiling level, and that presents suggestions of sections of code that could be annotated to enforce such techniques. The original ACCEPT framework applies a variety of approximation techniques, including hardware acceleration and loop perforation. This work extends ACCEPT in order to attack a broader range of applications and approximation techniques. We extend the framework to support function approximation and approximate loop parallelization. We evaluate the extended framework with seven benchmarks, showing the resulting speedup and quality degradation for each technique when applied in isolation and in combination with other approximation targets. We introduce these approximations without requiring application-level annotations. Our experiments with set benchmarks showed maximum speedups ranging from$1.22x$up to$634x$for the combined techniques, with quality degradation between 0% and 29%.
Lucas Reis, Lucas Francisco Wanner
SBAC-PAD2
2021 Design and evaluation of associative processing kernels
abstract
Associative Processing is a way to do Processing in Memory. Through an associative memory (Content-addressable memory - CAM) with special registers and lookup-tables, and applying comparisons and writings, this approach is able to compute in parallel. The potential of associative computing is applied to accelerate several applications including Neural Networks, DNA alignment, and Image processing. However, there is a lack of programming models for real applications and ways to measure performance impacts in a system with associative processing, making it difficult to adopt the approach. In this work, we present associative processing implementations for application kernels along with evaluation models for latency and energy for associative operations. We highlight implementations for matrix multiplication, 2D convolution, and the ReLU activation function, explaining how we built them to extract parallelism in the associative environment. We use a functional simulator to evaluate applications and apply its output to models for performance evaluation. The results show that associative processing can greatly improve performance over traditional CPU processing for parallel applications. For matrix multiplication, the associative processing model achieved a relative gain of 5x in latency, 4x in energy, and 25x in the number of load/store operations when compared to a CPU-only model.
Jonathas Silveira, Lucas Francisco Wanner
SBAC-PAD2
2020 ADeLe: A description language for approximate hardware
Isaías B. Felzmann, Matheus Martins Susin, Liana Dessandre Duenha, Rodolfo Azevedo, Lucas Francisco Wanner
Future Gener. Comput. Syst.5
2020 AxRAM: A lightweight implicit interface for approximate data access
João Fabrício Filho, Isaías B. Felzmann, Rodolfo Azevedo, Lucas Francisco Wanner
Future Gener. Comput. Syst.4
2020 Risk-5: Controlled Approximations for RISC-V
abstract
Approximate Computing offers enhanced energy efficiency by exploring quality relaxation on applications. Application-agnostic hardware-level techniques can provide high benefits under certain scenarios, but their integration on a general-purpose architecture presents novel control challenges. We present Risk-5, an extension of the RISC-V architecture that implements control mechanisms to orchestrate multiple coexisting approximation techniques within an architecture. In Risk-5, approximate hardware capabilities are exposed to software through identification registers, data structures, and drivers that describe the nature and configuration parameters for each approximate design. This allows the software stack to control what and how much is approximated in an application. Control options range from activating or deactivating a certain approximation (e.g., approximating ALU operations), to configuring allowable error levels (e.g., for a configurable FPU), and configuring operation parameters that may lead to probabilistic errors (e.g., setting the refresh rate for an approximate SDRAM). Approximations may be dynamically configured and combined at runtime, allowing for simplified design space exploration. Finally, supervisor- and machine-level control allows for the use of certain approximations without requiring changes to applications. In this article, we discuss the implementation of different classes of approximation techniques, detailing and evaluating how they interact with each other. Risk-5 and the selected approximations are demonstrated in the functional level in a RISC-V ISA simulator augmented with an approximate computing framework. Our experiments evaluate how six applications from different computing domains behave when subjected to a combination of approximation techniques. Our results show how Risk-5 can bridge the gap between software and hardware approximations, allowing designers to easily evaluate energy-quality tradeoffs.
Isaías B. Felzmann, João Fabrício Filho, Lucas Francisco Wanner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 ADeLe: Rapid Architectural Simulation for Approximate Hardware
abstract
Recent research has introduced approximate hardware units that produce incorrect outputs deterministically or probabilistically for some small subset of inputs but allow significantly higher throughput or lower power than their errorfree counterparts. The integration, validation, and evaluation of these approximate units in architectures and processors, however, remains challenging. In this paper, we introduce ADeLe, a high-level language for the description, configuration, and integration of approximate hardware units into processors. ADeLe reduces the design effort for approximate hardware by modeling approximations at a high level of abstraction and automatically injecting them into a processor model for architectural simulation. Approximations in ADeLe may modify or completely replace the functional behavior of instructions according to user-defined policies. Instructions may be approximated deterministically or probabilistically (e.g., based on operating voltage and frequency). To allow for controlled testing, approximations may be enabled and disabled from software. Energy is automatically accounted based on customizable models that consider the potential power savings of the approximations that are enabled in the system. ADeLe provides designers with a generic and flexible verification framework, allowing them to easily evaluate the energy-quality trade-offs of their designs in applications. We demonstrate the language and corresponding framework by introducing different approximation techniques into a processor model, on top of which we run selected applications. We demonstrate ADeLe using 6 approximate designs with 4 image processing and 2 floating point applications. Our experiments show how ADeLe may be used to generate approximate CPUs and to evaluate energy-quality trade-offs for different applications with reduced effort.
Isaías B. Felzmann, Matheus Martins Susin, Liana Dessandre Duenha, Rodolfo Azevedo, Lucas Francisco Wanner
SBAC-PAD5
2016 Speculative Precision Time Protocol: Submicrosecond clock synchronization for the IoT
abstract
Time synchronization is a keystone of Wireless Sensor Networks (WSN). It is fundamental to coordinate the action of nodes in a network and it is also a critical element of several security mechanisms. In this paper, we discuss and evaluate the time synchronization strategy behind the Trustful Space-Time Protocol (TSTP), which explores the protocol's cross-layer architecture to speculatively peek through the timestamps and geographic info present in message headers, implementing high-accuracy clock synchronization with minimal insertion of explicit messages. We evaluate the protocol analytically and experimentally. The analytic evaluation is based on the model defined by Schmid [15] for the Virtual High-resolution Time (VHT), while the experimental evaluation was performed on the IEEE 802.15.4-compliant EPOSMote platform running EPOS and TSTP. Our results demonstrate that nodes in the network can be consistently synchronized with sub-microsecond precision while exchanging far less messages than they would with an ordinary, non-speculative implementation, resulting in energy savings. Indeed, precision and energy savings are higher for networks with higher traffic, since more messages are available for peeking. In an experiment scenario in which messages were exchanged between devices every 15 seconds, nodes in the network achieved a synchronization error of approximately 15 microseconds in the worst case, while in a scenario in which messages were exchanged every 3 seconds, synchronization error was less than 0.5 microseconds in the worst case, and approximately 0.25 microseconds on average.
Davi Resner, Antônio Augusto Fröhlich, Lucas Francisco Wanner
ETFA3
2015 A Framework for Dynamic Real-Time Reconfiguration
abstract
In this work, we propose a framework capable of transparently switching between multiple hardware and software implementations of embedded system components to cope with and adapt to dynamic runtime characteristics such as power, throughput and quality of service. The reconfiguration process is decomposed into small steps such that it is preemptable, transparent, dynamic and compliant with real-time requirements. We present a Private Automatic Branch eXchange(PABX) system as a case study for the framework, and investigate the dynamic reconfiguration of three of its components: an ADPCM codec, a DTMF detector, and an AES core. Our results with this case study show how our framework is able to perform reconfiguration of hardware/software components in the order of a few milliseconds, without taking excessive system resources, and without disrupting the execution of application threads.
Joao Gabriel Reis, Lucas Francisco Wanner, Antônio Augusto Fröhlich
DSD2
2015 CAreDroid: Adaptation Framework for Android Context-Aware Applications
abstract
Context-awareness is the ability of software systems to sense and adapt to their physical environment. Many contemporary mobile applications adapt to changing locations, connectivity states, available computational and energy resources, and proximity to other users and devices. Nevertheless, there is little systematic support for context-awareness in contemporary mobile operating systems. Because of this, application developers must build their own context-awareness adaptation engines, dealing directly with sensors and polluting application code with complex adaptation decisions. In this paper, we introduce CAreDroid, which is a framework that is designed to decouple the application logic from the complex adaptation decisions in Android context-aware applications. In this framework, developers are required- only-to focus on the application logic by providing a list of methods that are sensitive to certain contexts along with the permissible operating ranges under those contexts. At run time, CAreDroid monitors the context of the physical environment and intercepts calls to sensitive methods, activating only the blocks of code that best fit the current physical context. CAreDroid is implemented as part of the Android runtime system. By pushing context monitoring and adaptation into the runtime system, CAreDroid eases the development of context-aware applications and increases their efficiency. In particular, case study applications implemented using CAre-Droid are shown to have: (1) at least half lines of code fewer and (2) at least 10× more efficient in execution time compared to equivalent context-aware applications that use only standard Android APIs.
Salma Hosni Emam Mohamed Elmalaki, Lucas Francisco Wanner, Mani Srivastava 0001
MobiCom2
2015 X-Ware: mutant computing substrates
abstract
In this paper we introduce X-Ware, a framework for computing whereby components used by an application can take different forms and characteristics across the lifetime of the system in order to adjust to dynamic application requirements. In particular, we explore two aspects of system mutability: dynamic choice of implementations for certain system components (e.g., leading to different trade-offs between quality and resource usage); and the change in non-functional characteristics of these components (e.g., speed and energy consumption) due to process and environmental variations. We demonstrate X-Ware with a library of mathematical APIs that can help components save up to 93% in energy and 94% in execution time by tolerating a small degradation in quality.
Joao Gabriel Reis, Antônio Augusto Fröhlich, Lucas Francisco Wanner
RSP3
2015 Runtime Optimization of System Utility with Variable Hardware
abstract
Increasing hardware variability in newer integrated circuit fabrication technologies has caused corresponding power variations on a large scale. These variations are particularly exaggerated for idle power consumption, motivating the need to mitigate the effects of variability in systems whose operation is dominated by long idle states with periodic active states. In systems where computation is severely limited by anemic energy reserves and where a long overall system lifetime is desired, maximizing the quality of a given application subject to these constraints is both challenging and an important step toward achieving high-quality deployments. This work describes VaRTOS, an architecture and corresponding set of operating system abstractions that provide explicit treatment of both idle and active power variations for tasks running in real-time operating systems. Tasks in VaRTOS express elasticity by exposing individual knobs —shared variables that the operating system can tune to adjust task quality and, correspondingly, task power, maximizing application utility both on a per-task and on a system-wide basis. We provide results regarding online learning of instance-specific sleep power, active power, and task-level power expenditure on simulated hardware with demonstrated effects for several prototypical applications. Our results on networked sensing applications, which are representative of a broader category of applications that VaRTOS targets, show that VaRTOS can reduce variability-induced energy expenditure errors from over 70% in many cases to under 2% in most cases and under 5% in the worst case.
Paul Martin 0008, Lucas Francisco Wanner, Mani Srivastava 0001
ACM Trans. Embed. Comput. Syst.2
2014 Distributed programming framework for fast iterative optimization in networked cyber-physical systems
abstract
Large-scale coordination and control problems in cyber-physical systems are often expressed within the networked optimization model. While significant advances have taken place in optimization techniques, their widespread adoption in practical implementations has been impeded by the complexity of internode coordination and lack of programming support for the same. Currently, application developers build their own elaborate coordination mechanisms for synchronized execution and coherent access to shared resources via distributed and concurrent controller processes. However, they typically tend to be error prone and inefficient due to tight constraints on application development time and cost. This is unacceptable in many CPS applications, as it can result in expensive and often irreversible side-effects in the environment due to inaccurate or delayed reaction of the control system. This article explores the design of a distributed shared memory (DSM) architecture that abstracts the details of internode coordination. It simplifies application design by transparently managing routing, messaging, and discovery of nodes for coherent access to shared resources. Our key contribution is the design of provably correct locality-sensitive synchronization mechanisms that exploit the spatial locality inherent in actuation to drive faster and scalable application execution through opportunistic data parallel operation. As a result, applications encoded in the proposed Hotline Application Programming Framework are error free, and in many scenarios, exhibit faster reactions to environmental events over conventional implementations. Relative to our prior work, this article extends Hotline with a new locality-sensitive coordination mechanism for improved reaction times and two tunable iteration control schemes for lower message costs. Our extensive evaluation demonstrates that realistic performance and cost of applications are highly sensitive to the prevalent deployment, network, and environmental characteristics. This highlights the importance of Hotline, which provides user-configurable options to trivially tune these metrics and thus affords time to the developers for implementing, evaluating, and comparing multiple algorithms.
Rahul Balani, Lucas Francisco Wanner, Mani Srivastava 0001
ACM Trans. Embed. Comput. Syst.2
2013 Towards analyzing and improving robustness of software applications to intermittent and permanent faults in hardware
abstract
Although a significant fraction of emerging failure and wearout mechanisms result in intermittent or permanent faults in hardware, their impact (as distinct from transient faults) on software applications has not been well studied. In this paper, we develop a distinguishing application characteristic, referred to as similarity from fundamental circuit-level understanding of the failure mechanisms. We present a mathematical definition and a procedure for similarity computation for practical software applications and experimentally verify the relationship between similarity and fault rate. Leveraging dependence of application robustness on the similarity metric, we present example architecture independent code transformations to reduce similarity and thereby the worst-case fault rate with minimal performance degradation. Our experimental results with arithmetic unit faults show as much as 74% improvement in the worst case fault rate on benchmark kernels, with less than 10% runtime penalty.
Joseph Sloan, Lucas Francisco Wanner, Salma Hosni Emam Mohamed Elmalaki, Mani Srivastava 0001, Puneet Gupta 0001
ICCD3
2013 Hardware Variability-Aware Duty Cycling for Embedded Sensors
abstract
Instance and temperature-dependent power variation has a direct impact on quality of sensing for battery-powered long-running sensing applications. We measure and characterize the active and leakage power for an ARM Cortex M3 processor and show that, across a temperature range of 20 -60, there is a 10% variation in active power, and a variation in leakage power. We introduce variability-aware duty cycling methods and a duty cycle (DC) abstraction for TinyOS which allows applications to explicitly specify the lifetime and minimum DC requirements for individual tasks, and dynamically adjusts the DC rates so that the overall quality of service is maximized in the presence of power variability. We show that variability-aware duty cycling yields a improvement in total active time over schedules based on worst case estimations of power, with an average improvement of across a wide variety of deployment scenarios based on the collected temperature traces. Conversely, datasheet power specifications fail to meet required lifetimes by 7%-15%, with an average 37 days short of the required lifetime of 1 year. Finally, we show that a target localization application using variability-aware DC yields a 50% improvement in quality of results over one based on worst case estimations of power consumption.
Lucas Francisco Wanner, Charwak Apte, Rahul Balani, Puneet Gupta 0001, Mani Srivastava 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Variability-aware duty cycle scheduling in long running embedded sensing systems
abstract
Instance and temperature-dependent leakage power variability is already a significant issue in contemporary embedded processors, and one which is expected to increase in importance with scaling of semiconductor technology. We measure and characterize this leakage power variability in current microprocessors, and show that variability aware duty cycle scheduling produces 7.1× improvement in sensing quality for a desired lifetime. In contrast, pessimistic estimations of power consumption leave 61% of the energy untapped, and datasheet power specifications fail to meet required lifetimes by 14%. Finally, we introduce a duty cycle abstraction for TinyOS that allows applications to explicitly specify lifetime and minimum duty cycle requirements for individual tasks, and dynamically adjusts duty cycle rates so that overall quality of service is maximized in the presence of power variability.
Lucas Francisco Wanner, Rahul Balani, Sadaf Zahedi, Charwak Apte, Puneet Gupta 0001, Mani Srivastava 0001
DATE1
2006 Operating Systems Portability: 8 bits and beyond
abstract
Embedded software often needs to be ported from one system to another. This may happen for a number of reasons among which are the need for using less expensive hardware or the need for extra resources. Application portability can be achieved through an architecture-independent software/hardware interface. This is not a straight-forward task in the realm of embedded systems, since they often have very specific platforms. This work shows how an application-oriented component-based operating system was developed to allow system and application portability. Case studies present two embedded applications running in different platforms, showing that application source code is totally free of architecture-dependencies
Hugo Marcondes, Arliones Hoeller, Lucas Francisco Wanner, Antônio Augusto Fröhlich
ETFA3
2006 Operating System Support for Data Acquisition in Sensor Networks
abstract
Due to modularity and heterogeneity in wireless sensor networks sensing devices, a sensor application developed for a given platform will seldom be portable to a different one, unless the run-time support systems on those platforms deliver mechanisms that abstract and encapsulate the sensor platform in an adequate manner. In this article we propose a software/hardware interface that is able to abstract families of sensing devices in an uniform fashion. We define classes of sensing devices based on their finality (e.g. sensing acceleration, sensing temperature), and establish a common substrate for each class. Each individual device in a class is able to describe itself and its properties, in a similar fashion to the IEEE 1451 standard sensors transducer electronic data sheet. A thin software layer adapts individual devices to fit the minimal requirements of its sensor class. Software-based self-description allows applications to use individual sensors' extended characteristics. We show that this strategy does not incur in excessive overhead, and presents a significant advantage with relation to solutions found in other operating system for sensor networks
Lucas Francisco Wanner, Arliones Hoeller, Augusto Born de Oliveira, Antônio Augusto Fröhlich
ETFA1