Javier Campos

dblp:13/6127 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
abstract
We present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.
Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck
ACM Trans. Reconfigurable Technol. Syst.24
2024 Reliable edge machine learning hardware for scientific applications
abstract
Extreme data rate scientific experiments create massive amounts of data that require efficient ML edge processing. This leads to unique validation challenges for VLSI implementations of ML algorithms: enabling bit-accurate functional simulations for performance validation in experimental software frameworks, verifying those ML models are robust under extreme quantization and pruning, and enabling ultra-fine-grained model inspection for efficient fault tolerance. We discuss approaches to developing and validating reliable algorithms at the scientific edge under such strict latency, resource, power, and area requirements in extreme experimental environments. We study metrics for developing robust algorithms, present preliminary results and mitigation strategies, and conclude with an outlook of these and future directions of research towards the longer-term goal of developing autonomous scientific experimentation methods for accelerated scientific discovery.
Tommaso Baldi, Javier Campos, Benjamin Hawks, Jennifer Ngadiuba, Daniel Diaz 0003, Javier M. Duarte, Ryan Kastner, Andres Meza 0001, Melissa Quinnan, Olivia Weng, Caleb Geniesse, Amir Gholami, Michael W. Mahoney, Vladimir Loncar, Philip C. Harris, Joshua Agar, Shuyu Qin
VTS2
2024 End-to-end codesign of Hessian-aware quantized neural networks for FPGAs
abstract
We develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.
Javier Campos, Jovan Mitrevski, Zhen Dong 0003, Amir Gholami, Michael W. Mahoney, Javier M. Duarte
ACM Trans. Reconfigurable Technol. Syst.1
2013 EDCA 802.11e Performance under Different Scenarios: Quantitative Analysis
abstract
The global throughput of an 802.11e WLAN is determined by EDCA (Enhanced Distributed Channel Access) parameters, among other aspects, that are usually configured with predetermined and static values. This study carefully evaluates the Quality of Service (QoS) of Wi-Fi with EDCA in several realistic scenarios with noise and a blend of wireless traffic (e.g., voice, video, and best effort, with Pareto distribution). The metrics of the benefits obtained in each case are compared, and the differentiated impact of network dynamics on each case is quantified. This study proposes a new experimental scenario based on the relative proportion of traffic present in the network. Stations have been implemented using HSANs (Hierarchical Stochastic Activity Networks) and simulated using the Möbius tool.
Santiago Pérez, Higinio Facchini, Gustavo Mercado, Luis Bisaro, Javier Campos
AINA5
2013 A Min-Max Problem for the Computation of the Cycle Time Lower Bound in Interval-Based Time Petri Nets
abstract
The time Petri net with firing frequency intervals (TPNF) is a modeling formalism used to specify system behavior under timing and frequency constraints. Efficient techniques exist to evaluate the performance of TPNF models based on the computation of bounds of performance metrics (e.g., transition throughput, place marking). In this paper, we propose a min-max problem to compute the cycle time of a transition under optimistic assumptions. That is, we are interested in computing the lower bound. We will demonstrate that such a problem is related to a maximization linear programming problem (LP-max) previously stated in the literature, to compute the throughput upper bound of the transition. The main advantage of the min-max problem compared to the LP-max is that, in addition to the optimal value, the optimal solutions provide useful feedback to the analyst on the system behavior (e.g., performance bottlenecks). We have implemented two solution algorithms, using CPLEX APIs, to solve the min-max problem, and have compared their performance using a benchmark of TPNF models, several of these being case studies. Finally, we have applied the min-max technique for the vulnerability analysis of a critical infrastructure, i.e., the Saudi Arabian crude-oil distribution network.
Simona Bernardi 0001, Javier Campos
IEEE Trans. Syst. Man Cybern. Syst.2
2011 Timing-Failure Risk Assessment of UML Design Using Time Petri Net Bound Techniques
abstract
Software systems that do not meet their timing constraints can cause risks. In this work, we propose a comprehensive method for assessing the risk of timing failure by evaluating the software design. We show how to apply best practises in software engineering and well-known Time Petri Net (TPN) modeling and analysis techniques, and we demonstrate the effectiveness of the method with reference to a case study in the domain of real-time embedded systems. The method customizes the Australian standard risk management process, where the system context is the UML-based software specification, enriched with standard MARTE profile annotations to capture nonfunctional system properties. During the risk analysis, a TPN is derived, via model transformation, from the software design specification and TPN bound techniques are applied to estimate the probability of timing failure. TPN bound techniques are also exploited, within the risk evaluation and treatment steps, to identify the risk causes in the software design.
Simona Bernardi 0001, Javier Campos, José Merseguer
IEEE Trans. Ind. Informatics2
2009 Computation of Performance Bounds for Real-Time systems using Time Petri Nets
abstract
Time Petri nets (TPNs) have been widely used for the verification and validation of real-time systems during the software development process. Their quantitative analysis consists in applying enumerative techniques that suffer the well known state space explosion problem. To overcome this problem, several methods have been proposed in the literature, that either provide rules to obtain equivalent nets with a reduced state space or avoid the construction of the whole state space. In this paper, we propose a method that consists in computing performance bounds to predict the average operational behavior of TPNs by exploiting their structural properties and by applying operational laws. Performance bound computation was first proposed for timed (Timed PNs) and stochastic Petri nets (SPNs). We generalize the results obtained for Timed PNs and SPNs to make the technique applicable to TPNs and their extended stochastic versions: TPN with firing frequency intervals (TPNFs) and extended TPNs (XTPNs). Finally, we apply the proposed bounding techniques on the case study of a robot-control application taken from the literature.
Simona Bernardi 0001, Javier Campos
IEEE Trans. Ind. Informatics2
2007 Approximate Throughput Computation of Stochastic Weighted T-Systems
abstract
A general iterative technique for approximate throughput computation of stochastic live and bounded weighted T-systems (WTS) is presented. It generalizes a previous technique on stochastic marked graphs. The approach has two basic foundations. First, a deep understanding of the qualitative behavior of WTS leads to a general decomposition technique. Second, after the decomposition phase, an iterative response-time approximation method is applied for the throughput computation. Existence of convergence points for the iterative approximation method can be proved. Experimental results generally have an error of less than 5%. The state space is usually reduced by more than one order of magnitude; therefore, the analysis of otherwise intractable systems is possible
Carlos J. Perez-Jimenez, Javier Campos, M. Silva
IEEE Trans. Syst. Man Cybern. Part A2
2004 Solving the mobile robot localization problem using string matching algorithms
abstract
In this paper we address the mobile robot localization using some techniques borrowed from the computational biology community. The specific problem studied here is also known as the kidnapped robot problem. Our proposal is to solve this problem by string matching algorithms, which have experienced a large advance in the last years due to (for example) the Genoma Project. The paper uses three different algorithms to solve the mentioned problem and shows their advantages, such as the robustness of the results and the memory and time efficiency. These results are validated by real experimentation using panoramic images of indoor buildings, and compared and discussed with existing techniques that have been used over the same test-bed.
Cándida González-Buesa, Javier Campos
IROS2
2003 Analysing Internet Software Retrieval Systems: Modeling and Performance Comparison
José Merseguer, Javier Campos, Eduardo Mena
Wirel. Networks2
2001 Performance analysis of internet based software retrieval systems using Petri Nets
abstract
Nowadays, there exist web sites that allow users to retrieve and install software in an easy way. The performance of these sites may be poor if they are used in wireless networks; the reason is the inadequate use of the net resources they need. If this kind of systems are designed using mobile agent technology the previous problem might be avoided. In this paper, we present a comparison between the performance of a software retrieval system especially designed to be used in wireless networks (e.g., mobile computers) and the performance of a software retrieval system similar to the well-known Tucows.com or Download.com web sites.
José Merseguer, Javier Campos, Eduardo Mena
MSWiM2
2000 Backlash Compensation in Discrete Time Nonlinear Systems Using Dynamic Inversion by Neural Networks
abstract
A dynamics inversion compensation scheme is designed for control of nonlinear discrete-time systems with input backlash. The compensator uses the backstepping technique with neural networks (NN) for inverting the backlash nonlinearity in the feedforward path. The technique provides a general procedure for using NN to determine the dynamics pre-inverse of an invertible discrete time dynamical system. A discrete-time tuning algorithm is given for the NN weights so that the backlash compensation scheme becomes adaptive, guaranteeing bounded tracking and backlash errors, and also bounded parameter estimates. A rigorous proof of stability and performance is given and a simulation example verifies the performance. Unlike standard discrete-time adaptive control techniques, no certainty equivalence assumption is needed.
Javier Campos, Frank L. Lewis, Rastko R. Selmic
ICRA1
1999 Deadzone compensation in discrete time using adaptive fuzzy logic
abstract
A fuzzy logic (FL) compensator is designed for control of nonlinear discrete-time systems with input deadzone. The classification property of FL systems makes them a natural candidate for the rejection of errors induced by the deadzone, which has regions in which it behaves differently. A discrete-time tuning algorithm is given for the FL parameters so that the deadzone compensation scheme becomes adaptive, guaranteeing bounded tracking errors and parameter estimates. A rigorous proof of stability and performance is given and a simulation example verifies performance. Unlike standard discrete-time adaptive control techniques, no certainty equivalence assumption is needed.
Javier Campos, Frank L. Lewis
IEEE Trans. Fuzzy Syst.1
1999 Structured Solution of Asynchronously Communicating Stochastic Modules
abstract
Asynchronously communicating stochastic modules (SAM) are Petri nets that can be seen as a set of modules that communicate through buffers, so they are not (yet another) Petri net subclass, but they complement a net with a structured view. This paper considers the problem of exploiting the compositionality of the view to generate the state space and to find the steady-state probabilities of a stochastic extension of SAM in a net-driven, efficient way. Essentially we give an expression of an auxiliary matrix, G, which is a supermatrix of the infinitesimal generator of a SAM. G is a tensor algebra expression of matrices of the size of the components for which it is possible to numerically solve the characteristic steady-state solution equation /spl pi//spl middot/G=0, without the need to explicitly compute G. Therefore, we obtain a method that computes the steady-state solution of a SAM without ever explicitly computing and storing its infinitesimal generator, and therefore without computing and storing the reachability graph of the system. Some examples of application of the technique are presented and compared to previous approaches.
Javier Campos, Susanna Donatelli, Manuel Silva 0001
IEEE Trans. Software Eng.1
1996 State machine reduction for the approximate performance evaluation of manufacturing systems modelled with cooperating sequential processes
abstract
We concentrate on a family of discrete event systems obtained from a simple modular design principle that include in a controlled way primitives to deal with concurrency, decisions, synchronization, blocking, and bulk movements of jobs. Due to the functional complexity of such systems, reliable throughput approximation algorithms must be deeply supported on a structure based decomposition technique. We present a decomposition technique and a fixed-point search iterative process based on response time preservation of subsystems. Extensive numerical experiments have shown that the error is less than 3%, and that the state space is usually reduced by more than one order of magnitude.
Carlos J. Perez-Jimenez, Javier Campos, M. Silva
ICRA2
1994 Approximate Throughput Computation of Stochastic Marked Graphs
abstract
A general iterative technique for approximate throughput computation of stochastic strongly connected marked graphs is presented. It generalizes a previous technique based on net decomposition through a single input-single output cut, allowing the split of the model through any cut. The approach has two basic foundations. First, a deep understanding of the qualitative behavior of marked graphs leads to a general decomposition technique. Second, after the decomposition phase, an iterative response time approximation method is applied for the computation of the throughput. Experimental results on several examples generally have an error of less than 3%. The state space is usually reduced by more than one order of magnitude; therefore, the analysis of otherwise intractable systems is possible.>
Javier Campos, José Manuel Colom, Hauke Jungnitz, Manuel Silva 0001
IEEE Trans. Software Eng.1
1993 Embedded Product-Form Queueing Networks and the Improvement of Performance Bounds for Petri Net Systems
Javier Campos, Manuel Silva 0001
Perform. Evaluation1
1991 Ergodicity and Throughput Bounds of Petri Nets with Unique Consistent Firing Count Vector
abstract
Ergodicity and throughput bound characterization are addressed for a subclass of timed and stochastic Petri nets, interleaving qualitative and quantitative theories. The nets considered represent an extension of the well-known subclass of marked graphs, defined as having a unique consistent firing count vector, independently of the stochastic interpretation of the net model. In particular, persistent and mono-T-semiflow net subclasses are considered. Upper and lower throughput bounds are computed using linear programming problems defined on the incidence matrix of the underlying net. The bounds proposed depend on the initial marking and the mean values of the delays but not on the probability distributions (thus including both the deterministic and the stochastic cases). From a different perspective, the considered subclasses of synchronized queuing networks; thus, the proposed bounds can be applied to these networks.>
Javier Campos, Giovanni Chiola, Manuel Silva 0001
IEEE Trans. Software Eng.1