Marisa López-Vallejo

dblp:l/MarisaLopezVallejo · also Marisa Luisa López-Vallejo, María Luisa López Vallejo · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-3833-524XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Ultra-Low-Power Flip-Flop With Near-Threshold Robust Operation and Redundant-Free Internal Clock Transitions
abstract
Optimizing the power consumption of flip-flops (FFs), as basic building blocks of sequential digital circuits, can substantially reduce the power of the whole system. This work proposes a new Robust Redundant-Free Flip-Flop (RRFF) for ultra-low-power purposes that addresses both aspects that have the potential to reduce power consumption: it has static and contention-free characteristics that provide the robustness required for supply voltage scalability down to the near-threshold region and, at the same time, it eliminates all internal redundant transitions, so that the circuit exclusively consumes dynamic energy when data changes. Measurement results from a test chip fabricated in a 65-nm CMOS technology show that the RRFF achieves 63.4% clock power reduction across a wide voltage range (0.4-1.1 V) and 15.0% power reduction at 100% activity for the same supply voltage range with minimal area overhead compared to the conventional transmission gate flip-flop (TGFF).
Borja Gutiérrez De Cabiedes, Javier De Mena Pacheco, Amadeo de Gracia Herranz, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Uniformly Distributed CORDIC
abstract
This paper presents the uniformly distributed (UD) coordinate rotation digital computer (CORDIC). It receives this name because the distance between its rotation angles is uniform. The paper presents two versions of the UD CORDIC: the binary UD CORDIC and the canonical signed digit (CSD) UD CORDIC. Both simplify the control part of the CORDIC by obtaining the micro-rotation directions straightforwardly from the binary representation of the rotation angle, with the particularity that, in the CSD UD CORDIC, the binary representation is transformed first into a CSD number. Compared to previous CORDIC approaches that simplify the control part, the proposed rotators further simplify the micro-rotation stages. This is achieved by using more efficient angles for the micro-rotations in the binary UD CORDIC and by merging consecutive micro-rotation stages in the CSD UD CORDIC. Thus, the simplification of the control part of the rotators, the careful selection of the micro-rotation angles, and their optimized shift-and-add implementation results in CORDIC architectures that require the smallest number of adders among pipelined CORDIC implementations for a given resolution so far. Additionally, experimental results of the proposed approach on field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) are provided, showing improvements with respect to previous approaches in multiple figures of merit, mostly area, clock frequency, and power consumption.
Mario Garrido, Daniel Medina, Pedro Paz, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A Highly Power- and Area-Efficient PMU for Cell-Size Autonomous Microsystems
abstract
Power Management Units (PMU) are present in most electronic devices today. However, for power and area constrained applications, their design becomes a real challenge, especially when cell-sized autonomous microsystems are targeted. This paper presents an LDO regulator and a voltage reference, which are usually a must in any PMU, for autonomous microsystems fabricated in 65 nm CMOS technology. This PMU has been designed to minimize quiescent power and area, consuming a minimum power of 15.6 nW and occupying only$391~\mu m^{2}$, while being able to deliver up to$40~\mu A$to the load. No significant degradation was observed, as the measured line regulation is 0.173%/V, the temperature coefficient is 128.4 ppm/°C over a very wide temperature range going from 0 °C up to 120 °C, and the PSR at low frequencies is −67.1 dB.
Javier De Mena Pacheco, Juan M. Carrillo, Tomás Palacios, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Serial Butterflies for Non-Power-of-Two FFT Architectures in 5G and Beyond
abstract
This paper presents new serial butterflies for non-power-of-two (NP2) fast Fourier transform (FFT) architectures. The paper considers radices 2, 3, 4, and 5, which are used in FFTs for 5G systems. Current designs for non-power-of-two FFTs are mostly based on the single-path delay feedback (SDF) architecture. This type of architecture processes data arriving in series. However, it uses butterflies with several parallel inputs. This results in low utilization, as the butterflies have to wait for all the inputs before they start to process them. Conversely, the proposed approach allows to calculate the butterflies on data that arrive in series. This removes waiting times and reduces the number of hardware components such as multipliers and adders. As a result, the proposed butterflies achieve high performance and provide a significant reduction in area and power consumption with respect to parallel butterflies. Thus, they are an efficient solution when data must be processed in series in the butterflies.
Víctor Manuel Bautista, Mario Garrido, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Reference-free power supply monitor with enhanced robustness against process and temperature variations
abstract
Power supply noise in current nanometer technologies represents a growing risk, specially because of the uncertainties it produces in the critical paths delays which can result in erroneous computations. Also, these very short variations can affect the functionality of analog circuits like a comparator or an ADC. To tackle with these issues and to have a better power management, on-line power supply monitors have become one common solution. Traditional approaches use an external reference, are very sensitive to temperature and process variations or present high latency. In this work we propose a 3 sigma monitor with a novel detector circuitry that employs a feedback loop that works without an external reference and it is hardened against temperature and process variations. The sensor was designed in the 40 nm CMOS technology node, operating at 1.1 V and has been validated for a temperature range of 0∘C to 85∘C covering all process corners. Also, considering the 3σ confidence, the sensor is able to detect over/undershoots voltage fluctuations up to 2 GHz in the range of -70 mV to +90 mV from the nominal voltage with a maximum latency of 1.1 ns and an energy consumption per measurement of 0.8 pJ.
Hernan Aparicio, Pablo Ituero, Marisa López-Vallejo
Integr.3
2022 Power-Efficient Implementation of Ternary Neural Networks in Edge Devices
abstract
There is a growing interest in pushing computation to the edge, especially the problem-solving abilities of artificial neural networks (ANNs). This article presents a simplified method to obtain a ternary neural network based on the multilayer perceptron. The method is focused on resource-constrained devices, where memory, computing power, and battery are some of the most relevant constraints. A dynamic threshold is estimated to perform ternarization, and a new pruning technique is proposed to obtain a drastic reduction in the ANN’s size, with the corresponding decrease in resource utilization and power consumption of the resulting hardware. In addition, a support framework has been developed to automate hardware design exploration and generation from the network trained in software. Experimental results show that the proposed method and architecture, when implemented in a field-programmable gate array (FPGA), provide excellent figures in power (0.11–0.13 W) and efficiency (1225–1448 kfps/W) with respect to state of the art, being its efficiency double than the maximum one reported previously.
Miguel Molina, Javier Mendez 0001, Diego Pedro Morales, Encarnación Castillo, Marisa López-Vallejo, Manuel P. Cuéllar
IEEE Internet Things J.5
2020 Time-domain writing architecture for multilevel RRAM cells resilient to temperature and process variations
Amadeo de Gracia Herranz, Marisa López-Vallejo
Integr.2
2019 A low power RFID based energy harvesting temperature resilient CMOS-only reference voltage
Asghar Bahramali, Marisa López-Vallejo
Integr.2
2017 Reconfigurable Writing Architecture for Reliable RRAM Operation in Wide Temperature Ranges
abstract
Resistive switching memories [resistive RAM (RRAM)] are an attractive alternative to nonvolatile storage and nonconventional computing systems, but their behavior strongly depends on the cell features, driver circuit, and working conditions. In particular, the circuit temperature and writing voltage schemes become critical issues, determining resistive switching memories performance. These dependencies usually force a design time tradeoff among reliability, device endurance, and power consumption, thereby imposing nonflexible functioning schemes and limiting the system performance. In this paper, we present a writing architecture that ensures the correct operation no matter the working temperature and allows the dynamic load of application-oriented writing profiles. Thus, taking advantage of more efficient configurations, the system can be dynamically adapted to overcome RRAM intrinsic challenges. Several profiles are analyzed regarding power consumption, temperature-variations protection, and operation speed, showing speedups near 700× compared with other published drivers.
Fernando García-Redondo, Pablo Royer, Marisa López-Vallejo, Hernan Aparicio, Pablo Ituero, Carlos A. López-Barrio
IEEE Trans. Very Large Scale Integr. Syst.3
2017 A 4096-Point Radix-4 Memory-Based FFT Using DSP Slices
abstract
This brief presents a novel 4096-point radix-4 memory-based fast Fourier transform (FFT). The proposed architecture follows a conflict-free strategy that only requires a total memory of size N and a few additional multiplexers. The control is also simple, as it is generated directly from the bits of a counter. Apart from the low complexity, the FFT has been implemented on a Virtex-5 field programmable gate array (FPGA) using DSP slices. The goal has been to reduce the use of distributed logic, which is scarce in the target FPGA. With this purpose, most of the hardware has been implemented in DSP48E. As a result, the proposed FPGA is efficient in terms of hardware resources, as is shown by the experimental results.
Mario Garrido, Miguel Angel Sánchez, Marisa López-Vallejo, Jesús Grajal
IEEE Trans. Very Large Scale Integr. Syst.3
2016 A temperature-independent PUF with a configurable duty cycle of CMOS ring oscillators
abstract
In this work we propose a Ring Oscillator PUF focused on the variability of the duty cycle instead of measuring the output frequency deviations. To achieve this goal, we replace the common ring oscillators, whose outputs are clock signals of 50% duty cycle, for ring oscillators with an asymmetric structure. The asymmetry confers the ability to configure the duty cycle of each individual node. Through the measurement of a relative value, such as the duty cycle, the robustness of the PUF is improved. For example, the output shift due to the temperature variation is decreased from 3% to less than 0,5%. Moreover, the potential input challenges are multiplied by the number of stages of each ring oscillator. Hence, with our design, the number of ring oscillators needed to build a robust PUF is decreased thanks to the addition of multiple and uncorrelated variables but with negligible area overhead.
Javier Agustin, Marisa López-Vallejo
ISCAS2
2016 A Performance Study of CUDA UVM versus Manual Optimizations in a Real-World Setup: Application to a Monte Carlo Wave-Particle Event-Based Interaction Model
abstract
The performance of a Monte Carlo model for the simulation of electromagnetic wave propagation in particle-filled atmospheres has been conducted for different CUDA versions and design approaches. The proposed algorithm exhibits a high degree of parallelism, which allows favorable implementation in a GPU. Practical implementation aspects of the model have been also explained and their impact assessed, such as the use of the different types of memories present in a GPU. A number of setups have been chosen in order to compare performance for manually optimized versus Unified Virtual Memory (UVM ) implementations for different CUDA versions. Features and relative performance impact of the different options have been discussed, extracting practical hints and rules useful to speed up CUDA programs.
Jose M. Nadal-Serrano, Marisa López-Vallejo
IEEE Trans. Parallel Distributed Syst.2
2015 A thermal adaptive scheme for reliable write operation on RRAM based architectures
abstract
Resistive RAMs (RRAMs) are one of the most promising alternatives to future storage and neuromorphic computing systems. However, the behavior of RRAM highly depends on voltage, crossbar design and operation temperature. Actually, the circuit temperature becomes one of the most critical issues in fast memories during writing operations. In this paper we propose a novel thermal-adaptive RRAM writing scheme, applicable to crossbar memories, whose smart operation is able to mitigate the writing errors induced by temperature variations. Using a sensing-acting scheme our system is able to improve the memory reliability without affecting the writing/reading performance. Moreover, the proposed architecture is compatible with most proposed write/read designs making achievable multibit storage, which requires extremely accurate operations.
Fernando García-Redondo, Marisa López-Vallejo, Pablo Ituero
ICCD2
2013 A Low-Area Reference-Free Power Supply Sensor
abstract
Power supply unpredictable fluctuations jeopardize the functioning of several types of current electronic systems. This work presents a power supply sensor based on a voltage divider followed by buffer-comparator cells employing just MOSFET transistors and providing a digital output. The divider outputs are designed to change more slowly than the thresholds of the comparators, in this way the sensor is able to detect voltage droops. The sensor is implemented in a 65nm technology node occupying an area of 2700 μm2and displaying a power consumption of 50 uW. It is designed to work with no voltage reference and with no clock and aiming to obtain a fast response.
Carlos Benito, Pablo Ituero, Marisa López-Vallejo
DSD3
2013 A low power 6t-SRAM using negative bit-line for variability tolerance beyond 22nm node
abstract
The expected large variations of electrical characteristics of sub-22 nm devices represents a limitation on future electronic circuits. This is particularly relevant on RAM memories that have to ensure both read stability and write ability of all the cells. In this paper we present a 6T-SRAM designed with 14nm FinFETs that makes use of the negative bit-line write assist technique allowing a reduction of the supply voltage without a degradation on neither speed nor stability. A new metric is introduced to quantify and control the drawbacks related to negative bit-line voltage. Additionally a new approach to predict the tails of the read current distribution under variability has been presented. Experimental results show that power consumption is reduced by 25% due to the decrease on the supply voltage.
Pablo Royer, Marisa López-Vallejo
ACM Great Lakes Symposium on VLSI2
2013 Floating-Point Exponentiation Units for Reconfigurable Computing
abstract
The high performance and capacity of current FPGAs makes them suitable as acceleration co-processors. This article studies the implementation, for such accelerators, of the floating-point power function x y as defined by the C99 and IEEE 754-2008 standards, generalized here to arbitrary exponent and mantissa sizes. Last-bit accuracy at the smallest possible cost is obtained thanks to a careful study of the various subcomponents: a floating-point logarithm, a modified floating-point exponential, and a truncated floating-point multiplier. A parameterized architecture generator in the open-source FloPoCo project is presented in details and evaluated.
Florent de Dinechin, Pedro Echeverría, Marisa López-Vallejo, Bogdan Pasca 0001
ACM Trans. Reconfigurable Technol. Syst.3
2012 Improving Hardware Reuse through XML-based Interface Encapsulation
Miguel Angel Sánchez, Marisa López-Vallejo, Carlos Angel Iglesias, Carlos A. López-Barrio
ICECCS2
2012 Hardware Reuse Improvement through the Domain Specific Language dHDL
abstract
The dHDL language has been defined to improve hardware design productivity. This is achieved through the definition of a better reuse interface (including parameters, attributes and macroports) and the creation of control structures that help the designer in the hardware generation process.
Miguel Angel Sánchez, Marisa López-Vallejo, Carlos Angel Iglesias
ISPA2
2011 On-chip Monitoring: A Light-Weight Interconnection Network Approach
abstract
Current nanometer technologies are subjected to several adverse effects that seriously impact the yield and performance of integrated circuits. Such is the case of within-die parameters uncertainties, varying workload conditions, aging, temperature, etc. Monitoring, calibration and dynamic adaptation have appeared as promising solutions to these issues and many kinds of monitors have been presented recently. In this scenario, where systems with hundreds of monitors of different types have been proposed, the need for light-weight monitoring networks has become essential. In this work we present a light-weight network architecture based on digitization resource sharing of nodes that require a time-to-digital conversion. Our proposal employs a single wire interface, shared among all the nodes in the network, and quantizes the time domain to perform the access multiplexing and transmit the information. It supposes a 16% improvement in area and power consumption compared to traditional approaches.
Pablo Ituero, Marisa López-Vallejo, Miguel A. Sánchez Marcos, Carlos Gómez Osuna
DSD2
2008 Designing Highly Parameterized Hardware using xHdl
abstract
Current submicron technologies allow a very high degree of integration, resulting in incredibly complex designs implemented in a single chip. The task of the hardware designer is every day more complicated since the data and parameters involved in a design are wider and with increasing complexity in the interfaces. Modular design is needed to deal with these large, regular and repetitive structures. With this purpose the meta-language xHDL was conceived, providing flexible and friendly mechanisms for component parameterization, customization, instantiation and interconnection. In this paper two case studies will be analyzed in depth to illustrate the advantages of using xHDL. Based on the lessons learnt when specifying the case studies the meta-language has been extended to deal with new advanced features such as instantiation of external VHDL components and automatic generation of libraries of components.
Miguel Angel Sánchez, Pedro Echeverría, Francisco Mansilla, Marisa López-Vallejo
FDL4
2008 Joint hardware-software leakage minimization approach for the register file of VLIW embedded architectures
David Atienza 0001, Praveen Raghavan, José Luis Ayala, Giovanni De Micheli, Francky Catthoor, Diederik Verkest, Marisa López-Vallejo
Integr.7
2007 Leakage-based On-Chip Thermal Sensor for CMOS Technology
abstract
Thermal characterization of ICs and on-chip temperature monitoring have become key tasks in electronic engineering. In this paper, we present the design of an on-chip CMOS temperature sensor based on the temperature dependent characteristics of the subthreshold current. The proposed sensor achieves high accuracy sensing (0.56°C maximum error), wide temperature range (25-90°C), and extremely low area (0.010 mm2) and power overhead (18μW). Our approach improves previous works on on-chip temperature sensors and is highly suitable for portable applications where temperature monitoring achieves great importance.
Pablo Ituero, José Luis Ayala, Marisa López-Vallejo
ISCAS3
2007 Reduction of Register File Delay Due to Process Variability in VLIW Embedded Processors
abstract
Process variation in future technologies can cause severe performance degradation since different parts of the shared register file (RF) in VLIW processors may operate at various speeds. In this paper we present a complete approach that handles speed variability of the RF proposing different compile-time and run-time design alternatives. The first alternative extends current RF architectures and uses a compile-time variability-aware register assignment algorithm. The second alternative presents a fully-adjustable pure run-time approach, which overcomes the variability loss as well, but at the extra cost of cycles and area. However, the savings achieved and the run-time management of the register delay variations without any support from the user, show a very promising application field. Our results in embedded system benchmarks show that variability can be tackled without significant performance penalty, and trade-offs between performance and area are possible thanks to the whole design spectrum provided by the two presented alternatives.
Praveen Raghavan, José Luis Ayala, David Atienza 0001, Francky Catthoor, Giovanni De Micheli, Marisa López-Vallejo
ISCAS6
2006 New Schemes in Clustered VLIW Processors Applied to Turbo Decoding
abstract
State-of-the-art communication standards make extensive use of Turbo codes. The complex and power consuming designs that currently implement the turbo decoder expose the need for innovative solutions. In recent years the area of application specific processors has attracted the attention of the research community and important advances have been made possible. This work introduces an ASIP architecture for SISO Turbo decoding based on a dual-clustered VLIW processor. The machine deals with instructions of up to 21 operands in an innovative way, the fetching and asserting of data is serialized and the addressing is automatized and transparent for the programmer. An optimized architecture is achieved, flexible enough to comply with leading edge standards and adaptable to demanding hardware constraints.
Pablo Ituero, Marisa López-Vallejo
ASAP2
2006 Automated design space exploration of FPGA-based FFT architectures based on area and power estimation
abstract
In this paper a tool aimed at generating fast Fourier transform (FFT) cores targeting FPGA platforms was presented. The tool is able to generate different pipelined architectures of the FFT that provide different points of the design space: from high performance to low area implementations. The user can select the most suitable architecture based on a broad set of configuration parameters, as they are the number of points, sample size, truncation, etc. Moreover, a set of accurate estimators has been implemented to allow the designer an early and quick design space exploration before synthesizing the core. Experimental results validate our approach and provide significant measurements about the accuracy of the estimation and the tool execution time
Miguel A. Sánchez Marcos, Mario Garrido, Marisa López-Vallejo, Carlos A. López-Barrio
FPT3
2004 Improving IP core reuse through the application of the meta-language xHDL
Marcos M. A. Sanchez, Fernandez Herrero, Marisa López-Vallejo
FDL3
2003 Energy Aware Register File Implementation through Instruction Predecode
abstract
The register file is a power-hungry device in modern architectures. Current research on compiler technology and computer architectures encourages the implementation of larger devices to feed multiple data paths and to store global variables. However, low power techniques are not able to appreciably reduce power consumption in this device without a time penalty. We introduce an efficient hardware approach to reduce the register file energy consumption by turning unused registers into a low power state. Bypassing the register fields of the fetch instruction to the decode stage allows the identification of registers required by the current instruction (instruction predecode) and allows the control logic to turn them back on. They are put into the low-power state after the instruction use. This technique achieves an 85% energy reduction with no performance penalty. The simplicity of the approach makes it an effective low-power technique for embedded processors.
José Luis Ayala, Marisa López-Vallejo, Alexander V. Veidenbaum, Carlos A. Lopez
ASAP2
2003 An Efficient Hash Table Based Approach to Avoid State Space Explosion in History Driven Quasi-Static Scheduling
Antonio G. Lomeña, Marisa López-Vallejo, Yosinori Watanabe, Alex Kondratyev
DATE2
2003 On the hardware-software partitioning problem: System modeling and partitioning techniques
abstract
This paper presents an in-depth study of several system partitioning procedures. It is based on the appropriate formulation of a general system model, being therefore independent of either the particular co-design problem or the specific partitioning procedure. The techniques under study are a knowledge-based system and three classical circuit partitioning algorithms (Simulated Annealing, Kernighan&Lin and Hierarchical Clustering). The former has been entirely proposed by the authors in previous works while the later have been properly extended to deal with system level issues. We will show how the way the problem is solved biases the results obtained, regarding both quality and convergence rate. Consequently it is extremely important to choose the most suitable technique for the particular co-design problem that is being confronted.
Marisa López-Vallejo, Juan Carlos López 0001
ACM Trans. Design Autom. Electr. Syst.1
2002 Block processing technique for low power turbo decoder design
abstract
We apply a block processing technique to the MAP algorithm used in turbo decoding. This new technique leads to a power-efficient way to access memory and to a reduced memory size. We introduce the "very long data word" (VLDW) memory architecture, which leads to a reduction in power consumption for memory access operations. The proposed architecture provides a low power implementation of the turbo decoder.
Inkyu Lee, Marisa López-Vallejo, Syed Aon Mujtaba
VTC Spring2
2000 Constraint-Driven System Partitioning
abstract
This paper describes how optimization techniques can be applied to efficiently solve the constrained co-design problem. This is performed by the formulation of different cost functions which will drive the hardware-software partitioning process. The use of complex cost functions allows us to capture more aspects of the design. Besides, the appropriate formulation of this kind of functions has a great impact on the results that can be obtained regarding both quality and algorithm convergence rate. A strong point of the proposed formulation is its generality. Therefore, it does not depend on the problem and can be easily extended for considering new design constraints.
Marisa López-Vallejo, Jesús Grajal, Juan Carlos López 0001
DATE1
1999 Hardware-Software Partitioning at the Knowledge Level
Marisa López-Vallejo, Juan Carlos López 0001, Carlos Angel Iglesias
Appl. Intell.1
1998 A Knowledge-based System for Hardware-Software Partitioning
abstract
This paper presents SHAPES, a tool for hardware-software partitioning. It is based on two main paradigms: the implementation of the partitioning tool by means of an expert system, and the use of fuzzy logic to model the parameters involved in the process.
Marisa López-Vallejo, Carlos Angel Iglesias, Juan Carlos López 0001
DATE1