Pedro Julián

dblp:53/2396 · also Pedro M. Julian · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-6308-4497ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-author · 3 since 2021Computer networks · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Partner Project: A Deep Learning Platform Targeting Embedded Hardware for Edge-AI Applications (NEUROKIT2E)
abstract
The goal of the NEUROKIT2E project is to create an open-source Deep Learning framework for edge and embedded AI built around an established European value chain. This framework, called AIDGE, supports a wide range of application areas that operate independently and serve a global user community. It provides easy and fast full-stack solutions from Neural Network design and optimization to AI application development all the way down to hardware implementations while enabling code generation for application-specific targets. This platform provides flexibility for academic users in the AI domain to explore and innovate while allowing them the possibility to prototype systems, ensuring their work aligns well with industrial needs. This paper presents the results and achievements of the first part of this three-year project, along with its roadmap and expected outcomes.
Rajendra Bishnoi, Mohammad Amin Yaldagard, Said Hamdioui, Kanishkan Vadivel, Manolis Sifalakis, Nicolás Rodríguez 0002, Pedro Julián, Lothar Ratschbacher, Maen Mallah, Yogesh Ramesh Patil, Fabian Chersi
DATE7
2023 System on Chip Testbed for Deep Neuromorphic Neural Networks
abstract
This paper describes a first prototype of a testbed System on chip (SoC) to design and evaluate different Neuromorphic Deep Neural Networks (NN) cores. The$1.25mm\times 1.25mm$SoC was fabricated in a 65nm CMOS technology and implements a system composed of an ARM based microprocessor, two memory banks of 32KB, a QSPI serial interface and two NN accelerators. The first one is a novel neuromorphic accelerator consisting of a$5\times 5$kernel Symmetrical Simplicial (SymSimp) core with a depthwise separable structure, which allows to efficiently implement multi-channel convolutional layers by breaking 3D kernels into 2D kernels. The second is a 3×3 conventional MAC engine to implement the fully connected layers. Experimental results show an energy efficiency of 0.49pJ/OP, which is competitive when compared to similar technology ICs, and extrapolated to the MobileNetworkV2 ImageNet represents a factor of 2 improvement with respect to NVIDIA Jetson Nano.
Nicolás Rodríguez 0002, Martin Villemur, Daniel Klepatsch, Diego Gigena Ivanovich, Pedro Julián
ISCAS5
2022 Embedded Processing Pipeline Exploration For Neuromorphic Event Based Perceptual Systems
abstract
Event-based vision cameras emulate the functionality of mamalian retina and promise to be a low-latency, energy efficient sensory front-end for machine perception. Despite the large-scale effort to deploy these sensors in a variety of scenarios, a proportionally small amount of effort has been devoted to the design and analysis of embedded architectures that process address events adjacent to the sensor. In this paper, a neuromorphic signal processing pipeline is reported which sparsifies the event stream thereby reducing energy consumption, increasing the signal-to-noise ratio, and improving downstream algorithm performance. It is integrated within a system-on-chip platform that will allow for the prototyping of different standards compliant, hardware modules within a embedded processing framework. We report two such modules which provides adaptive throughput management, spatiotemporal filtering, and programmable feature extraction.
Jonah Sengupta, Martin Villemur, Philippe O. Pouliquen, Pedro Julián, Andreas G. Andreou
ISCAS4
2018 Neuromorphic Cellular Neural Network Processor for Intelligent Internet-of-Things
abstract
We discuss the architecture, implementation and testing of a neuromorphic Cellular Neural Network (CNN) processor for intelligent IoT devices. The processor is based on a simplicial piecewise linear CNN architecture that allows implementation of linear and nolinear CNNs. A linear array of 64 processing element (PE) with column-shared computation resources, tightly coupled to two data memory caches was synthesized and fabricated in a 55nm CMOS technology using custom layout libraries. The fabricated chip achieves an overall performance of 2.95 TOPS/W with dynamic energy dissipation efficiency of 86.4fJ per OP at V=500mV. The processor can implement different types of processing on 2D data arrays, such as gray-scale morphology, gradient flow, median filters, and approximate Gaussian filters, among others.
Martin Villemur, Pedro Julián, Tomas Figliolia, Andreas G. Andreou
ISCAS2
2016 A true Random Number Generator using RTN noise and a sigma delta converter
abstract
The design of a true Bernoulli Random Number Generator (RNG) source with a true probability p = 0.5 is a challenging problem. In this work, we present a novel design of a True RNG (TRNG) that achieves a true E(p(n)) = 0.5. The architecture is based on the perturbation of a Sigma-Delta modulator using random telegraph noise (RTN).
Tomas Figliolia, Pedro Julián, Gaspar Tognetti, Andreas G. Andreou
ISCAS2
2010 PWL cores for nonlinear array processing
abstract
This paper presents an analysis of different alternatives for the realization of a VLSI cell in a nonlinear neuronal array, based on a simplicial piecewise linear (PWL) operation. Depending on the type of existing design constraints, namely, speed or density, different bus sizes can be used to broadcast the parameters stored in the memory, and in addition, row and column operations can be serialized. Based on a 90nm technology process, the different options will be analyzed and compared using simulations.
Martin Di Federico, Pedro Julián, Pablo Sergio Mandolesi, Andreas G. Andreou
ISCAS2
2007 A Simplicial PWL Integrated Circuit Realization
abstract
In this paper we present a mixed-signal integrated circuit in a standard CMOS 0.5 µm technology implementing a piecewise-linear (PWL) function with three inputs, where each input can be either analog or coded with 8 bits. The output of the circuit is a digital word with 8-bit precision, representing the value of the PWL function at the three-dimensional input. The circuit accesses also a 4 kB external memory, which is addressed with a 12-bit word. Experimental results are shown that demonstrate the circuit working up to 50 MHz with a maximum power consumption of 3.7 mW.
Martin Di Federico, Pedro Julián, Tomaso Poggi, Marco Storace
ISCAS2
2007 An Adaptive Cross-Correlation Derivative Algorithm for Ultra-Low Power Time Delay Measurement
abstract
In this paper, we report a low power integrated circuit that implements an adaptive version of the cross-correlation derivative algorithm for the estimation of inter-aural time difference. The architecture and logic structure as well as measured results reporting the performance of the IC -fabricated in a standard CMOS 0.5μm process- are shown.
F. N. Martin Pirchio, Pedro Julián, Pablo Sergio Mandolesi, Alfonso Chacón-Rodríguez
ISCAS2
2007 Bounded state space particle filter for network sensors
abstract
An estimation filter is presented based on the particle filter framework, which is suitable for detection of rare events-applications in sensor networks. This application has the property of low power consumption and implies low computation and RF transmission capabilities. For this reason a particle filter is implemented in each node, without the re-sampling stage; only bounds on the particles with high weights are transmitted. Experimental results are presented demonstrating the performance of the proposed scheme.
Silvana Sanudo, Favio R. Masson, Pedro Julián
ISCAS3
2006 A simplicial CNN visual processor in 3D SOI-CMOS
abstract
This paper presents the architecture for a SIMD digital visual processor unit (VPU) that is based on the simplicial CNN (S-CNN) algorithm. The system is designed for three dimensional CMOS integration in the three tier MITLL 3D SOI-CMOS 0.18 mum technology. The architecture includes input/output sub-systems, in the third tier, arithmetic logic units (ALU) and register files on the third and second tiers and instruction cache memory and a timing state machine on the first tier. The partition of the architecture exploits its physical realization in three dimensional CMOS. Parallel optical data input through an array of photodetectors and analog interface circuits in the third tier facilitate testing and characterization
Pablo Sergio Mandolesi, Pedro Julián, Andreas G. Andreou
ISCAS2
2006 VLSI implementation of an energy-aware wake-up detector for an acoustic surveillance sensor network
abstract
We present a low-power VLSI wake-up detector for a sensor network that uses acoustic signals to localize ground-based vehicles. The detection criterion is the degree of low-frequency periodicity in the acoustic signal, and the periodicity is computed from the “bumpiness” of the autocorrelation of a one-bit version of the signal. We then describe a CMOS ASIC that implements the periodicity estimation algorithm. The ASIC is fully functional and its core consumes 835 nanowatts. It was integrated into an acoustic enclosure and deployed in field tests with synthesized sounds and ground-based vehicles.
David H. Goldberg, Andreas G. Andreou, Pedro Julián, Philippe O. Pouliquen, Laurence Riddle, Rich Rosasco
ACM Trans. Sens. Networks3
2006 A low-power correlation-derivative CMOS VLSI circuit for bearing estimation
abstract
We present a CMOS integrated circuit (IC) for bearing estimation in the low-audio range that performs a correlation derivative approach in a 0.35-/spl mu/m technology. The IC calculates the bearing angle of a sound source with a mean variance of one degree in a 360/spl deg/ range using four microphones: one pair is used to produce the indication and the other to define the quadrant. An adaptive algorithm decides which pair to use depending on the direction of the incoming signal, in such a way to obtain the best estimate. The IC contains two blocks with 104 stages each. Every stage has a delay unit, a block to reduce the clock speed, and a 10-bit UP/DN counter. The IC measures 2 mm by 2.4 mm, and dissipates 600 /spl mu/W at 3.3 V and 200 kHz. It is purely digital and uses a one-bit quantization of the input signals.
Pedro Julián, Andreas G. Andreou, David H. Goldberg
IEEE Trans. Very Large Scale Integr. Syst.1
2004 A wake-up detector for an acoustic surveillance sensor network: algorithm and VLSI implementation
abstract
We describe a low-power VLSI wake-up detector for use in an acoustic surveillance sensor network. The detection criterion is based on the degree of low-frequency periodicity in the acoustic signal. To this end, we have developed a periodicity estimation algorithm that maps particularly well to a low-power VLSI implementation. The time-domain algorithm is based on the "bumpiness" of the autocorrelation of one-bit version of the signal. We discuss the relationship of this algorithm to the maximum-likelihood estimator for periodicity. We then describe a full-custom CMOS ASIC that implements this algorithm. This ASIC is fully functional and its core consumes 835 nano-Watts. The ASIC was integrated into an acoustic enclosure and tested outdoors on synthesized sounds. This unit was also deployed in a three-node sensor network and tested on ground-based vehicles.
David H. Goldberg, Andreas G. Andreou, Pedro Julián, Philippe O. Pouliquen, Laurence Riddle, Rich Rosasco
IPSN3
2002 The simplicial neural cell and its mixed-signal circuit implementation: an efficient neural-network architecture for intelligent signal processing in portable multimedia applications
abstract
This paper introduces a novel neural architecture which is capable of similar performance to any of the "classic" neural paradigms while having a very simple and efficient mixed-signal implementation which makes it a valuable candidate for intelligent signal processing in portable multimedia applications. The architecture and its realization circuit are described and the functional capabilities of the novel neural architecture called a simplicial neural cell are demonstrated for both regression and classification problems including nonlinear image filtering.
Radu Dogaru, Pedro Julián, Leon O. Chua, Manfred Glesner
IEEE Trans. Neural Networks2
2000 A model reduction procedure for high level canonical PWL functions
abstract
The canonical expression introduced previously has the degrees of freedom which are necessary to represent any PWL function defined over a compact rectangular set of arbitrary dimension, subdivided by a simplicial boundary configuration with grid step /spl delta/. However, in those cases where the domain dimension is large and the grid size is small, the number of parameters of the resulting canonical expression is, in general, very large. Accordingly, the purpose of this paper is to present a method to reduce the number of parameters of a given PWL function f, based on the use of orthonormal bases of PWL functions.
Pedro Julián, Belén D'Amico, Alfredo C. Desages
ISCAS1