EDBT 2026 Demo / reviewers in the wild / expert
David Castells-Rufas
dblp:23/6251
· DBLP profile ↗
15ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-7181-9705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Physics-Informed Neural Network Surrogate Model For Capacitive Touch Sensors By Solving Maxwell?s EquationsabstractCapacitive sensors based human-machine interfaces (HMI) have revolutionised in-cabin user interaction in automobiles. However, robust functionality and stability of the sensors in dynamic environmental conditions necessitate profound domain expertise coupled together with computationally intensive multi-physics simulations. As a promising alternative, this paper investigates the application of Physics-informed neural networks (PINNs) as a surrogate sensor model to capture the electrostatic behaviour of a capacitive sensor interacting with a conductor such as a human finger. Maxwell's equations are the fundamental governing laws for understanding the interplay of electric field interactions in a capacitive sensor. The PINN model solves these electrostatic equations for different positions of a finger interacting with the sensor. Given a finger position and a spatial coordinate in a 3D domain encompassing the finger, sensor, and PCB, the trained surrogate model is capable of predicting key electrostatic properties such as electric potential, electric field distribution, and charge density. The governing equations are incorporated into the neural network's loss function to capture the underlying physics. The performance of the model is evaluated on a wide range of unseen test scenarios, encompassing a diverse set of finger positions. Additionally, the capacity of the learnt PINN model to emulate a real-world sensor array setup is presented. Results demonstrate the significant potential of PINNs as surrogate models in electrostatics, thereby paving the way for a promising future in sensor design optimisation and degradation analysis. Ganyong Mo, Krishna Kumar Narayanan, David Castells-Rufas, Jordi Carrabina |
ECMS | 3 |
| 2025 | Eye side and orientation detection of iris images using lightweight textural descriptors for embedded systemsabstractIris recognition is a widely used biometric authentication technique due to its high accuracy and uniqueness. However, these systems are vulnerable to spoofing attacks, which can occur by rotating an iris image or an iris scanner during image acquisition. Additionally, correctly recognizing the eye side significantly reduces computational load and decreases the likelihood of false positives in biometric systems. This paper introduces a novel method for automatically detecting the correct left/right and upright/upside-down orientation of an iris image. The proposed method employs a lightweight feature extraction algorithm utilizing Local Binary Pattern (LBP) and Gray-Level Co-Occurrence Matrix (GLCM) to extract features from the iris image. A Support Vector Machine (SVM) classifier is then used to determine the eye side or orientation of the iris images. LBP captures the local texture details, whereas GLCM describes the statistical features of the iris images. By combining these textural features, the proposed method improves the ability to classify eye side or orientation. The efficient and precise texture descriptors allow for implementation on embedded systems, including IoT or mobile devices. Experimental results on benchmark datasets demonstrate that the proposed method outperforms existing baseline methods in both performance and speed. Chung Nguyen Tran, Jordi Carrabina, David Castells-Rufas, Minh Son Nguyen, Le-Anh Tran, Nhan Cach Dang |
IPAS | 3 |
| 2023 | GPU acceleration of Levenshtein distance computation between long stringsabstractComputing edit distance for very long strings has been hampered by quadratic time complexity with respect to string length. The WFA algorithm reduces the time complexity to a quadratic factor with respect to the edit distance between the strings. This work presents a GPU implementation of the WFA algorithm and a new optimization that can halve the elements to be computed, providing additional performance gains. The implementation allows to address the computation of the edit distance between strings having hundreds of millions of characters. The performance of the algorithm depends on the similarity between the strings. For strings longer than million characters, the performance is the best ever reported, which is above TCUPS for strings with similarities greater than 70% and above one hundred TCUPS for 99.9% similarity. David Castells-Rufas |
Parallel Comput. | 1 |
| 2022 | LoRaWAN Optimization using optimized Auto-Regressive algorithm, Support Vector Machine and Temporal Fusion Transformer for QoS ensuringabstractThe number of LoRaWAN networks have grown worldwide last years, offering a solution for the integration of the Internet of Things in rural and urban areas. After years of development, several performance issues and scalability limitations require to be enhanced for LoRa such as high collision rates and duty cycle limitations. Machine learning offers a chance for LoRaWAN to rise as the reference communication technology that offers the adequate communication performances for IoT. In this paper, our goal is to optimize the LoRaWAN network performances using detection mechanism and artificial intelli-gence to predict its behavior. first, we evaluate the full potential of the LoRaWAN factory setting, and we introduced a Quality of Service demanding application. Second, we constructed our proper database using available application criteria, we included a quality of service mechanism to simulate the effect of a new application connecting to a stable network and the perturbing causes. Then, we used two different methods one for classification, the second for prediction, and then the optimization. For clas-sification using Auto-Regressive and optimization it using burg algorithm and firefly algorithm, then we used the support vector machine for traffic classification, results are very promising, we were able to detect normal traffic, a normal surge, and an abnormal surge of network traffic with up to 99% accuracy. For prediction, we used a new algorithm developed by google Temporal Fusion Transformer. We were able to predict the network behaviour ahead with 14 days with 95% accuracy and up to 30 days with 80% accuracy. We were able to optimize the network to absorb the abnormal surge and return to normal in less than 60% of the normal time, uplifting the packet delivery ratio for uplink traffic by 20% and downlink traffic by 50%. Houssem Eddin Elbsir, Mohamed Kassab, Sami Bhiri, Mohamed Bedoui Hedi, David Castells-Rufas, Jordi Carrabina |
WiMob | 5 |
| 2022 | Continuous touch gesture recognition based on RNNs for capacitive proximity sensors
David Castells-Rufas, Juan Borrego-Carazo, Jordi Carrabina, Jordi Naqui, Ernesto Biempica |
Pers. Ubiquitous Comput. | 1 |
| 2021 | OpenCL-based FPGA Accelerator for Semi-Global Approximate String Matching Using Diagonal Bit-VectorsabstractAn FPGA accelerator for the computation of the semi-global Levenshtein distance between a pattern and a reference text is presented. The accelerator provides an important benefit to reduce the execution time of read-mappers used in short-read genomic sequencing. Previous attempts to solve the same problem in FPGA use the Myers algorithm following a column approach to compute the dynamic programming table. We use an approach based on diagonals that allows for some resource savings while maintaining a very high throughput of 1 alignment per clock cycle. The design is implemented in OpenCL and tested on two FPGA accelerators. The maximum performance obtained is 91.5 MPairs/s for 100 × 120 sequences and 47 MPairs/s for 300 × 360 sequences, the highest ever reported for this problem. David Castells-Rufas, Santiago Marco-Sola, Quim Aguado-Puig, Antonio Espinosa 0001, Juan C. Moure, Lluc Alvarez, Miquel Moretó |
FPL | 1 |
| 2021 | Speed Limits for Single-Beam Laser MarkingabstractLaser Marking is a very common part of many manufacturing processes usually performed at the last stages of production lines after the packaging of goods. The speed of production lines are key determinants of the output capacity of many industries, but their scalability could be hindered by some fundamental limits of current laser marking technologies. In this paper, we review related technologies, their limiting factors, and analyze how future systems could overcome them. David Castells-Rufas, Francesc Bravo-Montero, Jordi Carrabina |
IECON | 1 |
| 2021 | An interpretable assessment of sensor's orientation in the Quaternion domainabstractMagnetic and inertial measurement unit (MIMU) systems have been universally adopted in numerous navigation applications. These include domains such as aircraft, pedestrians, home automation, robots, etc. due to their advantages referring to price, size and accuracy. The fusion of the values recorded from magnetic and inertial sensors (magnetometer, accelerometer, gyroscope) can provide orientation with respect to the navigation path. Orientation can be given as either Euler angles or quaternions representing the rotation matrix associated with the orientation. The first is the commonest way since Euler angles can be easily interpreted in terms of yaw, pitch, and roll. However, their computation is ill-conditioned for some angulations due to a bad propagation of errors. Such intrinsic computational errors limit their use for free indoor, but equally affect the comparison and assessment of sensor fusion algorithms. In this paper, we present an assessment of orientation based on quaternion distances easy to interpret in terms of rotation axis and angle. We compare our approach to the standard assessment of orientation based on Euler angles in rotational trajectories around the three axes made using a Stäubli robotic arm. Results show the more superior reliability of the quaternion distance and the intrinsic artifacts of Euler angles for representing the whole space of rotations. Debora Gil Resina, Manuel Navarrete, Esmitt Ramírez, Carles Sánchez, Carlos García Calvo, David Castells-Rufas, Jordi Carrabina |
IPIN | 6 |
| 2020 | Capacitive-sensing module with dynamic gesture recognition for automotive applicationsabstractCapacitive sensing offers new possibilities for HMI product development. Its short range of interaction entails robustness against environmental noise and its flexibility for integration makes it a genuine technology for embedded systems. In the automotive context, capacitive sensing is explicitly devoted to driver interaction with car functionalities. However, the increasing complexity of captured signals and related interaction procedures impose severe difficulties for a classic modelling approach. Neural networks have demonstrated unbeatable performance in tasks with abundant data. Specifically, recurrent neural networks (RNNs) show excellent performance for tasks with inherent temporal structure. In this article, we develop a capacitive-sensing module that includes RNN-based dynamic gesture recognition, which has a suitable implementation size for embedded automotive applications. Juan Borrego-Carazo, David Castells-Rufas, Jordi Carrabina, Ernesto Biempica |
DDECS | 2 |
| 2020 | Extending SpArSe: Automatic Gesture Recognition Architectures for Embedded DevicesabstractNeural Architecture Search (NAS), which allows for automatically developing neural networks, has been mostly devoted to performance on a single metric, usually accuracy. New approaches have added more objectives, such as model size, in order to find networks suitable for resource-constrained platforms. SpArSe [1] is a multi-objective Bayesian optimization framework for automatically developing image classification convolutional neural networks (CNNs) for micro-controller units (MCUs). In this work, we first implement SpArSe and modify it to reduce search time, obtaining similar results regarding accuracy, model size, and maximum working memory but in less optimization time. Moreover, we extend the search space to include recurrent neural networks (RNNs) and add an inference latency objective for time-constrained tasks. Finally, we test our implementation in a gesture recognition task obtaining better results than previous manually tuned approaches for size and performance metrics, which validates the approach and its utility. Juan Borrego-Carazo, David Castells-Rufas, Jordi Carrabina, Ernesto Biempica |
ICMLA | 2 |
| 2013 | QoS-Driven Reconfigurable Parallel Computing for NoC-Based Clustered MPSoCsabstractReconfigurable parallel computing is required to provide high-performance embedded computing, hide hardware complexity, boost software development, and manage multiple workloads when multiple applications are running simultaneously on the emerging network-on-chip (NoC)-based multiprocessor systems-on-chip (MPSoCs) platforms. In these type of systems, the overall system performance may be affected due to congestion, and therefore parallel programming stacks must be assisted by quality-of-service (QoS) support to meet application requirements and to deal with application dynamism. In this paper, we present a hardware-software QoS-driven reconfigurable parallel computing framework, i.e., the NoC services, the runtime QoS middleware API and our ocMPI library and its tracing support which has been tailored for a distributed-shared memory ARM clustered NoC-based MPSoC platform. The experimental results show the efficiency of our software stack under a broad range of parallel kernels and benchmarks, in terms of low-latency interprocess communication, good application scalability, and most important, they demonstrate the ability to enable runtime reconfiguration to manage workloads in message-passing parallel applications. Jaume Joven, Akash Bagdia, Federico Angiolini, P. Strid, David Castells-Rufas, Eduard Fernandez-Alonso, Jordi Carrabina, Giovanni De Micheli |
IEEE Trans. Ind. Informatics | 5 |
| 2011 | Sharing FPUs in many-soft-coresabstractModern top of the line FPGAs can already host hundreds of simple soft-core processors. Because soft-cores often support floating point units through external interfaces this opens the door to explore the convenience for sharing the floating point units among a number of processors in many-soft-cores. We build two variants of a many-soft-core with 16 NIOSII cores to test if sharing the FPU gives an important area reduction and to test if the introduced time overhead is significant. We find out that area savings are a 30% of the non-shared FPU version for a 16 core system and that the overhead in clock cycles is almost inexistent for simple applications like matrix multiplication and below 2% for a parallel Mandelbrot application. However, if we consider the reduction of the maximum operational frequency that happens when the number of processors increase, we get that sharing among 8 processors is a very good option, and that it is not advisable to share among more than 12 processors because of the excessive time overhead. David Castells-Rufas, Eduard Fernandez-Alonso, Jordi Carrabina, Jaume Joven |
FPT | 1 |
| 2010 | Scalability of a Parallel JPEG Encoder on Shared Memory ArchitecturesabstractEmbedded multimedia systems are expected to fully embrace the future many-core wave. As a consequence parallel programming is being revamped as the only way to exploit the power of coming chips. While waiting for them we try to extrapolate some lessons learned from current multi-cores to influence future architectures and programming methods. In this paper we investigate the parallelism and scalability of a JPEG image encoder, which is a typical embedded application, on several shared memory machines using the OpenMP programming framework. We identify the Huffman coding as the bottleneck that blocks the application from scaling above a 7x factor. We propose a strategy to parallelize the Huffman coding, which introduces a small degradation in some parts of the image, allowing to reach higher speedup factors. A factor of 18.8x has been reached in SGI Altix 4700 using 22 threads. Contrasting these results with some previous works using message passing architectures we consider that the use of OpenMP on top of shared memory architectures should be reconsidered for future chips in favor of message passing architectures and programming models. David Castells-Rufas, Jaume Joven, Jordi Carrabina |
ICPP | 1 |
| 2008 | xENoC - An eXperimental Network-On-Chip Environment for Parallel Distributed Computing on NoC-based MPSoC ArchitecturesabstractThis paper describes xENoC, an automatic and component re-use HW-SW environment to build simulatable and synthesizable Network-on-Chip-based MPSoC architectures. xENoC is based on a tool, named NoCWizard, which uses an eXtensible Markup Language (XML) specification, and a set of modularized components and templates to generate many types of NoC instances by using Verilog HDL. This NoC models can be customized in terms of topology, tile location/mapping, RNIs generation, different types of routers, FIFO and packet/flit sizes, by simply modifying the XML specifications. Furthermore, xENoC is also composed of software components, i.e. RNI drivers and a parallel programming model, embedded Message Passing Interface (eMPI), which let us to carry out a complete HW-SW co-design methodology to design distributed-memory NoC-based MPSoCs parallel applications. Through xENoC different distributed-memory NoC-based MPSoCs designs have been created simulated and prototyped in physical platforms (e.g. FPGA boards), and some parallel multiprocessor test traffic applications are running there as system level demonstrators. Jaume Joven, Oriol Font-Bach, David Castells-Rufas, Lluís Terés, Jordi Carrabina |
PDP | 3 |
| 2007 | Jumble: A Hardware-in-the-Loop Simulation System for JHDLabstractThis paper presents a new verification system for FPGA based designs described in the JHDL hardware description language. The method consists of performing hardware emulation of designer selected blocks in a co-simulation environment. Although JHDL has a Hardware execution mode it does not provide a fine control of which blocks have to be executed in Hardware and it is based on Xilinx readback technology. In this paper we present a method to extend the simulation environment to add a fine control of the Hardware emulation system, a method to instrument the design for debug, and the process that automatically creates the interface to communicate the simulator with the emulated hardware block. The resulting system does not offer 100% observability and controllability of hardware blocks. Nevertheless its interactivity provides a solid basis for incremental verification while offering the possibility of substantial simulation speedups. David Castells-Rufas, Jordi Carrabina |
FCCM | 1 |