EDBT 2026 Demo / reviewers in the wild / expert
Cecilia Hernández
dblp:83/281
· DBLP profile ↗
18ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-1301-6987ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streaming algorithm and hardware accelerator for high-throughput entropy estimation of network flows in sliding windows
Yaime Fernández, Javier E. Soto, Carolina Gallardo-Pavesi, Yasmany Prieto, Cecilia Hernández, Miguel E. Figueroa |
Comput. Commun. | 5 |
| 2026 | Estimating the compressibility of raster data
Martita Muñoz, José Fuentes-Sepúlveda, Cecilia Hernández, Diego Seco Naveiras |
Inf. Syst. | 3 |
| 2025 | A streaming algorithm and hardware accelerator for top-K flow detection in network trafficabstractIdentifying the largest K flows in network traffic is an important task for applications such as flow scheduling and anomaly detection, which aim to improve network efficiency and security. However, accurately estimating flow frequencies is challenging due to the large number of flows and increasing network speeds. Hardware accelerators are often used in this endeavor due to their high computational power, but their limited amount of on-chip memory constrains their performance. Various sketch-based algorithms have been proposed to estimate properties of traffic such as frequency, with lower memory usage and theoretical bounds, but they often under perform with the skewed distribution of network traffic. In this work, we propose an algorithm for top- K identification using a modified TowerSketch and a priority queue array. Tested on real traffic traces, we identify the top- K flows, with K up to 32,768, with a precision of more than 0.94, and estimate their frequency with an average relative error under $1.96 \%$. We designed and implemented an accelerator for this algorithm on an AMD Virtex U280 UltraScale+ FPGA, which processes one packet per cycle at 392 MHz, reaching a minimum line rate of more than 200 Gbps. Carolina Gallardo-Pavesi, Yaime Fernández, Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
DSD | 4 |
| 2025 | Clustering-based compression for raster time seriesabstractAbstract A raster time series is a sequence of independent rasters arranged chronologically covering the same geographical area. These are commonly used to depict the temporal evolution of represented variables. The $T$-$k^{2}$-raster is a compact data structure that performs very well in practice for compact representations for raster time series. This structure classifies each raster as a snapshot or a log and encodes logs concerning their reference snapshots, which are the immediately preceding selected snapshots. An enhanced version of the $T$-$k^{2}$-raster, called Heuristic $T$-$k^{2}$-raster, incorporates a heuristic for automating the selection of snapshots. In this study, we investigate the optimality of the heuristic employed in Heuristic $T$-$k^{2}$-raster by comparing it with a dynamic programming (DP) approach. Our experimental evaluation demonstrates that Heuristic $T$-$k^{2}$-raster is a near-optimal solution, achieving compression performance almost identical to the DP method. These results indicate that variations of the structure that maintain the temporal order of the rasters are unlikely to significantly improve compression. Consequently, we explore an alternative approach based on clustering, where rasters are grouped according to their similarity, regardless of their temporal order. Our experimental evaluation reveals that this clustering-based strategy can enhance compression in scenarios characterized by cyclic behaviour. Martita Muñoz, José Fuentes-Sepúlveda, Cecilia Hernández, Gonzalo Navarro 0001, Diego Seco Naveiras, Fernando Silva-Coira |
Comput. J. | 3 |
| 2024 | A Hardware Accelerator for Quantile Estimation of Network Packet AttributesabstractMeasuring statistical properties of network traffic can improve our understanding of traffic distribution and help us detect short and long-term anomalies. However, computing the exact value of these properties requires significant storage and computation, which limits their application in high-speed networks. Hardware accelerators provide the computational power to process a large sequence of network packets with high throughput and low latency, but their performance is ultimately limited by the amount of on-chip memory available on the device. Consequently, researchers have proposed sketch-based algorithms to estimate properties of a data stream with sub linear memory and theoretical estimation error bounds. In this paper, we present a streaming algorithm and hardware accelerator for quantile estimation, which is based on the architecture of the KLL sketch. Implemented on an AMD Virtex XCU55 UltraScale+ FPGA, the accelerator operates at a clock frequency of 356 MHz, thereby achieving a minimum line rate of 182 Gbps and a maximum estimation latency of 4.33 µs. When processing a set of 10 real traffic traces of up to 123 million packets, the accelerator estimates 1000 packet-size quantiles per trace with a median error of 0.39% or less, and a maximum error of 1.3% or less across all traces. Carolina Gallardo-Pavesi, Yaime Fernández, Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
DSD | 4 |
| 2023 | A Sketch-Based Algorithm for Network-Flow Entropy Estimation on Programmable Switches Using P4abstractThe empirical Shannon entropy is a popular metric for anomaly detection in network traffic. However, computing its exact value in real time requires fast access to a large number of counters, which is unfeasible in high-speed networks. Approximate approaches using sketches can estimate the entropy with low memory usage. However, achieving good estimation accuracy still requires large data structures, making their implementation difficult in dedicated hardware and programable switches. In this paper, we present an entropy-estimation algorithm and its implementation in a programmable switch, which achieves good accuracy for large traffic traces with low memory usage. The algorithm uses sketches to track the packet count of only the most-frequent flows and models the rest of the traffic with a uniform distribution. The implementation operates within the restrictions imposed by the P4 switch programming language, achieving a 1.72% average estimation error on 12 real-world large traces from public repositories. Javier E. Soto, Sofía Vera, Yaime Fernández, Daniel Yunge, Cecilia Hernández, Miguel E. Figueroa |
DSD | 5 |
| 2023 | A streaming algorithm and hardware accelerator to estimate the empirical entropy of network flows
Yaime Fernández, Javier E. Soto, Sofía Vera, Yasmany Prieto, Cecilia Hernández, Miguel E. Figueroa |
Comput. Networks | 5 |
| 2023 | JACC-FPGA: A hardware accelerator for Jaccard similarity estimation using FPGAs in the cloud
Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
Future Gener. Comput. Syst. | 2 |
| 2021 | Compact structure for sparse undirected graphs based on a clique graph partition
Felipe Glaria, Cecilia Hernández, Susana Ladra, Gonzalo Navarro 0001, Lilian Salinas |
Inf. Sci. | 2 |
| 2020 | A hardware accelerator for entropy estimation using the top-k most frequent elementsabstractEstimating the empirical entropy of the elements in a dataset is an important task in data analysis. In particular, empirical entropy can be effectively used to detect anomalies in network traffic. However, computing the empirical entropy of a large dataset is computationally expensive and requires a large amount of memory. This is particularly important in high-speed network traffic analysis, where computing the entropy of a data flow in real time requires using hardware accelerators with restricted on-chip memory and arithmetic resources. In this work, we propose a method to estimate the entropy using a streaming algorithm with sublinear space requirements. Our approach uses a sketch to estimate the frequency of the elements in the stream, and a priority queue to store the top-k most frequent elements. We show that our method can provide a good approximation of the entropy of the dataset, and present the design of a hardware accelerator that can compute the entropy of the stream with a throughput of one packet per clock cycle. Implemented on a Xilinx Zynq UltraScale + MPSoC ZCU102 FPGA, our accelerator can operate at line rates above 181 Gbps, consuming 511 mW and using less than 24% of the resources available on the device. Javier E. Soto, Paulo Ubisse, Cecilia Hernández, Miguel E. Figueroa |
DSD | 3 |
| 2019 | Hardware Acceleration of k-Mer Clustering using Locality-Sensitive HashingabstractClustering is an essential operation in many data analysis applications. In particular, bioinformatics and genome analysis use clustering to group similar components in sequence data, in order to find important patterns such as DNA motifs. In this paper, we present an algorithm that clusters DNA data using locality-sensitive hashing with MinHash to group similar subsequences in large Chip-seq datasets. Tested on a standard mESC dataset, the algorithm builds clusters that contain subsequences with high-score matches to known DNA motifs. We also describe the architecture and implementation of a hardware accelerator on a Xilinx Kintex-7 XC7K325T FPGA, that exploits the parallelism of the algorithm to cluster data with a throughput of one k-mer per clock cycle at 350MHz. The accelerator achieves a speedup of 91 compared to a parallel software implementation of the algorithm on a 24-core server. Javier E. Soto, Thomas Krohmer, Cecilia Hernández, Miguel E. Figueroa |
DSD | 3 |
| 2018 | Heavy-Hitter Detection Using a Hardware Sketch with the Countmin-CU AlgorithmabstractWe present a custom hardware architecture for fast heavy hitter detection in large data streams. The architecture probabilistically estimates the frequency of each element in the data stream using the Countmin-CU sketch with the H3 family of hash functions. The sketch is stored in on-chip memory, and the architecture exploits the parallelism available in the data by simultaneously processing each row of the sketch. The hash functions map each element to a set of counters on the sketch, and the sketch increments the counters that hold the minimum value, which corresponds to the estimated frequency of the element. The hash functions and sorting network are implemented in hardware as fully pipelined circuits, in order to maximize their operating clock frequency. We show a prototype of the architecture running on a Xilinx Kintex-7 XC7K325T FPGA operating with a 300MHz clock, which can process a stream of 3,982,496 32-bit elements and detect the heavy hitters with an 4x16,384-element sketch in 13.27ms, achieving a speedup of 768 compared to a modern desktop computer. Antonio Saavedra, Cecilia Hernández, Miguel E. Figueroa |
DSD | 2 |
| 2014 | Compressed representations for web and social graphs
Cecilia Hernández, Gonzalo Navarro 0001 |
Knowl. Inf. Syst. | 1 |
| 2013 | Discovering Dense Subgraphs in Parallel for Compressing Web and Social Networks
Cecilia Hernández, Mauricio Marín |
SPIRE | 1 |
| 2013 | An Evaluation of Microwave Land Surface Emissivities Over the Continental United States to Benefit GPM-Era Precipitation AlgorithmsabstractPassive microwave (PMW) satellite-based precipitation over land algorithms rely on physical models to define the most appropriate channel combinations to use in the retrieval, yet typically require considerable empirical adaptation of the model for use with the satellite measurements. Although low-frequency channels are better suited to measure the emission due to liquid associated with rain, most techniques to date rely on high-frequency, scattering-based schemes since the low-frequency methods are limited to the highly variable land surface background, whose radiometric contribution is substantial and can vary more than the contribution of the rain signal. Thus, emission techniques are generally useless over the majority of the Earth's surface. As a first step toward advancing to globally useful physical retrieval schemes, an intercomparison project was organized to determine the accuracy and variability of several emissivity retrieval schemes. A three-year period (July 2004-June 2007) over different targets with varying surface characteristics was developed. The PMW radiometer data used includes the Special Sensor Microwave Imagers, SSMI Sounder, Advanced Microwave Scanning Radiometer (AMSR-E), Tropical Rainfall Measuring Mission (TRMM) Microwave Imager (TMI), Advanced Microwave Sounding Units, and Microwave Humidity Sounder, along with land surface model emissivity estimates. Results from three specific targets in North America were examined. While there are notable discrepancies among the estimates, similar seasonal trends and associated variability were noted. Because of differences in the treatment surface temperature in the various techniques, it was found that comparing the product of temperature and emissivity yielded more insight than when comparing the emissivity alone. This product is the major contribution to the overall signal measured by PMW sensors and, if it can be properly retrieved, will improve the utility of emission techniques for over land precipitation retrievals. As a more rigorous means of comparison, these emissivity time series were analyzed jointly with precipitation data sets, to examine the emissivity response immediately following rain events. The results demonstrate that while the emissivity structure can be fairly well characterized for certain surface types, there are other more complex surfaces where the underlying variability is more than can be captured with the PMW channels. The implications for Global Precipitation Measurement-era algorithms suggest that physical retrievals are feasible over vegetated land during the warm seasons. Ralph Ferraro, Christa D. Peters-Lidard, Cecilia Hernández, F. Joseph Turk, Filipe Aires, Catherine Prigent, Sid-Ahmed Boukabara, Fumie A. Furuzawa, Kaushik Gopalan, Kenneth W. Harrison, Fatima Karbou, Chuntao Liu, Hirohiko Masunaga, Leslie Moy, Sarah E. Ringerud, Gail M. Skofronick-Jackson, Yudong Tian, Nai-Yu Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2012 | Compressed Representation of Web and Social Networks via Dense Subgraphs
Cecilia Hernández, Gonzalo Navarro 0001 |
SPIRE | 1 |
| 2008 | A P2P Meta-index for Spatio-temporal Moving Object Databases
Cecilia Hernández, M. Andrea Rodríguez, Mauricio Marín |
DASFAA | 1 |
| 2008 | Complex Queries for Moving Object Databases in DHT-Based Systems
Cecilia Hernández, M. Andrea Rodríguez, Mauricio Marín |
Euro-Par | 1 |