EDBT 2026 Demo / reviewers in the wild / expert
Hendrik F. Hamann
dblp:17/6773
· DBLP profile ↗
11ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0001-9049-1330ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 9Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative MethodabstractClimate science studies the structure and dynamics of Earth's climate system and seeks to understand how climate changes over time, where the data is usually stored in the format of time series, recording the climate features, geolocation, time attributes, etc. Recently, much research attention has been paid to the climate benchmarks. In addition to the most common task of weather forecasting, several pioneering benchmark works are proposed for extending the modality, such as domain-specific applications like tropical cyclone intensity prediction and flash flood damage estimation, or climate statement and confidence level in the format of natural language. To further motivate the artificial intelligence development for climate science, in this paper, we first contribute a multi-modal climate benchmark, i.e., ClimateBench-M, which aligns (1) the time series climate data from ERA5, (2) extreme weather events data from NOAA, and (3) satellite image data from NASA HLS based on a unified spatial-temporal granularity. Second, under each data modality, we also propose a simple but strong generative method that could produce competitive performance in weather forecasting, thunderstorm alerts, and crop segmentation tasks in the proposed ClimateBench-M. The data and code of ClimateBench-M are publicly available at https://github.com/iDEA-iSAIL-Lab-UIUC/ClimateBench-M. Dongqi Fu, Yada Zhu, Zhining Liu 0002, Lecheng Zheng, Xiao Lin 0016, Zihao Li 0006, Liri Fang, Katherine Tieu, Onkar Bhardwaj, Komminist Weldemariam, Hanghang Tong, Hendrik F. Hamann, Jingrui He |
CIKM | 12 |
| 2024 | AIM: Attributing, Interpreting, Mitigating Data UnfairnessabstractData collected in the real world often encapsulates historical discrimination against disadvantaged groups and individuals. Existing fair machine learning (FairML) research has predominantly focused on mitigating discriminative bias in the model prediction, with far less effort dedicated towards exploring how to trace biases present in the data, despite its importance for the transparency and interpretability of FairML. To fill this gap, we investigate a novel research problem: discovering samples that reflect biases/prejudices from the training data. Grounding on the existing fairness notions, we lay out a sample bias criterion and propose practical algorithms for measuring and countering sample bias. The derived bias score provides intuitive sample-level attribution and explanation of historical bias in data. On this basis, we further design two FairML strategies via sample-bias-informed minimal data editing. They can mitigate both group and individual unfairness at the cost of minimal or zero predictive utility loss. Extensive experiments and analyses on multiple real-world datasets demonstrate the effectiveness of our methods in explaining and mitigating unfairness. Code is available at https://github.com/ZhiningLiu1998/AIM. Zhining Liu 0002, Ruizhong Qiu, Zhichen Zeng 0001, Yada Zhu, Hendrik F. Hamann, Hanghang Tong |
KDD | 5 |
| 2023 | TensorBank: Tensor Lakehouse for Foundation Model TrainingabstractStoring and streaming high dimensional data for foundation model training became a critical requirement with the rise of foundation models beyond natural language. In this paper we introduce TensorBank – a petabyte scale tensor lakehouse capable of streaming tensors from Cloud Object Store (COS) to GPU memory at wire speed based on complex relational queries. We use Hierarchical Statistical Indices (HSI) for query acceleration. Our architecture allows to directly address tensors on block level using HTTP range reads. Once in GPU memory, data can be transformed using PyTorch transforms. We provide a generic PyTorch dataset type with a corresponding dataset factory translating relational queries and requested transformations as an instance. By making use of the HSI, irrelevant blocks can be skipped without reading them as those indices contain statistics on their content at different hierarchical resolution levels. This is an opinionated architecture powered by open standards and making heavy use of open-source technology. Although, hardened for production use using geospatial-temporal data, this architecture generalizes to other use cases like computer vision, computational neuroscience, biological sequence analysis and more. Romeo Kienzler, Johannes Schmude, Naomi Simumba, Benedikt Blumenstiel, Marcus Freitag, Daiki Kimura, Zoltan Arnold Nagy, Michael Behrendt, Hendrik F. Hamann, S. Karthik Mukkavilli, Daniel Civitarese |
IEEE Big Data | 9 |
| 2019 | Learning and Recognizing Archeological Features from LiDAR DataabstractWe present a remote sensing pipeline that processes LiDAR (Light Detection And Ranging) data through machine & deep learning for the application of archeological feature detection on big geo-spatial data platforms such as e.g. IBM PAIRS Geoscope [1], [2].Today, archeologists get overwhelmed by the task of visually surveying huge amounts of (raw) LiDAR data in order to identify areas of interest for inspection on the ground. We showcase a software system pipeline that results in significant savings in terms of expert productivity while missing only a small fraction of the artifacts.Our work employs artificial neural networks in conjunction with an efficient spatial segmentation procedure based on domain knowledge. Data processing is constraint by a limited amount of training labels and noisy LiDAR signals due to vegetation cover and decay of ancient structures. We aim at identifying geo-spatial areas with archeological artifacts in a supervised fashion allowing the domain expert to flexibly tune parameters based on her needs. Conrad M. Albrecht, Chris Fisher, Marcus Freitag, Hendrik F. Hamann, Sharath Pankanti, Florencia Pezzutti, Francesca Rossi 0001 |
IEEE BigData | 4 |
| 2019 | N-dimensional geospatial data and analytics for critical infrastructure risk assessmentabstractThe assessment of the vegetation growth rate given remote sensing data is a challenging task in the Earth Observation sciences. LiDAR data acquisition is commonly used to extract height information at a given moment in time, however, the associated cost and complexity restrict continuous acquisitions. Frequently captured aerial imagery can be used to identify and separate vegetation from bare land, water, impervious surface, or built infrastructure. A combination of LiDAR data with aerial and radar imagery allows to track dynamic seasonal growth of vegetation around critical infrastructure such as power lines. We present a general framework that integrates tree identification and growth assessment around power lines with the goal to identify locations of high risk where trees potentially cause power outages. Levente J. Klein, Conrad M. Albrecht, Carlo Siebenschuh, Sharath Pankanti, Hendrik F. Hamann, Siyuan Lu 0003 |
IEEE BigData | 6 |
| 2017 | Event clustering & event series characterization on expected frequencyabstractWe present an efficient clustering algorithm applicable to one-dimensional data such as e.g. a series of times-tamps. Given an expected frequency ΔT-1, we introduce an O(N)-efficient method of characterizing N events represented by an ordered series of timestamps t1, t2,..., tN. In practice, the method proves useful to e.g. identify time intervals of missing data or to locate isolated events. Moreover, we define measures to quantify a series of events by varying ΔT to e.g. determine the quality of an Internet of Things service. Conrad M. Albrecht, Marcus Freitag, Theodore G. van Kessel, Siyuan Lu 0003, Hendrik F. Hamann |
IEEE BigData | 5 |
| 2017 | A low maintenance particle pollution sensing system using the Minimum Airflow Particle Counter (MAPC)abstractThe Minimum Airflow Particle Counter (MAPC) is a portable, low-power, low-cost, wireless optical counter which has been specifically designed for ultra-low-maintenance operation in heavily polluted environments. When exposed continuously to air with high particulate matter concentrations, the primary mode of failure for particle counters is a build-up of dust within the instrument. The MAPC circumvents this failure mode by severely restricting airflow through the system, enabling an estimated 5-year maintenance cycle. Such a long operational lifetime makes this instrument particularly suitable for IOT applications such as environmental air quality monitoring and pollutant source attribution using spatially distributed wireless sensor networks. Here, we present the theory of operation, instrument design, and collected data from a two-month field deployment in Beijing. We find that the MAPC performs comparably to other low-cost optical counters, but with a significantly enhanced maintenance-free operational lifetime. Theodore G. van Kessel, Ramachandran Muralidhar, Josephine B. Chang, Jun-Song Wang, Michael A. Schappert, Hendrik F. Hamann |
IEEE BigData | 6 |
| 2017 | Distributed wireless sensing for fugitive methane leak detectionabstractLarge scale environmental monitoring requires dynamic optimization of data transmission, power management, and distribution of the computational load. In this work, we demonstrate the use of a wireless sensor network for detection of chemical leaks on gas oil well pads. The sensor network consist of chemi-resistive and wind sensors and aggregates all the data and transmits it to the cloud for further analytics processing. The sensor network data is integrated with an inversion model to identify leak location and quantify leak rates. We characterize the sensitivity and accuracy of such system under multiple well controlled methane release experiments. It is demonstrated that even 1 hour measurement with 10 sensors localizes leaks within 1 m and determines leak rate with an accuracy of 40%. This integrated sensing and analytics solution is currently refined to be a robust system for long term remote monitoring of methane leaks, generation of alarms, and tracking regulatory compliance. Levente J. Klein, Theodore G. van Kessel, Dhruv Nair, Ramachandran Muralidhar, Nigel Hinds, Hendrik F. Hamann, Norma E. Sosa |
IEEE BigData | 6 |
| 2016 | IBM PAIRS curated big data service for accelerated geospatial data analytics and discoveryabstractIBM's Physical Analytics Integrated Data Repository and Services (PAIRS) is a geospatial Big Data service. PAIRS contains a massive amount of curated geospatial (or more precisely spatio-temporal) data from a large number of public and private data resources, and also supports user contributed data layers. PAIRS offers an easy-to-use platform for both rapid assembly and retrieval of geospatial datasets or performing complex analytics, lowering time-to-discovery significantly by reducing the data curation and management burden. In this paper, we review recent progress with PAIRS and showcase a few exemplary analytical applications which the authors are able to build with relative ease leveraging this technology. Siyuan Lu 0003, Xiaoyan Shao, Marcus Freitag, Levente J. Klein, Jason D. Renwick, Fernando J. Marianno, Conrad M. Albrecht, Hendrik F. Hamann |
IEEE BigData | 8 |
| 2016 | Solar irradiance forecasting by machine learning for solar car racesabstractSolar car race competitions offer realistic conditions to test and demonstrate the state-of-the-art technologies in multidisciplinary fields. In such races the solar panels mounted on the car produce the energy required to power the vehicle. A simulator runs during the race determines the optimal race speed based on the predicted availability of solar energy and other parameters as well as road conditions. The accuracy of the forecasts, especially the solar irradiance forecasts, has a significant impact on the race strategy. Here we report on the experience of providing irradiance forecasts for two races run by the University of Michigan Solar Car Team at the Bridgestone World Solar Challenge 2015 in Australia and at the American Solar Challenge 2016 from Ohio to South Dakota. The probabilistic forecasts of hourly solar irradiance generated from machine learning algorithms were deployed to optimally decide on the race strategy. This work showcases an example of real time decision making based on insights derived from machine learning utilizing big geospatial data — weather models and measurement data from weather station networks. Xiaoyan Shao, Siyuan Lu 0003, Theodore G. van Kessel, Hendrik F. Hamann, Leda Daehler, Jeffrey Cwagenberg, Alan Li |
IEEE BigData | 4 |
| 2015 | PAIRS: A scalable geo-spatial data analytics platformabstractGeospatial data volume exceeds hundreds of Petabytes and is increasing exponentially mainly driven by images/videos/data generated by mobile devices and high resolution imaging systems. Fast data discovery on historical archives and/or real time datasets is currently limited by various data formats that have different projections and spatial resolution, requiring extensive data processing before analytics can be carried out. A new platform called Physical Analytics Integrated Repository and Services (PAIRS) is presented that enables rapid data discovery by automatically updating, joining, and homogenizing data layers in space and time. Built on top of open source big data software, PAIRS manages automatic data download, data curation, and scalable storage while being simultaneously a computational platform for running physical and statistical models on the curated datasets. By addressing data curation before data being uploaded to the platform, multi-layer queries and filtering can be performed in real time. In addition, PAIRS offers a foundation for developing custom analytics. Towards that end we present two examples with models which are running operationally: (1) high resolution evapo-transpiration and vegetation monitoring for agriculture and (2) hyperlocal weather forecasting driven by machine learning for renewable energy forecasting. Levente J. Klein, Fernando J. Marianno, Conrad M. Albrecht, Marcus Freitag, Siyuan Lu 0003, Nigel Hinds, Xiaoyan Shao, Sergio Bermudez Rodriguez, Hendrik F. Hamann |
IEEE BigData | 9 |