Hendrik F. Hamann

dblp:17/6773 · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-9049-1330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 since 2021Databases, data management, data science and information retrieval · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Systems, architecture and hardware · 3Computer networks · 3 · 1 first-author
YearPublicationVenuePosition
2026 Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
abstract
Yuanchen Bei, Tianxin Wei, Xuying Ning, Yanjun Zhao, Zhining Liu, Xiao Lin, Yada Zhu, Hendrik Hamann, Jingrui He, Hanghang Tong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuanchen Bei, Tianxin Wei, Xuying Ning, Zhining Liu 0002, Xiao Lin 0016, Yada Zhu, Hendrik F. Hamann, Jingrui He, Hanghang Tong
ACL (1)8
2025 ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method
abstract
Climate science studies the structure and dynamics of Earth's climate system and seeks to understand how climate changes over time, where the data is usually stored in the format of time series, recording the climate features, geolocation, time attributes, etc. Recently, much research attention has been paid to the climate benchmarks. In addition to the most common task of weather forecasting, several pioneering benchmark works are proposed for extending the modality, such as domain-specific applications like tropical cyclone intensity prediction and flash flood damage estimation, or climate statement and confidence level in the format of natural language. To further motivate the artificial intelligence development for climate science, in this paper, we first contribute a multi-modal climate benchmark, i.e., ClimateBench-M, which aligns (1) the time series climate data from ERA5, (2) extreme weather events data from NOAA, and (3) satellite image data from NASA HLS based on a unified spatial-temporal granularity. Second, under each data modality, we also propose a simple but strong generative method that could produce competitive performance in weather forecasting, thunderstorm alerts, and crop segmentation tasks in the proposed ClimateBench-M. The data and code of ClimateBench-M are publicly available at https://github.com/iDEA-iSAIL-Lab-UIUC/ClimateBench-M.
Dongqi Fu, Yada Zhu, Zhining Liu 0002, Lecheng Zheng, Xiao Lin 0016, Zihao Li 0006, Liri Fang, Katherine Tieu, Onkar Bhardwaj, Komminist Weldemariam, Hanghang Tong, Hendrik F. Hamann, Jingrui He
CIKM12
2025 Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting
abstract
Time-series forecasting plays a critical role in many real-world applications. Although increasingly powerful models have been developed and achieved superior results on benchmark datasets, through a fine-grained sample-level inspection, we find that (i) no single model consistently outperforms others across different test samples, but instead (ii) each model excels in specific cases. These findings prompt us to explore how to adaptively leverage the distinct strengths of various forecasting models for different samples. We introduce TimeFuse, a framework for collective time-series forecasting with sample-level adaptive fusion of heterogeneous models. TimeFuse utilizes meta-features to characterize input time series and trains a learnable fusor to predict optimal model fusion weights for any given input. The fusor can leverage samples from diverse datasets for joint training, allowing it to adapt to a wide variety of temporal patterns and thus generalize to new inputs, even from unseen datasets. Extensive experiments demonstrate the effectiveness of TimeFuse in various long-/short-term forecasting tasks, achieving near-universal improvement over the state-of-the-art individual models. Code is available at https://github.com/ZhiningLiu1998/TimeFuse.
Zhining Liu 0002, Xiao Lin 0016, Ruizhong Qiu, Tianxin Wei, Yada Zhu, Hendrik F. Hamann, Jingrui He, Hanghang Tong
ICML7
2025 CLIMB: Class-imbalanced Learning Benchmark on Tabular Data
abstract
Class-imbalanced learning (CIL) on tabular data is important in many real-world applications where the minority class holds the critical but rare outcomes. In this paper, we present CLIMB, a comprehensive benchmark for class-imbalanced learning on tabular data. CLIMB includes 73 real-world datasets across diverse domains and imbalance levels, along with unified implementations of 29 representative CIL algorithms. Built on a high-quality open-source Python package with unified API designs, detailed documentation, and rigorous code quality controls, CLIMB supports easy implementation and comparison between different CIL algorithms. Through extensive experiments, we provide practical insights on method accuracy and efficiency, highlighting the limitations of naive rebalancing, the effectiveness of ensembles, and the importance of data quality. Our code, documentation, and examples are available at https://github.com/ZhiningLiu1998/imbalanced-ensemble.
Zhining Liu 0002, Zihao Li 0006, Tianxin Wei, Jian Kang 0008, Yada Zhu, Hendrik F. Hamann, Jingrui He, Hanghang Tong
NeurIPS7
2024 Climatic & Anthropogenic Hazards to the Nasca World Heritage: Application of Remote Sensing, AI, and Flood Modelling
abstract
Preservation of the Nasca geoglyphs at the UNESCO World Heritage Site in Peru is urgent as natural and human impact accelerates. More frequent weather extremes such as flash-floods threaten Nasca geoglyphs. We demonstrate that runoff models based on (sub-)meter scale, LiDAR-derived digital elevation data can highlight AI-detected geoglyphs that are in danger of erosion. We recommend measures of mitigation to protect the famous “lizard”, “tree”, and “hand” geoglyphs located close by, or even cut by the Pan-American Highway.
Masato Sakai, Marcus Freitag, Akihisa Sakurai, Conrad M. Albrecht, Hendrik F. Hamann
IGARSS5
2024 AIM: Attributing, Interpreting, Mitigating Data Unfairness
abstract
Data collected in the real world often encapsulates historical discrimination against disadvantaged groups and individuals. Existing fair machine learning (FairML) research has predominantly focused on mitigating discriminative bias in the model prediction, with far less effort dedicated towards exploring how to trace biases present in the data, despite its importance for the transparency and interpretability of FairML. To fill this gap, we investigate a novel research problem: discovering samples that reflect biases/prejudices from the training data. Grounding on the existing fairness notions, we lay out a sample bias criterion and propose practical algorithms for measuring and countering sample bias. The derived bias score provides intuitive sample-level attribution and explanation of historical bias in data. On this basis, we further design two FairML strategies via sample-bias-informed minimal data editing. They can mitigate both group and individual unfairness at the cost of minimal or zero predictive utility loss. Extensive experiments and analyses on multiple real-world datasets demonstrate the effectiveness of our methods in explaining and mitigating unfairness. Code is available at https://github.com/ZhiningLiu1998/AIM.
Zhining Liu 0002, Ruizhong Qiu, Zhichen Zeng 0001, Yada Zhu, Hendrik F. Hamann, Hanghang Tong
KDD5
2024 Temporal Graph Neural Tangent Kernel with Graphon-Guaranteed
abstract
_Graph Neural Tangent Kernel_ (GNTK) fuses graph neural networks and graph kernels, simplifies the process of graph representation learning, interprets the training dynamics of graph neural networks, and serves various applications like protein identification, image segmentation, and social network analysis. In practice, graph data carries complex information among entities that inevitably evolves over time, and previous static graph neural tangent kernel methods may be stuck in the sub-optimal solution in terms of both effectiveness and efficiency. As a result, extending the advantage of GNTK to temporal graphs becomes a critical problem. To this end, we propose the temporal graph neural tangent kernel, which not only extends the simplicity and interpretation ability of GNTK to the temporal setting but also leads to rigorous temporal graph classification error bounds. Furthermore, we prove that when the input temporal graph grows over time in the number of nodes, our temporal graph neural tangent kernel will converge in the limit to the _graphon_ NTK value, which implies the transferability and robustness of the proposed kernel method, named **Temp**oral **G**raph **N**eural **T**angent **K**ernel with **G**raphon-**G**uaranteed or **Temp-G$^3$NTK**. In addition to the theoretical analysis, we also perform extensive experiments, not only demonstrating the superiority of Temp-G$^3$NTK in the temporal graph classification task, but also showing that Temp-G^3NTK can achieve very competitive performance in node-level tasks like node classification compared with various SOTA graph kernel and representation learning baselines. Our code is available at https://github.com/kthrn22/TempGNTK.
Katherine Tieu, Dongqi Fu, Yada Zhu, Hendrik F. Hamann, Jingrui He
NeurIPS4
2023 TensorBank: Tensor Lakehouse for Foundation Model Training
abstract
Storing and streaming high dimensional data for foundation model training became a critical requirement with the rise of foundation models beyond natural language. In this paper we introduce TensorBank – a petabyte scale tensor lakehouse capable of streaming tensors from Cloud Object Store (COS) to GPU memory at wire speed based on complex relational queries. We use Hierarchical Statistical Indices (HSI) for query acceleration. Our architecture allows to directly address tensors on block level using HTTP range reads. Once in GPU memory, data can be transformed using PyTorch transforms. We provide a generic PyTorch dataset type with a corresponding dataset factory translating relational queries and requested transformations as an instance. By making use of the HSI, irrelevant blocks can be skipped without reading them as those indices contain statistics on their content at different hierarchical resolution levels. This is an opinionated architecture powered by open standards and making heavy use of open-source technology. Although, hardened for production use using geospatial-temporal data, this architecture generalizes to other use cases like computer vision, computational neuroscience, biological sequence analysis and more.
Romeo Kienzler, Johannes Schmude, Naomi Simumba, Benedikt Blumenstiel, Marcus Freitag, Daiki Kimura, Zoltan Arnold Nagy, Michael Behrendt, Hendrik F. Hamann, S. Karthik Mukkavilli, Daniel Civitarese
IEEE Big Data9
2019 Learning and Recognizing Archeological Features from LiDAR Data
abstract
We present a remote sensing pipeline that processes LiDAR (Light Detection And Ranging) data through machine & deep learning for the application of archeological feature detection on big geo-spatial data platforms such as e.g. IBM PAIRS Geoscope [1], [2].Today, archeologists get overwhelmed by the task of visually surveying huge amounts of (raw) LiDAR data in order to identify areas of interest for inspection on the ground. We showcase a software system pipeline that results in significant savings in terms of expert productivity while missing only a small fraction of the artifacts.Our work employs artificial neural networks in conjunction with an efficient spatial segmentation procedure based on domain knowledge. Data processing is constraint by a limited amount of training labels and noisy LiDAR signals due to vegetation cover and decay of ancient structures. We aim at identifying geo-spatial areas with archeological artifacts in a supervised fashion allowing the domain expert to flexibly tune parameters based on her needs.
Conrad M. Albrecht, Chris Fisher, Marcus Freitag, Hendrik F. Hamann, Sharath Pankanti, Florencia Pezzutti, Francesca Rossi 0001
IEEE BigData4
2019 N-dimensional geospatial data and analytics for critical infrastructure risk assessment
abstract
The assessment of the vegetation growth rate given remote sensing data is a challenging task in the Earth Observation sciences. LiDAR data acquisition is commonly used to extract height information at a given moment in time, however, the associated cost and complexity restrict continuous acquisitions. Frequently captured aerial imagery can be used to identify and separate vegetation from bare land, water, impervious surface, or built infrastructure. A combination of LiDAR data with aerial and radar imagery allows to track dynamic seasonal growth of vegetation around critical infrastructure such as power lines. We present a general framework that integrates tree identification and growth assessment around power lines with the goal to identify locations of high risk where trees potentially cause power outages.
Levente J. Klein, Conrad M. Albrecht, Carlo Siebenschuh, Sharath Pankanti, Hendrik F. Hamann, Siyuan Lu 0003
IEEE BigData6
2018 Closed Loop Controlled Precision Irrigation Sensor Network
abstract
A closed loop irrigation system is demonstrated that fully automates the delivery of irrigation and calculates the water requirement from satellite images. The system optimizes water delivery for 140 cells located across four hectares of land based on two independent objectives (e.g., maximizing yield and increasing water efficiency) and is continuously adapting irrigation scheduling to the local spatial–temporal variability of the vegetation across the growing season. Irrigation is controlled by a central computer that issues commands to 693 control nodes to start irrigation based on the analysis of satellite images. The control nodes are laid out to create 15 m$\times15$m cells and each cell can be addressed independently and can irrigate differentially. After two years of operation, this variable rate drip irrigation approach resulted in a 26% yield increase in the second year and an average increase of 16% in water use efficiency. This paper demonstrates that combining closed loop automation and advanced irrigation analytics can improve the water use efficiency and increase the yield on existing agricultural lands.
Levente J. Klein, Hendrik F. Hamann, Nigel Hinds, Supratik Guha, Luis Sanchez 0001, Brent S. Sams, Nick Dokoozlian
IEEE Internet Things J.2
2017 Event clustering & event series characterization on expected frequency
abstract
We present an efficient clustering algorithm applicable to one-dimensional data such as e.g. a series of times-tamps. Given an expected frequency ΔT-1, we introduce an O(N)-efficient method of characterizing N events represented by an ordered series of timestamps t1, t2,..., tN. In practice, the method proves useful to e.g. identify time intervals of missing data or to locate isolated events. Moreover, we define measures to quantify a series of events by varying ΔT to e.g. determine the quality of an Internet of Things service.
Conrad M. Albrecht, Marcus Freitag, Theodore G. van Kessel, Siyuan Lu 0003, Hendrik F. Hamann
IEEE BigData5
2017 A low maintenance particle pollution sensing system using the Minimum Airflow Particle Counter (MAPC)
abstract
The Minimum Airflow Particle Counter (MAPC) is a portable, low-power, low-cost, wireless optical counter which has been specifically designed for ultra-low-maintenance operation in heavily polluted environments. When exposed continuously to air with high particulate matter concentrations, the primary mode of failure for particle counters is a build-up of dust within the instrument. The MAPC circumvents this failure mode by severely restricting airflow through the system, enabling an estimated 5-year maintenance cycle. Such a long operational lifetime makes this instrument particularly suitable for IOT applications such as environmental air quality monitoring and pollutant source attribution using spatially distributed wireless sensor networks. Here, we present the theory of operation, instrument design, and collected data from a two-month field deployment in Beijing. We find that the MAPC performs comparably to other low-cost optical counters, but with a significantly enhanced maintenance-free operational lifetime.
Theodore G. van Kessel, Ramachandran Muralidhar, Josephine B. Chang, Jun-Song Wang, Michael A. Schappert, Hendrik F. Hamann
IEEE BigData6
2017 Distributed wireless sensing for fugitive methane leak detection
abstract
Large scale environmental monitoring requires dynamic optimization of data transmission, power management, and distribution of the computational load. In this work, we demonstrate the use of a wireless sensor network for detection of chemical leaks on gas oil well pads. The sensor network consist of chemi-resistive and wind sensors and aggregates all the data and transmits it to the cloud for further analytics processing. The sensor network data is integrated with an inversion model to identify leak location and quantify leak rates. We characterize the sensitivity and accuracy of such system under multiple well controlled methane release experiments. It is demonstrated that even 1 hour measurement with 10 sensors localizes leaks within 1 m and determines leak rate with an accuracy of 40%. This integrated sensing and analytics solution is currently refined to be a robust system for long term remote monitoring of methane leaks, generation of alarms, and tracking regulatory compliance.
Levente J. Klein, Theodore G. van Kessel, Dhruv Nair, Ramachandran Muralidhar, Nigel Hinds, Hendrik F. Hamann, Norma E. Sosa
IEEE BigData6
2016 IBM PAIRS curated big data service for accelerated geospatial data analytics and discovery
abstract
IBM's Physical Analytics Integrated Data Repository and Services (PAIRS) is a geospatial Big Data service. PAIRS contains a massive amount of curated geospatial (or more precisely spatio-temporal) data from a large number of public and private data resources, and also supports user contributed data layers. PAIRS offers an easy-to-use platform for both rapid assembly and retrieval of geospatial datasets or performing complex analytics, lowering time-to-discovery significantly by reducing the data curation and management burden. In this paper, we review recent progress with PAIRS and showcase a few exemplary analytical applications which the authors are able to build with relative ease leveraging this technology.
Siyuan Lu 0003, Xiaoyan Shao, Marcus Freitag, Levente J. Klein, Jason D. Renwick, Fernando J. Marianno, Conrad M. Albrecht, Hendrik F. Hamann
IEEE BigData8
2016 Solar irradiance forecasting by machine learning for solar car races
abstract
Solar car race competitions offer realistic conditions to test and demonstrate the state-of-the-art technologies in multidisciplinary fields. In such races the solar panels mounted on the car produce the energy required to power the vehicle. A simulator runs during the race determines the optimal race speed based on the predicted availability of solar energy and other parameters as well as road conditions. The accuracy of the forecasts, especially the solar irradiance forecasts, has a significant impact on the race strategy. Here we report on the experience of providing irradiance forecasts for two races run by the University of Michigan Solar Car Team at the Bridgestone World Solar Challenge 2015 in Australia and at the American Solar Challenge 2016 from Ohio to South Dakota. The probabilistic forecasts of hourly solar irradiance generated from machine learning algorithms were deployed to optimally decide on the race strategy. This work showcases an example of real time decision making based on insights derived from machine learning utilizing big geospatial data — weather models and measurement data from weather station networks.
Xiaoyan Shao, Siyuan Lu 0003, Theodore G. van Kessel, Hendrik F. Hamann, Leda Daehler, Jeffrey Cwagenberg, Alan Li
IEEE BigData4
2015 PAIRS: A scalable geo-spatial data analytics platform
abstract
Geospatial data volume exceeds hundreds of Petabytes and is increasing exponentially mainly driven by images/videos/data generated by mobile devices and high resolution imaging systems. Fast data discovery on historical archives and/or real time datasets is currently limited by various data formats that have different projections and spatial resolution, requiring extensive data processing before analytics can be carried out. A new platform called Physical Analytics Integrated Repository and Services (PAIRS) is presented that enables rapid data discovery by automatically updating, joining, and homogenizing data layers in space and time. Built on top of open source big data software, PAIRS manages automatic data download, data curation, and scalable storage while being simultaneously a computational platform for running physical and statistical models on the curated datasets. By addressing data curation before data being uploaded to the platform, multi-layer queries and filtering can be performed in real time. In addition, PAIRS offers a foundation for developing custom analytics. Towards that end we present two examples with models which are running operationally: (1) high resolution evapo-transpiration and vegetation monitoring for agriculture and (2) hyperlocal weather forecasting driven by machine learning for renewable energy forecasting.
Levente J. Klein, Fernando J. Marianno, Conrad M. Albrecht, Marcus Freitag, Siyuan Lu 0003, Nigel Hinds, Xiaoyan Shao, Sergio Bermudez Rodriguez, Hendrik F. Hamann
IEEE BigData9
2015 A testing platform for on-drone computation
abstract
This paper describes the development of a test bed for an on-drone computation system, in which the drone plays the game of ping-pong competitively (YCCD: The Yorktown Cognitive Competition Drone). Unlike other drone systems and demonstrators YCCD will be completely autonomous with no external support from cameras, servers, GPS etc. YCCD will have ultra-low power computation capabilities including on-drone real-time processing for vision and localization (non-GPS based). Architectural design and processing algorithms of the system are discussed in detail.
Dhruv Nair, Oki Gunawan, Theodore G. van Kessel, Hendrik F. Hamann
ICCD5
2015 From Smart Sensors to Smarter Solutions with Physical Analytics
abstract
While in the past most information on the internet was generated by humans or computers, with the emergence of the Internet of Things, vast amount of data is now being created by sensors from devices, machines etc, which are placed in the physical world. Here we present a series of example applications enabled by such sensor data and what we call "Physical Analytics", which provides the underlying intelligence using a combination of physical and statistical models. The smarter solutions, which are being presented in this talk, range from active energy management and optimization, environmental sensing and controls, precision agriculture to renewable energy forecasting. All these different applications have been built using a single platform, which is comprised of a set of "configurable" technologies components including ultra-low power sensing and communication, big data management technologies, numerical modeling for physical systems, machine learning based physical model blending, and physical analytics based automation and control.
Hendrik F. Hamann
SenSys1
2011 A unified approach to coordinated energy-management in data centers
Rajarshi Das, Srinivas Yarlanki, Hendrik F. Hamann, Jeffrey O. Kephart, Vanessa López
CNSM3
2011 Hotspot diagnosis on logical level
Bo Yang 0013, Hendrik F. Hamann, Jeffrey O. Kephart, Stephan Barabasi
CNSM2
2011 Smarter data center power monitoring and management
abstract
This demonstration presents a power panel level power monitoring and management (PMM) system developed at IBM Research. The ultimate goal of this project is to develop a low-cost, high accuracy, non-intrusive and retrofittable data center power management system.
Wael El-Essawy, Malcolm Allen-Ware, Karthick Rajamani, Juan C. Rubio, Michael A. Schappert, Tom W. Keller, Hendrik F. Hamann
SenSys8
2010 Power-efficient, reliable microprocessor architectures: modeling and design methods
abstract
Next generation system designs are challenged by multiple "walls": among them, the inter-related impediments offered by power dissipation limits and reliability are particularly difficult ones that all current chip/system design teams are grappling with. In this paper, we first describe the attendant challenges in integrated (multi-dimensional) pre-silicon modeling and the solution approaches being pursued. Later, we focus on leading edge solutions for power, thermal and failure-rate mitigation that have been proposed in our R&D work over the past decade.
Pradip Bose, Alper Buyuktosunoglu, Chen-Yong Cher, John A. Darringer, Meeta Sharma Gupta, Hendrik F. Hamann, Hans M. Jacobson, Prabhakar Kudva, Eren Kursun, Niti Madan, Indira Nair, Jude A. Rivers, Jeonghee Shin, Alan J. Weger, Victor V. Zyuban
ACM Great Lakes Symposium on VLSI6
2007 Thermal-aware task scheduling at the system software level
abstract
Power-related issues have become important considerations in current generation microprocessor design. One of these issues is that of elevated on-chip temperatures. This has an adverse effect on cooling cost and, if not addressed suitably, on chip reliability. In this paper we investigate the general trade-offs between temporal and spatial hot spot mitigation schemes and thermal time constants, workload variations and microprocessor power distributions. By leveraging spatial and temporal heat slacks, our schemes enable lowering of on-chip unit temperatures by changing the workload in a timely manner with Operating System(OS) and existing hardware support.
Jeonghwan Choi, Chen-Yong Cher, Hubertus Franke, Hendrik F. Hamann, Alan J. Weger, Pradip Bose
ISLPED4