VLDB 2026 Research / reviewers in the wild / expert
Conrad M. Albrecht
dblp:173/9302
· DBLP profile ↗
9ranked-venue papers in the field
3as first author
3since 2021 · last 2022
0009-0009-2422-7289ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8 (3 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Peaks Fusion assisted Early-stopping Strategy for Overhead Imagery Segmentation with Noisy LabelsabstractAutomatic label generation systems, which are capable to generate huge amounts of labels with limited human efforts, enjoy lots of potential in the deep learning era. These easy-to-come-by labels inevitably bear label noises due to a lack of human supervision and can bias model training to some inferior solutions. However, models can still learn some plausible features, before they start to overfit on noisy patterns. Inspired by this phenomenon, we propose a new Peaks fusion assisted EArly-Stopping (PEAS) approach for imagery segmentation with noisy labels, which is mainly composed of two parts. First, a fitting based early-stopping criterion is used to detect the turning phase from which models are about to mimic noise details. After that, a peaks fusion strategy is applied to select reliable models in the detection zone to generate final fusion results. Here, validation accuracies are utilized as indicators in model selection. The proposed method was evaluated on New York City dataset whose labels were automatically collected by a rule-based label generation system, thus noisy to some extent due to a lack of human supervision. The experimental results showed that the proposed PEAS method can achieve both promising statistical and visual results when trained with noisy labels. Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001 |
IEEE Big Data | 2 |
| 2022 | Deep Semantic Model Fusion for Ancient Agricultural Terrace DetectionabstractDiscovering ancient agricultural terraces in desert regions is important for the monitoring of long-term climate changes on the Earth’s surface. However, traditional ground surveys are both costly and limited in scale. With the increasing accessibility of aerial and satellite data, machine learning techniques bear large potential for the automatic detection and recognition of archaeological landscapes. In this paper, we propose a deep semantic model fusion method for ancient agricultural terrace detection. The input data includes aerial images and LiDAR generated terrain features in the Negev desert. Two deep semantic segmentation models, namely DeepLabv3+ and UNet, with EfficientNet backbone, are trained and fused to provide segmentation maps of ancient terraces and walls. The proposed method won the first prize in the International AI Archaeology Challenge. Codes are available at https://github.com/wangyi111/international-archaeologyai-challenge. Yi Wang 0072, Chenying Liu 0001, Arti Tiwari, Micha Silver, Arnon Karnieli, Xiao Xiang Zhu 0001, Conrad M. Albrecht |
IEEE Big Data | 7 |
| 2021 | AutoGeoLabel: Automated Label Generation for Geospatial Machine LearningabstractA key challenge of supervised learning is the availability of human-labeled data. We evaluate a big data processing pipeline to auto-generate labels for remote sensing data. It is based on rasterized statistical features extracted from surveys such as e.g. LiDAR measurements. Using simple combinations of the rasterized statistical layers, it is demonstrated that multiple classes can be generated at accuracies of ~ 0.9.As proof of concept, we utilize the big geo-data platform IBM PAIRS to dynamically generate such labels in dense urban areas with multiple land cover classes. The general method proposed here is platform independent, and it can be adapted to generate labels for other satellite modalities in order to enable machine learning on overhead imagery for land use classification and object detection. Conrad M. Albrecht, Fernando J. Marianno, Levente J. Klein |
IEEE BigData | 1 |
| 2020 | Map Generation from Large Scale Incomplete and Inaccurate Data LabelsabstractAccurately and globally mapping human infrastructure is an important and challenging task with applications in routing, regulation compliance monitoring, and natural disaster response management etc.. In this paper we present progress in developing an algorithmic pipeline and distributed compute system that automates the process of map creation using high resolution aerial images. Unlike previous studies, most of which use datasets that are available only in a few cities across the world, we utilizes publicly available imagery and map data, both of which cover the contiguous United States (CONUS). We approach the technical challenge of inaccurate and incomplete training data adopting state-of-the-art convolutional neural network architectures such as the U-Net and the CycleGAN to incrementally generate maps with increasingly more accurate and more complete labels of man-made infrastructure such as roads and houses. Since scaling the mapping task to CONUS calls for parallelization, we then adopted an asynchronous distributed stochastic parallel gradient descent training scheme to distribute the computational workload onto a cluster of GPUs with nearly linear speed-up. Conrad M. Albrecht, Wei Zhang 0022, Ulrich Finkler, David S. Kung 0001, Siyuan Lu 0003 |
KDD | 2 |
| 2019 | Learning and Recognizing Archeological Features from LiDAR DataabstractWe present a remote sensing pipeline that processes LiDAR (Light Detection And Ranging) data through machine & deep learning for the application of archeological feature detection on big geo-spatial data platforms such as e.g. IBM PAIRS Geoscope [1], [2].Today, archeologists get overwhelmed by the task of visually surveying huge amounts of (raw) LiDAR data in order to identify areas of interest for inspection on the ground. We showcase a software system pipeline that results in significant savings in terms of expert productivity while missing only a small fraction of the artifacts.Our work employs artificial neural networks in conjunction with an efficient spatial segmentation procedure based on domain knowledge. Data processing is constraint by a limited amount of training labels and noisy LiDAR signals due to vegetation cover and decay of ancient structures. We aim at identifying geo-spatial areas with archeological artifacts in a supervised fashion allowing the domain expert to flexibly tune parameters based on her needs. Conrad M. Albrecht, Chris Fisher, Marcus Freitag, Hendrik F. Hamann, Sharath Pankanti, Florencia Pezzutti, Francesca Rossi 0001 |
IEEE BigData | 1 |
| 2019 | N-dimensional geospatial data and analytics for critical infrastructure risk assessmentabstractThe assessment of the vegetation growth rate given remote sensing data is a challenging task in the Earth Observation sciences. LiDAR data acquisition is commonly used to extract height information at a given moment in time, however, the associated cost and complexity restrict continuous acquisitions. Frequently captured aerial imagery can be used to identify and separate vegetation from bare land, water, impervious surface, or built infrastructure. A combination of LiDAR data with aerial and radar imagery allows to track dynamic seasonal growth of vegetation around critical infrastructure such as power lines. We present a general framework that integrates tree identification and growth assessment around power lines with the goal to identify locations of high risk where trees potentially cause power outages. Levente J. Klein, Conrad M. Albrecht, Carlo Siebenschuh, Sharath Pankanti, Hendrik F. Hamann, Siyuan Lu 0003 |
IEEE BigData | 2 |
| 2017 | Event clustering & event series characterization on expected frequencyabstractWe present an efficient clustering algorithm applicable to one-dimensional data such as e.g. a series of times-tamps. Given an expected frequency ΔT-1, we introduce an O(N)-efficient method of characterizing N events represented by an ordered series of timestamps t1, t2,..., tN. In practice, the method proves useful to e.g. identify time intervals of missing data or to locate isolated events. Moreover, we define measures to quantify a series of events by varying ΔT to e.g. determine the quality of an Internet of Things service. Conrad M. Albrecht, Marcus Freitag, Theodore G. van Kessel, Siyuan Lu 0003, Hendrik F. Hamann |
IEEE BigData | 1 |
| 2016 | IBM PAIRS curated big data service for accelerated geospatial data analytics and discoveryabstractIBM's Physical Analytics Integrated Data Repository and Services (PAIRS) is a geospatial Big Data service. PAIRS contains a massive amount of curated geospatial (or more precisely spatio-temporal) data from a large number of public and private data resources, and also supports user contributed data layers. PAIRS offers an easy-to-use platform for both rapid assembly and retrieval of geospatial datasets or performing complex analytics, lowering time-to-discovery significantly by reducing the data curation and management burden. In this paper, we review recent progress with PAIRS and showcase a few exemplary analytical applications which the authors are able to build with relative ease leveraging this technology. Siyuan Lu 0003, Xiaoyan Shao, Marcus Freitag, Levente J. Klein, Jason D. Renwick, Fernando J. Marianno, Conrad M. Albrecht, Hendrik F. Hamann |
IEEE BigData | 7 |
| 2015 | PAIRS: A scalable geo-spatial data analytics platformabstractGeospatial data volume exceeds hundreds of Petabytes and is increasing exponentially mainly driven by images/videos/data generated by mobile devices and high resolution imaging systems. Fast data discovery on historical archives and/or real time datasets is currently limited by various data formats that have different projections and spatial resolution, requiring extensive data processing before analytics can be carried out. A new platform called Physical Analytics Integrated Repository and Services (PAIRS) is presented that enables rapid data discovery by automatically updating, joining, and homogenizing data layers in space and time. Built on top of open source big data software, PAIRS manages automatic data download, data curation, and scalable storage while being simultaneously a computational platform for running physical and statistical models on the curated datasets. By addressing data curation before data being uploaded to the platform, multi-layer queries and filtering can be performed in real time. In addition, PAIRS offers a foundation for developing custom analytics. Towards that end we present two examples with models which are running operationally: (1) high resolution evapo-transpiration and vegetation monitoring for agriculture and (2) hyperlocal weather forecasting driven by machine learning for renewable energy forecasting. Levente J. Klein, Fernando J. Marianno, Conrad M. Albrecht, Marcus Freitag, Siyuan Lu 0003, Nigel Hinds, Xiaoyan Shao, Sergio Bermudez Rodriguez, Hendrik F. Hamann |
IEEE BigData | 3 |