VLDB 2026 Research / reviewers in the wild / expert
Michael Kopp 0001
dblp:150/2685-1 · also Michael K. Kopp 0001
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0002-1385-1109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | xLSTM: Extended Long Short-Term MemoryabstractIn the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM). Since then, LSTMs have stood the test of time and contributed to numerous deep learning success stories, in particular they constituted the first Large Language Models (LLMs). However, the advent of the Transformer technology with parallelizable self-attention at its core marked the dawn of a new era, outpacing LSTMs at scale. We now raise a simple question: How far do we get in language modeling when scaling LSTMs to billions of parameters, leveraging the latest techniques from modern LLMs, but mitigating known limitations of LSTMs? Firstly, we introduce exponential gating with appropriate normalization and stabilization techniques. Secondly, we modify the LSTM memory structure, obtaining: (i) sLSTM with a scalar memory, a scalar update, and new memory mixing, (ii) mLSTM that is fully parallelizable with a matrix memory and a covariance update rule. Integrating these LSTM extensions into residual block backbones yields xLSTM blocks that are then residually stacked into xLSTM architectures. Exponential gating and modified memory structures boost xLSTM capabilities to perform favorably when compared to state-of-the-art Transformers and State Space Models, both in performance and scaling. Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp 0001, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter |
NeurIPS | 6 |
| 2024 | Dsfer-Net: A Deep Supervision and Feature Retrieval Network for Bitemporal Change Detection Using Modern Hopfield NetworksabstractChange detection, an essential application for high-resolution remote sensing (RS) images, aims to monitor and analyze changes in the land surface over time. Due to the rapid increase in the quantity of high-resolution RS data and the complexity of texture features, several quantitative deep learning-based methods have been proposed. These methods outperform traditional change detection (CD) methods by extracting deep features and combining spatial–temporal information. However, reasonable explanations for how deep features improve detection performance are still lacking. In our investigations, we found that modern Hopfield network (MHN) layers significantly enhance semantic understanding. In this article, we propose a deep supervision and feature retrieval network (Dsfer-Net) for bitemporal CD. Specifically, the highly representative deep features of bitemporal images are jointly extracted through a fully convolutional Siamese network. Based on the sequential geographical information of the bitemporal images, we designed a feature retrieval module to extract difference features and leverage discriminative information in a deeply supervised manner. In addition, we observed that the deeply supervised feature retrieval (DSFR) module provides explainable evidence of the semantic understanding of the proposed network in its deep layers. Finally, our end-to-end network establishes a novel framework by aggregating retrieved features and feature pairs from different layers. Experiments conducted on three public datasets (LEVIR-CD, WHU-CD, and CDD) confirm the superiority of the proposed Dsfer-Net over other state-of-the-art methods. Compared to the best-performing DSAMNet, Dsfer-Net demonstrates significant improvements, with$F1$scores increasing by 4.7%, 5.9%, and 2.3%. Furthermore, compared to our previous FrNet, Dsfer-Net also achieves noteworthy enhancements, with$F1$scores increasing by 2.0%, 1.4%, and 4.5% on three datasets. The code will be available online (https://github.com/ShizhenChang/Dsfer-Net). Shizhen Chang, Michael Kopp 0001, Pedram Ghamisi, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Txt2Img-MHN: Remote Sensing Image Generation From Text Using Modern Hopfield NetworksabstractThe synthesis of high-resolution remote sensing images based on text descriptions has great potential in many practical application scenarios. Although deep neural networks have achieved great success in many important remote sensing tasks, generating realistic remote sensing images from text descriptions is still very difficult. To address this challenge, we propose a novel text-to-image modern Hopfield network (Txt2Img-MHN). The main idea of Txt2Img-MHN is to conduct hierarchical prototype learning on both text and image embeddings with modern Hopfield layers. Instead of directly learning concrete but highly diverse text-image joint feature representations for different semantics, Txt2Img-MHN aims to learn the most representative prototypes from text-image embeddings, achieving a coarse-to-fine learning strategy. These learned prototypes can then be utilized to represent more complex semantics in the text-to-image generation task. To better evaluate the realism and semantic consistency of the generated images, we further conduct zero-shot classification on real remote sensing data using the classification model trained on synthesized images. Despite its simplicity, we find that the overall accuracy in the zero-shot classification may serve as a good metric to evaluate the ability to generate an image from text. Extensive experiments on the benchmark remote sensing text-image dataset demonstrate that the proposed Txt2Img-MHN can generate more realistic remote sensing images than existing methods. Code and pre-trained models are available online (https://github.com/YonghaoXu/Txt2Img-MHN). Yonghao Xu, Weikang Yu, Pedram Ghamisi, Michael Kopp 0001, Sepp Hochreiter |
IEEE Trans. Image Process. | 4 |
| 2023 | Metropolitan Segment Traffic Speeds From Massive Floating Car Data in 10 CitiesabstractTraffic analysis is crucial for urban operations and planning, while the availability of dense urban traffic data beyond loop detectors is still scarce. We present a large-scale floating vehicle dataset of per-street segment traffic information, Metropolitan Segment Traffic Speeds from Massive Floating Car Data in 10 Cities (MeTS-10), available for 10 global cities with a 15-minute resolution for collection periods ranging between 108 and 361 days in 2019–2021 and covering more than 1500 square kilometers per metropolitan area. MeTS-10 features traffic speed information at all street levels from main arterials to local streets for Antwerp, Bangkok, Barcelona, Berlin, Chicago, Istanbul, London, Madrid, Melbourne, and Moscow. The dataset leverages the industrial-scale floating vehicle Traffic4cast data with speeds and vehicle counts provided in a privacy-preserving spatio-temporal aggregation. We detail the efficient matching approach mapping the data to the OpenStreetMap (OSM) road graph. We evaluate the dataset by comparing it with publicly available stationary vehicle detector data (for Berlin, London, and Madrid) and the Uber traffic speed dataset (for Barcelona, Berlin, and London). The comparison highlights the differences across datasets in spatio-temporal coverage and variations in the reported traffic caused by the binning method. MeTS-10 enables novel, city-wide analysis of mobility and traffic patterns for ten major world cities, overcoming current limitations of spatially sparse vehicle detector data. The large spatial and temporal coverage offers an opportunity for joining the MeTS-10 with other datasets, such as traffic surveys in traffic planning studies or vehicle detector data in traffic control settings. Moritz Neun, Christian Eichenberger, Yanan Xin 0001, Nina Wiedemann, Henry Martin, Martin Tomko 0001, Lukas Ambühl, Luca Hermes, Michael Kopp 0001 |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2022 | CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIPabstractCLIP yielded impressive results on zero-shot transfer learning tasks and is considered as a foundation model like BERT or GPT3. CLIP vision models that have a rich representation are pre-trained using the InfoNCE objective and natural language supervision before they are fine-tuned on particular tasks. Though CLIP excels at zero-shot transfer learning, it suffers from an explaining away problem, that is, it focuses on one or few features, while neglecting other relevant features. This problem is caused by insufficiently extracting the covariance structure in the original multi-modal data. We suggest to use modern Hopfield networks to tackle the problem of explaining away. Their retrieved embeddings have an enriched covariance structure derived from co-occurrences of features in the stored embeddings. However, modern Hopfield networks increase the saturation effect of the InfoNCE objective which hampers learning. We propose to use the InfoLOOB objective to mitigate this saturation effect. We introduce the novel "Contrastive Leave One Out Boost" (CLOOB), which uses modern Hopfield networks for covariance enrichment together with the InfoLOOB objective. In experiments we compare CLOOB to CLIP after pre-training on the Conceptual Captions and the YFCC dataset with respect to their zero-shot transfer learning performance on other datasets. CLOOB consistently outperforms CLIP at zero-shot transfer learning across all considered architectures and datasets. Andreas Fürst, Elisabeth Rumetshofer, Johannes Lehner, Viet T. Tran, Hubert Ramsauer, David P. Kreil, Michael Kopp 0001, Günter Klambauer, Angela Bitto-Nemling, Sepp Hochreiter |
NeurIPS | 8 |
| 2022 | Sketched Multiview Subspace Learning for Hyperspectral Anomalous Change DetectionabstractIn recent years, multi-view subspace learning has been garnering increasing attention. It aims to capture the inner relationships of the data that are collected from multiple sources by learning a unified representation. In this way, comprehensive information from multiple views is shared and preserved for the generalization processes. As a special branch of temporal series hyperspectral image (HSI) processing, the anomalous change detection task focuses on detecting very small changes among different temporal images. However, when the volume of datasets is very large or the classes are relatively comprehensive, existing methods may fail to find those changes between the scenes, and end up with terrible detection results. In this paper, inspired by the sketched representation and multi-view subspace learning, a sketched multi-view subspace learning (SMSL) model is proposed for HSI anomalous change detection. The proposed model preserves major information from the image pairs and improves computational complexity by using a sketched representation matrix. Furthermore, the differences between scenes are extracted by utilizing the specific regularizer of the self-representation matrices. To evaluate the detection effectiveness of the proposed SMSL model, experiments are conducted on a benchmark hyperspectral remote sensing dataset and a natural hyperspectral dataset, and compared with other state-of-the-art approaches. The codes of the proposed method will be made available online1. Shizhen Chang, Michael Kopp 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide DetectionabstractThis study introducesLandslide4Sense, a reference benchmark for landslide detection from remote sensing. The repository features 3,799 image patches fusing optical layers from Sentinel-2 sensors with the digital elevation model and slope layer derived from ALOS PALSAR. The added topographical information facilitates an accurate detection of landslide borders, which recent researches have shown to be challenging using optical data alone. The extensive data set supports deep learning (DL) studies in landslide detection and the development and validation of methods for the systematic update of landslide inventories. The benchmark data set has been collected at four different times and geographical locations: Iburi (September 2018), Kodagu (August 2018), Gorkha (April 2015), and Taiwan (August 2009). Each image pixel is labelled as belonging to a landslide or not, incorporating various sources and thorough manual annotation. We then evaluate the landslide detection performance of 11 state-of-the-art DL segmentation models: U-Net, ResU-Net, PSPNet, ContextNet, DeepLab-v2, DeepLab-v3+, FCN-8s, LinkNet, FRRN-A, FRRN-B, and SQNet. All models were trained from scratch on patches from one quarter of each study area and tested on independent patches from the other three quarters. Our experiments demonstrate that ResU-Net outperformed the other models for the landslide detection task. We make the multi-source landslide benchmark data (Landslide4Sense) and the tested DL models publicly available at https://www.iarai.ac.at/landslide4sense, establishing an important resource for remote sensing, computer vision, and machine learning communities in studies of image classification in general and applications to landslide detection in particular. Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi, Michael Kopp 0001, David P. Kreil |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | High-resolution multi-channel weather forecasting - First insights on transfer learning from the Weather4cast Competitions 2021abstractWeather forecasting is both a high impact application as well as a complex Big Data modelling challenge. Recent advances in machine learning have already demonstrated the power of non-physical modelling approaches for the prediction of rainfall. The Weather4cast competitions now provide a unique multi-channel benchmark for the prediction of up to 8 hours of weather with high temporal and spatial resolutions (15 min, 4 km) for a diverse set of large regions across Earth. This diversity, for the first time, also permits a meaningful spatial transfer learning challenge in weather forecasting.Weather4cast introduces multi-channel weather ‘movies’ that encode temperature, rainfall, cloud properties, and turbulence as derived from the meteorological satellites by the EUMETSAT NWC SAF. Inspired by the Traffic4cast competitions at the NeurIPS conferences in 2019 and 2020, weather forecasting is thus presented as a video frame prediction task. As then, the U-Net based models developed for photographic image analysis intriguingly did well on these artificial videos. In contrast, however, the winning submission did not employ a U-Net but a recurrent convolutional network with residual units.Weather4cast introduces the first spatial transfer learning challenge in weather forecasting: only one-hour short snippets from spatial regions never seen before were provided as input to models. We can thus now present first insights from submissions to this spatial transfer learning challenge. Notably, models with better core prediction performance also generalized better. Moreover, the two top-ranked models – one RCN based, one U-Net based – were further ahead of the remaining top-ranked submissions for spatial transfer learning (+6%) than in the core prediction challenge (+1%).While submissions tested varying strategies for input data selection and training, there remains a wide range of additional complementary approaches to be explored in future analyses. The competition and its leaderboards remain available and open for new submissions on the weather4cast.ai website. Pedro Herruzo, Aleksandra Gruca, Llorenç Lliso, Xavier Calbet, Pilar Rípodas, Sepp Hochreiter, Michael Kopp 0001, David P. Kreil |
IEEE BigData | 7 |
| 2021 | CDCEO'21 - First Workshop on Complex Data Challenges in Earth ObservationabstractHigh-resolution remote sensing technology for Earth Observation (EO) has radically changed how we monitor the state of our planet around the clock. An effective interpretation of the resulting complex large-scale time series adopts the best machine learning techniques from signal processing, computer vision, pattern recognition, and artificial intelligence. The First Workshop on Complex Data Challenges in Earth Observation was open to both method development and advanced applications in a wide range of related topics, including image and signal processing, gap-filling, data fusion, feature extraction, prediction of spatio-temporal features, and the detection of rules underlying the observed state transitions and causal relationships. The full agenda, featuring keynotes and a selection of high quality contributed talks is available online at www.iarai.ac.at/cdceo21 Aleksandra Gruca, Pedro Herruzo, Pilar Rípodas, Andrzej Kucik, Christian Briese, Michael Kopp 0001, Sepp Hochreiter, Pedram Ghamisi, David P. Kreil |
CIKM | 6 |
| 2021 | Hopfield Networks is All You Need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David P. Kreil, Michael Kopp 0001, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter |
ICLR | 10 |
| 2019 | Towards Modeling Geographical Processes with Generative Adversarial Networks (GANs) (Short Paper)abstractRecently, Generative Adversarial Networks (GANs) have demonstrated great potential for a range of Machine Learning tasks, including synthetic video generation, but have so far not been applied to the domain of modeling geographical processes. In this study, we align these two problems and - motivated by the potential advantages of GANs compared to traditional geosimulation methods - test the capability of GANs to learn a set of underlying rules which determine a geographical process. For this purpose, we turn to Conway’s well-known Game of Life (GoL) as a source for spatio-temporal training data, and further argue for its (and simple variants of it) usefulness as a potential standard training data set for benchmarking generative geographical process models. David Jonietz, Michael Kopp 0001 |
COSIT | 2 |