EDBT 2026 Demo / reviewers in the wild / expert
Klaus Jung
dblp:77/8318
· DBLP profile ↗
29ranked-venue papers
4as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Exploration Graph: A Novel Approach for Efficient Nearest Neighbor Search in Evolving Multimedia Datasets
Nico Hezel, Kai Uwe Barthel, Bruno Schilling, Konstantin Schall, Klaus Jung |
MMM (1) | 5 |
| 2025 | Ten simple rules for effective research data management
Max J. Hassenstein, Klaus Jung |
PLoS Comput. Biol. | 2 |
| 2024 | An Exploration Graph with Continuous Refinement for Efficient Multimedia RetrievalabstractAs datasets and the dimensionality of feature vectors continue to grow, Approximate Nearest Neighbor Search (ANNS) in large multimedia databases becomes increasingly relevant. Graph-based approaches have demonstrated to offer the best trade-off between retrieval precision and search time. Despite their ability to deliver search times several orders of magnitude faster than exact search techniques, existing methods suffer from slow constructions speeds or high memory requirements. This paper presents a continuous refining Exploration Graph (crEG), a novel approach for rapidly constructing a compact exploration graph with state-of-the-art search performance. Additionally, it provides the ability to enhance its effectiveness even further through an optional edge optimization algorithm. Both algorithms are specifically designed to produce and operate on undirected graphs with even degrees and guarantee graph connectivity at any time - a property particularly valuable for exploratory search, where the query is part of the database elements. Although such queries provide an advantageous starting point for graph search algorithms, they have been rarely considered in the context of ANNS, yet are crucial for recommendation and exploration systems. Our experiments demonstrate high efficiency in ANNS does not necessarily translate to a good performance in exploratory search. Nico Hezel, Kai Uwe Barthel, Konstantin Schall, Klaus Jung |
ICMR | 4 |
| 2024 | Optimizing the Interactive Video Retrieval Tool Vibro for the Video Browser Showdown 2024
Konstantin Schall, Nico Hezel, Kai Uwe Barthel, Klaus Jung |
MMM (4) | 4 |
| 2024 | Adapting the Exploration Graph for High Throughput in Low Recall Regimes
Nico Hezel, Bruno Schilling, Kai Uwe Barthel, Konstantin Schall, Klaus Jung |
SISAP | 5 |
| 2024 | Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment
Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung |
SISAP | 4 |
| 2023 | navigu.net: NAvigation in Visual Image Graphs gets User-friendlyabstractDue to the size of today’s image collections it can be challenging to fully understand their content. Recent technological advances have enabled efficient visual search. These systems use joint visual and textual feature vectors to identify similar images based on image queries or text descriptions. Despite their effectiveness, high-dimensional feature vectors can lead to long search times for large collections. In this demonstration, we propose a solution that significantly reduces search times and increases the efficiency of the search system. By combining two separate image graphs, our method provides fast approximate nearest neighbor search and allows seamless visual exploration of the entire collection in real time through a standard web browser, using familiar navigation techniques such as zooming and dragging, common in systems like Google Maps. Kai Uwe Barthel, Nico Hezel, Konstantin Schall, Klaus Jung |
ICMR | 4 |
| 2023 | Improving Image Encoders for General-Purpose Nearest Neighbor Search and ClassificationabstractRecent advances in computer vision research led to large vision foundation models that generalize to a broad range of image domains and perform exceptionally well in various image based tasks. However, content-based image-to-image retrieval is often overlooked in this context. This paper investigates the effectiveness of different vision foundation models on two challenging nearest neighbor search-based tasks: zero-shot retrieval and k-NN classification. A benchmark for evaluating the performance of various vision encoders and their pre-training methods is established, where significant differences in the performance of these models are observed. Additionally, we propose a fine-tuning regime that improves zero-shot retrieval and k-NN classification through training with a combination of large publicly available datasets without specializing in any data domain. Our results show that the retrained vision encoders have a higher degree of generalization across different search-based tasks and can be used as general-purpose embedding models for image retrieval. Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung |
ICMR | 4 |
| 2023 | Vibro: Video Browsing with Semantic and Visual Image Embeddings
Konstantin Schall, Nico Hezel, Klaus Jung, Kai Uwe Barthel |
MMM (1) | 3 |
| 2022 | Efficient Search and Browsing of Large-Scale Video Collections with Vibro
Nico Hezel, Konstantin Schall, Klaus Jung, Kai Uwe Barthel |
MMM (2) | 3 |
| 2022 | PicArrange - Visually Sort, Search, and Explore Private Images on a Mac Computer
Klaus Jung, Kai Uwe Barthel, Nico Hezel, Konstantin Schall |
MMM (2) | 1 |
| 2022 | GPR1200: A Benchmark for General-Purpose Content-Based Image Retrieval
Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung |
MMM (1) | 4 |
| 2021 | Video Search with Sub-Image Keyword Transfer Using Existing Image Archives
Nico Hezel, Konstantin Schall, Klaus Jung, Kai Uwe Barthel |
MMM (2) | 3 |
| 2021 | Measuring reproducibility of virus metagenomics analyses using bootstrap samples from FASTQ-filesabstractMOTIVATION: High-throughput sequencing data can be affected by different technical errors, e.g. from probe preparation or false base calling. As a consequence, reproducibility of experiments can be weakened. In virus metagenomics, technical errors can result in falsely identified viruses in samples from infected hosts. We present a new resampling approach based on bootstrap sampling of sequencing reads from FASTQ-files in order to generate artificial replicates of sequencing runs which can help to judge the robustness of an analysis. In addition, we evaluate a mixture model on the distribution of read counts per virus to identify potentially false positive findings. RESULTS: The evaluation of our approach on an artificially generated dataset with known viral sequence content shows in general a high reproducibility of uncovering viruses in sequencing data, i.e. the correlation between original and mean bootstrap read count was highly correlated. However, the bootstrap read counts can also indicate reduced or increased evidence for the presence of a virus in the biological sample. We also found that the mixture-model fits well to the read counts, and furthermore, it provides a higher accuracy on the original or on the bootstrap read counts than on the difference between both. The usefulness of our methods is further demonstrated on two freely available real-world datasets from harbor seals. AVAILABILITY AND IMPLEMENTATION: We provide a Phyton tool, called RESEQ, available from https://github.com/babaksaremi/RESEQ that allows efficient generation of bootstrap reads from an original FASTQ-file. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Babak Saremi, Moritz Kohls, Pamela Liebig, Ursula Siebert, Klaus Jung |
Bioinform. | 5 |
| 2019 | Real-Time Visual Navigation in Huge Image Sets Using Similarity GraphsabstractNowadays stock photo agencies often have millions of images. Non-stop viewing of 20 million images at a speed of 10 images per second would take more than three weeks. This demonstrates the impossibility to inspect all images and the difficulty to get an overview of the entire collection. Although there has been a lot of effort to improve visual image search, there is little research and support for visual image exploration. Typically, users start "exploring" an image collection with a keyword search or an example image for a similarity search. Both searches lead to long unstructured lists of result images. In earlier publications, we introduced the idea of graph-based image navigation and proposed an efficient algorithm for building hierarchical image similarity graphs for dynamically changing image collections. In this demo we showcase real-time visual exploration of millions of images with a standard web browser. Subsets of images are successively retrieved from the graph and displayed as a visually sorted 2D image map, which can be zoomed and dragged to explore related concepts. Maintaining the positions of previously shown images creates the impression of an "endless map". This approach allows an easy visual image-based navigation, while preserving the complex image relationships of the graph. Kai Uwe Barthel, Nico Hezel, Konstantin Schall, Klaus Jung |
ACM Multimedia | 4 |
| 2019 | Visual Navigation of Large Image GraphsabstractIt is impossible to inspect or get an overview of image collections with millions of images. Users often start “exploring” images with a keyword or a similarity search. Both lead to long unstructured lists of result images. In this demo we present a graph-based system for visually exploring and navigating continuously changing sets of millions of images with a web browser. Subsets of images are successively retrieved from a image similarity graph and displayed as a visually sorted 2D image map, which can be zoomed and dragged to explore images from related concepts. Nico Hezel, Kai Uwe Barthel, Konstantin Schall, Klaus Jung |
MMSP | 4 |
| 2019 | Deep Aggregation of Regional Convolutional Activations for Content Based Image RetrievalabstractOne of the key challenges of deep learning based image retrieval remains in aggregating convolutional activations into one highly representative feature vector. Ideally, this descriptor should encode semantic, spatial and low level information. Even though off-the-shelf pre-trained neural networks can already produce good representations in combination with aggregation methods, appropriate fine tuning for the task of image retrieval has shown to significantly boost retrieval performance. In this paper we present a simple yet effective supervised aggregation method built on top of existing regional pooling approaches. In addition to the maximum activation of a given region, we calculate regional average activations of extracted feature maps. Subsequently, weights for each of the pooled feature vectors are learned to perform a weighted aggregation to a single feature vector. Furthermore, we apply our newly proposed NRA loss function for deep metric learning to fine tune the backbone neural network and to learn the aggregation weights. Our method achieves state-of-the-art results for the INRIA Holidays data set and competitive results for the Oxford Buildings and Paris data sets while reducing the training time significantly. Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung |
MMSP | 4 |
| 2019 | Deep Metric Learning using Similarities from Nonlinear Rank ApproximationsabstractIn recent years, deep metric learning has achieved promising results in learning high dimensional semantic feature embeddings where the spatial relationships of the feature vectors match the visual similarities of the images. Similarity search for images is performed by determining the vectors with the smallest distances to a query vector. However, high retrieval quality does not depend on the actual distances of the feature vectors, but rather on the ranking order of the feature vectors from similar images. In this paper, we introduce a metric learning algorithm that focuses on identifying and modifying those feature vectors that most strongly affect the retrieval quality. We compute normalized approximated ranks and convert them to similarities by applying a nonlinear transfer function. These similarities are used in a newly proposed loss function that better contracts similar and disperses dissimilar samples. Experiments demonstrate significant improvement over existing deep feature embedding methods on the CUB-200-2011, Cars196, and Stanford Online Products data sets for all embedding sizes. Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung |
MMSP | 4 |
| 2019 | Network meta-analysis correlates with analysis of merged independent transcriptome expression dataabstractBACKGROUND: Using meta-analysis, high-dimensional transcriptome expression data from public repositories can be merged to make group comparisons that have not been considered in the original studies. Merging of high-dimensional expression data can, however, implicate batch effects that are sometimes difficult to be removed. Removing batch effects becomes even more difficult when expression data was taken using different technologies in the individual studies (e.g. merging of microarray and RNA-seq data). Network meta-analysis has so far not been considered to make indirect comparisons in transcriptome expression data, when data merging appears to yield biased results. RESULTS: We demonstrate in a simulation study that the results from analyzing merged data sets and the results from network meta-analysis are highly correlated in simple study networks. In the case that an edge in the network is supported by multiple independent studies, network meta-analysis produces fold changes that are closer to the simulated ones than those obtained from analyzing merged data sets. Finally, we also demonstrate the practicability of network meta-analysis on a real-world data example from neuroinfection research. CONCLUSIONS: Network meta-analysis is a useful means to make new inferences when combining multiple independent studies of molecular, high-throughput expression data. This method is especially advantageous when batch effects between studies are hard to get removed. Christine Winter, Robin Kosch, Martin Ludlow, Albert Osterhaus, Klaus Jung |
BMC Bioinform. | 5 |
| 2018 | Fusing Keyword Search and Visual Exploration for Untagged Videos
Kai Uwe Barthel, Nico Hezel, Klaus Jung |
MMM (2) | 3 |
| 2018 | ImageX - Explore and Search Local/Private Images
Nico Hezel, Kai Uwe Barthel, Klaus Jung |
MMM (2) | 3 |
| 2017 | Visually Browsing Millions of Images Using Image GraphsabstractWe present a new approach to visually browse very large sets of untagged images. High quality image features are generated using transformed activations of a convolutional neural network. These features are used to model image similarities, from which a hierarchical image graph is build. We show how such a graph can be constructed efficiently. In our experiments we found best user experience for navigating the graph is achieved by projecting sub-graphs onto a regular 2D image map. This allows users to explore the image collection like an interactive map. Kai Uwe Barthel, Nico Hezel, Klaus Jung |
ICMR | 3 |
| 2017 | kmerPyramid: an interactive visualization tool for nucleobase and k-mer frequenciesabstractSUMMARY: Bioinformatics methods often incorporate the frequency distribution of nulecobases or k-mers in DNA or RNA sequences, for example as part of metagenomic or phylogenetic analysis. Because the frequency matrix with sequences in the rows and nucleobases in the columns is multi-dimensional it is hard to visualize. We present the R-package 'kmerPyramid' that allows to display each sequence, based on its nucleobase or k-mer distribution projected to the space of principal components, as a point within a 3-dimensional, interactive pyramid. Using the computer mouse, the user can turn the pyramid's axes, zoom in and out and identify individual points. Additionally, the package provides the k-mer frequency matrices of about 2000 bacteria and 5000 virus reference sequences calculated from the NCBI RefSeq genbank. The 'kmerPyramid' can particularly be used for visualization of intra- and inter species differences. AVAILABILITY AND IMPLEMENTATION: The R-package 'kmerPyramid' is available from the GitHub website at https://github.com/jkruppa/kmerPyramid. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jochen Kruppa, Erhard van der Vries, Wendy K. Jo, Alexander Postel, Paul Becher, Albert Osterhaus, Klaus Jung |
Bioinform. | 7 |
| 2017 | Automated multigroup outlier identification in molecular high-throughput data using bagplots and gemplotsabstractBACKGROUND: Analyses of molecular high-throughput data often lack in robustness, i.e. results are very sensitive to the addition or removal of a single observation. Therefore, the identification of extreme observations is an important step of quality control before doing further data analysis. Standard outlier detection methods for univariate data are however not applicable, since the considered data are high-dimensional, i.e. multiple hundreds or thousands of features are observed in small samples. Usually, outliers in high-dimensional data are solely detected by visual inspection of a graphical representation of the data by the analyst. Typical graphical representation for high-dimensional data are hierarchical cluster tree or principal component plots. Pure visual approaches depend, however, on the individual judgement of the analyst and are hard to automate. Existing methods for automated outlier detection are only dedicated to data of a single experimental groups. RESULTS: In this work we propose to use bagplots, the 2-dimensional extension of the boxplot, to automatically identify outliers in the subspace of the first two principal components of the data. Furthermore, we present for the first time the gemplot, the 3-dimensional extension of boxplot and bagplot, which can be used in the subspace of the first three principal components. Bagplot and gemplot surround the regular observations with convex hulls and observations outside these hulls are regarded as outliers. The convex hulls are determined separately for the observations of each experimental group while the observations of all groups can be displayed in the same subspace of principal components. We demonstrate the usefulness of this approach on multiple sets of artificial data as well as one set of gene expression data from a next-generation sequencing experiment, and compare the new method to other common approaches. Furthermore, we provide an implementation of the gemplot in the package 'gemPlot' for the R programming environment. CONCLUSIONS: Bagplots and gemplots in subspaces of principal components are useful for automated and objective outlier identification in high-dimensional data from molecular high-throughput experiments. A clear advantage over other methods is that multiple experimental groups can be displayed in the same figure although outlier detection is performed for each individual group. Jochen Kruppa, Klaus Jung |
BMC Bioinform. | 2 |
| 2015 | Comparative study on gene set and pathway topology-based enrichment methodsabstractBACKGROUND: Enrichment analysis is a popular approach to identify pathways or sets of genes which are significantly enriched in the context of differentially expressed genes. The traditional gene set enrichment approach considers a pathway as a simple gene list disregarding any knowledge of gene or protein interactions. In contrast, the new group of so called pathway topology-based methods integrates the topological structure of a pathway into the analysis. METHODS: We comparatively investigated gene set and pathway topology-based enrichment approaches, considering three gene set and four topological methods. These methods were compared in two extensive simulation studies and on a benchmark of 36 real datasets, providing the same pathway input data for all methods. RESULTS: In the benchmark data analysis both types of methods showed a comparable ability to detect enriched pathways. The first simulation study was conducted with KEGG pathways, which showed considerable gene overlaps between each other. In this study with original KEGG pathways, none of the topology-based methods outperformed the gene set approach. Therefore, a second simulation study was performed on non-overlapping pathways created by unique gene IDs. Here, methods accounting for pathway topology reached higher accuracy than the gene set methods, however their sensitivity was lower. CONCLUSIONS: We conducted one of the first comprehensive comparative works on evaluating gene set against pathway topology-based enrichment methods. The topological methods showed better performance in the simulation scenarios with non-overlapping pathways, however, they were not conclusively better in the other scenarios. This suggests that simple gene set approach might be sufficient to detect an enriched pathway under realistic circumstances. Nevertheless, more extensive studies and further benchmark data are needed to systematically evaluate these methods and to assess what gain and cost pathway topology information introduces into enrichment analysis. Both types of methods for enrichment analysis require further improvements in order to deal with the problem of pathway overlaps. Michaela Bayerlová, Klaus Jung, Frank Kramer 0001, Florian Klemm, Annalen Bleckmann, Tim Beißbarth |
BMC Bioinform. | 2 |
| 2014 | Adaption of the global test idea to proteomics data with missing valuesabstractMOTIVATION: Global test procedures are frequently used in gene expression analysis to study the relationship between a functional subset of RNA transcripts and an experimental group factor. However, these procedures have been rarely used for the analysis of high-throughput data from other sources, such as proteome expression data. The main difficulties in transferring global test procedures from genomics to proteomics data are the more complicated way of obtaining functional annotations and the handling of missing values in some types of proteomics data. RESULTS: We propose a simple mixed linear model in combination with a permutation procedure and missing values imputation to conduct global tests in proteomics experiments. This new approach is motivated by protein expression data obtained by means of 2-D gel electrophoresis within a mouse experiment of our current research. A simulation study yielded that power and testing level of the mixed model alone can be affected by missing values in the dataset. Imputation of missing values was able to correct for a bias in some simulation settings. Our new approach provides the possibility to rank Gene Ontology (GO) terms associated with protein sets. It is also helpful in the case in which a specific protein is represented by multiple spots on a 2-D gel by considering these spots also as a protein set. Analysis of our data points at correlations between the deficiency of the protein 'calreticulin' and protein sets related to biological processes in the heart muscle. AVAILABILITY AND IMPLEMENTATION: Our proposed approach is included in the R-package 'RepeatedHighDim', which already contains a global test procedure for gene expression data. The package can be retrieved from http://cran.r-project.org/. CONTACT: [email protected]. Klaus Jung, Hassan Dihazi, Asima Bibi, Gry H. Dihazi, Tim Beißbarth |
Bioinform. | 1 |
| 2011 | Comparison of global tests for functional gene sets in two-group designs and selection of potentially effect-causing genesabstractMOTIVATION: An important object in the analysis of high-throughput genomic data is to find an association between the expression profile of functional gene sets and the different levels of a group response. Instead of multiple testing procedures which focus on single genes, global tests are usually used to detect a group effect in an entire gene set. In a simulation study, we compare the power and computation times of four different approaches for global testing. The applicability of one of these methods to gene expression data is demonstrated for the first time. In addition, we propose an algorithm for the detection of those genes which might be responsible for a group effect. RESULTS: We could detect that the power of three of the approaches is comparable in many settings but considerable differences were detected in the computation times. Our proposed gene selection algorithm was able to detect potentially effect-causing genes in artificial sets with high power when many genes were altered with a small effect, while classical multiple testing was more powerful when few genes were altered with a large effect. AVAILABILITY: An R-package called 'RepeatedHighDim' which implements our new global test procedures is made available from http://cran.r-project.org/. Klaus Jung, Benjamin Becker, Edgar Brunner, Tim Beißbarth |
Bioinform. | 1 |
| 2011 | Reporting FDR analogous confidence intervals for the log fold change of differentially expressed genesabstractBACKGROUND: Gene expression experiments are common in molecular biology, for example in order to identify genes which play a certain role in a specified biological framework. For that purpose expression levels of several thousand genes are measured simultaneously using DNA microarrays. Comparing two distinct groups of tissue samples to detect those genes which are differentially expressed one statistical test per gene is performed, and resulting p-values are adjusted to control the false discovery rate. In addition, the expression change of each gene is quantified by some effect measure, typically the log fold change. In certain cases, however, a gene with a significant p-value can have a rather small fold change while in other cases a non-significant gene can have a rather large fold change. The biological relevance of the change of gene expression can be more intuitively judged by a fold change then merely by a p-value. Therefore, confidence intervals for the log fold change which accompany the adjusted p-values are desirable. RESULTS: In a new approach, we employ an existing algorithm for adjusting confidence intervals in the case of high-dimensional data and apply it to a widely used linear model for microarray data. Furthermore, we adopt a concept of different relevance categories for effects in clinical trials to assess biological relevance of genes in microarray experiments. In a brief simulation study the properties of the adjusting algorithm are maintained when being combined with the linear model for microarray data. In two cancer data sets the adjusted confidence intervals can indicate significance of large fold changes and distinguish them from other large but non-significant fold changes. Adjusting of confidence intervals also corrects the assessment of biological relevance. CONCLUSIONS: Our new combination approach and the categorization of fold changes facilitates the selection of genes in microarray experiments and helps to interpret their biological relevance. Klaus Jung, Tim Friede, Tim Beißbarth |
BMC Bioinform. | 1 |
| 2011 | Sequential Interim Analyses of Survival Data in DNA Microarray ExperimentsabstractBACKGROUND: Discovery of biomarkers that are correlated with therapy response and thus with survival is an important goal of medical research on severe diseases, e.g. cancer. Frequently, microarray studies are performed to identify genes of which the expression levels in pretherapeutic tissue samples are correlated to survival times of patients. Typically, such a study can take several years until the full planned sample size is available.Therefore, interim analyses are desirable, offering the possibility of stopping the study earlier, or of performing additional laboratory experiments to validate the role of the detected genes. While many methods correcting the multiple testing bias introduced by interim analyses have been proposed for studies of one single feature, there are still open questions about interim analyses of multiple features, particularly of high-dimensional microarray data, where the number of features clearly exceeds the number of samples. Therefore, we examine false discovery rates and power rates in microarray experiments performed during interim analyses of survival studies. In addition, the early stopping based on interim results of such studies is evaluated. As stop criterion we employ the achieved average power rate, i.e. the proportion of detected true positives, for which a new estimator is derived and compared to existing estimators. RESULTS: In a simulation study, pre-specified levels of the false discovery rate are maintained in each interim analysis, where reduced levels as used in classical group sequential designs of one single feature are not necessary. Average power rates increase with each interim analysis, and many studies can be stopped prior to their planned end when a certain pre-specified power rate is achieved. The new estimator for the power rate slightly deviates from the true power rate but is comparable to other estimators. CONCLUSIONS: Interim analyses of microarray experiments can provide evidence for early stopping of long-term survival studies. The developed simulation framework, which we also offer as a new R package 'SurvGenesInterim' available at http://survgenesinter.R-Forge.R-Project.org, can be used for sample size planning of the evaluated study design. Andreas Leha, Tim Beißbarth, Klaus Jung |
BMC Bioinform. | 3 |