EDBT 2026 Demo / reviewers in the wild / expert
Martin G. Herold
dblp:315/0899
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0002-1804-2842ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Broader View on Clustering under Cluster-Aware Norm ObjectivesabstractWe revisit the (\(f,q\))-Clustering problem that we introduced in a recent work [SODA’25]. Here, \(f\) and \(g\) are symmetric, monotone norms called inner and outer norms, respectively. The task is to partition a given set of points in a metric space into \(k\) clusters each represented by a cluster center. Each cluster is assigned a cluster cost, determined by the norm \(f\) applied to the vector of point-center distances in the cluster. The goal is to minimize the value of the norm \(g\) when applied to the vector of cluster costs. This problem subsumes fundamental clustering problems such as \(k\)-Center (i.e., \(\mathcal L_\infty, \mathcal L_\infty\)-Clustering), \(k\)-Median (i.e., \(\mathcal L_1, \mathcal L_1\)-Clustering), Min-Sum of Radii (i.e., \(\mathcal L_\infty, \mathcal L_1\)-Clustering), and Min-Load \(k\)-Clustering (i.e., \(\mathcal L_1, \mathcal L_\infty\)-Clustering). In our previous work, we focused on certain special cases of this problem for which we designed constant-factor approximation algorithms. Our bounds for more general settings left, however, large gaps to the known bounds for the basic problems they capture. Martin G. Herold, Evangelos Kipouridis, Joachim Spoerhase |
SODA | 1 |
| 2025 | Sublinear Data Structures for Nearest Neighbor in Ultra High DimensionsabstractGeometric data structures have been extensively studied in the regime where the dimension is much smaller than the number of input points. But in many scenarios in Machine Learning, the dimension can be much higher than the number of points and can be so high that the data structure might be unable to read and store all coordinates of the input and query points. Inspired by these scenarios and related studies in feature selection and explainable clustering, we initiate the study of geometric data structures in this ultra-high dimensional regime. Our focus is the approximate nearest neighbor problem. In this problem, we are given a set of n points C ⊆ ℝ^d and have to produce a small data structure that can quickly answer the following query: given q ∈ ℝ^d, return a point c ∈ C that is approximately nearest to q, where the distance is under 𝓁₁, 𝓁₂, or other norms. Many groundbreaking (1+ε)-approximation algorithms have recently been discovered for 𝓁₁- and 𝓁₂-norm distances in the regime where d≪ n. The main question in this paper is: Is there a data structure with sublinear (o(nd)) space and sublinear (o(d)) query time when d≫ n? This question can be partially answered from the machine-learning literature: - For 𝓁₁-norm distances, an Õ(log(n))-approximation data structure with Õ(n log d) space and O(n) query time can be obtained from explainable clustering techniques [Dasgupta et al. ICML'20; Makarychev and Shan ICML'21; Esfandiari, Mirrokni, and Narayanan SODA'22; Gamlath et al. NeurIPS'21; Charikar and Hu SODA'22]. - For 𝓁₂-norm distances, a (√3+ε)-approximation data structure with Õ(n log(d)/poly(ε)) space and Õ(n/poly(ε)) query time can be obtained from feature selection techniques [Boutsidis, Drineas, and Mahoney NeurIPS'09; Boutsidis et al. IEEE Trans. Inf. Theory'15; Cohen et al. STOC'15]. - For 𝓁_p-norm distances, a O(n^{p-1}log²(n))-approximation data structure with O(nlog(n) + nlog(d)) space and O(n) query time can be obtained from the explainable clustering algorithms of [Gamlath et al. NeurIPS'21]. An important open problem is whether a (1+ε)-approximation data structure exists. This is not known for any norm, even with higher (e.g. poly(n)⋅ o(d)) space and query time. In this paper, we answer this question affirmatively. We present (1+ε)-approximation data structures with the following guarantees. - For 𝓁₁- and 𝓁₂-norm distances: Õ(n log(d)/poly(ε)) space and Õ(n/poly(ε)) query time. We show that these space and time bounds are tight up to poly (log n/ε) factors. - For 𝓁_p-norm distances: Õ(n² log(d) (log log(n)/ε)^p) space and Õ (n(log log(n)/ε)^p) query time. Via simple reductions, our data structures imply sublinear-in-d data structures for some other geometric problems; e.g. approximate orthogonal range search (in the style of [Arya and Mount SoCG'95]), furthest neighbor, and give rise to a sublinear O(1)-approximate representation of k-median and k-means clustering. We hope that this paper inspires future work on sublinear geometric data structures. Martin G. Herold, Danupon Nanongkai, Joachim Spoerhase, Nithin Varma 0001, Zihang Wu |
SoCG | 1 |
| 2025 | Clustering to Minimize Cluster-Aware Norm ObjectivesabstractWe initiate the study of the following general clustering problem. We seek to partition a given set P of data points into k clusters by finding a set X of k centers and assigning each data point to one of the centers. The cost of a cluster, represented by a center x ∊ X, is a monotone, symmetric norm f (called inner norm) of the vector of distances of points assigned to x. The goal is to minimize a norm g (called outer norm) of the vector of cluster costs. This problem, which we call (f, g )-Clustering, generalizes many fundamental clustering problems such as k-Center (i.e., (𝓛∞, 𝓛∞)-Clustering), k-Median (i.e., (𝓛1, 𝓛1)-Clustering), Min-Sum of Radii (i.e., (𝓛∞, 𝓛1)-Clustering), and Min-Load k-Clustering (i.e., (𝓛1, L∞)-Clustering). A recent line of research (Byrka et al. [STOC’18], Chakrabarty, Swamy [ICALP’18, STOC’19], and Abbasi et al. [FOCS’23]) studies norm objectives that are oblivious to the cluster structure such as k-Median and k-Center. In contrast, our problem models cluster-aware objectives including Min-Sum of Radii and Min-Load k-Clustering. Martin G. Herold, Evangelos Kipouridis, Joachim Spoerhase |
SODA | 1 |
| 2023 | Upward Translation of Optimal and P-Optimal Proof Systems in the Boolean Hierarchy over NPabstractWe study the existence of optimal and p-optimal proof systems for classes in the Boolean hierarchy over $\mathrm{NP}$. Our main results concern $\mathrm{DP}$, i.e., the second level of this hierarchy: If all sets in $\mathrm{DP}$ have p-optimal proof systems, then all sets in $\mathrm{coDP}$ have p-optimal proof systems. The analogous implication for optimal proof systems fails relative to an oracle. As a consequence, we clarify such implications for all classes $\mathcal{C}$ and $\mathcal{D}$ in the Boolean hierarchy over $\mathrm{NP}$: either we can prove the implication or show that it fails relative to an oracle. Furthermore, we show that the sets $\mathrm{SAT}$ and $\mathrm{TAUT}$ have p-optimal proof systems, if and only if all sets in the Boolean hierarchy over $\mathrm{NP}$ have p-optimal proof systems which is a new characterization of a conjecture studied by Pudlák. Fabian Egidy, Christian Glaßer, Martin G. Herold |
MFCS | 3 |