EDBT 2026 Demo / reviewers in the wild / expert
Pan Peng 0001
dblp:08/9919-1
· DBLP profile ↗
59ranked-venue papers
14as first author
32since 2021 · last 2026
0000-0003-2700-5699ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 39 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sublinear-Time Algorithms for Diagonally Dominant Systems and Applications to the Friedkin-Johnsen Model
Weiming Feng 0001, Pan Peng 0001 |
COCOON | 3 |
| 2026 | An Exponential Lower Bound for Spectral Density Estimation on Unweighted GraphsabstractWe study lower bounds for estimating the spectral density of the normalized adjacency matrix of a graph. Previously, Cohen-Steiner et al. [KDD 2018] proposed an algorithm for $\varepsilon$-approximate spectral density estimation in the Wasserstein-1 distance, using $2^{O(1/\varepsilon)}$ random walks initiated from uniformly random nodes in the graph. Later, Jin et al. [COLT 2023] established a nearly matching exponential lower bound for \emph{weighted} graphs, assuming the algorithm has access to samples from random walks started at random nodes. It was left open whether this lower bound could be extended to \emph{unweighted} graphs. In this paper, we answer this question in the affirmative by proving an exponential lower bound for unweighted graphs. Specifically, we show that no algorithm can compute an $\varepsilon$-approximation to the spectrum of a normalized graph adjacency matrix with constant success probability, even when given the full transcripts of $2^{\Omega(1/\varepsilon^{1/6})}$ random walks, each of length $2^{\Omega(1/\varepsilon^{1/6})}$, started from uniformly random nodes. Pan Peng 0001, Joy Qiping Yang, Yichun Yang |
COLT | 1 |
| 2026 | Sublinear Algorithms for Estimating Single-Linkage Clustering CostsabstractSingle-linkage clustering (SLC) is a fundamental method for hierarchical data analysis. In the distance setting, a $k$-clustering produced by SLC can be obtained by computing a minimum spanning tree (MST) and deleting its $k-1$ heaviest edges. This naturally induces a cost profile for the SLC hierarchy: for each $k\in[n]$, we define $\mathrm{cost}_k$ to be the weight of the resulting $k$-component spanning forest, equivalently, the minimum total weight of any spanning forest with exactly $k$ connected components. The corresponding \emph{SLC cost profile} is $(\mathrm{cost}_1,\ldots,\mathrm{cost}_n)$, and the scalar quantity $\mathrm{cost}(G)=\sum_{k=1}^{n}\mathrm{cost}_k$ is the area under this profile. We study the problem of approximating these quantities in sublinear time. We assume that the input is a weighted graph $G$ of average degree $d$ with edge weights in $\{1,\dots,W\}$, accessed through adjacency-list queries; missing edges are treated as having infinite distance. Our main result is a sampling-based algorithm that outputs a succinct sketch of the entire SLC cost profile in the distance setting. The algorithm runs in $\widetilde{O}(d\sqrt{W}/\varepsilon^3)$ time and returns a sketch from which one can derive estimates $(\widehat{\mathrm{cost}}_1,\ldots,\widehat{\mathrm{cost}}_n)$ satisfying $\sum_{k=1}^{n}\bigl|\widehat{\mathrm{cost}}_k-\mathrm{cost}_k\bigr| \le \varepsilon\,\mathrm{cost}(G)$. Thus, we obtain an $\ell_1$ approximation to the full profile whose error is at most an $\varepsilon$-fraction of the area under the true profile. In particular, this yields a $(1\pm\varepsilon)$-approximation to $\mathrm{cost}(G)$ within the same running time. We also prove a nearly matching lower bound of $Ω(d\sqrt{W}/\varepsilon^2)$ queries for estimating $\mathrm{cost}(G)$. Pan Peng 0001, Christian Sohler |
ESA | 1 |
| 2026 | Local Computation Algorithms for (Minimum) Spanning Trees on Expander GraphsabstractWe study local computation algorithms (LCAs) for constructing spanning trees. In this setting, the goal is to determine locally, for each edge e ∈ E, whether it belongs to a spanning tree T of the input graph G, where T is defined implicitly by G and the randomness of the algorithm. It is known that sublinear-probe LCAs for spanning trees do not exist in general graphs, even for simple graph families. We identify a natural and well-studied class of graphs - expander graphs - that do admit sublinear-time LCAs for spanning trees. This is perhaps surprising, as previous work on expanders only succeeded in designing LCAs for sparse spanning subgraphs, rather than full spanning trees. We design an LCA with probe complexity O(√n ((log²n)/ϕ² + d)) for graphs with conductance at least ϕ and maximum degree at most d (not necessarily constant), which is nearly optimal when ϕ and d are constants, since Ω(√n) probes are necessary even for expanders. Next, we show that for the natural class of Erdős-Rényi graphs G(n, p) with np = n^δ for any constant δ > 0 (which are expanders with high probability), the √n lower bound can be bypassed. Specifically, we give an average-case LCA for such graphs with probe complexity Õ(√{n^{1 - δ}}). Finally, we extend our techniques to design LCAs for the minimum spanning tree (MST) problem on weighted expander graphs. Specifically, given a d-regular unweighted graph ̄{G} with sufficiently strong expansion, we consider the weighted graph G obtained by assigning to each edge an independent and uniform random weight from {1,…,W}, where W ≤ d/2 and W = o(log n). We show that there exists an LCA that is consistent with an exact MST of G, with probe complexity Õ(√nd²). Pan Peng 0001 |
ICALP | 1 |
| 2026 | Four-Cycle Counting in Low-Degeneracy Graph StreamsabstractWe study the problem of (1+ε)-approximating the number of four-cycles in graphs given as arbitrary order edge streams. We propose two new algorithms based on sampling induced subgraphs. Our first contribution is a two-pass algorithm that uses Õ(κm / √T) space, where m is the number of edges, T is the number of four-cycles, and κ is the graph's degeneracy. This algorithm improves upon existing theoretical bounds and is provably optimal for constant-degeneracy graphs, matching the known Ω(m/√T) lower bound up to lower-order factors. Our second contribution is a one-pass algorithm that remains accurate when four-cycles are not highly concentrated on individual nodes, edges, or wedges; this structural property is common in sparse social and collaboration networks. We evaluate both algorithms on a variety of real-world graph streams. The two-pass algorithm consistently outperforms state-of-the-art methods, using substantially less space to achieve a desired accuracy. The one-pass algorithm is competitive when four cycles are evenly distributed, matching our theoretical analysis. Unlike several recent works, our algorithms perform well even on non-bipartite graphs such as social networks. Sebastian Lüderssen, Stefan Neumann 0003, Pan Peng 0001 |
KDD (1) | 3 |
| 2026 | Near-Optimal Four-Cycle Counting in Graph StreamsabstractWe study four-cycle counting in arbitrary order graph streams. We present a 3-pass algorithm for \((1 + \varepsilon)\)-approximating the number of four-cycles using \(\tilde O(m/\sqrt T)\) space, where \(m\) is the number of edges and \(T\) the number of four-cycles in the graph. This improves upon a 3-pass algorithm by Vorotnikova using space \(\tilde O(m/T^{1/3})\) and matches a multi-pass lower bound of \(\Omega(m/\sqrt T)\) by McGregor and Vorotnikova. Sebastian Lüderssen, Stefan Neumann 0003, Pan Peng 0001 |
SODA | 3 |
| 2026 | Streaming algorithms for triangle counting: adversarial robustness and the weighted case
Yicheng Pan 0001, Pan Peng 0001 |
Frontiers Comput. Sci. | 3 |
| 2026 | Testing Cluster Structure of GraphsabstractWe study the problem of recognizing the spectral cluster structure of a graph in the framework of property testing in the bounded degree model. A graph is defined to be \((k,\phi_{\textrm{in}},\phi_{\textrm{out}})\) - clusterable , if it can be partitioned into no more than \( k \) parts, such that the (inner) conductance of the induced subgraph on each part is at least \(\phi_{\textrm{in}}\) and the (outer) conductance of each part is at most \(\phi_{\textrm{out}}\) . Our main result is a sublinear algorithm with the running time \(\widetilde{O}_{d,k}(\sqrt{n}\cdot\mathrm{poly}(\phi,1/\varepsilon))\) that takes as input an \( n \) -vertex graph with maximum degree bounded by \( d \) , parameters \( k \) , \(\phi\) , \(\varepsilon\) , and with probability at least \(\frac{2}{3}\) , accepts the graph if it is \((k,\phi,O_{d,k}(\varepsilon^{4}\phi^{2}))\) -clusterable, and rejects the graph if it is \(\varepsilon\) -far from \((k,\phi^{*},\psi^{*})\) -clusterable for \(\phi^{*}=O_{d,k}(\frac{\phi^{2}\varepsilon^{4}}{\log n})\) and any \(\psi^{*}\geq 0\) . By the lower bound of \(\Omega(\sqrt{n})\) on the number of queries needed for testing graph expansion, which corresponds to \(k=1\) in our problem, our algorithm is asymptotically optimal up to polylogarithmic factors. Artur Czumaj, Pan Peng 0001, Christian Sohler |
ACM Trans. Algorithms | 2 |
| 2025 | Testing Some First-Order Logic Properties on Sparse Graphs
Pan Peng 0001, Kefan Yu |
COCOON (2) | 1 |
| 2025 | Differentially Private Synthetic Graphs Preserving Triangle-Motif CutsabstractWe study the problem of releasing a differentially private (DP) synthetic graph $G’$ that well approximates the triangle-motif sizes of all cuts of any given graph $G$, where a motif in general refers to a frequently occurring subgraph within complex networks. Non-private versions of such graphs have found applications in diverse fields such as graph clustering, graph sparsification, and social network analysis. Specifically, we present the first $(\varepsilon,\delta)$-DP mechanism that, given an input graph $G$ with $n$ vertices, $m$ edges and local sensitivity of triangles $\ell_{3}(G)$, generates a synthetic graph $G’$ in polynomial time, approximating the triangle-motif sizes of all cuts $(S,V\setminus S)$ of the input graph $G$ up to an additive error of $\tilde{O}(\sqrt{m\ell_3(G)}n/\varepsilon^{3/2})$. Additionally, we provide a lower bound of $\Omega(\sqrt{mn}\ell_3(G)/\varepsilon)$ on the additive error for any DP algorithm that answers the triangle-motif size queries of all $(S,T)$-cut of $G$. Finally, our algorithm generalizes to weighted graphs, and our lower bound extends to any $K_h$-motif cut for any constant $h\geq 2$. Pan Peng 0001, Hangyu Xu |
COLT | 1 |
| 2025 | Average Sensitivity of Hierarchical k-Median ClusteringabstractHierarchical clustering is a widely used method for unsupervised learning with numerous applications. However, in the application of modern algorithms, the datasets studied are usually large and dynamic. If the hierarchical clustering is sensitive to small perturbations of the dataset, the usability of the algorithm will be greatly reduced.
In this paper, we focus on the hierarchical $k$ -median clustering problem, which bridges hierarchical and centroid-based clustering while offering theoretical appeal, practical utility, and improved interpretability. We analyze the average sensitivity of algorithms for this problem by measuring the expected change in the output when a random data point is deleted. We propose an efficient algorithm for hierarchical $k$-median clustering and theoretically prove its low average sensitivity and high clustering quality. Additionally, we show that single linkage clustering and a deterministic variant of the CLNSS algorithm exhibit high average sensitivity, making them less stable. Finally, we validate the robustness and effectiveness of our algorithm through experiments. Weiqiang He, Ruobing Bai, Pan Peng 0001 |
ICML | 4 |
| 2025 | Learning-Augmented Streaming Algorithms for Approximating MAX-CUTabstractWe study learning-augmented streaming algorithms for estimating the value of MAX-CUT in a graph. In the classical streaming model, while a 1/2-approximation for estimating the value of MAX-CUT can be trivially achieved with O(1) words of space, Kapralov and Krachun [STOC’19] showed that this is essentially the best possible: for any ε > 0, any (randomized) single-pass streaming algorithm that achieves an approximation ratio of at least 1/2 + ε requires Ω(n / 2^poly(1/ε)) space. We show that it is possible to surpass the 1/2-approximation barrier using just O(1) words of space by leveraging a (machine learned) oracle. Specifically, we consider streaming algorithms that are equipped with an ε-accurate oracle that for each vertex in the graph, returns its correct label in {-1, +1}, corresponding to an optimal MAX-CUT solution in the graph, with some probability 1/2 + ε, and the incorrect label otherwise. Within this framework, we present a single-pass algorithm that approximates the value of MAX-CUT to within a factor of 1/2 + Ω(ε²) with probability at least 2/3 for insertion-only streams, using only poly(1/ε) words of space. We also extend our algorithm to fully dynamic streams while maintaining a space complexity of poly(1/ε,log n) words. Yinhao Dong, Pan Peng 0001, Ali Vakilian |
ITCS | 2 |
| 2025 | Learning-Augmented Streaming Algorithms for Correlation ClusteringabstractWe study streaming algorithms for Correlation Clustering. Given a graph as an arbitrary-order stream of edges, with each edge labeled as positive or negative, the goal is to partition the vertices into disjoint clusters, such that the number of disagreements is minimized. In this paper, we give the first learning-augmented streaming algorithms for the problem on both complete and general graphs, improving the best-known space-approximation tradeoffs. Based on the works of Cambus et al. (SODA'24) and Ahn et al. (ICML'15), our algorithms use the predictions of pairwise distances between vertices provided by a predictor. For complete graphs, our algorithm achieves a better-than-$3$ approximation under good prediction quality, while using $\tilde{O}(n)$ total space. For general graphs, our algorithm achieves an $O(\log |E^-|)$ approximation under good prediction quality using $\tilde{O}(n)$ total space, improving the best-known non-learning algorithm in terms of space efficiency. Experimental results on synthetic and real-world datasets demonstrate the superiority of our proposed algorithms over their non-learning counterparts. Yinhao Dong, Pan Peng 0001 |
NeurIPS | 4 |
| 2024 | A Differentially Private Clustering Algorithm for Well-Clustered GraphsabstractWe study differentially private (DP) algorithms for recovering clusters in well-clustered graphs, which are graphs whose vertex set can be partitioned into a small number of sets, each inducing a subgraph of high inner conductance and small outer conductance. Such graphs have widespread application as a benchmark in the theoretical analysis of spectral clustering.
We provide an efficient ($\epsilon$,$\delta$)-DP algorithm tailored specifically for such graphs. Our algorithm draws inspiration from the recent work of Chen et al., who developed DP algorithms for recovery of stochastic block models in cases where the graph comprises exactly two nearly-balanced clusters. Our algorithm works for well-clustered graphs with $k$ nearly-balanced clusters, and the misclassification ratio almost matches the one of the best-known non-private algorithms. We conduct experimental evaluations on datasets with known ground truth clusters to substantiate the prowess of our algorithm. We also show that any (pure) $\epsilon$-DP algorithm would result in substantial error. Weiqiang He, Hendrik Fichtenberger, Pan Peng 0001 |
ICLR | 3 |
| 2024 | Sublinear-Time Opinion Estimation in the Friedkin-Johnsen ModelabstractOnline social networks are ubiquitous parts of modern societies and the discussions that take place in these networks impact people's opinions on diverse topics, such as politics or vaccination. One of the most popular models to formally describe this opinion formation process is the Friedkin--Johnsen (FJ) model, which allows to define measures, such as the polarization and the disagreement of a network. Recently, Xu, Bao and Zhang (WebConf'21) showed that all opinions and relevant measures in the FJ model can be approximated in near-linear time. However, their algorithm requires the entire network and the opinions of all nodes as input. Given the sheer size of online social networks and increasing data-access limitations, obtaining the entirety of this data might, however, be unrealistic in practice. In this paper, we show that node opinions and all relevant measures, like polarization and disagreement, can be efficiently approximated in time that is sublinear in the size of the network. Particularly, our algorithms only require query-access to the network and do not have to preprocess the graph. Furthermore, we use a connection between FJ opinion dynamics and personalized PageRank, and show that in d-regular graphs, we can deterministically approximate each node's opinion by only looking at a constant-size neighborhood, independently of the network size. We also experimentally validate that our estimation algorithms perform well in practice. Stefan Neumann 0003, Yinhao Dong, Pan Peng 0001 |
WWW | 3 |
| 2024 | On Testability of First-Order Properties in Bounded-Degree Graphs and Connections to Proximity-Oblivious TestingabstractAbstract. We study property testing of properties that are definable in first-order logic (FO) in the bounded-degree graph and relational structure models. We show that any FO property that is defined by a formula with quantifier prefix [Formula: see text] is testable (i.e., testable with constant query complexity), while there exists an FO property that is expressible by a formula with quantifier prefix [Formula: see text] that is not testable. In the dense graph model, a similar picture has long been known [N. Alon, E. Fischer, M. Krivelevich, and M. Szegedy, Combinatorica, 20 (2000), pp. 451–476] despite the very different nature of the two models. In particular, we obtain our lower bound by an FO formula that defines a class of bounded-degree expanders, based on zig-zag products of graphs. We expect this to be of independent interest. We then use our class of FO definable bounded-degree expanders to answer a long-standing open problem for proximity-oblivious testers (POTs). POTs are a class of particularly simple testing algorithms, where a basic test is performed a number of times that may depend on the proximity parameter, but the basic test itself is independent of the proximity parameter. In their seminal work, Goldreich and Ron [STOC 2009; SIAM J. Comput., 40 (2011), pp. 534–566] show that the graph properties that are constant-query proximity-oblivious testable in the bounded-degree model are precisely the properties that can be expressed as a generalized subgraph freeness (GSF) property that satisfies the non-propagation condition. It is left open whether the non-propagation condition is necessary. Indeed, calling properties expressible as a generalized subgraph freeness property GSF-local properties, they ask whether all GSF-local properties are non-propagating. We give a negative answer by showing that our FO definable property is GSF-local and propagating. Hence, in particular, our property does not admit a POT, despite being GSF-local. For this result we establish a new connection between FO properties and GSF-local properties via neighborhood profiles. Isolde Adler, Noleen Köhler, Pan Peng 0001 |
SIAM J. Comput. | 3 |
| 2023 | Effective Resistances in Non-Expander GraphsabstractEffective resistances are ubiquitous in graph algorithms and network analysis. In this work, we study sublinear time algorithms to approximate the effective resistance of an adjacent pair $s$ and $t$. We consider the classical adjacency list model for local algorithms. While recent works have provided sublinear time algorithms for expander graphs, we prove several lower bounds for general graphs of $n$ vertices and $m$ edges: 1.It needs $Ω(n)$ queries to obtain $1.01$-approximations of the effective resistance of an adjacent pair $s$ and $t$, even for graphs of degree at most 3 except $s$ and $t$. 2.For graphs of degree at most $d$ and any parameter $\ell$, it needs $Ω(m/\ell)$ queries to obtain $c \cdot \min\{d, \ell\}$-approximations where $c>0$ is a universal constant. Moreover, we supplement the first lower bound by providing a sublinear time $(1+ε)$-approximation algorithm for graphs of degree 2 except the pair $s$ and $t$. One of our technical ingredients is to bound the expansion of a graph in terms of the smallest non-trivial eigenvalue of its Laplacian matrix after removing edges. We discover a new lower bound on the eigenvalues of perturbed graphs (resp. perturbed matrices) by incorporating the effective resistance of the removed edge (resp. the leverage scores of the removed rows), which may be of independent interest. Dongrun Cai, Pan Peng 0001 |
ESA | 3 |
| 2023 | Massively Parallel Algorithms for the Stochastic Block ModelabstractLearning the community structure of a large-scale graph is a fundamental problem in machine learning, computer science, and statistics. Among others, the Stochastic Block Model (SBM) serves as a canonical model for community detection and clustering, and the Massively Parallel Computation (MPC) model is a mathematical abstraction of real-world parallel computing systems which provides a powerful computational framework for handling large-scale datasets. We study the problem of exactly recovering the communities in a graph generated from the SBM in the MPC model. Specifically, given kn vertices that are partitioned into k equal-sized clusters (i.e., each has size n), a graph on these kn vertices is randomly generated such that each pair of vertices is connected with probability p if they are in the same cluster and with probability q if not, where p > q > 0. We give an MPC algorithm that recovers the ground-truth clusters when (p-q)/√p ≥˜Ω(k1/2n(-1/2+1/(2r-2))) for any integer r ∈ [3, O(log n)] in O(kr/δ) rounds in the sublinear space MPC model, where each machine has local memory O(nδ) for some constant δ > 0. When (p-q)/√p≥ ˜Ω(k3/4n-1/4), we also give an MPC clustering algorithm that works in O(logs n) rounds in the s-space MPC model where each machine is only guaranteed to have memory s = Ω(log n). To implement the latter algorithm, we propose new algorithms for some basic graph operations in the s-space MPC model. Both algorithms significantly improve upon a recent result of Cohen-Addad et al. [PODC'22], who gave an algorithm that only works in the sublinear space MPC model with a much stronger condition on p, q, k. Our algorithms are based on collecting the r-step neighborhood of each vertex and comparing the difference of some statistical information generated from the local neighborhoods for each pair of vertices. Pan Peng 0001, Xianbin Zhu 0002 |
ESA | 2 |
| 2023 | An Optimal Separation Between Two Property Testing Models for Bounded Degree Directed Graphs
Pan Peng 0001 |
ICALP | 1 |
| 2023 | Recovering Unbalanced Communities in the Stochastic Block Model with Application to Clustering with a Faulty OracleabstractThe stochastic block model (SBM) is a fundamental model for studying graph clustering or community detection in networks. It has received great attention in the last decade and the balanced case, i.e., assuming all clusters have large size, has been well studied.
However, our understanding of SBM with unbalanced communities (arguably, more relevant in practice) is still limited. In this paper, we provide a simple SVD-based algorithm for recovering the communities in the SBM with communities of varying sizes.
We improve upon a result of Ailon, Chen and Xu [ICML 2013; JMLR 2015] by removing the assumption that there is a large interval such that the sizes of clusters do not fall in, and also remove the dependency of the size of the recoverable clusters on the number of underlying clusters. We further complement our theoretical improvements with experimental comparisons.
Under the planted clique conjecture, the size of the clusters that can be recovered by our algorithm is nearly optimal (up to poly-logarithmic factors) when the probability parameters are constant.
As a byproduct, we obtain an efficient clustering algorithm with sublinear query complexity in a faulty oracle model, which is capable of detecting all clusters larger than $\tilde{\Omega}({\sqrt{n}})$, even in the presence of $\Omega(n)$ small clusters in the graph. In contrast, previous efficient algorithms that use a sublinear number of queries are incapable of recovering any large clusters if there are more than $\tilde{\Omega}(n^{2/5})$ small clusters. Chandra Sekhar Mukherjee, Pan Peng 0001 |
NeurIPS | 2 |
| 2023 | A Sublinear-Time Spectral Clustering Oracle with Improved Preprocessing TimeabstractWe address the problem of designing a sublinear-time spectral clustering oracle for graphs that exhibit strong clusterability. Such graphs contain $k$ latent clusters, each characterized by a large inner conductance (at least $\varphi$) and a small outer conductance (at most $\varepsilon$). Our aim is to preprocess the graph to enable clustering membership queries, with the key requirement that both preprocessing and query answering should be performed in sublinear time, and the resulting partition should be consistent with a $k$-partition that is close to the ground-truth clustering. Previous oracles have relied on either a $\textrm{poly}(k)\log n$ gap between inner and outer conductances or exponential (in $k/\varepsilon$) preprocessing time. Our algorithm relaxes these assumptions, albeit at the cost of a slightly higher misclassification ratio. We also show that our clustering oracle is robust against a few random edge deletions. To validate our theoretical bounds, we conducted experiments on synthetic networks. Ranran Shen, Pan Peng 0001 |
NeurIPS | 2 |
| 2023 | Sublinear-Time Algorithms for Max Cut, Max E2Lin(q), and Unique Label Cover on ExpandersabstractWe show sublinear-time algorithms for MAX CUT and MAX E2LIN(q) on expanders in the adjacency list model that distinguishes instances with the optimal value more than 1 − ε from those with the optimal value less than 1 − ρ for ρ ≫ ε. The time complexities for MAX CUT and MAX 2LIN(q) are and , respectively, where m is the number of edges in the underlying graph and ϕ is its conductance. Then, we show a sublinear-time algorithm for UNIQUE LABEL COVER on expanders with ϕ ≫ ε in the bounded-degree model. The time complexity of our algorithm is Õd(2qO(1)·ϕ1/q·ε-1/2 · n1/2+qO(q)·ε41.5-q ·ϕ-2), where n is the number of variables. We complement these algorithmic results by showing that testing 3-colorability requires Ω(n) queries even on expanders. Pan Peng 0001, Yuichi Yoshida |
SODA | 1 |
| 2022 | Sublinear-Time Clustering Oracle for Signed GraphsabstractSocial networks are often modeled using signed graphs, where vertices correspond to users and edges have a sign that indicates whether an interaction between users was positive or negative. The arising signed graphs typically contain a clear community structure in the sense that the graph can be partitioned into a small number of polarized communities, each defining a sparse cut and indivisible into smaller polarized sub-communities. We provide a local clustering oracle for signed graphs with such a clear community structure, that can answer membership queries, i.e., “Given a vertex $v$, which community does $v$ belong to?”, in sublinear time by reading only a small portion of the graph. Formally, when the graph has bounded maximum degree and the number of communities is at most $O(\log n)$, then with $\tilde{O}(\sqrt{n}\operatorname{poly}(1/\varepsilon))$ preprocessing time, our oracle can answer each membership query in $\tilde{O}(\sqrt{n}\operatorname{poly}(1/\varepsilon))$ time, and it correctly classifies a $(1-\varepsilon)$-fraction of vertices w.r.t. a set of hidden planted ground-truth communities. Our oracle is desirable in applications where the clustering information is needed for only a small number of vertices. Previously, such local clustering oracles were only known for unsigned graphs; our generalization to signed graphs requires a number of new ideas and gives a novel spectral analysis of the behavior of random walks with signs. We evaluate our algorithm for constructing such an oracle and answering membership queries on both synthetic and real-world datasets, validating its performance in practice. Stefan Neumann 0003, Pan Peng 0001 |
ICML | 2 |
| 2022 | Approximately Counting Subgraphs in Data StreamsabstractEstimating the number of subgraphs in data streams is a fundamental problem that has received great attention in the past decade. In this paper, we give improved streaming algorithms for approximately counting the number of occurrences of an arbitrary subgraph H, denoted #H, when the input graph G is represented as a stream of m edges. To obtain our algorithms, we provide a generic transformation that converts constant-round sublinear-time graph algorithms in the query access model to constant-pass sublinear-space graph streaming algorithms. Using this transformation, we obtain the following results. Hendrik Fichtenberger, Pan Peng 0001 |
PODS | 2 |
| 2022 | Constant-time Dynamic (Δ +1)-ColoringabstractWe give a fully dynamic (Las-Vegas style) algorithm with constant expected amortized time per update that maintains a proper (Δ +1)-vertex coloring of a graph with maximum degree at most Δ. This improves upon the previous O (log Δ)-time algorithm by Bhattacharya et al. (SODA 2018). Our algorithm uses an approach based on assigning random ranks to vertices and does not need to maintain a hierarchical graph decomposition. We show that our result does not only have optimal running time but is also optimal in the sense that already deciding whether a Δ-coloring exists in a dynamically changing graph with maximum degree at most Δ takes Ω (log n ) time per operation. Monika Henzinger, Pan Peng 0001 |
ACM Trans. Algorithms | 2 |
| 2021 | GSF-Locality Is Not Sufficient For Proximity-Oblivious TestingabstractIn Property Testing, proximity-oblivious testers (POTs) form a class of particularly simple testing algorithms, where a basic test is performed a number of times that may depend on the proximity parameter, but the basic test itself is independent of the proximity parameter. In their seminal work, Goldreich and Ron [STOC 2009; SICOMP 2011] show that the graph properties that allow constant-query proximity-oblivious testing in the bounded-degree model are precisely the properties that can be expressed as a generalised subgraph freeness (GSF) property that satisfies the non-propagation condition. It is left open whether the non-propagation condition is necessary. Indeed, calling properties expressible as a generalised subgraph freeness property GSF-local properties, they ask whether all GSF-local properties are non-propagating. We give a negative answer by exhibiting a property of graphs that is GSF-local and propagating. Hence in particular, our property does not admit a POT, despite being GSF-local. We prove our result by exploiting a recent work of the authors which constructed a first-order (FO) property that is not testable [SODA 2021], and a new connection between FO properties and GSF-local properties via neighbourhood profiles. Isolde Adler, Noleen Köhler, Pan Peng 0001 |
CCC | 3 |
| 2021 | Towards a Query-Optimal and Time-Efficient Algorithm for Clustering with a Faulty OracleabstractMotivated by applications in crowdsourced entity resolution in database, signed edge prediction in social networks and correlation clustering, Mazumdar and Saha [NIPS 2017] proposed an elegant theoretical model for studying clustering with a faulty oracle. In this model, given a set of $n$ items which belong to $k$ unknown groups (or clusters), our goal is to recover the clusters by asking pairwise queries to an oracle. This oracle can answer the query that “do items $u$ and $v$ belong to the same cluster?”. However, the answer to each pairwise query errs with probability $\epsilon$, for some $\epsilon\in(0,\frac12)$. Mazumdar and Saha provided two algorithms under this model: one algorithm is query-optimal while time-inefficient (i.e., running in quasi-polynomial time), the other is time efficient (i.e., in polynomial time) while query-suboptimal. Larsen, Mitzenmacher and Tsourakakis [WWW 2020] then gave a new time-efficient algorithm for the special case of $2$ clusters, which is query-optimal if the bias $\delta:=1-2\epsilon$ of the model is large. It was left as an open question whether one can obtain a query-optimal, time-efficient algorithm for the general case of $k$ clusters and other regimes of $\delta$. In this paper, we make progress on the above question and provide a time-efficient algorithm with nearly-optimal query complexity (up to a factor of $O(\log^2 n)$) for all constant $k$ and any $\delta$ in the regime when information-theoretic recovery is possible. Our algorithm is built on a connection to the stochastic block model. Pan Peng 0001 |
COLT | 1 |
| 2021 | Local Algorithms for Estimating Effective ResistanceabstractEffective resistance is an important metric that measures the similarity of two vertices in a graph. It has found applications in graph clustering, recommendation systems and network reliability, among others. In spite of the importance of the effective resistances, we still lack efficient algorithms to exactly compute or approximate them on massive graphs. Pan Peng 0001, Daniel Lopatta, Yuichi Yoshida, Gramoz Goranci |
KDD | 1 |
| 2021 | On Testability of First-Order Properties in Bounded-Degree GraphsabstractWe study property testing of properties that are definable in first-order logic (FO) in the bounded-degree graph and relational structure models. We show that any FO property that is defined by a formula with quantifier prefix ∃∗∀∗ is testable (i.e., testable with constant query complexity), while there exists an FO property that is expressible by a formula with quantifier prefix ∀∗∃∗ that is not testable. In the dense graph model, a similar picture is long known (Alon, Fischer, Krivelevich, Szegedy, Combinatorica 2000), despite the very different nature of the two models. In particular, we obtain our lower bound by a first-order formula that defines a class of bounded-degree expanders, based on zig-zag products of graphs. We expect this to be of independent interest. We then prove testability of some first-order properties that speak about isomorphism types of neighbourhoods, including testability of 1-neighbourhood-freeness, and r-neighbourhood-freeness under a mild assumption on the degrees. Isolde Adler, Noleen Köhler, Pan Peng 0001 |
SODA | 3 |
| 2021 | Time Complexity Analysis of Randomized Search Heuristics for the Dynamic Graph Coloring ProblemabstractAbstract We contribute to the theoretical understanding of randomized search heuristics for dynamic problems. We consider the classical vertex coloring problem on graphs and investigate the dynamic setting where edges are added to the current graph. We then analyze the expected time for randomized search heuristics to recompute high quality solutions. The (1+1) Evolutionary Algorithm and RLS operate in a setting where the number of colors is bounded and we are minimizing the number of conflicts. Iterated local search algorithms use an unbounded color palette and aim to use the smallest colors and, consequently, the smallest number of colors. We identify classes of bipartite graphs where reoptimization is as hard as or even harder than optimization from scratch, i.e., starting with a random initialization. Even adding a single edge can lead to hard symmetry problems. However, graph classes that are hard for one algorithm turn out to be easy for others. In most cases our bounds show that reoptimization is faster than optimizing from scratch. We further show that tailoring mutation operators to parts of the graph where changes have occurred can significantly reduce the expected reoptimization time. In most settings the expected reoptimization time for such tailored algorithms is linear in the number of added edges. However, tailored algorithms cannot prevent exponential times in settings where the original algorithm is inefficient. Jakob Bossek, Frank Neumann 0001, Pan Peng 0001, Dirk Sudholt |
Algorithmica | 3 |
| 2021 | Constant-time dynamic weight approximation for minimum spanning forest
Monika Henzinger, Pan Peng 0001 |
Inf. Comput. | 2 |
| 2021 | Mixed-order spectral clustering for complex networks
Yan Ge 0002, Pan Peng 0001, Haiping Lu |
Pattern Recognit. | 2 |
| 2020 | Testable Properties in General Graphs and Random Order StreamingabstractWe consider the fundamental question of understanding the relative power of two important computational models: property testing and data streaming. We present a novel framework closely linking these areas in the setting of general graphs in the context of constant-query complexity testing and constant-space streaming. Our main result is a generic transformation of a one-sided error property tester in the random-neighbor model with constant query complexity into a one-sided error property tester in the streaming model with constant space complexity. Previously such a generic transformation was only known for bounded-degree graphs. Artur Czumaj, Hendrik Fichtenberger, Pan Peng 0001, Christian Sohler |
APPROX-RANDOM | 3 |
| 2020 | Augmenting the Algebraic Connectivity of GraphsabstractFor any undirected graph $G=(V,E)$ and a set $E_W$ of candidate edges with $E\cap E_W=\emptyset$, the $(k,γ)$-spectral augmentability problem is to find a set $F$ of $k$ edges from $E_W$ with appropriate weighting, such that the algebraic connectivity of the resulting graph $H=(V,E\cup F)$ is least $γ$. Because of a tight connection between the algebraic connectivity and many other graph parameters, including the graph's conductance and the mixing time of random walks in a graph, maximising the resulting graph's algebraic connectivity by adding a small number of edges has been studied over the past 15 years. In this work we present an approximate and efficient algorithm for the $(k,γ)$-spectral augmentability problem, and our algorithm runs in almost-linear time under a wide regime of parameters. Our main algorithm is based on the following two novel techniques developed in the paper, which might have applications beyond the $(k,γ)$-spectral augmentability problem. (1) We present a fast algorithm for solving a feasibility version of an SDP for the algebraic connectivity maximisation problem from [GB06]. Our algorithm is based on the classic primal-dual framework for solving SDP, which in turn uses the multiplicative weight update algorithm. We present a novel approach of unifying SDP constraints of different matrix and vector variables and give a good separation oracle accordingly. (2) We present an efficient algorithm for the subgraph sparsification problem, and for a wide range of parameters our algorithm runs in almost-linear time, in contrast to the previously best known algorithm running in at least $Ω(n^2mk)$ time [KMST10]. Our analysis shows how the randomised BSS framework can be generalised in the setting of subgraph sparsification, and how the potential functions can be applied to approximately keep track of different subspaces. Bogdan-Adrian Manghiuc, Pan Peng 0001, He Sun 0001 |
ESA | 2 |
| 2020 | More effective randomized search heuristics for graph coloring through dynamic optimizationabstractDynamic optimization problems have gained significant attention in evolutionary computation as evolutionary algorithms (EAs) can easily adapt to changing environments. We show that EAs can solve the graph coloring problem for bipartite graphs more efficiently by using dynamic optimization. In our approach the graph instance is given incrementally such that the EA can reoptimize its coloring when a new edge introduces a conflict. We show that, when edges are inserted in a way that preserves graph connectivity, Randomized Local Search (RLS) efficiently finds a proper 2-coloring for all bipartite graphs. This includes graphs for which RLS and other EAs need exponential expected time in a static optimization scenario. We investigate different ways of building up the graph by popular graph traversals such as breadth-first-search and depth-first-search and analyse the resulting runtime behavior. We further show that offspring populations (e. g. a (1 + λ) RLS) lead to an exponential speedup in λ. Finally, an island model using 3 islands succeeds in an optimal time of Θ(m) on every m-edge bipartite graph, outperforming offspring populations. This is the first example where an island model guarantees a speedup that is not bounded in the number of islands. Jakob Bossek, Frank Neumann 0001, Pan Peng 0001, Dirk Sudholt |
GECCO | 3 |
| 2020 | Sampling Arbitrary Subgraphs Exactly Uniformly in Sublinear TimeabstractGiven query access to an undirected graph $G$, we consider the problem of computing a $(1\pmε)$-approximation of the number of $k$-cliques in $G$. The standard query model for general graphs allows for degree queries, neighbor queries, and pair queries. Let $n$ be the number of vertices, $m$ be the number of edges, and $n_k$ be the number of $k$-cliques. Previous work by Eden, Ron and Seshadhri (STOC 2018) gives an $O^*(\frac{n}{n^{1/k}_k} + \frac{m^{k/2}}{n_k})$-time algorithm for this problem (we use $O^*(\cdot)$ to suppress $\poly(\log n, 1/ε, k^k)$ dependencies). Moreover, this bound is nearly optimal when the expression is sublinear in the size of the graph. Our motivation is to circumvent this lower bound, by parameterizing the complexity in terms of \emph{graph arboricity}. The arboricity of $G$ is a measure for the graph density "everywhere". We design an algorithm for the class of graphs with arboricity at most $α$, whose running time is $O^*(\min\{\frac{nα^{k-1}}{n_k},\, \frac{n}{n_k^{1/k}}+\frac{m α^{k-2}}{n_k} \})$. We also prove a nearly matching lower bound. For all graphs, the arboricity is $O(\sqrt m)$, so this bound subsumes all previous results on sublinear clique approximation. As a special case of interest, consider minor-closed families of graphs, which have constant arboricity. Our result implies that for any minor-closed family of graphs, there is a $(1\pmε)$-approximation algorithm for $n_k$ that has running time $O^*(\frac{n}{n_k})$. Such a bound was not known even for the special (classic) case of triangle counting in planar graphs. Hendrik Fichtenberger, Pan Peng 0001 |
ICALP | 3 |
| 2020 | Average Sensitivity of Spectral ClusteringabstractSpectral clustering is one of the most popular clustering methods for finding clusters in a graph, which has found many applications in data mining. However, the input graph in those applications may have many missing edges due to error in measurement, withholding for a privacy reason, or arbitrariness in data conversion. To make reliable and efficient decisions based on spectral clustering, we assess the stability of spectral clustering against edge perturbations in the input graph using the notion of average sensitivity, which is the expected size of the symmetric difference of the output clusters before and after we randomly remove edges. We first prove that the average sensitivity of spectral clustering is proportional to $łambda_2/łambda_3^2$, where $łambda_i$ is the i-th smallest eigenvalue of the (normalized) Laplacian. We also prove an analogous bound for k-way spectral clustering, which partitions the graph into k clusters. Then, we empirically confirm our theoretical bounds by conducting experiments on synthetic and real networks. Our results suggest that spectral clustering is stable against edge perturbations when there is a cluster structure in the input graph. Pan Peng 0001, Yuichi Yoshida |
KDD | 1 |
| 2020 | Robust Clustering Oracle and Local Reconstructor of Cluster Structure of GraphsabstractWe develop sublinear time algorithms for analyzing the cluster structure of graphs with noisy partial information. A graph G with maximum degree at most d is called (k, φin, φout)-clusterable, if it can be partitioned into at most k parts, such that each part has inner conductance at least φin and outer conductance at most φout, where d is assumed to be constant. A graph G is called to be an ∊-perturbation of a (k,φin, φout)-clusterable graph if there is partition of G with at most k parts (called clusters), such that one can insert/delete at most ϵdn intra-cluster edges to make it a (k,φin,φout)-clusterable graph. We are given query access to the adjacency list of such a graph. We show that one can construct in time a robust clustering oracle for a bounded-degree graph G that is an ∊-perturbation of a -clusterable graph. Using such an oracle, a typical clustering query (e.g., IsOutlier(s), SameCluster(s, t)) can be answered in time and the answers are consistent with a partition of G in which all but vertices belong to a good cluster, i.e., a set with inner conductance at least , and outer conductance . We also develop a local reconstruction algorithm that takes as input a graph as above, and on any query vertex v, outputs all its neighbors in the reconstructed graph G’, which is guaranteed to be -clusterable (with slightly boosting degree bound). The number of edges changed is at most . Furthermore, the algorithm runs in time (per query) and can answer consistently with the same G′ for any sequence of queries it gets. Pan Peng 0001 |
SODA | 1 |
| 2020 | Constant-Time Dynamic (Δ+1)-Coloring
Monika Henzinger, Pan Peng 0001 |
STACS | 2 |
| 2020 | Improved Guarantees for Vertex Sparsification in Planar GraphsabstractGraph sparsification aims at compressing large graphs into smaller ones while preserving important characteristics of the input graph. In this work we study vertex sparsifiers, i.e., sparsifiers whose goal is to reduce the number of vertices. We focus on the following notions: (1) Given a digraph $G=(V,E)$ and terminal vertices $K \subset V$ with $|K| = k$, a (vertex) reachability sparsifier of $G$ is a digraph $H=(V_H,E_H)$, $K \subset V_H$ that preserves all reachability information among terminal pairs. Let $|V_H|$ denote the size of $H$. In this work we introduce the notion of reachability-preserving minors (RPMs), i.e., we require $H$ to be a minor of $G$. We show any directed graph $G$ admits an RPM $H$ of size $O(k^3)$, and if $G$ is planar, then the size of $H$ improves to $O(k^{2} \log k)$. We complement our upper bound by showing that there exists an infinite family of grids such that any RPM must have $\Omega(k^{2})$ vertices. (2) Given a weighted undirected graph $G=(V,E)$ and terminal vertices $K$ with $|K|=k$, an exact (vertex) cut sparsifier of $G$ is a graph $H$ with $K \subset V_H$ that preserves the value of minimum cuts separating any bipartition of $K$. We show that planar graphs with all the $k$ terminals lying on the same face admit exact cut sparsifiers of size $O(k^{2})$ that are also planar. Our result extends to flow and distance sparsifiers. It improves the previous best-known bound of $O(k^22^{2k})$ for cut and flow sparsifiers by an exponential factor and matches an $\Omega(k^2)$ lower-bound for this class of graphs. Gramoz Goranci, Monika Henzinger, Pan Peng 0001 |
SIAM J. Discret. Math. | 3 |
| 2019 | Runtime analysis of randomized search heuristics for dynamic graph coloringabstractWe contribute to the theoretical understanding of randomized search heuristics for dynamic problems. We consider the classical graph coloring problem and investigate the dynamic setting where edges are added to the current graph. We then analyze the expected time for randomized search heuristics to recompute high quality solutions. This includes the (1+1) EA and RLS in a setting where the number of colors is bounded and we are minimizing the number of conflicts as well as iterated local search algorithms that use an unbounded color palette and aim to use the smallest colors and - as a consequence - the smallest number of colors. Jakob Bossek, Frank Neumann 0001, Pan Peng 0001, Dirk Sudholt |
GECCO | 3 |
| 2019 | Every Testable (Infinite) Property of Bounded-Degree Graphs Contains an Infinite Hyperfinite SubpropertyabstractOne of the most fundamental questions in graph property testing is to characterize the combinatorial structure of properties that are testable with a constant number of queries. We work towards an answer to this question for the bounded-degree graph model introduced in [GR02], where the input graphs have maximum degree bounded by a constant d. In this model, it is known (among other results) that every hyperfinite property is constant-query testable [NS13], where, informally, a graph property is hyperfinite, if for every δ > 0 every graph in the property can be partitioned into small connected components by removing δn edges. In this paper we show that hyperfiniteness plays a role in every testable property, i.e. we show that every testable property is either finite (which trivially implies hyperfiniteness and testability) or contains an infinite hyperfinite subproperty. A simple consequence of our result is that no infinite graph property that only consists of expander graphs is constant-query testable. Based on the above findings, one could ask if every infinite testable non-hyperfinite property might contain an infinite family of expander (or near-expander) graphs. We show that this is not true. Motivated by our counterexample we develop a theorem that shows that we can partition the set of vertices of every bounded degree graph into a constant number of subsets and a separator set, such that the separator set is small and the distribution of k-discs on every subset of a partition class, is roughly the same as that of the partition class if the subset has small expansion. Hendrik Fichtenberger, Pan Peng 0001, Christian Sohler |
SODA | 2 |
| 2019 | Dynamic Graph Stream Algorithms in o(n) Spaceabstractis necessary. Zengfeng Huang, Pan Peng 0001 |
Algorithmica | 2 |
| 2019 | Spectral concentration and greedy k-clustering
Tamal K. Dey, Pan Peng 0001, Alfred Rossi, Anastasios Sidiropoulos |
Comput. Geom. | 2 |
| 2018 | Dynamic Effective Resistances and Approximate Schur Complement on Separable GraphsabstractWe consider the problem of dynamically maintaining (approximate) all-pairs effective resistances in separable graphs, which are those that admit an $n^{c}$-separator theorem for some $c<1$. We give a fully dynamic algorithm that maintains $(1+\varepsilon)$-approximations of the all-pairs effective resistances of an $n$-vertex graph $G$ undergoing edge insertions and deletions with $\tilde{O}(\sqrt{n}/\varepsilon^2)$ worst-case update time and $\tilde{O}(\sqrt{n}/\varepsilon^2)$ worst-case query time, if $G$ is guaranteed to be $\sqrt{n}$-separable (i.e., it is taken from a class satisfying a $\sqrt{n}$-separator theorem) and its separator can be computed in $\tilde{O}(n)$ time. Our algorithm is built upon a dynamic algorithm for maintaining \emph{approximate Schur complement} that approximately preserves pairwise effective resistances among a set of terminals for separable graphs, which might be of independent interest. We complement our result by proving that for any two fixed vertices $s$ and $t$, no incremental or decremental algorithm can maintain the $s-t$ effective resistance for $\sqrt{n}$-separable graphs with worst-case update time $O(n^{1/2-δ})$ and query time $O(n^{1-δ})$ for any $δ>0$, unless the Online Matrix Vector Multiplication (OMv) conjecture is false. We further show that for \emph{general} graphs, no incremental or decremental algorithm can maintain the $s-t$ effective resistance problem with worst-case update time $O(n^{1-δ})$ and query-time $O(n^{2-δ})$ for any $δ>0$, unless the OMv conjecture is false. Gramoz Goranci, Monika Henzinger, Pan Peng 0001 |
ESA | 3 |
| 2018 | Estimating Graph Parameters from Random Order StreamsabstractWe develop a new algorithmic technique that allows to transfer some constant time approximation algorithms for general graphs into random order streaming algorithms. We illustrate our technique by proving that in random order streams with probability at least 2/3, the number of connected components of G can be approximated up to an additive error of εn using space, the weight of a minimum spanning tree of a connected input graph with integer edges weights from {1, …, W} can be approximated within a multiplicative factor of 1 + ε using space, the size of a maximum independent set in planar graphs can be approximated within a multiplicative factor of 1+ε using space . Pan Peng 0001, Christian Sohler |
SODA | 1 |
| 2017 | Improved Guarantees for Vertex Sparsification in Planar GraphsabstractGiven an edge-weighted graph G with a set Q of k terminals, a mimicking network is a graph with the same set of terminals that exactly preserves the sizes of minimum cuts between any partition of the terminals. A natural question in the area of graph compression is to provide as small mimicking networks as possible for input graph G being either an arbitrary graph or coming from a specific graph class. In this note we show an exponential lower bound for cut mimicking networks in planar graphs: there are edge-weighted planar graphs with k terminals that require 2^(k-2) edges in any mimicking network. This nearly matches an upper bound of O(k * 2^(2k)) of Krauthgamer and Rika [SODA 2013, arXiv:1702.05951] and is in sharp contrast with the O(k^2) upper bound under the assumption that all terminals lie on a single face [Goranci, Henzinger, Peng, arXiv:1702.01136]. As a side result we show a hard instance for the double-exponential upper bounds given by Hagerup, Katajainen, Nishimura, and Ragde [JCSS 1998], Khan and Raghavendra [IPL 2014], and Chambers and Eppstein [JGAA 2013]. Gramoz Goranci, Monika Henzinger, Pan Peng 0001 |
ESA | 3 |
| 2017 | The Power of Vertex Sparsifiers in Dynamic Graph AlgorithmsabstractWe introduce a new algorithmic framework for designing dynamic graph algorithms in minor-free graphs, by exploiting the structure of such graphs and a tool called vertex sparsification, which is a way to compress large graphs into small ones that well preserve relevant properties among a subset of vertices and has previously mainly been used in the design of approximation algorithms. Using this framework, we obtain a Monte Carlo randomized fully dynamic algorithm for (1 + epsilon)-approximating the energy of electrical flows in n-vertex planar graphs with tilde{O}(r epsilon^{-2}) worst-case update time and tilde{O}((r + n / sqrt{r}) epsilon^{-2}) worst-case query time, for any r larger than some constant. For r=n^{2/3}, this gives tilde{O}(n^{2/3} epsilon^{-2}) update time and tilde{O}(n^{2/3} epsilon^{-2}) query time. We also extend this algorithm to work for minor-free graphs with similar approximation and running time guarantees. Furthermore, we illustrate our framework on the all-pairs max flow and shortest path problems by giving corresponding dynamic algorithms in minor-free graphs with both sublinear update and query times. To the best of our knowledge, our results are the first to systematically establish such a connection between dynamic graph algorithms and vertex sparsification. We also present both upper bound and lower bound for maintaining the energy of electrical flows in the incremental subgraph model, where updates consist of only vertex activations, which might be of independent interest. Gramoz Goranci, Monika Henzinger, Pan Peng 0001 |
ESA | 3 |
| 2017 | Testable Bounded Degree Graph Properties Are Random Order StreamableabstractWe study which property testing and sublinear time algorithms can be transformed into graph streaming algorithms for random order streams. Our main result is that for bounded degree graphs, any property that is constant-query testable in the adjacency list model can be tested with constant space in a single-pass in random order streams. Our result is obtained by estimating the distribution of local neighborhoods of the vertices on a random order graph stream using constant space. We then show that our approach can also be applied to constant time approximation algorithms for bounded degree graphs in the adjacency list model: As an example, we obtain a constant-space single-pass random order streaming algorithms for approximating the size of a maximum matching with additive error epsilon n (n is the number of nodes). Our result establishes for the first time that a large class of sublinear algorithms can be simulated in random order streams, while Omega(n) space is needed for many graph streaming problems for adversarial orders. Morteza Monemizadeh, S. Muthukrishnan 0001, Pan Peng 0001, Christian Sohler |
ICALP | 3 |
| 2016 | Dynamic Graph Stream Algorithms in o(n) Space
Zengfeng Huang, Pan Peng 0001 |
ICALP | 2 |
| 2016 | Relating two property testing models for bounded degree directed graphsabstractWe study property testing algorithms in directed graphs (digraphs) with maximum indegree and maximum outdegree upper bounded by d. For directed graphs with bounded degree, there are two different models in property testing introduced by Bender and Ron (2002). In the bidirectional model, one can access both incoming and outgoing edges while in the unidirectional model one can only access outgoing edges. In our paper we provide a new relation between the two models: we prove that if a property can be tested with constant query complexity in the bidirectional model, then it can be tested with sublinear query complexity in the unidirectional model. A corollary of this result is that in the unidirectional model (the model allowing only queries to the outgoing neighbors), every property in hyperfinite digraphs is testable with sublinear query complexity. Artur Czumaj, Pan Peng 0001, Christian Sohler |
STOC | 2 |
| 2015 | On Constant-Size Graphs That Preserve the Local Structure of High-Girth GraphsabstractLet G=(V,E) be an undirected graph with maximum degree d. The k-disc of a vertex v is defined as the rooted subgraph that is induced by all vertices whose distance to v is at most k. The k-disc frequency vector of G, freq(G), is a vector indexed by all isomorphism types of k-discs. For each such isomorphism type Gamma, the k-disc frequency vector counts the fraction of vertices that have k-disc isomorphic to Gamma. Thus, the frequency vector freq(G) of G captures the local structure of G. A natural question is whether one can construct a much smaller graph H such that H has a similar local structure. N. Alon proved that for any epsilon>0 there always exists a graph H whose size is independent of |V| and whose frequency vector satisfies ||freq(G) - freq(G)||_1 <= epsilon. However, his proof is only existential and neither gives an explicit bound on the size of H nor an efficient algorithm. He gave the open problem to find such explicit bounds. In this paper, we solve this problem for the special case of high girth graphs. We show how to efficiently compute a graph H with the above properties when G has girth at least 2k+2 and we give explicit bounds on the size of H. Hendrik Fichtenberger, Pan Peng 0001, Christian Sohler |
APPROX-RANDOM | 2 |
| 2015 | Testing Small Set Expansion in General GraphsabstractWe consider the problem of testing small set expansion for general graphs. A graph G is a (k,\phi)-expander if every subset of volume at most k has conductance at least \phi. Small set expansion has recently received significant attention due to its close connection to the unique games conjecture, the local graph partitioning algorithms and locally testable codes. We give testers with two-sided error and one-sided error in the \adjacency list model that allows degree and neighbor queries to the oracle of the input graph. The testers take as input an n-vertex graph G, a volume bound k, an expansion bound \phi and a distance parameter \varepsilon>0. For the two-sided error tester, with probability at least 2/3, it accepts the graph if it is a (k,\phi)-expander and rejects the graph if it is \varepsilon-far from any (k^*,\phi^*)-expander, where k^*=\Theta(k\varepsilon) and \phi^*=\Theta(\frac{\phi^4}{\min\{\log(4m/k),\log n\}\cdot(\ln k)}). The query complexity and running time of the tester are \widetilde{O}(\sqrt{m}\phi^{-4}\varepsilon^{-2}), where m is the number of edges of the graph. For the one-sided error tester, it accepts every (k,\phi)-expander, and with probability at least 2/3, rejects every graph that is \varepsilon-far from (k^*,\phi^*)-expander, where k^*=O(k^{1-\xi}) and \phi^*=O(\xi\phi^2) for any 0<\xi<1. The query complexity and running time of this tester are \widetilde{O}(\sqrt{\frac{n}{\varepsilon^3}}+\frac{k}{\varepsilon \phi^4}). We also give a two-sided error tester in the \textit{rotation map} model that allows \textit{(neighbor, index)} queries and degree queries. This tester has asymptotically almost the same query complexity and running time as the two-sided error tester in the adjacency list model, but has a better performance: it can distinguish any $(k,\phi)$-expander from graphs that are $\varepsilon$-far from $(k^*,\phi^*)$-expanders, where $k^*=\Theta(k\varepsilon)$ and $\phi^*=\Theta(\frac{\phi^2}{\min\{\log(4m/k),\log n\}\cdot(\ln k)})$. In our analysis, we introduce a new graph product called \textit{non-uniform replacement product} that transforms a general graph into a bounded degree graph, and approximately preserves the expansion profile as well as the corresponding spectral property. Angsheng Li, Pan Peng 0001 |
STACS | 2 |
| 2015 | Testing Cluster Structure of GraphsabstractWe study the problem of recognizing the cluster structure of a graph in the framework of property testing in the bounded degree model. Given a parameter ε, a d-bounded degree graph is defined to be (k, φ)-clusterable, if it can be partitioned into no more than k parts, such that the (inner) conductance of the induced subgraph on each part is at least φ and the (outer) conductance of each part is at most cd,kε4φ2, where cd,k depends only on d,k. Our main result is a sublinear algorithm with the running time ~O(√n ⋅ poly(φ,k,1/ε)) that takes as input a graph with maximum degree bounded by d, parameters k, φ, ε, and with probability at least 2/3, accepts the graph if it is (k,φ)-clusterable and rejects the graph if it is ε-far from (k, φ*)-clusterable for φ* = c'd,kφ2 ε4}/log n, where c'd,k depends only on d,k. By the lower bound of Ω(√n) on the number of queries needed for testing graph expansion, which corresponds to k=1 in our problem, our algorithm is asymptotically optimal up to polylogarithmic factors. Artur Czumaj, Pan Peng 0001, Christian Sohler |
STOC | 2 |
| 2014 | Global core, and galaxy structure of networks
Yicheng Pan 0001, Pan Peng 0001, Jiankou Li, Angsheng Li |
Sci. China Inf. Sci. | 3 |
| 2013 | Detecting and Characterizing Small Dense Bipartite-Like Subgraphs by the Bipartiteness Ratio Measure
Angsheng Li, Pan Peng 0001 |
ISAAC | 2 |
| 2012 | A Local Algorithm for Finding Dense Bipartite-Like Subgraphs
Pan Peng 0001 |
COCOON | 1 |
| 2012 | The Small Community Phenomenon in Networks: Models, Algorithms and Applications
Pan Peng 0001 |
TAMC | 1 |
| 2012 | The small-community phenomenon in networksabstractWe investigate several geometric models of networks that simultaneously have some nice global properties, including the small-diameter property, the small-community phenomenon, which is defined to capture the common experience that (almost) everyone in society also belongs to some meaningful small communities, and the power law degree distribution, for which our result significantly strengthens those given in van den Esker (2008) and Jordan (2010). These results, together with our previous work in Li and Peng (2011), build a mathematical foundation for the study of both communities and the small-community phenomenon in various networks. In the proof of the power law degree distribution, we develop the method of alternating concentration analysis to build a concentration inequality by alternately and iteratively applying both the sub- and super-martingale inequalities, which seems to be a powerful technique with further potential applications. Angsheng Li, Pan Peng 0001 |
Math. Struct. Comput. Sci. | 2 |