VLDB 2026 Research / reviewers in the wild / expert
Xin Lyu 0003
dblp:183/8301-3
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cell-Probe Lower Bounds via Semi-Random CSP Refutation: Simplified and the Odd-Locality CaseabstractA recent work (Korten, Pitassi, and Impagliazzo, FOCS 2025) established an insightful connection between static data structure lower bounds, range avoidance of NC0 circuits, and the refutation of pseudorandom CSP instances, leading to improvements to some longstanding lower bounds in the cell-probe/bit-probe models. Here, we improve these lower bounds in certain cases via a more streamlined reduction to XOR refutation, coupled with handling the odd-arity case. Our result can be viewed as a complete derandomization of the state-of-the-art semi-random \(k\)-XOR refutation analysis (Guruswami, Kothari and Manohar, STOC 2022, Hsieh, Kothari and Mohanty, SODA 2023), which complements the derandomization of the even-arity case obtained by Korten et al. Venkatesan Guruswami, Xin Lyu 0003, Weiqiang Yuan 0002 |
SODA | 2 |
| 2025 | Trade-offs in Data Memorization via Strong Data Processing InequalitiesabstractRecent research demonstrated that training large language models involves memorization of a significant fraction of training data. Such memorization can lead to privacy violations when training on sensitive user data and thus motivates the study of data memorization’s role in learning. In this work, we develop a general approach for proving lower bounds on excess data memorization, that relies on a new connection between strong data processing inequalities and data memorization. We then demonstrate that several simple and natural binary classification problems exhibit a trade-off between the number of samples available to a learning algorithm, and the amount of information about the training data that a learning algorithm needs to memorize to be accurate. In particular, $\Omega(d)$ bits of information about the training data need to be memorized when $O(1)$ $d$-dimensional examples are available, which then decays as the number of examples grows at a problem-specific rate. Further, our lower bounds are generally matched (up to logarithmic factors) by simple learning algorithms. We also extend our lower bounds to more general mixture-of-clusters models. Our definitions and results build on the work of Brown et al. (2021) and address several limitations of the lower bounds in their work. Vitaly Feldman, Guy Kornowski, Xin Lyu 0003 |
COLT | 3 |
| 2025 | Fingerprinting Codes Meet Geometry: Improved Lower Bounds for Private Query Release and Adaptive Data Analysis
Xin Lyu 0003, Kunal Talwar |
STOC | 1 |
| 2023 | Time-Space Tradeoffs for Element Distinctness and Set Intersection via PseudorandomnessabstractIn the ELEMENT DISTINCTNESS problem, one is given an array a1,…, an of integers from [poly(n)] and is tasked to decide if {ai} are mutually distinct. Beame, Clifford and Machmouchi (FOCS 2013) gave a low-space algorithm for this problem that runs in space S(n) and time T(n) where T(n) ≤ Õ(n3/2/S(n)1/2), assuming a random oracle (i.e., random access to polynomially many random bits). A recent breakthrough by Chen, Jin, Williams and Wu (SODA 2022) showed how to remove the random oracle assumption in the regime S(n) = polylog(n) and T(n) = Õ(n3/2). They designed the first truly polylog(n)-space, Õ(n3/2)-time algorithm by constructing a small family of hash functions H ⊆ {h|h : [poly(n)] → [n]} with a certain pseudorandom property. In this paper, we give a significantly simplified analysis of the pseudorandom hash family by Chen et al. Our analysis clearly identifies the key pseudorandom property required to fool the BCM algorithm, allowing us to explore the full potential of this construction. Based on our new analysis, we show the following. • As our main result, we give a time-space tradeoff for ELEMENT DISTINCTNESS without random oracle. Namely, for every S(n),T(n) such that T ≈ Õ(n3/2/S(n)1/2), our algorithm can solve the problem in space S(n) and time T(n). Our algorithm also works for a related problem SET INTERSECTION, for which this tradeoff is tight due to a matching lower bound by Dinur (Eurocrypt 2020). • As a direct application of our technique, we show a more general pseudorandom property of the hash family, which we call the “c-connecting” property. It might be of independent interest. • The construction by Chen et al. needs O(log3 n log log n) random bits to sample the pseudorandom hash function. We slightly improve the seed length to O (log3 n). * The full version of the paper can be accessed at https://arxiv.org/abs/2210.07534 Xin Lyu 0003, Weihao Zhu |
SODA | 1 |
| 2022 | Range Avoidance for Low-Depth Circuits and Connections to Pseudorandomness
Venkatesan Guruswami, Xin Lyu 0003, Xiuhan Wang |
APPROX/RANDOM | 2 |