Dor Tsur

dblp:260/0302 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-6561-4965ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Optimized Couplings for Watermarking Large Language Models
abstract
Large-language models (LLMs) are now able to produce text that is indistinguishable from human-generated content. This has fueled the development of watermarks that imprint a “signal” in LLM-generated text with minimal perturbation of an LLM's output. This paper provides an analysis of text watermarking in a one-shot setting. Through the lens of hypothesis testing with side information, we formulate and analyze the fundamental trade-off between watermark detection power and distortion in generated textual quality. We argue that a key component in watermark design is generating a coupling between the side information shared with the watermark detector and a random partition of the LLM vocabulary. Our analysis identifies the optimal coupling and randomization strategy under the worst-case LLM next-token distribution that satisfies a minentropy constraint. We provide a closed-form expression of the resulting detection rate under the proposed scheme and quantify the cost in a max-min sense. Finally, we numerically compare the proposed scheme with the theoretical optimum.
Carol Xuan Long, Dor Tsur, Claudio Mayrink Verdun, Hsiang Hsu, Haim H. Permuter, Flávio P. Calmon
ISIT2
2025 HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy Distributions
abstract
Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate by changing the next-token predictions output by an LLM. The updated (i.e., watermarked) predictions depend on random side information produced, for example, by hashing previously generated tokens. LLM watermarking is particularly challenging in low-entropy generation tasks -- such as coding -- where next-token predictions are near-deterministic. In this paper, we propose an optimization framework for watermark design. Our goal is to understand how to most effectively use random side information in order to maximize the likelihood of watermark detection and minimize the distortion of generated text. Our analysis informs the design of two new watermarks: HeavyWater and SimplexWater. Both watermarks are tunable, gracefully trading-off between detection accuracy and text distortion. They can also be applied to any LLM and are agnostic to side information generation. We examine the performance of HeavyWater and SimplexWater through several benchmarks, demonstrating that they can achieve high watermark detection accuracy with minimal compromise of text generation quality, particularly in the low-entropy regime. Our theoretical analysis also reveals surprising new connections between LLM watermarking and coding theory.
Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana, Hsiang Hsu, Chun-Fu Chen 0001, Haim H. Permuter, Flávio P. Calmon
NeurIPS1
2024 Neural Estimation of Multi-User Capacity Regions Over Discrete Channels
abstract
This paper presents a data-driven methodology for estimating capacity regions in multi-user communication scenarios, focusing on channels with discrete alphabets, both with and without feedback. Prior research has successfully utilized neural networks for estimating capacity regions in continuous domains. However, the shift to discrete alphabets introduces a significant challenge due to the lack of end-to-end differentiability of the joint model. To tackle this issue, we first formulate the optimization problem of the causally conditioned directed information rate as a decentralized Markov decision process (MDP). Building on this formulation, we introduce a tractable optimization procedure specifically designed to estimate rate pairs that lie on the boundary of the capacity region. In addressing the inherent complexity of the MDP state space, we employ a reinforcement learning (RL) algorithm to learn optimal policies. We demonstrate the performance of our methodology by applying it to various communication scenarios, including the two-way channel and the multiple access channel (MAC). The results showcase the adaptability and performance of the proposed RL-based framework in estimating capacity regions without explicit knowledge of the underlying channel model, whether there is feedback or not.
Bashar Huleihel, Dor Tsur, Ziv Aharoni, Oron Sabag, Haim H. Permuter
ISIT2
2024 InfoMat: A Tool for the Analysis and Visualization Sequential Information Transfer
abstract
Despite the popularity of information measures in analysis of probabilistic systems, proper tools for their visualization are not common. This work develops a simple matrix representation of information transfer in sequential systems, termed information matrix (InfoMat). The simplicity of the InfoMat provides a new visual perspective on existing decomposition formulas of mutual information, and enables us to prove new relations between sequential information theoretic measures. We study various estimation schemes of the InfoMat, facilitating the visualization of information transfer in sequential datasets. By drawing a connection between visual pattern in the InfoMat and various dependence structures, we observe how information transfer evolves in the dataset. We then leverage this tool to visualize the effect of capacity-achieving coding schemes on the underlying exchange of information. We believe the InfoMat is applicable to any time-series task for a better understanding of the data at hand.
Dor Tsur, Haim H. Permuter
ISIT1
2024 Several Interpretations of Max-Sliced Mutual Information
abstract
Max-sliced mutual information (mSMI) was recently proposed as a data-efficient measure of dependence. This measure extends popular correlation-based methods and proves useful in various machine learning tasks. In this paper, we extend the notion of mSMI to discrete variables and investigate its role in popular problems of information theory and statistics. We use mSMI to propose a soft version of the Gacs-Korner common information, which, due to the mSMI structure, naturally extends to continuous domains and multivariate settings. We then characterize the optimal growth rate in a horse race with constrained side information. Additionally, we examine the error of independence testing under communication constraints. Finally, we study mSMI in communications. We characterize the capacity of discrete memoryless channels with constrained encoders and decoders, and propose an mSMI-based scheme to decode information obtained through remote sensing. These connections motivate the use of max-slicing in information theory, and benefit from its merits.
Dor Tsur, Haim H. Permuter, Ziv Goldfeld
ISIT1
2024 Data-Driven Optimization of Directed Information Over Discrete Alphabets
abstract
Directed information (DI) is a fundamental measure for the study and analysis of sequential stochastic models. In particular, when optimized over input distributions it characterizes the capacity of general communication channels. However, analytic computation of DI is typically intractable and existing optimization techniques over discrete input alphabets require knowledge of the channel model, which renders them inapplicable when only samples are available. To overcome these limitations, we propose a novel optimization framework for estimated DI over discrete spaces. We formulate DI optimization as a Markov decision process and leverage reinforcement learning techniques to optimize a deep generative model of the input process probability mass function (PMF). Combining this optimizer with the recently developed DI neural estimator, we obtain an alternating optimization algorithm which is applied to estimating the (feedforward and feedback) capacity of various discrete channels with memory. Furthermore, we demonstrate how to use the optimized PMF model to (i) obtain theoretical bounds on the feedback capacity of unifilar finite-state channels; and (ii) perform probabilistic shaping of constellations in the peak power-constrained additive white Gaussian noise channel.
Dor Tsur, Ziv Aharoni, Ziv Goldfeld, Haim H. Permuter
IEEE Trans. Inf. Theory1
2023 Neural Estimation of Multi-User Capacity Regions
abstract
In this paper, we introduce a data-driven methodology for estimating capacity regions of continuous channels in multi-user communication systems. Computing capacity regions is a long standing open problem, even in simple communication scenarios. Nevertheless, it is often possible to represent their capacity regions as the limit of an optimization problem (a multi-letter expression). In many cases, these multi-letter expressions can be expressed in terms of directed information (DI) rates. Accordingly, our approach utilizes neural networks to estimate capacity regions, leveraging the recent introduction of the directed information neural estimator (DINE). The main idea of our methodology involves training DINE-based models using samples of channel inputs and outputs, and using these models to estimate the DI rate terms that are intrinsic to the studied capacity region. To estimate the capacity region rates, we optimize the DI rates over the involved input distributions which are parameterized by a neural distribution transformer (NDT), and execute an alternating maximization procedure between the NDT models and DINE-based models until convergence is achieved. The methodology is suitable for the case where the channel is treated as a "black-box" and the designer can only gather observations of its inputs and outputs, lacking any knowledge of the explicit channel model. The performance of our proposed algorithm is shown via several well-known settings, including the Gaussian two-way channel and the two-user Gaussian multiple-access channel with and without feedback.
Bashar Huleihel, Dor Tsur, Ziv Aharoni, Oron Sabag, Haim H. Permuter
ISIT2
2023 Rate Distortion via Constrained Estimated Mutual Information Minimization
abstract
This paper proposes a novel methodology for the estimation of the rate distortion function (RDF) in both continuous and discrete reconstruction spaces. The approach is input-space agnostic and does not require prior knowledge of the source distribution, nor the distortion function, i.e., it treats them as "black box" models. Thus, our method is a general solution to the RDF estimation problem. The approach leverages neural estimation and optimization of information measures to optimize a generative model of the input distribution. In continuous spaces we learn a sample generating model and a PMF model is proposed for discrete spaces. Formal guarantees of the proposed method are explored and implementation details are discussed. We demonstrate the performance on both high dimensional and large alphabet synthetic data. This work has the potential to contribute to the fields of data compression and machine learning through the development of provably consistent and competitive compressors optimized for the fundamental limit of the RDF.
Dor Tsur, Bashar Huleihel, Haim H. Permuter
ISIT1
2023 Max-Sliced Mutual Information
abstract
Quantifying dependence between high-dimensional random variables is central to statistical learning and inference. Two classical methods are canonical correlation analysis (CCA), which identifies maximally correlated projected versions of the original variables, and Shannon's mutual information, which is a universal dependence measure that also captures high-order dependencies. However, CCA only accounts for linear dependence, which may be insufficient for certain applications, while mutual information is often infeasible to compute/estimate in high dimensions. This work proposes a middle ground in the form of a scalable information-theoretic generalization of CCA, termed max-sliced mutual information (mSMI). mSMI equals the maximal mutual information between low-dimensional projections of the high-dimensional variables, which reduces back to CCA in the Gaussian case. It enjoys the best of both worlds: capturing intricate dependencies in the data while being amenable to fast computation and scalable estimation from samples. We show that mSMI retains favorable structural properties of Shannon's mutual information, like variational forms and identification of independence. We then study statistical estimation of mSMI, propose an efficiently computable neural estimator, and couple it with formal non-asymptotic error bounds. We present experiments that demonstrate the utility of mSMI for several tasks, encompassing independence testing, multi-view representation learning, algorithmic fairness, and generative modeling. We observe that mSMI consistently outperforms competing methods with little-to-no computational overhead.
Dor Tsur, Ziv Goldfeld, Kristjan Greenewald
NeurIPS1
2023 Neural Estimation and Optimization of Directed Information Over Continuous Spaces
abstract
This work develops a new method for estimating and optimizing the directed information rate between two jointly stationary and ergodic stochastic processes. Building upon recent advances in machine learning, we propose a recurrent neural network (RNN)-based estimator which is optimized via gradient ascent over the RNN parameters. The estimator does not require prior knowledge of the underlying joint/marginal distributions and can be easily optimized over continuous input processes realized by a deep generative model. We prove consistency of the proposed estimation and optimization methods and combine them to obtain end-to-end performance guarantees. Applications for channel capacity estimation of continuous channels with memory are explored, and empirical results demonstrating the scalability and accuracy of our method are provided. When the channel is memoryless, we investigate the mapping learned by the optimized input generator.
Dor Tsur, Ziv Aharoni, Ziv Goldfeld, Haim H. Permuter
IEEE Trans. Inf. Theory1
2022 Density Estimation of Processes with Memory via Donsker Vardhan
abstract
Density estimation plays an important role in modeling random variables (RVs) with continuous alphabets. This work provides an algorithm that estimates the probability density function (PDF) of stationary and ergodic random processes using recurrent neural networks (RNNs). The main idea is to decompose the target PDF into a known auxiliary PDF and a likelihood ratio between the target and auxiliary PDFs. The algorithm focuses on estimating the likelihood ratio using the Donsker Vardhan (DV) variational formula of Kullback Leibler (KL) divergence. Together, the maximizer of the DV formula and the auxiliary PDF are used to construct the estimator of the target PDF in the form of a Gibbs density. The obtained estimator converges to the target PDF in total variation (TV) and in distribution. Also, we show that proposed estimator minimizes the cross entropy (CE) between the target and auxiliary distribution, and that with a proper choice of the auxiliary distribution, it defines a tight upper bound on the entropy rate. We demonstrate this approach by estimating the density of a Gaussian hidden Markov model.
Ziv Aharoni, Dor Tsur, Haim H. Permuter
ISIT2
2022 Optimizing Estimated Directed Information over Discrete Alphabets
abstract
Directed information (DI) is a fundamental measure for the study and analysis of sequential stochastic models. In particular, when optimized over the input distribution, it characterizes the capacity of general communication channels. However, existing optimization methods for discrete input alphabets assume full knowledge of the channel model, and are therefore not applicable when only samples are available. We derive a new method that overcomes this limitation and enables optimizing DI over unknown channels. To that end, we formulate the problem as a Markov decision process and leverage reinforcement learning techniques to optimize a deep generative model of the channel input probability mass function (PMF). Combining our optimizer with the DI neural estimator, we obtain an end-to-end estimation-optimization scheme which is applied for estimating the capacity of various discrete channels with memory. We provide empirical results that demonstrate the utility of the proposed framework and further show how to use the optimized PMF generator to obtain theoretical bounds on the feedback capacity for unifilar finite state channels.
Dor Tsur, Ziv Aharoni, Ziv Goldfeld, Haim H. Permuter
ISIT1
2020 Capacity of Continuous Channels with Memory via Directed Information Neural Estimator
abstract
Calculating the capacity (with or without feedback) of channels with memory and continuous alphabets is a challenging task. It requires optimizing the directed information (DI) rate over all channel input distributions. The objective is a multi-letter expression, whose analytic solution is only known for a few specific cases. When no analytic solution is present or the channel model is unknown, there is no unified framework for calculating or even approximating capacity. This work proposes a novel capacity estimation algorithm that treats the channel as a `black-box', both when feedback is or is not present. The algorithm has two main ingredients: (i) a neural distribution transformer (NDT) model that shapes a noise variable into the channel input distribution, which we are able to sample, and (ii) the DI neural estimator (DINE) that estimates the communication rate of the current NDT model. These models are trained by an alternating maximization procedure to both estimate the channel capacity and obtain an NDT for the optimal input distribution. The method is demonstrated on the moving average additive Gaussian noise channel, where it is shown that both the capacity and feedback capacity are estimated without knowledge of the channel transition kernel. The proposed estimation framework opens the door to a myriad of capacity approximation results for continuous alphabet channels that were inaccessible until now.
Ziv Aharoni, Dor Tsur, Ziv Goldfeld, Haim H. Permuter
ISIT2