Buu Phan

dblp:228/4651 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 62% Video understanding and tracking · 21% 3D vision · 10%
Theoretical computer science
1 paper
Coding theory · 56% Algorithmic game theory and mechanism design · 44%
Computer graphics and multimedia
1 paper
Image and video coding · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
byte-level language model
0.912025
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles · ICLR 2025
Natural language and speech › Language models and text generation › large language model
large language model ensemble
0.912025
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles · ICLR 2025
Natural language and speech › Language models and text generation
tokenization
0.912025
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles · ICLR 2025
Coding theory › channel coding
channel simulation
0.912025
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling · NeurIPS 2025
Algorithmic game theory and mechanism design › matching › algorithmic matching
distributed matching
0.912025
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling · NeurIPS 2025
Image and video coding › video compression
learned video compression
0.712023
On the choice of Perception Loss Function for Learned Video Compression · NeurIPS 2023
Image and video coding › rate-distortion analysis
rate-distortion-perception tradeoff
0.712023
On the choice of Perception Loss Function for Learned Video Compression · NeurIPS 2023
Image and video coding
video compression
0.712023
On the choice of Perception Loss Function for Learned Video Compression · NeurIPS 2023
Security and privacy of machine learning
adversarial attack
0.512021
Adversarial Imaging Pipelines · CVPR 2021
Computer vision › Video understanding and tracking
object tracking
0.412020
Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar · CVPR 2020
Computer vision › Video understanding and tracking › object tracking › robust tracking
occluded object tracking
0.412020
Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar · CVPR 2020
Coding theory › source coding
distributed compression
0.312025
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling · NeurIPS 2025
Robotics › Autonomous driving › perception
non-line-of-sight detection
0.112020
Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar · CVPR 2020
Robotics › Autonomous driving
perception
0.112020
Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar · CVPR 2020

Methods — techniques the papers use, named apart from their topics

rejection sampling · 0.9ensemble rejection sampling · 0.9information-theoretic analysis · 0.7deep learning · 0.7multi-task optimization · 0.5differentiable ISP approximation · 0.5temporal fusion · 0.4joint detection and tracking network · 0.4doppler radar · 0.4
YearPublicationVenuePosition
2026 One-Shot Broadcast Joint Source-Channel Coding with Codebook Diversity
abstract
We study a one-shot joint source-channel coding setting where the source is encoded once and broadcast to $K$ decoders through independent channels. Success is predicated on at least one decoder recovering the source within a maximum distortion constraint. We find that in the one-shot regime, utilizing disjoint codebooks at each decoder yields a codebook diversity gain, distinct from the channel diversity gain that may be expected when several decoders observe independent realizations of the channel's output but share the same codebook. Coding schemes are introduced that leverage this phenomenon, where first- and second-order achievability bounds are derived via an adaptation of the Poisson matching lemma which allows for multiple decoders using disjoint codebooks. We further propose a hybrid coding scheme that partitions decoders into groups to optimally balance codebook and channel diversity. Numerical results on the binary symmetric channel demonstrate that the hybrid approach outperforms strategies where the decoders' codebooks are either fully shared or disjoint.
Joseph Rowan, Buu Phan, Ashish Khisti
ISIT2
2025 Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
abstract
Tokenization is associated with many poorly understood shortcomings in language models (LMs), yet remains an important component for long sequence scaling purposes. This work studies how tokenization impacts model performance by analyzing and comparing the stochastic behavior of tokenized models with their byte-level, or token-free, counterparts. We discover that, even when the two models are statistically equivalent, their predictive distributions over the next byte can be substantially different, a phenomenon we term as ``tokenization bias''. To fully characterize this phenomenon, we introduce the Byte-Token Representation Lemma, a framework that establishes a mapping between the learned token distribution and its equivalent byte-level distribution. From this result, we develop a next-byte sampling algorithm that eliminates tokenization bias without requiring further training or optimization. In other words, this enables zero-shot conversion of tokenized LMs into statistically equivalent token-free ones. We demonstrate its broad applicability with two use cases: fill-in-the-middle (FIM) tasks and model ensembles. In FIM tasks where input prompts may terminate mid-token, leading to out-of-distribution tokenization, our method mitigates performance degradation and achieves 18\% improvement in FIM coding benchmarks, while consistently outperforming the standard token healing fix. For model ensembles where each model employs a distinct vocabulary, our approach enables seamless integration, resulting in improved performance up to 3.7\% over individual models across various standard baselines in reasoning, knowledge, and coding. Code is available at:https: //github.com/facebookresearch/Exact-Byte-Level-Probabilities-from-Tokenized-LMs.
Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew J. Muckley, Karen Ullrich
ICLR1
2025 On List Decoding With Importance Sampling
abstract
The Importance Matching Lemma (IML) [1] is a recently introduced importance sampling-based technique for distributed compression, offering a finite-proposals alternative to the Poisson Matching Lemma (PML). Unlike PML, which relies on an infinite proposals formulation and can encounter termination issues, IML operates directly on finite samples, ensuring practical feasibility in real-world scenarios. However, its current formulation is restricted to exact matching and does not extend to list decoding, a capability available under PML. In this work, we extend IML to support list decoding by leveraging results from the order statistics of exponential random variables with different rates. We provide a detailed theoretical analysis of the proposed extension and derive list-decoding guarantees in the context of channel coding. Our results demonstrate a novel application of importance sampling and list-decoding techniques, expanding the scope of IML to a wider range of coding scenarios.
Buu Phan, Ashish Khisti
ISIT1
2025 Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
abstract
We study channel simulation and distributed matching, two fundamental problems with several applications to machine learning, using a recently introduced generalization of the standard rejection sampling (RS) algorithm known as Ensemble Rejection Sampling (ERS). For channel simulation, we propose a new coding scheme based on ERS that achieves a near-optimal coding rate. In this process, we demonstrate that standard RS can also achieve a near-optimal coding rate and generalize the result of Braverman and Garg (2014) to the continuous alphabet setting. Next, as our main contribution, we present a distributed matching lemma for ERS, which serves as the rejection sampling counterpart to the Poisson Matching Lemma (PML) introduced by Li and Anantharam (2021). Our result also generalizes a recent work on importance matching lemma (Phan et al, 2024) and, to our knowledge, is the first result on distributed matching in the family of rejection sampling schemes where the matching probability is close to PML. We demonstrate the practical significance of our approach over prior works by applying it to distributed compression. The effectiveness of our proposed scheme is validated through experiments involving synthetic Gaussian sources and distributed image compression using the MNIST dataset.
Buu Phan, Ashish Khisti
NeurIPS1
2025 List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
abstract
We study a relaxation of the problem of coupling probability distributions — a list of samples is generated from one distribution and an *accept* is declared if any one of these samples is identical to the sample generated from the other distribution. We propose a novel method for generating samples, which extends the Gumbel-max sampling suggested in Daliri et al. (2025) for coupling probability distributions. We also establish a corresponding lower bound on the acceptance probability, which we call the *list matching lemma*. We next discuss two applications of our setup. First, we develop a new mechanism for multi-draft speculative sampling that is simple to implement and achieves performance competitive with baselines such as SpecTr and SpecInfer across a range of language tasks. Our method also guarantees a certain degree of *drafter invariance* with respect to the output tokens which is not supported by existing schemes. We also provide a theoretical lower bound on the token level acceptance probability. As our second application, we consider distributed lossy compression with side information in a setting where a source sample is compressed and available to multiple decoders, each with independent side information. We propose a compression technique that is based on our generalization of Gumbel-max sampling and show that it provides significant gains in experiments involving synthetic Gaussian sources and the MNIST image dataset.
Joseph Rowan, Buu Phan, Ashish Khisti
NeurIPS2
2024 Importance Matching Lemma for Lossy Compression with Side Information
abstract
We propose two extensions to existing importance sampling based methods for lossy compression. First, we introduce an importance sampling based compression scheme that is a variant of ordered random coding (Theis and Ahmed, 2022) and is amenable to direct evaluation of the achievable compression rate for a finite number of samples. Our second and major contribution is the \emph{importance matching lemma}, which is a finite proposal counterpart of the recently introduced {Poisson matching lemma} (Li and Anantharam, 2021). By integrating with deep learning, we provide a new coding scheme for distributed lossy compression with side information at the decoder. We demonstrate the effectiveness of the proposed scheme through experiments involving synthetic Gaussian sources, distributed image compression with MNIST and vertical federated learning with CIFAR-10.
Buu Phan, Ashish Khisti, Christos Louizos
AISTATS1
2023 On the choice of Perception Loss Function for Learned Video Compression
abstract
We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD, considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion compared to PLF-FMD and further makes it more difficult to recover from errors made in the earlier output frames. Although the choice of PLF decisively affects reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be {\em near universal} and can generate close to optimal reconstructions for either choice of PLF at the decoder. We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST and KTH datasets.
Sadaf Salehkalaibar, Buu Phan, Jun Chen 0005, Wei Yu 0001, Ashish Khisti
NeurIPS2
2021 Adversarial Imaging Pipelines
abstract
Adversarial attacks play a critical role in understanding deep neural network predictions and improving their robustness. Existing attack methods aim to deceive convolutional neural network (CNN)-based classifiers by manipulating RGB images that are fed directly to the classifiers. However, these approaches typically neglect the influence of the camera optics and image processing pipeline (ISP) that produce the network inputs. ISPs transform RAW measurements to RGB images and traditionally are assumed to preserve adversarial patterns. In fact, these low-level pipelines can destroy, introduce or amplify adversarial patterns that can deceive a downstream detector. As a result, optimized patterns can become adversarial for the classifier after being transformed by a certain camera ISP or optical lens system but not for others. In this work, we examine and develop such an attack that deceives a specific camera ISP while leaving others intact, using the same downstream classifier. We frame this camera-specific attack as a multi-task optimization problem, relying on a differentiable approximation for the ISP itself. We validate the proposed method using recent state-of-the-art automotive hardware ISPs, achieving 92% fooling rate when attacking a specific ISP. We demonstrate physical optics attacks with 90% fooling rate for a specific camera lens.
Buu Phan, Fahim Mannan, Felix Heide
CVPR1
2020 Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar
abstract
Conventional sensor systems record information about directly visible objects, whereas occluded scene components are considered lost in the measurement process. Non-line-of-sight (NLOS) methods try to recover such hidden objects from their indirect reflections - faint signal components, traditionally treated as measurement noise. Existing NLOS approaches struggle to record these low-signal components outside the lab, and do not scale to large-scale outdoor scenes and high-speed motion, typical in automotive scenarios. In particular, optical NLOS capture is fundamentally limited by the quartic intensity falloff of diffuse indirect reflections. In this work, we depart from visible-wavelength approaches and demonstrate detection, classification, and tracking of hidden objects in large-scale dynamic environments using Doppler radars that can be manufactured at low-cost in series production. To untangle noisy indirect and direct reflections, we learn from temporal sequences of Doppler velocity and position measurements, which we fuse in a joint NLOS detection and tracking network over time. We validate the approach on in-the-wild automotive scenes, including sequences of parked cars or house facades as relay surfaces, and demonstrate low-cost, real-time NLOS in dynamic automotive environments.
Nicolas Scheiner, Florian Kraus, Fangyin Wei, Buu Phan, Fahim Mannan, Nils Appenrodt, Werner Ritter, Jürgen Dickmann, Klaus Dietmayer, Bernhard Sick, Felix Heide
CVPR4
2018 An Automated Vehicle Safety Concept Based on Runtime Restriction of the Operational Design Domain
abstract
Automated vehicles need to operate safely in a wide range of environments and hazards. The complex systems that make up an automated vehicle must also ensure safety in the event of system failures. This paper proposes an approach and architectural design for achieving maximum functionality in the case of system failures. The Operational Design Domain (ODD) defines the domain over which the automated vehicle can operate safely. We propose modifying a runtime representation of the ODD based on current system capabilities. This enables the system to react with context-appropriate responses depending on the remaining degraded functionality. In addition to proposing an architectural design, we have implemented the approach to prove its viability. The proof of concept has shown promising directions for future work and moved our automated vehicle research platform closer to achieving level 4 automation.
Ian Colwell, Buu Phan, Shahwar Saleem, Rick Salay, Krzysztof Czarnecki 0001
Intelligent Vehicles Symposium2