VLDB 2026 Research / reviewers in the wild / expert
Tam Thuc Do
dblp:314/9673
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4632-0790ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Unrolling of Sparsity-Induced RDO for 3D Point Cloud Attribute CodingabstractWe study the problem of lossy attribute compression, given encoded 3D point cloud geometry available at the decoder, in a multi-resolution B-spline projection framework. A target continuous 3D attribute function is first projected onto a sequence of nested subspaces ${\mathcal {F}}^{(p)}_{l_{0}} \subseteq \cdots \subseteq {\mathcal {F}} ^{(p)}_{L}$ , where ${\mathcal {F}}^{(p)}_{l}$ is a family of functions spanned by a B-spline basis function of order $p$ at a chosen scale and its integer shifts. The projected low-pass coefficients $F_{l}^{*}$ are computed via variable-complexity unrolling of a rate-distortion (RD) optimization algorithm into a feed-forward network, where the rate term is the sparsity-promoting $\ell _{1}$ -norm. Thus, the projection operation is end-to-end differentiable. For a chosen coarse-to-fine predictor, the coefficients are then adjusted to account for the prediction from a lower-resolution to a higher-resolution, which is also optimized in a data-driven manner. Tam Thuc Do, Philip A. Chou, Gene Cheung |
IEEE Trans. Image Process. | 1 |
| 2024 | Volumetric 3d Point Cloud Attribute Compression: Learned Polynomial Bilateral Filter for PredictionabstractWe extend a previous study on 3D point cloud attribute compression scheme that uses a volumetric approach: given a target volumetric attribute function f : ℝ3↦ ℝ, we quantize and encode parameters θ that characterize f at the encoder, for reconstruction ${f_{\hat \theta }}({\mathbf{x}})$ at known 3D points x at the decoder. Specifically, parameters $\hat \theta $ are quantized coefficients of B-spline basis vectors Φl(for order p ≥ 2) that span the function space $\mathcal{F}_l^{(p)}$ at a particular resolution l, which are coded from coarse to fine resolutions for scalability. In this work, we focus on the prediction of finer-grained coefficients given coarser-grained ones by learning parameters of a polynomial bilateral filter (PBF) from data. PBF is a pseudo-linear filter that is signal-dependent with a graph spectral interpretation common in the graph signal processing (GSP) field. We demonstrate PBF’s predictive performance over a linear predictor inspired by MPEG standardization over a wide range of point cloud datasets. Tam Thuc Do, Philip A. Chou, Gene Cheung |
ICASSP | 1 |
| 2024 | Learned Nonlinear Predictor for Critically Sampled 3D Point Cloud Attribute CompressionabstractWe study 3D point cloud attribute compression via a volumetric approach: assuming point cloud geometry is known at both encoder and decoder, parameters $\theta$ of a continuous attribute function $f: \mathbb{R}^{3} \mapsto \mathbb{R}$ are quantized to $\hat{\theta}$ and encoded, so that discrete samples $f_{\hat{\theta}}\left(\mathbf{x}_{i}\right)$ can be recovered at known 3D points $\mathbf{x}_{i} \in \mathbb{R}^{3}$ at the decoder. Specifically, we consider a nested sequences of function subspaces ${\mathcal{F}}_{l_{0}}^{(p)} \subseteq \cdots \subseteq {\mathcal{F}}_{L}^{(p)}$, where ${\mathcal{F}}_{l}^{(p)}$ is a family of functions spanned by B-spline basis functions of order $p, f_{l}^{*}$ is the projection of f on ${\mathcal{F}}_{l}^{(p)}$ represented as low-pass coefficients $F_{l}^{*}$, and $g_{l}^{*}$ is the residual function in an orthogonal subspace ${\mathcal{G}}_{l}^{(p)}$ (where ${\mathcal{G}}_{l}^{(p)} \oplus {\mathcal{F}}_{l}^{(p)}=$ $\left.{\mathcal{F}}_{l+1}^{(p)}\right)$ represented as high-pass coefficients $G_{l}^{*}$. In this paper, to improve coding performance over [1], we study predicting $f_{l+1}^{*}$ at level $l+1$ given $f_{l}^{*}$ at level l and encoding of $G_{l}^{*}$ for the $p=1$ case (RAHT (1)). For the prediction, we formalize RAHT (1) linear prediction in MPEG-PCC in a theoretical framework, and propose a new nonlinear predictor using a polynomial of bilateral filter. We derive equations to efficiently compute the critically sampled high-pass coefficients $G_{l}^{*}$ amenable to encoding. We optimize parameters in our resulting feed-forward network on a large training set of point clouds by minimizing a rate-distortion Lagrangian. Experimental results show that our improved framework outperforms the MPEG G-PCC predictor by $11 \%-12 \%$ in bit rate. Tam Thuc Do, Philip A. Chou, Gene Cheung |
ICIP | 1 |
| 2024 | Constructing an Interpretable Deep Denoiser by Unrolling Graph Laplacian RegularizerabstractAn image denoiser can be used for a wide range of restoration problems via the Plug-and-Play (PnP) architecture. In this paper, we propose a general framework to build an interpretable graph-based deep denoiser (GDD) by unrolling a solution to a maximum a posteriori (MAP) problem equipped with a graph Laplacian regularizer (GLR) as signal prior. Leveraging a recent theorem showing that any (pseudo-)linear denoiser $\boldsymbol{\Psi}$, under mild conditions, can be mapped to a solution of a MAP denoising problem regularized using GLR, we first initialize a graph Laplacian matrix $\mathbf{L}$ via truncated Taylor Series Expansion (TSE) of $\Psi^{-1}$. Then, we compute the MAP linear system solution by unrolling iterations of the conjugate gradient (CG) algorithm into a sequence of neural layers as a feed-forward network—one that is amenable to parameter tuning. The resulting GDD network is “graph-interpretable”, low in parameter count, and easy to initialize thanks to $\mathbf{L}$ derived from a known well-performing denoiser $\boldsymbol{\Psi}$. Experimental results show that GDD achieves competitive image denoising performance compared to competitors, but employing far fewer parameters, and is more robust to covariate shift. Seyed Alireza Hosseini, Tam Thuc Do, Gene Cheung, Yuichi Tanaka 0001 |
ICIP | 2 |
| 2023 | Volumetric Attribute Compression for 3D Point Clouds Using Feedforward Network with Geometric AttentionabstractWe study 3D point cloud attribute compression using a volumetric approach: given a target volumetric attribute function f : ℝ3→ ℝ, we quantize and encode parameter vector θ that characterizes f at the encoder, for reconstruction ${f_{\hat \theta }}\left( {\mathbf{x}} \right)$ at known 3D points x’s at the decoder. Extending a previous work Region Adaptive Hierarchical Transform (RAHT) that employs piecewise constant functions to span a nested sequence of function spaces, we propose a feedforward linear network that implements higher-order B-spline bases spanning function spaces without eigen-decomposition. Feedforward network architecture means that the system is amenable to end-to-end neural learning. The key to our network is space-varying convolution, similar to a graph operator, whose weights are computed from the known 3D geometry for normalization. We show that the number of layers in the normalization at the encoder is equivalent to the number of terms in a matrix inverse Taylor series. Experimental results on real-world 3D point clouds show up to 2-3 dB gain over RAHT in energy compaction and 20-30% bitrate reduction. Tam Thuc Do, Philip A. Chou, Gene Cheung |
ICASSP | 1 |
| 2022 | Hybrid Model-Based / Data-Driven Graph Transform for Image CodingabstractTransform coding to sparsify signal representations remains crucial in an image compression pipeline. While the Karhunen-Loève transform (KLT) computed from an empirical covariance matrix ${\mathbf{\bar C}}$ is theoretically optimal for a stationary process, in practice, collecting sufficient statistics from a non-stationary image to reliably estimate ${\mathbf{\bar C}}$ can be difficult. In this paper, to encode an intra-prediction residual block, we pursue a hybrid model-based / data-driven approach: the first K eigenvectors of a transform matrix are derived from a statistical model, e.g., the asymmetric discrete sine transform (ADST), for stability, while the remaining N −K are computed from ${\mathbf{\bar C}}$ for data adaptivity. The transform computation is posed as a graph learning problem, where we seek a graph Laplacian matrix minimizing a graphical lasso objective inside a convex cone sharing the first K eigenvectors in a Hilbert space of real symmetric matrices. We efficiently solve the problem via augmented Lagrangian relaxation and proximal gradient (PG). Using open-source WebP as a baseline image codec, experimental results show that our hybrid graph transform achieved better coding performance than discrete cosine transform (DCT), ADST and KLT, and better stability than KLT. Saghar Bagheri, Tam Thuc Do, Gene Cheung, Antonio Ortega |
ICIP | 2 |