EDBT 2026 Demo / reviewers in the wild / expert
Jiahao Pang
dblp:133/4074
· DBLP profile ↗
41ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-8857-1152ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSEditor: Controllable mask-to-scene generation with diffusion model
Jiahao Pang, Zhiqiang Pu, Yanyan Liang 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Geometry Regularized Point Cloud AutoencoderabstractPoint cloud is a prevalent format in representing 3D geometry. Regardless of the recent advances, unsupervised learning for 3D point clouds remains arduous for various tasks due to its unorganized and sparsely distributed nature. To address this challenge, we propose a geometry regularized point cloud autoencoder, aiming to preserve local geometry structure. In particular, based on the Mahalanobis distance, we propose a point cloud geometry metric counting the local statistics. It endeavors to maximize the posterior probability of the reconstruction conditioned on the input point cloud. Our eigenspace analysis reveals the adaptivity of the developed metric—it behaves differently given different local structures. Moreover, a coarse-to-fine training strategy by varying the metric granularity is applied, leading to our proposed geometry regularized point cloud autoencoder. By applying our proposal to several off-the-shelf point cloud autoencoders, we show an improved point cloud reconstruction quality. In addition, the superior representability of our learned features is also demonstrated via an object classification task. Ritwik Sadhu, Jiahao Pang, Dong Tian |
ICIP | 2 |
| 2025 | UH-PCC: Unified Octree and Feature Coding for Hierarchical Point Cloud Geometry Compression
Muhammad Asad Lodhi, Jiahao Pang, Junghyun Ahn, Yuning Huang, Dong Tian |
PCS | 2 |
| 2024 | PIVOT-Net: Heterogeneous Point-Voxel-Tree-based Framework for Point Cloud CompressionabstractThe universality of the point cloud format enables many 3D applications, making the compression of point clouds a critical phase in practice. Sampled as discrete 3D points, a point cloud approximates 2D surface(s) embedded in 3D with a finite bit-depth. However, the point distribution of a practical point cloud changes drastically as its bit-depth increases, requiring different methodologies for effective consumption/analysis. In this regard, a heterogeneous point cloud compression (PCC) framework is proposed. We unify typical point cloud representations-pointbased, voxel-based, and tree-based representations-and their associated backbones under a learning-based framework to compress an input point cloud at different bit-depth levels. Having recognized the importance of voxel-domain processing, we augment the framework with a proposed context-aware upsampling for decoding and an enhanced voxel transformer for feature aggregation. Extensive experimentation demonstrates the state-of-the-art performance of our proposal on a wide range of point clouds. Jiahao Pang, Kevin Bui, Dong Tian |
3DV | 1 |
| 2024 | WrappingNet: Mesh Autoencoder Via Deep Sphere DeformationabstractThere have been recent efforts to learn more meaningful representations via fixed length codewords from mesh data, since a mesh serves as a complete model of underlying 3D shape compared to a point cloud. However, the mesh connectivity presents new difficulties when constructing a deep learning pipeline for meshes. Previous mesh unsupervised learning approaches typically assume category-specific templates, e.g., human face/body templates. It restricts the learned latent codes to only be meaningful for objects in a specific category, so the learned latent spaces are unable to be used across different types of objects. In this work, we present WrappingNet, the first mesh autoencoder enabling general mesh unsupervised learning over heterogeneous objects. It introduces a novel base graph in the bottleneck dedicated to representing mesh connectivity, which is shown to facilitate learning a shared latent space representing object shape. The superiority of WrappingNet mesh learning is further demonstrated via improved reconstruction quality and competitive classification compared to point cloud learning, as well as latent interpolation between meshes of different categories. The code is available at https://github.com/InterDigitalInc/WrappingNet. Eric Lei, Muhammad Asad Lodhi, Jiahao Pang, Junghyun Ahn, Dong Tian |
ICIP | 3 |
| 2024 | Towards Reproducible Learning-Based CompressionabstractA deep learning system typically suffers from a lack of reproducibility that is partially rooted in hardware or software implementation details. The irreproducibility leads to skepticism in deep learning technologies and it can hinder them from being deployed in many applications. In this work, the irreproducibility issue is analyzed where deep learning is employed in compression systems while the encoding and decoding may be run on devices from different manufacturers. The decoding process can even crash due to a single bit difference, e.g., in a learning-based entropy coder. For a given deep learning-based module with limited resources for protection, we first suggest that reproducibility can only be assured when the mismatches are bounded. Then a safeguarding mechanism is proposed to tackle the challenges. The proposed method may be applied for different levels of protection either at the reconstruction level or at a selected decoding level. Furthermore, the overhead introduced for the protection can be scaled down accordingly when the error bound is being suppressed. Experiments demonstrate the effectiveness of the proposed approach for learning-based compression systems, e.g., in image compression and point cloud compression. Jiahao Pang, Muhammad Asad Lodhi, Junghyun Ahn, Yuning Huang, Dong Tian |
MMSP | 1 |
| 2023 | Sparse Convolution Based Octree Feature Propagation for Lidar Point Cloud CompressionabstractWith the advent of new 3D scanning technologies, point clouds have become a crucial way to depict real and virtual objects/scenes. Point clouds represent the continuous sur-faces of underlying object/scene through a collection (usually millions) of discrete, irregular, and often sparsely distributed 3D samples on the surface of the objects, e.g. LiDAR scans. This nature of point cloud data presents a considerable challenge to not only store but also understand and extract the topology of object(s) from the point cloud data. In this regard, our work presents a point cloud compression procedure that leverages sparse 3D convolutions to extract features at various octree scales for lossless compression of octree representation of point clouds. For hierarchical flow of information between octree levels, our proposed method named SparseContextNet (SCN) also propagates features from a lower resolution scale to higher resolution scale via 3D upsampling convolutions. Our experiments with LiDAR datasets reveal competitive performance of our proposal compared to the state-of-the-art. Muhammad Asad Lodhi, Jiahao Pang, Dong Tian |
ICASSP | 2 |
| 2023 | DDA-Net: Deep Distribution-Aware Network for Point Cloud CompressionabstractDeep neural networks have been recently applied to point cloud compression (PCC). The features extracted via deep neural networks are essential for compression performance. Different from high level tasks such as point cloud classification or segmentation which homogenizes descriptors within same classes, PCC requires low level features discriminative for point-level 3D reconstructions. With this motivation, we first adopt Gaussian distribution to model the shape of feature elements. Then, we propose a deep distribution-aware network (DDA-Net) which manipulates distributions of feature elements on-the-fly to favor the point cloud reconstruction with high fidelity. Moreover, a residual network is integrated to enhance the modification of the Gaussian models. The proposed DDA-Net is incorporated into an end-to-end PCC system. Experimental results show that our DDA-Net significantly improves the compression performance across a wide range of point clouds. Junghyun Ahn, Jiahao Pang, Muhammad Asad Lodhi, Dong Tian |
ISCAS | 2 |
| 2022 | Graph-Based Depth Denoising & Dequantization for Point Cloud EnhancementabstractA 3D point cloud is typically constructed from depth measurements acquired by sensors at one or more viewpoints. The measurements suffer from both quantization and noise corruption. To improve quality, previous works denoise a point cloud a posteriori after projecting the imperfect depth data onto 3D space. Instead, we enhance depth measurements directly on the sensed images a priori, before synthesizing a 3D point cloud. By enhancing near the physical sensing process, we tailor our optimization to our depth formation model before subsequent processing steps that obscure measurement errors. Specifically, we model depth formation as a combined process of signal-dependent noise addition and non-uniform log-based quantization. The designed model is validated (with parameters fitted) using collected empirical data from a representative depth sensor. To enhance each pixel row in a depth image, we first encode intra-view similarities between available row pixels as edge weights via feature graph learning. We next establish inter-view similarities with another rectified depth image via viewpoint mapping and sparse linear interpolation. This leads to a maximum a posteriori (MAP) graph filtering objective that is convex and differentiable. We minimize the objective efficiently using accelerated gradient descent (AGD), where the optimal step size is approximated via Gershgorin circle theorem (GCT). Experiments show that our method significantly outperformed recent point cloud denoising schemes and state-of-the-art image denoising schemes in two established point cloud quality metrics. Xue Zhang 0008, Gene Cheung, Jiahao Pang, Yash Sanghvi, Abhiram Gnanasambandam, Stanley H. Chan |
IEEE Trans. Image Process. | 3 |
| 2022 | Graph Signal Processing for Geometric Data and Beyond: Theory and ApplicationsabstractGeometric data acquired from real-world scenes,e.g., 2D depth images, 3D point clouds, and 4D dynamic point clouds, have found a wide range of applications including immersive telepresence, autonomous driving, surveillance,etc. Due to irregular sampling patterns of most geometric data, traditional image/video processing methodologies are limited, while Graph Signal Processing (GSP)—a fast-developing field in the signal processing community—enables processing signals that reside on irregular domains and plays a critical role in numerous applications of geometric data from low-level processing to high-level analysis. To further advance the research in this field, we provide the first timely and comprehensive overview of GSP methodologies for geometric data in a unified manner by bridging the connections between geometric data and graphs, among the various geometric data modalities, and with spectral/nodal graph filtering techniques. We also discuss the recently developed Graph Neural Networks (GNNs) and interpret the operation of these networks from the perspective of GSP. We conclude with a brief discussion of open problems and challenges. Wei Hu 0003, Jiahao Pang, Xianming Liu 0005, Dong Tian, Chia-Wen Lin, Anthony Vetro |
IEEE Trans. Multim. | 2 |
| 2021 | TearingNet: Point Cloud Autoencoder To Learn Topology-Friendly RepresentationsabstractTopology matters. Despite the recent success of point cloud processing with geometric deep learning, it remains arduous to capture the complex topologies of point cloud data with a learning model. Given a point cloud dataset containing objects with various genera, or scenes with multiple objects, we propose an autoencoder, TearingNet, which tackles the challenging task of representing the point clouds using a fixed-length descriptor. Unlike existing works directly deforming predefined primitives of genus zero (e.g., a 2D square patch) to an object-level point cloud, our TearingNet is characterized by a proposed Tearing network module and a Folding network module interacting with each other iteratively. Particularly, the Tearing network module learns the point cloud topology explicitly. By breaking the edges of a primitive graph, it tears the graph into patches or with holes to emulate the topology of a target point cloud, leading to faithful reconstructions. Experimentation shows the superiority of our proposal in terms of reconstructing point clouds as well as generating more topology-friendly representations than benchmarks. Jiahao Pang, Duanshun Li, Dong Tian |
CVPR | 1 |
| 2021 | FESTA: Flow Estimation via Spatial-Temporal Attention for Scene Point CloudsabstractScene flow depicts the dynamics of a 3D scene, which is critical for various applications such as autonomous driving, robot navigation, AR/VR, etc. Conventionally, scene flow is estimated from dense/regular RGB video frames. With the development of depth-sensing technologies, precise 3D measurements are available via point clouds which have sparked new research in 3D scene flow. Nevertheless, it remains challenging to extract scene flow from point clouds due to the sparsity and irregularity in typical point cloud sampling patterns. One major issue related to irregular sampling is identified as the randomness during point set abstraction/feature extraction—an elementary process in many flow estimation scenarios. A novel Spatial Abstraction with Attention (SA2) layer is accordingly proposed to alleviate the unstable abstraction problem. Moreover, a Temporal Abstraction with Attention (TA2) layer is proposed to rectify attention in temporal domain, leading to benefits with motions scaled in a larger range. Extensive analysis and experiments verified the motivation and significant performance gains of our method, dubbed as Flow Estimation via Spatial-Temporal Attention (FESTA), when compared to several state-of-the-art benchmarks of scene flow estimation. Haiyan Wang 0019, Jiahao Pang, Muhammad Asad Lodhi, Yingli Tian, Dong Tian |
CVPR | 2 |
| 2020 | 3D Point Cloud Enhancement Using Graph-Modelled Multiview Depth MeasurementsabstractA 3D point cloud is often synthesized from depth measurements collected by sensors at different viewpoints. The acquired measurements are typically both coarse in precision and corrupted by noise. To improve quality, previous works denoise a synthesized 3D point cloud a posteriori, after projecting the imperfect depth data onto the 3D space. Instead, we enhance depth measurements on the sensed images a priori, exploiting inherent 3D geometric correlation across views, before synthesizing a 3D point cloud from the improved measurements. By enhancing closer to the actual sensing process, we benefit from optimization targeting specifically the depth image formation model, before subsequent processing steps that can further obscure measurement errors. Mathematically, for each pixel row in a pair of rectified viewpoint depth images, we first construct a graph reflecting inter-pixel similarities via metric learning using data in previous enhanced rows. To optimize left and right viewpoint images simultaneously, we write a non-linear mapping function from left pixel row to the right based on 3D geometry relations. We formulate a MAP optimization problem, which, after suitable linear approximations, results in an unconstrained convex and differentiable objective, solvable using fast gradient method (FGM). Experimental results show that our method noticeably outperforms recent denoising algorithms that enhance after 3D point clouds are synthesized. Xue Zhang 0008, Gene Cheung, Jiahao Pang, Dong Tian |
ICIP | 3 |
| 2020 | 3D Point Cloud Denoising Using Graph Laplacian Regularization of a Low Dimensional Manifold Modelabstract3D point cloud-a new signal representation of volumetric objects-is a discrete collection of triples marking exterior object surface locations in 3D space. Conventional imperfect acquisition processes of 3D point cloud-e.g., stereo-matching from multiple viewpoint images or depth data acquired directly from active light sensors-imply non-negligible noise in the data. In this paper, we extend a previously proposed low-dimensional manifold model for the image patches to surface patches in the point cloud, and seek self-similar patches to denoise them simultaneously using the patch manifold prior. Due to discrete observations of the patches on the manifold, we approximate the manifold dimension computation defined in the continuous domain with a patch-based graph Laplacian regularizer, and propose a new discrete patch distance measure to quantify the similarity between two same-sized surface patches for graph construction that is robust to noise. We show that our graph Laplacian regularizer leads to speedy implementation and has desirable numerical stability properties given its natural graph spectral interpretation. Extensive simulation results show that our proposed denoising scheme outperforms state-of-the-art methods in objective metrics and better preserves visually salient structural features like edges. Jin Zeng 0004, Gene Cheung, Michael Kwok-Po Ng, Jiahao Pang, Cheng Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2019 | Deep End-to-End Alignment and Refinement for Time-of-Flight RGB-D ModuleabstractRecently, it is increasingly popular to equip mobile RGB cameras with Time-of-Flight (ToF) sensors for active depth sensing. However, for off-the-shelf ToF sensors, one must tackle two problems in order to obtain high-quality depth with respect to the RGB camera, namely 1) online calibration and alignment; and 2) complicated error correction for ToF depth sensing. In this work, we propose a framework for jointly alignment and refinement via deep learning. First, a cross-modal optical flow between the RGB image and the ToF amplitude image is estimated for alignment. The aligned depth is then refined via an improved kernel predicting network that performs kernel normalization and applies the bias prior to the dynamic convolution. To enrich our data for end-to-end training, we have also synthesized a dataset using tools from computer graphics. Experimental results demonstrate the effectiveness of our approach, achieving state-of-the-art for ToF refinement. Di Qiu, Jiahao Pang, Wenxiu Sun, Chengxi Yang |
ICCV | 2 |
| 2018 | DSR: Direct Self-Rectification for Uncalibrated Dual-Lens CamerasabstractWith the developments of dual-lens camera modules, depth information representing the third dimension of the captured scenes becomes available for smartphones. It is estimated by stereo matching algorithms, taking as input the two views captured by dual-lens cameras at slightly different viewpoints. Depth-of-field rendering (also be referred to as synthetic defocus or bokeh) is one of the trending depth-based applications. However, to achieve fast depth estimation on smartphones, the stereo pairs need to be rectified in the first place. In this paper, we propose a cost-effective solution to perform stereo rectification for dual-lens cameras called direct self-rectification, short for DSR. It removes the need of individual offline calibration for every pair of dual-lens cameras. In addition, the proposed solution is robust to the slight movements, {\it e.g.}, due to collisions, of the dual-lens cameras after fabrication. Different with existing self-rectification approaches, our approach computes the homography in a novel way with zero geometric distortions introduced to the master image. It is achieved by directly minimizing the vertical displacements of corresponding points between the original master image and the transformed slave image. Our method is evaluated on both realistic and synthetic stereo image pairs, and produces superior results compared to the calibrated rectification or other self-rectification approaches. Ruichao Xiao, Wenxiu Sun, Jiahao Pang, Qiong Yan, Jimmy S. J. Ren |
3DV | 3 |
| 2018 | Single View Stereo MatchingabstractPrevious monocular depth estimation methods take a single view and directly regress the expected results. Though recent advances are made by applying geometrically inspired loss functions during training, the inference procedure does not explicitly impose any geometrical constraint. Therefore these models purely rely on the quality of data and the effectiveness of learning to generalize. This either leads to suboptimal results or the demand of huge amount of expensive ground truth labelled data to generate reasonable results. In this paper, we show for the first time that the monocular depth estimation problem can be reformulated as two sub-problems, a view synthesis procedure followed by stereo matching, with two intriguing properties, namely i) geometrical constraints can be explicitly imposed during inference; ii) demand on labelled depth data can be greatly alleviated. We show that the whole pipeline can still be trained in an end-to-end fashion and this new formulation plays a critical role in advancing the performance. The resulting model outperforms all the previous monocular depth estimation methods as well as the stereo block matching method in the challenging KITTI dataset by only using a small number of real training data. The model also generalizes well to other monocular depth estimation benchmarks. We also discuss the implications and the advantages of solving monocular depth estimation using stereo methods. Jimmy S. J. Ren, Mude Lin, Jiahao Pang, Wenxiu Sun, Hongsheng Li 0001, Liang Lin 0004 |
CVPR | 4 |
| 2018 | LSTM Pose MachinesabstractWe observed that recent state-of-the-art results on single image human pose estimation were achieved by multistage Convolution Neural Networks (CNN). Notwithstanding the superior performance on static images, the application of these models on videos is not only computationally intensive, it also suffers from performance degeneration and flicking. Such suboptimal results are mainly attributed to the inability of imposing sequential geometric consistency, handling severe image quality degradation (e.g. motion blur and occlusion) as well as the inability of capturing the temporal correlation among video frames. In this paper, we proposed a novel recurrent network to tackle these problems. We showed that if we were to impose the weight sharing scheme to the multi-stage CNN, it could be re-written as a Recurrent Neural Network (RNN). This property decouples the relationship among multiple network stages and results in significantly faster speed in invoking the network for videos. It also enables the adoption of Long Short-Term Memory (LSTM) units between video frames. We found such memory augmented RNN is very effective in imposing geometric consistency among frames. It also well handles input quality degradation in videos while successfully stabilizes the sequential outputs. The experiments showed that our approach significantly outperformed current state-of-the-art methods on two large-scale video pose estimation benchmarks. We also explored the memory cells inside the LSTM and provided insights on why such mechanism would benefit the prediction for video-based pose estimations.1 Jimmy S. J. Ren, Zhouxia Wang, Wenxiu Sun, Jinshan Pan, Jiahao Pang, Liang Lin 0004 |
CVPR | 7 |
| 2018 | Zoom and Learn: Generalizing Deep Stereo Matching to Novel DomainsabstractDespite the recent success of stereo matching with convolutional neural networks (CNNs), it remains arduous to generalize a pre-trained deep stereo model to a novel domain. A major difficulty is to collect accurate ground-truth disparities for stereo pairs in the target domain. In this work, we propose a self-adaptation approach for CNN training, utilizing both synthetic training data (with ground-truth disparities) and stereo pairs in the new domain (without ground-truths). Our method is driven by two empirical observations. By feeding real stereo pairs of different domains to stereo models pre-trained with synthetic data, we see that: i) a pre-trained model does not generalize well to the new domain, producing artifacts at boundaries and ill-posed regions; however, ii) feeding an up-sampled stereo pair leads to a disparity map with extra details. To avoid i) while exploiting ii), we formulate an iterative optimization problem with graph Laplacian regularization. At each iteration, the CNN adapts itself better to the new domain: we let the CNN learn its own higher-resolution output; at the meanwhile, a graph Laplacian regularization is imposed to discriminatively keep the desired edges while smoothing out the artifacts. We demonstrate the effectiveness of our method in two domains: daily scenes collected by smart-phone cameras, and street views captured in a driving car. Jiahao Pang, Wenxiu Sun, Chengxi Yang, Jimmy S. J. Ren, Ruichao Xiao, Jin Zeng 0004, Liang Lin 0004 |
CVPR | 1 |
| 2017 | Accurate Single Stage Detector Using Recurrent Rolling ConvolutionabstractMost of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are deep in context. We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available. Jimmy S. J. Ren, Xiaohao Chen, Wenxiu Sun, Jiahao Pang, Qiong Yan, Yu-Wing Tai, Li Xu 0001 |
CVPR | 5 |
| 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral CorrelationabstractColor filter array (CFA) interpolation, or three-band demosaicking, is a process of interpolating the missing color samples in each band to reconstruct a full color image. In this paper, we are concerned with the challenging problem of multispectral demosaicking, where each band is significantly undersampled due to the increment in the number of bands. Specifically, we demonstrate a frequency-domain analysis of the subsampled color-difference signal and observe that the conventional assumption of highly correlated spectral bands for estimating undersampled components is not precise. Instead, such a spectral correlation assumption is image dependent and rests on the aliasing interferences among the various color-difference spectra. To address this problem, we propose an adaptive spectral-correlation-based demosaicking (ASCD) algorithm that uses a novel anti-aliasing filter to suppress these interferences, and we then integrate it with an intra-prediction scheme to generate a more accurate prediction for the reconstructed image. Our ASCD is computationally very simple, and exploits the spectral correlation property much more effectively than the existing algorithms. Experimental results conducted on two data sets for multispectral demosaicking and one data set for CFA demosaicking demonstrate that the proposed ASCD outperforms the state-of-the-art algorithms. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Jiahao Pang, Klaus Mueller 0001, Oscar C. Au |
IEEE Trans. Image Process. | 4 |
| 2017 | Graph Laplacian Regularization for Image Denoising: Analysis in the Continuous DomainabstractInverse imaging problems are inherently underdetermined, and hence, it is important to employ appropriate image priors for regularization. One recent popular prior-the graph Laplacian regularizer-assumes that the target pixel patch is smooth with respect to an appropriately chosen graph. However, the mechanisms and implications of imposing the graph Laplacian regularizer on the original inverse problem are not well understood. To address this problem, in this paper, we interpret neighborhood graphs of pixel patches as discrete counterparts of Riemannian manifolds and perform analysis in the continuous domain, providing insights into several fundamental aspects of graph Laplacian regularization for image denoising. Specifically, we first show the convergence of the graph Laplacian regularizer to a continuous-domain functional, integrating a norm measured in a locally adaptive metric space. Focusing on image denoising, we derive an optimal metric space assuming non-local self-similarity of pixel patches, leading to an optimal graph Laplacian regularizer for denoising in the discrete domain. We then interpret graph Laplacian regularization as an anisotropic diffusion scheme to explain its behavior during iterations, e.g., its tendency to promote piecewise smooth signals under certain settings. To verify our analysis, an iterative image denoising algorithm is developed. Experimental results show that our algorithm performs competitively with state-of-the-art denoising methods, such as BM3D for natural images, and outperforms them significantly for piecewise smooth images. Jiahao Pang, Gene Cheung |
IEEE Trans. Image Process. | 1 |
| 2016 | Subpixel-Based Image Scaling for Grid-like Subpixel Arrangements: A Generalized Continuous-Domain Analysis ModelabstractSubpixel-based image scaling can improve the apparent resolution of displayed images by controlling individual subpixels rather than whole pixels. However, improved luminance resolution brings chrominance distortion, making it crucial to suppress color error while maintaining sharpness. Moreover, it is challenging to develop a scheme that is applicable for various subpixel arrangements and for arbitrary scaling factors. In this paper, we address the aforementioned issues by proposing a generalized continuous-domain analysis model, which considers the low-pass nature of the human visual system (HVS). Specifically, given a discrete image and a grid-like subpixel arrangement, the signal perceived by the HVS is modeled as a 2D continuous image. Minimizing the difference between the perceived image and the continuous target image leads to the proposed scheme, which we call continuous-domain analysis for subpixel-based scaling (CASS). To eliminate the ringing artifacts caused by the ideal low-pass filtering in CASS, we propose an improved scheme, which we call CASS with Laplacian-of-Gaussian filtering. Experiments show that the proposed methods provide sharp images with negligible color fringing artifacts. Our methods are comparable with the state-of-the-art methods when applied on the RGB stripe arrangement, and outperform existing methods when applied on other subpixel arrangements. Jiahao Pang, Lu Fang 0001, Jin Zeng 0004, Yuanfang Guo, Ketan Tang |
IEEE Trans. Image Process. | 1 |
| 2016 | Subpixel Image Quality Assessment Syncretizing Local Subpixel and Global Pixel FeaturesabstractThe subpixel rendering technology increases the apparent resolution of an LCD/OLED screen by exploiting the physical property that a pixel is composed of RGB individually addressable subpixels. Due to the intrinsic intercoordination between apparent luminance resolution and color fringing artifact, a common method of subpixel image assessment is subjective evaluation. In this paper, we propose a unified subpixel image quality assessment metric called subpixel image assessment (SPA), which syncretizes local subpixel and global pixel features. Specifically, comprehensive subjective studies are conducted to acquire data of user preferences. Accordingly, a collection of low-level features is designed under extensive perceptual validation, capturing subpixel and pixel features, which reflect local details and global distance from the original image. With the features and their measurements as the basis, the SPA is obtained, which leads to a good representation of the subpixel image characteristics. The experimental results justify the effectiveness and the superiority of the SPA. The SPA is also successfully adopted in a variety of applications, including content adaptive sampling and metric-guided image compression. Jin Zeng 0004, Lu Fang 0001, Jiahao Pang, Houqiang Li, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Optimal graph laplacian regularization for natural image denoisingabstractImage denoising is an under-determined problem, and hence it is important to define appropriate image priors for regularization. One recent popular prior is the graph Laplacian regularizer, where a given pixel patch is assumed to be smooth in the graph-signal domain. The strength and direction of the resulting graph-based filter are computed from the graph's edge weights. In this paper, we derive the optimal edge weights for local graph-based filtering using gradient estimates from non-local pixel patches that are self-similar. To analyze the effects of the gradient estimates on the graph Laplacian regularizer, we first show theoretically that, given graph-signal hDis a set of discrete samples on continuous function h(x; y) in a closed region Ω, graph Laplacian regularizer (hD)TLhDconverges to a continuous functional SΩintegrating gradient norm of h in metric space G-i.e., (∇h)TG-1(∇h)-over Ω. We then derive the optimal metric space G*: one that leads to a graph Laplacian regularizer that is discriminant when the gradient estimates are accurate, and robust when the gradient estimates are noisy. Finally, having derived G* we compute the corresponding edge weights to define the Laplacian L used for filtering. Experimental results show that our image denoising algorithm using the per-patch optimal metric space G* outperforms non-local means (NLM) by up to 1.5 dB in PSNR. Jiahao Pang, Gene Cheung, Antonio Ortega, Oscar C. Au |
ICASSP | 1 |
| 2015 | Image colorization via color propagation and rank minimizationabstractImage colorization aims to add colors to grayscale images, which used to be a time-consuming and tedious task that requires lots of human efforts. In this paper, we present a novel colorization method based on color propagation and rank minimization. Given a small portion of chrominance values and a grayscale image, we firstly propagate the known color values to other pixels to be colorized. As the colorized image after color propagation is not accurate, we then define a confidence matrix to measure the propagation fidelity. Finally, pixels that have propagated chrominance values with confidence are colorized by rank minimization, which exploits the redundancy of natural images. Experimental results on real data set show that our proposed method achieves state-of-the-art colorization quality. Yonggen Ling, Oscar C. Au, Jiahao Pang, Jin Zeng 0004, Yuan Yuan 0002, Amin Zheng |
ICIP | 3 |
| 2014 | Analysis of sampling pattern and Luma-Chroma filter design for subpixel-based image downsamplingabstractSubpixel-based image downsampling is attractive in that it produces higher apparent resolution of down-sampled images on LCD displays. However increased luminance resolution is achieved at the price of color fringing artifacts. In this paper, we propose an algorithm to find a pleasing balance between increased resolution and color fidelity. We separate the subpixel-based downsampling into two stages, shifting followed by downsampling with anti-aliasing filtering. In stage one, we find special characteristics of the luminance and chrominance spectra of the shifted image, based on which the optimal sampling pattern is found. In stage two, anti-aliasing filters for luminance and chrominance are designed respectively. Experimental results verify that the proposed method manages to suppress color artifacts while maintaining high luminance sharpness. Jin Zeng 0004, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Ketan Tang, Yonggen Ling |
ICASSP | 4 |
| 2014 | Self-similarity-based image colorizationabstractIn this work, we tackle the problem of coloring black-and-white images, which is image colorization. Existing image colorization algorithms can be categorized into two types: scribble-based colorization algorithms and example-based colorization algorithms. Differently, we propose a hybrid scheme that combines the advantages of both categories. Given the grayscale image to be colorized and a few color scribbles (or scattered color labels) as input, the proposed method manages to colorize the grayscale image with high quality. Similar to the mechanisms in example-based colorization methods, our algorithm firstly propagates chrominance information based on the assumption that similar image patches should have similar colors. Therefore colors of some pixels can be transferred from similar patches with known colors. After that, we apply scribble-based colorization algorithm to fully colorize the grayscale image, with different confidences assigned onto the transferred color labels. Experimental results show that, the proposed method effectively utilizes the known chrominance, and provides pleasant colorizations with very few user interventions. Jiahao Pang, Oscar C. Au, Yukihiko Yamashita, Yonggen Ling, Yuanfang Guo, Jin Zeng 0004 |
ICIP | 1 |
| 2014 | High bit-precision image acquisition and reconstruction by planned sensor distortionabstractWe present a novel framework for high bit-precision image acquisition and reconstruction. This framework is designed based on the inherent Markov property of image signals. In acquisition stage, we add planned sensor distortion (PSD) to the analog image signal before feeding it to A/D converters (or quantizers) in camera sensor. In reconstruction stage, the acquired quantized pixel values are jointly combined to get the reconstructed signal with reduced uncertainty range. Advantages of proposed PSD framework include 1) simplicity: it does not require any change to the core hardware of existing A/D converters; 2) effectiveness: experiment results demonstrate significant PSNR gain over traditional methods (up to 10 dB when quantizer bit-depth is relatively low); and 3) generality: this framework can also be applied for acquisition of other analog signals, including audio, video, etc. Pengfei Wan 0001, Oscar C. Au, Jiahao Pang, Ketan Tang |
ICIP | 3 |
| 2014 | Natural image matting via adaptive local and nonlocal sample clusteringabstractDigital image matting is the determination of foreground color, background color, and an opacity value of each pixel for an input image. Inherently, matting is a highly ill-posed and under-constrained problem. Thus, some assumptions need to be made to resolve it. Inspired by closed-form matting and color clustering matting, in this work, we first develop an adaptive sample clustering criterion to automatically assign either local or nonlocal neighborhood to each pixel. After that, in order to enhance matting accuracy, we improve the nonlocal clustering performance by introducing a new feature selection parameter to choose preferred feature space for different images in a fully automatic way. And finally we solve the problem using a closed form solution. Experimental results show that our algorithm achieves equal or even better performance among many state-of-the-art matting techniques. Oscar C. Au, Yuan Yuan 0002, Wenxiu Sun, Yonggen Ling, Jiahao Pang |
ICIP | 6 |
| 2014 | Intra prediction with adaptive CU processing order in HEVCabstractThe High Efficiency Video Coding (HEVC) utilizes Z-scan order to process coding units (CUs). For intra prediction, this order cannot fully exploit the spatial correlation between adjacent CUs. After transform and quantization, the residue still contains lots of energy along edges which consumes many bits for compression. To effectively reduce the residue energy along edges, a novel intra prediction approach is proposed, where the CU processing order is changed adaptively. Two additional orders are introduced in this paper besides traditional Z-scan order. Up to 1.9% bit saving is achieved in our experiments on HEVC test model. We also propose two fast order selection algorithms and the observed gains are obtained with 27% and 2% encoding time increase compared to HEVC, respectively. Amin Zheng, Oscar C. Au, Yuan Yuan 0002, Haitao Yang 0001, Jiahao Pang, Yonggen Ling |
ICIP | 5 |
| 2014 | Fast algorithm of arbitrary factor subpixel downsampling based on frequency analysisabstractSubpixel-based downsampling has shown its advantages over pixel-based downsampling in terms of preserving more spatial details along edges and generating sharper images, at the cost of certain amount of color-fringing artifacts in the downsampled image. To balance the sharpness and color-fringing artifacts, some algorithms are proposed to design optimal anti-aliasing (AA) filters, which are either image independent, or computationally too expensive. And all of the existing AA filters are designed for fixed downsampling factor, which makes them impractical for real applications. In this paper we propose two fast algorithms to design AA filter for arbitrary factor subpixel downsampling based on frequency analysis of the input image. The proposed algorithms generate image dependent AA filter which is as good as the state-of-the-art algorithm, but much faster. Ketan Tang, Oscar C. Au, Lu Fang 0001, Jiahao Pang, Yuanfang Guo |
ICME | 4 |
| 2014 | Photo album compression By leveraging temporal-spatial correlations and HEVCabstractThe advancing digital photography technology has resulted in a large number of photos stored in personal computers. Photo album compression algorithms aim to save storage space and efficiently manage photos. In this paper, a general forest structure model involving depth constrain for photo album compression is proposed, which further exploits the correlations between images in the photo album. We firstly represent the images as nodes in a graph and directed edges between them as predictive coding relationship. Affinity propagation is then applied to compute for a depth-constrained forest. Finally, we adopt depth-first search algorithm to generate the compression order according to forest structure and HEVC to compress the images with adaptive GOPs and reference list. Experimental results show that the proposed compression method provides much better rate-distortion performance compared to JPEG and significantly reduce the storage space. Yonggen Ling, Oscar C. Au, Ruobing Zou, Jiahao Pang, Amin Zheng |
ISCAS | 4 |
| 2013 | Image colorization using sparse representationabstractImage colorization is the task to color a grayscale image with limited color cues. In this work, we present a novel method to perform image colorization using sparse representation. Our method first trains an over-complete dictionary in YUV color space. Then taking a grayscale image and a small subset of color pixels as inputs, our method colorizes overlapping image patches via sparse representation; it is achieved by seeking sparse representations of patches that are consistent with both the grayscale image and the color pixels. After that, we aggregate the colorized patches with weights to get an intermediate result. This process iterates until the image is properly colorized. Experimental results show that our method leads to high-quality colorizations with small number of given color pixels. To demonstrate one of the applications of the proposed method, we apply it to transfer the color of one image onto another to obtain a visually pleasing image. Jiahao Pang, Oscar C. Au, Ketan Tang, Yuanfang Guo |
ICASSP | 1 |
| 2013 | Arbitrary factor image interpolation using geodesic distance weighted 2D autoregressive modelingabstractLeast square regression has been widely used in image interpolation. Some existing regression-based interpolation methods used ordinary least squares (OLS) to formulate cost functions. These methods usually have difficulties at object boundaries because OLS is sensitive to outliers. Weighted least squares (WLS) is then adopted to solve the outlier problem. Some weighting schemes have been proposed in the literature. In this paper we propose to use geodesic distance weighting in that geodesic distance can simultaneously measure both the spatial distance and color difference. Another contribution of this paper is that we propose an optimization scheme that can handle arbitrary factor interpolation. The idea is to separate the problem into two parts, an adaptive pixel correlation model and a convolution based image degradation model. Geodesic distance weighted 2D autoregressive model is used to model the pixel correlation which preserves local geometry. The convolution based image degradation model provides the flexibility to handle arbitrary interpolation factor. The entire problem is formulated as a WLS problem constrained by a linear equality. Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang |
ICASSP | 4 |
| 2013 | Arbitrary factor image interpolation by convolution kernel constrained 2-D autoregressive modelingabstractAmong existing interpolation methods, convolution-based methods are able to perform arbitrary factor interpolation but the results are usually blurry or jaggy, adaptive interpolation methods usually can reduce the blurry and jaggy artifacts but cannot handle arbitrary factor interpolation. In this paper we propose an arbitrary factor adaptive interpolation algorithm by combining 2-D piecewise autoregressive (PAR) modeling and convolution kernel constraint. PAR model ensures local geometries are well preserved thus the resultant image is not blurry or jaggy. Convolution kernel constraint ensures the recovered high resolution image consistent with the low resolution image, and also provides the flexibility to handle arbitrary interpolation factor. Experiment results show that our algorithm achieves state-of-the-art performance for any interpolation factor. Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Lu Fang 0001 |
ICIP | 4 |
| 2013 | Color clustering mattingabstractNatural image matting refers to the problem of extracting regions of interest such as foreground object from an image based on user inputs like scribbles or trimap. More specifically, we need to estimate the color information of background, foreground and the corresponding opacity, which is an ill-posed problem inherently. Inspired by closed-form matting and KNN matting, in this paper, we extend the local color line model which is based on the assumption of linear color clustering within a small local window, to nonlocal feature space neighborhood. New affinity matrix is defined to achieve better clustering. Further, we demonstrate that good clustering ensures better prediction of alpha matte. Experimental evaluations on benchmark datasets and comparisons show that our matting algorithm is of higher accuracy and better visual quality than some state-of-the-art matting algorithms. Yongfang Shi, Oscar C. Au, Jiahao Pang, Ketan Tang, Wenxiu Sun, Hong Zhang 0024, Luheng Jia |
ICME | 3 |
| 2013 | Data hiding in error diffused color halftone imagesabstractHalftone image watermarking has been explored and developed rapidly over the past decade. However, there are still issues to be studied. This paper presents a data hiding method called Data Hiding by Dual Color Conjugate Error Diffusion (DHDCCED) to hide a binary secret pattern into two error diffused color halftone images, such that when the two color halftone images are overlaid, the secret pattern will be revealed. The experimental results show that DHDCCED can significantly improve the performances when comparing both the correct decoding rate and the visual quality of the revealed secret pattern to the existing method Color Conjugate Error Diffusion (CCED). Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang, Wenxiu Sun, Lingfeng Xu |
ISCAS | 4 |
| 2013 | Hiding a Secret Pattern into Color Halftone Images
Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang |
IWDW | 4 |
| 2013 | Chroma Replacing and adaptive Chroma Blending for subpixel-based downsamplingabstractSubpixel-based downsampling generates images with higher apparent resolution with the expense of annoying color-fringing artifacts near strong edges. In this paper we propose two methods that find a balance in the tradeoff of apparent resolution and color-fringing artifacts. The first method is called Chroma Replacing in which the color-fringing artifacts are completely removed but the subpixel rendering effect is also removed. The second one is called Chroma Blending in which only the color-fringing artifacts that are strong enough to be noticed are removed, and also the subpixel rendering effect is retained. We also propose two objective measures for measuring the similarity of downsampled image to the original image. Experiment results show that the proposed methods are effective in removing color-fringing artifacts, without harming the high apparent resolution. Ketan Tang, Oscar C. Au, Lu Fang 0001, Yuanfang Guo, Jiahao Pang |
MMSP | 5 |
| 2013 | An analytical study of subpixel-based image down-sampling patterns in frequency domainabstractSubpixel-based image down-sampling is a class of methods that can provide improved apparent resolution of the down-scaled image compared to the pixel-based methods. The frequency characteristics of all possible subpixel-based down-sampling patterns for RGB vertical stripes are analytically studied in this paper. Our proposed algorithm reveals that there are merely seven equivalent energy distributions in the luminance frequency spectrum. To achieve higher luminance resolution, we then calculate and choose the optimal down-sampling pattern with anti-aliasing low-pass filter designed for it so as to maximize the energy of the luminance component within the cut-off shape. Experimental results show that the proposed method provides sharper images compared to the state-of-art subpixel-based methods, with little color distortion. Yonggen Ling, Oscar C. Au, Ketan Tang, Jiahao Pang, Jin Zeng 0004, Lu Fang 0001 |
VCIP | 4 |