VLDB 2026 Research / reviewers in the wild / expert
Jui-Chiu Chiang
dblp:34/1566
· DBLP profile ↗
40ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0003-1397-8393ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 9 first-author · 20 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS CompressionabstractBuilding on the success of 3D Gaussian Splatting (3DGS) in static 3D scene representation, its extension to dynamic scenes-commonly referred to as 4DGS or dynamic 3DGS- has attracted increasing attention. However, designing more compact, efficient deformation schemes together with rate-distortion-optimized compression strategies for dynamic 3DGS representations remains an underexplored area. Prior methods either rely on space-time 4DGS with overspecified, short-lived Gaussian primitives or on canonical 3DGS with deformation that lacks explicit temporal control. To address this, we present TED-4DGS, a temporally activated and embedding-based deformation scheme for rate-distortion- optimized 4DGS compression that unifies the strengths of both families. TED-4DGS is built on a sparse anchor-based 3DGS representation. Each canonical anchor is assigned with learnable temporal-activation parameters to specify its appearance and disappearance transitions over time, while a lightweight per-anchor temporal embedding queries a shared deformation bank to produce anchor-specific deformation. For rate-distortion compression, we incorporate an implicit neural representation (INR)-based hyperprior to model anchor attribute distributions, along with a channelwise autoregressive model to capture intra-anchor correlations. With these novel elements, our scheme achieves the state-of-the-art rate-distortion performance on several commonly used real-world datasets. To the best of our knowledge, this work represents one of the first attempts to pursue a rate-distortion-optimized compression framework for dynamic 3DGS representations. Cheng-Yuan Ho, Hebi Yang, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng |
WACV | 3 |
| 2026 | MEGA-PCC: A Mamba-based Efficient Approach for Joint Geometry and Attribute Point Cloud CompressionabstractJoint compression of point cloud geometry and attributes is essential for efficient 3D data representation. Existing methods often rely on post-hoc recoloring procedures and manually tuned bitrate allocation between geometry and attribute bitstreams in inference, which hinders end-to-end optimization and increases system complexity. To overcome these limitations, we propose MEGA-PCC, a fully end-to-end, learning-based framework featuring two specialized models for joint compression. The main compression model employs a shared encoder that encodes both geometry and attribute information into a unified latent representation, followed by dual decoders that sequentially reconstruct geometry and then attributes. Complementing this, the Mamba-based Entropy Model (MEM) enhances entropy coding by capturing spatial and channel-wise correlations to improve probability estimation. Both models are built on the Mamba architecture to effectively model long-range dependencies and rich contextual features. By eliminating the need for recoloring and heuristic bitrate tuning, MEGA-PCC enables data-driven bitrate allocation during training and simplifies the overall pipeline. Extensive experiments demonstrate that MEGA-PCC achieves superior rate-distortion performance and runtime efficiency compared to both traditional and learning-based baselines, offering a powerful solution for AI-driven point cloud compression. Kai Hsiang Hsieh, Monyneath Yim, Wen-Hsiao Peng, Jui-Chiu Chiang |
WACV | 4 |
| 2025 | CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compressionabstract3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application calls for rate-distortion-optimized 3DGS compression. How to quantize and entropy encode sparse Gaussian primitives in the 3D space remains largely unexplored. Few early attempts resort to the hyperprior framework from learned image compression. But, they fail to utilize fully the inter and intra correlation inherent in Gaussian primitives. Built on ScaffoldGS, this work, termed CAT-3DGS, introduces a context-adaptive triplane approach to their rate-distortion-optimized coding. It features multi-scale triplanes, oriented according to the principal axes of Gaussian primitives in the 3D space, to capture their inter correlation (i.e. spatial correlation) for spatial autoregressive coding in the projected 2D planes. With these triplanes serving as the hyperprior, we further perform channel-wise autoregressive coding to leverage the intra correlation within each individual Gaussian primitive. Our CAT-3DGS incorporates a view frequency-aware masking mechanism. It actively skips from coding those Gaussian primitives that potentially have little impact on the rendering quality. When trained end-to-end to strike a good rate-distortion trade-off, our CAT-3DGS achieves the state-of-the-art compression performance on the commonly used real-world datasets. Yu-Ting Zhan, Cheng-Yuan Ho, Hebi Yang, Yi-Hsin Chen, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng |
ICLR | 5 |
| 2025 | LiDAR Point Cloud Upsampling with Mamba ArchitectureabstractThis paper presents a range-image-based LiDAR point cloud upsampling method, utilizing the State Space Model (SSM) framework. The proposed method significantly reduces computational demands while maintaining high performance, improving efficiency in LiDAR point cloud upsampling tasks. We validated the effectiveness of our approach through experiments on the KITTI dataset. Results demonstrate that the SSM-based LiDAR upsampling method achieves a 69% reduction in FLOPs and a 72% decrease in number of parameters compared to the Swin Transformer model. This efficient solution offers enhanced LiDAR point cloud resolution, making it a promising approach for cost-effective applications in autonomous driving and related fields. Yu-Chi Chung, Che-Lu Chang, Jui-Chiu Chiang |
ISCAS | 3 |
| 2025 | Enhancing Inter-Frame Coding in V-DMC Through Dynamic Group Partitioning
Yu-Hsuan Yeh, Pin-Chang Pan, Jui-Chiu Chiang |
PCS | 3 |
| 2025 | HDR-CNF: single-image high dynamic range imaging based on conditional normalizing flows
Kai-Wei Peng, Jui-Chiu Chiang, Sau-Gee Chen, Yu-Shan Lin |
Multim. Tools Appl. | 2 |
| 2024 | Mamba-PCGC: Mamba-Based Point Cloud Geometry CompressionabstractPoint cloud compression plays a critical role in efficiently managing extensive 3D datasets, thereby facilitating their practical utilization. Robust feature extraction is essential for enabling learned codecs to attain optimal compression performance. Recent advancements in state space models (SSMs), coupled with efficient hardware-aware designs such as Mamba, have exhibited significant potential for effectively modeling long sequences. In light of these developments, we propose a novel approach: Mamba-based point cloud geometry compression (Mamba-PCGC), which not only achieves linear computational complexity but also preserves expansive receptive fields to better extract latent features. The comprehensive results from our experiments strongly demonstrate the effectiveness of our proposed approach. When compared against established point cloud coding standards, specifically G-PCC and V-PCC, Mamba-PCGC demonstrates remarkable bitrate savings of $50.0 \%$ and $34.3 \%$, respectively, in terms of D1-PSNR. Furthermore, in a comparison with PCGCv2, our approach consistently outperforms, showcasing an average bitrate reduction of $13.1 \%$ in terms of D1-PSNR. Monyneath Yim, Jui-Chiu Chiang |
ICIP | 2 |
| 2024 | BMT-PCGC: Point Cloud Geometry Compression with Bidirectional Mask Transformer Entropy ModelabstractPoint cloud compression plays a critical role in efficiently handling large 3D datasets, enabling their practical usage. An accurate entropy model is essential for learned codecs to achieve good compression performance. Although autoregressive entropy models can explore dependencies in large contexts, they are inefficient in terms of complexity and may miss information from the inverse direction. To improve coding efficiency while guaranteeing fast decoding, we propose a bidirectional mask Transformer entropy model (BMTEM) for point cloud geometry compression, by leveraging bidirectional self-attention in the mask Transformer. The pre-defined mask schedules facilitate group-wise autoregression, which improves parallel computation for faster inference compared to a fully autoregressive approach. Comparative evaluations against point cloud coding standards G-PCC and V-PCC reveal that BMTEM achieves significant bitrate savings of 53.22% and 47.18% in terms of D1-PSNR, respectively. When compared to other deep learning-based methods like PCGCv2 and ANFPCGC, our approach demonstrates average bitrate reductions of 20.01% and 17.67% in terms of D1-PSNR, respectively. Monyneath Yim, Bing-Han Wu, Jui-Chiu Chiang |
PCS | 3 |
| 2024 | Enhancing Global Tetris Packing in V-PCC Through Dynamic Group PartitioningabstractVideo-based point cloud compression (V-PCC) adopts a projection-based methodology that leverages established video compression techniques. To maintain temporal consistency within the packed images, V-PCC facilitates global patch packing solutions that aim to align corresponding patches across frames during the packing process. One such technique is global tetris packing (GTP), particularly efficient in scenarios with low-motion activity within point cloud sequences. However, as the length of the point cloud sequence grows, accumulated motion within the input sequence can lead to increasing inconsistencies between patches across images, degrading the effectiveness of GTP. In this paper, we present a novel approach to enhance GTP. Our method measures the similarity between point cloud frames and dynamically partitions the sequence into variable groups of frames (GOFs). This strategy enhances temporal consistency within each group, thereby improving coding efficiency. Experimental results validate the superiority of our proposed technique over the anchor, delivering average BD-rate savings of 3.62%, 3.24%, and 2.16% for D1-PSNR, D2-PSNR, and Y-PSNR, respectively. Yun-Chang Tsai, Jui-Chiu Chiang |
VCIP | 2 |
| 2024 | CRC-DPCGC: Conditional Residual Coding for Dynamic Point Cloud Geometry CompressionabstractDynamic point cloud compression faces significant challenges due to the inherent unstructured and sparse nature of point clouds, requiring efficient modeling of temporal context information. Current approaches often fall short in effectively capturing and utilizing inter-frame dependencies. This paper presents a novel conditional residual coding technique for dynamic point cloud geometry compression, specifically designed to address this challenge. Our method exploits spatial-temporal features by a predictor module and a Transformer-based entropy model. The predictor module transfers features from previous point cloud frames to the current frame, efficiently generating residual features. The Transformer entropy model employs a conditional coding strategy, utilizing cross-attention between the current frame and its predecessors, leading to a more precise bit estimation. Extensive experimental results demonstrate that our proposed method outperforms existing dynamic point cloud geometry compression techniques, including rule-based and learningbased approaches, showcasing its effectiveness in capturing and utilizing temporal information. Bing-Han Wu, Monyneath Yim, Jui-Chiu Chiang |
VCIP | 3 |
| 2024 | Enhanced Temporal Consistency for Global Patch Allocation in Video-Based Point Cloud CompressionabstractVideo-based point cloud compression (V-PCC) is a promising technique for compressing 3D point clouds. V-PCC projects the 3D point cloud into patches and encodes the generated 2D images using state-of-the-art video codecs. To maintain temporal consistency between frames, V-PCC supports global patch packing methods and one notable approach is Global Patch Allocation (GPA), which packs the global matched patches into the same location in each frame across the sequence. Additionally, frames are subdivided into groups (i.e., sub-contexts) to balance packing compactness and patch similarity within the groups. While video coding typically employs a Group of Picture (GOP) as the basic unit for encoding, GPA in V-PCC currently does not consider the reference relationship between images within or between GOPs, resulting in limited similarity between the current and the reference images, ultimately leading to reduced encoding efficiency. This paper presents an improved technique for GPA. We propose a dynamic sub-context and GOP determination technique, enhancing the similarity between images within the same GOP. Furthermore, we introduce a priority-based patch packing (PBPP) technique to reduce differences between frames in adjacent GOPs. Experimental results demonstrate the superiority of our proposed method over the anchor, achieving an average BD-rate savings of 3.09%, 3.04%, and 2.33% for D1-PSNR, D2-PSNR, and Y-PSNR, respectively. Jui-Chiu Chiang, Yu-Tze Wu, Hsin-Yun Hsieh, Yun-Chang Tsai |
IEEE Trans. Multim. | 1 |
| 2023 | Deep-Learning Technique for Risk-Based Action Prediction Using Extremely Low-Resolution Thermopile Sensor ArrayabstractEldering caring is important in today’s aging society, especially that accident anticipation/prevention plays an important role. In this paper, a novel approach to preventing elderly accidents based on a very low-resolution thermopile sensor array (TPA) (only$32\times 32$pixels) is proposed for prediction of bed-exit event that might lead to elderly falls in home caring. Low-resolution TPA sensor, capable of collecting far infrared energy, ensures cost-effective monitoring, no interference with user’s daily life, and most importantly privacy-preservation. Since most of the fall accidents occur when the elderly attempts to get off the bed without assistance, it is thus the focus of this paper to monitor his/her posture and action via TPA image sensor and then predict that an action of getting off the bed will occur in a near future (e.g.,$S$seconds later). Our system can raise an alarm to the caregivers so that they can intervene and offer the necessary assistance. A deep-learning model based on CNN-RNN (Convolutional neural network-Recurrent neural network) architecture was designed which is capable of predicting the elderly bed-exit intention by$S =5.78$seconds in advance of the action onset at an accuracy of 99.37% according to our dataset evaluation. Our system is also suitable for on-line real-time operation which will be helpful to elderly caring in our society. Igor Morawski, Wen-Nung Lie, Lee Aing, Jui-Chiu Chiang, Kuan-Ting Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | HDR-AGAN: Ghost-Free High Dynamic Range Imaging with Attention Guided Adversarial NetworkabstractCreating a high dynamic range (HDR) image from a set of low dynamic range (LDR) images with multiple exposures is always a challenging task when there is motion among the LDR image stacks. In this paper, we propose a generative adversarial network-based method with an attention module, called HDR- AGAN, to deal with the issue caused by object motions. Given three dynamic LDR images with different exposures, the well-exposed one is treated as the reference image, while the other two images are aligned to the reference one so that two additional inputs are obtained in the proposed system. An attention module operated between the reference image and the non-reference image is employed for extracting useful features. The generated HDR image is judged by the discriminator which operates on specified image areas derived by the ghost detection module. Experimental results show that, compared to existing algorithms, our method generates HDR images with less ghost artifacts and also yields better image quality in terms of several objective metrics. Jui-Chiu Chiang, Sau-Gee Chen |
ICIP | 2 |
| 2022 | 3D Head Pose Estimation Based on Graph Convolutional Network from A Single RGB ImageabstractResearch of head pose estimation in computer vision has been at the center of much attention. This work presents a framework based on adaptive graph convolution network (AGCN) to process both 2D and 3D facial landmarks extracted from the input RGB image. The network has a two-streams (teacher/3D-student/2D streams) architecture, trained with a 3D to 2D knowledge distillation training process, to transfer features of the 3D stream to the 2D stream for performance promotion. Several processing modules, such as depth-denoising for detected 3D landmarks, multi-stream fusion in inference, were also proposed for further increase of the prediction performance and robustness of our proposed method. In experiments, we follow standard protocols (in terms of datasets and metrices) to evaluate our performance. Three datasets 300W-LP, AFLW2000 and BIWI were used. The performance is measured in mean absolute error (MAE). We can achieve better performance compared to most of the state-of-the-art methods. Wen-Nung Lie, Monyneath Yim, Lee Aing, Jui-Chiu Chiang |
ICIP | 4 |
| 2022 | Remote PPG Estimation from RGB-NIR Facial Image Sequence for Heart Rate EstimationabstractThis paper presents a dual-modal (RGB-NIR) technique to estimate remote photoplethysmogram (rPPG) signal, i.e. the heart rate, from a facial image sequence. We developed denoising techniques with a modified amplitude selective filtering (ASF), wavelet decomposition and robust principal component analysis (RPCA), to enhance the uncovering of the rPPG signal through the well-known ICA algorithm. A new dataset built with RealSense RGB-D camera is considered in experiments: regular brightness, under-illumination, and face motion. Experimental results show that the proposed method has reached competitive performance among the state-of-the-art methods in motion and under-illuminated scenarios even at a shorter input video length (10 to 20 seconds). Dao-Quang Le, Jui-Chiu Chiang, Wen-Nung Lie |
ISCAS | 2 |
| 2022 | 3D Human Skeleton Estimation from Monocular Single RGB Image based on Multiple Virtual-View Skeleton Generationabstract3D human skeleton estimation from a single RGB image is one of the challenging problems in computer vision. Motivated by the advantage of the multi-views approach, we proposed a new two-stage approach. The 1ststage estimates a set of 3D heatmaps, by which 2D image coordinates and relative depth for each joint of a set of 2D skeletons can be derived. It consists of multi-streams, where 1ststream is to predict 2D+depth skeleton for the real-view and the other streams are to predict skeleton counterparts for virtual-views, thus called Multiple Virtual-View Skeleton Generator (MVSG) network. The 2ndstage contains: depth-denoising and fusion network, where the outputs of MVSG network are depth-denoised and then fused by concatenation for regression into the final 3D skeleton. Experiments show that our technique has achieved a performance of MPJPE=46.75 mm, which is comparable to the state-of-the-art methods. Wen-Nung Lie, Veasna Vann, Lee Aing, Jui-Chiu Chiang |
MMSP | 4 |
| 2022 | 3D Human Skeleton Estimation Based on RGB Image Sequence and Graph Convolution NetworkabstractWe propose a technique of 3D human skeleton estimation from RGB image sequence. Our method uses two-stages of deep learning networks. The first stage is to estimate enhanced 2D skeletons (2D image coordinates and relative-depths for all joints). The sequence of enhanced 2D skeletons is represented as a spatial-temporal graph (STG) which is then input to the second stage composed of Graph Convolutional Network (GCN) as the backbone. Techniques of high-order feature representations for joints, multi-stream feature adjustments, and denoising, were developed to further promote the accuracy performance for the estimated 3D skeletons. The Human3.6M dataset was used for training and testing. Experimental results show that our multi-stream GCN-based network can extract useful information from the input sequence efficiently. From experiments, the mean per joint position error (MPJPE) of the 3D skeletal joints is 47.27 mm when a sequence of 31 RGB frames are considered. Wen-Nung Lie, Pei-Hsuan Yang, Veasna Vann, Jui-Chiu Chiang |
MMSP | 4 |
| 2022 | Augmented Normalizing Flow for Point Cloud Geometry CodingabstractWith the increased popularity of immersive media, point clouds have become one of the popular data representations for presenting 3D scenes. The huge amount of point cloud data poses a great challenge on their storage and real-time transmission, which calls for efficient point cloud compression. This paper presents a novel point cloud geometry compression technique based on learning end-to-end an augmented normalizing flow (ANF) model to represent the occupancy status of voxelized data points. The higher expressive power of ANF than variational autoencoders (V AE) is leveraged for the first time to represent binary occupancy status. Compared to two coding standards developed by MPEG, namely G-PCC (geometry-based point cloud compression) and V-PCC (video-based point cloud compression), our method achieves more than 80% and 30% bitrate reduction, respectively. Compared to several learning-based methods, our method also yields better performance. Siao-Yu Li, Ji-Jin Chiu, Jui-Chiu Chiang, Wen-Hsiao Peng, Wen-Nung Lie |
VCIP | 3 |
| 2021 | High-Order Joint Information Input For Graph Convolutional Network Based Action RecognitionabstractGraph Convolution Network (GCN)-based networks for human action recognition, accepting 3D skeleton sequence as input, have gained much attention and good performances recently. In this paper, joint information enhanced with rich higher-order features/attributes is proposed to lift up their recognition performances. All joints in a spatio-temporal skeleton are described in terms of a set of 3-component vectors by referring up to 3 joint neighbors in the spatio-temporal domain. The referred joints are physically connected in spatial or corresponded in temporal domain. Our rich high-order joint information is fed as inputs to two kinds of GCN-based networks in two ways: early fusion and late fusion. Early fusion is to concatenate these 3-components vectors as different channels at input nodes and late fusion is to feed each 3-component vector to a multi-stream GCN network separately and then fuse the output from each stream for action recognition decision. We also propose to cascade a view-adaptive (VA) sub-network to further promote the performance. Experiments show that our approach is capable of boosting the accuracy of original GCN networks in both early or late fusion styles by up to 1.57% and 2.55%, respectively (in cross-subject (CS) protocol) when using NTU RGB-D 60 dataset for evaluations. Wen-Nung Lie, Yong-Jhu Huang, Jui-Chiu Chiang, Zhen-Yu Fang |
ICIP | 3 |
| 2021 | Action Prediction Using Extremely Low-Resolution Thermopile Sensor Array For Elderly MonitoringabstractAccident anticipation for monitoring the elderly is an important topic given the global issue of the rapidly aging population. In this work, we propose a novel approach to elderly accident prevention by using a low-resolution (32x32 pixels) infrared camera - thermopile sensor array (TPA) - for the action prediction task. Such a kind of sensor ensures that the monitoring system is cost-effective, does not interfere with daily life of the user, and most importantly fully preserves their privacy which makes it suitable for use in hospitals, nursing homes and private residences. As the majority of accidents involving the elderly occurs when a person attempts to exit the bed unassisted, we concentrate our efforts on predicting that an elderly person will attempt to get off the bed without asking for help. Our system raises an alarm in such a case and informs the caregiver so that they can intervene and offer assistance. Our designed deep-learning model can predict that the elderly patient monitored has the intention to get off the bed by 8.12 seconds at an accuracy of 96.51% (based on our own dataset collected), on average according to experiments, before the action onset is observed. Igor Morawski, Wen-Nung Lie, Jui-Chiu Chiang |
ICIP | 3 |
| 2021 | Faster and Finer Pose Estimation for Object Pool in a Single RGB ImageabstractPredicting/estimating the 6DoF pose parameters for multi-instance objects accurately in a fast manner is an important issue in robotic and computer vision. Even though some bottom-up methods have been proposed to be able to estimate multiple instance poses simultaneously, their accuracy cannot be considered as good enough when compared to other state-of-the-art top-down methods. Their processing speed still cannot respond to practical applications. In this paper, we present a faster and finer bottom-up approach of deep convolutional neural network to estimate poses of the object pool even multiple instances of the same object category present high occlusion/overlapping. Several techniques such as prediction of semantic segmentation map, multiple keypoint vector field, and 3D coordinate map, and diagonal graph clustering are proposed and combined to achieve the purpose. Experimental results and ablation studies show that the proposed system can achieve comparable accuracy at a speed of 24.7 frames per second for up to 7 objects by evaluation on the well-known Occlusion LINEMOD dataset. Lee Aing, Wen-Nung Lie, Jui-Chiu Chiang |
VCIP | 3 |
| 2021 | Saliency-driven rate-distortion optimization for 360-degree image coding
Jui-Chiu Chiang, Cheng-Yu Yang, Bhishma Dedhia, Yi-Fan Char |
Multim. Tools Appl. | 1 |
| 2021 | Rate-distortion optimization of multi-exposure image coding for high dynamic range image coding
Jui-Chiu Chiang, Wen-Hsien Shih, Jhih-You Deng |
Signal Process. Image Commun. | 1 |
| 2020 | Semi-automatic 2D-to-3D video conversion based on background sprite generation
Wen-Nung Lie, Shao-Ting Chiu, Yi-Kai Chen, Jui-Chiu Chiang |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Saliency Prediction for Omnidirectional Images Considering Optimization on Sphere DomainabstractThere are several formats to describe the omnidirectional images. Among them, equirectangular projection (ERP), represented as 2D image, is the most widely used format. There exist many outstanding methods capable of well predicting the saliency maps for the conventional 2D images. But these works cannot be directly extended to predict the saliency map of the ERP image, since the content on ERP is not for direct display. Instead, the viewport image on demand is generated after converting the ERP image to the sphere domain, followed by rectilinear projection. In this paper, we propose a model to predict the saliency maps of the ERP images using existing saliency predictors for the 2D image. Some pre-processing and post-processing are used to manage the problem mentioned above. In particular, a smoothing based optimization is realized on the sphere domain. A public dataset of omnidirectional images is used to perform all the experiments and competitive results are achieved. Bhishma Dedhia, Jui-Chiu Chiang, Yi-Fan Char |
ICASSP | 2 |
| 2019 | Fast intra mode decision and fast CU size decision for depth video coding in 3D-HEVC
Jui-Chiu Chiang, Kuan-Kai Peng, Chao-Chun Wu, Chih-You Deng, Wen-Nung Lie |
Signal Process. Image Commun. | 1 |
| 2018 | Perception-based High Dynamic Range Infrared Video CodingabstractInfrared imagery has been used in many applications. Since the dynamic range of the infrared camera is wider than that of the RGB camera, the infrared image is usually stored in high bit-depth format. This paper proposes a technique to encode the high bit-depth infrared video, where the perception for the high dynamic range content is taken into account. After realizing the modified rate-distortion optimization, the experimental results show that the proposed scheme can achieve up to 13% bitrate reduction while keeping comparable quality in terms of HDR-VDP-2, compared to the H.264/AVC FRExt scheme. Yan-Jhu Chen, Wen-Hsien Shih, Jui-Chiu Chiang, Wen-Nung Lie |
VCIP | 3 |
| 2018 | Dual-Layer Lossless Coding for Infrared VideoabstractConventional natural image/video is targeted for entertainment and some distortion is allowed for the decoded data. For infrared imagery, lossless coding is needed for some specified applications, where analysis over the content is always realized, such as military and medical applications. In this paper, we propose a bit-depth scalable coding scheme, where lossy compression is applied to the base layer while lossless compression is to the enhancement layer. Through the usage of deep residual learning, an improved inter-layer prediction is obtained. Experiment results show that the proposed scheme not only provides the flexibility by offering two kinds of content depending on the request, it also achieves up to 3.26% bitrate reduction compared to the single-layer lossless coding scheme. Hui-Shan Hsiao, Wen-Hsien Shih, Jui-Chiu Chiang, Wen-Nung Lie |
VCIP | 3 |
| 2016 | Low complexity depth intra coding combining fast intra mode and fast CU size decision in 3D-HEVCabstract3D-HEVC is the new coding standard dealing with both the texture and the associated depth video. In addition to some new coding tools designed for texture video with improved coding efficiency, some specified tools are devoted for depth video, such as depth modeling mode (DMM), segment-wise DC (SDC) mode and single depth intra mode. In this paper, we propose two techniques to speed up the encoding of depth video, including fast intra mode decision and fast CU size decision. For the fast intra mode decision, early termination is performed if the minimum rate-distortion (RD) cost of test candidate modes is smaller than the threshold computed from full mode search. For the fast CU size decision, smaller CU size will not be evaluated if the current CU presents some desired properties. The experimental results report that the proposed techniques achieve on average 37.6% time saving with 0.8% bitrate increase for the synthesized views under the all intra scenario. Kuan-Kai Peng, Jui-Chiu Chiang, Wen-Nung Lie |
ICIP | 2 |
| 2013 | Video retargeting based frame-compatible stereo video codingabstractDue to more and more request of visual entertainment with improved perceptual realism, stereo video offering enhanced realism and interactivity is highly desired. Usually, the stereo video is packed into a single video and each view preserves only half resolution to achieve an efficient delivery. In this paper, a novel frame-compatible stereo video coding technique based on content-aware video retargeting is presented. Despite of combing the stereo video into a single channel by uniform downsampling, we propose to sub-sample the stereo video in an unequal way, by taking the saliency map of the video into considering. During the reconstruction of the stereo video with original resolution, the regions with higher saliency attention are less distorted, relying on a smaller downsampling factor, instead of 2. The experimental results reveal that the proposed technique achieves up to 43% bitrate saving as compared to the most popular frame-compatible stereo formats. Siao-Wei Chen, Ming-Feng Tsai, Jui-Chiu Chiang |
ICASSP | 3 |
| 2013 | Intelligent exposure determination for high quality HDR image generationabstractHigh dynamic range (HDR) image generation had been studied for years. Due to the expense and rareness of HDR cameras, many works try to generate HDR images using several low dynamic range (LDR) images with different exposure conditions. To ensure high-quality HDR image generation, details of the scene should be retained in different LDR images and the exposure parameters of LDR image should be chosen carefully. In this paper, we propose a technique to dynamically determine the suitable exposure parameter for LDR images according to the property of each scene to be captured. Experimental results reveal that better HDR images are always generated with the LDR images captured by the determined exposure, compared to those created by the method with fixed exposures. Kun-Fang Huang, Jui-Chiu Chiang |
ICIP | 2 |
| 2012 | Multiview texture coding and free viewpoint image synthesis for mesh-based 3D video transmissionabstractIn this paper, the “3D mesh model plus texture” strategy is adopted to implement a 3D video system. First, the acquired multi-view videos are processed to construct dynamic 3D mesh models about the foreground subject. These 3D mesh models, together with the acquired texture information, are then compressed and transmitted via networks to receivers. Main contributions of our work lies on the proposals of data reduction and pre-processing of multi-view texture information before H.264/AVC encoding at transmitter and a robust occlusion test on synthesizing novel views at receiver. Our proposed data reduction and pre-processing schemes are capable of removing redundant texture information, while maintaining inter-frame correlation, to result in high coding efficiency. Experiment results show that the proposed occlusion test is capable of eliminating texture-rendering artifacts due to 3D model reconstruction errors, thus improving the viewing quality of 3D video at receiver. Besides, our texture encoder achieves a saving of 40% ~ 57% in transmission bit rate, compared with the traditional approach. Jui-Chiu Chiang, Ping-He Hou, Kai-Che Liu, Wen-Nung Lie |
ISCAS | 1 |
| 2012 | Efficient improvement of side information in GOB-based DVC systemabstractAmong the emerging video coding schemes, the effectual solution for the separate-encoding and joint-decoding architecture is the distributed source coding (DSC). In the DSC-based video system, the side information is available in the Wyner-Ziv (WZ) decoder and video reconstruction. Theoretically, the quality of side information (SI), usually referring to the difference between SI and source, dominates the coding efficiency. In order to improve the coding efficiency of DVC with temporal group of blocks, we proposed three SI generations after analyzing the correlation between the SI and original information. In contrast to the traditional improvement in the DVC decoder, the complexity of SI generator was dramatically reduced in our schemes. Experimental results show that the best method can reduce 0.7% bit-rate with less computation and no codec delay compared to the bi-linear interpolation (BLI) method. Tsung-Che Wu, Ji-Hua Hsu, Chang-Ming Lee, Jui-Chiu Chiang |
ISCAS | 4 |
| 2011 | Coding of Dynamic 3D Mesh Model for 3D Video Transmission
Jui-Chiu Chiang, Chun-Hung Chen, Wen-Nung Lie |
PSIVT (1) | 1 |
| 2010 | Region-of-interest based rate control scheme with flexible quality on demandabstractConventional rate control schemes focus on making output bit rate approach a target value and are deficient in ensuring a higher quality of ROI (Region of Interest) than others in a frame. In this paper, we propose a new scheme for H.264/AVC, aiming to allocate more bit resource for the encoding of ROI and still maintain the accuracy of the output bit rate. ROI-based rate control algorithms can find their specific advantages in video telephony and video surveillance applications. Our proposed scheme is based on the one implemented in H.264/AVC JM software, but enhanced with several features: ROI determination with saliency map, tunable quality factor for ROI, two-channel (ROI & non-ROI) rate control, and QP adjustment with constraints from temporal and spatial domains, as well as from ROI/non-ROI adjacency boundaries. Experiment results show both advantages in objective PSNR and subjective evaluations for ROI, while making the output bit rate accurate as before. Jui-Chiu Chiang, Cheng-Sheng Hsieh, Fan-Di Jou, Wen-Nung Lie |
ICME | 1 |
| 2010 | Block-based distributed video coding with variable block modesabstractIn this paper, a new block-based pixel domain distributed video coding scheme featured with variable block modes is proposed. In addition to intra mode and Wyner-Ziv mode employed in conventional block-based distributed video coding scheme, two supplementary block modes “SKIP mode” and “zero motion mode” are introduced in the proposed scheme to improve the overall coding efficiency, as well as to reduce the decoding complexity. Moreover, the channel coding is performed on macroblcok level to reduce the coding loss due to inserted information in the parity bits. The simulation results show that the proposed scheme outperforms both the conventional frame-based transform-domain and the block-based pixel-domain distributed coding schemes. Jui-Chiu Chiang, Kuan-Liang Chen, Chi-Ju Chou, Chang-Ming Lee, Wen-Nung Lie |
ISCAS | 1 |
| 2009 | Bit-depth scalable video coding using inter-layer prediction from high bit-depth layerabstractScalable video coding (SVC) is currently developed as an extension of H.264/AVC video coding standard. In this paper, we propose three H.264/AVC compliant bit-depth scalable video coding schemes, named LH mode (Low Bit-depth to High Bit-depth), HL mode (High Bit-depth to Low Bit-depth) and combined LH-HL mode for different applications. All these schemes efficiently exploit the high correlation between the high bit-depth layer and the low bit-depth layer on Macroblock level. Experimental results indicate that the HL mode outperforms the other two schemes and it achieves up to 7 dB improvement over the simulcast where the high bit-depth video and low bit-depth representations are 12-bit and 8-bit, respectively. Jui-Chiu Chiang, Wen-Ting Kuo |
ICASSP | 1 |
| 2009 | Enhanced Side Information Generator with Accurate Evaluations in Block-Based Wyner-Ziv Video Coding
Chang-Ming Lee, Jui-Chiu Chiang, Zhi-Heng Chiang, Kuan-Liang Chen, Wen-Nung Lie |
PSIVT | 2 |
| 2006 | Robust Video Transmission Over Mixed IP - Wireless Channels using Motion-Compensated Oversampled FilterbanksabstractRobust video coding has attracted increasing attention during the past few years. This paper proposes a joint source channel coding scheme able to resist transmission errors over mixed Internet-wireless channels. It involves motion-compensated oversampled filterbanks (OFBs). The redundancy introduced by the overcomplete representation in signals at the output of OFBs is employed for error correction of the motion compensated frames. The errors may be due to the wireless part of the channel (random noise), but also to the Internet part (packet losses). The performance of the proposed approach is illustrated for compressed streams transmitted through a packet erasure channel with an averaged packet loss of 6.25% followed by a binary symmetric channel with a crossover probability of 10-2 Jui-Chiu Chiang, Chang-Ming Lee, Michel Kieffer, Pierre Duhamel |
ICASSP (2) | 1 |
| 2005 | Motion-compensated oversampled filterbanks for robust video codingabstractDuring the past few years, much effort has been devoted for efficient coding of video. However, this high coding efficiency results in an increased sensitivity to transmission errors. This is the motivation for developing techniques allowing us altogether to transmit video sequences efficiently and to be robust to the transmission errors. This paper proposes a joint source-channel coding scheme using motion-compensated oversampled filterbanks (OFB). Previous work for robust transmission using OFB have concentrated on still images. Here, OFB are combined with motion compensation techniques. The redundancy introduced by the overcomplete representation is employed for error correction of the motion compensated frames. Experiments have shown that performance is acceptable even when the compressed stream is transmitted through a binary symmetric channel with crossover probability of 10/sup -2/. Jui-Chiu Chiang, Michel Kieffer, Pierre Duhamel |
ICASSP (5) | 1 |