Wen-Nung Lie

dblp:07/5659 · DBLP profile ↗
← Back
82ranked-venue papers
48as first author
12since 2021 · last 2026
0000-0002-8166-2844ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 38 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 2 since 2021Systems, architecture and hardware · 7 · 4 first-author · 1 since 2021Security and privacy · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Efficient 6DoF pose estimation for multi-instance objects from a single image
Wen-Nung Lie, Lee Aing
Image Vis. Comput.1
2024 Deep-learning-based head pose estimation from a single RGB image and its application to medical CROM measurement
abstract
Abstract For human beings, neck movement will be degraded due to aging, trauma, musculoskeletal disorders, or degenerative diseases. Cervical range of motion (CROM) measurement is one of the popular quantitative neck examinations. Despite radiography is considered as the gold standard, it suffers from invasiveness, radiation exposure, and expensiveness. Recently, vision-based methods have been applied for CROM measurement but achieve large errors and require depth camera. On the other hand, deep neural networks provide good performances on head pose estimation (HPE) from a single image, thus promising for medical CROM measurement. We propose to use CNN networks to extract pyramidal or multi-level image features, which are passed to cross-level attention modules for feature fusion and then to a modified ASPP module and a multi-bin classification/regression module for spatial-channel attention and Euler angle conversion/prediction, respectively. The proposed technique was evaluated on public datasets, such as 300W_LP, AFLW2000, and BIWI, to verify its superior performances (with mean MAE = 3.50°, 3.40°, and 2.31° for different experimental protocols) than state-of-the-art methods. Our pre-trained model was also evaluated with our own collected dataset from hospital for CROM measurement. It also achieved the lowest MAE of 4.58° among other methods and conformed with a medical standard of 5 degrees except the pitch angle (which has a MAE of 5.70°, larger than the standard and the yaw (MAE = 3.60°) and roll angles (MAE = 4.44°)). In general, HPE technique is feasible for CROM measurement and shows its advantages of speed, non-invasiveness, free of anatomical landmark and low cost of operation.
Panrasee Ritthipravat, Kittisak Chotikkakamthorn, Wen-Nung Lie, Worapan Kusakunniran, Pimchanok Tuakta, Paitoon Benjapornlert
Multim. Tools Appl.3
2023 Faster and finer pose estimation for multiple instance objects in a single RGB image
Lee Aing, Wen-Nung Lie, Guo-Shiang Lin
Image Vis. Comput.2
2023 Deep-Learning Technique for Risk-Based Action Prediction Using Extremely Low-Resolution Thermopile Sensor Array
abstract
Eldering caring is important in today’s aging society, especially that accident anticipation/prevention plays an important role. In this paper, a novel approach to preventing elderly accidents based on a very low-resolution thermopile sensor array (TPA) (only$32\times 32$pixels) is proposed for prediction of bed-exit event that might lead to elderly falls in home caring. Low-resolution TPA sensor, capable of collecting far infrared energy, ensures cost-effective monitoring, no interference with user’s daily life, and most importantly privacy-preservation. Since most of the fall accidents occur when the elderly attempts to get off the bed without assistance, it is thus the focus of this paper to monitor his/her posture and action via TPA image sensor and then predict that an action of getting off the bed will occur in a near future (e.g.,$S$seconds later). Our system can raise an alarm to the caregivers so that they can intervene and offer the necessary assistance. A deep-learning model based on CNN-RNN (Convolutional neural network-Recurrent neural network) architecture was designed which is capable of predicting the elderly bed-exit intention by$S =5.78$seconds in advance of the action onset at an accuracy of 99.37% according to our dataset evaluation. Our system is also suitable for on-line real-time operation which will be helpful to elderly caring in our society.
Igor Morawski, Wen-Nung Lie, Lee Aing, Jui-Chiu Chiang, Kuan-Ting Chen
IEEE Trans. Circuits Syst. Video Technol.2
2022 3D Head Pose Estimation Based on Graph Convolutional Network from A Single RGB Image
abstract
Research of head pose estimation in computer vision has been at the center of much attention. This work presents a framework based on adaptive graph convolution network (AGCN) to process both 2D and 3D facial landmarks extracted from the input RGB image. The network has a two-streams (teacher/3D-student/2D streams) architecture, trained with a 3D to 2D knowledge distillation training process, to transfer features of the 3D stream to the 2D stream for performance promotion. Several processing modules, such as depth-denoising for detected 3D landmarks, multi-stream fusion in inference, were also proposed for further increase of the prediction performance and robustness of our proposed method. In experiments, we follow standard protocols (in terms of datasets and metrices) to evaluate our performance. Three datasets 300W-LP, AFLW2000 and BIWI were used. The performance is measured in mean absolute error (MAE). We can achieve better performance compared to most of the state-of-the-art methods.
Wen-Nung Lie, Monyneath Yim, Lee Aing, Jui-Chiu Chiang
ICIP1
2022 Remote PPG Estimation from RGB-NIR Facial Image Sequence for Heart Rate Estimation
abstract
This paper presents a dual-modal (RGB-NIR) technique to estimate remote photoplethysmogram (rPPG) signal, i.e. the heart rate, from a facial image sequence. We developed denoising techniques with a modified amplitude selective filtering (ASF), wavelet decomposition and robust principal component analysis (RPCA), to enhance the uncovering of the rPPG signal through the well-known ICA algorithm. A new dataset built with RealSense RGB-D camera is considered in experiments: regular brightness, under-illumination, and face motion. Experimental results show that the proposed method has reached competitive performance among the state-of-the-art methods in motion and under-illuminated scenarios even at a shorter input video length (10 to 20 seconds).
Dao-Quang Le, Jui-Chiu Chiang, Wen-Nung Lie
ISCAS3
2022 3D Human Skeleton Estimation from Monocular Single RGB Image based on Multiple Virtual-View Skeleton Generation
abstract
3D human skeleton estimation from a single RGB image is one of the challenging problems in computer vision. Motivated by the advantage of the multi-views approach, we proposed a new two-stage approach. The 1ststage estimates a set of 3D heatmaps, by which 2D image coordinates and relative depth for each joint of a set of 2D skeletons can be derived. It consists of multi-streams, where 1ststream is to predict 2D+depth skeleton for the real-view and the other streams are to predict skeleton counterparts for virtual-views, thus called Multiple Virtual-View Skeleton Generator (MVSG) network. The 2ndstage contains: depth-denoising and fusion network, where the outputs of MVSG network are depth-denoised and then fused by concatenation for regression into the final 3D skeleton. Experiments show that our technique has achieved a performance of MPJPE=46.75 mm, which is comparable to the state-of-the-art methods.
Wen-Nung Lie, Veasna Vann, Lee Aing, Jui-Chiu Chiang
MMSP1
2022 3D Human Skeleton Estimation Based on RGB Image Sequence and Graph Convolution Network
abstract
We propose a technique of 3D human skeleton estimation from RGB image sequence. Our method uses two-stages of deep learning networks. The first stage is to estimate enhanced 2D skeletons (2D image coordinates and relative-depths for all joints). The sequence of enhanced 2D skeletons is represented as a spatial-temporal graph (STG) which is then input to the second stage composed of Graph Convolutional Network (GCN) as the backbone. Techniques of high-order feature representations for joints, multi-stream feature adjustments, and denoising, were developed to further promote the accuracy performance for the estimated 3D skeletons. The Human3.6M dataset was used for training and testing. Experimental results show that our multi-stream GCN-based network can extract useful information from the input sequence efficiently. From experiments, the mean per joint position error (MPJPE) of the 3D skeletal joints is 47.27 mm when a sequence of 31 RGB frames are considered.
Wen-Nung Lie, Pei-Hsuan Yang, Veasna Vann, Jui-Chiu Chiang
MMSP1
2022 Augmented Normalizing Flow for Point Cloud Geometry Coding
abstract
With the increased popularity of immersive media, point clouds have become one of the popular data representations for presenting 3D scenes. The huge amount of point cloud data poses a great challenge on their storage and real-time transmission, which calls for efficient point cloud compression. This paper presents a novel point cloud geometry compression technique based on learning end-to-end an augmented normalizing flow (ANF) model to represent the occupancy status of voxelized data points. The higher expressive power of ANF than variational autoencoders (V AE) is leveraged for the first time to represent binary occupancy status. Compared to two coding standards developed by MPEG, namely G-PCC (geometry-based point cloud compression) and V-PCC (video-based point cloud compression), our method achieves more than 80% and 30% bitrate reduction, respectively. Compared to several learning-based methods, our method also yields better performance.
Siao-Yu Li, Ji-Jin Chiu, Jui-Chiu Chiang, Wen-Hsiao Peng, Wen-Nung Lie
VCIP5
2021 High-Order Joint Information Input For Graph Convolutional Network Based Action Recognition
abstract
Graph Convolution Network (GCN)-based networks for human action recognition, accepting 3D skeleton sequence as input, have gained much attention and good performances recently. In this paper, joint information enhanced with rich higher-order features/attributes is proposed to lift up their recognition performances. All joints in a spatio-temporal skeleton are described in terms of a set of 3-component vectors by referring up to 3 joint neighbors in the spatio-temporal domain. The referred joints are physically connected in spatial or corresponded in temporal domain. Our rich high-order joint information is fed as inputs to two kinds of GCN-based networks in two ways: early fusion and late fusion. Early fusion is to concatenate these 3-components vectors as different channels at input nodes and late fusion is to feed each 3-component vector to a multi-stream GCN network separately and then fuse the output from each stream for action recognition decision. We also propose to cascade a view-adaptive (VA) sub-network to further promote the performance. Experiments show that our approach is capable of boosting the accuracy of original GCN networks in both early or late fusion styles by up to 1.57% and 2.55%, respectively (in cross-subject (CS) protocol) when using NTU RGB-D 60 dataset for evaluations.
Wen-Nung Lie, Yong-Jhu Huang, Jui-Chiu Chiang, Zhen-Yu Fang
ICIP1
2021 Action Prediction Using Extremely Low-Resolution Thermopile Sensor Array For Elderly Monitoring
abstract
Accident anticipation for monitoring the elderly is an important topic given the global issue of the rapidly aging population. In this work, we propose a novel approach to elderly accident prevention by using a low-resolution (32x32 pixels) infrared camera - thermopile sensor array (TPA) - for the action prediction task. Such a kind of sensor ensures that the monitoring system is cost-effective, does not interfere with daily life of the user, and most importantly fully preserves their privacy which makes it suitable for use in hospitals, nursing homes and private residences. As the majority of accidents involving the elderly occurs when a person attempts to exit the bed unassisted, we concentrate our efforts on predicting that an elderly person will attempt to get off the bed without asking for help. Our system raises an alarm in such a case and informs the caregiver so that they can intervene and offer assistance. Our designed deep-learning model can predict that the elderly patient monitored has the intention to get off the bed by 8.12 seconds at an accuracy of 96.51% (based on our own dataset collected), on average according to experiments, before the action onset is observed.
Igor Morawski, Wen-Nung Lie, Jui-Chiu Chiang
ICIP2
2021 Faster and Finer Pose Estimation for Object Pool in a Single RGB Image
abstract
Predicting/estimating the 6DoF pose parameters for multi-instance objects accurately in a fast manner is an important issue in robotic and computer vision. Even though some bottom-up methods have been proposed to be able to estimate multiple instance poses simultaneously, their accuracy cannot be considered as good enough when compared to other state-of-the-art top-down methods. Their processing speed still cannot respond to practical applications. In this paper, we present a faster and finer bottom-up approach of deep convolutional neural network to estimate poses of the object pool even multiple instances of the same object category present high occlusion/overlapping. Several techniques such as prediction of semantic segmentation map, multiple keypoint vector field, and 3D coordinate map, and diagonal graph clustering are proposed and combined to achieve the purpose. Experimental results and ablation studies show that the proposed system can achieve comparable accuracy at a speed of 24.7 frames per second for up to 7 objects by evaluation on the well-known Occlusion LINEMOD dataset.
Lee Aing, Wen-Nung Lie, Jui-Chiu Chiang
VCIP2
2020 Human Behavior Recognition from Multiview Videos
Yu-Ling Hsueh, Wen-Nung Lie, Guan-You Guo
Inf. Sci.2
2020 Semi-automatic 2D-to-3D video conversion based on background sprite generation
Wen-Nung Lie, Shao-Ting Chiu, Yi-Kai Chen, Jui-Chiu Chiang
J. Vis. Commun. Image Represent.1
2019 Fast intra mode decision and fast CU size decision for depth video coding in 3D-HEVC
Jui-Chiu Chiang, Kuan-Kai Peng, Chao-Chun Wu, Chih-You Deng, Wen-Nung Lie
Signal Process. Image Commun.5
2018 Perception-based High Dynamic Range Infrared Video Coding
abstract
Infrared imagery has been used in many applications. Since the dynamic range of the infrared camera is wider than that of the RGB camera, the infrared image is usually stored in high bit-depth format. This paper proposes a technique to encode the high bit-depth infrared video, where the perception for the high dynamic range content is taken into account. After realizing the modified rate-distortion optimization, the experimental results show that the proposed scheme can achieve up to 13% bitrate reduction while keeping comparable quality in terms of HDR-VDP-2, compared to the H.264/AVC FRExt scheme.
Yan-Jhu Chen, Wen-Hsien Shih, Jui-Chiu Chiang, Wen-Nung Lie
VCIP4
2018 Dual-Layer Lossless Coding for Infrared Video
abstract
Conventional natural image/video is targeted for entertainment and some distortion is allowed for the decoded data. For infrared imagery, lossless coding is needed for some specified applications, where analysis over the content is always realized, such as military and medical applications. In this paper, we propose a bit-depth scalable coding scheme, where lossy compression is applied to the base layer while lossless compression is to the enhancement layer. Through the usage of deep residual learning, an improved inter-layer prediction is obtained. Experiment results show that the proposed scheme not only provides the flexibility by offering two kinds of content depending on the request, it also achieves up to 3.26% bitrate reduction compared to the single-layer lossless coding scheme.
Hui-Shan Hsiao, Wen-Hsien Shih, Jui-Chiu Chiang, Wen-Nung Lie
VCIP4
2018 Key-Frame-Based Background Sprite Generation for Hole Filling in Depth Image-Based Rendering
abstract
In this paper, we propose a new depth image-based rending scheme for 3DTV applications, where a background sprite model is utilized for dis-occlusion/hole filling purpose. Dissimilar to traditional spatial (e.g., interpolation or inpainting) and temporal methods, our algorithm is capable of recovering the holes with true background information by incrementally integrating the spatial and temporal information of the video in a unified background sprite model. The technique of background sprite model construction in this paper is featured of resolving camera motions existing in most of the consumer videos, accurate registration of multiframe information, and efficient memory use and computation in realistic 3DTV application. To register/stitch each input frame to the background sprite model accurately, foreground removal considering color and depth information, and an adaptive key-frame-based scheme are developed for transform computation. Experimental results show that our proposed scheme has a large temporal reference distance and can retrieve true background information accurately, thus leading to better quality after novel view synthesis compared to existing spatial or spatio-temporal algorithms, especially for videos with significant camera motions or complex backgrounds.
Wen-Nung Lie, Chia-Yung Hsieh, Guo-Shiang Lin
IEEE Trans. Multim.1
2017 Using Sparse-Point Disparity Estimation and Spatial Propagation to Construct Dense Disparity Map for Stereo Endoscopic Images
Wen-Nung Lie, Hsi-Hung Huang, Shih-Wei Huang, Kai-Che Liu
PSIVT1
2017 Key-frame-based depth propagation for semi-automatic stereoscopic video conversion
Guo-Shiang Lin, Jian-Fa Huang, Wen-Nung Lie
J. Vis. Commun. Image Represent.3
2016 Low complexity depth intra coding combining fast intra mode and fast CU size decision in 3D-HEVC
abstract
3D-HEVC is the new coding standard dealing with both the texture and the associated depth video. In addition to some new coding tools designed for texture video with improved coding efficiency, some specified tools are devoted for depth video, such as depth modeling mode (DMM), segment-wise DC (SDC) mode and single depth intra mode. In this paper, we propose two techniques to speed up the encoding of depth video, including fast intra mode decision and fast CU size decision. For the fast intra mode decision, early termination is performed if the minimum rate-distortion (RD) cost of test candidate modes is smaller than the threshold computed from full mode search. For the fast CU size decision, smaller CU size will not be evaluated if the current CU presents some desired properties. The experimental results report that the proposed techniques achieve on average 37.6% time saving with 0.8% bitrate increase for the synthesized views under the all intra scenario.
Kuan-Kai Peng, Jui-Chiu Chiang, Wen-Nung Lie
ICIP3
2015 Fast encoding of 3D color-plus-depth video based on 3D-HEVC
abstract
3D-HEVC is the newest standard for compressing the Multi-View plus Depth (MVD) video. Inheriting from HEVC, 3D-HEVC presents a high encoding complexity by extra considering interview prediction. Under the encoding architecture of 3D-HEVC, we develop fast algorithms for early decisions of CU splitting/non-splitting for both texture and depth frame coding. High correlation between the texture and depth domains is exploited such that when encoding one domain of information (texture or depth), the coding efficiency (RD characteristics or speedup) can be improved by adding suitably augmenting information from the other domain, thus achieving the so-called depth-assisted and texture-assisted coding. In texture coding part, optical flow features and depth edges are combined to form a feature vector for searching similar CUs in previously coded block buffers and inheriting the CU splitting decision accordingly. In depth coding part, optical flows and depth map features are used as inputs to a neural classifier for fast decision of CU splitting/non-splitting. Compared to the original 3D-HEVC implementation, our texture coding algorithm achieves a 46.6% of time saving at only 0.4% of bit rate increase and 0.04 dB of quality degradation. On the other side, our depth coding algorithm saves 35.8% of the encoding time at 2.65% of bit rate reduction and 0.16 dB of PSNR degradation. In comparison to prior works [4][5][10], our algorithm achieves more time saving at comparable bit rate increase and PSNR degradation.
Wen-Nung Lie, Yan-Heng Lu
ICIP1
2015 All-Focus Image Fusion and Depth Image Estimation Based on Iterative Splitting Technique for Multi-focus Images
Wen-Nung Lie, Chia-Che Ho
PSIVT1
2015 Error concealment for the transmission of H.264/AVC-compressed 3D video in color plus depth format
Wen-Nung Lie, Guan-Hua Lin
J. Vis. Commun. Image Represent.1
2014 Rate control technique based on 3D quality optimization for 3D video encoding
abstract
This paper presents a new 3D video (in color plus depth format) encoding system, featured of 3D quality optimization and joint rate control between color and depth components. An SVR-based prediction model is pre-built for estimating the best bit rate allocation between color and depth for each frame by analyzing edge features in the color and the depth component images. We also modify rate control scheme in H.264/SVC JSVM reference software to be suitable for joint rate control of the color and depth sequences. Accordingly, bit rates can be dynamically and accurately allocated between color and depth components to achieve optimized 3D video quality. From the experiment results, it is shown that our proposed method is capable of achieving more accuracy in rate control and better 3D quality for visual perception.
Wen-Nung Lie, Yu-Peng Liao
ICIP1
2014 Sprite generation for hole filling in depth image-based rendering
abstract
In this paper, we propose a new depth image-based rending (DIBR) scheme for 3DTV applications, which is based on temporal hole filling with sprite generation. Dissimilar to traditional methods (e.g., spatial interpolation and inpainting), temporal information is utilized to recover the holes with true background information. The proposed scheme is composed of two parts: sprite generation and virtual view synthesis. To collect sufficiently temporal information, a key-frame-based sprite generation algorithm was developed to incrementally and accurately fuse the background color and depth information from successive frames. Then the holes in the synthesized virtual view can be recovered by referring to the constructed sprite models. Experiments demonstrate that our proposed scheme is capable of achieving good results and outperforms some existing methods subjectively and objectively.
Guo-Shiang Lin, Chia-Yung Hsieh, Wen-Nung Lie
ICIP3
2014 Motion Vector Recovery for Video Error Concealment by Using Iterative Dynamic-Programming Optimization
abstract
This paper proposes an error concealment technique for video transmission, focusing on motion vector (MV) recovery for both inter- and intra-coded frames, to improve video quality at decoder when video bit stream data incur transmission errors. The proposed algorithm considers slice (i.e., a row of macroblocks (MBs)) errors and uses DP (Dynamic Programming) optimization technique to estimate the lost MVs in a global manner, differing from the traditional Boundary Matching Algorithm (BMA) and others that recover MVs independently for individual MBs in an erroneous slice. We also propose an iterative DP process based on 8 × 8 pixels blocks to resolve finer motions (for 8 × 8, 8 × 16, and 16 × 8 pixels blocks) that will aid in the enhancement of reconstruction quality. Experiment results show that our algorithm outperforms the well-known BMA by up to 7.28 dB and the DMVE and another prior work by Qian by up to 1.0 dB at a packet loss rate of 15%. Subjective evaluation shows that our algorithm is especially promising in preserving line/curve features and motion details.
Wen-Nung Lie, Chang-Ming Lee, Chung-Hua Yeh, Zhi-Wei Gao 0001
IEEE Trans. Multim.1
2013 Semi-automatic 2D-to-3D video conversion based on depth propagation from key-frames
abstract
In this paper, we propose a new two-pass bi-directional keyframe depth propagation algorithm for semi-automatic 2D-to-3D video conversion. First, key-frames are selected from each video shot based on color compensation errors. Depths for key-frames are then manually assigned or rendered by using popular computer tools, which are then propagated to non-key-frames bounded by a pair of front and rear key-frames. Our two-pass bidirectional procedure is advantageous in solving the background occlusion/dis-occlusion problems that degrade traditional depth propagation Experiments show that our algorithm is capable of achieving better results in both subjective and objective evaluations.
Guo-Shiang Lin, Jian-Fa Huang, Wen-Nung Lie
ICIP3
2013 Error concealment for 3D video transmission
abstract
In this paper, an error concealment method for transmission of 3D video in color+depth dual stream format is proposed. Exploiting the correlation between 2D color video and its corresponding depth information, recovery of the lost information of one kind (color/depth) can be achieved with the aid of received information from the other kind (depth/color), thus getting better concealed results than prior works. Assuming separate encoding of the color and depth sequences and frame losses, our proposed method has a PSNR gain of up to 0.68 dB in color and a PSNR gain of up to 2.42 dB in depth. For subjective visual quality, our algorithm is especially effective in retaining object boundaries, which is also helpful in stereo or multi-view synthesis for 3D video display.
Wen-Nung Lie, Guan-Hua Lin
ISCAS1
2013 Quality enhancement based on retinex and pseudo-HDR synthesis algorithms for endoscopic images
abstract
In this paper, we present a quality enhancement scheme for endoscopic images. Traditional algorithms might be able to enhance the image contrast, but possible over-enhancement also lead to bad overall visual quality which prevents surgeons from accurate examination or operations of instruments in Minimal Invasive Surgery (MIS). Our proposed scheme integrates the well-known retinex algorithm with a pseudo-HDR (High Dynamic Range) synthesis process, designed to compose of three parts: multiscale retinex with gamma correction (MSR-G), local brightness range expansion (brightness diversity), and bilateral-filter-based HDR image fusion. Experiment results demonstrate that the proposed scheme is able to enhance image details and keep the overall visual quality good as well, with respect to other existing methods.
Jungle Chi-Hsiang Wu, Guo-Shiang Lin, Hsiao-Ting Hsu, You-Peng Liao, Kai-Che Liu, Wen-Nung Lie
VCIP6
2012 Super-resolution reconstruction of video sequences based on wavelet-domain spatial and temporal processing
Chang-Ming Lee, Chien-Jung Lee, Chia-Yung Hsieh, Wen-Nung Lie
ICPR4
2012 Adaptive support-window approximation to bilateral filtering
Guo-Shiang Lin, Chun-Ting Kuo, Wen-Nung Lie, Kai-Che Liu
ICPR4
2012 Multiview texture coding and free viewpoint image synthesis for mesh-based 3D video transmission
abstract
In this paper, the “3D mesh model plus texture” strategy is adopted to implement a 3D video system. First, the acquired multi-view videos are processed to construct dynamic 3D mesh models about the foreground subject. These 3D mesh models, together with the acquired texture information, are then compressed and transmitted via networks to receivers. Main contributions of our work lies on the proposals of data reduction and pre-processing of multi-view texture information before H.264/AVC encoding at transmitter and a robust occlusion test on synthesizing novel views at receiver. Our proposed data reduction and pre-processing schemes are capable of removing redundant texture information, while maintaining inter-frame correlation, to result in high coding efficiency. Experiment results show that the proposed occlusion test is capable of eliminating texture-rendering artifacts due to 3D model reconstruction errors, thus improving the viewing quality of 3D video at receiver. Besides, our texture encoder achieves a saving of 40% ~ 57% in transmission bit rate, compared with the traditional approach.
Jui-Chiu Chiang, Ping-He Hou, Kai-Che Liu, Wen-Nung Lie
ISCAS4
2012 3D human pose tracking based on depth camera and dynamic programming optimization
abstract
Depth camera has gained much more attention in applications of human computer interactivity. In this paper, we present an approach to tracking 3D human pose by constructing/estimating an articulated upper body joint model from depth data captured by Kinect camera. The system first calibrates the human pose parameters by processing the depth data of an initial pose. After that, the dynamic programming (DP) approach is performed for the tracking of joints, intending to optimize their observations in depth images and conformation to some physical constraints. Experiments show satisfactory tracking results. A performance comparison with the PrimeSense NITE program is also given. The current implementation presents a processing speed of over 30 frames per second.
Wen-Nung Lie, Hung-Wei Shiu, Chieh Huang
ISCAS1
2011 Coding of Dynamic 3D Mesh Model for 3D Video Transmission
Jui-Chiu Chiang, Chun-Hung Chen, Wen-Nung Lie
PSIVT (1)3
2011 2D to 3D Image Conversion Based on Classification of Background Depth Profiles
Guo-Shiang Lin, Han-Wen Liu, Wei-Chih Chen, Wen-Nung Lie, Sheng-Yen Huang
PSIVT (2)4
2010 Video error concealment by using iterative dynamic-programming optimization
abstract
This paper addresses an error concealment technique, focusing on motion vector (MV) recovery for inter-coded frames, to improve video quality at decoder when video bit stream data are incurred transmission errors. The proposed algorithm considers slice (i.e., a row of macroblocks (MBs)) errors and uses DP (Dynamic Programming) optimization technique to estimate the lost MVs in a combined manner, differing from the traditional Boundary Matching Algorithm (BMA) which recovers MV independently for each erroneous MB. We also consider MV recovery for blocks of 8×8, 8×16, and 16×8 pixels (rather than 16×16 pixels MBs only) and apply an iterative DP process to refine the estimated results. Experimental results show that our algorithm outperforms the well-known BMA by up to 1.36 dB in PSNR and a newly published algorithm by up to 0.56 dB at a pack loss rate of up to 15%. Subjective evaluation shows that our algorithm takes the advantages of moderate speed and accurate recovery of contiguous line/curve features in images.
Wen-Nung Lie, Chung-Hua Yeh, Zhi-Wei Gao 0001
ICIP1
2010 Region-of-interest based rate control scheme with flexible quality on demand
abstract
Conventional rate control schemes focus on making output bit rate approach a target value and are deficient in ensuring a higher quality of ROI (Region of Interest) than others in a frame. In this paper, we propose a new scheme for H.264/AVC, aiming to allocate more bit resource for the encoding of ROI and still maintain the accuracy of the output bit rate. ROI-based rate control algorithms can find their specific advantages in video telephony and video surveillance applications. Our proposed scheme is based on the one implemented in H.264/AVC JM software, but enhanced with several features: ROI determination with saliency map, tunable quality factor for ROI, two-channel (ROI & non-ROI) rate control, and QP adjustment with constraints from temporal and spatial domains, as well as from ROI/non-ROI adjacency boundaries. Experiment results show both advantages in objective PSNR and subjective evaluations for ROI, while making the output bit rate accurate as before.
Jui-Chiu Chiang, Cheng-Sheng Hsieh, Fan-Di Jou, Wen-Nung Lie
ICME5
2010 A 2D to 3D conversion scheme based on depth cues analysis for MPEG videos
abstract
In this article, we propose a 2D to 3D video conversion scheme for MPEG videos. The difficulty for 2D/3D conversion problem lies on depth estimation/ assignment with insufficient information. Our depth assignment is based on the analyses of multiple cues, e.g., motion parallax, atmospheric perspective, texture gradient, linear perspective, and relative height, separately for the foreground objects and the background area. To fit more kinds of videos, the proposed depth assignment scheme is content-adaptive by segmenting a video into shots and classifying each of them into three categories for different conversion schemes. Subjective experiments show that the 3D stereo video generated by using our depth assignment scheme and the Depth Image Based Rendering (DIBR) technique presents little difference to that created based on the depth ground truths.
Guo-Shiang Lin, Cheng-Ying Yeh, Wei-Chih Chen, Wen-Nung Lie
ICME4
2010 Block-based distributed video coding with variable block modes
abstract
In this paper, a new block-based pixel domain distributed video coding scheme featured with variable block modes is proposed. In addition to intra mode and Wyner-Ziv mode employed in conventional block-based distributed video coding scheme, two supplementary block modes “SKIP mode” and “zero motion mode” are introduced in the proposed scheme to improve the overall coding efficiency, as well as to reduce the decoding complexity. Moreover, the channel coding is performed on macroblcok level to reduce the coding loss due to inserted information in the parity bits. The simulation results show that the proposed scheme outperforms both the conventional frame-based transform-domain and the block-based pixel-domain distributed coding schemes.
Jui-Chiu Chiang, Kuan-Liang Chen, Chi-Ju Chou, Chang-Ming Lee, Wen-Nung Lie
ISCAS5
2010 Practical estimation of adaptive correlation noise model for Distributed Video Coding
abstract
In contrast with the traditional video compression system, Distributed Video Coding (DVC) architecture dramatically shifts the complexity from the encoder to the decoder. This low-cost encoding concept can be exploited in the emerging applications, e.g. wireless sensor networks. In order to increase the compression efficiency, improvement of side information generation and refinements of Correlation Noise Model (CNM) are main streams to improve DVC. However, most of these schemes are theoretical and expensive for the decoder. In order to retain low-cost and efficient system, a side information refinement with a practical CNM estimation is proposed. While maintaining the video quality, our proposed mechanism totally improves the system compression efficiency about 18% for the bit-rate with a low complexity decoder.
Chang-Ming Lee, Wen-Nung Lie
ISITA3
2010 A Framework of Enhancing Image Steganography With Picture Quality Optimization and Anti-Steganalysis Based on Simulated Annealing Algorithm
abstract
Picture quality and statistical undetectability are two key issues related to steganography techniques. In this paper, we propose a closed-loop computing framework that iteratively searches proper modifications of pixels/coefficients to enhance a base steganographic scheme with optimized picture quality and higher anti-steganalysis capability. To achieve this goal, an anti-steganalysis tester and an embedding controller-based on the simulated annealing (SA) algorithm with a proper cost function-are incorporated into the processing loop to conduct the convergence of searches. The cost function integrates several performance indices, namely, the mean square error, the human visual system (HVS) deviation, and the differences in statistical features, and guides a proper direction of searches during SA optimization. Our proposed framework is suitable for the kind of steganographic schemes that spreads each message information into multiple pixels/coefficients. We have selected two base steganographic schemes for implementation to show the applicability of the proposed framework. Experiment results show that the base schemes can be enhanced with better performances in image PSNR (by more than 5.0 dB), file-size variation, and anti-steganalysis pass-rate (by about 10% ~ 86%, at middle to high embedding capacities).
Guo-Shiang Lin, Yi-Ting Chang, Wen-Nung Lie
IEEE Trans. Multim.3
2009 Human pose estimation from monocular image captures
abstract
A human pose estimation method from monocular image captures is presented. The objective is to develop a human-computer interface (HCI) for virtual sport activities. In the proposed technique, a graphical 3D human model is first constructed. Its projection on a virtual image plane is then used to match the silhouettes obtained from the image sequence. By iteratively adjusting the 3D pose of the graphical 3D model with the physical and anatomic constraints of human motion, the human pose and the associate 3D motion parameters can be uniquely identified. Experimental results are presented with the real scene images.
Huei-Yung Lin, Ting-Wen Chen, Chih-Chang Chen, Chia-Hao Hsieh, Wen-Nung Lie
ICME5
2009 Enhanced Side Information Generator with Accurate Evaluations in Block-Based Wyner-Ziv Video Coding
Chang-Ming Lee, Jui-Chiu Chiang, Zhi-Heng Chiang, Kuan-Liang Chen, Wen-Nung Lie
PSIVT5
2008 Enhancing error resiliency for multi-hypothesis video coding techniques
abstract
In this paper, two techniques based on the optimization of end-to-end distortions in presence of channel errors, are proposed to enhance the error resiliency of the multi-hypothesis coded videos. First, an error resilient motion estimation algorithm is introduced for a given hypothesis-weighting vector to consider the error resiliency, in addition to the traditional coding efficiency. Second, based on the availability of MVs, the proposed adaptive hypothesis-weighting algorithm makes error resiliency adaptive to video contents, frame by frame. Experiment results show that both techniques are capable of improving PSNR performance by up to 1 dB when the packet loss rate is 15%.
Wen-Nung Lie, Zhi-Wei Gao 0001, Chih-Chang Chen
ICME1
2007 Quality-Optimized Image Steganography Subject to Anti-Steganalysis Constraint
abstract
In this paper, we propose an architecture that combines a quantization-based steganogrphic scheme with a steganalysis system, operated in a closed-loop manner with a cost function for minimization, to enhance the anti-steganalysis capability and image quality after data embedding. In this architecture, a controller based on the SA (simulated annealing) technique is adopted to guarantee fast and accurate convergence. Our proposed system is applied to the data embedding of JPEG-compressed images. Compared with the original embedding algorithm in (Seki, Y, et al., 2005), a better image quality (by an average improvement of 6.48 dB) can be achieved and simultaneously the anti-steganalysis capability is enhanced significantly.
Guo-Shiang Lin, Yi-Ting Chang, Wen-Nung Lie
ICASSP (2)3
2007 Dynamic Key Block Decision with Spatio-Temporal Analysis for Wyner-Ziv Video Coding
abstract
Wyner-Ziv coding has been recognized as the most popular method up to now. For traditional WZC, side information is generated from intra-coded frames for use in the decoding of WZ frames. The unit for intra-coding is a frame and the distance between key-frames is kept constant. In this paper, the unit for intra-coding is a block, and the temporal distance between two consecutive key blocks can varying with time. A block is assigned a mode (WZ or intra-coded), depending on the result of spatio-temporal analysis, and encoded in an alternative manner. This strategy improves the overall coding efficiency, while maintaining a low encoder complexity. The performance gain can achieve up to 6 dB with respect to the traditional pixel-domain WZC.
Dung-Chan Tsai, Chang-Ming Lee, Wen-Nung Lie
ICIP (6)3
2007 Error-Resilience Transcoding of H.264/AVC Compressed Videos
abstract
We propose an error resilience transcoder based on the intra refresh technique for H.264/AVC videos. Our algorithm focuses on making decisions on the number of inserted intra MBs, allocations of proper MBs for intra refresh, and the quantization parameter for each P-frame, according to varying channel condition and frame contents. We build source distortion, channel distortion, and bit-rate models for prediction and adopt the Lagrangian multiplier approach for constrained optimization. Existing algorithms are compared in experiments. The results show that our algorithm improves PSNR by 0.28~2.76 dB, compared to the "equal intra refresh" approach. Our algorithm also outperforms CBERC (Chen and Wang, 2001) by 0.1~1.13 dB, with the same number of inserted intra MBs, but different MB priorities in allocation
Wen-Nung Lie, Han-Ching Yeh, Zhi-Wei Gao 0001, Ping-Chang Jui
ISCAS1
2007 Prescription-based error concealment technique for video transmission on error-prone channels
Wen-Nung Lie, Tom C.-I. Lin
J. Vis. Commun. Image Represent.1
2006 Error Resilient Motion Estimation for Video Coding
Wen-Nung Lie, Zhi-Wei Gao 0001, Wei-Chih Chen, Ping-Chang Jui
PSIVT1
2006 Joint Source-Channel Video Coding Based on the Optimization of End-to-End Distortions
Wen-Nung Lie, Zhi-Wei Gao 0001, Tung-Lin Liu, Ping-Chang Jui
PSIVT1
2006 Rate-distortion-smoothness optimized rate allocation schemes for Spectral Fine Granular Scalable video coding technique
Wen-Nung Lie, Cheng-Hsiung Tseng, Tom C.-I. Lin, Ming-Yang Tseng, I-Cheng Ting
J. Vis. Commun. Image Represent.1
2006 Video Error Concealment by Integrating Greedy Suboptimization and Kalman Filtering Techniques
abstract
This paper addresses video error concealment techniques, focusing on motion vector recovery of P- and B-frames, to improve the decoded quality of videos when bit stream data incurs transmission errors. First, we propose a dynamic programming (DP) technique to optimize the path cost in a multistage topology and evaluate the goodness of boundary matching and side smoothness of recovered macroblocks. However, due to the high computational complexity of DP, a suboptimal alternative enhanced with an adaptive Kalman filtering algorithm is adopted instead. Experiments show that by considering the side smoothness between adjacent recovered MBs, the proposed algorithm improves the reconstructed video quality by about 0.4-0.9 dB with a packet loss rate up to 15%, compared to a traditional boundary matching algorithm. In addition, subjective image inspection demonstrates the efficiency of the proposed algorithm in retaining the continuity of lines and image details
Wen-Nung Lie
IEEE Trans. Circuits Syst. Video Technol.1
2006 Enhancing video error resilience by using data-embedding techniques
abstract
In this paper, error-resilient video coding schemes based on data-embedding techniques are proposed for the H.263+ codec. Data embedding, popularly applied to secret hiding and digital watermarking, is now used to convey error recovery information to the decoder via a covert channel, without causing significant increase in transmission bitrate. Our embedded information provides implicit macroblock (MB) delimiters for resynchronization in presence of channel errors. In this way, the decoder is capable of isolating erroneous MBs with the extracted information. A set of variational schemes is proposed, extensively analyzed, and compared to the competitive counterparts (e.g., the original H.263+ TMN8 and its synchronization-enhanced version) at the same bitrate. Experimental results show that our data embedding process decreases the average peak signal-to-noise ratio (PSNR) in error-free conditions by 0.3 dB (light data embedding) to 1.8 dB (heavy data embedding), but it is capable of achieving a significant PSNR improvement up to 2 and 9.5 dB when the bit error rate is 10/sup -5/ and 10/sup -3/, respectively. We also provide suggestions of how to adaptively apply the proposed schemes for different channel error conditions and different video contents.
Wen-Nung Lie, Tom C.-I. Lin, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.1
2006 Dual protection of JPEG images based on informed embedding and two-stage watermark extraction techniques
abstract
In this paper, the authors propose a watermarking scheme that embeds both image-dependent and fixed-part marks for dual protection (content authentication and copyright claim) of JPEG images. To achieve the goals of efficiency, imperceptibility, and robustness, a compressed-domain informed embedding algorithm, which incorporates the Lagrangian multiplier optimization approach and an adjustment procedure, is developed. A two-stage watermark extraction procedure is devised to achieve the functionality of dual protection. In the first stage, the semifragile watermark in each local channel is extracted for content authentication. Then, in the second stage, a weighted soft-decision decoder, which weights the signal detected in each channel according to the estimated channel condition, is used to improve the recovery rate of the fixed-part watermark for copyright protection. The experiment results manifest that the proposed scheme not only achieve dual protection of the image content, but also maintain higher visual quality (an average of 6.69 dB better than a comparable approach) for a specified level of watermark robustness. In addition, the overall computing load is low enough to be practical in real-time applications.
Wen-Nung Lie, Guo-Shiang Lin, Sheng-Lung Cheng
IEEE Trans. Inf. Forensics Secur.1
2006 Robust and high-quality time-domain audio watermarking based on low-frequency amplitude modification
abstract
This work proposes a method of embedding digital watermarks into audio signals in the time domain. The proposed algorithm exploits differential average-of-absolute-amplitude relations within each group of audio samples to represent one-bit information. The principle of low-frequency amplitude modification is employed to scale amplitudes in a group manner (unlike the sample-by-sample manner as used in pseudonoise or spread-spectrum techniques) in selected sections of samples so that the time-domain waveform envelope can be almost preserved. Besides, when the frequency-domain characteristics of the watermark signal are controlled by applying absolute hearing thresholds in the psychoacoustic model, the distortion associated with watermarking is hardly perceivable by human ears. The watermark can be blindly extracted without knowledge of the original signal. Subjective and objective tests reveal that the proposed watermarking scheme maintains high audio quality and is simultaneously highly robust to pirate attacks, including MP3 compression, low-pass filtering, amplitude scaling, time scaling, digital-to-analog/analog-to-digital reacquisition, cropping, sampling rate change, and bit resolution transformation. Security of embedded watermarks is enhanced by adopting unequal section lengths determined by a secret key.
Wen-Nung Lie, Li-Chun Chang
IEEE Trans. Multim.1
2005 News video classification based on multi-modal information fusion
abstract
A multi-modal information fusion technique integrating the closed caption, anchor's speech, and visual information for TV news video classification is presented. By recognizing closed-caption characters from video, phrases of single- and double-character are found for classification. On the other hand, content of the anchor's speech signal is not recognized, but instead, labeled with pre-trained cluster means by using a level-building DP (dynamic programming) algorithm. Visual information, including the color and motion features, is extracted from the news footage part for classification. The above three information is individually classified by using statistical relevance factor (RF) or SVM (support vector machine) technique, amounting to 7 different classifiers. Results of multiple classifiers are then combined to get fused outputs by using a modified Bayesian technique. Experiments show that the proposed fusion system is capable of increasing the classification rate by 14% with respect to the best single-modal system. Our Bayesian fusion rule also outperforms the best product rule presented in J. Kittler, et al (1998) by 3%.
Wen-Nung Lie, Chen-Kang Su
ICIP (1)1
2005 Error Resilient Coding Based on Reversible Data Embedding Technique for H.264/AVC Video
abstract
In this paper, a prescription-based error concealment (PEC) method is proposed. PEC relies on pre-analyses of the concealment error image (CEI) for I-frames and the optimal error concealment scheme for P-frames at encoder side. CEI is used to enhance the image quality at decoder side after error concealment (by spatial interpolation or zero motion) of the corrupted intra-coded MBs. A set of pre-selected error concealment methods is evaluated for each corrupted inter-coded MB to determine the optimal one for the decoder. Both the CEI and the scheme indices are considered as the prescriptions for decoder and transmitted along with the video bit stream based on a reversible data embedding technique. Experiments show that the proposed method is capable of achieving PSNR improvement of up to 1.48 dB, at a considerable bit-rate, when the packet loss rate is 20%
Wen-Nung Lie, Tom C.-I. Lin, Dung-Chan Tsai, Guo-Shiang Lin
ICME1
2005 Combining Caption and Visual Features for Semantic Event Classification of Baseball Video
abstract
In baseball game, an event is defined as the portion of video clip between two pitches, and a play is defined as a batter finishing his plate appearance. A play is a concatenation of many events, and a baseball game is formed by a series of plays. In this paper, only the event happened in the last pitch of a plate appearance is detected. It is then semantically classified to represent the corresponding play by using an algorithm integrating caption rule-inference and visual feature analysis. Our proposed system is capable of classifying each baseball play into eleven semantic categories, which are popular and familiar to most of the audiences. In an experiment of 260 testing plays, the classification rate achieves up to 87%.
Wen-Nung Lie, Sheng-Hsiung Shia
ICME1
2005 A robust dynamic programming algorithm to extract skyline in images for navigation
Wen-Nung Lie, Tom C.-I. Lin, Ting-Chih Lin, Keng-Shen Hung
Pattern Recognit. Lett.1
2005 A feature-based classification technique for blind image steganalysis
abstract
In contrast to steganography, steganalysis is focused on detecting (the main goal of this research), tracking, extracting, and modifying secret messages transmitted through a covert channel. In this paper, a feature classification technique, based on the analysis of two statistical properties in the spatial and DCT domains, is proposed to blindly (i.e., without knowledge of the steganographic schemes) to determine the existence of hidden messages in an image. To be effective in class separation, the nonlinear neural classifier was adopted. For evaluation, a database composed of 2088 plain and stego images (generated by using six different embedding schemes) was established. Based on this database, extensive experiments were conducted to prove the feasibility and diversity of our proposed system. It was found that the proposed system consists of: 1) a 90%/sup +/ positive-detection rate; 2) not limited to the detection of a particular steganographic scheme; 3) capable of detecting stego images with an embedding rate as low as 0.01 bpp; and 4) considering the test of plain images incurred low-pass filtering, sharpening, and JPEG compression.
Wen-Nung Lie, Guo-Shiang Lin
IEEE Trans. Multim.1
2004 Content-based retrieval of MP3 songs based on query by singing
abstract
With the growth of multimedia in the Internet, content analysis of multimedia plays an important role for humanistic management. We investigate the content-based retrieval of MP3 songs based on the interface of query by singing. MDCT (modified DCT) spectral coefficients are directly used to represent the tonic characteristics of a short-term sound. This spectral profile is used for detailed matching between two audio segments. Perceptual features are also computed from MDCT coefficients for audio classification. Two pre-stages based on SVM and k-means classifications are used to remove incorrect (or noisy) segment candidates and to speed up the subsequent matching process. On the other hand, exponential key-scaling schemes and time-warping techniques are developed to overcome key difference and tempo variation between different singers. Experiments show that the retrieval probability of our design can achieve up to 76% among the top 5 out of a total of 114 excerpts in the database.
Wen-Nung Lie, Chen-Kang Su
ICASSP (5)1
2004 Rate-distortion optimized DCT-domain video transcoder for bit-rate reduction of MPEG videos
abstract
In this paper, we propose a rate-distortion optimized video transcoder which converts MPEG videos into a similar form at lower bit-rates. Our transcoder design is characterized by two features: 1) transcoding is performed in the DCT domain and the motion vector information is re-used; and 2) the rate-distortion relationship is optimized in both the frame-level rate allocation and macroblock-level rate control, thus leading to performances even better than the direct encoding and re-encoding methods based on the well-known TM5. In the proposed algorithm, the Lagrangian multiplier plays not only its traditional role in macroblock-level optimization, but also a variable to be optimized in frame-level rate allocation. These two levels of optimization process are highly linked. Experiments show that the R-D optimization is effective in getting better video quality, even the drift errors are ignored. Several speedy schemes were developed to make our transcoder design suitable for real-time video transmission over heterogeneous networks.
Wen-Nung Lie, Ming-Lun Tsai, Tom C.-I. Lin
ICASSP (5)1
2004 Motion-based event detection and semantic classification for baseball sport videos
abstract
The techniques of event detection and semantic classification of baseball sport videos are investigated. Due to abundant motion information in sport videos, motion vectors are estimated, validated, and used to compute both the motion activity and camera motion parameters of a frame. Considering the domain-specific knowledge of baseball sport, behaviors of motion features in the spatial or temporal domain are analyzed for segmenting the whole baseball video into a lot of play events (defined as the time interval between two pitching shots), from which a key shot is identified for semantic classification. Taking motion features of key shots as the input of a neural network, our proposed system is capable of classifying segmented events into three semantic categories as "non-hitting", "in-field", and "out-field". According to experiments on more than 200 events, we can achieve a classification rate of 91.55%. The proposed technique will be helpful in applications of baseball video summary and retrieval.
Wen-Nung Lie, Ting-Chih Lin, Sheng-Hsiung Hsia
ICME1
2004 Error-resilient spectral fine granular scalable (SFGS) video coding for network streaming applications
abstract
As networking grows, video transmission becomes more popular. MPEG4 FGS is a video coding technique designed to cope with variability of bandwidth gracefully. Our previously proposed spectral FGS (SFGS) coding technique, a variation of traditional FGS, when in combination with a rate-distortion-based rate allocation scheme, was capable of making video quality smoother between pictures in a sequence. However, both FGS and SFGS-coded video streams, when transmitted through networks, may be partially corrupted by channel errors or packet loss. To resist errors, we add sync codes according to the spectrum-grouping property of SFGS to enhance error resilience capability. This scheme of adding sync codes is simple, does not cause heavy loads of the system, and can on average improve video quality by 1/spl sim/3 dB when errors are present.
Wen-Nung Lie, Cheng-Hshing Tseng, Ping-Chang Jui
ICME1
2004 Constant-quality rate allocation for spectral fine granular scalable (SFGS) video coding by using dynamic programming approach
abstract
Spectral FGS (SFGS), modified from MPEG-4 FGS, was proposed as a scalable video coding technique that has more even image quality and is more suitable for rate allocation for video smoothness. This paper proposed a unified rate allocation algorithm based on SFGS/SFGST to make good tradeoffs between image quality, video smoothness, and temporal scalability, according to users' preferences at the buffer fullness constraint. Our rate allocation algorithm was established based on the dynamic programming (DP) technique. It becomes easy to tune the preference by adjusting weighting parameters in the optimized cost function. Experiments show that our algorithm is capable of achieving the same video smoothness as our previous work (Wen-Nung Lie et al, Proc. IEEE Int. Symp. on Circuits and Systems, pp. 880-883, 2003), having more diversity in mode selection between SFGS and SFGST, and guaranteeing no buffer overflow or underflow.
Wen-Nung Lie, Cheng-Hsiung Tseng, Tom C.-I. Lin
ICME1
2004 A downstream algorithm based on extended gradient vector flow field for object segmentation
abstract
For object segmentation, traditional snake algorithms often require human interaction; region growing methods are considerably dependent on the selected homogeneity criterion and initial seeds; watershed algorithms, however, have the drawback of over segmentation. A new downstream algorithm based on a proposed extended gradient vector flow (E-GVF) field model is presented in this paper for multiobject segmentation. The proposed flow field, on one hand, diffuses and propagates gradients near object boundaries to provide an effective guiding force and, on the other hand, presents a higher resolution of direction than traditional GVF field. The downstream process starts with a set of seeds scored and selected by considering local gradient direction information around each pixel. This step is automatic and requires no human interaction, making our algorithm more suitable for practical applications. Experiments show that our algorithm is noise resistant and has the advantage of segmenting objects that are separated from the background, while ignoring the internal structures of them. We have tested the proposed algorithm with several realistic images (e.g., medical and complex background images) and gained good results.
Cheng-Hung Chuang, Wen-Nung Lie
IEEE Trans. Image Process.2
2003 Verification of image content integrity by using dual watermarking on wavelets domain
abstract
Fragile watermarking is designed to protect content integrity of digital data, meaning to detect any change in the content as well as localize areas that have been changed. In this paper, we propose a fragile watermarking scheme, which embeds dual watermarks (one is naturally fragile and the other is actually robust) in JPEG-2000 compressed images with ROI (region of interest) selection. In this scheme, the fragile watermark is embedded in ROI coefficients at the first decomposition layer, while the robust watermark embedded in ROB (region of backgrounds) coefficients at the third decomposition layer. The combinational working of them is not only capable of detecting malicious changes precisely and flexibly but also estimating the degree of alteration to discriminate intended attacks from unintended ones (e.g., such normal image processing as compression, low-pass filtering, sharpening, etc). Experiments show that our functionally fragile watermarking scheme is scalable in view of detecting attack's intension as well as accuracy in localizing the tampered areas.
Wen-Nung Lie, Tze-Liang Hsu, Guo-Shiang Lin
ICIP (2)1
2002 Multi-view representation and synthesis for 3D object movie
abstract
An object movie can provide viewers with a capability of seeing a 3-D object from any viewpoint. To achieve the best resolution, it needs many cameras to capture views simultaneously. This is however not cost-effective. To solve this problem, we take photographs around the object in a constant angle interval. After this, the object movie system is constructed in a manner of viewpoint synthesis from originally captured images. To further reduce the data requirements of the object movie, we propose a method to retain effective regions in each captured image. Only the retained regions are used to synthesize the resulting virtual views. Experiments show that our algorithm can synthesize virtual views better and faster than previous methods. The results also show that when an object movie is constructed from camera views with 10-degrees of angle interval, the effective region selection can reduce about 40% of pixels in multi-view representation while also keeping good synthesized quality.
Wen-Nung Lie, Bo-Er Wei
ICIP (2)1
2001 Region growing based on extended gradient vector flow field model for multiple objects segmentation
abstract
For image segmentation, traditional snake algorithms are often short of the requirement of human interaction and capability in processing multiple objects simultaneously. Watershed techniques however have the drawback of over-segmentation. A new region growing algorithm based on the extended gradient vector flow (E-GVF) field model is proposed for multiple object segmentation. The proposed force field propagates gradient information of object boundaries and provides a good feature for region growing. We perform scoring and selection of seeds by considering their local gradient direction information. This step is automatic and requires no human interaction, making our algorithm suitable for applications. Experiments show that our algorithm is noise-resistant and also resolves the abovementioned drawbacks for snakes and watershed methods. We have tested our algorithm in segmenting multiple objects from realistic and even medical CT images and gained good results.
Cheng-Hung Chuang, Wen-Nung Lie
ICIP (3)2
2001 Dynamic rate control for MPEG-2 bit stream transcoding
abstract
Most of the multimedia service over the Internet uses a pre-encoded video bit stream for transmission. This requires a transcoder capable of performing on-line bit rate conversion according to bandwidth variations to prevent network congestion and subsequently minimize cell data loss. In this paper, re-quantization of the DCT coefficients is adopted and rate control is performed at both the frame and macroblock levels. At the frame level, we prefer greater bit allocation for I-frames for video quality consideration. At the macroblock level, a feedback mechanism is devised to keep the transcoded output bit rate approximate to the varying channel capacity. Experiments show that our algorithm behaves well in its rate control capability and besides, keeps a nearly constant PSNR degradation of only 2 dB relative to the MPEG-2 encoder when a bit rate conversion ratio from 100% down to 50% is requested.
Wen-Nung Lie, Ying-Hsiang Chen
ICIP (1)1
2001 Tracking Moving Objects In Mpeg-Compressed Videos
Wen-Nung Lie, Ruey-Lung Chen
ICME1
2001 An Efficient Public Key Video Crypto-System
abstract
A new public key crypto-system for transmitting secured MPEG videos is proposed in this paper. Our proposed system retains the MPEG bit stream structure, but with the DC differential values and some AC coefficients of intra-coded macroblocks being encrypted, hence resulting in a content invisibility for all I-, P-, and B-frames. In this way, our algorithm is efficient in terms of processing speed and data conversion overhead. Parameters that control the generation of several pseudo-random binary patterns for encryption can be dynamically updated, encrypted by a public-key algorithm, and delivered to the receiver through the encrypted MPEG video by data embedding techniques. Since these parameters can be recovered for legal users by using a private key, transmission security can be thus achieved. Besides, our crypto-system is capable of changing these parameters dynamically and making them readable at the receiver side, it is difficult for illegal users to deduce the decryption parameters by some statistical methods.
Guo-Shiang Lin, Tom C.-I. Lin, Wen-Nung Lie, Ting-Chih Lin
ICME3
2000 Robust image watermarking on the DCT domain
abstract
This paper proposes two simple DCT-domain-based schemes to embed single or multiple watermarks into an image for copyright protection and data monitoring and tracking. The watermark data are essentially embedded in the middle band of the DCT domain to make a tradeoff between visual degradation and robustness. The proposed schemes are simple and no original host image is required for watermark extraction. The algorithm also features the capability of embedding multiple orthogonal watermarks into an image simultaneously. A set of systematic experiments, including Gaussian smoothing, JPEG compression, and image cropping are performed to prove the robustness of our algorithms.
Wen-Nung Lie, Guo-Shiang Lin, Chih-Liang Wu, Ta-Chun Wang
ISCAS1
2000 Multispectral satellite image compression based on multimode linear prediction
Wen-Nung Lie, Chun-Hung Chen, Chi-Fa Chen
VCIP1
1999 Data Hiding in Images with Adaptive Numbers of Least Significant Bits Based on the Human Visual System
abstract
We propose in this paper a novel method for embedding multimedia data (including audio, image, video, or text; compressed or non-compressed) into a host image. The classical LSB concept is adopted, but with the number of LSBs adapting to pixels of different graylevels. A piecewise mapping function according to human visual sensitivity of contrast is used so that adaptivity can be achieved without extra bits for overhead. The leading information for data decoding is few, no more than 3 bytes. Experiments show that a large amount of bit streams (nearly 30%-45% of the host image) can be embedded without sever degradation of the image quality (33-40 dB, depending on the volume of embedded bits).
Wen-Nung Lie, Li-Chun Chang
ICIP (1)1
1995 Automatic target segmentation by locally adaptive image thresholding
abstract
A locally adaptive thresholding algorithm, concerning the extraction of targets from a given field of background, is proposed. Conventional histogram-based or global-type methods are deficient in detecting small targets of possibly low contrast as well. The present research is notable for solving the mentioned problems by introducing (1) shape connectivity measure based on co-occurrence statistics for threshold evaluation; and (2) no-target identification procedure for modeling a local-processing paradigm. In this manner, thresholds are determined adaptively even in the presence of space-varying noise or clutter. Experiments show that the results are reliable and even outperform those that manual operations can achieve for global thresholding.
Wen-Nung Lie
IEEE Trans. Image Process.1
1994 A Fuzzy-Computing Method for Rotation-Invariant Image Tracking
abstract
A fuzzy tracking system is developed for rotation-invariant image tracking. The conventional rotation-invariant methods cost too much time on transformation for invariant-feature extraction (e.g., circular harmonic). They am not suitable for the implementation of real-time operation. In this paper, a dual-template strategy is proposed. Only two parallel matched filters and a simple fuzzy logic are employed to construct the novel fuzzy tracking system. The experiments show that the system can track the target more accurately and more rapidly.>
Wen-Nung Lie, Yung-Chang Chen
ICIP (1)2
1993 An efficient threshold-evaluation algorithm for image segmentation based on spatial graylevel co-occurrences
Wen-Nung Lie
Signal Process.1
1990 Robust line-drawing extraction for polyhedra using weighted polarized hough transform
Wen-Nung Lie, Yung-Chang Chen
Pattern Recognit.1
1990 Model-based recognition and positioning of polyhedra using intensity-guided range sensing and interpretation in 3-D space
Wen-Nung Lie, Ching-Wen Yu, Yung-Chang Chen
Pattern Recognit.1
1988 Moving object detection, inspection, and counting using image stripe analysis
Wen-Nung Lie, Yung-Chang Chen
Pattern Recognit. Lett.1