Daniel Pak-Kong Lun

dblp:73/5761 · DBLP profile ↗
← Back
76ranked-venue papers
13as first author
9since 2021 · last 2024
0000-0003-3891-1363ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 8 first-author · 5 since 2021Systems, architecture and hardware · 14 · 4 first-authorArtificial intelligence and machine learning · 9 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Solving the imbalanced dataset problem in surveillance image blur classification
Yikun Pan, Sik-Ho Tsang, Tom Tak-Lam Chan, Yui-Lam Chan, Daniel Pak-Kong Lun
Eng. Appl. Artif. Intell.5
2023 Optimized Quality Feature Learning for Video Quality Assessment
abstract
Recently, some transfer learning-based methods have been adopted in video quality assessment (VQA) to compensate for the lack of enormous training samples and human annotation labels. But these methods induce a domain gap between source and target domains, resulting in a sub-optimal feature representation that deteriorates the accuracy. This paper proposes the optimized quality feature learning via a multi-channel convolutional neural network (CNN) with the gated recurrent unit (GRU) for no-reference (NR) VQA. First, the multi-channel CNN is pre-trained on the image quality assessment (IQA) domain using non-human annotation labels, which is inspired by self-supervised learning. Then, semi-supervised learning is used to fine-tune CNN and transfer the knowledge from IQA to VQA while considering motion information for optimized quality feature learning. Finally, all frame quality features are extracted as the input of GRU to obtain the final video quality. Experimental results demonstrate that our model achieves better performance than state-of-the-art VQA approaches.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Daniel Pak-Kong Lun
ICASSP4
2023 Self-embedding reversible color-to-grayscale conversion with watermarking feature
S. K. Felix Yu, Yuk-Hee Chan, Kin-Man Lam 0001, Daniel Pak-Kong Lun
Signal Process. Image Commun.4
2023 Geometry-Aware Facial Expression Recognition via Attentive Graph Convolutional Networks
abstract
Learning discriminative representations with good robustness from facial observations serves as a fundamental step towards intelligent facial expression recognition (FER). In this article, we propose a novel geometry-aware FER framework to boost the FER performance based on both the geometric and appearance knowledge. Specifically, we propose an encoding strategy for facial landmarks, and adopt a graph convolutional network (GCN) to fully explore the structural information of the facial components behind different expressions. A convolutional neural network (CNN) is further applied to the whole facial observation to learn the global characteristics of different expressions. The features from these two networks are fused into a comprehensive high-semantic representation, which promotes the FER reasoning from both visual and structural perspectives. Moreover, to facilitate the networks to concentrate on the most informative facial regions and components, we introduce multi-level attention mechanisms into the proposed framework, which enhance the reliability of the learned representations for effective FER. Experiments on two challenging FER benchmarks demonstrate that the attentive graph-based learning on the facial geometry boosts the FER accuracy. Furthermore, the insensitivity of the geometric information to the appearance variations also improves the generalization of the proposed framework.
Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001
IEEE Trans. Affect. Comput.4
2023 Spatial-Temporal Graphs Plus Transformers for Geometry-Guided Facial Expression Recognition
abstract
Facial expression recognition (FER) is of great interest to the current studies of human-computer interaction. In this paper, we propose a novel geometry-guided facial expression recognition framework, based on graph convolutional networks and transformers, to perform effective emotion recognition from videos. Specifically, we detect and utilize facial landmarks to construct a spatial-temporal graph, based on both the landmark coordinates and local appearance, for representing a facial expression sequence. The graph convolutional blocks and transformer modules are employed to produce high-semantic emotion-related representations from the structured facial graphs, which facilitate the framework to establish both the local and non-local dependency between the vertices. Moreover, spatial and temporal attention mechanisms are introduced into graph-based learning to promote FER reasoning, via the emphasis on the most informative facial components and frames. Extensive experiments demonstrate that the proposed framework achieves promising performance for geometry-based FER and shows great generalization and robustness in real-world applications.
Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001
IEEE Trans. Affect. Comput.4
2022 Joint Spine Segmentation and Noise Removal From Ultrasound Volume Projection Images With Selective Feature Sharing
abstract
Volume Projection Imaging from ultrasound data is a promising technique to visualize spine features and diagnose Adolescent Idiopathic Scoliosis. In this paper, we present a novel multi-task framework to reduce the scan noise in volume projection images and to segment different spine features simultaneously, which provides an appealing alternative for intelligent scoliosis assessment in clinical applications. Our proposed framework consists of two streams: i) A noise removal stream based on generative adversarial networks, which aims to achieve effective scan noise removal in a weakly-supervised manner, i.e., without paired noisy-clean samples for learning; ii) A spine segmentation stream, which aims to predict accurate bone masks. To establish the interaction between these two tasks, we propose a selective feature-sharing strategy to transfer only the beneficial features, while filtering out the useless or harmful information. We evaluate our proposed framework on both scan noise removal and spine segmentation tasks. The experimental results demonstrate that our proposed method achieves promising performance on both tasks, which provides an appealing approach to facilitating clinical diagnosis.
Zixun Huang, Rui Zhao 0012, Frank H. F. Leung, Sunetra Banerjee, Timothy Tin-Yan Lee, De Yang, Daniel Pak-Kong Lun, Kin-Man Lam 0001, Sai-Ho Ling
IEEE Trans. Medical Imaging7
2021 Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection Images
abstract
Automatic spine segmentation, based on ultrasound volume projection imaging (VPI), is of great value in clinical applications to diagnose scoliosis in teenagers. In this paper, we propose a novel framework to improve the segmentation accuracy on spine images via structure-enhanced attentive learning. Since the spine bones contain strong prior knowledge of their shapes and positions in ultrasound VPI images, we propose to encode this information into the semantic representations in an attentive manner. We first revisit the self-attention mechanism in representation learning, and then present a strategy to introduce the structural knowledge into the key representation in self-attention. By this means, the network explores both the contextual and structural information in the learned features, and consequently improves the segmentation accuracy. We conduct various experiments to demonstrate that our proposed method achieves promising performance on spine image segmentation, which shows great potential in clinical diagnosis.
Rui Zhao 0012, Zixun Huang, Tianshan Liu, Frank H. F. Leung, Sai-Ho Ling, De Yang, Timothy Tin-Yan Lee, Daniel Pak-Kong Lun, Kin-Man Lam 0001
ICASSP8
2021 Improved Multiple-Image-Based Reflection Removal Algorithm Using Deep Neural Networks
abstract
When imaging through a semi-reflective medium such as glass, the reflection of another scene can often be found in the captured images. It degrades the quality of the images and affects their subsequent analyses. In this paper, a novel deep neural network approach for solving the reflection problem in imaging is presented. Traditional reflection removal methods not only require long computation time for solving different optimization functions, their performance is also not guaranteed. As array cameras are readily available in nowadays imaging devices, we first suggest in this paper a multiple-image based depth estimation method using a convolutional neural network (CNN). The proposed network avoids the depth ambiguity problem due to the reflection in the image, and directly estimates the depths along the image edges. They are then used to classify the edges as belonging to the background or reflection. Since edges having similar depth values are error prone in the classification, they are removed from the reflection removal process. We suggest a generative adversarial network (GAN) to regenerate the removed background edges. Finally, the estimated background edge map is fed to another auto-encoder network to assist the extraction of the background from the original image. Experimental results show that the proposed reflection removal algorithm achieves superior performance both quantitatively and qualitatively as compared to the state-of-the-art methods. The proposed algorithm also shows much faster speed compared to the existing approaches using the traditional optimization methods.
Tingtian Li, Yuk-Hee Chan, Daniel Pak-Kong Lun
IEEE Trans. Image Process.3
2021 Invertible Image Decolorization
abstract
Invertible image decolorization is a useful color compression technique to reduce the cost in multimedia systems. Invertible decolorization aims to synthesize faithful grayscales from color images, which can be fully restored to the original color version. In this paper, we propose a novel color compression method to produce invertible grayscale images using invertible neural networks (INNs). Our key idea is to separate the color information from color images, and encode the color information into a set of Gaussian distributed latent variables via INNs. By this means, we force the color information lost in grayscale generation to be independent of the input color image. Therefore, the original color version can be efficiently recovered by randomly re-sampling a new set of Gaussian distributed variables, together with the synthetic grayscale, through the reverse mapping of INNs. To effectively learn the invertible grayscale, we introduce the wavelet transformation into a UNet-like INN architecture, and further present a quantization embedding to prevent the information omission in format conversion, which improves the generalizability of the framework in real-world scenarios. Extensive experiments on three widely used benchmarks demonstrate that the proposed method achieves a state-of-the-art performance in terms of both qualitative and quantitative results, which shows its superiority in multimedia communication and storage systems.
Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001
IEEE Trans. Image Process.4
2020 NTGAN: Learning Blind Image Denoising without Clean Reference
Rui Zhao 0012, Daniel Pak-Kong Lun, Kin-Man Lam 0001
BMVC2
2020 Video Lightening with Dedicated CNN Architecture
abstract
Darkness brings us uncertainty, worry and low confidence. This is a problem not only applicable to us walking in a dark evening but also for drivers driving a car on the road with very dim or even without lighting condition. To address this problem, we propose a new CNN structure named as Video Lightening Network (VLN) that regards the low-light enhancement as a residual learning task, which is useful as reference to indirectly lightening the environment, or for vision-based application systems, such as driving assistant systems. The VLN consists of several Lightening Back-Projection (LBP) and Temporal Aggregation (TA) blocks. Each LBP block enhances the low-light frame by domain transfer learning that iteratively maps the frame between the low- and normal-light domains. A TA block handles the motion among neighboring frames by investigating the spatial and temporal relationships. Several TAs work in a multi-scale way, which compensates the motions at different levels. The proposed architecture has a consistent enhancement for different levels of illuminations, which significantly increases the visual quality even in the extremely dark environment. Extensive experimental results show that the proposed approach outperforms other methods under both objective and subjective metrics.
Li-Wen Wang, Wan-Chi Siu, Chu-Tak Li, Daniel Pak-Kong Lun
ICPR5
2020 Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature Sharing
abstract
Multi-task learning is an effective learning strategy for deep-learning-based facial expression recognition tasks. However, most existing methods take into limited consideration the feature selection, when transferring information between different tasks, which may lead to task interference when training the multi-task networks. To address this problem, we propose a novel selective feature-sharing method, and establish a multi-task network for facial expression recognition and facial expression synthesis. The proposed method can effectively transfer beneficial features between different tasks, while filtering out useless and harmful information. Moreover, we employ the facial expression synthesis task to enlarge and balance the training dataset to further enhance the generalization ability of the proposed method. Experimental results show that the proposed method achieves state-of-the-art performance on those commonly used facial expression recognition benchmarks, which makes it a potential solution to real-world facial expression recognition problems.
Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001
ICPR4
2020 Deep Lightening Network for Low-Light Image Enhancement
abstract
We propose a Deep Lightening Network (DLN) for low-light image enhancement. Inspire by the domain transfer study, we propose a novel cycle learning structure to learn the mapping relationship between low- and normal-light images. Each DLN consists of several Lightening Back-Projection (LBP) blocks that learn the residual between low- and normal-light images. To efficiently estimate the local and global information, we fuse the features from different LBP results. Experimental results on different datasets show that our proposed DLN approach outperforms other approaches in all objective and subjective measures.
Li-Wen Wang, Wan-Chi Siu, Daniel Pak-Kong Lun
ISCAS4
2020 Stereoscopic image reflection removal based on Wasserstein Generative Adversarial Network
abstract
Reflection removal is a long-standing problem in computer vision. In this paper, we consider the reflection removal problem for stereoscopic images. By exploiting the depth information of stereoscopic images, a new background edge estimation algorithm based on the Wasserstein Generative Adversarial Network (WGAN) is proposed to distinguish the edges of the background image from the reflection. The background edges are then used to reconstruct the background image. We compare the proposed approach with the state-of-the- art reflection removal methods. Results show that the proposed approach can outperform the traditional single-image based methods and is comparable to the multiple-image based approach while having a much simpler imaging hardware requirement.
Xiuyuan Wang 0006, Yikun Pan, Daniel Pak-Kong Lun
VCIP3
2020 A Framework of Reversible Color-to-Grayscale Conversion With Watermarking Feature
abstract
Reversible color-to-grayscale conversion (RCGC) is a method that embeds the chromatic information of a full color image into its grayscale version such that the original color image can be reconstructed in the future when necessary. In practical applications, it is required to provide a means to authenticate an information-embedded image such that its integrity can be guaranteed. However, none of the current RCGC algorithms take this factor into account. In this paper, to address this issue, we develop an information-embedding framework based on a vector quantization-based (VQ-based) RCGC algorithm recently proposed by us. Under this framework, we propose a RCGC algorithm that can embed both chromatic information and fragile watermark simultaneously into a grayscale image with the same technique to reduce the complexity and improve the efficiency. Like other VQ-based RCGC algorithms, the performance of the proposed RCGC algorithm highly relies on the palette it uses. We also propose a palette generation algorithm in this paper to support the information embedding process such that the visual quality of the color-embedded grayscale images and the reconstructed color images can be significantly improved.
Yuk-Hee Chan, Zi-Xin Xu, Daniel Pak-Kong Lun
IEEE Trans. Image Process.3
2020 Lightening Network for Low-Light Image Enhancement
abstract
Low-light image enhancement is a challenging task that has attracted considerable attention. Pictures taken in low-light conditions often have bad visual quality. To address the problem, we regard the low-light enhancement as a residual learning problem that is to estimate the residual between low- and normal-light images. In this paper, we propose a novel Deep Lightening Network (DLN) that benefits from the recent development of Convolutional Neural Networks (CNNs). The proposed DLN consists of several Lightening Back-Projection (LBP) blocks. The LBPs perform lightening and darkening processes iteratively to learn the residual for normal-light estimations. To effectively utilize the local and global features, we also propose a Feature Aggregation (FA) block that adaptively fuses the results of different LBPs. We evaluate the proposed method on different datasets. Numerical results show that our proposed DLN approach outperforms other methods under both objective and subjective metrics.
Li-Wen Wang, Wan-Chi Siu, Daniel Pak-Kong Lun
IEEE Trans. Image Process.4
2019 Image Reflection Removal Using the Wasserstein Generative Adversarial Network
abstract
Imaging through a semi-transparent material such as glass often suffers from the reflection problem, which degrades the image quality. Reflection removal is a challenging task since it is severely ill-posed. Traditional methods, while all require long computation time on minimizing different objective functions with huge matrices, do not necessarily give satisfactory performance. In this paper, we propose a novel deep-learning based method to allow fast removal of reflection. Similar to the traditional multiple-image approaches, the proposed algorithm first captures the multi-view images of a scene. Then the images are fed to a convolutional neural network to obtain the depth information along the edges of the image. It is sent to a Wasserstein generative adversarial networks (WGAN) for estimating the edges of the background. Finally, the background edges are used in another WGAN to reconstruct the background image. Experimental results show that the proposed method can achieve state-of-the-art performance, and is significantly faster than the traditional methods due to the use of the deep learning methods.
Tingtian Li, Daniel Pak-Kong Lun
ICASSP2
2019 Semi-Supervised Deep Vision-Based Localization Using Temporal Correlation Between Consecutive Frames
abstract
Vision-based localization is a temporal informative task in which we can obtain information about the ego-motion of a vehicle from the historical information via examining consecutive frames. Sufficient temporal information helps to reduce the search space of the next location. Hence, both efficiency and accuracy of the localization system can be enhanced. This paper presents a semi-supervised deep vision-based localization algorithm, using a novel tubing strategy to find the starting location of a vehicle. We group different number of consecutive frames as sets of tubes based on their temporal correlation to achieve pair searching with variable tube sizes. We also enhance an off-the-shelf network model with our modified training data generation method to improve the discrimination power of the features given by the model. Experimental results show that our proposed temporal correlation based initialization module can confidently localize the starting location of a vehicle (for a certain journey), and achieve 40% precision improvement over that of the conventional CNN approaches.
Chu-Tak Li, Wan-Chi Siu, Daniel Pak-Kong Lun
ICIP3
2019 Enhancement of a CNN-Based Denoiser Based on Spatial and Spectral Analysis
abstract
Convolutional neural network (CNN)-based image denoising methods have been widely studied recently, because of their high-speed processing capability and good visual quality. However, most of the existing CNN-based denoisers learn the image prior from the spatial domain, and suffer from the problem of spatially variant noise, which limits their performance in real-world image denoising tasks. In this paper, we propose a discrete wavelet denoising CNN (WDnCNN), which restores images corrupted by various noise with a single model. Since most of the content or energy of natural images resides in the low-frequency spectrum, their transformed coefficients in the frequency domain are highly imbalanced. To address this issue, we present a band normalization module (BNM) to normalize the coefficients from different parts of the frequency spectrum. Moreover, we employ a band discriminative training (BDT) criterion to enhance the model regression. We evaluate the proposed WDnCNN, and compare it with other state-of-the-art denoisers. Experimental results show that WDnCNN achieves promising performance in both synthetic and real noise reduction, making it a potential solution to many practical image denoising applications.
Rui Zhao 0012, Kin-Man Lam 0001, Daniel Pak-Kong Lun
ICIP3
2019 Deep Learning Based Period Order Detection in Structured Light Three-Dimensional Scanning
abstract
Fringe projection profilometry (FPP) is a popular optical 3-dimensional (3D) scanning method. However, existing FPP methods often suffer from the ambiguity problem that only the wrapped phase information can be measured while the true phase information is required for 3D measurement. Although various phase unwrapping methods were suggested to recover the wrapped phase in FPP methods, most of them will fail when the target objects have complex structures. To solve this problem, we propose in this paper to embed the fringe pattern with a set of textural patterns to encode the period order of the true phase information. During the offline phase, a convolutional neural network (CNN) is trained to learn a set of filters that will be activated when they see the code patterns. When the encoded fringe image is captured, the modified morphological component analysis is first performed to extract the code pattern. It is then decoded by the trained CNN to estimate the K-map, which contains the period order of the true phase information. Experimental results show that the proposed method can measure the 3D profile of objects with abrupt jumps in height profile, where the conventional approaches often fail to perform. It also has a much higher computational efficiency due to the effective utilization of GPU by CNN.
Budianto, Wicky Law, Daniel Pak-Kong Lun
ISCAS3
2019 Single-Image Reflection Removal via a Two-Stage Background Recovery Process
abstract
The reflection problem often occurs when imaging through a semitransparent material such as glass. It degrades the image quality and affects the subsequent analyses on the image. Traditional single-image based reflection removal methods assume the reflection is blurry. Deep neural networks (DNNs) are, then, used to identify the blurry reflection and remove it. However, it is often that the blurry reflection still contains strong edges. They will be treated as the background and kept in the image. In this letter, we propose a novel two-stage DNN based reflection removal algorithm. In the first stage, we include a new feature reduction term in the loss function when training the network. Due to its strong reflection suppression ability, the reflection components in the image can be more effectively suppressed. However, it will also attenuate the gradient values of the background image. For recovering the background, in the second stage, we first estimate a reflection gradient confidence map based on the initial estimation result and use it to identify the strong background gradients. Then, we use a generative adversarial network to reconstruct the background image from its gradients. Experimental results show that the proposed two-stage approach can give a superior performance compared with the state-of-the-art DNN based methods.
Tingtian Li, Daniel Pak-Kong Lun
IEEE Signal Process. Lett.2
2019 Robust Reflection Removal Based on Light Field Imaging
abstract
In daily photography, it is common to capture images in the reflection of an unwanted scene. This circumstance arises frequently when imaging through a semi-reflecting material such as glass. The unwanted reflection will affect the visibility of the background image and introduce ambiguity that perturbs the subsequent analysis on the image. It is a very challenging task to remove the reflection of an image since the problem is severely ill-posed. In this paper, we propose a novel algorithm to solve the reflection removal problem based on light field (LF) imaging. For the proposed algorithm, we first show that the strong gradient points of an LF epipolar plane image (EPI) are preserved after adding to the EPI of another LF image. We can then make use of these strong gradient points to give a rough estimation of the background and reflection. Rather than assuming that the background and reflection have absolutely different disparity ranges, we propose a sandwich layer model to allow them to have common disparities, which is more realistic in practical situations. Then, the background image is refined by recovering the components in the shared disparity range using an iterative enhancement process. Our experimental results show that the proposed algorithm achieves superior performance over traditional approaches both qualitatively and quantitatively. These results verify the robustness of the proposed algorithm when working with images captured from real-life scenes.
Tingtian Li, Daniel Pak-Kong Lun, Yuk-Hee Chan, Budianto
IEEE Trans. Image Process.2
2018 A Novel Reflection Removal Algorithm Using the Light Field Camera
abstract
In daily photography, it is common that the captured images are superimposed with undesired reflection of another scene. Such reflection does not only reduce the visual quality of the target background scene, but also affects the subsequent processing on the image. Removing the reflection without any prior information of the background is a very challenging task. Traditional approaches often come with various assumptions, which are difficult to fulfil in general. In this paper, we propose a novel method to remove the reflection based on light field (LF) imaging. Via the epipolar plane image, we firstly devise a method to retrieve the depth map of the scene from an LF image. According to the depth value, the gradient points of the image are classified to belong to the background, reflection or shared layer. Finally, the background image is reconstructed from the estimated gradient points in the background and shared layers using a sparse optimization process. The proposed algorithm does not have the assumptions of the existing methods; thus, it is more robust. Experimental results show that the proposed algorithm outperforms the existing approaches both qualitatively and quantitatively.
Tingtian Li, Daniel Pak-Kong Lun
ISCAS2
2018 Robust Single-Shot Fringe Projection Profilometry Based on Morphological Component Analysis
abstract
In a fringe projection profilometry (FPP) process, the captured fringe images can be modeled as the superimposition of the projected fringe patterns on the texture of the objects. Extracting the fringe patterns from the captured fringe images is an essential procedure in FPP; but traditional single-shot FPP methods often fail to perform if the objects have a highly textured surface. In this paper, a new single-shot FPP algorithm which allows the object texture and fringe pattern to be estimated simultaneously is proposed. The heart of the proposed algorithm is an enhanced morphological component analysis (MCA) tailored for FPP problems. Conventional MCA methods which use a uniform threshold in an iterative optimization process are inefficient to separate fringe-like patterns from image texture. We extend the conventional MCA by taking advantage of the low-rank structure of the fringe's sparse representation to enable an adaptive thresholding process. It ends up with a robust single-shot FPP algorithm that can extract the fringe pattern even if the object has a highly textured surface. The proposed approach has a side benefit that the object texture can be simultaneously obtained in the fringe pattern estimation process, which is useful in many FPP applications. Experimental results have demonstrated the improved performance of the proposed algorithm over the conventional single-shot FPP approaches.
Budianto, Daniel Pak-Kong Lun, Yuk-Hee Chan
IEEE Trans. Image Process.2
2016 Super-resolution imaging with occlusion removal using a camera array
abstract
In this paper, a novel algorithm which combines the super-resolution imaging and occlusion removal into a single and automatic procedure is proposed. By utilizing the visual parallax of objects at different depths and the sub-pixel information of the images captured by a camera array, we can estimate the shape of the occlusion and reconstruct the background at a higher resolution iteratively. The occlusion shape estimation is achieved by a new method called “seed growth”, which treats the detected feature points of the occlusion as “seeds”. These “seeds” will gradually grow until they reach the occlusion boundary. Experimental results show that the proposed algorithm can well remove the occlusion while super-resolving the background. It performs equally well when there are multiple occlusion objects or the object has irregular shape.
Tingtian Li, Daniel Pak-Kong Lun
ISCAS2
2016 Robust Fringe Projection Profilometry via Sparse Representation
abstract
In this paper, a robust fringe projection profilometry (FPP) algorithm using the sparse dictionary learning and sparse coding techniques is proposed. When reconstructing the 3D model of objects, traditional FPP systems often fail to perform if the captured fringe images have a complex scene, such as having multiple and occluded objects. It introduces great difficulty to the phase unwrapping process of an FPP system that can result in serious distortion in the final reconstructed 3D model. For the proposed algorithm, it encodes the period order information, which is essential to phase unwrapping, into some texture patterns and embeds them to the projected fringe patterns. When the encoded fringe image is captured, a modified morphological component analysis and a sparse classification procedure are performed to decode and identify the embedded period order information. It is then used to assist the phase unwrapping process to deal with the different artifacts in the fringe images. Experimental results show that the proposed algorithm can significantly improve the robustness of an FPP system. It performs equally well no matter the fringe images have a simple or complex scene, or are affected due to the ambient lighting of the working environment.
Budianto, Daniel Pak-Kong Lun
IEEE Trans. Image Process.2
2015 Inpainting for Fringe Projection Profilometry Based on Geometrically Guided Iterative Regularization
abstract
Conventional fringe projection profilometry methods often have difficulty in reconstructing the 3D model of objects when the fringe images have the so-called highlight regions due to strong illumination from nearby light sources. Within a highlight region, the fringe pattern is often overwhelmed by the strong reflected light. Thus, the 3D information of the object, which is originally embedded in the fringe pattern, can no longer be retrieved. In this paper, a novel inpainting algorithm is proposed to restore the fringe images in the presence of highlights. The proposed method first detects the highlight regions based on a Gaussian mixture model. Then, a geometric sketch of the missing fringes is made and used as the initial guess of an iterative regularization procedure for regenerating the missing fringes. The simulation and experimental results show that the proposed algorithm can accurately reconstruct the 3D model of objects even when their fringe images have large highlight regions. It significantly outperforms the traditional approaches in both quantitative and qualitative evaluations.
Budianto, Daniel Pak-Kong Lun
IEEE Trans. Image Process.2
2014 Efficient 3-dimensional model reconstruction based on marker encoded fringe projection profilometry
abstract
This paper presents a novel marker encoded fringe projection profilometry (FPP) scheme for efficient 3-dimensional (3D) model reconstruction. Traditional FPP schemes often have large error when reconstructing 3D model of objects with abruptly changing height profile. In the proposed scheme, markers are encoded in the projected fringe pattern to resolve the ambiguities in the fringe images due to that problem. Using the analytic complex wavelet, the marker cue information can be extracted from the fringe image, and is used to restore the order of the fringes. A series of simulations and experiments have been carried out to verify the proposed scheme. Experimental results show that the proposed method can accurately reconstruct the 3D model of objects with abruptly changing height profile. It is superior to the traditional FPP methods and facilitates real time 3D measurement for color object using only a single fringe image.
Budianto, Daniel Pak-Kong Lun
ICASSP2
2014 Speech enhancement based on L1 regularization in the cepstral domain
abstract
In this paper, a new speech enhancement algorithm using the L1regularization method in the cepstral domain is proposed. Since voiced speeches have a quasi-periodic nature that allows them to be compactly represented in the cepstral domain, the L1regularization technique can be applied to better control the optimization process required in speech enhancement applications. The proposed algorithm starts with the traditional temporal cepstral smoothing (TCS) method which gives the initial estimation of the power spectrum of the clean speech. It is then refined using a modified L1regularizer which imposes further constraint to the penalty function based on the feature of speech signals in the cepstral domain. A notable improvement of the proposed algorithm over the traditional method is its adaptability to the non-stationary noise. Performance of the proposed algorithm is evaluated using standard measures such as segSNR and PESQ based on a large quantity of speech signals. Our results show that a significant improvement is achieved as compared to the conventional approaches especially in the case that the noise is non-stationary.
Tak-Wai Shen, Daniel Pak-Kong Lun
ISCAS2
2014 A Novel Expectation-Maximization Framework for Speech Enhancement in Non-Stationary Noise Environments
abstract
Voiced speeches have a quasi-periodic nature that allows them to be compactly represented in the cepstral domain. It is a distinctive feature compared with noises. Recently, the temporal cepstrum smoothing (TCS) algorithm was proposed and was shown to be effective for speech enhancement in non-stationary noise environments. However, the missing of an automatic parameter updating mechanism limits its adaptability to noisy speeches with abrupt changes in SNR across time frames or frequency components. In this paper, an improved speech enhancement algorithm based on a novel expectation-maximization (EM) framework is proposed. The new algorithm starts with the traditional TCS method which gives the initial guess of the periodogram of the clean speech. It is then applied to an${L_1}$norm regularizer in the M-step of the EM framework to estimate the true power spectrum of the original speech. It in turn enables the estimation of the a-priori SNR and is used in the E-step, which is indeed a logmmse gain function, to refine the estimation of the clean speech periodogram. The M-step and E-step iterate alternately until converged. A notable improvement of the proposed algorithm over the traditional TCS method is its adaptability to the changes (even abrupt changes) in SNR of the noisy speech. Performance of the proposed algorithm is evaluated using standard measures based on a large set of speech and noise signals. Evaluation results show that a significant improvement is achieved compared to conventional approaches especially in non-stationary noise environment where most conventional algorithms fail to perform.
Daniel Pak-Kong Lun, Tak-Wai Shen, K. C. Ho 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Image enhancement for fringe projection profilometry
abstract
The fringe projection profilometry (FPP) is a popular approach for acquiring high-fidelity 3D model of objects. In the FPP, parallel fringe patterns are projected onto an object, and the deformed fringe as shown on the object surface is captured by a camera for 3D reconstruction. Due to the imperfection of imaging devices and operating environment, the captured fringe images however are inevitably to have many artifacts. They severely affect the robustness of the FPP and the quality of the reconstructed 3D model. In this paper, a new approach is proposed that successfully characterizes some common artifacts of fringe images, such as bias and noise, using the dual-tree complex wavelet transform. Effective methods are then devised to remove them from the fringe images. Experimental results show that the proposed algorithm is superior to the traditional methods and facilitates accurate reconstruction of objects' 3D model even with low quality fringe images.
William Wai-Lam Ng, Daniel Pak-Kong Lun
ISCAS2
2012 A new 3-phase design exploration methodology for video processor design
abstract
When making video processor design, conventional design exploration methodologies take extremely long time in parameter optimization but the final design may not necessarily meet the application requirements since the architecture cannot deviate too much from the initial design. To speed up the design process, statistical performance models were used to guide the simulation; however their accuracy is questionable. In this paper, a new 3-phase design exploration methodology for video processor is proposed. It makes use of an almost cycle-accurate performance model to provide information for refining the processor architecture. It can derive the optimal architecture in a much shorter period of time than the conventional methods. We successfully implemented a few video coding/decoding applications on the video processor derived from the proposed methodology. Simulation results show that it outperforms other video processors in both cost and performance perspectives.
Wing-Yee Lo, Daniel Pak-Kong Lun, Wan-Chi Siu
ISCAS2
2012 Improved speech presence probability estimation based on wavelet denoising
abstract
A reliable estimator for speech presence probability (SPP) can significantly improve the performance of many speech enhancement algorithms. Previous work showed that a good SPP estimator can be obtained by using a smooth a-posteriori signal to noise ratio (SNR) function, which can be achieved by reducing the noise variance when estimating the speech power spectrum. In this paper, a wavelet based denoising algorithm is proposed for such purpose. We first apply the wavelet transform to the periodogram of a noisy speech signal to generate an oracle for indicating the locations of the noise floor in the periodogram. We then make use of that oracle to selectively remove the wavelet coefficients of the noise floor in the log multitaper spectrum (MTS) of the noisy speech. The remaining wavelet coefficients are then used to reconstruct a denoised MTS and in turn generate a smooth a-posteriori SNR function. Simulation results show that the new SPP estimator outperforms the traditional approaches and enables a significantly improvement in the quality and intelligibility of the enhanced speeches.
Daniel Pak-Kong Lun, Tak-Wai Shen, Richard T. C. Hsung, K. C. Ho 0001
ISCAS1
2011 Zero spectrum removal using joint bilateral filter for Fourier transform profilometry
abstract
Traditional Fourier transform profilometry (FTP) removes the zero spectrum of the acquired fringe image using the linear filtering approach. Such approach assumes there is no aliasing between the zero spectrum and the higher harmonics of the image, which however is not true in general. Thus it cannot adapt to sharp illuminant changes in the fringe image. One practical solution is to exploit one more fringe pattern with π phase shift to eliminate the zero spectrum. This however complicates the system in that the projection of the fringe patterns should be synchronized with the capturing device. Motion or imperfect synchronization makes the subsequent 3D reconstruction erroneous. In this paper, we introduce a 2D joint bilateral filter to help in removing the zero spectrum. Since the bilateral filter is sensitive to abrupt changes in the fringe image, we can accurately extract the zero spectrum even if there are sharp spatial variations in the pixel magnitude. Experimental results show that the proposed algorithm is superior to the traditional methods and facilitate accurate reconstruction of objects' 3D model.
Richard T. C. Hsung, Daniel Pak-Kong Lun, William Wai-Lam Ng
VCIP2
2011 Improved SIMD Architecture for High Performance Video Processors
abstract
Single instruction multiple data (SIMD) execution is in no doubt an efficient way to exploit the data level parallelism in image and video applications. However, SIMD execution bottlenecks must be tackled in order to achieve high execution efficiency. We first analyze in this paper the implementation of two major kernel functions of H.264/AVC namely, SATD and subpel interpolation, in conventional SIMD architectures to identify the bottlenecks in traditional approaches. Based on the analysis results, we propose a new SIMD architecture with two novel features: 1) parallel memory structure with variable block size and word length support, and 2) configurable SIMD structure. The proposed parallel memory structure allows great flexibility for programmers to perform data access of different block sizes and different word lengths. The configurable SIMD structure allows almost “random” register file access and slightly different operations in ALUs inside SIMD. The new features greatly benefit the realization of H.264/AVC kernel functions. For instance, the fractional motion estimation, particularly the half to quarter pixel interpolation, can now be executed with minimal or no additional memory access. When comparing with the conventional SIMD systems, the proposed SIMD architecture can have a further speedup of 2.1X to 4.6X when implementing H.264/AVC kernel functions. Based on Amdahl's law, the overall speedup of H.264/AVC encoding application can be projected to be 2.46X. We expect significant improvement can also be achieved when applying the proposed architecture to other image and video processing applications.
Wing-Yee Lo, Daniel Pak-Kong Lun, Wan-Chi Siu, Jiqiang Song
IEEE Trans. Circuits Syst. Video Technol.2
2010 On optical phase shift profilometry based on dual tree complex wavelet transform
abstract
In optical phase shift profilometry, parallel fringe patterns are projected onto an object and the deformed fringes are captured using a digital camera. It is of particular interest because it enables reconstruction of the 3D shape of the object using just a few image captures, which facilitates real time applications. However, when using the approach in real life environment, it is noticed that the noise in the captured images can greatly affect the reconstruction quality. In this paper, we firstly analyze why the noisy fringe images can best be analyzed using the oriented 2D dual tree complex wavelet transform. We then suggest an effective yet simple method for enhancing the noisy fringe images. Both the simulation and experiment results show that the new approach can give good performance in reconstruction with fringe images even at high noise level.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICIP2
2010 Improved wavelet based a-priori SNR estimation for speech enhancement
abstract
To obtain a reliable estimate of the a-priori signal to noise (SNR) ratio is crucial to most frequency domain speech enhancement algorithms. Recently, the low variance multitaper spectrum (MTS) estimator with wavelet denoising was suggested for the estimation of the a-priori SNR However, traditional approach directly plugs in the wavelet shrinkage denoiser and adopts the universal threshold which is not fully optimized to the characteristic of the MTS of noisy signals. In this paper, a two-stage estimation algorithm is proposed. First, the log MTS components that are dominated by noise are detected and removed in the wavelet domain. Second, a modified SUREshrink scheme is applied to further remove the noise remained in the speech spectral peaks. The new estimator is applied to the traditional Wiener filter and log MMSE speech enhancement algorithms and leads to significantly better performance.
Daniel Pak-Kong Lun, Richard T. C. Hsung
ISCAS1
2010 Synchronized partial-body motion graphs
abstract
Motion graphs are regarded as a promising technique for interactive applications. However, the graphs are generated based on the distance metric of whole body, which produce a limit set of possible transitions. In this paper, we present an automatic method to construct a new data structure that specifies transitions and correlations between partial-body motions, called Synchronized Partial-body Motion Graphs (SPbMGs). We exploit the similarity between lower-body motions to create synchronization conditions with upper-body motions. Under these conditions, we generate all possible transitions between partial-body motions. The proposed graph representation not only maximizes the reusability of motion data, but also increases the connectivity of motion graphs while retaining the quality of motion.
William Wai-Lam Ng, Clifford S. T. Choy, Daniel Pak-Kong Lun, Lap-Pui Chau
SIGGRAPH ASIA (Sketches)3
2008 Speech enhancement based on adaptive wavelet denoising on multitaper spectrum
abstract
Classical speech enhancement algorithms often require a good estimation of the short-time power spectrum using, for instance, the periodogram methods. However, it is well known that traditional periodogram methods are prone to induce large variance, hence produces the “musical noise” after enhancement. To alleviate this problem, multitaper spectrum (MTS) estimators with wavelet denoising were proposed. In this paper, we investigate the properties of the MTS of noisy speech signals. We find that, in the log MTS domain, the variance of noise varies according to the magnitude of the underlying speech spectrum. It implies that when applying wavelet denoising to the log MTS, the constant threshold used in the traditional methods is not appropriate. Based on this observation, we further develop a wavelet denoising method with adaptive threshold for estimating power spectrum using multitaper. Simulation results show that the spectrum estimated using the proposed method is consistently more accurate than the traditional uniform thresholding methods. Hence, it further improves the current speech enhancement algorithms using the MTS approaches.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ISCAS2
2007 Denoising for Generalized Sidelobe Canceller
abstract
The Generalized Sidelobe Canceller (GSC) efficiently realizes the optimal Linearly Constrained Minimum Variance (LCMV) beamformer and delivers excellent beamforming results by reducing directional interference and noise. However, when the input signals are contaminated by other type of noises, such as background and diffused noises, the overall performance of GSC can be even worse than the traditional delay-and-sum beamformers. In this paper, we investigate the application of the spatially adaptive multiwavelet (MWT) denoising technique to the GSC in an environment with severe diffused noise. Comparing with the traditional scalar wavelets, the multiwavelets can better characterize the noise information in the signal such that better denoising performance can be achieved. Different approaches for integrating the GSC and the multiwavelet denoiser were studied. It is found that by adding two denoisers, one to the fixed constrained part and another to the final GSC output, an improvement of 2dB in SNR can be achieved as compared with the traditional GSC method.
Chun-Yat Ma, Richard T. C. Hsung, Daniel Pak-Kong Lun, K. C. Ho 0001, Hon Keung Kwan
ISCAS3
2007 Throughput optimization for video streaming proxy servers based on video staging
Wai Kong Cheuk, Daniel Pak-Kong Lun
Multim. Tools Appl.2
2006 A Generalized Orthogonal Symmetric Prefilter Banks for Discrete Multiwavelet Transforms
abstract
Prefilters are generally applied to the discrete multiwavelet transform (DMWT) for processing scalar signals. To fully utilize the benefit given by the multiwavelets, we have recently shown a maximally decimated orthogonal prefilter which preserves the linear phase property and the approximation power of the multiwavelets. However, such design requires the point of symmetry of each channel of the prefilter to match with the scaling functions of the target multiwavelet system. A compatible filter bank structure can be very difficult to find or simply does not exist, e.g. for multiplicity 2 multiwavelets. In this paper, we suggest a new DMWT structure in which the prefilter is combined with the first stage of DMWT. The advantage of the new structure is twofold: First, the computational complexity can be greatly reduced. Second, additional design freedom allows maximally decimated, orthogonal and symmetric prefilters even for low multiplicity. We evaluated the computational complexity and energy compaction capability of the new DMWT structure. Satisfactory results are obtained in comparing with the traditional approaches.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICIP2
2006 Design and implementation of contractual based real-time scheduler for multimedia streaming proxy server
Wai Kong Cheuk, Richard T. C. Hsung, Daniel Pak-Kong Lun
Multim. Tools Appl.3
2006 Orthogonal symmetric prefilter banks for discrete multiwavelet transforms
abstract
Traditional design of critically sampled prefilters for discrete multiwavelet transform ignores the preservation of the linear phase property, which is important for many applications, such as image coding and digital communications. Balanced multiwavelets solve this problem but make the filters longer. By using linear phase filter banks, we propose a simple algorithm for the design of orthogonal symmetric prefilter banks that can be used with the discrete multiwavelet transform. The prefilter bank resulted is orthogonal and critically sampled and can preserve the approximation power of the multiwavelet as well as the linear phase property. Experimental results show that the systems using the proposed symmetric prefilter banks give better performance as compared with using nonlinear phase prefilters.
Richard T. C. Hsung, Daniel Pak-Kong Lun, K. C. Ho 0001
IEEE Signal Process. Lett.2
2005 Content-based scalable H.263 video coding for road traffic monitoring
abstract
For sending video data through very low bit-rate mobile channels, video codec with high compression rate is the pre-requisite. Although the H.263 video codec is recommended as one of the candidates due to its simplicity and efficiency, it is generally believed that its compression efficiency can be further improved if the content-based scalable video coding technique can be applied. In this paper, we propose a modified H.263 encoder which supports real-time content-based scalable video coding. The proposed technique is applied to real-time video surveillance systems for road traffic monitoring. For the proposed approach, the moving objects, i.e. cars, are first extracted from the steady background. Their activities are then further classified as fast or slow by assessing the regularity of their motion. The information is then passed to a modified H.263 encoder to reduce the temporal and spatial redundancies in the video. As compared with the conventional H.263 encoder using for the same application, the proposed system has a 20% increase in compression rate with negligible visual distortion. The proposed system fully complies with the ITU H.263 standard hence the encoded bit stream is completely comprehensible to the conventional H.263 decoder.
Wallace Kai-Hong Ho, Wai Kong Cheuk, Daniel Pak-Kong Lun
IEEE Trans. Multim.3
2004 On optimal threshold selection for multiwavelet shrinkage [signal denoising applications]
abstract
Recent research found that multivariate shrinkage on multiwavelet transform coefficients further improves the traditional wavelet methods. It is because the multiwavelet transform, with appropriate initialization, provides better representation of signals so that their difference from noise can be clearly identified. In this paper, we consider the optimal threshold selection for multiwavelet denoising by using a multivariate shrinkage function. Firstly, we study the threshold selection using the Stein's unbiased risk estimator (SURE) for each resolution level when the noise structure is given. Then, we consider the method of generalized cross validation (GCV) when the noise structure is not known a priori. Simulation results show that the higher multiplicity (>2) wavelets usually give better denoising results. Besides, the proposed threshold estimators often suggest better thresholds as compared with the traditional estimators.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICASSP (2)2
2004 Video staging in video streaming proxy server
abstract
Video staging has been proposed as a mechanism for video streaming proxy servers to better exploit the temporal locality of requests made by clients. It divides a video stream into two parts, i.e. local storage and remote access. It opens a new possibility of scheduling in the video streaming proxy server, with the introduction of disk scheduling on top of network scheduling. However, a simple division cannot put all these resources into good use. We devise a novel video staging algorithm. It considers memory consumption, network bandwidth and disk access requirements so that these resources usage can be optimized accordingly.
Wai Kong Cheuk, Daniel Pak-Kong Lun
ICME2
2004 Precise Modeling of Design Patterns in UML
abstract
Prior research attempts to formalize the structure of object-oriented design patterns for a more precise specification of design patterns. It also allows automation support to be developed for user-defined design patterns in the future CASE tools. Targeting to a particular type of automation (e.g. verification of pattern instances), previous specification approaches over-specify pattern structures to a certain extend. Over-specification makes pattern specification ambiguous and disallows the specification language to be used for specifying compound patterns. In this paper, we present the structural properties of design patterns which reveal the true abstract nature of pattern structures. To support these properties so as to solve the over-specification problem, we propose an extension to UML 1.5 (basically UML 1.4 with Action semantics). The specialization and refining mechanism of UML provides also a smooth support for the instantiation, refinement and integration of pattern structures specified in UML. Our work makes no significant extension to the UML 1.5 meta-model but more in a UML Profile approach to ease the migration of our work to UML 2.0, which has not yet officially released by OMG during this work.
Jeffrey Ka-Hing Mak, Clifford S. T. Choy, Daniel Pak-Kong Lun
ICSE3
2004 Generalized cross validation for multiwavelet shrinkage
abstract
Traditional multiwavelet shrinkage denoising techniques require a priori knowledge of noise variance that may not be obtained in some practical situations. By using generalized cross validation (GCV), we propose in this paper a new level-dependent risk estimator for multiwavelet shrinkage that does not require such a priori information. Simulation results verify that the resulted risk estimator gives better indication on threshold selection comparing with the traditional GCV method. Improved denoising performance is then achieved particularly for higher multiplicity multiwavelet shrinkage.
Richard T. C. Hsung, Daniel Pak-Kong Lun
IEEE Signal Process. Lett.2
2004 Efficient blind image restoration using discrete periodic Radon transform
abstract
Restoring an image from its convolution with an unknown blur function is a well-known ill-posed problem in image processing. Many approaches have been proposed to solve the problem and they have shown to have good performance in identifying the blur function and restoring the original image. However, in actual implementation, various problems incurred due to the large data size and long computational time of these approaches are undesirable even with the current computing machines. In this paper, an efficient algorithm is proposed for blind image restoration based on the discrete periodic Radon transform (DPRT). With DPRT, the original two-dimensional blind image restoration problem is converted into one-dimensional ones, which greatly reduces the memory size and computational time required. Experimental results show that the resulting approach is faster in almost an order of magnitude as compared with the traditional approach, while the quality of the restored image is similar.
Daniel Pak-Kong Lun, Tommy C. L. Chan, Richard T. C. Hsung, David Dagan Feng, Yuk-Hee Chan
IEEE Trans. Image Process.1
2003 Impact of Computational Resource Reservation to the Communication Performance in the Hypercluster Environment
abstract
Emerging grid computing enables researchers, scientists and engineers to build so called hypercluster, cluster of cluster, from the available resources of the grid. However, hypercluster building from the resources of the grid inherits the unpredictable resource behaviors, which result in variable performance, even using the same set of the resources but in different time. This paper proposes a design of a reservation aware operating system, RAOS, which is able to stabilize the performance of applications running on the unpredictable grid environment. We show this by stabilizing the communication cost (connection setup time, message packing time), which is one of the important parameters required to optimize in parallel applications and distributed applications, under the loaded computation node in a cluster computer.
Jim Kai-Wing Tse, Daniel Pak-Kong Lun
CLUSTER2
2003 Precise Specification to Compound Patterns with ExLePUS
abstract
Prior research suggested modeling languages for precise specification of design pattern structures and behaviors. However, seldom has put effort on their integrations as well as their specifications. To provide a first class CASE support to the recognition, verification and application of design patterns as well as their compounds, a precise specification to their leitmotifs is critical. In this paper, we present the essentials of pattern integration and propose an extended version, we name it exLePUS, to a pattern specification language (LePUS) in order to support the specification of these essentials and thus compound patterns. A case study has illustrated how it is used a well-known compound pattern.
Jeffrey Ka-Hing Mak, Clifford S. T. Choy, Daniel Pak-Kong Lun
COMPSAC3
2003 Orthogonal discrete periodic Radon transform. Part I: theory and realization
Daniel Pak-Kong Lun, Richard T. C. Hsung, Tak-Wai Shen
Signal Process.1
2003 Orthogonal discrete periodic Radon transform. Part II: applications
Daniel Pak-Kong Lun, Richard T. C. Hsung, Tak-Wai Shen
Signal Process.1
2003 Improved MPEG-4 still texture image coding under noisy environment
abstract
This paper describes the performance of the MPEG-4 still texture image codec in coding noisy images. As will be shown, when using the MPEG-4 still texture image codec to compress a noisy image, increasing the compression rate does not necessarily imply reducing the peak-signal-to-noise ratio (PSNR) of the decoded image. An optimal operating point having the highest PSNR can be obtained within the low bit rate region. Nevertheless, the visual quality of the decoded noisy image at this optimal operating point is greatly degraded by the so-called "cross" shape artifact. In this paper, we analyze the reason for the existence of the optimal operating point and the "cross" shape artifact when using the MPEG-4 still texture image codec to compress noisy images. We then propose an adaptive thresholding technique to remove the "cross" shape artifact of the decoded images. It requires only a slight modification to the quantization process of the traditional MPEG-4 encoder while the decoder remains unchanged. Finally, an analytical study is performed for the selection and validation of the threshold value used in the adaptive thresholding technique. It is shown that, the visual quality and PSNR of the decoded images are much improved by using the proposed technique comparing with the traditional MPEG-4 still texture image codec in coding noisy images.
Tommy C. L. Chan, Richard T. C. Hsung, Daniel Pak-Kong Lun
IEEE Trans. Image Process.3
2002 Adaptive wavelet regularity scalable image coding
abstract
An adaptive regularity scalable wavelet image coding algorithm is proposed in this paper. The bitstream is generated in the order of regularity such that visually more important components of a decoded image can be obtained first, followed by the higher regularity components as more bits are received and decoded. The scalability is achieved by selecting different extents of wavelet coefficients at various regularity levels. These regularity levels are determined adaptively. Regularity of the image is estimated from the interscale ratios of the separable wavelet transform magnitude sums. Compared to the wavelet image coder generating resolution scalable bitstreams, significant improvement in visual quality and PSNR can be achieved.
Charlotte Yuk-Fan Ho, Richard T. C. Hsung, Daniel Pak-Kong Lun
ICASSP3
2002 Boundary filters design for multiwavelets
abstract
In this paper we consider the optimization of the boundary filters with perfect reconstruction property and moment conditions for multiwavelet, Since the filters coefficients are matrices, the moment conditions for multiwavelet do not relate to the filters coefficients in a simple way. We propose a scheme to formulate the moment conditions for the optimization of a given multiwavelet. We apply the proposed scheme on GHM multiwavelet and obtain a set of boundary filters, which possess the desired moment conditions and reduce the boundary artifacts up to 75% as compared to the truncated filters.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICASSP2
2001 Auditory model based speech recognition in noisy environment
abstract
The main purpose of this paper is to present how to raise the speech recognition performance in noisy environment. So far the most popularly used speech feature in speech recognition is probably the so-called MFCC. The recognition rate of speech recognition algorithm using MFCC and CDHMM is known to be very high in clean speech environment, but it deteriorates greatly in noisy environment, especially in the white noisy environment. In this paper, we propose a new speech feature, the ASBF speech feature based on the mathematical model of inner ear of human auditory system. This new speech feature is extracted using both mathematical model of inner ear and primary auditory nerve processing model of human auditory system, and it can track the speech formants effectively. In the experiment, the performance of MFCC and the ASBF are compared in both clean and noisy environments when using left-to-right CDHMM with 6 states and 5 Gaussian mixtures. The experimental result shows that the ASBF is much more robust to noise than MFCC. When only 5 dimension is used in ASBF vector, the recognition rate is approximately 38.6% higher than the traditional MFCC with 39 dimension in the condition of S/N=10dB with white noise.
Xiaoqing Yu, Daniel Pak-Kong Lun
INTERSPEECH3
2000 Adaptive Thresholding for Noisy MPEG-4 Still Texture Image
abstract
We show that when using the MPEG-4 still image codec to compress a noisy image, increasing the compression ratio does not necessarily imply reducing the peak-signal-to-noise ratio (PSNR) of the decoded image. An optimal operating point is noted where the peak PSNR can be obtained within the low bit rate region. Yet, the visual quality of the decoded noisy images in this region is greatly degraded by the so-called "cross" shape artefact. To deal with this problem, we analyze the cause of the optimal operating point and the "cross" shape artefact when using the MPEG-4 still image codec for coding noisy images. We then propose an adaptive thresholding technique to remove the "cross" shape artefact. An analytical study is performed for the selection and validation of the threshold value used. It is shown that, by using the proposed technique, the visual quality and PSNR of the decoded images are much improved compared with the traditional MPEG-4 still image codec in coding noisy images.
Tommy C. L. Chan, Daniel Pak-Kong Lun
ICIP2
1999 Embedded Singularity Detection Zerotree Wavelet Coding
abstract
We explore the wavelet coefficient selection and denoising by singularity detection (SD) for the embedded zero-tree wavelet (EZW) coding algorithm in this paper. The EZW coding algorithm exploits the relation between the multi-scale wavelet coefficients that finer scale wavelet coefficients are probably to vanish if the coarse scale wavelet coefficient vanishes. It is true for the parts of image that are regular but not the cases for noise-like features. In other words, the performance of the coding algorithm may be greatly degraded for the latter. In this paper, we investigate to arrange the wavelet coefficients according to the local regularity, by using the computed wavelet coefficients from the encoding filters. The advantage is two folds. For normal coding, we can make the encoded bit-stream first appear with the wavelet coefficients that correspond to the most regular part of the image, and the irregular one's follows. For noisy image encoding, we can remove noises before encoding hence increase the image quality as well as the coding efficiency.
Daniel Pak-Kong Lun, Richard T. C. Hsung, David Dagan Feng, Tommy C. L. Chan
ICIP (2)1
1998 Image deblocking by singularity detection
abstract
The blocking effect is considered as the most disturbing artifact of JPEG decoded images. Many researchers have suggested various methods to tackle this problem. The wavelet transform modulus maxima (WTMM) approach was proposed which gives a significant improvement over the previous methods in terms of the signal-to-noise ratio and visual quality. However, the WTMM deblocking algorithm is an iterative algorithm that requires a long computation time to reconstruct the processed WTMM to obtain the deblocked image. A new wavelet based algorithm for JPEG image deblocking is proposed. The new algorithm is based on the idea that, besides using the WTMM, the singularity of an image can also be detected by computing the sums of the wavelet coefficients inside the so-called "directional cone of influence" in different scales of the image. The new algorithm has the advantage as the WTMM approach that it can effectively identify the edge and the smooth regions of an image irrespective of the discontinuities introduced by the blocking effect. It is an improvement over the WTMM approach in that only a simple inverse wavelet transform is required to reconstruct the processed wavelet coefficients to obtain the deblocked image. As the WTMM approach, the new algorithm gives consistent and significant improvement over the previous methods for JPEG image deblocking.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICASSP2
1998 Non-invasive quantification of physiological processes with dynamic PET using blind deconvolution
abstract
Dynamic positron emission tomography (PET) has opened the possibility of quantifying physiological processes within the human body. On performing dynamic PET studies, the tracer concentration in blood plasma has to be measured, and acts as the input function for tracer kinetic modelling. In this paper, we propose an approach to estimate physiological parameters for dynamic PET studies without the need of taking blood samples. The proposed approach comprises two major steps. First, a wavelet denoising technique is used to filter the noise appeared in the projections. The denoised projections are then used to reconstruct the dynamic images using filtered backprojection. Second, an eigen-vector based blind deconvolution technique is applied to the reconstructed dynamic images to estimate the physiological parameters. To demonstrate the performance of the proposed approach, we carried out a Monte Carlo simulation using the fluoro-deoxy-2-glucose model, as applied to tomographic studies of human brain. The results demonstrate that the proposed approach can estimate the physiological parameters with an accuracy comparable to that of invasive approach which requires the tracer concentration in plasma to be measured.
Chi-Hoi Lau, Daniel Pak-Kong Lun, David Dagan Feng
ICASSP2
1998 On the efficient computation of 2-d image moments using the discrete radon transform
Tak-Wai Shen, Daniel Pak-Kong Lun, Wan-Chi Siu
Pattern Recognit.2
1998 A deblocking technique for block-transform compressed image using wavelet transform modulus maxima
abstract
In this work, we introduce a deblocking algorithm for Joint Photographic Experts Group (JPEG) decoded images using the wavelet transform modulus maxima (WTMM) representation. Under the WTMM representation, we can characterize the blocking effect of a JPEG decoded image as: 1) small modulus maxima at block boundaries oversmooth regions; 2) noise or irregular structures near strong edges; and 3) corrupted edges across block boundaries. The WTMM representation not only provides characterization of the blocking effect, but also enables simple and local operations to reduce the adverse effect due to this problem. The proposed algorithm first performs a segmentation on a JPEG decoded image to identify the texture regions by noting that their WTMM have small variation in regularity. We do not process the modulus maxima of these regions, to avoid the image texture being "oversmoothed"by the algorithm. Then, the singularities in the remaining regions of the blocky image and the small modulus maxima at block boundaries are removed. We link up the corrupted edges, and regularize the phase of modulus maxima as well as the magnitude of strong edges. Finally,the image is reconstructed using the projection onto convex set (POCS)technique on the processed WTMM of that JPEG decoded image.This simple algorithm improves the quality of a JPEG decoded image inthe senses of signal-to-noise ratio (SNR) as well as visual quality. We also compare the performance of our algorithm to the previous approaches,such as CLS and POCS methods. The most remarkable advantage of the WTMM deblocking algorithm is that we can directly process the edges and texture of an image using its WTMM representation.
Richard T. C. Hsung, Daniel Pak-Kong Lun, Wan-Chi Siu
IEEE Trans. Image Process.2
1998 Minimum Dynamic SPECT Image Acquisition Time Required for T1-201 Tracer Kinetic Modelling
Chi-Hoi Lau, Stefan Eberl, David Dagan Feng, Hidehiro Iida, Daniel Pak-Kong Lun, Wan-Chi Siu, Yoshikazu Tamura, George J. Bautovich, Yukihiko Ono
IEEE Trans. Medical Imaging5
1998 Dynamic Imaging and Tracer Kinetic Modeling for Emission Tomography Using Rotating Detectors
abstract
When performing dynamic studies using emission tomography the tracer distribution changes during acquisition of a single set of projections. This is particularly true for some positron emission tomography (PET) systems which, like single photon emission computed tomography (SPECT), acquire data over a limited angle at any time, with full projections obtained by rotation of the detectors. In this paper, an approach is proposed for processing data from these systems, applicable to either PET or SPECT. A method of interpolation, based on overlapped parabolas, is used to obtain an estimate of the total counts in each pixel of the projections for each required frame-interval, which is the total time to acquire a single complete set of projections necessary for reconstruction. The resultant projections are reconstructed using traditional filtered backprojection (FBP) and tracer kinetic parameters are estimated using a method which relies on counts integrated over the frame-interval rather than instantaneous values. Simulated data were used to illustrate the technique's capabilities with noise levels typical of those encountered in either PET or SPECT. Dynamic datasets were constructed, based on kinetic parameters for fluoro-deoxy-glucose (FDG) and use of either a full ring detector or rotating detector acquisition. For the rotating detector, use of the interpolation scheme provided reconstructed dynamic images with reduced artefacts compared to unprocessed data or use of linear interpolation. Estimates for the metabolic rate of glucose had similar bias to those obtained from a full ring detector.
Chi-Hoi Lau, David Dagan Feng, Brian F. Hutton, Daniel Pak-Kong Lun, Wan-Chi Siu
IEEE Trans. Medical Imaging4
1997 Region-of-interest tomography using multiresolution interpolation
abstract
The wavelet localization technique has been applied to region-of-interest tomography. It achieves a significant saving in the required projections if only a small region of a tomographic image is of interest. We first show that, with the same sampling requirement, a simple interpolation scheme applied on the samples can give a result at least as good as that achieved by using the wavelet localization approach. It means that we can use a much simpler approach to achieve the same performance. Second, we propose a new sampling scheme such that the required projections of each angle are further reduced in a multiresolution form. With this sampling scheme, more than 84% of projections are saved to reconstruct a 32/spl times/32 pixels region of a 256/spl times/256 pixels image. The signal-to-error ratio of the reconstructed region-of-interest is over 50 dB as compare with the case of full projection. Moreover, we also investigate the effect of applying the interlaced sampling scheme on the proposed method. It is seen that a further reduction in the sampling requirement can be achieved although a slight decrease in signal-to-error ratio may result.
Daniel Pak-Kong Lun, Richard T. C. Hsung
ICASSP1
1996 Fast algorithm for 2-D image moments via the Radon transform
abstract
A fast algorithm for the computation of the two-dimensional image moments is proposed. In our approach, a new discrete Radon transform (DRT) is used for the major part of the algorithm. The new DRT preserves an important property of the continuous Radon transform that the regular or geometric moments can be directly obtained from the projection data. With this property, the computation of a two-dimensional (2-D) image moments can be decomposed as a number of one-dimensional (1-D) ones, hence greatly reducing the computational complexity. Comparisons of the computational complexity and performance with some known methods are also given. It is shown that the proposed algorithm significantly reduces the complexity and computation time.
Tak-Wai Shen, Daniel Pak-Kong Lun, Wan-Chi Siu
ICASSP2
1996 A deblocking technique for JPEG decoded image using wavelet transform modulus maxima representation
abstract
In this paper, we introduce a local deblocking algorithm for JPEG decoded images using the wavelet transform modulus maxima (WTMM) representation. Under the WTMM representation, we can characterize the blocking effect as: 1) small modulus maxima at block boundaries over smooth regions; 2) noises or irregular structures near strong edges; 3) corrupted edges across block boundaries. The WTMM representation not only provides characterization of the blocking effect, but also enables simple and local operations on those singularities. The proposed algorithm first performs a segmentation to discriminate the texture regions of an image based on the WTMM local regularity variance. We then keep the modulus maxima of these regions, which are of low regularity variation, unchange to avoid the image texture being "over-smoothed" by the algorithm. Then, the singularities on the remaining regions of the blocky image and small modulus maxima at block boundaries are removed. Then, we link up the corrupted edges and regularize the phase of modulus maxima as well as the amplitude of strong edges. Finally, the image is reconstructed using the projection onto convex sets (POCS) technique on the processed WTMM of the JPEG decoded image. This simple algorithm improves the quality of JPEG decoded image in the sense of signal to noise ratio as well as visual quality. We also compare the performance of our algorithm with the previous approaches and show the superiority over them. The most remarkable advantage of the WTMM deblocking algorithm is that it incorporates direct edges and texture operations into the WTMM representation.
Richard T. C. Hsung, Daniel Pak-Kong Lun, Wan-Chi Siu
ICIP (2)2
1995 The wavelet transform of higher dimension and the Radon transform
abstract
Presents a fast algorithm for the computation of the wavelet transform in higher dimensional Euclidean space R/sup n/ with arbitrary shaped wavelets. The algorithm is a direct consequence of the convolution property of the Radon transform and shows significant improvement in speed. The authors also present a novel approach for the computation of the Daubechies type wavelet transform under the Radon transform domain where the n-dimensional multiresolution analysis (MRA) is reduced to one-dimensional MRA. They found applications of this approach on, for instance, multiresolution reconstruction of a tomographic image with the standard methods of denoising, where determination of wavelet coefficients is required under the Radon transform domain. Along with the possibility of reducing samples angularly with decreasing resolution, the efficiency can be further improved. Also, extra properties such as the "rotated" wavelet can be easily implemented with this algorithm.
Richard T. C. Hsung, Daniel Pak-Kong Lun
ICASSP2
1995 On the Convolution Property of a New Discrete Radon Transform and its Efficient Inversion Algorithm
Daniel Pak-Kong Lun, Richard T. C. Hsung, Wan-Chi Siu
ISCAS1
1994 On efficient software realization of the prime factor discrete cosine transform
abstract
The traditional approach in realizing the prime factor discrete cosine transform (PFDCT) often suffers from two problems. First, although only the Ruritanian mapping is used for input indexing, it requires to perform a series of complicated tests and additions which even outweigh the computational effort of the PFDCT. Second, the additions mentioned above are not carried out in an in-place form. This implies that an auxiliary data array is required to buffer the temporary results generated during the additions. Otherwise, erroneous results will be obtained. We propose an efficient indexing scheme for the computation of the PFDCT. By suitably swapping the data, all the additions can be carried out in an in-place form. Furthermore the number of tests required to perform on the indices of the data is greatly reduced. They are achieved by considering the special properties of the Ruritanian mapping.>
Daniel Pak-Kong Lun
ICASSP (3)1
1994 A Pipeline Design for the Realization of the Prime Factor Algorithm Using the Extended Diagonal Structure
abstract
In this brief contribution, an efficient pipeline architecture is proposed for the realization of the Prime Factor Algorithm (PFA) for digital signal processing. By using the extended diagonal feature of the Chinese Remainder Theorem (CRT) mapping, we show that the input data sequence can be directly loaded into a multidimensional array for the PFA computation without any permutation. Short length modules are modified such that an in-place and in-order computation is allowed. The computed results can then be directly restored back to the memory array without the need for further reordering. More importantly, the CRT mapping can also be used to represent the output data, hence we can utilize the extended diagonal feature of the CRT mapping to directly send the computed results to the outside world. As compared to the previous approaches, the present approach requires no shifting or rotation during the data loading and retrieval processes. In the case of multidimensional PFA computation, it does not require the computation to be split up into a number of two-dimensional computations. Hence, the overhead required for data loading and retrieval in each two-dimensional stage can be saved.
Daniel Pak-Kong Lun, Wan-Chi Siu
IEEE Trans. Computers1
1992 On inherent in-place and in-order features of the prime factor algorithm
abstract
For the computation of the prime factor algorithm (PFA), an in-place and in-order approach is always desirable because it reduces the memory requirement for the storage of the temporary results, and the computation time which is required to unscramble the output sequence to a proper order. In fact, the processing time required for this unscrambling process can take up as much as 50% of the overall computation time. It is shown that the PFA has an intrinsic property that allows it to be easily realized in an in-place and in-order form. No extra operation is required as in the previous propositions. Nevertheless, the sequence length of the PFA computation must be carefully selected. The conditions under which a particular sequence length is possible for a natural in-place and in-order PFA computation are analyzed. The result is useful to both the hardware and software realization of the PFA.>
Daniel Pak-Kong Lun, Wan-Chi Siu
ICASSP1
1992 Efficient mapping scheme for the prime factor discrete Hartley transform
abstract
Compared to the complexity for realizing the prime factor discrete Fourier transform (DFT), the prime factor discrete Hartley transform requires some extra arithmetic operations for the realization of the prime factor mapping. These extra arithmetic operations can take up as much as 40% of the total arithmetic operations required. A new prime factor mapping scheme which requires no extra arithmetic operations is proposed for the computation of the discrete Hartley transform. It is achieved by embedding all the extra arithmetic operations into the subsequent short length computations, whereas the arithmetic complexities of these embedded short length modules remain unchanged.>
Wan-Chi Siu, Daniel Pak-Kong Lun
ICASSP2
1990 Yet a faster address generation scheme for the computation of prime factor algorithms
abstract
An in-place, in-order address generation scheme is proposed for the realization of prime factor mapping (PFM). The new approach has the characteristic of forming systematic and regular structures. Hence it is suitable for realizations using both high-level and low-level languages. Furthermore, it requires very few modulo operations and no modulo inverse for its computation; such inverses often take up memory space for their storage and/or extra time for the computation in other address generation algorithms. The approach has been implemented using Fortran 77 and the assembly language of the 320C25 DSP. It shows that a maximum of 86% saving in address generation time can be achieved as compared to the conventional approach.>
Daniel Pak-Kong Lun, Wan-Chi Siu
ICASSP1