VLDB 2026 Research / reviewers in the wild / expert
Ja-Ling Wu
dblp:51/2694
· DBLP profile ↗
142ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-3631-1551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 15 · 3 since 2021Computer networks · 9 · 1 first-authorSecurity and privacy · 7Artificial intelligence and machine learning · 5Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SLIC: Secure Learned Image Codec through Compressed Domain Watermarking to Defend Image Manipulation
Chen-Hsiu Huang, Ja-Ling Wu |
MMAsia | 2 |
| 2024 | Joint Image Data Hiding and Rate-Distortion Optimization in Neural Compressed Latent Representations
Chen-Hsiu Huang, Ja-Ling Wu |
MMM (1) | 2 |
| 2023 | Image Data Hiding in Neural Compressed Latent RepresentationsabstractWe propose an end-to-end learned image data hiding framework that embeds and extracts secrets in the latent representations of a generic neural compressor. By leveraging a perceptual loss function in conjunction with our proposed message encoder and decoder, our approach simultaneously achieves high image quality and high bit accuracy. Compared to existing techniques, our framework offers superior image secrecy and competitive watermarking robustness in the compressed domain while accelerating the embedding speed by over 50 times. These results demonstrate the potential of combining data hiding techniques and neural compression and offer new insights into developing neural compression techniques and their applications. Chen-Hsiu Huang, Ja-Ling Wu |
VCIP | 2 |
| 2022 | Enhancing the Robustness of Deep Learning Based Fingerprinting to Improve Deepfake AttributionabstractArtificial Fingerprinting (AF or the so-called digital watermarking) is a technique that can be used to conduct Deepfake attribution by ensuring media authenticity. However, AF does not prioritize its robustness to certain kinds of distortions, making the embedded watermarks vulnerable to some standard image processing operations. Insufficient robustness reduces the practicality of digital watermarking techniques. To address this issue, we propose an enhanced distortion agnostic artificial fingerprinting (EDA-AF) framework which introduces a novel noise layer consisting of an attack booster followed by a convolutional network-based attacker. The attacker simulates various distortions by exploiting adversarial learning with AF for distortion agnostic robustness. Meanwhile, due to the modeling limitation of the convolutional network, we also employ the attack booster to apply a set of differentiable image distortions which cannot be well simulated by the attacker. Extensive experimental results show that the proposed approach improves the quality of the extracted fingerprints. EDA-AF can improve the bitwise accuracy by up to 36%, which takes another step forward on the road of Deepfake attribution. Chieh-Yin Liao, Chen-Hsiu Huang, Jun-Cheng Chen, Ja-Ling Wu |
MMAsia | 4 |
| 2021 | Attribute-Based Facial Image Manipulation on Latent SpaceabstractUsing machine learning to generate images has become more mature, especially the images produced using a Generative Adversarial Network. Unfortunately, the complicated architecture of those models makes it difficult for us to ensure the output images’ diversity and controllability without introducing little embarrassment in implementation. Therefore, some researchers try to edit the latent codes generated by a given learning model directly on the latent space for manipulating the output image by simply inputting the new latent codes into the original model without changing the model’s structure and learned parameters. However, the methods mentioned above faced the problems that the size of latent space cannot be too large or the trouble-some of features entanglement. In this work, we propose an approach to conquer the problems mentioned above, which is to compress the original latent space to better the applicability and usability of the methods limited by the size of the latent space. Compared with the existing methods, this method can be applied to more models and still reach the target of image manipulation. Yi-Lun Pan, Ja-Ling Wu |
AVSS | 3 |
| 2021 | JQF: Optimal JPEG Quantization Table Fusion by Simulated Annealing on Texture Images and Predicting TexturesabstractJPEG has been a widely used lossy image compression codec for nearly three decades. The JPEG standard allows to use customized quantization table; however, it's still a challenging problem to find an optimal quantization table within acceptable computational cost. This work tries to solve the dilemma of balancing between computational cost and image specific optimality by introducing a new concept of texture mosaic images. Instead of optimizing a single image or a collection of representative images, the simulated annealing technique is applied to texture mosaic images to search for an optimal quantization table for each texture category. We use pre-trained VGG-16 CNN model to learn those texture features and predict the new image's texture distribution, then fuse optimal texture tables to come out with an image specific optimal quantization table. On the Kodak dataset with the quality setting Q=95, our experiment shows a size reduction of 23.5% over the JPEG standard table with a slightly 0.35% FSIM decrease, which is visually unperceivable. The proposed JQF method achieves per image optimality for JPEG encoding with less than one second additional timing cost. Chen-Hsiu Huang, Ja-Ling Wu |
DCC | 2 |
| 2021 | A Multi-Factor Combinations Enhanced Reversible Privacy Protection System for Facial ImagesabstractWith the abuse of deepfake and other deep learning technologies, anonymization and deanonymization for face images have become one of the essential tasks for privacy protection. Thus, we propose a novel reversible privacy protection framework for facial images based on conditional encoder and de-coder framework. For the purpose to increase the diversity and controllability over the anonymized faces, we also introduce facial attributes and a style vector from a reference back-ground face dataset and pretrained face recognition model and thus name the proposed framework as the Multi-factor Modifier (MfM) to achieve multi-factor facial de/re-identification. Specifically, with the correct password, our method produces near-original reconstructed images. Otherwise, it can generate photo-realistic and diverse anonymized images. With extensive experiments, it shows that the proposed approach can successfully anonymize face images in high fidelity according to the given conditions as compared with other methods and deanonymize without altering the facial data distributions. Yi-Lun Pan, Jun-Cheng Chen, Ja-Ling Wu |
ICME | 3 |
| 2019 | K-Same-Siamese-GAN: K-Same Algorithm with Generative Adversarial Network for Facial Image De-identification with Hyperparameter Tuning and Mixed Precision TrainingabstractFor a data holder, such as a hospital or a government entity, who has a privately held collection of personal data, in which the revealing and/or processing of the personal identifiable data is restricted and prohibited by law. Then, “how can we ensure the data holder does conceal the identity of each individual in the imagery of personal data while still preserving certain useful aspects of the data after de-identification?” becomes a challenge issue. In this work, we propose an approach towards high-resolution facial image de-identification, called k-Same-Siamese-GAN, which leverages the k-Same-Anonymity mechanism, the Generative Adversarial Network, and the hyperparameter tuning methods. Moreover, to speed up model training and reduce memory consumption, the mixed precision training technique is also applied to make kSS-GAN provide guarantees regarding privacy protection on close-form identities and be trained much more efficiently as well. Finally, to validate its applicability, the proposed work has been applied to actual datasets - RafD and CelebA for performance testing. Besides protecting privacy of high-resolution facial images, the proposed system is also Justified for its ability in automating parameter tuning and breaking through the limitation of the number of adjustable parameters. Yi-Lun Pan, Min-Jhih Haung, Kuo-Teng Ding, Ja-Ling Wu, Jyh-Shing Roger Jang |
AVSS | 4 |
| 2017 | A Privacy-Preserving Cloud-Based Data Management System with Efficient Revocation SchemeabstractThere are lots of data management systems, according to various reasons, designating their high computational work-loads to public cloud service providers. It is well-known that once we entrust our tasks to a cloud server, we may face several threats, such as privacy-infringement with regard to users attribute information; therefore, an appropriate privacy preserving mechanism is a must for constructing a secure cloud-based data management system (SCBDMS). To design a reliable SCBDMS with server-enforced revocation ability is a very challenging task even if the server is working under the honest-but-curious mode. In existing data management systems, there seldom provide privacy-preserving revocation service, especially when it is outsourced to a third party. In this work, with the aids of oblivious transfer and the newly proposed stateless lazy re-encryption (SLREN) mechanism, a SCBDMS, with secure, reliable and efficient server-enforced attribute revocation ability is built. Comparing with related works, our experimental results show that, in the newly constructed SCBDMS, the storage-requirement of the cloud server and the communication overheads between cloud server and systems users are largely reduced, due to the nature of late involvement of SLREN. Shih-Chien Chang, Ja-Ling Wu |
PDCAT | 2 |
| 2016 | Emotion Prediction from User-Generated Videos by Emotion Wheel Guided Deep Learning
Che-Ting Ho, Yu-Hsun Lin, Ja-Ling Wu |
ICONIP (1) | 3 |
| 2015 | Reversible Data Hiding for Encrypted Audios by High Order Smoothness
Jing-Yong Qiu, Yu-Hsun Lin, Ja-Ling Wu |
IWDW | 3 |
| 2015 | Secure Client Side Watermarking with Limited Key Size
Jia-Hao Sun, Yu-Hsun Lin, Ja-Ling Wu |
MMM (1) | 3 |
| 2015 | Efficient human detection in crowded environment
Min-Chun Hu 0001, Wen-Huang Cheng, Chuan-Shen Hu, Ja-Ling Wu, Jhe-Wei Li |
Multim. Syst. | 4 |
| 2015 | Real-Time Human Movement Retrieval and Assessment With Kinect SensorabstractThe difficulty of vision-based posture estimation is greatly decreased with the aid of commercial depth camera, such as Microsoft Kinect. However, there is still much to do to bridge the results of human posture estimation and the understanding of human movements. Human movement assessment is an important technique for exercise learning in the field of healthcare. In this paper, we propose an action tutor system which enables the user to interactively retrieve a learning exemplar of the target action movement and to immediately acquire motion instructions while learning it in front of the Kinect. The proposed system is composed of two stages. In the retrieval stage, nonlinear time warping algorithms are designed to retrieve video segments similar to the query movement roughly performed by the user. In the learning stage, the user learns according to the selected video exemplar, and the motion assessment including both static and dynamic differences is presented to the user in a more effective and organized way, helping him/her to perform the action movement correctly. The experiments are conducted on the videos of ten action types, and the results show that the proposed human action descriptor is representative for action video retrieval and the tutor system can effectively help the user while learning action movements. Min-Chun Hu 0001, Chi-Wen Chen, Wen-Huang Cheng, Che-Han Chang, Jui-Hsin Lai, Ja-Ling Wu |
IEEE Trans. Cybern. | 6 |
| 2015 | Audio Musical Dice Game: A User-Preference-Aware Medley Generating SystemabstractThis article proposes a framework for creating user-preference-aware music medleys from users' music collections. We treat the medley generation process as an audio version of a musical dice game. Once the user's collection has been analyzed, the system is able to generate various pleasing medleys. This flexibility allows users to create medleys according to the specified conditions, such as the medley structure or the must-use clips. Even users without musical knowledge can compose medley songs from their favorite tracks. The effectiveness of the system has been evaluated through both objective and subjective experiments on individual components in the system. Yin-Tzu Lin, I-Ting Liu, Jyh-Shing Roger Jang, Ja-Ling Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2014 | A 3D HEVC Fast Mode Decision Algorithm Based on the Depth Information Guided Maximum Coding Levelabstract3D HEVC is one of the extensions of HEVC (High Efficiency Video Coding), which is the latest video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). The inherent high computing complexity of 3D HEVC handicaps its usage in practical applications. HEVC replaces the macroblock (MB) of H.264 with the largest coding unit (LCU), in which coding units (CUs) of sizes 64 × 64, 32 × 32, 16 × 16, or 8 × 8 pixels are used to build a coding tree unit (CTU). HEVC encoder will traverse the coding modes from the root (the coding level is 0) to the leaves (the coding level is 3) of the CTU. As a result, how to accelerate the encoding process of 3D HEVC with negligible loss of coding efficiency is a hot research topic in the field of video coding. Figure 1 shows the coding level distributions of I04 and I14, where Ikj represents the j-th frame in the k-th GOP. The POC's (picture order counts) of the two frame are P(I04) = 4 and P(I14) = 12 when GOP size is 8. From Figure 1 we observed that, in B-slices, the distributions of "optimal coding level = zero" for adjacent GOP's are almost the same. We found that the required coding level will be similar when the two CUs have similar depth values. We utilize this correlation to limit the required coding level visiting of a given CU which accelerates the encoding process. In this paper, a fast mode decision algorithm for 3D HEVC, based on the depth information related coding mode similarity, is proposed. According to the experimental results, the proposed algorithm can reduce up to 72% executing time of the overall encoding process, while the loss in coding efficiency is negligible. Ming-Chang Li, Yu-Hsun Lin, Yin-Tzu Lin, Yun-Chung Shen, Ja-Ling Wu |
DCC | 5 |
| 2014 | Binocular Perceptual Model for Symmetric and Asymmetric 3D Stereoscopic Image CompressionabstractThe objective approaches of 3D image quality assessment play a key role in the development of compression standards and various 3D multimedia applications. The quality assessment of 3D images faces many new challenges, e.g. asymmetric stereo compression, depth perception, and virtual view synthesis, as compared with its 2D counterparts. Moreover, the widely used 2D image quality metric (e.g. PSNR) cannot be directly applied to deal with these newly introduced challenges. This statement can be verified by the low correlation between the computed objective measures and the subjectively measured mean opinion scores (MOS), when 3D images are the tested targets. In order to meet these challenges, in this work, besides traditional 2D image metrics, two binocular behaviors - the binocular combination and the Binocular Frequency Integration (BFI), are utilized as the bases for measuring the quality of stereoscopic 3D images. The effectiveness of BFI-based metrics is verified by conducting subjective evaluations on a publicly available stereo image dataset. Experimental results show that significant consistency could be reached between the measured MOS and the BFI-based metrics, in which the correlation coefficient between them can go up to 0.89 even if the stereo images have been asymmetrically HEVC compressed. Based on the proposed quality metric, we find that asymmetric-stereo compression schemes outperform the corresponding symmetric ones in high bit rate scenarios, which opens up a new research direction for 3D stereo image/video compression studies. Yu-Hsun Lin, Ja-Ling Wu |
DCC | 2 |
| 2014 | Seam Carving for Color-Plus-Depth 3D ImageabstractColor-plus-Depth 3D images are booming up with the advance of depth-sensing camera (e.g., Kinect). This new 3D visual content imposes new challenges on image resizing since we have to resize both the color and the depth images simultaneously. In order to resolve these newly introduced challenges of color-plus-depth 3D images, we propose a new energy function with salient depth cues consideration for seam carving operation. We further incorporate super-pixel over-segmentation and depth remapping for achieving an object-based and 3D viewing comfort zone aware resizing framework. Experimental results show the proposed framework can effectively generate the resized images by maintaining the salient regions and providing a comfortable 3D visual perception, at the same time. Wei-Cih Jhou, Yu-Hsun Lin, Ja-Ling Wu |
ISM | 3 |
| 2014 | Bridging Music via Sound EffectsabstractThe prevalence of digital technologies allows people to easily create and share their own media contents, but sometimes we do not have handy tools to manipulate the media we want to create. For example, while creating personal films, a user may separately find the music segment that matches each part of the video, and then concatenate the segments to create the soundtrack that matches the visual contents along the timeline. However, one may lose the temporal coherence between consecutive music segments if the user just choose music segments that best match parts of the video contents. In this study, we focus on how to smoothly connect not-so-coherent music clips and make the transition natural and pleasant to hear. In particular, we improve the temporal smoothness by bar alignment and dual tempo adjustment. To further fit in with the transition between clips, we incorporate "sound effect insertion" which is a commonly used technique in popular song composition/remixing. In order to provide pleasant listening experience and systematically analyze the effectiveness of the proposed music bridging method, we have conducted specifically designed experiments to collect subjective opinions and reduce the cognitive loads of the participants. The experimental results indicate that with proper arrangement for creating smooth transition via tempo adjustment and sound effect insertion, the listening experience can be largely enhanced. Yin-Tzu Lin, Chuan-Lung Lee, Jyh-Shing Roger Jang, Ja-Ling Wu |
ISM | 4 |
| 2014 | When Specular Object Meets RGB-D Camera 3D Scanning: Color Image Plus Fragmented Depth Mapabstract3D scanning is an important technology since one can apply the scanning results of objects to numerous 3D applications (e.g. Artwork sculpture preserving, 3D animation and 3D printing). However, 3D laser scanners are too expensive to be adopted in daily usage. Therefore, building a 3D scanner by a low cost RGB-D camera (e.g. Kinect) is an emerging trend. For a 3D scanner, a specular object is one of the challenging 3D scanning targets where the specular surface will compromise the scanning (laser) lights and lead to scanning failure. In order to acquire the ground truth for 3D scanning of specular objects, we have to perform a non-specular painting process even a 3D laser scanner is used. In order to meet this challenge for an RGB-D camera based 3D scanner, we integrate the response of visual cues reflection and the depth scattering characteristics of specular surfaces to resolve the artifacts of 3D scanning results. Experimental results show our proposed system outperforms the traditional RGB-D 3D scanners in reconstruction quality while keeping the specular object intact (i.e., We need not to perform the pre-described non-specular painting process on the object before scanning). Shun-Xuan Wang, Yu-Hsun Lin, Ja-Ling Wu |
ISM | 3 |
| 2014 | Image Descriptor Based Digital Semi-blind Watermarking for DIBR 3D Images
Hsin Miao, Yu-Hsun Lin, Ja-Ling Wu |
IWDW | 3 |
| 2014 | MSVA: Musical Street View Animator: An Effective and Efficient Way to Enjoy the Street Views of Your JourneyabstractGoogle Maps with Street View (GSV) provides ways to explore the world but it lacks efficient ways to present a journey. Hyperlapse provides another ways for quick glimpsing the street-views along the route; however, its viewing experience is also not comfortable and could be tedious when the route is long-distance. In this paper, we provide an efficient and enjoyable way to present street view sequences of long journey. Street view journey video accompanied with locally listened music will be produced by the proposed approach. During the move between locations, we use the speed control techniques for animation production to improve the viewing experience. User evaluation results show that the proposed method increases the satisfaction of users in viewing the street view sequences. Yin-Tzu Lin, Po-Nien Chen, Chia-Hu Chang, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2014 | Event Detection in Broadcasting Video for Halfpipe SportsabstractIn this work, a low-cost and efficient system is proposed to automatically analyze the halfpipe (HP) sports videos. In addition to the court color ratio information, we find the player region by using salient object detection mechanisms to face the challenge of motion blurred scenes in HP videos. Besides, a novel and efficient method for detecting the spin event is proposed on the basis of native motion vectors existing in a compressed video. Experimental results show that the proposed system is effective in recognizing the hard-to-be-detected spin events in HP videos. Hao-Kai Wen, Wei-Che Chang, Chia-Hu Chang, Yin-Tzu Lin, Ja-Ling Wu |
ACM Multimedia | 5 |
| 2014 | TravelBuddy: Interactive Travel Route Recommendation with a Visual Scene Interface
Cheng-Yao Fu, Min-Chun Hu 0001, Jui-Hsin Lai, Hsuan Wang, Ja-Ling Wu |
MMM (1) | 5 |
| 2014 | Semantic Based Background Music Recommendation for Home Videos
Yin-Tzu Lin, Tsung-Hung Tsai, Min-Chun Hu 0001, Wen-Huang Cheng, Ja-Ling Wu |
MMM (2) | 5 |
| 2014 | Depth sculpturing for 2D paintings: A progressive depth map completion framework
Yu-Hsun Lin, Ming-Hung Tsai, Ja-Ling Wu |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Quality Assessment of Stereoscopic 3D Image Compression by Binocular Integration BehaviorsabstractThe objective approaches of 3D image quality assessment play a key role for the development of compression standards and various 3D multimedia applications. The quality assessment of 3D images faces more new challenges, such as asymmetric stereo compression, depth perception, and virtual view synthesis, than its 2D counterparts. In addition, the widely used 2D image quality metrics (e.g., PSNR and SSIM) cannot be directly applied to deal with these newly introduced challenges. This statement can be verified by the low correlation between the computed objective measures and the subjectively measured mean opinion scores (MOSs), when 3D images are the tested targets. In order to meet these newly introduced challenges, in this paper, besides traditional 2D image metrics, the binocular integration behaviors-the binocular combination and the binocular frequency integration, are utilized as the bases for measuring the quality of stereoscopic 3D images. The effectiveness of the proposed metrics is verified by conducting subjective evaluations on publicly available stereoscopic image databases. Experimental results show that significant consistency could be reached between the measured MOS and the proposed metrics, in which the correlation coefficient between them can go up to 0.88. Furthermore, we found that the proposed metrics can also address the quality assessment of the synthesized color-plus-depth 3D images well. Therefore, it is our belief that the binocular integration behaviors are important factors in the development of objective quality assessment for 3D images. Yu-Hsun Lin, Ja-Ling Wu |
IEEE Trans. Image Process. | 2 |
| 2013 | Angular Disparity Map: A Scalable Perceptual-Based Representation of Binocular DisparityabstractThis work addresses the data representation and the compression issues of angular disparity map following the way of HVS to perceive depth information. The continued fraction is utilized to represent the angular disparity map which enables the use of the state-of-the-art video codec (e.g. HEVC) to compress the data directly and maintains quality scalability properties. We observe that there is a non-monotonic phenomenon of the RD curves by applying HEVC compression to angular disparity map directly. This implies that the correlations among inter-layer (i.e., the neighboring integers in (2)) do not follow the traditional models of normal 2D video codecs. Of course, the detailed relationship between the sensitivities and the quantization errors of the newly proposed representation needs in depth further derivations. There are many interesting research issues may be introduced by the proposed data format (e.g., the sensitivities to quantization errors of θ and the rate-distortion optimization scheme for θ) which will, of course, be the research topics of our future work. We expect this work can be a bridge to connect the 3D perception and the 3D compression research fields. Yu-Hsun Lin, Ja-Ling Wu |
DCC | 2 |
| 2013 | Subsampling Input Based Side Information Creation in Wyner-Ziv Video CodingabstractSummary form only given. Distributed video coding (DVC) has been intensively studied in recent years. This new coding paradigm substantially differs from conventional prediction-based video codecs such as MPEG and H.26x, which are characterized by a complex encoder and simple decoder. The conventional DVC codec, e.g., DISCOVER codec, uses advanced frame interpolation techniques to create SI based on adjacent decoded reference frames. The quality of SI is a well-recognized factor in the RD performance of WZ video coding. A high SI quality implies a high correlation between the created SI and the original WZ frame, which then decreases the rate required to achieve a given decoded quality. Clearly, the performance of an SI creation process based on adjacent previously decoded frames is limited by the quality of the past and the future reference frames as well as the distance and motion behavior between them. The correlation between high-motion frames is low and vice versa. That is, SI quality in the conventional codecs depends on the temporal correlation of key frames, which affects the bitrate and PSNR of the compression process. In this work, a novel DVC architecture for dealing with the cases of high-motion and large GOP-size sequences is proposed to better the rate-distortion (RD) performance. For high-motion video sequences, the proposed architecture generates SI by using subsampled spatial information instead of interpolated temporal information. the proposed approach separates the video sequence into subsampled key frames and corresponding WZ frames, which changes the creation of SI. That is, all successive frames on the encoder side are downsized to sub-frames, which are then compressed by an H.264/AVC intra encoder. Experimental results reveal that the subsampling input based DVC codec can gain up to 1.47 dB in the RD measures and maintains the most important characteristic of the DVC codec, the encoder is lightweight, as compared with the conventional WZ codec, respectively. The novel DVC architecture evaluated in this study exploits spatial relations to create SI. The experimental results confirm that the RD performance of the proposed approach is superior to that of the conventional one for high-motion and/or large GOP-size sequences. The quality of spatial interpolation based SI is higher than that of the temporal interpolation one, which leads to a high-PSNR reconstructed WZ frame. The subsampled key frames are also decoded by LDPCA decoder to recover the information lost when H.264/AVC intra coding is used to increase PSNR gain. Since many spatial domain interpolation and super resolution schemes have been proposed for use in the fields of image processing and computer vision, the performance of the proposed DVC codec can be further enhanced by using better schemes to generate even better SI. Yun-Chung Shen, Ji-Ciao Luo, Ja-Ling Wu |
DCC | 3 |
| 2013 | A Novel Privacy Preserving Location-Based Service Protocol With Secret Circular Shift for k-NN SearchabstractLocation-based service (LBS) is booming up in recent years with the rapid growth of mobile devices and the emerging of cloud computing paradigm. Among the challenges to establish LBS, the user privacy issue becomes the most important concern. A successful privacy-preserving LBS must be secure and provide accurate query [e.g., -nearest neighbor (NN)] results. In this work, we propose a private circular query protocol (PCQP) to deal with the privacy and the accuracy issues of privacy-preserving LBS. The protocol consists of a space filling curve and a public-key homomorphic cryptosystem. First, we connect the points of interest (POIs) on a map to form a circular structure with the aid of a Moore curve. And then the homomorphism of Paillier cryptosystem is used to perform secret circular shifts of POI-related information (POI-info), stored on the server side. Since the POI-info after shifting and the amount of shifts are encrypted, LBS providers (e.g., servers) have no knowledge about the user's location during the query process. The protocol can resist correlation attack and support a multiuser scenario as long as the predescribed secret circular shift is performed before each query; in other words, the robustness of the proposed protocol is the same as that of a one-time pad encryption scheme. As a result, the security level of the proposed protocol is close to perfect secrecy without the aid of a trusted third party and simulation results show that the k-NN query accuracy rate of the proposed protocol is higher than 90% even when is large. I-Ting Lien, Yu-Hsun Lin, Jyh-Ren Shieh, Ja-Ling Wu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2013 | Relational term-suggestion graphs incorporating multipartite concept and expertise networksabstractTerm suggestions recommend query terms to a user based on his initial query. Suggesting adequate terms is a challenging issue. Most existing commercial search engines suggest search terms based on the frequency of prior used terms that match the leading alphabets the user types. In this article, we present a novel mechanism to construct semantic term-relation graphs to suggest relevant search terms in the semantic level. We built term-relation graphs based on multipartite networks of existing social media, especially from Wikipedia. The multipartite linkage networks of contributor-term, term-category, and term-term are extracted from Wikipedia to eventually form term relation graphs. For fusing these multipartite linkage networks, we propose to incorporate the contributor-category networks to model the expertise of the contributors. Based on our experiments, this step has demonstrated clear enhancement on the accuracy of the inferred relatedness of the term-semantic graphs. Experiments on keyword-expanded search based on 200 TREC-5 ad-hoc topics showed obvious advantage of our algorithms over existing approaches. Jyh-Ren Shieh, Ching-Yung Lin, Shun-Xuan Wang, Ja-Ling Wu |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2013 | Appearance-Based QR Code BeautifierabstractQuick Response (QR) code is a widely used matrix bar code with the increasing population of smart phones. However, QR code usually consists of random textures which are not suitable for incorporating with other visual designs (e.g. name card and business advertisement poster). In order to overcome the shortcomings of noise-like looks of QR codes, we propose a systematic QR code beautification framework where the visual appearance of QR code is composed of visually meaningful patterns selected by users, and more importantly, the correctness of message decoding is kept intact. Our work makes QR code from machine decodable only (i.e. standardized random texture) to a personalized form with human visual pleasing appearance. We expect the proposed QR code beautifier can inspire more visual-pleasant mobile multimedia applications. Yu-Hsun Lin, Yu-Pei Chang, Ja-Ling Wu |
IEEE Trans. Multim. | 3 |
| 2012 | Progressive Side Information Refinement with Non-local Means Based Denoising Process for Wyner-Ziv Video CodingabstractOn the basis of a non-local means denosing process, a novel progressive side information refinement framework for a transform domain Wyner-Ziv video codec is proposed, where the side information is progressively improved, as the decoding proceeds, by exploring both temporal-spatial similarities and already decoded DCT bands. Simulation results show up to 2.5dB in RD performance gain against that of the same codec without including the proposed side information refinement framework for sequences with high motion and/or large group-of-picture sizes. Moreover, the proposed framework adds no significant complexity to the decoder, and even results in a decoding speed-up due to the less required error correcting iterations, for most of the test sequences. Yun-Chung Shen, Pin-Shiang Wang, Ja-Ling Wu |
DCC | 3 |
| 2012 | Unseen visible watermarking for color plus depth map 3D imagesabstractUnseen visible watermarking (UVW) is a novel data hiding scheme that imitates real-world watermarks and maintains advantages of both visible and invisible watermarking. One important feature of UVW is that specific extraction module is not required during watermarking decoding. The UVW for 2D image is studied based on imaging functions, e.g. gamma-correction in LCD monitors. On the other hand, due to the great success of 3D movies and low-priced 3D display devices, 3D contents are booming up in recent years. In this work, we propose a UVW for color plus depth map 3D images in which the watermark extraction is realized by changing of the rendering conditions. Under normal rendering conditions, the watermarked cover can be perceived exactly the same as the original one. Limitations and future extensions of the proposed 3D UVW are also addressed in this work. Yu-Hsun Lin, Ja-Ling Wu |
ICASSP | 2 |
| 2012 | Sampling Technique Analysis of Nyström Approximation in Pixel-Wise Affinity MatrixabstractSpectral graph methods are widely employed in image segmentation, and they exhibit excellent performance. However, for high-resolution images, it is impractical to directly calculate the eigenvectors of the affinity matrix owing to the high computational requirements. The Nystrom method provides an efficient way to approximate the large-scale affinity matrix by low-rank approximation. In the machine learning field, previous studies have mainly focused on less data points with high dimensional features. To the best of our knowledge, this is the first study to discuss the performance of sampling methods for Nystrom approximation, in which we focus on the pixel-wise affinity matrix for a single image. In this paper, we propose a mean-shift segmentation-based Nystrom sampling technique for image analysis. The experimental results show that for images with simple compositions and backgrounds, k-means sampling performs better, whereas for images with more complicated compositions and backgrounds, the proposed method can perform better. Chieh-Chi Kao, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 3 |
| 2012 | Stable Pose Estimation with a Motion Model in Real-Time ApplicationabstractEstimation of a object pose from camera is a well-developing topic in computer vision. In theory, the pose from a calibrated camera can be uniquely determined. But in practice, most of the real-time pose estimation algorithms suffer from pose ambiguity due to low accuracy of the target object. We think that pose ambiguity¡Xtwo distinct local minima of the according error function¡Xexist because of the phenomenon of geometric illusions. Both of the ambiguous poses are plausible. After obtaining the solution of two minima (pose candidates), we develop a real-time algorithm for stable pose estimation of a target objects with a motion model. In the experimental results, the proposed algorithm diminish the significance of pose jumping and pose jittering effectively. To the best of our knowledge, this is the first work to solve the pose ambiguity problem with motion model in real-time application. Po-Chen Wu, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 3 |
| 2012 | Action tutor: real-time exemplar-based sequential movement assessment with kinect sensorabstractWith the aid of depth camera, such as Microsoft Kinect, the difficulty of vision-based posture estimation is greatly decreased, and human action analysis has achieved a wide range of applications. However, there is still much to do to develop effective movement assessment technique, which bridges the results of human posture estimation and the understanding of human action performance. In this work, we propose an action tutor system which enables the user to interactively retrieve the learning exemplar of the target action movement and to immediately acquire motion instructions while learning it in front of the Kinect. In the retrieval stage, non-linear time warping algorithms are designed to retrieve video segments similar to the query movement roughly performed by the user. In the learning stage, the user learns according to the selected video exemplar, and the motion assessment including both static and dynamic differences is presented to the user in a more effective and organized way, helping him/her to perform the action movement correctly. Chi-Wen Chen, Min-Chun Hu 0001, Wen-Huang Cheng, Che-Han Chang, Jui-Hsin Lai, Ja-Ling Wu |
ACM Multimedia | 6 |
| 2012 | U-Drumwave: An Interactive Performance System for Drumming
Yin-Tzu Lin, Shuen-Huei Guan, Yuan-Chang Yao, Wen-Huang Cheng, Ja-Ling Wu |
MMM | 5 |
| 2012 | Gait-Based Action Recognition via Accelerated Minimum Incremental Coding Length Classifier
Hung-Wei Lin, Min-Chun Hu 0001, Ja-Ling Wu |
MMM | 3 |
| 2012 | High efficient distributed video coding with parallelized design for LDPCA decoding on CUDA based GPGPU
Yu-Shan Pai, Yun-Chung Shen, Ja-Ling Wu |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Batch-pipelining for multicore H.264 decoding
Tang-Hsun Tu, Chih-Wen Hsueh, Ja-Ling Wu |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Recommendation in the end-to-end encrypted domainabstractIn recommendation systems, a central host typically requires access to user profiles in order to generate useful recommendations. This access, however, undermines user privacy; the more information is revealed to the host, the more the user's privacy is compromised. In this paper, we propose a novel end-to-end encrypted recommendation mechanism which encrypts sensitive private data at the user end, without ever exposing plaintext private data to the host server. Unlike previously proposed privacy-preserving recommendation mechanisms, the data in this proposed system are lossless - a pivotal feature to many applications, e.g., in health informatics, business analytics, cyber security, etc. We achieve this goal by developing encrypted-domain polynomial ring homomorphism cryptographic algorithms to compute similarity of encrypted scores on the server, so that collaborative recommendations can be computed in the encryption domain and only an authorized person can decrypt the exact results. We also propose a novel key management system to make sure private information retrieval and recommendation computations can be executed in the encrypted domain in practice. Our experiments show that the proposed scheme offers robust security and lossless accurate recommendation, as well as high efficiency. Our preliminary results show the recommendation accuracy is 21% better than the existing statistical lossy privacy-preserving mechanisms based on random perturbation and user profile distribution. This new approach can potentially be applied to various data mining and cloud computing environments and significantly alleviates the privacy concerns of users. Jyh-Ren Shieh, Ching-Yung Lin, Ja-Ling Wu |
CIKM | 3 |
| 2011 | Rendering Lossless Compression of Depth ImageabstractSummary form only given. In this work, we experimented on the compression efficiency of rendering lossless compression of depth images and found that the compression ratios can go up to 20.51 and 40 for Interview and Breakdancer test images, respectively, even if the parameter setting is in the worst case (i.e., set fdensity(Znear, Zfar) to its maximum value). This work is our first step toward exploring the performance of rendering lossless compression. It is our belief that, besides the rendering lossless quantization, there are a lot of different issues of depth image compression which are worthy of further exploitation. Yu-Hsun Lin, Ja-Ling Wu |
DCC | 2 |
| 2011 | An Efficient Distributed Video Coding with Parallelized Design for Concurrent ComputingabstractSummary form only given. In this paper, Wyner-Ziv (WZ) video coding is a particular case of distributed video coding (DVC). Although some works, with improved performance, have been made in recent years, the coding efficiency of state-of-the-art WZ codec is still far from that of the state-of-the-art prediction-based codec, especially for high and complex motion contents. Moreover, most reported WZ codecs have a high time delay in decoder, which hinders its practical application in real-time systems. The performance of the SI creation process based on adjacent previously decoded frames is limited by the quality of the past and the future reference frames as well as the distance and motion behavior between them. In this work, by combining coding tools developed in recent literatures on transform domain WZ coding with some newly developed modules on both encoding and decoding sides, an efficient and practical WZ video coding architecture, dubbed as Distributed video coding with PArallelized design for Concurrent computing (DISPAC), is proposed to better the rate-distortion (RD) performance. Another unique feature of DISPAC, lies in the parallelizability of the modules used by its WZ decoder which increased the decoding speed largely. Experimental results conducted on a concurrent computing environment (consisting of multi-core CPU and GPU processors) reveal that DISPAC codec can gain up to 2.8 dB in the RD measures and 14.35 times faster in the decoding speed as compared with the-state-of-art WZ video codec, respectively. By shifting the computational complexity from the encoder to the decoder and integrating with appropriate trascoding techniques, DVC has been expected to provide a video codec solution for Cloud computing mobile devices (such as mobile phones). Yun-Chung Shen, Han-Ping Cheng, Ja-Ling Wu |
DCC | 3 |
| 2011 | High efficient distributed video coding with parallelized design for cloud computingabstractIn this work, by combining coding tools developed in recent literatures on transform domain WZ video coding with some newly developed modules on both encoding and decoding sides, an efficient and practical WZ video coding architecture, dubbed as DIStributed video coding with PArallelized design for Cloud computing (DISPAC), is proposed to better the corresponding rate-distortion (RD) performance. Another unique feature of DISPAC, lies in the parallelizability of the modules used by its WZ decoder which increased the decoding speed largely. Experimental results conducted on an emulated Could computing environment reveal that DISPAC codec can gain up to 3.6 dB in the RD measures and 60.97 times faster in the decoding speed as compared with the-state-of-art WZ video codec, respectively Han-Ping Cheng, Yun-Chung Shen, Ja-Ling Wu, Kiyoharu Aizawa |
ACM Multimedia | 3 |
| 2011 | Real-time decoding for LDPC based distributed video codingabstractWyner-Ziv (WZ) video coding -- a particular case of distributed video coding (DVC), is well known for its low-complexity encoding and high-complexity decoding characteristics. Although some works have been made in recent years, especially for improving the coding efficiency, most reported WZ codecs have high time delay in the decoder, which hinders its practical values for applications with critical timing constraint. In this paper, a fully parallelized sum-product algorithm (SPA) for low density parity check accumulate (LDPCA) codes is proposed and realized through Compute Unified Device Architecture (CUDA) based on General-Purpose Graphics Processing Unit (GPGPU). Simulation results show that, through our work, QCIF (surveillance) videos can be decoded in real-time with extremely high quality and without rate-distortion (RD) performance loss. Tse-Chung Su, Yun-Chung Shen, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2011 | Distributed video coding with compressive measurementsabstractThis paper presents a novel distributed video coding (DVC) scheme using compressive sensing (CS) that achieves low-complexity for encoding and efficient signal sensing. Most CS recovery algorithms rely only on signal sparsity. Yet, under DVC architecture, additional statistical characterization of the signal is available, which offers the potential for more precise CS recovery. First, a set of random measurements are acquired and transmitted to the decoder. The decoder then exploits the statistical characterization of the signal and generates the side information (SI). Finally, utilizing the SI, a Bayesian inference using belief propagation (BP) decoding is performed for signal recovery. The proposed CS-DVC system offers a more direct way of signal acquisition and the potential for more precise estimation of the signal from random measurements. Experimental results indicate that SI can improve the signal reconstruction quality in comparison with a CS recovery algorithm that relies only on the sparsity. Hsiao-Yun Tseng, Yun-Chung Shen, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2011 | Interactive digital scrapbook generation for travel photos based on design principles of typographyabstractTo facilitate the photo management and sharing tasks, many application tools have been developed to generate pleasant photo slideshows, collages, or scrapbooks by applying simple templates/layouts and visual effects. In this work, we propose a convenient digital scrapbook generating system for travel photos, named as IS-Scrapbook, which keeps the virtues while dismisses the drawbacks of conventional digital photo presentation styles. The IS-Scrapbook system has three friendly attributes compared to other photo presentation tools. First, aiming to attract the audience, we highlight objects more meaningful or familiar to the viewer, e.g. the landmark and the protagonist in the photos. Second, five basic design principles of typography, i.e. proximity, contrast, balance, color harmony and repetition, are applied to produce more vivacious layouts. Third, the system automatically generates a digital scrapbook with a default layout, and the user can further adjust photo positions and enrich each page by sketching or inserting dialog bubbles with the aid of the developed user interface. User study shows that the proposed work enhances the experiences of photo browsing and gives a brand-new way of photo sharing. Jung-Yu Yeh, Min-Chun Hu 0001, Wen-Huang Cheng, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2011 | Robust Camera Calibration and Player Tracking in Broadcast Basketball VideoabstractWith the growth of fandom population, a considerable amount of broadcast sports videos have been recorded, and a lot of research has focused on automatically detecting semantic events in the recorded video to develop an efficient video browsing tool for a general viewer. However, a professional sportsman or coach wonders about high level semantics in a different perspective, such as the offensive or defensive strategy performed by the players. Analyzing tactics is much more challenging in a broadcast basketball video than in other kinds of sports videos due to its complicated scenes and varied camera movements. In this paper, by developing a quadrangle candidate generation algorithm and refining the model fitting score, we ameliorate the court-based camera calibration technique to be applicable to broadcast basketball videos. Player trajectories are extracted from the video by a CamShift-based tracking method and mapped to the real-world court coordinates according to the calibrated results. The player position/trajectory information in the court coordinates can be further analyzed for professional-oriented applications such as detecting wide open event, retrieving target video clips based on trajectories, and inferring implicit/explicit tactics. Experimental results show the robustness of the proposed calibration and tracking algorithms, and three practicable applications are introduced to address the applicability of our system. Min-Chun Hu 0001, Ming-Hsiu Chang, Ja-Ling Wu, Lin Chi |
IEEE Trans. Multim. | 3 |
| 2010 | Performance improvement of distributed video coding by using block mode selectionabstractBlock mode selection is one new way to improve the performance of distributed video coding (DVC). Since there are many factors influencing the correctness of the block mode selection, the decision of block mode is not an easy work. In this paper, a low complexity block mode selection model is proposed to select the block modes more correctly in the WZ (Wyner-Ziv) frame. The proposed block mode selection modules can increase the RD performance up to 2.8 dB as compared with the traditional DVC codecs; moreover, the subjective quality can also be improved by using the proposed deblocking filter when the objective quality is kept the same. Bo-Ruei Chiou, Yun-Chung Shen, Han-Ping Cheng, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2010 | Fast decoding for LDPC based distributed video codingabstractDistributed video coding (DVC) is a new coding paradigm targeting applications with the need for low-complexity encoding at the cost of a higher decoding complexity. In the DVC architecture based on a feedback channel, the high decoding complexity is mainly due to the request-decode operation with repetitively fixed step size (induced by Slepian-Wolf decoding). In this paper, a parallel message-passing decoding algorithm for low density parity check (LDPC) syndrome is applied through Compute Unified Device Architecture (CUDA) based on General-Purpose Graphics Processing Unit (GPGPU). Furthermore, we propose an approach to reduce the number of requests dubbed as Ladder Request Step Size (LRSS) which leads to more speedup gain. Experimental results show that, through our work, significant speedup in decoding time is achieved with negligible loss in rate-distortion (RD) performance. Yu-Shan Pai, Han-Ping Cheng, Yun-Chung Shen, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2010 | Virtual spotlighted advertising for tennis videos
Chia-Hu Chang, Kuei-Yi Hsieh, Ming-Che Chiang, Ja-Ling Wu |
J. Vis. Commun. Image Represent. | 4 |
| 2009 | Pushing information over acoustic channelsabstractWe propose a unidirectional communication system capable of pushing information from audio speakers to nearby audience via acoustic channels. Though constrained by short transmission distances and low data throughput of sound waves, information push services based on acoustic channels can be regarded as a cost-free message delivery scheme since music playback systems and handheld devices equipped with recording capabilities have become ubiquitous in our life. Furthermore, the proposed approach can provide better non-intrusiveness and more convenient user access than other auxiliary information delivery approaches like 2-D barcodes or watermarking schemes based on visual contents. To the best of our knowledge, our system is the first one that can effectively push information to receivers at feasible transmission throughput/robustness/distance within common noisy environments. Various interesting applications and new business models can be facilitated with the proposed scheme. Po-Wei Chen, Chun-Hsiang Huang, Yun-Chung Shen, Ja-Ling Wu |
ICASSP | 4 |
| 2009 | Billiards wizard: A tutoring system for broadcasting nine-ball billiards videosabstractIn this work, we propose a framework to build a billiards tutoring system based on broadcasting nine-ball video analysis. A robust table detection module is developed by mapping the displayed video frames to a predefined billiard table model. In addition, we detect balls and trace their positions at every time instant. The real-world spatial relationships between the table and the balls are used to provide the aiming and position play suggestions. Ball position information is also utilized to distinguish each play into corresponding event by a rule-based method. The experimental results are encouraging and are more comprehensive than existing works of billiards video analysis. Chen-Wei Chou, Min-Chun Hu 0001, Ja-Ling Wu |
ICASSP | 3 |
| 2009 | Real-time Gender Classification from Human Gait for Arbitrary View AnglesabstractIn this paper, we investigate an important but understudied problem, gender classification from human gaits. And we have proved the ability of using GEI (Gait Energy Image) as a representation of human gait for arbitrary view angles. Using GEI as a discriminative feature, we construct angle classifiers and gender classifiers from different approaches. Experiments show that our system achieved a good performance in real-time and is able to be applied to real-world application. Ping-Chieh Chang, Min-Chun Hu 0001, Ja-Ling Wu, Chuan-Shen Hu |
ISM | 3 |
| 2009 | WOW: wild-open warning for broadcast basketball video based on player trajectoryabstractIn basketball games, wild-open means that there is an offensive player not well defended by his/her opponents. The occurrence of wild-open usually implies the existence of a successful offense tactic. In this paper, a Wild-Open Warning (WOW) system is designed to assist basketball coaches/players in revealing possible tactics of their opponents through watching the broadcast game videos. The system automatically extracts semantic objects such as the court and the players in the video, and calibrates the players' positions to the real-world court coordinates. A robust and efficient algorithm for court detection and camera calibration is proposed for basketball videos. Wild-open is detected when the position of an offensive player satisfies three predefined criteria. In the mean time, the system will mark the wild-open players to warn the viewers such that they should keep attention to certain players. Ming-Hsiu Chang, Min-Chun Hu 0001, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2009 | Evolution-based virtual content insertionabstractThis demonstration presents an innovative framework for virtual content insertion in an interactive way. The virtual content is inserted into video with evolved animations according to predefined behaviors, and generates interactions with video contents. In order to reduce the intrusiveness and improve the impression for viewers, the evolution process is divided into distinct yet dependent phases, in which the virtual content evolves its appearances and behaviors according to the incremental interactions. Therefore, the augmented videos generated by the proposed system establish visually relevant connection between the inserted virtual content and the source videos, and increase the acceptability and the attractiveness at the same time. Moreover, the proposed system enables video owners to create entertaining personalized videos effectively, and engages viewers with the storyline. Ming-Che Chiang, Chia-Hu Chang, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2009 | Scalable computation for spatially scalable video coding using NVIDIA CUDA and multi-core CPUabstractThe scalable video coding (SVC), an extension of H.264/MPEG4-AVC (H.264), was standardized in 2007 by Joint Video Team (JVT). SVC provides spatial, temporal and SNR scalabilities. To achieve these scalabilities, SVC uses additional coding tools and coding modes based on H.264. The coding tools used by SVC and the variety coding modes decision make the corresponding coding complexity become extremely high, so real-time realization of SVC is nearly impossible by using software and single-core CPU only. One possible solution to generate SVC streams in real-time is to parallelize the whole encoding process. Currently, multi-core CPU and GPU are two popular kinds of parallel processing architectures. Not much research has been devoted to realize the parallel SVC encoders based on the co-work of these two architectures. In this paper, a scalable computation model for spatial SVC using multi-core CPU and GPGPU through NVIDIA CUDA is proposed. On the basis of the proposed computational model, a solution to solve the challenging data transition problem (will be detailed later) of this CPU-GPU co-work architecture is then provided. Simulation results show that, through our work, significant speed up gain in spatial SVC encoding can be achieved. Yen-Lin Huang, Yun-Chung Shen, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2009 | Sports wizard: sports video browsing based on semantic concepts and game structureabstractA convenient video browsing system for popular sports such as baseball, tennis and billiards is proposed in this work. A sports video is automatically segmented into video clips and annotated with semantic concepts in terms of event type. For sports with specific game structure, e.g. baseball and tennis, we further analyze the entire video of the match with the aid of caption information, webcast information, and the domain knowledge of the corresponding sport. Each segmented video clip with semantic annotation is then mapped to a certain part of the game structure. Finally, the proposed system, name as Sports Wizard, provides a favorable way for the user to browse sports videos based on semantic concepts or game structure. Min-Chun Hu 0001, Yin-Tzu Lin, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2009 | Interactive background blurringabstractPhotographers usually take a shallow focus image to highlight the main subject in the picture and blur the distractions in the background region. Softening these distractions can not only produce special vision effect but also keep the privacy for people in the background. In this work, we develop a brand new defocusing algorithm to simulate the blur effect in the background caused by a wide aperture camera. We also design an easy-to-use interactive interface for user to segment foreground/background objects, adjust the depth of field for images, and assign camera settings. Techniques including lazy snapping, alpha matting, depth information generation and defocusing are integrated in the system. The experimental results show the effectiveness of the proposed methodology. Chih-Yu Yan, Min-Chun Hu 0001, Ja-Ling Wu |
ACM Multimedia | 3 |
| 2009 | Video-Based Motion Capturing for Skeleton-Based 3D Models
Liang-Yu Shih, Bing-Yu Chen 0004, Ja-Ling Wu |
PSIVT | 3 |
| 2009 | Building term suggestion relational graphs from collective intelligenceabstractThis paper proposes an effective approach to provide relevant search terms for conceptual Web search. 'Semantic Term Suggestion' function has been included so that users can find the most appropriate query term to what they really need. Conventional approaches for term suggestion involve extracting frequently occurring key terms from retrieved documents. They must deal with term extraction difficulties and interference from irrelevant documents. In this paper, we propose a semantic term suggestion function called Collective Intelligence based Term Suggestion (CITS). CITS provides a novel social-network based framework for relevant terms suggestion with a semantic graph of the search term without limiting to the specific query term. A visualization of semantic graph is presented to the users to help browsing search results from related terms in the semantic graph. The search results are ranked each time according to their relevance to the related terms in the entire query session. Comparing to two popular commercial search engines, a user study of 18 users on 50 search terms showed better user satisfactions and indicated the potential usefulness of proposed method in real-world search applications. Jyh-Ren Shieh, Yung-Huan Hsieh, Yang-Ting Yeh, Tse-Chung Su, Ching-Yung Lin, Ja-Ling Wu |
WWW | 6 |
| 2009 | Fidelity-guaranteed robustness enhancement of blind-detection watermarking schemes
Chun-Hsiang Huang, Ja-Ling Wu |
Inf. Sci. | 2 |
| 2009 | Unseen visible watermarking: a novel methodology for auxiliary information delivery via visual contentsabstractA novel data hiding scheme, denoted as unseen visible watermarking (UVW), is proposed. In UVW schemes, hidden information can be embedded covertly and then directly extracted using the human visual system as long as appropriate operations (e.g., gamma correction provided by almost all display devices or changes in viewing angles relative to LCD monitors) are performed. UVW eliminates the requirement of invisible watermarking that specific watermark extractors must be deployed to the receiving end in advance, and it can be integrated with 2-D barcodes to transmit machine-readable information that conventional visible watermarking schemes fail to deliver. We also adopt visual cryptographic techniques to guard the security of hidden information and, at the same time, increase the practical value of visual cryptography. Since UVW can be alternatively viewed as a mechanism for visualizing patterns hidden with least-significant-bit embedding, its security against statistical steganalysis is proved by empirical tests. Limitations and other potential extensions of UVW are also addressed. Chun-Hsiang Huang, Shang-Chih Chuang, Yen-Lin Huang, Ja-Ling Wu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2009 | RoleNet: Movie Analysis from the Perspective of Social NetworksabstractWith the idea of social network analysis, we propose a novel way to analyze movie videos from the perspective of social relationships rather than audiovisual features. To appropriately describe role's relationships in movies, we devise a method to quantify relations and construct role's social networks, called RoleNet. Based on RoleNet, we are able to perform semantic analysis that goes beyond conventional feature-based approaches. In this work, social relations between roles are used to be the context information of video scenes, and leading roles and the corresponding communities can be automatically determined. The results of community identification provide new alternatives in media management and browsing. Moreover, by describing video scenes with role's context, social-relation-based story segmentation method is developed to pave a new way for this widely-studied topic. Experimental results show the effectiveness of leading role determination and community identification. We also demonstrate that the social-based story segmentation approach works much better than the conventional tempo-based method. Finally, we give extensive discussions and state that the proposed ideas provide insights into context-based video analysis. Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu |
IEEE Trans. Multim. | 3 |
| 2008 | Collusion-resistant video fingerprinting based on temporal oscillationabstractMost researches of collusion-resistant multimedia fingerprinting consider only the designing of the traceability code. In this paper, we not only discuss the traceability of the fingerprints, but also try to prevent collusion attacks by making the colluded version of the video useless. The proposed collusion-resistant video fingerprinting scheme embeds a curve fingerprint into the video by temporally oscillating the video frames. This curve is smooth enough so that the oscillation of the video is imperceptible. However, once the colluders average or copy-and-paste the video frames, the fluctuation of the video will be so severe that users will feel uncomfortable while watching the video. In addition, a shot-based fingerprinting is employed to embed a hierarchical traceability code to increase the collusion-resistance. Experimental results show good performances of the proposed scheme. That is, the proposed system not only can trace the colluders successfully, but also can discourage users from removing the traceability code by collusion attacks. Yu-Tzu Lin, Chun-Hsiang Huang, Ja-Ling Wu |
ICIP | 3 |
| 2008 | Traceable multimedia fingerprinting based on the multilevel user groupingabstractThis paper presents a traceable fingerprinting scheme based on a multilevel user grouping strategy. In the proposed code structure, coding efforts for groups of users are apportioned to different levels of the code. Instead of making efforts to find one “high-performance” traceability code, we try to derive a generalized framework of fingerprints construction which hierarchically constructs a traceability code having “high-performances” by composing several “moderate-performance” traceability codes together. Both theoretical analyses and experimental results are provided to show good performances of the proposed fingerprinting system, that is, it can serve a large amount of users with a much shorter code length and high traceability. Yu-Tzu Lin, Ja-Ling Wu |
ICME | 2 |
| 2008 | Using Semantic Graphs for Image SearchabstractIn this paper, we propose a Semantic Graphs for Image Search (SGIS) system, which provides a novel way for image search by utilizing collaborative knowledge in Wikipedia and network analysis to form semantic graphs for search-term suggestion. The collaborative article editing process of Wikipediapsilas contributors is formalized as bipartite graphs that are folded into networks between terms. When user types in a search term, SGIS automatically retrieves an interactive semantic graph of related terms that allow users easily find related images not limited to a specific search term. Interactive semantic graph then serves as an interface to retrieve images through existing commercial search engines. This method significantly saves userspsila time by avoiding multiple search keywords that are usually required in generic search engines. It benefits both naive user who does not possess a large vocabulary (e.g., students) and professionals who look for images on a regular basis. In our experiments, 85% of the participants favored SGIS system than commercial search engines. Jyh-Ren Shieh, Yang-Ting Yeh, Chih-Hung Lin, Ching-Yung Lin, Ja-Ling Wu |
ICME | 5 |
| 2008 | Event detection in tennis matches based on video data miningabstractThis paper proposes a mining-based method to achieve event detection for broadcasting tennis videos. Utilizing visual and aural information, we extract some high-level features to describe video segments. The audiovisual features are further transformed to symbolic streams and an efficient mining technique is applied to derive all frequent patterns that characterize tennis events. After mining, we categorize frequent patterns into several kinds of events and therefore achieve event detection for tennis videos by checking the correspondence between mined patterns and events. The experimental results show that the proposed approach is a promising way to detect events in broadcasting tennis video. Min-Chun Hu 0001, Yi-Tang Wang, Chen-Wei Chou, Kuei-Yi Hsieh, Wei-Ta Chu, Ja-Ling Wu |
ICME | 6 |
| 2008 | ViSA: virtual spotlighted advertisingabstractWith the constraint of limited intrusiveness, maximizing the advertising efficiency in sports video has been known as a challenging problem. In this paper, we propose a virtual advertisement system with novel advertising strategy, called Virtual Spotlighted Advertising (ViSA), for broadcasting tennis videos. We take psychology, advertising theory, and computational aesthetics into account for improving the effectiveness of advertising. ViSA system automatically detects the candidate insertion points in both temporal and spatial domains and estimates the maximum effective region for message communication. Then, the harmonically re-colored advertisements with foveation model based non-uniform transparency, are projected on the court in tennis videos. The experiments and evaluation results showed the effectiveness of ViSA, for sports video advertising, in terms of recall and recognition. Chia-Hu Chang, Kuei-Yi Hsieh, Ming-Che Chung, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2008 | SheepDog: group and tag recommendation for flickr photos by automatic search-based learningabstractOnline photo albums have been prevalent in recent years and have resulted in more and more applications developed to provide convenient functionalities for photo sharing. In this paper, we propose a system named SheepDog to automatically add photos into appropriate groups and recommend suitable tags for users on Flickr. We adopt concept detection to predict relevant concepts of a photo and probe into the issue about training data collection for concept classification. From the perspective of gathering training data by web searching, we introduce two mechanisms and investigate their performances of concept detection. Based on some existing information from Flickr, a ranking-based method is applied not only to obtain reliable training data, but also to provide reasonable group/tag recommendations for input photos. We evaluate this system with a rich set of photos and the results demonstrate the effectiveness of our work. Hong-Ming Chen, Ming-Hsiu Chang, Ping-Chieh Chang, Min-Chun Hu 0001, Winston H. Hsu, Ja-Ling Wu |
ACM Multimedia | 6 |
| 2008 | Photo navigatorabstractNowadays, travel has become a popular activity for people to relax their body and mind. Taking photos is then often an inevitable and frequent event during one's trip for recording the enjoyable experience. To help people to relive the wonderful travel experience they had recorded in photos, this paper presents a system, Photo Navigator, for enhancing the photo browsing experience by creating a new browsing style with a realistic feel to users as being into the scenes and taking a trip back in time to revisit the place. The proposed system is characterized by two main features. First, it better reveals the spatial relations among photos and offers a strong sense of space by taking users to fly into the scenes. Second, it is fully automatic and makes plausible for novice users to utilize the 3D technologies that are traditionally complex to manipulate. The proposed system is compared with two other photo browsing tools, ACDSee's photo slideshow and Microsoft's PhotoStory. User studies show that people would comparatively favor the browsing style we offer and appreciate the ease to create such a style. Chi-Chang Hsieh, Wen-Huang Cheng, Chia-Hu Chang, Yung-Yu Chuang, Ja-Ling Wu |
ACM Multimedia | 5 |
| 2008 | Video-based CPR analysis systemabstractIn this paper, we propose a system for automatically detecting cardiopulmonary resuscitation (CPR) events and analyzing CPR qualities in surveillance videos. The system is further applied to the training of healthcare providers in the emergency room. Instructors could more efficiently evaluate the CPR quality performed by a medical team with the aid of our system. We extract motion vectors in all image blocks, and take motion information as a clue to classify video sequences into CPR and non-CPR segments based on Support Vector Machine (SVM). We further analyze several indicators of CPR quality and attach CPR information to the video. In order not to infringe the privacy of the patient, we apply a mosaic mechanism to mask the skin color regions of the patient. Our system has been applied to several simulated CPR video sequences and the results show acceptable accuracy in CPR detection and CPR information measurement. Yi-Tang Wang, Min-Chun Hu 0001, Ja-Ling Wu, Chih-Wei Yang, Matthew Huei-Ming Ma |
ACM Multimedia | 3 |
| 2008 | Collaborative knowledge semantic graph image searchabstractIn this paper, we propose a Collaborative Knowledge Semantic Graphs Image Search (CKSGIS) system. It provides a novel way to conduct image search by utilizing the collaborative nature in Wikipedia and by performing network analysis to form semantic graphs for search-term expansion. The collaborative article editing process used by Wikipedia's contributors is formalized as bipartite graphs that are folded into networks between terms. When a user types in a search term, CKSGIS automatically retrieves an interactive semantic graph of related terms that allow users to easily find related images not limited to a specific search term. Interactive semantic graph then serve as an interface to retrieve images through existing commercial search engines. This method significantly saves users' time by avoiding multiple search keywords that are usually required in generic search engines. It benefits both naïve users who do not possess a large vocabulary and professionals who look for images on a regular basis. In our experiments, 85% of the participants favored CKSGIS system rather than commercial search engines. Jyh-Ren Shieh, Yang-Ting Yeh, Chih-Hung Lin, Ching-Yung Lin, Ja-Ling Wu |
WWW | 5 |
| 2008 | Explicit semantic events detection and development of realistic applications for broadcasting baseball videos
Wei-Ta Chu, Ja-Ling Wu |
Multim. Tools Appl. | 2 |
| 2008 | Two algorithms for constructing efficient huffman-code based reversible variable length codesabstractIn this paper, Huffman-code based reversible variable length code (RVLC) construction algorithms are studied. We use graph models to represent the prefix, suffix and Hamming distance relationships among RVLC candidate codewords. The properties of the so-obtained graphs are investigated in detail, based on which we present two efficient RVLC construction algorithms: algorithm 1 aims at minimizing the average codeword length while algorithm 2 jointly minimizes the average codeword length and maximizes the error detection probability at the same time. Chia-Wei Lin, Ja-Ling Wu, Yuh-Jue Chuang |
IEEE Trans. Commun. | 2 |
| 2008 | Semantic Analysis for Automatic Event Recognition and Segmentation of Wedding Ceremony VideosabstractWedding is one of the most important ceremonies in our lives. It symbolizes the birth and creation of a new family. In this paper, we present a system for automatically segmenting a wedding ceremony video into a sequence of recognizable wedding events, e.g., the couple's wedding kiss. Our goal is to develop an automatic tool that helps users to efficiently organize, search, and retrieve his/her treasured wedding memories. Furthermore, the obtained event descriptions could benefit and complement the current research in semantic video understanding. Based on the knowledge of wedding customs, a set of audiovisual features, relating to the wedding contexts of speech/music types, applause activities, picture-taking activities, and leading roles, are exploited to build statistical models for each wedding event. Thirteen wedding events are then recognized by a hidden Markov model, which takes into account both the fitness of observed features and the temporal rationality of event ordering to improve the segmentation accuracy. We conducted experiments on a collection of wedding videos and the promising results demonstrate the effectiveness of our approach. Comparisons with conditional random fields show that the proposed approach is more effective in this application domain. Wen-Huang Cheng, Yung-Yu Chuang, Yin-Tzu Lin, Chi-Chang Hsieh, Shao-Yen Fang, Bing-Yu Chen 0004, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2008 | Memory Efficient Hierarchical Lookup Tables for Mass Arbitrary-Side Growing Huffman Trees DecodingabstractThis paper addresses the optimization problem of minimizing the number of memory access subject to a rate constraint for any Huffman decoding of various standard codecs. We propose a Lagrangian multiplier based penalty-resource metric to be the targeting cost function. To the best of our knowledge, there is few related discussion, in the literature, on providing a criterion to judge the approaches of entropy decoding under resource constraint. The existing approaches which dealt with the decoding of the single-side growing Huffman tree may not be memory-efficient for arbitrary-side growing Huffman trees adopted in current codecs. By grouping the common prefix part of a Huffman tree, in stead of the commonly used single-side growing Huffman tree, we provide a memory efficient hierarchical lookup table to speed up the Huffman decoding. Simulation results show that the proposed hierarchical table outperforms previous methods. A Viterbi-like algorithm is also proposed to efficiently find the optimal hierarchical table. More importantly, the Viterbi-like algorithm obtains the same results as that of the brute-force search algorithm. Sung-Wen Wang, Ja-Ling Wu, Shang-Chih Chuang, Chih-Chieh Hsiao, Yi-Shin Tung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Digital-Invisible-Ink Data Hiding Based on Spread-Spectrum and Quantization TechniquesabstractA novel data-hiding methodology, denoted as digital invisible ink (DII), is proposed to implement secure steganography systems. Like the real-world invisible ink, secret messages will be correctly revealed only after the marked works undergo certain prenegotiated manipulations, such as lossy compression and processing. Different from conventional data-hiding schemes where content processing or compression operations are undesirable, distortions caused by prenegotiated manipulations in DII-based schemes are indispensable steps for revealing genuine secrets. The proposed scheme is carried out based on two important data-hiding schemes: spread-spectrum watermarking and frequency-domain quantization watermarking. In some application scenarios, the DII-based steganography system can provide plausible deniability and enhance the secrecy by taking cover with other messages. We show that DII-based schemes are indeed superior to existing plausibly deniable steganography approaches in many aspects. Moreover, potential security holes caused by deniable steganography systems are discussed. Chun-Hsiang Huang, Shang-Chih Chuang, Ja-Ling Wu |
IEEE Trans. Multim. | 3 |
| 2007 | Unseen Visible WatermarkingabstractA novel data-hiding methodology, denoted as unseen visible watermarking (UVW), is proposed. The proposed scheme is inspired by real-world watermarks and possesses advantages of both visible and invisible watermarking schemes. After watermark embedding, the differences between the original work and the stego work are imperceptible under normal viewing conditions. However, when the hidden message is to be extracted, no explicit watermark extracting module is required. Semantically-meaningful watermark patterns can be directly recognized from the stego work as long as common imaging-related functions, e.g. gamma-correction or even simply changing the user-viewing angle relative to the LCD monitor, are performed. The proposed scheme outperforms existing invisible watermarking methods in its capability to practically convey metadata to users of legacy display devices lacking renewal capability. On the other hand, it does not suffer from the annoying quality-degradation problem of visible watermarking schemes. Limitations and possible extensions of the proposed schemes are also addressed. We believe that many interesting new applications can be facilitated using such unseen visible watermarking schemes. Shang-Chih Chuang, Chun-Hsiang Huang, Ja-Ling Wu |
ICIP (3) | 3 |
| 2007 | Exploring Broadcasting Baseball Videos Based on Multimodal and Multidisciplinary StudyabstractThis demonstration presents a comprehensive work covering semantic event detection, summary/highlight generation, query answering, and efficient browsing, for broadcasting baseball videos. This system integrates the techniques of content-based analysis, key-phrase spotting, natural language processing, data mining, and human-computer interface. It presents a multimodal and multidisciplinary work to show an instance for a media-rich life. Wei-Ta Chu, Ja-Ling Wu |
ICME | 2 |
| 2007 | Movie Analysis Based on Roles' Social NetworkabstractRoles in a movie form a small society and their interrelationship provides clues for movie understanding. Based on this observation, we present a new viewpoint to perform semantic movie analysis. Through checking the co-occurrence of roles in different scenes, we construct a roles' social network to describe their relationships. We introduce the concept of social network analysis to elaborately identify leading roles and the hidden communities. With the results of community identification, we perform storyline detection that facilitates more flexible movie browsing and higher-level movie analysis. The experimental results show that the proposed community identification method is accurate and is robust to errors. Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu |
ICME | 3 |
| 2007 | A Parallel Algorithm for H.264/AVC Deblocking Filter Based on Limited Error Propagation EffectabstractIn this paper, we propose a parallel algorithm for H.264/AVC deblocking filter which is scalable to the number of processors. Unlike the conventional approach, which is limited by the independent data units, the designed algorithm allows issuing dependent data units concurrently to decrease the penalty from synchronization of data units. For the general-purpose dual-core processors, experimental results show that our method speeds up 1.72 and 1.39 times as compared with optimized sequential method and the well-known wavefront parallelizing method, respectively. Shu-Sian Yang, Sung-Wen Wang, Ja-Ling Wu |
ICME | 3 |
| 2007 | Film Narrative Exploration Through the Analysis of Aesthetic Elements
Chia-Wei Wang, Wen-Huang Cheng, Jun-Cheng Chen, Shu-Sian Yang, Ja-Ling Wu |
MMM (1) | 5 |
| 2007 | Two Algorithms for Constructing Efficient Huffman-Code-Based Reversible Variable Length CodesabstractIn this paper, Huffman-code-based reversible variable length code (RVLC) construction algorithms are studied. We use graph models to represent the prefix, suffix, and Hamming distance relationships among RVLC candidate code words. The properties of the so-obtained graphs are investigated in detail, based on which we present two efficient RVLC construction algorithms: Algorithm 1 aims at minimizing the average code word length while Algorithm 2 jointly minimizes the average code word length and maximizes the error-detection probability at the same time. Chia-Wei Lin, Ja-Ling Wu, Yuh-Jue Chuang |
IEEE Trans. Commun. | 2 |
| 2007 | Video Adaptation for Small Display Based on Content RecompositionabstractThe browsing of quality videos on small hand-held devices is a common scenario in pervasive media environments. In this paper, we propose a novel framework for video adaptation based on content recomposition. Our objective is to provide effective small size videos which emphasize the important aspects of a scene while faithfully retaining the background context. That is achieved by explicitly separating the manipulation of different video objects. A generic video attention model is developed to extract user-interest objects, in which a high-level combination strategy is proposed for fusing the adopted three types of visual attention features: intensity, color, and motion. Based on the knowledge of media aesthetics, a set of aesthetic criteria is presented. Accordingly, these objects are well reintegrated with the direct-resized background to optimally match the specific screen sizes. Experimental results demonstrate the efficiency and effectiveness of our approach Wen-Huang Cheng, Chia-Wei Wang, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Quality Enhancement of Frame Rate Up-Converted Video by Adaptive Frame Skip and Reliable Motion ExtractionabstractFrame rate up-conversion is a postprocessing tool to convert the frame rate from a lower number to a higher one. It is a useful technique for a lot of practical applications, such as display format conversion, low bit rate video coding and slow motion playback. Unlike traditional approaches, such as frame repetition or linear frame interpolation, motion-compensated frame interpolation (MCFI) technique is more efficient since it takes block motion into account. In this paper, by considering the deficiencies of previous works, new criteria and coding schemes for improving motion derivation and interpolation processes are proposed. Next, for video coding applications, adaptive frame skip is executed at the encoder side to maximize the power of MCFI so that the quality of interpolated frames is guaranteed. Experimental results show that our proposal effectively enhances the overall quality of the frame rate up-converted video sequence, both subjectively and objectively. Ya-Ting Yang, Yi-Shin Tung, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Digital invisible ink: revealing true secrets via attackingabstractA novel steganographic approach analogy to the real-world secret communication mechanism, in which secret messages are written on white papers using invisible ink like lemon juice and are revealed only after the papers are heated, is proposed. Carefully-designed informed embedders now play the role of "invisible ink"; some pre-negotiated attacks provided by common content-processing tools correspond to the required "heating" process. Theoretic models and feasible implementations of the proposed digital-invisible-ink watermarking approach based on both blind-detection spread-spectrum watermarking and quantization watermarking schemes are provided. The proposed schemes can prevent the supervisor from interpreting secret messages even when the watermark extractor and decryption tool, as well as session keys, are available to the supervisor. Furthermore, secret communication systems employing the proposed scheme can aggressively mislead the channel supervisor with fake watermarks and transmit genuine secrets at the same time. Chun-Hsiang Huang, Yu-Feng Kuo, Ja-Ling Wu |
AsiaCCS | 3 |
| 2006 | Extraction of Baseball Trajectory and Physics-Based Validation for Single-View Baseball Video SequencesabstractTo enrich the viewing experience of baseball games and provide some clues for enhancing pitcher's performance, we propose a Kalman filter-based approach to track ball trajectory from single-view pitching sequences. Without setting extraordinary equipments in stadiums or other sensing instruments, this approach robustly extracts ball trajectory for pitching sequences captured from TV channels or downloaded from the Internet. To validate the detected ball trajectories, we investigate the characteristics of ball trajectories on the basis of a baseball physical model. The effectiveness of ball trajectory extraction and ball position detection are presented Wei-Ta Chu, Chia-Wei Wang, Ja-Ling Wu |
ICME | 3 |
| 2006 | An Efficient Memory Construction Scheme for an Arbitrary Side Growing Huffman TableabstractBy grouping the common prefix of a Huffman tree, in stead of the commonly used single-side rowing Huffman tree (SGH-tree), we construct a memory efficient Huffman table on the basis of an arbitrary-side growing Huffman tree (AGH-tree) to speed up the Huffman decoding. Simulation results show that, in Huffman decoding, an AGH-tree based Huffman table is 2.35 times faster that of the Hashemian's method (an SGH-tree based one) and needs only one-fifth the corresponding memory size. In summary, a novel Huffman table construction scheme is proposed in this paper which provides better performance than existing construction schemes in both decoding speed and memory usage Sung-Wen Wang, Shang-Chih Chuang, Chih-Chieh Hsiao, Yi-Shin Tung, Ja-Ling Wu |
ICME | 5 |
| 2006 | A Colorization Based Animation Broadcast System with Traitor Tracing Capability
Chih-Chieh Liu, Yu-Feng Kuo, Chun-Hsiang Huang, Ja-Ling Wu |
IWDW | 4 |
| 2006 | Tiling slideshowabstractThis paper presents a new medium, called tiling slideshow, to display photos in a tile-like manner, coordinating with the pace of background music. In contrast to the conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Motivated by the concepts of technical writing, each displaying layout is composed of a larger topic photo and several small-size supportive photos. Based on this idea, the proposed tiling slideshow system consists of three major components: image clustering, music analyzer, and layout organizer. Given the limited displaying space, we consider the context and relationship between photos and model the layout organization as a constrainted optimization problem. Experiments on real consumer photograph collections show that the novel displaying method gives users more pleasant browsing experience than the methods that focus only on single photograph display. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 5 |
| 2006 | Audiovisual slideshow: present your journey by photosabstractThis demonstration presents a novel way to systematically display photos and enhance the viewing experience of photo browsing. In contrast to conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Moreover, the displaying pace is coordinated with the beat of the user-selected incidental music. To automatically generate the audiovisual slideshow, we develop a system that consists of three main components: photo analysis, music analysis, and audiovisual composition. Audiovisual content analysis and cross-media synchronization issues are addressed in this work. This novel demonstration is especially suitable to present photos taken in a journey. It vigorously presents the delights of traveling and helps us recall or experience the trip. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 5 |
| 2006 | Development of realistic applications based on explicit event detection in broadcasting baseball videosabstractThis paper presents a framework that explicitly detects events in broadcasting baseball videos and facilitates the development of various extended applications. Three phases are included: reliable shot classification, explicit event detection, and elaborate applications. In the shot classification stage, color and geometric information are utilized to classify shots into several canonical views. To explicitly detect semantic events, rule-based decision and model-based decision methods are developed. We emphasize that this system efficiently and exactly identifies what happened in baseball games rather than roughly finding some interesting parts. Based on explicit event detection, many accurate and practical applications such as box score generation and game summarization can be built. The evaluation results show the effectiveness of the proposed methods and demonstrate some insights about bridging semantic gaps for sports videos. Wei-Ta Chu, Ja-Ling Wu |
MMM | 2 |
| 2005 | A Hierarchical and Multi-Model Based Algorithm for Lead Detection and News Program Narrative ParsingabstractIn this paper, a hierarchical and multi-modal based news item detection algorithm, which can be viewed as a mid-stage solution between the single-modal and the semantic-based approaches, is proposed for parsing TV news program videos. We investigate the production model of TV news program first and then make use of the so-obtained domain knowledge to develop the proposed algorithm. With the add of multi-modal features, such as volume and zero crossing rate in audios and keyframe and human face in videos, the proposed algorithm showed rather satisfactory results in both precision and recall measures for parsing a 6-hour news program test video. Jin-Hau Kuo, Jen-Bin Kuo, Hsuan-Wei Chen, Ja-Ling Wu |
AINA | 4 |
| 2005 | Integration of rule-based and model-based decision methods for baseball event detectionabstractTo exactly detect what events occur in baseball games, a framework that integrates rule-based and model-based decision methods is proposed. The rule-based decision module infers what happened by checking the information changes in the caption. The model-based decision module further classifies the events that could not be explicitly determined by checking caption information only. Thirteen events, including hit, double, home run, and so on, are considered in this work. The promising experimental results show the effectiveness of the proposed framework and facilitate the development of advanced video applications. Wei-Ta Chu, Ja-Ling Wu |
ICME | 2 |
| 2005 | An adaptive edge detection based colorization algorithm and its applicationsabstractColorization is a computer-assisted process for adding colors to grayscale images or movies. It can be viewed as a process for assigning a three-dimensional color vector (YUV or RGB) to each pixel of a grayscale image. In previous works, with some color hints the resultant chrominance value varies linearly with that of the luminance. However, it is easy to find that existing methods may introduce obvious color bleeding, especially, around region boundaries. It then needs extra human-assistance to fix these artifacts, which limits its practicability. Facing such a challenging issue, we introduce a general and fast colorization methodology with the aid of an adaptive edge detection scheme. By extracting reliable edge information, the proposed approach may prevent the colorization process from bleeding over object boundaries. Next, integration of the proposed fast colorization scheme to a scribble-based colorization system, a modified color transferring system and a novel chrominance coding approach are investigated. In our experiments, each system exhibits obvious improvement as compared to those corresponding previous works. Yi-Chin Huang, Yi-Shin Tung, Jun-Cheng Chen, Sung-Wen Wang, Ja-Ling Wu |
ACM Multimedia | 5 |
| 2005 | Generative and Discriminative Modeling toward Semantic Context Detection in Audio TracksabstractSemantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. We propose a hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. For action movies we focused in this work, hidden Markov models (HMMs) are used to model four representative audio events, i.e. gunshot, explosion, car-braking, and engine sounds. At the semantic context level, generative (ergodic hidden Markov model) and discriminative (support vector machine, SVM) approaches are investigated to fuse the characteristics and correlations among various audio events, which provide cues for detecting gunplay and car-chasing scenes. The experimental results demonstrate the effectiveness of the proposed approaches and draw a sketch for semantic indexing and retrieval. Moreover, the differences between two fusion schemes are discussed to be the reference for future research. Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu |
MMM | 3 |
| 2005 | Toward semantic indexing and retrieval using hierarchical audio models
Wei-Ta Chu, Wen-Huang Cheng, Yung-Jen Hsu 0001, Ja-Ling Wu |
Multim. Syst. | 4 |
| 2005 | On error detection and error synchronization of reversible variable-length codesabstractReversible variable-length codes (RVLCs) are not only prefix-free but also suffix-free codes. Due to the additional suffix-free condition, RVLCs are usually nonexhaustive codes. When a bit error occurs in a sentence from a nonexhaustive RVLC, it is possible that the corrupted sentence is not decodable. The error is said to be detected in this case. We present a model for analyzing the error detection and error synchronization characteristics of nonexhaustive VLCs. Six indices, the error detection probability, the mean and the variance of forward error detection delay length, the error synchronization probability, the mean and the variance of forward error synchronization delay length are formulated based on this model. When applying the proposed model to the case of nonexhaustive RVLCs, these formulations can be further simplified. Since RVLCs can be decoded in backward direction, the mean and the variance of backward error detection delay length, the mean and the variance of backward error synchronization delay length are also introduced as measures to examine the error detection and error synchronization characteristics of RVLCs. In addition, we found that error synchronization probabilities of RVLCs with minimum block distance greater than 1 are 0. Chia-Wei Lin, Ja-Ling Wu |
IEEE Trans. Commun. | 2 |
| 2005 | A practical foveation-based rate-shaping mechanism for MPEG videosabstractFoveation is one of the nonuniform resolution properties of the human visual system. Recently, different foveation models are proposed and utilized for image and video coding, for the sake of bit-rate saving with no or minor perceptual quality distortion. In the first part of this paper, we propose an efficient and practical DCT-domain foveation model, which is deduced from existing experimental results. In the second part, we present a foveation-based rate-shaping mechanism for MPEG bitstreams, as an application example of the proposed foveation model. The rate shaper is based on eliminating DCT coefficients embedded in MPEG bitstreams. An efficient rate-shaping mechanism is developed to meet various bit-rate requirements. Our simulation confirmed that the proposed foveation model and the rate-shaping mechanism are practical for real-world usage. Chia-Chiang Ho, Ja-Ling Wu, Wen-Huang Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Software-Controlled Cache Architecture for Energy EfficiencyabstractPower consumption is an important design issue of current multimedia embedded systems. Data caches consume a significant portion of total processor power for multimedia applications because they are data intensive. In an integrated multimedia system, the cache architecture cannot be tuned specifically for an application. Therefore, a significant amount of cache energy is actually wasted. In this paper, we propose the software-controlled cache architecture that improves the energy efficiency of the shared cache in an integrated multimedia system on an application-specific base. Data types in an application are allocated to different cache regions. On each access, only the allocated cache regions need to be activated. We test the effectiveness of the software-controlled cache of the MPEG-2 software decoder. The results show up to 40% of cache energy reduction on an ARM-like cache architecture without sacrificing performance. Chia-Lin Yang, Hong-Wei Tseng, Chia-Chiang Ho, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | An importance measurement for video and its application to TV news items distillationabstractThe paper presents a method to distill the important frames and shots from long duration video clips. It is based on a voluntary choice by the subject (i.e. a top down approach). First, we introduce the concept of importance measurement. We use some predefined detectable events to represent the importance inferred from the semantics of the syntax. Since the event characteristics depend largely on the video sources, we choose news program videos as our analyzing focus. Then we calculate the importance measurement of frames and shots, respectively. Experiments show that a good subjective test result for news items distillation can be obtained, and the average distillation ratios of frames and shots are 13.18% and 41.23%, respectively. Jin-Hau Kuo, Chin-Wei Fang, Jen-Hao Yeh, Ja-Ling Wu |
ICIP | 4 |
| 2004 | A traceable content-adaptive fingerprinting for multimedia
Yu-Tzu Lin, Ja-Ling Wu |
ICIP | 2 |
| 2004 | A study of semantic context detection by using SVM and GMM approachesabstractSemantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. In this paper, we propose an hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. HMM are used to model audio events, and SVM and GMM are used to fuse the characteristics of various audio events related to some specific semantic concepts. The experimental results show that the approach is effective in detecting semantic context. The comparison between SVM- and GMM-based approaches is also studied Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu, Yung-Jen Hsu 0001 |
ICME | 3 |
| 2004 | Fidelity-Controlled Robustness Enhancement of Blind Watermarking Schemes Using Evolutionary Computational Techniques
Chun-Hsiang Huang, Chih-Hao Shen, Ja-Ling Wu |
IWDW | 3 |
| 2004 | Attacking visible watermarking schemesabstractVisible watermarking schemes are important intellectual property rights (IPR) protection mechanisms for digital images and videos that have to be released for certain purposes but illegal reproductions of them are prohibited. Visible watermarking techniques protect digital contents in a more active manner, which is quite different from the invisible watermarking techniques. Digital data embedded with visible watermarks will contain recognizable but unobtrusive copyright patterns, and the details of the host data should still exist. The embedded pattern of a useful visible watermarking scheme should be difficult or even impossible to be removed unless intensive and expensive human labors are involved. In this paper, we propose an attacking scheme against current visible image watermarking techniques. After manually selecting the watermarked areas, only few human interventions are required. For watermarks purely composed of thin patterns, basic image recovery techniques can completely remove the embedded patterns. For more general watermarks consisting of thick patterns, not only information in surrounding unmarked areas but also information within watermarked areas will be utilized to correctly recover the host image. Although the proposed scheme does not guarantee that the recovered images will be exactly identical to the unmarked originals, the structure of the embedded pattern will be seriously destroyed and a perceptually satisfying recovered image can be obtained. In other words, a general attacking scheme based on the contradictive requirements of current visible watermarking techniques is worked out. Thus, the robustness of current visible watermarking schemes for digital images is doubtful and needs to be improved. Chun-Hsiang Huang, Ja-Ling Wu |
IEEE Trans. Multim. | 2 |
| 2003 | A multi-modal-feature based algorithm for parsing news program videosabstractA multi-modal-feature based scene change detection algorithm (which can be viewed as a mid-stage solution between the single-modal-feature- and the semantic-based approaches) is proposed to parse the news programs, effectively. Hsuan-Wei Chen, Jin-Hau Kuo, Jen-Hao Yeh, Ja-Ling Wu |
ICASSP (3) | 4 |
| 2003 | MPEG-7 content-based analysis/retrieval system and its applications
Jin-Hau Kuo, Ja-Ling Wu |
VCIP | 2 |
| 2003 | Encoding strategies for realizing MPEG-4 universal scalable video coding
Yi-Shin Tung, Jin-Hau Kuo, Ja-Ling Wu, Wen-Huang Cheng, Ting-Jian Pan |
VCIP | 3 |
| 2002 | An efficient streaming and decoding architecture for stored FGS videoabstractFine granularity scalability (FGS) is the latest video-coding tool provided in Amendment 2 of the MPEG-4 standard. By taking advantage of bitplane coding of DCT residues, the compressed bitstream can tie truncated at any location to support the finest rate scalability in the enhancement layer. However, both frame buffer scanning several times in bitplane decoding and frame duplication simultaneously for base and enhancement layer decoding make FGS difficult in its implementation. We propose a corresponding pair of efficient streaming schedule and pipeline decoding architecture to deal with the prescribed problems. The design may be applied to the case of streaming stored FGS videos and benefit FGS-related applications. Yi-Shin Tung, Ja-Ling Wu, Po-Kang Hsiao, Kan-Li Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Fully scalable video codec
Shu-Wei Liu, Chi-Hui Huang, Ja-Ling Wu |
VCIP | 3 |
| 2001 | On constructing the Huffman-code-based reversible variable-length codesabstractIn this letter, we propose a generic and efficient algorithm that can construct both asymmetrical and symmetrical reversible variable-length codes (RVLCs). Starting from a given Huffman code, the construction is based on two developed codeword selection mechanisms, for the symmetrical case and the asymmetrical case, respectively; it is shown that the two mechanisms possess simple features and can generate efficient RVLCs easily. In addition, two new asymmetrical RVLCs are constructed and shown to be very efficient for further reducing the coding overheads in MPEG-4 when operating in the reversible decoding mode. Chien-Wu Tsai, Ja-Ling Wu |
IEEE Trans. Commun. | 2 |
| 2001 | Modified symmetrical reversible variable-length code and its theoretical boundsabstractReversible variable length codes (RVLCs) have been adopted in emerging video coding standards-H.263+ and MPEG-4-to enhance their error-resilience capabilities (which are important and essential) in error-prone environments. This study proposes an efficient algorithm to construct a symmetrical RVLC from a given Huffman code. In addition, theoretical bounds on the maximum codeword length for fixed-length Huffman codes, and on the optimal average codeword lengths for sources with exponential distribution are provided. Chien-Wu Tsai, Ja-Ling Wu |
IEEE Trans. Inf. Theory | 2 |
| 2000 | Polynomial transform based algorithms for computing two-dimensional generalized DFT, generalized DHT, and skew circular convolution
Yuh-Ming Huang, Ja-Ling Wu |
Signal Process. | 2 |
| 1999 | Hierarchical dictionary model and dictionary management policies for data compression
Chia-Lun Yu, Ja-Ling Wu |
Signal Process. | 2 |
| 1999 | Hidden digital watermarks in imagesabstractIn this paper, an image authentication technique by embedding digital "watermarks" into images is proposed. Watermarking is a technique for labeling digital pictures by hiding secret information into the images. Sophisticated watermark embedding is a potential method to discourage unauthorized copying or attest the origin of the images. In our approach, we embed the watermarks with visually recognizable patterns into the images by selectively modifying the middle-frequency parts of the image. Several variations of the proposed method are addressed. The experimental results show that the proposed technique successfully survives image processing operations, image cropping, and the Joint Photographic Experts Group (JPEG) lossy compression. Chiou-Ting Hsu, Ja-Ling Wu |
IEEE Trans. Image Process. | 2 |
| 1999 | Automatic facial feature extraction by genetic algorithmsabstractAn automatic facial feature extraction algorithm is presented. The algorithm is composed of two main stages: the face region estimation stage and the feature extraction stage. In the face region estimation stage, a second-chance region growing method is adopted to estimate the face region of a target image. In the feature extraction stage, genetic search algorithms are applied to extract the facial feature points within the face region. It is shown by simulation results that the proposed algorithm can automatically and exactly extract facial features with limited computational complexity. Chun-Hung Lin, Ja-Ling Wu |
IEEE Trans. Image Process. | 2 |
| 1998 | A lightweight genetic block-matching algorithm for video codingabstractA lightweight genetic search algorithm (LGSA) is proposed. Different evolution schemes are investigated, such that the control overheads are largely reduced. It is also shown that the proposed LGSA can be viewed as a novel expansion of the three-step search algorithm (TSS). It can be seen from the simulation results that the performance of LGSA is very similar to that of the full search algorithm (FSA), and the computational complexity is much lower than that of FSA and other previously proposed genetic motion estimation algorithms. Chun-Iiung Lin, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Range data acquisition using color structured lighting and stereo vision
Chu-Song Chen, Yi-Ping Hung, Chiann-Chu Chiang, Ja-Ling Wu |
Image Vis. Comput. | 4 |
| 1996 | Hidden signatures in imagesabstractAn image authentication technique by embedding each image with a signature so as to discourage unauthorized copying is proposed. The proposed technique could actually survive several kinds of image processing and the JPEG lossy compression. Chiou-Ting Hsu, Ja-Ling Wu |
ICIP (3) | 2 |
| 1996 | Multiresolution mosaicabstractMosaic techniques have been used to combine two or more signals into a new one with an invisible seam, and with as little distortion of each signal as possible. Multiresolution representation is an effective method for analyzing the information content of signals, and it also fits a wide spectrum of visual signal processing and visual communication applications. The wavelet transform is one kind of multiresolution representations, and has found a wide variety of application in many aspects, including signal analysis, image coding, image processing, computer vision and etc. Due to its characteristic of multiresolution signal decomposition, the wavelet transform is used for the image mosaic by choosing the width of the mosaic transition zone proportional to the frequency represented in the band. Both 1-D and 2-D signal mosaics are described, and some factors which affect the mosaics are discussed. Chiou-Ting Hsu, Ja-Ling Wu |
ICIP (3) | 2 |
| 1996 | Model-based object recognition using range images by combining morphological feature extraction and geometric hashingabstractThis paper proposes a new approach for model-based object recognition with range images by combining morphological feature extraction and geometric hashing. In low-level processing, range images are segmented into 3D-connected surface patches. In middle-level processing: each connected component is processed by using morphological operations to extract the skeletons of high-variation regions. These skeleton points can be viewed as invariant salient feature primitives. In high-level processing, geometric hashing is used to recognize objects. To reduce the number of spurious hypotheses, we propose a basis-similarity constraint. Experimental results have shown that the proposed method is effective and has great potential for model-based object recognition using range images. Chu-Song Chen, Yi-Ping Hung, Ja-Ling Wu |
ICPR | 3 |
| 1996 | MultiSync: A Synchronization Model for Multimedia SystemsabstractSynchronization among various media sources is one of the most important issues in multimedia communications and various audio/video (A/V) applications. For continuous playback (such as lip synchronization) under a time-sharing multiprocessing operating system (such as UNIX), the synchronization quality of traditional synchronization mechanisms employed on single processes may vary according to the workload of the system. When the system encounters an overload situation, the synchronization usually fails and, even worse, results in two fatal defects in human perception: the audio discontinuity (audio break) and the out-of-synchronization (synchronization anomaly). In order to overcome these problems, a novel media synchronization model employed on multiple processes (or threads) in a multiprocessing environment is proposed. The problem of asynchronism due to system overload is solved by assigning a higher priority to more important media and adopting a delay-or-drop policy to treat the lower priority ones. Some experimental results are presented to show the effectiveness of the proposed model and the implementation mechanisms under a UNIX, X-Windows environment. On the basis of the proposed model, a continuous media playback (CMP) module, which acted as the key component of some popular multimedia systems such as multimedia authoring system, multimedia E-mail system, multimedia bulletin board system (BBS), and video-on-demand (VoD) System, was implemented. Herng-Yow Chen, Ja-Ling Wu |
IEEE J. Sel. Areas Commun. | 2 |
| 1996 | A novel interpretation of the two-dimensional discrete Hartley transform
Ming-Chwen Yang, Ja-Ling Wu |
Signal Process. | 2 |
| 1996 | A three-dimensional muscle-based facial expression synthesizer for model-based image coding
Yuong-Wei Lei, Ja-Ling Wu, Ouhyoung Ming |
Signal Process. Image Commun. | 2 |
| 1996 | Class of majority decodable real-number codesabstractA majority decoding algorithm for a class of real-number codes is presented. Majority decoding has been a relatively simple and fast decoding technique for codes over finite fields. When applied to decode real-number codes, the robustness of the majority decoding to the presence of background noise, which is usually an annoying problem for existing decoding algorithms for real-number codes, is its most prominent property. The presented class of real-number codes has generator matrices similar to those of the binary Reed-Muller codes and is decoded by similar majority logic. Jiun Shiu, Ja-Ling Wu |
IEEE Trans. Commun. | 2 |
| 1995 | Real-number DFT codes for estimating a dispersive channelabstractThe utilization of real-number DFT codes for channel equalization is studied. As is shown, through real-number DFT codes, it is possible to deterministically calculate the dispersive parameters of a channel by introducing some redundancies into the transmitting data.> Jiun Shiu, Ja-Ling Wu |
IEEE Signal Process. Lett. | 2 |
| 1995 | On a Constant-Time, Low-Complexity Winner-Take-All Neural NetworkabstractA nearly cost-optimal winner-take-all (WTA) neural network derived from a constant-time sorting network is presented. The resultant WTA network has connection complexity O(n(2/sup s(2/sup s/-1))) where s is the depth of cascaded sorting networks. Application of the WTA network to other problems such as nonbinary majority is also included.> Yuen-Hsien Tseng, Ja-Ling Wu |
IEEE Trans. Computers | 2 |
| 1995 | Discrete cosine transform in error control codingabstractThe authors define a new class of real-number linear block codes using the discrete cosine transform (DCT). They also show that a subclass with a BCH-like structure can be defined and, therefore, encoding/decoding algorithms for BCH codes can be applied, A (16,10) DCT code is given as an example.> Ja-Ling Wu, Jiun Shin |
IEEE Trans. Commun. | 1 |
| 1994 | Echo cancellation with reference signal generator and reliable receiving schemes for intersymbol interferenceabstractFor adaptive echo cancellation (EC) in full-duplex two-wire data transmission, the problem of double-talk which disturbs adaptation procedure always exists since the far-end signal appears continuously. To provide reliable adaptation, an EC with a decision-directed reference signal generator (RSG) is proposed. Additionally, the intersymbol interference (ISI) is another problem. Fortunately, for an EC with far-end RSG, information on the pulse shape dispersion in the transmission channel is included in the RSG and is helpful for reducing the effect of the dispersive channel. Some methods which sufficiently make use of this information are proposed to provide optimal reception and decoding, and hence a lower bit error rate is obtained.> Hsiang-Feng Chi, Ja-Ling Wu |
ICASSP (3) | 2 |
| 1994 | Two-dimensional polynomial residue number system
Ming-Chwen Yang, Ja-Ling Wu |
Signal Process. | 2 |
| 1994 | Real-time software-based moving picture coding (SBMPC) system
Ho Chao Huang, Ja-Ling Wu |
Signal Process. Image Commun. | 2 |
| 1993 | Real-Time Software-Based Video Coder for Multimedia Communication SystemsabstractArticle Real-time software-based video coder for multimedia communication systems Share on Authors: Ho Chao Huang View Profile , Jau-Hsiung Huang View Profile , Ja-Ling Wu View Profile Authors Info & Claims MULTIMEDIA '93: Proceedings of the first ACM international conference on MultimediaSeptember 1993 Pages 65–73https://doi.org/10.1145/166266.166273Online:01 September 1993Publication History 7citation1,721DownloadsMetricsTotal Citations7Total Downloads1,721Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ho Chao Huang, Jau-Hsiung Huang, Ja-Ling Wu |
ACM Multimedia | 3 |
| 1993 | Real-time Software-based Video Codes for Multimedia Communication Systems
Ho Chao Huang, Jau-Hsiung Huang, Ja-Ling Wu |
Multim. Syst. | 3 |
| 1993 | Comments on "fixed-point error analysis of fast Hartley transform"
Guey-Shya Chen, Wei-Jou Duh, Ja-Ling Wu |
Signal Process. | 3 |
| 1993 | A novel scaling scheme for fast Hartley transform
Guey-Shya Chen, Ja-Ling Wu, Wei-Jou Duh, Lin-Shan Lee |
Signal Process. | 2 |
| 1992 | A novel modularized fast polynomial transform algorithm for two-dimensional convolutions
Ja-Ling Wu, Yuh-Ming Huang |
Signal Process. | 1 |
| 1992 | Two-variable modularized fast polynomial transform algorithm for 2-D discrete Fourier transformsabstractA novel two-variable modularized fast polynomial transform (FPT) algorithm is presented. In this method, only fast polynomial transforms and fast Fourier transforms of the same length are required. The modularity, regularity, and easy extensibility of the proposed algorithm make it of great practical value in computing multidimensional discrete Fourier transforms (DFTs).> Ja-Ling Wu, Yuh-Ming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1987 | Split vector radix 2D fast Fourier transformabstractSplit vector radix is used to develop a 2D fast Fourier transform algorithm, it is performed "in-place", and requires no matrix transpose operation; This method greatly improves the conventional vector radix 2D FFT, an overall saving of about 30% complex multiplications for a typical 1024 × 1024 array could be obtained. Soo-Chang Pei, Ja-Ling Wu |
ICASSP | 2 |
| 1985 | Exact fast digital convolution by using P-adic numbers and polynomial transformationsabstractIn this paper, an efficient method to do the digital convolution of rational numbers is proposed. The inefficient P-adic arithmetic is replaced by the integer arithmetic, in this new approach. Furthermore, since rational number are exactly representable, error-free results of the convolution of two rational sequences can be obtained by this method very efficiently. Soo-Chang Pei, Ja-Ling Wu |
ICASSP | 2 |
| 1984 | Pipeline Fast Biased Polynomial Transform Architecture for two Dimensional Convolutions
Soo-Chang Pei, Ja-Ling Wu |
ICC (1) | 2 |