Ming-Sui Lee

dblp:15/5265 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
7since 2021 · last 2023
0000-0002-6699-6694ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 End-to-end Video Matting with Trimap Propagation
abstract
The research of video matting mainly focuses on temporal coherence and has gained significant improvement via neural networks. However, matting usually relies on user-annotated trimaps to estimate alpha values, which is a labor-intensive issue. Although recent studies exploit video object segmentation methods to propagate the given trimaps, they suffer inconsistent results. Here we present a more robust and faster end-to-end video matting model equipped with trimap propagation called FTP-VM (Fast Trimap Propagation - Video Matting). The FTP-VM combines trimap propagation and video matting in one model, where the additional backbone in memory matching is replaced with the proposed lightweight trimap fusion module. The segmentation consistency loss is adopted from automotive segmentation to fit trimap segmentation with the collaboration of RNN (Recurrent Neural Network) to improve the temporal coherence. The experimental results demonstrate that the FTP-VM performs competitively both in composited and real videos only with few given trimaps. The efficiency is eight times higher than the state-of-the-art methods, which confirms its robustness and applicability in real-time scenarios. The code is available at https://github.com/csvt32745/FTP-VM.
Ming-Sui Lee
CVPR2
2023 BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression
abstract
The large amount of data collected by LiDAR sensors brings the issue of LiDAR point cloud compression (PCC). Previous works on LiDAR PCC have used range image representations and followed the predictive coding paradigm to create a basic prototype of a coding framework. However, their prediction methods give an inaccurate result due to the negligence of invalid pixels in range images and the omission of future frames in the time step. Moreover, their handcrafted design of residual coding methods could not fully exploit spatial redundancy. To remedy this, we propose a coding framework BIRD-PCC. Our prediction module is aware of the coordinates of invalid pixels in range images and takes a bidirectional scheme. Also, we introduce a deep-learned residual coding module that can further exploit spatial redundancy within a residual frame. Experiments conducted on SemanticKITTI and KITTI-360 datasets show that BIRD-PCC outperforms other methods in most bitrate conditions and generalizes well to unseen environments.
Chia-Sheng Liu, Jia-Fong Yeh, Hao Hsu, Hung-Ting Su, Ming-Sui Lee, Winston H. Hsu
ICASSP5
2023 LSR: A Light-Weight Super-Resolution Method
abstract
A light-weight super-resolution (LSR) method from a single image targeting mobile applications is proposed in this work. LSR predicts the residual image between the interpolated low-resolution (ILR) and high-resolution (HR) images using a self-supervised framework. To lower the computational complexity, LSR does not adopt the end-to-end optimization deep networks. It consists of three modules: 1) generation of a pool of rich and diversified representations in the neighborhood of a target pixel via unsupervised learning, 2) selecting a subset from the representation pool that is most relevant to the underlying super-resolution task automatically via supervised learning, 3) predicting the residual of the target pixel via regression. LSR has low computational complexity and reasonable model size so that it can be implemented on mobile/edge platforms conveniently. Besides, it offers better visual quality than classical exemplar-based methods in terms of PSNR/SSIM measures.
Wei Wang 0352, Xuejing Lei, Yueru Chen, Ming-Sui Lee, C.-C. Jay Kuo
ICIP4
2023 MT-DETR: Robust End-to-end Multimodal Detection with Confidence Fusion
abstract
Due to the trending need for autonomous driving, camera-based object detection has recently attracted lots of attention and successful development. However, there are times when unexpected and severe weather occurs in outdoor environments, making the detection tasks less effective and unexpected. In this case, additional sensors like lidar and radar are adopted to help the camera work in bad weather. However, existing multimodal detection methods do not consider the characteristics of different vehicle sensors to complement each other. Therefore, a novel end-to-end multimodal multistage object detection network called MT-DETR is proposed. Unlike the unimodal object detection networks, MT-DETR adds fusion modules and enhancement modules and adopts a hierarchical fusion mechanism. The Residual Fusion Module (RFM) and Confidence Fusion Module (CFM) are designed to fuse camera, lidar, radar, and time features. The Residual Enhancement Module (REM) reinforces each unimodal branch while a multistage loss is introduced to strengthen each branch’s effectiveness. The synthesis algorithm for generating camera-lidar data pairs in foggy conditions further boosts the performance in unseen adverse weather. Extensive experiments on various weather conditions of the STF dataset demonstrate that MT-DETR outperforms state-of-the-art methods. The generality of MT-DETR has also been confirmed by replacing the feature extractor in the experiments. The code and pre-trained models are available on https://github.com/Chushihyun/MT-DETR.
Shih-Yun Chu, Ming-Sui Lee
WACV2
2023 ABC-Norm Regularization for Fine-Grained and Long-Tailed Image Classification
abstract
Image classification for real-world applications often involves complicated data distributions such as fine-grained and long-tailed. To address the two challenging issues simultaneously, we propose a new regularization technique that yields an adversarial loss to strengthen the model learning. Specifically, for each training batch, we construct an adaptive batch prediction (ABP) matrix and establish its corresponding adaptive batch confusion norm (ABC-Norm). The ABP matrix is a composition of two parts, including an adaptive component to class-wise encode the imbalanced data distribution, and the other component to batch-wise assess the softmax predictions. The ABC-Norm leads to a norm-based regularization loss, which can be theoretically shown to be an upper bound for an objective function closely related to rank minimization. By coupling with the conventional cross-entropy loss, the ABC-Norm regularization could introduce adaptive classification confusion and thus trigger adversarial learning to improve the effectiveness of model learning. Different from most of state-of-the-art techniques in solving either fine-grained or long-tailed problems, our method is characterized with its simple and efficient design, and most distinctively, provides a unified solution. In the experiments, we compare ABC-Norm with relevant techniques and demonstrate its efficacy on several benchmark datasets, including (CUB-LT, iNaturalist2018); (CUB, CAR, AIR); and (ImageNet-LT), which respectively correspond to the real-world, fine-grained, and long-tailed scenarios.
Yen-Chi Hsu, Cheng-Yao Hong, Ming-Sui Lee, Davi Geiger, Tyng-Luh Liu
IEEE Trans. Image Process.3
2022 ScannerNet: A Deep Network for Scanner-Quality Document Images under Complex Illumination
Chih-Jou Hsu, Yu-Ting Wu 0001, Ming-Sui Lee, Yung-Yu Chuang
BMVC3
2021 Unsupervised video object segmentation with distractor-aware online adaptation
Ye Wang 0013, Jongmoo Choi, Yueru Chen, Siyang Li 0002, Qin Huang 0006, Kaitai Zhang, Ming-Sui Lee, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.7
2020 Query-Driven Multi-Instance Learning
abstract
We introduce a query-driven approach (qMIL) to multi-instance learning where the queries aim to uncover the class labels embodied in a given bag of instances. Specifically, it solves a multi-instance multi-label learning (MIML) problem with a more challenging setting than the conventional one. Each MIML bag in our formulation is annotated only with a binary label indicating whether the bag contains the instance of a certain class and the query is specified by the word2vec of a class label/name. To learn a deep-net model for qMIL, we construct a network component that achieves a generalized compatibility measure for query-visual co-embedding and yields proper instance attentions to the given query. The bag representation is then formed as the attention-weighted sum of the instances' weights, and passed to the classification layer at the end of the network. In addition, the qMIL formulation is flexible for extending the network to classify unseen class labels, leading to a new technique to solve the zero-shot MIML task through an iterative querying process. Experimental results on action classification over video clips and three MIML datasets from MNIST, CIFAR10 and Scene are provided to demonstrate the effectiveness of our method.
Yen-Chi Hsu, Cheng-Yao Hong, Ming-Sui Lee, Tyng-Luh Liu
AAAI3
2020 Activity Recognition Using First-Person-View Cameras Based on Sparse Optical Flows
abstract
First-person-view (FPV) cameras are finding wide use in daily life to record activities and sports. In this paper, we propose a succinct and robust 3D convolutional neural network (CNN) architecture accompanied with an ensemble-learning network for activity recognition with FPV videos. The proposed 3D CNN is trained on low-resolution (32 × 32) sparse optical flows using FPV video datasets consisting of daily activities. According to the experimental results, our network achieves an average accuracy of 90%.
Peng Yua Kao, Yan-Jing Lei, Chu-Song Chen, Ming-Sui Lee, Yi-Ping Hung
ICPR5
2020 VR Sickness Assessment with Perception Prior and Hybrid Temporal Features
abstract
Virtual reality (VR) sickness is one of the obstacles hindering the growth of the VR market. Different VR contents may cause various degree of sickness. If the degree of the sickness can be estimated objectively, it adds a great value and help in designing the VR contents. To address this problem, a novel content-based VR sickness assessment method which considers both the perception prior and hybrid temporal features is proposed. Based on the perception prior which assumes the user's field of view becomes narrower while watching videos, a Gaussian weighted optical flow is calculated with a specified aspect ratio. In order to capture the dynamic characteristics, hybrid temporal features including horizontal motion, vertical motion and the proposed motion anisotropy are adopted. In addition, a new dataset is compiled with one hundred VR sickness test samples and each of which comes along with the Discomfort Scores (DS) answered by the user and a Simulator Sickness Questionnaire (SSQ) collected at the end of test. A random forest regressor is then trained on this dataset by feeding the hybrid temporal features of both the present and the previous minute. Extensive experiments are conducted on the VRSA dataset and the results demonstrate that the proposed method is comparable to the state-of-the-art method in terms of effectiveness and efficiency.
Po-Chen Kuo, Li-Chung Chuang, Dong-Yi Lin, Ming-Sui Lee
ICPR4
2020 Video object tracking and segmentation with box annotation
Ye Wang 0013, Jongmoo Choi, Kaitai Zhang, Qin Huang 0006, Yueru Chen, Ming-Sui Lee, C.-C. Jay Kuo
Signal Process. Image Commun.6
2019 A Learning-Based Prediction Model for Baby Accidents
abstract
According to the statistics in the United Kingdom, more than two million babies and toddlers experienced accidents every year. Despite the places where accidents happened, most of the accidents could've been predicted and prevented. In order to avoid causing injuries by accident, a temporal-pyramid long short-term memory (TP-LSTM) network along with the temporal attention mechanism is proposed to predict whether an accident will happen in the future or not. The proposed network is capable of capturing important information of the video at different temporal resolution and selecting crucial frames that contribute to the accident most. Moreover, the proposed early exponential loss (EEL) function is incorporated to achieve better prediction. The baby video dataset (BVD) containing 670 videos is collected from several video-sharing websites. 320 of which are with accidents and the others are without accidents. The experimental results show that the proposed network attains average precision of 61.13% and the accidents are foreseen 4.196 seconds before the occurrence with 80% recall.
Peng-Jie Wang 0008, Shao-Fu Lien, Ming-Sui Lee
ICIP3
2018 Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation
Ye Wang 0013, Jongmoo Choi, Yueru Chen, Qin Huang 0006, Siyang Li 0002, Ming-Sui Lee, C.-C. Jay Kuo
ACCV (4)6
2018 Online object tracking via motion-guided convolutional neural network (MGNet)
Weihao Gan, Ming-Sui Lee, Chihao Wu 0001, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2018 Online CNN-based multiple object tracking with enhanced model updates and identity association
Weihao Gan, Xuejing Lei, Ming-Sui Lee, C.-C. Jay Kuo
Signal Process. Image Commun.4
2016 Object tracking with temporal prediction and spatial refinement (TPSR)
Weihao Gan, Ming-Sui Lee, Chihao Wu 0001, C.-C. Jay Kuo
Signal Process. Image Commun.2
2014 Opportunities for Persuasive Technology to Motivate Heavy Computer Users for Stretching Exercise
Yong-Xiang Chen, Siek-Siang Chiang, Shu-Yun Chih, Wen-Ching Liao, Shih-Yao Lin 0001, Shang-Hua Yang, Shun-Wen Cheng, Shih-Sung Lin, Yu-Shan Lin, Ming-Sui Lee, Jau-Yih Tsauo, Cheng-Min Jen, Chia-Shiang Shih, King-Jen Chang, Yi-Ping Hung
PERSUASIVE10
2013 IrotateGrasp: automatic screen rotation based on grasp of mobile devices
abstract
Automatic screen rotation improves viewing experience and usability of mobile devices, but current gravity-based approaches do not support postures such as lying on one side, and manual rotation switches require explicit user input. iRotateGrasp automatically rotates screens of mobile devices to match users' viewing orientations based on how users are grasping the devices. Our insight is that users' grasps are consistent for each orientation, but significantly differ between different orientations. Our prototype used a total of 44 capacitive sensors along the four sides and the back of an iPod Touch, and uses support vector machine (SVM) to recognize grasps at 25Hz. We collected 6-users' usage under 108 different combinations of posture, orienta-tion, touchscreen operation, and left/right/both hands. Our offline analysis showed that our grasp-based approach is promising, with 80.9% accuracy when training and testing on different users, and up to 96.7% if users are willing to train the system. Our user study (N=16) showed that iRo-tateGrasp had an accuracy of 78.8% and was 31.3% more accurate than gravity-based rotation.
Lung-Pan Cheng, Meng-Han Lee, Che-Yang Wu, Fang-I Hsiao, Yen-Ting Liu, Hsiang-Sheng Liang, Yi-Ching Chiu, Ming-Sui Lee, Mike Y. Chen
CHI8
2013 Spatially-varying super-resolution for HDTV
abstract
We propose a system to up-sample visual signals for high definition televisions (HDTVs). Although the original visual signals are degraded and limited, we still try to solve these problems by using super-resolution. First, we interpolate the visual signals with a edge taper to the desired size. Then, we combine L1 data term, L2 data term, Tikhonov-like regularizer and total variation regularizer with a saliency weighting for our L12TTV deblurring method. Adopting our L12TTV deblurring, we can remove the spatially-varying blurry conditions of the interpolated signals. Experimental results show that our system outperforms than comparisons in both image and video cases.
Chih-Tsung Shen, Hung-Hsun Liu, Ming-Sui Lee, Yi-Ping Hung, Soo-Chang Pei
ISCAS3
2013 Video Aesthetic Quality Assessment by Temporal Integration of Photo- and Motion-Based Features
abstract
This paper presents a new method for accessing the aesthetic quality of videos. It consists of two processes: aesthetic features construction and temporal integration. First, our method combines both photo-based and motion-based visual clues to extract the aesthetic features for each frame in a video. We introduce new motion-based features built from optical flow and salient region extraction, and show their effectiveness to enhance the estimation of aesthetic values. Then, a temporal-order-aware framework that integrates the frame-based features is presented to further improve the evaluation accuracy by taking the time-varying properties into consideration. The experimental results demonstrate that our approach can accomplish remarkable improvement for aesthetic quality assessment of videos.
Hsin-Ho Yeh, Ming-Sui Lee, Chu-Song Chen
IEEE Trans. Multim.3
2012 Human action recognition using Action Trait Code
Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Ming-Sui Lee, Yi-Ping Hung
ICPR4
2011 i - m - Breath: The Effect of Multimedia Biofeedback on Learning Abdominal Breath
Meng-Chieh Yu, Jin-Shing Chen, King-Jen Chang, Su-Chu Hsu, Ming-Sui Lee, Yi-Ping Hung
MMM (1)5
2011 A low-complexity upsampling technique for H.264
abstract
A hybrid up-sampling algorithm based on the predicted modes of H.264/AVC is proposed in this paper. Other than video codecs like MPEG group, H.264/AVC utilizes variable block size for motion estimation and motion compensation, which results in better precision and compression efficiency. According to the mode decision built in H.264/AVC, the macroblocks of each frame are divided into intra mode, skip mode and others. For intra-mode macroblocks which contain more details, they are up-sampled by MAP (maximum a posteriori) since this method has best performance among existing super resolution algorithms. For macroblocks coded as skip mode, they are assumed to be highly correlated to macroblocks in the reference frame. Thus those blocks are duplicated from those referenced blocks. For the rest of the macroblocks, they not only have correspondence with blocks in other frames but also contain relatively complicated content so that they are further analyzed into variable block sizes, say 16x 16, 8x16,16x8, 8x8, 8x4, 4x8 and 4χ4. By adopting different up- sampling methods adaptively with variable block sizes, the proposed method saves computational efforts of smoother blocks for complicated blocks so that the overall complexity can be successfully reduced. Comparing to traditional frame-based up-sampling methods, the experimental results demonstrated that the proposed algorithm provides a more efficient way to up- sample videos and is capable of preserving satisfactory visual quality.
Wei-Chi Chen, Ming-Sui Lee
VCIP2
2010 Touching the void: direct-touch interaction for intangible displays
abstract
In this paper, we explore the challenges in applying and investigate methodologies to improve direct-touch interaction on intangible displays. Direct-touch interaction simplifies object manipulation, because it combines the input and display into a single integrated interface. While traditional tangible display-based direct-touch technology is commonplace, similar direct-touch interaction within an intangible display paradigm presents many challenges. Given the lack of tactile feedback, direct-touch interaction on an intangible display may show poor performance even on the simplest of target acquisition tasks. In order to study this problem, we have created a prototype of an intangible display. In the initial study, we collected user discrepancy data corresponding to the interpretation of 3D location of targets shown on our intangible display. The result showed that participants performed poorly in determining the z-coordinate of the targets and were imprecise in their execution of screen touches within the system. Thirty percent of positioning operations showed errors larger than 30mm from the actual surface. This finding triggered our interest to design a second study, in which we quantified task time in the presence of visual and audio feedback. The pseudo-shadow visual feedback was shown to be helpful both in improving user performance and satisfaction.
Li-Wei Chan 0001, HuiShan Kao, Mike Y. Chen, Ming-Sui Lee, Yung-Jen Hsu 0001, Yi-Ping Hung
CHI4
2010 Image recovery of geometric distortion with multi-bit data embedding
abstract
Image transmission is sometimes accompanied with geometric distortions. A novel image recovery scheme with multi-bit binary message is proposed in this paper. In the proposed scheme, several predefined templates are introduced to an image in the discrete wavelet transform domain. A blind template detection algorithm is performed on the geometrically distorted image to extract locations of the templates and the hidden message, which are modeled probabilistically in a Bayesian network. Once the locations of the templates are successfully detected, they serve as the registration references in the recovering process. As a result, the image attacked by geometric distortions can be recovered according to the estimated displacements. The goal of this project is to develop a scheme to correct various geometric distortions with relatively lower bit error rate.
Ming-Sui Lee, Yu-Hsiang Bosco Chiu
ICME1
2010 i-m-Space: interactive multimedia-enhanced space for rehabilitation of breast cancer patients
abstract
This paper presents i-m-Space, an interactive multimedia rehabilitation space that helps the post-surgery recovery of breast cancer patients. Our goal is to improve patients' physical therapy and psychological relaxation experience through careful applications of multimedia technology. i-m-Space consists of three types of breathing-based relaxation and three types for interactive exercise-based rehabilitation.
Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Jin-Yao Lin, Szu-Wei Wu, Yi-Yu Chung, I-Ling Hu, Wei-Ting Peng, Shih-Yao Lin 0001, Chia-Han Chang, Pei-Hsuan Chou, King-Jen Chang, Mei-Lan Chang, Sue-Huei Chen, Jin-Shing Chen, Ming-Sui Lee, Mike Y. Chen, Yi-Ping Hung
ACM Multimedia17
2010 Transformational Breathing between Present and Past: Virtual Exhibition System of the Mao-Kung Ting
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Szu-Wei Wu, Yi-Yu Chung, Liang-Chun Lin, Ming-Sui Lee, Chu-Song Chen, Jiaping Wang, Quo-Ping Lin, I-Ling Liu
MMM11
2007 Gesture-based interaction for a magic crystal ball
abstract
Crystal balls are generally considered as media to perform divination or fortune-telling. These imaginations are mainly from some fantasy films and fiction, in which an augur can see into the past, the present, or the future through a crystal ball. With the distinct impressions, crystal ball has revealed itself as a perfect interface for the users to access and to manipulate visual media in an intuitive, imaginative and playful manner. We developed an interactive visual display system named Magic Crystal Ball (MaC Ball). MaC Ball is a spherical display system, which allows the users to see a virtual object/scene appearing inside a transparent sphere, and to manipulate the displayed content with barehanded interactions. Interacting with MaC Ball makes the users feeling acting with magic power. With MaC Ball, user can manipulate the display with touch and hover interactions. For instance, the user waves hands above the ball, causing clouds blowing from bottom of the ball, or slides fingers on the ball to rotate the displayed object. In addition, the user can press single finger to select an object or to issue a button. MaC Ball takes advantages on the impressions of crystal balls, allowing the users acting with visual media following their imaginations. For applications, MaC Ball has high potential to be used for advertising and demonstration in museums, product launches, and other venues.
Li-Wei Chan 0001, Yi-Fan Chuang, Meng-Chieh Yu, Yi-Liu Chao, Ming-Sui Lee, Yi-Ping Hung, Yung-Jen Hsu 0001
VRST5
2006 A Quad-Tree Decomposition Approach to Cartoon Image Compression
abstract
A quad-tree decomposition approach is proposed for cartoon image compression in this work. The proposed algorithm achieves excellent coding performance by using a unique quad-tree decomposition and shape coding method along with a GIF like color indexing technique to efficiently encode large areas of the same color, which appear in a cartoon-type image commonly. To reduce complexity, the input image is partitioned into small blocks and the quad-tree decomposition is independently applied to each block instead of the entire image. The LZW entropy coding method can be performed as a postprocessing step to further reduce the coded file size. It is demonstrated by experimental results that the proposed method outperforms several well-known lossless image compression techniques for cartoon images that contain 256 colors or less
Yi-Chen Tsai, Ming-Sui Lee, Mei-Yin Shen, C.-C. Jay Kuo
MMSP2
2005 A DCT-Domain Video Alignment Technique for MPEG Sequences
abstract
An image/video registration technique for multiple compressed video inputs such as MPEG sequences is investigated. The proposed technique is based on the matching of discrete cosine transform (DCT) coefficients and motion vectors. First, the I frame of each input sequence is separated into the background and moving objects. For the background, coarse edge features are extracted by applying edge detectors of different characteristics to the luminance DC coefficients. Each detector generates a difference map for a single background. A threshold is determined for each difference map to produce a binary map. Then, alignment parameters are determined using the binary maps of input images generated by the same detector. For the moving object, alignment parameters can be finetuned by the motion information of all frames in the same group of pictures (GOP). Finally, the actual displacement in the pixel domain is estimated by the weighted average of alignment parameters from all background detectors and refinement parameters from motion information. It is shown by experimental results that the proposed method reduces the computational cost of image/video registration significantly in comparison with the traditional pixel domain registration techniques while achieving certain quality of composition
Ming-Sui Lee, Mei-Yin Shen, C.-C. Jay Kuo
MMSP1
2004 Color matching techniques for video mosaic applications
abstract
Color matching techniques are proposed to merge two or more video inputs of smaller sizes into one single larger output video with a wider field of view for video mosaic applications. The main challenge is to remove the seam lines between image boundaries due to the different color tone of the inputs. In this paper, color differences between input images are first compensated using the polynomial-based contrast stretching technique. Then, a linear filtering technique is adopted to remove the seam lines. The algorithms are developed in both the pixel domain and the DCT (discrete cosine transform) domain. The second approach is attractive for its lower computational complexity. Experimental results demonstrate that the color-matching problem can be satisfactorily solved in the compressed domain even when the DCT blocks of original input images are not aligned and also for the images taken with camera movement as long as image registration is done in advance. The proposed approach is applicable for MPEG2 video.
Ming-Sui Lee, Mei-Yin Shen, C.-C. Jay Kuo
ICME1
2004 Pixel- and compressed-domain color matching techniques for video mosaic applications
abstract
Several color matching algorithms are proposed to merge two or more video inputs of smaller sizes into one single video output of a larger size with a wider field of view for the video mosaic application. The main challenge is to remove apparent seam lines between image boundaries. All developed algorithms share the same basic idea with different implementation details. That is, color differences between input images are first compensated using either histogram equalization or polynomial-based contrast stretching techniques. Then, a linear filtering technique is adopted to remove seam lines between image boundaries. The algorithms are developed in both the pixel and the DCT domains. The compressed-domain processing is attractive since it reduces the computational complexity. It is shown by experimental results that the color matching problem can be solved satisfactorily even in the compressed domain.
Ming-Sui Lee, Mei-Yin Shen, C.-C. Jay Kuo
VCIP1