Ning Zhang 0023

dblp:181/2597-23 · DBLP profile ↗
← Back
25ranked-venue papers
13as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction
abstract
In Multiple Object Tracking, objects often exhibit nonlinear motion of acceleration and deceleration, with irregular direction changes. Tacking-by-detection (TBD) trackers with Kalman Filter motion prediction work well in pedestrian-dominant scenarios but fall short in complex situations when multiple objects perform non-linear and diverse motion simultaneously. To tackle the complex non-linear motion, we propose a real-time diffusion-based MOT approach named DiffMOT. Specifically, for the motion predictor component, we propose a novel Decoupled Diffusion based Motion Predictor (D2MP). It models the entire distribution of various motion presented by the data as a whole. It also predicts an individual object's motion conditioning on an individual's historical motion information. Furthermore, it optimizes the diffusion process with much fewer sampling steps. As a MOT tracker, the DiffMOT is real-time at 22.7FPS, and also outperforms the state-of-the-art on DanceTrack[30] and SportsMOT[6] datasets with 62.3% and 76.2% in HOTA metrics, respectively. To the best of our knowledge, DiffMOT is the first to introduce a diffusion probabilistic model into the MOT to tackle non-linear motion prediction.
Weiyi Lv, Yuhang Huang 0006, Ning Zhang 0023, Ruei-Sung Lin, Dan Zeng 0001
CVPR3
2024 Deep Video Inverse Tone Mapping Based on Temporal Clues
abstract
Inverse tone mapping (ITM) aims to reconstruct high dynamic range (HDR) radiance from low dynamic range (LDR) content. Although many deep image ITM methods can generate impressive results, the field of video ITM is still to be explored. Processing video sequences by image ITM methods may cause temporal inconsistency. Besides, they aren't able to exploit the potentially useful information in the temporal domain. In this paper, we analyze the process of video filming, and then propose a Global Sample and Local Propagate strategy to better find and utilize temporal clues. To better realize the proposed strategy, we design a two-stage pipeline which includes modules named Incremental Clue Aggregation Module and Feature and Clue Propagation Module. They can align andfuseframes effectively under the condition of brightness changes and propagate features and temporal clues to all frames efficiently. Our temporal clues based video ITM method can recover realistic and temporal consistent results with high fidelity in over-exposed regions. Qualitative and quantitative experiments on public datasets show that the proposed method has significant advantages over existing methods. The code is available at https://github.com/ye3why/VITM-TC/.
Yuyao Ye, Ning Zhang 0023, Yang Zhao 0002, Hongbin Cao, Ronggang Wang
CVPR2
2024 One-Shot Multiple Object Tracking With Robust ID Preservation
abstract
Maintaining identity consistency and avoiding ID-switch during tracking is one of the primary focuses of multiple object tracking (MOT). One-shot MOT methods which jointly learn the detection and tracking models in one single network (hence namely, one-shot) have achieved promising results in tracking accuracy and speed. However, their capabilities of maintaining ID consistency are somehow weakened. The reason for this weakened ID consistency is two-fold: (1) the ID features learned by one-shot methods are not discriminative enough due to their heatmap-based single-location representation. (2) severe occlusion in the MOT scene leads to feature ambiguity and high ID-switch. In this paper, we propose a one-shot MOT system with strong ID consistency called PID-MOT (Preserved ID MOT). Specifically, we devise a visibility branch to predict the object occlusion level, and a predicted visibility map will be used in both Feature Refinement Model (FRM) and a visibility-guided two-stage association strategy (VGTAS). FRM is designed to strengthen the location-based features and enrich the identity information. VGTAS is proposed for tackling objects with high and low visibility separately. In addition, we initialize the parameters of our model by training on the recently emerged abundant synthetic MOTSynth dataset from scratch rather than the commonly used COCO dataset for full training. Finally, we carry out our method on the commonly used MOT datasets and the experimental results demonstrate that the proposed PID-MOT achieves especially good performances in ID F1 score (IDF1) and ID-Switch (IDS) compared with other state-of-the-art one-shot trackers, with comparable overall HOTA/MOTA performance. The code is available at https://github.com/Kroery/PIDMOT.
Weiyi Lv, Ning Zhang 0023, Junjie Zhang 0002, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Revisiting the Stack-Based Inverse Tone Mapping
abstract
Current stack-based inverse tone mapping (ITM) methods can recover high dynamic range (HDR) radiance by predicting a set of multi-exposure images from a single low dynamic range image. However, there are still some limitations. On the one hand, these methods estimate a fixed number of images (e.g., three exposure-up and three exposure-down), which may introduce unnecessary computational cost or reconstruct incorrect results. On the other hand, they neglect the connections between the up-exposure and down-exposure models and thus fail to fully excavate effective features. In this paper, we revisit the stack-based ITM approaches and propose a novel method to reconstruct HDR radiance from a single image, which only needs to estimate two exposure images. At first, we design the exposure adaptive block that can adaptively adjust the exposure based on the luminance distribution of the input image. Secondly, we devise the cross-model attention block to connect the exposure adjustment models. Thirdly, we propose an end-to-end ITM pipeline by incorporating the multi-exposure fusion model. Furthermore, we propose and open a multi-exposure dataset that indicates the optimal exposure-up/down levels. Experimental results show that the proposed method outperforms some state-of-the-art methods.
Ning Zhang 0023, Yuyao Ye, Yang Zhao 0002, Ronggang Wang
CVPR1
2023 Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach
abstract
Traditional image codecs prioritize signal fidelity and human perception, often neglecting machine vision tasks. Deep learning approaches have shown promising coding performance by leveraging rich semantic embeddings that can be optimized for both human and machine vision. However, these compact embeddings struggle to represent low-level details like contours and textures, leading to imperfect reconstructions. Additionally, existing learning-based coding tools lack scalability. To address these challenges, this paper presents a content-adaptive diffusion model for scalable image compression. The method encodes accurate texture through a diffusion process, enhancing human perception while preserving important features for machine vision tasks. It employs a Markov palette diffusion model with commonly-used feature extractors and image generators, enabling efficient data compression. By utilizing collaborative texture-semantic feature extraction and pseudo-label generation, the approach accurately learns texture information. A content-adaptive Markov palette diffusion model is then applied to capture both low-level texture and high-level semantic knowledge in a scalable manner. This framework enables elegant compression ratio control by flexibly selecting intermediate diffusion states, eliminating the need for deep learning model re-training at different operating points. Extensive experiments demonstrate the effectiveness of the proposed framework in image reconstruction and downstream machine vision tasks such as object detection, segmentation, and facial landmark detection. It achieves superior perceptual quality scores compared to state-of-the-art methods.
Sha Guo, Zhuo Chen 0006, Yang Zhao 0002, Ning Zhang 0023, Ling-Yu Duan
ACM Multimedia4
2022 A Real-Time Semi-Supervised Deep Tone Mapping Network
abstract
Tone mapping operators (TMOs) can compress the range of high dynamic range (HDR) images so that they can be displayed normally on the low dynamic range (LDR) devices. Recent TMOs based on deep neural networks can produce impressive results, but there are still some shortcomings. On the one hand, their supervised learning procedure requires a high-quality paired dataset which is hard to be accessed. On the other hand, they are too slow and heavy to meet the needs of practical applications. This paper proposes a real-time deep semi-supervised learning TMO to solve the above problems. The proposed method learns in a semi-supervised manner by combining the adversarial loss, cycle consistency loss, and the pixel-wise loss. The first two can simulate the image distributions in the real world from the unpaired LDR data and the latter can learn the guidance of paired LDR labels. In this way, the proposed method only requires HDR sources, unpaired high-quality LDR images, and a few well tone-mapped HDR-LDR pairs as training data. Furthermore, the proposed method divides tone mapping into luminance mapping and saturation adjustment and then processes them simultaneously. By this strategy, we can reconstruct each component more precisely. Based on the aforementioned improvements, we propose a lightweight tone mapping network that is efficient in tone mapping task (up to 5000x parameters-saving and 27x time-saving compared to the learning-based TMOs). Both quantitative and qualitative results demonstrate that the proposed method performs favorable against state-of-the-art TMOs.
Ning Zhang 0023, Yang Zhao 0002, Chao Wang 0037, Ronggang Wang
IEEE Trans. Multim.1
2021 Smart Director: An Event-Driven Directing System for Live Broadcasting
abstract
Live video broadcasting normally requires a multitude of skills and expertise with domain knowledge to enable multi-camera productions. As the number of cameras keeps increasing, directing a live sports broadcast has now become more complicated and challenging than ever before. The broadcast directors need to be much more concentrated, responsive, and knowledgeable, during the production. To relieve the directors from their intensive efforts, we develop an innovative automated sports broadcast directing system, called Smart Director, which aims at mimicking the typical human-in-the-loop broadcasting process to automatically create near-professional broadcasting programs in real-time by using a set of advanced multi-view video analysis algorithms. Inspired by the so-called “three-event” construction of sports broadcast [ 14 ], we build our system with an event-driven pipeline consisting of three consecutive novel components: (1) the Multi-View Event Localization to detect events by modeling multi-view correlations, (2) the Multi-View Highlight Detection to rank camera views by the visual importance for view selection, and (3) the Auto-Broadcasting Scheduler to control the production of broadcasting videos. To our best knowledge, our system is the first end-to-end automated directing system for multi-camera sports broadcasting, completely driven by the semantic understanding of sports events. It is also the first system to solve the novel problem of multi-view joint event detection by cross-view relation modeling. We conduct both objective and subjective evaluations on a real-world multi-camera soccer dataset, which demonstrate the quality of our auto-generated videos is comparable to that of the human-directed videos. Thanks to its faster response, our system is able to capture more fast-passing and short-duration events which are usually missed by human directors.
Yingwei Pan, Qian Bao, Ning Zhang 0023, Ting Yao 0003, Jingen Liu, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Robust Visual Object Tracking with Two-Stream Residual Convolutional Networks
abstract
The current deep learning based visual tracking approaches have been very successful by learning the target classification and/or estimation model from a large amount of supervised training data in offline mode. However, most of them can still fail in tracking objects due to some more challenging issues such as dense distractor objects, confusing background, motion blurs, and so on. Inspired by the human “visual tracking” capability which leverages motion cues to distinguish the target from the background, we propose a Two-Stream Residual Convolutional Network (TS-RCN) for visual tracking, which successfully exploits both appearance and motion features for model update. Our TS-RCN can be integrated with existing deep learning based visual trackers. To further improve the tracking performance, we adopt a “wider” residual network ResNeXt as its feature extraction backbone. To the best of our knowledge, TS-RCN is the first end-to-end trainable two-stream visual tracking system, which makes full use of both appearance and motion features of the target. We have extensively evaluated the TS-RCN on most widely used benchmark datasets including VOT2018, VOT2019, and GOT-10K. The experiment results have successfully demonstrated that our two-stream model can greatly outperform the appearance-based tracker, and achieves state-of-the-art performance. The tracking system can run at up to 38.1 FPS.
Ning Zhang 0023, Jingen Liu, Dan Zeng 0001, Tao Mei 0001
ICPR1
2020 AI-SAS: Automated In-match Soccer Analysis System
abstract
Real-time in-match soccer statistics provide continuous tracking of soccer ball and player positions and speeds, enabling advanced analytics. Currently, only elite soccer leagues have the luxury of tracking in-match soccer statistics operated with a large number of trained personnel. In this work, we present an Automated In-match Soccer Analysis System (AI-SAS), using a domain-knowledge-based multi-view global tracking. This system tracks player team, position, and speed automatically, providing real-time in-match team- and individual-level statistics and analyses. In comparison with the latest soccer analysis systems, AI-SAS is more scalable in streaming multiple video sources for real-time process and more flexible in hosting plug-and-play deep-learning-based tracking-by-detection algorithms. The global multi-view tracking also overcomes the single-view limitation and improves the tracking accuracy.
Ning Zhang 0023, Wei Zhang 0031, Dan Zeng 0001, Jingen Liu, Tao Mei 0001
ACM Multimedia1
2019 Deep tone mapping network in HSV color space
abstract
Tone mapping operators can convert high dynamic range (HDR) images to low dynamic range (LDR) images so that we can enjoy the informative contents of HDR images with LDR devices. However, current state-of-the-art tone mapping algorithms mainly focus on the luminance mapping while neglecting the color component. Meanwhile, they often suffer from halo artifacts and over-enhancement. In this paper, we propose a tone mapping network (TMNet) in Hue-Saturation-Value (HSV) color space to obtain better luminance and color mapping. We adopt the improved Wasserstein generative adversarial network (WGAN-GP) as the basic architecture and further introduce several improvements. A meticulously designed loss function is adopted to push tone mapped image to the natural image manifold. What’s more, we create a tone mapped image dataset in which the label images are manually adjusted by photographers. Compared with some state-of-the-art tone mapping methods, the proposed method can achieve better performance in both subjective and objective evaluations.
Ning Zhang 0023, Chao Wang 0037, Yang Zhao 0002, Ronggang Wang
VCIP1
2017 vConnect: perceive and interact with real world from CAVE
Xiaoming Nan, Ning Zhang 0023, Fei Guo 0002, Edward Rosales, Ling Guan
Multim. Tools Appl.4
2015 TapTell: Interactive visual search for mobile task recommendation
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
J. Vis. Commun. Image Represent.1
2012 Interactive mobile visual search for social activities completion using query image contextual model
abstract
Mobile devices are ubiquitous. People use their phones as a personal concierge not only discovering information but also searching for particular interest on-the-go and making decisions. This brings a new horizon for multimedia retrieval on mobile. While existing efforts have predominantly focused on understanding textual or a voice query, this paper presents a new perspective which understands visual queries captured by the built-in camera such that mobile-based social activities can be recommended for users to complete. In this work, a query image-based contextual model is proposed for visual search. A mobile user can take a photo and naturally indicate an object-of-interest within the photo via circle based gesture called “O” gesture. Both selected object-of-interest region as well as surrounding visual context in photo are used in achieving a search-based recognition by retrieving similar images based on a large-scale of visual vocabulary tree. Consequently, social activities such as visiting contextually relevant entities (i.e., local businesses) are recommended to the users based on their visual queries and GPS location. Along with the proposed method, an exemplary real application has been developed on Windows Phone 7 devices and evaluated with a wide variety of scenarios on million-scale image database. To test the performance of proposed mobile visual search model, extensive experimentation has been conducted and compared with state-of-the-art algorithms in content-based image retrieval (CBIR) domain.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
MMSP1
2012 A Generic Approach for Systematic Analysis of Sports Videos
abstract
Various innovative and original works have been applied and proposed in the field of sports video analysis. However, individual works have focused on sophisticated methodologies with particular sport types and there has been a lack of scalable and holistic frameworks in this field. This article proposes a solution and presents a systematic and generic approach which is experimented on a relatively large-scale sports consortia. The system aims at the event detection scenario of an input video with an orderly sequential process. Initially, domain knowledge-independent local descriptors are extracted homogeneously from the input video sequence. Then the video representation is created by adopting a bag-of-visual-words (BoW) model. The video’s genre is first identified by applying the k-nearest neighbor (k-NN) classifiers on the initially obtained video representation, and various dissimilarity measures are assessed and evaluated analytically. Subsequently, an unsupervised probabilistic latent semantic analysis (PLSA)-based approach is employed at the same histogram-based video representation, characterizing each frame of video sequence into one of four view groups, namely closed-up-view, mid-view, long-view, and outer-field-view. Finally, a hidden conditional random field (HCRF) structured prediction model is utilized for interesting event detection. From experimental results, k-NN classifier using KL-divergence measurement demonstrates the best accuracy at 82.16% for genre categorization. Supervised SVM and unsupervised PLSA have average classification accuracies at 82.86% and 68.13%, respectively. The HCRF model achieves 92.31% accuracy using the unsupervised PLSA based label input, which is comparable with the supervised SVM based input at an accuracy of 93.08%. In general, such a systematic approach can be widely applied in processing massive videos generically.
Ning Zhang 0023, Ling-Yu Duan, Lingfang Li, Qingming Huang, Wen Gao 0001, Ling Guan
ACM Trans. Intell. Syst. Technol.1
2011 TapTell: understanding visual intents on-the-go
abstract
This demonstration presents a mobile-based visual recognition and recommendation application on Windows Phone 7 called TapTell. This is different from other mobile-based visual search mechanisms which merely focus on the search process. TapTell firstly discovers and understands users' visual intents via a circle based natural user interaction called "O" gestures. Following, a Tap action is operated to choose the "O" gestured regions. The context-aware visual search mechanism is utilized for recognizing the intents and associating them with indexed metadata. Finally, the "Tell" action recommends relevant entities utilizing contextual information. The TapTell system has been evaluated at different scenarios on million scale images.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
ACM Multimedia1
2011 Tap-to-search: Interactive and contextual visual search on mobile devices
abstract
Mobile visual search has been an emerging topic for both researching and industrial communities. Among various methods, visual search has its merit in providing an alternative solution, where text and voice searches are not applicable. This paper proposes an interactive “tap-to-search” approach utilizing both individual's intention in selecting interested regions via “tap” actions on the mobile touch screen, as well as a visual recognition by search mechanism in a large-scale image database. Automatic image segmentation technique is applied in order to provide region candidates. Visual vocabulary tree based search is adopted by incorporating rich contextual information which are collected from mobile sensors. The proposed approach has been conducted on an image dataset with the scale of two million. We demonstrated that using GPS contextual information, such an approach can further achieve satisfactory results with the standard information retrieval evaluation.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
MMSP1
2010 An efficient framework on large-scale video genre classification
abstract
Efficient data mining and indexing is important for multimedia analysis and retrieval. In the field of large-scale video analysis, effective genre categorization plays an important role and serves one of the fundamental steps for data mining. Existing works utilize domain-knowledge dependent feature extraction, which is limited from genre diversification as well as data volume scalability. In this paper, we propose a systematic framework for automatically classifying video genres using domain-knowledge independent descriptors in feature extraction, and a bag-of-visualwords (BoW) based model in compact video representation. Scale invariant feature transform (SIFT) local descriptor accelerated by GPU hardware is adopted for feature extraction. BoW model with an innovative codebook generation using bottom-up two-layer K-means clustering is proposed to abstract the video characteristics. Besides the histogram-based distribution in summarizing video data, a modified latent Dirichlet allocation (mLDA) based distribution is also introduced. At the classification stage, a k-nearest neighbor (k-NN) classifier is employed. Compared with state of art large-scale genre categorization in, the experimental results on a 23-sports dataset demonstrate that our proposed framework achieves a comparable classification accuracy with 27% and 64% expansion in data volume and diversity, respectively.
Ning Zhang 0023, Ling Guan
MMSP1
2009 Automatic sports genre categorization and view-type classification over large-scale dataset
abstract
This paper presents a framework with two automatic tasks targeting large-scale and low quality sports video archives collected from online video streams. The framework is based on the bag of visual-words model using speeded-up robust features (SURF). The first task is sports genre categorization based on hierarchical structure. Following on the second task which is based on automatically obtained genre, views are classified using support vector machines (SVMs). As a consequence, the views classification result can be used in video parsing and highlight extraction. As compared with state-of-the-art methods, our approach is fully automatic as well as domain knowledge free and thus provides a better extensibility. Furthermore, our dataset consists of 14 sport genres with 6850 minutes in total. Both sport genre categorization and view type classification have more than 80% accuracy rates, which validate this framework's robustness and potential in web-based applications.
Lingfang Li, Ning Zhang 0023, Ling-Yu Duan, Qingming Huang, Ling Guan
ACM Multimedia2
2009 Context-based entropy coding in AVS video coding standard
Li Zhang 0006, Qiang Wang 0011, Ning Zhang 0023, Debin Zhao, Xiaolin Wu 0001, Wen Gao 0001
Signal Process. Image Commun.3
2007 Context-based Arithmetic Coding Reexamined for DCT Video Compression
abstract
This paper presents a new context modeling technique for arithmetic coding of DCT coefficients in video compression. A key feature of the new technique is the inclusion of all previously coded coefficient magnitudes in a DCT block in context modeling. This enables adaptive arithmetic coding to exploit the redundancy of the high-order Markov process in the DCT domain with a few conditioning states. In addition, a context weighting technique is used to further improve the coding efficiency. The complexity of the new arithmetic coding scheme is slightly lower than that of context-based adaptive binary arithmetic coding (CABAC) of H.264. Moreover, the scheme is made compatible to the AVS baseline profile. It achieves on average 13% improvement in compression ratio over context-based two dimension variable length coding (C2DVLC) designed for the DCT domain, and a similar coding efficiency as the CABAC technique in H.264.
Li Zhang 0006, Xiaolin Wu 0001, Ning Zhang 0023, Wen Gao 0001, Qiang Wang 0011, Debin Zhao
ISCAS3
2006 Lossless compression of color mosaic images
abstract
Lossless compression of color mosaic images poses a unique and interesting problem of spectral decorrelation of spatially interleaved R, G, B samples. We investigate reversible lossless spectral-spatial transforms that can remove statistical redundancies in both spectral and spatial domains and discover that a particular wavelet decomposition scheme, called Mallat wavelet packet transform, is ideally suited to the task of decorrelating color mosaic data. We also propose a low-complexity adaptive context-based Golomb-Rice coding technique to compress the coefficients of Mallat wavelet packet transform. The lossless compression performance of the proposed method on color mosaic images is apparently the best so far among the existing lossless image codecs.
Ning Zhang 0023, Xiaolin Wu 0001
IEEE Trans. Image Process.1
2005 On multirate optimality of JPEG2000 code stream
abstract
Arguably, the most important and defining feature of the JPEG2000 image compression standard is its R-D optimized code stream of multiple progressive layers. This code stream is an interleaving of many scalable code streams of different sample blocks. In this paper, we reexamine the R-D optimality of JPEG2000 scalable code streams under an expected multirate distortion measure (EMRD), which is defined to be the average distortion weighted by a probability distribution of operational rates in a given range, rather than for one or few fixed rates. We prove that the JPEG2000 code stream constructed by embedded block coding of optimal truncation is almost optimal in the EMRD sense for uniform rate distribution function, even if the individual scalable code streams have nonconvex operational R-D curves. We also develop algorithms to optimize the JPEG2000 code stream for exponential and Laplacian rate distribution functions while maintaining compatibility with the JPEG2000 standard. Both of our analytical and experimental results lend strong support to JPEG2000 as a near-optimal scalable image codec in a fairly general setting.
Xiaolin Wu 0001, Sorina Dumitrescu, Ning Zhang 0023
IEEE Trans. Image Process.3
2004 Lossless compression of color mosaic images
abstract
We present a low complexity algorithm for lossless compression of color mosaic images generated by a Bayer CCD color filter array. This algorithm is based on an interesting use of the integer wavelet transform followed by a fast adaptive context-based Golomb-Rice coding. The lossless compression performance of the proposed algorithm is apparently the best reported in the literature so far for color mosaic images.
Ning Zhang 0023, Xiaolin Wu 0001
ICIP1
2004 Primary-consistent soft-decision color demosaicking for digital cameras (patent pending)
abstract
Color mosaic sampling schemes are widely used in digital cameras. Given the resolution of CCD sensor arrays, the image quality of digital cameras using mosaic sampling largely depends on the performance of the color demosaicking process. A common problem with existing color demosaicking algorithms is an inconsistency of sample interpolations in different primary color channels, which is the cause of the most objectionable color artifacts. To cure the problem, we propose a new primary-consistent soft-decision framework (PCSD) of color demosaicking. In the PCSD framework, we make multiple estimates of a missing color sample under different hypotheses on edge or texture directions. The estimates are made via a primary consistent interpolation, meaning that all three primary components of a color are interpolated in the same direction. The final estimate of a color sample is obtained by testing different interpolation hypotheses in the reconstructed full-resolution color image and selecting the best via an optimal statistical decision or inference process. A concrete color demosaicking method of the PCSD framework is presented. This new method eliminates certain types of color artifacts of existing color demosaicking methods. Extensive experimental results demonstrate that the PCSD approach can significantly improve the image quality of digital cameras in both subjective and objective measures. In some instances, our gain over the competing methods can be as much as 7 dB.
Xiaolin Wu 0001, Ning Zhang 0023
IEEE Trans. Image Process.2
2003 Primary-consistent soft-decision color demosaic for digital cameras
abstract
Bayer color mosaic sampling scheme is widely used in digital cameras. Given the resolution of CCD sensor arrays, the image quality of digital cameras using Bayer sampling mosaic largely depends on the performance of the color demosaic process. A common and serious weakness shared by all existing color demosaic algorithms is an inconsistency of sample interpolations in different primary color components, which is the culprit for the most objectionable color artifacts. To cure the problem we propose a primary-consistent color demosaic algorithm. The performance of this algorithm is further enhanced by a soft-decision sample interpolation scheme. Experiments demonstrate that the proposed framework of primary-consistent soft-decision color demosaic can significantly improve the image quality of digital cameras.
Xiaolin Wu 0001, Ning Zhang 0023
ICIP (1)2