EDBT 2026 Demo / reviewers in the wild / expert
Zhengguo Li
dblp:46/4171 · also Zheng Guo Li
· DBLP profile ↗
174ranked-venue papers
41as first author
49since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 142 · 39 first-author · 34 since 2021Artificial intelligence and machine learning · 21 · 14 since 2021Systems, architecture and hardware · 11 · 4 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong BaselineabstractMetalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processes are often necessary to recover high-quality RGB images from metalens-captured inputs. While recent deep learning-based restoration methods show promise, they typically (1) blur or distort peripheral regions, or (2) fail entirely under unseen illumination conditions. To advance metalens image restoration, we introduce IlluMeta---the first and largest real-world, illumination-aware metalens image dataset—captured across diverse lighting environments. In addition, we propose a novel end-to-end restoration framework that directs attention to challenging regions and adaptively adjusts to varying illuminations via reinforcement learning. Experiments show that our method can be applied in a plug-and-play manner to enhance existing models, significantly improving image restoration quality, especially under unseen lighting conditions, paving the way for broader real-world deployment of metalens technologies. Fen Fang, Xinan Liang, Muli Yang, Jinghong Zheng 0001, Tobias Wilhelm W. Mass, Ying Sun 0001, Xulei Yang, Xuewu Xu, Zhengguo Li |
AAAI | 9 |
| 2026 | Next-Generation Metalens Vision System: Powered by AI and Applied to AIabstractMetalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical aberrations, blur, and illumination sensitivity, which degrade both visual quality and machine perception. In this demonstration, we present an end-to-end metalens vision system—from hardware sensing with a custom-built RGB metalens camera, to physics-informed imaging and real-time restoration, and finally to downstream vision applications such as object detection and depth estimation. By integrating spatially-aware attention enhancement and reinforcement learning-based illumination control into a real-time system, our solution transforms degraded raw captures into high-fidelity images that are both visually interpretable and functionally reliable for machine vision. This AI-powered pipeline highlights metalenses as a cornerstone for next-generation imaging, where advances in optics and machine intelligence jointly drive the future of visual perception. Fen Fang, Muli Yang, Henan Wang, Xinan Liang, Tobias Wilhelm W. Mass, Xuewu Xu, Xulei Yang, Zhengguo Li |
AAAI | 8 |
| 2026 | SRNet: Self-supervised structure regularization for stereo matching
Jun Cheng 0003, Zaiwang Gu, Weide Liu, Jiayuan Fan 0001, Zhengguo Li, Chuan-Sheng Foo |
Neurocomputing | 5 |
| 2026 | ADVersa: Abductive Driving Accident Video UnderstandingabstractUnderstanding traffic accident scenes is a long-standing research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video-based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model, which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis, and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding. Lei-Lei Li, Jianwu Fang, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | VDMamba: Vector Decomposition in Vision Mamba for Image Deraining and BeyondabstractImage deraining aims to remove rain perturbations from rainy images and restore clear backgrounds. Recent research has employed the Mamba technique for image restoration, achieving exceptional results due to its effectiveness and efficiency in modeling long-range sequence relationships. However, a significant challenge remains: developing a comprehensive framework that considers the intrinsic coupling characteristics between image deraining and the Mamba architecture is largely unexplored. We propose that introducing a 1D sequential representation of Mamba could enhance image deraining by characterizing the direction-aware distribution of rain perturbations. This motivates us to introduce a new vector decomposition-based vision Mamba approach (VDMamba). This method investigates vector decomposition within the context of vision Mamba, addressing the challenging task of image deraining and beyond in the frequency embedding space. The key innovation of VDMamba is the Mamba-based vector decomposition and synthesis module (VDSM). This module derives 1D basic vectors (vertical and horizontal) from the frequency components via vector decomposition and employs the single-direction scanning of Mamba to eliminate the direction-specific degradation perturbation. This transformation allows the incipient Mamba to explore directionspecific global relationships for accurate perturbation learning, without requiring an elaborate design of the Mamba scanning. Additionally, the vertical and horizontal components in VDSM are encoded jointly in a bidirectional coupling manner, enabling the exploration of complementary and redundant components for refinement. Experiments on various image enhancement tasks, including image deraining, raindrop removal, rain haze removal, image dehazing, low-light image enhancement, and underwater image enhancement, demonstrate that VDMamba delivers competitive performance compared to the NeRD method. Specifically, it achieves a 0.58 dB improvement in PSNR for the image deraining task while reducing model parameters by 94.3%, computational cost by 88.3%, and inference time by 77.5%. Kui Jiang, Junjun Jiang, Shiqi Wang 0001, Wenqi Ren, Chia-Wen Lin, Zhengguo Li |
IEEE Trans. Multim. | 6 |
| 2025 | IEBins: Iterative Elastic Bins for Monocular Depth Estimation and Completion
Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
Int. J. Comput. Vis. | 5 |
| 2025 | Efficient motion feature aggregation for optical flow via locality-sensitive hashing
Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li |
Neurocomputing | 5 |
| 2025 | Neural augmentation based panoramic high dynamic range stitching
Chaobing Zheng, Weihai Chen, Shiqian Wu, Zhengguo Li |
Neurocomputing | 6 |
| 2025 | CrossFlow: Learning cost volumes for optical flow by cross-matching local and non-local image features
Zimeng Liu, Xingming Wu, Weihai Chen, Zhong Liu 0005, Zhengguo Li |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | A Semantic-Aware Detail Adaptive Network for Image EnhancementabstractLow-light images often suffer from varying degrees of visual degradation. Current methods for recovering image texture details fail to rely on the self-adaptive correlation texture direction of the image itself, which leads the network to be unable to address the local texture characteristics of different images. To address this challenge, we propose a semantic-aware detail adaptive network (SDANet) that fully considers the image detail information. The network divides low-light images into high-frequency and low-frequency parts. Learning different forms of noise through a novel total variation regularization module with adaptive weights ensures that the final high-frequency part adequately integrates the texture information of the image. Simultaneously, a detail-adaptive module is incorporated to restore finer details in the resulting image. SDANet not only effectively suppresses noise in real low-light images while considering texture details but also effectively addresses the degradation of visible information, and it performs better than other state-of-the-art methods. The code is available athttps://github.com/cheer79/SDANet. Xuekai Wei, Mingliang Zhou 0001, Jielu Yan, Huayan Pu, Jun Luo 0003, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | An Attention-Locating Algorithm for Eliminating Background Effects in Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is a challenging task characterized by interclass similarity and intraclass diversity and has broad application prospects. Recently, several methods have adopted the vision Transformer (ViT) in FGVC tasks since the data specificity of the multihead self-attention (MSA) mechanism in ViT is beneficial for extracting discriminative feature representations. However, these works focus on integrating feature dependencies at a high level, which leads to the model being easily disturbed by low-level background information. To address this issue, we propose a fine-grained attention-locating vision Transformer (FAL-ViT) and an attention selection module (ASM). First, FAL-ViT contains a two-stage framework to identify crucial regions effectively within images and enhance features by strategically reusing parameters. Second, the ASM accurately locates important target regions via the natural scores of the MSA, extracting finer low-level features to offer more comprehensive information through position mapping. Extensive experiments on public datasets demonstrate that FAL-ViT outperforms the other methods in terms of performance, confirming the effectiveness of our proposed methods. The source code is available athttps://github.com/Yueting-Huang/FAL-ViT. Yueting Huang, Zhenzhe Hechen, Mingliang Zhou 0001, Zhengguo Li, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion ModelabstractOver the past few years, self-supervised monocular depth estimation has received widespread attention. Most efforts focus on designing different types of network architectures and loss functions or handling edge cases, for example, occlusion and dynamic objects. In this work, we take another path and propose a novel conditional diffusion-based generative framework for self-supervised monocular depth estimation, dubbed MonoDiffusion. Because the depth ground-truth is unavailable in a self-supervised setting, we develop a new pseudo ground-truth diffusion process to assist the diffusion for training. Instead of diffusing at a fixed high resolution, we perform diffusion in a coarse-to-fine manner that allows for faster inference time without sacrificing accuracy or even better accuracy. Furthermore, we develop a simple yet effective contrastive depth reconstruction mechanism to enhance the denoising ability of model. It is worth noting that the proposed MonoDiffusion has the property of naturally acquiring the depth uncertainty that is essential to be implemented in safety-critical cases. Extensive experiments on the KITTI, Make3D and DIML datasets indicate that our MonoDiffusion outperforms prior state-of-the-art self-supervised competitors. The source code will be publicly available upon the acceptance. Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | For Overall Nighttime Visibility: Integrate Irregular Glow Removal With Glow-Aware EnhancementabstractCurrent low-light image enhancement (LLIE) techniques truly enhance luminance but have limited exploration on another harmful factor of nighttime visibility, the glow effects with multiple shapes in the real world. The presence of glow is inevitable due to widespread artificial light sources, and direct enhancement can cause further glow diffusion. In the pursuit of Overall Nighttime Visibility Enhancement (ONVE), we propose a physical model guided framework ONVE to derive a Nighttime Imaging Model with Near-Field Light Sources (NIM-NLS), whose APSF prior generator is validated efficiently in six categories of glow shapes. Guided by this physical-world model as domain knowledge, we subsequently develop an extensible Light-aware Blind Deconvolution Network (LBDN) to face the blind decomposition challenge on direct transmission map D and light source map G based on APSF. Then, an innovative Glow-guided Retinex-based progressive Enhancement module (GRE) is introduced as a further optimization on reflection R from D to harmonize the conflict of glow removal and brightness boost. Notably, ONVE is an unsupervised framework based on a zero-shot learning strategy and uses physical domain knowledge to form the overall pipeline and network. Empirical evaluations on multiple datasets validate the remarkable efficacy of the proposed ONVE in improving nighttime visibility and performance of high-level vision tasks. Wanyu Wu, Wei Wang 0170, Zheng Wang 0007, Kui Jiang, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Gating Syn-to-Real Knowledge for Pedestrian Crossing Prediction in Safe DrivingabstractPedestrian crossing prediction (PCP) in driving scenes plays a critical role in ensuring the safe decision of intelligent vehicles. Due to the limited observations and annotations of pedestrian crossing behaviors in real situations, recent studies have begun to leverage synthetic data with flexible variation to boost prediction performance, employing domain adaptation frameworks. However, different domain knowledge has distinct cross-domain distribution gaps, which necessitates suitable domain knowledge adaption ways for PCP tasks. In this work, we propose a gated syn-to-real knowledge transfer approach for PCP (Gated-S2R-PCP), which has two aims: 1) designing the suitable domain adaptation ways for different kinds of crossing-domain knowledge, and 2) transferring suitable knowledge for specific situations with gated knowledge fusion. Specifically, we design a framework that contains three domain adaption methods including style transfer, distribution approximation, and knowledge distillation for various information, such as visual, semantic, depth, bounding boxes, etc. A learnable gated unit (LGU) is employed to fuse suitable cross-domain knowledge to boost pedestrian crossing prediction. We construct a new synthetic benchmark S2R-PCP-3181 with 3181 sequences (489,740 frames) which contains the pedestrian bounding boxes, RGB frames, semantic segmentation maps, and depth maps. With the synthetic S2R-PCP-3181, we transfer the knowledge to two real challenging datasets of PIE and JAAD, and superior PCP performance is obtained to the state-of-the-art methods. Jianwu Fang, Chen Lv 0001, Jianru Xue, Zhengguo Li |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video SynthesisabstractTraffic Accident Anticipation (TAA) in traffic scenes is a challenging problem for achieving zero fatalities in the future. Current approaches typically treat TAA as a supervised learning task needing the laborious annotation of accident occurrence duration. However, the inherent long-tailed, uncertain, and fast-evolving nature of traffic scenes has the problem that real causal parts of accidents are difficult to identify and are easily dominated by data bias, resulting in a background confounding issue. Thus, we propose an Attentive Video Diffusion (AVD) model that synthesizes additional accident video clips by generating the causal part in dashcam videos, i.e., from normal clips to accident clips. AVD aims to generate causal video frames based on accident or accident-free text prompts while preserving the style and content of frames for TAA after video generation. This approach can be trained using datasets collected from various driving scenes without any extra annotations. Additionally, AVD facilitates an Equivariant TAA (EQ-TAA) with an equivariant triple loss for an anchor accident-free video clip, along with the generated pair of contrastivepseudo-normalandpseudo-accidentclips. Extensive experiments have been conducted to evaluate the performance of AVD and EQ-TAA, and competitive performance compared to state-of-the-art methods has been obtained. Jianwu Fang, Lei-Lei Li, Zhedong Zheng, Hongkai Yu, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Multim. | 6 |
| 2025 | Transformer-Based and Structure-Aware Dual-Stream Network for Low-Light Image EnhancementabstractIn this article, we propose an end-to-end Transformer-based and structure-aware dual-stream network for low-light image enhancement. First, we divide the dual-stream network into a main stream and a structure stream. The main stream is used to recover an enhanced image and supply structural information to the structure stream, whereas the structure stream is designed to extract structural features from the main stream to provide rich structural information for the enhancement process. Second, we devise a structure-gated Transformer to balance the extraction of global and local features through parallel multihead self-attention with convolution operations following a multilayer perceptron, thus extracting sufficient structural information from the main stream in the encoder part of the dual-stream network. Finally, we develop a cross-attention-based feature fusion module that divides different window sizes into distinct feature fusion stages to achieve multistage, multiscale and multitype feature fusion in the decoder part of the dual-stream network. The experimental results show that our method not only works effectively but also offers favorable computational and storage overheads. Mingliang Zhou 0001, Shuqi Han, Jun Luo 0003, Xu Zhuang, Qin Mao, Zhengguo Li |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study
Yuecong Xu, Haozhi Cao, Jianxiong Yin, Zhenghua Chen, Xiaoli Li 0001, Zhengguo Li, Qianwen Xu 0001, Jianfei Yang 0001 |
Int. J. Comput. Vis. | 6 |
| 2024 | NDDepth: Normal-Distance Assisted Monocular Depth Estimation and CompletionabstractOver the past few years, monocular depth estimation and completion have been paid more and more attention from the computer vision community because of their widespread applications. In this paper, we introduce novel physics (geometry)-driven deep learning frameworks for these two tasks by assuming that 3D scenes are constituted with piece-wise planes. Instead of directly estimating the depth map or completing the sparse depth map, we propose to estimate the surface normal and plane-to-origin distance maps or complete the sparse surface normal and distance maps as intermediate outputs. To this end, we develop a normal-distance head that outputs pixel-level surface normal and distance. Afterthat, the surface normal and distance maps are regularized by a developed plane-aware consistency constraint, which are then transformed into depth maps. Furthermore, we integrate an additional depth head to strengthen the robustness of the proposed frameworks. Extensive experiments on the NYU-Depth-v2, KITTI and SUN RGB-D datasets demonstrate that our method exceeds in performance prior state-of-the-art monocular depth estimation and completion competitors. Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Vision-Based Traffic Accident Detection and Anticipation: A SurveyabstractTraffic accident detection and anticipation is an obstinate road safety problem and painstaking efforts have been devoted. With the rapid growth of video data, Vision-based Traffic Accident Detection and Anticipation (named Vision-TAD and Vision-TAA) become the last one-mile problem for safe driving and surveillance safety. However, the long-tailed, unbalanced, highly dynamic, complex, and uncertain properties of traffic accidents form the Out-of-Distribution (OOD) feature for Vision-TAD and Vision-TAA. Current AI development may focus on these OOD but important problems. What has been done for Vision-TAD and Vision-TAA? What direction we should focus on in the future for this problem? A comprehensive survey is important. We present the first survey on Vision-TAD in the deep learning era and the first-ever survey for Vision-TAA. The pros and cons of each research prototype are discussed in detail during the investigation. In addition, we also provide a critical review of 31 publicly available benchmarks and related evaluation metrics. Through this survey, we want to spawn new insights and open possible trends for Vision-TAD and Vision-TAA tasks. Jianwu Fang, Jiahuan Qiao, Jianru Xue, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Online Unsupervised Video Object Segmentation via Contrastive Motion ClusteringabstractOnline unsupervised video object segmentation (UVOS) uses the previous frames as its input to automatically separate the primary object(s) from a streaming video without using any further manual annotation. A major challenge is that the model has no access to the future and must rely solely on the history, i.e., the segmentation mask is predicted from the current frame as soon as it is captured. In this work, a novel contrastive motion clustering algorithm with an optical flow as its input is proposed for the online UVOS by exploiting the common fate principle that visual elements tend to be perceived as a group if they possess the same motion pattern. We build a simple and effective auto-encoder to iteratively summarize non-learnable prototypical bases for the motion pattern, while the bases in turn help learn the representation of the embedding network. Further, a contrastive learning strategy based on a boundary prior is developed to improve foreground and background feature discrimination in the representation learning stage. The proposed algorithm can be optimized on arbitrarily-scale data (i.e., frame, clip, dataset) and performed in an online fashion. Experiments on$\textit {DAVIS}_{\textit {16}}$, FBMS, and SegTrackV2 datasets show that the accuracy of our method surpasses the previous state-of-the-art (SoTA) online UVOS method by a margin of 0.8%, 2.9%, and 1.1%, respectively. Furthermore, by using an online deep subspace clustering to tackle the motion grouping, our method is able to achieve higher accuracy at$3\times $faster inference time compared to SoTA online UVOS method, and making a good trade-off between effectiveness and efficiency. Our code is available athttps://github.com/xilin1991/CluterNet. Lin Xi, Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Image Quality Assessment: Measuring Perceptual Degradation via Distribution Measures in Deep Feature SpacesabstractThis study aims to develop advanced and training-free full-reference image quality assessment (FR-IQA) models based on deep neural networks. Specifically, we investigate measures that allow us to perceptually compare deep network features and reveal their underlying factors. We find that distribution measures enjoy advanced perceptual awareness and test the Wasserstein distance (WSD), Jensen-Shannon divergence (JSD), and symmetric Kullback-Leibler divergence (SKLD) measures when comparing deep features acquired from various pretrained deep networks, including the Visual Geometry Group (VGG) network, SqueezeNet, MobileNet, and EfficientNet. The proposed FR-IQA models exhibit superior alignment with subjective human evaluations across diverse image quality assessment (IQA) datasets without training, demonstrating the advanced perceptual relevance of distribution measures when comparing deep network features. Additionally, we explore the applicability of deep distribution measures in image super-resolution enhancement tasks, highlighting their potential for guiding perceptual enhancements. The code is available on website. (https://github.com/Buka-Xing/Deep-network-based-distribution-measures-for-full-reference-image-quality-assessment). Xingran Liao, Xuekai Wei, Mingliang Zhou 0001, Zhengguo Li, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2024 | Graph-Represented Distribution Similarity Index for Full-Reference Image Quality AssessmentabstractIn this paper, we propose a graph-represented image distribution similarity (GRIDS) index for full-reference (FR) image quality assessment (IQA), which can measure the perceptual distance between distorted and reference images by assessing the disparities between their distribution patterns under a graph-based representation. First, we transform the input image into a graph-based representation, which is proven to be a versatile and effective choice for capturing visual perception features. This is achieved through the automatic generation of a vision graph from the given image content, leading to holistic perceptual associations for irregular image regions. Second, to reflect the perceived image distribution, we decompose the undirected graph into cliques and then calculate the product of the potential functions for the cliques to obtain the joint probability distribution of the undirected graph. Finally, we compare the distances between the graph feature distributions of the distorted and reference images at different stages; thus, we combine the distortion distribution measurements derived from different graph model depths to determine the perceived quality of the distorted images. The empirical results obtained from an extensive array of experiments underscore the competitive nature of our proposed method, which achieves performance on par with that of the state-of-the-art methods, demonstrating its exceptional predictive accuracy and ability to maintain consistent and monotonic behaviour in image quality prediction tasks. The source code is publicly available at the following website https://github.com/Land5cape/GRIDS. Wenhao Shen, Mingliang Zhou 0001, Jun Luo 0003, Zhengguo Li, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2024 | URCDC-Depth: Uncertainty Rectified Cross-Distillation With CutFlip for Monocular Depth EstimationabstractThis work aims to estimate a high-quality depth map from a single RGB image. Due to the lack of depth clues, making full use of the long-range correlation and local information is critical for accurate depth estimation. To this end, we introduce an uncertainty rectified cross-distillation between the Transformer and convolutional neural network (CNN) to achieve a comprehensive depth estimator. Specifically, we utilize the depth estimates from the Transformer branch and CNN branch as pseudo labels to teach each other. At the same time, the pixel-wise depth uncertainty is modeled to mitigate the negative impact of noisy pseudo labels. To avoid the large capacity gap induced by the strong Transformer branch deteriorating the cross-distillation, we transfer the feature maps from the Transformer to the CNN and develop coupling units to assist the weak CNN branch in leveraging the transferred features. Furthermore, we introduce CutFlip, a surprisingly simple yet highly effective data augmentation technique, which forces the model to focus on more valuable depth reasoning clues apart from the vertical image position. Extensive experiments demonstrate that our model, termedURCDC-Depth, exceeds in performance previous state-of-the-art approaches on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets, with no additional computational burden in the evaluation phase. The source code will be publicly available upon acceptance. The source code is available athttps://github.com/ShuweiShao/URCDC-Depth. Shuwei Shao, Zhongcai Pei, Weihai Chen, Zhong Liu 0005, Zhengguo Li |
IEEE Trans. Multim. | 6 |
| 2023 | NDDepth: Normal-Distance Assisted Monocular Depth EstimationabstractMonocular depth estimation has drawn widespread attention from the vision community due to its broad applications. In this paper, we propose a novel physics (geometry)-driven deep learning framework for monocular depth estimation by assuming that 3D scenes are constituted by piece-wise planes. Particularly, we introduce a new normal-distance head that outputs pixel-level surface normal and plane-to-origin distance for deriving depth at each position. Meanwhile, the normal and distance are regularized by a developed plane-aware consistency constraint. We further integrate an additional depth head to improve the robustness of the proposed framework. To fully exploit the strengths of these two heads, we develop an effective contrastive iterative refinement module that refines depth in a complementary manner according to the depth uncertainty. Extensive experiments indicate that the proposed method exceeds previous state-of-the-art competitors on the NYU-Depth-v2, KITTI and SUN RGB-D datasets. Notably, it ranks 1st among all submissions on the KITTI depth prediction online benchmark at the submission time. The source code is available at https://github.com/ShuweiShao/NDDepth. Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhengguo Li |
ICCV | 5 |
| 2023 | Neural Augmented Exposure Interpolation for HDR ImagingabstractBrightness order reversal usually appears when two large-exposure-ratio images of a high dynamic range scene are directly fused together by an existing multi-scale exposure fusion algorithm. To address the problem, a novel neural augmented framework is introduced to interpolate an image with the medium exposure by integrating physics-driven and data-driven approaches. The physics-driven method infers high-frequency information while the data-driven approach learns remaining information for the interpolated image. The interpolated image and two large-exposure-ratio images are fused together. Experimental results show that the proposed framework can indeed solve the brightness order reversal problem for the fusion of of two large-exposure-ratio images. Zhengguo Li, Chaobing Zheng, Jinghong Zheng 0001, Shiqian Wu |
ICIP | 1 |
| 2023 | Progressive Physics-Driven Deep Conversion Of sRGB to Raw ImagesabstractRaw images are commonly processed by image signal processing (ISP) algorithms to standard RGB (sRGB) images in order to save the storage space and provide a suitable format for human visual system. Distortion and information loss are introduced in this procedure which degrades the performance of following tasks with these sRGB images as the input. To leverage the benefits from raw images, reversing sRGB images to raw images is expected. In this paper, a progressive physics-driven deep learning algorithm is proposed for the conversion of sRGB images to higher quality raw images by fusing physical-driven and data-driven deep learning approaches. Based on an observation that deep convolution neural networks (CNNs) are biased towards learning low-frequency functions, the proposed framework includes a high-frequency aware guidance branch to provide progressive guidance for the reconstruction branch. Experimental results indicate that the proposed algorithm improves reconstruction quality of existing data-driven approaches. It also reduces the sensitivity of existing data-driven approaches to test data in the sense that those images that are hard to be restored by the existing approaches are improved much more. Haiyan Shu, Zhengguo Li, Jinghong Zheng 0001 |
IECON | 2 |
| 2023 | Monocular Depth Estimation: A SurveyabstractMonocular depth estimation is an ill-posed task in computer vision, which holds great significance in the fields such as artificial intelligence, virtual reality, augmented reality, path planning, unmanned driving, and navigation guidance. The primary objective of monocular depth estimation is to predict the depth value of each pixel or infer depth information, given just a single red-green-blue (RGB) image as input. Traditional monocular depth estimation methods rely on limited depth cues, such as strict scene conditions. With the significant advancements in computer vision and artificial intelligence, monocular depth estimation using deep learning has been extensively researched and has yielded substantial results. This paper presents a comprehensive survey of monocular depth estimation. Firstly, we give an overall introduction to monocular depth estimation and explain it from traditional and deep learning-based methods, respectively. To specify, supervised, self-supervised and semi-supervised models are described in detail in deep learning-based methods. Additionally, we introduce publicly available benchmark datasets and evaluation metrics commonly used in this field. Finally, we discuss the current challenges and promising prospects for the development of monocular depth estimation. Dong Wang 0051, Zhong Liu 0005, Shuwei Shao, Xingming Wu, Weihai Chen, Zhengguo Li |
IECON | 6 |
| 2023 | Physics-Driven Deep Panoramic Imaging for High Dynamic Range ScenesabstractDue to saturated regions of low dynamic range (LDR) images and large intensity changes among them, it is challenging to produce an information-enriched panoramic LDR image without visual artifacts from multiple geometrically synchronized LDR images with different exposures and piecewise overlapping fields of views for a high dynamic range (HDR) scene. Fortunately, the stitching of such images is innately a perfect scenario for the fusion of physics-driven and data-driven methods. Based on the insight, a novel neural augmented HDR panoramic stitching algorithm is proposed in this paper. Differently exposed panoramic LDR images are initialized by using a physics-driven method on top of the piecewise overlapping fields of views. They are then refined by a data-driven one, and finally merged together via a multi-scale exposure fusion algorithm to produce the desired panoramic LDR image. Experimental results validate the proposed algorithm11The source code and trained model will be publicly available upon the acceptance.. Chaobing Zheng, Weihai Chen, Shiqian Wu, Zhengguo Li |
IECON | 5 |
| 2023 | IEBins: Iterative Elastic Bins for Monocular Depth EstimationabstractMonocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem where a linear combination of probabilistic distribution and bin centers is used to predict depth. In this paper, we propose a novel concept of iterative elastic bins (IEBins) for the classification-regression-based MDE. The proposed IEBins aims to search for high-quality depth by progressively optimizing the search range, which involves multiple stages and each stage performs a finer-grained depth search in the target bin on top of its previous stage. To alleviate the possible error accumulation during the iterative process, we utilize a novel elastic target bin to replace the original target bin, the width of which is adjusted elastically based on the depth uncertainty. Furthermore, we develop a dedicated framework composed of a feature extractor and an iterative optimizer that has powerful temporal context modeling capabilities benefiting from the GRU-based architecture. Extensive experiments on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets demonstrate that the proposed method surpasses prior state-of-the-art competitors. The source code is publicly available at https://github.com/ShuweiShao/IEBins. Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 0005, Weihai Chen, Zhengguo Li |
NeurIPS | 6 |
| 2023 | Unsupervised Optical Flow Estimation for Differently Exposed Images in LDR DomainabstractDifferently exposed low dynamic range (LDR) images are often captured sequentially using a smart phone or a digital camera with movements. Optical flow thus plays an important role in ghost removal for high dynamic range (HDR) imaging. The optical flow estimation is based on the theory of photometric consistency, which assumes that the corresponding pixels between two images have the same intensity. However, the assumption is no longer valid for the differently exposed LDR images since a pixel’s intensity changes significantly inter images. To address the problem, an unsupervised optical flow estimation framework, is presented in this study. Intensity mapping functions (IMFs) are first adopted to alleviate the intensity changes between the LDR images. Then a novel IMF-based unsupervised learning objective is proposed to circumvent the need for ground truth optical flows when training the deep network. Experimental results and ablation studies on publicly available datasets show that our framework outperforms the state-of-the-art unsupervised optical flow methods, demonstrating the effectiveness of the IMF and the learning objective. Our code is available athttps://github.com/liuziyang123/LDRFlow. Zhengguo Li, Weihai Chen, Xingming Wu, Zhong Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multi-Source Video Domain Adaptation With Temporal Attentive Moment Alignment NetworkabstractMulti-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios, which relaxes the assumption in conventional Unsupervised Domain Adaptation (UDA) that source data are sampled from a single domain and match a uniform data distribution. The MSDA is more challenging due to the existence of different domain shifts between distinct domain pairs. When considering videos, the negative transfer would be provoked by spatial-temporal features and can be formulated into a more challenging Multi-Source Video Domain Adaptation (MSVDA) problem. In this paper, we address the MSVDA problem by proposing a novel Temporal Attentive Moment Alignment Network (TAMAN) which aims for effective feature transfer by dynamically aligning both spatial and temporal feature moments. The TAMAN further constructs robust global temporal features by attending to dominant domain-invariant local temporal features with high local classification confidence and low disparity between global and local feature discrepancies. To facilitate future research on the MSVDA problem, we introduce comprehensive benchmarks, covering extensive MSVDA scenarios. Empirical results demonstrate a superior performance of the proposed TAMAN across multiple MSVDA benchmarks. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhengguo Li, Zhenghua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | ALIKE: Accurate and Lightweight Keypoint Detection and Descriptor ExtractionabstractExisting methods detect the keypoints in a non-differentiable way, therefore they can not directly optimize the position of keypoints through back-propagation. To address this issue, we present a partially differentiable keypoint detection module, which outputs accurate sub-pixel keypoints. The reprojection loss is then proposed to directly optimize these sub-pixel keypoints, and the dispersity peak loss is presented for accurate keypoints regularization. We also extract the descriptors in a sub-pixel way, and they are trained with the stable neural reprojection error loss. Moreover, a lightweight network is designed for keypoint detection and descriptor extraction, which can run at 95 frames per second for 640x480 images on a commercial GPU. On homography estimation, camera pose estimation, and visual (re-)localization tasks, the proposed method achieves equivalent performance with the state-of-the-art approaches, while greatly reduces the inference time. Xiaoming Zhao 0003, Xingming Wu, Jinyu Miao, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Multim. | 6 |
| 2022 | Simultaneous Smoothing and Sharpening Using iWGIFabstractSmoothing and sharpening are two fundamental operations in image processing, and they are usually conducted separately. In this paper, an improved weighted guided image filter (iWGIF) is first proposed to enhance existing guided filters. A content adaptive smoothing and sharpening (CASS) algorithm is then presented by using coefficients of the iWGIF intelligently. A simple despeckling algorithm is finally introduced to simultaneously smoothen homogeneous regions and preserve heterogeneous regions of synthetic-aperture radar images. Experimental results validate all the proposed algorithms. Zhengguo Li, Jinghong Zheng 0001, J. Senthilnath 0001 |
ICIP | 1 |
| 2022 | Single Image Dehazing via Model-Based Deep-LearningabstractModel-based single image dehazing algorithms restore images with sharp edges and rich details at the expense of low PSNR values. Data-driven ones restore images with high PSNR values but with low contrast, and even some remaining haze. In this paper, a novel single image dehazing algorithm is introduced by integrating model-based and data-driven approaches. Both transmission map and atmospheric light are initialized by the model-based methods, and refined by deep learning based approaches which form a neural augmentation. Haze-free images are restored by using the transmission map and atmospheric light. Experimental results indicate that the proposed algorithm can remove haze well from real-world and synthetic hazy images. Zhengguo Li, Chaobing Zheng, Haiyan Shu, Shiqian Wu |
ICIP | 1 |
| 2022 | Adaptive weighted guided image filtering for depth enhancement in shape-from-focus
Zhengguo Li, Chaobing Zheng, Shiqian Wu |
Pattern Recognit. | 2 |
| 2022 | DSRGAN: Detail Prior-Assisted Perceptual Single Image Super-Resolution via Generative Adversarial NetworksabstractThe generative adversarial network (GAN) is successfully applied to study the perceptual single image super-resolution (SISR). However, since the GAN is data-driven, it has a fundamental limitation on restoring real high frequency information for an unknown instance (or image) during test. On the other hand, the conventional model-based methods have a superiority to achieve instance adaptation as they operate by considering the statistics of each instance (or image) only. Motivated by this, we propose a novel model-based algorithm, which can extract the detail layer of an image efficiently. The detail layer represents the high frequency information of image and it is constituted of image edges and fine textures. It is seamlessly incorporated into the GAN and serves as a prior knowledge to assist the GAN in generating more realistic details. The proposed method, named DSRGAN, takes advantages from both the model-based conventional algorithm and the data-driven deep learning network. Experimental results demonstrate that the DSRGAN outperforms the state-of-the-art SISR methods on perceptual metrics, meanwhile achieving comparable results in terms of fidelity metrics. Following the DSRGAN, it is feasible to incorporate other conventional image processing algorithms into a deep learning network to form a model-based deep SISR. Zhengguo Li, Xingming Wu, Zhong Liu 0005, Weihai Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Implicit Motion-Compensated Network for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (UVOS) aims at automatically separating the primary foreground object(s) from the background in a video sequence. Existing UVOS methods either lack robustness when there are visually similar surroundings (appearance-based) or suffer from deterioration in the quality of their predictions because of dynamic background and inaccurate flow (flow-based). To overcome the limitations, we propose an implicit motion-compensated network (IMCNet) combining complementary cues (i.e., appearance and motion) with aligned motion information from the adjacent frames to the current frame at the feature level without estimating optical flows. The proposed IMCNet consists of an affinity computing module (ACM), an attention propagation module (APM), and a motion compensation module (MCM). The light-weight ACM extracts commonality between neighboring input frames based on appearance features. The APM then transmits global correlation in a top-down manner. Through coarse-to-fine iterative inspiring, the APM will refine object regions from multiple resolutions so as to efficiently avoid losing details. Finally, the MCM aligns motion information from temporally adjacent frames to the current frame which achieves implicit motion compensation at the feature level. We perform extensive experiments on$\textit {DAVIS}_{\textit {16}}$and$\textit {YouTube-Objects}$. Our network achieves favorable performance while running at a faster speed compared to the state-of-the-art methods. Our code is available athttps://github.com/xilin1991/IMCNet. Lin Xi, Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Deep Joint Demosaicing and High Dynamic Range Imaging Within a Single ShotabstractSpatially varying exposure (SVE) is a promising choice for high-dynamic-range (HDR) imaging (HDRI). The SVE-based HDRI, which is called single-shot HDRI, is an efficient solution to avoid ghosting artifacts. However, it is very challenging to restore a full-resolution HDR image from a real-world image with SVE because: a) only one-third of pixels with varying exposures are captured by camera in a Bayer pattern, b) some of the captured pixels are over- and under-exposed. For the former challenge, a spatially varying convolution (SVC) is designed to process the Bayer images carried with varying exposures. For the latter one, an exposure-guidance method is proposed against the interference from over- and under-exposed pixels. Finally, a joint demosaicing and HDRI deep learning framework is formalized to include the two novel components and to realize an end-to-end single-shot HDRI. Experiments indicate that the proposed end-to-end framework avoids the problem of cumulative errors and surpasses the related state-of-the-art methods. Related codes and datasets will be provided athttps://github.com/yilun-xu/SVEHDRI/. Xingming Wu, Weihai Chen, Changyun Wen, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Probabilistic Spatial Distribution Prior Based Attentional Keypoints Matching NetworkabstractKeypoints matching is a pivotal component for many image-relevant applications such as image stitching, visual simultaneous localization and mapping (SLAM), and so on. Both handcrafted-based and recently emerged deep learning-based keypoints matching methods merely rely on keypoints and local features, while losing sight of other available sensors such as inertial measurement unit (IMU) in the above applications. In this paper, we demonstrate that the motion estimation from IMU integration can be used to exploit the spatial distribution prior of keypoints between images. To this end, a probabilistic perspective of attention formulation is proposed to integrate the spatial distribution prior into the attentional graph neural network naturally. With the assistance of spatial distribution prior, the effort of the network for modeling the hidden features can be reduced. Furthermore, we present a projection loss for the proposed keypoints matching network, which gives a smooth edge between matching and un-matching keypoints. Image matching experiments on visual SLAM datasets indicate the effectiveness and efficiency of the presented method. Xiaoming Zhao 0003, Jingmeng Liu, Xingming Wu, Weihai Chen, Fanghong Guo, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Dual-Scale Single Image Dehazing via Neural AugmentationabstractModel-based single image dehazing algorithms restore haze-free images with sharp edges and rich details for real-world hazy images at the expense of low PSNR and SSIM values for synthetic hazy images. Data-driven ones restore haze-free images with high PSNR and SSIM values for synthetic hazy images but with low contrast, and even some remaining haze for real-world hazy images. In this paper, a novel single image dehazing algorithm is introduced by combining model-based and data-driven approaches. Both transmission map and atmospheric light are first estimated by the model-based methods, and then refined by dual-scale generative adversarial networks (GANs) based approaches. The resultant algorithm forms a neural augmentation which converges very fast while the corresponding data-driven approach might not converge. Haze-free images are restored by using the estimated transmission map and atmospheric light as well as the Koschmieder's law. Experimental results indicate that the proposed algorithm can remove haze well from real-world and synthetic hazy images. Zhengguo Li, Chaobing Zheng, Haiyan Shu, Shiqian Wu |
IEEE Trans. Image Process. | 1 |
| 2021 | Non-Local Single Image DE-Raining Without DecompositionabstractIt is challenging to remove rain steaks from a single rainy image because the rain steaks are spatially varying in the rainy image. On top of a new insight in single image de-raining, a nonlocal de-raining algorithm is proposed in this paper to remove the rain streaks from the rainy image. The rainy image is not decomposed into different layers by the proposed algorithm. Experimental results validate the proposed algorithm. Chaobing Zheng, Zhengguo Li, Shiqian Wu |
ICASSP | 2 |
| 2021 | Multi-Scale Model Driven Single Image DehazingabstractModel driven single image dehazing was widely studied due to its broad applications. It is challenging to prevent noise from being amplified in sky region for the model driven dehazing algorithms. In this paper, a new multi-scale hazy image model is first built up by using the Laplacian pyramid of hazy image and the Gaussian pyramid of transmission map. A novel multi-scale dehazing algorithm is then proposed on top of the model to address the problem. Different haze removal and noise reduction approaches are applied to restore the scene radiance at different levels of the pyramid. The resultant pyramid is collapsed to restore a haze-free image. Experiment results demonstrate that the proposed algorithm outperform state of the art dehazing algorithms and the noise is indeed prevented from being amplified in the sky region. Zhengguo Li, Haiyan Shu |
ICIP | 1 |
| 2021 | Restoration of HDR Images for SVE-Based HDRI via a Novel DCNNabstractGhosting artifacts are believed to be the Achilles’ heel for high dynamic range (HDR) imaging (HDRI) via differently exposed images sequentially captured by a digital device. Spatially varying exposure (SVE)-based HDRI is an efficient solution to prevent the ghosting artifacts from appearing in a HDR image. However, it is challenging to restore a high-quality HDR image with the full resolution from a single raw Bayer image for the SVE-based HDRI. In this paper, a novel deep convolution neural network (DCNN) is proposed to address such a challenging problem. The proposed DCNN includes two distinctive components, a spatially varying convolution and an exposedness-aware compensation branch. The evaluations indicate that the quality of our results significantly surpasses several related algorithms. Related materials will be provided at https://github.com/yilun-xu/SVEHDRI/. Xingming Wu, Weihai Chen, Zhengguo Li |
ICME | 5 |
| 2021 | Deep Reinforcement Learning Boosted Partial Domain AdaptationabstractDomain adaptation is critical for learning transferable features that effectively reduce the distribution difference among domains. In the era of big data, the availability of large-scale labeled datasets motivates partial domain adaptation (PDA) which deals with adaptation from large source domains to small target domains with less number of classes. In the PDA setting, it is crucial to transfer relevant source samples and eliminate irrelevant ones to mitigate negative transfer. In this paper, we propose a deep reinforcement learning based source data selector for PDA, which is capable of eliminating less relevant source samples automatically to boost existing adaptation methods. It determines to either keep or discard the source instances based on their feature representations so that more effective knowledge transfer across domains can be achieved via filtering out irrelevant samples. As a general module, the proposed DRL-based data selector can be integrated into any existing domain adaptation or partial domain adaptation models. Extensive experiments on several benchmark datasets demonstrate the superiority of the proposed DRL-based data selector which leads to state-of-the-art performance for various PDA tasks. Keyu Wu 0002, Min Wu 0008, Jianfei Yang 0001, Zhenghua Chen, Zhengguo Li, Xiaoli Li 0001 |
IJCAI | 5 |
| 2021 | Cognitive Navigation for Indoor Environment Using FloorplanabstractThe recent years have seen the increasing importance of cognitive models for improved robot navigation. In this paper, a novel cognitive navigation package, which consists of topometric map representation and a three-level path planner, is proposed. The topometric maps are built from architectural floor plans with additional features within a limited number of selected regions. The inherent discrepancies between floor plans and the robot’s actual sensory readings are handled by the three-level path planner. The unique feature of this approach is that accurate localization is only required at the selected regions. At the other regions the robot will rely on the guiding directions towards the goal rather than on its accurate position on the map for navigation. Experiments show that our approach can endow a robot with capability for semantic interpretation and localization in unseen and dynamic environments. Jun Li 0005, Chee Leong Chan, Jian Le Chan, Zhengguo Li, Kong-Wah Wan, Weiyun Yau |
IROS | 4 |
| 2021 | S&CNet: A lightweight network for fast and accurate depth completion
Weihai Chen, Xingming Wu, Zhengguo Li |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Single Image Brightening via Multi-Scale Exposure Fusion With Hybrid LearningabstractA small ISO and a small exposure time are usually used to capture an image in back- or low-light condition which results in an image with negligible motion blur and small noise but looks dark. In this paper, a single image brightening algorithm is introduced to brighten such an image. The proposed algorithm includes a unique hybrid learning framework to generate two virtual images with large exposure times. The virtual images are first generated via intensity mapping functions (IMFs) which are computed using camera response functions (CRFs) and this is a model-driven approach. Both the virtual images are then enhanced by using a data-driven approach, i.e. a residual convolutional neural network to approach the ground truth images. The model-driven approach and the data-driven one compensate each other in the proposed hybrid learning framework. The final brightened image is obtained by fusing the original image and two virtual images via a multi-scale exposure fusion algorithm with properly defined weights. Experimental results show that the proposed brightening algorithm outperforms existing algorithms in terms of MEF-SSIM metric. Chaobing Zheng, Zhengguo Li, Yi Yang 0021, Shiqian Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Distributed Successive Convex Approximation for Nonconvex Economic Dispatch in Smart GridabstractThis article presents a distributed consensus-based successive convex approximation (DSCA) algorithm to solve nonconvex nondifferentiable economic dispatch (ED) problems. The ED model formulated incorporates generation constraints, valve-point effects, and multiple fuel types. A perturbation technique enables the proposed DSCA to tackle such a nondifferentiable and nonconvex optimization, which paves the way to solving more complicated optimization problems that occur in practical applications. The local generation constraint is taken care by a local surrogate convex optimization directly. The global equality constraint is handled based on a consensus protocol, where the local generation-demand mismatch among all dispatchable generators (DGs) is shared in a distributed manner. As a result, the power distribution of DGs is updated, and the generation cost is minimized. Several case studies show that the proposed DSCA algorithm can achieve superior ED solutions and computational efficiency over existing nonconvex optimization algorithms. Fanghong Guo, Wen-An Zhang 0001, Wei Wang 0016, Changyun Wen, Zhengguo Li |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | Multi-Scale Single Image Dehazing Using Laplacian and Gaussian PyramidsabstractModel-based single image dehazing was widely studied due to its extensive applications. Ambiguity between object radiance and haze and noise amplification in sky regions are two inherent problems of model-based single image dehazing. In this paper, a dark direct attenuation prior (DDAP) is proposed to address the former problem. A novel haze line averaging is proposed to reduce the morphological artifacts caused by the DDAP which enables a weighted guided image filter with a smaller radius to further reduce the morphological artifacts while preserve the fine structure in the image. A multi-scale dehazing algorithm is then proposed to address the latter problem by adopting Laplacian and Gaussian pyramids to decompose the hazy image into different levels and applying different haze removal and noise reduction approaches to restore the scene radiance at the different levels. The resultant pyramid is collapsed to restore a haze-free image. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art dehazing algorithms. Zhengguo Li, Haiyan Shu, Chaobing Zheng |
IEEE Trans. Image Process. | 1 |
| 2020 | Cross Image Cubic Interpolator for Spatially Varying ExposuresabstractSpatially varying exposures via rolling shutter is an efficient way to capture differently exposed images for high dynamic range (HDR) scenes. Neither camera movement nor moving objects is an issue for such a captured method. However, a possible issue is that the resolution of captured images is reduced. In this paper, we introduce a novel cross image cubic interpolator for the spatially varying exposures via the rolling shutter. Both intra correlation among pixels with the same exposure and inter correlation among pixels with the different exposures are utilized by the proposed interpolator. Experimental results show that quality of upsampled images is significantly improved. Zhengguo Li, Jinghong Zheng 0001, Shoulie Xie, Haiyan Shu |
ICASSP | 1 |
| 2020 | Exposure Interpolation Via Hybrid LearningabstractDeep learning based methods have become dominant solutions to many image processing problems. A natural question would be "Is there any space for conventional methods on these problems?" In this paper, exposure interpolation is taken as an example to answer this question and the answer is "Yes". A new hybrid learning framework is introduced to interpolate a medium exposure image for two large-exposure-ratio images from an emerging high dynamic range (HDR) video capturing device. The framework is set up by fusing conventional and deep learning methods. Experimental results indicate that the deep learning method can be used to improve the quality of interpolated image via the conventional method significantly. The conventional method can be adopted to increase the convergence speed of the deep learning method and to reduce the number of samples which is required by the deep learning method. They compensate each other. Chaobing Zheng, Zhengguo Li, Yi Yang 0021, Shiqian Wu |
ICASSP | 2 |
| 2020 | Pre-processing for single image dehazing
Minmin Yang, Jianchang Liu, Zhengguo Li, Shubin Tan |
Signal Process. Image Commun. | 3 |
| 2020 | Detail-Enhanced Multi-Scale Exposure Fusion in YUV Color SpaceabstractIt is recognized that existing multi-scale exposure fusion algorithms can be improved using edge-preserving smoothing techniques. However, the complexity of edge-preserving smoothing-based multi-scale exposure fusion is an issue for mobile devices. In this paper, a simpler multi-scale exposure fusion algorithm is designed in YUV color space. The proposed algorithm can preserve details in the brightest and darkest regions of a high dynamic range (HDR) scene and the edge-preserving smoothing-based multi-scale exposure fusion algorithm while avoiding color distortion from appearing in the fused image. The complexity of the proposed algorithm is about half of the edge-preserving smoothing-based multi-scale exposure fusion algorithm. The proposed algorithm is thus friendlier to the smartphones than the edge-preserving smoothing-based multi-scale exposure fusion algorithm. In addition, a simple detail-enhancement component is proposed to enhance fine details of fused images. The experimental results show that the proposed component can be adopted to produce an enhanced image with visibly enhanced fine details and a higher MEF-SSIM value. This is impossible for existing detail enhancement components. Clearly, the component is attractive for PC-based applications. Qiantong Wang, Weihai Chen, Xingming Wu, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Hazy Image Decolorization With Color Contrast RestorationabstractIt is challenging to convert a hazy color image into a gray-scale image because the color contrast field of a hazy image is distorted. In this paper, a novel decolorization algorithm is proposed to transfer a hazy image into a distortionrecovered gray-scale image. To recover the color contrast field, the relationship between the restored color contrast and its distorted input is presented in CIELab color space. Based on this restoration, a nonlinear optimization problem is formulated to construct the resultant gray-scale image. A new differentiable approximation solution is introduced to solve this problem with an extension of the Huber loss function. Experimental results show that the proposed algorithm effectively preserves the global luminance consistency while represents the original color contrast in gray-scales, which is very close to the corresponding ground truth gray-scale one. Wei Wang 0170, Zhengguo Li, Shiqian Wu, Liangcai Zeng |
IEEE Trans. Image Process. | 2 |
| 2018 | Reciprocal Collision Avoidance for Nonholonomic Mobile RobotsabstractIn this paper, reciprocal collision avoidance is studied for nonholonomic mobile robots to achieve an efficient navigation. Two strategies are proposed to respectively adjust the linear and angular velocities so that a collision-free navigation can be achieved. By characterizing the collision-free navigation as a set of changing ratios for linear velocities, it is shown that collision can be avoided if the changing ratio of linear velocities is inside this set. Moreover, the strategy of the adjusting angular velocities is established following TTC-based method. With the combination of these two strategies, a simulation is done for four robots crossing the intersection, which shows the effectiveness of the proposed method. Lei Wang 0059, Zhengguo Li, Changyun Wen, Fanghong Guo |
ICARCV | 2 |
| 2018 | Outlier Detection using Hierarchical Spatial Verification for Visual Place RecognitionabstractSpatial verification is a key step to remove outliers for accurate feature matching in visual place recognition. In this paper, we propose a novel method for outlier detection using a hierarchical spatial verification scheme. Given a set of putative correspondences between a pair of images, we convert the matching problem into a 4D transformation space and identify promising similarity transformations using Hough voting. In the hierarchical scheme, we first use a hypothesize-and-verify technique to identify groups of correspondences according to each similarity transformation. Second, the group with the largest number of correspondences serves as a standard to subsequently remove outliers in other groups by explicit geometric consistency checking. We have compared the proposed method with the state-of-the-art solutions on five popular public datasets to show that our method has better performance in place recognition and loop closure detection. Miaolong Yuan, Zhengguo Li, Kong-Wah Wan, Weiyun Yau |
ICARCV | 2 |
| 2018 | Detail Preserving Multi-Scale Exposure FusionabstractEdge-preserving smoothing based multi-scale exposure fusion is a state-of-the-art method to fuse differently exposed images of a high dynamic range (HDR) scene. However, its complexity could be an issue. In this paper, a novel multiscale exposure fusion algorithm is proposed by adopting an approximation method at the highest layer of the pyramid. Experimental results show that the proposed algorithm can be applied to fuse images with comparable or even better quality with the edge-preserving smoothing based multi-scale fusion algorithms. It simplifies the complexity of the edge-preserving smoothing based multi-scale exposure fusion algorithms significantly. Qiantong Wang, Weihai Chen, Xingming Wu, Zhengguo Li |
ICIP | 4 |
| 2018 | Lost Robot Self-Recovery via Exploration Using Hybrid Topological-Metric MapsabstractA robot might get lost due to abrupt wheel slippage, unsteady movements on uneven floors, collision with obstacles, blocked perception sensors or kidnapping. When this happens, the robot has to recover by itself. This is necessary but has not been satisfactorily resolved. In this paper, a generic lost robot self-recovery framework is proposed. It has self-exploration capability assisted by an efficient place recognition module using hybrid topological-metric maps. The metric map is used for path planning and navigation while the topological map is used for re-localization when lost. As soon as the robot detects that it is lost, self exploration is activated and starts to explore while performing place recognition using the topological map. Once the place is re-identified, the proximate global location is obtained and the robot performs fine localization using the metric map, recovery itself and continue to its destination. The proposed system has been implemented on a mobile robot operating in a typical office environment. Experiments conducted show a robot can be efficiently and reliably recover itself when it gets lost within or outside of the map. Miaolong Yuan, Weiyun Yau, Zhengguo Li |
TENCON | 3 |
| 2018 | Game theoretical security detection strategy for networked systems
Hao Wu 0008, Wei Wang 0016, Changyun Wen, Zhengguo Li |
Inf. Sci. | 4 |
| 2018 | Edge-preserving smoothing pyramid based multi-scale exposure fusion
Fei Kou, Zhengguo Li, Changyun Wen, Weihai Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Stretchy binary classification
Kar-Ann Toh, Zhiping Lin 0001, Lei Sun 0006, Zhengguo Li |
Neural Networks | 4 |
| 2018 | Multi-Scale Fusion of Two Large-Exposure-Ratio ImagesabstractExisting multiscale exposure fusion (MEF) algorithms cannot preserve relative brightness in an image fused from two large-exposure-ratio images if high-light regions in the dark image are darker than shadow regions in the bright image. In this letter, a strategy by synthesizing a virtual image with a medium exposure is presented to brighten the high-light regions in the dark image and to darken the darkest regions in the bright image. The virtual image is generated via intensity mapping functions. In order to avoid possible color distortion in the virtual image due to one-to-many mapping, two intermediate virtual images with the same exposure time are generated by the two input images, and then merged together to produce the desired virtual image using properly defined weights. The final image is obtained by fusing the original two input images and the virtual image via a state-of-the-art MEF algorithm. Experimental results show that the relative brightness is preserved much better and the MEF-SSIM is significantly improved by the proposed algorithm. Yi Yang 0021, Shiqian Wu, Zhengguo Li |
IEEE Signal Process. Lett. | 4 |
| 2018 | Local Inverse Tone Mapping for Scalable High Dynamic Range Image CodingabstractTone mapping operators (TMOs) and inverse TMOs (iTMOs) are important for scalable coding of high dynamic range (HDR) images. Because of the high nonlinearity of local TMOs, it is very difficult to estimate the iTMO accurately for a local TMO. In this letter, we present a two-layer local iTMO estimation algorithm using an edge-preserving decomposition technique. The low dynamic range (LDR) image is first linearized and then decomposed into a base layer and a detail layer via a fast edge-preserving decomposition method. The base layer of the HDR image is generated by subtracting the LDR detail layer from the HDR image. An iTMO function is finally estimated by solving a novel quadratic optimization problem formulated on the pair of base layers rather than the pair of HDR and LDR images as in existing methods. Experimental results show that the proposed two-layer iTMO can recover the HDR accurately so that it is possible to use these local TMOs in scalable HDR image coding schemes. Changyun Wen, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Single Image De-Hazing Using Globally Guided Image FilteringabstractLocal edge-preserving smoothing techniques such as guided image filtering (GIF) and weighted guided image filtering (WGIF) could not preserve fine structure. In this paper, a new globally guided image filtering (G-GIF) is introduced to overcome the problem. The G-GIF is composed of a global structure transfer filter and a global edge-preserving smoothing filter. The proposed filter is applied to study single image haze removal. Experimental results show that fine structure of the dehazed image is indeed preserved better by the proposed G-GIF and the dehazed images by the proposed G-GIF are sharper than those dehazed images by the existing GIF. Zhengguo Li, Jinghong Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Color Contrast-Preserving DecolorizationabstractDecolorization is to convert a color image into a gray scale image while preserve image features like salient structure and chrominance contrast. The sign of the color contrast is crucial for the decolorization algorithm and is usually determined in existing works by giving a strict defined color order or twomode weak order. In this paper, a fast computation on color order is achieved via a simple global mapping which is introduced in a linear parametric model using an extended structure transfer filter. The values of the parameters are obtained via an elegant approximation method. A local decolorization algorithm is finally designed on basis of the global linear mapping so that both color and spatial information are preserved robustly and accurately. Experimental results show that the proposed decolorization algorithms obtain a good performance among existing quality metrics for the decolorization. In addition, the proposed global decolorization algorithm is friendly to mobile devices with limited computational resource. Wei Wang 0170, Zhengguo Li, Shiqian Wu |
IEEE Trans. Image Process. | 2 |
| 2018 | Structure-Preserving Guided Retinal Image Filtering and Its Application for Optic Disk AnalysisabstractRetinal fundus photographs have been used in the diagnosis of many ocular diseases such as glaucoma, pathological myopia, age-related macular degeneration, and diabetic retinopathy. With the development of computer science, computer aided diagnosis has been developed to process and analyze the retinal images automatically. One of the challenges in the analysis is that the quality of the retinal image is often degraded. For example, a cataract in human lens will attenuate the retinal image, just as a cloudy camera lens which reduces the quality of a photograph. It often obscures the details in the retinal images and posts challenges in retinal image processing and analyzing tasks. In this paper, we approximate the degradation of the retinal images as a combination of human-lens attenuation and scattering. A novel structure-preserving guided retinal image filtering (SGRIF) is then proposed to restore images based on the attenuation and scattering model. The proposed SGRIF consists of a step of global structure transferring and a step of global edge-preserving smoothing. Our results show that the proposed SGRIF method is able to improve the contrast of retinal images, measured by histogram flatness measure, histogram spread, and variability of local luminosity. In addition, we further explored the benefits of SGRIF for subsequent retinal image processing and analyzing tasks. In the two applications of deep learning-based optic cup segmentation and sparse learning-based cup-to-disk ratio (CDR) computation, our results show that we are able to achieve more accurate optic cup segmentation and CDR measurements from images processed by SGRIF. Jun Cheng 0003, Zhengguo Li, Zaiwang Gu, Huazhu Fu, Damon Wing Kee Wong, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Intelligent Detail Enhancement for Exposure FusionabstractMultiscale exposure fusion is a fast approach to fuse several differently exposed images captured at the same high dynamic range (HDR) scene into a high-quality low-dynamic range (LDR) image. The fused image is expected to include all details of the input images. However the details in the brightest and darkest regions are usually not well preserved. Adding details that are extracted from the input images to the fused image is an efficient approach to overcome the problem. In this paper a new gradient domain weighted least square based image smoothing algorithm is proposed to extract the details in the brightest and darkest regions of the HDR scene. The extracted details are then added to an image that is produced using an edge-preserving smoothing pyramid based multiscale exposure fusion algorithm. Experimental results show that the proposed detail enhanced exposure fusion algorithm can preserve details in saturated regions especially the brightest regions better than the state-of-the-art multiscale exposure fusion algorithms. Fei Kou, Weihai Chen, Xingming Wu, Changyun Wen, Zhengguo Li |
IEEE Trans. Multim. | 6 |
| 2018 | Superpixel-Based Single Nighttime Image Haze RemovalabstractHaze removal is important to improve performance of outdoor vision systems. However, it is challenging to remove haze from a single nighttime haze image. In this paper, a novel superpixel-based single image haze removal algorithm is proposed for nighttime haze images. The input nighttime image is first decomposed into a glow image and a glow-free nighttime haze image using their relative smoothness. A superpixel-based method is then introduced to compute the value of the atmospheric light and dark channel for each pixel in the glow-free haze image. The transmission map is decomposed from the dark channel of the glow-free haze image by the weighted guided image filter. Since superpixels usually adhere to the boundaries of objects well, a smaller local window size can be selected. As such, details in areas of fine structures are preserved better. In addition, to avoid noticeable noise in the sky area, an adaptive threshold is added to the transmission map when the nighttime haze image is restored. Experiments show that our method produces better results than the existing haze removal algorithms for nighttime haze images. Minmin Yang, Jianchang Liu, Zhengguo Li |
IEEE Trans. Multim. | 3 |
| 2017 | Intelligent detail enhancement for differently exposed imagesabstractMulti-scale exposure fusion is a fast approach to fuse several differently exposed images captured at the same high dynamic range (HDR) scene into a high quality low dynamic range (LDR) image. The fused image is expected to include all details of the input images, however, the details in the brightest and darkest regions are usually not preserved well. Adding details that are extracted from the input images to the fused image is an efficient approach to overcome the problem. In this paper, a fast selectively detail enhancement algorithm is proposed to extract the details in the brightest and darkest regions of the HDR scene and add the extracted details to the fused image. Experimental results show that the proposed algorithm can enhance the details of the fused image much faster than the existing algorithms with comparable or even better visual quality. Fei Kou, Weihai Chen, Xingming Wu, Zhengguo Li |
ICIP | 4 |
| 2017 | Multi-scale exposure fusion via gradient domain guided image filteringabstractMulti-scale exposure fusion is an efficient way to fuse differently exposed low dynamic range (LDR) images of a high dynamic range (HDR) scene into a high quality LDR image directly. It can produce images with higher quality than single-scale exposure fusion, but has a risk of producing halo artifacts and cannot preserve details in brightest or darkest regions well in the fused image. In this paper, an edge-preserving smoothing pyramid is introduced for the multi-scale exposure fusion. Benefiting from the edge-preserving property of the filter used in the algorithm, the details in the brightest/darkest regions are preserved well and no halo artifacts are produced in the fused image. The experimental results prove that the proposed algorithm produces better fused images than the state-of-the-art algorithms both qualitatively and quantitatively. Fei Kou, Zhengguo Li, Changyun Wen, Weihai Chen |
ICME | 2 |
| 2017 | Detail-Enhanced Multi-Scale Exposure FusionabstractMulti-scale exposure fusion is an effective image enhancement technique for a high dynamic range (HDR) scene. In this paper, a new multi-scale exposure fusion algorithm is proposed to merge differently exposed low dynamic range (LDR) images by using the weighted guided image filter to smooth the Gaussian pyramids of weight maps for all the LDR images. Details in the brightest and darkest regions of the HDR scene are preserved better by the proposed algorithm without relative brightness change in the fused image. In addition, a new weighted structure tensor is introduced to the differently exposed images and it is adopted to design a detail extraction component for the proposed fusion algorithm, such that users are allowed to manipulate fine details in the enhanced image according to their preference. The proposed multi-scale exposure fusion algorithm is also applied to design a simple single image brightening algorithm for both low-light imaging and back-light imaging. Zhengguo Li, Changyun Wen, Jinghong Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Single image brightening via exposure fusionabstractIt is very challenging to capture images via cell phone cameras in low-lighting conditions due to possible motion blur, especially for a high dynamic range (HDR) scene. In this paper, a new single image brightening algorithm is introduced to support capturing an image with a small exposure time and a small ISO value for a low-lighting scene. There are negligible motion blur and over-exposed pixels in the captured image while both details in the darkest regions and the brightness of the image are reduced. The proposed algorithm is applied to brighten the under-exposed regions and to enhance details of the under-exposed regions with negligible increment on the brightness of the brightest areas. The proposed algorithm can also be adopted to brighten an image captured for an HDR scene at day time but with dark objects. Zhengguo Li, Jinghong Zheng 0001 |
ICASSP | 1 |
| 2016 | A mutual local-ternary-pattern based method for aligning differently exposed images
Shiqian Wu, Lingxian Yang, Wangming Xu, Jinghong Zheng 0001, Zhengguo Li, Zhijun Fang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2015 | Single image haze removal via a simplified dark channelabstractImages of outdoor scenes could be degraded by haze, fog, and smoke in the atmosphere. In this paper, we propose a novel single image haze removal algorithm by introducing a minimal color channel and a sky region compensation term. A simplified dark channel is computed via the minimal color channel. The transmission map is first estimated by using the simplified dark channel. To avoid amplifying noise in the sky, a non-negative sky region compensation term is proposed to adjust the transmission map in the sky. The map is then refined via a content adaptive guided image filter and is finally applied to recover the haze image. Experimental results on outdoor images with haze and without haze demonstrate that the proposed algorithm outperforms existing algorithms. Zhengguo Li, Jinghong Zheng 0001, Wei Yao 0001 |
ICASSP | 1 |
| 2015 | A hybrid edge-preserving image smoothing scheme for noise removalabstractIn this paper, we propose a new image denoising scheme that is an integration of a content-adaptive guided filter and a collaborative Wiener filter. The proposed scheme consists of two steps. First a content-adaptive guided filter, which smoothes image based on spatial similarity within a local window, is applied. The content-adaptive guided filter can efficiently preserve edges while smoothing noise. A preliminary estimation of noise-free image can be obtained by the content-adaptive guided filter. In the second step, a patch-grouping based collaborative Wiener filter is adopted to exploit non-local similarity, and outputs final denoised image. Compared to the state-of-the-art denoising scheme, BM3D, the proposed method is more efficient in computation. Moreover, simulation results have shown that the proposed method can achieve comparable PSNR values and better visual quality on denoising of textural images. Jinghong Zheng 0001, Zhengguo Li |
ICASSP | 2 |
| 2015 | Noise reduced high dynamic range tone mapping using information content weightsabstractIn this paper, we propose a noise reduced tone mapping method based on information content weights, where the perceptually unimportant pixels are smoothed during the decomposition in two steps. First, a saliency-based information content weight is introduced to give high fidelity to the data term based on the ratio of the local pixel power and the overall noise power in the base layer decomposition. Then, the detail layer is subtracted using the mutual information-based information content weight from the original image luminance and the clean base layer. Experiments show the effectiveness of the proposed method in the improvements of both signal-to-noise ratio and visual quality. Zhengguo Li, Shiqian Wu, Pasi Fränti |
ICASSP | 2 |
| 2015 | Superpixel based patch match for differently exposed images with moving objects and camera movementsabstractA challenging problem for high dynamic range (HDR) imaging is to reconstruct a ghosting-free HDR image for an HDR scene with moving objects. In this paper, a superpixel based patch match algorithm is proposed to synchronize differently exposed images with moving objects and camera movements according to a selected reference image. The concept of superpixel is adopted to divide the reference image into atomic regions, and a new bilateral weight is introduced for the computation of matching cost to reduce outliers. All the input images are synchronized by using the estimated optical flow. The synchronized images are further refined via a weighted guided image filter with the reference image as the guidance image. Experimental results show that the proposed algorithm can remove the scene variations in the presence of moving objects and camera movements. A ghost-free HDR image can then be reconstructed for differently exposed images with moving objects and camera movement by using an existing exposure fusion algorithm. Jinghong Zheng 0001, Zhengguo Li |
ICIP | 2 |
| 2015 | Content Adaptive Image Detail EnhancementabstractDetail enhancement is required by many problems in the fields of image processing and computational photography. Existing detail enhancement algorithms first decompose a source image into a base layer and a detail layer via an edge-preserving smoothing algorithm, and then amplify the detail layer to produce a detail-enhanced image. In this letter, we propose a newL0norm based detail enhancement algorithm which generates the detail-enhanced image directly. The proposed algorithm preserves sharp edges better than an existingL0norm based algorithm. Experimental results show that the proposed algorithm reduces color distortion in the detail-enhanced image, especially around sharp edges. Fei Kou, Weihai Chen, Zhengguo Li, Changyun Wen |
IEEE Signal Process. Lett. | 3 |
| 2015 | Instant Color Matching for Mobile Panorama ImagingabstractThis paper presents an efficient color matching approach to address the photometric inconsistency problem that commonly exists in panoramic images. Color correction, as the first step, is to adjust the color and luminance of source images so that the differences between adjacent images can be minimized. Color blending is used after the color correction to further smooth the color transition between adjacent images. With the first image being selected as a basis image, the proposed approach can start the color matching and stitching process once the second image is captured. The proposed approach is simple and is very suitable for panoramic imaging on mobile devices. Experimental results demonstrate that the color transitions between neighboring images are smooth without visible seams and the color tone of the final image is kept as close as the basis image. Wei Yao 0001, Zhengguo Li |
IEEE Signal Process. Lett. | 2 |
| 2015 | Gradient Domain Guided Image FilteringabstractGuided image filter (GIF) is a well-known local filter for its edge-preserving property and low computational complexity. Unfortunately, the GIF may suffer from halo artifacts, because the local linear model used in the GIF cannot represent the image well near some edges. In this paper, a gradient domain GIF is proposed by incorporating an explicit first-order edge-aware constraint. The edge-aware constraint makes edges be preserved better. To illustrate the efficiency of the proposed filter, the proposed gradient domain GIF is applied for single-image detail enhancement, tone mapping of high dynamic range images and image saliency detection. Both theoretical analysis and experimental results prove that the proposed gradient domain GIF can produce better resultant images, especially near the edges, where halos appear in the original GIF. Fei Kou, Weihai Chen, Changyun Wen, Zhengguo Li |
IEEE Trans. Image Process. | 4 |
| 2015 | Edge-Preserving Decomposition-Based Single Image Haze RemovalabstractSingle image haze removal is under-constrained, because the number of freedoms is larger than the number of observations. In this paper, a novel edge-preserving decomposition-based method is introduced to estimate transmission map for a haze image so as to design a single image haze removal algorithm from the Koschmiedars law without using any prior. In particular, weighted guided image filter is adopted to decompose simplified dark channel of the haze image into a base layer and a detail layer. The transmission map is estimated from the base layer, and it is applied to restore the haze-free image. The experimental results on different types of images, including haze images, underwater images, and normal images without haze, show the performance of the proposed algorithm. Zhengguo Li, Jinghong Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Weighted Guided Image FilteringabstractIt is known that local filtering-based edge preserving smoothing techniques suffer from halo artifacts. In this paper, a weighted guided image filter (WGIF) is introduced by incorporating an edge-aware weighting into an existing guided image filter (GIF) to address the problem. The WGIF inherits advantages of both global and local smoothing filters in the sense that: 1) the complexity of the WGIF is O(N) for an image with N pixels, which is same as the GIF and 2) the WGIF can avoid halo artifacts like the existing global smoothing filters. The WGIF is applied for single image detail enhancement, single image haze removal, and fusion of differently exposed images. Experimental results show that the resultant algorithms produce images with better visual quality and at the same time halo artifacts can be reduced/avoided from appearing in the final images with negligible increment on running times. Zhengguo Li, Jinghong Zheng 0001, Wei Yao 0001, Shiqian Wu |
IEEE Trans. Image Process. | 1 |
| 2014 | Camera noise model-based motion detection and blur removal for low-lighting images with moving objectsabstractIt is well known that modern CCD/CMOS digital cameras produce color images contaminated by mixed photon-electronic noise, which is a mixture of signal-dependent optical photon noise and signal-independent electronic noise. In statistical, variance of the mixed noise is a line function of mean intensity on the pixel. Based on this camera variance-mean model, we propose a fast and robust approach to generate a high quality image from a pair of noisy/blurred low-lighting images with moving objects along any directions. More precisely, camera noise variance model is employed to separate the effects of noise from moving objects on the images, followed by BM3D denoising method to reduce the noise of identified moving objects in the noisy image. Then motion blur in the blurred image is removed by a patching method, which is robust to object movements along any directions. We validate the effectiveness of our proposed approach on real images with moving objects in this paper. Shoulie Xie, Jinghong Zheng 0001, Zhengguo Li |
ICARCV | 3 |
| 2014 | Selectively detail-enhanced exposure fusion via a gradient domain content adaptive bilateral filterabstractBilateral filters suffer from halo artifacts when they are applied for image enhancement. In this paper, a new bilateral filter is proposed in gradient domain to address this problem. Both spatial similarity parameter and intensity similarity parameter of the proposed filter are spatially varying instead of being fixed as in the existing bilateral filters. As a result, it can preserve edges and smooth flat areas better than the existing bilateral filters. The proposed filter is then adopted to design a selectively detail-enhanced exposure fusion algorithm. Fine details of multiple differently exposed images are extracted simultaneously using the proposed filter. Instead of amplifying and adding all extracted fine details to an intermediate image which is fused by an existing exposure fusion algorithm, the fine details in all areas except flat ones are amplified and added to the intermediate image. The resultant algorithm can reduce halo artifacts and prevent noise in flat areas from being amplified in the final image. Therefore, the proposed algorithm fuses images with much better visual quality. Zhengguo Li, Jinghong Zheng 0001 |
ICASSP | 1 |
| 2014 | Content adaptive guided image filteringabstractLocal filtering based edge-preserving smoothing technique could suffer from halo artifacts when it is applied for image enhancement. In this paper, an adaptive guided image filter is proposed by incorporating edge aware weighting which is derived from normalized local variance of a guidance image into an existing guided filter. It is shown that the proposed filter preserves sharp edges better than the existing guided filter. With the observation, it is applied for tone mapping of high dynamic range images and fusion of differently exposed images. Experimental results show that the resultant tone mapping and exposure fusion algorithms can prevent halo artifacts from appearing in the final images. Zhengguo Li, Jinghong Zheng 0001 |
ICME | 1 |
| 2014 | Exposure-Robust Alignment of Differently Exposed ImagesabstractThis letter presents a novel exposure-robust method to align differently exposed images. First, a directional mapping approach is introduced to normalize differently exposed images so as to alleviate the effect of saturation. Then, a non-parametric local binary pattern (LBP) is employed to represent intensity-invariant features of these images. An efficient two-stage alignment is proposed for motion estimation. Experiments on a variety of synthesized and real image sequences demonstrate that the proposed method is less sensitive to the reference image, and robust to 12 exposure values (EV) increments, which is superior to existing methods. Shiqian Wu, Zhengguo Li, Jinghong Zheng 0001 |
IEEE Signal Process. Lett. | 2 |
| 2014 | Selectively Detail-Enhanced Fusion of Differently Exposed Images With Moving ObjectsabstractIn this paper, we introduce an exposure fusion scheme for differently exposed images with moving objects. The proposed scheme comprises a ghost removal algorithm in a low dynamic range domain and a selectively detail-enhanced exposure fusion algorithm. The proposed ghost removal algorithm includes a bidirectional normalization-based method for the detection of nonconsistent pixels and a two-round hybrid method for the correction of nonconsistent pixels. Our detail-enhanced exposure fusion algorithm includes a content adaptive bilateral filter, which extracts fine details from all the corrected images simultaneously in gradient domain. The final image is synthesized by selectively adding the extracted fine details to an intermediate image that is generated by fusing all the corrected images via an existing multiscale algorithm. The proposed exposure fusion algorithm allows fine details to be exaggerated while existing exposure fusion algorithms do not provide such an option. The proposed scheme usually outperforms existing exposure fusion schemes when there are moving objects in real scenes. In addition, the proposed ghost removal algorithm is simpler than existing ghost removal algorithms and is suitable for mobile devices with limited computational resource. Zhengguo Li, Jinghong Zheng 0001, Shiqian Wu |
IEEE Trans. Image Process. | 1 |
| 2013 | Perceptually relevant energy function for seam carvingabstractSeam carving, an image re-targeting method, works by progressively finding and removing connected paths of low energy pixels in an image until a desired image aspect ratio is reached. In this paper, we first cast the problem of minimizing an energy function as that of minimizing a distortion cost. We then leverage on our understanding of image quality metrics/ distortion metrics in proposing a perceptually relevant energy function. Experimental results show that our proposed energy function can generate more desirable resized images in which the original structures of the images are better preserved. Hui Li Tan, Yih Han Tan, Zhengguo Li, Susanto Rahardja, Chuohao Yeo |
ICASSP | 3 |
| 2013 | Residual DPCM for lossless coding in HEVCabstractIncorporating sample-based prediction during lossless coding can significantly improve coding performance. However, its use within a codec designed for lossy coding requires a modification of the available prediction scheme. When implementing the codec, two different prediction processes will have to be implemented. This paper describes a lossless coding scheme that delays the sample-based prediction till the residue coding stage of the codec and carries out prediction in the residual domain. In this way, the prediction scheme of the lossy coder can be retained while realizing the coding gains associated with sample-based prediction. The proposed scheme improves lossless intra coding performance in HEVC Main Profile by an average of 6.5%. Yih Han Tan, Chuohao Yeo, Zhengguo Li |
ICASSP | 3 |
| 2013 | Intra Coding With Adaptive Partial ReconstructionabstractIntra prediction improves coding performance by reducing inter pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra coding modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in the key technical area software when the new modes are incorporated in the current 8$\,\times\,$8 prediction modes. Since the new coding modes (8$\,\times\,$8) are designed with transform size smaller than coding block size, the modes can also be useful when the source block is larger than the maximum transform size. Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Dynamic Range Analysis in High Efficiency Video Coding Residual Coding and ReconstructionabstractWe present a method for analyzing dynamic range along the residual coding and reconstruction pathways during video coding by deriving bounds on the maximum absolute value of intermediary data using a simple combination of triangle inequality and reformulation using Kronecker products. The proposed method is then applied toward analyzing the residual coding and reconstruction process in the emerging High Efficiency Video Coding (HEVC) standard. Our analysis shows that, for an input residual with a bitdepth of (B+1) that uses a uniform quantizer, the dynamic range of quantized levels and dequantized coefficients are no more than (B+7) bits and 17 bits, respectively. Furthermore, a 16 bits transpose buffer is sufficient, while up to 5 bits of bitdepth expansion can occur in the reconstructed residual. The analysis is validated by simulation results with both randomly generated residual and encoding/decoding test video sequences using the HEVC reference software. Chuohao Yeo, Yih Han Tan, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | A Perceptually Relevant MSE-Based Image Quality MetricabstractImage quality metrics (IQMs), such as the mean squared error (MSE) and the structural similarity index (SSIM), are quantitative measures to approximate perceived visual quality. In this paper, through analyzing the relationship between the MSE and the SSIM under an additive noise distortion model, we propose a perceptually relevant MSE-based IQM, MSE-SSIM, which is expressed in terms of the variance of the source image and the MSE between the source and distorted images. Evaluations on three publicly available databases (LIVE, CSIQ, and TID2008) show that the proposed metric, despite requiring less computation, compares favourably in performance to several existing IQMs. In addition, due to its simplicity, MSE-SSIM is amenable for the use in a wide range of image and video tasks that involve solving an optimization problem. As an example, MSE-SSIM is used as the objective function in designing a Wiener filter that aims at optimizing the perceptual visual quality of the output. Experimental results show that the images filtered with a MSE-SSIM-optimal Wiener filter have better visual quality than those filtered with a MSE-optimal Wiener filter. Hui Li Tan, Zhengguo Li, Yih Han Tan, Susanto Rahardja, Chuohao Yeo |
IEEE Trans. Image Process. | 2 |
| 2013 | Hybrid Patching for a Sequence of Differently Exposed Images With Moving ObjectsabstractIt is very challenging to synthesize a high dynamic range (HDR) image from multiple differently exposed low dynamic range images when there are moving objects in the images. This is due to the fact that the moving objects will cause ghosting artifacts to appear in the synthesized HDR image. To prevent such artifacts, a patching algorithm is required to correct motion regions such that all the moving objects are synchronized in the differently exposed images. In this paper, a new optimization problem is formulated to correct the motion regions of the multiple differently exposed images by considering both spatial and temporal consistencies. The resultant scheme is a hybrid patching scheme composed of a correction method which is an intensity mapping function at pixel level, and a hole-filling method that uses block-level template matching. The proposed patching scheme is not only robust to large intensity changes in these input images, but also at regions that are over- or underexposed. Experimental results show that the proposed method is able to prevent ghosting artifacts from appearing in the final synthesized HDR image. Jinghong Zheng 0001, Zhengguo Li, Shiqian Wu, Susanto Rahardja |
IEEE Trans. Image Process. | 2 |
| 2013 | Model-Based Online Learning With KernelsabstractNew optimization models and algorithms for online learning with Kernels (OLK) in classification, regression, and novelty detection are proposed in a reproducing Kernel Hilbert space. Unlike the stochastic gradient descent algorithm, called the naive online Reg minimization algorithm (NORMA), OLK algorithms are obtained by solving a constrained optimization problem based on the proposed models. By exploiting the techniques of the Lagrange dual problem like Vapnik's support vector machine (SVM), the solution of the optimization problem can be obtained iteratively and the iteration process is similar to that of the NORMA. This further strengthens the foundation of OLK and enriches the research area of SVM. We also apply the obtained OLK algorithms to problems in classification, regression, and novelty detection, including real time background substraction, to show their effectiveness. It is illustrated that, based on the experimental results of both classification and regression, the accuracy of OLK algorithms is comparable with traditional SVM-based algorithms, such as SVM and least square SVM (LS-SVM), and with the state-of-the-art algorithms, such as Kernel recursive least square (KRLS) method and projectron method, while it is slightly higher than that of NORMA. On the other hand, the computational cost of the OLK algorithm is comparable with or slightly lower than existing online methods, such as above mentioned NORMA, KRLS, and projectron methods, but much lower than that of SVM-based algorithms. In addition, different from SVM and LS-SVM, it is possible for OLK algorithms to be applied to non-stationary problems. Also, the applicability of OLK in novelty detection is illustrated by simulation results. Guoqi Li 0002, Changyun Wen, Zhengguo Li, Feng Yang 0011, Kezhi Mao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2012 | A local intensity adaptive structural similarity indexabstractExisting structural similarity (SSIM) index comprises of one term on luminance comparison and the other term on contrast and structure comparison. In this paper, the SSIM index is first improved by introducing three weighting factors to the second term such that it is adaptive to local intensities of two images to be compared. The improved SSIM (iSSIM) index is further extended for two images with possibly different exposures. Experimental results show that the proposed indices are more robust to large intensity changes of two images from the same scene and more sensitive to two images from different scenes than the existing SSIM index. Zhengguo Li, Chuohao Yeo, Yih Han Tan, Susanto Rahardja |
ICASSP | 1 |
| 2012 | A bilateral filter in gradient domainabstractIn this paper, a bilateral filter in gradient domain is first proposed. It is then applied to study detail enhancement via multi-light images and noise reduction of differently exposed low dynamic range images. These two applications show that the proposed filter can be applied to extract fine details from a set of images simultaneously and to provide flexibility for noise reduction from selected areas of an image. Zhengguo Li, Jinghong Zheng 0001, Shiqian Wu, Susanto Rahardja |
ICASSP | 1 |
| 2012 | Noise reduction for differently exposed imagesabstractFor scenes under low lighting condition, cameras are usually set to a high sensitivity (ISO) mode to reduce motion blur at the cost of increased image noise. When multiple differently exposed images are used to generate a high dynamic range (HDR) image, the high ISO noise from each low dynamic range (LDR) image can be further amplified by the HDR synthesis algorithm which would result in severely degradation of visual quality. This paper proposes an intensity mapping function based noise reduction method for differently exposed images with high ISO noise. The proposed method does not require any knowledge on either camera response functions or exposure times. In addition, the method is simple yet effective for noise removal from the LDR images without introducing any blurring or other artifacts. Wei Yao 0001, Zhengguo Li, Susanto Rahardja |
ICASSP | 2 |
| 2012 | Anti-ghost of differently exposed images with moving objectsabstractIn a typical image synthesis where multiple differently exposed images are captured for processing, it is important to design an anti-ghost algorithm so as to prevent ghosting artifacts from appearing in the final image. An anti-ghost algorithm is usually composed of a detection module and a correction module. In this paper, a new detection module is proposed to detect non-consistent pixels of all input images without predefining any initial reference image. The proposed module is suitable when an interactive mode is desired. In addition, a bidirectional approach is introduced to correct the non-consistent pixels in the correction module. Compared with existing unidirectional correction methods, the proposed bidirectional correction approach uses information from two adjacent images of a detected image to correct its non-consistent pixels. This leads to a quality improvement in the final image. Zhengguo Li, Shiqian Wu, Shoulie Xie, Susanto Rahardja |
ICIP | 1 |
| 2012 | Single-Pass Rate Control With Texture and Non-Texture Rate-Distortion ModelsabstractOne of the challenges in video rate control lies in determining a quantization parameter (Qp) that will be used for both the rate-distortion (R-D) optimization process and the quantization of transform coefficients. In this paper, we attempt to achieve effective rate control with a different approach. By modeling the relationships of distortion, texture bits, non-texture bits, and Qp, we derive the Qp required for both R-D optimization and quantization through Lagrangian optimization. From experiments with several video sequences, we found that our rate control scheme is capable of effective rate control with only a few model updates during encoding. The proposed rate control scheme adapts quickly to the characteristics of the source data and is particularly effective at controlling the rate of videos with high and unpredictable motion content. Yih Han Tan, Chuohao Yeo, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Mode-Dependent Transforms for Coding Directional Intra Prediction ResidualsabstractThe use of mode-dependent transforms for coding directional intra prediction residuals has been previously shown to provide coding gains, but the transform matrices have to be derived from training. In this paper, we derive a set of separable mode-dependent transforms by using a simple separable, directional, and anisotropic image correlation model. Our analysis shows that only one additional transform, the odd type-3 discrete sine transform (ODST-3), is required for the optimal implementation of mode-dependent transforms. In addition, the four-point ODST-3 also has a structure that can be exploited to reduce the operation count of the transform operation. Experimental results show that in terms of coding efficiency, our proposed approach matches or improves upon the performance of a mode-dependent transforms approach that uses transform matrices obtained through training. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Detail-Enhanced Exposure FusionabstractIn a typical processing chain of image enhancement, an exposure fusion scheme can be used to synthesize a more detailed low dynamic range (LDR) image directly from a set of differently exposed LDR images, without generation of an intermediate high dynamic range image. In this brief, we introduce a new quadratic optimization-based method to extract fine details from a vector field. The new method extracts fine details from a set of differently exposed LDR images simultaneously. The extracted fine details are then added to an intermediate LDR image which is fused by simply using an existing exposure fusion scheme. With this, the proposed scheme can enhance fine details to produce sharper images. Zhengguo Li, Jinghong Zheng 0001, Susanto Rahardja |
IEEE Trans. Image Process. | 1 |
| 2011 | Structural similarity indices for high dynamic range imagingabstractIn this paper, a structural similarity index is first proposed for two images with possibly different dynamic ranges and intensities as well as possibly small rotation and translation. The proposed index is then extended by dividing two images into local windows, and the similarity is detected by checking all pairs of local windows. It is shown by experimental results that the proposed indices are more robust to large intensity and dynamic range changes of two images from the same scene than the structural similarity (SSIM) index in. Zhengguo Li, Susanto Rahardja |
ICASSP | 1 |
| 2011 | Quadratic optimization based small scale details extractionabstractIn many image processing problems, it is required to extract small scale details from an image or a set of images. In this paper, we introduce a new framework for extracting small scale details from a single input image or a set of input images. We then show how to apply the framework to address several important problems in the field of image processing including tone mapping of high dynamic range images, de-noising of a non-flash image with a pair of non flash and flash images, as well as details enhancement via multi light images and a single input image. Experimental results show that the proposed framework outperforms existing methods. Zhengguo Li, Jinghong Zheng 0001, Chuohao Yeo, Susanto Rahardja |
ICASSP | 1 |
| 2011 | Fast movement detection for high dynamic range imagingabstractWhen a high dynamic range image is synthesized by using a set of differently exposed low dynamic range (LDR) images, it is important to detect moving objects so as to remove ghosting from the final HDR image. A pixel level movement detection scheme was recently proposed in [8]. It included a pixel level similarity index for differently exposed LDR images, an adaptive threshold for the classification of pixels and an approach that utilizes intensity mapping (IMF) function for patching invalid regions. In this paper, we first propose a new adaptive threshold and a new patching approach to improve the scheme in [8]. Then, a sub-sampling method is introduced to simplify the improved movement detection scheme. Experimental results show that the improved movement detection scheme indeed outperforms the scheme in [8]. In addition, the speed is significantly improved by the proposed fast movement detection scheme. Zhengguo Li, Susanto Rahardja |
ICIP | 1 |
| 2011 | Intra-prediction with adaptive sub-samplingabstractIntra-prediction improves coding performance by reducing inter-pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra-prediction modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in KTA when the new modes are incorporated in the current 8×8 prediction modes. Since the new coding modes (8×8) are designed with transform size smaller than coding block size, the modes can also be useful when the prediction unit is larger than the maximum transform size. The use of smaller transform sizes can potentially lead to reduction of decoder complexity and implementation costs. Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja |
ICIP | 3 |
| 2011 | Low-complexity mode-dependent KLT for block-based intra codingabstractApplying mode-dependent separable transforms, e.g., mode- dependent directional transform (MDDT), is an effective method for improving transform coding of intra prediction residuals. However, two transform matrices typically need to be stored for each intra prediction mode. By using a simple image correlation mode, we have previously derived and proposed a simplified mode-dependent separable transforms scheme that uses a combination of two well-known trans- forms: Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST). In this paper, we propose an orthogonal 4-point integer DST that has a multiplier-less implementation consisting of only adds and bit-shifts. We also propose a simple set of mode-dependent scans for coefficient coding that can be used on top of mode-dependent transforms. Our experimental results on the current HEVC reference software show that in terms of coding efficiency, our proposed approach has comparable performance to MDDT. More importantly, compared to MDDT, our approach requires no training and has lower computational and storage costs. Chuohao Yeo, Yih Han Tan, Zhengguo Li |
ICIP | 3 |
| 2011 | Chroma intra prediction using template matching with reconstructed luma componentsabstractIntra coding in the current H.264/AVC video coding standard achieves high compression efficiency, in part due to the highly effective intra prediction process that exploits spatial directional correlation. However, intra prediction of chroma components in YUV 4:2:0 videos uses a limited set of possible predictions available for coding of luma components. Furthermore, coding of chroma components proceeds somewhat independently of luma components, and ignores any possible correlation between them. In this paper, we show a way of using reconstructed luma pixels to help with intra prediction of chroma pixels. By making use of the reconstructed co-located luma block to perform template matching in the luma plane, we are able to use as predictors the co-located chroma blocks of the matched luma blocks. Simulations results indicate that the proposed approach is able to obtain up to 33% chroma bit-rate reduction and up to 8% overall bit-rate reduction over H.264/AVC. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
ICIP | 3 |
| 2011 | De-ghosting of HDR images with double-credit intensity mappingabstractGhosting artifacts are usually caused by moving object when composing a high dynamic range image from multiple differently exposed conventional images. In this paper, a robust de-ghosting algorithm is proposed based on a double-credit intensity mapping function (IMF) and an adaptive threshold model derived from statistical training. The double-credit IMF is estimated using both pixel intensity distribution and spatial correlation. A statistical threshold model is trained from the image database, and the key parameters are determined on the fly with variance vector calculated during the IMF estimation to adapt to different scenarios. Optimal bidirectional comparison is used for further improves the detection accuracy. The experiments show the effectiveness of the proposed de-ghosting method. Zhengguo Li, Susanto Rahardja, Pasi Fränti |
ICIP | 2 |
| 2011 | Mode-dependent fast separable KLT for block-based intra codingabstractIn this paper, we derive separable KLTs for coding H.264/AVC intra prediction residuals, using a simple image correlation model. Our analysis shows that for some intra prediction modes, we can in fact just use the DCT for performing either the row-wise or column-wise transform. Furthermore, we also compute the KLT that should be used based on the image correlation model, which happens to have sinosuidal terms. The 4×4 transform also has a structure that can be exploited to reduce the operation count of the transform operation. In our simplified implementation of mode-dependent directional transforms (MDDT), we only need to make use of two matrices: the DCT and the derived KLT. Our experimental results show that in terms of coding efficiency, our proposed approach has similar performance when compared with MDDT. More importantly, compared to MDDT, our approach requires no training and has lower computational and storage costs. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
ISCAS | 3 |
| 2011 | On residual quad-tree coding in HEVCabstractIn the current working draft of HEVC, residual quad-tree (RQT) coding is used to encode prediction residuals in both Intra and Inter coding units (CU). However, the rationale for using RQT as a coding tool is different in the two cases. For Intra prediction units, RQT provides an efficient syntax for coding a number of sub-blocks with the same intra prediction mode. For Inter CUs, RQT adapts to the spatial-frequency variations of the CU, using as large a transform size as possible while catering to local variations in residual statistics. While providing coding gains, effective use of RQT currently requires an exhaustive search of all possible combinations of transform sizes within a block. In this paper, we exploit our insights to develop two fast RQT algorithms, each designed to meet the needs of Intra and Inter prediction residual coding. Yih Han Tan, Chuohao Yeo, Hui Li Tan, Zhengguo Li |
MMSP | 4 |
| 2010 | Robust generation of high dynamic range imagesabstractA robust scheme is proposed to generate an anti-ghosting high dynamic range (HDR) image from a set of low dynamic range (LDR) images with different exposure times. Three major contributions of this paper are 1) a bi-directional prediction method; 2) an adaptive threshold for the classification of pixels; 3) Bayes estimator based methods for the on-line updating of predicted values and the synthesis of pixels to fill in the regions of moving objects to preserve their dynamic ranges. The proposed scheme is suitable for both static and dynamic scenes. Zhengguo Li, Shoulie Xie, Shiqian Wu, Susanto Rahardja |
ICASSP | 1 |
| 2010 | Half-quadratic regularization based de-noising for high dynamic range image synthesisabstractIt is possible to synthesis a high dynamic range (HDR) image by combining differently exposed low dynamic range (LDR) images of the same scene into one single image. For an HDR scene under low light condition, captured LDR images tend to be noisy, and the noise could usually be amplified during the synthesis process, causing severe degradation of the image quality. In order to reduce the noise during the HDR synthesis process, a de-noising scheme based on half-quadratic regularization is proposed for the synthesis of HDR images in this paper. By taking into consideration of two unique HDR image features, the proposed scheme effectively removes the noise, while the edges still being well preserved in the synthesized HDR images. Wei Yao 0001, Zhengguo Li, Susanto Rahardja, Susu Yao, Jinghong Zheng 0001 |
ICASSP | 2 |
| 2010 | Collaborative image processing algorithm for detail refinement and enhancement via multi-light imagesabstractIn this paper, we introduce a new collaborative image processing algorithm to enhance contours and surface details of a scene through combination of multi-light images, which capture the same scene with fixed view point but different lighting conditions. Firstly, a detail layer that contains all contents of the multi-light images is generated through a gradient domain method and a quadratic filter. A new shadow detection algorithm is introduced to remove the artifacts from the detail layer. Subsequently, a base layer is constructed by using one input image. To further increase the visibility of details in dark areas, a simple tool is presented to brighten all dark areas of the base layer. Finally, the detail layer is multiplied to the base layer to synthesize the enhanced image that contains the desirable details. This enhanced image can present an elaborate description of the real scene. The proposed scheme also gives some interactivities that allow users to easily change the appearance of the enhanced image according to their preferences. Jinghong Zheng 0001, Zhengguo Li, Susanto Rahardja, Susu Yao, Wei Yao 0001 |
ICASSP | 2 |
| 2010 | Movement detection for the synthesis of high dynamic range imagesabstractIn this paper, we propose an intensity mapping function (IMF) based scheme to detect moving objects in a set of low dynamic range (LDR) images with different known exposure times. The objective is to remove ghosting artifacts from the eventual high dynamic range (HDR) image. Our contributions include a bidirectional similarity detection method, an adaptive threshold for movement detection, and an IMF based method for the synthesis of pixels to fill in the regions of moving objects. Experimental results show that the proposed scheme outperforms existing schemes. Zhengguo Li, Susanto Rahardja, Shoulie Xie, Shiqian Wu |
ICIP | 1 |
| 2010 | A robust and fast anti-ghosting algorithm for high dynamic range imagingabstractThis paper presents a robust and fast algorithm for automatically generating high dynamic range (HDR) images in presence of camera movement and moving objects. This scheme comprises five modules: 1) image alignment, 2) estimation of camera response function (CRF) in dynamic scenes, 3) moving object detection, 4) progressive image correction, and 5) construction of HDR images. The key advantage of the algorithm is the ability to generate HDR images without ghost artifact. The proposed algorithm is fast as it is a one-shot solution without iterative computation and post-processing or even manual operation. Experimental results demonstrate that the proposed method outperforms the existing commercial products. Shiqian Wu, Shoulie Xie, Susanto Rahardja, Zhengguo Li |
ICIP | 4 |
| 2010 | Detecting and composing near-identical HDR images without exposure informationabstractIn high dynamic range (HDR) imaging, two essential problems are to compose HDR image from conventional image set without any prior information about their exposures, and to access the synthesis result. To solve these problems, we first develop an exposure ratio estimation algorithm based on intensity mapping function (IMF). Then, we introduce an HDR image comparison method to verify whether two HDR images are from the same scene by using their log histogram similarity. Even though the images carrying the same information, their similarity cannot be detected by pixel-wise comparisons. We name such a pair of HDR images as near-identical images. According to experiments, our detection method is able to identify near-identical HDR images effectively, and our synthesis algorithm is able to recover the correct exposure ratios and compose near-identical HDR images. Susanto Rahardja, Zhengguo Li, Pasi Fränti |
ICIP | 3 |
| 2009 | High dynamic range compression by half quadratic regularizationabstractThis paper presents a new adaptive tone mapping for high dynamic range (HDR) images. In the proposed scheme, the luminance of an HDR image is decomposed into a base layer with large gradients and a detail layer with small gradients by using an adaptive half quadratic regularization method. The base layer is compressed by a novel global mapping to reduce the dynamic range while the detail layer can be amplified to enhance the local contrasts. With the proposed scheme, the generated low dynamic range (LDR) images look more natural and local details are also preserved very well. Zhengguo Li, Susanto Rahardja, Susu Yao, Jinghong Zheng 0001, Wei Yao 0001 |
ICIP | 1 |
| 2008 | Early detection of all-zero block in H.264 with new rate-quantization modelsabstractThis paper presents a new algorithm for detecting all-zero DCT coefficient blocks (AZB) prior to DCT and quantization for H.264. The early detection criterion is derived from a new rate-quantization model which is established by considering the unique features of quantization process in H.264. The proposed algorithm aims to eliminate the redundant computations in AZB and consequently speed up the encoding process of H.264 video codec. Simulation results show that the proposed algorithm achieves a high detection ratio up to 95.88% while maintain a very low false detection ratio. The results confirm that a more effective AZB detection for H.264 can be achieved by taking the unique feature of quantization into consideration. Wei Yao 0001, Zhengguo Li, Susanto Rahardja |
ISCAS | 2 |
| 2008 | Rate Control of H.264/AVC Scalable ExtensionabstractThis paper presents a rate control scheme for H.264/AVC scalable extension. Based on our previous work on H.264/AVC rate control, a switched model is proposed to predict the mean absolute difference (MAD) of the residual texture from the available MAD information of the previous frame in the same layer and the same frame in its ldquobase layer.rdquo Thus, abrupt MAD fluctuations could be predicted properly in the enhancement layer. Moreover, a bit allocation scheme is proposed for the hierarchical B-frames structure by taking into consideration the relative importance of each frame. With our algorithm, the rate control for all the coarse-grain-scalability, spatial, temporal and combined enhancement layer could be realized, and the target bit rate for each layer can be achieved. Our method encodes the sequence only once and the buffer is well controlled to prevent it from overflowing and under flowing. Yang Liu 0028, Zhengguo Li, Yeng Chai Soh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Region-of-Interest Based Resource Allocation for Conversational Video Communication of H.264/AVCabstractDue to the complexity of H.264/AVC, it is very challenging to apply this standard to design a conversational video communication system. This problem is addressed in this paper by using region-of-interest (ROI) based bit allocation and computational power allocation schemes. In our system, the ROI is first detected by using the direct frame difference and skin-tone information. Several coding parameters including quantization parameter, candidates for mode decision, the number of referencing frames, accuracy of motion vectors and the search range of motion estimation are adaptively adjusted at the macroblock (MB) level according to the relative importance of each MB. Subsequently, the encoder could allocate more resources such as bits and computational power to the ROI, and the decoding complexity is also optimized at the encoder side by utilizing an ROI based rate-distortion-complexity (R-D-C) cost function. The encoder is thus simplified and decoding-friendly, and the overall subjective visual quality can also be improved. Yang Liu 0028, Zhengguo Li, Yeng Chai Soh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Simplified Motion-Refined Scheme for Fine-Granularity ScalabilityabstractIn this paper, we introduce a low-complexity fine-granularity scalable (FGS) video encoder that refines both residue and motion information in the quality layers. The current scalable video coding (SVC) draft shows that significant gains can be achieved when each enhancement layer undergoes the motion estimation/motion compensation (ME/MC) process with its own motion vector field (MVF). However, given the high computational cost of ME/MC, a motion-refined FGS scheme can be expensive to implement. The proposed scheme controls the macroblock (MB) mode allowed in the base layer and channels computational resources to refine motion in enhancement layers. Through a proper selection of Lagrangian factor for the generation of the first MVF, it is possible to design a low-complexity FGS encoder that has good overall coding performance. A simplified motion-refinement scheme is also adopted for selected MBs in enhancement layers by exploiting the correlation of MB-type information between successive layers to further reduce the complexity of the FGS encoder. Meanwhile, a framework of rate-distortion-complexity optimization is proposed for the FGS by considering the interpolation complexity during the ME/MC in each layer. The FGS decoder can be simplified through the reduction in the number of interpolations. Yih Han Tan, Zhengguo Li, Keng-Pang Lim, Susanto Rahardja |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Adaptive Decoder Complexity Reduction for Coarse Granular ScalabilityabstractThe on-going scalable video coding (SVC) standard is an extension of H.264/AVC. It enables significantly improved compression performance at the expense of greatly increased computational complexity at both the encoder and the decoder sides. This paper presents an adaptive algorithm to reduce the complexity of decoder at the encoder side for coarse granular scalability. Hence, lightweight bitstreams are generated at the encoder which requires significantly less decoding complexity. The experimental results show that the proposed scheme provides significant reduction in the complexity of decoder with acceptable coding loss and minor impact on the encoder complexity. Zhengguo Li, Changyun Wen |
ICASSP (1) | 2 |
| 2007 | Fast Mode Decision for Coarse Granular Scalability via Switched Candidate Mode SetabstractThis paper presents an improved fast mode decision algorithm for coarse grain scalability (CGS). The modified fast mode decision algorithm extends previous work to provide better encoder complexity reduction with insignificant degradation in picture quality. The candidate mode set is adaptive to the quantization parameter difference between a layer and its "base layer". Furthermore, the proposed scheme fully utilizes the statistics of mode transition between a layer and its "base layer" when the quantization parameter difference between these two layers is small. Simulation results demonstrate that the proposed scheme provides up to 64% of time saving compared with the original SVC encoder. Zhengguo Li, Changyun Wen, Shoulie Xie |
ICME | 2 |
| 2007 | Wyner-Ziv Image Coding from Random ProjectionsabstractIn this paper, we present a Wyner-Ziv coding based on random projections for image compression with side information at the decoder. The proposed coder consists of random projections (RPs), nested scalar quantization (NSQ), and Slepian-Wolf coding (SWC). Most of natural images are compressible or sparse in the sense that they are well-approximated by a linear combination of a few coefficients taken from a known basis, e.g., FFT or Wavelet basis. Recent results show that it is surprisingly possible to reconstruct compressible signal to within very high accuracy from limited random projections by solving a simple convex optimization program. Nested quantization provides a practical scheme for lossy source coding with side information at the decoder to achieve further compression. SWC is lossless source coding with side information at the decoder. In this paper, ideal SWC is assumed, thus rates are conditional entropies of NSQ quantization indices. Recently theoretical analysis shows that for the quadratic Gaussian case and at high rate, NSQ with ideal SWC performs the same as conventional entropy-coded quantization with side information available at both the encoder and decoder. We note that the measurements of random projects for a natural large-size image can behave like Gaussian random variables because most of random measurement matrices behave like Gaussian ones if their sizes are large. Hence, by combining random projections with NSQ and SWC, the tradeoff between compression rate and distortion will be improved. Simulation results support the proposed joint codec design and demonstrate considerable performance of the proposed compression systems. Shoulie Xie, Susanto Rahardja, Zhengguo Li |
ICME | 3 |
| 2007 | Rate Control for Spatial/CGS Scalable Extension of H.264/AVCabstractThis paper presents a rate control scheme for H.264/AVC spatial/ coarse-gain-SNR (CGS) scalable extension. A switched model is proposed to predict mean absolute difference (MAD) either from the previous frame of the current layer or from the current frame of the previous layer. This way, abrupt MAD fluctuations could be predicted properly in the enhancement layer. Moreover, a sum bits R-Q model is formulated to describe the relationship between the total amount of bits for texture and non-texture information and quantization parameter (QP) so as to reduce the negative effect caused by inaccurate non-texture bits estimation of the previous rate control schemes. With the relation between PSNR and QP value of H.264/AVC, our proposed sum bits R-Q model could further optimize the QP calculation at the MB level. With our algorithm, the rate control for all the coarse-grain-scalability (CGS) and spatial enhancement layer could be realized, and the encoder could achieve fixed bitrate encoding for each scalable layer. Our method encodes the sequence only once and the buffer is well controlled to prevent it from overflowing and underflowing. Yang Liu 0028, Yeng Chai Soh, Zhengguo Li |
ISCAS | 3 |
| 2007 | An OWE-based Algorithm for Line Scratches Restoration in Old MoviesabstractLine scratch is the primary artifact in old films. In this paper a new algorithm for line scratch detection and removal is proposed. First, we establish an effective model to represent line scratches in the domain of over-complete wavelet expansion (OWE), which offers more precise position description for scratches than traditional downsampled wavelet transform. We then use it to detect and locate line scratches accurately. After that, an adaptive restoration method is adopted to remove line scratches by replacing the wavelet coefficients in line scratches of each scale with new interpolated wavelet coefficients of corresponding width computed from the line scratch model. Experiments show that the proposed method can detect more line artifacts with less false detection and remove the line scratches effectively as compared with the classic algorithm. Jinghuo Guan, Jun Sun 0005, Guangtao Zhai, Zhengguo Li |
ISCAS | 6 |
| 2007 | Image Quality Assessment using Foveated Wavelet Error Sensitivity and Isotropic ContrastabstractIn this paper, we propose a new image quality evaluation method, which is based on foveation spatial error sensitivity model and wavelet error sensitivity model. The image quality is measured by isotropic contrast and the standard deviation of the errors of the wavelet coefficients in subbands with the weighting factors that are determined by two models. Experiments on a test database comprising JPEG and JPEG 2000 compressed images have shown that the proposed metric can achieve very good correlation with subjective evaluation. Susu Yao, Weisi Lin, Ee Ping Ong, Zhongkang Lu, Mei Hwan Loke, Zhengguo Li |
ISCAS | 6 |
| 2007 | Balanced Inter-Layer Prediction for Combined Coarse Granular Scalability and Spatial ScalabilityabstractAn inter-layer prediction scheme with two base layers and a new concept of auxiliary layer are introduced in this paper. The scheme achieves a good balance among all layers for the combined coarse granular scalability and spatial scalability. The objective is to improve the coding efficiency of layers with higher resolution. Meanwhile, a price based scheme is proposed to determine the auxiliary layer. With the proposed scheme, a new element of price can be integrated into scalable video coding which in turn justifies the necessity of scalable coding Wei Yao 0001, Zhengguo Li, Susanto Rahardja |
ISCAS | 2 |
| 2007 | Analysis of Monotonic Responsive Functions for Congestion Control
Zhengguo Li, Yeng Chai Soh |
MMM (2) | 2 |
| 2007 | Motion refined medium granular scalabilityabstractIn this paper, we propose an interesting scheme to obtain a good tradeoff between motion information and residual information for medium granular scalability (MGS). In this scheme, both motion information and residual information are refined at enhancement layers when the scalable bit rate range is wide, whereas only residual information is refined when the range is narrow. In other words, for the case of wide bit rate range, there can be more than one motion vector fields (MVFs) where one is generated at base layer and others are generated at enhancement layers. When it is narrow, only one MVF is necessary. The layers can either share one MVF or have its own, depending on the bit rate range cross layers. Unlike Coarse Granular Scalability (CGS), the correlation between two adjacent MVFs in MGS is very strong. Hence MGS can be provided in the most important bit rate range to achieve a better tradeoff between motion and residual information and a finer granularity in that range. CGS can be applied in less important bit rate ranges to give a coarse granularity. Experimental results show that the coding efficiency can be improved by up to 1dB compared with existing SNR scalability scheme at high bit rate. Zhengguo Li, Wei Yao 0001, Susanto Rahardja |
VCIP | 1 |
| 2007 | A Novel Rate Control Scheme for Low Delay Video Communication of H.264/AVC StandardabstractThis paper presents a novel rate control scheme for low delay video communication of H.264/AVC standard. A switched mean-absolute-difference (MAD) prediction scheme is introduced to enhance the traditional temporal MAD prediction model, which is not suitable for predicting abrupt MAD fluctuations. Our new model could reduce the MAD prediction error by up to 69%. Furthermore, an accurate linear rate-quantization (R-Q) model is also formulated to describe the relationship between the total amount of bits for both texture and nontexture information and the quantization parameter (QP), so that the negative effect caused by the inaccurate estimation of nontexture bits is removed. By exploring the relationship between peak signal-to-noise ratio and QP value, the proposed linear R-Q model could further optimize QP calculation at the macroblock level. When compared with the rate control scheme JVT-G012 which is adopted by the latest JVT H.264/AVC reference model JM9.8, the proposed rate control algorithm could reduce the mismatch between actual bits and target ones by up to 75%. To meet the low delay requirement, the buffer is better controlled to prevent overflowing and underflowing. The average luminance PSNR of reconstructed video is increased by up to 1.13 dB at low bit rates, and the subjective video quality is also improved Yang Liu 0028, Zhengguo Li, Yeng Chai Soh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | MRF: a framework for source and destination based bandwidth differentiation service
Zhengguo Li, Yeng Chai Soh |
IEEE/ACM Trans. Netw. | 2 |
| 2006 | A Practical Chaotic Secure Communication Scheme Based on Chen's SystemabstractIn this paper, a practical impulsive synchronization scheme is proposed for the synchronization of two Chen's chaotic systems with parametric uncertainty and mismatch. With this scheme, the error of the systems can converge to a given bound Chengjin Zhang, Changyun Wen, Zhengguo Li |
ICARCV | 4 |
| 2006 | Fast Mode Decision for Coarse Grain SNR Scalable Video CodingabstractScalable Video Coding (SVC) is an on-going standard and the current working draft (WD) is an extension of H.264/AVC. In the WD, exhaustive search technique is employed to select the best coding mode for each macroblock (MB). This technique achieves highest possible coding efficiency, but it results in higher computational complexity. To overcome this, we propose a novel fast mode decision scheme for coarse grain SNR scalability (CGS) in SVC. In this scheme, the mode distribution relationships between the base layer and enhancement layers are employed to reduce the candidate mode set at enhancement layers. The experimental results show that the proposed scheme provides significant reduction in computational complexity with negligible coding loss. Zhengguo Li, Changyun Wen |
ICASSP (2) | 2 |
| 2006 | Adaptive Mad Prediction and Refined R-Q Model for H.264/AVC Rate ControlabstractThis paper presents an improved rate control scheme for the H.264/AVC video coding scheme. By analyzing the relationship between direct mean absolute difference (MAD) and actual MAD, a new MAD prediction scheme is introduced to enhance traditional linear MAD prediction model, which is unable to predict abrupt MAD fluctuations. Our proposed adaptive model could reduce MAD prediction error by up to 34%. One simple sum bit quadratic R-Q model is also presented to solve the problem caused by inaccurate texture bits estimation of H.264/AVC. With the new MAD prediction model and R-Q model, our proposed scheme could reduce the mismatch between actual frame bits and target frame bits by up to 32%, and the buffer occupancy is much closer to the ideal status. Meanwhile, reconstructed video quality is also improved by up to 0.21 dB at low bitrate Yang Liu 0028, Zhengguo Li, Yeng Chai Soh |
ICASSP (2) | 2 |
| 2006 | Conversational Video Communication of H.264/AVC with Region-of-Interest ConcernabstractAlthough region-of-interest (ROI) based video coding has been well studied for some other video coding standards, its application for H.264 is still of significant interest because there exists a dilemma in the detection of the ROI due to the rate distortion optimization. In this paper, the ROI is first detected by using the direct MAD that is determined without motion information. In this way, the ROI-motion dilemma of H.264/AVC is solved. The rate control scheme with ROI concern can then adjust the quantization parameter (QP) to allocate more bits to the ROI, so the overall subjective visual quality is improved. Yang Liu 0028, Zhengguo Li, Yeng Chai Soh, Mei Hwan Loke |
ICIP | 2 |
| 2006 | Fast mode decision for spatial scalable video codingabstractScalable video coding (SVC) is an on-going standard and the current working draft (WD) is an extension of H.264/AVC. It provides scalability at the bit stream level with good compression efficiency, and allowing free combinations of spatial, temporal and SNR scalability. In the WD, exhaustive search technique is employed to select the best coding mode for each macroblock. This technique achieves highest possible coding efficiency, but it results in higher computational complexity. To overcome this, we propose a novel fast mode decision scheme for spatial scalability in SVC. In this scheme, the mode distribution relationship between base layer and enhancement layers is employed to reduce the candidate mode set at enhancement layers. The experimental results show that the proposed scheme provides significant reduction in computational complexity without any noticeable coding loss. Zhengguo Li, Changyun Wen, Lap-Pui Chau |
ISCAS | 2 |
| 2006 | Adaptive rate control for H.264
Zhengguo Li, Wen Gao 0001, Feng Pan 0002, S. W. Ma, Keng-Pang Lim, G. N. Feng, Xiao Lin 0001, Susanto Rahardja, H. Q. Lu, Yan Lu 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2006 | Fast Mode Decision Algorithm for Inter-Frame Coding in Fully Scalable Video CodingabstractScalable video coding is an ongoing standard, and the current working draft (WD) is an extension of H.264/AVC. In the WD, an exhaustive search technique is employed to select the best coding mode for each macroblock. This technique achieves the highest possible coding efficiency, but it results in extremely large encoding time which obstructs it from practical use. This paper proposes a fast mode decision algorithm for inter-frame coding for spatial, coarse grain signal-to-noise ratio, and temporal scalability. It makes use of the mode-distribution correlation between the base layer and enhancement layers. Specifically, after the exhaustive search technique is performed at the base layer, the candidate modes for enhancement layers can be reduced to a small number based on the correlation. Experimental results show that the fast mode decision scheme reduces the computational complexity significantly with negligible coding loss and bit-rate increases Zhengguo Li, Changyun Wen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Implicit Bit Allocation for Combined Coarse Granular Scalability and Spatial ScalabilityabstractA new type of implicit bit allocation (IBA) is studied for the combined coarse granular scalability (CGS) and spatial scalability. Given a region which is a quality level at a spatial resolution and an input of weighting factor in the region that is determined by customers' interest, the IBA is formulated as a multiple objective optimization problem. The IBA exhibits a distinguished feature that allows bits allocation to each region being fixed and only tradeoff between motion information and residual information in each region can be properly set such that coding efficiency of each region is guaranteed in order according to the weighting factor. Due to the nonlinearity of the combined CGS and spatial scalability, this optimization problem is very complex. In this paper, a simple solution to the optimization problem is proposed by using the conventional approach of "Divide and Conquer" or "Hierarchy." Two cross-layer motion estimation/motion compensation (ME/MC) schemes that were introduced for the CGS and spatial scalability are further adopted to support the solution. It is shown that a combined CGS and spatial scalability scheme together with adaptability to customer composition allows the solution to achieve a customer oriented scalable tradeoff (COST) Zhengguo Li, Susanto Rahardja, Hanwu Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | A novel SNR refinement scheme for scalable video codingabstractCross SNR layer motion estimation/motion compensation (ME/MC) scheme is proposed in this paper for the SNR scalability of scalable video coding. Both the motion information and residual information are refined at the enhancement layers by our cross SNR layer ME/MC scheme. A good trade-off between motion information and residual information can be obtained at all bit rates. The coding efficiency is improved by up to more than 1 dB at high bit rate. Zhengguo Li, Yang Liu 0028, Yeng Chai Soh |
ICIP (3) | 1 |
| 2005 | Customer Adaptive Combined SNR and Spatial ScalabilityabstractA new type of region of interest (ROI) is proposed in this paper for scalable video coding with the whole region be the combined SNR and spatial scalability, and a sub-region be a specific choice of spatial resolution and bit rate range. The ROI is applied to design a combined SNR and spatial scalability scheme that is adaptive to customer composition such that an optimal customer oriented scalable tradeoff (COST) can be achieved. The profit can thus be maximized Zhengguo Li, Susanto Rahardja, Xiao Lin 0001, Wei Yao 0001 |
MMSP | 1 |
| 2005 | Geometrically determining the leaky bucket parameters for video streaming over constant bit-rate channels
Ping Li 0002, Weisi Lin, Susanto Rahardja, Xiao Lin 0001, Xiaokang Yang 0001, Zhengguo Li |
Signal Process. Image Commun. | 6 |
| 2005 | Fast mode decision algorithm for intraprediction in H.264/AVC video codingabstractThe H.264/AVC video coding standard aims to enable significantly improved compression performance compared to all existing video coding standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode and reference frame for each macroblock. As a result, the complexity and computation load increase drastically. This paper presents a fast mode decision algorithm for H.264/AVC intraprediction based on local edge information. Prior to intraprediction, an edge map is created and a local edge direction histogram is then established for each subblock. Based on the distribution of the edge direction histogram, only a small part of intraprediction modes are chosen for RDO calculation. Experimental results show that the fast intraprediction mode decision scheme increases the speed of intracoding significantly with negligible loss of peak signal-to-noise ratio. Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | Fast intermode decision in H.264/AVC video codingabstractThe new video coding standard, H.264/MPEG-4 AVC, uses variable block sizes ranging from 4/spl times/4 to 16/spl times/16 in interframe coding. This new feature has achieved significant coding gain compared to coding a macroblock (MB) using fixed block size. However, this feature results in extremely high computational complexity when brute force rate distortion optimization (RDO) algorithm is used. This paper proposes a fast intermode decision algorithm to decide the best mode in intercoding. It makes use of the spatial homogeneity and the temporal stationarity characteristics of video objects. Specifically, spatial homogeneity of a MB is decided based on the MB's edge intensity, and temporal stationarity is decided by the difference of the current MB and it colocated counterpart in the reference frame. Based on the homogeneity and stationarity of the video objects, only a small number of intermodes are selected in the RDO process. The experimental results show that the fast intermode decision algorithm is able to reduce on the average 30% encoding time, with a negligible peak signal-to-noise ratio loss of 0.03 dB or, equivalently, a bit rate increment of 0.6%. Dajun Wu, Feng Pan 0002, Keng-Pang Lim, Si Wu 0004, Zhengguo Li, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | An unequal packet loss resilience scheme for video over the InternetabstractWe present an unequal packet loss resilience scheme for robust transmission of video over the Internet. By jointly exploiting the unequal importance existing in different levels of syntax hierarchy in video coding schemes, GOP-level and Resynchronization-packet-level Integrated Protection (GRIP) is designed for joint unequal loss protection (ULP) in these two levels using forward error correction (FEC) across packets. Two algorithms are developed to achieve efficient FEC assignment for the proposed GRIP framework: a model-based FEC assignment algorithm and a heuristic FEC assignment algorithm. The model-based FEC assignment algorithm is to achieve optimal allocation of FEC codes based on a simple but effective performance metric, namely distortion-weighted expected length of error propagation, which is adopted to quantify the temporal propagation effect of packet loss on video quality degradation. The heuristic FEC assignment algorithm aims at providing a much simpler yet effective FEC assignment with little computational complexity. The proposed GRIP together with any of the two developed FEC assignment algorithms demonstrates strong robustness against burst packet losses with adaptation to different channel status. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Xiao Lin 0001, Nam Ling |
IEEE Trans. Multim. | 3 |
| 2004 | Geometrically determining leaky bucket parameters for video streaming over constant bit-rate channelsabstractFor streaming of pre-encoded bitstreams over constant bit-rate (CBR) channels, the channel bandwidth, the receiver buffer capacity as well as the latency requirement vary greatly from application to application. We propose an algorithm to determine the minimum buffer size and the minimum start-up delay required for streaming a pre-encoded bitstream over CBR channels at any specific bit rate. The proposed method employs geometric operations to derive the optimal determination for low or high bit rates and sub-optimal determination for medium bit rates. The results have been compared with the H.264/AVC hypothetical reference decoder. The proposed approach provides a theoretical insight and a simple but effective algorithm for determining the leaky bucket parameters for video streaming over CBR channels. Ping Li 0002, Weisi Lin, Susanto Rahardja, Xiao Lin 0001, Xiaokang Yang 0001, Zhengguo Li |
ICASSP (3) | 6 |
| 2004 | Block INTER mode decision for fast encoding of H.264abstractThe paper presents a fast block INTER mode decision algorithm to improve significantly the time efficiency of the encoder in H.264. It makes use of the spatial homogeneity of a video object's textures and the temporal stationarity characteristics inherent in video sequences. Specifically, the homogeneity decision of a block is based on edge information, and MB differencing is used to judge whether the MB is time-stationary. Based on the above analysis, only parts of inter prediction modes are chosen for RDO (rate distortion optimization) calculation. Experimental results show that the new scheme is able to achieve a reduction of 30% encoding time on average, with a negligible average PSNR loss of only 0.03 dB and a mere 0.6% bit rate increase compared with the original H.264 reference software. Dajun Wu, Si Wu 0004, Keng-Pang Lim, Feng Pan 0002, Zhengguo Li, Xiao Lin 0001 |
ICASSP (3) | 5 |
| 2004 | Adaptive rate control for H.264abstractThis paper presents a rate control scheme for H.264 by introducing the concept of basic unit and a linear prediction model. The basic unit can be a macroblock (MB), a slice, or a frame. It can be used to obtain a trade-off between the overall coding efficiency and the bits fluctuation. The linear model is used to solve the chicken and egg dilemma existing in the rate control of H.264. Both constant bit rate (CBR) and variable bit rate (VBR) cases are studied. Our scheme has been adopted by H.264. Zhengguo Li, Feng Pan 0002, Keng-Pang Lim, Xiao Lin 0001, Susanto Rahardja |
ICIP | 1 |
| 2004 | Motion compensated temporal filtering with optimal temporal distance between each motion compensation pairabstractA novel motion compensated temporal filtering (MCTF) scheme is proposed in this paper by properly using forward and backward motion compensations. The mean, the second moment and the maximum value for the temporal distance of all motion compensation pairs (MCPs) in a group of frames (GOF) are minimized such that the number of "unconnected" pixels is minimized. The overall coding efficiency is improved by up to 1.5 dB when compared to the scheme provided in [?] while the total number of motion estimation remains the same. Yang Liu 0028, Zhengguo Li, Yeng Chai Soh |
ICIP | 2 |
| 2004 | Fast intra mode decision algorithm for H.264/AVC video codingabstractThe emerging H.264-AVC video coding standard aims to significantly improve compression performance compared to all existing video coding standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode and reference frame for each macroblock. As a result, the complexity and computation load increase drastically. This paper presents a fast mode decision algorithm for H.264 intra prediction based on local edge information. Prior to intra prediction, an edge map is created and a local edge direction histogram is then established for each sub-block. Based on the distribution of the edge direction histogram, only a small part of intra prediction modes are chosen for RDO calculation. Experimental results show that the last intra mode decision scheme increases the speed of intra coding significantly with negligible loss of PSNR. Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004, C. All, W. Ye, Z. Liang |
ICIP | 5 |
| 2004 | An iterative method for hypothetical reference decoderabstractWe propose a simple iterative method to verify whether a coded bitstream conforms to a hypothetical reference decoder (HRD). A concept of maximum tolerated delay (MTD) is introduced to study the possible low-delay operation, such that we can bound the degree of incorrect motion rendition caused by the variation in end to end delay in the neighborhood of big pictures. Our method can be used in the design of a rate control algorithm to improve the possibility of a coded bitstream that conforms to the HRD. Zhengguo Li, Nam Ling, Susanto Rahardja, Xiao Lin 0001, Ping Li 0002 |
ICME | 1 |
| 2004 | Complexity adaptive quantization for intra-frames in very low bit rate video codingabstractConventional rate control schemes focus on the problem of finding an optimal quantization value for P- and B-frames. No rate control is available for the encoding of I-frames. This could pose big problems due to the large number of bits an I-frame could generate, and due to the fact that the number of bits varies drastically from sequence to sequence, depending on their image complexity. This problem becomes severe especially at very low bit rate. Therefore, a mechanism to allocate data bits to an I-frame according to its complexity is indispensable in order to have constant coding quality. This work presents a mechanism to establish for I-frames a generic relationship between the quantization value, data bits, and their image complexity. Experimental results show that this generic relationship provides a fairly accurate estimation of quantization value for an I-frame at given data bits and image complexity, and is very useful in controlling the data bits generated by an I-frame. Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, Xiao Lin 0001, Susanto Rahardja, Dajun Wu, Si Wu 0004 |
ICME | 2 |
| 2004 | A directional field based fast intra mode decision algorithm for H.264 video codingabstractThe H.264 video coding standard can achieve considerably higher coding efficiency than previous standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode for each macroblock. As a result. the encoder complexity is increased considerably. This paper presents a directional field based fast intra mode decision algorithm to improve the encoder's efficiency. Prior to intra prediction, the directional field is calculated for all the block size to decide the dominant edge direction in the blocks. Based on the edge direction information, only a small number of prediction modes are chosen for RDO calculation. Experimental results show that the fast intra mode decision algorithm increases the speed of intra coding significantly with negligible loss of PSNR Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li |
ICME | 5 |
| 2004 | Proactive frame-skipping decision scheme for variable frame rate video codingabstractMany rate control algorithms focus on the adjustment of quantisation values to retain a certain buffer level, and arbitrary frame-skipping is often needed to keep the buffer from overflow at very low bit-rates. A content adaptive rate control is proposed to optimise the balance between spatial and temporal quality via active frame-skipping. The occurrence of frame-skipping is jointly dependent on the temporal and spatial quality of the video, and on the fullness of the buffer. This helps to achieve a consistent spatial and temporal quality and to enhance the overall perceptual quality. Experimental results show that the new scheme is simple but very effective, with large average PSNR gains and consistently improved visual quality, and the improvement in perceptual quality is much more significant than that in average PSNRs. Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004 |
ICME | 5 |
| 2004 | A new rate-distortion model for video transmission using multiple logarithmic functionsabstractIn this letter, we propose a new logarithmic rate distortion model for DCT-based video encoders. The proposed model can fit into the actual rate curve by combining several logarithmic functions, where the boundary of each function is adaptive to the statistics of video sequence. Experimental results demonstrate that the new rate distortion model provides a more accurate estimation of the bit rate than existing models. Boon-Hee Soong, Zhengguo Li |
IEEE Signal Process. Lett. | 3 |
| 2003 | Adaptive frame layer rate control for H.264abstractThis paper proposes an adaptive frame layer rate control scheme for H.264 by introducing a linear model to predict the mean absolute difference (MAD) of current frame by that of previous one. The target bit rate for each frame is computed by adopting a fluid flow traffic model and linear tracking theory. The corresponding quantization parameter is computed by using a quadratic rate-distortion model. The rate distortion optimization (RDO) is then performed for all macroblocks (MBs) in the current frame by the quantization parameter. Both constant bit rate (CBR) and variable bit rate (VBR) cases are studied. The average PSNR is improved up to 0.75 dB compared to an encoder with fixed quantization parameter. Zhengguo Li, Feng Pan 0002, Keng-Pang Lim, Genan Feng, Xiao Lin 0001, Susanto Rahardja, Dajun Wu |
ICME | 1 |
| 2003 | Video streaming on embedded devices through GPRS networkabstractWe introduce a PDA-based live video streaming system on GPRS network based on MPEG-4 video compression standard. Due to the limited computational resources of PDA, all the key modules of MPEG-4 codec are efficiently implemented and optimized such as multithreading, buffer design, wireless communication, encoder and decoder. Several novel techniques are developed in the coding, streaming as well as the post- processing stages of the system. Keng-Pang Lim, Dajun Wu, Si Wu 0004, Susanto Rahardja, Xiao Lin 0001, Lijun Jiang, Rongshan Yu, Feng Pan 0002, Zhengguo Li, Susu Yao, Genan Feng, Chi Chung Ko |
ICME | 9 |
| 2003 | An adaptive rate control algorithm for video coding over personal digital assistants (PDA)abstractWith the recent development of third-generation communication technologies, encoding live video using a PDA and sharing it among friends has become a reality. However, the embedded processor inside a PDA is still not powerful enough and there are two major hurdles to overcome: (1) video coding needs to meet the rigorous constraint of the available computation capacity of a PDA; (2) In a PDA the computing power allocated to video coding may vary drastically (in bursts). In this paper, a new adaptive rate control algorithm is proposed for video coding over a PDA. This adaptive rate control scheme takes into account the time constraint of a PDA, and its bit allocation depends not only on the available data bits, but more importantly, on the available coding time. Experimental results show that, compared to the existing rate control scheme, the new algorithm can always achieve the maximum frame rate, maximize the utilization of the available bandwidth and computing power, increase the average PSNR, and improve the subjective perceptual quality of the reconstructed video. Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, Dajun Wu, Rongshan Yu, Genan Feng |
ICME | 2 |
| 2003 | Unequal loss protection for robust transmission of motion compensated video over the internet
Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Xiao Lin 0001, Zhengguo Feng, Si Wu 0004, Nam Ling |
Signal Process. Image Commun. | 3 |
| 2003 | A new chaotic secure communication systemabstractThe paper proposes a digital chaotic secure communication by introducing a magnifying glass concept, which is used to enlarge and observe minor parameter mismatch so as to increase the sensitivity of the system. The encryption method is based on a one-time pad encryption scheme, where the random key sequence is replaced by a chaotic sequence generated via a Chua's circuit. We make use of an impulsive control strategy to synchronize two identical chaotic systems embedded in the encryptor and the decryptor, respectively. The lengths of impulsive intervals are piecewise constant and, as a result, the security of the system is further improved. Moreover, with the given parameters of the chaotic system and the impulsive control law, an estimate of the synchronization time is derived. The proposed cryptosystem is shown to be very sensitive to parameter mismatch and hence the security of the chaotic secure communication system is greatly enhanced. Zhengguo Li, Changyun Wen, Yeng Chai Soh |
IEEE Trans. Commun. | 1 |
| 2003 | A unified architecture for real-time video-coding systemsabstractThis paper presents a unified architecture for a live video over the Internet with emphasis on solving some challenging problems such as network bandwidth adaptation for rate and congestion, loss packet recovery, joint source and channel coding, and packetization. In our architecture, a time-varying bit rate for the source coding and time-varying ratios for the channel coding are simultaneously computed by a new congestion-control protocol. An adaptive rate-control scheme is then proposed to calculate quantization parameters and to determine the number of skipping frames corresponding to the bit rate. An adaptive unequal error-control scheme is also provided to protect the bitstream. Furthermore, a simple and MPEG-4 standard compatible algorithm is designed to packetize generated bitstream at the SyncLayer by using the existing resynchronization marker approach. With the proposed architecture, the coding efficiency and the robustness of the whole system are improved greatly. Zhengguo Li, Ce Zhu, Nam Ling, Xiaokang Yang 0001, Genan Feng, Si Wu 0004, Feng Pan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | A study of MPEG-4 rate control scheme and its improvementsabstractThis paper discusses some practical issues in implementing the MPEG-4 Q2 rate-control scheme, and proposes a number of ways to improve it. The improved algorithm has the following main features: (1) the bits allocated to each P-frame or B-frame are in proportion to its distance from the end this GOP. i.e., more bits are allocated to the frames that are nearer to their reference I-frame; (2) the target buffer level is a function of the frame position in the GOP, so that it will be achieved gracefully at the end of a GOP; and (3) the quantization value of an I-frame is decided based on its spatial complexity. Experimental results show that the improved rate-control scheme has significantly reduced the occurrence of frame skipping, increased the average PSNR by up to 0.6 dB, and improved the perceptual quality of the reconstructed video. Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, Genan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | A novel rate control scheme for video over the internetabstractIn this paper, we present a rate control scheme for video over the Internet by adopting a fluid-flow traffic model and a new quadratic rate-distortion (R-D) model. Some simple control theory rather than a heuristic method is used to compute the target rate for each frame. Our scheme is better than some other existing schemes in the sense that our scheme can adapt itselfin time to the variation of channel bandwidth, the number of skipped frames is reduced and the average PSNR is usually improved. Thus, our scheme is very attractive for video over the Internet. Zhengguo Li, Pan Feng |
ICASSP | 1 |
| 2002 | Reducing frame skipping in MPEG-4 rate control schemeabstractFrame skipping in low bit video coding could significantly reduce the visual quality of the reconstructed video. This paper analyzes the' main causes of frame skipping in current MPEG-4 frame rate control scheme, and presents a number of ways to reduce the occurrence of frame skipping. The main features of the new algorithm include an adaptive thresholding technique, an adaptive P-frame target bits allocation, and an adaptive I-frame target bits allocation. Experimental results show that the new rate control scheme has significantly reduced the risks of frame skipping without sacrificing the overall PSNR of the reconstructed video. Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, G. N. Feng |
ICASSP | 2 |
| 2002 | Unequal error protection for motion compensated video streaming over the InternetabstractThis paper presents an unequal error protection scheme for motion compensated video over the Internet. The forward error correction (FEC) codes are optimally assigned to different frames in group of picture (GOP) by exploiting the temporal dependency among frames. To achieve optimal allocation of FEC assigned in GOP, we propose a performance criterion for the measurement of allocation, namely the expected length of error propagation (ELEP), which makes sense intuitively, as fewer frames corrupted implies better quality of reconstruction. Experiment results show the proposed scheme is robust to burst packet loss in the Internet. More importantly, graceful degradation of video quality is achieved by the proposed scheme as the packet loss probability of an Internet connection increases. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling |
ICIP (2) | 3 |
| 2002 | Degressive error protection algorithm for MPEG-4 FGS video streamingabstractThis paper presents a novel degressive error protection (DEP) algorithm adaptive to time varying packet loss rate for robustly transporting MPEG-4 fine granularity scalability (FGS) bit-stream over the packet erasure networks. By exploiting the embedded nature of FGS enhancement-layer bit-stream, a rate-distortion (R-D) optimization framework is developed to optimally partition the FGS enhancement-layer bit-stream into blocks and then to apply DEP to blocks by forward error correction (FEC). Experimental results show that this algorithm achieves graceful degradation of video quality in a wide range of packet loss rates. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling |
ICIP (3) | 3 |
| 2002 | A router based unequal error control scheme for video over the InternetabstractThis paper presents a router based unequal error protection scheme for video over the Internet. The proposed scheme classifies the whole bitstream into two priorities according to their importance and packetizes them into packets with two priorities. To reduce the loss ratio of video packets with higher priority, an active queue management scheme is designed at each router to provide more chance for these packets to transmit over the Internet. Compared to an end-to-end based approach, in which some redundancy packets are generated to protect the packets with higher priority, our proposed scheme utilizes the whole bandwidth for the source coding. Therefore, the final picture quality is improved. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling, Feng Pan 0002 |
ICIP (2) | 3 |
| 2002 | Adaptive unequal error control for video over the InternetabstractWe propose an adaptive unequal error protection scheme to help protect a video stream over the Internet. The data partition approach and the resynchronization marker are applied to generate two types of packets with different importance. Our scheme is very adaptive to network traffic conditions and can be easily implemented via software. With the proposed scheme, more protection is provided for the more important packets. The final quality can therefore be improved at a higher packet loss ratio. Zhengguo Li, Nam Ling, Ce Zhu, Xiaokang Yang 0001, Genan Feng, Si Wu 0004, Feng Pan 0002 |
ICME (1) | 1 |
| 2002 | A congestion control strategy for multipoint videoconferencingabstractWe formulate a congestion control problem for multipoint videoconferencing by introducing the concept of generalized fairness. A numerical method is provided to solve such a congestion control problem. The proposed method is much simpler than an earlier proposed method. This is desirable because congestion control is executed in real time. Zhengguo Li, Xiao Lin 0001, Ce Zhu, Xiaokang Yang 0001, Genan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | A Congestion Control Strategy For Multipoint Videoconferencing
Zhengguo Li, Xiao Lin 0001 |
ICME | 1 |
| 2001 | A Novel Prediction Scheme for Lossless Compression of Audio WaveformabstractFor compression of audio waveform, prediction is one of the important key components. We propose a novel multi-stage adaptive linear predictor (MSALP) for the prediction of high fidelity audio waveform. The MSALP achieves higher prediction gain compared with conventional linear predictor (LP). Besides, the MSALP uses less number of coefficients. The MSALP is embedded into a lossless audio waveform compression system. As a result, the system compression ratio is obviously improved. By selecting known lossless audio waveform compression algorithms, some comparison results are presented. Xiao Lin 0001, Li Gang, Zhengguo Li, Chia Thien King, Yoh Ai Ling |
ICME | 3 |
| 2001 | A switched priority scheduling mechanism for ATM switches with multi-class output buffers
Zhengguo Li, X. J. Yuan, Changyun Wen, Boon-Hee Soong |
Comput. Networks | 1 |
| 2001 | A switching mechanism for ATM ABR traffic control
X. J. Yuan, Zhengguo Li, Boon-Hee Soong, Changyun Wen |
Comput. Commun. | 2 |